Data Engineer - Azure Databricks & Application Development
Your role as a Data Engineering & Application Development
- Design, develop, test, and maintain scalable data engineering solutions on Microsoft Azure and Databricks.
- Build robust and reusable ETL/ELT pipelines using Python, PySpark, Scala, and SQL.
- Develop data processing applications for ingestion, transformation, enrichment, validation, and serving of analytical datasets.
- Implement reliable batch and streaming data workflows using Azure Databricks, Delta Lake, and modern data lakehouse patterns.
- Contribute to application design, code quality, performance optimization, and maintainability of data engineering solutions.
- Develop reusable libraries, utilities, frameworks, and technical components to accelerate data product delivery.
Azure / Databricks Data Platform
- Develop solutions leveraging Azure andDatabricks data services, including:
- Azure Databricks
- Delta Lake / Lakehouse architecture
- Azure Data Lake Storage
- Azure Data Factory / Synapse pipelines
- Azure Event Hubs / streaming ingestion patterns
- Azure Key Vault
- Azure DevOps
- Implement data models, transformation layers, and curated datasets supporting analytics, reporting, AI, and application use cases.
- Apply best practices for data partitioning, performance tuning, schema evolution, data quality, and operational monitoring.
- Collaborate with architecture and platform teams to ensure solutions are secure, scalable, cost-efficient, and aligned with enterprise standards.
ETL, Data Pipelines & Software Engineering
- Develop and maintain production-grade data pipelines using Python, PySpark, Scala, and SQL.
- Implement automated testing, validation, logging, error handling, and monitoring for data processing applications.
- Optimize Spark jobs for performance, scalability, memory usage, and cost efficiency.
- Apply software engineering practices such as modular design, clean code, version control, code reviews, and documentation.
- Build CI/CD pipelines for data applications using Azure DevOps YAML pipelines.
- Package, deploy, and maintain data engineering code across multiple environments.
Infrastructure & DevOps Collaboration
- Collaborate with Cloud, DevOps, and Platform teams on deployment, environment configuration, and operational readiness.
- Contribute to infrastructure automation where needed, with a working understanding of Terraform and cloud deployment principles.
- Support the integration of application code with cloud resources, security configurations, and CI/CD deployment processes.
- Follow DevOps practices for release management, environment promotion, and production support.
Collaboration & Technical Leadership
- Work cross-functionally with Architecture, Data, AI, Cloud, CI/CD, and application teams.
- Translate business and analytical requirements into scalable data engineering solutions.
- Provide technical guidance on Databricks, Spark, Python/PySpark development, and ETL implementation patterns.
- Support less experienced engineers through code reviews, technical coaching, and knowledge sharing.
- Contribute to design documentation, development standards, and reusable engineering practices.
Required Qualifications
- 5+ years of hands-on experience in data engineering, application development, or cloud-based data platform development.
- Strong hands-on development experience with:
- Python
- PySpark
- Scala
- SQL
- ETL/ELT development
- Solid experience with the Azure Databricks ecosystem, including:
- Databricks notebooks and jobs
- Spark clusters
- Delta Lake
- Lakehouse architecture
- Performance tuning and optimization
- Experience designing, developing, and operating production data pipelines.
- Good knowledge of Azure data services such as:
- Azure Data Lake Storage
- Azure Data Factory
- Azure Key Vault
- Azure DevOps
- Azure monitoring/logging capabilities
- Experience with CI/CD pipelines, preferably using Azure DevOps YAML pipelines.
- Strong understanding of software engineering practices, including version control, automated testing, modular development, and code reviews.
- Understanding of cloud security, access management, observability, and cost-aware development practices.
- Working knowledge of Terraform or infrastructure-as-code concepts is a plus, but not the primary focus of the role.
Preferred Qualifications
- Experience with large-scale distributed data processing using Spark.
- Experience with streaming data pipelines and real-time ingestion patterns.
- Experience building reusable data engineering frameworks, libraries, or shared components.
- Knowledge of data quality frameworks, metadata management, lineage, and governance practices.
- Experience integrating data pipelines with AI, machine learning, or analytics use cases.
- Familiarity with MLOps or AI engineering workflows is an advantage.
- Experience working in enterprise environments with multiple teams, environments, and delivery governance.
Preferred Certifications
Azure Certifications
- Microsoft Certified: Azure Data Engineer Associate
- Microsoft Certified: Azure Developer Associate
- Microsoft Certified: Azure Solutions Architect Expert
- Microsoft Certified: Azure DevOps Engineer Expert
Databricks Certifications
- Databricks Certified Data Engineer Associate
- Databricks Certified Data Engineer Professional
- Databricks Certified Developer for Apache Spark
Soft Skills
- Strong communication and documentation skills.
- Ability to translate business and analytical needs into practical data engineering solutions.
- Strong problem-solving mindset with a focus on automation, reliability, and maintainability.
- Team-oriented, proactive, and comfortable working in a cross-functional environment.
- Ability to explain technical data engineering concepts to diverse technical audiences.
- Continuous improvement mindset and willingness to adopt evolving data platform technologies.