Key Responsibilities
- Design and develop scalable Databricks data pipelines based on the Medallion Architecture (Bronze, Silver, and Gold layers) using Delta Lake.
- Lead the adoption of Spec-Driven Development (SDD) by creating machine-readable specifications and data contracts before implementation.
- Build, monitor, and optimize workflow orchestration pipelines using Apache Airflow.
- Develop, test, and maintain modular data transformation models with dbt, integrated with the Databricks SQL and Spark ecosystem.
- Optimize PySpark workloads, manage cluster configurations, and control compute and storage costs.
- Implement data governance, granular access controls, and data lineage tracking through Unity Catalog.
Qualifications
- 5+ years of hands-on data engineering experience, including extensive production experience with Databricks.
- Strong expertise in Apache Airflow for workflow scheduling, dependency management, and error handling.
- Advanced knowledge of dbt modelling, testing practices, and macro development.
- Practical experience with Spec-Driven Development (SDD) and metadata-driven data platform design.
- Expert-level proficiency in Python (PySpark) and SQL.
- Familiarity with Git, CI/CD processes, and infrastructure-as-code practices.
Multi-Cloud Experience (AWS, Azure, GCP)
- AWS: Configure and manage integrations with Amazon S3, IAM roles, and VPCs to support secure and scalable data platforms.
- Azure: Implement secure data storage using Azure Data Lake Storage Gen2 (ADLS Gen2), integrate with Microsoft Entra ID, and deliver data to analytics platforms such as Power BI.
- GCP: Build and operate data workflows using Google Cloud Storage (GCS) while enabling integration with BigQuery and Vertex AI for downstream analytics and AI use cases.