Role Overview
We are seeking an experienced Senior Databricks Engineer with deep, hands-on data engineering expertise to independently design, build and own production-grade pipelines on our Databricks Lakehouse. In this role you will develop complex transformations at scale in PySpark and Spark SQL, and deliver curated Bronze/Silver/Gold data products on the medallion architecture that power analytics, reporting and machine learning. You will own your pipelines end to end - design, build, test, deploy, optimize and support - working closely with data architecture, analytics, data science and platform engineering teams to deliver reliable, performant and cost-efficient data products across development, test and production.
Key Responsibilities & Skillsets
- Independently design, build and own production-grade data pipelines on Databricks, from source ingestion through curated delivery, using PySpark, Spark SQL, Delta Live Tables and Databricks Workflows.
- Build and maintain Bronze/Silver/Gold data products on the medallion architecture - raw ingestion, cleansing and conformance, and business-ready aggregates modelled for analytics and ML consumption.
- Implement complex transformations at scale - multi-source joins, SCD Type 1/2 history, deduplication, late-arriving and out-of-order data, CDC merges, windowing and business-rule logic.
- Engineer batch and streaming ingestion from files, databases, APIs and event streams using Auto Loader, Structured Streaming, Delta Lake MERGE and change data feed.
- Tune Spark and Delta performance and cost - partitioning, liquid clustering, Z-ordering, OPTIMIZE and VACUUM, caching, broadcast strategy, skew and spill remediation, Photon and right-sized compute.
- Model curated data products with consuming teams - dimensional and star schemas, semantic layers, and serving through Databricks SQL warehouses and Delta Sharing.
- Engineer data quality and observability into every pipeline - expectations, schema enforcement and evolution, reconciliation and threshold checks, freshness and volume SLAs, and actionable alerting.
- Develop pipelines as software - modular Python packages, unit and integration tests, code review, Git branching, CI/CD with Azure DevOps or GitHub Actions, and deployment via Databricks Asset Bundles.
- Work within Unity Catalog governance - catalogs, schemas, volumes, lineage, tagging and fine-grained access - applying the standards set by the platform administration team.
- Own production support for your pipelines - triage, root cause analysis, backfills and reprocessing, and continuous hardening; mentor junior engineers and set code, design and documentation standards.
Candidate Profile
- A Bachelor's or Master's degree in Computer Science, Information Systems, Engineering or a similar discipline.
- Strong hands-on experience building and running production data pipelines on Databricks on Microsoft Azure (AWS or Google Cloud Platform exposure a plus), with a proven ability to deliver independently end to end.
- Expert-level PySpark and Spark SQL, plus advanced SQL - window functions, complex joins, CTEs and performance-oriented query design.
- Deep Delta Lake expertise - ACID transactions, MERGE and upsert patterns, time travel, schema evolution, change data feed, OPTIMIZE, VACUUM and liquid clustering.
- Proven track record delivering Bronze/Silver/Gold (medallion) data products, including data modelling for analytics, BI and ML consumption.
- Hands-on experience with Databricks Workflows, Delta Live Tables, Auto Loader and Structured Streaming for batch and near-real-time pipelines.
- Strong Spark performance tuning and cost optimization skills - reading query plans and the Spark UI, and diagnosing skew, spill, shuffle and small-file problems.
- Production-grade software engineering practice - modular Python, testing with pytest, Git and code review, CI/CD (Azure DevOps or GitHub Actions), and Databricks Asset Bundles or Terraform.
- Working knowledge of Unity Catalog and the surrounding Azure data ecosystem - ADLS Gen2, Azure Data Factory, Key Vault, Event Hubs or Kafka, Synapse or Microsoft Fabric, and Microsoft Purview.
- Excellent problem-solving, documentation and stakeholder communication skills, with experience mentoring junior engineers; Databricks certifications (Data Engineer Associate/Professional) and Azure certifications (DP-203) are a plus.