Responsibilities
- Develop scalable data pipelines using Databricks, PySpark and SQL.
- Build and maintain Delta Lake tables using the Medallion Architecture.
- Develop batch and real-time data processing pipelines.
- Implement data ingestion using Auto Loader, APIs, Kafka or cloud storage.
- Optimize Spark jobs, SQL queries and Delta tables.
- Implement data governance and security using Unity Catalog.
- Develop Databricks Workflows/Jobs and monitor production pipelines.
- Integrate Databricks with AWS/Azure/Google Cloud Platform data services.
- Implement CI/CD and source-code management using Git and DevOps tools.
- Collaborate with data architects, analysts, ML engineers and business teams.
IR35 Status:
Not specified
Industry:
Data & Analytics