Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks.
- Implement ETL/ELT workflows for structured and unstructured data.
- Develop and optimize Spark jobs for performance and cost efficiency.
- Build and maintain data models, schemas, and database structures supporting analytical and operational use cases.
- Integrate Databricks solutions with Azure Data Factory and other Azure cloud services.
- Work with Azure Data Lake Storage and data warehouse solutions.
- Implement data validation and quality checks to ensure data accuracy, consistency, and reliability.
- Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging.
- Implement data security measures, including encryption, access controls, and auditing.
- Support compliance with applicable regulations, security requirements, and industry best practices.
- Automate deployments using CI/CD pipelines, DevOps practices, and version control systems.
- Work with Databricks notebooks, clusters, jobs, and Delta Lake.
- Utilize Unity Catalog and/or Delta Lake to support data quality, governance, and security.
- Troubleshoot and debug data pipelines, Spark applications, and related technical issues.
- Collaborate with data scientists, data analysts, stakeholders, and cross-functional teams.
- Work effectively within Agile and multicultural environments.
Required Qualifications
- 4+ years of experience implementing ETL/ELT workflows for structured and unstructured data.
- 4+ years of experience automating deployments using CI/CD tools.
- 4+ years collaborating with data scientists, analysts, stakeholders, and cross-functional teams.
- 4+ years designing and maintaining data models, schemas, and database structures.
- 4+ years working with data storage solutions, including Azure Data Lake Storage and data warehouses.
- 4+ years implementing data validation and data quality checks.
- 4+ years contributing to data governance, metadata management, data lineage, and data cataloging.
- 4+ years implementing data security measures, including encryption, access controls, and auditing.
- 4+ years of proficiency in Python and R programming languages.
- 4+ years of strong SQL querying and data manipulation experience.
- 4+ years of experience with the Microsoft Azure cloud platform.
- 4+ years of experience with DevOps, CI/CD pipelines, and version control systems.
- 4+ years working in Agile and multicultural environments.
- 4+ years of strong troubleshooting and debugging capabilities.
- 3+ years designing and developing scalable data pipelines using Apache Spark on Databricks.
- 3+ years optimizing Spark jobs for performance and cost efficiency.
- 3+ years integrating Databricks with Azure Data Factory.
- 3+ years ensuring data quality, governance, and security using Unity Catalog or Delta Lake.
- 3+ years of strong understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL.
- 3+ years of hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake.
Preferred Qualifications
- Knowledge of machine learning libraries such as:
- MLflow
- Scikit-learn
- TensorFlow
- Databricks Certified Associate Developer for Apache Spark certification.
- Microsoft Certified: Azure Data Engineer Associate certification.
Role Overview:
The position involves designing, developing, and optimizing scalable data pipelines and big data solutions using Azure, Databricks, and Apache Spark.
The role also involves data quality, governance, security, CI/CD, and collaboration with data engineering, analytics, and business teams.