Responsibilities
- Implement ETL/ELT workflows for both structured and unstructured data
- Collaborate with cross-functional teams including data scientists, analysts, and stakeholders
- Design and maintain data models, schemas, and database structures to support analytical and operational use cases
- Implement data validation and quality checks to ensure accuracy and consistency
- Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging
- Implement data security measures, including encryption, access controls, and auditing; ensure compliance with regulations and best practices
- Working in agile, multicultural environments
- Automate deployments using CI/CD tools
- Evaluate and implement appropriate data storage solutions, including data lakes (Azure Data Lake Storage) and data warehouses
- Proficiency in Python and R programming languages
- Strong SQL querying and data manipulation skills
- Experience with Azure cloud platform
- Experience with DevOps, CI/CD pipelines, and version control systems
- Strong troubleshooting and debugging capabilities
- Design and develop scalable data pipelines using Apache Spark on Databricks
- Optimize Spark jobs for performance and cost efficiency
- Integrate Databricks solutions with cloud services (Azure Data Factory)
- Ensure data quality, governance, and security using Unity Catalog or Delta Lake
- Deep understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL
- Hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake
- Build AI accelerators and reusable components to boost development teams productivity.
- Implement Gen AI / LLM application and utilities (Agentic AI, Harness / Prompt engineering, RAG)
- Implement AI application governance, observability and evaluation framework
- Design & Implement vector stores for knowledge bases and agent memory
- Integrate AI applications / utils / tools with enterprise traceability and SIEM tools
- Knowledge of ML libraries (MLflow, Scikit-learn, TensorFlow)
- Databricks Certified Associate Developer for Apache Spark
- Azure Data Engineer Associate
Industry:
Data & Analytics
Seniority Level:
Not Specified