Job Summary
We are seeking an experienced Senior AI/ML Data Engineer with strong expertise in Azure Databricks, Python, PySpark, Apache Spark, Azure Data Factory, and AI/ML technologies. The ideal candidate will have hands-on experience building enterprise-scale data platforms, developing AI/ML data pipelines, implementing MLOps practices, and supporting Generative AI and Large Language Model (LLM) solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Azure Databricks, PySpark, Apache Spark, and Azure Data Factory.
- Build and optimize enterprise data platforms using Delta Lake, Lakehouse Architecture, and ADLS Gen2.
- Develop data pipelines supporting Machine Learning, predictive analytics, and AI-driven applications.
- Implement MLflow for experiment tracking, model lifecycle management, and MLOps best practices.
- Build scalable data pipelines for Generative AI, Azure OpenAI, LLM, and RAG-based applications.
- Process structured and unstructured data from APIs, relational databases, cloud storage, JSON, XML, CSV, and Parquet files.
- Optimize Spark workloads, SQL queries, and distributed data processing for high performance.
- Develop feature engineering pipelines for AI/ML models.
- Implement CI/CD pipelines using GitHub Actions and Azure DevOps.
- Collaborate with Data Scientists, ML Engineers, Product Owners, and business stakeholders to deliver enterprise AI solutions.
- Mentor junior engineers and establish best practices for cloud-native data engineering.
Required Qualifications
- 10+ years of experience in Data Engineering, ETL/ELT, and Cloud Data Platforms.
- Strong hands-on experience with Azure Databricks, Python, PySpark, Apache Spark, and Spark SQL.
- Experience with Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Gen2), and Delta Lake.
- Strong knowledge of Lakehouse Architecture and distributed data processing.
- Experience implementing Machine Learning pipelines, MLflow, and MLOps practices.
- Experience working with Azure OpenAI, Generative AI, LLMs, and RAG frameworks.
- Strong SQL development, performance tuning, and data modeling experience.
- Experience with GitHub Actions, Azure DevOps, Docker, Kubernetes, and CI/CD pipelines.
- Excellent analytical, troubleshooting, and communication skills.
Preferred Qualifications
- Experience with Terraform and Infrastructure as Code (IaC).
- Knowledge of Azure Functions and Azure Service Bus.
- Experience with Snowflake and enterprise cloud data platforms.
- Experience working in Agile/Scrum environments.
- Strong leadership and mentoring experience.
Required Skills
- Azure Databricks
- Python
- PySpark
- Apache Spark
- Spark SQL
- Azure Data Factory (ADF)
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Delta Lake
- Lakehouse Architecture
- ETL/ELT
- SQL
- Machine Learning
- MLflow
- MLOps
- Azure OpenAI
- Generative AI
- Large Language Models (LLMs)
- Retrieval-Augmented Generation (RAG)
- GitHub Actions
- Azure DevOps
- Docker
- Kubernetes
- CI/CD
- Data Engineering
- Data Modeling
- Data Architecture