About the Role
We are seeking an experienced AWS Lakehouse Data Engineer to design, build, and operate a modern AWS-native data platform supporting analytics, reporting, visualization, and AI/ML workloads.
The ideal candidate will have strong hands-on experience with AWS data services, Amazon S3, Apache Iceberg, Python, PySpark, SQL, ETL/ELT pipelines, and data governance.
Responsibilities
- Design and build scalable batch and streaming data pipelines using Python and PySpark.
- Build and maintain an AWS-native lakehouse on Amazon S3 using Apache Iceberg and Parquet.
- Develop ETL/ELT workflows, incremental processing, CDC, schema validation, and reusable transformations.
- Use services such as AWS Glue, Athena, EMR, Lake Formation, and Redshift for data processing, querying, and governance.
- Implement metadata management, lineage, data quality, encryption, and fine-grained access controls.
- Automate AWS infrastructure using Infrastructure as Code (IaC) and support CI/CD pipelines across development, test, and production environments.
- Monitor and optimize platform performance, reliability, security, and cost.
- Collaborate with data, analytics, AI/ML, security, and cloud teams and maintain technical documentation.
Required Qualifications
- Hands-on experience building AWS data lake/lakehouse architectures.
- Strong Python, PySpark, and SQL skills with production ETL/ELT experience.
- Experience with Apache Iceberg, including ACID transactions, schema evolution, snapshots, partitioning, and optimization.
- Experience with AWS services including S3, Glue, Athena, EMR, Lake Formation, IAM, KMS, and Redshift.
- Experience with data governance, metadata, lineage, data quality, and fine-grained access controls.
- Experience with IaC, CI/CD, AWS security, and distributed data-processing environments.
Nice to Have
- Databricks or Delta Lake experience, including migration to AWS-native platforms.
- Experience with Terraform, CloudFormation/CDK, Docker, GitHub Actions, Jenkins, Step Functions, MWAA, Kinesis, Lambda, DMS, or MSK.
- Experience with QuickSight, Tableau, or Power BI.
- Knowledge of AI-assisted development, Knowledge Graphs, and Graph RAG, including ontology modeling, entity resolution, and hybrid graph/vector retrieval.