Job Description
We are looking for an experienced AWS Data Engineer with strong hands-on expertise in AWS Glue, Amazon Redshift, S3, Python, PySpark, and SQL to design, develop, and maintain scalable cloud data pipelines and enterprise data warehouse solutions.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using AWS Glue, Python, PySpark, and SQL.
- Develop AWS Glue Jobs, Crawlers, Workflows, and Data Catalog solutions for automated data ingestion, transformation, and processing.
- Build and optimize data pipelines to ingest structured and semi-structured data from multiple sources into Amazon S3 and Amazon Redshift.
- Design and develop Amazon Redshift data warehouse solutions, including tables, views, stored procedures, and analytical data models.
- Develop complex SQL queries for data transformation, validation, reconciliation, and analytical reporting.
- Optimize Redshift performance using appropriate distribution styles, sort keys, table design, workload optimization, and query tuning techniques.
- Implement incremental and full-load ETL strategies, including data validation, error handling, reconciliation, and restartable processing.
- Integrate AWS Glue, S3, Redshift, Athena, and other AWS services to build scalable cloud data platforms.
- Develop PySpark-based transformations for large-volume datasets and optimize Spark jobs for performance and reliability.
- Implement data quality checks to ensure data completeness, accuracy, consistency, and integrity across source and target systems.
- Work with AWS Glue Data Catalog and crawlers to manage metadata and support schema discovery for enterprise datasets.
- Monitor and troubleshoot production data pipelines, Glue jobs, and Redshift workloads to ensure high availability and reliability.
- Support migration of legacy ETL and data warehouse workloads to AWS cloud-native platforms.
- Implement security and access controls using AWS IAM roles, policies, and encryption mechanisms.
- Collaborate with Data Architects, BI Developers, Business Analysts, and other engineering teams to translate business requirements into scalable data solutions.
- Participate in CI/CD and infrastructure automation using tools such as Git, GitHub Actions, Jenkins, Terraform, or CloudFormation.
Required Skills
- 5+ years of experience in Data Engineering
- Strong hands-on experience with AWS Glue
- Strong experience with Amazon Redshift
- Proficiency in Python and PySpark
- Advanced SQL skills
- Experience with Amazon S3
- Experience with AWS Glue Data Catalog and Crawlers
- Experience with AWS Athena
- Strong understanding of ETL/ELT concepts and data warehousing
- Experience with dimensional modeling, fact/dimension tables, and data marts
- Experience with performance tuning and troubleshooting of large-scale data pipelines
Preferred Skills
- Experience with Databricks / Apache Spark
- Experience with AWS EMR
- Experience with Apache Airflow
- Experience with Snowflake
- Experience with Terraform
- Experience with Power BI or Tableau
- Experience with cloud migration and legacy ETL modernization
- Experience working with healthcare, finance, telecom, or other enterprise data environments