Responsibilities
- Build and operate data pipelines (batch and streaming) from various sources including APIs, relational databases, file drops, event streams, and external partners.
- Design, implement, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning.
- Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
- Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
- Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
- Implement SQL-like table reliability features including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities.
- Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
- Optimize performance and cost efficiency through partitioning, compaction, file sizing, caching, lifecycle policies, and efficient compute/storage separation.
- Establish standardized development, testing, and production environments with consistent configuration and controlled promotion across stages.
- Implement data governance and fine-grained access control utilizing AWS-native services like AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS, and related security tools.
- Create a managed metadata repository for dataset cataloging, ownership, tagging, classification, and discoverability.
- Support end-to-end data lineage for source, transformation, and consumption to facilitate auditability and impact analysis.
- Apply security policies such as least privilege access, data classification, encryption, retention, and secure data handling.
- Build operational data quality checks for metrics such as freshness, completeness, validity, and anomaly detection, along with publishing SLAs/SLOs.
- Implement automated AWS provisioning through Infrastructure as Code (IaC) to ensure consistent, secure environments.
- Enhance CI/CD pipelines for data workflows and lakehouse components, including automated testing, security validation, packaging, deployment, promotion, and rollback.
- Maintain observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident response procedures.
- Continually evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
- Collaborate closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission-critical requirements.
- Maintain high-quality documentation including architecture diagrams, SOPs, data models, interface specs, and operational runbooks.
- Present technical findings, trade-offs, risks, and recommendations clearly to stakeholders.
Requirements
- Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or four (4) years of equivalent practical experience.
- Six (6) years of relevant hands-on experience in data engineering.
- Extensive experience designing, implementing, and operating AWS-native data lake or lakehouse architectures with Amazon S3, AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
- Proven ability developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, tuning, and error handling.
- Hands-on experience with Apache Iceberg covering ACID transactions, snapshots, schema evolution, and query optimization.
- Advanced SQL skills supporting analytical queries, reporting, and data visualization workloads.
- Demonstrated experience with data governance, cataloging, lineage, ownership, classification, and access controls.
- Knowledge of AWS security fundamentals including IAM, encryption (KMS), secrets management, and network security.
- Experience provisioning resources via Infrastructure as Code (IaC) and managing multi-environment platforms.
- Skilled in building and maintaining CI/CD pipelines for data workflows with automation, testing, and rollback strategies.
- Strong troubleshooting skills for distributed data workloads with focus on performance, reliability, and cost management.
Industry:
Data & Analytics