Role Summary
As a Lead AI/ML Engineer you are an experienced individual contributor who independently drives one or more AI/ML workstreams end-to-end.
You own delivery outcomes for your assigned workstream, make key technical decisions, and interface directly with customer technical leads and stakeholders.
You bring deep expertise in ML Ops and time-series forecasting to architect and implement scalable solutions on AWS.
Key Responsibilities
- Design, develop, and deploy machine learning models for latent capacity prediction across pipeline systems (weather, gas turbine HP, compressor flow, line pack, equipment performance)
- Build and maintain ML Ops infrastructure on AWS SageMaker including model registry, versioning, CI/CD pipelines, and multi-environment endpoints (Dev, Pre-Prod, Prod)
- Conduct exploratory data analysis and feature engineering for time-series forecasting of pipeline operational data (SCADA, performance curves, hydraulic models)
- Develop ensemble model strategies combining LSTM, Prophet, and XGBoost for improved prediction accuracy of latent capacity
- Build automated model deployment workflows, training/evaluation pipelines, and retraining frameworks
- Design real-time and batch inference architectures with API Gateway integration and model monitoring
- Collaborate with data engineering team on feature stores and data pipeline integration from Bronze/Silver/Gold data lakehouse layers
- Support hydraulic model integration with Gregg Engineering NextGen software for automated scenario generation
- Independently own delivery of assigned ML workstream, driving technical decisions and ensuring quality
- Mentor L4/L5 team members on ML best practices, code reviews, and architectural patterns
- Interface directly with customer technical leads to align on requirements, review progress, and resolve technical blockers
Required Skills & Qualifications
- Strong experience with AWS SageMaker, including SageMaker Pipelines, Model Registry, and Feature Store
- Proficiency in time-series forecasting (LSTM, Prophet, XGBoost, ensemble methods)
- Experience with ML Ops practices: CI/CD for ML, model monitoring, automated retraining
- Python (NumPy, Pandas, scikit-learn, TensorFlow/PyTorch)
- Experience with real-time and batch inference architectures
- Knowledge of data lakehouse architectures (S3, Glue, Redshift)
- Understanding of industrial/operational data (SCADA, IoT sensors) is a plus
- Experience in Energy & Utilities domain preferred
- Demonstrated ability to independently lead technical workstreams and make architectural decisions
- Experience mentoring junior engineers or consultants
- Strong communication skills for customer-facing interactions