Job Description:
Client is seeking a Senior ML Deployment Engineer to own the end-to-end machine learning model lifecycle from post-training through production. This is a production engineering role, not an ML research position. You will work closely with ML researchers, platform teams, and software engineers to deploy, benchmark, monitor, and optimize ML models in a hybrid cloud environment.
Required Skills:
- 10+ years of IT experience with strong expertise in ML deployment and inference.
- Hands-on experience with Google Cloud Platform (Google Cloud Platform) for model deployment and serving.
- Strong knowledge of TensorFlow, PyTorch, JAX, or similar ML frameworks.
- Experience with ML model benchmarking, performance tuning, and quality testing.
- Ability to quickly learn and adapt to new technologies and evolving tech stacks.
- Strong problem-solving skills with an end-to-end ownership mindset.
- Understanding of distributed systems and debugging production environments.
Nice to Have:
- Exposure to Java/JVM-based applications.
- Experience with streaming data pipelines and hybrid cloud environments.
- Familiarity with deployment automation and production monitoring.
Responsibilities:
- Deploy, benchmark, and optimize ML models for production on Google Cloud Platform.
- Integrate ML models into Java-based streaming applications.
- Design and execute performance, inference, and quality tests.
- Monitor production model performance and troubleshoot issues.
- Collaborate with ML researchers to evaluate models and improve deployment strategies.
- Drive deployment automation and ensure reliable production serving.
Education:
Bachelor's or Master's degree in Computer Science, Engineering, Mathematics, or a related field.