All Jobs Vacancy

Site Reliability Engineer - Snowflake Cortex + AWS Bedrock @ Remote

Posted 1 week ago by Momento USA LLC

Role: Site Reliability Engineer - Snowflake Cortex + AWS Bedrock

Location: Remote

Job Description

We are looking for an experienced Site Reliability Engineer (SRE) to support and operate highly available Snowflake and AWS-based data and AI platforms, with hands-on exposure to Snowflake Cortex and Amazon Bedrock.

The engineer will be responsible for platform reliability, monitoring, incident management, automation, performance optimization, security, and operational support for data and Generative AI workloads.

Key Responsibilities

  • Design, deploy, monitor, and maintain highly available AWS, Snowflake, and GenAI platforms.
  • Support Snowflake Cortex capabilities for AI/ML and Generative AI workloads.
  • Work with Amazon Bedrock to integrate and operate foundation models and GenAI applications.
  • Monitor infrastructure, applications, data pipelines, APIs, and AI workloads.
  • Define and monitor SLIs, SLOs, and SLAs for critical services.
  • Implement observability using CloudWatch, logs, metrics, traces, and alerting.
  • Troubleshoot production incidents and perform Root Cause Analysis (RCA).
  • Participate in on call / production support and incident management.
  • Automate repetitive operational activities using Python, Shell scripting, Terraform, or AWS services.
  • Build and maintain CI/CD pipelines for application, data, and AI workloads.
  • Optimize Snowflake compute usage, query performance, warehouse utilization, and cost.
  • Monitor and troubleshoot Snowflake data pipelines and workloads.
  • Support AWS services such as S3, Lambda, IAM, CloudWatch, ECS/EKS, API Gateway, and Secrets Manager.
  • Implement security, access controls, encryption, secrets management, and least-privilege IAM.
  • Monitor the reliability and performance of LLM/GenAI applications using Amazon Bedrock.
  • Track AI application metrics such as latency, throughput, errors, token usage, and cost.
  • Establish automated health checks, alerts, recovery mechanisms, and disaster-recovery procedures.
  • Collaborate with Data Engineers, ML Engineers, Developers, Architects, Product Managers, and business stakeholders.
  • Continuously improve platform reliability through automation, observability, capacity planning, and performance engineering.
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
IT
Seniority Level:
Not Specified

Take-Home Pay

Not Available

Visit calculators for additional details

Share job