- 5+ years of relevant experience
- Experience in planning of systems and development deployment and meeting software compliance standards
- Experience evaluating hardware/software interfaces, operational requirements, and overall system characteristics
Technical Skills:
AWS: EC2, S3, IAM, KMS, Lambda, ECS/EKS, Step Functions, CloudWatch, EventBridge, VPC networking
Azure: Commercial Azure compute, storage, and networking
Databricks: cluster/workspace administration, Workflows, GPU clusters
Containers & orchestration: Docker, Kubernetes (EKS)
Infrastructure-as-code & config management: Terraform, CloudFormation, Ansible, Git
Observability: CloudWatch, centralized logging, alerting/incident management tools
Languages: Python, Bash/PowerShell; integration via REST APIs
- BS in Computer Science or related field
- Education substitution: HS/GED + 6 yrs, HS/GED + cert + 4 yrs, AA + 4 yrs, or AA + cert + 2 yrs additional experience
- Able to read, write, speak, and understand English
- Building and scaling AI infrastructure: GPU optimization, containerized model serving (ECS/EKS), and elastic compute for training and inference
- Designing architectures for agentic AI applications (agent orchestration, tool integration, secure API access)
- Using AIOps for predictive monitoring, anomaly detection, and automated incident triage to meet uptime SLOs
- Using AI coding assistants (e.g., AWS Kiro, Copilot, Codex) for infrastructure-as-code and system automation