The Role
You will help design and implement the standards, automation and infrastructure required to provide consistent observability across a large, multi-account AWS environment.
Key areas will include:
- Designing scalable telemetry collection and distribution across AWS and EKS.
- Developing standards for metrics, logs, alerting and telemetry.
- Building and integrating observability platforms, including storage, routing and retention.
- Working with OpenTelemetry, Prometheus-compatible tooling and cloud-native telemetry.
- Integrating observability with service ownership, inventory and PagerDuty.
- Creating automated onboarding and self-service capabilities for engineering teams.
- Applying observability standards to both new and existing cloud workloads.
- Automating repetitive operational processes and reducing engineering toil.
- Designing resilient architectures that avoid circular dependencies and remain diagnosable when the observability platform itself has issues.
- Working closely with Cloud, Platform, SRE and application engineering teams.
Technology
The environment includes: AWS, EKS/Kubernetes, OpenTelemetry, CloudWatch, Prometheus, VictoriaMetrics, VictoriaLogs, Alertmanager, Grafana, PagerDuty, Terraform, Python and Go.
You don't need experience with every technology listed. My client is more interested in strong production experience and an understanding of observability fundamentals than expertise with a particular vendor.
What I'm Looking For
You should have experience building or significantly improving production observability or SRE capabilities at scale.
You'll ideally bring: Strong AWS and Kubernetes experience.
Solid observability and distributed systems knowledge.
Infrastructure-as-Code and automation experience.
Experience with telemetry pipelines, metrics, logging and alerting.
Strong Python and/or Go skills.
A practical, problem-solving mindset.
The ability to investigate unfamiliar systems and work through ambiguity.
Experience building platforms that engineering teams actually adopt.
The client is looking for someone who owns outcomes rather than tickets — someone who can understand the underlying problem, challenge assumptions and build reliable solutions.
The Opportunity
This is an opportunity to join while the client's cloud observability capability is still being shaped.
You'll have genuine influence over architecture, standards, automation and platform adoption, working alongside experienced Cloud, Platform and Engineering teams.
Role: Senior Site Reliability - Observability Specialist
Engagement: Contract
Environment: AWS / EKS / Kubernetes / Observability