Key Responsibilities
- Own and evolve the Datadog observability platform.
- Design and maintain synthetic monitoring for critical API and UI workflows.
- Build continuous production validation covering business-critical customer journeys.
- Integrate monitoring, testing, dashboards, and alerting into Azure DevOps and GitHub pipelines.
- Develop monitoring-as-code and testing-as-code practices using Terraform.
- Create actionable dashboards, SLOs, SLIs, alerts, and anomaly detection.
- Integrate Datadog with Azure, Cloudflare, and modern SaaS architectures.
- Drive reliability, performance, and root-cause analysis across production systems.
Required Experience
- Strong hands-on Datadog expertise, including: Synthetic Monitoring, APM, RUM, Log Management, SLOs and Alerting
- Experience operating large-scale global SaaS platforms.
- Deep Azure experience.
- Experience integrating Cloudflare services.
- Strong CI/CD experience with Azure DevOps and GitHub.
- Expertise in API, integration, and browser-based testing.
- Infrastructure as Code experience using Terraform.
- Experience with distributed systems, microservices, and cloud-native architectures.
Desirable
- Datadog certifications.
- Azure certifications.
- Cloudflare administration experience.
- Background in Site Reliability Engineering (SRE) or Platform Engineering leadership roles.