Job Description
We are looking for an experienced Lead Site Reliability Engineer (SRE) – Observability to join our team and drive the design, implementation, and support of enterprise-scale observability platforms. The ideal candidate will have strong expertise in Splunk, Elasticsearch (ELK), Grafana, Prometheus, OpenTelemetry, Kafka, Terraform, and Kubernetes, with a solid background in Site Reliability Engineering, DevOps, and cloud technologies.
Required Skills
- 7+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, or DevOps.
- Hands-on experience with Splunk Enterprise and/or Splunk Cloud administration.
- Strong proficiency in Splunk SPL.
- Experience with Elasticsearch (ELK Stack), Kibana, Prometheus, Grafana, Grafana Tempo, and OpenTelemetry.
- Experience implementing distributed tracing, monitoring, logging, and alerting solutions.
- Strong knowledge of Kafka and observability pipelines.
- Hands-on experience with Terraform and Infrastructure as Code (IaC).
- Experience with Kubernetes, Docker, and Linux environments.
- Strong scripting skills using Python, Go, Ruby, or Bash.
- Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
Responsibilities
- Design, deploy, and maintain enterprise observability platforms.
- Administer Splunk infrastructure, including Search Head Clusters, Indexers, Heavy Forwarders, and Deployment Servers.
- Build and manage Elasticsearch clusters for large-scale log analytics.
- Develop dashboards, alerts, and monitoring solutions using Splunk, Grafana, and Kibana.
- Implement distributed tracing using OpenTelemetry and Grafana Tempo.
- Automate infrastructure deployments using Terraform.
- Troubleshoot production issues and improve platform reliability, scalability, and performance.
- Collaborate with development and infrastructure teams to enhance monitoring and operational excellence.
Preferred Qualifications
- Splunk Certification.
- Experience with Ansible, Consul, CI/CD pipelines, and Service Mesh technologies.
- Experience working in FedRAMP or other regulated environments.