Key Responsibilities
- Perform GitLab administration, including installation, configuration, upgrades, patching, and ongoing maintenance.
- Manage GitLab Geo replication, backup, restore, and disaster recovery processes.
- Design and maintain infrastructure using Infrastructure as Code (IaC) principles.
- Develop and maintain automation using Ansible and Terraform.
- Administer Linux/RHEL environments supporting GitLab infrastructure.
- Manage containerized environments using Docker and Podman.
- Support and troubleshoot Kubernetes and OpenShift environments.
- Build, maintain, and troubleshoot GitLab Shared Runner infrastructure.
- Configure and maintain monitoring and observability using Prometheus, Grafana, and Datadog.
- Monitor system health, capacity, availability, and performance of GitLab infrastructure.
- Participate in on-call rotations and provide timely response to production incidents.
- Perform incident investigation, troubleshooting, and Root Cause Analysis (RCA).
- Implement preventive measures to minimize recurring incidents and improve platform reliability.
- Collaborate with DevOps, Platform Engineering, Security, and Development teams.
- Automate operational tasks and continuously improve infrastructure reliability and efficiency.
- Maintain technical documentation, operational procedures, and disaster recovery runbooks.
Required Skills
- GitLab administration: upgrades, Geo replication, backup/restore
- Ansible + Terraform (Infrastructure as Code)
- Linux (RHEL), Docker/Podman, K8s/OpenShift
- Prometheus, Grafana, Datadog
- Shared Runner infrastructure
- On-call, incident response, RCA
Seniority Level:
Not Specified