Key Outcomes & Responsibilities
- Incident Resolution: Own and resolve P1-P3 incidents within SLA; collaborate across teams for swift recovery.
- Problem Management: Perform root cause analysis for recurring issues and implement workarounds or fixes.
- Change Delivery: Execute approved changes, participate in CAB processes, and support release deployments.
- Environment Management: Ensure availability of non-production environments; resolve environment-related tickets.
- Monitoring & Reliability: Enhance monitoring, alerts, and observability to improve service stability.
- Service Reporting: Ensure SLAs are met and update stakeholders on ticket progress.
- Knowledge Sharing: Create runbooks, work instructions, and upskill junior engineers.
- Shift Left: Automate or document repeatable tasks to move work to L1/L2 teams.
Essential Skills (Must Have)
- Strong technical troubleshooting across applications and infrastructure.
- Experience with logs, monitoring tools, Java, JavaScript and API/microservices debugging.
- Good understanding of cloud platforms (AWS).
- Familiarity with CI/CD pipelines and deployment workflows.
- Working knowledge of Scripting (Python/Shell/PowerShell).
Desirable Skills (Nice to Have)
- Knowledge of container platforms (Docker, Kubernetes).
- Infrastructure as Code familiarity (Terraform, CloudFormation).
- Experience with automated testing or performance analysis.
- Ability to contribute to release planning and DevOps automation.
NOTE: Hybrid (Once/twice per week).