Job Description: Core Skills & Experience
- Microsoft Azure Expert
- Azure Administration and Architecture
- Networking and Connectivity
- Virtual Machines and Platform Services
- Azure Monitor and Log Analytics
- Azure Identity and Access Management
- Azure Networking, NSGs, Firewalls, Load Balancing
- Azure Backup and Disaster Recovery
- Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
- Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)
- Exposure and understanding of building, deploying and managing API Gateways
- Strong understanding of Azure security controls, governance, and compliance frameworks.
- Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.
- Strong FinOps expertise
- Scripting skills (PowerShell, Bash, Python, React).
- Strong understanding of Devops practices, tooling, and SDLC methods
- Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.
- Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
- Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
- Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
- Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
- Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
- Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
- Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
- Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
- Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.
Thanks & Regards
Saravanan
DMinds Solutions Inc.