
Senior DevOps Engineer
GoKwik
Posted 2026-07-23
INR 2,500,000 - INR 3,500,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Senior DevOps Engineer focused on Site Reliability Engineering for a fast‑growing eCommerce platform. Own on‑call, observability and automation to keep production systems resilient and performant.
Role Snapshot
- Lead SRE practices and reliability roadmap
- Own on‑call and incident response
- Design self‑healing, scalable infrastructure
- Build and evolve observability (metrics, logs, traces, SLIs/SLOs)
- Automate CI/CD pipelines with Terraform & Kubernetes
- Mentor junior engineers on reliability best practices
- Drive adoption of new reliability tools and processes
Must-Have Requirements
- SRE / Production Engineering experience
- Incident management & on‑call operations
- Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
- AWS or GCP cloud infrastructure
- Kubernetes
- Terraform
- CI/CD pipelines
- Shell/Python/Go scripting
- SRE / Production Engineering
- incident management
- observability
Nice-to-Have Signals
- Chaos engineering
- HA/DR setups
- Distributed tracing & APM
- Cloud networking knowledge
- chaos engineering
- cloud networking
Work Setup
- Location: Gurugram, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Lead SRE practices for scaling and performance of production systems
- Run on‑call duties, troubleshoot incidents and conduct blameless postmortems
- Create automated, self‑healing infrastructure using Terraform and Kubernetes
- Implement advanced observability (metrics, logs, traces, SLIs/SLOs, APM)
- Support and improve CI/CD pipelines and automation workflows
- Mentor junior engineers in reliability and DevOps best practices
- Partner with engineering teams on resilient architecture reviews
- Champion new tools and processes to increase infrastructure reliability
Good Fit If You Have
- Experience debugging distributed systems in production
- Familiarity with cloud networking, HA/DR setups
- Comfort balancing short‑term firefighting with long‑term reliability goals
Skills
- Site Reliability Engineering (SRE)
- Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
- Kubernetes orchestration
- Terraform infrastructure as code
- CI/CD pipeline automation
- AWS / GCP cloud services
- Shell / Python / Go scripting
- Incident management & blameless postmortems
- Chaos engineering (preferred)
- HA/DR design (preferred)