
Senior DevOps Engineer
GoKwik
Posted 2026-06-02
INR 2,500,000 - INR 3,500,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
The Senior DevOps Engineer will own Site Reliability Engineering for GoKwik’s production platform, driving observability, incident response, and automation. The role blends deep SRE practice with DevOps enablement to keep e‑commerce services highly available and performant.
Role Snapshot
- Lead SRE practices
- Own on‑call and incident response
- Build observability systems
- Automate infrastructure with Terraform & Kubernetes
- Support CI/CD pipelines
- Mentor junior engineers
- Partner on resilient architecture reviews
- Drive post‑mortem and prevention culture
Must-Have Requirements
- SRE / Production Engineering experience
- Incident management and on‑call operations
- Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
- AWS or GCP cloud infrastructure
- Kubernetes
- Terraform
- CI/CD pipeline tooling
- Scripting (Shell, Python, Go)
- SRE / Production Engineering
- incident management
- observability
Nice-to-Have Signals
- Chaos engineering
- Resiliency testing
- Mentoring junior engineers
- chaos engineering
- mentoring
- eCommerce
Work Setup
- Location: Gurugram, India
- Work mode: ONSITE
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Define and evolve SRE practices for reliability, scaling and performance
- Lead on‑call rotations and manage incident response to minimize impact
- Design and implement advanced observability (metrics, logs, traces, APM, SLIs/SLOs)
- Automate self‑healing, scalable infrastructure using Terraform and Kubernetes
- Support and improve CI/CD pipelines and infrastructure automation
- Mentor junior engineers on incident handling and DevOps best practices
- Collaborate with engineering teams on resilient architecture reviews
- Conduct blameless post‑mortems and evolve incident playbooks
Good Fit If You Have
- Proven track record in incident management and debugging distributed systems
- Hands‑on experience with Kubernetes and Terraform in production
- Deep familiarity with observability stacks such as Prometheus or Datadog
- Ability to mentor engineers and champion reliability culture
- Comfort balancing urgent firefighting with long‑term reliability initiatives
Skills
- Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
- Kubernetes
- Terraform
- CI/CD pipeline tooling
- AWS / GCP cloud infrastructure
- Incident management & debugging
- Scripting (Shell, Python, Go)
- Chaos engineering & resiliency testing
- SLI / SLO design