Ovii Job Board

Senior DevOps Engineer

GoKwik

Gurugram, India • Onsite - Gurugram, India • Full-Time • 4-7 years

Posted 2026-07-23 INR 2,500,000 - INR 3,500,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior DevOps Engineer focused on Site Reliability Engineering for a fast‑growing eCommerce platform. Own on‑call, observability and automation to keep production systems resilient and performant.

Role Snapshot

  • Lead SRE practices and reliability roadmap
  • Own on‑call and incident response
  • Design self‑healing, scalable infrastructure
  • Build and evolve observability (metrics, logs, traces, SLIs/SLOs)
  • Automate CI/CD pipelines with Terraform & Kubernetes
  • Mentor junior engineers on reliability best practices
  • Drive adoption of new reliability tools and processes

Must-Have Requirements

  • SRE / Production Engineering experience
  • Incident management & on‑call operations
  • Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
  • AWS or GCP cloud infrastructure
  • Kubernetes
  • Terraform
  • CI/CD pipelines
  • Shell/Python/Go scripting
  • SRE / Production Engineering
  • incident management
  • observability

Nice-to-Have Signals

  • Chaos engineering
  • HA/DR setups
  • Distributed tracing & APM
  • Cloud networking knowledge
  • chaos engineering
  • cloud networking

Work Setup

  • Location: Gurugram, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Lead SRE practices for scaling and performance of production systems
  • Run on‑call duties, troubleshoot incidents and conduct blameless postmortems
  • Create automated, self‑healing infrastructure using Terraform and Kubernetes
  • Implement advanced observability (metrics, logs, traces, SLIs/SLOs, APM)
  • Support and improve CI/CD pipelines and automation workflows
  • Mentor junior engineers in reliability and DevOps best practices
  • Partner with engineering teams on resilient architecture reviews
  • Champion new tools and processes to increase infrastructure reliability

Good Fit If You Have

  • Experience debugging distributed systems in production
  • Familiarity with cloud networking, HA/DR setups
  • Comfort balancing short‑term firefighting with long‑term reliability goals

Skills

  • Site Reliability Engineering (SRE)
  • Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Kubernetes orchestration
  • Terraform infrastructure as code
  • CI/CD pipeline automation
  • AWS / GCP cloud services
  • Shell / Python / Go scripting
  • Incident management & blameless postmortems
  • Chaos engineering (preferred)
  • HA/DR design (preferred)