Ovii Job Board

Senior DevOps Engineer

GoKwik

Gurugram, India • Onsite - Gurugram, India • Full-Time • 5-8 years

Posted 2026-06-02 INR 2,500,000 - INR 3,500,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Senior DevOps Engineer will own Site Reliability Engineering for GoKwik’s production platform, driving observability, incident response, and automation. The role blends deep SRE practice with DevOps enablement to keep e‑commerce services highly available and performant.

Role Snapshot

  • Lead SRE practices
  • Own on‑call and incident response
  • Build observability systems
  • Automate infrastructure with Terraform & Kubernetes
  • Support CI/CD pipelines
  • Mentor junior engineers
  • Partner on resilient architecture reviews
  • Drive post‑mortem and prevention culture

Must-Have Requirements

  • SRE / Production Engineering experience
  • Incident management and on‑call operations
  • Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
  • AWS or GCP cloud infrastructure
  • Kubernetes
  • Terraform
  • CI/CD pipeline tooling
  • Scripting (Shell, Python, Go)
  • SRE / Production Engineering
  • incident management
  • observability

Nice-to-Have Signals

  • Chaos engineering
  • Resiliency testing
  • Mentoring junior engineers
  • chaos engineering
  • mentoring
  • eCommerce

Work Setup

  • Location: Gurugram, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Define and evolve SRE practices for reliability, scaling and performance
  • Lead on‑call rotations and manage incident response to minimize impact
  • Design and implement advanced observability (metrics, logs, traces, APM, SLIs/SLOs)
  • Automate self‑healing, scalable infrastructure using Terraform and Kubernetes
  • Support and improve CI/CD pipelines and infrastructure automation
  • Mentor junior engineers on incident handling and DevOps best practices
  • Collaborate with engineering teams on resilient architecture reviews
  • Conduct blameless post‑mortems and evolve incident playbooks

Good Fit If You Have

  • Proven track record in incident management and debugging distributed systems
  • Hands‑on experience with Kubernetes and Terraform in production
  • Deep familiarity with observability stacks such as Prometheus or Datadog
  • Ability to mentor engineers and champion reliability culture
  • Comfort balancing urgent firefighting with long‑term reliability initiatives

Skills

  • Observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Kubernetes
  • Terraform
  • CI/CD pipeline tooling
  • AWS / GCP cloud infrastructure
  • Incident management & debugging
  • Scripting (Shell, Python, Go)
  • Chaos engineering & resiliency testing
  • SLI / SLO design