Ovii Job Board

Senior Site Reliability Engineer

Blitzy

Kirtane Baugh, Magarpatta, Hadapsar, Pune, India • Onsite - Kirtane Baugh, Magarpatta, Hadapsar, Pune, India • Full-Time • 5+ years

Posted 2026-05-26 INR 4,800,000 - INR 8,000,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Blitzy seeks a Senior Site Reliability Engineer to design, build, and operate highly available, scalable infrastructure for its AI‑driven development platform. The role blends software engineering with operations, requiring deep ownership of reliability, observability, and automation in a fast‑moving, on‑site environment.

Role Snapshot

  • Design & operate fault‑tolerant multi‑cloud infrastructure
  • Define and enforce SLOs/SLAs, lead blameless postmortems
  • Build and maintain CI/CD and deployment automation
  • Own observability stack (logging, metrics, tracing, alerts)
  • Partner with engineering teams to embed reliability practices
  • Capacity planning, performance benchmarking, cost optimization
  • Champion security best practices in infrastructure

Must-Have Requirements

  • Major cloud platform (AWS preferred)
  • Kubernetes & container orchestration
  • Infrastructure-as-Code (Terraform, Pulumi, or equivalent)
  • Observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Scripting (Python, Go, Bash or similar)
  • Incident management & on‑call practices
  • SLO/SLAs & error‑budget definition
  • Site Reliability Engineering
  • Cloud platform expertise
  • Kubernetes
  • Infrastructure as Code

Nice-to-Have Signals

  • AI/ML or GPU‑accelerated workload support
  • High‑growth startup experience
  • eBPF or service‑mesh (Istio, Linkerd) familiarity
  • Open‑source SRE/DevOps contributions
  • Global multi‑region infrastructure experience
  • AI/ML workload support
  • Startup high‑growth environment
  • eBPF / service‑mesh familiarity

Work Setup

  • Location: Pune, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design, build, and operate scalable, fault‑tolerant infrastructure across major cloud providers.
  • Define, monitor, and enforce service level objectives, error budgets, and lead blameless postmortems.
  • Create and maintain robust CI/CD pipelines and automated deployment workflows.
  • Develop and sustain observability systems including logging, metrics, tracing, and alerting.
  • Collaborate with software engineering teams to embed reliability into the development lifecycle.
  • Perform capacity planning, performance benchmarking, and cost‑optimization initiatives.
  • Drive security hardening and best‑practice adoption across infrastructure layers.
  • Lead major reliability initiatives from concept through production.

Good Fit If You Have

  • Experience supporting AI/ML or GPU‑accelerated workloads.
  • Background in high‑growth startup environments with broad responsibilities.
  • Familiarity with eBPF, service‑mesh technologies (Istio, Linkerd) or advanced networking.
  • Contributions to open‑source SRE/DevOps tooling or communities.
  • Experience building global, multi‑region infrastructure with strict latency targets.

Skills

  • AWS (or GCP/Azure) cloud platforms
  • Kubernetes & container orchestration
  • Infrastructure as Code (Terraform, Pulumi, etc.)
  • Observability tools (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Scripting (Python, Go, Bash or similar)
  • CI/CD pipeline automation
  • Reliability engineering practices (SLO/SLAs, incident management)
  • Security best practices for cloud infrastructure