Ovii Job Board

Infrastructure SRE - HPC

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5+ years

Posted 2026-06-26 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam seeks an Infrastructure SRE specialist to own the reliability of a large multi‑vendor GPU fleet that powers AI training and inference. The role blends deep expertise in one HPC focus area with broad fluency across storage, networking, Kubernetes, and workload reliability.

Role Snapshot

  • GPU fleet reliability
  • On‑call incident response
  • Distributed storage & fabric expertise
  • Kubernetes platform stewardship
  • Internal tooling development
  • Collaboration with ML teams

Must-Have Requirements

  • Python or Go
  • Kubernetes fluency
  • Parallel filesystem expertise
  • InfiniBand/RDMA networking knowledge
  • infrastructure or site reliability engineering
  • operating GPU clusters at scale

Nice-to-Have Signals

  • Slurm experience
  • On‑premise GPU deployment coordination
  • Multi‑tenant GPU isolation (MIG, MPS)
  • Familiarity with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud
  • deep domain expertise in storage or fabric

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Remote eligibility
  • Education requirement
  • Certifications

What You'll Likely Work On

  • Operate the GPU fleet end‑to‑end – provisioning, monitoring, capacity, and health.
  • Maintain on‑call rotation, author runbooks, and lead post‑mortems that drive lasting fixes.
  • Develop internal tooling to automate fleet operations and incident handling.
  • Partner with ML and platform teams to keep large training runs alive and inference latency predictable.
  • Troubleshoot and route incidents across storage, fabric, GPU, Kubernetes, and workload domains.
  • Manage parallel filesystems, RDMA fabrics, driver/firmware lifecycles, and Kubernetes GPU operator stack.

Good Fit If You Have

  • Depth in at least one focus area (storage, fabric, GPU systems, K8s platform, or workload reliability).
  • Proven on‑call ownership with measurable post‑mortem impact.
  • Ability to communicate across specialties and guide incident triage.

Skills

  • Python or Go
  • Kubernetes
  • Parallel filesystems (Lustre, GPFS, WEKA, BeeGFS)
  • InfiniBand / RDMA networking
  • GPU cluster management
  • Slurm (optional)
  • Automation & observability
  • Capacity planning