Ovii Job Board

Infrastructure SRE - HPC

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5+ years

Posted 2026-06-26 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam seeks an Infrastructure SRE specialist to own the reliability of a large multi‑vendor GPU fleet that powers both massive training jobs and latency‑critical inference services. The role blends deep expertise in one HPC domain with broad fluency across storage, networking, and orchestration, delivering tooling, capacity planning, and on‑call incident response.

Role Snapshot

  • GPU fleet reliability
  • On‑call incident ownership
  • Internal tooling development
  • Capacity & observability
  • Cross‑team ML platform partnership
  • Deep expertise in storage/fabric/network

Must-Have Requirements

  • Python
  • Go
  • GPU fleet operation
  • On‑call ownership
  • Runbook and post‑mortem creation
  • Internal tooling development
  • infrastructure/site reliability engineering
  • GPU cluster operations at scale

Nice-to-Have Signals

  • Slurm
  • Kubernetes hybrid environments
  • On‑prem GPU deployment
  • InfiniBand cabling coordination
  • Multi‑tenant GPU isolation (MIG, MPS)
  • Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud
  • deep storage or fabric expertise
  • on‑prem GPU deployment

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Operate the end‑to‑end GPU fleet, handling provisioning, monitoring, capacity planning, and health checks for both training and inference workloads.
  • Participate in a rigorous on‑call rotation, author runbooks, and lead post‑mortems that drive lasting fixes.
  • Design and build internal tooling to automate fleet management and incident response.
  • Collaborate with ML and platform teams to keep large training runs alive and serving latency predictable.
  • Diagnose issues across storage, networking, and compute layers, routing incidents to the appropriate specialist.

Good Fit If You Have

  • Experience with Slurm or hybrid Kubernetes environments.
  • Hands‑on work with on‑prem GPU deployments, power/cooling coordination, or InfiniBand cabling.
  • Familiarity with Indian cloud providers (NCPs) or large‑scale GPU solutions such as DGX SuperPOD, Lambda, CoreWeave, NeevCloud.

Skills

  • Python
  • Go
  • GPU cluster operations
  • Parallel filesystems
  • RDMA fabrics
  • NCCL troubleshooting
  • Distributed training systems
  • Slurm (preferred)
  • Kubernetes hybrid environments (preferred)
  • On‑prem GPU deployment (preferred)
  • InfiniBand cabling coordination (preferred)
  • Multi‑tenant GPU isolation (MIG/MPS) (bonus)