Job Description
Ovii's Interpretation of the Role
Sarvam seeks an Infrastructure SRE specialist to own the reliability of a large multi‑vendor GPU fleet that powers both massive training jobs and latency‑critical inference services. The role blends deep expertise in one HPC domain with broad fluency across storage, networking, and orchestration, delivering tooling, capacity planning, and on‑call incident response.
Role Snapshot
- GPU fleet reliability
- On‑call incident ownership
- Internal tooling development
- Capacity & observability
- Cross‑team ML platform partnership
- Deep expertise in storage/fabric/network
Must-Have Requirements
- Python
- Go
- GPU fleet operation
- On‑call ownership
- Runbook and post‑mortem creation
- Internal tooling development
- infrastructure/site reliability engineering
- GPU cluster operations at scale
Nice-to-Have Signals
- Slurm
- Kubernetes hybrid environments
- On‑prem GPU deployment
- InfiniBand cabling coordination
- Multi‑tenant GPU isolation (MIG, MPS)
- Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud
- deep storage or fabric expertise
- on‑prem GPU deployment
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Operate the end‑to‑end GPU fleet, handling provisioning, monitoring, capacity planning, and health checks for both training and inference workloads.
- Participate in a rigorous on‑call rotation, author runbooks, and lead post‑mortems that drive lasting fixes.
- Design and build internal tooling to automate fleet management and incident response.
- Collaborate with ML and platform teams to keep large training runs alive and serving latency predictable.
- Diagnose issues across storage, networking, and compute layers, routing incidents to the appropriate specialist.
Good Fit If You Have
- Experience with Slurm or hybrid Kubernetes environments.
- Hands‑on work with on‑prem GPU deployments, power/cooling coordination, or InfiniBand cabling.
- Familiarity with Indian cloud providers (NCPs) or large‑scale GPU solutions such as DGX SuperPOD, Lambda, CoreWeave, NeevCloud.
Skills
- Python
- Go
- GPU cluster operations
- Parallel filesystems
- RDMA fabrics
- NCCL troubleshooting
- Distributed training systems
- Slurm (preferred)
- Kubernetes hybrid environments (preferred)
- On‑prem GPU deployment (preferred)
- InfiniBand cabling coordination (preferred)
- Multi‑tenant GPU isolation (MIG/MPS) (bonus)