Job Description
Ovii's Interpretation of the Role
Sarvam seeks a Platform Engineer to design and ship the control‑plane, scheduling, scaling and self‑service layers that power its sovereign AI GPU fleet. The role blends deep systems work in Kubernetes, Go/Python, and GPU‑specific constraints with a product mindset for internal ML engineers.
Role Snapshot
- Build AI platform control‑plane services
- Design scheduler & autoscaling for GPU workloads
- Implement multi‑tenant RBAC, quota and isolation
- Create CLI, SDK and APIs for ML engineers
- Develop observability, cost and networking tooling
- Provision clusters via IaC (Terraform / Crossplane)
Must-Have Requirements
- Go
- Python
- Kubernetes operators/controllers
- GPU platform constraints (MIG, gang scheduling, topology‑aware placement)
- building control‑plane services
- Kubernetes operator development
- GPU‑specific platform constraints
Nice-to-Have Signals
- Serving/inference platform experience
- GPU scheduling systems (Kueue, Volcano, Slurm)
- Deep Kubernetes networking (CNI, RDMA, SR‑IOV)
- On‑prem multi‑vendor GPU platform work
- Open‑source contributions to Kubernetes or GPU projects
- Terraform or Crossplane
- Observability stack (metrics, logging, tracing)
- serving/inference platform at scale
- multi‑tenant isolation mechanisms
- deep networking (CNI, RDMA)
- on‑prem GPU platform deployments
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Create the serving platform that turns model artifacts into scalable, multi‑tenant endpoints with rollout and traffic‑splitting capabilities
- Build autoscaling and elasticity layers for both training and inference, handling burst traffic, preemption and efficient GPU bin‑packing
- Develop the scheduler layer (Kueue, Volcano, Slurm or custom controllers) with gang scheduling, priority, quota enforcement and topology‑aware placement
- Implement multi‑tenant isolation, RBAC, quota and self‑service tiers for MIG/MPS and time‑slicing
- Engineer platform‑level networking, including CNI configuration, RDMA/SR‑IOV exposure and multi‑cluster connectivity
- Provide fleet‑scale observability and cost tooling for metrics, logs and traces
- Design storage abstractions over parallel filesystems, caching tiers and data‑locality‑aware volume provisioning
- Deliver developer experience via CLI, SDK and APIs, plus documentation for self‑service job submission
Good Fit If You Have
- Prior experience building a serving or inference platform at scale
- Hands‑on work with GPU scheduling systems such as Kueue, Volcano or Slurm
- Deep knowledge of Kubernetes networking internals and RDMA/SR‑IOV
- Experience with on‑prem multi‑vendor GPU environments
Skills
- Go
- Python
- Kubernetes operators & controllers
- GPU scheduling (MIG, gang scheduling, topology‑aware placement)
- Autoscaling & elasticity
- Infrastructure‑as‑code (Terraform, Crossplane)
- Observability (metrics, logging, tracing)
- CNI / RDMA / SR‑IOV networking