Ovii Job Board

Platform Engineer - AI Infrastructure

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5+ years

Posted 2026-06-29 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam seeks a Platform Engineer to design and ship the control‑plane, scheduling, scaling and self‑service layers that power its sovereign AI GPU fleet. The role blends deep systems work in Kubernetes, Go/Python, and GPU‑specific constraints with a product mindset for internal ML engineers.

Role Snapshot

  • Build AI platform control‑plane services
  • Design scheduler & autoscaling for GPU workloads
  • Implement multi‑tenant RBAC, quota and isolation
  • Create CLI, SDK and APIs for ML engineers
  • Develop observability, cost and networking tooling
  • Provision clusters via IaC (Terraform / Crossplane)

Must-Have Requirements

  • Go
  • Python
  • Kubernetes operators/controllers
  • GPU platform constraints (MIG, gang scheduling, topology‑aware placement)
  • building control‑plane services
  • Kubernetes operator development
  • GPU‑specific platform constraints

Nice-to-Have Signals

  • Serving/inference platform experience
  • GPU scheduling systems (Kueue, Volcano, Slurm)
  • Deep Kubernetes networking (CNI, RDMA, SR‑IOV)
  • On‑prem multi‑vendor GPU platform work
  • Open‑source contributions to Kubernetes or GPU projects
  • Terraform or Crossplane
  • Observability stack (metrics, logging, tracing)
  • serving/inference platform at scale
  • multi‑tenant isolation mechanisms
  • deep networking (CNI, RDMA)
  • on‑prem GPU platform deployments

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Create the serving platform that turns model artifacts into scalable, multi‑tenant endpoints with rollout and traffic‑splitting capabilities
  • Build autoscaling and elasticity layers for both training and inference, handling burst traffic, preemption and efficient GPU bin‑packing
  • Develop the scheduler layer (Kueue, Volcano, Slurm or custom controllers) with gang scheduling, priority, quota enforcement and topology‑aware placement
  • Implement multi‑tenant isolation, RBAC, quota and self‑service tiers for MIG/MPS and time‑slicing
  • Engineer platform‑level networking, including CNI configuration, RDMA/SR‑IOV exposure and multi‑cluster connectivity
  • Provide fleet‑scale observability and cost tooling for metrics, logs and traces
  • Design storage abstractions over parallel filesystems, caching tiers and data‑locality‑aware volume provisioning
  • Deliver developer experience via CLI, SDK and APIs, plus documentation for self‑service job submission

Good Fit If You Have

  • Prior experience building a serving or inference platform at scale
  • Hands‑on work with GPU scheduling systems such as Kueue, Volcano or Slurm
  • Deep knowledge of Kubernetes networking internals and RDMA/SR‑IOV
  • Experience with on‑prem multi‑vendor GPU environments

Skills

  • Go
  • Python
  • Kubernetes operators & controllers
  • GPU scheduling (MIG, gang scheduling, topology‑aware placement)
  • Autoscaling & elasticity
  • Infrastructure‑as‑code (Terraform, Crossplane)
  • Observability (metrics, logging, tracing)
  • CNI / RDMA / SR‑IOV networking