ML Engineer (Training Infra), Foundational Models
Sarvam
Posted 2026-05-21
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Sarvam seeks an ML Engineer to own and evolve the distributed training platform for its next generation of foundational models. The role blends deep systems engineering, GPU kernel work, and reliability engineering to keep large‑scale training runs fast and stable.
Role Snapshot
- Own distributed training infrastructure for foundational models
- Design parallelism strategies across GPU clusters
- Optimize GPU kernel performance and training throughput
- Ensure reliability of long‑running training jobs
- Collaborate with researchers and data teams
Must-Have Requirements
- Distributed training frameworks
- GPU kernel development (CUDA/Triton)
- Deep knowledge of GPU architecture
- PyTorch internals
- BS or MS in Computer Science or related field
- ML training infrastructure
- large‑scale distributed systems
- BS or MS in Computer Science or a closely related technical field
Nice-to-Have Signals
- Custom CUDA/Triton kernel development with measurable wins
- Training 10B+ parameter models or 1000+ GPU clusters
- Cluster orchestration (Slurm/Kubernetes)
- Mixed‑precision or quantization‑aware training
- First‑author papers on training systems
- Familiarity with cluster orchestration and job scheduling
- Experience with mixed precision (BF16, FP8)
- training 10B+ models
- operating 1000+ GPU clusters
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
What You'll Likely Work On
- Build and maintain a high‑performance distributed training stack across large GPU clusters
- Create and evaluate data, tensor, pipeline, sequence, and expert parallelism strategies
- Profile end‑to‑end training pipelines and optimize kernel performance, communication overlap, memory layout, checkpointing, and data loading
- Write custom CUDA/Triton kernels when off‑the‑shelf solutions fall short
- Own fault tolerance, checkpoint integrity, deterministic restarts, and automated detection of slow or corrupted nodes
- Partner with researchers to translate architectural ideas into efficient training runs and with data engineers to eliminate pipeline bottlenecks
Good Fit If You Have
- Open‑source contributions to Megatron, DeepSpeed, PyTorch, vLLM, Triton, or NCCL
- Experience training models >10 B parameters or on >1 000 GPU clusters
- Familiarity with large‑scale cluster orchestration and job scheduling
- Hands‑on work with mixed‑precision (BF16, FP8) or quantization‑aware training
- First‑author papers on training systems or model efficiency
Skills
- Distributed training frameworks (Megatron‑LM, DeepSpeed, FSDP, NeMo)
- GPU kernel development (CUDA, Triton)
- GPU architecture & profiling tools (Nsight, PyTorch profiler)
- PyTorch internals
- Large‑scale systems design
- Cluster orchestration (Slurm, Kubernetes)
- Mixed‑precision and quantization‑aware training
- Open‑source contributions to training ecosystem