Job Description
Ovii's Interpretation of the Role
Sarvam seeks an MLOps Engineer to own the full model lifecycle for strategic deployments, building and operating serving infrastructure, CI/CD pipelines, monitoring, and incident response. The role works on‑prem and cloud environments, collaborating closely with data scientists and deployment engineers.
Role Snapshot
- Own model lifecycle across deployments
- Design & operate serving infrastructure
- Build CI/CD pipelines for model updates
- Monitor latency, accuracy drift, and throughput
- Create evaluation harnesses and A/B testing tools
- Manage containerised serving in edge/air‑gapped settings
- Write runbooks and lead incident response
- Partner with data scientists on eval pipelines
Must-Have Requirements
- Model serving platforms (vLLM, Triton, TGI)
- Containerisation (Docker, Kubernetes)
- Monitoring/observability (Prometheus, Grafana)
- Python programming
- CI/CD tooling for ML (GitHub Actions, ArgoCD, DVC)
- ML engineering
- MLOps
- production LLM systems
Nice-to-Have Signals
- Quantized model formats (GGUF, AWQ, GPTQ)
- Edge & air‑gapped deployment experience
- Evaluation infrastructure (A/B testing, harnesses)
- Lightweight k8s alternatives (K3s, K0s)
- quantized model formats
- edge/air‑gapped environments
Work Setup
- Location: Delhi, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design and operate model serving infrastructure for on‑prem and cloud environments
- Build and maintain CI/CD pipelines that support model updates, rollbacks, and evaluation‑gated releases
- Monitor production model performance—latency, accuracy drift, throughput—and build proactive alerting
- Develop evaluation harnesses, A/B testing frameworks, and model comparison tooling
- Manage containerised model serving in constrained, air‑gapped, and edge environments
- Author runbooks and operational playbooks for field deployment engineers
- Own incident response for model‑layer failures across all active deployments
- Collaborate with data scientists to own the underlying evaluation infrastructure
Good Fit If You Have
- Hands‑on experience keeping a production LLM system running under load
- Proactive mindset: builds systems that surface issues before they impact clients
- Strong documentation habits that are used by non‑author teammates
Skills
- Model serving platforms (vLLM, Triton, TGI)
- Containerisation (Docker, Kubernetes, K3s/K0s)
- CI/CD for ML (GitHub Actions, ArgoCD, DVC)
- Observability (Prometheus, Grafana)
- Python programming
- Quantized model formats (GGUF, AWQ, GPTQ)
- Edge & air‑gapped deployment
- Evaluation infrastructure (A/B testing, harnesses)