Ovii Job Board

ML Ops Engineer, Chanakya

Sarvam

Delhi, India • Onsite - Delhi, India • Full-Time • 3-5 years

Posted 2026-04-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam seeks an MLOps Engineer to own the full model lifecycle for strategic deployments, building and operating serving infrastructure, CI/CD pipelines, monitoring, and incident response. The role works on‑prem and cloud environments, collaborating closely with data scientists and deployment engineers.

Role Snapshot

  • Own model lifecycle across deployments
  • Design & operate serving infrastructure
  • Build CI/CD pipelines for model updates
  • Monitor latency, accuracy drift, and throughput
  • Create evaluation harnesses and A/B testing tools
  • Manage containerised serving in edge/air‑gapped settings
  • Write runbooks and lead incident response
  • Partner with data scientists on eval pipelines

Must-Have Requirements

  • Model serving platforms (vLLM, Triton, TGI)
  • Containerisation (Docker, Kubernetes)
  • Monitoring/observability (Prometheus, Grafana)
  • Python programming
  • CI/CD tooling for ML (GitHub Actions, ArgoCD, DVC)
  • ML engineering
  • MLOps
  • production LLM systems

Nice-to-Have Signals

  • Quantized model formats (GGUF, AWQ, GPTQ)
  • Edge & air‑gapped deployment experience
  • Evaluation infrastructure (A/B testing, harnesses)
  • Lightweight k8s alternatives (K3s, K0s)
  • quantized model formats
  • edge/air‑gapped environments

Work Setup

  • Location: Delhi, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design and operate model serving infrastructure for on‑prem and cloud environments
  • Build and maintain CI/CD pipelines that support model updates, rollbacks, and evaluation‑gated releases
  • Monitor production model performance—latency, accuracy drift, throughput—and build proactive alerting
  • Develop evaluation harnesses, A/B testing frameworks, and model comparison tooling
  • Manage containerised model serving in constrained, air‑gapped, and edge environments
  • Author runbooks and operational playbooks for field deployment engineers
  • Own incident response for model‑layer failures across all active deployments
  • Collaborate with data scientists to own the underlying evaluation infrastructure

Good Fit If You Have

  • Hands‑on experience keeping a production LLM system running under load
  • Proactive mindset: builds systems that surface issues before they impact clients
  • Strong documentation habits that are used by non‑author teammates

Skills

  • Model serving platforms (vLLM, Triton, TGI)
  • Containerisation (Docker, Kubernetes, K3s/K0s)
  • CI/CD for ML (GitHub Actions, ArgoCD, DVC)
  • Observability (Prometheus, Grafana)
  • Python programming
  • Quantized model formats (GGUF, AWQ, GPTQ)
  • Edge & air‑gapped deployment
  • Evaluation infrastructure (A/B testing, harnesses)