Ovii Job Board

Backend Engineer, API Team

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time

Posted 2026-04-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam is hiring a Backend Engineer for its API team to build high‑performance, cloud‑agnostic Python APIs that serve AI/ML models at scale. The role involves designing low‑latency, fault‑tolerant services, integrating data stores and streaming platforms, and ensuring secure, observable production deployments.

Role Snapshot

  • Build low‑latency Python APIs for ML inference
  • Design fault‑tolerant, cloud‑agnostic backend services
  • Implement authentication, rate limiting, and security
  • Integrate PostgreSQL, Redis, ClickHouse, and streaming
  • Collaborate on CI/CD, canary releases, and observability

Must-Have Requirements

  • Python
  • FastAPI/Django/Flask (Python web frameworks)
  • HTTP/WebSockets/gRPC protocols
  • Low‑latency distributed backend design
  • PostgreSQL/Redis/ClickHouse
  • Docker
  • Kubernetes
  • CI/CD pipelines
  • Authentication & authorization security
  • At least one major cloud platform (Azure preferred)
  • building low‑latency distributed backend systems
  • Python API development for ML models
  • cloud‑agnostic infrastructure

Nice-to-Have Signals

  • Canary deployments / progressive rollouts
  • Observability tools (Prometheus, Grafana, OpenTelemetry)
  • Prior ML inference serving experience
  • Open‑source contributions / GitHub portfolio
  • Kafka or Redis Streams (familiarity)
  • ML inference serving
  • canary or progressive rollouts
  • observability and monitoring

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Relocation
  • Notice period

What You'll Likely Work On

  • Design, develop and optimise Python‑based APIs for serving ML models at scale
  • Build communication layers using HTTP, WebSockets, and gRPC
  • Architect low‑latency, fault‑tolerant, secure backend systems for real‑time inference
  • Implement authentication, rate limiting, prioritisation, and secure coding practices
  • Integrate voice‑agent, LLM, and other AI SDKs
  • Manage data storage and performance with PostgreSQL, Redis, and ClickHouse
  • Create event‑driven streaming pipelines using Kafka or Redis Streams
  • Support canary deployments, feature rollouts, and CI/CD pipelines
  • Ensure observability, reliability, and vendor‑agnostic operation across Azure, AWS, GCP, and on‑prem

Good Fit If You Have

  • Prior experience with ML model serving or inference infrastructure
  • Familiarity with canary or progressive rollout techniques
  • Experience using observability tools such as Prometheus, Grafana, or OpenTelemetry
  • Contributions to open‑source backend projects or a strong GitHub portfolio
  • Hands‑on work with Azure (preferred) or other cloud platforms

Skills

  • Python
  • FastAPI/Django/Flask (Python web frameworks)
  • HTTP, WebSockets, gRPC (protocols)
  • Low‑latency distributed backend design
  • PostgreSQL, Redis, ClickHouse (databases)
  • Docker & Kubernetes (container orchestration)
  • CI/CD pipelines
  • Azure (or other major cloud platform)
  • Kafka or Redis Streams (event streaming)
  • API authentication & authorization