Job Description
Ovii's Interpretation of the Role
Sarvam is hiring a Backend Engineer for its API team to build high‑performance, cloud‑agnostic Python APIs that serve AI/ML models at scale. The role involves designing low‑latency, fault‑tolerant services, integrating data stores and streaming platforms, and ensuring secure, observable production deployments.
Role Snapshot
- Build low‑latency Python APIs for ML inference
- Design fault‑tolerant, cloud‑agnostic backend services
- Implement authentication, rate limiting, and security
- Integrate PostgreSQL, Redis, ClickHouse, and streaming
- Collaborate on CI/CD, canary releases, and observability
Must-Have Requirements
- Python
- FastAPI/Django/Flask (Python web frameworks)
- HTTP/WebSockets/gRPC protocols
- Low‑latency distributed backend design
- PostgreSQL/Redis/ClickHouse
- Docker
- Kubernetes
- CI/CD pipelines
- Authentication & authorization security
- At least one major cloud platform (Azure preferred)
- building low‑latency distributed backend systems
- Python API development for ML models
- cloud‑agnostic infrastructure
Nice-to-Have Signals
- Canary deployments / progressive rollouts
- Observability tools (Prometheus, Grafana, OpenTelemetry)
- Prior ML inference serving experience
- Open‑source contributions / GitHub portfolio
- Kafka or Redis Streams (familiarity)
- ML inference serving
- canary or progressive rollouts
- observability and monitoring
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Relocation
- Notice period
What You'll Likely Work On
- Design, develop and optimise Python‑based APIs for serving ML models at scale
- Build communication layers using HTTP, WebSockets, and gRPC
- Architect low‑latency, fault‑tolerant, secure backend systems for real‑time inference
- Implement authentication, rate limiting, prioritisation, and secure coding practices
- Integrate voice‑agent, LLM, and other AI SDKs
- Manage data storage and performance with PostgreSQL, Redis, and ClickHouse
- Create event‑driven streaming pipelines using Kafka or Redis Streams
- Support canary deployments, feature rollouts, and CI/CD pipelines
- Ensure observability, reliability, and vendor‑agnostic operation across Azure, AWS, GCP, and on‑prem
Good Fit If You Have
- Prior experience with ML model serving or inference infrastructure
- Familiarity with canary or progressive rollout techniques
- Experience using observability tools such as Prometheus, Grafana, or OpenTelemetry
- Contributions to open‑source backend projects or a strong GitHub portfolio
- Hands‑on work with Azure (preferred) or other cloud platforms
Skills
- Python
- FastAPI/Django/Flask (Python web frameworks)
- HTTP, WebSockets, gRPC (protocols)
- Low‑latency distributed backend design
- PostgreSQL, Redis, ClickHouse (databases)
- Docker & Kubernetes (container orchestration)
- CI/CD pipelines
- Azure (or other major cloud platform)
- Kafka or Redis Streams (event streaming)
- API authentication & authorization