Ovii Job Board

Backend Engineer - Studio Media Platform

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 4-6 years

Posted 2026-06-01 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Sarvam seeks a Backend Engineer to build and operate scalable services for its Studio media platform, enabling AI‑driven dubbing, live translation, and other multilingual media capabilities. The role combines production‑grade FastAPI services, async pipelines, Kubernetes deployments, and ML model integration for enterprise customers.

Role Snapshot

  • Build and scale production FastAPI services
  • Design async, distributed pipelines for audio/media processing
  • Manage Kubernetes/Helm deployments and cloud infrastructure
  • Integrate ML models and maintain SDKs for internal consumption
  • Ensure observability, testing, and CI/CD best practices

Must-Have Requirements

  • Python
  • FastAPI or similar async web framework
  • Async programming (asyncio, concurrency patterns)
  • Distributed task systems (Celery or similar)
  • PostgreSQL with async ORM (SQLAlchemy preferred)
  • Docker
  • Kubernetes & Helm
  • Cloud platform (AWS/GCP/Azure)
  • Testing discipline and CI/CD pipelines
  • ML model integration (REST/gRPC, PyTorch, ONNX Runtime)
  • backend engineering
  • production services at scale
  • async programming
  • distributed task systems
  • PostgreSQL

Nice-to-Have Signals

  • Speech/NLP systems (ASR, TTS, machine translation)
  • Model serving infrastructure (Triton, TorchServe)
  • LLM orchestration for structured output
  • Real‑time streaming (WebSocket, WebRTC, SSE)
  • Observability tooling (Prometheus, Grafana, OpenTelemetry)
  • Video processing pipelines
  • Open‑source contributions in backend/audio/NLP
  • Audio/media processing libraries (FFmpeg, librosa, soundfile)
  • speech/NLP systems
  • model serving infrastructure
  • LLM orchestration
  • real‑time streaming

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design and optimize FastAPI services for dubbing and live translation, including task orchestration and rate‑limiting
  • Build distributed worker architectures with independent scaling and automatic recovery
  • Own the data layer: async ORM models, schema migrations, and PostgreSQL query optimization
  • Implement real‑time features such as WebSocket job tracking and streaming audio pipelines
  • Manage Kubernetes deployments via Helm charts, secrets, and ingress configuration
  • Extend core dubbing library across audio extraction, VAD, ASR, translation, QC, TTS, and video stitching
  • Integrate and serve ML models both remotely and locally, and maintain LLM integration layers
  • Develop and maintain a shared SDK with middleware for authentication, billing, workspace isolation, and input validation
  • Create media storage abstractions, rate limiting, metering, audit trails, and request history
  • Build observability foundations with OpenTelemetry, structured logging, and metrics collection

Good Fit If You Have

  • Experience with speech/NLP systems (ASR, TTS, translation) is a plus
  • Familiarity with real‑time streaming protocols (WebSocket, WebRTC, SSE) is advantageous
  • Background in building platform middleware such as auth, billing, or multi‑tenant isolation

Skills

  • Python
  • FastAPI / async web framework
  • Async programming (asyncio, concurrency patterns)
  • Distributed task systems (Celery or similar)
  • PostgreSQL with async ORM (SQLAlchemy)
  • Docker & Kubernetes (Helm)
  • Cloud platforms (AWS/GCP/Azure)
  • Testing discipline & CI/CD pipelines
  • ML model integration (PyTorch, ONNX Runtime)
  • Audio processing (FFmpeg, librosa) – optional
  • Observability tooling (OpenTelemetry, Prometheus, Grafana) – preferred