Job Description
Ovii's Interpretation of the Role
We need a Machine Learning Engineer to own the integration layer for dubbing and live translation, building end‑to‑end speech pipelines that connect ASR, translation, TTS and voice‑cloning. The role is on‑site in Bengaluru, full‑time, and requires strong Python, PyTorch and real‑time streaming expertise.
Role Snapshot
- Own ML integration layer for dubbing & live translation
- Build production pipelines linking ASR, translation, TTS, voice‑cloning
- Optimize latency for real‑time speech‑to‑speech
- Design fan‑out architectures for multi‑listener streams
- Implement automated quality‑control loops
Must-Have Requirements
- Python
- PyTorch
- Asyncio / async Python
- Speech model integration (ASR/TTS)
- Real‑time streaming systems
- Audio signal processing fundamentals
- Undergraduate degree in technical discipline
- Python development
- speech model integration
- real‑time streaming systems
- audio signal processing
- Undergraduate degree in Computer Science, Electrical Engineering, Statistics, Physics, or equivalent
Nice-to-Have Signals
- Multilingual / Indic speech systems
- Voice cloning or speaker adaptation
- Real‑time media protocols (WebRTC, RTMP, SRT, HLS)
- Multi‑participant audio systems (diarization, concurrent routing)
- FFmpeg or GStreamer
- Familiarity with model serving infrastructure (Triton, TorchServe, ONNX Runtime)
- Open‑source speech/audio contributions
- multilingual Indian speech systems
- voice cloning techniques
- media protocol integration
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
What You'll Likely Work On
- Build and optimise real‑time speech‑to‑speech translation pipelines, including streaming ASR, low‑latency translation and TTS synthesis
- Design fan‑out architectures that serve multiple concurrent listeners with personalised translated audio
- Implement voice‑cloning for streaming and batch contexts, handling speaker identity and utterance length
- Reduce end‑to‑end latency across the ASR‑translation‑TTS chain via buffering, segmentation and inference optimisation (quantisation, batching, acceleration)
- Integrate ML pipelines with real‑time media infrastructure (WebRTC, RTMP, SRT) and manage fallbacks, retries and quality routing
- Create automated QC loops and evaluation harnesses (WER/CER, tempo, pronunciation) to monitor speech quality
- Develop and maintain audio data pipelines for segment extraction, filtering, deduplication and QA
- Debug production speech systems, addressing latency, audio artifacts, code‑mixed content, dialects and edge cases
Good Fit If You Have
- Comfortable working with ambiguous, evolving roadmaps
- Passion for AI‑first products with high impact
- Experience with multilingual or Indic speech systems is a plus
- Open‑source audio/speech contributions valued
- Collaborative mindset across research, engineering and product teams
Skills
- Python & Asyncio
- PyTorch
- Speech model integration (ASR/TTS)
- Real‑time streaming (WebSocket, low‑latency pipelines)
- Audio signal processing
- Model serving (Triton/TorchServe/ONNX)
- Media protocols (WebRTC, RTMP, SRT)
- Voice cloning & speaker adaptation
- Evaluation & quality metrics (WER/CER)