Ovii Job Board

Architect, On-Device Inference

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 8+ years

Posted 2026-06-04 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Architect, On‑Device Inference will own the end‑to‑end technical architecture for Sarvam’s edge AI products, driving latency, footprint and accuracy across diverse chipsets. This senior engineering role blends hands‑on implementation with player‑coach leadership, partnering directly with OEMs such as Qualcomm and Apple.

Role Snapshot

  • Own on‑device inference architecture
  • Define latency & footprint budgets
  • Lead runtime‑selection strategy
  • Drive model export pipelines
  • Technical liaison to OEM partners
  • Mentor and guide senior engineers

Must-Have Requirements

  • ML systems engineering
  • On‑device inference experience
  • TensorRT/CUDA
  • OpenVINO
  • QNN/SNPE
  • CoreML/ANE
  • ONNX Runtime EP development
  • Streaming inference (KV cache, chunked attention)
  • xPU/runtime selection system
  • ML systems
  • on‑device inference

Nice-to-Have Signals

  • Prior collaboration with Qualcomm, Intel, NVIDIA or Apple
  • OS‑layer integration (Windows IME, macOS Accessibility, Android InputMethodService)
  • ASR‑specific optimization (Whisper, Conformer, RNN‑T)
  • Indic‑language ML systems
  • OEM collaboration
  • OS‑layer integration
  • ASR optimization
  • Indic‑language ML

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Set and enforce latency, memory and footprint budgets across NPUs, CPUs and GPUs
  • Design and maintain a chipset‑specific runtime selection matrix (e.g., OpenVINO vs ONNX Runtime, TensorRT vs CUDA)
  • Build and own the model export pipeline from PyTorch to each target runtime
  • Implement xPU selector logic with capability probing and graceful fallback paths
  • Serve as the primary technical contact for OEM partners (Qualcomm, Intel, NVIDIA, AMD, Apple)
  • Lead the optimization team: hiring bar, roadmap, architecture reviews and mentorship of senior engineers
  • Ensure products meet published latency, footprint and accuracy targets on all supported hardware

Good Fit If You Have

  • Proven cross‑team technical leadership without becoming a bottleneck
  • Experience influencing OEM roadmaps or driver releases
  • Passion for building high‑impact AI products for the Indian market

Skills

  • ML systems engineering
  • On‑device inference
  • TensorRT / CUDA
  • OpenVINO
  • QNN / SNPE
  • CoreML / ANE
  • ONNX Runtime EP development
  • Streaming inference (KV cache, chunked attention)
  • xPU/runtime selection systems
  • Technical leadership & architecture