Ovii Job Board

Performance Engineer, On-Device Inference

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 3+ years

Posted 2026-06-04 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Performance Engineer will take research‑grade models to production‑ready, on‑device artifacts across multiple chipsets. Responsibilities include quantization, benchmarking, integration with app teams, and maintaining the performance benchmark suite.

Role Snapshot

  • ML performance engineering
  • On‑device inference optimization
  • Model quantization & validation
  • Cross‑chipset deployment
  • Benchmark harness maintenance

Must-Have Requirements

  • PyTorch
  • ONNX export
  • Model quantization in production
  • Comfort with at least two inference runtimes (ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, LiteRT)
  • Performance profiling on at least one platform
  • ML systems

Nice-to-Have Signals

  • Custom op authoring in any runtime

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility

What You'll Likely Work On

  • Quantize models, validate accuracy, benchmark performance, and produce deployment documentation.
  • Own end‑to‑end delivery of 1‑2 model‑chipset pairs, writing deployment workbooks for each.
  • Collaborate part‑time with consuming application teams to debug performance and accuracy issues.
  • Maintain and extend the internal benchmark harness for ongoing performance tracking.

Good Fit If You Have

  • Enjoys deep technical debugging alongside product teams.
  • Thrives in a fast‑moving, high‑ownership environment.
  • Has a strong interest in edge AI and on‑device inference.

Skills

  • PyTorch
  • ONNX export
  • Model quantization
  • Inference runtimes (ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, LiteRT)
  • Performance profiling