Performance Engineer, On-Device Inference
Sarvam
Posted 2026-06-04
Tech & Engg
Job Description
Ovii's Interpretation of the Role
The Performance Engineer will take research‑grade models to production‑ready, on‑device artifacts across multiple chipsets. Responsibilities include quantization, benchmarking, integration with app teams, and maintaining the performance benchmark suite.
Role Snapshot
- ML performance engineering
- On‑device inference optimization
- Model quantization & validation
- Cross‑chipset deployment
- Benchmark harness maintenance
Must-Have Requirements
- PyTorch
- ONNX export
- Model quantization in production
- Comfort with at least two inference runtimes (ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, LiteRT)
- Performance profiling on at least one platform
- ML systems
Nice-to-Have Signals
- Custom op authoring in any runtime
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
What You'll Likely Work On
- Quantize models, validate accuracy, benchmark performance, and produce deployment documentation.
- Own end‑to‑end delivery of 1‑2 model‑chipset pairs, writing deployment workbooks for each.
- Collaborate part‑time with consuming application teams to debug performance and accuracy issues.
- Maintain and extend the internal benchmark harness for ongoing performance tracking.
Good Fit If You Have
- Enjoys deep technical debugging alongside product teams.
- Thrives in a fast‑moving, high‑ownership environment.
- Has a strong interest in edge AI and on‑device inference.
Skills
- PyTorch
- ONNX export
- Model quantization
- Inference runtimes (ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, LiteRT)
- Performance profiling