Senior Performance Engineer, Intel Stack
Sarvam
Posted 2026-06-04
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Senior Performance Engineer driving end‑to‑end deployment and optimization of AI models on Intel NPU, GPU and CPU stacks. Owns OpenVINO build, quantization and CI pipelines to meet defined SLAs for edge AI workloads.
Role Snapshot
- Senior Performance Engineer
- Intel inference stack ownership
- OpenVINO model deployment
- CPU & GPU performance tuning
- Edge AI integration
- CI regression management
Must-Have Requirements
- ML deployment (5+ years)
- Intel inference stack experience (2+ years)
- Production OpenVINO experience (model conversion, accuracy validation, driver‑version pinning)
- ONNX Runtime EP knowledge
- x86 CPU profiling and optimization (VTune, perf, assembly‑level analysis)
- ML deployment
- Intel inference stack
Nice-to-Have Signals
- AVX‑512 or AMX intrinsics (strong plus)
- Direct prior interaction with OpenVINO team
- Custom OpenVINO operator authoring
- AVX‑512 or AMX intrinsics
- OpenVINO ecosystem interaction
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Deploy edge AI models on Intel NPU, dGPU and iGPU within defined SLAs
- Maintain the OpenVINO build and quantization recipe, handling driver‑version compatibility
- Optimize x86 CPU fallback paths using AVX‑512, AMX and threading strategies
- Manage the Intel device CI pool and detect regressions across OpenVINO upgrades
- Collaborate with Intel’s ecosystem and the OpenVINO team to ensure seamless integration
Good Fit If You Have
- 2+ years experience with Intel inference stacks
- Familiarity with AVX‑512 or AMX intrinsics is a strong plus
- Prior direct interaction with the OpenVINO team or Intel partners
Skills
- OpenVINO
- ONNX Runtime
- Intel NPU / GPU
- x86 CPU optimization (AVX‑512, AMX)
- Model quantization & driver compatibility
- Performance profiling (VTune, perf)
- CI/CD for inference pipelines
- Machine Learning deployment