Ovii Job Board

Machine Learning Engineer, Vision

Sarvam

Bengaluru, India • Onsite - Bengaluru, India • Full-Time

Posted 2026-04-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Machine Learning Engineer, Vision will design, train, and ship large vision‑language models on GPU clusters. You’ll build multimodal data pipelines, create evaluation harnesses, and deliver production‑grade solutions for client use‑cases such as document processing and visual search.

Role Snapshot

  • Full‑cycle vision‑language model development
  • GPU‑accelerated training & fine‑tuning
  • Multimodal data pipeline engineering
  • Client‑facing solution delivery
  • Production‑grade inference optimisation

Must-Have Requirements

  • Python
  • PyTorch
  • experience training or fine‑tuning large models
  • experience building data pipelines at scale
  • solid grounding in transformer architectures
  • hands‑on training/fine‑tuning of large models
  • building data pipelines at scale
  • Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent)

Nice-to-Have Signals

  • vision‑language or multimodal model experience
  • distributed training frameworks (FSDP, DeepSpeed, Megatron‑LM)
  • post‑training methods (RLHF, DPO, alignment)
  • inference optimisation (quantisation, distillation, serving)
  • open‑source contributions or strong GitHub portfolio
  • comfort with ambiguous roadmaps
  • secure coding practices
  • vision‑language or multimodal systems
  • distributed training
  • post‑training alignment techniques
  • AI
  • vision‑language
  • multimodal

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Coding test

What You'll Likely Work On

  • Design and run training/fine‑tuning pipelines for large vision‑language models
  • Build and maintain multimodal data pipelines (ingestion, filtering, synthetic generation, QA)
  • Implement research‑driven model architectures and training techniques
  • Create evaluation harnesses, benchmarks, and automated regression tracking
  • Optimise models for inference through quantisation, batching, and serving infrastructure
  • Develop robust integrations that expose vision model capabilities to end users
  • Translate client problems into scoped ML tasks with appropriate data and evaluation
  • Own end‑to‑end delivery for client use‑cases such as document processing and visual search

Good Fit If You Have

  • Comfort working with ambiguous roadmaps and evolving research directions
  • Strong focus on code quality, security, and system reliability
  • Interest in collaborating directly with enterprise clients

Skills

  • Python
  • PyTorch
  • Transformer architectures
  • Large‑scale data pipelines
  • GPU cluster training
  • Model quantisation & batching
  • Secure coding & system reliability