Job Description
Ovii's Interpretation of the Role
The Machine Learning Engineer, Vision will design, train, and ship large vision‑language models on GPU clusters. You’ll build multimodal data pipelines, create evaluation harnesses, and deliver production‑grade solutions for client use‑cases such as document processing and visual search.
Role Snapshot
- Full‑cycle vision‑language model development
- GPU‑accelerated training & fine‑tuning
- Multimodal data pipeline engineering
- Client‑facing solution delivery
- Production‑grade inference optimisation
Must-Have Requirements
- Python
- PyTorch
- experience training or fine‑tuning large models
- experience building data pipelines at scale
- solid grounding in transformer architectures
- hands‑on training/fine‑tuning of large models
- building data pipelines at scale
- Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent)
Nice-to-Have Signals
- vision‑language or multimodal model experience
- distributed training frameworks (FSDP, DeepSpeed, Megatron‑LM)
- post‑training methods (RLHF, DPO, alignment)
- inference optimisation (quantisation, distillation, serving)
- open‑source contributions or strong GitHub portfolio
- comfort with ambiguous roadmaps
- secure coding practices
- vision‑language or multimodal systems
- distributed training
- post‑training alignment techniques
- AI
- vision‑language
- multimodal
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Coding test
What You'll Likely Work On
- Design and run training/fine‑tuning pipelines for large vision‑language models
- Build and maintain multimodal data pipelines (ingestion, filtering, synthetic generation, QA)
- Implement research‑driven model architectures and training techniques
- Create evaluation harnesses, benchmarks, and automated regression tracking
- Optimise models for inference through quantisation, batching, and serving infrastructure
- Develop robust integrations that expose vision model capabilities to end users
- Translate client problems into scoped ML tasks with appropriate data and evaluation
- Own end‑to‑end delivery for client use‑cases such as document processing and visual search
Good Fit If You Have
- Comfort working with ambiguous roadmaps and evolving research directions
- Strong focus on code quality, security, and system reliability
- Interest in collaborating directly with enterprise clients
Skills
- Python
- PyTorch
- Transformer architectures
- Large‑scale data pipelines
- GPU cluster training
- Model quantisation & batching
- Secure coding & system reliability