Job Description
Ovii's Interpretation of the Role
Researcher, Vision at Sarvam will drive end‑to‑end development of vision‑language models, from data curation to training, evaluation, and production. The role blends deep research with practical engineering to advance multilingual multimodal AI for India.
Role Snapshot
- Full‑cycle vision‑language model research
- Design and run experiments in PyTorch
- Create multilingual data and evaluation pipelines
- Collaborate with engineers to scale prototypes
- Publish and share findings with the research community
Must-Have Requirements
- Deep understanding of vision‑language models
- Strong PyTorch skills for end‑to‑end experiments
- Rigorous experimental design
- vision-language model research
- experimental design
Nice-to-Have Signals
- PhD or Master's in ML, Computer Vision, NLP
- Experience with multilingual or low‑resource language modeling
- Familiarity with document understanding, OCR, or structured visual prediction
- Experience with large‑scale data curation
- multilingual modeling
- large‑scale data curation
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility
- Education requirement
- Certifications
What You'll Likely Work On
- Explore and prototype new vision‑language architectures and fusion mechanisms
- Develop training strategies (pre‑training, SFT, RLHF, DPO) for multilingual VLMs
- Design data mixtures, quality signals, and synthetic data pipelines to improve model performance
- Build benchmarks and evaluation tools for Indic multimodal tasks
- Analyze failure modes, robustness, and interpretability of models
- Partner with engineering teams to turn research ideas into scalable prototypes
- Contribute to open‑source projects and engage with the broader research community
Good Fit If You Have
- Track record of research impact via publications or shipped AI features
- Comfort with rapid iteration across data, training, and evaluation problems
- Interest in high‑ownership, high‑impact AI research for large‑scale populations
Skills
- Vision‑language model research
- PyTorch experimentation
- Multilingual modeling
- Data curation & synthetic data
- Evaluation framework design
- Robustness & interpretability analysis