Job Description
Ovii's Interpretation of the Role
As a Data Scientist on Sarvam’s Evaluation team, you’ll design and implement domain‑specific evaluation frameworks to assess LLM‑driven AI outputs, ensuring quality before and after deployment.
Role Snapshot
- Design domain‑specific evaluation frameworks
- Define quality metrics with experts and clients
- Run pre‑ and post‑deployment evaluation cycles
- Automate evaluation pipelines with MLOps
- Create and curate domain datasets and annotation workflows
- Publish findings to guide product and engineering roadmap
Must-Have Requirements
- Python
- pandas
- NumPy
- statistics
- evaluation framework design
- LLM evaluation tooling
- prompt engineering
- unstructured data handling
- data science
- ML research
- applied AI
- LLM production
Nice-to-Have Signals
- high‑stakes domain experience
- red‑teaming
- adversarial evaluation
- high‑stakes domains
- healthcare
- legal
- defence
- finance
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Build and iterate evaluation harnesses for tasks such as document comprehension, command summarisation, and geospatial reasoning
- Translate operational requirements into measurable quality metrics
- Automate evaluation pipelines, integrating them with deployment events
- Detect failure modes, edge cases, and distribution shifts in production
- Manage domain‑specific datasets and annotation processes
- Produce internal quality reports that influence product and engineering roadmaps
Good Fit If You Have
- Proven experience designing custom evaluation metrics and red‑teaming AI models
- Background in high‑impact domains where model errors have real consequences
- Ability to collaborate across product, engineering, and domain‑expert teams
Skills
- Python (pandas, NumPy)
- Statistical analysis & probability
- Evaluation framework design
- LLM evaluation tooling (RAGAS, Eval Harness, LangSmith)
- Prompt engineering & model fine‑tuning
- Unstructured data handling (PDFs, transcripts)
- Dashboard & reporting
- Collaboration & stakeholder communication
- High‑stakes domain knowledge (healthcare, finance, legal, defence)