Job Description
Ovii's Interpretation of the Role
Sarvam seeks an ML Engineer (Data) to own petabyte‑scale data infrastructure for foundational models, building pipelines, quality filters, and mixture design that directly impact model training.
Role Snapshot
- Own petabyte‑scale data pipelines
- Design data mixture and curriculum
- Build model‑based quality filtering
- Create tooling for data analysis & debugging
- Ensure data provenance, licensing, and multilingual support
- Partner with research and training‑infra teams
Must-Have Requirements
- Python
- Distributed data processing frameworks (Spark, Ray, Beam, Dask or equivalent)
- Large‑scale data pipeline engineering
- Data curation & filtering for LLM training
- Data mixture design & curriculum planning
- Open‑source contributions to data tooling
- petabyte‑scale data pipelines
- distributed processing frameworks
- LLM data curation and filtering
- BS or MS in Computer Science or a closely related technical field
Nice-to-Have Signals
- Multilingual data collection and normalization
- Model‑based data quality classifiers / contamination detection
- Tokenization research and practical implications
- First‑author research papers on data curation
- multilingual data handling
- open‑source tooling contributions
- research publications on data quality
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Design and implement petabyte‑scale ingestion, parsing, normalization, filtering, deduplication, tokenization, and packing pipelines.
- Develop and iterate on model‑based quality classifiers and contamination detection systems.
- Define and maintain data mixture, curriculum, and annealing strategies in close collaboration with researchers.
- Build internal tools that let engineers slice, attribute, and debug data used for pre‑training.
- Scale pipelines to handle multilingual corpora, code, math, and licensed datasets while tracking provenance and licensing.
- Work with the training infrastructure team to keep data flow from becoming a training bottleneck.
Good Fit If You Have
- Significant open‑source contributions to data‑tooling ecosystems.
- Experience with multilingual data collection, normalization, and quality scoring.
- Hands‑on work with model‑based data quality classifiers or contamination detection.
- First‑author papers or technical reports on data curation or pre‑training mixtures.
Skills
- Distributed data processing (Spark, Ray, Beam, Dask)
- Python programming
- Large‑scale data pipeline engineering
- Data curation & filtering for LLM training
- Tokenization, sharding, and packing
- Open‑source contributions to data tooling