Job Description
Ovii's Interpretation of the Role
Sarvam seeks an Embedded Data Scientist to work on‑site with clients, turning heterogeneous, multimodal data into structured ontologies and knowledge‑graph representations that power the company’s sovereign AI platform.
Role Snapshot
- Embedded Data Scientist
- Client‑site data modeling
- Ontology & knowledge‑graph design
- Multimodal data pipelines
- AI reasoning enablement
- Python & ML expertise
Must-Have Requirements
- Python (pandas, NumPy)
- LLM/NLP tooling
- Experience with large unstructured datasets
- Design of data ontologies, schemas, and metadata frameworks
- data science
- applied machine learning
- large‑scale data analysis
Nice-to-Have Signals
- Knowledge‑graph or semantic data modeling
- Multimodal dataset experience
- Operating in constrained or air‑gapped environments
- Familiarity with vector search or retrieval pipelines
- knowledge‑graph development
- multimodal data handling
- constrained environment operations
Work Setup
- Location: Delhi, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Map and analyze client data sources across documents, images, audio, geospatial and structured records
- Create domain ontologies, entity hierarchies, and metadata frameworks that capture client‑specific concepts
- Define document segmentation and chunking strategies that preserve semantic meaning for retrieval
- Specify indexing, embedding, and linking approaches for multimodal datasets
- Partner with Strategic Deployment Engineers to build operational data ingestion pipelines
- Assess AI system retrieval and reasoning performance and iteratively refine semantic structures
- Collaborate with the models team to set benchmarks and evaluation criteria for real‑world deployments
- Translate data‑derived insights into structured signals for product and engineering teams
Good Fit If You Have
- Comfort operating autonomously in client environments without a dedicated data team
- Ability to bridge domain knowledge, data modeling, and AI system design
- Experience turning messy, unstructured data into rigorous, reusable schemas
Skills
- Python (pandas, NumPy)
- LLM/NLP tooling
- Large unstructured data handling
- Data ontology & schema design
- Metadata & entity modeling
- Vector search & retrieval concepts
- Multimodal data integration