
Job Description
Ovii's Interpretation of the Role
We are looking for a hands‑on Data Scientist to drive an enterprise‑wide data transformation effort. The role blends Generative AI, traditional machine learning, and data engineering to deliver production‑ready data quality and metadata solutions.
Role Snapshot
- Build AI‑powered data quality solutions
- Develop LLM‑based metadata enrichment tools
- Create scalable Python & SQL pipelines
- Design human‑in‑the‑loop validation workflows
- Collaborate with business, governance, and engineering teams
- Deploy production‑grade ML models at scale
Must-Have Requirements
- LLMs & prompt engineering
- RAG architectures
- Python
- SQL
- ML frameworks (Scikit‑learn, LangChain, LlamaIndex)
- Supervised & unsupervised ML techniques (XGBoost, Random Forest, clustering, PCA, Isolation Forest, anomaly detection)
- Data governance & metadata management knowledge
- Strong analytical and communication skills
- AI/ML solution delivery in enterprise
- data quality and governance
- Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, Statistics, Mathematics, or related field
Work Setup
- Location: Gurgaon, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Travel requirement
- Security clearance
- Coding test
What You'll Likely Work On
- Create LLM‑driven pipelines for metadata enrichment, semantic classification, and summarization
- Build and optimise RAG architectures that leverage enterprise metadata and rule libraries
- Develop ML models for anomaly detection, entity resolution, clustering, and predictive analytics
- Engineer scalable Python and SQL pipelines for profiling, data onboarding, and quality monitoring
- Design Human‑in‑the‑Loop workflows to validate and audit AI outputs
- Partner with business, governance, and technical stakeholders to operationalise AI/ML solutions in an enterprise setting
Good Fit If You Have
- Proven track record delivering AI/ML solutions in large‑scale enterprise environments
- Strong analytical, problem‑solving and communication abilities
- Comfort translating advanced research into production‑ready implementations
- Familiarity with data quality, governance and metadata concepts
Skills
- Large Language Models (LLMs)
- Prompt engineering
- Retrieval Augmented Generation (RAG)
- Python
- SQL
- Scikit‑learn / LangChain / LlamaIndex
- XGBoost / Random Forest / Clustering
- Anomaly detection
- Data governance & metadata management
- Responsible AI