
Job Description
Ovii's Interpretation of the Role
Lead Data Engineer driving AI‑centric data pipelines, retrieval systems, and ML/LLM ops on cloud platforms. Own end‑to‑end data infrastructure that powers conversational agents, dashboards, and predictive models for enterprise AI products.
Role Snapshot
- Lead data engineering initiatives
- Design and operate production‑grade AI pipelines
- Build retrieval and vector‑store infrastructure
- Ensure data governance, security, and compliance
- Drive ML/LLM ops foundations
- Mentor engineering best practices
Must-Have Requirements
- Python
- SQL
- PySpark
- Kafka
- Snowflake/DataBricks
- Delta Lake
- AWS (S3, Glue, Kinesis, EKS, Redshift)
- Docker
- Kubernetes
- GitHub Actions
- data engineering
- cloud services
- AI/ML data infrastructure
Nice-to-Have Signals
- LangChain
- LlamaIndex
- LLM APIs (OpenAI, Bedrock, Claude, HuggingFace)
- Pinecone
- FAISS
- ChromaDB
- OpenSearch
- MLflow
- FastAPI
- Neo4j
- prompt engineering
- RLHF dataset preparation
- LLM fine‑tuning workflows
- vector store retrieval
- LLM fine‑tuning
Work Setup
- Location: India, India
- Work mode: REMOTE
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility details
- Travel requirements
What You'll Likely Work On
- Build, test, and maintain batch and real‑time data pipelines on Snowflake, PySpark, Delta Lake, and Kafka.
- Create end‑to‑end retrieval infrastructure, including document ingestion, embedding pipelines, and vector store management.
- Implement semantic layers, feature stores, and knowledge graphs to serve both ML models and LLM applications.
- Develop ML/LLM data pipelines, experiment tracking, dataset versioning, and automated evaluation for model quality.
- Provide data APIs, tool schemas, and agent observability for autonomous agents and conversational AI consumers.
- Enforce data governance, RBAC, PII masking, and audit logging to meet compliance and security standards.
Good Fit If You Have
- Enjoy building production‑grade pipelines at scale (batch & streaming).
- Comfort with cloud‑native data architectures on AWS/Azure.
- Interest in AI/LLM‑driven products and retrieval systems.
- Experience with data governance, quality frameworks, and compliance.
- Collaborative remote‑first mindset.
Skills
- Python
- SQL
- PySpark
- Kafka
- Snowflake / Databricks
- Delta Lake
- AWS / Azure cloud services
- Docker & Kubernetes
- GitHub Actions / CI‑CD
- MLflow
- Vector stores (Pinecone, FAISS, ChromaDB, OpenSearch)
- Knowledge graphs (Neo4j)