
Senior Software Engineer, Data Engineering
Valgenesis
Posted 2026-04-24
Tech & Engg
Job Description
Ovii's Interpretation of the Role
ValGenesis seeks a senior data engineer to design, build, and operate scalable data ingestion and transformation pipelines on Azure, delivering ML‑ready datasets and analytics for life‑science customers.
Role Snapshot
- Design and develop batch & real‑time data pipelines
- Build and optimize Azure Lakehouse architectures
- Implement data quality, lineage, and governance
- Enable self‑service analytics with BI tools
- Deploy data APIs and ML models in production
- Ensure performance, scalability, and observability
- Collaborate across engineering, product, and business teams
Must-Have Requirements
- Python
- SQL
- Compiled language (C#/Java/Scala)
- Azure Data Factory
- Databricks
- Azure Synapse / Delta Lake
- Kafka / Event Hubs / Service Bus
- Power BI / Superset / Tableau
- Docker
- Kubernetes
- CI/CD (GitHub Actions / Azure DevOps)
- Relational databases (SQL Server, PostgreSQL, MySQL)
- NoSQL databases (MongoDB, Cosmos DB)
- Data Engineering
- Azure data platform
Nice-to-Have Signals
- Knowledge graph or semantic search solutions
- LLM‑based data retrieval (RAG) patterns
- Data mesh or data fabric concepts
- MLflow, Delta Live Tables, or Databricks Unity Catalog
- Familiarity with TensorFlow or PyTorch
- Exposure to MLOps concepts
- knowledge graph
- semantic search
- LLM retrieval
Work Setup
- Location: Chennai, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Create and maintain end‑to‑end data ingestion, transformation, and orchestration pipelines
- Architect and tune Azure Lakehouse solutions using Synapse, Delta Lake, and Databricks
- Integrate diverse data sources—including SQL, NoSQL, files, and IoT streams—into unified storage
- Partner with data scientists to deliver ML‑ready datasets and operationalize models via Azure ML and Kubernetes
- Establish data quality, lineage, and governance standards across all pipelines
- Build self‑service analytics layers with Power BI, Superset, or Tableau for business users
- Implement monitoring, automation, and CI/CD for reliable data workflows
Good Fit If You Have
- Experience with knowledge graphs or semantic search is a plus
- Familiarity with LLM‑based retrieval (RAG) patterns is advantageous
- Exposure to data mesh, data fabric, or domain‑oriented architectures is beneficial
Skills
- Python & SQL programming
- Compiled language (C#/Java/Scala)
- Azure Data Platform (Data Lake, Synapse, Databricks)
- ETL/ELT tools (Data Factory, Airflow, dbt)
- Streaming (Kafka, Event Hubs, Service Bus)
- Visualization (Power BI, Superset, Tableau)
- Containerization & orchestration (Docker, Kubernetes)
- CI/CD (GitHub Actions, Azure DevOps)
- Relational & NoSQL databases