
Job Description
Ovii's Interpretation of the Role
WebEngage seeks a Data Engineer to design, build, and maintain production ETL pipelines and data models that power analytics, ML, and real‑time personalization. The role blends Python/SQL development, data‑warehouse architecture, and stakeholder collaboration in a fast‑moving product environment.
Role Snapshot
- Own end‑to‑end ETL/ELT pipelines
- Design dimensional data models
- Ensure data quality and observability
- Develop Python and SQL transformation code
- Create BI dashboards with Streamlit
- Collaborate with product, analytics, and data science teams
Must-Have Requirements
- Strong SQL (complex queries, performance optimization, cost‑efficient design on BigQuery/Redshift)
- Strong Python scripting for data ingestion and transformation
- End‑to‑end ETL/ELT pipeline ownership
- Dimensional and transactional data modeling
- Data engineering
- ETL pipeline development
- Data modeling
- Bachelor’s degree in Computer Science, Engineering, Mathematics, Statistics, or related quantitative field
Nice-to-Have Signals
- Airflow
- dbt
- Docker
- CI/CD (GitHub Actions / GitLab CI)
- GCP/AWS
- Streamlit and visualization libraries
- Git version‑control workflows
- Agile development practices
- Airflow orchestration
- dbt transformations
- Docker containerization
Work Setup
- Location: Mumbai, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, build, and maintain production‑grade ETL/ELT pipelines ingesting APIs, databases, event streams (Kafka/Pub‑Sub) and flat files into BigQuery or Redshift
- Implement idempotent, incremental loads with retry logic, dead‑letter queues and SLA‑based alerting
- Set up pipeline observability – data freshness checks, row‑count validation, schema‑drift detection and anomaly alerts using Great Expectations or dbt tests
- Translate business requirements into clean dimensional models, SCDs, bridge and fact tables, and manage partitioning, clustering and materialised views for cost‑effective performance
- Write modular, well‑tested Python and SQL code, follow DRY principles, use Git for version control and participate in peer reviews
- Build reusable transformation frameworks with dbt (or equivalent) and containerise services with Docker for CI/CD deployment
- Develop interactive dashboards and analytical tools in Streamlit, and create semantic BI layers for self‑service reporting
- Partner with product managers, analysts and data scientists to capture data needs, document lineage and SLAs, and share knowledge across engineering guilds
Good Fit If You Have
- Familiarity with Airflow, dbt, Docker or CI/CD pipelines is advantageous
- Experience with GCP or AWS cloud environments is a plus
- Comfort presenting data insights to both technical and non‑technical audiences
Skills
- SQL (BigQuery, Redshift)
- Python (Pandas, SQLAlchemy)
- ETL/ELT pipeline development
- Dimensional data modeling (star/snowflake, SCD)
- dbt or equivalent testing framework
- Airflow (preferred)
- Docker (preferred)
- CI/CD (GitHub Actions / GitLab CI, preferred)
- GCP/AWS cloud services (preferred)
- Streamlit & data visualization (preferred)
- Git version control (optional)
- Agile development practices (optional)