Ovii Job Board

4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer

Noida, India • Onsite - Noida, India • Full-Time • 5+ years

Posted 2026-08-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior Data Engineer on Innovaccer's Lakehouse team, building and operating end‑to‑end data pipelines that ingest raw healthcare files with Spark and serve analytics via Trino on Apache Iceberg tables. The role combines hands‑on development, performance tuning, and production reliability for a large‑scale on‑premise platform.

Role Snapshot

  • Senior Data Engineer – Lakehouse team
  • Build & operate Spark ingestion pipelines
  • Develop Trino SQL transform workflows
  • Maintain Apache Iceberg tables
  • Optimize query and pipeline performance
  • Collaborate across data platform stakeholders

Must-Have Requirements

  • Apache Spark
  • SQL
  • Python or Java
  • Airflow or equivalent
  • S3‑compatible storage
  • Parquet
  • data engineering
  • production pipelines at scale
  • B.E., B.Tech., or M.Sc. in Computer Science or related technical field

Nice-to-Have Signals

  • Apache Iceberg
  • Trino/Presto
  • Healthcare data formats (HL7, CCDA, claims)
  • Delta Lake or Hudi
  • healthcare data formats
  • regulated environment

Work Setup

  • Location: Noida, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Travel requirement
  • Security clearance
  • Coding test

What You'll Likely Work On

  • Build Spark ingestion jobs that land high‑volume raw files into Iceberg tables with schema handling and idempotent replay.
  • Develop Trino SQL pipelines for validation, business‑rule transforms, MERGE‑based deduplication, and aggregations.
  • Port existing warehouse SQL workloads to Trino and Spark SQL dialects, verifying output correctness.
  • Automate Iceberg table maintenance tasks such as compaction, snapshot expiry, and orphan‑file cleanup.
  • Tune performance through partitioning strategy, file sizing, statistics collection, and resource‑group configuration.
  • Instrument pipelines with data‑quality checks, reconciliation reports, alerting, and support rollout validation.

Good Fit If You Have

  • Experience with healthcare data formats (HL7, CCDA, claims) is a plus.
  • Familiarity with regulated‑environment data handling.
  • Strong problem‑solving skills for performance and reliability tuning.

Skills

  • Apache Spark
  • SQL
  • Python / Java
  • Trino (Presto) or comparable distributed SQL engine
  • Apache Iceberg (preferred)
  • Airflow or equivalent orchestration
  • S3‑compatible object storage
  • Parquet columnar format
  • Healthcare data formats (HL7, CCDA, claims)