Ovii Job Board

Principal Software Engineer - Data Platform (Iceberg/Trino)

Innovaccer

United States • Remote - United States • Full-Time • 12+ years

Posted 2026-08-13 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Principal Software Engineer will own the on‑premise lakehouse architecture that powers Innovaccer’s healthcare data platform, designing and operating Apache Iceberg tables, Trino query clusters, and Spark‑based transforms. The role blends deep distributed‑SQL expertise with platform‑wide standards and mentorship of senior engineers.

Role Snapshot

  • Own lakehouse reference architecture (Iceberg, Trino, Spark, catalog service)
  • Design on‑premise equivalents for cloud‑warehouse capabilities
  • Define table layout, partitioning, file sizing, and maintenance standards
  • Lead SQL dialect strategy for Trino and Spark SQL
  • Mentor senior engineers across data workstreams
  • Partner with platform engineering on storage sizing and capacity planning

Must-Have Requirements

  • Distributed SQL engine expertise (Trino/Presto or Spark SQL)
  • Apache Iceberg deep knowledge
  • Java and/or Python programming
  • REST catalog service experience
  • S3‑compatible object storage familiarity
  • Large‑scale data platform design
  • Distributed systems engineering
  • B.E., B.Tech., or M.Sc. in Computer Science or related technical field
  • Must have 12+ years industry experience
  • Must hold a B.E., B.Tech., or M.Sc. in Computer Science or related field
  • Must be authorized to work in the United States (E‑Verify)

Nice-to-Have Signals

  • Experience with Delta Lake or Hudi
  • Healthcare data domain exposure
  • Regulated or air‑gapped environment experience
  • Healthcare data handling
  • Regulated/air‑gapped environments

Work Setup

  • Location: United States
  • Work mode: REMOTE
  • Remote scope: COUNTRY_RESTRICTED
  • Remote countries: United States
  • Employment type: Full-Time

Eligibility Gates

  • Work authorization: Must be authorized to work in the United States
  • Visa sponsorship: no

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Coding test

What You'll Likely Work On

  • Architect and evolve the on‑premise lakehouse engine, including Iceberg tables, Trino clusters, and Spark transform compute
  • Create on‑premise replacements for cloud‑warehouse features such as CDC streams, scheduled tasks, and write‑back paths
  • Validate catalog and query engine prototypes at production‑scale data volumes and set placement triggers (VM vs Kubernetes)
  • Establish platform‑wide standards for table layout, partitioning, file sizing, compaction, snapshot expiry, and orphan‑file cleanup
  • Define and drive the SQL dialect migration strategy from existing warehouses to Trino and Spark SQL
  • Mentor senior engineers, review designs, and raise engineering quality across data workstreams
  • Collaborate with platform engineering on storage sizing, isolation, and capacity planning for the lakehouse footprint

Good Fit If You Have

  • Hands‑on experience with cloud data warehouses (Snowflake, BigQuery, Redshift) is a plus
  • Familiarity with Delta Lake or Hudi adds value
  • Background in healthcare data or regulated environments is advantageous

Skills

  • Distributed SQL engines (Trino/Presto, Spark SQL)
  • Apache Iceberg table management
  • Java and/or Python development
  • REST catalog services (Polaris, Nessie, Hive Metastore)
  • S3‑compatible object storage
  • Performance engineering & query planning
  • Capacity planning & resource isolation
  • Healthcare data domain knowledge (preferred)
  • Regulated/air‑gapped environment experience (preferred)

Remote Eligibility

  • United States