Ovii Job Board

Tech Lead Manager, Data Infrastructure

Cartesia

San Francisco, United States • Onsite - San Francisco, United States • Full-Time

Posted 2026-07-09 USD 250,000 - USD 375,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

We need a Tech Lead Manager to own Cartesia's multimodal data strategy and build the high‑throughput pipelines that feed our foundation models. You will lead a data‑engineering team, partner with research and inference groups, and ensure data quality at scale.

Role Snapshot

  • Lead and grow a data‑engineering team
  • Define multimodal data strategy for pre‑ and post‑training
  • Design scalable pipelines for text, audio, and video
  • Collaborate with research and inference engineers
  • Set and enforce data‑quality standards
  • Manage external data‑vendor relationships and budgets

Must-Have Requirements

  • ML data infrastructure (training pipelines, dataset versioning, large‑scale loading)
  • Audio‑centric multimodal data processing
  • High‑throughput distributed data pipelines
  • Modern software engineering practices
  • Team leadership & engineering management
  • ML data infrastructure
  • multimodal audio data
  • engineering team leadership

Nice-to-Have Signals

  • Familiarity with building/evaluating datasets for generative models
  • generative model dataset evaluation

Work Setup

  • Location: San Francisco, United States
  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: yes

Not Specified in JD

  • Salary range
  • Equity details
  • Remote eligibility
  • Education requirement
  • Certifications
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Craft and execute a multi‑modal data roadmap covering human, synthetic, and web‑scale sources
  • Guide a team of data engineers in building and operating ingestion, preprocessing, augmentation, versioning, and loading pipelines
  • Co‑design data systems alongside model training and serving infrastructure to ensure GPU‑aware loading and efficient evaluation
  • Implement rigorous data‑quality metrics and feedback loops that directly influence model behavior
  • Source novel datasets, negotiate with external vendors, and oversee related budgeting
  • Partner with research scientists to align data pipelines with experimental goals

Good Fit If You Have

  • Proven track record leading high‑impact engineering teams in research‑driven settings
  • Hands‑on experience building ML‑focused data pipelines at scale
  • Deep familiarity with audio data formats, preprocessing, and streaming patterns

Skills

  • ML data infrastructure (training pipelines, dataset versioning, large‑scale loading)
  • Audio‑centric multimodal data processing
  • High‑throughput distributed data pipelines
  • Modern software engineering practices (clean code, testing, tool selection)
  • Team leadership & engineering management
  • Cloud storage & streaming for large datasets