
Tech Lead Manager, Data Infrastructure
Cartesia
Posted 2026-07-09
USD 250,000 - USD 375,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
We need a Tech Lead Manager to own Cartesia's multimodal data strategy and build the high‑throughput pipelines that feed our foundation models. You will lead a data‑engineering team, partner with research and inference groups, and ensure data quality at scale.
Role Snapshot
- Lead and grow a data‑engineering team
- Define multimodal data strategy for pre‑ and post‑training
- Design scalable pipelines for text, audio, and video
- Collaborate with research and inference engineers
- Set and enforce data‑quality standards
- Manage external data‑vendor relationships and budgets
Must-Have Requirements
- ML data infrastructure (training pipelines, dataset versioning, large‑scale loading)
- Audio‑centric multimodal data processing
- High‑throughput distributed data pipelines
- Modern software engineering practices
- Team leadership & engineering management
- ML data infrastructure
- multimodal audio data
- engineering team leadership
Nice-to-Have Signals
- Familiarity with building/evaluating datasets for generative models
- generative model dataset evaluation
Work Setup
- Location: San Francisco, United States
- Work mode: ONSITE
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: yes
Not Specified in JD
- Salary range
- Equity details
- Remote eligibility
- Education requirement
- Certifications
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Craft and execute a multi‑modal data roadmap covering human, synthetic, and web‑scale sources
- Guide a team of data engineers in building and operating ingestion, preprocessing, augmentation, versioning, and loading pipelines
- Co‑design data systems alongside model training and serving infrastructure to ensure GPU‑aware loading and efficient evaluation
- Implement rigorous data‑quality metrics and feedback loops that directly influence model behavior
- Source novel datasets, negotiate with external vendors, and oversee related budgeting
- Partner with research scientists to align data pipelines with experimental goals
Good Fit If You Have
- Proven track record leading high‑impact engineering teams in research‑driven settings
- Hands‑on experience building ML‑focused data pipelines at scale
- Deep familiarity with audio data formats, preprocessing, and streaming patterns
Skills
- ML data infrastructure (training pipelines, dataset versioning, large‑scale loading)
- Audio‑centric multimodal data processing
- High‑throughput distributed data pipelines
- Modern software engineering practices (clean code, testing, tool selection)
- Team leadership & engineering management
- Cloud storage & streaming for large datasets