Ovii Job Board

Applied Researcher, Audio Post-Training

Cartesia

San Francisco, United States • Onsite - San Francisco, United States • Full-Time

Posted 2026-07-20 USD 200,000 - USD 350,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Applied Researcher, Audio Post‑Training will design, build, and evaluate capabilities for Cartesia’s generative audio models, turning customer‑driven problems into research plans and production‑ready features. The role blends deep ML research with end‑to‑end system ownership and close collaboration with product and customer teams.

Role Snapshot

  • Audio generative model research
  • End‑to‑end model pipeline ownership
  • Customer‑driven capability development
  • Cross‑functional collaboration
  • Evaluation design & data quality
  • Production failure analysis

Must-Have Requirements

  • Machine learning fundamentals
  • Software engineering fundamentals
  • Large multilingual dataset creation
  • Training and debugging generative models (speech, text, multimodal)
  • Supervised Fine‑Tuning (SFT)
  • Reinforcement Learning (RL)
  • Synthetic data generation
  • Human and automated evaluation
  • building large multilingual datasets
  • training and debugging generative audio models
  • designing and running SFT/RL experiments

Nice-to-Have Signals

  • Native proficiency in additional languages
  • native proficiency in other languages

Work Setup

  • Location: San Francisco, United States
  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: yes

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility

What You'll Likely Work On

  • Partner with product teams to translate customer requests into concrete research objectives.
  • Design and run experiments across data processing, synthetic data, SFT, RL, and evaluation pipelines.
  • Build and maintain high‑quality multilingual datasets for audio model training.
  • Debug and root‑cause failures in production‑grade generative models.
  • Determine readiness of new features and capabilities for public release.
  • Communicate research outcomes and model improvements to product and stakeholder audiences.

Good Fit If You Have

  • Enjoys turning ambiguous customer complaints into clear research plans.
  • Thrives in fast‑moving environments where speed and quality are both critical.
  • Bonus if fluent in additional languages beyond English.

Skills

  • Machine Learning fundamentals
  • Generative audio model training & debugging
  • Large multilingual dataset creation
  • Supervised Fine‑Tuning (SFT)
  • Reinforcement Learning (RL)
  • Synthetic data generation
  • Human & automated evaluation methods
  • Software engineering (Python/C++)