Ovii Job Board

Research Engineer (Reinforcement Learning)

LiveKit

NAMER • Remote - NAMER • Full-Time

Posted 2026-08-18 USD 135,000 - USD 300,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

LiveKit seeks a Research Engineer to design, build, and ship reinforcement‑learning models for voice and text agents. The role owns synthetic data pipelines, training experiments, and evaluation suites, delivering production‑ready AI behavior.

Role Snapshot

  • Research Engineer – RL focus
  • Remote, full‑time
  • Python‑centric development
  • End‑to‑end model lifecycle
  • Synthetic data pipeline ownership

Must-Have Requirements

  • Python
  • GPU computing
  • end‑to‑end model development
  • synthetic data pipeline ownership
  • training experiment execution

Nice-to-Have Signals

  • Reinforcement learning
  • Fine‑tuning frameworks (TRL, verl, OpenRLHF)
  • Fast rollout tools (vLLM, SGLang)
  • Multi‑GPU training (FSDP)
  • post‑training fine‑tuning
  • reinforcement‑learning research

Work Setup

  • Location: NAMER
  • Work mode: REMOTE
  • Remote scope: UNSPECIFIED
  • Remote countries: NAMER, APJ
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Remote eligibility specifics

What You'll Likely Work On

  • Design and implement training environments and verifiers
  • Own the end‑to‑end synthetic data pipeline and quality gates
  • Run full training experiments, analyze results, and iterate
  • Create release‑level evaluation suites for model validation
  • Select and adapt open‑weight base models for voice and text agents
  • Deploy trained models to production and monitor real‑world performance

Good Fit If You Have

  • Enjoys collaborating in a fully remote, senior‑engineer team
  • Values data as a product—coverage, diversity, and leakage awareness
  • Proactive about model reward exploitation and mitigation

Skills

  • Python
  • GPU computing
  • Synthetic data generation
  • Model lifecycle management
  • Reinforcement learning (preferred)
  • Fine‑tuning frameworks (TRL, verl, OpenRLHF) (preferred)
  • Fast rollout tools (vLLM, SGLang) (preferred)
  • Multi‑GPU training (FSDP) (preferred)

Remote Eligibility

  • NAMER
  • APJ