Ovii Job Board

Research Engineer, Post-Training

Harvey

San Francisco, USA • Hybrid - San Francisco, USA • Full-Time

Posted 2026-06-26 USD 231,000 - USD 340,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

We need a Research Engineer to own post‑training experiments that improve our AI agents for legal work. You will design training recipes, grading systems, and evaluation pipelines while collaborating with internal and external researchers.

Role Snapshot

  • Research engineer for post‑training AI models
  • Hybrid role based in San Francisco
  • Full‑time individual contributor
  • Focus on LLM fine‑tuning and RLHF
  • Python‑centric research engineering

Must-Have Requirements

  • Model training techniques (SFT, RLHF/RLAIF, reward modeling, distillation)
  • Python programming
  • Research‑engineering ability
  • Model behavior analysis
  • Self‑management of applied research projects and clear communication
  • hands‑on model‑training experience
  • strong Python and research‑engineering skills
  • ability to self‑manage applied research projects

Nice-to-Have Signals

  • Data/evaluation infrastructure (dataset pipelines, experiment tracking, dashboards)
  • Distributed training, inference systems, GPU workloads
  • Research publications, open‑source contributions, shipped ML work
  • building data/evaluation infrastructure
  • distributed training and large‑scale ML experimentation
  • research publications or open‑source contributions

Work Setup

  • Location: San Francisco, United States
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Relocation
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Run and iterate post‑training experiments, balancing performance, cost, latency, security and governance
  • Optimize agent harnesses—skills, tools, sub‑agents, retrieval and validation loops—for long‑horizon legal tasks
  • Design grading and reward systems that are reliable, efficient and meet high‑stakes legal standards
  • Analyze agent traces to surface success patterns and turn insights into new training data or evals
  • Partner with internal researchers and external collaborators to define experiments, evaluate methods and drive model improvements

Good Fit If You Have

  • Enjoys ambiguous research problems and can self‑direct projects
  • Strong judgment about model outputs and failure modes
  • Clear communicator with engineers, product, domain experts and external partners

Skills

  • Model training (SFT, RLHF/RLAIF, reward modeling, distillation)
  • Python programming
  • Research‑engineering (experiment debugging, reliable code)
  • Model behavior analysis
  • Self‑managed applied research projects
  • Cross‑functional communication
  • Data/evaluation infrastructure (datasets, tracking, dashboards) – preferred
  • Distributed training & GPU workloads – preferred
  • Research publications / open‑source contributions – preferred