Ovii Job Board

Backend Engineer

Bespoke Labs

Mountain View, United States • Hybrid - Mountain View, United States • Full-Time

Posted 2026-05-27 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

We need a Backend/Infrastructure Engineer to build the execution layer for large‑scale RL environments. The role blends systems engineering, sandboxing, and cloud‑scale performance to deliver reliable, cost‑effective platforms for research and enterprise customers.

Role Snapshot

  • Design sandboxing & execution layer
  • Own performance, latency, and cost at scale
  • Build snapshot/restore tooling
  • Collaborate with research and enterprise teams
  • Drive observability and reliability

Must-Have Requirements

  • Production systems development
  • Distributed systems
  • Container & sandboxing (namespaces, cgroups, VMs)
  • Python programming
  • Cloud platforms (GCP, AWS)
  • building production systems at scale
  • distributed systems design
  • container and sandboxing expertise

Nice-to-Have Signals

  • Rust/Go/C++ programming
  • RL training/evaluation infrastructure experience
  • Checkpoint/snapshot‑restore systems (CRIU)
  • High‑throughput, low‑latency execution background
  • Open‑source contributions
  • RL training or evaluation infrastructure
  • checkpoint/snapshot‑restore systems

Work Setup

  • Location: Mountain View, United States
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design and own sandboxing and execution infrastructure for long‑running RL environments
  • Implement snapshot and restore capabilities for disk, process, memory, and accelerator state
  • Detect and remediate failure modes early, enabling rollback and continued execution
  • Optimize throughput, latency, and cost‑per‑rollout across distributed, multi‑node deployments
  • Profile and eliminate bottlenecks from container startup to environment teardown
  • Build observability tools to monitor thousands of concurrent rollouts
  • Create frameworks for packaging, deploying, and debugging RL environments
  • Document and ship reproducible workflows for internal teams and external customers

Good Fit If You Have

  • Enjoys operating in ambiguous, research‑driven settings
  • Strong communicator who can bridge research needs and production requirements
  • Passionate about building reliable, high‑performance infrastructure

Skills

  • Production systems development
  • Distributed systems
  • Container & sandboxing (namespaces, cgroups, VMs)
  • Python programming
  • Cloud platforms (GCP, AWS)
  • Rust/Go/C++ (optional)
  • RL training/evaluation infrastructure (optional)
  • Checkpoint‑restore systems (CRIU) (optional)
  • High‑throughput, low‑latency execution (optional)
  • Open‑source contributions (optional)