
Job Description
Ovii's Interpretation of the Role
We need a Backend/Infrastructure Engineer to build the execution layer for large‑scale RL environments. The role blends systems engineering, sandboxing, and cloud‑scale performance to deliver reliable, cost‑effective platforms for research and enterprise customers.
Role Snapshot
- Design sandboxing & execution layer
- Own performance, latency, and cost at scale
- Build snapshot/restore tooling
- Collaborate with research and enterprise teams
- Drive observability and reliability
Must-Have Requirements
- Production systems development
- Distributed systems
- Container & sandboxing (namespaces, cgroups, VMs)
- Python programming
- Cloud platforms (GCP, AWS)
- building production systems at scale
- distributed systems design
- container and sandboxing expertise
Nice-to-Have Signals
- Rust/Go/C++ programming
- RL training/evaluation infrastructure experience
- Checkpoint/snapshot‑restore systems (CRIU)
- High‑throughput, low‑latency execution background
- Open‑source contributions
- RL training or evaluation infrastructure
- checkpoint/snapshot‑restore systems
Work Setup
- Location: Mountain View, United States
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design and own sandboxing and execution infrastructure for long‑running RL environments
- Implement snapshot and restore capabilities for disk, process, memory, and accelerator state
- Detect and remediate failure modes early, enabling rollback and continued execution
- Optimize throughput, latency, and cost‑per‑rollout across distributed, multi‑node deployments
- Profile and eliminate bottlenecks from container startup to environment teardown
- Build observability tools to monitor thousands of concurrent rollouts
- Create frameworks for packaging, deploying, and debugging RL environments
- Document and ship reproducible workflows for internal teams and external customers
Good Fit If You Have
- Enjoys operating in ambiguous, research‑driven settings
- Strong communicator who can bridge research needs and production requirements
- Passionate about building reliable, high‑performance infrastructure
Skills
- Production systems development
- Distributed systems
- Container & sandboxing (namespaces, cgroups, VMs)
- Python programming
- Cloud platforms (GCP, AWS)
- Rust/Go/C++ (optional)
- RL training/evaluation infrastructure (optional)
- Checkpoint‑restore systems (CRIU) (optional)
- High‑throughput, low‑latency execution (optional)
- Open‑source contributions (optional)