Ovii Job Board

Software Engineer, Infrastructure

Anyscale

San Francisco, San Francisco • Hybrid - San Francisco, San Francisco • Full-Time • 3+ years

Posted 2026-07-10 USD 200,000 - USD 240,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Software Engineer on Anyscale's Infrastructure team building the control‑plane and data‑plane services that power scalable AI workloads. The role blends cloud‑native engineering, distributed systems design, and on‑call support for a hybrid San Francisco environment.

Role Snapshot

  • Build and scale control‑plane services for Ray clusters
  • Design cloud‑native infrastructure on AWS, Azure, GCP
  • Develop Kubernetes orchestration and scheduling
  • Implement observability, reliability, and performance features
  • Provide on‑call support and troubleshoot infrastructure issues
  • Collaborate with ML and distributed‑systems experts

Must-Have Requirements

  • Go
  • Python
  • Kubernetes
  • AWS/Azure/GCP
  • Networking & security
  • Observability (Prometheus, Grafana)
  • Linux kernel / container fundamentals
  • Accelerator integration (GPUs, TPUs)
  • production‑grade code
  • highly available distributed systems
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience

Work Setup

  • Location: San Francisco, USA
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility

What You'll Likely Work On

  • Design, build, and scale services that orchestrate Ray clusters across cloud and on‑prem environments
  • Optimize control‑plane components for large‑scale AI/ML workloads
  • Create intelligent scheduling and resource‑management systems for heterogeneous compute clusters
  • Enhance reliability, performance, scalability, and observability of managed Ray workloads
  • Integrate and optimize accelerators such as GPUs and TPUs
  • Manage container images and dependency resolution for distributed workloads
  • Participate in code reviews and architecture discussions
  • Provide on‑call support and work closely with customer and field teams

Good Fit If You Have

  • Familiarity with Linux kernel internals and file‑system concepts
  • Experience integrating hardware accelerators (GPUs, TPUs)
  • Comfort with open‑source contributions and community collaboration
  • Interest in AI/ML infrastructure and distributed computing

Skills

  • Go
  • Python
  • Kubernetes
  • Public cloud platforms (AWS, Azure, GCP)
  • Distributed systems design
  • Observability (Prometheus, Grafana)
  • Networking & security
  • Container technologies