
Software Engineer, Infrastructure
Anyscale
Posted 2026-07-10
USD 200,000 - USD 240,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Software Engineer on Anyscale's Infrastructure team building the control‑plane and data‑plane services that power scalable AI workloads. The role blends cloud‑native engineering, distributed systems design, and on‑call support for a hybrid San Francisco environment.
Role Snapshot
- Build and scale control‑plane services for Ray clusters
- Design cloud‑native infrastructure on AWS, Azure, GCP
- Develop Kubernetes orchestration and scheduling
- Implement observability, reliability, and performance features
- Provide on‑call support and troubleshoot infrastructure issues
- Collaborate with ML and distributed‑systems experts
Must-Have Requirements
- Go
- Python
- Kubernetes
- AWS/Azure/GCP
- Networking & security
- Observability (Prometheus, Grafana)
- Linux kernel / container fundamentals
- Accelerator integration (GPUs, TPUs)
- production‑grade code
- highly available distributed systems
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Work Setup
- Location: San Francisco, USA
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
What You'll Likely Work On
- Design, build, and scale services that orchestrate Ray clusters across cloud and on‑prem environments
- Optimize control‑plane components for large‑scale AI/ML workloads
- Create intelligent scheduling and resource‑management systems for heterogeneous compute clusters
- Enhance reliability, performance, scalability, and observability of managed Ray workloads
- Integrate and optimize accelerators such as GPUs and TPUs
- Manage container images and dependency resolution for distributed workloads
- Participate in code reviews and architecture discussions
- Provide on‑call support and work closely with customer and field teams
Good Fit If You Have
- Familiarity with Linux kernel internals and file‑system concepts
- Experience integrating hardware accelerators (GPUs, TPUs)
- Comfort with open‑source contributions and community collaboration
- Interest in AI/ML infrastructure and distributed computing
Skills
- Go
- Python
- Kubernetes
- Public cloud platforms (AWS, Azure, GCP)
- Distributed systems design
- Observability (Prometheus, Grafana)
- Networking & security
- Container technologies