
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
Anyscale
Posted 2026-06-17
USD 215,000 - USD 275,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
We are seeking a Senior Site Reliability Engineer to own and evolve the platform infrastructure that powers Anyscale's distributed AI cloud. You will design, build, and operate highly available services that orchestrate Ray clusters, ensuring performance, reliability, and observability for our customers.
Role Snapshot
- Senior Site Reliability Engineer
- Platform Infrastructure
- Hybrid – San Francisco
- Go & Python
- Kubernetes & Cloud
- Distributed AI workloads
- On‑call support
Must-Have Requirements
- Go
- Python
- Kubernetes
- AWS
- Azure
- GCP
- Distributed systems
- Linux kernel / containers
- Networking
- Security
- production code
- distributed systems
- cloud‑native infrastructure
- Bachelor's degree in Computer Science, Engineering, or equivalent
Nice-to-Have Signals
- Observability stacks (Prometheus, Grafana)
- observability stacks
Work Setup
- Location: San Francisco, United States
- Work mode: HYBRID
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Relocation
- Travel
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, build, and scale services that orchestrate Ray clusters across cloud and on‑prem environments
- Optimize control‑plane components for large‑scale AI/ML workloads
- Create intelligent scheduling and resource‑management systems for heterogeneous compute clusters
- Develop features that improve reliability, performance, scalability, and observability of Ray workloads
- Integrate and optimize accelerators such as GPUs and TPUs
- Manage container images and dependency resolution for distributed jobs
- Provide on‑call support and troubleshoot infrastructure issues with customers and field teams
Good Fit If You Have
- Strong communication skills for customer and field‑team interaction
- Passion for open‑source and distributed systems
- Experience collaborating with ML and systems experts
Skills
- Go
- Python
- Kubernetes
- AWS / Azure / GCP
- Distributed systems
- Observability (Prometheus, Grafana)
- Linux kernel & containers
- Networking & security
- Resource scheduling