
Distributed LLM Inference Engineer
Anyscale
Posted 2026-05-27
USD 170,000 - USD 245,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
The Distributed LLM Inference Engineer builds and optimizes high‑throughput, low‑latency inference pipelines for large language models using Ray and related open‑source tools. The role collaborates closely with product teams and the open‑source community to deliver batch and online inference solutions at scale.
Role Snapshot
- Build high‑throughput LLM inference pipelines
- Optimize distributed ML workloads
- Integrate Ray Data and LLM engines
- Collaborate with product and open‑source communities
- Deliver batch and online inference solutions
- Contribute to open‑source projects
Nice-to-Have Signals
- Distributed systems
- Large‑scale ML inference
- Ray
- PyTorch
- Solid understanding of ML inference challenges
- vLLM
- CUDA / GPU
- TensorRT‑LLM
- Contributions to deep learning frameworks
- Contributions to deep learning compilers
- large‑scale ML inference
- distributed systems
Work Setup
- Location: San Francisco, USA
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility
What You'll Likely Work On
- Iterate rapidly with product teams to ship end‑to‑end batch and online inference solutions at high scale.
- Integrate Ray Data and LLM engines, delivering cost‑effective large‑scale inference.
- Work with open‑source projects such as vLLM and contribute improvements back to the community.
- Stay current with state‑of‑the‑art research and apply best practices to inference performance.
Good Fit If You Have
- Enjoy fast iteration and shipping production‑ready features.
- Comfortable engaging with open‑source communities and contributors.
- Passionate about cutting‑edge AI infrastructure and performance engineering.
Skills
- Distributed systems
- Large‑scale ML inference
- Ray
- vLLM
- PyTorch
- CUDA / GPU
- TensorRT‑LLM
- Performance optimization
- Batch & online inference