Ovii Job Board

Distributed LLM Inference Engineer

Anyscale

San Francisco, San Francisco • Hybrid - San Francisco, San Francisco • Full-Time

Posted 2026-05-27 USD 170,000 - USD 245,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Distributed LLM Inference Engineer builds and optimizes high‑throughput, low‑latency inference pipelines for large language models using Ray and related open‑source tools. The role collaborates closely with product teams and the open‑source community to deliver batch and online inference solutions at scale.

Role Snapshot

  • Build high‑throughput LLM inference pipelines
  • Optimize distributed ML workloads
  • Integrate Ray Data and LLM engines
  • Collaborate with product and open‑source communities
  • Deliver batch and online inference solutions
  • Contribute to open‑source projects

Nice-to-Have Signals

  • Distributed systems
  • Large‑scale ML inference
  • Ray
  • PyTorch
  • Solid understanding of ML inference challenges
  • vLLM
  • CUDA / GPU
  • TensorRT‑LLM
  • Contributions to deep learning frameworks
  • Contributions to deep learning compilers
  • large‑scale ML inference
  • distributed systems

Work Setup

  • Location: San Francisco, USA
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Remote eligibility

What You'll Likely Work On

  • Iterate rapidly with product teams to ship end‑to‑end batch and online inference solutions at high scale.
  • Integrate Ray Data and LLM engines, delivering cost‑effective large‑scale inference.
  • Work with open‑source projects such as vLLM and contribute improvements back to the community.
  • Stay current with state‑of‑the‑art research and apply best practices to inference performance.

Good Fit If You Have

  • Enjoy fast iteration and shipping production‑ready features.
  • Comfortable engaging with open‑source communities and contributors.
  • Passionate about cutting‑edge AI infrastructure and performance engineering.

Skills

  • Distributed systems
  • Large‑scale ML inference
  • Ray
  • vLLM
  • PyTorch
  • CUDA / GPU
  • TensorRT‑LLM
  • Performance optimization
  • Batch & online inference