
AI Platform Engineer, Training and Inference
Saviynt
Posted 2026-05-18
USD 274,000 - USD 304,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
The AI Platform Engineer will design, build, and operate distributed training pipelines and LLM inference services for Saviynt's identity platform. The role spans Ray ecosystem management, model lifecycle automation, and performance optimization on GPU clusters.
Role Snapshot
- Own Ray ecosystem on GKE
- Operate distributed training pipelines
- Build LLM inference mesh
- Design model routing and promotion lifecycle
- Develop RL training infrastructure
- Optimize GPU utilization and inference performance
Must-Have Requirements
- Ray ecosystem (KubeRay, Train, Serve, Data)
- PyTorch
- Distributed training concepts (DDP, FSDP, NCCL, mixed precision)
- LLM serving engines (vLLM, SGLang, NVIDIA Triton)
- Flyte orchestration
- Vector databases (Pgvector, Qdrant)
- Python programming
- Model lifecycle tools (MLflow)
- ML platform engineering
- distributed training
- LLM serving
- Bachelor's degree in Computer Science, Engineering, or related field
Nice-to-Have Signals
- Quantization techniques (INT8/INT4/FP8)
- quantization
- reinforcement learning
Work Setup
- Location: Milpitas, USA
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Relocation
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Manage the full Ray stack on GKE, including KubeRay, scheduling, and object store configuration
- Run multi‑node GPU training jobs with Ray Train, handling checkpointing, spot‑preemption recovery, and fine‑tuning workflows
- Create a unified LLM inference service using Ray Serve, vLLM, SGLang, and NVIDIA Triton with zero‑copy memory sharing
- Tune inference performance through fractional GPU allocation, continuous batching, autoscaling, and KV‑cache sizing
- Implement capability‑, version‑, and tenant‑based model routing with cost‑aware fallback
- Build reinforcement‑learning pipelines in Flyte, integrating RLlib or custom PPO loops
- Run the end‑to‑end model promotion process from shadow mode through A/B testing to full rollout
- Automate retraining pipelines with drift detection, warm‑start, and quality‑gate checks
- Add Retrieval‑Augmented Generation using vector search (Pgvector/Qdrant) for context‑aware LLM responses
Good Fit If You Have
- Familiarity with post‑training quantization techniques (INT8/INT4/FP8)
- Experience applying reinforcement‑learning algorithms such as PPO or RLHF
- Exposure to vector‑search databases and ANN indexing
Skills
- Ray (KubeRay, Train, Serve, Data)
- PyTorch
- LLM serving (vLLM, SGLang, NVIDIA Triton)
- Distributed training (DDP, FSDP, NCCL, mixed precision)
- Flyte orchestration
- Vector databases (Pgvector, Qdrant)
- Python
- MLflow model registry
- Quantization (INT8/INT4/FP8) – nice to have