Ovii Job Board

Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation

Ambient.AI

Redwood City, United States • Hybrid - Redwood City, United States • Full-Time • 4+ years

Posted 2026-07-14 USD 168,000 - USD 205,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior Software Engineer building and scaling AI infrastructure that powers real‑time computer‑vision, LLM and multimodal inference at Ambient.ai. The role blends production ML systems, distributed engineering, and performance optimization to deliver reliable, low‑latency AI services.

Role Snapshot

  • Senior AI infrastructure engineer
  • Hybrid (Redwood City) full‑time
  • 4+ years production ML experience
  • Python‑centric development
  • GPU‑heavy model serving

Must-Have Requirements

  • Python
  • Distributed systems
  • Machine‑learning platform engineering
  • Model‑serving frameworks (vLLM, Triton)
  • GPU inference optimization
  • Container orchestration
  • Evaluation framework development
  • Infrastructure engineering
  • Production ML platforms
  • Inference optimization
  • BS/MS in Computer Science or related technical field

Nice-to-Have Signals

  • Computer vision experience
  • CUDA / NCCL / PyTorch / TensorRT / ONNX
  • Large‑scale GPU infrastructure operation
  • Model compression, distillation, pruning
  • Retrieval‑augmented generation, vector databases
  • Familiarity with RAG pipelines
  • Experience with prompt evaluation or human‑in‑the‑loop testing
  • Computer vision
  • Large‑scale GPU infrastructure
  • Model compression techniques

Work Setup

  • Location: Redwood City, United States
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility

What You'll Likely Work On

  • Design and operate scalable AI infrastructure for real‑time video and sensor data
  • Build and optimize model‑serving stacks, including batching, caching, routing and quantization
  • Create robust evaluation harnesses and continuous‑learning pipelines
  • Collaborate with research scientists to productionize vision‑language, LLM and multimodal models
  • Develop observability, monitoring and debugging tools for AI services
  • Engineer data engines for training‑data collection and feedback loops

Good Fit If You Have

  • Enjoys solving performance bottlenecks in GPU‑heavy workloads
  • Thrives in a fast‑paced, cross‑functional environment
  • Strong communication skills for partnering with research and product teams

Skills

  • Python programming
  • Distributed systems design
  • Machine‑learning platform engineering
  • Model‑serving frameworks (vLLM, Triton)
  • GPU inference optimization (batching, quantization, parallelism)
  • Container orchestration (Docker/Kubernetes)
  • Evaluation & benchmarking pipelines
  • Cloud infrastructure for AI workloads