Ovii Job Board

AI Systems & Platform Internals - Technical Architect

accellor

San Francisco, United States • Onsite - San Francisco, United States • Full-Time • 10-12 years

Posted 2026-08-04 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior Technical Architect driving the design, scaling, and reliability of AI inference and platform internals at Accellor. Owns end‑to‑end architecture across GPU performance, distributed serving, context engineering, cost optimization, and production safety for large‑scale LLM workloads.

Role Snapshot

  • Senior Technical Architect
  • AI inference platform
  • Large‑scale distributed systems
  • GPU & accelerator performance
  • Cost & performance engineering
  • Context engineering
  • Technical leadership

Must-Have Requirements

  • Python
  • C++ / Go / Rust / Java / TypeScript
  • Distributed systems design
  • GPU programming (CUDA, Triton)
  • ML frameworks (PyTorch, JAX, TensorFlow)
  • Model serving stacks (Triton, vLLM, Ray)
  • Kubernetes & cloud infrastructure
  • software engineering
  • systems architecture
  • ML infrastructure

Nice-to-Have Signals

  • LLM inference experience
  • Multimodal inference experience
  • Agentic systems experience
  • Tensor/pipeline parallelism and model sharding
  • Cost optimization dashboards
  • GPU profiling tools (Nsight, rocprof)
  • Release gates and canary validation
  • Large‑scale distributed training
  • Reinforcement‑learning infrastructure
  • LLM inference platforms
  • multimodal inference
  • agentic AI systems
  • Information Technology and Services
  • Healthcare
  • Life Sciences
  • Telecom
  • Retail
  • Financial Services
  • Technology

Work Setup

  • Location: San Francisco, United States
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test

What You'll Likely Work On

  • Design and evolve large‑scale AI systems that power ChatGPT, OpenAI API, Codex, and multimodal workloads.
  • Architect high‑throughput, low‑latency inference and model‑serving pipelines across GPU clusters.
  • Analyze and improve GPU kernel performance, memory bandwidth, and distributed execution.
  • Build context‑engineering platforms for prompt structuring, retrieval‑augmented generation, and long‑context management.
  • Create cost‑optimization frameworks that reduce token usage, redundant calls, and infrastructure spend.
  • Collaborate on training and research infrastructure for large‑scale model training and post‑training workflows.
  • Define release safety, validation, and evaluation gates to ensure safe, regression‑free platform updates.
  • Design reliability, observability, and production‑operations tooling (telemetry, alerts, runbooks).
  • Support agentic and multimodal platform internals, including tool use, memory, and workflow orchestration.
  • Provide technical leadership, mentorship, and cross‑team architecture guidance.

Good Fit If You Have

  • Hands‑on experience with LLM or multimodal inference platforms.
  • Familiarity with GPU profiling tools such as Nsight Systems or Compute.
  • Background in building cost‑aware AI infrastructure or token‑budgeting systems.

Skills

  • Python
  • C++ / Go / Rust / Java / TypeScript
  • Distributed systems design
  • GPU programming (CUDA, Triton)
  • ML frameworks (PyTorch, JAX, TensorFlow)
  • Model serving stacks (Triton, vLLM, Ray)
  • Kubernetes & cloud infrastructure
  • Observability (Prometheus, Grafana, OpenTelemetry)
  • LLM inference (preferred)
  • Multimodal inference (preferred)
  • Agentic systems (preferred)
  • Cost optimization dashboards (preferred)