Ovii Job Board

Site Reliability Engineer - Public Sector

Blitzy

United States • Remote - United States • Full-Time • 3+ years

Posted 2026-08-05 USD 140,000 - USD 170,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Blitzy seeks a remote Site Reliability Engineer to own the deployment, operation, and observability of its self‑hosted AI platform for a public‑sector customer. The role blends deep technical ownership with direct partnership with customer security and infrastructure teams.

Role Snapshot

  • Site Reliability Engineer
  • Public‑sector focus
  • Remote (U.S.)
  • Kubernetes‑based deployments
  • Embedded engineer on customer account
  • Observability & incident management

Must-Have Requirements

  • Kubernetes
  • Infrastructure-as-Code (Terraform / Pulumi)
  • Cloud platform (AWS / Azure / GCP)
  • Observability tooling
  • Incident management & on‑call
  • Scripting (Python, Go, Bash)
  • U.S. citizenship
  • Site Reliability Engineering
  • Kubernetes orchestration
  • Regulated network environments
  • U.S. citizenship required

Nice-to-Have Signals

  • Experience with government‑accredited cloud environments
  • AI/ML workload infrastructure experience
  • Embedded/forward‑deployed engineer background
  • Startup high‑growth experience
  • Familiarity with regulated‑industry security frameworks
  • Government‑accredited cloud deployments
  • AI/ML workload support
  • AI/ML

Work Setup

  • Location: United States
  • Work mode: REMOTE
  • Remote scope: COUNTRY_RESTRICTED
  • Remote countries: United States
  • Travel: occasional travel for key customer workshops
  • Employment type: Full-Time

Eligibility Gates

  • Work authorization: U.S. citizenship required
  • Visa sponsorship: no
  • Security clearance: none
  • Background check: customer background/badging process

Not Specified in JD

  • Visa sponsorship
  • Relocation
  • Notice period
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub

What You'll Likely Work On

  • Deploy, operate, and maintain Blitzy's self‑hosted platform inside a customer‑controlled, secure cloud environment.
  • Own the Kubernetes deployment lifecycle, including releases, upgrades, capacity planning, and performance benchmarking for AI workloads.
  • Design and run observability pipelines—logging, metrics, tracing, and alerting—fully within the customer's security boundary.
  • Partner with customer infrastructure, security, and governance teams on provisioning, reviews, documentation, and escalations.
  • Handle sensitive data per strict security requirements and champion best‑practice security controls.
  • Feed operational lessons back into Blitzy's product and infrastructure roadmap for future public‑sector deployments.

Good Fit If You Have

  • Strong communication with customer engineering, security, and governance stakeholders.
  • Experience operating in highly regulated or restricted network environments.
  • Background supporting AI/ML workloads or related infrastructure.
  • Prior forward‑deployed or embedded‑engineer experience at an enterprise site.
  • Adaptability from high‑growth startup environments.

Skills

  • Kubernetes
  • Infrastructure‑as‑Code (Terraform / Pulumi)
  • Cloud platform (AWS / Azure / GCP)
  • Python / Go / Bash scripting
  • Observability tooling (logging, metrics, tracing, alerting)
  • Incident response & on‑call practices
  • Secure/regulated cloud environments
  • Capacity planning & performance benchmarking

Remote Eligibility

  • United States