Ovii Job Board

Senior Site Reliability Engineer

CertifyOS

United States • Remote - United States • Full-Time • 5+ years

Posted 2026-06-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior Site Reliability Engineer responsible for end‑to‑end reliability of a cloud‑native healthcare data platform. Owns infrastructure design, automation, observability, and incident response at scale.

Role Snapshot

  • Own full lifecycle of production systems
  • Drive reliability standards and SLIs/SLOs
  • Automate infrastructure with IaC
  • Scale GCP workloads efficiently
  • Mentor teams on reliability best practices

Must-Have Requirements

  • Linux systems administration
  • Incident response & root cause analysis
  • GCP (GKE, Cloud Run)
  • Terraform or Pulumi
  • CI/CD pipelines (GitHub Actions)
  • Python/Bash/Go scripting
  • Observability platforms (Google Cloud Monitoring, Datadog, Grafana, Prometheus)
  • SLI/SLO/error‑budget design
  • Autoscaling & resource optimization
  • SRE/DevOps/Platform Engineering
  • Operating production systems at scale

Nice-to-Have Signals

  • Large‑scale distributed systems experience
  • Healthcare domain familiarity
  • AI‑assisted observability tooling
  • NodeJS/TypeScript/Java/React familiarity
  • Large‑scale distributed systems
  • Healthcare data platforms

Work Setup

  • Location: United States
  • Work mode: REMOTE
  • Remote scope: COUNTRY_RESTRICTED
  • Remote countries: United States
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test

What You'll Likely Work On

  • Design and operate highly available GCP‑based services handling millions of provider records
  • Build and maintain IaC pipelines to provision and update infrastructure
  • Create and manage SLIs, SLOs, error budgets, and alerting to reduce noise
  • Improve autoscaling behavior and resource utilization across distributed workloads
  • Lead incident response, root‑cause analysis, and post‑mortem processes
  • Instrument data freshness and health metrics for the provider data platform
  • Collaborate with engineering teams to embed reliability into deployment workflows

Good Fit If You Have

  • Enjoys preventing incidents more than reacting to them
  • Comfortable influencing operational standards across multiple teams
  • Strong written and verbal communication for cross‑functional audiences

Skills

  • Linux systems administration
  • GCP (GKE, Cloud Run)
  • Terraform / Pulumi
  • CI/CD (GitHub Actions)
  • Python / Bash / Go scripting
  • Observability (Prometheus, Grafana, Datadog, Cloud Monitoring)
  • SLI / SLO design
  • Autoscaling & resource optimization

Remote Eligibility

  • United States
  • US