
Job Description
Ovii's Interpretation of the Role
Senior Site Reliability Engineer responsible for end‑to‑end reliability of a cloud‑native healthcare data platform. Owns infrastructure design, automation, observability, and incident response at scale.
Role Snapshot
- Own full lifecycle of production systems
- Drive reliability standards and SLIs/SLOs
- Automate infrastructure with IaC
- Scale GCP workloads efficiently
- Mentor teams on reliability best practices
Must-Have Requirements
- Linux systems administration
- Incident response & root cause analysis
- GCP (GKE, Cloud Run)
- Terraform or Pulumi
- CI/CD pipelines (GitHub Actions)
- Python/Bash/Go scripting
- Observability platforms (Google Cloud Monitoring, Datadog, Grafana, Prometheus)
- SLI/SLO/error‑budget design
- Autoscaling & resource optimization
- SRE/DevOps/Platform Engineering
- Operating production systems at scale
Nice-to-Have Signals
- Large‑scale distributed systems experience
- Healthcare domain familiarity
- AI‑assisted observability tooling
- NodeJS/TypeScript/Java/React familiarity
- Large‑scale distributed systems
- Healthcare data platforms
Work Setup
- Location: United States
- Work mode: REMOTE
- Remote scope: COUNTRY_RESTRICTED
- Remote countries: United States
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Design and operate highly available GCP‑based services handling millions of provider records
- Build and maintain IaC pipelines to provision and update infrastructure
- Create and manage SLIs, SLOs, error budgets, and alerting to reduce noise
- Improve autoscaling behavior and resource utilization across distributed workloads
- Lead incident response, root‑cause analysis, and post‑mortem processes
- Instrument data freshness and health metrics for the provider data platform
- Collaborate with engineering teams to embed reliability into deployment workflows
Good Fit If You Have
- Enjoys preventing incidents more than reacting to them
- Comfortable influencing operational standards across multiple teams
- Strong written and verbal communication for cross‑functional audiences
Skills
- Linux systems administration
- GCP (GKE, Cloud Run)
- Terraform / Pulumi
- CI/CD (GitHub Actions)
- Python / Bash / Go scripting
- Observability (Prometheus, Grafana, Datadog, Cloud Monitoring)
- SLI / SLO design
- Autoscaling & resource optimization
Remote Eligibility
- United States
- US