
Job Description
Ovii's Interpretation of the Role
We need a Senior Site Reliability Engineer to own the full lifecycle of our cloud‑native healthcare data platform, driving reliability, automation, and observability at scale.
Role Snapshot
- Own end‑to‑end operational lifecycle
- Design and enforce reliability standards
- Build and maintain IaC & CI/CD pipelines
- Lead incident response and post‑mortems
- Optimize autoscaling and cost efficiency
Must-Have Requirements
- GCP (GKE, Cloud Run)
- Terraform / Pulumi
- Docker / Kubernetes
- CI/CD (GitHub Actions)
- Python / Bash / Go
- Linux systems administration
- Observability platforms (Cloud Monitoring, Datadog, Grafana, Prometheus)
- SLI / SLO design
- Incident response & root‑cause analysis
- SRE
- DevOps
- Platform Engineering
- Infrastructure Engineering
Nice-to-Have Signals
- Healthcare domain familiarity
- AI‑assisted observability tooling
- Large‑scale distributed systems experience
- NodeJS / TypeScript / Java / React familiarity
- Healthcare provider data
- Large‑scale distributed systems
- Healthcare
Work Setup
- Location: United States
- Work mode: REMOTE
- Remote scope: COUNTRY_RESTRICTED
- Remote countries: United States
- Employment type: Full-Time
What You'll Likely Work On
- Define and implement reliability standards, SLIs, SLOs, and alerting across the platform
- Automate infrastructure provisioning and configuration using Terraform or Pulumi
- Develop and maintain CI/CD pipelines that deliver containerized workloads to GKE and Cloud Run
- Monitor system health, reduce alert fatigue, and create actionable dashboards
- Lead incident response, root‑cause analysis, and post‑mortem processes
- Improve autoscaling policies and resource utilization for cost‑effective scaling
- Mentor engineering teams on reliability best practices and security hygiene
- Collaborate with product and compliance stakeholders to ensure data‑privacy requirements
Good Fit If You Have
- Strong written and verbal communication for cross‑functional audiences
- Experience influencing operational standards and mentoring teams
- Background with regulated or PII‑sensitive data environments
Skills
- GCP (GKE, Cloud Run)
- Terraform / Pulumi
- Docker & Kubernetes
- CI/CD (GitHub Actions / Cloud Build)
- Python / Bash / Go scripting
- Linux systems administration
- Observability (Cloud Monitoring, Datadog, Grafana, Prometheus)
- SLI / SLO / error‑budget design
- Security, secrets, and compliance management
Remote Eligibility
- United States
- US