
Job Description
Ovii's Interpretation of the Role
Supabase seeks a senior Release Engineer (SRE) to own the reliability of its deployment and release systems. You will define SLOs, build observability, drive disaster‑recovery readiness, and lead on‑call incident response for a globally distributed, fully remote team.
Role Snapshot
- Own deployment reliability and control plane
- Define and track SLOs, error budgets, DORA metrics
- Build health monitoring and synthetic testing
- Lead on‑call, blameless postmortems, and automation
- Standardize CI/CD pipelines and auditability
- Document runbooks, break‑glass procedures, and access models
Must-Have Requirements
- SRE practices (SLAs, SLOs, error budgets, DORA metrics)
- Observability tooling (Prometheus, Grafana, Alertmanager)
- Incident management (PagerDuty, Opsgenie, incident.io)
- AWS cloud operations
- Infrastructure as code (Terraform, Pulumi)
- Kubernetes
- Scripting/automation
- CI/CD pipeline automation
- SRE
- production operations
- platform engineering
- release engineering
Work Setup
- Location: Remote, Anywhere
- Work mode: REMOTE
- Remote scope: GLOBAL
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility details
- Education requirement
- Certifications
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Own the reliability of Supabase's deployment and release systems against defined SLOs and error budgets
- Standardize and instrument fragmented deployment workflows to create trustworthy pre‑production signals
- Design and operate health monitoring, synthetic tests, and DORA metrics for critical user flows
- Reduce mean‑time‑to‑detect and mean‑time‑to‑recover for deploy‑related incidents
- Participate in on‑call rotation, lead blameless postmortems, and create runbooks, alerts, and automation to eliminate toil
- Improve deployment observability, auditability, and access controls, including break‑glass and scoped self‑service workflows
- Partner with product and platform teams to align release practices with reliability targets
Good Fit If You Have
- Enjoys async, globally distributed collaboration
- Comfortable navigating ambiguity and iterating on systems
- Strong communicator with both infrastructure specialists and product engineers
Skills
- SRE practices (SLAs, SLOs, error budgets, DORA metrics)
- Observability tooling (Prometheus, Grafana, Alertmanager)
- Incident management (PagerDuty, Opsgenie, incident.io)
- AWS cloud operations (IAM, VPC, multi‑account)
- Infrastructure as code (Terraform, Pulumi)
- Kubernetes orchestration
- Scripting/automation (Python, Bash, Go)
- CI/CD pipeline automation