
Job Description
Ovii's Interpretation of the Role
Supabase is seeking a senior Site Reliability Engineer to embed within Service Operations and raise reliability across all engineering teams. The role focuses on defining SLIs/SLOs, driving operational readiness, and building automation that reduces toil while influencing without direct authority.
Role Snapshot
- Drive reliability practices across engineering teams
- Define and operationalize SLIs, SLOs, and error budgets
- Lead Operational Readiness Review process
- Build automation to eliminate operational toil
- Partner with services on architecture and resilience design
- Shape on‑call and incident response practices
- Report org‑wide operational maturity
Must-Have Requirements
- SRE methodology (SLIs/SLOs, error budgets)
- Software engineering (coding & tooling)
- Incident response & postmortem facilitation
- SRE
- Production engineering
- Reliability‑focused roles
Nice-to-Have Signals
- AWS cloud infrastructure
- Infrastructure as Code (Pulumi, Terraform, CDK)
- Observability tooling (OpenTelemetry, Grafana, VictoriaMetrics)
- Kubernetes platform operations
- Developer‑facing reliability tooling
- PostgreSQL / managed database familiarity
- Large‑scale multi‑tenant systems
- Managed database platforms
Work Setup
- Location: Remote, Anywhere
- Work mode: REMOTE
- Remote scope: GLOBAL
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility specifics
- Education requirement
- Certifications
- Travel
- Coding test
What You'll Likely Work On
- Collaborate with service teams to craft meaningful SLIs/SLOs and enforce error‑budget policies
- Own and continuously improve the Operational Readiness Review (ORR) workflow for new services and major changes
- Translate postmortem findings into systemic fixes and reduce repeat failure patterns
- Provide reliability expertise for architecture reviews, failure‑mode analysis, and resilience design
- Identify operational toil and champion automation to streamline processes
- Design sustainable on‑call practices, improving alert quality and runbook coverage
- Measure and communicate organization‑wide operational maturity, driving remediation initiatives
Good Fit If You Have
- Thrives in async, globally distributed environments
- Strong communicator who can influence without authority
- Passionate about enabling other teams to own reliability
- Deep experience with large‑scale, multi‑tenant platforms
- Enjoys building developer‑focused reliability tooling
Skills
- SRE methodology (SLIs/SLOs, error budgets)
- Software engineering (coding & tooling)
- Incident response & postmortem facilitation
- Cloud infrastructure (AWS)
- Infrastructure as Code (Pulumi, Terraform, CDK)
- Observability tooling (OpenTelemetry, Grafana, VictoriaMetrics)
- Kubernetes platform operations
- Developer‑facing reliability tooling
- Large‑scale multi‑tenant system experience
- PostgreSQL / managed database familiarity