
Senior Site Reliability Engineer
Blitzy
Posted 2026-06-10
USD 160,000 - USD 180,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Blitzy is hiring a Senior Site Reliability Engineer to own reliability, scalability, and operational excellence of its AI‑driven development platform. The role blends software engineering with infrastructure, designing fault‑tolerant cloud systems, building observability pipelines, and partnering with engineering teams to embed SRE practices.
Role Snapshot
- Own reliability of AI development platform
- Design and operate cloud infrastructure (AWS/GCP/Azure)
- Build observability and incident response systems
- Define SLOs/SLAs and drive post‑mortem improvements
- Automate CI/CD pipelines and deployment infrastructure
- Lead capacity planning, performance benchmarking, and cost optimization
- Champion security best practices in infrastructure
Must-Have Requirements
- Major cloud platform (AWS preferred)
- Kubernetes & container orchestration
- Infrastructure as Code (Terraform, Pulumi)
- Observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
- Scripting (Python, Go, Bash)
- Incident management & on‑call practices
- CI/CD pipeline automation
- Site Reliability Engineering
- Infrastructure as Code
- Observability
- Must work on‑site in Cambridge, MA
Nice-to-Have Signals
- AI/ML workload support
- High‑growth startup experience
- Global multi‑region infrastructure design
- eBPF & service‑mesh (Istio, Linkerd)
- Open‑source SRE/DevOps contributions
- High‑growth startup environment
Work Setup
- Location: Cambridge, USA
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, build, and operate scalable, fault‑tolerant infrastructure across major cloud providers.
- Define and enforce SLOs, SLAs, and error budgets; lead blameless post‑mortems and drive systemic improvements.
- Develop and maintain robust CI/CD pipelines, release automation, and deployment infrastructure.
- Own the observability stack—logging, metrics, tracing, and alerting using Prometheus, Grafana, Datadog, OpenTelemetry.
- Partner with software engineering teams to embed reliability practices and improve incident response, reducing MTTR.
- Conduct capacity planning, performance benchmarking, and cost optimization across the platform.
- Champion security best practices within infrastructure and deployment layers.
Good Fit If You Have
- Experience supporting AI/ML or GPU‑accelerated workloads.
- Background in fast‑growth startup environments wearing multiple hats.
- Familiarity with eBPF and service‑mesh technologies such as Istio or Linkerd.
- Contributions to open‑source SRE/DevOps tooling or communities.
- Experience building global, multi‑region infrastructure with strict latency and availability requirements.
Skills
- Cloud platforms (AWS, GCP, Azure)
- Kubernetes & container orchestration
- Infrastructure as Code (Terraform, Pulumi)
- Observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
- Scripting (Python, Go, Bash)
- Incident management & on‑call practices
- CI/CD pipeline automation