Ovii Job Board

Incident Operations Specialist

zapier

San Francisco, United States • Remote - San Francisco, United States • Full-Time

Posted 2026-07-10 USD 119,000 - USD 178,500 per year BPO & Shared Services

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Incident Operations Specialist ensures Zapier’s incident response tooling runs smoothly, builds AI‑driven automation, and turns incident data into actionable improvements. The role blends operational ownership with hands‑on technical work, requiring strong incident‑management background and AI‑tool fluency.

Role Snapshot

  • Incident tooling ownership
  • AI‑powered workflow automation
  • Incident analysis & program improvement
  • Data reporting & dashboard maintenance
  • Community coaching for responders
  • Cross‑functional communication

Must-Have Requirements

  • AI‑native tool proficiency
  • Incident management platform experience
  • SQL / Databricks querying
  • Slack workflow automation
  • incident response and reliability
  • AI‑driven workflow automation
  • Must have AI fluency (use AI‑native tools as default)
  • Must have incident response / reliability experience

Nice-to-Have Signals

  • Data visualization tools (Grafana, Looker)
  • API integration

Work Setup

  • Location: San Francisco, United States
  • Work mode: REMOTE
  • Remote scope: UNSPECIFIED
  • Remote countries: United States, NAMER
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility specifics
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Own and maintain incident.io, PagerDuty, on‑call rotations, escalation paths, and Slack‑based workflows.
  • Design, build, and sustain AI‑driven workflows for summarisation, post‑mortem drafting, triage, severity classification, and data hygiene.
  • Participate in incident response and post‑mortems, identify recurring patterns, and recommend program‑level fixes.
  • Create, troubleshoot, and monitor dashboards and reports using Databricks, Grafana, and Looker while ensuring data quality.
  • Grow and coach the community of Incident Commanders and Support Leads, sharing best practices and playbooks.
  • Maintain clear, usable documentation and templates that stay reliable under pressure.

Good Fit If You Have

  • Proven experience using AI‑native tools as your default work environment.
  • Hands‑on incident response, on‑call rotation, and escalation experience.
  • Ability to translate technical details into plain language for support, GTM, and leadership audiences.
  • Comfort building lightweight AI agents, API integrations, and Slack workflows.
  • Self‑starter who prioritises work ruthlessly and pushes back on off‑program requests.

Skills

  • AI‑native tool proficiency (Cursor, Claude, Copilot, etc.)
  • Incident management platforms (incident.io, PagerDuty)
  • SQL & Databricks querying
  • Slack workflow automation
  • Data visualization & reporting (Grafana, Looker)
  • API integration & lightweight AI agents
  • Async‑first communication

Remote Eligibility

  • United States
  • NAMER