Ovii Job Board

Staff Engineer – Observability Platform

postman

Onsite • Full-Time • 10+ years

Posted 2026-06-24 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Postman seeks a Staff Engineer to own the technical vision of its Observability Platform. The role drives architecture, builds scalable telemetry solutions, and partners across infra, security, and product teams to improve reliability and developer productivity.

Role Snapshot

  • Define observability platform architecture
  • Design scalable metrics, logging, and tracing solutions
  • Lead cross‑team technical initiatives
  • Mentor senior engineers
  • Establish standards and instrumentation frameworks
  • Collaborate with infrastructure, security, and data teams

Must-Have Requirements

  • distributed systems
  • cloud‑native architectures
  • observability domains (metrics, logging, tracing)
  • Go/Java/Python/Node.js
  • OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb
  • incident management
  • software engineering
  • observability

Nice-to-Have Signals

  • building internal developer platforms
  • AIOps / intelligent alerting / anomaly detection
  • reliability initiatives across multiple orgs
  • SRE / Platform Engineering background
  • AIOps and intelligent alerting

Work Setup

  • Work mode: ONSITE
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement

What You'll Likely Work On

  • Drive the technical vision and architecture for the Observability Platform
  • Design and implement scalable telemetry pipelines for metrics, logs, and traces
  • Investigate complex production incidents, perform root‑cause analysis, and define long‑term fixes
  • Partner with engineering teams to raise service reliability, availability, and performance
  • Create and enforce observability standards, best practices, and instrumentation across the company
  • Build automation and tooling to accelerate incident detection, diagnosis, and remediation
  • Use telemetry data to identify performance bottlenecks, capacity risks, and reliability gaps
  • Lead cross‑functional initiatives that improve platform health and engineering productivity
  • Mentor senior engineers and elevate the technical bar organization‑wide

Good Fit If You Have

  • Experience building internal developer or observability platforms at scale
  • Exposure to AIOps, intelligent alerting, or AI‑powered operational tooling
  • Background in SRE, Platform Engineering, or Infrastructure Engineering

Skills

  • Distributed systems & cloud‑native architectures
  • Observability domains (metrics, logging, tracing, alerting)
  • Go / Java / Python / Node.js
  • OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, Honeycomb
  • Incident management & production operations
  • Performance optimization & capacity analysis
  • Automation tooling for incident detection