
Technical Lead, Platform Engineering (Observability)
mercari
Posted 2026-08-14
Tech & Engg
Job Description
Ovii's Interpretation of the Role
We are seeking a Technical Lead to own Mercari's observability platform, driving reliability, AI‑powered incident detection, and self‑service tooling for 500+ microservices. The role blends deep engineering expertise with mentorship and cross‑team collaboration.
Role Snapshot
- Technical lead for observability platform
- Own end‑to‑end observability stack
- Mentor and shape engineering culture
- Drive AI‑enabled incident detection
- Define observability standards and SLOs
- Collaborate with SRE, security, and product teams
Must-Have Requirements
- Go or Python
- Kubernetes
- Terraform
- GCP or AWS
- Observability platforms (Datadog, Prometheus, Grafana)
- Metrics, logging, distributed tracing
- SLO/SLI frameworks
- building scalable production systems
- observability platform development
Nice-to-Have Signals
- AI/ML for anomaly detection
- OpenTelemetry
- Cost optimization of observability data
- Open‑source observability contributions
- AI for observability
- large‑scale distributed systems
Work Setup
- Location: Bengaluru, India
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, build, and operate a scalable observability platform covering metrics, logs, traces, and alerts
- Reduce Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) through AI‑powered anomaly detection and alert correlation
- Create self‑service tooling that enables product engineers to instrument, monitor, and alert on their services independently
- Define and champion observability standards, best practices, and SLO frameworks across the organization
- Collaborate with SRE, security, and product engineering teams to ensure comprehensive system visibility
- Automate operational workflows and improve platform efficiency
- Lead technical decisions, mentor engineers, and foster a strong engineering culture within the observability team
Good Fit If You Have
- Enjoys turning noisy alerts into actionable insights
- Passionate about improving developer experience through platform tooling
- Thrives in a fast‑moving, AI‑enabled environment
Skills
- Go or Python
- Kubernetes
- Terraform
- GCP / AWS
- Prometheus / Datadog / Grafana
- Metrics, logs, distributed tracing
- SLO / SLI frameworks
- AI/ML for anomaly detection
- OpenTelemetry
- Developer tooling & self‑service dashboards