
Sr Site Reliability Engineer
SigNoz
Posted 2026-06-23
INR 5,000,000 - INR 7,000,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
The Senior Site Reliability Engineer will own reliability, scalability, and operability of SigNoz's cloud observability platform. You will design SLO/SLI frameworks, scale petabyte‑scale ingest pipelines, and automate infrastructure using Kubernetes and IaC. The role is fully remote within India and requires deep production‑grade SRE experience.
Role Snapshot
- Own cloud platform reliability
- Scale petabyte‑scale ingest pipelines
- Operate and tune ClickHouse
- Drive Kubernetes at scale
- Define SLO/SLI and incident response
- Build IaC and CI/CD tooling
Must-Have Requirements
- Kubernetes (deep fluency)
- Distributed systems failure analysis
- Capacity planning
- Incident management / SLO‑SLI
- Infrastructure‑as‑Code / CI‑CD
- SRE
- Kubernetes
- distributed systems
- capacity planning
Nice-to-Have Signals
- ClickHouse operation and tuning
- Go programming
- OpenTelemetry familiarity
- Kafka or similar high‑throughput data systems
- Open‑source contributions
- ClickHouse
- Go
- OpenTelemetry
- observability
- startup environment
Work Setup
- Location: India
- Work mode: REMOTE
- Remote scope: COUNTRY_RESTRICTED
- Remote countries: India
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Travel requirement
- Security clearance
- Coding test
What You'll Likely Work On
- Define and maintain SLOs/SLIs, error budgets, and on‑call practices
- Scale and harden the ingest path for burst traffic while keeping data fresh
- Plan capacity and enable auto‑scalability for a petabyte‑scale SaaS system
- Operate, upgrade, and automate Kubernetes clusters across multi‑tenant workloads
- Tune ClickHouse and the data layer for performance and cost efficiency
- Develop and maintain IaC, CI/CD pipelines, and internal tooling
- Write clear runbooks, technical documentation, and communicate trade‑offs
Good Fit If You Have
- Experience on SRE teams at Series B+ startups
- Hands‑on work with Kafka or similar high‑throughput systems
- Contributions to open‑source projects
- Strong written communication for runbooks and docs
- Comfortable in a fast‑moving, remote‑first environment
Skills
- Kubernetes (advanced)
- Distributed systems failure analysis
- Capacity planning
- Incident management (SLO/SLI)
- Go programming (preferred)
- ClickHouse operations (preferred)
- OpenTelemetry (preferred)
- CI/CD & Infrastructure‑as‑Code
- Observability concepts
Remote Eligibility
- India