
Site Reliability Engineer
techdome
Posted 2026-08-15
INR 600,000 - INR 1,200,000 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Techdome seeks a Site Reliability Engineer to own production reliability for Healthcare, FinTech, AI, and SaaS services. You will build zero‑downtime pipelines, manage multi‑cloud environments, and lead incident response on‑call. The role demands strong cloud, IaC, and observability expertise.
Role Snapshot
- Own production reliability
- Build zero‑downtime CI/CD pipelines
- Manage multi‑cloud (AWS, Azure, GCP)
- Drive observability & alerting
- Automate incident workflows
- Participate in on‑call rotation
Must-Have Requirements
- 2+ years as SRE/DevOps/Platform/Cloud Engineer
- Hands‑on production with AWS, Azure, or GCP
- Docker and Kubernetes experience
- Infrastructure‑as‑Code (Terraform, Ansible)
- CI/CD pipelines built from scratch (Jenkins, GitHub Actions, GitLab CI)
- Strong Linux, networking, and distributed‑systems fundamentals
- Scripting in Python, Go, or Bash
- Experience with Blue‑Green, Canary, and Rolling deployments
- Define and own SLIs, SLOs, error budgets
- SRE/DevOps/Platform production
- Incident response
- Cloud operations
Nice-to-Have Signals
- Background in FinTech, Payments, or Healthcare
- Regular use of AI tools (Copilot, Claude, ChatGPT, etc.)
- AI‑powered ops workflow experience
- Fluent SRE terminology and chaos engineering
- FinTech/Payments/Healthcare domain
- AI‑assisted operational workflows
- FinTech
- Payments
- Healthcare
- AI
- SaaS
Work Setup
- Location: Hyderabad, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Keep production services available, reliable, and performant across all environments
- Manage and optimize cloud platforms (AWS, Azure, GCP)
- Design and operate zero‑downtime CI/CD pipelines using Blue‑Green, Canary, and Rolling strategies
- Codify infrastructure with Terraform and Ansible
- Implement observability stacks with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry
- Define and own SLIs, SLOs, and error budgets
- Lead incident response, root‑cause analysis, and post‑mortems
- Optimize cloud spend and plan capacity
- Automate operational workflows, including AI‑powered alert triage and incident summarization
- Participate in the on‑call rotation
Good Fit If You Have
- Experience in FinTech, Payments, or Healthcare domains
- Regular use of AI coding assistants (Copilot, Claude, ChatGPT, etc.)
- Built AI‑powered operational workflows
- Familiarity with chaos engineering concepts
Skills
- AWS / Azure / GCP
- Terraform & Ansible
- Docker & Kubernetes
- CI/CD (Jenkins, GitHub Actions, GitLab CI)
- Prometheus, Grafana, ELK, Datadog, OpenTelemetry
- Python / Go / Bash scripting
- Linux, networking & distributed‑systems fundamentals
- SLO/SLI & error‑budget management
- Cloud cost optimization
- AI‑assisted ops tools