
Site Reliability Engineer (SRE) – AWS / Kubernetes / DevOps
techdome
Posted 2026-06-19
INR 5 - INR 10 per year
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Techdome seeks a Site Reliability Engineer to own the availability, performance and scalability of its cloud‑based payments and platform services. The role blends automation, observability, CI/CD and incident response, with a focus on AWS, Kubernetes and AI‑assisted ops tooling.
Role Snapshot
- Ensure high‑availability and performance of production systems
- Build automation and IaC pipelines
- Implement observability, metrics and alerting
- Lead incident response and on‑call rotation
- Optimize cloud cost and capacity planning
Must-Have Requirements
- Site Reliability Engineering (3+ years)
- AWS, GCP or Azure cloud experience
- Docker and Kubernetes
- Terraform and Ansible
- Python, Go or Bash scripting
- Linux, networking and distributed‑systems fundamentals
- CI/CD pipeline tools (Jenkins, GitHub Actions, GitLab CI)
- Observability tools (Prometheus, Grafana, ELK, Datadog)
- Site Reliability Engineering
- DevOps
- Platform engineering
Nice-to-Have Signals
- AI / LLM‑powered tooling for ops automation
- Payments / fintech production experience
- SLO‑driven reliability and on‑call process improvement
- Payments / fintech production
- AI / LLM ops tooling
- payments
- fintech
Work Setup
- Location: Indore, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Maintain high availability, performance and scalability of cloud production services
- Create automation for deployment, monitoring and incident response
- Deploy observability stack – metrics, logging, tracing and alerting
- Define and manage SLIs, SLOs and error budgets
- Build and maintain CI/CD pipelines and infrastructure as code
- Lead incident response, root‑cause analysis and post‑mortems
- Conduct capacity planning and cloud cost optimization
- Participate in an on‑call rotation for production support
Good Fit If You Have
- Prior experience with payments or fintech production systems
- Familiarity with AI/LLM tools for operational automation
- Experience driving SLO‑based reliability improvements
- Comfort with multi‑cloud environments (AWS, GCP, Azure)
- Ability to thrive in a fast‑growth, collaborative team
Skills
- AWS / GCP / Azure
- Kubernetes & Docker
- Terraform & Ansible
- Python / Go / Bash scripting
- Linux, networking & distributed‑systems fundamentals
- Prometheus, Grafana, ELK, Datadog
- CI/CD (Jenkins, GitHub Actions, GitLab CI)
- AI / LLM‑powered ops tooling (preferred)
- Payments / fintech domain experience (preferred)