Ovii Job Board

Site Reliability Engineer (SRE) – AWS / Kubernetes / DevOps

techdome

Indore, India • Onsite - Indore, India • Full-Time • 3-7 years

Posted 2026-06-19 INR 5 - INR 10 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Techdome seeks a Site Reliability Engineer to own the availability, performance and scalability of its cloud‑based payments and platform services. The role blends automation, observability, CI/CD and incident response, with a focus on AWS, Kubernetes and AI‑assisted ops tooling.

Role Snapshot

  • Ensure high‑availability and performance of production systems
  • Build automation and IaC pipelines
  • Implement observability, metrics and alerting
  • Lead incident response and on‑call rotation
  • Optimize cloud cost and capacity planning

Must-Have Requirements

  • Site Reliability Engineering (3+ years)
  • AWS, GCP or Azure cloud experience
  • Docker and Kubernetes
  • Terraform and Ansible
  • Python, Go or Bash scripting
  • Linux, networking and distributed‑systems fundamentals
  • CI/CD pipeline tools (Jenkins, GitHub Actions, GitLab CI)
  • Observability tools (Prometheus, Grafana, ELK, Datadog)
  • Site Reliability Engineering
  • DevOps
  • Platform engineering

Nice-to-Have Signals

  • AI / LLM‑powered tooling for ops automation
  • Payments / fintech production experience
  • SLO‑driven reliability and on‑call process improvement
  • Payments / fintech production
  • AI / LLM ops tooling
  • payments
  • fintech

Work Setup

  • Location: Indore, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Travel
  • Security clearance
  • Coding test

What You'll Likely Work On

  • Maintain high availability, performance and scalability of cloud production services
  • Create automation for deployment, monitoring and incident response
  • Deploy observability stack – metrics, logging, tracing and alerting
  • Define and manage SLIs, SLOs and error budgets
  • Build and maintain CI/CD pipelines and infrastructure as code
  • Lead incident response, root‑cause analysis and post‑mortems
  • Conduct capacity planning and cloud cost optimization
  • Participate in an on‑call rotation for production support

Good Fit If You Have

  • Prior experience with payments or fintech production systems
  • Familiarity with AI/LLM tools for operational automation
  • Experience driving SLO‑based reliability improvements
  • Comfort with multi‑cloud environments (AWS, GCP, Azure)
  • Ability to thrive in a fast‑growth, collaborative team

Skills

  • AWS / GCP / Azure
  • Kubernetes & Docker
  • Terraform & Ansible
  • Python / Go / Bash scripting
  • Linux, networking & distributed‑systems fundamentals
  • Prometheus, Grafana, ELK, Datadog
  • CI/CD (Jenkins, GitHub Actions, GitLab CI)
  • AI / LLM‑powered ops tooling (preferred)
  • Payments / fintech domain experience (preferred)