Ovii Job Board

Site Reliability Engineer

techdome

Hyderabad, India • Onsite - Hyderabad, India • Full-Time • 2+ years

Posted 2026-08-15 INR 600,000 - INR 1,200,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Techdome seeks a Site Reliability Engineer to own production reliability for Healthcare, FinTech, AI, and SaaS services. You will build zero‑downtime pipelines, manage multi‑cloud environments, and lead incident response on‑call. The role demands strong cloud, IaC, and observability expertise.

Role Snapshot

  • Own production reliability
  • Build zero‑downtime CI/CD pipelines
  • Manage multi‑cloud (AWS, Azure, GCP)
  • Drive observability & alerting
  • Automate incident workflows
  • Participate in on‑call rotation

Must-Have Requirements

  • 2+ years as SRE/DevOps/Platform/Cloud Engineer
  • Hands‑on production with AWS, Azure, or GCP
  • Docker and Kubernetes experience
  • Infrastructure‑as‑Code (Terraform, Ansible)
  • CI/CD pipelines built from scratch (Jenkins, GitHub Actions, GitLab CI)
  • Strong Linux, networking, and distributed‑systems fundamentals
  • Scripting in Python, Go, or Bash
  • Experience with Blue‑Green, Canary, and Rolling deployments
  • Define and own SLIs, SLOs, error budgets
  • SRE/DevOps/Platform production
  • Incident response
  • Cloud operations

Nice-to-Have Signals

  • Background in FinTech, Payments, or Healthcare
  • Regular use of AI tools (Copilot, Claude, ChatGPT, etc.)
  • AI‑powered ops workflow experience
  • Fluent SRE terminology and chaos engineering
  • FinTech/Payments/Healthcare domain
  • AI‑assisted operational workflows
  • FinTech
  • Payments
  • Healthcare
  • AI
  • SaaS

Work Setup

  • Location: Hyderabad, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Keep production services available, reliable, and performant across all environments
  • Manage and optimize cloud platforms (AWS, Azure, GCP)
  • Design and operate zero‑downtime CI/CD pipelines using Blue‑Green, Canary, and Rolling strategies
  • Codify infrastructure with Terraform and Ansible
  • Implement observability stacks with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry
  • Define and own SLIs, SLOs, and error budgets
  • Lead incident response, root‑cause analysis, and post‑mortems
  • Optimize cloud spend and plan capacity
  • Automate operational workflows, including AI‑powered alert triage and incident summarization
  • Participate in the on‑call rotation

Good Fit If You Have

  • Experience in FinTech, Payments, or Healthcare domains
  • Regular use of AI coding assistants (Copilot, Claude, ChatGPT, etc.)
  • Built AI‑powered operational workflows
  • Familiarity with chaos engineering concepts

Skills

  • AWS / Azure / GCP
  • Terraform & Ansible
  • Docker & Kubernetes
  • CI/CD (Jenkins, GitHub Actions, GitLab CI)
  • Prometheus, Grafana, ELK, Datadog, OpenTelemetry
  • Python / Go / Bash scripting
  • Linux, networking & distributed‑systems fundamentals
  • SLO/SLI & error‑budget management
  • Cloud cost optimization
  • AI‑assisted ops tools