Ovii Job Board

Principal Site Reliability Engineer I

Zeta

Hyderabad, India • Onsite - Hyderabad, India • Full-Time • 10-15 years

Posted 2026-06-30 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Zeta seeks a Principal Site Reliability Engineer to own reliability, scalability, and automation of its cloud‑native banking platforms. The role focuses on infrastructure as code, incident response, performance tuning, and mentoring within an engineering‑first culture.

Role Snapshot

  • Principal SRE (individual contributor)
  • 10‑15 years of SRE experience
  • Onsite in Hyderabad
  • Full‑time employment
  • Banking technology domain

Must-Have Requirements

  • Programming languages (Python, Go, Shell/Bash)
  • Automation tools (Ansible, Puppet, Chef, custom scripts)
  • Infrastructure as Code (Terraform)
  • Containerization (Docker) and orchestration (Kubernetes)
  • Cloud platforms (AWS, Azure, GCP)
  • Monitoring & logging (Prometheus, Grafana, ELK)
  • Networking concepts
  • Security best practices
  • CI/CD pipeline implementation
  • Version control (Git)
  • site reliability engineering
  • B.Tech/M.Tech in computer science, information technology or related field
  • Must be able to work onsite in Hyderabad

Nice-to-Have Signals

  • Experience in a product‑focused organization
  • Cloud certifications (AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, Microsoft Certified)
  • product organization experience

Work Setup

  • Location: Hyderabad, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Remote eligibility
  • Certifications (required)
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design and operate scalable, highly available infrastructure for banking services
  • Build automation tools and IaC pipelines to reduce manual operational work
  • Lead incident response, root‑cause analysis, and post‑mortem processes
  • Plan capacity and performance tuning for population‑scale systems
  • Implement monitoring, alerting, and logging to ensure system health
  • Collaborate with security teams to embed hardening practices
  • Develop disaster recovery strategies and test recovery procedures
  • Mentor junior engineers and promote SRE best practices

Good Fit If You Have

  • Enjoys deep technical problem solving at scale
  • Thrives in a collaborative, ownership‑driven environment
  • Has product‑focused mindset and can translate business needs into reliable systems

Skills

  • General‑purpose programming (Python, Go, Shell/Bash)
  • Automation & configuration management (Ansible, Puppet, Chef, scripts)
  • Infrastructure as Code (Terraform)
  • Containerization & orchestration (Docker, Kubernetes)
  • Cloud platforms (AWS, Azure, GCP)
  • Monitoring & logging (Prometheus, Grafana, ELK)
  • Networking fundamentals
  • Security best practices
  • CI/CD pipeline implementation
  • Version control (Git)