Job Description
Ovii's Interpretation of the Role
Zeta seeks a Principal Site Reliability Engineer to own reliability, scalability, and automation of its cloud‑native banking platforms. The role focuses on infrastructure as code, incident response, performance tuning, and mentoring within an engineering‑first culture.
Role Snapshot
- Principal SRE (individual contributor)
- 10‑15 years of SRE experience
- Onsite in Hyderabad
- Full‑time employment
- Banking technology domain
Must-Have Requirements
- Programming languages (Python, Go, Shell/Bash)
- Automation tools (Ansible, Puppet, Chef, custom scripts)
- Infrastructure as Code (Terraform)
- Containerization (Docker) and orchestration (Kubernetes)
- Cloud platforms (AWS, Azure, GCP)
- Monitoring & logging (Prometheus, Grafana, ELK)
- Networking concepts
- Security best practices
- CI/CD pipeline implementation
- Version control (Git)
- site reliability engineering
- B.Tech/M.Tech in computer science, information technology or related field
- Must be able to work onsite in Hyderabad
Nice-to-Have Signals
- Experience in a product‑focused organization
- Cloud certifications (AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, Microsoft Certified)
- product organization experience
Work Setup
- Location: Hyderabad, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility
- Certifications (required)
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design and operate scalable, highly available infrastructure for banking services
- Build automation tools and IaC pipelines to reduce manual operational work
- Lead incident response, root‑cause analysis, and post‑mortem processes
- Plan capacity and performance tuning for population‑scale systems
- Implement monitoring, alerting, and logging to ensure system health
- Collaborate with security teams to embed hardening practices
- Develop disaster recovery strategies and test recovery procedures
- Mentor junior engineers and promote SRE best practices
Good Fit If You Have
- Enjoys deep technical problem solving at scale
- Thrives in a collaborative, ownership‑driven environment
- Has product‑focused mindset and can translate business needs into reliable systems
Skills
- General‑purpose programming (Python, Go, Shell/Bash)
- Automation & configuration management (Ansible, Puppet, Chef, scripts)
- Infrastructure as Code (Terraform)
- Containerization & orchestration (Docker, Kubernetes)
- Cloud platforms (AWS, Azure, GCP)
- Monitoring & logging (Prometheus, Grafana, ELK)
- Networking fundamentals
- Security best practices
- CI/CD pipeline implementation
- Version control (Git)