
Lead Site Reliability Engineer - India
Juniper Square
Posted 2026-04-27
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Lead Site Reliability Engineer driving reliability, scalability, and security for Juniper Square's cloud platform. Own architecture, incident response, and AI‑enhanced observability while partnering with product and engineering leadership.
Role Snapshot
- Own reliability posture and SLO/SLAs
- Define infrastructure architecture and design patterns
- Lead incident response and on‑call rotation
- Drive medium‑to‑large SRE projects as DRI
- Partner with engineering and product managers on roadmap
- Mentor junior and mid‑level engineers
- Integrate AI/LLM tooling into operational workflows
- Enforce cloud security and compliance
Must-Have Requirements
- AWS cloud services
- Kubernetes
- Terraform / CDK / CloudFormation
- Linux production environments
- PostgreSQL
- Python or Go programming
- Observability tooling
- CI/CD pipelines
- Cloud security (IAM, secrets, network)
- AWS Well‑Architected Framework
- Site Reliability Engineering
- AWS cloud
- Linux production
- Infrastructure-as-Code
Nice-to-Have Signals
- AI/LLM integration for reliability
- Service mesh & API gateway patterns
- Microservices architecture
- GitOps best practices
- AI workflow automation (alert triage, runbook automation)
- AI/LLM integration
- service mesh
Work Setup
- Location: India
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility details
- Education requirement
- Certifications
What You'll Likely Work On
- Set technical direction and architecture for infrastructure systems, balancing reliability, scalability, and cost.
- Design and review distributed systems, expand design patterns, and reduce technical debt.
- Establish and enforce security best practices across team‑owned services.
- Own reliability posture: define SLOs, monitor SLAs, and hold teams accountable.
- Lead complex incident response, root‑cause analysis, and preventive measures.
- Define logging, monitoring, and operational standards; drive observability adoption.
- Act as Directly Responsible Individual for multi‑team SRE projects, managing scope and risk.
- Collaborate with engineering managers and product managers to scope and deliver roadmap initiatives.
Good Fit If You Have
- Experience applying AI/LLM tools to improve reliability and incident response.
- Familiarity with service mesh, API gateways, and microservices patterns.
- Proven ability to lead technical discussions and influence cross‑functional stakeholders.
Skills
- AWS cloud services
- Kubernetes
- Infrastructure‑as‑Code (Terraform, CDK, CloudFormation)
- Linux production environments
- PostgreSQL performance & recovery
- Python / Go scripting
- Observability (metrics, logging, tracing)
- CI/CD pipelines
- Cloud security (IAM, secrets, network)
- AI/LLM reliability tooling (preferred)
- Service mesh & API gateway patterns (preferred)
- Microservices architecture (preferred)