
Senior Site Reliability Engineer (Night Shift)
resilic
Posted 2026-07-07
Tech & Engg
Job Description
Ovii's Interpretation of the Role
We need a senior SRE to own night‑shift reliability for a globally‑used supply‑chain platform. You’ll design, automate, and operate Azure‑based, Kubernetes‑driven services while reducing incident MTTR. The role is fully remote from India and works closely with development teams to ensure high availability and security.
Role Snapshot
- Senior Site Reliability Engineer
- Night‑shift coverage for US business hours
- Remote – India based
- Azure & Kubernetes focus
- Automation & incident ownership
- Large‑scale distributed systems
Must-Have Requirements
- Azure Cloud services
- Kubernetes
- Docker
- Linux system administration
- GitHub Actions
- Helm Charts
- Grafana
- Kafka
- Redis
- PostgreSQL
- Hadoop
- Cloudflare
- SRE/DevOps
- Azure Cloud
Nice-to-Have Signals
- Terraform
- Ansible
- Databricks
- Clickhouse
- MLOps
- large‑scale distributed systems
- infrastructure automation
Work Setup
- Location: India, India
- Work mode: REMOTE
- Remote scope: COUNTRY_RESTRICTED
- Remote countries: India
- Employment type: Full-Time
- Shift: Night shift
Eligibility Gates
- Visa sponsorship: unknown
Not Specified in JD
- Salary range
- Visa sponsorship
- Security clearance
- Notice period
- Travel requirement
- Coding test
What You'll Likely Work On
- Design, implement, and manage scalable Azure services with high availability
- Operate and optimize Kubernetes clusters and container workloads
- Build and maintain CI/CD pipelines using GitHub Actions
- Automate infrastructure deployment with Helm charts and IaC tools
- Configure Cloudflare for performance, security, and traffic routing
- Set up monitoring, alerting, and observability with Grafana
- Perform root‑cause analysis and drive preventive automation
- Collaborate with developers to improve reliability, security, and deployment practices
Good Fit If You Have
- Experience with large‑scale distributed systems
- Familiarity with Terraform or Ansible for automation
- Exposure to Databricks, Clickhouse, or MLOps
- Strong problem‑solving and troubleshooting abilities
Skills
- Azure Cloud services
- Kubernetes & Docker
- CI/CD (GitHub Actions)
- Infrastructure as code (Helm, Terraform)
- Observability (Grafana, Cloudflare)
- Distributed messaging (Kafka, Redis)
- Databases (PostgreSQL, Hadoop)
- Linux system administration
Remote Eligibility
- India