Ovii Job Board

Senior Site Reliability Engineer (Night Shift)

resilic

India, India • Remote - India, India • Full-Time • 6-12 years

Posted 2026-07-07 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

We need a senior SRE to own night‑shift reliability for a globally‑used supply‑chain platform. You’ll design, automate, and operate Azure‑based, Kubernetes‑driven services while reducing incident MTTR. The role is fully remote from India and works closely with development teams to ensure high availability and security.

Role Snapshot

  • Senior Site Reliability Engineer
  • Night‑shift coverage for US business hours
  • Remote – India based
  • Azure & Kubernetes focus
  • Automation & incident ownership
  • Large‑scale distributed systems

Must-Have Requirements

  • Azure Cloud services
  • Kubernetes
  • Docker
  • Linux system administration
  • GitHub Actions
  • Helm Charts
  • Grafana
  • Kafka
  • Redis
  • PostgreSQL
  • Hadoop
  • Cloudflare
  • SRE/DevOps
  • Azure Cloud

Nice-to-Have Signals

  • Terraform
  • Ansible
  • Databricks
  • Clickhouse
  • MLOps
  • large‑scale distributed systems
  • infrastructure automation

Work Setup

  • Location: India, India
  • Work mode: REMOTE
  • Remote scope: COUNTRY_RESTRICTED
  • Remote countries: India
  • Employment type: Full-Time
  • Shift: Night shift

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Security clearance
  • Notice period
  • Travel requirement
  • Coding test

What You'll Likely Work On

  • Design, implement, and manage scalable Azure services with high availability
  • Operate and optimize Kubernetes clusters and container workloads
  • Build and maintain CI/CD pipelines using GitHub Actions
  • Automate infrastructure deployment with Helm charts and IaC tools
  • Configure Cloudflare for performance, security, and traffic routing
  • Set up monitoring, alerting, and observability with Grafana
  • Perform root‑cause analysis and drive preventive automation
  • Collaborate with developers to improve reliability, security, and deployment practices

Good Fit If You Have

  • Experience with large‑scale distributed systems
  • Familiarity with Terraform or Ansible for automation
  • Exposure to Databricks, Clickhouse, or MLOps
  • Strong problem‑solving and troubleshooting abilities

Skills

  • Azure Cloud services
  • Kubernetes & Docker
  • CI/CD (GitHub Actions)
  • Infrastructure as code (Helm, Terraform)
  • Observability (Grafana, Cloudflare)
  • Distributed messaging (Kafka, Redis)
  • Databases (PostgreSQL, Hadoop)
  • Linux system administration

Remote Eligibility

  • India