
Job Description
Ovii's Interpretation of the Role
Litmus seeks a Site Reliability Engineer to own the reliability, security, and performance of its Azure‑hosted platform for enterprise customers. The role spans provisioning, monitoring, incident response, and automation while maintaining a 99.9% uptime SLA.
Role Snapshot
- Site Reliability Engineer
- Azure cloud infrastructure
- Kubernetes operations
- Infrastructure‑as‑code automation
- On‑call incident response
- Customer‑facing environment
Must-Have Requirements
- Azure services (AKS, VMs, VNet, Load Balancer, Application Gateway, Azure Database for PostgreSQL/MySQL, Key Vault, Azure Monitor, Log Analytics)
- Kubernetes production workload administration
- Infrastructure‑as‑code (Terraform, Bicep, ARM templates)
- Scripting/automation with Python, Bash, PowerShell
- Networking fundamentals (DNS, TLS/SSL, load balancing, firewalls/NSGs)
- Incident management and on‑call rotation
- Security baseline implementation (vulnerability, secrets, certificate management)
- Site Reliability Engineering
- Azure cloud infrastructure
- Kubernetes production operations
- Microsoft Certified: DevOps Engineer Expert
Nice-to-Have Signals
- Identity federation / SSO protocols (OIDC, SAML, Keycloak, Okta)
- MQTT or other IoT data protocols
- Manufacturing / industrial domain experience
- Manufacturing / industrial IoT environments
- Identity federation / SSO
- MQTT protocol
- Microsoft Certified: Azure Administrator Associate
- Azure Solutions Architect Expert
- Certified Kubernetes Administrator (CKA)
- Manufacturing
- Industrial IoT
Work Setup
- Location: Pune, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Salary range
- Visa sponsorship
- Remote eligibility
- Equity
- Bonus
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Provision, configure, and maintain Azure compute, container, networking, and database services.
- Design and operate monitoring, alerting, and dashboards for system health.
- Participate in on‑call rotation, incident response, root‑cause analysis, and post‑mortem reviews.
- Implement security controls, vulnerability management, secrets handling, and certificate lifecycle.
- Configure identity/SSO integration and troubleshoot authentication issues.
- Manage high‑throughput networking and data flows across distributed sites.
- Build IaC and CI/CD pipelines to automate provisioning and deployments.
- Handle backup, disaster recovery, and periodic recovery validation.
Good Fit If You Have
- Microsoft Certified: DevOps Engineer Expert (Azure DevOps) certification.
- Experience with identity federation protocols (OIDC/SAML) and SSO platforms.
- Familiarity with MQTT or other IoT/industrial data protocols.
- Background supporting manufacturing or industrial IoT environments.
Skills
- Azure services (AKS, VMs, networking, databases)
- Kubernetes
- Terraform / Bicep / ARM templates
- Python / Bash / PowerShell scripting
- Monitoring & alerting (Azure Monitor, Log Analytics)
- Security (vulnerability, secrets, certificates)
- CI/CD pipeline automation
- Networking fundamentals (DNS, TLS/SSL, load balancing)
- Incident management & on‑call rotation
- Capacity planning
- Identity federation / SSO (OIDC, SAML, Okta, Keycloak)
- MQTT / IoT data protocols