Ovii Job Board

Site Reliability Engineer

Ema

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 4+ years

Posted 2026-06-10 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

As a Site Reliability Engineer at Ema, you’ll keep the Agentic AI platform highly available, targeting 99.9%+ uptime for enterprise customers. You’ll design cloud infrastructure, automate deployments, and lead incident response across GCP, Azure, and AWS environments.

Role Snapshot

  • Own platform reliability and uptime
  • Design and provision cloud infrastructure
  • Automate CI/CD and deployment pipelines
  • Monitor, alert, and resolve production incidents
  • Document runbooks and knowledge transfer
  • Collaborate with engineering and customer stakeholders

Must-Have Requirements

  • Cloud platforms (GCP, Azure, AWS)
  • Infrastructure-as-Code (Terraform, Ansible)
  • CI/CD pipelines (GitLab CI, Jenkins)
  • Observability tooling (Prometheus, Grafana, Datadog, Splunk)
  • Incident troubleshooting and root‑cause analysis
  • DevOps
  • Infrastructure
  • Deployment Engineering

Nice-to-Have Signals

  • Containerization & orchestration (Docker, Kubernetes)
  • Fast‑paced startup experience
  • fast‑paced startup
  • containerization

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Design and provision secure, scalable cloud infrastructure on GCP, Azure, or AWS for customer environments
  • Build and maintain IaC using Terraform or Ansible and automate end‑to‑end deployment workflows
  • Operate and improve CI/CD pipelines (GitLab CI, Jenkins) to enable reliable SaaS releases
  • Monitor logs, metrics, and alerts; maintain dashboards and ensure SLA‑level uptime
  • Respond to incidents, perform root‑cause analysis, and implement permanent fixes
  • Create and maintain runbooks, troubleshooting guides, and knowledge‑transfer documentation
  • Partner with engineering and customer stakeholders to communicate system health and reliability status

Good Fit If You Have

  • Experience in fast‑paced, high‑growth startup environments
  • Hands‑on with containerization and orchestration tools such as Docker and Kubernetes
  • Strong communication skills for interacting with engineering teams and customers

Skills

  • Cloud platforms (GCP, Azure, AWS)
  • Infrastructure-as-Code (Terraform, Ansible)
  • CI/CD pipelines (GitLab CI, Jenkins)
  • Observability tools (Prometheus, Grafana, Datadog, Splunk)
  • Incident troubleshooting and root‑cause analysis
  • Containerization & orchestration (Docker, Kubernetes)