
Job Description
Ovii's Interpretation of the Role
As a Site Reliability Engineer at Ema, you’ll keep the Agentic AI platform highly available, targeting 99.9%+ uptime for enterprise customers. You’ll design cloud infrastructure, automate deployments, and lead incident response across GCP, Azure, and AWS environments.
Role Snapshot
- Own platform reliability and uptime
- Design and provision cloud infrastructure
- Automate CI/CD and deployment pipelines
- Monitor, alert, and resolve production incidents
- Document runbooks and knowledge transfer
- Collaborate with engineering and customer stakeholders
Must-Have Requirements
- Cloud platforms (GCP, Azure, AWS)
- Infrastructure-as-Code (Terraform, Ansible)
- CI/CD pipelines (GitLab CI, Jenkins)
- Observability tooling (Prometheus, Grafana, Datadog, Splunk)
- Incident troubleshooting and root‑cause analysis
- DevOps
- Infrastructure
- Deployment Engineering
Nice-to-Have Signals
- Containerization & orchestration (Docker, Kubernetes)
- Fast‑paced startup experience
- fast‑paced startup
- containerization
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design and provision secure, scalable cloud infrastructure on GCP, Azure, or AWS for customer environments
- Build and maintain IaC using Terraform or Ansible and automate end‑to‑end deployment workflows
- Operate and improve CI/CD pipelines (GitLab CI, Jenkins) to enable reliable SaaS releases
- Monitor logs, metrics, and alerts; maintain dashboards and ensure SLA‑level uptime
- Respond to incidents, perform root‑cause analysis, and implement permanent fixes
- Create and maintain runbooks, troubleshooting guides, and knowledge‑transfer documentation
- Partner with engineering and customer stakeholders to communicate system health and reliability status
Good Fit If You Have
- Experience in fast‑paced, high‑growth startup environments
- Hands‑on with containerization and orchestration tools such as Docker and Kubernetes
- Strong communication skills for interacting with engineering teams and customers
Skills
- Cloud platforms (GCP, Azure, AWS)
- Infrastructure-as-Code (Terraform, Ansible)
- CI/CD pipelines (GitLab CI, Jenkins)
- Observability tools (Prometheus, Grafana, Datadog, Splunk)
- Incident troubleshooting and root‑cause analysis
- Containerization & orchestration (Docker, Kubernetes)