
Job Description
Ovii's Interpretation of the Role
Senior Systems Software Engineer on the SRE team, building and shipping production code to improve reliability, performance, and observability of critical services. Collaborates with US engineering teams, profiles systems, and contributes to incident response and cost‑optimization initiatives.
Role Snapshot
- Write and ship production‑grade code
- Profile and optimize services
- Design infrastructure solutions
- Improve observability and monitoring
- Collaborate with US teams (late‑evening IST overlap)
- Participate in on‑call and incident response
- Drive cost‑optimization efforts
Must-Have Requirements
- Python
- Go
- PostgreSQL
- Docker
- Kubernetes
- Observability platforms (Datadog, CloudWatch, Prometheus)
- Strong coding and production‑code delivery
- backend development
- SRE / infrastructure
Nice-to-Have Signals
- AWS
- Terraform/Terragrunt
- Familiarity with cloud platforms (AWS preferred)
- IaC experience
- IaC
- AWS cloud
Work Setup
- Location: Bangalore, India
- Work mode: HYBRID
- Remote scope: UNSPECIFIED
- Timezone: Late evening IST overlap with US teams
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, develop, and ship code to fix bugs, boost performance, and raise reliability across multiple services
- Profile APIs under load, identify bottlenecks, and implement fixes at code, database, or configuration level
- Evaluate and document infrastructure trade‑offs, including container orchestration and cloud platform choices
- Build tooling, automation, and test frameworks for traffic simulation and regression prevention
- Enhance logs, metrics, dashboards, alerts, and tracing to improve system observability
- Define and measure SLIs/SLOs/SLAs, aligning reliability goals with business outcomes
- Contribute to incident response runbooks, on‑call rotation, and post‑mortem analysis
- Support cost‑optimization and explore new monitoring technologies
Good Fit If You Have
- Experience with IaC tools such as Terraform or Terragrunt
- Familiarity with AWS or other cloud platforms
- Comfortable working late‑evening IST hours to sync with US teams
- Strong analytical and problem‑solving mindset
- Structured execution approach for ambiguous problems
Skills
- Python / Go
- PostgreSQL
- Docker & Kubernetes
- Observability platforms (Datadog, CloudWatch, Prometheus)
- AWS (preferred)
- Terraform / Terragrunt (optional)
- SLI / SLO / SLA knowledge
- Distributed systems