
Job Description
Ovii's Interpretation of the Role
Senior Infrastructure Engineer building and operating a multi‑tenant, cloud‑native platform for an enterprise AI product. Owns end‑to‑end design, reliability, observability, and DevOps practices across GCP, Azure, and AWS.
Role Snapshot
- Design & evolve multi‑tenant Kubernetes platforms
- Build core services in Go & Python
- Define reliability contracts (SLIs/SLOs, error budgets)
- Drive IaC, GitOps, and CI/CD pipelines
- Own observability stack and incident response
Must-Have Requirements
- Golang
- Python
- Docker
- Kubernetes
- Terraform
- Helm
- gRPC
- protobuf
- Cloud provider (GCP/Azure/AWS)
- Distributed systems fundamentals
- Database expertise (NoSQL, graph)
- CI/CD pipelines
- Platform engineering
- Infrastructure design
- Backend services
- Cloud platforms
- Distributed systems
- Bachelor's degree in Computer Science or related field
Nice-to-Have Signals
- Multi‑cloud experience
- Service mesh (Istio/Linkerd)
- Observability tools (Prometheus, Grafana, OpenTelemetry)
- GitOps (ArgoCD/Flux)
- Autoscaling (KEDA, HPA, VPA)
- Vector databases (pgvector, Pinecone, Milvus)
- Graph databases (Neo4j, Neptune)
- Auth & security (Vault, mTLS, RBAC, OIDC/SAML)
- Open‑source contributions
- Message queues (Kafka, Pulsar, NATS, PubSub)
- Multi‑cloud deployments
- Security & auth
- Vector/graph databases
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Relocation
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Architect scalable, multi‑tenant microservices on Kubernetes across GCP, Azure, and AWS
- Implement core platform components in Go and Python with strict latency and throughput SLOs
- Manage service‑to‑service communication using gRPC, protobuf, and service‑mesh technologies
- Define and monitor reliability contracts, error budgets, and autoscaling policies
- Build and maintain observability pipelines with Prometheus, Grafana, and OpenTelemetry
- Drive IaC (Terraform), Helm charts, and GitOps workflows for consistent deployments
- Profile performance, conduct load testing, and optimize cost per request
- Participate in on‑call rotations, lead incident response, and perform root‑cause analysis
Good Fit If You Have
- Enjoys deep dives into service‑mesh internals and database internals
- Thrives in a fast‑moving, AI‑driven product environment
- Has a track record of building platforms that other teams adopt
Skills
- Go & Python development
- Kubernetes & Docker
- Terraform & Helm
- gRPC / protobuf APIs
- Multi‑cloud (GCP, Azure, AWS)
- Distributed systems fundamentals
- Database design (NoSQL, graph)
- Observability (Prometheus, Grafana, OpenTelemetry)
- GitOps (ArgoCD / Flux)
- CI/CD pipelines