Ovii Job Board

Senior Infrastructure Engineer

Ema

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5+ years

Posted 2026-06-19 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Senior Infrastructure Engineer building and operating a multi‑tenant, cloud‑native platform for an enterprise AI product. Owns end‑to‑end design, reliability, observability, and DevOps practices across GCP, Azure, and AWS.

Role Snapshot

  • Design & evolve multi‑tenant Kubernetes platforms
  • Build core services in Go & Python
  • Define reliability contracts (SLIs/SLOs, error budgets)
  • Drive IaC, GitOps, and CI/CD pipelines
  • Own observability stack and incident response

Must-Have Requirements

  • Golang
  • Python
  • Docker
  • Kubernetes
  • Terraform
  • Helm
  • gRPC
  • protobuf
  • Cloud provider (GCP/Azure/AWS)
  • Distributed systems fundamentals
  • Database expertise (NoSQL, graph)
  • CI/CD pipelines
  • Platform engineering
  • Infrastructure design
  • Backend services
  • Cloud platforms
  • Distributed systems
  • Bachelor's degree in Computer Science or related field

Nice-to-Have Signals

  • Multi‑cloud experience
  • Service mesh (Istio/Linkerd)
  • Observability tools (Prometheus, Grafana, OpenTelemetry)
  • GitOps (ArgoCD/Flux)
  • Autoscaling (KEDA, HPA, VPA)
  • Vector databases (pgvector, Pinecone, Milvus)
  • Graph databases (Neo4j, Neptune)
  • Auth & security (Vault, mTLS, RBAC, OIDC/SAML)
  • Open‑source contributions
  • Message queues (Kafka, Pulsar, NATS, PubSub)
  • Multi‑cloud deployments
  • Security & auth
  • Vector/graph databases

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Relocation
  • Travel
  • Security clearance
  • Coding test

What You'll Likely Work On

  • Architect scalable, multi‑tenant microservices on Kubernetes across GCP, Azure, and AWS
  • Implement core platform components in Go and Python with strict latency and throughput SLOs
  • Manage service‑to‑service communication using gRPC, protobuf, and service‑mesh technologies
  • Define and monitor reliability contracts, error budgets, and autoscaling policies
  • Build and maintain observability pipelines with Prometheus, Grafana, and OpenTelemetry
  • Drive IaC (Terraform), Helm charts, and GitOps workflows for consistent deployments
  • Profile performance, conduct load testing, and optimize cost per request
  • Participate in on‑call rotations, lead incident response, and perform root‑cause analysis

Good Fit If You Have

  • Enjoys deep dives into service‑mesh internals and database internals
  • Thrives in a fast‑moving, AI‑driven product environment
  • Has a track record of building platforms that other teams adopt

Skills

  • Go & Python development
  • Kubernetes & Docker
  • Terraform & Helm
  • gRPC / protobuf APIs
  • Multi‑cloud (GCP, Azure, AWS)
  • Distributed systems fundamentals
  • Database design (NoSQL, graph)
  • Observability (Prometheus, Grafana, OpenTelemetry)
  • GitOps (ArgoCD / Flux)
  • CI/CD pipelines