
Job Description
Ovii's Interpretation of the Role
Ema seeks a senior Platform Engineer to own and evolve the backend infrastructure that powers its Agentic AI platform. You will design multi‑tenant, cloud‑native microservices, define reliability contracts, and drive DevOps practices across GCP, Azure, and AWS.
Role Snapshot
- Senior Platform Engineer
- Backend infrastructure ownership
- Multi‑tenant microservices design
- Cloud‑native (GCP/Azure/AWS)
- Reliability & observability focus
Must-Have Requirements
- Golang
- Python
- Docker
- Kubernetes
- Microservices architecture
- gRPC/protobuf
- Service mesh (Istio/Linkerd)
- Terraform
- Helm
- GitOps (ArgoCD/Flux)
- CI/CD pipelines
- Observability stack (Prometheus, Grafana, OpenTelemetry)
- Major cloud provider (GCP/Azure/AWS)
- NoSQL and graph databases
- Message queues (Kafka/Pulsar/NATS/PubSub)
- Platform engineering
- Infrastructure design
- Backend services
- Distributed systems
- Cloud-native microservices
- Bachelor's degree in Computer Science or related field
- 5+ years of Platform/Infrastructure/Backend experience
Nice-to-Have Signals
- Multi‑cloud experience
- Auth & security tooling (Vault, mTLS, OIDC/SAML)
- Vector databases (pgvector, Pinecone, Milvus)
- Specific graph databases (Neo4j, Neptune)
- Open‑source contributions to infrastructure projects
- High‑scale operating‑system experience (high QPS, large data volumes)
- Multi‑cloud deployments
- Auth and security engineering
- Vector and graph database work
- Open‑source infrastructure contributions
Work Setup
- Location: Bengaluru, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Equity details
- Bonus details
- Coding test
- Portfolio
What You'll Likely Work On
- Design and evolve scalable, multi‑tenant microservices on Kubernetes across GCP, Azure, and AWS.
- Build core platform components in Golang and Python, meeting latency and throughput SLOs.
- Define and maintain gRPC/protobuf contracts, service‑mesh routing, load‑balancing, retries, and circuit‑breaking.
- Make architectural trade‑offs around partitioning, consistency, caching, and build‑vs‑buy decisions.
- Establish reliability contracts: SLIs/SLOs, error budgets, capacity planning, autoscaling, and graceful degradation.
- Implement and operate the observability stack (Prometheus, Grafana, OpenTelemetry, distributed tracing).
- Drive DevOps practices using Terraform, Helm, GitOps (ArgoCD/Flux), and CI/CD pipelines.
- Profile performance, conduct load testing, and optimize cost per request.
- Participate in on‑call rotations, lead incident response, and perform root‑cause analysis.
Good Fit If You Have
- Enjoys deep technical work on service‑mesh and database internals.
- Thrives in a fast‑moving, AI‑driven product environment.
- Comfortable documenting architectural decisions and reliability contracts.
Skills
- Golang
- Python
- Docker & Kubernetes
- Microservices architecture
- gRPC / protobuf APIs
- Service mesh (Istio or Linkerd)
- Terraform IaC
- Helm & GitOps (ArgoCD / Flux)
- CI/CD pipelines
- Observability stack (Prometheus, Grafana, OpenTelemetry)
- Major cloud provider (GCP, Azure, AWS)
- NoSQL & graph databases