Ovii Job Board

Principal Site Reliability Engineer, Google Cloud

Saviynt

Atlanta, United States • Hybrid - Atlanta, United States • Full-Time • 1+ years

Posted 2026-07-01 USD 240,000 - USD 250,000 per year Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Principal Site Reliability Engineer will own reliability strategy for Saviynt's SaaS platform, building shared infrastructure, tooling, and observability across multi‑cloud environments. This hands‑on role requires deep expertise in Kubernetes, Go, and Google Cloud, with a focus on scalable, fault‑tolerant services.

Role Snapshot

  • Principal SRE – reliability strategy
  • Build shared platform services
  • Kubernetes & multi‑cloud infrastructure
  • Automation & tooling in Go
  • Observability & monitoring pipelines

Must-Have Requirements

  • Go programming
  • Python programming
  • Kubernetes production expertise
  • Google Cloud Platform (GCP)
  • GitLab CI
  • Event‑driven architecture (Kafka, Pub/Sub)
  • Observability tools (Prometheus, Grafana, ELK, Datadog)
  • Service Mesh (Istio, Envoy)
  • RESTful API design
  • Relational databases (MySQL, PostgreSQL)
  • Principal SRE responsibilities
  • building tools and services for other engineers
  • Bachelor's degree in Computer Science, Engineering, or related field
  • Advanced Professional Google Cloud Platform Certification
  • Advanced Professional GCP Certification required
  • Bachelor's degree or equivalent required

Nice-to-Have Signals

  • Multi‑cloud experience (AWS, Azure)
  • Building abstractions over multiple clouds
  • Familiarity with NATS or RabbitMQ
  • multi‑cloud platform abstractions

Work Setup

  • Location: Atlanta, United States
  • Work mode: HYBRID
  • Remote scope: UNSPECIFIED
  • Employment type: Full-Time

Eligibility Gates

  • Visa sponsorship: unknown

Not Specified in JD

  • Visa sponsorship
  • Remote eligibility
  • Relocation
  • Travel
  • Security clearance

What You'll Likely Work On

  • Design and operate highly available Kubernetes platforms as a service for internal teams
  • Create reusable, scalable infrastructure components and automation tools using Go
  • Build and maintain event‑driven messaging services (Kafka, Google Pub/Sub) and shared data platforms
  • Develop and run CI/CD pipelines (GitLab CI, ArgoCD) to standardize deployments
  • Implement observability, monitoring, and service‑mesh solutions for global, multi‑region clouds
  • Define clear RESTful APIs for internal infrastructure services
  • Collaborate with product teams to translate their needs into reliable platform services
  • Participate in on‑call rotations supporting the shared infrastructure

Good Fit If You Have

  • Enjoy solving hard reliability problems at scale
  • Strong customer‑centric mindset for internal developer users
  • Excellent communication of complex technical concepts

Skills

  • Go & Python programming
  • Kubernetes (production)
  • Google Cloud Platform (GCP)
  • Multi‑cloud (AWS, Azure) abstractions
  • Event‑driven architecture (Kafka, Pub/Sub)
  • GitLab CI / ArgoCD pipelines
  • Observability (Prometheus, Grafana, ELK, Datadog)
  • Service Mesh (Istio, Envoy)
  • RESTful API design
  • Relational DB (MySQL, PostgreSQL)