
Job Description
Ovii's Interpretation of the Role
Signzy seeks a Site Reliability Engineer to design, operate, and automate reliable cloud and Kubernetes systems. The role blends infrastructure provisioning, observability tooling, and incident response to keep production services stable and performant.
Role Snapshot
- Site Reliability Engineer
- Cloud & Kubernetes infrastructure
- Automation & IaC
- Monitoring & observability
- Incident response
- CI/CD pipeline support
Must-Have Requirements
- AWS / Azure / GCP
- Linux production systems
- Kubernetes
- Terraform / Terragrunt / Crossplane
- CI/CD pipeline tools (GitHub Actions, GitLab CI, Jenkins, Azure DevOps)
- Monitoring & observability tools (Prometheus, Grafana, ELK/EFK, Datadog)
- Bash or Python scripting
- Cloud networking concepts
- Relational database experience
- On‑call rotation and incident response
- Understanding of SLIs, SLOs, error budgets
- Strong problem‑solving and debugging skills
- Effective verbal and written communication
- cloud infrastructure management
- Linux production
- Kubernetes orchestration
- Infrastructure as Code
- CI/CD pipeline development
- monitoring & alerting
- on‑call incident response
Nice-to-Have Signals
- High‑availability system operation
- On‑prem or hybrid infrastructure exposure
- Data platform, analytics, or AI/ML workload experience
- Familiarity with GitOps tools (Argo CD, Flux)
- Exposure to NoSQL or other data platforms
- high‑availability systems
- on‑prem or hybrid environments
- data platforms / AI/ML workloads
- financial services
- Fintech
Work Setup
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Design, deploy, and operate scalable cloud and Kubernetes services.
- Automate infrastructure provisioning, deployments, and operational workflows with IaC tools.
- Build and maintain tooling for deployment, monitoring, and system operations.
- Monitor health and performance, proactively identifying improvement opportunities.
- Troubleshoot issues across development, test, and production environments.
- Participate in on‑call rotations, incident response, root‑cause analysis, and reliability enhancements.
- Collaborate with engineering teams to improve operability and deployment safety.
- Support large‑scale, data‑intensive or AI‑driven workloads.
Good Fit If You Have
- Experience with high‑availability or large‑scale systems.
- Exposure to on‑prem or hybrid infrastructure environments.
- Familiarity with data platforms, analytics, or AI/ML workloads.
- Strong ownership mindset for production services.
- Clear verbal and written communication during incidents.
Skills
- AWS / Azure / GCP
- Kubernetes
- Terraform / Terragrunt / Crossplane
- GitHub Actions / GitLab CI / Jenkins / Azure DevOps
- Prometheus / Grafana / Datadog / ELK
- Bash / Python scripting
- Linux production systems
- GitOps (Argo CD / Flux)
- Cloud networking & load balancing
- Relational databases