
Customer Success Engineer (HQ2)
Blitzy
Posted 2026-06-30
INR 3,000,000 - INR 5,000,000 per year
Customer Success & Delivery
Job Description
Ovii's Interpretation of the Role
Blitzy seeks a technically‑savvy Customer Success Engineer to own end‑to‑end platform deployments, upgrades, and day‑to‑day operations for enterprise clients. You will troubleshoot distributed systems on Kubernetes, Docker, and major clouds, turning incidents into clear runbooks and proactive monitoring. The role is onsite in Pune and aligns with US business hours, including on‑call rotations.
Role Snapshot
- L2 technical support engineer
- Kubernetes & Docker expertise
- Multi‑cloud (GCP, AWS, Azure) operations
- Distributed‑systems debugging
- Observability & monitoring (Datadog or equivalent)
- Customer incident escalation & communication
- Runbook and dashboard creation
- On‑call coverage for US time zones
Must-Have Requirements
- Kubernetes
- Docker
- Major cloud provider (GCP/AWS/Azure)
- Distributed systems debugging
- Observability stack (Datadog or equivalent)
- Kubernetes & Docker operations
- Multi‑cloud platform management
- Observability and monitoring
- On‑site work in Pune, India
- Support US business hours
- Participate in on‑call rotations
Nice-to-Have Signals
- Python
- Redis
- Message queueing (Redis/rq)
- Networking & WebSockets
- PostgreSQL
- GitHub / Azure DevOps / GitLab
- CI/CD pipelines
- Helm
- ArgoCD
- Vault
- Linux / Windows OS
- Incident management (Jira)
- Prior SRE/on‑call experience
- Python scripting
- Redis and message‑queue handling
- CI/CD and Helm deployments
- Secrets management with Vault
- Customer‑facing support or SRE experience
Work Setup
- Location: Pune, India
- Work mode: ONSITE
- Timezone: Must align with US business hours (ET–PT)
- Employment type: Full-Time
- Shift: US business hours with on‑call rotation
Not Specified in JD
- Visa sponsorship
- Remote eligibility
What You'll Likely Work On
- Deploy and install the platform into customer environments and resolve installation issues
- Support ongoing upgrades and keep customer environments stable
- Triage and resolve customer‑reported incidents alongside L1, escalating when needed
- Diagnose failures across compute, networking, storage, and services using read‑only diagnostics
- Reproduce issues safely in multi‑tenant environments before making changes
- Build and maintain dashboards, alerts, and runbooks to accelerate issue resolution
- Write evidence‑backed escalation notes and post‑incident reports
- Communicate status and resolutions to customers clearly and on time
- Participate in a rotating on‑call schedule covering US business hours
Good Fit If You Have
- Prior SRE or on‑call experience
- Comfort handling secrets and credentials (Vault preferred)
- Experience with incident‑management tools such as Jira
- Methodical, evidence‑first troubleshooting mindset
- Ability to work US business hours from Pune
Skills
- Kubernetes
- Docker
- GCP / AWS / Azure
- Distributed systems debugging
- Observability stack (Datadog or equivalent)
- Python
- Redis
- Message queueing (Redis/rq)
- Networking & WebSockets
- PostgreSQL
- GitHub / Azure DevOps / GitLab
- CI/CD pipelines