
Job Description
Ovii's Interpretation of the Role
The AI Evaluation Engineer will design and operate evaluation harnesses, test infrastructure, and release gates for AI‑driven products. Working onsite in Bangalore, you’ll ensure AI assistants behave reliably across hosts while driving continuous production quality.
Role Snapshot
- Build AI evaluation harnesses
- Own API test infrastructure
- Design host‑behavior probes
- Gate releases via UAT & regression analysis
- Drive production reliability
- On‑site Bangalore
Must-Have Requirements
- Python
- TypeScript
- Playwright
- Cypress
- API‑level testing
- SDET/QA automation background
- Ownership of test or evaluation infrastructure
- SDET/QA automation
- ML evaluation
- Must work onsite in Bangalore; remote not allowed
Nice-to-Have Signals
- Familiarity with LLM applications
- Highly autonomous workflow definition
- LLM applications
Work Setup
- Location: Bangalore, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Notice period
- Travel
- Security clearance
- Coding test
- Portfolio
- GitHub
- Writing sample
- Cover letter
What You'll Likely Work On
- Create and maintain evaluation harnesses that measure capture, retrieval, and guidance quality for AI‑facing features
- Develop and run end‑to‑end API test suites using Playwright or Cypress to support a daily release cadence
- Build scripted probes that simulate user sessions across multiple AI assistants and verify expected behavior
- Conduct user‑acceptance testing, regression analysis, and generate quality reports to gate production releases
- Define autonomous workflows and processes for evaluation and verification
- Collaborate with founders and engineers to translate ambiguous problems into concrete test solutions
Good Fit If You Have
- Enjoy turning messy problems into concrete solutions
- Thrive in fast‑moving, low‑process environments
- Passionate about production engineering (monitoring, reliability, latency)
- Highly autonomous and self‑directed
Skills
- Python/TypeScript
- Playwright/Cypress
- API‑level testing
- SDET/QA automation
- LLM evaluation
- Continuous release testing
- Quality reporting
- Production monitoring