
Job Description
Ovii's Interpretation of the Role
We are seeking a senior Technical Architect to design and lead an AI agent testing and evaluation platform for a large enterprise client. The role combines deep AI/ML expertise with client‑facing technical leadership and end‑to‑end system architecture.
Role Snapshot
- Lead AI agent testing platform design
- Define evaluation methodology and safety rules
- Integrate testing framework with CI/CD on Google Cloud
- Guide and review distributed delivery team
- Client‑facing technical leadership
Must-Have Requirements
- Google Cloud AI stack (Vertex AI, Gemini, Agent Development Kit, Dialogflow CX)
- LLM evaluation, prompt engineering, agent architectures
- Software engineering / machine learning (8+ years)
- software engineering
- machine learning
- LLM/AI agent system development
Nice-to-Have Signals
- LLM‑as‑judge evaluation
- RAG systems
- Agent observability
- Production AI systems in regulated or large‑scale consumer environment
- agent observability
- production AI in regulated environment
Work Setup
- Location: US
- Work mode: ONSITE
- Remote scope: UNSPECIFIED
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Certifications
- Relocation
- Travel
- Security clearance
- Coding test
What You'll Likely Work On
- Own the end‑to‑end technical design of an agent testing framework covering deterministic checks, statistical LLM‑based evaluation, and adversarial testing.
- Define datasets, safety rules, and calibration methods for AI agent evaluation.
- Design integration with the client’s CI/CD pipeline and cloud environment to enforce quality gates.
- Lead technical discussions with client architects and engineering leadership, establishing architectural standards.
- Guide, mentor, and review work of a distributed delivery team.
Good Fit If You Have
- Proven ability to lead client‑facing technical engagements.
- Experience guiding distributed engineering teams.
- Familiarity with AI safety rules and evaluation dataset design.
- Exposure to RAG systems and agent observability (preferred).
Skills
- Google Cloud AI stack (Vertex AI, Gemini, Dialogflow CX)
- LLM evaluation & prompt engineering
- AI agent architecture
- CI/CD pipeline integration
- Deterministic, statistical & adversarial testing
- Production AI systems (regulated/large‑scale) – preferred
- RAG systems – preferred
- Agent observability – preferred