Ovii Job Board

Staff Software Development Test Engineer - AI Evaluation

Tekion

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5-8 years

Posted 2026-07-09 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

We are hiring a Staff Software Development Test Engineer to own the AI evaluation platform for Tekion's automotive AI agents. You will design datasets, build automated scoring pipelines, and define quality metrics that ensure trustworthy AI across multiple product lines.

Role Snapshot

  • Own AI evaluation platform as a shared service
  • Design and curate evaluation datasets
  • Build automated scoring pipelines (LLM‑as‑judge, rubric‑based)
  • Define accuracy, relevance, safety, and other quality metrics
  • Integrate evaluation gates into CI/CD
  • Create dashboards for AI quality visibility
  • Collaborate with ML, data science, and product teams

Must-Have Requirements

  • Python
  • LLM evaluation methods
  • Dataset curation
  • Statistical analysis
  • CI/CD integration
  • SDET
  • quality engineering
  • ML engineering
  • data science

Nice-to-Have Signals

  • Eval frameworks/tools (Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM)
  • Online evaluation and production model monitoring
  • Experiment tracking (MLflow, Weights & Biases)
  • Responsible AI / safety evaluation familiarity
  • Red‑teaming experience
  • experience with eval frameworks/tools

Work Setup

  • Location: Bangalore, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Salary range
  • Visa sponsorship
  • Remote eligibility
  • Education requirement
  • Certifications

What You'll Likely Work On

  • Define and build a shared AI evaluation infrastructure used across ML teams
  • Create and maintain evaluation datasets and golden‑ground‑truth sets for multiple domains
  • Develop quality metrics such as accuracy, relevance, faithfulness, consistency, safety, and task success
  • Implement automated scoring pipelines, including LLM‑as‑judge, rubric‑based, and reference‑based methods
  • Detect hallucinations, bias, and edge‑case failures and design targeted eval suites
  • Run offline benchmarking and online production monitoring with A/B testing and drift detection
  • Embed evaluation gates into CI/CD to quality‑check model, prompt, or data changes before release
  • Build dashboards that surface AI quality signals for ML and product stakeholders

Good Fit If You Have

  • Passionate about trustworthy AI and responsible‑AI practices
  • Strong communicator who can translate quality metrics into product decisions
  • Hands‑on experience with LLM‑as‑judge or similar evaluation frameworks
  • Comfort building platform services that serve multiple engineering teams

Skills

  • Python
  • LLM evaluation methods
  • Dataset curation & labeling
  • Statistical analysis
  • CI/CD integration
  • Dashboarding & reporting
  • Responsible AI best practices
  • Experiment tracking (MLflow, Weights & Biases)