Ovii Job Board

Senior Software Development Test Engineer

Tekion

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 5+ years

Posted 2026-07-17 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

The Senior Software Development Test Engineer will design and operate automated quality frameworks for generative AI and large‑language‑model products on Tekion's cloud‑native automotive platform. The role blends AI data validation, LLM evaluation, and MLOps CI/CD integration to ensure reliable, bias‑free AI services.

Role Snapshot

  • AI/GenAI test automation
  • LLM evaluation & hallucination detection
  • Data quality & statistical validation
  • MLOps CI/CD integration
  • Python & SQL expertise
  • Vector DB and embedding testing
  • Performance & load testing

Must-Have Requirements

  • Python
  • SQL
  • Pytest
  • RAGAS/DeepEval/Promptflow
  • LangChain/LangSmith/LlamaIndex
  • OpenAI/Anthropic/HuggingFace APIs
  • Vector DB testing
  • Pandas/NumPy
  • Docker
  • GitHub Actions/Jenkins
  • Grafana/Kibana/OpenTelemetry
  • MLflow
  • Software Development Engineer in Test
  • AI/GenAI testing

Nice-to-Have Signals

  • AWS Bedrock/Azure OpenAI/GCP Vertex AI
  • Kubeflow/Weights & Biases/Feast
  • Scikit‑learn/TensorFlow/PyTorch
  • Terraform
  • Playwright/Cypress
  • Locust/JMeter
  • Statistical hypothesis testing
  • Synthetic data generation
  • Cloud AI service testing
  • MLOps platform experience

Work Setup

  • Location: Bangalore, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Education requirement
  • Certifications
  • Relocation
  • Notice period
  • Travel
  • Security clearance
  • Coding test
  • Portfolio
  • GitHub
  • Writing sample
  • Cover letter

What You'll Likely Work On

  • Build automated suites to detect hallucinations, bias, toxicity, and prompt‑injection in LLM‑powered products
  • Implement RAG evaluation pipelines measuring relevance, groundedness, and answer faithfulness
  • Design test beds for multi‑agent workflows, tool‑calling accuracy, and autonomous decision loops
  • Create scripted and synthetic conversation simulations to stress‑test multi‑turn dialog flows
  • Develop prompt regression frameworks to monitor output consistency across model changes
  • Statistically validate AI data outputs, monitor precision/recall, and analyze error patterns
  • Audit data ingestion and feature‑store pipelines for schema drift and corruption
  • Validate vector DB indexing, embedding similarity, and retrieval latency
  • Maintain automated ML metric suites (Precision, Recall, F1, ROC‑AUC) across model versions
  • Integrate AI quality checks into MLOps pipelines so failures block releases

Good Fit If You Have

  • Experience with cloud AI services such as AWS Bedrock, Azure OpenAI, or GCP Vertex AI
  • Familiarity with MLOps platforms like Kubeflow, Weights & Biases, or Feast
  • Knowledge of ML frameworks (Scikit‑learn, TensorFlow, PyTorch)
  • Comfort with infrastructure‑as‑code tools (Docker, Kubernetes, Terraform)
  • Background in UI automation (Playwright or Cypress) or performance testing (Locust, JMeter)

Skills

  • Python (expert)
  • SQL
  • Pytest
  • LLM evaluation frameworks (RAGAS, DeepEval, Promptflow)
  • Agent workflow tools (LangChain, LangSmith, LlamaIndex)
  • LLM endpoint APIs (OpenAI, Anthropic, HuggingFace)
  • Vector database testing
  • Data analysis libraries (Pandas, NumPy)
  • Docker & containerization
  • CI/CD automation (GitHub Actions, Jenkins)
  • Observability (Grafana, Kibana, OpenTelemetry)
  • MLflow