Ovii Job Board

Senior AI Engineer

metaforms

Bengaluru, India • Onsite - Bengaluru, India • Full-Time • 4+ years

Posted 2026-06-24 Tech & Engg

Apply on employer site

Job Description

Ovii's Interpretation of the Role

Metaforms seeks a senior AI engineer to design, build, and continuously improve production‑grade AI agents that power large‑scale market‑research surveys. The role blends applied LLM work with systems engineering, demanding high ownership of agent harnesses, evaluation pipelines, and reliability tooling.

Role Snapshot

  • Senior AI Engineer
  • On‑site in Bengaluru
  • Full‑time
  • High‑ownership AI systems
  • Production LLM/agent pipelines
  • Python‑centric development

Must-Have Requirements

  • Python
  • Frontier model APIs (Anthropic, OpenAI, Gemini)
  • Context engineering
  • Debugging complex non‑deterministic failures
  • production AI agent systems
  • LLM/agent pipelines

Nice-to-Have Signals

  • LLM observability and eval tooling (Braintrust, Langfuse, LangSmith, Weave, promptfoo)
  • Semantic parsing / DSLs
  • Computer‑use or browser agents
  • Human‑in‑the‑loop workflows
  • Go
  • TypeScript
  • LLM observability tooling
  • semantic parsing and DSLs
  • computer‑use agents

Work Setup

  • Location: Bengaluru, India
  • Work mode: ONSITE
  • Employment type: Full-Time

Not Specified in JD

  • Visa sponsorship
  • Salary range
  • Remote eligibility
  • Coding test
  • Portfolio

What You'll Likely Work On

  • Own and evolve the agent harness that orchestrates planning, tool use, and self‑repair
  • Research and implement long‑context handling and cascading‑error mitigation in multi‑step pipelines
  • Define structured evaluation rubrics and build continuous monitoring, tracing, and failure‑mode analysis for agents in production
  • Create tooling for domain experts to update skill files, eval sets, and knowledge bases
  • Develop regression suites, golden datasets, and LLM‑as‑judge pipelines to catch regressions before deployment
  • Design human‑in‑the‑loop review workflows for high‑stakes survey outputs

Good Fit If You Have

  • Enjoys fast‑paced environments with zero‑bureaucracy
  • Thrives when decisions have direct architectural impact
  • Comfortable collaborating with domain experts and leadership

Skills

  • Python
  • Frontier LLM APIs (Anthropic, OpenAI, Gemini)
  • Context engineering
  • Debugging non‑deterministic systems
  • Go or TypeScript (plus)
  • LLM observability & eval tooling
  • Semantic parsing / DSLs
  • Computer‑use / browser agents
  • Human‑in‑the‑loop workflows