
Job Description
Ovii's Interpretation of the Role
Staff Software Engineer – AI Engineer at Tekion builds and operates the LLM control plane and AI platform that powers real‑time dealer intelligence. The role combines large‑scale data, MLOps, knowledge‑graph retrieval and safety‑first AI services for the Automotive Retail and Enterprise clouds.
Role Snapshot
- Build LLM control plane & gateway
- Design unified APIs & SDKs
- Implement safety, privacy, and PII redaction
- Create MLOps pipelines and model monitoring
- Scale graph‑based knowledge retrieval
- Mentor engineers and champion AI safety
Must-Have Requirements
- Python
- Java/Scala/Go
- Docker
- Kubernetes
- Airflow
- Kubeflow
- MLflow
- CI/CD for models
- Spark/Flink/Kafka
- Knowledge graph (Neo4j/Neptune/TigerGraph)
- GraphQL
- Vector search (pgvector/Qdrant/Milvus)
- REST/gRPC
- Observability (traces, logs, metrics)
- LLM integration
- Agentic systems
- MLOps
- large‑scale data/ML platforms
- distributed systems
- LLM control plane
- knowledge graphs
Nice-to-Have Signals
- AWS
- platform‑as‑product mindset
- AI safety
- cost‑aware engineering
Work Setup
- Location: Bangalore, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Education requirement
- Coding test
What You'll Likely Work On
- Build and run the LLM control plane/gateway with routing, rate‑limits, failover, and cost tracking
- Ship a unified REST/gRPC API and SDKs with caching, schema normalization, and full observability
- Enforce safety and privacy defaults, including content filtering, prompt validation, and PII redaction
- Enable multi‑model, multi‑vendor LLM usage with automated canarying, versioning, and quota management
- Own the agent runtime: tool registry, permissions, function calling, grounding, and retrieval
- Design orchestration patterns and manage long‑running agent workflows
- Develop MLOps pipelines for classical and deep models, standardize experiment tracking, and monitor drift
- Evolve the domain graph, build reliable ingestion pipelines, and serve real‑time context to agents
Good Fit If You Have
- Platform‑as‑product mindset focused on developer experience and SLAs
- System‑thinking with observability, fallback, and access‑control baked in
- Passion for AI safety and cost‑aware engineering
Skills
- Python
- Java/Scala/Go
- Docker & Kubernetes
- AWS (preferred)
- Airflow / Kubeflow
- MLflow
- Spark / Flink / Kafka
- Knowledge graphs (Neo4j, Neptune, TigerGraph)
- Vector search (pgvector, Qdrant, Milvus)
- REST & gRPC APIs
- Observability (traces, logs, metrics)
- LLM integration & agentic systems