
Senior AI/ML Solution Architect - Generative AI & Agentic Systems
Flentas
Posted 2026-08-06
Tech & Engg
Job Description
Ovii's Interpretation of the Role
Flentas seeks a senior AI/ML Solution Architect to design and deliver enterprise‑scale generative AI and agentic systems. The role blends deep LLM/SLM expertise with solution‑architecture leadership across cloud, edge and on‑prem environments.
Role Snapshot
- Design enterprise‑scale generative AI solutions
- Architect agentic systems using LLMs and SLMs
- Lead integration via APIs, event‑driven pipelines and Model Context Protocol
- Optimize model performance, cost and latency across cloud, edge and on‑prem
- Define fine‑tuning, compression and quantization strategies
- Collaborate with product and engineering stakeholders
Must-Have Requirements
- LLMs (GPT‑4, Claude, LLaMA)
- SLMs (Phi‑3, Gemma, TinyLlama)
- Agent frameworks (LangChain, LangGraph, Semantic Kernel, Agno)
- Retrieval‑Augmented Generation
- Fine‑tuning techniques (LoRA, QLoRA, DoRA, prompt‑tuning)
- Model compression & quantization
- Cloud AI platforms (AWS, GCP, Azure)
- API development (REST, gRPC, GraphQL)
- Event‑driven architecture (webhooks, message buses)
- Security & authentication (SSO/OIDC)
- ML frameworks (TensorFlow, PyTorch, Hugging Face Transformers)
- Containerization & orchestration (Docker, Kubernetes)
- Technology and software development
- Generative AI (LLMs/SLMs)
- Enterprise solution architecture
Nice-to-Have Signals
- Master's or PhD in Computer Science, AI, ML or related field
- Published research or open‑source contributions
- Multi‑modal or cross‑modal model experience
- MLOps and model lifecycle management
- Regulatory compliance knowledge (GDPR, AI Act)
- Cloud AI certifications (AWS/GCP/Azure)
- Federated learning experience
- Few‑shot / zero‑shot learning techniques
- Academic research or open‑source contributions
- Multi‑modal AI
- Regulatory compliance
- Master's or PhD in Computer Science, AI, Machine Learning or related field
- AWS/GCP/Azure AI/ML certifications
Work Setup
- Location: Pune, India
- Work mode: ONSITE
- Employment type: Full-Time
Not Specified in JD
- Visa sponsorship
- Salary range
- Remote eligibility
- Certifications
- Travel
- Coding test
- Portfolio
- Cover letter
What You'll Likely Work On
- Design and architect scalable agentic solutions leveraging advanced LLM capabilities
- Implement Model Context Protocol integrations to connect AI models with enterprise services
- Build multi‑agent orchestration and context‑memory management for complex workflows
- Develop and optimize Retrieval‑Augmented Generation pipelines for fast knowledge access
- Create end‑to‑end fine‑tuning, compression and quantization workflows for both LLMs and SLMs
- Engineer cloud‑native inference pipelines and edge/on‑prem deployment strategies
- Design secure, standards‑based APIs (REST/gRPC/GraphQL) and event‑driven architectures
- Establish model selection, evaluation and cost‑optimization frameworks for large‑scale AI deployments
Good Fit If You Have
- Proven record delivering AI solutions at enterprise scale
- Strong communication and stakeholder‑management abilities
- Experience navigating AI regulatory compliance (GDPR, AI Act)
- Ability to evaluate and select appropriate models for varied workloads
- Comfort working across cross‑functional teams and rapid‑iteration environments
Skills
- LLMs (GPT‑4, Claude, LLaMA)
- SLMs (Phi‑3, Gemma, TinyLlama)
- Agent frameworks (LangChain, LangGraph, Semantic Kernel, Agno)
- Retrieval‑Augmented Generation (RAG)
- Fine‑tuning & adaptation (LoRA, QLoRA, DoRA, prompt‑tuning)
- Model compression & quantization (pruning, INT8/INT4, distillation)
- Cloud AI platforms (AWS, GCP, Azure)
- API & event‑driven integration (REST, gRPC, GraphQL, webhooks)
- Security & auth (SSO/OIDC)
- ML frameworks (TensorFlow, PyTorch, Hugging Face Transformers)
- Containerization & orchestration (Docker, Kubernetes)