
A support lead at a mid-size lender asks the company's new AI assistant a simple question: what is the prepayment penalty on a loan signed before March? The assistant answers instantly, in confident, well-written prose. It is also wrong. It quoted the current policy, not the one that applied to that loan, because the retrieval step pulled the most recently indexed document instead of the right one. Nobody notices until a customer disputes the figure.
That failure has nothing to do with the language model. The model did exactly what it was asked: write a fluent answer from the context it was handed. The context was wrong. This is the pattern behind most disappointing retrieval-augmented generation projects, and it is why 2026 has become the year enterprises started rebuilding the RAG systems they rushed into production a year or two earlier.
This guide covers what changed, what a production-grade RAG application looks like now, where agentic retrieval helps and where it just burns money, what's different about building for India and Gurgaon (including Hindi and code-mixed content), and how to scope a first build. If you want the short version of what Akoode builds in this space, the RAG application development page has it.
The clearest signal comes from a VentureBeat Pulse survey of enterprise respondents published in April 2026. Between the end of 2025 and Q1 2026, enterprise intent to adopt hybrid retrieval — combining dense vector search with keyword search and a reranking layer — tripled, from 10.3% to 33.3% of respondents in a single quarter. In the same survey, 22% of qualified enterprise respondents reported having no production RAG system at all. VentureBeat's framing was blunt: the RAG architecture most enterprises built to scale is not the one they expect to be running by year-end, because first-generation designs hit a wall once agentic workloads and real access-control requirements arrived.
That sits inside a wider honesty gap the industry keeps rediscovering. McKinsey's 2025 global survey found 71% of organizations using generative AI regularly, while only 17% attributed more than 5% of EBIT to it. Demos are cheap. Production value is not. And in RAG specifically, the distance between the two is usually the retrieval layer.
A word on how to read the numbers you'll see circulating. Many figures about RAG accuracy come from vendor blogs and practitioner write-ups, and they depend heavily on the baseline being compared against. Industry analysis commonly cites naive retrieve-then-generate pipelines failing at the retrieval step roughly 40% of the time in production. Practitioner benchmarks commonly report hybrid retrieval improving recall by 15-30% on enterprise corpora. Reported results for agentic and graph-based pipelines are larger still: a May 2026 MLOps Community benchmark across 47 production deployments, as relayed by secondary sources, put the hallucination reduction from agentic RAG paired with a knowledge graph at roughly 62% versus naive setups, and a June 2026 preprint reported hallucinations falling from 14.1% to 4.9% on a 9,000-question financial-compliance set at about 220 milliseconds of added latency. Treat these as directional, not guarantees. They are good evidence of where the technique helps, and weak evidence of what it will do on your documents. The only number that matters is the one you measure on your own questions, which is a point we'll return to.
Retrieval-augmented generation makes an AI system look something up in your documents or data before answering, instead of answering from whatever the model absorbed in training. It is the difference between a colleague who answers from memory and one who opens the file first.
It is not a vector database bolted onto a chatbot. A working RAG application has an ingestion pipeline that decides how your documents are broken up, an index that decides what can be found, a retrieval step that decides what the model actually sees, a generation step that decides what to do with it (including what to say when nothing relevant turns up), and an evaluation loop that tells you how often all of that is right. Get the retrieval step wrong and the whole system inherits the hallucination problem, with extra infrastructure.
It also isn't a replacement for fine-tuning in every case; the two solve different problems. Fine-tuning changes how a model writes, reasons, or formats. RAG changes what the model knows at the moment it answers. For knowledge that changes, such as policies, prices, contracts, and procedures, RAG is almost always the right tool, because updating the knowledge base means updating documents rather than retraining anything. For tone, structure, and domain vocabulary, fine-tuning helps, and some of the strongest deployments layer RAG over a lightly tuned model. A related and fast-moving question is whether very long context windows make retrieval unnecessary. For small, stable document sets that fit comfortably in context, sometimes yes. For large, changing, permission-sensitive collections, no: you still need to decide what gets loaded, who is allowed to see it, and how to trace an answer back to its source.
Pure vector search captures meaning but routinely misses exact strings: policy numbers, account IDs, product codes, error messages, clause references. Keyword search (BM25) captures exact terms but misses paraphrase. Production systems in 2026 combine both, often fusing the two result lists with reciprocal rank fusion, then passing the merged set to a reranker that scores each candidate against the actual question. If you built your first RAG system on vector search alone in 2024 or 2025, this is the single most common rebuild.
A modular production pipeline typically rewrites the user's query (expanding abbreviations, resolving references to earlier turns), retrieves with hybrid search, reranks, assembles the context, and only then generates. Each stage is a place where accuracy is won or lost, and each can be measured separately. That separability matters: when an answer is wrong, you want to know whether retrieval failed or generation failed, because the fixes are completely different.
In agentic RAG, the system decides how to retrieve rather than running one fixed pass. It can break a question into sub-questions, search a different source, check whether the evidence actually answers the question, and search again if it doesn't. For multi-hop questions — "which of our supplier contracts renew in Q1 and have a price-escalation clause?" — this is genuinely better. But the economics are unforgiving. Practitioner analyses suggest that something like 60-70% of production queries in a typical enterprise system are simple enough for single-hop retrieval, while agentic loops can add several times the per-query cost and push latency from a couple of seconds to ten or more. The sensible design routes simple questions down the cheap path and reserves the agentic loop for the queries that need it. Teams that run every query through an agent pay a latency and cost tax on questions a simple lookup would have handled.
A RAG system that retrieves a document the asking user isn't allowed to see has just leaked it, regardless of how polished the answer sounds. Enterprise buyers now ask directly whether retrieval respects existing document permissions, and the answer has to be architectural: access filters applied at retrieval time, against the same identity and entitlement systems the rest of the business uses, not a promise that the model "won't mention" restricted material. This is also why large, permission-sensitive document sets are where generic RAG tutorials tend to fall apart.
The shift in mature teams is measuring retrieval quality directly, before judging the final answer. Useful metrics include precision at K (are the retrieved passages relevant), provenance coverage (does every claim trace to a source), faithfulness (does the answer stay within the retrieved evidence), end-to-end latency, and hallucination rate on a fixed set of real questions. Some practitioners treat a faithfulness score above roughly 0.85 as a baseline for regulated settings; the exact threshold is a judgment call, but having a threshold at all is the point. A system with no measurement is a system whose accuracy you will learn about from a customer.
Behind those five changes is a fairly consistent architecture:
Ingestion and chunking. Documents are parsed (including the ugly ones: scanned PDFs, tables, slide decks) and split into retrievable units sized to preserve meaning. A chunking strategy that works for short FAQs will cut a long contract clause in half. Metadata, such as document version, effective date, owner, and access group, is attached here, because it is what later lets the system pick the right version of a policy rather than the newest one.
Embedding and indexing. An embedding model is chosen to fit the content and languages involved, and an index is built that supports hybrid retrieval. There is no single best embedding model; benchmark results don't transfer cleanly, and the only evaluation that counts is on your own corpus.
Retrieval and reranking. Query rewriting, hybrid search, metadata filtering (including permissions), and a reranker produce a short, high-quality context.
Generation with abstention. The model is instructed to answer only from the supplied context, to cite its sources, and, critically, to say it doesn't know when the retrieved evidence doesn't support an answer. Designing that refusal path is as important as designing the answer path.
Orchestration for harder questions. A router decides when a question needs the cheap single-pass path versus a multi-step or multi-source path, including questions that need both structured data (a database) and unstructured content (documents) in the same answer.
Evaluation and monitoring. Retrieval is logged so any wrong answer can be traced to a specific document or query pattern, and quality is re-measured as documents change, because a knowledge base that isn't reviewed drifts out of date.
Frameworks such as LlamaIndex and LangChain are widely deployed for this, with LangGraph increasingly used when the control flow is stateful, and a common production pairing uses LlamaIndex for the retrieval tooling and LangGraph for orchestration. The framework matters less than the discipline: choose based on your content and scale, not on which has the best marketing this quarter.
RAG is infrastructure rather than a product, so the useful question is which problems it unlocks:
Internal knowledge assistants that answer from SOPs, wikis, past decisions, and policy documents, replacing the shared-drive search that everyone does by hand.
Grounded customer support that answers from the actual catalogue, terms, and account data, with a citation, instead of a plausible guess that generates a follow-up ticket.
Contract and compliance Q&A where an answer without a traceable clause is unusable, which is most of finance, insurance, and the public sector.
Technician and field assistants that pull the right section of an equipment manual or service bulletin during a live issue.
Multi-source analytics questions that combine a database lookup with the documents that explain it.
The knowledge layer for AI agents. Agents that act, not just answer, need grounded access to policy and data to act correctly, which is why RAG is increasingly the foundation beneath agentic systems and the enterprise agent architecture built on top of them.
If the work you actually want automated is a specific process, such as invoice intake or claims routing, rather than question answering, that is a different scope and closer to AI automation services. If generation itself is the product, such as drafting contracts or producing marketing copy at volume, see generative AI development. Many production systems combine all three, and the retrieval layer is usually what keeps the other two honest.
The dangerous failures aren't the ones that crash; they're the ones that launch, sound fine, and are wrong more often than anyone notices. The recurring causes are consistent:
Stale or conflicting documents. The system faithfully retrieves last year's policy because nobody tagged versions or effective dates.
Chunking that destroys context. A clause is split from its exception, so the retrieved fragment is accurate and misleading at once.
No abstention. The model is never taught to say "I couldn't find that," so it fills gaps with fluent guesses.
No permission model. The system works beautifully until someone discovers it will summarise a document they shouldn't be able to open.
No evaluation set. Quality is judged by whoever tried it last, with whatever question came to mind.
Agentic everywhere. Every query is routed through a multi-step agent, raising cost and latency without improving the answers to simple questions.
Every item on that list is an engineering decision, not a model limitation, which is the encouraging part: they can be designed out.
India adds three requirements that most globally written RAG guides skip.
Multilingual and code-mixed content is the default, not the edge case. Real Indian enterprise content mixes English with Hindi and other languages, often in the same sentence and sometimes in Roman script. Retrieval quality on Indic text is a genuine research area, with dedicated benchmarks such as Hindi-BEIR for Hindi retrieval and ongoing work on code-mixed Hindi-English evaluation, which tells you the problem is real and not yet solved by default settings. On the engineering side, open-weight multilingual embedding models such as BGE-M3, which supports dense, sparse, and multi-vector retrieval across 100+ languages under a permissive license, give teams a strong starting point, and translation models covering India's 22 scheduled languages (IndicTrans2 and Sarvam-Translate among them) make cross-lingual retrieval practical. The practical advice is the same as everywhere but matters more here: build a test set from your real queries in your real languages, including the code-mixed ones, and measure retrieval on it before choosing a model.
Data protection shapes the architecture. For India-based deployments, the Digital Personal Data Protection Act's obligations, which are phasing in through 2027, mean questions about where documents and embeddings live, who can retrieve them, and how access is logged should be answered at design time. Keeping the vector store and source documents within your own infrastructure or an Indian region is a common requirement and a reasonable default for sensitive content.
The ecosystem is moving in your favor. The IndiaAI Mission has built national compute capacity past 34,000 GPUs, and Haryana's Global Artificial Intelligence Centre, a pillar of the state's AI Mission, is being established in Gurugram. For a RAG builder, the local relevance is practical: Gurgaon is a dense cluster of enterprises, consulting firms, global capability centres, and BPOs, all of them document-heavy businesses where an assistant that answers accurately from internal policy, contracts, and case history has an immediate use. Akoode's own team sits in Gurugram, in Sector 49, and works across India, the UK, and the USA, so the same-time-zone, DPDP-aware build is the default rather than an add-on.
For international deployments the questions shift but the architecture does not. Under GDPR, the location and lawful basis for processing personal data inside your index matter as much as the answers it produces. The EU AI Act's transparency and traceability expectations make source citation and retrieval logging more than a nicety for higher-risk uses. US sector rules (HIPAA for health data, financial recordkeeping and model-risk expectations for banks) point to the same underlying requirement: be able to show what the system retrieved and why. Teams serving several regions often need region-pinned indexes so that a European user's query never retrieves from, or logs into, an index hosted elsewhere.
Not every organization needs the full stack on day one. A sensible sequence:
Start with a data audit, not a model. Find out what documents actually exist, how current they are, which versions conflict, and who is allowed to see what. Retrieval quality is capped by the quality of what's being retrieved from.
Build the evaluation set first. Collect fifty to a few hundred real questions with known good answers and source passages. Everything else gets measured against it.
Ship hybrid retrieval with reranking and abstention. This is the minimum viable production design in 2026.
Add permissions before you add users. Retrofitting access control onto a live system that has already indexed sensitive content is far harder than designing it in.
Add agentic routing only where the evaluation set shows single-pass retrieval failing. Let measured failures, not enthusiasm, justify the extra cost and latency.
Monitor and review on a cadence. Documents change and retrieval quality drifts; plan the review before launch.
Akoode's published guidance for a focused RAG application on a well-organised document set is roughly five to eight weeks from data audit to deployment, and ten to fourteen weeks where the content is messier or permission-aware retrieval across multiple sources is required. Both depend heavily on how much cleanup the documents need.
Akoode hasn't published a case study built under the RAG label, and it would be misleading to relabel existing work. What does transfer is engineering discipline on the parts of RAG that fail most often. The AI-powered quantity-takeoff platform built for Qualis Construction runs entirely offline with zero cloud dependency, because the client's drawings were too sensitive to leave their environment, which is the same constraint that drives many on-premise and in-region retrieval deployments. The AI-powered advertisement catalogue generator is a multi-stage pipeline where deterministic, reviewable output was non-negotiable, the same mindset a trustworthy generation step needs. Neither is a RAG system; both show the habits RAG depends on. If a RAG-specific reference matters to your decision, ask for a live walkthrough rather than relying on a case study from an adjacent problem.
A few questions separate an engineering team that understands retrieval from one that wires a vector store to a chatbot:
Do they ask to see your documents and your real questions before proposing an architecture?
Can they explain how they'll measure retrieval accuracy separately from answer quality, and on what test set?
Do they build hybrid search and reranking by default, and can they say why for your content type?
How does the system behave when nothing relevant is retrieved? Ask to see it decline to answer.
Is permission-aware retrieval part of the first build, tied to your existing identity and access systems?
How do they handle document versions and effective dates, so the system doesn't confidently cite a superseded policy?
Can they say when agentic retrieval is worth its cost for your use case, and when it isn't?
For India, can they show retrieval working on Hindi and code-mixed queries, not just English?
RAG did not stop working; the first-generation version of it stopped being enough. The teams getting reliable value in 2026 are the ones who treat retrieval as its own engineering discipline: hybrid search, reranking, abstention, permissions, and a measured evaluation loop, with agentic routing added deliberately rather than by default.
Akoode builds RAG applications, pipelines, and enterprise retrieval systems for teams in India and abroad, with a 4.9 Google rating from 110+ reviews and a 97% client retention rate. If you already have an LLM feature that sounds right and isn't, we'll audit the retrieval layer and tell you what's broken before you rebuild anything. Book time with Akhil Verma, Founder & CEO of Akoode, and bring a short brief or a handful of the questions your system keeps getting wrong.
If you're mapping where retrieval fits in a wider AI roadmap, our guides on AI agents for finance and banking, AI agents for healthcare, AI agents for insurance, and AI agents for public sector and government cover the sectors where grounded, traceable answers matter most, and how to build an AI app covers the broader build path. For custom agent work in the region, see custom AI agent development in Gurgaon, and for the wider capability set, the artificial intelligence and AI agent development service pages.
What is RAG application development? RAG application development is building an AI system that retrieves relevant information from your own documents or data before generating an answer, so responses are grounded in your actual business knowledge rather than a model's general training. It covers the ingestion pipeline, retrieval layer, generation behavior, permissions, and evaluation, not just the chat interface.
What is the difference between RAG and fine-tuning? Fine-tuning changes the model itself, which helps with tone, format, and domain vocabulary but is expensive to repeat whenever your information changes. RAG leaves the model as is and retrieves current information at answer time, so updating knowledge means updating documents. For changing facts such as policies, prices, and contracts, RAG is usually the better fit, and many strong systems combine the two.
Do long context windows make RAG unnecessary? For small, stable document sets that fit comfortably in context, sometimes. For large, frequently changing, or permission-sensitive collections, retrieval is still needed to decide what gets loaded, enforce who can see what, control cost and latency, and trace each answer back to a source.
What is hybrid search in RAG, and why does it matter? Hybrid search combines vector search, which captures meaning, with keyword search such as BM25, which captures exact terms like account numbers and product codes, usually followed by a reranker. It has become the production baseline because each method misses what the other finds, and practitioner benchmarks commonly report recall gains in the 15-30% range on enterprise content.
When should I use agentic RAG? Use it for multi-step or multi-source questions where single-pass retrieval measurably fails. Practitioner analyses suggest most production queries are simple enough for single-hop retrieval, and agentic loops add cost and latency, so a router that sends only the harder queries down the agentic path is usually the better design.
How do I reduce hallucinations in a RAG system? Improve retrieval first with hybrid search, reranking, and clean, versioned documents. Then require citations, design the model to abstain when evidence is missing, add a grading step for high-stakes answers, and measure faithfulness and hallucination rate continuously on a fixed set of real questions. Reported reductions vary widely with the baseline, so measure on your own data.
Can RAG work with Hindi and other Indian languages? Yes, but it needs deliberate engineering. Multilingual embedding models such as BGE-M3 and translation models covering India's 22 scheduled languages make it practical, and code-mixed Hindi-English content should be part of your evaluation set from the start rather than discovered after launch.
How secure is a RAG system? Security depends on design: permission-aware retrieval tied to your existing access controls, in-region or on-premise hosting of the index where required, and logging of what was retrieved for whom. For India-based deployments, DPDP Act obligations should shape where documents and embeddings are stored.
Subscribe to the Akoode newsletter for carefully curated insights on AI, digital intelligence, and real-world innovation. Just perspectives that help you think, plan, and build better.