Skip to main content
Blog

RAG Mechanics: How Retrieval-Augmented Generation Actually Works

Last updated AI

Retrieval-augmented generation is usually described in one sentence: look things up, then ask the model to answer using what you found. That sentence is accurate and it hides the entire problem, which is that "look things up" contains at least six decisions, each of which can quietly cap the quality of every answer the system will ever produce. A model given the wrong passages will write a fluent, confident, wrong response — and no improvement to the model fixes it.

This guide takes the pipeline apart. What happens to a document before it can be retrieved, what an embedding actually is, why pure vector search fails on the queries that matter most, what reranking does and why it produces the largest single improvement, and how to tell which stage is responsible when an answer is wrong.


What you will learn
  • The full pipeline, ingestion through generation, and what each stage decides
  • Chunking strategies and why they matter more than they appear to
  • What embeddings are, in plain terms, and what they cannot capture
  • Why hybrid search beats pure semantic search on real queries
  • Reranking: the highest-return stage and the most commonly omitted
  • How to diagnose which stage caused a bad answer
In this article
  1. The problem it solves
  2. The pipeline in full
  3. Ingestion and extraction
  4. Chunking
  5. Embeddings
  6. The vector index
  7. Why pure vector search is not enough
  8. Query processing
  9. Reranking
  10. Context packing
  11. Generation and citation
  12. Permissions and multi-tenancy
  13. Freshness and updates
  14. Evaluating each stage
  15. Diagnosing a wrong answer
  16. Advanced patterns
  17. Twelve mistakes
  18. A worked example: one question, traced
  19. Frequently asked questions

1. The problem it solves

A language model knows what was in its training data, up to a cut-off, with no notion of which parts are true and no ability to cite anything. For questions about your policies, your products, your customers or anything that changed last week, that is useless — and worse than useless, because the model will produce a plausible answer anyway.

Retrieval augmentation changes the task from recall to summarisation. Instead of "what do you know about our refund policy", the model is asked "here are three passages from the refund policy; answer using only these". Models are dramatically more reliable at the second task than the first.

Three properties follow, and they are the reason this pattern dominates enterprise applications:

  • Currency. Update a document and the answer changes immediately. No retraining.
  • Attribution. Every claim can point at a source, which makes fabrication detectable rather than merely regrettable.
  • Access control. Retrieval can be filtered by permission, so a user only sees answers derived from material they could read anyway.

The governing constraint is equally simple: the answer can only be as good as the passages retrieved. Every improvement effort should start by asking whether the right material was found, because if it was not, nothing downstream can recover.

2. The pipeline in full

Two paths, running at different times.

Ingestion happens continuously in the background: documents are collected, their text extracted, split into chunks, converted to embeddings, and written into an index alongside their metadata.

Query happens per request: the question is processed and possibly rewritten, candidates are retrieved by several methods, filtered by permission, reranked against the actual question, packed into the context budget, and sent to the model with instructions to answer from them and cite.

Each stage constrains the next. A chunk that split a table in half cannot be retrieved usefully no matter how good the search is. A retrieval that misses the right passage cannot be rescued by reranking. Context that is packed carelessly buries the relevant passage where the model attends to it least. Understanding this ordering is what makes the system debuggable.

3. Ingestion and extraction

The least glamorous stage and the one that most often silently caps quality.

Documents arrive in formats designed for humans reading them, not for machines indexing them. A PDF is a description of where marks appear on a page — it has no concept of a paragraph, a heading or a table. Extracting usable text from one is a genuine problem, and doing it badly produces chunks that are grammatically broken, with headers interleaved into body text and tables flattened into unreadable sequences of numbers.

What to preserve during extraction, because it is expensive to recover later:

  • Structure. Which heading a passage sits under is enormously informative and is the first thing naive extraction discards.
  • Tables. Flattened tables are worse than useless — they retrieve on their numbers and mean nothing out of context. Convert them to a readable form or keep them whole.
  • Metadata. Source, author, date, document type, and any classification. All of it becomes filter criteria later.
  • Position. Page and section, so a citation can point somewhere a person can verify.

Two practical notes. Scanned documents need optical character recognition, and its error rate propagates into everything downstream — a misrecognised policy number is a wrong answer waiting to happen. And content written for browsing frequently needs restructuring to be retrievable at all: a page whose meaning depends on a heading four sections above will not survive chunking, and the fix is editorial rather than technical.

4. Chunking

Splitting documents into retrievable pieces, and the decision with the widest quality impact per unit of effort.

The tension is straightforward. Small chunks are precise — a retrieved passage is mostly relevant — but may lack the context needed to be understood alone. Large chunks carry context and dilute relevance, because the embedding represents an average of several topics and matches nothing strongly.

StrategyHow it splitsSuits
Fixed sizeEvery N characters, ignoring meaningNothing, really — the baseline everyone starts with
Recursive by separatorParagraphs, then sentences, then charactersGeneral prose; a reasonable default
StructuralBy heading, section or listDocumentation and policies — usually the best choice
SemanticWhere the topic shifts, detected by embedding similarityUnstructured long-form text
Whole documentNo splittingShort documents where splitting loses more than it gains

Three techniques improve almost any strategy:

Overlap. Adjacent chunks sharing some text prevents a sentence spanning a boundary from being lost to both. Modest overlap costs storage and prevents a real failure.

Contextual prefixing. Prepend the document title and the heading path to each chunk before embedding. A chunk reading "must be submitted within 30 days" is meaningless alone; the same chunk prefixed with "Refunds › Time limits" is retrievable and interpretable. This is a small change with a disproportionate effect.

Parent-child retrieval. Embed small chunks for precise matching, but return the larger surrounding section for context. You get the retrieval accuracy of small chunks and the comprehensibility of large ones.

The honest guidance on size: there is no universal answer, and the common defaults are a starting point rather than a recommendation. Structure your chunks around the document's own boundaries where they exist, and evaluate rather than assume.

5. Embeddings

An embedding is a list of numbers representing a piece of text's meaning, produced by a model trained so that texts about similar things end up with similar lists. Similarity is measured by the angle between them, which is why two passages saying the same thing in different words score highly.

Three properties matter in practice:

They capture topic better than they capture precision. "The refund window is 30 days" and "the refund window is 14 days" are nearly identical embeddings, because they are about the same thing. This is the fundamental limitation and the reason exact matching still matters.

The model must match at both ends. Documents and queries must be embedded by the same model. This sounds obvious and is violated regularly — usually when a model is upgraded and only new documents are re-embedded, which silently degrades everything.

Changing the model means re-embedding everything. Vectors from different models are not comparable at all. This is the single most operationally expensive fact about embeddings, and it should be planned for rather than discovered.

On choosing one: multilingual models matter if your content or your users are. Domain-specific models can outperform general ones on specialised vocabulary. Larger dimensions capture more nuance at higher storage and search cost. As always, measure on your own content — public benchmark rankings correlate loosely with performance on a specific corpus.

6. The vector index

Comparing a query against every stored vector is exact and impractically slow past a modest scale. Approximate nearest neighbour indexes trade a small amount of recall for an enormous speed improvement, and the trade is almost always worth making.

Two structures dominate. Graph-based indexes connect vectors to their neighbours and traverse towards the query, giving excellent recall and speed at high memory cost. Partition-based indexes cluster vectors and search only the nearest clusters, using less memory with somewhat lower recall.

The parameters worth understanding are all versions of the same trade: search harder, get better recall, take longer. Defaults are usually reasonable, and tuning them is worth doing only after establishing that recall is your problem.

Metadata filtering is the feature that matters most in production and is worth checking carefully when choosing a store. Filtering after retrieval means requesting many more candidates than you need and hoping enough survive — which fails unpredictably for narrow filters. Filtering during the search is correct and not universally supported.

On whether you need a dedicated vector database: for most workloads, a relational database with a vector extension or a search engine with vector support is entirely sufficient and spares you a new system with its own backup, access control and operational story. Move to a dedicated store when scale, filtering complexity or latency genuinely demand it.

7. Why pure vector search is not enough

Semantic search excels at paraphrase and fails at precision, which is exactly backwards for a large class of real queries.

The failures are specific and predictable:

  • Identifiers. A part number, an error code, an invoice reference. Semantically these are near-meaningless, and the search returns things about the general topic instead of the specific item.
  • Rare terms. A product name or an internal acronym the embedding model never saw during training.
  • Negation and small distinctions. "Refunds are permitted" and "refunds are not permitted" embed almost identically.
  • Numbers. Thresholds, dates and amounts are poorly represented.

Keyword search handles all four well and fails at paraphrase, which is what semantic search handles. The answer is to run both and combine — hybrid search — which is the single most reliable improvement available to a struggling retrieval pipeline.

Combining results from two rankings is usually done by fusing the ranks rather than the scores, since the two systems produce scores on incomparable scales. The result is a ranking where a document strong in either method surfaces, and one strong in both surfaces higher.

8. Query processing

The question a user types is frequently not the question that should be searched, and a small amount of processing pays for itself.

Rewriting for self-containment. In a conversation, "what about for business accounts?" means nothing alone. Resolving it against the conversation into a standalone question is essential for multi-turn systems and is a common omission.

Expansion. Adding synonyms and expanded acronyms improves keyword recall, particularly for internal vocabulary that differs from what customers say.

Decomposition. A question containing two questions retrieves poorly for both. Splitting, retrieving separately and combining handles compound queries that otherwise fail silently.

Hypothetical answer generation. Generating a plausible answer and embedding that rather than the question sometimes retrieves better, because answers resemble documents more than questions do. Worth testing; not universally better.

Filter extraction. Pulling structured constraints out of natural language — a date range, a product, a region — and applying them as metadata filters rather than hoping the semantic search handles them.

Each of these costs latency, and none should be added without measuring whether it helps on your evaluation set. Query rewriting for conversation is close to mandatory; the others are situational.

9. Reranking

The stage that consistently produces the largest accuracy improvement per unit of effort, and the one most commonly missing from first implementations.

The reason it works is structural. Retrieval embeds documents and queries separately — a document's vector is computed without any knowledge of the question — which is what makes searching millions of documents fast. A reranker looks at the query and a candidate together and scores their actual relationship, which is far more accurate and far too slow to run over a whole corpus.

So the pipeline uses both: retrieve fifty candidates quickly with a method that is fast and imprecise, then rerank them carefully and keep the best five. The expensive operation runs fifty times rather than a million.

Two practical points. The improvement is largest when initial retrieval is mediocre, which is most of the time — reranking rescues the case where the right passage was retrieved at position thirty and would otherwise never have reached the model. And the number of candidates to rerank is a genuine trade: more candidates means better recall and more latency, and the useful range is usually smaller than intuition suggests.

10. Context packing

The final retrieval decision: which passages go into the prompt, in what order, and with what surrounding information.

Order matters. Models attend more strongly to the beginning and end of a long context than to the middle. Placing the most relevant passages at the edges and less relevant ones in the middle measurably improves answers, and is a free improvement.

Deduplicate. Overlapping chunks and near-identical passages from different documents waste budget and can make the model over-weight a repeated claim.

Attach identifiers. Each passage needs a reference the model can cite and your code can verify. This is what makes citation checking possible.

Include structural context. The document title and section heading alongside each passage, so the model can distinguish a policy from an example.

Budget deliberately. More context is not better: it costs money, adds latency, and dilutes attention. Five well-chosen passages routinely outperform twenty mediocre ones, and the twenty cost four times as much.

11. Generation and citation

The instruction that most improves reliability is the one telling the model what to do when the passages do not contain the answer. Without it, a model will produce something plausible, because that is what it was trained to do. Instructed to state explicitly that the provided material does not answer the question, it produces a signal your code and your interface can act on.

Require citations at the claim level, referencing the identifiers you supplied. Then verify them in code: every cited identifier must appear in the passages actually retrieved for that request. A citation to something not in the context is a fabrication, and catching it is a trivial check that eliminates the most damaging failure mode.

Return the grounded status to the caller, so the interface can decide whether to show the answer confidently, hedge it, or escalate. An honest "I could not verify this, here are the two most relevant documents" is more useful than a confident paragraph, and users trust a system that admits uncertainty considerably more than one that never does.

12. Permissions and multi-tenancy

The failure with the worst consequences, and it is entirely preventable.

The rule: filter by permission during retrieval, in code, not by instructing the model. A prompt asking the model not to reveal certain documents is not access control — it is a suggestion, and it fails against a determined user and occasionally against an ordinary one.

Implementation requires that every chunk carries the access metadata of its source document, that permission filters are applied as part of the search rather than after it, and that the filter is derived from the authenticated caller's actual entitlements at query time rather than from anything the client supplied.

Two related traps. Caches must be keyed by entitlement, or a cached answer computed for one user is served to another — this is a genuine cross-tenant leak with a mundane cause. And permissions change, so an index built when someone had access must be filtered by current permissions rather than by what was true at ingestion.

Worth testing explicitly: an automated probe that attempts to retrieve another tenant's content on every deployment. It is a small test suite and it covers the failure nobody can afford.

13. Freshness and updates

Documents change, and a system answering from stale material is confidently wrong in a way that is hard to notice.

The mechanics: detect changes at the source, re-extract and re-chunk the changed document, re-embed the affected chunks, and replace them in the index. Deletion matters as much as addition — a removed document whose chunks remain indexed will be cited in answers, and a policy that was withdrawn is exactly the kind of thing that must not be.

Two operational practices worth adopting. Store a document version with each chunk, so you can identify and remove stale material and explain which version an answer was based on. And expose document age in the interface, so a user can see that an answer draws on something from two years ago.

Re-embedding the whole corpus is occasionally necessary — when the embedding model changes, or when chunking strategy changes. Plan for it: it is expensive, it takes time, and doing it partially leaves an index containing incomparable vectors, which degrades everything silently.

14. Evaluating each stage

Evaluate retrieval separately from generation, because they fail differently and mixing them makes both undiagnosable.

StageMetricQuestion it answers
RetrievalRecall at kWas the right passage retrieved at all?
RankingPosition of the correct passageDid it reach the model?
RerankingImprovement in positionIs this stage earning its latency?
GenerationFaithfulness to the passagesDid it use what it was given?
GenerationAnswer correctnessIs the answer actually right?
End to endGrounded rateWhat share of answers are fully traceable?

Recall at k is the most important number in the system, because it is the ceiling. If the correct passage is not in the top k, no model can produce a correct answer, and every effort spent on prompting is wasted. Measuring it requires an evaluation set of questions with the passages that should answer them — which is tedious to build and is the highest-value asset the project will produce.

15. Diagnosing a wrong answer

A structured sequence, because guessing at this is expensive.

  1. Is the information in the corpus at all? Search directly. Surprisingly often the answer is no, and the fix is content rather than engineering.
  2. Was it retrieved? Look at what came back. If the right passage is absent, the problem is chunking, embedding or search — and nothing downstream matters.
  3. Where did it rank? If it was retrieved at position forty and you keep five, the problem is ranking, and reranking is the answer.
  4. Was it in the final context? If it ranked well and was cut, the problem is packing or budget.
  5. Did the model use it? If the right passage was present and the answer ignored it, the problem is generation — prompt, position in context, or a conflicting passage.
  6. Is the answer wrong but grounded? Then the source document is wrong, which is a content problem and a genuinely useful discovery.

The distribution of causes in real systems is heavily weighted towards the first three. Teams that begin by adjusting prompts are usually working on the last stage of a problem that occurred in the first.

16. Advanced patterns

Worth knowing, and worth adding only after the basics are measured.

Multi-hop retrieval for questions requiring information from several documents in sequence: retrieve, generate an intermediate question, retrieve again. Powerful and it multiplies latency and cost, so bound the hops.

Agentic retrieval where the model decides what to search for and can search repeatedly. More capable on complex questions, considerably harder to evaluate and to bound.

Graph-augmented retrieval, which builds a graph of entities and relationships alongside the text and traverses it. Genuinely better for questions about relationships between things; substantial additional machinery.

Summary indexes, where document-level summaries are indexed alongside chunks so a broad question retrieves an overview rather than a fragment.

Long-context alternatives. As context windows grow, putting whole documents in becomes possible — and it is expensive, slower, and attention across very long contexts remains imperfect. Retrieval and long context are complementary: retrieve well, then use the larger budget to include more of what you found.

17. Twelve mistakes

  1. Fixed-size chunking that ignores structure. Split tables and orphaned sentences.
  2. No contextual prefix on chunks. Passages that are meaningless without their heading.
  3. Pure vector search. Fails on identifiers, rare terms, negation and numbers.
  4. No reranking. The largest available improvement, omitted.
  5. Different embedding models for documents and queries. Silent, total degradation.
  6. Permission filtering after retrieval, or in the prompt. Not access control.
  7. Cache keys without tenant and entitlement. Cross-tenant leakage with a mundane cause.
  8. No citation verification. Fabricated references pass through unnoticed.
  9. Too much context. Dilutes attention, costs money, buries the relevant passage.
  10. Deletions not propagated to the index. Withdrawn policies still cited.
  11. Evaluating end to end only. No way to know which stage failed.
  12. Tuning prompts to fix a retrieval problem. Working on the last stage of a first-stage failure.

18. A worked example: one question, traced

Consider a question that a support system gets wrong, and follow it through the pipeline to see where the failure actually occurred.

The question: "Can a business account get a refund on order 88213 after 40 days?" The system answers that refunds are available within 30 days, which is true for consumer accounts and wrong for business accounts, which have a 60-day window under a separate policy.

Step one: is it in the corpus? Yes — there is a business account policy document containing the 60-day rule. So this is not a content gap.

Step two: was it retrieved? No. The retrieved passages are all from the consumer refunds policy. This immediately tells us the problem is in retrieval, and that no amount of prompt adjustment would have helped.

Step three: why not? Two causes, both instructive. The business policy document was chunked by fixed size, and the chunk containing "60 days" begins mid-sentence and never mentions "business account" — that phrase appears in a heading two chunks earlier. Without contextual prefixing, the chunk is semantically about time limits in general, not about business accounts. And the query embedded semantically retrieves passages about refunds and time limits, which the consumer policy expresses more directly and more often.

The order number reveals a second problem. The identifier in the question is semantically meaningless, so pure vector search ignores it entirely. Hybrid search with a keyword component would have surfaced the order record, which carries the account type — the single piece of information that determines the correct answer.

The fixes are structural, not clever. Chunk by heading so each passage retains its section context. Prefix every chunk with the document title and heading path, so "Business Accounts › Refunds › Time limits" is part of what gets embedded. Add keyword search alongside semantic search. Extract the order identifier as a filter and retrieve the order record, so the account type becomes a metadata filter on the policy search.

After the changes, the correct passage retrieves at position three. Reranking moves it to position one. The answer becomes correct, and it cites the business policy document with a page reference an agent can verify in one click.

The instructive part is the ordering. Four changes, all in the first half of the pipeline, none of them touching the prompt or the model. A team that had started by rewriting the prompt would have spent a week producing marginal changes to how a wrong passage was summarised.

19. Frequently asked questions

What chunk size should we use?

There is no universal answer, and the common defaults are a starting point rather than a recommendation. Split on the document's own structure where it exists — headings and sections — rather than on a character count. Use overlap, prefix chunks with their heading path, and evaluate recall on your own questions. The structural choice matters considerably more than the size.

Do we need a dedicated vector database?

Usually not at first. A relational database with a vector extension, or a search engine with vector support, handles millions of chunks comfortably and spares you a new system to operate. It also frequently gives you hybrid search in one place, which is an advantage. Move to a dedicated store when scale, filtering complexity or latency genuinely demand it.

How many passages should we retrieve?

Retrieve broadly, then narrow. Something like fifty candidates from search, reranked down to three to five for the context. The final number should be small: five well-chosen passages routinely outperform twenty mediocre ones and cost a quarter as much. Tune both numbers against your evaluation set rather than adopting a default.

Is reranking worth the latency?

Almost always. It typically produces the largest single accuracy improvement in the pipeline, at a cost of tens to low hundreds of milliseconds. If you add one thing to a struggling system, add this. Measure the position improvement it produces — if it is zero, your initial retrieval is already excellent, which would be unusual.

How do we handle documents that change frequently?

Detect changes at the source, re-process only the affected documents, and ensure deletions remove chunks from the index. Store a document version with each chunk so you can identify stale material and explain which version an answer used. Expose document age in the interface, because a user seeing that an answer draws on a two-year-old document can judge it appropriately.

Will larger context windows make retrieval unnecessary?

No, for three reasons: cost scales with tokens, latency scales with tokens, and attention across very long contexts remains imperfect so material in the middle is under-weighted. Larger windows make retrieval more forgiving — you can include more of what you found — but selecting the right material remains the thing that determines answer quality.

How do we stop it citing documents a user cannot see?

Filter by permission during the search, using the authenticated caller's current entitlements, with access metadata stored on every chunk. Never instruct the model to withhold documents — that is a suggestion, not a control. Key any cache by tenant and entitlement, and add an automated test that attempts cross-tenant retrieval on every deployment.

Where should we start when quality is poor?

Measure recall at k on a set of real questions with their correct passages. That single number tells you whether the problem is retrieval or generation, and it is almost always retrieval. Then add hybrid search and reranking, which together resolve a large share of retrieval problems. Prompt work comes last, because it can only affect how well the model uses what it was already given.

Key takeaways

  • Retrieval quality caps answer quality. Measure recall first; it is the ceiling on everything else.
  • Chunk on structure, not size, and prefix each chunk with its heading path.
  • Hybrid search, always. Semantic search fails on identifiers, rare terms, negation and numbers.
  • Reranking is the highest-return stage and the most commonly omitted.
  • Filter permissions in retrieval, in code. A prompt instruction is not access control.
  • Verify citations against the retrieved context. A two-line check that catches fabrication.

The pipeline is not complicated, but it has many places to be quietly wrong, and almost all of them are upstream of the model. Diagnose in pipeline order, fix the earliest failing stage, and measure each one separately — which turns a system that produces confident nonsense into one whose every answer can be traced back to a document somebody wrote.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch