RAG: retrieval augmented generation explained
Retrieval augmented generation (RAG) is a method where relevant text is looked up from an outside source and placed in the prompt so the model can use it when answering.
What RAG is
A language model only knows what was in its training data and what is in the current prompt. RAG adds a step before generation: search a collection of documents for passages related to the question, then include the best ones in the prompt. The model answers using that text. It is how assistants answer questions about your files, and in roleplay it is how a frontend recalls old chat or setting documents.
How it works
- Chunking. Documents are split into pieces, typically a few hundred tokens each, sometimes with overlap so sentences are not cut in half.
- Indexing. Each chunk is embedded into a vector by an embedding model, and often also indexed for keyword search.
- Retrieval. The query is embedded the same way and the nearest chunks are found. Hybrid systems combine vector and keyword search; some add a reranking model to reorder results.
- Generation. The top chunks are inserted into the prompt with instructions to use them, and the chat model writes the answer.
RAG in roleplay
- Vector memory retrieves old chat messages.
- SillyTavern's Data Bank lets you attach documents (setting bibles, rules, character histories) that are chunked and retrieved.
- A lorebook is a simpler cousin: retrieval by keyword rather than embedding.
RAG suits material that is too large to send every turn but only partly relevant at any moment, such as a long campaign setting.
Common mistakes
- Bad chunks. Chunks that are too small lose context; too large and retrieval gets vague and expensive.
- Trusting retrieval blindly. If the right chunk is not retrieved, the model answers from memory or invents details, often confidently.
- Too many chunks. Stuffing ten loosely related passages in the prompt distracts the model.
- Mismatched embeddings. Indexing with one embedding model and querying with another returns nonsense.
RAG versus a long context
With 512K to 1M token context windows, as on Wild West API models, you can sometimes skip RAG and send whole documents. That is simpler and avoids retrieval misses, but you pay for every token on every request, and recall of details deep in a very long prompt is not perfect. Prompt caching, where it applies, can reduce the cost of a large fixed document at the start of the prompt. For very large or frequently changing collections, RAG is still the practical option. A model with tool calling can also run retrieval itself by calling a search function you provide.
FAQ
What is the difference between RAG and fine-tuning?
RAG supplies information in the prompt at request time. Fine-tuning changes the model's weights through training. RAG is easier to update and shows the source text.
Does RAG need a vector database?
Not always. Small collections can be searched in memory, and keyword search alone works for some uses. Vector databases help at larger scale.