Vector memory: recalling old messages by meaning
Vector memory is a technique that stores past messages as embeddings and automatically re-inserts the ones most similar to the current conversation.
What vector memory is
When a chat grows past what fits in the prompt, old messages drop out. Vector memory brings back the relevant ones. If you mention the locket from chapter one, the original scene with the locket can be retrieved and placed in the prompt again, even though it is far outside the recent history. It is retrieval augmented generation applied to your own chat.
How it works
- Embedding. Each message (or chunk of one) is passed through an embedding model, which turns text into a vector of numbers. Texts with similar meaning get vectors that point in similar directions.
- Storage. The vectors are saved in a small index alongside the chat.
- Query. Before each generation, the last few messages are embedded and compared with the stored vectors, usually by cosine similarity.
- Insertion. The top matches are inserted into the prompt at a set position, often with a short header so the model knows they are past events.
Unlike a lorebook, this matches by meaning, not keywords. "The silver pendant" can retrieve a message about "the locket".
In SillyTavern
The Vector Storage extension handles chat messages and the Data Bank (documents attached to a chat, character or globally). Key settings are the embedding source (a local model running in SillyTavern, or an external embedding API), how many messages to query with, how many to retrieve, the minimum similarity score, chunk size, and insertion position. Changing the embedding source requires re-vectorizing, since vectors from different models are not comparable.
Limits and mistakes
- Similarity is not importance. Retrieved messages are the most similar ones, not the most relevant to the plot.
- Out of order. Retrieved fragments arrive without surrounding context and can confuse the timeline.
- Noise. Too many retrieved messages, or a low score threshold, fill the prompt with loosely related text.
- Weak embedding models give poor matches on fiction.
Most people combine vector memory with summarization for the overall arc.
Vector memory with Wild West API
Vector memory runs in the frontend and needs its own embedding source. SillyTavern can use a local embedding model, so you do not need an embeddings endpoint from your chat provider. The retrieved text is sent to the chat model as ordinary prompt tokens. With 512K to 1M token context windows you may simply keep more history instead; vector memory pays off mainly on very long campaigns or to keep per-request cost down.
FAQ
Is vector memory the same as a lorebook?
No. A lorebook inserts entries when exact keywords appear. Vector memory retrieves past text by similarity of meaning, with no keywords needed.
Do I need an embeddings API for vector memory?
Not necessarily. SillyTavern can run a small local embedding model for Vector Storage.