The roleplay glossary.
Every setting in the connection panel and every term in the guides, explained in plain words: what it does, what it does to a story, and what to set it to.
AbliterationAbliteration is a technique that removes a language model's tendency to refuse by finding the internal direction tied to refusal and erasing it from the weights.
API keyAn API key is a secret string that identifies your account to an API and authorizes requests, so anyone holding it can use and spend your account.
Author's noteAn author's note is a short block of instructions or context that a roleplay frontend inserts near the end of the chat history, where the model pays the most attention.
Base URLA base URL is the root address of an API that a client library or app adds endpoint paths to when it sends requests.
Character cardA character card is a file, usually a PNG image with embedded JSON, that stores everything a roleplay frontend needs to play a character.
Chat completion vs text completionChat completion APIs take a list of role-tagged messages and format them server side, while text completion APIs take one raw prompt string that you format yourself.
Context windowThe context window is the maximum number of tokens a model can take into account in one request, covering both the input and the reply.
DRY samplerDRY (Don't Repeat Yourself) is a sampler that penalizes a token when choosing it would continue a sequence of text that already appeared earlier in the context.
Example dialogueExample dialogue is a set of sample exchanges in a character card that shows the model how the character talks and how replies should be formatted.
First messageThe first message, or greeting, is the opening reply from the character that starts every new chat and sets the style the model will imitate.
Frequency penaltyFrequency penalty is an OpenAI API setting that lowers a token's score in proportion to how many times it has already appeared.
Instruct templateAn instruct template is the pattern of special tokens and markers that wraps each message so the model can tell system, user and assistant turns apart.
Jailbreak promptA jailbreak prompt is text added to a request to persuade a safety-tuned model to answer things it would normally refuse.
Logit biasLogit bias is an API setting that adds a fixed amount to the scores of specific token IDs, making those tokens more likely, less likely, or effectively banned.
LorebookA lorebook, called World Info in SillyTavern, is a set of entries that are added to the prompt only when their trigger keywords show up in the recent chat.
Max tokensMax tokens is the upper limit on how many tokens the model is allowed to generate in a single response.
Min PMin P is a truncation sampler that removes every token whose probability is less than a set fraction of the most likely token's probability.
MirostatMirostat is an adaptive sampler that adjusts itself token by token to keep the text's average surprise close to a target value.
Mixture of experts (MoE)A mixture of experts (MoE) model contains many parallel sub-networks called experts and uses a router to send each token through only a few of them.
OpenAI-compatible APIAn OpenAI-compatible API is a service that accepts requests and returns responses in the same format as OpenAI's API, so tools built for OpenAI can use it with only a URL and key change.
PersonaA persona is the name and description of the character you play in a roleplay, which the frontend sends to the model so it knows who it is talking to.
PrefillPrefill is the technique of supplying the opening text of the model's reply so that the model continues from those words instead of starting fresh.
Presence penaltyPresence penalty is an OpenAI API setting that lowers the score of any token that has already appeared, by the same amount no matter how often.
Prompt cachingPrompt caching is a feature where a model server saves the computed state of a prompt's opening section so later requests that start with the same text can skip that work.
QuantizationQuantization is the process of storing a model's weights with fewer bits per number so the model uses less memory and runs faster, at some cost in accuracy.
RAGRetrieval augmented generation (RAG) is a method where relevant text is looked up from an outside source and placed in the prompt so the model can use it when answering.
Rate limitsRate limits are caps an API places on how many requests or tokens a key can use within a time window, enforced by rejecting excess requests with an error.
Reasoning modelsA reasoning model is a language model trained to write out a step-by-step thinking process before giving its final answer, which tends to improve accuracy on hard problems.
Refusal directionThe refusal direction is a single direction in a language model's internal activation space that, when present, makes the model refuse a request.
Repetition penaltyRepetition penalty is a sampler that makes tokens already present in the recent context less likely to be chosen again.
Stop sequencesA stop sequence is a string that, when the model produces it, immediately ends the response so the text stops at that point.
StreamingStreaming is an API mode where the model's reply is sent piece by piece while it is being generated, instead of all at once at the end.
SummarizationSummarization is a memory technique where older parts of a chat are condensed into a short recap that is inserted into the prompt so the model keeps track of the story.
SwipesA swipe is an alternative version of the latest AI reply, generated on request, that you can flip between to pick the one you want to keep.
System promptA system prompt is a block of instructions placed before the conversation that tells the model how to behave for the whole chat.
TemperatureTemperature is a sampling setting that makes a language model's word choices more predictable when lowered and more varied when raised.
TokenizerA tokenizer is the component that converts text into the numbered tokens a language model works with, and converts the model's tokens back into text.
TokensA token is a small chunk of text, often a word or part of a word, that a language model reads and writes one at a time.
Tool callingTool calling, also called function calling, lets a model respond with a structured request to run one of your functions, whose result you then send back to it.
Top KTop K is a sampler that only lets the model choose its next token from the K most probable candidates.
Top PTop P, or nucleus sampling, limits each token choice to the smallest group of candidates whose combined probability reaches a set threshold.
Uncensored fine-tuneAn uncensored fine-tune is a language model that has been trained further on data without refusals or moralizing so that it answers requests a safety-tuned model would decline.
Vector memoryVector memory is a technique that stores past messages as embeddings and automatically re-inserts the ones most similar to the current conversation.
Vision modelsA vision model is a language model that can take images as input along with text and describe, reason about, or respond to what it sees.
XTC samplerXTC (Exclude Top Choices) is a sampler that occasionally removes the most predictable token candidates so the model picks a less obvious but still plausible one.