Wild West API

SillyTavern with uncensored AI models: setup guide

SillyTavern is the frontend. It needs a model behind it. This guide connects SillyTavern to uncensored AI models in about two minutes, then covers the settings that hold up over a long story.

Updated 7 October 2026

Why SillyTavern needs an API at all

SillyTavern does not contain a language model. It is the stage: character cards, chat history, lorebooks, group chats and a very long list of settings. Every reply you see is written by a model running somewhere else, and SillyTavern reaches it through an API. That is why the same install can feel brilliant with one connection and useless with another.

Most mainstream APIs are built for assistants, and they show it. Ask for a villain with real menace, a horror scene that actually frightens, or adult fiction between adults, and you get a soft refusal, a lecture, or a character who suddenly talks like a help desk. An uncensored API removes that layer, so the model plays the part the card describes. Wild West API serves only uncensored and abliterated models, through the same OpenAI-compatible format SillyTavern already speaks.

What you need

  • SillyTavern installed and running. Any recent release works.
  • A Wild West API key, which starts with sk-ww-.
  • A few dollars of credit. The smallest top-up is $5.00.

Connect it, step by step

  1. Get a key. Sign up with your email, top up a few dollars from the dashboard, and create a key under API. Copy it, it is only shown once.
  2. Open API Connections. In SillyTavern, click the plug icon at the top of the screen.
  3. Choose Chat Completion. Set API to Chat Completion. Text Completion will not connect.
  4. Choose the custom source. Set Chat Completion Source to Custom (OpenAI-compatible).
  5. Paste the base URL. In Custom Endpoint (Base URL), paste https://wildwestapi.com/v1 with no trailing slash.
  6. Paste your key. Put your sk-ww- key in Custom API Key.
  7. Connect and pick a model. Press Connect. The model list fills in from the API. Pick one, or type its id into the model box.
  8. Send a test message. Use Test Message, or just open a character and say hello.

The two values that matter: base URL https://wildwestapi.com/v1, and your key. Everything else is SillyTavern's own settings.

Which model to start with

Start cheap and step up when a scene deserves it. Switching is one click in the connection panel, and your chat carries over, so you can write the slow middle of a story on a fast model and bring in the big one for the finale.

  • Everyday roleplay: a Flash model. Fast, cheap and a long memory.
  • Big moments and complex casts: the full-size model. Slower and pricier, better at keeping a large cast straight.
  • On a budget: the smallest model on the line.

Live prices for each are on the model list, and the guide to the best uncensored LLM for roleplay goes deeper into choosing.

Settings that work well

SettingTryWhy
Context Size32K to 64K to beginThe models hold far more, but every token of context is sent and billed on every message. Raise it when the story needs older history.
Max Response Length600 to 1,000Room for a full reply, plus thinking on reasoning models.
StreamingOnText appears as it is written.
Temperature0.8 to 1.0Lower is steadier, higher is looser. Drop it if characters start to ramble.

Keep the character card tight. A long card is resent with every message, so a 3,000 token card costs more than a 600 token one for the whole chat. Lorebooks help here, because entries only go in when a keyword triggers them.

What a long chat costs

Say a scene has grown to 6,000 tokens of character card and history, and the model writes a 300 token reply, on a model priced at $0.40 per million input tokens and $1.60 per million output tokens. Input costs $0.0024, output costs $0.00048, so the message costs about $0.003, or roughly 350 messages for a dollar. Because the whole history goes back with every message, cost per message climbs as the chat grows. Capping Context Size, trimming old messages, or using a summary keeps it flat. Each key can also carry a monthly cap in the dashboard, which is a good idea if you share a key across devices.

Group chats and long campaigns

Group chats send every active character's card, so they grow fast. With a million token model you can let history run long, but the bill grows with it. SillyTavern's Summarize extension is a good middle path: it keeps a running summary of older events and lets you trim the raw context without losing the plot. For multi-session campaigns, put the world, factions and key places in a lorebook rather than in each card.

Writing better with an uncensored model

Removing refusals does not make a model a better writer by itself. Example dialogue on the card does more for consistency than any setting. Author's Note is the right place for tone instructions you want close to the end of the prompt, such as pacing or point of view. If a model keeps steering towards happy endings, say plainly in the card or the note that the story is dark and should stay that way; an uncensored model will follow that instruction instead of fighting it.

Browser-only apps

SillyTavern calls the API from your own machine, which is why it works. Some websites call an API straight from the browser tab, and that needs cross-origin support the API does not offer today. If a web app shows a network or CORS error, use SillyTavern for that chat.

Troubleshooting

SillyTavern says the connection failed. What now?

Check three things. The base URL ends in /v1 with nothing after it. The key has no space at the start or end. Your balance is above zero. A 402 error means credit, a 401 means the key.

The model list is empty.

Press Connect again after pasting the key. If it stays empty, type a model id by hand, like glm-5.3-flash-xploded. Ids are on the model list.

Replies cut off halfway.

Raise Max Response Length in the AI Response Configuration panel. Reasoning models think before they write, and the thinking counts towards that limit.

I get an error about message roles or order.

Some models are strict about alternating user and assistant turns. In the connection panel, set Prompt Post-Processing to one of the stricter options and try again.

Can I use this on my phone?

Yes. SillyTavern runs on Android through Termux, and you can also run it on a home computer and open it from your phone on the same network. The API settings are the same either way.

Keep reading