Use an uncensored API in AnythingLLM with Generic OpenAI
AnythingLLM is a desktop and Docker app for chatting with documents and models. Its Generic OpenAI provider accepts any OpenAI-compatible API, and its built-in server makes the call, so the Wild West API works.
Set it up
- Sign in at wildwestapi.com/signup/ with the six-digit code sent to your email, open the dashboard and create an API key. It starts with sk-ww- and is shown in full once, so copy it somewhere safe.
- Open settings and go to AI Providers, then LLM.
- Choose Generic OpenAI as the LLM provider.
- Set Base URL to https://wildwestapi.com/v1.
- Paste your sk-ww- key into API Key.
- Type outlaw-1 into the model field (shown as Selected Model, placeholder "Model id used for chat requests").
- Fill in Model context window and Max Tokens (see below), then save.
Does it work with the Wild West API?
Yes. Both the desktop app and the Docker image run AnythingLLM's own Node server, and that server calls the LLM provider. The browser never talks to the API directly, so CORS does not apply.
Fill in the two "optional" fields
AnythingLLM's Generic OpenAI form has a Model context window field (placeholder "Content window limit (eg: 4096)") and a Max Tokens field (placeholder "Max tokens per request (eg: 1024)"). Left empty, they fall back to small defaults, which trims history and cuts long replies short.
- Model context window: 65536 is a sensible start for roleplay. The models here go to 512K or 1M tokens, so you can go higher if you accept the cost of resending more history each turn. See context window.
- Max Tokens: 1024 to 2048 so scenes are not cut off. See max tokens.
Other settings for long roleplay
- Each workspace has its own system prompt and chat history settings. Put your character or narrator rules in the workspace prompt; see system prompt.
- Workspace chat mode should be Chat rather than Query if you are writing fiction, so the model is not forced to answer only from documents.
- The Wild West API serves chat models, not embeddings. Keep AnythingLLM's built-in embedder for documents.
Troubleshooting AnythingLLM
- 401: the key is missing, mistyped or disabled. Paste it again with no trailing space and check it starts with
sk-ww-. A key from another provider will never work here. - 402: your balance is too low for the request. Top up in the dashboard; see pricing. A huge context on every turn burns credit faster than you might expect.
- 404 or 400 with an unknown model: the model id is wrong. Use an exact id from the model list, such as
outlaw-1, with no provider prefix and no spaces. - 429: you are rate limited. Wait a few seconds and send again, and turn off any auto-retry loop that fires instantly.
- 5xx: trouble upstream. Retry; if it keeps happening, switch to another model id for a while.
- Network error, "Failed to fetch" or a CORS message: this should not happen, since the AnythingLLM server makes the request. Check the Base URL ends in /v1 and that the machine running AnythingLLM can reach wildwestapi.com.
- Replies stop short or forget the start: Model context window or Max Tokens is empty or too small.
FAQ
Can I use this API for embeddings in AnythingLLM?
No. Use the built-in embedder or a separate embedding provider. The Generic OpenAI setting here is for chat.
Why are my replies cut off at around 1000 tokens?
Max Tokens is empty and falling back to a small default. Set it to 1024 or more.
Does the desktop version behave differently from Docker?
No. Both run the same server, which makes the API call, so the setup is identical.