Wild West API

Leaving KoboldCpp? Uncensored AI models instead

KoboldCpp is the simplest way to run a model on your own machine. The alternatives are the other local runners below, or a hosted endpoint for the models your hardware cannot hold, which is what Wild West API is.

What KoboldCpp is, and where it stops

KoboldCpp runs GGUF models locally from a single executable, with the KoboldAI Lite interface and an OpenAI-compatible API built in, and filters nothing. The ceiling is your hardware: the largest models will not fit, and the server only answers while your machine is on.

Other options people compare with KoboldCpp

OptionWhat it is
OllamaLocal runner with a one-line install and a model library. Serves an OpenAI-compatible API on localhost.
LM StudioDesktop app for downloading and running local models, with a built-in OpenAI-compatible server.
Text Generation WebUIOobabooga's web UI for loading local weights with a choice of loaders and extensions. Fully manual, fully local.
AI HordeFree, volunteer-run network that KoboldAI Lite and other front ends can send requests to. Availability depends on donated hardware.
SillyTavernFree, open-source front end you run yourself: character cards, lorebooks, group chats, and a connection to almost any API, including any OpenAI-compatible one. No model of its own.
Worth being straight about Nothing here is a jailbreak, and none of it makes an illegal act legal. A loosely aligned model answers more questions; it does not change what you are allowed to do with the answer, and it is frequently more confidently wrong than the model you left.

What Wild West API carries instead

Wild West API sells 5 uncensored models on one OpenAI-compatible endpoint, every one able to call tools. Ordered for this particular job rather than by a single ranking: MiMo V2.6 Flash Xploded leads because it carries the longest window of the set at 1.05M, which is what keeps a long session coherent.

ModelContextInOutTools
MiMo V2.6 Flash Xploded
mimo-v2.6-flash-xploded
1.05M $1.00 $3.00 Yes
Outlaw 1 Xploded
outlaw-1-xploded
1.05M $0.300 $1.00 Yes
GLM 5.3 Flash Xploded
glm-5.3-flash-xploded
1.05M $0.400 $1.60 Yes
GLM 5.3 Xploded
glm-5.3-xploded
1.05M $2.50 $4.50 Yes

Switching takes one base URL

Anything that speaks the OpenAI chat completions format works unchanged. Point it here, use a Wild West API key, and name the model from the table.

curl https://wildwestapi.com/v1/chat/completions \
  -H "Authorization: Bearer sk-ww-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mimo-v2.6-flash-xploded",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

FAQ

Is Wild West API actually uncensored, or does it just say so?

Wild West API does not moderate requests. It prices the call, holds the money against your key cap and passes the request to the model. What comes back is whatever that model does: an uncensored build refuses far less than a frontier model but can still decline, and only an abliterated build has the refusal direction removed outright. Neither tag is a promise that any given answer will arrive.

What does it cost compared with KoboldCpp?

You pay per token at the rates in the table above, with no subscription. $1.00 in and $3.00 out per million tokens for MiMo V2.6 Flash Xploded.

Can I cap what a key is allowed to spend?

Yes, and the cap is enforced before the request goes upstream rather than reconciled afterwards. The reply cannot physically cost more than the hold placed before it was sent.

Will it work with my existing client?

If the client takes an OpenAI-compatible base URL, yes: point it at https://wildwestapi.com/v1 with a Wild West API key. Claude Code and the Anthropic SDKs use the Anthropic format, which is served too.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.