Wild West API or self-hosted LiteLLM?
LiteLLM Proxy is the build-it-yourself version of what Wild West API sells: an open-source, OpenAI-compatible gateway with virtual keys and budgets. If you want full control and no margin, self-hosting wins. What it costs you is the ops and the upstream keys.
Facts checked September 2026 against self-hosted LiteLLM’s own site, docs and policies.
| Wild West API | self-hosted LiteLLM | |
|---|---|---|
| What it is | Hosted gateway for one uncensored, tool capable model line (Xploded), at twice the upstream price. | An open-source (MIT) OpenAI-compatible proxy you run yourself, in front of your own provider keys, with virtual keys and budgets. |
| Pricing | Pay as you go, 2x upstream price, $5 minimum top-up, no subscription required. | Free and open source. You pay your upstream providers directly at their list price, no gateway margin. Some enterprise features (SSO, audit logs) are paid. |
| OpenAI-compatible | Yes | Yes |
| Anthropic-compatible | Yes, /v1/messages (verified with Claude Code) | Yes, an Anthropic-format pass-through is available |
| Uncensored / abliterated | Four uncensored models, the Xploded line: GLM 5.3, GLM 5.3 Flash, MiMo V2.6 Flash and Qwen3.8 27B. Every one calls tools. | None by itself: you point it at whatever upstreams you configure, so the uncensored catalogue is on you to source |
| Per-key spend caps | Lifetime and monthly per-key caps, enforced by reserving the worst case and clamping max_tokens | Per-virtual-key max_budget with a budget_duration reset window, plus rate limits; worst-case budget reservation is on by default |
Where self-hosted LiteLLM is the better choice
- No margin and full privacy: it is free, and prompts go only to the upstreams you configure, from a server you control. Nothing passes through a third-party gateway.
- Total control and real budgets: your own key policies, routing, logging and model list, with virtual-key budgets that reserve the worst case before the call, on by default. Its spend controls are as capable as Wild West API’s or more.
Where Wild West API is different
- You run it: hosting, upgrades, uptime, secrets and the upstream provider contracts are all yours, and you must source uncensored models yourself. Wild West API is that work already done, with the uncensored catalogue included.
The uncensored models Wild West API sells
If you are comparing self-hosted LiteLLM because you want an uncensored or abliterated model API, these are the models on Wild West API: Outlaw 1 abliterated, GLM 5.3 uncensored, GLM 5.3 Flash uncensored, MiMo V2.6 Flash abliterated and Qwen3.8 27B uncensored. Every one calls tools, works in Claude Code through the Anthropic endpoint, and is priced at twice what we pay upstream, per million tokens.
| Model | What it is | Context | Price in / out | Supports |
|---|---|---|---|---|
| Outlaw 1 Xploded outlaw-1-xploded | Outlaw 1 abliterated API | 1M | $0.30 / $1.00 | tools, vision, reasoning |
| GLM 5.3 Xploded glm-5.3-xploded | GLM 5.3 uncensored API | 1M | $2.50 / $4.50 | tools, reasoning |
| GLM 5.3 Flash Xploded glm-5.3-flash-xploded | GLM 5.3 Flash uncensored API | 1M | $0.40 / $1.60 | tools, vision, reasoning |
| MiMo V2.6 Flash Xploded mimo-v2.6-flash-xploded | MiMo V2.6 Flash abliterated API | 1M | $1.00 / $3.00 | tools, vision, reasoning |
| Qwen3.8 27B Xploded qwen3.8-27b-xploded | Qwen3.8 27B uncensored API | 524K | $0.30 / $2.40 | tools, vision, reasoning |
FAQ
Which uncensored models does Wild West API sell?
Wild West API sells Outlaw 1 Xploded, GLM 5.3 Xploded, GLM 5.3 Flash Xploded, MiMo V2.6 Flash Xploded, Qwen3.8 27B Xploded. The table above has the price and context for each.
Are these models abliterated or uncensored?
Outlaw 1 Xploded and MiMo V2.6 Flash Xploded are abliterated: the refusal direction was removed from the weights. GLM 5.3 Xploded, GLM 5.3 Flash Xploded, Qwen3.8 27B Xploded are uncensored: the refusal was trained back out. The abliterated vs uncensored page explains the difference.
Which one answers fastest?
MiMo V2.6 Flash Xploded usually starts answering in one to two seconds. The GLM 5.3 Flash model thinks before it answers, so its first word can take much longer.