Wild West API or Shannon AI?
Shannon AI is the cheapest place we found to call an uncensored Kimi K3, and it runs on its own GPUs rather than reselling. Two things are worth knowing: the K3 is heavily compressed, which Shannon’s own docs say costs accuracy, and "uncensored" there means no filter in front of the model rather than refusal removed from the weights.
Facts checked October 2026 against Shannon AI’s own site, docs and policies.
| Wild West API | Shannon AI | |
|---|---|---|
| What it is | Hosted gateway for one uncensored, tool capable model line (Xploded), at twice the upstream price. | Uncensored chat and API gateway with 21 open-weight models served from its own GPUs, aimed at red team and security work. |
| Pricing | Pay as you go, 2x upstream price, $5 minimum top-up, no subscription required. | Token quota valued at $5 per million quota tokens. Kimi K3 3BIT-REAP is $3.83 in / $19.12 out per million, cached input at 25%. |
| OpenAI-compatible | Yes | Yes, /v1/chat/completions and /v1/responses |
| Anthropic-compatible | Yes, /v1/messages (verified with Claude Code) | Yes, /v1/messages |
| Uncensored / abliterated | Four uncensored models, the Xploded line: GLM 5.3, GLM 5.3 Flash, MiMo V2.6 Flash and Qwen3.8 27B. Every one calls tools. | Kimi K3 3BIT-REAP, DeepSeek V4, MiniMax, Nemotron and more, served with no input gate and no output classifier. It does not claim abliteration, so trained-in refusals may remain. |
| Per-key spend caps | Lifetime and monthly per-key caps, enforced by reserving the worst case and clamping max_tokens | Not documented |
Where Shannon AI is the better choice
- Kimi K3, which Wild West API does not serve, with 262K context.
- Its own GPUs, Tor and I2P endpoints, and a broader catalogue of 21 models.
- An unusually frank model card about what compression costs.
Where Wild West API is different
- Shannon’s K3 is 3-bit quantized and expert pruned (REAP, about 5.6TB down to 1.05TB); its docs say compression costs accuracy, mostly in long-tail facts and hard maths.
- No filter in front of stock weights is not the same as abliteration. A model can still refuse on its own training. Xploded models have refusal taken out of the weights: MiMo by abliteration, the GLM and Qwen models by retraining.
- Wild West API states its metadata window and enforces hard per-key caps. Shannon documents neither.
The uncensored models Wild West API sells
If you are comparing Shannon AI because you want an uncensored or abliterated model API, these are the models on Wild West API: Outlaw 1 abliterated, GLM 5.3 uncensored, GLM 5.3 Flash uncensored, MiMo V2.6 Flash abliterated and Qwen3.8 27B uncensored. Every one calls tools, works in Claude Code through the Anthropic endpoint, and is priced at twice what we pay upstream, per million tokens.
| Model | What it is | Context | Price in / out | Supports |
|---|---|---|---|---|
| Outlaw 1 Xploded outlaw-1-xploded | Outlaw 1 abliterated API | 1M | $0.30 / $1.00 | tools, vision, reasoning |
| GLM 5.3 Xploded glm-5.3-xploded | GLM 5.3 uncensored API | 1M | $2.50 / $4.50 | tools, reasoning |
| GLM 5.3 Flash Xploded glm-5.3-flash-xploded | GLM 5.3 Flash uncensored API | 1M | $0.40 / $1.60 | tools, vision, reasoning |
| MiMo V2.6 Flash Xploded mimo-v2.6-flash-xploded | MiMo V2.6 Flash abliterated API | 1M | $1.00 / $3.00 | tools, vision, reasoning |
| Qwen3.8 27B Xploded qwen3.8-27b-xploded | Qwen3.8 27B uncensored API | 524K | $0.30 / $2.40 | tools, vision, reasoning |
FAQ
Which uncensored models does Wild West API sell?
Wild West API sells Outlaw 1 Xploded, GLM 5.3 Xploded, GLM 5.3 Flash Xploded, MiMo V2.6 Flash Xploded, Qwen3.8 27B Xploded. The table above has the price and context for each.
Are these models abliterated or uncensored?
Outlaw 1 Xploded and MiMo V2.6 Flash Xploded are abliterated: the refusal direction was removed from the weights. GLM 5.3 Xploded, GLM 5.3 Flash Xploded, Qwen3.8 27B Xploded are uncensored: the refusal was trained back out. The abliterated vs uncensored page explains the difference.
Which one answers fastest?
MiMo V2.6 Flash Xploded usually starts answering in one to two seconds. The GLM 5.3 Flash model thinks before it answers, so its first word can take much longer.