Wild West API or Hyperbolic?
Hyperbolic used to sell per-token inference, but its own docs now say that API and the playground have been retired, and it sells GPU rental while it builds a new inference product. If you want to run any open model on hardware you control, Hyperbolic is the better fit; if you want to call an uncensored model by API without running a server, Wild West API is.
Facts checked October 2026 against Hyperbolic’s own site, docs and policies.
| Wild West API | Hyperbolic | |
|---|---|---|
| What it is | Hosted gateway for one uncensored, tool capable model line (Xploded), at twice the upstream price. | An AI cloud selling On-Demand GPUs, Reserved capacity and Private Cloud (H100, H200, B200, B300) for training, fine-tuning and serving models. |
| Pricing | Pay as you go, 2x upstream price, $5 minimum top-up, no subscription required. | Per GPU hour. On-Demand is billed hourly with no minimum commitment, but your balance must cover an hour of runtime before an instance launches. Reserved is a discounted rate paid up front for a fixed term from 1 week. A one-time deposit of at least $5 is needed before creating instances, and credits never expire. A free tier is not documented. |
| OpenAI-compatible | Yes | No, not today. Hyperbolic's FAQ says the serverless inference API has been retired and requests to the old models fail even with a valid key. You can run your own OpenAI-compatible server, for example vLLM, on a rented GPU. |
| Anthropic-compatible | Yes, /v1/messages (verified with Claude Code) | Not documented. |
| Uncensored / abliterated | Four uncensored models, the Xploded line: GLM 5.3, GLM 5.3 Flash, MiMo V2.6 Flash and Qwen3.8 27B. Every one calls tools. | They no longer host models. On a rented GPU you can run whatever open-weight model you choose, uncensored finetunes included, but you deploy and serve it yourself. |
| Per-key spend caps | Lifetime and monthly per-key caps, enforced by reserving the worst case and clamping max_tokens | Per-key spend caps are not documented. Spending is bounded by your prepaid balance, and instances can be terminated automatically when it runs out. |
Where Hyperbolic is the better choice
- Full root access to dedicated H100, H200, B200 or B300 GPUs, so you can run any open-weight model, finetune or serving stack you like, not a fixed list.
- For sustained heavy use, a reserved GPU at a discounted hourly rate can work out cheaper than paying per token, with no reseller markup.
Where Wild West API is different
- Wild West API is per token and serverless: an OpenAI-compatible endpoint and an Anthropic-compatible /v1/messages endpoint for the Xploded models (GLM 5.3 Xploded, GLM 5.3 Flash Xploded, MiMo V2.6 Flash Xploded, Qwen3.8 27B Xploded), with no GPU to provision. It resells upstream providers at 2x the upstream price.
- Each Wild West API key can carry hard lifetime and monthly spend caps, enforced by reserving the worst-case cost and clamping max_tokens before the call, pay as you go from a $5 top-up with no subscription.
The uncensored models Wild West API sells
If you are comparing Hyperbolic because you want an uncensored or abliterated model API, these are the models on Wild West API: Outlaw 1 abliterated, GLM 5.3 uncensored, GLM 5.3 Flash uncensored, MiMo V2.6 Flash abliterated and Qwen3.8 27B uncensored. Every one calls tools, works in Claude Code through the Anthropic endpoint, and is priced at twice what we pay upstream, per million tokens.
| Model | What it is | Context | Price in / out | Supports |
|---|---|---|---|---|
| Outlaw 1 Xploded outlaw-1-xploded | Outlaw 1 abliterated API | 1M | $0.30 / $1.00 | tools, vision, reasoning |
| GLM 5.3 Xploded glm-5.3-xploded | GLM 5.3 uncensored API | 1M | $2.50 / $4.50 | tools, reasoning |
| GLM 5.3 Flash Xploded glm-5.3-flash-xploded | GLM 5.3 Flash uncensored API | 1M | $0.40 / $1.60 | tools, vision, reasoning |
| MiMo V2.6 Flash Xploded mimo-v2.6-flash-xploded | MiMo V2.6 Flash abliterated API | 1M | $1.00 / $3.00 | tools, vision, reasoning |
| Qwen3.8 27B Xploded qwen3.8-27b-xploded | Qwen3.8 27B uncensored API | 524K | $0.30 / $2.40 | tools, vision, reasoning |
FAQ
Does Hyperbolic still have a serverless inference API?
No. Hyperbolic's own FAQ says the serverless inference API and the playground have been retired, the models it used to list are no longer available, and a new inference product has no release date yet. It suggests renting an on-demand GPU and running your own serving stack.
Can I use Hyperbolic with Claude Code?
Not directly. No Anthropic-compatible endpoint is documented and the hosted inference API is retired, so you would serve a model yourself on a rented GPU and put a compatible proxy in front of it. Wild West API offers an Anthropic-compatible /v1/messages endpoint that works with Claude Code.
Is renting a Hyperbolic GPU cheaper than paying per token?
It can be if the GPU stays busy, because you pay by the hour whatever you send. For light or bursty use you pay for idle hours too, plus the work of running the server, so per-token billing is usually simpler.
Which uncensored models does Wild West API sell?
Wild West API sells Outlaw 1 Xploded, GLM 5.3 Xploded, GLM 5.3 Flash Xploded, MiMo V2.6 Flash Xploded, Qwen3.8 27B Xploded. The table above has the price and context for each.
Are these models abliterated or uncensored?
Outlaw 1 Xploded and MiMo V2.6 Flash Xploded are abliterated: the refusal direction was removed from the weights. GLM 5.3 Xploded, GLM 5.3 Flash Xploded, Qwen3.8 27B Xploded are uncensored: the refusal was trained back out. The abliterated vs uncensored page explains the difference.
Which one answers fastest?
MiMo V2.6 Flash Xploded usually starts answering in one to two seconds. The GLM 5.3 Flash model thinks before it answers, so its first word can take much longer.