Run GLM 5.3 Flash Xploded in Cline
Run glm-5.3-flash-xploded in Cline through Wild West API: one key, a hard spend cap, and a model that answers instead of refusing.
| Model id | glm-5.3-flash-xploded |
|---|---|
| Context window | 1,048,576 tokens (1M) |
| Tool calling | Supported |
| Price per 1M tokens | $0.40 input, $0.14 on a cache hit, $1.60 output |
| Intelligence | 42 on the Artificial Analysis Intelligence Index v4.3.2, for the base model |
| Web search | Add :online to the id |
Set it up
- Open the Cline settings (the gear icon) and choose OpenAI Compatible as the API Provider.
- Base URL https://wildwestapi.com/v1, your sk-ww-... key, Model ID glm-5.3-flash-xploded.
- Under Model Configuration set Context Window 1,048,576 and Max Output Tokens 32,768.
API Provider: OpenAI Compatible Base URL: https://wildwestapi.com/v1 API Key: sk-ww-... Model ID: glm-5.3-flash-xploded Context Window: 1048576 Max Output Tokens: 32768
What a session costs
For a 25-turn agent task (about 750K input tokens re-read across turns, 85% of them cache hits, 25K output), GLM 5.3 Flash Xploded comes to about $0.17 at today's prices; the cache price does most of the work, since re-read context is billed at $0.14 rather than $0.40 per million.
FAQ
Can GLM 5.3 Flash Xploded edit files and run commands in Cline?
Yes. Cline acts through tool calls and this model supports tool calling, so it can read, edit and run, not just chat.
What does a session cost?
About $0.17 for a 25-turn agent task (about 750K input tokens re-read across turns, 85% of them cache hits, 25K output), at $0.40 in, $1.60 out and $0.14 per million on cache hits. Put a spend cap on the key so a long run cannot overrun it.
Is it uncensored?
A million tokens of context, tool calling and vision, with no refusals. Point a coding agent at it.