Fixing a 402 "Account balance is too low" error
A 402 means your key is valid but the account does not have enough balance to cover the request it was about to send. Nothing was generated or charged.
What the error says
The API answers HTTP 402 with type insufficient_quota:
{
"error": {
"message": "Account balance is too low. Top up at /app.",
"type": "insufficient_quota",
"code": "insufficient_credit",
"param": null
}
}Top-ups are done from the dashboard, by card or crypto.
How the balance check works
Before calling the model, the API estimates the prompt size, works out the most output it could produce, and places a temporary hold on your balance for that worst case. When the reply finishes, the hold is released and you are charged what the reply actually cost. A 402 with insufficient_credit means that hold could not be placed.
Two details matter here. First, holds from requests still in flight count against your balance, so several parallel generations (for example group chats or swipes fired quickly) can trip it even when one request alone would fit. Second, before refusing outright the API tries to shrink the reply: it lowers max_tokens to what the balance can afford. A 402 only appears when even a minimal reply will not fit.
How it shows up
SillyTavern shows an error toast with the message text. The OpenAI Python SDK raises openai.APIStatusError with status_code 402; there is no dedicated class for 402, and the SDKs do not retry it. Anthropic SDKs on /v1/messages see a 402 billing_error, also not retried.
Step by step fix
- Open /dashboard/ and check the balance.
- Top up by card or crypto. See pricing for per-token rates by model.
- If you run several requests at once, wait for the others to finish, since their holds are still reserved.
- Retry the message. No settings change is needed.
Telling it apart from similar errors
Three different 402s exist. insufficient_credit is the account balance. key_budget_exceeded is a spend cap you set on one key; see key spend cap reached. budget_exceeded with "will not cover this request" means the tightest of balance and caps cannot pay even for a minimum reply; see spend limit reached. If replies arrive but are oddly short, the balance is close to empty and the reply length was trimmed; see replies cut short.
FAQ
Was I charged for the failed request?
No. The request is refused before the model is called, so there is no cost.
Why did it work for one message and fail on the next?
Each request holds enough balance for its worst case. A long chat sends a bigger prompt each turn, so the hold grows, and in-flight requests also count against what is free.