Rate limits: why requests get 429 errors
Rate limits are caps an API places on how many requests or tokens a key can use within a time window, enforced by rejecting excess requests with an error.
What rate limits are
Every API that serves models has finite capacity, and rate limits divide it fairly. Limits are usually expressed per key or per account as requests per minute, tokens per minute, or concurrent requests (how many can be in progress at once). Separate from rate limits, a request can also fail because the upstream model is overloaded, or because your account has run out of credit.
How they show up
A rate-limited request returns HTTP status 429 Too Many Requests. The response often includes a Retry-After header saying how many seconds to wait, and sometimes headers with your remaining quota. Other statuses to tell apart:
401: bad or missing API key.402or a billing message: out of credit.429: too many requests, slow down.500,502,503: server or upstream trouble, usually temporary.
Handling them in code
- On a 429, wait for the
Retry-Aftervalue if given. - Otherwise retry with exponential backoff: wait 1 second, then 2, 4, 8, with some random jitter, up to a maximum.
- Cap the number of retries and surface the error after that.
- Limit concurrency in batch jobs rather than firing hundreds of requests at once.
Do not retry 400 or 401 errors in a loop; they will not fix themselves. Be careful with automatic retries on long, expensive prompts: a client that retries five times on a server error can multiply your input cost if some of those attempts were in fact processed.
When streaming, a limit or upstream error can also arrive after the stream has started, as an error event or a stream that ends without a finish reason. Treat a stream that closes early as a failed request rather than a finished reply.
How roleplay frontends hit limits
- Rapid swipes and regenerations.
- Extensions firing extra requests in the background: summaries, captions, expression classification.
- Group chats with auto-mode, where several characters reply in a row.
- Several browser tabs or devices using the same key.
- Very long prompts counting heavily against token-per-minute limits.
In SillyTavern, the error text shown in the notification is usually the API's own message. Turning off streaming temporarily can make the full error easier to read.
Rate limits on Wild West API
Limits on Wild West API depend on your account and the upstream model, and can change. Check the docs for current values and error formats. Treat 429 responses as a signal to back off, and keep one key per app so you can tell which program is using your quota.
FAQ
What does error 429 mean?
Too many requests. You have hit a rate limit and should wait, ideally for the time given in the Retry-After header, before trying again.
Do rate limits count tokens or requests?
It depends on the API. Many count both requests per minute and tokens per minute, and some also limit concurrent requests.