Fixing 429 Too Many Requests when calling the API
Paid requests on this API are not answered with a 429 for sending too often. If you are seeing a 429, it is worth finding out where it really comes from.
How this API handles rate limits
The chat completions path does not apply a per-key request-rate limit of its own. What limits you is money: the balance and any caps on your key, which come back as 402s, not 429s.
The models behind the API can be busy, and they sometimes answer with a 429. The API treats that as a brief outage: it retries after about 0.6 and 1.8 seconds within a 60 second budget, and if it still fails you get a 503 Model overloaded. Please try again in a minute. The one exception is a 429 about billing upstream, which is not retried, since repeating it does not help; that is also reported as the 503.
More on the general idea in rate limits.
Where a 429 actually comes from
- Another provider. SillyTavern keeps settings per Chat Completion source. If the source is not Custom (OpenAI-compatible), you are talking to a different service with its own limits.
- A proxy in between. Some hosted front ends, shared proxies and gateways enforce their own limits and return their own 429.
- The free chat on the website, which allows one free message and then returns 429
You have used your free message. Sign in to keep chatting.That has nothing to do with API keys.
How to check
- Look at the error body. This API's errors are JSON with an
error.message; a 429 page in HTML or with unfamiliar wording comes from somewhere else. - In SillyTavern, confirm the source is Custom (OpenAI-compatible) and the endpoint is
https://wildwestapi.com/v1. - Send one request with curl straight to the API. If that works, the limit lives in whatever sits in between.
If you are sending many requests at once
Parallel requests are fine, but each one holds part of your balance until it finishes. Many at once can run into a 402 even with enough money for each one alone. Spread them out, or keep more balance on the account. If you get 503s under heavy load, add backoff and retry.
Telling it apart from similar errors
A 503 Model overloaded is what an upstream rate limit looks like here; see 503. A 402 is about money; see 402.
FAQ
Do the OpenAI SDKs retry 429 and 503?
Yes, both are retried twice by default with backoff before the error is raised.
Is there a requests-per-minute limit on my key?
Not on the chat completions path. Spending is limited by balance and per-key caps.