Proxium has limits on size, time, retries, spend and names. This page says what each limit does and what happens when a call reaches it.
Size
| Limit | At the limit |
|---|
| Request body | 413, plain-text body |
| Vendor answer | The call fails. Proxium tries the next model |
| Stored copy of an answer | Proxium cuts the stored copy. Your app gets the full answer |
| Model id of a media call | 400 bad_model |
A media model id can hold only letters, digits, ., _, - and :.
Time
| Limit | At the limit |
|---|
x-proxium-timeout-ms | A value above the maximum is lowered to the maximum |
| All calls of one request, with the header | 503 upstream_failed |
| All calls of one request, without the header | No limit |
| One call, not streamed | Timeout. Proxium tries the next model |
| Connect to a vendor | The call fails. Proxium tries the next model |
| Silence in a stream | Proxium ends the stream |
- Proxium ignores an
x-proxium-timeout-ms that is not a positive number.
- When the header time ends, the message of the
503 starts with client deadline exceeded.
- For a stream, the header time ends at the first piece. After that, only the silence limit applies. A stream has no limit on its total time.
- The silence limit also runs while the vendor prepares its first piece.
The response cache
A short cache is available, so a loop of the same request is cached. Each project can turn its cache off, or keep answers for a shorter or a longer time, in Settings › Cache. See The response cache.
Retries and failover
- A plan holds a few models. When all of them fail, the call gets
503 upstream_failed.
- After a server error, a timeout or a "warming" answer, Proxium can call the same model again, after a short wait. Then it tries the next model.
- After a
429, the next call goes to a model at another vendor of the plan.
- Proxium honors a vendor's
Retry-After up to a limit.
- When a vendor keeps failing, Proxium pauses it, then sends a trial call. Good trial answers bring the vendor back.
- When a model keeps failing, Proxium retires it for your project for a while.
- Proxium retires a vendor in the same way when all its models keep failing. It also retires a vendor that says your own key has no credits or is rejected.
- A retired model or vendor comes back by itself.
- A
429 never pauses a vendor and never counts toward a retirement.
- Your project sets the retries, the waits and the retirement rules in its failover policy, on Routing › When a call fails. See Set your failover policy.
- Proxium does not skip a fallback model for its price. A fallback can cost more than your first model.
Spend and rate
| Limit | At the limit |
|---|
| App: Calls per hour | 429 source_capped |
| App: Dollars per day | 429 source_capped |
| Key: calls per hour | 429 tenant_capped |
| Key: requests and tokens per minute | 429 tenant_capped |
| Key: dollars per day | 429 tenant_capped |
- A limit that you do not set does not apply.
- Each
429 carries a retry-after header.
- Each limit counts a rolling window: the last hour, the last day or the last minute.
- You set the app limits on Overview › Per-app burn, and the key limits on Keys. See Set budgets and limits.
Names and lists
| Limit | At the limit |
|---|
| Key name, on Keys | A name that is too long is dropped. The key has no name |
| Endpoint name | 400 bad_name |
| Vendor name | 400 bad_name |
| Models of one vendor | 400 too_many_models |
A vendor name holds only lower-case letters, digits and -, with no - at the start or the end.
How soon a change applies
| Change | Applies |
|---|
| Routing, app limits, the cache price | At once on one server, and soon after on all |
| The model list of each vendor | Proxium reads it again from time to time |
Related pages