Set budgets and limits
Many apps can share one key. If one of them loops, or gets a traffic spike, it can spend money for all of them. To stop that, give the app a limit. Above its limit, Proxium refuses the calls of that app. The other apps keep running.
You can also give a key its own limits. Above them, Proxium refuses every call of that key.
What Proxium checks before each call
- 1Did the key reach one of its limits?The limits that you set for the key on Keys. A new key has noneyesRefused: 429 tenant_cappedno
- 2Did support-bot reach its own limit?The limit that you set for this app on OverviewyesRefused: 429 source_cappedno
A refused call never reaches a vendor, so it costs nothing.
Set a limit for one app
An app is the name that it sends in the x-proxium-source header. You can give each app two limits:
| Limit | Means |
|---|---|
| Calls per hour | The most calls in the last hour |
| Dollars per day | The most spend in the last day |
You set both on the console screen Overview, in the panel Per-app burn. There is no API for them.
To give support-bot its limits:
- Make sure the app sends
x-proxium-source: support-bot. Track spend shows how. - Open Overview, and select the row support-bot in Per-app burn.
- Type the limits in Calls per hour and Dollars per day.
- Select Set ceiling.
The Ceiling column of the row shows the new limits. 0 in a field means no limit. To remove both limits, select the row, then Remove ceiling.
An app shows in Per-app burn only after it made a call in the selected time window. If you do not see it, make one call, or select a longer time window.
The limits apply at once on the server that saved them, and soon after on the others.
Set the limits of a key
A key can have these limits. A new key has none.
| Limit | Means |
|---|---|
| Calls per hour | The most calls in the last hour |
| Requests per minute | The most calls in the last minute |
| Tokens per minute | The most tokens in the last minute |
| Dollars per day | The most spend in the last day |
To give a key its limits:
- Open Keys.
- On the row of the key, select Limits.
- Type the limits that you want. Leave the other fields empty.
- Select Save limits.
The Limits column of the row shows the new limits. An empty field means no limit. The key of an endpoint is on Keys too.
The limits apply to the next call of the key.
When an app reaches its limit
Proxium answers the call with HTTP 429. The retry-after header says how long to wait. The body says which app, and which limit:
{
"error": {
"message": "source budget exceeded for 'support-bot' (calls)",
"type": "source_capped",
"code": "source_capped"
}
}
(calls) means the calls of the last hour. (cost) means the spend of the last day.
Your app should wait, then try again. With the OpenAI SDK:
- Python
- Node
import time
import openai
try:
response = client.chat.completions.create(model="standard", messages=messages)
except openai.RateLimitError as e:
wait = int(e.response.headers["retry-after"])
time.sleep(wait)
response = client.chat.completions.create(model="standard", messages=messages)
async function ask() {
return client.chat.completions.create({ model: "standard", messages });
}
let response;
try {
response = await ask();
} catch (e) {
if (e.status !== 429) throw e;
const wait = Number(e.headers?.["retry-after"]);
await new Promise((r) => setTimeout(r, wait * 1000));
response = await ask();
}
On /v1/messages, the body has the Anthropic shape, with the type rate_limit_error. The message is the same.
How the limits are counted
- Calls per hour counts each call that Proxium let through in the last hour. A refused call does not count.
- Dollars per day adds the cost of the calls in the last day.
- Proxium checks the spend before a call, not during it. So the last call that starts under the limit can end a little above it.
- A free cache hit counts toward neither limit. See The response cache.
- All Proxium servers count together. Two servers cannot both let through the last allowed call.
Check the key before a call
GET https://proxium.tech/status/peek checks your key, and says if your next call will be refused. It does not count as a call, and it costs nothing. Use it before a batch of work.
curl https://proxium.tech/status/peek \
-H "Authorization: Bearer $PROXIUM_KEY"
A wrong or revoked key gets 401 invalid_key. A valid key gets this answer:
{
"tenant_capped": false,
"budget": { "calls_pct": 0, "cost_pct": 0 },
"credits": { "over": false },
"providers": [{ "provider": "t.acme.groq", "state": "closed" }],
"any_provider_available": true
}
| Field | Means |
|---|---|
tenant_capped | true: the key reached a limit, and the next call gets 429 tenant_capped |
budget.calls_pct, budget.cost_pct | How much of the limits of the key is used, in percent. 0 for a key with no limit |
credits.over | true: the project reached a monthly cap. No project has one today |
providers | Each vendor of the project: closed is working, open is skipped. See Routing |
any_provider_available | true: at least one vendor is not skipped, or the list is empty |
/status/peek does not see the limits of an app. A call can get 429 source_capped while tenant_capped is false.
A check before a batch:
import os
import requests
state = requests.get(
"https://proxium.tech/status/peek",
headers={"Authorization": f"Bearer {os.environ['PROXIUM_KEY']}"},
timeout=5,
).json()
if state["tenant_capped"] or not state["any_provider_available"]:
raise RuntimeError("Proxium will refuse this call now. Hold the work.")
Limits that Proxium does not have
Every limit of Proxium belongs to a key or to an app. These limits do not exist:
| You want | State | What to do |
|---|---|---|
| A spend limit for the whole project | Not available. Proxium counts the spend of the project, but no limit refuses a call on it | Give each key a Dollars per day limit. Add a spend level for the project to get an alert |
| A limit for a tier or a model | Not available | Give the apps that use the tier their own key or app limit. To spend less on a tier, put a cheaper model first on Routing |
| A limit for a month | Not available. Dollars per day counts the last day | Add a spend level for a month. It sends an alert, and it refuses no call |
Related pages
- Monitor your project: watch the spend and the failures.
- Track spend: see what each app spent.
- Errors: every 429.