Route requests to models
The model field of a request decides which model answers. You send one of two things:
| You send | Example | Proxium calls |
|---|---|---|
| A model id | t.<project>.<vendor>/<model> | That model first |
| A tier name | standard | The models of the tier, in order |
A tier is a name for an ordered list of models. Proxium calls the first model. It calls the next one only when the model before it fails. To change the models of your apps, you change the tier on the Routing screen. Your code stays the same.
Use a model id
A model id is a vendor of your project and one of its models: t.<project>.<vendor>/<model>. To list the ids, send GET /v1/models, or open Providers and read the Models column.
To add a model to a vendor, see Add vendors and your own keys.
Set the models of a tier
The Routing screen lists the tiers, for example standard and heavy.
- Open Routing in the console.
- In Text routing, find the row of the tier, and select Change.
- Under Pick the model to try first, select the first model.
- Under Pick a fallback, add each next model, in order.
- Select Save for this project.
The row shows custom, and Resolves to shows your models. To go back, select Use default on the row. Each Proxium server applies the change soon after.
Proxium suggests a ready-made routing for the vendors of your project. It holds the models of each tier, and a classifier that chooses the tier for auto. Proxium writes nothing until you choose it, and a tier that you set yourself keeps your models. See Automatic routing.
Set the reasoning level of a model
Some models can reason before they answer. You can set how much, for each model in a tier.
- Open Routing, and select Change on the row of the tier.
- Beside the model, select a level:
reasoning: none,minimal,low,mediumorhigh. The list shows only for a model that its vendor reports as a reasoning model. - Select Save for this project.
Proxium sends the level as reasoning_effort to that model. It does this only when your request has no reasoning_effort of its own: your value wins. Proxium sends the level only to a vendor that uses the OpenAI API format. reasoning: model default sends no level.
Use ready-made routing
- Open Routing. The Ready-made routing section shows the preset of the project, its version and the models of each tier.
- To change the preset, select it in the list and select Use this preset. The list holds the presets of your vendors only.
- To make the models your own, select Copy and edit. A new version of the preset then no longer changes them.
- To go back to the preset, select Reset to the preset.
Reset to the preset replaces the models of every tier of the preset, the tiers that you set yourself included.
The API has the same actions: GET /api/teams/{slug}/routing/presets, and POST /api/teams/{slug}/routing/preset with {"action": "use", "preset": "anthropic"}, {"action": "copy"} or {"action": "reset", "preset": "anthropic"}.
Give one app its own models
An app names itself with the x-proxium-source header. A rule for that name gives the app its own models for a tier. The console calls this rule A sender.
- Open Routing. In Scoped rules, select Add a rule.
- In Applies to, select A sender.
- In Sender name, type the name that the app sends, for example
inbox-worker. - Select the tier, pick the models, and select Save rule.
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-H "x-proxium-source: inbox-worker" \
-d '{"model": "standard", "messages": [{"role": "user", "content": "hello"}]}'
To give one key its own models, select A key in Applies to. An endpoint makes such a rule for you.
Which models a tier name gets
Proxium asks these questions in order. The first "yes" decides:
Example: the app inbox-worker sends model: "standard" with the key of the endpoint chat-cheap. Proxium asks these questions in order, and stops at the first "yes":
- Does the app have a rule for
standard? A rule for the value ofx-proxium-source, on Routing › Scoped rules › A sender. No. - Does the key have a rule for
standard? An endpoint makes one, or A key on Routing › Scoped rules. Yes: the call gets the models of the key rule,small-modelthenlarge-model. Proxium stops here. - Does the project have its own models for
standard? Set on Routing › Change. Asked only when no rule above matched. - Is
standarda default tier of Proxium, such asstandard,heavyorcode? Asked last.
If every answer is no, Proxium uses standard as a model id.
The tier auto
A call with model: "auto", axon/auto, or no model asks Proxium to choose the tier. auto has no models of its own:
- If the project has a classifier, the classifier names a tier, and the call gets the models of that tier.
- If the project has no classifier, the call gets the models of
standard.
A project has a classifier with a preset, with its own vendor in the classifier tier, or with the grant of the classifier of Proxium. Proxium refuses a chain for the chat tier auto. In a project with the grant, the classifier also chooses the tier of a call that names a text tier. A call that names a model keeps it. See Automatic routing.
What happens when a model fails
Proxium makes a plan of a few models: the models of the tier, then other models of your project. It calls them in order, until one answers. Your failover policy sets how this works. The example uses the default policy.
Example of one request, with the default policy:
- The first call goes to the first model, which answers
503. A 5xx, 408 or "warming" answer gets one more call to the same model. - The next call goes to the first model again, after a short wait. It answers
503again. - The next call goes to the second model of the plan. It answers
200. - Your app gets the answer of the second model. Each call is a row of the Attempt log, numbered in Try.
Your policy sets how many models a plan holds, and how many extra calls each model gets. If every call fails, your app gets 503 upstream_failed.
| Question | Answer with the default policy |
|---|---|
| How many times does Proxium call one model? | Once. After a 5xx, 408 or "warming" answer, once more |
| How long does it wait before that call? | A short time, longer for each next call |
| What after a 429 of the vendor? | The next call goes to a model at another vendor of the plan. The other models of that vendor stay in the plan, after it. With no other vendor in the plan, Proxium waits for the Retry-After of the vendor, up to a limit, and calls the same model again |
| What after a 400 or a 401 of the vendor? | No second call. Proxium goes to the next model |
What if the answer is empty, cut at your max_tokens? | Your app gets it as it is, with finish_reason: "length", and Proxium calls no other model: each model gets the same budget. Raise max_tokens. A reasoning model can spend the whole budget before it writes any content |
| What if the answer is empty for another reason? | Proxium calls the model once more, then the next model |
| What if every call fails? | Your app gets 503 upstream_failed. The message gives the reason of the last call |
| Does Proxium send the request again later? | No. Your app can send it again |
The image, speech, transcription and video routes use the same plan, the same retirements and x-proxium-timeout-ms. They call each model of the plan once, with no extra call to the same model. See Generate images, speech and video.
| Can a stream move to another model? | Only before its first piece reaches your app |
| Can I limit the time of all the calls? | Yes. Send x-proxium-timeout-ms |
The Attempt log on Requests shows each vendor call as a row, numbered in Try. See Debug a failed call.
With the catalog fill on, the plan can reach a model that is not in your tier, at its own price. Check the price of each model in Routing › Model catalog, or turn the fill off in your failover policy.
Set your failover policy
Each project has a failover policy. It decides how many times Proxium calls one model, how long it waits, and what a 429 does. It also sets the size of a plan, and when a model or a whole vendor is retired. Until you save your own, your project uses the defaults. The policy sets:
- Extra calls to the same model, after a 5xx, a 408 or a timeout of the vendor, and the wait before each one.
- What a 429 does: the models of other vendors first, or wait and call the same model.
- The longest wait for the
Retry-Afterof a vendor. - How many models a plan holds.
- Whether the plan is filled with catalog models outside the tier. On by default.
- Whether a model that keeps failing is retired, and for how long. On by default.
- Whether a vendor is retired at once when your own key for it has no credits or is rejected. On by default.
- Whether a vendor whose models keep failing is retired, and for how long. On by default.
- Whether each retirement sends an alert. On by default.
To change the policy:
- Open Routing, and find When a call fails.
- Change the settings.
- Select Save. Each Proxium server applies the policy soon after.
- To go back to the defaults, select Reset to default.
An agent can read and change the policy with the MCP tools proxium_read_failover_policy and proxium_set_failover_policy. See Connect an MCP client.
With the catalog fill off, a call whose tier has no model that your project can use answers 400 no_route. Proxium checks this before the cache and the limits, so the call uses no call slot.
Retire a model or a vendor
A retired model or vendor is skipped by the calls of your project for a while. The next model of the plan answers. When the time ends, the model or vendor is back, and its first call is the test call. Proxium does not change your tiers: Improve names what was retired, and you remove, replace or fix it.
Your failover policy retires:
| What | When |
|---|---|
| A model | When it keeps failing |
| A vendor | At once, when the vendor answers that your own key has no credits or is rejected |
| A vendor | When its models keep failing |
A failed call is one of these:
- a 5xx, a 408, a 404 or a 410;
- a connection error, or a timeout of the vendor;
- a model name that the vendor does not list;
- an empty answer: no content and no tool call. An empty answer that the vendor cut at your
max_tokensdoes not count: the budget is yours.
These do not count: a 429, a "warming" answer, and the end of your own x-proxium-timeout-ms. A success starts the count of the model and of its vendor again. The counts and the retirements are per project, and every Proxium server shares them.
"No credits" or "key rejected" retires a vendor only for a key of your own. That is your key for a known vendor, or a vendor that you added. On a key of Proxium, the problem is Proxium's, and the operator gets the alert.
If every model of the plan is retired, Proxium calls them as if none is retired. A project always has a model to call, and a retirement never answers no_route.
To retire a model or a vendor by hand, or to bring it back:
- Open Requests, and find Model health.
- On the row of a model, select Retire for the model, or Retire vendor for every model of its vendor. Proxium retires it for the minutes of your policy.
- To bring it back at once, select Bring back on its row in Retired now.
An agent can do the same with the MCP tools proxium_retire_model, proxium_restore_model, proxium_retire_vendor and proxium_restore_vendor. The console and the tools call these routes, as the person who signed in:
| Route | Body | Does |
|---|---|---|
GET /api/teams/{slug}/failover | The policy, its defaults and limits, what is retired now, and the recent retirements | |
PUT /api/teams/{slug}/failover | The whole policy, as GET answers it | Saves the policy |
DELETE /api/teams/{slug}/failover | Goes back to the default policy | |
POST /api/teams/{slug}/failover/retire | provider, and optional model and minutes | Retires one model. Without model, the whole vendor. Without minutes, until it is brought back |
POST /api/teams/{slug}/failover/restore | provider, and optional model | Brings one model back, or the whole vendor without model, and starts its count again |
When a vendor answers 410 Gone, Proxium reads the model list of the vendor again. Your tiers stay as you set them. The failure counts toward the policy, and Improve names the model to remove or replace.
When Proxium skips a vendor key
Each vendor key has a circuit breaker: each key of Proxium, and each key of your own. It works below your failover policy. On a key of Proxium, it protects every project from a vendor that keeps failing. On a key of your own, it parks one rejected key at the first failure:
| The vendor | Proxium |
|---|---|
| Keeps failing | Skips that key for a short time, then sends a trial call |
| Rejects the key, or says that the quota is used up | Skips that key for a longer time, then sends a trial call |
| Answers the trial calls | Calls it again as normal |
| Answers 429 | Does not skip it. A 429 never opens the breaker |
A skipped key shows as circuit_open in the Outcome column of Requests. To read the state of each vendor before a call, send GET https://proxium.tech/status/peek. See Set budgets and limits.