The request flow
This page follows one chat call through Proxium, and explains why the steps come in this order. It applies to /v1/chat/completions and to /v1/messages. Proxium translates a Messages call to the chat shape at the start, and the answer back at the end.
The steps
One chat call passes these steps in order:
- Find the project from the virtual key, and read the body. A wrong key stops here with
401, and a body that is not JSON with415or400 bad_json. - Scan the request for sensitive data, and flag, redact or block it. A block stops here with
403 guardrail_blocked. - Add memories, only with
x-proxium-memory: recall. A wrong value of the header stops here with400 invalid_memory_mode. - For
model: "auto", or a call with nomodel, a classifier chooses a tier. A project with no classifier uses the tierstandard. In a project with the grant of the classifier of Proxium, a call that names a text tier is classified too. - Turn the tier into a plan of models: app rule, key rule, project, default, then the catalog fill. Proxium leaves out the models and vendors that your project retired. A request can send an image, audio or video, or ask for audio or an image. Then the catalog fill skips the models that cannot do it. A call with no model stops here with
400 no_route, and uses no call slot. - Look in the response cache. A hit answers now with the stored answer, with no vendor call.
- Check the limits of the key and the app. Over a limit stops here with
429. - Call the models of the plan, by your failover policy, until one answers. If all fail,
503 upstream_failed. - Record the cost, and store the answer in the cache. The answer goes to your app.
| Step | What Proxium does | A failure here gets |
|---|---|---|
| 1. Authenticate | Reads the virtual key, finds its project, and reads the body | 401 missing_key or 401 invalid_key; 415 or 400 bad_json for the body |
| 2. Scan for sensitive data | Flags, redacts or blocks each class that the project set on Data protection | 403 guardrail_blocked |
| 3. Add memories | With x-proxium-memory: recall, adds the matching memories to the messages | 400 invalid_memory_mode for a wrong header value. A recall error stops nothing: the call continues without memories |
4. Classify auto | For model: "auto" or no model, in a project with a classifier, a classifier chooses a tier. With the grant, also for a named text tier | Nothing. A call for auto uses the tier standard, and a call that named a tier keeps it |
| 5. Build the plan | Turns model into an ordered plan of models, and leaves out the retired models and vendors | 400 no_route, when no model exists for the call. A retirement never causes it: when every model of the plan is retired, the plan keeps them |
| 6. Look up the cache | Looks for the same request in the response cache | Nothing. An error counts as a miss |
| 7. Check the ceilings | Reserves one call in the windows of the key and the application | 429 tenant_capped or 429 source_capped |
| 8. Call the models | Calls the models of the plan by the project's failover policy, until one answers | 503 upstream_failed |
| 9. Record the cost | Prices the tokens of the answer, stores the answer in the cache | Nothing. The answer still goes to your app |
Why this order
The key comes first
Proxium reads the key before it reads the body. A caller with no valid key gets 401, not an error about the body. The project comes from the key only. No header can name a different project.
The scan comes before everything that stores or sends
The scan runs before the memories, the cache, the stored prompts and the vendor. A redacted or blocked value is therefore never written anywhere. See Sensitive data scanning.
Memories come before the cache
A recall changes the messages of the call. The cache key holds the messages, so the memories must be in them before the lookup. Two end users with different memories therefore never share a cached answer.
The classifier comes before the chain
The classifier turns auto into a tier name, such as heavy. Then the chain of that tier resolves like any other tier. A project with no classifier skips this step, and auto becomes standard. The classifier is a real model call, and Proxium records its cost against your project. A key that is already at a limit skips it, because Proxium refuses that call anyway.
The plan comes before the cache and the ceilings
Proxium builds the plan before it looks in the cache and before it takes a call slot. A call for which no model exists answers 400 no_route at once. It gets no cached answer and uses no call slot.
The cache key holds the first model of the resolved tier. It also holds the whole body of the request, with the model field that you sent. A cached answer serves that request, whichever model of the plan made it. A retirement does not remove cached answers: each one expires on its own.
The cache comes before the ceilings
A cache hit calls no vendor, so by default it costs nothing and uses no call. A project at its ceiling still gets its cache hits.
The operator of Proxium can set a price for the hits of a project, as a part of the saved cost. Then a hit goes through the ceilings too, because it is a paid request. See The response cache.
The ceilings come before the vendor call
A refused call costs nothing at the vendor. All Proxium servers share the windows, and the check and the reservation are one step. Two servers cannot both let through a call that only one slot allows.
The vendor key comes last
Proxium adds your vendor key on its server, at each attempt. Your app never holds the vendor key. For each model of the plan, Proxium checks three things before the call:
| Check | Result |
|---|---|
The time limit of x-proxium-timeout-ms ran out | Proxium stops and answers 503 |
| The circuit breaker of the vendor key is open | Proxium skips that model |
| The key may not use the model | Proxium drops it from the plan |
| The model or its vendor is retired for your project | Proxium left it out of the plan at step 5 |
A plan has a few models, by your failover policy. Route requests to models shows the extra calls and the fallback.
The cost comes from the answer
Proxium prices the token counts that the answer reports, with the price of the model that answered. After a failover, that is the second model, not the model you asked for. A streamed call gets its cost when the stream ends. See Track spend.
Other routes
/v1/embeddings, /v1/responses, /v1/moderations and /v1/rerank take a shorter path:
| Step | Chat and Messages | The other four |
|---|---|---|
| Authenticate | Yes | Yes |
| Memories | recall adds memories | No |
| Response cache | Yes | No |
Plan, and 400 no_route | Before the cache lookup | Before the ceilings |
| Ceilings | After the cache lookup | After the plan |
| Failover | Yes | Yes |
| Cost record | Yes | Yes |
Not active on proxium.tech
The code has four more steps that proxium.tech does not run today:
| Step | What it would do |
|---|---|
| Near-match cache | Serve the answer of a similar request, found by an embedding |
| Failover price guard | Skip a fallback model that costs much more than the first |
| Monthly credits | At step 7, refuse a project that used its monthly credits, with 429 credits_exhausted. No project on proxium.tech has a credit grant |
| Plan check | At step 7, refuse a project with no paid plan, with 402 subscription_required. proxium.tech is in open beta |
Without the price guard, a fallback model can cost more than your first model.