Skip to main content

The request flow

This page follows one chat call through Proxium, and explains why the steps come in this order. It applies to /v1/chat/completions and to /v1/messages. Proxium translates a Messages call to the chat shape at the start, and the answer back at the end.

The steps​

Fig. 1 · one chat call, step by step

One chat call passes these steps in order:

  1. Find the project from the virtual key, and read the body. A wrong key stops here with 401, and a body that is not JSON with 415 or 400 bad_json.
  2. Scan the request for sensitive data, and flag, redact or block it. A block stops here with 403 guardrail_blocked.
  3. Add memories, only with x-proxium-memory: recall. A wrong value of the header stops here with 400 invalid_memory_mode.
  4. For model: "auto", or a call with no model, a classifier chooses a tier. A project with no classifier uses the tier standard. In a project with the grant of the classifier of Proxium, a call that names a text tier is classified too.
  5. Turn the tier into a plan of models: app rule, key rule, project, default, then the catalog fill. Proxium leaves out the models and vendors that your project retired. A request can send an image, audio or video, or ask for audio or an image. Then the catalog fill skips the models that cannot do it. A call with no model stops here with 400 no_route, and uses no call slot.
  6. Look in the response cache. A hit answers now with the stored answer, with no vendor call.
  7. Check the limits of the key and the app. Over a limit stops here with 429.
  8. Call the models of the plan, by your failover policy, until one answers. If all fail, 503 upstream_failed.
  9. Record the cost, and store the answer in the cache. The answer goes to your app.
StepWhat Proxium doesA failure here gets
1. AuthenticateReads the virtual key, finds its project, and reads the body401 missing_key or 401 invalid_key; 415 or 400 bad_json for the body
2. Scan for sensitive dataFlags, redacts or blocks each class that the project set on Data protection403 guardrail_blocked
3. Add memoriesWith x-proxium-memory: recall, adds the matching memories to the messages400 invalid_memory_mode for a wrong header value. A recall error stops nothing: the call continues without memories
4. Classify autoFor model: "auto" or no model, in a project with a classifier, a classifier chooses a tier. With the grant, also for a named text tierNothing. A call for auto uses the tier standard, and a call that named a tier keeps it
5. Build the planTurns model into an ordered plan of models, and leaves out the retired models and vendors400 no_route, when no model exists for the call. A retirement never causes it: when every model of the plan is retired, the plan keeps them
6. Look up the cacheLooks for the same request in the response cacheNothing. An error counts as a miss
7. Check the ceilingsReserves one call in the windows of the key and the application429 tenant_capped or 429 source_capped
8. Call the modelsCalls the models of the plan by the project's failover policy, until one answers503 upstream_failed
9. Record the costPrices the tokens of the answer, stores the answer in the cacheNothing. The answer still goes to your app

Why this order​

The key comes first​

Proxium reads the key before it reads the body. A caller with no valid key gets 401, not an error about the body. The project comes from the key only. No header can name a different project.

The scan comes before everything that stores or sends​

The scan runs before the memories, the cache, the stored prompts and the vendor. A redacted or blocked value is therefore never written anywhere. See Sensitive data scanning.

Memories come before the cache​

A recall changes the messages of the call. The cache key holds the messages, so the memories must be in them before the lookup. Two end users with different memories therefore never share a cached answer.

The classifier comes before the chain​

The classifier turns auto into a tier name, such as heavy. Then the chain of that tier resolves like any other tier. A project with no classifier skips this step, and auto becomes standard. The classifier is a real model call, and Proxium records its cost against your project. A key that is already at a limit skips it, because Proxium refuses that call anyway.

The plan comes before the cache and the ceilings​

Proxium builds the plan before it looks in the cache and before it takes a call slot. A call for which no model exists answers 400 no_route at once. It gets no cached answer and uses no call slot.

The cache key holds the first model of the resolved tier. It also holds the whole body of the request, with the model field that you sent. A cached answer serves that request, whichever model of the plan made it. A retirement does not remove cached answers: each one expires on its own.

The cache comes before the ceilings​

A cache hit calls no vendor, so by default it costs nothing and uses no call. A project at its ceiling still gets its cache hits.

The operator of Proxium can set a price for the hits of a project, as a part of the saved cost. Then a hit goes through the ceilings too, because it is a paid request. See The response cache.

The ceilings come before the vendor call​

A refused call costs nothing at the vendor. All Proxium servers share the windows, and the check and the reservation are one step. Two servers cannot both let through a call that only one slot allows.

The vendor key comes last​

Proxium adds your vendor key on its server, at each attempt. Your app never holds the vendor key. For each model of the plan, Proxium checks three things before the call:

CheckResult
The time limit of x-proxium-timeout-ms ran outProxium stops and answers 503
The circuit breaker of the vendor key is openProxium skips that model
The key may not use the modelProxium drops it from the plan
The model or its vendor is retired for your projectProxium left it out of the plan at step 5

A plan has a few models, by your failover policy. Route requests to models shows the extra calls and the fallback.

The cost comes from the answer​

Proxium prices the token counts that the answer reports, with the price of the model that answered. After a failover, that is the second model, not the model you asked for. A streamed call gets its cost when the stream ends. See Track spend.

Other routes​

/v1/embeddings, /v1/responses, /v1/moderations and /v1/rerank take a shorter path:

StepChat and MessagesThe other four
AuthenticateYesYes
Memoriesrecall adds memoriesNo
Response cacheYesNo
Plan, and 400 no_routeBefore the cache lookupBefore the ceilings
CeilingsAfter the cache lookupAfter the plan
FailoverYesYes
Cost recordYesYes

Not active on proxium.tech​

The code has four more steps that proxium.tech does not run today:

StepWhat it would do
Near-match cacheServe the answer of a similar request, found by an embedding
Failover price guardSkip a fallback model that costs much more than the first
Monthly creditsAt step 7, refuse a project that used its monthly credits, with 429 credits_exhausted. No project on proxium.tech has a credit grant
Plan checkAt step 7, refuse a project with no paid plan, with 402 subscription_required. proxium.tech is in open beta

Without the price guard, a fallback model can cost more than your first model.