From your app to the answer

Every call passes these steps. A step that stops a bad call stops it before a vendor bills you for it.

  1. Your key

    The virtual key finds the project. A vendor key never leaves Proxium.

    401 wrong key
  2. Scan

    Card numbers and secrets are redacted or blocked before anything stores or sends them.

    403 blocked
  3. Route

    The tier becomes an ordered list of models. With auto, a classifier picks the tier.

  4. Cache

    The same question again is answered from the cache, with no call to a vendor.

    no vendor call
  5. Limits

    The limits of the key and of the app are checked before any vendor call.

    429 limit
  6. Call

    Proxium adds your vendor key and calls the models in order, until one answers.

    503 all failed
  7. Answer

    Your app gets one answer, in the API it called.

The right model for each request

Keep the model names your app sends today: each call goes to that model, with failover behind it. To let Proxium choose, send auto as the model. A classifier reads each request and picks the tier: a greeting goes to a cheap model, a long coding task to a strong one. Choose a ready-made routing for your vendors, or set the models of each tier yourself.

"model": "auto"
→classifier reads the request
→tier heavy
→the models of heavy, in order

One vendor down is not your app down

When a model fails, Proxium calls it once more, then the next model of the plan. After a 429, a model at another vendor goes first. Proxium also reads each failure and retires what keeps failing: a model after 3 failed calls in a row, a whole vendor at once when your key for it has no credits or is rejected. Every number and rule is yours to change, and you can retire or bring back a model by hand. Your app sees one answer.

call 1·model A→503
call 2·model A→503
call 3·model B→200
→your app gets the answer of model B
vendor key rejected→vendor retired

The cost of every call, and a limit for each app

Proxium reads the tokens from each vendor answer and prices them at the model’s list price, for each app and each key. A repeated question is answered from the cache, with no call to a vendor. Give an app a limit in calls per hour or dollars per day, and Proxium refuses its calls above it, before a vendor is called.

support-bot·1,200 tokens→$0.00036
the same question→cache·no vendor call
limit·$5 a day→429 above it

Secrets stop at Proxium

Every request is scanned before it leaves: card numbers, bank accounts, national IDs and API keys. Choose for each kind what Proxium does: flag it, redact it or block the call. A redacted value never reaches the cache, the stored prompts or the vendor.

you send
Refund card 4242 4242 4242 4242
the vendor gets
Refund card [REDACTED:credit_card:1]

No code changes

Point your AI endpoint at https://proxium.tech/v1 with your Proxium key, and Proxium takes care of the rest. Each call goes to any of your providers.

from openai import OpenAI client = OpenAI( base_url="https://proxium.tech/v1", api_key=os.environ["PROXIUM_KEY"], )

Where it runs

CloudSTANDARD

Fully managed — the standard way to run Proxium. We manage the infrastructure for you.

Self-hostedENTERPRISE ONLY

On request, for enterprises that need it in their own infrastructure — any provider, including local models.

Start right away.

Point your existing SDK at one URL. Everything else — routing, caching, failover — is already on.