Skip to main content

Get started with Proxium

Who calls Proxium, and where each call goes:

  • Who calls: a developer with curl or the OpenAI SDK, and your app with the OpenAI or Anthropic SDK. A coding agent such as Claude Code, Codex or Cursor calls over MCP.
  • Proxium: one base URL and one virtual key for all of them.
  • LLM APIs, with your keys: OpenAI, Anthropic, Groq, OpenRouter, Mistral, DeepSeek, Together AI, Google Gemini, or your own server.
  • Memory: coding agents read and write the project memory over MCP.
  • MCP tools: agents call the tools of your project's MCP servers through the same MCP URL, with one sign-in. See Connect MCP servers.

Proxium is an LLM gateway. Your apps call it with the OpenAI API or the Anthropic API, and one virtual key. Proxium calls your provider with your own provider key. If a model fails, Proxium tries the next model of the chain. It records the cost of each call.

Setup​

To connect an app, set the base URL to https://proxium.tech/v1 and the API key to your virtual key.

app.py
import os
from openai import OpenAI

client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
)

answer = client.chat.completions.create(
model=os.environ["PROXIUM_MODEL"],
messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(answer.choices[0].message.content)

PROXIUM_KEY is a virtual key of your project. PROXIUM_MODEL is a model id from GET /v1/models, or a tier name of your routing policy. The Quickstart shows how to get both.

Choose your path​

What Proxium does with a chat call​

  1. It finds the project of the virtual key.
  2. It resolves the model field to a chain of models. See Routing.
  3. It looks for the answer in the response cache. A cache hit makes no provider call and costs nothing by default. See The response cache.
  4. It counts the call against the ceilings that you set for the calling application. See Budgets and limits.
  5. It calls the models of the chain in order, with your provider key, until one answers.
  6. It records the cost: the token counts from the answer, times the price of the model. See Track spend.

The full order, with the reason for each step, is in The request flow.

See your spend and your failures​

The console turns each call into dashboards:

ScreenPanels
OverviewSpent, Remaining, Saved by cache, Spend over time, Model split, Per-app burn, Per-key burn
RequestsAttempt log, What failed, Who is unreliable, Model health, Cache
ImproveWhat to change: a cheaper model of the same vendor, a vendor that fails often, an app that the cache does not help

Track spend explains each number. To get an alert in Slack, Discord, Telegram or a webhook when the spend passes a level, see Get spend alerts.