Use the OpenAI SDK
Proxium speaks the OpenAI API, so an app that uses the OpenAI SDK needs no new library. You change three values:
| Value | Set it to |
|---|---|
| Base URL | https://proxium.tech/v1 |
| API key | A virtual key of your project, from Keys in the console |
| Model | A model id or a tier name of your project. Route requests to models explains both |
The examples read the key from PROXIUM_KEY and the model from PROXIUM_MODEL. The Quickstart sets both.
Connect your app
- Python
- Node
- curl
import os
from openai import OpenAI
client = OpenAI(
base_url="https://proxium.tech/v1",
api_key=os.environ["PROXIUM_KEY"],
)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://proxium.tech/v1",
apiKey: process.env.PROXIUM_KEY,
});
curl https://proxium.tech/v1/models \
-H "Authorization: Bearer $PROXIUM_KEY"
Send a request
A request is a list of messages. The model reads them and writes one reply.
- Python
- Node
- curl
response = client.chat.completions.create(
model=os.environ["PROXIUM_MODEL"],
messages=[
{"role": "system", "content": "You answer in one sentence."},
{"role": "user", "content": "What is a virtual key?"},
],
)
print(response.choices[0].message.content)
const response = await client.chat.completions.create({
model: process.env.PROXIUM_MODEL,
messages: [
{ role: "system", content: "You answer in one sentence." },
{ role: "user", content: "What is a virtual key?" },
],
});
console.log(response.choices[0].message.content);
curl https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"messages\": [{\"role\": \"system\", \"content\": \"You answer in one sentence.\"}, {\"role\": \"user\", \"content\": \"What is a virtual key?\"}]}"
The answer is the vendor's answer, in the OpenAI shape:
| Field | Holds |
|---|---|
choices[0].message.content | The reply |
model | The model that answered |
usage | The token counts. Proxium prices the request from them |
If the first model fails, Proxium sends the request to the next model of its chain. Route requests to models shows the order.
Stream the reply
Without streaming, your app waits until the model has written the whole reply. With streaming, the reply arrives in small pieces while the model writes it. Use streaming to show the reply as it grows, for example in a chat window.
To stream, set stream to true. Proxium sends the pieces as server-sent events, in the OpenAI format.
- Python
- Node
- curl
stream = client.chat.completions.create(
model=os.environ["PROXIUM_MODEL"],
messages=[{"role": "user", "content": "Write a haiku about queues."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()
const stream = await client.chat.completions.create({
model: process.env.PROXIUM_MODEL,
messages: [{ role: "user", content: "Write a haiku about queues." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
process.stdout.write("\n");
curl -N https://proxium.tech/v1/chat/completions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$PROXIUM_MODEL\", \"stream\": true, \"messages\": [{\"role\": \"user\", \"content\": \"Write a haiku about queues.\"}]}"
The last piece can carry only the token counts, with an empty choices list. Proxium asks the vendor for them, to price the request. The examples check choices before they read it.
Handle errors
When Proxium refuses a request, or every model fails, your app gets an error. The SDK raises it as an exception. The body names the reason in code:
{
"error": {
"message": "invalid or revoked key",
"type": "invalid_key",
"code": "invalid_key"
}
}
The errors you meet first:
| Status | code | Cause | Fix |
|---|---|---|---|
| 401 | missing_key | The request has no Authorization: Bearer header | Set the API key of the client |
| 401 | invalid_key | The key is wrong or revoked | Create a new key on Keys |
| 400 | no_route | Your project has no vendor that can serve the request. The message is no models available for route | Add a vendor on Providers |
| 429 | source_capped | Your app reached a ceiling that you set | Wait the seconds in the retry-after header |
| 503 | upstream_failed | Every model of the chain failed. The message gives the last reason | Check the vendor on Requests, in What failed |
- Python
- Node
import openai
try:
response = client.chat.completions.create(
model=os.environ["PROXIUM_MODEL"],
messages=[{"role": "user", "content": "Hello"}],
)
except openai.APIStatusError as e:
print(e.status_code, e.code, e.message)
try {
await client.chat.completions.create({
model: process.env.PROXIUM_MODEL,
messages: [{ role: "user", content: "Hello" }],
});
} catch (e) {
console.log(e.status, e.code, e.message);
}
Errors lists every code.
Next steps
- Route requests to models: model ids, tiers and the fallback order.
- Track spend: name each app, and see what it costs.