Generate images, speech and video, and transcribe audio
Proxium serves media APIs with your virtual key. When the model or the tier has more than one vendor, Proxium calls the next one if a vendor fails.
| Route | Makes | Answer |
|---|---|---|
POST /v1/images/generations | One image | JSON, with the image as base64 in data[0].b64_json |
POST /v1/audio/speech | Speech from text | The audio bytes, audio/mpeg by default |
POST /v1/audio/transcriptions | Text from audio | JSON, the provider's answer |
POST /v1/videos | A video | A video object to read until it is finished |
POST /video/generations | A video | An operation to poll |
The video routes have no /v1. Their URL is https://proxium.tech/video/….
Before you start
Your project needs a vendor of the right request type: Images, Speech, Transcription or Video. To add one, select that Request type when you add a server. See Add vendors and your own keys.
In model, send a model id of that vendor, such as t.acme.my-images/flux.1-schnell. A tier name also works, if Routing has a chain for it.
Generate an image
curl https://proxium.tech/v1/images/generations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.my-images/flux.1-schnell", "prompt": "a lighthouse at dawn", "size": "1024x1024"}' \
| jq -r '.data[0].b64_json' | base64 -d > lighthouse.png
| Field | Default |
|---|---|
prompt | — |
size | 1024x1024 |
n | 1 |
Proxium returns the image in this shape for each vendor.
Generate speech
curl https://proxium.tech/v1/audio/speech \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.my-voice/tts-1", "input": "Your order has shipped.", "voice": "alloy"}' \
--output order.mp3
voice is alloy when you leave it out.
Transcribe audio
curl https://proxium.tech/v1/audio/transcriptions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-F model="t.acme.my-voice/whisper-1" \
-F file="@call.wav"
| Field | Default |
|---|---|
file | — |
model | — |
language | — |
prompt | — |
response_format | json |
temperature | — |
Proxium records the seconds of audio that the vendor reports. If the model has a price for each minute of audio, the cost of the call comes from those seconds. A file over the body limit gets 413, and a body that is not valid multipart form data gets 400 bad_multipart.
Generate a video
A video takes minutes. Proxium serves the OpenAI video API, so the OpenAI SDKs work with your Proxium key:
from openai import OpenAI
client = OpenAI(base_url="https://proxium.tech/v1", api_key=PROXIUM_KEY)
video = client.videos.create_and_poll(
model="t.acme.my-video/veo-3.0-generate-001",
prompt="waves on a rocky shore",
seconds="8",
size="1280x720",
)
client.videos.download_content(video.id).write_to_file("waves.mp4")
The video vendor can be a Veo server or a video job API, such as the one that serves the video models of an aggregator. A model id of a job API can hold a / after the vendor, for example t.acme.my-videos/vendor/model-1.
| Route | Does |
|---|---|
POST /v1/videos | Starts a video. The answer has id and status queued |
GET /v1/videos/{id} | Reads the video. status is queued, in_progress, completed or failed |
GET /v1/videos/{id}/content | Downloads a finished video |
GET /v1/videos | Lists the videos of your project |
DELETE /v1/videos/{id} | Removes a video from the list |
| Field | Default |
|---|---|
prompt | — |
model | The default video vendor |
seconds | 8 |
size | 1280x720. Also 720x1280, 1920x1080 and 1080x1920 |
A video is recorded for the project that made it. Another project gets 404 for it, also with its id. When a video is finished, Proxium records its cost once, from its seconds and the price of the model.
Proxium does not serve input_reference, remix or edits yet.
Each read of an unfinished video and each download uses one call of your ceilings. Do not poll more often than you need.
The video routes with no /v1
The older routes do the same job in three steps:
- Start the video:
curl https://proxium.tech/video/generations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "video", "prompt": "waves on a rocky shore", "aspect_ratio": "16:9", "duration": "8"}'
The answer has operation and provider.
- Poll the operation with both values, until it is done:
curl https://proxium.tech/video/operations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"operation": "<operation>", "provider": "<provider>"}'
When done is true, the answer has video_uri.
- Download the video:
curl https://proxium.tech/video/download \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"uri": "<video_uri>", "provider": "<provider>"}' \
--output waves.mp4
A poll or a download of an operation or a file that your project did not start gets 404 not_found. The provider value names the vendor that started the video. Proxium refuses an operation or a uri that is not on a Google API host or on the vendor's own host, with 400 forbidden_target.
How the media routes behave
| Media routes | |
|---|---|
| Failover to another model | Yes. Proxium calls the models of the tier or of the routing rule in order, until one answers. Your failover policy sets the size of the plan and retires a model that keeps failing |
x-proxium-timeout-ms | Yes. It limits all the calls of the plan together |
| Allowed models of the key | Yes. A model that the key may not use is left out of the plan |
| Response cache | No |
| Row in the Attempt log of Requests | Yes, one row for each call. A failed call also shows in What failed. The drill-down stores the request and an error text, never the image, the audio or the video |
| Cost | From the usage block of the vendor, when it sends one. Else the call is recorded with no cost |
Related pages
- API reference: every field of each route.
- Errors:
no_route,bad_modeland the other media errors.