Skip to main content

Generate images, speech and video, and transcribe audio

Proxium serves media APIs with your virtual key. When the model or the tier has more than one vendor, Proxium calls the next one if a vendor fails.

RouteMakesAnswer
POST /v1/images/generationsOne imageJSON, with the image as base64 in data[0].b64_json
POST /v1/audio/speechSpeech from textThe audio bytes, audio/mpeg by default
POST /v1/audio/transcriptionsText from audioJSON, the provider's answer
POST /v1/videosA videoA video object to read until it is finished
POST /video/generationsA videoAn operation to poll

The video routes have no /v1. Their URL is https://proxium.tech/video/….

Before you start​

Your project needs a vendor of the right request type: Images, Speech, Transcription or Video. To add one, select that Request type when you add a server. See Add vendors and your own keys.

In model, send a model id of that vendor, such as t.acme.my-images/flux.1-schnell. A tier name also works, if Routing has a chain for it.

Generate an image​

Terminal
curl https://proxium.tech/v1/images/generations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.my-images/flux.1-schnell", "prompt": "a lighthouse at dawn", "size": "1024x1024"}' \
| jq -r '.data[0].b64_json' | base64 -d > lighthouse.png
FieldDefault
prompt—
size1024x1024
n1

Proxium returns the image in this shape for each vendor.

Generate speech​

Terminal
curl https://proxium.tech/v1/audio/speech \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "t.acme.my-voice/tts-1", "input": "Your order has shipped.", "voice": "alloy"}' \
--output order.mp3

voice is alloy when you leave it out.

Transcribe audio​

Terminal
curl https://proxium.tech/v1/audio/transcriptions \
-H "Authorization: Bearer $PROXIUM_KEY" \
-F model="t.acme.my-voice/whisper-1" \
-F file="@call.wav"
FieldDefault
file—
model—
language—
prompt—
response_formatjson
temperature—

Proxium records the seconds of audio that the vendor reports. If the model has a price for each minute of audio, the cost of the call comes from those seconds. A file over the body limit gets 413, and a body that is not valid multipart form data gets 400 bad_multipart.

Generate a video​

A video takes minutes. Proxium serves the OpenAI video API, so the OpenAI SDKs work with your Proxium key:

video.py
from openai import OpenAI

client = OpenAI(base_url="https://proxium.tech/v1", api_key=PROXIUM_KEY)
video = client.videos.create_and_poll(
model="t.acme.my-video/veo-3.0-generate-001",
prompt="waves on a rocky shore",
seconds="8",
size="1280x720",
)
client.videos.download_content(video.id).write_to_file("waves.mp4")

The video vendor can be a Veo server or a video job API, such as the one that serves the video models of an aggregator. A model id of a job API can hold a / after the vendor, for example t.acme.my-videos/vendor/model-1.

RouteDoes
POST /v1/videosStarts a video. The answer has id and status queued
GET /v1/videos/{id}Reads the video. status is queued, in_progress, completed or failed
GET /v1/videos/{id}/contentDownloads a finished video
GET /v1/videosLists the videos of your project
DELETE /v1/videos/{id}Removes a video from the list
FieldDefault
prompt—
modelThe default video vendor
seconds8
size1280x720. Also 720x1280, 1920x1080 and 1080x1920

A video is recorded for the project that made it. Another project gets 404 for it, also with its id. When a video is finished, Proxium records its cost once, from its seconds and the price of the model.

Proxium does not serve input_reference, remix or edits yet.

warning

Each read of an unfinished video and each download uses one call of your ceilings. Do not poll more often than you need.

The video routes with no /v1​

The older routes do the same job in three steps:

Fig. 1 · The three steps of a video
1POST /video/generationsProxium starts the video. You get operation and provider
2POST /video/operationsSend it again until done is true. Then you get video_uri
3POST /video/downloadYou get the video bytes
  1. Start the video:
Terminal
curl https://proxium.tech/video/generations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "video", "prompt": "waves on a rocky shore", "aspect_ratio": "16:9", "duration": "8"}'

The answer has operation and provider.

  1. Poll the operation with both values, until it is done:
Terminal
curl https://proxium.tech/video/operations \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"operation": "<operation>", "provider": "<provider>"}'

When done is true, the answer has video_uri.

  1. Download the video:
Terminal
curl https://proxium.tech/video/download \
-H "Authorization: Bearer $PROXIUM_KEY" \
-H "Content-Type: application/json" \
-d '{"uri": "<video_uri>", "provider": "<provider>"}' \
--output waves.mp4

A poll or a download of an operation or a file that your project did not start gets 404 not_found. The provider value names the vendor that started the video. Proxium refuses an operation or a uri that is not on a Google API host or on the vendor's own host, with 400 forbidden_target.

How the media routes behave​

Media routes
Failover to another modelYes. Proxium calls the models of the tier or of the routing rule in order, until one answers. Your failover policy sets the size of the plan and retires a model that keeps failing
x-proxium-timeout-msYes. It limits all the calls of the plan together
Allowed models of the keyYes. A model that the key may not use is left out of the plan
Response cacheNo
Row in the Attempt log of RequestsYes, one row for each call. A failed call also shows in What failed. The drill-down stores the request and an error text, never the image, the audio or the video
CostFrom the usage block of the vendor, when it sends one. Else the call is recorded with no cost
  • API reference: every field of each route.
  • Errors: no_route, bad_model and the other media errors.