status starts at queued, moves to running, and ends at succeeded, failed or canceled. Only a succeeded generation is charged.
Start one
POST /v1/generations takes the model and its input. Every model takes the same request shape; only input differs, as the model’s input schema says.
201 Created with the generation, queued, and its path in the Location header.
Retry safely
Send anidempotency-key you generate for each new request, such as a UUID. If the network drops and you retry with the same key and body, you get the same generation back, marked with Idempotent-Replayed: true, and no second run starts. Reusing a key with a different body is a 409 with the code conflict. Keys belong to your workspace and don’t expire, so never reuse one for a new request.
Get the result
Pick by how long the model takes and where your code runs:Wait in the request
prefer: wait=N holds the request open for up to N seconds (at most 60) until the run finishes, and the Preference-Applied header echoes the wait used. Most images and speech finish inside the wait. If the run is still going when the wait ends, you get it back queued or running, and you carry on with a long-poll.
Long-poll
GET /v1/generations/{id} takes the same prefer: wait=N. It answers as soon as the run reaches succeeded, failed or canceled, or after N seconds with the run as it is, whichever comes first. Either way the status is 200: read status to tell which. Loop until it’s final.
429 and 5xx:
Webhooks
To hear when a run ends without holding anything open, name a URL on the request:generation.succeeded, generation.failed or generation.canceled event to that URL, with the generation inside. Each request in a batch can name its own webhook. To get every run’s events at one URL instead, set up a webhook endpoint. Webhooks covers both, and how to verify the signature. Keep a long-poll as a fallback, to catch a run whose event you missed.
Poll
Some clients can’t hold a request open for a minute. They can read the generation withoutprefer and get an answer at once. While the run isn’t final, the answer carries a Retry-After header with the seconds to wait before the next read: 2 for images and audio, 5 for video and 3D. A final answer has none. The answer to POST /v1/generations carries it too.
Polling counts against your rate limit like any request, so follow Retry-After rather than reading faster.
The generation object
When a run fails
A run that fails is still a200 when you read it: its status is failed, error.message says why (the model refused the prompt, or the provider had an error), and you aren’t charged. Check status before you read output.
What to do next depends on the reason:
- When the model’s safety filter turned the run down,
error.messagesays so and names what to change: the prompt, or the photo you sent. The same input would most likely be turned down again, so change it, or try another model. - Any other failure, such as a provider error or a run that took too long, may succeed if you start it again as it was.
idempotency-key. Sending the old key with the same body returns the same failed generation, and no new run starts.
Cancel
canceled, its hold is released, and if the provider finishes it anyway, the result is thrown away and nothing is charged. A run that has already ended comes back as it is.
Two kinds of run can’t be stopped once they’ve started, and get a 409 with the code conflict: a run at a provider that bills it either way, and a preset run that has already shown its preview still. Each runs to its end and is charged only if it succeeds.
List your runs
Batches
POST /v1/batches starts up to 50 generations in one request. They’re priced and held together, so either all start or none do, and each then runs as its own generation.
201 with {"id": "bat_...", "object": "batch", "data": [...]}, holding each generation in the order you sent them. Each one carries the batch’s ID in batch_id; read it at GET /v1/generations/{id} like any other generation.
Read the whole batch with GET /v1/batches/{id}. It takes prefer: wait=N too, and then answers once every generation in it is final, or after N seconds. While any isn’t final, the answer carries Retry-After. A webhook endpoint also gets one batch.completed event when every run in the batch has ended.