Skip to main content
POST
Create a generation
Generates images, videos, voice-overs, music or 3D models and charges credits. Pick a model_id from List models. Generation is asynchronous: the response contains the created assets with status queued or processing. Set wait_seconds (max 55) to wait for the result, or poll Get a generation. Required permission: generate. MCP tool: generate.
string
required
image, video, audio or 3d.
string
required
Model id from List models. Must belong to the modality (and track).
string
required
What to create (max 5000 characters). For voice: the exact text to speak.
string
Audio only: voice (default) or music.
integer
default:"1"
Number of outputs, images only (1 to 4). Each output is charged.
string
Image or video aspect ratio, e.g. 1:1, 9:16, 16:9.
string
Image resolution (e.g. 1k, 2k) or video resolution (e.g. 720p, 1080p).
integer
Video duration in seconds (see options.durations_seconds).
boolean
Video: generate native sound when the model supports it.
string
Voice id for text to speech (see options.voices).
string
Voice language code, e.g. fr, en.
integer[]
Library image ids used as references (max 4).
string[]
Public https image URLs or data: URIs used as references (max 4, JPEG/PNG/WEBP, 10 MB).
integer
default:"0"
Wait up to N seconds (0 to 55) for completion before responding.