Audio Understanding
Use the returned task ID to query the task for the final result.
Authorizations
All endpoints require Bearer Token authentication. Add to the request header:
Authorization: Bearer YOUR_API_KEY
YOUR_API_KEY is the API Token (sk-... format).
Body
"gemini-2.5-pro"
Audio source. Accepts one of the following two forms:
- Publicly reachable HTTP/HTTPS URL
data:audio/<type>;base64,<payload>data URI (base64 inline)
Audio format support per family (the specific available models are driven by channel configuration):
- Gemini family (e.g.
gemini-*): wav/mp3/aiff/aac/ogg/flac/m4a; total request body (prompt + system + inline files) ≤ 20 MB
Inline base64 (asynchronous mode only, when sync is omitted or false): Base64 recognized by the gateway as inline media (data: URI, or at least 4096 encoded characters whose decoded content matches a known media signature) is uploaded to file storage and replaced with a URL before the request is persisted or submitted upstream; the original base64 is not stored. Shorter raw base64 or content with an unrecognized signature is passed through unchanged and is not size-validated. Limits apply only to inline base64 items (public URLs are not limited): decoded size ≤ 5 MB per item. This field can contain at most one inline item, which still counts toward the shared per-request limits of ≤ 10 MB total and ≤ 20 inline items. An invalid data: prefix or payload that decodes to zero bytes returns 422; exceeding a limit also returns 422. Use /v1/files/upload first and pass the resulting URL instead.
When sync: true, the request is not stored as a task and the rules above do not apply: inline base64 is submitted upstream unchanged without size or count validation.
1"https://storage.googleapis.com/cloud-samples-tests/speech/brooklyn.flac"
User prompt. When omitted, defaults to 'Please transcribe this audio file', aligning with the transcription scenario.
100000"Identify the speakers and emotion in this audio."
Synchronous mode. When true, the endpoint blocks until the upstream completes and returns the full response (if stream=true at the same time, returns an SSE stream); when false, the endpoint returns the task ID immediately, and results are fetched via GET /v1/tasks/{task_id} or the SSE endpoint.
false
Whether to stream. When true, the Submit response includes stream.url pointing to the SSE subscription path; streaming chunks are unified as the OpenAI chat.completion.chunk format.
false
Generation token limit. Optional.
x >= 1256
Sampling temperature, range [0, 2]. Optional.
0 <= x <= 2System instruction. Optional.
10000Whether to include reasoning tokens. Some thinking models require this to be set to true.
Response
Task created
Submit response, conforming to the unified task standard shape. results / error are fixed at null during submit; they are returned via GET /v1/tasks/{task_id} after the task completes or fails.
Task ID, formatted as task-llm-{timestamp}-{8random}.
"task-llm-1776874565-yq3szvcu"
llm.generation.task "llm.generation.task"
llm "llm"
The model name submitted by the client (echoed verbatim)
"gemini-2.5-pro"
pending "pending"
0
1776874565
Returns {url: ...} when stream=true; null when stream=false.
Fixed at null during submit; returned via GET /v1/tasks/{task_id} after the task completes — results[0] is the full OpenAI ChatCompletion response (audio transcription / understanding output is in message.content).
null
Fixed at null during submit; returned via GET /v1/tasks/{task_id} when the task fails.
null