Skip to main content
POST

Authorizations

Authorization
string
header
required

All endpoints require Bearer Token authentication. Add to the request header:

Authorization: Bearer YOUR_API_KEY

YOUR_API_KEY is the API Token (sk-... format).

Body

application/json
model
string
required
Example:

"gemini-2.5-pro"

audio_url
string
required

Audio source. Accepts one of the following two forms:

  • Publicly reachable HTTP/HTTPS URL
  • data:audio/<type>;base64,<payload> data URI (base64 inline)

Audio format support per family (the specific available models are driven by channel configuration):

  • Gemini family (e.g. gemini-*): wav/mp3/aiff/aac/ogg/flac/m4a; total request body (prompt + system + inline files) ≤ 20 MB

Inline base64 (asynchronous mode only, when sync is omitted or false): Base64 recognized by the gateway as inline media (data: URI, or at least 4096 encoded characters whose decoded content matches a known media signature) is uploaded to file storage and replaced with a URL before the request is persisted or submitted upstream; the original base64 is not stored. Shorter raw base64 or content with an unrecognized signature is passed through unchanged and is not size-validated. Limits apply only to inline base64 items (public URLs are not limited): decoded size ≤ 5 MB per item. This field can contain at most one inline item, which still counts toward the shared per-request limits of ≤ 10 MB total and ≤ 20 inline items. An invalid data: prefix or payload that decodes to zero bytes returns 422; exceeding a limit also returns 422. Use /v1/files/upload first and pass the resulting URL instead.

When sync: true, the request is not stored as a task and the rules above do not apply: inline base64 is submitted upstream unchanged without size or count validation.

Minimum string length: 1
Example:

"https://storage.googleapis.com/cloud-samples-tests/speech/brooklyn.flac"

prompt
string | null

User prompt. When omitted, defaults to 'Please transcribe this audio file', aligning with the transcription scenario.

Maximum string length: 100000
Example:

"Identify the speakers and emotion in this audio."

sync
boolean
default:false

Synchronous mode. When true, the endpoint blocks until the upstream completes and returns the full response (if stream=true at the same time, returns an SSE stream); when false, the endpoint returns the task ID immediately, and results are fetched via GET /v1/tasks/{task_id} or the SSE endpoint.

Example:

false

stream
boolean
default:false

Whether to stream. When true, the Submit response includes stream.url pointing to the SSE subscription path; streaming chunks are unified as the OpenAI chat.completion.chunk format.

Example:

false

max_tokens
integer | null

Generation token limit. Optional.

Required range: x >= 1
Example:

256

temperature
number | null

Sampling temperature, range [0, 2]. Optional.

Required range: 0 <= x <= 2
system_prompt
string | null

System instruction. Optional.

Maximum string length: 10000
reasoning
boolean | null

Whether to include reasoning tokens. Some thinking models require this to be set to true.

Response

Task created

Submit response, conforming to the unified task standard shape. results / error are fixed at null during submit; they are returned via GET /v1/tasks/{task_id} after the task completes or fails.

id
string
required

Task ID, formatted as task-llm-{timestamp}-{8random}.

Example:

"task-llm-1776874565-yq3szvcu"

object
enum<string>
required
Available options:
llm.generation.task
Example:

"llm.generation.task"

type
enum<string>
required
Available options:
llm
Example:

"llm"

model
string
required

The model name submitted by the client (echoed verbatim)

Example:

"gemini-2.5-pro"

status
enum<string>
required
Available options:
pending
Example:

"pending"

progress
integer
required
Example:

0

created
integer
required
Example:

1776874565

stream
object · null · null

Returns {url: ...} when stream=true; null when stream=false.

results
object[] | null

Fixed at null during submit; returned via GET /v1/tasks/{task_id} after the task completes — results[0] is the full OpenAI ChatCompletion response (audio transcription / understanding output is in message.content).

Example:

null

error
object | null

Fixed at null during submit; returned via GET /v1/tasks/{task_id} when the task fails.

Example:

null