Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
InputAudio | components.STTInputAudio | :heavy_check_mark: | Base64-encoded audio to transcribe | { “data”: “UklGRiQA…”, “format”: “wav” } |
Language | *string | :heavy_minus_sign: | ISO-639-1 language code (e.g., “en”, “ja”). Auto-detected if omitted. | en |
Model | string | :heavy_check_mark: | STT model identifier | openai/whisper-large-v3 |
Provider | *components.STTRequestProvider | :heavy_minus_sign: | Provider-specific passthrough configuration | |
ResponseFormat | *components.STTRequestResponseFormat | :heavy_minus_sign: | Output format. “json” (default) returns { text, usage }. “verbose_json” additionally returns task, language, duration, and segment-level timestamps; only supported by OpenAI-compatible providers. | json |
Temperature | *float64 | :heavy_minus_sign: | Sampling temperature for transcription | 0 |
TimestampGranularities | []components.STTTimestampGranularity | :heavy_minus_sign: | Timestamp detail levels to include when response_format is “verbose_json”. “segment” returns segment-level timestamps; “word” additionally returns word-level timestamps in the words array. Ignored unless response_format is “verbose_json”. | [ “segment” ] |