Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
channel | Optional[int] | :heavy_minus_sign: | Zero-based audio channel index for the word, present when the provider transcribes channels separately | 0 |
confidence | Optional[float] | :heavy_minus_sign: | Provider confidence for the word from 0 to 1, present when the provider returns per-word confidence | 0.98 |
end | float | :heavy_check_mark: | Word end time in seconds | 0.4 |
speaker | Optional[int] | :heavy_minus_sign: | Speaker index for the word, present when the provider returns diarization data | 0 |
speaker_label | Optional[str] | :heavy_minus_sign: | Provider speaker label for the word, present when the provider labels speakers with a string | speaker_0 |
start | float | :heavy_check_mark: | Word start time in seconds | 0 |
type | Optional[components.STTWordType] | :heavy_minus_sign: | Kind of entry; omitted or “word” for spoken words, “audio_event” for non-speech sounds the provider tags with timestamps | word |
word | str | :heavy_check_mark: | The transcribed word, or the event tag such as “(laughter)” when type is audio_event | Hello |