The exact JSON shape
The file is a single object. It includes the video metadata (ID, title, channel), the caption language code, a total_segments count, and a segments array. Each segment has the caption text, start and duration in seconds (rounded to two decimals), and a human-readable startFormatted timestamp.
Sample JSON output (original example text, not from a real video)
{
"video": {
"videoId": "VIDEO_ID",
"title": "Example Video Title",
"channelName": "Example Channel"
},
"language": "en",
"total_segments": 3,
"segments": [
{
"text": "First, write down the question you want the video to answer.",
"start": 0,
"duration": 4,
"startFormatted": "00:00"
},
{
"text": "Next, search the transcript for a distinctive phrase.",
"start": 4,
"duration": 4.5,
"startFormatted": "00:04"
},
{
"text": "Finally, check the matching moment in the source video.",
"start": 8.5,
"duration": 5,
"startFormatted": "00:08"
}
]
}Common ways developers use it
- RAG and embeddings: chunk the segments array by time or word count and keep start with each chunk, so an answer can cite the moment it came from.
- Search indexes: index text with start as a field to build “jump to where they said X” search over your own video library.
- Subtitle tooling: start and duration map directly onto cue timing if you want to generate a custom subtitle format.
- Analysis: load it with pandas (pd.json_normalize(data["segments"])) or JSON.parse in Node for word frequency, pacing, or topic analysis.
Notes on the timing fields
- Times are seconds as numbers, not strings, so you can do arithmetic on them directly.
- The end of a segment is start + duration. Caption tracks sometimes overlap slightly, so do not assume one segment ends exactly when the next begins.
- For many videos at once, or to fetch programmatically, the REST API returns the same data without the download step.