What the VTT file looks like
A WebVTT file starts with the line WEBVTT, then one cue per caption segment: a time range in HH:MM:SS.mmm format (a period before the milliseconds, unlike SRT’s comma) followed by the caption text. Cues are not numbered.
Timing is taken from the caption track. Each cue ends when its segment ends, or earlier if the next caption begins first, so cues never overlap. Very short cues are held for at least half a second.
Sample VTT output (original example text, not from a real video)
WEBVTT 00:00:00.000 --> 00:00:04.000 First, write down the question you want the video to answer. 00:00:04.000 --> 00:00:08.500 Next, search the transcript for a distinctive phrase. 00:00:08.500 --> 00:00:13.500 Finally, check the matching moment in the source video.
Adding VTT captions to a web page
WebVTT is the caption format built into HTML5 video. Place the .vtt file next to your video and reference it with a track element:
HTML5 example
<video controls src="lecture.mp4"> <track kind="captions" src="lecture.vtt" srclang="en" label="English" default> </video>
Where VTT is the better choice
- Web players and course platforms (LMS systems such as Moodle or Canvas, and many hosted video players) generally expect VTT.
- Serve the file with the text/vtt content type; some browsers ignore caption tracks served with the wrong type.
- If the video is hosted on a different domain from the page, the caption file needs CORS headers, or the browser will block the track.
- For desktop editors and offline players, SRT is usually the safer choice. Both formats carry the same text and timing.