More than the words
Most transcription APIs read the audio only. But a video also shows information: the name of the speaker in a lower third, the slide behind them, or the score in the corner. Video Context puts the speech and the on-screen text on one timeline.
What you get
- Segments with timestamps. Each segment has a start time, an end time, and the text.
- Speaker labels. Each segment shows who speaks.
- On-screen text. Names, titles, slides, captions, and scores, with the times when they show.
- One schema. The transcript uses the same timeline as scenes, shots, and objects. See the video understanding API.
Where teams use it
- News. Find every quote from one person across a week of briefings. See Video Context for news.
- Search. Make an archive searchable by what people say. See the video search API.
- Captions and translation. Turn the segments into SRT or WebVTT files.