Why agents need video context
A language model cannot watch a two-hour video every time it needs one fact. It needs the video as context: a compact, structured record that it can read, search, and cite.
What video context includes
- Scenes and shots with timestamps
- A transcript with speakers
- On-screen text
- Objects, people, and actions
- Style: pacing, captions, colour, and framing
- Structure: sections and roles, for templates
The Video Context API makes this record once per video, and serves it to apps and agents through REST and MCP.