---
title: 'What is video context?'
description: 'Video context is the structured, timestamped description of a video that software and AI agents use in place of the raw file.'
url: 'https://videocontextapi.com/glossary/video-context/'
---

# Video context

> Video context is the structured, timestamped description of a video, including its scenes, shots, transcript, on-screen text, objects, style, and structure, that software and AI agents read in place of the raw video file.

## Why agents need video context

A language model cannot watch a two-hour video every time it needs one fact. It needs the video as **context**: a compact, structured record that it can read, search, and cite.

## What video context includes

- Scenes and shots with timestamps
- A transcript with speakers
- On-screen text
- Objects, people, and actions
- Style: pacing, captions, colour, and framing
- Structure: sections and roles, for templates

The Video Context API makes this record once per video, and serves it to apps and agents through REST and MCP.

Related terms: https://videocontextapi.com/glossary/video-understanding/, https://videocontextapi.com/glossary/video-embeddings/, https://videocontextapi.com/glossary/shot-detection/
