---
title: 'What is video understanding?'
description: 'Video understanding is the process that turns a video''s frames and audio into structured facts, such as scenes, speech, objects, and events, with timestamps.'
url: 'https://videocontextapi.com/glossary/video-understanding/'
---

# Video understanding

> Video understanding is the process that turns the frames and audio of a video into structured, timestamped facts, such as scenes, speech, on-screen text, objects, and events, that software can read and reason about.

## How video understanding works

A video understanding system usually combines several models:

1. **Shot detection** splits the video where the camera cuts.
2. **Speech recognition** makes a transcript with timestamps and speakers.
3. **OCR** reads text on screen.
4. **Vision models** describe objects, people, and actions in sampled frames.
5. **A language model** joins these signals into scenes, chapters, and a summary.

## Why it matters

Video is the largest type of data on the internet, but software cannot read it directly. Video understanding makes video searchable, editable, and usable by AI agents.

Video Context provides video understanding as an API. See the [Video Understanding API](/video-understanding-api/).

Related terms: https://videocontextapi.com/glossary/temporal-grounding/, https://videocontextapi.com/glossary/shot-detection/, https://videocontextapi.com/glossary/video-context/
