---
title: 'Video Understanding Glossary'
description: 'Definitions for the words of video understanding.'
url: 'https://videocontextapi.com/glossary/'
---

# Video understanding glossary

- [Semantic video search](https://videocontextapi.com/glossary/semantic-video-search/): Semantic video search is a way to find moments inside video by their meaning, so a natural-language query such as "someone opens a laptop" returns the matching timestamps even when nobody says those words.
- [Shot detection](https://videocontextapi.com/glossary/shot-detection/): Shot detection, also called shot boundary detection, is the process that finds each point in a video where one continuous camera shot ends and the next shot starts, including hard cuts, fades, and dissolves.
- [Temporal grounding](https://videocontextapi.com/glossary/temporal-grounding/): Temporal grounding is the task of finding the exact start and end time in a video that match a natural-language description, such as "the moment the speaker shows the pricing slide".
- [Video context](https://videocontextapi.com/glossary/video-context/): Video context is the structured, timestamped description of a video, including its scenes, shots, transcript, on-screen text, objects, style, and structure, that software and AI agents read in place of the raw video file.
- [Video embeddings](https://videocontextapi.com/glossary/video-embeddings/): Video embeddings are numeric vectors that represent the visual, audio, and text content of a video segment, so that segments with similar meaning sit close together and software can search, cluster, and compare them.
- [Video understanding](https://videocontextapi.com/glossary/video-understanding/): Video understanding is the process that turns the frames and audio of a video into structured, timestamped facts, such as scenes, speech, on-screen text, objects, and events, that software can read and reason about.
