---
title: 'What is video embeddings?'
description: 'Video embeddings are vectors that represent the meaning of a video segment, so similar moments sit close together for search and clustering.'
url: 'https://videocontextapi.com/glossary/video-embeddings/'
---

# Video embeddings

> Video embeddings are numeric vectors that represent the visual, audio, and text content of a video segment, so that segments with similar meaning sit close together and software can search, cluster, and compare them.

## Frame, clip, and scene embeddings

- **Frame embeddings** represent one still image.
- **Clip embeddings** represent a few seconds and capture motion.
- **Scene embeddings** combine many clips and the speech in them.

The level that you choose changes what search can find. Actions need clip embeddings. Topics need scene embeddings that include the transcript.

## What embeddings are used for

Semantic search, deduplication, recommendations, and clustering a catalog. See the [video lakehouse](/video-lakehouse/).

Related terms: https://videocontextapi.com/glossary/semantic-video-search/, https://videocontextapi.com/glossary/video-context/
