Frame, clip, and scene embeddings
- Frame embeddings represent one still image.
- Clip embeddings represent a few seconds and capture motion.
- Scene embeddings combine many clips and the speech in them.
The level that you choose changes what search can find. Actions need clip embeddings. Topics need scene embeddings that include the transcript.
What embeddings are used for
Semantic search, deduplication, recommendations, and clustering a catalog. See the video lakehouse.