What video understanding means
Video understanding is the process that turns pixels and audio into facts that software can use. A person watches a video and knows that the speaker changed topic at 3:12, that a price shows on screen at 1:28, and that the best moment is near the end. Software does not know these things until something extracts them.
Video Context does that extraction once, and stores the result as context: a structured, timestamped description of the video. Your product and your agents read the context, not the raw file.
What you get
- Scenes and shots. The boundaries of each shot, grouped into scenes with a label.
- Transcript. Words with timestamps and speaker labels.
- On-screen text. Titles, prices, slides, and captions, with the time that they show.
- Objects and people. What is in the frame, and when.
- Summary and chapters. A short description and a chapter list.
Built for agents
Every result uses the same schema. An agent can call the API through the Video Context MCP server, read the JSON, and decide what to do next: cut a clip, write a caption, or tag a catalog.