Private beta: now booking design-partner calls
Video Context
Video understanding

Every video,
fully understood.

A video understanding API takes a video file or URL and returns structured data about it, such as scenes, shots, a transcript, on-screen text, objects, and a summary, each with timestamps, so software and AI agents can reason about the video without watching it.

REQUESTPOST /v1/understand
const context = await vc.videos.understand({
  url: 'https://cdn.example.com/launch-keynote.mp4',
  include: ['scenes', 'transcript', 'ocr', 'objects', 'summary'],
})
● ILLUSTRATIVE RESPONSEPREVIEW API
{
  "video_id": "vid_8f2c",
  "duration": 1843.2,
  "summary": "Product launch keynote for the v4 release.",
  "scenes": [
    { "start": 0.0, "end": 41.6, "label": "Cold open, city b-roll" },
    { "start": 41.6, "end": 212.9, "label": "CEO on stage, intro" }
  ],
  "ocr": [{ "t": 88.4, "text": "Pro plan ยท $49/mo" }]
}

This is the planned API design. Video Context is in private beta, and beta partners help set the final design of each endpoint.

What video understanding means

Video understanding is the process that turns pixels and audio into facts that software can use. A person watches a video and knows that the speaker changed topic at 3:12, that a price shows on screen at 1:28, and that the best moment is near the end. Software does not know these things until something extracts them.

Video Context does that extraction once, and stores the result as context: a structured, timestamped description of the video. Your product and your agents read the context, not the raw file.

What you get

  • Scenes and shots. The boundaries of each shot, grouped into scenes with a label.
  • Transcript. Words with timestamps and speaker labels.
  • On-screen text. Titles, prices, slides, and captions, with the time that they show.
  • Objects and people. What is in the frame, and when.
  • Summary and chapters. A short description and a chapter list.

Built for agents

Every result uses the same schema. An agent can call the API through the Video Context MCP server, read the JSON, and decide what to do next: cut a clip, write a caption, or tag a catalog.

✦ FAQs

Frequently asked
questions

Common questions about video understanding with Video Context.

What does a video understanding API return?+

It returns one JSON object per video. The object holds scenes, shots, a transcript with speakers, on-screen text, detected objects, and a summary. Every item has a start and an end time in seconds.

How is this different from sending a video to Gemini or GPT?+

A general model answers one question about one video. Video Context stores the result as a stable, indexed schema, so you can query it again, search across many videos, and feed it to an editor or an agent without another model call.

Is the Video Context API available now?+

Video Context is in private beta. Book a call with the team to get access and to shape the first version.

Give your agent
eyes on video.

Private beta for teams that build video editors, agents, and catalogs. Book 30 minutes with the founders.