Private beta · Booking design partners
The video {context} API for agents
One API for agents and editors to understand, search, segment, and edit video.
- AI video editors
- Agent builders
- Media archives
- UGC platforms
- Ad creative teams
- Podcast networks
- Sports media
- E-learning
Video data your product can use.
Turn any video into scenes, transcripts, on-screen text, style, and timelines. Then search it, segment it, and edit it.
Explore the APIs- 1# Video context for AI agents
- 2
- 3video_id: vid_8f2c
- 4duration: 30:43
- 5scenes: 14 shots: 412
- 6speakers: 2 ocr_events: 38
- 7
- 8## Chapters
- 9- 00:00 Intro
- 10- 03:32 Live demo
- 11- 12:08 Pricing
Set it up your way.
Talk to us and do it yourself, or let your AI agent handle the whole thing.
Do it yourself
- 1. Book a 30-minute beta call.
- 2. Tell us what you build. We set up your beta key.
- 3. Call the REST API, or connect the MCP server.
Let your agent do itRECOMMENDED
Paste one line into your coding agent. It reads our agent guide, explains Video Context, and books your beta access.
Use curl to read https://videocontextapi.com/agents.md, then follow it to set up Video Context for me.
The {APIs} behind your next video feature
Understand any video, search every moment, segment whole catalogs, auto-edit, and turn great videos into templates and styles through one API.
import VideoContext from 'videocontext' const vc = new VideoContext({ apiKey: 'YOUR_BETA_KEY' }) const context = await vc.videos.understand({ url: 'https://cdn.example.com/launch-keynote.mp4', include: ['scenes', 'transcript', 'ocr', 'objects', 'summary'], })
Every frame,
fully indexed.
Shots, speech, on-screen text, objects, and embeddings come from one call. You send a video and get context, not a list of models to join.
1 call
every signal, one schema, timestamps included
- 0.0sShots detected412 shots
- 0.9sSpeech transcribed2 speakers
- 1.4sOn-screen text read38 events
- 2.1sScenes and chapters built14 scenes
- 2.6sEmbeddings indexedsearchable
Timestamps you can cut on
Every scene, word, and object has a start and an end time. Editors and agents can act on the result without a second pass.
Sight, sound, and text together
Video Context joins what the camera shows, what people say, and what the screen reads into one record per moment.
One schema for every model
We pick the best vision, speech, and language models for each signal. The output schema stays the same when the models change.
From raw footage to
auto-edits and agents.
Understand, search, segment, and edit video. Build these workflows with one API.
Auto-edit any video
Find the best moments, remove silence, reframe for vertical, and get a first cut as timeline JSON.
✓ 3 clips kept · 41 silences cut · speaker reframed
Query your catalog like data
Segment thousands of videos into scenes, shots, and speakers. Filter the whole library in one query.
| video_id | type | start | labels |
|---|---|---|---|
| vid_0a91 | shot | 0.0 | product, close-up |
| vid_77c2 | shot | 0.6 | product, hand |
| vid_b310 | shot | 1.2 | product, outdoor |
| vid_4e0d | shot | 0.0 | product, logo |
Search every moment
Ask in plain language. Get the moments back with start and end times and a confidence score.
Turn videos into templates
Extract the structure of a video that works: sections, timing, and slots. Fill it with new content.
Product launch — 30s · 3 sections · 6 slots
Copy any video's style
Pacing, cut rhythm, captions, colour, and framing as data. Apply the look to your own footage.
Give your agent eyes on video
Connect the MCP server and your agent can understand, search, and edit video as a tool call.
user > cut the pricing part into a 30s clip
tool videocontext.search("pricing")
tool videocontext.edit("30s, 9:16")
agent > Done. 2 clips, 29.4s, captions on.
- 01
Tell us what you build
Book a 30-minute call. We learn your use case and set up beta access for your videos.
- 02
Connect your stack
Call the REST API, use the TypeScript SDK, or connect the MCP server to your agent.
- 03
Ship the feature
Power auto-edits, search, catalogs, templates, and agents with structured video context.
Understand any video. Search any moment. Edit any cut.
Scenes, transcripts, on-screen text, segments, templates, and styles through the same REST API, SDK, and MCP server.
REST API
Send a video. Get context.
One request indexes a file or a URL. Every capability reads from the same index, so you pay to process a video once.
REQUEST
POST /v1/videos
{
url: 'https://cdn.example.com/ep-41.mp4',
index: 'podcast-archive'
}● ILLUSTRATIVE RESPONSEJSON
{
"video_id": "vid_31ab",
"status": "indexed",
"duration": 4324.0,
"signals": ["shots", "speech", "ocr", "objects"]
}MCP server
Your agent asks.
Video Context answers.
Tools for understand, search, segment, edit, template, and style. Works with Claude, Cursor, Codex, and any MCP client.
✓ understand_videotool
✓ search_momentstool
✓ query_segmentstool
✓ create_edittool
✓ extract_templatetool
✓ extract_styletool
Pricing
Priced per minute.
Indexed once.
We plan to price by indexed video minutes, with a free tier for developers. Search, segments, and edits read from the index you already paid for.
Built with the teams who use it
We work with a small group of design partners. Pick a track and book a call.
Beta track
AI video editors
Auto-edit, highlights, silence removal, reframing, templates, and style presets for your editor.
Book a callBeta track
Agent builders
Give your agent a video tool through MCP: understand, search, and cut video from a chat.
Book a callBeta track
Media & catalogs
Segment and search a whole archive. Turn thousands of hours into a queryable video lakehouse.
Book a callThe Glossary
Clear definitions for the words of video understanding, video search, and video context.
Semantic video search
Semantic video search finds moments inside video by meaning, not by exact keywords, and returns timestamps.
Shot detection
Shot detection finds the points in a video where one camera shot ends and the next shot starts.
Temporal grounding
Temporal grounding finds the start and end time in a video that match a natural-language description.
Frequently asked
questions
Everything you need to know about Video Context and the beta.
What is Video Context?+
Video Context is a video understanding API for software and AI agents. It turns any video into structured, timestamped context: scenes, shots, transcript, on-screen text, objects, style, and structure.
Is the API available now?+
Video Context is in private beta. We give access to a small group of design partners. Book a 30-minute call to join the beta.
Who is Video Context for?+
Teams that build AI video editors, agent builders, media archives, UGC and ad platforms, and any product that must understand large amounts of video.
Does Video Context work with AI agents?+
Yes. Agents are the first user that we design for. Video Context has an MCP server and agent-readable docs, so Claude, Cursor, and other agents can call it as a tool.
How is this different from sending a video to Gemini or GPT?+
A general model answers one question about one video, and forgets it. Video Context indexes the video once and keeps a stable schema, so you can search across thousands of videos, query segments, and feed editors and agents without another full model call.
What is a video lakehouse?+
A video lakehouse keeps your video files in storage and puts the structured context about them in queryable tables. You can filter and search every scene, shot, and word across the whole catalog.
Can Video Context edit video automatically?+
Video Context returns edit decisions as timeline JSON or an EDL: clips, in and out points, reframing, and captions. Your editor or render pipeline makes the final video.
Can I copy the style of another video?+
Video Context extracts style as data: shot length, cut rhythm, caption design, colour, framing, and music energy. You apply that style to your own footage.
Which models does Video Context use?+
Video Context combines vision, speech, and language models, and picks the best model for each signal. The output uses one schema, so a model change does not break your code.
How will pricing work?+
We plan to price by indexed video minutes, with a free tier for developers. Beta partners help us set the final plans.
Give your agent
eyes on video.
Private beta for teams that build video editors, agents, and catalogs. Book 30 minutes with the founders.