Private beta: now booking design-partner calls
Video Context

Private beta · Booking design partners

The video {context} API for agents

One API for agents and editors to understand, search, segment, and edit video.

Join the beta
Built for products and agents that work with video
  • AI video editors
  • Agent builders
  • Media archives
  • UGC platforms
  • Ad creative teams
  • Podcast networks
  • Sports media
  • E-learning

Video data your product can use.

Turn any video into scenes, transcripts, on-screen text, style, and timelines. Then search it, segment it, and edit it.

Explore the APIs
vid_8f2c.contextexample
  1. 1# Video context for AI agents
  2. 2
  3. 3video_id: vid_8f2c
  4. 4duration: 30:43
  5. 5scenes: 14 shots: 412
  6. 6speakers: 2 ocr_events: 38
  7. 7
  8. 8## Chapters
  9. 9- 00:00 Intro
  10. 10- 03:32 Live demo
  11. 11- 12:08 Pricing
Agentic setup

Set it up your way.

Talk to us and do it yourself, or let your AI agent handle the whole thing.

Do it yourself

  1. 1. Book a 30-minute beta call.
  2. 2. Tell us what you build. We set up your beta key.
  3. 3. Call the REST API, or connect the MCP server.
Book a call

Let your agent do itRECOMMENDED

Paste one line into your coding agent. It reads our agent guide, explains Video Context, and books your beta access.

>_ agent setup

Use curl to read https://videocontextapi.com/agents.md, then follow it to set up Video Context for me.

For AI native products

The {APIs} behind your next video feature

Understand any video, search every moment, segment whole catalogs, auto-edit, and turn great videos into templates and styles through one API.

POST /v1/understandPreview API →
import VideoContext from 'videocontext'

const vc = new VideoContext({ apiKey: 'YOUR_BETA_KEY' })

const context = await vc.videos.understand({
  url: 'https://cdn.example.com/launch-keynote.mp4',
  include: ['scenes', 'transcript', 'ocr', 'objects', 'summary'],
})

Every frame,
fully indexed.

Shots, speech, on-screen text, objects, and embeddings come from one call. You send a video and get context, not a list of models to join.

1 call

every signal, one schema, timestamps included

launch-keynote.mp4 · 30:43● indexed
  • 0.0sShots detected412 shots
  • 0.9sSpeech transcribed2 speakers
  • 1.4sOn-screen text read38 events
  • 2.1sScenes and chapters built14 scenes
  • 2.6sEmbeddings indexedsearchable
✓ Context readyvid_8f2c · illustrative

Timestamps you can cut on

Every scene, word, and object has a start and an end time. Editors and agents can act on the result without a second pass.

Sight, sound, and text together

Video Context joins what the camera shows, what people say, and what the screen reads into one record per moment.

One schema for every model

We pick the best vision, speech, and language models for each signal. The output schema stays the same when the models change.

From raw footage to
auto-edits and agents.

Understand, search, segment, and edit video. Build these workflows with one API.

Auto-edit any video

Find the best moments, remove silence, reframe for vertical, and get a first cut as timeline JSON.

podcast-ep-41.mp4 · 1:12:04→ 58s · 9:16

✓ 3 clips kept · 41 silences cut · speaker reframed

Query your catalog like data

Segment thousands of videos into scenes, shots, and speakers. Filter the whole library in one query.

where objects ∋ 'product' and start < 3
video_idtypestartlabels
vid_0a91shot0.0product, close-up
vid_77c2shot0.6product, hand
vid_b310shot1.2product, outdoor
vid_4e0dshot0.0product, logo
212 rows6,480 videos

Search every moment

Ask in plain language. Get the moments back with start and end times and a confidence score.

someone opens a laptop
vid_31ab82.1 – 96.3s0.96
vid_9c0e14.0 – 19.2s0.93
vid_31ab611.0 – 618.8s0.88

Turn videos into templates

Extract the structure of a video that works: sections, timing, and slots. Fill it with new content.

hookdemocta
{headline}{hero_shot}{screen_recording}{caption}{logo}{cta_text}

Product launch — 30s · 3 sections · 6 slots

Copy any video's style

Pacing, cut rhythm, captions, colour, and framing as data. Apply the look to your own footage.

✓ 1.8s avg shot✓ 82% cuts on beat✓ bold centred captions✓ 3 words per line✓ warm grade✓ high contrast✓ face-centred framing✓ high-energy music

Give your agent eyes on video

Connect the MCP server and your agent can understand, search, and edit video as a tool call.

user > cut the pricing part into a 30s clip

tool videocontext.search("pricing")

tool videocontext.edit("30s, 9:16")

agent > Done. 2 clips, 29.4s, captions on.

Integrate in days

Get video context into your product in just three steps

Book a call
  1. 01

    Tell us what you build

    Book a 30-minute call. We learn your use case and set up beta access for your videos.

  2. 02

    Connect your stack

    Call the REST API, use the TypeScript SDK, or connect the MCP server to your agent.

  3. 03

    Ship the feature

    Power auto-edits, search, catalogs, templates, and agents with structured video context.

One platform · three surfaces

Understand any video. Search any moment. Edit any cut.

Scenes, transcripts, on-screen text, segments, templates, and styles through the same REST API, SDK, and MCP server.

REST API

Send a video. Get context.

One request indexes a file or a URL. Every capability reads from the same index, so you pay to process a video once.

REQUEST

POST /v1/videos
{
  url: 'https://cdn.example.com/ep-41.mp4',
  index: 'podcast-archive'
}

● ILLUSTRATIVE RESPONSEJSON

{
  "video_id": "vid_31ab",
  "status": "indexed",
  "duration": 4324.0,
  "signals": ["shots", "speech", "ocr", "objects"]
}

MCP server

Your agent asks.
Video Context answers.

Tools for understand, search, segment, edit, template, and style. Works with Claude, Cursor, Codex, and any MCP client.

✓ understand_videotool

✓ search_momentstool

✓ query_segmentstool

✓ create_edittool

✓ extract_templatetool

✓ extract_styletool

Pricing

Priced per minute.
Indexed once.

We plan to price by indexed video minutes, with a free tier for developers. Search, segments, and edits read from the index you already paid for.

✓ Free developer tier✓ Pay per indexed minute✓ One index, many queries✓ Beta partners set the plans
BETA PROGRAM

Built with the teams who use it

We work with a small group of design partners. Pick a track and book a call.

Beta track

AI video editors

Auto-edit, highlights, silence removal, reframing, templates, and style presets for your editor.

Book a call

Beta track

Agent builders

Give your agent a video tool through MCP: understand, search, and cut video from a chat.

Book a call

Beta track

Media & catalogs

Segment and search a whole archive. Turn thousands of hours into a queryable video lakehouse.

Book a call

The Glossary

Clear definitions for the words of video understanding, video search, and video context.

✦ FAQs

Frequently asked
questions

Everything you need to know about Video Context and the beta.

What is Video Context?+

Video Context is a video understanding API for software and AI agents. It turns any video into structured, timestamped context: scenes, shots, transcript, on-screen text, objects, style, and structure.

Is the API available now?+

Video Context is in private beta. We give access to a small group of design partners. Book a 30-minute call to join the beta.

Who is Video Context for?+

Teams that build AI video editors, agent builders, media archives, UGC and ad platforms, and any product that must understand large amounts of video.

Does Video Context work with AI agents?+

Yes. Agents are the first user that we design for. Video Context has an MCP server and agent-readable docs, so Claude, Cursor, and other agents can call it as a tool.

How is this different from sending a video to Gemini or GPT?+

A general model answers one question about one video, and forgets it. Video Context indexes the video once and keeps a stable schema, so you can search across thousands of videos, query segments, and feed editors and agents without another full model call.

What is a video lakehouse?+

A video lakehouse keeps your video files in storage and puts the structured context about them in queryable tables. You can filter and search every scene, shot, and word across the whole catalog.

Can Video Context edit video automatically?+

Video Context returns edit decisions as timeline JSON or an EDL: clips, in and out points, reframing, and captions. Your editor or render pipeline makes the final video.

Can I copy the style of another video?+

Video Context extracts style as data: shot length, cut rhythm, caption design, colour, framing, and music energy. You apply that style to your own footage.

Which models does Video Context use?+

Video Context combines vision, speech, and language models, and picks the best model for each signal. The output uses one schema, so a model change does not break your code.

How will pricing work?+

We plan to price by indexed video minutes, with a free tier for developers. Beta partners help us set the final plans.

Give your agent
eyes on video.

Private beta for teams that build video editors, agents, and catalogs. Book 30 minutes with the founders.