Private beta: we are looking for design partners
Video Context

Comparison · 2026

The best video understanding
APIs in 2026.

A fair comparison of the main APIs that turn video into data: what each one does, where it runs, how it is priced, and what it is best for. We build Video Context, so it is on this list too. We checked every fact on the vendors’ own pages on 1 October 2026.

8 APIs compared

  • Video Context
  • Twelve Labs
  • Gemini API (Google)
  • Azure AI Video Indexer
  • Amazon Rekognition Video
  • Mixpeek
  • OpenAI API
  • Google Cloud Video Intelligence API
Changes in 2026: Google deprecated the Cloud Video Intelligence API on 14 September 2026, and it shuts down on 14 September 2027 (see the migration guide). Amazon Rekognition closed streaming video analysis to new customers on 30 April 2026.

At a glance

FeatureVideo ContextPrivate betaTwelve Labs Gemini Azure Rekognition Mixpeek OpenAI Google VIDeprecated
Search a whole library in plain languageYesYes—Not listedPartly—Not listedYes—Not listed—Not listed
Ready-made labels, shots, and on-screen textYesPartlyPartlyYesYesYes—Not listedYes
Speech transcriptionYesPartlyYesYes—Not listedYesPartlyYes
Answers and summaries from a promptPartlyYesYesPartly—Not listed—Not listedPartly—Not listed
Runs in your cloud or on your serversYesPartly—Not listedYes—Not listedYes—Not listed—Not listed
Official MCP server for agentsYesYes—Not listed—Not listed—Not listedYes—Not listed—Not listed
Pricing modelPer minute (planned)Per hour + per queryPer tokenPer minutePer minute, per featurePlan + usagePer tokenPer minute

“Partly” means that the feature is limited, available on request, or that you build part of it yourself. A dash means that the vendor does not list the feature.

How to choose

You need to

Ask questions about one video at a time

A multimodal model is flexible and quick to start. You write the prompts.

You need to

Search a large library

Pick an API that indexes video once and searches it many times.

You need to

Ready-made labels, faces, or moderation

These two have the widest catalogues of ready-made insights.

You need to

Video that cannot leave your network

Check deployment first. Twelve Labs offers private deployment on request.

You need to

Agents that use video as a tool

Look for an official MCP server.

Private beta

Video Context

A video understanding API that turns video into structured, actionable data, and lets you automate what happens next. It is built for teams that want one API for video data and workflows, with agents as a first-class user, and the option to run in their own cloud or on their own servers.

  • People, objects, speech, on-screen text, shots, and events, with timestamps, in one schema
  • Search for moments in plain language
  • Automations: tag, categorise, route, and flag
  • Agent tools for ChatGPT, Claude, Gemini, and MCP clients
Where it runs
Our cloud, your cloud account, or your own servers.
Pricing
Planned: per indexed video minute, with a free developer tier. Beta partners help set the plans.
Agents / MCP
MCP server and agent guide for beta partners.

The other options

Twelve Labs

Video foundation models: Marengo for search and embeddings, and Pegasus for text generation from video.

Best for: Semantic search, embeddings, and generated text from video, from one video-specific vendor.

Show capabilities, pricing, and sources
  • ✓Search with text or image queries
  • ✓Summaries, captions, and prompt-based answers (Analyze)
  • ✓Video, audio, image, and text embeddings
  • ✓Jockey agent system (research preview)
Where it runs
Twelve Labs cloud and Amazon Bedrock. Private or self-managed deployment on request.
Pricing
Usage-based: per hour of video indexed and analysed, per 1,000 search queries, and per token for embeddings. Free plan with 600 minutes.
Agents / MCP
Official MCP server (Jockey), with OAuth sign-in.
Note
Pegasus 1.2 is now a legacy model.

Sources: Pricing · Docs · MCP server

Gemini API (Google)

Gemini multimodal models that take video as a direct input to a prompt.

Best for: Flexible questions and reasoning about individual videos, when you write the prompts.

Show capabilities, pricing, and sources
  • ✓Summaries and Q&A with timestamps
  • ✓Audio transcription
  • ✓Process part of a video by start and end offset, at a custom frame rate
  • ✓Up to 10 videos in one request on recent models
Where it runs
Google cloud: the Gemini Developer API or Vertex AI.
Pricing
Per token. Video is sampled at 1 frame per second by default, about 100 to 300 tokens per second of video.
Agents / MCP
No official video MCP server found.
Note
Up to about 3 hours of video per request at low resolution, or 1 hour at high resolution, on 1M-context models. Google names Gemini as the replacement for the Video Intelligence API.

Sources: Video understanding docs

Azure AI Video Indexer

A Microsoft service that runs many audio and video models together and builds an index of insights.

Best for: Media and enterprise archives that need a wide set of ready-made insights, with an edge or on-premises option.

Show capabilities, pricing, and sources
  • ✓Transcription in 50+ languages, translation, captions, and speaker identification
  • ✓OCR, labels, objects, scenes, shots, and keyframes
  • ✓Faces and celebrities (limited access, application required)
  • ✓Topics, keywords, sentiment, and summaries with Azure OpenAI
Where it runs
Azure cloud, private endpoints, and an Azure Arc extension for edge and on-premises Kubernetes.
Pricing
Per input minute, by preset. See the Azure pricing page. A trial account includes free minutes.
Agents / MCP
No official MCP server.
Note
Files up to 6 hours (12 hours for the Basic Audio preset).

Sources: Overview · Pricing

Amazon Rekognition Video

An AWS computer-vision service for stored and streaming video.

Best for: AWS teams that need structured labels, faces, and moderation on stored video.

Show capabilities, pricing, and sources
  • ✓Object, scene, and activity labels
  • ✓Content moderation and text detection
  • ✓Face detection, face search, and celebrity recognition
  • ✓Shot detection and person pathing
Where it runs
AWS cloud.
Pricing
Per minute of video, per feature.
Agents / MCP
No official MCP server found.
Note
Streaming video analysis is closed to new customers from 30 April 2026. Stored video analysis is not affected.

Sources: Features · Availability changes · Pricing

Mixpeek

A semantic retrieval layer that turns video, images, audio, and documents into searchable features.

Best for: Teams that want to configure a multimodal retrieval pipeline on top of third-party models.

Show capabilities, pricing, and sources
  • ✓Feature extractors: scenes, keyframes, transcripts, faces, and embeddings
  • ✓Semantic, face, object, and transcript search
  • ✓Multi-stage retrieval pipelines
  • ✓Reverse video search
Where it runs
Managed (multi-tenant or dedicated), or in your own AWS, GCP, or Azure account.
Pricing
Monthly plans plus processing, or per stored vector. See the Mixpeek pricing page.
Agents / MCP
Official MCP servers for ingestion, retrieval, and admin.

Sources: Pricing · MCP server

OpenAI API

General language models with image input. The official docs list image input, not video input.

Best for: Teams that already use OpenAI and can build their own frame sampling and transcription.

Show capabilities, pricing, and sources
  • ✓Answers about frames that you extract and send as images
  • ✓Long context for many frames at once
  • ✓Transcription through separate audio models
Where it runs
OpenAI cloud, and Azure OpenAI.
Pricing
Per token. Each frame is billed as an image.
Agents / MCP
Not specific to video.
Note
There is no built-in shot detection, video index, or video search. OpenAI’s cookbook samples frames with OpenCV.

Sources: Vision guide · Video cookbook

Google Cloud Video Intelligence API

deprecated

A classic computer-vision API with labels, shots, OCR, logos, objects, and transcription.

Best for: Existing users only, while they migrate.

Show capabilities, pricing, and sources
  • ✓Labels, shot changes, and explicit content
  • ✓Speech transcription, OCR, and logos
  • ✓Object tracking, faces, and people
Where it runs
Google Cloud.
Pricing
Per minute, per feature, with free monthly minutes.
Agents / MCP
No MCP server.
Note
Deprecated on 14 September 2026. It shuts down on 14 September 2027. Google recommends Gemini instead.

Sources: Deprecation notice

We checked these facts on the vendors’ own pages on 1 October 2026. Prices, models, and limits change often, so check each vendor’s page before you decide. Logos belong to their owners and show here only to identify each product. If you see a mistake, email us and we will fix it.

FAQ

Questions, answered.

Short answers about video understanding APIs.

What is the best video understanding API?+

It depends on the job. Twelve Labs is strong for video search and embeddings. Gemini is flexible for questions about a single video. Azure AI Video Indexer has a wide set of ready-made insights. Amazon Rekognition suits AWS teams that need labels, faces, and moderation. Video Context is built for teams that want structured video data, automations, and agent tools in one API.

Is the Google Cloud Video Intelligence API being shut down?+

Yes. Google deprecated the Video Intelligence API on 14 September 2026, and it shuts down on 14 September 2027. Google recommends moving to Gemini models.

Can the OpenAI API understand video?+

The OpenAI docs list image input, not video input. To use OpenAI models on video, you extract frames yourself, send them as images, and transcribe the audio separately.

Which video APIs can run on my own infrastructure?+

Azure AI Video Indexer has an Azure Arc extension for edge and on-premises use. Mixpeek can run in your own cloud account. Twelve Labs offers private deployment on request. Video Context can run in your cloud account or on your own servers.

Which video APIs work with AI agents through MCP?+

Twelve Labs and Mixpeek have official MCP servers. Video Context gives beta partners an MCP server and an agent setup guide.

Your videos contain
valuable data.
Start using it.

Analyse video at scale. Find exactly what you need. Give your products and AI the context to act on it.