# The best video understanding APIs in 2026

> A comparison of video understanding APIs, checked on the vendors’ own pages on 1 October 2026. Video Context publishes this page and is on the list.

Changes in 2026: Google deprecated the Cloud Video Intelligence API on 14 September 2026 (shutdown 14 September 2027). Amazon Rekognition closed streaming video analysis to new customers on 30 April 2026.

## Video Context

A video understanding API that turns video into structured, actionable data, and lets you automate what happens next.

- People, objects, speech, on-screen text, shots, and events, with timestamps, in one schema
- Search for moments in plain language
- Automations: tag, categorise, route, and flag
- Agent tools for ChatGPT, Claude, Gemini, and MCP clients

- **Best for:** Teams that want one API for video data and workflows, with agents as a first-class user, and the option to run in their own cloud or on their own servers.
- **Where it runs:** Our cloud, your cloud account, or your own servers.
- **Pricing:** Planned: per indexed video minute, with a free developer tier. Beta partners help set the plans.
- **Agents / MCP:** MCP server and agent guide for beta partners.
- **Note:** Private beta. Access after a call with the team.

## Twelve Labs

Video foundation models: Marengo for search and embeddings, and Pegasus for text generation from video.

- Search with text or image queries
- Summaries, captions, and prompt-based answers (Analyze)
- Video, audio, image, and text embeddings
- Jockey agent system (research preview)

- **Best for:** Semantic search, embeddings, and generated text from video, from one video-specific vendor.
- **Where it runs:** Twelve Labs cloud and Amazon Bedrock. Private or self-managed deployment on request.
- **Pricing:** Usage-based: per hour of video indexed and analysed, per 1,000 search queries, and per token for embeddings. Free plan with 600 minutes.
- **Agents / MCP:** Official MCP server (Jockey), with OAuth sign-in.
- **Note:** Pegasus 1.2 is now a legacy model.
- **Sources:** [Pricing](https://www.twelvelabs.io/pricing), [Docs](https://docs.twelvelabs.io/docs/get-started/introduction), [MCP server](https://docs.twelvelabs.io/docs/advanced/model-context-protocol)

## Gemini API (Google)

Gemini multimodal models that take video as a direct input to a prompt.

- Summaries and Q&A with timestamps
- Audio transcription
- Process part of a video by start and end offset, at a custom frame rate
- Up to 10 videos in one request on recent models

- **Best for:** Flexible questions and reasoning about individual videos, when you write the prompts.
- **Where it runs:** Google cloud: the Gemini Developer API or Vertex AI.
- **Pricing:** Per token. Video is sampled at 1 frame per second by default, about 100 to 300 tokens per second of video.
- **Agents / MCP:** No official video MCP server found.
- **Note:** Up to about 3 hours of video per request at low resolution, or 1 hour at high resolution, on 1M-context models. Google names Gemini as the replacement for the Video Intelligence API.
- **Sources:** [Video understanding docs](https://ai.google.dev/gemini-api/docs/video-understanding)

## Azure AI Video Indexer

A Microsoft service that runs many audio and video models together and builds an index of insights.

- Transcription in 50+ languages, translation, captions, and speaker identification
- OCR, labels, objects, scenes, shots, and keyframes
- Faces and celebrities (limited access, application required)
- Topics, keywords, sentiment, and summaries with Azure OpenAI

- **Best for:** Media and enterprise archives that need a wide set of ready-made insights, with an edge or on-premises option.
- **Where it runs:** Azure cloud, private endpoints, and an Azure Arc extension for edge and on-premises Kubernetes.
- **Pricing:** Per input minute, by preset. See the Azure pricing page. A trial account includes free minutes.
- **Agents / MCP:** No official MCP server.
- **Note:** Files up to 6 hours (12 hours for the Basic Audio preset).
- **Sources:** [Overview](https://learn.microsoft.com/en-us/azure/azure-video-indexer/video-indexer-overview), [Pricing](https://azure.microsoft.com/en-us/pricing/details/video-indexer/)

## Amazon Rekognition Video

An AWS computer-vision service for stored and streaming video.

- Object, scene, and activity labels
- Content moderation and text detection
- Face detection, face search, and celebrity recognition
- Shot detection and person pathing

- **Best for:** AWS teams that need structured labels, faces, and moderation on stored video.
- **Where it runs:** AWS cloud.
- **Pricing:** Per minute of video, per feature.
- **Agents / MCP:** No official MCP server found.
- **Note:** Streaming video analysis is closed to new customers from 30 April 2026. Stored video analysis is not affected.
- **Sources:** [Features](https://aws.amazon.com/rekognition/video-features/), [Availability changes](https://docs.aws.amazon.com/rekognition/latest/dg/rekognition-availability-changes.html), [Pricing](https://aws.amazon.com/rekognition/pricing/)

## Mixpeek

A semantic retrieval layer that turns video, images, audio, and documents into searchable features.

- Feature extractors: scenes, keyframes, transcripts, faces, and embeddings
- Semantic, face, object, and transcript search
- Multi-stage retrieval pipelines
- Reverse video search

- **Best for:** Teams that want to configure a multimodal retrieval pipeline on top of third-party models.
- **Where it runs:** Managed (multi-tenant or dedicated), or in your own AWS, GCP, or Azure account.
- **Pricing:** Monthly plans plus processing, or per stored vector. See the Mixpeek pricing page.
- **Agents / MCP:** Official MCP servers for ingestion, retrieval, and admin.
- **Sources:** [Pricing](https://mixpeek.com/pricing), [MCP server](https://www.mixpeek.com/docs/integrations/developer-tools/mcp-server)

## OpenAI API

General language models with image input. The official docs list image input, not video input.

- Answers about frames that you extract and send as images
- Long context for many frames at once
- Transcription through separate audio models

- **Best for:** Teams that already use OpenAI and can build their own frame sampling and transcription.
- **Where it runs:** OpenAI cloud, and Azure OpenAI.
- **Pricing:** Per token. Each frame is billed as an image.
- **Agents / MCP:** Not specific to video.
- **Note:** There is no built-in shot detection, video index, or video search. OpenAI’s cookbook samples frames with OpenCV.
- **Sources:** [Vision guide](https://developers.openai.com/api/docs/guides/images-vision), [Video cookbook](https://developers.openai.com/cookbook/examples/gpt_with_vision_for_video_understanding)

## Google Cloud Video Intelligence API

A classic computer-vision API with labels, shots, OCR, logos, objects, and transcription.

- Labels, shot changes, and explicit content
- Speech transcription, OCR, and logos
- Object tracking, faces, and people

- **Best for:** Existing users only, while they migrate.
- **Where it runs:** Google Cloud.
- **Pricing:** Per minute, per feature, with free monthly minutes.
- **Agents / MCP:** No MCP server.
- **Note:** Deprecated on 14 September 2026. It shuts down on 14 September 2027. Google recommends Gemini instead.
- **Sources:** [Deprecation notice](https://docs.cloud.google.com/video-intelligence/docs/deprecations)

## FAQ

### What is the best video understanding API?

It depends on the job. Twelve Labs is strong for video search and embeddings. Gemini is flexible for questions about a single video. Azure AI Video Indexer has a wide set of ready-made insights. Amazon Rekognition suits AWS teams that need labels, faces, and moderation. Video Context is built for teams that want structured video data, automations, and agent tools in one API.

### Is the Google Cloud Video Intelligence API being shut down?

Yes. Google deprecated the Video Intelligence API on 14 September 2026, and it shuts down on 14 September 2027. Google recommends moving to Gemini models.

### Can the OpenAI API understand video?

The OpenAI docs list image input, not video input. To use OpenAI models on video, you extract frames yourself, send them as images, and transcribe the audio separately.

### Which video APIs can run on my own infrastructure?

Azure AI Video Indexer has an Azure Arc extension for edge and on-premises use. Mixpeek can run in your own cloud account. Twelve Labs offers private deployment on request. Video Context can run in your cloud account or on your own servers.

### Which video APIs work with AI agents through MCP?

Twelve Labs and Mixpeek have official MCP servers. Video Context gives beta partners an MCP server and an agent setup guide.
