Private beta: we are looking for design partners
Video Context

Video transcription

Transcripts that know
what is on screen.

A video transcription API turns the speech in a video into text with timestamps. Video Context also labels each speaker and adds the text that shows on screen, such as names, slides, and scores, so one JSON file holds everything that the video says.

REQUESTPOST /v1/transcripts
const transcript = await vc.transcripts.create({
  video: 'https://example.com/press-briefing.mp4',
  speakers: true,
  onScreenText: true,
})
● ILLUSTRATIVE RESPONSEPREVIEW API
{
  "video_id": "vid_7c02",
  "segments": [
    {
      "start": 12.4,
      "end": 18.9,
      "speaker": "S1",
      "text": "We will open the new road in March."
    }
  ],
  "on_screen_text": [
    { "start": 11.0, "end": 20.0, "text": "Minister of Transport" }
  ]
}

This is the planned API design. Video Context is in private beta, and beta partners help set the final design of each endpoint.

More than the words

Most transcription APIs read the audio only. But a video also shows information: the name of the speaker in a lower third, the slide behind them, or the score in the corner. Video Context puts the speech and the on-screen text on one timeline.

What you get

  • Segments with timestamps. Each segment has a start time, an end time, and the text.
  • Speaker labels. Each segment shows who speaks.
  • On-screen text. Names, titles, slides, captions, and scores, with the times when they show.
  • One schema. The transcript uses the same timeline as scenes, shots, and objects. See the video understanding API.

Where teams use it

  • News. Find every quote from one person across a week of briefings. See Video Context for news.
  • Search. Make an archive searchable by what people say. See the video search API.
  • Captions and translation. Turn the segments into SRT or WebVTT files.

FAQ

Questions, answered.

Common questions about video transcription with Video Context.

What is a video transcription API?+

A video transcription API takes a video file or URL and returns the speech as text, with timestamps. Software can then search, caption, translate, or summarise the video.

How is this different from a speech-to-text API?+

A speech-to-text API reads only the audio. Video Context also reads the picture, so the result includes the speaker labels and the text on screen, such as lower thirds, slides, and scoreboards, on the same timeline.

Can I get captions from the transcript?+

Yes. The segments have start and end times, so you can turn them into SRT or WebVTT captions.

Is the Video Context API available now?+

Video Context is in private beta. Book a call with the team to get access.

Your videos contain
valuable data.
Start using it.

Analyse video at scale. Find exactly what you need. Give your products and AI the context to act on it.