Why AI chat apps need frames
Most AI chat apps understand images much better than video. Claude does not accept video files. ChatGPT can sometimes take a video, but the upload limits are small and the support changes often. Gemini accepts video, but a long file can use a large part of the context.
A frame grid solves this. It puts many frames from the video into one image, in time order, with a timestamp on each frame. A model can read the grid like a storyboard: what happens, in what order, and at what time.
How many frames to use
| Video length | Frames | Why |
|---|---|---|
| Under 1 minute | 9 to 12 | Enough to see each step |
| 1 to 5 minutes | 16 to 24 | One frame every 10 to 20 seconds |
| 5 to 30 minutes | 24 to 48 | Use one frame per scene, so each shot appears once |
Each grid holds up to 16 frames. If you choose more frames, the tool makes more than one grid. Models read a few large grids better than many small images.
Even spacing or one frame per scene
- Even spacing is best for screen recordings, lectures, and videos that change slowly.
- One frame per scene is best for edited videos: ads, trailers, sports highlights, and social clips. The tool finds each cut with the same method as our free scene detection tool.
What is in the context pack
The context pack is a ZIP file with:
grid-1.png,grid-2.png, and so on: the frame grids.prompt.txt: a ready prompt that tells the model how to read the grids.context.json: the file name, length, resolution, and the time of each frame, for agents and scripts.
Tips for better answers
- Paste the grids before the prompt.
- Ask about times: “What happens between 0:30 and 0:45?” The model can read the timestamps.
- If the speech matters, add a transcript. See the video transcription API.
- For a long video, ask the model to describe each grid first, then ask your question.