Why temporal grounding is hard
A model can often say what happens in a video. It is much harder to say when it happens, to the second. Temporal grounding needs the model to connect language to a precise span of frames and audio.
Where it is used
- Search. Return moments, not files.
- Editing. Set in and out points for a clip.
- Agents. Cite the exact time in a video as evidence.
Video Context returns a start and an end time for every result, so agents and editors can act on it.