video-temporal-reasoning · activeReasoning about time in video
Tracking motion, ordering, causality, and change across video frames rather than treating video as isolated stills.
Tags: reasoning, context
A model can describe a single frame well and still fail at what makes video video: what happened before what, how something moved, whether an action completed. This capability is about the temporal dimension specifically — grounding an answer in a span of time rather than a moment.
What counts as this capability
Scope boundary used when deciding whether a paper is really about this capability, rather than merely mentioning it.
In scope when the difficulty is temporal: event ordering, tracking across frames, before/after reasoning, long-video retrieval where position in time matters. NOT in scope: per-frame recognition that happens to use video as a source of images, or captioning benchmarks where a single keyframe would answer the question.
Claims
Techniques
None yet.
Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.