video-understanding-instruction-contains-several-conditional-branches-model-must-pick
observationsingle paperpending review

When a video-understanding instruction contains several conditional branches and the model must pick the branch matching what the video shows, both proprietary and open multimodal models pick correctly far more often when the correct branch is listed first, and accuracy falls as the correct branch moves later in the list.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Following an unfamiliar procedure

Observed on

Multimodal LLMs given video plus a selection-style instruction with three conditions plus an else branch; measured as task pass rate with the branch set held fixed and only the cor.

Sources

Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: Running the same controlled position swap on a broader set of models and finding pass rate flat across branch positions, or higher for later positions.