vlm-agents-must-verify-claims-against-interactive-visualizations-where
observationsingle paperpending review

When VLM agents must verify claims against interactive visualizations where no claim is answerable from the initial viewport, giving them an interaction budget of ten actions does not reliably beat answering immediately from the first screenshot — some models score slightly lower with interaction — because partial, unplanned exploration leaves the model with enough evidence to abandon its prior but not enough to replace it.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Checking claims against evidence

Observed on

Screenshot-only agents (no DOM), multi-view Vega-Lite notebooks, three-way True/False/NEI labels, 500-claim stratified subset; ablation run at budgets of 1, 10 and 25 actions for o.

Sources

  • Measured on 500 claims across fourteen model configurations. Gemini 3.5 Flash: 44.2% at one action, 41.6% at ten, 50.0% at twenty-five; Qwen 3.5 27B 46.2% vs 45.0%; Gemma 4 31B 45.4% vs 47.2%. Best overall was GPT-5.5 at 57.2% against a 33.3% chance baseline. The budget ablation covers only three models, and gaps are within a few points, so the dip is suggestive rather than established.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: low. Falsifier as drafted: A comparable evaluation where the same models gain consistently in accuracy as soon as any interaction is allowed, with no non-monotonic dip at intermediate budgets. Proposed technique, not catalogued: Interaction Efficiency Score (accuracy weighted by fraction of actions that change the observed state).