forecasting-underperforms-humans-without-scaffolding
observationsingle paper

Without added retrieval infrastructure, language models underperform human experts at forecasting real-world events from a benchmark of actual forecasting-tournament questions, though accuracy improves with model scale and access to relevant news context.

Capability: Predicting future events · Reasoning

Observed on

2022, pre-GPT-4-era models. General.

Sources

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimsteadily cited in the last 12 months33 in 12mo · 68 total — Forecasting Future World Events with Neural Networks
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Not contested by the 2024 scaffolded result above — different eras and setups, showing a real progression over time rather than a disagreement. A clean example of why claims need a date and an era, not a single timeless verdict.