forecasting-underperforms-humans-without-scaffoldingobservationsingle paper
Without added retrieval infrastructure, language models underperform human experts at forecasting real-world events from a benchmark of actual forecasting-tournament questions, though accuracy improves with model scale and access to relevant news context.
Capability: Predicting future events · Reasoning
Observed on
2022, pre-GPT-4-era models. General.
Sources
- Introduced the Autocast forecasting benchmark; models underperform human experts, but accuracy scales with model size and access to relevant news.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- A retrieval-augmented forecasting system built on a GPT-4-class model — not the base model alone — approaches, and sometimes exceeds, the accuracy of competitive human forecasters on real prediction-market questions dated after the model's training cutoff.Predicting future events
- In partially observable text environments (ALFWorld, ScienceWorld), letting an LLM agent query an externally maintained state estimate that keeps an explicit distribution over unobserved object locations raises task success more than querying a deterministic memory of what has already been observed, and the gain shrinks to near zero on a frontier model that already nearly saturates the benchmark.Tracking state through a long task · unreviewed
- When research-idea forecasts are scored by whether a later paper matches them under an LLM judge rubric, higher scores partly reflect broader, less specific ideas: the backbone with the higher hit rate also produced measurably more general forecasts and had more retrieved candidate papers pass the gate, and tightening the specificity threshold does not separate the two.Whether the measurement made the finding · unreviewed
Notes
Not contested by the 2024 scaffolded result above — different eras and setups, showing a real progression over time rather than a disagreement. A clean example of why claims need a date and an era, not a single timeless verdict.