forecasting-scaffolded-system-approaches-human-crowd-accuracyobservationsingle paper
A retrieval-augmented forecasting system built on a GPT-4-class model — not the base model alone — approaches, and sometimes exceeds, the accuracy of competitive human forecasters on real prediction-market questions dated after the model's training cutoff.
Capability: Predicting future events · Reasoning
Observed on
2024, GPT-4 class with retrieval scaffolding. General.
Sources
- Achieved accuracy comparable to, and sometimes exceeding, human crowd forecasts when tested against real forecasting-tournament questions with dates after training cutoff.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Without added retrieval infrastructure, language models underperform human experts at forecasting real-world events from a benchmark of actual forecasting-tournament questions, though accuracy improves with model scale and access to relevant news context.Predicting future events
- When research-idea forecasts are scored by whether a later paper matches them under an LLM judge rubric, higher scores partly reflect broader, less specific ideas: the backbone with the higher hit rate also produced measurably more general forecasts and had more retrieved candidate papers pass the gate, and tightening the specificity threshold does not separate the two.Whether the measurement made the finding · unreviewed
Notes
Requires real infrastructure (search, retrieval, aggregation) built around the model — not an out-of-the-box capability. Worth reading alongside the earlier, lower-scaffolding result below to see the progression.