grounded-reflection-improves-multistep-tasks
mechanismsingle paper

When a model gets a real environment signal after acting — a test result, a tool error, a task-success indicator — reflecting on that signal in words and retrying substantially improves success rates on multi-step coding and decision-making tasks over a single attempt.

Capability: Fixing its own mistakes · Agentic, Coding agent

Observed on

2023, GPT-3.5/GPT-4 class agents. Coding agent.

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 5075 total — Reflexion: Language Agents with Verbal Reinforcement Learning
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Worth weighing against how modest real-world resolution rates stayed on repository-scale coding tasks even with test feedback available in the same era (2023 models resolved only a few percent of real GitHub issues in SWE-bench) — the mechanism is real, but its magnitude on toy or narrow benchmarks likely overstates what it does on messier real tasks.