external-feedback-repair-works-only-with-real-grounding
mechanismreplicatedcontested

Reflect-and-retry raises task success when the feedback in the loop is a genuine external signal — a failing test, a compiler error, an environment outcome — but not when the "feedback" is the model's own unaided critique. The grounding, not the reflection step, is what does the work.

Capability: Fixing its own mistakes · Agentic, Coding agent

Observed on

2023, GPT-3.5/GPT-4 class.

Sources

Suspected axis of disagreement(a guess, not verified)

Probably task family and how "improvement" was measured. The contesting result is strongest on open-ended generation, where better is a fuzzy judgment a second pass can nudge; the supporting results are on checkable tasks with a pass/fail signal. A secondary possibility is baseline setup — whether the no-revision comparison got the same inference budget. I have not reconciled the experimental setups myself.

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 5075 total — Reflexion: Language Agents with Verbal Reinforcement Learning1000+ in 12mo · 4530 total — Self-Refine: Iterative Refinement with Self-Feedback503 in 12mo · 1164 total — Large Language Models Cannot Self-Correct Reasoning Yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Filed as an efficacy claim rather than as a property of the technique, so the scope condition ("only with real grounding") and the contest are both first-class. Marked replicated because two independent papers converge on grounding being the active ingredient from opposite directions, not because any single result was rerun.