self-feedback-improves-generation-quality
mechanismsingle papercontested

Prompting a model to generate its own feedback on a draft output and revise accordingly improves quality on open-ended generation tasks (dialogue, code, review-writing) over a single-shot attempt.

Capability: Fixing its own mistakes · Reasoning

Observed on

2023, GPT-3.5/GPT-4 class.

Sources

Suspected axis of disagreement(a guess, not verified)

Likely how rigorously the baseline (no-revision) condition was set up and evaluated, and possibly task family — open-ended generation tasks (where "better" is fuzzier and easier to nudge with a second pass) vs. the more checkable reasoning benchmarks in the contesting paper. I haven't reconciled the two papers' exact experimental setups myself.

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 4530 total — Self-Refine: Iterative Refinement with Self-Feedback503 in 12mo · 1164 total — Large Language Models Cannot Self-Correct Reasoning Yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

This is the pair I'd most want to actually dig into further — the disagreement axis above is a first guess, not a verified reconciliation.