functional-correctness-testing-is-the-standard-for-grading-code
mechanismsingle paper

Grading generated code by whether it actually passes held-out functional tests, rather than by surface similarity to a reference solution, is what makes code generation gradeable at scale — and it's how the field has graded it since the earliest large benchmarks.

Capability: Generating and editing working code · Reasoning

Observed on

2021, Codex/GPT-3 class. Coding agent.

Sources

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 11293 total — Evaluating Large Language Models Trained on Code
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims