static-code-benchmarks-saturate-need-continuous-refresh
mechanismsingle paper

Static, fixed-problem coding benchmarks saturate and leak into training data over time, so getting an honest read on current code-generation ability requires a benchmark that continuously adds newly-published problems rather than reusing an older fixed set like HumanEval.

Evidence for: Goodhart's law (holds)

Capability: Generating and editing working code · Reasoning

Observed on

2024, evaluating 52 models on 2023-2024 problems. Coding agent.

Sources

Status: activeLast checked: 2026-09-04Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 2073 total — LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Pairs with self-repair's SWE-bench claim (arxiv-2310-06770) for the repo-scale, real-issue-resolution side of code competence, versus isolated-problem generation here.