measured-rd-uplift-is-below-the-self-sustaining-thresholdOn the best current calibration, AI is not yet accelerating its own development in a self-sustaining way: the modelled threshold is that a one-unit gain in model capability must buy at least 15% higher AI R&D productivity, and the back-of-envelope figure from reported engineer uplift since coding agents launched is about 9% — below it, but rising, so the gap is a current reading rather than a ceiling.
Capability: Whether the measurement made the finding · Evaluation
Observed on
2026, calibrated on the Epoch Capabilities Index and reported engineer uplift.
Sources
- Derives the self-sustaining acceleration condition from the elasticities around the core feedback loop, then calibrates it: the condition is met at roughly 15% R&D productivity gain per unit of capability, and the reported figure is around 9%. States plainly that the return is not constant and has been increasing, so the finding is that we are not there yet, not that we cannot get there.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- When a calculator-using model is trained with reinforcement learning on verifiable final-answer rewards for an arithmetic search task (Countdown), the gain lands almost entirely at low k — pass@1 rises sharply while pass@16 moves little or falls — because the update can only reinforce correct trajectories the starting policy already samples, and prompt groups with no correct sample supply no gradient.Using the tools it is given · unreviewed
- In agent pipelines that generate whole domain-specific ML codebases (medical imaging) from scratch, execution/assembly validation and domain knowledge bases fix different failures: removing runtime and assembly testing mainly costs autonomy (more human interventions to debug), while removing the curated domain knowledge base mainly costs output quality (syntactically fine but domain-inappropriate pipelines) with little change in intervention count.Generating and editing working code · unreviewed
- A large share of an agent benchmark score is attributable to the harness around the model rather than the model itself: holding the model fixed and changing only guides, sensors, retries and state management is reported to move the same model from 30.9% to 74.6% on GAIA and 30th to 5th on Terminal Bench — so a score reported for "a model" on an agentic benchmark is a score for a model-plus-harness, and comparisons across labs with different harnesses are not comparisons of models.Whether the measurement made the finding · unreviewed
- Reporting how much a lab's models are used in its own research — tokens, lines of code, inference compute, experiments per researcher — cannot show that the models are accelerating the research, because every one of those is an input. Establishing the loop needs an outcome variable over time: algorithmic efficiency as a function of capability, which is the edge that turns a pipeline into a feedback loop.Whether the measurement made the finding
- The large reported gains from fine-tuning on tool-call trajectories come from small open models starting far behind — distilling a few hundred trajectories from a stronger model moves them a long way. Whether the same procedure helps a model that is already good at tool use is a separate question this evidence does not answer.Using the tools it is given
Notes
Filed as an observation, not a mechanism. The 15% threshold comes from a model whose structure could be wrong, and the 9% is a back-of-envelope from self-reported uplift — both will move. The durable part is the shape of the argument, which is held separately in input-telemetry-cannot-establish-self-improvement. Relevant to this project rather than merely interesting: docs/ambitions.md proposes that this system improve itself from its own catalog. This is the serious attempt to say what would have to be true for that to be happening, and its answer is a ratio nobody currently publishes.