input-telemetry-cannot-establish-self-improvement
mechanismmechanism reasoning

Reporting how much a lab's models are used in its own research — tokens, lines of code, inference compute, experiments per researcher — cannot show that the models are accelerating the research, because every one of those is an input. Establishing the loop needs an outcome variable over time: algorithmic efficiency as a function of capability, which is the edge that turns a pipeline into a feedback loop.

Capability: Whether the measurement made the finding · Evaluation

Sources

  • Sets out the causal graph and marks each node by whether it was disclosed. Inference and experimental compute are reported; algorithmic improvements, algorithmic efficiency and capabilities are not; and the one edge marked "what we need to measure" is capability feeding back into the rate of improvement. Effort spent is not effect achieved.
  • The underlying work. The diagram in the post is this paper's causal graph, and the paper publishes the data wish list the post is asking OpenAI to fill: growth rate of algorithmic efficiency inside labs, how R&D spend splits across inputs, and direct measures of how much models contribute to research.
Status: activeLast checked: 2026-09-07Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Filed as mechanism-reasoning, not single-paper. The argument is structural and holds on its own — an input cannot evidence an output — and the source is a post rather than a study, so it can originate the claim but not carry measurement behind it. That ceiling is deliberate for kind: post. Recorded here because this catalog is subject to the same test. Its ambitions include improving itself from its own contents, and the honest measure of that is not tokens spent or drafts produced but a held-out outcome that moves — which is what npm run backtest exists to be. Anything else is our own green boxes.