ui-to-code-generation-where-rendered-screenshot-gives-genuine-external-feedback
mechanismsingle paperpending review

In UI-to-code generation, where a rendered screenshot gives genuine external feedback, iterative self-critique still drifts below the initial draft over ten rounds, because a code edit changes the rendering non-locally and a fix for one region breaks regions that were already faithful; restricting each round to one typed, scoped repair target and carrying past targets forward as an avoid-list reduces that drift.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Generating and editing working code

Observed on

HTML/CSS webpage reproduction from a screenshot, three frontier and three open Qwen VLMs, ten refinement rounds, VLM-judge visual-fidelity scoring; effects on open models are mixed.

Sources

  • Measured across 6 models x 3 benchmarks (18 settings): scoped rubrics beat naive self-evolution in 15/18 final-round and 14/18 best-round settings, average +1.20 overall judge points; naive refinement fell below the initial draft on Design2Code for both GPT models. A 60-sample human preference study on one benchmark for two models backs the main comparison. Scoring is by VLM judge, which the authors note is unstable on subtle differences; the coupling mechanism is argued and illustrated rather than directly measured.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafter linked technique "external-feedback-repair", which does not list this capability in addresses -- recorded, not asserted. Drafted confidence: medium. Falsifier as drafted: Show that naive self-refinement with rendered feedback improves monotonically over rounds on these benchmarks, or that restricting each round to a single scoped repair target gives no advantage over free-form critique at equal round budget. Proposed technique, not catalogued: Rubric-scoped iterative repair with history.