code-generation-prompt-combines-prompt-enforced-output-format-json-or-xml
observationsingle paperpending review

When a code-generation prompt combines a prompt-enforced output format (JSON or XML), a persona, and urgency framing at once, pass@1 on function-level Python problems can fall well below the sum of each constraint's individual effect, even when each constraint alone is neutral or helpful — observed in the GPT-4o family and absent in the GPT-4.1 family and o3-mini.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Generating and editing working code

Observed on

164 HumanEval+ problems, greedy decoding, five OpenAI models; format enforced by prompt instruction rather than API-level constrained decoding; prompt length confounded with constr.

Sources

Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: A factorial replication on other models and benchmarks where triple-constraint pass@1 matches the additive prediction within noise, or where the GPT-4o-family gaps vanish once prompt-token length is held constant. Proposed technique, not catalogued: Factorial compound-prompt reliability testing.