scientific-tasks-where-domain-specific-constraint-general-presentation-constraint-letter
observationsingle paperpending review

On scientific tasks where a domain-specific constraint and a general presentation constraint (letter case, output format, structure) are both stated in the same prompt, multimodal models satisfy the domain-specific constraint more often than the general formatting one, so domain competence does not imply full instruction compliance.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Following an unfamiliar procedure

Observed on

2,527 constraint-injected samples built from 13 existing scientific datasets across chemistry, geography, biology, materials science and physics; four closed-source and seven open-.

Sources

  • Measured with per-constraint DRFR split by constraint domain and group: the reported gap for the top closed model is roughly 89% on scientific versus 75% on general constraints, and letter constraints sit near 51%. Constraints were injected into existing problems, so difficulty of the two constraint types is not matched — the gap could partly reflect that general constraints here are harder in absolute terms, not that they are neglected.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: A benchmark of the same kind where models score equal or higher on general formatting constraints than on injected discipline-specific constraints, across model families. Drafted stance toward verifiable-instructions-such-length-format-constraints-strong-models: supports -- The paper reports that adherence falls as constraint counts rise and that general format/letter constraints are the weakest group even for the strongest model.