building-preference-data-teach-7-8b-instruct-models-follow-multi-constraint
observationsingle paperpending review

When building preference data to teach 7-8B instruct models to follow multi-constraint instructions, perturbing each atomic constraint in several directions so the perturbed instruction's valid-response set contains, partly overlaps, or is disjoint from the original's — and pairing responses drawn from each sub-region — raises constraint adherence more than rejection sampling or teacher-correction pairs, and the gain holds on perturbed instruction variants the baselines barely improve on.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Following an unfamiliar procedure

Observed on

Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, DPO/KTO/online-DPO, 12k pairs, constraints that are checkable by executable code; tested on IFEval, IFBench, a perturbed set built by.

Sources

  • Two backbones, three optimization methods, four test sets, fixed 12k training size, with ablations removing each relationship type and each pair region — each removal lowers accuracy. One of the four test sets (Perturbed-IF) is generated by the authors' own perturbation method, which favors the method; the data-scaling curve and the fixed budget do argue the gain is not just more data. Single group, no independent replication.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: medium. Falsifier as drafted: Matching the same 12k pair budget with disjoint-only or random perturbations yields the same constraint accuracy, or the advantage disappears on a held-out benchmark not built with the same perturbation procedure. Proposed technique, not catalogued: Cross-relational preference pairs over perturbed constraints.