raw-llm-judges-show-position-bias-gpt-4observationreplicated
As a raw pairwise judge with no order-swap mitigation, GPT-4 can have its verdict flipped simply by swapping which candidate answer is shown first; applying a position-swap protocol restores agreement with careful human raters to close to human-human agreement levels.
Capability: Biased when judging other outputs
Observed on
GPT. 2023, GPT-4 class. LLM as judge.
Sources
- Swapping the order of two candidate answers alone could flip GPT-4's preference between them.
- Identifies the same position bias (plus verbosity and self-enhancement bias) in strong judges, and shows that mitigating it restores high agreement with human raters.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Swapping candidate order alone can flip the judge's preference, so an evaluator can be gamed by ordering.Biased when judging other outputs
- GPT-4 could reliably name a well-known celebrity's parent, but was far less reliable naming the celebrity when given the parent — the same directional-recall asymmetry demonstrated in smaller fine-tuned models, showing up in a deployed model that wasn't fine-tuned for the test.Not generalizing "A is B" to "B is A"