raw-llm-judges-show-position-bias-gpt-4
observationreplicated

As a raw pairwise judge with no order-swap mitigation, GPT-4 can have its verdict flipped simply by swapping which candidate answer is shown first; applying a position-swap protocol restores agreement with careful human raters to close to human-human agreement levels.

Capability: Biased when judging other outputs

Observed on

GPT. 2023, GPT-4 class. LLM as judge.

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 10936 total — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena506 in 12mo · 1242 total — Large Language Models are not Fair Evaluators
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims