identifies-position-verbosity-self-enhancement-biases-strong-judges-while
mechanismsingle paper

Identifies position, verbosity, and self-enhancement biases in strong judges, while also showing high agreement with humans once mitigated.

Capability: Biased when judging other outputs

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months1000+ in 12mo · 10936 total — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.