models-readily-adopt-single-counter-memory-passage-when-coherentmechanismsingle paper
Models readily adopt a single counter-memory passage when it is coherent, and when sources conflict they follow the majority and show confirmation bias toward their own beliefs.
Capability: Checking claims against evidence
Sources
- Models readily adopt a single counter-memory passage when it is coherent, and when sources conflict they follow the majority and show confirmation bias toward their own beliefs.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- Training the model to decide when to retrieve and to critique whether passages support its output improves factuality and citation accuracy.Checking claims against evidence
- For persistent-memory agents, any memory-write rule that decides using only recency and provenance is stuck on a single tradeoff — accepting more genuine preference updates means admitting more poisoned ones — because an adversary who can launder a claim through the user's own channel matches the statistics a genuine revision produces; conditioning the write decision on the inferred authenticity of the claim given the interaction history moves both axes at once.Remembering across sessions · unreviewed
- A model's tendency to shift political stance toward a bare demographic identity label is a separate vulnerability from its tendency to shift toward an explicitly stated user opinion: on open-ended US policy prompts, the models most moved by an identity label are among the least moved by a stated opinion, so a benchmark using only stated opinions misses identity-driven shift.Telling the user what they want to hear · unreviewed
- Models endorse widely held falsehoods, showing weak verification against what they know.Checking claims against evidence
- In multi-agent LLM systems where one agent starts with an erroneous shared belief, broadcasting every agent's full context at every step raises the rate of false assertions above doing no synchronization at all, because the error propagates to agents that were previously correct — and the harm appears only in tasks where one wrong fact cascades across semantically linked dimensions (destination to airport to weather), not where agent contexts are largely orthogonal.Stating false facts confidently · unreviewed