llm-agent-must-reconcile-user-supplied-identity-credentials-against-database
observationsingle paperpending review

When an LLM agent must reconcile user-supplied identity credentials against database records before a sensitive read or write, frontier and open models frequently skip the cross-field consistency check and act anyway, and the failure rate barely moves whether the request is simple or has several parallel sub-requests and whether the forged field is visually near-identical to the true one or completely unrelated.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Checking claims against evidence

Observed on

Multi-turn simulated retail customer-service tasks with stateful tools and a domain policy requiring credential verification; nine models, one trial per task, greedy decoding, poli.

Sources

  • Measured on 93 retail tasks within a 238-task benchmark, scored by an assertion model plus human trajectory verification; best retail score was 0.7312 (Claude Haiku 4.5) with several models far lower. The complexity and similarity insensitivity is reported qualitatively without per-condition numbers, so that part is argued more than quantified. Single trial per task limits variance estimates.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: low. Falsifier as drafted: Showing that detection and clarification rates rise substantially as the forged credential becomes more dissimilar from the true one, or as the number of parallel requests drops, under the same harness. Proposed technique, not catalogued: conflict-injection agent benchmark.