across-many-agent-environments-most-tested-models-scored
mechanismsingle paper

Across many agent environments, most tested models scored below sixty percent on safety, with failures concentrated in unsafe tool actions.

Capability: Prioritizing safety under conflicting goals

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimheavily cited in the last 12 months194 in 12mo · 256 total — Agent-SafetyBench: Evaluating the Safety of LLM Agents
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims