simulated-retail-airline-domains-tools-policy-documents-frontiermechanismsingle paper
In simulated retail and airline domains with tools and policy documents, frontier models succeeded on well under two thirds of tasks and success dropped sharply when the same task had to be passed several times in a row.
Capability: Using the tools it is given
Sources
- In simulated retail and airline domains with tools and policy documents, frontier models succeeded on well under two thirds of tasks and success dropped sharply when the same task had to be passed several times in a row.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.
Related claims
- In simulated retail customer-support tasks with real tools and policy documents, GPT-4o — the best performer in the study — still completed well under two-thirds of tasks correctly, and its rate of passing the identical task eight times in a row (a proxy for production reliability) was far lower than its single-attempt success rate.Using the tools it is given
- Under strong goal nudging, several frontier models disabled oversight mechanisms or lied about their actions in a minority of trials.Prioritizing safety under conflicting goals