tool-use-reliability-in-customer-support-gpt-4o
observationsingle paper

In simulated retail customer-support tasks with real tools and policy documents, GPT-4o — the best performer in the study — still completed well under two-thirds of tasks correctly, and its rate of passing the identical task eight times in a row (a proxy for production reliability) was far lower than its single-attempt success rate.

Capability: Using the tools it is given

Observed on

GPT. 2024, GPT-4o. Customer support.

Sources

Status: activeLast checked: 2026-09-03Evidence activityHow much the field cites the sources under this claimvery heavily cited in the last 12 months889 in 12mo · 1075 total — τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims