tool-use-reliability-in-customer-support-gpt-4oobservationsingle paper
In simulated retail customer-support tasks with real tools and policy documents, GPT-4o — the best performer in the study — still completed well under two-thirds of tasks correctly, and its rate of passing the identical task eight times in a row (a proxy for production reliability) was far lower than its single-attempt success rate.
Capability: Using the tools it is given
Observed on
GPT. 2024, GPT-4o. Customer support.
Sources
- GPT-4o succeeded on well under two thirds of retail tasks, with a far lower rate of passing the same task eight times in a row.
Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.