An AI agent passed 77% of runs but went five for five on only 53% of tasks
IBM researchers tested a ReAct agent powered by GPT-4.1 on 168 AppWorld tasks. They ran every task five times. The agent passed 77.4 percent of all runs, but it passed all five runs on only 53.0 percent of the tasks. The researchers call the 24.4-point difference the consistency gap.
Verified 12:36 AM PDT · 2 original sources
Their method reviews the agent's recorded steps and reruns individual decision points. It flags decisions that are likely to change on another run. The method then stores short guidelines about those weak points in memory.
On the same tasks, the method raised the share that passed all five runs by 16 points, to 69.0 percent. It raised that share by 13 points on similar tasks. The average pass rate also increased.
The authors tested one benchmark, two model backends and one agent pattern. No independent team has reproduced IBM's research claim. The paper is a preprint, and the same researchers built and tested the method.
AppWorld uses controlled API tasks rather than live customer work. Requiring five successful runs may be stricter than some production settings need. It still catches failures that a one-run benchmark can miss.
Independent teams could repeat the test on browser, coding and enterprise agents. Production teams could also rerun the same task over time, after model updates and during real tool failures.
Teams could publish average success rates alongside repeat success rates. Adding repeatability scores to model cards would help buyers compare reliability without depending on one successful run.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
- arXivClosing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course ↗
- Hugging Face and IBM ResearchYour Agent Aced the Task. Will It Do It Again? ↗
Continue the edition