AI2 and the University of Washington conducted an experiment on Harness Evolution, a popular direction in Agent research, comparing it with simple test-time scaling methods (such as "running multiple times") under the same feedback and similar inference budgets. The experimental results indicate that the complex Harness Evolution method does not consistently outperform simple methods, and the improvements in Harness evolved by it on new tasks are limited. Some of the gains are closer to task-specific adaptations rather than optimizations of a general Agent architecture. The study also reveals that the performance improvements of some current Harness Evolution methods may primarily stem from increased computational resources, rather than optimizations of the Harness itself. Therefore, when evaluating Agent systems in the future, it is essential to use similar inference budgets and external feedback as standards and explore task benchmarks that better reflect the value of Harness.
