The research team from the Hong Kong University of Science and Technology (HKUST) has introduced LexAgentHallu, a new benchmark designed to assess hallucinations—instances where AI generates plausible but incorrect or irrelevant information—in AI agents operating within the legal domain. This innovative benchmark shifts the focus from merely evaluating the final outputs of AI agents to scrutinizing their entire decision-making processes, or "execution trajectories."
LexAgentHallu incorporates a structured error classification system that takes into account both the legal content accuracy and the procedural integrity of the AI agent's actions. It includes a dataset of 3,414 cases spanning 17 legal categories and six distinct task types, providing a comprehensive framework for evaluation.
In testing 18 different configurations of AI agents, both proprietary and open-source, the researchers discovered that even the top-performing models exhibited hallucinations in 89% of their execution trajectories. Notably, among the trajectories that produced correct answers, an average of 68% still contained hallucinations related to legal content.
The study also revealed that errors were not randomly distributed. Mistakes made early in the process, such as misinterpretations of legal source hierarchies or missed procedural deadlines, often led to a cascade of subsequent errors. By dissecting hallucinations into identifiable fault types, LexAgentHallu offers a robust mechanism for enhancing the reliability of legal AI agents. It supports defining boundaries before content generation, verifying evidence during execution, and ensuring consistency before final output.
The findings underscore a critical point: the credibility of legal AI agents cannot be established based solely on the accuracy of their final answers. Instead, the entire reasoning chain must be rigorously examined to prevent the propagation of errors throughout the decision-making process.
