Stanford Study Reveals Alarming Hallucination Rates in Legal AI Tools
The latest study from Stanford University's Institute for Human-Centered Artificial Intelligence (HAI) presents troubling findings about the reliability of specialized AI tools designed for legal research. The results show that even the most advanced legal AI systems produce inaccurate or unsupported information in one out of every six queries, posing a serious risk to legal practice.
How the Study Was Conducted
The study, conducted by Stanford HAI and RegLab, focused on evaluating leading AI tools for legal research, with an emphasis on their tendency to "hallucinate"—that is, to generate false or unsupported legal information. The researchers tested two major legal AI tools: Lexis+ AI from LexisNexis and Westlaw AI-Assisted Research/Ask Practical Law AI from Thomson Reuters, comparing their results with general-purpose models such as GPT-4. The study's methodology was carefully designed to cover a broad range of legal queries. The researchers manually compiled a dataset containing more than 200 open-ended legal queries designed to test different types of questions. The categories tested included general research questions concerning legal doctrine, court decisions, and bar exam questions. The study also focused on questions specific to a particular jurisdiction or time period, including differences between judicial circuits and recent changes in the law. Questions based on false premises, which simulated user misunderstandings, were also tested, as were factual-recall queries concerning objective legal facts.
Troubling Study Results
The study's results were deeply troubling, particularly given the critical importance of accuracy in the legal field. Lexis+ AI and Ask Practical Law AI produced hallucinated information in more than 17 percent of the test queries, meaning approximately one incorrect or unsupported result out of every six queries. Westlaw AI-Assisted Research performed even worse, with a hallucination rate exceeding 34 percent. Although these rates are lower than those of general-purpose models such as GPT-4, they still pose a substantial risk in a field where accuracy is critical. The study defines a hallucination as either directly incorrect information or a false claim that a cited source supports a particular assertion when it actually does not. This definition is essential to understanding the severity of the problem, because hallucinations can take different forms, all of which pose a potential risk to legal decision-making.
The findings raise fundamental questions about vendors' claims regarding the reliability of their tools. As the researchers noted, hallucinations in legal AI tools remain "substantial, widespread, and potentially insidious." Although the tools do reduce errors compared with general-purpose AI models such as GPT-4, which represents an improvement, even these specialized legal AI tools still hallucinate at an alarming rate. The study's results have far-reaching implications for the legal profession. Given that nearly three-quarters of lawyers plan to use generative AI in their work, it is essential for the legal community to be aware of these tools' limitations. The study highlights the need for ongoing benchmarking, transparency, and caution when deploying AI for high-stakes legal research.
Expansion of the Study
Stanford researchers are considering expanding the study in response to questions raised about its methodology and the fairness of the testing. However, the central finding remains unchanged: even specialized legal AI tools are not free from hallucinations. This finding is critically important for legal practice, where a single error can have serious consequences for both clients and lawyers themselves. The study also reveals a need to better educate lawyers about the limitations of AI technology. Many lawyers may not understand how often these tools produce incorrect information, which can lead to overreliance on AI outputs without proper verification. Legal education and continuing professional development should include topics on the responsible use of AI and the importance of verifying AI-generated information.
The Future of AI in the Legal Sector
The future of legal AI will likely depend on the development of more reliable systems and better practices for using them. Technology vendors will need to communicate their products' limitations more transparently and invest in research and development aimed at reducing hallucination rates. At the same time, the legal community will need to develop standards and best practices for integrating AI into legal practice in a way that minimizes risks and maximizes benefits.
One of the best-known cases of AI misuse in legal practice is the 2023 case involving New York lawyer Steven A. Schwartz. Schwartz, from the law firm Levidow, Levidow & Oberman, used ChatGPT to conduct legal research in a personal injury case filed in federal court. The AI generated six entirely fictitious court decisions with fabricated citations and commentary, which Schwartz included in his legal filing without verification. When Judge P. Kevin Castel noticed these fake cases and confronted Schwartz about them, the lawyer admitted to using ChatGPT but initially defended his actions by saying that he had been unaware that the content could be false. Schwartz even submitted records of his conversations with ChatGPT in which he asked the AI whether the cases were real, and the AI falsely confirmed that they were legitimate and "can be found in reputable legal databases such as LexisNexis and Westlaw." Schwartz and his colleague Peter LoDuca were ultimately fined $5,000 for misleading the court and were required to send apologies to all the judges mentioned in the fabricated citations.



