Jul 2026· Proceedings of the National Academy of Sciences of the United States of America· Vol 123· 1 citation· 58 references
Medicine
Abstract
Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations”—wherein models make up facts, law, and precedent—leading Chief Justice Roberts to spotlight this risk in his annual report on the judiciary. We argue that there is a need for public AI benchmarking in law. First, relative to other AI application domains, the legal AI ecosystem lacks legibility—there is little information about the design and performance of many commercial legal AI systems. Legal AI has not benefited from the types of benchmarking that have catalyzed, measured, and informed AI innovation and responsible use in other domains. Second, we articulate the challenges of the institutional design of benchmarking. We illustrate how benchmarks can be captured, watered down, and abused. Careful institutional design around the why, who, what, and how of benchmarking will be critical to navigate difficult tradeoffs of transparency, objectivity, expertise, and resources. Third, addressing legal AI’s illegibility requires matching institutional models to available resources and constraints. Rather than advocating for a single “best” approach to benchmarking, we show how benchmarking strategies depend on available resources.
Designed for judges, prosecutors, lawyers, judicial training institutions, and legal educators, the Global Toolkit on AI & the Rule of Law for the Judiciary is a practical gateway into the fast‑evolving world of artificial intelligence in justice. It explains AI in clear, accessible terms, showing how algorithms are already shaping investigations, evidence assessment, case management, and even judicial decision‑making. Through modular lessons, real‑world examples, and guided reflections, it helps readers recognize both the promise of AI for efficiency and access to justice, and the dangers it poses to equality, transparency, privacy, and due process if left unchecked.
The Toolkit can be used directly in courses and workshops, inspiring new activities, teaching methods, and discussions that facilitate and improve the learning experience in judicial training institutions. It equips justice actors to ask the right questions, spot risks like bias and opacity, and uphold human rights standards whenever AI enters the courtroom.
Through this Toolkit and a growing portfolio of trainings, partnerships, and policy support, the AI &Rule of Law programme is empowering judiciaries worldwide to steer AI in service of justice.
M. Stankovich, Ivana Feldfeber, Y. Quiroga et al.· 5 citations
It is argued that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Iris Marion Young's theories of oppression and structural injustice.
Abstract While Artificial Intelligence (AI) may transform legal research within hours and deliver results in just a few seconds, it can also fabricate cases, a process known as “hallucination.” This article offers a legal practice perspective in the Philippines under the Supreme Court’s Code of Professional Responsibility and Accountability (CPRA) and within the context of ASEAN Governance. The CPRA reaffirms that professional responsibility in the practice of law cannot be delegated to digital devices. The Supreme Court of the Philippines approved the Governance Framework on the Use of Human-Centered Augmented Intelligence in the Judiciary through a resolution dated February 18, 2026 (A.M. No. 25-11-28-SC), based on documented instances of AI-generated fake citations from foreign jurisdictions, and in response to a growing number of U.S. courts issuing standing orders requiring disclosure and human verification of all AI-generated content in court submissions. Alongside the ASEAN Guide on AI Governance and Ethics (2024) and the article’s proposed three-step verification workflow (existence-context-status), the following governance recommendations are suggested to anchor responsible one-click practice. Empirical research, however, shows that mistakes may still be made, and most importantly, Retrieval-Augmented Generation (RAG) systems do not necessarily prevent hallucinations since they do not ensure the accuracy of the text but rather rely on a retrieval and generation mechanism, which provides context for the text. As a result, it is the duty of lawyers in the Philippines to use their professional judgement and common sense in the use of any technological tool. The following recommendations are based on the need to ensure the integrity, credibility and truthfulness of the justice system in a technology rich world.
Rizal Thaddeus Acas, Jeanne Alejo-Abitago· International Journal of Dig...· 0 citations
Lawyers and self-represented litigants are already using artificial intelligence (AI) to draft legal documents, and courts are responding with rules. After more than 1,500 cases involving AI hallucinations, lawyers have been instructed to perform careful, independent review of AI-assisted filings. Discharging these duties requires what the human-computer interaction (HCI) literature calls ``appropriate reliance,''which cannot be calibrated without evidence on how often, how badly, and how detectably these tools fail at legal work. Existing research barely describes any of the three. We analyze the official record of the New York court system. The documents repeatedly call for evidence that does not exist (e.g., error rates, do-not-use lists). In its place they invoke procedure, including training mandates, checklists, and uncalibrated human review. The burden falls hardest on those least equipped to bear it: legal aid programs are told to track their own error rates, and judges are left to improvise their own tests. The paper makes four contributions: (1) a mapping from the legal duties to concepts in HCI; (2) a set of requirements elicited from the official record; (3) an analysis of how the legal system substitutes process for evidence; and (4) a research agenda for computing, including task taxonomies, shared error metrics, maintained benchmarks, and test harnesses for evaluations on private data. The computing community must supply what the justice system lacks; in doing so, it can help close, rather than widen, the justice gap.
Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized'precedent overfitting'bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.
Angel Mary John, Vipin Kumar Singh, J. T. Panachakel· 0 citations
This article argues for a re-evaluation of the current orthodoxy in artificial intelligence (AI) regulation, driven by the inconvenient reality that existing mechanisms for measuring AI capabilities are inadequate for the task of evaluating general-purpose AI. Lawmakers and legal scholars alike have drastically overestimated our understanding of frontier AI systems and how they work. Nascent governance frameworks often assume a basic capacity for evaluating general-purpose AI that is simply not supported by the technical literature – they are built on a house of cards. The rapid pace of progress on the AI frontier has taken general-purpose AI past the point where a human expert can reliably interpret or interrogate their behaviour, particularly to the legal standards that lawmakers have codified to date. This article explores how technological complexities within general-purpose AI create sui generis challenges for regulators and the law. Faced with the prospect of an anthropocentric ceiling to efforts to understand frontier AI systems, this article proposes a new paradigm for lawmakers, Sentinel Governance, grounded in governance-oriented innovation and experimentation to supplement human oversight of AI. New mechanisms for evaluation and enforcement are needed to avoid AI regulations subsiding into a checkbox exercise.
Matt Bartlett· The Cambridge Law Journal· 0 citations