The promise of arbitration has always rested on a simple, seductive value proposition: it is faster, cheaper, and more final than the labyrinthine procedures of national courts. Parties are sold a streamlined dispute resolution mechanism — a private justice optimized for commercial efficiency. Yet as legal practice moves deeper into the age of generative artificial intelligence, a subset of the profession is actively sabotaging that value proposition, submitting briefs riddled with hallucinations: fabricated case law, phantom statutes, and invented evidentiary records generated by large language models.
When caught, the defense is almost always the same: a plea of ignorance, a blaming of the “black box,” a claim that the technology misled them. These excuses are no longer tenable. In 2026, the submission of hallucinated legal authority is not a technological glitch; it is a fundamental breach of the ethical duty of competence. It is a self-inflicted wound that destroys the credibility of counsel and imposes a verification tax on the entire system, driving up the very costs arbitration was designed to curtail.
§ 01 · The death of the “black box” defense
Start by dispensing with the myth that AI hallucinations are mysterious accidents. They are a feature, not a bug, of how large language models operate. These models are probabilistic engines, not truth machines. When a lawyer prompts an LLM to find a case where an arbitrator was removed for undisclosed conflicts in the energy sector, the model does not search a database of verified law. It predicts the most statistically likely sequence of words that looks like such a case. A lawyer who uses a general-purpose LLM for legal research without understanding this mechanism is akin to a pilot flying a plane without understanding the difference between an altimeter and a fuel gauge.
The standard for technological competence is already set in stone. Model Rule of Professional Conduct 1.1 and its Comment 8 explicitly require lawyers to keep abreast of changes in the law and its practice, including the benefits and risks associated with relevant technology. This is a non-delegable duty. You cannot outsource your professional judgment to a summer associate, and you certainly cannot outsource it to a chatbot.
§ 02 · The verification tax
The most pernicious effect of hallucinations in arbitration is economic. In a typical arbitration the tribunal is paid by the hour, and every hour spent chasing a citation that does not exist is an hour billed to the parties. Consider the procedural fallout of a single hallucinated citation. Opposing counsel spends hours attempting to locate the case, then must draft correspondence or applications to the tribunal pointing out the discrepancy. The arbitrator, bound by a duty to fully investigate the law, must personally verify the non-existence of the authority. The inevitable result is a show-cause order or a motion for sanctions — and the main dispute, the breach of contract or the intellectual property theft, is put on hold while the parties litigate the conduct of the lawyers.
This is the verification tax: a surcharge levied on the process by the incompetence of one party. Recent federal cases illustrate the scale of the waste — in 2025, one court ordered sanctioned counsel to pay over $26,000 to reimburse the opposing party for legal fees incurred investigating fake citations. In arbitration, where the loser-pays principle is often the default, a party that submits hallucinations is effectively writing a blank check to its opponent for costs.
The deeper damage is to trust, the primary efficiency driver of arbitration. An arbitrator who finds one fake case in a submission will instinctively distrust every other factual assertion and legal argument in that brief. The benefit of the doubt is lost; the tribunal is forced into a forensic, skeptical posture that slows the drafting of the award and increases the hours required to complete the mandate. And the stakes run past the fee. Under the New York Convention, an award must be enforceable. A tribunal that relies, even inadvertently, on a hallucinated legal principle renders its award vulnerable to set-aside proceedings — the losing party can argue the award is contrary to public policy, or that it was unable to present its case against fictitious law. Counsel who introduce hallucinations into the record are planting a poison pill in their own client's victory.
§ 03 · The case law of consequence
Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023), is the case that started it all. Plaintiff's counsel filed a brief citing at least six nonexistent cases, and when the court and opposing counsel could not find them, the lawyers doubled down — submitting “copies” of the fake opinions generated by the AI. The lesson: the court imposed a $5,000 sanction, but the reputational damage far exceeded it. The judge observed that there is nothing inherently improper about using a reliable artificial intelligence tool for assistance, but emphasized that existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.
Park v. Kim, 91 F.4th 610 (2d Cir. 2024), turned the warning into an execution. An attorney used ChatGPT to generate a reply brief containing a nonexistent citation and, when confronted, admitted the error — but the Second Circuit found her conduct part of a pattern of sustained and willful intransigence. The lesson: the court affirmed dismissal of the entire action. That is the nuclear risk in arbitration, where a tribunal empowered to manage the proceedings may strike the submission or draw adverse inferences severe enough to amount to the same thing.
United States v. Cohen (S.D.N.Y. 2023) showed that no one is immune. Michael Cohen, seeking an early end to his supervised release, submitted citations to three nonexistent cases generated by Google Bard. The lesson, in the presiding judge's words: it is not the court's job to wade through the record to determine which citations are real and which are hallucinations. In arbitration the sentiment is amplified. Arbitrators are paid service providers, and a party that forces them to traverse garbage is likely to receive an award that reflects their frustration.
§ 04 · The layer nobody checks: metadata
The risk does not stop at the visible text. When AI generates a document, it may quietly populate the hidden metadata fields embedded in the file — author, company, creation date, version history — with fictitious or misleading information. Faced with uncertain or nonexistent data for the fields its training says a “complete” document should contain, the model fabricates them. A team can review every sentence and every citation, file the brief, and still hand the court a document whose properties name an author who does not exist or a creation date that predates the engagement. Metadata is not a technical curiosity; it is evidence. Courts and opposing counsel routinely rely on it to authenticate documents and establish chain of custody, and forensic tools report what a file contains, not whether it is truthful. Hallucinated metadata can pass undetected unless it is carefully cross-checked against other evidence.
In the final analysis, the “black box” is a mirror. When it produces garbage, and that garbage is filed, it reflects the competence of the filer.
§ 05 · A verification protocol
The exposure is manageable with discipline. A workable internal protocol has three phases. (1) The sandbox restriction: general-purpose LLMs are prohibited for final legal research. They may assist with brainstorming, summarizing, or drafting non-legal text, but they never generate citations — only legal-grade, retrieval-grounded tools that link directly to primary law may do that. (2) The click-through requirement: no AI-generated citation goes into a brief until the drafting attorney has opened the primary source and verified that the case exists, that the pin-cite is accurate, and that the proposition of law is actually supported by the text of the opinion — with a source-verification log confirming every case has been checked. (3) The hallucination check: before filing, run the table of authorities through a traditional legal database, and halt the filing if any case fails to resolve until the source is manually retrieved. Given the metadata risk, inspect the document's file properties before it goes out the door as well.
The integration of AI into legal practice is inevitable and, effectively managed, desirable. Retrieval-augmented tools grounded in verified databases are genuine force multipliers. But they multiply judgment; they do not replace it. The lawyer who files a hallucinated brief is not a victim of bad software — they have chosen the speed of the draft over the integrity of the submission. The mandate for the modern practitioner is clear. Trust, but verify. Anything less is malpractice.
Adapted from “The Hallucination Tax: Why AI Negligence Is a Self-Inflicted Wound on Arbitral Efficiency,” co-authored with Hon. Leo M. Gordon (Ret.), and from Daniel's work on AI metadata hallucination with Jennifer Deutsch and Morgan Ward.