Sunday, August 30

Josh Morrow, a partner at Lehotsky Cohn LLP and a graduate of Harvard Law School, has published findings suggesting that AI authorship in federal appellate opinions is no longer a theoretical concern: his analysis of roughly 2,250 published circuit court opinions from January through early August 2026 identified more than 50 showing signs of AI-generated prose.

Morrow used Pangram, an AI-detection tool, to run the full set of this year’s published opinions from the regional courts of appeals. The results ranged from under 1% to more than 50% AI-written text within individual opinions, with most clustering near the lower end of that range.

What the Data Shows About AI Authorship in Federal Appellate Opinions

To establish a baseline, Morrow tested all published circuit opinions from January 2022, a period predating widespread large language model use. Pangram returned a result of 0.000000% AI-generated text for every single one of those opinions. The contrast with 2026 was stark.

Pangram itself claims a false positive rate of just 0.0041%, or roughly one false positive for every 24,000 documents. That figure has external support. An independent evaluation by the University of Chicago Booth School of Business verified a false positive rate of 1 in 10,000 overall, dropping to 1 in 25,000 on an academic essay dataset comparable to Turnitin’s own evaluation set.

In August 2025, Booth researchers Brian Jabarian and Alex Emi published a paper titled ‘Artificial Writing and Automated Detection,’ concluding that Pangram is ‘the only detector that meets a stringent policy cap (False Positive Rates ≤ 0.005) without compromising the ability to accurately detect AI text.’ That standard is the same threshold applied under institutional AI-integrity policies.

The tool’s latest iteration reinforces the case for its reliability. Pangram 4 correctly identifies 99.66% of AI-generated documents and achieves a false negative rate of 0.3396%, representing nearly six times fewer false negatives and fourteen times fewer false positives than its predecessor. Domain-specific performance data shows false positive rates of 0.01% for creative writing, 0.02% for academic writing, and 0.01% for biomedical writing, with the tool performing less reliably on niche genres such as poetry.

Morrow’s methodology involved uploading the full text of each opinion to Pangram, then re-running those with substantial AI signals from a separate account to guard against noise. The duplicate results were consistent, which he treats as evidence that Pangram’s classifications were stable rather than random.

A Profession Already Moving Towards AI, With or Without Guidance

Morrow’s findings land in a context where AI adoption in the federal judiciary is already well under way. A Northwestern University survey of 502 randomly selected federal judges, conducted in December 2025 and published by the Sedona Conference in March 2026, found that over 60% of responding judges use at least one AI tool in their chambers, primarily for legal research and document review. Nearly half reported receiving no AI training from their court.

Separately, over 300 federal judges across every circuit have issued standing orders or local rules addressing AI use in filings, with approaches generally falling into disclosure requirements or cautious-use guidelines, according to a survey of current court AI rules. Those rules, however, govern litigants’ submissions. They say nothing about what judges and law clerks may use when drafting the opinions themselves.

Morrow declines to identify any of the judges whose opinions Pangram flagged. His position is that AI can improve legal writing and reasoning, and he does not suggest that the opinions are deficient as a matter of law or analysis. What he does raise is a more structural concern: that writing an opinion is itself part of the judicial function, not merely a record of a conclusion already reached. If extended passages arrive essentially finished from a language model and remain largely unedited, a question arises about whether the discipline that writing imposes on legal reasoning has been displaced.

He is also careful about the limits of his evidence. Pangram analyses only final text; AI used for research, outlining, or reviewing a draft leaves little trace in the published opinion. Human editing can erode the signals Pangram relies on. The tool cannot identify how AI-generated language entered the drafting process, nor reconstruct whatever judicial review followed.

The normative questions he surfaces are ones the profession has not yet resolved: when, if ever, should AI involvement in judicial drafting be disclosed, and to what degree does apparent AI authorship affect an opinion’s persuasive or precedential weight? With at least two federal district judges already having acknowledged AI-generated text entering their opinions, and litigants raising the issue in state proceedings, those questions are no longer abstract. The more pressing is whether the appellate courts will develop disclosure frameworks before detection tools make the conversation unavoidable.

Share.
Law News | AI Authorship in Federal Appellate Opinions Detected Across Dozens of 2026 Rulings

Catherine Sadler practised law for fourteen years before she started writing about it. She trained at a City firm, qualified into commercial litigation, and spent the bulk of her career at a mid-sized practice handling regulatory disputes, professional negligence, and the kind of cases that are dull to describe and expensive to lose. She writes about court judgments, regulatory enforcement, legal reform, and the cases that set precedent without making the evening news. She can read a judgment and explain what it actually means for the people who were not in the courtroom. Catherine lives in Oxfordshire. She reads the Law Gazette out of habit and considers the phrase 'access to justice' to be doing a lot of unsupported work.

Comments are closed.