
All rights reserved, Habeas 2026
See our Privacy Policy
See our Privacy Policy

At some point during a Federal Circuit and Family Court proceeding in July 2026, a judge decided to conduct his own experiment. A self-represented litigant had filed submissions that purported to summarise one of his own rulings. The judge opened ChatGPT and asked it about the case directly. The chatbot admitted, on the record, that it had fabricated those details.
The exchange was reported by Lawyerly on 23 July 2026. The predictable commentary followed: another AI hallucination, another warning about the dangers of generative tools in legal proceedings. Cautionary tales, calls for vigilance, a few more column inches about the risks of technology moving faster than regulation.
All of that is correct, and none of it is the story.
We have known for years that large language models confabulate. The technical literature is unambiguous. The product warnings are prominent, if not always read. Researchers, practitioners, and AI developers alike have said, at length and in public, that these tools produce fluent nonsense with the same confident tone they use for accurate outputs. What happened in that courtroom was a demonstration of a known phenomenon, not a discovery of a new one.
The story is that the only person who checked was the judge.
Australian courts have moved with real seriousness on AI guidance over the past two years. The Federal Court's GPN-AI, the NSW and Victorian civil procedure notes, the Fair Work Commission's position: these instruments share a coherent framework. Disclose AI use. Verify citations before filing. Accept professional responsibility for every document that carries your signature. A practitioner who files a hallucinated authority is now in breach of a specific, named obligation, not a general duty of candour.
That framework is sensible, and it will do real work for practitioners who are subject to it.
Self-represented litigants are not subject to it. They do not have professional obligations that translate disclosure requirements into routine practice. They do not have a supervising principal, a colleague to sense-check a draft, or a firm policy on AI use. They have whatever they can find online, and increasingly that means a general-purpose chatbot that produces authoritative-sounding answers about a jurisdiction it was not built for, citing cases that do not exist, in the register of careful legal reasoning.
Practice notes are addressed to lawyers because lawyers are the ones courts can regulate. The population of people who turn up to federal family law proceedings without representation is large, often under serious stress, and making decisions with real consequences for themselves and their families. Telling that population to verify their AI outputs is advice addressed to the wrong layer of the problem.
There is also a structural point that practice notes do not address: the adversarial system usually does some of this work. When both parties are legally represented, an opposing solicitor who receives a submission citing an authority that cannot be located is likely to raise it. The error may not reach the bench at all. In proceedings where one party is unrepresented, that check collapses. The represented party's lawyer may not have read the submission closely enough to interrogate a source that looked facially plausible. The fabricated authority travels, unchallenged, toward the record.
What should concern us about the July incident is the nature of the circuit-breaker. A judge happened to be curious enough to interrogate the tool himself, in the moment, on his own initiative. That is luck, not design.
There is a particular irony in the specific failure here that is worth naming. The chatbot was not hallucinating an obscure authority in a jurisdiction the judge was unfamiliar with. It was hallucinating the judge's own ruling, back to him. A fabrication that audacious was only discoverable because the judge knew the answer before he asked the question. A submission citing an invented authority from a different court, in a matter the judge had not personally decided, would have offered no such shortcut.
A different judge, a less conspicuous error, a slightly more polished submission, and the fabricated summary travels unchallenged into the record. The self-represented party may never know their source was invented. The other party may not have the means to expose it. And the hallucination, having passed unchallenged, becomes, functionally, a submitted fact.
The access-to-justice dimension here is not incidental. When people cannot afford legal representation, they increasingly turn to the same consumer AI tools that are demonstrably unreliable for legal research. The gap between what those tools produce and what a practitioner would catch is, in the ordinary case, invisible; there is no one positioned to close it.
Family law adds a further layer. These are proceedings where the stakes are not abstract: where the court is deciding where children live, how assets are divided, what contact arrangements look like for years. People in those proceedings are often in acute personal crisis, processing the end of a relationship alongside navigating an unfamiliar court system. Methodical citation verification is not high on the list of things they have bandwidth for. They need answers that work. Consumer AI gives them answers that sound like they work. The gap between those two things is where the July incident lives.
We want to be careful about what we claim here, because anti-hype is a design commitment for us, not a marketing position.
Habeas was built specifically for Australian legal practitioners, grounded in Australian primary law, with citations that trace to their source documents. The Search Engine scans over 300,000 Australian cases and pieces of legislation in seconds, from a closed dataset of legitimate Australian legal sources, so results are verifiable and traceable rather than generated from pattern-matching across the open internet. When a barrister or a GC uses the platform to locate an authority, they can follow the citation back. If it is wrong, they will know before it reaches the court.
That design does not solve the problem the July incident exposes. Habeas is built for practitioners, and we are not claiming otherwise. But the incident does clarify why citation-first design matters, and why the distinction between a tool built around verified Australian sources and a general-purpose chatbot asked legal questions is a substantive one.
The chatbot in that courtroom was doing what it does: predicting fluent text. It generated details that sounded plausible because plausibility is what its architecture produces. Asked to summarise a judgment, it produced a summary in the register of a judgment summary. Whether the underlying case existed was not a question the tool was built to answer.
A tool that grounds every output in a closed corpus of real Australian legal sources, and surfaces the source document alongside the answer, operates on a different basis entirely. The confidence interval is different. The failure mode is different. The practitioner who checks the citation can find the document; the one who relies on a consumer chatbot cannot. That difference is architectural: no amount of prompting or instruction changes what a model optimised for fluency does when it encounters a gap in its knowledge.
Courts will keep refining their guidance. The Federal Court symposium planned for later this year will likely sharpen the verification obligations further as the judiciary gathers more data on how AI is being used in proceedings. That is the right institutional response.
What no practice note will fix is a self-represented litigant relying on a tool that will, under a small amount of interrogation, admit it invented the case law it cited. The fix for that is harder, and it involves the whole ecosystem of who builds tools for legal use, what claims those tools make about their accuracy, and whether the architecture is designed around verification or around fluency.
We think the July incident will be cited for years as an early, on-the-record illustration of the difference. The question worth sitting with is how many times the same admission would have been true before a judge happened to ask.
If you want to see citation-first research at work, book a demo at habeas.ai.
If you want to try for yourself or get in contact, book a demo with us here. We also offer the capacity for self-serve individuals to sign up, and subscribe or register a free trial at app.habeas.ai.
The legal research in this article was conducted and every citation verified using Habeas, the Australian legal AI research platform.
Hero image: KATRIN BOLOVTSOVA on Pexels
