The CPS Hallucination and the Limits of "Verify Everything"

The CPS filed AI-hallucinated case citations in a UK High Court extradition appeal.
Female judge reviewing legal documents at the bench with a gavel, illustrating the courtroom impact of AI-hallucinated case citations.

Somewhere in the preparation for a UK extradition appeal, an AI research tool generated case citations that sounded exactly right. Proper case names, the structure you would expect, the kind of reference that belongs in a High Court submission. The Crown Prosecution Service filed them. Then a barrister checked. Then another, independently, did the same. The cases didn't exist.

In July 2026, the CPS admitted that AI-hallucinated citations had made their way into High Court extradition proceedings. The admission drew the usual commentary: AI is risky, lawyers must take care, supervision is essential. All of that is true. But if the lesson stops there, it is also incomplete. The conditions that produced the CPS incident (time pressure, plausible output, no source document to check against) are present in every practice that has adopted AI research tools without asking what traceability actually requires.

The detail that deserves more attention is how the problem was caught. A human went and checked. Then another human, independently, went and checked too. That sequence worked, and it worked because both barristers were willing to do the independent research the AI tool was supposed to reduce. It should make practitioners ask a prior question: what does it mean to verify an AI citation in the first place?

The standard advice following any hallucinated-citation incident is predictable. Verify all AI output before relying on it. That advice is correct. It is also built on an assumption that breaks down quietly.

Verification assumes you have something to verify against. A real case citation points to a real judgment. You pull it up, read the relevant passage, confirm it supports the proposition. The process has a target. The whole exercise takes a few minutes and either holds or doesn't. A hallucinated case name gives you none of that. It has the right structure. It looks like a citation. There is no judgment behind it. Discovering this requires independent research, which is what the barristers in the CPS matter did. Calling that a verification workflow overstates what it is: doing the research twice, once by the AI incorrectly, and once by a human correcting the AI's invention.

There is a structural tension buried here that guidance rarely names. AI research tools are adopted precisely when practitioners are under time pressure. The "verify everything" instruction requires time. For an invented citation, it requires the same time the AI was supposed to save. A practitioner working to a filing deadline who receives a plausible case name is not well-positioned to run an independent authority search. The failure mode is most likely to occur exactly when the tool is most likely to be used.

Australian courts have moved on disclosure. The Federal Court's GPN-AI, which came into force in April 2026, requires practitioners to verify that cited legal authorities exist and support the stated proposition before signing any document filed with the court. The obligation is real, binding, and appropriate. It also runs downstream of the research process in a way that matters.

Consider what the GPN-AI's attestation actually requires a practitioner to certify. Signing a document that relies on an AI-generated citation is a representation to the court that the cited authority exists and supports the stated proposition. A practitioner who cannot confirm that by pulling up the source document cannot make that certification in good conscience. The practice note doesn't create a new standard; it makes explicit what professional responsibility already required. But a tool whose citations don't trace to a real document places the practitioner in a position where completing the attestation demands parallel research they may not have time to do. The CPS barristers were not in an unusual situation. They were in the situation the guidance anticipates, doing what the guidance requires, spending hours they had not budgeted on confirming that an AI's output corresponded to reality. The practice note defines the obligation. It doesn't supply the infrastructure to meet it.

The practice notes also require practitioners to account for their AI use if the court asks: what tool, how used, for what purpose. "I used it and then independently verified each authority" is an answer. It describes a workflow in which the AI contributed something less than nothing: a set of plausible fictions that required additional labour to expose.

This is the question the CPS incident actually raises, and the one most commentary sidesteps. The problem is not that practitioners failed to verify. The problem is what verification requires when the citations aren't retrieved from real sources but generated from a model's sense of what a case name looks like. Treating those two situations as equivalent (a draft that needs checking versus an invention that needs exposing) is the category error most AI guidance quietly commits.

When Habeas returns a citation, a real document sits at the end of it. The search engine draws from over 300,000 Australian cases and pieces of legislation in a closed corpus of verified primary law. Pull up the case, read the passage, confirm it supports the proposition. That sequence works because there is actually something at the end of it. For practitioners who have spent a morning running parallel verification on an AI's inventions (doing the research twice to catch what should never have been returned) the felt difference is not subtle. Foundational research that used to take a full morning can now be completed in minutes, and the citations that come back are traceable to source, not generated to plausibility.

There is a further consequence of this architecture that matters practically. When Habeas cannot find an authority for a proposition, it returns no result. That is a different failure mode entirely. A practitioner who gets no results knows they need to look harder or reconsider the proposition. A practitioner who gets a confident, well-formed citation has no reason to doubt it. Hallucination fails silently; an honest absence of results fails loudly, and that is the preferable outcome. The GPN-AI's verification obligation is one the practitioner can actually complete when there is a real document to complete it against.

We are not claiming to have eliminated all risk of error. Habeas is a first-line intelligence layer, not a substitute for the lawyer's own analysis, and professional judgment over any research output remains with the practitioner. What we do claim is that a citation Habeas returns points to something real in the corpus.

The CPS admission will recede from legal news. Incidents like it are becoming regular enough that they no longer carry the shock they did two years ago. But the pattern they expose is consistent: an AI tool produces a plausible output, the practitioner relies on it, the plausibility isn't backed by anything real, and downstream someone finds out.

For Australian practitioners, the GPN-AI and practice notes like it set the governance floor. Disclose your AI use. Verify what you file. Accept professional responsibility for everything that reaches the court. These are the right baseline obligations. What sits above that floor is a design question about the tools that feed into the workflow. A tool built on a general-purpose model, where citation is a prediction rather than a retrieval, converts verification into independent research. A tool built from the start on a closed corpus of verifiable Australian legal sources changes that task into something a practitioner can do in the time they have: check the citation against the document it came from.

The CPS barristers caught the error because they checked. The checking worked because it eventually resolved against reality. The verification habit matters. So does the architecture that makes verification possible, and those are not the same thing.

Related reading

If you want to try for yourself or get in contact, book a demo with us here. We also offer the capacity for self-serve individuals to sign up, and subscribe or register a free trial at app.habeas.ai.

The legal research in this article was conducted and every citation verified using Habeas, the Australian legal AI research platform.

Hero image: khezez | خزاز on Pexels

Other blog posts

see all

Experience the Future of Law