The Kill Switch and the Citation

Albanese wants a US-China AI treaty with kill switches, but who verifies what the system outputs?
Grand parliament hall with a domed ceiling and intricate architecture, reflecting the governmental arena where AI oversight and kill-switch policy decisions are debated.

Anthony Albanese this month called for a US-China AI agreement modelled on the nuclear non-proliferation treaty, with "kill switches" for frontier AI systems. Australia, alongside more than twenty other countries, co-signed a statement ahead of the UN General Assembly declaring that AI "must remain under human direction, oversight and control."

That is the register you reach for at the UN. Non-proliferation treaties, kill switches, human direction: the vocabulary of existential risk, a technology that might one day exceed the capacity of its creators to contain it. Frontier models are being deployed at a pace that no single government has fully grappled with, and the idea that the United States and China might agree on baseline safeguards is worth pursuing.

But that register describes almost none of what has gone wrong with AI in the legal system. A kill switch is a safeguard against a system that has run out of control. The failures that have reached courts and tribunals involve systems doing what they were built to do, in a context where what they were built to do is dangerous.

From non-proliferation to citation verification

A practitioner who relies on the output without checking it has made a verification error. A person's liberty can depend on whether the case law cited to a parole board exists. So can a solicitor's professional standing, and the integrity of evidence before a court. The person relying on the output had no easy way to tell the difference between a real citation and a plausible one. Plausibility is enough to fool someone who is busy, under-resourced, or trusting, and the legal system has no shortage of people in that position.

The most widely reported example remains Mata v Avianca, a 2023 personal injury case in the US District Court for the Southern District of New York. A lawyer representing the plaintiff used ChatGPT to research whether the Montreal Convention governed the claim. The model returned case names, citations, and summaries that read convincingly. None of the cases existed. The lawyer filed them in a submission, asked the model to confirm they were real, and received further confident assurances that they were. The court sanctioned the lawyers involved, and the episode became the reference point for every subsequent discussion of AI-generated legal research.

The mechanism behind that failure is the design itself. A large language model predicts the next token in a sequence based on patterns in its training data. Case names, citation formats, and judicial reasoning all follow predictable linguistic patterns. A model that has ingested enough legal text can produce a citation that conforms to the correct format, names the correct jurisdiction, and summarises a holding in plausible terms, without ever consulting a case database or knowing whether the case exists.

The difficulty for the profession is that the failure mode looks like competence. A fabricated citation does not announce itself. It arrives in the correct format, with a judge's name, a paragraph number, and a summary that engages with the legal issue at hand. The practitioner who reads it sees what looks like a properly researched answer. The temptation, especially under time pressure, is to treat the fluency of the output as evidence that the underlying research was done. The model was generating text, not consulting authorities. A practitioner who understands this distinction can catch the problem, but catching it requires a verification step that the tool itself provides no incentive to perform.

What verification looks like

The safeguard that should catch a fabricated citation is a step where someone opens the reported cases, checks the citation, and confirms the authority stands for what it is being asked to support. If that step doesn't happen, the failure is cheap to cause and devastating when it lands. That safeguard requires a tool that grounds its answers in primary sources and shows you where each citation comes from, and a professional culture that treats verification as a default rather than an afterthought.

The legal profession has always had verification norms. Junior lawyers are trained to check every citation in a senior's draft, to pull the reported series, to read the headnotes, to confirm the proposition is supported by the actual judgment rather than the summary alone. What changed with generative AI is that the output arrives looking finished. A draft from a senior partner at least carries the authority of someone who read the cases. A draft from a chatbot has the surface of that authority without any of the underlying work. The verification instinct the profession developed over decades was calibrated to catch errors made by humans who had done the research but missed something. It was not calibrated to catch outputs from a system that did no research at all.

The Federal Court's Generative AI Practice Note already gestures in this direction. It requires practitioners to personally confirm that cited authorities exist and support the propositions they are cited for. The obligations extend to evidence referenced in submissions, facts stated in pleadings, and the accuracy of chronologies. For expert reports and affidavits, the requirements are stricter still: AI can assist with structure and drafting, but it cannot substitute for the deponent's own recollection or the expert's own process of reasoning. The Court assumed that lawyers using AI could tell it what tool they used, how, and for what purpose. That is human oversight and control, written for a courtroom rather than a UN assembly.

The practice note also addresses confidentiality, and here the distinction between tools becomes sharper. Entering privileged or confidential information into an open AI system may breach the implied undertaking that governs documents produced in litigation. The Court draws a line between open tools, where information may become accessible to third parties, and closed or ringfenced systems where the risk is lower. A practitioner using a general-purpose chatbot to research a matter has no way to know where their input goes or who might see it next. A practitioner using a system built on a closed dataset of Australian legal sources is working within a controlled environment.

A practice note only works if the tools practitioners use are built to make verification possible. A general-purpose AI chatbot cannot tell you where its citations come from because, in a meaningful sense, they don't come from anywhere. The model predicts the next token in a sequence. Sometimes that sequence is a real case name. Sometimes it isn't. The system has no internal mechanism for distinguishing the two, because it was designed for fluency, not citation accuracy, and fluency is what makes the output dangerous in a legal context. It sounds right, so it gets trusted.

A tool that generates text about the law and a tool that searches the law belong to different categories of software. The first produces language that resembles legal research, useful for drafting and summarising. The second retrieves primary sources, ranks them by relevance, and presents them with enough provenance that a practitioner can exercise judgment over what they mean. That second category is what you need before you cite anything in a submission or advise a client on the strength of their position.

We built Habeas around that distinction. Our Search Engine scans over 300,000 Australian cases and pieces of legislation in seconds, with results grounded in a closed dataset of legitimate Australian legal sources, so they are verifiable and traceable, never hallucinated. Foundational research processes that used to take a full morning can now be completed in minutes. Every citation resolves to a primary source a practitioner can open, read, and confirm before it reaches a court, a tribunal, or a client. A General Counsel at a Series A startup described the experience as having "a law firm in your pocket," a first-line legal intelligence tool that earns its place in the workflow because the answers are grounded.

This is the unglamorous version of human oversight, far from the vocabulary of kill switches and non-proliferation treaties. A practitioner runs a search, the system returns authorities from a closed Australian corpus, and the practitioner confirms the citation before filing. It would have prevented fabricated case law from reaching a parole board, a solicitor from filing submissions with invented authorities, and whatever else is happening in matters we haven't heard about yet.

If the principle is human oversight, the question is what tool makes oversight possible. Book a demo at habeas.ai.

Related reading

If you want to try for yourself or get in contact, book a demo with us here. We also offer the capacity for self-serve individuals to sign up, and subscribe or register a free trial at app.habeas.ai.

The legal research in this article was conducted and every citation verified using Habeas, the Australian legal AI research platform.

Hero image: Czapp Árpád on Pexels

Other blog posts

see all

Experience the Future of Law