The Wrong Question

The Claude watermark tells you a model touched the text. It doesn't tell you whether the text is right. Verification is the lawyer's job either way.
The Wrong question

On watermarks, the model debate, and what lawyers are trusting

A lawyer should ask of any AI tool not whether the text it produces can be detected as machine-authored. The useful question is what the system does when it is wrong, and what it costs the practitioner to find out. Everything else, including the present anxiety about statistical watermarks, is downstream of that.

On 2 August 2026, Anthropic confirmed that Claude models launched from that date embed an imperceptible statistical watermark in generated text, with C2PA provenance metadata attached to generated files. The measure was rolled out worldwide to comply with the EU AI Act's Article 50 and the Code of Practice on Transparency of AI-Generated Content. Coverage in the legal press has asked what this means for practitioners who use AI to draft. The honest answer is: very little, and not in the way the framing suggests.

The reason is not that watermarks are technically trivial. The practitioner's obligation, the obligation the watermark is imagined to serve, already exists, independent of any detection mechanism. On 15 April 2026, the Federal Court of Australia issued its General Practice Note on AI (GPN-AI), which requires practitioners to confirm before filing that cited authorities exist and support the propositions stated, and mandates disclosure where a tool summarised material a witness relied on. That obligation does not turn on whether the underlying text carries a detectable mark. The practitioner's duty is verification.

The framing that has dominated the legal AI conversation — which model is smartest, which model leaves a mark, which vendor benchmarked best — is in this light a category error. It treats the inference layer as the whole system, and detection as a substitute for verification.

What the Watermark Cannot Tell You

Anthropic has been direct about the limits of its own measure. The watermark cannot tell you how much of a given document is AI-authored. It cannot tell you whether the text was heavily edited after generation. It cannot tell you whether the document is any good. It can tell you, with some probability and within some confidence interval, that a Claude model touched the text at some point in its production. That is a narrow claim, narrower than the legal-industry discussion has tended to assume.

This narrowness matters because the practitioner's duty under GPN-AI is a duty of verification. The Federal Court does not ask whether counsel's submissions were drafted with AI assistance. It asks whether the authorities cited exist and whether they support the propositions for which they are cited. A watermark answers neither question. A document can carry no detectable mark and still cite a hallucinated authority; a document can carry a mark and still be accurate. The mark is orthogonal to the risk the Court's practice note is concerned with.

A document can carry no detectable mark and still cite a hallucinated authority; a document can carry a mark and still be accurate.

This is the first sense in which the watermark debate is misdirected attention. It treats the practitioner's problem as a problem of provenance, whether the text was written by a machine. The practitioner's problem has always been a problem of truth: does this text assert what the underlying sources support? Those are different questions, and answering the first does not help answer the second. A detection layer is a story about platform accountability, useful to regulators and platform operators. Practitioner risk is a different story altogether.

For a lawyer evaluating AI tools, the implication is straightforward. A watermark does not reduce the verification burden. It does not make the work more reliable. It does not protect the practitioner who files an unverified citation. At best, it provides an ex post signal that may be relevant to disciplinary or evidentiary enquiries after something has already gone wrong. By the time that signal matters, the practitioner's failure has already occurred. The watermark arrives too late to be a meaningful part of the daily decision about whether to file.

Generate-then-Cite vs. Retrieval-before-Generation

There is a second, deeper reason the watermark framing misses what matters. Even if detection were perfect, even if every AI-authored passage could be identified with certainty, the mark would tell you nothing about the architecture of the system that produced it. The architecture is where the practitioner's actual exposure sits.

Consider two designs.

Design one: generate-then-cite

The system produces fluent prose and then attaches citations to support claims the model has already made. The lawyer receiving the output must establish, for each citation, whether the source exists, whether it says what the model claims, and whether it supports the proposition. The fluent prose gives no signal about which claims are grounded and which are not. A hallucinated authority and a real one look identical in the text. This is the failure mode the Federal Court's practice note is designed to surface.

Design two: retrieval-before-generation

The system locates the actual document before generating anything. It reasons only over what it has retrieved. It cites to the subparagraph, not to a proposition it has already committed to. The output is synthesis constrained to real, retrieved primary sources, rather than free-floating generation with citations bolted on afterward.

Retrieval-before-generation lowers the cost of catching an error because the output is anchored to specific, locatable passages the practitioner can check. The duty to verify remains; only the burden changes.

This is the failure mode a watermark cannot reach. A mark tells you that AI touched the text. It does not tell you whether the text was produced under the discipline of retrieval or as free generation with citations attached. That distinction is where the practitioner's actual risk lives. A lawyer who treats the watermark as the relevant signal will examine the wrong layer of the document for the wrong kind of error.

Model Quality Is Not the Measure

There is a third dimension to the misdirection, and it is the one most worth naming. Much of the benchmark culture in legal AI assumes the underlying model is the only thing of significance. The framing is implicit but pervasive: compare legal AI tools by which model powers them, upgrade the model, and the product improves. The inference layer is treated as the product.

This is a category error. The model is one layer of a stack that also includes document search, document processing, retrieval quality, agent architecture, citation discipline, and the corpus over which the system reasons. A stronger model in a worse stack can produce worse outcomes than a weaker model in a better stack, because the model can only reason over what the stack gives it.

A stronger model in a worse stack can produce worse outcomes than a weaker model in a better stack, because the model can only reason over what the stack gives it.

If a system relies on the inference layer alone to be right, it masks flaws in the rest of its architecture. If retrieval returns the wrong passage, a smarter model will produce more fluent, confident reasoning over the wrong passage. If the agent architecture commits to a proposition before retrieving the supporting authority, no amount of model quality will save the practitioner from the resulting hallucination.

The practical implication is that the model-swapping benchmark measures something, but it does not measure what a lawyer is trusting when they file. The lawyer is trusting the stack:

  • that document processing was faithful,
  • that retrieval surfaced the right authorities,
  • that the agent did not commit to a proposition before having the sources to support it,
  • and that the synthesis across those sources was disciplined.

None of these properties are captured by a model benchmark. Several of them are not properties of the model at all.

This is why the watermark story is misdirected in a second sense. Even if detection were perfect, it would tell you the system was used, not whether it was reliable. The right comparator between legal AI systems is what the system does when the model is wrong, and how expensive it is for the practitioner to find out. That question is not answered by benchmarks. It is answered by the architecture of the stack, the discipline of retrieval, and the verification burden the design imposes on the lawyer who signs the work.

What You Are Trusting

The watermark is a story about platform accountability, and a useful one within those limits. Practitioner risk sits elsewhere. The model-quality debate is a story about one layer of a stack, and a misleading one when treated as the whole.

The question a lawyer should ask of any AI tool, whether it is a general-purpose model deployed in a firm, a benchmark-leading agent, or a retrieval-first system, is the same: what does this system do when it is wrong, and what does it cost me to find out?

That question does not yield to detection. Architecture answers it. A system built on the assumption that the lawyer verifies everything regardless, that locates the primary source before generating, that cites to the subparagraph, and that treats the verification duty as the premise, is built around the right question.

Habeas was built on this premise. The retrieval-first architecture is designed to show the practitioner where the system fails, and what they have to work with when it does.

The watermark will tell you that a model touched the text. The architecture will tell you what the text is worth.

Lawyers should worry about the architecture.

What Habeas Does With This

Habeas is built on retrieval-before-generation, not generate-then-cite. The system retrieves from a closed corpus of Australian case law and legislation before it writes a word, and every citation in the output traces to a subparagraph in a real, retrieved source. You can open the passage the citation points to and check it in the time it takes to read a paragraph, not the time it takes to run a full research pass.

This is not a claim about the model. Swap the model underneath and the retrieval discipline stays the same, because the discipline sits in the architecture, not the inference layer. A lawyer working under GPN-AI or NSW SC Gen 23 needs a system that lowers the cost of verification. Habeas was built to answer that need directly: locate the source, constrain the output to what the source says, and put the citation where you can check it.

The watermark tells you a model touched the page. Habeas tells you where the page came from.

Other blog posts

see all

Experience the Future of Law