Hallucinated Scepticism: General-Purpose AI Fabricates Critique the Same Way It Fabricates Case Citations

General-purpose AI fabricates critique as confidently as it invents case citations.
Man with glasses reading in a sophisticated library, illustrating how legal AI research requires trustworthy, authoritative sources.

A junior lawyer is finishing a research memo at 9pm, the night before a directions hearing. She has run the client's argument through a general-purpose AI and received back a numbered list of weaknesses: structured, cross-referenced, each point delivered with apparent authority. There is no time to push back five rounds and wait for something more honest to emerge. She trusts the output, because it reads like critique.

A Lawyers Weekly op-ed published this week named what she experienced. The author calls it "hallucinated scepticism." Feed the model a document and ask it to find weaknesses in the argument. It obliges: structured points, hedged qualifications, identified weaknesses. Then push back on one of the critiques. The model softens. Push again, it reverses. After three to five rounds of sustained pressure, something closer to honest engagement finally surfaces.

The author's framing is generous: here is how you work around the problem. We think the problem is more fundamental than the workaround can reach. The better standard to build toward is the shift from managing output quality through interrogation to expecting verifiable grounding by default.

The Bug Is the Same Bug

When a general-purpose model returns a confident citation to a case that doesn't exist, the failure is not that it lied. Models don't lie. The failure is that fluency and accuracy are, by design, separate things. The model learned to produce outputs that read like legal research. Whether those outputs correspond to anything in the world is a different question, and one the model has no reliable mechanism to answer.

Hallucinated scepticism is the same mechanism. The model learned to produce outputs that read like rigorous critique. Structured objections, hedged qualifications, identified weaknesses. The output has the shape of critical analysis. Whether anything supports those critiques is, again, a different question.

The sycophancy the author documents, the model's willingness to soften and reverse under pressure, is a predictable consequence of how these models are trained. Reinforcement learning from human feedback optimises for approval: humans consistently rate confident, well-structured answers more favourably than tentative ones, so the training process systematically rewards outputs that look authoritative over outputs that are authoritative. The adversarial prompting the author recommends works by temporarily reversing that incentive within the conversation, asking the model to hold a position under pressure. But the pressure has to come from somewhere, and the user has to know to apply it.

The practitioner who received a confident citation to a non-existent case at least had a clear failure signal: the case wasn't there. The practitioner who receives hallucinated critique has no equivalent check available on the first answer. The critique sounds like critique. It reads the way a thoughtful reader reads. The model has, in effect, learned what rigour looks like without learning what rigour requires.

What the Fix Tells Us

The author's three-to-five-round solution is honest about its own limits in a way worth dwelling on. He is describing a process in which the model's first answer is, by implication, unreliable, and where sustained adversarial pressure is the mechanism for extracting something better. That is a candid description of a workflow that requires the user to already know the first answer may be wrong, and to persist through multiple rounds before trusting the output.

Realistically, who will do this?

A junior lawyer finishing a research memo at 9pm before a directions hearing. A general counsel working through a contract question between two back-to-back meetings. A barrister handed a brief on a tight timeline who needs to identify the key issues before the conference. None of them will run five rounds of adversarial prompting on every output they receive. This is a description of how legal work operates under time pressure, not a comment on the people doing it.

There is a further directional risk that the op-ed doesn't name directly. Hallucinated scepticism doesn't only generate false weaknesses in a position; it may equally fail to surface real ones while producing plausible-sounding alternative concerns. A practitioner who receives five confident-looking critiques and works through them may feel the analysis is complete, when the actual vulnerability in the client's argument was never raised. The model's output creates a false floor: the appearance of thorough critical engagement where the actual gap remains open. That risk is harder to see than a fabricated citation, and harder to catch under time pressure.

Prompting discipline is not a substitute for a tool that shows its working on the first answer. If three to five rounds are required to surface honest engagement, every first answer from the model is suspect, not occasionally, not in edge cases, but on every query, until proven otherwise.

What the Alternative Looks Like

There is a different architecture. The author arrives at his fix by reasoning within a general-purpose AI framework, where output quality is a function of how hard you press the model. That framework is the wrong starting point for legal research.

The structural answer to hallucinated output, whether citations or critique, is grounding. Tying what the model returns to a source it can point to, rather than a source it plausibly inferred from training data, is what separates a verifiable answer from a fluent one. When an output is traceable to a primary document that exists and says what the output claims it says, the user doesn't need to run five rounds of pushback to establish whether the answer is real. The citation resolves, or it doesn't.

A traceable citation also does something the prompting ritual cannot: it opens a door. When a research output points to an actual judgment, the practitioner can read the passage in context, see whether the authority was applied at trial level or on appeal, check whether it was followed or distinguished in subsequent decisions, and form a view about whether the proposition it supports is as strong as the summary suggests. Grounding doesn't prove the source exists and stop there; it makes genuine legal judgment possible. An output with nothing underneath it closes that door at the moment it arrives.

This is why the distinction between general-purpose AI and a purpose-built legal research platform matters more than the surface feature comparison suggests. Both can produce fluent, authoritative-sounding answers. Whether the answer is tethered to something verifiable at the moment it is returned, and whether the practitioner can see exactly where each output came from without first interrogating the model's confidence, is a different question entirely.

One general counsel put it plainly: the Australian-law focus and the depth of nuance in the answers is a major differentiator, one that materially changes confidence and speed when forming legal views. That is a different relationship between the output and the underlying law, and it doesn't depend on a prompting ritual before the practitioner can decide whether to rely on what came back.

The Underlying Problem the Op-Ed Names

The Lawyers Weekly piece is valuable precisely because a practitioner wrote it from experience, and landed on a limitation that the AI industry's own communications consistently avoid acknowledging. General-purpose AI is very good at producing outputs that look like rigorous work. Producing rigorous work is a different capability, and the gap between them creates a specific kind of risk in legal practice: confident, well-formatted advice with nothing underneath it.

The author's three-to-five-round mitigation is, to his credit, a real attempt to close that gap. We don't think it closes it in a way that holds under the conditions where legal AI is used.

Practitioners who find value in AI-assisted research tend to be those who have stopped trying to prompt their way to rigour and started demanding that the tool show its working from the first answer. Managing output quality through interrogation is the harder path; expecting verifiable grounding by default is the better standard. The Lawyers Weekly author named a real problem. The answer is a tool that shows its working on the first answer.

Habeas returns citations that resolve to their source document, across a closed corpus of over 300,000 Australian cases and pieces of legislation, without a prompting ritual in between. habeas.ai

Related reading

If you want to try for yourself or get in contact, book a demo with us here. We also offer the capacity for self-serve individuals to sign up, and subscribe or register a free trial at app.habeas.ai.

The legal research in this article was conducted and every citation verified using Habeas, the Australian legal AI research platform.

Hero image: Mazen Tumi on Pexels

Other blog posts

see all

Experience the Future of Law