An AI Agent Beat a Barrister at the FWC.

A computing academic beat Macquarie's lawyers at the FWC with ChatGPT and Claude. What does this mean for AI-assisted litigation in Australia?
Historic neoclassical council chambers facade with columns and ornate detailing, illustrating the tribunal setting where AI-assisted self-representation was tested.

Greg Baker turned up to the Fair Work Commission on a penny-farthing. His words, not ours. A computing academic at Macquarie University, Baker represented himself against the university's legal team, which included veteran employment barrister Leigh Howard, using ChatGPT Pro and Claude. On 12 August, the FWC handed down its judgment: Baker was entitled to permanent part-time status, in the first ruling to cite the federal government's casual conversion legislation, which took effect in 2024. Macquarie is considering an appeal.

Jeannie Paterson, who heads the Centre for AI and Digital Ethics at the University of Melbourne, called the result the first publicised court victory by an AI agent. That framing captures the headline, not what Baker built or what his win reveals about the gap between a capable private assistant and work product a tribunal can trust.

What Baker Built

Baker had been teaching at Macquarie's school of computing in a casual capacity since 2023. In October 2025, his AI tool suggested he could shift to a part-time role that would permit him to apply for research grants. Macquarie denied the request. Baker took the case to the FWC, motivated by a desire to push back against the university sector's "permanent casual system" and, as he confessed, by "the lulz" of orchestrating a faceoff between AI agents and seasoned lawyers.

He uploaded his emails, conversations, and meeting records to the AI tools and trained them to validate their own output. His agents cross-checked citations and references against that record. As he described it, the system reviewed his submissions each day, digging up cases that would normally require a team of paralegals.

This is document-based, source-tracing work. Baker fed his own evidentiary record into the system and asked the AI to reason against it and verify its own claims. After Macquarie filed a late submission on the eve of the hearing, Baker turned the same agents on the university's arguments. "It tore apart the university's arguments, found all the cases, presented it to me, and I don't think I would have even spent two hours on that," he said. He described the result as a 50-to-nil football game.

The workflow was good enough to beat a veteran employment barrister acting for a major university. Baker is also honest about the technology's limits. He admits that ChatGPT Pro struggled with legal inquiries and needed significant support before it became useful around June. The US$200-per-month tool was not ready out of the box. It required a computing academic, working for months, to shape it into something that could handle the task.

The penny-farthing metaphor works in a way Baker may not have intended. He arrived at the FWC with a tool that took months of retraining before it functioned reliably, on a course short enough for that to be enough. It got him there. That is not the same as equipment anyone could pick up and ride.

The Unevenness Is the Story

Baker's experience makes a familiar pattern concrete: the technology is powerful but uneven. Its value depends on the quality of the sources it can access and the rigor with which it verifies its own output. Baker spent months training his agents because the default behavior of these tools, when pointed at legal questions, falls short.

A skilled practitioner spent months building a workflow that compensated for the tool's weaknesses, and that workflow happened to be good enough for a specific matter at a specific tribunal. It is not generalisable in the way the coverage implies. Most litigants do not have a computing academic's ability to train and debug AI agents over several months. Most will pick up a consumer chatbot, ask it a question, and take the answer at face value. That is the scenario tribunals should be preparing for, and it is a different proposition from Baker's engineered setup.

Baker argues that as this technology diffuses, the standard of legal argument will rise, and individuals will gain access to capabilities reserved for those who can afford lawyers on tap. He foresees a world, perhaps by 2031, where an AI agent as capable as his setup sits standard on every desktop. That distribution-of-power argument is fair. Whether it plays out the way Baker imagines depends on a question his case does not answer.

The Question the Case Leaves Open

Baker trained his agents to chase citations, and for the person who built the system, the result is a convincing assistant. The reporting on the case does not describe a system where every claim links back to a specific, verifiable passage of Australian primary law that a tribunal member or opposing counsel could check independently. Baker's agents validated their own output against his uploaded documents and whatever sources they found. A tribunal member reading a submission needs to be able to follow each assertion to its source. That traceable chain from claim to authority is the difference between a private research tool and work product you can stake your name on.

This distinction matters more in tribunals than anywhere else. The FWC, like most tribunals, handles a high volume of matters involving unrepresented litigants. If those litigants start using AI agents to prepare their cases, the tribunal will face a practical question: Can a member reading a submission click through to the source and verify the claim being made? Baker himself predicts that tribunal workloads will worsen as AI-mediated disputes multiply. He suggests, only half in jest, that parties may "put our bots in a room to argue with each other for 1000 years of human time."

A tribunal member who cannot verify a citation has no way to distinguish a genuine legal argument from a confident hallucination. The cost of that uncertainty falls on the tribunal, not the litigant.

Where the Gap Is

Baker can verify his own agents' output because he built them, trained them for months, and cross-checked their results against his own uploaded records. Producing auditable, source-linked work product that someone other than the builder can verify is a different order of problem.

Habeas was built to close that gap for Australian law. Our search engine scans over 300,000 Australian cases and pieces of legislation in seconds, with results grounded in a closed dataset of legitimate Australian legal sources. Every answer is traceable to its source. A practitioner or tribunal member reading a Habeas output can follow the citation back to the specific passage of primary law that supports it. Foundational research that used to take a full morning can now be completed in minutes, but the time saving is secondary to the fact that the output is verifiable.

Baker's agents chased citations autonomously. The question his case leaves open for the tribunal system is whether the results of that kind of work can be checked by someone who did not build the system. We built Habeas to answer that question, because legal AI that cannot be audited is not yet ready for the work it is being asked to do.

The Practical Upshot for Practitioners

Consumer-grade AI, in the hands of a skilled user, can do meaningful legal work: that is what Baker's win shows. Practitioners and tribunal members who face that reality need to understand what it changes.

The same technology that helped Baker win required months of training before it became useful, and the output it produced could only be verified by Baker himself. The FWC has not yet faced a flood of AI-assisted litigants, but it will. When it does, the question of whether outputs are auditable will move from an academic concern to an operational one.

Habeas is an Australian legal AI research platform. Every output is grounded in verified Australian legal sources, with citations traceable to the source document. See for yourself at habeas.ai.

Related reading

If you want to try for yourself or get in contact, book a demo with us here. We also offer the capacity for self-serve individuals to sign up, and subscribe or register a free trial at app.habeas.ai.

The legal research in this article was conducted and every citation verified using Habeas, the Australian legal AI research platform.

Hero image: Sonny Sixteen on Pexels

Other blog posts

see all

Experience the Future of Law