AI answers beat most law students in Wollongong exam study
Three Australian legal academics have repeated an experiment first run in 2023: give generative AI a real university law exam, mix its answers in with the students' papers, and have them marked. The follow-up was published in the journal Law, Technology and Humans on 23 September 2026, and the authors described it in The Conversation on 29 September. In 2023 the AI papers averaged 52.5% and sat around the 22nd percentile. This time they averaged 76.3% in criminal law and 66.0% in tort law, and seven of the 18 papers ranked in the top 10% of the class. The authors say the result calls into question marks from assessments that are not supervised.

Summary
- The study generated 18 exam answers in June 2025 using AI models from five providers, nine for a Criminal Law exam and nine for a Tort Law exam at the University of Wollongong.
- On average the AI papers outperformed 82.5% of students in Criminal Law and 61.0% in Tort Law. Seven of the 18 ranked at or above the 90th percentile.
- The authors recommend a mix of supervised AI-free assessment, assessment that openly includes AI, and a two-stage 'relay' format.
Key takeaways
- The students sat these exams in person, under supervision, with no digital devices. The AI answers were inserted by the researchers. The finding is about what AI could do on unsupervised work, not about cheating in this exam.
- Results were uneven. Two models that scored above the 90th percentile in Criminal Law fell to the 28th and 38th percentiles in Tort Law.
- The study covers two subjects at one university, with one generated answer per question, and the models were tested in mid-2025.
What the study did
The paper is Legal Minds Vs. Neural Networks: Law Schools Need to Confront Generative AI's Increasing Sophistication in Legal Reasoning by Armin Alimardani of Western Sydney University and Melissa Porter and Warwick Gullett of the University of Wollongong. It is open access and follows Alimardani's 2023 study, which tested AI on a Criminal Law exam at the same university.
- Subjects: Criminal Law and Procedure A, a first-year subject, and Law of Torts, which the paper says is usually taken in third year. Both are compulsory in the Bachelor of Laws.
- The exams: three-hour, in-person, open-book exams under invigilation. Students could bring printed materials and could not use digital devices.
- The AI papers: 18 in total, nine per subject, generated in June 2025. The paper lists Claude Opus 4, OpenAI's o3, GPT-4o and GPT-4.5 with Deep Research, Google's Gemini 2.5 Pro, xAI's Grok 3 in two configurations, and two Perplexity configurations.
- What the models were given: the exam questions and a prompt. They had internet access and were not given lecture notes, textbooks or other subject materials.
- Marking: the AI papers were mixed with student papers. Twelve of the 18 were marked by tutors who did not know AI was involved. The other six were marked by the two subject coordinators, who are authors of the study.
- Comparison group: anonymised grades from 302 Criminal Law students and 273 Tort Law students.
Six of the nine papers per subject were copied out by hand by the coordinators and three were printed, which the paper says mirrors the adjustments available to some students. Hyperlinks were removed from the AI answers before marking. The paper says errors and invented references were left in.
What it found
| Measure | 2023 study (Criminal Law) | New study: Criminal Law | New study: Tort Law |
|---|---|---|---|
| Mean AI mark | 52.5% | 76.3% | 66.0% |
| Share of students outperformed, on average | 22.1% | 82.5% | 61.0% |
| AI papers above the 90th percentile | None | 5 of 9 | 2 of 9 |
| Best AI paper | About the 86th percentile | Claude Opus 4, above 97% of students | o3, 99th percentile |
All figures in the table are from the paper. The authors caution that the two studies used different prompting methods and that the 2023 study gave some models subject-specific material, which the new one did not.
The 2023 study found AI was much weaker on problem questions, where a student has to apply the law to a set of facts, than on essays. The new paper reports that this gap has closed in Criminal Law: the AI papers averaged 76.0% on the problem question and 76.8% on the essay questions, ahead of the students by 16.4 and 13.2 percentage points. In Tort Law the picture was different. On problem questions the AI lead was less than one percentage point. On the essay question it was 43.2 points.
Where the AI answers fell down, and the study's limits
The paper does not describe AI as a reliable lawyer. Writing in The Conversation, the authors say the models have "jagged capabilities": strong analysis could sit beside poor citations, weak choice of sources or fabricated authorities.
- Inconsistency between subjects: Claude Opus 4 scored 85 and 83 across the two exams. Two Perplexity configurations scored above the 90th percentile in Criminal Law, then 48 and 53 in Tort Law, which put them at the 28th and 38th percentiles.
- Invented cases: the paper reports that Claude Opus 4, the second-best model for citing case law in Criminal Law, also had the second-highest rate of hallucinated case citations.
- Search helps, with a catch: the three Deep Research configurations produced one hallucination between them in Criminal Law. The authors note that part of this is because those models cited fewer sources.
The authors list their own limitations. Each AI paper is a single run, and output varies between runs, so a model's mark is one sample and not a precise estimate. Only two subjects at one law school were tested. Marking in law is subjective. The AI answered each question separately, without the time pressure, fatigue or handwriting a student faces, which the paper says may modestly favour the AI. And the models tested were those available in June 2025.
What the authors recommend
The paper argues against two simple answers: making everything a supervised exam, and prohibiting AI on unsupervised work. On the second, it says that where an assessment is not invigilated and AI is prohibited, enforcement will often be unreliable. It proposes combining three formats.
| Format | How it works | What it checks |
|---|---|---|
| AI-free and invigilated | Supervised exams and in-class tasks, including mid-semester work. The authors say these should form the bulk of first-year assessment. | That the student has the underlying knowledge and can reason without help |
| AI as part of the task | Students are expected to use AI. Marking looks at the process: what was prompted, which sources were checked, which errors were caught. | Whether the student can direct and correct an AI tool |
| Relay | Two stages marked separately. Either the student drafts under supervision and then improves the draft with AI, or the AI drafts and the student critiques and fixes it under supervision without AI. | Both independent knowledge and the ability to work with AI |
What this means for students and teachers
- Students: expect more supervised, device-free assessment, especially early in a degree. The study's authors recommend it, and other universities have moved the same way, as in the two-lane model at the University of Bath.
- Students: a fluent AI answer is not a checked one. The best-scoring model in Criminal Law was also among the most likely to cite cases that do not exist.
- Teachers: this is evidence that a take-home essay mark may no longer show what a student can do alone. It is one study in one discipline, so treat it as a prompt to test your own assessments, not as a universal figure.
- Teachers: if AI is prohibited on an unsupervised task, consider how that rule would be enforced. The authors' view is that such tasks should either be supervised or be redesigned to include AI openly.
Sources
- Armin Alimardani, Melissa Porter and Warwick Gullett, Legal Minds Vs. Neural Networks: Law Schools Need to Confront Generative AI's Increasing Sophistication in Legal Reasoning, Law, Technology and Humans, advance online publication, 23 September 2026. DOI 10.5204/lthj.4896.
- Armin Alimardani, Melissa Porter and Warwick Gullett, Our research shows how AI is getting better at exams. Here's how unis can respond, The Conversation, 29 September 2026.
- University of Wollongong, Our research shows how AI is getting better at exams. Here's how unis can respond, media centre republication, 2026.
- HyperAI News, Rising AI Exam Scores Prompt Universities to Overhaul Assessments, summary credited to Phys.org, accessed 4 October 2026.
- Armin Alimardani, Generative artificial intelligence vs. law students: an empirical study on criminal law exam performance, Law, Innovation and Technology, 2024 (the 2023 baseline study).
- Jerome Doraisamy, Can GenAI outperform Australian law students?, Lawyers Weekly, 24 September 2024 (report on the 2023 study).
FAQ
Can AI pass a university law exam?
In this study, yes, and comfortably. Eighteen AI-written papers averaged 76.3% in Criminal Law and 66.0% in Tort Law at the University of Wollongong. Seven of them ranked at or above the 90th percentile of students. Results varied a lot between models and between the two subjects.
Did students use AI to cheat in these exams?
No. The student exams were in person, supervised and device-free. The researchers generated the AI answers themselves and mixed them in with student papers for marking. The study measures what AI could produce, which matters for unsupervised assessments.
Which AI models were tested?
The paper lists Claude Opus 4, OpenAI's o3, GPT-4o and GPT-4.5 with Deep Research, Google's Gemini 2.5 Pro, xAI's Grok 3 in two configurations, and two Perplexity configurations. The answers were generated in June 2025, so newer models were not included.
Did the markers know which papers were written by AI?
Twelve of the 18 AI papers were marked by tutors who were not told that any papers were AI-generated. Six were marked by the subject coordinators, who are authors of the study and knew, after some tutors did not consent to their marks being used.
What do the authors say universities should do?
Use three kinds of assessment together: supervised AI-free tasks, tasks that openly include AI and mark the process, and a two-stage relay format that marks independent work and AI-assisted work separately.
What happens next
The study's own limits are clear: two subjects, one university, one run per question, and models from mid-2025. The authors call for repeated generations per model and say results may differ in other subjects such as contract and constitutional law. We will update this article if a replication in another discipline is published.
Is using AI plagiarism?
A plain guide to when AI help on coursework is allowed, when it is not, and why the answer depends on your institution's rules.


