AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Pakistani Judges Give Their Verdict on JudgeGPT

Collected Oct 1, 2026

A first-of-its-kind trial in Pakistan found that a custom AI tool for judges, together with appropriate training, increased the number of cases resolved by 6.3 percent with no obvious drop in the quality of judgments. The study focused on Pakistan's trial courts, which faced a backlog of 2.26 million cases and fewer than two judges per 100,000 people, compared with 22 in the EU and eight in Brazil.

Economist Sultan Mehmood of the New Economic School in Moscow and collaborators built JudgeGPT, combining OpenAI's GPT-4 large language model with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes, to assist with legal research and drafting judgments. They began offering the tool in 2024 to 1,559 trial judges, roughly half the country's justices. The tool used retrieval-augmented generation to query a database of 128,292 Pakistani judicial opinions and 943 statutes, with responses including footnotes linking to cases and laws.

The team put 1,197 judges through six 90-minute Zoom training sessions developed in collaboration with Pakistan's Federal Judicial Academy, covering how LLMs work, their limitations, risks of bias and hallucinations, and the importance of verifying outputs. Another 180 judges received only general technology training, and a final group got no training. By the time 487 judges had completed the program, the median district saw a 6.3 percent jump in resolved cases, and more trained judges in a district produced a bigger effect. Appeal rates also fell slightly.

Mehmood said the researchers found an increase in cases resolved with no corresponding decrease in decision quality. The researchers do not report hallucination rates. To assess quality, the team asked OpenAI's GPT-5-mini to choose between pairs of judgments from the same judge before and after training; it chose post-training judgments 59 percent of the time. Two experienced Pakistani lawyers agreed with GPT-5-mini 70.6 percent of the time, compared to 73 percent agreement with each other.

Training proved vital. JudgeGPT-trained judges logged in 56 times and sent 212 prompts on average over the study period, compared to 10 logins and 25 prompts after generic training. Those with no training tended to use the tool for about a month and then drop off. The researchers calculated that a trained judge resolved 38.5 more cases a month than baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool.

David Autor, an economics professor at MIT, called the study a notable large-scale field experiment. John Zeleznikow, professor of law and technology at La Trobe University, said it was not clear whether the quality of justice improved. The working paper's authors found roughly a fifth of participants' prompts involved "substantial AI delegation," though training lowered the proportion of inappropriate delegation. The story was updated on 13 August 2026 to clarify that the training course was developed in coordination with Pakistan's Federal Judicial Academy.

Read at IEEE Spectrum · AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Judges around the world have made headlines for illicitly using generative AI in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology—together with appropriate training–boosted the number of cases resolved by 6.3 percent with no obvious drop in the quality of judgments. With a backlog of 2.26 million cases and fewer than two judges per 100,000 people—compared to 22 in the EU and eight in Brazil—Pakistan’s judiciary was in sore need of help. So, in consulta