TLDRocket
Sign in

Pakistani Judges Give Their Verdict on JudgeGPT

IEEE Spectrum Edd Gent

Pakistan tested JudgeGPT with judges, and it helped them clear 6.3% more cases. The twist: training mattered, and the tool didn’t seem to make rulings worse.

Based on reporting by IEEE Spectrum, Edd Gent — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Pakistan’s courts are drowning. The country has 2.26 million pending cases and fewer than two judges per 100,000 people, a number that looks brutal beside 22 in the EU and 8 in Brazil. So when a team led by economist Sultan Mehmood set out to test AI in the judiciary, the idea wasn’t flashy. It was triage.

The result, though, is more interesting than a simple productivity story. The researchers built JudgeGPT, a custom system based on OpenAI’s GPT-4 and a legal database of 128,292 Pakistani judicial opinions plus 943 statutes. Judges could use it for legal research and drafting, and every answer came with footnotes pointing back to the underlying cases and laws. That matters, because the commercial chatbots some judges were already trying had a habit of making things up.

The trial started in 2024 and reached 1,559 trial judges, about half of Pakistan’s justices. Another 1,197 judges went through six 90-minute Zoom training sessions on how large language models work, where they fail, and why verification still counts. A smaller group got generic tech training, and another group got none. The pattern was clear: the more training judges received, the more they used the tool, and the more cases their districts resolved. By the time 487 judges had completed the program, the median district was clearing 6.3% more cases.

That gain did not come with an obvious quality hit. Appeal rates fell slightly, and the team’s checks on judgment quality suggested post-training decisions were favored. The researchers also saw a big gap in usage: trained judges logged in 56 times and sent 212 prompts on average, while judges with only generic training used it far less, and those with no training tended to abandon it after about a month. In other words, software alone wasn’t the story. Habits were.

The economics are hard to ignore. The team says a trained judge was resolving 38.5 more cases a month than baseline, which they estimate at about US $38.50 in judicial cost savings for every dollar spent running the tool. But the study also found that about a fifth of prompts involved what the authors call substantial AI delegation, where judges asked the system to suggest the decision or draft reasoning with little input. That is the part to watch: not whether AI can speed up courts, but whether courts can keep their hands on the wheel while it does it.

My take — AI-written commentary, not fact-checked reporting

This is the sensible AI story everyone keeps pretending doesn’t exist: use the tool, but chain it to citations and training, or enjoy the hallucinations. The real scandal isn’t judges using AI; it’s courts being so overloaded that bad software starts looking like relief. Europe loves rules, America loves demos, and Pakistan just needed a system that works without acting like a genius in a suit.

Read more about this at: IEEE Spectrum

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.