AI trained judges cleared 1,848 more cases per district in Pakistan trial
Curated by the Inblix editorial team
Pakistan’s courts are drowning. With fewer than two judges per 100,000 residents and 2.26 million pending cases, the system barely functions. Now, a massive randomized trial suggests AI can throw it a lifeline — but only if you show people how to use it properly.
Researchers from ETH Zurich, Imperial College London, and the New Economic School gave 1,559 judges access to JudgeGPT, a custom tool built on GPT-4 that searches a database of 129,235 Pakistani legal documents. The catch? Simply handing over the keys did almost nothing. Judges who got six 90-minute training lectures from ETH Professor Elliott Ash used the tool four times as much as those who only attended a general tech seminar.
Those trained judges didn’t just play with the tech — they moved cases. Districts with moderate exposure to trained judges resolved about 1,848 extra cases per year, a 6.3 percent productivity bump. The economics are hard to argue with. Researchers peg the return at $38.50 for every dollar spent. Even their most conservative estimate lands at a 10x return, based on what it would cost to hire enough judges to match the output.
Quality didn’t suffer for speed. Appeal rates dipped slightly. A review of 4,000 judgments found readability and legal rigor held steady, and a review by Pakistani lawyers actually showed a quality improvement. The study found no evidence AI amplified gender or religious bias. Crucially, training shaped how judges used the tool — steering them toward editing and summarizing, where language models are reliable, and away from asking it to draft legal reasoning from scratch. The researchers are clear: this isn’t about replacing judges. It’s about making the ones we have dramatically more effective.
💡 Key Takeaways
- Simply giving judges access to AI did nothing; hands-on training that taught them which tasks suit the tool quadrupled usage and drove all productivity gains.
- The estimated $38.50 return per dollar invested is based on the cost of hiring enough judges to achieve the same 6.3% output increase, making this a dirt-cheap intervention by government standards.
- Training steered judges toward safe tasks like editing and summarization, with only 20% of prompts involving risky requests for AI to draft legal reasoning on its own.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.