Modern states have become bottlenecks for economic growth, trapped in endless backlogs and chronic labor shortages. However, a large-scale experiment by researchers from ETH Zurich and Imperial College London proves that the solution lies not in expanding payrolls, but in a marriage of heavy-duty AI and rigorous training. In Pakistan, where 2.26 million cases are gathering dust, the implementation of JudgeGPT has shown that technology becomes a force multiplier precisely where traditional scaling is impossible. This is not a story about automating paperwork; it is about technological leapfrogging—when a system without basic digitalization adopts cutting-edge solutions, bypassing decades of intermediary bureaucracy.
An extreme testing ground
Pakistan serves as an ideal "red zone" for testing institutional resilience. There are fewer than two judges per 100,000 people, compared to 22 in the EU. This talent drought meant that by the end of 2024, 82% of all cases were hopelessly stuck in courts of first instance. The participants' technological background was near zero: 75% of the 1,559 judges had never even heard of large language models. Researchers equipped them with JudgeGPT, based on GPT-4 and enhanced with a Retrieval-Augmented Generation (RAG) mechanism to search a database of 129,000 documents, including nearly a thousand Pakistani laws and a massive archive of judicial rulings.
The failure of "naked" access
The data confirms a healthy skepticism: simply providing access to advanced tools is a waste of money. The experiment divided judges into groups, and the results were binary. Those who were merely given a login and password barely saw any improvement. In contrast, judges who underwent an intensive course—six 90-minute lectures by Professor Elliott Ash—used the AI four times more frequently. After 40 weeks, the trained lawyers were generating an average of 200 prompts, while their untrained colleagues couldn't even reach 50.
In districts with trained judges, 1,848 additional cases were closed annually, representing a 6.3% increase in productivity.
Crucially, this speed did not turn justice into a low-quality assembly line. An LLM-based audit, validated by local attorneys, showed that the quality of decisions for the "upgraded" judges actually improved, while the appeal rate per 1,000 cases saw a slight decrease.
The $38.50 multiplier
For architects of government systems and business leaders, the most sobering figure is the ROI. Researchers estimated a return of $38.50 for every dollar invested, comparing software costs to the expense of hiring human assistants to achieve the same output. Even under conservative estimates, the return does not drop below 10-to-1. This efficiency stems from a radical shift in the delegation model: AI takes over the grunt work that, in developing economies, there is simply no one else to perform.
Since JudgeGPT ran on GPT-4, these performance metrics are likely a floor rather than a ceiling. In districts with a high concentration of trained personnel, the system cleared an extra 1,848 cases per year. Even the most sluggish regions in the bottom quartile managed to close 616 more cases than they would have without AI. We are seeing a classic example of how resource scarcity forces a jump over developmental stages: while Western courts spend years debating the ethics of implementation, the Pakistani case demonstrates that the alternative to an AI assistant isn't a "perfect human judge," but the total absence of justice as a service.