Widespread adoption of generative AI tools has delivered an intoxicating illusion of instant productivity. While software assistants generate immediate spikes in output volume, measuring actual cognitive mastery requires testing what remains once those automated crutches are pulled away. A major empirical investigation now demonstrates how delegated problem-solving builds a brittle operational facade that steadily erodes unassisted capability.
The 27,000-Student Experiment
To quantify the structural toll of generative tools on skill acquisition, David Stromberg of Stockholm University, alongside Victor Lei and Wu Yanhui of the University of Hong Kong, tracked a cohort of 27,000 pupils aged 12 to 18 across China. Roughly 80% of the students routinely offloaded assignments to LLMs like Doubao and DeepSeek, while the remaining 20% abstained, serving as a clean baseline control. This scale reflects a broader global shift highlighted by the researchers: edtech platform Chegg reports an 80% adoption rate among undergraduates in developed economies, matched by 94% in Britain and 93% in Germany.
Over a six-month evaluation period, the data exposed a stark divergence between assisted throughput and raw competence. As reported by The Economist, students using generative models saw average homework marks climb 18% across all subjects. Yet, when evaluated under strict, closed-book exam conditions without AI access, those same students plummeted 20% below peers who completed their coursework unaided.
Students who used AI saw their average homework scores rise by 18% across all subjects over a six-month period, but scored 20% below classmates on unassisted exams.
This collapse mirrors earlier findings from a 2024 University of Pennsylvania trial evaluating math students using ChatGPT and specialized AI tutoring software: short-term practice gains vanished the moment pupils faced an unassisted testing environment.
Masked Capability and Corporate Blindspots
The gap between routine task execution and unassisted evaluation reveals how algorithmic automation disguises foundational decay. As educators noted formulaic patterns in submitted work, cognitive offloading emerged as the core failure mode. Outsourcing synthesis to an LLM produces a polished end product while bypassing the friction and neural effort required to internalize complex principles.
Guidance from the Brookings Institution underscores this risk, warning that substituting AI for critical thinking and creative friction stunts core cognitive architecture. For engineering leaders, CTOs, and L&D managers, this dynamic carries immediate operational risks. Measuring junior developers or analysts solely by the velocity of their copilot-assisted output creates a dangerous false signal of mastery. Building durable organizational capability demands instituting rigorous, unassisted baseline audits before trusting teams to supervise automated workflows.
While the full preprint from Stockholm University and the University of Hong Kong awaits final peer review, the empirical signal is unmistakable: relying entirely on AI-generated artifacts inflates superficial output metrics while gutting independent problem-solving capacity. Engineering resilience requires closing the feedback loop with unassisted evaluations to verify genuine baseline competence before deploying automated pipelines.