In 2024, Laurence Holt of the XQ Institute published an essay titled “The 5 Percent Problem,” a term that quickly became common parlance within the ed-tech zeitgeist. Holt argued that when it comes to measuring the student-learning outcomes of various ed-tech programs, the results often only apply to the five percent of students who “used the program as intended. The other 95 percent see minimal gains, if any.”
Holt’s primary focus was on online math programs, and the central concern running throughout his essay involves whether and how students might “get the recommended dosage” of using various ed-tech tools. Whether the problem arises from unmotivated students or unenthused teachers or disparate access to tech or some combination thereof, the challenge when framed this way is getting students to opt in to using a particular technology.
Generative AI does not suffer from this problem. Uniquely perhaps in the history of ed-tech, AI has been broadly embraced by students worldwide, so much so that the phrase “AI is inevitable” has become commonplace in the education discourse. There are pockets of resistance, of course, and over the last six months we’ve seen mounting opposition to AI across multiple vectors, most prominently with data centers. Nonetheless, we all know that students are using AI frequently. “Dosage” is not an issue.
In the nearly four years of time that’s passed since ChatGPT was commercially deployed, however, we have suffered from a lack of high-quality empirical research on the impact of students using generative AI. There are notable exceptions—such as these—but a recent research landscape analysis out of Stanford indicated that, out of more than 800 studies of AI in education, a mere 20 employed true causal measures. Meanwhile, at least one prominent meta-analysis purporting to show massive learning gains stemming from AI has been retracted, though not before being viewed 400,000 times. The field is a mess.
Perhaps this is because studying AI in “the real world” poses serious challenges. We know that AI tools are free and widely available to students. We also know that many are using them after school hours for various purposes, including to help with their schoolwork. But quantifying this is very difficult, because it’s hard to peer into the home lives of children. And it’s equally if not more challenging to connect out-of-school behavior to measurable learning outcomes, the data sets are not easy to link. Accordingly, research on AI’s education impact has tended toward lab studies or small-scale evaluations.
No longer.
Last week, a study titled “The Generative AI Learning Penalty: Evidence from Chinese Secondary Education,” authored by David Strömberg, Victor Lei, and Yanhui Wu, went viral after The Economist published data visualizations from the research paper. The preprint came out in June—not sure how I missed it—and uses data from approximately 27,000 Chinese students in grades seven to 12 to explore a straightforward question that has preoccupied me for several years: “How does self-directed use of generative AI affect cumulative learning over time in ordinary school settings?”
We will get into the details shortly, but here’s the headline summary of the results:
Between 2023 and 2025, approximately 80 percent of all students started using generative AI, and approximately 50 percent of all students engaged in “full homework outsourcing,” meaning, they essentially stopped doing their homework independently and used AI instead.
Over time, this caused significant learning loss as measured both on monthly closed-book exams and via the comprehensive high-stakes exams that China uses to determine high school and college placement for students. On those tests, by the end of June 2025, overall student performance fell by 24% on the former and 18% on the latter, a massive decline.
What’s more, there is no evidence—none—of any corresponding learning benefit arising from students using AI.
As such, the researchers state plainly that “our findings show that generative AI, which is likely to become a prevalent technology for education, has a substantial negative impact on student learning.” (My emphasis)
Put simply, we now have rigorous, empirical evidence on the real world, long-term impact of students using AI. This research indicates that AI substantially harms learning for half of all students, what the researchers call the “Generative AI penalty,” and provides no discernible benefit to the other half. Again, by the time this study concluded, 50 percent of students had fully outsourced their homework to generative AI. For these students, when at home, they just stopped thinking about their schoolwork.
I hereby dub this the 50 Percent Problem. Unlike the Five Percent Problem, the 50 Percent Problem is a measure of harm stemming from widespread adoption of AI by students, rather than their failure to use this technology. It is my conservative estimate of how many students are being directly harmed by using these tools.1 As I’ve said repeatedly for several years, generative AI is a tool of cognitive automation. It’s both predictable and tragic that students are learning less because of it.
This new empirical research provides us with a clear measure of the degree of AI’s educational harm.
Let’s turn to the data itself. In what follows, we’ll be looking at the performance of three different groups of students across three different learning activities.
As to the students, we first have the baseline group of student results from the period prior to the introduction of generative AI—these students are labelled “Pre AI,” and charted in green. Next, we have the admirable group of students, all 19 percent of them, who did not use generative AI at any point between 2023 and June 2025 when the study concluded—these are labelled “Never AI,” and charted in blue. Finally, we have the remaining 81 percent of students who adopted AI and used it for their schoolwork—these are labelled “Used AI,” and charted in red. Of note, these students did not all start using AI at the same time, adoption phased in over time (which allowed the researchers to uncover some interesting things that we won’t go into here).2
As to the learning activities, we’ll start by looking at the amount of time that students spent on their homework; the researchers were able to track this because students had to log in online to download their homework and upload it upon completion. Then, we’ll look at how students scored on this same homework. Last but definitely not least, we’ll examine how student performance changed on monthly closed-book exams, as well as China’s very high-stakes high school and college entrance exams.
Both the underlying research paper and The Economist present a series of helpful data visualizations of this, but without the underlying raw numbers. So I contacted the researchers who conducted the study and David Strömberg (the lead author) graciously agreed to provide me with the underlying histogram data I used to create the three animated GIFs you’re about to see.
With that as backdrop, we can explore three specific questions.
1. What is the impact of AI on the amount of time that students spend on their homework?
The results here are unsurprising—AI significantly reduces the time students spend doing homework. On average, both the Pre AI and Never AI groups spent about an hour doing it, across a range of 50 to 80 minutes. In contrast, within the Used AI group, most students spent around 45 minutes, and some even cruised through in 25, presumably the minimum amount of time it takes to cut-and-paste answers out of a chatbot.
This is AI as tool of cognitive automation working exactly as intended.
2. What is the impact of AI on how students score on their homework?
Here again we see that Pre AI students and Never AI students have near-identical results, as we’d expect. But not so with the Used AI students—now, we see a huge shift to the right in purportedly “positive” outcomes, with scores far higher than even the most studious Never AI student managed to achieve. (Note the scores here were normalized to make the average score a 100 (not the maximum), so a score of 130 means “30 percent above the average,” essentially.)
To restate, the Used AI students studied less yet scored better than their peers. That’s a pretty sweet deal for them, but it’s worth reflecting on the broader implications of this within schools. When I talk to students about AI, one thing I hear them say frequently is that they don’t want to be played for chumps (my term, not theirs). Meaning, if they know their classmates are using AI and getting better grades as a result, even those inclined to resist AI may feel trapped into using it, just to keep pace.
It’s a cognitive race to the bottom.
I’ll also add that homework scores are obviously a very imperfect measure of student learning. In my view, a great deal of education research, and certainly the studies often promoted by ed-tech vendors, use comparable “point in time” data of supposed learning akin to what we see here. This is understandable, to a degree—it’s very difficult to information on long-term learning outcomes. But what ultimately matters in education is building durable student knowledge. The seductive danger of AI exposed here is that, by using AI, students may falsely have believed everything was proceeding swimmingly. In this sense, AI fosters a mental masquerade, scores go up as actual learning goes down.
Now to remove the mask.
3. What is the impact of AI on how students perform on closed-book exams?
Here is where the proverbial rubber meets the road.
Consistent with the two previous data sets, we again see near-perfect alignment between the Pre AI and Never AI students on the monthly closed-book exams administered to students. But now the adverse impact of AI is laid bare—just look at that shift to the left. The Used AI students are scoring at levels far lower than their Never AI peers, indeed, many score far lower than anything recorded prior to AI existing.
I’ve labeled this the Deadweight Learning Loss to underscore the volume of harm caused by the generative AI learning penalty. What’s we’re seeing here is unambiguous evidence of a decline in overall student performance that grew over time as students adopted AI. In fact, the average decline was 20 percent with a 1.4 standard deviation (SD). Although it’s far from an apples-to-apples comparison, this vastly exceeds the estimated learning loss in the US after the pandemic (approximately .25 SD in math and .13 SD in reading).
What’s more, over time this Deadweight Learning Loss had significant adverse consequences for students on the comprehensive high-stakes high school and college entrance exams that China administers to determine student placement.3 For students who adopted generative AI two or more years prior to being tested, the estimated negative effects are 24 percent (1.5 SD) for the high school exam, and 18 percent (1.3 SD) for the college exam.
If you are a parent with a child in junior high or high school who uses generative AI, this data should terrify you. This isn’t about academic integrity, it’s about a generation of kids being told AI is “inevitable” and “the future” and acting accordingly, in a societies where we’ve yet to develop firm norms about what’s acceptable to do with these tools within education. The upshot is that students are using AI in ways that are harming their cognitive development and foreclosing their life opportunities. Harms that are not easily remediated, if at all.
This is the 50 percent problem. Generative AI is cognitive cancer and we are doing next to nothing to stop it from spreading.
I’ll now address a few objections.
First, I want to credit The Economist for putting this research on my radar, and for raising broader attention to its harrowing findings. Yet, remarkably, in the brief article accompanying the research data we’ve just covered, the anonymous magazine author suggests that “AI can boost learning productivity but only for those who use the technology intelligently.”4 In similar fashion, Blake Richards, a researcher at Google who works “on the intersection of machine learning and neuroscience,” argued on social media that AI is not really the problem here, because “if you control for how long students spend studying, then students using AI actually perform equal or better” than those that did not.
Motivated reasoning is a powerful force. Here, of course, “controlling” for how long students spend studying erases the major findings of this research. But let’s leave that aside, and probe whether this argument can be justified on its own terms—does this study suggest that AI can boost learning if used “intelligently”?
As best I can tell, this claim is premised on this graph that correlates homework completion time with exam results:
If we ignore all that pesky data on the left side of this chart (which of course we shouldn’t), and just focus on the overlap between the Generative AI students who continued to study for the same duration as the No Generative AI students, it’s true there’s no major gap between them. Of course, note that there’s no discernible benefit to using AI either—apart that is from the spike at the very top end.
So what’s the deal with the spike? Might we cling to it as proof of the “promise” of AI? Well, here’s a good lesson on why one should always be careful eyeballing charts without access to the relevant underlying data, because when I asked Strömberg how many Used AI students persisted in studying for at least 75 minutes, he told me this (via email, with my emphasis):
It is very rare for students who have adopted AI to spend 75 minutes on homework. This occurs for only 20 students, and for each of them only in a single month (0.2% of AI student-month observations). More than five months after AI adoption, we never observe students spending 75 minutes on homework. In fact, in this group, only four students spend more than 65 minutes on homework.
Let’s get real. Given this study involved almost 27,000 students, indexing on this vanishingly small number of studious AI-using students isn’t just putting lipstick on a pig, it’s smearing its body in Revlon from snout to tail. We need to stop pretending there is some massive benefit to learning if only we trained kids properly on how to use AI. It’s fundamentally harmful.
A different and more sophisticated counterargument might proceed along the following lines: This study tracked student usage of AI stemming from its earliest days, when the tools were less capable than they are today. Perhaps relatedly, the researchers here found evidence that it was the early AI adopters most harmed by generative AI, insofar as “the estimated AI learning penalty fell from around 25 percent in early 2023 to around 16 percent by June 2025.” As such, perhaps this study constitutes the “high-water mark” of AI-induced harm, and perhaps the AI learning penalty will continue to drop. Or so we might hope, anyway.
But as the cliche goes, hope is not a strategy, and there are some problems with this counterclaim. For one thing, the rapid evolution of AI cuts both ways—we now have AI companies explicitly marketing “AI agents” to students to complete their tasks for them. Agents are even worse than chatbots from a cognitive development standpoint—at least the latter require an interaction of some sort to produce output. For another, even if the magnitude of the learning penalty continues to diminish, there’s a question of volume, too. Recall that 20 percent of students managed to resist using AI prior to June 2025. Do you think that number has gone up or down since then? I know my bet.
Finally, I can imagine someone saying I’m placing too much weight on this research—it’s just one study from China, after all. On that front, and as a self-sanity check, I asked three PhD education researchers to review the methodology employed, and all three came back with positive reviews (“it’s very good work,” said one). And China is surely the one country in the world that can match the US for AI adoption and enthusiasm, though it may surprise you to learn that AI companies in China disable their products completely during the high-stakes testing periods. (OpenAI, in revealing contrast, heavily promotes ChatGPT on college campuses during finals week in the US.)
Perhaps more importantly, if you are an AI-in-Education Enthusiast, I feel confident in saying that you will not be able to produce research of comparable rigor that shows positive long-term education impact of generative AI in real-world conditions comparable to those here. If there were such evidence, my inbox would be filled with people jamming it down my throat, trust me. And no, this single study from 2024 involving roughly 150 physics students at Harvard is not going to cut it, sorry.
The 50 Percent Problem will not disappear of its own accord. The Edu-Cognoscenti continues to chatter about learning loss related to school closures during the pandemic. Well, the learning loss stemming from generative AI appears more substantial, it’s happening right now, and it may endure for far longer. So to all the policymakers and philanthropists and so-called thought leaders who purport to take education evidence seriously, I ask you: What are you doing to prevent ongoing educational harms of generative AI? Are you doing anything at all?
We are in cognitive crisis.
I’ll close with this. As I was drafting this essay, my friend Dan Willingham published his perspective about when students should use AI. Please read it. Although he’s a tad more measured (or perhaps just realistic) about its role in education, his core contention echoes what I’ve been arguing for several years, and stems from a basic understanding of human cognition:
[T]he point of assignments is the mental processes required to complete them, and the point of the mental processes is learning. That seems to suggest a simple litmus test for the use of AI. Artificial Intelligence tools should not substitute for tasks wherein students would benefit from doing the mental work themselves. Only use AI for what you already know how to do.
A tool that should only be used once you already know something is not a learning tool. We must stop gesturing at AI’s imagined potential, and focus our efforts instead on mitigating the 50 Percent Problem it’s created.
How, you might reasonably ask? It won’t be easy. But I have some emerging ideas, stemming from recent investigations into communities of technological refusal. It’s been quite the learning journey. More soon.
My thanks to David Strömberg for providing his underlying data and to my anonymous academic friends who reviewed the study and this essay—all errors mine and mine alone, of course.
In my view, 50 percent is a conservative estimate because the researchers themselves describe the core problem in more expansive terms: “The negative effects on learning outcomes appear to be mostly driven by the 81 percent of AI-using students, who spend less time on homework than even the fastest non-AI student, receive high homework scores matching the capability of generative AI tools they are using, and yet very low exam scores.” So I was tempted to call this the 81 Percent Problem, but as discussed above, not all of the 81 percent of AI-using students suffered the AI learning penalty (but nor did they gain any meaningful benefit). Ultimately, it’s the half of students using AI to complete their homework that I’m most worried about.
Alert, nerdy readers may be wondering why “student observations” is listed on the Y axis rather than just “students.” Answer: because the students groups were not static and evolved in composition as students adopted AI, the researchers used “student-month” as their unit of analysis, meaning, the data for students was parsed by month—e.g., a single individual student might produce 12 separate “observations” over a year within a single subject. I know, it’s wonky.
Of course, whether China or any other education system should employ high-stakes testing to this degree is a highly charged topic, but I’m not interested in having that conversation right now, please and thank you.
The Economist also cited this recent study involving the use of chatbots with roughly 200 undergraduates at Middlebury College, conducted over two sessions approximately one week apart. This is a perfect example of a point-in-time education study that bears no resemblance to the reality of how most students are actually using AI.








Great essay, thanks! What struck me is that the "Used AI" group do not have a normal distribution for any of the plots. That seems unusual, since most educational variables tens to result in normal distributions. That made me think there's two populations of students within this group. Is there anything in the data to suggest what it might be? If you look at the time spent on homework graph there's like a second peak near the same place as the "never AI" peak. I wonder if this group corresponds to older cohort of students who had access to older, less capable models? Will be curious to look through the data if you can share it. Thanks!
Great and helpful write-up!
It is a huge challenge right now, but part of addressing the challenge requires acknowledging what’s going on, and this kind of evidence is crucial for that.