24 Comments
User's avatar
Alok Bharadwaj's avatar

Great essay, thanks! What struck me is that the "Used AI" group do not have a normal distribution for any of the plots. That seems unusual, since most educational variables tens to result in normal distributions. That made me think there's two populations of students within this group. Is there anything in the data to suggest what it might be? If you look at the time spent on homework graph there's like a second peak near the same place as the "never AI" peak. I wonder if this group corresponds to older cohort of students who had access to older, less capable models? Will be curious to look through the data if you can share it. Thanks!

Benjamin Riley's avatar

Thanks Alok, and you ask insightful questions. Couple of comments. First, as noted in the footnote, the unit of analysis here was what the researchers call "student-month," meaning they were looking at data collected across the cohort by month across seven different subjects. Second, the Used AI and Never AI groups were not static -- students moved from the latter into the former over time. Third, with the homework scores in particular, I think it's very telling that the data doesn't fall into a normal distribution, but instead clumps to the far tail of positive outcomes -- which is what we might expect from cutting-and-pasting AI, there's no real variation.

I don't think I should share data myself but if you contact Stromberg I suspect he'll be happy to send along to you as well. There's quite a bit of detail in the research paper, too.

Sean Trott's avatar

Great and helpful write-up!

It is a huge challenge right now, but part of addressing the challenge requires acknowledging what’s going on, and this kind of evidence is crucial for that.

Benjamin Riley's avatar

Thanks Sean, much appreciated. And I keep wondering -- do we have any way collecting data on this in the US? We could use SAT and ACT scores, but I don't know how to capture the AI usage at home.

I know this isn't your field, but...any ideas?

Sean Trott's avatar

Good question...

So just to work through the study's logic:

1) if I understand correctly, it looks like they were able to use a pretty massive 2025 survey asking students about (among other things) their AI usage, including student reports of when (month/year) they started using AI.

2) They were then able to merge these data with data about homework/test performance provided by local educational bureaus.

3) The study then shows that *pre-AI*, there are no differences between the students that would go on to use AI (ever-AI) and the students who didn't go on to use AI (never-AI), which is crucial for estimating causality...and then differences show up between the groups once AI is actually adopted.

With the caveat that I don't know too much about this kind of study design: in principle, I don't see clear obstacles to conducting such a survey in the US. The more challenging part (I'd think) would be *linking* responses at the individual student-level to performance measures, which I'm assuming are more fragmented across districts than it sounds like they are in China. I.e., step (2) above. But I could be wrong! Maybe SAT/ACT scores, as you say, though one issue here is that they're only taken a limited number of times per student (often once), so you get less within-student comparison.

Benjamin Riley's avatar

Right, exactly, your understanding accords with my own. Point (2) definitely poses challenges, though (thinking aloud) I wonder if a charter network might be game to provide data. If we had a functioning federal Department of Education that would help too (laughs maniacally). China's investment in its own research infrastructure is really visible to me in this study.

Alok Bharadwaj's avatar

Thanks for the kind reply! Yeah, I guess if the students moved from Never AI to Used AI over time, it makes sense why the Used AI group gets distorted. I expect the early switchers among this group to suffer more from cognitive offloading than later ones. Probably it's there in the paper, so I'm curious to check it out.

Benjamin Riley's avatar

Your astute suspicion is borne out in the data (and I mention this briefly in the objection section at the end), the early adopters have indeed suffered greater learning loss.

Marla Simpson's avatar

So grateful to you for sharing information about this study! I had the same reaction: that this was a methodologically sound study that would be hard to wave away. For those of us dealing with AI intrusions on a daily basis, it's nice to have the empirical validation. I really think that the US should be looking to Australia for inspiration about how to fix this problem. They are moving to a very concrete set of standards for assessment. I wish we were.

Benjamin Riley's avatar

Thanks Marla, I'm glad you found the study as compelling as I did. I've done some work in Australia and have been tracking their social media ban for children, but not sure what you're referring to with standards for assessment -- tell me more?

jwr's avatar

Really appreciate your discussion of this! One of the things that really stands out to me is the way that the negative effects appear to have accumulated over time: the longer students had been using AI, the more likely they were to outsource/fully outsource their homework, the greater the learning loss the study saw. Not great!

One question I have, having recently worked with a student who did her project on the Chinese education system, is how far conditions specific to that system may have shaped the dynamics seen in the study. As my student explained to me, the experience of students ramping up for the Zhongkao/Gaokao is very different from the experience of students in the US. To what degree does that matter for the way they used AI and the way it affected their learning?

Benjamin Riley's avatar

Thanks for the kind words, and being a regular reader/liker (I've noticed!). On first point, you're absolutely right about the learning loss accumulating over time, it's one of the reasons I found this research so striking -- rarely do we see 30-month data sets.

On point two, I am not an expert on the Chinese education system by any means. There is however a brief discussion in the paper about how student study habits intensify in the high-stakes testing years. My hunch (just a hunch): in countries that don't have "higher intensity" study years as China does, the learning loss might be even greater, because there's no offset when students go hard on exam prep. But this would be a fruitful area for empirical research -- see my dialogue with Sean Trott above.

Norman Fischer's avatar

Thanks!

Not surprised but good to get hard data to back up common sense knowledge that practice makes people better at things

Conn McQuinn's avatar

At the cellular level, learning happens through repetition. The connections made between neurons are made stronger through repeated firings, especially if done spaced out over time. The converse is true; learning that is not reinforced will degrade, or may be pruned away entirely.

As to even the more-conservative idea of “only use it for what you already know,” talk to any professional, be it a musician, athlete, artist, or other person engaged in a skill, and ask them when they got good enough to give up practicing.

Benjamin Riley's avatar

Indeed, it is unfortunate that generative AI offers such a convenient off-ramp for students to skip the sustained practice required by their homework assignments.

Ted Frick's avatar

This is not surprising, given all the past research on Academic Learning Time (ALT) as being a good predictor of student learning achievement. Decades of empirical research going back 50+ years document that the more time students are engaged successfully in tasks similar to those they will be later assessed on, the better they tend to do on those assessments. "Past research has also shown that ALT is strongly and significantly correlated with student academic achievement as measured by standardized tests (cf. Berliner 1991; Brown and Saks 1986; Fisher et al. 1978; Kuh et al. 2006)." (https://bio.tedfrick.me/etrd/FrickTALQ2009.pdf ).

This article provides a succinct overview of ALT: https://www.edpsycinteractive.org/topics/process/ALT.html .

The data you have summarized in your post indicate that students who are using AI more often appear to be successfully engaged less often in the tasks they are expected to learn to do themselves. If more use of AI is reducing ALT, then students who use AI to do their homework would be expected to do less well on tests that measure learning achievement in that content area.

i code for joy's avatar

Love the article - thanks for writing and for the visualizations. I am intrigued by the change in the shape of the 'Used AI' histograms - Pre-AI and non users seem to be normal but the AI user distribution has some skew across the three visualizations...

Benjamin Riley's avatar

See my comment to Alok for more detail.

Benjamin Riley's avatar

(And also, thanks for the kind words! Forgot to include that, much appreciated. )

Craig Yirush's avatar

So Chinese students are falling behind because of AI use; meanwhile we’re are being told by so-called experts that China is eating our lunch in all areas, including AI.

Benjamin Riley's avatar

From what I can glean, China is taking this problem much more seriously than we are. There is a bit of Catch 22 that the places where AI is most advanced are most vulnerable to the problem of students using it, that's for sure.

Craig Yirush's avatar

More seriously in what way? Was the study you discussed not using China as an example of the problem.

Benjamin Riley's avatar

Well, as noted in the essay, the major AI companies actually disable their products during the high-stakes testing period. There are also new policies designed to curtail student use throughout the school year: https://www.cnbc.com/2025/05/15/key-ai-hub-china-restricts-schoolchildrens-use-of-the-tech.html

Stefano's avatar

Basically between smartphones, social media and generative AI we're going to create generations of retarded insecure goldfishes.

Technology is so friggin amazing.

We should promote gen AI usage with the catchphrase: it makes you retarded.

What a disaster technology is turning out to be.

I have a feeling this study is going to get replicated.

Thanks for the writeup and making it available!