AI Assistance Reduces Persistence and Hurts Independent Performance

Artificial Intelligence

, a preprint by a team from Carnegie Mellon University, University of Oxford, Massachusetts Institute of Technology and UCLA offers a well-designed "randomized controlled experiment (N = 354) on fraction-solving tasks. Participants were randomly assigned to an AI condition or a control condition. In the AI condition, participants solved 12 fraction problems with an AI assistant (GPT-5) available in a sidebar. The AI was then removed without warning, and all participants solved 3 additional test problems independently. Participants in the AI condition had a significantly lower solve rate (mean 0.57 vs. 0.73; p < 0.001, Cohen's d = −0.42) and higher skip rate (mean 0.20 vs. 0.11; p = 0.031, Cohen's d = 0.25) than control participants." The authors took care of obvious confounders such as skill levels and similar effects were replicated for reading comprehension. The conclusion stated in the title seems warranted and supported by the evidence. One might argue that a 10-minute test where persistence is measured is of limited applicability, and indeed it is when taken literally (for example an interesting question would be do the effects persist after 1 week?). But I think the more interesting result is in the middle of the paper, when the authors analyze how the AI was used: "Persistence costs were concentrated among participants who prompted AI to solve tasks for them directly. Using AI for hints or clarifications did not produce significant impairments." In other words, and I think it matters a lot, how the AI was used was the key differentiator, more than AI vs not-AI, which goes to show that all these half-baked experiments that do not monitor or control how the AI is used and generate "AI is destroying your brain" headlines are doing the whole field a disservice. Using AI is not a sufficiently accurate description of an experiment. The biggest work that remains to be done is precisely to design the best way to use AI: here we see that using it correctly does not impede abilities.

Not me (but I wish it were)

I want to address persistent rumors: my company DID NOT spend $500M on Anthropic tokens accidentally. Not even on purpose. Now that we have cleared that misunderstanding, I also want to say that I would be very happy to have $500m to spend by mistake on tokens (although my mistakes would probably be more diversified). I'm raising if that's appealing to any token-pilled investor.