Tutorial

How to Learn With AI: The Practice Went Great, the Test Didn't

You want to understand something new. A tool your team is adopting, a part of the business you have never worked in, a technique everyone assumes you know. So you ask an AI to explain it, and it does, clearly and patiently, for as long as you like.

An hour later you feel like you understand it. There is now good evidence that the feeling and the understanding come apart.

The experiment

Researchers ran a The strongest common study design. You assign people to groups at random, so the groups differ only by the thing being tested and not by who chose what. Most of the AI-at-work research this site covers is survey data, where people report on themselves and you can never fully rule out that the keen ones simply answered differently. A trial does not have that problem, which is why one is worth several surveys. across roughly a thousand high school students, in about fifty classes, over four 90-minute math sessions. Three groups:

  • Control. Notes and textbook. No devices.
  • GPT Base. A standard ChatGPT-style assistant.
  • GPT Tutor. The same underlying model, set up by their teachers to give hints rather than answers.

Then everyone sat the same exam, closed book, with no AI.

During practice, the AI groups pulled away, exactly as you would expect. GPT Base scored 48% higher than the control group. GPT Tutor scored 127% higher.

On the exam, that reversed:

  • GPT Base scored 17% lower than the control group, a statistically significant drop. Worse than the students who had nothing.
  • GPT Tutor scored the same as the control group. Not better. The difference was too small to distinguish from zero.

The group that looked most impressive while practicing, more than twice the control group’s score, walked into the exam with nothing extra to show for it.

Nobody in the room could feel it

Here is the finding that matters most for you, because it removes the defense you are probably reaching for.

The students in the GPT Base arm did not perceive that they had performed worse or learned less. They had spent the sessions getting more problems right than anyone else. Everything about the experience seemed to be going well.

This is the ordinary shape of an The gap between how well you feel you know something and how well you can actually produce it unaided. It is reliably produced by anything that makes material feel easy while you are studying it: rereading, highlighting, following along with a worked solution. Ease of understanding gets mistaken for durability, and the two are close to unrelated. , and AI is an unusually good machine for producing it. A clear explanation delivered on demand is the most fluent possible version of the material, which makes it the most misleading possible version of how much of it is yours.

It is not that the AI got things wrong

The obvious explanation is errors. It is not the explanation.

GPT Base was correct on only 51% of the math problems, so there were plenty of errors to blame. The researchers checked, and found that students did worse on the specific practice problems where the model erred, but that this did not carry through to the matching questions on the exam. Bad answers were not what depressed the exam scores.

What did was how the tool was used. In the GPT Base group, 67% of messages across a session were superficial: restating the question, or straight out asking for the answer. In the GPT Tutor group that fell to 37%, and students increasingly asked for help and then attempted the problem themselves. The researchers describe the first pattern as using the model as a “crutch.”

This matters for the objection you should be forming, which is that the study used a 2023 model and the models are much better now. Model accuracy was tested and cleared as the cause. A more capable model does not remove the crutch. It produces a more correct, more fluent answer, which is a more tempting one to copy, so the mechanism plausibly gets stronger rather than weaker. That reading is mine rather than the authors’, but it follows from which mechanism they ruled out.

The realistic target is breaking even

Most coverage of this study stops at “guardrails fix it.” Look again at what the good arm actually achieved.

GPT Tutor did not beat the control group. It matched it. The best-performing setup in a carefully designed experiment produced the same exam result as having no AI at all, while feeling enormously more productive along the way.

So set expectations accordingly. Nobody has yet shown that AI makes you learn a thing faster. What the evidence supports is that AI used carelessly makes you learn it worse, and that effort deliberately added back gets you to roughly where you would have been anyway, with better company.

That is still worth having, because the alternative you are actually choosing between is not “AI or a good teacher.” It is AI or nothing, at 9pm, on a subject nobody is going to teach you.

Four ways to make it withhold

Every one of these reintroduces effort the tool would otherwise remove. That is the entire point, and it will feel worse than the version that does not work.

  • Ask for the question, not the answer. Open with something like: “I want to understand X. Do not explain it yet. Ask me what I already think is going on, then correct me.” You are forcing yourself to produce a first attempt, which is the step copying skips.
  • Explain it back before you move on. Type your own version, without scrolling up, then ask it to mark you and name specifically what you left out. Producing it from memory is the part that does the work. Recognizing a correct explanation is not.
  • Make it test you tomorrow, not today. Ask for six questions on what you covered, then close the chat and answer them the next day from memory. This is Pulling something out of your own head rather than reading it again. Decades of memory research find it produces far more durable learning than restudying, even though it feels harder and less productive at the time. The effect is well replicated for facts and concepts, and more mixed for complex problem-solving, so it is a strong default rather than a law. , and it is the single highest-value habit here.
  • Do one unaided. Whatever the skill is, do one instance of it with the chat closed before you decide you have learned it. The exam in the study was closed book for a reason: it is the only measurement that is not flattered.

Remember, if a study session felt smooth and quick, you probably did not learn much. The version that works feels like effort, because it is the effort that is doing it.

When none of this applies

Most of what you do with AI is not learning.

If you need the output and not the skill, take the answer. Formatting a table, drafting a note, summarizing a document you will never need to recall: there is no future exam, so there is nothing to protect. The question is whether you will need to do this unaided later. If the honest answer is no, use the crutch. That is what it is for.

The trap is only the middle case, where you tell yourself you are learning something while a machine does it beside you.

The takeaway

Practice performance and learning are different things, and AI moves the first one hard.

If you are trying to learn, make the tool withhold. Ask it to question you before it explains, say it back from memory, get tested a day later, and do one unaided. In the only trial that has measured this properly, the careful version broke even and the careless version came out 17% behind people who had no AI at all, and nobody in either group could feel which one they were in.


Sources: the trial is Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman, “Generative AI without guardrails can harm learning: Evidence from high school mathematics” (opens in a new tab) (PNAS, 25 June 2025, 122(26):e2422633122), a randomized controlled trial with nearly 1,000 students across around fifty 9th to 11th grade classes at one large high school in Turkey, over four 90-minute sessions covering roughly 15% of the math curriculum. Scope worth holding onto: high school mathematics is not adult knowledge work, one school is one school, and four sessions is short. The mechanism transfers more readily than the percentages do. The 51% model accuracy, the 67% and 37% superficial-message shares, and the finding that students did not perceive learning less are all from the paper. The argument that a more capable model would not fix this, and the four habits, are mine, reasoned from the mechanism the authors ruled in rather than measured by them. Retrieval practice is a long-standing finding in memory research rather than anything this paper tested, and its benefits are better established for facts and concepts than for complex problem-solving.

Related: AI Skill Erosion: How to Use It Without Getting Worse at Your Job is the same problem across a career rather than a single session, and its central finding, that trust in the tool predicts how hard you think, is what this trial demonstrates experimentally. How to Learn AI at Work When Nobody’s Training You is the version aimed at learning the tool itself.