A tool can help someone finish a problem. The next problem tells you more about what the person carried forward.
What changed
A revised preprint from Liu and colleagues reports three randomized online experiments with 1,222 recruited participants. They worked on fractions or reading-comprehension questions in sessions lasting roughly 10 to 15 minutes. Some participants could use GPT-5 in a sidebar during the assisted phase. Then the sidebar disappeared before three final problems.
The result repeated across all three experiments: participants in the AI groups solved fewer unsupported problems than control participants who had worked without AI. The first experiment also found more skipping, although the authors identified a possible skill-selection confound. The second experiment addressed that issue with a pretest and a comparable sidebar transition for the control group. Its later solve-rate gap remained. Skipping was directionally higher after AI use, but the aggregate difference was not statistically significant. The reading-comprehension experiment found both a lower solve rate and a higher skip rate after assistance ended.
One detail deserves careful handling. Within the second experiment's AI group, people who said they used the assistant for direct answers later solved less and skipped more than people who used it for hints. Participants chose how to use the tool. That comparison cannot show that direct answers caused the difference.
Why it matters
This makes the evaluation window important. Completion rate, answer quality, and time saved describe the assisted session. They leave open whether the person retained the method, transferred it to a new problem, or kept working through difficulty.
A learning product may look effective while the assistant is present and tell a different story on the next unsupported attempt.
But the study supports that warning only within a narrow setting. It does not establish durable deskilling or a universal effect of AI assistance. The tasks were short and controlled, removal was abrupt, one assistant setup was used, and skipping was a study-specific proxy for persistence. The paper is a revised preprint, and its results were not independently reproduced for this card.
Watch next
Longer studies in natural settings could test whether the pattern survives ordinary use. Randomized comparisons between direct solutions and scaffolded help could clarify whether assistance style changes later performance. Follow-up measures taken days or weeks later could separate a brief session effect from changes in retention or transfer.
So product teams have a bounded check available now: add a fresh problem after assistance is removed and record what happens. That result will not settle the wider question. It will show whether the current evaluation ends before the most revealing attempt.