What’s trending in AI on 11 October 2026: The AI headline people keep sharing this weekend is not a launch or a lawsuit. It is a psychology result. A study co-authored by UC Berkeley’s Brian Christian with researchers from Carnegie Mellon, MIT, Oxford and UCLA, presented this month at the Conference on Language Modeling (COLM) and publicized by Berkeley on 9 October, found that about 10 to 15 minutes of working with an AI assistant made people less accurate and quicker to give up once the assistant was taken away. Across three randomized trials involving 1,222 people, the pattern held for both maths and reading. The most useful detail for managers sits in the fine print: people who used AI for hints held up; people who asked it for answers did not. Below: what the study did, what it does and does not prove, why it matters for anyone who relies on humans to check AI, and a traffic-light rubric for deciding when your team should take the answer and when it should take the hint.
Key takeaways
- Short exposure, measurable effect. In all three experiments, people who had an AI assistant solved fewer problems on their own once it was removed. In the first trial, 57% versus 73% for the group that never had AI.
- Giving up is part of the story. In the first and third experiments, former AI users skipped noticeably more problems; in the second the gap pointed the same way but was not statistically significant.
- How you use AI matters most. About 61% of AI-group participants mostly asked for direct answers. Their unaided score fell 10 points from their own pretest; hint-seekers did not decline.
- It echoes earlier work. An Anthropic trial with 52 developers and a Microsoft and Carnegie Mellon survey of 319 knowledge workers point in the same direction.
- It is a security issue too. If people checking AI work lose the habit of struggling with hard problems, “human in the loop” becomes a rubber stamp.
1. What the study actually did
The paper, “AI Assistance Reduces Persistence and Hurts Independent Performance,” is by Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel Bakker and Rachit Dubey. A first draft appeared on arXiv in April; the peer-reviewed version was presented at COLM in early October, and Berkeley News covered it on 9 October. Christian is a research fellow at Berkeley’s Center for Human-Compatible AI and the author of The Alignment Problem.
The design is simple, which is why the result travels. Online participants in the US, recruited through Prolific, were randomly split. One group worked alone. The other had an AI assistant in a sidebar (GPT-5, according to the paper) that they could use however they liked, including asking for the answer outright. After a block of practice problems, the assistant disappeared and everyone faced three final problems unaided. Wrong answers carried no penalty, so the researchers treated skipping a problem as a sign of giving up. Sessions took roughly 13 to 15 minutes.
- Experiment 1 (fractions): 354 recruited, 307 analyzed. Twelve practice problems, then three test problems without AI.
- Experiment 2 (fractions, replication): 667 recruited, 585 analyzed. Added a pretest so the team could compare each person with themselves, and asked AI users how they had used it.
- Experiment 3 (SAT-style reading): 201 recruited, 168 analyzed. Answers given in under five seconds counted as skips.
While the assistant was available, the AI group did better, as you would expect. The interesting part is what happened next.
Christian summed it up by saying the tools often end up helping people “in ways that are kind of unhelpful.” The effect sizes are modest (the paper reports Cohen’s d of roughly 0.2 to 0.4 for solve rates), but they appeared after a quarter of an hour, on basic tasks, with a general-purpose assistant. That is a lot closer to an ordinary workday than most lab studies.
2. The detail that matters: answers versus hints
Experiment 2 asked AI-group participants how they had used the assistant. Of 308 people, 61% said they mainly asked for direct answers, 27% used it for hints or clarification, and 12% hardly used it at all. All groups had scored about the same on the pretest, so the differences afterward are worth looking at, with one caution: people chose their own usage style, so this part shows a correlation rather than proof of cause.
| How people used the AI (Exp. 2) | Share of AI group | Unaided test solve rate | Test skip rate | Change vs. own pretest |
|---|---|---|---|---|
| Mostly asked for direct answers | 61% | 65% | 13% | Solve rate down 10 points; skips up |
| Mostly asked for hints or clarification | 27% | 76% | 5% | No decline |
| Barely used it | 12% | 89% | 7% | No decline |
| No AI at all (control group) | n/a | 77% | 7% | Roughly flat (+1 point) |
The answer-seekers are the story. They were the majority, they were the only group that got worse than their own starting point, and they skipped more. Hint users finished roughly level with people who never had AI and skipped the fewest problems of any group. The authors’ recommendation follows from this: assistants should be built to protect people’s long-term ability as well as finish the task, and a good collaborator, they write, should “know when not to help.”
3. Why a few minutes could change behaviour
The researchers offer two explanations. The first is an effort reset: once an instant answer is one click away, slow, effortful work starts to feel unreasonably expensive, so when the shortcut disappears, quitting feels like the sensible choice. The second is lost productive struggle: wrestling with a problem is how people build skill and learn what they are capable of, and outsourcing that step removes both.
Neither idea is new to teachers, but the study is notable for showing the effect causally, quickly and with the kind of assistant millions of people use at work. It also lands as chat products are designed to be more instantly helpful. When assistants build the whole answer for you, as we described in our look at GPT-6 turning answers into apps, the gap between asking and having gets even smaller.
4. Not a one-off: what earlier studies found
Two earlier pieces of research point in the same direction, and both came from companies that sell AI tools.
Anthropic’s coding trial (early 2026). Anthropic researchers randomly assigned 52 mostly junior engineers to learn Trio, an unfamiliar Python async library, with or without AI help. The AI group finished about two minutes faster, a difference that was not statistically significant, but averaged 50% on a follow-up quiz versus 67% for the manual group, with the biggest gap on debugging. Developers who scored well asked follow-up questions and requested explanations; those who scored lowest handed the whole job to the AI, InfoQ reported.
Microsoft Research and Carnegie Mellon (CHI 2025). A survey of 319 knowledge workers who use generative AI weekly found that the more people trusted the AI on a task, the less critical thinking they reported applying, while people confident in their own skills checked more. Because it was self-reported, it shows association rather than cause, which is exactly the gap the Berkeley-led trials help fill.
Put together, the message is consistent: AI raises output in the moment, and whether it builds or erodes skill depends on whether people stay mentally engaged. That fits the pattern in our analysis of why individual AI gains don’t always reach the team, and the warning in our piece on the verification gap: generating is cheap, checking is the scarce skill.
5. Why this is a security and oversight problem
Most organizations’ answer to AI risk includes a person. A human approves the agent’s payment, reviews the AI-written code, checks the generated contract clause or signs off on the triaged security alert. That safeguard assumes the reviewer will dig in when something looks plausible but wrong. The persistence study suggests that a reviewer who has spent the morning accepting AI answers may be measurably less inclined to dig in at exactly that moment.
The timing is pointed. On 10 October, Microsoft CEO Satya Nadella published an essay arguing that companies should treat frontier models as insider risks, keeping identity, least privilege, logging and an authorized human’s ability to stop a model mid-task (his “emergency brake”) with the deploying organization rather than the model vendor, according to FourWeekMBA’s summary. Brakes only work if the person holding them is alert. The same week brought fresh reminders of what agents do when no one is watching closely, from an AI agent filing a fake police tip to agents getting their own Workspace accounts as coworkers.
Three practical implications for security and risk teams:
- Approval fatigue is now a measurable risk. The habit that makes people click “Allow Always” on agent permissions, covered in our agent permissions explainer, is close kin to the habit of accepting the answer.
- Design reviews so the reviewer has to think. Ask reviewers to state what they checked, or to predict an outcome before seeing the AI’s version, rather than clicking approve.
- Keep unassisted skills alive in critical roles. Incident responders, auditors and approvers of payments should regularly practise without AI, the same way pilots still practise manual landings.
6. What the study does not show
It would be easy to oversell this, and the authors themselves warn against concluding that AI lowers persistence on every task. Keep these limits in mind:
- Short tasks, short sessions. Fractions and SAT reading passages over about 15 minutes are not a quarterly forecast or a code review. Whether effects build up over months is the key open question, and the paper does not measure it.
- Small test sets. Each result rests on three final problems, and one of the six headline comparisons (the Experiment 2 skip rate) was not statistically significant.
- One model, one population. All AI use involved GPT-5, and participants were paid US online workers.
- Usage styles were self-chosen. The answer-versus-hint split is correlational; people who prefer hints may simply be more persistent to begin with.
None of this cancels the finding. It means the right response is to change how AI is used, not to ban it.
7. The traffic-light rubric: when to take the answer and when to take the hint
Not every task deserves a struggle. Formatting a table or drafting a routine email is fine to hand over. The question is whether the skill involved is one your team needs to keep. Use this rubric to sort work into three lanes, then set up the tools to match.
Then put the rubric into practice with five moves, in this order:
- Label your top ten tasks. In a 30-minute team session, list the ten things people most often ask AI to do and assign each a colour. Disagreements are useful; they reveal which skills people think they can afford to lose.
- Ship a “hint-first” prompt for amber work. Add a saved instruction or custom assistant that tells the model to ask what the person has tried, offer the next step rather than the full solution and explain its reasoning. Several assistants now offer study or tutor modes; turn them on where they exist.
- Add a “your view first” field to red-lane approvals. Before an approver sees the AI recommendation, ask them to record their own call in a sentence. It takes seconds and stops the review collapsing into agreement.
- Schedule unplugged practice for critical roles. Once a month, run one incident drill, code review or reconciliation without AI and compare the result with the AI-assisted version.
- Measure skill, not just speed. Alongside time saved, track error rates found in review and how often reviewers send AI work back. A falling send-back rate with steady error rates is a warning sign, not a win.
Put the rubric in your AI acceptable-use policy too. If staff are already using unapproved chatbots, as we found in our look at shadow AI, a policy that only says what not to paste misses half the risk. For the people side of rollout, our guide to workplace AI adoption covers training and change management.
8. What to watch next
- Long-term studies. The authors flag cumulative effects as the big open question. Expect follow-ups over weeks and months, and in real workplaces rather than online tasks.
- Product defaults. If more evidence points the same way, pressure will grow on AI makers to make tutor-style help a default in education and training products, not a buried setting.
- Oversight rules for agents. Regulators and frameworks that rely on human review will need to say what makes oversight meaningful. Nadella’s emergency-brake argument is one marker of where enterprise practice is heading.
- Education policy. Fractions and reading are foundational skills. Schools adopting AI tutors will be asked whether their tools give hints or answers.
Frequently asked questions
What did the Berkeley AI persistence study find?
Across three randomized trials with 1,222 recruited participants, people who used an AI assistant for about 10 to 15 minutes solved fewer problems and often skipped more once the AI was removed, compared with people who never had it. The effect appeared on both fraction problems and SAT-style reading questions.
Who conducted the study and where was it published?
The paper, “AI Assistance Reduces Persistence and Hurts Independent Performance,” is by Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel Bakker and Rachit Dubey, with researchers from UC Berkeley, Carnegie Mellon, MIT, Oxford and UCLA. It is on arXiv (2604.04721) and was presented at COLM in October 2026.
Does using AI make you worse at your job?
Not necessarily. In the study, people who used AI for hints or explanations did not decline from their own pretest, while people who mostly asked for direct answers did. The risk comes from handing over the thinking on skills you still need, not from using AI at all.
Which AI model was used in the study?
According to the paper, all AI assistance came from GPT-5 in a sidebar next to the task. Participants were US-based online workers recruited through Prolific. Whether results differ with other models or populations has not yet been tested.
Why does AI reliance matter for cybersecurity and human oversight?
Many AI safeguards depend on a person reviewing outputs, approving agent actions or stopping a model. If AI-assisted work makes reviewers less inclined to persist on hard checks, that human-in-the-loop control weakens, so reviews should require independent judgment and critical staff should practise without AI.
How can teams use AI without losing skills?
Sort tasks by the skill they use. Let AI do routine work, use hint-first prompts for skills people need to keep, and require people to form their own view before seeing AI recommendations on high-stakes decisions. Track review quality, not just time saved.
Bottom line: AI made people faster and more accurate while it was there, and less persistent once it was gone. That is not a reason to switch it off. It is a reason to decide, task by task, which skills you are willing to rent and which you need to own, and to set up your tools so the default for the second group is a hint, not an answer.
Sources
- Berkeley News: Using AI for just 10 minutes erodes your ability to persist at hard things (9 Oct 2026)
- Liu, Christian, Dumbalska, Bakker, Dubey: AI Assistance Reduces Persistence and Hurts Independent Performance (arXiv 2604.04721)
- EurekAlert: Using AI for just 10 minutes erodes your ability to persist at hard things
- The Decoder: Ten minutes of using AI as an answer machine can erode problem-solving skills
- InfoQ: Anthropic study on AI coding assistance and skill formation (Feb 2026)
- The Register: Microsoft and Carnegie Mellon study on generative AI and critical thinking (Feb 2025)
- FourWeekMBA: Nadella says treat frontier AI models like insider risks (10 Oct 2026)
