---
title: "Claim Ledger: Use AI to Learn Without Getting Worse at It"
description: "Answer-giving AI left learners worse off once it was gone. Three setups built to keep the thinking with you, and a way to check what stuck."
author: "Harry Floyd"
publication: "The Durability Curve"
canonical: "https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/"
essay: "https://durabilitycurve.com/blog/learn-with-ai-without-getting-worse/"
published: "2026-09-24"
ledger_date: "2026-09-27"
claims: 70
removed_before_publication: 0
---

# Claim Ledger: Use AI to Learn Without Getting Worse at It

*Answer-giving AI left learners worse off once it was gone. Three setups built to keep the thinking with you, and a way to check what stuck.*

Essay: https://durabilitycurve.com/blog/learn-with-ai-without-getting-worse/  
Ledger (canonical, cite this): https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/  
Published 2026-09-24 · last verified 2026-09-27

## The evidence, row by row

Labels: VERIFIED = checked against the primary source itself by a checker who did not draft the piece, usually with the exact words, where they sit and the date read · EXECUTED = a run or observation done for the piece (code, a query, a count, or a check made live in an app); the claim is what it returned; run records are not published · CHECKED = checked against its source by a separate checker in the older ledger format, which recorded no quote, locator or date · REPORTED = not verified word for word against a primary source (secondary, not openable in full, or supported only in part); the essay words it accordingly · STRUCK = drafted, checked and removed before publication · EXCLUDED = considered and deliberately left out.

### B1. In a Turkish high-school trial, a plain ChatGPT-style assistant raised practice scores by 48% and a hint-only tutor by 127%

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B1
- Quote: "GPT Base and GPT Tutor would increase performance on the assisted practice sessions by 48% and 127%, respectively"
- Source: Bastani et al., PNAS 2025, Main Results; Table 1 col (1)
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B2. Once the AI was taken away, the plain-assistant group scored 17% lower on the exam than students who never had it

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B2
- Quote: "GPT Base diminished the average control student's performance on the unassisted exam by 17%."
- Source: Bastani et al., PNAS 2025, Main Results; Table 1 col (2)
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/

### B3. The analysis the authors registered in advance finds a smaller exam harm than the headline (−0.035, CI −0.057 to −0.012), still negative

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B3
- Quote: "GPT Base vs. Control ... −0.035 [−0.057, −0.012] 0.003"
- Source: Bastani et al., PNAS 2025, SI, SI Appendix A.6, Table A.2
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12232635/

### B4. The hint-only tutor removed the exam harm but produced no exam gain

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B4
- Quote: "this negative effect is essentially eradicated in the GPT Tutor arm, though we still do not observe a positive effect."
- Source: Bastani et al., PNAS 2025, Main Results; Introduction
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B5. The tutor's prompt carried the correct solution and told it not to give the whole solution away

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B5
- Quote: "GPT-4 is given a prompt including the solution to each problem (to mitigate hallucinations) as well as instructions to avoid giving away the entire solution"
- Source: Bastani et al., PNAS 2025, Footnote †; Experimental Design
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B6. Without the answer key, the plain assistant got the practice problems right only 51% of the time

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B6
- Quote: "GPT Base gives a correct answer only 51% of the time on average"
- Source: Bastani et al., PNAS 2025, "GPT Errors vs. Student Performance"; Fig. 2; SI C.1
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B7. Students used the plain assistant as a crutch, asking for and copying solutions

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B7
- Quote: "students often use GPT Base as a "crutch" by asking for and copying solutions"
- Source: Bastani et al., PNAS 2025, Introduction; Potential Mechanism
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B8. The trial covered nearly 1,000 students

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B8
- Quote: "We conducted four 90-min sessions for about fifty 9th, 10th, and 11th-grade classes, comprising nearly 1,000 students."
- Source: Bastani et al., PNAS 2025, Experimental Design
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### B9. The trial measured short-term outcomes only

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-B9
- Quote: "we focus on short-term outcomes due to limitations imposed by our partner school"
- Source: Bastani et al., PNAS 2025, Discussion, final paragraph
- URL: https://www.pnas.org/doi/10.1073/pnas.2422633122

### K1. In a Harvard physics trial, students' median post-test was 4.5 with the AI tutor against 3.5 in an active-learning class

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K1
- Quote: "higher median (M) post-score (M = 4.5, N = 142) compared to those in the in-class active learning group (M = 3.5, N = 174)"
- Source: Kestin et al., Scientific Reports 2025, Results § Learning gains
- URL: https://www.nature.com/articles/s41598-025-97652-6

### K2. Median learning gains with the AI tutor were more than double the class's

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K2
- Quote: "in the AI-tutored group were over double those for students in the in-class active learning group"
- Source: Kestin et al., Scientific Reports 2025, Results § Learning gains
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12179260/

### K3. The regression effect size was 0.63

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K3
- Quote: "While the linear regression suggests an effect size of 0.63, this is an underestimation due to ceiling effect"
- Source: Kestin et al., Scientific Reports 2025, Results § Linear regression model; Table S1
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12179260/

### K4. The Harvard tutor's instructions released one step at a time and kept replies brief

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K4
- Quote: "Only give away ONE STEP AT A TIME, DO NOT give away the full solution in a single message" / "Keep responses BRIEF (a few sentences or less) but helpful."
- Source: Kestin et al., Supplementary Material 1, Supplement, system prompt
- URL: https://www.nature.com/articles/s41598-025-97652-6

### K5. The Harvard team loaded each question with its step-by-step worked answer

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K5
- Quote: "we enriched our prompts with comprehensive, step-by-step answers"
- Source: Kestin et al., Scientific Reports 2025, Discussion § Designing successful student-AI interactions
- URL: https://www.nature.com/articles/s41598-025-97652-6

### K6. 194 students took part

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K6
- Quote: "Of the 233 enrolled students, 194 were eligible for inclusion in the study."
- Source: Kestin et al., Scientific Reports 2025, Methods § Study population
- URL: https://www.nature.com/articles/s41598-025-97652-6

### K9. Each student learned one lesson with the AI tutor and another in the active-learning class (crossover)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K9
- Quote: "Given that our experiment is a crossover design in which each student experiences both conditions"
- Source: Kestin et al., Scientific Reports 2025, Results (regression controls); Methods § Study design
- URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12179260/

### A1. In an Anthropic trial, 52 developers learned a new Python library, half with an AI assistant

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A1
- Quote: "In our main study, 52 participants completed the task, 26 for each of the control and treatment groups."
- Source: Shen & Tamkin, arXiv 2601.20245, §5.2.1
- URL: https://arxiv.org/abs/2601.20245

### A2. The AI group scored 4.15 points lower on a 27-point quiz, about two grade points

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A2
- Quote: "There is a 4.15 point difference between the means ... For a 27-point quiz, this translates into a 17% score difference or 2 grade points."
- Source: Shen & Tamkin, arXiv 2601.20245, §5.2.2
- URL: https://arxiv.org/html/2601.20245v2

### A3. Nobody could use AI during the quiz

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A3
- Quote: "All participants were not allowed to use AI in the comprehension check."
- Source: Shen & Tamkin, arXiv 2601.20245, Figure 4 caption
- URL: https://arxiv.org/html/2601.20245v2

### A4. Using AI did not make them significantly faster

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A4
- Quote: "While using AI to complete our coding task did not significantly improve task completion time, the level of skill formation gained ... is significantly reduced."
- Source: Shen & Tamkin, arXiv 2601.20245, §1 / §5.2.2
- URL: https://arxiv.org/html/2601.20245v2

### A5. High-scoring ways of using the AI averaged 65% to 86% on the quiz; low-scoring ways 24% to 39% (post-hoc, small groups)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A5
- Quote: "high-scoring interaction patterns (65%-86% quiz score) vs low-scoring interaction patterns (24%-39% quiz score)"
- Source: Shen & Tamkin, arXiv 2601.20245, §6; Figure 11
- URL: https://arxiv.org/abs/2601.20245

### A6. Asking the AI conceptual questions was a high-scoring pattern, and the fastest of them

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A6
- Quote: "On average, this mode was the fastest among high-scoring patterns and second fastest overall after the AI Delegation mode."
- Source: Shen & Tamkin, arXiv 2601.20245, §6; Figure 11 (Conceptual Inquiry)
- URL: https://arxiv.org/abs/2601.20245

### N1. In a World Bank trial in Nigeria, students had twelve 90-minute sessions guided by teachers

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N1
- Quote: "Those assigned to the intervention attended twelve 90-minute sessions in computer labs, engaging in curriculum-aligned activities guided by teachers."
- Source: De Simone et al., World Bank PRWP 11125, §1, p. 2
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### N2. English scores rose 0.238 standard deviations, which the authors equate to 1.5 years of ordinary schooling

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N2
- Quote: "Our ITT effect of 0.238 standard deviation in English is equivalent to increasing 1.5 years of 'business-as-usual' schooling in Nigeria"
- Source: De Simone et al., World Bank PRWP 11125, §4.1, p. 18
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### N3. In the sample week's starting prompt, the chatbot answered the student's question, then set exercises on which it was told to hint after a wrong reply and give the answer only if the reply was still wrong

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N3
- Quote: "provide encouraging words and provide a hint" … "if my reply is still incorrect"
- Source: De Simone et al., World Bank PRWP 11125, Figure 11, PDF p. 51 (Week 7 starting prompt, read from rendered image)
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### N4. The control group got no programme at all, so extra time and attention are not separated from the AI

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N4
- Quote: "the control group, which did not receive any intervention but continued their regular learning in the classroom"
- Source: De Simone et al., World Bank PRWP 11125, §2.2, p. 9
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### N5. The authors credit the whole package (prompts plus teacher guidance), not the chatbot alone

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N5
- Quote: "we interpret that the intervention as a whole -which includes the interaction with the LLM and teacher guidance with specific prompts- is driving the results."
- Source: De Simone et al., World Bank PRWP 11125, §4.3, p. 21
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### N6. Nigeria's third-term exam (effect 0.206 SD) came after the intervention on content not limited to it; it ran the day after the last session, so it is not a test of whether the learning lasted

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-N6
- Quote: "the third term exam score, with an effect size of 0.206 standard deviation (SE = 0.067), although this exam was not limited to the intervention's specific content"
- Source: De Simone et al., World Bank PRWP 11125, §4.1; timing from Appendix timeline (pilot sessions end 7/11/24, third-term exam 7/12/24)
- URL: https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf

### O1. OpenAI's own trial ran over 300 college students through study sessions of a nominal 40 minutes

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-O1
- Quote: "we ran a randomized study with over 300 college students preparing for neuroscience and microeconomics exams" / "the nominal 40 minute sessions"
- Source: OpenAI, "New tools for understanding AI and learning outcomes" (2026-03-04), "Origins and early research"; "Study design"
- URL: https://openai.com/index/understanding-ai-and-learning-outcomes/

### O2. Study mode scored roughly 15% higher in microeconomics

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-O2
- Quote: "roughly a 15% higher score relative"
- Source: OpenAI (2026-03-04), "Findings", bullet 2
- URL: https://openai.com/index/understanding-ai-and-learning-outcomes/

### O3. In neuroscience, study mode was not distinguishable from studying with ordinary online resources

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-O3
- Quote: "results were not distinguishable from students studying with traditional online resources"
- Source: OpenAI (2026-03-04), "Findings", bullet 1
- URL: https://openai.com/index/understanding-ai-and-learning-outcomes/

### O4. The comparison group used ordinary online resources, not ordinary ChatGPT

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-O4
- Quote: "a control group studied using traditional online resources such as Google Search and YouTube, with AI generated overview features disabled"
- Source: OpenAI (2026-03-04), "Study design", para 1
- URL: https://openai.com/index/understanding-ai-and-learning-outcomes/

### O5. OpenAI calls the results early; analysis is still underway

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-O5
- Quote: "While analysis is still underway, early results give us confidence"
- Source: OpenAI (2026-03-04), "Origins and early research"
- URL: https://openai.com/index/understanding-ai-and-learning-outcomes/

### G1. Google's own trial of Gemini Guided Learning in Sierra Leone raised maths scores 0.258 standard deviations (CI 0.027 to 0.488)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G1
- Quote: "yielding a gain of +0.258 standard deviations across the Guided Learning classrooms (intent to treat; 95% confidence interval [0.027, 0.488]"
- Source: LearnLM Team, Google & Fab AI, tech report (May 2026), p. 3; Table C.4
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### G2. It enrolled 1,763 students in 48 grade 7 and 8 classrooms

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G2
- Quote: "The trial enrolled 𝑁= 1763 students aged 13 or older in 48 grades 7 and 8 classrooms."
- Source: LearnLM Team tech report, p. 2
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### G3. Students shared devices in pairs, in lessons teachers ran

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G3
- Quote: "Students accessed the Gemini app on tablets or desktop computers, sharing at a 2:1 student-to-device ratio."
- Source: LearnLM Team tech report, p. 2
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### G4. The gain came from grade 8; the grade 7 estimate was slightly negative (−0.078; grade 8 interaction 0.429)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G4
- Quote: "Arm: Treatment –0.078* (0.037)"; "Arm: Treatment × Grade: Grade 8 0.429** (0.142)"
- Source: LearnLM Team tech report, Table C.14, p. 28
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### G5. Gemini gave a direct solution in 2.1% of its messages (Gemini-classified)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G5
- Quote: "far more frequently than providing direct solutions (2.1%)"
- Source: LearnLM Team tech report, p. 4
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### G6. The report has not been peer reviewed

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-G6
- Quote: "Please note works submitted as a preprint have not undergone a peer review process."
- Source: LearnLM Team tech report, p. 8
- URL: https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf

### D1. OpenAI's own help page says Study mode may still give a direct answer

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D1
- Quote: "there may be times when it gives a direct answer"
- Source: OpenAI Help Center, "Using study mode in ChatGPT", Article body
- URL: https://help.openai.com/en/articles/11780217-using-study-mode-in-chatgpt

### D2. Study mode is not available inside ChatGPT Projects

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D2
- Quote: "The Study option is not available in Temporary Chats, GPTs, or Projects."
- Source: OpenAI Help Center, Article body
- URL: https://help.openai.com/en/articles/11780217-using-study-mode-in-chatgpt

### D3. A ChatGPT scheduled task made inside a project cannot read that project's files

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D3
- Quote: "If you create a task in a project, it cannot access uploaded files or files stored in that project."
- Source: OpenAI Help Center, Tasks, Article body
- URL: https://help.openai.com/en/articles/10291617

### D4. Claude's free plan allows up to five Projects

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D4
- Quote: "Free users can create a maximum of five projects."
- Source: Claude Help Center, Article body
- URL: https://support.claude.com/en/articles/9519177

### D5. Gemini Guided Learning sits under Add files, then More tools

- Status: VERIFIED (read 2026-09-27)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D5
- Quote: "In the text box, click Add Files." / "At the bottom, click More tools" … "Guided Learning"
- Source: Google Gemini Help, "Use Guided Learning", steps 2–3
- URL: https://support.google.com/gemini/answer/16448384

### X1. Study appears as a composer mode in the ChatGPT web app (Plus)

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X1
- Quote: The author's screenshot (not published): composer shows "Study" mode chip
- Source: First-party: ChatGPT web app (Plus), the author's run

### X2. In 8 runs of the final Level 2 prompt (7 across Claude, Gemini and GPT, 1 in the ChatGPT app), the tutor withheld the answer on a direct demand and, after the second wrong try, gave the answer (some models per step) and named the missed step; once (GPT, code) it described the fix in words while withholding the code

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X2
- Quote: See run record
- Source: First-party: python3 run.py, run_gemini.py, run_codex.py + the author's ChatGPT run

### X3. In a simulated project (instructions plus file given to the model directly), the Level 3 setup stuck to the file, said when the file did not cover a question, and scored a no-feedback quiz correctly on all three model families; the Claude and Gemini apps themselves were not tested

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X3
- Quote: See run record
- Source: First-party: python3 l3.py, l3x.py

### X4. In the ChatGPT app, our first Level 3 instructions answered a "just tell me" question at once from general knowledge; the revised instructions held it back and answered from the file

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X4
- Quote: See run record
- Source: First-party: ChatGPT web (Plus), Project "Cold Check Test", the author's run (v2) + Claude via Claude in Chrome (v3)

### X5. In the ChatGPT app, the cold-check prompt held all feedback to the end and scored a 3-of-6 quiz correctly against the file

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X5
- Quote: See run record
- Source: First-party: ChatGPT web (Plus), Project "Cold Check Test", driven by Claude via Claude in Chrome

### D6. In Claude.ai, the old Styles menu is gone; Anthropic now offers an official "learn" skill (Customize → Skills → Discover → learn → Add) whose stated goal is to help the learner answer it themselves, and which does not trigger on coding tasks or factual lookups

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D6
- Quote: "The goal is not to answer the learner's question but to help them be able to answer it themselves" / "Don't trigger for: Tasks: coding, writing, calculation, translation, factual lookup"
- Source: Anthropic "learn" skill (updated Sep 14), SKILL.md + Overview, read in the live Claude.ai app, Skill Overview + Contents › SKILL.md "Learning Mode"
- URL: https://claude.ai/new#customize/skills/id/learn

### A7. They were Python programmers recruited through a crowd-work platform; 29 of the 52 had seven or more years of coding experience

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A7
- Quote: Paper Table 1 (experience bands; 7+ years = 29 of 52) / "the small group size of the 1-3 year participant group (n=4)"
- Source: Shen & Tamkin, arXiv 2601.20245, Table 1; §5.2.2
- URL: https://arxiv.org/abs/2601.20245

### S2. Only two of the seven studies went through journal peer review (PNAS and Scientific Reports); the others are two preprints, a working paper, a vendor tech report and a vendor blog post

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-S2
- Quote: PNAS 122(26); Scientific Reports 15, 17458; arXiv 2601.20245 and arXiv 2409.09047 (preprints); World Bank Policy Research Working Paper 11125; Google tech report "have not undergone a peer review process"; OpenAI "While analysis is still underway"
- Source: The seven studies' own publication records

### D7. Microsoft Copilot has a "Study and learn" mode, chosen from the Quick response menu under the prompt

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D7
- Quote: "Press (or tap) 'Quick response' under your prompt and select an appropriate mode"
- Source: Microsoft Support, "Conversation modes in Microsoft Copilot", Article body
- URL: https://support.microsoft.com/en-us/microsoft-copilot/conversation-modes-in-microsoft-copilot

### D8. Claude Code has a Learning output style, switched on with /output-style learning

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D8
- Quote: "/output-style learning" (Learning style: Claude "stops and waits" at TODO(human) markers)
- Source: Claude Code docs, Output styles, Built-in styles section
- URL: https://code.claude.com/docs/en/output-styles

### K7. The Harvard tutor's instructions allowed it to give the answer if the student demanded it

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K7
- Quote: "DO NOT not tell them the answer UNLESS they demand you to give them the answer"
- Source: Kestin et al., Supplementary Material 1, Supplement, system prompt
- URL: https://www.nature.com/articles/s41598-025-97652-6

### K8. The Harvard tutor's instructions asked students to try first

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-K8
- Quote: "encourage them to give it a try first"
- Source: Kestin et al., Supplementary Material 1, Supplement, system prompt
- URL: https://www.nature.com/articles/s41598-025-97652-6

### A8. The highest-scoring pattern (2 people) had the AI generate code and then asked it questions to understand it, averaging 86%; asking conceptual questions averaged 65%; the lowest patterns, which delegated or relied on it, averaged 24% to 39%

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A8
- Quote: "high-scoring interaction patterns (65%-86% quiz score) vs low-scoring interaction patterns (24%-39% quiz score)"
- Source: Shen & Tamkin, arXiv 2601.20245, §1.1; §6; Figure 11 (Generation-Then-Comprehension n=2 86%; Conceptual Inquiry n=7 65%)
- URL: https://arxiv.org/abs/2601.20245

### A10. The low-scoring patterns were delegation, progressive reliance and iterative AI debugging (relying on the AI to debug or verify their code)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A10
- Quote: "Iterative AI Debugging (n=4): Participants in this group relied on AI to debug or verify their code."
- Source: Shen & Tamkin, arXiv 2601.20245v2, §6, interaction patterns
- URL: https://arxiv.org/html/2601.20245v2

### L1. In pre-registered lab experiments on learning to code, giving students an AI chatbot had no overall effect on learning

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-L1
- Quote: "we find no effect of LLMs on overall learning outcomes"
- Source: Lehmann, Cornelius & Sting, arXiv 2409.09047v2, Abstract; Tables 5, 7
- URL: https://arxiv.org/abs/2409.09047

### L2. Students who used it to substitute for their own work (e.g. generating solutions to exercises) covered more topics but understood less; students who used it to complement their work (e.g. asking for explanations) understood more (exploratory, not randomised)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-L2
- Quote: "Students who substitute some of their learning activities with LLMs (e.g., by generating solutions to exercises) increase the volume of topics" / "Students who complement their learning activities with LLMs (e.g., by asking for explanations) do not increase topic volume but do increase their understanding."
- Source: Lehmann et al., arXiv 2409.09047v2, Abstract; §9
- URL: https://arxiv.org/abs/2409.09047

### L3. 42% of messages asking for a solution were sent without a single attempt

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-L3
- Quote: "42% of messages asking for a solution were sent without a single attempt"
- Source: Lehmann et al., arXiv 2409.09047v2, §9.1 (Study 3)
- URL: https://arxiv.org/abs/2409.09047

### L4. The chatbot raised how much students felt they had learned by more than their actual learning explains

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-L4
- Quote: "LLMs increase perceived learning by more than can be explained by actual differences in learning"
- Source: Lehmann et al., arXiv 2409.09047v2, §9.3
- URL: https://arxiv.org/abs/2409.09047

### D9. Claude Skills need code execution switched on (Settings > Capabilities > Code execution and file creation on Free, Pro and Max)

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D9
- Quote: "This feature requires code execution to be enabled."
- Source: Claude Help Center, "Use Skills in Claude", Article body
- URL: https://support.claude.com/en/articles/12512180-use-skills-in-claude

### D10. ChatGPT's Study mode is chosen by typing @study or from the + menu, and is on every plan

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D10
- Quote: "Available across ChatGPT plans globally on web, iOS, and Android"
- Source: OpenAI Help Center, "Using study mode in ChatGPT", Article body
- URL: https://help.openai.com/en/articles/11780217-using-study-mode-in-chatgpt

### X6. The cold check with a topic line kept to the named topics and scored correctly in the ChatGPT app (3 of 6, the planted misses exactly), and kept to topic on 17 of 18 questions across Claude, Gemini and GPT

- Status: EXECUTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-X6
- Quote: See run record
- Source: First-party: ChatGPT web (Plus) via Claude in Chrome + l3_cc2.py, l3x_cc2.py

### P1. Tutor prompts built to make students work before giving answers were published by Ethan and Lilach Mollick in 2023, and the Nigeria programme borrowed from them

- Status: VERIFIED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-P1
- Quote: "Some of the prompt structures were derived from Mollick and Mollick (2023a)"
- Source: De Simone et al., World Bank PRWP 11125, fn 3; Mollick & Mollick, "Assigning AI" (SSRN 4475995), §2.1, pp. 7 to 8 (main text)
- URL: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4475995

### A9. The no-AI group averaged about 65% on the quiz

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-A9
- Quote: Figure 6 plots about 65.4% for the control group (read from marker positions); the Anthropic blog gives 67%
- Source: Shen & Tamkin, arXiv 2601.20245, Figure 6; Anthropic blog, Figure 6
- URL: https://arxiv.org/abs/2601.20245

### D11. ChatGPT project: New project in the sidebar; instructions via the project's settings

- Status: VERIFIED (read 2026-09-27)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D11
- Quote: "Select New project in the sidebar." / "Select the more options menu (•••), then select Project settings to add instructions for the project."
- Source: OpenAI Help Center, Projects in ChatGPT, "Create a project"; "Add project instructions"
- URL: https://help.openai.com/en/articles/10169521

### D12. Claude project: Projects, then New Project, with instructions and knowledge files

- Status: VERIFIED (read 2026-09-27)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D12
- Quote: "click “Projects,”" … "Click "+ New Project" in the upper right corner." / "Anything you upload to this space will be used across all of your chats within that project." / "Click on "Set project instructions.""
- Source: Claude Help Center, projects article, "How to create a project"; "Add content to project knowledge"; "Add project instructions"
- URL: https://support.claude.com/en/articles/9519177

### D13. Gemini: a Gem (Gems, then New Gem) with files added under Knowledge

- Status: REPORTED (read 2026-09-24)
- Row link: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-D13
- Quote: Help-page path: Gems → New Gem → Knowledge → Add files
- Source: Google Gemini Help, Article body
- URL: https://support.google.com/gemini/answer/15146780

## Cite

- A claim: name the original source (and its locator) first, then the row it was checked in, e.g. "<source>, <locator>. Checked in: Claim Ledger, "Use AI to Learn Without Getting Worse at It", The Durability Curve, row <id>, https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/#c-<id>."
- The essay: Floyd, Harry (2026). Use AI to Learn Without Getting Worse at It. The Durability Curve. https://durabilitycurve.com/blog/learn-with-ai-without-getting-worse/
- This ledger: Floyd, Harry (2026). Claim Ledger: Use AI to Learn Without Getting Worse at It. The Durability Curve. https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse/
- Downloads: https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse.csv · https://durabilitycurve.com/claims/learn-with-ai-without-getting-worse.json

Licence: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Reuse it, quote it or train on it, with credit to The Durability Curve and a link. Open to AI (robots.txt: search=yes, ai-input=yes, ai-train=yes).
