What Makes Revision Actually Stick – High school students can look adequately prepared on everyday math practice right up until a proctored exam exposes the gap. A June 2026 working paper, “Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build,” documented exactly this pattern: after ChatGPT’s release, students spent 31% less time on math word problems, while their supervised exam accuracy on those same problems fell from roughly 80% to roughly 60%. Performance on graphing problems—far harder to delegate to AI—showed no comparable decline. The damage didn’t register in daily practice.

It only surfaced when conditions demanded genuinely independent work.

For IB Maths candidates, the gap between feeling ready and actually being ready follows the same logic. It opens whenever revision sidesteps what the exam actually demands: working through problems without access to the answer, then evaluating completed reasoning against examiner marking criteria under timed conditions. Most students have no deliberate framework for deciding whether a revision resource generates that kind of feedback or merely a reassuring sense of progress. Judging resources by the quality of feedback they produce—rather than by their volume or visual polish—is the practical skill that separates revision that improves exam performance from revision that just consumes time.

When Better Scores Mask Worse Learning

Surface performance indicators holding steady while actual exam-condition competence erodes is a structurally predictable outcome. When the effort of independently working through a problem gets bypassed—however that happens—daily scores can remain positive precisely because the conditions that require genuine unaided performance haven’t been triggered yet. There’s no warning built into the system.

The working paper “Faster Completion, Less Learning” establishes that the gap tracks the mechanism of bypassing rather than a general skill decline. Accuracy fell specifically on word problems, which are susceptible to AI outsourcing, and held on graphing problems, which are not. That divergence points to effortful independent engagement as a key factor in whether practice transfers to supervised conditions.

A convergent pattern appears in CEPR Discussion Paper DP21577, “The Generative AI Learning Penalty: Evidence from a Large-Scale Longitudinal Study,” a roughly 30-month study of 26,811 students in grades 7–12 using a staggered adoption difference-in-differences design. Among students who began using generative AI independently, homework scores rose 18% while time spent on homework fell. Monthly exam results dropped 20% after five months, and high-stakes entrance exam scores fell by 18% and 24% respectively. The largest penalties concentrated among usage patterns consistent with homework outsourcing—very short completion times paired with high homework scores. The authors caution against directly extrapolating effect sizes to other settings, and the design’s conclusions depend on standard parallel-trends assumptions. Both studies converge on the same mechanism. Bypassing independent submission of one’s own working yields weaker exam-condition competence, and everyday indicators don’t register the decline until supervised conditions expose it.

when better scores mask worse learning

Feedback That Actually Diagnoses

Feedback that only tells a student whether their final answer is right or wrong leaves the most important question unanswered: at which step did the reasoning diverge, and what’s the correct move at that point? Most revision resources—worked examples watched after the fact, answer-key checks, rereading notes—operate at the outcome layer. A student learns whether their answer matched. They don’t learn where their logic broke down or why the examiner’s solution takes a different route.

Shute’s 2008 peer-reviewed synthesis “Focus on Formative Feedback,” published in Review of Educational Research, finds that feedback effectiveness depends on design features such as specificity, actionability, and fit to the learner’s current state—not simply on the presence of feedback. That’s why step-targeted guidance is instructionally different from right/wrong confirmation: specificity about where reasoning diverged is the design feature that makes correction possible. Those findings cover general feedback mechanisms, which sets up a sharper question when the marking criteria are unusually explicit—as they are in IB Maths.

Mathspace, an online mathematics program developed by Mathspace Pty Ltd, illustrates what step-level diagnostic design looks like in practice. The platform evaluates every intermediate algebraic or procedural step a student submits, distinguishing among a correct step on the recommended path, a mathematically valid but off-path step, and an incorrect entry. Hints and short lesson videos attach to the exact step in progress rather than to the problem generically, mirroring how an expert marker follows a working script to locate precisely where a student’s reasoning diverged.

A tool that evaluates only the final answer can report whether a student is right, but cannot locate the step at which the reasoning broke down—step-aware design addresses a different, more granular problem. Mathspace’s step-level structure closes that gap: there’s a real difference between “your answer is wrong” and “your reasoning went wrong here”—the first is a verdict, the second is diagnostic feedback in its most actionable form. It does not, however, train the separate skill of evaluating completed reasoning against an externally defined marking standard after the problem is finished—a distinct cognitive act the IB exam also demands.

The IB-Specific Layer—Markschemes and Exam Conditions

IB Maths candidates face a scoring structure that makes post-attempt markscheme comparison a necessary and distinct practice skill. The IBO document Mathematics: Analysis & Approaches Specimen Papers + Markschemes (first assessment 2021) PDF makes the structure explicit: markschemes assign M (method), A (accuracy/answer), and R (reasoning) marks separately, and examiner instructions state, “Do not automatically award full marks for a correct answer; all working must be checked.” Candidate instructions reinforce this, warning that full marks are not necessarily awarded for a correct answer with no working shown. A student can reach the right final result and still lose marks if the demonstrated method and reasoning don’t match what the markscheme requires at each step.

Panadero and Jonsson’s 2013 peer-reviewed review “The use of scoring rubrics for formative assessment purposes revisited: A review,” published in Educational Research Review, finds that rubrics support learning primarily through criteria clarity and self-assessment—students comparing completed work against an explicit standard to identify what needs to change in future attempts. In IB Maths, the markscheme is that external standard, and learning to read finished reasoning against what actually earns credit at each step is the practical skill at stake. What the review documents—that making criteria explicit enables more accurate self-audit—is precisely the mechanism IB’s step-structured marking operationalizes.

Closing that gap—between reaching a correct answer and demonstrating the method and reasoning that actually earn marks—is the specific problem that Revision Village, an online revision platform used by more than 350,000 IB students from over 135 countries, is structured to address. A Revision Village review centres on the platform’s Questionbank: syllabus-tagged, exam-style questions paired with written markschemes and step-by-step video solutions produced by experienced IB educators, including examiners. After completing a question independently, students compare their finished reasoning against those markscheme criteria, seeing precisely where their logic diverged from what examiners reward. Timed practice exams add the pressure and structure of the real assessment, training performance under exam conditions rather than in an open-ended environment. For IB candidates, understanding the material and performing against examiner criteria under time pressure are related but distinct skills—and the distance between them tends to surface only when conditions are real. Revision Village’s design addresses both: markscheme comparison builds criterion awareness, and timed practice develops the capacity to perform when those conditions apply.

One Question Worth Asking of Every Revision Tool

One high-leverage question to ask of any revision resource is whether it forces you to produce an answer independently and then shows exactly where your reasoning broke down—or whether it mainly re-exposes you to correct material. Dunlosky et al.’s 2013 review “Improving Students’ Learning With Effective Learning Techniques,” published in Psychological Science in the Public Interest, rates practice testing and distributed practice as “high utility” while rating rereading and highlighting as “low utility,” and notes that many students rely heavily on the lower-utility strategies. Roediger and Karpicke (2006) further show that prior testing can outperform restudying on delayed tests, even when restudying feels more reassuring in the moment. Together, they make a reliable case for “forces independent attempts and reveals where reasoning broke down” as a selection criterion—one that holds broadly across learners and settings when implementation quality, prior knowledge, and time-on-task are all working in the same direction.

Run that check on a typical revision stack and it’s often heavier on re-exposure than on retrieval—which is precisely the imbalance that proctored exams have a way of surfacing at the worst possible time.

Choosing Revision Tools by Their Feedback

Feedback quality—specifically whether a resource reveals reasoning gaps when conditions demand genuinely independent work—is one of the most practical ways to judge revision tools.

That pattern—where unsupervised practice looks adequate but supervised performance exposes a gap—captures the structural risk of tools that mainly reduce visible effort. Volume and polish are not reliable proxies for the feedback quality that most strongly influences whether practice transfers to the exam room.

The governing question to ask of any revision resource is whether it forces independent reasoning and then shows, specifically, where that reasoning diverged from what examiners award. Step-level diagnostic feedback—the kind Mathspace is designed to deliver—addresses one part of that: locating exactly where reasoning breaks down. Alignment with IB marking criteria under timed conditions, the practice that a Revision Village review covers, addresses the other: training post-attempt comparison against what actually earns marks. Students who apply that question before exam season find out early. Students who don’t may only find out on results day.