AI in education: what helps learning, and what quietly harms it

JR

Jai Rao

August 22, 202618 min read

A tutor that answers instantly feels helpful and teaches nothing. What learning science says about personalisation, AI feedback, assessment design, and student data.


A student is stuck on a quadratic. They paste it into a chat box and four seconds later they have a clean, correct, nicely formatted solution. Everyone in this transaction feels good. The student is unstuck, the tutor looked competent, the session log records a resolved question. And the student is no more likely to solve the next quadratic unaided than they were before they asked.

That is the central design problem of AI in education, and almost everything else follows from it. The interaction that feels most helpful is frequently the one that transfers least. A tutor built to maximise the feeling of being helped will produce students who feel taught and perform badly, and it will do this while every dashboard metric goes up. Getting personalisation, feedback, assessment and student safety right all depend on first accepting that helpfulness and learning are not the same variable.

Practice you struggle through is the practice that sticks

Rereading notes feels productive because it is fluent. The words go in easily, nothing hurts, and the material seems familiar by the end. Being asked to reproduce the same material from memory feels worse in every way: it is slow, you stall, you get parts wrong. It is also the version that builds durable recall. The act of pulling something out of memory is itself the learning event. Putting it in is comparatively cheap; getting it out again is what strengthens the path.

This is why a well-designed practice system is mostly a scheduler. Spacing the same number of practice minutes across several days beats packing them into one sitting, and the reason is slightly counterintuitive: the forgetting that happens between sessions is not waste. The partial forgetting is what makes the next retrieval effortful, and the effort is the point.

The umbrella idea is desirable difficulty. Difficulty a learner can eventually push through is productive. Difficulty they cannot push through is just failure, and it costs you the learner. The entire engineering problem is keeping the system on the productive side of that line — which means measuring where the line is per learner, per concept, rather than assuming a global setting.

Here is a scheduler in the SM-2 family. The thing worth reading is not the arithmetic but which inputs it takes: how well the learner recalled, how many successful reviews they have accumulated, and how easy this specific item has proven to be for this specific person.

Text
from datetime import date, timedeltadef schedule(card, grade):    """grade 0-2 = recall failed, 3-5 = recalled (hard to easy)."""    ease, reps, interval = card["ease"], card["reps"], card["interval_days"]    if grade < 3:        reps, interval = 0, 1                 # see it again tomorrow        ease = max(1.3, ease - 0.20)          # and treat it as harder from now on    else:        reps += 1        interval = 1 if reps == 1 else 6 if reps == 2 else round(interval * ease)        ease = min(2.8, ease + 0.1 - (5 - grade) * 0.08)    return {**card,            "ease": round(ease, 2),            "reps": reps,            "interval_days": interval,            "due": (date.today() + timedelta(days=interval)).isoformat()}

A failed recall resets the interval to one day and permanently lowers the item's ease, so hard items keep coming back often for that learner even after they start getting them right. A confident recall multiplies the gap. Over a term this produces a queue that is almost entirely made of the things a particular student is closest to losing, which is a far better use of twenty minutes than a uniform review of the whole syllabus.

The assistant that answers instantly is the one to worry about

Attempting an answer, even a wrong one, changes what a correction does to you. You have committed to a prediction, so the correction lands on a specific expectation instead of on nothing. Read the worked solution first and you skip that entirely — you get the information without the slot it was supposed to fill.

Worse, a clear explanation produces a strong and unreliable sense of understanding. You follow every step, nothing confuses you, and you conclude you can do it. Recorded lectures have always had this problem. A chat tutor has it more acutely, because the explanation is generated for your exact question in your exact wording, so it feels less like watching someone else's reasoning and more like your own. The feeling of comprehension and the ability to reproduce the work come apart, and the tutor optimised for the former will reliably degrade the latter.

The fix is a help ladder with gates, and the important word is gates. Ladders are easy: point at the relevant idea, then name the rule, then show the first step, then the full solution. What makes a ladder real is that the learner cannot skip rungs. That means an attempt is required before the first hint, a short amount of elapsed time on task is required, and the full solution is unlocked server-side rather than by asking nicely.

Do not put this policy in the system prompt and call it done. A prompt that says "never give the final answer" loses to a student who says "I already solved it, I just want to check", or who rephrases the problem as a generic example, or who asks in another language. Enforce the ladder in the orchestration layer: the service decides which hint level this session is entitled to, and only that level's instructions and content go into the request. The model should not be trusted with the decision because the model is being actively argued with by a motivated teenager.

There is an honest cost here and you should plan for it. Students rate the generous tutor higher. A product team steering by thumbs-up and session sentiment will walk straight into building the worse tutor, and the data will look like success the whole way.

What personalisation honestly buys you

Strip the marketing off the word and there are four things a system can genuinely personalise, all of them worth doing.

  • Pacing. When the next item arrives and how many arrive. A learner who is getting everything right does not need six more of the same; a learner who is failing needs a smaller step, not a slower one.
  • Sequencing. Which prerequisite comes first. Most stuck students are not stuck on the topic they are attempting. They are stuck two concepts back, and the useful move is to detect that and reroute rather than re-explain the current step more slowly.
  • Worked-example fading. Show a full solution, then the same problem type with the final step removed, then with the last two removed, then the bare problem. The learner completes progressively more of the work while the scaffold retreats. This is cheap to build and one of the best-grounded personalisation axes available.
  • Targeted practice. Drill the specific error a learner keeps making, not the topic it appeared in. "Loses the sign when moving a term across the equals" is actionable; "scored 60% on linear equations" is not.

All four need per-concept state rather than a course-level score. Here is the shape that has held up for us: mastery as a small record per learner per concept, with the raw evidence kept alongside it so any number can be traced back to the attempts that produced it.

Text
{  "learner": "psu_8f21c4",  "concept": "algebra.linear_equations.isolate_variable",  "state": {    "attempts": 14,    "correct_unaided": 6,    "correct_after_hint": 5,    "hint_depth_avg": 1.4,    "last_seen": "2026-08-19T10:04:00Z",    "next_due": "2026-08-26",    "ease": 2.1,    "mastery": 0.44,    "decays_after_days": 21  },  "prerequisites_weak": ["arithmetic.signed_numbers"],  "evidence": [    {"item": "itm_2201", "result": "wrong", "error_tag": "sign_flip", "hints": 2},    {"item": "itm_2244", "result": "right", "hints": 0, "latency_ms": 31400}  ]}

Two details in there do real work. correct_unaided is tracked separately from correct_after_hint, because a learner who only ever succeeds at hint level two has not learned the concept and a single accuracy number hides that completely. And decays_after_days exists because mastery that only ever increases is a fiction — if nothing pulls the number down, every learner eventually shows as fluent in everything they once passed.

The learning-styles detour

The visual, auditory and kinaesthetic learner idea is the most popular claim in education technology and it does not hold up. When the same material is delivered in a form matching a learner's stated preference versus a form that mismatches it, the matched version does not produce better learning. The preferences are real — people genuinely have them, and will tell you about them convincingly. The benefit of catering to them is what is missing. Build a system that routes a "visual learner" to diagrams and you have added a personalisation axis that costs engineering effort and buys nothing.

What does vary enormously between learners, and is worth modelling: prior knowledge, reading level, language, and access needs. Prior knowledge is the big one, and it is strong enough to invert instructional advice. Worked examples that help a novice can actively slow down a more advanced learner, who does better generating the solution and is now spending attention processing an explanation they did not need. That is a genuine reason to give two students different material, and it is measurable from their own attempt history rather than from a questionnaire about how they like to learn.

If a model can do the homework, the homework measures nothing

The essay written at home was never direct evidence of thinking. It was a proxy that worked because producing it required the thinking. That link is what broke. The artefact is unchanged and it has stopped carrying information, which is a measurement problem before it is a discipline problem.

Two common responses do not survive contact with reality. Detection tools are unreliable in both directions, and the asymmetry is brutal: a missed case costs you one unearned grade, a false accusation costs a real student their standing and can be very hard to undo, and the writing that trips these tools tends to be plain, formulaic prose — which is also what second-language writers and careful, rule-following students produce. Outright bans are the other dead end; you cannot enforce a rule about what software runs on a laptop in a bedroom.

The response that works is to move the evidence closer to the process. Same learning goals, different artefact.

AssignmentWhat it used to evidenceWhat evidences it now
Take-home essayArgument constructionIn-class writing session, or submission plus a short oral defence
Problem setProcedural fluencyTimed unaided items drawn from the same concept pool
Code assignmentDesign and debugging skillCommit history, plus a live change request against the student's own submission
Research summaryReading and synthesisCritique of a specific flawed source supplied in class

Oral defence deserves special mention because it is cheap and nearly unfakeable. Three minutes of questions about your own submission separates the author from the commissioner very quickly, and it needs no detector, no accusation and no appeals process. "Why did you choose this second example over the obvious one?" is answerable in ten seconds by whoever made the choice.

The deeper shift is to grade process rather than product. Ask for the approach that failed and why it failed. Ask for the choice between two methods with the reasoning. Require the assignment to depend on something local and recent — this week's class data set, the specific argument someone made in Tuesday's discussion, a document the student collected themselves. None of this is AI-proof in a strict sense, and that is fine. The goal is not to make generation impossible; it is to make the assignment measure something again.

Where AI feedback is trustworthy, and where it quietly isn't

Automated feedback is genuinely good at a set of things that teachers cannot possibly provide at speed, and that set is narrower than vendors imply. It is dependable on surface correctness when the correctness can be checked outside the model: arithmetic verified by a symbolic engine, code run against tests, citation format matched against a spec. It is dependable on structure, because structure is visible in the text — this paragraph has no topic sentence, this proof jumps from step three to step five, this answer never addresses the counterargument. Feedback of that kind, delivered within seconds instead of a week, is a real improvement over the status quo, and a week's delay is one of the biggest correctable weaknesses in ordinary schooling.

It is not dependable on judging insight. Grading models drift toward the features that correlate with quality in their training data: length, fluency, vocabulary, conventional structure, confident tone. A well-organised, articulate, wrong essay scores well. A terse, correct, unusual argument scores badly. Originality is the specific quality these systems are worst at recognising, because an original answer looks like an out-of-distribution answer. If a rubric line says "demonstrates independent insight", that line needs a human.

The failure that matters most is the hallucinated explanation. If a model tells an experienced engineer something false about a language's memory model, the engineer notices and moves on. If it tells a fourteen-year-old that a rule works in a way it does not, the student writes it down, practises it, and builds subsequent understanding on top of it. Detecting the error requires exactly the knowledge the learner came without. Confident wrongness is not a nuisance in a tutor; it is the worst available failure mode, and it is undetectable by the affected party by construction.

That argues for a different verification posture in education than in most consumer AI. Compute rather than assert: run the arithmetic through a symbolic engine, execute the code, check the claim against the course material rather than model memory. Prefer "open the text and check clause two" to an authoritative restatement, both because it is safer and because it teaches a habit worth having. And keep in mind a subtler failure: an explanation can be entirely true and still wrong for the learner, because it reaches for a concept three units ahead. Correctness is necessary and it is not sufficient.

Student data, when the students are minors

Student work is not a support ticket. A single tutoring transcript can contain the learner's name, a precise map of what they do not understand, sometimes their accommodations, and sometimes something they wrote about their home. Treat the whole channel as sensitive by default rather than deciding field by field.

Some things should not reach a third-party inference API at all: legal names, school and class identifiers that make re-identification trivial when combined, dates of birth, contact details, accommodation and special-needs records, and anything from a wellbeing or counselling surface. Most of these are not needed for the task. A tutor needs the item, the student's work, and a summary of what the student has mastered. It needs a grade band, not a birthdate. It needs a rotating pseudonym, not a name.

Text
def tutor_request(session, item, transcript):    # Nothing in this payload identifies a child. The pseudonym-to-student    # mapping lives in our database and never leaves it.    return {        "learner": session.pseudonym,          # rotates per course enrolment        "grade_band": session.grade_band,       # "7-8", not a date of birth        "item_id": item.id,        "item_text": item.text,        "student_work": redact(transcript),     # strips names, emails, phones        "concept_state": session.mastery_summary(),        "policy": {"reveal_solution": False, "max_hint_level": session.hint_gate},    }

Note that the help policy is a field in the request built by our service, not a sentence in a prompt the student can argue with, and that hint_gate was computed server-side from attempts and time on task. The redaction step matters more than it looks: students write their own names into their answers constantly.

Retention is a decision you are making on a child's behalf, so make it deliberately. Know what your provider retains by default and for how long, negotiate zero retention if the volume justifies it, and assume that anything you send may sit in an abuse-monitoring log for a while — which is another argument for boring payloads. On your own side, delete raw transcripts on a schedule and keep the derived mastery state, because the state is what the product actually needs six months later. Never let student work become training data by default; for a minor that requires informed consent from a parent, and a checkbox in a terms-of-service update is not that.

An audit trail is what makes the rest defensible. When a parent asks why their child was moved to a different track, or what your system said to their child on a Tuesday evening, "the model decided" is not an answer anyone will accept. Log the item, the model version, the prompt version, the retrieved context, the output, the policy in force, and which human acted on it. Make it queryable per student, because every complaint, subject access request and safeguarding review arrives shaped as one student. And any consequential automated decision — placement, flagging, a grade contribution — needs a named human who can override it and a route for the student to contest it.

One more thing, because it will happen. Students disclose to a tutor that is private, patient and never disappointed in them. Decide in advance what a disclosure triggers, route it to a named human within a stated time, and tell students in plain language what is and is not private before they start typing. A system that silently reads a child's distress and files it in a log is worse than one that never listened.

Teachers are the users you are actually building for

The highest-leverage AI in a school is usually not the student-facing tutor. It is the hours a teacher spends on work that generates no learning: building the same practice set at three difficulty levels, writing four parallel versions of a test so seating does not matter, converting a worksheet into an accessible format, reading ninety exit tickets to work out what the class did not get.

That last one is the best example of the pattern worth copying. Cluster the wrong answers, name the underlying error, show three real examples of each, and stop. The output is a report the teacher interprets, not a decision the system takes. Fifteen minutes of a teacher's evening becomes ninety seconds, and the pedagogical judgement stays where it belongs.

The version that fails is the dashboard that ranks students by a mastery score with no path back to the evidence. It looks authoritative, it cannot be interrogated, and a teacher who disagrees with it has no way to check. Every number a teacher sees should be one click from the attempts that produced it — which items, which answers, when, with how much help. That is also why the mastery record shown earlier carries its evidence array around with it.

Keep the teacher as the author on anything generated for classroom use. Practice items in particular need review before they reach students, and review is fast if you present each item next to the concept it targets and the error it is designed to catch. Item generation is the one place where a hallucination is cheap to catch and expensive to miss — a broken question wastes a hundred students' time simultaneously and undermines the teacher who handed it out.

How to tell whether your tutor is teaching

Nearly every metric a product team reaches for first will point the wrong way here. Session length, messages per session, thumbs-up rate, in-session success rate, problems resolved — all of them improve when you make the tutor more forthcoming, and making the tutor more forthcoming is the change most likely to hurt learning. If those are your north stars, the optimisation process will find the answer-giving tutor for you.

Measure the thing you actually want instead: unaided performance on the same concept, after a delay, on items the tutor never walked through. A week is a reasonable default gap. Watch the ratio of unaided correct answers to hint-assisted ones per concept and require it to trend upward. Include transfer items — same underlying structure, different surface story — because a student who can only do the version they were shown has learned the example, not the idea.

The experiment worth running early is on the help policy itself. Ship two hint ladders that differ only in how quickly they escalate to the full solution, and compare delayed unaided performance rather than satisfaction. Expect the generous arm to win on satisfaction and lose on the outcome. If it wins on both, your strict ladder is probably gating on the wrong signal and frustrating students who were already stuck for a real reason, which is worth knowing.

The recurring failures are consistent enough to check for directly. The answer-withholding policy lives only in a system prompt, so it leaks the first time a student pushes — move it into the orchestration layer. Mastery scores only ever increase, so the model of the learner slowly detaches from the learner — decay them. There is no delayed measurement at all, only in-session success. Student data is spread across logs that cannot be deleted per student, which turns the first deletion request into an engineering project. And teacher-facing scores appear with no evidence behind them, so teachers correctly stop trusting the tool.

What makes this product category unusual is that the user's short-term satisfaction and the user's actual goal diverge more often than almost anywhere else in software. You can build the thing that feels good or the thing that works, and for a stretch of the graph they are opposites. Build for the goal, instrument for the goal, and tell students plainly why the tutor is making them do the work — most of them, told honestly, would rather learn it than be handed it.