Edutopica

Using AI to Close Achievement Gaps at Scale

AI tutoring systems are closing learning gaps three times faster than traditional interventions.

Senior Writer · · 10 min read
Cover illustration for “Using AI to Close Achievement Gaps at Scale”
AI in Education · August 18, 2026 · 10 min read · 2,320 words

Let's start with the number that made me raise an eyebrow. A randomized controlled trial published in Scientific Reports back in June 2025 found AI tutoring beat regular classroom instruction by 0.73 to 1.3 standard deviations. Students scored higher, and they got there faster.

I want you to sit with that 0.73 for a second, because in education research, that's not a normal number. Most interventions schools spend real money on (new curricula, smaller class sizes, extra homework) land under 0.5. When you see 1.3, the honest reaction isn't excitement. It's suspicion. Researchers who study this stuff for a living check the methodology twice before they believe it.

Then there's a separate program working with disadvantaged students that found 0.3 standard deviation gains in six weeks. Translate that into plain terms: roughly two years of normal learning, compressed into a month and a half. It beat 80% of the education interventions ever studied in developing countries. And the girls who started out struggling the most saw the biggest jumps, gains that held up for as long as researchers kept watching.

Add a Coursera survey from February 2026 (over 4,200 students and educators, five countries) where four out of five students said AI helped their grades, and 70% expected it to help their exam results too. Self-reported survey data is softer evidence than a controlled trial, sure. But when a rigorous RCT and a big multi-country survey both point the same direction, that's worth taking seriously.

What we still don't know: does this stick a year later? Does it work this well for every kind of student, or mostly the ones already primed to benefit? And what happens the moment you take the AI away?

Diagram: AI Tutoring vs. Conventional Interventions: Effect Sizes. Visualizes: Show a ranked magnitude comparison of education intervention effect sizes to make viscerally clear how unusual the AI tutoring results are.

What makes AI personalization structurally different from adaptive quizzing

Quick gut check: is a quiz app that gives you a harder question after you get one right actually "personalizing" your learning?

Not really. It's adjusting a dial. That's useful, but it's not thinking about you.

Real personalization is a system constantly asking three questions about one specific kid: does this student need to go back and fix something right now, are they ready to speed up, or do they need to branch backward into a prerequisite they never actually nailed? That's a different job entirely from turning a difficulty knob.

Think about how pacing works in a normal classroom. It's set by the average. The kids who get it fast get bored and check out. The kids who need more time get left behind, because the class has already moved to the next unit whether they're ready or not. AI sidesteps that problem completely: every student moves at their own speed, and nobody's waiting on anybody else.

Here's the part that actually surprised me: the system gets sharper the more a student uses it. It's not just teaching the subject, it's learning the learner. That means gaps get caught the moment they happen, not three weeks later on a unit test, after a small misunderstanding has calcified into something much harder to undo. That's the whole pitch behind "AI as a tutor for everyone": the same diagnostic instinct a great human tutor has, minus the requirement that a human be physically in the room.

Venn diagram: AI Tutoring vs. Traditional Instruction. Compares AI Tutoring and Traditional Classroom; overlap: Shared Goals.

Where AI adoption stands in schools right now, and who's getting left behind

The adoption numbers move fast enough to give you whiplash. Most K-12 teachers now use AI tools in class, a sharp jump from a year ago, and the global AI-in-education market reached multiple billions of dollars in 2025. Most districts are on pace to have AI training in place for teachers by the end of the year.

For teachers using it weekly, the payoff is real: about 5.9 hours a week back. Add that up over a school year and you're close to six full working weeks, mostly clawed back from grading, planning, and busywork.

But here's the part that should bother you. Low-poverty districts are way ahead in adoption. By fall 2024, 67% of low-poverty districts had trained teachers on AI, compared to a much smaller share of high-poverty districts. That gap showed up in 2023, and it hasn't closed since. The districts serving kids who need the most help are the last ones getting the tools.

So, is AI just going to become another version of the same inequity it was supposed to fix? That's a live question, and the answer depends entirely on whether anyone deliberately routes these tools toward high-poverty districts first, rather than letting the market sort it out on its own timeline.

One thing worth untangling here: district rollout and individual access are two different problems. A kid in a district that hasn't trained a single teacher can still open a laptop tonight and use a tool directly. That's where something like Passionfruit fits in, a student-facing tool that reaches a kid who needs help right now, without waiting on a district-wide plan that might take another two years.

How AI detects the specific gaps that hold students back, not just what they got wrong, but why

Knowing a student got a question wrong is easy. Knowing why is the actual job.

Khan Academy's internal research put real numbers on this. Show the AI a student's recent problem-solving history (what they attempted, what they got right, what they missed) and next-item accuracy improved by 3.4%, across 608,000 tutoring threads. Show it which prerequisite skills a student hadn't mastered before handing them a harder problem, and accuracy improved by 2.7% across 1.36 million threads, with a 98.5% chance that having that information beat not having it. Include the full prior conversation, and cognitive engagement jumped 5.09%, with a 99.4% chance of a real improvement.

Those percentages look small sitting on a page by themselves. But multiply them across millions of interactions and you're looking at a meaningful shift in what students actually learn. This isn't abstract. Gap detection changes the very next problem a kid sees.

Where's this heading next? Researchers are now mining student-AI chat logs to catch knowledge gaps by spotting patterns in the conversation itself, the kind of phrasing that gives away a missing concept before a wrong answer ever shows up (Fu, Wu & Williams, 2025). The tutoring session becomes the diagnostic tool. That's the idea Passionfruit is built around: catching exactly where a student's reasoning breaks down, not just logging which answers were wrong. A checkmark tells you nothing. Knowing where the thinking went sideways tells you everything.

What the AP exam landscape reveals about where targeted gap detection matters most

Diagram: AP Exam Gap: Participation vs. Passing. Visualizes: Visualize the widening distance between AP participation and qualifying scores using the 2025 data.

AP exams make a good stress test, because more kids are showing up, but pass rates haven't kept up with them. In the class of 2025, 1,307,781 students (37.0% of U.S. public high school graduates) took at least one AP exam, up from 34.3% in 2015.

Qualifying scores haven't kept pace with that growth. In 2025, 24.8% of graduates (875,778 students) scored a 3 or higher on at least one exam, up from 20.7% in 2015. More kids taking the exam, sure, but the gap between "took it" and "earned credit for it" is still wide open.

There's genuine progress on who's showing up: 497,799 Black, Hispanic/Latino, and American Indian/Alaska Native students graduated in 2025 having taken an AP exam, up 167,412 from 2015. That matters, and it's worth saying plainly.

But look at the pass rates and the story gets harder: AP Latin, 58.6% scoring a 3 or higher. AP Statistics, 60.3%. AP Music Theory, 60.5%. AP Computer Science Principles, 61.9%. AP Human Geography, 64.7%. AP World History, 64.3%. AP Physics 1, 67.3%. These aren't obscure electives. Huge numbers of students take these courses, and more than a third of them walk away with no credit to show for it. Overall participation grew 7% from 2024 to 2025 across 36 AP courses. Scale is going up right when pass rates are the real problem.

Here's what makes this hard to fix with human tutors alone: a kid taking five AP courses needs expert-level help in five completely different subjects, at the same time. Hire five human tutors at $60 to $150 an hour each, and you've priced out most families before you even start. This is exactly where AI's ability to give gap-aware, subject-specific help across multiple courses, without a per-session invoice, actually matters. It's not a nice-to-have. It's the only version of this that scales to everyone instead of just the families who can afford five tutors.

How the Digital SAT changed what effective prep looks like, and what AI tools can now do

The Digital SAT quietly rewrote the rules, and a lot of families still haven't caught up. It's computer-adaptive: how you do on Module 1 decides how hard Module 2 gets, and a harder Module 2 is the only way to unlock a higher score.

A static prep book with the same fixed sections in the same fixed order can't simulate that. It just can't. AI-adaptive tools match the format of the actual test in a way a book never will, because the test itself is adaptive.

The numbers back this up. Per the College Board's Digital SAT Impact Report from January 2026, students using AI-adaptive prep in 2025 scored 90 points higher on average than students using static practice books. A joint study between the College Board and Khan Academy found students who finished the recommended AI practice plan on Khan Academy gained 120 points on average.

And the cost gap is almost embarrassing. Premium SAT prep courses run $600 to over $2,000. Solid AI-powered options exist for under $150, some completely free. That's not a minor discount. That's closing a gap that used to make good prep a straightforward function of family income.

A few names worth knowing: Khan Academy is the official College Board SAT partner, strong on concept mastery, and free. Passionfruit is built specifically for AP and SAT mastery, with unlimited practice problems, AI grading, and gap detection aimed at explaining why a student missed something, not just flagging that they did.

Now, the catch, because there's always a catch. Some AI-generated practice questions mimic the SAT's style without testing the actual reasoning skills the College Board cares about. Experts who reviewed at least one AI-generated practice test found questions that didn't line up with the real skill targets at all. A student can grind through 300 of those and walk into the real exam with confidence that has nothing behind it. So the quality of the feedback matters as much as the quantity of the practice. More questions isn't the goal. Better-targeted questions are.

How AI grading gives teachers back the time they need to teach

About 70% of a teacher's non-teaching time gets swallowed by grading, planning, and admin work. That's time that could otherwise go toward sitting next to one kid and figuring out exactly where they got stuck.

A teacher with 150 students across five classes cannot give each one personalized written feedback on every assignment. The math doesn't work, no matter how much they care. AI grading is starting to change that math. Tools like CoGrader take a teacher's own rubric and grade a whole class in a fraction of the time, with the teacher reviewing and signing off rather than starting from zero on every paper.

EssayGrader matched teacher scores, or landed within one point of them, the vast majority of the time, benchmarked against tens of thousands of essays from The Learning Agency. That's accurate enough to actually use, not just accurate enough for a sales demo.

Teachers in high-poverty schools tend to carry the heaviest caseloads with the least support. AI grading has a real shot at raising the floor for those teachers specifically, not just handing extra convenience to teachers who already had room to breathe. The global AI grading market reached hundreds of millions of dollars in 2025, which tells you this isn't a fringe experiment anymore.

But here's the line that matters: AI grading is only worth something if it buys teachers more time for human interaction, not less. If it just moves paper faster without changing what happens in the room, that's a productivity trick wearing an equity costume. Generic, slapped-together feedback closes zero gaps. It just speeds up the shuffling. The feedback quality is the entire point.

What it actually takes for AI to close gaps rather than just widen them differently

Here's the part nobody likes saying out loud. If AI tools show up first, and best, in the districts that already have everything, they don't close the gap. They widen it, just with better technology this time. The tool itself isn't the problem. Who gets it first is.

What decides which way this goes is the quality of the system, not the fact that it's "AI." A tool that tracks wrong answers without ever asking why is just a faster, cheaper version of a mediocre tutor. Speed without insight doesn't help anybody, it just fails faster.

So what actually separates a gap-closing tool from one that just looks impressive? Three things, and none of them are optional:

  • Real prerequisite mapping. The system has to know what a student needs to understand before the next concept has any shot at landing.
  • Feedback that explains the reasoning, not just the correctness. A student needs to know where their thinking veered off, not just that the final answer missed.
  • Memory across sessions. It should build on what it learned about a student last time, instead of starting cold every single time.

What made a good private tutor worth the money was never how much material they had memorized. It was a person watching a kid think, catching the exact second understanding slipped, and redirecting right then. That's the thing worth building at scale. Not more content. Not more questions. That instinct, handed to every student, whether their family can afford a $150-an-hour tutor or not.

Sources

  1. engageli.com
Filed underAI in Education

More in AI in Education