AI Tutoring Versus Human Tutoring Outcomes

Start with the number that should have changed everything. In 1984, psychologist Benjamin Bloom showed that one-on-one tutoring with a mastery-based approach pushed students two full standard deviations above their classroom-taught peers. The average tutored student outperformed the vast majority of students in traditional instruction.
Bloom's finding wasn't an indictment of teaching. It was an indictment of cost. High-quality tutoring has largely worked. It's just consistently been expensive.
That gap between what's possible and what's affordable is exactly the problem AI tutoring was supposed to solve. So does it?
A 2025 Harvard randomized controlled trial tested 194 students, giving each group one week with a traditional active-learning classroom and one week with an AI tutor. Median test scores climbed from 2.75 to 4.5 on a six-point scale. Roughly double the learning per hour. Sounds like a headline.
But here's the part that gets quietly dropped. This wasn't someone handing students a free ChatGPT subscription and wishing them luck. The AI system was GPT-4 integrated with faculty-designed pedagogical scaffolds, step-by-step reasoning logic, and explicit guardrails against hallucination. That result belongs to a carefully engineered system. Attributing it to "AI tutoring" as a broad category is a bit like crediting "cars" for a Formula 1 lap time.
A 2025 Stanford meta-analysis covering 73 studies and tens of thousands of students found an average effect size of 0.47 standard deviations compared to traditional classroom instruction. Call it roughly half a year of additional learning per school year.
A 2025 systematic review complicates that number further. Those effect sizes shrink significantly when AI tools go up against other modern tutoring methods rather than passive classroom instruction. "Better than sitting in a lecture" is a much lower bar than "better than a good tutor." The comparison group is doing a lot of quiet work in most of these studies, and a lot of the headlines are picking the most flattering one available.
Where human tutors still outperform AI in the research
The gap isn't about content knowledge. GPT-4 scored 1460 on the SAT, landing in the 96th percentile. The model knows the material. So why does this comparison still favor humans in meaningful ways?
The gap is about what happens when a student is confused in ways that don't show up in a wrong answer.
Research from Zheng and Li in 2025 found that AI tutors followed predictable response patterns and struggled to adjust when students needed redirection or more scaffolding. Human tutors used richer questions, adaptive follow-ups, and what researchers call metacognitive prompts. Prompts that help students think about their own thinking, not just arrive at the right answer.
A human tutor notices hesitation. They hear the "I think?" at the end of an answer. They catch frustration before it hardens into avoidance. AI largely cannot do this reliably yet. And that's not a small gap to paper over with more practice problems.
A 2024 meta-analysis of 282 randomized controlled trials confirmed that high-impact tutoring remains one of the most effective academic supports available, and that quality depends heavily on that human element.
There's also a structural problem in many AI tutoring tools that doesn't get enough airtime. Evaluations have found that current systems often deviate from the course content they're supposed to cover, lack any sustained strategy across a session, and weren't built with real subject-matter expertise baked into the curriculum design.
That last one is the one I'd push on. AI optimizes for correctness signals. Human tutors optimize for understanding. Those are not the same target. Building a tool around the wrong one produces a lot of confident-looking wrong turns, and students often can't tell the difference until they're sitting in front of an exam.
How these differences play out in SAT preparation specifically
Cost shapes everything here, so start there.
One-on-one SAT tutoring runs from $45 per hour at the low end to well over $1,000 per hour for elite providers. A meaningful score improvement, roughly 23%, typically requires 36 to 48 hours of one-on-one instruction, more than 25 hours of homework, and 6 to 9 full-length practice tests spread over several months. Run that math and Bloom's access gap from 1984 looks pretty much intact.
So what can AI actually deliver on the SAT? One platform reports 80 to 150 point score improvements across more than 127,000 students. Those are platform self-reported figures, not independently verified results. But the direction is consistent enough to take seriously rather than dismiss.
A randomized trial adds a useful wrinkle. AI tools significantly helped lower-performing students but showed mixed or even negative effects for higher-proficiency learners. Outcomes diverged sharply rather than rising uniformly.
Why does that happen? Think about what AI tutoring is structurally good at. Volume of practice. Instant feedback. Patience with repetition. No cost-per-hour pressure. Those advantages compound for a student who has foundational understanding and just needs reps. For a student with a foundational gap they can't even articulate yet, more volume doesn't help. It reinforces confusion at scale, which is its own kind of problem.
Of the more than 2 million students from the Class of 2025 who took the SAT, only 39% met both the Reading/Writing and Math college readiness benchmarks. The preparation gap is large. But the question of which tool closes which part of it is a different question for different students, and that distinction matters more than any aggregate number.
What the AP exam context adds to the picture
Over 1.3 million public high school graduates took more than 4.8 million AP Exams in 2025, with participation growing 7% from the prior year across 36 of 40 AP courses.
More students taking AP does not automatically mean more students passing. AP Statistics had only 60.3% of students score a 3 or higher in 2025, and plenty of subjects carry significant failure rates.
There's also a format layer that gets overlooked. Twenty-eight of 36 AP subjects now use the College Board's digital Bluebook platform, with 16 fully digital and 12 hybrid. More than 90% of students found the interface easy to use. But familiarity with format under timed pressure is a real variable, and it's one that doesn't show up until it matters most.
This is actually where AI has a genuine structural advantage. A well-designed AI tutoring tool can shift modes within a single session: Socratic questioning to uncover a conceptual gap, then timed drills to build speed, then format walkthroughs to reduce exam-day surprises. Human tutors can do all of this too. The cost-per-hour clock just doesn't stop running while they do it.
One honest caveat on the AP AI evidence specifically. After two years of Khanmigo operating at scale, rigorous independently controlled outcome data tied to learning gains is still pending. What internal data does show is a 6.1% improvement in next-item correctness across more than 15 million tutoring threads in roughly a six-month window. Directionally interesting. Not a controlled outcome study.
A mixed-methods study of 69 undergraduate students comparing Khanmigo to Google search found significant learning gains in both conditions, but no statistically significant difference between them. Which is a useful reality check on the bolder platform claims.
AP subjects also vary enormously in what kind of help actually matters. A student struggling with AP Calculus BC free-response format needs something different from a student with conceptual gaps in AP Biology. The format of the tutoring has to match the subject's demands, whether that tutor is human or AI.
Why the mechanism difference — not the effect size — is the most useful frame for students


Most comparisons treat this like a horse race. AI got 0.47 standard deviations. Bloom's human tutoring hit 2.0. Human tutoring wins, full stop. Move on.
But what do those numbers actually tell you about what to do next Tuesday, three weeks before your AP exam? Not much.
The more useful question is what each approach does well, and why.
AI's mechanism: available any time, consistent feedback quality, unlimited patience with repetition, no anxiety about asking the same question three times, no cost-per-session pressure. These advantages compound for students who are self-directed and already have foundational understanding. The student close to mastery who needs volume and reinforcement gets a lot out of a good AI tool.
Human tutoring's mechanism: adaptive real-time dialogue that reads between the lines, metacognitive scaffolding, recognition of emotional and motivational state, curriculum authority from someone who deeply knows the subject and its specific demands. These advantages matter most when a student is stuck in a way they can't articulate. When the problem isn't a gap in practice reps but a gap in understanding that more practice will not fix.
Research is also pretty clear that unsupervised AI use for simplification or answer generation produces limited or inconsistent benefits. The pedagogical design of the system drives results. Not AI in general.
But what if a student doesn't know which category they're in? That's the messier, more honest version of this question. A student who thinks they need more practice might actually need someone to tell them their foundational model of a concept is just wrong. More volume won't surface that. Only someone asking the right questions will. And that someone doesn't have to be human, but the tool has to be built to ask those questions in the first place.
So the real question isn't "AI or human tutor?" It's a sequencing question: which mechanism does this specific student need right now?
What to look for in an AI tutoring tool if you're using one for AP or SAT prep
Pedagogical design over raw model power. The Harvard result came from a system with faculty-designed scaffolds and guardrails, not from GPT-4 alone. Ask whether the tool was built with real subject-matter expertise or assembled quickly on top of a general model. That difference shows up in results, sometimes badly.
Explanation over answer revelation. A tool that tells you an answer is wrong is less valuable than one that tells you why, and walks you through the reasoning. The mechanism driving learning gains in the research is explanation and reasoning. Not performance logging.
Gap detection over problem volume. A tool that tracks what you actually understand versus what you only think you understand is categorically different from one that logs right/wrong ratios. Drilling more problems without gap detection just reinforces what you already know. It feels productive. It often isn't.
Alignment to College Board standards. For SAT and AP specifically, content needs to be reviewed against actual exam standards. Non-aligned AI prep is a real risk, especially for AP free-response format and the SAT's specific question structures.
Passionfruit is built around these criteria: AI-powered grading with substantive feedback, unlimited practice problems aligned to AP and SAT content, and a knowledge model that maps what a student actually understands rather than just logging performance.
Where a human tutor still earns the cost: when a student has a foundational gap they can't self-diagnose, when motivation or test anxiety is the real barrier, or when a subject requires the kind of adaptive, real-time dialogue that current AI can't reliably replicate.
Bloom's two-sigma problem was never really about finding a better teacher. It was about access, consistency, and scale. AI doesn't replace the best human tutors. For a student who never had access to those tutors in the first place, a well-designed AI tool represents a meaningfully different situation than anything that existed ten years ago. Whether that's enough depends entirely on which problem that student is actually trying to solve.


