StudentBench: AI Tutoring Matches a Human GRE Tutor — Time to Grade the Teaching, Not the Answering
AI tutoring now delivers statistically equivalent GRE learning gains to expert human tutors at a fraction of the cost.
The StudentBench study, submitted to arXiv on 23 Sep 2026, evaluates whether large language models can replicate the learning outcomes of human instruction. Researchers collected over 175,000 student-AI messages from 2,383 participants tackling Quantitative and Verbal GRE questions. The data shows AI tutoring produces learning gains statistically equivalent to expert human tutoring (p = .015). In five of the seven tested GRE domains, the top-performing AI tutor exceeded the average performance of the human tutor.
Cost efficiency represents the most significant divergence between the two modalities. One AI tutor achieved parity with human instruction (p = .044) while operating at 918 times lower cost. The expenditure per percentage point of learning gain was USD 0.0052 for the AI system compared to USD 4.81 for the human tutor. A second study involved expert human tutors conducting 2,008 pairwise rubric evaluations of LLM-generated lesson plans and practice problems. These evaluations allowed the researchers to isolate performance across five distinct dimensions: lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement.
Operational metrics within the AI sessions revealed specific drivers of efficacy. For Quantitative GRE sessions, faster AI reply times correlated with increased student message volume. Higher message counts led to more correct practice attempts, which in turn drove larger learning gains. All three correlations registered statistical significance with p < .002. The authors have released the StudentBench platform publicly to facilitate further large-scale data collection and evaluation of AI teaching capabilities.