AI Education

Stanford Law Study: AI Tutor Answers Beat Professors 75%

A Stanford Law blind study found professors preferred AI contract-law tutoring answers in 75% of head-to-head matchups, and flagged AI answers as harmful less often than peers'.

Stanford Law Study: AI Tutor Answers Beat Professors 75% — article cover
On this page6 SECTIONS
  1. How the Study Worked
  2. The Core Numbers
  3. Which Systems Were Tested
  4. Limits and the Authors’ Own Caveats
  5. What It Means for EdTech
  6. Sources

On June 1, 2026, Stanford Law School published a blind-evaluation study, “Law Professors Prefer AI Over Peer Answers.” Sixteen law professors across U.S. law schools answered 40 contract-law questions, then blindly graded anonymized responses without knowing whether they came from AI systems or fellow professors. Across nearly 3,000 blind comparisons, AI won 75% of head-to-head matchups.

This is not another “AI beats a benchmark” story. The graders were law professors — a population professionally trained to pick apart arguments. The task was tutoring, not exam scoring. And AI answers were flagged as pedagogically harmful or misleading only 3.5% of the time, versus 12% for professor-written answers. The study was led by Stanford Law professor Julian Nyarko, head of the Legal Innovation through Frontier Technology Lab (liftlab), with liftlab researcher Alejandro Salinas as first author and co-authors including Sarath Sanga of Yale Law plus colleagues from NYU, the University of Chicago, and other institutions.

How the Study Worked

Three design choices carry the weight. First, the questions were deliberately mundane: 40 contract-law prompts of the kind students actually ask after class or during office hours — not competition problems. Second, the grading was blind: professors wrote their own answers, then scored anonymized responses with no indication of whether an answer came from AI or a colleague. Third, for fairness, AI responses were calibrated to match the length and structure of human answers, so graders could not guess the source from formatting, and multiple evaluation methods were used as cross-checks.

Sixteen professors and roughly 3,000 anonymized comparisons is not a huge sample, but the blind design controls for brand halo and identity bias — exactly what most “AI beats X” studies are missing. The length-and-structure calibration matters too: without it, graders tend to reward verbosity or penalize formatting they do not recognize, which would tilt the comparison toward whichever side happened to write longer.

The Core Numbers

  • AI beat professor-written answers in 75% of head-to-head matchups
  • AI performed comparably to the best human instructor in the study
  • AI answers were flagged as pedagogically harmful or misleading 3.5% of the time, versus 12% for peer-written answers

The third number gets the least attention and may matter most. When professors graded blind, they had fewer safety concerns about AI answers than about their colleagues’ answers. That runs directly against the popular intuition that AI will mislead students.

Which Systems Were Tested

The study examined commercial tutoring systems and Google’s NotebookLM, with varying performance levels across systems. The press release does not name a full model list; the paper is available as a PDF and via SSRN (abstract 6849678). The notable detail: these are off-the-shelf products, not specially tuned models. Any student can use the same tools today, which makes the finding feel a lot less hypothetical than a lab evaluation.

Limits and the Authors’ Own Caveats

Nyarko was careful to stress that the study evaluated answer quality only — “how to implement these tools to most effectively improve student learning is still an open question.” In plain terms: AI answers well; that does not mean students using AI learn more. The blind design also has boundaries. All 40 questions sit in a single subject, contracts; 16 professors cannot represent all of legal education, let alone other disciplines.

The study landed hard on Hacker News after publication (417 points, 356 comments), and most of the argument there was about the same gap: answer quality versus educational outcome.

What It Means for EdTech

Three observations. First, introductory-level Q&A is likely the first place AI gets adopted at scale in higher education — it is high-frequency, time-consuming for instructors, and has a blind-testable quality signal. Second, once AI answers beat the instructor’s own in blind grading, the institutional question shifts from “should we ban this” to “how do we redesign tutoring and assessment” — a ban will not stop tools every student already has. Third, the appearance of NotebookLM — a system grounded in course materials — in a rigorous study means “AI plus your own materials” now has measurable evidence behind it, not just marketing copy. The study does not settle whether AI belongs in the classroom, and its authors say so plainly. What it does is move the burden of proof: from here on, the case against AI tutoring has to be argued with data, not intuition.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL