An empirical analysis of six generative artificial intelligence models evaluated against human assessment standards
Marked by the Machine
We tested major AI models on their ability to accurately grade student work. Pulling from a dataset of the ASAP corpus — real student essays and short answers in science and English, scored against the same official rubrics used by human raters in the Hewlett Foundation’s original grading competition — we found that the leading AI models are effectively able to grade both long and short student materials. While it can’t and shouldn’t replace the personalized feedback that teachers can give their students, it can reduce the grading burden by automating quantitative tasks like how many errors were present? or did the student cite this specific fact?
Crucially, AI is not a single tool, and lumping all models together pollutes the data. Free or lower-tier models (like GPT-4o mini, Gemini 2.5 Flash, and Claude Haiku) completely fail at grading subjective essays, consistently grading them several points too low. For subjective essay grading, only premium flagship models (like Claude Opus and GPT-5.5) work, and they work exceptionally well, matching human standards. On the other hand, for objective content checking, almost any model can be used.
How to read the data: each model shows two scores: objective content checking (black) and subjective essay grading (coral). The horizontal lines show how well two human teachers agree with each other. All six models cluster tightly on objective content — but tight clustering is not the same as matching humans. On essays, only the two paid flagship models get close to the human line; cheaper models collapse.
Accuracy,calibration,andreliability:threedistinctwaystomeasureamodel.
Before you can trust a model with a single grade, judge it on three separate questions. A good model needs to be accurate (close to humans), calibrated (not systematically too harsh or too soft), and reliable (giving the same score to the same work twice). A model can excel in one area while struggling in another — this is what separates the models that work from the ones that don’t.
1. Accuracy
Accuracy is simply how close a model’s grade lands to a human grader’s. We measure it with quadratic weighted kappa (QWK), where 1.0 is perfect agreement and 0 is random.
- Claude Opus 4.80.66
- GPT-5.50.58
- Claude Haiku 4.50.53
- Gemini 2.5 Pro0.49
- GPT-4o mini0.47
- Gemini 2.5 Flash0.46
Overall agreement (QWK) with human graders, by model. The coral line marks human-to-human agreement (0.79) — the ceiling no AI model reaches. The best model, Claude Opus, closes about two-thirds of that gap.
2. Calibration
Calibration is direction, not distance: does a model land in the right place on average, or does it consistently grade too harsh or too lenient, even when its overall accuracy looks fine?
- Claude Opus 4.8+0.11
- GPT-5.5+0.02
- Claude Haiku 4.5-0.95
- Gemini 2.5 Pro-1.44
- GPT-4o mini-0.70
- Gemini 2.5 Flash-1.11
Calibration bias by model: bars left of center grade harsher than humans; bars to the right are more lenient. The two paid flagships sit close to zero. Every cheaper model is meaningfully harsh — Gemini 2.5 Flash by more than a full point on average.
3. Reliability
Reliability measures how consistently a model repeats its own grade when it sees the same submission more than once. It is not the same as accuracy — a model can be perfectly consistent and consistently wrong — but an assistant that changes its mind every time you ask it is not one you can build a workflow around.
- Claude Opus 4.893%
- GPT-5.584%
- Claude Haiku 4.590%
- Gemini 2.5 Pro71%
- GPT-4o mini81%
- Gemini 2.5 Flash69%
Share of trials where a model repeated the exact same score on the same submission, across three runs. Claude is the steadiest at 93%; Gemini 2.5 Pro is the least steady at 69% — and its figures here are provisional, since it hit API rate limits and missed four short-answer sets.
Twodistinctgradingtasksyieldtwocompletelydifferentresults.
Objective grading asks simple questions: is the required content there? This includes checking if a student correctly named the steps of a cell cycle or cited specific evidence. There is a clear right answer to verify.
Subjective grading asks complex questions: how good is the writing? This involves assessing essay structure, tone, and narrative flow. These are qualitative calls where even experienced teachers can disagree.
Look closely at the graph in the hero section. For objective tasks, even the most affordable models align closely with each other, landing in a tight band around 0.52 to 0.66. While this is highly consistent, it is important to note that none of them actually reach the 0.832 human-to-human ceiling.
This means that while the models are highly consistent with each other, they are still not quite matching human standards. This gap exists because models have a major blind spot: they tend to over-credit short answers, awarding points for correct keywords even when the student's reasoning is flawed.
Short-Answer Disagreement Breakdown (Objective Grading)
Based on 1,310 disagreement cases between AI models and human teachers. When models disagreed with a human on short-answer content questions, they awarded unearned points (over-credited) 68% of the time, usually by accepting correct keywords in flawed or contradictory student reasoning.
AIishighlyreliableasacontentchecker,butlessconsistentasanessayjudge.
When you line up tasks from simple to complex, a clear trend emerges: the more direct and structured the answer, the closer AI gets to a human. As tasks require wider, more holistic judgment, the model scores begin to drift.
Note that the essay grading scores in this chart combine all six models, which lowers the overall average. This combines both premium flagship models and cheaper tiers, masking the strong performance of premium models shown above.
Short content answers
objective0.58QWK · 61% exact · +0.2 lenientAll six configurations demonstrate closely clustered agreement scores ranging from 0.51 to 0.63, indicating high baseline consistency and a slight tendency toward liberal calibration.
Source-based reading
objective0.57QWK · 44% exact · -0.3 harshThis category measures reading comprehension tied to passage evidence, showing reliable performance across all model tiers, with lower exact agreement scores resulting from the wider point scales utilized.
Holistic essays
subjective0.35QWK · 28% exact · -1.2 harshWhile the combined average appears low, this macro figure obscures a significant divergence between premium flagship architectures, which achieve agreement scores up to 0.68, and lower-tier models, which range between 0.08 and 0.25.
Trait-summed essays
subjective0.44QWK · 6% exact · -4.6 harshEvaluating high-range multidimensional scales leads to extremely low exact agreement percentages, with automated ratings exhibiting a conservative calibration bias of approximately 4.6 points.
Workflow recommendation matrix
Short-answer content assessments, factual rubrics, and structured laboratory conclusions
All model tiers demonstrate approximately 60% exact agreement and high trial reliability, making these tasks highly suitable for automated baseline scoring.
Source-grounded reading comprehension and text analyses
These tasks are viable for automated assistance, although agreement scores exhibit wider variance and show a slight conservative calibration bias on larger scales.
Holistic writing essays and qualitative composition evaluations
Advanced flagship architectures are recommended for these tasks, and only as a secondary reference, as lower-tier systems show significant evaluation variance and systematic under-evaluation bias.
Priceandaccuracydonotscaleonthesameaxis.
You do not need to pay for the most expensive model for every task. Cheap models handle objective content grading exceptionally well. Paying for premium flagship models is only necessary when you need balanced calibration and essay evaluation.
Price vs. accuracy: the horizontal axis shows grading cost per assignment in cents on a linear scale; the vertical axis shows accuracy. The free and budget models bunch near zero — all under a quarter of a cent — while the two flagships sit out to the right at roughly 1.3¢ and 1.6¢, meaning higher cost mainly purchases essay evaluation and better calibration rather than basic content checking.
At the per-assignment level the absolute costs remain marginal: the most expensive flagship grades a single response for approximately 1.6¢, while the most economical model does so for roughly 0.02¢ — under two cents per assignment even at the upper bound.
Verbatimexplanationsshowhowthemodelsarriveattheirscores.
Numeric scores only tell part of the story. When automated models disagree with human teachers, they follow distinct patterns: they are consistently too harsh on essays by over-penalizing mechanical errors, while over-crediting short-answer content by accepting key terms without sound reasoning.
These four real samples show student responses alongside the actual explanations written by the models, offering a transparent look at what happens when you delegate grading tasks.
When models are wrong, in which direction?
- Trait-summed essays100% harsh · avg gap -10.53
- Holistic essays100% harsh · avg gap -3.19
- Source-based reading67% harsh · avg gap -0.48
- Short content answers32% harsh · avg gap +0.42
Direction of error: of the times a model disagreed with a human, the coral bar shows how often it graded too low. Models are almost always too harsh on essays, but they lean more lenient and over-credit short content answers.
Four real evaluations in the models' own words
Biology · protein-synthesis short answer
“Prophase, metaphase, anaphase, and telephase.”
AI’s reasoning
“The response lists stages of mitosis rather than steps in protein synthesis. It does not include any rubric criteria such as mRNA exiting the nucleus, traveling to a ribosome, codon/anticodon matching, or amino acids being linked into a protein.”
What this shows · For objective assessments, automated systems demonstrate high reliability in identifying irrelevant responses, aligned with human reasoning and specific rubric indicators.
ELA · “computers” persuasive essay
“Dear [editor], I understand you think computers don’t have an effect on people. Computers do have an effect on people. For one reason the computer helps people find out information about anything. Also, it helps people do work. Plus, you can communicate with anyone you want…”
AI’s reasoning
“Takes a clear position and is well organized with an introduction, three reasons (information, work, communication), supporting details, and a conclusion, showing good audience awareness. However, elaboration is somewhat repetitive and general rather than deeply developed, placing it in the competent middle range rather than the strongest score points.”
What this shows · A premium flagship model aligns perfectly with human holistic evaluations, detailing qualitative trade-offs with high academic clarity.
ELA · “patience” narrative essay
“[Look out!] [someone] squealed. The sound of crunching metal was deafening. There was no way out. My eyes seemed sealed shut, but they were as wide as could be. The eerie silence frightened me…”
AI’s reasoning
“Somewhat focused ideas about patience but lacks specific details — 2 points for Ideas. The organization is logical but connections could be clearer — 2 points. Language use is adequate — 2 points for Style…”
What this shows · Lower-tier models can systematically under-evaluate narrative complexity, highlighting the need to avoid budget configurations for holistic writing assessments.
Science · dog-house lid-color item
“white :: [They] would like the dog to stay warm in his dog house and white doesn’t allow that much air to come in.”
AI’s reasoning
“The student chose a valid color (white), satisfying one criterion, but the reasoning is flawed — they claim white keeps the dog warm, which contradicts the experimental data showing white had the lowest temperature (42°C)… it does not earn full marks.”
What this shows · Under certain conditions, a model may over-credit a response that contains a correct keyword despite flawed reasoning, which highlights the importance of periodic evaluation checks.
Apragmaticchecklistforintegratingautomatedgrading.
These findings do not suggest avoiding automated grading altogether. Instead, they outline a balanced strategy: deploy AI where it earns trust, and keep your hands on evaluations that require professional human judgment.
Deploy automated grading for objective content verification
Any model configuration is highly reliable for short-answer items, checklists, and laboratory conclusions, which allows educators to establish rapid baseline assessments.
Reserve holistic essay grading for flagship models with human moderation
Qualitative writing evaluation requires premium model architectures to align with human standards, and should only serve as a secondary review to support final grading.
Account for conservative calibration bias in lower-tier models
Free or lower-tier models systematically under-evaluate essay composition, meaning that lower grades often reflect model calibration rather than student proficiency.
Moderate evaluations for potential over-crediting in short answers
Automated systems may occasionally award points to short answers with correct keywords despite flawed logical reasoning, which requires periodic validation of high scores.
Reclaiming grading hours is not about grading more; it is about spending that saved time where only you can make a difference—feedback, direct conferencing, and planning. This is the heart of the 3X Framework, and you can explore detailed step-by-step procedures in the Grading Assistant Guide.
How we ran the study, and where it is limited
Six models graded real student work against official rubrics, and we compared every score to human consensus. Five parameters define the scope and limits of these findings.
Standardized student work and official rubrics
This study utilized the public ASAP corpus, which includes student essays and short answers across grades seven to ten in science and English, evaluated using official prompts and rubrics identical to those used by human raters.
Multi-trial model configurations
Six configurations representing free and premium tiers from three major providers evaluated every item over three independent trials, measuring both raw performance and self-consistency.
Model tiers as product proxies
The study assessed core model architectures rather than consumer-facing chat interfaces, meaning that these results reflect underlying model performance rather than specific product optimizations.
Holistic evaluation aggregation
To establish a clear comparative metric, models were prompted to generate a single overall score even where rubrics specified multiple traits, which supports holistic consistency but increases variance on larger scales.
Identified coverage constraints
Gemini 2.5 Pro experienced API request limitations on four short-answer subsets, which makes those specific metrics preliminary and calculated solely on completed evaluations.
Reading the numbers
- QWK (quadratic weighted kappa)
- The standard statistical metric for measuring evaluation agreement, where a score of 1 indicates perfect agreement and 0 represents random alignment, with human-to-human agreement in this study achieving 0.79.
- Exact %
- The percentage of trials where the model's score was identical to the human score, which is a highly stringent metric on wide evaluation scales.
- Bias
- The average direction of evaluation variance, where negative values indicate a conservative bias that grades lower than humans and positive values indicate a liberal bias.
- Self-consistency
- The reliability metric measuring how frequently a model generates the identical score across three separate evaluations of the same student submission.
- Human ceiling
- The level of agreement achieved between two independent human evaluators, representing the practical upper limit of assessment consistency.
What the study can and can’t support
Data & sources
- ASAP-AES — Automated Student Assessment Prize (Hewlett Foundation / Kaggle, 2012)
- ASAP-SAS — Short Answer Scoring (Hewlett Foundation / Kaggle, 2012)
- Quadratic weighted kappa — the ASAP scoring metric
Total API cost across all 6 models, every item graded three times: $35.81.
AI can lighten the grading load — but the judgment that matters most, on a student’s writing, is still yours. Use it as a reader, not a replacement.