Skip to main content

An empirical analysis of six generative artificial intelligence models evaluated against human assessment standards

Marked by the Machine

We tested major AI models on their ability to accurately grade student work. Pulling from a dataset of the ASAP corpus — real student essays and short answers in science and English, scored against the same official rubrics used by human raters in the Hewlett Foundation’s original grading competition — we found that the leading AI models are effectively able to grade both long and short student materials. While it can’t and shouldn’t replace the personalized feedback that teachers can give their students, it can reduce the grading burden by automating quantitative tasks like how many errors were present? or did the student cite this specific fact?

Crucially, AI is not a single tool, and lumping all models together pollutes the data. Free or lower-tier models (like GPT-4o mini, Gemini 2.5 Flash, and Claude Haiku) completely fail at grading subjective essays, consistently grading them several points too low. For subjective essay grading, only premium flagship models (like Claude Opus and GPT-5.5) work, and they work exceptionally well, matching human standards. On the other hand, for objective content checking, almost any model can be used.

Fig. 01
00.250.50.751content ceiling 0.832essay ceiling 0.651.55.67GPT-5.5paid.66.67Claudepaid.54.26GPT-4ofree.60.26Claudefree.58.24Geminipaid.52.24Geminifree
Objective — is the content there?Subjective — how good is the writing?

How to read the data: each model shows two scores: objective content checking (black) and subjective essay grading (coral). The horizontal lines show how well two human teachers agree with each other. All six models cluster tightly on objective content — but tight clustering is not the same as matching humans. On essays, only the two paid flagship models get close to the human line; cheaper models collapse.

Scroll

Accuracy,calibration,andreliability:threedistinctwaystomeasureamodel.

Before you can trust a model with a single grade, judge it on three separate questions. A good model needs to be accurate (close to humans), calibrated (not systematically too harsh or too soft), and reliable (giving the same score to the same work twice). A model can excel in one area while struggling in another — this is what separates the models that work from the ones that don’t.

1. Accuracy

Accuracy is simply how close a model’s grade lands to a human grader’s. We measure it with quadratic weighted kappa (QWK), where 1.0 is perfect agreement and 0 is random.

How to read it · Accuracy
Fig. 02
best in testflagship (paid)budget / free

Overall agreement (QWK) with human graders, by model. The coral line marks human-to-human agreement (0.79) — the ceiling no AI model reaches. The best model, Claude Opus, closes about two-thirds of that gap.

2. Calibration

Calibration is direction, not distance: does a model land in the right place on average, or does it consistently grade too harsh or too lenient, even when its overall accuracy looks fine?

How to read it · Calibration
Fig. 03
← harsherhuman-calibratedmore lenient →
  • Claude Opus 4.8
    +0.11
  • GPT-5.5
    +0.02
  • Claude Haiku 4.5
    -0.95
  • Gemini 2.5 Pro
    -1.44
  • GPT-4o mini
    -0.70
  • Gemini 2.5 Flash
    -1.11
0 · matches humans

Calibration bias by model: bars left of center grade harsher than humans; bars to the right are more lenient. The two paid flagships sit close to zero. Every cheaper model is meaningfully harsh — Gemini 2.5 Flash by more than a full point on average.

3. Reliability

Reliability measures how consistently a model repeats its own grade when it sees the same submission more than once. It is not the same as accuracy — a model can be perfectly consistent and consistently wrong — but an assistant that changes its mind every time you ask it is not one you can build a workflow around.

How to read it · Reliability
Fig. 04
flagship (paid)budget / free

Share of trials where a model repeated the exact same score on the same submission, across three runs. Claude is the steadiest at 93%; Gemini 2.5 Pro is the least steady at 69% — and its figures here are provisional, since it hit API rate limits and missed four short-answer sets.

Twodistinctgradingtasksyieldtwocompletelydifferentresults.

Objective grading asks simple questions: is the required content there? This includes checking if a student correctly named the steps of a cell cycle or cited specific evidence. There is a clear right answer to verify.

Subjective grading asks complex questions: how good is the writing? This involves assessing essay structure, tone, and narrative flow. These are qualitative calls where even experienced teachers can disagree.

Look closely at the graph in the hero section. For objective tasks, even the most affordable models align closely with each other, landing in a tight band around 0.52 to 0.66. While this is highly consistent, it is important to note that none of them actually reach the 0.832 human-to-human ceiling.

This means that while the models are highly consistent with each other, they are still not quite matching human standards. This gap exists because models have a major blind spot: they tend to over-credit short answers, awarding points for correct keywords even when the student's reasoning is flawed.

Fig. 05

Short-Answer Disagreement Breakdown (Objective Grading)

68%of disagreements gave the student unearned credit
Over-credited (lenient)68%Harsh32%

Based on 1,310 disagreement cases between AI models and human teachers. When models disagreed with a human on short-answer content questions, they awarded unearned points (over-credited) 68% of the time, usually by accepting correct keywords in flawed or contradictory student reasoning.

AIishighlyreliableasacontentchecker,butlessconsistentasanessayjudge.

When you line up tasks from simple to complex, a clear trend emerges: the more direct and structured the answer, the closer AI gets to a human. As tasks require wider, more holistic judgment, the model scores begin to drift.

Note that the essay grading scores in this chart combine all six models, which lowers the overall average. This combines both premium flagship models and cheaper tiers, masking the strong performance of premium models shown above.

Fig. 06
  1. Short content answers

    objective
    0.58QWK · 61% exact · +0.2 lenient

    All six configurations demonstrate closely clustered agreement scores ranging from 0.51 to 0.63, indicating high baseline consistency and a slight tendency toward liberal calibration.

  2. Source-based reading

    objective
    0.57QWK · 44% exact · -0.3 harsh

    This category measures reading comprehension tied to passage evidence, showing reliable performance across all model tiers, with lower exact agreement scores resulting from the wider point scales utilized.

  3. Holistic essays

    subjective
    0.35QWK · 28% exact · -1.2 harsh

    While the combined average appears low, this macro figure obscures a significant divergence between premium flagship architectures, which achieve agreement scores up to 0.68, and lower-tier models, which range between 0.08 and 0.25.

  4. Trait-summed essays

    subjective
    0.44QWK · 6% exact · -4.6 harsh

    Evaluating high-range multidimensional scales leads to extremely low exact agreement percentages, with automated ratings exhibiting a conservative calibration bias of approximately 4.6 points.

Workflow recommendation matrix

Short-answer content assessments, factual rubrics, and structured laboratory conclusions

All model tiers demonstrate approximately 60% exact agreement and high trial reliability, making these tasks highly suitable for automated baseline scoring.

Source-grounded reading comprehension and text analyses

These tasks are viable for automated assistance, although agreement scores exhibit wider variance and show a slight conservative calibration bias on larger scales.

Holistic writing essays and qualitative composition evaluations

Advanced flagship architectures are recommended for these tasks, and only as a secondary reference, as lower-tier systems show significant evaluation variance and systematic under-evaluation bias.

Priceandaccuracydonotscaleonthesameaxis.

You do not need to pay for the most expensive model for every task. Cheap models handle objective content grading exceptionally well. Paying for premium flagship models is only necessary when you need balanced calibration and essay evaluation.

Fig. 07
00.250.50.7500.5¢1.5¢average cost per graded assignment (¢)two humans 0.792Claude Opus 4.8GPT-5.5Claude Haiku 4.5Gemini 2.5 ProGPT-4o miniGemini 2.5 Flash

Price vs. accuracy: the horizontal axis shows grading cost per assignment in cents on a linear scale; the vertical axis shows accuracy. The free and budget models bunch near zero — all under a quarter of a cent — while the two flagships sit out to the right at roughly 1.3¢ and 1.6¢, meaning higher cost mainly purchases essay evaluation and better calibration rather than basic content checking.

At the per-assignment level the absolute costs remain marginal: the most expensive flagship grades a single response for approximately 1.6¢, while the most economical model does so for roughly 0.02¢ — under two cents per assignment even at the upper bound.

Verbatimexplanationsshowhowthemodelsarriveattheirscores.

Numeric scores only tell part of the story. When automated models disagree with human teachers, they follow distinct patterns: they are consistently too harsh on essays by over-penalizing mechanical errors, while over-crediting short-answer content by accepting key terms without sound reasoning.

These four real samples show student responses alongside the actual explanations written by the models, offering a transparent look at what happens when you delegate grading tasks.

Fig. 08

When models are wrong, in which direction?

graded too harshtoo lenient
  • Trait-summed essays100% harsh · avg gap -10.53
  • Holistic essays100% harsh · avg gap -3.19
  • Source-based reading67% harsh · avg gap -0.48
  • Short content answers32% harsh · avg gap +0.42

Direction of error: of the times a model disagreed with a human, the coral bar shows how often it graded too low. Models are almost always too harsh on essays, but they lean more lenient and over-credit short content answers.

Four real evaluations in the models' own words

✓ agreesobjective·GPT-5.5 · paid

Biology · protein-synthesis short answer

AI 0/human 00 to 3 scale indicating exact agreement
Prophase, metaphase, anaphase, and telephase.

AI’s reasoning

The response lists stages of mitosis rather than steps in protein synthesis. It does not include any rubric criteria such as mRNA exiting the nucleus, traveling to a ribosome, codon/anticodon matching, or amino acids being linked into a protein.

What this shows · For objective assessments, automated systems demonstrate high reliability in identifying irrelevant responses, aligned with human reasoning and specific rubric indicators.

✓ agreessubjective·Claude Opus 4.8 · paid

ELA · “computers” persuasive essay

AI 8/human 82 to 12 scale indicating exact agreement
Dear [editor], I understand you think computers don’t have an effect on people. Computers do have an effect on people. For one reason the computer helps people find out information about anything. Also, it helps people do work. Plus, you can communicate with anyone you want…

AI’s reasoning

Takes a clear position and is well organized with an introduction, three reasons (information, work, communication), supporting details, and a conclusion, showing good audience awareness. However, elaboration is somewhat repetitive and general rather than deeply developed, placing it in the competent middle range rather than the strongest score points.

What this shows · A premium flagship model aligns perfectly with human holistic evaluations, detailing qualitative trade-offs with high academic clarity.

⚠ disagreessubjective·GPT-4o mini · free

ELA · “patience” narrative essay

AI 14/human 2430-point scale indicating conservative variance of 10 points
[Look out!] [someone] squealed. The sound of crunching metal was deafening. There was no way out. My eyes seemed sealed shut, but they were as wide as could be. The eerie silence frightened me…

AI’s reasoning

Somewhat focused ideas about patience but lacks specific details — 2 points for Ideas. The organization is logical but connections could be clearer — 2 points. Language use is adequate — 2 points for Style…

What this shows · Lower-tier models can systematically under-evaluate narrative complexity, highlighting the need to avoid budget configurations for holistic writing assessments.

⚠ disagreesobjective·Claude Opus 4.8 · paid

Science · dog-house lid-color item

AI 1/human 00 to 2 scale indicating liberal variance of 1 point
white :: [They] would like the dog to stay warm in his dog house and white doesn’t allow that much air to come in.

AI’s reasoning

The student chose a valid color (white), satisfying one criterion, but the reasoning is flawed — they claim white keeps the dog warm, which contradicts the experimental data showing white had the lowest temperature (42°C)… it does not earn full marks.

What this shows · Under certain conditions, a model may over-credit a response that contains a correct keyword despite flawed reasoning, which highlights the importance of periodic evaluation checks.

Apragmaticchecklistforintegratingautomatedgrading.

These findings do not suggest avoiding automated grading altogether. Instead, they outline a balanced strategy: deploy AI where it earns trust, and keep your hands on evaluations that require professional human judgment.

1

Deploy automated grading for objective content verification

Any model configuration is highly reliable for short-answer items, checklists, and laboratory conclusions, which allows educators to establish rapid baseline assessments.

2

Reserve holistic essay grading for flagship models with human moderation

Qualitative writing evaluation requires premium model architectures to align with human standards, and should only serve as a secondary review to support final grading.

3

Account for conservative calibration bias in lower-tier models

Free or lower-tier models systematically under-evaluate essay composition, meaning that lower grades often reflect model calibration rather than student proficiency.

4

Moderate evaluations for potential over-crediting in short answers

Automated systems may occasionally award points to short answers with correct keywords despite flawed logical reasoning, which requires periodic validation of high scores.

Reclaiming grading hours is not about grading more; it is about spending that saved time where only you can make a difference—feedback, direct conferencing, and planning. This is the heart of the 3X Framework, and you can explore detailed step-by-step procedures in the Grading Assistant Guide.

How we ran the study, and where it is limited

Six models graded real student work against official rubrics, and we compared every score to human consensus. Five parameters define the scope and limits of these findings.

Standardized student work and official rubrics

This study utilized the public ASAP corpus, which includes student essays and short answers across grades seven to ten in science and English, evaluated using official prompts and rubrics identical to those used by human raters.

Multi-trial model configurations

Six configurations representing free and premium tiers from three major providers evaluated every item over three independent trials, measuring both raw performance and self-consistency.

Model tiers as product proxies

The study assessed core model architectures rather than consumer-facing chat interfaces, meaning that these results reflect underlying model performance rather than specific product optimizations.

Holistic evaluation aggregation

To establish a clear comparative metric, models were prompted to generate a single overall score even where rubrics specified multiple traits, which supports holistic consistency but increases variance on larger scales.

Identified coverage constraints

Gemini 2.5 Pro experienced API request limitations on four short-answer subsets, which makes those specific metrics preliminary and calculated solely on completed evaluations.

Reading the numbers

QWK (quadratic weighted kappa)
The standard statistical metric for measuring evaluation agreement, where a score of 1 indicates perfect agreement and 0 represents random alignment, with human-to-human agreement in this study achieving 0.79.
Exact %
The percentage of trials where the model's score was identical to the human score, which is a highly stringent metric on wide evaluation scales.
Bias
The average direction of evaluation variance, where negative values indicate a conservative bias that grades lower than humans and positive values indicate a liberal bias.
Self-consistency
The reliability metric measuring how frequently a model generates the identical score across three separate evaluations of the same student submission.
Human ceiling
The level of agreement achieved between two independent human evaluators, representing the practical upper limit of assessment consistency.

What the study can and can’t support

Automated systems grade objective, content-based short answers with a level of consistency that aligns with baseline task constraints across all tested models.
Premium flagship models evaluate holistic essay quality with a level of consistency that matches a secondary human evaluator.
Lower-tier models exhibit significant evaluation variance on holistic essays and systematically under-evaluate submissions.
No single model matched the overall consensus agreement score achieved between two human raters.
Evaluators can assume uniform performance across all automated assessment tools.
Free or lower-tier automated systems can be trusted to evaluate holistic writing quality.
Automated models can entirely replace human evaluation on subjective writing tasks.

AI can lighten the grading load — but the judgment that matters most, on a student’s writing, is still yours. Use it as a reader, not a replacement.