Explore state by state cost analysis of US colleges in an interactive article

Can AI Assess Critical Thinking? Accuracy, Fairness, Limits

Sep 2, 2026
9 minute read

Can AI assess critical thinking? Accuracy, fairness, limits

A DistilBERT-based essay-scoring model recently matched human raters closely enough to post a macro-F1 score of .87, a measure of how well the model's predicted score categories lined up with the categories human graders actually assigned. On test data spanning eight essay topics, its scores correlated with human scores at levels between .65 and .85, meaning the AI's ranking of essays tracked closely, though not perfectly, with how human raters ranked the same essays (PLOS One, last month). That kind of result is why "Can AI assess critical thinking?" has stopped being a theoretical question for college instructors, instructional designers, and writing-program administrators currently evaluating AI scoring tools.

The same study found something less reassuring. When researchers compared human and AI scores across different essay topics, they found indications of fairness violations, meaning the tool didn't treat every topic the same way (PLOS One, last month). A separate paper, "Should Machines Get to Judge?," published three months ago, goes further, describing bias, opacity, reliability gaps, and accountability voids as persistent, unresolved problems in how these systems get deployed.

Several of the studies behind this article focus specifically on higher education and on written responses, not oral exams, group projects, class discussion, or K-12 assessment, so the findings below apply most directly to that context. This piece breaks the accuracy-and-fairness question into its working parts: what these tools are actually measuring, how well that measurement holds up under scrutiny, whether it's fair across different students and topics, and what role AI should realistically play in evaluating critical thinking right now.

Advertisement

How does AI assess critical thinking skills?

Automated essay scoring works by training a model, in this case a DistilBERT architecture, on large sets of essays that human raters have already scored. The model learns statistical patterns associated with each score level and applies those patterns to new, unseen essays. Recent advances in large language models have made this kind of pattern-based evaluation feasible for open-ended, non-numerical writing (PLOS One, last month).

The same DistilBERT study shows how uneven that pattern-matching gets once you look past the headline number. Individual topic-level F1 scores in the test set ranged from .75 to .95, even though the model's overall macro-F1 landed at .87 (PLOS One, last month). That spread is an early sign that "accurate overall" and "accurate on every essay prompt" aren't the same claim, a distinction the fairness section below digs into.

The scoring problem gets harder to pin down from there. Standardized ability tests are built from discrete items, each with a defined right answer or scoring key. Essays don't work that way. There's no single "item" to validate against, which creates evaluation challenges that don't apply to a multiple-choice exam or a personality questionnaire (PLOS One, last month). A model can predict the score a human rater would assign without anyone being fully sure what feature of the essay is driving that prediction: sentence complexity, vocabulary range, topic familiarity, or actual reasoning quality.

Consider a persuasive essay arguing that cities should ban single-use plastics. It cites two sources, raises a counterargument about cost to small businesses, and closes with a clear thesis restatement. Structurally, it checks the boxes a rubric might list for "uses evidence" and "addresses counterarguments." Whether the student's counterargument actually engages the strongest version of the opposing view, or just gestures at it before dismissing it, is a distinction a pattern-matching model may not reliably capture. That gap between structural completeness and reasoning quality sits close to the center of the debate over whether these tools measure critical thinking or something that just resembles it on the page.

For a teacher or program evaluating a scoring platform, the practical move is asking what the model was trained on and whether its scoring categories map to your rubric's definition of critical thinking, not just to general writing quality. A tool trained to reward five-paragraph structure and topic sentences isn't the same as a tool trained to reward sound inference.

How accurate is AI-based critical thinking assessment?

Raw agreement with human scores is only part of the picture. Researchers increasingly evaluate AI scores against two separate testing standards: reliability and validity (PLOS One, last month). Reliability refers to the precision of measurement, whether the tool produces consistent scores without large measurement error. Validity asks a different question: does evidence actually support the interpretation you want to draw from the score, in this case, that a high score reflects strong critical thinking rather than something else?

Advertisement

On reliability, the PLOS One model performed well, showing high internal consistency with Spearman-Brown coefficients, a measure of internal consistency, ranging from .77 to .92. The study's authors described the model as strong and reported evidence for "several important aspects of validity" (PLOS One, last month). Those are the researchers' own conclusions about their own model, worth noting as such rather than as an independent, field-wide verdict.

A much larger review tells a more cautious story. A 74-study systematic review of generative AI in second-language writing, feedback, and assessment identifies construct validity, the question of what a scoring system is actually measuring, as a central, unresolved tension running through the field (Indian Journal of Language and Linguistics, 2 months ago). The review also flags recurring methodological weaknesses across the literature: small samples, short observation windows, and thin theoretical grounding, all of which limit how far any single accuracy finding should be generalized beyond its own study population.

Put those findings together and a pattern emerges. A tool can be statistically accurate at reproducing scores a human grader would have assigned anyway, while still failing to demonstrate it's measuring critical thinking rather than a correlated stand-in, like sentence length, vocabulary sophistication, or how closely an essay's topic matches material the model saw during training. High agreement with human raters is necessary evidence. It isn't sufficient proof on its own.

Instructors evaluating any AI grading tool should ask the vendor directly what validity evidence exists beyond correlation with human raters, and whether that evidence has gone through independent peer review rather than appearing only in the product's own materials.

Can AI score essays for critical thinking fairly?

Fairness, in testing terms, isn't a vague impression of even-handedness. It's a defined standard: a fair test "reflects the same construct(s) for all test takers," and doesn't advantage or disadvantage people based on traits that have nothing to do with what's actually being measured (PLOS One, last month).

This is where the PLOS One findings get complicated, because the study reports two different results depending on how the data is sliced. Looking at aggregate predictions overall, the researchers found no systematic bias. But when they compared human and AI scores across different essay topics specifically, they found indications of fairness violations (PLOS One, last month). A tool can show strong agreement with human ratings overall while still producing unequal results across topics. Aggregate fairness and contextual fairness are not the same claim, and a school relying only on an overall accuracy number could miss a topic-level problem sitting underneath it.

Advertisement

A separate review of AI-mediated classroom assessment frames this as more than a technical glitch. Its authors describe bias, opacity, reliability gaps, and accountability voids as persistent structural issues, arguing that handing evaluative judgment to language models isn't a neutral efficiency upgrade. It represents a redistribution of power within education, shifting who decides what counts as good thinking away from the people directly accountable to students, the same review argues.

The L2-writing review adds a dimension easy to overlook in a purely statistical conversation: generative AI assessment tools raise equity concerns specifically for multilingual learners and students from under-resourced settings, alongside risks to individual writer voice and identity when scoring rewards conformity to a particular writing style over a student's own voice (Indian Journal of Language and Linguistics, 2 months ago).

The confirmed finding here is specific: one study found indications of fairness violations when comparing human and AI scores across essay topics. It does not establish that a given topic, dialect, or writing style is misjudged every time it appears. What it does establish is that schools should independently test relevant prompts and student populations before using AI scores for anything consequential, rather than trusting an aggregate accuracy figure to stand in for topic-level or population-level fairness.

What role should AI play in assessing critical thinking right now?

A systematic review of how generative AI affects critical thinking in higher education, based on records from Scopus and Web of Science covering 2020 through 2025, identifies two distinct patterns of use (AIS eLibrary systematic review, 4 months ago).

The first, a "cognitive crutch" pathway, involves answer-first, product-oriented AI use. It's associated with cognitive offloading and overreliance, reduced effort in checking one's own reasoning, and weaker verification practices, with downstream risks to argument quality and source evaluation. The second, a "cognitive coach" pathway, involves process-first designs: Socratic questioning, counterargument routines, verification prompts, and reflective checkpoints. This pattern aligns more consistently with stronger epistemic vigilance, better calibration, and higher-quality justification and error detection (AIS eLibrary, 4 months ago).

That distinction matters for anyone designing how AI gets used in a classroom, but the same review adds a caution worth carrying forward: the specific mechanisms behind these effects, automation bias, calibration, verification behavior, are rarely measured directly in the current literature. That gap limits how confidently anyone can claim AI reliably improves critical thinking, rather than simply correlating with settings where improvement happens for other reasons.

Advertisement

Researchers behind the same "Should Machines Get to Judge?" review propose a specific design principle: treat AI's role in assessment as ongoing human-and-AI negotiation, prioritizing pedagogical value and human oversight rather than letting a system issue automated, final decisions. UNESCO's public position lines up with that framing from a policy angle: critical thinking and creativity remain skills that no technology can replace, a stance that's less an empirical finding than a boundary UNESCO is drawing around how schools should think about AI's role in evaluating those exact skills (UNESCO, 6 weeks ago).

The evidence supports a fairly specific recommendation: AI currently works best as a formative tool, generating counterargument prompts, feedback drafts, or verification questions for students to work through, with a human reviewing anything that affects a grade, placement, or credential. That's narrower than "AI grades critical thinking," but it's the role the research actually backs.

What students should ask when AI scores their work

If an AI tool is grading or scoring an assignment tied to a grade, students and families have a reasonable set of questions to bring to a teacher, professor, or program before assuming the score works like a human grade:

  • Is the AI score final, or does a human instructor review it and can override it?
  • Does the assignment prompt or course syllabus disclose that AI is involved in scoring?
  • Is there a way to request a human re-read if the score seems off, especially on an essay covering an unfamiliar or personal topic?
  • Does using outside help or drafting with a writing tool count differently than the AI scoring process itself, and where is that boundary written down?

These are fair questions for any course policy, not challenges to a teacher's judgment. Asking them helps confirm how the grade was actually determined.

Questions to ask before adopting an AI-scored assignment

Before a school, program, or individual instructor adopts an AI-scored assignment, a few direct questions separate a defensible tool from a marketing claim:

  • Was the tool validated on students and writing prompts comparable to the ones it will actually score, not just on a different sample or subject area?
  • Can a human reviewer see and override any score that affects a grade, placement, or credential decision?
  • Do students know when and how AI is involved in scoring their work?
  • Does the institution's academic policy permit AI-assisted grading for this type of assignment, confirmed rather than assumed?

Any single "no" is reason enough to slow down before treating an AI score as equivalent to a human one.

Advertisement

The decision this research actually supports

AI can approximate human scoring with real statistical strength on written responses, correlations up to .85 and a macro-F1 of .87 in one PLOS One study, but statistical agreement with human raters is not proof that a system measures critical thinking itself. The same research documents fairness violations across essay topics, unresolved construct-validity questions, and equity concerns for multilingual and under-resourced students. "Accurate on average" and "fair for every student" are separate questions that deserve separate answers.

The decision rule that follows from this evidence is straightforward. Use AI for practice, drafting feedback, and Socratic-style prompting, where errors are low-stakes and revisable. Reserve consequential scoring, grades, placement, credentials, for tools with documented, context-specific validity and fairness testing plus a human reviewer in the loop. Research on subgroup fairness, model drift, and authentic authorship, what happens when students use AI to help write the essay being graded, remains thin enough that a vendor's accuracy claim shouldn't be trusted outside the exact context it was tested in.

Before adopting or trusting an AI-scored assignment, ask the platform provider for its documented reliability, validity, and fairness testing, and confirm with the relevant department or program whether AI-assisted grading is used only for formative feedback or actually counts toward a final grade. That conversation, not a vendor's accuracy percentage, is what should settle the decision.

Sponsored
The Classroom Logo

The Classroom provides honest, relatable, step-by-step guidance for high schoolers applying to college and first-time undergraduate students.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.