During my school years, I was often called a student who “could do better.” I took feedback seriously and tried to act on it. But one teacher had already made up his mind about me. He thought I was lazy and naughty. There was a reason for it, but somehow that single incident became my identity, and it showed up in my grades. I kept wondering how to get out of his bad books. I never did.
Psychology calls this the horn effect: a cognitive bias where one negative trait distorts everything else you see in a person. It shapes classrooms as much as any other space, and human bias in assessment is still a daily reality. I used to wonder whether there was a way out, some objective tool that could set the bias aside.
Melissa Peplinski and Haley Gaudreau offer one answer in their 2023 entry for the Fordham Institute’s annual Wonkathon, “How we can use AI to increase access and equity in science education.” They propose that AI grading tools evaluate student work against objective standards, keeping a teacher’s implicit biases out of the score. However, they caution that AI models carry their own risks of bias and misinformation if left unverified.
Now, working in ed-tech, I see how alive this question is. Visit almost any school and someone asks, “Who grades better, a computer or a teacher?” I’ve come to think that’s the wrong question. Both can be wrong, and both can be unfair. What differs is the kind of mistake each one makes.
Before we get to the design choice, we have to acknowledge the complexity of human behavior. Teachers are brilliant, and they are the backbone of K–12 education. They are primarily human beings who are prone to decision fatigue and mental fog. They make mistakes, and that’s OK.
On a day-to-day basis, when a teacher works through 40 assignments, the fatigued brain reaches for shortcuts by the time it gets to the bottom of the stack. Skewed standards of checking become more prominent due to implicit biases, unintentionally and unconsciously. Without quality rest, physical and mental, the cycle simply repeats itself the next week.
AI does not carry these particular human burdens. It reads the hundredth assignment with exactly the same attention it gave the first, applies exactly the same rules every time, and works at a speed no human can match. Is that enough to grade assessments then?
To be honest, being steady and fast is worthless if the rules are wrong. As Peplinski and Gaudreau point out, no technology is free of bias or error. AI is trained on human data. So if the rules are flawed to begin with, consistency and speed only help you be unfair/biased faster!
You can actually see this on a marked paper. AI graders like long answers. A long, polished one can end up beating a shorter answer that was more accurate. Grammar is another one. A student writing in their second language might have the right idea and still lose marks for how it's written, which doesn't seem fair to me. Then the training data is mostly Western and mostly English. So if a student uses an example from their own culture, the AI may not give the credit a teacher would have given. Same with a correct answer that got there in an unusual way or method—it can lose marks just for not looking like the “typical” good answer. To be honest, this next part bothers me the most: Sometimes the AI reacts to who the student seems to be! When researchers gave ChatGPT identical essays with different descriptions of the student, the scores shifted. An analysis by ETS of more than 13,000 essays found GPT-4o marked some student groups down more harshly than expert human raters did. It reminded me of my teacher who decided I was lazy: the same kind of judgment but coming from a tool this time.
This is why schools have to be deliberate about how they deploy these tools. Some of these biases live in the models themselves; others creep in through how the tool is set up. Schools can’t retrain the models, but they can design around both.
The most significant part of the process is for teachers to sit down together and agree on the ground rules. The team must be on the same page about rubrics, which papers get the perfect grade, which papers qualify as weak ones. Once this aspect is settled, the team can train the AI to mirror that marking style, so it reflects how a particular department wants the assignment or paper to be assessed. Teachers set the standard, in other words, and AI just holds the line.
This does not quite guarantee an unbiased AI. Training will always need human oversight. Two simple steps help here. Hide student names before the AI marks anything, so it can’t be swayed by who the student seems to be. And when showing the AI what a strong answer looks like, include good answers from students writing in a second language, so it doesn’t learn to mistake polished English for good thinking. An automated grader can easily disadvantage students who are still learning the nuances of the concept. Only a teacher who spends a large amount of instructional time with the student can understand a student’s true learning level. Only a teacher knows that a particular student underperformed on a summative because they were shaky on the topic, or simply nervous on the day.
The best design uses both. Think of AI as a very capable assistant: It handles the fast, repetitive checking, holds the rubric steady, and buys back hours of a teacher’s week.
But the teacher stays in charge. When a response is ambiguous or sits somewhere in the middle, it goes back to the teacher for a human read. The teacher always has the final say and can override any grade. Nothing reaches a report card until a real person signs off on it.
Maybe we can override the horn effect and fatigue without blaming any one stakeholder. The way forward is to support the teachers and allow them more time to be humane.