The study, published in PNAS Nexus and involving a randomised experiment with more than 1300 teachers in Greece, has major implications for schools.
Researchers measured how teachers responded to unfairly harsh grades given to students’ work by an AI system compared with another teacher. They found teachers were more likely to accept an unfair grade when it had been assigned by AI.
The research was conducted by Dr Sofoklis Goulas, from Yale University, Monash University’s Professor Rigissa Megalokonomou, and Curtin University’s Dr Panagiotis Sotirakopoulos.
Megalokonomou, from the Monash Business School, says the findings challenge the assumption that human oversight is enough to prevent AI mistakes.
“Teachers are deferring to AI outputs without sufficiently questioning them, and errors are quietly going uncorrected,” Megalokonomou, whose most recent research has also focussed on female students, disruption, and STEM outcomes in disadvantaged schools, says.
“Multiply that across thousands of students and thousands of classrooms, and you start to see how quickly this becomes a serious problem.”
Megalokonomou says the findings present challenges for schools, with AI tools increasingly used for lesson planning, feedback, identifying struggling students and administrative tasks.
According to the research, nearly half of teachers (48 per cent) used AI tools at least weekly for lesson preparation. However, only 16 per cent actively urged fellow teachers to adopt the technology.
The study also found that younger teachers, those with postgraduate qualifications, and those who described themselves as technologically confident were the least likely to challenge a harsh AI-generated grade.
“The very people we might expect to be the most capable and critical users of AI tools turned out to be the most likely to defer to a harsh AI grade,” Megalokonomou says.
“That’s counterintuitive and concerning because the people pushing AI integration forward in schools may be the least likely to catch its errors.”

The report says treating educators as co-designers and critical partners – not passive end users – will help ensure that AI supports rather than supplants professional judgment and advances decision quality.
Researchers are now developing a training program to help teachers use AI critically.
“It’s not enough to just tell people AI can be wrong,” Megalokonomou says.
“You need to show them specifically how and when their judgment is likely to go astray and build the habits to push back on AI.”
The study found that while teachers are open to AI as a planning aid, they remain skeptical about delegating evaluative authority to machines.
“This skepticism is grounded not only in abstract concerns about fairness or control, but also in the nuanced realities of classroom life,” the report stated.
At the end of the survey, teachers were invited to respond to an open-ended prompt: “Is there anything you would like to add or explain?”
Some used this opportunity to elaborate on their views about the role of AI in grading, the research stated.
“Their responses emphasised that grading often requires contextual sensitivity that AI systems lack.
“Teachers cited examples involving illness, learning difficulties, and students from vulnerable backgrounds – situations in which discretion and empathy are essential.”
One respondent asked, “Can AI perceive special circumstances, such as a student’s absence due to illness, a lack of teaching due to teacher absences, and therefore insufficient instruction?”
Another stated, “Learning difficulties influence grading,” while a third explained that “teachers take into account parameters such as social conditions (eg. how do I grade a child from Egypt who, if held back, will be married off at 12), and family environment (abused children, children of drug addicts, etc.).”
Others pointed to subject-specific concerns, arguing that AI may be better suited for technical tasks than interpretive ones.
“One teacher remarked, ‘The artificial intelligence system lacks emotional flexibility; it is fair and accurate in science subjects. In humanities subjects, I think it just helps [as a tutor], it’s not [good] in grading.’“
The researchers say the results have direct implications for education policy and the responsible use of AI in classrooms.
“First, algorithmic tools should be designed to signal both competence and accountability, so that teachers treat AI advice with the same critical scrutiny they apply to human guidance,” the report’s conclusion read.
“Second, professional development programs may help prepare teachers to audit and, when necessary, override algorithmic recommendations.
“Training in bias detection, error patterns, and explainability can calibrate trust and improve the accuracy of human oversight.
“Treating educators as co-designers and critical partners – not passive end users – will help ensure that AI supports rather than supplants professional judgment and advances decision quality.”
The research team is exploring future projects to investigate whether AI helps reduce or amplify the biases teachers bring to grading, whether AI training improves teachers’ day-to-day productivity, and whether simple, low-cost interventions can change how educators engage with AI tools.
“I hope this research reaches the people who are making decisions right now about AI in schools, including policymakers, school leaders and education departments,” Megalokonomou says.
To read the research paper titled ‘Why do experts miss AI’s errors? Evidence from a randomized labeling experiment ‘ click here.