Skip to main content

Mobile Navigation

  • National
    • Policy
      • High Expectations
      • Quality Choices
      • Personalized Pathways
    • Research
    • Commentary
      • Gadfly Newsletter
      • Flypaper Blog
    • Events
    • Scholars Program
  • Ohio
    • Policy
      • Priorities
      • Media & Testimony
    • Research
    • Commentary
      • Ohio Education Gadfly Biweekly
      • Ohio Gadfly Daily
  • Charter Authorizing
    • Application
    • Sponsored Schools
    • Resources
  • About
    • Mission
    • Board
    • Staff
    • Career
Home
Home
Advancing Educational Excellence

Main Navigation

  • National
  • Ohio
  • Charter Authorizing
  • About

National Menu

  • Topics
    • Accountability & Testing
    • Advanced Education
    • Career & Technical Education
    • Charter Schools
    • Curriculum & Instruction
    • ESSA
    • Evidence-Based Learning
    • Facilities
    • Governance
    • Personalized Learning
    • Private School Choice
    • School Finance
    • Standards
    • Teachers & School Leaders
    • Think Again
  • Research
  • Commentary
    • Gadfly Newsletter
    • Flypaper Blog
    • Gadfly Podcast
  • Events
  • Scholars Program
Flypaper

Should AI be used for teacher evaluation?

Kim Marshall
3.26.2026
Teacher teaching students in a classroom
Getty Images/Unaihuiziphotography
Listen to this article
Loading the Elevenlabs Text to Speech AudioNative Player...

For teachers, artificial intelligence can be a boon as well as a challenge.

K–12 administrators are also taking advantage of AI, and I’ve been curious about how it’s being applied to teacher evaluation. I just did a survey of my Marshall Memo subscribers around the world and got 1,280 responses. The sample is admittedly unscientific (readers who chose to respond in the middle of a busy week), but the results jibe with what I’ve been seeing in my school visits and in conversations with front-line educators. The headlines:

  • Lots of supervisors are using AI to save time drafting, organizing, summarizing, and editing their notes on classroom visits.

  • A smaller number are using AI tools to make audio or video transcriptions of lessons and draw conclusions about teachers.

  • A few are feeding the Danielson or another evaluation rubric into AI and generating Distinguished, Proficient, Basic, and Unsatisfactory ratings of teachers based on their notes or transcriptions.

The survey also asked for written comments and suggestions, and Claude helped me analyze the 522 responses. Some insights:

  • The use of AI for teacher evaluation is largely unregulated, with very few guardrails, policies, and norms and minimal training.

  • Many respondents felt strongly about maintaining human judgment in teacher evaluation. “Supervision is relational and cannot be automated,” was one comment. “It should assist, not replace,” said another.

  • Transparency was an issue, captured in this comment: “Teachers should know if AI is involved.”

  • There was concern about protecting the privacy of teachers’ and students’ information. (Teachers unions are watching this issue closely.)

As a former principal, I can certainly understand administrators using AI to get their evaluation paperwork done more quickly and efficiently, especially if the process is bureaucratic and burdensome. But there’s a real danger of skipping conversations after classroom visits (supervisors will “just send feedback,” worried one survey respondent), with teachers getting impersonal, robotic evaluations containing errors and misunderstandings. As AI is used more often in the years ahead, there’s a strong possibility that this downside will become more pervasive.

But what worries me most is AI transcriptions (audio or video) being converted directly into rubric ratings of teachers. I’ve heard about this in schools and seen it touted on commercial AI websites and believe it’s doubly flawed. Even when done by a competent human, giving rubric scores after a single classroom visit is not a good idea, especially after a short visit. And even the cleverest AI tool is incapable of capturing the complex dynamics of a classroom and producing accurate high-stakes evaluative ratings. Rubrics are best used at the end of the school year, summing up a teacher’s performance by drawing on information from multiple interactions and including an in-person comparison with the teacher’s self-assessment.

I’m also worried that using AI tools will reduce the number of teacher-supervisor conversations after classroom visits. Face-to-face debriefs, when handled well, are a chance to affirm and encourage effective teaching, give the teacher a chance to explain the broader context, engage them in non-defensive reflection, build relationships and trust, and improve classroom practices a little at a time—with significant cumulative impact.

There is some good news, however: Some enterprising educators have developed AI tools with great potential for improving teaching and learning. Three examples:

  • Lesson analysis for teachers. Inform(Ed) by Tommy Mulvoy transcribes a lesson and immediately gives the teacher charts with detailed, confidential information on the ratio of teacher-student talk time, the type of questions asked and answered, wait time, the sophistication of discourse, and more. This platform is designed not for evaluation but to get teachers reflecting on lessons and planning specific improvements. Teachers might choose to look at the data with a critical friend or a supervisor, likely adding even more value.

  • Preparing supervisors for post-visit conversations. The Leverage Leadership App by Paul Bambrick-Santoyo, to be released in June, helps administrators get ready for conversations after classroom visits. It synthesizes classroom notes, student work, assessment results, and previous teacher/supervisor interactions; suggests the best way to start the meeting, affirm what’s working, and draw the teacher out with an open-ended question (“When do you think the most learning was taking place?” “Did you get your intended results?”); and lists several ideas for a leverage point.

  • Summaries of teacher/supervisor debriefs. The Conversation Summarizer by Will and Matt Krasnow transcribes the post-visit conversation (with the teacher’s permission) and immediately produces a confidential summary of about 150 words. The AI is trained to follow the sequence of an effective conversation: appreciation of what went well, context supplied by the teacher, the coaching point (if there is one), and closure. Teacher and supervisor read the summary on the spot, make edits, sign off, and receive an electronic copy. The whole conversation—lesson debrief and summary—takes about 15 minutes.

All three of these tools use AI for what it does best—quickly synthesizing and summarizing lots of information and saving educators time—while supporting the human work of focusing on what’s going well in classrooms and what can be improved. By streamlining the observation and evaluation process, they increase the likelihood that teachers and administrators will make time for face-to-face conversations (ideally in the teacher’s classroom when students aren’t there) and do the essential work of affirming and continuously improving instruction.

An important proviso: To implement the second and third tools, schools need to modify the traditional teacher evaluation process, specifically:

  • Shift from annual, full-lesson, high-stakes evaluations to frequent, short, informal visits.

  • Liberate supervisors from using checklists and rubrics during classroom visits.

  • Require a brief face-to-face conversation after each evaluation visit.

  • Save rubric ratings for the end of the school year, combining insights from classroom observations, debriefs, other points of contact, and the teacher’s self-assessment.

It’s also vital to guarantee that information within the AI platform is confidential.

Given these conditions (which a number of schools have embraced), well-crafted AI tools can rescue the teacher evaluation process from bureaucratic purgatory, build trust and collaboration, strengthen the quality of day-to-day instruction, and improve learning for all students.

Kim Marshall, formerly a Boston teacher and administrator, coaches principals, consults and speaks on school leadership and evaluation, and publishes the weekly Marshall Memo www.marshallmemo.com. He’s the author of Rethinking Teacher Supervision and Evaluation (Jossey-Bass, 3rd edition, 2024).

Policy Priority:
High Expectations
Topics:
Curriculum & Instruction
Teachers & School Leaders
Tags: Artificial intelligence Boston

Kim Marshall, formerly a Boston teacher and administrator, coaches principals, consults and speaks on school leadership and evaluation, and publishes the weekly Marshall Memo www.marshallmemo.com.

Related Content

view
Charter performance - 2025-26 blog image
School Choice

Ohio charter schools’ performance in 2025–26

Aaron Churchill 10.2.2026
OhioOhio Gadfly Daily
view
Ohio charter news logo
School Choice

Ohio Charter News Weekly – 10.2.26

Jeff Murray 10.2.2026
OhioOhio Gadfly Daily
view
Gadfly Bites logo
School Funding

Gadfly Bites 10/2/26—Two weird tricks

Jeff Murray 10.2.2026
OhioOhio Gadfly Daily
Fordham Logo

© 2026 The Thomas B. Fordham Institute
Privacy Policy
Usage Agreement

National

P.O. Box 110
Burke, VA 22009

202.223.5452

[email protected]

Ohio

P.O. Box 82291
Columbus, OH 43202

614.223.1580

[email protected]

Sponsorship

130 West Second Street, Suite 410
Dayton, Ohio 45402

937.227.3368

[email protected]