Concern is growing over the impacts of AI use on students’ ability to write complete, coherent, logical essays on their own and to learn from the recursive process of drafting, reviewing, and rewriting. But what if AI’s role is limited to that of the editor, rather than the author or assistant? A new report looks at the impact of an Automated Writing Evaluation (AWE) program on students’ writing performance.
Nine researchers, led by Joshua Wilson of the University of Delaware, use a randomized controlled trial to compare outcomes for middle school students using an AWE system called MI Write to those of a similar group of students who learned via more traditional methods. (Note: MI Write was developed by Measurement Incorporated, a firm that sells numerous assessment and evaluation services to schools nationwide, and two of its staffers are included on the research team that produced this report.) Data come from three unnamed school districts in the mid-Atlantic and southeastern regions of the United States during the 2021–22 school year. One is suburban, one is urban, and one is rural. All three were chosen because they serve a population in which at least 50 percent of students belong to one or more “priority populations”—that is, Black, Hispanic, and/or low-income students, determined by eligibility for free or reduced-price lunch—as defined by the priorities of the Gates Foundation, who funded the study. A total of 37 seventh- and eighth-grade English language arts teachers across those districts opted to participate, and 19 were randomly chosen as the treatment group (n = 1,260 students), while 18 teachers were randomly assigned to the control group (n = 1,277 students).
Control group classrooms experienced “business as usual” ELA instruction. In one district, this consisted of an in-house curriculum that taught one type of writing per quarter and included narrative writing, expository research, and literary analysis. The other two districts used a web-based curriculum called StudySync whose writing component covered a number of writing styles—including argumentative, informative, explanatory, literacy analysis, and narrative—at various times throughout the year.
Treatment group classrooms incorporated MI Write, which uses AI to guide students through recursive cycles of planning, drafting, revising, and editing and can be tailored for any style of writing. Students draft essays in response to MI Write prompts, and while the platform facilitates peer and teacher review, its central feature is an automated feedback system designed to provide almost-immediate analysis of submitted essays across multiple drafts.
Feedback reports include a score of 1.0 to 5.0 on each of six writing traits (development of ideas, organization, style, word choice, sentence fluency, and conventions of a given style), descriptive evaluation and feedback statements, questions for consideration ahead of revision, recommended interactive lessons based on the score for each trait, and text-embedded feedback on grammar and spelling to support editing. There’s much more to it, and interested readers should check out this video demo for additional details.
The researchers examined how students using MI Write compared to their control group peers based on the various grading processes connected to the differing curricula, and whether MI Write students reported greater self-efficacy, greater enjoyment of writing, and more positive beliefs about writing as a recursive process. The two groups of students were broadly similar in grade level (more eighth- than seventh-graders) and gender breakdown (slightly more girls than boys), were plurality Hispanic (more than 42 percent of both groups), and comprised more than 77 percent of “priority population” students.
Overall, the regression analysis showed no statistically significant impact of treatment on students’ writing quality when measured either by human raters or via the MI Write scoring system. Same for all of the survey-based outcomes. However, additional analysis indicated that the three districts implemented MI Write with varying levels of fidelity and frequency. In the district identified as rural, for example, every lesson that research team members observed used MI Write. In that district, small but significant positive impacts were seen on all outcomes. The other two districts were classified as “rarely used” (suburban) and “intermittently used” (urban) and registered null or slight negative impacts (although still significant) across the board. Even in the most-committed district, the report provides the anecdote of a teacher who “routinely struggled with classroom management, reducing effective writing time,” with concomitant lower scores than other classrooms in the same district. Meanwhile, some control group students “seldom wrote beyond brief worksheet answers, and no essays were observed to be drafted” during class. There are more such anecdotes included.
What does all this tell us? Anyone hoping for AI-powered curricula to be a revolution in education is going to run into the same implementation problems that have plagued the much more well-regarded (and much more traditional) science of reading revolution. As Robert Pondiscio puts it: even the highest quality curriculum doesn’t teach itself. And even with the highest fidelity, a revolution is unlikely to occur. Teachers and students are not machines—even when working and learning with machines.
SOURCE: Joshua Wilson et al., “Impact of Automated Writing Evaluation (AWE) on Middle School Students’ Writing Outcomes,” American Educational Research Journal (November 2025).