The types, amounts, and modes of testing done in our schools are subject to debate and analysis, with “too much,” “too often,” “too complex,” and “not helpful” among the many complaints. A new report looks to address these questions by focusing on optional assessments called “benchmark modules.” They appear to counter some of the criticisms of testing—proving popular with teachers and seeming to provide positive impact on students’ end-of-year exam scores.
Data come from Utah, where the state board of education launched benchmark modules for instructional use in 2016. These are available for grades 3–8 in English language arts (ELA), math, and writing, and for grades 4–8 in science. Use was—and remains—optional from the state’s perspective and is typically determined at the school or classroom level. (Parents can also opt students out of these or nearly any test in Utah.) While the state doesn’t report who uses them, that information is collected, and researchers were given access to it down to the classroom level. State policymakers see the modules as a “productivity tool”: short, targeted interim assessments for teachers to understand how students are performing in relation to Utah’s state standards during the course of instruction.
Researcher Kyla McClure of the University of Colorado Boulder looks at administrative data for all benchmark modules assigned by Utah teachers from the 2020–21 through the 2022–23 school years in grades 3–8. Each observation identifies the student, teacher, class, and school, as well as scores on every module, thus allowing those results to be linked to end-of-year test scores. McClure runs fixed effects regressions to estimate the effect of completing benchmark modules on end-of-year outcomes.
Use of benchmark modules more than doubled during the relatively brief period under study, from 486,000 in 2020–21 to 1.3 million in 2022–23. The number of students in Utah also increased over this period, as did the percentage of eligible classrooms and teachers assigning modules (from 26 percent to 50 percent), which accounted for some of that growth. Also relevant was that teachers assigned a growing number of benchmark modules per student per class.
In 2022–23, most teachers who used the modules did not begin assigning them until March, indicating primary usage as preparation for end-of-year state testing. The second largest group included teachers who only assigned modules at or near the start of the school year, perhaps as diagnostic tools. The third largest group comprised teachers who used modules continuously through the year—and the number of such users increased by over 400 percent since 2020–21. Across all three years of the study, teachers gave an average of 2.4 to 4.2 benchmark modules per student, with end-of-year users giving the fewest. Continuous users gave almost triple the number of benchmark modules as the other groups, on average.
McClure finds a small, positive, and statistically-significant relationship between having taken a benchmark module and end-of-year summative test scores. Students who took at least one module in a grade and subject scored about 0.02 to 0.05 standard deviations higher than their peers in the same school or with the same teacher. Students who took an additional benchmark module scored about 0.006 to 0.012 standard deviations higher, on average, compared to students who took one fewer module—a smaller but still statistically significant impact as the number of modules rose. The effect was higher on end-of-year math tests than on ELA tests but near zero on science tests. Scores were higher in classes where modules were assigned continuously throughout the year, versus start-of-year and end-of-year classes.
In the end, this report presents an interesting snapshot. No causal effects can be determined, and mechanisms to explain the observations can only be conjectured at. Are kids responding to the feedback they get via benchmark modules? Do teachers make changes in their instructional approach based on module scores? Does experiencing multiple modules sharpen a student’s test-taking ability? Do school leaders review and address module data with teachers? Utah’s schools and classrooms remain black boxes, but the data emanating from them—especially in terms of this optional interim testing regime—are intriguing.
SOURCE: Kyla N. McClure, “Benchmark Modules: A Better Interim Assessment? Evidence From Statewide Use of Benchmark Modules in Utah,” AERA Open Journal (October 2025).