In a recent Flypaper post, Dale Chu argues that the promise of a single assessment system doing multiple things well is “too good to be true.” He says that through-year assessments (TYA)—tests administered periodically throughout the school year that culminate in a summative score—are trying to do too much. But his critique relies on theoretical exaggerations, ignores the glaring failures of our current testing landscape, and dismisses the very real potential of through-year assessments to reduce testing while actually helping teachers.
The foundation of his argument rests on a philosophical, rather than empirical, claim: that a test cannot simultaneously be instructionally useful for teachers and provide valid accountability data for policymakers. In reality, state assessment systems should serve two perfectly legitimate purposes: instructional improvement and system-level accountability. When state assessments are designed to serve instructional improvement, they can also serve system-level accountability. It’s the reverse that is not true. The dozen states redesigning their state end-of-year tests into through-year tests to support educators to improve learning also can help districts and states see whether all students are being served well, and where the system needs to act.
Dale also selectively points out the potential “unintended effects” of through-year assessments, while conveniently avoiding the massive unintended negative effects created by the current system. The modern assessment landscape is an incoherent pile of redundant, unhelpful tests. The average American student takes seven standardized tests a year, and in the highest-testing districts, students take 88 by the end of eighth grade.
Because state end-of-year tests fail to give teachers timely feedback, a market failure has emerged. Districts spend millions layering on commercial interim assessments (like NWEA or STAR) throughout the year. District leaders see these interim scores, assume students are on track, then learn otherwise from the state summative results.
This pile of expensive, unhelpful tests littering the school calendar surely isn’t what advocates for system accountability intend.
Perhaps most frustrating is Dale’s assumption that states adopting TYA will inevitably spend more money and “add to students’ total testing burden.” This assumes the worst case and ignores how TYA actually functions. If a state-coordinated through-year system replaces the disjointed commercial interims districts currently buy, total testing time goes down, not up. In fact, during Indiana’s first year of ILEARN implementation, more than 60 percent of surveyed districts reported they are eliminating commercial interims like MAP or STAR because ILEARN provides what they need. Texas soon will explicitly prohibit districts from administering locally mandated benchmark tests that are not aligned to what’s taught as it implements its new through-year assessment system. This should result in fewer tests overall.
Through-year assessments are not a magic bullet, but they hold the distinct promise of making summative tests far more meaningful to educators and families. By breaking up the end-of-year testing monolith and aligning assessments more closely with what is actually being taught, TYA provides timely, relevant feedback that lets teachers make real-time course corrections during the school year. Here are facts about leading models:
- Montana’s MAST allows districts to choose when to administer “testlets” according to their own local scope-and-sequence, counts the testlets’ results toward the summative score, and includes student misconceptions on the math reports that educators are nearly universally praising as immediately actionable and helpful. MAST was designed first for instructional utility and is in federal peer review now to serve as the state’s accountability assessment. And the worry that this funnels every school onto a single course sequence does not hold up: The state built a configurator that lets each district match the testlets to its own curriculum and pacing. This year it ran more than 800 configurations across 52 curricula, built on the most commonly used high-quality instructional materials (HQIM). (See Arthur VanderVeen’s post for more about Montana’s model.)
- Indiana’s ILEARN is a series of formative “checkpoints,” with “second chance opportunities” throughout the year that feed predictive modeling, though the summative score comes only from the end-of-year test. It uses an application programming interface (API) to sync state assessment data automatically with locally adopted HQIM and intervention platforms, which eliminates manual data tracking and generates personalized study and remediation plans in tools students already use. Indiana is also moving to require those vendors to align to the state’s Performance Level Descriptors, so the rigor bar holds no matter which materials a district adopts.
Dale also raises the specter of kids and teachers facing three high-stakes tests a year instead of one. We would frame the stakes differently. The stakes can be higher, but in a way, we should want them to be. When a test gives a teacher honest, specific information about what a student knows, where they are stuck, and why—and does so while that student is still in the room—it pushes adults to act on the results during the year instead of reading about them after the fact. That is the kind of accountability people have been asking for, when they complain that schools are judged on results they never had a real chance to do anything about.
We’re excited to see what Kansas, Texas, and Missouri will contribute to continued state innovation and learning. At last month’s National Conference on Student Assessment, Texas Education Agency Commissioner Mike Morath indicated the new state-developed beginning- and middle-of-year interim assessments will use “warm reads,” which are texts related to the concepts and vocabulary in students’ curriculum. Warm reads, described here by science of reading guru Natalie Wexler, help more accurately measure students’ reading comprehension. Texas is going a step further and barring districts from layering other non-instructionally relevant tests on top of its system and setting criteria for what qualifies. We are for it, on one condition: The bar for quality has to be high, and the state must face down the vendor lobby when their products cannot clear it.
Given all of this potential, we wish Dale and others would stop arguing against the theoretical versions of through-year assessment, rather than evaluating whether specific models actually outperform the broken status quo. We need more research, evaluation, and innovation to understand how well TYA measures provide support for both instruction and accountability and fewer theoretical critiques.
The question we should be asking is not “can one test do everything flawlessly?” (Because, duh, no.) The real question is whether a thoughtfully designed, state-coordinated through-year system can work better than the expensive, redundant, and instructionally useless status quo we have today.