New York City Schools piloted an external qualitative review (QR) in a subset of its schools for the first time in 2005, intended to expand the definition of school quality via a school inspection-type model and to broaden the scope of educational accountability. Along with traditionally examined outcomes like test scores, disciplinary actions, attendance, and graduation rates, outside analysts aimed to capture schools’ quality across three additional dimensions: their instructional core, school culture, and structures in place for continuous improvement. New research looks at whether such an expansion of accountability improved schools.
Due to concerns over cost and potential disruption to school operations (external reviewers were on site for two or three days), the New York City Department of Education both delayed and staggered the rollout of QR across the district, beginning in the 2009–10 school year. The first year prioritized schools that were low-performing or in danger of becoming low-performing based on the most recent district-issued Progress Reports. Higher-performing schools were entered into a lottery system to receive their first qualitative review as soon as possible after 2009–10. The low performers received annual reviews; higher performers were on a three- or four-year schedule depending on how highly they were initially rated.
American University researcher Robert Shand exploits all of these implementation quirks in his analysis, covering 2009–10 through 2013–14. His main goal: determining how more frequent exposure to external qualitative evaluation impacted school practice and learning outcomes, including test scores, graduation rates, attendance, and the various QR measures. For just a taste of these new quality measures: Instructional core indicators included measures of a rigorous, engaging curriculum aligned with learning standards; school culture measures captured a positive learning environment (or not) and a climate of high expectations (or not); and structures for improvement concerned processes like institutional goal-setting and teacher support systems.
The publicly available data Shand uses came from the New York City Department of Education, including whether and when a school was subject to QR, scores for all QRs conducted, and measures of teacher practice from the annual Learning Environment Survey/School Survey. Student outcomes included scores on third through eighth grade math and English language arts state tests, four-year and six-year high school graduation rates, and student attendance, all compiled by NYU’s Research Alliance for New York City Schools. Shand employs a regression discontinuity design to examine the effects of additional evaluation and feedback provided via QR on student outcomes and the new quality measures.
First, the bad news: No statistically significant impacts from a quality review were seen in any school receiving a review, and even the small nonsignificant positive improvements on test scores observed were driven by a handful of schools, appeared two to three years after QR, and faded by the following year. Whether reviews were annual or intermittent made no difference.
On the more positive side, some additional dimensions of school quality examined via QR did show improvement following a review. These include the number of teachers reporting that they use data to guide decision-making and those working in collaboration with their peers, as well as measures of overall learning environment quality. Even better: Higher overall QR scores predicted modest test score gains in a given school—albeit not for a year or two after the review—with Shand suggesting that QR’s primary value is as a leading indicator rather than a diagnostic tool.
“Taken together,” Shand concludes, “the findings support a cautiously optimistic but qualified view of qualitative accountability as one component of a broader school improvement strategy…but insufficient on its own to drive sustained gains in student learning.” In other words, even a rigorous on-site review doesn’t negate the need to keep a sharp eye on more traditional measures of quality and for schools to change course when, say, test scores fall or absenteeism rises.
SOURCE: Robert Shand, “The Effects of Qualitative Accountability on School Improvement Efforts: Evidence from New York City,” Leadership and Policy in Schools (June 2026).