We read with great interest the study by Arnold and colleagues evaluating alternative retinopathy of prematurity (ROP) screening criteria in a large, racially diverse United States cohort. Their analysis importantly showed that attempts to reduce screening workload may compromise sensitivity in selected populations and that high-risk subgroups, particularly infants of Pacific ancestry, require careful consideration.
We agree that modifications to ROP screening criteria should not be adopted solely on the basis of efficiency. However, their findings also raise a broader issue: external validation should assess not only diagnostic performance within a new dataset, but also the transportability of screening criteria across different health systems and neonatal-care environments.
This distinction is clinically important. Arnold et al. acknowledge that populations with less consistent neonatal intensive care may require broader screening thresholds, and they show that even within a United States cohort, subgroup risk may influence the safety of reduced-sensitivity strategies. The next step, therefore, may be to frame ROP screening not as a search for a universally superior algorithm, but as a process requiring local validation and population-sensitive adaptation.
Recent international data support this view. In the prospective multicenter TR-ROP 2 study from Türkiye, three infants requiring treatment would have been missed without risk-factor-based extension of screening criteria, despite broader national thresholds. In the United Kingdom, Aulakh and colleagues reported six treated infants who would have fallen outside the updated national screening criteria, emphasizing the limitations of narrowing gestational-age thresholds without incorporating systemic risk factors such as poor postnatal weight gain. In China, Lu et al. found that several prediction models performed differently in a bicentric validation cohort and concluded that screening algorithms require adjustment according to local population characteristics. South African registry data further illustrate that screening coverage, completion, and local neonatal-care conditions are themselves critical determinants of program safety.
Together, these observations suggest that the key question is not simply whether a proposed criterion maintains sensitivity in a single external cohort, but whether it remains safe when transported across populations with different disease epidemiology, survival patterns, neonatal-care quality, and follow-up infrastructure.
Thus, Arnold et al.’s study should be viewed not only as an external validation of proposed United States screening modifications, but also as a reminder of the limits of generalizability. Future ROP screening research should prioritize prospective local validation, explicit evaluation of subgroup risk, and incorporation of population-specific clinical factors alongside gestational age and birth weight. Such an approach may better preserve patient safety while still allowing rational reduction of unnecessary examinations.
CRediT authorship contribution statement
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree