We read with great interest the study by Azizi et al. , applying the machine-to-machine deep learning model to estimate retinal nerve fiber layer thickness from fundus photographs in the Canadian Longitudinal Study on Aging cohort.
The model’s mean absolute error exceeds the between-group differences attributed to sex and ethnicity
The authors reported that the machine-to-machine model had a mean absolute error of 7.4 µm in predicting retinal nerve fiber layer thickness from fundus photographs. Within the same analysis, the sex-related difference in predicted thickness was 1.0 µm for females vs males, and the ethnic group differences between White participants and those identifying as Black, Latin American, or Middle Eastern ranged from approximately 3.4 to 4.1 µm. The authors noted this measurement noise floor in their limitations section, correctly observing that these differences were smaller than the model’s mean absolute error. However, the clinical implications deserve sharper framing than those provided in this study. When the magnitude of a between-group difference is substantially smaller than the instrument’s measurement error , the difference cannot be reliably attributed to a biological signal rather than prediction noise. The sex and ethnicity findings are presented with statistically significant p -values derived from very large sample sizes (14,339 females and 13,775 males), where the statistical power is sufficient to detect mean differences far smaller than the measurement error with high confidence. Statistical significance in this context does not establish that the model resolves a true biological difference ; it establishes only that the model’s outputs differ systematically between groups at a magnitude that may reflect differential input features exploited by the algorithm rather than actual retinal nerve fiber layer differences . The authors’ reassurance that such small differences “would not be clinically relevant” does not address the deeper inferential problem, which is that findings falling below the instrument’s noise floor cannot be meaningfully interpreted, regardless of their p -values.
SELF-REPORTED GLAUCOMA INTRODUCES SYSTEMATIC MISCLASSIFICATION THAT OPERATES ASYMMETRICALLY ACROSS THE AGE SPECTRUM
The 8.8 µm mean difference in model-predicted thickness between participants with and without self-reported glaucoma was the most clinically consequential finding of the study. The authors acknowledge that self-reported diagnoses are subject to recall bias and lack of diagnostic confirmation and note that the observed 4.8% prevalence aligns with population-based estimates . However, beyond prevalence alignment, the direction and magnitude of misclassification were not random across the cohort’s age range . In the 45 to 54-year age group, the prevalence of self-reported glaucoma was 1.2%, whereas among those aged 75 years and older, it reached 11.4%. Early glaucoma in younger participants is more frequently undiagnosed and, therefore, more likely to be misclassified as normal, whereas older participants with a longstanding diagnosis are more likely to correctly self-report. This asymmetry means that the model-predicted thickness difference between the self-reported glaucoma and normal groups is almost certainly an underestimate of the true structural difference, and the age-by-glaucoma interaction term showing accelerated thinning of 0.25 µm per year in those with glaucoma vs 0.11 µm per year in those without glaucoma, conflates true biological acceleration with the progressive improvement in diagnostic self-awareness that occurs with age and disease duration. However, the cross-sectional design of this study cannot disentangle these contributions.
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree