We thank Aitha and colleagues for their thoughtful comments regarding our article. Several points warrant clarification.
The first critique raises an important issue regarding the distinction between individual-level prediction error and the precision of group-level estimates. The reported MAE of 7.4 µm characterizes individual-level prediction accuracy, not the standard error (SE) of a group mean. Because the SE of a group mean scales with σ/√N, the large Canadian Longitudinal Study on Aging (CLSA) sample size allows relatively precise estimation of average group differences. For example, with 14,339 women and 13,775 men, the SE of the sex-difference estimate is approximately 0.1 µm, an order of magnitude smaller than the observed 1.0 µm difference. Similar reasoning applies to ethnicity/race comparisons, although those findings should be interpreted cautiously given smaller subgroup sizes and limited cohort diversity. Notably, OCT measurements also have non-negligible test-retest variability; in glaucomatous eyes imaged with Cirrus HD-OCT, one study reported an intervisit tolerance limit of approximately 3.9 µm for average peripapillary RNFL thickness. Applying this reasoning to OCT would imply that OCT-based studies could not detect group-level differences smaller than test-retest variability. Yet OCT-based studies have reported sex differences of approximately 1 to 2 µm and ethnicity/race-related differences of several micrometers in global RNFL thickness, , on the same order of magnitude as the differences estimated by the M2M model in the CLSA cohort.
We agree that self-reported glaucoma may introduce misclassification and that age-related diagnostic awareness could influence the age × glaucoma interaction. Self-reported glaucoma in population cohorts has been shown to have reasonable specificity, though imperfect sensitivity. However, if some participants with undiagnosed glaucoma were classified as non-glaucomatous, they would be expected to make the non-glaucoma group appear thinner, thereby reducing rather than exaggerating the observed difference between groups. Thus, the observed glaucoma-associated RNFL difference is likely conservative. Additionally, the consistently thinner M2M-predicted RNFL and greater inter-eye asymmetry among participants with self-reported glaucoma support that the model captured a meaningful structural signal. Further support comes from our subsequent longitudinal analysis of the same cohort. Among participants reporting no glaucoma at baseline, faster M2M-predicted RNFL loss was independently associated with incident self-reported glaucoma over three years of follow-up (HR, 1.125 per 1 µm/year; 95% CI, 1.070-1.183; P <.001). Because structural change was quantified before the incident diagnosis, these findings reduce the likelihood that the observed association is explained solely by age-related diagnostic awareness and support the biological plausibility of the observed cross-sectional association.
The suggested sensitivity analysis using intraocular pressure (IOP) or cup-to-disc ratio (CDR) as alternative ascertainment has important limitations. IOP is an imperfect surrogate for glaucoma status: treated patients may have normal pressures, and normal-tension glaucoma occurs despite statistically normal IOP, which could introduce non-random misclassification related to treatment status and disease subtype. Similarly, CDR estimation from photographs would reintroduce the subjectivity that the M2M model was designed to reduce. Adjudicated CDR grading at this scale would also be impractical, reflecting the very constraint that motivated the use of a scalable, fundus-based deep learning approach for structural screening.
CRediT authorship contribution statement
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree