W e read with great interest the article by Swaminathan and Medeiros, which discusses key considerations in interpreting large database studies and their direct clinical implications. The authors eloquently delineate the appropriate use cases for big data and the biases that can undermine their interpretation. We sincerely appreciate the many deliberations brought out in the paper, including the lack of longitudinal imaging and residual unmeasured confounding. We found the points regarding data quality and measuring variability to be of particular interest in that clinical care often does not follow a standard protocol as a prospective study would. As such, readers of database studies should maintain a level of moderation in interpreting the conclusions of said papers. Many existing database papers have included some of these considerations in their Limitations sections; however, this paper expands upon these implications with further detail, and we are in agreement. We write to extend one dimension of their discussion that we believe warrants explicit attention: the need for structured mentorship and training infrastructure for trainees and junior investigators engaging with these platforms.
The pitfalls that the authors describe—residual confounding, reliance on billing codes, and informative missingness—require not only awareness but also practiced judgment to navigate responsibly. However, most trainees encounter these data sets without formal methodological training specific to their unique limitations. Guidelines analogous to the Strengthening the Reporting of Observational Studies in Epidemiology for observational studies would be helpful for ophthalmic database studies. Explicit checklist items addressing phenotype validation strategies, handling of missing data, multiplicity correction, and effect size interpretation would directly address the concerns raised by Swaminathan and Medeiros, while also providing peer reviewers with clearer evaluative criteria. Within TriNetX studies alone, we see methodological variability. For instance, in propensity score matching, different authors defer to different definitions of adequately matched cohorts with variable standardized mean difference values. Some use a value of >0.1 to denote inadequate matching, whereas others categorize it into buckets of 0 to 0.2, 0.2 to 0.5, 0.5 to 0.8, and ≥0.8 to denote the degree of difference between cohorts. Such variability in the interpretation of in-platform statistics can introduce a level of bias into the studies, not accounted for by a standard reporting guideline. We propose that academic ophthalmology training programs should consider integrating dedicated big data methodology curricula into their research training pathways, ensuring that the next generation of clinician-scientists develops fluency not only in using these platforms but also in understanding their inherent constraints.
We commend Swaminathan and Medeiros for opening up this important discussion and articulating these principles with clarity. Their perspective provides an important foundation upon which the field can develop more standardized and methodologically rigorous approaches to large-scale ophthalmic research. We hope that these discussions can catalyze broader conversations among researchers, educators, trainees, and editors about how the ophthalmology community can use these tools responsibly in both reporting and interpretation.
CRediT authorship contribution statement
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree