PURPOSE
To compare the epidemiology and phenotypic characteristics of 2 populations of patients with retinopathy of prematurity (ROP) using a quantitative artificial intelligence (AI)–derived vascular severity score (VSS).
METHODS
We retrospectively analyzed 2 data sets of ROP images: (1) The iROP data set including 1019 babies and 6552 eye-level examinations developed as part of a large US multicenter cohort study from 2012 to 2020. (2) 8281 consecutive ROP telemedicine examinations from 2363 babies in the Indian Aravind Eye Care System (AECS), from 2019 to 2020. We analyzed all eye examinations using the iROP DL system, which outputs a VSS from 1 to 9. The main outcome was the relationship between the VSS and the zone, stage, and extent of stage 3 in each data set, assessed using generalized estimating equations to account for the correlation from bilateral and repeated examinations.
RESULTS
In both data sets, higher VSS was associated with higher stage ( P <.001 for all stages in both data sets), higher extent of stage 3 (i-ROP, P =.004, and AECS, P <.001), and a more posterior zone for a given stage of ROP (i-ROP: stage 1, P =.04, stage 2, P <.001, and stage 3, P <.001; AECS: stage 1, P <.001, and stage 3, P <.001). For any given stage, the VSS was lower in the AECS data set than in the i-ROP data set (all P <.001).
CONCLUSIONS
In both populations, the VSS was associated with more posterior zone and higher ROP stage/extent, though on average the VSS was lower in the AECS population. A correlation between measured VSS and intraocular vascular endothelial growth factor concentration might explain both findings.
INTRODUCTION
A HISTORICAL LOOK AT RETINOPATHY OF PREMATURITY
Understanding the regional epidemiology of retinopathy of prematurity (ROP) has always depended on identifying a given region’s evolution in the broader epidemiologic transition of the modern era, specifically as it reflects neonatal mortality and the survival of prematurely born infants. , In regions and times where premature infants do not survive, the incidence of ROP is low. As neonatal care, including resuscitation of earlier preterm infants, increases, the incidence of ROP rises significantly. Finally, as modern neonatal care practices mature, including careful monitoring and titration of oxygen, minimization of infection, and optimization of nutrition, the incidence of severe ROP levels out with the risk of ROP being primarily related to the degree of preterm birth. ,
Looking back at the evolution of ROP in the 20th century, these transitions have been described in terms of 3 “epidemics.” Although these are not true epidemics in the epidemiologic sense, the schema has become so ingrained in our understanding of ROP that we will continue to use the term here. The first epidemic of what we now call ROP occurred with the advent of modern neonatal care in the mid-20th century in high-income regions and was characterized by babies developing “retrolental fibroplasia” (RLF) with total retinal detachments after surviving premature birth. , In the next few decades, as our understanding of the spectrum of disease and its risk factors, including exposure to oxygen, infection, etc., improved, the incidence of severe ROP decreased dramatically, before any proven treatment or formal disease classification. The major lesson from the first epidemic of ROP was the key role of primary prevention in reducing the number of children with retrolental fibroplasia.
The second epidemic of ROP evolved as modern neonatal care pushed the boundaries of viability to lower and lower gestational ages (GAs). In the first epidemic, babies born at 34 weeks might develop RLF, and babies born at less than 30 weeks were not surviving. Optimizing neonatal care and our understanding of ROP physiology changed both of these things; in the second epidemic, the 34-week baby was no longer developing severe ROP, but babies at younger and younger GAs were surviving. Conceptually, we remain in this phase of the disease and will remain so for the foreseeable future in higher-income countries, where babies as young as 22 weeks are now surviving. The 2 largest risk factors in this era for severe ROP are the degree of prematurity and the exposure to oxygen necessary for survival. ,
The so-called third epidemic of ROP reflects the broader epidemiologic transition that is occurring throughout low- and middle-income countries worldwide. ,, In these regions, not only is the rate of premature birth higher, but resource constraints limit the ability of emerging neonatal care units to provide optimal staffing and monitoring of infants, and the lack of screening means that many babies progress to retinal detachment and blindness, despite the advancements in ROP care in the last 50 years which have demonstrated that this outcome is nearly always avoidable. India is a country of >1 billion people and has one of the highest prematurity rates in the world. As health care conditions have improved in the last 3 decades, ophthalmologists in India have begun to see what might rightly be described as an emerging epidemic of retrolental fibroplasia. ,
In many ways, the third epidemic and the first epidemic of ROP share a number of similarities; fewer premature babies were (are) developing severe ROP, in large part due to access to unregulated oxygen. As described previously, the solution to the third epidemic is multifaceted—first and foremost, improving primary prevention will reduce the population at risk and the incidence of severe ROP for any given GA. But secondarily, systems need to be developed to provide secondary prevention through ophthalmology-led ROP screening, either in person or via telemedicine, in each region where babies are surviving. There are a number of ways to accomplish this, but first it is important to take a step back and look at how we diagnose and classify ROP and how this has changed as the epidemiology of ROP has evolved.
DISEASE PHENOTYPES, CLASSIFICATION, AND TREATMENT
The International Classification of ROP (ICROP) was first defined in the 1980s, in conjunction with the planning and execution of the first randomized clinical trial of cryotherapy (CRYO-ROP). , The ICROP defined the language that enabled clinicians to interpret and apply treatment consistently based on the ophthalmoscopic examination findings. ICROP defined the “zone,” which is a method for describing the extent of retinal vascularization from the optic nerve to the ora serrata—a process that normally completes by term birth but is in various degrees of immaturity with premature birth. The “stage” of disease was defined as the degree of pathology at the vascular-avascular border. Plus disease was defined as marked dilation and tortuosity of the posterior retinal blood vessels, occasionally associated with neovascularization of the iris and/or an enlarged persistent tunica vasculosa lentis. Interestingly, this finding, now called plus disease, was one of the earliest recognized risk factors for subsequent retrolental fibroplasia in the earliest descriptions of the disease. The CRYO-ROP study used the definitions from ICROP to evaluate whether treatment of “threshold ROP” would lead to improved anatomical outcomes (fewer retinal detachments). Threshold ROP was defined as 5 contiguous clock hours of stage 3 or 8 noncontiguous clock hours of stage 3, in zone I or II, with plus disease. Subsequently, cryotherapy became the standard of care after the publication of this study in 1988.
The treatment “threshold” was re-evaluated in the Early Treatment for ROP (ETROP) study, published in 2003. ETROP evaluated whether treatment of “pre-threshold” ROP might improve outcomes and found that treatment of “type 1” prethreshold ROP was advantageous over waiting for threshold disease to develop. This changed the treatment criteria to zone I or II ROP, any stage, with plus disease, or zone I, stage 3 (with or without plus disease). Notably, this shift removed the only “measurable” or quantifiable biomarker for disease severity—the number of clock hours of stage 3 that defined threshold ROP—and replaced it with the simpler but more subjective heuristic that (apart from zone I, stage 3), treatment depended only on whether plus disease was present or absent.
The ICROP was revisited shortly after the ETROP was published and introduced several new concepts, including increased recognition of the importance of zone I ROP, the introduction of aggressive posterior ROP (APROP), and the concept of “pre-plus” disease, all of which are important phenotypic characteristics that have certain epidemiologic and prognostic implications. Zone I ROP was recognized to be at higher risk for progression and was related to the main demographic risk factor—GA. The lower the GA, the more posterior the zone of ROP is, in general, and the more oxygen is required for resuscitation, on average. APROP was most common in zone I eyes and was defined both by its phenotypic appearance (flat neovascularization rather than the more traditional “staged” appearance and plus disease “out of proportion” to the degree of apparent stage) and functionally by its faster progression to retinal detachment. Finally, the introduction of pre-plus reflected the earliest acknowledgment that vascular changes associated with plus disease were not a binary disease characteristic. ,
Therapies have evolved over time, both with and without clinical trials. Between CRYO-ROP and ETROP, indirect laser photocoagulation replaced cryotherapy as the standard of care, without a pivotal study comparing the 2 treatment modalities. Subsequent to ETROP, there have been several pivotal clinical trials for anti-vascular endothelial growth factor (anti-VEGF) treatments. ,, Although these have not changed our indications for treatment, anti-VEGF has become increasingly the first line for the highest-risk eyes (e.g., zone I and APROP eyes), with multiple studies demonstrating improved anatomic outcomes with anti-VEGF compared with laser in this population. ,,
The most recent iteration of ICROP was published in 2021 with several key goals: to be more broadly representative of the spectrum of pathology found both in high- and low-income regions (second and third epidemics), to introduce new concepts unique to the anti-VEGF era, such as reactivation, and to further elaborate on the spectrum of plus disease. APROP was renamed aggressive ROP (AROP) because, in the third epidemic regions, AROP can occur in zone II eyes with high oxygen exposure. This often leads to phenotypic appearances that differ from traditional zone I “APROP.” , Whereas the 2005 ICROP introduced the concept of pre-plus, the 2021 edition formally acknowledged that dilation and tortuosity reflect a spectrum of disease severity, a finding that was largely the result of a decade of literature identifying discrepancies in the diagnosis of plus disease, in part due to systematic bias and the challenges of consistently applying a reference standard image to clinical care.
THE IMPACT OF DIGITAL IMAGING
The incorporation of digital imaging has provided a number of opportunities not only to improve ROP care but also to advance ROP research. ,,,,, Clinically, it has enabled the ability to photo-document disease for the medical record and to compare over time, which has proven to be an important capability. , It has also enabled remote “tele-ROP” monitoring systems, such as SUNDROP in the United States, KIDROP, and the Aravind Eye Care System (AECS) ROP telemedicine programs in India, which have been force multipliers in enabling ROP screening especially in rural and low-income regions. ,,
Imaging has also led to a better understanding of the patterns of ROP diagnosis (and misdiagnosis). Although a full review of this area is beyond the scope of this paper, the early work exploring interobserver variability in plus disease is relevant. Because ETROP changed practice and elevated the importance of plus disease, it became more important to understand how this diagnosis was made in practice. Despite a photographic standard for plus disease, analysis of expert diagnostic patterns revealed a high degree of disagreement. , Subsequent qualitative analysis yielded insights that not only did clinicians disagree on the outcome, they often disagreed on the process of how to make the diagnosis.
Over time, larger imaging data sets enabled larger studies of the pattern of expert behavior, and it became clear that a key underlying factor was that plus disease was a spectrum, with some experts over- and underdiagnosing compared with their colleagues, without any objective standard to apply. , It further became apparent that this bias was leading to real-world treatment differences, , complicated efforts to train ophthalmologists, and may be subject to temporal drift—that is, the level of disease at which the average clinician diagnoses “plus” might be changing over time. The presence of large data sets of ROP images also emerged as interest in machine learning and artificial intelligence (AI) in medicine was rapidly evolving, which led to attempts to classify and quantify plus disease.
APPLICATIONS OF AN ROP VASCULAR SEVERITY SCORE
There have been multiple groups and methodological approaches focused on using AI for the diagnosis of plus disease. ,,, The ROPTool was one of the earliest approaches that used a semi-autonomous method integrating user input and output quantitative assessments of tortuosity and dilation, which could then be compared with clinical diagnoses. , Subsequent work using feature extraction methods led to a classifier that could diagnose plus disease at least and clinical experts, but it required manual segmentation of the retinal vessels as input. In 2018, using the Imaging and Informatics in ROP (i-ROP) data set, the i-ROP deep learning (i-ROP-DL) algorithm used a convolutional neural network (CNN) to classify 3-level plus disease, including as international experts, without requiring manual segmentation by training a separate network to extract the vessel segmentation.
The i-ROP-DL algorithm was developed based on reference standard input images for no plus, pre-plus, and plus and outputted not only a diagnosis but also a 3-level probabilistic output that summed to 1. Using this output, the team converted the 3 probabilities (P) into a weighted “vascular severity score” (VSS) from 1 to 9, based on the rough equation 1 x P (“no plus”) + 5 x P (“pre-plus”) + 9 x P (“plus”). This resulted in a 1 to 9 scale where 1 translated to thin, narrow, and straight vessels and 9 reflected severe plus disease, typically AROP. , Validation studies demonstrated not only that it corresponded to the spectrum of plus disease among experts, but also that it increased with higher stage, higher extent of stage 3, and more posterior zone. As a result, a single photo could yield a VSS that corresponded to the overall ICROP classification of an eye. ,, Subsequently, the VSS concept has been used for a number of different applications, including assistive diagnosis, autonomous screening, longitudinal disease monitoring, , and integrated risk prediction. ,
For the purposes of this study, 2 specific applications of the VSS will be analyzed. First, although it is not surprising that the VSS correlates with plus disease diagnostic labels—indeed, it was trained on them—it was a novel finding to recognize that it correlated with more posterior zone and higher disease stage, all biomarkers of more severe ROP. However, the original description of this finding was published based on the same data set from which the algorithm itself was developed and reflected only a North American population. In this study, we chose to analyze a second data set of images from the AECS, Coimbatore ROP telemedicine program that reflects some of the phenotypic variability of both second and third epidemic ROP.
Second, because the VSS reflects a surrogate of the level of ROP in an infant, it also works for population-level epidemiologic assessment by averaging individual baby-level data into neonatal care unit (NCU) level or regional data. ,, Previous work demonstrated the value of an “institutional” VSS that could be used to compare the severity of ROP geographically and over time. ,, In the United States, these differences were associated with underlying demographic risk (mean GA and birthweight of various NCUs), whereas in an older AECS population (2015-2017), we found stronger associations with local NCU resources, such as the presence and use of oxygen blenders and pulse oximeters. We further found that the population-level VSS correlated with changing ROP epidemiology over time in South India between 2015 to 2017 and 2019 to 2020, strongly suggesting improvements in neonatal care in that region over a relatively short period of time.
In this study, we use an eye-level VSS to compare differences in the quantitative relationship between VSS and ICROP classification across the 2 populations: the i-ROP cohort and the more recent AECS cohort. The central hypothesis is that the underlying relationship between vascular severity and higher disease stage and more posterior disease zone remains intact, despite differences in underlying ROP epidemiology and phenotypic characteristics.
METHODS
ETHICAL APPROVALS
This study followed the tenets of the Declaration of Helsinki and was approved by the institutional review boards (IRB) at Oregon Health & Science University and participating i-ROP sites, as well as by the Aravind Eye Care System (AECS) IRB before the study began for prospective data collection. This specific study is a retrospective investigation of the existing data set and was covered by the original IRB at all sites. Analyses for this study were performed on deidentified clinical and imaging-derived variables. Informed consent was obtained from the parents of all babies in both countries, and no financial incentives were provided.
DATA SETS
We analyzed 2 independent data sets: the i-ROP data set and the AECS data set.
i-ROP
From January 2012 to July 2020, the i-ROP Consortium assembled a longitudinal data set of serial retinal fundus images from infants undergoing routine ROP screening (GA < 31 weeks, birth weight [BW] <1501 g). Images were acquired using a RetCam3 (Natus) with 5 standard fields of view captured per session (posterior, nasal, temporal, inferior, and superior). Bedside ophthalmoscopic examinations were performed in parallel with standardized image review by expert readers, and a reference-standard diagnosis was derived by integrating these assessments, as previously described. Images judged not acceptable for diagnosis by reader consensus, examinations diagnosed with stage 4 or 5 ROP, and post-treatment examinations were excluded.
AECS
The AECS data set was collected through a telemedicine screening program in India. Between August 2015 and October 2017, tele-medical examinations were performed using a platform designed for diabetic retinopathy tele-screening. In 2018, iTelegen, a telemedicine and data management system specifically designed for ROP, was developed and implemented, with a full transition occurring in 2019. Therefore, the AECS data set was split into 2 time periods based on which system was used: August 2015 to October 2017 and February 2019 to December 2020, the latter of which was used for the primary analysis and labeled “AECS,” unless otherwise specified. The 2015-2017 and 2019-2020 data sets screened infants across 18 and 48 neonatal care units (NCUs), respectively. Trained technicians acquired images using either a RetCam Shuttle (Natus) or a Forus 3nethra (Forus) from infants meeting the Indian screening guideline criteria (GA < 34 weeks, BW < 2000 g), capturing multiple retinal fields of view. For this data set, the telemedical diagnoses were used as the ground truth. The exclusion criteria were analogous to those of i-ROP.
QUANTITATIVE VSS ANALYSIS
Vascular severity was quantified using an AI-derived VSS from the i-ROP-DL algorithm, using methods previously published and illustrated in Figure 1 . Specifically, after automated image quality assessment, we averaged the image-level VSS output from all available images for each eye examination, yielding a single eye-level VSS per examination.
Process for creation of an eye-level VSS. All images from an eye examination are input into the algorithm. A quality step is performed as a preprocessing step that identifies the presence of an optic nerve in the image. This ensures the media are clear enough to visualize the disc, the image reflects the back of the eye, and the field of view is relatively posterior. A segmentation step is applied to images that pass the quality check. The algorithm then assigns probabilities to no plus, pre-plus, and plus, which are converted to a 1 to 9 scale for each image. All images are then averaged for an eye-level VSS. VSS = vascular severity score.
CROSS-SECTIONAL ANALYSIS
Within each data set, we tested cross-sectional associations between examination-level VSS and ICROP definitions for zone, stage, plus, and the extent of stage 3. We evaluated associations of VSS with plus disease (normal, pre-plus, plus) and ROP category (no ROP, mild ROP, type 2 ROP, or any ROP with pre-plus, TR-ROP), and stage (0-3) after stratifying by zone (I-II). We also evaluated associations of VSS with clock-hour extent of stage 3 ROP, categorized as <3, 3 to 6, and >6 hours. Extent was measured during the ophthalmoscopic examination in the i-ROP data set and by manual image review by a physician with expertise in diagnosing ROP in India via telemedicine (SN).
LONGITUDINAL ANALYSIS
Within each data set, we evaluated longitudinal patterns in examination-level VSS across PMA. PMA was binned into 2-week intervals from 30 to 40 weeks. Longitudinal models included PMA bin, treatment status (treatment-requiring [TR] vs non-TR), and their interaction to assess whether VSS trajectories differed over time between treated and untreated eyes.
STATISTICAL ANALYSIS
All analyses were performed using R (version 4.5.1). The unit of analysis was either a patient or an eye exam. Baseline demographic characteristics were compared between cohorts at the infant level using Wilcoxon rank sum tests and 2-sample proportion tests, where appropriate. All inferential analyses used generalized estimating equations (GEE) to account for correlation from repeated exams within infants (and both eyes when available). Gaussian family models with an identity link and robust (sandwich) standard errors were fit, clustering on infant. Omnibus Wald tests and prespecified pairwise contrasts estimated from model-based marginal means with Holm adjustment for multiple comparisons were used. Statistical significance was defined as P <.05.
RESULTS
PATIENT-LEVEL DEMOGRAPHICS AND EXAMINATION-LEVEL FINDINGS
Tables 1 and 2 highlight the patient-level and examination-level demographics in the 2 data sets for both TR and non-TR infants. The i-ROP data set included images from 6552 examinations from 1019 babies, and the AECS data set included images from 8281 exams from 2363 babies. Table 1 demonstrates substantial between-cohort differences in examination-level disease severity and topography. Overall, eyes in the AECS cohort were more likely to be in zone III on examination ( P <.001), whereas i-ROP had a higher proportion of zone I eyes ( P <.001). The proportion of eyes with TR-ROP was higher in i-ROP (201/6552 [3.2%]) compared with AECS (123/8281 [1.5%], P <.001). Table 2 demonstrates the demographic variables for TR and non-TR babies in both data sets. There are several notable differences. The mean GA and BW were higher in the AECS data set in both TR and non-TR groups and were significantly lower in TR babies in both data sets.
TABLE 1
Examination-Level Distribution of Zone, Stage, Plus, and Stage 3 Extent in i-ROP and AECS
| Diagnosis | i-ROP, n (%) | AECS, n (%) |
|---|---|---|
| Category | ||
| None | 2683 (42.3) | 6843 (83.0) |
| Mild | 2328 (36.7) | 981 (11.9) |
| Moderate | 1138 (17.9) | 296 (3.6) |
| TR | 201 (3.2) | 123 (1.5) |
| AROP | 46 (0.7) | 84 (1.0) |
| Plus | ||
| Normal | 5260 (80.3) | 7931 (95.8) |
| Pre-Plus | 1086 (16.6) | 306 (3.7) |
| Plus | 206 (3.1) | 44 (0.5) |
| Stage by zone | ||
| Zone I | 335 (5.1) | 88 (1.1) |
| Stage 0 | 49 (0.7) | 42 (0.5) |
| Stage 1 | 79 (1.2) | 5 (0.1) |
| Stage 2 | 104 (1.6) | 13 (0.2) |
| Stage 3 | 103 (1.6) | 28 (0.3) |
| Zone II | 6106 (93.2) | 3398 (41.0) |
| Stage 0 | 2624 (40.0) | 2066 (24.9) |
| Stage 1 | 1545 (23.6) | 541 (6.5) |
| Stage 2 | 1571 (24) | 682 (8.2) |
| Stage 3 | 366 (5.6) | 109 (1.3) |
| Zone III | 111 (1.7) | 4795 (57.9) |
| Stage 0 | 90 (1.4) | 4696 (56.7) |
| Stage 1 | 19 (0.3) | 87 (1.1) |
| Stage 2 | 2 (0.1) | 12 (0.1) |
| Stage 3 | 0 (0.0) | 0 (0.0) |
| Extent of stage 3 | ||
| <3 clock hours | 253 (53.9) | 75 (54.7) |
| 3–6 clock hours | 191 (40.7) | 58 (42.3) |
| >6 clock hours | 25 (5.3) | 4 (2.9) |
TABLE 2
Demographics of TR ROP in Both Populations
| Not Treated | Treated | |||||
|---|---|---|---|---|---|---|
| i-ROP | AECS | p value | i-ROP | AECS | p value | |
| Patients, n (%) | 865 (84.9) | 2289 (96.9) | <0.001 | 154 (15.1) | 74 (3.1) | <0.001 |
| Gestational age , mean weeks ± SD | 27.7 ± 2.3 | 33.5 ± 2.8 | <0.001 | 24.9 ± 1.3 | 30.4 ± 2.5 | <0.001 |
| Birthweight , mean weeks ± SD | 1017 ± 320 | 1732 ± 438 | <0.001 | 661 ± 173 | 1291 ± 291 | <0.001 |
AECS = Aravind Eye Care System.
CROSS-SECTIONAL ANALYSIS
Figure 2 demonstrates the relationship between the eye-level VSS and the ICROP classification of zone, stage, extent of stage 3 disease, plus disease, and overall category. There was a direct relationship between the VSS and the clinical diagnosis of stage in both data sets ( Figure 2 , A, all P <.001). In i-ROP, we found that for stages 1 ( P =.044), 2 ( P <.001), and 3 ( P <.001), VSS was higher in more posterior zones ( Figure 2 , B ). This finding held for stages 1 ( P <.001) and 3 ( P <.001) in the AECS data set. Traditionally, ROP classification has recorded the “extent” (number of clock hours) of each stage. In these data sets, we only had labels for the extent of stage 3 but found that the VSS increased with a higher extent of stage 3 in both i-ROP ( P =.004) and AECS ( P <.001), as found in Figure 2 , C . Unsurprisingly, there was a relationship between the VSS and the clinical diagnosis of plus disease ( Figure 2 , D , indeed, the VSS algorithm was developed from the plus labels of the i-ROP data set), but confirmation in a new data set is novel. Finally, when we evaluated VSS across the overall category of disease (none, mild, type 2/preplus, or type 1/treatment requiring), VSS increased with worsening ROP in both data sets ( Figure 2 , E , both P <.001). Interestingly, at all levels of comparison, the VSS was lower in the AECS data set compared with i-ROP, including stage (all P <.001, Figure 2 , A), plus disease (all P <.001, Figure 2 , D ), and category (“None” P =.002, all other P <.001, Figure 2 , E ).
Cross-sectional comparisons of i-ROP and AECS data sets across ROP risk factors. (A) VSS increased with worsening ROP stage in both data sets (all P <.001), with consistently lower VSS in AECS than i-ROP at a given stage (all P <.001). (B) For a given stage, VSS tended to be lower in more posterior zones. In i-ROP, this was significant for stage 1 ( P =.044), 2 ( P <.001), and 3 ( P <.001), and in AECS for stages 1 ( P <.001) and 3 ( P <.001). (C) Among stage 3 examinations, VSS increased with greater clock-hour extent of disease (i-ROP P =.004, AECS P <.001). (D) VSS increased with plus disease severity (all P < 0.001) and (E) overall ROP severity, with i-ROP demonstrating higher VSS than AECS within severity strata (all P <.001). AECS = Aravind Eye Care System; VSS = vascular severity score.
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree