Topic
Minimally invasive glaucoma surgery (MIGS) has been widely adopted as a surgical option for glaucoma, yet concerns remain regarding the robustness of the evidence supporting many procedures. The fragility index (FI) and fragility quotient (FQ) provide complementary measures to traditional p values by quantifying how dependent statistically significant results are on a small number of outcome events.
Clinical Relevance
Understanding the statistical fragility of MIGS trials is important for clinicians, guideline developers, and policymakers when interpreting reported efficacy and safety outcomes, particularly in a field characterised by small trials, heterogeneous endpoints, and commercially sensitive interventions.
Methods
A systematic review was conducted in accordance with PRISMA guidelines and registered on PROSPERO. Electronic databases were searched from inception to August 1, 2025, to identify randomized controlled trials evaluating MIGS procedures in patients with glaucoma. Trials reporting at least one statistically significant binary outcome with sufficient raw data were included. The FI was calculated using iterative Fisher exact testing, and the FQ was derived by normalizing FI to total sample size. Risk of bias was assessed using the Revised Cochrane Risk of Bias Tool (RoB 2.0).
Results
Sixteen randomized controlled trials encompassing 4,562 participants met inclusion criteria. The median sample size was 236.5 (interquartile range [IQR], 88–511), with a median follow-up of 24 months. The mean FI was 9.4 ± 10.0 and the mean FQ was 0.031 ± 0.025, indicating that approximately 3% of participants would require a different outcome to render results non-significant. Three trials (18.8%) had an FI of zero. In six studies (37.5%), the FI was less than or equal to the number of participants lost to follow-up, suggesting vulnerability to attrition bias. Fragility varied across device categories, with larger multicenter trials demonstrating higher FI values but consistently low FQ values.
Conclusion
Statistically significant outcomes in RCTs of MIGS exhibit variable statistical robustness when assessed using fragility metrics, often due to small numbers of events. These findings highlight the need for cautious interpretation of reported significance and support the incorporation of fragility metrics alongside conventional statistical measures when appraising the MIGS evidence base.
Introduction
G laucoma remains one of the leading causes of irreversible blindness worldwide, affecting more than 75 million individuals and projected to exceed 110 million by 2040. The foundation of management continues to be reducing intraocular pressure (IOP), which remains the only modifiable risk factor shown to reduce disease progression. Traditional filtering operations such as trabeculectomy and aqueous shunts achieve significant IOP reduction but are linked to potentially vision-threatening complications, prolonged recovery, and decreasing success over time. Over the past decade, there has been a paradigm shift towards procedures that require less surgical time and balance effectiveness with safety, collectively known as minimally invasive glaucoma surgery (MIGS).
MIGS procedures aim to achieve meaningful IOP reduction through micro-incisional or ab interno approaches that reduce tissue disruption and postoperative follow-up. By targeting established outflow pathways such as trabecular, canalicular, subconjunctival, or suprachoroidal these procedures offer an alternative to traditional filtration surgery for patients with mild to moderate disease, or for those in whom long-term medication burden or intolerance limits medical treatment. These devices have broadened the surgical options, reshaping expectations around safety, recovery, and the timing of surgical intervention.
Although the adoption of MIGS has been rapid, the strength and consistency of its evidence base remain uncertain. Many randomized controlled trials (RCTs) have been underpowered with short follow-up, and frequently industry-sponsored. , Across studies, there is marked heterogeneity in inclusion criteria, comparator arms, surgical technique, and the definition of surgical success, complicating meaningful comparison across device categories. Several systematic reviews have noted that the overall certainty of evidence supporting MIGS remains moderate to low, mainly due to imprecision, lack of masking, incomplete allocation concealment, and high attrition rates. ,,, Furthermore, a large proportion of trials report marginally significant outcomes without proper adjustment for multiple testing or loss of follow-up, raising concerns about the reproducibility of these findings. Interpreting statistical significance in this context is particularly challenging, as P -values alone do not indicate how reliant a reported result may be on a small number of events.
The fragility index (FI) has emerged as a complementary measure to address this limitation. Introduced by Walsh and colleagues in 2014, the FI represents the minimum number of event status changes required to convert a statistically significant result ( P <.05) into a non-significant one. A low FI indicates that the reported significance is highly dependent on only a few outcome events and, therefore, statistically fragile. The fragility quotient (FQ) expands on this concept by adjusting the FI for the total sample size, enabling comparisons across different trial scales. Together, these metrics offer an intuitive and quantitative way to assess the robustness of trial results.
The aim of this systematic review is to quantify the FI and FQ of statistically significant outcomes reported in MIGS RCTs. By assessing the robustness of current evidence across various device categories and comparator groups, this study aims to provide a clearer understanding of the reliability of MIGS trial data and to identify areas where methodological improvements are needed. In doing so, it offers a framework for more critical interpretation of statistical significance in glaucoma surgical research and to improve the quality and transparency of evidence on the use of MIGS.
METHODS
Protocol and Registration
This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. The review protocol was developed and the literature search conducted in August 2025, with formal PROSPERO registration completed in October 2025 (CRD420251166050). Data screening, extraction, and analysis were performed only following registration. The methodology employed was from established approaches from previous fragility analyses in surgical and ophthalmic trials. ,, As this study is a systematic review of previously published data, Institutional Review Board approval and informed consent were not applicable. The study adhered to the principles of the Declaration of Helsinki.
Search Strategy
A comprehensive electronic search of PubMed/MEDLINE, Embase (Ovid), and the Cochrane Central Register of Controlled Trials (CENTRAL) was performed from database inception to 1st August 2025, without language or date restrictions. The search strategy combined Medical Subject Headings (MeSH) and free-text terms; the full list of search terms used can be found in Appendix 1. Reference lists of included studies and relevant systematic reviews were manually screened to identify additional eligible trials. Duplicates were removed using Rayyan, EndNote and manual review. ,
Eligibility Criteria
Studies were included if they were randomized controlled trials (parallel arm or two-by-two factorial design) enrolling patients diagnosed with any form of glaucoma. Eligible interventions comprised any MIGS procedure, including but not limited to trabecular, canalicular, subconjunctival, or suprachoroidal approaches. Comparators included medical therapy, sham surgery, other MIGS procedures, or conventional glaucoma surgeries such as trabeculectomy or aqueous shunt implantation. Trials were required to report at least one statistically significant binary outcome, such as surgical success or failure, ≥20% IOP reduction, the need for reoperation, or the occurrence of complications. Only studies providing sufficient raw binary data, including the number of events and total sample size in each arm, to enable calculation of the FI were considered eligible for inclusion.
Trials were excluded if they reported only continuous outcomes, were quasi-randomized, or lacked adequate randomization or allocation concealment. When multiple publications arose from the same randomized trial, each was analyzed as a separate study because the populations, follow-up periods, and outcomes often differed. Long-term extensions frequently involved only subsets of the original cohort, assessed distinct endpoints or had different outcomes.
Study Selection
All titles and abstracts were screened independently and in duplicate by 2 reviewers (A.S.A, M.J.) using Rayyan. Full-text articles of potentially eligible studies were subsequently retrieved and reviewed. Discrepancies were resolved through consensus. The selection process was documented using a PRISMA flow diagram ( Figure 1 ).
PRISMA flow diagram of included studies
Data Extraction
Data extraction was carried out independently by 2 reviewers (A.S.A, M.J) using a standardized, pilot-tested Excel sheet. The following information was collected from each study: first author, year of publication, journal, country or geographic region, and whether the trial was single-center or multicenter in design. Details of the study population and inclusion criteria were recorded, along with information on the intervention and comparator arms, the type of MIGS device evaluated, and the definitions of primary and secondary outcomes. The total sample size and number of events in each treatment arm were extracted, together with reported P -values, CIs, and the statistical tests employed. Data on follow-up duration, attrition rate, and funding source (industry vs non-industry) were also collected. Any discrepancies between reviewers were resolved through consensus, and where information was incomplete or unclear, corresponding authors were contacted for clarification.
Risk of Bias Assessment
Risk of bias was independently assessed by both reviewers using the Revised Cochrane Risk of Bias Tool for Randomized Trials (RoB 2). The following domains were evaluated: (1) randomization process, (2) deviations from intended interventions, (3) missing outcome data, (4) measurement of outcomes, and (5) selection of reported results. Each domain was rated as “low risk,” “some concern,” or “high risk,” with an overall judgment made accordingly. Disagreements were resolved through discussion and consensus.
Statistical Analysis and Fragility Index Calculation
Descriptive statistics were used to summarize study characteristics and fragility metrics. Continuous variables were reported as medians with IQRs (IQR), and categorical variables as counts and percentages. All statistical analyses were performed using R (version 4.3.2; R Foundation for Statistical Computing) and Stata (version 18.0; StataCorp LLC).
The FI was calculated for each eligible trial following the methods described by Walsh et al. Event and non-event counts from each treatment arm were entered into a 2 × 2 contingency table using a publicly available online calculator. In summary, an event was iteratively added to the group with the smaller number of events, with a corresponding non-event removed from the same group to maintain the total sample size. After each iteration, the 2-sided P -value from Fisher’s exact test was recalculated. The FI was defined as the minimum number of non-events that required conversion to events for the p -value to reach non-significance (i.e., ≥0.05). The FQ was calculated by dividing the FI by the total sample size of the trial (FQ = FI/ N), allowing for comparison across studies of different sizes. Trials that originally reported P -values derived from χ² (Chi-squared) tests were recalculated using Fisher’s exact test to ensure methodological consistency. In multi-arm trials, pairwise comparisons were performed between each intervention and the control arm when both demonstrated statistical significance. For cluster-randomized trials, the effective sample size was used to account for within-cluster correlation. When trials reported multiple statistically significant binary outcomes, a single outcome was selected per trial for the primary analysis using a pre-specified hierarchy: (1) the trial designated primary binary outcome; (2) complete surgical success; (3) qualified surgical success; (4) need for secondary intervention; (5) achievement of a pre-specified IOP reduction threshold, in instances where the outcome did not meet this criteria we reported the first binary outcome reported in the results.
Exploratory analyses were planned to assess potential associations between FI and study-level factors such as journal impact factor, sample size, and event count. However, the small number of eligible trials rendered the review underpowered for meaningful inferential statistics, and these analyses were not pursued.
Sensitivity analyses examined whether trial fragility was influenced by loss to follow-up or study characteristics. Trials with FI less than or equal to the number of participants lost to follow-up were identified as particularly vulnerable to attrition bias. Subgroup trends by funding source, sample size, and study design were also reviewed qualitatively, though formal statistical comparisons were not undertaken due to limited study numbers and heterogeneity of outcomes.
RESULTS
Study Selection and Characteristics
Our search identified 531 records, including 467 from electronic databases and 64 from citation searching and prior systematic reviews on MIGS RCTs. After removing 231 duplicates, 300 unique records remained for title and abstract screening. Of these, 248 were excluded due to incorrect study population, design, or absence of a binary outcome of interest. Fifty-two full-text articles were reviewed in detail, and 36 were excluded for not being RCTs, not evaluating MIGS, or lacking extractable binary outcome data. A total of 16 randomized controlled trials met the inclusion criteria and were included in the final analysis ( Figure 1 ). ,,,,,,,,,,,,,,,
These 16 RCTs involved 4562 participants, with individual sample sizes from 52 to 1184 eyes and a median of 236.5 (IQR: 88–511). Twelve trials (75%) were multicenter, and 4 (25%) were single-center. The median follow-up duration was 24 months (range: 6–36 months). Most studies focused on adults with primary open-angle glaucoma or ocular hypertension, although 4 also included secondary glaucomas. Among the included RCTs, 7 (43.8%) employed intention-to-treat (ITT) analysis, 7 (43.8%) used modified intention-to-treat (mITT) analysis, and 2 (12.5%) used per-protocol (PP) analysis for their primary efficacy endpoints. Full characteristics of trials included can be seen in Table 1 .
Table 1
Characteristics of Included Studies n = 16.
| Authors (Year) | Total sample size [n (n arm1, n arm2)] | No. lost to follow-up (%) | p -value | Journal Impact factor | Fragility index | Fragility Quotient | Overall ROB |
|---|---|---|---|---|---|---|---|
| Baker et al., 2021 | 527 (395, 132) | 21 (3.98%) | <.01 | 9.5 | 37 | 0.07 | Low |
| Gupta et al., 2025 | 52 (22, 30) | Not reported | .009 | 4.2 | 0 | 0 | Some concern |
| Ye et al., 2025 | 52 (26, 26) | 2 (3.85%) | .0127 | 2.8 | 2 | 0.04 | Some concern |
| Vold et al., 2016 | 505 (332, 116) | 25 (4.95%) | .001 | 9.5 | 9 | 0.02 | Some concern |
| Pfeiffer et al., 2015 | 100 (50, 50) | 7 (7%) | .0008 | 9.5 | 8 | 0.08 | Low |
| Dada et al., 2025 | 30 (15, 15) | 6 (20%) | .05 | 1.8 | 0 | 0 | Some concern |
| Samuelson et al., 2019 | 556 (369, 187) | 17 (3.06%) | <.001 | 9.5 | 22 | 0.04 | Some concern |
| Craven et al., 2012 | 240 (116, 123) | 7 (2.92%) | .036 | 3.2 | 0 | 0 | Some concern |
| Panarelli et al., 2024 | 527 (395, 132) | 9 (1.71%) | .005 | 9.5 | 15 | 0.03 | low |
| Chen et al., 2020 | 32 (16, 16) | 0 (0%) | .02 | 5.6 | 1 | 0.03 | low |
| Falkenberry et al., 2020 | 164 (79, 78) | 7 (4.27%) | .04 | 3.2 | 1 | 0.01 | Some concern |
| Samuelson et al., 2019 | 505 (380, 118) | 18 (3.56%) | .005 | 9.5 | 6 | 0.01 | low |
| Samuelson et al., 2011 | 233 (117, 123) | 7 (3%) | <.001 | 9.5 | 11 | 0.05 | low |
| Jones et al., 2019 | 331 (219, 112) | 10 (3.02%) | <.001 | 3.2 | 16 | 0.05 | low |
| Ahmed et al., 2020 | 152 (73, 75) | 4 (2.63%) | <.001 | 9.5 | 8 | 0.05 | low |
| Ahmed et al., 2022 | 556 (308, 134) | 15 (2.7%) | <.001 | 9.5 | 15 | 0.03 | Some concern |
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree