Deep Learning-Based Diagnostic Model for Ocular Surface Neoplastic Diseases

Purpose

To develop a deep learning (DL) model for diagnosing ocular surface tumors and evaluating its diagnostic performance.

Setting

Development of a deep learning diagnosis algorithm.

Methods

A total of 1491 ocular surface images representing 7 diseases—nevus (28 eyes), limbal dermoid (144), MALT lymphoma (20), ocular surface squamous neoplasia (OSSN; 138), melanoma (14), pinguecula (29), and pterygium (1,118)—were captured using slit-lamp microscopy. A YOLOv5-based DL model was trained using 5-fold cross-validation. Diagnostic performance was compared using 299 external validation images assessed by 8 corneal specialists, 7 board-certified ophthalmologists, and 8 residents.

Results

The model achieved a positive predictive value (PPV) of 96.0%, outperforming the corneal specialists (95.0 ± 2.1%), board-certified ophthalmologists (82.6 ± 11.9%), and residents (81.9 ± 10.7%). Disease-specific PPVs were: nevus 75.0%, limbal dermoid 93.5%, MALT lymphoma 57.9%, OSSN 87.0%, melanoma 38.5%, pinguecula 82.1%, and pterygium 96.7%. The area under the curve (AUC) was: nevus 0.897 (95% confidence interval [CI], 0.810-0.983), limbal dermoid 0.998 (95% CI, 0.996-1.000), MALT lymphoma 0.894 (95% CI, 0.794-0.993), OSSN 0.954 (95% CI, 0.933-0.975), melanoma 0.966 (95% CI, 0.919-1.000), pinguecula 0.954 (95% CI, 0.912-0.995), and pterygium 0.984 (95% CI, 0.976-0.992). Sensitivities were: nevus 0.643, limbal dermoid 0.993, MALT lymphoma 0.550, OSSN 0.681, melanoma 0.357, pinguecula 0.793, and pterygium 0.991. Specificities were: nevus 0.996, limbal dermoid 0.993, MALT lymphoma 0.995, OSSN 0.990, melanoma 0.995, pinguecula 0.997, and pterygium 0.898.

Conclusions

The deep learning model demonstrated high diagnostic accuracy for common ocular surface tumors such as pterygium and limbal dermoid, while diagnostic performance for rare malignancies, including melanoma and MALT lymphoma, remains limited and requires further refinement.

Synopsis

We developed a deep learning model that demonstrated promising performance in identifying ocular surface neoplastic diseases, suggesting its potential as a supportive diagnostic tool in ophthalmic practice.

Artificial intelligence (AI) has emerged as a transformative technology across various domains of healthcare, driven by remarkable progress in computational power and machine learning algorithms. In the medical field, AI has already shown high performance in image-based diagnostics, aiding in the detection of diseases such as cancer, skin disorders, cardiovascular conditions, and ophthalmic diseases. ,,,, Within ophthalmology, the application of AI—particularly deep learning (DL) techniques—has expanded rapidly in recent years. Much of this progress has focused on posterior segment diseases, where AI models have achieved expert-level accuracy in detecting diabetic retinopathy (DR), age-related macular degeneration (AMD), and glaucoma. These developments have been widely validated and integrated into clinical workflows and telemedicine platforms.

Recently, a DL model known as “CorneAI” was developed to classify various corneal diseases into nine anterior segment categories. While this represents an important advancement, the classification remains broad, and disease-specific diagnosis within each category has yet to be fully addressed. Current research efforts are focused on building more refined DL models tailored to individual conditions, such as infectious keratitis, immunological keratitis (Shijo et al. under review), and corneal deposits. Of the nine categories in CorneAI, ocular surface tumors have not been extensively investigated in the context of AI-driven diagnostics. These conditions—ranging from benign lesions like pterygium to life-threatening malignancies such as ocular surface squamous neoplasia (OSSN) and malignant melanoma—require precise and early diagnosis. Furthermore, diagnosing ocular surface tumors can be sometimes challenging even for board-certified ophthalmologists. Several reports have demonstrated efficient differentiation between OSSN and normal conditions using DL models. ,, Recently, Li et al. reported a domain-specific pretrained model to classify ocular surface lesions into malignant, premalignant and benign categories; however, comprehensive multi-class DL models for the diagnosis of ocular surface tumors remain limited. In this study, we developed a comprehensive DL model to classify anterior segment neoplastic lesions into 7 categories: conjunctival nevus, limbal dermoid, MALT lymphoma, OSSN, melanoma, pinguecula, and pterygium.

MATERIALS AND METHODS

Ethical Considerations

Ethical approval was granted by the Institutional Review Board of the Japanese Ophthalmological Society (Protocol No. R03-108). All study procedures were conducted in accordance with the ethical standards of the Declaration of Helsinki and the Japanese Guidelines for Research on Human Subjects in Life Sciences and Medicine.

Data Source and Image Collection

We retrospectively collected a dataset of 1652 anterior segment images from the Japan Ocular Imaging Registry (JOIR), incorporating data from 23 tertiary ophthalmology centers and 6 affiliated hospitals in Japan: Tsukuba University, Tokyo Dental College Ichikawa General Hospital, Kyoto Prefectural University of Medicine, Tottori University Hospital, Juntendo University Hospital, Miyata Eye Hospital and Fukushima Medical University. Anterior segment photographs were obtained using slit-lamp microscopes with built-in cameras, captured under diffuse lighting and set at either 10 × or 16 × magnification. Slit-lamp microscopy used included models from Haag-Streit (Köniz, Switzerland), ZEISS (Oberkochen, Germany), and Takagi (Nagano, Japan). All images were stored in JPEG format. Images were excluded if they met one or more of the following conditions: (1) obtained after corneal transplantation; (2) displayed slit-light artifacts or fluorescein dye staining; (3) exhibited poor focus, improper centering, or insufficient lighting. A total of 161 images were excluded, yielding 1491 images for final analysis ( Figure 1 A).

Figure 1

Study design and datasets. A. Selection process and representative images of 7 ocular surface tumors. A total of 1652 images collected from Japan Ocular Imaging Registry (JOIR) and other sources; 161 excluded for poor quality; 1491 included. B. The dataset comprised nevus (28 eyes), limbal dermoid (144 eyes), MALT lymphoma (20 eyes), ocular surface squamous neoplasia (OSSN; 138 eyes), melanoma (14 eyes), pinguecula (29 eyes), and pterygium (1118 eyes). C. Confusion matrix illustrating the predictive performance of the YOLOv5 model. Each element ( i, j ) denotes the empirical probability that a lesion with true label i is predicted as class j by the model.

Image Annotation and Disease Classification

Corneal specialists (T.S., Y.U., H.M., D.M., H.F., R.N., K.M., T.I., H.M., Y.O. and T.Y.) at each institution reviewed and categorized the images into 1 of 7 ocular surface tumor and tumor-like lesion types: nevus, limbal dermoid, MALT lymphoma, OSSN, melanoma, pinguecula, or pterygium. Classification decisions were based on comprehensive clinical information, including medical history, examination findings, disease progression, and response to treatment. Although diagnoses of melanoma, nevus, OSSN, and MALT lymphoma were confirmed histopathologically, those of pterygium, limbal dermoid, and pinguecula were based on clinical findings. To ensure reliable ground truth labeling, the diagnoses of all images were reviewed, discussed and determined by 2 corneal specialists (T.S. and T.Y.), and discrepancies were resolved through consensus. This procedure aimed to minimize interobserver variability and ensure label accuracy, thereby improving the reliability of the training dataset. The final dataset included: 28 images of nevus, 144 of limbal dermoid, 20 of MALT lymphoma, 138 of OSSN, 14 of melanoma, 29 of pinguecula, and 1118 of pterygium ( Figure 1 B).

YOLOv5 object detection algorithm

A DL model based on YOLOv5s was employed for image classification. , YOLOv5 is an end-to-end DL architecture for object detection. The model consists of 3 components: backbone, neck, and head modules ( Figure 2 ). The backbone module applies convolution layers to the input image to extract multiscale image features. The module outputs image features of different scales from 3 layers at different network depths. The final part of the backbone is a spatial pyramid pooling (SPP) module, which combines multiscale information into feature maps. The 3 outputs from the module contain rich multiscale image features with multiscale contextual information. The neck module integrates multiscale image features extracted from the backbone module and refines them to obtain global, high-level image features. High-level image features contain not only information about the target object but also about surrounding objects and their relationships. High-level image features significantly contribute to the output of the YOLOv5 head module. The head module makes the output using 3 prediction heads. The use of 3 prediction heads enables the detection of objects ranging from small to large in the image. The prediction heads output detection results, bounding boxes, category labels and likelihood scores, described above.

Figure 2

Architecture of YOLOv5: backbone, neck, and head modules. The backbone extracts multiscale image features using convolutional layers and a spatial pyramid pooling (SPP) module. The neck integrates and refines multiscale features to obtain high-level representations. The head employs 3 prediction heads to detect objects at multiple scales.

Deep Learning Model Development

A total of 1491 slit-lamp images were used, and 5-fold cross-validation was applied to divide the dataset for training and validation. The hyperparameter values of YOLOv5 were set as follows: initial learning rate was 0.01, final learning rate was 0.1, warmup epochs were 3, class binary cross-entropy (BCE) loss weight was 1.0, and object BCE loss weight was 1.0. During the training process, data augmentations were applied to images for training to improve the generalization ability of the YOLOv5 model. The data augmentations include random color change along HSV color channels, random scaling, random translation, and left-right flipping. The model was trained for 400 epochs with a mini-batch size of 16. To reduce the risk of overfitting, cross-validation was combined with a likelihood threshold of 0.05 during the validation phase. For each image, the YOLOv5 detection and classification model generated multiple candidate bounding boxes, each associated with a predicted disease label and confidence score (likelihood). For evaluation, only the bounding box with the highest score was considered from each image. The likelihood of the AI model was derived using the sigmoid function.

s b , c ( x b , c ) = 1 1 + e x b , c ,

Specifically, in the DL-based model, the sigmoid function is applied to the feature values entering the final output layer, as expressed by the following equation: where s b,c denotes the predicted score, x b,c is the feature value supplied to the final layer, b represents the index of the predicted bounding boxes, and c = 1,…,4 indicates the index of the categories.

Comparison with ophthalmologists

To cf the model’s diagnostic performance with human examiners, a subset of 299 validation images from one of the 5 folds was used as a test dataset. Twenty-three ophthalmologists participated: 8 corneal specialists, 7 board-certified ophthalmologists, and 8 residents with 1 to 4 years of ophthalmology experience. Each physician reviewed and diagnosed the same image set, and their performance was compared with that of the DL model.

Grad-CAM++ Explainability Analysis

To visualize the regions influencing the model’s classification decisions, Gradient-weighted Class Activation Mapping (Grad-CAM), and its generalized variant Grad-CAM++, were applied. The YOLOv5 head module comprises 3 prediction heads at layers 17, 20, and 23. A prior study demonstrated that model attention was predominantly concentrated in the deepest prediction head (layer 23), with minimal activation in the shallower heads (layers 17 and 20); this was confirmed in the present study. Therefore, heatmaps were generated from the final convolutional layer of the deepest prediction head (layer 23, cv3 convolution). To quantify model attention, the area of interest at the 0.5 threshold (AOI₅₀) was defined as the proportion of pixels within the detected bounding box where the Grad-CAM++ saliency value exceeded 0.5, and was compared across the 7 disease classes.

Statistical Analysis

All statistical analyses were conducted using R ver. 4.4.3 (R Foundation for Statistical Computing, Vienna, Austria) and EZR version 1.68 for Windows. Receiver operating characteristic (ROC) curves were generated based on the likelihood scores for the 7 disease categories, and the area under the curve (AUC) with 95% CIs (CIs) was calculated using the DeLong method. The model’s diagnostic performance was further evaluated by calculating the positive predictive value (PPV), sensitivity, and specificity, along with their corresponding 95% CIs. Confusion matrices were generated using Python version 3.12.5. AOI₅₀ scores were compared across the 7 disease classes using the Kruskal-Wallis test. Pairwise comparisons were performed using the Mann-Whitney U test with Bonferroni correction.

RESULTS

Diagnostic Accuracy of the Deep Learning Model

The classification outcomes of the DL model across the 7 anterior segment diseases are summarized in the confusion matrix shown in Figure 1 C. Based on this matrix, performance metrics including PPV, sensitivity, specificity, F1 score, and Youden Index were calculated and are summarized in Table 1 . Figure 3 shows the ROC curves for the DL algorithm for each disease. The AUC for each disease were as follows: nevus, 0.897 (95% CI, 0.810-0.983); limbal dermoid, 0.998 (95% CI, 0.996-1.000); MALT lymphoma, 0.894 (95% CI, 0.794-0.993); OSSN, 0.954 (95% CI, 0.933-0.975); malignant melanoma, 0.966 (95% CI, 0.919-1.000); pinguecula, 0.954 (95% CI, 0.912-0.995); and pterygium, 0.984 (95% CI, 0.976-0.992).

Table 1

Performance of YOLO v5 for 7 Categories

PPV (95% CI) Sensitivity (95% CI) Specificity (95% CI) F1 Score Youden Index
Nevus 0.750 (0.533-0.902) 0.643 (0.441-0.814) 0.996 (0.991-0.998) 0.643 0.639
Limbal dermoid 0.935 (0.883-0.968) 0.993 (0.962-1.000) 0.993 (0.986-0.996) 0.962 0.986
MALT
lymphoma
0.579 (0.335-0.797) 0.550 (0.315-0.769) 0.995 (0.989-0.998) 0.564 0.545
OSSN 0.870 (0.792-0.927) 0.681 (0.596-0.758) 0.990 (0.983-0.994) 0.765 0.671
Malignant melanoma 0.385 (0.139-0.684) 0.357 (0.128-0.649) 0.995 (0.989-0.998) 0.371 0.352
Pinguecula 0.821 (0.631-0.939) 0.793 (0.603-0.920) 0.997 (0.992-0.999) 0.807 0.790
Pterygium 0.967 (0.955-0.976) 0.991 (0.984-0.996) 0.898 (0.863-0.927) 0.979 0.889

PPV: positive predictive value, CI: confidence interval.

Figure 3

Diagnostic performance of the deep learning model in classifying 7 ocular surface tumors . Receiver operating characteristic (ROC) curves depict the discriminative capacity of the YOLOv5-based model for each tumor type.

Comparison of Diagnostic Performance Between DL Model and Ophthalmologists

Figure 4 A-D presents the confusion matrices for YOLOv5 and the median performance of each group of human examiners. The PPVs were 96.0% for YOLOv5, 95.0 ± 2.1% for corneal specialists, 82.6 ± 11.9% for board-certified ophthalmologists, and 81.9 ± 10.7% for residents ( Figure 4 E).

Figure 4

Confusion matrices for YOLOv5 and ophthalmologists. Confusion matrices for YOLOv5 (A), corneal specialists (B), board-certified ophthalmologists (C), and residents (D). E. Comparison of positive predictive value (PPV) between the YOLOv5 model and ophthalmologists.

Representative Cases of DL Model Performance

Figure 5 A-D presents examples in which the DL model correctly predicted the diagnosis, whereas many ophthalmologists misclassified the lesions. In the OSSN case ( Figure 5 A), the DL model achieved a correct diagnosis with a likelihood of 0.93, while many ophthalmologists misclassified it as MALT lymphoma or pinguecula. In the limbal dermoid case ( Figure 5 B), the DL model provided a correct prediction with a likelihood of 0.92, whereas many ophthalmologists diagnosed it as OSSN. For the pterygium case ( Figure 5 C), the DL model achieved a correct prediction with a likelihood of 0.92, while many board-certified ophthalmologists and residents misclassified the case. In the melanoma case ( Figure 5 D), the DL model correctly diagnosed the lesion with a likelihood of 0.80; however, more than half of the ophthalmologists misclassified it as a nevus. In contrast, Figure 5 E–H show examples misclassified by the DL model. These included an OSSN misclassified as pterygium ( Figure 5 E), a conjunctival nevus misclassified as pterygium ( Figure 5 F), a melanoma incorrectly predicted as pterygium ( Figure 5 G), and a pinguecula erroneously diagnosed as a limbal dermoid ( Figure 5 H).

Sep 20, 2026 | Posted by in OPHTHALMOLOGY | Comments Off on Deep Learning-Based Diagnostic Model for Ocular Surface Neoplastic Diseases

Full access? Get Clinical Tree

Get Clinical Tree app for offline access