Tipo de contenido material para medios audiovisuales:
Comienzo del material para medios audiovisuales:
Duración del material para medios audiovisuales:
Conventional CNN models have demonstrated strong performance in single-task ophthalmic diagnoses, yet they carry substantial limitations: heavy reliance on large-scale expert-annotated data, limited generalizability, and an inability to integrate multiple imaging modalities with clinical text—a prerequisite for real-world diagnostic reasoning. Foundation models address these constraints through self-supervised learning, acquiring generalizable representations from massive unlabeled datasets. RETFound validated the feasibility of this approach but does not support cross-modal fusion. Subsequent models extended the frontier: VisionFM supports eight imaging modalities via a modality-agnostic decoder; EyeCLIP integrates 11 imaging modalities with clinical text; EyeFM combines five imaging modalities with a large language model; MIRAGE focuses on paired OCT and scanning laser ophthalmoscopy learning. The review notes that multimodal fusion demonstrates performance advantages over single-modal counterparts in few-shot and zero-shot settings.
Clinical validation depth varies considerably. The strongest evidence comes from EyeFM’s double-blind randomized controlled trial: the AI-assisted group achieved higher diagnostic accuracy than the independent physician group (92.2 % vs. 75.4 %), with an average 63.3-second reduction in report writing time. This is the only RCT-based validation in the field to date. The other four models have been evaluated exclusively on retrospective datasets. The review explicitly notes that, aside from EyeFM, the remaining models lack prospective, multicenter study designs.
The review also summarizes exploratory results on systemic disease prediction. RETFound achieved AUROCs of 0.737, 0.794, and 0.754 for predicting myocardial infarction, heart failure, and ischemic stroke, respectively. VisionFM estimated 38 systemic biomarkers from ophthalmic images with a mean accuracy of 78.6 %. These results represent statistical associations reflecting model performance on specific datasets; the review does not extrapolate them as causal evidence. Key limitations remain: training data lack demographic diversity; most models use only 2D OCT slices; interpretability relies on post-hoc methods that do not reveal causal mechanisms; and nearly all models lack prospective validation. The review concludes that the core value of ophthalmic foundation models is shifting from technical performance to clinical integration, with systematic prospective validation being the prerequisite for clinical translation.
The work titled “Foundation Models in Diagnosis of Ophthalmic Diseases: Construction and Application”, was published on Eye & ENT Research (published on June 4, 2026).
DOI:10.1002/eer3.70041
Regions: Asia, China, North America, United States
Keywords: Science, Life Sciences