Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Inference of Visual Acuity and Postoperative Outcomes from Routine Biometry Images Using Pretrained Embedding Architectures and Machine Learning Techniques: A feasibility Study
ABSTRACT
Background:
As the world’s most frequent surgical procedure, cataract surgery is essential for restoring vision and quality of life. However, routine diagnostic data remains underutilized e.g. for predicting functional outcomes, which is vital for managing patient expectations. This study uses deep learning to leverage such data, providing an objective basis for personalized surgical counseling.
Objective:
To evaluate the feasibility of a deep learning framework for objectively inferring current visual acuity (VA) and forecasting postoperative VA improvement by leveraging latent anatomical data from routine biometry imaging.
Methods:
A machine learning pipeline was developed using the PyCaret library to process high-dimensional embeddings. Features were extracted from standardized image types using three state-of-the-art architectures for comparison: OpenAI’s CLIP (ViT-B/32), Meta’s DinoV2 (ViT-S/14), and the YOLO11n-cls classification backbone. Extracted features were concatenated and classified into discrete VA and logMAR gain categories. The study analyzed images from the ZEISS IOLMaster 700 (anterior segment OCT, frontal eye photographs, and foveal fixation OCT) across two cohorts: a large-scale same-day inference dataset consisting of 15,074 eyes and a longitudinal postoperative dataset of 650 eyes.
Results:
The YOLO11n-cls architecture consistently yielded the most clinically relevant feature space. For same-day VA inference, a YOLO-based Logistic Regression model achieved a macro-averaged AUC of 0.80, peaking at 0.85 for high-acuity cases. In the postoperative task, a YOLO-XGBoost configuration achieved an AUC of 0.84 for predicting significant visual gains (≥0.4 logMAR improvement).
Conclusions:
Routine biometry images contain a robust biological signal for visual potential. YOLO-based architectures demonstrate superior efficacy in feature extraction for standardized ophthalmic imaging, providing a viable pathway for objective surgical counseling and preoperative decision support. The combination with ensemble learning techniques seems to be a good way of extracting clinically relevant features.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.