Chapter 4 · complete English translation
AI in Medicine and Ophthalmology
Chapter 4: Applications of Artificial Intelligence in Medicine
- Introduction
- Examples of AI techniques in clinical practice
- Examples of AI techniques in ophthalmology
- Examples of AI for the detection of keratoconus
4.1 Introduction
The use of artificial intelligence in medicine aims to unlock hidden, useful information from data and support clinical decisions. AI can assist in diagnosis and treatment selection, assess risks, classify diseases, reduce medical errors, and improve productivity.
Possible data sources include demographic information, notes from healthcare providers, medical images, laboratory results, genetic tests, and records from medical or wearable devices such as smartwatches. The easy availability of these data in electronic health records and intelligent devices equipped with sensors, network connectivity, and cloud storage opens up possibilities for the management of medical information — from the patient, physician, and hospital to health policy and decision-makers.

4.2 Examples of AI Applications in Clinical Practice
4.2.1 As a Screening Tool
- Analysis of radiographic images (projection or CT images), estimation of the probability of disease, and marking of findings for radiologists to interpret. An important example is AI screening of chest X-rays for COVID-19 (Minaee, Kafieh, Sonka, Yazdani, & Jamalipour Soufi, 2020; Shi et al., 2020).
- Analysis of fundus images for detecting vision-threatening findings requiring referral to ophthalmology. The FDA-approved IDx-DR system examines retinal images and identifies individuals requiring referral and vision-threatening diabetic retinopathy (Van Der Heijden et al., 2018).
- In the United Kingdom, an AI-based chat application was deployed that distinguishes between individuals who only need reassurance and those who need to be examined by a physician. The aim is to relieve the healthcare system and direct resources to those with genuine need (W. Wang & Siau, 2018).
- Skin tumors such as melanomas can be diagnosed with high, expert-level accuracy and distinguished from nevi (Esteva et al., 2017).
4.2.2 As a Prognosis Assessment Tool
- Estimation of survival time after treatment of uveal melanoma (Damato, Eleuteri, Fisher, Coupland, & Taktak, 2008).
- Estimation of survival time and recurrence rate in breast tumors (Cirkovic, Cvetkovic, Ninkovic, & Filipovic, 2015).
4.2.3 As Therapy Support
- Support in planning radiotherapy to minimize the exposure of healthy tissue (C. Wang, Zhu, Hong, & Zheng, 2019).
- Support in selecting optimal therapeutic strategies for sepsis in the intensive care unit. When no clear protocol exists, reinforcement learning can select an appropriate approach (Komorowski, Celi, Badawi, Gordon, & Faisal, 2018).
4.2.4 As a Replacement for a Healthcare Provider
It is unlikely that AI will fully replace physicians in the foreseeable future. However, it can perform certain tasks more consistently, faster, and more reproducibly than humans, such as bone age estimation from radiographs (Tajmir et al., 2019), diagnosis of certain retinal diseases in OCT images (De Fauw et al., 2018), or quantification of vascular stenoses and other measurements in cardiac imaging (Slomka et al., 2017). These tasks are not necessarily complex but can consume time that the healthcare provider could devote to more demanding tasks.
4.2.5 As Healthcare Provider Support
Several studies show that the synergy between AI and the healthcare provider delivers better results than either alone. It improves the ability to support clinical decisions in real time and thereby enhance efforts toward precision healthcare (Sitapati et al., 2017).
4.3 Examples of AI Applications in Ophthalmology

Below, important studies on ophthalmic conditions in which AI has been used are presented. Figure 38 organizes these by the number of studies conducted.

4.3.1 Diabetic Retinopathy
IDx-DR is one of the most important practical applications. The system was commercially introduced with FDA approval for screening for diabetic retinopathy in individuals over 21 years of age in primary care settings, without an ophthalmologist being present. The examination is non-invasive and does not require pupillary dilation. In practice, the system achieved an accuracy of approximately 90% in detecting referral-worthy cases (Van Der Heijden et al., 2018).
The system was trained with 128,175 images and tested with 9,963 images. The training images were evaluated by 54 US board-certified or fourth-year ophthalmology residents, with a mean of 3–7 physicians per image. For the test group, the seven physicians with the highest agreement were selected. Artificial neural networks and the Inception-v3 network pre-trained on ImageNet with batch normalization for accelerated learning were used.

4.3.2 Glaucoma
- Li et al. examined the detection of glaucomatous optic neuropathy. The AI model was trained with a fundus database of 48,116 images, including 8,000 test images. Accuracy reached up to 98.6%. The most common causes of false-negative results were pathological and high myopia; false-positive results were associated with myopia and enlarged physiological cupping (Li et al., 2018).
- Muhammad et al. examined the suspicion of open-angle glaucoma using swept-source OCT with multiple maps. The RNFL map achieved 93.1% accuracy and was more accurate than conventional OCT and conventional visual field examination (Muhammad et al., 2017).
- Asaoka et al. distinguished preperimetric from normal visual fields. The database comprised 171 visual field images; mean deviation, pattern deviation, and standard deviation served as inputs of a feedforward neural network. Accuracy in detecting preperimetric visual fields was 92.6% (Asaoka, Murata, Iwase, & Araie, 2016).
4.3.3 Age-Related Macular Degeneration
Studies included diagnosis based on OCT images (Treder, Lauermann, Eter, & Ophthalmology, 2018), detection of active neovascularization (Chakravarthy et al., 2016), prediction of future need for repeated injections (Bogunović et al., 2017), and assessment of the current need for antiangiogenic injections (Prahs et al., 2018).
4.3.4 Cataract
Gao developed a neural network for diagnosis and grading of senile cataract with high accuracy (Gao, Lin, & Wong, 2015). An AI screening for congenital cataracts could be particularly significant due to the prevention of avoidable blindness through early diagnosis (Liu et al., 2017).
AI also plays a role in the new generation of formulas for calculating intraocular lens power, such as Hill-RBF, Kane, and PEARL-DGS (Cheng et al., 2020). Hill-RBF is the most well-known formula and is available online. Inputs include the axial length, anterior chamber depth, corneal curvature values, and axis. To improve accuracy, central corneal thickness, lens thickness, and white-to-white distance can be added. Accuracy within 0.5 dpt is 71.2% and comparable to or better than third- and fourth-generation formulas (Darcy et al., 2020).
4.3.5 Various Applications
Poplin et al. trained an AI network that could predict age, sex, smoking status, systemic blood pressure, and a history of cardiac disease from fundus images (Poplin et al., 2018). Zhou et al. developed an AI system that can detect retinal changes associated with Alzheimer's disease that are not identifiable by humans; a patent was filed for this (Zhou, Sinai, Moore, & Wong, 2006).
This is only a small selection of AI applications in ophthalmology. The most important applications in keratoconus diagnostics are presented separately below.
4.4 AI Applications for the Detection of Keratoconus
Below, important studies on AI applications in keratoconus are described. Tables 4 and 5 summarize the algorithms, samples, results, and advantages and disadvantages of the studies.
4.4.1 Keratoconus Diagnostics with Biomechanical Properties and Regression Algorithms
Corvis ST records the response of the cornea to a defined air puff with a high-resolution Scheimpflug camera. It captures 4,300 images per second and enables precise measurement of corneal thickness and intraocular pressure as well as biomechanical properties (Roberts, 2016).
Two important indices for detecting early keratoconus stages are the Corvis Biomechanical Index (CBI), developed exclusively from Corvis data using regression, and the Tomography and Biomechanical Index (TBI), which combines regression and random forest and integrates Corvis with Pentacam data (Figure 40; Ambrósio et al., 2017; Renato, 2016; Riccardo, 2016).
The TBI is more accurate than the CBI (98.5% versus 88.2%). However, both indices can produce false-negative and false-positive results. At values within the thresholds of 0.5 and 0.29, respectively, the development of ectasia after refractive surgery cannot be excluded with 100% certainty (Fernández, Rodríguez-Vallejo, & Piñero, 2019).

4.4.2 Keratoconus Diagnostics with Support Vector Machine
The support vector machine is a supervised learning algorithm. It classifies training points in an n-dimensional space by a separating hyperplane with n-1 dimensions such that the distance between the nearest points of different classes is maximized. Two point groups in a two-dimensional space are separated, for example, by a straight line (Cortes & Vapnik, 1995).
SIRIUS uses this method to assign images to four groups based on the indices SIf, SIb, RBFf, BCVf, BCVb, RMS(HOA), and THKmin: keratoconus-compatible, suspect, normal, and abnormal. Classification accuracy exceeded 97% (Arbelaez et al., 2012). A study compared 25 algorithms for classifying corneal images from an AS-OCT device (CASIA SS-1000, Tomey). After selection of the eight most discriminating indices, the SVM achieved the best accuracy of 93.6% (Lavric, Popa, Takahashi, & Yousefi, 2020).
4.4.3 Keratoconus Diagnostics with Convolutional Neural Networks
KeratoDetect was developed at the Ștefan cel Mare University in Romania. The model was trained with 3,000 artificially generated corneal images, not with images from real patients. This is a weakness; additionally, only curvature maps were used. Accuracy in classifying normal versus keratoconic was 99.33% (Lavric & Valentin, 2019).
A study from Kitasato University in Japan used six maps captured with CASIA AS-OCT SS-1000 (Tomey): anterior and posterior curvature, anterior and posterior elevation, total refractive power, and thickness. Six ResNet-18-based models were trained with 304 keratoconic and 239 normal images. Mean overall accuracy was 99.1%. The posterior elevation map was most accurate at 99.3%, followed by the posterior curvature map at 99.1% (Kamiya et al., 2019).
At National Taiwan University, 354 images from 206 patients were examined using a TMS-4 videokeratoscope (Tomey). Three pre-trained models (VGG16, InceptionV3, and ResNet152) were used. The images were divided into normal, keratoconic, and subclinical groups; the first two were used for training, the third for testing. ResNet152 achieved 95.8%, the other two 93.1%. Prediction of subclinical cases was unsatisfactory at 28.5% with a probability threshold of 50%. Pixel-wise discriminating features and class-specific heatmaps were created to understand the network's operation and draw the examiner's attention to conspicuous image regions (Figure 41; Kuo et al., 2020).

A study from Assiut University (Egypt) trained with 2,574 and tested with 644 Pentacam images. Anterior and posterior elevation maps, anterior sagittal curvature map, thickness map, and a combined image of these four maps were used. The model consisted of two convolutional layers and a four-layer neural network before the output layer. The combined four-map image achieved 98.9%, followed by the posterior elevation map at 97.7%; heatmaps were also examined (Abdelmotaal et al., 2020).
Table 4 – Comparison of earlier studies (excluding CNN):
| Study/Algorithm | Method and Sample | Results | Advantages | Disadvantages |
|---|---|---|---|---|
| CBI (Vinciguerra et al., 2016), Regression algorithms | Retrospective distinction of normal (478) and keratoconic (180) corneas using the CORVIS device | Sensitivity 94.3%, specificity 97.5%, AUC 97.7%, accuracy 88.2% | CORVIS is a unique device | Retrospective; suspect corneas not examined; false-positive and false-negative findings in some studies |
| TBI (Ambrósio et al., 2017), Regression + Random Forest | Retrospective distinction of normal (480), keratoconic (204), unilateral keratoconic (72), and forme fruste keratoconic (72) corneas | Threshold 0.29 for detecting FFKC; AUC 98.5%, sensitivity 90.4%, specificity 96% | Integrates the capabilities of CORVIS and PENTACAM | Retrospective; false-positive and false-negative findings in some studies |
| SIRIUS-SVM (Arbelaez et al., 2012), Support Vector Machine | Retrospective classification of keratoconic (877), normal (1,259), subclinical (426), and post-refractive surgery (940) corneas | KC: accuracy 99.3%, sensitivity 98.2%, specificity 95%; subclinical: accuracy 97.3%, sensitivity 92%, specificity 97.7% | Training based on anterior and posterior corneal surface features; SIRIUS (Placido + Scheimpflug) | Retrospective; relatively few features used to avoid overfitting — a limitation of the SVM |
Table 5 – Comparison of earlier CNN studies:
| Study/Model | Method and Sample | Results | Advantages | Disadvantages |
|---|---|---|---|---|
| KeratoDetect (Lavric & Valentin, 2019), CNN | Artificially generated images: 1,500 normal and 1,500 KC; the maps used are not clearly described and were presumably limited to curvature maps | Accuracy 99.33% | — | No images from real patients; map types unclear, presumably only curvature maps |
| Kamiya et al. (2019), CNN/ResNet-18 | Retrospective with AS-OCT, six maps; 239 normal and 304 KC | Overall accuracy 99.1%; posterior elevation map 99.3%, followed by posterior curvature 99.1% | AS-OCT; six maps; pre-trained ResNet-18 | Suspect cases not examined; use of a pre-trained model |
| Kuo et al. (2020), CNN (VGG16, InceptionV3, ResNet152) | Retrospective with videokeratoscope; training: 170 KC and 156 normal, testing: 28 subclinical | ResNet152: 95.8% in training; 28.5% in subclinical test cases | Pre-trained models; discriminating pixel feature and class-specific heatmaps for model understanding | Videokeratoscope does not provide accurate information on the posterior corneal surface; only anterior curvature map; subclinical cases were inadequately detected |
| Abdelmotaal et al. (2020), CNN | Retrospective with PENTACAM: 1,038 KC, 1,108 normal, 1,072 subclinical or forme fruste; four individual maps and a combined image | Combined four-map image 98.9%, posterior elevation map 97.7% | Four maps individually and combined; PENTACAM; heatmaps to examine model operation | No independent test group; no true suspect cases, only subclinical/forme fruste cases |
| Present study | Retrospective collection/training and independent test group with real images; training: 559 KC, 1,217 normal, 167 suspect; testing: 13 KC, 366 normal, 43 suspect | Best networks: anterior elevation map, followed by posterior elevation map, anterior refractive power, equivalent refractive power; accuracy/F1-score 94.5–94.3% | Novel network architecture; transfer learning; data augmentation; training with real images; large sample; independent test group; SIRIUS; heatmaps; reading with support | Training, validation, and test groups unbalanced; AS-OCT considered more accurate than SIRIUS; retrospective design |