Vision Transformer for Cataract Detection in Anterior Eye Images: Achieving Optimal Performance with Limited Data
DOI:
https://doi.org/10.21512/commit.v20i2.13529Abstract
Convolutional Neural Networks (CNNs) continue to face several challenges in cataract classification using image data, particularly due to limitations in dataset size and variability in color, shape, and position. These constraints arise because CNNs primarily process local features within each layer. The research uses Vision Transformer (ViT) to capture global spatial relationships in front-eye images. ViT can detect cataracts more accurately than CNN-based owing to its capability of modeling long-range dependencies. The research aims to evaluate ViT’s performance on four publicly available cataract datasets using training time, accuracy, precision, recall, and F1-score. To assess the model’s reliability when working with a small number of training datasets, ViT’s performance is also be compared with ResNet-50 and EfficientNet-B7. The results indicate that ViT outperforms ResNet-50 in accuracy by 10%–31% and exceeds EfficientNet-B7 by 30%–41%. However, ViT requires approximately 16 seconds longer to train than ResNet-50 and EfficientNet-B7 due to its deeper architecture. Although ViT’s accuracy is 2.47% lower than previous studies using hybrid deep learning approaches, its parameter structure is simpler. Overall, the findings indicate that, compared with CNN-based models, ViT performs more effectively when trained on smaller datasets. Future research may incorporate segmentation techniques to further validate cataract detection in anterior eye images.
References
[1] L. Khan, N. Shaheen, Q. Hanif, S. Fahad, and M. Usman, “Genetics of congenital cataract, its diagnosis and therapeutics,” Egyptian Journal of Basic and Applied Sciences, vol. 5, no. 4, pp. 252–257, 2018.
[2] J. Wang et al., “A transformer-based knowledge distillation network for cortical cataract grading,” IEEE Transactions on Medical Imaging, vol. 43, no. 3, pp. 1089–1101, 2023.
[3] R. Azad et al., “Advances in medical image analysis with vision transformers: A comprehensive review,” Medical Image Analysis, vol. 91, 2024.
[4] A. F. Adinegoro, G. N. Sutapa, A. A. N. Gunawan, N. K. N. Anggarani, P. Suardana, and I. G. A. Kasmawan, “Classification and segmentation of brain tumor using EfficientNet-B7 and u-net,” Asian Journal of Research in Computer Science, vol. 15, no. 1, pp. 1–9, 2023.
[5] J. Wang et al., “Prediction of postoperative visual acuity in patients with age-related cataracts using macular optical coherence tomography-based deep learning method,” Frontiers in Medicine, vol. 10, pp. 1–10, 2023.
[6] I. Santoso, A. M. Manurung, and E. R. Subhiyakto, “Comparison of ResNet-50, EfficientNet-B1, and VGG-16 algorithms for cataract eye image classification,” Journal of Applied Informatics and Computing, vol. 9, no. 2, pp. 284–294, 2025.
[7] A. K. Khanra, M. Kumar, A. Mandal, and S. Chatterjee, “EfficientNetB0 vs. VGG16 vs. ResNet-50: Classification of various skin diseases using deep learning,” in 2025 International Conference on Next Generation of Green Information and Emerging Technologies (GIET). Gunupur, India: IEEE, Aug. 8–9, 2025, pp. 1–5.
[8] B. V. K. Golusu, N. Rajana, R. R. Arji, N. Killamsetty, and P. Naveena, “Revolutionizing image duplicate detection: Harnessing ResNet-50 with spatial transformers for unprecedented precision,” International Journal of Novel Research and Development, vol. 9, no. 4, pp. 521–530, 2024.
[9] S. B. Thakare, “Knowledge distillation of the ResNet50 model for ocular diseases analysis,” Master’s thesis, National College of Ireland, 2022.
[10] S. S. Mahmood, S. Chaabouni, and A. Fakhfakh, “Improving automated detection of cataract disease through transfer learning using ResNet50,” Engineering, Technology & Applied Science Research, vol. 14, no. 5, pp. 17 541–17 547, 2024.
[11] H. Imaduddin, I. C. Utomo, and D. A. Anggoro, “Fine-tuning ResNet-50 for the classification of visual impairments from retinal fundus images,” International Journal of Electrical & Computer Engineering, vol. 14, no. 4, pp. 4175–4182, 2024.
[12] R. Comle, “Fusion of vision transformer, Inception-V3 and ResNet50 for efficient eye disease detection,” International Journal of Engineering & Management Informatics, vol. 2, no. 2, pp. 10–22, 2026.
[13] M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning. Long Beach, California, USA: PMLR, June 9–15, 2019, pp. 6105–6114.
[14] A. A. Rakib, M. M. Billah, A. S. Ahamed, H. M. Imamul, and M. S. A. Masum, “EfficientNetbased model for automated classification of retinal diseases using fundus images,” European Journal of Computer Science and Information Technology, vol. 12, no. 8, pp. 48–61, 2024.
[15] J. H. L. Goh et al., “Artificial intelligence for cataract detection and management,” Asia-Pacific Journal of Ophthalmology, vol. 9, no. 2, pp. 88–95, 2020.
[16] Y. C. Tham et al., “Detecting visually significant cataract using retinal photograph-based deep learning,” Nature Aging, vol. 2, no. 3, pp. 264–271, 2022.
[17] A. C. I. Ardison, M. J. R. Hutagalung, R. Chernando, and T. W. Cenggoro, “Observing pretrained Convolutional Neural Network (CNN) layers as feature extractor for detecting bias in image classification data,” CommIT (Communication and Information Technology) Journal, vol. 16, no. 2, pp. 149–158, 2022.
[18] D. Philippi, K. Rothaus, and M. Castelli, “A vision transformer architecture for the automated segmentation of retinal lesions in spectral domain optical coherence tomography images,” Scientific Reports, vol. 13, no. 1, pp. 1–14, 2023.
[19] B. Hassan, T. Hassan, R. Ahmed, N. Werghi, and J. Dias, “SIPFormer: Segmentation of multiocular biometric traits with transformers,” IEEE Transactions on Instrumentation and Measurement, vol. 72, pp. 1–14, 2023.
[20] J. W. Kusno and A. Chowanda, “Modeling emotion recognition system from facial images using convolutional neural networks,” CommIT (Communication and Information Technology) Journal, vol. 18, no. 2, pp. 251–259, 2024.
[21] B. Zhang et al., “SegViT: Semantic segmentation with plain vision transformers,” Advances in Neural Information Processing Systems, vol. 35, pp. 4971–4982, 2022.
[22] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” in ICLR 2021: The Ninth International Conference on Learning Representations, Virtual, May 3–7, 2021.
[23] Q. Liu, C. Kaul, J. Wang, C. Anagnostopoulos, R. Murray-Smith, and F. Deligianni, “Optimizing vision transformers for medical image segmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, June 4–10, 2023, pp. 1–5.
[24] D. Kumar, B. Bakariya, C. Verma, and Z. Illes, “Cataract disease identification using transformer and convolution neural network: A novel framework,” in 2023 3rd International Conference on Technological Advancements in Computational Sciences (ICTACS). Tashkent, Uzbekistan: IEEE, Nov. 1–3, 2023, pp. 1230–1235.
[25] T. Zhang, W. Xu, B. Luo, and G. Wang, “Depthwise convolutions in vision transformers for efficient training on small datasets,” Neurocomputing, vol. 617, pp. 1–11, 2025.
[26] Z. Chen, L. Xie, J. Niu, X. Liu, L. Wei, and Q. Tian, “Visformer: The vision-friendly transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision. Virtual: Computer Vision Foundation, Oct. 11–17, 2021, pp. 589–598.
[27] X. Mao et al., “Towards robust vision transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, Louisiana: Computer Vision Foundation, 2022, pp. 12 042–12 051.
[28] S. M¨uller et al., “Artificial intelligence in cataract surgery: A systematic review,” Translational Vision Science & Technology, vol. 13, no. 4, 2024.
[29] I. Sonata, Y. Heryadi, A. Wibowo, and W. Budiharto, “End-to-end steering angle prediction for autonomous car using vision transformer,” CommIT (Communication and Information Technology) Journal, vol. 17, no. 2, pp. 221–234, 2023.
[30] D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations 2015, San Diego, California, United States, 2015.
[31] P. B. Singh et al., “Glaucoma classification using light vision transformer,” EAI Endorsed Trans Pervasive Health and Technology, vol. 9, pp. 1–17, 2023.
[32] J. Vuppalapati, T. R. Kotha, K. P. Battula, and P. Kavali, “Eye disease classification using vision transformer: A deep learning approach,” 2024. [Online]. Available: http://bit.ly/4xcWG7N
[33] Roboflow, “Cataract v01 computer vision dataset,” 2022. [Online]. Available: https://universe.roboflow.com/cataract/cataract-v01/dataset/7
[34] A. Ramakrishnan, “Cataract classification dataset,” 2024. [Online]. Available: https://www.kaggle.com/datasets/akshayramakrishnan28/ cataract-classification-dataset
[35] A. Abdulkhaliq, “Cataract,” 2023. [Online]. Available: https://www.kaggle.com/datasets/kershrita/cataract
[36] N. Padia, A. Hirpara, D. Jani, and S. Patel, “Cataract dataset,” 2023. [Online]. Available: https://www.kaggle.com/datasets/nandanp6/cataract-image-dataset/data
[37] S. Park et al., “Features of long-standing Korean type 2 diabetes mellitus patients with diabetic retinopathy: A study based on standardized clinical data,” Diabetes & Metabolism Journal, vol. 41, no. 5, 2017.
[38] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, NV, USA: IEEE, June 27–30, 2016, pp. 770–778.
[39] S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning. Lille, France: PMLR, July 7–9, 2015, pp. 448–456.
[40] A. K. Annamraju, “Dataset pre-processing and artificial augmentation, network architecture and training parameters used in appropriate training of convolutional neural networks for classification based computer vision applications: A survey,” International Journal of Advanced Engineering, Management and Science, vol. 2, no. 9, 2016.
[41] A. Nazarkar, H. Kuchulakanti, C. S. Paidimarry, and S. Kulkarni, “Impact of various data splitting ratios on the performance of machine learning models in the classification of lung cancer,” in Proceedings of the Second International Conference on Emerging Trends in Engineering (ICETE 2023), vol. 223. Hyderabad, India: Atlantis Press, April 28–30, 2023, pp. 96–104.
[42] R. Strudel, R. Garcia, I. Laptev, and C. Schmid, “Segmenter: Transformer for semantic segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision. Virtual: Computer Vision Foundation, Oct. 11–17, 2021, pp. 7262–7272.
[43] E. Simon, “Fine-tuning a vision transformer model for image classification into IAB taxonomy categories,” 2023. [Online]. Available: https://www.blend360.com/thought-leadership/fine-tuning-a-vision-transformer
[44] C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to fine-tune BERT for text classification?” in China National Conference on Chinese Computational Linguistics. Kunming, China: Springer, Oct. 18–20, 2019, pp. 194–206.
[45] R. M. Zur, Y. Jiang, L. L. Pesce, and K. Drukker, “Noise injection for training artificial neural networks: A comparison with weight decay and early stopping,” Medical Physics, vol. 36, no. 10, pp. 4810–4818, 2009.
[46] J. Zhang, F. Li, X. Zhang, H. Wang, and X. Hei, “Automatic medical image segmentation with vision transformer,” Applied Sciences, vol. 14, no. 7, pp. 1–18, 2024.
[47] I. Markoulidakis and G. Markoulidakis, “Probabilistic confusion matrix: A novel method for machine learning algorithm generalized performance analysis,” Technologies, vol. 12, no. 7, pp. 1–23, 2024.
[48] O. Gorokhovatskyi and O. Peredrii, “Image pair comparison for near-duplicates detection,” International Journal of Computing, vol. 22, no. 1, pp. 51–57, 2023.
[49] J. Olaniyan, D. Olaniyan, I. C. Obagbuwa, B. M. Esiefarienrhe, and M. Odighi, “Transformative transparent hybrid deep learning framework for accurate cataract detection,” Applied Sciences, vol. 14, no. 21, pp. 1–20, 2024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nina Sevani, Westlee Matthew Agustinus, Edy Kristianto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Â
USER RIGHTS
All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows: Creative Commons Attribution-Share Alike (CC BY-SA)

















