Abstract

Diseases of the retina continue to be a leading cause of blindness and visual impairment around the world. In the field of medical image analysis, specifically retinal disease identification, deep learning techniques, such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), have showed remarkable potential. In this paper, we present a unique method for detecting retinal diseases by combining the advantages of the Inception-V3, ResNet-50, and Vision Transformer architectures into a single model called a Cascade CNN-ViT. The suggested Cascade CNN-ViT model extracts local features from retinal pictures by leveraging the spatial hierarchy learning capabilities of Inception-V3 and ResNet-50. The Vision Transformer takes these regional characteristics and uses self-attention mechanisms to pick up global context information and long-range interdependence. The model successfully combines fine-grained local information with semantically significant global contextual cues by merging the output representations from the CNNs and Vision Transformer. undertaking comprehensive experiments on a large and varied dataset of multimodal retinal pictures to evaluate the performance of the proposed technique. Cascade CNN-ViT model outperforms standalone CNNs and Vision Transformers, as shown by the experimental findings. The model is also resilient across all classes of retinal diseases and is able to successfully deal with the complications introduced by using multiple picture types. Overall, the power of cascading Inception-V3, ResNet-50, and Vision Transformer topologies for improved retinal illness diagnosis has been demonstrated. Potentially improving the management of retinal illnesses and preserving visual health, the proposed approach could have important consequences for early detection and timely intervention.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call