Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Enhancing Digital Art Style Recognition via a Hybrid Vision Transformer and Lightweight CNN with Attention Mechanisms

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas.

Similar Papers
  • Research Article
  • 10.1504/ijica.2025.149778
Enhancing digital art style recognition via a hybrid vision transformer and lightweight CNN with attention mechanisms
  • Jan 1, 2025
  • International Journal of Innovative Computing and Applications
  • Ying Zhang + 1 more

Current methods for art style recognition often struggle to capture local details and balance global and texture features, leading to vague style representation during multi-scale fusion. To address limitations in capturing local details and balancing global and texture features in digital art style classification, we propose a hybrid model based on vision transformer and lightweight CNN. The model adopts a multi-scale attention-weighted fusion strategy and multi-task learning to optimise classification and style reconstruction simultaneously. Experimental results show the proposed model achieves a Kappa coefficient of 0.96, significantly outperforming baselines (0.78 and 0.74), with a classification accuracy of 96.3% and an F1 score of 97.8%. These findings demonstrate the model's strong performance in digital art style recognition, promoting intelligent applications in cultural and creative industries. The proposed model significantly improves accuracy and feature representation in digital art style recognition, supporting the intelligent development of the cultural and creative industry and advancing deep learning applications in art analysis.

  • Research Article
  • 10.3390/buildings15020176
Innovative Framework for Historical Architectural Recognition in China: Integrating Swin Transformer and Global Channel–Spatial Attention Mechanism
  • Jan 9, 2025
  • Buildings
  • Jiade Wu + 3 more

The digital recognition and preservation of historical architectural heritage has become a critical challenge in cultural inheritance and sustainable urban development. While deep learning methods show promise in architectural classification, existing models often struggle to achieve ideal results due to the complexity and uniqueness of historical buildings, particularly the limited data availability in remote areas. Focusing on the study of Chinese historical architecture, this research proposes an innovative architectural recognition framework that integrates the Swin Transformer backbone with a custom-designed Global Channel and Spatial Attention (GCSA) mechanism, thereby substantially enhancing the model’s capability to extract architectural details and comprehend global contextual information. Through extensive experiments on a constructed historical building dataset, our model achieves an outstanding performance of over 97.8% in key metrics including accuracy, precision, recall, and F1 score (harmonic mean of the precision and recall), surpassing traditional CNN (convolutional neural network) architectures and contemporary deep learning models. To gain deeper insights into the model’s decision-making process, we employed comprehensive interpretability methods including t-SNE (t-distributed Stochastic Neighbor Embedding), Grad-CAM (gradient-weighted class activation mapping), and multi-layer feature map analysis, revealing the model’s systematic feature extraction process from structural elements to material textures. This study offers substantial technical support for the digital modeling and recognition of architectural heritage in historical buildings, establishing a foundation for heritage damage assessment. It contributes to the formulation of precise restoration strategies and provides a scientific basis for governments and cultural heritage institutions to develop region-specific policies for conservation efforts.

  • Research Article
  • Cite Count Icon 34
  • 10.1016/j.compbiomed.2023.106606
One-stage and lightweight CNN detection approach with attention: Application to WBC detection of microscopic images
  • Jan 23, 2023
  • Computers in Biology and Medicine
  • Zhenggong Han + 7 more

One-stage and lightweight CNN detection approach with attention: Application to WBC detection of microscopic images

  • Conference Article
  • 10.1117/12.2640778
Toward more efficient iris recognition using a lightweight CNN framework with attention mechanism
  • Oct 3, 2022
  • Qinhong Zou + 4 more

Iris recognition is considered as one of the most promising biometrics due to its discriminative features and friendly acquisition methods. Herein, a deep learning-based method is proposed to achieve more accurate and efficient iris recognition. The proposed framework Iris Attention Network (IrisAttenNet) integrates the attention mechanism into a lightweight CNN to extract iris features more specifically. In the process of feature learning, the channel features with more information that contribute to the recognition result will attract more attention and be given higher weights, which is similar to the human visual perception mechanism. The performance of the proposed framework is evaluated by four publicly available datasets representing different intra-class variations: CASIA_Iris_V4 Interval, Lamp, Thousand and UBIRIS.v1. The experimental results have demonstrated that the approach based on the IrisAttenNet shows higher accuracy, stronger generalization and less computational cost. The intermediate outcomes heat maps have proved that the key contribution of the attention module through visualization of the feature areas of images.

  • Research Article
  • 10.14569/ijacsa.2026.0170142
Attention-Guided Lightweight MobileNetV2 for Real-Time Driver Drowsiness Classification on Edge-IoT Systems
  • Jan 1, 2026
  • International Journal of Advanced Computer Science and Applications
  • Yo Ceng Giap + 6 more

Driver drowsiness is a major cause of traffic accidents, so Edge-IoT platforms with limited resources need to be able to accurately and quickly detect when drivers are drowsy. This study examines attention-guided lightweight CNN design predicated on MobileNetV2 for real-time driver drowsiness detection. The authors compare a SE-enhanced MobileNetV2 to the baseline model and a structurally optimized version that uses Depthwise Separable Convolution (DSC), Bottleneck blocks, and Expansion layers. Experiments on 500 images demonstrate that channel attention enhances feature discrimination, whereas structural optimization yields the most resilient trade-off between accuracy and latency. Statistical validation employing 95% confidence intervals and two-proportion Z-tests substantiates the significance of these enhancements. The proposed models support real-time inference despite their small size (about 2.6 million parameters and 315 million FLOPs). These findings suggest structural optimization is more important than attention mechanisms in designing lightweight CNNs for embedded driver monitoring.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/hpcc-dss-smartcity-dependsys53884.2021.00145
A Lightweight VK-Net Based on Motion Profiles for Hazardous Driving Scenario Identification
  • Dec 1, 2021
  • Zhen Gao + 6 more

Hazardous driving scenarios are critical for the development and testing of highly autonomous vehicles. Recent studies have proposed new approaches for extracting hazardous scenarios from driving videos, which greatly improve the detection accuracy compared with traditional approaches that extract information from structured kinematic data. However, direct processing on driving videos is costly, and the videos need to be uploaded and processed in a centralized manner. In this study, we propose a novel approach for reducing the cost of video analysis while maintaining high accuracy in hazardous scenario identification. Firstly, the videos are significantly compressed into motion profile images that contain motion trajectories of objects. Secondly, a lightweight CNN with attention mechanism has been designed to extract risky trajectory features of the conflicting traffic objects. Furthermore, data augmentation has been applied to increase the generalizability of the model. The proposed approach achieves significantly lightweight network size, and meanwhile, experimental results have demonstrated that the approach achieves high accuracy with sensitivity 0.95, specificity 0.85 and AUC 0.95, which outperforms state-of-the-art hazardous scenario identification approaches. The efficient execution of identification implies that it is suitable to be distributed and deployed at each vehicle, and consequently, videos could be processed via edge computing, which would improve the diversity of hazardous scenarios and the efficiency of scenario collection for further research work.

  • Preprint Article
  • 10.21203/rs.3.rs-4536797/v1
BDR-Net: Digital Recognition Network for Billet Surface Based on Flow Alignment and Attention Mechanism
  • Jun 20, 2024
  • Research Square
  • Jinyu Xu + 2 more

The surface characteristics of billets are crucial for subsequent traceability, yet the production process generates intricate digital features on their surfaces. This paper introduces BDR-Net, a novel billet surface digit recognition network. Drawing inspiration from Inception, the network adopts a ResNext-like architecture as its primary framework. It uniformly distributes output in dimension, extracts positional and scale features separately, and introduces a mixed dilated convolution block to reduce parameters while expanding the sensory field. To address the challenge of lost up-sampled features during fusion, an innovative stream alignment-based up-sampled feature fusion algorithm is proposed. Additionally, to enhance the network's focus on extracting salient spatial and channel features, a mixed-dimensional attention mechanism (scSE) is integrated into the alignment-based upsampling feature fusion module. Experimental results showcase BDR-Net's outstanding performance, achieving an impressive 95.6\% accuracy in digitally classifying billet surfaces, surpassing the ResNext50\_32x4d benchmark model by 4.3\% in recognition accuracy. Moreover, compared to current classification networks, this model exhibits significant accuracy improvements. Furthermore, the mAP@0.95 metric reaches 0.897, surpassing current classification networks. These findings underscore the remarkable performance of the model in billet surface digit recognition, offering an effective solution for digit recognition on billet surfaces in steel mills.

  • Research Article
  • Cite Count Icon 54
  • 10.1109/lgrs.2020.3031593
Convolutional Neural Network With Attention Mechanism for SAR Automatic Target Recognition
  • Nov 2, 2020
  • IEEE Geoscience and Remote Sensing Letters
  • Ming Zhang + 5 more

Synthetic aperture radar automatic target recognition (SAR ATR) is a key technique of remote-sensing image recognition, which has many potential applications in the fields of military surveillance, national defense, civil application, and so on. With the development of science and technology, deep convolutional neural network (DCNN) has been widely applied for SAR ATR. However, it is difficult to use deep learning to train models with limited ray SAR images. To resolve this problem, we proposed an effectively lightweight attention mechanism CNN (AM-CNN) model for SAR ATR. Extensive experimental results on the Moving and Stationary Target Acquisition and Recognition (MSTAR) data set illustrate that the AM-CNN model can achieve a superior recognition performance, and the average recognition accuracy can reach 99.35% on the classification of 10 class targets. Compared with the traditional CNN and the state-of-the-art method, our model is significantly superior to improve performance and efficiency.

  • Conference Article
  • 10.1109/incacct65424.2025.11011390
Towards Smart Waste Sorting: Lightweight CNN with Attention Mechanisms for Enhanced Classification
  • Apr 17, 2025
  • Rashi Chauhan + 3 more

Managing the growing challenges of municipal solid waste is crucial for environmental sustainability, and effective waste management plays a key role in achieving that goal. One promising solution is automating the sorting of refuse, which can greatly improve recycling and overall waste management. However, many current deep learning models that rely on transfer learning suffer from high computational demands and limited adaptability in resource-constrained environments. In this work, we introduce a lightweight convolutional neural network (CNN) specifically designed to overcome these limitations in trash classification. The model incorporates advanced attention mechanism, namely Squeeze-and-Excitation (SE) and Spatial Attention (SA) blocks along with the Mish activation function, collectively augmenting feature representation and optimizing gradient flow. Evaluated on the Kaggle Garbage Dataset, our proposed model achieves a very high classification accuracy of 96.54%, far beyond conventional transfer learning models including VGG16, ResNet50, and EfficientNet B0. With just over 0.8 million parameters and an inference time of only 15 milliseconds per image, the model demonstrates both high computational efficiency and excellent scalability. These results highlight the model’s potential for real-time deployment in resource-limited settings, offering a sustainable and effective alternative for automated waste management. Ultimately, our approach promises to improve garbage classification systems and support more environmentally friendly waste management strategies by leveraging more efficient tools.

  • Research Article
  • Cite Count Icon 10
  • 10.1109/tim.2025.3563050
A Robust Anchor-Free Detection Method for SAR Ship Targets With Lightweight CNN
  • Jan 1, 2025
  • IEEE Transactions on Instrumentation and Measurement
  • Yisheng Hao + 3 more

To address the challenges of compromised detection accuracy caused by near-shore clutter in SAR (synthetic aperture radar) ship detection and the limited deployability of complex algorithms on embedded systems, this paper proposes L-YOLOX, a lightweight SAR target detection algorithm optimized for terminal devices. Firstly, we devise a new feature extraction module based on the MobileNetV3 block to reduce the parameters of traditional YOLOX while strengthening feature representation. Additionally, we incorporate a cross-channel local connection structure to construct an efficient lightweight feature extraction backbone, which is beneficial to improving the network’s ability to fuse SAR ship target information. Next, we develop a multi-scale detection block by using a feature pyramid architecture and dilated convolution to improve the network’s multi-scale detection performance. Finally, we integrate a lightweight convolutional attention mechanism into YOLOX’s Neck structure to enhance the expression of important target detail information and propose the Alpha-AIoU loss function to optimize the gradient propagation process and the network’s weight update. Ablation experimental results on the SSDD dataset show that our network achieves an AP of 90.8%, outperforming Baseline YOLOX with a 70.1% reduction in parameters and a 46.9% decrease in computational cost. Our network also demonstrates a marked enhancement in robustness, validating the effectiveness of our innovations. Some comparative experiments with other state-of-the-art algorithms on SSDD and HRSID further confirm the advantages of our network in terms of SAR image lightweight detection performance and generalization capacity.

  • Research Article
  • Cite Count Icon 57
  • 10.1109/jsen.2023.3244833
Fire Sensor and Surveillance Camera-Based GTCNN for Fire Detection System
  • Apr 1, 2023
  • IEEE Sensors Journal
  • P Sridhar + 3 more

Fire accident is a disaster that can happen anytime anywhere due to accidental causes. In existing works, sensor- and computer vision-based approaches have been used for developing the fire detection model, but they fail to attain the accurate results. The sensor-based methods need more time to detect the fire locations and detection coverage also less. The camera sometimes will consider heavy sunlight as fire and it leads to false positive result, which degrades the accuracy. To overcome the above problems, in this research, a novel optimized Gaussian probability-based threshold convolutional neural network (GTCNN) model has been proposed for detecting the fire accidents using various sensors and surveillance camera-based video (SV). Sensor features map has been calculated from various fire sensors and frames/images from SV are preprocessed using a multiscale retinex algorithm. In addition, the Gaussian threshold (GT) logically integrates with the feature map to increase fire pixel count in low-resolution images. The probability results from sensors and SV camera are optimized by multiobjective mayfly optimization (MOMO) algorithm that normalizes the network parameters, which gives the accurate result. The performance of the proposed optimized GTCNN net is different from the existing deep learning networks in terms of multifeature processing. The result of the proposed work attains the detection accuracy of 98.23%. The proposed optimized GTCNN improves the overall accuracy of 3.25%, 3.79%, and 0.21% better than the channel attention mechanism, lightweight CNN, and you only look once (YOLOv5m), respectively.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/ctisc54888.2022.9849794
A Lightweight CNN for Large-scale Chinese Character Recognition
  • Apr 22, 2022
  • Junwei Zhou + 2 more

The ancient Chinese characters appear in various historical documents and poetry. People tend to use optical character recognition tools to understand these uncommon characters. The current Chinese text recognition interface is restricted to a limited character set, such as GB2312-80 and GB18010-2005 standard. However, the newest HanYu Dictionary contains over 55K characters, much more than the commonly-used character set. This work proposes a compact deep network (HYD-CNet) composed of depthwise separable convolutional blocks and co-ordinate attention mechanism to recognize the ancient Chinese characters. It can achieve efficient retrieval and low-storage need for large-scale character recognition on mobile devices. We build a Chinese character database (HYDDB) using the HanYu Dictionary to evaluate the model performance, containing 55,360 character images. The experiment demonstrates that the proposed HYD-CNet has fewer model parameters at a similar accuracy to mainstream lightweight CNNs.

  • Research Article
  • 10.3389/frai.2026.1824067
A lightweight CNN\u2013transformer hybrid architecture with channel attention for real-time hazardous acoustic event detection
  • Jan 1, 2026
  • Frontiers in Artificial Intelligence
  • Aigerim Altayeva + 1 more

IntroductionHazardous acoustic event detection is critically important for intelligent surveillance, emergency response systems, and public safety monitoring applications. Accurate and real-time identification of dangerous sound events such as explosions, alarms, screaming, and weapon-related sounds can significantly improve situational awareness and accelerate emergency response in safety-critical environments.MethodsThis study proposes a lightweight deep learning architecture for hazardous sound classification based on convolutional feature extraction and channel attention mechanisms. The proposed framework utilizes log-mel spectrogram representations as input and incorporates a TinyCNN backbone enhanced with squeeze-and-excitation channel attention modules to improve discriminative spectral feature learning while preserving computational efficiency. A custom balanced dataset consisting of eight hazardous acoustic classes, including crying, dog barking, emergency alarm, explosion, fire, glass breaking, screaming, and weapon-related sounds, was constructed with one thousand audio samples per class. The model was evaluated using accuracy, precision, recall, and F1-score metrics.ResultsExperimental results demonstrate that the proposed architecture achieves strong multi-class classification performance while maintaining real-time inference capability suitable for edge deployment scenarios. Quantitative evaluations confirm the effectiveness of the lightweight framework for hazardous acoustic event detection. Additional ablation studies indicate that the integration of channel attention mechanisms and spectrogram-based augmentation strategies substantially improves model robustness, feature discrimination, and generalization performance.DiscussionThe obtained findings demonstrate that the proposed lightweight channel-attention-enhanced architecture provides an efficient and reliable solution for real-time hazardous sound detection in intelligent monitoring and public safety systems. The combination of computational efficiency and robust classification performance highlights the suitability of the proposed framework for deployment in resource-constrained and edge-based environments.

  • Research Article
  • 10.71451/istaer2511
Methods for Ground Target Recognition from an Aerial Camera on a Helicopter Using the MISU-YOLOv8 Model in Dark and Foggy Environments
  • Mar 5, 2025
  • International Scientific Technical and Economic Research
  • Houbin Wang + 5 more

Helicopters are critical aerial platforms, and their operational capability in complex environments is crucial. However, their performance in dark and foggy conditions is limited, particularly in ground target recognition using onboard cameras due to poor visibility and lighting conditions. To address this issue, we propose a YOLOv8-based model enhanced to improve ground target recognition in dark and foggy environments. The MS block is a multi-scale feature fusion module that enhances generalization by extracting features at different scales. The improved Residual Mobile Block (iRMB) incorporates attention mechanisms to enhance feature representation. SCINet, a spatial-channel attention-based network, adaptively adjusts feature map weights to improve robustness. UnfogNet, a defogging algorithm, enhances image clarity by removing fog. This integrated approach significantly improves ground target recognition capabilities. Unlike traditional models, AOD-Net generates clean images via a lightweight CNN, making it easily integrable into other deep models. Our MISU-YOLOv8 model outperforms recent state-of-the-art real-time object detectors, including YOLOv7 and YOLOv8, with fewer parameters and FLOPs, improving YOLOv8's Average Precision (AP) from 37% to over 41%. This work can also serve as a plug-and-play module for other YOLO models, this advancement provides robust technical support for helicopter reconnaissance missions in complex environments. **************** ACKNOWLEDGEMENTS**************** Thanks for the data support provided by National-level Innovation Program Project Fund "Research on Seedling Inspection Robot Technology Based on Multi-source Information Fusion and Deep Network" (No.: 202410451009); Jiangsu Provincial Natural Science Research General Project (No.: 20KJB530008); China Society for Smart Engineering "Research on Intelligent Internet of Things Devices and Control Program Algorithms Based on Multi-source Data Analysis" (No.: ZHGC104432); China Engineering Management Association "Comprehensive Application Research on Intelligent Robots and Intelligent Equipment Based on Big Data and Deep Learning" (No.: GMZY2174); Key Project of National Science and Information Technology Department Research Center National Science and Technology Development Research Plan (No.: KXJS71057); Key Project of National Science and Technology Support Program of Ministry of Agriculture (No.: NYF251050).

  • Research Article
  • Cite Count Icon 11
  • 10.1016/j.bspc.2025.108425
ShallowMRI: A novel lightweight CNN with novel attention mechanism for Multi brain tumor classification in MRI images
  • Jan 1, 2026
  • Biomedical Signal Processing and Control
  • Saif Ur Rehman Khan + 5 more

ShallowMRI: A novel lightweight CNN with novel attention mechanism for Multi brain tumor classification in MRI images

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant