Articles published on Faster R-CNN
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
3638 Search results
Sort by Recency
- New
- Research Article
- 10.1080/14942119.2026.2686905
- Jun 24, 2026
- International Journal of Forest Engineering
- Amir Savadkouhi + 2 more
ABSTRACT Forest road surfaces are subject to continuous deterioration from environmental factors, even in the absence of regular traffic, creating significant challenges for safety and maintenance planning. Traditional manual inspections are labor-intensive, costly, and often impractical for extensive networks. This study introduces an Artificial Intelligence (AI)-powered framework, leveraging advanced computer vision techniques, to automate the visual inspection of forest road surfaces using high-resolution UAV imagery. We evaluated three state-of-the-art deep learning models, YOLOv5, YOLOv8, and Faster R-CNN, on a custom dataset of potholes, rutting, and stone protrusions from a mountainous forest road in Iran. The YOLOv5 model demonstrated superior performance, achieving a mean Average Precision (mAP@0.50) of 0.56 and an overall recall of 0.459. It proved highly effective at detecting common distresses like rutting and potholes, while also offering the most robust performance for challenging stone protrusions. These findings validate the integration of AI and computer vision as a transformative tool, providing forest engineers with a scalable, objective, and data-driven methodology to enhance safety, optimize maintenance budgets, and support sustainable forest management.
- New
- Research Article
- 10.1080/10589759.2026.2690158
- Jun 24, 2026
- Nondestructive Testing and Evaluation
- Jianfeng Wei + 3 more
ABSTRACT Wind turbine blade crack detection is a critical measure for ensuring the stable operation of wind turbines. Facing the challenges of complex crack shape, irregular crack characteristics and insufficient detection accuracy, this study proposes a multi-scale feature enhancement detection model for irregular cracks of wind turbine blade based on YOLOv5, named CHBS-YOLO. To capture the local edge and global morphology of irregular blade cracks, a scheme integrating the C3TR module and the Haar wavelet transform is proposed, which enhances the focus on the crack locations and feature information of irregular cracks. On this basis, an embedding scheme integrating the Bidirectional Feature Pyramid Network (BiFPN) and the Shuffle Attention (SA) mechanism is designed to further capture crack features. By fusing multi-scale crack feature information, the proposed scheme enhances the detection model’s focus on crack features. Experimental results demonstrate that on the wind turbine blade crack dataset, the mean Average Precision (mAP) of CHBS-YOLO reaches 81.74%, outperforming classic models such as EfficientDet, SSD, Faster R-CNN, and YOLOv5s by 17.05%, 6.75%, 21.44%, and 1.47%, respectively. These results validate the effectiveness of the detection model, suggesting its applicability to the detection of irregular cracks in wind turbine blades.
- Research Article
- 10.1080/01431161.2026.2684034
- Jun 9, 2026
- International Journal of Remote Sensing
- Shaojie Zhang + 4 more
ABSTRACT Small-target detection in underwater sensing systems is particularly challenging. For the side-scan sonar imagery, the problem is aggravated by low spatial resolution and strong acoustic noise. Speckle noise introduces spurious highlights around targets, and ambient noise can mask weak acoustic backscatter. These factors lead to missed detections and false alarms. The problem becomes more severe after backbone downsampling, and small targets may occupy only a few pixels. Although recent deep-learning approaches have explored super-resolution techniques to enhance image quality, existing methods often fail to achieve a suitable trade-off between detection accuracy and computational efficiency. This paper proposes a target detection model named Impro-YOLO-S for small-target detection in low-resolution imageries. Impro-YOLO-S is composed of a Perceptual and Frequency Loss Enhanced SRGAN and a Re-parameterized YOLO detection network. The SRGAN module includes a global-local receptive field loss enhanced by a Swin-transformer branch and a ResNet branch. Additionally, this module integrates a fast Fourier transform frequency loss to improve frequency domain representation. Specifically, we compute a frequency-domain loss via fast Fourier transform (FFT) and apply inverse FFT to project the spectral discrepancy back to the image space. Impro-YOLO-S is trained end-to-end, and the detection loss is backpropagated to the SR generator to drive task-relevant reconstruction and alleviate SR–detector feature mismatch. The Re-parameterized YOLO detection network employs a structural optimization strategy that mathematically fuses convolution and batch normalization parameters into YOLO detection network. The experimental results show that Impro-YOLO-S achieves an mAP of 74.45%, which represents improvements of 14.37%, 9.93%, 57.87%, and 68.36% compared to YOLOv8, YOLOv5, RetinaNet, and Faster R-CNN, respectively. We also compare with wavelet energy-based methods KWE and SWE. Compared with wavelet energy-based methods, Impro-YOLO-S generates more accurate and stable bounding boxes with correct semantic labels across challenging scenarios. Moreover, the GFLOPS of Impro-YOLO-S reaches 39.9, thus meeting the requirements for real-time search.
- Research Article
- 10.3390/s26113573
- Jun 4, 2026
- Sensors (Basel, Switzerland)
- Ziyi Yang + 5 more
Stainless steel pipes are critical components in industrial systems such as oil and gas transportation and nuclear power cooling. Surface defects can severely degrade their mechanical performance and operational safety. However, existing inspection methods still face challenges including difficult feature extraction, strong reflection interference, and limited accuracy in small-target detection. To address these issues, this paper proposes an improved detection algorithm termed YOLOv8s-BISW (incorporating BiFPN, SGE attention, and WIoU loss), which introduces multidimensional optimizations based on the YOLOv8s baseline. First, an image enhancement module combining Gamma correction and Contrast Limited Adaptive Histogram Equalization (CLAHE) is designed to mitigate uneven illumination and blurred defect imaging. Second, a Bidirectional Feature Pyramid Network (BiFPN) structure is introduced to strengthen multi-scale feature fusion and improve adaptability to defects of different sizes. Meanwhile, a Spatial Group-wise Enhance (SGE) attention module is embedded into the backbone to enhance defect feature representation while suppressing background interference. Furthermore, the Wise Intersection over Union (WIoU) loss function replaces Complete IoU (CIoU) to improve bounding box regression for irregular defects. Experimental results show that the proposed model achieves an mAP of 0.979 on a self-constructed Stainless-steel Tube Flaw (STF) dataset. Compared with the original YOLOv8s, precision, recall, and mAP are improved by 0.007, 0.010, and 0.033, respectively, while the average detection time per image is only 3.7 ms, achieving a favorable balance between accuracy and real-time performance. Compared with mainstream algorithms such as SSD, YOLOv3, and Faster R-CNN, the proposed method demonstrates superior overall performance, providing reliable technical support for automated surface defect detection of stainless steel pipes and offering practical value for intelligent manufacturing quality control.
- Research Article
- 10.3390/diagnostics16111705
- Jun 2, 2026
- Diagnostics
- Chiao-Hua Lee + 8 more
Background/Objectives: To evaluate the effect of detector architecture and dataset characteristics on intracranial hemorrhage (ICH) subtype localization on noncontrast head CT, with emphasis on bidirectional cross-dataset generalization. Methods: This retrospective study analyzed two publicly available datasets: the Brain Hemorrhage Extended (BHX) dataset and the RSNA 2019+ dataset. Models were trained and internally validated on one dataset and externally tested on the other dataset in both directions: BHX-to-RSNA+ and RSNA+-to-BHX. Six representative deep learning detectors, including CNN-based one-stage and two-stage detectors and a Swin Transformer-based RT-DETR (Swin-RT-DETR) variant, were evaluated. Localization performance was assessed using mean average precision at a bounding-box intersection-over-union threshold of 0.5 (mAP@50), bounding-box Dice similarity coefficient (BB-DSC), and bounding-box intersection-over-union (BB-IoU). Image-level and patient-level analyses were performed, with Bonferroni correction applied for statistical comparisons. Dataset characterization analyses were performed to compare subtype prevalence, bounding-box geometry, lesion burden, annotation density, and spatial distribution. Results: Under internal validation, Swin-RT-DETR achieved competitive or superior performance across several ICH subtypes, but its advantage was subtype-dependent rather than uniform. Faster R-CNN with a ResNeXt101 backbone achieved comparable IVH performance and higher IPH BB-DSC and BB-IoU, whereas Swin-RT-DETR performed better for SAH, SDH, and EDH. External validation showed substantial performance degradation across architectures, subtypes, and validation directions. Absolute BB-DSC reductions for Swin-RT-DETR ranged from approximately 0.54–0.79 in the BHX-to-RSNA+ direction and 0.17–0.74 in the RSNA+-to-BHX direction. Similar degradation patterns were observed at the patient level. Statistical comparisons showed fewer significant model-level differences under external validation, suggesting attenuation of architecture-specific advantages under domain shift. Dataset characterization analysis demonstrated differences in subtype distribution, bounding-box geometry, lesion burden, annotation density, and spatial localization patterns between BHX and RSNA+. Conclusions: ICH subtype localization performance is strongly influenced by dataset characteristics, annotation heterogeneity, and domain shift. Although Transformer-based hierarchical feature extraction showed subtype-dependent advantages under internal validation, these advantages diminished under bidirectional external validation. These findings highlight the need for dataset characterization, external validation, patient-level evaluation, and task-specific clinical benchmarks before automated ICH localization models can be considered for real-world clinical integration.
- Research Article
- 10.3390/s26113508
- Jun 2, 2026
- Sensors (Basel, Switzerland)
- Qianqian Chen + 2 more
HighlightsThis study proposes a vision–radar fusion-based dynamic object detection method for maritime navigation scenarios, addressing challenges such as occlusion, scale variation, and multi-target interference. By introducing a cross-modal feature mapping mechanism and an augmented Lagrangian-based fusion strategy, the method effectively integrates complementary visual and radar information. An improved Faster R-CNN framework with optimized RPN and multi-scale training further enhances detection performance. Experimental results on the MVRD demonstrate that the proposed approach achieves high detection accuracy and strong robustness across diverse conditions, including sunny, strong glare, foggy, and densely populated environments, highlighting its potential for reliable environmental perception in intelligent maritime navigation systems.What are the main findings?A vision–radar fusion-based dynamic object detection method is proposed, integrating cross-modal feature mapping and augmented Lagrangian optimization to enhance feature representation and consistency.The proposed method achieves superior detection performance on the MVRD, reaching accuracies of 88.93%, 76.86%, 74.47%, and 83.01% under sunny, strong glare, foggy, and dense traffic scenarios, respectively.What are the implications of the main findings?The proposed fusion framework significantly improves the robustness and reliability of dynamic object detection in complex maritime environments compared with single-sensor approaches.The method provides an effective solution for intelligent navigation perception systems, supporting safer and more reliable autonomous maritime operations.With the rapid development of intelligent navigation technologies, accurate dynamic object detection in complex maritime environments remains a critical challenge due to occlusion, scale variation, and multi-target interference. To address these issues, this study proposes a vision–radar fusion-based dynamic object detection method. A cross-modal feature mapping mechanism is developed to achieve deep integration of visual and radar information, and an augmented Lagrangian optimization strategy is introduced to enhance feature consistency and representation capability. Furthermore, an improved Faster R-CNN framework is designed by optimizing the region proposal network and incorporating a multi-scale training strategy to improve detection performance for objects of varying scales. Experimental results on a self-constructed MVRD show that the proposed method achieves detection accuracies of 88.93%, 76.86%, 74.47%, and 83.01% under sunny, strong illumination, foggy, and crossing-waterway conditions, respectively. These results demonstrate that the proposed approach exhibits strong robustness and stability in complex maritime environments. Overall, the method significantly improves dynamic object detection accuracy and provides effective support for reliable environmental perception in intelligent navigation systems.
- Research Article
- 10.1088/1748-0221/21/06/p06008
- Jun 1, 2026
- Journal of Instrumentation
- Changdong Wu + 1 more
Substation equipment is a critical component of the power system. Accurately detecting substation equipment is essential. In this paper, we propose a substation equipment detection algorithm based on improved Faster R-CNN. Firstly, it introduces the Feature Pyramid Network (FPN) structure on the basis of ResNet-50 and replaces the Region of Interest (ROI) pooling with ROI Align. Then, the anchor box sizes are reconfigured using the K-means clustering algorithm to better match real-world scenarios. Finally, the improved Focal Loss (FL) function is employed to replace the traditional cross-entropy loss. Which can address the poor performance of conventional methods when dealing with the complex background and imbalanced class distributions. Experiments are performed on an infrared image dataset of substation to detect six types of electrical equipment. The robustness tests of our model are carried out as well. The experimental results of detecting the actual substation equipment show that mean Average Precision (mAP) reaches 90.07%, a 7.75% higher over the original network, the Average Recall (AR) reaches 78.0%, a 9.3% higher over the original network, and the processing speed of 57.95 frames per second, satisfying the real-time detection of substation electrical equipment. In addition, our model is robust to occlusion and noise variations. It indicates that the method is an effective substation equipment detection method.
- Research Article
- 10.25077/ajeeet.v6i1.218
- Jun 1, 2026
- Andalas Journal of Electrical and Electronic Engineering Technology
- Palman + 3 more
Workplace accidents remain one of the major issues in industrial environments and are often caused by low compliance with the use of Personal Protective Equipment (PPE), particularly safety helmets. Manual supervision of PPE usage tends to be inefficient and prone to human error. This study aims to develop an intelligent computer-vision-based system capable of automatically and real-time monitoring helmet compliance. The proposed system employs the Faster Region-Convolutional Neural Network (Faster R-CNN) algorithm to detect and classify workers who are wearing and not wearing helmets. The dataset was obtained from CCTV video recordings in industrial areas, which were converted into image frames for training and testing processes. The experimental results show that the system achieved an accuracy of 90% for helmet-wearing workers and 87% for non-helmet-wearing workers during daytime conditions, and 97% and 91% respectively at night. With an average computation time of 0.1 seconds per frame, the system is capable of real-time detection at up to 10 frames per second. These results indicate that the Faster R-CNN method is effective in detecting PPE compliance and has the potential to be implemented as an automated safety-support system in industrial environments.
- Research Article
1
- 10.1016/j.hcc.2025.100359
- Jun 1, 2026
- High-Confidence Computing
- Zhuo Yan + 7 more
FNon R-CNN: A multi-scale ground object detection and recognition network
- Research Article
- 10.1016/j.comnet.2026.112305
- Jun 1, 2026
- Computer Networks
- Junkang Ren + 3 more
ProField R-CNN: A network protocol format extraction algorithm leveraging Fast R-CNN
- Research Article
- 10.3390/jimaging12060245
- May 29, 2026
- Journal of imaging
- Yawen Su + 1 more
Low-light object detection remains challenging because insufficient illumination obscures visual features and increases the discrepancy between training and testing conditions. Existing approaches often rely on detector redesign, image enhancement, or target-domain data, which may introduce additional complexity during training or inference. This paper presents Exposure-Aware Training (EAT), a lightweight degradation-based training strategy that applies illumination attenuation and additive Gaussian noise to normal-light images during training. The degradation parameters are estimated from real low-light image pairs, while the detector architecture remains unchanged. Our experimental results show that moderate degradation consistently improves low-light detection performance, whereas excessively strong degradation may damage semantic information, especially for small objects. Under both cross-domain and mixed-training settings, EAT achieves stable improvements on YOLOv8 and Faster R-CNN, with more noticeable gains for illumination-sensitive categories. These results indicate that exposing detectors to task-oriented illumination degradation during training can effectively improve low-light detection performance without additional inference overhead.
- Research Article
- 10.3390/bioengineering13050591
- May 21, 2026
- Bioengineering
- Rana Gursoy + 7 more
Dermatophytosis is commonly assessed using potassium hydroxide (KOH) microscopy, yet accurate recognition of fungal hyphae is hindered by preparation-related artifacts, heterogeneous keratin clearance, and notable inter-observer variability. This study presents a transformer-based object detection framework using the RT-DETR architecture for precise, query-driven localisation of fungal structures in high-resolution KOH images. A dataset of 2540 routinely acquired microscopy images was manually annotated using a multi-class strategy that explicitly distinguishes fungal elements from confounding artifacts, enabling the model to actively suppress false detections arising from visually similar mimics. To assess architectural trade-offs, RT-DETR was benchmarked against two CNN-based detectors (YOLOv11 and Faster R-CNN) under identical training and inference conditions. Five-fold stratified cross-validation was performed, and each fold-level model was evaluated on the same independent held-out test set (n = 254). Across the five evaluations, RT-DETR achieved a mean AP@0.50 of , a mean recall of , and a mean precision of . At the image level, the model achieved a mean sensitivity of on the independent test set, with a mean of missed positive cases across the five evaluations. These results demonstrate the technical feasibility of a transformer-based artificial intelligence (AI) system as a decision-support aid for fungal region detection in KOH microscopy, pending prospective multi-center validation to establish clinical generalisability.
- Research Article
- 10.12182/20260560204
- May 20, 2026
- Sichuan da xue xue bao. Yi xue ban = Journal of Sichuan University. Medical science edition
- Sisi Zhang + 5 more
By introducing an attention enhancement mechanism and improving the post-processing strategy, the Faster Region-based Convolutional Neural Network (Faster R-CNN) is enhanced. The resulting improved Faster R-CNN method for wound detection (WD-IFRCNN) increases the accuracy and stability of wound detection. ① Algorithm construction: The 50-layer residual network (ResNet-50) is used as the backbone. In the fourth residual stage (convolutional stage 4, Conv4_x) and fifth residual stage (convolutional stage 5, Conv5_x), the fuzzy mask attention module (FMAM) and convolutional block attention module (CBAM) are embedded together to enhance the model's ability to represent features in key areas and blurry regions of the wound surface. At the same time, soft non-maximum suppression (Soft-NMS) replaces the traditional non-maximum suppression strategy to reduce missed detections in overlapping target scenarios. ② Algorithm validation: The experimental dataset consists of open-source wound images and clinically collected images, totaling 740 original images. After data augmentation, the dataset expands to 5920 images and is evaluated using ten-fold cross-validation. Model performance is assessed through internal validation, external validation, hyperparameter tuning, and ablation experiments, using metrics such as precision, recall, average precision, and F1 score. ① Internal validation showed that the model performs best when CBAM is embedded in both the Conv4_x and Conv5_x stages of ResNet-50. When ResNet-50-CBAM is used as the backbone, the model's detection performance surpasses that of VGG16 and ResNet-50. ② External validation showed that WD-IFRCNN achieves an precision of 92.31%, recall of 93.95%, average precision of 92.33%, and F1 score of 0.93. The average precision was 3.21%, 2.30%, 1.39%, 0.86%, 0.82%, 0.37%, and 0.63% higher than SSD, YOLOv4, YOLOv5, YOLOv8, DETR, RT-DETR, and FR-CNN-FPN, respectively. ③ Ablation experiments showed that CBAM, FMAM, and Soft-NMS each positively impact model performance, with the best results achieved when all are used together. WD-IFRCNN effectively enhances the model's ability to represent features in key areas and blurry regions of the wound surface, improves the accuracy and stability of wound detection, and demonstrates good adaptability to complex wound scenarios. It can provide technical support for clinical wound assessment and auxiliary diagnosis.
- Research Article
- 10.1038/s41598-026-51304-5
- May 18, 2026
- Scientific reports
- Parul Priya + 1 more
Inspection of power line insulators is crucial to the safe and reliable operation of power transmission networks. Manual inspection techniques are deficient in terms of efficiency, accuracy, availability, security, and cost. To overcome these problems, this article proposes an improved version, Enhanced Faster R-CNN I2D-Net, which uses faster region-based convolutional neural network (Faster R-CNN) with some quality of feature integration and attention-based features to achieve consistent insulator defects detection in complicated scenes. The suggested architecture is an integration of a residual network backbone (ResNet-50) which extracts features, a bidirectional feature pyramid network (BiFPN) which learns fusion weights to achieve adaptive multiscale fusion of features and receptive field attention plus (RFA+) which optimises the features and improves the overall detection performance. The model also has an Improved Context Perception Module (I-CPM) that comes after the Region Proposal Network (RPN). Also employs dilated convolution to achieve better categorization and bounding-box regression. Additionally, Gradient-weighted Class Activation Mapping (Grad-CAM) is utilized on the feature maps produced by the RFA+ module to illustrate defect-associated areas and create heatmaps that pinpoint defect evidence. This work is new because it successfully combines multi-scale feature fusion, attention mechanisms, and contextual modeling into one framework for strong insulator defect detection. The suggested method works better than the ones that are already out there. It has a mean Average Precision (mAP) of 98.03%, a precision of 99.05%, and a recall of 98.96% for finding faults.
- Research Article
- 10.1038/s41598-026-50240-8
- May 14, 2026
- Scientific reports
- D Ramana Kumar + 6 more
Brain tumor segmentation from multi-modal MRI scans remains challenging due to the heterogeneity of tumors, intensity variations, and different protocols. In this study, a Switchable Normalization-based Faster R-CNN (SNFRC) framework is proposed. Furthermore, by employing the region proposal network (RPN), which uses switchable normalization (SN), heterogeneous data distributions can be tackled effectively to enhance feature consistency. A detection-based segmentation strategy is applied for the explicit localization of the tumor before creating the pixel-wise mask for accurate identification of irregular small tumor regions. Also, a composite loss function was introduced that adds the Dice loss, L2 loss, and Kullback-Leibler divergence losses for jointly optimising all the spatial overlaps, intensity reconstruction, and probabilistic regularisation. To evaluate the performance of the proposed architecture, the subject-wise splits of three benchmark datasets (BraTS 2018, 2019, and 2020) are used. The experimental results demonstrate that the proposed SNFRC outperforms CNN, DenseNet, and even the baseline Faster R-CNN consistently, with Dice scores above 93% and Hausdorff distance decreased by almost 1.5 pixels. Subsequent ablation and normalization experiments validate the efficacy of switchable normalization, yielding improvements of six to seven% over regular normalization. Given the training convergence analysis, this model optimizes faster and more stably than the baseline. The framework shows good performance over the dataset it was trained on, and it is computationally efficient enough that it can be used in real-time. However, achieving cross-dataset generalisation and conducting clinical validation are two important areas for future work. Overall, the SNFRC framework offers a reliable and efficient approach for the automated segmentation of brain tumors in multi-modal MRI.
- Research Article
- 10.1038/s41598-026-48621-0
- May 12, 2026
- Scientific reports
- José Duarte Pereira + 2 more
While automation has transformed many areas inside clinical laboratories, microbiology still relies heavily on manual tasks, particularly the culture of samples on agar plates and their subsequent manual review for microorganism identification and antibiotic susceptibility profiling. Bacterial colony detection and classification require trained professionals, making the process time-consuming and prone to human error. Developing deep learning models to automate these tasks could improve microbiology workflows and accelerate clinical decision-making. In this study we trained and evaluated five object detection architectures (Faster R-CNN and RetinaNet with ResNet-50 and ResNet-101 backbones, and YOLOv8) on the Annotated Germs for Automated Recognition (AGAR) dataset for bacterial colony classification. Transfer learning, cross-subset generalization, and Weighted Box Fusion (WBF) ensemble methods were applied to enhance and characterize performance. Additionally, we created and publicly released a curated dataset of 165 agar plate images containing colonies of S.aureus, P.aeruginosa, and E.coli cultured across four distinct culture media. YOLOv8m achieved a mean Average Precision (mAP) of 69.0% on the AGAR dataset, outperforming the best Detectron2 model (Faster R-CNN ResNet-101, 63.1%) by 5.9 percentage points. A four-model WBF ensemble combining both architectures reached 70.5% mAP (95% CI: 68.4-71.7). Cross-subset evaluation showed that a single model trained on the full dataset generalizes well to individual imaging conditions, making subset-specific fine-tuning largely unnecessary. On the curated dataset, a mixed ensemble reached 58.7% mAP (95% CI: 57.1-63.7). These results demonstrate that architecture choice and training data diversity are the primary drivers of performance for colony detection on agar plates.
- Research Article
- 10.1038/s41598-026-52802-2
- May 12, 2026
- Scientific reports
- Deepthi G Pai + 2 more
Sustainable agriculture has been greatly challenged by the problems of dense canopies that result in extreme occlusion, scale disparities and imbalances of classes, making it difficult to detect small and overlapping weeds in real field conditions. Actual blackgram fields in Karnataka, India, were used to generate a dataset that featured a high level of crop to weed contest, and exhibited a significant amount of scale diversity, with 75 percent of objects comprising less than 0.8 percent of the image space. To address these challenges, the paper presents a better Faster R-CNN framework that adds five modules, i.e. Spatial Attention (SA), Multi-Scale Fusion (MSF), Context-Aware RoI, Shape-Aware Prediction and Adaptive Non-Maximum Suppression (ANMS). The research shows that with larger architectural complexity, there is no guarantee of better performance through an extensive ablation study involving 32 model configurations. The best two-module combination (SA + ANMS) obtained the highest F1 -score of 0.9547, precision of 0.9439, recall of 0.9658, and a mean Intersection over Union (mIoU) of 0.9445, and inference speed of 4.42 FPS, which was higher than the full five-module version. It is important to note that only 12.5% of configurations were better than the baseline, highlighting the fact that too much module stacking may hurt performance. The analysis shows that there is a high positive synergy between SA and ANMS and that the MSF module competes slightly on features. The novelty of this work lies in the systematic exploration of how various enhancement modules can be interacted to function as a single detection framework and show that the selective combination of modules can produce a better outcome than random stacking. Especially, the finding of a complementarity, optimum minimum combination of SA + ANMS confirms that the efficiency-based design is used, not complexity-based expansion. It provides informative insight on designing efficient detection system for precision agriculture applications by revealing the impact of wise module selection instead of architecture complexity on obtaining consistent weed detection against occlusion and inter-class inconsistencies.
- Research Article
- 10.47392/irjaem.2026.0189
- May 5, 2026
- International Research Journal on Advanced Engineering and Management (IRJAEM)
- Ms Ashwini G Gaikwad + 4 more
Accurate detection of blood cells from microscopic images plays an essential role in clinical hematology, disease diagnosis, and laboratory automation. Traditional microscopic examination of peripheral blood smears requires expert pathologists and is time-consuming, especially when large volumes of samples must be analyzed. Recent developments in computer vision and deep learning have enabled automated systems capable of detecting and classifying blood cells with high accuracy. This research presents a comparative evaluation of two advanced object detection architectures—YOLOv9 and Faster R-CNN—for automated blood cell detection in microscopic images. The proposed framework aims to identify and localize three primary blood components: red blood cells (RBC), white blood cells (WBC), and platelets. The system utilizes preprocessing techniques, deep learning-based feature extraction, and bounding-box detection to accurately locate cells within smear images. Experimental evaluation is performed using annotated microscopy datasets. YOLOv9 provides faster inference speed suitable for real-time clinical applications, while Faster R-CNN demonstrates strong localization capability due to its region proposal network. Comparative results indicate that YOLOv9 achieves superior detection speed and competitive accuracy, making it suitable for automated laboratory systems. The proposed system contributes toward improving diagnostic efficiency and reducing manual workload in hematology laboratories.
- Research Article
- 10.1186/s12891-026-09845-3
- May 5, 2026
- BMC musculoskeletal disorders
- Baisen Chen + 6 more
Differentiating acute from chronic wedge-shaped thoracolumbar vertebral deformities on conventional lateral lumbar radiographs remains clinically challenging, especially when osteoporosis status also needs to be considered. This study aimed to develop and evaluate a You Only Look Once (YOLO)v8n framework for vertebral-level detection and classification of thoracolumbar fractures with osteoporosis-related stratification on lateral lumbar radiographs. We retrospectively collected 1352 lateral lumbar radiographs from 1352 patients, with one radiograph per patient. A total of 1774 vertebral fracture segments were manually annotated. Lumbar magnetic resonance imaging (MRI) and dual-energy X-ray absorptiometry (DXA) were used as reference standards to stratify vertebral targets into three categories: acute fracture with osteoporosis, acute fracture without osteoporosis and chronic fracture with osteoporosis. The dataset was divided into training and validation subsets at the patient level. A YOLOv8n detector was trained as the primary model. To strengthen methodological rigor, additional baseline comparison experiments were conducted under the same patient-level training/validation split using YOLOv5n and Faster R-CNN. Detection performance was assessed using precision, recall, F1-score, mean average precision (mAP) 50 and mAP50-95. On the validation set, the YOLOv8n model achieved a precision of 0.495, recall of 0.482, F1-score of 0.490, mAP50 of 0.506, and mAP50-95 of 0.397. In comparative experiments, YOLOv5n achieved a precision of 0.451, recall of 0.549, F1-score of 0.495, mAP50 of 0.494, and mAP50-95 of 0.367, whereas Faster R-CNN achieved a precision of 0.273, recall of 0.814, F1-score of 0.409, mAP50 of 0.300, and mAP50-95 of 0.217. These findings indicate that YOLOv8n provided the most balanced overall detection performance in the present dataset. The proposed YOLOv8n framework demonstrated preliminary feasibility for automated vertebral-level detection and classification of thoracolumbar fractures with osteoporosis-related stratification on lateral lumbar radiographs. However, given the moderate overall performance and lack of external validation, the current model should be regarded as an assistive screening tool rather than a standalone diagnostic system.
- Research Article
- 10.61310/mjst.v24i1.2522
- May 4, 2026
- Mindanao Journal of Science and Technology
- Chelsie D Bajamunde + 3 more
This study evaluated three object detection models for estimating traffic density on Elias Angeles St., Naga City, using mean average precision (mAP). The object detection model classified vehicles into five classes: private cars, jeepneys, trucks, motorcycles, and tricycles. Closed-circuit television (CCTV) footage was subjected to adaptive background subtraction and morphological opening to produce 320px × 320px images for use as a dataset for object detection models. Using Common Objects in Context means Average Precision at Intersection over Union (COCO mAP at IoU=50 as the metric for mAP. Faster Region-Convolutional Neural Network (Faster R-CNN) achieved the highest mAP of 92.54%, compared with You Only Look Once version 3 (YOLOv3) and Single Shot MultiBox Detector (SSD). In the traffic density estimation, vehicle size was accounted for; consequently, a private car was used as the standard vehicle type.