Articles published on Focal Loss
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
2984 Search results
Sort by Recency
- Research Article
- 10.1016/j.phymed.2026.158218
- Jul 1, 2026
- Phytomedicine : international journal of phytotherapy and phytopharmacology
- Peirong Cai + 7 more
Narciclasine alleviates zearalenone-induced focal adhesion damage and anoikis in rat Sertoli cells through PI3K signaling pathway.
- Research Article
- 10.59543/jidmis.v3.441
- Jun 30, 2026
- Journal of Intelligent Decision Making and Information Science
- Tarannum Rafiq Shaikh, Maheshwari Biradar
Breast cancer subtype classification from histopathological images is challenging due to severe class imbalance and subtle inter-class morphological variations. While deep learning methods have shown success in binary diagnosis, robust multi-class subtype discrimination under imbalanced conditions remains insufficiently addressed. This study proposes an imbalance-aware attention-guided hybrid deep learning framework for multi-class breast cancer subtype classification using the BreakHis dataset. The proposed pipeline integrates overlapping patch extraction with brightness-based filtering, preprocessing optimization using a Bilateral–Unsharp enhancement strategy, synthetic minority augmentation via Wasserstein GAN with Gradient Penalty (WGAN-GP), controlled majority-class down sampling, and focal loss optimization. A dual-branch DenseNet–Xception architecture with Convolutional Block Attention Module (CBAM) refinement is employed to capture complementary multi-scale representations. Experimental results demonstrate progressive improvement from 73.22% testing accuracy on raw patches to 91.42% after preprocessing and further to 92.97% following imbalance-aware synthetic augmentation, achieving a macro-F1 score of 0.9314 and macro-AUC of 0.997. Comparative evaluation against pretrained, custom, and alternative hybrid models confirms the superiority of the proposed framework. The findings highlight the importance of integrating preprocessing optimization, imbalance-aware training, and hybrid attention-guided architectures for reliable multi-class breast cancer subtype classification.
- Research Article
- 10.7507/1001-5515.202411017
- Jun 25, 2026
- Sheng wu yi xue gong cheng xue za zhi = Journal of biomedical engineering = Shengwu yixue gongchengxue zazhi
- Xuelian Gu + 6 more
Cervical intraepithelial neoplasia is the primary type of cervical precancerous lesion; however, manual clinical diagnosis is prone to bias and has limited grading accuracy. To achieve precise automated grading of CIN, this paper proposes a multimodal fusion Swin Transformer model and develops a corresponding computer-aided diagnosis system. This method employs three-channel fusion of raw images, cervical mask images, and directional gradient histogram features to enhance lesion texture and location information. Within the Swin Transformer backbone, an atrous spatial pyramid pooling module channel attention module and a convolutional feature extraction module are embedded to balance global semantic and local detail features. A focal loss function is adopted to address class imbalance in the dataset and improve the model's ability to identify difficult-to-classify samples. On a dataset of 3 915 clinical colposcopy images, the model achieved an overall accuracy of 90.01%, precision of 87.55%, recall of 86.17%, F1 score of 89.13%, outperforming baseline models such as VGG, ResNet, and Swin Transformer. The developed system integrates image quality screening, lesion identification, and three-level classification functions, providing an effective tool for the rapid and objective screening of clinical cervical precancerous lesions.
- Research Article
- 10.7507/1001-5515.202508042
- Jun 25, 2026
- Sheng wu yi xue gong cheng xue za zhi = Journal of biomedical engineering = Shengwu yixue gongchengxue zazhi
- Hubin Yan + 3 more
The evaluation of disability grades in traffic accidents is a professional forensic clinical appraisal matter, and its results directly affect the fairness of judicial compensation. In the construction of automated disability grade evaluation models, the imbalanced distribution of disability cases leads to low recognition accuracy for minority categories, becoming a key bottleneck restricting the technology's implementation. In response, this paper proposes an imbalanced data classification method based on a hybrid parameter scaling weight optimization mechanism. First, a loss weight calculation model is constructed based on category proportion, category sparsity, and category diversity. Second, the loss weight calculation model is designed by integrating the focal loss function's ability to focus on hard samples with the cross-entropy loss function's global gradient stability advantage. Then, at the early stages of training, the model proposed in this paper aligns sensitivity to imbalanced categories and constructs a low-computational-demand hybrid parameter scaling weight optimization mechanism. Experimental results show that, compared with the best-performing baseline methods, the proposed method significantly improves both accuracy and macro-F1 score on the traffic accident disability grade dataset. It can effectively enhance the classification performance of minority grade categories in imbalanced data and help improve the accuracy of automated appraisal in judicial identification of traffic accident disability grades.
- Research Article
- 10.1002/mdc3.70715
- Jun 19, 2026
- Movement disorders clinical practice
- Antonio Costa + 8 more
Sporadic Progressive Ataxia with Palatal Tremor (PAPT) is an extremely rare movement disorder syndrome with only three autopsy reports published in the literature to date. Previously described cases showed hypertrophic olivary degeneration with tau-positive neuronal inclusions, although differences were noted in tau isoforms and the presence of additional neurodegenerative pathology. We report an additional autopsy case with morphological similarities in pathology but without tau inclusions. To describe clinically, radiologically and neuropathologically a case of PAPT. Retrospective clinical data collection from electronic records, authorized video material and postmortem neuropathology study. A 74-year-old women presented with a gait ataxia and parkinsonism at the age 67, followed three years later by rhythmic chin and palatal tremor. Probable REM sleep behavior disorder was reported approximately ten years before symptom onset. MRI demonstrated bilateral T2/T2-FLAIR hyperintensity and enlargement of the inferior olivary nuclei. The clinical progression was dominated by the ataxia, and the tremor remained throughout the disease course. Neuropathological findings showed hypertrophic olivary degeneration with glomeruloid bodies, without tau pathology and occasional neuronal and dendritic p62 immunoreactivity. Focal neuronal loss and gliosis were observed in cerebellar dentate nucleus. There was substantia nigra neuronal loss with Lewy pathology in neocortical stage (Braak stage 5), but without involvement of inferior olivary nucleus or cerebellum. Our findings confirmed hypertrophic olivary degeneration as the pathologic substrate of PAPT. The absence of tau pathology distinguishes it from previously reported cases, indicating heterogeneity in the underlying pathophysiological disease mechanisms. Concomitant Lewy body pathology may have contributed to parkinsonism and REM sleep behavior disorder.
- Research Article
- 10.1097/pas.0000000000002568
- Jun 17, 2026
- The American journal of surgical pathology
- Kyung Park + 9 more
Papillary renal neoplasm with reverse polarity (PRNRP) has been proposed as a distinct subtype of renal cell neoplasm with recurrent KRAS mutations and indolent behavior. However, its epigenetic landscape is poorly understood. In this study, 12 PRNRPs and a PRNRP initially diagnosed as "papillary adenoma" were analyzed. All 13 cases underwent targeted next-generation sequencing for driver mutations. Eleven PRNRPs were profiled using the Illumina MethylationEPIC array and compared with a reference cohort of 71 common renal cell tumors. KRAS mutations were identified in 12 of 13 (92%) cases of PRNRP. Copy-number analysis from methylation profiling showed that 9 of 11 (82%) PRNRPs lacked copy-number changes. Two cases showed a focal loss of chromosome 8 and a gain of chromosome 16, respectively. Unsupervised clustering based on methylation data showed that PRNRPs form a distinct epigenetic group, separate from papillary renal cell carcinomas (pRCCs) and other major renal tumors, but with the closest affinity to clear cell papillary renal cell tumors. In addition, DNA methylation analysis suggested PRNRP may arise from the distal nephron, in contrast to pRCC, which appears to recapitulate proximal tubules. These findings support PRNRP as a subtype of renal cell neoplasm with a distinct epigenetic signature.
- Research Article
- 10.1094/pdis-07-25-1440-re
- Jun 16, 2026
- Plant disease
- Xinlong Li + 5 more
The precise and timely identification of apple leaf diseases play a key role in targeted pesticide application in orchards. Conventional deep learning techniques encounter issues like the substantial size of model parameters and low detection accuracy across various disease scales in natural environments. To overcome these limitations, this paper presents YOLO-APLD, a lightweight algorithm for detecting apple leaf diseases, utilizing the improved YOLOv8n model. The proposed model incorporates four key improvements to enhance its detection performance. First, an EP-C2f enhancement module is embedded at the output of the backbone to strengthen the representation of local and structural features of damaged area, thereby achieving significant improvements in the recognition of morphologically complex diseases such as rust. Additionally, spatial intersection over union (SIoU) loss and focal loss are combined to form Focal-SIoU loss, which simultaneously optimizes bounding box regression and classification, thus enhancing the detection stability for hard-to-distinguish samples and few-shot categories including mosaic and brown spot. Meanwhile, a bidirectional feature pyramid network is adopted in the neck for efficient multiscale feature fusion, which strengthens the perceptual capability for both large-scale damaged area (powdery mildew and scab) and small-scale damaged area (Alternaria blotch and gray spot). Finally, a Slim-neck structure is employed to simplify the feature fusion architecture, reducing model size and accelerating inference speed. Comprehensive experiments demonstrate that YOLO-APLD achieves excellent performance while maintaining real-time capability, with precision, recall, mean average precision, and F1-score reaching 88.5, 84.3, 88.5, and 86.4%, respectively. Compared with YOLOv8n, these metrics show respective improvements of 1.7, 1.5, 0.8, and 1.6%. Meanwhile, floating point operations, parameter count, and model size are reduced by 22.2, 23.3, and 17.5%, respectively. The detection frame rate on edge computing devices reaches 90.3 f/s, indicating significantly accelerated inference speed. Additionally, testing performance on grape and tomato datasets further validates the generality of the proposed method. In summary, YOLO-APLD exhibits strong detection performance in the field of apple leaf disease detection and can provide practical technical support for precision pesticide application in orchards and on-site disease monitoring.
- Research Article
- 10.1021/acs.jcim.6c00985
- Jun 16, 2026
- Journal of chemical information and modeling
- Shencheng Zhou + 2 more
Accurate identification of space groups from powder X-ray diffraction (pXRD) is essential for understanding crystal structures and accelerating materials discovery. However, this task remains highly challenging due to inherent peak overlap, experimental noise, and the complexity of the 230-class classification problem. To address the critical issues of class imbalance and data scarcity, we first design a general physics-informed data augmentation pipeline. We then propose a dual-channel fusion uncertainty-aware network (DFUN) for automated space group classification. The DFUN architecture integrates two complementary feature representations: convolutional features extracted directly from raw diffraction profiles and domain-specific peak descriptors. These distinct representations are adaptively fused through a gating mechanism. Furthermore, to mitigate the inherent long-tailed distribution of crystallographic data, we employ a hybrid loss function that combines Focal Loss with Label Smoothing. Finally, we incorporate Monte Carlo Dropout to provide predictive uncertainty estimation, thereby enabling not only accurate classification but also a crucial assessment of the model's reliability. Evaluated on large-scale simulated data and two public data sets (opXRD and RRUFF), DFUN outperforms the evaluated baseline methods across the reported metrics. The framework also provides uncertainty-aware predictions, establishing DFUN as a robust and interpretable solution for high-throughput automated crystallographic analysis from powder diffraction.
- Research Article
- 10.5414/np301727
- Jun 15, 2026
- Clinical neuropathology
- Jeroen J De Vries + 2 more
Episodic ataxias (EAs) are rare autosomal-dominant channelopathies presenting with recurrent attacks of ataxia and variable neurological features. While MRI studies suggest mild cerebellar atrophy in some patients, detailed neuropathological descriptions are lacking. We examined postmortem brains and spinal cords from two genetically confirmed patients: EA1 (KCNA1 Val174Phe mutation) and EA2 (CACNA1A A253Y mutation). Routine histology and immunohistochemistry were performed. In EA1, the cerebellar vermis and hemispheres showed narrowing of the folia, segmental Purkinje cell loss, Bergmann gliosis, and scattered torpedoes. The dentate nucleus contained focal lipofuscin-laden macrophages and Rosenthal fibers. Concurrently, α-synuclein pathology (Braak stage 3) was also present, along with sparse age-related tau tangles and minimal vascular amyloid. In EA2, only focal Purkinje cell loss was detected, but novel p62-positive axonal and dendritic inclusions were observed in the cerebellar cortex and inferior olive. No additional α-synuclein, tau, or amyloid pathology was found. A concurrent glioblastoma was identified as the immediate cause of death. These are the first detailed autopsy reports of EA1 and EA2. Both revealed cerebellar pathology, with EA1 showing more pronounced Purkinje cell degeneration, while EA2 demonstrated previously undescribed p62-positive inclusions. These findings expand the clinicopathological spectrum of EAs and highlight the value of postmortem studies in elucidating underlying disease mechanisms.
- Research Article
- 10.1038/s41598-026-55382-3
- Jun 10, 2026
- Scientific reports
- R Gokul + 8 more
Cardiac arrhythmias are a leading cause of cardiovascular mortality worldwide, and auto- mated analysis of electrocardiogram (ECG) recordings is an unmet clinical need. This paper introduces a hybrid 1D Convolutional Neural Network (CNN) architecture for automated ECG heartbeat classification that captures both multi-scale morphological and long-range temporal dependencies. The architecture combines Dynamic Kernel Switching (DKS) with adaptive gating over kernel sizes {3, 5, 7}, a Self-Attention module, and Dilated Convolutions at dilation rates 2 and 4. For low-power edge deployment on Neural Processing Units (NPUs) and FPGA/ASIC platforms, three quantization paths are evaluated: (1) Dynamic Post-Training Quantization (PTQ) with INT8 weights and FP32 I/O, (2) Full-Integer PTQ with INT8 weights and activations, and (3) Quantization-Aware Training (QAT). With a corrected stratified 72/8/20 train/validation/test protocol, focal loss (γ = 2.0), AdamW optimization and cosine learning rate decay, the proposed model achieves 98.99% test accuracy and a macro F1-score of 0.9484 on the MIT-BIH Arrhythmia Database (16,219 held-out test beats, five AAMI classes). Five-fold stratified cross-validation confirms statistical robustness: 98.93% ± 0.10% accuracy with 95% confidence interval [98.84%, 99.02%]. A comprehensive seven-variant ablation study covering Self-Attention, CBAM, CSA, and SimSA demonstrates that the proposed DKS + Self-Attention combination outperforms all alternatives. Generalization is further validated on five independently withheld MIT-BIH patient recordings (10,444 beats), achieving 92.86% accuracy and macro F1 of 0.8219. The model comprises only 331,733 parameters and reduces to ≈0.32MB under Dynamic PTQ, enabling direct TFLite/INT8 conversion for RTL-level FPGA deployment.
- Research Article
- 10.1038/s41598-026-55588-5
- Jun 10, 2026
- Scientific Reports
- Sadia Munawar + 4 more
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and accurate histopathological classification is essential for timely diagnosis and treatment planning. This study presents a Contrastive Language Image Pretraining (CLIP)-based framework for multiclass lung histopathology classification, designed to distinguish among benign lung tissue, lung adenocarcinoma, and lung squamous cell carcinoma. The proposed approach leverages a pretrained CLIP ViT-B/32 backbone, domain-specific prompt engineering, multimodal image text pairing, and similarity-based classification within a shared embedding space. To strengthen convergence and robustness during fine-tuning, the training pipeline incorporates data augmentation, Focal Loss, AdamW optimization, OneCycle learning rate scheduling, mixed-precision training, gradient clipping, and early stopping. The dataset is organized into separate training, validation, and testing splits, with the reported training and validation partitions containing 3,500 and 500 images per class, respectively. Experimental training on a Tesla T4 GPU demonstrated steady performance improvement across epochs, with the best validation accuracy reaching 95.20%, accompanied by a macro AUC of 0.9870 and a micro AUC of 0.9877, before early stopping was triggered at epoch 23. These findings indicate that integrating CLIP with pathology-specific text prompts provides a strong and reliable framework for automated lung cancer histopathology classification, with promising potential for future intelligent digital pathology systems.
- Research Article
- 10.48084/etasr.17576
- Jun 6, 2026
- Engineering, Technology & Applied Science Research
- Chaithra Reddy + 1 more
The increasing presence of pesticide residues in vegetables poses a major threat to public health, and there is an urgent need to develop efficient, accurate, and scalable detection methods. Traditional analytical techniques such as Gas Chromatography (GC) and Liquid Chromatography–Mass Spectrometry (LC–MS) offer high sensitivity but are expensive, labor-intensive, and unsuitable for large-scale or real-time screening applications. Recent advances in spectroscopy, machine vision, and Machine Learning (ML) show promise; however, most existing models rely on handcrafted or shallow features and fail to capture the complex spatial, textural, and thermal variations associated with multi-residue contamination. In this direction, the present study proposes a data-driven hybrid deep learning framework for multi-residue risk classification in vegetables using thermal imaging. This framework integrates ΔT-based thermal image preprocessing, Gray-Level Co-occurrence Matrix (GLCM) texture descriptors, and deep feature embeddings extracted from InceptionV3, thereby forming a comprehensive hybrid feature vector. This fused representation is classified by a custom neural network trained with categorical focal loss to mitigate class imbalance and optimized using a cosine-decay learning rate to enhance convergence stability. Experimental evaluation on a custom thermal vegetable image dataset resulted in 84.97% validation accuracy and a loss of 0.0373, outperforming conventional Convolutional Neural Networks (CNNs) and other shallow classifiers. The model demonstrated good generalization with balanced precision and recall on contamination classes, supported by a well-converged training–validation performance and confusion matrix analysis. These results highlight the efficacy of this framework for non-destructive, real-time, and scalable pesticide contamination risk classification, pointing to its potential for deployment in automated food safety monitoring and smart agricultural inspection systems.
- Research Article
- 10.1016/j.jos.2026.04.010
- Jun 4, 2026
- Journal of orthopaedic science : official journal of the Japanese Orthopaedic Association
- Gong Long + 5 more
Medial Tibial Plateau depth is a key morphological determinant of cartilage lesion severity in anteromedial knee osteoarthritis.
- Research Article
- 10.1186/s12880-026-02477-y
- Jun 4, 2026
- BMC medical imaging
- Hamidreza Rokhsati + 3 more
Gastric intestinal metaplasia (GIM) is often visually inconspicuous on routine endoscopy, while many artificial intelligence systems rely on dense supervision, lack calibrated probabilities, or provide limited evidence of transfer across datasets and devices. We developed a single-frame, four-class endoscopic classifier that jointly models anatomic context and metaplasia status. We propose a dual-stream architecture that combines an RGB Swin Transformer backbone with SHAP-guided lesion-aware multi-scale auxiliary tokens. The two streams are fused through class-token attention to obtain a compact and interpretable representation without relying on video context or pixel-level masks. To address class imbalance, training combines class-balanced focal loss, balanced-softmax/logit adjustment, class-aware sampling, and validation-tuned per-class thresholds. The model was evaluated on an internal four-class cohort of 666 endoscopic still frames and externally assessed, without retuning, on an unseen public endoscopy dataset recast as a binary normal-versus-abnormal task. On the internal cohort, the proposed model achieved macro-AUROC 0.950, macro-AUPRC 0.920, macro-F1 0.926, and accuracy 0.954, with expected calibration error 0.034 and inference latency of approximately 182 ms per 224 × 224 frame. On the unseen external dataset, the model retained AUROC 0.940, AUPRC 0.900, and F1 0.890 using frozen operating thresholds. Comparative and ablation analyses indicated that lesion-aware tokenization and token fusion contributed more strongly to performance gains than backbone choice alone, while calibration quality also improved. A dual-stream, single-frame token-fusion model can provide accurate, calibrated, and interpretable classification of gastric intestinal metaplasia while remaining compatible with low-latency edge-oriented inference. Although broader multicenter validation is still required, the results support the feasibility of deployment-oriented AI assistance for endoscopic GIM triage.
- Research Article
- 10.3390/bios16060326
- Jun 3, 2026
- Biosensors
- Abdullah + 5 more
Electrocardiogram (ECG) analysis is a cornerstone non-invasive diagnostic technique for detecting cardiac arrhythmias, which remain a leading cause of mortality worldwide. While recent advances in deep learning have significantly improved automated arrhythmia classification, the current literature lacks systematic, fair comparisons of fundamental neural architectures under unified experimental conditions, and very few studies provide model interpretability. This study addresses these gaps by first providing a rigorous comparative analysis of three representative architectures-Artificial Neural Network (ANN), Convolutional Neural Network (CNN), and Residual Network (ResNet)-on the MIT-BIH Arrhythmia Database under identical preprocessing, training, and evaluation protocols. We then propose an efficient Fine-Tuned CNN (FT-CNN) optimized for ECG signal characteristics through adaptive kernel sizing for P-QRS-T morphological extraction, multi-faceted regularization including L2, dropout, and batch normalization, cosine annealing learning rate, and a custom loss function combining weighted categorical cross-entropy with focal loss with gamma equal to 2.0 to address severe class imbalance. The FT-CNN achieves an accuracy of 98.51%, outperforming fourteen benchmark models, including standard CNN with an accuracy of 97.20%, ResNet with 96.88%, LSTM with 96.50%, GRU with 96.30%, and traditional classifiers. Comprehensive ablation studies confirm an improvement of 6.17% over the baseline. Class-wise analysis reveals excellent performance for normal beats with an F1-score of 0.99, ventricular ectopic beats with 0.95, and unknown beats with 0.98, while supraventricular ectopic beats with an F1-score of 0.79 and fusion beats with 0.70 remain challenging. Unlike most prior works, we integrate Grad-CAM and Integrated Gradients for explainability, quantitatively evaluating attribution faithfulness, sanity checks, and noise robustness.
- Research Article
- 10.1038/s41598-026-56149-6
- Jun 3, 2026
- Scientific reports
- Qingming Hou + 4 more
Defective sealing of outer packaging can result in squeezed, deformed, or fallen-off products, inevitably leading to losses for businesses or individuals. Thus, detecting tape-sealing defects in external packaging is crucial in industrial manufacturing. Systems for detecting surface defects based on visual perception are popular in industrial quality inspection. Nevertheless, traditional surface defect detection methods demonstrate limited precision and slow performance. Multi-scale, multi-type defects, numerous small targets, and complex background interference characterize surface defects in industrial products. Detecting minor, multi-scale defects amidst complex background interference is significantly challenging. Therefore, developing algorithms that can accurately detect industrial defects remains a challenging problem. Consequently, we propose a model that uses YOLOv8 to detect tape-sealing defects of cigarette carton. Using diverse branch blocks (DBB), The backbone extracts multi-scale representations and enlarges the effective receptive field. Secondly, implementing a combined module (CCFM+C2fiRMA) in the Neck enhances the detection accuracy. Also, we adopt focal loss to further enhance the detection of tiny objects. Finally, experimental results show that our model achieves 98.1% mAP50 and 67.8% mAP50-95 on the QZU-DET dataset, outperforming YOLOv8n by 3.2 and 2.2% points, respectively. In addition, our method attains an F1 score of 91.6%, indicating strong detection reliability. Overall, the proposed model enables efficient and accurate tape-sealing defect detection in practical industrial scenarios.
- Research Article
- 10.3390/bios16060320
- Jun 2, 2026
- Biosensors
- Yea-Jin Park + 1 more
Herbal medicines represent a significant global market, yet food safety remains threatened by counterfeit products morphologically resembling authentic samples. Models trained on limited datasets are prone to shortcut learning, relying on superficial features rather than intrinsic morphological characteristics. This study identified size-based shortcut learning as a critical factor degrading the classification of Ziziphus jujuba Mill. var. spinosa and its counterfeit Ziziphus mauritiana Lam., and demonstrated that focal loss alone can effectively mitigate this issue. Models trained on the internal dataset were evaluated on an external dataset acquired with the Herb-X. On the internal test set, all configurations achieved high classification accuracies (≥98%), thereby obscuring meaningful differences in external generalization. However, consistent performance degradation was observed on the external dataset. The cross-entropy model trained on background-removed data dropped to 82.08 ± 10.97%, while size-normalized models recovered to 84.17 ± 10.15% (upsizing) and 88.94 ± 6.76% (downsizing), confirming that suppressing size shortcuts improves external generalization. The focal loss model, without any preprocessing, achieved 90.88 ± 2.71%, reducing the internal-external generalization gap from 16.18 to 8.11 percentage points. Grad-CAM++ and loss analyses confirmed that the focal loss model attended to intrinsic morphological features rather than object size. This study provides a practical, preprocessing-free approach for reliable herbal-medicine authentication in field conditions.
- Research Article
- 10.1088/1748-0221/21/06/p06008
- Jun 1, 2026
- Journal of Instrumentation
- Changdong Wu + 1 more
Substation equipment is a critical component of the power system. Accurately detecting substation equipment is essential. In this paper, we propose a substation equipment detection algorithm based on improved Faster R-CNN. Firstly, it introduces the Feature Pyramid Network (FPN) structure on the basis of ResNet-50 and replaces the Region of Interest (ROI) pooling with ROI Align. Then, the anchor box sizes are reconfigured using the K-means clustering algorithm to better match real-world scenarios. Finally, the improved Focal Loss (FL) function is employed to replace the traditional cross-entropy loss. Which can address the poor performance of conventional methods when dealing with the complex background and imbalanced class distributions. Experiments are performed on an infrared image dataset of substation to detect six types of electrical equipment. The robustness tests of our model are carried out as well. The experimental results of detecting the actual substation equipment show that mean Average Precision (mAP) reaches 90.07%, a 7.75% higher over the original network, the Average Recall (AR) reaches 78.0%, a 9.3% higher over the original network, and the processing speed of 57.95 frames per second, satisfying the real-time detection of substation electrical equipment. In addition, our model is robust to occlusion and noise variations. It indicates that the method is an effective substation equipment detection method.
- Research Article
- 10.1186/s44342-026-00074-7
- Jun 1, 2026
- Genomics & Informatics
- Rashmi Siddalingappa + 6 more
PurposePrecision oncology depends on identifying cancer driver genes and linking them to targeted therapies. Current methods using curated gene sets or generic classifiers often miss biologically relevant patterns in complex gene interaction networks.MethodsWe developed the Precision Medicine Gene Network Analyser, integrating network topology analysis with machine learning for cancer gene identification. The dataset included 699 cancer driver genes (COSMIC Cancer Gene Census) and 15,050 background genes, mapped to high-confidence protein–protein interaction networks from STRING (456,300 edges, 15,749 nodes). Network features such as degree, betweenness, PageRank, k-core, and clustering coefficients were extracted. Imbalance Aware Network Integrator (IANI) was proposed to address class imbalance, where balanced resampling and ensemble models (logistic regression, random forest, gradient boosting) were combined with deep neural networks using focal loss, optimising thresholds for maximum F1-score. Hub genes were defined using a statistical cutoff of mean outdegree + 2 × SD (standard deviation).ResultsOn a test set of 3150 samples (140 cancer, 3010 non-cancer genes), the optimised ensemble improved ROC-AUC from 0.84 to 0.96, precision from 0.78 to 0.90, and recall from 0.42 to 0.81 (F1 = 0.85) at a threshold of 0.466. Hub analysis identified 689 hubs with fourfold enrichment of cancer genes (16.1% vs. 4.4%, p < 10 − 20), showing higher betweenness centrality (p < 0.001). Key features such as degree (0.32), betweenness (0.24), and PageRank (0.19) contributed 75% of the model’s performance. Top hubs (TP53: 758, EGFR: 512, AKT1: 415 connections) showed 60–67% cancer gene enrichment, with pathway clustering in p53 signalling (75%) and cell cycle regulation (67.7%).ConclusionIntegrating protein interaction topology with imbalance-aware machine learning achieved 96% discrimination accuracy. This work forms a base for the upcoming phases of drug-gene mapping and patient-specific therapy prediction within the Precision Medicine Gene Network Analyser.Supplementary InformationThe online version contains supplementary material available at 10.1186/s44342-026-00074-7.
- Research Article
- 10.1016/j.aiia.2026.03.006
- Jun 1, 2026
- Artificial Intelligence in Agriculture
- Daniel Alexander Méndez + 4 more
Computer vision offers significant potential for the continuous, stress-free, and cost-effective monitoring of animal behavior, yet its application in goat farming remains limited. In this study, a Multi-Object Detection (MOD) model was developed to classify goat behavior into four categories: eating, standing, drinking, and lying. Zenithal videos were recorded under both light and dark conditions on an experimental goat farm, resulting in 10,740 labelled annotations used to train, validate, and test 13 models leveraging Transformer- and CNN-based pretrained architectures. YOLO-based models (YOLOv8 and YOLOX) achieved the highest overall performances across both large and lightweight versions, demonstrating high detection capability and potential suitability for hardware-constrained scenarios. YOLOX-based MOD model is preferred for goat behavior detection due to its superior classification accuracy, fast inference speed, and fully open-source license, enabling flexible customization, deployment, and reproducibility. Other models, particularly DAB-DETR and H-DINO, underperformed, especially in detecting drinking behavior, which represents the most challenging class due to its visual similarity with standing, class imbalance, and fisheye distortion effects that affect the frame regions where drinkers are located. Mitigation strategies, including focal loss and distortion correction, improved detection accuracy for this class and reduced performance variability. The developed MOD model can be deployed for continuous group-level monitoring of goats, paving the way for scalable and efficient solutions for advanced behavioral analyses. Future works will focus on integrating tracking algorithms for animal-level insights, as well as on evaluating model generalizability across different farming conditions and goat breeds. • Computer vision applications for goat farming are still very limited in literature. • Various models for goat behavior detection were developed and tested. • YOLOX achieved a mAP @ 0.50 of 0.957 with strong performances across the classes. • Drinking detection is hard due to visual similarity, class imbalance, distortion. • Frame undistortion improves drinking detection performance and reduces variability.