Articles published on Ensemble Learning Approach
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
1575 Search results
Sort by Recency
- New
- Research Article
- 10.1108/jqme-10-2025-0120
- Jul 1, 2026
- Journal of Quality in Maintenance Engineering
- Pauline Ong + 5 more
Purpose Rolling bearing failures remain a primary cause of induction motor breakdowns, creating significant reliability and maintenance challenges. This study aims to enhance fault diagnosis robustness under variable-speed conditions by addressing the limitations of fixed-speed datasets and single-model deep learning approaches. Design/methodology/approach An ensemble deep learning model (EDLM) is developed by integrating convolutional neural networks (CNN), deep belief networks (DBN) and stacked autoencoders (SAE). The framework applies a weighted-fusion strategy to exploit the complementary strengths of the base learners. Diagnostic performance is evaluated using three input modalities: raw vibration signals, spectrograms and infrared thermal images. Findings The EDLM consistently outperformed individual models across all metrics. Among the modalities, infrared thermal images achieved the highest diagnostic accuracy (98%), demonstrating superior capability in capturing subtle fault features under variable-speed conditions. Originality/value To the best of current knowledge, this study presents the first ensemble of CNN, DBN and SAE for bearing fault diagnosis under variable-speed conditions. By validating multi-sensor inputs and highlighting the diagnostic advantage of infrared thermography, the work provides a reliable and scalable solution for real-world machinery health monitoring.
- New
- Research Article
- 10.1016/j.jconhyd.2026.104973
- Jul 1, 2026
- Journal of contaminant hydrology
- Jiani Wang + 6 more
Unveiling the drivers of groundwater quality in an industrial plain: An integrated hydrogeochemical and stacking ensemble learning approach.
- Research Article
- 10.1038/s41598-026-58932-x
- Jun 23, 2026
- Scientific reports
- S Prejesh + 2 more
Modern day healthcare has seen an increase in polypharmacy, which is the prescription of multiple drugs as medication to treat illnesses simultaneously. Therefore there is an increased risk of adverse drug reactions resulting from drug-drug interactions. Existing techniques in the field of pharmacovigilance suffer from many drawbacks. Many machine learning approaches using single models face difficulties in identifying complex interaction patterns. Some methods also overlook the fact that a single adverse event may contain a large number of drugs. In terms of interpretability in a clinical context, mere predicting a combination to be risky or not may not provide a clear enough picture. In an effort to address these challenges, in this study, we propose an ensemble learning approach to effectively predict adverse drug combinations using data obtained from the FDA Adverse Event Reporting System (FAERS). Our proposed framework also makes use of DrugBank for mapping drugs and incorporates binary feature vector representations to handle the complexities of the pharmacovigilance data. The ensemble model developed in this study composed of logistic regression, random forest, and CatBoost algorithms proved to be effective compared to several existing techniques in detecting drug interactions with an accuracy of 93.6%, recall of 97.9%, ROC-AUC of 97.5% and PR-AUC of 96%. In addition to achieving strong predictive performance, the model also calculates a confidence score representing the risk associated with specific drug combinations. These results show how ensemble learning can help to enhance the detection of adverse drug reactions and serve as a clinical decision support tool.
- Research Article
- 10.1038/s41598-026-57514-1
- Jun 22, 2026
- Scientific reports
- Parna Chaudhury + 3 more
In this paper, several audio features are leveraged to improve detection accuracy through an advanced ensemble learning approach. For the extraction of diverse signal characteristics, the method uses Mel-frequency cepstral coefficients (MFCCs), statistical features, and short-time Fourier transform (STFT) features. Specialized neural networks are utilized to handle each set of features: MLPs categorize speech according to statistical descriptions, RNNs detect temporal anomalies from MFCC, and CNNs examine spectral patterns from STFT. These are finally combined by a meta-learning model through the stacking ensemble method, enhancing robustness and decreasing misclassification. Compared to individual models, the system is more effective when evaluated against the Fake-or-Real dataset. Based on experimental findings, deepfake detection is significantly enhanced and highly accurate when spectral, temporal, and statistical data are fused. In addition, the model is hosted using Gradio on Hugging Face, providing real-time deepfake detection.
- Research Article
- 10.1080/02533839.2026.2673014
- Jun 12, 2026
- Journal of the Chinese Institute of Engineers
- Amit Thakur + 1 more
ABSTRACT The research presented here develops an advanced prediction framework that incorporates various analytical methods to address the critical clinical need for early post-COVID recovery diagnosis. The research establishes the need for reliable predictive and diagnostic models to identify chest infections. To address this problem, a novel method is devised: Input X-ray images are pre-processed using the TV-L1 norm. The TV-L1 norm regularization plays a dual role in noise reduction and preservation of critical structures. Further, a hybrid feature extraction strategy is devised that combines Quaternion Wavelet Transforms (QWT) and texture features extracted using Maximally Stable Extremal Regions (MSER). The research also significantly emphasizes feature selection to streamline the dataset and improve model efficiency. Aquila Optimization is introduced for this purpose, helping select the most informative features. An ensemble learning approach is used for the predictive modeling phase, and a Spiking Neural Network (SNN) is fine-tuned using Sine Cosine Optimization (SCO). A combination of algorithms yields a highly accurate chest infection diagnostic model, achieving 99.53% accuracy. The results of this research hold immense promise for accurate diagnosis and prediction of chest infections in patients.
- Research Article
- 10.1371/journal.pone.0349122
- Jun 12, 2026
- PLOS One
- Hong-Jae Choi + 8 more
BackgroundEndovascular aneurysm repair (EVAR) for abdominal aortic aneurysm (AAA) is associated with risks such as endoleaks and late aneurysm rupture, highlighting the importance of long-term survival prediction. Despite recent advancements in machine learning (ML), predictive models utilizing time-to-event analysis remain limited for AAA patients undergoing EVAR. We aimed to develop a stacking ensemble ML model to predict long-term outcomes in EVAR-treated AAA patients.MethodsFrom 2002 to 2019, a total of 12,312 patients underwent EVAR. The primary outcome was AAA-related mortality, with follow-up until December 31, 2019. Using 5 ML algorithms, we developed a model comprising 34 variables. Model performance was assessed using the time-dependent C-index and Brier score. Variable importance was evaluated through permutation-based and partial dependent plots.ResultsThe stacking ensemble model showed the best predictive performance among the tested models (time-dependent C-index: 0.759 at 30 days, 0.716 at 365 days). The time-dependent Brier scores generally increased slightly over time but remained stable across all ML algorithms. Important predictors included age, smoking status, duration between diagnosis and surgery, household income, renal function, and blood pressure. Variable importance differed over time, and each predictor presented a nonlinear relationship with AAA-related mortality risk.ConclusionThe stacking ensemble ML model for time-to-event prediction identified dynamic, time-varying changes in predictor importance, providing improved risk stratification and phase-specific management after EVAR.
- Research Article
- 10.1038/s41598-026-54224-6
- Jun 11, 2026
- Scientific reports
- Fan Yu + 5 more
The inherent performance conflicts among the acoustic, mechanical, and hydraulic properties of pervious concrete represent a core obstacle to its application as a low-noise pavement material. To address this challenge, this paper proposes a multi-objective synergistic optimization method based on Stacking ensemble learning and the NSGA-II algorithm to proactively optimize mix proportions, thereby achieving a balance and enhancement of multiple performance metrics. A comprehensive database, comprising both proprietary experimental data and data from the literature, was first established to systematically train and construct a high-precision predictive model for the sound absorption performance of pervious concrete. Subsequently, this model was combined with previously established models for compressive strength and permeability to serve as the fitness functions for the NSGA-II genetic algorithm, which performed a multi-objective search for optima. The accuracy and reliability of the optimization results were then confirmed through experimental validation. Results indicate that aggregate gradation has a significant impact on the sound absorption of pervious concrete, with a relative performance improvement of 95.7% between the optimal and poorest gradations. The constructed Stacking ensemble learning model achieved a coefficient of determination (R2) of 0.97, outperforming all individual models with minimal fluctuation. The proposed multi-objective optimization framework successfully resolved the intrinsic conflict between permeability, compressive strength, and sound absorption. The optimized mix proportion solution (O3) not only satisfied the standards for permeability and strength but also achieved superior sound absorption performance that surpassed all single-sized aggregate groups, with an error of only 8.9% between the model's prediction and the experimental value.
- Research Article
- 10.2174/0115734056420925251127020640
- Jun 10, 2026
- Current medical imaging
- V M Manesh + 3 more
Pancreatic Adenocarcinoma (PDAC) is the fourth leading cause of cancer-related mortality, predominantly affecting individuals over 45 years. Conventional diagnostic imaging modalities such as CT, MRI, PET, and ultrasound remain limited in detecting early-stage disease. Advances in deep learning (DL) hold promise for enhanced segmentation and early detection of pancreatic tumors. This study proposes an AI-based framework for the early detection of PDAC from CT images. Noise reduction was performed using a Boosted Adaptive Diffusion Filter (BADF). Tumor segmentation was achieved with a modified U-Net model, optimized using the Adam optimizer. Classification was conducted through an ensemble learning approach. The framework was validated on publicly available CT datasets, with performance assessed using standard evaluation metrics. The proposed model achieved 98.7% accuracy, 98.7% precision, 97.92% specificity, and 99.63% AUC in distinguishing pancreatic cancers, including tumors smaller than 2 cm. These results demonstrate superior performance compared to existing DL-based approaches. Integration of BADF preprocessing, U-Net-based segmentation, and ensemble classification enhanced model robustness and detection accuracy. The framework addresses challenges in small tumor detection, a critical factor in improving clinical outcomes. The proposed method demonstrates significant potential for early PDAC detection using CT images. By combining advanced preprocessing, segmentation, and ensemble learning, the framework enhances diagnostic accuracy and reliability, supporting clinical decision-making and contributing to improved patient care.
- Research Article
- 10.1016/j.metabol.2026.156600
- Jun 1, 2026
- Metabolism: clinical and experimental
- Bingtao Weng + 7 more
Identifying high-risk individuals for cardiovascular and all-cause mortality among individuals with cardiovascular-kidney-metabolic (CKM) syndrome stage 0-3 can guide the implementation of targeted interventions. This study aimed to evaluate the predictive value of plasma proteins for future cardiovascular and all-cause mortality. This study included 39,007 participants from the UK Biobank (UKB) with CKM stage 0-3 and available proteomic data. Associations between plasma proteins and future risks of cardiovascular and all-cause mortality were assessed using Cox proportional hazards models. Key proteins were identified through an ensemble machine learning approach integrating support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGBoost) algorithms. Subsequently, Cox models were applied to evaluate the incremental predictive value of these key proteins and their ability to enhance risk stratification for mortality outcomes. Furthermore, temporal trajectories of protein levels were examined in the years preceding death. During a median follow-up of 15.2years, 505 participants died from cardiovascular causes and 3368 from any cause. 56 and 269 out of 2911 plasma proteins were significantly associated with cardiovascular and all-cause mortality, respectively (Bonferroni-adjusted P<0.05). Incorporating seven and eight key proteins into conventional model significantly improved long-term predictive performance (C-statistics: 0.812 versus 0.782 for cardiovascular mortality; 0.772 versus 0.739 for all-cause mortality; both P<0.001), and also provided incremental predictive value for 5- and 10-year mortality risks. Notably, participants died during follow-up exhibited markedly elevated certain protein levels over a decade before deaths, with progressively increasing trajectories over time. Stratification based on optimal predicted risk thresholds further revealed distinct cumulative mortality risks across groups. In individuals with CKM stage 0-3, plasma proteins combined with traditional risk factors may predict future cardiovascular and all-cause mortality.
- Research Article
- 10.1016/j.biteb.2026.102688
- Jun 1, 2026
- Bioresource Technology Reports
- Chaeyeon Park + 5 more
Stress condition optimization for maximizing polyhydroxybutyrate accumulation in Cupriavidus necator via experimental and ensemble machine learning approaches
- Research Article
- 10.1016/j.health.2025.100444
- Jun 1, 2026
- Healthcare Analytics
- Zahra Gharibi
An ensemble learning approach for predicting hospital stay in transplant patients
- Research Article
- 10.1016/j.egyr.2026.109136
- Jun 1, 2026
- Energy Reports
- Sultan Mesfer Aldossary + 1 more
Accurate solar irradiance forecasting is critical for optimizing photovoltaic systems, enhancing grid stability, and supporting energy storage management, particularly in regions with high climatic variability like Saudi Arabia. However, existing models often lack either the precision to handle dynamic atmospheric conditions or the interpretability needed for operational trust. This study addresses these challenges by developing a modular solar forecasting dashboard that balances accuracy and transparency for five climatically diverse Saudi cities: Riyadh, Jeddah, Abha, Dammam, and Qassim. The proposed framework combines Random Forest regression with physics-informed feature engineering (e.g., hour-cloud interactions) and Shapley Additive Explanations to ensure interpretable predictions. Empirical results demonstrated robust performance, with determination coefficients exceeding 0.90 across all regions and root mean squared errors as low as 92.4 watts per square meter in Dammam. A residual-based alert system flagged significant deviations ( ± 100 watts per square meter), revealing Riyadh’s high alert rate (32.2%) due to dust-induced variability. Feature analysis confirmed the dominance of temporal-cloud interactions, accounting for 38%–42% of prediction variance. The study provides a scalable, interpretable solution for solar forecasting, aligning with Saudi Vision 2030’s renewable energy goals. Future work will integrate real-time satellite data and adaptive alert thresholds to further enhance operational utility. • Achieved coefficient of determination above 0.90 and error as low as 92 W/m 2 . • Feature analysis showed temporal-cloud interactions drive 38%–42% of variance. • Alerts identified high-variability days, with Riyadh showing a 32% alert rate. • Coastal cities (e.g., Dammam) had more stable forecasts than mountainous regions. • Physics-informed features improved accuracy by 12% over traditional inputs.
- Research Article
- 10.30574/ijsra.2026.19.2.1054
- May 31, 2026
- International Journal of Science and Research Archive
- Pranjali Tewari + 2 more
Cardiovascular diseases (CVDs) are also considered as some of the most common causes of death in the world and are usually diagnosed up to the late stages, thus early diagnosis is much important in preventing those diseases. The current paper introduces the SmartHeart AI, an intelligent multi-modal machine learning model that will support early cardiovascular disease prediction based on structured clinical and lifestyle data, and be extended to include unstructured medical data and physiological observations. I propose a framework that uses information processing, feature engineering, and the use of ensemble learning approaches in modeling complex interrelationship between risk factors. An approach using the Random Forest approach has a better result, recording a high accuracy of about 85% and a ROC-AUC value of 0.90 which represents a great ability to predict. Probabilistic risk estimation is also provided to present patient-specific risk scores which improves its applicability in clinical decision support. There is also multi-modal data fusion design to provide future capability of integrating clinical text, wearable sensor and imaging modalities. The model has been experimentally tested to be robust and generalize well and the ability to interpret the features is supported by feature importance analysis. The suggested system brings out the opportunities of smart, data-driven solutions in enhancing early detection and preventative health services of cardiovascular diseases.
- Research Article
- 10.1038/s41598-026-55482-0
- May 30, 2026
- Scientific reports
- Most Nusrat Jahan Resma + 4 more
Diabetes mellitus (DM) is an escalating global public health concern, with a rapidly increasing burden in low- and middle-income countries, including Bangladesh. Despite its growing prevalence and associated complications such as cardiovascular disease, kidney failure and stroke, comprehensive evidence on its determinants and predictive modeling at the population level remains limited. This study aimed to predict the DM and identify its associated risk factors using ensemble machine learning (EML) approaches among adults in northern Bangladesh. A community-based cross-sectional study was conducted among 1408 adults in Dinajpur district between March 25 and June 5, 2025, using structured and pilot-tested questionnaires administered through face-to-face interviews. Feature selection was performed using Recursive Feature Elimination, Random Forest importance and Best First Search methods. Six machine learning models were developed, followed by a stacking ensemble model to enhance predictive performance. Model evaluation was based on accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC). Model interpretability was assessed using SHAP analysis, and findings were validated using multivariable logistic regression. The prevalence of DM was 15.1% in the study population. Among individual models, LightGBM demonstrated the highest performance (accuracy: 89.44%; AUC: 0.958 [95% CI 0.945-0.973]), followed by XGBoost (accuracy: 88.69%; AUC: 0.955 [95% CI 0.945-0.972]). The stacking ensemble model outperformed all base learners, achieving an accuracy of 91.67% and an AUC of 0.967 (95% CI 0.957-0.981). SHAP analysis identified age, family history of diabetes, BMI, weight, dietary behaviors (particularly low vegetable intake and added salt/sugar), family income, and gender as key predictors. Multivariable logistic regression confirmed these findings, showing that advancing age especially 51-60years, female gender, family history of diabetes, hypertension, kidney disease and low vegetable consumption were independently associated with DM. Therefore, stacking-based ensemble learning significantly improves the predictive accuracy of DM while enabling robust identification of key risk factors. The consistency between machine learning and traditional statistical approaches strengthens the validity of the findings. These results highlight the importance of integrating advanced analytical methods into public health research to support early detection, targeted prevention, and evidence-based decision-making in resource-constrained settings such as northern Bangladesh.
- Research Article
- 10.1038/s41598-026-53037-x
- May 28, 2026
- Scientific Reports
- Andre Jatmiko Wijaya + 2 more
Horizontal gene transfer (HGT) is widely recognized as a major driver of antimicrobial resistance (AMR) dissemination, with genomic islands (GIs) as one of the drivers facilitating the spread. Detecting GIs is essential for improving AMR surveillance. Numerous computational approaches have been developed for GIs detection, including recent advances in machine learning (ML). Several studies in other fields have shown that ML model performance depends on data representations. Combining multiple data representations in ensemble learning has been shown to improve performance in other genomics tasks. However, this approach has not yet been evaluated for GIs detection. To this end, we investigate the efficacy of integrating diverse data representations in ensemble learning for GIs detection, particularly for classification task. Then, we assess its applicability to localizing GIs, which are clusters of genes acquired through HGT, in a genomic sequence. We implemented a two-stage ensemble selection strategy to determine the optimal combination of data representations. Our ensemble selection strategy reveals that combining low-correlated data representations in an ensemble classifier yields a slightly higher Recall than individual representation for the classification task, but the improvement is not statistically significant. Nevertheless, the ensemble classifier could not localize GIs better, suggesting that the cross-task generalizability remains constrained. This finding presents an opportunity for future research to advance the field by redefining the problem formulation of GIs detection.
- Research Article
- 10.1038/s41598-026-50480-8
- May 18, 2026
- Scientific reports
- Achin Jain + 10 more
This paper explores the use of optimized convolutional neural networks (CNNs) to classify diseases affecting potato leaves using TensorFlow-2. The dataset, sourced from Kaggle's Plant Village repository, includes 152 images of healthy potato leaves and 1000 images each of early and late blight. The methodology covers data preparation, model architecture design, training, evaluation, and deployment. During data preparation, the data set was split into training sets (80%) and testing sets (20%), with images resized to 128x128 pixels. The Deep Learning (DL) models built using CNN with 4 different optimizers (ADAM, SGD, RMSPROP, and ADAMAX) and trained using a sparse categorical cross-entropy loss function, include multiple convolutional and pooling layers for feature extraction, and fully connected layers for classification. Early stopping was used to prevent overfitting. Model performance was assessed using accuracy, loss curves, confusion matrix, ROC curve, precision recall curve, classification report, and F1 score. In addition, we have used data augmentation to balance the dataset by increasing healthy potato leaves 6 times and the use of Ensemble Deep Learning (EDL). EDL10 which contains DL1 (CNN + ADAM), DL2 (CNN + SGD), DL3 (CNN + RMSPROP) and DL4 (CNN + ADAMX) performs best with a accuracy score of 97.0%. This highlights the importance of data balancing and the use of the ensemble classification approach for the detection of blight in Potato Leaves.
- Research Article
- 10.34133/research.1283
- May 13, 2026
- Research
- Augustin Tillement + 13 more
The elemental composition of brains changes progressively with age, yet these metallome alterations remain largely unexplored as diagnostic biomarkers in neurological disease. Here, we present a comprehensive analysis of 24 inorganic elements in paired cerebrospinal fluid and serum samples from 1,608 individuals spanning healthy aging through 14 neurological conditions, representing the largest systematically standardized cohort for neurological metallomics. Uniquely, our unselected, consecutively admitted clinical cohort captures the full heterogeneity of neurological presentations, overcoming the limitations of traditional case–control designs focused on isolated disease entities. Machine learning analysis reveals that aging is associated with distinct cerebrospinal fluid elemental signatures independent of peripheral blood changes, primarily reflecting blood–brain barrier permeability alterations that correlate with established albumin quotient measurements. We identify 2 predominant patterns of neurological elemental dysregulation: one mainly consistent with passive barrier-mediated leakage in inflammatory conditions, and another mainly indicative of disease-intrinsic perturbations of metal homeostasis in neurodegenerative disorders. Age-stratified analysis reveals that elemental signatures evolve differently across the lifespan for distinct pathological processes. The integration of elemental signatures with routine clinical parameters through ensemble learning approaches enhances diagnostic accuracy across all tested neurological categories, establishing metallomics as a complementary biomarker class that captures orthogonal pathophysiological information. These findings establish brain metallomics as an emerging field where artificial intelligence reveals complex multi-element interactions present in neurological aging, opening new avenues for precision medicine in age-related neurological disorders.
- Research Article
- 10.1186/s12874-026-02858-5
- May 12, 2026
- BMC medical research methodology
- Tony Zbysinski + 8 more
Missing data is a challenge in clinical research, especially in real-world data (RWD), where complete case analysis can bias results and reduce power. Ensemble learning approaches like Super Learner (SL) show strong numerical performance for prediction problems, but their use for missing value imputation (MVI) in oncology datasets is unexplored. We sought to develop and evaluate a novel SL-based imputation function that can impute multiple variables and quantify observation-specific uncertainty. We analyzed two independent cohorts of acute myeloid leukemia patients (n = 1641), 546 patients from the University of Colorado and 1095 from an external real-world cohort. The SL-based MVI function includes data processing, predictor selection, binary and continuous variable pipelines, and automatic performance measurement. Ensembles for both binary and continuous variables integrate diverse base learners, such as generalized linear models, random forests, and neural networks, via a meta learner that optimizes predictive accuracy. The binary variable pipeline's SL ensemble was optimized using area under the curve (AUC), while the continuous variable pipeline's SL ensemble was optimized via non-negative least squares. Performance was compared to multiple imputation by chained equations (MICE) using balanced accuracy, F1-score, root mean square error (RMSE), and visualizations. Observation-specific uncertainty was quantified for all imputations of both binary and continuous variables, with both additionally having lower and upper resampling-based potential imputation values. The SL cross-validation loop, SL ensemble trained for imputation, and resampling all supported parallelization. Clinically significant features of the cohorts were selected a priori based on prior literature. In a numerical experiment with 9 clinically important binary features, the proposed MVI function imputed and achieved higher balanced accuracy than MICE for 7/9 variables (mean balanced accuracy 89.04% vs. 80.75%) with comparable performance for the other 2 variables. The continuous variable SL ensemble, across 4 variables, showed an average 24.45% lower RMSE than MICE. On average, the SL ensemble trained for prediction took 145.02s to process for binary targets. This study demonstrates that the SL-based imputation function has improved performance over MICE in high-dimensional RWD while providing novel, observation-level uncertainty quantification.
- Research Article
- 10.1038/s41598-026-46764-8
- May 12, 2026
- Scientific Reports
- Golshid Fathi + 6 more
Accurate identification of Allium seed genotypes is essential for cultivar authentication, breeding, and fraud prevention, yet remains challenging due to morphological similarities. This study evaluates the potential of a visible and near-infrared (Vis–NIR) spectrometer and a hyperspectral camera for non-destructive classification of seven closely related Allium genotypes, including shallot, red, white, and yellow onions, bon-sorkh, and two leek varieties. A total of 700 spectra and 70 images were acquired using the Vis–NIR spectrometer and hyperspectral camera, respectively, under controlled conditions and spectral preprocessing was applied to enhance signal quality. For spectrometer data, classification models were developed using soft independent modelling of class analogy (SIMCA), artificial neural networks (ANN), and histogram-based gradient boosting (HisGB). For hyperspectral data, pixel-level spectra were used to train ANN, HisGB, and deep convolutional neural networks (1D and 2D CNNs). Among the spectrometer models, the combination of second derivative preprocessing with HisGB achieved the highest performance (F1-score: 98.52%). For HSI, HisGB yielded the highest pixel-level classification accuracy (F1-score: 97.83%; error: 2.49%), followed by 1D CNN (F1-score: 96.85%). Spatial analysis revealed that HisGB and 1D CNN produced consistent classification maps across genotypes, whereas ANN and 2D CNN exhibited higher misclassification rates, particularly for morphologically similar classes such as shallot and bon-sorkh. At image level, the hyperspectral camera outperformed the Vis–NIR spectrometer, achieving perfect classification across all models. These results demonstrate the potential of hyperspectral imaging, especially when combined with ensemble and deep learning approaches, for high-throughput, non-destructive seed sorting and genotype purity assessment. The study also emphasizes the trade-off between the lower cost but reduced precision of the Vis–NIR spectrometer and the superior accuracy offered by the hyperspectral camera.
- Research Article
- 10.1186/s40562-026-00481-2
- May 5, 2026
- Geoscience Letters
- Kieu Anh Nguyen + 2 more
cyclones act as dominant triggers due to the combination of intense rainfall, strong winds, and steep terrain (Jones et al. 2022) . Countries across Southeast Asia, East Asia, and the Pacific Islands are particularly vulnerable, with typhoon-induced landslides regularly resulting in severe disasters (Chiang and Chang 2011). Taiwan represents one of the most exposed regions globally; its steep mountains, fragile geology, and frequent typhoon landfalls create highly landslide-prone conditions (Wu and Lin 2021; Lo et al. 2014) . Over the past decades, numerous typhoon events in Taiwan have triggered widespread slope failures, highlighting the urgent need for accurate