Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Machine learning or morphometric scaling? A systematic review of methodological confounds and the generalizability of sex classification in neuroimaging

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Background: This systematic review critically evaluates whether machine learning (ML) identifies biologically meaningful sex-related brain architecture or merely exploits methodological artifacts and allometric scaling. While ML models achieve high classification accuracies, it remains unclear if these reflect stable, mechanistically informative dimorphism or are driven by confounds such as total intracranial volume (TIV) and site-specific noise. We examine how imaging modalities, algorithms, and population strata influence both classification outcomes and biological interpretability. Methods: Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, we searched Web of Science, PubMed, and Scopus through January 2024. Included studies [healthy humans, 3T magnetic resonance imaging (MRI), ML-based sex classification] were assessed for risk of bias, focusing on data leakage, validation strategies, and confound management. Results: Thirty-five studies (n > 110,000) were included. While reported accuracies reached 98.06% for T1-weighted MRI, 96.0% for diffusion MRI (dMRI), and 94.72% for functional MRI (fMRI), performance was highly dependent on population characterization and age. Deep learning consistently outperformed traditional ML (TML) but showed high sensitivity to methodological artifacts. Notably, studies failing to correct for TIV reported potentially inflated accuracies, suggesting that many models identify physical scale rather than intrinsic neuroanatomical dimorphism. Discussion: High classification accuracies are often bolstered by methodological confounds and a lack of cross-site validation. There is a significant discrepancy between ML-driven predictive power and biological inference validity. Current pipelines do not yet allow for robust, generalizable inference about brain sex. To move beyond statistical separation toward mechanistic understanding, the field must prioritize TIV-corrected benchmarks and diverse non-WEIRD (Western, Educated, Industrialized, Rich, Democratic) datasets. We conclude that while ML is a powerful pattern detector, its results must be interpreted with caution regarding biological dimorphism.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 17
  • 10.3389/fphys.2023.1281506
MRI radiomics-based decision support tool for a personalized classification of cervical disc degeneration: a two-center study
  • Jan 3, 2024
  • Frontiers in Physiology
  • Jun Xie + 11 more

Objectives: To develop and validate an MRI radiomics-based decision support tool for the automated grading of cervical disc degeneration.Methods: The retrospective study included 2,610 cervical disc samples of 435 patients from two hospitals. The cervical magnetic resonance imaging (MRI) analysis of patients confirmed cervical disc degeneration grades using the Pfirrmann grading system. A training set (1,830 samples of 305 patients) and an independent test set (780 samples of 130 patients) were divided for the construction and validation of the machine learning model, respectively. We provided a fine-tuned MedSAM model for automated cervical disc segmentation. Then, we extracted 924 radiomic features from each segmented disc in T1 and T2 MRI modalities. All features were processed and selected using minimum redundancy maximum relevance (mRMR) and multiple machine learning algorithms. Meanwhile, the radiomics models of various machine learning algorithms and MRI images were constructed and compared. Finally, the combined radiomics model was constructed in the training set and validated in the test set. Radiomic feature mapping was provided for auxiliary diagnosis.Results: Of the 2,610 cervical disc samples, 794 (30.4%) were classified as low grade and 1,816 (69.6%) were classified as high grade. The fine-tuned MedSAM model achieved good segmentation performance, with the mean Dice coefficient of 0.93. Higher-order texture features contributed to the dominant force in the diagnostic task (80%). Among various machine learning models, random forest performed better than the other algorithms (p < 0.01), and the T2 MRI radiomics model showed better results than T1 MRI in the diagnostic performance (p < 0.05). The final combined radiomics model had an area under the receiver operating characteristic curve (AUC) of 0.95, an accuracy of 89.51%, a precision of 87.07%, a recall of 98.83%, and an F1 score of 0.93 in the test set, which were all better than those of other models (p < 0.05).Conclusion: The radiomics-based decision support tool using T1 and T2 MRI modalities can be used for cervical disc degeneration grading, facilitating individualized management.

  • Research Article
  • Cite Count Icon 65
  • 10.1016/j.brs.2019.06.015
Real-time estimation of electric fields induced by transcranial magnetic stimulation with deep neural networks
  • Jun 17, 2019
  • Brain Stimulation
  • Tatsuya Yokota + 7 more

Real-time estimation of electric fields induced by transcranial magnetic stimulation with deep neural networks

  • Research Article
  • Cite Count Icon 2
  • 10.2214/ajr.25.33759
Artificial Intelligence for CT and MRI Protocoling: A Meta-Analysis of Traditional Machine Learning, BERT, and Large Language Models.
  • Oct 29, 2025
  • AJR. American journal of roentgenology
  • Ethan Sacoransky + 2 more

BACKGROUND. Examination protocoling is a resource-intensive task. Various artificial intelligence (AI) approaches have been investigated to automate this process. OBJECTIVE. The purpose of this study was to evaluate performance of traditional machine learning (ML) models, bidirectional encoder representations from transformers (BERT) models, and large language models (LLMs) for automated CT and MRI protocoling. EVIDENCE ACQUISITION. MEDLINE, Embase, Scopus, Web of Science, IEEE Xplore, and Google Scholar databases were searched through July 2025 for studies reporting the performance of an AI-based technique in assigning protocols for CT or MRI requisitions. Accuracy results were separately extracted for all models tested in each study and pooled using a random-effects meta-analysis. AI approaches were compared using Welch t tests. Common sources of error were qualitatively summarized. EVIDENCE SYNTHESIS. The final analysis included 23 studies, comprising 1,196,259 imaging requisitions. Requisition subspecialties included body imaging (n = 4), musculoskeletal imaging (n = 3), neuroradiology (n = 6), thoracic imaging (n = 1), and multiple subspecialties (n = 9). Sixteen studies evaluated traditional ML models, eight evaluated BERT models, and five evaluated LLMs. Task-specific model fine-tuning was performed in three studies for traditional ML models, all studies for BERT models, and one study for LLMs. The overall pooled protocoling accuracy was 85% (95% CI, 83-87%). The pooled accuracy was 83% (95% CI, 80-85%) for traditional ML models, 87% (95% CI, 85-89%) for BERT models, and 86% (95% CI, 83-89%) for LLMs; these pooled accuracies were not significantly different between any pairwise combination of the three AI approaches (all p > .05). Among 30 distinct models (14 traditional ML models, nine BERT models, seven LLMs), the top-10 performing models comprised two traditional ML models, six BERT models (including the top performing model [BioBERT, a biomedical-domain BERT; accuracy, 93%]), and two LLMs. Common sources of error included ambiguous requisition text, data imbalance yielding incorrect protocol assignments for low-volume protocols, the presence of multiple clinically reasonable protocols for given requisitions, and difficulty handling requisitions containing terms strongly associated with disparate protocols. CONCLUSION. The top-performing AI models for automated CT and MRI protocoling included predominantly fine-tuned BERT models. CLINICAL IMPACT. AI tools show strong potential to help streamline radiologist workflows, possibly through hybrid AI-radiologist approaches. Fine-tuned LLMs warrant further exploration. TRIAL REGISTRATION. PROSPERO identifier CRD420251088671.

  • Research Article
  • Cite Count Icon 9
  • 10.1186/s40537-024-00998-3
A systematic literature review of neuroimaging coupled with machine learning approaches for diagnosis of attention deficit hyperactivity disorder
  • Oct 4, 2024
  • Journal of Big Data
  • Imran Ashraf + 3 more

ProblemAttention deficit hyperactivity disorder (ADHD) is the most commonly found neurodevelopmental condition among children with an estimated 2.5% to 9% global prevalence. While ADHD has been regarded as a lifelong condition, its early diagnosis elevates the probability of recovery and normal life for children. Despite clinical diagnosis being the primary one, substantial developments in ADHD diagnosis have been made during the past decade.AimSeveral imaging-based technologies and approaches have been presented in the existing literature including magnetic resonance imaging (MRI) and functional MRI (fMRI). In addition, the deployment of machine learning and deep learning models has paved the way for automated diagnosis of ADHD which increases both accuracy and robustness of ADHD detection. A comprehensive and systematic literature review (SLR) of imaging technologies and machine learning approaches is highly desired to comprehend the current status of such approaches concerning their potential and challenges to outline future directions. Although a substantial body of literature exists on imaging-based ADHD diagnosis, comprehensive SRL on such approaches is scarce. This SLR aims to provide a comprehensive overview of imaging-based ADHD diagnosis with emphasis on machine learning approaches, reveals their pros and cons, and provides potential future research directions thereby contributing to the scientific community to accelerate further research for ADHD diagnosis.MethodsThis SRL focuses on analyzing recently published studies between 2010 and 2023. For this purpose, preferred reporting items for systematic review and meta-analyses (PRISMA) approach is performed in this study. Five eminent academic databases Web of Science, ACM, Springerlink, Elsevier, and PubMed are selected for article search. The SLR follows a systematic methodology comprising article search, selection based on inclusion and exclusion criteria, and rigorous assessment for categorization.ResultsIt is found that MRI and fMRI are the dominant approaches integrated with machine learning models for ADHD detection and function near-infrared spectroscopy is also adopted by a few studies. Predominantly, the ATHENA preprocessing approach is used to preprocess MRI data before model training. Due to the public availability of the ADHD-200 dataset, it is widely used in the existing literature while a few studies utilized their self-collected datasets. Machine learning models are the choice of the majority of the studies, particularly, the support vector machines model has been widely used for ADHD detection. Feature fusion is observed to be a better choice for obtaining more accurate results.ConclusionMachine learning and deep learning models provide automated ADHD detection with better accuracy and robustness, however, such models are not generalizable and their performance varies concerning the locality of data used for experiments. In addition, the heterogeneity of data collection from various devices is a challenge, and the use of a standard device may provide better solutions. The lack of labeled data also adds difficulties to training models. Besides the use of MRI and fMRI, other novel technologies should be explored for better ADHD detection performance.

  • Research Article
  • 10.1177/2325967124s00238
Poster 271: Improving Non-Invasive Diagnosis and Grading of Cartilage Defects in the Knee - Accuracy of Ultra High Field 7-Tesla MRI as Compared with Arthroscopy
  • Jul 1, 2024
  • Orthopaedic Journal of Sports Medicine
  • Andrew George + 4 more

Objectives: Osteoarthritis (OA) is a joint disease characterized by degradation of articular cartilage and subchondral bone. Cartilage injury plays an important role in the pathogenesis of osteoarthritis, and disease modifying treatments aiming to preserve and possibly regenerate cartilage are an area of active interest. Early diagnosis is crucial, as the disease may be more responsive to treatment at an early stage, and there is a critical unmet need for improved non-invasive morphologic cartilage imaging techniques. Standard clinical MRI is performed on a 1.5T or 3T machine. Ultra-high field (UHF) 7-Tesla (7T) magnetic resonance imaging (MRI) is a new technology that offers 2.3x improved signal-to noise ratio (SNR) compared to 3T MRI and 2.8x SNR compared to 1.5T MRI. The purpose of this study was to estimate the accuracy of 7T MRI for the detection and grading of cartilage lesions in the knee. The hypothesis was that 7T images would offer better detection and more accurate grading of chondral lesions than standard clinical grade MRI scans. Methods: In this prospective, paired, blinded study, patients who had undergone a 1.5 or 3T standard of care (SOC) MRI and were scheduled for knee arthroscopy were enrolled (10/2019 to 08/2021) and a study intervention 7T MRI was performed prior to surgery. The study was approved by our institutional review board, and it was funded by the Methodist-Siemens Collaboration. Exclusion criteria included contraindication to MRI (e.g., pacemaker), weight less than 30 kg, and pregnancy. Scans were reviewed by three independent fellowship-trained radiologists who were blinded to clinical and arthroscopic data. Scans were reviewed over 5 days, with a 4 week “washout” period between review of SOC and 7T scans to limit recall bias. Readers were blinded to clinical and arthroscopic data and graded all six articular surfaces of the knee (patella, trochlea, lateral condyle, medial condyle, lateral tibia, medial tibia) using a modified Outerbridge classification system. At the time of arthroscopy, each articular surface was graded by the operating surgeon according to a modified Outerbridge system similar to above, with the exception that Grade 1 = softening of the articular cartilage. The surgeon was blinded to 7T images at the time of arthroscopy, although they had access to the SOC MRI, which was necessary for patient care. Using arthroscopy as the gold standard, we calculated sensitivity and specificity of SOC MRI and 7T MRI. An Outerbridge grade of 0 was classified as negative, while a grade of 1-4 was classified as positive. A secondary analysis of sensitivity and specificity was performed on a per articular surface basis, with 6 articular surfaces per patient after correction for within subject clustered data. A Mann-Whitney U test was used to compare diagnostic scores and ratings between instruments (SOC vs. 7T). Coefficients of variation between observers within each instrument were compared for each variable using a paired sample t-test. Type-I error was set at alpha=0.05 Results: A total of 100 patients (age: 43 ± 14, 54 female, 46 male) were enrolled. 7T MRI resulted in improved sharpness (defined by visibility of nerve fascicles) and shading (based on artifacts) compared to SOC MRI (Figure 1, p &lt; 0.001 for both). There was improved contrast between fluid and cartilage with 7T MRI (using confidence rating at axial mid-patella, p = 0.003). 7T MRI had a higher sensitivity in detecting cartilage lesions for five of the six articular surfaces when using the arthroscopic gold standard, but with lower specificity for all surfaces (Figure 2). Finally, there was improved inter-observer reliability in detecting and grading cartilage defects with 7T compared to SOC MRI (p &lt; 0.05). Conclusion: Based on our findings thus far, 7T MRI appears to result in improved measurement ratings, sensitivity, and inter-observer reliability compared to SOC in detecting cartilage lesions in the knee. Further work is needed to evaluate cost effectiveness of 7T MRI.

  • Research Article
  • Cite Count Icon 76
  • 10.1093/neuros/nyab103
Machine Learning for the Prediction of Molecular Markers in Glioma on Magnetic Resonance Imaging: A Systematic Review and Meta-Analysis.
  • Jul 1, 2021
  • Neurosurgery
  • Anne Jian + 5 more

Molecular characterization of glioma has implications for prognosis, treatment planning, and prediction of treatment response. Current histopathology is limited by intratumoral heterogeneity and variability in detection methods. Advances in computational techniques have led to interest in mining quantitative imaging features to noninvasively detect genetic mutations. To evaluate the diagnostic accuracy of machine learning (ML) models in molecular subtyping gliomas on preoperative magnetic resonance imaging (MRI). A systematic search was performed following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analysis) guidelines to identify studies up to April 1, 2020. Methodological quality of studies was assessed using the Quality Assessment for Diagnostic Accuracy Studies (QUADAS)-2. Diagnostic performance estimates were obtained using a bivariate model and heterogeneity was explored using metaregression. Forty-four original articles were included. The pooled sensitivity and specificity for predicting isocitrate dehydrogenase (IDH) mutation in training datasets were 0.88 (95% CI 0.83-0.91) and 0.86 (95% CI 0.79-0.91), respectively, and 0.83 to 0.85 in validation sets. Use of data augmentation and MRI sequence type were weakly associated with heterogeneity. Both O6-methylguanine-DNA methyltransferase (MGMT) gene promoter methylation and 1p/19q codeletion could be predicted with a pooled sensitivity and specificity between 0.76 and 0.83 in training datasets. ML application to preoperative MRI demonstrated promising results for predicting IDH mutation, MGMT methylation, and 1p/19q codeletion in glioma. Optimized ML models could lead to a noninvasive, objective tool that captures molecular information important for clinical decision making. Future studies should use multicenter data, external validation and investigate clinical feasibility of ML models.

  • Research Article
  • 10.25147/ijcsr.2017.001.1.224
Empirical Analysis of the State-of-the-Art Models for Handling Polarity Shifts Due to Implicit Negation in Mobile Phone Reviews
  • Jan 1, 2025
  • International Journal of Computing Sciences Research
  • Millicent Murithi + 2 more

Purpose–This paper presents a comprehensive empirical analysis focusing on sentiment flux within state-of-the-art models designed for handling polarity shifts due to implicit negation in Amazon mobile phones' reviews. Method–The research evaluates diverse models across five categories: traditional machine learning (ML), deep learning (DL), and hybrid models combining both approaches. Various feature extraction, feature selection, and data augmentation techniques are tested on Amazon mobile phone reviews dataset. BERT and LSTM are used for deep learning while SVM and Naive Bayes are used for traditional ML. ANOVA is used to identify the presence or absence of significant differences and interactions among these entities. Results –DL shows superior performance compared to traditional ML models. ANOVA analysis shows significant performance differences between conventional ML and DL models. Traditional ML models interact significantly with feature extraction and selection techniques while DL models do not. Traditional ML models do not interact significantly with data augmentation methods while DL models do. FastText extraction outperforms word2vec; Back translation outperforms synonym replacement while recursive feature selection (RFE) surpasses TF-IDF (Term Frequency-Inverse Document Frequency). The BERT and LSTM exhibit one of the strongest performances. Conclusion –The study concludes that DL models are more effective. Data augmentation techniques significantly impact the performance of DL models, with back translation showing superior performance over synonym replacement. This provides a leverage point in developing an improved model in the future. Recommendations –Future research should focus on developing a hybrid model for Enhanced Polarity Shift Management of Mobile Phone Reviews using Contextual Back Translation Augmented by Seq2seq Perturbations. This aims at leveraging contextual back translation and Seq2seq perturbations to generate a diverse interpretation that consequently improves the model's ability to handle nuanced expressions of sentiments due to implicit negation with enhanced accuracy, generalizability, robustness to polarity shifts, and contextual understanding. Research Implications –The findings provide valuable insights into the development of state-of-the-art models, offering a promising direction for further research in sentiment analysis. Keywords –empirical analysis, hybrid, perturbations, implicit negation, sentiment flux

  • Research Article
  • Cite Count Icon 5
  • 10.1097/ju.0000000000003224.05
MP09-05 AUTOMATED PROSTATE GLAND AND PROSTATE ZONES SEGMENTATION USING A NOVEL MRI-BASED MACHINE LEARNING FRAMEWORK AND CREATION OF SOFTWARE INTERFACE FOR USERS ANNOTATION
  • Apr 1, 2023
  • Journal of Urology
  • Masatomo Kaneko + 17 more

MP09-05 AUTOMATED PROSTATE GLAND AND PROSTATE ZONES SEGMENTATION USING A NOVEL MRI-BASED MACHINE LEARNING FRAMEWORK AND CREATION OF SOFTWARE INTERFACE FOR USERS ANNOTATION

  • Research Article
  • Cite Count Icon 28
  • 10.1097/bpo.0b013e3182619181
Accuracy of 3-Tesla Magnetic Resonance Imaging for the Diagnosis of Intra-articular Knee Injuries in Children and Teenagers
  • Dec 1, 2012
  • Journal of Pediatric Orthopaedics
  • David L Schub + 5 more

Magnetic resonance imaging (MRI) is a commonly used tool for the diagnosis of intra-articular knee pathologies. Although many studies have reported the accuracy of MRI in the adult population, fewer studies have investigated these tests in younger patients. Furthermore, these studies have shown a higher variability in both the sensitivity and the specificity of MRI for these knee injuries in this age group. Advancements in MRI technology, such as the 3-Tesla (3T) MRI magnet, have shown promising results for musculoskeletal injury diagnosis in adults. This study aims to evaluate 3 T MRI for the diagnosis of intra-articular knee pathologies in a pediatric and adolescent patient population. The records of 116 patients (119 knees) under the age of 20 years who underwent 3 T MRI studies of the knee and subsequent knee arthroscopy were reviewed retrospectively. The MRI report from the musculoskeletal radiology staff, the interpretation from the staff orthopedic surgeon, and the operative note dictations were compared, with a focus on meniscus and anterior cruciate ligament (ACL) pathologies. Seventeen orthopedic staff reads were not obtainable. Arthroscopy was used as the gold standard for diagnosis. The average age at MRI exam was 16.0 years and at surgery was 16.2 years. Using the musculoskeletal radiologist interpretation, the sensitivity and the specificity of 3 T MRI were 81.0% and 90.9% for medial meniscus injuries, 68.8% and 93% for lateral meniscus injuries, and 97.9% and 98.6% for ACL injuries, respectively. The orthopedic surgeon's interpretation of 3 T MRI had a sensitivity and specificity of 75.7% and 92.4% for medial meniscus injuries, 69.8% and 98.3% for lateral meniscus injuries, and 100% and 98.6% for ACL injuries, respectively. Posterior horn tears had the greatest discrepancies. When performed on pediatric and adolescent patients, newer 3 T MRI studies have excellent accuracy for diagnosing ACL tears. These studies also show a higher accuracy for the diagnosis of medial meniscal tears than lateral meniscal tears. Diagnostic study--Level 2.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.3389/fninf.2023.1266713
OFVSD: a Python package of optimized forward variable selection decoder for high-dimensional neuroimaging data
  • Sep 26, 2023
  • Frontiers in Neuroinformatics
  • Tung Dang + 2 more

The complexity and high dimensionality of neuroimaging data pose problems for decoding information with machine learning (ML) models because the number of features is often much larger than the number of observations. Feature selection is one of the crucial steps for determining meaningful target features in decoding; however, optimizing the feature selection from such high-dimensional neuroimaging data has been challenging using conventional ML models. Here, we introduce an efficient and high-performance decoding package incorporating a forward variable selection (FVS) algorithm and hyper-parameter optimization that automatically identifies the best feature pairs for both classification and regression models, where a total of 18 ML models are implemented by default. First, the FVS algorithm evaluates the goodness-of-fit across different models using the k-fold cross-validation step that identifies the best subset of features based on a predefined criterion for each model. Next, the hyperparameters of each ML model are optimized at each forward iteration. Final outputs highlight an optimized number of selected features (brain regions of interest) for each model with its accuracy. Furthermore, the toolbox can be executed in a parallel environment for efficient computation on a typical personal computer. With the optimized forward variable selection decoder (oFVSD) pipeline, we verified the effectiveness of decoding sex classification and age range regression on 1,113 structural magnetic resonance imaging (MRI) datasets. Compared to ML models without the FVS algorithm and with the Boruta algorithm as a variable selection counterpart, we demonstrate that the oFVSD significantly outperformed across all of the ML models over the counterpart models without FVS (approximately 0.20 increase in correlation coefficient, r, with regression models and 8% increase in classification models on average) and with Boruta variable selection algorithm (approximately 0.07 improvement in regression and 4% in classification models). Furthermore, we confirmed the use of parallel computation considerably reduced the computational burden for the high-dimensional MRI data. Altogether, the oFVSD toolbox efficiently and effectively improves the performance of both classification and regression ML models, providing a use case example on MRI datasets. With its flexibility, oFVSD has the potential for many other modalities in neuroimaging. This open-source and freely available Python package makes it a valuable toolbox for research communities seeking improved decoding accuracy.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 22
  • 10.3389/frai.2024.1365777
Application of machine learning for lung cancer survival prognostication-A systematic review and meta-analysis.
  • Apr 5, 2024
  • Frontiers in Artificial Intelligence
  • Alexander J Didier + 5 more

Machine learning (ML) techniques have gained increasing attention in the field of healthcare, including predicting outcomes in patients with lung cancer. ML has the potential to enhance prognostication in lung cancer patients and improve clinical decision-making. In this systematic review and meta-analysis, we aimed to evaluate the performance of ML models compared to logistic regression (LR) models in predicting overall survival in patients with lung cancer. We followed the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) statement. A comprehensive search was conducted in Medline, Embase, and Cochrane databases using a predefined search query. Two independent reviewers screened abstracts and conflicts were resolved by a third reviewer. Inclusion and exclusion criteria were applied to select eligible studies. Risk of bias assessment was performed using predefined criteria. Data extraction was conducted using the Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modeling Studies (CHARMS) checklist. Meta-analytic analysis was performed to compare the discriminative ability of ML and LR models. The literature search resulted in 3,635 studies, and 12 studies with a total of 211,068 patients were included in the analysis. Six studies reported confidence intervals and were included in the meta-analysis. The performance of ML models varied across studies, with C-statistics ranging from 0.60 to 0.85. The pooled analysis showed that ML models had higher discriminative ability compared to LR models, with a weighted average C-statistic of 0.78 for ML models compared to 0.70 for LR models. Machine learning models show promise in predicting overall survival in patients with lung cancer, with superior discriminative ability compared to logistic regression models. However, further validation and standardization of ML models are needed before their widespread implementation in clinical practice. Future research should focus on addressing the limitations of the current literature, such as potential bias and heterogeneity among studies, to improve the accuracy and generalizability of ML models for predicting outcomes in patients with lung cancer. Further research and development of ML models in this field may lead to improved patient outcomes and personalized treatment strategies.

  • Research Article
  • Cite Count Icon 1
  • 10.1089/neur.2023.0001
7T MRI Versus 3T MRI of the Brain in Professional Fighters and Patients With Head Trauma.
  • Mar 1, 2023
  • Neurotrauma Reports
  • Jonathan K Lee + 7 more

Many studies have investigated the imaging sequelae of repetitive head trauma with mixed results, particularly with regard to the detection of intracranial white matter changes (WMCs) and cerebral microhemorrhages (CMHs) on ≤3 Tesla (T) field magnetic resonance imaging (MRI). 7T MRI, which has recently been approved for clinical use, is more sensitive at detecting lesions associated with multiple neurological diagnoses. In this study, we sought to determine whether 7T MRI would detect more WMCs and CMHs than 3T MRI in 19 professional fighters, 16 patients with single TBI, versus 82 normal healthy controls (NHCs). Fighters and patients with TBI underwent both 3T and 7T MRI; NHCs underwent either 3T (n = 61) or 7T (n = 21) MRI. Readers agreed on the presence/absence of WMCs in 88% (84 of 95) of 3T MRI studies (Cohen's kappa, 0.76) and in 93% (51 of 55) of 7T MRI studies (Cohen's kappa, 0.79). Readers agreed on the presence/absence of CMHs in 96% (91 of 95) of 3T MRI studies (Cohen's kappa, 0.76) and in 96% (54 of 56) of 7T MRI studies (Cohen's kappa, 0.88). The number of WMCs detected was greater in fighters and patients with TBI than NHCs at both 3T and 7T. Moreover, the number of WMCs was greater at 7T than at 3T for fighters, patients with TBI, and NHCs. There was no difference in the number of CMHs detected with 7T MRI versus 3T MRI or in the number of CMHs observed in fighters/patients with TBI versus NHCs. These initial findings suggest that fighters and patients with TBI may have more WMCs than NHCs and that the improved voxel size and signal-to-noise ratio at 7T may help to detect these changes. As 7T MRI becomes more prevalent clinically, larger patient populations should be studied to determine the cause of these WMCs.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 14
  • 10.7759/cureus.38373
Harnessing Machine Learning in Early COVID-19 Detection and Prognosis: A Comprehensive Systematic Review.
  • May 1, 2023
  • Cureus
  • Rufaidah Dabbagh + 11 more

During the early phase of the COVID-19 pandemic, reverse transcriptase-polymerase chain reaction (RT-PCR) testing faced limitations, prompting the exploration of machine learning (ML) alternatives for diagnosis and prognosis.Providing a comprehensive appraisal of such decision support systems and their use inCOVID-19 management can aid the medical community in making informed decisions during the risk assessment of their patients, especially in low-resource settings. Therefore, the objective of this study was to systematically review the studies that predicted the diagnosis of COVID-19 or the severity of the disease using ML. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA), we conducted a literature search of MEDLINE (OVID), Scopus, EMBASE, and IEEE Xplorefrom January 1 to June 31, 2020. The outcomes were COVID-19 diagnosis or prognostic measures such as death, need for mechanical ventilation, admission, and acute respiratory distress syndrome. We included peer-reviewed observational studies, clinical trials, research letters, case series, and reports. We extracted data about the study's country, setting, sample size, data source, dataset, diagnostic or prognostic outcomes, prediction measures, type of ML model, and measures of diagnostic accuracy. Bias was assessed using the Prediction model Risk Of Bias ASsessmentTool (PROBAST). This study was registered in the International Prospective Register of Systematic Reviews (PROSPERO), with the number CRD42020197109. The final records included for data extraction were 66. Forty-three (64%) studies used secondary data. The majority of studies were from Chinese authors (30%).Most of the literature (79%) relied on chest imaging for prediction, while the remainder used various laboratory indicators, including hematological, biochemical, and immunological markers. Thirteen studies explored predicting COVID-19 severity, while the rest predicted diagnosis. Seventy percent of the articles used deep learning models, while 30% used traditional ML algorithms. Most studies reported high sensitivity, specificity, and accuracy for the ML models (exceeding 90%). The overall concern about the risk of bias was "unclear" in 56% of the studies. This was mainly due to concerns about selection bias. ML may help identify COVID-19 patients in the early phase of the pandemic, particularly in the context of chest imaging. Although these studies reflect that these ML models exhibit high accuracy, the novelty of these models and the biases in dataset selection make using them as a replacement for the clinicians' cognitive decision-making questionable.Continued research is needed to enhance the robustness and reliability of ML systems in COVID-19 diagnosis and prognosis.

  • Research Article
  • Cite Count Icon 10
  • 10.12688/f1000research.154680.2
Role of Artificial intelligence model in prediction of low back pain using T2 weighted MRI of Lumbar spine.
  • Oct 10, 2024
  • F1000Research
  • Ali Muhaimil + 8 more

Low back pain (LBP), the primary cause of disability, is the most common musculoskeletal disorder globally and the primary cause of disability. Magnetic resonance imaging (MRI) studies are inconclusive and less sensitive for identifying and classifying patients with LBP. Hence, this study aimed to investigate the role of artificial intelligence (AI) models in the prediction of LBP using T2 weighted MRI image of the lumbar spine. This was a prospective case-control study. A total of 200 MRI patients (100 cases and controls each) referred for lumbar spine and whole spine screening were included. The scans were performed using 3.0 Tesla MRI (United Imaging Healthcare). T2 weighted images of the lumbar spine were segmented to extract radiomic features. Machine learning (ML) models, such as random forest, decision tree, logistic regression, K-nearest neighbors, adaboost, and deep learning methods (DL), such as ResNet and GoogleNet, were used, and performance measures were calculated. Our study showed that Random forest and AdaBoost are the most reliable ML models for predicting LBP. Random forest showed high performance with area under curve (AUC) values from 0.83 to 0.88 across all lumbar vertebrae and L2-L3, L3-L4, and L4-L5 intervertebral discs (IVDs), with AUCs of 0.88 the highest at L5-S1 IVD (0.92). Adaboost demonstrated high performance at the L2-L5 vertebrae with AUC values of 0.82 to 0.90, with the highest AUC (0.97) at the L5-S1 IVD. Among the DL models, GoogleNet outperformed the other models at 30 epochs with an accuracy of 0.85, followed by ResNet 18 (30 epochs) with an accuracy of 0.84. The study demonstrated that ML and DL models can effectively predict LBP from MRI T2 weighted image of the lumbar spine. ML and DL models could also enhance the diagnostic accuracy of LBP, potentially leading to better patient management and outcomes.

  • Research Article
  • Cite Count Icon 2
  • 10.12688/f1000research.154680.1
Role of Artificial intelligence model in prediction of low back pain using T2 weighted MRI of Lumbar spine
  • Sep 10, 2024
  • F1000Research
  • Ali Muhaimil + 8 more

Background Low back pain (LBP), the primary cause of disability, is the most common musculoskeletal disorder globally and the primary cause of disability. Magnetic resonance imaging (MRI) studies are inconclusive and less sensitive for identifying and classifying patients with LBP. Hence, this study aimed to investigate the role of artificial intelligence (AI) models in the prediction of LBP using T2 weighted MRI image of the lumbar spine. Methods This was a prospective case-control study. A total of 200 MRI patients (100 cases and controls each) referred for lumbar spine and whole spine screening were included. The scans were performed using 3.0 Tesla MRI (United Imaging Healthcare). T2 weighted images of the lumbar spine were segmented to extract radiomic features. Machine learning (ML) models, such as random forest, decision tree, logistic regression, K-nearest neighbors, adaboost, and deep learning methods (DL), such as ResNet and GoogleNet, were used, and performance measures were calculated. Results Our study showed that Random forest and AdaBoost are the most reliable ML models for predicting LBP. Random forest showed high performance with area under curve (AUC) values from 0.83 to 0.88 across all lumbar vertebrae and L2-L3, L3-L4, and L4-L5 intervertebral discs (IVDs), with AUCs of 0.88 the highest at L5-S1 IVD (0.92). Adaboost demonstrated high performance at the L2-L5 vertebrae with AUC values of 0.82 to 0.90, with the highest AUC (0.97) at the L5-S1 IVD. Among the DL models, GoogleNet outperformed the other models at 30 epochs with an accuracy of 0.85, followed by ResNet 18 (30 epochs) with an accuracy of 0.84. Conclusion The study demonstrated that ML and DL models can effectively predict LBP from MRI T2 weighted image of the lumbar spine. ML and DL models could also enhance the diagnostic accuracy of LBP, potentially leading to better patient management and outcomes.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant