The role of artificial intelligence in predicting cardiovascular outcomes: a systematic review and meta-analysis
This systematic review and meta-analysis evaluated the performance, validation strategies, and clinical applicability of AI-based models for cardiovascular outcome prediction. Following PRISMA 2020 guidelines, seven databases were searched for studies published between 2015 and 2025. Studies reporting predictive performance of AI models were included. A bivariate random-effects model was used to pool sensitivity, specificity, and area under the curve (AUC), while heterogeneity, publication bias, and risk of bias (PROBAST) were assessed. Seventeen studies met the inclusion criteria. AI models, including DL, ensemble learning, and traditional ML methods, showed strong performance, with a pooled AUC of 0.86, sensitivity of 83.5%, and specificity of 81.7%. DL models achieved the highest performance with an AUC of 0.89. Models trained on large datasets and those externally validated demonstrated better generalizability. Moderate heterogeneity (I2 = 48%) reflected differences in study design, populations, and methodologies. Nearly half of the studies showed high risk of bias, mainly due to overfitting and limited calibration. Funnel plot asymmetry suggested possible publication bias. Explainability tools improved interpretability. Overall, AI models show strong potential but require more rigorous validation and transparent reporting before widespread clinical implementation. Abbreviations: AI: artificial intelligence; AUC: area under the curve; CNN: convolutional neural network; CVD: cardiovascular disease; DL: deep learning; EHR: electronic health record; F1-score: harmonic mean of precision and recall; Grad-CAM: gradient-weighted class activation mapping; K-fold: K-fold cross-validation; LIME: local interpretable model-agnostic explanations; MACE: major adverse cardiovascular events; ML: machine learning; MLP: multi-layer perceptron; PROBAST: prediction model risk of bias assessment tool; PRISMA: preferred reporting items for systematic reviews and meta-analyses; SHAP: Shapley additive explanations; SROC: summary receiver operating characteristic; TRIPOD-AI: transparent reporting of a multivariable prediction model for individual prognosis or diagnosis – artificial intelligence; CONSORT-AI: consolidated standards of reporting trials – artificial intelligence; XGBoost: extreme gradient boosting
- Research Article
- 10.64898/2026.02.06.26345251
- Feb 11, 2026
- medRxiv : the preprint server for health sciences
Artificial intelligence (AI) has emerged as a promising tool for interpreting 12-lead electrocardiograms (ECGs), with the potential to enhance diagnostic accuracy for arrhythmia detection. However, published studies vary widely in methodology and validation strategy, warranting a quantitative synthesis of diagnostic performance. A systematic review and meta-analysis was conducted according to the PRISMA-DTA 2018 guidelines and registered in PROSPERO (CRD420251027264). Searches were performed in MEDLINE, Embase, and Cochrane Library through September 2025 without language restrictions. Studies evaluating AI algorithms for arrhythmia detection using 12-lead ECGs were included. Data on sensitivity, specificity, and area under the curve (AUC) were extracted. Pooled estimates were generated using a bivariate random-effects model. Risk of bias was assessed with QUADAS-2, and the certainty of evidence was quantified using GRADE. 20 studies were included in the meta-analysis, encompassing over 5.5 million ECGs. The pooled sensitivity, specificity, and AUC for AI-based arrhythmia detection were 94.0% (95% CI 90.8-96.2; I2 = 96.9%), 98.7% (95% CI 97.3-99.3; I2 = 98.3%), and 0.982 (95% CI 0.965-0.986), respectively. Detection of atrial fibrillation (AF) yielded a sensitivity of 92.6% (95% CI 86.4-96), a specificity of 99.1% (95% CI 98.4-99.5), and an AUC of 0.988. Convolutional neural networks (CNN) specifically demonstrated a sensitivity of 97.6%, specificity of 98.7%, and an AUC of 0.982 for overall arrhythmia detection. When limited to external validation (n=6), the sensitivity was 96.9% (95% CI 89.2-99.1), specificity was 95.6% (95% CI 77.6-99.3), and AUC was 0.983. No significant publication bias was detected, and the overall certainty of evidence was rated as high. AI models applied to 12-lead ECGs demonstrate excellent diagnostic performance for arrhythmia detection. Findings support potential integration into clinical workflows, particularly in settings with limited cardiology expertise. Given substantial heterogeneity, standardized datasets and multicenter prospective validation are essential to ensure effective and equitable implementation.
- Supplementary Content
- 10.22037/aaem.v13i1.2720
- Dec 30, 2025
- Archives of Academic Emergency Medicine
Introduction:Missed or delayed diagnosis of pulmonary embolism (PE) is associated with increased morbidity, mortality, and longer hospitalizations. This study aimed to evaluate the diagnostic accuracy of Artificial Intelligence (AI) models in detecting PE across imaging.Methods:We systematically searched PubMed/MEDLINE, Scopus, Embase and Web of Science from inception to 1 January 2025 without language or regional limits. After removing duplicate results, the remaining records were screened through titles/abstracts, and two reviewers independently assessed full texts. Risk of bias was evaluated in duplicate with the QUADAS-2 tool. Pooled sensitivity, specificity, positive and negative likelihood ratios, diagnostic odds ratio and area under the ROC curve were calculated with random-effects models in STATA 17. Heterogeneity was quantified with Cochran’s Q and I², while we explored its sources using subgroup analyses (for categorical moderators) and meta-regression (for continuous moderators). Publication bias was assessed with Deeks’ funnel plot and trim-and-fill, and we examined robustness through leave-one-out sensitivity analyses. Results:A total of 1,432 records were identified through database searches, with 654 duplicates removed. After screening titles and abstracts of 787 articles, 256 full-text articles were assessed for eligibility, and 28 studies met the inclusion criteria. Internal validation phases included 43,330 participants (4,866 PE-positive, 38,463 PE-negative), while external validation phases comprised 3,588 participants (1,699 PE-positive, 1,889 PE-negative). In the internal validation phase, the pooled sensitivity and specificity of AI in PE diagnosis across imaging were 0.91 (95% confidence interval (CI): 0.88-0.95; I²=9%) and 0.94 (95% CI: 0.86-0.98; I²=99.78%), respectively. The positive likelihood ratio (PLR) was 16.08, and the negative likelihood ratio (NLR) was 0.09, both statistically significant (P < 0.001). The pooled diagnostic odds ratio (DOR) was 163.55 (95% CI: 71.30-375.14, I2: 96.1), and the area under the curve (AUC) was 0.95 (95% CI: 0.93 to 0.97), indicating excellent accuracy. In external validation, the pooled sensitivity and specificity were slightly lower at 0.89 (95% CI: 0.79-0.95; I²=95.60%) and 0.88 (95% CI: 0.80-0.93; I²=91.48%), respectively. The DOR was 59.65 (95% CI: 23.53 to 151.17, I2: 89.6) and AUC was 0.94 (95% CI: 0.92 to 0.96, I2: 89.6). There was no significant publication bias detected. Conclusion:AI models achieved high diagnostic accuracy in detecting PE through imaging. However, this performance tends to decrease from internal to external validation, highlighting limitations in generalizability. Additionally, substantial heterogeneity was observed across studies, as indicated by high I² values, which should be considered when interpreting the pooled estimates.
- Supplementary Content
- 10.3390/tomography12050062
- Apr 28, 2026
- Tomography
Prostate cancer (PCa) is the second most commonly diagnosed cancer in men globally. Radiation oncologists often find PCa tumor characterization and outcome prediction challenging. Therefore, the potential for artificial intelligence (AI) implementation in radiation oncology has increased in recent years. This systematic review aims to evaluate the efficacy of AI algorithms in characterizing PCa tumors and predicting post-therapy outcomes. A total of 2055 studies were identified through a comprehensive search across PubMed and Scopus, then exported to Covidence. Inclusion criteria focused on prospective and retrospective cohort studies as well as randomized clinical trials (RCTs) published between 2015 and 2024 that explored the implementation of AI in tumor characterization and outcome prediction of PCa. Two independent reviewers evaluated each paper, and evaluation metrics such as specificity, sensitivity, accuracy, and area under the curve (AUC) were analyzed. The Risk of Bias in Non-randomized Studies of Interventions, Version 2 (ROBINS-I V2) tool was used to assess the risk of bias (ROB). Across the 19 studies analyzed, there was no significant difference in model performance between machine learning (ML) and deep learning (DL) models. AI models using multi-input strategies (e.g., radiomics with clinical markers) generally performed better than single-input models. Of the imaging modalities used for radiomic feature extraction, multiparametric MRI (mpMRI)-trained AI models consistently achieved the highest performance. AI displays considerable potential for integration into clinical workflows for PCa management. However, further studies utilizing larger datasets and external cohorts independent of the sample population are needed to validate clinical utility and improve model transparency for reliable implementation.
- Research Article
8
- 10.1016/j.ejrad.2025.111948
- Mar 1, 2025
- European journal of radiology
To perform a systematic literature review of the efficacy of different AI models to predict HCC treatment response to transarterial chemoembolization (TACE), including overall survival (OS) and time to progression (TTP). This systematic review was performed according to the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines until May 2, 2024. The systematic review included 23 studies with 4,486 HCC patients. The AI algorithm receiver operator characteristic (ROC) area under the curve (AUC) for predicting HCC response to TACE based on mRECIST criteria ranged from 0.55 to 0.97. Radiomics-models outperformed non-radiomics models (AUCs: 0.79, 95%CI: 0.75-0.82 vs. 0.73, 0.61-0.77, respectively). The best ML methods used for the prediction of TACE response for HCC patients were CNN, GB, SVM, and RF with AUCs of 0.88 (0.79-0.97), 0.82 (0.71-0.89), 0.8 (0.60-0.87) and 0.8 (0.55-0.96), respectively. Of all predictive feature models, those combining clinic-radiologic features (ALBI grade, BCLC stage, AFP level, tumor diameter, distribution, and peritumoral arterial enhancement) had higher AUCs compared with models based on clinical characteristics alone (0.79, 0.73-0.89; p=0.04 for CT+clinical, 0.81, 0.75-0.88; p=0.017 for MRI+clinical versus 0.6, 0.55-0.75 in clinical characteristics alone). Integrating clinic-radiologic features enhances AI models' predictive performance for HCC patient response to TACE, with CNN, GB, SVM, and RF methods outperforming others. Key predictive clinic-radiologic features include ALBI grade, BCLC stage, AFP level, tumor diameter, distribution, and peritumoral arterial enhancement. Multi-institutional studies are needed to improve AI model accuracy, address heterogeneity, and resolve validation issues.
- Supplementary Content
1
- 10.2196/85319
- Mar 10, 2026
- JMIR Mental Health
BackgroundIn recent years, advances in wearable sensor technology and artificial intelligence (AI) have provided new possibilities for detecting and monitoring depression.ObjectiveThis study systematically reviewed and meta-analyzed the diagnostic and predictive performance of wearable device–based AI models for detecting depression and predicting depressive episodes and explored factors influencing outcomes.MethodsFollowing PRISMA-DTA (Preferred Reporting Items for a Systematic Review and Meta-Analysis of Diagnostic Test Accuracy) guidelines, the PubMed, Embase, Web of Science, and PsycINFO databases were searched from inception to May 27, 2025. Eligible studies used AI algorithms on wearable device data for depression detection or episode prediction. Sensitivity, specificity, diagnostic odds ratio, and area under the curve (AUC) were pooled using a bivariate random effects model. Risk of bias was assessed using Prediction Model Risk of Bias Assessment Tool plus artificial intelligence (PROBAST+ AI), and certainty of evidence was assessed using the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) tool.ResultsWe included 16 studies (32 datasets) with 1189 patients and 13,593 samples. For depression detection, pooled sensitivity and specificity were 0.89 (95% CI 0.83‐0.93) and 0.93 (95% CI 0.87‐0.96), with a diagnostic odds ratio of 110.47 (95% CI 33.33‐366.17) and AUC of 0.96 (95% CI 0.94‐0.98). Random forest models showed the best performance (sensitivity=0.89, specificity=0.91, AUC=0.97). Subgroup analyses indicated that study design, AI method, reference standard, and input type significantly affected diagnostic accuracy (P<.05). For depressive episode prediction (3 datasets), pooled sensitivity was 0.86 (95% CI 0.80‐0.91), and pooled specificity was 0.65 (95% CI 0.59‐0.71). The overall risk of bias was low to moderate, with no evidence of publication bias.ConclusionsWearable device–based AI models achieved high accuracy for detecting depression and moderate utility in predicting episodes. However, heterogeneity, reliance on retrospective and public datasets, and lack of standardized methods limited generalizability.
- Research Article
- 10.1161/svi270000_381
- Nov 1, 2025
- Stroke: Vascular and Interventional Neurology
Introduction Intracranial aneurysms (IAs) are a major cause of subarachnoid hemorrhage, yet accurate detection on CT angiography (CTA) remains challenging, particularly for small or complex lesions. While artificial intelligence (AI) algorithms have demonstrated high standalone accuracy, their true clinical value lies in augmenting human expertise. This systematic review and meta‐analysis evaluates the diagnostic accuracy of human‐AI collaboration compared with human‐only performance in detecting IAs. Methods We systematically reviewed meta‐analyses, multicenter retrospective validations, and prospective clinical trials reporting clinician performance with and without AI assistance. Eligible studies required reference standards from digital subtraction angiography (DSA) or expert consensus. Data on sensitivity, specificity, and area under the curve (AUC) were extracted. A pooled analysis was performed using random‐effects models to quantify improvements attributable to AI assistance. Secondary outcomes included reading time and clinician acceptance. Results Four representative studies encompassing more than 20,000 patients were analyzed. Across all datasets, AI assistance consistently enhanced clinician performance. In Gu et al. (2022), pooled clinician sensitivity increased from 0.84 to 0.92 with AI, accompanied by a rise in AUC from 0.85 to 0.93 and reduced interpretation time by ∼7 seconds per case. Din et al. (2023) reported similar improvements, with sensitivity rising from 0.83 to 0.90 and AUC from 0.88 to 0.91. Prospective real‐world validation by Hu et al. (2024) confirmed these gains: sensitivity improved from 0.59 to 0.825 and AUC from 0.787 to 0.909, with reading time reduced by ∼5 seconds per case. Wei et al. (2024) demonstrated external validation where clinician‐AI collaboration achieved an AUC of 0.93 versus 0.91 for radiology reports, while maintaining efficiency with a mean processing time of 1.7 minutes per scan. Pooled analysis across studies yielded a mean sensitivity improvement of +8% (95% CI: +6‐10%) and AUC gain of +0.06, both statistically significant. Clinician adoption rates exceeded 90% in prospective settings, underscoring strong acceptance of AI tools. Conclusion Human‐AI collaboration significantly improves the diagnostic accuracy of intracranial aneurysm detection on CTA, with consistent gains in sensitivity, AUC, and efficiency across retrospective and prospective studies. AI assistance reduces missed diagnoses while maintaining or improving workflow efficiency. These findings support the integration of AI as a diagnostic aid rather than a standalone tool. Future work should focus on harmonized reporting standards, prospective multicenter trials, and evaluation of long‐term clinical outcomes to fully establish AI's role in routine neuroradiological practice.
- Supplementary Content
1
- 10.3390/jcm14238466
- Nov 28, 2025
- Journal of Clinical Medicine
Background/Objectives: To evaluate the diagnostic accuracy of artificial intelligence (AI)-based imaging techniques for liver fibrosis and metabolic dysfunction-associated steatotic liver disease (MASLD). Materials and Methods: We performed a comprehensive search in PubMed, Embase, Cochrane Library, and Web of Science until August 2025. A total of 15 studies (mean age of patients 56 years, 60% male) were included. The risk of bias in the included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Diagnostic performance metrics were calculated using a random-effects bivariate model, including the area under the curve (AUC), sensitivity, specificity, positive and negative likelihood ratios, and diagnostic odds ratio. Meta-regression analysis was conducted to investigate potential sources of heterogeneity when I2 was ≥50%. A p-value < 0.05 was considered statistically significant. Results: For liver fibrosis, pooled sensitivity was 0.85, specificity was 0.81, and AUC was 0.92. For MASLD, sensitivity was 0.86, specificity was 0.95, and AUC was 0.99. Different imaging modalities and AI classifiers caused significant study heterogeneity. To avoid misleading pooled estimates across varied datasets, imaging modality and AI model subgroup analyses were performed. Only three studies were used to estimate MASLD; therefore, considerable between-study heterogeneity should be considered. Conclusions: AI-based imaging modalities demonstrate promising diagnostic accuracy for liver fibrosis and MASLD, warranting further standardization to enhance diagnostic consistency.
- Research Article
6
- 10.7759/cureus.82461
- Apr 17, 2025
- Cureus
Obesity is a global health crisis, with bariatric surgery considered a highly effective intervention for sustained weight loss and resolution of associated health conditions. Despite its benefits, some patients experience postoperative complications, emphasizing the importance of accurate risk prediction. Traditional models often lack the capacity to manage complex clinical data. Artificial intelligence (AI) offers transformative potential for improving the prediction of surgical complications. This systematic review synthesizes existing research on AI's role in forecasting complications following bariatric surgery. The review followed PRISMA 2020 guidelines, with searches conducted across PubMed, Scopus, Web of Science, and IEEE Xplore for studies examining AI applications in this context. Seven retrospective cohort studies were included, and data were extracted on study design, AI algorithms, and outcomes. Risk of bias was assessed using PROBAST, and a narrative synthesis was conducted due to study heterogeneity. The included studies showed variability in AI model performance, with ensemble methods and neural networks generally performing better than traditional logistic regression. Reported area under the curve (AUC) values varied widely, with higher accuracy noted for predicting specific complications such as diabetes and leaks. Key challenges included overfitting, data imbalance, and limited generalizability, especially in deep learning models. Most studies were conducted in Sweden and the United States, utilizing large datasets that may introduce regional biases. Overall, AI shows promise in enhancing complication prediction in bariatric surgery, though methodological limitations highlight the need for prospective, multicenter validation. Future research should focus on addressing data imbalance, refining feature selection, and facilitating the clinical integration of AI through decision-support systems to improve patient care.
- Research Article
1
- 10.5812/mejrh-158082
- Apr 29, 2025
- Middle East Journal of Rehabilitation and Health Studies
Context: Odontogenic keratocysts (OKCs) are aggressive jaw cysts characterized by a high recurrence rate, making accurate diagnosis critical for effective treatment. Recent advances in artificial intelligence (AI) have demonstrated potential for enhancing diagnostic accuracy in histopathology. However, the effectiveness of AI in diagnosing OKCs has not yet been systematically reviewed. Objectives: This study aims to evaluate the diagnostic and prognostic performance of AI models in detecting OKCs in histopathologic images. Methods: This systematic review was conducted in accordance with the preferred reporting items for systematic reviews and meta-analyses (PRISMA) guidelines. A comprehensive literature search was performed across PubMed, Scopus, Embase, Google Scholar, and ScienceDirect to identify studies that utilized AI models for diagnosing OKCs from histopathologic images. Studies were eligible for inclusion if they addressed the PICO (patient/population, intervention, comparison, and outcomes) framework, specifically investigating whether AI models (I) can enhance diagnostic and prognostic accuracy (O) for OKCs in histopathologic images (P). A meta-analysis was performed to pool the diagnostic performance of AI models across studies, and Egger’s test was conducted to assess publication bias. Results: A total of eight studies were included in the review. The risk of bias (ROB) across the included studies was generally low, with a few exceptions. The pooled area under the curve (AUC) for AI models in diagnosing OKCs was 0.967 (95% CI: 0.957 - 0.978). The pooled sensitivity ranged from 0.89 to 0.92, and the pooled specificity ranged from 0.88 to 0.94. The summary receiver operating characteristic (sROC) curve demonstrated an AUC of 0.93. Egger’s test for publication bias yielded a P-value of 0.522, indicating no significant evidence of publication bias. The review also highlighted several limitations, including small sample sizes, lack of external validation, and limited interpretability of the AI models. Conclusions: Artificial intelligence models, particularly deep learning architectures, demonstrate high diagnostic accuracy in detecting OKCs from histopathologic images.
- Research Article
7
- 10.1016/j.acra.2025.05.049
- Nov 1, 2025
- Academic radiology
Accuracy of CT-Based Radiomics Models for Preoperative Grading of Clear Cell Renal Cell Carcinoma: A Systematic Review and Meta-analysis.
- Preprint Article
- 10.2196/preprints.67922
- Oct 24, 2024
BACKGROUND Emerging evidence underscores the potential application of artificial intelligence (AI) in discovering noninvasive blood biomarkers. However, the diagnostic value of AI-derived blood biomarkers for ovarian cancer (OC) remains inconsistent. OBJECTIVE We aimed to evaluate the research quality and the validity of AI-based blood biomarkers in OC diagnosis. METHODS A systematic search was performed in the MEDLINE, Embase, IEEE Xplore, PubMed, Web of Science, and the Cochrane Library databases. Studies examining the diagnostic accuracy of AI in discovering OC blood biomarkers were identified. The risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies–AI tool. Pooled sensitivity, specificity, and area under the curve (AUC) were estimated using a bivariate model for the diagnostic meta-analysis. RESULTS A total of 40 studies were ultimately included. Most (n=31, 78%) included studies were evaluated as low risk of bias. Overall, the pooled sensitivity, specificity, and AUC were 85% (95% CI 83%-87%), 91% (95% CI 90%-92%), and 0.95 (95% CI 0.92-0.96), respectively. For contingency tables with the highest accuracy, the pooled sensitivity, specificity, and AUC were 95% (95% CI 90%-97%), 97% (95% CI 95%-98%), and 0.99 (95% CI 0.98-1.00), respectively. Stratification by AI algorithms revealed higher sensitivity and specificity in studies using machine learning (sensitivity=85% and specificity=92%) compared to those using deep learning (sensitivity=77% and specificity=85%). In addition, studies using serum reported substantially higher sensitivity (94%) and specificity (96%) than those using plasma (sensitivity=83% and specificity=91%). Stratification by external validation demonstrated significantly higher specificity in studies with external validation (specificity=94%) compared to those without external validation (specificity=89%), while the reverse was observed for sensitivity (74% vs 90%). No publication bias was detected in this meta-analysis. CONCLUSIONS AI algorithms demonstrate satisfactory performance in the diagnosis of OC using blood biomarkers and are anticipated to become an effective diagnostic modality in the future, potentially avoiding unnecessary surgeries. Future research is warranted to incorporate external validation into AI diagnostic models, as well as to prioritize the adoption of deep learning methodologies. CLINICALTRIAL PROSPERO CRD42023481232; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023481232
- Supplementary Content
11
- 10.2196/67922
- Mar 24, 2025
- Journal of Medical Internet Research
BackgroundEmerging evidence underscores the potential application of artificial intelligence (AI) in discovering noninvasive blood biomarkers. However, the diagnostic value of AI-derived blood biomarkers for ovarian cancer (OC) remains inconsistent.ObjectiveWe aimed to evaluate the research quality and the validity of AI-based blood biomarkers in OC diagnosis.MethodsA systematic search was performed in the MEDLINE, Embase, IEEE Xplore, PubMed, Web of Science, and the Cochrane Library databases. Studies examining the diagnostic accuracy of AI in discovering OC blood biomarkers were identified. The risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies–AI tool. Pooled sensitivity, specificity, and area under the curve (AUC) were estimated using a bivariate model for the diagnostic meta-analysis.ResultsA total of 40 studies were ultimately included. Most (n=31, 78%) included studies were evaluated as low risk of bias. Overall, the pooled sensitivity, specificity, and AUC were 85% (95% CI 83%-87%), 91% (95% CI 90%-92%), and 0.95 (95% CI 0.92-0.96), respectively. For contingency tables with the highest accuracy, the pooled sensitivity, specificity, and AUC were 95% (95% CI 90%-97%), 97% (95% CI 95%-98%), and 0.99 (95% CI 0.98-1.00), respectively. Stratification by AI algorithms revealed higher sensitivity and specificity in studies using machine learning (sensitivity=85% and specificity=92%) compared to those using deep learning (sensitivity=77% and specificity=85%). In addition, studies using serum reported substantially higher sensitivity (94%) and specificity (96%) than those using plasma (sensitivity=83% and specificity=91%). Stratification by external validation demonstrated significantly higher specificity in studies with external validation (specificity=94%) compared to those without external validation (specificity=89%), while the reverse was observed for sensitivity (74% vs 90%). No publication bias was detected in this meta-analysis.ConclusionsAI algorithms demonstrate satisfactory performance in the diagnosis of OC using blood biomarkers and are anticipated to become an effective diagnostic modality in the future, potentially avoiding unnecessary surgeries. Future research is warranted to incorporate external validation into AI diagnostic models, as well as to prioritize the adoption of deep learning methodologies.Trial RegistrationPROSPERO CRD42023481232; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023481232
- Research Article
- 10.1186/s12903-026-07770-4
- Feb 3, 2026
- BMC oral health
Artificial intelligence (AI) has shown increasing potential in dental diagnostics, yet its accuracy for binary classification of dental caries across different imaging modalities remains unclear. This study aimed to systematically evaluate the diagnostic performance of AI models using clinical intraoral images and dental radiographs. Following the PRISMA-DTA guidelines, PubMed, Embase, Scopus, Web of Science, and IEEE Xplore were systematically searched for studies published between January 2015 and June 2025. Eligible studies applied AI models for caries diagnosis with extractable sensitivity and specificity. Data on dentition, dataset, analysis unit, caries prevalence in test dataset, and preprocessing methods were extracted. Reporting quality and risk of bias were assessed using CLAIM and QUADAS-2. Pooled estimates were calculated with a bivariate random-effects model, with subgroup analyses by image type and analytical unit. 25 studies met the criteria, and 13 were included in the meta-analysis. Pooled sensitivity, specificity, and the area under the curve (AUC) were 0.86, 0.91 and 0.94, respectively. Intraoral image-based models achieved higher sensitivity (0.88) and AUC (0.95), while radiograph-based models showed higher specificity (0.92). Tooth-level analyses yielded stable, clinically relevant performance (0.87/0.91). High heterogeneity (I² > 90%) was partly explained by image type, model architecture, reference standard variation, and test set caries prevalence. AI models showed good diagnostic accuracy for caries detection across imaging modalities and analytical units. However, given the substantial heterogeneity and limitations in study quality and reference standards, these summary estimates should be interpreted with caution. AI-based systems may serve as complementary decision-support tools in clinical practice, but further standardization, external validation, and high-quality multicenter studies are required before broad clinical implementation.
- Research Article
- 10.1093/humrep/deaf097.404
- Jun 1, 2025
- Human Reproduction
Study question What is the performance and robustness of an interpretable artificial intelligence model in embryo selection across two different time-lapse systems? Summary answer Internal and external validation showed that the interpretable AI system provided stable clinical pregnancy rate predictions across two incubator brands. What is known already Embryo assessment is vital for successful single embryo transfers. With advances in time-lapse imaging and AI, we can now continuously monitor embryos and analyze images for better outcome predictions. The necessity for thorough internal and external validation of AI systems to demonstrate their reliability are increasingly recognized. While in most studies, AI models were developed and validated on the same brand time-lapse system, despite significant differences in imaging quality and photographic parameters across different incubators that may affect AI model performance. Study design, size, duration A total of 7917 embryos from 787 patients who underwent IVF with D5 single blastocyst transfer in fresh cycles at the Center for Reproductive Medicine and Obstetrics and Gynecology, Nanjing Drum Tower Hospital, from September 2016 to October 2024 were analyzed in this study.The embryos were cultured in CCM-iBIS(ASTEC, Japan) and EmbryoScope(Vitrolife, Sweden) time-lapse incubators.The CCM-iBIS group had 397 pregnancies and 272 non-pregnant cases, while the EmbryoScope group had 88 pregnancies and 30 non-pregnant cases. Participants/materials, setting, methods A t-test was used to select clinical features to train logistic regression to predict pregnancy outcomes along with AI scores. We developed this model using Bootstrap resampling on the internal dataset of CCM-iBIS, which was split in the ratio of 80:20 and EmbryoScope data for external validation.Accuracy(ACC) and Area Under the Curve(AUC) were calculated for two datasets respectively, and the performance of the model was compared using the Delong test and Chi-Square test. Main results and the role of chance In this study, we developed a two-stage multimodal AI model for intelligent blastocyst assessment, integrating deep learning-based interpretable embryo image analysis and clinical characterization data. A logistic regression model was trained to predict the likelihood of clinical pregnancy, using inputs such as assessment scores, left antral follicle count (L-AFC), and right antral follicle count (R-AFC) derived from the interpretable AI. On the internal validation set, the model achieved an ACC of 0.703 and an AUC of 0.697. Bootstrap sampling was performed 1000 times, with each sampling conducted separately from the training set. The mean AUC of the model was then calculated, resulting in a mean AUC of 0.64 ± 0.08 for the internal validation. The 95% confidence interval for the AUC difference was [-0.131, 0.115]. When tested externally, the model achieved an ACC of 0.721 and an AUC of 0.627. These results indicate that the model performance is not significantly different from the internal validation (pACC=0.952, pACC&gt;0.05, pAUC=0.104, pAUC&gt;0.05). Limitations, reasons for caution The current cross-brand validation only involves time-lapse imaging systems from two brands. Future studies should include more brands and consider additional outcome predictions, such as embryo ploidy models, to ensure comprehensive validation. Wider implications of the findings The AI blastocyst assessment system, powered by interpretable models, demonstrated consistent AUC performance across CCM-iBIS and EmbryoScope datasets. It emphasizes enhancing reliability, cross-brand applicability, and trustworthiness in embryo selection, providing a valuable reference for developing more generalized models. Trial registration number No
- Research Article
1
- 10.1016/j.crad.2025.107001
- Jan 1, 2026
- Clinical radiology
Artificial intelligence in CT for predicting lymph node metastasis in rectal cancer patients: a meta-analysis.