The Enviromic marker
Abstract We formalize “enviromic markers” as modeling units parallel to DNA markers, but herein for genotype-environment (G × E) prediction. Four operational premises (linearity; site potential; heterogeneous favorability; and envirotypic covariates (ECs)-genotype-dependence) are presented to enable their use in linear mixed models and also to motivate four construction strategies: (i) using raw environmental covariates as linear markers; (ii) applying transformations to capture mild nonlinearities; (iii) deriving ecophysiological functions; and (iv) engineering markers with Artificial Intelligence (AI) models which learn nonlinear environment → phenotype mappings for linear downstream use. Environmental data quality control is detailed, including checks of spatial coverage and resolution, variance within the TPE, collinearity control, and spatial/temporal validation without leakage. Envirome data are linked with GIS to compute environmental kernels, quantify covariate shifts, and deliver pixel-level predictions with uncertainty diagnostics. The framework clarifies assumptions and standardizes the use of enviromic markers for predictive breeding analyses.
- Research Article
10
- 10.1016/j.jacr.2021.06.025
- Feb 1, 2022
- Journal of the American College of Radiology
Real-World Surveillance of FDA-Cleared Artificial Intelligence Models: Rationale and Logistics.
- Research Article
50
- 10.1016/j.fertnstert.2020.10.040
- Nov 1, 2020
- Fertility and Sterility
Predictive modeling in reproductive medicine: Where will the future of artificial intelligence research take us?
- Research Article
14
- 10.1101/2025.03.14.25323836
- Mar 17, 2025
- medRxiv : the preprint server for health sciences
This study evaluates the diagnostic performance of several AI models, including Deepseek, in diagnosing corneal diseases, glaucoma, and neuro□ophthalmologic disorders. We retrospectively selected 53 case reports from the Department of Ophthalmology and Visual Sciences at the University of Iowa, comprising 20 corneal disease cases, 11 glaucoma cases, and 22 neuro□ophthalmology cases. The case descriptions were input into DeepSeek, ChatGPT□4.0, ChatGPT□01, and Qwens 2.5 Max. These responses were compared with diagnoses rendered by human experts (corneal specialists, glaucoma attendings, and neuro□ophthalmologists). Diagnostic accuracy and interobserver agreement, defined as the percentage difference between each AI model's performance and the average human expert performance, were determined. DeepSeek achieved an overall diagnostic accuracy of 79.2%, with specialty-specific accuracies of 90.0% in corneal diseases, 54.5% in glaucoma, and 81.8% in neuro□ophthalmology. ChatGPT□01 outperformed the other models with an overall accuracy of 84.9% (85.0% in corneal diseases, 63.6% in glaucoma, and 95.5% in neuro□ophthalmology), while Qwens exhibited a lower overall accuracy of 64.2% (55.0% in corneal diseases, 54.5% in glaucoma, and 77.3% in neuro□ophthalmology). Interobserver agreement analysis revealed that in corneal diseases, DeepSeek differed by -3.3% (90.0% vs 93.3%), ChatGPT□01 by -8.3%, and Qwens by -38.3%. In glaucoma, DeepSeek outperformed the human expert average by +3.0% (54.5% vs 51.5%), while ChatGPT□4.0 and ChatGPT□01 exceeded it by +12.1%, and Qwens was +3.0% above the human average. In neuro□ophthalmology, DeepSeek and ChatGPT□4.0 were 9.1% lower than the human average, ChatGPT□01 exceeded it by +4.6%, and Qwens was 13.6% lower. ChatGPT□01 demonstrated the highest overall diagnostic accuracy, especially in neuro□ophthalmology, while DeepSeek and ChatGPT□4.0 showed comparable performance. Qwens underperformed relative to the other models, especially in corneal diseases. Although these AI models exhibit promising diagnostic capabilities, they currently lag behind human experts in certain areas, underscoring the need for a collaborative integration of clinical judgment. This study evaluated how well several artificial intelligence (AI) models diagnose eye diseases compared to human experts. We tested four AI systems across three types of eye conditions: diseases of the cornea, glaucoma, and neuro-ophthalmologic disorders. Overall, one AI model, ChatGPT-01, performed the best, correctly diagnosing about 85% of cases, and it excelled in neuro-ophthalmology by correctly diagnosing 95.5% of cases. Two other models, DeepSeek and ChatGPT-4.0, each achieved an overall accuracy of around 79%, while the Qwens model performed lower, with an overall accuracy of about 64%. When compared with human experts, who achieved very high accuracy in corneal diseases (93.3%) and neuro-ophthalmology (90.9%) but lower in glaucoma (51.5%), the AI models showed mixed results. In glaucoma, for instance, some AI models even outperformed human experts slightly, while in corneal diseases, all AI models were less accurate than the experts. These findings indicate that while AI shows promise as a supportive tool in diagnosing eye conditions, it still needs further improvement. Combining AI with human clinical judgment appears to be the best approach for accurate eye disease diagnosis. Why carry out this study? With the rising burden of eye diseases and the inherent diagnostic challenges for complex conditions like glaucoma and neuro-ophthalmologic disorders, there is an unmet need for innovative diagnostic tools to support clinical decision-making. What did the study ask? This study evaluated the diagnostic performance of four AI models across three ophthalmologic subspecialties, testing the hypothesis that advanced language models can achieve accuracy levels comparable to human experts. What was learned from the study? Our results showed that ChatGPT-01 achieved the highest overall accuracy (84.9%), excelling in neuro-ophthalmology with a 95.5% accuracy, while DeepSeek and ChatGPT-4.0 each achieved 79.2%, and Qwens reached 64.2%. What specific outcomes were observed? In glaucoma, AI model accuracies ranged from 54.5% to 63.6%, with some models slightly surpassing the human expert average of 51.5%, underscoring the diagnostic difficulty of this condition. What has been learned and future implications? These findings highlight the potential of AI as a valuable adjunct to clinical judgment in ophthalmology, although further research and the integration of multimodal data are essential to optimize these tools for routine clinical practice.
- Research Article
63
- 10.1007/s00535-022-01849-9
- Jan 1, 2022
- Journal of Gastroenterology
BackgroundUltrasonography (US) is widely used for the diagnosis of liver tumors. However, the accuracy of the diagnosis largely depends on the visual perception of humans. Hence, we aimed to construct artificial intelligence (AI) models for the diagnosis of liver tumors in US.MethodsWe constructed three AI models based on still B-mode images: model-1 using 24,675 images, model-2 using 57,145 images, and model-3 using 70,950 images. A convolutional neural network was used to train the US images. The four-class liver tumor discrimination by AI, namely, cysts, hemangiomas, hepatocellular carcinoma, and metastatic tumors, was examined. The accuracy of the AI diagnosis was evaluated using tenfold cross-validation. The diagnostic performances of the AI models and human experts were also compared using an independent test cohort of video images.ResultsThe diagnostic accuracies of model-1, model-2, and model-3 in the four tumor types are 86.8%, 91.0%, and 91.1%, whereas those for malignant tumor are 91.3%, 94.3%, and 94.3%, respectively. In the independent comparison of the AIs and physicians, the percentages of correct diagnoses (accuracies) by the AIs are 80.0%, 81.8%, and 89.1% in model-1, model-2, and model-3, respectively. Meanwhile, the median percentages of correct diagnoses are 67.3% (range 63.6%–69.1%) and 47.3% (45.5%–47.3%) by human experts and non-experts, respectively.Conclusion The performance of the AI models surpassed that of human experts in the four-class discrimination and benign and malignant discrimination of liver tumors. Thus, the AI models can help prevent human errors in US diagnosis.
- Research Article
- 10.1093/humrep/deae108.541
- Jul 3, 2024
- Human Reproduction
Study question What is the performance of an image-based artificial intelligence (AI) model for ranking blastocyst stage embryos compared to embryologists using traditional morphology? Summary answer The AI was non-inferior to manual embryo selection. The AI showed significant improvement in clinical pregnancy when there was disagreement between AI and manual selection. What is known already In previous work, we developed an image-based AI model that predicts the likelihood of clinical pregnancy by analyzing a single static image of a blastocyst captured prior to biopsy or freeze. This model was trained on data from over 8,000 single-blastocyst transfer cycles from multiple U.S. IVF clinics performed between 2014 to 2021. Study design, size, duration We performed a retrospective, double-blinded, comparative reader study. The study included data from 438 single-blastocyst transfers from 10 different IVF clinics in U.S. that were not part of previous model development or testing. Using this data, a set of 1,257 virtual patient panels were created. Each virtual patient panel included between 2-5 embryos that were matched by age (18 - 29, 30 - 34, 35 - 37, ≥38), race (white, non-white, and unknown) and PGT-status. Participants/materials, setting, methods A group of 5 embryologists (readers) with varying levels of experience were asked to select their top embryo for transfer for each virtual patient panel (control arm) based on morphology grades. The AI model was also used to select a top embryo for transfer from each patient panel (treatment arm). The clinical pregnancy rates of the top-selected embryos were calculated and compared. Main results and the role of chance There was disagreement on the top pick embryo amongst the five embryologists 34.6% of the time, which increased to 43.9% when there were 3 or more embryos to choose from, supporting the need for a tool to standardize this decision. The clinical pregnancy rate of the control arm (average of embryologist readers) was 61.0% (individual rates of 58.9%, 59.6%, 61.5%, 61.6%, and 63.3%), and the clinical pregnancy rate of the treatment arm (AI model) was 62.3% (demonstrating non-inferiority with p<.001). The pregnancy rate of random embryo selection was 53.2%. All 5 of the embryologist readers agreed on the top-pick embryo 65% of the time. When all 5 readers agreed, the AI model disagreed with that consensus 31% of the time, and in these cases the AI model pregnancy rate was significantly higher by 8.6% (AI: 63.1%, Embryologist Consensus: 54.5% (p < 0.05)). Limitations, reasons for caution While data for this study was collected prospectively, the analysis was done retrospectively. Furthermore, readers were provided morphology grades from retrospective data rather than grading the embryos themselves. However, it is common for an embryologist to make an embryo selection based on morphology grades previously assigned by a different embryologist. Wider implications of the findings The AI model was able to select the top embryo for transfer with performance comparable to experienced embryologists. Such a model could allow for automated and objective embryo selection using a single static image of blastocyst stage embryos. Trial registration number NA
- Research Article
1
- 10.54364/aaiml.2024.43159
- Jan 1, 2024
- Advances in Artificial Intelligence and Machine Learning
Introduction The accurate prediction of mandibular bone growth is crucial in orthodontics and maxillofacial surgery, impacting treatment planning and patient outcomes. Traditional methods often fall short due to their reliance on linear models and clinician expertise, which are prone to human error and variability. Artificial intelligence (AI) and machine learning (ML) offer advanced alternatives, capable of processing complex datasets to provide more accurate predictions. This systematic review examines the efficacy of AI and ML models in predicting mandibular growth compared to traditional methods. Method. A systematic review was conducted following the PRISMA guidelines, focusing on studies published up to July 2024. Databases searched included PubMed, Embase, Scopus, and Web of Science. Studies were selected based on their use of AI and ML algorithms for predicting mandibular growth. A total of 31 studies were identified, with 6 meeting the inclusion criteria. Data were extracted on study characteristics, AI models used, and prediction accuracy. The risk of bias was assessed using the QUADAS-2 tool. Results. The review found that AI and ML models generally provided high accuracy in predicting mandibular growth. For instance, the LASSO model achieved an average error of 1.41 mm for predicting skeletal landmarks. However, not all AI models outperformed traditional methods; in some cases, deep learning models were less accurate than conventional growth prediction models. Discussion. The variability in datasets and study designs across the included studies posed challenges for comparing AI models’ effectiveness. Additionally, the complexity of AI models may limit their clinical applicability. Despite these challenges, AI and ML show significant promise in enhancing predictive accuracy for mandibular growth. Conclusion. AI and ML models have the potential to revolutionize mandibular growth prediction, offering greater accuracy and reliability than traditional methods. However, further research is needed to standardize methodologies, expand datasets, and improve model interpretability for clinical integration.
- Research Article
2
- 10.2196/72815
- Jul 8, 2025
- JMIR Formative Research
BackgroundQualitative research appraisal is crucial for ensuring credible findings but faces challenges due to human variability. Artificial intelligence (AI) models have the potential to enhance the efficiency and consistency of qualitative research assessments.ObjectiveThis study aims to evaluate the performance of 5 AI models (GPT-3.5, Claude 3.5, Sonar Huge, GPT-4, and Claude 3 Opus) in assessing the quality of qualitative research using 3 standardized tools: Critical Appraisal Skills Programme (CASP), Joanna Briggs Institute (JBI) checklist, and Evaluative Tools for Qualitative Studies (ETQS).MethodsAI-generated assessments of 3 peer-reviewed qualitative papers in health and physical activity–related research were analyzed. The study examined systematic affirmation bias, interrater reliability, and tool-dependent disagreements across the AI models. Sensitivity analysis was conducted to evaluate the impact of excluding specific models on agreement levels.ResultsResults revealed a systematic affirmation bias across all AI models, with “Yes” rates ranging from 75.9% (145/191; Claude 3 Opus) to 85.4% (164/192; Claude 3.5). GPT-4 diverged significantly, showing lower agreement (“Yes”: 115/192, 59.9%) and higher uncertainty (“Cannot tell”: 69/192, 35.9%). Proprietary models (GPT-3.5 and Claude 3.5) demonstrated near-perfect alignment (Cramer V=0.891; P<.001), while open-source models showed greater variability. Interrater reliability varied by assessment tool, with CASP achieving the highest baseline consensus (Krippendorff α=0.653), followed by JBI (α=0.477), and ETQS scoring lowest (α=0.376). Sensitivity analysis revealed that excluding GPT-4 increased CASP agreement by 20% (α=0.784), while removing Sonar Huge improved JBI agreement by 18% (α=0.561). ETQS showed marginal improvements when excluding GPT-4 or Claude 3 Opus (+9%, α=0.409). Tool-dependent disagreements were evident, particularly in ETQS criteria, highlighting AI’s current limitations in contextual interpretation.ConclusionsThe findings demonstrate that AI models exhibit both promise and limitations as evaluators of qualitative research quality. While they enhance efficiency, AI models struggle with reaching consensus in areas requiring nuanced interpretation, particularly for contextual criteria. The study underscores the importance of hybrid frameworks that integrate AI scalability with human oversight, especially for contextual judgment. Future research should prioritize developing AI training protocols that emphasize qualitative epistemology, benchmarking AI performance against expert panels to validate accuracy thresholds, and establishing ethical guidelines for disclosing AI’s role in systematic reviews. As qualitative methodologies evolve alongside AI capabilities, the path forward lies in collaborative human-AI workflows that leverage AI’s efficiency while preserving human expertise for interpretive tasks.
- Research Article
4
- 10.1093/jas/skae234.078
- Sep 13, 2024
- Journal of Animal Science
Quantum computing (QC) is not a futuristic notion in agriculture, though its full potential has yet to be realized. QC is an emerging field at the intersection of physics and computer science that holds immense potential to revolutionize various sectors, including agriculture production and artificial intelligence (AI) modeling. While QC is still in the early stages of development and practical applications within agriculture are not yet widespread, researchers are actively exploring its potential benefits in various agricultural domains, including crop optimization, livestock breeding, and environmental monitoring. QC harnesses the principles of quantum mechanics to perform computations using quantum bits or qubits, which can exist in multiple states simultaneously. Unlike classical computers, which rely on binary bits representing 0 or 1, quantum computers exploit phenomena such as superposition and entanglement to process information in parallel, potentially offering exponential speedup for certain types of problems. In agriculture production, particularly in animal science, QC offers promising avenues for optimizing processes and enhancing productivity. Quantum algorithms can analyze vast amounts of genomic data to improve breeding programs, leading to the development of more resilient and productive livestock breeds. Furthermore, QC can facilitate precision farming techniques by modeling complex environmental factors and animal behavior to optimize feeding strategies, disease management, and overall farm management practices. Moreover, QC can significantly benefit AI modeling by accelerating computations and enabling more efficient training of AI models. Quantum algorithms can enhance the performance of AI algorithms in various tasks, including pattern recognition, natural language processing, and predictive analytics. By leveraging quantum-enhanced optimization algorithms, AI models can achieve better convergence and accuracy, leading to more effective decision-making and problem-solving capabilities. While hybrid intelligent models also represent a novel frontier in agriculture, QC has the potential to expedite the merging of mechanistic and AI modeling paradigms, facilitating a more holistic understanding of complex systems in agriculture and beyond. By integrating mechanistic models, which describe the underlying physical processes, with AI models, which learn patterns from data, quantum computing can enable comprehensive simulations and predictions of agricultural systems. This fusion of modeling paradigms can lead to more accurate and robust predictions of crop yields, livestock performance, and environmental impacts, facilitating informed decision-making for farmers and policymakers. The application of QC in agriculture, however, requires interdisciplinary collaborations between physicists, computer scientists, agronomists/animal scientists, and AI researchers. These collaborations can drive the development of quantum algorithms tailored to agricultural applications, the integration of quantum-enhanced AI techniques into existing modeling frameworks, and the deployment of QC resources in real-world agricultural systems. Ultimately, harnessing the power of QC holds the potential to revolutionize agriculture production practices, including regenerative agriculture, and advance AI modeling capabilities, paving the way for a more sustainable and efficient agricultural industry.
- Research Article
- 10.1158/1538-7445.advbc23-b078
- Feb 1, 2024
- Cancer Research
Introduction: To enhance reproducibility and robustness in mammographic density assessment, various artificial intelligence (AI) models have been proposed to automatically classify mammographic images into BI-RADS density categories. Despite their promising performances, so far density AI models have been assessed primarily in traditional full-field digital mammography (FFDM) images. Our study aims to assess the potential of AI in breast density assessment in FFDM versus the newer synthetic mammography (SM) images acquired with digital breast tomosynthesis. Methods: We retrospectively analyzed negative (BI-RADS 1 or 2) routine mammographic screening exams (Selenia or Selenia Dimensions; Hologic) acquired at sites within the Barnes-Jewish/Christian (BJC) Healthcare network in St. Louis, MO from 2015 to 2018. BI-RADS breast density assessments of radiologists were obtained from BJC’s mammography reporting software (Magview 7.1). For each mammographic imaging modality, a balanced dataset of 4,000 women was selected so there were equal numbers of women in each of the four BI-RADS density categories, and each woman had at least one mediolateral oblique (MLO) and one craniocaudal (CC) view per breast in that mammographic imaging modality. Previously validated pre-processing steps were applied to all FFDM and SM images to standardize image orientation and intensity. Images were then split into training, validation, and test sets at ratios of 80%, 10%, and 10%, respectively, while maintaining the distribution of breast density categories and ensuring that all images of the same woman appear only in one set. Our AI model was based on the widely used ResNet50 architecture and was designed to accept as an input a mammographic image and predict the BI-RADS breast density category that the image belongs to. Our AI model was optimized, trained, and evaluated separately for each mammographic imaging modality. We report on the AI model’s predictive accuracy on the test set for each mammographic imaging modality, for both views as well as separately for CC and MLO; accuracy differences in FFDM versus SM were assessed via bootstrapping. Results: A batch size of 32, learning rate of e-6, and Adam optimizer were chosen as the optimal hyperparameters for our AI model. Using the same hyperparameters, the AI model demonstrated substantially higher accuracy on the test set for FFDM than for SM (FFDM: accuracy = 71% ± 4.5% versus SM: accuracy = 66% ± 4.2%; p-value&lt;0.001 for comparison). Similar conclusion held when CC and MLO views were evaluated separately (accuracy = 72% ± 4.6% versus 66% ± 4.3% for CC; accuracy = 69% ± 4.5% versus 62% ± 4.3% for MLO; p-value&lt;0.001 for both comparisons). Conclusions: AI performance in BI-RADS breast density assessment was significantly higher on FFDM versus SM, even under the same AI model design, dataset size and training process. Our preliminary findings suggest that further AI optimizations and adaptations may be needed as we translate AI models from FFDM to the newer SM format acquired with digital breast tomosynthesis. Citation Format: Krisha Anant, Juanita Hernandez Lopez, Debbie Bennett, Aimilia Gastounioti. Artificial-intelligence-driven breast density assessment in the transition from full-field digital mammograms to digital breast tomosynthesis [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Breast Cancer Research; 2023 Oct 19-22; San Diego, California. Philadelphia (PA): AACR; Cancer Res 2024;84(3 Suppl_1):Abstract nr B078.
- Research Article
15
- 10.1186/s12884-023-06046-x
- Oct 10, 2023
- BMC Pregnancy and Childbirth
BackgroundTo study the validity of an artificial intelligence (AI) model for measuring fetal facial profile markers, and to evaluate the clinical value of the AI model for identifying fetal abnormalities during the first trimester.MethodsThis retrospective study used two-dimensional mid-sagittal fetal profile images taken during singleton pregnancies at 11–13+ 6 weeks of gestation. We measured the facial profile markers, including inferior facial angle (IFA), maxilla-nasion-mandible (MNM) angle, facial-maxillary angle (FMA), frontal space (FS) distance, and profile line (PL) distance using AI and manual measurements. Semantic segmentation and landmark localization were used to develop an AI model to measure the selected markers and evaluate the diagnostic value for fetal abnormalities. The consistency between AI and manual measurements was compared using intraclass correlation coefficients (ICC). The diagnostic value of facial markers measured using the AI model during fetal abnormality screening was evaluated using receiver operating characteristic (ROC) curves.ResultsA total of 2372 normal fetuses and 37 with abnormalities were observed, including 18 with trisomy 21, 7 with trisomy 18, and 12 with CLP. Among them, 1872 normal fetuses were used for AI model training and validation, and the remaining 500 normal fetuses and all fetuses with abnormalities were used for clinical testing. The ICCs (95%CI) of the IFA, MNM angle, FMA, FS distance, and PL distance between the AI and manual measurement for the 500 normal fetuses were 0.812 (0.780–0.840), 0.760 (0.720–0.795), 0.766 (0.727-0.800), 0.807 (0.775–0.836), and 0.798 (0.764–0.828), respectively. IFA clinically significantly identified trisomy 21 and trisomy 18, with areas under the ROC curve (AUC) of 0.686 (95%CI, 0.585–0.788) and 0.729 (95%CI, 0.621–0.837), respectively. FMA effectively predicted trisomy 18, with an AUC of 0.904 (95%CI, 0.842–0.966). MNM angle and FS distance exhibited good predictive value in CLP, with AUCs of 0.738 (95%CI, 0.573–0.902) and 0.677 (95%CI, 0.494–0.859), respectively.ConclusionsThe consistency of fetal facial profile marker measurements between the AI and manual measurement was good during the first trimester. The AI model is a convenient and effective tool for the early screen for fetal trisomy 21, trisomy 18, and CLP, which can be generalized to first-trimester scanning (FTS).
- Research Article
1
- 10.3390/cancers14122915
- Jun 13, 2022
- Cancers
Simple SummaryThe need for predictive and prognostic biomarkers in colorectal carcinoma (CRC) brought us to an era where the use of artificial intelligence (AI) models is increasing. We investigated the expression of Claudin-7, a tight junction component, which plays a crucial role in maintaining the integrity of normal epithelial mucosa, and its potential prognostic role in advanced CRCs by drawing a parallel between statistical and AI algorithms. Claudin-7 immunohistochemical expression was evaluated in the tumor core and invasion front of CRCs and correlated with clinicopathological parameters and survival using statistical and AI algorithms. The Kaplan–Meier univariate survival analysis showed that the immunohistochemical overexpression of Claudin-7 in the tumor invasive front may represent a poor prognostic factor in advanced stages of CRCs. On the contrary, AI models could not predict the same outcome, probably because of the small number of patients included in our cohort.Aim: The need for predictive and prognostic biomarkers in colorectal carcinoma (CRC) brought us to an era where the use of artificial intelligence (AI) models is increasing. We investigated the expression of Claudin-7, a tight junction component, which plays a crucial role in maintaining the integrity of normal epithelial mucosa, and its potential prognostic role in advanced CRCs, by drawing a parallel between statistical and AI algorithms. Methods: Claudin-7 immunohistochemical expression was evaluated in the tumor core and invasion front of CRCs from 84 patients and correlated with clinicopathological parameters and survival. The results were compared with those obtained by using various AI algorithms. Results: the Kaplan–Meier univariate survival analysis showed a significant correlation between survival and Claudin-7 intensity in the invasive front (p = 0.00), a higher expression being associated with a worse prognosis, while Claudin-7 intensity in the tumor core had no impact on survival. In contrast, AI models could not predict the same outcome on survival. Conclusion: The study showed through statistical means that the immunohistochemical overexpression of Claudin-7 in the tumor invasive front may represent a poor prognostic factor in advanced stages of CRCs, contrary to AI models which could not predict the same outcome, probably because of the small number of patients included in our cohort.
- Research Article
3
- 10.1111/cts.70353
- Oct 23, 2025
- Clinical and Translational Science
A Comparison of AI and Population PK Models to Predict the Concentrations of Antiepileptic Drugs Using Therapeutic Drug Monitoring Records
- Research Article
37
- 10.1007/s41870-023-01635-7
- Jan 2, 2024
- International Journal of Information Technology
The big Artificial General Intelligence models inspire hot topics currently. The black box problems of Artificial Intelligence (AI) models still exist and need to be solved urgently, especially in the medical area. Therefore, transparent and reliable AI models with small data are also urgently necessary. To build a trustable AI model with small data, we proposed a prior knowledge-integrated transformer model. We first acquired prior knowledge using Shapley Additive exPlanations from various pre-trained machine learning models. Then, we used the prior knowledge to construct the transformer models and compared our proposed models with the Feature Tokenization Transformer model and other classification models. We tested our proposed model on three open datasets and one non-open public dataset in Japan to confirm the feasibility of our proposed methodology. Our results certified that knowledge-integrated transformer models perform better (1%) than general transformer models. Meanwhile, our proposed methodology identified that the self-attention of factors in our proposed transformer models is nearly the same, which needs to be explored in future work. Moreover, our research inspires future endeavors in exploring transparent small AI models.
- Research Article
4
- 10.1016/j.ajo.2025.08.046
- Dec 1, 2025
- American journal of ophthalmology
Development and Evaluation of an Artificial Intelligence Model to Set Target IOP for Glaucoma.
- Research Article
12
- 10.1016/j.echo.2024.06.016
- Jul 2, 2024
- Journal of the American Society of Echocardiography
Heart failure with preserved ejection fraction (HFpEF) accounts for approximately 50% of diagnoses of heart failure (HF) and frequently leads to hospitalization. Clinical algorithms developed for diagnosis have been applied to stratify risk for HF hospitalization or death.1Heidenreich P.A. Bozkurt B. Aguilar D. et al.2022 AHA/ACC/HFSA guideline for the Management of heart failure: executive summary: a report of the American College of Cardiology/American Heart Association Joint Committee on clinical practice guidelines.J Am Coll Cardiol. 2022; 79: 1757-1780Crossref PubMed Scopus (392) Google Scholar, 2Selvaraj S. Myhre P.L. Vaduganathan M. et al.Application of diagnostic algorithms for heart failure with preserved ejection fraction to the community.JACC Heart Fail. 2020; 8: 640-653Crossref PubMed Scopus (80) Google Scholar, 3Verbrugge F.H. Reddy Y.N.V. Sorimachi H. et al.Diagnostic scores predict morbidity and mortality in patients hospitalized for heart failure with preserved ejection fraction.Eur J Heart Fail. 2021; 23: 954-963Crossref PubMed Scopus (29) Google Scholar Deep learning has been applied to the automated interpretation of echocardiograms, but limited information exists regarding the potential of the learning models for predicting clinical outcomes.4Lau E.S. Di Achille P. Kopparapu K. et al.Deep learning-enabled assessment of left heart structure and function predicts cardiovascular outcomes.J Am Coll Cardiol. 2023; 82: 1936-1948Crossref Scopus (8) Google Scholar An artificial intelligence (AI) model was recently developed to identify patients with HFpEF using a single apical four-chamber video clip from a standard transthoracic echocardiographic examination.5Akerman A.P. Porumb M. Scott C.G. et al.Automated echocardiographic detection of heart failure with preserved ejection fraction using artificial intelligence.JACC: Advances. 2023; 2100452Crossref Scopus (20) Google Scholar A convolutional neural network was applied to the video clip. The model comprised a series of three-dimensional convolutional layers designed to operate on two-dimensional videos over two in-plane spatial dimensions within the image frames and across the time dimension. The present study was conducted to assess the association between the model output and other HF biomarkers, risk for HF hospitalization and cardiac mortality, and to compare its performance with two clinical scores: H2FPEF (heavy, hypertensive, atrial fibrillation, pulmonary hypertension, elder, and filling pressure)6Reddy Y.N.V. Carter R.E. Obokata M. et al.A simple, evidence-based approach to help guide diagnosis of heart failure with preserved ejection fraction.Circulation. 2018; 138: 861-870Crossref PubMed Scopus (713) Google Scholar and HFA-PEFF (Heart Failure Association pretest assessment, echocardiography and natriuretic peptide score, functional testing, and final etiology).7Pieske B. Tschope C. de Boer R.A. et al.How to diagnose heart failure with preserved ejection fraction: the HFA-PEFF diagnostic algorithm: a consensus recommendation from the Heart Failure Association (HFA) of the European Society of Cardiology (ESC).Eur Heart J. 2019; 40: 3297-3317Crossref PubMed Scopus (951) Google Scholar This retrospective, multisite study was approved by our institutional review board. The model was developed to classify patients with HFpEF vs individuals without HFpEF (control subjects). Patients with HFpEF were defined according to guidelines and included a diagnosis by the treating physician1Heidenreich P.A. Bozkurt B. Aguilar D. et al.2022 AHA/ACC/HFSA guideline for the Management of heart failure: executive summary: a report of the American College of Cardiology/American Heart Association Joint Committee on clinical practice guidelines.J Am Coll Cardiol. 2022; 79: 1757-1780Crossref PubMed Scopus (392) Google Scholar within 1 year of an echocardiographic examination demonstrating elevated left ventricular filling pressure. Control subjects were patients undergoing clinically indicated echocardiography who lacked these features (Supplemental Appendix). All patients had left ventricular ejection fractions ≥50%. The present analysis used the second version of the AI model (Supplemental Appendix). In the previously described independent test population consisting of 646 patients with HFpEF and 638 control subjects, the updated model produced 95 uncertain outputs (7.4%); in the remaining 607 patients and 582 control subjects, sensitivity was 89.8% (95% CI, 87.5%-92.5%), specificity was 86.3% (95% CI, 83.6%-89.7%), negative predictive value was 89.0% (95% CI, 87.0%-91.7%), and positive predictive value was 87.2% (95% CI, 84.7%-89.8%). Incident HF hospitalization was obtained from electronic health record chart review using standardized definitions, using the first event after the echocardiographic examination. Mortality was obtained from the National Death Index, and causes of cardiac deaths were manually reviewed. End points were adjudicated by investigators blinded to AI analysis results. Cardiac mortality and HF hospitalization were plotted accounting for death as a competing risk. The method of Fine and Gray was used to estimate the hazard ratios (HRs) adjusted for differences in age and sex. Among 1,284 patients followed for a median of 3.4 years (interquartile range, 1.7-6.5 years), there were 252 HF hospitalizations and 540 deaths. Figure 1 demonstrates the risk for HF hospitalization on the basis of HF categorical AI output (top) and quartiles of continuous probability output (bottom). After adjustment for age and sex, positive AI output was associated with a higher risk for HF hospitalization than negative output (HR, 3.76; 95% CI, 2.71-5.21; P < .001) and likewise for uncertain output (HR, 2.79; 95% CI, 1.60-4.62; P < .001). Cardiac deaths (n = 135) were attributable to HF in 63 patients (47%), to coronary artery disease in 55 (41%), to valve disease in five (4%), to arrhythmia in five (4%), and to other causes in seven (5%). Again adjusting for age and sex, cardiac mortality was higher in patients with positive output (HR, 5.55; 95% CI, 3.28-9.37; P < .001); patients with an uncertain output tended to have a higher mortality (HR, 2.22; 95% CI, 0.94-5.24; P = .07). Patients with higher continuous probability outputs demonstrated incrementally higher risk for cardiac mortality (fourth quartile vs first quartile: HR, 11.65; 95% CI, 4.65-29.20; P < .0001). Figure 2 demonstrates the risk for HF hospitalization on the basis of clinical H2FPEF score6Reddy Y.N.V. Carter R.E. Obokata M. et al.A simple, evidence-based approach to help guide diagnosis of heart failure with preserved ejection fraction.Circulation. 2018; 138: 861-870Crossref PubMed Scopus (713) Google Scholar (top). The clinical score differentiated high and low risk for HF hospitalization, but 776 of 1,284 patients (60%) were indeterminate. Application of the AI model to the nondiagnostic H2FPEF outputs (bottom) allowed the classification of all but 68 of the 776 patients (8.8%). The AI model demonstrated a similar relationship between output and risk for HF hospitalization in patients with and those without diagnostic H2FPEF output. Findings were similar when patients were stratified according to HFA-PEFF score. HFA-PEFF score, brain natriuretic peptide, and N-terminal pro–brain natriuretic peptide also differed according to the AI model prediction (Table). Few patients underwent exercise testing; differences in exercise capacity were not significantly different.Figure 2Risk for HF hospitalization was higher in patients with positive H2FPEF outputs (top), but many (n = 776) had nondiagnostic outputs. Application of the AI model was able to reclassify 708 (91%) of the nondiagnostic H2FPEF outputs (bottom).View Large Image Figure ViewerDownload Hi-res image Download (PPT)TableAdditional testing within 1 year of the qualifying echocardiographic studyAI model predictionPNegative (n = 564)Positive (n = 625)Uncertain (n = 95)H2FpEFF category, n (%)<.0001∗Chi-Square P value. Prediction negative161 (28.5)6 (1.0)6 (6.3) Prediction positive50 (8.9)264 (42.2)21 (22.1) Nondiagnostic353 (62.6)355 (56.8)68 (71.6)HFA-PEFF category, n (%)<.0001∗Chi-Square P value. Prediction negative292 (51.8)27 (4.3)26 (27.4) Prediction positive20 (3.5)207 (33.1)11 (11.6) Nondiagnostic252 (44.7)391 (62.6)58 (61.1)BNP, pg/mL<.0001†Kruskal-Wallis P value. Median (IQR)63.0 (22-148)352 (200-636)115 (27-211) n4911410NT-proBNP, pg/mL<.0001†Kruskal-Wallis P value. Median (IQR)210 (74-572)1,941 (697-5,866)675 (206-3,306) n8728426Exercise test workload, METs Mean ± SD9.1 ± 2.87.4 ± 3.211.1 ± 2.2.13‡Analysis of variance P value. n4983Exercise test FAC, % Mean ± SD100.8 ± 28.986.0 ± 28.0128.0 ± 25.5.11‡Analysis of variance P value. n4373BNP, Brain natriuretic peptide; FAC, Functional aerobic capacity; IQR, interquartile range; METs, metabolic equivalents; NT-proBNP, N-terminal pro–brain natriuretic peptide.∗ Chi-Square P value.† Kruskal-Wallis P value.‡ Analysis of variance P value. Open table in a new tab BNP, Brain natriuretic peptide; FAC, Functional aerobic capacity; IQR, interquartile range; METs, metabolic equivalents; NT-proBNP, N-terminal pro–brain natriuretic peptide. In this study we assessed the ability of a novel, HFpEF AI model using a single echocardiographic video clip to identify patients at increased risk for HF hospitalization and cardiac mortality. In summary, (1) positive model output was associated with higher risks for HF hospitalization and cardiac mortality, (2) patients with uncertain outputs demonstrated intermediate risks for these end points, (3) HF hospitalization and cardiac mortality risk were incrementally associated with higher model probability output scores, and (4) the AI model reclassified HF hospitalization risk in nondiagnostic clinical scores, including 91% for H2FPEF outputs and 92% for HFA-PEFF. This is the first AI echocardiographic model to produce outputs discriminating a specific disease (HFpEF) that are incrementally associated with risk for HF hospitalization and cardiac mortality. Prospective studies are required to confirm these retrospective results, to externally validate the AI model's outputs in other echocardiographic laboratories, and to understand the implications for patient management. Studies using a broad representation of HFpEF phenotypes should be undertaken to understand the generalizability of this model in a naturally heterogeneous clinical syndrome. Given her role as JASE Editor-in-Chief, Patricia A. Pellikka, MD, had no involvement in the peer review of this article and has no access to information regarding its peer review. Full responsibility for the editorial process for this article was delegated to Federico M. Asch, MD. Drs. Akerman, Porumb, Hawkes, Woodward, and Upton are employed by Ultromics. Download .pdf (.13 MB) Help with pdf files Supplemental Appendix