Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A bio-ecological model for early screening of developmental coordination disorder.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

To develop and externally validate a bio-ecological model for early screening of developmental coordination disorder (DCD) using maternal and environmental risk factors from electronic health records, aimed at improving early detection in children under 5 years. This was a prospective study that examined data from 150 948 preschool children in China. Perinatal and sociodemographic predictors were integrated using logistic regression and random forest algorithms. The model was internally validated on split training and testing subsets and externally validated on an independent clinical sample of 1359 children aged 3 to 10 years, including confirmed diagnoses of DCD. Model performance was evaluated using the area under the curve (AUC), sensitivity, specificity, and accuracy. In the group aged 3 to 5 years, the model achieved an AUC of 0.70, sensitivity of 71.43%, accuracy of 77.61%, and specificity of 78.00%. In the group aged 6 to 10 years, performance was moderate (AUC = 0.58; sensitivity = 54.88%; accuracy = 61.50%; specificity = 62.28%). This bio-ecological model offers a scalable, cost-effective tool to support the early identification of DCD using electronic health record data. It performs well in early childhood and maintains moderate accuracy in older children, supporting its utility for longer-term risk prediction. The model could enhance existing screening systems by enabling earlier triage and intervention. Further validation across diverse health care settings is warranted.

Similar Papers
  • Abstract
  • 10.1002/alz70855_101257
Detection and prediction of Alzheimer's and related dementia using administrative data
  • Dec 1, 2025
  • Alzheimer's & Dementia
  • Korin Reid

BackgroundThere are more than 5.8 million Americans who suffer from Alzheimer's disease 1. By 2060 13.8 million Americans will be living with Alzheimer's Dementia 2. As of 2022, the financial burden of Alzheimer's Dementia is $321 billion and is expected to reach $1 trillion. 2 Detection and prediction of Alzheimer's and related dementia using widely available data such as electronic health record and electronic claim data could help identify populations at risk for development the disease.MethodWe used electronic health record and claim data to predict Alzheimer's and Related Dementias (ADRD) up to 10 years in advance using electronic health record data from the Mimic‐IV dataset, and electronic claim data form a 5% sample of medicare data from various care settings. For this modeling effort, we used machine learning methods. We also explored techniques to ascertain that models are explainable and work effectively across demographic gorups.ResultUsing the Mimic‐IV dataset, we achieved an AUC of 0.92 predict ADRD up to 10 years in advance. Using a 5% sample of Medicare claim data from diverse healthcare settings, we have achieved an Area Under the Curve (AUC) of 0.86, 0.85, and 0.83 respectively for predicting ADRD 1,3, and 5 years in advance respectively.ConclusionMachine learning methods can be utilized to predict ADRD using widely available administrative data such as electronic health record and electronic claim data.

  • Abstract
  • 10.1002/alz70862_109956
Detection and prediction of Alzheimer's and related dementia using administrative data
  • Dec 1, 2025
  • Alzheimer's & Dementia
  • Korin Reid

BackgroundThere are more than 5.8 million Americans who suffer from Alzheimer’s disease1. By 2060 13.8 million Americans will be living with Alzheimer's Dementia 2. As of 2022, the financial burden of Alzheimer's Dementia is $321 billion and is expected to reach $1 trillion.2 Detection and prediction of Alzheimer's and related dementia using widely available data such as electronic health record and electronic claim data could help identify populations at risk for development the disease.MethodWe used electronic health record and claim data to predict Alzheimer's and Related Dementias (ADRD) up to 10 years in advance using electronic health record data from the Mimic‐IV dataset, and electronic claim data form a 5% sample of medicare data from various care settings. For this modeling effort, we used machine learning methods. We also explored techniques to ascertain that models are explainable and work effectively across demographic gorups.ResultUsing the Mimic‐IV dataset, we achieved an AUC of 0.92 predict ADRD up to 10 years in advance. Using a 5% sample of Medicare claim data from diverse healthcare settings, we have achieved an Area Under the Curve (AUC) of 0.86, 0.85, and 0.83 respectively for predicting ADRD 1,3, and 5 years in advance respectively.ConclusionMachine learning methods can be utilized to predict ADRD using widely available administrative data such as electronic health record and electronic claim data.

  • Research Article
  • Cite Count Icon 60
  • 10.1111/ajt.14099
Big Data, Predictive Analytics, and Quality Improvement in Kidney Transplantation: A Proof of Concept.
  • Jan 4, 2017
  • American Journal of Transplantation
  • T.R Srinivas + 9 more

Big Data, Predictive Analytics, and Quality Improvement in Kidney Transplantation: A Proof of Concept.

  • Research Article
  • 10.1200/jco.2021.39.15_suppl.1510
Augmenting machine learning algorithms to predict mortality using patient-reported outcomes in oncology.
  • May 20, 2021
  • Journal of Clinical Oncology
  • Ravi B Parikh + 9 more

1510 Background: Machine learning (ML) algorithms based on electronic health record (EHR) data have been shown to accurately predict mortality risk among patients with cancer, with areas under the curve (AUC) generally greater than 0.80. While patient-reported outcomes (PROs) may also predict mortality among patients with cancer, it is unclear whether routinely-collected PROs improve the predictive performance of EHR-based ML algorithms. Methods: This cohort study included 8600 patients with cancer who had an outpatient encounter at one of 18 medical oncology practices in a large academic health system between July 1st, 2019 and January 1st, 2020. 4692 (54.9%) patients completed assessments of symptoms, performance status, and quality of life from the PRO version of the Common Terminology Criteria for Adverse Events and the Patient-Reported Outcomes Measurement Information System Global v.1.2 scales. We hypothesized that ML models predicting 180-day all-cause mortality based on EHR + PRO data would improve AUC compared to ML models based on EHR data alone. We assessed univariate and adjusted associations between each PRO and 180-day mortality. To train the EHR-only model, we fit a Least Absolute Shrinkage and Selection Operator (LASSO) regression using 192 EHR demographic, comorbidity, and laboratory variables. To train the EHR + PRO model, we used a two-phase approach to fit a model using EHR data for all patients and PRO data for those who completed assessments. To test our hypothesis, we compared the bootstrapped AUC, area under the precision-recall curve (AUPRC), and sensitivity at a 20% risk threshold for both models. Results: 464 (5.4%) patients died within 180 days of the encounter. Decreased quality of life, functional status, and appetite were associated with greater 180-day mortality (Table). Compared to the EHR-only model, the EHR + PRO model significantly improved AUC (0.86 [95% CI 0.85-0.86] vs. 0.80 [95% CI 0.80-0.81]), AUPRC (0.40 [95% CI 0.37-0.42] vs. 0.30 [95% CI 0.28-0.32]), and sensitivity (0.45 [95% CI 0.42-0.48] vs. 0.33 [95% CI 0.30-0.35]). Conclusions: Routinely collected PROs augment EHR-based ML mortality risk algorithms. ML algorithms based on EHR and PRO data may facilitate earlier supportive care for patients with cancer. Association of PROs with 180-day mortality.[Table: see text]

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 3
  • 10.1371/journal.pone.0312193
Comparing the performance of screening surveys versus predictive models in identifying patients in need of health-related social need services in the emergency department.
  • Nov 20, 2024
  • PloS one
  • Olena Mazurenko + 5 more

Health-related social needs (HRSNs), such as housing instability, food insecurity, and financial strain, are increasingly prevalent among patients. Healthcare organizations must first correctly identify patients with HRSNs to refer them to appropriate services or offer resources to address their HRSNs. Yet, current identification methods are suboptimal, inconsistently applied, and cost prohibitive. Machine learning (ML) predictive modeling applied to existing data sources may be a solution to systematically and effectively identify patients with HRSNs. The performance of ML predictive models using data from electronic health records (EHRs) and other sources has not been compared to other methods of identifying patients needing HRSN services. A screening questionnaire that included housing instability, food insecurity, transportation barriers, legal issues, and financial strain was administered to adult ED patients at a large safety-net hospital in the mid-Western United States (n = 1,101). We identified those patients likely in need of HRSN-related services within the next 30 days using positive indications from referrals, encounters, scheduling data, orders, or clinical notes. We built an XGBoost classification algorithm using responses from the screening questionnaire to predict HRSN needs (screening questionnaire model). Additionally, we extracted features from the past 12 months of existing EHR, administrative, and health information exchange data for the survey respondents. We built ML predictive models with these EHR data using XGBoost (ML EHR model). Out of concerns of potential bias, we built both the screening question model and the ML EHR model with and without demographic features. Models were assessed on the validation set using sensitivity, specificity, and Area Under the Curve (AUC) values. Models were compared using the Delong test. Almost half (41%) of the patients had a positive indicator for a likely HRSN service need within the next 30 days, as identified through referrals, encounters, scheduling data, orders, or clinical notes. The screening question model had suboptimal performance, with an AUC = 0.580 (95%CI = 0.546, 0.611). Including gender and age resulted in higher performance in the screening question model (AUC = 0.640; 95%CI = 0.609, 0.672). The ML EHR models had higher performance. Without including age and gender, the ML EHR model had an AUC = 0.765 (95%CI = 0.737, 0.792). Adding age and gender did not improve the model (AUC = 0.722; 95%CI = 0.744, 0.800). The screening questionnaire models indicated bias with the highest performance for White non-Hispanic patients. The performance of the ML EHR-based model also differed by race and ethnicity. ML predictive models leveraging several robust EHR data sources outperformed models using screening questions only. Nevertheless, all models indicated biases. Additional work is needed to design predictive models for effectively identifying all patients with HRSNs.

  • Research Article
  • 10.1136/bmjsit-2025-000387
Neural network models for predicting readmission among patients undergoing peripheral vascular intervention using electronic health record data and clinical registry data
  • Jun 1, 2025
  • BMJ Surgery, Interventions, & Health Technologies
  • Jialin Mao + 6 more

ObjectivesTo determine whether neural network models based on electronic health record (EHR) data can match and augment the performance of models based on clinical registry data in predicting readmission after peripheral vascular intervention (PVI).DesignObservational cohort study.SettingVascular Quality Initiative registry and INSIGHT Clinical Research Network EHR data from multiple academic institutions in New York City.ParticipantsPatients undergoing PVI during January 1, 2013 to September 30, 2021.Main outcome measuresOur outcome variable was 90-day readmission. We developed logistic regression (LR), multilevel perceptron (MLP), and recurrent neural network (RNN) models using registry alone, EHR data alone, and combined registry-EHR data. EHR data were evaluated using derived variables to match registry variables (EHR-derived data) and clinically meaningful code aggregation (EHR-direct data). Models were evaluated using area under the curve (AUC) for discrimination, Spiegelhalter z score for calibration, and Brier score for overall performance.ResultsThe analytical cohort included 2348 patients undergoing PVI (mean age: 69.9±11.5 years). 832 (35%) patients were readmitted within 90 days. LR to predict 90-day readmission based on registry data alone had an AUC of 0.710, Spiegelhalter z score of 1.021, and Brier score of 0.211. MLP based on registry data alone had similar performance. MLP and RNN based on EHR-direct data (MLP: AUC=0.742, Spiegelhalter z=0.933, Brier=0.204; RNN: AUC=0.737, Spiegelhalter z=1.026, Brier=0.206) and registry+EHR-direct data (MLP: AUC=0.756, Spiegelhalter z=0.794, Brier=0.199; RNN: AUC=0.751, Spiegelhalter z=1.057, Brier=0.200) had improved performances. LR based on EHR-direct data and combined registry+EHR-direct data had worse performances.ConclusionsEHR data, when used with neural network models, can be useful to establish readmission predictive models or augment clinical registry data. EHR-based models can be potentially embedded in the clinical workflow, but model performance may be constrained by the absence of certain information in clinical encounters, such as social determinants of health.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/ichi54592.2022.00032
Identify Cancer Patients at Risk for Heart Failure using Electronic Health Record and Genetic Data
  • Jun 1, 2022
  • Zehao Yu + 6 more

Heart failure is a critical side effect of many cancer treatments. Identifying cancer patients at a high risk of cardiotoxicity before cancer treatment is a critical step towards early detection and possible prevention. This study seeks to examine how genetic data can be used with Electronic Health Record (EHR) data to identify cancer patients at risk of treatment-related heart failure. We explored four machine learning models, including Logistic Regression (LR), Support Vector Machines (SVMs), Random Forest (RF), and Gradient Boost (GB) for heart failure prediction using EHR data linked with genetic data from the UK Biobank. We first identified the best machine learning model using only EHR data and then further added genetic data to the model using three different strategies. The experimental results show that the GB model combining EHR data and genetic data achieved the best area under the curve (AUC) score of 0.7781, outperforming the machine learning models using only EHR data. Among the three strategies of including genetic data, the genome-wide association study (GWAS) method achieved the best performance. Our study shows that genomic data can be used to improve the performance of heart failure prediction among cancer patients.

  • Abstract
  • 10.1016/j.jval.2016.03.179
PND3 - Predictive Model Of Parkinson’s Disease In Large Electronic Health Records Database
  • May 1, 2016
  • Value in Health
  • S Kabadi + 3 more

PND3 - Predictive Model Of Parkinson’s Disease In Large Electronic Health Records Database

  • Research Article
  • Cite Count Icon 3
  • 10.1136/lupus-2024-001170
Development and validation of a risk scoring system to identify patients with lupus nephritis in electronic health record data
  • May 1, 2024
  • Lupus Science & Medicine
  • Zara Izadi + 4 more

ObjectiveAccurate identification of lupus nephritis (LN) cases is essential for patient management, research and public health initiatives. However, LN diagnosis codes in electronic health records (EHRs) are underused, hindering efficient...

  • Research Article
  • Cite Count Icon 2
  • 10.1136/bmjopen-2024-097016
Predictive machine-learning model for screening iron deficiency without anaemia: a retrospective cohort study
  • Aug 1, 2025
  • BMJ Open
  • Orly Efros + 6 more

ObjectivesThis study aimed to develop and validate a machine-learning (ML) model to predict iron deficiency without anaemia (IDWA) using routinely collected electronic health record (EHR) data. The primary hypothesis was that an ML model could achieve better accuracy in identifying low ferritin levels (<30 ng/mL) in non-anaemic patients compared with traditional methods.Design A retrospective cohort study.SettingData were derived from secondary and tertiary care facilities within the eight-hospital Mount Sinai Health System, an urban academic health system.ParticipantsThe study included 211 486 adult patients (aged ≥18 years) with normal haemoglobin levels (≥130 g/L for men and ≥120 g/L for women) and recorded ferritin measurements.Primary and secondary outcome measuresThe primary outcome was the prediction of low ferritin levels (<30 ng/mL) using extreme gradient-boosted decision trees, an ML algorithm suited for structured clinical data. Secondary outcomes included subgroup analyses stratified by sex and age to evaluate model performance in different populations.Data from 211 486 Mount Sinai Health System patients with normal haemoglobin levels and ferritin testing were analysed. The model used demographic data, blood count indices and chemistry results to identify low ferritin levels (<30 ng/mL).ResultsOf the 211 486 patients analysed, 19.56% (n=41 368) of the patients had low ferritin levels. In the low ferritin group, the mean age was 41.28 years with 89.64% females. In contrast, the normal ferritin group had a mean age of 50.14 years with 62.02% females. The model achieved an area under the curve (AUC) of 0.814. At a sensitivity threshold of 70%, the model had a specificity of 75.85%, with a positive predictive value of 37.6% and a negative predictive value of 92.41%. The model outperformed an alternative model based only on complete blood count indices (AUC 0.814 vs 0.741). Subgroup analysis showed that model accuracy varied by sex and age, with lower performance in premenopausal women (AUC 0.736) compared with postmenopausal women (AUC 0.793) and men (AUC of 0.832 in those under 60 years and 0.806 in those aged 60 and above).ConclusionsThe ML model provides an effective approach to screening for IDWA using readily available EHR data. Implementing this tool in clinical settings may facilitate early diagnosis of IDWA.

  • Research Article
  • Cite Count Icon 118
  • 10.1177/1932296815620200
Reverse Engineering and Evaluation of Prediction Models for Progression to Type 2 Diabetes: An Application of Machine Learning Using Electronic Health Records.
  • Dec 20, 2015
  • Journal of Diabetes Science and Technology
  • Jeffrey P Anderson + 10 more

Application of novel machine learning approaches to electronic health record (EHR) data could provide valuable insights into disease processes. We utilized this approach to build predictive models for progression to prediabetes and type 2 diabetes (T2D). Using a novel analytical platform (Reverse Engineering and Forward Simulation [REFS]), we built prediction model ensembles for progression to prediabetes or T2D from an aggregated EHR data sample. REFS relies on a Bayesian scoring algorithm to explore a wide model space, and outputs a distribution of risk estimates from an ensemble of prediction models. We retrospectively followed 24 331 adults for transitions to prediabetes or T2D, 2007-2012. Accuracy of prediction models was assessed using an area under the curve (AUC) statistic, and validated in an independent data set. Our primary ensemble of models accurately predicted progression to T2D (AUC = 0.76), and was validated out of sample (AUC = 0.78). Models of progression to T2D consisted primarily of established risk factors (blood glucose, blood pressure, triglycerides, hypertension, lipid disorders, socioeconomic factors), whereas models of progression to prediabetes included novel factors (high-density lipoprotein, alanine aminotransferase, C-reactive protein, body temperature; AUC = 0.70). We constructed accurate prediction models from EHR data using a hypothesis-free machine learning approach. Identification of established risk factors for T2D serves as proof of concept for this analytical approach, while novel factors selected by REFS represent emerging areas of T2D research. This methodology has potentially valuable downstream applications to personalized medicine and clinical research.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 28
  • 10.1186/s12874-020-00923-1
The application of unsupervised deep learning in predictive models using electronic health records
  • Feb 26, 2020
  • BMC Medical Research Methodology
  • Lei Wang + 4 more

BackgroundThe main goal of this study is to explore the use of features representing patient-level electronic health record (EHR) data, generated by the unsupervised deep learning algorithm autoencoder, in predictive modeling. Since autoencoder features are unsupervised, this paper focuses on their general lower-dimensional representation of EHR information in a wide variety of predictive tasks.MethodsWe compare the model with autoencoder features to traditional models: logistic model with least absolute shrinkage and selection operator (LASSO) and Random Forest algorithm. In addition, we include a predictive model using a small subset of response-specific variables (Simple Reg) and a model combining these variables with features from autoencoder (Enhanced Reg). We performed the study first on simulated data that mimics real world EHR data and then on actual EHR data from eight Advocate hospitals.ResultsOn simulated data with incorrect categories and missing data, the precision for autoencoder is 24.16% when fixing recall at 0.7, which is higher than Random Forest (23.61%) and lower than LASSO (25.32%). The precision is 20.92% in Simple Reg and improves to 24.89% in Enhanced Reg. When using real EHR data to predict the 30-day readmission rate, the precision of autoencoder is 19.04%, which again is higher than Random Forest (18.48%) and lower than LASSO (19.70%). The precisions for Simple Reg and Enhanced Reg are 18.70 and 19.69% respectively. That is, Enhanced Reg can have competitive prediction performance compared to LASSO. In addition, results show that Enhanced Reg usually relies on fewer features under the setting of simulations of this paper.ConclusionsWe conclude that autoencoder can create useful features representing the entire space of EHR data and which are applicable to a wide array of predictive tasks. Together with important response-specific predictors, we can derive efficient and robust predictive models with less labor in data extraction and model training.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 11
  • 10.3389/fvets.2023.1189157
Machine learning-based risk prediction model for canine myxomatous mitral valve disease using electronic health record data.
  • Aug 31, 2023
  • Frontiers in Veterinary Science
  • Yunji Kim + 5 more

Myxomatous mitral valve disease (MMVD) is the most common cause of heart failure in dogs, and assessing the risk of heart failure in dogs with MMVD is often challenging. Machine learning applied to electronic health records (EHRs) is an effective tool for predicting prognosis in the medical field. This study aimed to develop machine learning-based heart failure risk prediction models for dogs with MMVD using a dataset of EHRs. A total of 143 dogs with MMVD between May 2018 and May 2022. Complete medical records were reviewed for all patients. Demographic data, radiographic measurements, echocardiographic values, and laboratory results were obtained from the clinical database. Four machine-learning algorithms (random forest, K-nearest neighbors, naïve Bayes, support vector machine) were used to develop risk prediction models. Model performance was represented by plotting the receiver operating characteristic (ROC) curve and calculating the area under the curve (AUC). The best-performing model was chosen for the feature-ranking process. The random forest model showed superior performance to the other models (AUC = 0.88), while the performance of the K-nearest neighbors model showed the lowest performance (AUC = 0.69). The top three models showed excellent performance (AUC ≥ 0.8). According to the random forest algorithm's feature ranking, echocardiographic and radiographic variables had the highest predictive values for heart failure, followed by packed cell volume (PCV) and respiratory rates. Among the electrolyte variables, chloride had the highest predictive value for heart failure. These machine-learning models will enable clinicians to support decision-making in estimating the prognosis of patients with MMVD.

  • Research Article
  • Cite Count Icon 4
  • 10.1159/000518187
The Development and Use of an EHR-Linked Database for Glomerular Disease Research and Quality Initiatives
  • Jul 5, 2021
  • Glomerular Diseases
  • Richard N Eikstadt + 11 more

Background and Objective: The use of electronic health record (EHR) data can facilitate efficient research and quality initiatives. The imprecision of ICD-10 codes for kidney diagnoses has been an obstacle to discrete data-defined diagnoses in the EHR. This manuscript describes the Kidney Research Network (KRN) registry and database that provide an example of a prospective, real-world data glomerular disease registry for research and quality initiatives. Methods: KRN is a multicenter collaboration of patients, physicians, and scientists across diverse health-care settings with a focus on improving treatment options and outcomes for patients with glomerular disease. The registry and data warehouse amasses retrospective and prospective data including EHR, active research study, completed clinical trials, patient reported outcomes, and other relevant data. Following consent, participating sites enter the patient into KRN and provide a physician-confirmed primary kidney diagnosis. Kidney biopsy reports are redacted and uploaded. Site programmers extract local EHR data including demographics, insurance type, zip code, diagnoses, encounters, laboratories, procedures, medications, dialysis/transplant status, vitals, and vital status monthly. Participating sites transform data to conform to a common data model prior to submitting to the Data Analysis and Coordinating Center (DACC). The DACC stores and reviews each site’s EHR data for quality before loading into the KRN database. Results: As of January 2021, 1,192 patients have enrolled in the registry. The database has been utilized for research, clinical trial design, clinical trial end point validation, and supported quality initiatives. The data also support a dashboard allowing enrolling sites to assist with clinical trial enrollment and population health initiatives. Conclusion: A multicenter registry using EHR data, following physician- and biopsy-confirmed glomerular disease diagnosis, can be established and used effectively for research and quality initiatives. This design provides an example which may be readily replicated for other rare or common disease endeavors.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.jstrokecerebrovasdis.2023.107255
Development and validation of a model predicting mild stroke severity on admission using electronic health record data
  • Jul 18, 2023
  • Journal of Stroke and Cerebrovascular Diseases
  • Kimberly J Waddell + 9 more

Development and validation of a model predicting mild stroke severity on admission using electronic health record data

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant