Stroke Outcome Measurements From Electronic Medical Records: Cross-sectional Study on the Effectiveness of Neural and Nonneural Classifiers.

Bruna Stella Zanotto,Ana Paula Beck Da Silva Etges,Ana Claudia De Souza,Renata Vieira,Marcos André Gonçalves,Claudio M V Andrade,Renata Ruschel,Avner Dal Bosco,Sergio Canuto,Eduardo Gabriel Cortes,Felipe Viegas,Washington Luiz,Carisi Polanczyk,Sheila Ouriques Martins

doi:10.2196/29120

Abstract

BackgroundWith the rapid adoption of electronic medical records (EMRs), there is an ever-increasing opportunity to collect data and extract knowledge from EMRs to support patient-centered stroke management.ObjectiveThis study aims to compare the effectiveness of state-of-the-art automatic text classification methods in classifying data to support the prediction of clinical patient outcomes and the extraction of patient characteristics from EMRs.MethodsOur study addressed the computational problems of information extraction and automatic text classification. We identified essential tasks to be considered in an ischemic stroke value-based program. The 30 selected tasks were classified (manually labeled by specialists) according to the following value agenda: tier 1 (achieved health care status), tier 2 (recovery process), care related (clinical management and risk scores), and baseline characteristics. The analyzed data set was retrospectively extracted from the EMRs of patients with stroke from a private Brazilian hospital between 2018 and 2019. A total of 44,206 sentences from free-text medical records in Portuguese were used to train and develop 10 supervised computational machine learning methods, including state-of-the-art neural and nonneural methods, along with ontological rules. As an experimental protocol, we used a 5-fold cross-validation procedure repeated 6 times, along with subject-wise sampling. A heatmap was used to display comparative result analyses according to the best algorithmic effectiveness (F1 score), supported by statistical significance tests. A feature importance analysis was conducted to provide insights into the results.ResultsThe top-performing models were support vector machines trained with lexical and semantic textual features, showing the importance of dealing with noise in EMR textual representations. The support vector machine models produced statistically superior results in 71% (17/24) of tasks, with an F1 score >80% regarding care-related tasks (patient treatment location, fall risk, thrombolytic therapy, and pressure ulcer risk), the process of recovery (ability to feed orally or ambulate and communicate), health care status achieved (mortality), and baseline characteristics (diabetes, obesity, dyslipidemia, and smoking status). Neural methods were largely outperformed by more traditional nonneural methods, given the characteristics of the data set. Ontological rules were also effective in tasks such as baseline characteristics (alcoholism, atrial fibrillation, and coronary artery disease) and the Rankin scale. The complementarity in effectiveness among models suggests that a combination of models could enhance the results and cover more tasks in the future.ConclusionsAdvances in information technology capacity are essential for scalability and agility in measuring health status outcomes. This study allowed us to measure effectiveness and identify opportunities for automating the classification of outcomes of specific tasks related to clinical conditions of stroke victims, and thus ultimately assess the possibility of proactively using these machine learning techniques in real-world situations.

Highlights

BackgroundStroke is the second leading cause of mortality and disability-adjusted life years globally [1,2]
Advances in information technology capacity are essential for scalability and agility in measuring health status outcomes
For reasons of practical application and as a research exercise, as a secondary analysis, we compared the OWL technique with the machine learning (ML) model ranked as the best based on the Friedman test. This analysis allowed us to identify the Discussions with experts in the stroke care pathway allowed us to define 30 tasks that were considered feasible to extract from electronic medical records (EMRs)

Summary

Introduction

BackgroundStroke is the second leading cause of mortality and disability-adjusted life years globally [1,2]. There has been an increasing interest in the use of automated machine learning (ML) techniques to track stroke outcomes, with the hope that such methods could make use of large, routinely collected data sets and deliver accurate, personalized prognoses [3]. Few studies have addressed the unstructured textual portion of electronic medical records (EMRs) as the primary source of information. The information technology (IT) gap between automated data collection from EMRs and improving the quality of care has been described in the literature as a decelerator of value initiatives [15,16,17,18]. With the rapid adoption of electronic medical records (EMRs), there is an ever-increasing opportunity to collect data and extract knowledge from EMRs to support patient-centered stroke management

Objectives

Methods

Results

Discussion

Conclusion