Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

The Natural Language of Finance

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

We summarize the wide array of natural language processing (NLP) tools used in financial economics research. These tools empower researchers to incorporate rich but subjective textual data into advanced empirical analysis. NLP tools have pros and cons, and some are better suited to certain research agendas. Research using these tools has exploded in prevalence over the past ten years, and we document the major contributions in corporate finance, asset pricing, and beyond. These tools offer the flexibility to test hypotheses that were not possible before their advent, while also offering improvements in the clarity of identification and the ability to separate hypotheses that purport to explain a set of findings. Finally, we identify challenges and directions for future work.

Similar Papers
  • Discussion
  • Cite Count Icon 29
  • 10.1161/circoutcomes.115.002125
Natural Language Processing and the Promise of Big Data: Small Step Forward, but Many Miles to Go.
  • Aug 18, 2015
  • Circulation: Cardiovascular Quality and Outcomes
  • Thomas M Maddox + 1 more

The promise of big data has captured healthcare’s imagination. Although the term lacks a consensus definition, it generally refers to electronic health data sets characterized by the 3 Vs: volume, variety, and velocity.1,2 Volume refers to the sheer amount of healthcare data currently generated by clinical operations, administration, and patients themselves. By one estimate, ≈25 000 petabytes of healthcare data will be available by 2020—an amount that could fill 500 billion file cabinets.2 Variety refers to the wide range of healthcare data formats. For example, electronic health records (EHRs) contain both structured and unstructured (or free-text) data, diagnostic images come in a variety of multimedia formats, and patient data are generated from wearables, mobile devices, medical devices, and social media—each with its own format. Velocity refers to the rapidity with which new data are generated, and thus the speed at which it needs incorporation into data sets and analyses to provide real-time insights into health care. Article see p 477 The potential of such data is enormous. Insights from big data could fuel innovation and improvement in clinical operations, research and development, and public health.1 However, the potential of big data to realize these lofty aspirations is matched by the challenge of organizing, analyzing, and generating actionable insights from it. One of the biggest challenges in realizing the potential of big data is in abstracting it. With the passage of the HITECH (The Health Information Technology for Economic and Clinical Health) Act in 2009, the adoption of EHRs in clinical practice has accelerated, and now over half of office-based practices and hospitals are using some form of EHR.3,4 As a result, more point-of-care clinical data, previously inaccessible in its paper format, is potentially available. However, the variety aspect of EHR data—its mix …

  • Research Article
  • 10.1016/j.nlp.2025.100187
Trusted knowledge extraction for operations and maintenance intelligence
  • Dec 1, 2025
  • Natural Language Processing Journal
  • Kathleen P Mealey + 4 more

• Released OMIn, the first open-source dataset for maintenance knowledge extraction. • Evaluated 16 NLP tools for zero-shot performance on real-world maintenance data. • Identified domain-specific challenges limiting NLP tool accuracy and robustness. • Assessed tool readiness and trust for mission-critical maintenance applications. Deriving operational intelligence from organizational data repositories is a key challenge due to the dichotomy of data confidentiality vs data integration objectives, as well as the limitations of Natural Language Processing (NLP) tools relative to the specific knowledge structure of domains such as operations and maintenance. In this work, we discuss Knowledge Graph construction and break down the Knowledge Extraction process into its Named Entity Recognition, Coreference Resolution, Named Entity Linking, and Relation Extraction functional components. We then evaluate sixteen NLP tools in concert with or in comparison to the rapidly advancing capabilities of Large Language Models (LLMs). We focus on the operational and maintenance intelligence use case for trusted applications in the aircraft industry. A baseline dataset is derived from a rich public domain US Federal Aviation Administration dataset focused on equipment failures or maintenance requirements. We assess the zero-shot performance of NLP and LLM tools that can be operated within a controlled, confidential environment (no data is sent to third parties). Based on our observation of significant performance limitations, we discuss the challenges related to trusted NLP and LLM tools as well as their Technical Readiness Level for wider use in mission-critical industries such as aviation. We conclude with recommendations to enhance trust and provide our open-source curated dataset to support further baseline testing and evaluation.

  • Research Article
  • Cite Count Icon 52
  • 10.2196/jmir.4612
Automatically Detecting Failures in Natural Language Processing Tools for Online Community Text
  • Aug 31, 2015
  • Journal of Medical Internet Research
  • Albert Park + 4 more

BackgroundThe prevalence and value of patient-generated health text are increasing, but processing such text remains problematic. Although existing biomedical natural language processing (NLP) tools are appealing, most were developed to process clinician- or researcher-generated text, such as clinical notes or journal articles. In addition to being constructed for different types of text, other challenges of using existing NLP include constantly changing technologies, source vocabularies, and characteristics of text. These continuously evolving challenges warrant the need for applying low-cost systematic assessment. However, the primarily accepted evaluation method in NLP, manual annotation, requires tremendous effort and time.ObjectiveThe primary objective of this study is to explore an alternative approach—using low-cost, automated methods to detect failures (eg, incorrect boundaries, missed terms, mismapped concepts) when processing patient-generated text with existing biomedical NLP tools. We first characterize common failures that NLP tools can make in processing online community text. We then demonstrate the feasibility of our automated approach in detecting these common failures using one of the most popular biomedical NLP tools, MetaMap.MethodsUsing 9657 posts from an online cancer community, we explored our automated failure detection approach in two steps: (1) to characterize the failure types, we first manually reviewed MetaMap’s commonly occurring failures, grouped the inaccurate mappings into failure types, and then identified causes of the failures through iterative rounds of manual review using open coding, and (2) to automatically detect these failure types, we then explored combinations of existing NLP techniques and dictionary-based matching for each failure cause. Finally, we manually evaluated the automatically detected failures.ResultsFrom our manual review, we characterized three types of failure: (1) boundary failures, (2) missed term failures, and (3) word ambiguity failures. Within these three failure types, we discovered 12 causes of inaccurate mappings of concepts. We used automated methods to detect almost half of 383,572 MetaMap’s mappings as problematic. Word sense ambiguity failure was the most widely occurring, comprising 82.22% of failures. Boundary failure was the second most frequent, amounting to 15.90% of failures, while missed term failures were the least common, making up 1.88% of failures. The automated failure detection achieved precision, recall, accuracy, and F1 score of 83.00%, 92.57%, 88.17%, and 87.52%, respectively.ConclusionsWe illustrate the challenges of processing patient-generated online health community text and characterize failures of NLP tools on this patient-generated health text, demonstrating the feasibility of our low-cost approach to automatically detect those failures. Our approach shows the potential for scalable and effective solutions to automatically assess the constantly evolving NLP tools and source vocabularies to process patient-generated text.

  • Research Article
  • Cite Count Icon 4
  • 10.1145/3711091
Thoughtful Adoption of NLP for Civic Participation: Understanding Differences Among Policymakers
  • May 2, 2025
  • Proceedings of the ACM on Human-Computer Interaction
  • Jose A Guridi + 2 more

Natural language processing (NLP) tools have the potential to boost civic participation and enhance democratic processes because they can significantly increase governments' capacity to gather and analyze citizen opinions. However, their adoption in government remains limited, and harnessing their benefits while preventing unintended consequences remains a challenge. While prior work has focused on improving NLP performance, this work examines how different internal government stakeholders influence NLP tools' thoughtful adoption. We interviewed seven politicians (politically appointed officials as heads of government institutions) and thirteen public servants (career government employees who design and administrate policy interventions), inquiring how they choose whether and how to use NLP tools to support civic participation processes. The interviews suggest that policymakers across both groups focused on their needs for career advancement and the need to showcase the legitimacy and fairness of their work when considering NLP tool adoption and use. Because these needs vary between politicians and public servants, their preferred NLP features and tool designs also differ. Interestingly, despite their differing needs and opinions, neither group clearly identifies who should advocate for NLP adoption to enhance civic participation or address the unintended consequences of a poorly considered adoption. This lack of clarity in responsibility might have caused the governments' low adoption of NLP tools. We discuss how these findings reveal new insights for future HCI research. They inform the design of NLP tools for increasing civic participation efficiency and capacity, the design of other tools and methods that ensure thoughtful adoption of AI tools in government, and the design of NLP tools for collaborative use among users with different incentives and needs.

  • Research Article
  • Cite Count Icon 4
  • 10.1002/cae.22778
The underlying potential of NLP for microcontroller programming education
  • Aug 14, 2024
  • Computer Applications in Engineering Education
  • André Rocha + 3 more

The trend for an increasingly ubiquitous and cyber‐physical world has been leveraging the use and importance of microcontrollers (C) to unprecedented levels. Therefore, microcontroller programming (CP) becomes a paramount skill for electrical and computer engineering students. However, CP poses significant challenges for undergraduate students, given the need to master low‐level programming languages and several algorithmic strategies that are not usual in “generic” programming. Moreover, CP can be time‐consuming and complex even when using high‐level languages. This article samples the current state of CP education in Portugal and unveils the potential support of natural language processing (NLP) tools (such as chatGPT). Our analysis of CP curricular units from seven representative Portuguese engineering schools highlights a predominant use of AVR 8‐bit C and project‐based learning. While NLP tools emerge as strong candidates as students' C companion, their application and impact on the learning process and outcomes deserve to be understood. This study compares the most prominent NLP tools, analyzing their benefits and drawbacks for CP education, building on both hands‐on tests and literature reviews. By providing automatic code generation and explanation of concepts, NLP tools can assist students in their learning process, allowing them to focus on software design and real‐world tasks that the C is designed to handle, rather than on low‐level coding. We also analyzed the specific impact of chatGTP in the context of a CP course at ISEP, confirming most of our expectations, but with a few curiosities. Overall, this work establishes the foundations for future research on the effective integration of NLP tools in CP courses.

  • Research Article
  • Cite Count Icon 12
  • 10.1007/s11135-020-01067-6
What’s in a text? Bridging the gap between quality and quantity in the digital era
  • Nov 11, 2020
  • Quality & Quantity
  • Roberto Franzosi

The digital era has not only given us the world of big data, but the tools to deal with this mostly unstructured, mostly textual data: Natural Language Processing (NLP) tools. Yet, most humanists and social scientists do not work with big data. They do not deal with millions of documents. Literary critics’ corpora are the handful of works produced by an author. The historians’ primary source documents number in the tens, perhaps hundreds. Social scientists deal with tens or hundreds of transcripts of focus groups and in-depth interviews, or at most a few thousand media articles. And they analyze these data either qualitatively or quantitatively with a variety of manual or computer-assisted methodologies, from content analysis to frame analysis, discourse analysis, quantitative narrative analysis. But, once developed, at least some of the NLP tools of automatic textual analysis and the data analytics visualization tools, can be applied not just to big data but to small data as well. This paper illustrates how some of these tools can be used by focusing on a short first-person narrative. And the NLP tools reveal patterns of language use perhaps not immediately discernible, thus proving useful in the analysis of even small data. But understanding and interpreting these patterns requires knowledge way beyond the NLP tools themselves. Humanists and social scientists need not fear computer scientists; rather, they need to learn to take advantage of them. NLP tools lay a bridge between quality and quantity, with much to be gained from a constant interaction between distant and close reading.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/iacs.2014.6841958
Using Continuous Integration to organize and monitor the annotation process of domain specific corpora
  • Apr 1, 2014
  • Marc Schreiber + 2 more

Applications in the World Wide Web aggregate vast amounts of information from different data sources. The aggregation process is often implemented with Extract, Transform and Load (ETL) processes. Usually ETL processes require information for aggregation available in structured formats, e. g. XML or JSON. In many cases the information is provided in natural language text which makes the application of ETL processes impractical. Due to the fact that information is provided in natural language, Information Extraction (IE) systems have been evolved. They make use of Natural Language Processing (NLP) tools to derive meaning from natural language text. State-of-the-art NLP tools apply Machine Learning methods. These NLP tools perform on newspapers with good quality, but they drop accuracy in other domains. However, to improve the quality for IE systems in specific domains often NLP tools are trained on domain specific text which is a time consuming process. This paper introduces an approach using a Continuous Integration pipeline for organizing and monitoring the annotation process on domain specific corpora.

  • Book Chapter
  • 10.1163/9789004260122_003
Computational Resources and Tools for Latin
  • Jan 1, 2014
  • Barbara Mcgillivray

In spite of important precedent, after an early phase computational and corpus research has been focusing on modern languages, especially English. Krauwer (1998) introduces idea of a Basic Language Resource Kit (BLARK), which he defines as the minimal set of language resources that is necessary to do any precompetitive research and education at all. The first step towards developing computational resources for Latin consists of collecting textual material in a format that can be handled by Natural Language Processing (NLP) tools. Those texts were not originally produced digitally, which means that a conversion from paper to electronic format is necessary. Morphological features of words can be included either through a manual annotation or through automatic methods. The chapter presents a brief overview on existing computational resources and tools for Latin covered annotated corpora, various NLP tools, and lexical databases.Keywords: annotating morphology; Basic Language Resource Kit (BLARK); computational resources; Krauwer; Latin corpora; Natural Language Processing (NLP) tools; semantic resources

  • Research Article
  • Cite Count Icon 35
  • 10.1017/s0261444812000547
Advancing research in second language writing through computational tools and machine learning techniques: A research agenda
  • Feb 22, 2013
  • Language Teaching
  • Scott A Crossley

This paper provides an agenda for replication studies focusing on second language (L2) writing and the use of natural language processing (NLP) tools and machine learning algorithms. Specifically, it introduces a range of the available NLP tools and machine learning algorithms and demonstrates how these could be used to replicate seminal studies in L2 writing that concentrate on longitudinal writing development, predicting essay quality, examining differences between L1 and L2 writers, the effects of writing topics, and the effects of writing tasks. The paper concludes with implications for the recommended replication studies in the field of L2 writing and the advantages of using NLP tools and machine learning algorithms.

  • Research Article
  • Cite Count Icon 1
  • 10.1200/jco.2022.40.6_suppl.072
Diagnosis codes overestimate the burden of prostate cancer cases.
  • Feb 20, 2022
  • Journal of Clinical Oncology
  • Tori Anglin-Foote + 5 more

72 Background: Identifying cancer cases within the electronic health record (EHR) or claims data can be challenging because diagnosis codes are often entered into patient records during routine screenings or as “rule out” diagnosis codes when the patient is referred to a procedure. To improve accuracy of prostate cancer (PCa) case ascertainment, we compared algorithms that used diagnoses codes to natural language processing (NLP) tools applied to clinical notes and pathology reports to identify Veterans with prostate cancer (PCa). Methods: This is a retrospective observational cohort study using VA EHR data to identify veterans diagnosed with PCa between 2000 and 2020. Using International Classification of Diseases (ICD-10 CM or ICD-9 CM) diagnosis and procedure codes, we identified veterans who may have PCa. We deployed validated NLP tools to identify the presence of Gleason score, metastatic PCa, and castration sensitivity to identify evidence of PCa within the notes. We conducted a descriptive analysis to compare the results of algorithms that relied exclusively on diagnosis codes compared to use of NLP tools. Results: From 2000 through 2020,1,031,296 veterans had one or more PCa diagnosis code. This number decreased by 11% for each additional PCa diagnosis code required. When we required 4 or more PCa diagnosis codes to be present, only 746,350 veterans had PCa. When we deployed NLP tools to identify mention of a Gleason score or an indicator of mPCa, only 685,847 Veterans had these indicators of PCa, a 35% decrease in the number of PCa cases with a single diagnosis code. Chart review of patients with their first PCa diagnosis codes in 2019 and 4 or more codes in their records illustrated no evidence of Gleason score or mPCa disease in their EHR. Analysis of their pathology reports revealed that these patients had prostatic intraepithelial neoplasia or atypical small acinar proliferation and had not yet developed prostate cancer. Conclusions: Accurate ascertainment of PCa using EHR and claims data requires using NLP tools and clinical notes combined with structured data sources such as diagnosis codes. Relying on ICD diagnosis codes alone will overestimate the burden of PCa up to 30%.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.2196/45534
Understanding Views Around the Creation of a Consented, Donated Databank of Clinical Free Text to Develop and Train Natural Language Processing Models for Research: Focus Group Interviews With Stakeholders
  • May 3, 2023
  • JMIR Medical Informatics
  • Natalie K Fitzpatrick + 6 more

BackgroundInformation stored within electronic health records is often recorded as unstructured text. Special computerized natural language processing (NLP) tools are needed to process this text; however, complex governance arrangements make such data in the National Health Service hard to access, and therefore, it is difficult to use for research in improving NLP methods. The creation of a donated databank of clinical free text could provide an important opportunity for researchers to develop NLP methods and tools and may circumvent delays in accessing the data needed to train the models. However, to date, there has been little or no engagement with stakeholders on the acceptability and design considerations of establishing a free-text databank for this purpose.ObjectiveThis study aimed to ascertain stakeholder views around the creation of a consented, donated databank of clinical free text to help create, train, and evaluate NLP for clinical research and to inform the potential next steps for adopting a partner-led approach to establish a national, funded databank of free text for use by the research community.MethodsWeb-based in-depth focus group interviews were conducted with 4 stakeholder groups (patients and members of the public, clinicians, information governance leads and research ethics members, and NLP researchers).ResultsAll stakeholder groups were strongly in favor of the databank and saw great value in creating an environment where NLP tools can be tested and trained to improve their accuracy. Participants highlighted a range of complex issues for consideration as the databank is developed, including communicating the intended purpose, the approach to access and safeguarding the data, who should have access, and how to fund the databank. Participants recommended that a small-scale, gradual approach be adopted to start to gather donations and encouraged further engagement with stakeholders to develop a road map and set of standards for the databank.ConclusionsThese findings provide a clear mandate to begin developing the databank and a framework for stakeholder expectations, which we would aim to meet with the databank delivery.

  • Book Chapter
  • Cite Count Icon 1
  • 10.14705/rpnet.2019.38.1036
SimpleApprenant: a platform to improve French L2 learners’ knowledge of multiword expressions
  • Dec 9, 2019
  • Amalia Todirascu + 1 more

We present SimpleApprenant, a platform aiming to improve French L2 learners’ knowledge of Multi Word Expressions (MWEs). SimpleApprenant integrates an MWE database annotated with the Common European Framework of Reference for languages (CEFR) level and several Natural Language Processing (NLP) tools: a spelling checker, a parser, and a set of transformation rules. NLP tools and resources are used to build training and writing exercises to improve MWE knowledge and writing skills of French L2 learners. We present the user scenarios, the platform’s architecture, as well as the preliminary evaluation of its NLP tools.

  • Conference Article
  • Cite Count Icon 1
  • 10.3115/v1/w14-5205
Significance of Bridging Real-world Documents and NLP Technologies
  • Jan 1, 2014
  • Tadayoshi Hara + 3 more

Most conventional natural language processing (NLP) tools assume plain text as their input, whereas real-world documents display text more expressively, using a variety of layouts, sentence structures, and inline objects, among others. When NLP tools are applied to such text, users must first convert the text into the input/output formats of the tools. Moreover, this awkwardly obtained input typically does not allow the expected maximum performance of the NLP tools to be achieved. This work attempts to raise awareness of this issue using XML documents, where textual composition beyond plain text is given by tags. We propose a general framework for data conversion between XML-tagged text and plain text used as input/output for NLP tools and show that text sequences obtained by our framework can be much more thoroughly and efficiently processed by parsers than naively tag-removed text. These results highlight the significance of bridging real-world documents and NLP technologies.

  • Research Article
  • 10.1055/a-2796-1975
Optimizing the Accuracy of Natural Language Processing Tools for Pulmonary Embolism Detection Through Integration with Claims Data: The PE-EHR+ Study.
  • Feb 9, 2026
  • Thrombosis and haemostasis
  • Sina Rashedi + 20 more

Rule-based natural language processing (NLP) tools can identify pulmonary embolism (PE) via radiology reports. However, their external validity remains uncertain.In this cross-sectional study, 1,712 hospitalized patients (with and without PE) at Mass General Brigham (MGB) hospitals (2016-2021) were analyzed. Two previously published NLP algorithms were applied to radiology reports to identify PE. Chart review by two physicians was the reference standard. We tested three approaches: (A) NLP applied to all patients; (B) NLP limited to radiology reports of patients with principal or secondary International Classification of Diseases 10th revision (ICD-10) PE discharge codes; and (C) NLP applied to patients with PE discharge codes or a Present-on-Admission (POA) indicator ("Y") for PE. All others were assumed PE-negative in Approaches B and C to minimize NLP false positives. Weighted estimates were derived from the MGB hospitalized cohort (n = 381,642) to calculate F1 scores (as the harmonic mean of sensitivity and positive predictive value [PPV]).In Approach A, both NLP tools showed high sensitivity (82.5%, 93.0%) and specificity (98.9%, 98.7%) but low PPV (60.3%, 59.6%). Approach B improved PPV (95.2%, 94.9%) but reduced sensitivity (74.1%, 76.2%), while Approach C preserved both high sensitivity (82.5%, 93.0%) and PPV (95.6%, 95.8%). Approach C demonstrated the best performance, yielding significantly higher F1 scores for both NLP tools (88.6%, 94.4%) compared with Approach A (69.7%, 72.6%) and Approach B (83.3%, 84.5%) (P < 0.001).The accuracy of PE detection improves when rule-based NLP algorithms are operationalized using administrative claims data in addition to radiology reports.

  • Research Article
  • 10.1016/j.jvsvi.2025.100302
Using natural language processing to extract carotid stenosis severity from clinical notes to create a nationwide veteran cohort
  • Oct 31, 2025
  • JVS-vascular insights
  • Kyung Min Lee + 17 more

Objective:The prevalence of moderate to severe asymptomatic carotid stenosis (ie, atherosclerotic narrowing of the extracranial carotid arteries) is generally approximately 6% and 2%, respectively. Most prior studies of carotid stenosis risk factors have been small. This study describes the development and validation of a natural language processing (NLP) tool to identify carotid stenosis and uses it to identify significant risk factors, presence, and severity of carotid stenosis.Methods:We created an NLP tool to extract the ratio of peak systolic velocity of the internal carotid artery to the common carotid artery (ICA/CCA ratio) in veterans receiving carotid duplex ultrasound examinations in the Veteran’s Health Administration from 2001 to 2020. Among those who had at least one valid ICA/CCA ratio, we identified carotid stenosis severity (<50%, 50%−69%, ≥70%) based on the ICA/CCA ratio (<2, ≥2 to <4, ≥4) and assessed the association between presence and severity of carotid stenosis and clinical and demographic characteristics, including age, sex, self-identified race and ethnicity, smoking status, body mass index, systolic and diastolic blood pressures, indicator variables for pre-existing hypertension, coronary heart disease, and type 2 diabetes, and selected laboratory measures (ie, hemoglobin A1c, low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, triglyceride, and creatinine).Results:The harmonic F1 score of the NLP tool was 0.907 for the right value, 0.882 for the left value, and 0.920 for the maximum value. Among the 290,517 veterans in the cohort, the median age was 68.2 years. Black patients had 16% decreased risk of more severe carotid stenosis (odds ratio: 0.84, 95% confidence interval: 0.81–0.87, P < .001). All patient-level risk factors except high-density lipoprotein cholesterol were significantly associated with carotid stenosis severity.Conclusions:The NLP tool performed well, and the study performed with our NLP-created cohort largely validates the risk factors identified by previous smaller studies, demonstrating the utility of big data and NLP in carotid stenosis research. (JVS-Vascular Insights 2025;3:100302.)

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant