Context, Language Modeling, and Multimodal Data in Finance

Sanjiv Das,Connor Goggins,Rob Van Dusen,Sandeep Krishnamurthy,George Karypis,Mitali Mahajan,Shuai Zheng,Nagpurnanand Prabhala,Shenghua Yue,Sheng Zha,Dylan Slack,John He

doi:10.3905/jfds.2021.1.063

Abstract

The authors enhance pretrained language models with Securities and Exchange Commission filings data to create better language representations for features used in a predictive model. Specifically, they train RoBERTa class models with additional financial regulatory text, which they denote as a class of RoBERTa-Fin models. Using different datasets, the authors assess whether there is material improvement over models that use only text-based numerical features (e.g., sentiment, readability, polarity), which is the traditional approach adopted in academia and practice. The RoBERTa-Fin models also outperform generic bidirectional encoder representations from transformers (BERT) class models that are not trained with financial text. The improvement in classification accuracy is material, suggesting that full text and context are important in classifying financial documents and that the benefits from the use of mixed data, (i.e., enhancing numerical tabular data with text) are feasible and fruitful in machine learning models in finance. <b>TOPICS:</b>Quantitative methods, big data/machine learning, legal/regulatory/public policy, information providers/credit ratings <b>Key Findings</b> ▪ Machine learning based on multimodal data provides meaningful improvement over models based on numerical data alone. ▪ Context-rich models perform better than context-free models. ▪ Pretrained language models that mix common text and financial text do better than those pretrained on financial text alone.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Context, Language Modeling, and Multimodal Data in Finance

Abstract

Talk to us

Similar Papers

More From: The Journal of Financial Data Science

Lead the way for us

Journal: The Journal of Financial Data Science	Publication Date: Jun 1, 2021
Citations: 1

Similar Papers

Retracted] Analysis and Risk Assessment of Corporate Financial Leverage Using Mobile Payment in the Era of Digital Technology in a Complex Environment
Wenjing Wei ... Bingxiang Li
Journal of Mathematics | VOL. 2022
Wenjing Wei, et. al.Wenjing Wei ... Bingxiang Li
01 Jan 2021
Journal of Mathematics | VOL. 2022

Tibetan Sentence Boundaries Automatic Disambiguation Based on Bidirectional Encoder Representations from Transformers on Byte Pair Encoding Word Cutting Method
Fenfang Li ... Han Deng
Applied Sciences | VOL. 14
Fenfang Li, et. al.Fenfang Li ... Han Deng
02 Apr 2024
Applied Sciences | VOL. 14

Bidirectional Encoder Representations from Transformers (BERT) Language Model for Sentiment Analysis task: Review

-

19 Apr 2021
19 Apr 2021

MenuNER: Domain-Adapted BERT Based NER Approach for a Domain with Limited Dataset and Its Application to Food Menu Domain
Muzamil Hussain Syed ... Sun-Tae Chung
Applied Sciences | VOL. 11
Muzamil Hussain Syed, et. al.Muzamil Hussain Syed ... Sun-Tae Chung
28 Jun 2021
Applied Sciences | VOL. 11

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Context, Language Modeling, and Multimodal Data in Finance

Abstract

Talk to us

Similar Papers

More From: The Journal of Financial Data Science