LINSPECTOR: Multilingual Probing Tasks for Word Representations

Gözde Gül Şahin,Clara Vania,Iryna Gurevych,Ilia Kuznetsov

doi:10.1162/coli_a_00376

Gözde Gül Şahin, Clara Vania + Show 2 more

Open Access

https://doi.org/10.1162/coli_a_00376

Copy DOI

Abstract

Despite an ever-growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the community to get an estimate of the downstream task performance, as well as to design more informed neural architectures, while avoiding extensive experimentation that requires substantial computational resources not all researchers have access to. A recent development in NLP is to use simple classification tasks, also called probing tasks, that test for a single linguistic feature such as part-of-speech. Existing studies mostly focus on exploring the linguistic information encoded by the continuous representations of English text. However, from a typological perspective the morphologically poor English is rather an outlier: The information encoded by the word order and function words in English is often stored on a subword, morphological level in other languages. To address this, we introduce 15 type-level probing tasks such as case marking, possession, word length, morphological tag count, and pseudoword identification for 24 languages. We present a reusable methodology for creation and evaluation of such tests in a multilingual setting, which is challenging because of a lack of resources, lower quality of tools, and differences among languages. We then present experiments on several diverse multilingual word embedding models, in which we relate the probing task performance for a diverse set of languages to a range of five classic NLP tasks: POS-tagging, dependency parsing, semantic role labeling, named entity recognition, and natural language inference. We find that a number of probing tests have significantly high positive correlation to the downstream tasks, especially for morphologically rich languages. We show that our tests can be used to explore word embeddings or black-box neural models for linguistic cues in a multilingual setting. We release the probing data sets and the evaluation suite LINSPECTOR with https://github.com/UKPLab/linspector .

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Computational Linguistics	Publication Date: Jun 1, 2020
Citations: 24	License type: cc-by-nc-nd

R Discovery Prime

R Discovery Prime

LINSPECTOR: Multilingual Probing Tasks for Word Representations

Abstract

Talk to us

Similar Papers

More From: Computational Linguistics

Lead the way for us

Similar Papers

LINSPECTOR WEB: A Multilingual Probing Suite for Word Representations
Max Eichler ... Gözde Gül Şahin
-
Max Eichler, et. al.Max Eichler ... Gözde Gül Şahin
01 Jan 2019
01 Jan 2019

Probing What Different NLP Tasks Teach Machines about Function Word Comprehension
Najoung Kim ... Ellie Pavlick
-
Najoung Kim, et. al.Najoung Kim ... Ellie Pavlick
01 Jan 2019
01 Jan 2019

MACEDONIZER - The Macedonian Transformer Language Model
Jovana Dobreva ... Stojancho Tudzarski
-
Jovana Dobreva, et. al.Jovana Dobreva ... Stojancho Tudzarski
01 Jan 2021
01 Jan 2021

Evaluation and Analysis of Word Embedding Vectors of English Text Using Deep Learning Technique
Jaspreet Singh ... Prithvipal Singh
-
Jaspreet Singh, et. al.Jaspreet Singh ... Prithvipal Singh
01 Jan 2018
01 Jan 2018

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

LINSPECTOR: Multilingual Probing Tasks for Word Representations

Abstract

Talk to us

Similar Papers

More From: Computational Linguistics