Natural Language Inference with Transformer Ensembles and Explainability Techniques

Isidoros Perikos,Spyro Souli

doi:10.3390/electronics13193876

Abstract

Natural language inference (NLI) is a fundamental and quite challenging task in natural language processing, requiring efficient methods that are able to determine whether given hypotheses derive from given premises. In this paper, we apply explainability techniques to natural-language-inference methods as a means to illustrate the decision-making procedure of its methods. First, we investigate the performance and generalization capabilities of several transformer-based models, including BERT, ALBERT, RoBERTa, and DeBERTa, across widely used datasets like SNLI, GLUE Benchmark, and ANLI. Then, we employ stacking-ensemble techniques to leverage the strengths of multiple models and improve inference performance. Experimental results demonstrate significant improvements of the ensemble models in inference tasks, highlighting the effectiveness of stacking. Specifically, our best-performing ensemble models surpassed the best-performing individual transformer by 5.31% in accuracy on MNLI-m and MNLI-mm tasks. After that, we implement LIME and SHAP explainability techniques to shed light on the decision-making of the transformer models, indicating how specific words and contextual information are utilized in the transformer inferences procedures. The results indicate that the model properly leverages contextual information and individual words to make decisions but, in some cases, find difficulties in inference scenarios with metaphorical connections which require deeper inferential reasoning.

Full Text