Explain and improve: LRP-inference fine-tuning for image captioning models

Jiamei Sun,Sebastian Lapuschkin,Wojciech Samek,Alexander Binder

doi:10.1016/j.inffus.2021.07.008

Abstract

This paper analyzes the predictions of image captioning models with attention mechanisms beyond visualizing the attention itself. We develop variants of Layer-wise Relevance Propagation (LRP) and gradient-based explanation methods, tailored to image captioning models with attention mechanisms. We compare the interpretability of attention heatmaps systematically against the explanations provided by explanation methods such as LRP, Grad-CAM, and Guided Grad-CAM. We show that explanation methods provide simultaneously pixel-wise image explanations (supporting and opposing pixels of the input image) and linguistic explanations (supporting and opposing words of the preceding sequence) for each word in the predicted captions. We demonstrate with extensive experiments that explanation methods (1) can reveal additional evidence used by the model to make decisions compared to attention; (2) correlate to object locations with high precision; (3) are helpful to “debug” the model, e.g. by analyzing the reasons for hallucinated object words. With the observed properties of explanations, we further design an LRP-inference fine-tuning strategy that reduces the issue of object hallucination in image captioning models, and meanwhile, maintains the sentence fluency. We conduct experiments with two widely used attention mechanisms: the adaptive attention mechanism calculated with the additive attention and the multi-head attention mechanism calculated with the scaled dot product.

Highlights

Image captioning is a setup that aims at generating text descriptions from image representations
Attentions are usually visualized as attention heatmaps, indicating which parts of the image are related to the generated words
Attention heatmaps are usually considered as the qualitative evaluations of image captioning models in addition to the quantitative evaluation metrics such as BLEU [16], METEOR [17], ROUGE-L [18], CIDEr [19], SPICE [20]

Summary

INTRODUCTION

Image captioning is a setup that aims at generating text descriptions from image representations. To gain more insights into the image captioning models, we adapt layer-wise relevance propagation (LRP) and gradientbased explanation methods (Grad-CAM, Guided Grad-CAM [21], and GuidedBackpropagation [22]) to explain image captioning predictions with respect to the image content and the words of the sentence generated so far. We quantitatively measure and compare the properties of explanation methods and attention mechanisms, including tasks of finding the related features/evidence for model decisions, grounding to image content, and the capability of debugging the models (in terms of providing possible reasons for object hallucination and differentiating hallucinated words). We propose an LRP-inference fine-tuning strategy that reduces object hallucination and guides the models to be more precise and grounded on image evidence when predicting frequent object words.

Image Captioning

Towards de-biasing visual-language models

Explanation-guided training

Notations for image captioning models

Attention mechanisms used in this study

EXPLANATION METHODS FOR IMAGE CAPTIONING

Model preparation and implementation details

Explanation results and evaluation

Grad-CAM

Reducing object hallucination with explanation

Discussion and outlook

CONCLUSION

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Information Fusion	Publication Date: Jul 31, 2021
Citations: 22	License type: cc-by

R Discovery Prime

R Discovery Prime

Explain and improve: LRP-inference fine-tuning for image captioning models

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Information Fusion

Lead the way for us

Similar Papers

Spatio-temporal learning and explaining for dynamic functional connectivity analysis: Application to depression
Jinlong Hu ... Gangqiang Hou
Journal of Affective Disorders | VOL. 364
Jinlong Hu, et. al.Jinlong Hu ... Gangqiang Hou
11 Aug 2024
Journal of Affective Disorders | VOL. 364

SLRP: Improved heatmap generation via selective layer‐wise relevance propagation
Yeon‐Jee Jung ... Seung‐Ho Han
Electronics Letters | VOL. 57
Yeon‐Jee Jung, et. al.Yeon‐Jee Jung ... Seung‐Ho Han
20 Apr 2021
Electronics Letters | VOL. 57

Attention-based GCN integrates multi-omics data for breast cancer subtype classification and patient-specific gene marker identification.
Hui Guo ... Xiang Lv
Briefings in functional genomics | VOL. 22
Hui Guo, et. al.Hui Guo ... Xiang Lv
26 Apr 2023
Briefings in functional genomics | VOL. 22

An Approach for Estimating Explanation Uncertainty in fMRI dFNC Classification
Charles A Ellis ... Robyn L Miller
-
Charles A Ellis, et. al.Charles A Ellis ... Robyn L Miller
01 Nov 2022
01 Nov 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Explain and improve: LRP-inference fine-tuning for image captioning models

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Information Fusion