ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax.

Zachary Huemann,Xin Tie,Junjie Hu,Tyler J Bradshaw

doi:10.1007/s10278-024-01051-8

Abstract

Radiology narrative reports often describe characteristics of a patient's disease, including its location, size, and shape. Motivated by the recent success of multimodal learning, we hypothesized that this descriptive text could guide medical image analysis algorithms. We proposed a novel vision-language model, ConTEXTual Net, for the task of pneumothorax segmentation on chest radiographs. ConTEXTual Net extracts language features from physician-generated free-form radiology reports using a pre-trained language model. We then introduced cross-attention between the language features and the intermediate embeddings of an encoder-decoder convolutional neural network to enable language guidance for image analysis. ConTEXTual Net was trained on the CANDID-PTX dataset consisting of 3196 positive cases of pneumothorax with segmentation annotations from 6 different physicians as well as clinical radiology reports. Using cross-validation, ConTEXTual Net achieved a Dice score of 0.716±0.016, which was similar to the degree of inter-reader variability (0.712±0.044) computed on a subset of the data. It outperformed vision-only models (Swin UNETR: 0.670±0.015, ResNet50 U-Net: 0.677±0.015, GLoRIA: 0.686±0.014, and nnUNet 0.694±0.016) and a competing vision-language model (LAVT: 0.706±0.009). Ablation studies confirmed that it was the text information that led to the performance gains. Additionally, we show that certain augmentation methods degraded ConTEXTual Net's segmentation performance by breaking the image-text concordance. We also evaluated the effects of using different language models and activation functions in the cross-attention module, highlighting the efficacy of our chosen architectural design.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax.

Abstract

Talk to us

Similar Papers

More From: Journal of imaging informatics in medicine

Lead the way for us

Journal: Journal of imaging informatics in medicine	Publication Date: Mar 14, 2024
Citations: 5

Similar Papers

Neural Transfer Learning For Vietnamese Sentiment Analysis Using Pre-trained Contextual Language Models
An Pha Le ... Tran Vu Pham
-
An Pha Le, et. al.An Pha Le ... Tran Vu Pham
16 Dec 2021
16 Dec 2021

On the Power of Pre-Trained Text Representations
Yu Meng ... Jiawei Han
-
Yu Meng, et. al.Yu Meng ... Jiawei Han
14 Aug 2021
14 Aug 2021

Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
Anoop K ... Lajish V L
-
Anoop K, et. al. Anoop K ... Lajish V L
01 Jan 2021
01 Jan 2021

Effectiveness of Pre-Trained Language Models for the Japanese Winograd Schema Challenge
Keigo Takahashi ... Mamoru Komachi
Journal of Advanced Computational Intelligence and Intelligent Informatics | VOL. 27
Keigo Takahashi, et. al.Keigo Takahashi ... Mamoru Komachi
20 May 2023
Journal of Advanced Computational Intelligence and Intelligent Informatics | VOL. 27

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax.

Abstract

Talk to us

Similar Papers

More From: Journal of imaging informatics in medicine