A Unified Transformer Framework for Group-Based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection

Yukun Su,Qingyao Wu,Ruizhou Sun,Guosheng Lin,Hanjing Su,Jingliang Deng

doi:10.1109/tmm.2023.3264883

Abstract

Humans tend to mine objects by learning from a group of images or several frames of video since we live in a dynamic world. In the computer vision area, many researchers focus on co-segmentation (CoS), co-saliency detection (CoSD) and video salient object detection (VSOD) to discover the co-occurrent objects. However, previous approaches design different networks for these similar tasks separately, and they are difficult to apply to each other. Besides, they fail to take full advantage of the cues among inter- and intra-feature within a group of images. In this paper, we introduce a unified framework to tackle these issues from a unified view, term as UFGS (Unified Framework for Group-based Segmentation). Specifically, we first introduce a transformer block, which views the image feature as a patch token and then captures their long-range dependencies through the self-attention mechanism. This can help the network to excavate the patch-structured similarities among the relevant objects. Furthermore, we propose an intra-MLP learning module to produce self-mask to enhance the network to avoid partial activation. Extensive experiments on four CoS benchmarks (PASCAL, iCoseg, Internet and MSRC), three CoSD benchmarks (Cosal2015, CoSOD3k, and CocA) and five VSOD benchmarks (DAVIS <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$_{16}$</tex-math></inline-formula> , FBMS, ViSal, SegV2 and DAVSOD) show that our method outperforms other state-of-the-arts on three different tasks in both accuracy and speed by using the same network architecture, which can reach 140 FPS in real-time. Code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/suyukun666/UFO</uri>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Unified Transformer Framework for Group-Based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia

Lead the way for us

Journal: IEEE Transactions on Multimedia	Publication Date: Jan 1, 2024
Citations: 29

Similar Papers

Transformer-based Cross Reference Network for video salient object detection
Kan Huang ... Jingyong Su
Pattern Recognition Letters | VOL. 160
Kan Huang, et. al.Kan Huang ... Jingyong Su
01 Aug 2022
Pattern Recognition Letters | VOL. 160

Collaborative spatial-temporal video salient object detection with cross attention transformer
Yuting Su ... Peiguang Jing
Signal Processing | VOL. 224
Yuting Su, et. al.Yuting Su ... Peiguang Jing
09 Jul 2024
Signal Processing | VOL. 224

Motion Guided Attention for Video Salient Object Detection
Haofeng Li ... Yizhou Yu
-
Haofeng Li, et. al.Haofeng Li ... Yizhou Yu
01 Oct 2019
01 Oct 2019

Shifting More Attention to Video Salient Object Detection
Deng-Ping Fan ... Wenguan Wang
-
Deng-Ping Fan, et. al.Deng-Ping Fan ... Wenguan Wang
01 Jun 2019
01 Jun 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Unified Transformer Framework for Group-Based Segmentation: Co-Segmentation, Co-Saliency Detection and Video Salient Object Detection

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia