Fine-Grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection.

Yanxin Long,Runhui Huang,Chunjing Xu,Yi Zhu,Xiaodan Liang,Hang Xu,Jianhua Han

doi:10.1109/tnnls.2023.3293484

Abstract

Inspired by the success of vision-language methods (VLMs) in zero-shot classification, recent works attempt to extend this line of work into object detection by leveraging the localization ability of pretrained VLMs and generating pseudolabels for unseen classes in a self-training manner. However, since the current VLMs are usually pretrained with aligning sentence embedding with global image embedding, the direct use of them lacks fine-grained alignment for object instances, which is the core of detection. In this article, we propose a simple but effective fine-grained visual-text prompt-driven self-training paradigm for open-vocabulary detection (VTP-OVD) that introduces a fine-grained visual-text prompt adapting stage to enhance the current self-training paradigm with a more powerful fine-grained alignment. During the adapting stage, we enable VLM to obtain fine-grained alignment using learnable text prompts to resolve an auxiliary dense pixelwise prediction task. Furthermore, we propose a visual prompt module to provide the prior task information (i.e., the categories need to be predicted) for the vision branch to better adapt the pretrained VLM to the downstream tasks. Experiments show that our method achieves the state-of-the-art performance for open-vocabulary object detection, e.g., 31.5% mAP on unseen classes of COCO.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Fine-Grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems

Lead the way for us

Journal: IEEE transactions on neural networks and learning systems	Publication Date: Jan 1, 2024
Citations: 2

Similar Papers

A YOLOv5 Baseline for Underwater Object Detection
Hao Wang ... Shixin Sun
-
Hao Wang, et. al.Hao Wang ... Shixin Sun
20 Sep 2021
20 Sep 2021

Zero-shot Fine-grained Classification by Deep Feature Learning with Semantics
Ao-Xue Li ... Li-Wei Wang
International Journal of Automation and Computing | VOL. 16
Ao-Xue Li, et. al.Ao-Xue Li ... Li-Wei Wang
15 May 2019
International Journal of Automation and Computing | VOL. 16

Efficient object detection based on selective attention
Huapeng Yu ... Yafei Wang
Computers and Electrical Engineering | VOL. 40
Huapeng Yu, et. al.Huapeng Yu ... Yafei Wang
17 Oct 2013
Computers and Electrical Engineering | VOL. 40

Visual Language Based Succinct Zero-Shot Object Detection
Ye Zheng ... Xi Huang
-
Ye Zheng, et. al.Ye Zheng ... Xi Huang
17 Oct 2021
17 Oct 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Fine-Grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems