VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning

Tangfei Liao,Tao Wang,Li Zhao,Xiaoqin Zhang,Guobao Xiao

doi:10.1609/aaai.v38i4.28123

Abstract

Correspondence pruning aims to find correct matches (inliers) from an initial set of putative correspondences, which is a fundamental task for many applications. The process of finding is challenging, given the varying inlier ratios between scenes/image pairs due to significant visual differences. However, the performance of the existing methods is usually limited by the problem of lacking visual cues (e.g., texture, illumination, structure) of scenes. In this paper, we propose a Visual-Spatial Fusion Transformer (VSFormer) to identify inliers and recover camera poses accurately. Firstly, we obtain highly abstract visual cues of a scene with the cross attention between local features of two-view images. Then, we model these visual cues and correspondences by a joint visual-spatial fusion module, simultaneously embedding visual cues into correspondences for pruning. Additionally, to mine the consistency of correspondences, we also design a novel module that combines the KNN-based graph and the transformer, effectively capturing both local and global contexts. Extensive experiments have demonstrated that the proposed VSFormer outperforms state-of-the-art methods on outdoor and indoor benchmarks. Our code is provided at the following repository: https://github.com/sugar-fly/VSFormer.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Similar Papers

The Role of Global and Local Contexts in Pronoun Comprehension
Bing Gao
Acta Psychologica Sinica | VOL. 40
Bing GaoBing Gao
19 Sep 2008
Acta Psychologica Sinica | VOL. 40

The Effects of Global and Local Stimulus Context on Auditory Frequency Discrimination
I Tsaliach, ... K Banai,
Journal of Basic and Clinical Physiology and Pharmacology | VOL. 21
I Tsaliach,, et. al.I Tsaliach, ... K Banai,
01 Jun 2010
Journal of Basic and Clinical Physiology and Pharmacology | VOL. 21

Multi-Level Joint Feature Learning for Person Re-Identification
Shaojun Wu ... Ling Gao
Algorithms | VOL. 13
Shaojun Wu, et. al.Shaojun Wu ... Ling Gao
29 Apr 2020
Algorithms | VOL. 13

Influential Global and Local Contexts Guided Trace Representation for Fault Localization
Zhuo Zhang ... Yue Yu
ACM Transactions on Software Engineering and Methodology | VOL. 32
Zhuo Zhang, et. al.Zhuo Zhang ... Yue Yu
26 Apr 2023
ACM Transactions on Software Engineering and Methodology | VOL. 32

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence