Multiple Scenes Research Articles

AbstractLiDAR‐based 3D place recognition is an essential component of simultaneous localization and mapping systems in multi‐scene robotic applications. However, extracting discriminative and generalizable global descriptors of point clouds is still an open issue due to the insufficient use of the information contained in the LiDAR scans in existing approaches. In this paper, we propose a novel spatial‐temporal point cloud encoding network for multiple scenes, dubbed STM‐Net, to fully fuse the multi‐view spatial information and temporal information of LiDAR point clouds. Specifically, we first develop a spatial feature encoding module consisting of the single‐view transformer and multi‐view transformer. The module learns the correlation both within a single view and between two views by utilizing the multi‐layer range images generated by spherical projection and multi‐layer bird's eye view images generated by top‐down projection. Then in the temporal feature encoding module, we exploit the temporal transformer to mine the temporal information in the sequential point clouds, and a NetVLAD layer is applied to aggregate features and generate sub‐descriptors. Furthermore, we use a GeM pooling layer to fuse more information along the time dimension for the final global descriptors. Extensive experiments conducted on unmanned ground/surface vehicles with different LiDAR configurations indicate that our method (1) achieves superior place recognition performance than state‐of‐the‐art algorithms, (2) generalizes well to diverse sceneries, (3) is robust to viewpoint changes, (4) can operate in real‐time, demonstrating the effectiveness and satisfactory capability of the proposed approach and highlighting its promising applications in multi‐scene place recognition tasks.

Read full abstract

Abstract. Retrieving UAV images that lack POS information with georeferenced satellite orthoimagery is challenging due to the differences in angles of views. Most existing methods rely on deep neural networks with a large number of parameters, leading to substantial time and financial investments in network training. Consequently, these methods may not be well-suited for downstream tasks that have high timeliness requirements. In this work, we propose a cross-view remote sensing image retrieval method based on transformer and visual foundation model. We investigated the potential of visual foundation model for extracting common features from cross-view images. Training is only conducted on a small, self-designed retrieval head, alleviating the burden of network training. Specifically, we designed a CVV module to optimize the features extracted from the visual foundation model, making these features more adept for cross-view image retrieval tasks. And we designed an MLP head to achieve similarity discrimination. The method is verified on a publicly available dataset containing multiple scenes. Our method shows excellent results in terms of both efficiency and accuracy on 15 sub-datasets (10 or 50 scene categories) derived from the public dataset, which holds practical value in engineering applications with streamlined scene categories and constrained computational resources. Furthermore, we initiated a comprehensive discussion and conducted ablation experiments on the network design to validate its efficacy. Additionally, we analyzed the presence of overfitting within the network and deliberated on the limitations of our study, proposing potential avenues for future enhancements.

Read full abstract

Multiple Scenes Research Articles

Related Topics

Articles published on Multiple Scenes

SCARF: Scalable Continual Learning Framework for Memory‐efficient Multiple Neural Radiance Fields

A transient network model for cross-regional layered fire smoke diffusion in high-rise buildings

BEW-YOLOv8: A deep learning model for multi-scene and multi-scale flood depth estimation

Rapid motion planning of manipulator in three-dimensional space under multiple scenes

Crowd Scenes in Péter Nádas’ Parallel Stories

LiDAR‐based place recognition for mobile robots in ground/water surface multiple scenes

Deep Edge-Based Fault Detection for Solar Panels.

MLA-MFL: A Smartphone Indoor Localization Method for Fusing Multisource Sensors Under Multiple Scene Conditions

Transformer-based berm detection for automated bulldozer safety in edge dumping

Research on a Near-Field Millimeter Wave Imaging Algorithm and System Based on Multiple-Input Multiple-Output Sparse Sampling

Few-Shot Remote Sensing Novel View Synthesis with Geometry Constraint NeRF

The Batman and the redemption of first responders: The culmination of a rhetoric of rescue during COVID-19, Trump and protest

A Transformer and Visual Foundation Model-Based Method for Cross-View Remote Sensing Image Retrieval

PCINet: a Prototype- and Concept-based Interpretable Network for Mutli-scene Recognition

Research on the countermeasures of the single-phase-to-earth faults in flexibly grounded distribution networks

OF-DFN: Optical flow prediction network for different perspective image fusion

Visual Recognition Memory of Scenes Is Driven by Categorical, Not Sensory, Visual Representations.

Learning single and multi-scene camera pose regression with transformer encoders

Dark Light Image-Enhancement Method Based on Multiple Self-Encoding Prior Collaborative Constraints

Photoelectricity Theory-Based Concrete Crack Image Segmentation and Optimal Exposure Interval Research

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Multiple Scenes Research Articles

Related Topics

Articles published on Multiple Scenes

SCARF: Scalable Continual Learning Framework for Memory‐efficient Multiple Neural Radiance Fields

A transient network model for cross-regional layered fire smoke diffusion in high-rise buildings

BEW-YOLOv8: A deep learning model for multi-scene and multi-scale flood depth estimation

Rapid motion planning of manipulator in three-dimensional space under multiple scenes

Crowd Scenes in Péter Nádas’ Parallel Stories

LiDAR‐based place recognition for mobile robots in ground/water surface multiple scenes

Deep Edge-Based Fault Detection for Solar Panels.

MLA-MFL: A Smartphone Indoor Localization Method for Fusing Multisource Sensors Under Multiple Scene Conditions

Transformer-based berm detection for automated bulldozer safety in edge dumping

Research on a Near-Field Millimeter Wave Imaging Algorithm and System Based on Multiple-Input Multiple-Output Sparse Sampling

Few-Shot Remote Sensing Novel View Synthesis with Geometry Constraint NeRF

The Batman and the redemption of first responders: The culmination of a rhetoric of rescue during COVID-19, Trump and protest

A Transformer and Visual Foundation Model-Based Method for Cross-View Remote Sensing Image Retrieval

PCINet: a Prototype- and Concept-based Interpretable Network for Mutli-scene Recognition

Research on the countermeasures of the single-phase-to-earth faults in flexibly grounded distribution networks

OF-DFN: Optical flow prediction network for different perspective image fusion

Visual Recognition Memory of Scenes Is Driven by Categorical, Not Sensory, Visual Representations.

Learning single and multi-scene camera pose regression with transformer encoders

Dark Light Image-Enhancement Method Based on Multiple Self-Encoding Prior Collaborative Constraints

Photoelectricity Theory-Based Concrete Crack Image Segmentation and Optimal Exposure Interval Research