DeLiVoTr: Deep and light-weight voxel transformer for 3D object detection

Gopi Krishna Erabati,Helder Araujo

doi:10.1016/j.iswa.2024.200361

Abstract

The image-based backbone (feature extraction) networks downsample the feature maps not only to increase the receptive field but also to efficiently detect objects of various scales. The existing feature extraction networks in LiDAR-based 3D object detection tasks follow the feature map downsampling similar to image-based feature extraction networks to increase the receptive field. But, such downsampling of LiDAR feature maps in large-scale autonomous driving scenarios hinder the detection of small size objects, such as pedestrians. To solve this issue we design an architecture that not only maintains the same scale of the feature maps but also the receptive field in the feature extraction network to aid for efficient detection of small size objects. We resort to attention mechanism to build sufficient receptive field and we propose a Deep and Light-weight Voxel Transformer (DeLiVoTr) network with voxel intra- and inter-region transformer modules to extract voxel local and global features respectively. We introduce DeLiVoTr block that uses transformations with expand and reduce strategy to vary the width and depth of the network efficiently. This facilitates to learn wider and deeper voxel representations and enables to use not only smaller dimension for attention mechanism but also a light-weight feed-forward network, facilitating the reduction of parameters and operations. In addition to model scaling, we employ layer-level scaling of DeLiVoTr encoder layers for efficient parameter allocation in each encoder layer instead of fixed number of parameters as in existing approaches. Leveraging layer-level depth and width scaling we formulate three variants of DeLiVoTr network. We conduct extensive experiments and analysis on large-scale Waymo and KITTI datasets. Our network surpasses state-of-the-art methods for detection of small objects (pedestrians) with an inference speed of 20.5 FPS.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

DeLiVoTr: Deep and light-weight voxel transformer for 3D object detection

Abstract

Talk to us

Similar Papers

More From: Intelligent Systems with Applications

Lead the way for us

Journal: Intelligent Systems with Applications	Publication Date: Mar 19, 2024
License type: cc-by-nc-nd

Similar Papers

Small object intelligent Detection method based on Adaptive Cascading Context
Jie Zhang ... Fengxian Wang
ACM Journal on Autonomous Transportation Systems | VOL. -
Jie Zhang, et. al.Jie Zhang ... Fengxian Wang
23 May 2024
ACM Journal on Autonomous Transportation Systems | VOL. -

Small Object Detection using Multi-scale Feature Fusion and Attention
Baokai Liu ... Shiqiang Du
-
Baokai Liu, et. al.Baokai Liu ... Shiqiang Du
25 Jul 2022
25 Jul 2022

Multi-scale Non-local Feature Enhancement Network for Robust Small-object Detection
Jun Ho Choi ... Byung Cheol Song
IEIE Transactions on Smart Processing & Computing | VOL. 9
Jun Ho Choi, et. al.Jun Ho Choi ... Byung Cheol Song
31 Aug 2020
IEIE Transactions on Smart Processing & Computing | VOL. 9

Efficient Small Object Detection with an Improved Region Proposal Networks
Dong Wen Ma ... Honghong Yang
IOP Conference Series: Materials Science and Engineering | VOL. 533
Dong Wen Ma, et. al.Dong Wen Ma ... Honghong Yang
01 May 2019
IOP Conference Series: Materials Science and Engineering | VOL. 533

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

DeLiVoTr: Deep and light-weight voxel transformer for 3D object detection

Abstract

Talk to us

Similar Papers

More From: Intelligent Systems with Applications