FFAVOD: Feature fusion architecture for video object detection

Hughes Perreault,Maguelonne Héritier,Guillaume-Alexandre Bilodeau,Nicolas Saunier

doi:10.1016/j.patrec.2021.09.002

Hughes Perreault, Maguelonne Héritier + Show 2 more

Open Access

https://doi.org/10.1016/j.patrec.2021.09.002

Copy DOI

Abstract

• We designed a novel architecture for video object detection that capitalizes on temporal information. • We designed a novel fusion module to merge feature maps coming from several temporally close frames. • We proposed an improvement to the SpotNet attention module. • We trained and evaluated our architecture with three different base detectors on two traffic surveillance datasets. • We demonstrated a consistent and significant improvement of our model over the three baselines. A significant amount of redundancy exists between consecutive frames of a video. Object detectors typically produce detections for one image at a time, without any capabilities for taking advantage of this redundancy. Meanwhile, many applications for object detection work with videos, including intelligent transportation systems, advanced driver assistance systems and video surveillance. Our work aims at taking advantage of the similarity between video frames to produce better detections. We propose FFAVOD, standing for feature fusion architecture for video object detection. We first introduce a novel video object detection architecture that allows a network to share feature maps between nearby frames. Second, we propose a feature fusion module that learns to merge feature maps to enhance them. We show that using the proposed architecture and the fusion module can improve the performance of three base object detectors on two object detection benchmarks containing sequences of moving road users. Additionally, to further increase performance, we propose an improvement to the SpotNet attention module. Using our architecture on the improved SpotNet detector, we obtain the state-of-the-art performance on the UA-DETRAC public benchmark as well as on the UAVDT dataset. Code is available at https://github.com/hu64/FFAVOD .

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

FFAVOD: Feature fusion architecture for video object detection

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition Letters

Lead the way for us

Journal: Pattern Recognition Letters	Publication Date: Nov 1, 2021
Citations: 14

Similar Papers

Lightweight Hardware Architecture for Object Detection in Driver Assistance Systems
Bhaumik Vaidya ... Chirag Paunwala
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 36
Bhaumik Vaidya, et. al.Bhaumik Vaidya ... Chirag Paunwala
06 Apr 2022
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 36

Video Object Detection Using Object’s Motion Context and Spatio-Temporal Feature Aggregation

-

29 Dec 2020
29 Dec 2020

Video Object Detection Using Object's Motion Context and Spatio-Temporal Feature Aggregation
Jaekyum Kim ... Jun Won Choi
-
Jaekyum Kim, et. al.Jaekyum Kim ... Jun Won Choi
10 Jan 2021
10 Jan 2021

Visual Feature Learning on Video Object and Human Action Detection: A Systematic Review.
Dengshan Li ... Rujing Wang
Micromachines | VOL. 13
Dengshan Li, et. al.Dengshan Li ... Rujing Wang
31 Dec 2021
Micromachines | VOL. 13

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

FFAVOD: Feature fusion architecture for video object detection

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition Letters