CAA-Net: Conditional Atrous CNNs With Attention for Explainable Device-Robust Acoustic Scene Classification

Zhao Ren,Mark D Plumbley,Jing Han,Qiuqiang Kong,Bjorn W Schuller

doi:10.1109/tmm.2020.3037534

Abstract

Acoustic Scene Classification (ASC) aims to classify the environment in which the audio signals are recorded. Recently, Convolutional Neural Networks (CNNs) have been successfully applied to ASC. However, the data distributions of the audio signals recorded with multiple devices are different. There has been little research on the training of robust neural networks on acoustic scene datasets recorded with multiple devices, and on explaining the operation of the internal layers of the neural networks. In this article, we focus on training and explaining device-robust CNNs on multi-device acoustic scene data. We propose conditional atrous CNNs with attention for multi-device ASC. Our proposed system contains an ASC branch and a device classification branch, both modelled by CNNs. We visualise and analyse the intermediate layers of the atrous CNNs. A time-frequency attention mechanism is employed to analyse the contribution of each time-frequency bin of the feature maps in the CNNs. On the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 ASC dataset, recorded with three devices, our proposed model performs significantly better than CNNs trained on single-device data.

Highlights

W ITH the development of computer audition [1], Acoustic Scene Classification (ASC) has become a major research field, aiming to automatically recognise acoustic environments [2], [3]
Many deep learning structures have been proposed for ASC, including Convolutional Neural Networks (CNNs) [7] and Recurrent Neural Networks (RNNs) [14]
We propose to employ atrous CNNs and an attention mechanism for visualisation

Summary

Introduction

W ITH the development of computer audition [1], Acoustic Scene Classification (ASC) has become a major research field, aiming to automatically recognise acoustic environments [2], [3]. The goal of ASC is to identify acoustic scenes in an audio stream, using computational approaches such as signal processing [4], [5], machine learning [6], and deep learning [7], [8]. Deep learning approaches have shown good performance in ASC [7], [12]. Many deep learning structures have been proposed for ASC, including Convolutional Neural Networks (CNNs) [7] and Recurrent Neural Networks (RNNs) [14]. Log mel spectrograms have been successfully utilised in ASC [7], [17]. In this regard, we extract log mel spectrograms as the inputs of the CNNs

Methods

Results

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE Transactions on Multimedia	Publication Date: Nov 12, 2020
Citations: 18	License type: cc-by

R Discovery Prime

R Discovery Prime

CAA-Net: Conditional Atrous CNNs With Attention for Explainable Device-Robust Acoustic Scene Classification

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia

Lead the way for us

Similar Papers

Sound Context Classification based on Joint Learning Model and Multi-Spectrogram Features
Dat Ngo ... Lam Pham
International Journal of Computing | VOL. -
Dat Ngo, et. al.Dat Ngo ... Lam Pham
30 Jun 2022
International Journal of Computing | VOL. -

Acoustic Scene Classification Using Reduced MobileNet Architecture
Jun-Xiang Xu ... Tzu-Ching Lin
-
Jun-Xiang Xu, et. al.Jun-Xiang Xu ... Tzu-Ching Lin
01 Dec 2018
01 Dec 2018

Late Fusion of Convolutional Neural Network with Wavelet-Based Ensemble Classifier for Acoustic Scene Classification
Cheng Siong Chin ... Jianhua Zhang
-
Cheng Siong Chin, et. al.Cheng Siong Chin ... Jianhua Zhang
04 Aug 2021
04 Aug 2021

Shallow Convolutional Neural Networks for Acoustic Scene Classification
Lu Lu ... Yuzhi Jiang
Wuhan University Journal of Natural Sciences | VOL. 23
Lu Lu, et. al.Lu Lu ... Yuzhi Jiang
19 Mar 2018
Wuhan University Journal of Natural Sciences | VOL. 23

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

CAA-Net: Conditional Atrous CNNs With Attention for Explainable Device-Robust Acoustic Scene Classification

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia