Pansharpening by Convolutional Neural Networks
A new pansharpening method is proposed, based on convolutional neural networks. We adapt a simple and effective three-layer architecture recently proposed for super-resolution to the pansharpening problem. Moreover, to improve performance without increasing complexity, we augment the input by including several maps of nonlinear radiometric indices typical of remote sensing. Experiments on three representative datasets show the proposed method to provide very promising results, largely competitive with the current state of the art in terms of both full-reference and no-reference metrics, and also at a visual inspection.
- Research Article
30
- 10.1109/access.2020.2971502
- Jan 1, 2020
- IEEE Access
Pan-sharpening is a significant task that aims to generate high spectral- and spatial- resolution remote-sensing image by fusing multi-spectral (MS) and panchromatic (PAN) image. The conventional approaches are insufficient to protect the fidelity both in spectral and spatial domains. Inspired by the robust capability and outstanding performance of convolutional neural networks (CNN) in natural image super-resolution tasks, CNN-based pan-sharpening methods are worthy of further exploration. In this paper, a novel pan-sharpening method is proposed by introducing a multi-scale channel attention residual network (MSCARN), which can represent features accurately and reconstruct a pan-sharpened image comprehensively. In MSCARN, the multi-scale feature extraction blocks comprehensively extract the coarse structures and high-frequency details. Moreover, the multi-residual architecture guarantees the consistency of feature learning procedure and accelerates convergence. Specifically, we introduce a channel attention mechanism to recalibrate the channel-wise features by considering interdependencies among channels adaptively. The extensive experiments are implemented on two real-datasets from GaoFen series satellites. And the results show that the proposed method performs better than the existing methods both in full-reference and no-reference metrics, meanwhile, the visual inspection displays in accordance with the quantitative metrics. Besides, in comparison with pan-sharpening by convolutional neural networks (PNN), the proposed method achieves faster convergence rate and lower loss.
- Research Article
24
- 10.1038/s41598-021-97195-6
- Sep 9, 2021
- Scientific Reports
Artificial neural networks (ANN) which include deep learning neural networks (DNN) have problems such as the local minimal problem of Back propagation neural network (BPNN), the unstable problem of Radial basis function neural network (RBFNN) and the limited maximum precision problem of Convolutional neural network (CNN). Performance (training speed, precision, etc.) of BPNN, RBFNN and CNN are expected to be improved. Main works are as follows: Firstly, based on existing BPNN and RBFNN, Wavelet neural network (WNN) is implemented in order to get better performance for further improving CNN. WNN adopts the network structure of BPNN in order to get faster training speed. WNN adopts the wavelet function as an activation function, whose form is similar to the radial basis function of RBFNN, in order to solve the local minimum problem. Secondly, WNN-based Convolutional wavelet neural network (CWNN) method is proposed, in which the fully connected layers (FCL) of CNN is replaced by WNN. Thirdly, comparative simulations based on MNIST and CIFAR-10 datasets among the discussed methods of BPNN, RBFNN, CNN and CWNN are implemented and analyzed. Fourthly, the wavelet-based Convolutional Neural Network (WCNN) is proposed, where the wavelet transformation is adopted as the activation function in Convolutional Pool Neural Network (CPNN) of CNN. Fifthly, simulations based on CWNN are implemented and analyzed on the MNIST dataset. Effects are as follows: Firstly, WNN can solve the problems of BPNN and RBFNN and have better performance. Secondly, the proposed CWNN can reduce the mean square error and the error rate of CNN, which means CWNN has better maximum precision than CNN. Thirdly, the proposed WCNN can reduce the mean square error and the error rate of CWNN, which means WCNN has better maximum precision than CWNN.
- Research Article
36
- 10.1109/access.2019.2943927
- Jan 1, 2019
- IEEE Access
In order to mine information from medical health data and develop intelligent application-related issues, the multi-modal medical health data feature representation learning related content was studied, and several feature learning models were proposed for disease risk assessment. In the aspect of medical text feature learning, a medical text feature learning model based on convolutional neural network is proposed. The convolutional neural network text analysis technology is applied to the disease risk assessment application. The medical data feature representation adopts the deep learning method. The learning and extraction of different disease characteristics use the same method to realize the versatility of the model. A simple preprocessing of the experimental data samples, including its power frequency denoising and lead convolution regularization, constructs a convolutional neural network for medical data feature advancement and intelligent recognition. On the basis of it, several sets of experiments were carried out to discuss the influence of the convolution kernel and the choice of learning rate on the experimental results. In addition, comparative experiments with support vector machine, BP neural network and RBF neural network are carried out. The results show that the convolutional neural network used in this paper shows obvious advantages in recognition rate and training speed compared with other methods. In the aspect of time series data feature learning, a multi-channel convolutional self-encoding neural network is proposed. Analyze the connection between fatigue and emotional abnormalities and define the concept of emotional fatigue. The proposed multi-channel convolutional neural network is used to learn the data features, and the convolutional self-encoding neural network is used to learn the facial image data features. These two characteristics and the collected physiological data are combined to perform emotional fatigue detection. An emotional fatigue detection demonstration platform for multi-modal data feature fusion is established to realize data acquisition, emotional fatigue detection and emotional feedback. The experimental results verify the validity, versatility and stability of the model.
- Research Article
13
- 10.1016/j.neucom.2020.04.003
- May 17, 2020
- Neurocomputing
A convolutional fuzzy min-max neural network
- Research Article
197
- 10.1016/j.neucom.2019.01.111
- Apr 24, 2019
- Neurocomputing
Brain tumor segmentation with deep convolutional symmetric neural network
- Research Article
31
- 10.1038/s41598-020-80713-3
- Jan 14, 2021
- Scientific Reports
Most speech separation studies in monaural channel use only a single type of network, and the separation effect is typically not satisfactory, posing difficulties for high quality speech separation. In this study, we propose a convolutional recurrent neural network with an attention (CRNN-A) framework for speech separation, fusing advantages of two networks together. The proposed separation framework uses a convolutional neural network (CNN) as the front-end of a recurrent neural network (RNN), alleviating the problem that a sole RNN cannot effectively learn the necessary features. This framework makes use of the translation invariance provided by CNN to extract information without modifying the original signals. Within the supplemented CNN, two different convolution kernels are designed to capture information in both the time and frequency domains of the input spectrogram. After concatenating the time-domain and the frequency-domain feature maps, the feature information of speech is exploited through consecutive convolutional layers. Finally, the feature map learned from the front-end CNN is combined with the original spectrogram and is sent to the back-end RNN. Further, the attention mechanism is further incorporated, focusing on the relationship among different feature maps. The effectiveness of the proposed method is evaluated on the standard dataset MIR-1K and the results prove that the proposed method outperforms the baseline RNN and other popular speech separation methods, in terms of GNSDR (gloabl normalised source-to-distortion ratio), GSIR (global source-to-interferences ratio), and GSAR (gloabl source-to-artifacts ratio). In summary, the proposed CRNN-A framework can effectively combine the advantages of CNN and RNN, and further optimise the separation performance via the attention mechanism. The proposed framework can shed a new light on speech separation, speech enhancement, and other related fields.
- Research Article
90
- 10.1002/ctm2.102
- Jun 1, 2020
- Clinical and Translational Medicine
Deep learning-based classification and mutation prediction from histopathological images of hepatocellular carcinoma.
- Research Article
2
- 10.1149/ma2021-01541317mtgabs
- May 30, 2021
- ECS Meeting Abstracts
The electronic nose is a gas detection instrument using the bionic olfactory mechanism, which usually consists of the gas sensor array and the gas classification algorithm. Furthermore, the capability of the gas classification algorithm is critical to the reliable accuracy of gas recognition. The process of gas classification usually involves pattern recognition of multiple time-related gas sensor response curves. Traditional gas classification algorithms are mainly machine learning methods, such as PCA, LDA, ICA, SVM, KNN, etc. These algorithms are relatively cumbersome because we need to extract handcrafted features before using them. Deep learning also has applications in electronic nose gas classification algorithms[1-4] which improve the accuracy of classification result, but most of deep neural networks have complex structures and consume huge computing resources. The spiking neural network is the third generation of artificial neural network, and its spiking neuron model which is more bionic than previous artificial neurons can process spike sequence signals[5]. The spiking neural network model has a simple structure with higher computational efficiency, and takes up less computational resources. Moreover, its time attribute makes it more suitable for processing information about time series. In order to simultaneously take advantage of the efficient feature extraction of the convolutional neural network and the high computational efficiency of the spiking neural network, our team converted the convolutional neural network into a convolutional spiking neural network(CSNN) and applied it to gas classification. The activation function layer in the traditional convolutional layer was replaced with the spiking neuron layer which used the IF or LIF spiking neuron model to transform the continuous values passed by the convolutional later into discrete values so as to achieve the transmission of spikes between layers. The first convolutional spiking layer was used as a spiking encoder, so the spiking encoding method such as Gaussian encoding was not used. The spike-firing-frequency output by the last layer of neurons was calculated to obtain the classification result. The probability that a gas sample belonged to a certain class was proportional to the spike-firing-frequency of the corresponding neuron of the class. Our team built a convolutional spiking neural network model with 9 layers of convolutional spiking layer and 2 layers of fully connected spiking layer, and used the food spoilage gas dataset collected by us and open source gas mixtures dataset[6] to evaluate the capability of our model. With regard to the gas mixtures dataset, ethylene, methane, CO and their mixed state need to be classified. After training, CSNN achieved the test accuracy of 92.6%, and the other algorithms’ test accuracy were 92.9% of ResNet-18, 91.2% of one-dimensional deep convolutional neural network(1D-DCNN) and 88.5% of SVM. As for the food spoilage gas dataset, 30 types of spoiled meat, vegetables, fruits and their mixed state samples were measured. The first task was to classify the major categories of spoiled food, furtuer, the 30 types of spoiled food odor samples were going to be divided into 4 categories: fresh food, spoiled meat, spoiled vegetables and spoiled fruits. After training, CSNN achieved the test accuracy of 81.4% which had a certain accuracy improvement comparing with 80.6% of ResNet-18, 80.1% of 1D-DCNN and 77.3% of SVM. The second task was to classify the subcategories of spoiled fruits, that is, 8 classes of spoiled fruit odor samples should be classified. After training, CSNN could achieve high test accuracy of 90.7%, and the accuracy of other algorithms was 88.8% of ResNet-18, 87.1% of 1D-DCNN and 77.9% of PCA+ANN. The CSNN output of a spoiled watermelon sample is shown in Figure 1.In conclusion, CSNN had similar odor classification performance to ResNet-18, but the computing resources occupied by CSNN was only 1/5 of ResNet-18. This research shows that the spiking neural network has the advantages of high odor classification accuracy, great calculation efficiency and occupying few computing resources. It is suitable as a gas classification algorithm of electronic nose and for further development.
- Book Chapter
2
- 10.1007/978-3-031-28076-4_51
- Jan 1, 2023
The quality of classification is crucial in medical applications. Especially when it comes to confirm that the patient does not have a malignant tumor. An example of such an application is a binary classification of breast tumor malignancy based on histopathological images. This paper explains the most popular attention mechanism in convolution neural networks as follows. Convolutional Block Attention Module, Attention Augmented Convolution, and Attention Guided Convolutional Neural Networks. Four neural networks are built and compared. Each is evaluated in the classification problem of histopathological images of breast cancer. On the basis of the results, it is clear that some attentional neural networks can outperform standard convolutional networks in the classification of breast cancer. In our investigation, the convolution networks reached an accuracy level of 90% and an AUC-ROC of 95.9%. It is worse compared to the Convolutional Block Attention Module Network (accuracy 90.7%, AUC-ROC 96.9%) and the Attention-Guided Convolutional Network (accuracy 91.2%, AUC-ROC 96.6%). Attention-augmented convolution remains behind the standard convolutional network, achieving 88.9% accuracy and 94.8% AUC-ROC. The Attention-Guided Convolution Network was the best network of all four. We also compared precision, NPV, sensitivity, specificity, and \(F_{1}\)-score. We came to the conclusion that the Convolutional Block Attention Module network has the highest NPV (90.8%) and sensitivity (96.2%), while the Attention-Guided Convolutional Neural Network scored the highest in precision (92.4%) and specificity (82.9%).KeywordsConvolutional neural networksAttentional neural networksBreast cancerHistopathologic imagesConvolutional block attention moduleAttention-augmented convolutionAttention-guided convolutional neural network
- Research Article
2
- 10.34229/2707-451x.21.3.6
- Sep 30, 2021
- Cybernetics and Computer Technologies
Introduction. The implementation of information technologies in various spheres of public life dictates the creation of efficient and productive systems for entering information into computer systems. In such systems it is important to build an effective recognition module. At the moment, the most effective method for solving this problem is the use of artificial multilayer neural and convolutional networks. The purpose of the paper. This paper is devoted to a comparative analysis of the recognition results of handwritten characters of the Azerbaijani alphabet using neural and convolutional neural networks. Results. The analysis of the dependence of the recognition results on the following parameters is carried out: the architecture of neural networks, the size of the training base, the choice of the subsampling algorithm, the use of the feature extraction algorithm. To increase the training sample, the image augmentation technique was used. Based on the real base of 14000 characters, the bases of 28000, 42000 and 72000 characters were formed. The description of the feature extraction algorithm is given. Conclusions. Analysis of recognition results on the test sample showed: as expected, convolutional neural networks showed higher results than multilayer neural networks; the classical convolutional network LeNet-5 showed the highest results among all types of neural networks. However, the multi-layer 3-layer network, which was input by the feature extraction results; showed rather high results comparable with convolutional networks; there is no definite advantage in the choice of the method in the subsampling layer. The choice of the subsampling method (max-pooling or average-pooling) for a particular model can be selected experimentally; increasing the training database for this task did not give a tangible improvement in recognition results for convolutional networks and networks with preliminary feature extraction. However, for networks learning without feature extraction, an increase in the size of the database led to a noticeable improvement in performance. Keywords: neural networks, feature extraction, OCR.
- Research Article
105
- 10.1016/j.jrmge.2021.09.004
- Dec 1, 2021
- Journal of Rock Mechanics and Geotechnical Engineering
Tunnel boring machine vibration-based deep learning for the ground identification of working faces
- Research Article
49
- 10.1007/s10489-020-02015-5
- Nov 27, 2020
- Applied Intelligence
Convolutional neural network (CNN) is recognized as state of the art of deep learning algorithm, which has a good ability on the image classification and recognition. The problems of CNN are as follows: the precision, accuracy and efficiency of CNN are expected to be improved to satisfy the requirements of high performance. The main work is as follows: Firstly, wavelet convolutional neural network (wCNN) is proposed, where wavelet transform function is added to the convolutional layers of CNN. Secondly, wavelet convolutional wavelet neural network (wCwNN) is proposed, where fully connected neural network (FCNN) of wCNN and CNN are replaced by wavelet neural network (wNN). Thirdly, image classification experiments using CNN, wCNN and wCwNN algorithms, and comparison analysis are implemented with MNIST dataset. The effect of the improved methods are as follows: (1) Both precision and accuracy are improved. (2) The mean square error and the rate of error are reduced. (3) The complexitie of the improved algorithms is increased.
- Conference Article
2468
- 10.1109/cvpr.2018.00685
- Jun 1, 2018
The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved significantly. However, to date, conventional research has only explored relatively shallow 3D architectures. We examine the architectures of various 3D CNNs from relatively shallow to very deep ones on current video datasets. Based on the results of those experiments, the following conclusions could be obtained: (i) ResNet-18 training resulted in significant overfitting for UCF-101, HMDB-51, and ActivityNet but not for Kinetics. (ii) The Kinetics dataset has sufficient data for training of deep 3D CNNs, and enables training of up to 152 ResNets layers, interestingly similar to 2D ResNets on ImageNet. ResNeXt-101 achieved 78.4% average accuracy on the Kinetics test set. (iii) Kinetics pretrained simple 3D architectures outperforms complex 2D architectures, and the pretrained ResNeXt-101 achieved 94.5% and 70.2% on UCF-101 and HMDB-51, respectively. The use of 2D CNNs trained on ImageNet has produced significant progress in various tasks in image. We believe that using deep 3D CNNs together with Kinetics will retrace the successful history of 2D CNNs and ImageNet, and stimulate advances in computer vision for videos. The codes and pretrained models used in this study are publicly available1.
- Conference Article
7
- 10.1109/igarss.2019.8897928
- Jul 1, 2019
Recently, convolutional neural network (CNN) has achieved great results in pansharpening. Most pansharpening methods with CNN are based on PNN [1] inspired by super-resolution methods with CNN and learn the pansharpening of downsampled images. In this work, we presented a novel framework for pansharpening based on two desired property of pansharpened images: downsampled pansharpened images become low-resolution multi-spectral images (spectral consistency) and panchromatic images are approximated by weighted addition of each bands of pansharpened images (spatial consistency). Our framework train CNN to learn this spectral-spatial consistency. The advantage of our framework is that there is no scale mismatch between training and test data. We applied our method to Landsat-8 images and compared it with some previous methods.
- Research Article
43
- 10.1080/10106049.2020.1740950
- Mar 18, 2020
- Geocarto International
Identification of crops is an important topic in the agricultural domain. Hyperspectral remote sensing data are very useful for crop feature extraction and classification. Remote sensing data is an unstructured data and Convolutional Neural Network (CNN) can work on unstructured data efficiently. This paper presents an evaluation of CNN for crop classification using the Indian Pines standard dataset obtained from the AVIRIS sensor and the study area dataset obtained from the EO-1hyperion sensor. Optimized CNN has been tuned by training the model on different parameters. It has been compared with two classification algorithms: Deep Neural Network (DNN) and Convolutional Autoencoder. According to the test results, the proposed optimized CNN model provided better results as compared to the other two methods. CNN has given 97 ± 0.58% overall accuracy for the Indian Pines standard dataset and 78 ± 2.43% for our study area dataset.