SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
- Book Chapter
2
- 10.5772/4752
- Jun 1, 2007
The later conculsions can be obtained from the experimental results of the fourth section. 1. The three parts of speech recognition are conjunct one another and exist the relation restricted among themselves. Bark wavelet was used in improving the feature of ZCPA and MFCC the latter effect is obviously better than the former.It illustrates that Bark wavelet and the speech character described by MFCC feature are more closer than Bark wavelet and the speech character described by ZCPA feature. The fact is also as such. Bark wavelet is constructed directly according to the hearing perception of human ear, and MFCC is the cepstrum coefficents on the basis of Mel frequency. While Mel frequency is just the hearing frequency of human ear. Though the frequency bins of ZCPA are divided according to the hearing perception, the zero-crossing rate and peak amplitude are time-domain parameters which are transformed nonlinearly mapping to frequency bin. This kind of nonlinear transform may affect the consistency of ZCPA and hearing frequency, that results in decreasing in function. If the selection of training or recognition network is different, they have different effect on the results. Furthermore, the function of recogntion network has direct relationship with front-end filter and feature extracted. This point can be seen from the experimental results of combination mode1 FIR+ZCPA+HMM and mode 2 FIR+ZCPA+WNN . Comparing the two modes, the former two parts are same and the third part is different from using HMM or WNN the results obtained have much more different.The wavelet neural network has bright foreground for speech recognition.Its training speed is fast, which is good for implementation in real time. Further,it has also good recognition rates under no noise or noise environment and the number of recogintion words is larger. The paper researched some kinds combination modes aiming to the three parts of speech recognition system in Fig. 1. For other combination modes, such as Bark+MFCC+WNN Bark+ZCPA+WNN and so on, we will research them in later work. Which of combination ever is optimal? This needs considering practical application case. We hope the research can be refered by interesting researcher and get to the purpose of communication mutually and progress.
- Research Article
64
- 10.1016/j.neucom.2020.03.111
- May 15, 2020
- Neurocomputing
Deep historical long short-term memory network for action recognition
- Research Article
329
- 10.1016/0893-6080(95)00061-5
- Jan 1, 1995
- Neural Networks
Artificial convolution neural network for medical image pattern recognition
- Research Article
17
- 10.1109/tim.2004.827057
- Jun 1, 2004
- IEEE Transactions on Instrumentation and Measurement
This paper presents the neuro-fuzzy Takagi-Sugeno-Kang (TSK) network for the recognition and classification of flavor. The important role in this process fulfills the self-organizing process used for the creation of the inference rules. The self-organizing neurons perform the role of clustering data into fuzzy groups with different membership values (the preprocessing stage). Applying the automatic control of clusters, we have the optimal size of the TSK network. The developed measuring system has been applied for the recognition of flavor of different brands of beer. The fuzzy neural network is used for processing signals obtained from the semiconductor sensor array. The results of numerical experiments have confirmed the excellent performance of such solutions.
- Conference Article
2
- 10.1109/icpr48806.2021.9413343
- Jan 10, 2021
In recent years, scene text recognition is always regarded as a sequence-to-sequence problem. Connectionist Temporal Classification (CTC) and Attentional sequence recognition (Attn) are two very prevailing approaches to tackle this problem while they may fail in some scenarios respectively. CTC concentrates more on every individual character but is weak in text semantic dependency modeling. Attn based methods have better context semantic modeling ability while tends to overfit on limited training data. In this paper, we elaborately design a Rectified Attentional Double Supervised Network (ReADS) for general scene text recognition. To overcome the weakness of CTC and Attn, both of them are applied in our method but with different modules in two supervised branches which can make a complementary to each other. Moreover, effective spatial and channel attention mechanisms are introduced to eliminate background noise and extract valid foreground information. Finally, a simple rectified network is implemented to rectify irregular text. The ReADS can be trained end-to-end and only word-level annotations are required. Extensive experiments on various benchmarks verify the effectiveness of ReADS which achieves state-of-the-art performance.
- Research Article
6
- 10.1007/s11571-024-10092-2
- Mar 18, 2024
- Cognitive neurodynamics
Electroencephalogram (EEG) emotion recognition plays an important role in human-computer interaction, and a higher recognition accuracy can improve the user experience. In recent years, domain adaptive methods in transfer learning have been used to construct a general emotion recognition model to deal with domain difference among different subjects and sessions. However, it is still challenging to effectively reduce domain difference in domain adaptation. In this paper, we propose a Multiple-Source Distribution Deep Adaptive Feature Norm Network for EEG emotion recognition, which reduce domain difference by improving the transferability of task-specific features. In detail, the domain adaptive method of our model employs a three-layer network topology, inserts Adaptive Feature Norm to self-supervised adjustment between different layers, and combines a multiple-kernel selection approach to mean embedding matching. The method proposed in this paper achieves the best classification performance in the SEED and SEED-IV datasets. In SEED dataset, the average accuracy of cross-subject and cross-session experiments is 85.01 and 91.93%, respectively. In SEED-IV dataset, the average accuracy is 58.81% in cross-subject experiments and 59.51% in cross-session experiments. The experimental results demonstrate that our method can effectively reduce the domain difference and improve the emotion recognition accuracy.
- Research Article
552
- 10.1016/j.patcog.2019.01.020
- Jan 15, 2019
- Pattern Recognition
MORAN: A Multi-Object Rectified Attention Network for scene text recognition
- Conference Article
26
- 10.1109/cvprw.2018.00012
- Jun 1, 2018
Recently, deep learning based approaches have yielded a significant improvement in face recognition in the wild. However," disguised face" recognition is still a challenging task that needs to be investigated, and the Disguised Faces in the Wild (DFW) competition is designed for this task. In this paper, we propose a two-stage training approach to utilize the small-scale training data provided by the DFW competition. Specifically, in the first stage, we train Deep Convolutional Neural Networks (DCNNs) for generic face recognition. In the second stage, we use Principal Components Analysis (PCA) based on the DFW training set to find the best transformation matrix for identity representation of disguised faces. We evaluate our model on the DFW testing dataset and it shows better performance over the state-of-the-art generic face recognition methods. It also achieves the best results on the DFW competition - Phase 1.
- Research Article
- 10.47059/alinteri/v36i1/ajas21088
- Jun 29, 2021
- Alinteri Journal of Agriculture Sciences
Aim: The study aims to identify or recognize the alphabets using neural networks and fuzzy classifier/logic. Methods and materials: Neural network and fuzzy classifier are used for comparing the recognition of characters. For each classifier sample size is 20. Character recognition was developed using MATLAB R2018a, a software tool. The algorithm is again compared with the Fuzzy classifier to know the accuracy level. Results: Performance of both fuzzy classifier and neural networks are calculated by the accuracy value. The mean value of the fuzzy classifier is 82 and the neural network is 77. The recognition rate (accuracy) with the data features is found to be 98.06%. Fuzzy classifier shows higher significant value of P=0.002 < P=0.005 than the neural networks in recognition of characters. Conclusion: The independent tests for this study shows a higher accuracy level of alphabetical character recognition for Fuzzy classifier when compared with neural networks. Henceforth, the fuzzy classifier shows higher significant than the neural networks in recognition of characters.
- Research Article
11
- 10.1109/tpami.2023.3279378
- Oct 1, 2023
- IEEE Transactions on Pattern Analysis and Machine Intelligence
Face recognition has always been courted in computer vision and is especially amenable to situations with significant variations between frontal and profile faces. Traditional techniques make great strides either by synthesizing frontal faces from sizable datasets or by empirical pose invariant learning. In this paper, we propose a completely integrated embedded end-to-end Lie algebra residual architecture (LARNeXt) to achieve pose robust face recognition. First, we explore how the face rotation in the 3D space affects the deep feature generation process of convolutional neural networks (CNNs), and prove that face rotation in the image space is equivalent to an additive residual component in the feature space of CNNs, which is determined solely by the rotation. Second, on the basis of this theoretical finding, we further design three critical subnets to leverage a soft regression subnet with novel multi-fusion attention feature aggregation for efficient pose estimation, a residual subnet for decoding rotation information from input face images, and a gating subnet to learn rotation magnitude for controlling the strength of the residual component that contributes to the feature learning process. Finally, we conduct a large number of ablation experiments, and our quantitative and visualization results both corroborate the credibility of our theory and corresponding network designs. Our comprehensive experimental evaluations on frontal-profile face datasets, general unconstrained face recognition datasets, and industrial-grade tasks demonstrate that our method consistently outperforms the state-of-the-art ones.
- Research Article
194
- 10.1109/tip.2019.2929447
- Jul 29, 2019
- IEEE Transactions on Image Processing
Recently, food recognition has received more and more attention in image processing and computer vision for its great potential applications in human health. Most of the existing methods directly extracted deep visual features via convolutional neural networks (CNNs) for food recognition. Such methods ignore the characteristics of food images and are, thus, hard to achieve optimal recognition performance. In contrast to general object recognition, food images typically do not exhibit distinctive spatial arrangement and common semantic patterns. In this paper, we propose a multi-scale multi-view feature aggregation (MSMVFA) scheme for food recognition. MSMVFA can aggregate high-level semantic features, mid-level attribute features, and deep visual features into a unified representation. These three types of features describe the food image from different granularity. Therefore, the aggregated features can capture the semantics of food images with the greatest probability. For that solution, we utilize additional ingredient knowledge to obtain mid-level attribute representation via ingredient-supervised CNNs. High-level semantic features and deep visual features are extracted from class-supervised CNNs. Considering food images do not exhibit distinctive spatial layout in many cases, MSMVFA fuses multi-scale CNN activations for each type of features to make aggregated features more discriminative and invariable to geometrical deformation. Finally, the aggregated features are more robust, comprehensive, and discriminative via two-level fusion, namely multi-scale fusion for each type of features and multi-view aggregation for different types of features. In addition, MSMVFA is general and different deep networks can be easily applied into this scheme. Extensive experiments and evaluations demonstrate that our method achieves state-of-the-art recognition performance on three popular large-scale food benchmark datasets in Top-1 recognition accuracy. Furthermore, we expect this paper will further the agenda of food recognition in the community of image processing and computer vision.
- Conference Article
2
- 10.1109/ijcnn.1999.833552
- Jul 10, 1999
This paper provides a brief review of the state-of-the-art of neural networks in off-line text recognition. We discuss the role that neural networks have played in text recognition. We also assess the state of the art of neural networks in character and word recognition. Despite the success of neural networks in character and word recognition, there are still many challenging problems.
- Research Article
11
- 10.3390/mi12060670
- Jun 8, 2021
- Micromachines
Activity recognition is a fundamental and crucial task in computer vision. Impressive results have been achieved for activity recognition in high-resolution videos, but for extreme low-resolution videos, which capture the action information at a distance and are vital for preserving privacy, the performance of activity recognition algorithms is far from satisfactory. The reason is that extreme low-resolution (e.g., 12 × 16 pixels) images lack adequate scene and appearance information, which is needed for efficient recognition. To address this problem, we propose a super-resolution-driven generative adversarial network for activity recognition. To fully take advantage of the latent information in low-resolution images, a powerful network module is employed to super-resolve the extremely low-resolution images with a large scale factor. Then, a general activity recognition network is applied to analyze the super-resolved video clips. Extensive experiments on two public benchmarks were conducted to evaluate the effectiveness of our proposed method. The results demonstrate that our method outperforms several state-of-the-art low-resolution activity recognition approaches.
- Research Article
- 10.13053/rcs-78-1-2
- Dec 31, 2014
- Research in Computing Science
In this paper an approach based on modelling of recogni- tion error with Articial Neural Networks is presented to increase face emotion recognition in absence of any pre-processing or enhancement technique for feature extraction. This approach consists of two stages: in the rst stage an ANN structure is dened for the recognition task by means of a Genetic Algorithm (Recognition ANN - ReANN). Then this structure is used to perform recognition on a test set Y to estimate classication error probabilities. In the second stage an additional ANN is dened to associate these error patterns with correct classication pat- terns using the same test set (Corrective ANN - CoANN). The composite ANN system then is tested with a dierent set Z. In recognition tasks performed with the ReANN it was observed that some emotions were more likely to be incorrectly classied than others. This was further corroborated with perceptual data. With the integrated ANN system (ReANN plus CoANN) it was observed that some of these emotions could be recognised more accurately. In general overall recognition was increased from 75% to 85% with this approach.
- Research Article
43
- 10.1007/s42979-020-00325-6
- Jan 1, 2020
- Sn Computer Science
Current state-of-the-art models for automatic facial expression recognition (FER) are based on very deep neural networks that are effective but rather expensive to train. Given the dynamic conditions of FER, this characteristic hinders such models of been used as a general affect recognition. In this paper, we address this problem by formalizing the FaceChannel, a light-weight neural network that has much fewer parameters than common deep neural networks. We introduce an inhibitory layer that helps to shape the learning of facial features in the last layer of the network and, thus, improving performance while reducing the number of trainable parameters. To evaluate our model, we perform a series of experiments on different benchmark datasets and demonstrate how the FaceChannel achieves a comparable, if not better, performance to the current state-of-the-art in FER. Our experiments include cross-dataset analysis, to estimate how our model behaves on different affective recognition conditions. We conclude our paper with an analysis of how FaceChannel learns and adapts the learned facial features towards the different datasets.