Learning Deep Representation for Face Alignment with Auxiliary Attributes.
In this study, we show that landmark detection or face alignment task is not a single and independent problem. Instead, its robustness can be greatly improved with auxiliary information. Specifically, we jointly optimize landmark detection together with the recognition of heterogeneous but subtly correlated facial attributes, such as gender, expression, and appearance attributes. This is non-trivial since different attribute inference tasks have different learning difficulties and convergence rates. To address this problem, we formulate a novel tasks-constrained deep model, which not only learns the inter-task correlation but also employs dynamic task coefficients to facilitate the optimization convergence when learning multiple complex tasks. Extensive evaluations show that the proposed task-constrained learning (i) outperforms existing face alignment methods, especially in dealing with faces with severe occlusion and pose variation, and (ii) reduces model complexity drastically compared to the state-of-the-art methods based on cascaded deep model.
- Book Chapter
1663
- 10.1007/978-3-319-10599-4_7
- Jan 1, 2014
Facial landmark detection has long been impeded by the problems of occlusion and pose variation. Instead of treating the detection task as a single and independent problem, we investigate the possibility of improving detection robustness through multi-task learning. Specifically, we wish to optimize facial landmark detection together with heterogeneous but subtly correlated tasks, e.g. head pose estimation and facial attribute inference. This is non-trivial since different tasks have different learning difficulties and convergence rates. To address this problem, we formulate a novel tasks-constrained deep model, with task-wise early stopping to facilitate learning convergence. Extensive evaluations show that the proposed task-constrained learning (i) outperforms existing methods, especially in dealing with faces with severe occlusion and pose variation, and (ii) reduces model complexity drastically compared to the state-of-the-art method based on cascaded deep model [21].
- Research Article
5
- 10.1088/1742-6596/887/1/012079
- Aug 1, 2017
- Journal of Physics: Conference Series
In order to solve face alignment more effectively, we proposed a multi-task deep convolution network for face alignment, which achieves good performance even in the case of large pose variations and severe occlusion. Instead of dealing with face alignment as a single task, we jointly trained the auxiliary task of pose estimation together with face alignment to guide the distribution of facial points. By doing so we are able to 1) avoid trapping in the local optimum due to the inaccurate face boxes, 2) improve the robustness in dealing with faces with pose variation and severe occlusion. Compared with the traditional methods, our method also improves the accuracy by providing better initialization instead of mean shape. Extensive experiments show that our method has great performance on various benchmarks.
- Conference Article
11
- 10.1109/icpr.2018.8545898
- Aug 1, 2018
In the past decades, face alignment has been studied widely, but it has long been impeded by the problem of pose variation. Recent studies show that pose information used as additional source of information can help address the above problem. In this paper, we adopt a multi-task cascaded CNNs based framework for simultaneous face detection, dense face alignment and fine head pose estimation. Especially, our framework exploits the inherent correlation between face alignment and fine head pose estimation to boost up landmark detection robustness in the case of various poses. Experiments show that our method not only demonstrates real-time performance for face detection, dense face alignment and fine head pose estimation, but also outperforms most state-of-the-art methods for face alignment on the challenging 300-W benchmark. Especially in the case of large pose variations, it achieves outstanding results.
- Conference Article
8
- 10.1109/icb45273.2019.8987289
- Jun 1, 2019
Face detection and alignment are considered as two independent tasks and conducted sequentially in most face applications. However, these two tasks are highly related and they can be integrated into a single model. In this paper, we propose a novel single-shot detector for joint face detection and alignment, namely FLDet, with remarkable performance on both speed and accuracy. Specifically, the FLDet consists of three main modules: Rapidly Digested Backbone (RDB), Lightweight Feature Pyramid Network (LFPN) and Multi-task Detection Module (MDM). The RDB quickly shrinks the spatial size of feature maps to guarantee the CPU real-time speed. The LFPN integrates different detection layers in a top-down fashion to enrich the feature of low-level layers with little extra time overhead. The MDM jointly performs face and landmark detection over different layers to handle faces of various scales. Besides, we introduce a new data augmentation strategy to take full usage of the face alignment dataset. As a result, the proposed FLDet can run at 20 FPS on a single CPU core and 120 FPS using a GPU for VGA-resolution images. Notably, the FLDet can be trained end-to-end and its inference time is invariant to the number of faces. We achieve competitive results on both face detection and face alignment benchmark datasets, including AFW, PASCAL FACE, FDDB and AFLW.
- Conference Article
11
- 10.1109/crv.2019.00027
- May 1, 2019
ComunicaciĂł presentada a: 16th Conference on Computer and Robot Vision (CRV) celebrat del 29 al 31 de maig de 2019 a Kingston, CanadĂ .
- Conference Article
138
- 10.1109/iccvw.2017.188
- Oct 1, 2017
We show how a simple convolutional neural network (CNN) can be trained to accurately and robustly regress 6 degrees of freedom (6DoF) 3D head pose, directly from image intensities. We further explain how this FacePoseNet (FPN) can be used to align faces in 2D and 3D as an alternative to explicit facial landmark detection for these tasks. We claim that in many cases the standard means of measuring landmark detector accuracy can be misleading when comparing different face alignments. Instead, we compare our FPN with existing methods by evaluating how they affect face recognition accuracy on the IJB-A and IJB-B benchmarks: using the same recognition pipeline, but varying the face alignment method. Our results show that (a) better landmark detection accuracy measured on the 300W benchmark does not necessarily imply better face recognition accuracy. (b) Our FPN provides superior 2D and 3D face alignment on both benchmarks. Finally, (c), FPN aligns faces at a small fraction of the computational cost of comparably accurate landmark detectors. For many purposes, FPN is thus a far faster and far more accurate face alignment method than using facial landmark detectors.
- Research Article
48
- 10.1109/tip.2020.2991549
- Jan 1, 2020
- IEEE Transactions on Image Processing
Facial expression recognition, face synthesis, and face alignment are three coherently related tasks and can be solved in a joint framework. To achieve this goal, in this paper, we propose a novel end-to-end deep learning model by exploiting the expression code, geometry code and generated data jointly for simultaneous pose-invariant facial expression recognition, face image synthesis, and face alignment. The proposed deep model enjoys several merits. First, to the best of our knowledge, this is the first work to address these three tasks jointly in a unified deep model to complement and enhance each other. Second, the proposed model can effectively disentangle the global and local identity representation from different expression and geometry codes. As a result, it can automatically generate facial images with different expressions under arbitrary geometry codes. Third, these three tasks can further boost their performance for each other via our model. Extensive experimental results on three standard benchmarks demonstrate that the proposed deep model performs favorably against state-of-the-art methods on the three tasks.
- Research Article
4
- 10.1016/j.imavis.2016.04.012
- May 3, 2016
- Image and Vision Computing
Robust face alignment and tracking by combining local search and global fitting
- Book Chapter
4
- 10.1007/978-3-030-03000-1_3
- Dec 15, 2018
Face alignment in a video is an important research area in computer vision and can provides strong support for video face recognition, face animation, etc. It is different from face alignment in a single image where each face is regarded as an independent individual. For the latter, lack of amount of information makes the face alignment an under-determined problem although good results have been obtained by using prior information and auxiliary models. For the former, temporal and spatial relations are among faces in a video. These relations can impose constraints among multiple face images each other and help to improve alignment performance. In the chapter, definition of face alignment in a video and its significance are described. Methods for face alignment in a video are divided into three kinds: face alignment using image alignment algorithms, joint alignment of face images, and face alignment using temporal and spatial continuities. The first kind of face alignment is studied and some of surveys have described the work. The chapter will mainly focus on joint face alignment and face alignment using temporal and spatial continuities. Herein, some representative methods are described, and some factors influencing alignment performance are analyzed. Then the state-of-the-art methods are described and the future trends of face alignment in a video are discussed.
- Conference Article
114
- 10.1109/cvpr.2016.373
- Jun 1, 2016
Face alignment or facial landmark detection plays an important role in many computer vision applications, e.g., face recognition, facial expression recognition, face animation, etc. However, the performance of face alignment system degenerates severely when occlusions occur. In this work, we propose a novel face alignment method, which cascades several Deep Regression networks coupled with De-corrupt Autoencoders (denoted as DRDA) to explicitly handle partial occlusion problem. Different from the previous works that can only detect occlusions and discard the occluded parts, our proposed de-corrupt autoencoder network can automatically recover the genuine appearance for the occluded parts and the recovered parts can be leveraged together with those non-occluded parts for more accurate alignment. By coupling de-corrupt autoencoders with deep regression networks, a deep alignment model robust to partial occlusions is achieved. Besides, our method can localize occluded regions rather than merely predict whether the landmarks are occluded. Experiments on two challenging occluded face datasets demonstrate that our method significantly outperforms the state-of-the-art methods.
- Research Article
13
- 10.1016/j.neucom.2018.10.068
- Nov 5, 2018
- Neurocomputing
Face alignment by Component Adaptive Mechanism
- Conference Article
- 10.1109/cds52072.2021.00040
- Jan 1, 2021
Face Alignment, focusing on detecting facial landmarks of input faces, has been widely applied in criminal investigation and security verification system. However, face alignment keeps an incredibly challenging task in computer version due to occlusion, pose variation and unsuited initial shape. Many related methods have been proposed to tackle these problems. Many traditional methods, such as Face Alignment by Coarse-to-Fine Shape Searching, focus on initial face shape optimization. Recent studies pay more attention to making advantages of deep learning approaches such as cascaded classifiers, convolutional neural networks and multitask learning methods, which have already achieved outstanding performance on prevalent datasets. In this paper, we summarize the mainstream face alignment models and analyze the corresponding advantages and shortages. Then, we also discuss further application and development of face alignment algorithms in the future.
- Book Chapter
2
- 10.1007/978-3-319-70090-8_31
- Jan 1, 2017
Most face alignment approaches perform landmark detection over the entire face. However, it has been shown that the difficulty for landmark detection is unbalanced among different facial parts. Thus, in this paper, we propose a novel region-based facial landmark detection algorithm based on a two-level convolutional neural networks (CNNs). In the first level, we partition the whole face into four regions including three facial components (eyebrow-eyes, nose, and mouth) and the face contour. Regions are detected through an improved CNN model which is incorporated with a feature fusion scheme. To simultaneously detect three facial components and face contour landmarks, a novel weighted loss function combining bounding box regression with landmark localization is presented. In the second level, the landmarks are separately detected for three facial components. Experimental results on the public benchmarks demonstrate the superiority of the proposed algorithm over several state-of-the-art face alignment algorithms.
- Book Chapter
42
- 10.1007/978-3-319-46454-1_50
- Jan 1, 2016
Face alignment, which is the task of finding the locations of a set of facial landmark points in an image of a face, is useful in widespread application areas. Face alignment is particularly challenging when there are large variations in pose (in-plane and out-of-plane rotations) and facial expression. To address this issue, we propose a cascade in which each stage consists of a mixture of regression experts. Each expert learns a customized regression model that is specialized to a different subset of the joint space of pose and expressions. The system is invariant to a predefined class of transformations (e.g., affine), because the input is transformed to match each expert's prototype shape before the regression is applied. We also present a method to include deformation constraints within the discriminative alignment framework, which makes our algorithm more robust. Our algorithm significantly outperforms previous methods on publicly available face alignment datasets.
- Conference Article
73
- 10.5244/c.29.130
- Jan 1, 2015
In this paper we propose a supervised initialization scheme for cascaded face alignment based on explicit head pose estimation. We first investigate the failure cases of most state of the art face alignment approaches and observe that these failures often share one common global property, i.e. the head pose variation is usually large. Inspired by this, we propose a deep convolutional network model for reliable and accurate head pose estimation. Instead of using a mean face shape, or randomly selected shapes for cascaded face alignment initialisation, we propose two schemes for generating initialisation: the first one relies on projecting a mean 3D face shape (represented by 3D facial landmarks) onto 2D image under the estimated head pose; the second one searches nearest neighbour shapes from the training set according to head pose distance. By doing so, the initialisation gets closer to the actual shape, which enhances the possibility of convergence and in turn improves the face alignment performance. We demonstrate the proposed method on the benchmark 300W dataset and show very competitive performance in both head pose estimation and face alignment.