AnyFace++: A Unified Framework for Free-style Text-to-Face Synthesis and Manipulation.

Jianxin Sun,Muyi Sun,Yunfan Liu,Qiyao Deng,Qi Li,Zhenan Sun

doi:10.1109/tpami.2023.3345866

Abstract

Human faces contain rich semantic information that could hardly be described without a large vocabulary and complex sentence patterns. However, most existing text-to-image synthesis methods could only generate meaningful results based on limited sentence templates with words contained in the training set, which heavily impairs the generalization ability of these models. In this paper, we define a novel 'free-style' text-to-face generation and manipulation problem, and propose an effective solution, named AnyFace++, which is applicable to a much wider range of open-world scenarios. The CLIP model is involved in AnyFace++ for learning an aligned language-vision feature space, which also expands the range of acceptable vocabulary as it is trained on a large-scale dataset. To further improve the granularity of semantic alignment between text and images, a memory module is incorporated to convert the description with arbitrary length, format, and modality into regularized latent embeddings representing discriminative attributes of the target face. Moreover, the diversity and semantic consistency of generation results are improved by a novel semi-supervised training scheme and a series of newly proposed objective functions. Compared to state-of-the-art methods, AnyFace++ is capable of synthesizing and manipulating face images based on more flexible descriptions and producing realistic images with higher diversity.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

AnyFace++: A Unified Framework for Free-style Text-to-Face Synthesis and Manipulation.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence

Lead the way for us

Journal: IEEE transactions on pattern analysis and machine intelligence	Publication Date: Jan 1, 2024
Citations: 1

Similar Papers

CMAFGAN: A Cross-Modal Attention Fusion based Generative Adversarial Network for attribute word-to-face synthesis
Xiaodong Luo ... Xinyue Tan
Knowledge-Based Systems | VOL. 255
Xiaodong Luo, et. al.Xiaodong Luo ... Xinyue Tan
24 Aug 2022
Knowledge-Based Systems | VOL. 255

AnyFace: A Data-Centric Approach For Input-Agnostic Face Detection
Askat Kuzdeuov ... Huseyin Atakan Varol
-
Askat Kuzdeuov, et. al.Askat Kuzdeuov ... Huseyin Atakan Varol
01 Feb 2023
01 Feb 2023

Which CNNs and Training Settings to Choose for Action Unit Detection? A Study Based on a Large-Scale Dataset
Mina Bishay ... Mohammad Mavadati
-
Mina Bishay, et. al.Mina Bishay ... Mohammad Mavadati
15 Dec 2021
15 Dec 2021

Are open set classification methods effective on large-scale datasets?
Ayesha Gonzales ... Christopher Kanan
-
Ayesha Gonzales, et. al.Ayesha Gonzales ... Christopher Kanan
04 Sep 2020
04 Sep 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

AnyFace++: A Unified Framework for Free-style Text-to-Face Synthesis and Manipulation.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on pattern analysis and machine intelligence