S3 - Visual Conceptual Blending with Large-scale Language and Vision Models

Songwei Ge

doi:10.48448/kx4k-1p61

S3 - Visual Conceptual Blending with Large-scale Language and Vision Models

Songwei Ge

https://doi.org/10.48448/kx4k-1p61

Copy DOI

Publication Date: Sep 5, 2021

#Language Models #Image Generation Models + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description of the blend of the two using a language model. We then generate a visual depiction of the blend using a text-based image generation model. Quantitative and qualitative evaluations demonstrate the superiority of language models over classical methods for conceptual blending, and of recent large-scale image generation models over prior models for the visual depiction.

Full Text