Contrastive learning for hierarchical topic modeling

Pengbo Mao,Hegang Chen,Yanghui Rao,Haoran Xie,Fu Lee Wang

doi:10.1016/j.nlp.2024.100058

Abstract

Topic models have been widely used in automatic topic discovery from text corpora, for which, the external linguistic knowledge contained in Pre-trained Word Embeddings (PWEs) is valuable. However, the existing Neural Topic Models (NTMs), particularly Variational Auto-Encoder (VAE)-based NTMs, suffer from incorporating such external linguistic knowledge, and lacking of both accurate and efficient inference methods for approximating the intractable posterior. Furthermore, most existing topic models learn topics with a flat structure or organize them into a tree with only one root node. To tackle these limitations, we propose a new framework called as Contrastive Learning for Hierarchical Topic Modeling (CLHTM), which can efficiently mine hierarchical topics based on inputs of PWEs and Bag-of-Words (BoW). Experiments show that our model can automatically mine hierarchical topic structures, and have a better performance than the baseline models in terms of topic hierarchical rationality and flexibility.

Full Text