Topic Models for Unsupervised Cluster Matching

Tomoharu Iwata,Naonori Ueda,Tsutomu Hirao

doi:10.1109/tkde.2017.2778720

Abstract

We propose topic models for unsupervised cluster matching, which is the task of finding matching between clusters in different domains without correspondence information. For example, the proposed model finds correspondence between document clusters in English and German without alignment information, such as dictionaries and parallel sentences/documents. The proposed model assumes that documents in all languages have a common latent topic structure, and there are potentially infinite number of topic proportion vectors in a latent topic space that is shared by all languages. Each document is generated using one of the topic proportion vectors and language-specific word distributions. By inferring a topic proportion vector used for each document, we can allocate documents in different languages into common clusters, where each cluster is associated with a topic proportion vector. Documents assigned into the same cluster are considered to be matched. We develop an efficient inference procedure for the proposed model based on collapsed Gibbs sampling. The effectiveness of the proposed model is demonstrated with real data sets including multilingual corpora of Wikipedia and product reviews.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Topic Models for Unsupervised Cluster Matching

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering

Lead the way for us

Journal: IEEE Transactions on Knowledge and Data Engineering	Publication Date: Apr 1, 2018
Citations: 40

Similar Papers

Cluster Matching for Discrete Data with Multiple Domains with out Alignment of Information
-
International Journal of Innovative Technology and Exploring Engineering | VOL. 8
--
26 Jul 2019
International Journal of Innovative Technology and Exploring Engineering | VOL. 8

Constrained common cluster based model for community detection in temporal and multiplex networks
Pengfei Jiao ... Wenjun Wang
Neurocomputing | VOL. 275
Pengfei Jiao, et. al.Pengfei Jiao ... Wenjun Wang
12 Sep 2017
Neurocomputing | VOL. 275

Quality Control for Hierarchical Classification with Incomplete Annotations
Masafumi Enomoto ... Takeshi Okadome
-
Masafumi Enomoto, et. al.Masafumi Enomoto ... Takeshi Okadome
01 Jan 2020
01 Jan 2020

Interlayer Link Prediction in Multiplex Social Networks Based on Multiple Types of Consistency Between Embedding Vectors.
Rui Tang ... Shuyu Jiang
IEEE Transactions on Cybernetics | VOL. 53
Rui Tang, et. al.Rui Tang ... Shuyu Jiang
01 Apr 2023
IEEE Transactions on Cybernetics | VOL. 53

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Topic Models for Unsupervised Cluster Matching

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering