Logo image
Using PageRank for Characterizing Topic Quality in LDA
Conference proceeding

Using PageRank for Characterizing Topic Quality in LDA

Sujatha Das Gollapalli, Xiao-li Li and ACM
Proceedings of the 2018 ACM SIGIR International Conference on Theory of Information Retrieval, pp.115-122
ACM Conferences
ICTIR '18: The 2018 ACM SIGIR International Conference on the Theory of Information Retrieval
10/09/2018

Abstract

Information systems -- Information retrieval -- Document representation -- Document topic models
Topic models based on Latent Dirichlet Allocation (LDA) are employed effectively in various information retrieval and data mining tasks. Despite their popularity and wide-spread application, the question of assessing the quality of topics extracted by LDA models is still not completely resolved. While various measures have been proposed to quantify the thematic coherence and interpretability of a topic extracted by LDA, they do not address this problem sufficiently. We observe that existing quality measures select top topic words based on their topic-word co-occurrence without considering word co-occurrences within the same context. We incorporate precisely this information by constructing topic-specific graphs capturing neighborhood of words in an LDA modeled corpus. Next, the PageRank algorithm is applied on these graphs to assign word importance scores based on centrality. We propose two measures to compute topic quality: (1) the Aggregate PageRank of Top-words of a topic and (2) the PageRank Centralization Index of a topic-specific word graph. Our experiments across three datasets show that unlike existing quality measures, our proposed measures are able to identify topics that are discriminative as well as interpretable and yield superior performance on both classification and intruder word identification tasks.

Metrics

1 Record Views

Details

Logo image