Robust Estimation of Topic Summaries Leveraging Word Frequency and Exclusivity

Authored by: Edoardo M. Airoldi , David M. Blei , Elena A. Erosheva , Stephen E. Fienberg , Jonathan M. Bischof , Edoardo M. Airoldi

Handbook of Mixed Membership Models and Their Applications

Print publication date:  November  2014
Online publication date:  November  2014

Print ISBN: 9781466504080
eBook ISBN: 9781466504097
Adobe ISBN:

10.1201/b17520-18

 Download Chapter

 

Abstract

An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. However, the current practice in mixed membership models of text (Blei et al., 2003) of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential use of words across topics. Words that are both common and exclusive to a theme are more effective at characterizing the topical content of such a theme. We consider a setting where professional editors have annotated documents to a collection of topic categories, organized into a tree, in which leaf-nodes correspond to the most specific topics. Each document is annotated to multiple categories, at different levels of the tree. We introduce hierarchical Poisson convolution (HPC) as a model to analyze annotated documents in this setting. The model leverages the structure among categories defined by professional editors to infer a clear semantic description for each topic in terms of words that are both frequent and exclusive. We develop a parallelized Hamiltonian Monte Carlo sampler that allows the inference to scale to millions of documents.

 Cite
Search for more...
Back to top

Use of cookies on this website

We are using cookies to provide statistics that help us give you the best experience of our site. You can find out more in our Privacy Policy. By continuing to use the site you are agreeing to our use of cookies.