Language Modeling Approach Research Articles

We argue that expert finding is sensitive to multiple document features in an organizational intranet. These document features include multiple levels of associations between experts and a query topic from sentence, paragraph, up to document levels, document authority information such as the PageRank, indegree, and URL length of documents, and internal document structures that indicate the experts’ relationship with the content of documents. Our assumption is that expert finding can largely benefit from the incorporation of these document features. However, existing language modeling approaches for expert finding have not sufficiently taken into account these document features. We propose a novel language modeling approach, which integrates multiple document features, for expert finding. Our experiments on two large scale TREC Enterprise Track datasets, i.e., the W3C and CSIRO datasets, demonstrate that the natures of the two organizational intranets and two types of expert finding tasks, i.e., key contact finding for CSIRO and knowledgeable person finding for W3C, influence the effectiveness of different document features. Our work provides insights into which document features work for certain types of expert finding tasks, and helps design expert finding strategies that are effective for different scenarios. Our main contribution is to develop an effective formal method for modeling multiple document features in expert finding, and conduct a systematic investigation of their effects. It is worth noting that our novel approach achieves better results in terms of MAP than previous language model based approaches and the best automatic runs in both the TREC2006 and TREC2007 expert search tasks, respectively.

Read full abstract

In many probabilistic modeling approaches to Information Retrieval we are interested in estimating how well a document model "fits" the user's information need (query model). On the other hand in statistics, goodness of fit tests are well established techniques for assessing the assumptions about the underlying distribution of a data set. Supposing that the query terms are randomly distributed in the various documents of the collection, we actually want to know whether the occurrences of the query terms are more frequently distributed by chance in a particular document. This can be quantified by the so-called goodness of fit tests. In this paper, we present a new document ranking technique based on Chi-square goodness of fit tests. Given the null hypothesis that there is no association between the query terms q and the document d irrespective of any chance occurrences, we perform a Chi-square goodness of fit test for assessing this hypothesis and calculate the corresponding Chi-square values. Our retrieval formula is based on ranking the documents in the collection according to these calculated Chi-square values. The method was evaluated over the entire test collection of TREC data, on disks 4 and 5, using the topics of TREC-7 and TREC-8 (50 topics each) conferences. It performs well, outperforming steadily the classical OKAPI term frequency weighting formula but below that of KL-Divergence from language modeling approach. Despite this, we believe that the technique is an important non-parametric way of thinking of retrieval, offering the possibility to try simple alternative retrieval formulas within goodness-of-fit statistical tests' framework, modeling the data in various ways estimating or assigning any arbitrary theoretical distribution in terms.

Read full abstract

Language Modeling Approach Research Articles

Related Topics

Articles published on Language Modeling Approach

Integrating multiple document features in language models for expert finding

Context-Aware Person Identification in Personal Photo Collections

An empirical study of gene synonym query expansion in biomedical information retrieval

A language modeling framework for expert finding

Fast exact maximum likelihood estimation for mixture of language model

Smoothing document language models with probabilistic term count propagation

Ordinal Regression for Information Retrieval

Language‐modeling kernel based approach for information retrieval

In this issue

Machine learning for Asian language text classification

A novel dependency language model for information retrieval

Generating gene summaries from biomedical literature: A study of semi-structured summarization

An empirical study of query expansion and cluster-based retrieval in language modeling approach

Table extraction for answer retrieval

N-gram Adaptation with Dynamic Interpolation Coefficient Using Information Retrieval Technique

Parsimonious translation models for information retrieval

A goodness of fit test approach in information retrieval

Incorporating context within the language modeling approach for ad hoc information retrieval

Backoff hierarchical class n-gram language models: effectiveness to model unseen events in speech recognition

Two-stage statistical language models for text database selection

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Language Modeling Approach Research Articles

Related Topics

Articles published on Language Modeling Approach

Integrating multiple document features in language models for expert finding

Context-Aware Person Identification in Personal Photo Collections

An empirical study of gene synonym query expansion in biomedical information retrieval

A language modeling framework for expert finding

Fast exact maximum likelihood estimation for mixture of language model

Smoothing document language models with probabilistic term count propagation

Ordinal Regression for Information Retrieval

Language‐modeling kernel based approach for information retrieval

In this issue

Machine learning for Asian language text classification

A novel dependency language model for information retrieval

Generating gene summaries from biomedical literature: A study of semi-structured summarization

An empirical study of query expansion and cluster-based retrieval in language modeling approach

Table extraction for answer retrieval

N-gram Adaptation with Dynamic Interpolation Coefficient Using Information Retrieval Technique

Parsimonious translation models for information retrieval

A goodness of fit test approach in information retrieval

Incorporating context within the language modeling approach for ad hoc information retrieval

Backoff hierarchical class n-gram language models: effectiveness to model unseen events in speech recognition

Two-stage statistical language models for text database selection