DE eng

Search in the Catalogues and Directories

Hits 1 – 2 of 2

1
A Probabilistic Model for Latent Semantic Indexing
In: http://crd.lbl.gov/~cding/papers/lsilong6.pdf (2005)
BASE
Show details
2
A Probabilistic Model for Dimensionality Reduction in Information Retrieval and Filtering
In: http://www.nersc.gov/research/SCG/cding/papers_ps/jsis2.ps (2001)
Abstract: Dimension reduction methods, such as Latent Semantic Indexing (LSI), when applied to semantic spaces built upon text collections, improve information retrieval, information filtering and word sense disambiguation. A new dual probability model based on similarity concepts is introduced to explain the observed success. Semantic associations can be quantitatively characterized by their statistical significance, the likelihood. Semantic dimensions containing redundant and noisy information can be separated out and should be ignored because their contribution to the overall statistical significance is negative, giving rise to LSI: LSI is the optimal solution of the model. The peak in likelihood curve indicates the existence of an intrinsic semantic dimension. The importance of LSI dimensions follows the Zipf-distribution, indicating that LSI dimensions represent latent concepts. Document frequency of words follow the Zipf distribution, and the number of distinct words follows log-normal distribution. Experiments on four standard document collections both confirm and illustrate the results and concepts presented here.
Keyword: retrieve information by exactly matching query keywords to; such as Internet search engines
URL: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.21.7962
http://www.nersc.gov/research/SCG/cding/papers_ps/jsis2.ps
BASE
Hide details

Catalogues
0
0
0
0
0
0
0
Bibliographies
0
0
0
0
0
0
0
0
0
Linked Open Data catalogues
0
Online resources
0
0
0
0
Open access documents
2
0
0
0
0
© 2013 - 2024 Lin|gu|is|tik | Imprint | Privacy Policy | Datenschutzeinstellungen ändern