Abstract
For computational purposes documents or other objects are most often represented by a collection of individual attributes that may be strings or numbers. Such attributes are often called features and success in solving a given problem can depend critically on the nature of the features selected to represent documents. Feature selection has received considerable attention in the machine learning literature. In the area of document retrieval we refer to feature selection as indexing. Indexing has not traditionally been evaluated by the same methods used in machine learning feature selection. Here we show how indexing quality may be evaluated in a machine learning setting and apply this methodology to results of the Indexing Initiative at the National Library of Medicine.
Full text
PDF




Selected References
These references are in PubMed. This may not be the complete list of references from this article.
- Aronson A. R., Bodenreider O., Chang H. F., Humphrey S. M., Mork J. G., Nelson S. J., Rindflesch T. C., Wilbur W. J. The NLM Indexing Initiative. Proc AMIA Symp. 2000:17–21. [PMC free article] [PubMed] [Google Scholar]
- Aronson A. R. The effect of textual variation on concept based information retrieval. Proc AMIA Annu Fall Symp. 1996:373–377. [PMC free article] [PubMed] [Google Scholar]
- Bodenreider O., Nelson S. J., Hole W. T., Chang H. F. Beyond synonymy: exploiting the UMLS semantics in mapping vocabularies. Proc AMIA Symp. 1998:815–819. [PMC free article] [PubMed] [Google Scholar]
- Haynes R. B., McKibbon K. A., Walker C. J., Ryan N., Fitzgerald D., Ramsden M. F. Online access to MEDLINE in clinical settings. A study of use and usefulness. Ann Intern Med. 1990 Jan 1;112(1):78–84. doi: 10.7326/0003-4819-112-1-78. [DOI] [PubMed] [Google Scholar]
- Hersh W. R., Hickam D. H., Haynes R. B., McKibbon K. A. A performance and failure analysis of SAPHIRE with a MEDLINE test collection. J Am Med Inform Assoc. 1994 Jan-Feb;1(1):51–60. doi: 10.1136/jamia.1994.95236136. [DOI] [PMC free article] [PubMed] [Google Scholar]
- McCray A. T., Nelson S. J. The representation of meaning in the UMLS. Methods Inf Med. 1995 Mar;34(1-2):193–201. [PubMed] [Google Scholar]
- Salton G. Developments in automatic text retrieval. Science. 1991 Aug 30;253(5023):974–980. doi: 10.1126/science.253.5023.974. [DOI] [PubMed] [Google Scholar]
- Wilbur W. J. Boosting naïve Bayesian learning on a large subset of MEDLINE. Proc AMIA Symp. 2000:918–922. [PMC free article] [PubMed] [Google Scholar]
