Skip to main content
Proceedings of the AMIA Symposium logoLink to Proceedings of the AMIA Symposium
. 1998:765–769.

A randomized controlled trial of automated term composition.

P L Elkin 1, K R Bailey 1, C G Chute 1
PMCID: PMC2232145  PMID: 9929322

Abstract

OBJECTIVE: To compare the ability of an Automated Term Composition (ATC) algorithm with non-compositional mappings to provide coverage (exact mappings to a controlled vocabulary) for a randomly selected set of free text entries which were entered as headings to the Impression section of the clinical notes system at the Mayo Foundation. We also compare the results of four evaluators to determine the inter-observer variability and the variance between term sets, with respect to the accuracy of the mappings and the reliability of the failure analysis. METHODS: From a corpus of approximately 1,000,000 unique terms entered into the Impression/Report/Plan section of the clinical notes system in the calendar year 1997, we randomly selected 1,000 terms. We then further randomized these 1,000 terms into two groups of 500 (Sets A and B). We constructed two copies of the same term matching interface, one without ATC (alpha) and one with ATC (beta). We took four expert Indexers and assigned them to one of the following tasks. The first reviewer (R1) compared set A using the alpha program and then set B using the beta program (R1(Aalpha + Bbeta)). The second compared set A using the alpha program and then set B using the alpha program (R2(A + B) alpha). The third compared set B using the beta program and then set A using the beta program (R3(B + A) beta). The fourth compared set A using the beta program and then set B using the alpha program (R4(Abeta + Balpha)). RESULTS: The program with Automated Term Composition mapped 540 out of the 1,000 Concepts correctly (54.0%). The same program without ATC mapped only 276 out of the 1,000 Concepts correctly (27.6%). Therefore the program with ATC was significantly more effective at matching concepts in our problem lists than the same search engine without ATC (p < 0.0001; McNemar Method). These figures result from the comparison of the alpha program with the beta program by reviewers one and four. Failure analysis showed that with the alpha version 425 out of the 724 mismatches were because a base concept was missing from the retrieval set (58.7%) and 299 mismatches were from missing qualifiers or modifiers or both (41.3%). In the beta version of the program (with ATC) 340 out of the 460 mismatches were secondary to there being a missing base concept in the retrieval set (73.9%) and only 120 mismatches due to missing modifiers and or qualifiers (26.1%). CONCLUSIONS: Automated term composition provided significantly better coverage of a randomly chosen set of patient problems, diagnosed at the Mayo Clinic during the 1997 calendar year, when compared with the same information retrieval system without ATC. We believe that these results speak further to the excellent content coverage provided by the UMLS metathesaurus. These authors believe that increased structure, normalization of UMLS content and semantics, and better tools to make use of the currently available content such as automated term composition, are what is needed to leverage the production of commercially viable tools that provide access to controlled vocabularies for medicine.

Full text

PDF
765

Selected References

These references are in PubMed. This may not be the complete list of references from this article.

  1. Baud R. H., Rassinoux A. M., Scherrer J. R. Natural language processing and semantical representation of medical texts. Methods Inf Med. 1992 Jun;31(2):117–125. [PubMed] [Google Scholar]
  2. Chute C. G., Elkin P. L. A clinically derived terminology: qualification to reduction. Proc AMIA Annu Fall Symp. 1997:570–574. [PMC free article] [PubMed] [Google Scholar]
  3. Cimino J. J., Clayton P. D., Hripcsak G., Johnson S. B. Knowledge-based approaches to the maintenance of a large controlled medical terminology. J Am Med Inform Assoc. 1994 Jan-Feb;1(1):35–50. doi: 10.1136/jamia.1994.95236135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Cimino J. J., Elkin P. L., Barnett G. O. As we may think: the concept space and medical hypertext. Comput Biomed Res. 1992 Jun;25(3):238–263. doi: 10.1016/0010-4809(92)90041-8. [DOI] [PubMed] [Google Scholar]
  5. Elkin P. L., Mohr D. N., Tuttle M. S., Cole W. G., Atkin G. E., Keck K., Fisk T. B., Kaihoi B. H., Lee K. E., Higgins M. C. Standardized problem list generation, utilizing the Mayo canonical vocabulary embedded within the Unified Medical Language System. Proc AMIA Annu Fall Symp. 1997:500–504. [PMC free article] [PubMed] [Google Scholar]
  6. Evans D. A., Cimino J. J., Hersh W. R., Huff S. M., Bell D. S. Toward a medical-concept representation language. The Canon Group. J Am Med Inform Assoc. 1994 May-Jun;1(3):207–217. doi: 10.1136/jamia.1994.95236153. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Humphreys B. L., McCray A. T., Cheh M. L. Evaluating the coverage of controlled health data terminologies: report on the results of the NLM/AHCPR large scale vocabulary test. J Am Med Inform Assoc. 1997 Nov-Dec;4(6):484–500. doi: 10.1136/jamia.1997.0040484. [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Kushniruk A., Patel V., Cimino J. J., Barrows R. A. Cognitive evaluation of the user interface and vocabulary of an outpatient information system. Proc AMIA Annu Fall Symp. 1996:22–26. [PMC free article] [PubMed] [Google Scholar]
  9. Lindberg D. A., Humphreys B. L., McCray A. T. The Unified Medical Language System. Methods Inf Med. 1993 Aug;32(4):281–291. doi: 10.1055/s-0038-1634945. [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Rector A. L., Nowlan W. A. The GALEN project. Comput Methods Programs Biomed. 1994 Oct;45(1-2):75–78. doi: 10.1016/0169-2607(94)90020-5. [DOI] [PubMed] [Google Scholar]
  11. Tuttle M. S., Sherertz D. D., Erlbaum M. S., Sperzel W. D., Fuller L. F., Olson N. E., Nelson S. J., Cimino J. J., Chute C. G. Adding your terms and relationships to the UMLS Metathesaurus. Proc Annu Symp Comput Appl Med Care. 1991:219–223. [PMC free article] [PubMed] [Google Scholar]
  12. Yang Y., Chute C. G. A schematic analysis of the Unified Medical Language System. Proc Annu Symp Comput Appl Med Care. 1991:204–208. [PMC free article] [PubMed] [Google Scholar]

Articles from Proceedings of the AMIA Symposium are provided here courtesy of American Medical Informatics Association

RESOURCES