Abstract
One of the important procedures during the operation of an Institutional Repository System (IRS) is to categorize and index the submitted digital objects. Based on current practice, human catalogers are frequently involved in this process to make accurate categorization. Funded by NLM development grant, we are developing an UMLS-based indexing tool. The tool will be integrated or plug-in into most IRS, and categorize and assign MeSH headings and keywords to text-based digital objects automatically. The project has been supported under NLM Knowledge Management & Applied Informatics Grants (1 G08 LM008704-01)
Project Website: http://umlsresearch.welch.jhmi.edu
Description
The operation of institutional repository applications includes indexing and categorizing submitted digital objects in order for users to find them through appropriate search engines. These digital objects include images, full text, data sets, and applications. Current institutional repository systems (IRS) such as DSpace, Fedora, and others provide proper interface for users to type in their favorite keywords during their submission, or, human catalogers are involved in the indexing process after the submission. Although these IRS systems provide options to select standard thesaurus or controlled vocabulary systems, none of them provide the interfaces to integrate with or connect to standard thesaurus systems.
NLM started a project called Medical Text Indexer (MTI) in 1997 and developed an effective system for recommending MeSH terms to provide assistance to human indexers in automatic or semi-automatic fashions.
The operation of MTI would be involved in fourteen servers and implementation of Prolog, C and Java. To avoid reinventing the wheel and take advantage of what NLM has achieved, the project team has digested and analyzed the MTI application, and focused our effort on developing a practical application for the purpose of enhancing current IRS search functions with MeSH keywords capability. There are several differences between MTI and the indexing tool we are currently working on:
Open-sources based
We have been using Java as the primary software tool exclusively during the development, thus, it can be easily exported and implemented on other systems;
Simplification
In order to improve the system performance, we have simplified some procedures during the indexing process, for example, at the stage of “Restricted to MeSH”, we have used plain database query searching methods, instead using the UMLS semantic network. We will make sure the indexing outcome will not be suffered too much by these simplifications;
Remote and Plug-in operation
One important goal for this project is to make sure the application being developed could be added into current operational institutional repository applications such as DSpace and Fedora from other health or medical libraries. Thus, current operations won’t be affected by the addition and indexing and searching would be greatly improved. The indexing process itself would be involved in interacting UMLS database, PubMed citation database, and other resources and sub-processes. We are anticipating the future system could be running as an indexing center and be used remotely by other sites through Internet.

