Abstract
Published clinical case reports are a valuable yet underutilized source of evidence for drug repurposing. However, systematically identifying relevant reports remains a challenge due to the volume of literature and the diversity of candidate compounds. We present TheraMind, an AI system that leverages large language models (LLMs) to automate the identification and analysis of case reports supporting potential drug repurposing for non-small cell lung cancer (NSCLC). Our system screened 10,023 PubMed-indexed case reports across 18 candidate drugs using coordinated data extraction and standardized four-question prompts assessing diagnosis, drug administration, discontinuation, and clinical outcomes. We employed three evaluation strategies, rule-based classifiers, single-model validators, and a majority-vote ensemble integrating GPT-40-mini, Gemini-2.0-Flash, and LLaMA-3-8B. The ensemble approach achieved 92% recall and 99.7% specificity in detecting clinically relevant reports. Structured outputs included patient demographics, therapeutic responses, and case summaries. This LLM-driven framework offers a scalable approach to accelerate drug repurposing by mining real-world evidence from unstructured clinical literature.
Subject terms: Non-small-cell lung cancer, Virtual screening
Introduction
Non-small-cell lung cancer (NSCLC) is the leading cause of cancer mortality worldwide, accounting for over 1.7 million deaths annually1. While advances in targeted therapies and immunotherapy have improved outcomes for some patients, the five-year survival rate for metastatic NSCLC remains dismal at ~3%2. The traditional drug development pipeline is slow, costly, and high-risk, often requiring over a decade, $2–3 billion per approved compound, and facing failure rates nearing 90% in oncology trials3. In this context, drug repurposing offers a promising and efficient strategy to accelerate therapeutic discovery for NSCLC1. Given NSCLC’s intrinsic heterogeneity and the plasticity of cancer stem cells, which contribute to therapy resistance, repurposing existing compounds may help overcome key barriers to durable treatment response4. Moreover, repurposing FDA-approved drugs potentially offers safer therapeutic options5, especially for treating refractory patients with failed prior therapies.
Recent research from our group and others has identified potential repurposing drugs for treating NSCLC. Our previous studies established an AI pipeline to identify multi-omics gene signatures combined with CRISPR-Cas9/RNAi and drug screening data to discover potential new or repurposed drugs for treating lung or breast cancer6–12. Network-based methods using gene co-expression and drug-gene interaction networks have identified novel repurposing candidates13. MAPK pathway targeting approaches have shown therapeutic potential14. Inflammatory pathway modulation has emerged as another promising avenue, with anti-inflammatory drugs showing potential for repurposing in NSCLC treatment15. Clinical trials of repurposed agents have demonstrated encouraging results, with drugs like selinexor showing efficacy in KRAS-mutant NSCLC16. While these approaches accelerate the drug screening process, concrete clinical evidence substantiated by published case reports is often needed for oncologists to adopt a repurposed drug in the clinic or launch a clinical trial to test its efficacy. Nevertheless, efficient retrieval of relevant reports is a daunting task due to the volume of the literature and candidate compounds generated from recent AI screening.
Biomedical text mining has evolved rapidly, with recent advances in information extraction from clinical records17–19. Clinical natural language processing systems have demonstrated effectiveness in extracting structured information from unstructured clinical texts20–22. Large language models (LLMs) have demonstrated remarkable capabilities in biomedical text analysis and clinical decision support23–25. Recent comprehensive reviews have highlighted their transformative potential in healthcare, with applications spanning clinical documentation, medical education, and diagnostic support26–28. LLMs encode substantial clinical knowledge, with models like Med-PaLM achieving state-of-the-art performance on medical question answering benchmarks29. Medical domain-specific models have shown particular promise, with biomedical language models demonstrating enhanced performance on clinical natural language processing tasks30–32. Entity extraction pipelines using LLMs have shown promising results for mining clinical research data from textual records33. However, systematic evaluations reveal limitations in clinical decision-making applications, including inaccurate diagnoses and failure to follow clinical guidelines34. Currently, there is no LLM-driven AI tool to determine whether a case report is relevant to repositioning a drug for a new indication.
Multi-modal and ensemble approaches have proven particularly effective in decision making, including biomedical applications35–37. Ensemble learning techniques have demonstrated superior performance in disease prediction and clinical decision support38–40. Recent studies show that multimodal LLMs can integrate diverse data types to facilitate comprehensive patient assessments41. However, current systems face significant challenges, including hallucinations, bias, and limited transparency42–44.
To address these challenges, we developed an LLM-ensembled based framework (TheraMind) to retrieve case reports that support the repositioning of a drug for treating NSCLC. TheraMind includes the following steps: (1) extract all published case reports from PubMed based on a keyword search of drug candidates and new indications. (2) Scraped abstracts and case reports (if available) are input into our multi-LLM ensemble architecture employing GPT-40-mini, Gemini-2.0-Flash45, and Llama-3-8B46 in parallel evaluation schemes anchored by few-shot examples. Each report is interrogated using four targeted prompts examining: (i) patient NSCLC diagnosis, (ii) use of the study drug in treating NSCLC, (iii) premature discontinuation of therapy, and (iv) a favorable outcome. (3) Based on these responses, we implemented three classification strategies to determine if a case report supports the repositioning of a drug for treating NSCLC: rule-based decision trees, individual LLM validators, and majority-vote ensembles. Finally, for relevant cases, TheraMind extracts structured patient demographics and generates concise summaries, providing complete evaluation trails throughout the entire process.
Results
Using the Methods-described pipeline, we retrieved 10,023 unique PubMed citations across 18 candidate drugs. After screening for coverage and data availability, 10 compounds had published case reports and advanced to full-text analysis and LLM-based processing. The distribution of retrieved abstracts was heavily skewed toward Daunorubicin (n = 8265), followed by Ivermectin (n = 1095), Aztreonam (n = 197), Disulfiram (n = 371), Midostaurin (n = 59), Puromycin (12), Penfluridol (n = 9), Anisomycin (n = 3), and Lestaurtinib and Homosalate (1 each; Table 1). The system identified 26 case reports as clinically relevant for NSCLC repurposing from this corpus.
Table 1.
Summary of the 18 candidate drugs identified in our previous studies, along with the relevance and distribution of case reports on these drugs retrieved from PubMed by TheraMind
| Drug | Any relevant abstract/report extracted | Number of abstracts retrieved |
|---|---|---|
| Daunorubicin | Yes | 8265 |
| Ivermectin | Yes | 1095 |
| Aztreonam | Yes | 197 |
| Disulfiram | Yes | 371 |
| Midostaurin | Yes | 59 |
| Puromycin | Yes | 12 |
| Penfluridol | Yes | 9 |
| Anisomycin | Yes | 3 |
| Lestaurtinib | Yes | 1 |
| Homosalate | Yes | 1 |
| Trichostatin | No | 0 |
| FK 888 | No | 0 |
| U-0126 | No | 0 |
| PD-198306 | No | 0 |
| PQ-401 | No | 0 |
| ZM-306416 | No | 0 |
| Piperlongumine | No | 0 |
| BX-912 | No | 0 |
As shown, Daunorubicin yielded the largest number of abstracts (n = 8265), followed by Ivermectin (n = 1095), Aztreonam (n = 197), and Disulfiram (n = 371).
Worked example: structured prompting of a report
To illustrate the system’s reasoning step-by-step, we provide a worked example47 from an ivermectin case report retrieved by TheraMind (Fig. 1). After extraction, the structured prompts were applied sequentially to the text (for full prompts refer to Supplementary File 2).
Fig. 1. Presented is a sample case report excerpt and structured prompt–response framework used to evaluate model outputs.
The excerpt highlights a case where ivermectin was combined with dichloroacetate, omeprazole, and tamoxifen in cancer treatment. Structured prompts queried four dimensions, diagnosis confirmation, drug administration, treatment discontinuation, and clinical outcome, with corresponding model responses shown. This framework illustrates how TheraMind systematically interrogates case reports to extract clinically meaningful insights.
This worked example47 demonstrates how TheraMind applies its structured prompts and reasoning to a single case report. In the next section, we discuss the approaches used to classify the relevance of case reports for NSCLC based on these responses.
Determining the relevance of case reports for NSCLC drug repurposing
Approach 1: decision tree classification
The first classification approach employs a transparent, rule-based decision framework that systematically evaluates case reports based on the responses to the four prompts. Each case report is assessed for evidence of NSCLC diagnosis, drug administration for NSCLC treatment, treatment discontinuation status, and clinical outcome. Reports are classified as relevant for drug repurposing only when they satisfy all four requirements: confirmed NSCLC diagnosis, documented drug administration, absence of premature treatment discontinuation, and favorable patient outcome. The decision tree method in Fig. 2 was implemented as “hard-coded” reasoning logic for the LLM-based agent.
Fig. 2. Detailed decision flow for NSCLC case report classification.
This flowchart outlines the sequential decision-making process used to evaluate the relevance of a clinical case report for drug repositioning in NSCLC treatment. Starting with the collection of responses from multiple LLMs, the process asks key clinical questions: whether the patient was diagnosed with NSCLC, whether the drug was used in NSCLC treatment, whether the drug was discontinued, and if the patient’s outcome was favorable. Based on the yes/no responses to these prompts, the system determines if the case study is relevant for repositioning the drug for NSCLC treatment or not, using a decision-tree logic.
This deterministic approach provides complete transparency and reproducibility, as every classification decision can be traced to specific clinical criteria. The method serves as a robust baseline for comparison with more sophisticated classification techniques, offering a clear clinical rationale for each inclusion or exclusion decision.
The decision tree approach achieved 77.8% precision with 80.8% recall when using Gemini-based responses, outperforming GPT and Llama. Gemini-based method exhibited near-perfect specificity (99.9%), indicating exceptional ability to correctly identify non-relevant cases while maintaining strong performance in capturing truly relevant clinical evidence (Table 2).
Table 2.
Model performance metrics grouped by approach
| Classifier/Model | TP | FP | FN | TN | Precision | Recall | Sensitivity | Specificity | F1 Score | Accuracy |
|---|---|---|---|---|---|---|---|---|---|---|
| Approach 1: decision tree | ||||||||||
| Decision tree/Gemini | 21 | 6 | 5 | 7515 | 77.8 | 80.8 | 80.8 | 99.9 | 79.2 | 99.9 |
| Decision tree/GPT | 13 | 12 | 13 | 7509 | 52.0 | 50.0 | 50.0 | 99.8 | 51.0 | 99.7 |
| Decision tree/Llama | 17 | 45 | 9 | 7476 | 27.4 | 65.4 | 65.4 | 99.4 | 38.6 | 99.3 |
| Approach 2: individual model classifier | ||||||||||
| Gemini-based/Gemini | 23 | 13 | 3 | 7508 | 63.9 | 88.5 | 88.5 | 99.8 | 74.2 | 99.8 |
| Gemini-based/GPT | 14 | 21 | 12 | 7500 | 40.0 | 53.8 | 53.8 | 99.7 | 45.9 | 99.6 |
| Gemini-based/Llama | 20 | 38 | 6 | 7483 | 34.5 | 76.9 | 76.9 | 99.5 | 47.6 | 99.4 |
| GPT-based/Gemini | 24 | 21 | 2 | 7500 | 53.3 | 92.3 | 92.3 | 99.7 | 67.6 | 99.7 |
| GPT-based/GPT | 17 | 22 | 9 | 7499 | 43.6 | 65.4 | 65.4 | 99.7 | 52.3 | 99.6 |
| GPT-based/Llama | 22 | 49 | 4 | 7472 | 31.0 | 84.6 | 45.4 | 99.3 | 45.4 | 99.3 |
| Llama-based/Gemini | 22 | 16 | 4 | 7505 | 57.9 | 84.6 | 84.6 | 99.8 | 68.8 | 99.7 |
| Llama-based/GPT | 10 | 12 | 16 | 7509 | 45.5 | 38.5 | 38.5 | 99.8 | 41.7 | 99.6 |
| Llama-based/Llama | 21 | 59 | 5 | 7462 | 26.3 | 80.8 | 80.8 | 99.2 | 39.6 | 99.2 |
| Approach 3: majority vote ensemble | ||||||||||
| Majority Vote/Gemini | 24 | 20 | 2 | 7501 | 54.5 | 92.3 | 92.3 | 99.7 | 68.6 | 99.7 |
| Majority Vote/GPT | 15 | 21 | 11 | 7500 | 41.7 | 57.7 | 57.7 | 99.7 | 48.4 | 99.6 |
| Majority Vote/Llama | 22 | 44 | 4 | 7477 | 33.3 | 84.6 | 84.6 | 99.4 | 47.8 | 99.4 |
Approach 2: individual model classifier
Next, we tested the reasoning capability of LLMs without “hard-coding” the decision tree logic. The second classification strategy utilizes advanced language model capabilities to conduct comprehensive evaluations of case report relevance. As illustrated in Fig. 3, this approach integrates multiple sources of clinical evidence, including initial diagnostic assessments, treatment protocols, discontinuation patterns, and patient outcomes, to form a holistic evaluation framework for each case report.
Fig. 3. Illustrated is the detailed decision flow for case report classification using individual LLMs.
This flowchart outlines the sequential decisio3n-making process used to evaluate the relevance of a clinical case report for drug repositioning in NSCLC treatment. Starting with the collection of responses from multiple LLMs, the process asks key clinical questions: whether the patient was diagnosed with NSCLC, whether the drug was used in NSCLC treatment, whether the drug was discontinued, and whether the patient’s outcome was favorable. Based on the yes/no responses to these prompts, the system determines if the case study is relevant for repositioning the drug for NSCLC treatment or not.
The methodology incorporates an efficient screening protocol that prioritizes computational resources while maintaining classification accuracy. Cases without a confirmed NSCLC diagnosis are systematically excluded early in the process, ensuring focused analysis on clinically relevant reports. Subsequently, qualified reports undergo detailed evaluation where specialized LLMs assess the totality of clinical evidence to render definitive relevance determinations. This approach maintains complete auditability through comprehensive documentation of all evaluation steps and clinical rationale. Every classification decision preserves the underlying evidence trail, enabling thorough review and validation of the assessment process.
Performance analysis revealed considerable variation across different model configurations, as detailed in Table 2. The optimal configuration was a GPT-based classifier for Gemini responses, achieving 92.3% recall with 53.3% precision, demonstrating strong capability for identifying relevant clinical cases while maintaining acceptable accuracy standards. All models achieved specificity exceeding 99.2%, reflecting robust performance in correctly identifying non-relevant cases. Compared with Approach 1, individual LLMs can classify the relevance of case reports in NSCLC drug repositioning with higher recall and comparable specificity, indicating that LLMs can reason based on the syntax of prompts similarly to the decision tree logic.
Approach 3: majority vote approach
To further enhance the performance of LLMs, the third classification strategy implements a consensus-based decision framework that synthesizes judgments from multiple independent classifiers to improve overall classification accuracy. As demonstrated in Fig. 4, this ensemble methodology treats each classifier’s assessment as an equal contribution to the final relevance determination, requiring agreement from at least two out of three classifiers to establish case report relevance for NSCLC drug repurposing applications.
Fig. 4. Illustrated is the majority vote ensemble classification workflow for NSCLC case reports.
This diagram illustrates the multi-model approach for determining the relevance of a case study in repositioning a drug for NSCLC treatment. Individual classifiers based on GPT, Gemini, and Llama generate preliminary classifications, which are then aggregated through a decision module. A relevance-threshold check, requiring at least two classifiers to agree, determines whether the case study is deemed relevant or not for drug repositioning in NSCLC.
The consensus mechanism effectively mitigates individual classifier limitations while leveraging collective analytical strengths. Cases achieving majority support (at least 2 out of 3 models flagging the same case as relevant) are retained for further analysis, while those lacking sufficient consensus are systematically excluded. This threshold-based approach maintains optimal balance between sensitivity and specificity, ensuring comprehensive capture of relevant clinical evidence while minimizing false positive classifications.
Performance evaluation demonstrated that the Majority Vote Ensemble approach successfully balanced the strengths observed in previous methodologies, as shown in Table 2. The optimal configuration based on Gemini responses achieved 92.3% recall with 54.5% precision, indicating strong capability for identifying relevant cases while maintaining reasonable accuracy standards. Specificity remained consistently above 99.4% across all configurations, demonstrating robust performance in correctly identifying non-relevant clinical reports.
Given the research priority of maximizing clinical case capture, the Majority Vote Ensemble approach represents the optimal balance for drug repurposing applications, outperforming both Approaches 1 and 2. This configuration ensures comprehensive identification of relevant NSCLC clinical evidence while maintaining manageable candidate volumes for subsequent expert clinical review.
To illustrate this process, consider the ivermectin case report47 example previously described. While Llama misclassified this report, both GPT and Gemini produced correct classification (refer to github link in data availability section Data/03_RelevanceClassification/MajorityVote/pooled_gpt_links.csv). Under the majority vote framework, consensus from GPT and Gemini was sufficient to retain the case as relevant, demonstrating how the ensemble approach mitigates individual model errors and leverages collective strength to improve overall reliability.
Data structuring for patient demographics
The relevant case reports identified in our previous analyses undergo structured metadata extraction and concise summarization. Our objective in this final step is twofold: (1) to harvest key patient demographics in a machine-readable JSON schema and (2) to distill each report’s primary purpose into a pithy statement of no more than fifteen words. By transforming narrative case studies into standardized data objects and terse synopses, we enable downstream analyses such as cohort characterization or hypothesis generation to proceed without manual parsing of unstructured text.
To accomplish these tasks, each model receives two fixed-form prompts. The first instructs the LLM to emit a JSON object containing exactly six fields: age, gender, race, history, condition, and causative_agent, with no extraneous text, commentary, or formatting. The second prompt requests a one-sentence summary of the report’s primary purpose, capped at fifteen words. This dual-prompt design ensures both structural consistency for demographic attributes and brevity in summarization (Fig. 5).
Fig. 5. Displayed are the inputs, evaluation rubric, and performance outcomes for the metadata-extraction stage.
Panel 1 presents the two prompts instructing the LLMs to (i) extract six patient attributes in a fixed JSON schema and (ii) provide a one-sentence summary of the paper’s primary purpose. Panel 2 shows the five-level scoring rubric (0 = incorrect, 0.25 = mostly incorrect, 0.50 = partially correct, 0.75 = mostly correct, 1.0 = fully correct), applied independently to accuracy and completeness. Panel 3 reports evaluator mean ± SD scores (0–1 scale) for accuracy (green) and completeness (red) across GPT, Gemini, and Llama, highlighting GPT’s consistently superior metadata quality relative to the other models. Error bars occasionally exceed 1 due to high means combined with residual variance. Strong inter-rater agreement was confirmed by Cohen’s κ (0.50–1.00), and paired t-tests indicated that GPT significantly outperformed Gemini (p < 0.01) and Llama (p < 0.001). Panel 4 provides example output showing the extracted drug name, PubMed ID, and concise summary, illustrating the type and format of metadata used in the evaluation.
The performance of the LLMs was assessed by two independent evaluators with relevant background (L.L., Biology and V.M., Computer Science). Specifically, GPT and Gemini achieved perfect inter-rater agreement for accuracy (Cohen’s kappa κ = 1.00 each), while Llama achieved substantial agreement (κ = 0.78). For completeness, agreement was moderate to substantial (κ = 0.50 for GPT, κ = 0.75 for Gemini, κ = 0.80 for Llama). GPT achieved the highest metadata-extraction performance, with mean accuracy 0.96 (±0.07) and completeness 0.91 (±0.11). Gemini followed with accuracy 0.75 (±0.20) and completeness 0.70 (±0.11). Llama-3-8B trailed with accuracy 0.53 (±0.27) and completeness 0.56 (±0.22). Paired t-tests confirmed that GPT significantly outperformed Gemini (p < 0.01) and Llama (p < 0.001). These findings underscore GPT’s superior ability to extract structured patient information and craft concise study summaries, validating its role as the primary engine for metadata curation in our NSCLC case-report workflow.
Exploration of TheraMind’s scope beyond lung cancer
To further examine TheraMind’s adaptability, we extended our analysis beyond lung cancer and applied the updated structured prompts to breast cancer (see Supplementary File 2 for the full breast cancer prompt set). We evaluated five candidate compounds of interest, ZM-306416, PD19830, BRD-K12244279, pilocarpine, and tremorine, which were previously discovered as potential new/repositioning drugs for treating breast cancer by our group9,10. Despite comprehensive querying, TheraMind did not retrieve any case reports linking these drugs to breast cancer treatment. This outcome reflects the absence of published evidence in the biomedical literature, not a limitation of the platform. It also demonstrates that TheraMind faithfully reports when no case-level support is available. These findings indicate that the system is readily extensible to breast cancer, although in this instance we could not perform downstream validation due to the lack of reported case data.
Discussion
Most NSCLC patients acquire resistance to first-line chemotherapy or targeted therapies within 4–12 months48–52. Despite second-line therapy with newer medications, the 5-year survival rate is 6–7% for metastatic lung cancer53–57. Oncologists do not know how to treat refractory NSCLC patients after exhausting all treatment options. There is an unmet clinical need to inform oncologists of safe, repositioning drugs for treating refractory NSCLC patients. An advantage of repositioning drugs is the known safety, which is a top priority when oncologists choose a therapeutic option for refractory cancer patients. Since the toxicity profiles of the existing drugs are well established, oncologists can bypass preclinical studies and phase I clinical trials and go directly to phase II clinical trials to test the efficacy of the new indications of these drugs.
Recent network-based approaches have demonstrated promising results in identifying repurposing candidates through systematic analysis of molecular interactions58, Computational frameworks have shown success in predicting drug-disease associations59. Our group established an AI pipeline to analyze multi-omics networks combined with CRISPR-Cas9/RNAi and drug screening data. Using this AI pipeline, we identified 18 potential new or repurposed drugs that can reverse gene signatures associated with poor patient prognosis, drug resistance, and proliferation in NSCLC6–8,11,12. In early discussions with oncologists, a consistent message emerged: trust in computational drug-repurposing insights hinges on evidence of real-world clinical success. Specifically, they emphasized the need for documented case reports demonstrating that a repurposed drug was administered to a patient with NSCLC, yielded tangible benefits, and caused minimal adverse effects. This feedback informed our decision to systematically mine, curate, and audit clinical case reports, prioritizing real-world evidence over in silico predictions or preclinical data alone.
The conventional approach to identifying relevant case reports relies on keyword-based searches in PubMed, followed by manual review and curation, a process that is both time- and labor-intensive. TheraMind addresses this evidentiary bottleneck by functioning as an intelligent agent that bridges the gap between hypothesis-generating algorithms and actionable clinical evidence. Given a list of candidate drugs and a target indication (e.g., NSCLC), TheraMind automatically retrieves case reports using an integrated scraping engine, then evaluates their relevance to drug repurposing through structured prompts. For each relevant report, it generates an auditable dossier summarizing key clinical details, including patient demographics, treatment course, and outcomes. This capability transforms drug repurposing from a theoretical exercise into a clinician-ready evidence package. In this study, TheraMind retrieved 10,023 case reports from PubMed across 18 candidate drugs, identifying 26 reports as relevant to NSCLC. By automating literature triage and evidence synthesis, TheraMind offers a scalable approach to accelerate clinical trial planning and inform therapeutic strategies for patients with refractory cancers who lack standard treatment options.
A central innovation of TheraMind is the multi-LLM ensemble architecture. Prior biomedical text systems frequently rely on a single LLM, leading to propagation of errors and hallucinations60–62. Recent reviews have highlighted the predominance of single-model approaches in medical applications, with limited exploration of ensemble methods63. Our approach aligns with emerging evidence that ensemble learning can significantly improve reliability in healthcare applications64–66. Our Majority Vote Ensemble method, leveraging GPT-40-mini, Gemini-2.0-Flash, and Llama-3-8B, integrated with in-depth therapeutically relevant prompts, curbs hallucinations and yields a consensus recall of 92% with 99.7% specificity. This performance exceeds benchmarks reported in recent systematic evaluations of LLMs in healthcare67–69.
The decision to adopt a multi-LLM ensemble architecture was not merely for redundancy, but to capitalize on the complementary strengths and failure modes of each model. No single LLM can robustly handle all the linguistic nuance, domain complexity, and noise typical of biomedical case reports. For instance, GPT-4’s left-to-right token processing excels at structured reasoning and crisp logic but may miss contextual cues embedded in convoluted syntax; Gemini’s bidirectional transformer better captures embedded clauses and elliptical references; while Llama-3’s lightweight footprint enables faster, privacy-conscious processing at the edge, albeit with slightly reduced reliability. By integrating these agents under a coordinated ensemble strategy, with early exit logic, rule-based pruning, and majority-vote aggregation, we not only elevate recall and precision but also enhance system fault tolerance. Each model serves as both a validator and counterbalance to the others, enabling the system to suppress individual hallucinations while rescuing borderline true positives that a single-agent design might discard. In high-stakes clinical workflows, such layered consensus is not just preferable, it is essential.
The integration of rule-based and machine learning approaches represents another key innovation. Once each report is distilled to four binary flags (diagnosis, drug use, discontinuation, outcome), even a deterministic decision tree approaches LLM-level accuracy. This hybrid approach addresses key concerns about AI reliability in clinical settings70–72. Recent work has emphasized the importance of explainable AI in healthcare, particularly for high-stakes clinical decisions73–75. Our transparent decision-making process enables full auditability while maintaining high performance.
Although LLMs are often treated as black boxes, our study demonstrates that LLMs can be decomposed into reasoning frameworks that mirror traditional modular pipelines. Each binary judgment, whether confirming diagnosis or assessing treatment outcome, acts as a discrete reasoning unit, and the sequence of prompt-based inferences effectively mimics the logic gates of classical decision trees. By structuring prompts to elicit both categorical labels and brief justifications, we create an interpretable audit trail that resembles intermediate variables in a deterministic program. Moreover, when LLMs are consistently guided by few-shot examples and enforced output formats, their behavior approximates that of symbolic rule-followers, capable of chaining evidence across a multi-step decision path. In this way, LLMs do not replace structured reasoning; rather, they instantiate it with far greater flexibility, scaling across varied writing styles and idiosyncratic clinical descriptions without sacrificing interpretability or auditability.
Our calibrated ensemble converts unstructured NSCLC case reports into clinician-grade evidence, satisfying the real-world standard of “show me the patient” while retaining 92% of relevant studies and filtering noise with 99.7% specificity. This performance compares favorably with recent benchmarks for LLMs in clinical applications76–78. The system addresses key limitations identified in systematic reviews of AI applications in medicine, including lack of transparency, limited clinical validation, and insufficient consideration of real-world deployment challenges79–81.
The demonstrated effectiveness of TheraMind in mining published case reports suggests significant potential for expansion to more comprehensive datasets. Future applications could extend beyond published literature to include analysis of electronic medical records (EMRs) and private patient databases, enabling more robust identification of successful drug-patient combinations and treatment outcomes. Such expansion would provide unprecedented granularity in documenting real-world therapeutic responses, potentially capturing nuanced treatment patterns and rare but clinically significant drug effects that are unnoticed without powerful AI mining. This comprehensive approach could fundamentally transform evidence generation for drug repurposing by creating exhaustive repositories of patient-level therapeutic responses across diverse clinical contexts.
However, the integration of private patient data introduces substantial privacy and regulatory challenges that must be carefully addressed. Current regulatory frameworks, including the General Data Protection Regulation (GDPR)82 and the Health Insurance Portability and Accountability Act (HIPAA)83, impose strict limitations on sharing sensitive patient information with third-party LLM service providers. Large-scale language models have demonstrated the potential to memorize and inadvertently disclose fragments of training data84–86, creating significant risks for patient confidentiality. Simple de-identification techniques are frequently insufficient, as re-identification attacks have demonstrated that anonymized data can be linked with external datasets to re-establish individual identities87,88. Consequently, more rigorous privacy-preserving methods, such as Differential Privacy89, must be employed to safely integrate patient health records without compromising individual confidentiality. Recent advances in generating differentially private prompts for in-context learning90,91 offer promising approaches, though generating high-quality data that are simultaneously factually accurate and privacy-compliant remains challenging. Additional security considerations include protecting proprietary prompts from extraction attacks72,86,92, which could compromise intellectual property and operational logic. Potential mitigation strategies include deploying locally hosted open-source models within trusted computing environments and implementing robust defensive measures such as intention analysis and goal prioritization93,94.
Our team has developed a web-portal (https://sostos-ai.com) for our discovered candidate repositioning drugs, with information on mechanisms of action, evidence on published case reports, FDA-approved administration routes and dosage, and genomic characteristics of patient response to the candidate drugs, for oncologists to select drug candidates for treating refractory NSCLC patients. Next steps include extending the framework to additional cancer types and incorporating multimodal data sources. The growing field of AI-driven clinical decision support presents opportunities for expanding our approach to broader clinical applications85–87. Future directions include coupling our curated evidence layer with our AI pipeline, utilizing a multi-omics network approach for biomarker/drug discovery6,7,11, to establish a fully integrated bench-to-bedside repurposing pipeline.
TheraMind demonstrates the potential of multi-LLM ensemble systems to comprehensively mine published case reports for drug repurposing evidence in NSCLC. Our ensemble approach achieved 92% recall and 99.7% specificity in identifying clinically relevant reports, providing oncologists with structured, auditable evidence to support therapeutic decision-making. By transforming unstructured clinical literature into standardized datasets with patient demographics and outcome summaries, this framework addresses the critical gap between computational drug discovery and clinical adoption. The transparent, scalable methodology offers a practical solution for accelerating evidence-based drug repurposing across cancer types, potentially reducing the time and cost barriers that limit therapeutic innovation for patients with limited treatment options.
Methods
System overview of TheraMind
TheraMind transforms a simple list of candidate drugs into a vetted collection of NSCLC case reports ready for repurposing analysis. As depicted in Fig. 6, the workflow begins with large-scale web scraping (Step 1), routes the retrieved text to several language models that generate responses with the help of task-specific prompts (Step 2), and then applies a dedicated classification layer that keeps only those reports whose answers confirm drug use and a favorable NSCLC outcome (Step 3). These records are compiled into a structured dataset that includes key patient demographics and a concise statement of study purpose (Step 4). Every processing block is self-contained, allowing any segment to be rerun without repeating earlier steps and ensuring both traceability and ease of maintenance. The outcome is a transparent, end-to-end system that converts raw literature into clinician-ready evidence for therapeutic decision-making.
Fig. 6. Overview of the data extraction and decision module in TeraMind.
This diagram illustrates the integrated pipeline for evaluating case studies in terms of their relevance to drug repositioning for NSCLC. Data from diverse sources—including PubMed, FDA records, and clinical trials—is first extracted and used to generate model responses via multiple LLMs. These responses are then processed through a classifier framework that combines ensemble methods and decision trees to deliver a final classification on whether a given case study is relevant. Further, TheraMind can also generate structured data, including patient demographics and a summary of the relevant case reports.
Data sources and case selection
To construct a corpus of real-world evidence for NSCLC, we compiled a list of 18 systemic agents that we previously identified as potential new or repurposed drugs for treating NSCLC. For each compound, authoritative chemical identity and synonym information were obtained from PubChem88,89. Regulatory context—including approved dosage forms, marketing status, and proprietary (brand) names—was gathered from the U.S. FDA Orange Book90 and Drugs@FDA databases91. Evidence of investigational use was captured by cross-referencing the same compounds against ClinicalTrials.gov (completed and active studies) and the FDA Orphan Drug Designations catalogue; these registries were used solely to characterize the clinical footprint of each agent and did not contribute primary text for model analysis.
Using the consolidated synonym sheet (generic and exact-match brand names), we searched PubMed with the built-in Case Reports article type filter to identify reports with full or partial English content published between January 2000 and December 2024. Articles were included if they reported administration of a study drug, while reviews, animal studies, and meta-analyses were excluded. The search yielded 10,023 unique citations, of which case reports linked to ten drugs met the inclusion criteria and were advanced to full-text retrieval and subsequent large language model interrogation.
Data extraction from PubMed
Our data extraction pipeline (Fig. 7) retrieved a corpus providing diverse clinical contexts for systematic evidence mining from PubChem, PubMed, FDA, and patent databases. Table 3 summarizes the characteristics of each data source, primary endpoints extracted, and key filtering criteria applied during selection. In this study, TheraMind identified 10,023 unique PubMed citations across the 18 candidate drugs identified from our previous studies6–8,11,12. Following assessment of citation coverage and data availability, 10 compounds were reported in published case reports and advanced to full-text analysis and LLM-based processing (Fig. 7). This dataset formed the foundation for subsequent multi-model analysis and classification workflows.
Fig. 7. Outlined is the five-stage drug-centric sourcing workflow in our case-report retrieval system.
1 Standardize Drug List: Candidate drugs are matched to their unique PubChem compound identifiers. 2 Add Regulatory Context: U.S. FDA approval status and marketed brand names are appended to each compound record. 3 Collect Evidence “Hints”: additional evidence is harvested from ClinicalTrials.gov and the FDA Orphan-Drug database to prioritize under-explored indications. 4 Build Synonym Sheet: A comprehensive synonym table covering generic and brand names as well as spelling variants is generated to maximize recall when querying biomedical literature. 5 Search PubMed Case Reports: the synonym sheet drives automated PubMed searches that return titles, snippets, and URLs for all matching clinical case reports, forming the initial corpus for downstream LLM analysis.
Table 3.
Overview of data sub-corpora, primary endpoints, and key filters
| Sub-corpus | Primary endpoint | Key filters |
|---|---|---|
| PubChem drug lookup | Unique PubChem CID; IUPAC name; summary text; PNG structure | Clue-IO Broad IDs matched via PubChem REST API |
| FDA approvals | NDA/ANDA numbers; dosage form/route; marketing status; sponsor | Include “Active” only; drop discontinued entries |
| ClinicalTrials.gov | NCT identifiers mentioning NSCLC and a listed drug | Bulk CSV download; regex match against drug list |
| Orphan-drug designations | Year of designation; indication text | FDA OPD list (US only); deduplicate by drug |
| Published Case reports (PubMed) | Article title; snippet; URL | Query “drug AND ‘Case Reports’”; paginated at 200/page |
| Patent mentions | Count of US and WO patent families | Retrieved via PubChem patents endpoint; collapse duplicates |
Text extraction and preprocessing
Abstracts and HTML full-text pages were fetched over HTTPS using Python’s requests library (version 2.31.0), with a 1-second inter-request delay to respect server policies. We parsed content with BeautifulSoup 4 (version 4.12.2), isolating the “Abstract” section and subsequent case-report narrative. All text was normalized (whitespace collapse, removal of non-ASCII characters) and tokenized using the Natural Language Toolkit (NLTK, version 3.8.1); intermediate outputs were checkpointed in CSV format via pandas (version 2.0.3) to enable safe restart and audit trails.
Model selection and configuration
We employed three large language models in parallel: OpenAI GPT-40-mini (GPT-40-mini-2024-04-09), Google Gemini-2.0-Flash (gemini-1.0-2.0-Flash), and Meta Llama-3-8B (Meta-Llama-3-8B-Instruct).
Prompt engineering for LLM queries
To obtain consistent judgements across GPT-4, Gemini-2.0-Flash, and Llama-3-8B, each case report is presented through a streamlined prompt suite built around four yes-or-no questions that probe (i) confirmed NSCLC diagnosis, (ii) use of the study drug in treating NSCLC, (iii) premature discontinuation of therapy, and (iv) a favorable outcome (Fig. 8). Every response is followed by a single-sentence justification, encouraging the models to supply both a machine-readable label and a brief clinical rationale. For few-shot prompting, we employed two curated examples: one clearly positive and one clearly negative. Negative cases were defined as case reports failing any of the core criteria (confirmed NSCLC diagnosis, drug usage specific to NSCLC, absence of premature discontinuation, and favorable outcome). We deliberately excluded ambiguous or borderline examples from the few-shot context to maintain consistency across models and ensure reproducibility. This decision also reflected technical constraints: since Llama 3–8B has a shorter context window than other models, we standardized on two shots for all systems to preserve fairness and comparability. For full prompts please refer to Supplementary File 2.
Fig. 8. Presented are the simplified prompts supplied to the LLMs.
four rigorously worded yes/no questions with justifications that probe (1) NSCLC diagnosis, (2) drug administration, (3) treatment discontinuation, and (4) favorable outcome. Each prompt instructs models to answer solely “Yes” or “No” + Justification (if needed), establishing a standardized framework for comparing LLM performance in medical text extraction.
Multi‑model few‑shot inference for generating ground-truth
Using the same template of the prompts and few-shots, each clinical narrative (abstract plus full case report when available) is submitted in parallel to the three LLMs. The few-shot examples serve as calibration points, helping the models make consistent decisions and reducing variation in how they interpret different writing styles. Because the prompt and context are held constant, the resulting “Yes” or “No” answers and single-line justifications can be evaluated side by side, providing an immediate view of agreement or divergence among the models.
The questions are queried in successive order, beginning with only the diagnosis prompt (“Was the patient diagnosed with NSCLC?”) to Gemini, GPT-4, and Llama-3-8B simultaneously. This parallel call serves as an efficient filter: if all three models unanimously return “No,” indicating that the narrative lacks any evidence of NSCLC diagnosis, the pipeline immediately abandons further questioning for that record. In practice, this early-exit gating conserves nearly 90% of total API calls without diminishing sensitivity for true positives, since out-of-scope documents rarely contain subsequent treatment details or outcome statements.
For narratives where at least one model affirms NSCLC diagnosis, we proceed to the remaining three prompts: drug administration, discontinuation status, and treatment outcome, submitting them sequentially to each model using the same two-shot context. By maintaining an identical prompt structure across all questions, we ensure that each model’s responses are directly comparable, facilitating downstream ensembling and discrepancy analysis. Figure 9 displays a visual overview of this workflow, and a detailed algorithm is included as Algorithm 1 in Supplementary File 1. Here, all LLM-generated responses for case reports with an NSCLC diagnosis were verified by three individuals with a background in biological sciences (L.L.), biomedical engineering (S.M.) and computer science (V.M.).
Fig. 9. Outlined is the multi-model workflow for ground truth generation.
This diagram begins with input prompts and scraped content, which are augmented with two-shot examples before being passed to GPT, Gemini, and Llama. Guided by contextual information, the models generate responses regarding NSCLC diagnosis, drug usage, discontinuation, and treatment outcomes, ultimately contributing to a comprehensive evaluation of each case study.
Model querying infrastructure
GPT-4 and Gemini were accessed via their respective REST APIs with exponential back-off to handle rate limits. The Llama model was hosted locally on a Debian 12 server (AMD EPYC 74F3 CPU, 503 GiB RAM, four NVIDIA RTX 6000 Ada GPUs) and accessed via FastAPI over an SSH-tunneled connection. This hybrid cloud-local approach reduced overall runtime by ~70% compared to sequential querying while maintaining reproducibility.
Large language model usage documentation
Large language models were used specifically for: (1) Text classification: Determining relevance of case reports for NSCLC drug repurposing based on four predefined clinical criteria. (2) Metadata extraction: Extracting structured patient demographics (age, gender, race, medical history, condition, causative agent) from unstructured case report text. (3) Summary generation: Creating concise one-sentence study purpose statements (≤15 words). All LLM outputs were logged verbatim and subjected to human evaluation. Models were not used for data interpretation, statistical analysis, or manuscript writing.
Classification methodology
Approach 1: decision tree classification
We implemented a transparent, deterministic decision tree operating on binary judgments from the multi-model extraction stage. Reports were classified as relevant only if they exactly matched the pattern (NSCLC diagnosis: yes, drug administration: yes, treatment discontinuation: no, favorable outcome: yes). Refer to Algorithm 2 in the Supplementary File 1. This approach provided a zero-parameter, fully auditable baseline for comparison.
Approach 2: individual model classifier
A dedicated LLM performed a comprehensive review of candidate case reports by evaluating the collected binary responses and free-text justifications from the multi-model extraction stage. Reports lacking a confirmed NSCLC diagnosis across all models were immediately labeled “Not Relevant” (refer to Algorithm 3 in the Supplementary File 1).
Approach 3: majority vote ensemble
Individual judgments from the three independent classifiers were fused using majority voting. Reports were deemed relevant only if at least two of the three models agreed on relevance, balancing precision and recall while suppressing individual model errors (refer to Algorithm 4 in the Supplementary File 1).
Performance evaluation and statistical analysis
Two independent evaluators, one with a biology background and one with a computer science background, assessed model outputs. For each case report, evaluators first established the ideal answer, then judged the similarity of model responses using a standardized 0–1 scale. Accuracy was defined as the correctness of the result, while completeness measured the presence of all relevant details. Inter-rater reliability was assessed using Cohen’s kappa coefficient. Mean values and standard deviations for both metrics were calculated to summarize performance across the corpus.
Evaluation procedure
Two independent human evaluators scored every output: one with a biology background (L.L.) and one with a computer science background (V.M.). Each evaluated both the JSON and summary responses. Accuracy was defined as factual correctness of each extracted attribute, and completeness as the presence of all required elements. Summaries were assessed for factual accuracy, grammaticality, and adherence to the ≤15-word limit. Disagreements were resolved through discussion after initial scoring.
Scoring rubric and scale
Outputs were rated on a five-level rubric: 0 (incorrect/missing), 0.25 (mostly incorrect), 0.50 (partially correct), 0.75 (mostly correct), and 1.0 (fully correct). Results are reported as mean ± SD on the 0–1 scale.
Inter-rater agreement
Cohen’s kappa was calculated to quantify agreement. Agreement was strong overall, ranging from κ = 0.50–1.00 across tasks.
Ethics statement
Because all data were retrieved from publicly available de-identified case reports, institutional review board approval and patient consent were not required. We extracted key demographics (age, sex, race/ethnicity) and clinical history directly from published text, logging them in a structured spreadsheet to permit subgroup analyses and to assess representativeness.
Supplementary information
Acknowledgements
This study is supported by the NSF 2444759 and SUNY Empire Innovation Program (to N.L.G). We thank the developers and maintainers of the open-source tools that made this research possible, including the Python scientific computing ecosystem (pandas, numpy, scipy), BeautifulSoup for web scraping, and the Natural Language Toolkit (NLTK) for text processing. We acknowledge OpenAI, Google, and Meta for providing API access to GPT-4-turbo, Gemini-Pro, and Llama-3-8B, respectively.We are grateful to the authors of the case reports analyzed in this study for sharing their clinical experiences through publication, which enables research such as ours to advance cancer care.We thank Grady King at West Virginia University for developing the scraper to extract publicly available information.
Author contributions
V.M.: Conceptualization, methodology development, software implementation, data curation, formal analysis, validation, visualization, project administration, writing – original draft, writing – review and editing. L.U.: Methodology development, data analysis, validation, literature review, writing – review and editing. Z.D.: Writing – original draft, methodology review, writing – review and editing. Z.X.: Writing – original draft, methodology validation, writing – review and editing. S.M.: Data analysis, validation, literature review. N.L.G: Supervision, conceptualization, methodology oversight, resources, project leadership, writing – original draft, writing – review and editing, funding acquisition. All authors discussed the results, contributed to the interpretation of findings, and approved the final manuscript.
Data availability
All data supporting the findings of this study are provided in the Data directory within the TheraMind folder: TheraMind Data Repository (https://github.com/nancylanguo1/CATOS-AI/tree/main/TheraMind). File descriptions and usage notes are included in the repository.
Code availability
The code used in this study is available in the TheraMind folder: TheraMind GitHub Repository ((https://github.com/nancylanguo1/CATOS-AI/tree/main/TheraMind)). The repository includes scripts and documentation to reproduce the analyses.
Competing interests
N.L.G is the founder and CEO of Sostos Inc.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
The online version contains supplementary material available at 10.1038/s41698-025-01265-1.
References
- 1.Gou, Q., Gou, Q., Gan, X. & Xie, Y. Novel therapeutic strategies for rare mutations in non-small cell lung cancer. Sci. Rep.14, 10317 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Siegel, R. L., Miller, K. D., Wagle, N. S. & Jemal, A. Cancer statistics, 2023. CA Cancer J. Clin.73, 17–48 (2023). [DOI] [PubMed] [Google Scholar]
- 3.Doumat, G. et al. Drug repurposing in non-small cell lung carcinoma: old solutions for new problems. Curr. Oncol.30, 704–719 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Mamdani, H., Matosevic, S., Khalid, A. B., Durm, G. & Jalal, S. I. Immunotherapy in lung cancer: current landscape and future directions. Front. Immunol.13, 823618 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Mohi-Ud-Din, R. et al. Repurposing approved non-oncology drugs for cancer therapy: a comprehensive review of mechanisms, efficacy, and clinical prospects. Eur. J. Med. Res.28, 345 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ye, Q. et al. A multi-omics network of a seven-gene prognostic signature for non-small cell lung cancer. Int. J. Mol. Sci.23, 219 (2021). [DOI] [PMC free article] [PubMed]
- 7.Ye, Q. & Guo, N. L. Single B cell gene co-expression networks implicated in prognosis, proliferation, and therapeutic responses in non-small cell lung cancer bulk tumors. Cancers14, 3123 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Ye, Q. et al. MicroRNA, mRNA, and proteomics biomarkers and therapeutic targets for improving lung cancer treatment outcomes. Cancers15, 2294 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ye, Q. et al. MicroRNA-based discovery of biomarkers, therapeutic targets, and repositioning drugs for breast cancer. Cells12, 1917 (2023). [DOI] [PMC free article] [PubMed]
- 10.Ye, Q. et al. Expression-based diagnosis, treatment selection, and drug development for breast cancer. Int. J. Mol. Sci.24, 10561 (2023). [DOI] [PMC free article] [PubMed]
- 11.Ye, Q. et al. Multi-omics immune interaction networks in lung cancer tumorigenesis, proliferation, and survival. Int. J. Mol. Sci.23, 14978 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Ye, Q., Singh, S., Qian, P. R. & Guo, N. L. Immune-omics networks of CD27, PD1, and PDL1 in non-small cell lung cancer. Cancers13, 4296 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.MotieGhader, H. et al. Drug repositioning in non-small cell lung cancer (NSCLC) using gene co-expression and drug-gene interaction networks analysis. Sci. Rep.12, 9417 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Jain, A. S. et al. Everything old is new again: drug repurposing approach for non-small cell lung cancer targeting MAPK signaling pathway. Front. Oncol.11, 741326 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Rajasegaran, T., How, C. W., Saud, A., Ali, A. & Lim, J. C. W. Targeting inflammation in non-small cell lung cancer through drug repurposing. Pharmaceuticals16, 451 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.von Itzstein, M. S. et al. Phase I/II trial of exportin 1 inhibitor selinexor plus docetaxel in previously treated, advanced KRAS-mutant non–small cell lung cancer. Clin. Cancer Res.31, 639–648 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Percha, B. Modern clinical text mining: a guide and review. Annu. Rev. Biomed. Data Sci.4, 165–187 (2021). [DOI] [PubMed] [Google Scholar]
- 18.Bazoge, A., Morin, E., Daille, B. & Gourraud, P.-A. Applying natural language processing to textual data from clinical data warehouses: systematic review. JMIR Med. Inf.11, e42477 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zong, H. et al. Advancing Chinese biomedical text mining with community challenges. J. Biomed. Inf.157, 104716 (2024). [DOI] [PubMed] [Google Scholar]
- 20.Kreimeyer, K. et al. Natural language processing systems for capturing and standardizing unstructured clinical information: a systematic review. J. Biomed. Inf.73, 14–29 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Sim, J.-A. et al. Natural language processing with machine learning methods to analyze unstructured patient-reported outcomes derived from electronic health records: a systematic review. Artif. Intell. Med.146, 102701 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Eguia, H., Sánchez-Bocanegra, C. L., Vinciarelli, F., Alvarez-Lopez, F. & Saigí-Rubió, F. Clinical decision support and natural language processing in medicine: systematic literature review. J. Med. Internet Res.26, e55315 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Doneva, S. E. et al. Large language models to process, analyze, and synthesize biomedical texts: a scoping review. Discov. Artif. Intell.4, 107 (2024). [Google Scholar]
- 25.Liu, F. et al. Application of large language models in medicine. Nat. Rev. Bioeng. 3, 445–464 (2025).
- 26.Park, Y.-J. et al. Assessing the research landscape and clinical utility of large language models: a scoping review. BMC Med. Inf. Decis. Mak.24, 72 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Meng, X. et al. The application of large language models in medicine: a scoping review. iScience27, 109713 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Shool, S. et al. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med. Inf. Decis. Mak.25, 117 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med.31, 943–950 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Alsentzer, E. et al. Publicly available clinical BERTembeddings. In Proceedings of the 2nd clinical natural languageprocessing workshop, 72–78 (2019).
- 31.Wang, B. et al. Pre-trained language models in biomedical domain: a systematic survey. ACM Comput. Surv.56, 55:51–55:52 (2023). [Google Scholar]
- 32.Cho, H. N. et al. Task-specific transformer-based language models in health care: scoping review. JMIR Med Inf.12, e49724 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Wang, L., Ma, Y., Bi, W., Lv, H. & Li, Y. An entity extraction pipeline for medical text records using large language models: analytical study. J. Med Internet Res26, e54580 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med.30, 2613–2622 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Liu, F. et al. A medical multimodal large language model for future pandemics. npj Digit. Med.6, 1–15 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.AlSaad, R. et al. Multimodal large language models in health care: applications, challenges, and future outlook. J. Med. Internet Res.26, e59505 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Cascella, M. et al. The breakthrough of large language models release for medical applications: 1-year timeline and perspectives. J. Med. Syst.48, 22 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Mahajan, P., Uddin, S., Hajati, F. & Moni, M. A. Ensemble learning for disease prediction: a review. Healthcare11, 1808 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Vincent, A. C. S. R. & Sengan, S. Edge computing-based ensemble learning model for health care decision systems. Sci. Rep.14, 26997 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Naderalvojoud, B. & Hernandez-Boussard, T. Improving machine learning with ensemble learning on observational healthcare data. AMIA Annu. Symp. Proc.2023, 521–529 (2024). [PMC free article] [PubMed] [Google Scholar]
- 41.Acosta, J. N., Falcone, G. J., Rajpurkar, P. & Topol, E. J. Multimodal biomedical AI. Nat. Med.28, 1773–1784 (2022). [DOI] [PubMed] [Google Scholar]
- 42.Lu, Z. et al. Large language models in biomedicine and health: current research landscape and future directions. J. Am. Med. Inf. Assoc.31, 1801–1811 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Jin, Q., Leaman, R. & Lu, Z. PubMed and beyond: biomedical literature search in the age of artificial intelligence. EBioMedicine100, 104988 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Tam, T. Y. C. et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digit. Med.7, 1–20 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Team, G. et al. Gemini: a family of highly capable multimodal models. Preprint at https://arxiv.org/abs/2312.11805 (2025).
- 46.Grattafiori, A. et al. The Llama 3 herd of models. Preprint at https://arxiv.org/abs/2407.21783 (2024).
- 47.Ishiguro, T., Ishiguro, R. H., Ishiguro, M., Toki, A. & Terunuma, H. Synergistic anti-tumor effect of dichloroacetate and ivermectin. Cureus14, e21884 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Kim, E. S. Chemotherapy resistance in lung cancer. In Lung Cancer and Personalized Medicine: Current Knowledge and Therapies (eds. Ahmad, A. & Gadgeel, S.) 189–209 (Springer International Publishing, 2016).
- 49.Camidge, D. R., Pao, W. & Sequist, L. V. Acquired resistance to TKIs in solid tumours: learning from lung cancer. Nat. Rev. Clin. Oncol.11, 473–481 (2014). [DOI] [PubMed] [Google Scholar]
- 50.Shaw, A. T. et al. Crizotinib versus chemotherapy in advanced ALK-positive lung cancer. N. Engl. J. Med.368, 2385–2394 (2013). [DOI] [PubMed] [Google Scholar]
- 51.Piotrowska, Z. et al. Heterogeneity underlies the emergence of EGFRT790 wild-type clones following treatment of T790M-positive cancers with a third-generation EGFR inhibitor. Cancer Discov.5, 713–722 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Oxnard, G. R. et al. Association between plasma genotyping and outcomes of treatment with osimertinib (AZD9291) in advanced non-small-cell lung cancer. J. Clin. Oncol.34, 3375–3382 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Oxnard, G. R., Binder, A. & Jänne, P. A. New targetable oncogenes in non-small-cell lung cancer. J. Clin. Oncol.31, 1097–1104 (2013). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Cancer Stat Facts: Lung and Bronchus Cancer. Avaialbel at: https://seer.cancer.gov/statfacts/html/lungb.html (Accessed on January 19, 2026).
- 55.Gettinger, S. et al. Five-year follow-up of nivolumab in previously treated advanced non-small-cell lung cancer: results from the CA209-003 Study. J. Clin. Oncol.36, 1675–1684 (2018). [DOI] [PubMed] [Google Scholar]
- 56.Garon, E. B. et al. Pembrolizumab for the treatment of non-small-cell lung cancer. N. Engl. J. Med.372, 2018–2028 (2015). [DOI] [PubMed] [Google Scholar]
- 57.Hellmann, M. D. et al. Nivolumab plus ipilimumab as first-line treatment for advanced non-small-cell lung cancer (CheckMate 012): results of an open-label, phase 1, multicohort study. Lancet Oncol.18, 31–41 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Li, J. et al. A survey of current trends in computational drug repositioning. Brief. Bioinform17, 2–12 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Pushpakom, S. et al. Drug repurposing: progress, challenges and recommendations. Nat. Rev. Drug Discov.18, 41–58 (2019). [DOI] [PubMed] [Google Scholar]
- 60.Tang, L. et al. Evaluating large language models on medical evidence summarization. npj Digit. Med.6, 1–8 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Wei, Q. et al. Evaluation of ChatGPT-generated medical responses: a systematic review and meta-analysis. J. Biomed. Inf.151, 104620 (2024). [DOI] [PubMed] [Google Scholar]
- 62.Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med.30, 1134–1142 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Mehandru, N. et al. Evaluating large language models as agents in the clinic. npj Digit. Med.7, 1–3 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Dietterich, T. G. Ensemble Methods in Machine Learning, 1–15 (Springer, 2000).
- 65.Zhou, Z.-H. Ensemble Methods: Foundations and Algorithms (Chapman & Hall/CRC, 2012).
- 66.Hossen, M. J., Ramanathan, T. T. & Al Mamun, A. An ensemble feature selection approach-based machine learning classifiers for prediction of COVID-19 disease. Int. J. Telemed. Appl.2024, 8188904 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Gonzalez, G. H., Tahsin, T., Goodale, B. C., Greene, A. C. & Greene, C. S. Recent advances and emerging applications in text and data mining for biomedical discovery. Brief. Bioinform17, 33–42 (2016). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Zhao, S., Su, C., Lu, Z. & Wang, F. Recent advances in biomedical literature mining. Brief. Bioinform22, bbaa057 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Scherbakov, D., Hubig, N., Jansari, V., Bakumenko, A. & Lenert, L. A. The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review. J. Am. Med. Inform. Assoc.32, 1071–1086 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Chen, T. & Guestrin, C. XGBoost: A Scalable Tree Boosting System, 785–794 (Association for Computing Machinery, 2016).
- 71.Scott, I. A., Cook, D., Coiera, E. W. & Richards, B. Machine learning in clinical practice: prospects and pitfalls. Med. J. Aust.211, 203–205.e201 (2019). [DOI] [PubMed] [Google Scholar]
- 72.Qiao, H., Chen, Y., Qian, C. & Guo, Y. Clinical data mining: challenges, opportunities, and recommendations for translational applications. J. Transl. Med.22, 185 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Holzinger, A., Biemann, C., Pattichis, C. S. & Kell, D. B. What do we need to build explainable AI systems for the medical domain? Preprint at https://arxiv.org/abs/1712.09923 (2017).
- 74.Guidotti, R. et al. A survey of methods for explaining black box models. ACM Comput. Surv.51, 93:91–93:42 (2018). [Google Scholar]
- 75.Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell.1, 206–215 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Wu, W.-T. et al. Data mining in clinical big data: the frequently used databases, steps, and methodological models. Mil. Med. Res.8, 44 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 77.Joseph, N. et al. Automated data extraction of electronic medical records: Validity of data mining to construct research databases for eligibility in gastroenterological clinical trials. Ups J. Med. Sci.127, e8260 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Lee, C., Britto, S. & Diwan, K. Evaluating the impact of artificial intelligence (AI) on clinical documentation efficiency and accuracy across clinical settings: a scoping review. Cureus16, e73994 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Lucas, H. C., Upperman, J. S. & Robinson, J. R. A systematic review of large language models and their implications in medical education. Med Educ.58, 1276–1285 (2024). [DOI] [PubMed] [Google Scholar]
- 80.Gong, E. J. et al. Large language models in gastroenterology: systematic review. J. Med. Internet Res.26, e66648 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Elhaddad, M. & Hamam, S. AI-driven clinical decision support systems: an ongoing pursuit of potential. Cureus16, e57728 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation) OJ 2016 L 119/1 (2016).
- 83.Health Insurance Portability and Accountability Act of 1996 (1996).
- 84.Carlini, N. et al. Extracting Training Data from Diffusion Models. 5253–5270 (USENIX Association, 2023).
- 85.Priyanshu, A., Vijay, S., Kumar, A., Naidu, R. & Mireshghallah, F. Are chatbots ready for privacy-sensitive applications? An investigation into input regurgitation and prompt-induced sanitization. Preprint at https://arxiv.org/abs/2305.15008 (2023).
- 86.Zhang, Y., Carlini, N. & Ippolito, D. Effective prompt extraction from language models. Preprint at https://arxiv.org/abs/2307.06865 (2024).
- 87.Sweeney, L. K-anonymity: a model for protecting privacy. Int. J. Uncertain. Fuzziness Knowl. Based Syst.10, 557–570 (2002). [Google Scholar]
- 88.Narayanan, A. & Shmatikov, V. Robust de-anonymization of large sparse datasets. In 2008 IEEE Symposium on Security and Privacy (SP 2008) 111–125 (IEEE, 2008).
- 89.Dwork, C., McSherry, F., Nissim, K. & Smith, A. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography (eds. Halevi, S. & Rabin, T.) 265–284 (Springer, 2006).
- 90.Tang, X. et al. Privacy-preserving in-context learning with differentially private few-shot generation. Preprint at https://arxiv.org/abs/2309.11765 (2024).
- 91.Wu, T., Panda, A., Wang, J. T. & Mittal, P. Privacy-preserving in-context learning for large language models. Preprint at https://arxiv.org/abs/2305.01639 (2023).
- 92.Agarwal, D. et al. Prompt Leakage effect and defense strategies for multi-turn LLM interactions. Preprint at https://arxiv.org/abs/2404.16251 (2024).
- 93.Zhang, Z. et al. Defending large language models against jailbreaking attacks through goal prioritization. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics, Vol. 1: Long Papers, 8865–8887 (2024).
- 94.Zhang, Y., Ding, L., Zhang, L. & Tao, D. Intention analysis makes LLMs a good jailbreak defender. In Proc. 31st International Conference on Computational Linguistics, 2947–2968 (Association for Computational Linguistics, 2025).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
All data supporting the findings of this study are provided in the Data directory within the TheraMind folder: TheraMind Data Repository (https://github.com/nancylanguo1/CATOS-AI/tree/main/TheraMind). File descriptions and usage notes are included in the repository.
The code used in this study is available in the TheraMind folder: TheraMind GitHub Repository ((https://github.com/nancylanguo1/CATOS-AI/tree/main/TheraMind)). The repository includes scripts and documentation to reproduce the analyses.









