Summary:
CAR T cells have demonstrated curative potential in hematologic cancers and increasing efficacy in solid tumors and non-malignant diseases. However, target identification remains a major bottleneck. We developed an AI-driven approach for CAR T cell target discovery by integrating single-cell RNA sequencing datasets from human skin cancer and healthy tissue. Candidates were refined using public datasets to optimize for tumor composition, tissue specificity, and clinical feasibility. Large language models were applied to prioritize and nominate targets with therapeutic promise. Glycoprotein non-metastatic melanoma protein B (GPNMB) was the most frequently nominated target. We validated its expression across hematologic and solid tumors. We engineered a human GPNMB-directed CAR T cell, which showed potent antitumor activity in mouse models of monoblastic leukemia, melanoma, and colorectal adenocarcinoma. These findings establish a scalable pipeline for CAR T cell target discovery and support the translation of GPNMB-directed CAR T cells as a multi-cancer therapeutic.
Graphical Abstract

In brief:
An AI-guided framework integrating high-resolution cellular atlases, public knowledge databases, and large language models enables a modular framework for CAR T cell target discovery. GPNMB emerged as a versatile candidate antigen, enabling the development of GPNMB CAR T cells with potent activity in multiple xenograft mouse models of cancer.
Introduction
Chimeric antigen receptor (CAR) T cell therapy has ushered in a new paradigm in medicine, with curative outcomes in a subset of patients with hematologic cancers1,2. Its expansion into solid tumors and non-malignant diseases is underway, but remains substantially constrained by the scarcity of safe target antigens3. The highly personalized and complex nature of CAR T cell therapy further limits the reach of these therapies through the clinical pipeline. These regulatory limitations highlight the utility of CD19 CAR T cells, which have demonstrated efficacy in both leukemia and lymphoma, leading to enormous clinical impact4. It is this same product that is now demonstrating early but remarkable effects in refractory systemic autoimmune disease5,6. Thus, developing a CAR that is applicable across multiple indications would be highly desirable.
Existing antigen discovery pipelines require the generation or analysis of databases for disease-specific enrichment of the target. Then, depending on data/tissue availability, subsequent subpopulation analysis to assay if the deleterious population expresses the target antigen3,7. Typically, databases are then manually probed to search for candidates that are known or predicted to localize to the cell surface so as to be suitable for CAR T cell therapy. Following these analyses, further examination is done to determine the potential tolerability of antigen profile in normal tissues as on-target killing of vital tissues has emerged as a severe toxicity of CAR T cell therapy8. Finally, preference may be given to antigens that have been targeted either via drugs or biologics to facilitate translation of the CAR T cell product. This process has historically been time and resource intensive with variable success rates3.
A scalable platform to systematically identify safe and potent targets would accelerate the expansion of CAR T cell therapies across a wide range of cancers and other diseases. Recent advances in artificial intelligence (AI), particularly in the form of large language models (LLMs), have opened new avenues for biomedical research9,10. These models can rapidly synthesize, reason across, and prioritize vast amounts of complex biological data11. While there is growing interest in applying these models across biomedicine, their application to CAR T cell therapy has been limited. We hypothesized that integrating LLM-based reasoning with high-resolution cellular atlases could create a scalable, systematic framework for CAR T cell target nomination.
We developed an AI-driven framework to systematically identify optimal CAR T cell targets. We integrated multiple human skin cancer single-cell datasets with a healthy human reference dataset to identify genes compositionally enriched in the malignant compartment. Leveraging knowledge from a wide range of publicly available resources, we extracted features for each target related to subcellular localization, specificity, potency, and translational feasibility. We then utilized three LLMs to prioritize and nominate targets based on these features. Our computational pipeline identified glycoprotein non-metastatic melanoma protein B (GPNMB) as the top candidate, and we subsequently developed and validated a CAR T cell targeting GPNMB. Intriguingly this antigen is expressed not only in skin cancer but also across a wide spectrum of cancers. GPNMB CAR T cells demonstrated potent antitumor efficacy in humanized models of leukemia, melanoma, and colorectal cancer. Overall, our experimentally validated AI-driven approach offers a generalizable strategy for CAR T cell target discovery and highlights GPNMB as a versatile candidate for multi-cancer therapy.
Results:
Identifying optimal CAR T cell targets requires balancing multiple critical features, including tumor expression, absence in healthy tissues, cell surface localization, and clinical translatability. We developed a three-pronged strategy—integrating data from multiple studies, augmenting knowledge from public resources, and applying LLMs to systematically nominate targets for CAR T cell development (Fig 1A).
Figure 1. Data, knowledge, and reasoning driven AI framework identifies GPNMB as a CAR T cell target in skin cancer.

(A) Schematic of CAR T target discovery pipeline: data integration, knowledge augmentation, LLM-based prioritization and nomination, followed by experimental validation.
(B) UMAP of integrated single-cell RNA-seq data from skin cancer samples, healthy skin from Tabula Sapiens, and the final malignant-plus-healthy atlas.
(C) Overview of the four feature categories and their subcomponents used for target evaluation.
(D) Feature weights for “cancer composition” feature across 1000 simulations from the three LLMs used (OpenAI GPT-4o, Anthropic Claude-3.7, and Google Gemini-2.5-Pro).
(E) Radar plot showing the average weight score given to each feature by the three independent LLMs.
(F) Stacked bar plot of the top 15 genes nominated cumulatively across all models.
(G) Alluvial plots showing relative model nomination of each target (middle) and the reason that a target was nominated (right).
(H) UMAP of GPNMB expression across the integrated scRNA-seq atlas and dotplot with the quantified expression of targets in malignant and healthy cells in the integrated atlas.
(I) The Cancer Genome Atlas (TCGA) bulk RNAseq analysis of skin cutaneous melanoma samples of nominated targets. Tumor samples shown in red and normal skin tissue from GTEx shown in blue. Statistical significance was calculated using one-way ANOVA, *P<0.01
(J) Immunofluorescence of melanoma cells (SK-MEL-5) shows intracellular and surface expression of respective candidates (shown in red) with DAPI counterstain. IgG control shown on left. Scale bar is 20 μm.
(K) Non-permeabilized flow cytometry of melanoma cells (SK-MEL-5) shows surface expression of several candidates nominated by this computational framework. Grey histograms show cells stained with isotype/secondary antibodies only; whereas red shows anti-target + secondary antibody.
Constructing a single-cell skin cancer and healthy tissue atlas to identify CAR T cell targets
Given the severity of refractory skin cancer and the clear utility of immunotherapies in this disease, we focused on identifying CAR T cell targets for skin cancer12–15. Known complications of CAR T therapy include tumor antigen heterogeneity and on target-off tumor toxicity, so we assessed antigen heterogeneity at single cell resolution. We constructed an atlas by integrating four publicly available scRNAseq datasets from skin cancer patients, capturing disease stages and subtypes, including basal, squamous, and varying stages of melanoma including metastases (Fig. 1B)16–19. Treatment histories varied; some patients were untreated, while others had received immunotherapy or targeted therapy prior to sampling. Following batch effect correction using Harmony, we integrated these datasets into a common standardized atlas (Supp Fig. 1A)20. The malignant populations from the cancer datasets were extracted and integrated with a healthy human skin dataset from Tabula Sapiens21. This atlas was used to identify genes compositionally prevalent in the malignant compartment and served as the foundation for downstream target prioritization
Augmenting potential CAR T cell targets with safety, tractability, and feasibility features to prioritize discovery
Following dataset integration, we extracted additional features from publicly available resources for each potential target in our atlas. This included databases such as UniProt, The Human Protein Atlas, GTEx, Open Targets, clinicaltrials.gov, and ProteomicsDB (Fig. 1A)22–25. These features broadly encompassed four critical categories for CAR T cell development: subcellular localization, tumor specificity, cancer composition, and clinical translatability (Fig 1C). Subcellular location was of significance since CAR T cells rely on surface expressed targets. Specificity of the target to the disease tissue is critical due to the potency of CAR T cells and their ability to target and eliminate healthy tissue leading to severe toxicity. We subclustered this feature into four categories based on tissue type and data modality: vital tissue-RNA, vital tissue-protein, non-vital tissue-RNA, and non-vital tissue-protein (Fig. 1C). We also generated a cancer composition score to assess therapeutic potential. We quantified the proportion of malignant cells expressing each target across the tumor population. Since intratumoral heterogeneity can drive resistance and relapse in cancers like skin cancer, this approach prioritized targets broadly expressed across malignant cells, seeking to mitigate clonal escape and increase durability. Finally, to assess translatability, we weighed the amount of clinical development surrounding a target. While this feature is not required for a CAR T cell target, we reasoned that a target that has been clinically explored has stronger rationale for clinical development. For each of these features we developed an expert-guided quantitative score based on the relevant information available (Supp Fig 1C). When we prioritize our candidate antigens for any single feature (i.e. translatability, surface localization, etc.), this deprioritizes other critical features for the optimal CAR T cell target (Supp Fig. 1D). To remedy this, we deployed a human-in-the-loop approach to generate a weighting system for each feature using LLMs. We leveraged our group’s clinical and developmental experience with CAR T cells and generated a score from first principles for each feature. We gave particularly strong weight to cancer composition, cell surface localization, and limited target expression in vital tissues. Using these base weights we then queried three independent LLMs: ChatGPT-4o, Claude-3.7, and Gemini-2.5-Pro, across 1000 independent simulations to optimize the weight for each target feature (Fig 1D). We observed consensus on the relative importance of key features centered around cancer composition, subcellular localization, and vital tissue expression (Supp Fig 1E–G). Despite the differences between these models, they largely converged upon the relative importance of each feature weight (Supp Fig 1H). From this we established a robust composite scoring system to prioritize the multifactorial features of the ideal CAR T cell target (Fig. 1E).
Large language models reason-driven nomination of CAR T cell targets
Using this optimized weighting scheme, we prioritized the top 100 candidates for further evaluation. Leveraging the LLMs, we prompted each model 1000 times to nominate the most promising skin cancer CAR T cell targets for development (Supp Fig 2A). Each model was tasked with selecting 10 genes from this 100 gene shortlist for CAR T cell development and to provide explicit reasoning for nomination based on five distinct possibilities: tumor specificity, multi-cancer utility, clinical development potential, low vital tissue expression, and relevance to patient subpopulations (Fig. 1G). Among the three LLMs used we observed differences in the targets nominated and the reasons provided (Supp Fig 2B–C). However, strong consensus did emerge among the models around the top hits for CAR T cell development (Fig 1F–G, Supp Fig. 2D). The reason associated with a target’s nomination diverged significantly as expected (Supp Fig 2E). For example, PMEL—a known melanoma enriched target—was always nominated for the reason of tumor specificity while rarely nominated for its multi-cancer potential. IGF1R in contrast was rarely nominated for tumor specificity but ranked highly for its multi-cancer potential. GPNMB, the most frequently nominated target, was consistently selected for tumor specificity but also showed a strong signal for multi-cancer utility. Our nominated targets consistently showed strong enrichment in the malignant population of our scRNAseq atlas (Fig. 1H).
To further demonstrate the stability of this framework we re-ran the pipeline an additional two iterations totaling 3,000 independent simulations per model (Supp Fig 3N). This systematic human-in-the-loop pipeline led to the identification and prioritization of several targets for CAR T cell therapy in skin cancer.
Robustness of target nomination with and without domain expertise
To assess the influence of human domain expertise on target prioritization, we removed our expert-defined base weights and re-prompted each LLM to assign feature weights across 1,000 independent simulations. While modest shifts in feature rankings were observed (Supp Fig 3A), the target lists generated by each approach were highly concordant. Specifically, 94 of the top 100 genes overlapped between expert and non-expert weighting strategies (Supp Fig. 3B–C), and 12 of the top 15 final nominees were shared (Supp Fig. 3D–F). These findings underscore the robustness of our framework and suggest this framework is accessible even to groups without deep domain expertise.
Impact of input data resolution on target nomination
Though single-cell data is becoming increasingly abundant and offers crucial cellular resolution for CAR T cell targeting, bulk RNAseq is more common in historic discovery efforts. We interrogated if bulk RNAseq data could be leveraged to yield similar results to single-cell resolution data. To assess the consequences of this difference, we re-ran our full discovery pipeline on the bulk RNAseq the Cancer Genome Atlas (TCGA) melanoma cohort. In this analysis we replaced the single-cell cancer composition score with a tumor expression z-score for each target.
This substitution led to substantial divergence in prioritization outcomes. The LLM-derived feature weightings shifted due to the absence of cell-type–specific prevalence information (Supp Fig. 3G). Only 21 of the top 100 targets were shared between the scRNAseq and bulk RNAseq-based shortlists (Supp Fig. 3H–I). We then performed downstream model-based nomination using this 100 gene shortlist over 1000 simulations. Only 2 of the top 15 targets shared across these approaches (Supp Fig. 3J–L). One reason for this discrepancy is the reduced transcript coverage of single-cell datasets, as only 48.2% of transcripts present in bulk RNA-seq were represented in the integrated scRNAseq atlas. Another reason for this difference is the lack of cell-type discrimination in bulk data. For instance, TIGIT, CTLA4, and CD83 all are nominated in the top 15 from the bulk analysis (Supp Fig. 3J). However, when examined at single cell resolution TIGIT and CTLA4 are expressed in the T cell compartment and CD83 is expressed in myeloid cells (Supplementary Fig. 3M), making them unsuitable for tumor-directed CAR targeting despite their enrichment in bulk tumor tissue, presumably due to the presence of tumor infiltrating lymphocytes.
These results suggest that bulk RNAseq is not interchangeable with scRNAseq for CAR T cell target discovery, as the loss of cellular resolution significantly alters the prioritization process. As single-cell dataset quality and coverage continue to expand, we expect this modality to become the preferred data foundation for CAR T cell target discovery.
Validating the real-world relevance of AI-nominated targets in human skin cancer
To assess the biological relevance of our AI-driven CAR T cell target prioritization framework, we performed orthogonal validation approaches incorporating bulk transcriptomic data, immunofluorescence microscopy, and surface-level protein confirmation via flow cytometry. Analysis of TCGA melanoma dataset showed that several of our top targets (GPNMB, ERBB3, IGF1R) were transcriptionally upregulated in melanoma samples compared to normal tissue (Fig. 1I; Supp. Fig. 2G).
To validate the presence of the target protein in melanoma we performed immunofluorescence of top targets in four separate melanoma cell lines. We observed strong target expression at the plasma membrane and intracellularly in all of these cells (Fig. 1J, Supp Fig 4A). Given that CAR T cell therapy requires surface expression for effective targeting, we validated the surface presentation of these antigens using non-permeabilized flow cytometry. Analysis of all four melanoma cell lines confirmed surface expression of multiple of our hits (GPNMB, ERBB3, IGF1R) at the surface of melanoma cells (Fig 1K, Supp Fig. 4B).
Collectively, these findings validate our target discovery pipeline, identifying several promising targets in human melanoma including ERBB3 and IGF1R. Moreover, these results provided us with compelling rationale to explore the preclinical development of GPNMB-directed CAR T cells.
GPNMB is a tractable target across a variety of cancers
Glycoprotein non-metastatic melanoma protein B (GPNMB) is a type I transmembrane glycoprotein implicated in cell adhesion, migration, tissue repair, and immunosuppression26,27. We evaluated the expression of GPNMB across multiple cancer types. Bulk RNA sequencing data from TCGA revealed elevated GPNMB expression not only in melanoma, but also across a wide range of malignancies (Fig. 2A).
Figure 2. GPNMB is a broadly expressed tumor antigen and enables potent CAR T cell targeting of hematologic malignancies.

(A) GPNMB expression across cancer types based on TCGA bulk RNA-seq data compared to normal tissue from GTEx. Tumor samples shown in red and normal tissue from GTEx shown in blue. Only cancers with significant enrichment are shown. Abbreviations: CHOL-Cholangiocarcinoma; DLBC-Lymphoid Neoplasm Diffuse Large B-cell Lymphoma; GBM-Glioblastoma multiforme; HNSC-Head and Neck squamous cell carcinoma; KICH-Kidney Chromophobe; KIRP-Kidney renal papillary cell carcinoma; LIHC-Liver hepatocellular carcinoma; LUSC-Lung squamous cell carcinoma; PAAD-Pancreatic adenocarcinoma; SKCM-Skin Cutaneous Melanoma; STAD-Stomach adenocarcinoma; TGCT-Testicular Germ Cell Tumors; THYM-Thymoma. Statistical significance was calculated using one-way ANOVA, *P<.01
(B) Representative IHC images from tumor or control biopsies stained for GPNMB protein. GPNMB staining is shown in brown at 40x magnification.
(C) Flow cytometry confirming GPNMB surface expression on tumor cell lines under non-permeabilized conditions. Grey histograms show cells stained with secondary antibody only; whereas red shows anti-GPNMB + secondary antibody.
(D) Schematic of second generation GPNMB CAR construct containing a glembatumumab vedotin–derived scFv, 4–1BB costimulatory, and CD3ζ signaling domains.
(E) Non-permeabilized flow cytometry showing GPNMB CAR surface expression on primary human T cells post-transduction in comparison to control T cells from the same human donor.
(F) Non-permeabilized flow cytometry showing surface expression of GPNMB in monocytic leukemia (ML) and acute myeloid leukemia (AML), but absent in B-cell acute lymphoblastic leukemia (B-ALL).
(G) IFN-γ secretion by CAR T cells following co-culture with GPNMB+ tumor cells but not GPNMB− targets. Statistical significance was calculated using unpaired two-tailed t-tests with Welch’s correction. **P< 0.01.
(H) GPNMB CAR T cells proliferate in response to GPNMB+ tumor cells but not GPNMB− targets. Proliferation measured by CTV dilution.
(I) GPNMB CAR T cells exhibit antigen-dependent killing of tumor cells. Cytotoxicity was assessed after 48–68hr co-culture via flow cytometry. Statistical significance was calculated using unpaired two-tailed t-tests. ****P<0.0001
(J) In vivo schematic of humanized mouse model of monoblastic leukemia. At D-2 1×10e5 tumor cells were injected I.V. At D0 5×10e6 GPNMB CAR T cells were infused. Pooled data from 2 T cell donors. n = 7–8 mice/group/donor.
(K) CD45+ cell count (=infused T cells) analyzed via flow cytometry from peripheral blood at indicated time point (n=14–15/cohort, pooled from two donors). Average and individual values are shown. Mann-Whitney U Test was used for statistical analyses. ****P<0.0001.
(L) Kaplan-Meier survival curve (n=15 CAR T; n=14 Control T, pooled from two donors). Statistical significance was calculated using the log-rank Mantel-Cox test. ****P<0.0001.
All data representative of at least 2 independent experiments or biological replicates.
To validate GPNMB’s presence at the protein level, we performed immunohistochemistry (IHC) on patient tumor biopsies. Consistent with transcriptomic data, GPNMB protein was expressed in tumor tissues in melanoma and across cancers (Fig. 2B). Consistently the majority of melanoma samples showed strong staining for GPNMB (Supp Fig 5A). Whereas other cancer types demonstrated a range of tumor positive staining (Supp Fig 5A). Evaluation of GPNMB expression revealed low or absent expression across a range of healthy tissues (Supp Fig. 5A). Importantly for a CAR T-centric approach, surface expression of GPNMB was confirmed across a wide range of cancers (Fig. 2C).
Developing a GPNMB targeted chimeric antigen receptor
To target GPNMB-expressing tumors, we engineered a second-generation CAR using the antigen-binding domains of glembatumumab vedotin, an antibody-drug conjugate that underwent significant clinical development28–34. Glembatumumab vedotin demonstrated that targeting surface GPNMB was feasible, though its development was discontinued when the Phase 2b trial failed to meet its primary endpoint. CAR T cells can show exceptional potency in a variety of indications where patients did not respond to antibody-based therapies3,5,35,36. This enhanced potency supported the rationale of creating a GPNMB CAR T cell.
We incorporated the antibody’s light and heavy chains into a 2nd generation CAR construct featuring a CD8α hinge and transmembrane domain, a 4–1BB costimulatory domain, and a CD3ζ signaling domain (Fig. 2D). Primary human T cells transduced with a lentiviral vector encoding this CAR successfully expressed surface CAR (Fig 2E). They also demonstrated robust activation and expansion across multiple donors (Supp. Fig 5B).
GPNMB CAR T cells show potent anti-tumor efficacy in hematologic malignancies
Given the striking clinical successes of CAR T cells in blood cancers, we first GPNMB CAR T cells to recognize and eliminate GPNMB+ blood malignancies. B cell acute lymphoblastic leukemia (B-ALL) cells showed no surface expression of GPNMB whereas cells derived from monoblastic leukemia (ML), or acute myeloid leukemia (AML) showed variable surface expression of GPNMB (Fig. 2F). To test the specificity and potency of GPNMB CAR T cells we performed a triad of tests to examine activation, effector function, and killing. GPNMB CAR T cells showed antigen specific activation–only proliferating when co-cultured with GPNMB+ malignancies, and correspondingly these CAR T cells only showed significantly increased cytokine production when co-cultured with GPNMB+ tumor cells (Fig 2G–H). Most importantly, GPNMB CAR T cells showed strong anti-tumor potency against ML and AML while sparing antigen negative malignancies (Fig. 2I). Collectively, these findings demonstrate the development of a functional human GPNMB-targeting CAR exhibiting specificity and potency.
To assess the in vivo efficacy of GPNMB CAR T cells in hematologic malignancies, we used a xenograft model of human ML. Immunodeficient mice were intravenously injected with tumor cells. After two days mice were randomized into treatment groups and either control T cells or GPNMB CAR T cells were administered (Fig 2J). To evaluate the antigen dependent expansion of CAR T cells in vivo, we assessed the peripheral blood of these mice 13 days following T cell infusion. Flow cytometry revealed a significant expansion of human CD45+ cells in the peripheral blood of mice treated with GPNMB CAR T cells consistent with target engagement (Fig. 2K). Peripheral blood was taken to monitor tumor growth; however, no tumor burden was seen in the peripheral blood, likely due to tumor engraftment and propagation in the bone marrow37,38. However, mice receiving control T cells exhibited rapid disease progression, reaching humane endpoints within weeks post-engraftment. In sharp contrast groups treated with GPNMB CAR T cells survived for months with no signs of disease (Fig 2L). These mice showed no overt signs of toxicity with stable body weights throughout the experiment (Supp Fig. 5E). Together these findings establish the robust in vivo efficacy of GPNMB CAR T cells against ML highlighting their therapeutic potential against GPNMB-expressing hematologic malignancies.
GPNMB CAR T cells show potent anti-tumor efficacy against melanoma
We then tested the utility of GPNMB CAR T cells in solid tumors. Our AI-driven framework provided strong rationale to test GPNMB CAR T cells against melanoma. We evaluated the ability of GPNMB CAR T cells to target and eliminate melanoma cells in vitro. Again, GPNMB CAR T cells showed strong cytotoxic activity against melanoma cells in a dose dependent fashion (Fig. 3A).
Figure 3. GPNMB CAR T cells show robust anti-tumor response in solid tumors.

(A) Schematic of experimental design and in vitro cytotoxicity of GPNMB CAR T cells against the two melanoma cell lines (SK-MEL-3 and SK-MEL-5) at respective effector to target ratios (E:T). Target cell lysis assessed using xCELLigence Real-Time Cell Analyzer System. Average values with SD shown. (B)-(F) Results from in vivo SK-MEL-3 melanoma flank tumor model.
(B) Schematic of experimental design.
(C) CD45+ cell count (=infused T cells) analyzed via flow cytometry from peripheral blood at indicated time point (n=10/cohort, pooled from two donors). Average and individual values are shown. Mann-Whitney U Test was used for statistical analyses. ****P<0.0001.
(D) Cytokine levels in mouse serum (n=10/cohort, pooled from two donors). Average and individual values are shown. Mann-Whitney U Test was used for statistical analyses. ***P<0.001; ****P<0.0001.
(E) Changes in tumor volume over time. Left: Average and individual values shown with n=16 (CAR T) and n=14 (Control T), pooled from two donors. Ordinary two-way ANOVA was used for statistical analysis. ****P<0.0001. Right: Individual values with n=8 (CAR T) and n=7 (Control T) for each donor ND607 and ND657. CR, complete remission.
(F) Kaplan-Meier survival curve (n=16 CAR T; n=14 Control T, pooled from two donors). Statistical significance was calculated using the log-rank Mantel-Cox test. ****P<0.0001.
(G) Schematic of experimental design and in vitro cytotoxicity of GPNMB CAR T cells against indicated tumor cell lines at respective effector to target ratios (E:T). Target cell lysis assessed using xCELLigence Real-Time Cell Analyzer System. Average values with SD shown.
(H)-(L) Results from in vivo SNU-503 colorectal flank tumor model.
(H) Schematic of experimental design.
(I) CD45+ cell count (=infused T cells) analyzed via flow cytometry from peripheral blood at indicated time point (n=10/cohort, pooled from two donors). Average and individual values are shown. Mann-Whitney U Test was used for statistical analyses. ****P<0.0001.
(J) Cytokine levels in mouse serum (n=10/cohort, pooled from two donors). Average and individual values are shown. Mann-Whitney U Test was used for statistical analyses. ns (non-significant) P>0.05; *P<0.05; ****P<0.0001.
(K) Changes in tumor volume over time. Left: Average and individual values shown with n=16 (CAR T) and n=14 (Control T), pooled from two donors. Ordinary two-way ANOVA was used for statistical analysis. ****P<0.0001. Right: Individual values with n=8 (CAR T) and n=7 (Control T) for each donor ND607 and ND657. CR, complete remission.
(L) Kaplan-Meier survival curve (n=16 CAR T; n=14 Control T, pooled from two donors). Statistical significance was calculated using the log-rank Mantel-Cox test. ****P<0.0001.
All data representative of at least 2 independent experiments or biological replicates.
To evaluate the in-vivo therapeutic potential of GPNMB CAR T cells in melanoma, we used a xenograft model of human melanoma engrafted into immunodeficient mice. Mice were subcutaneously inoculated with melanoma cells, 22 days later they were randomized to receive either control T cells or GPNMB CAR T cells (Fig. 3B). Mice treated with GPNMB CAR T cells demonstrated significant expansion of human T cells in the peripheral blood of mice 21 days following T cell infusion compared to control T cell groups (Fig. 3C). This finding is consistent with target engagement and proliferation. Serum analysis of the blood showed significantly increased cytokine production in mice treated with GPNMB CAR T cells (Fig. 3D, Supp Fig 6B)
Tumor growth was longitudinally monitored in these animals and mice treated with GPNMB CAR T cells showed a striking reduction in tumor volume in comparison to mice treated with control T cells. Notably, the majority of mice treated with GPNMB CAR T cells underwent a complete remission (Fig. 3E). Significantly improved overall survival was observed with mice treated with GPNMB CAR T cells (Fig. 3F). All mice treated with control T cells reached humane endpoints and required sacrifice due to the progression of the flank tumor.
Together, these data demonstrate that GPNMB CAR T cells exert potent and sustained anti-melanoma activity in vivo, supporting their potential for clinical translation for patients with melanoma.
GPNMB CAR T cells show potent anti-tumor efficacy in colorectal cancer
Given the broad expression of GPNMB in various solid cancers we sought to assess the potential of this therapy for other solid cancers. GPNMB was expressed on the surface of many cancer cell lines from a variety of tumor origins (Fig. 2C). GPNMB CAR T cells consistently show anti-tumor lysis of GPNMB+ cells, including cells derived from patients with colorectal cancer, pancreatic cancer, and renal cancer (Fig. 3G).
We evaluated the in vivo potential of GPNMB CAR T cells against colorectal adenocarcinoma. Immunodeficient mice were engrafted with colorectal adenocarcinoma cells. 22 days after mice were randomized to either receive control T cells or GPNMB CAR T cells (Fig 3H). Similarly to our melanoma studies, mice treated with GPNMB CAR T cells showed significantly increased expansion of human T cells in the peripheral blood and increased cytokine production in the serum, consistent with target engagement and effector function (Fig 3I–J). Consistent with this on-target activation, mice treated with GPNMB CAR T cells showed a strong reduction in tumor volume when compared to mice treated with control T cells. All mice underwent a complete remission following treatment with GPNMB CAR T cells (Fig 3K). Mice that were treated with GPNMB CAR T cells showed significantly improved survival compared to mice treated with control T cells (Fig. 3L). All mice treated with control T cells reached humane endpoints and required sacrifice due to progression of the flank tumor.
Although GPNMB CAR T cells demonstrated potent anti-tumor activity in both solid tumor models, tumor recurrence was observed in a small minority of mice in the melanoma cohort, whereas no relapses were observed in the colorectal cancer model. Further analysis revealed that despite identical CAR T cell dosing across cohorts, the colorectal model exhibited markedly greater T cell expansion on day 21, accompanied by a pronounced CD8+ skew in the peripheral blood (Fig 3C,I; Supp Fig 6F). These findings suggest that tumor-intrinsic factors may differentially shape GPNMB CAR T cell expansion and efficacy across models.
Together, our data shows that GPNMB CAR T cells can result in sustained responses in colorectal cancer and could have broad therapeutic effects against a wide range of solid tumor malignancies.
Discussion:
The expansion of CAR T cell therapies throughout medicine hinges on the discovery of safe, target antigens. Current approaches remain slow, fragmented, and vulnerable to bias, underscoring the need for systematic and scalable solutions. Here, we present an AI-guided framework that integrates single-cell transcriptomics, public knowledge resources, and large language model–driven reasoning to rationally prioritize CAR T cell targets. This framework can lower the barrier for discovery in the CAR T field. Intriguingly, a similar framework could be applied to identify cancer specific targets leveraging orthogonal data types including proteomics, glycomics, or intracellular peptides for alternative T cell based modalities.
This study represents one of the first uses of large language models in the context of CAR T therapy. LLMs offer a scalable framework for accelerating target discovery. Our strategy to ground these models with real world data and public knowledge may be a way to reduce the risk of hallucinations. As LLMs continue to evolve, particularly through the emergence of agentic AI systems, their capacity to generate biologically informed reasoning will deepen. Beyond target identification, similar strategies could be extended to key challenges in the field—including CAR construct optimization, logic gating design, prioritization of genetic multiplexing, and clinical trial formulation. The convergence of artificial intelligence and engineered cell therapies signals a broader shift in how immunotherapies may be conceived, evaluated, and realized.
The tumor-specific effect on GPNMB CAR T cells is an important area for further exploration as well as the role of the local microenvironment across a range of malignancies highlighted by the differential complete responses seen in the colorectal tumor model and the melanoma tumor model. GPNMB’s expression across diverse malignancies may reflect its regulation by stress-response transcription factors within the MiT/TFE family that activate the CLEAR network, coordinating lysosomal biogenesis and autophagy39,40. These pathways are frequently co-opted in cancer to sustain growth, invasion, and immune evasion, potentially explaining GPNMB upregulation across a diversity of tumor lineages41,42. Understanding the mechanistic link may provide a biological basis for GPNMB’s broad distribution.
Expression of GPNMB across diverse tumor types, low presence in vital tissues, and the availability of clinically validated targeting domains from glembatumumab vedotin highlight the translational potential of this therapeutic. CAR T cells have demonstrated potency in patients refractory to antibody-based therapies. This may be due to the ability for CAR T cells to proliferate, penetrate, and persist. Thus, the fact that surface GPNMB-directed antibody therapy showed clinical feasibility, and some signs of efficacy provides strong rationale for the deployment of GPNMB-centric T cell therapy. Work to establish the long-term kinetics of a GPNMB CAR T cells could help to elucidate these potential effects. These results position GPNMB CAR T cells as a promising candidate for broad-spectrum immunotherapy.
While our analysis of GPNMB expression in healthy tissues suggests a favorable safety profile there remains the serious concern of on-target, off-tumor toxicity. The safety demonstrated in a clinical trial by an antibody-centric modality supports the safety of targeting this antigen28; however, CAR T cells show superior potency on the same target, so questions remain. During the final stages of this work a case report emerged clinically deploying GPNMB CAR T cells for alveolar soft part sarcoma (ASPS)43. Authors report that in this patient GPNMB CAR T cells show signs of activity, with no severe toxicity. We find this clinical data encouraging. While further safety characterization is warranted, the use of GPNMB CAR T cells in the clinic reinforces the strength of our platform and the multi-cancer potential of this therapeutic.
In conclusion, our study demonstrates the power of integrating advanced computational approaches with robust experimental validation to identify and develop promising CAR T cell therapies. The identification of GPNMB as a versatile target expressed across multiple cancer types, coupled with the impressive efficacy of GPNMB CAR T cells in preclinical models, provides strong rationale for the ongoing clinical translation. More broadly, our AI-guided discovery framework offers a scalable, generalizable approach for uncovering CAR T cell targets, with potential to extend across cancer and other diseases.
Limitations of the study:
This work focuses on human CAR T cells evaluated in immunodeficient xenograft models, which limits assessment of T-cell–extrinsic mechanisms, tumor-microenvironmental influences, and on-target, off-tumor toxicities. The GPNMB CAR construct is specific for the human antigen and does not recognize murine GPNMB, precluding direct safety testing in mouse tissues; however, immunohistochemistry across healthy human organs showed low expression in vital tissues. Recurrent melanoma tumors were unavailable for mechanistic analysis, and leukemia monitoring was confined to peripheral blood, where no circulating tumor cells were detected. In addition, in the target discovery stage, our scoring is limited by the number of studies currently available at single-cell resolution and will benefit from larger number of samples. Finally, our AI framework employed a generalized prompt structure to ensure cross-model applicability, which, while enhancing reproducibility, may underrepresent performance gains achievable through model-specific prompt optimization. Future work using humanized mouse models and refined prompting strategies will further strengthen translational and computational generalizability.
Resource availability
Lead contact:
Further information and requests for resources should be directed to the lead contact, Carl H June (cjune@upenn.edu).
Materials availability:
Materials generated in this study are available from the lead contact upon request.
Data and code availability:
Data generated from skin cancer dataset integration in this study can be accessed via the Zenodo repository: https://zenodo.org/records/19071268. The source code for this AI-driven discovery framework is available at https://github.com/hayatlab/cart_llms. Additional information required to analyze the data reported in this paper is available from the lead contact upon request.
STAR Methods:
Experimental Model Details
Cell lines and culture conditions
SK-MEL-5, HT-144, A-375, SK-MEL-28, SK-MEL-3, NALM-6, U-937, THP-1, and SW-156 were obtained from the American Type Culture Collection (ATCC). KLM-1 cells were obtained from RIKEN BioResource Research Center. SNU-503 cells were obtained from AddexBio. Cells underwent authentication from Arizona Genetics Core and routine testing for mycoplasma contamination by the Department of Genetics at the University of Pennsylvania (MycoAlert Mycoplasma Detection Kit, Lonza). All tumor cells were cultured in RPMI 1640 (gibco) supplemented with 10% heat-inactivated fetal bovine serum (FBS, Seradigm), 1% 1 M HEPES buffer solution (gibco), 1% 100X glutaMAX (gibco), and 1% Penicillin-Streptomycin (gibco).
Mouse models
Animal experiments were performed according to protocols approved by the Institutional Animal Care and Use Committee of the University of Pennsylvania. Five- to ten-week-old female NOD/scid/IL2rγ−/− (NSG) were bred in the vivarium at the University of Pennsylvania in pathogen-free conditions. Mice were maintained under pathogen free conditions. See below for tumor model specific details.
Method Details
Computational Methods
Single-cell atlas construction and data integration
Individual skin cancer single-cell datasets were obtained from Biermann et al. (2022), Ji et al. (2020), Li et al. (2019), and Yost et al. (2019)16–19. Basic filtering was applied to the individual datasets to remove low-quality cells and genes. The datasets were integrated using Harmony, with sample ID as the batch-effect variable, to obtain a unified single-cell atlas of skin cancer datasets20. The final skin cancer atlas consists of 285557 cells from 91 samples. From this dataset, the cells labeled as malignant (n=106834) were extracted from these integrated datasets and added to normal skin cells obtained from Tabula Sapiens (The Tabula Sapiens Consortium, 2022)21. The malignant and normal tabula sapiens data were again corrected for technical batch effects using Harmony, with sample ID as the batch variable. The final skin cancer atlas used in this study consists of 124620 cells from 111 samples (n = 106834 cancer, n = 17786 normal).
Extraction of target-relevant features from publicly available datasets:
To augment the target list derived from single-cell transcriptomics analyses, we programmatically integrated annotations from six public resources: ClinicalTrials.gov API v2 (version 2.15.0), OpenTargets (release 25.03), GTEx bulk RNA-seq medians (GTEx version 8), ProteomicsDB protein expression via its OData service (version 1.1), the Human Protein Atlas (HPA version 24.0), and UniProt (release 2025_02). ClinicalTrials.gov was queried for interventional trials per gene (expanded with gene synonyms harvested from HPA), OpenTargets was used to derive tractability modality classes, and HPA and UniProt provided independent subcellular localization evidence. Normal-tissue expression was summarized at the protein level from ProteomicsDB (iBAQ-derived normalized intensities) and at the transcript level from GTEx medians, each split into “vital” (lung, heart, brain) and “non-vital” tissues. Single-cell selectivity was estimated from publicly available scRNA-seq anndata objects22–25. All programmatic extraction was performed in Python 3.9.20 (conda-forge build; GCC 13.3.0) using scanpy 1.10.3, pandas 2.2.3, numpy 1.26.4, seaborn 0.13.2, matplotlib 3.6.3, requests 2.32.3, pyspark 3.5.3, and in R 4.2.3 using gtexr 0.1.0, dplyr 1.1.4, and tidyr 1.3.1 on Rocky Linux 8.10 (Green Obsidian), 4 CPU, 128GB. API endpoints and versioned archives are specified below to facilitate reproducibility. All code to reproduce this analysis is provided in the GitHub associated with this manuscript.
Empirical weighting of feature weights:
Clinical tractability
Clinical trial evidence was used as a proxy for tractability. For each gene, we queried ClinicalTrials.gov API v2 (endpoint https://clinicaltrials.gov/api/v2/studies, parameters query.term=[gene symbol], pageSize=1000, format=json) and restricted to interventional studies. To increase recall, search terms combined the HPA-reported gene symbol with its listed synonyms. For each returned study, we parsed protocolSection modules (identificationModule, statusModule, descriptionModule, conditionsModule, designModule, oversightModule, armsInterventionsModule, outcomesModule). Studies were retained if the gene or any synonym appeared in the title and/or text fields (title, brief summary, detailed description; case-insensitive). Each study received a phase score and a status score, to which we added a disease-relevance and a modality component. Phase scores were: Phase 4, 6 points; Phase 3, 5; Phase 2, 4; Phase 1, 3; Early Phase 1, 2; and N/A, 0. Status scores were: completed, 3 points; active not recruiting/recruiting/not yet recruiting, 2; unknown, 1; terminated/withdrawn/suspended, 0. A cancer relevance score of 10 points was added if the conditions string contained cancer-related terms (“cancer”, “melanoma”, “hcc”, “carcinoma”, “sarcoma”, “leukemia”, “lymphoma”, “tumor”, “neoplasm”, “oncology”, “metastatic”, “malignant”; case-insensitive), and 0 otherwise. Modality was derived from OpenTargets (release 25.03) tractability “modality” flags, loaded from JSON and parsed with PySpark. We assigned 10 points for antibody-based modalities (“AB”), 5 points for small molecules (“SM”), and 1 point for PROTAC/other chemical modalities (“PR” (Proteolysis Targeting Chimeras) or “OC” (other clinical modalities)); the modality score per gene was the sum of the maximum achievable points for each present modality class. The tractability score for a study was the sum of phase, status, cancer, and modality scores. Per gene, we retained only the single study with the highest combination of phase and status scores (ties resolved by the highest resulting tractability score), thereby avoiding score inflation from multiple low-evidence trials.
Subcellular localization
Orthogonal evidence for cell-surface localization was aggregated from HPA (v24.0) and UniProt (release 2025_02). From HPA, we derived “predicted membrane” status from “Protein class” and “plasma membrane” annotation from “Subcellular location.” From UniProt, we retrieved XML records via https://www.uniprot.org/uniprot/[UniProtID].xml and searched for “Plasma membrane” or “Cell membrane” terms in curated subcellular location fields. We scored each gene according to concordance and evidence strength as follows: known cell surface in both HPA and UniProt, 10 points; UniProt annotated cell surface and HPA “predicted membrane,” 8 points; cell surface annotated in only one resource (HPA “plasma membrane” or UniProt membrane annotation alone), 7 points; HPA “predicted membrane” without UniProt annotation, 5 points; and all other combinations, 0 points. These rules were implemented deterministically as: HPA=2 and UniProt=2 → 10; UniProt=2 and HPA=1 → 8; HPA=2 and UniProt=0 → 7; UniProt=2 and HPA=0 → 7; HPA=1 and UniProt=0 → 5; else 0, where HPA=2 indicates explicit “plasma membrane,” HPA=1 indicates “predicted membrane,” and UniProt=2 indicates a curated plasma/cell membrane location.
Protein abundance in normal tissues
We summarized protein expression using ProteomicsDB (version 1.1) via its OData service (service root https://www.proteomicsdb.org/proteomicsdb/logic/api/proteinexpression.xsodata; collections “InputParams” and “CA_EXPRESSION_PROFILE_API”). For each UniProt accession (mapped from gene symbols using the UniProt REST API at https://rest.uniprot.org/uniprotkb/search?query=gene:[SYMBOL]+AND+reviewed:true+AND+organism_id:9606), we executed the OData query InputParams(PROTEINFILTER=‘[UniProtID]’, MS_LEVEL=1, TISSUE_ID_SELECTION=‘’, TISSUE_CATEGORY_SELECTION=‘tissue; fluid’, SCOPE_SELECTION=1, GROUP_BY_TISSUE=1, CALCULATION_METHOD=0, EXP_ID=−1)/Results?$select= UNIQUE_IDENTIFIER, TISSUE_ID, TISSUE_NAME, NORMALIZED_INTENSITY&$format=json. NORMALIZED_INTENSITY values (iBAQ-based normalization from MaxQuant results, as provided by ProteomicsDB) were parsed per tissue, coerced to float, and averaged across “vital” tissues (lung, heart, brain) and across all remaining non-vital tissues. We recorded Avg_Vital_Expression, Avg_NonVital_Expression, and their ratio for downstream scoring.
Bulk RNA expression in normal tissues
We used GTEx medians to quantify gene expression burden in healthy tissues. Gene symbols were first mapped to GENCODE identifiers via gtexr::get_genes, and median expression per tissue was retrieved via gtexr::get_median_gene_expression([itemsPerPage setting, reference panel, units]) in R 4.3.3. After pivoting genes × tissues, we computed VitalAverage as the mean of tissues whose names matched “Lung|Heart|Brain” (case-insensitive) and NonVitalAverage as the mean of all other tissues for each gene. Where a gene had multiple rows per tissue, we used the mean to collapse duplicates. These bulk metrics were inverted during normalization (see below), as lower expression in normal tissues is preferred.
Cancer composition
The cancer composition score was computed in Python using scanpy. The integrated skin atlas was normalized to 1e6 total counts per cell (sc.pp.normalize_total(target_sum=1e6)). Cells annotated as “Malignant” in adata.obs[‘cell_type’] were considered malignant; all others were considered non-malignant. For each gene, we calculated the percentages of malignant and non-malignant cells with expression >0. Z-scores were then computed for these percentages across all genes within each population using , where is the gene’s percentage, is the across-gene mean, and is the across-gene standard deviation. The enrichment Z-score per gene was defined as Zmalignant – Znon-malignant, quantifying over-representation in malignant relative to non-malignant cells. This enrichment statistic was used as the single-cell feature in downstream normalization and weighting. The analysis was conducted using Python 3.9.20, scanpy 1.10.3, numpy 1.26.4, and pandas 2.2.3.
Normalization and aggregation
Prior to aggregation, each feature was normalized to [0,1] across genes using min–max scaling; features in which lower values indicate desirability, namely GTEX-based bulk RNA expression in normal tissues, and ProteomicsDB-based protein abundance in normal tissues, were inverted. To mitigate duplicate entries across sources, per-dataset gene-level scores were deduplicated by taking the maximum available score per gene. Because features did not share identical gene universes, we aligned the union of all gene sets; missing feature values were imputed as 0. The final prioritization score for each gene was the weighted sum of normalized features using the llm-retrieved weighting. These weights were applied uniformly across genes without further scaling, and the resulting composite scores were used to rank targets. All code paths for normalization, inversion, deduplication, alignment, and weighted aggregation are provided in the accompanying repository/notebook, with file paths to intermediate CSVs and JSONs recorded in the scripts; exact software versions and hardware details should be inserted as follows: requests 2.32.3, pyspark 3.5.3, R 4.2.3, gtexr 0.1.0, dplyr 1.1.4, Rocky Linux 8.10 (Green Obsidian), 4 CPU, 128GB.
Selection of optimal feature weights using LLMs:
To identify an empirically grounded weighting scheme for integrating tractability, accessibility, and safety features into CAR T target prioritization, we implemented an automated LLM-driven evaluation pipeline in Python that repeatedly queried independent frontier models with a fixed prompt and parsed their outputs. The pipeline programmatically assembled candidate gene lists produced under alternative weighting regimes and summarized each feature definition (Clinical tractability, subcellular localization, cancer composition, GTEx bulk expression in vital and non-vital tissues, and ProteomicsDB protein abundance in vital and non-vital tissues). A standardized, model-agnostic instruction began with a mandatory machine-readable header to force structured output, namely: “RECOMMENDED WEIGHTS: Clinical trials (X.X), HPA (X.X), Heatmap (X.X), Vital (X.X), Non-vital (X.X), Protein vital (X.X), Protein non-vital (X.X)”.
The full prompt using the single cell data and all weighting schemes was:
“““As an expert in cancer biology and computational oncology, analyze these weighted scoring results for potential CAR T targets.
The scoring features include:
Clinical Trials: Scores based on clinical trial progression (phase reached), availability of antibodies/drugs, and number of cancer-related studies
HPA: Cell surface localization scores from Human Protein Atlas and UniProt (10=confirmed surface in both, 8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Heatmap: Z-score difference between expression in malignant vs. normal cells from single-cell data
Vital/Non-Vital: GTEx bulk RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Vital/Protein Non-Vital: Protein expression levels in vital vs. non-vital tissues from ProteomicsDB
Current weights are:
Base weight: Clinical trials (1.0), HPA (2.0), Heatmap (5.0), Vital (2.0), Non-vital (1.0), Protein vital (2.0), Protein non-vital (1.0)
Clinical trials emphasis: Clinical trials (5.0), HPA (2.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
HPA emphasis: Clinical trials (1.0), HPA (5.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Heatmap emphasis: Clinical trials (1.0), HPA (2.0), Heatmap (10.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Expression emphasis: Clinical trials (1.0), HPA (1.0), Heatmap (2.0), Vital (5.0), Non-vital (2.0), Protein vital (5.0), Protein non-vital (2.0)
Balanced approach: Clinical trials (2.0), HPA (2.0), Heatmap (2.0), Vital (2.0), Non-vital (1.0), Protein vital (2.0), Protein non-vital (1.0)
Equal weights: All features weighted equally (1.0)
Based on your expertise in CAR T development, what would be the optimal weighting scheme to prioritize genes that are:
Highly specific to tumor cells vs. normal tissue (high heatmap Z-score)
Accessible on the cell surface (high HPA score)
Limited expression in vital tissues (low vital/protein vital scores)
Have existing clinical evidence supporting targetability (clinical trials score)
IMPORTANT: Start your response with a clear specification of your recommended weights in this exact format:
“RECOMMENDED WEIGHTS: Clinical trials (X.X), HPA (X.X), Heatmap (X.X), Vital (X.X), Non-vital (X.X), Protein vital (X.X), Protein non-vital (X.X)”
Then explain your reasoning for the recommended weights and how they align with established principles for identifying ideal CAR T targets. Which of the top-ranked genes from your recommended weighting scheme appear most promising for further investigation?”””
The full prompt for the version that excluded the expert weights was:
“““As an expert in cancer biology and computational oncology, analyze these weighted scoring results for potential CAR T targets.
The scoring features include:
Clinical Trials: Scores based on clinical trial progression (phase reached), availability of antibodies/drugs, and number of cancer-related studies
HPA: Cell surface localization scores from Human Protein Atlas and UniProt (10=confirmed surface in both, 8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Heatmap: Z-score difference between expression in malignant vs. normal cells from single-cell data
Vital/Non-Vital: GTEx bulk RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Vital/Protein Non-Vital: Protein expression levels in vital vs. non-vital tissues from ProteomicsDB
Current weights are:
Clinical trials emphasis: Clinical trials (5.0), HPA (2.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
HPA emphasis: Clinical trials (1.0), HPA (5.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Heatmap emphasis: Clinical trials (1.0), HPA (2.0), Heatmap (10.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Expression emphasis: Clinical trials (1.0), HPA (1.0), Heatmap (2.0), Vital (5.0), Non-vital (2.0), Protein vital (5.0), Protein non-vital (2.0)
Balanced approach: Clinical trials (2.0), HPA (2.0), Heatmap (2.0), Vital (2.0), Non-vital (1.0), Protein vital (2.0), Protein non-vital (1.0)
Equal weights: All features weighted equally (1.0)
Based on your expertise in CAR T development, what would be the optimal weighting scheme to prioritize genes that are:
Highly specific to tumor cells vs. normal tissue (high heatmap Z-score)
Accessible on the cell surface (high HPA score)
Limited expression in vital tissues (low vital/protein vital scores)
Have existing clinical evidence supporting targetability (clinical trials score)
IMPORTANT: Start your response with a clear specification of your recommended weights in this exact format:
“RECOMMENDED WEIGHTS: Clinical trials (X.X), HPA (X.X), Heatmap (X.X), Vital (X.X), Non-vital (X.X), Protein vital (X.X), Protein non-vital (X.X)”
Then explain your reasoning for the recommended weights and how they align with established principles for identifying ideal CAR T targets. Which of the top-ranked genes from your recommended weighting scheme appear most promising for further investigation?”””
The zero-shot prompt was, so excluding each predetermined weighting scheme and gene short lists was:
“““As an expert in cancer biology and computational oncology, analyze these weighted scoring results for potential CAR T targets.
The scoring features include:
Clinical Trials: Scores based on clinical trial progression (phase reached), availability of antibodies/drugs, and number of cancer-related studies
HPA: Cell surface localization scores from Human Protein Atlas and UniProt (10=confirmed surface in both, 8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Heatmap: Z-score difference between expression in malignant vs. normal cells from single-cell data
Vital/Non-Vital: GTEx bulk RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Vital/Protein Non-Vital: Protein expression levels in vital vs. non-vital tissues from ProteomicsDB
Based on your expertise in CAR T development, what would be the optimal weighting scheme to prioritize genes that are:
Highly specific to tumor cells vs. normal tissue (high heatmap Z-score)
Accessible on the cell surface (high HPA score)
Limited expression in vital tissues (low vital/protein vital scores)
Have existing clinical evidence supporting targetability (clinical trials score)
IMPORTANT: Start your response with a clear specification of your recommended weights in this exact format:
“RECOMMENDED WEIGHTS: Clinical trials (X.X), HPA (X.X), Heatmap (X.X), Vital (X.X), Non-vital (X.X), Protein vital (X.X), Protein non-vital (X.X)”
Then explain your reasoning for the recommended weights and how they align with established principles for identifying ideal CAR T targets. Which of the top-ranked genes from your recommended weighting scheme appear most promising for further investigation?”””
The full prompt for the bulk version was:
“““As an expert in cancer biology and computational oncology, analyze these weighted scoring results for potential CAR T targets.
The scoring features include:
Clinical Trials: Scores based on clinical trial progression (phase reached), availability of antibodies/drugs, and number of cancer-related studies
HPA: Cell surface localization scores from Human Protein Atlas and UniProt (10=confirmed surface in both, 8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Heatmap: Z-score difference between log fold change in malignant vs. normal cells from bulk rna data
Vital/Non-Vital: GTEx bulk RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Vital/Protein Non-Vital: Protein expression levels in vital vs. non-vital tissues from ProteomicsDB
Current weights are:
Base weights: Clinical trials (1.0), HPA (2.0), Heatmap (5.0), Vital (2.0), Non-vital (1.0), Protein vital (2.0), Protein non-vital (1.0)
Clinical trials emphasis: Clinical trials (5.0), HPA (2.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
HPA emphasis: Clinical trials (1.0), HPA (5.0), Heatmap (3.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Heatmap emphasis: Clinical trials (1.0), HPA (2.0), Heatmap (10.0), Vital (1.0), Non-vital (0.5), Protein vital (1.0), Protein non-vital (0.5)
Expression emphasis: Clinical trials (1.0), HPA (1.0), Heatmap (2.0), Vital (5.0), Non-vital (2.0), Protein vital (5.0), Protein non-vital (2.0)
Balanced approach: Clinical trials (2.0), HPA (2.0), Heatmap (2.0), Vital (2.0), Non-vital (1.0), Protein vital (2.0), Protein non-vital (1.0)
Equal weights: All features weighted equally (1.0)
Based on your expertise in CAR T development, what would be the optimal weighting scheme to prioritize genes that are:
Highly specific to tumor cells vs. normal tissue (high heatmap Z-score)
Accessible on the cell surface (high HPA score)
Limited expression in vital tissues (low vital/protein vital scores)
Have existing clinical evidence supporting targetability (clinical trials score)
IMPORTANT: Start your response with a clear specification of your recommended weights in this exact format:
“RECOMMENDED WEIGHTS: Clinical trials (X.X), HPA (X.X), Heatmap (X.X), Vital (X.X), Non-vital (X.X), Protein vital (X.X), Protein non-vital (X.X)”
Then explain your reasoning for the recommended weights and how they align with established principles for identifying ideal CAR T targets. Which of the top-ranked genes from your recommended weighting scheme appear most promising for further investigation?”””
We used multiple weighting schemes that focused on different scoring features, and our experts provided one weighting scheme. Using each scheme, a top-100 gene list was extracted and provided to the LLMs, including the weighting scheme itself. To mitigate bias, we did not state to the LLMs that an expert provided one of the schemes and named that weighting scheme the base weight.
We queried three independent models via their official SDKs: OpenAI GPT-4o (model “gpt-4o”, OpenAI Python client version 1.66.5, Anthropic Claude 3.7 Sonnet (model “claude-3-7-sonnet-20250219”, anthropic 0.49.0), and Google Gemini 2.5 Pro (model “gemini-2.5-pro-preview-03-25”, google-generativeai 0.8.4). Unless stated below, SDK defaults were retained. We performed 1,000 independent queries per model, varying the sampling temperature from 0.7 to 0.85 to sample the response distribution while preserving formatting fidelity. For GPT-4o and Claude, the temperature was set within this range, with other parameters at their defaults; for Claude and Gemini, we explicitly constrained the maximum output tokens to 4,096 and 4,000, respectively, to match the observed GPT-4o response lengths and mitigate truncation. For Gemini, we used top_k = 40 and reduced top_p to 0.9 after piloting revealed occasional formatting lapses at higher top_p settings; these settings eliminated header omissions without altering content qualitatively. All requests and responses were logged with timestamps, model identifiers, and approximate token accounting; raw responses were persisted per run as JSON for full reproducibility.
For downstream analyses, we used a tiered, regex-based parser to extract numeric weights from free text implemented in Python. The primary pattern required an exact, case-insensitive match to the canonical header line with seven features in fixed order and decimal values in parentheses. If absent, a second pattern accepted Markdown-bold headers with the same order and formatting. Failing these, we applied a label-wise flexible extractor that matched either integer or decimal values adjacent to each label, allowed optional parentheses, and used a negative lookbehind on “Vital” to prevent collisions with “Protein vital.” If at least four of seven features were recovered with high confidence at this stage, we returned the partial dictionary; otherwise, a final fallback scanned the response line-by-line for occurrences of each label and captured the first decimal on that line. Each run thus yielded a record containing the model, temperature, and a dictionary of parsed weights; runs with fewer than four fields remained in the raw archive but were excluded from summary statistics.
We aggregated extracted weights across runs and models in pandas (2.2.3) and computed per-feature and per-model summary statistics (mean, median, standard deviation, minima, maxima) and the empirical frequency of weight “modes,” defined as tuples of the seven weights rounded to one decimal place. For visualization, we used seaborn (0.13.2) and matplotlib (3.10.1) to produce box-and-violin distributions by feature and model, heatmaps of model-specific mean weights, radar plots of feature profiles, and stacked bar charts of relative feature importance (weights normalized to unit sum within run). We further assessed between-model differences in recommended weights via one-way ANOVA per feature and model using scipy (1.15.2), and we quantified any systematic temperature effects using ordinary least-squares regressions of weight on temperature per model-feature pair, correcting p-values across tests using Benjamini–Hochberg FDR in statsmodels (0.14.4). To facilitate auditability, we saved the full weight matrix, summary tables, top combination tallies, and figure assets (PNG/SVG) to a versioned results directory; run-wise line plots provided a visual check for drift or mode-switching over time. All analyses ran on Python 3.10.16 in a Linux environment Rocky Linux 8.10 (Green Obsidian), with numpy 2.2.3, regex 2024.11.6, requests 2.32.3, tqdm 4.67.1, and standard logging.
This LLM-based procedure yields a reproducible ensemble of weight recommendations conditioned on identical evidence summaries and candidate lists, while the strict output header and multi-stage parser minimize formatting-induced data loss. By design, the approach complements our data-driven scoring by externalizing expert-like trade-offs between tumor selectivity (single-cell enrichment), accessibility (cell-surface localization), on-target/off-tumor risk (vital-tissue burden in bulk RNA and protein), and translational feasibility (clinical trial tractability). The consolidated distributions and consensus statistics from the three model families were used downstream to select a final weighting scheme for target ranking, as reported in the Empirical weighting section, with all intermediate artifacts and code paths recorded to enable full reruns on future model snapshots.
Simulation of target prioritization reasoning using LLMs:
To emulate expert reasoning for melanoma CAR T target nomination, we implemented an automated, multi-model large language model (LLM) pipeline that submitted a standardized prompt to three commercial APIs (OpenAI GPT-4o, Anthropic Claude-3.7 (Sonnet 20250219), and Google Gemini-2.5-Pro (preview 03–25)) and parsed their structured outputs for further downstream analyses. The prompt contained the fixed set of 100 candidate genes from the empirical weighting stage and mandated a two-part output: (i) a CSV-formatted “Top 10 Candidate Selection” using exactly five nomination reason phrases, and (ii) structured narrative assessments for each top gene. The prompt using single cell data was:
“““As an expert in cancer immunotherapy and molecular oncology, analyze this list of 100 genes identified through computational screening for CAR T targets, focusing specifically on melanoma.
STRUCTURED EVALUATION REQUEST:
Review the following genes, considering the provided scoring features:
[GENES LIST: [ERBB3, RNF43, PMEL, MC1R, GPNMB, LZTS1, HPS4, IGF1R, SORT1, TNFRSF14, NRG2, MAK, ROBO1, NRG4, PTCH1, ADIPOR2, PKD2, CDH1, BLNK, TCHP, LTK, TNFSF13B, EPS8, CDH3, DDR2, LAT, FHIT, EPOR, PIGF, NKTR, BMPR2, ABCB4, MLLT10, DST, CD226, CHRNA5, NEO1, IFNAR2, MYO10, CTNS, F2RL1, ERBB2, PDCD1LG2, HTR2B, APC, TGFBR1, TRPM7, CNIH3, BCL2L12, RAD51B, EPHA2, WDPCP, PDE4D, FUT2, HCAR2, ADAM12, GPER1, PTK2, TNK2, ITGA6, FCGR2B, NPIPA1, GPR55, WLS, PPL, AXIN2, IL6R, HAS3, P2RX7, SLC22A1, RAF1, LAPTM4B, SLC38A6, BRAF, ZNRF3, RASGRP1, ATP10A, NF1, IL1RAP, PMP22, STIL, INSR, CATSPER2, CADM1, GRK4, TULP3, TMEM241, TMEM117, TRPV2, ABCA7, CTTN, SLC16A1, SLC39A4, LEPR, CRCP, SLC47A1, WDR83OS, MCOLN3, MR1, MME]
SCORING CRITERIA (Previously Used):
Clinical Trials: Scores based on clinical trial progression (phase reached), antibody/drug availability, and number of cancer-related studies
Cell Surface Localization: Human Protein Atlas and UniProt data (10=confirmed in both, 8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Expression Difference: Z-score difference between malignant vs. normal cells from single-cell data
Vital vs. Non-Vital Tissue Expression: RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Expression: Protein levels in vital vs. non-vital tissues from ProteomicsDB
REQUIRED OUTPUT:
Part 1: Top 10 Candidate Selection
Present the 10 most promising genes for CAR T development in melanoma in a clearly ranked order (1–10).
For each gene, provide:
Rank (#1–10)
Gene Symbol
Brief description (1–2 sentences on function/role)
Primary nomination reason (select ONE from these categories only):
* High tumor-specific expression (>X% of melanoma cells)
* Potential utility across multiple cancer types
* Strong existing clinical development profile
* Minimal expression in healthy vital tissues
* Expression in significant patient subpopulations
Part 2: Individual Assessments
For each of the top 10 genes, provide a structured evaluation:
-
MELANOMA RELEVANCE:
Expression pattern in melanoma
Functional significance in melanoma pathology
Clinical correlations (if known)
-
SAFETY ASSESSMENT:
Expression in vital normal tissues
Known essential functions
Potential on-target/off-tumor toxicity concerns
-
TECHNICAL FEASIBILITY:
Cell surface localization confidence
Antibody/binding domain availability
Previous CAR development (if any)
-
BROADER APPLICABILITY:
Evidence for utility in other cancer types
Known mutations or variants affecting targetability
Base all assessments on established scientific knowledge and clinical evidence. Indicate any areas of uncertainty.”””
The prompt for using bulk expression data was:
“““As an expert in cancer immunotherapy and molecular oncology, analyze this list of 100 genes identified through computational screening for CAR T targets, focusing specifically on melanoma.
STRUCTURED EVALUATION REQUEST:
Review the following genes, considering the provided scoring features:
[GENES LIST: [PRAME, IL13RA2, MICB, ICOS, SLCO1A2, GPR143, GRPR, NTSR1, TNFRSF8, RET, LY6K, RHD, LTK, SLC30A8, CCR4, TNFSF13B, MC1R, CCR5, HTR3A, MC5R, TYR, MCHR1, MAL, CTLA4, PSCA, SLC45A2, SSTR5, EPOR, NTRK1, HTR2B, CHRNA5, CRHR2, BLNK, CD52, TDO2, DRD1, BTLA, HTR2C, RNF43, FASLG, MPL, CSF3R, CXCR6, PTCH1, FFAR4, ABCB5, LZTS1, CXCR1, NRG2, ADORA3, SIGLEC6, OPRL1, ABCB4, CCR8, TIGIT, OSCAR, CTXN1, ABCG5, SLC5A4, TAS1R3, PVRIG, DIO2, DRD4, ROS1, CXCR3, TFR2, GPR35, NRG4, BOC, PTGDR, LHCGR, CD1D, HPN, TRPV1, PMEL, PMP22, ALK, DLL3, EPHA5, CD83, XCR1, CHRNA1, FCGR2B, CD226, STAC, ERBB3, CXCR2, GNRHR, AVPR2, BCL2L12, SSTR1, HCRTR1, HCST, ADAM12, DCT, GPR55, DLL4, HTR3B, L1CAM, CHRNA7]
SCORING CRITERIA (Previously Used):
Clinical Trials: Scores based on clinical trial progression (phase reached), antibody/drug availability, and number of cancer-related studies
Cell Surface Localization: Human Protein Atlas and UniProt data (10=confirmed in both,8=confirmed in UniProt/predicted in HPA, 7=confirmed in one source only, 5=predicted in HPA only)
Expression Difference: Z-score difference between malignant vs. normal cells from differential gene epression in bulk data
Vital vs. Non-Vital Tissue Expression: RNA expression in vital tissues (brain, lung, heart) vs. non-vital tissues
Protein Expression: Protein levels in vital vs. non-vital tissues from ProteomicsDB
REQUIRED OUTPUT:
Part 1: Top 10 Candidate Selection in CSV Format
Present the 10 most promising genes for CAR T development in melanoma in a clearly ranked order (1–10).
For each gene, provide:
Rank (#1–10)
Gene Symbol
Brief description (1–2 sentences on function/role)
ALL applicable nomination reasons from ONLY these five categories (list all that apply, between 1–5 reasons max):
* High tumor-specific expression (>X% of melanoma cells)
* Potential utility across multiple cancer types
* Strong existing clinical development profile
* Minimal expression in healthy vital tissues
* Expression in significant patient subpopulations
IMPORTANT:
Please use EXACTLY these five reason phrases and do not create additional reasons. Do not modify the wording.
Your output for Part 1 should be structured in a CSV-like format with the following columns:
Rank, Gene, Description, Reason1, Reason2, Reason3, Reason4, Reason5
Example:
GENE1, Description of gene function, High tumor-specific expression, Minimal expression in healthy vital tissues,,
GENE2, Description of gene function, Strong existing clinical development profile, Potential utility across multiple cancer types, Expression in significant patient subpopulations,,
Part 2: Individual Assessments
For each of the top 10 genes, provide a structured evaluation:
-
MELANOMA RELEVANCE:
Expression pattern in melanoma
Functional significance in melanoma pathology
Clinical correlations (if known)
-
SAFETY ASSESSMENT:
Expression in vital normal tissues
Known essential functions
Potential on-target/off-tumor toxicity concerns
-
TECHNICAL FEASIBILITY:
Cell surface localization confidence
Antibody/binding domain availability
Previous CAR development (if any)
-
BROADER APPLICABILITY:
Evidence for utility in other cancer types
Known mutations or variants affecting targetability
Base all assessments on established scientific knowledge and clinical evidence. Indicate any areas of uncertainty.”””
To minimize instruction-following variance and facilitate deterministic parsing, we required that Part 1 be formatted as Rank, Gene, Description, Reason1, Reason2, Reason3, Reason4, Reason5 and that nomination reasons be drawn verbatim from a predefined five-category lexicon (“High tumor-specific expression (>X% of melanoma cells)”, “Potential utility across multiple cancer types”, “Strong existing clinical development profile”, “Minimal expression in healthy vital tissues”, “Expression in significant patient subpopulations”). No additional reason phrasings were accepted during parsing.
Programmatic access was implemented in Python 3.10.16 using openai 1.66.5, anthropic 0.49.0, google-generativeai 0.8.4 (and google-ai-generativelanguage 0.6.15), pandas 2.2.3, numpy 2.2.3, regex 2024.11.6, tqdm 4.67.1, matplotlib 3.10.1, seaborn 0.13.2, matplotlib-venn 1.1.2, upsetplot 0.9.0, and plotly 6.0.1 with kaleido 0.2.1 for static export. API keys were provided via environment variables, and all requests were logged to llm_queries.log (timestamp, model, run_id, temperature, [HTTP status]). For each model, we executed 1,000 independent runs. To induce modest sampling variability while preserving comparability, we cycled temperatures across runs as 0.70, 0.75, 0.80, and 0.85, where temperature = 0.7 + 0.2 × (run_id mod 4)/4. For GPT-4o (model=gpt-4o), we adjusted temperature only; for Claude-3.7 (model=claude-3-7-sonnet-20250219), we additionally set max_tokens=4,000 to improve CSV compliance; for Gemini-2.5-Pro (model=gemini-2.5-pro-preview-03-25), we used the default output length with temperature matching the schedule above. Each provider was queried through the official SDKs with exponential backoff on transient failures or rate-limit responses.
To ensure high-throughput yet compliant usage, a thread-safe rate limiter governed both concurrency and approximate token budgets per provider. For OpenAI, Anthropic, and Gemini, we set tokens_per_min to 28,000, 90,000, and 50,000, and max_parallel_requests to 3, 5, and 4, respectively. A per-request semaphore limited in-flight calls, and a token counter reset every 60 s; prospective requests waited when estimated usage would exceed the per-minute budget. Token usage was computed exactly for GPT-4o using the response.usage (prompt + completion tokens) and approximated for Claude and Gemini as floor(len(prompt + response)/4). Requests were submitted in batches (OpenAI=20, Claude=30, Gemini=25) with short inter-batch pauses (5 s, 3 s, 4 s). Raw responses were persisted run-by-run as JSON files (openai_run_[id].json, claude_run_[id].json, gemini_run_[id].json) and periodically consolidated into model-level archives. The driver respected existing outputs to support restartable execution; missing runs were scheduled until n=1,000 per model was reached. No deterministic random seeds were available to constrain generative sampling; therefore, each run should be regarded as an independent stochastic draw from the same prompt-model configuration.
Model outputs were parsed by a deterministic, rule-based extractor written in Python that prioritizes the CSV in Part 1 and falls back to robust pattern matching. First, if a fenced CSV block was detected—common in Gemini responses—the CSV section was parsed line-wise after removing the header, enforcing that ranks were integers 1–10 and that gene symbols matched a predefined whitelist of the 100 candidates by word boundary. If no fenced block was present, a permissive CSV-like regex captured lines of the form “rank, gene, description [,reason(s)]”. When both CSV strategies failed, a flexible fallback scanned the entire response for gene mentions, inferred ranks from local patterns (e.g., “Rank 1”, “#1”, “1) Gene”), and extracted brief descriptions from the trailing clause. Candidate nomination reasons were then assigned by searching a fixed context window centered on the gene mention (approximately 300 characters upstream and 700 downstream) for exact phrase matches to the five permitted reasons; if no exact phrase was found, we allowed controlled keyword anchoring to the closest canonical reason (e.g., “overexpression/specific expression” → “High tumor-specific expression”, “other cancers” → “Potential utility across multiple cancer types”, “clinical trial/antibody available” → “Strong existing clinical development profile”, “low/limited in vital tissues” → “Minimal expression in healthy vital tissues”, “subset/patient population” → “Expression in significant patient subpopulations”). Each gene–reason pair was counted at most once per run. Across all strategies, parsing stopped after the first 10 unique genes per run in ascending rank order.
Parsed outputs were normalized to a per-run, per-model long format that retained, for each ranked gene, its position (1–10), terse description, and the set of assigned nomination reasons. For summary statistics, we collapsed to unique (model, run_id, gene) and (model, run_id, gene, reason) tuples to prevent double-counting within a run. Gene nomination frequency and average rank were computed per model and across models, and nomination reason distributions were summarized per model. Co-occurrence of reasons for the same gene within a run was quantified by run-level set intersections and visualized via co-occurrence heatmaps and Venn diagrams (top three reasons). To analyze the intersectionality and heterogeneity of reasoning across models, we generated UpSet plots using boolean indicators per run for the five reasons, stacked by model, and constructed an alluvial (Sankey) diagram linking model → gene → reason with link widths proportional to unique run counts. For presentation-ready figures, we exported vector graphics (SVG) with editable fonts (rcParams: svg.fonttype=none; pdf/ps fonttype=42; sans-serif family prioritized Arial/Helvetica with DejaVu fallback) and provided summary bar plots of the top genes per model and stacked distributions of reason usage with run-to-run variability shown as standard-deviation error bars. All visualization scripts are included in the analysis driver and write to prompt2_visualizations_revised; consolidated CSVs include per-model gene rankings, per-run reason assignments, average ranks, and cross-model comparisons.
Reproducibility notes - The LLM prompt, model identifiers, temperature schedule, and parsing rules were held constant across runs, but stochastic decoding introduces irreducible variation in content and formatting. The rule-based parser enforces a fixed ontology of five nomination reasons. The Claude max_tokens setting (4,000) was required to maintain consistent CSV emission; GPT-4o and Gemini did not require an explicit output-length cap in our experiments. We verified that minor variations in temperature within 0.70–0.85 did not qualitatively alter the aggregate ranking or the distributions of dominant reasons at n=1,000 runs/model.
Target filtering and ranking
The resulting targets were first filtered by only including genes that have at least a single nomination by each model. The targets were then ranked by the total number of nominations across the three models. Visualization modules produced heatmaps, barplots, and UpSet diagrams to characterize both the reproducibility and diversity of LLM-driven target nomination. This approach enabled a systematic comparison of LLM-generated prioritization rationales, highlighting both convergent and model-specific reasoning patterns for CAR T target selection.
Bulk RNAseq vs. Single Cell RNAseq as Input
We integrated TCGA-SKCM primary tumor RNA-seq raw counts (STAR–Counts) with GTEx sun-exposed skin (lower leg) read counts, intersecting Ensembl genes after removing version suffixes and filtering to genes with ≥10 reads in ≥10% of samples. Differential expression was performed with DESeq2 v1.46.0 (R 4.4.3) using design ~ condition with Normal as reference, yielding Tumor-vs-Normal log2 fold-changes under DESeq2’s size-factor normalization; results were extracted without independent filtering or Cook’s cutoff to retain the full gene set. We standardized these log2 fold-changes across genes to obtain Z-scores, which replaced the prior composition-based Z-score in the downstream pipeline, and updated the prompt to “Z-score of DESeq2 log2 fold-change (TCGA SKCM tumors vs GTEx sun-exposed skin).”
Bulk RNAseq TCGA analysis
http://gepia.cancer-pku.cn/index.html was used to analyze the expression of candidate genes for various malignancies. Differential gene expression was analyzed looking at malignant vs TCGA normal+GTEx normal. The method for differential analysis is one-way ANOVA, using disease state (Tumor or Normal) as variable for calculating differential expression.
In vitro studies
Immunofluorescence
Cancer cells were fixed using 4% PFA at room temperature (RT) for 15 minutes. Samples were then permeabilized by 0.2% Triton X-100/PBS and incubated with blocking buffer (1% BSA, 5% goat serum in PBS) at RT for 1 hour. Primary antibodies anti-GPNMB (R&D, AF2550, 1:100), PE conjugated anti-HER-3 (Biolegend, 324706, 1:100), PE conjugated anti-IGF-1R (R&D, FAB391P, 1:100), and anti-IgG (R&D, F0102B, 1:100) were incubated overnight at 4 °C, followed with staining of secondary antibody (Invitrogen, PA1-29953, 1:200, BD Pharmingen, 568275, 1:200 or Thermo Fisher Scientific, A32727, 1:200) at RT for 2 hours if required. Samples were stained with DAPI solution for nuclei, then mounted by antifade mountant (Thermo Fisher Scientific, P36935), and examined by fluorescence microscope (Keyence).
Non-permeabilized flow cytometry of cells
All cells were stained in 1X DPBS (gibco) containing 5% heat-inactivated FBS (Seradigm). Staining was conducted under non permeabilized conditions to assay surface expression of target of interest. Antibodies used in this study include: GPNMB (R&D, AF2550. 1:50), ERBB3 (Biolegend, 324706, 1:100), IGF1R (R&D Systems, FAB391P, 1:100), anti-human IgG F(ab’)2 fragment specific (Jackson ImmunoResearch Labs, AB_2337634, 1:50). Secondary antibodies that were used include: anti-Goat IgG (H+L) PE (Invitrogen, PA1-29953, 1:100), anti-PE Mouse IgG2a (BioLegend, 400214, 1:100), mouse IgG1 PE (R&D Systems, IC002P, 1:100), PE Streptavidin (BioLegend, 405204, 1:100). BD FACSymphony™ A5 SE Cell Analyzer (BD Biosciences) equipped with the FACSDiva software (BD Biosciences) was used for flow cytometry.
Immunohistochemistry
Tissue microarray slide (MC5004b and BN1021a) was obtained from tissuearray.com. Staining was done using antibody against GPNMB (clone E4D7P, Cell Signaling, cat# 38313, diluted at 1:250). Staining was done on a Leica Bond-IIITM instrument using the Bond Polymer Refine Detection System (Leica Microsystems DS9800). Heat-induced epitope retrieval was done for 20 minutes with ER2 solution (Leica Microsystems AR9640). Experiments were done at room temperature. Slides were washed three times between each step with a bond wash buffer or water.
CAR Design and Construction
VL and VH sequences of CR011, a fully human GPNMB antibody (WO2010135547A1), were combined with a (G4S)3 linker and inserted into a CAR framework containing a Kozak sequence, CD8α leader, CD8α hinge and transmembrane domains, 4–1BB and CD3zeta endodomains, and cloned into pTRPE, a third-generation self-inactivating lentiviral vector.
Lentiviral vector production
HEK293T cells were transfected with lentiviral CAR and packaging plasmids using Lipofectamine 3000 (Invitrogen) following the manufacturer’s protocol. Lentiviral supernatants were collected at 24 post-transfection and concentrated using high-speed ultracentrifugation. To generate the lentiviral stocks, the resulting concentrated lentivirus batches were resuspended in cold R10 media and stored at −80°C.
Primary human CAR T cell expansion and transduction
Primary human CD4+ T and CD8+ T cells from normal donors were provided by the University of Pennsylvania Human Immunology Core. CAR T cells were generated as previously described44. Briefly, CD4+ and CD8+ T at 1:1 ratio at 1×10e6 cells/ml were activated with Dynabeads® CD3/CD28 CTS™ (Thermofisher) at a 3:1 bead-to-cell ratio. Approximately 24hrs later, T cells were transduced at a multiplicity of infection (MOI) of 3 to 5. At day 5 beads were removed from cultures. T cell cultures were maintained at 8 × 10e5 cells/ml. Cell number and volume were monitored daily using a Multisizer 4 Coulter Counter (Beckman). Transduced T cells were cryopreserved when they reached the resting state, as determined by cell size. T cells were thawed and immediately used for in vitro and in vivo experiments
Cytokine supernatant analysis
Cytokine concentrations were quantified using the BioLegend Human Th1/Th2/Th17 Cytometric Bead Array (CBA) Kit following the manufacturer’s protocol (BioLegend, Cat. No. 560484) as per the manufacturer’s instructions. Briefly, samples were incubated with a mixture of fluorescent beads coated with capture antibodies specific for human IL-2, IL-4, IL-6, IL-10, TNF-α, IFN-γ, and IL-17A, followed by the addition of a PE-conjugated detection antibody. Samples were then washed and analyzed using a flow cytometer. BD FACSymphony™ A5 SE Cell Analyzer (BD Biosciences) equipped with the FACSDiva software (BD Biosciences) was used for flow cytometry. Standard curves were generated for each cytokine using provided recombinant standards. Cytokine concentrations were calculated using GraphPad Prism version 10.4.1 with a log-transformed standard curve fit. Comparative statistical analyses between groups were conducted using unpaired two-tailed t-tests with Welch’s correction to account for potential unequal variances. Statistical significance was set at p<0.05 (*), p < 0.01 (**).
Cell proliferation assay
T cells were stained with Cell Trace Violet (Thermofisher, C34557) per manufacturer instructions, then co-cultured with 1×10e4 target cells at 1:1 effector to target ratio. 96 hours later flow cytometry was performed for the presence of CTV dilution.
Cytotoxicity assays
Cytotoxic killing of suspension target cells was assessed using flow cytometry. Briefly, target cells were labelled with Cell Trace Violet, they were then seeded at 1×10e4/well in a 96 well plate. Following control T cells or GPNMB CAR T cells were added at the respective effector to target ratios. 48–72hrs following co-culture wells were harvested. Counting beads were spiked into these tubes and flow cytometry was performed to assess the presence of CTV+ cells normalized to beads. Average presence of tumor cells in control T cell conditions across biological replicates were used to normalize % target lysis. Cytotoxic killing of adherent target cells was assessed using a real-time, impedance-based assay with xCELLigence Real-Time Cell Analyzer System (ACEA Biosciences). Briefly, 1–5×10e4 target cells were seeded to the 96-well E-plate. Following seeding, control T cells or GPNMB CAR T cells were added at the respective effector to target ratios. Tumor killing was monitored every 20 min over the rest of the assay.
In vivo models
Human monoblastic leukemia in NSG mice.
Female NSG mice received 1×10e5 U937 cells via intravenous injection. 2 days following tumor injection mice received either 5×10e6 GPNMB CAR T cells or the relative number of untransduced T cells. Body weight was monitored weekly with routine assessment of the peripheral blood for T cell expansion. Animal wellbeing was scored twice weekly according to the IACUC body-condition system.
Human solid tumor xenografts in NSG mice.
Female NSG mice received subcutaneous implants of 5×10e5 SK-MEL-3 melanoma, or 5×10e5 SNU-503 colorectal cancer cells resuspended in 100 μL of a 1 : 1 mixture of matrigel (Corning) and 1x DPBS (Gibco). When tumors reached approximately 100 mm^3 (= day 22), animals were randomized into treatment groups. Mice were either treated with intravenous tail-vein injections of 4×10e6 GPNMB-specific CAR T cells, or with the relative amount of untransduced T cells as control, diluted in 200uL of 1x DPBS (Gibco). Tumor growth and body weight were recorded at least weekly with digital calipers and a precision scale, respectively. Body weight values were plotted relative to the day −1 baseline (= day prior to T cell infusion). In a subset of animals, peripheral blood was collected for flow cytometry and serum collection at day 21 following T cell infusion (for details see below). Animal wellbeing was scored twice weekly according to the IACUC body-condition system.
Tumor growth assessment by caliper.
Subcutaneous tumors were measured with a digital caliper, recording the longest diameter (length) and the perpendicular diameter (width). Volumes were estimated with the modified ellipsoid equation . Resulting values were plotted relative to the baseline measurement obtained on day −1 (= day prior to T cell infusion) in all main and supplementary figures.
T cell counts in peripheral blood.
Whole blood was collected by retroorbital bleeding into EDTA tubes. Samples were stained and absolute cell numbers were quantified with BD TruCount tubes (BD Biosciences) per manufacturer’s instructions. Brilliant Violet 605 anti-human CD45 antibody was used to detect human T cells in samples (1/80 dilution; clone 2D1; CAT: 368524; BioLegend). BD FACSymphony™ A5 SE Cell Analyzer (BD Biosciences) equipped with the FACSDiva software (BD Biosciences) was used for flow cytometry. Data were analyzed in FlowJo v10 (BD Biosciences).
Cytokine analyses from mouse serum.
To obtain mouse serum, blood was collected via retro-orbital bleeding in Eppendorf tubes, left at room temperature for approximately 20 minutes, and centrifuged to remove clots. Serum samples were frozen, and thawed at a later time point for cytokine analyses using the BD Human Th1/Th2/Th17 CBA Kit (BD Biosciences) as per the manufacturer’s instructions. BD FACSymphony™ A5 SE Cell Analyzer (BD Biosciences) equipped with the FACSDiva software (BD Biosciences) was used for flow cytometry. Data were analyzed in FlowJo v10 using the FlowJo CBA plug-in (BD Biosciences).
Quantification and Statistical Analysis
Statistics were performed in Prism 10 (GraphPad Software). Details of the test applied are provided in each individual figure legends. Survival curves were generated by the Kaplan–Meier method and compared with the log-rank (Mantel–Cox) test. Mann–Whitney U tests compared two groups. Ordinary two-way ANOVA was used for pairwise multiple comparisons. Significance thresholds: ns > 0.05; p ≤ 0.05; *p ≤ 0.01; **p ≤ 0.001; ***p ≤ 0.0001; ****p ≤ 0.00001.
For quantification of GPNMB IHC on tissue micro-arrays, a blinded reviewer gave a score of 1–5 for staining intensity.
Figures were prepared in Python version 3.10.16 using matplotlib version 3.10.1 and seaborn version 0.13.2, Prism version 10 (GraphPad Software), FlowJo version 10, Adobe Illustrator (Adobe), and created with BioRender.com.
Supplementary Material
Key Resource Table.
| REAGENT or RESOURCE | SOURCE | IDENTIFIER |
|---|---|---|
| Antibodies | ||
| Human Osteoactivin/GPNMB Antibody | R&D Systems | AF2550 |
| PE anti-human erbB3/HER-3 Antibody | BioLegend | 1B4C3 |
| Human IGF-IR/IGF1R PE-conjugated Antibody | R&D Systems | 33255 |
| Mouse F(ab)2 IgG (H+L) PE-conjugated Antibody | R&D Systems | F0102B |
| Donkey anti-Goat IgG (H+L) Secondary Antibody, PE | Invitrogen | PA1-29953 |
| Goat anti-Mouse IgG (H+L) Highly Cross-Adsorbed Secondary Antibody, Alexa Fluor™ Plus 555 | Invitrogen | A32727 |
| PE Mouse IgG2a, κ Isotype Ctrl (FC) Antibody | BioLegend | MOPC-173 |
| GPNMB (E4D7P) XP® Rabbit mAb | Cell Signaling Technology | 38313 |
| Biotin-SP (long spacer) AffiniPure® F(ab')2 Fragment Goat Anti-Human IgG, F(ab')2 fragment specific | Jackson ImmunoResearch Labs | AB_2337634 |
| PE Streptavidin | BioLegend | 405204 |
| Brilliant Violet 605 anti-human CD45 antibody | BioLegend | 368524 |
| Biological samples | ||
| High-density multiple organs tumor with normal tissue microarray | tissuearray.com | MC5004b |
| Multi-organ (20 types from 9 systems) normal tissue microarray | tissuearray.com | BN1021a |
| T lymphocytes from human healthy donors | UPenn Human Immunology Core | ND607, (female) ND657, (female) |
| Deposited data | ||
| Skin cancer dataset | Biermann et al. (2022) | GEO: GSE185386 |
| Skin cancer dataset | Ji et al. (2020) | GEO: GSE144240 |
| Skin cancer dataset | Li et al. (2019) | GEO: GSE123139 |
| Skin cancer dataset | Yost et al. (2019) | GEO: GSE123814 |
| Normal Skin dataset | The Tabula Sapiens Consortium (2022) | GEO: GSE201333 |
| GTEX skin sun exposed lower leg | The Adult Genotype Tissue Expression (GTEx) Project | GTEx V10 |
| TCGA Bulk Data | National Cancer Institute GDC Data Portal | https://portal.gdc.cancer.gov/analysis_page?app= |
| Experimental models: Cell lines | ||
| Human (female) SK-MEL-5 | ATCC | HTB-70 |
| Human (female) SK-MEL-3 | ATCC | HTB-69 |
| Human (male) HT-144 | ATCC | HTB-63 |
| Human (male) SK-MEL-28 | ATCC | HTB-72 |
| Human (female) A-375 | ATCC | CRL-1619 |
| Human (male) U-937 | ATCC | CRL-1593.2 |
| Human (male) NALM-6 | ATCC | CRL-3273 |
| Human (male) THP-1 | ATCC | TIB-202 |
| Human (male) SW-156 | ATCC | CRL-2175 |
| Human (male) SNU-503 | AddexBio | C0009015 |
| Human (male) KLM-1 | RIKEN BioResource Research Center | RCB2138 |
| Experimental models: Organisms/strains | ||
| NOD/scid/IL2rγ-/- (NSG) | Jackson Laboratory | Cat# 5557 |
| Software and algorithms | ||
| Unique code | This paper | https://github.com/hayatlab/cart_llms |
| Python 3.10.16 | Python Software Foundation | https://www.python.org/ |
| OpenAI 1.66.5 | OpenAI | https://github.com/openai/openai-python |
| Google-generativeai 0.8.4 | Alphabet Inc. | https://github.com/google-gemini/deprecated-generative-ai-python |
| Anthropic 0.49.0 | Anthropic | https://github.com/anthropics/anthropic-sdk-python |
| Scanpy 1.10.3 | Wolf et. al. (2018) | https://github.com/scverse/scanpy |
| Pandas 2.2.3 | McKinney et. al. (2010) | https://github.com/pandas-dev/pandas |
| Numpy 2.2.3 | Harris et. al. (2020) | https://github.com/numpy/numpy |
| Requests 2.32.3 | N/A | https://github.com/psf/requests |
| R 4.2.3 | CRAN | https://cran.r-project.org/ |
| gtexr 0.1.0 | Warwick et. al. (2025) | https://github.com/ropensci/gtexr/ |
| TCGAbiolinks 2.34.1 | Colaprico et. al. (2016) | https://github.com/BioinformaticsFMRP/TCGAbiolinks |
| DESeq2 1.46.0 | Love et. al. (2014) | https://github.com/thelovelab/DESeq2 |
| FlowJo™ v10.8 Software | BD Life Sciences | https://www.flowjo.com |
Highlights:
LLMs used with biological data nominate tractable CAR T cell targets.
GPNMB is expressed on the surface across a broad range of tumor types.
GPNMB CAR T cells eliminate blood and solid tumors in vivo
Acknowledgments:
We thank Carolyn Shaw, Jian Li, and Yujie Ma for their logistical support. We also highlight Donna Gonzales and Sebastian Atoche for assistance with animal studies. D.J.B was supported by was supported by the Norman and Selma Kron Endowed Fellowship and the Centurion Foundation Innovation Fund. U.U. was supported by a Mildred-Scheel-Postdoctoral Fellowship of the German Cancer Aid. S.H. is funded by RWTH Aachen START (ID 692308), CRU344, and a Leducq Immuno-Fib HF seed grant. This work was supported by CA248315 (Z.A.). This work was supported by the Centurion Foundation Innovation Fund, 1P01CA214278 and R01CA226983 (C.H.J.), and the Parker Institute for Cancer Immunotherapy.
Footnotes
Declaration of interests:
D.J.B., L.F., U.U, K.K.P., D.Z., N.W.E., J.M.G., W.Z., S.I.K., L.S., C.R., J.A.E., and Z.A. declare no competing interests.
P.C.R. is an inventor on patents and patent applications licensed to Kite Pharma and receives license revenue from such licenses.
R.M.Y., C.H.J. is an inventor on patents and/or patent applications licensed to Novartis Institutes of Biomedical Research and receives license revenue from such licenses. R.M.Y., C.H.J. is an inventor on patents and/or patent applications licensed to Kite Gilead and may receive revenue from such licenses.
S.H. is a co-founder and shareholder of Sequantrix GmbH and reports research funding from AskBio and Novo Nordisk.
C.H.J is a scientific founder of Tmunity Therapeutics, Bluewhale Bio and Capstan Therapeutics. C.H.J is a member of the scientific advisory boards of AC Immune, BluesphereBio, Cabaletta, Cartography, Cellares, Celldex, Danaher, Decheng, Kite Gilead, Verismo, ViTToria Bio, and WIRB-Copernicus.
Declaration of generative-AI and AI-assisted technologies in the writing process:
Authors acknowledge the careful use of large language models in the preparation of this manuscript. The authors have reviewed and edited this content as needed and take responsibility for the content of this publication.
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References:
- 1.Baker DJ, Levine BL, and June CH (2025/March/01). Assessing the oncogenic risk: the long-term safety of autologous chimeric antigen receptor T cells. The Lancet 405. 10.1016/S0140-6736(25)00039-X. [DOI] [PubMed] [Google Scholar]
- 2.Ruella M, Korell F, Porazzi P, Maus MV, Ruella M, Korell F, Porazzi P, and Maus MV (2023-October-31). Mechanisms of resistance to chimeric antigen receptor-T cells in haematological malignancies. Nature Reviews Drug Discovery 2023 22:12 22. 10.1038/s41573-023-00807-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Baker DJ, Arany Z, Baur JA, Epstein JA, June CH, Baker DJ, Arany Z, Baur JA, Epstein JA, and June CH (2023-07-26). CAR T therapy beyond cancer: the evolution of a living drug. Nature 2023 619:7971 619. 10.1038/s41586-023-06243-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Brudno JN, Maus MV, and Hinrichs CS (2024/December/10). CAR T Cells and T-Cell Therapies for Cancer. JAMA 332. 10.1001/jama.2024.19462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Müller F, Taubmann J, Bucci L, Wilhelm A, Bergmann C, Völkl S, Aigner M, Rothe T, Minopoulou I, Tur C, et al. (2024-February-22). CD19 CAR T-Cell Therapy in Autoimmune Disease — A Case Series with Follow-up. New England Journal of Medicine 390. 10.1056/NEJMoa2308917. [DOI] [PubMed] [Google Scholar]
- 6.Baker DJ, and June CH (2022/November/23). CAR T therapy extends its reach to autoimmune diseases. Cell 185. 10.1016/j.cell.2022.10.026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Braun DA, and Wu CJ (March/April 2017). Antigen Discovery and Therapeutic Targeting in Hematologic Malignancies. The Cancer Journal 23. 10.1097/PPO.0000000000000257. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Flugel CL, Majzner RG, Krenciute G, Dotti G, Riddell SR, Wagner DL, Abou-el-Enein M, Flugel CL, Majzner RG, Krenciute G, et al. (2022-November-23). Overcoming on-target, off-tumour toxicity of CAR T cell therapy for solid tumours. Nature Reviews Clinical Oncology 2022 20:1 20. 10.1038/s41571-022-00704-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Gao S, Fang A, Huang Y, Giunchiglia V, Noori A, Schwarz JR, Ektefaie Y, Kondic J, and Zitnik M (2024/October/31). Empowering biomedical discovery with AI agents. Cell 187. 10.1016/j.cell.2024.09.022. [DOI] [PubMed] [Google Scholar]
- 10.Cui H, Tejada-Lapuerta A, Brbić M, Saez-Rodriguez J, Cristea S, Goodarzi H, Lotfollahi M, Theis FJ, Wang B, Cui H, et al. (2025-April-16). Towards multimodal foundation models in molecular cell biology. Nature 2025 640:8059 640. 10.1038/s41586-025-08710-y. [DOI] [PubMed] [Google Scholar]
- 11.Bzdok D, Thieme A, Levkovskyy O, Wren P, Ray T, and Reddy S (2024/March/06). Data science opportunities of large language models for neuroscience and biomedicine. Neuron 112. 10.1016/j.neuron.2024.01.016. [DOI] [PubMed] [Google Scholar]
- 12.Jilani S, Saco JD, Mugarza E, Pujol-Morcillo A, Chokry J, Ng C, Abril-Rodriguez G, Berger-Manerio D, Pant A, Hu J, et al. (2024-February-09). CAR-T cell therapy targeting surface expression of TYRP1 to treat cutaneous and rare melanoma subtypes. Nature Communications 2024 15:1 15. 10.1038/s41467-024-45221-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Hackett CS, Hirschhorn D, Tang MS, Purdon TJ, Marouf Y, Piersigilli A, Agaram NP, Liu C, Schad SE, Stanchina E.d., et al. (2024/September/19). TYRP1 directed CAR T cells control tumor progression in preclinical melanoma models. Molecular Therapy: Oncology 32. 10.1016/j.omton.2024.200862. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Ribas A, and Wolchok JD (2018-March-23). Cancer immunotherapy using checkpoint blockade. Science 359. 10.1126/science.aar4060. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Klebanoff CA, Chandran SS, Baker BM, Quezada SA, Ribas A, Klebanoff CA, Chandran SS, Baker BM, Quezada SA, and Ribas A (2023-October-27). T cell receptor therapeutics: immunological targeting of the intracellular cancer proteome. Nature Reviews Drug Discovery 2023 22:12 22. 10.1038/s41573-023-00809-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Biermann J, Melms JC, Amin AD, Wang Y, Caprio LA, Karz A, Tagore S, Barrera I, Ibarra-Arellano MA, Andreatta M, et al. (2022/July/07). Dissecting the treatment-naive ecosystem of human melanoma brain metastasis. Cell 185. 10.1016/j.cell.2022.06.007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Li H, Leun A.M.v.d., Yofe I, Lubling Y, Gelbard-Solodkin D, Akkooi A.C.J.v., Braber M.v.d., Rozeman EA, Haanen JBAG, Blank CU, et al. (2019/February/07). Dysfunctional CD8 T Cells Form a Proliferative, Dynamically Regulated Compartment within Human Melanoma. Cell 176. 10.1016/j.cell.2018.11.043. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ji AL, Rubin AJ, Thrane K, Jiang S, Reynolds DL, Meyers RM, Guo MG, George BM, Mollbrink A, Bergenstråhle J, et al. (2020/July/23). Multimodal Analysis of Composition and Spatial Architecture in Human Squamous Cell Carcinoma. Cell 182. 10.1016/j.cell.2020.05.039. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Yost KE, Satpathy AT, Wells DK, Qi Y, Wang C, Kageyama R, McNamara KL, Granja JM, Sarin KY, Brown RA, et al. (2019-July-29). Clonal replacement of tumor-specific T cells following PD-1 blockade. Nature Medicine 2019 25:8 25. 10.1038/s41591-019-0522-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Korsunsky I, Millard N, Fan J, Slowikowski K, Zhang F, Wei K, Baglaenko Y, Brenner M, Loh PR, and Raychaudhuri S (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nat Methods 16, 1289–1296. 10.1038/s41592-019-0619-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Consortium*, T.T.S., Jones RC, Karkanias J, Krasnow MA, Pisco AO, Quake SR, Salzman J, Yosef N, Bulthaup B, Brown P, et al. (2022-May-13). The Tabula Sapiens: A multiple-organ, single-cell transcriptomic atlas of humans. Science 376. 10.1126/science.abl4896. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Consortium TU, Bateman A, Martin M-J, Orchard S, Magrane M, Adesina A, Ahmad S, Bowler-Barnett EH, Bye-A-Jee H, Carpentier D, et al. (2025/January/06). UniProt: the Universal Protein Knowledgebase in 2025. Nucleic Acids Research 53. 10.1093/nar/gkae1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Uhlén M, Fagerberg L, Hallström BM, Lindskog C, Oksvold P, Mardinoglu A, Sivertsson Å, Kampf C, Sjöstedt E, Asplund A, et al. (2015-January-23). Tissue-based map of the human proteome. Science 347. 10.1126/science.1260419. [DOI] [PubMed] [Google Scholar]
- 24.Lautenbacher L, Samaras P, Muller J, Grafberger A, Shraideh M, Rank J, Fuchs ST, Schmidt TK, The M, Dallago C, et al. (2022/January/07). ProteomicsDB: toward a FAIR open-source resource for life-science research. Nucleic Acids Research 50. 10.1093/nar/gkab1026. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Buniello A, Suveges D, Cruz-Castillo C, Llinares, Manuel B, Cornu H, Lopez I, Tsukanov K, Roldán-Romero, Juan M, Mehta C, Fumis L, et al. (2025/January/06). Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Research 53. 10.1093/nar/gkae1128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Saade M, Araujo de Souza G, Scavone C, and Kinoshita PF (2021/May/12). Frontiers | The Role of GPNMB in Inflammation. Frontiers in Immunology 12. 10.3389/fimmu.2021.674739. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Lazaratos A-M, Annis MG, Siegel PM, Lazaratos A-M, Annis MG, and Siegel PM (2022-September-01). GPNMB: a potent inducer of immunosuppression in cancer. Oncogene 2022 41:41 41. 10.1038/s41388-022-02443-2. [DOI] [PubMed] [Google Scholar]
- 28.Bendell J, Saleh M, Rose AAN, Siegel PM, Hart L, Sirpal S, Jones S, Green J, Crowley E, Simantov R, et al. (2014-November-10). Phase I/II Study of the Antibody-Drug Conjugate Glembatumumab Vedotin in Patients With Locally Advanced or Metastatic Breast Cancer. Journal of Clinical Oncology 32. 10.1200/JCO.2013.52.5683. [DOI] [PubMed] [Google Scholar]
- 29.Hasanov M, Rioth MJ, Kendra K, Hernandez-Aya L, Joseph RW, Williamson S, Chandra S, Shirai K, Turner CD, Lewis K, et al. (2020-August-13). A Phase II Study of Glembatumumab Vedotin for Metastatic Uveal Melanoma. Cancers 2020, Vol. 12, Page 2270 12. 10.3390/cancers12082270. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Kopp LM, Malempati S, Krailo M, Gao Y, Buxton A, Weigel BJ, Hawthorne T, Crowley E, Moscow JA, Reid JM, et al. (2019/November/01). Phase II trial of the glycoprotein non-metastatic B-targeted antibody–drug conjugate, glembatumumab vedotin (CDX-011), in recurrent osteosarcoma AOST1521: A report from the Children’s Oncology Group. European Journal of Cancer 121. 10.1016/j.ejca.2019.08.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Ott PA, Hamid O, Pavlick AC, Kluger H, Kim KB, Boasberg PD, Simantov R, Crowley E, Green JA, Hawthorne T, et al. (2014-November-10). Phase I/II Study of the Antibody-Drug Conjugate Glembatumumab Vedotin in Patients With Advanced Melanoma. Journal of Clinical Oncology 32. 10.1200/JCO.2013.54.8115. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Ott PA, Pavlick AC, Johnson DB, Hart LL, Infante JR, Luke JJ, Lutzky J, Rothschild NE, Spitler LE, Cowey CL, et al. (2019/April/01). A phase 2 study of glembatumumab vedotin, an antibody-drug conjugate targeting glycoprotein NMB, in patients with advanced melanoma. Cancer 125. 10.1002/cncr.31892. [DOI] [PubMed] [Google Scholar]
- 33.Vahdat LT, Schmid P, Forero-Torres A, Blackwell K, Telli ML, Melisko M, Möbus V, Cortes J, Montero AJ, Ma C, et al. (2021-May-20). Glembatumumab vedotin for patients with metastatic, gpNMB overexpressing, triple-negative breast cancer (“METRIC”): a randomized multicenter study. npj Breast Cancer 2021 7:1 7. 10.1038/s41523-021-00244-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Yardley DA, Weaver R, Melisko ME, Saleh MN, Arena FP, Forero A, Cigler T, Stopeck A, Citrin D, Oliff I, et al. (2015-May-10). EMERGE: A Randomized Phase II Study of the Antibody-Drug Conjugate Glembatumumab Vedotin in Advanced Glycoprotein NMB–Expressing Breast Cancer. Journal of Clinical Oncology 33. 10.1200/JCO.2014.56.2959. [DOI] [PubMed] [Google Scholar]
- 35.Grupp SA, Kalos M, Barrett D, Aplenc R, Porter DL, Rheingold SR, Teachey DT, Chew A, Hauck B, Wright JF, et al. (2013-April-18). Chimeric Antigen Receptor–Modified T Cells for Acute Lymphoid Leukemia. New England Journal of Medicine 368. 10.1056/NEJMoa1215134. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Porter DL, Levine BL, Kalos M, Bagg A, and June CH (2011-August-25). Chimeric Antigen Receptor–Modified T Cells in Chronic Lymphoid Leukemia. New England Journal of Medicine 365. 10.1056/NEJMoa1103849. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Saland E, Boutzen H, Castellano R, Pouyet L, Griessinger E, Larrue C, Toni F.d., Scotland S, David M, Danet-Desnoyers G, et al. (2015. Mar 20). A robust and rapid xenograft model to assess efficacy of chemotherapeutic agents for human acute myeloid leukemia. Blood Cancer Journal 5. 10.1038/bcj.2015.19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Zabel BA, Wang Y, Lewén S, Berahovich RD, Penfold MET, Zhang P, Powers J, Summers BC, Miao Z, Zhao B, et al. (2009/September/01). Elucidation of CXCR7-Mediated Signaling Events and Inhibition of CXCR4-Mediated Tumor Cell Transendothelial Migration by CXCR7 Ligands. The Journal of Immunology 183. 10.4049/jimmunol.0900269. [DOI] [PubMed] [Google Scholar]
- 39.Sardiello M, Palmieri M, Ronza A.d., Medina DL, Valenza M, Gennarino VA, Malta CD, Donaudy F, Embrione V, Polishchuk RS, et al. (2009-July-24). A Gene Network Regulating Lysosomal Biogenesis and Function. Science 325. 10.1126/science.1174447. [DOI] [PubMed] [Google Scholar]
- 40.Loftus SK, Antonellis A, Matera I, Renaud G, Baxter LL, Reid D, Wolfsberg TG, Chen Y, Wang C, Program NCS, et al. (2008. Nov 1). Gpnmb is a Melanoblast-Expressed, MITF-Dependent Gene. Pigment cell & melanoma research 22. 10.1111/j.1755-148X.2008.00518.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Xia H, Green DR, Zou W, Xia H, Green DR, and Zou W (2021–03–23). Autophagy in tumour immunity and therapy. Nature Reviews Cancer 2021 21:5 21. 10.1038/s41568-021-00344-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Debnath J, Gammoh N, Ryan KM, Debnath J, Gammoh N, and Ryan KM (2023. –03–02). Autophagy and autophagy-related pathways in cancer. Nature Reviews Molecular Cell Biology 2023 24:8 24. 10.1038/s41580-023-00585-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Zemp FJ, Breckenridge Z, Song H, Gill GS, Liu H, Suh Y, Guignard L, Pyczek J, John C, Mah LK, et al. (2025-February-27). Development and first-in-human CAR T therapy against the pathognomonic MiT-fusion driven protein GPNMB. medRxiv. 10.1101/2025.02.26.24319604. [DOI] [Google Scholar]
- 44.Good CR, Aznar MA, Kuramitsu S, Samareh P, Agarwal S, Donahue G, Ishiyama K, Wellhausen N, Rennels AK, Ma Y, et al. (2021/December/09). An NK-like CAR T cell transition in CAR T cell dysfunction. Cell 184. 10.1016/j.cell.2021.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Data generated from skin cancer dataset integration in this study can be accessed via the Zenodo repository: https://zenodo.org/records/19071268. The source code for this AI-driven discovery framework is available at https://github.com/hayatlab/cart_llms. Additional information required to analyze the data reported in this paper is available from the lead contact upon request.
