Skip to main content
Computational and Structural Biotechnology Journal logoLink to Computational and Structural Biotechnology Journal
. 2025 Jun 11;27:2566–2573. doi: 10.1016/j.csbj.2025.06.018

Unsupervised cell line embedding using pairwise drug response correlation

Yutae Kim 1, Doheon Lee 1,⁎
PMCID: PMC12205321  PMID: 40586099

Abstract

Human cell line models are essential for understanding diseases and cellular functions. They are particularly emphasized in drug discovery because these models enable the systematic screening of chemical compounds and their effects. However, the heterogeneity in measurement techniques and the fragmented characterization of cell lines in chemical screening and omics data pose significant challenges to their optimal utilization. To address this, we introduce an unsupervised deep learning model based on contrastive learning that integrates heterogeneous drug response screening data into a unified cell line embedding. Utilizing the resulting embedding enhances the performance of drug-cell line-related downstream machine learning tasks to varying degrees. We used drug response data from 1,136 cell lines to train an embedding model and subsequently embedded 537 additional cell lines that were not included in the training, thereby completing the full set of 1,673 cancer cell lines from the Cancer Dependency Map (DepMap) that have corresponding gene expression data. We demonstrate that incorporating the embedding into various drug response-related tasks improves machine learning performance, including predicting drug synergy and drug response in cell lines. Furthermore, we applied SHapley additive explanations (SHAP) to identify genes with significant contributions to the embedding and found that these genes are strongly associated with drug resistance of various cancers and multiple types of cancer.

Keywords: Cancer cell lines, GDSC, PRISM, CTD2, Contrastive learning, Cell line embedding, Drug response

Graphical abstract

graphic file with name gr001.jpg

Highlights

  • •

    Created a cell line embedding model using the correlation of drug responses between two cells.

  • •

    The resulting embedded vector improves the performances of drug related machine learning tasks.

  • •

    Genes that contribute the most to the embedding are associated with different drug resistance genes.

1. Introduction

Cancer cell lines are indispensable preclinical tools in biological research, providing a reproducible platform to investigate the molecular mechanisms of different diseases, including cancers [1], [2]. This importance has driven extensive research efforts, making these cell lines the most studied and well-characterized human cell models [3]. A key focus of these models is measuring responses to genetic and chemical perturbations through large-scale screening, leading to the development of datasets such as Project SCORE [4], Connectivity Map [5], and Cancer Dependency Map [6] (DepMap).

The DepMap consortium has characterized human cell line models, from sequencing and profiling different cell line omics to quantifying gene perturbation results through RNA interference screens [6]. These modalities facilitated the characterization of in vitro cell models derived from different tumors through mutations [7], expressions [8], epigenetic changes [9], proteomics [10], and physiological profiling [11]. In particular, chemical and genetic perturbation datasets have led researchers to systematically evaluate the effects of therapeutics on cellular phenotypes, thereby informing the development of new drugs [1].

Due to the availability of large-scale perturbation data on cell lines, many researchers have applied various machine learning methodologies to improve the understanding of the mechanisms underlying human cell line model responses, be it in the form of drug sensitivity prediction [12], drug synergy prediction [13], or the discovery of therapeutic targets [4].

However, despite advances and successes in the application of machine learning, several limitations persist. Although the expansion of cell line perturbation datasets provides critical resources for machine learning, the vast number of potential drug combinations, along with the diverse modalities and methodologies used in generating these screening datasets, have resulted in high heterogeneity and fragmentation, complicating integration efforts [14].

Two widely accessible datasets that measure the sensitivity of specific drugs in specific cell lines, Profiling Relative Inhibition Simultaneously in Mixtures (PRISM) [15] and Genomics of Drug Sensitivity in Cancer (GDSC) [16], remain distinct and not easily interchangeable. This discrepancy arises from differences in data representation and screening methodologies: PRISM data are provided as the log-fold change in fluorescence measured by a luminex assay after specific drug and dosage combinations, whereas GDSC reports fitted dose-response curves from a different screening assay, with drug sensitivity quantified by the half-maximal inhibitory concentration (IC50) for each cell line and drug combination. Another example of this heterogeneity is the presence of four competing metrics that measure drug synergy: Highest Single Agent (HSA), Bliss Independence Model (Bliss), Zero Interaction Potency (ZIP) [17], and Loewe Additivity Model (Loewe). While commonly used, these metrics differ in their underlying assumptions and interpretations, leading to variability in synergy assessment across studies, as different research efforts adopt different metrics [18], [19].

In addition to the heterogeneous nature of screening data from cell line models, only 955 out of the 1,749 model cell lines listed in the Cell Model Passport [20] (https://cellmodelpassports.sanger.ac.uk/) have complete information on mutations, gene expression, copy number variation, and methylation. This incompleteness poses a significant challenge in developing methods that effectively integrate all four data types while leveraging the full scope of available screening datasets.

To integrate these heterogeneous datasets and enable predictive modeling despite missing cell line modalities, cell line representations were inferred from available data sources. Although the embedding process was not always explicitly defined, several studies applied their models to cell lines lacking direct perturbation screening data by using basal gene expression profiles as surrogate features [13], [21], [22], [23]. Some studies used linear dimensionality reduction [24] for this purpose, while others utilized deep learning approaches such as variational autoencoders (VAEs) [14]. In addition, multi-omics integration has been explored in some studies to improve the fidelity of cell line representations [18].

Another, more generalized approach is utilizing large knowledge graphs to represent cell lines. For example, Fernández-Torras et al. built a knowledge graph from various data sources including different biological entities including cell lines [25]. But due to the generalized nature in which cell lines are represented from just using one drug response dataset and gene expression, it potentially shows reduced precision in the context of fine-grained cellular prediction. Obtaining a comprehensive and fully integrated view remains challenging due to the fragmented nature of the available datasets.

To address the scarcity of cell line perturbation data resulting from heterogeneity and data compartmentalization, we present a deep learning-based cell line embedding method that integrates multiple sources of chemical perturbation data. Rather than directly training for higher performance in direct drug response predictions [13], [21], [22], [23], we create an informative embedding that performs well in drug related tasks. Once trained, this approach requires only gene expression values for inference. By directly comparing differences in chemical perturbation responses between cell lines and incorporating these differences into the embedding process, we generated an embedding that effectively captures both similarities and dissimilarities in drug response across cell lines. We demonstrate that incorporating this embedding improves performance across various predictive tasks.

2. Materials and methods

2.1. Gene expression and drug response data

Gene expression data (log2⁡(TPM+1) transformed) for 1,673 cancer cell lines (DepMap 24Q4 release) were downloaded from the DepMap portal [26]. Drug response data were obtained from the Cancer Therapeutics Response Portal (CTD2; accessed November 2024) [27] and the PRISM primary and secondary screens (DepMap 24Q2) [15]. GDSC1 and GDSC2 datasets (dated October 27, 2023) [16] were downloaded from the GDSC portal (https://www.cancerrxgene.org/). After removing duplicate genes based on gene IDs, 19,121 unique genes were retained for all downstream analyses.

2.2. Data processing

From each drug response database (Table 1), cell line pairs that share more than 3 drug responses in common were selected, and Pearson's correlation coefficient (PCC) was calculated. The number of drugs tested per cell line is not uniform across the database, so the number of drugs used for the calculation of the correlation value was also kept in a record. A total of 1,087,682 unique combinations of cell line pairs and drug response databases were used to train the embedding model.

Table 1.

Data used for embedding. Each database was parsed and used in a pairwise context. A total of 1,136 unique cell lines were used to embed the data.

Database Cell lines Raw Data Count Pair Data Count
CTD2 818 11,475,654 334,107
PRISM 901 4,149,998 399,578
PRISM secondary 477 5,107,832 113,526
GDSC 451 339,071 240,471

2.3. Embedding model

The architecture of the embedding model is a neural network model with a mixture of deep learning techniques: convolution layers, attention, bottlenecking, and attention pooling layers (Supp. Fig. A.4). This model reduces the gene expression data of 19,121 genes for each cell line to an embedded vector of size 1,024. The model was trained for 100 epochs but stopped early if the loss converged, so loss reduction per epoch was lower than 0.001. The learning rate was 0.0001 and optimized using the Adam optimizer [28]. The model was built using PyTorch (v2.1.0).

2.4. Convolution layers and blocks

The convolution layer performs linear operations by applying a kernel that slides over input features. A 1-D convolution layer can be represented as follows:

xt+1=∑i=0k−1wi⁎xt

Where xt+1 is the output of the weighted layer; xt is the input, and wi is the sliding windows' weights. In our model, the process of the 1-D convolution layer, following the activation function of the Gaussian error linear unit (GELU) and batch normalization, is grouped and labeled as a convolution block and utilized multiple times.

Conv Block=Conv1D(σ(BatchNorm(x)))

σ is the activation function.

2.5. Attention layer

The attention layer computes attention scores, which capture the relationship between the input elements. In the model, we utilized the self-attention layer to calculate the values.

headh=Attention(Q,K,V)=softmax(QKTdk)V,h∈{1,2,⋯,H}MultiHead(Q,K,V)=Concat(head1,⋯,headH)W0

2.6. Attention pooling

Attention pooling [29] aggregates information from features using learnable attention weights to focus on the most relevant parts, compared to average pooling or max pooling, which aggregates information by averaging or using the maximum value.

2.7. Bottlenecking

Bottlenecking temporarily reduces the dimensionality of feature representations in the intermediate layers of the model to improve the efficiency and learn more compact and informative representations.

2.8. Residual connection

Residual or skip connections add an additional connection that passes information directly across layers. A residual connection is represented as follows.

y=F(x)+x

F(x) is output of the deep learning layer; y is the output, and x is the input.

2.9. Contrastive learning in latent space

Contrastive learning enforces the similarity of the embedding vectors for similar inputs while pushing dissimilar inputs apart. The contrastive loss used for embedding is as follows.

L=∑i,jClog⁡(nij)⋅(yij⋅‖zi−zj‖2+(1−yij)⋅max(0,2−‖zi−zj‖2))

C is all possible combinations of two different cell lines i,j in a given database. yij is the PCC of the overall drug responses between cell line i,j. nij is the number of data points used to calculate the PCC between cell lines i,j. zi,zj is the model output embedding of cell line i,j.

2.10. Downstream task scheme

To evaluate the utility of the cell line embedding vector, we devised three different downstream tasks that predict three different values associated with the responses of a cell line to drugs. Deep learning techniques such as convolutional layers, attention mechanisms, attention-based pooling, and bottleneck architectures with residual connections were also incorporated into these models (Supp. Fig. A.4). For each downstream task, the embedding vectors were provided as additional inputs alongside task-specific features and concatenated with the drug SMILES-derived feature vectors and processed gene expression to form the complete input for each model. To prevent data leakage, we trained task-specific embeddings by computing the PCC between cell lines after excluding drug perturbation data associated with that task. These embeddings were optimized using the Adam optimizer [28] with a learning rate of 0.0001.

The model was initially trained on drug response correlations derived from 1,136 cell lines. To expand coverage for downstream tasks, we used the pretrained model to embed an additional 537 cell lines that lacked drug screening data from the four training datasets but had available gene expression profiles, resulting in a total of 1,673 embedded cell lines.

The models were evaluated using a five-fold cross-validation scheme. For each fold, we used embedding vectors generated from five independent runs of the embedding model with different random weight initializations, resulting in 25 evaluation sets per task. All models were trained for 100 epochs using the Adam optimizer [28] with a learning rate of 0.0001. For each training run, the model weights corresponding to the lowest validation loss were selected for final evaluation.

2.11. Drug representations

Simplified molecular input line entry system (SMILES) string was used to represent drugs in downstream tasks, downloaded using PubChemPy (v1.04; https://github.com/mcs07/PubChemPy) from PubChem [30]. TheSMILES string was then processed to float tensors using ChemBERTa-77M-MTR [31], with pretrained weights from HuggingFace (https://huggingface.co/).

2.12. Downstream tasks

For the first task, we built a model that predicts the ZIP synergy score of two drugs in a cell line. Drugcomb data (v1.5) [32] was downloaded from the DrugComb webpage (https://drugcomb.fimm.fi/). The downstream task model was trained to predict the ZIP score of two drugs in a given cell line. The task model takes the cell lines' gene expression, the embedded cell line vector, and the SMILES string of the two drugs as the input. Single drug response data within the database were not used. In the case of replicate or duplicate data, the values were normalized and removed following the procedures of Khili et al. [13]. Only samples with a standard deviation of the ZIP score that did not exceed 0.1 were kept, and the median value was used. Of the remaining 340,223 unique combinations of cell lines and two drugs with a median ZIP score of -0.726, a mean of 1.334, and a standard deviation of 12.769, we removed five outlier data points with ZIP scores over 1,000 for training stability.

For the second task, we utilized the same model architecture from the first task that predicts the synergy score of two drugs in a cell line. The data from Supplementary Table 6 in Nair et al. [19] were used. The downstream task model was trained to predict the “synergy_second_best” value from the table, the synergy score based on the Bliss model of the drug synergy calculation [19]. The model's input is the cell lines' gene expression, embedded vector, and SMILES string of the two drugs. Replicate data were averaged.

For the third task, we built a model that predicts the growth inhibition rate (GR) of a drug, its dosage, and a cell line. The gCSI data were downloaded from the gCSI webpage (http://research-pub.gene.com/gCSI_GRvalues2019/), which is an extension of the data published in Haverty et al. [33]. The downstream task model is trained to predict the GR [34] value, using cell line gene expression, embedded cell line vector, drug dosage, and the SMILES string of the drug. Data were averaged for replicates.

2.13. Embedding contribution analysis

SHapley Additive exPlanations (SHAP; v0.46.0) [35], using the GradientExplainer method [36], was used to estimate the contribution of each gene to the cell line embedding by measuring changes in the model gradients with respect to the input gene expression. For each target cell line, SHAP values were computed using 500 randomly selected background cell lines, and this process was repeated five times from the embedding done with all 4 datasets. The output of the SHAP computation was a matrix of the shape (Ngenes,Nembedding), representing the contribution of each gene to each embedding dimension. To obtain a single contribution score per gene, the absolute values of the contributions were summed across the embedding dimensions. These per-gene scores were then averaged across all cell lines to estimate the overall gene importance in the embedding process.

3. Results

3.1. Embedding results

Utilizing the DepMap gene expression dataset and a well-known set of drug perturbation data, including CTD2, GDSC, and PRISM, we generated an embedding vector from the gene expression profiles. This embedding was derived from the similarity of the cell lines' drug response patterns (Fig. 1a), enabling the construction of a latent representation that is independent of tissue specificity (Fig. 1b).

Fig. 1.

Fig. 1

a) Overview of the embedding process and the model. b) Uniform manifold approximation and projection (UMAP) visualization of the embedding [37]. c) Compartmentalized cell line data available for each dataset. A total of 1,136 unique cell lines that have all five transcriptome data, PRISM primary and secondary data, GDSC, and CTD2 data were used in the training of the model.

While gene expression is an informative and widely used feature in many drug-related machine learning models [13], [12], several studies have demonstrated the benefit of incorporating additional biological information such as protein–protein interaction (PPI) networks [38] and pathway-level data [39] to improve the predictive performance. These findings suggest that gene expression alone may be insufficient to fully capture the complex mechanisms underlying drug response. In our approach, we integrate drug response similarity across cell lines to account for inter-cell line variability, thereby enriching the representation beyond what gene expression alone can provide.

Our objective was to construct a unified embedding framework that integrates multiple modalities of perturbation data while also accommodating cell lines not explicitly present in the drug screening datasets (Fig. 1c). To achieve this, we adopted a contrastive learning approach rather than conventional dimensionality reduction techniques. This choice allowed us to guide the unsupervised embedding process using similarity and dissimilarity signals (Fig. 1a), enabling generalization to unseen cell lines with the available gene expression data.

3.2. Downstream tasks

To evaluate whether the embedding captures the drug response similarity in a way that benefits downstream drug-related tasks, we trained and compared two models for each task: one with and one without the embedding. Performance was assessed across three distinct regression tasks using separate datasets (Table 2): (1) a ZIP score prediction task using DrugComb [32] to assess drug synergy; (2) a synergy ratio prediction task based on the Bliss model in non-small cell lung cancer cell lines [19]; and (3) a growth rate (GR) inhibition prediction task that evaluates drug sensitivity across varying dosages [33]. Median loss of the training and test dataset can be seen from Supplementary Fig. B.5.

Table 2.

Data used in three different downstream tasks. The numbers listed are the count of the data used in the training and evaluation.

Downstream Task Dataset Cell lines Data Count
ZIP score prediction Drugcomb [32] 151 340,223
Synergy prediction Nair et al.[19] 72 295,680
GR prediction gCSI [33] 488 111,308

To measure the performance, we calculated the mean squared error (MSE), and the PCC between the predicted values from the trained models and the values listed in the database (Table 3). Including the embedding led to significant but varying performance improvement across different tasks (Fig. 2a). Specifically, the ZIP score prediction task in DrugComb showed a 0.4% increase, and synergy ratio prediction from the non-small cell lung cancer data showed an 11% increase in the PCC. Notably, the GR prediction task showed the most substantial gain at a 42% increase in the PCC.

Table 3.

Detailed results of the downstream tasks. All three tasks show statistically significant improved performance when the embedding vector is added to the tasks. Standard deviation is mentioned within the parentheses.

Task MSE PCC p-value
ZIP score 44.26 (±1.12) 0.854 (±0.00) 1.38e-04
Without Embedding 45.42 (±1.50) 0.850 (±0.00)



Synergy ratio 0.0320 (±0.0031) 0.59 (±0.048) 1.17e-12
Without Embedding 0.0329 (±0.0014) 0.48 (±0.068)



GR 0.019 (±0.000) 0.950 (±0.002) 2.01e-53
Without Embedding 0.137 (±0.000) 0.529 (±0.003)

Fig. 2.

Fig. 2

a) PCC of three downstream tasks with distinct datasets, with and without embedding. Pairwise t-test was done to calculate the p-values. b) The PCC performance of the top 10%, middle 10%, and bottom 10% performing cell lines based on the PCC values for each downstream task. c) The PCC between the reconstructed and database values of each dataset with regards to each cell line.

To further investigate how the embedding contributed to the performance, we grouped cell lines into the top 10%, middle 10%, and bottom 10% based on the PCC values between the predicted and true values for each downstream task (Fig. 2b). While the inclusion of the embedding improved model performance across all three percentile groups, the gains were particularly notable in the bottom 10% group, indicating that the embedding was especially beneficial for cell lines that were more difficult to predict.

To assess which cell lines exhibited a strong or weak predictive performance, we evaluated model outputs using fold-specific weights that excluded each target cell line during training for all three downstream tasks (Fig. 2c). In the ZIP score prediction task, skin melanoma cell lines, such as WM-115, demonstrated a particularly high predictive performance. Notably, skin-derived cell lines were significantly overrepresented among those in the top 10% of the PCC values (Fisher's exact test, p = 1.37×10−14). In contrast, the metastatic status of the cell lines did not show a statistically significant association with the predictive performance across any of the three tasks. For the synergy prediction task using non-small cell lung cancer cell lines, despite shared tissue origin, the predictive performance varied considerably among these cell lines, suggesting that additional molecular or phenotypic factors contribute to drug response variability.

3.3. Interpretation of the embedding

To understand which particular set of genes has the most impact on the embedding, we used the SHAP [35] algorithm to calculate the contribution of each gene to the embedding space. We found CNRIP1 to have the highest contribution to genes, followed by LRATD2 and ELL2 (Fig. 3a). Top contributing genes had associations with multiple specific cancers and drug resistances (Table 4), with nine out of the fifteen genes with literature evidence of drug resistances and five associated with multiple cancers. CD40 and KRT17 are listed in DGIdb [40] as druggable genes.

Fig. 3.

Fig. 3

a) Top 15 genes with the highest SHAP contributions to embedding. Average contribution value per gene is 5.477 × 10−3b) Gene set enrichment analysis of the top 50 contributing genes using Enrichr [42]. The top 50 genes show significant association with cancer associated genes and epithelial-to-mesenchymal transition (EMT).

Table 4.

Top 15 genes with the highest contributions to embedding with associations of different types of cancer and drug responses found in the literature. If association with cancer drug resistance and gene was found, associations with cancers were not listed.

Genes Related Literature
CNRIP1 different cancers [43], [44], [45]
LRATD2 platinum drug resistance [46]
ELL2 different cancers [47], [48], [49]
S100A3 cancer drug resistance [50], [51], [52]
KRT17 cancer drug resistance [53], [54]
POLR2L cisplatin resistance [55]
KIF7 cancer drug resistance [56], [57]
SLC50A1 breast cancer biomarker [58], [59]
STC2 cancer drug resistance [60], [61]
BBOF1 -
ARHGDIB different cancers [62], [63], [64]
ESRP1 cancer drug resistance [65], [66]
TRIM62 chemotherapy resistance [67]
CD40 cancer drug resistance [68], [69], [70]
PLBD1 Bruton tyrosine kinase inhibitor resistance [71]

Despite top fifteen genes having on-average twenty-fold more contributions than the contributions of an average gene (Fig. 3a), these genes are not the sole contributors, implying much more complex mechanisms of drug response which may include a total array of genes.

However, despite the roles of multiple genes in embedding, gene set enrichment analysis of the top 50 genes of contribution with the Molecular Signatures Database (MSigDB) hallmark gene set [41], demonstrated significant correlation (p-value < 0.005), with Epithelial to mesenchymal transition (EMT) reported to have a key impact in drug resistance and metastasis [10]. A detailed list can be found in C.5.

4. Discussion

Drug response prediction and drug synergy modeling are central to drug discovery and the study of drug–cell interactions and have been the focus of extensive modeling efforts in recent years [13], [23], [22], [72]. While various deep learning algorithms have been developed to extract informative representations of cell lines using diverse data types and increasingly complex model architectures, their progress has been constrained by limited sample sizes and fragmentation across datasets. In this study, we propose a novel approach that leverages multiple modalities within chemical perturbation datasets by distilling them into pairwise PCCs, capturing similarity relationships between cell lines. This framework enables a more tractable integration of heterogeneous data sources. The resulting embedding improved the performance across three distinct downstream tasks, despite being trained solely on similarity and dissimilarity signals derived from perturbation responses. To our knowledge, this is the first method to construct a dedicated embedding space for cell lines based on drug response characteristics.

Our approach only utilizes gene expression, yet we integrated heterogeneous drug response databases and demonstrated utility in improving the performances of downstream tasks and one-drug or two-drug response prediction of cell lines. Results from SHAP analyses suggest many or most genes affect the drug response, although some seem to contribute more than others. These contributing genes are supported by literature-based evidence linking them to resistance to specific drugs and a close association with the epithelial–mesenchymal transition (EMT) process. Be it in deciphering the mechanisms of drug response or as additional features usable in predicting drug responses, our cell line embeddings can potentially serve as a valuable tool for cancer research. Future studies should further the integration of additional databases for better and more accurate embeddings and application to clinical samples.

Ethics statement

This study is based on publicly available datasets. No human participants, patient data, or animal subjects were directly involved in this research. All data used in this study were obtained from sources that comply with ethical and legal guidelines for data sharing. No additional ethical approval was required.

CRediT authorship contribution statement

Yutae Kim: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Resources, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Doheon Lee: Supervision, Funding acquisition.

Declaration of Competing Interest

There is no conflict of interest.

Footnotes

Appendix

Supplementary material related to this article can be found online at https://doi.org/10.1016/j.csbj.2025.06.018.

Appendix. Supplementary material

The following is the Supplementary material related to this article.

MMC

Model architecture schematic, training diagnostics, and gene contribution table. This file includes a detailed schematic of the neural network architecture used in the study, training and validation set loss plots for each downstream task, and a ranked table of genes with the highest contribution to the model output based on feature attribution analysis.

mmc1.pdf (304.2KB, pdf)

Data availability

Embedded vectors, associated code, and full list of gene contributions can be downloaded from https://github.com/Atalasia/CellEmb.

References

  • 1.Sharma S.V., Haber D.A., Settleman J. Cell line-based platforms to evaluate the therapeutic efficacy of candidate anticancer agents. Nat Rev Cancer. 2010;10:241–253. doi: 10.1038/nrc2820. [DOI] [PubMed] [Google Scholar]
  • 2.Horvath P., Aulner N., Bickle M., Davies A.M., Nery E.D., Ebner D., et al. Screening out irrelevant cell-based models of disease. Nat Rev, Drug Discov. 2016;15:751–769. doi: 10.1038/nrd.2016.175. [DOI] [PubMed] [Google Scholar]
  • 3.Trastulla L., Noorbakhsh J., Vazquez F., McFarland J., Iorio F. Computational estimation of quality and clinical relevance of cancer cell lines. Mol Syst Biol. 2022;18 doi: 10.15252/msb.202211017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Behan F.M., Iorio F., Picco G., Gonçalves E., Beaver C.M., Migliardi G., et al. Prioritization of cancer therapeutic targets using crispr-cas9 screens. Nature. 2019;568:511–516. doi: 10.1038/s41586-019-1103-9. [DOI] [PubMed] [Google Scholar]
  • 5.Subramanian A., Narayan R., Corsello S.M., Peck D.D., Natoli T.E., Lu X., et al. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell. 2017;171:1437–1452.e17. doi: 10.1016/j.cell.2017.10.049. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Tsherniak A., Vazquez F., Montgomery P.G., Weir B.A., Kryukov G., Cowley G.S., et al. Defining a cancer dependency map. Cell. 2017;170:564–576.e16. doi: 10.1016/j.cell.2017.06.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Garnett M.J., Edelman E.J., Heidorn S.J., Greenman C.D., Dastur A., Lau K.W., et al. Systematic identification of genomic markers of drug sensitivity in cancer cells. Nature. 2012;483:570–575. doi: 10.1038/nature11005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Garcia-Alonso L., Iorio F., Matchan A., Fonseca N., Jaaks P., Peat G., et al. Transcription factor activities enhance markers of drug sensitivity in cancer. Cancer Res. 2018;78:769–780. doi: 10.1158/0008-5472.CAN-17-1679. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Iorio F., Knijnenburg T.A., Vis D.J., Bignell G.R., Menden M.P., Schubert M., et al. A landscape of pharmacogenomic interactions in cancer. Cell. 2016;166:740–754. doi: 10.1016/j.cell.2016.06.017. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Gonçalves E., Poulos R.C., Cai Z., Barthorpe S., Manda S.S., Lucas N., et al. Pan-cancer proteomic map of 949 human cell lines. Cancer Cell. 2022;40:835–849.e8. doi: 10.1016/j.ccell.2022.06.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Li H., Ning S., Ghandi M., Kryukov G.V., Gopal S., Deik A., et al. The landscape of cancer cell line metabolism. Nat Med. 2019;25:850–860. doi: 10.1038/s41591-019-0404-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Hsu Y.-C., Chiu Y.-C., Lu T.-P., Hsiao T.-H., Chen Y. Predicting drug response through tumor deconvolution by cancer cell lines. Patterns. 2024;5 doi: 10.1016/j.patter.2024.100949. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.El Khili M.R., Memon S.A., Emad A. Marsy: a multitask deep-learning framework for prediction of drug combination synergy scores. Bioinformatics. 2023;39 doi: 10.1093/bioinformatics/btad177. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Cai Z., Apolinário S., Baião A.R., Pacini C., Sousa M.D., Vinga S., et al. Synthetic augmentation of cancer cell line multi-omic datasets using unsupervised deep learning. Nat Commun. 2024;15 doi: 10.1038/s41467-024-54771-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Corsello S.M., Nagari R.T., Spangler R.D., Rossen J., Kocak M., Bryan J.G., et al. Discovering the anticancer potential of non-oncology drugs by systematic viability profiling. Nat Cancer. 2020;1:235–248. doi: 10.1038/s43018-019-0018-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Yang W., Soares J., Greninger P., Edelman E.J., Lightfoot H., Forbes S., et al. Genomics of drug sensitivity in cancer (gdsc): a resource for therapeutic biomarker discovery in cancer cells. Nucleic Acids Res. 2012;41:D955–D961. doi: 10.1093/nar/gks1111. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Yadav B., Wennerberg K., Aittokallio T., Tang J. Searching for drug synergy in complex dose–response landscapes using an interaction potency model. Comput Struct Biotechnol J. 2015;13:504–513. doi: 10.1016/j.csbj.2015.09.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Jaaks P., Coker E.A., Vis D.J., Edwards O., Carpenter E.F., Leto S.M., et al. Effective drug combinations in breast, colon and pancreatic cancer cells. Nature. 2022;603:166–173. doi: 10.1038/s41586-022-04437-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Nair N.U., Greninger P., Zhang X., Friedman A.A., Amzallag A., Cortez E., et al. A landscape of response to drug combinations in non-small cell lung cancer. Nat Commun. 2023;14 doi: 10.1038/s41467-023-39528-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.van der Meer D., Barthorpe S., Yang W., Lightfoot H., Hall C., Gilbert J., et al. Cell model passports-a hub for clinical, genetic and functional datasets of preclinical cancer models. Nucleic Acids Res. 2019;47:D923–D929. doi: 10.1093/nar/gky872. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Kuru H.I., Cicek A.E., Tastan O. From cell lines to cancer patients: personalized drug synergy prediction. Bioinformatics. 2024;40 doi: 10.1093/bioinformatics/btae134. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Park S., Lee H. Molecular data representation based on gene embeddings for cancer drug response prediction. Sci Rep. 2023;13 doi: 10.1038/s41598-023-49003-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Chawla S., Rockstroh A., Lehman M., Ratther E., Jain A., Anand A., et al. Gene expression based inference of cancer drug sensitivity. Nat Commun. 2022;13:5680. doi: 10.1038/s41467-022-33291-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Argelaguet R., Arnol D., Bredikhin D., Deloro Y., Velten B., Marioni J.C., et al. Mofa+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol. 2020;21 doi: 10.1186/s13059-020-02015-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Fernández-Torras A., Duran-Frigola M., Bertoni M., Locatelli M., Aloy P. Integrating and formatting biomedical data as pre-calculated knowledge graph embeddings in the bioteque. Nat Commun. 2022;13:5304. doi: 10.1038/s41467-022-33026-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Broad-DepMap Depmap 24q4 public. 2024. https://plus.figshare.com/articles/dataset/DepMap_24Q4_Public/27993248/1https://doi.org/10.25452/FIGSHARE.PLUS.27993248.V1 Available from:
  • 27.Seashore-Ludlow B., Rees M.G., Cheah J.H., Cokol M., Price E.V., Coletti M.E., et al. Harnessing connectivity in a large-scale small-molecule sensitivity dataset. Cancer Discov. 2015;5:1210–1223. doi: 10.1158/2159-8290.CD-15-0235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Kingma D.P., Ba J. Adam: a method for stochastic optimization. 2017. arXiv:1412.6980 Available from:
  • 29.Avsec v., Agarwal V., Visentin D., Ledsam J.R., Grabska-Barwinska A., Taylor K.R., et al. Effective gene expression prediction from sequence by integrating long-range interactions. Nat Methods. 2021;18:1196–1203. doi: 10.1038/s41592-021-01252-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Kim S., Chen J., Cheng T., Gindulyte A., He J., He S., et al. Pubchem 2023 update. Nucleic Acids Res. 2022;51:D1373–D1380. doi: 10.1093/nar/gkac956. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Chithrananda S., Grand G., Ramsundar B. Chemberta: large-scale self-supervised pretraining for molecular property prediction. 2020. arXiv:2010.09885 Available from:
  • 32.Zagidullin B., Aldahdooh J., Zheng S., Wang W., Wang Y., Saad J., et al. Drugcomb: an integrative cancer drug combination data portal. Nucleic Acids Res. 2019;47:W43–W51. doi: 10.1093/nar/gkz337. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Haverty P.M., Lin E., Tan J., Yu Y., Lam B., Lianoglou S., et al. Reproducible pharmacogenomic profiling of cancer cell line panels. Nature. 2016;533:333–337. doi: 10.1038/nature17987. [DOI] [PubMed] [Google Scholar]
  • 34.Hafner M., Niepel M., Sorger P.K. Alternative drug sensitivity metrics improve preclinical cancer pharmacogenomics. Nat Biotechnol. 2017;35:500–502. doi: 10.1038/nbt.3882. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Lundberg S., Lee S.-I. A unified approach to interpreting model predictions. 2017. arXiv:1705.07874 Available from:
  • 36.Sundararajan M., Taly A., Yan Q. Axiomatic attribution for deep networks. 2017. arXiv:1703.01365 Available from:
  • 37.McInnes L., Healy J., Saul N., Großberger L. Umap: uniform manifold approximation and projection. J Open Source Softw. 2018;3:861. doi: 10.21105/joss.00861. [DOI] [Google Scholar]
  • 38.Yang J., Li A., Li Y., Guo X., Wang M. A novel approach for drug response prediction in cancer cell lines via network representation learning. Bioinformatics. 2018;35:1527–1535. doi: 10.1093/bioinformatics/bty848. [DOI] [PubMed] [Google Scholar]
  • 39.Tang Y.-C., Gottlieb A. Explainable drug sensitivity prediction through cancer pathway enrichment. Sci Rep. 2021;11 doi: 10.1038/s41598-021-82612-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Cannon M., Stevenson J., Stahl K., Basu R., Coffman A., Kiwala S., et al. DGIdb 5.0: rebuilding the drug-gene interaction database for precision medicine and drug discovery platforms. Nucleic Acids Res. 2024;52:D1227–D1235. doi: 10.1093/nar/gkad1040. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Liberzon A., Birger C., Thorvaldsdóttir H., Ghandi M., Mesirov J.P., Tamayo P. The molecular signatures database hallmark gene set collection. Cell Syst. 2015;1:417–425. doi: 10.1016/j.cels.2015.12.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Kuleshov M.V., Jones M.R., Rouillard A.D., Fernandez N.F., Duan Q., Wang Z., et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 2016;44:W90–W97. doi: 10.1093/nar/gkw377. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Bethge N., Lothe R.A., Honne H., Andresen K., Trøen G., Eknæs M., et al. Colorectal cancer DNA methylation marker panel validated with high performance in non-Hodgkin lymphoma. Epigenetics. 2013;9:428–436. doi: 10.4161/epi.27554. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44.Zhang T., Cui G., Yao Y.-L., Wang Q.-C., Gu H.-G., Li X.-N., et al. Value of cnrip1 promoter methylation in colorectal cancer screening and prognosis assessment and its influence on the activity of cancer cells. Arch Med Sci. 2017;6:1281–1294. doi: 10.5114/aoms.2017.65829. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Sun X., Chen D., Jin Z., Chen T., Lin A., Jin H., et al. Genome-wide methylation and expression profiling identify methylation-associated genes in colorectal cancer. Epigenomics. 2019;12:19–36. doi: 10.2217/epi-2019-0133. [DOI] [PubMed] [Google Scholar]
  • 46.Zhang Y., Li Q., Yu S., Zhu C., Zhang Z., Cao H., et al. Long non-coding RNA fam84b-as promotes resistance of gastric cancer to platinum drugs through inhibition of fam84b expression. Biochem Biophys Res Commun. 2019;509:753–762. doi: 10.1016/j.bbrc.2018.12.177. [DOI] [PubMed] [Google Scholar]
  • 47.Zang Y., Pascal L.E., Zhou Y., Qiu X., Wei L., Ai J., et al. Ell2 regulates DNA non-homologous end joining (nhej) repair in prostate cancer cells. Cancer Lett. 2018;415:198–207. doi: 10.1016/j.canlet.2017.11.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Yang T., Jing Y., Dong J., Yu X., Zhong M., Pascal L.E., et al. Regulation of ell2 stability and polyubiquitination by eaf2 in prostate cancer cells. Prostate. 2018;78:1201–1212. doi: 10.1002/pros.23695. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49.Wang Z., Pascal L.E., Chandran U.R., Chaparala S., Lv S., Ding H., et al. Ell2 is required for the growth and survival of ar-negative prostate cancer cells. Cancer Manag Res. 2020;12:4411–4427. doi: 10.2147/cmar.s248854. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Gianni M., Terao M., Kurosaki M., Paroni G., Brunelli L., Pastorelli R., et al. S100a3 a partner protein regulating the stability/activity of rarα and pml-rarα in cellular models of breast/lung cancer and acute myeloid leukemia. Oncogene. 2018;38:2482–2500. doi: 10.1038/s41388-018-0599-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Świerczewska M., Klejewski A., Wojtowicz K., Brązert M., Iżycki D., Nowicki M., et al. New and old genes associated with primary and established responses to cisplatin and topotecan treatment in ovarian cancer cell lines. Molecules. 2017;22:1717. doi: 10.3390/molecules22101717. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Shao X., Chen Y., Wang W., Du W., Zhang X., Cai M., et al. Blockade of deubiquitinase yod1 degrades oncogenic pml/rarα and eradicates acute promyelocytic leukemia cells. Acta Pharm Sin B. 2022;12:1856–1870. doi: 10.1016/j.apsb.2021.10.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Li J., Chen Q., Deng Z., Chen X., Liu H., Tao Y., et al. Krt17 confers paclitaxel-induced resistance and migration to cervical cancer cells. Life Sci. 2019;224:255–262. doi: 10.1016/j.lfs.2019.03.065. [DOI] [PubMed] [Google Scholar]
  • 54.Wu L., Ding W., Wang X., Li X., Yang J. Interference krt17 reverses doxorubicin resistance in triple-negative breast cancer cells by wnt/β-catenin signaling pathway. Genes Genomics. 2023;45:1329–1338. doi: 10.1007/s13258-023-01437-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 55.Zhou D., Li X., Zhao H., Sun B., Liu A., Han X., et al. Combining multi-dimensional data to identify a key signature (gene and mirna) of cisplatin-resistant gastric cancer. J Cell Biochem. 2018;119:6997–7008. doi: 10.1002/jcb.26908. [DOI] [PubMed] [Google Scholar]
  • 56.Jenks A.D., Vyse S., Wong J.P., Kostaras E., Keller D., Burgoyne T., et al. Primary cilia mediate diverse kinase inhibitor resistance mechanisms in cancer. Cell Rep. 2018;23:3042–3055. doi: 10.1016/j.celrep.2018.05.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57.Zou J.X., Duan Z., Wang J., Sokolov A., Xu J., Chen C.Z., et al. Kinesin family deregulation coordinated by bromodomain protein ancca and histone methyltransferase mll for breast cancer cell growth, survival, and tamoxifen resistance. Mol Cancer Res. 2014;12:539–549. doi: 10.1158/1541-7786.MCR-13-0459. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Wang Y., Shu Y., Gu C., Fan Y. The novel sugar transporter slc50a1 as a potential serum-based diagnostic and prognostic biomarker for breast cancer. Cancer Manag Res. 2019;11:865–876. doi: 10.2147/cmar.s190591. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Zhang Q., Fang Y., She C., Zheng R., Hong C., Chen C., et al. Diagnostic and prognostic significance of slc50a1 expression in patients with primary early breast cancer. Exp Ther Med. 2022;24 doi: 10.3892/etm.2022.11553. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Cheng H., Wu Z., Wu C., Wang X., Liow S.S., Li Z., et al. Overcoming stc2 mediated drug resistance through drug and gene co -delivery by phb-pdmaema cationic polyester in liver cancer cells. Mater Sci Eng C. 2018;83:210–217. doi: 10.1016/j.msec.2017.08.075. [DOI] [PubMed] [Google Scholar]
  • 61.Liu Y., Tsai M., Wu S., Chang T., Tsai T., Gow C., et al. Acquired resistance to egfr tyrosine kinase inhibitors is mediated by the reactivation of stc2/jun/axl signaling in lung cancer. Int J Cancer. 2019;145:1609–1624. doi: 10.1002/ijc.32487. [DOI] [PubMed] [Google Scholar]
  • 62.Liu L., Cui J., Zhao Y., Liu X., Chen L., Xia Y., et al. Kdm6a-arhgdib axis blocks metastasis of bladder cancer by inhibiting rac1. Mol Cancer. 2021;20 doi: 10.1186/s12943-021-01369-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63.Zhu J., Tian Z., Li Y., Hua X., Zhang D., Li J., et al. Atg7 promotes bladder cancer invasion via autophagy-mediated increased arhgdib mrna stability. Adv Sci. 2019;6 doi: 10.1002/advs.201801927. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 64.Wang X., Bi X., Huang X., Wang B., Guo Q., Wu Z. Systematic investigation of biomarker-like role of arhgdib in breast cancer. Cancer Biomark. 2020;28:101–110. doi: 10.3233/CBM-190562. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Yu S., Wang M., Zhang H., Guo X., Qin R. Circ_0092367 inhibits emt and gemcitabine resistance in pancreatic cancer via regulating the mir-1206/esrp1 axis. Genes. 2021;12:1701. doi: 10.3390/genes12111701. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Zheng M., Niu Y., Bu J., Liang S., Zhang Z., Liu J., et al. Esrp1 regulates alternative splicing of carm1 to sensitize small cell lung cancer cells to chemotherapy by inhibiting tgf-β/smad signaling. Aging. 2021;13:3554–3572. doi: 10.18632/aging.202295. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Zhang M., Luo S. Gene expression profiling of epithelial ovarian cancer reveals key genes and pathways associated with chemotherapy resistance. Genet Mol Res. 2016;15 doi: 10.4238/gmr.15017496. [DOI] [PubMed] [Google Scholar]
  • 68.Memon D., Schoenfeld A.J., Ye D., Fromm G., Rizvi H., Zhang X., et al. Clinical and molecular features of acquired resistance to immunotherapy in non-small cell lung cancer. Cancer Cell. 2024;42:209–224.e9. doi: 10.1016/j.ccell.2023.12.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69.Li X., Wang S., Xie Y., Jiang H., Guo J., Wang Y., et al. Deacetylation induced nuclear condensation of hp1γ promotes multiple myeloma drug resistance. Nat Commun. 2023;14 doi: 10.1038/s41467-023-37013-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Qian Y., Galan-Cobo A., Guijarro I., Dang M., Molkentine D., Poteete A., et al. Mct4-dependent lactate secretion suppresses antitumor immunity in lkb1-deficient lung adenocarcinoma. Cancer Cell. 2023;41:1363–1380.e7. doi: 10.1016/j.ccell.2023.05.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 71.Ding Y., Huang K., Sun C., Liu Z., Zhu J., Jiao X., et al. A Bruton tyrosine kinase inhibitor-resistance gene signature predicts prognosis and identifies trip13 as a potential therapeutic target in diffuse large b-cell lymphoma. Sci Rep. 2024;14 doi: 10.1038/s41598-024-72121-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72.Partin A., Brettin T.S., Zhu Y., Narykov O., Clyde A., Overbeek J., et al. Deep learning methods for drug response prediction in cancer: predominant and emerging trends. Front Med (Lausanne) 2023;10 doi: 10.3389/fmed.2023.1086097. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

MMC

Model architecture schematic, training diagnostics, and gene contribution table. This file includes a detailed schematic of the neural network architecture used in the study, training and validation set loss plots for each downstream task, and a ranked table of genes with the highest contribution to the model output based on feature attribution analysis.

mmc1.pdf (304.2KB, pdf)

Data Availability Statement

Embedded vectors, associated code, and full list of gene contributions can be downloaded from https://github.com/Atalasia/CellEmb.


Articles from Computational and Structural Biotechnology Journal are provided here courtesy of AAAS Science Partner Journal Program

RESOURCES