Abstract
Biomarker research for cancer diagnosis and prognosis has rapidly expanded technologically and thematically, along with advancements in molecular diagnostics, liquid biopsy, immunotherapy, and artificial intelligence (AI)-based technologies. However, few studies have systematically structured these technological transitions and research framework changes using time-series analyses. Accordingly, this study treated cancer biomarker research as a large-scale biological knowledge system and aimed to examine the structural transitions and patterns of technological evolution through a temporal network-based analysis. Using the Web of Science database, 149,419 papers on cancer diagnosis and prognostic biomarkers were collected over three periods (2006–2011, 2012–2017, and 2018–2023). For each period, 500 core keywords were extracted using the weighted PageRank algorithm, and keyword co-occurrence networks were analyzed to identify temporal changes in the network structure indicators. Clustering was performed based on integrated keywords and evolutionary patterns were analyzed through research intensity and openness measurements. An attribute-based overlay analysis was conducted, along with analyses of keyword retention/turnover rates. Biomarker research has evolved from pathology-based diagnostic technology (2006–2011) and liquid biopsy-based noninvasive diagnostic technology (2012–2017) to AI- and immune-based precision medicine (2018–2023). Network structure analysis revealed an increased density and clustering of keyword connections over time, with expanded inter-cluster connections, indicating research topic convergence and increased structural complexity. Cluster 3 (Immune/AI-based Precision Medicine) exhibited the highest keyword turnover rate (0.545), emerging as a recent research focus, whereas Cluster 1 (Solid Tumor Pathology/Molecular Mechanisms) maintained stability with high keyword retention and a low change rate. This study introduces a temporal keyword co-occurrence network framework for detecting structural transitions in cancer biomarker research, providing an analytical perspective that complements conventional frequency-based bibliometric approaches. These findings offer quantitative insights into the technological complexity and convergent patterns of research development, reflected in increasing inter-cluster connectivity and integrated research trajectories in cancer diagnostic and prognostic biomarkers.
Supplementary Information
The online version contains supplementary material available at 10.1038/s41598-026-52746-7.
Keywords: Cancer biomarkers, Keyword co-occurrence network, Temporal network analysis, Technological transition, Research evolution
Subject terms: Biomarkers, Cancer, Computational biology and bioinformatics
Introduction
Cancer remains a significant global health challenge, necessitating new approaches for its early detection, precise diagnosis, and personalized treatment. With the spread of precision medicine, biomarkers have emerged as central pillars for personalized cancer diagnosis and prognosis1–3. Meanwhile, artificial intelligence (AI) has rapidly established itself as a transformative technological element in modern oncology, supporting sophisticated analyses and decision-making across various cancer treatment fields4–6. Advancements in AI have also been applied to the biomarker field for cancer diagnosis and treatment, quickly altering the technological landscape with predictive modeling, image analysis, and multiomics data integration technologies alongside immunotherapy and liquid biopsy1,7,8. Consequently, there is a growing need for research that provides a macroscopic perspective on the complex and rapidly evolving trends in biomarker research.
Recently, bibliometric analyses have been performed to identify changes in biomarker-related research architectures and technological trends. Bibliometrics provides insight into the distribution, changes, and development patterns of literature in a specific field by quantitatively analyzing scientific literature using mathematical and statistical techniques9,10. Li et al. quantitatively analyzed research trends over 20 years focusing on exosomes in prostate cancer (PCa)11. They collected 995 documents from the Web of Science Core Collection from 2003 to 2022, and analyzed the relationships and collaborations among countries/regions, institutions, authors, journals, references, and keywords. This confirmed the emergence of ‘Liquid biopsy,’ ‘microRNAs,’ ‘Tumor-derived exosomes,’ and ‘Identification’ as key keywords. Qiao et al. analyzed growth trends in biomarker research in the field of cancer immunotherapy9. They conducted a bibliometric analysis of 2,686 papers published from 1993 to 2023 and observed a significant increase in the number of papers related to cancer immunotherapy from 2015. Additionally, research topics such as ‘Programmed cell death-1 receptor (PD-1),’ ‘Programmed cell death-ligand 1 (PD-L1),’ ‘Cytotoxic T lymphocyte antigen-4 (CTLA-4),’ ‘Gene expression,’ ‘Tumor mutational burden (TMB),’ and ‘Gut microbiota’ were emerging. In particular, recent keywords related to ‘Artificial intelligence (AI)’ and ‘Machine learning’ were identified as key research topics. Additionally, Huang et al. analyzed the immune regulatory mechanisms within the tumor microenvironment, focusing on the CD39-CD73-eAdo/A2aR axis and its related trends in cancer immunotherapy research12. They analyzed the collaboration structures between countries, institutions, and authors as well as keyword networks and citation analysis, targeting 1,721 papers published from 2015 to 2024. By identifying the major institutions, high-frequency keywords, and highly cited papers in CD39-CD73-based cancer immunotherapy research, they identified research hotspots.
These prior studies are meaningful in that they quantitatively identified the developmental patterns of biomarker-related technologies at the national, institutional, and researcher levels, and confirmed recently emerging research topics. However, these are mostly limited to specific types of cancers or their mechanisms, making it difficult to view the overall technological landscape of biomarkers. Additionally, they failed to adequately explain transitions in technology over time or the evolution of technological aspects by research topic.
Rather than viewing keywords as simple descriptors of research topics, this study conceptualizes author keywords as biological knowledge entities whose co-occurrence networks reflect latent structures and transitions in biomarker research13. To address these limitations, Keyword Co-occurrence Network (KCN) analysis has been widely employed across various fields as a quantitative approach for modeling the relationships among research themes and exploring their structural evolution over time. Previous studies have demonstrated that KCN-based analyses effectively capture thematic reorganization, technological transitions, and knowledge diffusion across diverse scientific domains13–17.
You et al. empirically investigated the trends of convergence and reorganization of topics over time by analyzing a vast number of papers in the field of physics and comparing the network structures of keyword co-occurrence over different periods18. They quantitatively analyzed the roles played by specific topics within the research community by examining changes in cluster density and centrality indices. Catone et al. demonstrated that the field of social research methods is transitioning from a traditional empirical methodology focused on digital technology-based methodology by analyzing intertopic structures through a KCN analysis19. These studies demonstrate that keyword network analysis is effective for identifying the evolution of thematic structures and points of technological transition in various fields. In particular, Li et al. identified the structure of technological evolution and knowledge diffusion in the field of digital twin technology through a time-series analysis based on a KCN20. Araújo et al. compared KCNs across seven academic fields using the keyword ‘complexity,’ analyzing the evolution of thematic structures and the increasing levels of complexity over time within each field21. They used minimum-spanning trees to examine changes in structural distances among core keywords and patterns of centrality diffusion. Additionally, Kim et al. quantitatively analyzed technology evolution in the field of Resistive Random-Access Memory (ReRAM) by combining KCN analysis and time-series clustering techniques22. They extracted and refined keywords from 12,025 papers using a natural language processing-based tokenizer. Through Louvain clustering and PageRank centrality analysis, they identified major technology groups and conducted a time-series analysis of keyword changes within each cluster. These studies demonstrate that the KCN analysis is effective for exploring the structural evolution of research topics, points of technological transition, and patterns of knowledge diffusion across various academic fields.
Recently, an analysis method has been proposed that expands the KCN analysis to integrally consider not only co-occurrence relationships between keywords but also co-citation and co-reference relationships. Zhang et al. confirmed the evolutionary structure of knowledge by using this integrated analysis method23. In addition to KCN analysis, other well-established bibliometric methods for expressing the relationship between literature or keywords as a network include co-citation and bibliographic coupling (or co-reference)24,25. Co-citation involves analyzing the relationships between documents commonly cited by later researchers, allowing for the identification of core papers and knowledge bases in a relevant field. However, one disadvantage is that citations require time to accumulate, resulting in delays in reflection on the most recent research. Conversely, bibliographic coupling is based on the relationships between documents that share the same references. This swiftly reflects the latest research trends immediately after publication. However, there is a risk of distorting the network with excessive connections, particularly in cases involving documents with numerous citations, such as review articles or classical literature18,19. Although various network-based analysis methods exist, selecting an appropriate approach is crucial depending on the research objectives and data characteristics.
In this study, KCN analysis was adopted to examine structural changes in biomarker research, as it is relatively less affected by citation bias and is sensitive to emerging thematic relationships. Accordingly, a large-scale analysis was conducted to examine temporal structural changes and shifts in research themes across the entire biomarker research landscape without restricting the scope to specific diseases or targets.
Unlike conventional bibliometric studies that primarily focus on publication counts, citation patterns, or keyword frequency trends, this study applies a temporal keyword co-occurrence network framework to examine how the relational structure among research topics evolves over time. By analyzing structural network metrics such as network density, clustering coefficient, and inter-cluster connectivity, the proposed approach enables the detection of structural transitions in large-scale biomedical research fields. Through this approach, the analysis reveals three structurally distinct evolutionary patterns: consolidation of Cluster 1 (Solid Tumor Pathology/Molecular Mechanisms), sustained cross-cluster connectivity in Cluster 2 (Liquid Biopsy/Circulating Biomarkers), and rapid expansion with increasing internal cohesion in Cluster 3 (Immune/AI-based Precision Medicine). These structural patterns are difficult to identify using frequency-based bibliometric analyses alone. To support this analysis, a large-scale corpus of 149,419 publications related to cancer diagnosis and prognosis published between 2006 and 2023 was compiled from the Web of Science Core Collection. A KCN was constructed using author keywords that concisely represented the core research themes. Subsequently, the relationships between the research topics and the evolution of knowledge structures were examined through clustering, time-series comparisons, network structure analysis, and keyword change analysis.
Accordingly, this study aimed to uncover the latent structural transitions and convergent evolution patterns in cancer biomarker research by applying a temporal KCN-based bio data mining approach. Specifically, we examined the time-series changes in network structure, cluster-level growth and openness, and keyword retention and turnover to elucidate the differentiated evolutionary mechanisms across major biomarker domains. By moving beyond simple keyword frequency counts, this analysis provides empirical evidence for understanding the technological complexity of the biomarker field, where multiple diagnostic technologies emerge simultaneously and offers insights for future research planning and policymaking.
Methods
Data collection and keyword extraction
Data were collected and analyzed from papers published between 2006 and 2023 in the Web of Science Core Collection database. A keyword-based search query was used for data collection, a method widely applied to effectively select meaningful literature related to a research topic from large databases11,12. To analyze biomarker research related to cancer diagnosis and prognosis, a search was conducted within the Web of Science Core Collection. The query was applied to abstracts (AB) and author keywords (AK), with the publication period limited to 2006–2023. Document types were restricted to Articles and Reviews from journals indexed in the SCIE (Science Citation Index Expanded), SSCI (Social Sciences Citation Index), and AHCI (Arts & Humanities Citation Index) editions.
The specific search string used was as follows:
(AB = (cancer OR tumor* OR oncology) AND AB = (biomarker* OR “bio-marker*” OR “bio marker*”))
OR (AK = (cancer OR tumor* OR oncology) AND AK = (biomarker* OR “bio-marker*” OR “bio marker*”)).
A total of 149,419 papers were identified through the search and author keywords from their bibliographic information were extracted and refined for use in this analysis (102,950 unique keywords).
Time series-based keyword change analysis
To examine the evolution of research themes and changes in the thematic structure of biomarker research, this study analyzed the annual influx patterns and relational structures of author keywords representing research topics. First, to analyze the pattern of new keyword influx by year, unique author keywords identified by publication year and the number of new author keywords that had not appeared in previous years were counted. In addition, the proportion of newly introduced keywords for each year was calculated as the ratio of the number of new author keywords appearing in that year to the total number of unique author keywords observed in the same year. This proportion represents the relative share of new keywords within the yearly keyword vocabulary and is reported in Fig. 1B. This analysis was conducted to identify the timing of the emergence of new research topics and trends in technological change.
Fig. 1.
Yearly changes in keyword inflow. (A) Yearly changes in the total number of keywords and number of newly introduced keywords. (B) Yearly changes in the proportion of newly introduced keywords.
To elucidate the trends in technological change, the entire analysis period (2006–2023) was divided into six-year intervals, resulting in three distinct periods: P1, 2006–2011; P2, 2012–2017; and P3, 2018–2023. The six-year equal-interval segmentation was adopted as a methodological choice to ensure comparable observation windows across periods and to reduce the possibility that observed differences in network structure reflect unequal aggregation windows rather than underlying structural changes in the research landscape. This segmentation does not imply that technological transitions occur precisely at these boundary years. Rather, the intervals were used as a practical analytical framework for longitudinal comparison of network structures. To assess the robustness of the findings to alternative temporal boundaries, an additional sensitivity analysis was conducted using an alternative temporal segmentation with different boundary years (A1: 2006–2013; A2: 2014–2018; A3: 2019–2023). The same analytical pipeline was applied to this alternative segmentation, and the resulting network metrics are reported in Supplementary Table S1.
Each author keyword was assigned to the corresponding time period (P1–P3) according to the publication year of the paper in which it was published. If the same author keyword appeared in multiple papers, it was counted independently based on the period of each paper. Additionally, to explain thematic changes in biomarker technologies over time, core keywords representing technological content for each period (P1, P2, and P3) were extracted. A co-occurrence network was constructed using author keywords for each period, and the weighted PageRank algorithm was applied to derive the top 500 core keywords for each period. The threshold of 500 keywords was selected to balance two methodological objectives: capturing the dominant co-occurrence structure of the keyword network while maintaining comparability of network size across periods. The cumulative frequency proportion of the top 500 keywords (top 500) in the entire network was calculated to examine whether the derived core keywords represented the technological content of the corresponding period. To remove the noise caused by items with low keyword frequencies, the analysis was conducted on keywords that appeared 10 times or more among all keywords for each period. Consequently, the cumulative proportions of the top 500 keywords in PageRank were 90.15% for P1 (2006–2011), 77.52% for P2 (2012–2017), and 72.13% for P3 (2018–2023) (Table 1). The declining coverage across periods reflects increasing semantic dispersion of the keyword space over time, consistent with the expansion and diversification of biomarker research. Nevertheless, because the selected keywords are ranked by PageRank centrality and account for a substantial cumulative proportion of keyword occurrences, the top 500 keywords can capture the structurally central portion of each period’s network. To assess whether the main findings are sensitive to this threshold choice, a sensitivity analysis using alternative thresholds of 300, 500, and 800 keywords was conducted; results are reported in Supplementary Table S2. The analysis showed that the overall structural trends remained consistent across different thresholds. Therefore, the PageRank top 500 keywords for each period were used as the core keyword sets for the subsequent analyses. Using this core keyword set, comparative analyses of keyword retention, disappearance, and new entries across periods were conducted, and the relationships among keywords were summarized using a keyword overlap table across periods (Table 2).
Table 1.
Cumulative frequency proportion of the PageRank top 500 keywords.
| Period | Number of keywords with frequency ≥ 10 | PageRank top 500 coverage (%) |
|---|---|---|
| P1 | 756 | 90.15% |
| P2 | 1900 | 77.52% |
| P3 | 3908 | 72.13% |
Table 2.
Overlap of core keywords across the three temporal periods (P1–P3).
| Keyword group | Periods present | Number of keywords |
|---|---|---|
| Persistent across all periods | P1 + P2 + P3 | 294 |
| Shared between P1 and P2 only | P1 + P2 | 78 |
| Shared between P2 and P3 only | P2 + P3 | 81 |
| Shared between P1 and P3 only | P1 + P3 | 9 |
| Unique to P1 | P1 only | 119 |
| Unique to P2 | P2 only | 47 |
| Unique to P3 | P3 only | 116 |
| Total | - | 744 |
Note: Persistent keywords appear in all three periods (P1–P3). Pairwise shared keywords appear in two of the three periods, and unique keywords appear only within a single period.
Network visualization and attribute overlay
An overlay analysis was conducted to intuitively understand the relationships between the core keywords derived for each period and to visually verify attributes such as research concentration based on keywords, academic influence, novelty, and international collaboration. A subnetwork was constructed and visualized using the core keyword sets derived for each period, including only their co-occurrence relationships. This study utilized VOSviewer (version 1.6.20), developed by the Centre for Science and Technology Studies (CWTS) at Leiden University, The Netherlands26, to visualize the following three attributes using overlay visualization:
(1) Normalized Citation Count (ref_count_norm): The value obtained by averaging the citation counts of papers containing the keyword adjusted for citation variance by year of publication (citations of an individual paper/average citations of papers from the same year).
(2) Average Year of Publication (pubyear_mean): The average year of publication of papers containing the keyword (a measure of recency).
(3) Country Diversity (country_unique): The number of unique countries derived from the author affiliations of papers containing the keyword (an indicator of international collaboration). Multiple affiliations were tallied using the full count method.
Clustering and theme identification
In this study, cluster analysis was conducted to systematically distinguish the major technological areas of biomarker research and to consistently interpret the research structure by theme. When clustering is performed independently using the core keyword network for each period, a direct comparison between clusters across different periods is challenging because of the changes in node composition and algorithm variability. Therefore, to consistently interpret the core keyword structures of the three periods, the top 500 core keywords from P1 to P3 were integrated, resulting in 744 unique keywords after removing duplicates. These were used to perform clustering using the built-in clustering algorithm of VOSviewer (based on the Louvain method). The resolution parameter (γ) of VOSviewer was set to 0.75 to ensure that the structure was neither too fragmented nor excessively integrated for interpreting the research area. Within each derived cluster, the Total Link Strength (TLS) of each keyword was calculated, and the top 20 keywords based on TLS were selected as major keywords for each cluster. Subsequently, a large language model (ChatGPT-4o, version 2024-11-20) was used to derive the research topics that captured the core concepts of each cluster. The language model was used solely for semantic labeling and interpretative summarization of clusters without influencing the clustering process or network structure.
Analysis of network metrics and keyword retention and change rates
(1) Calculation of network metrics.
To quantitatively evaluate the changes in research structure over different periods, the following network metrics were calculated:
Number of edges: The scale of the network connections.
Total Link Strength (TLS): The sum of the co-occurrence frequency weights for all edges in the network, representing the overall strength of the connections between keywords. An increase in TLS reflects not only an increase in the number of connections, but also a rise in the frequency of co-occurrences between keywords, indicating the extent of substantive convergence and integration among research topics.
Network density: The actual number of edges divided by the maximum possible number of edges, indicating how closely the research topics are connected.
Average clustering coefficient: The average level of interconnectedness among neighboring keywords for each keyword, representing the degree of local cluster formation within the network (unweighted).
Inter-cluster edge ratio (entire network level): The proportion of edges connecting different clusters out of all edges in the network, indicating the degree of interdisciplinary research integration.
The aforementioned network metrics were calculated using Python NetworkX (version 3.4.2), with the results presented in Table 5. For clarity, the inter-cluster edge ratio in Table 5 denotes the proportion of edges connecting different clusters relative to the total number of edges in the entire network. In contrast, the inter-cluster edge ratio in Table 6 is a cluster-specific metric. It represents the proportion of a particular cluster’s connections that link to external clusters, out of the total edges associated with that cluster. Consequently, these two metrics differ in their calculation methods and interpretive focus: the former measures overall network integration, while the latter measures the external openness of an individual cluster.
Table 5.
Structural changes in the keyword co-occurrence networks by period.
| Metric | P1 (2006–2011) | P2 (2012–2017) | P3 (2018–2023) | Change (P1→P3) |
|---|---|---|---|---|
| Number of edges | 14,153 | 25,267 | 39,669 | + 180.3% |
| Total Link Strength (TLS) | 34,346 | 101,304 | 255,579 | + 644.1% |
| Network density | 0.113 | 0.203 | 0.318 | + 181.4% |
| Average clustering coefficient | 0.424 | 0.482 | 0.553 | + 30.3% |
| Inter-cluster edge ratio (%) | 35.1 | 42.2 | 51.7 | + 16.6%p |
Note: Calculations were performed using co-occurrence networks where the top 500 keywords based on PageRank in each period were selected as nodes. TLS (Total Link Strength) represents the sum of the connection strengths of all edges. The inter-cluster edge ratio is a network-level metric representing the proportion of edges connecting different clusters relative to the total number of edges.
Table 6.
Network characteristics and openness by cluster.
| Cluster | Metric | P1 (2006–2011) | P2 (2012–2017) | P3 (2018–2023) | Change (P1→P3) |
|---|---|---|---|---|---|
| Cluster 1 | Number of nodes | 371 | 339 | 291 | −21.6% |
| TLS per node | 38.6 | 100.0 | 174.6 | + 352.7% | |
| Inter-cluster edge ratio (%) | 36.4 | 43.7 | 56.6 | + 20.3%p | |
| Cluster 2 | Number of nodes | 100 | 99 | 95 | −5.0% |
| TLS per node | 31.8 | 78.1 | 178.1 | + 459.8% | |
| Inter-cluster edge ratio (%) | 77.0 | 79.8 | 81.9 | + 4.9%p | |
| Cluster 3 | Number of nodes | 29 | 62 | 114 | + 293.1% |
| TLS per node | 7.0 | 64.3 | 319.7 | + 4,489.9% | |
| Inter-cluster edge ratio (%) | 96.4 | 89.2 | 79.6 | −16.7%p |
Note: Number of nodes refers to the number of keywords within each cluster among the top 500 PageRank keywords in each period. The TLS per node is calculated by dividing the internal cluster connection strength (TLS) by the number of nodes, excluding keywords used for the literature search (cancer, tumor, oncology, and biomarker), and reflects the research intensity and cohesion within each cluster. The inter-cluster edge ratio (%) indicates the proportion of edges connected to other clusters among all edges involving the cluster, illustrating the openness and interdisciplinary links of each cluster.
(2) Indicators of internal connection strength and openness by cluster.
The following metrics were utilized to assess the research intensity and technological openness of each cluster, as presented in Table 6:
TLS per node: TLS within a cluster divided by the number of nodes in the cluster. This reflects the intensity and cohesion of research within the cluster. However, keywords used for paper searches (cancer, tumor, oncology, and biomarker) were excluded when counting the number of nodes because they did not represent the research topics of individual clusters. This approach aims to measure the substantive research intensity of each cluster more accurately.
Inter-cluster edge ratio (cluster level): The proportion of edges connected to other clusters out of all edges associated with a particular cluster. This indicates the openness of the cluster and the extent of its linkages with external technologies.
(3) Analysis of keyword retention/turnover rates and new/extinct keywords by cluster.
The turnover concept proposed by Bentley et al.27 was applied to quantitatively measure the shifts in keyword composition across clusters over time.
Keyword turnover rate: Defined as the proportion of newly introduced keywords in the current period that were absent in the preceding period. For instance, if two out of ten keywords in a cluster are newly introduced, the turnover rate is 20%. A higher turnover rate generally indicates more active technological transitions or research shifts within a specific subject area27,28.
Retention rate: Represents the proportion of keywords from the previous period that persist into the current period, serving as an indicator of thematic stability28.
To further identify specific research trends, the core keyword sets from P1 (2006–2011) and P3 (2018–2023) were compared. This comparative analysis allowed for the identification of newly emerged keywords and extinct keywords within each cluster.
Results
Research into biomarkers for cancer diagnosis and prognosis has expanded substantially in recent decades, accompanied by an increase in technological diversity. In this section, we quantitatively characterize this technological evolution using author keyword-based network analyses, focusing on temporal changes in research themes and structural patterns.
Yearly keyword inflow patterns
In this study, the entire analysis period (2006–2023, 18 years) was divided into three intervals of equal length to analyze the temporal changes: P1 (2006–2011), P2 (2012–2017), and P3 (2018–2023). The six-year equal-interval segmentation was adopted to ensure comparable observation windows across periods and to reduce the possibility that observed differences in network structure reflect unequal aggregation windows rather than underlying structural changes in the research landscape.
Analyzing annual patterns in keyword inflows enables the identification of emerging research trends and periods of technological transition. Figure 1 presents the quantitative analysis results of the yearly inflow patterns of author keywords from 2006 to 2023. The total number of keywords was defined as the number of unique author keywords that appeared in a given year. Newly introduced keywords were counted as the number of keywords that did not appear in any previous year.
As shown in Fig. 1A, the total number of keywords has consistently increased since 2006, with a particularly sharp increase observed after 2018. This pattern reflects the rapid expansion in the diversity of research topics during this period. The number of newly introduced keywords also generally increased, but the upward trend was more moderate than the overall growth in total keywords. Figure 1B shows that the share of newly introduced keywords within the yearly keyword vocabulary declined over the analysis period. This pattern reflects the faster expansion of the overall keyword vocabulary relative to the introduction of entirely new terms. In other words, although new keywords continue to appear, the total vocabulary of keywords grows even more rapidly as the literature expands. A temporary stabilization is observed around 2018–2021, after which the declining trend resumes. During this period, newly introduced keywords increased in absolute terms (Fig. 1A); however, because the total keyword vocabulary also expanded substantially, the proportional change remained limited. These results indicate that new research topics continued to emerge while the overall keyword vocabulary expanded rapidly as the field developed.
Changes and distribution of keywords
In this study, keyword analysis by period was conducted to identify the flow of technological change, as well as to examine the entry of new keywords and the persistence of existing technologies. To this end, the core keyword sets for each of the three intervals (P1, P2, and P3) were extracted and keywords that appeared only in a specific period and common keywords that appeared continuously throughout the entire period were identified (Table 2). The detailed contents of the top 15 keywords by PageRank for each interval are presented in Table 3. Keywords that appeared only during 2006–2011 can be interpreted as reflecting research themes that were prominent in earlier phases, but became less central in subsequent periods. Keywords that appeared exclusively during 2012–2017 correspond to technological areas that experienced active exploration and diffusion during that interval. In contrast, keywords that first appeared during 2018–2023 represent research topics that have emerged rapidly in more recent years.
Table 3.
Changes in the appearance of keywords by period (top 15 by PageRank in each interval).
| only_2006_2011 (P1) | only_2012_2017 (P2) | only_2018_2023 (P3) | common_all (P1 + P2+P3) |
common_2006_2011 and_2012_2017 (P1 + P2) |
common_2012_2017 and_2018_2023 (P2 + P3) |
common_2006_2011 and_2018_2023 (P1 + P3) |
|---|---|---|---|---|---|---|
| RT-PCR | MET | Immune cell infiltration | Biomarker | Chemoprevention | Long non-coding RNA | Dendritic cell |
| DNA adduct | iTRAQ | Machine learning | Breast cancer | Cetuximab | Liquid biopsy | Malignancy |
| SELDI-ToF-MS | Everolimus | Bioinformatic analysis | Cancer | Predictive marker | Pancreatic ductal adenocarcinoma | Endometrial carcinoma |
| Loss of heterozygosity | PET/CT | Radiomics | Prognosis | Matrix metalloproteinase | Programmed death-ligand 1 | Wnt |
| Comet assay | Microvesicles | Tumor mutation burden | Immunohistochemistry | Bcl-2 | Extracellular vesicle | Transcriptomics |
| Differential diagnosis | Vaccine | Pan-cancer | Prostate cancer | Gefitinib | Programmed cell death protein 1 | Disease |
| Barrett esophagus | In situ hybridization | Gut microbiota | Colorectal cancer | Differentiation | Precision medicine | Cancer detection |
| Neoplasia | Ipilimumab | WGCNA | Proteomics | Dysplasia | Circulating microRNA | Interleukin |
| Two-Dimensional Gel Electrophoresis (2D-GE) | c-MET | Deep learning | Apoptosis | FISH | Cell-free DNA | Relapse |
| Phosphorylation | Cancer vaccine | Competing endogenous RNA | Diagnosis | Erlotinib | The Cancer Genome Atlas | - |
| Medulloblastoma | PCA3 | Artificial intelligence | Lung cancer | Surviving | microRNA-21 | - |
| Aromatase inhibitor | Graphene | Ferroptosis | Inflammation | Proteome | Tumor infiltrate lymphocyte | - |
| p63 | Molecular target therapy | Tumor immune microenvironment | Tumor marker | Lipid peroxidation | Human epididymis protein 4 | - |
| Cancer marker | Tissue | Immune microenvironment | microRNA | Progesterone receptor | Expression | - |
| Hepatitis C virus | HIF-1α | COVID-19 | Ovarian cancer | LC-MS/MS | Heterogeneity | - |
During the entire period (2006–2023), 294 common keywords remained in the field of biomarkers, accounting for approximately 60% of the 500 core keywords identified in each period. This indicates a high degree of continuity in research topics in this field.
A total of 119 keywords appeared exclusively during 2006–2011, reflecting the relatively traditional direction of biomarker research. Keywords such as ‘RT-PCR,’ ‘DNA adduct,’ ‘Loss of heterozygosity,’ ‘Comet assay,’ and ‘Differential diagnosis’ indicated that biomarker technologies during this period were focused on traditional genetic diagnostic techniques and pathological approaches.
There were 47 keywords that appeared exclusively during 2012–2017, including ‘Ipilimumab,’ ‘PET/CT,’ ‘Cancer vaccine,’ ‘Microvesicles,’ and ‘Graphene.’ These reflect the transitional changes that occurred with the introduction of immunotherapy concepts and emergence of new diagnostic platforms. This period can be interpreted as the point at which convergent approaches, such as immune-based diagnostics and nanotechnology, began to gradually spread.
In total, 116 keywords appeared exclusively during 2018–2023, including terms such as ‘Immune cell infiltration,’ ‘Machine learning,’ ‘Radiomics,’ ‘Tumor mutation burden,’ ‘Deep learning,’ and ‘Artificial intelligence,’ reflecting a significant technological shift. This demonstrates that recent biomarker research is actively incorporating new analytical techniques and diagnostic approaches such as AI, multiomics, and the immune microenvironment.
Additionally, 78 keywords that were commonly found only between 2006 and 2011 and 2012 and 2017 likely represent topics that were prominent at the time but subsequently moved away from mainstream research. In contrast, the 81 keywords that appeared in both 2012–2017 and 2018–2023 remain central to current research. Keywords such as ‘Long non-coding RNA,’ ‘Programmed death-ligand 1,’ and ‘Liquid biopsy’ continue to play pivotal roles in cutting-edge immune biomarker and noninvasive diagnostic technology research.
Meanwhile, keywords that appeared exclusively in both 2006–2011 and 2018–2023 (e.g., ‘Dendritic cell,’ ‘Wnt,’ ‘Cancer detection,’ and ‘Interleukin’) likely represent reemerging topics that received relatively less attention after the initial research but regained prominence when combined with the latest analytical technologies and new clinical contexts.
This analysis demonstrates that biomarker research is characterized by both long-term continuity and stability as well as a structure in which new keywords are rapidly introduced in tandem with technological advancements over time. In particular, the recent spread of keywords related to AI, imaging-based analysis, and immunology is remarkable, indicating that biomarker technologies are rapidly shifting from traditional pathology-centered diagnostics to data-driven precision medicine.
Keyword network visualization and characteristic analysis
To comprehensively interpret the structural characteristics and technological changes in biomarker research from various perspectives, this study assessed research concentration, academic influence, recency, and degree of international collaboration based on keywords. To accomplish this, co-occurrence networks were constructed by integrating the core keyword sets from the three periods, and normalized citation count, average publication year, and number of collaborating countries were used as the attribute values. Visualization was performed as an overlay map using VOSviewer (version 1.6.20) (Fig. 2). The size of a node represents the frequency of occurrence of the corresponding keyword. In terms of color, higher values are indicated by more intense red hues.
Fig. 2.
Overlay visualization of the research network. (A) Overlay map based on normalized citation count. (B) Overlay map based on average publication year. (C) Overlay map based on country diversity.
The citation count serves as an indirect indicator of research influence, and the normalized citation count was used to account for annual variations in citation rates. Keywords such as ‘Microvesicles,’ ‘Nanoparticles,’ ‘Drug delivery,’ ‘COVID-19,’ ‘Exosome,’ ‘circRNA,’ ‘Non-coding RNA,’ ‘Programmed cell death protein 1,’ ‘Biosensor,’ ‘Tumor mutation burden,’ ‘Microbiome,’ ‘Artificial intelligence (AI),’ ‘Immune checkpoint,’ and ‘Immunotherapy’ were found to be research topics that possess high academic influence according to normalized citation counts, despite having a low frequency of occurrence (Fig. 2A). This may indicate research topics that rapidly garnered attention or became prominent during periods of technological transition.
The overlay map based on the average publication year illustrates the recency of the research (Fig. 2B). Research topics related to ‘Single cell RNA sequencing,’ ‘Pan-cancer,’ ‘Pyroptosis,’ ‘Immune cell infiltration,’ ‘Multiomics,’ ‘COVID-19,’ ‘Artificial intelligence (AI),’ ‘Deep learning,’ ‘Immune checkpoint inhibitor,’ ‘Prognostic model,’ ‘Tumor mutation burden,’ ‘Radiomics,’ ‘Precision medicine,’ ‘Tumor microenvironment,’ ‘Immunotherapy,’ ‘circRNA,’ ‘Microbiome,’ ‘Liquid biopsy,’ ‘Circulating tumor DNA,’ and ‘Bioinformatics’ have recently been actively emerging. This suggests that biomarker technology is being integrated with the latest advances, such as AI-based analysis, immune microenvironment, and precision medicine applications.
Country diversity refers to the examination of authors’ national affiliations in articles containing relevant keywords from the perspective of global collaboration. Research outcomes produced by a wider range of countries are considered to have higher potential for international dissemination. These keywords can be interpreted as representing topics that are the focus of shared scientific interests and strategic investments at the global level. ‘Pyroptosis,’ ‘Pan-cancer,’ ‘Immune cell infiltration,’ ‘Weighted Gene Co-expression Network Analysis (WGCNA),’ ‘Nomogram,’ ‘Non-small cell lung cancer,’ ‘Immune microenvironment,’ ‘Proliferation,’ ‘Lung adenocarcinoma,’ ‘Migration,’ ‘Circular RNA (circRNA),’ ‘The Cancer Genome Atlas,’ ‘Long non-coding RNA,’ ‘Bioinformatics,’ ‘Gastric cancer,’ ‘Overall survival,’ and ‘Prognostic biomarker’ exhibited high country diversity (Fig. 2C).
Analysis of network structure and cluster changes by period
Although the overlay map provides an attribute-focused analysis at the keyword level, clustering analysis was conducted to break down the overall research structure and examine trends across subject areas. Clustering groups thematically similar keywords by considering their connection strengths based on the co-occurrence frequency of keywords. This approach allows the identification of closely related research streams or subtopics within biomarker research. Clustering analysis was conducted using VOSviewer and three major clusters were identified. For each cluster, core keywords were extracted based on TLS, and the semantic context of each cluster was analyzed using a large language model (ChatGPT-4o) to determine the research topics (Table 4). The language model was used only to generate concise semantic labels for clusters based on representative keywords and did not influence the clustering procedure. Additional details of the prompt template, keyword inputs, and generated outputs are provided in Supplementary Section S1.
Table 4.
Major keywords and proposed topics by cluster.
| Cluster | Representative Keywords | Topic label (LLM-assisted) |
|---|---|---|
|
Cluster 1 (Red) |
Biomarker, Breast cancer, Colorectal cancer, Cancer, Hepatocellular carcinoma, Metastasis, Immunohistochemistry, Inflammation, Apoptosis, Epithelial-mesenchymal transition, Chemotherapy, DNA methylation, Proliferation | Molecular and Pathological Mechanisms of Solid Tumor Progression and Biomarkers |
|
Cluster 2 (Green) |
microRNA, Diagnosis, Prostate cancer, Lung cancer, Liquid biopsy, Exosome, Proteomics, Circulating tumor cells, Extracellular vesicles, Circulating tumor DNA, Metabolomics, Early detection, Cell-free DNA, Mass spectrometry | Liquid Biopsy and Circulating Biomarkers for Noninvasive Cancer Diagnosis |
|
Cluster 3 (Blue) |
Prognosis, Immunotherapy, Non-small cell lung cancer, Long non-coding RNA, Gastric cancer, Tumor microenvironment, Programmed death-ligand 1 (PD-L1), Immune checkpoint inhibitor, Immune cell infiltration, Bioinformatics, Prognostic biomarker, The Cancer Genome Atlas, Machine learning | AI- and Omics-based Prognostic and Immunotherapy Response Prediction |
Cluster 1 includes a large number of keywords related to various solid tumors, such as ‘Biomarker,’ ‘Breast cancer,’ ‘Colorectal cancer,’ ‘Cancer,’ and ‘Hepatocellular carcinoma.’ Additionally, keywords such as ‘Metastasis,’ ‘Immunohistochemistry,’ ‘Epithelial-mesenchymal transition,’ ‘Apoptosis,’ ‘DNA methylation,’ and ‘Proliferation,’ which are related to the pathological mechanisms of tumors, basic molecular biological research, and the mechanisms of cancer metastasis, are central to this cluster. In particular, DNA methylation is closely associated with ‘Apoptosis’, ‘Programmed cell death,’ and ‘Chemotherapy’29. This cluster represents research focused on elucidating the pathological and molecular characteristics of various solid tumors, such as breast, colorectal, and liver cancers, and identifying biomarkers.
Cluster 2 includes keywords such as ‘microRNA,’ ‘Liquid biopsy,’ ‘Exosome,’ ‘Circulating tumor cells (CTCs),’ ‘Cell-free DNA (cfDNA),’ ‘Early detection,’ and ‘Diagnosis.’ These represent liquid biopsy technologies and circulating biomarker techniques that utilize blood or body fluids. Additionally, keywords such as ‘Metabolomics,’ ‘Proteomics,’ and ‘Mass spectrometry’ refer to omics-based analytical technologies that are relevant and applicable to liquid biopsies30–32. In other words, this cluster focuses on research concerning cancer diagnosis and early detection technologies that utilize liquid biopsies and circulating biomarkers.
Cluster 3 consists of keywords such as ‘Immunotherapy,’ ‘Immune checkpoint inhibitor,’ ‘Programmed death-ligand 1 (PD-L1),’ ‘Immune cell infiltration,’ ‘Tumor microenvironment,’ ‘Bioinformatics,’ ‘Machine learning,’ and ‘Long non-coding RNA.’ This cluster encompasses research trends in immunotherapy and precision medicine. In particular, it appears to include AI-based data analysis technologies and genome analysis-based prognostic biomarker technologies, such as ‘The Cancer Genome Atlas,’ ‘Bioinformatics,’ ‘Prognostic biomarker,’ and ‘Machine learning.’ This aligns with recent technological trends, in which research on immune microenvironmental regulatory factors, omics-based integrative analysis, and the development of AI-based prognostic models are expanding into the biomarker field1,33.
Visual changes in network structure by period
To examine how the structural complexity and thematic diffusion of biomarker research manifest across clusters over time, KCNs (left) and density maps (right) were analyzed for the three periods (P1, P2, and P3) (Fig. 3). Overall, from P1 to P3, the network expanded outward from the center and became more densely interconnected.
Fig. 3.
Keyword co-occurrence networks (left) and density-based visualizations (right) by period (P1-P3). In the density map, yellow denotes regions of high keyword density and blue indicates low density. During P1 (2006–2011), the network exhibits a centralized structure. As the period progresses to P3 (2018–2023), keyword interconnections become more complex, and the spread into peripheral regions becomes more prominent, signifying technological diversification and convergence. Node colors represent thematic clusters: red for Cluster 1 (solid tumor pathology), green for Cluster 2 (liquid biopsy), and blue for Cluster 3 (immunology/AI).
In P1 (2006–2011), a relatively concentrated network structure is observed, with research primarily focused on the classification of cancer types (e.g., breast cancer and colorectal cancer) and pathological mechanism studies centered around ‘Biomarker.’ The density map also indicates that the research was conducted primarily around core keywords located at the center of the keyword network.
In P2 (2012–2017), the network structure expanded moderately with increased connectivity between clusters. During this period, research on noninvasive diagnostic technologies and circulating biomarker-based methods, such as ‘Liquid biopsy,’ ‘Exosome,’ ‘Circulating tumor cells (CTCs),’ ‘Cell-free DNA (cfDNA),’ ‘microRNA,’ and ‘Circulating tumor DNA (ctDNA),’ spread (lower left), and immune-related keywords such as ‘Immunotherapy’ and ‘Programmed death-ligand 1 (PD-L1)’ began to appear (lower right). Additionally, keywords such as ‘Long non-coding RNA’ and ‘The Cancer Genome Atlas (TCGA)’ emerged (upper right), highlighting the increasing diversity and technological differentiation of research. The density map also reveals newly emerging high-density clusters in peripheral areas outside the network core.
In P3 (2018–2023), the complexity of the network reached its highest point, with keywords related to immunotherapy, such as ‘Immunotherapy,’ ‘Immune checkpoint inhibitor,’ and ‘Tumor mutation burden,’ emerging as a new central area (lower right). In addition, keywords such as ‘Long non-coding RNA,’ ‘circRNA,’ ‘Competing endogenous RNA (ceRNA),’ ‘Hub gene,’ ‘Tumor microenvironment,’ ‘Immune cell infiltration,’ ‘Differentially expressed gene (DEG),’ ‘Bioinformatic analysis,’ and ‘Prognostic model’ appeared, illustrating that research on the regulation of the tumor immune microenvironment and genome-based prognosis prediction is being actively conducted (upper right). In addition, keywords such as ‘Artificial intelligence,’ ‘Machine learning,’ ‘Deep learning,’ ‘Next-generation sequencing (NGS),’ and ‘Radiomics’ were combined with the network core, confirming the rapid expansion of AI-based precision medicine research. The density map also reveals that numerous high-density clusters are distributed at the periphery, indicating that, while connections to established core topics are maintained, research subjects are simultaneously diversifying and becoming more sophisticated.
Quantitative analysis of network structure by period
To quantitatively verify the structural changes in the networks observed in Fig. 3, the structural metrics of the co-occurrence networks were examined for each period using the 500 core keywords as nodes (Table 5).
The number of network edges increased by 180.3% in P3 compared to P1 (from 14,153 to 39,669), and the TLS increased by 644.1% (from 34,346 to 255,579). This indicates not only an increase in the number of connections but also a substantial increase in connection strength (keyword co-occurrence frequency). This pattern is consistent with the significantly denser network structure observed in the P3 KCN (Fig. 3, left) compared to P1.
Network density increased by 181.4%, from 0.113 to 0.318, indicating that the proportion of actual connections among all possible connections increased. This demonstrates that biomarker research topics have become more closely integrated over time. Furthermore, the average clustering coefficient increased by 30.3%, from 0.424 to 0.553, indicating that the keywords were grouped more densely at the local level. This local cohesion is also evident in the density map (right) in Fig. 3: in P1, high-density areas (yellow) are concentrated in the central region, whereas in P3, these high-density regions expand into the peripheral areas of the network.
The most noteworthy indicator was the change in the inter-cluster edge ratio. In P1, only 35.1% of all the connections were between different clusters; however, this increased to 51.7% in P3, with more than half of the connections occurring between clusters. A chi-square test confirmed that the composition of inter- and intra-cluster edges differed significantly across periods (χ²(2) = 1335.63, p < 0.001), statistically supporting the observed increase in inter-cluster connectivity. This pattern is also visually evident in Fig. 3, where connections between clusters become more prominent in P3 compared with P1. These results indicate that interactions among research topics have intensified and that the boundaries between thematic areas have become increasingly permeable. Consequently, research on solid tumor pathology/basic studies (Cluster 1), liquid biopsy/circulating biomarker research (Cluster 2), and immune/AI-based precision medicine research (Cluster 3) appear increasingly interconnected rather than developing independently.
Analysis of research structure and keywords by cluster
As shown in the previous section, the overall network structure of biomarker research has evolved over time, with increasing density and stronger inter-cluster connections. In this section, the growth patterns and changes in the keyword composition of each cluster are analyzed to identify the differences between clusters.
Cluster-level analysis of network structure
To assess the scale and openness of the research network within each cluster, the number of nodes, TLS per node, and inter-cluster edge ratio were analyzed by cluster (Table 6). The number of nodes represents the diversity of the research topics included in the cluster, whereas the TLS per node indicates the average co-occurrence strength per keyword, reflecting both the degree of connection between topics and internal cohesion. The inter-cluster edge ratio refers to the proportion of a cluster’s total connections that link to other clusters and serves as an indicator to evaluate the openness and degree of technological convergence within the cluster.
Cluster 1 (Solid Tumor Pathology/Molecular Mechanisms) experienced a 21.6% decrease in the number of nodes from P1 to P3 (371→291); however, the TLS per node increased by 352.7% (38.6→174.6), indicating that research has become increasingly concentrated around specific core mechanisms. The marked increase in the inter-cluster edge ratio from 36.4% to 56.6% (+ 20.3%) indicates that a structure previously centered on pathological and molecular mechanisms is now rapidly converging with external technologies such as liquid biopsy and AI-based analysis.
Cluster 2 (Liquid Biopsy/Circulating Biomarkers) showed little change in the number of nodes (100→95); however, the TLS per node increased by 459.8% (31.8→178.1), indicating strengthened connectivity among topics and greater research intensity. The inter-cluster edge ratio remained consistently high, increasing from 77.0% to 81.9%, demonstrating that liquid biopsy technologies are closely associated with multidisciplinary techniques, such as genomics, proteomics, imaging, and AI-based analysis.
Cluster 3 (Immune/AI-based Precision Medicine) exhibits the most remarkable changes. The number of nodes increased by 293.1% (29→114) and the TLS per node increased by 4,489.9% (7.0→319.7), indicating substantial growth in both scale and internal cohesion. The inter-cluster edge ratio decreased from 96.4% to 79.6%, which does not signify a reduced openness. This suggests that this emerging technological domain, which initially relied predominantly on connections with external clusters, rapidly developed dense internal connections over time. Consequently, the relative proportion of external linkages declined as a cohesive internal knowledge structure was established.
Keyword retention and turnover rates by cluster
While Table 5 examines the structural openness of the network, this section analyzes the retention and turnover rates (proportion of newly introduced keywords) for each cluster over time to assess the stability or dynamism of cluster topics by period (Table 7). This analysis can serve not only to evaluate the persistence of topics within each cluster, but also as an indicator of the openness of the knowledge system and the rate of topic transitions27.
Table 7.
Keyword composition turnover rate by cluster.
| Cluster | Interval | Number of previous keywords | Number of current keywords | Number of retained keywords | Number of newly introduced keywords | Keyword retention rate |
Keyword turnover rate |
Average keyword retention rate | Average keyword turnover rate | |
|---|---|---|---|---|---|---|---|---|---|---|
| Cluster 1 | P1 → P2 | 371 | 339 | 281 | 58 | 0.757 | 0.171 | 0.742 | 0.163 | |
| P2 → P3 | 339 | 291 | 246 | 45 | 0.726 | 0.155 | ||||
| Cluster 2 | P1 → P2 | 100 | 99 | 65 | 34 | 0.650 | 0.343 | 0.694 | 0.288 | |
| P2 → P3 | 99 | 95 | 73 | 22 | 0.737 | 0.232 | ||||
| Cluster 3 | P1 → P2 | 29 | 62 | 26 | 36 | 0.897 | 0.581 | 0.900 | 0.545 | |
| P2 → P3 | 62 | 114 | 56 | 58 | 0.903 | 0.508 | ||||
Note: The keyword retention rate denotes the fraction of core keywords from the preceding period that remained in the current period’s core set. The keyword turnover rate represents the proportion of newly emerged keywords in the current period relative to the total number of core keywords in the same period, serving as a proxy for technological dynamism27. Statistical comparison of keyword dynamics across clusters was conducted using a chi-square test pooling both transition intervals (P1→P2 and P2→P3), confirming significant differences among all cluster pairs (Cluster 1 vs. 2: p_adj = 0.0003; Cluster 1 vs. 3: p_adj < 0.001; Cluster 2 vs. 3: p_adj < 0.001; Bonferroni correction). Consistent results were obtained when intervals were tested separately (P1→P2: χ²(2) = 51.11, p < 0.001; P2→P3: χ²(2) = 55.00, p < 0.001).
Cluster 1 exhibits a high average keyword retention rate (0.742) and a low average keyword turnover rate (0.163) across all intervals. These patterns indicate that pathology-based diagnostics and biomarker research focusing on solid tumors, such as breast, colorectal, and lung cancers, exhibit characteristics that are consistent with relatively mature and stable research structures.
Cluster 2 displays moderate keyword retention (average = 0.694) and turnover (average = 0.288) rates, indicating transitional characteristics in which gradual diffusion occurs. In fact, Cluster 2 consistently incorporated keywords such as ‘Exosome,’ ‘Circulating tumor cells (CTCs),’ ‘Liquid biopsy,’ ‘microRNA,’ ‘Circulating tumor DNA (ctDNA),’ and ‘Cell-free DNA (cfDNA).’
Cluster 3 exhibits the highest keyword turnover rate (average = 0.545), reflecting the most dynamic changes28. Differences in keyword dynamics across clusters were statistically significant (χ²(2) = 101.59, p < 0.001, Cramér’s V = 0.319; see Table 7 Note for details). Despite the high turnover rate, Cluster 3 also showed a high retention rate (0.9), likely reflecting the characteristics of a recently emerging cluster that was initially small but experienced a rapid influx of new keywords. Notably, as the period transitioned from P2 (2012–2017) to P3 (2018–2023), keywords such as ‘Immune checkpoint inhibitor,’ ‘Tumor mutation burden,’ ‘Tumor immune microenvironment,’ ‘Pan-cancer analysis,’ ‘Bioinformatic analysis,’ ‘Deep learning,’ and ‘Artificial Intelligence’ emerged rapidly, indicating the occurrence of remarkable technological transitions and thematic innovations in recent years.
Analysis of new and extinct keywords by cluster
To examine how certain technologies or thematic areas have been reduced or replaced, and to identify the directions of technological evolution, the top 10 newly emerged or extinct keywords (based on PageRank) were identified for each cluster by comparing P1 (2006–2011) and P3 (2018–2023) (Table 8).
Table 8.
Newly emerged and extinct keywords by cluster (P1 vs. P3).
| Cluster | Top 10 newly emerged keywords | Top 10 extinct keywords |
|---|---|---|
| Cluster 1 | Gut microbiota, Microbiome, COVID-19, Cancer associate fibroblast, Systematic review, Glycolysis, Heterogeneity, Microbiota, Microenvironment, Stemness | Chemoprevention, Cetuximab, Predictive marker, Matrix metalloproteinase, Bcl-2, Gefitinib, Differentiation, Dysplasia, RT-PCR, FISH |
| Cluster 2 | Liquid biopsy, Extracellular vesicle, Circulating tumor DNA, Pancreatic ductal adenocarcinoma, Cell-free DNA, Surface-enhanced Raman scattering, Diagnostic biomarker, Electrochemical biosensor, Point of care test, Circulating microRNA | SELDI-TOF-MS, Proteome, LC-MS/MS, Secretome, Two-dimensional gel electrophoresis, Glycoprotein, Osteopontin, Validation, Cancer marker |
| Cluster 3 | Long non-coding RNA, Programmed cell death-ligand 1, Immune checkpoint inhibitor, Immune cell infiltration, The Cancer Genome Atlas, circRNA, Machine learning, Bioinformatic analysis, Precision medicine, Programmed cell death protein 1 | Expression profile, Image analysis |
Keywords that disappeared from Cluster 1 include ‘Chemoprevention,’ ‘Cetuximab,’ ‘Predictive marker,’ ‘RT-PCR,’ and ‘FISH,’ which were important technologies in early cancer research but have recently been replaced by other techniques or have declined in significance. In contrast, the newly emerged keywords are primarily related to cancer biological mechanisms, the tumor microenvironment, and metabolism. In particular, ‘Gut microbiota’ and ‘Microbiome’ reflect the recent surge in research on the relationship between cancer and the intestinal microbiota. Additionally, ‘COVID-19’ appears to have gained prominence owing to its relevance to cancer research during the pandemic period, while ‘Glycolysis,’ ‘Heterogeneity,’ and ‘Microenvironment’ reflect the growing interest in the metabolic characteristics of cancer cells and the tumor microenvironment. In other words, whereas traditional research focused primarily on gene expression and protein marker-based pathological diagnosis, recent attention has shifted toward analyses that consider more complex biological environments, such as the tumor microenvironment, immune-microbiome interactions, and associations with infectious diseases.
In Cluster 2, classical protein analysis technique keywords such as ‘SELDI-TOF-MS,’ ‘Proteome,’ ‘LC-MS/MS,’ and ‘Two-dimensional gel electrophoresis’ disappeared, and the newly emerged keywords were ‘Liquid biopsy,’ ‘Circulating tumor DNA,’ ‘Cell-free DNA,’ and ‘Circulating microRNA,’ which are related to liquid biopsy and circulating biomarker technologies. This indicates that the central axis of cancer diagnostic technologies have shifted from traditional protein expression analysis to noninvasive diagnostic and monitoring techniques.
In Cluster 3, traditional technology keywords (‘Expression profile’ and ‘Image analysis’) have recently shown a slight decrease in importance, while immunotherapy technologies such as ‘Immune checkpoint inhibitor,’ ‘Programmed death-ligand 1 (PD-L1),’ and ‘Immune cell infiltration,’ as well as data-driven precision medicine technologies such as ‘Machine learning,’ ‘Bioinformatics,’ and ‘Precision medicine,’ are on the rise. Recent developments have centered on AI-based analyses and immunotherapy technologies, indicating that biomarker research is evolving beyond simple diagnosis to prognosis prediction, therapeutic response monitoring, and personalized treatment design.
Discussion
We performed a time-series-based keyword analysis and network clustering analysis of biomarker papers related to cancer diagnosis and prognosis published in the Web of Science database from 2006 to 2023. Using these methods, this study quantitatively analyzed the technological evolution and changes in the thematic structure of the field. In particular, by employing co-occurrence network analysis and clustering based on core keywords for each period as well as analyzing the rates of keyword change, this study examined the evolutionary trajectory of biomarker technologies and research trends from multiple perspectives.
The analysis revealed that biomarker research has generally maintained a high degree of thematic continuity while simultaneously incorporating new keywords and expanding its research scope in accordance with technological changes over time. In particular, since 2018, keywords related to AI, radiomics, advanced imaging analysis, and immunotherapy have become prominent, indicating that biomarker technologies have shifted from pathology-based diagnostics (2006–2011) to liquid biopsy-based diagnostics (2012–2017) and, more recently, to a focus on immune/AI-based precision medicine (2018–2023) (Table 3; Fig. 3).
Network structure analysis showed that these shifts in research themes are not just changes in subject matter, but are also accompanied by changes in how different research topics are interconnected (Fig. 3; Table 5). Over time, the number of connections between keywords increased quantitatively, and the strength of co-occurrence intensified rapidly (Table 5). This indicates that biomarker research is evolving not through the simple introduction of new keywords but through the substantive integration of new concepts with existing keywords. In particular, the increase in the proportion of inter-cluster connections demonstrates that traditional pathology-based diagnostics, liquid biopsy, AI, and immunotherapy research are not developing independently but are converging with other subject areas to form an integrated knowledge structure.
This distinction can be further understood by contrasting our findings with those obtained from a conventional frequency-based bibliometric approach. A frequency-based analysis would primarily capture the emergence and increasing prominence of individual keywords, such as those related to AI, machine learning, or immunotherapy. While such approaches highlight the growing importance of these topics, they do not directly reveal the structural mechanisms through which these topics become integrated with established domains such as pathology-based diagnostics or liquid biopsy.
In contrast, the temporal keyword co-occurrence network analysis employed in this study makes such structural mechanisms visible. For example, while a frequency-based approach would identify AI and immunotherapy as increasingly prominent topics, it would not capture the fact that the inter-cluster edge ratio increased from 35.1% to 51.7% across the study period, suggesting that these emerging domains are becoming increasingly interconnected with, rather than simply growing alongside, established research clusters. Furthermore, the framework reveals that this integration follows differentiated trajectories: consolidation in Cluster 1, sustained inter-cluster connections in Cluster 2, and rapid expansion in Cluster 3. These patterns of structural convergence would be difficult to identify using frequency-based approaches alone.
Distinct evolutionary patterns were also observed at the cluster level. Cluster 1 (Solid Tumor Pathology/Molecular Mechanisms) showed a gradual decrease in the number of research topics, whereas the connection strength between keywords increased (Table 6). In addition, owing to high keyword retention and low turnover rates (Table 7), the cluster exhibited the typical characteristics of mature technology, with research consolidating around relatively stable core concepts (Table 6). Simultaneously, increased connectivity with other clusters suggests that the research significance of traditional pathology-based biomarker studies is being reemphasized through their convergence with new technologies. Cluster 2 (Liquid Biopsy/Circulating Biomarkers) exhibited little change in the number of nodes, whereas the connection strength was enhanced. This demonstrates that, while maintaining a relatively stable thematic composition, research intensity has continued to increase (Table 6). The sustained high proportion of inter-cluster connections throughout the period indicates that liquid biopsy technology is not limited to specific cancer types or analytical methods. Instead, it connects various research domains, such as pathological diagnosis, omics analysis, and prognosis prediction. Additionally, the keyword turnover rate was moderate, indicating transitional characteristics (Table 7). In the analysis of newly emerged and extinct keywords, a decline in protein analysis-based keywords and an increase in noninvasive diagnostic keywords, such as ‘Circulating tumor DNA,’ ‘Cell-free DNA,’ and ‘Circulating microRNA’ were also observed (Table 8). The most pronounced change was observed in Cluster 3 (Immune/AI-based Precision Medicine). This cluster exhibited a marked increase in both the number of nodes and the TLS per node (Table 6), indicating a rapid expansion in research scale and intensity within a short period. Additionally, the high keyword turnover rate (Table 7), together with the emergence of keywords such as ‘Immune checkpoint inhibitor,’ ‘Programmed death-ligand 1 (PD-L1),’ ‘Immune cell infiltration,’ ‘Machine learning,’ ‘Bioinformatics,’ and ‘Precision medicine,’ indicates that the recent convergence of AI and immunotherapy technologies is driving a paradigm shift in biomarker research toward greater expansion and diversification of research themes (Table 8). Although the proportion of inter-cluster connections decreased slightly, it remained high (Table 6). This suggests that, whereas nearly all connections initially depended on other clusters, a dense internal network was rapidly established. Concurrently, active integration with other clusters continues, indicating that this area is evolving toward a more independent internal structure while maintaining strong integration with other research domains (Tables 6 and 7).
Additionally, through overlay map analysis, qualitative attributes, such as citation impact, recency, and degree of international collaboration, were examined from multiple perspectives beyond the frequency of keyword occurrences. This analysis revealed that topics such as ‘circRNA,’ ‘Tumor mutation burden,’ and ‘Immune checkpoint’ had low frequency but high citation counts, indicating strong academic impact, whereas keywords such as ‘Immunotherapy,’ ‘Deep learning,’ and ‘Precision medicine’ were central in recent research. Additionally, topics such as ‘Pyroptosis,’ ‘Pan-cancer,’ and ‘The Cancer Genome Atlas’ were identified as core global subjects that are being actively researched through collaboration among various countries (Fig. 2). These findings demonstrate the value of a multidimensional thematic evolution analysis that incorporates qualitative attributes and extends beyond simple frequency-based analyses.
The technological change trends identified in this study closely align with the arguments of recent reviews1, which state that biomarker technology is expanding through ctDNA-based diagnostics, immune profiling, AI-based prognostic prediction, and noninvasive precision diagnostics. Thus, the present study provides quantitative empirical evidence supporting these patterns of change in biomarker technology.
However, this study analyzed author keywords from Web of Science article data, and thus has inherent limitations owing to potential data biases, such as the arbitrariness of keyword expressions or the possibility of omission. For example, differences in the expression of similar concepts, such as ‘Immune checkpoint inhibitor’ and ‘PD-1 blockade,’ can influence the interpretation of the network structure. Additionally, because only the Web of Science database was used, articles indexed in other databases such as Scopus and PubMed may have been excluded, and the focus on English-language publications means that research trends in non-English-speaking regions may not have been sufficiently reflected. Furthermore, while the observed increase in inter-cluster edge ratio is interpreted as evidence of growing integration across research domains, it should be noted that expanding semantic diversity over time may also contribute to this trend. As the volume of published research grows, new and more varied keywords enter the literature, potentially increasing connections across clusters independently of genuine thematic convergence. The present analysis cannot fully disentangle these two mechanisms, and this interpretive limitation should be borne in mind when evaluating the inter-cluster connectivity trends reported in Table 5. Additionally, the division of the study period into three equal six-year intervals represents a pragmatic analytical choice. To assess whether the main findings are sensitive to this segmentation, a sensitivity analysis was conducted using an alternative domain-informed segmentation (A1: 2006–2013; A2: 2014–2018; A3: 2019–2023). Under the alternative segmentation, all primary network-level metrics—network density, average clustering coefficient, and inter-cluster edge ratio—showed a consistent monotone increase across periods, mirroring the pattern observed under the main segmentation (Supplementary Table S1). These results confirm that the reported structural trends are robust to reasonable alternative temporal boundaries. Future research should address the limitations of author keywords by employing natural language processing-based analyses of abstracts and full texts, and should also seek to expand the analysis by integrating heterogeneous data, such as clinical trials and patents.
This study goes beyond the limitations of previous bibliometric research in the biomarker field, which has primarily focused on frequency analysis or network visualization, by quantitatively analyzing the technological evolution stages and patterns of thematic convergence through time-series changes in network structure metrics (such as density, clustering coefficient, and inter-cluster edge ratio). By simultaneously considering the research intensity and openness at the cluster level, this study provides an analytical framework for structurally interpreting the differentiated evolutionary mechanisms of biomarker technologies.
Conclusions
This study introduces a temporal keyword co-occurrence network framework for detecting structural transitions in cancer biomarker research, providing an analytical perspective that complements conventional bibliometric approaches by capturing the structural integration and interconnections among research topics that are not readily observable through frequency-based analyses alone. By employing time-series network clustering and keyword turnover analysis, we systematically captured the paradigm shift from traditional pathology-based diagnostics to liquid biopsy, and most recently to AI-assisted analysis, immunotherapy, and precision medicine. Our findings reveal that the biomarker field is evolving through a substantive integration of new technologies with existing knowledge structures, rather than through isolated advancements. This framework enables the systematic identification of thematic convergence and differentiated evolutionary patterns across major research domains within large-scale biomedical literature. These results provide empirical evidence for assessing the maturity of current biomarker technologies and identifying emerging research trends. Furthermore, the multidimensional analytical framework proposed in this study can serve as a useful analytical reference for researchers and policymakers seeking to formulate R&D priorities in the rapidly diversifying landscape of precision medicine.
Electronic Supplementary Material
Below is the link to the electronic supplementary material.
Abbreviations
- AI
Artificial Intelligence
- ceRNA
Competing endogenous RNA
- cfDNA
Cell-free DNA
- circRNA
Circular RNA
- CTCs
Circulating tumor cells
- ctDNA
Circulating tumor DNA
- DEG
Differentially expressed gene
- FISH
Fluorescence in situ hybridization
- HIF-1α
Hypoxia-inducible factor 1-alpha
- iTRAQ
Isobaric tag for relative and absolute quantitation
- KCN
Keyword co-occurrence network
- NGS
Next‑generation sequencing
- PCa
Prostate cancer
- PCA3
Prostate cancer antigen 3
- PD‑1
Programmed death‑1
- PD‑L1
Programmed death ligand‑1
- PET
Positron emission tomography
- RT-PCR
Reverse transcription-polymerase chain reaction
- SELDI-ToF-MS
Surface-enhanced laser desorption/ionization time-of-flight mass spectrometry
- TCGA
The Cancer Genome Atlas
- TLS
Total link strength
- TMB
Tumor mutation burden
- WGCNA
Weighted gene co-expression network analysis
Author contributions
JH was responsible for conceptualization, methodology, and formal analysis, and was a major contributor in writing, reviewing, and editing the manuscript, as well as providing supervision. SK and HK performed data curation and contributed to the methodology. HK was also responsible for software implementation, and SK conducted validation. JS contributed to the methodology, performed formal analysis, and interpreted the data. All authors read and approved the final manuscript.
Funding
This research was supported by Korea Institute of Science and Technology Information (KISTI). (No. K25L4M2C4)
Data availability
The datasets supporting the conclusions of this study were obtained from the Web of Science Core Collection database. The data are subject to license restrictions and are therefore not publicly available. The processed data and analysis results generated in the current study are available from the corresponding author upon request.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Passaro, A. et al. Cancer biomarkers: Emerging trends and clinical implications for personalized treatment. Cell187 (7), 1617–1635. 10.1016/j.cell.2024.02.041 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Gadade, D. D., Jha, H., Kumar, C. & Khan, F. Unlocking the power of precision medicine: exploring the role of biomarkers in cancer management. Futur J. Pharm. Sci.1010.1186/s43094-023-00573-2 (2024).
- 3.Panagopoulou, M. et al. BRCA1 & BRCA2 methylation as a prognostic and predictive biomarker in cancer: Implementation in liquid biopsy in the era of precision medicine. Clin. Epigenetics. 16 (1), 178. 10.1186/s13148-024-01787-8 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Tiwari, A., Mishra, S. & Kuo, T. R. Current AI technologies in cancer diagnostics and treatment. Mol. Cancer. 24 (1), 159. 10.1186/s12943-025-02369-9 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Vyas, A. et al. Advancing the frontier of artificial intelligence on emerging technologies to redefine cancer diagnosis and care. Comput. Biol. Med.191, 110178. 10.1016/j.compbiomed.2025.110178 (2025). [DOI] [PubMed] [Google Scholar]
- 6.Wang, Z. et al. Revolutionizing gastrointestinal cancer research with artificial intelligence: from precision patient stratification to real-world evidence. World J. Gastrointest. Oncol.17 (10), 111339. 10.4251/wjgo.v17.i10.111339 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Prelaj, A. et al. Artificial intelligence for predictive biomarker discovery in immuno-oncology: a systematic review. Ann. Oncol.35 (1), 29–65. 10.1016/j.annonc.2023.10.125 (2024). [DOI] [PubMed] [Google Scholar]
- 8.Garemilla, S. S. S., Kadambala, M. C., Gampa, S. C., Chinthala, S. & Garimella, S. V. Cancer metastasis: therapeutic challenges and opportunities. Med. Oncol.42 (11), 518. 10.1007/s12032-025-03072-x (2025). [DOI] [PubMed] [Google Scholar]
- 9.Qiao, Y., Xie, D., Li, Z., Cao, S. & Zhao, D. Global research trends on biomarkers for cancer immunotherapy: Visualization and bibliometric analysis. Hum. Vaccin Immunother. 21 (1), 2435598. 10.1080/21645515.2024.2435598 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Xu, P. et al. Association between intestinal microbiome and infammatory bowel disease: insights from bibliometric analysis. Comput. Struct. Biotechnol. J.20, 1716–1725. 10.1016/j.csbj.2022.04.006 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Li, Y. et al. Knowledge mapping of exosomes in prostate cancer from 2003 to 2022: a bibliometric analysis. Discov Oncol.15 (1), 307. 10.1007/s12672-024-01183-x (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Huang, T. et al. Current perspectives and trends of CD39-CD73-eAdo/A2aR research in tumor microenvironment: a bibliometric analysis. Front. Immunol.15, 1427380. 10.3389/fimmu.2024.1427380 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Jian, Y. et al. Mapping the evolving trend of research on efferocytosis: a comprehensive data-mining-based study. BioData Min.18, 58. 10.1186/s13040-025-00475-4 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Kumar, R., Romano, J. D. & Ritchie, M. D. Network-based analyses of multiomics data in biomedicine. BioData Min.18, 37. 10.1186/s13040-025-00452-x (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Cobo, M. J., López-Herrera, A. G., Herrera‐Viedma, E. & Herrera, F. Science mapping software tools: Review, analysis, and cooperative study among tools. J. Am. Soc. Inf. Sci. Technol.62 (7), 1382–1402. 10.1002/asi.21525 (2011). [Google Scholar]
- 16.Radhakrishnan, S., Erbis, S., Isaacs, J. A. & Kamarthi, S. Novel keyword co-occurrence network-based methods to foster systematic reviews of scientific literature. PLoS ONE. 12 (3), e0172778. 10.1371/journal.pone.0172778 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Su, H. N. & Lee, P. C. Mapping knowledge structure by keyword co-occurrence: a first look at journal papers in Technology Foresight. Scientometrics85, 65–79. 10.1007/s11192-010-0259-8 (2010). [Google Scholar]
- 18.You, T., Yoon, J., Kwon, O. H. & Jung, W. S. Tracing the evolution of physics with a keyword co-occurrence network. J. Korean Phys. Soc.78, 236–243. 10.1007/s40042-020-00051-5 (2021). [Google Scholar]
- 19.Catone, M. C., Diana, P. & Giordano, G. Keywords co-occurrence analysis to map new topics and recent trends in social research methods. In Advanced Information Networking and Applications. AINA 2020. Advances in Intelligent Systems and Computing Vol. 1151 (eds Barolli, L. et al.) (Springer, ). 10.1007/978-3-030-44041-1_93.
- 20.Li, W., Zhou, H., Lu, Z. & Kamarthi, S. Navigating the evolution of Digital Twins research through keyword co-occurrence network analysis. Sensors24, 1202. 10.3390/s24041202 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Araújo, T., Abreu, A. & Louçã, F. The evolution of complexity co-occurring keywords: bibliometric analysis and network approach. Published online 2023. 10.48550/arXiv.2308.00992
- 22.Kim, H. et al. A keyword-based approach to analyzing scientific research trends: ReRAM present and future. Sci. Rep.15, 12011. 10.1038/s41598-025-93423-5 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhang, X., Xie, Q., Song, C., Li, Y. & Zhao, K. Mining the evolutionary process of knowledge through multiple relationships between keywords. Scientometrics127 (4), 2023–2053. 10.1007/s11192-022-04272-2 (2022). [Google Scholar]
- 24.Donthu, N., Kumar, S., Mukherjee, D., Pandey, N. & Lim, W. M. How to conduct a bibliometric analysis: An overview and guidelines. J. Bus. Res.133, 285–296 (2021). [Google Scholar]
- 25.Song, Y., Lei, L., Wu, L. & Chen, S. Studying domain structure: a comparative analysis of bibliographic coupling analysis and co-citation analysis considering all authors. Online Inf. Rev.47, 123–137 (2022). [Google Scholar]
- 26.van Eck, N. J. & Waltman, L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics84, 523–538. 10.1007/s11192-009-0146-3 (2010). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Bentley, R. A. Random drift versus selection in academic vocabulary. PLoS ONE, 2008;3(8):e3057. (2008). 10.1371/journal.pone.0003057 [DOI] [PMC free article] [PubMed]
- 28.Iñiguez, G., Pineda, C., Gershenson, C. & Barabási, A. L. Dynamics of ranking. Nat. Commun.13 (1), 1646. 10.1038/s41467-022-29256-x (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Choi, S. J. et al. Alteration of DNA methylation in gastric cancer with chemotherapy. J. Microbiol. Biotechnol.27 (8), 1367–1378. 10.4014/jmb.1704.04035 (2017). [DOI] [PubMed] [Google Scholar]
- 30.McCartney, A. et al. Metabolomics in breast cancer: A decade in review. Cancer Treat. Rev.67, 88–96 (2018). [DOI] [PubMed] [Google Scholar]
- 31.Ding, Z., Wang, N., Ji, N., Li, X. & Chen, Y. Proteomics technologies for cancer liquid biopsies. Mol. Cancer. 21 (1), 53. 10.1186/s12943-022-01526-8 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Lu, Y., Ma, Y., Liu, Q. & Luo, D. Recent progress in mass spectrometry-based liquid biopsy for cancer detection and analysis: A comprehensive review. TRAC Trends Anal. Chem.190, 118291. 10.1016/j.trac.2025.118291 (2025). [Google Scholar]
- 33.Alum, E. U. AI-driven biomarker discovery: enhancing precision in cancer diagnosis and prognosis. Discov Oncol.16 (1), 313. 10.1007/s12672-025-02064-7 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets supporting the conclusions of this study were obtained from the Web of Science Core Collection database. The data are subject to license restrictions and are therefore not publicly available. The processed data and analysis results generated in the current study are available from the corresponding author upon request.



