Skip to main content
Frontiers in Artificial Intelligence logoLink to Frontiers in Artificial Intelligence
. 2026 Mar 9;9:1737530. doi: 10.3389/frai.2026.1737530

Unveiling patterns in clinical data: exploring the role of large language models and clustering algorithms

Abbas S Ali 1,*, Subi Gandhi 2, Syed H Jafri 3, Mohammed M Ali 4, Syed Y Raza 5, Sulaiman Samian 6, James Mehaffey 6
PMCID: PMC13006407  PMID: 41877738

Abstract

Objective

Large Language Models (LLMs) have shown exceptional performance in natural language processing, yet their utility in structured clinical data analysis remains relatively underexplored. This pilot study investigates whether LLM-generated embeddings can preserve the structural integrity of clinical datasets and enhance predictive modeling, particularly in resource-constrained settings.

Methods

We applied dimensionality reduction techniques such as Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), and k-means clustering to compare original data structures with those derived from LLM embeddings. Evaluation metrics included cosine similarity, area under the curve (AUC), and R2, applied across 100 synthetic datasets and two real-world clinical datasets: the UCI medical database and endocarditis patient records. We assessed multiple LLM architectures, including BERT, RoBERTa, Llama 2, and E5-small, focusing on predictive accuracy and computational efficiency.

Results

LLM embeddings closely mirrored original data structures, with BERT achieving a cosine similarity of 0.95 on linear datasets and Llama 2 (30B) reaching 0.85 on quadratic datasets, albeit with higher computational costs. Predictive performance improved significantly across the board with increases in subject variable ratio (SVR), three groups were identified similar performance, assisted better and assisted significantly better. These groups differed based upon the equation used to generate synthetic data.

Discussion

These findings highlight the potential of LLMs to enhance structured data analysis by identifying optimal conditions, such as SVR thresholds, for their practical use. The trade-off between computational cost and performance across different LLM architectures is also emphasized, suggesting the need for context-specific model selection.

Conclusion

LLMs can be effectively leveraged to repurpose existing clinical datasets for individualized clinical questions, such as optimizing surgical timing for patients with infective endocarditis and embolic stroke. This approach advances precision medicine and supports data-driven clinical decision-making.

Keywords: endocarditis, large language models, medical informatics, natural language processing, precision medicine, predictive Modeling

1. Introduction

The integration of artificial intelligence (AI) into healthcare is transforming clinical research, with large language models (LLMs) such as ChatGPT playing a pivotal role (Alum and Ugwu, 2025; Parekh et al., 2023; Nzeako et al., 2025; Ali et al., 2023). These models offer powerful tools for interpreting complex clinical data and uncovering nuanced patterns. However, their application in healthcare is tempered by known limitations, including issues with accuracy, contextual relevance, and hallucination (Li et al., 2023; Shah et al., 2024). To mitigate these risks, effective prompt engineering, setting the temperature of the LLM and strategic querying are essential, along with rigorous evaluation of LLM outputs against domain expertise to ensure reliability (Shah et al., 2024).

Traditionally, researchers have used statistical methods such as univariate analysis, stepwise regression, and multivariate analysis to identify and evaluate cardiovascular disease risk factors. These approaches have been instrumental in uncovering associations, such as the link between high cholesterol and heart disease. With the advent of greater computational power, machine learning (ML) methods, both supervised and unsupervised, have enabled the analysis of high-dimensional datasets, revealing hidden patterns (Nadif and Role, 2021; Tirthyani et al., 2024; Vinci, 2024). However, ML models often lack transparency, which makes their decision-making processes challenging to interpret (Golovin, 2024; Patil, 2024).

To address the challenge of interpreting complex data, dimensionality reduction techniques such as Principal Component Analysis (PCA) and t-distributed Stochastic Neighbor Embedding (t-SNE) are commonly used (Ray et al., 2021). PCA projects data into lower dimensions by capturing the greatest variance through eigenvectors, typically highlighting linear relationships. In contrast, t-SNE excels at revealing non-linear patterns and local clusters by mapping data into a probabilistic space and is commonly used for visualization. Although t-SNE does not preserve global structure and can be computationally intensive, it is often preceded by PCA to improve efficiency and clustering quality (Shah and Silwal, 2019; Song et al., 2013).

In addition to dimensionality reduction, pattern recognition techniques like K-Nearest Neighbors (KNN) and k-means clustering are widely used (Zriaa and Amali, 2021). KNN is a simple yet effective classification method that identifies the ‘k’ closest data points to a query point and uses them to determine its class. It is often confused with k-means clustering, which instead relies on centroids and iterative repositioning to form clusters. Unlike k-means, KNN does not involve centroids or iterative updates and is purely distance-based (Cunningham and Delany, 2020) (Figure 1).

Figure 1.

Flowchart showing synthetic data processed via dimensionality reduction, followed by KNN clustering using distance and cosine similarity. LLM embeddings are used to reproduce data concepts, which are then applied to UCI data.

Comprehensive framework for evaluating LLM-assisted analysis.

Modern large language models (LLMs), such as OpenAI’s GPT series, are built on the Transformer architecture introduced by Vaswani et al. (2017) (Yenduri et al., 2024). These models generate coherent and contextually relevant text by predicting the next word in a sequence based on prior context (Alomari et al., 2023; Chafi et al., 2023). Central to their design is the self-attention mechanism, which captures global dependencies and enables parallel processing for improved scalability (Ahmed et al., 2023). A key innovation—multi-head attention—allows multiple attention heads to operate simultaneously, capturing diverse linguistic features including syntactic and semantic relationships. These outputs are concatenated and transformed to produce rich, high-dimensional embeddings that encode complex semantic relationships. Variations in training data, tokenization, attention mechanisms, and directionality contribute to differences among models (Passi et al., 2024; Devlin et al., 2018; Raiaan et al., 2024).

In healthcare, Transformer-based models have demonstrated potential in uncovering hidden determinants, supporting personalized care, and enhancing predictive modeling (Alum and Ugwu, 2025). This study leverages embeddings from advanced LLMs to represent tabular patient data as natural language, withholding outcome variables to preserve analytic integrity. Structured clinical data often suffers from sparsity and multicollinearity, particularly with high-dimensional categorical variables, where traditional preprocessing methods like one-hot encoding can inflate feature space and reduce interpretability (Zhao et al., 2017).

By modeling inter-variable relationships through attention mechanisms, we aimed to reduce preprocessing, preserve semantic structure, and improve both accuracy and interpretability. Specifically, we aimed to evaluate the feasibility of LLMs in structured clinical data analysis by addressing sparsity and multicollinearity, modeling complex relationships, assessing embedding integrity via clustering metrics, identifying optimal performance conditions, and supporting precision medicine through interpretable insights.

2. Methods

2.1. Comparative evaluation framework for transformer-based embedding models

2.1.1. Selection of transformer architectures and preprocessing pipeline

We evaluated Transformer models including BERT, RoBERTa, DistilBERT, ERNIE, T5, XLNet, GatorTron, MiniLM, E5-small-cluster, and LLaMA2-30B, selected for computational feasibility, architectural diversity, and performance across varied data relationships (Devlin et al., 2018; Raiaan et al., 2024; Zhang et al., 2019; Raffel et al., 2019; Yang et al., 2019; Yang et al., 2022; Wang et al., 2020; Wang et al., 2022; Shindo et al., 2025)Performance was assessed using geometric (centroid distances, cosine similarity) and predictive metrics (R2, AUC), with ROC-based AUC for kNN cluster prediction mitigating high-dimensional degradation (Salem et al., 2025; Aggarwal et al., 2001; Saito and Rehmsmeier, 2015). Embedding time and comparisons with traditional ML models were recorded, using R2, AUC, or MAE depending on outcome type. Hyperparameters were tuned via 5-fold Grid Search Cross-Validation.

Preprocessing included low-variance feature removal (Figure 2), median/mode imputation, and one-hot encoding. SMOTE was applied post–cross-validation to address class imbalance. An ensemble Voting Classifier (Random Forest, XGBoost, SVM) was used, with hyperparameters tuned to prioritize precision and recall over accuracy for balanced clinical performance.

Figure 2.

Flowchart in blue and light blue boxes showing a data analysis process: starting with KNN clustering, followed by supervised machine learning to predict clusters, identifying top variables, then applying machine learning to predict CAD or mortality, and identifying top predictive variables.

Steps involving integration of LLM, clustering, and ML techniques.

2.1.2. Clustering methodology and embedding comparison

To analyze high-dimensional data, K-means clustering was applied using Euclidean distance and cosine similarity (Uhryn and Kalancha, 2024; Mussabayev, 2024). Given the limitations of these metrics in high-dimensional spaces (Zahariah and Midi, 2024), PCA and t-SNE were used to preserve global and local data structures, respectively. Cluster relationships were assessed using cosine similarity, Euclidean distance, and Spearman correlation to evaluate how well LLM-generated embeddings preserved original data geometry. Optimal cluster counts were determined via mean pooling, elbow method, and silhouette scores. Feature selection improved clustering coherence, as evidenced by higher silhouette scores, which validated the approach.

2.1.3. Dataset design, SVR analysis, and model performance

We generated 100 synthetic datasets (500 rows each) with SVRs ranging from 10 to 100, incorporating exponential, cubic, quadratic, and linear relationships. These datasets included continuous and categorical variables with embedded collinearity, allowing us to assess how SVR and variable types influence model performance (Supplementary Figure S1).

Model performance was evaluated using linear regression, random forest, and gradient boosting, with metrics including R2, RMSE, MAE, and AUC (Supplementary Figures S2, S3, and Figure 3). LLM-assisted models using eight embeddings outperformed unassisted models in 79% of binary tasks. SHAP scores identified top predictors, and paired t-tests confirmed significance. Fisher’s Exact Test (p = 0.0001; Supplementary Table S1) showed a strong association between model performance and dataset type. While linear and cubic datasets showed similar results across models, exponential and quadratic datasets benefited most from LLM assistance—none favored unassisted models.

Figure 3.

Scatter plot with trend lines showing the relationship between subject variable ratio (unassisted model, x-axis) and R-squared difference (assisted minus unassisted, y-axis), grouped by performance: light blue for "Assisted Better," gray for "Similar Performance," and dark green for "Assisted Significantly Better." Each group has a separate linear regression line with shaded confidence intervals; positive correlation is seen for dark green, negative for light blue, and near zero for gray. Legends identify each group and trend line style.

Illustration of mean R2 and 95% CI across linear models: Optimal LLM performance at SVR 35–40 and 15–20 categorical variables.

2.1.4. Categorical complexity and predictive reliability

Supplementary Figure S4 shows that lower SVRs (<20) and high categorical complexity (>20 variables) reduce model reliability. Residual plots and the Breusch-Pagan test (Supplementary Figure S4) revealed heteroscedasticity at low SVRs. LLM assisted models performed better than LLM unassisted models with fewer subjects and higher number of variables (Figure 3).

2.1.5. Embedding fidelity and concept capture

We assessed whether LLM embeddings preserved original data structure using cosine similarity (Supplementary Figures S6–8 and Figure 4). BERT achieved the highest similarity (0.95), followed by E5 and LLaMA 2 30B. Ensemble models performed best in high-quality clusters, while logistic regression excelled in RoBERTa-derived clusters but struggled with lower-performing embeddings like T5. Supplementary Figure S6 illustrates the reconstruction of the original outcome variable used to generate the synthetic datasets. When cluster assignments derived from different LLMs were incorporated as predictors, the resulting AUC values varied across supervised ML algorithms, reflecting differences in how each model leveraged the LLM-based cluster structure. Supplementary Figures S9 and Figure 5 show confidence intervals for cosine similarity and AUC, reinforcing the superior performance of BERT and E5-small.

Figure 4.

Box plot graph titled “Cosine Similarity by Function Type and Model” compares cosine similarity for four function types (cubic, quadratic, linear, exponential) across eight models, with roberta, bert, gatortron, minilm (blue shades) generally performing higher than t5, llama, e5_small, ernie (red shades).

Boxplot of cosine similarity scores.

Figure 5.

Bar graph comparing mean AUC values with ninety-five percent confidence intervals across eight LLM clusters: E5, LLAMA 2 30B, MiniLLM, Ernie, GatorTron, RoBERTa, BERT, and T5. E5, LLAMA 2 30B, RoBERTa, and T5 display red stars, denoting statistically significant AUC differences. BERT has the highest variance while T5 shows the lowest mean AUC. Confidence intervals and star indicators are noted in a legend and caption.

Comparison of mean AUC scores: LLM embeddings vs. K-means clusters for binary outcome (median split).

2.1.6. Feature selection and SVR balance

Feature selection improved generalizability by reducing dimensionality and overfitting. Maintaining SVRs above 20 and limiting categorical variables to fewer than 20 enhanced model accuracy and interpretability (Figure 3).

2.2. Clinical variable context

2.2.1. Categories of clinical data

This study utilized clinical data spanning key categories for predictive modeling, including demographics (age, gender, socioeconomic status) from the UCI dataset of 76 variables across 303 patients. Medical history (e.g., hypertension, diabetes, smoking), lab values (blood sugar, cholesterol), and diagnostic data (ECG abnormalities, ST depression, thallium stress test results) were included to assess cardiac risk and disease severity. Each variable was selected for its relevance to improving model accuracy and informing patient prognosis.

2.2.2. Relevance to clinical prediction

Variables included in the clinical prediction models were selected for their relevance and impact on predictive accuracy. By incorporating demographics, medical history, lab results, and diagnostic findings, the models were designed to reflect real-world clinical decision-making. This careful selection enhanced both model interpretability and clinical relevance.

2.2.3. Variable processing and missing data

The endocarditis dataset variables were grouped into three types—continuous, categorical, and ordinal—each contributing uniquely to predictive modeling. Continuous variables (e.g., BMI, age, EF, troponin) offered quantifiable health indicators. Categorical variables (e.g., valve type, pulmonary dysfunction, IV drug use, stroke) captured discrete clinical traits for risk stratification. Ordinal variables (e.g., valvular insufficiency severity, surgical urgency, vegetation size) reflected ranked assessments of disease severity and intervention urgency. To manage missing data, median imputation was used for continuous variables and mode imputation for categorical ones (Occhipinti, 2024; Memon et al., 2023). To address class imbalance—especially for rare outcomes like mortality, the Synthetic Minority Oversampling Technique (SMOTE) was applied after cross-validation (Dou et al., 2024; Elreedy and Atiya, 2019), improving model robustness. A summary of preprocessing steps and their purposes is provided in Table 1.

Table 1.

Summary of steps, techniques, and purpose.

Step Technique Purpose
Handling missing data Multiple imputation Fill missing values (continuous variables)
Numeric preprocessing Standard scaler Rescale variables to zero mean/unit variance
Categorical preprocessing One-hot encoding Convert categorical variables to binary features
Data balancing SMOTE Generates synthetic samples to balance imbalanced datasets and improve model performance.

2.2.4. Sentence construction as a bridge between structured data and LLMs

Structured data rows—excluding outcome variables—were converted into natural language sentences using a Python script (see Supplemental Folder and GitHub repository: https://github.com/aliabbasmd/response_reviewer). These were processed by LLMs to generate embeddings, which were clustered via k-means. Cluster assignments served as categorical features in LLM-assisted models, while unassisted models omitted this step. In clinical datasets, clusters were treated as intermediate outcomes; SHAP analysis identified key predictors of cluster membership, informing a refined model for the primary outcome. This two-step approach enhanced both performance and interpretability.

2.3. Resource and scalability considerations

2.3.1. Resource requirements and scalability challenges

LLMs, though optimized for text (Bommasani et al., 2021), can still reveal model behavior through embeddings, even with proprietary weights (Lundberg and Lee, 2017). We used SHAP and LIME to assess alignment with clinical judgment in complex cases like endocarditis with stroke (Lundberg and Lee, 2017; Ribeiro et al., 2016; Bender et al., 2021). Interpretability is essential for safe decision-making.

LLMs require GPU-based parallel processing, with deployment dependent on tools like CUDA, Torch, and Hugging Face (Moore, 2023). Quantized models reduce memory requirements and improve accessibility, but do so at the cost of some loss in numerical precision and model accuracy. In contrast, full-precision (non-quantized) models demand extremely large amounts of VRAM to perform high-precision calculations. These requirements frequently exceed available hardware capacity, causing systems to crash or terminate processes due to ‘memory explosion.’ Furthermore, their extreme sensitivity allows them to perfectly memorize data noise, causing the loss function to collapse to zero through overfitting rather than true learning. Frequent software updates and infrastructure demand challenge reproducibility, highlighting the need for institutional support and skilled, prompt engineers to ensure scalable, impactful use.

2.3.2. Description of datasets

Three datasets were used to evaluate the performance of LLM-assisted models in structured clinical data analysis. Table 2 summarizes their characteristics and rationale. The first two datasets supported conceptual development, while the third focused on real-world clinical application.

Table 2.

Characteristics of the three datasets used in the study.

Number of dataset Number of rows Number of variables Non-zero data entries (%)
Synthetic data 500 2 random continuous, 1 calculated, 2 to 40 categorical 100
UCI data 303 13 (7 categorical) 91.6
West Virginia University ‘Raw’ Endo-caritas Data (de-identified data) 442 58 (51 categorical) 15
2.3.2.1. Synthetic data

To assess the sensitivity and robustness of our analytical pipeline to variations in data structure, we generated 100 synthetic datasets, each containing 500 observations. Within these datasets, we systematically manipulated the ratio of samples to variables (500 samples and 2–40 categorical variables) to represent settings with differing levels of dimensional complexity. We introduced controlled multicollinearity by allowing selected categorical predictors to exert direct influence on the outcome variable. This experimental design enabled us to evaluate pipeline performance across a broad spectrum of feature dependencies, redundancy patterns, and structural complexities that often characterize real-world clinical data.

By simulating such diverse conditions, we were able to rigorously interrogate the stability and reliability of our natural language processing (NLP) workflow. In particular, this approach allowed us to test whether the model retained consistent behavior under challenging scenarios including high feature overlap, dominance of specific categorical attributes, and varying degrees of noise. As various LLMs have diverse training data and methods our process also teased out the particular kind of LLM best for a particular data structure. Collectively, these experiments provided a controlled environment to evaluate whether the pipeline remained robust when confronted with data characteristics known to degrade model interpretability and clustering performance (Supplementary Table S1).

Each row of every synthetic dataset was transformed into a natural-language sentence using a custom text-generation function. The resulting sentences were stored in a dedicated “combined” column, creating a structured text corpus in which each row served as an independent qualitative unit. To convert these sentences into numerical representations suitable for downstream analysis, we generated text embeddings using the final hidden state of the [CLS] token from multiple Transformer architectures. Both general-purpose models (RoBERTa, DistilBERT) and domain-specific biomedical models (e.g., GatorTron) were employed to capture theoretical variability in semantic representation. Larger models were processed using batch-based inference to accommodate memory constraints, resulting in a final embedding matrix for each model type.

Clustering analysis was performed using K-means to evaluate how effectively each model’s embeddings grouped semantically similar entries. The optimal number of clusters was determined using the Elbow Method, which identifies the inflection point at which additional clusters yield diminishing returns in explained variance. Cluster quality and cohesion were then assessed using established internal validation metrics, including the Silhouette Score and the Calinski–Harabasz Index. These metrics allowed systematic comparison of embedding quality across models and informed selection of the most effective representation for subsequent analysis.

To further characterize cluster separability and global embedding geometry, high-dimensional embeddings were visualized using a two-stage dimensionality reduction approach. Principal Component Analysis (PCA) was first applied to preserve maximal variance while reducing dimensionality. Subsequently, t-Distributed Stochastic Neighbor Embedding (t-SNE) was used to project the data into two dimensions. This combined approach provided a qualitative assessment of the semantic structure of the embeddings and facilitated visual inspection of cluster boundaries, local coherence, and overall geometric organization.

The project’s GitHub repository contains not only the primary analysis scripts but also detailed, step-by-step walkthroughs located in the [GitHub] and Ahmed et al. (2023). These examples illustrate the full workflow for generating T5-based text embeddings and conducting unsupervised clustering to identify groups of clinical phenotypes. This supplemental material provides users with concrete demonstrations of the modeling process, facilitating hands-on exploration and replication.

A streamlined README.md file integrates these components and provides clear guidance on repository structure, dependencies, and execution steps. Executing the python files sequentially allows for reproduction and insight into the synthetic data. Worked out UCI examples illustrate clinical insights. Together, these resources support transparency, reproducibility, and knowledge transfer, making the workflow accessible for both research replication and educational use in clinical informatics and training environments.

2.3.2.2. West Virginia University endocarditis data

We used de-identified clinical data from patients with endocarditis (Supplementary Table S2) who had undergone life-saving heart valve surgery (number of variables = 517; sample size = 442), with a subject-to-variable ratio of 0.85 in order to implement and assess our technique (Figure 3). The state of West Virginia is plagued by illicit drug use (Merino et al., 2019; Bhandari et al., 2022), and the data set was representative of this population in the hospital data. Life-threatening infections such as IE are on the rise in West Virginia (WV) in conjunction with the injection drug use epidemic (Raiaan et al., 2024; Zhao et al., 2017).

We followed the steps outlined in section 2.3.2.1 and used the same transformers to identify the underlying clinical concepts for the WVU IE clinical data. Figures 3 and 4 illustrate the process of analyzing the WVU-HVI data and its characteristics.

The West Virginia University Endocarditis dataset cannot be publicly released due to privacy and confidentiality restrictions associated with the underlying clinical data and can be made available on request.

2.3.2.3. University of California Irvine data

In this step, we used the University of California Irvine (UCI) clinical data repository (Ahmed et al., 2023), which is a robust dataset for CAD, a leading cause of mortality in the United States (Passi et al., 2024). Data from this repository (Supplementary Table S2) were used to assess whether LLM-assisted ML identifies features associated with the outcome variable similar to conventional ML. Using this repository, we demonstrated nine perspectives on CAD data using various LLM transformers, including Llama2, BERT, RoBERTa, DistilBERT, T5, Ernie, Gatorton, and GatortonS, and XLNet.

A similar analysis involving dimensionality reduction techniques was performed using UCI data (Devlin et al., 2018). The data and Python code available on their website were downloaded and processed. This analysis aimed to identify the most significant features in the data that contribute to the clustering of the embeddings. Because directly obtaining this information from t-SNE is not feasible, we used a different clustering technique, namely K-means. The K-means-generated clusters were subsequently superimposed onto the t-SNE clusters to evaluate their similarity. Next, we thoroughly assessed the features associated with the K-means clusters using ML techniques to identify a list of refined variables. Subsequently, we used these variables to construct supervised ML models to predict the outcome of interest. The UCI dataset, a widely published resource, included clinical and diagnostic features such as stress test results and coronary angiography, serving as a benchmark for comparing variable distributions in confirmed versus unconfirmed coronary artery disease cases (Supplementary Table S3).

The Python code used to run the analyses is publicly available on GitHub (link above).

None of the synthetic data used in this study were part of any language model’s pretraining corpus, especially the synthetic and endocarditis data.

LLM-assisted and Random Forest models showed similar performance in predicting coronary artery disease (AUCs of 0.88 and 0.89). The LLM model prioritized heart rate, ST depression, age, cholesterol, and blood pressure, while Random Forest emphasized thallium imaging and exercise-induced angina. These results align with SVR and categorical variable counts, confirming both models performed similarly under these conditions.

3. Results

3.1. Performance metrics

Cosine similarity was used to assess how well LLM embeddings preserved the structure of the original datasets. Despite requiring 24 h to process 100 datasets, LLaMA 2 30B underperformed compared to smaller, faster models like BERT and E5-small, which completed the same task in minutes. Table 3 summarizes the trade-off between performance (cosine similarity) and resource utilization (processing time), highlighting the efficiency of smaller models.

Table 3.

Performance-resource tradeoff assessment.

Name of LLM Cosine distance for data generated using exponential equations (Mean ± Confidence Interval) Processing time (min) for 100 Files
bert 0.716 ± 0.012 Less than 5 min
minillm 0.715 ± 0.014 Around 30 min
roberta 0.711 ± 0.015 Less than 5 min
gatortron 0.707 ± 0.013 Less than 5 min
ernie 0.703 ± 0.014 Less than 5 min
t5 0.700 ± 0,015 Less than 5 min
Llama 2 30b 0.684 ± 0.016 24 h (1,440 min)
T5 small 0.665 ± 0.021 Less than 5 min

3.2. Comparative evaluation against established methods

Supplementary Figures S2, and Figure 3 illustrate scenarios where LLM-assisted models outperform traditional models, particularly under favorable SVR and categorical complexity conditions. Supplementary Figure S7 shows that specific LLMs perform better depending on the underlying data-generating function. For example, Supplementary Figure S8 demonstrates that BERT significantly outperforms E5-small for exponential functions, while Figure 5 highlights AUC differences across models, p < 0.05 marked by red stars.

3.2.1. Statistical validation of findings

To validate robustness, we ran multiple iterations across LLMs, modeling each LLM-derived cluster as the target (Y) and using other variables as predictors (X). SHAP analysis identified the top 20 features, with Random Forest outperforming Gradient Boosting in efficiency and consistency. AUC values remained stable, and the top 10 SHAP features were consistent across runs—suggesting LLM embeddings did not add predictive value beyond structured features. If deeper latent patterns were captured, we would expect distinct SHAP profiles or improved AUCs. However, the consistency across LLM-assisted and unassisted models suggests embeddings may effectively capture the same core concepts. Variables such as age, BMI, ejection fraction, embolism, and cardiac arrest consistently ranked as top predictors of hospital mortality. Notably, “Days from stroke to surgery” emerged as a key variable (Supplementary Figure S9), with shorter intervals linked to better outcomes—highlighting the clinical importance of timely surgical intervention. In the clinical context of a young patient usually a drug user with infection of left sided heart valves with a clot of high likelihood to embolize to the brain, early surgery even after a stroke is associated with improved in-hospital mortality. Delay in such patients increases death from endocarditis.

3.2.2. Benchmarking against traditional feature engineering

The consistent performance of various LLMs suggests that their embeddings effectively capture core data concepts, offering a reliable framework for benchmarking model interpretability. Deviations in future experiments may signal either deeper conceptual understanding or limitations in capturing key patterns—highlighting the influence of proprietary model weights versus embeddings.

This consistency across architectures indicates that LLM embeddings generalize well, regardless of training objectives. In more complex datasets, such as unstructured clinical notes, performance differences may become more pronounced. This framework can then help identify which models best capture meaningful data aspects, using both predictive metrics (e.g., AUC) and conceptual alignment (e.g., SHAP feature similarity).

3.3. Clinical applications

3.3.1. Empirical validation of clinical relevance

Supplementary Figure S9 presents SHAP scores from the LLM-unassisted model, with postoperative ejection fraction (EF) emerging as the strongest predictor of hospital mortality. Lower EF values (indicated by red SHAP points) correlate with increased mortality risk, consistent with clinical understanding that impaired cardiac function elevates surgical risk (Topkara et al., 2005; Kim et al., 2024).

Additional variables reinforce known clinical patterns. Female gender appears protective, with lower SHAP values (blue points clustered near zero), while male gender is associated with slightly higher mortality risk. Older age is associated with an increased risk, as indicated by the red points on the positive SHAP axis. Extremes in body mass index (BMI)—both underweight and obese—also contribute to poor outcomes. A history of cardiac arrest or intraoperative arrest significantly raises mortality risk (Fielding-Singh et al., 2020), reflected in high SHAP contributions. Larger vegetations (>1 cm) predict worse outcomes due to increased embolic risk and procedural complexity. “Days from stroke to surgery” is a critical variable: longer delays (red, high feature values) are associated with higher mortality, while shorter delays (blue) may be neutral or protective, suggesting earlier intervention improves outcomes. Selection bias cannot be excluded here that is to say the healthier patients underwent surgery earlier and had better outcomes.

Supplementary Figure S10 shows SHAP scores from the LLM-assisted model, highlighting intravenous drug use (IVDU) as a significant predictor of hospital mortality in endocarditis patients (Tan et al., 2020; Citro et al., 2022; Nguemeni Tiako et al., 2019). This aligns with evidence linking IVDU to tricuspid valve endocarditis (Stolear et al., 2024; Pietrolungo and Gandler, 2023), often complicated by septic pulmonary emboli (SPEs). SHAP plots for SPEs show moderate contributions to mortality risk, with red points (high feature values) indicating higher risk and blue points (low feature values) suggesting minimal impact. This separation underscores SPEs as clinically meaningful indicators of disease severity.

These findings suggest that LLM-assisted models can integrate detailed, clinically relevant information, potentially enhancing predictive performance and interpretability in complex cases.

3.3.2. Real-world applicability

SHAP plots from LLM-assisted models [e.g., E5-small, DistilBERT, GatorTron (Wu, 2023; Akkur, 2023)] reveal consistent influence from core features like age, BMI, and preoperative EF, while each model emphasizes different clinical variables—GatorTron highlights IVDU and catheterization, DistilBERT focuses on surgical and valvular conditions. These differences suggest that each LLM captures distinct aspects of patient profiles, which can be leveraged for subpopulation-specific modeling and personalized care. One-hot encoded features support precise stratification, helping clinicians select models that best reflect disease mechanisms. Patient-specific predictions visualized using tools like LIME, help bridge statistical outputs and clinical interpretability. As shown in Supplementary Figure S11, a case from the high-risk tertile of the endocarditis dataset illustrates a 75% survival prediction for a 28-year-old male, emphasizing the importance of completing antibiotic therapy. Such visualizations enhance clinician-patient communication and support shared decision-making by clarifying how individual variables influence outcomes.

3.4. Strengths

This study shows that LLMs with lower computational demands can perform comparably to more resource-intensive models using geometric metrics, making them suitable for deployment on modest hardware. Synthetic data enabled pre-deployment testing across varying sample-to-variable ratios and class imbalances, improving efficiency. LLM-assisted models, combined with interpretability tools like SHAP and LIME, produced clinically meaningful insights in real-world datasets and are scalable to larger applications. The methodology is adaptable across healthcare domains, with clustering and embedding techniques revealing actionable patterns. Deployment guidelines include using de-identified data locally, managing platform licensing, and ensuring secure environments. Interpretability was enhanced through dimensionality reduction techniques, and benchmarking on structured data provided a novel way to assess model alignment and feature relevance.

A key strength of LLMs lies in their ability to leverage attention mechanisms to identify clinically relevant information within complex inputs. Through the Query–Key–Value framework, the model calculates similarity scores between the query and each key, applies scaling and softmax normalization, and generates a weighted sum that privileges the most contextually important features. This enables the model to process nuanced clinical questions—for example, estimating the likelihood of tricuspid valve endocarditis in a pregnant patient with substance use—without reducing them to binary classifications. Instead, the model distributes probabilities across potential interpretations, reflecting the inherent uncertainty of real-world clinical reasoning.

LLMs also integrate probabilistic reasoning akin to Bayesian inference, allowing them to operate effectively under conditions of incomplete, noisy, or uncertain data. In this framework, attention determines which components of the clinical presentation warrant emphasis, while probabilistic inference estimates the degree of certainty associated with a conclusion. Such reasoning parallels the implicit cognitive processes used by experienced clinicians and provides a systematic means of scaling these insights across populations and settings.

A further strength lies in the high-dimensional embedding space through which LLMs represent knowledge. Rather than storing discrete facts, LLMs encode related clinical concepts as clustered vectors in latent space. This facilitates recognition of patterns such as the relationships between endocarditis, intravenous drug use, pregnancy-related physiological changes, and in-hospital mortality risk. In this manner, LLM-assisted analysis can support pattern discovery, risk stratification, and clinical decision-making. Moreover, just as large models can be distilled into smaller, deployable versions, clinician expertise can be analogously aggregated and transmitted via model-guided processes, offering a scalable mechanism to disseminate experiential knowledge.

3.5. Limitations

Several limitations warrant consideration. First, evaluating transformer-based models remains challenging due to a lack of consensus on optimal performance metrics. Traditional ML metrics, geometric measures in embedding space, and task-specific clinical validity assessments each capture different properties, and no unified framework currently exists.

Second, constraints related to pretraining data introduce uncertainty. While proprietary clinical datasets and synthetic endocarditis data were excluded from pretraining, complete assurance regarding the presence or absence of publicly available datasets, such as those from the UCI repository, cannot be guaranteed. This complicates claims regarding model naïveté or independence from training exposure.

Computational demands constitute an additional limitation. Training and deploying LLMs require substantial hardware resources and cloud-based infrastructure, contributing to high costs and limiting accessibility for smaller institutions, particularly in rural or resource-limited healthcare systems. Introduction of new hardware designed for inference will likely mitigate inference issues.

Methodologically, preprocessing steps may introduce bias that affect clustering reliability, and the regional origin of the dataset may impose demographic or epidemiologic skew. Such limitations can restrict generalizability beyond the study population. Technical challenges also emerged in modeling efforts, including instability in support of vector regression (SVR) following one-hot encoding, which expanded the feature space and complicated optimization.

Integrating unstructured clinical data, such as provider notes, imaging reports, and free-text histories, remains a major area for future work. Implicit issues in such data include use of templates and autofill in clinical charts resulting in large amounts of copy forward and duplicate data of less clincal relevance with a few sentences devoted to the key clinical issue at hand. Annotation complexity and interpretive variability limit comprehensive incorporation of these high-value data sources.

Finally, although synthetic datasets allowed systematic evaluation and were accompanied by rigorous diagnostic assessments (e.g., residual analyses to assess model assumptions), they cannot fully replicate the complex noise structures, heteroscedasticity, and irregular error patterns found in real-world clinical data. Consequently, model behavior observed in synthetic environments may overestimate robustness when applied to authentic clinical settings.

3.6. Conclusion

This study highlights the transformative potential of integrating LLMs with machine learning for clinical data analysis. Using geometric evaluation and interpretability tools like SHAP and LIME, we show that even resource-efficient models can match or outperform more complex architectures. This paves the way for scalable, cost-effective, and privacy-preserving healthcare, especially when synthetic data protects patient confidentiality.

However, challenges remain. Variability in model performance, computational demands, and demographic biases limit generalizability. SVR constraints further complicate modeling in high-dimensional, categorical-rich datasets. Future work integrating unstructured clinical text could enhance model depth and interpretability, bridging structured and unstructured data to advance precision medicine through transparent, adaptable, and ethically grounded AI.

Acknowledgments

We are grateful to the University of Virginia Vascular Institute for their support of this study.

Funding Statement

The author(s) declared that financial support was not received for this work and/or its publication.

Footnotes

Edited by: Fahim Sufi, Monash University, Australia

Reviewed by: Marc Leon, Stanford University, United States

Tingyi Wanyan, UTSouthwestern Medical Center, United States

Data availability statement

The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.

Ethics statement

The studies involving humans were approved by West Virginia University Health Sciences Institutional Review Board approval was obtained with waiver of consent (Protocol #1709755537, Approved 6/23/22). The studies were conducted in accordance with the local legislation and institutional requirements. The ethics committee/institutional review board waived the requirement of written informed consent for participation from the participants or the participants’ legal guardians/next of kin because research involves minimal risk and cannot be practicably conducted without the waiver, such as in data reviewed was anonymous.

Author contributions

AA: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. SG: Investigation, Methodology, Supervision, Writing – original draft, Writing – review & editing. SJ: Conceptualization, Methodology, Supervision, Writing – original draft, Writing – review & editing. MA: Formal analysis, Methodology, Project administration, Writing – original draft, Writing – review & editing. SR: Methodology, Writing – original draft, Writing – review & editing. SS: Conceptualization, Supervision, Writing – original draft. JM: Methodology, Resources, Supervision, Writing – original draft, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/frai.2026.1737530/full#supplementary-material

References

  1. Aggarwal C. C., Hinneburg A., Keim D. A. (2001). “On the surprising behavior of distance metrics in high dimensional space” in Database theory — ICDT 2001 (Lecture notes in computer science. Berlin, Heidelberg: Springer Berlin Heidelberg; ), 420–434. doi: 10.1007/3-540-44503-x_27 [DOI] [Google Scholar]
  2. Ahmed S., Nielsen I. E., Tripathi A., Siddiqui S., Ramachandran R. P., Rasool G. (2023). Transformers in time-series analysis: a tutorial. Circuits Syst. Signal Process 42, 7433–7466. doi: 10.1007/s00034-023-02454-8 [DOI] [Google Scholar]
  3. Akkur E. (2023). Prediction of cardiovascular disease based on voting ensemble model and SHAP analysis. Sakarya Univ. J. Comput. Inform. Sci. 6, 226–238. doi: 10.35377/saucis...1367326 [DOI] [Google Scholar]
  4. Ali M. M., Gandhi S., Sulaiman S., Jafri S. H., Ali A. S. (2023). Mapping the heartbeat of America with ChatGPT-4: unpacking the interplay of social vulnerability, digital literacy, and cardiovascular mortality in county residency choices. J. Pers. Med. 13, 1–22. doi: 10.3390/jpm13121625, [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Alomari G., Aljrah I., Aljarrah M., Aljarah A., Aljarah B. (2023). Transforming text generation in NLP: deep learning with GPT models and 2023 twitter corpus using transformer architecture. Int. J. Recent Innov. Trends Comput. Commun. 2023, 3139–3143. doi: 10.17762/ijritcc.v11i9.9463 [DOI] [Google Scholar]
  6. Alum E. U., Ugwu O. P.-C. (2025). Artificial intelligence in personalized medicine: transforming diagnosis and treatment. Discov. Appl. Sci. 7:193. doi: 10.1007/s42452-025-06625-x [DOI] [Google Scholar]
  7. Bender E. M., Gebru T., Mcmillan-Major A., Shmitchell S. (2021). “On the dangers of stochastic parrots: can language models be too big?” in Proceedings of the proceedings of the 2021 ACM conference on fairness, accountability, and transparency (New York, NY: ACM; ). [Google Scholar]
  8. Bhandari R., Alexander T., Annie F. H., Kaleem U., Irfan A., Balla S., et al. (2022). Steep rise in drug use-associated infective endocarditis in West Virginia: characteristics and healthcare utilization. PLoS One 17:e0271510. doi: 10.1371/journal.pone.0271510, [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, et al. (2021). On the opportunities and risks of foundation models. Machine Learning. 1, 1–214. arXiv:2108.07258. doi: 10.48550/arXiv.2108.07258 [DOI] [Google Scholar]
  10. Chafi S., Kabil M., Kamouss A. (2023). Text generator using GPT2 model. Int. J. Sci. Res. Eng. Manag. 7. doi: 10.55041/ijsrem17751 [DOI] [Google Scholar]
  11. Citro R., Chan K.-L., Miglioranza M. H., Laroche C., Benvenga R. M., Furnaz S., et al. (2022). Clinical profile and outcome of recurrent infective endocarditis. Heart 108, 1729–1736. doi: 10.1136/heartjnl-2021-320652, [DOI] [PubMed] [Google Scholar]
  12. Cunningham P, Delany SJ. (2020). K-nearest neighbour classifiers: 2nd edition (with Python examples). doi: 10.1145/3459665 [DOI]
  13. Devlin J, Chang M-W, Lee K, Toutanova K. (2018). BERT: pre-training of deep bidirectional transformers for language understanding. doi: 10.48550/arXiv.1810.04805 [DOI]
  14. Dou J., Wei G., Song Y., Zhou D., Li M. (2024). Switching triple-weight-SMOTE in empirical feature space for imbalanced and incomplete data. IEEE Trans. Autom. Sci. Eng. 21, 1850–1866. doi: 10.1109/tase.2023.3240759 [DOI] [Google Scholar]
  15. Elreedy D., Atiya A. F. (2019). “A novel distribution analysis for SMOTE oversampling method in handling class imbalance” in Lecture notes in computer science. Lecture notes in computer science (Cham: Springer International Publishing; ), 236–248. [Google Scholar]
  16. Fielding-Singh V., Willingham M. D., Fischer M. A., Grogan T., Benharash P., Neelankavil J. P. (2020). A population-based analysis of intraoperative cardiac arrest in the United States. Anesth. Analg. 130, 627–634. doi: 10.1213/ANE.0000000000004477, [DOI] [PubMed] [Google Scholar]
  17. Golovin K. S. (2024). The need for legal regulation of the black box of artificial intelligence. Bull. Kostroma State Univ. 30, 290–297. doi: 10.34216/1998-0817-2024-30-3-290-297 [DOI] [Google Scholar]
  18. Kim H., Lee K. Y., Choo E. H., Hwang B.-H., Kim J. J., Kim C. J., et al. (2024). Long-term risk of cardiovascular death in patients with mildly reduced ejection fraction after acute myocardial infarction: a multicenter, prospective registry study. J. Am. Heart Assoc. 13:e034870. doi: 10.1161/JAHA.124.034870, [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Li J, Cheng X, Zhao WX, Nie J-Y, Wen J-R. (2023). HaluEval: a large-scale hallucination evaluation benchmark for large language models. doi: 10.48550/arXiv.2305.11747 [DOI]
  20. Lundberg S, Lee S-I. (2017). A unified approach to interpreting model predictions. doi: 10.48550/arXiv.1705.07874 [DOI]
  21. Memon S., Wamala R., Kabano I. H. (2023). A comparison of imputation methods for categorical data. doi: 10.1016/j.imu.2023.101382 [DOI]
  22. Merino R., Bowden N., Katamneni S., Coustasse A. (2019). The opioid epidemic in West Virginia. Health Care Manag. 38, 187–195. doi: 10.1097/HCM.0000000000000256 [DOI] [PubMed] [Google Scholar]
  23. Moore SK. (2023). The secret to Nvidia’s AI success. IEEE Spectr. Available online at: https://spectrum.ieee.org/nvidia-gpu (Accessed March 1, 2025).
  24. Mussabayev R. (2024). Optimizing Euclidean distance computation. Mathematics 12:3787. doi: 10.3390/math12233787 [DOI] [Google Scholar]
  25. Nadif M., Role F. (2021). Unsupervised and self-supervised deep learning approaches for biomedical text mining. Brief. Bioinform. 22, 1592–1603. doi: 10.1093/bib/bbab016, [DOI] [PubMed] [Google Scholar]
  26. Nguemeni Tiako M. J., Mori M., Bin Mahmood S. U., Shioda K., Mangi A., Yun J., et al. (2019). Recidivism is the leading cause of death among intravenous drug users who underwent cardiac surgery for infective endocarditis. Semin. Thorac. Cardiovasc. Surg. 31, 40–45. doi: 10.1053/j.semtcvs.2018.07.016 [DOI] [PubMed] [Google Scholar]
  27. Nzeako T. R., Elendu C., Echefu G., Olanisa O., Kiladejo A., Bob-Manuel E. D. (2025). Artificial intelligence in interventional cardiology: a review of its role in diagnosis, decision-making, and procedural precision. Ann. Med. Surg. (Lond.). 5720–5734. doi: 10.1097/ms9.0000000000003602 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Occhipinti S. M. (2024). Advanced research methods for. Appl. Psychol., 211–223. doi: 10.4324/9781003362715-19 [DOI] [Google Scholar]
  29. Parekh A.-D. E., Shaikh O. A., Simran M. S., Hasibuzzaman M. A. (2023). Artificial intelligence (AI) in personalized medicine: AI-generated personalized therapy regimens based on genetic and medical history: short communication. Ann. Med. Surg. (Lond.) 85, 5831–5833. doi: 10.1097/MS9.0000000000001320, [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Passi N, Raj M, Shelke NA. (2024) A review on transformer models: applications, taxonomies, open issues and challenges. doi: 10.1109/ASIANCON62057.2024.10838047 [DOI]
  31. Patil D. (2024). Explainable artificial intelligence (XAI): enhancing transparency and trust in machine learning models.
  32. Pietrolungo C., Gandler A. J. (2023). Now you see it, now you don’t: a case of well-tolerated large septic pulmonary embolism in tricuspid valve endocarditis. Chest 164, A640–A641. doi: 10.1016/j.chest.2023.07.481 [DOI] [Google Scholar]
  33. Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. (2019). Exploring the limits of transfer learning with a unified text-to-text transformer. doi: 10.48550/arXiv.1910.10683 [DOI]
  34. Raiaan M. A. K., Mukta M. S. H., Fatema K., Fahad N. M., Sakib S., Mim M. M. J., et al. (2024). A review on large language models: architectures, applications, taxonomies, open issues and challenges. IEEE Access 12, 26839–26874. doi: 10.1109/access.2024.3365742 [DOI] [Google Scholar]
  35. Ray P., Reddy S. S., Banerjee T. (2021). Various dimension reduction techniques for high dimensional data analysis: a review. Artif. Intell. Rev. 54, 3473–3515. doi: 10.1007/s10462-020-09928-0 [DOI] [Google Scholar]
  36. Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?: explaining the predictions of any classifier.,” Proceedings of the proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. ACM; New York, NY: (2016) [Google Scholar]
  37. Saito T., Rehmsmeier M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One 10:e0118432. doi: 10.1371/journal.pone.0118432, [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Salem M, Mohamed A, Shaalan K. (2025). Transformer models in natural language processing: a comprehensive review and prospects for future development. Proceedings of the 11th International Conference on Advanced Intelligent Systems and Informatics (AISI 2025). Lecture Notes on Data Engineering and Communications Technologies. eds. Hassanien A. E., Rizk R. Y., Darwish A., Alshurideh M. T. R., Snášel V., Tolba M. F. (Springer; ), pp 463–472. doi: 10.1007/978-3-031-81308-5_42 [DOI] [Google Scholar]
  39. Shah R, Silwal S. (2019). Using dimensionality reduction to optimize t-SNE. doi: 10.48550/arXiv.1912.01098 [DOI]
  40. Shah K., Xu A. Y., Sharma Y., Daher M., McDonald C., Diebo B. G., et al. (2024). Large language model prompting techniques for advancement in clinical medicine. J. Clin. Med. 13:5101. doi: 10.3390/jcm13175101, [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. Shindo H, Pfanschilling V, Dhami DS, Kersting K. (2025). Learning differentiable logic programs for abstract visual reasoning. doi: 10.48550/arXiv.2307.00928 [DOI]
  42. Song B., Zhang G., Wang H., Zhu W., Liang Z. (2013). “A dimension reduction strategy for improving the efficiency of computer-aided detection for CT colonography” in Medical imaging 2013: computer-aided diagnosis. eds. Novak C. L., Aylward S. (Piscataway, New Jersey, USA: SPIE; ). [Google Scholar]
  43. Stolear A., Dulgher M., Kaminsky L., Ramponi F., Lancaster G. (2024). Crossroads of care: navigating injection drug use-associated endocarditis. Cureus 16:e62490. doi: 10.7759/cureus.62490, [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Tan C., Shojaei E., Wiener J., Shah M., Koivu S., Silverman M. (2020). Risk of new bloodstream infections and mortality among people who inject drugs with infective endocarditis. JAMA Netw. Open 3:e2012974. doi: 10.1001/jamanetworkopen.2020.12974, [DOI] [PMC free article] [PubMed] [Google Scholar]
  45. Tirthyani D., Kumar S., Vats S. (2024). “Clustering and unsupervised learning” in Advances in systems analysis, software engineering, and high performance computing (Hershey, Pennsylvania, USA: IGI Global; ), 119–139. [Google Scholar]
  46. Topkara V. K., Cheema F. H., Kesavaramanujam S., Mercando M. L., Cheema A. F., Namerow P. B., et al. (2005). Coronary artery bypass grafting in patients with low ejection fraction. Circulation 112, I344–I350. doi: 10.1161/CIRCULATIONAHA.104.526277 [DOI] [PubMed] [Google Scholar]
  47. Uhryn D. I., Kalancha A. D. (2024). Comparison of text information from information sources based on the cosine similarity algorithm. Inf. Culture Technol. 1, 173–177. doi: 10.15276/ict.01.2024.25 [DOI] [Google Scholar]
  48. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. (2017). Attention is all you need. doi: 10.48550/arXiv.1706.03762 [DOI]
  49. Vinci G. U. (2024). Statistical methods in epilepsy. Boca Raton: CRC Press. [Google Scholar]
  50. Wang W, Wei F, Dong L, Bao H, Yang N, Zhou M. (2020). MiniLM: deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv. http://arxiv.org/abs/2002.10957
  51. Wang L, Yang N, Huang X, Jiao B, Yang L, Jiang D, et al. (2022). Text embeddings by weakly-supervised contrastive pre-training. doi: 10.48550/arXiv.2212.03533 [DOI]
  52. Wu L. “Interpretable prediction of heart disease based on random Forest and SHAP.,” In: Sheng H, Dong H, editors. Proceedings of the eighth international conference on electronic technology and information science. Leeds UK: SPIE; (2023). [Google Scholar]
  53. Yang Z, Dai Z, Yang Y, Carbonell J, Salakhutdinov R, Le QV. (2019). XLNet: generalized autoregressive pretraining for language understanding. doi: 10.48550/arXiv.1906.08237 [DOI]
  54. Yang X., Pour Nejatian N., Shin H. C., Smith K., Parisien C., Compas C., et al. (2022). GatorTron: a large clinical language model to unlock patient information from unstructured electronic health records. doi: 10.48550/arXiv.2203.03540 [DOI]
  55. Yenduri G., Ramalingam M., Selvi G. C., Supriya Y., Srivastava G., Maddikunta P. K. R., et al. (2024). GPT (generative pre-trained transformer)— a comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions. IEEE Access 12, 54608–54649. doi: 10.1109/access.2024.3389497 [DOI] [Google Scholar]
  56. Zahariah S., Midi H. (2024). “Identification of high leverage points in high dimensional sparse and non-sparse data” in Statistical outliers and related topics (Boca Raton: CRC Press; ), 414–440. doi: 10.1201/9781003379881-21 [DOI] [Google Scholar]
  57. Zhang Z, Han X, Liu Z, Jiang X, Sun M, Liu Q. (2019) Enhanced language representation with informative entities. doi: 10.48550/arXiv.1905.07129 [DOI]
  58. Zhao J., Papapetrou P., Asker L., Boström H. (2017). Learning from heterogeneous temporal data in electronic health records. J. Biomed. Inform. 65, 105–119. doi: 10.1016/j.jbi.2016.11.006 [DOI] [PubMed] [Google Scholar]
  59. Zriaa R., Amali S. (2021). “A comparative study between K-nearest neighbors and K-means clustering techniques of collaborative filtering in e-learning environment” in Innovations in smart cities applications volume 4. Lecture notes in networks and systems (Cham: Springer International Publishing; ), 268–282. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Data Availability Statement

The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.


Articles from Frontiers in Artificial Intelligence are provided here courtesy of Frontiers Media SA

RESOURCES