Skip to main content
AMIA Annual Symposium Proceedings logoLink to AMIA Annual Symposium Proceedings
. 2026 Feb 14;2025:929–938.

Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models

Aditya Nagori 1,2, Ayush Gautam 1,3, Matthew O Wiens 4,5,6, Vuong Nguyen 4, Nathan Kenya Mugisha 8, Jerome Kabakyenga 9,10, Niranjan Kissoon 4,6,7, John Mark Ansermino 4,5,6, Rishikesan Kamaleswaran 1,2
PMCID: PMC12919534  PMID: 41726472

Abstract

The clustering of patient subgroups is essential for personalized care and efficient use of resources. Traditional clustering methods struggle with high-dimensional heterogeneous healthcare data and lack contextual understanding. This study evaluates clustering based on the Large Language Model (LLM) against classical methods using a pediatric sepsis dataset from a low-income country (LIC), containing 2,686 records with 28 numerical variables and 119 categorical variables. Patient records were serialized into text with and without a clustering objective. Embeddings were generated using quantized LLAMA 3.1 8B, DeepSeek-R1-Distill-Llama-8B with low-rank adaptation(LoRA), and Stella-En-400M-V5 models. K-means clustering was applied to these embeddings. Classical comparisons included K-Medoids clustering on UMAP and FAMD-reduced mixed data. Silhouette scores and statistical tests evaluated the quality and distinctiveness of the cluster. Stella-En-400M-V5 achieved the highest Silhouette Score (0.86). LLAMA 3.1 8B with the clustering objective performed better with a higher number of clusters, identifying subgroups with distinct nutritional, clinical, and socioeconomic profiles. LLM-based methods outperformed classical techniques by capturing richer context and prioritizing key features. These results highlight the potential of LLMs for contextual phenotyping and informed decision making in resource-limited settings.

1. Introduction

Precision medicine aims to tailor healthcare by accounting for individual variability in genes, environment, and lifestyle.1 The key to this approach is to identify patient phenotypes that respond differently to treatments, allowing targeted therapies.2 Unsupervised machine learning, particularly clustering algorithms, is instrumental in discovering these phenotypes within complex biomedical data.3 By grouping patients with similar clinical and molecular profiles, clustering supports tailored treatment plans and advances personalized care.4

Clustering in healthcare spans from patient stratification by disease subtype to categorizing medical literature for evidence synthesis.5 However, classical algorithms such as K-means and hierarchical clustering face challenges when applied to healthcare data. Medical datasets often include mixed numerical (e.g., lab results) and categorical (e.g., diagnosis codes) variables, and traditional methods usually require data transformation or dissimilarity measures that may not capture true relationships.6, 7 Furthermore, high-dimensional data with numerous patient variables exacerbate the “curse of dimensionality,” diminishing the meaningfulness of distance metrics and increasing computational complexity.8, 9, 10 Additionally, classical clustering may yield clusters that are difficult to interpret for clinical decision-making.11

Large Language Models (LLMs), such as GPT-4, have shown impressive capabilities in understanding and generating human-like text.12 Their ability to process unstructured data and generate embeddings offers new opportunities for analyzing complex healthcare information.13 By converting mixed-type data into text and producing embeddings, LLMs can unify numerical and categorical data into a single space that captures inter-variable relationships.14 They also reduce dimensionality by distilling high-dimensional data into lower-dimensional representations that retain essential information.15 Furthermore, LLMs capture nonlinear interactions often missed by classical methods and can enhance interpretability through explainable AI techniques.16, 17

Despite these advantages, applying traditional clustering algorithms directly to LLM-generated embeddings in health-care presents challenges.18 Issues include managing sensitive patient information, ensuring privacy, and maintaining clinical interpretability. Transparent clustering is vital for clinical translation; clinicians must understand what distinguishes patient subgroups based on clinical features or trajectories. Classical methods often act as “black boxes” without clear explanations, limiting adoption where clusters must be both clinically meaningful and statistically robust.

To address these gaps, we propose a novel pipeline that leverages LLM-based dynamic embedding generation, context-aware dimensionality reduction, and phenotyping algorithms. Unlike static embeddings, our approach emphasizes features relevant to clustering outcomes. We demonstrate the feasibility of our pipeline using a pediatric sepsis cohort, a compelling case given its clinical heterogeneity, temporal complexity, and need for early intervention. Our pipeline serializes mixed-type EHR data in text, generates embeddings with LLMs, and employs dimensionality reduction to preserve semantic structure. We then evaluate clustering performance and the clinical coherence of the resulting phenotypes against traditional methods, advancing the integration of foundation models into clinical data science.

2. Methods

2.1. Study Population and Data Collection

We used a synthetic dataset19 based on a prospective multisite observational cohort study of children aged 6 to 60 months with suspected sepsis admitted to Ugandan hospitals between 2017 and 202020 The synthetic dataset was generated using the synthpop package in R, applying the nonparametric classification and regression tree (CART) method for variable synthesis.21 The dataset included demographics, vital signs, lab values, symptoms, comorbidities, medications, socio-environmental factors, and outcomes such as mortality and length of stay.

2.2. Study Design

This secondary analysis aimed to phenotype pediatric patients with sepsis using advanced machine learning techniques, including transformer-based models and clustering methods.

2.3. Data Serialization

To convert structured data into a format compatible with LLM input, data serialization was performed. Let X be a dataset with n samples (patient records) and m features (variables). Each patient record xi is represented as:

xi=(xi1,xi2,,xim),

where xij represents the value of the j-th feature for the i-th patient.

To serialize the data, each record xi is converted into a serialized text format si:

si=serialize(xi)=concatenate(xi1,xi2,,xim).

2.4. Embedding Generation Using Language Models

To represent patient data in a high-dimensional latent space that encodes semantic and clinical patterns, we utilized three advanced language models to generate embeddings from serialized patient data. The first two models, the quantized LLama 3.1 8B model22 and the DeepSeek-R1-Distill-Llama-8B model,23, 24 were used with 4-bit quantization and accessed via the unsloth Python library with a LoRA25 adapter. Let the serialized patient record for the i-th patient be denoted as si; the corresponding embedding is generated as:

ei,llm=fllm(si;θllm),

where fllm(·; θllm) represents either the LLama 3.1 8B or DeepSeek-R1-Distill-Llama-8B model, and ei,llm ∈ ℝ4096 is the high-dimensional embedding for the i-th patient. Quantization improves computational efficiency, while the LoRA adapter minimizes the number of parameters updated during fine-tuning.

The third model, Stella-en-400M-v5, was accessed via the sentence_transformers Python library. Denoting its function as fStella(·; θStella), the embedding for the i-th patient is generated as:

ei,Stella=fStella(si;θStella),

with ei,Stella ∈ ℝ1024. Together, these models provide complementary embeddings that capture both task-specific and general semantic features of patient data.

2.5. Incorporation of Clustering Objective

To prioritize critical features for clustering, we appended the clustering objective to each serialized patient record. The clustering objective was: “Generate an embedding for clustering patients based on their physiological severity and prioritize features indicative of critical conditions.” This was developed through expert consultation with pediatric clinicians involved in the original study. This objective guided the language model to generate embeddings tailored for clustering tasks. Formally, let the serialized patient record for the i-th patient be si, and let the clustering objective be represented as O. The modified serialized record si is:

si=concatenate(si,O).

The sentence transformer model f (·; θ), parameterized by θ, generates the embedding ei for the modified record si :

ei=f(si;θ).

Here si is the original serialized text of patient data. O is the clustering objective appended to the serialized text. f (·; θ) is the language model that generates embeddings. eid is the resulting task-specific embedding in a d-dimensional space.

By incorporating O into si, the embedding process explicitly emphasizes features critical for clustering tasks, aligning the generated embeddings with the clustering objectives.

2.6. Workflow Summary

Patient records si were serialized into text representations. Embeddings ei,llm and ei,Stella were generated with and without the clustering objective appended. UMAP reduced embeddings ei to lower-dimensional representations zi. K-means clustering was applied to {zi}i=1n to partition the dataset into k clusters. The clustering performance was evaluated using silhouette scores.

2.7. Classical Clustering Approaches for Comparison

2.7.1. Clustering on Mixed Data UMAP Embeddings

To reduce dimensionality while preserving the global structure of the data, we applied Uniform Manifold Approximation and Projection (UMAP) separately to numerical and categorical features.26 For categorical data, UMAP used the Dice similarity metric. The resulting numeric and categorical embeddings were then concatenated to form a unified representation. K-Medoids clustering was applied to this combined embedding using Euclidean distance, and silhouette scores were calculated for cluster counts ranging from 2 to 9.

2.7.2. Clustering on FAMD Embeddings

To provide a classical low-rank representation of patient profiles while accommodating both numerical and categorical variables, we applied Factor Analysis of Mixed Data (FAMD) using the Python prince library.27, 28 This dimensionality-reduced embedding was then clustered using K-Medoids, and silhouette scores were calculated for cluster counts ranging from 2 to 9.

2.8. Descriptive Statistics and Statistical Testing

The median and interquartile range (IQR) were calculated for continuous variables (e.g., age, height, weight, MUAC, heart rate (HR)). Frequencies and percentages were calculated for categorical variables (e.g., symptoms, comorbidities, vaccination status). We performed the Kruskal-Wallis test29 to compare continuous variables and the chi-square test to compare categorical variables between clusters, applying a Bonferroni correction to control for multiple comparisons and to determine the statistical significance of observed differences.

3. Results

3.0.1. Performance Comparison of Clustering Algorithms

The clustering performance of various algorithms was assessed using Silhouette Scores across different numbers of clusters. The algorithms evaluated included Stella-en-400M-v5, Llama 3.1 8b, DeepSeek-R1-Distil-Llama-8b and their variants with incorporated clustering objectives, and K-Medoids clustering applied in UMAP and FAMD embeddings.

Llama 3.1 8b performed well, with a Silhouette Score of 0.74 with two clusters (Fig. 1b, Fig. 2) , decreasing to 0.52 with nine clusters. When incorporating the clustering objective (Llama 3.1 8b with the objective), the algorithm showed improved performance at higher cluster numbers (Fig. 1a) (Fig. 2). The Silhouette score peaked at 0.71 with five clusters (Fig. 2).

Fig. 1.

Fig. 1.

Clustering Performance Metrics Across Different Models Fig 1a). Clusters on Llama 3.1 8b model embeddings with objectives inserted in the serialized data. Fig 1b). Clusters on Llama 3.1 8b model embeddings. Fig 1c). Clusters on Stella model embeddings with objectives inserted in the serialized data. Fig 1d). Clusters on Stella model embeddings. Fig 1e). Clusters on DeepSeek-R1-Distill-Llama-8b model embeddings with objectives inserted in the serialized data. Fig 1f). Clusters on DeepSeek-R1-Distill-Llama-8b model embeddings. Fig 1g.) K-Medoids clusters on mixed data Umap Embeddings. Fig 1h.) K-Medoids clusters on FAMD Embeddings.

Fig. 2.

Fig. 2.

Silhouette score vs number of clusters

Stella-en-400M-v5 achieved the highest Silhouette Score of 0.86 with two clusters (Fig.1d, Fig.2), indicating well-defined and cohesive clustering. However, the Silhouette Score decreased as the number of clusters increased, reaching 0.39 at nine clusters (Fig. 2).

Similarly, Stella-en-400M-v5 with a clustering objective (Fig.1c) maintained high Silhouette Scores across cluster numbers, starting at 0.85 for two clusters and decreasing to 0.44 for nine clusters (Fig. 2). Although slightly lower than it’s without objective variant at two clusters, it demonstrated better performance at higher cluster counts, suggesting that incorporating clustering objectives helps maintain cluster integrity as complexity increases.

DeepSeek-R1-Distil-Llama-8b produced a Silhouette Score of 0.82 with two clusters and 0.43 at nine clusters (Fig.1f, Fig.2). With the clustering objective added, the model peaked at two clusters with a Silhouette score of 0.83 (Fig.1e, Fig.2), showing no major improvement over it’s without objective variant.

The K-Medoids clustering on UMAP embeddings (Fig. 1g) resulted in moderate Silhouette Scores ranging from 0.37 to 0.34 across two to nine clusters (Fig. 2). On FAMD embeddings (Fig. 1h), K-Medoids showed slightly better results, with the highest Silhouette Score of 0.44 at four clusters (Fig. 2), suggesting improved cluster compactness and separation at this cluster number.

Overall, the LLM-based embeddings, particularly when combined with clustering objectives, performed better than the K-Medoids methods on both UMAP and FAMD embeddings (Fig. 1g, 1h). Incorporating clustering objectives into the embedding generation process enhanced the ability of the models to produce task-specific embeddings, leading to better defined clusters. Stella-en-400M-v5 and Llama 3.1 8b with objective demonstrated superior clustering performance, with the former excelling at lower cluster numbers Fig. 1c) and the latter showing optimal results at five clusters Fig. 1a).

3.0.2. Descriptive Statistics and Group Comparisons

The clustering analysis reveals five distinct clusters using LLama 3.1 8b model embeddings (Cluster 0 to Cluster 4) characterized by various demographic, clinical and socioeconomic variables. Statistical significance among clusters was assessed using the Kruskal-Wallis test for continuous variables and the Chi-Square test for categorical variables.

Cluster 0 includes children with a median (IQR) age at admission of 19.05 (13.3) months (p-value(p) = 0). The growth z-scores i.e. Weight-for-Length (WFL) z-score: -1.03 (2.27), p = 0, Body Mass Index (BMI) z-score: -0.93 (2.46), p = 0, Weight-for-Age (WFA) z-score: -1.28 (2.08), p = 0, indicating slightly underweight children for their age and height. A notable characteristic is the high usage of intravenous ampicillin or amoxicillin antibiotics, with 99.73% of children receiving these medications (p = 0). Families in cluster 0 have moderate access to municipal water (46.29%, p = 0), and a higher percentage use water purification methods (80.49%, p = 0), indicating slightly better socioeconomic conditions. The median hematocrit level is higher at 36% (p = 0), which may reflect better overall health. Vital signs are stable, with a median (IQR) heart rate (HR) of 142 (29) bpm and a respiratory rate (RR) of 44 (20) breaths per minute (brpm) (both p = 0).

Cluster 1 represents younger children, median (IQR) age of 12.10 (13.9) months (p = 0). They have poorer nutrition, i.e. WFL z-score (-0.99 (2.7), p = 0), BMI z-score (-0.98 (2.65), p = 0), WFA z-score (-1.13 (2.31), p =0) and higher vital signs (HR 149 (31) bpm, RR 47 (20) brpm, all p = 0). This cluster has the highest in-hospital mortality rate at 6.56% (p = 0.001). Clinical symptoms such as severe respiratory distress are significantly more prevalent (24.29%, p = 0), and there is a higher incidence of coma (4.96%, p = 0). Measles vaccination rates are significantly lower, with only 24.82% vaccinated (p = 0), indicating a gap in preventive healthcare. Mothers in cluster 1 have lower educational levels (p = 0.0001), which can contribute to poorer health outcomes in children.

Cluster 2 has the oldest children, median (IQR) age 21.30 (16.75) months (p = 0). Their z-scores, i.e. WFA z-score: -0.99 (1.84), p = 0, BMI z-score: -0.81 (2.2), p = 0, and WFL z-score: -0.89 (2.1), p = 0 are closest to normal. Vital signs are most stable (HR 139 (28) bpm, RR 43 (17) brpm, p=0). Significantly, cluster 2 has the highest percentage of malaria-positive cases (36.42%, p = 0). The primary water source for families in this cluster is boreholes (31.56%, p = 0). Additionally, children in this cluster have longer durations of exclusive breastfeeding, with 50.38% exclusively breastfed for 6 months (p = 0), indicating better early-life nutrition practices.

Cluster 3 had the youngest children, median (IQR) age 10.45 (10.65) months (p = 0), and the poorest nutritional status: i.e. WFL z-score: -1.26 (2.23), p = 0, BMI z-score: -1.31 (2.38), p=0, WFA z-score: -1.55 (2.50), p =0. The cluster exhibits higher rates of prolonged diarrhea (41.57%, p = 0.0008) and symptoms like rash (23.37%, p = 0), indicating more severe and perhaps chronic health issues. Mothers in Cluster 3 have the lowest maternal education levels (p = 0.0001) and higher rates of unknown age at first pregnancy (p = 0). There is also a significant percentage of premature births (4.41%, p = 0.026) and low birth weight infants (4.60%, p = 0.0025). These factors contribute to the overall vulnerability of children in this cluster.

Cluster 4 has median (IQR) age at admission of 18.40 (15.6) months (p = 0), similar to Cluster 0. Nutritional status is moderate, with the growth WFL z-score: -0.96 (2.19), p = 0, BMI z-score: -0.93 (2.20), p =0, and WFA z-score: -1.13 (1.92), p = 0. A distinctive feature of Cluster 4 is a higher incidence of symptoms, such as changes in urine color (19.72%, p = 0.0001). Additionally, there is a slightly higher usage of HIV medications (0.94%, p = 0.0039). Families in Cluster 4 have moderate access to municipal water (44.13%, p = 0) and water purification methods (75.59%, p = 0). The median respiratory rate is the lowest among clusters at 43 (20) breaths per minute (p = 0), indicating stable vital signs.

4. Discussion

The use of Large Language Models (LLMs), particularly the LLAMA 3.1 8B model, in clustering structured health-care data represents a significant advance over traditional methods that struggle with high-dimensional, heterogeneous datasets.9 LLMs leverage deep architectures and non-linear activations to capture complex patterns and relationships,30, 31, 32 which is crucial for healthcare data with mixed types and intricate interactions.

Our pipeline incorporated the dynamic embedding generation by incorporation of clustering objectives. We embedded clustering objectives into the prompts used for generating embeddings, ensuring that the embeddings are task-specific and capture relevant features for clustering. We introduced the task-Specific embeddings with LLMs. By utilizing an LLM, we generated embeddings that unify mixed data types and capture complex relationships, addressing challenges with traditional methods. We resolved the scalability and efficiency issue by implementing mini-batch processing and parallel computing to handle large datasets efficiently, addressing computational challenges.

LLMs are pre-trained on vast amounts of data, allowing them to incorporate contextual understanding into their analysis.33 We used a quantized Llama 3.1 8B and DeepSeek-R1-Distill-Llama-8b Models with LoRA Adapter, the these model were chosen for its open-access weights and state-of-the-art performance. We applied quantization techniques to reduce the model size and computational requirements. Low-Rank Adaptation (LoRA) was used to fine-tune the model efficiently on our dataset without updating all model parameters.25 The model generated embeddings for each serialized record with and without the clustering objective appended. We also used the Stella-En-400M-V5 Model, the Stella-En-400M-V5 model was selected due to its high performance on the Massive Text Embedding Benchmark (MTEB) leaderboard and smaller model size. Importantly, it is trained based on Alibaba-NLP/gte-large-en-v1.5 and Alibaba-NLP/gte-Qwen2-1.5B-instruct which is trained on 75 languages, including Ugandan languages, making it well-suited for our dataset.34 Like the LLAMA model, embeddings were generated with and without the clustering objective. Although our dataset is structured, the LLAMA model’s ability to recognize patterns and relationships beyond numerical proximity enhances clustering results. This contextual awareness can lead to more meaningful and clinically relevant clusters.

The LLAMA 3.1 8B model identified five distinct patient clusters with statistically significant differences in demo-graphic, clinical, and socioeconomic variables. Cluster 1 consisted of the youngest patients with the highest in-hospital mortality rate. They exhibited poorer nutritional status and more severe clinical symptoms, such as severe respiratory distress and coma. These findings align with previous studies indicating that younger children with malnutrition are at higher risk of adverse outcomes.35, 36 Cluster 2 included older children with better nutritional indicators but a high incidence of malaria-positive cases. This suggests that while nutritional status is better, exposure to malaria remains a significant health concern, consistent with epidemiological data from malaria-endemic regions.37 Cluster 3 featured young children with the poorest nutritional indicators and higher prevalence of health issues, such as prolonged diarrhea and rash. Maternal factors, including lower education levels and higher rates of premature births, may contribute to the health challenges in this cluster, echoing findings that maternal education significantly impacts child health outcomes.38, 39, 40 Cluster 0 represented children with stable clinical presentations. The high usage of standard antibiotics and better access to municipal water indicate favorable socioeconomic conditions, which are known to correlate with improved health outcomes.40 Cluster 4 was similar to Cluster 0 but exhibited specific clinical concerns, such as higher incidences of changes in urine color and potential HIV exposure. This underscores the importance of monitoring for renal conditions and addressing HIV-related healthcare needs.41

These results suggest that LLM-based clustering not only simplifies the analysis of complex healthcare data but also provides actionable insights for targeted interventions and resource allocation.42 Observed associations—such as those between maternal education and child health—underscore the importance of educational initiatives.43 However, high computational requirements and the “black-box” nature of transformer models may limit clinical transparency and acceptance.11

4.1. Conclusion

LLM-driven representation learning and clustering demonstrate strong potential for phenotyping pediatric sepsis in low-resource settings by effectively managing high-dimensional, heterogeneous data and capturing nonlinear, contextrich patterns. The identified patient subgroups reveal distinct clinical, nutritional, and socioeconomic profiles, offering practical implications for real-time triage, treatment prioritization, and equitable resource allocation. The use of a clustering objective in guiding embedding generation, developed with clinical input, enhanced the relevance of clusters to physiological severity. While results are promising, future validation on real-world clinical data is essential to ensure generalizability. Further efforts should also focus on improving model efficiency, interpretability, and integration into clinical workflows to support decision-making in resource-constrained environments.

Figures & Tables

Table 1:

Cluster-level summary of clinical, laboratory, and socioeconomic characteristics in a pediatric sepsis cohort. Data for each variable are presented across five patient clusters with corresponding p-values (p) indicating statistical significance.

Variable Cluster 0 Cluster 1 Cluster 2 Cluster 3 Cluster 4 p-value
Anthropometrics
Age at Admission (months) 19.05 (13.30) 12.10 (13.90) 21.30 (16.75) 10.45 (10.65) 18.40 (15.60) 0
Height (cm) 78.90 (11.62) 73.35 (14.50) 82.00 (13.00) 71.90 (11.20) 80.00 (15.00) 0
Weight (kg) 9.50 (3.00) 8.70 (2.80) 10.00 (2.65) 7.80 (2.79) 9.50 (3.00) 0
MUAC (mm) 140.00 (18.00) 137.00 (21.00) 142.00 (18.00) 134.00 (20.00) 140.00 (16.00) 0
Vital Signs
Heart Rate (bpm) 142.00 (29.00) 149.00 (31.00) 139.00 (28.00) 146.50 (33.00) 143.00 (27.00) 0
Respiratory Rate (breaths/min) 44.00 (20.00) 47.00 (20.00) 43.00 (17.00) 47.50 (19.25) 43.00 (20.00) 0
Hematocrit (%) 36.00 (10.00) 32.00 (13.75) 32.00 (16.00) 34.00 (8.00) 34.00 (10.00) 0
Oxygen Measures
Oxygen Saturation (1st measure) 97.00 (5.00) 97.00 (7.00) 97.00 (4.75) 97.00 (8.00) 97.00 (5.00) 1e-04
Oxygenation Index (sqi2 perc oxi adm) 88.00 (23.75) 84.50 (33.25) 89.00 (27.00) 85.00 (31.00) 86.00 (30.25) 2e-04
SpO2 Measured on Oxygen 9.07% 25.35% 10.02% 14.94% 12.21% 0
O2 available & used 14.84% 29.43% 14.57% 22.61% 15.96% 0
O2 available & not used 80.91% 65.78% 81.34% 72.99% 78.87% 0
Severe respiratory distress 10.99% 24.29% 9.56% 19.16% 10.33% 0
Capillary refill time 7.55% 14.72% 14.57% 11.69% 10.33% 2e-04
Blantyre Coma Scale (BCS)
BCS Eye: fails to watch/follow 0.96% 22.87% 0.91% 1.15% 6.10% 0
BCS Eye: watches/follows 99.04% 77.13% 99.09% 98.85% 93.90% 0
BCS Motor: localizes pain 97.39% 73.40% 96.66% 95.40% 93.43% 0
BCS Motor: no/inappropriate response 0.00% 6.21% 0.15% 0.00% 0.47% 0
BCS Motor: withdraws from pain 2.61% 20.39% 3.19% 4.60% 6.10% 0
BCS Verbal: appropriate cry/speech 99.31% 74.47% 99.54% 99.04% 96.24% 0
BCS Verbal: moan/abnormal cry 0.69% 18.97% 0.46% 0.96% 3.76% 0
BCS Verbal: no response 0.00% 6.56% 0.00% 0.00% 0.00% 0
Antibiotics on Admission
IV Ampicillin/Amoxicillin 99.73% 13.83% 0.00% 94.83% 51.64% 0
PO Ampicillin/Amoxicillin 0.14% 3.90% 1.97% 1.34% 0.00% 0
IV Penicillin 0.41% 24.29% 25.64% 0.77% 15.02% 0
IV Ceftriaxone 3.98% 55.50% 63.28% 4.60% 30.99% 0
IV Gentamicin 91.07% 56.56% 47.04% 92.34% 66.67% 0
IV/PO Antimalarial 18.68% 30.32% 33.84% 15.52% 19.72% 0
Vaccination Status
Measles vacc unknown 0.00% 0.35% 0.00% 10.73% 2.35% 0
Measles vacc yes 100.00% 24.82% 99.09% 4.02% 65.26% 0
Source measles vacc (none) 0.00% 75.18% 0.91% 95.98% 34.74% 0
Source measles vacc: card 16.62% 3.90% 16.54% 0.57% 8.92% 0
Source measles vacc: self-report 83.38% 20.92% 82.55% 3.45% 56.34% 0
Pneumo vacc 0 doses 0.27% 8.51% 0.61% 6.90% 1.88% 0
Pneumo vacc 1 dose 0.27% 5.32% 0.61% 6.13% 3.29% 0
Pneumo vacc 2 doses 7.42% 14.89% 5.77% 13.22% 11.27% 0
Pneumo vacc 3 doses 92.03% 68.97% 93.02% 60.34% 79.81% 0
Pneumo vacc unknown 0.00% 2.30% 0.00% 13.41% 3.76% 0
Source pneumo vacc (none) 0.00% 10.64% 0.15% 20.50% 5.63% 0
Source pneumo vacc: self-report 83.38% 69.50% 83.61% 63.22% 81.22% 0
DPT/Penta vacc 0 doses 0.27% 4.79% 0.46% 6.13% 1.41% 0
DPT/Penta vacc 1 dose 0.27% 5.32% 0.30% 5.17% 2.82% 0
DPT/Penta vacc 2 doses 7.42% 15.25% 5.61% 13.22% 11.27% 0
DPT/Penta vacc 3 doses 92.03% 73.05% 93.63% 63.03% 81.22% 0
DPT/Penta vacc unknown 0.00% 1.60% 0.00% 12.45% 3.29% 0
Source DPT/Penta vacc (none) 0.00% 6.21% 0.00% 18.20% 4.69% 0
Source DPT/Penta vacc: self-report 83.79% 74.11% 83.76% 65.52% 81.69% 0
Recent Medication Usage
Antibiotics used in past week: don’t know 0.00% 6.38% 0.00% 0.00% 0.00% 0
Antibiotics used in past week: no 51.10% 41.31% 51.90% 42.72% 45.07% 1e-04
Antimalarials used in past week: don’t know 0.00% 5.67% 0.00% 0.00% 0.00% 0
Antimalarials used in past week: no 72.25% 58.33% 65.10% 69.16% 63.85% 0
Symptoms
Rash 7.28% 15.60% 7.44% 23.37% 10.80% 0
Urine color changes 10.71% 17.02% 18.36% 12.84% 19.72% 1e-04
Coma 0.14% 4.96% 0.30% 0.38% 0.00% 0
Co-morbid dx unknown 1.65% 5.32% 1.21% 6.51% 1.41% 0
Feeding/Breastfeeding
Feeding poorly 77.75% 62.23% 75.72% 75.67% 74.65% 0
Not feeding 3.57% 21.63% 5.16% 4.98% 4.69% 0
No cough/choke with liquids 47.80% 38.30% 44.16% 38.12% 51.17% 1e-04
6 months exclusive BF 54.40% 45.21% 50.38% 40.80% 51.17% 0
Unknown exclusive BF 1.10% 1.42% 1.21% 7.28% 2.82% 0
Total BF >12 months 41.48% 24.11% 48.71% 19.35% 38.50% 0
Still being breastfed 38.19% 59.57% 30.35% 60.34% 40.38% 0
Unknown BF duration 0.69% 0.71% 1.06% 3.45% 0.94% 2e-04
Travel
Motorcycle travel 55.77% 42.02% 41.43% 53.45% 49.77% 0
Private vehicle travel 1.92% 2.66% 5.92% 1.92% 5.63% 0
Taxi/special hire travel 39.01% 49.29% 48.56% 40.61% 42.25% 2e-04
Other Clinical/Maternal Info
Good health prior to illness 87.64% 83.69% 86.49% 80.08% 92.49% 0
Mother’s age known 97.53% 97.34% 96.36% 91.00% 95.31% 0
Mother’s age at first pregnancy known 95.19% 91.31% 92.26% 84.48% 92.02% 0
Mother’s education unknown 0.69% 1.24% 0.00% 2.87% 0.94% 1e-04
Maternal HIV unknown 4.95% 5.32% 5.92% 11.69% 7.51% 0
Household/Environment
Borehole water source 18.96% 28.37% 31.56% 19.54% 22.07% 0
Uses water purification 80.49% 61.35% 64.34% 74.52% 75.59% 0
Malaria
Malaria test positive 25.14% 31.74% 36.42% 20.50% 26.76% 0
Admission/Nutritional Indices
Length of admission (days) 4.00 (3.00) 4.00 (4.00) 4.00 (4.00) 5.00 (4.75) 4.00 (4.00) 0
Weight-for-length Z-score -1.03 (2.27) -0.99 (2.70) -0.89 (2.10) -1.26 (2.23) -0.96 (2.19) 0
BMI Z-score -0.93 (2.46) -0.98 (2.65) -0.81 (2.20) -1.31 (2.38) -0.93 (2.20) 0
Weight-for-age Z-score -1.28 (2.08) -1.13 (2.31) -0.99 (1.84) -1.55 (2.50) -1.13 (1.92) 0

References

  • 1.Collins FS, Varmus H. A new initiative on precision medicine. New England Journal of Medicine. 2015;372(9):793–5. doi: 10.1056/NEJMp1500523. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Ashley EA. Towards precision medicine. Nature Reviews Genetics. 2016;17(9):507–22. [Google Scholar]
  • 3.Kourou K, Exarchos TP, Exarchos KP, Karamouzis MV, Fotiadis DI. Machine learning applications in cancer prognosis and prediction. Computational and Structural Biotechnology Journal. 2015;13:8–17. doi: 10.1016/j.csbj.2014.11.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Shah NH, Tenenbaum JD. The coming age of data-driven medicine: translational bioinformatics’ next frontier. Journal of the American Medical Informatics Association. 2012 Jun;19(e1):e2–4. doi: 10.1136/amiajnl-2012-000969. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wang Y, Wang L, Rastegar-Mojarad M, Moon S, Shen F, Afzal N, et al. Clinical information extraction applications: A literature review. Journal of Biomedical Informatics. 2018 Jan;77:34–49. doi: 10.1016/j.jbi.2017.11.011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Boriah S, Chandola V, Kumar V. Similarity measures for categorical data: A comparative evaluation. Proceedings of the 8th SIAM International Conference on Data Mining. SIAM. 2008:p. 243–54. [Google Scholar]
  • 7.Ahmad A, Dey L. A k-mean clustering algorithm for mixed numeric and categorical data. Data & Knowledge Engineering. 2007 Nov;63(2):503–27. [Google Scholar]
  • 8.Bellman R. Princeton Legacy Library. Princeton, NJ: Princeton University Press; 1961. Adaptive Control Processes: A Guided Tour. [Google Scholar]
  • 9.Aggarwal CC, Hinneburg A, Keim DA. Database Theory — ICDT 2001. vol. 1973 of Lecture Notes in Computer Science. Springer; 2001. On the surprising behavior of distance metrics in high dimensional space; pp. p. 420–34. [Google Scholar]
  • 10.Dean J, Ghemawat S. MapReduce: simplified data processing on large clusters. Communications of the ACM. 2008 Jan;51(1):107–13. [Google Scholar]
  • 11.Doshi-Velez F, Kim B. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:170208608. 2017.
  • 12.Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems. 2020;33:p. 1877–901. [Google Scholar]
  • 13.Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of NAACL-HLT (1) 2019:p. 4171–86. [Google Scholar]
  • 14.Zerveas G, Jayaraman S, Patel D, Bhamidipaty A, Eickhoff C. A Transformer-Based Framework for Multivariate Time Series Representation Learning. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD ‘21). KDD ‘21. 2021:p. 2114–24. [Google Scholar]
  • 15.Mikolov T, Chen K, Corrado G, Dean J. Efficient Estimation of Word Representations in Vector Space. 1st International Conference on Learning Representations (ICLR) Workshop Track. 2013.
  • 16.Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS ‘17) 2017;30:p. 5998–6008. [Google Scholar]
  • 17.Ribeiro MT, Singh S, Guestrin C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16) 2016:p. 1135–44. [Google Scholar]
  • 18.Esteva A, Robicquet A, Ramsundar B, Kuleshov V, DePristo M, Chou K, et al. A Guide to Deep Learning in Healthcare. Nature Medicine. 2019;25(1):24–9. [Google Scholar]
  • 19.Huxford C, Rafiei A, Nguyen V, Wiens MO, Ansermino JM, Kissoon N, et al. The 2024 Pediatric Sepsis Challenge: Predicting In-Hospital Mortality in Children With Suspected Sepsis in Uganda. Pediatric Critical Care Medicine. 2024 Nov;25(11):1047–50. doi: 10.1097/PCC.0000000000003556. Epub ahead of print June 21, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Wiens MO, Bone JN, Kumbakumba E, Businge S, Tagoola A, Sherine SO, et al. Mortality after hospital discharge among children younger than 5 years admitted with suspected sepsis in Uganda: a prospective, multisite, observational cohort study. The Lancet Child & Adolescent Health. 2023 Aug;7(8):555–66. doi: 10.1016/S2352-4642(23)00052-4. Epub May 11, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Nowok B, Raab GM, Dibben C. synthpop: Bespoke Creation of Synthetic Data in R. Journal of Statistical Software. 2016 Oct;74(11):1–26. [Google Scholar]
  • 22.Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA 3.1: Open and Efficient Foundation Language Models. 2024. Accessed: 2024-11-28. https://github.com/facebookresearch/llama.
  • 23.deepseek-ai/DeepSeek-R1-Distill-Llama-8B – Hugging Face. Published: 3 months ago. https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B.
  • 24.Guo D, Yang D, Zhang H, Song J, Zhang R, et al. DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:250112948. 2025 Jan. Available from: https://arxiv.org/abs/2501.12948.
  • 25.Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, et al. LoRA: Low-Rank Adaptation of Large Language Models. arXiv preprint arXiv:210609685. 2021 Jun.
  • 26.McInnes L, Healy J, Melville J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv preprint arXiv:180203426. 2018 Feb.
  • 27.Halford M. MIT License; 2024. Prince: Multivariate Exploratory Data Analysis in Python—PCA, CA, MCA, MFA, FAMD, GPA. original release 2016-10-22. https://github.com/MaxHalford/prince. [Google Scholar]
  • 28.Pagès J. Boca Raton, FL: Chapman & Hall/CRC; 2014. Multiple Factor Analysis by Example Using R. [Google Scholar]
  • 29.Kruskal WH, Wallis WA. Use of Ranks in One-Criterion Variance Analysis. Journal of the American Statistical Association. 1952 Dec;47(260):583–621. doi: 10.1080/01621459.1952.10483441. Available from: [DOI] [Google Scholar]
  • 30.Naveed H, Khan AU, et al. A Comprehensive Overview of Large Language Models. arXiv preprint arXiv:2307.06435. 2024. Available from: http://arxiv.org/abs/2307.06435.
  • 31.Ester M, Kriegel HP, Sander J, Xu X. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining (KDD ‘96) 1996:226–31. [Google Scholar]
  • 32.LeCun Y, Bengio Y, Hinton G. Deep Learning. Nature. 2015 May;521(7553):436–44. doi: 10.1038/nature14539. Available from: [DOI] [PubMed] [Google Scholar]
  • 33.Shah K, Xu AY, Sharma Y, Daher M, McDonald C, Diebo BG, et al. Large Language Model Prompting Techniques for Advancement in Clinical Medicine. Journal of Clinical Medicine. 2024 Aug;13(17):5101. doi: 10.3390/jcm13175101. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Zhang X, Zhang Y, Long D, Xie W, Dai Z, Tang J, et al. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. arXiv preprint arXiv:2407.19669. 2024. Available from: http://arxiv.org/abs/2407.19669.
  • 35.Black RE, Victora CG, Walker SP, Bhutta ZA, Christian P, de Onis M, et al. Maternal and Child Undernutrition and Overweight in Low-Income and Middle-Income Countries. The Lancet. 2013 Aug;382(9890):427–51. doi: 10.1016/S0140-6736(13)60937-X. Available from: [DOI] [Google Scholar]
  • 36.Pelletier DL, Frongillo EA. The Effects of Malnutrition on Child Mortality in Developing Countries. Bulletin of the World Health Organization. 1995;73(4):443–8. [PMC free article] [PubMed] [Google Scholar]
  • 37.World Health Organization. World Health Organization; 2020. World Malaria Report 2020. Available from: https://www.who.int/teams/global-malaria-programme/reports/world-malaria-report-2020. [Google Scholar]
  • 38.Caldwell JC. Education as a Factor in Mortality Decline: An Examination of Nigerian Data. Population Studies. 1979 Nov;33(3):395–413. doi: 10.1080/0032472031000143635. Available from: [DOI] [Google Scholar]
  • 39.Cleland JG, Van Ginneken J. Maternal Education and Child Survival in Developing Countries: The Search for Pathways of Influence. Social Science & Medicine. 1988 Dec;27(12):1357–68. doi: 10.1016/0277-9536(88)90201-8. Available from: https://doi.org/ 10.1016/0277-9536(88)90253-0. [DOI] [PubMed] [Google Scholar]
  • 40.Victora CG, Adair L, Fall C, Hallal PC, Martorell R, Richter L, et al. Maternal and Child Undernutrition: Consequences for Adult Health and Human Capital. The Lancet. 2008 Jan;371(9609):340–57. doi: 10.1016/S0140-6736(07)61692-4. Available from: [DOI] [Google Scholar]
  • 41.Gibb DM, Duong T, Tookey PA, Sharland M, Tudor-Williams G, Lyall H, et al. Decline in Mortality, AIDS, and Hospital Admissions in Perinatally HIV-1 Infected Children in the United Kingdom and Ireland. BMJ. 2003 Oct;327(7422):1019. doi: 10.1136/bmj.327.7422.1019. Available from: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42.Haines A, Kuruvilla S. Bridging the Implementation Gap Between Knowledge and Action for Health. Bulletin of the World Health Organization. 2004;82:724–31. [PMC free article] [PubMed] [Google Scholar]
  • 43.Frost MB, Forste R. Maternal Education and Child Nutritional Status in Bolivia: Finding the Links. Social Science & Medicine. 2005;60(2):395–407. doi: 10.1016/j.socscimed.2004.05.010. [DOI] [PubMed] [Google Scholar]

Articles from AMIA Annual Symposium Proceedings are provided here courtesy of American Medical Informatics Association

RESOURCES