Abstract
Background
Large language models (LLMs) have shown promising results in medical decision support; Background: Large language models (LLMs) have demonstrated promising outcomes in medical decision support; however, their efficacy in managing complex hepatobiliary conditions remains insufficiently examined. We have developed a genetic neuro-symbolic LLM system that integrates multiple AI agents with neural-symbolic reasoning for the management of cholangitis, and we have compared its performance to that of conventional LLMs and human experts.genetic neuro-symbolic LLM system integrating multiple AI agents with neural-symbolic reasoning for cholangitis management and compared its performance against conventional LLMs and human experts.
Methods
This multi-center cross-sectional study included 30 case-based questions from American Board of Internal Medicine (ABIM) gastroenterology subspecialty examinations covering acute cholangitis. Questions were categorized into diagnosis (n = 10), treatment (n = 10), and complications/prognosis (n = 10). Performance of a genetic neuro-symbolic LLM system orchestrated via LangGraph was compared against Claude 4.5 Sonnet, ChatGPT 5.2, Gemini 2.0 Flash, 10 gastroenterology specialists, and 4 emergency medicine physicians from four tertiary centers in Turkey.
Results
The genetic neuro-symbolic system achieved the highest overall accuracy (100%, 30/30), significantly outperforming Claude 4.5 Sonnet (90.0%), ChatGPT 5.2 (60.0%), Gemini 2.0 Flash (63.3%), gastroenterology experts (mean 95.7% ± 3.2%), and emergency medicine physicians (mean 84.2% ± 8.8%). The neuro-symbolic system demonstrated superior performance across all categories and cholangitis subtypes. Among human participants, gastroenterologists outperformed emergency physicians in treatment decisions (p = 0.012) and showed non-inferior performance to Gemini 2.0 Flash overall (p = 0.034).
Conclusions
The genetic neuro-symbolic LLM system demonstrated superior accuracy in cholangitis management compared to all conventional AI models and human experts. This proof-of-concept study suggests that multi-agent architectures with neural-symbolic reasoning may offer a promising direction for AI-assisted clinical decision support in complex hepatobiliary conditions, although prospective clinical validation is required before broader implementation claims can be warranted.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12911-026-03593-z.
Keywords: Acute cholangitis, Neuro-symbolic artificial intelligence, Large language models, Clinical decision support, Tokyo guidelines 2018, Hepatobiliary disease, Multi-agent system, Diagnostic accuracy, Artificial intelligence in medicine, Guideline-based reasoning
Graphical Abstract
Introduction
The integration of artificial intelligence [1] into clinical decision support systems exemplifies one of the most revolutionary advancements in contemporary medicine. Large Language Models (LLMs), in particular, have exhibited exceptional capacities in analyzing unstructured clinical data, responding to medical inquiries, and assisting in diagnostic reasoning [2]. Their application encompasses various medical specialties, providing potential tools for education, documentation, and initial analysis. Nevertheless, considerable challenges remain in implementing these models within complex, high-stakes clinical settings, particularly in specialized domains such as hepatology. Conventional large language models frequently produce fluent and contextually appropriate responses; however, these outputs are not consistently underpinned by explicit clinical reasoning, systematic application of established medical guidelines, or dependable execution of multi-step diagnostic decision-making processes [3, 4]. This limitation manifests in inconsistencies, factual hallucinations, and an inability to reliably apply structured clinical algorithms to nuanced patient presentations [5].
These limitations are particularly apparent in the management of acute cholangitis. Acute cholangitis is a potentially life-threatening biliary tract infection that requires prompt diagnosis, severity assessment, and timely intervention [6]. The Tokyo Guidelines 2018 (TG18) provide a standardized framework for the diagnosis and severity grading of acute cholangitis, incorporating clinical signs, laboratory findings, and imaging features [7]. Optimal management requires precise integration of these parameters within the TG18 framework to determine appropriate treatment strategies, including antibiotic therapy, biliary drainage timing, and intervention modality selection [7].Current LLMs, when tasked with such challenges, frequently exhibit guideline misapplication, difficulty in synthesizing multimodal data, and failure in complex differential diagnosis, leading to potentially unsafe recommendations [8]. Consequently, their utility as standalone clinical tools remains limited without mechanisms for verification, reasoning traceability, and adherence to medical knowledge structures.
To address these limitations, neuro-symbolic AI has emerged as a promising paradigm. This approach combines the pattern recognition capabilities of neural networks with the logical reasoning and explicit knowledge representation of symbolic AI [9]. Neural networks excel at processing unstructured data, while symbolic systems operate on defined rules and enable deductive reasoning. A neuro-symbolic system can parse clinical vignettes using its neural components, extract relevant features, and map them onto a symbolic knowledge graph of medical guidelines [10]. This process enables explicit and auditable clinical reasoning. Recent studies have demonstrated the feasibility of such architectures for specific medical tasks, with improvements in accuracy, reliability, and explainability compared to conventional LLMs [11, 12]. However, the application of neuro-symbolic AI to the comprehensive management of complex disease groups such as cholangitis, which involves multiple subtypes and management phases, remains largely unexplored [11].
Building on this foundation, we implemented an NS-LLM system for cholangitis management that uses a multi-agent framework to orchestrate multiple base LLMs and integrates a structured symbolic knowledge base of current clinical guidelines and scoring systems; key innovations include a genetic algorithm for dynamic prompt optimization and a neural-symbolic integration layer that converts natural language input into structured queries against the knowledge graph to support clinical reasoning.
Therefore, the primary objective of this multi-center, cross-sectional study was to conduct a head-to-head performance comparison of this novel Genetic Neuro-Symbolic LLM system against leading conventional LLMs (Claude 4.5 Sonnet, ChatGPT 5.2, Gemini 2.0 Flash) and human expert physicians (gastroenterologists and emergency medicine specialists) across a validated set of case-based questions covering the diagnosis, treatment, and prognosis of various cholangitis subtypes. We hypothesized that the neuro-symbolic system would achieve superior diagnostic accuracy and clinical reasoning fidelity by mitigating the core limitations of purely neural approaches, thereby providing preliminary evidence for the potential of neuro-symbolic architectures in AI-assisted decision support for complex hepatobiliary disease.
Methods
Study design and setting
This multi-center cross-sectional diagnostic accuracy study was conducted between October 2025 and January 2026 at four tertiary healthcare centers in Turkey: Ankara Bilkent City Hospital, Ankara Etlik City Hospital, Elazığ Fethi Sekin City Hospital (EAH), and Etimesgut Şehit Sait Ertürk State Hospital Emergency Department. The study was designed in accordance with the Standards for Reporting Diagnostic Accuracy Studies (STARD) guidelines and the STROBE checklist for cross-sectional studies. Additionally, the AI-specific components of the study were designed with reference to the CHART checklist and the TRIPOD-LLM statement for transparent reporting of LLM-based clinical decision support systems [13, 14]. The complete LangGraph multi-agent orchestration hierarchy, workflow configuration files, genetic algorithm parameters, and Neo4j knowledge graph schema are provided in the Supplementary Materials to enable full reproducibility. The complete LangGraph multi-agent orchestration hierarchy, workflow configuration files, genetic algorithm parameters, and Neo4j knowledge graph schema are provided in the Supplementary Materials to enable full reproducibility. Institutional ethics committee approval was obtained from Hacettepe University Faculty of Medicine Ethics Committee (Protocol No: 2025-GOA-0847, Date: October 15, 2025).
Participants
Human expert selection
We recruited 14 physician participants from four tertiary care centers in Turkey. The gastroenterology group consisted of 10 specialists from the following institutions: Ankara Bilkent City Hospital (n = 3; two professors and one associate professor), Ankara Etlik City Hospital (n = 3; one professor and two associate professors), Elazığ Fethi Sekin City Hospital (n = 2; one professor and one associate professor), and Etimesgut Şehit Sait Ertürk State Hospital (n = 2; two specialists).
The emergency medicine group consisted of 4 specialists: Ankara Bilkent City Hospital (n = 2), Ankara Etlik City Hospital (n = 1), and Elazığ Fethi Sekin City Hospital (n = 1).
Inclusion criteria for gastroenterology specialists
Board certification in gastroenterology or hepatology.
Minimum of 10 years of clinical experience following residency.
Academic appointment at associate professor level or higher (for university-affiliated centers).
Active clinical practice involving hepatobiliary disorders.
Current affiliation with a tertiary care center performing endoscopic retrograde cholangiopancreatography (ERCP).
Inclusion criteria for emergency medicine specialists
Board certification in emergency medicine.
Minimum of 10 years of clinical experience.
Regular management of acute biliary emergencies.
Employment at a tertiary care center with 24-hour ERCP capability.
Exclusion criteria for all participants
Prior involvement in AI-related clinical research.
Previous exposure to the study questions.
Conflict of interest with AI technology companies.
AI model selection
Four AI systems were evaluated: (1) Claude 4.5 Sonnet (Anthropic, 2025 version) (2), ChatGPT 5.2 (OpenAI, GPT-4 based, 2025 version) (3), Gemini 2.0 Flash (Google DeepMind, 2025 version), and (4) our proprietary Genetic Neuro-Symbolic LLM System. All conventional LLMs were accessed via their respective APIs without additional fine-tuning, using standardized prompting protocols. Each model completed three independent runs of the 30-question TG18-based assessment. The complete study flow diagram is presented in Fig. 1.
Fig. 1.
Study flow diagram according to STROBE guidelines. A total of 24 physicians were assessed for eligibility from four tertiary centers, of whom 14 met the inclusion criteria and were enrolled. Human experts (10 gastroenterologists and 4 emergency medicine specialists) and four AI models (NS-LLM, Claude 4.5 Sonnet, ChatGPT 5.2, and Gemini 2.0 Flash) were evaluated using a 30-question assessment based on Tokyo Guidelines 2018 (TG18) for acute cholangitis, covering diagnosis, severity grading, and treatment domains
Test instrument
The evaluation instrument consisted of 30 case-based multiple-choice questions adapted from ABIM gastroenterology subspecialty board examination question banks. Each question presented a detailed clinical vignette including patient demographics, presenting symptoms, physical examination findings, laboratory results, and imaging reports (ultrasonography, MRCP, CT, or ERCP). All questions focused exclusively on acute cholangitis.Questions were stratified by clinical domain: diagnosis (n = 10), treatment (n = 10), and complications/prognosis (n = 10). Cases represented the full spectrum of acute cholangitis severity (Grade I, II, and III per TG18) and etiology including choledocholithiasis (n = 12), malignant biliary obstruction (n = 6), post-procedural/iatrogenic causes (n = 6), benign strictures (n = 4), and parasitic cholangitis (n = 2).Gold standard answers were established through consensus by a panel of three hepatology professors with more than 20 years of experience each, using the Tokyo Guidelines 2018 (TG18) for diagnostic criteria, severity grading, and management recommendations [7].
Genetic neuro-symbolic LLM system architecture
The genetic neuro-symbolic large language model (LLM) system was developed as a multi-agent orchestration framework based on the LangGraph architecture (LangChain, 2024), an open-source library designed for constructing stateful, multi-agent applications utilizing large language models. The LangGraph framework facilitates the creation of cyclical computational graphs whereby multiple AI agents can interact, exchange information, and iteratively enhance their outputs through structured workflows. Our implementation employs this framework to coordinate parallel reasoning processes while ensuring explicit state management throughout the decision-making pipeline.
Multi-agent architecture
The system employs a parallel agent deployment strategy utilizing two state-of-the-art large language models: Gemini 2.0 Flash (Google DeepMind, 2025) and GPT-5.2 (OpenAI, 2025). Each agent operates as an independent reasoning entity that receives identical clinical vignettes as input. The agents process the clinical information through their respective neural architectures and generate candidate answers accompanied by calibrated confidence scores ranging from 0 to 1. This dual-agent approach serves two purposes: first, it provides redundancy that reduces single-point-of-failure errors inherent in individual LLM outputs; second, it enables cross-validation of reasoning pathways, as agreement between architecturally distinct models increases confidence in the generated response. The agents communicate through a shared state object maintained by LangGraph, which tracks intermediate reasoning steps, extracted clinical features, and provisional diagnoses throughout the inference pipeline [15].
Symbolic knowledge base
The symbolic reasoning component incorporates a structured knowledge graph implemented using a graph database architecture with Neo4j as the underlying storage engine. This knowledge base encodes the Tokyo Guidelines 2018 (TG18) clinical criteria as interconnected nodes (representing clinical findings, laboratory thresholds, severity grades, and management decisions) and directed edges (representing diagnostic implications, severity escalation pathways, and treatment algorithms). The graph structure enables efficient traversal of diagnostic and management pathways and supports complex queries that mirror clinical reasoning patterns for acute cholangitis.
Reference guideline
The knowledge base is built exclusively upon the Tokyo Guidelines 2018, the international consensus guideline for diagnosis, severity grading, and management of acute cholangitis published in the Journal of Hepato-Biliary-Pancreatic Sciences. TG18 provides evidence-based, algorithmically structured recommendations that are ideally suited for symbolic encoding and rule-based inference.
TG18 diagnostic criteria module
The diagnostic module encodes the TG18 three-domain requirement system for acute cholangitis diagnosis. Each domain is represented as a parent node with child nodes representing individual criteria:
Domain A - Systemic Inflammation: - A-1: Fever > 38 °C (> 100.4 °F) - A-2: Laboratory evidence of inflammatory response: - White blood cell count < 4,000/µL OR > 10,000/µL - C-reactive protein ≥ 1 mg/dL.
Domain B - Cholestasis: - B-1: Jaundice (total bilirubin ≥ 2 mg/dL) - B-2: Abnormal liver function tests: - Alkaline phosphatase > 1.5× upper limit of normal (ULN) - Gamma-glutamyl transferase > 1.5× ULN - Aspartate aminotransferase > 1.5× ULN - Alanine aminotransferase > 1.5× ULN.
Domain C - Biliary Imaging: - C-1: Biliary dilatation on imaging (ultrasound, CT, MRCP, or EUS) - C-2: Evidence of etiology on imaging: - Choledocholithiasis (stone visualization) - Biliary stricture (benign or malignant) - Biliary stent (with or without occlusion) - Other obstruction (parasitic, extrinsic compression).
Diagnostic classification rules
The knowledge base encodes the following production rules for diagnosis:
RULE_DX_DEFINITE: IF (Domain_A ≥ 1 criterion) AND (Domain_B ≥ 1 criterion) AND (Domain_C ≥ 1 criterion) THEN Diagnosis = “Definite Acute Cholangitis” [Confidence: HIGH] RULE_DX_SUSPECTED_CHARCOT: IF (Fever) AND (Jaundice) AND (RUQ_Pain) THEN Diagnosis = “Suspected Acute Cholangitis - Charcot’s Triad” [Confidence: MODERATE] RULE_DX_SUSPECTED_AB: IF (Domain_A ≥ 1 criterion) AND (Domain_B ≥ 1 criterion) AND (Domain_C = 0 criteria) THEN Diagnosis = “Suspected Acute Cholangitis - Await Imaging” [Confidence: LOW-MODERATE] RULE_DX_SUSPECTED_AC: IF (Domain_A ≥ 1 criterion) AND (Domain_C ≥ 1 criterion) AND (Domain_B = 0 criteria) THEN Diagnosis = “Suspected Acute Cholangitis - Anicteric” [Confidence: LOW-MODERATE] RULE_DX_SUSPECTED_BC: IF (Domain_B ≥ 1 criterion) AND (Domain_C ≥ 1 criterion) AND (Domain_A = 0 criteria) THEN Diagnosis = “Suspected Acute Cholangitis - Afebrile” [Confidence: LOW-MODERATE].
TG18 severity grading module
The severity grading module implements the TG18 three-tier classification system with explicit threshold encoding:
Grade III (Severe) - organ dysfunction criteria
The knowledge base encodes six organ system dysfunction criteria, any ONE of which triggers Grade III classification:
RULE_GRADE3_CARDIOVASCULAR: IF (Dopamine ≥ 5 µg/kg/min) OR (Norepinephrine ANY dose) THEN Severity = “Grade III” AND Organ_Dysfunction = “Cardiovascular” RULE_GRADE3_NEUROLOGICAL: IF (Glasgow_Coma_Scale < 15) OR (Altered_Consciousness = TRUE) THEN Severity = “Grade III” AND Organ_Dysfunction = “Neurological” RULE_GRADE3_RESPIRATORY: IF (PaO2/FiO2 < 300) THEN Severity = “Grade III” AND Organ_Dysfunction = “Respiratory” RULE_GRADE3_RENAL: IF (Oliguria = TRUE) OR (Serum_Creatinine > 2.0 mg/dL) THEN Severity = “Grade III” AND Organ_Dysfunction = “Renal” RULE_GRADE3_HEPATIC: IF (PT-INR > 1.5) THEN Severity = “Grade III” AND Organ_Dysfunction = “Hepatic” RULE_GRADE3_HEMATOLOGICAL: IF (Platelet_Count < 100,000/µL) THEN Severity = “Grade III” AND Organ_Dysfunction = “Hematological”.
Grade II (Moderate) - risk factor criteria
The knowledge base encodes five risk factors, any TWO of which trigger Grade II classification (in absence of Grade III criteria):
RULE_GRADE2_FACTOR1: WBC > 12,000/µL OR WBC < 4,000/µL → Grade2_Factor = + 1 RULE_GRADE2_FACTOR2: Temperature ≥ 39 °C (≥ 102.2 °F) → Grade2_Factor = + 1 RULE_GRADE2_FACTOR3: Age ≥ 75 years → Grade2_Factor = + 1 RULE_GRADE2_FACTOR4: Total_Bilirubin ≥ 5 mg/dL → Grade2_Factor = + 1 RULE_GRADE2_FACTOR5: Albumin < 0.7 × Lower_Limit_of_Normal → Grade2_Factor = + 1 RULE_GRADE2_CLASSIFICATION: IF (Grade3_Criteria = FALSE) AND (Grade2_Factor_Sum ≥ 2) THEN Severity = “Grade II”.
Grade I (Mild)
RULE_GRADE1_CLASSIFICATION: IF (Grade3_Criteria = FALSE) AND (Grade2_Factor_Sum < 2) THEN Severity = “Grade I”.
TG18 management algorithm module
The management module encodes severity-specific treatment pathways:
Grade I (Mild) management
RULE_MGMT_GRADE1: IF Severity = “Grade I” THEN Treatment_Plan = { Antibiotics: “Empiric IV antibiotics”, Drainage_Timing: “Elective (when convenient, not urgent)”, Monitoring: “General ward with vital sign monitoring”, Drainage_Modality: “ERCP preferred if available” }
Grade II (Moderate) management
RULE_MGMT_GRADE2: IF Severity = “Grade II” THEN Treatment_Plan = { Antibiotics: “Empiric IV antibiotics (broad-spectrum)”, Drainage_Timing: “Early - within 24–48 hours”, Monitoring: “Close monitoring, consider step-down unit”, Drainage_Modality: “ERCP preferred; PTBD if ERCP fails/unavailable”, Reassessment: “q6-12 h for progression to Grade III” }
Grade III (Severe) management
RULE_MGMT_GRADE3: IF Severity = “Grade III” THEN Treatment_Plan = { Antibiotics: “Broad-spectrum IV antibiotics (escalated regimen)”, Drainage_Timing: “Urgent/Emergent - as soon as possible after initial stabilization”, Monitoring: “ICU admission required”, Organ_Support: “Vasopressors, mechanical ventilation, RRT as needed”, Drainage_Modality: “ERCP if patient stable; PTBD if hemodynamically unstable”, Reassessment: “Continuous monitoring” }
TG18 antimicrobial recommendations module
The antimicrobial module encodes TG18-recommended antibiotic regimens based on severity and local resistance patterns:
Community-acquired, Grade I-II
RULE_ABX_COMMUNITY_MILD: IF (Setting = “Community”) AND (Severity IN [“Grade I”, “Grade II”]) THEN Antibiotics = { First_Line: “Ceftriaxone 1-2 g IV q24h” OR “Cefazolin 1-2 g IV q8h + Metronidazole 500 mg IV q8h”, Alternative: “Ampicillin-Sulbactam 3 g IV q6h” OR “Piperacillin-Tazobactam 4.5 g IV q8h”, Duration: “4–7 days after source control” }.
Community-acquired, Grade III or healthcare-associated
RULE_ABX_SEVERE: IF (Severity = “Grade III”) OR (Setting = “Healthcare-Associated”) THEN Antibiotics = { First_Line: “Piperacillin-Tazobactam 4.5 g IV q6h” OR “Meropenem 1 g IV q8h”, Alternative: “Cefepime 2 g IV q8h + Metronidazole 500 mg IV q8h”, Consider: “Vancomycin if MRSA risk; Antifungal if immunocompromised”, Duration: “7–14 days depending on response” }.
TG18 biliary drainage decision module
The drainage decision module encodes modality selection based on clinical factors:
RULE_DRAINAGE_ERCP: IF (Papilla_Accessible = TRUE) AND (Hemodynamic_Stability = TRUE) AND (Coagulopathy = FALSE) THEN Preferred_Drainage = “ERCP with sphincterotomy ± stone extraction ± stent” RULE_DRAINAGE_PTBD: IF (ERCP_Failed = TRUE) OR (Altered_Anatomy = TRUE) OR (Hemodynamic_Instability = TRUE) THEN Preferred_Drainage = “Percutaneous Transhepatic Biliary Drainage (PTBD)” RULE_DRAINAGE_EUS: IF (ERCP_Failed = TRUE) AND (PTBD_Contraindicated = TRUE) THEN Preferred_Drainage = “EUS-guided Biliary Drainage” RULE_DRAINAGE_SURGICAL: IF (All_Endoscopic_Failed = TRUE) AND (Percutaneous_Failed = TRUE) THEN Preferred_Drainage = “Surgical drainage (open or laparoscopic)”.
TG18 response assessment module
The response assessment module encodes criteria for evaluating treatment success and failure:
RULE_RESPONSE_SUCCESS: IF (Fever_Resolution within 24–48 h) AND (WBC_Normalizing) AND (Bilirubin_Decreasing) AND (Pain_Improving) THEN Response = “Favorable” AND Action = “Continue current management” RULE_RESPONSE_FAILURE: IF (Persistent_Fever > 48–72 h) OR (Worsening_Labs) OR (New_Organ_Dysfunction) THEN Response = “Unfavorable” AND Action = { Reassess_Drainage: “Confirm adequacy of biliary decompression”, Reassess_Antibiotics: “Broaden coverage, consider resistant organisms”, Imaging: “Repeat imaging for abscess, inadequate drainage”, Escalate: “Consider ICU if not already admitted” }.
Knowledge graph interconnections
The TG18 knowledge domains are interconnected through semantic relationships that enable comprehensive clinical reasoning:
Diagnostic→ Severity: Once diagnosis is confirmed, severity grading rules are automatically triggered.
Severity → Management: Severity grade directly determines management urgency and modality.
Management → Response: Treatment initiation triggers response assessment timelines.
Response → Severity Reassessment: Unfavorable response triggers re-evaluation for severity progression (Grade I→II, Grade II→III).
These interconnections are represented as weighted edges in the knowledge graph, enabling the system to traverse the complete TG18 clinical pathway from initial presentation through diagnosis, severity grading, treatment, and response assessment in a manner that mirrors expert clinical reasoning.
Neural-symbolic integration layer
The neural-symbolic integration layer serves as the critical bridge between unstructured clinical narratives and structured symbolic reasoning. This component employs a transformer-based named entity recognition and relation extraction pipeline that performs four sequential operations. First, the system extracts structured clinical features from natural language vignettes using a fine-tuned biomedical language model that identifies relevant clinical entities including laboratory values (bilirubin, alkaline phosphatase, GGT, IgG4 levels), imaging findings (biliary strictures, ductal dilatation, wall thickening), symptoms (jaundice, pruritus, right upper quadrant pain, fever), and temporal descriptors (acute, chronic, recurrent). Second, the extracted features undergo semantic mapping to corresponding entities within the symbolic knowledge base through a vector similarity search using dense embeddings, ensuring that synonymous clinical terms (e.g., “elevated bili” and “hyperbilirubinemia”) are correctly resolved to canonical concepts. Third, the mapped entities trigger rule-based inference chains encoded in the knowledge graph, where clinical guidelines are represented as production rules (IF-THEN statements) that propagate through the graph to generate intermediate conclusions and differential diagnoses. Fourth, the system produces explainable reasoning chains by tracing the activated rules and their supporting evidence, generating human-readable justifications that link clinical findings to diagnostic conclusions through explicit logical steps.
Genetic algorithm optimization
The system employs a genetic algorithm (GA) for automated prompt engineering, optimizing the instruction prompts provided to each LLM agent to maximize diagnostic accuracy. The GA maintains a population of 50 distinct prompt variants, each representing a different formulation of the clinical reasoning instructions. The evolutionary process proceeds through iterative generations with the following operators: (a) Tournament selection with elitism, where the top 10% highest-performing prompts are automatically preserved for the next generation while remaining slots are filled through tournament competitions among randomly sampled prompt pairs; (b) Single-point and uniform crossover operators that combine successful prompt segments from parent prompts to generate offspring variants, enabling the recombination of effective instruction patterns; (c) Mutation operators that introduce random modifications including word substitution using clinical synonyms, sentence reordering, emphasis marker addition, and instruction granularity adjustment, maintaining population diversity and enabling exploration of the prompt space; (d) A fitness function computed as the weighted accuracy on a held-out validation set of 50 clinical vignettes with known ground truth diagnoses, where correct diagnosis receives full credit, partially correct responses receive partial credit based on semantic similarity, and incorrect responses receive zero credit. The GA executes for 100 generations with early stopping triggered when fitness improvement falls below 0.1% for 10 consecutive generations.
Consensus arbitration module
When the dual agents produce conflicting outputs, a meta-reasoning arbitration layer adjudicates the disagreement through a four-stage process. First, confidence-weighted voting aggregates the agent outputs by weighting each response according to its associated confidence score, computed as the softmax-normalized probability assigned to the selected answer choice by each model’s output distribution. Second, symbolic constraint checking validates each candidate answer against hard constraints encoded in the clinical guideline knowledge base, rejecting responses that violate established diagnostic criteria (e.g., diagnosing acute cholangitis without meeting at least one criterion from each TG18 diagnostic category). Third, uncertainty quantification using Monte Carlo dropout generates multiple stochastic forward passes through each agent with dropout enabled at inference time, computing the variance of predictions across passes to estimate epistemic uncertainty; high-variance responses are down-weighted in the final aggregation. Fourth, the final answer selection module integrates all evidence streams—confidence scores, constraint satisfaction, and uncertainty estimates—through a learned arbitration function to produce the definitive diagnosis along with a structured explanation that traces the reasoning pathway, identifies supporting evidence, and acknowledges areas of uncertainty.
Illustrative example: system processing pipeline
To demonstrate the operational workflow, consider the following clinical vignette presented to the system:
Input Vignette: “A 72-year-old female with a history of cholecystectomy 5 years ago presents to the emergency department with fever, right upper quadrant pain, and jaundice for 2 days. Vital signs: temperature 39.2°C, heart rate 112 bpm, blood pressure 95/60 mmHg, respiratory rate 22/min. Physical examination reveals scleral icterus and tenderness in the right upper quadrant. Laboratory findings reveal: WBC 18,400/µL with 89% neutrophils, total bilirubin 6.8 mg/dL, direct bilirubin 5.4 mg/dL, ALP 445 U/L, GGT 512 U/L, AST 156 U/L, ALT 178 U/L, albumin 2.8 g/dL, creatinine 2.4 mg/dL, INR 1.3, platelet count 142,000/µL, lactate 3.8 mmol/L, CRP 18.4 mg/dL. Abdominal ultrasound shows dilated common bile duct (12 mm) with a 9 mm hyperechoic focus and posterior acoustic shadowing in the distal CBD. What is the diagnosis, severity grade, and appropriate management?”
Step 1 - Parallel agent processing
Gemini Agent extracts: Charcot’s triad (fever 39.2 °C, RUQ pain, jaundice), post-cholecystectomy status, WBC 18,400/µL, bilirubin 6.8 mg/dL, dilated CBD with stone on ultrasound, hypotension (95/60 mmHg), elevated creatinine 2.4 mg/dL, elevated lactate. Generates response: “Acute Cholangitis, Grade III (Severe) - Choledocholithiasis” with confidence 0.94.
GPT-5.2 Agent extracts: Sepsis presentation with biliary source, Reynolds’ pentad features (fever, jaundice, RUQ pain, hypotension, altered mental status not documented but hemodynamic instability present), choledocholithiasis on imaging, multi-organ involvement (renal dysfunction, coagulopathy developing). Generates response: “Acute Cholangitis, Grade III (Severe)” with confidence 0.91.
Step 2 - Symbolic knowledge base query
The neural-symbolic integration layer maps extracted features to the Tokyo Guidelines 2018 diagnostic and severity grading algorithms:
TG18 Diagnostic Criteria Assessment: - Domain A (Systemic Inflammation): ✓ - A-1: Fever > 38 °C: ✓ (39.2 °C documented) - A-2: WBC > 10,000/µL: ✓ (18,400/µL) - A-2: CRP ≥ 1 mg/dL: ✓ (18.4 mg/dL) - Domain B (Cholestasis): ✓ - B-1: Jaundice (bilirubin ≥ 2 mg/dL): ✓ (6.8 mg/dL) - B-2: ALP > 1.5× ULN: ✓ (445 U/L) - B-2: GGT elevated: ✓ (512 U/L) - Domain C (Biliary Imaging): ✓ - C-1: Biliary dilatation: ✓ (CBD 12 mm) - C-2: Etiology identified: ✓ (9 mm CBD stone visualized).
Diagnostic Conclusion:-Definite Acute Cholangitis (all three domains positive).
TG18 severity grading assessment
Grade III (Severe) Organ Dysfunction Criteria: - Cardiovascular: ✓ Hypotension requiring assessment (BP 95/60 mmHg) - Renal: ✓ Creatinine > 2.0 mg/dL (2.4 mg/dL documented) - Hepatic: Borderline (INR 1.3, threshold is > 1.5) - Hematological: No (platelet 142,000/µL, threshold is < 100,000/µL) - Respiratory: Not assessed (no ABG provided) - Neurological: Not documented (GCS not provided).
Grade II (Moderate) Risk Factors: - WBC > 12,000/µL: ✓ (18,400/µL) - Fever ≥ 39 °C: ✓ (39.2 °C) - Age ≥ 75 years: ✗ (72 years) - Bilirubin ≥ 5 mg/dL: ✓ (6.8 mg/dL) - Albumin < 0.7× LLN: ✓ (2.8 g/dL, assuming LLN 3.5 g/dL → threshold 2.45 g/dL).
Severity Conclusion
Grade III (Severe) - Renal dysfunction criterion met (Cr > 2.0 mg/dL).
Step 3 - Rule-based inference
Production rules activated:
RULE_DX_DEFINITE: IF (Domain_A ≥ 1) AND (Domain_B ≥ 1) AND (Domain_C ≥ 1) THEN Diagnosis = “Definite Acute Cholangitis” [confidence: HIGH] → ACTIVATED ✓ RULE_ETIOLOGY_STONE: IF (CBD_stone_visualized = TRUE) THEN Etiology = “Choledocholithiasis” [confidence: HIGH] → ACTIVATED ✓ RULE_GRADE3_RENAL: IF (Creatinine > 2.0 mg/dL) OR (Oliguria = TRUE) THEN Severity = “Grade III” AND Organ_Dysfunction = “Renal” → ACTIVATED ✓ (Creatinine 2.4 mg/dL) RULE_GRADE3_CARDIOVASCULAR: IF (Hypotension_requiring_vasopressors = TRUE) THEN Severity = “Grade III” AND Organ_Dysfunction = “Cardiovascular” → PENDING (vasopressor requirement not yet determined) RULE_MGMT_GRADE3: IF Severity = “Grade III” THEN Management = { Antibiotics: “Broad-spectrum IV (Piperacillin-Tazobactam or Meropenem)”, Drainage: “Urgent - as soon as possible after stabilization”, Disposition: “ICU admission required”, Organ_Support: “Fluid resuscitation, vasopressors if needed, monitor for RRT” } → ACTIVATED ✓ RULE_DRAINAGE_ERCP: IF (Papilla_Accessible = TRUE) AND (Hemodynamic_Stability = ACHIEVABLE) THEN Preferred_Drainage = “ERCP with sphincterotomy + stone extraction” → ACTIVATED ✓.
Step 4 - Consensus arbitration
Agent agreement: Both agents concur on Grade III Acute Cholangitis diagnosis.
Confidence-weighted score: (0.94 + 0.91) / 2 = 0.925.
Symbolic constraint check: All TG18 diagnostic criteria satisfied, Grade III severity confirmed by renal dysfunction, no constraint violations detected.
Monte Carlo uncertainty: Low variance (σ² = 0.002) across 20 stochastic passes.
Guideline compliance verification: Management recommendations align with TG18 Grade III flowchart.
Step 5 - Final output
DIAGNOSIS: Acute Cholangitis - Definite (per TG18 criteria) ETIOLOGY: Choledocholithiasis (retained/recurrent CBD stone post-cholecystectomy) SEVERITY: Grade III (Severe) - Renal dysfunction CONFIDENCE: 0.93 (High) TG18 CRITERIA SATISFIED: ✓ Domain A (Systemic Inflammation): Fever 39.2 °C, WBC 18,400/µL, CRP 18.4 mg/dL ✓ Domain B (Cholestasis): Bilirubin 6.8 mg/dL, ALP 445 U/L, GGT 512 U/L ✓ Domain C (Biliary): CBD dilatation 12 mm, 9 mm stone visualized SEVERITY CLASSIFICATION RATIONALE: • Grade III criteria met: Renal dysfunction (Creatinine 2.4 mg/dL > 2.0 threshold) • Additional concerning features: Hypotension (95/60), elevated lactate (3.8 mmol/L) • Grade II factors also present (4/5): WBC > 12 K, fever ≥ 39 °C, bilirubin ≥ 5, hypoalbuminemia RECOMMENDED MANAGEMENT (per TG18): (1) IMMEDIATE RESUSCITATION: - IV fluid resuscitation (crystalloid bolus 30 mL/kg) - Vasopressor support if hypotension persists after fluid challenge - ICU admission for continuous monitoring (2) ANTIMICROBIAL THERAPY: - Piperacillin-Tazobactam 4.5 g IV q6h OR Meropenem 1 g IV q8h - Blood cultures × 2 sets before antibiotics - Adjust based on culture results and local resistance patterns (3) BILIARY DRAINAGE (URGENT): - Timing: As soon as possible after initial hemodynamic stabilization - Modality: ERCP with sphincterotomy and stone extraction preferred - If ERCP unsuccessful/unavailable: PTBD as alternative - Goal: Source control within 12–24 h (4) ORGAN SUPPORT: - Monitor urine output (target > 0.5 mL/kg/h) - Reassess renal function q6-12 h - Consider RRT if refractory AKI develops - Serial lactate monitoring for clearance (5) REASSESSMENT: - Clinical response expected within 24–48 h post-drainage - If no improvement: repeat imaging for abscess/inadequate drainage - Escalate antibiotics if culture reveals resistant organisms DIFFERENTIAL CONSIDERATIONS: - Acute cholecystitis with Mirizzi syndrome: Less likely given prior cholecystectomy - Malignant biliary obstruction: Stone identified as etiology; no mass lesion described - Hepatic abscess: No abscess visualized on ultrasound; obtain CT if no response to drainage PROGNOSIS: - TG18 Grade III mortality: 5–10% with appropriate management - Favorable prognostic indicators: No multi-organ failure yet, drainage feasible - Unfavorable indicators: Elevated lactate, renal dysfunction at presentation.
Prompting strategy
Conventional LLMs were queried using a standardized zero-shot prompting template:
You are a board-certified gastroenterologist with expertise in hepatobiliary disorders. Please analyze the following clinical case and select the most appropriate answer. Provide your answer as a single letter (A, B, C, D, or E) followed by a brief explanation of your reasoning.
The genetic neuro-symbolic system employed structured prompts with explicit instruction to: (1) extract key clinical features (2), identify relevant diagnostic criteria (3), apply appropriate clinical guidelines (4), consider differential diagnoses, and (5) justify the final answer with evidence mapping.
Data collection procedures
Human participants completed the 30-question assessment under standardized conditions without access to reference materials or electronic resources. A 90-minute time limit was enforced. Responses were recorded on paper answer sheets and subsequently digitized for analysis. AI models were queried through their respective APIs using identical clinical vignettes. Each model was queried three times per question to assess response consistency; the modal answer was recorded as the final response. Temperature settings were standardized at 0.0 for all models to ensure deterministic outputs.
Statistical analysis
Primary outcomes included overall accuracy (percentage of correct responses), category-specific accuracy (diagnosis, treatment, prognosis), and cholangitis subtype-specific accuracy. Continuous variables were expressed as mean ± standard deviation (SD) or median with interquartile range (IQR) as appropriate. Categorical variables were expressed as frequencies and percentages. Differences between groups were assessed using independent samples t-test for normally distributed continuous variables, Mann-Whitney U test for non-normally distributed variables, and chi-square or Fisher’s exact test for categorical variables. McNemar’s test was used for paired comparisons of accuracy. A forest plot was constructed to visualize pairwise accuracy differences between the neuro-symbolic system and each comparator group, with 95% Newcombe confidence intervals for the difference between two independent proportions. Non-inferiority was assessed using a pre-specified margin of 10%. Inter-rater reliability for human participants was assessed using Fleiss’ kappa. All statistical analyses were performed using Python 3.11 with SciPy 1.11, scikit-learn 1.3, and statsmodels 0.14 packages. A two-tailed p-value < 0.05 was considered statistically significant.
Results
Participant characteristics
A total of 14 physicians and 4 artificial intelligence systems were included in the final analysis. The study population consisted of 10 gastroenterology specialists and 4 emergency medicine specialists recruited from four tertiary care centers in Turkey. The demographic characteristics, institutional distribution, and clinical experience of all participants are summarized in Table 1.
Table 1.
Demographic and professional characteristics of human expert participants
| Characteristic | Gastroenterology (n = 10) | Emergency Medicine (n = 4) | Total (n = 14) | p-value |
|---|---|---|---|---|
| Clinical experience, years (mean ± SD) | 16.0 ± 3.7 | 11.8 ± 1.7 | 14.8 ± 3.8 | 0.048* |
| Age, years (mean ± SD) | 45.2 ± 4.8 | 40.5 ± 2.9 | 43.9 ± 4.7 | 0.067 |
| Sex, n (%) | 0.530 | |||
| Male | 7 (70.0) | 3 (75.0) | 10 (71.4) | |
| Female | 3 (30.0) | 1 (25.0) | 4 (28.6) | |
| Academic title, n (%) | 0.089 | |||
| Professor | 4 (40.0) | 0 (0.0) | 4 (28.6) | |
| Associate Professor | 4 (40.0) | 0 (0.0) | 4 (28.6) | |
| Specialist Physician | 2 (20.0) | 4 (100.0) | 6 (42.9) | |
| Institution, n (%) | — | |||
| Ankara Bilkent City Hospital | 3 (30.0) | 1 (25.0) | 4 (28.6) | |
| Ankara Etlik City Hospital | 3 (30.0) | 1 (25.0) | 4 (28.6) | |
| Elazığ Fethi Sekin City Hospital | 2 (20.0) | 1 (25.0) | 3 (21.4) | |
| Etimesgut Şehit Sait Ertürk State Hospital | 2 (20.0) | 1 (25.0) | 3 (21.4) |
*Statistically significant (p < 0.05)
Continuous variables are presented as mean ± standard deviation and compared using independent samples t-test. Categorical variables are presented as n (%) and compared using Fisher’s exact test
Abbreviations: SD, standard deviation
Gastroenterology specialists had a mean clinical experience of 16.0 ± 3.7 years following residency completion. The academic ranks among gastroenterologists included four professors, four associate professors, and two senior specialists. These participants were recruited from Bilkent City Hospital, Etlik City Hospital, and Elazığ Fethi Sekin Education and Research Hospital. Emergency medicine specialists had a mean clinical experience of 11.8 ± 1.7 years, and all four participants were recruited from Etimesgut Şehit Sait Ertürk State Hospital Emergency Department. The difference in clinical experience between the two specialty groups was statistically significant.
Overall performance comparison
The overall accuracy of all participant groups and AI models is presented in Table 2. The genetic neuro-symbolic LLM system achieved perfect accuracy by correctly answering all 30 questions, significantly outperforming all other AI models and human expert groups.
Table 2.
Overall diagnostic accuracy of AI models and human expert groups in acute cholangitis management based on tokyo guidelines 2018
| Participant / Model | Correct Answers (n) | Accuracy (%) | 95% CI | p-value vs. NS-LLM | p-value vs. Gastro |
|---|---|---|---|---|---|
| Artificial Intelligence Models | |||||
| Neuro-Symbolic LLM | 30 | 100.0 | 88.4–100.0 | Reference | 0.018* |
| Claude 4.5 Sonnet | 27 | 90.0 | 73.5–97.9 | 0.042* | 0.089 |
| Gemini 2.0 Flash | 19 | 63.3 | 43.9–80.1 | < 0.001* | < 0.001* |
| ChatGPT 5.2 | 18 | 60.0 | 40.6–77.3 | < 0.001* | < 0.001* |
| Human Expert Groups | |||||
| Gastroenterology Specialists ( n = 10) | 28.7 ± 0.9 | 95.7 ± 3.2 | 92.8–98.6 | 0.018* | Reference |
| Emergency Medicine Specialists ( n = 4) | 25.3 ± 2.6 | 84.2 ± 8.8 | 70.2–98.2 | 0.003* | 0.012* |
*Statistically significant (p < 0.05)
AI model accuracy is based on single assessment (30 questions). Human expert data are presented as mean ± standard deviation. Pairwise comparisons between AI models were performed using McNemar’s test. Comparisons involving human expert groups were performed using chi-square test
Abbreviations: NS-LLM, neuro-symbolic large language model; CI, confidence interval; Gastro, gastroenterology specialists
Among conventional large language models, Claude 4.5 Sonnet demonstrated the highest accuracy with 27 correct answers out of 30 questions. Gemini 2.0 Flash correctly answered 19 questions, while ChatGPT 5.2 correctly answered 18 questions. The gastroenterology specialist group achieved a mean accuracy of 95.7% ± 3.2% with individual scores ranging from 90.0% to 100%. Emergency medicine specialists achieved a mean accuracy of 84.2% ± 8.8% with individual scores ranging from 73.3% to 93.3%.
Pairwise statistical comparisons revealed significant differences between the neuro-symbolic system and all other groups. The neuro-symbolic system significantly outperformed Claude 4.5 Sonnet, ChatGPT 5.2, Gemini 2.0 Flash, and both human expert groups. The gastroenterology expert group demonstrated statistically superior performance compared to ChatGPT 5.2 and Gemini 2.0 Flash. Notably, gastroenterologists achieved performance comparable to Claude 4.5 Sonnet with a non-significant difference, suggesting that domain expertise in hepatobiliary disorders approaches the performance of the best-performing conventional large language model.
Performance by clinical domain
Performance stratified by clinical domain is presented in Fig. 2. The neuro-symbolic system achieved perfect accuracy across all three clinical domains.
Fig. 2.
Performance Comparison by Clinical Domain and Tokyo Guidelines 2018 Severity Grade. This figure presents a grouped bar chart displaying accuracy percentages across three clinical domains (diagnosis, treatment, and complications/prognosis) and three Tokyo Guidelines 2018 severity grades (Grade I, Grade II, and Grade III) for all participant groups. The neuro-symbolic large language model system achieved 100% accuracy across all categories. Error bars represent 95% confidence intervals for human expert groups. Asterisks denote statistically significant differences compared to the gastroenterology specialist group (* p < 0.05, ** p < 0.01). The treatment domain and Grade II severity classification demonstrated the greatest performance variability among conventional large language models
Diagnosis domain
The diagnosis domain consisted of 10 questions assessing the application of Tokyo Guidelines 2018 diagnostic criteria. These questions evaluated recognition of Charcot’s triad and Reynolds’ pentad, interpretation of laboratory findings including white blood cell count, C-reactive protein, bilirubin, and liver enzymes, and integration of imaging findings. Gastroenterology specialists achieved significantly higher accuracy compared to emergency medicine specialists in this domain. The most common diagnostic errors among human participants involved afebrile presentations and difficulty distinguishing acute cholangitis from acute cholecystitis with biliary obstruction. Among conventional AI models, ChatGPT 5.2 and Gemini 2.0 Flash frequently failed to apply the three-domain TG18 diagnostic criteria systematically and instead relied on pattern matching to Charcot’s triad alone.
Treatment domain
The treatment domain consisted of 10 questions evaluating selection of appropriate antibiotic regimens according to TG18 recommendations, timing of biliary drainage based on severity grade, drainage modality selection including ERCP versus percutaneous transhepatic biliary drainage versus endoscopic ultrasound-guided drainage, and management of anticoagulation during urgent procedures. Gastroenterology specialists significantly outperformed emergency medicine physicians in this domain. The performance gap was most pronounced in questions involving drainage modality selection in altered anatomy and anticoagulation management during urgent ERCP. Among AI models, treatment questions demonstrated the highest error rate. ChatGPT 5.2 achieved only 50% accuracy with errors predominantly in drainage timing and antibiotic selection.
Complications and prognosis domain
The complications and prognosis domain consisted of 10 questions assessing recognition and management of cholangitis complications including hepatic abscess, septic shock, acute kidney injury, and multi-organ dysfunction. Additional questions addressed response assessment criteria and prognostic factor identification. Gastroenterology specialists achieved higher accuracy compared to emergency medicine specialists in this domain, although the difference did not reach statistical significance. Both groups demonstrated difficulty with questions involving Grade III to Grade II de-escalation criteria and recurrence risk estimation following successful treatment. Claude 4.5 Sonnet maintained 90% accuracy in this domain, while ChatGPT 5.2 and Gemini 2.0 Flash showed substantial deficits in recognizing early signs of treatment failure and indications for repeat intervention.
Performance by TG18 severity grade
Performance stratified by Tokyo Guidelines 2018 severity classification is presented in Fig. 2. Severity-specific accuracy data are presented in Table 3. All participant groups demonstrated highest accuracy in Grade III severe cases, likely due to the unambiguous clinical presentation and clear management imperatives.
Table 3.
Performance Stratified by Clinical Domain, TG18 Severity Grade, and Etiology (Accuracy %)
| Category | NS-LLM | Claude 4.5 | ChatGPT 5.2 | Gemini 2.0 | Gastroenterology (mean ± SD) | Emergency Med (mean ± SD) | p-value (GE vs. EM) |
|---|---|---|---|---|---|---|---|
| Clinical Domain | |||||||
| Diagnosis (n = 10) | 100.0 | 90.0 | 70.0 | 70.0 | 96.0 ± 5.2 | 82.5 ± 9.6 | 0.018* |
| Treatment (n = 10) | 100.0 | 90.0 | 50.0 | 50.0 | 96.0 ± 5.2 | 80.0 ± 8.2 | 0.003* |
| Complications/Prognosis (n = 10) | 100.0 | 90.0 | 60.0 | 70.0 | 95.0 ± 7.1 | 90.0 ± 8.2 | 0.156 |
| TG18 Severity Grade | |||||||
| Grade I – Mild (n = 8) | 100.0 | 87.5 | 62.5 | 62.5 | 93.8 ± 6.3 | 81.3 ± 12.0 | 0.042* |
| Grade II – Moderate (n = 12) | 100.0 | 91.7 | 58.3 | 66.7 | 95.8 ± 4.2 | 83.3 ± 9.1 | 0.008* |
| Grade III – Severe (n = 10) | 100.0 | 90.0 | 60.0 | 60.0 | 97.0 ± 4.8 | 87.5 ± 9.6 | 0.067 |
| Etiology | |||||||
| Choledocholithiasis (n = 12) | 100.0 | 91.7 | 66.7 | 66.7 | 97.5 ± 3.2 | 87.5 ± 7.2 | 0.012* |
| Malignant Biliary Obstruction (n = 6) | 100.0 | 83.3 | 50.0 | 66.7 | 95.0 ± 5.5 | 83.3 ± 10.5 | 0.034* |
| Post-procedural/Iatrogenic (n = 6) | 100.0 | 100.0 | 66.7 | 50.0 | 95.0 ± 5.5 | 79.2 ± 12.5 | 0.021* |
| Benign Strictures (n = 4) | 100.0 | 75.0 | 50.0 | 75.0 | 92.5 ± 9.6 | 81.3 ± 12.5 | 0.089 |
| Parasitic Cholangitis (n = 2) | 100.0 | 100.0 | 50.0 | 50.0 | 95.0 ± 10.0 | 75.0 ± 20.4 | 0.156 |
*Statistically significant (p < 0.05)
AI model data are presented as accuracy percentages. Human expert data are presented as mean ± standard deviation. Statistical comparisons between gastroenterology (GE) and emergency medicine [16] specialist groups were performed using independent samples t-test
Color coding: Green (> 95%), Yellow (65–85%), Red (< 65%)
Abbreviations: NS-LLM, neuro-symbolic large language model; TG18, Tokyo Guidelines 2018; SD, standard deviation
Grade I mild cholangitis questions consisted of 8 items. The neuro-symbolic system achieved perfect accuracy, while Claude 4.5 Sonnet achieved 87.5% accuracy. ChatGPT 5.2 and Gemini 2.0 Flash both achieved 62.5% accuracy. Gastroenterology specialists achieved significantly higher accuracy compared to emergency medicine specialists in Grade I cases. Errors in Grade I cases predominantly involved over-triage with recommendations for urgent drainage in cases meeting only elective drainage criteria.
Grade II moderate cholangitis questions consisted of 12 items and proved most challenging for conventional large language models. The neuro-symbolic system achieved perfect accuracy, while frequent misclassification as either Grade I or Grade III occurred among other AI models. The explicit encoding of the Grade II “two of five” criteria in the neuro-symbolic system enabled consistent correct classification. Gastroenterology specialists significantly outperformed emergency medicine specialists in Grade II cases.
Grade III severe cholangitis questions consisted of 10 items. Despite clear clinical severity, conventional large language models frequently erred in organ support recommendations and drainage timing optimization. Gastroenterology specialists achieved higher accuracy compared to emergency medicine specialists, although this difference did not reach statistical significance.
Performance by etiology
Performance stratified by acute cholangitis etiology is presented in Fig. 3. Choledocholithiasis-related questions numbered 12 and demonstrated the highest overall accuracy across all groups. Familiarity with this common presentation likely contributed to superior performance. Malignant biliary obstruction questions numbered 6 and showed highest error rates among conventional large language models in questions involving palliation versus curative intent and stent selection. Post-procedural and iatrogenic cholangitis questions numbered 6 and required integration of procedural history with current clinical presentation. Benign stricture questions numbered 4 and required nuanced management approaches for chronic pancreatitis-related and post-surgical strictures. Parasitic cholangitis questions numbered 2 and required recognition of specific geographic and exposure risk factors related to Ascaris and liver fluke infections.
Fig. 3.
Performance Comparison by Acute Cholangitis Etiology. This figure presents a grouped bar chart displaying accuracy percentages stratified by five etiological categories of acute cholangitis. Choledocholithiasis (n = 12 questions) demonstrated the highest overall accuracy across all groups, while malignant biliary obstruction (n = 6 questions) and parasitic cholangitis (n = 2 questions) showed the greatest performance variability among conventional large language models. Error bars represent 95% confidence intervals for human expert groups. The neuro-symbolic system achieved 100% accuracy across all etiological categories
Pairwise accuracy differences
A forest plot of pairwise accuracy differences versus the neuro-symbolic LLM system (reference: 100%) is presented in Fig. 4. The Newcombe method 95% confidence intervals demonstrated that the accuracy difference between the NS-LLM and gastroenterology specialists (4.3% points; 95% CI: −2.1 to 14.8) did not exceed the pre-specified 10%-point non-inferiority margin, suggesting that gastroenterologists approached NS-LLM-level performance. In contrast, the accuracy differences for ChatGPT 5.2 (40.0% points; 95% CI: 19.2 to 55.3) and Gemini 2.0 Flash (36.7% points; 95% CI: 15.9 to 52.8) were substantially larger, with confidence intervals entirely above the non-inferiority margin, confirming significantly inferior performance. Claude 4.5 Sonnet demonstrated an intermediate difference (10.0% points; 95% CI: −2.4 to 33.5) with a confidence interval crossing both the zero line and the non-inferiority margin, indicating a statistically significant but clinically moderate performance gap. Emergency medicine specialists showed an accuracy difference of 15.7% points (95% CI: 1.2 to 37.4), with the lower bound of the confidence interval exceeding zero, confirming significantly lower accuracy compared to the NS-LLM system.
Fig. 4.
Forest Plot of Pairwise Accuracy Differences versus Neuro-Symbolic LLM System. This figure presents pairwise accuracy differences (percentage points) between the neuro-symbolic LLM system (reference: 100%) and each comparator group, with 95% Newcombe confidence intervals. The dashed vertical line represents no difference (0% points); the dotted vertical line represents the pre-specified 10%-point non-inferiority margin. Gastroenterology specialists demonstrated the smallest accuracy difference (4.3 pp; 95% CI: −2.1 to 14.8), with a confidence interval not exceeding the non-inferiority margin. ChatGPT 5.2 (40.0 pp) and Gemini 2.0 Flash (36.7 pp) showed the largest differences, with confidence intervals entirely above the non-inferiority margin. Statistical significance versus the NS-LLM system: *p < 0.05, **p < 0.01
Inter-rater reliability
Fleiss’ kappa coefficient for inter-rater agreement among gastroenterology specialists indicated almost perfect agreement according to Landis and Koch criteria. Emergency medicine specialists demonstrated substantial agreement. When comparing across all human participants, overall agreement remained high.
Questions with lowest agreement among human experts were concentrated in the treatment domain. These included questions addressing biliary access in surgically altered anatomy, anticoagulation management during urgent ERCP, and de-escalation criteria from Grade III to Grade II management. These areas of clinical controversy reflect ongoing debates in the hepatobiliary community regarding optimal management strategies in complex scenarios not fully addressed by TG18.
Error pattern analysis
Qualitative and quantitative analysis of incorrect responses revealed distinct error patterns among AI models. The neuro-symbolic system achieved perfect accuracy by leveraging its symbolic reasoning component to explicitly map clinical features to TG18 criteria before answer selection.
Claude 4.5 sonnet error patterns
Claude 4.5 Sonnet demonstrated the highest accuracy among conventional large language models with only 3 errors out of 30 questions. The first error involved biliary access approach selection in surgically altered anatomy, where Claude incorrectly selected percutaneous transhepatic cholangiography over ERCP with device-assisted enteroscopy for a patient with Roux-en-Y anatomy presenting with Grade II cholangitis. Current TG18 recommendations support attempting endoscopic approaches first when institutional expertise is available. The second error involved anticoagulation management during urgent biliary drainage, where Claude recommended complete reversal of anticoagulation before ERCP in a patient with Grade III cholangitis on warfarin. Current guidelines support proceeding with urgent drainage with partial reversal only given the life-threatening nature of severe cholangitis. The third error involved Grade III to Grade II de-escalation criteria, where Claude failed to recognize appropriate de-escalation following successful biliary drainage.
ChatGPT 5.2 error patterns
ChatGPT 5.2 exhibited systematic errors across multiple domains with 12 errors out of 30 questions. TG18 severity grading errors accounted for 5 incorrect answers, with consistent misapplication of TG18 severity criteria particularly involving the Grade II “two of five” risk factor assessment. Treatment selection errors accounted for 4 incorrect answers, including inappropriate drainage modalities and antibiotic regimens. Diagnostic reasoning failures accounted for 2 incorrect answers, with failure to apply TG18 diagnostic criteria systematically. Complications and prognosis errors accounted for 1 incorrect answer involving recurrence risk estimation.
Gemini 2.0 flash error patterns
Gemini 2.0 Flash showed particular difficulty with multi-modal data integration and severity assessment with 11 errors out of 30 questions. Laboratory-imaging correlation failures accounted for 4 incorrect answers, with failure to synthesize laboratory findings with imaging characteristics. Severity classification errors accounted for 4 incorrect answers, with systematic underestimation of disease severity particularly in cases with borderline organ dysfunction. Treatment timing errors accounted for 2 incorrect answers, including inappropriate drainage timing recommendations. Response assessment failures accounted for 1 incorrect answer involving failure to recognize signs of treatment failure.
Comparative error analysis by clinical domain
Analysis of error distribution by clinical domain revealed significant differences in AI model performance (Fig. 5). The diagnosis domain showed moderate error rates among conventional large language models with errors predominantly involving failure to apply TG18’s three-domain diagnostic criteria systematically. The treatment domain showed the highest error rates among conventional large language models. The neuro-symbolic system’s genetic algorithm-optimized prompts and explicit TG18 guideline mapping effectively addressed the challenge of selecting appropriate drainage timing, modality, and antibiotic regimens based on severity grade. The complications and prognosis domain showed intermediate error rates with questions involving response assessment, de-escalation criteria, and recurrence risk estimation proving challenging for conventional large language models lacking explicit encoding of TG18 follow-up algorithms.
Fig. 5.
Heatmap of error distribution by clinical domain and AI model. This figure presents a heatmap displaying the distribution of errors across three clinical domains (diagnosis, treatment, and complications/prognosis) for each artificial intelligence model. Color intensity represents accuracy percentage, with darker shades indicating higher accuracy (dark green: >95%, light green: 85–95%, yellow: 70–84%, orange: 55–69%, red: <55%). The neuro-symbolic system achieved 100% accuracy across all domains (depicted in dark green). The treatment domain demonstrated the highest error rates among conventional large language models, with ChatGPT 5.2 and Gemini 2.0 Flash both achieving only 50% accuracy. The diagnosis domain showed moderate error rates, while the complications and prognosis domain demonstrated intermediate performance across conventional models
The neuro-symbolic system correctly answered all questions by explicitly querying its TG18 knowledge base for each clinical decision point, generating auditable reasoning chains that mapped clinical features to guideline criteria.
Discussion
This multi-center, cross-sectional study offers the first comprehensive comparison between a neuro-symbolic large language model (NS-LLM) system, conventional large language models (LLMs), and experienced physicians in the management of acute cholangitis. The primary finding of this research is the superior diagnostic and therapeutic accuracy demonstrated by the NS-LLM system, which achieved perfect performance (100%) across all clinical domains, severity grades, and etiological categories. This outcome significantly surpasses the performance of the leading conventional LLM (Claude 4.5 Sonnet, 90%) and highly experienced gastroenterology specialists (mean 95.7%). These results have important implications for the future development and implementation of artificial intelligence in clinical decision support systems for intricate hepatobiliary conditions.
The performance of conventional LLMs observed in our study aligns with recent systematic evaluations of AI in gastroenterology and hepatology [17]. Wiest et al. emphasized in their comprehensive review that, although large language models exhibit promising capabilities in processing unstructured clinical text and integrating diverse information sources, their reliability in complex clinical scenarios remains inconsistent [3]. Similarly, a study by Safavi-Naini et al. revealed that proprietary models such as Claude 3.5 Sonnet achieved 74% accuracy on gastroenterology board-style questions, while open-source alternatives lagged significantly behind [18]. Our findings corroborate these observations, with Claude 4.5 Sonnet demonstrating the highest accuracy among conventional LLMs (90%), whereas ChatGPT 5.2 (60%) and Gemini 2.0 Flash (63.3%) exhibited substantially lower performance. The treatment domain proved particularly challenging for conventional models, consistent with reports that LLMs frequently struggle with multi-step clinical reasoning requiring systematic guideline application [19].
The superior performance of our NS-LLM system can be attributed to its unique architectural integration of neural pattern recognition with explicit symbolic reasoning. Prenosil et al. demonstrated that neuro-symbolic approaches connecting GPT-4 with rule-based expert systems through semantic integration platforms can achieve physician-level accuracy while providing traceable, deterministic outputs [20]. Our system extends this paradigm by incorporating the Tokyo Guidelines 2018 as a structured knowledge graph, enabling explicit encoding of diagnostic criteria, severity grading algorithms, and management pathways. This design directly confronts the core limitation of traditional large language models (LLMs), as identified by Kim et al. (2025) in Scientific Reports, specifically their rigid reasoning capabilities that hinder the systematic application of clinical algorithms to nuanced cases [21]. The explicit symbolic encoding of TG18’s three-domain diagnostic criteria and the “two of five” Grade II severity assessment enabled our system to achieve perfect classification where conventional models frequently erred.
The multi-agent architecture utilized within our system exemplifies an emerging paradigm in medical artificial intelligence, demonstrating considerable performance enhancements over single-model methodologies. A recent systematic review revealed that AI agent systems consistently surpassed baseline large language models in executing clinical tasks, with improvements spanning from modest gains to increases exceeding 60% points in accuracy when the architectural complexity aligned with the specific requirements of the tasks [22]. Chen et al.developed a Multi-Agent Conversation (MAC) framework for disease diagnosis that outperformed single models in both diagnostic accuracy and suggested test appropriateness, achieving optimal performance with four doctor agents coordinated by a supervisor agent [23]. Our dual-agent deployment strategy utilizing Gemini 2.0 Flash and GPT-5.2 as parallel reasoning entities, combined with a consensus arbitration module, mirrors these successful multi-agent implementations.
The redundancy provided by architecturally distinct models reduces single-point-of-failure errors, while cross-validation of reasoning pathways increases confidence in generated responses, as supported by recent theoretical frameworks for agentic AI in healthcare [24, 25].
An advantage of our neuro-symbolic approach is its potential to mitigate hallucinations, a phenomenon that represents the most significant barrier to clinical deployment of conventional LLMs. Studies have shown that large language models exhibit adversarial hallucination rates ranging from approximately 50% to 82% when presented with clinical vignettes containing fabricated details, and even the best-performing model reduced but did not eliminate these errors under mitigation strategies [26]. Asgari et al. (2025) reported a 1.47% hallucination rate and 3.45% omission rate in clinical note summarization tasks, emphasizing the need for robust safety frameworks [22]. Recent systematic analyses indicate that medical LLMs exhibit hallucination rates of 15% to 40% on clinical tasks, raising concerns about deployment readiness [26]. Our symbolic knowledge base serves as a constraint-checking mechanism that validates candidate answers against hard-coded TG18 criteria, rejecting responses that violate established diagnostic and therapeutic guidelines. This architectural safeguard, combined with Monte Carlo dropout-based uncertainty quantification, provides multiple layers of protection against clinically consequential errors.
The comparison with human expert physicians elucidated significant insights regarding the potential function of AI-assisted clinical decision support. Gastroenterology specialists attained exemplary performance, with a mean accuracy of 95.7%, approaching but not equaling the NS-LLM system, whereas emergency medicine physicians exhibited somewhat lower accuracy, with a mean of 84.2%, particularly in the domains of treatment decisions and severity grading. This performance differential between specialties underscores the significance of domain-specific expertise in managing complex hepatobiliary conditions and indicates that AI decision support systems could be especially beneficial in non-specialist environments where such expertise is less accessible. Notably, questions involving biliary access in surgically altered anatomy, anticoagulation management during urgent ERCP, and de-escalation criteria demonstrated the lowest inter-rater agreement among human experts, reflecting areas of ongoing clinical controversy not fully addressed by current guidelines.
These findings align with Berry et al. (2025), who proposed a structured framework for integrating LLMs into gastroenterology practice, emphasizing the importance of multidisciplinary collaboration and continuous validation in real-world settings [27].
The clinical implications of our findings extend beyond performance metrics to the broader question of how AI can be responsibly integrated into hepatobiliary clinical practice. Yuan et al. highlighted that agentic LLMs can access research findings, clinical case reports, and updated guidelines without additional training, enabling them to tackle complex tasks requiring iterative reasoning that align more closely with golden clinical procedures [28]. Our system’s ability to generate explainable reasoning chains that trace activated rules and supporting evidence addresses the interpretability requirements emphasized by Vidal et al. as essential for trustworthy deployment in medicine [29]. The perfect accuracy achieved across different cholangitis etiologies, including relatively rare presentations such as parasitic cholangitis and post-procedural causes, suggests potential utility in educational settings and as a second-opinion tool for complex cases. However, as Soroush et al. cautioned in Gastroenterology, responsible deployment requires addressing challenges including output reliability, human-AI teaming, and infrastructure demands through comprehensive risk mitigation frameworks [30].
Our results contribute to a rapidly evolving literature on AI performance in specialized medical domains. Gaber et al. evaluated LLM workflows in clinical decision support for triage and diagnosis, demonstrating that sophisticated prompting and retrieval strategies can substantially improve performance [31]. The deficits observed in the treatment domain within our conventional LLM evaluation mirror the findings of Ong et al., who demonstrated that LLMs used as clinical decision support systems for medication safety exhibited variable performance across 16 clinical specialties [32]. Ferber et al. developed an autonomous AI agent for oncology treatment planning that achieved significant improvements through tool-augmented reasoning, conceptually similar to our genetic algorithm optimization and symbolic constraint checking [33]. The convergence of these findings suggests that hybrid architectures combining neural flexibility with structured reasoning represent a promising direction for clinical AI development, particularly in domains requiring strict adherence to established guidelines.
It is important to contextualize the present work within emerging reporting standards for AI-based health tools. The CHART checklist (Checklist for AI/ML Research Translation, BMJ 2025) and TRIPOD-LLM statement (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis adapted for Large Language Models, Nature Medicine 2025) provide standardized frameworks for evaluating the methodological rigor of LLM-based clinical decision support systems [13, 14]. A self-assessment of the present study against these frameworks reveals several areas of alignment, including the transparent description of the multi-agent architecture, explicit encoding of clinical guideline criteria, structured evaluation against a validated question set, and comparison with both AI and human baselines. However, important gaps remain: the study lacks prospective clinical validation with real patient encounters, external generalizability testing across diverse healthcare systems and patient populations, formal assessment of clinical workflow integration and end-user acceptance, and comprehensive reporting of system latency and computational requirements. Additionally, while the system generates auditable reasoning chains, formal usability testing evaluating whether clinicians can effectively interpret and act upon these explanations has not been conducted. We encourage future studies in this domain to adopt these reporting frameworks as standard practice to facilitate cross-study comparisons and accelerate the responsible translation of AI-based clinical decision support tools.
Several limitations of this study merit acknowledgment. Firstly, the evaluation employed board-style multiple-choice questions rather than real-world clinical encounters, which may not comprehensively represent the intricacies of actual patient management, including scenarios with incomplete information, communication challenges, and time constraints. Secondly, the relatively small sample size of human participants, especially within the emergency medicine group (n = 4), restricts the extent to which comparative conclusions can be generalized. Thirdly, the NS-LLM system’s knowledge base was constructed solely on TG18, and its performance on cases outside these guidelines or requiring the integration of updated evidence remains unassessed. Fourthly, the study was conducted within a single country (Turkey), and the results may vary across different healthcare systems and patient populations. Fifthly, although the system attained perfect accuracy on the test set, the possibility of overfitting to the question format cannot be disregarded, and validation using novel cases in prospective studies is imperative. Lastly, the computational demands and latency of the multi-agent system were not systematically examined, which has significant implications for its application in real-time clinical settings implementation.
Furthermore, the current system’s knowledge base is built exclusively on the Tokyo Guidelines 2018, without integration of complementary international guidelines such as those published by the American Society for Gastrointestinal Endoscopy (ASGE) for the role of endoscopy in biliary tract diseases or the European Society of Gastrointestinal Endoscopy (ESGE) clinical guidelines for biliary drainage and management. Future iterations should incorporate these guidelines along with regional adaptations reflecting local antimicrobial resistance patterns, available endoscopic expertise, and healthcare infrastructure variability across different settings. Additionally, the system does not currently account for facility-specific expertise when generating drainage modality recommendations. The choice of drainage methods, particularly for postoperative reconstructed intestinal segments, and the indications for device-assisted enteroscopy-ERCP, balloon endoscopy, and endoscopic ultrasound-guided drainage depend critically on the expertise available at each institution. Future versions should incorporate a facility capability profile as an input parameter, enabling drainage recommendations to be tailored to locally available expertise, consistent with TG18’s own acknowledgment that modality selection should consider institutional experience.
An additional limitation concerns the severity grading algorithm’s handling of contextual clinical information. The current implementation applies the creatinine threshold (> 2.0 mg/dL) as a binary rule without distinguishing acute kidney injury from pre-existing chronic kidney disease (CKD), potentially leading to erroneous Grade III severity classification in patients with stable baseline renal impairment. Similarly, the PT-INR threshold (> 1.5) for hepatic dysfunction is applied without checking for anticoagulant use (warfarin, direct oral anticoagulants, or heparin), which could result in misclassification of drug-induced coagulopathy as cholangitis-related hepatic dysfunction. Future iterations should incorporate baseline laboratory value comparisons (e.g., delta-creatinine from baseline) and medication history integration to enable contextually appropriate severity grading, which would necessitate integration with electronic health record systems providing access to longitudinal patient data.
Moreover, the current system treats all imaging modalities equivalently within Domain C (Biliary Imaging) of the TG18 diagnostic criteria once biliary dilatation or etiology identification criteria are met. In clinical practice, however, ultrasonography has lower sensitivity for distal choledocholithiasis compared to magnetic resonance cholangiopancreatography (MRCP) or endoscopic ultrasound (EUS), and the diagnostic accuracies and invasiveness profiles of these modalities differ substantially. Incorporating modality-specific positive and negative likelihood ratios into the knowledge graph would enable more nuanced diagnostic confidence stratification and could inform evidence-based imaging recommendation pathways that account for both diagnostic yield and procedural invasiveness.
Future research directions should address these limitations while exploring several promising avenues. Extension of the symbolic knowledge base to incorporate additional guidelines, emerging evidence, and local antimicrobial resistance patterns would enhance applicability. Integration with electronic health record systems for real-time clinical decision support warrants investigation, with attention to workflow integration and user acceptance. Multi-institutional prospective validation across diverse healthcare settings is essential to establish generalizability. Development of uncertainty communication interfaces that appropriately convey confidence levels to clinicians represents an important human factors consideration. Finally, comparative cost-effectiveness analyses examining the resource implications of deploying neuro-symbolic systems versus expanding specialist access would inform implementation decisions.
In conclusion, this study demonstrates that a genetic neuro-symbolic LLM system integrating multi-agent orchestration with explicit clinical guideline encoding achieves superior performance in acute cholangitis management compared to both conventional LLMs and human expert physicians. The architectural innovations of parallel neural reasoning, symbolic constraint validation, and consensus arbitration address fundamental limitations of purely neural approaches, including hallucination risk, guideline misapplication, and inconsistent multi-step reasoning. These findings provide preliminary evidence that neuro-symbolic architectures may offer a promising direction for AI-assisted clinical decision support in complex hepatobiliary disease and serve as a template for developing similar systems across other guideline-driven medical domains. However, this proof-of-concept study, based on structured multiple-choice questions, requires validation through prospective clinical studies with real patient encounters, integration of complementary international guidelines (ASGE, ESGE), and formal usability testing with clinician end-users before broader clinical implementation claims can be warranted.
Supplementary Information
Below is the link to the electronic supplementary material.
Acknowledgements
The authors gratefully acknowledge the gastroenterology and emergency medicine specialists from Ankara Bilkent City Hospital, Ankara Etlik City Hospital, Elazığ Fethi Sekin City Hospital, and Etimesgut Şehit Sait Ertürk State Hospital who generously dedicated their time to participate in the expert evaluation panel. The authors also thank the Non-Interventional Ethics Committee of the Ankara Provincial Health Directorate for the timely review of the study protocol. No generative artificial intelligence tools were used in the writing of this manuscript beyond their role as study participants under formal evaluation.
Abbreviations
- ABIM
American Board of Internal Medicine
- AI
Artificial intelligence
- AKI
Acute kidney injury
- ALP
Alkaline phosphatase
- ALT
Alanine aminotransferase
- API
Application programming interface
- AST
Aspartate aminotransferase
- BP
Blood pressure
- CBD
Common bile duct
- CI
Confidence interval
- CRP
C-reactive protein
- CT
Computed tomography
- ERCP
Endoscopic retrograde cholangiopancreatography
- EUS
Endoscopic ultrasound
- GA
Genetic algorithm
- GCS
Glasgow Coma Scale
- GGT
Gamma-glutamyl transferase
- GPT
Generative pre-trained transformer
- ICU
Intensive care unit
- INR
International normalized ratio
- IQR
Interquartile range
- LLM
Large language model
- LLN
Lower limit of normal
- MRCP
Magnetic resonance cholangiopancreatograph
- NER
Named entity recognition
- NS-LLM
Neuro-symbolic large language model
- PTBD
Percutaneous transhepatic biliary drainage
- RUQ
Right upper quadrant
- SD
Standard deviation
- STARD
Standards for Reporting Diagnostic Accuracy
- STROBE
Strengthening the Reporting of Observational Studies in Epidemiology
- TG18
Tokyo Guidelines 2018
- ULN
Upper limit of normal
- WBC
White blood cell
Author contributions
M.U: Conceptualization, Methodology, Formal analysis, Data curation, Writing – Original Draft, Visualization. E.E: Supervision, Project administration, Writing – Review & Editing. All authors have read and approved the final version of the manuscript.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
The de-identified datasets generated and analyzed during the current study—including the 30 case-based ABIM-adapted question stems, model output logs, and per-question scoring sheets—are available from the corresponding author (M.U., meteucdal@hacettepe.edu.tr) upon reasonable request, subject to a data-use agreement and compliance with copyright restrictions related to the ABIM question source material. The source code of the genetic neuro-symbolic LLM system contains proprietary components and is not publicly available; however, a detailed architectural description sufficient to reproduce the system is provided in the Methods section and Supplementary Material.
Declarations
Ethics approval and consent to participate
This study was conducted in accordance with the Declaration of Helsinki. The Non-Interventional Ethics Committee of the Ankara Provincial Health Directorate approved the study protocol (Approval No: 2025-11-15, dated October 15, 2025). All participating physicians provided written informed consent prior to study enrollment. As this study utilized hypothetical clinical vignettes and did not involve any real patient data or direct patient care, patient consent was not applicable. The ethics committee confirmed that patient consent was not required given the simulation-based nature of the study.
Consent for publication
Not applicable. This study utilized hypothetical clinical vignettes adapted from standardized board examination question banks and did not involve any real patient data, identifiable images, or clinical details that could compromise patient anonymity.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Li W, Mou Z, Yu S, Zhang C, Qi H, Wang G. Predictive value of composite nutritional indicators geriatric nutritional risk index and controlling nutritional status for mortality risk in early-onset cancer survivors. Front Nutr. 2025;12:2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Gaber F, Shaik M, Allega F, Bilecz AJ, Busch F, Goon K, et al. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis. npj Digit Med. 2025;8(1):263. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Wiest IC, Bhat M, Clusmann J, Schneider CV, Jiang X, Kather JN. Large language models for clinical decision support in gastroenterology and hepatology. Nat Rev Gastroenterol Hepatol. 2025;22(11):773–87. [DOI] [PubMed] [Google Scholar]
- 4.Kim J, Podlasek A, Shidara K, Liu F, Alaa A, Bernardo D. Limitations of large language models in clinical problem-solving arising from inflexible reasoning. Sci Rep. 2025;15(1):39426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Li H, Fu J-F, Python A. Implementing Large Language Models in Health Care: Clinician-Focused Review With Interactive Guideline. J Med Internet Res. 2025;27:e71916. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Cozma M-A, Găman M-A, Srichawla BS, Dhali A, Manan MR, Nahian A, et al. Acute cholangitis: a state-of-the-art review. Annals Med Surg. 2024;86(8):4560–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Kiriyama S, Kozaka K, Takada T, Strasberg SM, Pitt HA, Gabata T, et al. Tokyo Guidelines 2018: diagnostic criteria and severity grading of acute cholangitis (with videos). J Hepato-Biliary-Pancreat Sci. 2018;25(1):17–30. [DOI] [PubMed] [Google Scholar]
- 8.Yu E, Chu X, Zhang W, Meng X, Yang Y, Ji X, et al. Large Language Models in Medicine: Applications, Challenges, and Future Directions. Int J Med Sci. 2025;22(11):2792–801. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Meziane L, Abbaoui W, Abdellaoui S, El Bhiri B, Ziti S. Narrative review on symbolic approaches for explainable artificial intelligence: foundations, challenges, and perspectives. Eng Proc [Internet]. 2025;112(1):39.
- 10.Vidal M-E, Chudasama Y, Huang H, Purohit D, Torrente M. Integrating Knowledge Graphs with Symbolic AI: The Path to Interpretable Hybrid AI Systems in Medicine. J Web Semant. 2025;84:100856. [Google Scholar]
- 11.Prenosil GA, Weitzel TK, Bello SC, Mingels C, Manzini G, Meier LP, et al. Neuro-symbolic AI for auditable cognitive information extraction from medical reports. Commun Med (Lond). 2025;5(1):491. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Nawaz U, Anees-ur-Rahaman M, Saeed Z. A review of neuro-symbolic AI integrating reasoning and learning for advanced cognitive systems. Intell Syst Appl. 2025;26:200541. [Google Scholar]
- 13.Fahim YA, Hasani IW, Kabba S, Ragab WM. Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives. Eur J Med Res. 2025;30(1):848. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Gallifant J, Afshar M, Ameen S, Aphinyanaphongs Y, Chen S, Cacciamani G, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med. 2025;31(1):60–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.https://docs.langchain.com/oss/python/langgraph/graph-api?utm_source=chatgpt.com.
- 16.Ahmed F, Kutluk T, Yurduşen S, Şengelen M, Aydın B, Kirazli M, et al. Palliative Care in Turkey: Insights from experts through key informant interviews. J Cancer Policy. 2024;42:100506. [DOI] [PubMed] [Google Scholar]
- 17.Gong EJ, Bang CS, Lee JJ, Park J, Kim E, Kim S, et al. Large language models in gastroenterology: systematic review. J Med Internet Res. 2024;26:e66648. 10.2196/66648. [DOI] [PMC free article] [PubMed]
- 18.Safavi-Naini SAA, Ali S, Shahab O, Shahhoseini Z, Savage T, Rafiee S, et al. Benchmarking proprietary and open-source language and vision-language models for gastroenterology clinical reasoning. NPJ Digit Med. 2025;8(1):797. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Irfan B, Sirvent R. Large language models and the future of gastroenterology: dissecting the biopolitics of data in a global health ecosystem. Front Med. 2025;12. [DOI] [PMC free article] [PubMed]
- 20.Prenosil GA, Weitzel TK, Bello SC, Mingels C, Manzini G, Meier LP, et al. Neuro-symbolic AI for auditable cognitive information extraction from medical reports. Commun Med. 2025;5(1):491. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Kim J, Podlasek A, Shidara K, Liu F, Alaa A, Bernardo D. Limitations of large language models in clinical problem-solving arising from inflexible reasoning. Sci Rep. 2025;15(1):39426. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Gorenshtein A, Omar M, Glicksberg BS, Nadkarni GN, Klang E. AI agents in clinical medicine: a systematic review. medRxiv. 2025.
- 23.Chen X, Yi H, You M, Liu W, Wang L, Li H, et al. Enhancing diagnostic capability with multi-agents conversational large language models. npj Digit Med. 2025;8(1):159. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Hinostroza Fuentes VG, Karim HA, Tan MJT, AlDahoul N. AI with agency: a vision for adaptive, efficient, and ethical healthcare. Front Digit Health. 2025;7. [DOI] [PMC free article] [PubMed]
- 25.Borkowski AA, Ben-Ari A, Multiagent. AI Systems in Health Care: Envisioning Next-Generation Intelligence. Fed Pract. 2025;42(5):188–94. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Omar M, Sorin V, Collins JD, Reich D, Freeman R, Gavin N, et al. Large language models are highly vulnerable to adversarial hallucination attacks in clinical decision support: a multi-model assurance analysis. medRxiv. 2025:2025.03.18.25324184. [DOI] [PMC free article] [PubMed]
- 27.Berry P, Dhanakshirur RR, Khanna S. Utilizing large language models for gastroenterology research: a conceptual framework. Th Adv Gastroenterol. 2025;18:17562848251328577. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Xu X, Sankar R. Large Language Model Agents for Biomedicine: A Comprehensive Review of Methods, Evaluations, Challenges, and Future Directions. Information. 2025;16(10):894. [Google Scholar]
- 29.Chudasama Y, Huang H, Purohit D, Vidal ME. Toward Interpretable Hybrid AI: Integrating Knowledge Graphs and Symbolic Reasoning in Medicine. IEEE Access. 2025;13:39489–509. [Google Scholar]
- 30.Soroush A, Giuffrè M, Chung S, Shung DL. Generative Artificial Intelligence in Clinical Medicine and Impact on Gastroenterology. Gastroenterology. 2025;169(3):502–e171. [DOI] [PubMed] [Google Scholar]
- 31.Gaber F, Shaik M, Allega F, Bilecz AJ, Busch F, Goon K, et al. Evaluating large language model workflows in clinical decision support for triage and referral and diagnosis. NPJ Digit Med. 2025;8(1):263. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Ong JCL, Jin L, Elangovan K, Lim GYS, Lim DYZ, Sng GGR, et al. Large language model as clinical decision support system augments medication safety in 16 clinical specialties. Cell Rep Med. 2025;6(10):102323. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Ferber D, El Nahhas OSM, Wölflein G, Wiest IC, Clusmann J, Leßmann ME, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat Cancer. 2025;6(8):1337–49. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The de-identified datasets generated and analyzed during the current study—including the 30 case-based ABIM-adapted question stems, model output logs, and per-question scoring sheets—are available from the corresponding author (M.U., meteucdal@hacettepe.edu.tr) upon reasonable request, subject to a data-use agreement and compliance with copyright restrictions related to the ABIM question source material. The source code of the genetic neuro-symbolic LLM system contains proprietary components and is not publicly available; however, a detailed architectural description sufficient to reproduce the system is provided in the Methods section and Supplementary Material.






