Skip to main content
Frontiers in Oncology logoLink to Frontiers in Oncology
. 2026 Aug 19;16:1785951. doi: 10.3389/fonc.2026.1785951

Prospective multi-centre evaluation of guideline−based artificial intelligence to streamline multidisciplinary tumour board

Oleksandr Ivashchuk 1,*, Serhiy Hovornyan 1
PMCID: PMC13533744  PMID: 42688024

Abstract

Introduction

Multidisciplinary tumour board (MTB) or tumour board (TB) is the “gold standard” for providing oncological patients’ diagnosis and treatment. Often, MTBs are time-intensive, capacity-constrained, and absent or inconsistent in many routine hospitals. Despite that, in MTB, a few experienced specialists in the oncological field, sometimes the final decision is mistaken (4-47%), the necessity of treatment plan changing (12-64%), etc. Using artificial intelligence (AI) for cancer patients’ treatment is a modern approach, but the main idea of our investigation should be to possibly replace MTB. On the other hand, the choice of correct treatment in oncological patients for some doctors, like surgeons, gynaecologists, and family medicine doctors, is a big trouble. They need to address the guidelines, which are so complicated quick changing. We have offered our AI algorithm, which connects doctors with patients’ data, guidelines (only official sources), and then uses LLM for decision-making. Main research question – assess the role of our AI algorithm in diagnosis and treatment for cancer patients, compare with or together with MTB.

Methods

Prospective, multicentre study at four specialised regional oncological hospitals. Trail period 2023-2025. Additionally, we included 37 doctors (DR) from general and rural hospitals (15 surgeons, 11 gynaecologists, 11 family medicine doctors). 728 cases in eight tumour groups were included in the study: breast 18.13% (132), colorectal 22.12% (161), lung 30.08% (219), prostate 8.79% (64), gynaecological 8.1% (59), gastric 6.59% (48), urothelial 3.99% (29), head & neck 2.2% (16). They were divided into 4 groups: 1) Diagnosis (D) and treatment plan (TP) with MTB only - 226 patients, 2) D and TP with AI only - 206 patients, 3) D and TP with MTB + AI - 147 patients, 4) D and TP with DR + AI - 149 patients. NCCN guidelines were used for all groups. Patients’ data were ingested from the e-health system, de-identified, and normalized prior to inference. Evaluation criteria. Immediate - preparation time (PT), decision time (DT), time to recommendation (TTR), where TTR = PT + DT; Midterm - change diagnosis, change treatment plan (1 month, 6 months). We have collected MTB reports and for group N4 - questionnaires.

Results

Our proposed AI algorithm works as an intermediary between the patient, the doctor, MTB, guidelines, and LLMs. It processes data at each stage for structuring, unification, conversion, machine-readable representation, creates appropriate Python classes, and ultimately a structured supervisory layer with a decision threshold 80%. For immediate criteria we have next results (min): group 1 - PT – 21.3 ± 3.6; DT – 15.6 ± 6.2; TTR – 36.9 ± 7.8; group 2 - PT – 19.7 ± 4.1; DT – 2.1 ± 0.02; TTR – 22.4 ± 3.3; group 3 - PT – 26.9 ± 4.2; DT – 4.9 ± 0.9; TTR – 29.2 ± 3.1; for midterm criteria (change D) results were: on 1 month – group 1 (6.79%), group 2 (5.36%), group 3 (2.78%); on 6 month – group 1 (18.22%), group 2 (19.8%), group 3 (12.4%); midterm criteria (change TP) results were: on 1 month – group 1 (11.31%), group 2 (8.78%), group 3 (4.17%); on 6 month – group 1 (25.12%), group 2 (22.84%), group 3 (15.5%); midterm criteria (change D + TP) results were: on 1 month – group 1 (13.12%), group 2 (8.78%), group 3 (4.86%); on 6 month – group 1 (28.08%), group 2 (24.3%), group 3 (14.73%). The combination of the MTB with AI gives the most effective result in reducing the change in D and TP from 28.08% to 14.73% within 6 months. For group 4, we have found next results. Before the implementation of AI, the time for TP decision in 78.4% of cases was from 30 to 60 minutes, and in 21.6% - more than 60 minutes. At 6 months, DR has time for TP decision, no more than 30 minutes in 67.7%, and among 30 and 60 minutes, 32.3%. Overall assessment using AI for TP in 3 months (good and very good) – 82.8%, in 6 months – 97.1%. Convenience AI in 3 months (good and very good) – 54.2%, in 6 months – 94.1%. Use AI after the end of the study in 3 months (good and very good) – 54.3%, in 6 months – 54.7%. Use in clinical practice the combination of MTB and AI, optimize TTR, decrease the percentage of change D and TP in 1,6 months, decrease the negative sides of single MTB and AI. We can use this combination for first-detected cases without comorbidities at a basic level. AI cannot replace MTB, but it empowers non-oncological clinicians and DR to generate evidence-based, auditable recommendations in settings with limited specialist capacity.

Discussion

The decrease in the percentage of change D and TP after MTB using AI in the first detected cases highlights the importance of conducting further studies to optimize MTB performance. The positive results of using DR assistance in the form of AI indicate the need to create other practical tools, such as an app, a pocket version, etc. To reach expert-level autonomy, further work is required: a local guideline adaptation, multimodal data ingestion, and longitudinal validation across tumour types, tumour recurrence, multiple cancers, comorbidity, and assessment of the financial efficiency of implementing AI in the work of the MTB.

Keywords: artificial intelligence, cancer, cancer diagnostics, cancer treatment, guideline, LLM, machine learning, tumour board

1. Introduction

Oncology is one of the most complicated fields in modern science of human health. Multidisciplinary tumour board (MTB) (in different countries, it has another name) is the best solution for correct diagnosis (D) and treatment plan (TP) for cancer patients. In MTB, experienced oncologists, oncogynaecologists, breast cancer specialists, radiologists, chemotherapists, pathomorphologists, etc., are included. In spite of that, there are many problems connected with correct D and TP (1–3). Let’s go into more detail.

First of all, incomplete or incorrect data in the patient cards during preparation for the MTB meeting (lack of images, histology, CT/MRI/PET data), different interpretation of images/pathology between specialists, differences in local protocols or resources (regional restrictions on access to therapies/technologies), patient preferences or clinical management steps after the meeting that change the implementation of the recommendation, low quality of evidence for certain clinical situations - leads to variability in recommendations (4, 5).

Secondly, large variability between centres and countries - discrepancies in diagnosis/staging and recommendations are often related to organizational differences, diagnostic availability, data availability, and local protocols. Recent studies show that discordance in recommendations can vary widely depending on the study design and type of comparison, from a few percent to ~50% in narrow comparisons (depending on criteria and quality of evidence) (6, 7).

Thirdly, some criteria for correct assessment in trials, which concern MTB efficacy - different definitions of “change” (what is considered significant) and different protocols for recording changes, different study designs (prospective vs retrospective; single-country vs multi-country), differences in case reporting practices (data completeness, availability of images and pathology), and tumour type. Unfortunately, we don’t have a universal approach for MTB assessment and comparison (4, 8, 9).

Today, the need to change the D or TP occurs in 4 to 67% of cases at different times after MTB (depending on different cancer types) in modern oncological centres. This leads to complications, reduced treatment effectiveness, deterioration in the patient’s quality of life, prolonged treatment duration, and an additional financial burden (10, 11).

On the other hand, if the issue of determining D and TP is assigned to an ordinary doctor (surgeon, gynaecologist, family doctor, thoracic surgeon, urologist, etc.), then the results here are significantly worse compared to MTB. This is due to the fact that for this doctor, oncological patients are one of many others that they deal with. It is much more difficult for them to follow clinical recommendations and their updates, and to treat cancer in certain localizations (11, 12).

How can we increase results for MTB and other doctors, who are responsible for cancer diagnosis and treatment? The most important approach that we can offer is to try to use AI algorithms. Nowadays, AI is widely involved in different areas in oncology: diagnosis, treatment, developing new drugs, etc. It means that AI is rapidly entering clinical oncology as a tool for clinical delivery support (CDS), personalization of therapy (based on images, molecular profiles, and genetic features), patient selection for clinical trials, and prediction of response/outcome. Preliminary reviews and systematic reviews from 2024–2025 show an increase in both publications and early clinical applications, but at the same time, there are notable challenges with validation, interpretability, and regulatory issues (13, 14).

Main areas of application of AI in determining treatment tactics: clinical decision support systems — ranking of therapeutic options, assistance in interpreting molecular tests, recommendations on targeted drugs/immunotherapy (regulatory documents determine when such software is a medical device); radiomics and digital pathology — DL-models extract predictive biomarkers from CT/MR/PET and pathomorphological sections to predict response to chemo/radio/immunotherapy; genomics/omics approaches — models based on tumour sequencing (WES/RNA-seq) for selection of targeted therapies or prediction of sensitivity; Cancer Genome Atlas is a key resource for training models; clinical trial recruitment and protocol optimization — algorithms find suitable patients and predict responses/risks. There are examples of implemented projects and initial research (15, 16).

Among the known, the most common are the following methods and architectures: classical machine learning - logistic regression, random forests, XGBoost — often for tabular clinical data; deep learning - CNN for images, transformers and integrative models for multi-omic data; federated learning and synthetic data — for privacy and sample expansion. An important detail: models that work well on retrospective sets often generalize poorly without external validation (17–19).

Limitations and risks: insufficient external validation and replication of results; difference between correlation and causation - the model can predict, but not substantiate, the mechanism of action of the therapy; risk of overfitting, especially with high-dimensional omics data and small cohorts (20, 21).

Evidence base and validation for AI: mostly retrospective studies and internal cross-validation, there is a growing number of prospective trials and clinical studies, but randomized controlled trials are still limited; regulators (FDA) are issuing guidance for the use of AI in clinical and regulatory decisions; for clinical decision support systems, there are documents that delineate when software is regulated as a medical device; this is critical when developing systems that advise on therapy (17, 22).

Ethical, legal and practical issues that are important when using AI in oncology should not be forgotten: interpretability and clinician trust - “black boxes” reduce adoption, explanatory mechanisms (explainable AI) are needed; bias and inequalities - data containing predominantly patients from a single geography/ethnicity yield biased predictions; liability for erroneous recommendations - an unresolved issue in many jurisdictions; data privacy and security - the need for federated approaches and thorough de-identification (17–19, 23).

Another thing that AI developers for oncology encounter is the widespread use of anthropomorphism among doctors. Anthropomorphism – the integration of artificial AI into oncology has coincided with an expanding use of human-centred terminology. Such anthropomorphism - when algorithms are described as if they possessed cognition, intent, or moral agency - risks distorting conceptual clarity and weakening scientific rigor. Globally, oncological AI literature exhibits a gradual shift toward linguistic precision, privileging technical over metaphorical framing and reflecting a maturing scientific discourse. Reducing reliance on anthropomorphic language may enhance conceptual rigor, prevent diffusion of responsibility, and support more accountable and balanced discourse in AI-driven oncology research (24, 25).

Thus, to date, there is no correct AI tool to address current issues in clinical oncology. Research needs to develop new AI algorithms that would satisfy both highly specialised and non-specialised physicians.

2. Materials and methods

2.1. Study design

For receiving correct results, we have used a prospective, multi-centre study at four specialised regional oncological hospitals, where experienced MTB functioning (consisting of an oncosurgeon, oncogynaecologist, breast cancer specialist, pathologist, radiologist, radiotherapy specialist, chemotherapy specialist, secretary, technical assistant, and others as needed). All specialists in MTB had at least 15 years of experience in their field. For analysis, we collected MTB reports during the period 01.01.2023-01.10.2025. Additionally, we included 37 doctors (DR) from general and rural hospitals (15 surgeons, 11 gynaecologists, 11 family medicine doctors), who are involved in the diagnosis and treatment of cancer patients. They filled out questionnaires in Google Forms. Before starting the study, we had provided 2 hours of training webinars or live lessons on how to use the AI algorithm and collect data for study participants. The training period was 1 month.

728 cases in eight tumour groups were included in the study: breast 18.13% (132), colorectal 22.12% (161), lung 30.08% (219), prostate 8.79% (64), gynaecological 8.1% (59), gastric 6.59% (48), urothelial 3.99% (29), head & neck 2.2% (16). Inclusion criteria: age 20–60 years (37,6 ± 8.3), regardless of gender, first detected cases. Exclusion criteria: moderate and severe concomitant disease, previously treated cancer, relapse of cancer, and patient (or his family) doubts about the use of AI.

They were divided into 4 groups:

  1. D and TP with MTB only - 226 patients (the algorithm is presented in Figure 1)

  2. D and TP with AI only - 206 patients (the algorithm is presented in Figure 2)

  3. D and TP with MTB + AI - 147 patients (the algorithm is presented in Figure 3)

  4. D and TP with DR + AI - 149 patients (the algorithm is presented in Figure 4)

Figure 1.

Flowchart illustrating decision-making: “Patient plus patient's data” leads to “MTB (extract from guidelines, based on own experience and knowledge), close source,” followed by “Decision.” Caption explains MTB analyzes patient information and makes decisions using guidelines and personal expertise.

D and TP decision in 1st group algorithm. It means, that MTB analyze whole patient's information, including assessment of his condition. After discussion, MTB makes decision, which based on own experience and presented data.

Figure 2.

Flowchart showing two inputs, “Patient’s data” and “Extract from guidelines, close source,” both leading to “AI algorithm,” then “LLM model (GPT, Gemini, You, Deep Seek, etc.),” another “AI algorithm,” and final “Decision.” Below, explanatory text states the AI process combines patient data and guidelines via LLMs to make decisions.

D and TP decision in 2nd group algorithm. In this case, Al takes only patient's data, downloaded guidelines and address directly to external LLM model, which gives decision (based on all information and questions).

Figure 3.

Flowchart on a blue background showing two decision paths: one begins with “Patient + patient’s data” leading to “MTB (with proper order)”, while the other starts with “Patient’s data” leading to “AI algorithm (with proper order)”. Both paths reach “Decision”, then “Concordance”. If concordant, the process ends; if not, it leads to “MTB (expert level)” before reaching “Final decision”. A note below clarifies that in non-concordance cases, MTB must be involved again for the final decision.

D and TP decision in 3rd group algorithm. In 3 group we unite paths of 1 and 2 group. Main difference, that in case of non-concordance MTB and Al decision, we must attract MTB again for final decision.

Figure 4.

Flowchart illustrating a decision-making algorithm with the sequence: patient and patient's data, doctor, artificial intelligence with proper order, doctor, and final decision. Additional notes clarify the use of this process for a single doctor in alignment with a prior group, emphasizing that only the doctor makes the decision in this case.

D and TP decision in 4th group algorithm. Algorithm for use by single doctor same to group 2, but here only doctor makes decision.

NCCN guidelines were used for all groups. Patients’ data were ingested from the e-health system, de-identified, and normalized prior to inference. All patients included in the study signed informed consent. Also, the local bioethics committee was permitted to study.

Reports from MTB were inserted in the proper Google Forms.

For effectiveness assessment, we used evaluation criteria: immediate (better understanding of how AI can influence diagnosis) - preparation time (PT), decision time (DT), time to recommendation (TTR), where TTR = PT + DT; midterm (for correct results after MTB, AI, and both decisions) - change diagnosis, change treatment plan (1 month, 6 month).

Questionnaire for DR includes next questions: overall assessment using AI for D, overall assessment using AI for TP, convenience AI, ease of use AI, use AI decreases the time for D and TP decisions, recommendation to colleagues for AI use, and use AI after the end of the study. Answers were (on 3rd and 6th month of study): 0 – very bad or strictly negative, 1 - bad or negative, 2 - moderate or not sure, 3 – good or positive, 4 - very good or strictly positive.

2.2. Statistical analysis

Descriptive analyses for ordinally scaled data (e.g., rankings or numerical scores) used mean, standard deviation (SD), median, and 95% confidence intervals (CI). Changes versus baseline were either expressed in absolute or relative terms for scores as well as proportions. Significance was set at p<0.05 (two-sided).

3. Results

3.1. Technical aspects of our own AI algorithm implementation

Our AI algorithm was created with one important reason – to help decrease errors for MTB and any doctor in their routine work with D and TP for cancer patients. It can be provided by building an AI algorithm that unites all the needed parts.

The hierarchy of our AI includes many components, but strategically we can distinguish three main stages that are performed sequentially.

3.1.1. Introduction to the three-stage AI decision pipeline

The proposed AI algorithm operates within a decision-support workflow designed to approximate the reasoning of a multidisciplinary tumour board while remaining technically tractable, auditable, and modular. Clinicians interact with the system through a web interface by providing three heterogeneous inputs: clinical practice guidelines (as PDF documents), patient-specific data (in structured form), and free-text clinical questions. These inputs differ substantially in structure, reliability, and semantic complexity, and cannot be processed safely as a single undifferentiated stream. To manage this complexity, the AI workflow is explicitly decomposed into three sequential stages.

The first reason for this staging is the separation of concerns. Converting noisy, human-facing artefacts into machine-readable representations is a fundamentally different task from generating recommendations, and both are distinct from validating those recommendations against safety and concordance criteria. By isolating these tasks, each stage can use specialised prompting strategies, external language models, and rule-based logic optimised for its role, while communicating via stable, typed Python classes(e.g., GuidelineRepresentation, PatientClinicalProfile, ParsedRequest). This improves maintainability and allows individual components to be updated or replaced without destabilising the overall system.

A second reason is traceability and governance. The staged design forces the algorithm to expose intermediate representations and decisions instead of collapsing all reasoning into a single opaque model call. This enables explicit inspection of what the system has “understood” about the guideline, the patient, and the clinician’s request before any recommendation is issued. It also creates natural checkpoints for logging, error detection, and performance evaluation at each stage.

Finally, the third reason is safety and calibration. Preliminary decisions produced from structured inputs are not delivered directly to clinicians; they must pass through a dedicated supervisory stage that assesses concordance with guideline logic and patient-specific constraints. This separation between decision formation and decision validation is essential in a clinical research context, where the system must support, rather than replace, expert judgement and where failures must be detectable, explainable, and correctable.

3.1.2. AI algorithm Stage 1: ingestion of inputs and conversion to machine-readable format

In Stage 1, the AI algorithm transforms three heterogeneous input streams into a unified, machine-readable representation that can be consumed by downstream components. The system receives: a guideline document as a PDF file, patient data as JSON-formatted text, and the user request as unstructured natural language. Each of these inputs is routed to a dedicated module—Guideline Cleaner, Patient Data Processor, and Request Parser, respectively—whose primary function is to interface with outsourced large language models (LLMs), enforce structured output, and encapsulate the result in typed Python classes (GuidelineRepresentation, PatientClinicalProfile, and ParsedRequest) (Figure 5).

Figure 5.

Flowchart illustrating a system's Stage 1 AI algorithm for ingesting guidelines, patient data, and user requests, processing each into machine-readable JSON and prompts, then interfacing with large language models including Llama 3.3 70B, GPT 5.1 mini, and GPT-5.1.

Al algorithm stage 1: ingestion of inputs and conversion to machine-readable format.

The Guideline Cleaner module operates on the guideline PDF, which is first converted into raw text using standard document parsing tools. The extracted text is segmented into logical units (sections, paragraphs, tables) that are then formatted into prompts for an external LLM. These prompts instruct the model to identify and extract atomic clinical recommendations, eligibility criteria (e.g. tumour site, TNM stage, performance status constraints), and levels of evidence, and to return the results in a predefined JSON schema. The system enforces this schema by rejecting or reparsing non-conforming outputs and, where necessary, applying post-hoc rule-based corrections. The validated JSON is then deserialised into the GuidelineRepresentation Python class, which captures both structured recommendations and references to the underlying text segments for subsequent use.

The Patient Data Processor ingests patient data supplied as JSON from the user-facing interface. Although structurally formatted, this input typically reflects user-interface and local documentation conventions rather than a canonical oncologic representation. The module, therefore, constructs prompts that ask an outsourced LLM to normalise the data onto a domain-specific schema, including primary diagnosis (tumour site, histology, TNM, and stage group), performance status, comorbidities, prior treatments, and salient laboratory or functional parameters. The LLM returns a JSON object that is constrained to a predefined structure; deterministic rules are then applied to derive secondary attributes such as frailty status, treatment contraindications (e.g., cisplatin ineligibility), and other clinically relevant flags. The final, validated representation is stored in the PatientClinicalProfile class, which serves as the canonical “snapshot” of the patient used throughout subsequent stages.

The Request Parser processes the clinician’s free-text question. Its goal is to convert a natural-language request (e.g., “What is the recommended primary treatment for this T2N1M0 oral cancer in a cachectic patient?”) into a formal query specification. The module constructs prompts that instruct the outsourced LLM to infer the task type (diagnostic, treatment, follow-up), tumour site, stage or stage hypotheses, treatment intent, and explicit constraints or preferences stated in the question. The LLM is required to respond using a strict JSON schema, from which the system instantiates the ParsedRequest class. This object encodes what the system believes is being asked, independent of the particular wording, and includes uncertainty flags if key elements cannot be reliably inferred.

Across all three modules, a common pattern is applied: inputs are decomposed into prompts targeting outsourced LLMs, expected outputs are defined as JSON schemas, and any deviations from these schemas are captured through validation and error-handling routines. The resulting Python classes—GuidelineRepresentation, PatientClinicalProfile, and ParsedRequest—constitute the complete machine-readable state at the end of Stage 1. This staged conversion isolates the noisy, human-facing aspects of the problem in a single phase, ensuring that subsequent stages operate exclusively on well-typed, semantically aligned artefacts rather than raw text or ad hoc data structures.

3.1.3. AI algorithm Stage 2: structured prompt generation and preliminary decision formation

Stage 2 processes the core decision-making step of the AI algorithm by transforming the structured representations from Stage 1 into a preliminary recommendation. At this point, all human-facing heterogeneity has been abstracted into three typed Python classes: GuidelineRepresentation, PatientClinicalProfile, and ParsedRequest. The Final Generator module is responsible for integrating these artefacts into a single structured prompt, submitting this prompt to an outsourced large language model (GPT-5.1 Thinking), and translating the model’s response into an internal decision object (DecisionSupportAnswer) (Figure 6).

Figure 6.

Flowchart diagram illustrating how an AI algorithm processes data through multiple Python classes to generate structured prompts, interacts with a large language model labeled GPT-5.1, and receives output in JSON format for decision support within a healthcare system context.

Al algorithm stage 2: structured prompt generation and preliminary decision formation.

The process begins with context assembly. The Final Generator receives instances of GuidelineRepresentation, PatientClinicalProfile, and ParsedRequest and performs a selective transformation into a compact prompt payload. From the guideline representation, it extracts only those recommendations and text segments whose eligibility criteria match the patient’s stage, tumour site, and key constraints as encoded in the patient profile. From the patient profile, it constructs a succinct summary of the diagnosis (site, TNM, stage group), performance status, comorbidities, prior treatments, and derived flags (e.g., frailty, treatment contraindications). From the parsed request, it retrieves the task type, treatment intent, and any explicitly stated constraints or user preferences.

These elements are then assembled into a highly structured system prompt that specifies to GPT-5.1 Thinking both the decision problem and the output contract. The prompt includes: a machine-readable block describing the patient; a block summarising relevant guideline recommendations and evidence annotations; and a block describing the clinical question and required outputs (e.g., primary recommendation, alternatives, non-recommended options, and rationale). Critically, the prompt also embeds a formal JSON schema that defines the expected structure of the model’s response, including fields for the primary recommendation, alternative options, non-recommended strategies with reasons, guideline alignment references, and uncertainties or open issues.

GPT-5.1 Thinking is then invoked as an outsourced reasoning engine, with the Final Generator acting as a strict contract enforcer. The model’s output is accepted only if it conforms to the predefined JSON schema; non-conforming outputs trigger automatic re-prompting or rejection, thereby reducing the risk of malformed responses. Once a valid JSON object is obtained, it is deserialised into the DecisionSupportAnswer Python class. This class encapsulates the preliminary decision, including: a structured description of the primary recommended strategy and its intent; a set of alternative strategies with indications, advantages, and disadvantages; a list of options that are explicitly not recommended for this patient and the reasons for exclusion; and a guideline alignment section that links each proposed option to specific recommendations contained in GuidelineRepresentation.

Stage 2 is deliberately limited to decision formation, not validation. While the Final Generator uses guideline context and patient constraints as conditioning information, it does not independently enforce safety thresholds or concordance metrics. Instead, its role is to produce the best possible candidate answer under full visibility of the structured inputs and to expose its reasoning in a transparent, machine-readable format. The resulting DecisionSupportAnswer object, together with the original guideline and patient representations, is then passed to Stage 3, where a dedicated supervisory process evaluates the consistency, safety, and guideline concordance of the preliminary decision before any output is returned to the user.

3.1.4. AI algorithm Stage 3: supervisor module validation and final decision delivery

Stage 3 implements a structured supervisory layer that evaluates the preliminary recommendation before it is exposed to the end user. At this point, the AI algorithm has already produced a candidate decision encoded in the DecisionSupportAnswer class and retains full access to the structured representations of guideline knowledge (GuidelineRepresentation) and patient status (PatientClinicalProfile). The Supervisor Module uses these three artefacts to perform an independent consistency and concordance assessment, leveraging an outsourced LLM (GPT-5.1 Thinking) as a critic rather than as a primary decision-maker (Figure 7).

Figure 7.

Flowchart diagram illustrating an AI algorithm system that receives input from a web interface, uses a supervisor module to combine outputs from three Python classes, validates decisions based on concordance level, and outsources additional reasoning to an LLM via prompt and JSON exchange.

Al algorithm stage 3: supervisor module validation and final decision delivery.

The validation process begins with the construction of a supervisory prompt that combines: a compact representation of the patient profile, including diagnosis, stage, performance status, comorbidities, and derived flags; the subset of guideline recommendations relevant to this case, as encoded in GuidelineRepresentation; and the full content of the DecisionSupportAnswer, including primary recommendation, alternatives, non-recommended options, and guideline alignment claims. The prompt explicitly instructs GPT-5.1 Thinking to assess, in a structured manner, whether the proposed decisions are consistent with guideline logic, compatible with patient-specific constraints, and internally coherent.

The LLM is required to output a JSON object conforming to a predefined schema that includes: (a) a concordance score between 0 and 100%, reflecting the degree of alignment between the decision and the guideline–patient combination; (b) a list of identified issues, categorised by severity (e.g. minor omission vs. major contraindication violation); and (c) an optional set of suggested modifications to the decision, including justifications and references to specific guideline elements. The Supervisor Module validates the returned JSON against this schema and, if necessary, re-prompts until a structurally valid response is obtained.

Once a valid supervisory response is available, the algorithm applies a decision threshold. If the concordance score is at or above a predefined value of 80%, the Supervisor Module integrates any suggested minor corrections into the original DecisionSupportAnswer and instantiates a ReviewedDecision class that represents the final, vetted recommendation. This object includes the refined decision, the final concordance score, a summary of modifications relative to the preliminary answer, and a structured list of residual uncertainties that should be highlighted to the clinician.

If the concordance score falls below the threshold, or if the supervisory output flags critical safety issues (e.g., recommendation of a modality that conflicts with a major contraindication encoded in the PatientClinicalProfile), the algorithm does not deliver a recommendation. Instead, it generates an error status that is returned upstream, together with a concise explanation of the detected issues and, where appropriate, suggestions for revising the input (e.g., clarifying missing staging information or correcting inconsistent patient data). In this failure mode, the user is required to revise the case and re-run the algorithm, maintaining a human-in-the-loop model of oversight.

All supervisory interactions, including the concordance score, issue list, and any applied modifications, are logged alongside the underlying inputs and the preliminary decision. This architecture ensures that Stage 3 functions as a transparent, auditable control layer that enforces guideline adherence and patient-specific safety constraints, while still enabling the use of powerful but fallible LLMs as both decision engines and structured critics within a controlled pipeline.

3.1.5. Summary of the staged AI decision pipeline

The proposed three-stage AI pipeline formalises guideline-based decision support as a sequence of well-defined transformations from heterogeneous clinical inputs to a vetted, machine-generated recommendation. In Stage 1, guideline documents, patient data, and clinician queries are ingested and converted into harmonised, domain-specific representations (GuidelineRepresentation, PatientClinicalProfile, ParsedRequest). This isolates the inherently noisy interface between human and machine and ensures that all subsequent reasoning operates on semantically aligned artefacts rather than raw text or ad hoc data structures.

Stage 2 leverages these representations to generate a preliminary recommendation via a dedicated Final Generator module. By conditioning GPT-5.1 Thinking on a structured description of the patient, the applicable guideline content, and the formalised clinical question, and by constraining outputs to the DecisionSupportAnswer schema, the system produces decisions that are both context-aware and amenable to downstream analysis.

Stage 3 then introduces an explicit supervisory layer that re-evaluates the preliminary decision against the original guideline and patient profile, using GPT-5.1 Thinking as a structured critic. The Supervisor Module quantifies guideline concordance, identifies safety or consistency issues, and either promotes the decision to a ReviewedDecision or blocks it when predefined thresholds are not met.

Collectively, this staged design combines the expressive power of large language models with principled abstractions, explicit schemas, and threshold-based control. It creates a technically transparent and auditable framework within which AI-generated recommendations can be prospectively evaluated against multidisciplinary tumour board decisions, while preserving a clear role for human oversight and clinical responsibility.

3.2. Clinical aspects of our own AI algorithm implementation

There are several key aspects in the work of the MTB that affect the quality of its work - these are urgent and midterm criteria, which served as the subject of our research.

Assessment begun from immediate criteria, namely time, spending for 1 patient in different groups (Table 1; Figure 8).

Table 1.

Immediate criteria (M ± SD), min.

Group PT DT TTR
D and TP with MTB only, n=226 21.3 ± 3.6 15.6 ± 6.2 36.9 ± 7.8
D and TP with AI only, n=206 19.7 ± 4.1 2.1 ± 0.02* 22.4 ± 3.3*
D and TP with MTB + AI, n=147 26.9 ± 4.2*,** 4.9 ± 0.9*,** 29.2 ± 3.1*,**

* statistically significant results in comparison to group - D and TP with MTB only.

** statistically significant results in comparison to group - D and TP with AI only.

Figure 8.

Bar chart comparing immediate criteria in minutes for PT, DT, and TTR across three methods: D and TP with MTB only, D and TP with AI only, and D and TP with MTB + AI. D and TP with AI only consistently results in the shortest times, especially for DT, while MTB + AI tends to take the longest.

Immediate criteria in different groups.

The higher PT results in group 3 indicate that more time is needed to load patient data and interpret it accordingly. At the same time, the DT results in group 3 are the lowest, and 7 times less than in group 1.

Group 3 (MTB + AI) has higher scores TTR than in group 2, but still significantly lower than group 1. It means, that adding AI to MTB not only does not extend the TTR, but also reduces it on 21%. The implementation of AI does not prolong the operation of MTB, which allows TTR to be accepted for a larger number of patients per unit of time.

The next stage of our research was to study the impact of AI implementation on midterm criteria, namely whether the implementation of AI in MTB work affected the quality of D and TP, and possible changes in different terms.

Table 2 presents the results of D change in the terms of patient observation 1 and 6 months after MTB.

Table 2.

Midterm criteria (change D).

Group 1 month 6 months
D and TP with MTB only, n=226 n1 = 221
n1x=15
6.79%
n6 = 203
n6x=37
18.22%
D and TP with AI only, n=206 n1 = 205
n1x=12
5.36%
n6 = 197
n6x=39
19.8%
D and TP with MTB + AI, n=147 n1 = 144
n1x=4
2.78%*,**
n6 = 129
n6x=16
12.4%*,**

n – number of patients studied on MTB.

n1= overall number of patients in 1 month after MTB.

n1x= number of patients with change D in 1 month after MTB.

n6= overall number of patients in 6 months after MTB.

n6x= number of patients with change D in 6 months after MTB.

* statistically significant results in comparison to group - D and TP with MTB only.

** statistically significant results in comparison to group - D and TP with AI only.

It should be taken into account that a change in D does not always mean a change in TP, and vice versa, when leaving the previous D, TP should be changed. Therefore, we investigated the role of AI in changing the D and TP (Tables 3, 4; Figure 9).

Table 3.

Midterm criteria (change TP).

Group 1 month 6 months
D and TP with MTB only, n=226 n1 = 221
n1x=25
11.31%
n6 = 203
n6x=51
25.12%
D and TP with AI only, n=206 n1 = 205
n1x=18
8.78%
n6 = 197
n6x=45
22.84%
D and TP with MTB + AI, n=147 n1 = 144
n1x=6
4.17%*,**
n6 = 129
n6x=20
15.5%*,**

n – number of patients studied on MTB.

n1= overall number of patients in 1 month after MTB.

n1x= number of patients with change TP in 1 month after MTB.

n6= overall number of patients in 6 months after MTB.

n6x= number of patients with change TP in 6th months after MTB.

* statistically significant results in comparison to group - D and TP with MTB only.

** statistically significant results in comparison to group - D and TP with AI only.

Table 4.

Midterm criteria (change D + TP).

Group 1 month 6 months
D and TP with MTB only, n=226 n1 = 221
n1x=29
13.12%
n6 = 203
n6x=57
28.08%
D and TP with AI only, n=206 n1 = 205
n1x=18
8.78%*
n6 = 197
n6x=48
24.3%
D and TP with MTB + AI, n=147 n1 = 144
n1x=7
4.86%*,**
n6 = 129
n6x=19
14.73%*,**

n – number of patients studied on MTB,.

n1= overall number of patients in 1 month after MTB.

n1x= number of patients with change TP in 1 month after MTB.

n6= overall number of patients in 6 months after MTB.

n6x= number of patients with change TP in 6 months after MTB.

* statistically significant results in comparison to group - D and TP with MTB only.

** statistically significant results in comparison to group - D and TP with AI only.

Figure 9.

3D bar chart illustrating midterm criteria, showing percentage change in D and TP for three groups: MTB only, AI only, and MTB plus AI. For each group, two bars represent increases at one month and six months. All groups show higher percentages at six months compared to one month, with AI only and MTB only reaching above 30 percent at six months and the combined group just below 20 percent.

Midterm criteria in different groups.

The generalized assessment shows that the introduction of AI into the work of the MTB significantly reduces the percentage of change in D and TP. At 1 month of the study, the difference between group 3 and 1 was 2.7 times, and at 6 months 1.9 times. The use of AI alone shows better results with TP, but this difference is not statistically significant.

The combination of the MTB with AI gives the most effective result in reducing the change in D and TP from 28.08% to 14.73% within 6 months (13.12% vs 4.86% in 1 month).

A number of excluded patients from study explained by next reasons: refusal to participate in the study, occurrence of exclusion criteria, change of residence.

The next step was to study the effectiveness of using AI in the work of DR (group 4, Table 5). The comparison group was taken as indicators of D and TP by doctors before the implementation of AI.

Table 5.

Technical aspects of AI use for DR.

Criteria <15 min 15–30 min 30–60 min >60 min
time for TP decision, n = 37 – – 29
(78.4%)
8
(21.6%)
time for TP decision with AI, n = 37 – 6
(16.2%)
28
(75.7%)
3
(8.1%)
3 month, n3 = 35
time for TP decision with AI 1
(2.9%)
13
(37.1%)
21
(60%)
–
6 month, n6 = 34
time for TP decision with AI 3
(8.8%)
20
(58.9%)
11
(32.3%)
–

n – number of DR on study start.

n3 = number of DR in 3 months after D and TP decision.

n6= number of DR in 6 months after D and TP decision.

Before the introduction of AI, the time to make a decision on the D and TP for all 100% of doctors was more than 30 minutes. At the beginning of the study, 16.2% of doctors managed to establish a D and prescribe a TP in less than 30 minutes, and the percentage of doctors who needed more than 60 minutes to do this decreased by 2.7 times.

In the period from 3 to 6 months, the decision-making time of doctors does not exceed 60 minutes in any case. If at 3 months 40% of doctors have a decision-making time of no more than 30 minutes, then at 6 months this percentage is 67.7%. This indicates the accumulation of experience and the facilitation of work with AI.

The next important question that interested us during the study was the perception of AI by doctors and their impressions (Table 6). The survey was conducted at 3 and 6 months, without taking into account their opinions at the start of the study, given the complexity of the initial stage, developing skills for working with AI.

Table 6.

Assessment AI algorithm by DR, n = 37.

Criteria 0 1 2 3 4
3 month, n3 = 35
overall assessment using AI for D – 2
(5.7%)
6
(17.1%)
17
(48.6%)
10
(28.6%)
overall assessment using AI for TP – 1
(2.9%)
5
(14.3%)
23
(65.7%)
6
(17.1%)
convenience AI 1
(2.9%)
1
(2.9%)
14
(40%)
15
(42.8%)
4
(11.4%)
ease of use of AI – 2
(5.7%)
9
(25.7%)
19
(54.3%)
5
(14.3%)
use AI decrease time for D and TP decision 2
(5.7%)
4
(11.4%)
7
(20%)
16
(45.8%)
6
(17.1%)
can you recommend AI for other colleagues? – – 7
(20%)
24
(68.6%)
4
(11.4%)
use AI after end of study 2
(5.7%)
3
(8.6%)
11
(31.4%)
14
(40%)
5
(14.3%)
6 month, n6 = 34
overall assessment using AI for D – – 5
(14.7%)
13
(38.2%)
16
(47.1%)
overall assessment using AI for TP – – 1
(2.9%)
20
(58.9%)
13
(38.2%)
convenience AI – – 2
(5.9%)
28
(82.3%)
4
(11.8%)
ease of use of AI – – 1
(2.9%)
14
(41.2%)
19
(55.9%)
use AI decrease time for D and TP decision 1
(2.9%)
4
(11.8%)
9
(26.5%)
14
(41.2%)
6
(17.6%)
can you recommend AI for other colleagues? – – 3
(8.8%)
12
(35.3%)
19
(55.9%)
use AI after end of study – 2
(5.9%)
10
(29.4%)
17
(50%)
5
(14.7%)

n – number of DR on study start.

n3 = number of DR in 3 months after D and TP decision.

n6= number of DR in 6 months after D and TP decision.

0 – very bad or strictly disagree.

1 - bad or disagree.

2 - moderate or not sure.

3 – good or agree.

4 - very good or strictly agree.

When evaluating the use of AI for D at 3 and 6 months, doctors answered well and very well in 77.2% and 85.3% of cases, respectively. It is more important for the doctor to state the stage of the disease in this case.

Another important factor for the doctor is determining the correct TP. We see that the percentage of good and very good ratings increased from 82.8% at month 3 to 97.1% at month 6.

Convenience and ease of use seem like similar terms, but for doctors, the former indicates more of a psychological and mental perception, while the latter reflects the technical side of use. Convenience with a rating of good or very good increases from 54.2% to 94.1%, and ease of use from 68.6% to 97.1%, indicating the importance of ease and simplicity in using AI.

When asked whether there is a reduction in the time to make a decision on D and TP, we see that agree and strictly agree at month 3 are 62.9% and 58.8% at month 6. This may indicate mastery of AI skills, in contrast to the previous table (Table 5), where doctors note an overall reduction in the time to make a decision on D and TP.

The ability to recommend AI to other colleagues is evidence of understanding the importance of its use for D and TP in cancer patients. 80% and 91.2% agree and strictly agree with the absence of any negative response may also indicate the literacy of doctors in terms of AI.

However, the percentage of 54.3% and 64.7% of doctors who want to continue using AI indicates a certain conservatism, since AI is not mandatory for their clinical practice, which involves other diagnosis, complex clinical cases, etc. This indicates the importance of explanatory work, promotion of AI, and involvement in congresses and conferences dedicated to AI and oncology.

4. Discussion

4.1. MTB – pros and cons

Treatment of cancer patients in comparison with the vast majority of other somatic diseases has significant differences. This is the presence of surgical, therapeutic, immunological, hormonal, targeted, radiation and other treatment methods, which are used both in mono-mode and sequentially or in parallel. In addition, the presence of treatment in the neoadjuvant or adjuvant mode significantly complicates the treatment itself (26).

Therefore, in most specialised medical clinics, MTB has been introduced into the work, the main principle of which is based on the consolidated decision of specialists in various fields involved in the treatment of cancer patients. This has significantly improved the results of setting D and determining TP in patients of various localizations (27, 28).

However, despite the widespread implementation of MTB, today the results of the latter’s activities leave much to be desired. According to the literature, the percentage of change in D and TP remains high (up to 67%). This applies not only to differences between continents, countries, but also to centres within one country (29).

Our results indicate that the percentage of change in D and TP after MTB increases over time and at 6 months is 28.08%, although at 1 month it was 13.12%. In our opinion, the percentage of change in D and TP is high and requires close attention in order to develop measures to reduce it (30–32).

According to the researchers, the main reasons for this are – incomplete or incorrect data in the medical documentation in preparation for the MTB meeting (lack of images, histology, CT/PET data, etc.); different interpretation of images/pathology between specialists; differences in local protocols or resources (regional restrictions in access to therapies/technologies); low quality of evidence for certain clinical situations leads to variability in recommendations; patient preferences or clinical management steps after the MTB meeting that change the implementation of recommendations. Thus, there is a dire need to improve the work of MTB (29, 31).

In our work, we did not set ourselves the task of analysing the reasons leading to the change in D and TP, since our attention is focused only on the efficiency of MTB work.

4.2. AI and oncology: ethical, legal, moral, technical, validation, limitation and risk issues

The ethical and legal legitimacy of artificial intelligence in oncology hinges on whether its deployment improves patient-centred outcomes while preserving accountability, transparency, and equity. In this prospective multicentre comparison of multidisciplinary tumour board recommendations against a guideline-based artificial intelligence approach, the central ethical tension is not whether machines should replace clinicians, but whether the decision pathway can be made more reliable, auditable, and consistently aligned with current standards of care without introducing new forms of harm. Contemporary guidance emphasizes that real-world deployment must be governed as a socio-technical intervention: stakeholders define the clinical objective and acceptable trade-offs, the system is evaluated in context, and accountability for decisions remains explicit throughout the lifecycle. A guideline-based system can strengthen moral accountability because it can expose the rule base, the evidence provenance, and the rationale chain linking patient features to recommendations, thereby enabling contestability and structured review. However, this only delivers ethical value if the guideline representation is demonstrably up to date, context-aware (including local access constraints), and paired with human oversight that actively mitigates automation bias (20, 33–46).

From a regulatory perspective, oncology-facing clinical software is increasingly scrutinized for pre-market evidence quality and post-deployment safety monitoring. Analyses of approvals highlight that many medical artificial intelligence products reach practice with limited prospective evaluation and variable external validation, raising concerns about hidden vulnerabilities after implementation (36, 41). This is directly relevant to our study design: prospective, multicentre evaluation is an essential risk-control mechanism, because it surfaces heterogeneity in case-mix, documentation completeness, imaging and pathology interpretation, and resource availability that can systematically shift performance. Dataset shift is not an edge case in oncology; it is a predictable operational reality driven by evolving protocols, changing diagnostics, and population differences, and it can degrade model performance unless monitoring and recalibration are built into governance (42). Therefore, the organizational “operating model” should include performance surveillance, drift detection, and a clear escalation route when discordance emerges between tumour board decisions and guideline-based outputs.

Validation standards must also be aligned with contemporary reporting and evaluation frameworks. For interventional systems, protocol and trial reporting extensions specify how the artificial intelligence component, human-system interaction, error modes, and clinical workflow integration should be documented to enable reproducibility and appraisal. For prediction or risk models often used to triage or prioritise oncology pathways, reporting remains poor across the literature, and high risk of bias is common, undermining clinical readiness. Early-stage decision support evaluations similarly require transparent specification of intended use, clinical context, and human factors, because safety risks frequently arise from interaction failures rather than pure algorithmic error (35). In the present work, a key technical-ethical advantage of guideline-based artificial intelligence is interpretability by construction: it can reduce opaque statistical inference and instead foreground traceable mapping to guideline logic. Nonetheless, it remains vulnerable to mis-specification (for example, incomplete patient data, ambiguous staging descriptors, or discordant pathology inputs) and to guideline limitations, including evidence gaps and slow update cycles. These limitations reintroduce moral questions about fairness and beneficence: consistent decisions are not ethically preferable if they are consistently wrong for under-represented subgroups, complex multimorbidity, or atypical tumour biology (20, 33, 34, 37, 38).

Clinician trust and patient legitimacy are also decisive. Oncologists report that explainability, consent, responsibility allocation, and medico-legal liability are major barriers to adoption, particularly when outputs influence high-stakes treatment selection. This implies that implementation should not treat artificial intelligence as a standalone product, but as a governance-managed capability with explicit communication norms: what the system can and cannot infer, when escalation to tumour board is mandatory, and how disagreements are resolved and documented. Recent oncology reviews converge on the view that the main near-term risk is not speculative “superhuman” behaviour but mundane failure modes: data quality defects, biased training sources, inadequate external validation, workflow mismatch, and uncritical reliance on generated recommendations. In this context, the comparative framing of our study is strategically valuable: discordance between tumour board and guideline-based outputs becomes an auditable signal to priorities root-cause analysis (missing inputs, interpretive disagreement, local resource constraints, or guideline ambiguity) and to design targeted remediation, rather than attributing variability to individual clinician preference (39, 43, 44).

Finally, a forward-looking risk posture requires enterprise-level governance. Practical models proposed within comprehensive cancer centres emphasize structured intake, risk stratification, lifecycle monitoring, and multidisciplinary oversight spanning legal, ethics, adoption, and performance domains (46). Policy-oriented guidance similarly highlights that regulation and institutional controls will continue to evolve, and oncology programmers should plan for adaptive compliance, documentation traceability, and continuous quality assurance rather than one-time “go-live” certification (45). Taken together, the ethical, legal, technical, and validation agenda for our guideline-based artificial intelligence approach is clear: maximize transparency and reproducibility, operationalize prospective multicentre evaluation, embed monitoring for dataset shift and performance drift, and maintain accountable human governance so that clinical responsibility is preserved even as decision processes become more standardized and scalable (20, 33–47).

4.3. AI and MTB are both better ways for a clinical decision support system. Is not?

Thus, having considered the role of AI and MTB separately in decision-making regarding D and TP, we see that they have both positive and negative sides. Many researchers have made attempts to combine AI and MTB, but the algorithm and results remain unsatisfactory.

Our approach differs from others in that we have created a decision-making algorithm regarding D and TP, which introduces parallel paths of AI and MTB to the “Concordance” level, after which a decision is made if there is a match, and if there is a mismatch, MTB is re-evaluated.

The implementation of this approach allowed us to reduce TTR by 20.9%, which makes it possible to process a larger number per unit of time and reduce the financial burden.

If the changes in D and TP when making a decision only to use MTB were 13.12% (1 month), 28.08% (6 months); then in the group of patients with a combination of MTB and AI, 4.86% and 14.73%, respectively. The latter indicates that when combining MTB and AI, we managed to obtain good results. This was achieved by reducing the negative aspects of MTB and AI and enhancing the positive ones.

Of course, in our study we included patients with newly diagnosed cancer, without significant comorbidities, and without cancer recurrence. Therefore, our results may differ from other studies of this type.

As for the use of AI by DR who make decisions on D and TP alone, the situation is ambiguous. The results of treating patients with oncological pathology by DR are significantly lower, because they work in non-specialised medical hospitals, they encounter such patients less often, it is more difficult for them to follow guidelines, their theoretical and practical skills need to be improved, etc.

We proposed a DR tool that should significantly help in working with cancer patients, especially for D and TP. Before starting the study, we had doubts, given the certain conservatism of DR in using various models, algorithms, and AI tools, especially in oncology.

If before the implementation of AI in DR work, the time to make a decision was more than 30 minutes in 100% of cases, then after 6 months the time was less than 30 minutes in 67.7% of cases.

After 6 months of using AI, characteristics such as - overall assessment using AI for TP, convenience AI, ease of use of AI, recommendation AI for other colleagues reach more than 90%.

This demonstrates interest and understanding of the use of AI in everyday work with cancer patients. The psychological wariness that exists at the beginning of AI research is changing to pragmatism when DRs see concrete results and effectiveness.

Unlike MTB, a DR does not have colleagues with whom he can exchange ideas. Therefore, evaluating the solutions provided by AI is significantly more difficult and requires the development of appropriate skills.

The literature presents a number of serious problems when working with DR and AI, which cannot always be solved.

4.4. Conclusions

Thus, the analysis of the results obtained in the article indicates their certain multidirectionality.

The AI model we proposed, which differs from existing ones in the use of official guidelines, works as an intermediary between patients, guidelines, and LLM, determining the threshold of correct decisions and a high degree of patient de-identification.

Use of MTB and AI combination to optimize TTR, decrease the percentage of change D and TP for 1,6 months, decrease the negative sides of separate MTB and AI. This highlights the positive results both in the early and long term. The study in four oncology centres allowed to expand the geography of patients and reach a larger number of doctors.

Of course, it’s the first step of our work, when we can use this combination for the first detected cases at a basic level. Our patients were without comorbidities. At the same time, a significant part of oncological patients has comorbidities and multiple cancers. In this case, given the increase in the number of tasks, it is necessary to develop an AI agent based on other logistical approaches using orchestration.

Our attempt to help doctors who do not work in specialised institutions was successful. AI cannot replace MTB, but it empowers non-oncological clinicians and DR to generate evidence-based, auditable recommendations in settings with limited specialist capacity. It turned out to be interesting that a certain number of doctors reacted rather “coolly” to the introduction of AI into their work, namely, the lack of desire to use AI after the study was completed, and to recommend the use of AI to other colleagues. This indicates the need for active educational work among all doctors before introducing AI.

The introduction of AI into the work of MTB and other doctors can reduce the financial burden on the healthcare system by establishing the correct diagnosis and choosing the optimal treatment tactics. However, this requires separate research in the future.

5. Future trajectory

To reach expert-level autonomy, further work is required a local guideline adaptation, multimodal data ingestion and longitudinal validation across tumour types.

Assessment of the financial efficiency of implementing AI in the work of the MTB.

Funding Statement

The author(s) declared that financial support was not received for this work and/or its publication.

Footnotes

Edited by: Pradakshina Sharma, The University of Tennessee, United States

Reviewed by: Ovidiu Pop, University of Oradea, Romania

Maaruf Ali, Metropolitan University of Tirana, Albania

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The studies involving humans were approved by Ethics committee of the Bukovinian State Medical University. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

OI: Writing – original draft. SH: Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1. Specchia ML, Frisicale EM, Carini E, Di Pilla A, Cappa D, Barbara A, et al. The impact of tumor board on cancer care: Evidence from an umbrella review. BMC Health Serv Res. (2020) 20:73. doi:  10.1186/s12913-020-4930-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Walraven JEW, van der Hel OL, van der Hoeven JJM, Lemmens VEPP, Verhoeven RHA, Desar IME. Factors influencing the quality and functioning of oncological multidisciplinary team meetings: Results of a systematic review. BMC Health Serv Res. (2022) 22:829. doi:  10.1186/s12913-022-08112-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Taberna M, Gil Moncayo F, Jané-Salas E, Antonio M, Arribas L, Vilajosana E, et al. The multidisciplinary team (MDT) approach and quality of care. Front Oncol. (2020) 10:85. doi:  10.3389/fonc.2020.00085 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Walraven JEW, Verhoeven RHA, van der Meulen R, van der Hoeven JJM, Lemmens VEPP, Hesselink G, et al. Facilitators and barriers to conducting an efficient, competent and high-quality oncological multidisciplinary team meeting. BMJ Open Qual. (2023) 12:e002130. doi:  10.1136/bmjoq-2022-002130 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Law NLW, Hong LW, Tan SSN, Foo CJ, Lee D, Voon PJ. Barriers and challenges of multidisciplinary teams in oncology management: A scoping review protocol. BMJ Open. (2024) 14:e079559. doi:  10.1136/bmjopen-2023-079559 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Rasmussen TR, Gouliaev A, Jakobsen E, Hjorthaug K, Larsen LU, Meldgaard P, et al. Impact of multidisciplinary team discrepancies on comparative lung cancer outcome analyses and treatment equality. BMC Cancer. (2024) 24:1423. doi:  10.1186/s12885-024-13188-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Tirotta F, Hodson J, Alcorn D, Al-Mukhtar A, Ayre G, Barlow A, et al. Assessment of inter-centre agreement across multidisciplinary team meetings for patients with retroperitoneal sarcoma. Br J Surg. (2023) 110:1189–96. doi:  10.1093/bjs/znad157 [DOI] [PubMed] [Google Scholar]
  • 8. Kočo L, Weekenstroo HHA, Lambregts DMJ, Sedelaar JPM, Prokop M, Fütterer JJ, et al. The effects of multidisciplinary team meetings on clinical practice for colorectal, lung, prostate and breast cancer: A systematic review. Cancers. (2021) 13:4159. doi:  10.3390/cancers13164159 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Khalafallah AM, Jimenez AE, Romo CG, Kamson DO, Kleinberg L, Weingart J, et al. Quantifying the utility of a multidisciplinary neuro-oncology tumor board. J Neurosurg. (2021) 135:87–92. doi:  10.3171/2020.5.JNS201299 [DOI] [PubMed] [Google Scholar]
  • 10. Bortot L, Targato G, Noto C, Giavarra M, Palmero L, Zara D, et al. Multidisciplinary team meeting proposal and final therapeutic choice in early breast cancer: Is there an agreement? Front Oncol. (2022) 12:885992. doi:  10.3389/fonc.2022.885992 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Qureshi S, Abbasi WA, Jalil HA, Ahmed R, Iqbal M, Saiyed H, et al. Evaluating treatment plan modifications from surgeons’ initial recommendations to multidisciplinary tumor board consensus for cancer care in a resource-limited setting. Curr Oncol. (2025) 32:310. doi:  10.3390/curroncol32060310 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Freytag M, Herrlinger U, Hauser S, Bauernfeind FG, Gonzalez-Carmona MA, Landsberg J, et al. Higher number of multidisciplinary tumor board meetings per case leads to improved clinical outcome. BMC Cancer. (2020) 20:355. doi:  10.1186/s12885-020-06809-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Riaz IB, Khan MA, Osterman TJ. Artificial intelligence across the cancer care continuum. Cancer. (2025) 131:e70050. doi:  10.1002/cncr.70050 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Liu Y, Yu W, Dillon T. Regulatory responses and approval status of artificial intelligence medical devices with a focus on China. NPJ Digital Med. (2024) 7:Article 255. doi:  10.1038/s41746-024-01254-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Liñares-Blanco J, Pazos A, Fernandez-Lozano C. Machine learning analysis of TCGA cancer data. PeerJ Comput Sci. (2021) 7:e584. doi:  10.7717/peerj-cs.584 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Azenkot T, Rivera DR, Stewart MD, Patel SP. Artificial intelligence and machine learning innovations to improve design and representativeness in oncology clinical trials. Am Soc Clin Oncol Educ Book. (2025) 45:e473590. doi:  10.1200/EDBK-25-473590 [DOI] [PubMed] [Google Scholar]
  • 17. Corti C, Cobanaj M, Dee EC, Criscitiello C, Tolaney SM, Celi LA, et al. Artificial intelligence in cancer research and precision medicine: Applications, limitations and priorities to drive transformation in the delivery of equitable and unbiased care. Cancer Treat Rev. (2023) 112:102498. doi:  10.1016/j.ctrv.2022.102498 [DOI] [PubMed] [Google Scholar]
  • 18. Sheller MJ, Edwards B, Reina GA, Martin J, Pati S, Kotrotsou A, et al. Federated learning in medicine: Facilitating multi-institutional collaborations without sharing patient data. Sci Rep. (2020) 10:12598. doi:  10.1038/s41598-020-69250-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Qian Z, Callender T, Cebere B, Janes SM, Navani N, van der Schaar M. Synthetic data for privacy-preserving clinical risk prediction. Sci Rep. (2024) 14:25676. doi:  10.1038/s41598-024-72894-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. (2024) 385:e078378. doi:  10.1136/bmj-2023-078378 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Wang J-W, Meng M, Dai M-W, Liang P, Hou J. Correlation does not equal causation: The imperative of causal inference in machine learning models for immunotherapy. Front Immunol. (2025) 16:1630781. doi:  10.3389/fimmu.2025.1630781 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Singh V, Cheng S, Kwan AC, Ebinger J. United States Food and Drug Administration regulation of clinical software in the era of artificial intelligence and machine learning. Mayo Clinic Proceedings: Digital Health. (2025) 3:100231. doi:  10.1016/j.mcpdig.2025.100231 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Rosenbacke R, Melhus Å, McKee M, Stuckler D. How explainable artificial intelligence can increase or decrease clinicians’ trust in AI applications in health care: Systematic review. JMIR AI. (2024) 3:e53207. doi:  10.2196/53207 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Peter S, Riemer K, West JD. The benefits and dangers of anthropomorphic conversational agents. PNAS. (2025) 122:e2415898122. doi:  10.1073/pnas.2415898122 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Ivashchuk O, Hovornyan S. 397P Anthropomorphism vs. technically correct terminology in AI oncological literature: Toward a mature scientific approach. ESMO Real World Data Digital Oncol. (2025) 10:100593. doi:  10.1016/j.esmorw.2025.100593 38826717 [DOI] [Google Scholar]
  • 26. Loibl S, André F, Bachelot T, Barrios CH, Bergh J, Burstein HJ, et al. Early breast cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol. (2024) 35:159–82. doi:  10.1016/j.annonc.2023.11.016 [DOI] [PubMed] [Google Scholar]
  • 27. de Castro G, Jr., Souza FH, Lima J, Bernardi LP, Teixeira CHA, Prado GF. Does multidisciplinary team management improve clinical outcomes in NSCLC? A systematic review with meta-analysis. JTO Clin Res Rep. (2023) 4:100580. doi:  10.1016/j.jtocrr.2023.100580 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Di Pilla A, Cozzolino MR, Mannocci A, Carini E, Spina F, Castrini F, et al. The impact of tumor boards on breast cancer care: Evidence from a systematic literature review and meta-analysis. Int J Environ Res Public Health. (2022) 19:14990. doi:  10.3390/ijerph192214990 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. He C. Multidisciplinary team meetings: Barriers to implementation in cancer care. Oncol (Williston Park). (2024) 38:339–44. doi:  10.46883/2024.25921026 [DOI] [PubMed] [Google Scholar]
  • 30. Sassé B, Shaya S, Nimmo J, Cao K, Day D, Evans K, et al. Evaluating the impact of a tertiary multidisciplinary meeting in metastatic breast cancer: A prospective study. Breast. (2024) 79:103861. doi:  10.1016/j.breast.2024.103861 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Ichikawa M, Uematsu K, Yano N, Yamada M, Ono T, Kawashiro S, et al. Implementation rate and effects of multidisciplinary team meetings on decision making about radiotherapy: An observational study at a single Japanese institution. BMC Med Inf Decis Making. (2022) 22:Article 111. doi:  10.1186/s12911-022-01849-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Ayub B, Qureshi FA, Hassan NH, Shaukat F, Qureshi TA. Optimising head and neck cancer patient management: The crucial contributions of multidisciplinary tumour board decision-making. ecancermedicalscience. (2024) 18:1710. doi:  10.3332/ecancer.2024.1710 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Hendrickx J-J, Mennega T, Uppelschoten JM, Leemans CR. Changes in multidisciplinary team decisions in a high volume head and neck oncological center following those made in its preferred partner. Front Oncol. (2023) 13:1205224. doi:  10.3389/fonc.2023.1205224 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Rivera SC, Liu X, Chan A-W, Denniston AK, Calvert MJ, The SPIRIT-AI and CONSORT-AI Working Group . Guidelines for clinical trial protocols for interventions involving artificial intelligence: The SPIRIT-AI extension. BMJ. (2020) 370:m3210. doi:  10.1136/bmj.m3210 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. DECIDE-AI Expert Group . Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. (2022) 28:924–33. doi:  10.1038/s41591-022-01772-9 [DOI] [PubMed] [Google Scholar]
  • 36. Wu E, Wu K, Daneshjou R, Ouyang D, Ho DE, Zou J. How medical AI devices are evaluated: Limitations and recommendations from an analysis of Food and Drug Administration approvals. Nat Med. (2021) 27:582–4. doi:  10.1038/s41591-021-01312-x [DOI] [PubMed] [Google Scholar]
  • 37. Dhiman P, Ma J, Andaur Navarro CL, Speich B, Bullock G, Damen JA, et al. Reporting of prognostic clinical prediction models based on machine learning methods in oncology needs to be improved. J Clin Epidemiol. (2021) 138:60–72. doi:  10.1016/j.jclinepi.2021.06.024 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Dhiman P, Ma J, Andaur Navarro CL, Speich B, Bullock G, Damen JA, et al. Risk of bias of prognostic models developed using machine learning: A systematic review in oncology. Diagn Progn Res. (2022) 6:13. doi:  10.1186/s41512-022-00126-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Hantel A, Clancy DD, Kehl KL, Marron JM, Van Allen EM, Abel GA. A process framework for ethically deploying artificial intelligence in oncology. J Clin Oncol. (2022) 40:3907–11. doi:  10.1200/JCO.22.01113 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Hantel A, Walsh TP, Marron JM, Kehl KL, Sharp R, Van Allen E, et al. Perspectives of oncologists on the ethical implications of using artificial intelligence for cancer care. JAMA Netw Open. (2024) 7:e244077. doi:  10.1001/jamanetworkopen.2024.4077 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Muehlematter UJ, Daniore P, Vokinger KN. Approval of artificial intelligence and machine learning-based medical devices in the USA and Europe (2015–20): A comparative analysis. Lancet Digital Health. (2021) 3:e195–203. doi:  10.1016/S2589-7500(20)30292-2 [DOI] [PubMed] [Google Scholar]
  • 42. Finlayson SG, Subbaswamy A, Singh K, Kohane IS. The clinician and dataset shift in artificial intelligence. N Engl J Med. (2021) 385:283–6. doi:  10.1056/NEJMc2104626 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Kolla L, Parikh RB. Uses and limitations of artificial intelligence for oncology. Cancer. (2024) 130:2101–7. doi:  10.1002/cncr.35307 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Elemento O, Khozin S, Sternberg CN. The use of artificial intelligence for cancer therapeutic decision-making. NEJM AI. (2025) 2:10.1056/AIra2401164. doi:  10.1056/AIra2401164 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Shah A, Mitchell S, Coad G, Michels D. Safe and responsible use of artificial intelligence in health care: Current regulatory landscape and considerations for regulatory policy. JCO Clin Cancer Inf. (2025) 9:e2500123. doi:  10.1200/CCI-25-00123 [DOI] [PubMed] [Google Scholar]
  • 46. Stetson PD, Choy J, Summerville N, Baldwin-Medsker A, Mak J, Chatterjee A, et al. Responsible Artificial Intelligence governance in oncology. NPJ Digital Med. (2025) 8:Article 407. doi:  10.1038/s41746-025-01794-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Liu X, Rivera SC, Moher D, Calvert MJ, Denniston AK, The CONSORT-AI and SPIRIT-AI Working Group . Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. BMJ. (2020) 370:m3164. doi:  10.1136/bmj.m3164 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.


Articles from Frontiers in Oncology are provided here courtesy of Frontiers Media SA

RESOURCES