Abstract
Background
Traditional face-to-face mental health treatments are often limited by time and space. Thanks to the development of advanced large language models (LLMs), digital mental health treatments can provide personalized advice to patients and improve compliance. However, in the field of CBT-I, specialized, real-time interactive dialogue platforms have not been fully developed.
Methods
Our research team construct an eCBT-I intelligent dialogue system based on the RAG architecture, aiming to provide an example of the deep integration of CBT-I knowledge graphs and large language models. Furthermore, in order to optimize the performance of the system’s core language generation module on the insomnia dialogue dataset, we systematically include eight mainstream large language models (ChatGLM2-6b, ChatGLM3-6b, Baichuan-7b, Baichuan-13b, Qwen-7b, Qwen2-7b, Llama-2-7b-chat-hf, and Llama-2-13b-chat-hf) and three adaptation strategies (LoRA, QLoRA, and Freeze). We screen the suitability of the three adaptation strategies for the eight major language models in the group, and thus determine the best adaptation method for each language model to maximize performance improvement. The eight best-adapted language models are then evaluated in three dimensions to compare their performance on the small sample sleep dialogue dataset and the C-eval dataset. All subjects that evaluated under experimental conditions are historical medical records and patients who did not exhibit delirium and had normal language expression abilities.
Results
Through the matching of model characteristics to adaptation strategies and the horizontal evaluation of multiple models, we compare the contribution of different fine-tuning strategies to the performance improvement of different language models on the small insomnia dialogue dataset, and finally determine that Qwen2-7b (Freeze) is the model with the best performance on the insomnia dialogue dataset.
Conclusions
This study effectively integrates the CBT-I knowledge graph with the large language model through the RAG architecture, which improves the professionalism of the eCBT-I intelligent dialogue system. The systematic fine-tuning method selection process and the confirmation of the optimal model not only improve the adaptability of the large language model in the CBT-I task, but also provide a useful paradigm for AI applications in medical subfields with resource constraints and difficulties in data collection, laying a solid foundation for more accurate and efficient digital CBT-I clinical practice in the future.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12967-025-06871-y.
Keywords: Mental health, eCBT-I, RAG architecture, Large language models, Adaptation strategy
Introduction
Insomnia, a prevalent disorder affecting up to 10% of the adult population, has been shown to worsen with age [1]. When insomnia symptoms worsen, insomnia disorder syndrome may develop, which induces a series of negative physical conditions such as somatic complaints, psychological distress, mental health, physical fatigue, and impaired quality of life [2, 3]. The financial burden of insomnia on healthcare systems is significant, with annual expenditures in the United States reported to exceed $100 billion [4]. In addition to conventional pharmaceutical interventions, international guidelines advocate cognitive behavioral therapy for insomnia (CBT-I) as a primary treatment modality [5]. CBT-I includes a multifaceted approach, targeting behavioral, cognitive, and physiological factors to modify maladaptive behaviors and erroneous beliefs concerning sleep and insomnia [6]. While medication has been shown to be effective in the short term, CBT-I has been demonstrated to offer more sustained benefits in addressing insomnia as a chronic condition [7]. However, given the limited availability of individual face-to-face (F2F) CBT-I in health care systems, alternative Internet-delivered cognitive-behavioral therapy for insomnia (eCBT-I) settings (e.g., group or internet-based CBT-I) have been proposed, and evidence for their efficacy exists [8, 9].
The non-response rate of CBT-I treatment in a F2F environment is high, and patients are easily unable to adhere to long-term treatment due to space and time constraints [10]. The eCBT-I retains the combinations of sleep hygiene, sleep restriction, stimulus control, relaxation therapy/mindfulness, and cognitive therapy, and can be used remotely through a website and downloadable apps, overcoming the limitations of traditional F2F treatment. Several previous studies have confirmed that eCBT-I could significantly reduce insomnia symptoms while improving patients’ mental health. Jennifer N. Felder and others tested and compared the effectiveness of eCBT-I with standard CBT-I in pregnant women with insomnia. They found that, compared with pregnant women randomized to standard care, those randomized to eCBT-I not only had a statistically significant improvement in insomnia symptom severity (difference = -0.36; 95% CI, -0.48 to -0.23; χ2 = 29.8; P < 0.001; d = -1.03), but all secondary outcomes also showed statistically significant improvements [11]. In addition, eCBT-I could improve overall sleep quality in diverse populations, including the elderly [12], breast cancer survivors [13], and patients with COVID-19 [14].Therefore, it has strong scalability and application potential.
Artificial intelligence algorithms have already demonstrated their powerful potential in a number of fields, such as machine learning evaluation [15, 16], medical image segmentation [17, 18], treatment effect prediction [19–22] and so on. Large language models (LLMs) are a class of complex neural network models based on extensive pre-training and transfer learning [23]. They achieve real-time dynamic adaptation of user input and natural language output through deep learning on massive amounts of text data. In recent years, the intersection of psychiatry and large-scale language models has become a rapidly developing field with potential applications ranging from diagnostic support to treatment interventions. For example, LLMs have already shown extraordinary potential in identifying mental health problems and assessing suicide risk, providing valuable insights to clinicians [24]. They could effectively support the diagnostic process of mental illness, with up to 86.9% accuracy in identifying clinical symptoms associated with depression and anxiety [25]. The application of LLMs in the field of CBT-I also shows great potential. By quickly integrating CBT-related professional information and providing personalized advice, LLM can not only reduce the workload of psychotherapists, but also improve users’ engagement, expand treatment accessibility and reduce related costs. In 2024, Zhang et al. constructed the CBT-BENCH evaluation benchmark and systematically evaluated the performance of LLMs in CBT tasks. The results showed that LLM performed well in memorizing and reproducing basic CBT knowledge and answering multiple-choice questions related to CBT theory, but its performance still needs to be improved when dealing with specific questions and answers in patients’ complex real-life situations [26]. Meanwhile, Talha et al. used three adaptation strategies (Mistral 7b v0.3, Qwen 2.5 7b and Llama 3.1 8b) to improve the performance of LLMs, the results showed that the model fine-adapted specifically for CBT significantly outperformed the model fine-adapted with instructions only, with an average improvement of 11. 33 points (P
0.001) on the Clinical Treatment Rating Scale (CTRS) total score [27]. However, it still had deficiencies in contextual coherence and depth exploration.
Against this background, we construct an eCBT-I dialog system for insomnia treatment based on LLMs. To improve the performance of the LLMs, which is the core component of the system, we systematically incorporate eight common LLMs and three adaptation methods. These models differ significantly in terms of parameter size and ecological support, providing us with an ideal platform to evaluate the diversity of adaptation methods and the impact of model parameters. Meanwhile, the application of large language models in the medical domain is often limited by the amount of data due to issues such as patient privacy, ethics, and instrumentation specifications. Thus, by comparing the contribution of different adaptation strategies to model performance improvement in a small-data environment, we can test whether the adapted small-parameter model can match or even exceed the performance of the large-parameter model in a given task. If this positive effect is confirmed, this study will not only provide valuable experience for the selection and adaptation methods of LLMs in the field of sleep, but also lay a solid theoretical and practical foundation for efficient model training under limited data conditions in clinical and other fields. The sleep LLM is just one example of the application of small-sample adaptation, and its methodology and experience are expected to be further extended to more professional fields.
Materials and methods
Ethics and quality standards
This study was approved by the Medical Ethics Review Committee of the First Affiliated Hospital of Wenzhou Medical University for the use of patient cases in the study (approval number YS2024 No. 597). The institutional review board granted a waiver of informed consent given the retrospective nature of the study and the anonymity and unidentifiability of the data from the dialog process. Our methods and reporting adhere strictly to the Standards for Quality Improvement Reporting Excellence (SQUIRE) guidelines, which meet the ethical and quality standards for this quality improvement study.
CBT-I dialogue dataset
The data collection period for the CBT-I treatment of insomnia dialogues is from November 28, 2023 to August 14, 2024. During this period, a total of 22,780 CBT-I dialogues are recorded, and 2387 dialogues that met our research criteria are selected for inclusion in the study. The scope of the collection includes sleep guides, professional books, insomnia treatment-related literature, and transcripts of therapist interviews during the treatment of insomnia using CBT-I therapy. These dialogic processes and knowledge content are primarily captured in text format and contain key information about the treatment process.
Inclusion and exclusion criteria
The data cleaning process consists of two main parts: initial AI screening and manual re-screening. Initial AI screening: (1) Conversations with semantic ambiguity or logical breaks are screened out based on the BERT semantic coherence score. Conversations with a score of less than 0.6 are excluded; (2) Dialog fragments with high sleep relevance are screened using the open-source large language model Qwen2.5-72B-Instruct. Manual re-screening: (1) Recommendations that conflict with or are irrelevant to the CBT-I principles are screened out. (2) Dialog fragments containing personally identifiable health information (PHI) and duplicate dialog texts are removed. After rigorous screening, we finally obtain 2387 pieces of dialog data that meets the research criteria, providing a solid foundation for subsequent model training.
The labeling of the dataset is performed from September 1, 2024 to November 24, 2024. The labeling of the conversations mainly includes two types of labels: instruction and output, which represent the patient’s questions and the professional’s responses based on CBT-I, respectively. The evaluation and screening of the relevance of the conversations in the dataset to CBT-I is performed by five people with medical training and extensive experience under the guidance of medical experts. These medical experts are all licensed physicians in China who have obtained their medical practice licenses and are currently employed as doctors at the First Affiliated Hospital of Wenzhou Medical University.
Dataset division and external verification
To ensure the scientific nature of the research and the reliability of the model training results, the dialog dataset is randomly divided into three parts to ensure that there is no overlap of patients between subsets: the training set accounts for 80%, containing 1909 dialogues; the validation set accounts for 10%, containing 239 dialogues; and the test set accounts for 10%, containing 239 dialogues. The test set is isolated in a separate section and maintained independently by experts. This separation strategy aims to maintain the representativeness and balance of the data to avoid overfitting or underfitting problems during model training. In addition, the C-eval dataset is a publicly available external dataset. It contains 13,948 multiple-choice questions in 52 disciplines and four levels of difficulty, and can be used to evaluate the coverage of model knowledge in different domains, as well as the reasoning ability and comprehensive decision-making level when faced with complex problems [28].
Algorithm design and model selection
Construction of the eCBT-I dialogue system
In this study, we carefully design the RAG (Retrieval-Augmented Generation) architecture as the basis for this intelligent system. The technical architecture of the RAG model consists of two main parts: the retrieval module (Retriever) and the generation module (Generator). The retrieval module uses a dual-encoder model to perform efficient vectorized retrieval, mapping user queries and documents in the large CBT-I knowledge base to the same vector space, and accurately filtering relevant content through similarity calculation to provide high-quality input for the generation module. The generation module relies on a generative language model to generate coherent, accurate, and informative responses based on the retrieved knowledge documents, ensuring the system’s professionalism and reliability in the therapeutic dialog task.
Adaptation strategy based on model characteristics
To achieve an accurate adaptation of the large language model to the CBT-I task, we construct a systematic adaptation strategy selection framework. Firstly, we rigorously select eight different open-source models (ChatGLM2-6b, ChatGLM3-6b, Baichuan-7b, Baichuan-13b, Qwen-7b, Qwen2-7b, Llama-2-7b-chat-hf, Llama-2-13b-chat-hf) with parameters ranging from 6B to 13B and covering different architecture types (ChatGLM, Llama, Qwen, Baichuan). Each model has three possible adaptation strategies to choose from: Low-rank adaptation (LoRA), Quantized LoRA (QLoRA), and Parameter freeze (Freeze). In particular, these three adaptation methods can achieve different levels of coverage of adaptation strategies: LoRA adjusts only a small number of trainable parameters by adding a low-rank adaptation matrix to a specific layer, thereby improving model adaptability while maintaining computational efficiency. QLoRA introduces 4-bit quantization on top of LoRA to further reduce memory requirements, allowing large models to be adapted efficiently on ordinary consumer-grade GPUs. Freeze is the lightest method, freezing most of the pre-trained parameters and adapting only certain layers to reduce the computational overhead and improve the generalization ability. In summary, these three adaptation strategies provide optimization solutions for models in terms of parameter update methods, computational resource requirements, applicable scenarios, and other aspects, allowing them to adapt to different framework characteristics.
During the experiment, we implement these three adaptation strategies for each of the eight candidate models and optimize the model parameters using Adam optimizer. This optimizer combines the advantages of Adam with weight decay regularization to address the problem of overfitting and ensure stable convergence during training. The training process spans 450 epochs, during which the model is iteratively adapted over multiple rounds on our dataset to fully adapt to the requirements of the CBT-I task and improve the quality of generation.
Selection of the best-adapted language models
After determining the best adaptation strategy for each model, we systematically evaluate the inference and parameter verification of the eight best-adapted large language models. During the specific training process, we first use an independent validation set to screen the hyperparameters of each model and selected the parameter configurations that performed best on the validation set by monitoring key indicators such as BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L. We then use this optimal configuration to infer on a pre-divided internal test set of the same origin to evaluate the model’s performance under familiar data distributions. At the same time, we also infer on a completely independent external test set C-eval to measure the model’s coverage of domain knowledge and its reasoning ability when faced with complex problems. Based on these systematic inference and evaluation processes, we comprehensively compare the inference efficiency, generation quality, and task adaptability of each model to determine the best-adapted language model that performed best on the CBT-I task.
Indicator setting and model evaluation
To investigate the adaptability of different adaptation strategies to different architectures, we use four evaluation indicators commonly used in the field of large language models. They are Bilingual Evaluation Understudy-4 (BLEU-4) [29], Recall-Oriented Understudy for Gisting Evaluation-1 (ROUGE-1), Recall-Oriented Understudy for Gisting Evaluation-2 (ROUGE-2), and Recall-Oriented Understudy for Gisting Evaluation-L (ROUGE-L) [30]. Specifically, ROUGE-1 and ROUGE-2 focus on measuring the overlap between the generated text and the reference text, while ROUGE-L generally assigns higher weights to reflect the overall coherence and logical consistency of the generated content. BLEU-4, on the other hand, mainly reflects the accuracy and linguistic fluency of the generated text at the 4-gram level. 4-gram refers to the process of breaking text into sequences of four consecutive words. Together, these indicators provide a multi-faceted assessment of the semantic coherence, accuracy, and information retrieval capabilities of the text generated by the language model. Without considering the architecture, Bonferroni-corrected Wilcoxon Signed-Rank Test is used to examine the scores of each of the three adaptation methods on each of the four metrics to determine whether the adaptation methods themselves are superior or inferior. To quantify the combined performance of the models in terms of their performance in adopting different adaptation strategies, we weight and sum the total scores of the standardized scores on these four indicators and averaged the scores for each adaptation method. The fine-tuning strategy with the highest total score for each model architecture is determined to be the best adaptation strategy for that model. The average score for each adaptation strategy is also assigned weights based on expert opinion.
After determining the optimal adaptation strategy for each model, we implement a three-dimensional verification framework to evaluate the clinical applicability of the model in the CBT-I dialog task. In terms of the model’s semantic fidelity, the original BLEU-4 and ROUGE series (L/1/2) indicators can quantify the degree of lexical overlap and sentence structure matching between the generated text and the standard response to evaluate the semantic and structural consistency between the model-generated text and the reference text. In terms of the model’s logical reasoning, the newly introduced C-eval indicator is specifically designed for large-scale Chinese language models and can evaluate the model’s performance in logical reasoning and expertise acquisition. In addition, we also measure the technical feasibility of the model by evaluating the model’s inference time, response time, and deployment parameters. The three-dimensional comprehensive evaluation is based on the statistical and analytical methods in the open-source statistical tool Python.
Since large language models may generate illusory or harmful responses to user questions in certain special circumstances, and there are no automated indicators for assessing harmfulness [31], we use a 5-point Likert scale (1 = Strongly Agree (Extremely harmful), 2 = Agree, 3 = Neutral, 4 = Disagree, 5 = Strongly Disagree) to manually assess their potential harmfulness [32]. Here, harmfulness refers to answers that may cause physical or psychological harm, or unintended changes in treatment or adherence due to misinterpretation of information. All manual assessments are made by experts, with the process strictly recorded and reasons for the scores provided.
Results
Construction of CBT-I dialogue dataset and quality control
From November 28, 2023, to August 14, 2024, our research team collected a substantial dataset comprising 22,780 records through five distinct methods: 21,796 records were obtained from the database, 64 were sourced from the sleep guide, 519 were drawn from the book, 38 were retrieved from sleep-related literature, and 363 were acquired from the therapist’s CBT-I treatment process records. To ensure the integrity and relevance of the conversational data set, a meticulous cleaning process is carried out by five experienced quality controllers, each of whom has extensive expertise in mental health counseling. The cleaning of the data set is mainly divided into two parts: AI initial screening and manual re-screening. AI initial screening: (1) We screen for semantic coherence based on BERT scores, excluding a total of 487 dialogue sequences with semantic ambiguity or logical breaks (BERT score < 0.6) (Fig. 1a); (2) We use the currently powerful open-source large language model Qwen2.5-72B-Instruct to screen for sleep-related dialogue segments, excluding a total of 7461 dialogues. Manual re-screening: (1) We then proceed to screen the remaining conversations for relevance to CBT-I treatment, a process that ultimately led to the exclusion of 13,267 conversations. This stringent step is taken to ensure that the text content is highly relevant to CBT-I therapy. (2) The conversation fragments containing sensitive personally identifiable health information (PHI) and repeated conversation texts are removed, totaling 309. This step is designed to screen out conversation data that is not related to the research objectives and ensure that the text content is relevant to the context of potential sleep disorders. Through a rigorous data filtering process, 2387 high-quality samples are included from the initial collection of 22,780 conversations for model training (Fig. 2). The C-eval dataset, which contains 13,948 multiple-choice questions, is categorized into four broad subject categories and 52 subcategories, with four levels of difficulty. This comprehensive evaluation of the model’s knowledge and logical reasoning abilities is instrumental in ensuring the validity and reliability of the study’s findings. It is imperative to note that all dialogue data utilized in this study obtained ethical approval from the Ethics Review Committee of the First Affiliated Hospital of Wenzhou Medical University, ensuring compliance with established ethical standards.
Fig. 1.
Structure of the AI-assisted CBT-I dialogue system and language model. (a) BERT score for the final dataset included in the study. (b) The system processes the user’s questions and returns a response. (c) Comparisons of model parameters and reasoning times for eight common language models on the market. (d) Four evaluation indicators and weight allocation for model performance. (e) Structure of the Qwen2-7b
Fig. 2.
Data entry and exit group diagram
To facilitate systematic analysis, the final data set of 2387 dialogues is randomly divided into three subsets: 80% for the training set (1909 dialogues), 10% for the validation set (239 dialogues), and 10% for the test set (239 dialogues). The distribution of dialogues within each subset has been meticulously designed to ensure randomness and balance, thereby contributing to the robustness of the analysis. An external dataset, C-eval, which is isolated and maintained separately, is utilized for objective evaluation of model performance.
Deployment of eCBT-I dialogue systems
The intelligent dialog system developed in this study innovatively combines the Retrieval-Augmented Generation (RAG) framework with a structured CBT-I knowledge graph to achieve efficient and accurate therapeutic dialog support. The technical architecture of the RAG model consists of two main parts: the retriever and the generator. The retrieval module adopts the dual-encoder model, which converts the user’s query and the relevant documents in the knowledge base into vector representations by pre-trained encoders, and then uses cosine similarity to achieve efficient vector retrieval. The final output is the documents or paragraphs related to the query as context input to the generation module. The generation module is based on a Transformer decoder, which integrates the retrieved knowledge information and generates coherent and informative responses through a self-attention mechanism and layer normalization (Fig. 1b).
The performance of the language model is closely related to the text generation capability of the generation module, which ensures that the system can output coherent, accurate, and informative responses. We evaluate the eight commonly used benchmark language models ChatGLM2-6b, ChatGLM3-6b, Baichuan-7b, Baichuan-13b, Qwen-7b, Qwen2-7b, Llama-2-7b-chat-hf, and Llama-2-13b-chat-hf in terms of number of parameters and inference time to measure their training costs in our local deployment. The order of magnitude of the parameters of ChatGLM2-6b and ChatGLM3-6b is about 6 billion; the order of magnitude of the parameters of Baichuan-7b, Qwen-7b, Qwen2-7b, Llama-2-7b-chat-hf has a parameter order of about 7 billion, while Baichuan-13b and Llama-2-13b-chat-hf have a parameter order of 13 billion (Fig. 1c). The hierarchical structure of the core Transformer module of the large language model includes, from bottom to top, a text input layer, an embedding layer, a decoding layer, and a text output layer. The text input layer first tokenizes and preprocesses the raw input, breaking it down into discrete word tokens. The embedding layer then maps these word tokens to continuous vector representations and injects positional information through positional encoding to capture the relative positions of words in the sequence. The subsequent decoding layer consists of several transformer blocks, each of which contains a multi-head self-attention mechanism, a feed-forward neural network, a residual connection, and a layer normalization operation. These components work together to extract deep contextual semantics and long-range dependencies. Finally, the text output layer converts the vector representation processed by the decoding layer into a probability distribution on the vocabulary table to generate the final natural language text (Fig. 1e).
Selection of adaptation methods for the eight large language models
In the adaptation strategy based on the characteristics of the architecture, our candidate is a combination of three adaptation methods and eight large language models. Four key indicators (BLEU-4 and ROUGE-1, 2, L) are used in the study to screen the performance of the adapted models. In addition, to quantify the overall performance of the model, we also calculate the total score by weighting the scores of the four indicators. The weight distributions of BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L were 0.1, 0.3, 0.3, and 0.3, respectively (Fig. 1d).
The eight candidate models are evaluated on the training, validation, and test sets after 450 training cycles for each of the three adaptation strategies. The heat map shows the changes in the four indicators (BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators) for the eight language models under the three adaptation strategies (Fig. 3a). In addition, we evaluate the performance of the three adaptation strategies on the four indicators without considering the model architecture, and find that there is no significant difference between the different adaptation strategies for each indicator (p > 0.05), which also indicates that the adaptation strategy itself does not have a better or worse effect, and its adaptation effect is closely related to the architectural characteristics of the model (Fig. 3b). After 450 training rounds, ChatGLM2-6b (LoRA) achieves 0.018587, 0.131443, 0.016057, and 0.082684 for the BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators, respectively. ChatGLM2-6b (QLoRA) achieves 0.019731, 0.139191, 0.019402, and 0. 085231 on the BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L indicators respectively. ChatGLM2-6b (Freeze) achieves 0.059165, 0.208244, 0.055209, 0.157865 on the indicators BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L, UGE-2, and ROUGE-L. It can be seen that ChatGLM2-6b adapted by Freeze shows the optimal dialog performance of the model, and its four indicators are higher than those of the other two adaptation methods, significantly improved. ChatGLM3-6b (LoRA) reaches 0.033459, 0.150845, 0.031321, and 0.11858 in BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L, respectively; ChatGLM3-6b (QLoRA) reaches 0.031888, 0.136302, 0.026611, and 0. 106,968 on the indicators BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L respectively. ChatGLM3-6b (Freeze) reaches 0.052632, 0.188763, 0.049453, 0.147771 on the indicators BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L, UGE-2, and ROUGE-L. It can be seen that ChatGLM3-6b adapted by Freeze shows the optimal dialog performance of the model, and its four indicators are higher than those of the other two adaptation methods, significantly improved. Baichuan-7b (LoRA) gets 0.06214, 0.213111, 0.058696, and 0.148064 in BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L, respectively. Baichuan-7b (QLoRA) gets 0.05539, 0.205105, 0.054908, 0. 144,452 on the BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L indicators, respectively. Baichuan-7b (Freeze) gets 0.009481, 0.0478, 0.000664, 0.042857 on the BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators. Baichuan-7b adapted by LoRA shows the optimal dialog performance of the model, with all four indicators significantly higher than the other two adaptation methods. Using the same comparison method, we tentatively selected the best adaptation strategy for each language model: except for Baichuan-13b, whose best adaptation strategy is QLoRA, and Baichuan-7b, whose best adaptation strategy is LoRA, the best adaptation strategy for the remaining models is Freeze.
Fig. 3.
Selection of the best adaptation method for each language model. (a) Comparison of the three adaptation methods (LoRA, QLoRA and Freeze) for the eight language models using four evaluation indicators. (b) Performance of the three adaptation methods under each of the four indicators. (c) Selection of the normalized total score for each of the three adaptation methods for the eight language models
To further quantify the overall performance of each model, we calculate a weighted total score based on the four indicators evaluated as a whole, and the weighted total scores and contribution ratios of each model are displayed in the form of a stacked column chart (Fig. 3b). The “*” mark represents the best adaptation strategy for each model. The average scores of the eight models for LoRA, QLoRA and Freeze are 0.352212, 0.379534, and 0.418145. This systematic evaluation not only comprehensively reflects the model’s advantages in capturing long and short phrases, maintaining semantic coherence, and overall generation quality, but also provides a solid theoretical and practical basis for selecting the best-adapted model in the CBT-I task.
Three-dimensional evaluations of the best-adapted models
After determining the best adaptation strategy for each language model, the performance of the eight best-adapted language models is evaluated on training, validation, and test sets. The four indicators in the line graph illustrate the evolution of the performance during training. After optimization over 450 training cycles, the performance of all eight models improves (Fig. 4a). For example, ChatGLM2-6b (Freeze) achieves BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0.071818, 0.212498, 0.065676, and 0.166059, respectively, on the training set. ChatGLM2-6b (Freeze) gets BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0. 054809, 0.197145, 0.049382, and 0.15058, respectively. ChatGLM2-6b (Freeze) attains BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0.0591 65, 0.208244, 0.055209, 0.157865. It is worth noting that Qwen2-7b (Freeze) on the training set achieves BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0.239775, 0.347811, 0.222394, and 0.315154. Qwen2-7b (Freeze) on the validation set gets BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0. 165,097, 0.283705, and 0.154503. Qwen2-7b (Freeze) on the testing set attains BLEU-4, ROUGE-1, ROUGE-2, and ROUGE-L indicators of 0.209688, 0.326662, 0.195259, and 0.291442, respectively. Although in general larger parameters often indicate higher model accuracy, these results show that with an appropriate adaptation strategy, the Qwen2-7b (Freeze) model with smaller parameters not only performs well during training, but also maintains fairly stable and close performance on the validation and test sets, reflecting its strong semantic fidelity.
Fig. 4.
Eight best-adapted language model performance. (a) Changes in BLEU-4, ROUGE-1, ROUGE2 and ROUGE-L indicators in the training 450 epochs, training set and test set. (b) External dataset for models’ C-eval scores. (c) Comparisons of model parameters and reasoning times for eight adjusted language model
In addition, after optimizing the model to extend its task-specific functionality, we evaluate it using the external dataset C-eval. The test provides a comparison of the average accuracy of the model: Qwen-7b (Freeze) reaches 0.6062, while Baichuan-13b (QLoRA) and Qwen2-7b (Freeze) achieves 0.5156 and 0.8076 in this test. High scores on C-eval set mean that the model has better knowledge reserves and logical reasoning capabilities.
The technical feasibility of the model is also a major focus of our evaluation. Specifically, ChatGLM2-6b (Freeze) has 6.244 billion parameters and an inference response time of 5.71 s, while ChatGLM3-6b (Freeze) has the same number of parameters but a response time that is extended to 22.24 s. Baichuan-7b (LoRA) has 7.018 billion parameters and takes only 1.54 s to respond. In contrast, Baichuan-13b (QLoRA) has a response time of 4.17 s with 13.293 billion parameters. Llama-2-7b-chat-hf (Freeze) and Llama-2-13b-chat-hf (Freeze) have 6.738 billion and 13.016 billion parameters, respectively, with response times of 12.41 s and 5.07 s. In addition, Qwen-7b (Freeze) and Qwen2-7b (Freeze) reach response times of 18.76 s and 6.16 s with 7.721 billion and 7.616 billion parameters, respectively (Fig. 4b and c). Overall, these data highlight the significant differences in technical feasibility and inference efficiency between different architectures and adaptation strategies, providing an important reference for practical deployment.
To further assess the potential harmfulness of the model’s generated responses, eCBT-I Dialogue System conducts 180 simulated dialogue sessions with insomnia patients. The harmfulness of the model’s responses is manually evaluated using a Likert scale, with scores ranging from 1 to 5 representing different levels (1 = Strongly Agree (Extremely harmful), 2 = Agree, 3 = Neutral, 4 = Disagree, 5 = Strongly Disagree). The final statistical results show that the average score for the 200 rounds of dialogue is 4.89, with 0% Strongly Agree (Extremely harmful), 0% Agree (Harmful), 2.2% Neutral, 6.6% Disagree, and 91.2% Strongly Disagree (Table S4). We find that 2.2% of neutral responses are often related to the use of sleeping pills, which may be due to a lack of context or professional advice from doctors.
Discussion and conclusion
Although international guidelines have already endorsed CBT-I as the first-line treatment for insomnia, traditional F2F CBT-I relies mostly on standardized treatment plans and is limited by patient access to medical care, which cannot fully meet the individual needs of different patients [10]. The emergence of eCBT-I has completely changed the way insomnia is treated, and the use of online digital platforms has made this treatment more accessible, effective, and engaging for patients [33–35]. In this work, we design and implement an intelligent dialog system based on the RAG architecture. Through dynamic interaction with users, it actively obtains key sleep information and accordingly generates customized CBT-I treatment recommendations to help users improve their sleep quality. In order to improve the capability of the RAG generation module in the system on the sleep dialog dataset, we rigorously investigate and evaluate the LLMs, which is the core module in the generation module. Firstly, we determine the adaptation strategies for each language model based on its characteristics, and then conduct a three-dimensional evaluation of the best-adapted models to select the model and its adaptation method that is most suitable for the sleep disorder dialog dataset.
Our results show the great performance of Freeze strategy to most LLMs under small samples. We suppose that the Freeze strategy can effectively prevent overfitting in small samples by updating only the topmost layer or a very small number of parameters, and maximize the general-purpose language capability of the pre-trained models. In addition, the Freeze strategy also significantly reduces the computation and memory overheads, which makes the training more efficient and stable when the resources are limited. However, for Baichuan-13B and Baichuan-7B with a large number of parameters, QLoRA/LoRA with an appropriate number of trainable parameters can better activate the model potential. In the end, Qwen2-7b (Freeze) has the best performance among the sleep language models, especially in knowledge mastery and logical reasoning tasks. For example, in the C-eval evaluation, its average accuracy is as high as 0.8076, significantly ahead of other models (such as Qwen-7b (Freeze) 0.6062 and Baichuan-13b (QLoRA) 0.5156). This advantage is not only reflected in the accuracy, but also in the consistency and semantic coherence of the generated text, as measured by indicators such as BLEU and ROUGE. Further analysis shows that Qwen2-7b (Freeze) achieves a better parameter usage efficiency with its Freeze adaptation strategy. Its parameter size of about 76.16 billion provides a good balance with high precision output, while its inference response time (about 6.16 s) is also at an excellent level. This shows that the model effectively adapts to downstream tasks while freezing most of the pre-trained weights, and avoids the risk of overfitting and wasting computational resources caused by full parameter updates by adapting only a small number of parameters. This gives it a clear advantage in terms of engineering deployment and real-time response. Overall, Qwen2-7b (Freeze) performed well in multidimensional evaluations, demonstrating its overall strength in integrating knowledge, reasoning, and generative capabilities. Our work fully demonstrates that with the strategy of freezing most of the pre-trained weights and adaptation only a small number of parameters, low-parameter models in small sample vertical fields have the potential to match or exceed the performance of models with larger parameters on specific tasks, thereby showing significant advantages in terms of engineering deployment and real-time response [36].
Although eCBT-I has demonstrated efficacy in the general population, recent studies suggested that its effectiveness may be limited in populations with specific neurological or cognitive vulnerabilities. For example, Malarkey et al. (2024) conducted an evaluation of an eCBT-I program for patients with traumatic brain injury (TBI), noting that while treatment outcomes were significant, adherence rates among patients were low [37]. These findings may stem from difficulties patients face in maintaining attention, processing complex information, or sustaining treatment motivation in the absence of human support. Leveraging the flexibility and scalability of digital platforms, personalized eCBT-I platforms can be developed to provide treatment and improve adherence for diverse populations, including older adults, cancer survivors, and frontline workers [38]. For example, ShUTi OASIS is specifically designed for older adults [12], while SleepCare is tailored for nurses working shift schedules [39].
Compared with existing digital platforms, the RAG intelligent dialogue system we proposed not only achieves technological innovation in system architecture, but also provides empirical evidence through comprehensive evaluation of eight language models and three adaptation methods under large-scale workloads. Additionally, the system incorporates a user feedback mechanism. Real-time user feedback is recorded by the system and utilized to optimize future dialogue workflows and recommendation generation strategies, thereby enhancing the system’s practical value and ensuring continuous improvement of its recommendations. In particular, the RAG system can be combined with a quality control framework protocol to synchronize the monitoring of the potential harmfulness and hallucinogenicity of generated text [40]. In previous digital mental health studies, machine system response times of 15–20 s were generally considered acceptable for asynchronous or semi-synchronous human-machine interactions [41]. Kocaballi et al. conducted dialogue design on ChatGPT and found that within a dialogue context of 4,000 tokens, this conversational AI was able to understand natural language prompts well, maintain contextual relevance in the dialogue, and respond quickly [42]. To further reduce system latency and improve response speed, the following latency reduction strategies can be adopted during deployment: (1) Model distillation or quantization to create lightweight, low-latency versions of large models [43]; (2) Loading caches and reusing common question responses; (3) Edge deployment or GPU inference acceleration to reduce server load and round-trip latency. The eCBT-I system demonstrated in our research proves that small-parameter models fine-tuned with limited data have the potential to achieve or even surpass the performance of large-parameter models [44], providing a solid theoretical and practical foundation for efficient fine-tuning and real-time applications in small-sample vertical domains in the future. Our workflow not only optimizes the adaptability of intelligent dialogue systems in CBT-I tasks but also provides a reference paradigm for broader medical AI applications. The core ideas in our research, such as fine-grained adaptation strategy selection, multi-model horizontal evaluation, and user feedback-driven optimization, can be extended to other clinical decision support, medical education, and patient self-management systems, particularly in resource-constrained medical subfields [45–47].
However, there are several limitations in our research. First, we only include three common adaptation methods and eight major language models, and we have not yet extended the screening to more objects. As of the publication of this paper, emerging LLMs such as Qwen 3 and Llama-4, released in April 2025, have demonstrated significant potential in long-context reasoning, multi-language instruction following, and multimodal alignment. However, due to resource and deployment constraints, these models are not included in this study. Therefore, the scope of this study is limited to the eight large language models and three adaptation methods included, and there may be language models and adaptation methods that perform better on sleep and insomnia dialogue datasets. Future research should include more methods and large language models for evaluation and screening. Besides, we only compare the models’ performance on local datasets. Future studies should extend the RAG intelligent dialogue system to insomnia patients across multiple centers and conduct randomized controlled trials to validate its therapeutic efficacy.
Additionally, Traditional face-to-face cognitive behavioral therapy (CBT-I) relies not only on verbal content but also on nonverbal cues, including facial expressions, eye contact, body posture, and tone of voice, which convey important emotional and behavioral cues to the therapist. These signals help establish therapeutic relationships, enhance empathy, and enable timely adjustments to treatment strategies. Future AI-based CBT-I systems could be enhanced by integrating multimodal capabilities, such as facial emotion recognition via webcams, emotion detection from voice intonation, and analysis of body movements using visual sensors. Incorporating these nonverbal channels could enable more sensitive detection of users’ emotional states, stress levels, and engagement, thereby allowing the system to adjust responses and interventions more effectively. However, these additional features raise important challenges related to privacy, real-time processing, and ethical deployment, particularly in clinical and home settings. Future research must address these issues.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
We thank our colleagues for helpful discussions and their helpful guidance in refining the illustrations.
Abbreviations
- CBT-I
Cognitive Behavioral Therapy for Insomnia
- F2F
Face-to-Face
- eCBT-I
Internet-delivered Cognitive-Behavioral Therapy for Insomnia
- LLMs
Large Language Models
- CTRS
Clinical Treatment Rating Scale
- PHI
Personally identifiable Health Information
- RAG
Retrieval-Augmented Generation
- LoRA
Low-rank Adaptation
- QLoRA
Quantized Low-rank Adaptation
- Freeze
Parameter Freeze
- BLEU-4
Bilingual Evaluation Understudy-4
- ROUGE-1
Recall-Oriented Understudy for Gisting Evaluation-1
- ROUGE-2
Recall-Oriented Understudy for Gisting Evaluation-2
- ROUGE-L
Recall-Oriented Understudy for Gisting Evaluation-L
- BERT
Bidirectional Encoder Representations from Transformers
Author contribution
X.Y.B.: Conceptualization, Methodology, Writing – original draft. X.Y.Z.: Investigation, Formal analysis, Data curation. D.R.Y.: Software, Validation, Visualization. H.L.: Project administration, Supervision. R.Y.W.: Resources, Investigation. Y.T.W.: Methodology, Writing – review & editing. W.H.L.: Investigation, Visualization. Y.X.: Resources, Formal analysis. L.Z.: Software, Validation. Y.Y.P.: Writing – review & editing, Funding acquisition. X.Q.W.: Methodology, Resources. X.Z.: Formal analysis, Visualization. C.L.: Funding acquisition, Resources. Y.H.L.: Supervision, Writing – review & editing. Y.Z.: Supervision, Project administration. Q.Z.: Conceptualization, Supervision, Writing – review & editing. M.Y.: Conceptualization, Supervision, Funding acquisition.
Funding
This work is supported by Zhejiang Provincial Natural Science Foundation of China (Grant No.LY21H050006), Zhejiang Provincial Medical and Health Science and Technology Plan (Grant Nos.2025KY1000 and 2024KY1262), Fundamental Research Funds for the Liaoning Universities (Grant No. LJ212410146026).
Data availability
Not applicable.
Declarations
Conflict of interest
The authors declare that the research was conducted without any commercial or financial relationships that could be construed as potential conflicts of interest.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Xueying Bao, Xingyu Zhu, Dongren Yang, Hao Lou and Ruoyun Wang contributed equally to this work.
Contributor Information
Yan Zhang, Email: zhangyan@wmu.edu.cn.
Qi Zhao, Email: zhaoqi@lnu.edu.cn.
Mei Yang, Email: yangmei1@wmu.edu.cn.
References
- 1.Riemann D, Benz F, Dressle RJ, et al. Insomnia disorder: state of the science and challenges for the future. J Sleep Res. 2022;31(4):e13604. [DOI] [PubMed] [Google Scholar]
- 2.Taylor DJ, Lichstein KL, Durrence HH. Insomnia as a health risk factor. Behav Sleep Med. 2003;1(4):227–47. [DOI] [PubMed] [Google Scholar]
- 3.Dolsen EA, Asarnow LD, Harvey AG. Insomnia as a transdiagnostic process in psychiatric disorders. Curr Psychiatry Rep. 2014;16(9):471. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Wickwire EM, Shaya FT, Scharf SM. Health economics of insomnia treatments: the return on investment for a good night’s sleep. Sleep Med Rev. 2016;30:72–82. [DOI] [PubMed] [Google Scholar]
- 5.Chan NY, Chan JWY, Li SX, et al. Non-pharmacological approaches for management of insomnia. Neurotherapeutics. 2021;18(1):32–43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Morin CM, Bootzin RR, Buysse DJ, et al. Psychological and behavioral treatment of insomnia: update of the recent evidence (1998–2004). Sleep. 2006;29(11):1398–414. [DOI] [PubMed] [Google Scholar]
- 7.Mitchell MD, Gehrman P, Perlis M, et al. Comparative effectiveness of cognitive behavioral therapy for insomnia: a systematic review. BMC Fam Pract. 2012;13(1):40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Simon L, Steinmetz L, Feige B, et al. Comparative efficacy of onsite, digital, and other settings for cognitive behavioral therapy for insomnia: a systematic review and network meta-analysis. Sci Rep. 2023;13(1):1929. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zachariae R, Lyby MS, Ritterband LM, et al. Efficacy of internet-delivered cognitive-behavioral therapy for insomnia – a systematic review and meta-analysis of randomized controlled trials. Sleep Med Rev. 2016;30:1–10. [DOI] [PubMed] [Google Scholar]
- 10.Steinmetz L, Simon L, Baumeister H, et al. Treatment effect heterogeneity of cognitive behavioral therapy for insomnia – a meta-analysis. Sleep Med Rev. 2024;77:101966. [DOI] [PubMed] [Google Scholar]
- 11.Felder JN, Epel ES, Neuhaus J, et al. Efficacy of digital cognitive behavioral therapy for the treatment of insomnia symptoms among pregnant women: a randomized clinical trial. JAMA Psychiatry. 2020;77(5):484–92. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Kyle SD, Hurry ED, Emsley R, et al. The effects of digital cognitive behavioral therapy for insomnia on cognitive function: a randomized controlled trial. Sleep. 2020;43(9):zsaa034. [DOI] [PubMed] [Google Scholar]
- 13.Starling CM, Greenberg D, Lewin D, et al. Voice-activated cognitive behavioral therapy for insomnia: a randomized clinical trial. JAMA Netw Open. 2024;7(9):e2435011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Cheng P, Casement MD, Kalmbach DA, et al. Digital cognitive behavioral therapy for insomnia promotes later health resilience during the coronavirus disease 19 (COVID-19) pandemic. Sleep. 2021;44(4):zsaa258. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Chen H, Chen F, Wang Y, et al. A machine learning model for diagnosing opportunistic infections in HIV patients: broad applicability across infection types. J Cell Mol Med. 2025;29(6):e70497. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Chen H, Ye H, Chen F, et al. Revolutionizing infection risk scoring: an ensemble from weak to strong deduction strategy and enhanced point-of-care testing tools. Adv Intell Syst. 2023;5(11):2300224. [Google Scholar]
- 17.Wu W, Huang J, Zhang M, et al. MSA-MaxNet: multi-scale attention enhanced multi-axis vision transformer network for medical image segmentation. J Cell Mol Med. 2024;28(24):e70315. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Chen X, Chen F, Liang C, et al. MRI advances in the imaging diagnosis of tuberculous meningitis: opportunities and innovations. Front Microbiol. 2023;14:1308149. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Yang X, Wang Y, Lin Y et al. A multi-task self-supervised strategy for predicting molecular properties and FGFR1 inhibitors. Adv Sci 2025;12(13):e2412987. [DOI] [PMC free article] [PubMed]
- 20.Liu L, Wei Y, Zhang Q, et al. SSCRB: predicting circRNA-RBP interaction sites using a sequence and structural feature-based attention model. IEEE J Biomedical Health Inf. 2024;28(3):1762–72. [DOI] [PubMed] [Google Scholar]
- 21.Liang C, Pan S, Wu W, et al. Glucocorticoid therapy for sepsis in the AI era: a survey on current and future approaches. Comput Struct Biotechnol J. 2024;24:292–305. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Yin S, Xu P, Jiang Y, et al. Predicting the potential associations between circrna and drug sensitivity using a multisource feature-based approach. J Cell Mol Med. 2024;28(19):e18591. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. 2023;330(9):866–9. [DOI] [PubMed] [Google Scholar]
- 24.Omar M, Soffer S, Charney AW, et al. Applications of large language models in psychiatry: a systematic review. Front Psychiatry. 2024;15:1422807. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Xu S, Yan Y, Ding Y et al. Identifying psychiatric manifestations in outpatients with depression and anxiety: a large language model-based approach. medRxiv. 10.1101/2025.01.03.24318117v1 (2025). Accessed 03 Jan 2025. https://www.medrxiv.org/content/
- 26.Zhang M, Yang X, Zhang X et al. CBT-Bench: evaluating large language models on assisting cognitive behavior therapy. arXiv. (2025). Accessed 26 Jan 2025. https://arxiv.org/abs/2410.13218
- 27.Tahir T. Fine tuning large language models to deliver CBT for depression. arXiv. (2024). Accessed 29 Nov 2024. https://arxiv.org/abs/2412.00251
- 28.Huang Y, Bai Y, Zhu Z et al. C-Eval: a multi-level multi-discipline chinese evaluation suite for foundation models. arXiv. (2023). Accessed 15 May2023. https://arxiv.org/abs/2305.08322
- 29.Kishore P, Salim R, Todd W et al. Bleu: a method for automatic evaluation of machine translation. Association Comput Linguistics 2002:311–8.
- 30.Lin CY. ROUGE: A package for automatic evaluation of summaries. Association Comput Linguistics 2004:74–81.
- 31.Alexander R, Fabbri, Wojciech K et al. SummEval: re-evaluating summarization Evaluation. arXiv. (2021). Accessed 1 Feb 2021. https://arxiv.org/abs/2007.12626
- 32.Tang L, Sun Z, Idnay B, et al. Evaluating large Language models on medical evidence summarization. NPJ Digit Med. 2023;6(1):158. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Sweetman A, Reynolds C, Richardson C. O035 digital CBT-i versus digital sleep education control in an Australian community-based cohort: A randomised controlled trail. SLEEP Adv. 2023;4(1):A12. [DOI] [PubMed] [Google Scholar]
- 34.Benz F, Grolig L, Hannibal S, et al. Investigating non-inferiority of internet-delivered versus face-to-face cognitive behavioural therapy for insomnia (CBT-I): a randomised controlled trial (iSleep well). Trials. 2024;25(1):371. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Espie CA, Henry AL. Disseminating cognitive behavioural therapy (CBT) for insomnia at scale: capitalising on the potential of digital CBT to deliver clinical guideline care. J Sleep Res. 2023;32(6):e14025. [DOI] [PubMed] [Google Scholar]
- 36.Han J, Kong T, Liu J. PepNet: an interpretable neural network for anti-inflammatory and antimicrobial peptides prediction using a pre-trained protein Language model. Commun Biology. 2024;7(1):1198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Malarkey ME, Fu AJ, Mannan N, et al. Internet-Guided cognitive behavioral therapy for insomnia among patients with traumatic brain injury: A randomized clinical trial. JAMA Netw Open. 2024;7(7):e2420090. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Luik AI, van der Zweerde T, van Straten A, et al. Digital delivery of cognitive behavioral therapy for insomnia. Curr Psychiatry Rep. 2019;21(7):50. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Ell J, Brückner HA, Johann AF, et al. Digital cognitive behavioural therapy for insomnia reduces insomnia in nurses suffering from shift work disorder: A randomised-controlled pilot trial. J Sleep Res. 2024;33(6):e14193. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Shusterman R, Waters AC, O’Neill S, et al. An active inference strategy for prompting reliable responses from large language models in medical practice. NPJ Digit Med. 2025;8(1):119. [DOI] [PMC free article] [PubMed]
- 41.Timothy W, Bickmore, Rosalind W. Picard. Establishing and maintaining long-term human-computer relationships. Association Comput Mach. 2005;12(2):293–327. [Google Scholar]
- 42.Kocaballi B, Conversational. AI-Powered Design: ChatGPT as Designer, User, and Product. arXiv. https://arxiv.org/abs/2302.07406 (2023). Accessed 15 Feb 2023.39987335 [Google Scholar]
- 43.Zhu J, Zeng L, Mo Z, et al. LMCD-OR: a large-scale, multilevel categorized diagnostic dataset for oral radiography. J Translational Med. 2024;22(1):930. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Yang X, Sun J, Jin B, et al. Multi-task aquatic toxicity prediction model based on multi-level features fusion. J Adv Res. 2025;68:477–89. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Wei S, Lu Y, Wang P, et al. Investigation of cell development and tissue structure network based on natural language processing of scRNA-seq data. J Translational Med. 2025;23(1):264. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Wu Y, Yu T, Zhang M, et al. Design and implementation of a radiomic-driven intelligent dental hospital diversion system utilizing multilabel imaging data. J Translational Med. 2024;22(1):1123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Zhai Y, Hai D, Zeng L, et al. Artificial intelligence-based evaluation of prognosis in cirrhosis. J Translational Med. 2024;22(1):933. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
Not applicable.




