Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Aug 4;29(8):e70816. doi: 10.1111/1756-185x.70816

Systematic Review on Artificial Intelligence in Rheumatology Practice: From Implementation Concerns to Imaging and Clinical Application

Raffaele Barile 1,✉, Cinzia Rotondo 1, Giulio Giancaspro 1, Nicola Maruotti 1, Francesco Paolo Cantatore 1, Addolorata Corrado 1
PMCID: PMC13436147  PMID: 42549989

ABSTRACT

Background

The integration of artificial intelligence (AI) into healthcare has shown significant promise in addressing complex diagnostic and therapeutic challenges in rheumatology. This review examines the current state of AI applications across rheumatological practice.

Objective

To systematically evaluate AI applications in rheumatology, assess their clinical performance and identify future research directions.

Methods

We conducted a comprehensive literature review of AI applications in rheumatology, focusing on diagnostic imaging, clinical decision support and disease monitoring across major rheumatic conditions.

Results

AI demonstrates promising performance across multiple domains, with diagnostic accuracies frequently exceeding 80%–90% for imaging interpretation and disease classification. Applications span from automated radiographic scoring to real‐time disease monitoring.

Conclusions

While AI shows significant potential in rheumatology, successful clinical implementation requires addressing challenges related to data quality, algorithm transparency and clinical integration.

Keywords: artificial intelligence, deep learning, diagnostic imaging, digital health, disease monitoring, machine learning, rheumatology

Practitioner Points

  • AI demonstrates strong diagnostic performance across multiple rheumatologic applications, including automated radiographic scoring (AUC 0.71–0.97), ultrasound‐based synovitis quantification, and MRI interpretation for inflammatory lesions, often matching or exceeding expert‐level accuracy. However, most studies exhibit a moderate‐to‐high risk of bias and limited adherence to reporting standards, emphasizing the need for rigorous external validation before clinical adoption.

  • Successful clinical integration requires addressing practical implementation challenges beyond algorithmic performance, including seamless integration into electronic health record systems, mitigation of alert fatigue, adequate clinician training and workflow compatibility. Current implementations remain in pilot phases, with substantial barriers related to interoperability, user acceptance and organizational change management.

  • Evidence linking AI implementation to improved patient outcomes remains limited. While AI tools show promise for earlier disease detection, flare prediction and treatment optimization, prospective studies demonstrating measurable improvements in clinical endpoints, such as disease progression rates, functional outcomes, quality of life, or healthcare utilization are absent. Future research must prioritize patient‐centered outcome measures alongside technical performance metrics.

1. Introduction

Rheumatology comprises a heterogeneous group of autoimmune and inflammatory disorders that pose substantial diagnostic and therapeutic challenges. The variability in clinical presentation and disease course demands advanced methods for diagnosis, monitoring and treatment optimization. Artificial intelligence (AI), particularly machine learning (ML), offers powerful capabilities for integrating multimodal data, such as imaging, biomarkers and electronic health records, to improve diagnostic accuracy, enable earlier detection and support personalized therapeutic strategies in rheumatology [1, 2] The field is especially well‐suited for AI applications due to the availability of rich multimodal datasets, the chronic nature of many conditions requiring longitudinal assessment and the standardized imaging protocols that provide clear targets for algorithm development.

2. Scope and Objectives

This review provides a comprehensive assessment of current AI applications across rheumatology, spanning diagnostic imaging, laboratory analysis, clinical decision support and patient monitoring. It synthesizes evidence on AI use in the major rheumatic diseases, including rheumatoid arthritis (RA), vasculitis and connective tissue disorders. It highlights emerging applications in specialized domains such as capillaroscopy and ultrasound. The analysis addresses both technical considerations, such as accuracy metrics and validation strategies, and practical challenges related to clinical integration. The strengths and limitations of existing AI methodologies are evaluated, and unmet needs in the literature are identified, offering direction for future research and pathways toward effective clinical implementation.

3. Methodological Approach

The review methodology adhered to the Preferred Reporting Items for Systematic Reviews and Meta‐Analyses guidelines and followed established methodological standards for evidence synthesis in digital health and medical artificial intelligence research. The research question was formulated according to the Population–Intervention–Comparator–Outcome framework. The population comprised patients affected by inflammatory and systemic rheumatic diseases; the intervention corresponded to AI‐based analytical approaches, including machine learning, deep learning, radiomics and natural language processing models; the comparator consisted of conventional clinical evaluation or expert interpretation; and the outcomes included diagnostic performance, predictive accuracy, risk stratification capability and clinical decision‐support effectiveness. A comprehensive literature search was performed in PubMed, Scopus and Web of Science databases to identify relevant studies published between January 2020 and March 2025. The search strategy combined Medical Subject Headings and keywords related to artificial intelligence technologies and rheumatology. Search terms included ‘artificial intelligence,’ ‘machine learning,’ ‘deep learning,’ ‘neural networks,’ ‘radiomics,’ and ‘natural language processing,’ combined with disease‐specific terminology such as ‘rheumatoid arthritis,’ ‘vasculitis,’ ‘spondyloarthritis,’ ‘systemic sclerosis,’ ‘psoriatic arthritis,’ ‘connective tissue diseases,’ and imaging‐related terms including ‘ultrasound,’ ‘magnetic resonance imaging,’ ‘computed tomography,’ and ‘capillaroscopy.’ Study selection was conducted using a sequential screening process involving title review, abstract assessment and full‐text evaluation. Two independent reviewers screened all records for eligibility. Discrepancies were resolved through discussion and consensus. When agreement could not be reached, a third reviewer adjudicated the decision. Studies were eligible for inclusion if they investigated the application of AI techniques within rheumatology and reported original quantitative data derived from clinical cohorts, imaging datasets, registries, or electronic health records. Eligible studies were required to report objective performance metrics, accuracy, sensitivity, specificity, precision and recall. Studies were excluded if they lacked complete methodological description or had insufficient methodological transparency, preventing quality appraisal. Two reviewers independently performed data extraction to ensure the accuracy and reproducibility of collected information. Figures 1 and 2.

FIGURE 1.

FIGURE 1

Study selection process according to PRISMA 2020 guidelines. Records were identified through database searching (PubMed, Scopus, and Web of Science; n = 2047) and manual reference screening (n = 47). After duplicate removal, records were screened by title and abstract, and 75 studies were included in the qualitative synthesis.

FIGURE 2.

FIGURE 2

Eligibility criteria for study selection. Inclusion criteria comprised peer‐reviewed, English‐language studies (2020–2025) investigating artificial intelligence applications in rheumatology with original quantitative data and reported performance metrics. Studies lacking methodological transparency, clinical applicability, or quantitative outcomes were excluded. Discrepancies were resolved by consensus, and unresolved cases were adjudicated by a third reviewer.

4. What's the Matter?

A variety of ML methodologies have been widely adopted in the medical domain.

4.1. Linear Regression

Linear regression, one of the most fundamental ML models, approximates the relationship between one or more numerical inputs and a continuous output using a linear equation; in multivariate regression each input feature carries a coefficient reflecting its contribution to the output [2], and the discrepancy between predicted and actual values is captured by residuals.

4.2. Logistic Regression

Despite its name, logistic regression is used primarily for classification, estimating the probability of a categorical outcome from one or more input features [2]. Rather than a linear function, it applies the sigmoid (logistic) function, mapping inputs onto an S‐shaped curve between 0 and 1; this probabilistic output suits binary classification and can be extended to multiclass problems.

4.3. Decision Trees and Random Forests

Decision trees are a supervised method used for classification but also for regression [2]. From a root node, the data are recursively partitioned by the feature that best separates the classes, producing a hierarchy of decision and terminal (leaf) nodes that maximizes within‐subset homogeneity and assigns the final label. Random forests extend this idea through ensemble learning: many trees built on different feature subsets and samples vote on the prediction, improving generalizability and reducing overfitting.

4.4. Artificial Neural Networks (ANNs)

Artificial Neural Networks (ANNs) are inspired by the architecture of biological neural systems [3]. Interconnected nodes (analogous to neurons) transmit signals through weighted connections whose strengths are adjusted, following Hebbian principles, according to the correlation of node activations to improve predictive accuracy.

4.5. Convolutional Neural Networks

For image data, Convolutional Neural Networks (CNNs) are a specialized class of ANNs that preserve spatial relationships between pixels [3]. Rather than treating each pixel independently, CNNs apply localized filters (convolutional kernels) across small patches of the image; each filter detects features such as edges or patterns through element‐wise matrix multiplication, producing a feature map [4]. Successive convolutional layers extract increasingly abstract features, which are then passed to a feedforward ANN for final classification [3]; this multi‐layered process is a hallmark of DL [2].

5. From Diagnosis to Treatment Optimization

5.1. Arthritis

The diagnosis of RA continues to pose significant challenges, despite advancements in clinical methodologies. Symptoms such as arthralgia, morning stiffness and joint swelling frequently overlap with those of other musculoskeletal conditions, complicating differential diagnosis. Current diagnostic strategies typically combine clinical evaluation, serological assays and imaging modalities such as radiographs and ultrasonography. Nevertheless, these tools are not without limitations. Serological tests may yield false‐negative results in the early disease phase, although imaging findings do not always show the state of inflammation [1, 5]. Furthermore, variability in clinical experience and the subjective interpretation of imaging contribute to inconsistencies in diagnostic accuracy. Despite updated diagnostic frameworks, delays and misdiagnoses remain prevalent, particularly in atypical or early‐stage presentations. The management of RA necessitates continuous monitoring and dynamic treatment modifications in response to disease activity or adverse effects [6]. These persistent challenges have spurred growing interest in AI to improve the speed and accuracy of diagnostic and therapeutic decisions. Early diagnosis refers to the identification of a pathological condition at its initial phase before the manifestation of advanced symptoms or irreversible sequelae. In RA, early detection is critical for preventing structural joint damage and functional decline [7]. Research has shown that ML algorithms can synthesize clinical data, including patient‐reported outcomes, serological tests and inflammatory biomarkers, to assess the likelihood of RA development in individuals with undifferentiated arthritis [8]. In RA, continuous monitoring is essential to optimize therapeutic efficacy and prevent complications. AI‐enabled digital health platforms, including smartphone applications and wearable technologies, allow for real‐time data acquisition on symptomatology, physical activity and inflammatory status [9]. For example, wearable devices equipped with biosensors can measure joint mobility and capture self‐reported pain metrics. AI algorithms analyzed data streams to generate automated, individualized disease activity reports that can be transmitted to healthcare providers [10]. Moreover, digital tools such as chatbots and virtual health assistants can enhance patient self‐management by supporting adherence to treatment regimens and facilitating symptom tracking [11, 12]. O'Neil et al. used ML to detect proteomic signatures linked to clinical remission and imminent flares in RA patients. Within the RETRO cohort of 130 individuals, 1307 serum proteins were quantified at baseline. Although clustering alone did not accurately predict flare, an XGBoost model trained on the baseline proteomic data achieved an AUC of 0.80 for predicting relapse 12 months post‐medication withdrawal [13]. Meanwhile, deep learning (DL) algorithms applied to EHRs produced moderate performance (AUC = 0.7) in predicting flare severity, highlighting challenges posed by data inconsistencies across electronic records [14]. Selecting optimal RA treatments remains complex due to the extensive range of therapeutic options, including glucocorticoids, csDMARDs, tsDMARDs and bDMARDs, as well as heterogeneous patient responses and a lack of robust comparative data. AI tools hold promises for improving the prediction of both therapeutic efficacy and adverse event risks. Surendran et al. used ML to predict methotrexate‐induced elevation of liver enzyme levels from EHR data, achieving high accuracy (F1 = 0.87), with baseline high‐normal transaminase levels and elevated lymphocyte/neutrophil counts being major predictors [15]. A XGBoost classifier predicted 6‐month non‐remission in 222 DMARD‐naïve RA patients initiating methotrexate, using a super learner approach combining elastic net, RnFr and SVM. This model reported AUCs of 0.75–0.76 and sensitivities of 0.77–0.81, with the Rheumatoid Arthritis Impact of Disease score identified as the strongest baseline predictor [16]. Kalweit et al. employed DL to stratify RA patients and predict their b/ts DMARD responses. Their analysis revealed that males, patients on ≥ 2 csDMARDs plus prednisone, and those with lower disease activity responded better to tocilizumab than adalimumab. Furthermore, seronegative females and seropositive females with high disease burden were less likely to respond to golimumab [17]. But a systematic review of 29 studies evaluating ML model performance and transparency in RA treatment response prediction found that adherence to TRIPOD reporting standards was moderate, with most models exhibiting moderate to elevated risk of bias. These findings stress the need for improved calibration, transparency and handling of missing data in RA‐related ML research [18].

5.2. Vasculitides

Vasculitides are a group of systemic, immune‐mediated inflammatory disorders that affect blood vessels and present significant diagnostic complexities due to their heterogeneous clinical manifestations [19]. Early identification and therapeutic intervention are essential to optimize patient prognosis. In this context, AI has emerged as a promising technology to support clinical decision‐making in the management of vasculitides. In fact, prompt diagnosis is vital to minimize irreversible tissue injury and improve clinical outcomes [19]. AI methodologies, including both traditional ML algorithms, such as logistic regression, Random Forest and support vector machines (SVMs) and more advanced DL architectures, such as CNNs. They offer innovative opportunities to extract insights from complex biomedical data and support early detection efforts [20]. More recently, natural language processing (NLP) tools and large language models (LLMs) have demonstrated the capacity to interpret unstructured clinical narratives. AI‐based tools can integrate and analyze multimodal data, including clinical documentation, laboratory parameters and imaging findings, thereby assisting with diagnosis, prognostication and treatment individualization [21]. Although AI has yielded encouraging results in other autoimmune contexts, its application to rare conditions like vasculitis remains underexplored. This underutilization is attributed to their low prevalence, which contributes to diagnostic delays and unfamiliarity among healthcare providers [22]. In diagnostic applications, 13 studies employed ML models to classify various forms of vasculitis. Some of these focused on Kawasaki's Disease (KD): Tsai et al. utilized XGBoost on 74 641 pediatric cases to distinguish KD from other febrile illnesses, achieving 92.5% sensitivity and 97.3% specificity using laboratory variables such as pyuria, urinary white cell count, alanine transaminase and C‐reactive protein [23] and Li et al. employed LASSO regression and an SVM to differentiate KD from sepsis in 608 individuals, attaining an AUC of 0.878 with comparable clinical and biochemical features [24]. Other forms of vasculitides were similarly investigated. Vries et al. employed radiomics extracted from 18F‐FDG PET scans to distinguish GCA from atherosclerosis, reaching 84% classification accuracy [25]. In Behçet disease, Güler et al. analyzed ophthalmic arterial Doppler signals from 106 participants, achieving 96.43% accuracy in healthy controls and 93.75% in patients using a multilayer perceptron [26]. Nie et al. applied XGBoost to distinguish abdominal Henoch–Schönlein purpura from acute appendicitis in 6965 pediatric cases, attaining 82.4% accuracy with 53 laboratory metrics [27]. Broader‐spectrum studies were also conducted; Ryyppö et al. implemented ML models to identify vasculitis, myositis and glomerulonephritis across 114 897 individuals, including 2919 vasculitis cases, achieving a true positive rate of 76.7% and true negative rate of 98.4% via XGBoost [28]. In forecasting complications in GCA, Venerito et al. used Random Forest to identify flares after glucocorticoid tapering in 107 GCA patients, achieving 71.4% accuracy and an AUC of 0.76, surpassing logistic regression and decision trees [29]. For Behçet‐related vision complications, Hammam et al. used XGBoost on 1094 patients, achieving an AUC of 0.85, 85% sensitivity and 86% specificity [30].

5.2.1. AI in Dermatological Manifestations

5.2.1.1. Psoriasis and Automated Assessment

Over the past two decades, research has focused on developing AI platforms to alleviate the clinical burden of psoriasis assessment. Significant advancements have occurred in integrating image segmentation with DL architecture. A 2019 investigation demonstrated that CNN, combined with automated lesion segmentation, can accurately distinguish psoriatic from non‐psoriatic skin, achieving 94.8% accuracy even with confounding features such as hair or ill‐defined lesion borders [31].

5.2.1.2. Systemic Sclerosis Applications

In Systemic Sclerosis (SSc), AI has been deployed to enhance diagnostic stratification. Early applications involved neural network‐based algorithms to quantify cutaneous thickening in morphea, improving diagnostic sensitivity [32]. AI‐driven tools show promise for resource‐constrained environments, with pilot projects training CNNs on smartphone‐derived facial images to identify phenotypic features indicative of SSc, supporting early referral and triage by non‐specialists [33].

5.2.2. Radiological Applications

5.2.2.1. Conventional Radiography

Conventional hand radiographs remain fundamental for identifying radiographic abnormalities in rheumatology. Artificial neural networks (ANNs) and DL approaches have shown promise in replicating expert‐level scoring while reducing the use of manual interpretation. Izumi et al. designed a deep neural network model to detect wrist subluxation and ankylosis as part of an automated modified total Sharp score (mTSS) evaluation system, achieving a high AUC of 97% despite training on a limited dataset [34]. Radke et al. enhanced erosion detection by training RetinaNet models using adaptively adjusted Intersection over Union (IoU) thresholds, achieving 94% classification accuracy and mAP of 0.81, markedly outperforming static IoU models [35]. The RA2‐DREAM Challenge facilitated the development of ML models capable of reliably quantifying joint damage based on radiographs, achieving high concordance with expert‐assessed SvH scores (ranging from 0.71 to 0.82) [36]. Venäläinen et al. introduced and externally validated the AuRA (Automated RA Scoring Algorithm), developed during the same challenge. Trained on 367 radiographs and tested on an independent cohort of 205, AuRA yielded superior predictive accuracy compared to the two top RA2‐DREAM models and demonstrated robust longitudinal correlation with expert scores [37].

5.2.2.2. Computed Tomography

AI‐driven classification and quantification methods in CT imaging are expected to outperform traditional visual assessment by improving reproducibility and enabling quantification of subtle lesions [38]. Recent years have witnessed increasing AI implementation in interstitial lung diseases (ILDs), with substantial literature addressing CT‐based diagnosis [39, 40], quantification of parenchymal abnormalities [41] and prognostication via quantitative CT metrics [42]. These imaging biomarkers are pivotal for tracking disease progression [43]. Connective tissue diseases (CTDs) are systemic autoimmune disorders characterized by aberrant immune activation, leading to chronic inflammation and damage to the body's connective tissues. This group includes conditions such as RA, systemic lupus erythematosus (SLE), Sjögren's syndrome (SS), idiopathic inflammatory myopathies, namely dermatomyositis (DM) and polymyositis (PM), SSc and mixed connective tissue disease (MCTD) [44]. Walsh et al. introduced SOFIA, a DL tool for identifying UIP patterns on CT. Using CT annotations from multiple radiologists as reference, they compared SOFIA's agreement with 91 radiologists, reporting concordance rates of 73.3% for AI and 70.7% for human readers, suggesting AI parity with expert interpretation [45]. Given that ILD diagnosis is grounded in peripheral lung histopathology, attempts have been made to analyze corresponding CT regions (i.e., virtual wedge resections); Shaish et al. used HRCT images from 221 ILD cases, extracting approximately 500 virtual wedges per patient. Using a 16.5% threshold of wedges classified as UIP yielded 74% sensitivity and 58% specificity, highlighting potential diagnostic variability based on disease extent [46]. Some ILDs exhibit progressive fibrotic behavior and are referred to as progressive fibrosing ILD, progressive phenotype, or progressive pulmonary fibrosis [47]. The efficacy of antifibrotics has been demonstrated in trials that included fibrosis > 10% on CT and radiologic progression over time [48], necessitating accurate, reproducible fibrosis quantification. Given the limitations of visual HRCT assessment, such as inter‐reader variability, various AI‐based quantification tools using texture analysis, support vector machines and DL have emerged [41] Limitations inherent to AI include its task‐specific nature‐models cannot generalize beyond their training. Online AI systems may produce inaccurate results due to flawed input data. The ‘black box’ issue difficulty in interpreting AI decision logic, complicates error tracing and resolution. Interpretability thus remains a critical hurdle for AI integration into clinical workflows [1]. Presently, physicians retain ultimate responsibility for AI‐assisted diagnoses, though this paradigm may evolve as technologies mature.

5.2.2.3. Ultrasonography

Ultrasound (US), recognized as an adaptable and environmentally sustainable imaging modality, has become an essential diagnostic tool in clinical settings. Its distinctive benefits, including the absence of ionizing radiation, portability, affordability and the ability to acquire real‐time images, have made it increasingly favored in routine practice. In particular, the continuous advancement of ultrasonographic technologies, such as high‐frequency ultrasound, has significantly broadened their clinical applications. As a result, US is now widely adopted as one of the primary imaging modalities in rheumatology [49]. The enhanced soft‐tissue resolution of US has led to its increasing utility in evaluating the musculoskeletal system. DL‐based ultrasound approaches have been applied to a variety of tasks, including the diagnosis of myopathies [50] and the segmentation of muscular structures [51]. For instance, Burlina et al. introduced a CNN‐based framework to classify inflammatory muscle disorders, yielding improved diagnostic performance for neuromuscular conditions [50]. When compared with traditional ML techniques, their CNN model demonstrated superior accuracy in classifying myositis subtypes: 76.2% ± 3.1% versus 72.3% ± 3.3% (normal vs. affected including inclusion body myositis, dermatomyositis and polymyositis); 86.6% ± 2.4% versus 84.3% ± 2.3% (normal vs. inclusion body myositis); and 74.8% ± 3.9% versus 68.9% ± 2.5% (inclusion body myositis vs. dermatomyositis/polymyositis) [50]. Temporal artery ultrasound (TAUS) has emerged as a promising technique for addressing this issue, with the European Alliance of Associations for Rheumatology recommending it as the preferred initial imaging modality for evaluating suspected GCA cases [52]. Although TAUS demonstrates reasonable diagnostic performance, its clinical utility is limited by only moderate interobserver reliability and the need for substantial operator training—factors that hinder its broader implementation [53]. In this context, DL methods offer a compelling opportunity to enhance diagnostic consistency and potentially reduce the learning curve associated with TAUS interpretation. Roncato et al. introduced a CNN approach specifically designed to detect the halo sign [54]. The initial phase of their model involved semantic segmentation of both transverse and longitudinal views from color Doppler and power Doppler ultrasound scans of temporal arteries. Regions of interest were delineated around arterial segments and adjacent soft tissues, with each pixel annotated as either halo‐positive or halo‐negative. To implement this segmentation task, the researchers employed the U‐Net architecture, a CNN model specifically optimized for biomedical image segmentation tasks [54]. Final image classification (positive or negative for the halo sign) was determined based on the proportion of pixels within the predefined region labeled as halo positive. A higher density of such pixels correlated with an increased likelihood of a true halo sign being present [54]. The model's performance was evaluated on two datasets: the first consisting of images acquired by a single operator using a standardized imaging protocol, and the second comprising images collected by multiple operators employing diverse acquisition parameters. The CNN demonstrated superior accuracy on the standardized dataset (AUC 0.95) compared to the more heterogeneous dataset (AUC 0.82) [54]. These findings suggest that although training the model on a wider range of image qualities and acquisition techniques may improve its robustness, DL can also play a pivotal role in guiding real‐time image acquisition. Such computer‐assisted acquisition systems could help clinicians and sonographers obtain standardized and diagnostically valuable views during live ultrasound examinations. In power Doppler ultrasound imaging, the number of existing automated techniques for quantifying synovitis remains limited when compared to other imaging modalities. This is attributable to the inherent technical complexity of interpreting ultrasound images, primarily due to the high noise levels. A method combining traditional image processing approaches and ML has been proposed for the automated identification of anatomical structures such as skin, bone and synovial tissue, with the latter quantified based on region area [55]. More recently, DL models have been introduced to quantify synovitis across the entire ultrasound image. These models typically involve training CNNs using visual scoring systems as ground truth references [56].

5.3. Magnetic Resonance

MRI has emerged as a crucial non‐invasive imaging modality, offering unparalleled visualization of anatomical and pathological alterations associated with rheumatic diseases. It serves an essential function in the diagnosis and clinical management of conditions such as RA and spondyloarthritis (SpA), among others. MRI facilitates evaluation of joint and periarticular soft tissue involvement, inflammation and structural damage without the need for invasive procedures. However, interpreting MRI images remains labor‐intensive, time‐consuming and subject to variability both within and between observers. Furthermore, the high operational costs and limited access to MRI scanners impede their widespread use in routine rheumatological practice [57]. Advancements in AI present significant opportunities to revolutionize rheumatology by enhancing and automating various aspects of MRI analysis [58]. SpA encompasses a spectrum of inflammatory disorders affecting the axial and peripheral skeleton. Early identification, classification and disease monitoring are imperative to prevent irreversible joint damage and functional decline. AI applications in SpA management show considerable promise, spanning early detection, disease characterization and longitudinal tracking [59]. Zheng et al. developed a DL model capable of detecting multifocal inflammatory lesions in the hips of SpA patients on MRI, achieving diagnostic performance on par with expert radiologists [60]. Bressem et al. validated a DL‐based diagnostic tool in a multicenter retrospective cohort of 593 patients, achieving AUCs of 0.94 for inflammatory lesions, 0.88 for ASAS‐defined inflammatory features, and 0.89 for structural changes. Sensitivity and specificity ranged from 71% to 88%, demonstrating generalizability across imaging protocols and scanner types [61]. Collectively, DL architectures have demonstrated robust utility in SpA management, frequently matching or exceeding the diagnostic accuracy of clinical experts. Schlereth et al. designed and validated an automated CNN‐based scoring system to evaluate bone erosions, osteitis and synovitis in hand MRI scans of patients with RA and psoriatic arthritis (PsA) [62]. The model was trained on MRI data from 211 patients and externally validated on an independent cohort of 220 patients. It demonstrated robust diagnostic performance, achieving macro‐level AUCs of 92% for erosions, 91% for osteitis and 85% for synovitis, along with Spearman correlation coefficients of 90%, 78% and 69%, respectively, when compared with expert annotations [62]. Folle et al. applied ResNet‐based deep neural networks to distinguish inflammatory MRI patterns. among seropositive RA, seronegative RA and PsA, reporting area under the receiver operating characteristic curves (AUROCs) of 75%, 74% and 67%, respectively. The inclusion or exclusion of contrast‐enhanced sequences and clinical metadata had minimal impact on classification performance. Interestingly, when this model was tested on psoriasis patients without clinical arthritis, it frequently classified them as PsA, implying the potential to detect early PsA‐like MRI signatures before clinical manifestation [63].

5.4. Capillaroscopy

Nailfold capillaroscopy (NFC) is a key diagnostic tool in SSc, enabling early detection of microvascular abnormalities characteristic of the capillaroscopic scleroderma pattern [61]. Structural changes, including capillary enlargement, giant capillaries, capillary loss and microhemorrhages, reflect progressive microangiopathy and determine the established early, active and late NFC stages [62]. NFC abnormalities correlate with disease activity, severity and prognosis, underscoring the importance of accurate image evaluation [63]. However, assessment remains limited by operator dependence, variability and time‐consuming analysis [64]. To address these challenges, automated NFC image analysis has gained momentum. Systems such as AUTOCAPI enable fully automated evaluation within seconds, requiring no pre‐processing [65], and supporting more consistent clinical and research assessments [64]. ML‐based approaches have further advanced the field: the Manchester system uses a random forest model to extract quantitative vascular metrics within minutes [66], while state‐of‐the‐art DL architectures such as Vision Transformers (ViT) [67] demonstrate excellent performance in detecting giant capillaries, microhemorrhages, capillary loss and capillary enlargement [68]. Notably, ViT models show performance comparable to human experts while offering far greater consistency and rapid processing (0.19 s per image) [68]. These tools hold promise for broad clinical deployment and may extend beyond SSc to other diseases with microvascular involvement [69].

5.5. Human Factor

Clinicians, as the primary users of AI technologies, bring with them a range of personal experiences, cognitive styles, decision‐making habits and varying levels of clinical expertise. These human elements introduce the potential for cognitive biases and reasoning errors, which must be considered during the development and deployment of AI‐based diagnostic tools [64]. The dual process theory of cognition [28, 29] offers a useful framework for understanding how such biases arise [65]. This theory distinguishes between two modes of thinking: System 1, which is fast, automatic and intuitive, and System 2, which is deliberate, slower and analytical [66]. While each system serves specific cognitive functions, System 2 can, to a certain extent, oversee and regulate the output of System 1 [66]. Errors in clinical reasoning may occur during either type of cognitive processing [66]. However, it has been argued that many cognitive failures stem from excessive reliance on System 1 and an under‐engaged or inactive System 2 [66]. Common biases linked to intuitive (System 1) reasoning include confirmation bias, anchoring, overconfidence, availability bias and framing effects [66]. To evaluate an individual's capacity to override intuitive but incorrect responses, the Cognitive Reflection Test (CRT) was developed [67]. This three‐question assessment contrasts immediate, instinctive answers with correct responses that require analytical reasoning. Individuals who score highly on the CRT are believed to rely more on System 2 processing and to scrutinize intuitive judgments more critically [66]. Brennan et al. explored how physicians interacted with an ML model designed to assess preoperative risk and employed the CRT to characterize participants' cognitive styles [68]. The algorithm outperformed clinicians in accuracy, and physicians' decision quality improved following interaction with the model. Although not statistically significant, results suggested that those with higher CRT scores (indicating a more analytic style) were more likely to revise their decisions based on the algorithm's input than more intuitive clinicians [68]. Cognitive errors are considered the leading cause of diagnostic mistakes in medicine [69]. Many of these errors are attributed to flawed reasoning or biased judgment [69]. Over 30 cognitive biases and heuristics have been implicated in medical errors [48]. AI‐based diagnostic systems may interact with human cognitive biases in many ways [70]. Some biases could be amplified; others diminished or remain unaffected depending on how the AI is employed and on the users' cognitive style. For instance, if an AI offers a diagnosis that differs from the clinician's own, it may prompt more analytical reflection [70]. Conversely, if the AI's suggestion precedes the clinician's assessment, it could result in anchoring or inhibit reconsideration of an erroneous diagnosis. If the AI confirms a pre‐existing opinion, the clinician may fall victim to confirmation bias or premature diagnostic closure. Alternatively, some clinicians may disregard AI outputs altogether [70]. Nonetheless, AI introduces its own risks, notably automation bias, the undue reliance on automated systems, which may lead to acceptance of incorrect outputs and neglect of contradictory evidence [64].

6. Limitations

ML in medicine faces several technical limitations that affect model performance, arising from algorithmic constraints to issues related to data quality and availability [3]. A major challenge is data scarcity, particularly in rare autoimmune rheumatic diseases, where low prevalence restricts the development of statistically robust models; multicenter collaborations and data‐sharing platforms offer potential solutions to this limitation [3]. Another key concern is the lack of representativeness within training datasets, as models trained on non‐diverse populations often fail to generalize across heterogeneous clinical settings. Ensuring demographic and clinical diversity is therefore essential, while institutional differences in clinical practice further compromise external validity [71]. Additionally, real‐world data, unlike rigorously curated clinical trial datasets, are characterized by variability, structural heterogeneity, noise and missing values, all of which increase the risk of bias and reduce the reliability of ML outputs [72]. Crucially, the external validity and generalizability of most published models remain insufficiently established. The majority of algorithms are developed and tested within single‐center cohorts, and their reported accuracy frequently deteriorates when applied to external populations because of distribution shift arising from differences in scanner hardware, acquisition and reconstruction protocols, patient case‐mix, disease prevalence and local labeling conventions. Robust generalizability therefore cannot be inferred from internal cross‐validation alone: it requires prospective, multicenter external validation on geographically and demographically distinct cohorts, ideally followed by prospective evaluation within the intended clinical workflow. Adherence to established reporting and appraisal frameworks, such as TRIPOD‐AI for the reporting of prediction models and PROBAST for risk‐of‐bias assessment, would improve transparency, temper the optimistic performance estimates that characterize much of the current literature, and enable meaningful comparison across studies [73, 74].

7. Expert Considerations

AI development can be likened to medical training: high‐quality data and structured guidance are essential to achieve reliable performance, particularly in rare fields such as rheumatology. Establishing centralized, standardized data‐sharing infrastructures would improve algorithm robustness, diagnostic accuracy, clinical management and trial design. However, beyond data availability, adequate development time is crucial; rapid creation of high‐performing, bias‐free AI systems is unrealistic (Table 1).

TABLE 1.

Summary of key AI applications in rheumatology.

Clinical domain Study AI method Application Cohort size Performance metrics
Rheumatoid arthritis
Flare prediction O'Neil et al. [10] XGBoost Predict relapse 12 months after medication withdrawal using proteomic signatures 130 patients AUC 0.80
Treatment adverse events Surendran et al. [12] Machine learning Predict methotrexate‐induced liver enzyme elevation using EHR data 569 patients F1 score 0.87
Treatment response G. Li et al. [13] Super learner (Elastic Net, Random Forest, SVM) Predict 6‐month non‐remission in DMARD‐naïve patients 222 patients AUC 0.75–0.76; Sensitivity 0.77–0.81
Treatment stratification Kalweit et al. [14] Deep learning Predict b/tsDMARD response and patient stratification 3516 patients Stratification based on demographics and disease activity
Vasculitis
Kawasaki disease Tsai et al. [20] XGBoost Differentiate Kawasaki disease from febrile illnesses 74 641 pediatric cases Sensitivity 92.5%; Specificity 97.3%
Kawasaki disease Li et al. [21] LASSO + SVM Differentiate Kawasaki disease from sepsis 608 patients AUC 0.878
Giant cell arteritis Vries et al. [22] PET radiomics Distinguish GCA from atherosclerosis 238 patients Accuracy 84%
Behçet disease Güler et al. [23] Multilayer perceptron Analyze ophthalmic arterial Doppler signals 106 patients Accuracy 96.43%
Henoch–Schönlein purpura Nie et al. [24] XGBoost Differentiate abdominal HSP from acute appendicitis 6965 pediatric cases Accuracy 82.4%
Multi‐vasculitis classification Ryyppö et al. [25] XGBoost Identify vasculitis, myositis and glomerulonephritis 114 897 patients (2919 vasculitis cases) TPR 76.7%; TNR 98.4%
GCA complications Venerito et al. [26] Random forest Predict flares after glucocorticoid tapering 107 GCA patients Accuracy 71.4%; AUC 0.76

Equally important is the recognition that AI outputs cannot be interpreted in isolation from clinical judgment. In rheumatology, diagnostic and classification criteria, definitions of disease activity and treatment recommendations are not entirely value‐neutral or purely objective; they derive, at least in part, from expert consensus and collective clinical judgment [73, 75]. Consequently, the reference standards against which AI models are trained and evaluated inherit this interpretive, consensus‐based character and their outputs require critical appraisal rather than uncritical acceptance. The clinician's role remains essential in contextualizing algorithmic predictions within the individual patient's history, comorbidities and preferences, in recognizing when a model is operating outside the data distribution on which it was trained, and in retaining ultimate accountability for the clinical decision. Rather than replacing expert reasoning, AI is therefore best positioned as a decision‐support adjunct that augments, but does not supplant, human clinical judgment, thereby guarding against automation bias while preserving the interpretive flexibility that consensus‐based rheumatologic practice demands [75].

8. Conclusion

Artificial intelligence is progressively transforming rheumatology by enhancing diagnostic accuracy, supporting prognostic stratification and enabling personalized therapeutic strategies. Evidence demonstrates robust performance in imaging analysis and predictive modeling; however, translation into routine clinical practice remains limited. Robust external validation, standardized reporting, transparent model development and integration within clinical workflows are essential prerequisites for sustainable implementation. Future research should prioritize prospective validation studies and the evaluation of patient‐centered clinical outcomes to establish the AI's impact in rheumatology. These tools should be deployed as decision‐support adjuncts that assist rather than replace expert clinical judgment, which remains indispensable for the critical appraisal, contextualization and accountable use of AI‐generated outputs [73, 75].

Author Contributions

Conceptualization: R.B., Methodology: A.C., writing – original draft: R.B., writing – review and editing: C.R., G.G., N.M., supervision: F.P.C., A.C. All authors have read and agreed to the published version of the manuscript.

Funding

The authors have nothing to report.

Conflicts of Interest

The authors declare no conflicts of interest.

Data Availability Statement

The authors have nothing to report.

References

  • 1. Vlad A. L., Popazu C., Lescai A. M., Voinescu D. C., and Baltă A. A. Ș., “The Role of Artificial Intelligence in the Diagnosis and Management of Rheumatoid Arthritis,” Medicina 61, no. 4 (2025): 689, 10.3390/medicina61040689. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Chen C. W., Tsai H. H., Yeh C. Y., et al., “Application of Artificial Intelligence in Rheumatic Disease Classification: An Example of Ankylosing Spondylitis Severity Inspection Model,” Annals of Medicine 57, no. 1 (2025): 2512131, 10.1080/07853890.2025.2512131. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Géron A., Hands‐on Machine Learning With Scikit‐Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems (O'Reilly Media, Inc., 2019). [Google Scholar]
  • 4. Zhou T., Cheng Q. R., Lu H. L., Li Q., Zhang X. X., and Qiu S., “Deep Learning Methods for Medical Image Fusion: A Review,” Computers in Biology and Medicine 160 (2023): 106959, 10.1016/j.compbiomed.2023.106959. [DOI] [PubMed] [Google Scholar]
  • 5. Walter P., Kau C. K., and Chen C. W., “A Platform‐Based Approach to Assisting Rheumatoid Arthritis Management,” Journal of Rheumatic Diseases 26, no. 6 (2023): 1019, 10.1111/1756-185X.14688. [DOI] [PubMed] [Google Scholar]
  • 6. Brown P., Pratt A. G., and Hyrich K. L., “Therapeutic Advances in Rheumatoid Arthritis,” BMJ 384 (2024), 10.1136/bmj-2022-070856. [DOI] [PubMed] [Google Scholar]
  • 7. Pavlov‐Dolijanovic S., Bogojevic M., Nozica‐Radulovic T., Radunovic G., and Mujovic N., “Elderly‐Onset Rheumatoid Arthritis: Characteristics and Treatment Options,” Medicina 59 (2023): 1878, 10.3390/medicina59101878. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Madrid‐García A., Merino‐Barbancho B., Rodríguez‐González A., Fernández‐Gutiérrez B., Rodríguez‐Rodríguez L., and Menasalvas‐Ruiz E., Understanding the Role and Adoption of Artificial Intelligence Techniques in Rheumatology Research: An In‐Depth Review of the Literature (W.B. Saunders, 2023), 10.1016/j.semarthrit.2023.152213. [DOI] [PubMed] [Google Scholar]
  • 9. Knitza J., Gupta L., and Hügle T., “Rheumatology in the Digital Health Era: Status Quo and Quo Vadis?,” Nature Reviews 20 (2024): 747–759, 10.1038/s41584-024-01177-7. [DOI] [PubMed] [Google Scholar]
  • 10. Creagh A. P., Hamy V., Yuan H., et al., “Digital Health Technologies and Machine Learning Augment Patient Reported Outcomes to Remotely Characterise Rheumatoid Arthritis,” npj Digital Medicine 7, no. 1 (2024): 33, 10.1038/s41746-024-01013-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Jung S., Kim S., Yoo B. Y., Park M., and Yoon H. J., “Development of a Voice‐Interactive Chatbot for Personalized Health Management Using a Lifestyle Recommendation Knowledge Base,” Studies in Health Technology and Informatics 327 (2025): 1017–1018, 10.3233/SHTI250534. [DOI] [PubMed] [Google Scholar]
  • 12. Chen C. W., Walter P., and Wei J. C.‐C., “Using ChatGPT‐Like Solutions to Bridge the Communication Gap Between Patients With Rheumatoid Arthritis and Health Care Professionals,” JMIR Medical Education 10 (2024): e48989, 10.2196/48989. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. O'Neil L. J., Hu P., Liu Q., et al., “Proteomic Approaches to Defining Remission and the Risk of Relapse in Rheumatoid Arthritis,” Frontiers in Immunology 12 (2021): 729681, 10.3389/fimmu.2021.729681. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14. Chandran U., Reps J., Stang P. E., and Ryan P. B., “Inferring Disease Severity in Rheumatoid Arthritis Using Predictive Modeling in Administrative Claims Databases,” PLoS One 14, no. 12 (2019): e0226255, 10.1371/journal.pone.0226255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Surendran S., Gilvaz V., Manyam P. K., Panicker K., Pradeep M., and Manyam P. K., “Prediction of Liver Enzyme Elevation Using Supervised Machine Learning in Patients With Rheumatoid Arthritis on Treatment With Methotrexate,” Cureus 16 (2024): 1, 10.7759/cureus.52110. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Li G., Kolan S. S., Grimolizzi F., et al., “Development of Machine Learning Models for Predicting Non‐Remission in Early RA Highlights the Robust Predictive Importance of the RAID Score‐Evidence From the ARCTIC Study,” Frontiers in Medicine (Lausanne) 12 (2025): 1526708, 10.3389/fmed.2025.1526708. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Kalweit M., Burden A. M., Boedecker J., Hügle T., and Burkard T., “Patient Groups in Rheumatoid Arthritis Identified by Deep Learning Respond Differently to Biologic or Targeted Synthetic DMARDs,” PLoS Computational Biology 19, no. 6 June (2023): e1011073, 10.1371/journal.pcbi.1011073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Mendoza‐Pinto C., Sánchez‐Tecuatl M., Berra‐Romani R., et al., “Machine Learning in the Prediction of Treatment Response in Rheumatoid Arthritis: A Systematic Review,” Seminars in Arthritis and Rheumatism 68 (2024): 152501, 10.1016/j.semarthrit.2024.152501. [DOI] [PubMed] [Google Scholar]
  • 19. Marvisi C., Delvino P., Di Cianni F., et al., “Systemic Vasculitis: One Year in Review 2026,” Clinical and Experimental Rheumatology 44, no. 4 (2026): 593–606, 10.55563/clinexprheumatol/bdck2x. [DOI] [PubMed] [Google Scholar]
  • 20. Hügle M., Omoumi P., van Laar J. M., Boedecker J., and Hügle T., “Applied Machine Learning and Artificial Intelligence in Rheumatology,” Rheumatology Advances in Practice 4 (2021): 1, 10.1093/rap/rkaa005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Alowais S. A., Alghamdi S. S., Alsuhebany N., et al., “Revolutionizing Healthcare: The Role of Artificial Intelligence in Clinical Practice,” BMC Medical Education 23 (2023): 689, 10.1186/s12909-023-04698-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Domaradzki J. and Walkowiak D., “Knowledge and Attitudes of Future Healthcare Professionals Toward Rare Diseases,” Frontiers in Genetics 12 (2021): 639610, 10.3389/fgene.2021.639610. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Tsai C. M., Lin C. H. R., Kuo H. C., et al., “Use of Machine Learning to Differentiate Children With Kawasaki Disease From Other Febrile Children in a Pediatric Emergency Department,” JAMA Network Open 6, no. 4 (2023): e237489, 10.1001/jamanetworkopen.2023.7489. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Li C., Liu Y. C., Zhang D. R., et al., “A Machine Learning Model for Distinguishing Kawasaki Disease From Sepsis,” Scientific Reports 13, no. 1 (2023): 12553, 10.1038/s41598-023-39745-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Vries H. S., van Praagh G. D., Nienhuis P. H., Alic L., and Slart R. H. J. A., “A Machine Learning Model Based on Radiomic Features as a Tool to Identify Active Giant Cell Arteritis on [18F]FDG‐PET Images During Follow‐Up,” Diagnostics 15, no. 3 (2025): 367, 10.3390/diagnostics15030367. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Güler I. and Übeyli E. D., “Detection of Ophthalmic Arterial Doppler Signals With Behcet Disease Using Multilayer Perceptron Neural Network,” Computers in Biology and Medicine 35, no. 2 (2005): 121–132, 10.1016/j.compbiomed.2003.12.007. [DOI] [PubMed] [Google Scholar]
  • 27. Nie D., Zhan Y., Xu K., et al., “Artificial Intelligence Differentiates Abdominal Henoch‐Schönlein Purpura From Acute Appendicitis in Children,” International Journal of Rheumatic Diseases 26, no. 12 (2023): 2534–2542, 10.1111/1756-185X.14956. [DOI] [PubMed] [Google Scholar]
  • 28. Ryyppö R., Häyrynen S., Joutsijoki H., Juhola M., and Seppänen M., “Comparison of Machine Learning Methods in the Early Identification of Vasculitides, Myositides and Glomerulonephritides,” Computer Methods and Programs in Biomedicine 243 (2024): 107917, 10.1016/j.cmpb.2023.107917. [DOI] [PubMed] [Google Scholar]
  • 29. Venerito V., Emmi G., Cantarini L., et al., “Validity of Machine Learning in Predicting Giant Cell Arteritis Flare After Glucocorticoids Tapering,” Frontiers in Immunology 13 (2022): 860877, 10.3389/fimmu.2022.860877. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Hammam N., Bakhiet A., el‐Latif E. A., et al., “Development of Machine Learning Models for Detection of Vision Threatening Behçet's Disease (BD) Using Egyptian College of Rheumatology (ECR)–BD Cohort,” BMC Medical Informatics and Decision Making 23, no. 1 (2023): 37, 10.1186/s12911-023-02130-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Gong X., Hu M., Basu M., and Zhao L., “Heterogeneous Treatment Effect Analysis Based on Machine‐Learning Methodology,” CPT: Pharmacometrics & Systems Pharmacology 10, no. 11 (2021): 1433–1443, 10.1002/psp4.12715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Jia K., Li H., Wu X., Xu C., and Xue H., “The Value of High‐Resolution Ultrasound Combined With Shear‐Wave Elastography Under Artificial Intelligence Algorithm in Quantitative Evaluation of Skin Thickness in Localized Scleroderma,” Computational Intelligence and Neuroscience 2022 (2022): 1–9, 10.1155/2022/1613783. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Fryxelius A., “POS0202‐Pare Chronically Active ‐ A Patient Empowerment Program With Peer Support,” Annals of the Rheumatic Diseases 82 (2023): 326–327, 10.1136/annrheumdis-2023-eular.2159. [DOI] [Google Scholar]
  • 34. Izumi K., Suzuki K., Hashimoto M., et al., “Detecting Hand Joint Ankylosis and Subluxation in Radiographic Images Using Deep Learning: A Step in the Development of an Automatic Radiographic Scoring System for Joint Destruction,” PLoS One 18, no. 2 February (2023): e0281088, 10.1371/journal.pone.0281088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Radke K. L., Kors M., Müller‐Lutz A., et al., “Adaptive IoU Thresholding for Improving Small Object Detection: A Proof‐Of‐Concept Study of Hand Erosions Classification of Patients With Rheumatic Arthritis on X‐Ray Images,” Diagnostics 13, no. 1 (2023): 104, 10.3390/diagnostics13010104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Sun D., Nguyen T. M., Allaway R. J., et al., “A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis,” JAMA Network Open 5, no. 8 (2022): e2227423, 10.1001/jamanetworkopen.2022.27423. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Venalainen M. S., Biehl A., Holstila M., Kuusalo L., and Elo L. L., “Deep Learning Enables Automatic Detection of Joint Damage Progression in Rheumatoid Arthritis—Model Development and External Validation,” Rheumatology 64, no. 3 (2025): 1068–1076, 10.1093/rheumatology/keae215. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Uyama M., Handa T., Uozumi R., et al., “Prognostic Value of a Composite Physiologic Index Developed by Adding Bronchial and Hyperlucent Volumes Quantified via Artificial Intelligence Technology,” Respiratory Research 25, no. 1 (2024): 442, 10.1186/s12931-024-03075-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Handa T., “The Potential Role of Artificial Intelligence in the Clinical Practice of Interstitial Lung Disease,” Respiratory Investigation 61 (2023): 702–710, 10.1016/j.resinv.2023.08.006. [DOI] [PubMed] [Google Scholar]
  • 40. Park S., Hwang H. J., Yun J., et al., “Influence of Content‐Based Image Retrieval on the Accuracy and Inter‐Reader Agreement of Usual Interstitial Pneumonia CT Pattern Classification,” European Radiology 35 (2025): 7199–7212, 10.1007/s00330-025-11689-9. [DOI] [PubMed] [Google Scholar]
  • 41. Chassagnon G., Vakalopoulou M., Régent A., et al., “Deep Learning–Based Approach for Automated Assessment of Interstitial Lung Disease in Systemic Sclerosis on Ct Images,” Radiology: Artificial Intelligence 2, no. 4 (2020): 1–10, 10.1148/ryai.2020190006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Chung J. H., Chelala L., Pugashetti J. V., et al., “A Deep Learning‐Based Radiomic Classifier for Usual Interstitial Pneumonia,” Chest 165, no. 2 (2024): 371–380, 10.1016/j.chest.2023.10.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Lee J. S., Kim G. H. J., Ha Y. J., et al., “The Extent and Diverse Trajectories of Longitudinal Changes in Rheumatoid Arthritis Interstitial Lung Diseases Using Quantitative Hrct Scores,” Journal of Clinical Medicine 10, no. 17 (2021): 3812, 10.3390/jcm10173812. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Joy G. M., Arbiv O. A., Wong C. K., et al., “Prevalence, Imaging Patterns and Risk Factors of Interstitial Lung Disease in Connective Tissue Disease: A Systematic Review and Meta‐Analysis,” European Respiratory Review 32 (2023): 220210, 10.1183/16000617.0210-2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Walsh S. L. F., Mackintosh J. A., Calandriello L., et al., “Deep Learning‐Based Outcome Prediction in Progressive Fibrotic Lung Disease Using High‐Resolution Computed Tomography,” American Journal of Respiratory and Critical Care Medicine 206, no. 7 (2022): 883–891, 10.1164/rccm.202112-2684OC. [DOI] [PubMed] [Google Scholar]
  • 46. Walsh S. L. F., Calandriello L., Silva M., and Sverzellati N., “Deep Learning for Classifying Fibrotic Lung Disease on High‐Resolution Computed Tomography: A Case‐Cohort Study,” Lancet Respiratory Medicine 6, no. 11 (2018): 837–845, 10.1016/S2213-2600(18)30286-8. [DOI] [PubMed] [Google Scholar]
  • 47. Rajan S. K., Cottin V., Dhar R., et al., “Progressive Pulmonary Fibrosis: An Expert Group Consensus Statement,” European Respiratory Journal 61, no. 3 (2023): 2103187, 10.1183/13993003.03187-2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48. Vally Z. I., Khammissa R. A. G., Feller G., Ballyram R., Beetge M., and Feller L., “Errors in Clinical Diagnosis: A Narrative Review,” Journal of International Medical Research 51, no. 8 (2023): 03000605231162798, 10.1177/03000605231162798. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Filippou G., Pellegrino M. E., Sorce A., et al., “Updates in Ultrasound in Rheumatology,” Radiologic Clinics of North America 52 (2024): 363–380, 10.1016/j.rcl.2024.02.012. [DOI] [PubMed] [Google Scholar]
  • 50. Burlina P., Billings S., Joshi N., and Albayda J., “Automated Diagnosis of Myositis From Muscle Ultrasound: Exploring the Use of Machine Learning and Deep Learning Methods,” PLoS One 12, no. 8 (2017): e0184059, 10.1371/journal.pone.0184059. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Chen X., Xie C., Chen Z., and Li Q., “Automatic Tracking of Muscle Cross‐Sectional Area Using Convolutional Neural Networks With Ultrasound,” Journal of Ultrasound in Medicine 38, no. 11 (2019): 2901–2908, 10.1002/jum.14995. [DOI] [PubMed] [Google Scholar]
  • 52. Dejaco C., Ramiro S., Bond M., et al., “EULAR Recommendations for the Use of Imaging in Large Vessel Vasculitis in Clinical Practice: 2023 Update,” Annals of the Rheumatic Diseases 83, no. 6 (2024): 741–751, 10.1136/ard-2023-224543. [DOI] [PubMed] [Google Scholar]
  • 53. Luqmani R., Lee E., Singh S., et al., “The Role of Ultrasound Compared to Biopsy of Temporal Arteries in the Diagnosis and Treatment of Giant Cell Arteritis (TABUL): A Diagnostic Accuracy and Cost‐Effectiveness Study,” Health Technology Assessment (Rockville, Md.) 20, no. 90 (2016): 1–270, 10.3310/hta20900. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54. Roncato C., Perez L., Brochet‐Guégan A., et al., “Colour Doppler Ultrasound of Temporal Arteries for the Diagnosis of Giant Cell Arteritis: A Multicentre Deep Learning Study,” Clinical and Experimental Rheumatology 38 (2020): 120–125. [PubMed] [Google Scholar]
  • 55. Cupek R. and Ziebiński A., “Automated Assessment of Joint Synovitis Activity From Medical Ultrasound and Power Doppler Examinations Using Image Processing and Machine Learning Methods,” Reumatologia 54, no. 5 (2016): 239–242, 10.5114/reum.2016.63664. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 56. Andersen J. K. H., Pedersen J. S., Laursen M. S., et al., “Neural Networks for Automatic Scoring of Arthritis Disease Activity on Ultrasound Images,” RMD Open 5, no. 1 (2019): e000891, 10.1136/rmdopen-2018-000891. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 57. Stoel B., “Use of Artificial Intelligence in Imaging in Rheumatology‐Current Status and Future Perspectives,” RMD Open 6 (2020): e001063, 10.1136/rmdopen-2019-001063. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58. Kubassova O., Boesen M., Cimmino M. A., and Bliddal H., “A Computer‐Aided Detection System for Rheumatoid Arthritis MRI Data Interpretation and Quantification of Synovial Activity,” European Journal of Radiology 74, no. 3 (2010): e67–e72, 10.1016/j.ejrad.2009.04.010. [DOI] [PubMed] [Google Scholar]
  • 59. Van Den Berghe T., Babin D., Chen M., et al., “Neural Network Algorithm for Detection of Erosions and Ankylosis on CT of the Sacroiliac Joints: Multicentre Development and Validation of Diagnostic Accuracy,” European Radiology 33, no. 11 (2023): 8310–8323, 10.1007/s00330-023-09704-y. [DOI] [PubMed] [Google Scholar]
  • 60. Zheng Y., Bai C., Zhang K., et al., “Deep‐Learning Based Quantification Model for Hip Bone Marrow Edema and Synovitis in Patients With Spondyloarthritis Based on Magnetic Resonance Images,” Frontiers in Physiology 14 (2023): 1132214, 10.3389/fphys.2023.1132214. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61. Bressem K. K., Adams L. C., Proft F., et al., “Deep Learning Detects Changes Indicative of Axial Spondyloarthritis at MRI of Sacroiliac Joints,” Radiology 305, no. 3 (2022): 655–665, 10.1148/radiol.212526. [DOI] [PubMed] [Google Scholar]
  • 62. Schlereth M., Mutlu M. Y., Utz J., et al., “Deep Learning‐Based Classification of Erosion, Synovitis and Osteitis in Hand MRI of Patients With Inflammatory Arthritis,” RMD Open 10, no. 2 (2024): e004273, 10.1136/rmdopen-2024-004273. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 63. Folle L., Bayat S., Kleyer A., et al., “Advanced Neural Networks for Classification of MRI in Psoriatic Arthritis, Seronegative, and Seropositive Rheumatoid Arthritis,” Rheumatology 61, no. 12 (2022): 4945–4951, 10.1093/rheumatology/keac197. [DOI] [PubMed] [Google Scholar]
  • 64. Sujan M., Furniss D., Grundy K., et al., “Human Factors Challenge for the Safe Use of Artificial Intelligence in Patient Care,” BMJ Health & Care Informatics 26, no. 1 (2019): e100081, 10.1136/bmjhci-2019-100081. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65. Evans J. S., “In Two Minds: Dual‐Process Accounts of Reasoning,” Trends in Cognitive Sciences 7 (2003): 454–459, 10.1016/j.tics.2003.08.012. [DOI] [PubMed] [Google Scholar]
  • 66. Khalil R. and Brüne M., “Adaptive Decision‐Making ‘Fast’ and ‘Slow’: A Model of Creative Thinking,” European Journal of Neuroscience 61 (2025): e70024, 10.1111/ejn.70024. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67. Mertens J. F., Kempen T. G. H., Koster E. S., Deneer V. H. M., Bouvy M. L., and van Gelder T., “Cognitive Processes in Pharmacists' Clinical Decision‐Making,” Research in Social and Administrative Pharmacy 20, no. 2 (2024): 105–114, 10.1016/j.sapharm.2023.10.007. [DOI] [PubMed] [Google Scholar]
  • 68. Brennan M., Puri S., Ozrazgat‐Baslanti T., et al., “Comparing Clinical Judgment With the MySurgeryRisk Algorithm for Preoperative Risk Assessment: A Pilot Usability Study,” Surgery 165, no. 5 (2019): 1035–1045, 10.1016/j.surg.2019.01.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 69. Van Den Berge K. and Mamede S., “Cognitive Diagnostic Error in Internal Medicine,” European Journal of Internal Medicine 24 (2013): 525–529, 10.1016/j.ejim.2013.03.006. [DOI] [PubMed] [Google Scholar]
  • 70. Felmingham C. M., Adler N. R., Ge Z., Morton R. L., Janda M., and Mar V. J., “The Importance of Incorporating Human Factors in the Design and Implementation of Artificial Intelligence for Skin Cancer Diagnosis in the Real World,” American Journal of Clinical Dermatology 22 (2021): 233–242, 10.1007/s40257-020-00574-4. [DOI] [PubMed] [Google Scholar]
  • 71. Futoma J., Simons M., Panch T., Doshi‐Velez F., and Celi L. A., “The Myth of Generalisability in Clinical Research and Machine Learning in Health Care,” Lancet Digital Health 2, no. 9 (2020): e489–e492, 10.1016/S2589-7500(20)30186-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 72. Kim H. S., Lee S., and Kim J. H., “Real‐World Evidence Versus Randomized Controlled Trial: Clinical Research Based on Electronic Medical Records,” Journal of Korean Medical Science 33, no. 34 (2018): e213, 10.3346/jkms.2018.33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 73. Moon J., Jadhav P., and Choi S., “Deep Learning Analysis for Rheumatologic Imaging: Current Trends, Future Directions, and the Role of Human,” Journal of Rheumatic Diseases 32, no. 2 (2025): 73–88, 10.4078/jrd.2024.0128. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 74. Benavent D., Carmona L., García Llorente J. F., et al., “Artificial Intelligence to Predict Treatment Response in Rheumatoid Arthritis and Spondyloarthritis: A Scoping Review,” Rheumatology International 45, no. 4 (2025): 91, 10.1007/s00296-025-05825-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75. Chin‐Yee B. and Upshur R., The Impact of Artificial Intelligence on Clinical Judgment: A Briefing Document (AMS Healthcare, 2019). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The authors have nothing to report.


Articles from International Journal of Rheumatic Diseases are provided here courtesy of Wiley

RESOURCES