Skip to main content
Journal of Rural Medicine : JRM logoLink to Journal of Rural Medicine : JRM
. 2026 Jul 11;21(3):269–278. doi: 10.2185/jrm.2025-079

Communication features of ChatGPT responses and care-seeking intention for syphilis in Japan: a controlled dialogue experiment

Fuka Nagano 1, Emi Furukawa 2, Shinya Ito 1
PMCID: PMC13389713  PMID: 42487889

Abstract

Objective: Syphilis is a major re-emerging infectious disease, and pre-consultation support may be particularly important in rural and remote areas where access to confidential care is limited and stigma can delay help-seeking. We examined how communication features in ChatGPT-generated responses influence perceived consultation satisfaction and intention to seek professional care in a simulated syphilis-related anxiety scenario in Japan.

Patient/Materials and Methods: We conducted a controlled dialogue experiment in Japanese using standardized interactions between a simulated patient concerned about syphilis and ChatGPT (GPT-4o mini). Across 48 three-turn sessions, artificial intelligence (AI) responses varied by requested communication stance and personal information disclosure. Supportive communication features, stance alignment, personalization consistency, readability (jReadability), and information volume (character count) were coded. Three health professional evaluators rated consultation satisfaction using the Consultation and Relational Empathy (CARE) Measure and rated care-seeking intention using a Readiness Ruler.

Results: CARE scores showed moderate inter-rater reliability, whereas intention ratings showed low agreement. Two AI responses contained medical misinformation. In hierarchical regression analyses, longer responses and higher readability were associated with higher satisfaction, and among supportive communication behaviors, praise expressions independently associated with satisfaction. Satisfaction was strongly associated with care-seeking intention (β=0.80, P<0.001).

Conclusion: In Japanese-language simulated syphilis consultations, information volume, readability, and praise expressions were associated with higher perceived consultation satisfaction, which in turn was associated with greater intention to seek professional care. These findings suggest that conversational AI may serve as a low-threshold, anonymous pre-consultation support tool to encourage timely help-seeking, particularly in rural and remote settings where confidential access to sexual health services is limited. Low inter-rater agreement for intention ratings and the presence of misinformation underscore the need for robust outcome measurement and safety safeguards.

Keywords: syphilis, ChatGPT, conversational AI, consultation satisfaction, care-seeking intention

Introduction

Syphilis has re-emerged globally as a major public health concern, with an estimated eight million new infections among individuals aged 15–49 years in 20221). In Japan, notifications have risen sharply over the past decade, reaching a record 14,895 cases in 2023 and remaining high in 20242). Untreated syphilis can cause severe complications, including adverse pregnancy outcomes as well as neurological and cardiovascular disorders, making early consultation and testing essential. However, stigma, shame, and fear surrounding sexually transmitted infections (STIs) continue to impede timely care-seeking3,4,5). These psychosocial barriers can delay access to testing and professional support, contributing to ongoing transmission. Syphilis-related stigma, in particular, has been shown to reduce testing uptake and discourage engagement with healthcare services.

In rural and regional areas of Japan, access to healthcare can be constrained by geographic distance and persistent workforce shortages, contributing to well-documented rural–urban disparities in access to and quality of care6). In addition, the geographical maldistribution of obstetricians and gynecologists—specialists who often provide sexual and reproductive health services—may further limit local consultation options7). In such contexts, opportunities for confidential consultation are reduced, which is particularly consequential for conditions associated with social stigma. Prior research has shown that STI-related stigma can act as a barrier to testing and timely care-seeking, especially when concerns about privacy and social judgment are present8). Accordingly, in the context of syphilis and other STIs, fear of being recognized within small communities and anticipated stigma may further increase hesitancy to seek professional advice, potentially leading to delays in accessing appropriate care in rural settings.

Given these barriers, people may turn to low-threshold, anonymous sources of information before seeking in-person care. Digital platforms have become central sources of sexual health information. Social media and artificial intelligence (AI) chatbots are increasingly used to obtain STI-related information9, 10). However, the quality of digital content remains uneven. A recent evaluation of Japanese-language YouTube Shorts found that 80% of videos showed low reliability, and 20% contained misinformation or stigmatizing content11). Similar concerns have been raised regarding large language models (LLMs), which may provide variable or inaccurate medical information12, 13). AI chatbots such as ChatGPT have therefore attracted attention as potential pre-consultation support tools. Prior research has shown that AI chatbots are feasible and acceptable in mental health contexts14), and ChatGPT responses may even be rated as more empathetic and higher in quality than clinician replies in online medical forums15). Reviews suggest benefits for user engagement but also highlight challenges related to empathy consistency, personalization, and safety16). Despite growing public use, little is known about how AI-generated messages influence satisfaction or intention to seek care in STI-related contexts, particularly for syphilis. Whether variations in user prompts (such as requesting listening versus advice) or disclosure of personal attributes shapes these outcomes also remains unclear.

Guided by the communication model developed by Street et al., which posits that communication behaviors influence proximal outcomes such as satisfaction and subsequently affect care-seeking behaviors17), this study systematically examined AI–user interactions. Within this framework, patient-centered communication encompasses core functions such as information exchange (e.g., providing explanations or advice) and responding to emotions (e.g., listening and empathy). In this study, these two functions were operationalized as distinct consultation stances (advice-oriented versus listening-oriented), alongside a condition without a stance request. We conducted a controlled dialogue experiment in which a simulated patient expressing anxiety about syphilis consulted ChatGPT. We evaluated how supportive communication features, stance alignment, personalization, readability, and information volume influenced satisfaction and readiness to seek medical consultation.

Therefore, this study examined how communication features in ChatGPT-generated responses influence perceived consultation satisfaction and subsequent intention to seek professional care in a simulated syphilis-related anxiety scenario in Japan. By systematically varying the requested communication stance and personal information disclosure in controlled Japanese-language dialogues, we aimed to identify which response characteristics are associated with higher satisfaction and care-seeking intention.

Patient/Materials and Methods

Study design

This study employed a controlled dialogue experiment conducted entirely in Japanese, reflecting typical online health-seeking behavior among Japanese-speaking individuals concerned about STIs. All dialogues were generated between a simulated patient and ChatGPT (OpenAI, San Francisco, CA, USA, GPT-4o mini) under controlled conditions. The simulated patient’s utterance were fixed across all sessions, and only the AI-generated responses varied according to experimental manipulations. One nursing-student co-author acted as the simulated patient across all sessions and selected the pre-specified second and third turns from Japanese Dialogue-Act-based templates (Supplementary Table 1)18). As this study analyzed AI-generated dialogue text and did not collect identifiable personal data from patients or the public, institutional ethical review was not required. Evaluators were members of the research team and provided ratings as part of the study procedures.

Experimental conditions and dialogue structure

Two experimental factors were manipulated: 1) the communication stance requested by the simulated patient; and 2) the presence or absence of personal information disclosure. Three stance versions of the initial message were prepared: a listening-oriented request (“I would just be grateful if you could listen to me”); an advice-seeking request (“I would appreciate any advice you can give”); and a stance-neutral expression (“Thank you in advance”).

For the disclosure manipulation, four attributes—sex (male or female), age group (20s, 40s, or 60s), occupation (healthcare worker or office worker), and living situation (living alone or with family)—were randomly assigned and added to the initial message in the disclosure condition; no attributes were included in the non-disclosure condition. The study did not aim to measure socioeconomic status per se. Instead, to examine whether personalization in AI responses to disclosed information might influence satisfaction, we focused on two attributes likely to prompt changes in response content: whether the user was a healthcare worker and whether they lived with a partner. Combining these factors yielded six conditions (3 stances × 2 disclosure levels), with eight dialogue sessions per condition (total=48). Each session consisted of three turns (patient → AI → patient → AI → patient → AI). The simulated patient’s utterances were held constant within each Dialogue-Act template to enable direct comparison of AI responses (Supplementary Table 2).

Simulated patient utterances and Dialogue-Act framework

Simulated patient utterance templates were constructed using the Dialogue-Act Classification framework described by Malhotra et al18). Of the 12 original categories, four unsuitable for early consultation phases were excluded: Greeting; Positive Answer; Negative Answer; and Clarification Delivery. The remaining eight categories—Information Request; Clarification Request; Yes/No Question; Opinion Request; Information Delivery; Opinion Delivery; Acknowledgement; and Grounding Chat—were used to develop candidate utterances for the second and third turns. For each category, the research team prepared natural-sounding Japanese options (e.g., for Clarification Request: “What exactly does the test involve?” or “Would it be done just with a blood test?”). The simulated patient selected the corresponding pre-defined template in each session.

ChatGPT response settings

To maintain consistency across sessions, ChatGPT was given a uniform system instruction to respond as an experienced nurse who acknowledges emotional concerns, avoids excessive medical jargon, and provides clear and polite explanations. This ensured that stylistic features remained consistent across all sessions.

Coding of supportive communication features

A coding framework was developed to identify supportive communication elements in AI responses. The first author reviewed all responses and coded expressions related to behavioral recommendations, de-stigmatization, emotional validation, facilitation of disclosure or listening, presence/support, empowerment, positive appraisal, praise, and anxiety reduction. These ten subcategories were grouped under four higher-order categories conceptually aligned with the communication functions defined by Street et al.17): emotional response; uncertainty management; relationship building; and decision support. Each subcategory was coded as present (1) or absent (0) for each session. Coding was conducted at the session level. Since each session contained three AI responses, a supportive communication feature was coded as present if it appeared in any of the three responses within that session.

Outcome measures

Consultation satisfaction

Consultation satisfaction was assessed using the Consultation and Relational Empathy (CARE) Measure, which has been validated both internationally19) and in Japan20). Three evaluators with medical or health science backgrounds independently reviewed all 48 sessions and rated the 10 CARE items on a 0–4 scale (total score range, 0–40).

Healthcare-seeking intention

Intention to seek professional consultation was assessed using a single-item Readiness Ruler adapted from Rollnick et al.21): “After completing this interaction with the AI; how willing are you to consult a healthcare professional or another person?” Responses were recorded on a scale from 0 (“not at all willing”) to 10 (“very willing”).

Response alignment

Stance alignment assessed whether AI responses matched the requested communicative stance. Personalization consistency assessed whether disclosed personal attributes (when present) were appropriately addressed. Both variables were coded as 1 (aligned) or 0 (not aligned).

Readability metrics

The three AI-generated responses in each session were concatenated and analyzed using jReadability22), producing a readability score (0.5–1.4=“very difficult”, 1.5–2.4=“difficult”, 2.5–3.4=“somewhat difficult”, 3.5–4.4=“neutral”, 4.5–5.4=“readable”, and 5.5–6.4=“very readable”) and total character count. Readability was treated as an index of ease of comprehension, while character count represented information volume.

Statistical analysis

First, descriptive statistics were calculated for all variables. Each evaluator independently rated all dialogue sessions, and for statistical analyses, each rater’s score was treated as an independent observation (48 sessions × 3 raters=144 ratings). Next, inter-rater reliability for CARE scores and intention ratings was assessed using intraclass correlation coefficients (ICCs) based on a two-way random-effects model with absolute agreement. Pearson correlation coefficients were then computed to examine associations among variables. Hierarchical multiple regression analyses were conducted to identify predicators of consultation satisfaction, with the CARE score as the dependent variable. In Step 1, stance type, personal information disclosure, stance alignment, personalization consistency, total character count, and jReadability score were entered. In Step 2, the ten supportive communication subcategories were added. The number of predictors included in the regression models was determined to ensure that the number of observations exceeded the commonly recommended minimum of 10 observations per predictor variable. Multicollinearity was assessed using variance inflation factors, with values below 10 considered acceptable.

The unit of analysis for outcome ratings was the rater–session pair, resulting in 144 observations (48 sessions × 3 raters). CARE scores and intention to seek help were treated as rater-specific ratings. Stance alignment and personalization consistency were also evaluated by each rater based on the three AI responses within each session and were therefore treated as rater-level variables. In contrast, supportive communication features were coded at the session level by the first author. As each session contained three AI responses, a feature was coded as present if it appeared in any of the responses. Text characteristics, including readability and total character count, were computed at the session level by concatenating the three AI responses within each session. These session-level variables were then merged with each rater’s ratings for statistical analyses. Finally, a simple linear regression model evaluated whether CARE scores predicted intention to seek medical consultation. All analyses were conducted using IBM SPSS Statistics version 25 (IBM Corp., Tokyo, Japan), with significance set at P<0.05.

Results

Inter-rater reliability among the three evaluators demonstrated moderate agreement for the CARE Measure, with an ICC of 0.468 (95% confidence interval [CI]=0.167–0.677). In contrast, agreement for the Readiness Ruler was low (ICC=0.137, 95% CI=−0.318–0.468). Assessment of the medical accuracy of all 48 ChatGPT responses by a physician identified two instances of misinformation: one concerning an inaccurate explanation of syphilis testing and treatment procedures, and the other involving an incorrect description of symptoms.

Regarding the linguistic and alignment characteristics of the AI-generated responses, readability of the Japanese text was rated “Somewhat difficult” in 28 sessions (58.3%) and “Neutral” in 20 sessions (41.7%) (Table 1). Stance alignment—defined as concordance between the communication stance requested by the simulated patient and the tone of the AI response—was present in 97 responses (67.4%). Personalization consistency, referring to the appropriate reflection of disclosed personal attributes in AI responses, was observed in 35 responses (24.3%). Across all sessions, the mean CARE score was 30.7 (standard deviation [SD]=6.3). The average total number of Japanese characters generated by the AI was 2,118.9 (SD=773.2), and the mean jReadability score was 3.4 (SD=0.3). The mean intention to seek medical consultation score was 6.0 (SD=1.8). Analysis of supportive communication elements showed that emotional validation, listening or disclosure facilitation, and positive appraisal appeared in all AI responses. Since positive appraisal appeared in all sessions and therefore showed no variance, it is reported descriptively (Table 2) but was excluded from regression analyses. Other subcategories appeared less consistently: de-stigmatization was present in 33 responses (68.8%), while empowerment and praise appeared in 22 responses (45.8%).

Table 1. Descriptive characteristics of the dialogue conditions, AI response features, and outcome measures.

Communication stance (n, %)
Listening 16 33.3
Advising 16 33.3
No specific stance 16 33.3
Disclosure of personal information (n, %)
Disclosed 24 50.0
Not disclosed 24 50.0
jReadability (n, %)
Very readable 0 0.0
Readable 0 0.0
Neutral 20 41.7
Somewhat difficult 28 58.3
Difficult 0 0.0
Very difficult 0 0.0
Others (n, %)
Response aligned with requested stance 97 67.4
Response aligned with disclosed personal information 35 24.3
CARE score (M, SD) 30.7 6.3
jReadability metrics (M, SD)
Total characters of AI responses 2,118.9 773.2
jReadability score 3.4 0.3
Healthcare-seeking intention (M, SD) 6.0 1.8

CARE: consultation and relational empathy; M: mean.

Table 2. Frequencies of supportive communication categories in ChatGPT responses.

n %
Action facilitation
Behavioral recommendation 40 83.3
De-stigmatization 33 68.8
Empathy
Emotional validation 48 100.0
Listening / Disclosure facilitation 48 100.0
Support & Presence 44 91.7
Praise & Empowerment
Empowerment 22 45.8
Positive appraisal 48 100.0
Praise 22 45.8
Reassurance
Anxiety reduction 43 89.6
Positive appraisal 45 93.8

“Positive appraisal” under Reassurance refers to reassuring affirmations, whereas “Positive appraisal” under Praise & Empowerment refers to positive evaluative statements; the variables used in regression analyses follow the definitions specified in the coding framework.

Correlation analyses (Table 3) indicated that CARE scores were significantly associated with intention to seek medical consultation (r=0.76, P<0.01), personalization consistency (r=0.17, P<0.05), total character count (r=0.35, P<0.01), jReadability score (r=0.26, P<0.01), de-stigmatization (r=0.28, P<0.01), empowerment (r=0.28, P<0.01), and praise (r=0.43, P<0.01). Hierarchical regression analyses examining predictors of CARE scores are shown in Table 4. Model 1, which included the requested communication stance and the presence or absence of personal information disclosure, did not significantly predict CARE scores (R2=0.027, adjusted R2=0.006). In Model 2, adding stance alignment, personalization consistency, total character count, and readability revealed significant associations for total character count (β=0.40, P<0.001) and jReadability score (β=0.22, P=0.006), increasing the explained variance to R2=0.224 (adjusted R2=0.184). In Model 3, which additionally included all supportive communication subcategories, praise emerged as the only significant independent predictor of CARE scores (β=0.24, P=0.016), with the model explaining R2=0.289 (adjusted R2=0.212). No other supportive communication subcategories were significant. Finally, simple linear regression analysis evaluating whether CARE scores were associated with intention to seek medical consultation (Table 5) demonstrated a strong positive association (β=0.80, P<0.001). The model accounted for 57.0% of the variance (R2=0.570, adjusted R2=0.567).

Table 3. Correlations among communication features, readability metrics, supportive communication categories, and outcomes.

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
1 CARE score 0.76** −0.01 −0.06 0.07 0.15 0.15 0.17* 0.35** 0.26** 0.11 0.28** 0.01 0.28** 0.43** 0.12 −0.03
2 Healthcare-seeking intention −0.13 −0.01 0.14 0.12 0.08 0.21* 0.33** 0.11 0.09 0.15 0.00 0.24** 0.32** 0.07 0.02
3 Listening −0.50** −0.50** 0.00 0.08 −0.09 −0.44** 0.04 −0.16 0.19* −0.11 −0.03 −0.03 0.10 −0.18*
4 Advising −0.50** 0.00 0.27** −0.06 0.07 0.00 0.08 −0.19* 0.21* 0.15 −0.03 −0.34** 0.18*
5 No specific stance 0.00 −0.36** 0.15 0.37** −0.05 0.08 0.00 −0.11 −0.12 0.06 0.24** 0.00
6 Disclosure of personal information 0.10 0.57** 0.21* 0.02 0.22** 0.14 0.30** 0.08 0.33** 0.21* −0.09
7 Response aligned with requested stance 0.15 −0.02 0.02 0.13 0.01 0.11 0.11 0.19* −0.14 −0.06
8 Response aligned with disclosed personal information 0.22** 0.04 0.21* 0.10 0.17* 0.19* 0.36** 0.14 −0.19*
9 Total characters of AI responses 0.10 0.41** 0.25** 0.30** 0.33** 0.48** 0.08 0.12
10 jReadability score 0.28** 0.60** 0.03 0.27** 0.21* 0.20* 0.02
11 Behavioral recommendation 0.30** 0.47** 0.08 0.19* −0.15 0.12
12 De-stigmatization 0.29** 0.26** 0.17* 0.21* 0.01
13 Support & Presence −0.03 0.13 −0.10 0.23**
14 Empowerment 0.41** −0.10 0.07
15 Praise 0.18* −0.11
16 Anxiety reduction −0.09
17 Positive appraisal

Communication variables: Supportive-communication subcategories were included as 0/1 dummy variables; items that appeared in 100% of all AI responses were excluded because they had no variance. CARE: consultation and relational empathy; AI: artificial intelligence.

Table 4. Hierarchical regression analysis predicting patient satisfaction.

β Standardized β t-value P-value 95% CI VIF

Lower Upper
Model 1 (R2=0.027, adjusted R2=0.006)
Constant 29.6 28.1 0.000 27.5 31.7
Advising −0.4 0.0 −0.3 0.772 −2.9 2.2 1.3
No specific stance 0.8 0.1 0.6 0.562 −1.8 3.3 1.3
Disclosure of personal information 1.8 0.1 1.8 0.081 −0.2 3.9 1.0
Model 2 (R2=0.224, adjusted R2=0.184)
Constant 6.3 1.0 0.319 −6.1 18.7
Advising −2.3 −0.2 −1.8 0.071 −4.7 0.2 1.5
No specific stance −1.2 −0.1 −0.9 0.396 −4.0 1.6 1.9
Disclosure of personal information 0.3 0.0 0.2 0.815 −2.1 2.6 1.5
Response aligned with requested stance 2.2 0.2 1.9 0.056 −0.1 4.4 1.2
Response aligned with disclosed personal information 0.6 0.0 0.5 0.646 −2.1 3.4 1.6
Total characters of AI responses 0.0 0.4 4.1 0.000 0.0 0.0 1.4
jReadability score 5.0 0.2 2.8 0.006 1.5 8.6 1.0
Model 3 (R2=0.289, adjusted R2=0.212)
Constant 15.8 1.9 0.066 −1.1 32.7
Advising −1.0 −0.1 −0.7 0.488 −3.8 1.8 2.0
No specific stance −0.4 0.0 −0.2 0.804 −3.4 2.6 2.3
Disclosure of personal information 0.4 0.0 0.3 0.741 −2.0 2.8 1.7
Response aligned with requested stance 1.8 0.1 1.6 0.117 −0.5 4.0 1.3
Response aligned with disclosed personal information 0.2 0.0 0.1 0.911 −2.7 3.0 1.8
Total characters of AI responses 0.0 0.3 2.3 0.021 0.0 0.0 2.5
jReadability score 3.0 0.1 1.2 0.232 −2.0 8.0 2.1
Behavioral Recommendation −1.7 −0.1 −1.0 0.304 −5.1 1.6 1.8
De-stigmatization 1.9 0.1 1.3 0.211 −1.1 5.0 2.3
Support & Presence −2.8 −0.1 −1.2 0.223 −7.4 1.7 1.8
Empowerment 0.0 0.0 0.0 0.991 −2.4 2.4 1.7
Praise 3.2 0.2 2.4 0.016 0.6 5.7 1.9
Anxiety reduction −0.8 0.0 −0.4 0.674 −4.5 2.9 1.5
Positive appraisal 0.6 0.0 0.3 0.772 −3.6 4.8 1.2

Reference category: In the regression models, listening stance was used as the reference category for stance-type predictors; Communication variables: Supportive-communication subcategories were included as 0/1 dummy variables; items that appeared in 100% of all AI responses were excluded because they had no variance; Model structure: Models 1–3 represent a hierarchical regression, with communication-category variables added in Model 3. VIF: variance inflation factor; AI: artificial intelligence.

Table 5. Regression analysis predicting healthcare-seeking intention.

β Standardized β t-value P-value 95% CI

Lower Upper
Constant −0.4 −0.9 0.363 −1.4 0.5
CARE 0.2 0.8 13.7 0.000 0.2 0.2

R2=0.570, adjusted R2=0.567. CARE: consultation and relational empathy.

Discussion

The aim of this study was to examine how communication features in ChatGPT-generated responses influence perceived consultation satisfaction and, subsequently, the intention to seek medical care in a simulated syphilis-related anxiety scenario. The findings showed that longer AI-generated responses and the inclusion of praise-related expressions (e.g., “That is a great question”) were associated with higher CARE scores. In turn, higher consultation satisfaction was strongly associated with a greater willingness to consult a healthcare professional. These results partially support the communication pathway proposed by Street et al.17), in which proximal outcomes such as satisfaction mediate downstream behavioral intentions.

An important methodological consideration concerns the low inter-rater reliability observed for the Readiness Ruler (ICC=0.137). Behavioral intention is inherently tied to subjective constructs such as personal values, health beliefs, and perceived vulnerability. Third-party evaluators may thus struggle to infer intention solely from written dialogues. The low ICC likely reflects this conceptual difficulty rather than simple measurement error. Nevertheless, the variability among evaluators indicates uncertainty in intention estimates, warranting caution in interpreting the strength of the association between satisfaction and behavioral intention.

Contrary to our expectations, supportive communication categories traditionally linked to behavior change (such as Action Facilitation, Empathy, and Reassurance) were not significant predictors of satisfaction. This pattern suggests that ChatGPT responses may converge toward a relatively uniform empathic and supportive style due to the assigned “experienced nurse” role, thereby reducing variability in qualitative communication features. Previous research has shown that ChatGPT often provides more uniformly empathic and polite responses than human clinicians15). This homogenization may have limited the extent to which differences in qualitative communication could influence satisfaction scores in the present study. Similarly, neither the requested communication stance (listening vs. advising vs. unspecified) nor the disclosure of personal information influenced satisfaction, likely because the role instruction constrained the AI to maintain a stable tone across conditions.

Instead, formal linguistic characteristics—specifically response length and readability—emerged as significant predictors of satisfaction. This aligns with prior work showing that users value ChatGPT for its structured, detailed, and easily comprehensible explanations in medical contexts15, 23). Studies of LLM-generated medical information have likewise reported that coherent, readable explanations enhance understanding16). Users may thus prioritize the clarity and quantity of information over nuanced emotional support when interacting with AI, especially in anonymous consultations about stigmatized conditions such as STIs. The professional backgrounds of the evaluators (physician, epidemiologist, or nursing student) may also have biased ratings toward valuing informational completeness and clarity, potentially attenuating the perceived impact of qualitative relational features. Taken together, these findings indicate that, in AI-mediated consultations, user satisfaction and subsequent care-seeking intention may be shaped more by the clarity, amount, and perceived usefulness of information than by nuanced relational communication features. In anonymous and low-stakes interactions—such as simulated consultations for stigmatized conditions like sexually transmitted infections—users may prioritize feeling sufficiently informed and reassured over experiencing individualized empathic responses. This behavioral pathway provides an important lens through which the potential role of conversational AI can be considered in broader healthcare access contexts.

In rural and remote settings, access to specialized sexual health services and confidential consultation may be limited, and individuals seeking care for sexually transmitted infections may face compounded barriers related to service availability and stigma3). These barriers may discourage timely testing and consultation, potentially leading to delayed diagnosis or untreated infection. The present findings—showing that clearer, more readable, and information-rich AI-generated responses were associated with higher consultation satisfaction and greater care-seeking intention—suggest that conversational AI may help lower access barriers by providing an anonymous and low-threshold entry point to information and remote consultation pathways12). Although such systems are not intended to diagnose or replace clinical care, they may function as an initial gateway that encourages timely professional evaluation. This function may be particularly relevant in rural and underserved settings, where opportunities for confidential sexual health consultation are limited. By emphasizing user satisfaction and behavioral intention rather than diagnostic accuracy, this study adopts a perspective central to rural and remote health practice. Interventions that facilitate early engagement with healthcare services may be especially valuable in contexts where delayed presentation can result in worse individual and public health outcomes. Importantly, since the present experiment did not directly include rural participants or healthcare settings, these implications for rural and remote contexts should be interpreted as conceptual and hypothesis-generating rather than empirically demonstrated rural-specific effects.

Several limitations should be acknowledged. First, all dialogues were constructed using predefined templates for the simulated patient and therefore cannot fully capture the variability of real-world patient expressions, emotions, and iterative questioning. The short, three-turn dialogue structure also limits the ecological validity of the interaction and may not allow the model to demonstrate adaptive or personalized behaviors. Second, as patient utterances were standardized based on Dialogue-Act categories, opportunities for ChatGPT to adjust responses dynamically were restricted. While beneficial for experimental control, this approach limits insight into how AI might respond to more naturalistic conversational cues such as hesitation, confusion, or changing emotional states.

Third, satisfaction and behavioral intention were evaluated by only three raters with health-related professional backgrounds. The shared emphasis on clinical accuracy and completeness may thus have shaped the evaluation of communication quality. Future work involving more diverse raters (including lay users) may better capture general perceptions of AI-delivered consultation. The exceptionally low ICC for behavioral intention suggests significant instability in this outcome measure. Future studies should therefore consider multi-item intention scales, self-reported intention from actual users, or behavioral follow-up (e.g., whether participants ultimately seek testing).

Fourth, the findings are model-specific. This study used ChatGPT-4o mini, and the communicative properties of LLMs evolve rapidly across versions. The reproducibility of these findings across models or future iterations thus remains uncertain. Fifth, two instances of medically inaccurate information were identified in the AI responses. In these cases, the AI generated statements that were not fully consistent with current clinical knowledge regarding the management of sexually transmitted infections. Although the number of such inaccuracies was small, they could potentially mislead users or delay appropriate help-seeking if interpreted as authoritative medical advice. These observations illustrate persistent risks associated with AI-generated medical information and highlight the importance of safety monitoring and expert validation before integrating AI chatbots into healthcare pathways. In addition, the use of conversational AI to encourage consultation with healthcare professionals may raise legal and ethical considerations. Because AI-generated responses may contain inaccuracies, such systems should not be interpreted as providing medical diagnosis or treatment but rather as informational support, and appropriate safety monitoring and expert oversight will be important for responsible implementation. Sixth, behavioral intention was used as a proxy and does not directly measure actual healthcare-seeking behavior. Longitudinal studies are needed to determine whether AI-assisted consultations can meaningfully influence real-world consultation or testing uptake. Finally, this study did not specifically sample rural or remote populations; therefore, the implications for rural and remote settings should be interpreted as hypothesis-generating rather than directly generalizable.

Conclusion

This study examined how the communicative characteristics of ChatGPT influence consultation satisfaction and the intention to seek medical care in a simulated syphilis-related counseling scenario. Formal linguistic features—information volume, readability, and praise expressions—were associated with higher perceived satisfaction, which in turn predicted greater willingness to seek testing or professional consultation. These findings highlight the potential role of conversational AI as a low-threshold, anonymous “first step” that may help address access and privacy barriers to sexual health consultation. Although rural and remote communities represent one plausible context in which such barriers are prominent, the relevance to rural settings should be interpreted as suggestive rather than directly generalizable.

Conflict of interest

The authors declare no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

Funding information

This work was supported by JSPS KAKENHI Grant Numbers JP24K06389, JP24K23676, and JP25K20501.

Ethical considerations

This study analyzed AI-generated dialogue text and did not collect identifiable personal data from patients or the public. Ratings were provided by health-professional evaluators for research purposes only. Therefore, ethical approval was not required.

Consent to participate

Consent was not required because no identifiable personal data were collected and the study analyzed only AI-generated text; evaluators provided ratings as part of research procedures.

Consent for publication

All authors consent to the publication of this manuscript in the Journal of Rural Medicine.

Author contributions

Fuka Nagano contributed to the study design, constructed the simulated patient utterances, generated the dialogue data, and participated in data coding and interpretation. Emi Furukawa contributed to the study design, supervised the assessment of medical accuracy, and supported data analysis and interpretation. Shinya Ito conceived the study, designed the methodology, supervised all analytic procedures, performed the statistical analyses, and drafted and revised the manuscript. All authors critically reviewed the manuscript, approved the final version, and agreed to be accountable for all aspects of the work.

Supplementary Material

jrm-21-3-269-s001.pdf (154.7KB, pdf)

Funding Statement

This work was supported by JSPS KAKENHI Grant Numbers JP24K06389, JP24K23676, and JP25K20501.

Data availability

All data used in this study consist of AI-generated dialogue texts. These materials are available from the corresponding author upon reasonable request.

References

  • 1.World Health Organization. Syphilis fact sheet. 2025. https://www.who.int/news-room/fact-sheets/detail/syphilis.
  • 2.National Institute of Infectious Diseases, Infectious Disease Epidemiology Center. Syphilis surveillance in Japan (infectious diseases weekly report). 2025. https://id-info.jihs.go.jp/surveillance/idwr/article/syphilis/010/index.html.
  • 3.Malone S, Counts L, Zabotka L, et al. RESPECT TeambRESPECT Team.Stigma measurement in health: a systematic review. EClinicalMedicine 2025; 86: 103360. doi: 10.1016/j.eclinm.2025.103360 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.The Lancet child adolescent health. Youth STIs: an epidemic fuelled by shame. Lancet Child Adolesc Health 2022; 6: 353. doi: 10.1016/S2352-4642(22)00128-6 [DOI] [PubMed] [Google Scholar]
  • 5.World Health Organization. Global health sector strategies on, respectively, HIV, viral hepatitis and sexually transmitted infections for the period 2022–2030. Global strategy 2022.
  • 6.Kaneko M, Ohta R, Mathews M. Rural and urban disparities in access and quality of healthcare in the Japanese healthcare system: a scoping review. BMC Health Serv Res 2025; 25: 667. doi: 10.1186/s12913-025-12848-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Matsumoto K, Seto K, Hayata E, et al. The geographical maldistribution of obstetricians and gynecologists in Japan. PLoS One 2021; 16: e0245385. doi: 10.1371/journal.pone.0245385 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Lee ASD, Cody SL. The stigma of sexually transmitted infections. Nurs Clin North Am 2020; 55: 295–305. doi: 10.1016/j.cnur.2020.05.002 [DOI] [PubMed] [Google Scholar]
  • 9.Yalamanchili A, Sengupta B, Song J, et al. Quality of large language model responses to radiation oncology patient care questions. JAMA Netw Open 2024; 7: e244630. doi: 10.1001/jamanetworkopen.2024.4630 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Zack T, Lehman E, Suzgun M, et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit Health 2024; 6: e12–e22. doi: 10.1016/S2589-7500(23)00225-X [DOI] [PubMed] [Google Scholar]
  • 11.Furukawa E, Okuhara T, Ito S, et al. Hidden misinformation in YouTube short videos on syphilis: a mixed-methods study. PEC Innov 2025; 7: 100428. doi: 10.1016/j.pecinn.2025.100428 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Huo B, Boyle A, Marfo N, et al. Large language models for chatbot health advice studies: a systematic review. JAMA Netw Open 2025; 8: e2457879. doi: 10.1001/jamanetworkopen.2024.57879 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Nickel B, Moynihan R, Gram EG, et al. Social media posts about medical tests with potential for overdiagnosis. JAMA Netw Open 2025; 8: e2461940–e2461940. doi: 10.1001/jamanetworkopen.2024.61940 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Casu M, Triscari S, Battiato S, et al. AI chatbots for mental health: a scoping review of effectiveness, feasibility, and applications. Appl Sci (Basel) 2024; 14: 5889. doi: 10.3390/app14135889 [DOI] [Google Scholar]
  • 15.Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med 2023; 183: 589–596. doi: 10.1001/jamainternmed.2023.1838 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Hindelang M, Sitaru S, Zink A. Transforming health care through chatbots for medical history-taking and future directions: comprehensive systematic review. JMIR Med Inform 2024; 12: e56628. doi: 10.2196/56628 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Street RL, Jr, Makoul G, Arora NK, et al. How does communication heal? Pathways linking clinician-patient communication to health outcomes. Patient Educ Couns 2009; 74: 295–301. doi: 10.1016/j.pec.2008.11.015 [DOI] [PubMed] [Google Scholar]
  • 18.Malhotra G, Waheed A, Srivastava A, et al. Speaker and time-aware joint contextual learning for dialogue-act classification in counselling conversations. proceedings of the fifteenth ACM international conference on web search and data mining 2022; 735–745. [Google Scholar]
  • 19.Mercer SW, Maxwell M, Heaney D, et al. The consultation and relational empathy (CARE) measure: development and preliminary validation and reliability of an empathy-based consultation process measure. Fam Pract 2004; 21: 699–705. doi: 10.1093/fampra/cmh621 [DOI] [PubMed] [Google Scholar]
  • 20.Aomatsu M, Abe H, Abe K, et al. Validity and reliability of the Japanese version of the CARE measure in a general medicine outpatient setting. Fam Pract 2014; 31: 118–126. doi: 10.1093/fampra/cmt053 [DOI] [PubMed] [Google Scholar]
  • 21.Rollnick S, Mason P, Butler C. Health behavior change: a guide for practitioners. Churchill Livingstone, 1999. [Google Scholar]
  • 22.Lee J. Readability Research for Japanese Language Education. Waseda Studies in Japanese Language Education 2016: 1-16. [Google Scholar]
  • 23.Zaretsky J, Kim JM, Baskharoun S, et al. Generative artificial intelligence to transform inpatient discharge summaries to patient-friendly language and format. JAMA Netw Open 2024; 7: e240357. doi: 10.1001/jamanetworkopen.2024.0357 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

jrm-21-3-269-s001.pdf (154.7KB, pdf)

Data Availability Statement

All data used in this study consist of AI-generated dialogue texts. These materials are available from the corresponding author upon reasonable request.


Articles from Journal of Rural Medicine : JRM are provided here courtesy of Japanese Association of Rural Medicine

RESOURCES