Abstract
Background
Social media platforms such as X (formerly Twitter) are increasingly used by journals, authors, and institutions to promote newly published research. Well-designed posts can enhance visibility, accelerate knowledge translation, and increase altmetric attention. However, creating accurate and policy-compliant content is time-intensive. Large language models (LLMs) offer a potential solution, yet systematic evaluations of their performance in post-publication promotion remain limited.
Methods
We conducted a blinded, crossed, offline evaluation of four LLMs: GPT-5 (OpenAI), Gemini 2.5 Pro (Google DeepMind), Grok-3 (xAI), and Perplexity Pro (Perplexity AI), tasked with generating X-style posts (≤ 260 characters) for 36 open access articles from The Lancet Public Health, The Lancet Planetary Health, and Annual Review of Public Health. Posts were generated using a standardized system and user prompt. A single blinded rater scored outputs using a five-domain rubric (factual accuracy, clarity, policy compliance, call-to-action quality, structure/metadata; maximum score 10). Secondary measures included character count, hashtag use, and readability (Flesch–Kincaid Grade Level). General linear models with Bonferroni-adjusted post hoc tests and non-parametric analyses were applied.
Results
All four models achieved perfect factual accuracy and no policy violations. Mean total quality scores differed significantly by model, P < 0.001. GPT-5 (9.60) and Perplexity Pro (9.60) performed best, followed by Gemini 2.5 Pro (9.47), while Grok-3 scored lower (8.80). Domain analyses showed Grok-3 underperformed in call-to-action quality (1.40 vs. ≥1.97 in other models, P < 0.001) and produced significantly shorter posts (median 194 characters, P < 0.001). Perplexity Pro scored highest for policy compliance, while GPT-5 and Gemini 2.5 Pro achieved superior structural scores. Readability varied: GPT-5 8.9 (7.3–9.2) and Perplexity Pro 7.3 (6.5–8.8) generated more complex outputs, whereas Gemini 2.5 Pro 5.1 (4.8–6.5) and Grok-3 4.5 (3.6–6.3) produced more accessible posts.
Conclusion
LLMs can reliably generate accurate and policy-compliant social media posts for research promotion, with differences in style and readability that may inform audience targeting. GPT-5, Gemini 2.5 Pro, and Perplexity Pro produced high-quality outputs, while Grok-3 underperformed across several domains. These findings highlight the potential of LLMs as scalable first-draft tools for post-publication promotion, capable of improving the reach and accessibility of scientific research. Careful model selection, tailored to audience and communication goals, together with human oversight, remains essential.
Keywords: Publishing, Social Media, Large Language Models, Scholarly Communication, Public Health, Knowledge Management
Graphical Abstract
INTRODUCTION
Digital platforms such as social media now serve as key channels for disseminating scientific knowledge, enabling rapid and wide-reaching engagement with scholarly work beyond traditional academic audiences. In public health and medicine, these channels, most notably X (formerly Twitter), LinkedIn, and other professional networks are increasingly used by journals, authors, and institutions to promote newly published research, highlight key findings, and stimulate scholarly discussion. This process, often referred to as post-publication promotion, complements conventional dissemination routes and has been linked to increased article visibility, higher readership, and improved alternative metrics (“altmetrics”) such as the Altmetric Attention Score and PlumX captures. Unlike citation-based metrics, altmetrics capture early attention from practitioners, policymakers, journalists, and the public, making them especially relevant for translating public health evidence into practice.
Recent years have seen a surge of research into the relationship between social media promotion and the visibility or impact of scientific publications. Multiple observational studies have found that articles actively promoted on platforms such as X and LinkedIn accrue higher Altmetric Attention Scores, greater download counts, and, in some cases, increased citation rates compared to non-promoted articles.1,2,3 Randomized trials in biomedical journals have demonstrated that coordinated Twitter campaigns can significantly increase article page views and PDF downloads within days of publication.4,5 While the association between altmetrics and long-term citation counts remains mixed,6,7,8 early online attention has been linked to broader audience engagement, including policy uptake and media coverage.9 In public health specifically, social media has been shown to facilitate rapid dissemination of practice-relevant findings to clinicians, researchers, and decision-makers.10,11 These findings underscore the potential of strategically crafted posts to extend the reach of scholarly work, making post-publication promotion a valuable complement to traditional dissemination pathways.
However, the creation of high-quality, engaging, and policy-compliant social media content is time-intensive and requires both subject expertise and an understanding of platform-specific communication strategies.12 In many research settings, these tasks fall to authors or editorial staff, who may face constraints in time, communication training, or language proficiency. As a result, the potential of social media for accelerating knowledge translation is often underutilized.
The rise of large language models (LLMs) offers an opportunity to streamline this process. These models can rapidly generate text tailored to specific audiences and formats, including short-form social media posts. In theory, they could assist in drafting concise, accurate, and engaging summaries of scholarly articles, thereby reducing the workload for authors and editors and expanding the reach of scientific publications.13,14,15 Early reports and limited descriptive studies suggest that AI-generated outputs can produce coherent content with minimal prompting.16 However, concerns remain regarding factual accuracy, overclaiming, adherence to journal or platform policies, and the inclusion of necessary metadata such as journal names and digital object identifiers (DOIs).
Despite the growing interest in AI-assisted science communication, there is a paucity of structured evaluations comparing how different LLMs perform when tasked with generating posts for real-world scholarly articles in public health and medicine. Previous work has typically focused on either theoretical discussions or single-model demonstrations, without standardized quality assessment, or cross-model comparisons. Moreover, little is known about whether AI models can produce policy-compliant outputs that maintain accuracy while preserving the clarity and conciseness required for social media dissemination.
The aim of this study was to conduct a blinded, structured evaluation of multiple contemporary LLMs in generating X (Twitter)-style posts promoting recently published scholarly articles in public health and medicine. Specifically, we sought to compare models’ performance on predefined quality domains: factual accuracy, clarity, policy compliance, call-to-action quality, and metadata inclusion.
METHODS
Study design
We conducted a blinded, crossed, offline evaluation of social media posts generated by several LLMs for a fixed set of scholarly articles in public health and medicine. The evaluation compared models’ ability to generate concise, accurate, and policy-compliant posts suitable for post-publication promotion on X (formerly Twitter).
Article selection
A total of 36 articles published between [01/2025] and [07/2025] were selected from The Lancet Planetary Health, Annual Review of Public Health, and The Lancet Public Health. All selected articles were open access, ensuring unrestricted availability for evaluation and reproducibility. Articles were stratified by journal and article type (original research, review, editorial). Within each stratum, articles were chosen by simple random sampling from the eligible pool. For each article metadata were recorded, including title, journal name, DOI, article type, and publication date.
Language models and configurations
We evaluated four LLMs, each in the further public release versions at the time of generation: GPT-5 (OpenAI), Gemini 2.5 Pro (Google DeepMind), Grok-3 (xAI), Perplexity Pro (Perplexity AI).
Prompts
A system prompt was used to establish the communication role, and a user prompt template specified the task requirements. The exact wording of the prompts was preserved to ensure reproducibility and is provided verbatim.
System prompt
You are a scholarly communication assistant specialized in public health and medicine. Your role is to generate accurate, concise, and policy-compliant social media posts promoting published research articles. Always remain neutral, avoid medical advice or unverified claims, and ensure posts are suitable for a professional, multidisciplinary audience.
User prompt template
Write a single post for X (formerly Twitter), ≤ 260 characters, to promote the article provided below.
Incorporate a brief version of the article title (no more than 10–12 words); the journal name; a direct DOI link in the format https://doi.org/.... ; Language: clear, accurate, and neutral. Tone: engaging but professional; suitable for a scholarly and policy-oriented audience. Policy: do not exaggerate, hype, or imply causality unless the article reports a randomized controlled trial. Hashtags: include at most two, selected from relevant and widely used scholarly tags (e.g., #PublicHealth, #MedTwitter, #OpenScience). Do not invent hashtags. Alt-text: provide one suggested alt-text line if the article contains a figure or table (describe content, not interpretation). Constraints: strictly avoid patient-specific advice, prescriptive recommendations, or fabricated statistics. Article information: Title: <insert title> Journal: <insert journal name> DOI: <insert DOI> Article type: <Original Research / Review / Editorial> Open Access: <Yes/No>
Posts exceeding the 260-character limit or violating explicit policy instructions were regenerated once; these instances were documented and flagged.
Blinding and presentation to rater
Generated posts were anonymized by removing metadata indicating their model of origin. Posts were presented in random order using a computer-generated sequence, with no two posts derived from the same article appearing consecutively.
Evaluation rubric
A single rater, blinded to model identity, assessed each generated post using a structured five-item rubric, with each domain scored from 0–2 points (maximum total score = 10). The rater held PhD-level training in Medicine, had prior experience in rubric-based content evaluation, and received specific instruction on applying the study’s multi-domain scoring framework. The rubric domains were:
1) Factual accuracy — absence of invented data; correct inclusion of journal name and DOI.
2) Clarity and relevance — clear, concise, and directly aligned with the article content.
3) Policy compliance — no medical advice, cure claims, or hype inconsistent with the study design.
4) Call-to-action quality — encourages readers to engage with the article without exaggeration.
5) Structure and metadata — inclusion of journal name and DOI, ≤ 2 hashtags, and adherence to the 260-character limit.
In addition to rubric scores, the rater documented three secondary features for each post: character count, number of hashtags, and readability grade level calculated using the Flesch–Kincaid Grade Level (FKGL) formula.
The primary outcome was the mean total quality score (0–10) per post. Secondary outcomes included: 1) the proportion of posts with policy violations, 2) the proportion including both journal name and DOI, 3) character count (reported as median with interquartile range, Q1–Q3), 4) hashtag count (median, Q1–Q3), and 5) readability grade level (FKGL; median, Q1–Q3).
Statistical analysis
All statistical analyses were conducted in IBM SPSS Statistics (version 27; IBM Corp., Armonk, NY, USA). Descriptive statistics were first generated for each outcome, with results summarized as means and 95% confidence intervals (CIs) for normally distributed quality scores, and medians with interquartile ranges (Q1–Q3) for secondary outcomes (character count, hashtag number, and readability scores).
To evaluate differences in primary outcomes (domain-specific quality scores and total quality score), univariate general linear models (GLMs) were fitted with “Model” (GPT-5, Gemini 2.5 Pro, Grok-3, Perplexity Pro) as a fixed factor. Effect sizes were expressed as partial η2 and interpreted according to conventional thresholds (small ≥ 0.01, medium ≥ 0.06, large ≥ 0.14). Pairwise comparisons between models were conducted using Bonferroni-adjusted post hoc tests to control for Type I error inflation. Accuracy scores exhibited no variance (all models scored 2.0), precluding inferential testing for this domain.
For secondary outcomes, character count, hashtag use, and readability (Flesch–Kincaid Grade Level, FKGL) were assessed for normality. Because distributional assumptions were not consistently met, non-parametric tests were applied. A Kruskal–Wallis H test compared median values across models, followed by pairwise Mann–Whitney U tests with Bonferroni correction. Medians and interquartile ranges (Q1–Q3) are reported to provide robust measures of central tendency and dispersion for these skewed variables. All statistical tests were conducted as two-tailed, with P < 0.05 considered the threshold for significance.
RESULTS
A univariate GLM demonstrated a significant effect of model on total quality scores, P < 0.001, indicating that the choice of LLM explained a substantial proportion of variance in post quality. The mean scores were highest for GPT-5 (9.60; 95% CI, 9.40–9.81) and Perplexity Pro (9.60; 95% CI, 9.40–9.81), followed closely by Gemini 2.5 Pro (9.47; 95% CI, 9.26–9.67). Grok-3 scored markedly lower (8.80; 95% CI, 8.60–9.01). Pairwise Bonferroni-adjusted comparisons confirmed that Grok-3 performed significantly worse than GPT-5 (mean difference = –0.80, P < 0.001), Perplexity Pro (–0.80, P < 0.001), and Gemini 2.5 Pro (–0.67, P < 0.001). No statistically significant differences were observed among GPT-5, Gemini 2.5 Pro, and Perplexity Pro (Table 1). These findings highlight Grok-3 as a consistent underperformer, while the other three models achieved high-quality outputs.
Table 1. Domain-specific and total scores across large language models with post-hoc pairwise comparisons.
| Domain | GPT-5 (mean, 95% CI) | Gemini 2.5 Pro (mean, 95% CI) | Grok-3 (mean, 95% CI) | Perplexity Pro (mean, 95% CI) | ANOVA F(df) | P | Partial η2 | Pairwise comparisons (P values) |
|---|---|---|---|---|---|---|---|---|
| Accuracy | 2.00 (2.00–2.00) | 2.00 (2.00–2.00) | 2.00 (2.00–2.00) | 2.00 (2.00–2.00) | - | - | - | All identical (no post-hoc possible) |
| Clarity | 1.73 (1.60–1.87) | 1.70 (1.57–1.83) | 1.83 (1.70–1.97) | 1.90 (1.77–2.03) | 1.88 (3,56) | 0.144 | 0.091 | GPT-5 vs. Gemini 2.5 Pro: P = 0.999; GPT-5 vs. Grok-3: P = 1.000; GPT-5 vs. Perplexity Pro: P = 0.503; Gemini 2.5 Pro vs. Grok-3: P = 0.988; Gemini 2.5 Pro vs. Perplexity Pro: P = 1.000; Grok-3 vs. Perplexity Pro: P = 0.235 |
| Policy | 1.80 (1.68–1.92) | 1.80 (1.68–1.92) | 1.60 (1.48–1.72) | 1.90 (1.78–2.02) | 4.43 (3,56) | 0.007 | 0.192 | GPT-5 vs. Gemini 2.5 Pro: P = 1.000; GPT-5 vs. Grok-3: P = 0.129; GPT-5 vs. Perplexity Pro: P = 1.000; Gemini 2.5 Pro vs. Grok-3: P = 0.129; Gemini 2.5 Pro vs. Perplexity Pro: P = 1.000; Grok-3 vs. Perplexity Pro: P = 0.005 |
| CTA | 2.00 (1.89–2.11) | 1.97 (1.86–2.07) | 1.40 (1.29–1.51) | 2.00 (1.89–2.11) | 31.3 (3,56) | < 0.001 | 0.626 | GPT-5 vs. Gemini 2.5 Pro: P = 1.000; GPT-5 vs. Grok: P < 0.001; GPT-5 vs. Perplexity Pro: P = 1.000; Gemini 2.5 Pro vs. Grok-3: P < 0.001; Gemini 2.5 Pro vs. Perplexity Pro: P = 1.000; Grok-3 vs. Perplexity Pro: P < .001 |
| Structure | 2.00 (1.90–2.10) | 2.00 (1.90–2.10) | 1.97 (1.87–2.07) | 1.80 (1.70–1.90) | 3.61 (3,56) | 0.019 | 0.162 | GPT-5 vs. Gemini 2.5 Pro: P = 1.000; GPT-5 vs. Grok-3: P = 1.000; GPT-5 vs. Perplexity Pro: P = 0.041; Gemini 2.5 Pro vs. Grok-3: P = 1.000; Gemini 2.5 Pro vs. Perplexity Pro: P = 0.041; Grok-3 vs. Perplexity Pro: P = 0.138 |
| Total | 9.60 (9.40–9.80) | 9.47 (9.26–9.67) | 8.80 (8.60–9.01) | 9.60 (9.40–9.80) | 13.95 (3,56) | < 0.001 | 0.428 | GPT-5 vs. Gemini 2.5 Pro: P = 1.000; GPT-5 vs. Grok-3: P < 0.001; GPT-5 vs. Perplexity Pro: P = 1.000; Gemini 2.5 Pro vs. Grok-3: P < 0.001; Gemini 2.5 Pro vs. Perplexity Pro: P = 1.000; Grok-3 vs. Perplexity Pro: P < 0.001 |
Bold font highlights statistically significant results.
CI = confidence interval, ANOVA = analysis of variance.
All four models achieved the maximum possible accuracy score (mean = 2.00, standard deviation = 0.00), with no variation across articles or between models. Consequently, the GLM detected no differences, and all pairwise comparisons yielded zero mean differences. This indicates that the models were uniformly accurate in reproducing factual information, with perfect consistency in reporting the correct journal name and DOI. Furthermore, no policy violations were identified in any of the outputs, underscoring a uniformly high level of compliance with editorial and ethical standards (Fig. 1).
Fig. 1. Comparison of large language models on five quality domains and total score, presented as estimated marginal means.
CTA = call-to-action.
For clarity and relevance, mean scores ranged from 1.70 (Gemini 2.5 Pro) to 1.90 (Perplexity Pro) on a 2-point scale. Although descriptive variation suggested slightly higher clarity for Perplexity Pro and Grok-3, the effect did not reach statistical significance, P = 0.144, and no pairwise differences were detected. Thus, clarity did not significantly distinguish model performance.
For policy compliance, a significant main effect of model was observed, P = 0.007. Perplexity Pro (1.90; 95% CI, 1.78–2.02) scored significantly higher than Grok (1.60; 95% CI, 1.48–1.72; p = 0.005), whereas GPT-5 (1.80; 95% CI, 1.68–1.92) and Gemini 2.5 Pro (1.80; 95% CI, 1.68–1.92) performed comparably. These findings suggest that while all models adhered to policy requirements, Perplexity Pro achieved marginally stronger compliance than Grok.
For call-to-action (CTA) quality, a strong main effect was detected, P < 0.001. GPT-5 and Perplexity Pro consistently achieved maximum CTA scores, with Gemini 2.5 Pro performing comparably (Table 1). In contrast, Grok-3 performed substantially worse (1.40; 95% CI, 1.29–1.51), with pairwise comparisons confirming Grok-3’s underperformance relative to all other models (all P < 0.001).
For structural quality, a significant effect was observed (P = 0.019). GPT-5 and Gemini 2.5 Pro consistently achieved maximum structure scores (2.00; 95% CI, 1.90–2.10). Grok performed nearly equivalently (1.97; 95% CI, 1.87–2.07), whereas Perplexity Pro scored lower (1.80; 95% CI, 1.70–1.90). Pairwise comparisons confirmed Perplexity Pro was significantly weaker than GPT-5 and Gemini 2.5 Pro (P = 0.041), while Grok did not differ significantly from either.
A one-way ANOVA further revealed a significant main effect of model on post length (P < 0.001). Post hoc Bonferroni comparisons demonstrated that Grok generated significantly shorter posts than GPT-5 (P < 0.001), Gemini 2.5 Pro (P < 0.001), and Perplexity Pro (P < 0.001). No differences were observed between GPT-5, Gemini 2.5 Pro, and Perplexity Pro (all P > 0.10).
Analysis of secondary outcomes showed significant differences in character count (P < 0.001) and readability (P < 0.001), but not in hashtag use (Table 2). Grok-3 produced the shortest posts (median 194 [181–202]) compared with GPT-5, Gemini 2.5 Pro, and Perplexity Pro. All models consistently used two hashtags. Readability (Flesch–Kincaid grade level) differed markedly, with GPT-5 and Perplexity Pro generating more complex, academic-style text, while Gemini 2.5 Pro and Grok-3 produced more accessible outputs.
Table 2. Secondary outcomes of X (Twitter) post quality by large language model.
| Model | Median character count (Q1–Q3) | Median hashtags (Q1–Q3) | Median readability grade (Q1–Q3) |
|---|---|---|---|
| GPT-5 | 235 (222–248) a | 2 (2–2) | 8.9 (7.3–9.2) b |
| Gemini 2.5 Pro | 210 (204–218) c | 2 (2–2) | 5.1 (4.8–6.5) d |
| Grok-3 | 194 (181–202) | 2 (2–2) | 4.5 (3.6–6.3) e |
| Perplexity Pro | 222 (216–254) f | 2 (2–2) | 7.3 (6.5–8.8) g |
Bold font highlights statistically significant results.
Median character count: aGPT-5 > Grok-3, P < 0.001; cGemini 2.5 Pro > Grok-3, P < 0.001; fPerplexity Pro > Grok-3; P < 0.001 Readability: bGPT-5 > Gemini 2.5 Pro, P < 0.001; eGPT-5 > Grok-3, P = 0.004; dPerplexity Pro > Gemini 2.5 Pro, P < 0.001; gPerplexity Pro > Grok-3, P < 0.001.
DISCUSSION
This study provides a structured, blinded evaluation of four LLMs for generating X (Twitter) posts that promote scholarly articles in public health and medicine. Across all models, factual accuracy was perfect. Every post included the correct journal name and DOI; no policy violations were observed. It underscores the baseline reliability of LLMs for metadata reproduction under standardised prompting. The structured prompt design and reliance on metadata supplied likely minimised opportunities for factual errors. However, this finding may not generalise to more open-ended tasks. Beyond this uniform compliance, however, meaningful stylistic and structural differences were evident. GPT-5, Gemini 2.5 Pro, and Perplexity Pro consistently produced high-quality posts, whereas Grok-3 underperformed on several dimensions, particularly call-to-action quality, clarity, and post length. Notably, readability diverged across models: GPT-5 and Perplexity Pro generated more academically complex text, while Gemini 2.5 Pro and Grok-3 produced more accessible outputs. Because readability determines whether content resonates with academic, policy, or public audiences, overly complex scientific communication risks undermining accessibility and impact.17
Our results also intersect with the literature on scholarly dissemination via social media. Early work demonstrated that Twitter activity can track or even anticipate scholarly impact (altmetrics and subsequent citations), reinforcing the value of timely, well-structured posts that include durable identifiers (e.g., DOIs) and clear calls to engage.18 Within that context, the consistent metadata accuracy and strong CTA performance we observed for most models are pragmatic advantages for editors and authors seeking to amplify reach while maintaining compliance.
From an implementation perspective, these results reinforce the value of LLMs as first-draft tools for authors, editors, and communication teams. When combined with structured prompts and minimal oversight (checking DOI inclusion, journal name, hashtags, and character limits), LLMs can produce policy-compliant posts with minimal revision, potentially reducing workload and standardising quality. This aligns with recent recommendations from the World Health Organisation, which emphasise that LLMs should be used in a fit-for-purpose manner, with human oversight and clear governance safeguards rather than full automation.19 At the same time, existing literature cautions that model drift, calibration differences across vendors, and the risk of subtle “hallucinations” remain persistent challenges, particularly when models are tasked with content beyond structured metadata.20,21,22 These risks highlight the ongoing need for human-in-the-loop review.
The broader implications for post-publication promotion and accessibility of science are particularly noteworthy. Social media is now recognised as a valuable tool for disseminating research, with evidence showing that Twitter activity correlates with altmetric attention and, in some cases, future citations.23,24 Therefore, effective use of LLMs could amplify the reach of research findings, improving awareness among clinicians, policymakers, and the public. LLMs may help accelerate the uptake of new knowledge into practice and decision-making by reliably producing concise, policy-compliant posts that include durable identifiers. Moreover, by lowering the barrier to generating professional-quality promotional content, these tools may support the visibility of research from less-resourced scientific environments, where limited communication infrastructure and editorial support have historically constrained integration into global discourse.25 At the same time, the relationship between the quality of posts and actual engagement outcomes remains underexplored. Readability and CTA strength are likely contributors to engagement, but evidence suggests that simplified language alone does not guarantee comprehension or knowledge transfer.26
Strengths of this study include a standardised prompt applied to all models, blinded independent scoring, a multi-domain rubric (accuracy, clarity, policy compliance, CTA quality, structure/metadata), and evaluation on the same curated article set, enabling direct, controlled comparisons across models. At the same time, several limitations should be acknowledged. First, only a single blinded rater evaluated all posts. While blinding is commendable, the absence of multiple raters precluded inter-rater reliability testing, and this limitation may introduce subjective bias. Although the rater possessed relevant expertise and was trained in rubric-based evaluation, future research should involve multiple independent raters and cross-validation procedures to strengthen methodological rigour. Second, the study was restricted to English-language posts in public health and medicine; the implications for other languages remain uncertain. Third, the evaluation focused on offline quality scores rather than real-world engagement outcomes such as clicks, retweets, or altmetric indicators. As a result, the practical correlation between rubric-based quality and dissemination impact could not be assessed. Fourth, vendor model versioning and rapid release cycles present challenges for reproducibility, since outputs may shift over time. Fifth, readability was assessed using theFKGL, a useful but limited metric that does not capture higher-order comprehension, design, or layout features that may influence audience uptake. Finally, we must acknowledge the snapshot nature of rapidly evolving vendor models (versioning and release cadence can affect reproducibility). Considering ongoing debates about LLM safety and reliability in health contexts, rigorous reporting (model/version/date), governance consistent with WHO recommendations, and human-in-the-loop editorial oversight remain essential.19
Future work should 1) test whether cross-model differences in clarity, CTA strength, and readability translate into measurable engagement (click-throughs, saves, altmetric signals) and, more importantly, downstream knowledge uptake; 2) evaluate multilingual outputs and non-Latin scripts; 3) incorporate retrieval-augmented or template-constrained generation to mitigate risk further; 4) include multi-rater designs with inter-rater reliability testing, and 5) extend evaluation to accessibility (e.g., alt-text quality) and equity (e.g., jargon, cultural references).
In this blinded, structured evaluation of four LLMs, all systems demonstrated perfect factual accuracy and policy compliance when generating X (Twitter)-style posts for scholarly articles in public health and medicine. Nonetheless, significant differences were observed in stylistic and structural domains. GPT-5, Gemini 2.5 Pro, and Perplexity Pro consistently achieved high overall quality, while Grok-3 underperformed, particularly in call-to-action strength, clarity, and post length. Readability patterns revealed that GPT-5 and Perplexity Pro tended to produce more academically complex outputs. In contrast, Gemini 2.5 Pro and Grok-3 generated more accessible text, underscoring the importance of model selection in tailoring content to specific audiences.
Our findings suggest that LLMs are not yet replacements for human communicators but can serve as scalable, supportive tools that improve consistency and reduce the burden of post-publication promotion. For journals and institutions, selective adoption of LLMs may help optimise dissemination strategies: models that generate more complex text may be suited for professional or academic communities. At the same time, those producing accessible outputs may be preferable for broader public engagement.
ACKNOWLEDGEMENTS
The language editing of this manuscript was conducted using Grammarly.
Footnotes
Disclosure: The authors have no potential conflicts of interest to disclose.
Data Availability Statement: The data supporting the findings of this study are available from the corresponding author upon reasonable request.
- Conceptualization: Doskaliuk B.
- Data curation: Yessirkepov M, Mukhamediyarov M.
- Formal analysis: Doskaliuk B.
- Methodology: Doskaliuk B, Zimba O.
- Visualization: Yessirkepov M.
- Writing - original draft: Doskaliuk B.
- Writing - review & editing: Zimba O, Yessirkepov M, Mukhamediyarov M.
References
- 1.Peoples BK, Midway SR, Sackett D, Lynch A, Cooney PB. Twitter predicts citation rates of ecological research. PLoS One. 2016;11(11):e0166570. doi: 10.1371/journal.pone.0166570. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Luc JGY, Archer MA, Arora RC, Bender EM, Blitz A, Cooke DT, et al. Does tweeting improve citations? One-year results from the TSSMN prospective randomized trial. Ann Thorac Surg. 2021;111(1):296–300. doi: 10.1016/j.athoracsur.2020.04.065. [DOI] [PubMed] [Google Scholar]
- 3.Zimba O, Gasparyan AY. Social media platforms: a primer for researchers. Reumatologia. 2021;59(2):68–72. doi: 10.5114/reum.2021.102707. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Fox CS, Bonaca MA, Ryan JJ, Massaro JM, Barry K, Loscalzo J. A randomized trial of social media from circulation. Circulation. 2015;131(1):28–33. doi: 10.1161/CIRCULATIONAHA.114.013509. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Hawkins CM, Hunter M, Kolenic GE, Carlos RC. Social media and peer-reviewed medical journal readership: a randomized prospective trial. J Am Coll Radiol. 2017;14(5):596–602. doi: 10.1016/j.jacr.2016.12.024. [DOI] [PubMed] [Google Scholar]
- 6.Costas R, Zahedi Z, Do Wouters P. “altmetrics” correlate with citations? Extensive comparison of altmetric indicators with citations from a multidisciplinary perspective. J Assoc Inf Sci Technol. 2015;66(10):2003–2019. [Google Scholar]
- 7.Doskaliuk B, Yatsyshyn R, Klishch I, Zimba O. COVID-19 from a rheumatology perspective: bibliometric and altmetric analysis. Rheumatol Int. 2021;41(12):2091–2103. doi: 10.1007/s00296-021-04987-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Zimba O, Korkosz M, Alnaimat F, Fragoulis GE, Yessirkepov M, Baimukhamedov C, et al. Global practice guidelines in rheumatology: a cross-sectional altmetric and citation analysis. Rheumatol Int. 2025;45(6):150. doi: 10.1007/s00296-025-05899-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Bornmann L. Alternative metrics in scientometrics: A meta-analysis of research into three altmetrics. Scientometrics. 2015;103(3):1123–1144. [Google Scholar]
- 10.Zimba O, Gasparyan AY, Qumar AB. Ethics for disseminating health-related information on YouTube. J Korean Med Sci. 2024;39(7):e93. doi: 10.3346/jkms.2024.39.e93. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Gupta L, Gasparyan AY, Misra DP, Agarwal V, Zimba O, Yessirkepov M. Information and misinformation on COVID-19: a cross-sectional survey study. J Korean Med Sci. 2020;35(27):e256. doi: 10.3346/jkms.2020.35.e256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zimba O, Radchenko O, Strilchuk L. Social media for research, education and practice in rheumatology. Rheumatol Int. 2020;40(2):183–190. doi: 10.1007/s00296-019-04493-4. [DOI] [PubMed] [Google Scholar]
- 13.Doskaliuk B, Zimba O. Beyond the keyboard: academic writing in the era of ChatGPT. J Korean Med Sci. 2023;38(26):e207. doi: 10.3346/jkms.2023.38.e207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Doskaliuk B, Zimba O, Yessirkepov M, Klishch I, Yatsyshyn R. Artificial intelligence in peer review: enhancing efficiency while preserving integrity. J Korean Med Sci. 2025;40(7):e92. doi: 10.3346/jkms.2025.40.e92. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Pidvalna U, Holota S, Kryshchyshyn-Dylevych A, Telishevska U, Lonchyna V, Zayachkivska O, et al. The role of social media in shaping scientific medical journals. Proc Shevchenko Sci Soc Med Sci. 2025;77(1):1–6. [Google Scholar]
- 16.Chen B, Zhang Z, Langrené N, Zhu S. Unleashing the potential of prompt engineering for large language models. Patterns (N Y) 2025;6(6):101260. doi: 10.1016/j.patter.2025.101260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Plavén-Sigray P, Matheson GJ, Schiffler BC, Thompson WH. The readability of scientific texts is decreasing over time. eLife. 2017;6:e27725. doi: 10.7554/eLife.27725. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Eysenbach G. Can tweets predict citations? Metrics of social impact based on Twitter and correlation with traditional metrics of scientific impact. J Med Internet Res. 2011;13(4):e123. doi: 10.2196/jmir.2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.World Health Organizaion. WHO guidance on ethics and governance of AI for health. 2023. [Updated 2023]. [Accessed August 27, 2025]. https://www.who.int/publications/i/item/9789240029200 .
- 20.Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198. doi: 10.1371/journal.pdig.0000198. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Thorp HH. ChatGPT is fun, but not an author. Science. 2023;379(6630):313. doi: 10.1126/science.adg7879. [DOI] [PubMed] [Google Scholar]
- 22.Sanderson K. GPT-4 is here: what scientists think. Nature. 2023;615(7954):773. doi: 10.1038/d41586-023-00816-5. [DOI] [PubMed] [Google Scholar]
- 23.Karim MA. Influence of social media on the dissemination and uptake of cardiology research. Cureus. 2025;17(6):e86509. doi: 10.7759/cureus.86509. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Gholampour S, Lim WM, Lund BD, Noruzi A, Elahi A, Saboury AA, et al. Does social media contribute to research impact? An altmetric study of highly-cited marketing research. Total Qual Manage Bus Excell. 2024;35(13-14):1671–1701. [Google Scholar]
- 25.Doskaliuk B, Zimba O, Yatsyshyn R, Kovalenko V. Rheumatology in Ukraine. Rheumatol Int. 2020;40(2):175–182. doi: 10.1007/s00296-019-04504-4. [DOI] [PubMed] [Google Scholar]
- 26.Kirchner GJ, Kim RY, Weddle JB, Bible JE. Can artificial intelligence improve the readability of patient education materials? Clin Orthop Relat Res. 2023;481(11):2260–2267. doi: 10.1097/CORR.0000000000002668. [DOI] [PMC free article] [PubMed] [Google Scholar]


