Skip to main content
. 2026 Jul 7;26:1889. doi: 10.1186/s12903-026-09174-w

Table 1.

Comparison of correct response rates within and between models

ChatGPT 5.2 Gemini 3 DeepSeek V3.2 Test statistics p
Value (%) p CI 95%-Wald
binominal
Standard
error
Value (%) p CI 95%-Wald
binominal
Standard
error
Value (%) p CI 95%-Wald
binominal
Standard
error
Group A 108 (90) aA 84,632: 95,368 0,027 106 (88,3) abA 82,590: 94,077 0,029 93 (77,5) b 70,029: 84,971 0,038 8,253 0,016x
Group B 107 (89,2) A 83,606: 94,728 0,028 98 (81,7) A 74,744: 88,590 0,035 95 (79,2) 71,900: 86,433 0,037 4,833 0,109x
Group C 104 (86,7) A 80,585: 92,749 0,031 107 (89,2) A 83,606: 94,728 0,028 106 (88,3) 82,590: 94,077 0,029 0,387 0,880x
Group D 119 (99,2) aB 97,540: 100,000 0,008 120 (100) aB 00,000: 100,000 0,000 103 (85,8) b 79,594: 92,072 0,032 29,225 < 0,001x
Group E 94 (78,3) aA 70,962: 85,704 0,038 108 (90) bA 84,632: 95,368 0,027 103 (85,8) ab 79,594: 92,072 0,032 6,299 0,048x
Group F 108 (90) aA 84,632: 95,368 0,027 108 (90) aA 84,632: 95,368 0,027 91 (75,8) b 68,174: 83,493 0,039 11,845 0,001x
Total 640 (88,9) a 86,593: 91,184 0,012 647 (89,9) a 87,656: 92,066 0,011 591 (82,1) b 79,282: 84,884 0,014 22,783 < 0,001y
Test statistics 30,936 29,824 11,023
p < 0,001x < 0,001x 0,058x

x Fisher’s Exact Test with Monte Carlo correction

y Pearson’s chi-square test; frequency (percentage)

a–b No significant difference between language models sharing the same letter

A–B No significant difference between groups sharing the same letter. Bold values indicate statistically significant differences (p < 0.05).