Table 1.
Comparison of correct response rates within and between models
| ChatGPT 5.2 | Gemini 3 | DeepSeek V3.2 | Test statistics | p | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Value (%) p | CI 95%-Wald binominal |
Standard error |
Value (%) p | CI 95%-Wald binominal |
Standard error |
Value (%) p | CI 95%-Wald binominal |
Standard error |
|||
| Group A | 108 (90) aA | 84,632: 95,368 | 0,027 | 106 (88,3) abA | 82,590: 94,077 | 0,029 | 93 (77,5) b | 70,029: 84,971 | 0,038 | 8,253 | 0,016x |
| Group B | 107 (89,2) A | 83,606: 94,728 | 0,028 | 98 (81,7) A | 74,744: 88,590 | 0,035 | 95 (79,2) | 71,900: 86,433 | 0,037 | 4,833 | 0,109x |
| Group C | 104 (86,7) A | 80,585: 92,749 | 0,031 | 107 (89,2) A | 83,606: 94,728 | 0,028 | 106 (88,3) | 82,590: 94,077 | 0,029 | 0,387 | 0,880x |
| Group D | 119 (99,2) aB | 97,540: 100,000 | 0,008 | 120 (100) aB | 00,000: 100,000 | 0,000 | 103 (85,8) b | 79,594: 92,072 | 0,032 | 29,225 | < 0,001x |
| Group E | 94 (78,3) aA | 70,962: 85,704 | 0,038 | 108 (90) bA | 84,632: 95,368 | 0,027 | 103 (85,8) ab | 79,594: 92,072 | 0,032 | 6,299 | 0,048x |
| Group F | 108 (90) aA | 84,632: 95,368 | 0,027 | 108 (90) aA | 84,632: 95,368 | 0,027 | 91 (75,8) b | 68,174: 83,493 | 0,039 | 11,845 | 0,001x |
| Total | 640 (88,9) a | 86,593: 91,184 | 0,012 | 647 (89,9) a | 87,656: 92,066 | 0,011 | 591 (82,1) b | 79,282: 84,884 | 0,014 | 22,783 | < 0,001y |
| Test statistics | 30,936 | 29,824 | 11,023 | ||||||||
| p | < 0,001x | < 0,001x | 0,058x | ||||||||
x Fisher’s Exact Test with Monte Carlo correction
y Pearson’s chi-square test; frequency (percentage)
a–b No significant difference between language models sharing the same letter
A–B No significant difference between groups sharing the same letter. Bold values indicate statistically significant differences (p < 0.05).