Table 5.
Standard deviation (σ) of diagnostic accuracy across temperature settings (0.0–1.0) for LLMs
| Dataset | Model | Vanilla σ | Cypher RAG σ | Vector RAG σ | RAG σ Reduction (%) |
|---|---|---|---|---|---|
| Publication set selective test | GPT-3.5-turbo | 2.53 | 1.47 | 1.60 | 39.24% |
| GPT-4-turbo | 2.81 | 1.17 | 1.03 | 60.83% | |
| GPT-4o | 2.35 | 0.75 | 0.75 | 67.90% | |
| Publication set non-selective test | GPT-3.5-turbo | 2.64 | 1.26 | 2.50 | 28.62% |
| GPT-4-turbo | 4.55 | 1.41 | 0.98 | 73.63% | |
| GPT-4o | 2.26 | 0.41 | 0.75 | 74.29% | |
| GMDB set selective test | GPT-3.5-turbo | 2.53 | 1.86 | 1.47 | 34.11% |
| GPT-4-turbo | 2.35 | 0.89 | 0.75 | 64.88% | |
| GPT-4o | 2.28 | 1.37 | 1.03 | 47.40% | |
| GMDB set non-selective test | GPT-3.5-turbo | 2.64 | 1.26 | 1.17 | 53.89% |
| GPT-4-turbo | 2.26 | 1.03 | 0.75 | 60.47% | |
| GPT-4o | 2.07 | 1.41 | 0.98 | 41.97% |
Publication set results show Cypher-RAG achieves the lowest variability (σ = 0.41 for GPT-4o in non-selective), while Vector-RAG demonstrates consistent stability (σ ≤ 1.60 in selective). GMDB set exhibits similar patterns. All RAG LLMs show significantly reduced variance compared to Vanilla LLMs (average RAG σ Reduction: 53.94%), suggesting that they perform with less effect of temperature parameter.