Skip to main content
. 2025 Aug 24;8:543. doi: 10.1038/s41746-025-01955-x

Table 5.

Standard deviation (σ) of diagnostic accuracy across temperature settings (0.0–1.0) for LLMs

Dataset Model Vanilla σ Cypher RAG σ Vector RAG σ RAG σ Reduction (%)
Publication set selective test GPT-3.5-turbo 2.53 1.47 1.60 39.24%
GPT-4-turbo 2.81 1.17 1.03 60.83%
GPT-4o 2.35 0.75 0.75 67.90%
Publication set non-selective test GPT-3.5-turbo 2.64 1.26 2.50 28.62%
GPT-4-turbo 4.55 1.41 0.98 73.63%
GPT-4o 2.26 0.41 0.75 74.29%
GMDB set selective test GPT-3.5-turbo 2.53 1.86 1.47 34.11%
GPT-4-turbo 2.35 0.89 0.75 64.88%
GPT-4o 2.28 1.37 1.03 47.40%
GMDB set non-selective test GPT-3.5-turbo 2.64 1.26 1.17 53.89%
GPT-4-turbo 2.26 1.03 0.75 60.47%
GPT-4o 2.07 1.41 0.98 41.97%

Publication set results show Cypher-RAG achieves the lowest variability (σ = 0.41 for GPT-4o in non-selective), while Vector-RAG demonstrates consistent stability (σ ≤ 1.60 in selective). GMDB set exhibits similar patterns. All RAG LLMs show significantly reduced variance compared to Vanilla LLMs (average RAG σ Reduction: 53.94%), suggesting that they perform with less effect of temperature parameter.