Skip to main content
. 2025 Aug 24;8:543. doi: 10.1038/s41746-025-01955-x

Table 3.

Diagnostic accuracy on the publication set

Dataset Model Vanilla LLMs Cypher RAG LLMs Vector RAG LLMs

Publication set

selective test

GPT-3.5-turbo 67.60 92.40 90.60
GPT-4-turbo 90.40 97.00 95.00
GPT-4o 90.20 95.40 95.00
Claude-3-opus 85.20 94.40 94.00
Claude-3-sonnet 63.60 91.20 92.20
Claude-3-haiku 75.60 90.60 91.00
Gemini-pro 54.40 47.60 92.40
LLaMA-70b 72.80 94.40 93.20
Average 74.98 87.88 92.93

Publication set

non-selective test

GPT-3.5-turbo 49.00 87.60 86.20
GPT-4-turbo 67.40 95.00 92.40
GPT-4o 69.20 95.00 93.40
Claude-3-opus 71.20 93.40 92.80
Claude-3-sonnet 54.00 73.80 81.80
Claude-3-haiku 54.60 71.40 81.20
Gemini-pro 25.20 1.20 86.00
LLaMA-70b 46.20 88.20 87.80
Average 54.60 75.70 87.70

Results demonstrate that both Cypher RAG and Vector RAG significantly outperform vanilla LLMs, improving average accuracy by 12.90% and 17.95% in selective tests, and 21.10% and 33.10% in non-selective tests, respectively. The bold values represent the best performance among these three methods.