Table 3.
Diagnostic accuracy on the publication set
| Dataset | Model | Vanilla LLMs | Cypher RAG LLMs | Vector RAG LLMs |
|---|---|---|---|---|
|
Publication set selective test |
GPT-3.5-turbo | 67.60 | 92.40 | 90.60 |
| GPT-4-turbo | 90.40 | 97.00 | 95.00 | |
| GPT-4o | 90.20 | 95.40 | 95.00 | |
| Claude-3-opus | 85.20 | 94.40 | 94.00 | |
| Claude-3-sonnet | 63.60 | 91.20 | 92.20 | |
| Claude-3-haiku | 75.60 | 90.60 | 91.00 | |
| Gemini-pro | 54.40 | 47.60 | 92.40 | |
| LLaMA-70b | 72.80 | 94.40 | 93.20 | |
| Average | 74.98 | 87.88 | 92.93 | |
|
Publication set non-selective test |
GPT-3.5-turbo | 49.00 | 87.60 | 86.20 |
| GPT-4-turbo | 67.40 | 95.00 | 92.40 | |
| GPT-4o | 69.20 | 95.00 | 93.40 | |
| Claude-3-opus | 71.20 | 93.40 | 92.80 | |
| Claude-3-sonnet | 54.00 | 73.80 | 81.80 | |
| Claude-3-haiku | 54.60 | 71.40 | 81.20 | |
| Gemini-pro | 25.20 | 1.20 | 86.00 | |
| LLaMA-70b | 46.20 | 88.20 | 87.80 | |
| Average | 54.60 | 75.70 | 87.70 |
Results demonstrate that both Cypher RAG and Vector RAG significantly outperform vanilla LLMs, improving average accuracy by 12.90% and 17.95% in selective tests, and 21.10% and 33.10% in non-selective tests, respectively. The bold values represent the best performance among these three methods.