Table 3.
Comparison between closed- and open-source vanilla models with zero-shot prompts.
| Models | Prompt | Top 5 | Top 10 | |||||||||||||||||
|
|
|
F1-score (relaxed), mean (SD) | F1-score (Jaccard), mean (SD) | MRRa (SD), relaxed | MRR (SD), Jaccard | F1-score (relaxed), mean (SD) | F1-score (Jaccard), mean (SD) | MRR, SD (relaxed) | MRR, SD (Jaccard) | |||||||||||
| Closed-source LLMs b | ||||||||||||||||||||
|
|
GPT-5.2 | |||||||||||||||||||
|
|
|
General | 0.492 (0.068)c | 0.307 (0.071) | 0.578 (0.045)c | 0.476 (0.124)c | 0.496 (0.058)c | 0.31 (0.077) | 0.578 (0.045)c | 0.467 (0.117)c | ||||||||||
|
|
|
Structured | 0.332 (0.077) | 0.206 (0.048) | 0.513 (0.153) | 0.359 (0.123) | 0.336 (0.073) | 0.211 (0.047) | 0.513 (0.153) | 0.359 (0.123) | ||||||||||
|
|
GPT-5-mini | |||||||||||||||||||
|
|
|
General | 0.46 (0.073) | 0.129 (0.048) | 0.561 (0.156) | 0.196 (0.11) | 0.505 (0.063) | 0.145 (0.045) | 0.552 (0.151) | 0.196 (0.11) | ||||||||||
|
|
|
Structured | 0.387 (0.055) | 0.12 (0.046) | 0.575 (0.165) | 0.24 (0.17) | 0.387 (0.075) | 0.127 (0.052) | 0.566 (0.153) | 0.24 (0.17) | ||||||||||
| Open-source LLMs | ||||||||||||||||||||
|
|
Mistral 7B | |||||||||||||||||||
|
|
|
General | 0.372 (0.085) | 0.323 (0.081) | 0.451 (0.175) | 0.453 (0.172) | 0.401 (0.119) | 0.345 (0.115) | 0.432 (0.167) | 0.444 (0.177) | ||||||||||
|
|
|
Structured | 0.351 (0.083) | 0.243 (0.056) | 0.582 (0.108) | 0.455 (0.112) | 0.349 (0.081) | 0.236 (0.055) | 0.573 (0.106) | 0.455 (0.112) | ||||||||||
|
|
Llama 3.1 8B | |||||||||||||||||||
|
|
|
General | 0.347 (0.054) | 0.333 (0.055) | 0.426 (0.112) | 0.392 (0.108) | 0.368 (0.063) | 0.35 (0.056) | 0.422 (0.113) | 0.392 (0.108) | ||||||||||
|
|
|
Structured | 0.371 (0.056) | 0.188 (0.051) | 0.521 (0.14) | 0.318 (0.112) | 0.358 (0.059) | 0.186 (0.055) | 0.517 (0.139) | 0.318 (0.112) | ||||||||||
|
|
BioMistral 7B | |||||||||||||||||||
|
|
|
General | 0.21 (0.076) | 0.017 (0.032) | 0.353 (0.141) | 0.033 (0.071) | 0.211 (0.078) | 0.02 (0.04) | 0.353 (0.141) | 0.033 (0.071) | ||||||||||
|
|
|
Structured | 0.304 (0.121) | 0.017 (0.018) | 0.507 (0.152) | 0.056 (0.056) | 0.304 (0.123) | 0.017 (0.018) | 0.507 (0.152) | 0.056 (0.056) | ||||||||||
|
|
DeepSeek 8B | |||||||||||||||||||
|
|
|
General | 0.42 (0.077) | 0.383 (0.053)c | 0.416 (0.113) | 0.398 (0.13) | 0.443 (0.077) | 0.409 (0.068)c | 0.414 (0.113) | 0.397 (0.13) | ||||||||||
|
|
|
Structured | 0.326 (0.026) | 0.211 (0.053) | 0.467 (0.131) | 0.4 (0.158) | 0.328 (0.048) | 0.218 (0.059) | 0.467 (0.131) | 0.4 (0.158) | ||||||||||
aMRR: mean reciprocal rank.
bHighest score for each metric (column) among the compared models.
cLLM: large language model.