Table 4.
Comparison between closed- and open-source vanilla models with few-shot prompts.
| Models and prompts | Top 5 | Top 10 | |||||||||||||||||
|
|
F1-score, mean (SD), relaxed | F1-score, mean (SD), Jaccard | MRRa (SD), relaxed | MRR (SD), Jaccard | F1-score, mean (SD), relaxed | F1-score, mean (SD), Jaccard | MRR (SD), relaxed | MRR (SD), Jaccard | |||||||||||
| Closed-source LLMsb | |||||||||||||||||||
|
|
GPT-5.2 | ||||||||||||||||||
|
|
|
General | 0.475 (0.068)c | 0.323 (0.06) | 0.561 (0.098)c | 0.536 (0.092)c | 0.478 (0.07)c | 0.326 (0.058) | 0.561 (0.098)c | 0.536 (0.092)c | |||||||||
|
|
|
Structured | 0.354 (0.069) | 0.233 (0.054) | 0.536 (0.174) | 0.387 (0.166) | 0.364 (0.076) | 0.237 (0.052) | 0.536 (0.173) | 0.387 (0.165) | |||||||||
|
|
GPT-5-mini | ||||||||||||||||||
|
|
|
General | 0.475 (0.077) | 0.243 (0.087) | 0.543 (0.138) | 0.4 (0.133) | 0.492 (0.056) | 0.245 (0.069) | 0.543 (0.138) | 0.4 (0.133) | |||||||||
|
|
|
Structured | 0.348 (0.068) | 0.161 (0.041) | 0.508 (0.14) | 0.321 (0.172) | 0.347 (0.087) | 0.163 (0.052) | 0.495 (0.131) | 0.323 (0.174) | |||||||||
| Open-source LLMs | |||||||||||||||||||
|
|
Mistral 7B | ||||||||||||||||||
|
|
|
General | 0.387 (0.046) | 0.338 (0.039) | 0.498 (0.141) | 0.503 (0.11) | 0.383 (0.05) | 0.335 (0.045) | 0.498 (0.141) | 0.503 (0.11) | |||||||||
|
|
|
Structured | 0.363 (0.067) | 0.335 (0.056) | 0.534 (0.104) | 0.533 (0.101) | 0.356 (0.073) | 0.332 (0.065) | 0.534 (0.104) | 0.533 (0.101) | |||||||||
|
|
Llama 3.1 8B | ||||||||||||||||||
|
|
|
General | 0.372 (0.078) | 0.335 (0.093) | 0.521 (0.145) | 0.557 (0.126) | 0.369 (0.073) | 0.333 (0.096) | 0.521 (0.145) | 0.547 (0.132) | |||||||||
|
|
|
Structured | 0.368 (0.116) | 0.346 (0.111)c | 0.442 (0.066) | 0.502 (0.109) | 0.374 (0.126) | 0.353 (0.119)c | 0.442 (0.066) | 0.502 (0.109) | |||||||||
|
|
BioMistral 7B | ||||||||||||||||||
|
|
|
General | 0.208 (0.063) | 0.08 (0.036) | 0.459 (0.115) | 0.3 (0.132) | 0.207 (0.067) | 0.081 (0.033) | 0.459 (0.115) | 0.3 (0.132) | |||||||||
|
|
|
Structured | 0.227 (0.085) | 0.098 (0.042) | 0.435 (0.072) | 0.318 (0.136) | 0.225 (0.084) | 0.097 (0.041) | 0.435 (0.072) | 0.318 (0.136) | |||||||||
|
|
DeepSeek 8B | ||||||||||||||||||
|
|
|
General | 0.344 (0.058) | 0.335 (0.058) | 0.494 (0.125) | 0.485 (0.115) | 0.344 (0.066) | 0.334 (0.064) | 0.485 (0.116) | 0.476 (0.102) | |||||||||
|
|
|
Structured | 0.327 (0.105) | 0.299 (0.116) | 0.48 (0.173) | 0.466 (0.188) | 0.332 (0.105) | 0.304 (0.113) | 0.48 (0.173) | 0.466 (0.188) | |||||||||
aMRR: mean reciprocal rank.
bHighest score for each metric (column) among the compared models.
cLLM: large language model.