Skip to main content
. 2026 Jul 17;5:e75561. doi: 10.2196/75561

Table 4.

Comparison between closed- and open-source vanilla models with few-shot prompts.

Models and prompts Top 5 Top 10

F1-score, mean (SD), relaxed F1-score, mean (SD), Jaccard MRRa (SD), relaxed MRR (SD), Jaccard F1-score, mean (SD), relaxed F1-score, mean (SD), Jaccard MRR (SD), relaxed MRR (SD), Jaccard
Closed-source LLMsb

GPT-5.2


General 0.475 (0.068)c 0.323 (0.06) 0.561 (0.098)c 0.536 (0.092)c 0.478 (0.07)c 0.326 (0.058) 0.561 (0.098)c 0.536 (0.092)c


Structured 0.354 (0.069) 0.233 (0.054) 0.536 (0.174) 0.387 (0.166) 0.364 (0.076) 0.237 (0.052) 0.536 (0.173) 0.387 (0.165)

GPT-5-mini


General 0.475 (0.077) 0.243 (0.087) 0.543 (0.138) 0.4 (0.133) 0.492 (0.056) 0.245 (0.069) 0.543 (0.138) 0.4 (0.133)


Structured 0.348 (0.068) 0.161 (0.041) 0.508 (0.14) 0.321 (0.172) 0.347 (0.087) 0.163 (0.052) 0.495 (0.131) 0.323 (0.174)
Open-source LLMs

Mistral 7B


General 0.387 (0.046) 0.338 (0.039) 0.498 (0.141) 0.503 (0.11) 0.383 (0.05) 0.335 (0.045) 0.498 (0.141) 0.503 (0.11)


Structured 0.363 (0.067) 0.335 (0.056) 0.534 (0.104) 0.533 (0.101) 0.356 (0.073) 0.332 (0.065) 0.534 (0.104) 0.533 (0.101)

Llama 3.1 8B


General 0.372 (0.078) 0.335 (0.093) 0.521 (0.145) 0.557 (0.126) 0.369 (0.073) 0.333 (0.096) 0.521 (0.145) 0.547 (0.132)


Structured 0.368 (0.116) 0.346 (0.111)c 0.442 (0.066) 0.502 (0.109) 0.374 (0.126) 0.353 (0.119)c 0.442 (0.066) 0.502 (0.109)

BioMistral 7B


General 0.208 (0.063) 0.08 (0.036) 0.459 (0.115) 0.3 (0.132) 0.207 (0.067) 0.081 (0.033) 0.459 (0.115) 0.3 (0.132)


Structured 0.227 (0.085) 0.098 (0.042) 0.435 (0.072) 0.318 (0.136) 0.225 (0.084) 0.097 (0.041) 0.435 (0.072) 0.318 (0.136)

DeepSeek 8B


General 0.344 (0.058) 0.335 (0.058) 0.494 (0.125) 0.485 (0.115) 0.344 (0.066) 0.334 (0.064) 0.485 (0.116) 0.476 (0.102)


Structured 0.327 (0.105) 0.299 (0.116) 0.48 (0.173) 0.466 (0.188) 0.332 (0.105) 0.304 (0.113) 0.48 (0.173) 0.466 (0.188)

aMRR: mean reciprocal rank.

bHighest score for each metric (column) among the compared models.

cLLM: large language model.