Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Jan 27;16:6396. doi: 10.1038/s41598-026-37115-8

Expert assignment system based on natural language processing for Marie Sklodowska-Curie actions

Elena Álvarez-García 1, Daniel García-Costa 1, Ilse De Waele 2, Ana Marusic 3, Francisco Grimaldo 1,✉
PMCID: PMC12909779  PMID: 41588151

Abstract

Assigning experts to project proposals is a critical process in research evaluation. Traditional Information Retrieval (IR) methods, such as the Single Evaluation Platform (SEP) used by the European Research Executive Agency, automatically assign experts based on keyword matching, but these assignments are subsequently reviewed and corrected by Vice Chairs (VCs) to ensure suitability. To address the limitations of keyword-based systems and enhance semantic relevance, we developed a novel expert assignment system leveraging Natural Language Processing with Large Language Models (LLMs). Our approach integrates dynamic retrieval of expert publications via ORCID with GALACTICA, a specialized scientific LLM, to compute fine-grained semantic similarity between publications and proposal abstracts. Using a dataset of 48 experts and 181 proposals, we evaluated three similarity aggregation strategies: Sum, Product, and Maximum. The Maximum similarity approach most closely replicated VCs-reviewed assignments, achieving an AUC of 0.82, significantly outperforming the traditional SEP system (AUC = 0.75), Sum (AUC = 0.69), and Product (AUC = 0.57). These results demonstrate that focusing on the single most relevant match effectively captures human decision-making, highlighting the potential of LLM-based semantic matching to provide a more accurate and scalable alternative to existing IR systems. Furthermore, unlike SEP’s discrete affinity scores, our aggregation strategies produce highly discriminative, fine-grained ratings, allowing for more nuanced differentiation among candidate experts.

Subject terms: Computer science, Scientific data, Software

Introduction

The assignment of reviewers to project proposals (RAP) is a core component of project evaluation and scientific assessment. Its objective is to ensure that proposals are reviewed by experts whose knowledge is well aligned with the project topics, guaranteeing accurate and informed reviews. Traditionally, this task has been carried out manually or through keyword matching, which is based on the search for matching terms between proposals and reviewer profiles. However, these approaches are often ineffective and tend to be biased, especially when dealing with large volumes of proposals or when evaluating projects in newly emerging fields due to the insufficient set of predefined keywords1.

Marie Skłodowska Curie Actions (MSCA) postdoctoral fellowships are an example of the challenges founded by matching experts in large-scale competitive environments. The MSCA program, which is part of the Horizon Europe framework program, is a European Union initiative that supports career development and excellence in research2. With a budget of €6.6 billion for the period 2021-2027, MSCA aims to enhance the international, interdisciplinary, and inter-sectoral mobility of researchers3,4. Each year, between 8,000 and 10,000 proposals are submitted, requiring evaluation by multiple scientific experts.

The evaluation process is managed by the European Research Executive Agency (REA) using the Single Evaluation Platform (SEP), currently known as the Commission’s Evaluation Tool or Evaluation Service, which automatically assigns experts to proposals. Although the internal workings of the SEP algorithm are not publicly documented, the official guidelines suggest that it is based mainly on the correspondence between keywords in the proposal descriptions and the profiles of experts’ profiles5. The SEP output is an initial list of findings (LIF), which includes three experts automatically assigned to each proposal, together with a ranked list of up to 25 alternative experts and their corresponding Affinity scores, represented as a discrete numerical value, reflecting the estimated suitability of the system for the proposal.

Nevertheless, these assignments are later reviewed by a panel of Vice Chairs (VCs), who are experienced assessors responsible for ensuring the suitability of the matches. The final LIF represents the version revised after the VCs’ intervention. In practice, VCs make significant manual corrections to the assignments generated by the SEP algorithm, which implies a substantial workload and additional work time.

A quick check on the full 2020 LIF panel confirms the magnitude of this intervention: Total changes: 40.18% changed; Changes per proposal: none (19.96%), one (43.08%), two (33.41%), all three (3.55%). This high manual adjustment rate highlights the limitations of the SEP algorithm and the heavy dependence on the expertise of the Vice Chairs’ to guaranty accurate matching. The process is not only time-consuming, but also lacks transparency and scalability, especially given the increasing number of submissions.

Based on these observations, our study proposes an expert assignment system designed to streamline and automate this process, reducing the workload of vise chairs while maintaining the quality of assignments. Our work makes three main contributions:

  • ORCID-integrated retrieval We introduce a system that collects experts’ publications through ORCID, offering a dynamic and comprehensive representation of their expertise.

  • LLM-based semantic similarity Using GALACTICA, a large language model trained specifically on scientific text, we calculate semantic similarity between expert publications and project proposals. This approach reduces the dependence on predefined keywords and removes the limitations of systems such as SubSift6 and the Microsoft conference management toolkit.

  • Evaluation of aggregation strategies We compare three strategies (Sum, Product, and Maximum) to aggregate similarity scores across expert publications, and we provide empirical evidence that the Maximum strategy best aligns with the current semi-automatic assignments performed in the MSCA process.

Related work

The problem of reviewer assignment (RAP) has been addressed from multiple perspectives in the literature. Early approaches were based on keywords, using predefined lists or checklists to match reviewers with manuscripts7–10. Although these methods are effective in some cases, they require static and precise representations of expertise, which often fail when dealing with new or interdisciplinary topics.

In addition to keyword-based approaches, RAP has also been designed as a classification task, with the goal of predicting if a particular reviewer is appropriate for a given manuscript. Zhao et al.11 proposed a classification-based framework that uses Word Mover distance to calculate textual similarity between articles and reviewers, in combination with constructive coverage algorithms for simultaneous reviewer and manuscript assignment. Similarly, Zhang et al.12 modeled RAP as a multi-label classification problem, predicting multiple matches of potential reviewers based on a hierarchical representation of the expertise.

Other studies try to deal with RAP as a problem of thematic coverage, with the aim of ensuring that the selected reviewers would cover all relevant topics in a submission. Kou et al.13 explored weighted coverage strategies with topic-based optimization. This highlights the balance between the individual expertise of the reviewers and the overall coverage of the manuscript’s subject matter.

The most recent approaches to RAP adopt information retrieval (IR) paradigms and recommendation systems. They deal with manuscripts as queries and reviewers as documents in a large collection. The task consists of identifying the most relevant reviewers whose previous publications are more related to the work submitted. Some research incorporates other information among reviewer publications topics to assess suitability, such as author’s publication history, co-authorship networks, and distribution of subject expertise9,14. Yue et al.15 introduced reviewer assignment based on factors such as research activity, academic affiliations, and conflicts of interest. Similarly, Marco-Tordera et al.16 proposed cognitive similarity methods based on bibliometric links and text-based measures. These methods increase the quality of assignment by combining textual similarity with structural and relational information between researchers.

Recently, large language models (LLMs) have emerged as powerful tools for RAP and related retrieval tasks. Systems such as SemRank17 and Mitrov et al.18 improve the retrieval of scientific papers and the ranking of documents by combining embeddings with LLM-guided prompts, which can be adapted to reviewer assignment at scale. For instance, studies in conference management systems have utilized LLM-based recommendation systems for the assignment problem, validating the efficacy of transformer-based similarity for expertise matching19. Moreover, researchers have explored combining content-based methods with LLM-generated information to enrich expert profiles and enhance reviewer recommendation accuracy20. At the same time, there is a growing amount of work on the application of AI to peer review in different disciplines, with initiatives on an organizational level (e.g., the AAAI’s AI-based peer review system)21 and specific studies in biomedicine22, social sciences23 and clinical medicine24. In addition, Okasa et al.25 proposed a supervised machine learning approach to evaluate the quality of peer review reports for grants. While their work focuses on review texts, our study addresses a complementary challenge: the automatic assignment of experts to project proposals using NLP. Both approaches highlight how AI can improve the efficiency and transparency of evaluation processes, but from different perspectives.

In summary, RAP has been addressed through various approaches: keyword-based, classification, multi-label prediction, topic coverage, and IR. In all these perspectives, there is a clear tendency towards semantic representations and similarity based on embedding using LLMs, which capture richer textual and contextual information than traditional keyword matching or older topic modeling approaches. Following these advances, our work makes a contribution to this emerging scenario by introducing an ORCID-integrated, LLM-based system designed specifically for MSCA proposal evaluation, which combines semantic similarity with aggregation strategies to provide a scalable and more accurate alternative to existing semi-automated methods.

Results

To evaluate the proposed assignment system, we compared our results with a list of 181 expert-proposal sets for validation, as these were the project proposals that were assigned to one of the experts on our list and which can be used to validate our system.

The main objective of this evaluation was to determine the accuracy of our content similarity-based system in reproducing the MSCA semi-automated assignment system. For each proposal, we generated a ranked list of experts based on similarity scores based on three different aggregation strategies: Sum, Product and Maximum .

First, to quantify the overall performance of each assignment strategy, we computed the area under the receiver operating characteristic curve (AUC) for the predicted expert rankings. The AUC is a commonly used metric in classification and ranking tasks26, as it represents the probability that a relevant item chosen randomly (here, it will be a final expert selected by the Vice Chairs) is ranked higher than an irrelevant item selected randomly (an expert not selected). Values closer to 1 indicate stronger agreement between the predicted ranking and the ground truth (VCs selection), while values near 0.5 correspond to random performance27,28. In other words, higher values of AUC will reflect more consistent prioritization of the most suitable experts.

In this analysis, for each proposal we consider a pool of 48 experts and evaluate whether each expert appears among the 3 final experts selected by the Vice Chairs (VCs). Using this binary ground truth, we generate ROC curves for the different aggregation strategies (Maximum, Sum, Product) and for the SEP system.

Figure 1 summarizes the ROC curves, showing the trade-off between true positive rate (experts correctly included in the top 3 VCs selection) and false positive rate across all proposals. We can see that the Maximum strategy shows the strongest alignment with human decisions with and AUC of 0.82, suggesting that reviewers prioritize the presence of at least one highly relevant match between a proposal and an expert’s prior work, rather than taking into account the overall similarity across all publications. The Sum based strategy with an AUC of 0.69 also captures relevant correspondences but it distributes similarity in a more uniform way, which may dilute the influence of the most distinctive work matches. The Product based approach with AUC = 0.57 performs notably worse, likely because low similarity values for some publications heavily penalize the aggregate score, underestimating experts with broad yet pertinent research profiles. The SEP system (AUC = 0.75) performs slightly below the Maximum strategy, reinforcing that emphasizing the strongest content-based match better approximates human patterns.

Fig. 1.

Fig. 1

Receiver Operating Characteristic (ROC) curves comparing the three aggregation strategies (Maximum, Sum, and Product) with the Single Evaluation Platform (SEP). The Area Under the Curve (AUC) values indicate the degree of agreement with the manual assignments made by vice-chairs, where higher values denote better discrimination between assigned and non-assigned experts. Among the tested methods, the Maximum approach achieved the best performance (AUC = 0.82), followed by SEP (AUC = 0.75), Sum (AUC = 0.69), and Product (AUC = 0.57).

To evaluate the ranking performance of our strategies, we analyzed the 48 candidate experts in our dataset against the final selections made by the VCs for 181 proposals. Although the VCs assigned a median of 3 experts per proposal in the original scenario, our study focused on the median of 1 expert per proposal that were available within our 48 candidate pool. Figure 2 illustrates the distribution of rank positions for the Maximum, Sum and Product strategies.

Fig. 2.

Fig. 2

Ranking distribution and dispersion analysis of expert assignments. The figure compares the three aggregation strategies (Maximum, Sum, and Product) ordered from highest to lowest AUC value. The x-axis represents the ranking positions assigned by each strategy exclusively to the subset of VC-selected experts available within our 48 candidate pool. While VCs assign a median of 3 experts per proposal, only a median of 1 of those selected experts were available within our candidate pool. The upper panel’s internal labels show the exact frequency of assignments at each rank. The Maximum strategy demonstrates the best performance, with a median rank of 4.0, meaning that the target expert is consistently placed among the first 4 recommendations. In contrast, the Product strategy shows a significantly higher Median of 28.0.

The results show that the Maximum similarity strategy consistently places VC-selected experts in higher ranks (closer to rank 1) than the other approaches, achieving a median rank of 4.0. In our search space of 48 candidates, this indicates that the max system consistently places the ground-truth expert among the top four recommendations. This confirms that prioritizing the most relevant match closely replicates human decision-making. The Sum-based strategy performs moderately with a median rank of 11.0, capturing cumulative similarity, but diluting the impact of the most relevant publications. In contrast, the Product-based approach performs worst, reaching a median rank of 28.0. This suggests that multiplicative aggregation heavily penalizes experts lacking a single standout match, even if they have several moderately relevant publications.

These trends are reinforced by the total sum of ranking positions across proposals, where lower values indicate greater alignment with the VCs final selections: Maximum = 1535, Sum = 2544, Product = 4496. Overall, the Maximum similarity strategy demonstrates the strongest concordance with human-reviewed assignments, highlighting its effectiveness for accurate and interpretable expert recommendation.

To complete the study of our aggregation strategies, we evaluated the numerical granularity of the SEP Affinity scores and our three similarity strategies (Maximum, Sum, Product) using standard objective metrics of resolution and dispersion. We calculated the following metrics: First, the Shannon entropy, to measure the uncertainty or information content of the distribution, where higher values reflect richer discrimination29. Secondly, the Minimum resolution in order to capture the smallest difference between adjacent unique values, reflecting the fineness of the scale. Then the Relative variance, defined as variance normalized by the mean, which quantifies the dispersion of values compared to their mean30. And Finally, the Mean distance between adjacent values.

Table 1 shows that our similarity based strategies are more accurate and fine-grained to analyze the relevance of expert-proposal. As our approaches semantically compare a full set of abstracts for each expert’s publications with the proposal applying for the MSCA, they capture more matches across content in contrast of the algorithm used in the SEP platform, which is primarily based on keyword affinity scores. Our aggregation strategies are better discriminating, as we can see focusing on the entropy values (all around 9.1 in contrast to SEP = 3.61) allowing the system to differentiate better among experts and providing richer and more precise information for then creating the ranking.

Table 1.

Granularity metrics for Single Evaluation Platform (SEP) affinity and the three similarity aggregation strategies (Maximum, Sum and Product).

Variable Entropy Min_resolution Relative_variance Mean_distance
SEP_Affinity 3.6105 1.00e+00 4.5186 1.4808e+00
Maximum_Similarity 9.0962 4.86e-09 0.0260 8.5735e-05
Sum_Similarity 9.0972 1.74e-08 29.3062 3.1034e-02
Product_Similarity 9.0968 1.67e-308 0.2758 7.6692e-05

Entropy = Shannon entropy of value distribution; Min resolution = smallest gap between adjacent unique values; Relative variance = Inline graphic; Mean distance = mean distance between unique adjacent values.

Between our different strategies, the Sum presents a very high relative variance (=29), exaggerating differences for experts with more experience and scientific career, and Maximum achieves a balanced combination of high resolution and stable dispersion, allowing for nuanced discrimination without numerical instability.

In conclusion, our results show that the different aggregation strategies can capture effectively the logic underlying the VCs assignment process. The Maximum strategy (AUC = 0.82) aligns with a best-fit approach, emphasizing the strongest single match between a proposal and an expert’s work. The Sum strategy (AUC = 0.69) reflects a cumulative expertise logic, taking into account overall similarity across all their publications, while the Product strategy (AUC = 0.57) implies a stricter requirement for consistent similarity across multiple contributions. The SEP system, which provides automated initial assignments by the platform based on keyword–expert correlation, performs comparably to the Maximum strategy but slightly below it, suggesting that the VCs panel also prioritizes top individual matches. Overall, these results indicate that the VCs’s final expert selections are more consistent with the best-fit logic (modifying the SEP’s automated assignment), that our Maximum similarity strategy can replicate. Furthermore, since our similarity ranking operates at a much finer granularity, at the proposal level, than the keyword-based matching of the SEP, it captures more nuanced correspondences between expert expertise and project content, adding interpretive richness.

Discussion and conclusion

The analysis conducted allowed us to evaluate the effectiveness of the expert assignment system proposed in this study, compared to manual assignment. The three similarity aggregation strategies (Maximum, Sum, and Product) had an impact on the position of the experts in the similarity ranking of each proposal, with the Maximum strategy having the highest accuracy (0.82). This indicates that the Maximum strategy is closest to the manual assignment of experts to proposals made by the Vice Chairs, which is the standard approach in MSCA expert assignment. It should be noted that this manual assignment usually involves an initial automated suggestion by the system, followed by a final evaluation and decision by the experts. We compared our system with this final and manual selection performed by the MSCA semi-automatic system.

This finding does not imply that the other two strategies lose value or are invalid. Each strategy provides unique information that can be valuable in different circumstances.

On the one hand, the sum strategy helps to weight the number of relevant publications of each expert, accumulating similarities to generate an assignment that reflects each reviewer’s extent of knowledge on the specific topic of the proposal. This summative approach can be especially useful when an expert has multiple relevant publications that together provide a high degree of semantic alignment with a proposal. However, it may be disadvantageous for new experts with a smaller volume of publications.

However, the Product aggregation provided a stricter view by reducing the scores in the presence of any publications with low similarity. This approach focuses on assignments where there is consistency in the similarity of all the expert’s publications. It allows identifying cases where the relationship between the reviewer’s publications and the proposal was strong across all the papers analyzed, enhancing thematic consistency. However, this strategy may be less effective when experts have a large volume of publications, as any low similarity significantly reduces the overall result.

Finally, the Maximum strategy proved to be useful in contexts where a single highly relevant publication could justify the assignment of the expert to a proposal. This approach was shown to be effective in providing a quick and efficient methodology to select experts on specialized topics. However, its success is highly dependent on the variability in the quality of similarity scores between abstracts and does not take into account the full set of publications of each reviewer.

In conclusion, this study presents a content similarity-based system based on Natural Language Processing (NLP) for the assignment of experts in the context of MSCA postdoctoral fellowships, integrating ORCID-based publication retrieval with semantic similarity, calculated using the GALACTICA model. The evaluation of three aggregation strategies demonstrated that the Maximum similarity approach achieved the highest concordance with the assignments of the final MSCA semi-automated system (AUC = 0.82) after the VCs intervention. This result suggests that human decision-making often prioritizes a single highly relevant publication over cumulative or multiplicative evidence. Furthermore, our similarity-based ratings provide much finer granularity and more precise and accurate differentiation between experts compared to the discrete Affinity scores made by the SEP platform, allowing our system to capture sensitive semantic matches across all publication records.

The Maximum aggregation provides the most robust and interpretable results, closely reflecting the VC’s ”best-fit” selection thought while maintaining meaningful consistency with SEP’s automatic assignments. It offers a reliable, scalable, and interpretable framework for expert assignment that could reduce manual workload and improve transparency in the reviewer selection process.

The results highlight the potential of LLMs to improve expert matching by reducing the use of predefined keywords and also provide a more qualitative assessment of expertise. Although Sum and Product strategies also capture valuable aspects of reviewer expertise, the Maximum approach showed closer alignment with current practice.

We note several limitations in our study. First, Our validation dataset includes 190 proposals and a limited list of 50 experts from the LIF panel, as we followed the transparency and data-sharing agreements made with the REA internal ethics and data-protection teams. Consequently, some experts assigned to proposals by SEP or modified by the VCs were not included in our dataset and could not be evaluated. Although the size is limited, this dataset represents a real-world, validated MSCA ground truth, which provides a robust basis for internal evaluation. Moreover, SEP initially assigns three experts per proposal and calculates affinities for all candidates, while VCs modify approximately 39% of these assignments. Since the VCs start from SEP’s assignment recommendations, many proposals keep one or more of the originally assigned experts, resulting in a moderate number of changes in general.

As a future step, it would be useful to test our three strategies in collaboration with Vice Chairs to determine which approach would be most effective for the final assignment of experts. This collaborative evaluation could provide valuable information to refine the system and better support the decision-making process in the assignment of experts.

As future work, our aim is to conduct a human evaluation by systematically assigning proposals using multiple systems (e.g., SEP, VC manual assignment, our method) and having independent reviewers assess the suitability of each assignment in a blind review mode. Furthermore, while this work successfully utilized GALACTICA, an LLM specifically trained on scientific text, future research will explore the potential impact of leveraging other contemporary, larger-scale LLMs (such as advanced generalist or domain-specific models, or their state-of-the-art embedding representations). This analysis is crucial for assessing the transferability and impact of the latest language models on the system’s expert assignment performance. Although our system provides precise, fine-grained similarity-based rankings to streamline the review process, we will evaluate across broader panels and disciplines in order to confirm generalization. This approach would provide a comparative benchmark and strengthen the generalization of our findings.

Overall, this work provides a novel and transparent empirical methodology for expert assignment, offering a scalable tool to support more accurate and efficient evaluation workflows in large-scale research funding programs. By comparing our content-based approach with the MSCA’s semi-automated assignments, we demonstrate its potential to complement existing procedures and provide data-driven insights to optimize reviewer assignment.

Methods

The methodology for this study is focused on implementing an expert assignment pipeline based on the semantic similarity between the abstracts of experts’ research articles and project proposals. The overall process involved three main phases: (1) Data Construction (extraction of ORCID identifiers and publication corpus), (2) Semantic Similarity Computation (using a Large Language Model for content comparison), and (3) Evaluation of Aggregation Strategies to produce the final expert rankings.

Figure 3 illustrates the workflow followed to obtain the final dataset for this study. The evaluation is based on the Life Sciences (LIF) panel, part of the Individual Fellowships (IF) program under the Horizon 2020 (H2020) framework. LIF proposals span a wide range of research areas, including molecular and cell biology, neurosciences, immunology, and public health. Experts are selected to cover the diversity of these disciplines, ensuring proportional representation across the seven LIF sub-panels.

Fig. 3.

Fig. 3

Flowchart of the validation dataset creation. Process used to create the final dataset for expert assignment in evaluating initial list of findings (LIF) proposals under the H2020 call. Starting with 50 semi-automatically assigned experts, we retrieved ORCID identifiers and publication records for 48 experts to assess content similarity with project proposals. After filtering, the final dataset included 181 unique proposals assigned to experts with accessible publication data.

The evaluation of proposals within the Life Sciences panel is coordinated by the European Research Executive Agency (REA) using the Single Evaluation Platform (SEP), also known as the Commission’s Evaluation Tool or Evaluation Service. SEP generates preliminary expert assignments by automatically matching proposal abstracts to expert profiles, mainly relying on keyword overlaps. For each proposal, SEP provides an initial List of Findings (LIF), which consists of three assigned experts along with a ranked list of up to 25 alternative candidates and their respective affinity scores. These preliminary assignments are then reviewed and adjusted by a panel of Vice Chairs (VCs) that verify the suitability of the proposed matches. The final LIF incorporates the VCs-approved assignments, which can differ substantially from the original SEP output due to manual corrections and refinements.

In this study, we concentrated on a curated set of 50 experts who were initially semi-automatically allocated to evaluate LIF proposals. Publication data and ORCID identifiers were successfully retrieved for 48 of these experts, enabling the computation of semantic similarity scores between their published work and proposal abstracts. To maintain a consistent comparison, both the SEP and VCs datasets were filtered to retain only the proposals associated with these 48 experts, ensuring that subsequent analyses of ranking performance reflect the same subset of candidates.

Our sample dataset shows a similar trend as the complete H2020 assignments, where 38.95% of the SEP assignments were modified by Vice Chairs: Total proposals: 190; Total allocations: 570; Total changes: 222 (38.95% changed); Changes per proposal: none (22.11%), one (41.05%), two (34.74%), all three (2.11%).

This curated dataset was used to validate our semantic similarity-based system while ensuring compliance with personal data protection regulations (e.g., GDPR European Commission, 2018)31. By restricting the analysis to this subset, we were able to compare our system’s rankings with both the SEP-assigned experts and the VCs-reviewed experts, allowing an evaluation of alignment with the automatic assignments and the human-corrected final assignments.

Data construction

First of all, in order to make comparisons between project proposals and the experts we needed to extract the abstracts of all papers from the expert list. For this purpose, we implemented a technique for the extraction of ORCID identifiers from the authors’ names. The ORCID identifier is a standard and reliable tool that facilitates the unique identification between researchers and their respective publications, avoiding confusion caused by similar or variations in names32.

The extraction process used a combination of name matching techniques and similarity filters to ensure the accuracy of the extracted ORCIDs. Initially, the search was performed through the ORCID API, using normalized first and last names, which allows retrieving the set of matches by these two parameters within the ORCID platform. When the query generated a single match, the retrieved ORCID was added to our expert dataset. However, when multiple matches were obtained, we applied an additional filter based on the institutional affiliation of the expert. For these cases, our system calculates the cosine similarity between the name of the expert’s institution and the affiliations registered in ORCID providing a list with the similarities. This analysis allowed the identification of the most similar profiles among which we can manually select the corresponding ORCID.

At the end of the process, the data was saved in separate files, classifying those with unique matches and those that required additional manual verification. This approach allowed us to retrieve ORCIDs for 48 of the total 50 experts (96%).

Once we obtained the ORCID identifiers of the list of experts, we had to create the corpus of scientific articles that would help us to assign experts to the project proposals. To collect the abstracts of the papers of each of the experts on the list, we used the ORCID API to obtain a list of all the DOIs of the papers of each expert. Once the DOIs were obtained, we made requests to the CrossRef API through its Python library to obtain the abstracts of each paper (DOI).

In case the library did not return the abstract of an article, we used web scraping techniques to retrieve it. We used tools such as BeautifulSoup together with requests to extract the content directly from the article web pages (making requests directly to the DOIs). However, we encountered more complex situations where the structure of the web pages was dynamic and a simple request could not be made because the content was not loaded. For this we use Selenium with a headless browser to simulate the interaction with the site and extract the necessary data. These combined techniques allowed us to obtain a total of 2,806 unique abstracts, which formed the database used to calculate the similarity of content between project proposals and scientific articles.

Semantic similarity computation

Finally, for the evaluation of the similarity between the abstracts of the articles and the proposals, we used the LLM GALACTICA, developed by Facebook AI. We chose GALACTICA instead of other well-known LLMs, such as BERT or GPT, because of its unique specialization in the scientific domain. Although general-purpose models such as BERT33 and GPT34 are strong in a wide range of natural language processing tasks, their training is mainly based on diverse and non-specialized corpus. In contrast, GALACTICA was specifically trained on a large corpus of scientific articles, research datasets and technical content35. This specific training allows GALACTICA to better understand the language, terminology, and semantic complexities of academic literature, which is a significant advantage when analyzing similarity between scientific publications and research-focused texts, such as project proposals.

To calculate the similarity between expert abstracts and project proposals, we applied GALACTICA to both of them, generating embeddings for each of the documents. These embeddings transform the textual data into high-dimensional numerical vectors, allowing us to directly compare the semantic content. For the comparison, we use the cosine similarity metric, which measures the distance between two vectors in the embedding space. This metric produces a score between 0 and 1, where a value close to 1 indicates a higher degree of semantic similarity. Thanks to this method, we can ensure a reliable and contextual evaluation of the relationship between experts’ publications and project proposals.

The advantages of this approach are as follows: first of all, GALACTICA’s domain improves the accuracy of the text representation, capturing semantic nuances that may be missed by non-specialist models; and secondly, the use of cosine similarity provides an accurate numerical assessment of the semantic alignment between the proposals and the experts’ work. Together, these factors help to improve the reliability and accuracy of inter-expert agreement in the evaluation process.

Aggregation strategies for expert ranking

Once we have obtained the similarity between abstracts and proposals, we proceed to obtain the expert or experts who were most in line with each of the proposals. To do so, we proposed three alternatives, taking into account the set of similarities that each author will have, which varied according to the total number of published articles. The 3 proposals of aggregation of similarities for each author were: the Sum, the Product and the Maximum.

Firstly, the Sum strategy consisted of aggregating all similarity scores of the abstracts authored by an expert, which allowed for a cumulative approach where a higher volume of similarities could indicate a higher overall relevance of the assignment for a specific expert. Alternatively, the Product strategy multiplied similarity scores, looking at those assignments where there was consistent alignment across all individual similarities. Since these scores are in the range of [0,1], the product generated lower scores in cases where any individual similarity was low, which helped to identify assignments with lower consistency. Finally, the Maximum strategy selected the highest similarity score from the abstracts authored by each expert, under the assumption that if at least one of the publications had considerable similarity, this would be enough to rank the expert on the topic.

Acknowledgements

This work was partially supported by the Capgemini-University of Valencia Chair for Innovation in Software Development.

Disclaimer

All views expressed in this article are strictly those of the authors and may in no circumstances be regarded as an official position of the Research Executive Agency or the European Commission.

Author contributions

EA-G designed the study, built the dataset, designed and executed the analysis and revised the manuscript. DG-C designed the study, built the dataset, executed the analysis and revised the manuscript. IW provided data and revised the manuscript. AM coordinated part of the data collection, designed the study and revised the manuscript. FG coordinated and designed the study, revised the analysis and revised the manuscript.

Funding

EA-G. is supported by the Spanish Ministry of Science, Innovation and Universities through the FPU doctoral fellowship program [grant number FPU21/00570].

Data availability

The data for replication are available at this link: https://dataverse.harvard.edu/previewurl.xhtml?token=9ddf76ed-c05f-4b96-a2da-f1f4ad79dfcd

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Zhao, X. & Zhang, Y. Reviewer assignment algorithms for peer review automation: A survey. Inf. Process. Manag.59, 103028. 10.1016/j.ipm.2022.103028 (2022). [Google Scholar]
  • 2.Commission, E. Marie skłodowska-curie actions work programme 2018-2020 (2018). https://ec.europa.eu/research/participants/data/ref/h2020/wp/2018-2020/main/h2020-wp1820-msca_en.pdf.
  • 3.Baumert, P., Cenni, F. & Ten Antonkine, M. L. simple rules for a successful eu marie skłodowska-curie actions postdoctoral (msca) fellowship application. PLoS Comput. Biol.18, e1010371. 10.1371/journal.pcbi.1010371 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Commission, E. Postdoctoral fellowships - marie skłodowska-curie actions.
  • 5.European Commission, R. & Portal, I. F. T. Evaluate a proposal—it how to—evaluation tool. https://webgate.ec.europa.eu/funding-tenders-opportunities/display/IT/Evaluate+a+proposal (2021).
  • 6.Flach, P. A. et al. Novel tools to streamline the conference review process: Experiences from sigkdd’09. ACM SIGKDD Explor. Newsl.11, 63–67. 10.1145/1809400.1809413 (2010). [Google Scholar]
  • 7.Protasiewicz, J. A support system for selection of reviewers. In 2014 IEEE International Conference on Systems, Man, and Cybernetics 3062–3065. 10.1109/SMC.2014.6974408 (IEEE, 2014).
  • 8.Di Mauro, N., Basile, T. M. A. & Ferilli, S. Grape: An expert review assignment component for scientific conference management systems. In 18th International Conference on Innovations in Applied Artificial Intelligence 789–798 (2005).
  • 9.Karimzadehgan, M., Zhai, C. & Belford, G. Multi-aspect expertise matching for review assignment. In 17th ACM Conference on Information and Knowledge Management (CIKM) 1113–1122 (2008).
  • 10.Charlin, L. & Zemel, R. The toronto paper matching system: An automated paper-reviewer assignment system (In ICML Workshop on Peer Reviewing and Publishing Models (2013).
  • 11.Zhao, S. et al. A novel classification method for paper-reviewer recommendation. Scientometrics115, 1293–1313. 10.1007/s11192-018-2726-6 (2018). [Google Scholar]
  • 12.Zhang, D. et al. A multi-label classification method using a hierarchical and transparent representation for paper-reviewer recommendation. arXiv preprint arXiv:1912.08976 (2019).
  • 13.Kou, N. M., Mamoulis, N. & Gong, Z. Weighted coverage based reviewer assignment. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data 2031–2046. 10.1145/2723372.2723727 (Association for Computing Machinery, 2015).
  • 14.Mirzaei, M., Sander, J. & Stroulia, E. Multi-aspect review-team assignment using latent research areas. Inf. Process. Manag.56, 858–878. 10.1016/j.ipm.2019.01.007 (2019). [Google Scholar]
  • 15.Yue, M., Tian, K. & Ma, T. An accurate and impartial expert assignment method for scientific project review. J. Data Inf. Sci.2, 65–80. 10.1515/jdis-2017-0020 (2017). [Google Scholar]
  • 16.Marco-Tordera, L., García-Costa, D. & Grimaldo, F. Cognitive Similarity Through Bibliometric Analysis (IOS Press, 2023).
  • 17.Zhang, Y., Yang, R., Jiao, S., Kang, S. & Han, J. Scientific paper retrieval with llm-guided semantic-based ranking. arXiv preprint arXiv:2505.2181510.48550/arXiv.2505.21815 (2025).
  • 18.Mitrov, G., Stanoev, B., Gievska, S., Mirceva, G. & Zdravevski, E. Combining semantic matching, word embeddings, transformers, and llms for enhanced document ranking: Application in systematic reviews. Big Data Cognit. Comput.8, 110. 10.3390/bdcc8090110 (2024). [Google Scholar]
  • 19.Stergiopoulos, V., Vassilakopoulos, M., Tousidou, E., et al. & Corral, A. Conference management system utilizing an llm-based recommendation system for the reviewer assignment problem. In 27th International Conference on Enterprise Information Systems (ICEIS 2025). 10.5220/0013482600003929 (2025).
  • 20.Bagheri, F., Buscaldi, D. & Recupero, D. R. A study on content-based reviewer assignment in the semantic web and computer science domains. Computación y Sistemas28, 10.13053/cys-28-4-5299 (2024).
  • 21.AAAI. Aaai launches ai-powered peer review assessment system. https://aaai.org/aaai-launches-ai-powered-peer-review-assessment-system/ (2024).
  • 22.Doskaliuk, B., Zimba, O., Yessirkepov, M., Klishch, I. & Yatsyshyn, R. Artificial intelligence in peer review: Enhancing efficiency while preserving integrity. J. Korean Med. Sci.40, e92. 10.3346/jkms.2025.40.e92 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Horbach, S. P. J. M. Pandemic publishing: Medical journals strongly speed up their publication process for covid-19. Quant. Sci. Stud.1, 1056–1067. 10.1162/qss_a_00076 (2020). [Google Scholar]
  • 24.Perlis, R. H. et al. Artificial intelligence in peer review. JAMA10.1001/jama.2025.15827 (2025). [DOI] [PubMed]
  • 25.Okasa, G. et al. A supervised machine learning approach for assessing grant peer review reports. Quant. Sci. Stud.10.1162/qss.a.23 (2025). [Google Scholar]
  • 26.Hoo, Z. H., Candlish, J. & Teare, D. What is an roc curve? (2017). [DOI] [PubMed]
  • 27.Fawcett, T. An introduction to roc analysis. Pattern Recogn. Lett.27, 861–874. 10.1016/j.patrec.2005.10.010 (2006). [Google Scholar]
  • 28.Hanley, J. A. & McNeil, B. J. The meaning and use of the area under a receiver operating characteristic (roc) curve. Radiology143, 29–36. 10.1148/radiology.143.1.7063747 (1982). [DOI] [PubMed] [Google Scholar]
  • 29.Shannon, C. E. A mathematical theory of communication. Bell Syst. Tech. J.27, 379–423. 10.1002/j.1538-7305.1948.tb01338.x (1948). [Google Scholar]
  • 30.Manning, C. D., Raghavan, P. & Schütze, H. Introduction to Information Retrieval (Cambridge University Press, 2008).
  • 31.Commission, E. General Data Protection Regulation (gdpr): Regulation (eu) 2016/679 of the European Parliament and of the Council (Official Journal of the European Union, 2018).
  • 32.Haak, L. L., Fenner, M., Paglione, L., Pentz, E. & Ratner, H. Orcid: a system to uniquely identify researchers. Learned Publ.25, 259–264. 10.1087/20120404 (2012).
  • 33.Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. 10.48550/ARXIV.1810.04805 (2018).
  • 34.Language models are few-shot learners. 10.48550/ARXIV.2005.14165 (2020).
  • 35.Galactica: A large language model for science. 10.48550/ARXIV.2211.09085 (2022).

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data for replication are available at this link: https://dataverse.harvard.edu/previewurl.xhtml?token=9ddf76ed-c05f-4b96-a2da-f1f4ad79dfcd


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES