Abstract
Background
Knowledge discovery in scientific literature is hindered by the increasing volume of publications and the scarcity of extensive annotated data. To tackle the challenge of information overload, it is essential to employ automated methods for knowledge extraction and processing. Finding the right balance between the level of supervision and the effectiveness of models poses a significant challenge. While supervised techniques generally result in better performance, they have the major drawback of demanding labeled data. This requirement is labor-intensive, time-consuming, and hinders scalability when exploring new domains.
Methods and Results
In this context, our study addresses the challenge of identifying semantic relationships between biomedical entities (e.g., diseases, proteins, medications) in unstructured text while minimizing dependency on supervision. We introduce a suite of unsupervised algorithms based on dependency trees and attention mechanisms and employ a range of pointwise binary classification methods. Transitioning from weakly supervised to fully unsupervised settings, we assess the methods’ ability to learn from data with noisy labels. The evaluation on four biomedical benchmark datasets explores the effectiveness of the methods, demonstrating their potential to enable scalable knowledge discovery systems less reliant on annotated datasets.
Conclusion
Our approach tackles a central issue in knowledge discovery: balancing performance with minimal supervision which is crucial to adapting models to varied and changing domains. This study also investigates the use of pointwise binary classification techniques within a weakly supervised framework for knowledge discovery. By gradually decreasing supervision, we assess the robustness of these techniques in handling noisy labels, revealing their capability to shift from weakly supervised to entirely unsupervised scenarios. Comprehensive benchmarking offers insights into the effectiveness of these techniques, examining how unsupervised methods can reliably capture complex relationships in biomedical texts. These results suggest an encouraging direction toward scalable, adaptable knowledge discovery systems, representing progress in creating data-efficient methodologies for extracting useful insights when annotated data is limited.
Keywords: Knowledge discovery, Relation extraction, Unsupervised learning, Weakly supervised learning, Biomedical text
Introduction
The exponential growth in scientific publications designates knowledge discovery as a critical research area, as the sheer volume of new findings presents ongoing challenges for researchers and practitioners to stay updated on developments. Automated knowledge extraction and processing methods are essential to address this information overload. In knowledge discovery, one of the core tasks is to determine whether there is a semantic relationship between entities, such as diseases, proteins, genes, and medications, in unstructured text. Relation extraction represents an essential building block for downstream knowledge discovery pipelines that integrate extracted relations into structured knowledge schemes or support higher-level reasoning tasks. Balancing supervision level and model performance is a primary challenge, as supervised methods typically yield higher performance, but they come with the significant drawback of requiring annotated data, which is time-consuming, resource-intensive, and lacks scalability when applied to new domains. The creation of gold-standard datasets is cumbersome and further limits model adaptability and generalizability to emerging areas, thus underscoring the need for methods that maintain performance with reduced supervision.
Given these constraints, reducing the reliance on supervision while maintaining high levels of performance is a pivotal research necessity. Unsupervised and weakly supervised methods offer promising alternatives, especially in domains where annotated data is scarce or difficult to acquire. In this paper, we introduce a suite of algorithms designed to minimize supervision in knowledge discovery. We hypothesize that grammatical and syntactic structures play a crucial role in defining semantic relationships between entities. To test this, we explore approaches that leverage dependency trees [1] to model these relationships. Manning et al. [2] emphasize that deep contextual language models capture linguistic hierarchies and syntactic relationships without explicit supervision. By introducing methods to detect these emergent structures within neural networks, they demonstrate that specific model components target syntactic dependencies and anaphoric references. By applying a linear transformation to the models’ learned embeddings, they successfully approximate parse tree distances, allowing for a credible reconstruction of sentence structures traditionally outlined by linguists [2]. Inspired by this research path, we utilize the inherent capabilities of language models (LMs) to capture semantic relations by incorporating attention mechanisms [3] to predict entity interconnections.
Additionally, we frame knowledge discovery as a pointwise binary classification task [4] and progressively reduce the level of supervision, transitioning from weakly supervised to fully unsupervised approaches. Our methods are evaluated on four biomedical benchmark datasets, showing promising results that suggest the feasibility of minimizing supervision while maintaining a robust performance. This research highlights the potential for more scalable and adaptable knowledge discovery systems that are less dependent on annotated data, paving the way for broader applicability across scientific domains. In summary, the key paper’s contributions are:
Introduction of Novel Unsupervised Algorithms: We develop dependency tree and attention-based algorithms for identifying semantic relations between biomedical entities, reducing the reliance on annotated data.
Pointwise Binary Classification for Knowledge Discovery: To the best of our knowledge, our paper introduces the use of diverse pointwise binary classification methods in a weakly supervised knowledge discovery setting.
Pivoting from Weakly Supervised to Unsupervised Setups: We demonstrate a transition from weakly supervised to unsupervised learning, showcasing the methods’ capacity to handle data with noisy labels.
Robust Validation across Biomedical Datasets: We rigorously validate our methods across four biomedical datasets, revealing insights and demonstrating their effectiveness and robustness under different learning paradigms.
The paper is structured as follows: Section "Related work" reviews related work, including distant supervision, multi-instance learning, rule-based methods, and zero-shot and few-shot learning approaches. Section "Methods" outlines the study’s methods and formalizes the knowledge discovery task. Section "Experiments" describes the experimental setup, detailing the datasets, implementation, and results. Section "Discussion" discusses the findings, providing insights into methods and the different learning paradigms of the study.
Related work
Rule-based methods. Rule-based approaches to relation extraction involve defining rules [5–13] as regular expressions over words or part-of-speech (POS) tags, either manually or automatically learned from training data. These rules help identify relations between entities. For example, early research [5] extracts gene-gene interactions using manually crafted linguistic patterns. A pattern like "gene product modifies gene" could match sentences like "Eg1 protein modifies BicD.", where "Eg1" and "BicD" are identified as arguments of the predicate "modifies". Similarly, Ono et al. [6] design rules based on syntactic features to handle complex sentences, incorporating negation handling to improve performance. Blaschke and Valencia [14] enhance rule-based methods by assigning probability scores to rules based on reliability and incorporating factors such as negation and the distance between protein names. In another example, the PPInterFinder tool [15] utilizes rule-based patterns to extract human PPIs from biomedical texts.
However, manually defining rules requires significant human effort and potentially lacks generalizability to different domains. It is also impractical to enumerate rules to cover all possible descriptions of PPIs in text. Consequently, researchers attempt to learn rules automatically from data incorporating supervision. For instance, Phuong et al. [7] use link grammar parsers and heuristics to automatically derive extraction rules, while Huang et al. [8] employ dynamic programming to learn PPI patterns based on POS tags. Liu et al. [11] use PATRICIA trees to store training sentences and extract potential interaction patterns. Additionally, Thomas et al. [13] infer a large set of linguistic patterns using information about interacting proteins, refining patterns based on shallow linguistic features and dependency semantics. Inspired by rule-based approaches, we propose a dependency-based method that establishes a strong unsupervised baseline for our study, using three universal, coarse-grained assumptions. The design of the method reduces reliance on domain-specific tuning while maintaining robust performance. However, pattern rules have difficulties in capturing linguistic patterns exhaustively and often ignore contextual information that is necessary to detect the biomedical relationships presented in the texts, which is also empirically shown by the results of our unsupervised baseline.
Distant supervision and multi-instance learning. Initial research on distant supervision for relation extraction (RE) [16, 17] follows the assumption that if two entities are related, all sentences mentioning them should convey that relationship. However, this assumption is limiting and is subsequently relaxed by Riedel et al. [18], who introduce multi-instance learning (MIL). MIL suggests that if two entities are related, at least one sentence mentioning them could express that relationship. Building upon this, Hoffmann et al. [19] further enhance the approach by accommodating overlapping relations. Zeng et al. [20] improve distant supervision by combining MIL with a piecewise convolutional neural network (PCNN), while Lin et al. [21] introduce an attention mechanism that focuses on relevant information within a collection of sentences. This attention-based MIL approach for sentence-level relation extraction spurred many subsequent studies [22–24]. Han et al. [23] propose a joint model combining a knowledge graph with MIL and an attention mechanism to enhance relation extraction. Dai et al. [25] extend the work of Han et al. [23] into the biomedical domain, using a PCNN to encode sentences.
Further improvements are introduced by Amin et al. [26], who employ BioBERT [27] for encoding sentences in relation extraction tasks. They utilize MIL with entity-marking strategies inspired by R-BERT [28], achieving the best performance when aligning the direction of extracted triples with the UMLS knowledge graph [29]. Hogan et al. [30] introduce abstractified multi-instance learning (AMIL), which significantly improves performance for finding uncommon relationships. In our approach, we do not rely on either of the two common assumptions: that every sentence mentioning two entities expresses their relationship [16, 17], or that at least one sentence mentioning both entities does [18]. In addition, we do not rely on external knowledge sources such as a knowledge graph and do not use this extra supervision to improve the results making our method more generally applicable.
A limitation of MIL approaches is their assumption that a minimum number of sentences (multiple instances) must contain both entities. This restricts the efficient inclusion of rare interconnections and poses challenges in new, unexplored domains. Another limitation of MIL is the inability to determine the specific sentence in which the relation exists. This increases the effort required for potential human or expert-in-the-loop evaluation, as instead of providing individual sentences, an entire group of sentences must be reviewed. Unlike MIL methods, we focus on predicting whether a relationship exists between two entities in every sentence that includes the entity, rather than across a group of sentences, with an emphasis on reducing the level of supervision required.
Zero-shot and few-shot learning. The emergence of large language models (LLMs) brings improvements and potential to RE, including zero-shot few-shot learning with domain-specific applications like biomedical RE. Traditional RE approaches rely on supervised learning methods [31–33], where models are trained to predict the relationships between tagged entity spans. In this paper, our goal is to reduce supervision and even propose an unsupervised approach. Recent research explores the potential of treating RE as a sequence-to-sequence (seq2seq) task, where relations are linearized and generated as output strings conditioned on the input text. Again, this approach requires the supervision of sufficient annotated training examples to learn a mapping between input and output. Wadhwa et al. [34] revisit RE by leveraging the potential of large language models that capture a larger context and are trained on a large amount of data, such as GPT-3 [35] and Flan-T5 [36]. They demonstrate that using LLMs in a seq2seq framework could achieve near state-of-the-art (SOTA) results under few-shot settings. Gao et al. [37] introduce the hierarchical prototype optimization (HPO) approach. They construct prototypes from a small number of samples, and address prototype bias in few-shot relational triple extraction by leveraging prompt learning and hierarchical contrastive learning [38–40].
In the biomedical domain, studies [41, 42] highlight that few-shot in-context learning using LLMs with prompt engineering presents quite low performance compared to classification-based methods. Prompts need to be carefully designed and do not guarantee applicability across subject domains. Zhang et al. [43] explore the capabilities of GPT
3.5-turbo and GPT-4 [35, 44] in zero-shot and one-shot learning setups. Although using LLM in zero-shot or few-shot learning setups can reduce the need for supervision, challenges remain regarding the stability and hallucinations of the models [45–48]. In this study, we recognize that language models possess factual relational knowledge [49] and propose algorithms that use attention scores to perform RE in an unsupervised way. Furthermore, we approach the knowledge discovery task as a pointwise binary classification problem, gradually decreasing the level of supervision and transitioning from a weakly supervised to an unsupervised setting.
Methods
Structuring raw textual information into actionable relational knowledge is a crucial step in knowledge discovery pipelines, supporting downstream reasoning tasks. We formally define the task as follows:
Task formulation. Let the set of sentences, each containing two identified entities (e.g. protein, gene, disease, drug, etc.)
and
, be
and the label set be
, where the positive label (+1) indicates that there is a semantic relation between
and
and the negative label (-1) that there is no relation. The dataset is defined as
where each instance
is independently sampled from the joint distribution with density p(s, y), which includes a sentence
and a label
. The goal is to determine a binary classifier
.
Dependency-based method
According to Jurafsky and Martin [50], a sentence’s syntactic structure consists of words connected by directed, labeled arcs representing binary grammatical relations. This typed dependency structure uses labels from a predefined set and includes a root node marking the tree’s head. In the resulting directed graph, the root has no incoming arcs, each other node has exactly one, and a unique path exists from the root to every node. To identify semantic relations between entities within a sentence, we focus on the shortest dependency path (SDP) between two target words, which is extracted from the dependency tree. The motivation for focusing on SDP is based on the observation that the SDP between entities usually contains the necessary information to identify their semantic relationship [51]. We formulate the following assumptions regarding the existence of a semantic relation between entities:
Root Verb: The root word should be a verb and appear in SDP.
Root Word: The root word should appear in SDP.
Verb: A verb should appear in SDP.
Purpura et al. [52] introduced the first assumption to mine relationships in biomedical texts, particularly between diseases and symptoms. We exclusively incorporate the two additional assumptions to reduce the strictness level. The second assumption hypothesizes that the structural head of the sentence plays a role in connecting the entities, and the third focuses on the importance of verbs in describing actions and relationships. The assumptions provide a structured and syntactic way to assess whether two entities are semantically related within a sentence. Building on these assumptions, we define two heuristics:
Two entities are considered semantically related if the SDP satisfies the given assumption or if there is a direct link between the entities in the SDP.
The first heuristic should hold and the SDP must not contain any “conjunction” words.
The second assumption is based on the hypothesis that conjunctions (e.g. and, but) often connect clauses or phrases with distinct meanings, and their presence may indicate that the entities belong to separate, unrelated parts of the sentence. The conjunctions are detected using a POS tagger. Considering the assumptions and the heuristics, we formally introduce the SDP-based Algorithm for Relation Detection (SARD).
Algorithm 1.
SARD
SARD. Given a sentence sent with two defined entities
and
, an assumption id
and a heuristic id
, predict if there is a semantic relation between the entities (Algorithm 1). The assumption id
can have a value of 1, 2, or 3, referring to the corresponding assumption that is introduced in this subsection. Accordingly, the heuristic is set to 1 or 2 based on the chosen heuristic. Extract the dependency parsing tree dpt of the sentence sent (
). Get the POS [50] tags pos of each word in the sentence sent (
). Find the SDP
between the entities
and
using the dependency parsing tree dpt (
). Extract the POS tags
of the words that are part of the SDP
(
). Check if there is a direct link in the dependency parsing tree dpt between the entities
and
(
) and store the boolean value
. Check if the given assumption
holds (
) and store the boolean value a. For example, if
then the function
checks if the root word is a verb, using the POS tags
, and appears in the SDP
. If the assumption does not hold and there is no direct link between
and
then predict that there is no semantic relation between the entities; otherwise, apply the given heuristic
to predict the label.
SARD, acting as a baseline method in our study, is referred to as an unsupervised algorithm for knowledge discovery because, although it relies on a POS tagger and a dependency parser, which themselves are trained using supervised methods, it does not use any direct supervision and annotated data specifically for the knowledge discovery task. We acknowledge that the algorithm’s reliance on external tools could be a limitation in specific domains where syntactic parsing may be less reliable, although POS taggers and dependency parsers generally achieve strong performance [53–58].
Attention-based methods
ATLOP [59], a SOTA model in document-level relation extraction [60, 61], introduces the concept of localized context embeddings to address the variations of relevant mentions and contextual information for different entity pairs
. To accomplish this, the localized context embeddings exploit the idea of localized context pooling that utilizes the attention patterns of a pre-trained language model to pinpoint and gather relevant context essential for understanding the relationship between entity pairs
. The estimation of the localized context distribution (
) is a key step in the computation of the localized context embeddings. We utilize a variation of
, focusing on the context of the sentence, that is formally described as follows:
Localized Context Distribution (
). The tokens of a sentence S,
, are encoded via a transformer layer [3]
with
of a pre-trained language model (LM), where k the number of the encoding layers of the LM, as follows:
![]() |
1 |
where
and
represent the token embeddings with dimension d and the average attention weights of all attention heads from the jth encoding layer, respectively. As the attention mechanism in the LM captures the significance of each token within the context, it can be leveraged to identify the context that is most relevant for the two detected entities
and
, consisting of
and
tokens of the sentence S, respectively. The significance of each token can be derived from the cross-token dependencies matrix
as obtained in Eq. (1). We average the attention scores of the tokens
to calculate the collective attention
of
as follows:
![]() |
2 |
where
and
are the indexes of the first and the last token of the entity
. Analogously, we calculate the attention
of
. Afterward, the relevance of each token for a given entity pair
, represented as
, is calculated using
and
as follows:
![]() |
3 |
where
represents the Hadamard product [62]. Therefore,
illustrates a normalized distribution with range [0, 1] that indicates the importance of each token for
.
Unlike ATLOP, which uses L to enhance hidden representations for relation extraction, our method investigates whether L itself can serve as a signal for detecting semantic relationships in a fully unsupervised setting. We leverage the
distribution to introduce three algorithms for semantic relation detection. The Pick Most Important (PicMI) algorithm (Algorithm 2) focuses on the most significant token of
. The Pick Most Important and Upraise (PicMI-Up) algorithm (Algorithm 3) makes the prediction using the attention scores of each entity to the most significant token. PicMI and PicMI-Up hypothesize that the most important token has a central role in capturing the intrinsic interconnection of the two entities. The Contrast and Examine (ConEx) algorithm (Algorithm 4) relaxes the hypothesis of PicMI and PicMI-Up and assumes that multiple tokens can be important to reflect the associations of the entities. In the case that the attention mechanism of the LM is not informative then the
distribution should be identical to the discrete uniform distribution
as each token of the sentence is equally significant. ConEx is conceptualized based on the assumption that the divergence of
from
indicates that the attention mechanism illustrates dependencies between the entities that bind them semantically. The perspective of the utilization of the attention mechanism is inspired by the findings of Manning et al. [2], who demonstrate that deep contextual LMs encode rich linguistic structures, such as syntactic and hierarchical information, even without explicit supervision. Building on this insight, we argue that attention patterns in such models are not arbitrary, but encode meaningful dependency information that can be leveraged to infer semantic relations.
PicMI. Given a sentence sent with two defined entities
and
, a pre-trained LM lm, a transformer layer id
of the LM, and a threshold t, predict if there is a semantic relation between the entities (Algorithm 2). Calculate the localized context distribution
(
—Eqs. 1, 2, 3). Get the maximum value
of
. If
is above or equal to the threshold then predict that there is a semantic relation between the entities
and
.
Algorithm 2.
PicMI
PicMI-Up. Given a sentence sent with two defined entities
and
, a pre-trained LM lm, a transformer layer id
of the LM, and a threshold t, predict if there is a semantic relation between the entities (Algorithm 3). Calculate the localized context distribution
(
—Eqs. 1, 2, 3). Find the index in of the token with the maximum value of
. Compute the average attention scores
and
, across the heads, of the entities
and
to the token with index in (
—Eq. (2)). If the average of
and
is above or equal to the threshold then predict that there is a semantic relation between the entities
and
.
Algorithm 3.
PicMI-Up
Algorithm 4.
ConEx
ConEx. Given a sentence sent with two defined entities
and
, a pre-trained LM lm, a transformer layer id
of the LM, and a threshold t, predict if there is a semantic relation between the entities (Algorithm 4). Calculate the localized context distribution
(
—Eqs. 1, 2, 3). Find the length len of sent (
) and define the discrete uniform attention score distribution
(
). Compute the Kullback–Leibler (KL) divergence [63, 64]
(
) to estimate the statistical distance and similarity of the two distributions
and
. If
is above or equal to the threshold then predict that there is a semantic relation between the entities
and
.
The threshold in the attention-based algorithms serves as a flexible parameter that can be tuned based on the specific requirements of the task. It enables adaptability by allowing the decision boundary to be adjusted to prioritize either precision or recall. A higher threshold creates a more stringent decision boundary, enhancing precision by ensuring that only the most confident predictions are classified as positive. Conversely, a lower threshold broadens the decision scope, favoring recall and capturing more potential relationships. This aspect provides the versatility needed to optimize the algorithms for diverse use cases and datasets.
Pointwise classification methods
Binary Classification. In the supervised setup, the goal is to train the classifier
by minimizing the classification risk R(f) as follows:
![]() |
4 |
where
is a binary loss function,
and
refer to the positive and negative class prior probability, respectively. The class-conditional probability density of the positive and negative data is defined as
and
, respectively.
Unlabeled-Unlabeled (UU) Classification. Lu et al. [65] and Lu et al. [66] prove that it is possible to train a binary classifier using two unlabeled datasets with different class priors. Under these conditions, Lu et al. [65] demonstrated that the classification risk R(f) can be formulated as follows:
![]() |
5 |
where
and
denote the different class priors of two unlabeled datasets, and
and
correspond to the densities of the unlabeled datasets. The risk estimator of UU classification
is general for binary classification in weakly supervised scenarios.
Data generation mechanism
To frame the knowledge discovery task as an UU classification problem, the data typically needs to be modeled in a pairwise format. A consistent data generation mechanism is necessary to produce pairwise comparison data [4, 67, 68], which consists of pairs of unlabeled instances, where one instance has a higher likelihood of being labeled as positive. Formally, in the pairwise setup, the provided dataset
, where
have the unavailable gold labels
and are expected to satisfy
.
Feng et al. [4] followed the assumption that weakly supervised examples are initially drawn from the gold data distribution, but only the labeler has access to the labels, not the classifier [69]. Hence, the weakly supervised pairwise information is only provided to the classifier. The labeler considers a pair
to be a valid pairwise comparison based on the gold labels
, sampling from
pairs of data whose labels fall into one of three categories:
. The condition
is violated if
[4]. Since the labeler has access to the gold data distribution, we refer to this process as Gold Data Generation (GoDaG). We introduce Soft Data Generation (SoDaG), where the labeler uses silver labels obtained through the unsupervised methods proposed in this paper. SoDaG reduces supervision by transitioning from a weakly supervised to a fully unsupervised setup, where neither the labeler nor the classifier has access to the gold data distribution.
The end goal is to perform pointwise binary classification. Therefore, we divide
as
and
, representing the sets of positive and negative instances, with probability densities
and
, respectively. Feng et al. (2021) [4] proved that pointwise instances in
and
are independently drawn from
and
(Theorem 2 in Feng et al. [4]), indicating that starting from pairwise data, we can independently obtain pointwise instances. The probability densities
and
are formally defined as follows:
![]() |
6 |
Pcomp classification
Feng et al. [4] demonstrated in Theorem 3 of the official paper that the classification risk R(f) (Eq. 4) can be expressed as:
![]() |
7 |
and a classifier can be trained by minimizing the empirical approximation of
as follows:
![]() |
8 |
Lu et al. [66] observed that complex models trained by minimizing
tend to suffer from overfitting [70] due to the issue of negative risk (negative terms in Eq. 8). They proposed the use of consistent correction functions for alleviating the problem. Feng et al. [4] applied these functions and proposed the following alternation of the empirical approximation:
![]() |
9 |
where g(x) is a non-negative function such as the rectified linear unit (ReLU) function
[71] and the absolute value function
.
Noisy-label Learning Perspective. The data sampled from
and
can be considered as noise positive and noise negative data, respectively. Feng et al. [4] discussed noisy-label learning methods [72, 73] and proposed a variation of the RankPruning method [73], which incorporates consistency regularization [74], inspired by the Mean Teacher approach introduced in semi-supervised learning [75].
In summary, this paper evaluates the following methods using the GoDaG process and the newly introduced SoDaG process (silver labels obtained through SARD and ConEx):
Binary-Biased/BER Minimization [76], which minimizes R(f) (Eq. 4) by applying binary classification, using the data from
and
as positive and negative data, respectively.Noisy-Unbiased, which signifies the noisy-label learning approach that reduces the empirical approximation proposed by Natarajan et al. [72].
RankPruning, which is a method for learning from noisy labels introduced by Northcutt et al. [73] and achieves reliable noise estimation.
Pcomp-ReLU [4], which minimizes
(Eq. 9) using ReLU as the risk correction function.Pcomp-ABS [4], which minimizes
(Eq. 9) using the absolute value function as the risk correction function.Pcomp-Teacher [4], which is a variation of the RankPruning method and imposes consistency regularization.
Figure 1 presents the overall workflow of the pointwise classification setup.
Fig. 1.
The diagram illustrates the workflow of the pointwise classification setup. Given a sampled pair of sentences
, each containing two identified entities, the labeler checks whether the condition
holds. The Gold Data Generation (GoDaG) labeler has access to the gold labels, while the Soft Data Generation (SoDaG) labeler uses silver labels obtained through the unsupervised methods proposed in this paper (SARD - algorithm 1 or ConEx - algorithm 4). The labeler constructs a pairwise dataset by selecting pairs that satisfy the condition. Using Theorem 2 by Feng et al. [4] (Eq. 6), the pointwise instances are obtained, starting from the pairwise data. Finally, the resulting pointwise dataset is used to train the pointwise binary classification methods
Experiments
Datasets
We evaluate the methods on four benchmark datasets, including ReDReS, ReDAD [77], GAD [78], and BioInfer [79]. The ReDReS and ReDAD datasets consist of sentences related to Rett Syndrome [80] and Alzheimer’s disease [81, 82], respectively, and incorporate entities with up to 82 different entity types (Table 1). The relation annotations incorporate positive (direct semantic connection), negative (negative semantic connection where negative words like "no" and "absence" are present), complex (semantic connection with complex reasoning), and no relation labels [77]. In the binary setup, the positive, negative, and complex labels are grouped under the relation label. Hence, in the task formulation of the paper, the instances with the relation and no relation labels are considered positive (+1 label) and negative (-1 label), respectively. We use the official splits of the 5-fold cross-validation setup [77].
Table 1.
Statistics of the benchmark datasets
| Dataset | # Instances | # Entity Types | # Relations | # No Relations |
|---|---|---|---|---|
| ReDReS | 5,259 | 73 | 3,314 (63.1%) | 1,945 (36.9%) |
| ReDAD | 8,565 | 82 | 5,495 (64.2%) | 3,070 (35.8%) |
| GAD | 5,330 | 2 | 2,801 (52.5%) | 2,529 (47.5%) |
| Train set | 4,261 | 2 | 2,227 (52.3%) | 2,034 (47.7%) |
| Dev. set | 535 | 2 | 293 (54.8%) | 242 (45.2%) |
| Test set | 534 | 2 | 281 (52.6%) | 253 (47.4%) |
| BioInfer | 9,595 | 6 | 2,516 (26.2%) | 7,079 (73.8%) |
The Genetic Association Database (GAD) corpus [78] is created using a semi-automated approach based on the Genetic Association Archive, which includes lists of gene-disease associations and corresponding sentences from PubMed1 abstracts [83] describing these associations. Bravo et al. [78] employ a biomedical named entity recognition (NER) tool to detect mentions of genes and diseases within the text. Positive samples are derived from sentences with annotated gene-disease associations, while negative samples are generated from gene-disease co-occurrences that are not annotated in the archive (Table 1). For our experiments, we use a preprocessed version of the GAD dataset, along with its training, development, and test split, as provided by Lee et al. [27]. This version is widely used and made available through the Biomedical Language Understanding and Reasoning Benchmark (BLURB) [84].
The BioInfer dataset is a protein-protein interaction (PPI) corpus that employs ontologies to define detailed types of protein entities, such as protein family or group and protein complex as well as their relationships. It consists of sentences with annotations, covering full dependency structures, dependency types, and detailed information on biological entities and their interactions (Table 1). During preprocessing, instances with overlapping entities are excluded. As the dataset lacks predefined training, development, and test splits, we perform 5-fold cross-validation2 for evaluation.
Implementation details
For the dependency-based method (SARD, Algorithm 1), we utilize scispaCy [56], a specialized version of spaCy3 designed for processing biomedical, scientific, and clinical texts. Specifically, we use the en_core_sci_scibert pipeline, which incorporates SciBERT (base version) [31] as the underlying transformer model to extract both dependency parsing and POS tags from the input text. To construct the dependency tree and extract the shortest dependency path (SDP) between the two target entities in a sentence, we employ the NetworkX library [85].
For the attention-based methods (PicMI, PicMI-Up, ConEx - Algorithms 2, 3, 4), we utilize BiomedBERT (base version) [84, 86] as the pre-trained LM, accessed via HuggingFace’s Transformers library [87]. BiomedBERT is specifically designed for biomedical text, having been pre-trained on the PubMed4 corpus. In the probing experiments, Theodoropoulos et al. [77] reveal that the localized context vector, extracted from the 10th and 11th encoding layer, provides informative representations for the relation detection task. Hence, our analysis focuses on these particular encoding layers of BiomedBERT.
For the pointwise classification methods, we define the classifier using BiomedBERT (base version) as the backbone language model LM. Let the tokens of a given sentence S be
. Let the tokens of the detected entities
and
of the sentence be
and
with the corresponding index spans
and
. The classifier f is defined as follows:
![]() |
10 |
![]() |
11 |
![]() |
12 |
![]() |
13 |
![]() |
14 |
where
represent the token embeddings extracted from the last encoding layer of BiomedBERT, and the entity representations
and
are condensed via average pooling of the token embeddings that correspond to each entity. Following, the representation
is defined by the concatenation (||) of the entity representations
and
and passed through a dropout layer d() [88]. The final output of the classifier y is extracted from a fully connected layer with a weight vector
and a bias term
. We apply batch normalization [89] after the extraction of the LM embeddings (Eq. 13).
We set the dropout probability to 0.3 and use logistic loss
as the binary loss function and Adam [90] as the optimizer with learning rate of 10-3 and mini-batch size set to 256. We train the models for 50 epochs, retaining the best scores based on the performance on the development set (15% of the train set) and conducting the experiments on a NVIDIA RTX 3090 GPU 24GB. During training, BiomedBERT is kept frozen, with only the final encoding layer being trainable to maintain a controlled experimental setup. We implement the methods using PyTorch [91].
Feng et al. [4] mention that the positive class prior
can be estimated according to the GoDaG data generation process. Specifically,
can be determined by counting the fraction of collected pairwise comparison data in all sampled pairs of data, allowing for exact estimation of the true class priors if we know whether
is larger than
. However, this exact estimation requires additional knowledge about the dataset’s gold distribution, specifically which class prior is dominant. In contrast, we don’t make this assumption in our approach, hypothesizing that we do not have access to such information about the data distribution, as we are also evaluating the SoDaG process, where only silver labels generated by unsupervised methods are available.
We experiment with different values for
, selecting from the set {0.3, 0.4, 0.5, 0.6} to evaluate performance across varying assumptions about the class prior. Additionally, unlike Feng et al. [4], we do not assume that the test set has a distribution similar to or identical to the train set. For every run, we evaluate the performance on the original test set without any sampling. This introduces a more challenging setup, where the robustness of the different methods is evaluated against potentially varying data distributions. By testing multiple class priors, we provide a more comprehensive analysis of how different methods generalize under uncertain and shifting data conditions. When employing the SoDaG process, we utilize the silver labels provided by the best-performing unsupervised approaches of the dependency-based and attention-based methods: the SARD algorithm (with assumption 3. and heuristic 1.) and the ConEx algorithm.
Time Complexity. We clarify that the unsupervised algorithms - SARD, PicMI, PicMI-Up, and ConEx - do not involve any training phase, and their complexity is solely related to the inference step. For the attention-based algorithms (PicMI, PicMI-Up, and ConEx), the primary computational cost arises from the forward pass through the BiomedBERT-base model, used as the backbone language model. Therefore, the total time complexity for these algorithms is
, where L is the number of transformer layers (12 for BiomedBERT-base), n is the sequence length (number of tokens), and d is the hidden embedding dimension size (768 for BiomedBERT-base). For the SARD algorithm, generating contextualized embeddings via the en_core_sci_scibert transformer-based pipeline also has a complexity of
. The additional part-of-speech tagging and dependency parsing steps have a linear complexity of
. Thus, the overall time complexity for SARD remains dominated by the
term. For the pointwise classification methods, the inference phase has the same time complexity,
, since it relies on the BiomedBERT forward pass (Eq. 10). The training phase remains efficient, as BiomedBERT’s parameters (110 million) are kept frozen, with only the final encoding layer being trained, significantly reducing the number of trainable parameters and the overall training cost.
Overall, we emphasize that the efficient time complexity of the proposed methods supports their effective usability in practical applications. We highlight that all experiments are conducted using a user-accessible NVIDIA RTX 3090 GPU (24 GB), without the need for specialized or high-end server-grade GPUs.
Results
For the experimental evaluation, we use Precision, Recall, and F1-Score. We note that F1-Score is widely used for relation extraction evaluation in the literature, but it may not fully capture the balance between precision and recall needed to assess a method’s effectiveness. Table 2 presents the results of the SARD algorithm using different assumptions and heuristics. The best-performing setting of SARD defines the baseline performance. We establish the upper-performance boundary with a fully supervised model. Theodoropoulos et al. [77] conduct extensive benchmarking, identifying LaMReDA, utilizing BiomedBERT-base, as the top-performing model, using the relation representation M for ReDReS and K for ReDAD. To provide a strong upper limit also for BioInfer and GAD, we train LaMReDA with the relation representation M, achieving performance competitive with SOTA models [33, 92].
Table 2.
Results (%) of SARD (Algorithm 1) on the four benchmark datasets
| Data | Assumption | Heuristic | Precision | Recall | ![]() |
|---|---|---|---|---|---|
| ReDReS | 1 | 1 | 57.24/57.8 | 46.81/46.85 | 51.5/51.6 |
| 2 | 66.7/67.19 | 36.95/36.94 | 47.56/47.51 | ||
| 2 | 1 | 58.31/58.91 | 54.43/54.54 | 56.31/56.48 | |
| 2 | 66.63/67.13 | 41.11/41.18 | 50.85/50.86 | ||
| 3 | 1 | 60.48/60.77 | 79.69/79.82 | 68.76/68.93 | |
| 2 | 70.58/70.92 | 57.93/58.05 | 63.63/63.67 | ||
| ReDAD | 1 | 1 | 59.25/60.42 | 47.22/47.31 | 52.56/52.55 |
| 2 | 70.92/71.7 | 33.38/33.46 | 45.39/45.36 | ||
| 2 | 1 | 60.74/61.52 | 55.32/55.34 | 57.9/57.93 | |
| 2 | 71.75/72.23 | 37.85/37.92 | 49.56/49.53 | ||
| 3 | 1 | 60.89/62.01 | 75.12/75.31 | 67.26/67.47 | |
| 2 | 74.42/74.96 | 47.75/48.02 | 58.18/58.19 | ||
| GAD | 1 | 1 | 52.8/56.94 | 41.13/42.35 | 46.24/48.57 |
| 2 | 52.21/58.43 | 32.45/34.52 | 40.03/43.4 | ||
| 2 | 1 | 52.06/54.55 | 47.41/46.98 | 49.63/50.48 | |
| 2 | 51.63/55.79 | 37.34/37.72 | 43.34/45.01 | ||
| 3 | 1 | 53.68/53.17 | 75.79/77.58 | 62.85/63.1 | |
| 2 | 52.92/54.35 | 57.98/62.28 | 55.33/58.04 | ||
| BioInfer | 1 | 1 | 27.88/28.01 | 57.15/57.07 | 37.48/37.53 |
| 2 | 29.99/30.13 | 47.69/47.61 | 36.82/36.85 | ||
| 2 | 1 | 27.82/27.84 | 64.11/64.01 | 38.81/38.75 | |
| 2 | 30.37/30.43 | 51.71/51.58 | 38.26/38.24 | ||
| 3 | 1 | 28.98/28.96 | 78.93/78.77 | 42.39/42.35 | |
| 2 | 32.49/32.51 | 60.45/60.31 | 42.27/42.2 |
Each cell shows the performance on the full datasets and the average performance on the test set of the 5-fold cross-validation setup for ReDReS, ReDAD, and BioInfer. For the GAD dataset, the second value of each cell corresponds to the performance on the official test set. The best performance is highlighted in bold.
For the PicMI and PicMI-Up algorithms, we use threshold values ranging from [0.3, 0.7] and [0.2, 0.6], respectively, with a step size of 0.05. For the ConEx algorithm, we experiment with decision boundaries based on the KL divergence within the range of [0.05, 0.14], incrementing by 0.01. This stepwise approach enables a systematic examination of how changes in the decision boundary affect the trade-off between precision and recall, offering insights into the optimal value for maximizing performance while balancing the detection of relevant instances (Recall) and the reduction of false positives (Precision) (Figs. 2, 3, and 4).
Fig. 2.
The figure shows the results (%) of PicMI (Algorithm 2), across the four benchmark datasets, illustrating the Precision, Recall, and F1-score achieved using attention scores from the 10th and 11th encoding layers. The baseline is established by SARD, using the third assumption and first heuristic, while the upper boundary of performance is defined by the LaMReDA model
Fig. 3.
The figure shows the results (%) of PicMI-Up (Algorithm 3), across the four benchmark datasets, illustrating the Precision, Recall, and F1-score achieved using attention scores from the 10th and 11th encoding layers. The baseline is established by SARD, using the third assumption and first heuristic, while the upper boundary of performance is defined by the LaMReDA model
Fig. 4.
The figure shows the results (%) of ConEx (Algorithm 4), across the four benchmark datasets, illustrating the Precision, Recall, and F1-score achieved using attention scores from the 10th and 11th encoding layers. The baseline is established by SARD, using the third assumption and first heuristic, while the upper boundary of performance is defined by the LaMReDA model
The experimental results for the pointwise classification methods (Pcomp-Unbiased, Pcomp-ReLU, Pcomp-ABS, Pcomp-Teacher, Binary-Biased, Noisy-Unbiased, and RankPruning) are illustrated in Figs. 5, 6, and 7. The setups differ in their data generation processes. In Fig. 5, the GoDaG process is used, where the labeler has access to the gold data distribution. Hence, the classifier operates under a weakly supervised paradigm. In Figs. 6 and 7, the SoDaG process is employed, which uses silver labels generated by the SARD and ConEx algorithms instead of gold labels. This approach shifts the classifier to a fully unsupervised setup. In detail, the silver labels are provided by the SARD algorithm in Fig. 6, specifically using the third assumption and first heuristic. In Fig. 7, the silver labels are generated by the ConEx algorithm, with specific configurations for each dataset to optimize the balance between precision and recall. The configurations for the ConEx setup are as follows:
ReDReS: 10th layer, decision boundary: 0.08
ReDAD: 10th layer, decision boundary: 0.07
GAD: 11th layer, decision boundary: 0.07
BioInfer: 10th layer, decision boundary: 0.08
Fig. 5.
The figure shows the performance results (%) of different methods (Pcomp-Unbiased, Pcomp-ReLU, Pcomp-ABS, Pcomp-Teacher, Binary-Biased, Noisy-Unbiased, and RankPruning) across the four benchmark datasets, using the GoDaG mechanism for data generation. The baseline is set by SARD with the third assumption and first heuristic, while LaMReDA defines the upper-performance limit
Fig. 6.
The figure shows the results (%) of different methods (Pcomp-Unbiased, Pcomp-ReLU, Pcomp-ABS, Pcomp-Teacher, Binary-Biased, Noisy-Unbiased, and RankPruning) across the four benchmark datasets, using the SoDaG mechanism for data generation and the silver labels acquired by SARD with the third assumption and first heuristic (baseline performance)
Fig. 7.
The figure shows the results (%) of different methods (Pcomp-Unbiased, Pcomp-ReLU, Pcomp-ABS, Pcomp-Teacher, Binary-Biased, Noisy-Unbiased, and RankPruning) across the four benchmark datasets, using the SoDaG mechanism for data generation and the silver labels acquired by ConEx (baseline performance)
Discussion
Initially, we analyze the results for the algorithms SARD, PicMI, PicMI-Up, and ConEx provide valuable insights into the effectiveness and limitations of each method for semantic relation detection across multiple benchmark datasets. Then, we discuss the experimental results using the GoDaG data generation process and evaluate the performance of different classification methods under the weakly supervised setup. Next, we observe the experiments using the SoDaG data generation process, highlighting the impact of label quality on model performance and revealing patterns regarding the effectiveness of various methods under the unsupervised setup. Finally, we demonstrate a comparative analysis between the different learning paradigms.
SARD. The results indicate that using the third assumption, which involves the presence of a verb in the SDP, consistently improves performance across all datasets (Table 2). This suggests that verbs serve as strong indicators of semantic relationships between entities. In contrast, the stricter first and second assumptions, which emphasize the root word in the SDP, prioritize precision over recall without yielding substantial gains in the F1-score, indicating limited benefit in focusing solely on root words. The adoption of the first heuristic leads to better results than the second heuristic, which is more restrictive. The findings imply that conjunctions in the SDP do not indicate a lack of semantic connection between entities, making the first heuristic more effective for enhancing both precision and recall.
PicMI. The PicMI algorithm struggles to find an optimal balance between precision and recall (Fig. 2). Increasing the decision threshold improves precision but significantly reduces recall. Despite this challenge, PicMI outperforms the baseline on several datasets (excluding BioInfer) under specific conditions, mainly driven by the very high recall when the decision threshold is relatively low. This indicates that the algorithm is particularly effective when prioritizing recall, although this may come at the cost of precision.
PicMI-Up. Compared to PicMI, PicMI-Up demonstrates better robustness in managing the precision-recall trade-off, suggesting that predicting based on the attention scores of each entity to the most significant token of
is a more appropriate strategy. A decision threshold within the range of [0.25, 0.35] generally offers a balanced performance, although the optimal threshold varies by dataset (Fig. 3).
ConEx. The ConEx algorithm exhibits higher performance compared to both PicMI and PicMI-Up, indicating that the relaxation of PicMI’s and PicMI-Up’s assumptions, along with the consideration of multiple tokens in the context distribution, enhances the algorithm’s flexibility and robustness. In summary, while PicMI and PicMI-Up demonstrate the value of focused attention mechanisms, ConEx highlights the benefits of greater contextual flexibility, strengthening our overall methodological contribution. This is particularly evident in the smoother changes in precision and recall as the KL divergence threshold increases, as shown in Fig. 4. ConEx achieves higher gains over the baseline across more settings, highlighting its robustness to different decision boundary values. Moreover, the algorithm’s performance is good for a wider range of threshold values when using the 11th encoding layer, suggesting that this layer provides a more consistent solution, making the choice of the decision boundary less critical.
Experimentation with GoDaG data generation process. The Pcomp-Unbiased and Noisy-Unbiased methods consistently achieve the best performance across the four benchmark datasets (Fig. 5). For ReDReS, ReDAD, and GAD, these methods demonstrate performance comparable to LaMReDA, the supervised upper boundary, indicating that restricting the classifier’s access to the gold data distribution does not substantially compromise the results. This finding suggests that weak supervision can still be effective for relation classification, particularly when GoDaG allows the labeler to perform sampling, reflecting the gold distribution.
The stability of Pcomp-Unbiased and Noisy-Unbiased methods across different prior probabilities indicates that these approaches are robust to variations in the data distribution. This is promising since exact prior distribution estimation requires knowledge of the dataset’s gold distribution, including which class (positive or negative) dominates. The results suggest that even a weak prior signal can be sufficient to achieve competitive performance. However, an exception is noted with the BioInfer dataset, where higher
values lead to a significant performance drop. This outcome may point to characteristics unique to the BioInfer dataset, such as a more imbalanced class distribution, which affects the robustness of the methods when the likelihood of sampling positive data points is increased.
As expected, increasing the
value generally favors recall over precision across all methods tested (Fig. 5). This trend aligns with the theoretical expectation that higher probabilities of sampling positive data points increase the likelihood of detecting relevant instances (higher recall), albeit sometimes at the cost of more false positives (lower precision). The Binary-Biased method shows the lowest F1-score, indicating that binary classification is not an efficient approach for pairwise comparison-based relation classification (Fig. 5). In contrast with the findings of Feng et al. [4], the Pcomp-Teacher method does not perform well across the datasets, which challenges the incorporation of consistency regularization in this context. The effectiveness of the Pcomp-Teacher approach seems to depend on the teacher model’s performance, and when the teacher is not highly reliable, consistency regularization fails to yield performance improvements.
The Pcomp-ReLU and Pcomp-ABS methods, which use consistent correction functions to prevent negative empirical risk values, are less effective than Pcomp-Unbiased (Fig. 5). This suggests that enforcing non-negative risk can lead to underfitting, causing the classifier to fail to capture the nuances of the data. By avoiding negative values, these methods may impose conservative adjustments, reducing the classifier’s ability to adapt to the variability in the data and ultimately limiting performance.
Experimentation with SoDaG data generation process. When the silver labels are generated using the SARD algorithm with the third assumption and first heuristic, the Rank-Pruning and Pcomp-ReLU methods achieve the best results across all four benchmark datasets (Fig. 6). Pcomp-ABS also performs well, particularly at higher
values. These methods seem better suited for handling noisy labels compared to others like Pcomp-Unbiased and Noisy-Unbiased, which showed strong results in the GoDaG experiments but struggled with the SoDaG-generated data. The better performance of Rank-Pruning implies that this method’s ability to filter out noise and focus on more reliable predictions gives it an advantage in noisy label environments.
The experiments show that while many methods can boost recall across the datasets, precision remains relatively low. In some cases, however, precision levels still surpass the baseline, indicating that even with noisy labels, certain methods can enhance performance. Across all datasets, several methods manage to outperform the SARD baseline, indicating that training a classifier using the SoDaG process with silver labels can improve unsupervised performance (Fig. 6). This is especially evident in the GAD and BioInfer datasets, where almost all methods (except for Pcomp-Teacher and Binary-Biased, respectively) show an increase in F1-score compared to the baseline. These results suggest that even when using less accurate silver labels, the classifiers can still learn meaningful patterns in the data, potentially boosting their ability to generalize. The observation of performance enhancement beyond the labeler’s (SARD) performance constitutes a notable research finding.
The utilization of silver labels generated by the ConEx algorithm reveals analogous performance patterns. Rank-Pruning, Pcomp-ReLU, and Pcomp-ABS outperform the baseline or closely match it in the ReDReS and ReDAD datasets (Fig. 7). However, Pcomp-ABS displays more sensitivity to the definition of the
value, with performance declining when a lower prior is used. Overall, the ConEx algorithm establishes a stronger baseline than SARD and we notice that the benefits of training the classifier using SoDaG-generated labels are more pronounced in the GAD and BioInfer datasets.
The overall results from the SoDaG experiments stress the importance of silver label and priors quality. Compared to gold labels (GoDaG), using silver labels introduces noise, which affects the performance of the methods tested. The different classification approaches do not show a very high tolerance to label noise in this study, suggesting that accessing high-quality silver labels is crucial for strong performance. These findings emphasize that the method for creating silver labels can significantly influence the success of unsupervised learning approaches.
Comparison of learning paradigms. In the unsupervised setting, ConEx outperforms other methods such as SARD, PicMI, and PicMI-Up across the benchmark datasets (Tables 3 and 4). Notably, for the ReDReS and ReDAD datasets, ConEx achieves the best F1-score performance, indicating its robustness in identifying semantic relations without supervision. The results suggest that ConEx’s attention-based approach, which allows for the consideration of multiple important tokens in context, offers a significant advantage over the more rigid assumptions of the other algorithms. For the GAD and BioInfer datasets, Pcomp-ABS and RankPruning, using the SoDaG data generation process, achieve the highest F1-score in the unsupervised setup. This highlights the potential benefits of training classifiers using silver labels with good priors. The ability to utilize noisy silver labels for data sampling provides a valuable alternative when clean, annotated datasets are unavailable. These results suggest that with the right methods, the use of silver labels can lead to improvements even in challenging scenarios.
Table 3.
Best results (%) of the different methods on ReDReS and ReDAD
| Data | Method/Model | Type | Precision | Recall | ![]() |
|---|---|---|---|---|---|
| ReDReS | SARD
|
Unsupervised | 60.77 | 79.82 | 68.93 |
PicMI
|
Unsupervised | 63.06 | 98.64 | 76.87 | |
PicMI-Up
|
Unsupervised | 62.88 | 97.39 | 76.34 | |
ConEx
|
Unsupervised | 63.92 | 98.02 | 77.3 | |
Pcomp-ABS
|
Unsupervised | 63.64 | 96.71 | 76.66 | |
Pcomp-Unbiased
|
Weakly Supervised | 84.5 | 91.37 | 87.77 | |
| LaMReDA | Supervised | 88.94 | 91.71 | 90.27 | |
| ReDAD | SARD
|
Unsupervised | 62.01 | 75.31 | 67.47 |
PicMI
|
Unsupervised | 64.83 | 99.63 | 78.2 | |
PicMI-Up
|
Unsupervised | 66.23 | 96.38 | 78.32 | |
ConEx
|
Unsupervised | 68.78 | 96 | 80.08 | |
Pcomp-ABS
|
Unsupervised | 69.09 | 91.96 | 78.67 | |
Noisy-Unbiased
|
Weakly Supervised | 84.56 | 91.6 | 87.82 | |
| LaMReDA | Supervised | 88.94 | 93.22 | 91.01 |
Each cell presents the average performance on the test set of the 5-fold cross-validation setup. The best performance across the unsupervised methods is highlighted in bold.
Assumption 3. and heuristic 1
11th encoding layer, threshold: 0.35
10th encoding layer, threshold: 0.2
10th encoding layer, threshold: 0.05

: 0.6, SoDaG process - silver labels provided by ConEx (10th layer, threshold: 0.07)

: 0.5, GoDaG process
11th encoding layer, threshold: 0.4
10th encoding layer, threshold: 0.25

: 0.6, SoDaG process - silver labels provided by SARD

: 0.6, GoDaG process
Table 4.
Best results (%) of the different methods on GAD and BioInfer
| Data | Method/Model | Type | Precision | Recall | ![]() |
|---|---|---|---|---|---|
| GAD | SARD
|
Unsupervised | 53.17 | 77.58 | 63.1 |
PicMI
|
Unsupervised | 54.2 | 87.19 | 66.85 | |
PicMI-Up
|
Unsupervised | 52.58 | 97.86 | 68.41 | |
ConEx
|
Unsupervised | 53.58 | 98.58 | 69.42 | |
Pcomp-ABS
|
Unsupervised | 56.82 | 94.77 | 70.91 | |
Pcomp-Unbiased
|
Weakly Supervised | 73.95 | 89 | 80.69 | |
| LaMReDA | Supervised | 77.43 | 87.4 | 82.01 | |
| BioInfer | SARD
|
Unsupervised | 28.96 | 78.77 | 42.35 |
PicMI
|
Unsupervised | 27.39 | 91.93 | 42.15 | |
PicMI-Up
|
Unsupervised | 31.09 | 70.77 | 43.17 | |
ConEx
|
Unsupervised | 33.61 | 70.68 | 45.49 | |
RankPruning |
Unsupervised | 33.82 | 84.49 | 48.25 | |
Pcomp-Unbiased |
Weakly Supervised | 72.2 | 70.09 | 71.06 | |
| LaMReDA | Supervised | 79.16 | 77.7 | 78.22 |
Each cell presents the average performance on the test set of the 5-fold cross-validation setup for BioInfer. For the GAD dataset, each cell corresponds to the performance on the official test set. The best performance across the unsupervised methods is highlighted in bold.
Assumption 3. and heuristic 1
11th encoding layer, threshold: 0.4
11th encoding layer, threshold: 0.25
11th encoding layer, threshold: 0.06

: 0.5, SoDaG process - silver labels provided by ConEx (11th layer, threshold: 0.07)

: 0.3, GoDaG process
10th encoding layer, threshold: 0.5
10th encoding layer, threshold: 0.3
10th encoding layer, threshold: 0.07

: 0.4, SoDaG process - silver labels provided by SARD

: 0.4, GoDaG process
In the weakly supervised learning setup, some methods approach the performance of the fully supervised LaMReDA model, except in the BioInfer dataset (Tables 3 and 4). This indicates that weak supervision can be effective in bridging the gap between unsupervised and fully supervised learning. The larger performance gap observed in the BioInfer dataset suggests that it poses a greater challenge, possibly due to higher data complexities. This limitation demonstrates that while weak supervision can significantly enhance performance, it may not be sufficient for all datasets, especially those with more intricate or ambiguous relationships. Future research that addresses ambiguity through higher-level reasoning, potentially incorporating LLMs, may extend the potential of reduced-supervision approaches. One notable observation is the performance decline of ConEx in the ReDAD dataset, where the F1-score decreases by 8.8% and 12% compared to its weakly supervised and supervised counterparts, respectively. While there is a performance gap, the relatively small decline suggests that ConEx’s fully unsupervised approach remains competitive even when compared to methods with varying degrees of supervision.
Conclusion
In this study, we introduce a set of unsupervised algorithms based on dependency trees and attention mechanisms, aimed at reducing reliance on annotated data for identifying semantic relationships between biomedical entities. This approach addresses a key challenge in knowledge discovery: balancing performance with minimized supervision, which is essential for adapting models across diverse and evolving domains. Our work also explores the applications of pointwise binary classification methods in a weakly supervised context for knowledge discovery. By progressively reducing the level of supervision, we test the robustness of these methods in handling noisy labels, demonstrating the potential of transitioning from weakly supervised to fully unsupervised setups. The extensive benchmarking conducted on four biomedical datasets provided valuable insights into the performance of these methods, discussing the adaptability and reliability of unsupervised approaches in capturing complex relationships in biomedical text. These findings highlight a promising pathway toward scalable, adaptable knowledge discovery systems, marking a step forward in developing data-efficient approaches capable of extracting critical insights in low annotated data resource scenarios.
Appendix A: Zero-shot learning
We extend our experiments by employing a zero-shot learning approach using additional LLMs following the paradigm of Wadhwa et al. [34] and Gao et al. [37]. Specifically, we utilize the open-source model BioMistral 7B [94], a medical adaptation of Mistral 7B Instruct v0.1 [95] trained on the PMC Open Access Subset.5 This model is selected for its superior performance across various medical learning tasks compared to other open-source medical LLMs [94].
To investigate potential performance gains from incorporating general-purpose reasoning, we include three merged models, BioMistral 7B TIES, BioMistral 7B DARE, and BioMistral 7B SLERP, developed by Labrak et al. [94] through specific merging techniques that combine Mistral 7B Instruct and BioMistral 7B models. SLERP [96] uses Spherical Linear Interpolation for smooth parameter transitions, minimizing information loss compared to direct weight averaging. TIES [97] extracts unique contributions by generating sparse "task vectors" by subtracting a common base model such as Mistral 7B Instruct and reduces interference with a method of sign consensus. TIES combines models by generating "task vectors" from each model, extracting distinct contributions by subtracting a common base model such as Mistral 7B Instruct. These vectors are subsequently averaged with the base model. The primary enhancement over earlier techniques is achieved by diminishing model interference through the use of sparse vectors and a method of sign consensus [94]. DARE [98] builds on TIES by pruning redundant parameters while preserving or enhancing model performance. In our setup, we set the temperature parameter to zero (or do_sample to False) to reduce creativity and ensure deterministic text outputs.
Building on the prompt engineering of [37], we use the following prompts per dataset:
-
ReDReS and ReDAD: TASK: The task is to classify relations between two entities in a sentence.
INPUT: The input is a sentence where the two entities are included with the special tokens [ent] and [/ent]. The start token is [ent]. The end token is [/ent].
OUTPUT: Your task is to select only one relation type (0, 1) for the two entities:
0, when the sentence conveys no relationship between the entities
1, when the sentence conveys a positive, negative, or complex relationship between the entities
Select only one relation type (0, 1) for the two entities in the provided sentence.
Sentence: sentence
Relation:

-
GAD:
TASK: The task is to classify relations between a gene and a disease in a sentence.
INPUT: The input is a sentence where the gene is labeled as @GENE and the disease is labeled as @DISEASE.
OUTPUT: Your task is to select only one relation type (0, 1) for the two entities:
0, when the sentence conveys no relation between the gene and disease
1, when the sentence when the sentence conveys a positive or negative relation between the gene and disease
Select only one relation type (0, 1) for the two entities in the provided sentence.
Sentence: sentence
Relation:

-
BioInfer:
TASK: The task is to classify relations between two proteins in a sentence.
INPUT: The input is a sentence where the two proteins are included with the special tokens [ent] and [/ent]. The start token is [ent]. The end token is [/ent].
OUTPUT: Your task is to select only one relation type (0, 1) for the two entities:
0, when the sentence conveys no protein-protein interaction between two proteins
1, when the sentence when the sentence conveys a protein-protein interaction between two proteins
Select only one relation type (0, 1) for the two entities in the provided sentence.
Sentence: sentence
Relation:

The results of our study reveal that the best-performing unsupervised method outperforms all BioMistral 7B variants in terms of F1-score across every benchmark dataset (Table 5). This performance gap is most profound on the GAD dataset, followed by BioInfer, highlighting the competitiveness and effectiveness of our proposed methods. This observation holds significant computational implications, as the unsupervised methods of the study (ConEx, RunkPruning, etc.) leverage BiomedBERT (base version) [84, 86] as the backbone LM, that has 110 million parameters, constituting only approximately 1.57% of BioMistral’s size (7 billion parameters). Additionally, it is noteworthy that BioMistral exhibits consistency challenges in generation, due to hallucination issues that are common in LLMs [99, 100], and the fact that minor, seemingly inconsequential modifications in prompt structure, such as reordering of relational descriptions, may lead to different outputs. The absence of determinism is, on one hand, a notable advantage of LLMs for generating text with complexity, yet on the other hand, it may hinder their integration into systems with strictly defined guidelines and protocols. Nevertheless, the use of BioMistral 7B in a zero-shot setup shows promise, as it achieves a decent level of performance overall. A comparison between the merged versions and the original BioMistral 7B indicates that the incorporation of general-purpose reasoning is not beneficial and may, in some cases, lead to a performance decline, as observed on the GAD dataset.
Table 5.
Results (%) of the different BioMistral 7B versions on the four benchmark datasets in the zero-shot setting
| Data | Method/Model | Type | Precision | Recall | ![]() |
|---|---|---|---|---|---|
| ReDReS | BioMistral 7B | zero-shot | 64.01 | 96.76 | 76.97 |
| BioMistral 7B DARE | zero-shot | 63.22 | 98.21 | 76.84 | |
| BioMistral 7B TIES | zero-shot | 63.15 | 97.81 | 76.67 | |
| BioMistral 7B SLERP | zero-shot | 63.4 | 97.06 | 76.6 | |
ConEx
|
Unsupervised | 63.92 | 98.02 | 77.3 | |
| ReDAD | BioMistral 7B | zero-shot | 65.97 | 98.57 | 78.9 |
| BioMistral 7B DARE | zero-shot | 64.82 | 99.48 | 78.15 | |
| BioMistral 7B TIES | zero-shot | 64.85 | 98.5 | 77.92 | |
| BioMistral 7B SLERP | zero-shot | 64.83 | 99.41 | 78.13 | |
ConEx
|
Unsupervised | 68.78 | 96 | 80.8 | |
| GAD | BioMistral 7B | zero-shot | 52.03 | 86.83 | 65.07 |
| BioMistral 7B DARE | zero-shot | 50.64 | 70.46 | 58.93 | |
| BioMistral 7B TIES | zero-shot | 51.28 | 78.29 | 61.97 | |
| BioMistral 7B SLERP | zero-shot | 51.24 | 66.19 | 57.76 | |
Pcomp-ABS
|
Unsupervised | 56.82 | 94.77 | 70.91 | |
| BioInfer | BioMistral 7B | zero-shot | 27.62 | 93.05 | 42.57 |
| BioMistral 7B DARE | zero-shot | 28.97 | 90.3 | 43.83 | |
| BioMistral 7B TIES | zero-shot | 27.91 | 94.05 | 43.02 | |
| BioMistral 7B SLERP | zero-shot | 31.41 | 77.93 | 44.71 | |
RankPruning
|
Unsupervised | 33.82 | 84.49 | 48.25 |
Each cell presents the average performance on the test set of the 5-fold cross-validation setup for ReDReS, ReDAD, and BioInfer. For the GAD dataset, each cell corresponds to the performance on the official test set. The last row of each dataset presents the best performance based on F1-score of the unsupervised methods (Tables 3, 4). The best performance is highlighted in bold.
10th encoding layer, threshold: 0.05

: 0.5, SoDaG process - silver labels provided by ConEx (11th layer, threshold: 0.07)

: 0.4, SoDaG process - silver labels provided by SARD
Author contributions
Christos Theodoropoulos designed the work, implemented the code, conducted the experiments, and wrote the manuscript. Andrei Catalin Coman, James Henderson, and Marie-Francine Moens contributed to the conceptualization of the research direction, reviewed and revised the paper. Marie-Francine Moens contributed to the interpretation of the results. All authors have approved the submitted version and have agreed both to be personally accountable for the author’s own contributions and to ensure that questions related to the accuracy or integrity of any part of the work, even ones in which the author was not personally involved, are appropriately investigated, resolved, and the resolution documented in the literature.
Funding
This work is supported by the Research Foundation—Flanders (FWO) and Swiss National Science Foundation (SNSF), through the grants 200021E_189458 and G094020N.
Availability of data and materials
The GAD dataset [78] used in the current study is available in the BioBERT repository [27], https://github.com/dmis-lab/biobert and in the PPI-Relation-Extraction repository [93], https://github.com/BNLNLP/PPI-Relation-Extraction. The BioInfer dataset [79] used in the current study is available in the PPI-Relation-Extraction repository [93], https://github.com/BNLNLP/PPI-Relation-Extraction and in the Hugging Face dataset repository, https://huggingface.co/datasets/bigbio/bioinfer. The ReDReS and ReDAD datasets [77] are available in the Enhancing-Biomedical-Knowledge-Discovery-for-Diseases repository, https://github.com/christos42/Enhancing-Biomedical-Knowledge-Discovery-for-Diseases.
Declarations
Ethics approval and consent to participate
Not applicable.
Consent for publication
Not applicable.
Competing interests
The authors declare that they have no competing interests.
Footnotes
The exact splits will be released for fair comparison.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Kübler S, McDonald R, Nivre J. Dependency parsing, 2009;11–20.
- 2.Manning CD, Clark K, Hewitt J, Khandelwal U, Levy O. Emergent linguistic structure in artificial neural networks trained by self-supervision. Proc Natl Acad Sci. 2020;117:30046–54. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Vaswani A, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
- 4.Feng L, et al. Pointwise binary classification with pairwise confidence comparisons. In: Meila M, Zhang T editors. Proceedings of the 38th international conference on machine learning, Vol. 139, Proceedings of machine learning research. p. 3252–62. PMLR; 2021. https://proceedings.mlr.press/v139/feng21d.html.
- 5.Proux D, Rechenmann F, Julliard L. A pragmatic information extraction strategy for gathering data on genetic interactions. In: Bourne PE, et al. editors. Proceedings of the 8th international conference on intelligent systems for molecular biology. p.279–85. AAAI Press, San Diego; 2000. http://www.aaai.org/Library/ISMB/2000/ismb00-029.php. [PubMed]
- 6.Ono T, Hishigaki H, Tanigami A, Takagi T. Automated extraction of information on protein-protein interactions from the biological literature. Bioinformatics. 2001;17:155–61. [DOI] [PubMed] [Google Scholar]
- 7.Phuong TM, Lee D, Lee KH. Learning rules to extract protein interactions from biomedical text. In: Whang KY, Jeon J, Shim K, Srivastava J editors. Proceedings of the 7th Pacific-Asia conference on knowledge discovery and data mining, vol. 2637. Heidelberg: Springer; 2003. p. 148–58.
- 8.Huang M, et al. Discovering patterns to extract protein-protein interactions from full texts. Bioinformatics. 2004;20:3604–12. [DOI] [PubMed] [Google Scholar]
- 9.Hao Y, Zhu X, Huang M, Li M. Discovering patterns to extract protein-protein interactions from the literature: Part ii. Bioinformatics. 2005;21:3294–300. [DOI] [PubMed] [Google Scholar]
- 10.Chun HW, Hwang YS, Rim HC. Unsupervised event extraction from biomedical literature using co-occurrence information and basic patterns. In: Su KY, Tsujii J, Lee JH, Kwong OY editors. Proceedings of the 1st international conference on natural language processing. Berlin Heidelberg: Springer; 2004. p. 777–86.
- 11.Liu H, Blouin C, Kešelj V. Identifying interaction sentences from biological literature using automatically extracted patterns. In: Cohen KB et al. editors. Proceedings of the BioNLP 2009 workshop. p. 133–41. Association for Computational Linguistics, Boulder, Colorado; 2009. https://aclanthology.org/W09-1317.
- 12.Hakenberg J, et al. Efficient extraction of protein-protein interactions from full-text articles. IEEE/ACM Trans Comput Biol Bioinf. 2010;7:481–94. [DOI] [PubMed] [Google Scholar]
- 13.Thomas P, Pietschmann S, Solt I, Tikk D, Leser U. Not all links are equal: Exploiting dependency types for the extraction of protein-protein interactions from text. In: Cohen KB, et al. editors. Proceedings of BioNLP 2011 workshop. p. 1–9. Association for Computational Linguistics, Portland, Oregon; 2011. https://aclanthology.org/W11-0201.
- 14.Blaschke C, Valencia A. The frame-based module of the SUISEKI information extraction system. IEEE Intell Syst. 2002;17:14–20. [Google Scholar]
- 15.Raja K, Subramani S, Natarajan J. PPInterFinder-a mining tool for extracting causal relations on human proteins from literature. Database. 2013;2013. [DOI] [PMC free article] [PubMed]
- 16.Craven M, Kumlien J. Constructing biological knowledge bases by extracting information from text sources. In: Lengauer T, et al. editors. Proceedings of the 7th international conference on intelligent systems for molecular biology. p. 77–86. AAAI Press; 1999. [PubMed]
- 17.Bunescu R, Mooney R. Learning to extract relations from the web using minimal supervision. In: Zaenen A, van den Bosch A editors. Proceedings of the 45th annual meeting of the association of computational linguistics. p. 576–83. Association for Computational Linguistics, Prague; 2007. https://aclanthology.org/P07-1073.
- 18.Riedel S, Yao L, McCallum A. Modeling relations and their mentions without labeled text. In: Balcázar JL, Bonchi F, Gionis A, Sebag M editors. Machine learning and knowledge discovery in databases: European conference, ECML PKDD 2010, Proceedings, Part III 21. New York: Springer; 2010. p. 148–163.
- 19.Hoffmann R, Zhang C, Ling X, Zettlemoyer L, Weld DS. Knowledge-based weak supervision for information extraction of overlapping relations. In: Lin D, Matsumoto Y, Mihalcea R editors. Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies. p. 541–550. Association for Computational Linguistics, Portland; 2011. https://aclanthology.org/P11-1055.
- 20.Zeng D, Liu K, Chen Y, Zhao J. Distant supervision for relation extraction via piecewise convolutional neural networks. In: Màrquez L, Callison-Burch C, Su J editors. Proceedings of the 2015 conference on empirical methods in natural language processing. p. 1753–62. Association for Computational Linguistics, Lisbon; 2015. https://aclanthology.org/D15-1203.
- 21.Lin Y, Shen S, Liu Z, Luan H, Sun M. Neural relation extraction with selective attention over instances. In: Erk K, Smith NA editors. Proceedings of the 54th annual meeting of the association for computational linguistics (Vol. 1 Long Papers). p. 430–9. Association for Computational Linguistics, Vancouver; 2017.
- 22.Luo B, et al. Learning with noise: enhance distantly supervised relation extraction with dynamic transition matrix. In: Barzilay R, Kan MY editors. Proceedings of the 55th annual meeting of the association for computational linguistics (Vol. 1: Long Papers), p. 430–9. Association for Computational Linguistics, Vancouver; 2017. https://aclanthology.org/P17-1040.
- 23.Han X, Liu Z, Sun M. Neural knowledge acquisition via mutual attention between knowledge graph and text. In: McIlraith SA, Weinberger KQ editors. Proceedings of the 32nd AAAI conference on artificial intelligence. vol. 32. 2018.
- 24.Alt C, Hübner M, Hennig L. Fine-tuning pre-trained transformer language models to distantly supervised relation extraction. In: Korhonen A, Traum D, Màrquez L editors. Proceedings of the 57th annual meeting of the association for computational linguistics. p. 1388–98. Association for Computational Linguistics, Florence; 2019. https://aclanthology.org/P19-1134.
- 25.Dai Q, Inoue N, Reisert P, Takahashi R, Inui K. Distantly supervised biomedical knowledge acquisition via knowledge graph based attention. In: Nastase V, Roth B, Dietz L, McCallum A editors. Proceedings of the workshop on extracting structured knowledge from scientific publications. p. 1–10. Association for Computational Linguistics, Minneapolis; 2019. https://aclanthology.org/W19-2601.
- 26.Amin S, Dunfield KA, Vechkaeva A, Neumann G. A data-driven approach for noise reduction in distantly supervised biomedical relation extraction. In: Demner-Fushman D, Cohen KB, Ananiadou S, Tsujii J editors. Proceedings of the 19th SIGBioMed workshop on biomedical language processing. p. 187–94. Association for Computational Linguistics; 2020. https://aclanthology.org/2020.bionlp-1.20.
- 27.Lee J, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36:1234–40. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Wu S, He Y. Enriching pre-trained language model with entity information for relation classification. In: Zhu W et al. editors. Proceedings of the 28th ACM international conference on information and knowledge management, CIKM ’19. p. 2361–64. Association for Computing Machinery, New York; 2019. 10.1145/3357384.3358119. [DOI]
- 29.Bodenreider O. The unified medical language system (umls): integrating biomedical terminology. Nucleic Acids Res. 2004;32:D267–70. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Hogan WP et al. Abstractified multi-instance learning (AMIL) for biomedical relation extraction. 2021.
- 31.Beltagy I, Lo K, Cohan A. SciBERT: a pretrained language model for scientific text. In: Inui K, Jiang J, Ng V, Wan X editors. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). p. 3615–3620. Association for Computational Linguistics, Hong Kong; 2019. https://aclanthology.org/D19-1371.
- 32.Asada M, Miwa M, Sasaki Y. Using drug descriptions and molecular structures for drug-drug interaction extraction from literature. Bioinformatics. 2021;37:1739–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Yasunaga M, Leskovec J, Liang P. LinkBERT: Pretraining language models with document links. In: Muresan S, Nakov P, Villavicencio A editors. Proceedings of the 60th annual meeting of the association for computational linguistics. Vol. 1: Long Papers. p. 8003–16. Association for Computational Linguistics, Dublin; 2022. https://aclanthology.org/2022.acl-long.551.
- 34.Wadhwa S, Amir S, Wallace B. Revisiting relation extraction in the era of large language models. In: Rogers A, Boyd-Graber J, Okazaki N editors. Proceedings of the 61st annual meeting of the association for computational linguistics. Vol. 1: Long Papers. p. 15566–15589. Association for Computational Linguistics, Toronto; 2023. https://aclanthology.org/2023.acl-long.868. [DOI] [PMC free article] [PubMed]
- 35.Brown TB. Language models are few-shot learners. 2020.
- 36.Chung HW, et al. Scaling instruction-finetuned language models. J Mach Learn Res. 2024;25:1–53. [Google Scholar]
- 37.Gao C, et al. Few-shot relational triple extraction with hierarchical prototype optimization. Pattern Recogn. 2024;156: 110779. [Google Scholar]
- 38.Khosla P, et al. Supervised contrastive learning. Adv Neural Inf Process Syst. 2020;33:18661–73. [Google Scholar]
- 39.Theodoropoulos C, Henderson J, Coman AC, Moens MF. Imposing relation structure in language-model embeddings using contrastive learning. In: Bisazza A, Abend O editors. Proceedings of the 25th conference on computational natural language learning. p. 337–348. Association for Computational Linguistics; 2021. https://aclanthology.org/2021.conll-1.27.
- 40.Liu X, et al. Self-supervised learning: generative or contrastive. IEEE Trans Knowl Data Eng. 2021;35:857–76. [Google Scholar]
- 41.Chen Q, et al. An extensive benchmark study on biomedical text generation and mining with ChatGPT. Bioinformatics. 2023;39:btad557. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Asada M, Fukuda K. Enhancing relation extraction from biomedical texts by large language models. In: Degen H, Ntoa S editors. Proceedings of the 5th international conference of artificial intelligence in human-computer interaction, held as part of the 26th human-computer interaction international conference. p. 3–14. New York: Springer; 2024.
- 43.Zhang J, et al. A study of biomedical relation extraction using GPT models. AMIA Summits Transl Sci Proc. 2024;2024:391. [PMC free article] [PubMed] [Google Scholar]
- 44.Achiam J, et al. Gpt-4 technical report 2023.
- 45.Zhang K, Jimenez Gutierrez B, Su Y, Rogers A. Aligning instruction tasks unlocks large language models as zero-shot relation extractors. In: Rogers A, Boyd-Graber J, Okazaki N editors. Findings of the association for computational linguistics: ACL 2023. p. 794–812. Association for Computational Linguistics, Toronto; 2023. https://aclanthology.org/2023.findings-acl.50.
- 46.Ji Z, et al. Survey of hallucination in natural language generation. ACM Comput Surveys. 2023. 10.1145/3571730. [DOI]
- 47.Thirunavukarasu AJ, et al. Large language models in medicine. Nat Med. 2023;29:1930–40. [DOI] [PubMed] [Google Scholar]
- 48.Ceballos-Arroyo AM, et al. Open (clinical) LLMs are sensitive to instruction phrasings. In: Demner-Fushman D, Ananiadou S, Miwa M, Roberts K, Tsujii J editors. Proceedings of the 23rd workshop on biomedical natural language processing. p. 50–71. Association for Computational Linguistics, Bangkok; 2024. https://aclanthology.org/2024.bionlp-1.5.
- 49.Petroni F, et al. Language models as knowledge bases? In: Inui K, Jiang J, Ng V, Wan X editors. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). p. 2463–2473. Association for Computational Linguistics, Hong Kong; 2019.
- 50.Jurafsky D, Martin JH. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. 1st ed. USA: Prentice Hall PTR; 2000. [Google Scholar]
- 51.Bunescu R, Mooney R. A shortest path dependency kernel for relation extraction. In: Mooney R, Brew C, Chien LF, Kirchhoff K editors. Proceedings of human language technology conference and conference on empirical methods in natural language processing. p. 724–731. Association for Computational Linguistics, Vancouver, British Columbia; 2005. https://aclanthology.org/H05-1091.
- 52.Purpura A, Bonin F, Bettencourt-silva J. Accelerating the discovery of semantic associations from medical literature: mining relations between diseases and symptoms. In: Li Y, Lazaridou A editors. Proceedings of the 2022 conference on empirical methods in natural language processing: industry track. p. 77–89. Association for computational linguistics, Abu Dhabi; 2022. https://aclanthology.org/2022.emnlp-industry.6.
- 53.Toutanvoa K, Manning CD. Enriching the knowledge sources used in a maximum entropy part-of-speech tagger. In: Toutanvoa K, Manning CD editors. 2000 Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora. Association for Computational Linguistics, Hong Kong; 2000.
- 54.Toutanova K, Klein D, Manning CD, Singer Y. Feature-rich part-of-speech tagging with a cyclic dependency network. In: Toutanova K, Klein D, Manning CD, Singer Y editors. Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology - Volume 1, NAACL ’03. p. 173–80. Association for Computational Linguistics; 2003. 10.3115/1073445.1073478. [DOI]
- 55.Dozat T, Manning CD. Deep biaffine attention for neural dependency parsing. In: Dozat T, Manning CD editors. International conference on learning representations. 2017. https://openreview.net/forum?id=Hk95PK9le.
- 56.Neumann M, King D, Beltagy I, Ammar W. ScispaCy: fast and robust models for biomedical natural language processing. In: Demner-Fushman D, Cohen KB, Ananiadou S, Tsujii J editors. Proceedings of the 18th BioNLP workshop and shared task. p. 319–327. Association for Computational Linguistics, Florence; 2019. https://aclanthology.org/W19-5034.
- 57.Akbik A, et al. FLAIR: an easy-to-use framework for state-of-the-art NLP. In: Ammar W, Louis A, Mostafazadeh N editors. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics (Demonstrations). p. 54–9. Association for Computational Linguistics, Minneapolis; 2019. https://aclanthology.org/N19-4010.
- 58.Amini A, Liu T, Cotterell R. Hexatagging: Projective dependency parsing as tagging. In: Rogers A, Boyd-Graber J, Okazaki N editors. Proceedings of the 61st annual meeting of the association for computational linguistics. Vol. 2: Short Papers. p. 1453–64. Association for Computational Linguistics, Toronto; 2023. https://aclanthology.org/2023.acl-short.124.
- 59.Zhou W, Huang K, Ma T, Huang J. Document-level relation extraction with adaptive thresholding and localized context pooling. 2021.
- 60.Zhao X, et al. A comprehensive survey on relation extraction: recent advances and new frontiers. ACM Comput Surv. 2024;56:1–39. [Google Scholar]
- 61.Coman A, Theodoropoulos C, Moens M-F, Henderson J. GADePo: graph-assisted declarative pooling transformers for document-level relation extraction. In: Yu W, et al. editors. Proceedings of the 3rd workshop on knowledge augmented methods for NLP. p. 1–14. Association for Computational Linguistics, Bangkok; 2024. https://aclanthology.org/2024.knowledgenlp-1.1.
- 62.Horn RA, Johnson CR. Matrix analysis. Cambridge University Press; 2012.
- 63.Kullback S, Leibler RA. On information and sufficiency. Ann Math Stat. 1951;22:79–86. [Google Scholar]
- 64.Kullback S. Information theory and statistics. Courier Corporation; 1997.
- 65.Lu N, Niu G, Menon AK, Sugiyama M. On the minimal supervision for training any binary classifier from only unlabeled data. 2019.
- 66.Lu N, Zhang T, Niu G, Sugiyama M. Mitigating overfitting in supervised classification from two unlabeled datasets: a consistent risk correction approach. In: Chiappa S, Calandra R. Proceedings of the 23rd international conference on artificial intelligence and statistics. Vol. 108 of Proceedings of machine learning research. p. 1115–25. PMLR; 2020. https://proceedings.mlr.press/v108/lu20c.html.
- 67.Xu Y, Zhang H, Miller K, Singh A, Dubrawski A. Noise-tolerant interactive learning using pairwise comparisons. In: Guyon I, et al. editors. Advances in neural information processing systems. 30 (Curran Associates, Inc.; 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/e11943a6031a0e6114ae69c257617980-Paper.pdf.
- 68.Xu L, Honda J, Niu G, Sugiyama M. Uncoupled regression from pairwise comparison data. In: Wallach H, et al. editors. Advances in neural information processing systems. vol. 32. Curran Associates, Inc.; 2019. https://proceedings.neurips.cc/paper_files/paper/2019/file/6832a7b24bc06775d02b7406880b93fc-Paper.pdf.
- 69.Cui Z, Charoenphakdee N, Sato I, Sugiyama M. Classification from triplet comparison data. Neural Comput. 2020;32:659–81. [DOI] [PubMed] [Google Scholar]
- 70.Hawkins DM. The problem of overfitting. J Chem Inf Comput Sci. 2004;44:1–12. [DOI] [PubMed] [Google Scholar]
- 71.Nair V, Hinton GE. Rectified linear units improve restricted boltzmann machines. In: Fürnkranz, J, Joachims T editors. Proceedings of the 27th international conference on machine learning. p. 807–814. 2010.
- 72.Natarajan N, Dhillon IS, Ravikumar PK, Tewari A. Learning with noisy labels. In: Burges C, Bottou L, Welling M, Ghahramani Z, Weinberger K editors. Advances in neural information processing systems. vol. 26. Curran Associates, Inc.; 2013. https://proceedings.neurips.cc/paper_files/paper/2013/file/3871bd64012152bfb53fdf04b401193f-Paper.pdf.
- 73.Northcutt CG, Wu T, Chuang IL. Learning with confident examples: rank pruning for robust classification with noisy labels. In: Zhalama JZ, Eberhardt F, Mayer W editors. Proceedings of the 33rd conference on uncertainty in artificial intelligence (UAI). 2017. https://auai.org/uai2017/proceedings/papers/35.pdf.
- 74.Laine S, Aila T. Temporal ensembling for semi-supervised learning. 2017. https://openreview.net/forum?id=BJ6oOfqge.
- 75.Tarvainen A, Valpola H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In: Guyon I, et al. editors. Advances in neural information processing systems. vol. 30. Curran Associates, Inc.; 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/68053af2923e00204c3ca7c6a3150cf7-Paper.pdf.
- 76.Menon A, Van Rooyen B, Ong CS, Williamson B. Learning from corrupted binary labels via class-probability estimation. In: Bach F, Blei D editors. Proceedings of the 32nd international conference on machine learning. p. 125–34. PMLR; 2015.
- 77.Theodoropoulos C, Coman AC, Henderson J, Moens M-F. Enhancing biomedical knowledge discovery for diseases: an open-source framework applied on RETT syndrome and alzheimer’s disease. IEEE Access. 2024;12:180652–73. [Google Scholar]
- 78.Bravo À, Piñero J, Queralt-Rosinach N, Rautschka M, Furlong LI. Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research. BMC Bioinf. 2015;16:1–17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79.Pyysalo S, et al. BioInfer: a corpus for information extraction in the biomedical domain. BMC Bioinf. 2007;8:1–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 80.Petriti U, Dudman DC, Scosyrev E, Lopez-Leon S. Global prevalence of RETT syndrome: systematic review and meta-analysis. Syst Rev. 2023;12:5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Scheltens P, et al. Alzheimer’s disease. Lancet. 2021;397:1577–90. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 82.Trejo-Lopez JA, Yachnis AT, Prokop S. Neuropathology of Alzheimer’s disease. Neurotherapeutics. 2023;19:173–85. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Lu Z. Pubmed and beyond: a survey of web tools for searching biomedical literature. Database. 2011;2011, baq036. [DOI] [PMC free article] [PubMed]
- 84.Gu Y, et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Trans Comput Healthcare (HEALTH). 2021;3:1–23. [Google Scholar]
- 85.Hagberg A, Swart PJ, Schult DA. Exploring network structure, dynamics, and function using networkX. In: Varoquaux G, Vaught T, Millman J editors. Proceedings of the 7th python in science conference. p. 11–5. Pasadena; 2008.
- 86.Tinn R, et al. Fine-tuning large neural language models for biomedical natural language processing. Patterns. 2023;4. [DOI] [PMC free article] [PubMed]
- 87.Wolf T, et al. Transformers: state-of-the-art natural language processing. In: Liu Q, Schlangen D editors. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. p. 38–45. Association for Computational Linguistics, Online; 2020. https://aclanthology.org/2020.emnlp-demos.6.
- 88.Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R. Dropout: a simple way to prevent neural networks from overfitting. J Mach Learn Res. 2014;15:1929–58. [Google Scholar]
- 89.Ioffe S, Szegedy C. Batch normalization: accelerating deep network training by reducing internal covariate shift. In: Bach F, Blei D editors. Proceedings of the 32nd international conference on machine learning. Vol. 37 of Proceedings of machine learning research. p. 448–56. PMLR, Lille; 2015. https://proceedings.mlr.press/v37/ioffe15.html.
- 90.Kingma DP, Ba J. Adam: A method for stochastic optimization; 2014.
- 91.Paszke A, et al. Pytorch: an imperative style, high-performance deep learning library. In: Wallach H et al. Advances in neural information processing systems. vol. 32. Curran Associates, Inc.; 2019. https://proceedings.neurips.cc/paper_files/paper/2019/file/bdbca288fee7f92f2bfa9f7012727740-Paper.pdf.
- 92.Yuan Z, Liu Y, Tan C, Huang S, Huang F. Improving biomedical pretrained language models with knowledge. In: Demner-Fushman D, Cohen KB, Ananiadou S, Tsujii J editors. Proceedings of the 20th workshop on biomedical language processing. p. 180–90. Association for Computational Linguistics, Online; 2021. https://aclanthology.org/2021.bionlp-1.20.
- 93.Park G, McCorkle S, Soto C, Blaby I, Yoo S. Extracting protein-protein interactions (PPIs) from biomedical literature using attention-based relational context information. In: Tsumoto S et al. editors. Proceedings of the 10th IEEE international conference on big data (big data), p. 2052–61. IEEE; 2022.
- 94.Labrak Y, et al. BioMistral: a collection of open-source pretrained large language models for medical domains. In: Ku LW, Martins A, Srikumar V editors. Findings of the association for computational linguistics ACL 2024. p. 5848–64. Association for computational linguistics, Bangkok, Thailand and virtual meeting; 2024. https://aclanthology.org/2024.findings-acl.348.
- 95.Jiang AQ, et al. Mistral 7b 2023.
- 96.Shoemake K. Animating rotation with quaternion curves. ACM SIGGRAPH Comput Graph. 1985;19:245–54. 10.1145/325165.325242. [Google Scholar]
- 97.Yadav P, Tam D, Choshen L, Raffel CA, Bansal M. Ties-merging: resolving interference when merging models. In: Oh A, et al. editors. Advances in neural information processing systems. vol. 36, p. 7093–115. Curran Associates, Inc.; 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/1644c9af28ab7916874f6fd6228a9bcf-Paper-Conference.pdf.
- 98.Yu L, Yu B, Yu H, Huang F, Li Y. Language models are super mario: Absorbing abilities from homologous models as a free lunch. In: Salakhutdinov R, et al. editors. Proceedings of the 41st international conference on machine learning. 2024. https://openreview.net/forum?id=fq0NaiU8Ex.
- 99.Pal A, Umapathi LK, Sankarasubbu M. Med-HALT: Medical domain hallucination test for large language models. In: Jiang J, Reitter D, Deng S editors. Proceedings of the 27th conference on computational natural language learning (CoNLL). p. 314–34. Association for Computational Linguistics, Singapore; 2023. https://aclanthology.org/2023.conll-1.21.
- 100.Huang L, et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans Inf Syst. 2025;43. 10.1145/3703155.
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The GAD dataset [78] used in the current study is available in the BioBERT repository [27], https://github.com/dmis-lab/biobert and in the PPI-Relation-Extraction repository [93], https://github.com/BNLNLP/PPI-Relation-Extraction. The BioInfer dataset [79] used in the current study is available in the PPI-Relation-Extraction repository [93], https://github.com/BNLNLP/PPI-Relation-Extraction and in the Hugging Face dataset repository, https://huggingface.co/datasets/bigbio/bioinfer. The ReDReS and ReDAD datasets [77] are available in the Enhancing-Biomedical-Knowledge-Discovery-for-Diseases repository, https://github.com/christos42/Enhancing-Biomedical-Knowledge-Discovery-for-Diseases.


























































