Skip to main content
EPA Author Manuscripts logoLink to EPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Feb 1.
Published in final edited form as: Comput Toxicol. 2023 Feb;25:1–15. doi: 10.1016/j.comtox.2022.100258

Towards systematic read-across using Generalised Read-Across (GenRA)

Grace Patlewicz 1, Imran Shah 1
PMCID: PMC10483627  NIHMSID: NIHMS1921927  PMID: 37693774

Abstract

Read-across continues to be a popular data gap filling technique within category and analogue approaches. One of the main issues hindering read-across acceptance is the notion of addressing and reducing uncertainties. Frameworks and formats have been created to help facilitate read-across development, evaluation, and residual uncertainties. However, read-across remains an expert-driven approach with each assessment decided on its own merits with no objective means of evaluating performance or quantifying uncertainties. Here, the underlying motivation of creating an algorithmic approach to read-across, namely the Generalised Read-Across (GenRA) approach, is described. The overall objectives of the approach were to quantify performance and uncertainty. Progress made in quantifying the impact of each similarity context commonly relied upon as part of read-across assessment are discussed. The framework underpinning the approach, the software tools developed to date and how GenRA can be used to make and interpret predictions as part of a screening level hazard assessment decision context are illustrated. Future directions and some of the overarching issues still needed in this field and the extent to which GenRA might facilitate those needs are discussed.

Keywords: Read-Across, GenRA, New Approach Methods (NAMs), ToxCast

1.0. Introduction

1.1. Background

Read-across describes the method of filling a data gap whereby a chemical with existing data can be used to make a prediction for a “similar” chemical that is lacking that data [1]. It is an approach that has been in widespread use for over 20 years but since it became called upon to be applied in different regulatory contexts, notably under REACH [2], the level of scrutiny and mechanistic justification significantly increased [3]. The evolution of read-across development and application has been well described in previous publications (readers are referred to [4-7]). That said, some brief recap of read-across definitions is merited here by way of background context. Read-across is a data gap filling technique used within analogue and category approaches. A chemical of interest, the target, has a data gap that needs to be addressed from which one or more source analogues that do have relevant data can be used to make the prediction. The source analogue is typically a substance that is structurally similar but ‘similar’ is placed in quotation marks, since similarity is a relative concept and depending on endpoint to be filled, the mechanistic understanding surrounding that endpoint, etc., a number of other considerations come into play in terms of identifying and evaluating relevant analogues. Structural similarity is usually the starting point for identifying candidate analogues, a so-named ‘unsupervised’ approach but it is by no means the only option. The term “unsupervised” is coined here to emphasise that the structural characteristics used to identify candidate analogues have not been specifically tuned for the endpoint data gap under study, thus they are uninformed. Contrast this with a ‘supervised’ approach where a custom set of features characterising electrophilic reactivity for instance that might be used to identify candidate analogues for a genotoxicity or skin sensitisation endpoint.

Over the years, much technical guidance has been drafted to facilitate the development of read-across [1,8], yet regulatory acceptance is still an issue, not helped by the fact that read-across remains a subjective expert-driven assessment performed on a case-by-case basis. It has been postulated that one specific issue that has thwarted acceptance relates to the uncertainty of the read-across prediction [9,10]. This is not altogether surprising – after all how does one make an objective determination of the validity and robustness of the read-across prediction. Herein lies the issue, indeed many of these same questions surfaced when QSAR models and their predictions were first discussed during the development of technical guidance for REACH [8] but the burden of characterising read-across by the same principles as for QSAR was spared possibly to its later detriment given the challenges that continue to plague read-across acceptance.

Rather than re-frame read-across in the same manner as QSARs despite the fact they share a common foundation, efforts have instead focused on identifying the many different sources of uncertainty in read-across – both from the point of view of the underlying data supporting a read-across (i.e., the toxicity data available for the source analogue(s) themselves) as well as the uncertainty underpinning the similarity rationale forming the basis of the grouping being performed. The motivation was that if the sources of potential uncertainty could be characterised in a structured fashion, then practical strategies to address and reduce those uncertainties could be readily undertaken, i.e., if the rationale was that a target and source chemical were related by metabolic transformation, empirical or in silico evidence could be generated to substantiate that hypothesis.

A number of frameworks were proposed by various researchers to facilitate the identification, and documentation of uncertainties associated with read-across inferences and predictions. An early example of these was the analogue identification and suitability framework published by Wu et al [11]. An extension to the framework by Wu et al [11] was subsequently published by his colleagues in 2014 [12]. This extension outlined considerations around the consistency and quality of the underlying data for the source analogues including the concordance in effects and potency across endpoints. All these considerations were integrated together to assign an overall uncertainty category (low, moderate, high) from which different assessment factors could be applied as appropriate. Complementary to that effort was the strategy for structuring and reporting templates developed by Schultz et al. [13]. This aimed to structure confidence levels associated with elements of a read-across – both from the point of view of an assessment of the similarity of the source analogues themselves as well as an assessment of the mechanistic relevance and completeness of the read-across. In the run up to the 2nd revision of the OECD grouping guidance, work proposed by Cefic LRI and ECETOC led to further considerations of read-across development and acceptance that were subsequently published in Patlewicz et al. [14,15]. A joint workshop with ECHA convened in 2012 discussed some of the ongoing challenges to read-across acceptance. It was also the venue where an early version of ECHA’s Read-across Assessment Framework, the RAAF, was first communicated. The RAAF provides a framework for assessing read-across based on one of 6 scenarios. Each scenario comprises a number of scientific considerations which are scored and where a minimum score is needed for a read-across to be taken up for subsequent decision making.

In any of these aforementioned cases, the scope of the data underlying the read-across or used to substantiate the read-across have been primarily limited to conventional in vivo or in vitro toxicity data. Whilst in the OECD 2014 guidance update [16,17], the notion of using so-called New Approach Methods (NAM) data (where NAM data includes high throughput (HT) and high content (HC) screening data) was starting to evolve, the context was very much aligned with Adverse Outcome Pathways (AOPs), as that was considered at the time to provide the best means of interpreting NAM data in the most appropriate biological context. Indeed, in Patlewicz et al [18], a scientific confidence framework (SCF) for AOPs which was itself an extension of one previously developed for HTS data [19] was intended to provide a means to evaluate and qualify the relevance of NAM data in an appropriate decision context. The key examples developed at the time were focused on estrogenicity in light of the Endocrine Disruption Screening Programme (EDSP) in the US as well as skin sensitisation since this was the first AOP that the OECD developed [20]. In both cases the linear construct of the AOP provided a convenient means by which knowledge associated with the different steps culminating in estrogenicity or skin sensitisation induction could be readily aligned to the key events in the AOP construct summarising the different levels of biological organisation.

In an earlier paper [7], many of these frameworks for read-across were discussed in the context of how they assisted in the development of read-across or its subsequent evaluation. The intent was to provide some means of harmonising the terms and concepts into one hybrid workflow which would illustrate where NAM data could be best leveraged. In addition, scenarios where data driven approaches could be leveraged to move towards quantification of uncertainties were also described. Read-across was also positioned within the broader construct of an Integrated Approaches to Testing and Assessment (IATA) such that depending on purpose, endpoint, number of data gaps, the most efficient data gap filling approach might be applied which need not be a read-across. Examples could include Defined Approaches such as those developed for skin sensitisation [21] or valid QSAR models for specific environmental fate and physicochemical properties.

1.2. The motivations behind Generalised Read-Across (GenRA)

The underlying question that motivated the research to develop GenRA was whether an objective assessment of read-across performance and the associated uncertainty of the predictions could be undertaken. The initial effort [22] focused on establishing a baseline in performance, exploiting the approaches already applied in the QSAR domain and more routinely in the machine learning field to provide some basis for comparison. The approach was coined Generalised Read-Across, GenRA, which used a simple framework of deriving a similarity weighted activity from source analogues (nearest neighbours). For a given target substance, what source analogues could be identified on the basis of (structural) similarity and could a weighted average of their toxicity be used in making a many-to-one read-across prediction. The weighted average would be dependent on the pairwise similarities (specifically a Jaccard (Tanimoto) similarity index). There was a recognition that this baseline relying on structural characteristics alone would be limited in predicting in vivo endpoints. However, by establishing an initial baseline, the impact of other similarity rationales in improving performance could be evaluated in an objective and transparent manner. The GenRA workflow was aligned with established steps in a generalised read-across workflow yet would be sufficiently flexible to incorporate future refinements. The initial workflow relied on analogues identified by chemical structure to make similarity weighted predictions of toxicity using binary representations of in vivo toxicity. The uncertainty assessment relied upon a y-randomisation process whereby the toxicity outcomes were permuted relative to the chemical identifiers in order to derive a distribution of performance scores (using the Area Under the Receiver Operator Characteristic (ROC) curve) from which a p-value could be computed to demonstrate how likely that performance might be arrived at by chance.

In parallel with the approach being developed and mindful of the frameworks (as described in existing technical guidance documents), a review of publicly available software tools to facilitate read-across [23] was undertaken. The goal in implementing GenRA as a tool was to offer the scientific community novel functionality that was not already captured in existing tools. GenRA was formulated as a web-based application closely aligned to a general read-across workflow described in Patlewicz et al (2017). The GenRA web application was first released in 2018 as an integrated workflow that could be launched from within the EPA CompTox Chemicals Dashboard [24] for a specific target substance. The decision context was that an end-user would be approaching read-across from the perspective of a specific discrete substance in a bottom-up approach (see Fig. 1) and after exhausting options within the Dashboard or from its related links to locate relevant toxicity information, source analogues with associated toxicity data could be used to make a read-across prediction. The source analogues would be identified using one of several fingerprint options – these would either be representations of chemical structure using conveniently generated chemical fingerprints from the open-source python cheminformatics package, RDKit (rdkit.org, Landrum) or a fixed set of ToxPrints derived using the public Chemotyper tool ([25]; chemotyper.org). In addition to chemical fingerprints, biologically similar chemicals based on their tested profile in ToxCast or Tox21 assays could also be used to identify source analogues. These candidate analogues would be returned in order of decreasing pairwise similarities relative to the target substance and filtered by availability of in vivo toxicity where the source of toxicity data was reliant on the curated Toxicity Reference Database (ToxRefDB) v1 [26] aggregated by study type (one of 10 study types) and toxicity effect. The GenRA workflow was embedded into a series of interactive panels to offer the end-user limited options whether it be changing the number of analogues returned, the fingerprint type or deselecting analogues from being used in the read-across prediction itself. The first implementation was described in brief in Helman et al [27].

Figure 1.

Figure 1.

Bottom-up or Top-down approaches to category development

1.3. Current and ongoing research

Whilst this first implementation mirrored the baseline approach [22] largely by incorporating one representation of in vivo toxicity as well as chemical and biological fingerprints, research continued to investigate the role and impact of other similarity contexts that are considered important in evaluating analogue relevance and suitability. The main similarity rationales critical to read-across are discussed in Patlewicz et al [6,15] and may be categorised as ‘general’ or ‘specific’. General considerations capture physicochemical similarity, metabolic similarity, reactivity similarity and mechanistic similarity. These considerations are general in that they frame aspects that are not necessarily pertinent to the endpoint under evaluation for which a read-across is being considered. For these “general” considerations, several systematic analyses to evaluate the impact of these in predicting in vivo toxicity have been undertaken. The motivation was firstly to be able to quantify the contribution that these similarity aspects play in read-across performance and secondly to understand to what extent such insights might be generalisable. Are specific insights only applicable to specific chemistries or toxicities or are they more broadly applicable to larger more diverse chemical sets or toxicity outcomes? This could be potentially helpful in refining guidance for developing and applying read-across itself.

The general considerations investigated to date have included the following: physicochemical similarity [28], metabolic similarity [29], reactivity similarity [30] and transcriptomics similarity [31]. The depth within these analyses have varied. In Helman et al [28], the extent to which physicochemical similarity was important in read-across using specific physicochemical parameters as surrogates for modelling bioavailability was investigated. Identifying analogues based on physicochemical parameters and structural parameters at the same time (termed a search expansion) resulted in better read-across performance than merely filtering structural analogues on the basis of their physicochemical features. The relative contributions of structural and physicochemical characteristics could be adjusted to modify the candidate source analogues identified. The search expansion approach performed at least as well as a baseline GenRA using only chemical fingerprints but showed up to a 9% improvement in read-across performance for at least 10 of the 50 organs considered.

Nelms et al., [30] focused on foundational work to better understand selected publicly available structural alert schemes for reactivity that could play a role in identifying or refining neighbourhoods identified by structural similarity. The intention remains to leverage this understanding to implement a means of addressing similarity in reactivity profile.

In Tate et al., [31], a systematic analysis was performed to compare and contrast targeted transcriptomics data generated in HepaRG cells relative to structural fingerprints or a combination of both (hybrid) in predicting binary in vivo toxicity outcomes. There were only modest improvements in performance (based on AUC score) using hybrid descriptors across all endpoints aside for liver specific endpoints where AUC performance did improve by up to 17% over chemical structural descriptors alone.

Boyce et al. [29] described a proof-of-concept study to understand the scope and performance of different liver metabolic prediction systems against a limited set of chemicals to inform subsequent analysis. An initial custom fingerprint (FP) was also proposed that could more generically capture the transformations relative to a parent substance rather than relying on a named transformation pathway provided by the prediction software itself. A case study complemented the performance assessment to illustrate how metabolic transformations from one tool represented as a fingerprint could be used to compare 2 or more analogues.

In addition to investigating ways in which similarity contexts could be codified objectively, a transition from predicting in vivo toxicity as a classification problem to a regression problem using potency was made. Two analyses were undertaken to compare and contrast the use of GenRA in predicting Lowest Observed Adverse Effect Level (LOAEL) values from repeated dose studies from ToxRefDB v2 [32] and the 50% Lethal Dose (LD50) from acute oral rodent studies [33]. The aim was to evaluate the extent to which the similarity-weighted approach was applicable in a regression scenario and compare and contrast the approach to other machine learning techniques. A follow-up manuscript is currently in preparation (Tate et al., in prep) which addresses this question more fully in addition to exploring the imbalance in datasets from which read-across predictions are made.

Despite this concerted progress over the last 6 years, there are still many remaining questions to investigate. Questions such as how to integrate the different similarity contexts and provide tentative recommendations on their relative contributions depending on endpoint of interest and type of chemical structural features, for example, should the contribution from metabolism similarity be larger relative to reactivity similarity or structural similarity itself? To what extent would different representations or even targeted representations result in improved read-across performance (e.g. use of an overall HTS fingerprint or a focused HTS fingerprint) as well as how algorithmic read-across predictions compare with previous read-across examples that have been performed by expert judgement. An area of current focus is to develop a compendium of expert-driven read-across examples. Preliminary work in this area was presented at QSAR 2021 [34].

In addition to considering similarity contexts, another aspect of interest is how to transition from binary representation of chemical or biological information towards other representations for source substances. For instance, how might representations be combined into custom fingerprints? How to consider the role of potency in bioactivity or chemical representations of substances? Or how to incorporate other continuous chemical descriptors beyond fingerprints? Are there ways in which fingerprints or even descriptors can be superseded by graph convolutional network representations or borrow from approaches that are used routinely in natural language processing? Currently the bioactivity data represented within GenRA has been limited to ToxCast and Tox21 hitcall data but our Centre, CCTE has invested in generating data using other HT technologies that go beyond HTS assays, including high throughput phenotypic profiling (HTPP) [35] and high throughput transcriptomics (HTTR) [36]. These two technologies will be useful in providing a broader profile of general bioactivity or toxicity activity, in contrast to some of the ToxCast/Tox21 assays which are more specific for certain effects e.g. Bioseek assays to provide an indication of immunotoxicity activity or ACEA assays to provide indications of estrogenic effects. Indeed, a proof of concept fingerprint representation of phenotypic profiling data has been developed which will be disseminated as part of a future release. Other opportunities could consider the role of AOP or AOP-like information to better link assays or data streams to known or purported pathways that are mechanistically driven though this also relates to the specific endpoint considerations discussed earlier.

2. Software implementations: genra-py

The first release of the GenRA web application permitted a read-across to be conducted on a specific chemical as part of a generalised read-across workflow. In contrast, an alternative and programmatic batch means of using GenRA is available through genra-py. This is a standalone python library that enables user-specific datasets to be analysed – see [37] and the associated code repository available at https://github.com/i-shah/genra-py. The standalone genra-py is not supplied with any databases, thus source analogues are identified from within the end user’s dataset of chemicals. The genra-py package is targeted towards data scientists with familiarity of using the python programming language. The genra-py implementation conforms to the scikit-learn [38] machine learning library’s estimator design pattern, which allows a data scientist to compare the GenRA approach relative to other machine learning algorithms when building prediction models for a specific toxicity outcome. Scikit-learn is the most commonly used python package for machine learning. A repository illustrating how to use the python package without the need for any installation using the acute toxicity example outlined in Helman et al., [33] is available at https://github.com/patlewig/UNC_Rax. The genra-py package is available as an easy installation from the PyPi repository. A Docker image which has genra-py installed within a Jupyter Lab stack is also available from Dockerhub – see https://hub.docker.com/r/patlewig/genra-py.

2.1. Software applications: GenRA

Since the first release of the GenRA web application, underlying databases including the DSSTox inventory [39] have been updated, extended, and refined further. To that end, efforts over the last 2 years have focused on updating these data streams so that they are consistent with what has been publicly released. In addition, the GenRA web application was decoupled from the main CompTox Chemicals Dashboard to facilitate continued development. The intent was to ensure that GenRA’s underlying data streams were updated in a scheduled manner.

Version 2 of the GenRA web application was released in November 2021. Whilst the same workflow was retained, the web application was rebuilt, and the underlying databases supporting the tool were updated. Additional chemical fingerprints were computed to factor in the new substances that had been registered in the DSSTox database since 2018. The ability to search for analogues without any prefiltering on the basis of ToxRefDB data was also added. The latter was added as a feature to provide end-users with a perspective of what source analogues might be more similar but not necessarily associated with ToxRefDB data. The main entry point to access GenRA is now from the comptox.epa.gov portal using the GenRA tile (as shown in Fig. 2). Secondary entry points include using the predictions menu within the main Dashboard application or from within the landing page of a specific substance within the Dashboard as was available in the initial release. The direct url link is https://comptox.epa.gov/genra/.

Figure 2.

Figure 2.

Main entry point for GenRA

3. Using the GenRA workflow: Version 2

The main workflow is described in brief herein. An end-user clicks on the GenRA tile within the portal to launch the application which appears as an empty workflow interface. A search box is available in the top right-hand side of the screen whereby a search query can be made to retrieve a target substance by using chemical name, DTXSID, CAS or other identifier to initialise the workflow. If there is a match within the underlying DSSTox database, a radial plot is returned as a first panel (Panel 1) with the target substance appearing in the centre of the plot with the 10 most similar source substances (clockwise in decreasing order of similarity) on the basis of Morgan chemical fingerprints pre-filtered by availability of ToxRefDB v2 data [40]. The defaults of returning the top 10 similar analogues on the basis of Morgan fingerprints with associated in vivo toxicity data is purely arbitrary. The number of analogues returned can be modified to return only 1 analogue or up to 20 analogues. The fingerprint representation can also be changed from Morgan fingerprints to another chemical or bioactivity fingerprint. Once an end user is satisfied with the selection, Next is clicked which in turn launches 2 further panels in the workflow (Panels 2 and 3) which summarise the availability of information from different data streams for the target substances and the candidate source analogues as a fingerprint count. Panel 2 summarises the representations overall – number of data records per substances, with a different colour gradient dependent on each data stream. The number reflected within each cell in this panel reflects the granularity and breadth of information for the set of substances. Large numbers for the chemical fingerprints reveal how specific and granular the representations of chemistry are for the substances, here the number reflects the number of bits present in a chemical structure e.g. the total bit length for the ToxPrint fingerprint is 729 but for a substance such as Bisphenol A, only 10 bits are present. The higher number for ToxCast and ToxRefDB data highlight the extent to which substances have been tested, e.g. for Bisphenol A, the number of tox_txrf records is 97 which reflects the total number of study-toxicity effects combinations measured or the 323 in the bio_txct column for all the vendor-assay combinations tested. In contrast, Panel 3 represents the toxicity information by default in an expanded grid view in terms of which types of toxicity effect information are available across the analogues. Upon exploration of the data landscape, a generate data matrix button can be clicked to launch a data matrix in Panel 4. Panel 4 summarises the binary in vivo toxicity data in terms of positive and negative responses across the analogues relative to the target. This view allows an evaluation of the concordance and consistency of the analogues between each other and across different toxicity endpoints. The substances are ordered by decreasing similarity. Substances can be deselected from consideration based on expert judgement when considering the toxicity data across analogues and toxicity effects. This flexibility is particularly important since the ways in which candidate analogues are identified within GenRA remains limited to chemical and bioactivity similarity alone. Toxicity outcomes can be filtered to limit predictions to specific toxicity effects or study types and different criteria can be incorporated into the prediction to mandate a specific number of positive/negative outcomes that will inform the uncertainty and performance assessment afterwards. If data are available 9e.g. from ToxRefDB) for the target substance which is the first column in the data matrix then this is also presented. The run read-across button can then be clicked to start the GenRA prediction using the GenraPred engine by default. (Note: GenraPred is the original read-across prediction engine that was first released in 2018. Since it was the only engine, there was no drop down menu visible to the end user to denote a name. GenraPred simply references the function call used in code repository). The first column for the target is updated with prediction information, coloured by the same red and blue but with a change in transparency based on the p-value which provides some measure of confidence for the AUC performance metric. Results can be downloaded as a flat file or as a excel workbook for additional review and analysis. The excel workbook format provides additional meta data that specifies parameters chosen such as the type of fingerprint, the number of analogues specified, the number of positive and negative outcomes, whether the read-across prediction was performed or whether the initial Panel 4 data matrix was simply selected for download. Targeted help is available for each Panel by clicking on the icon button. This provides a brief synopsis of the goal and functionality within each panel. The About link at the top of the application includes a short user manual, link to the original Shah et al., (2016) publication as well as a contact email in case of specific questions.

Figure 3 showcases the overall set of Panels captured in the main GenRA tool using Bisphenol A as an example chemical.

Figure 3.

Figure 3.

Overview of the main GenRA tool when all panels have been launched.

3.1. Using the GenRA workflow: Version 3

A Version 3 of the GenRA web application was then released in February 2022. The user interface was rebuilt using a different technology that provided more out of the box interactivity rather than having to hard code features, e.g., end-user ability to sort columns and deselect substance columns for the different panels within the workflow. This technology was consistent with the look and feel of the existing CompTox Chemicals Dashboard.

In addition to the standard chemical and bioactivity fingerprints, an additional option was provided – a custom fingerprint. This would provide the user with the ability to create their own custom hybrid fingerprint using combinations of the existing fingerprints. An end-user would be able to identify analogues based on up to 3 different fingerprint sets and adjust the weightings in terms of their relative contributions. Selection of weights is up to the end-user to decide, though a future release of GenRA may provide some guiding principles. In Fig. 4, the interface for creating a custom fingerprint is shown. The slider allows the relative contributions for each fingerprint to be adjusted, in this case resulting in a 25% contribution from Morgan and ToxPrint fingerprints and the remaining 50% from the ToxCast (bioactivity) fingerprint.

Figure 4.

Figure 4.

Illustration for the functionality to demonstrate creation of a custom fingerprint.

Another additional feature enabled the end-user to draw substances that were not already registered within the Dashboard so that predictions using chemical fingerprints could be made in real-time using a Ketcher drawing palette (Fig. 5). An end-user can draw a structure using the drawing palette or alternatively import a mol file or copy/paste a SMILES representation to introduce a new substance into the GenRA workflow. Predictions generated are not stored and any substances drawn are not automatically incorporated into underlying databases supporting GenRA. End-users can request substances to be registered into the DSSTox database [39] with a support ticket should they wish them to appear as potential candidate analogues in future GenRA releases.

Figure 5.

Figure 5.

Using the Ketcher drawing palette to submit a chemical to GenRA

3.2. New features and functionalities within GenRA Version 3.1

Version 3.1 was released September 2022 with several significant new features:

1) ability to download the radial plot view and top 100 most similar analogues.

2) ability to explore physicochemical similarity by exploring the distribution of specific properties across analogues.

3) a network tool to enable an exploration of neighbourhoods of source analogues and compare them across different FP types.

4) prediction of in vivo toxicity potencies rather than binary toxicity predictions initially reliant on dose values aggregated in ToxRefDB (ToxValDB will be incorporated in a future release) making use of the genra-py library [37].

5) prediction of ToxCast hitcall outcomes – on a per assay basis.

3.2.1. Additional download options

In addition to the download options as part of Panel 4, two new download options are now available within Panel 1. The Top 100 most similar analogues can be downloaded as a CSV file. Downloading such a list could then be used in a batch search within the Dashboard to access additional information. The other option is the ability to download the radial plot as an image to capture the structural representations of the source analogues.

3.2.2. Physicochemical property explorer

As part of Panel 1 and the radial plot view, a toggle button named Physchem Data is now visible. This physicochemical data explorer presents the distribution of different predicted properties as a series of boxplots overlaid with swarm plots. The physicochemical data streams are molecular weight, melting point, boiling point, LogP (octanol-water) partition coefficient, vapour pressure, water solubility and Henry’s Law constant. The values that are plotted are stored predictions that have been generated using the OPERA suite of models [41]. Whilst the physicochemical properties are plotted within Panel 1 of the application; the estimated values are also tabulated as part of the data matrix view in Panel 4 (Fig. 6 shows an example of the distributions). The plot is interactive so that hovering over any of the data points will show the exact value for a given property and the associated chemical name.

Figure 6.

Figure 6.

Boxplot and Swarm plots for chemical structural analogues for target Fluconazole.

3.2.3. Neighbourhood exploration

The network exploration graph tool is a beta exploration tool that is accessible by clicking on the button in Panel 1. This results in a pop-up window which can be resized for ease of viewing. A help button is available to provide some initial instruction of how to navigate the graph representation presented. Different fingerprint representations can be showcased simultaneously on the graph. The graph depicts the first 3 source analogues (shown as squares) for a target substance (shown in red) but in addition the subsequent neighbours of those source analogues (next generation analogues) are shown. The intent of the exploration tool is twofold: 1) to show the immediate landscape surrounding the target substance to provide some context of whether the target resides in a sparse or dense part of the chemical landscape, and 2) to show the overlap and commonality of the top 3 analogues and their next generation analogues using different FP representations, i.e., are the same analogues identified by different FP representations. These two factors are provided to end users for additional context with respect to the initial analogues identified. The graph returned is prefiltered on the basis of availability of toxicity data and/or bioactivity data as appropriate. The depiction shown can be zoomed in or out and the specific FP selections made can be exported as a json download for subsequent data analysis in other 3rd party tools. By default, two fingerprint representations are returned – Morgan FP (red) and ToxCast (green) FP. Other FPs can be selected using a slider button and clicking on the the green Update button to refresh the graph shown. The network (graph) is effectively a collection of nodes (target and source analogues) whereas the connections or edges between them are represented by the pairwise similarities (weighted by the actual pairwise similarity). Higher pairwise similarities correspond with thicker edges. Nodes (chemicals) can be moved on the panel interactively as the graph is explored interactively. A reset button will return the view to the last original state in an effort to ‘clean’ the depiction. Clicking on any node will centre the graph and present the user with additional information for that specific substance including its chemical identifiers as well as its structural depiction. Hovering over the source analogues in the graph will show the DTXSID identifier and chemical name. Clicking on a substance within the graph will also provide the option to start the GenRA workflow with that substance by clicking on the green GenRA button. The short case study in Section 4.0 provides some screenshots to illustrate how to interpret the graph.

3.2.4. Potency predictions using ToxRefDB

As another proof-of-concept functionality, the potency values (minimum doses) that are presented for ToxRefDB can be used as inputs to a read-across prediction. In this case, the workflow is slightly modified in that only positive predictions are made on the dose information available. Panels 3 and 4 reflect the main changes and functionality. In Panel 3, a different toxicity fingerprint representation (Tox fingerprint Dosage) is selected under the By dropdown filter to update the panel. Clicking on the generate data matrix button will return a different data matrix representation (Fig. 7). Instead of the red and blue colour coded cells, the matrix is filled with a coloured bar scaled by the range of potencies observed across all analogues for a given study type. The actual dose for the source substance or target is depicted by a circle. Predictions can then be made to update the target column using the genra-py engine. The prediction returned (depicted as a triangle) is presented in both −log molar and mg/kg-bw/day units.

Figure 7:

Figure 7:

Panel 4 Data matrix with potency information depicted.

3.2.5. Hitcall predictions for ToxCast assays

ToxCast hitcall outcomes can be predicted based on different chemical fingerprint representations. In this case, similar analogues are identified in the same way but are pre-filtered on the basis of the availability of ToxCast data in Panel 1. This automatically updates the matrix shown in Panel 3, such that analogues can be summarised by the hitcall information from ToxCast. The Group By dropdowns reflect ToxCast and ToxCast fingerprint rather than ToxRef or ToxRef fingerprint. When Panel 4 is launched, GenRA predictions can be made for ToxCast hitcall outcomes on an per assay basis. Note assays in ToxCast are performed on a sample level, such that the data matrix depiction reflects an aggregated maximum value across samples to enable per substance outcome. Hit call outcomes have not been prefiltered based on cytotoxicity information or analytical information. Since the prediction is binary – a hitcall being active or inactive, the corresponding data matrix in Panel 4 mirrors the format used for ToxRefDB binary toxicity predictions.

4. Bringing it together with a case study example: Bisphenol A

Although Bisphenol A [DTXSID7020182] is a data rich substance, it presents a convenient example to illustrate some of the new functionalities when applying GenRA.

Michalowicz [42] presented a mini review to summarise the main concerns surrounding Bisphenol A in terms of its toxicity profile. Here, GenRA is used to perform an exploratory evaluation on Bisphenol A to inform analogue identification and selection.

Bisphenol A was searched for within the GenRA application and the default radial plot of the top 10 analogues on the basis of Morgan fingerprints were returned (Fig 8).

Figure 8.

Figure 8.

Radial plot for Bisphenol A on the basis of Morgan Fingerprints.

The most similar analogue to Bisphenol A with associated ToxRefDB v2 data was found to be DTXSID8021771 4-(2-methylbutan-2-yl)phenol which had a pairwise similarity of 0.48. Clicking on the Physchem Data button launches the pop-up window which reveals the interactive distributions plot described in Section 3.22 (Fig. 9).

Figure 9.

Figure 9.

Physical properties distribution plots

On the plot, Bisphenol A [DTXSID7020182] is observed to have a predicted boiling point (Bpt) of 343 deg C whereas phenol [DTXSID5021124] which is shown at the tail of the distribution has a predicted Bpt of 182 deg C. For predicted LogKow, 3,3’,5,5’-Tetrabromobisphenol A [DTXSID1026081] is at one end of the spectrum with a predicted value of 6.66 vs 3.32 for Bisphenol A. Based on the analogues returned with available toxicity data, Phenol and 3,3’,5,5’-Tetrabromobisphenol A might be deselected from consideration as source analogues since their predicted physicochemical characteristics appear quite different than that of Bisphenol A. The target chemical is depicted by a diamond, whereas the source analogues are colour coded and ordered by decreasing similarity as captured in the associated legend. Phenol is one of the least similar analogues in the set returned. It has a pairwise similarity of 0.26 although the range is limited given the 3,3’,5,5’-Tetrabromobisphenol A, the 5th most similar analogue only has a pairwise similarity of 0.29. Both of these similarities are quite low reflecting the lack of very structurally similar analogues with in vivo data within ToxRefDB.

Alternatively, a different FP representation could be selected to identify and evaluate the source analogues returned. In Fig. 10, ToxPrints were used to identify candidate analogues. The pairwise similarities were generally higher, the most similar analogue had a pairwise similarity of 0.73. The set of analogues returned shows some overlap, but the order of substances differs.

Figure 10.

Figure 10.

Radial plot of Bisphenol A using ToxPrints

On the basis of these analogues, the corresponding physicochemical distributions also differ (Fig. 11). 3,3’,5,5’-Tetrabromobisphenol A [DTXSID1026081] still appears as an outlier but 4-(2-methylbutan-2-yl)phenol [DTXSID8021771], clorophene [DTXSID5020154] and tert-butylhydroquinone [DTXSID6020220] look most similar in terms of their properties.

Figure 11.

Figure 11.

Physchem data for the analogues returned on the basis of ToxPrints

Overall, the physchem distribution provides an initial view on how comparable the properties are across the analogues and whether a balance of greater or fewer analogues might be merited to carry forward into the remaining panels.

The neighbourhood exploration tool provides a complementary means to explore and evaluate the analogues both in terms of chemical landscape density but also to navigate commonalities in analogues returned from different FPs overlaid on the same plot. Fig. 12 shows a portion of the network on the basis of ToxPrints (in blue) and Morgan FPs (in red) filtered on the basis of availability in ToxRefDB data to help showcase the overlap in the top 3 analogues and their next generation analogues. One of the source analogues, 4-(1,1,3,3-Tetramethylbutyl)phenol [DTXSID9022360] is in common between these 2 FP types as denoted by the loop in the graph. The first analogue [DTXSID8021771] on the basis of Morgan FP shares a link with the second analogue on the basis of ToxPrints. Beyond this commonality across the first two analogues, the neighbourhoods from these 2 FPs appear otherwise quite distinct.

Figure 12.

Figure 12.

ToxPrint and Morgan top 3 analogues and their associated neighbours. FP types are colour coded blue for ToxPrints and red for Morgan fingerprints.

Overlaying what analogues might be returned on the basis of ToxCast FPs (in green), as shown in Fig. 13, finds one substance to be in common between ToxCast FPs and Morgan FPs; and a different substance is shared between ToxCast FPs and ToxPrints. DTXSID3020465 is the second closest analogue with respect to ToxCast FPs and the fourth closest on the basis of Morgan FPs whereas DTXSID5020154 is shared between ToxCast FPs (third closest) and ToxPrints (seventh closest).

Figure 13.

Figure 13.

Network graph view updated with ToxCast FP based analogues (green)

Analogue selection is challenging though the network view permits an initial exploration of commonality between the closest analogues, and how their similarity metrics vary (based on the edge thickness). The analogues that are returned on the basis of ToxCast FPs are structurally dissimilar from Bisphenol A in terms of their scaffold but are enriched by substances that exhibit estrogenic activity (assuming the ToxCast summary data is either probed using the Dashboard or the ToxCast assay level Panel 4 data matrix is explored in more detail). In light of this, exploring the concordance of effects within reproductive or developmental toxicity study types may be helpful in prioritising analogue selection for read-across predictions for Bisphenol A. This is consistent in terms of what has been reported as far as the toxicity profile of Bisphenol A.

Using the network view and the apparent interconnectivity between analogues on the basis of different contexts, identifying analogues on the basis of a custom FP using both structural representations in conjunction with a ToxCast FP representation may give rise to a more representative set of analogues.

A custom hybrid fingerprint of 33% ToxCast, 33% ToxPrints and 33% Morgan FP was used to identify source analogues (Fig. 14).

Figure 14.

Figure 14.

Radial plot for analogues returned based on a hybrid fingerprint

From here, the data summary from Panels 2 and 3 (Fig. 15) can be explored in more detail.

Fig. 15.

Fig. 15.

Panels 2 and 3 for Bisphenol and its analogues.

Hovering over the column headers in Panel 2 will return a description of how that data stream was generated. These 2 panels provide a perspective of data sparsity across the source analogues relative to the target before Panel 4’s data matrix is generated. Using data matrix within Panel 4, the physicochemical characteristics can be explored further to make a determination on whether certain analogues should be deselected from further consideration using the tick option next to the pairwise similarity metric shown (Fig. 16). Ideally this would be undertaken in concert with an evaluation of the concordance of the toxicity outcomes across the study types that are more relevant for the assessment, i.e., developmental (DEV) and reproductive (REP) study types.

Fig. 16.

Fig. 16.

Deselection of analogues on the basis of physicochemical characteristics.

Upon deselection of several analogues (in this case, physchem data was used to de-select 2 analogues), the remaining analogues may be carried forward into the GenRA run read-across calculation itself using the default GenraPred engine. Clicking on Run Read-across will start the prediction such that the first column is updated with predictions (Fig. 17).

Fig. 17.

Fig. 17.

Generating predictions for Bisphenol A using analogues identified by a custom FP and deselecting those analogues that have considerably different physicochemical property profiles.

To ensure more meaningful AUC calculations, it is preferable to adjust the minimum and maximum positives/negatives to at least 2 before running the read-across calculation. In the hover-over, a similarity weighted activity, denoted ACT reflects the actual prediction. A designation of False Positive (FP), True Positive (TP) etc. reflects the agreement of the prediction with the empirical data that might have been available for Bisphenol A to start with. AUC and p-value provides some performance metrics for the read-across prediction. Ideally a high AUC (close to 1) and low p-value (less than 0.05) would be consistent with a more robust read-across prediction though it should be noted that often times even with a custom FP, predictions are not likely to be robust due to the challenges of predicting complex toxicity endpoints. Manipulating the download file with predictions would allow some sorting and ranking of predictions to be performed.

5. Conclusions

Read-across continues to be a popular data gap filling technique. Our research has focused on developing an algorithmic data driven approach to read-across named GenRA to facilitate objective and reproducible predictions. The baseline approach relied primarily on chemical fingerprints to make binary predictions of toxicity. Over the last few years, a concerted effort has been made to transition to potency predictions as well as quantify the impact and contribution that other similarity contexts play in predicting toxicity. To actualise our research in practice, two implementations of the GenRA have been developed – one that facilitates programmatic access to making read-across predictions using specific data sets (genra-py; [37]). Concurrently considerable effort has been made to update and extend the functionality of GenRA as a webapp. Herein, we highlight some of the most recent features from incorporating the ability to predictions of substances not in DSSTox but through the use of a drawing palette, which returns analogues on the basis of custom hybrid fingerprints. The newest functionalities provide a means to array analogues with respect to their physicochemical profile based on predicted properties, compare and contrast the top few analogues across several FP at the same time and transition to potency predictions of toxicity as well as predictions of ToxCast assay hitcall outcomes. We have attempted to provide some context to these functionalities by walking through the outcomes using Bisphenol A as a prototypical chemical for illustration. Despite this progress, much more research is still underway that is yet to be implemented for broader use by the scientific community. Our current focus is on characterising metabolic similarity and evaluating the impact of different relative weightings using custom fingerprints. The latter should inform guiding principles on what optimal weights to use in the custom hybrid fingerprint depending on the chemical and toxicity outcome under consideration.

Acknowledgements

The authors wish to thank the rest of the GenRA Development Team for their continued efforts in implementing and deploying the GenRA web application. Special thanks go to Terry N Brown, EPA (scientific analyst, API development, QA and all round point person), Kenta Baron-Furuyama, EPA (API development), Sean Hamilton, GDIT (UI development), Carl (Freddie) Valone, EPA (UI development) and Phu (PK) Do, EPA (Scrum master).

Footnotes

Disclaimer

The authors declare that they have no competing interests. This manuscript has been reviewed by the Center for Computational Toxicology and Exposure, Office of Research and Development, U.S. Environmental Protection Agency, and approved for publication. Approval does not signify that the contents reflect the views or policy of the Agency, nor does mention of trade names or commercial products constitute endorsement or recommendation for use.

Data availability

GenRA is accessible from https://comptox.epa.gov/genra/.

References

  • [1].OECD, Guidance on Grouping of Chemicals, Second Edition ∣ en ∣ OECD, (2014). https://www.oecd.org/publications/guidance-on-grouping-of-chemicals-second-edition-9789264274679-en.htm (accessed August 10, 2021). [Google Scholar]
  • [2].Regulation (EC) No 1907/2006 of the European Parliament and of the Council of 18 December 2006 concerning the Registration, Evaluation, Authorisation and Restriction of Chemicals (REACH), establishing a European Chemicals Agency, amending Directive 1999/45/EC and repealing Council Regulation (EEC) No 793/93 and Commission Regulation (EC) No 1488/94 as well as Council Directive 76/769/EEC and Commission Directives 91/155/EEC, 93/67/EEC, 93/105/EC and 2000/21/EC, 2006. http://data.europa.eu/eli/reg/2006/1907/oj/eng (accessed September 18, 2022). [Google Scholar]
  • [3].Schultz TW, Richarz A-N, Cronin MTD, Assessing uncertainty in read-across: Questions to evaluate toxicity predictions based on knowledge gained from case studies, Computational Toxicology. 9 (2019) 1–11. 10.1016/j.comtox.2018.10.003. [DOI] [Google Scholar]
  • [4].Tier G, Gallegos SA, Pavan M, Worth A, Benigni R, Aptula A, Bassan A, Bossa C, Falk-Filipsson A, Gillet V, Jeliazkova N, Mcdougal A, Mestres J, Munro A, Netzeva T, Safford B, Simon-Hettich B, Tsakovska I, Wallén M, Chemical Similarity and Threshold of Toxicological Concern (TTC) Approaches: Report of an ECB Workshop held in Ispra, November 2005, JRC Publications Repository. (2007). https://publications.jrc.ec.europa.eu/repository/handle/JRC35474 (accessed September 18, 2022). [Google Scholar]
  • [5].Enoch S. j., Chemical Category Formation and Read-Across for the Prediction of Toxicity, in: Puzyn T, Leszczynski J, Cronin MT (Eds.), Recent Advances in QSAR Studies: Methods and Applications, Springer Netherlands, Dordrecht, 2010: pp. 209–219. 10.1007/978-1-4020-9783-6_7. [DOI] [Google Scholar]
  • [6].Patlewicz G, Ball N, Booth ED, Hulzebos E, Zvinavashe E, Hennes C, Use of category approaches, read-across and (Q)SAR: General considerations, Regulatory Toxicology and Pharmacology. 67 (2013) 1–12. 10.1016/j.yrtph.2013.06.002. [DOI] [PubMed] [Google Scholar]
  • [7].Patlewicz G, Cronin MTD, Helman G, Lambert JC, Lizarraga LE, Shah I, Navigating through the minefield of read-across frameworks: A commentary perspective, Computational Toxicology. 6 (2018) 39–54. 10.1016/j.comtox.2018.04.002. [DOI] [Google Scholar]
  • [8].ECHA, Guidance on information requirements and chemical safety assessment Chapter R.6: QSARs and grouping of chemicals, (2008). https://echa.europa.eu/documents/10162/17224/information_requirements_r6_en.pdf/77f49f81-b76d-40ab-8513-4f3a533b6ac9?t=1322594777272.
  • [9].Patlewicz G, Ball N, Becker RA, Booth ED, Cronin MTD, Kroese D, Steup D, van Ravenzwaay B, Hartung T, Read-across approaches--misconceptions, promises and challenges ahead, ALTEX. 31 (2014) 387–396. 10.14573/altex.1410071. [DOI] [PubMed] [Google Scholar]
  • [10].Ball N, Cronin MTD, Shen J, Blackburn K, Booth ED, Bouhifd M, Donley E, Egnash L, Hastings C, Juberg DR, Kleensang A, Kleinstreuer N, Kroese ED, Lee AC, Luechtefeld T, Maertens A, Marty S, Naciff JM, Palmer J, Pamies D, Penman M, Richarz A-N, Russo DP, Stuard SB, Patlewicz G, van Ravenzwaay B, Wu S, Zhu H, Hartung T, Toward Good Read-Across Practice (GRAP) guidance, ALTEX. 33 (2016) 149–166. 10.14573/altex.1601251. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Wu S, Blackburn K, Amburgey J, Jaworska J, Federle T, A framework for using structural, reactivity, metabolic and physicochemical similarity to evaluate the suitability of analogs for SAR-based toxicological assessments, Regul Toxicol Pharmacol. 56 (2010) 67–81. 10.1016/j.yrtph.2009.09.006. [DOI] [PubMed] [Google Scholar]
  • [12].Blackburn K, Stuard SB, A framework to facilitate consistent characterization of read across uncertainty, Regul Toxicol Pharmacol. 68 (2014) 353–362. 10.1016/j.yrtph.2014.01.004. [DOI] [PubMed] [Google Scholar]
  • [13].Schultz TW, Amcoff P, Berggren E, Gautier F, Klaric M, Knight DJ, Mahony C, Schwarz M, White A, Cronin MTD, A strategy for structuring and reporting a read-across prediction of toxicity, Regulatory Toxicology and Pharmacology. 72 (2015) 586–601. 10.1016/j.yrtph.2015.05.016. [DOI] [PubMed] [Google Scholar]
  • [14].Patlewicz G, Roberts DW, Aptula A, Blackburn K, Hubesch B, Workshop: use of “read-across” for chemical safety assessment under REACH, Regul Toxicol Pharmacol. 65 (2013) 226–228. 10.1016/j.yrtph.2012.12.004. [DOI] [PubMed] [Google Scholar]
  • [15].Patlewicz G, Ball N, Boogaard PJ, Becker RA, Hubesch B, Building scientific confidence in the development and evaluation of read-across, Regulatory Toxicology and Pharmacology. 72 (2015) 117–133. 10.1016/j.yrtph.2015.03.015. [DOI] [PubMed] [Google Scholar]
  • [16].Ankley GT, Bennett RS, Erickson RJ, Hoff DJ, Hornung MW, Johnson RD, Mount DR, Nichols JW, Russom CL, Schmieder PK, Serrrano JA, Tietge JE, Villeneuve DL, Adverse outcome pathways: a conceptual framework to support ecotoxicology research and risk assessment, Environ Toxicol Chem. 29 (2010) 730–741. 10.1002/etc.34. [DOI] [PubMed] [Google Scholar]
  • [17].OECD, Guidance Document for the Use of Adverse Outcome Pathways in Developing Integrated Approaches to Testing and Assessment (IATA), Organisation for Economic Co-operation and Development, Paris, 2017. https://www.oecd-ilibrary.org/environment/guidance-document-for-the-use-of-adverse-outcome-pathways-in-developing-integrated-approaches-to-testing-and-assessment-iata_44bb06c1-en;jsessionid=qIxTrvRIM6C5cT-QZyfB0GFgUAChc_ZMpz9Tt5GK.ip-10-240-5-4 (accessed September 18, 2022). [Google Scholar]
  • [18].Patlewicz G, Simon TW, Rowlands JC, Budinsky RA, Becker RA, Proposing a scientific confidence framework to help support the application of adverse outcome pathways for regulatory purposes, Regul Toxicol Pharmacol. 71 (2015) 463–477. 10.1016/j.yrtph.2015.02.011. [DOI] [PubMed] [Google Scholar]
  • [19].Patlewicz G, Simon T, Goyak K, Phillips RD, Rowlands JC, Seidel SD, Becker RA, Use and validation of HT/HC assays to support 21st century toxicity evaluations, Regul Toxicol Pharmacol. 65 (2013) 259–268. 10.1016/j.yrtph.2012.12.008. [DOI] [PubMed] [Google Scholar]
  • [20].OECD, The Adverse Outcome Pathway for Skin Sensitisation Initiated by Covalent Binding to Proteins, Organisation for Economic Co-operation and Development, Paris, 2014. https://www.oecd-ilibrary.org/environment/the-adverse-outcome-pathway-for-skin-sensitisation-initiated-by-covalent-binding-to-proteins_9789264221444-en (accessed September 18, 2022). [Google Scholar]
  • [21].Guidance Document on the Reporting of Defined Approaches to be Used Within Integrated Approaches to Testing and Assessment ∣ en ∣ OECD, (n.d.). https://www.oecd.org/publications/guidance-document-on-the-reporting-of-defined-approaches-to-be-used-within-integrated-approaches-to-testing-and-assessment-9789264274822-en.htm (accessed March 17, 2021). [Google Scholar]
  • [22].Shah I, Liu J, Judson RS, Thomas RS, Patlewicz G, Systematically evaluating read-across prediction and performance using a local validity approach characterized by chemical structure and bioactivity information, Regulatory Toxicology and Pharmacology. 79 (2016) 12–24. 10.1016/j.yrtph.2016.05.008. [DOI] [PubMed] [Google Scholar]
  • [23].Grace P, George H, Prachi P, Imran S, Navigating through the minefield of read-across tools: A review of in silico tools for grouping, Comput Toxicol. 3 (2017) 1–18. 10.1016/j.comtox.2017.05.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Williams AJ, Grulke CM, Edwards J, McEachran AD, Mansouri K, Baker NC, Patlewicz G, Shah I, Wambaugh JF, Judson RS, Richard AM, The CompTox Chemistry Dashboard: a community data resource for environmental chemistry, J Cheminform. 9 (2017) 61. 10.1186/s13321-017-0247-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Yang C, Tarkhov A, Marusczyk J, Bienfait B, Gasteiger J, Kleinoeder T, Magdziarz T, Sacher O, Schwab CH, Schwoebel J, Terfloth L, Arvidson K, Richard A, Worth A, Rathman J, New publicly available chemical query language, CSRML, to support chemotype representations for application to data mining and modeling, J Chem Inf Model. 55 (2015) 510–528. 10.1021/ci500667v. [DOI] [PubMed] [Google Scholar]
  • [26].Martin MT, Judson RS, Reif DM, Kavlock RJ, Dix DJ, Profiling Chemicals Based on Chronic Toxicity Results from the U.S. EPA ToxRef Database, Environ Health Perspect. 117 (2009) 392–399. 10.1289/ehp.0800074. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Helman G, Shah I, Williams AJ, Edwards J, Dunne J, Patlewicz G, Generalized Read-Across (GenRA): A workflow implemented into the EPA CompTox Chemicals Dashboard, ALTEX. 36 (2019) 462–465. 10.14573/altex.1811292. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Helman G, Shah I, Patlewicz G, Extending the Generalised Read-Across approach (GenRA): A systematic analysis of the impact of physicochemical property information on read-across performance, Comput Toxicol. 8 (2018) 34–50. 10.1016/j.comtox.2018.07.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Boyce M, Meyer B, Grulke C, Lizarraga L, Patlewicz G, Comparing the performance and coverage of selected in silico (liver) metabolism tools relative to reported studies in the literature to inform analogue selection in read-across: A case study, Computational Toxicology. 21 (2022) 100208. 10.1016/j.comtox.2021.100208. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [30].Nelms MD, Lougee R, Roberts DW, Richard A, Patlewicz G, Comparing and contrasting the coverage of publicly available structural alerts for protein binding, Computational Toxicology. 12 (2019) 100100. 10.1016/j.comtox.2019.100100. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [31].Tate T, Wambaugh J, Patlewicz G, Shah I, Repeat-dose toxicity prediction with Generalized Read-Across (GenRA) using targeted transcriptomic data: A proof-of-concept case study, Computational Toxicology. 19 (2021) 100171. 10.1016/j.comtox.2021.100171. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Helman G, Patlewicz G, Shah I, Quantitative prediction of repeat dose toxicity values using GenRA, Regulatory Toxicology and Pharmacology. 109 (2019) 104480. 10.1016/j.yrtph.2019.104480. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Helman G, Shah I, Patlewicz G, Transitioning the Generalised Read-Across approach (GenRA) to quantitative predictions: A case study using acute oral toxicity data, Comput Toxicol. 12 (2019). 10.1016/j.comtox.2019.100097. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Jenkins W, Tate T, Shah I, Patlewicz G. Building a compendium of expert driven read-across (EDRA) cases to investigate the utility of New Approach Methodology (NAM) data in Generalized Read-Across. Poster presentation at QSAR 2021. [Google Scholar]
  • [35].Nyffeler J, Willis C, Lougee R, Richard A, Paul-Friedman K, Harrill JA, Bioactivity screening of environmental chemicals using imaging-based high-throughput phenotypic profiling, Toxicol Appl Pharmacol. 389 (2020) 114876. 10.1016/j.taap.2019.114876. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Harrill JA, Everett LJ, Haggard DE, Sheffield T, Bundy JL, Willis CM, Thomas RS, Shah I, Judson RS, High-Throughput Transcriptomics Platform for Screening Environmental Chemicals, Toxicol Sci. 181 (2021) 68–89. 10.1093/toxsci/kfab009. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Shah I, Tate T, Patlewicz G, Generalised Read-Across prediction using genra-py, Bioinformatics. (2021) btab210. 10.1093/bioinformatics/btab210. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay É, Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res 12 (2011) 2825–2830. [Google Scholar]
  • [39].Grulke CM, Williams AJ, Thillanadarajah I, Richard AM, EPA’s DSSTox database: History of development of a curated chemistry resource supporting computational toxicology research, Comput Toxicol. 12 (2019). 10.1016/j.comtox.2019.100096. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [40].Watford S, Ly Pham L, Wignall J, Shin R, Martin MT, Friedman KP, ToxRefDB version 2.0: Improved utility for predictive and retrospective toxicology analyses, Reprod Toxicol. 89 (2019) 145–158. 10.1016/j.reprotox.2019.07.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [41].Mansouri K, Grulke CM, Judson RS, Williams AJ, OPERA models for predicting physicochemical properties and environmental fate endpoints, J Cheminform. 10 (2018) 10. 10.1186/s13321-018-0263-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [42].Michałowicz J, Bisphenol A--sources, toxicity and biotransformation, Environ Toxicol Pharmacol. 37 (2014) 738–758. 10.1016/j.etap.2014.02.003. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

GenRA is accessible from https://comptox.epa.gov/genra/.

RESOURCES