Skip to main content
PLOS One logoLink to PLOS One
. 2024 Apr 3;19(4):e0299993. doi: 10.1371/journal.pone.0299993

EndoGeneAnalyzer: A tool for selection and validation of reference genes

Eliel Barbosa Teixeira 1,*, André Salim Khayat 1, Paulo Pimentel Assumpção 1, Samir Mansour Casseb 1, Caroline Aquino Moreira-Nunes 1,2,3, Fabiano Cordeiro Moreira 1
Editor: Karel Sedlar4
PMCID: PMC10990236  PMID: 38568963

Abstract

The selection of proper reference genes is critical for accurate gene expression analysis in all fields of biological and medical research, mainly because there are many distinctions between different tissues and specimens. Given this variability, even in known classic reference genes, demands of a comprehensive analysis platform is needed to identify the most suitable genes for each study. For this purpose, we present an analysis tool for assisting in decision-making in the analysis of reverse transcription-quantitative polymerase chain reaction (RT-qPCR) data. EndoGeneAnalyzer, an open-source web tool for reference gene analysis in RT-qPCR studies, was used to compare the groups/conditions under investigation. This interactive application offers an easy-to-use interface that allows efficient exploration of datasets. Through statistical and stability analyses, EndoGeneAnalyzer assists in the select of the most appropriate reference gene or set of genes for each condition. It also allows researchers to identify and remove unwanted outliers. Moreover, EndoGeneAnalyzer provides a graphical interface to compare the evaluated groups, providing a visually informative differential analysis.

Introduction

Reverse transcription-quantitative polymerase chain reaction (RT-qPCR or qPCR) is a highly sensitive and specific technique used to study gene expression in many research fields, such as human disease, because of its capacity to detect rare transcripts and observe small variations in gene expression [14].

RT-qPCR is a technique widely used for quantifying gene expression levels. By quantifying the the RNA molecules present in a sample, RT-qPCR provides valuable insights into the expression of specific genes. To ensure accurate and reliable results, reference genes are used in the normalization process. These reference genes are stably expressed under various experimental conditions and serve as internal controls to normalize the gene expression data. Normalization with reference genes allows for a more accurate comparison of gene expression levels across different samples or experimental conditions, eliminating potential variations caused by factors such as RNA quality, sample quantity, or technical variations [5]. According to Chervoneva et al [6], among the essential criteria for choosing a good reference gene are: the level of expression unaffected by experimental factors, minimal variability in its expression between tissues and physiological states of the organism, and, preferably, that the gene has a quantification cycle (Cq) value similar that of the target gene. The Cq value indicates the position of the amplification curve with respect to the cycle axis [7].

According to the MIQE guidelines, the selection and the number of reference genes are essential, especially because they need to be experimentally validated for each specific sample type and study condition [8, 9]. Thus, ideal reference genes should have a minimum intersubject variation in terms of quantification cycle (Cq) values, and it is recommended that this variation of the reference genes between samples be less than 1 Cq [8]. This characteristic is crucial in the data normalization process for expression comparisons, as it ensures accurate mRNA concentration measurements and reliable conclusions [10]. The normalization of data using reference genes involves correcting errors that arise from the initial concentration of RNA/cDNA, and the most common method used in RT-qPCR assays involves the use of one or more reference genes [11].

Studies have demonstrated the variability of commonly used reference genes, such as GAPDH, β2M, and 18S, under various conditions [1215]. In aging studies, variability in GAPDH expression, when used as an internal control, interferes with the detection of subtle variations in the target genes under investigation [16]. Selecting an unstable reference gene for qPCR normalization can compromise experimental accuracy. Commonly used reference genes such as 18S, ACTB, and GAPDH may not always be suitable for this purpose and should be validated for stability in a specific study context [11]. Choosing an inappropriate reference gene leads to inaccurate normalization and misleading conclusions. This approach may also introduce variability and bias, hindering comparisons between samples [10].

In this context, algorithms have been developed to help identify the most appropriate reference genes from a given set of candidate genes. NormFinder [17], geNorm [18], BestKeeper [19], RefGenes [9], and RefFinder [20] are some software tools available for this purpose, with RefFinder being the only web-based tool currently available.

In this study, we developed the EndoGeneAnalyzer tool, available at https://npobioinfo.shinyapps.io/endogeneanalyzer/. The open-source code can be found at https://github.com/MoreiraFC/EndoGeneAnalyzer. This tool is a dynamic web-based tool for comparing and selecting the most stable set of reference genes from a dataset derived from RT-qPCR experiments. It also integrates NormFinder software [17]. Unlike existing algorithms, this tool allows the identification of variations by group/condition and the removal of outliers present in reference genes, a step often overlooked in most gene expression studies. Furthermore, the tool provides ability to analyze the differential expression of the target genes across different groups/conditions, allowing the investigation of differences in the expression of the gene of interest and the identification of potential associations with experimental conditions.

Materials and methods

EndoGeneAnalyzer platform

EndoGeneAnalyzer is a dynamic web-based platform that simplifies and assists in selecting reference genes in scientific studies and performing differential gene expression analysis for RT-qPCR data. With interactive interfaces and a statistical approach, the tool facilitates the identification of the best reference genes for the investigated groups or conditions. The EndoGeneAnalyzer workflow is illustrated in Fig 1.

Fig 1. EndoGeneAnalyzer tool workflow.

Fig 1

Users import data and select target genes for analysis, and the tool allows outlier removal. It also calculates stability metrics are calculated, the best reference gene is identified, and differential expression analysis between groups or conditions is enabled.

The tool has been developed to be intuitive, interactive, and efficient, with several steps to guide the user. In the first step, the user enters the data with the option to choose between the supported file formats:.xls/.xlsx or.txt/.csv. This flexibility makes it easy to import data from different sources. After loading the data, the user selects the targets of interest for analysis (nonreference genes). The next step focuses on evaluating the best reference set of genes based on mean variation using descriptive statistical data such as gene standard deviation, the sum of squared differences between the mean of each group and the gene mean, and the sum of squared differences between the standard deviation of each group and the gene standard deviation. Stability metrics are also calculated using NormFinder, which helps to identify the genes that best fit the study conditions. One of the critical innovations of EndoGeneAnalyzer is the ability to analyze the stability of reference genes. In this step, the tool allows the user to identify and remove outliers, which are samples with ΔCq mean values above or below a user-defined threshold (default = 2 standard deviations), providing flexibility in the analysis.

Finally, the EndoGeneAnalyzer can perform differential expression analysis using the target ΔCq and the mean ΔCq of the set of reference genes. This step allows accurate and efficient comparisons between different groups or conditions, further delivering a fold change result.

Data upload

This step is critical to ensure that correct information is used during the analysis. The input file must contain the following columns: i) the first column with the sample names; ii) the following columns with the mean Ct values of specific targets and reference genes for the sample; and iii) the last column with information about the groups or conditions to which the samples belong.

The data can be imported in two ways: i) the Excel tables (.xls/.xlsx) option, in which the tool does not require modification of the decimal separator; and ii) text tables (.txt/.csv) option, in which the default decimal separator is dot(.) and in which it is necessary to configure the text delimiter.

Finally, after verifying the correct formatting of the table, the user needs to click on the "Confirm Data Table" button to proceed with the analysis process. This final step ensures that the tool correctly recognizes and validates the data provided.

Data summary

The selection of target genes (nonreference) is essential because it is crucial to guide the analysis; these genes are related to the research objectives. Once the target-gene(s) have been identified, the user must click the "Update Target Gene" button to confirm the selection. This step ensures that the selected genes are processed and included in the analysis.

Gene reference samples

Outliers are atypical data values that can be identified in RT-qPCR data, with experimental errors being the leading cause of these occurrences. These errors are related to environmental conditions, instrument calibration problems, or other sources of uncontrolled variation that may occur during the experiment [21]. It is essential to be aware of outliers and to understand the potential impact of their removal on results.

EndoGeneAnalyzer identifies outliers per group for each gene and their removal can be easily performed using an available function. By default, the tool considers a sample as an outlier if the mean ΔCq > |2| standard deviation from the mean of the group/condition to which the sample belongs for the reference gene. This value can be configured according to the user’s preferences. The tool offers two methods of outlier removal: i) removal of all outliers and ii) removal of those that directly interfere with the mean Cq values of the reference genes.

The user can choose which outlier to show and remove by clicking on the "Choose which outlier to remove" radio button. "Only Mean" will show and remove only outliers in the mean of the reference genes. "All Outliers" will show and remove outliers in each gene individually. This second option tends to result in the removal of more outliers and, consequently, more samples from the analysis. Removing outliers is an interactive process since it decreases the group standard deviation and may reveal additional outliers. In addition, the tool’s dynamic interface facilitates the restoration of outliers as part of the analysis process.

Gene reference analysis

This is a crucial step in the tool’s operation, as it provides information about the reference genes and their variation between the different groups or conditions studied. Significant changes in the reference genes are observed at this stage, especially in the mean values between the analyzed groups.

The first table generated is the "Gene Reference by group", which presents information about the variation observed between the groups or conditions studied for each reference gene or the averages of the group of reference genes. The statistical tests used are the Wilcoxon-Mann-Whitney test (2 groups) or Kruskall-Wallis/Dunn test (3 or more groups). At this stage, it is expected that there will be no significant changes (p-value > 0.05) in reference genes between the studied groups or conditions.

The tool also provides the "Gene Reference Descriptive Statistics" table, which presents three fundamental values for assessing the reference genes: the gene standard deviation, the sum of squared differences between the mean of each group and the gene mean, and the sum of squared differences between the standard deviation of each group and the gene standard deviation. The formulas used to calculate the sum of squared differences are as follows: n is the number of groups, μi is the mean Cq of the group, μg is the mean Cq of the gene, σi is the standard deviation of the group and σg is the standard deviation of the gene.

sum.mean.square.diff=i=0n(μiμg)2
sum.SD.square.diff=i=0n(σiσg)2

In addition, NormFinder provides information on the stability and suitability of the reference, since this software is integrated into our tool interface. The NormFinder software employs an ANOVA-based model to account for intra- and intergroup variations. The generated table consists of four defined columns: the first column is the gene name, the second column (GroupDif) represents a measure of the difference between the groups, the third column (GroupSD) is the common standard deviation within a group, and the fourth column (Stability) provides the stability value. Thus, genes with lower stability values exhibit less variable expression and maintain a consistently stable expression pattern, while genes with higher stability values exhibit variable expression and uphold a less stable expression pattern [17].

Differential expression analysis

EndoGeneAnalyzer allows for comparison gene expression differences among the investigated groups using ΔCq. ΔCq was calculated as the difference between the target gene and the mean of the reference genes. For a given 2 groups, two statistical tests are integrated into the tool: the Pearson t test and the Wilcoxon-Mann-Whitney Rank Sum test. For the comparisons between 3 or more groups, ANOVA/Tukey’s teste and Kruskall-Wallis/Dunn test are applied.

The system also calculates the Shapiro test for the normality of each group and the Fold-Change using the formula 2-ΔΔCq [22]. This metric quantifies the difference in expression between a given two groups, considering the relative variation in ΔCq values.

Development of the EndoGeneAnalyzer web tool

The EndoGeneAnalyzer tool was developed with the Shiny framework in R studio software (v.1.7.4, https://shiny.rstudio.com/) [23] that transforms regular R code into an interactive environment that can follow and “react” to remote user instructions. The tool is compatible with all commonly used internet browsers.

The interactive visualizations and tables were rendered using the ggplot2 (v.3.4.2, https://ggplot2.tidyverse.org), knitr (v1.42, https://cran.r-project.org/web/packages/knitr), and DT (v.0.19, https://CRAN.R-project.org/package=DT) R packages. Kruskal-Wallis and Dunn’s tests were performed using the dunn.test R package (https://cran.r-project.org/web/packages/dunn.test/), and all the other statistical tests were performed using R basic statisitcs.

Results

To illustrate the usefulness of EndoGeneAnalyzer, we employed unpublished RT-qPCR data from our laboratory, which can be accessed on the "Tutorial" tab at the following web address: https://npobioinfo.shinyapps.io/endogeneanalyzer/. After confirming the table submission, the data were loaded, and the first analysis panel titled "Data Summary" will be available. In this panel, the loaded Cq averages were displayed in a graph highlighting the conditions specified during the upload (Fig 2), which initially allowed us to observe of the dynamics of the reference and target genes under the studied conditions. The user also selects the target gene for further analysis at this stage.

Fig 2. Analysis of specific reference genes.

Fig 2

This figure provides a comprehensive side-by-side analysis of specific reference genes. Each gene is visually represented by a distinct square, allowing for a clear and concise depiction of its individual attributes. Within each square, box plots showcase the mean Cq values for each condition examined in the study. The implementation of vivid colors throughout the figure serves to effectively differentiate the diverse groups or conditions under investigation, enriching our comprehension of their unique characteristics and trends.

Subsequently, as previously mentioned regarding outliers and their impact on RT-qPCR data analysis, EndoGeneAnalyzer allows handling these values to be handled. In the “Gene Reference Samples” panel, the user can identify and remove these values in two ways: i) remove "All outliers" or ii) remove those outliers that affect the reference gene averages "Only mean", as shown in Fig 3.

Fig 3. Graphical verification of potential outliers separated by analyzed groups.

Fig 3

This figure shows four rectangles, each representing the sample distribution for a single reference gene. The icons (dot, triangle, square, and cross) symbolize the conditions, with red indicating non-outlier data and blue icons indicating outlier data. At the top, the standard deviation can be adjusted, and the user can remove outliers that affect the mean or all outliers.

Furthermore, in the "Gene Reference Analysis" tab, statistical reports are generated for the selected reference genes based on ΔCt and the mean of the reference genes. These reports provide information that supports the choice of the most stable set of reference genes (Fig 4). The first report is the result of nonparametric tests, with the choice of the test being conditional on the number of groups to be compared. To select the best reference genes, it is ideal that there is no significant variation among the groups, especially in the MeanRef column. The second report provides a descriptive analysis of the reference genes (standard deviation; sum.SD.square.diff and sum.mean.square.diff), with lower values indicating better reference gene(s). The last reports are additional analyses generated by the NormFinder software integrated into EndoGeneAnalyzer [16].

Fig 4. Gene Reference Descriptive Statistics by EndoGeneAnalyzer.

Fig 4

(1) Analysis of reference genes separated by study groups. (2) Descriptive statistics of the analyzed reference genes. (3) Results generated from the NormFinder tool. The first table displays the statistical results obtained using the Kruskal-Wallis test for each reference gene and the mean of all genes. Statistically significant values are highlighted in red. The second table presents descriptive data for each examined group, indicating lower values for superior gene expression or a more favorable set of genes. The third table shows the stability data generated by the NormFinder software, providing insights into the reliability of the identified reference genes for accurate gene expression analysis.

In the example provided, the use of the GAPDH gene alone was not a good internal control (< 0.05), mainly when used in combination with ABL, TBP and RPLPO under the conditions investigated (MeanRef ≤ 0.05). However, removing GAPDH from the analysis was demonstrated to be a viable alternative for the conditions investigated (MeanRef > 0.05), given that it is a gene with high variability (Standard.Deviation = 3.35 and Stability = 1.71). In addition, since NormFinder analyses determine expression stability by assessing intra- and intergroup variation, it is possible to identify which reference gene varies between groups and exclude it from the analyses. RPLPO, TBP and ABL (0.49, 0.90, and 0.98, respectively) were the genes with the greatest stability; that is, they exhibited a consistently stable expression pattern, suggesting that they are excellent internal controls. The user can perform differential expression analysis in the last panel of the "Differential Expression Analysis" tool by selecting the target gene and comparison groups (Fig 5). In addition to the graphs generated based on ΔCq, fold change values, normality tests results, and statistical test tables are also displayed.

Fig 5. Graphical visualization of differential expression analysis separated by target and study groups.

Fig 5

The figure summarizes the differential gene expression among the examined groups using ΔCq values in box plots. The color-coded conditions and tables showing the fold change values and Shapiro-Wilk normality test results provided comprehensive information. The user interface allows the selection of target genes and groups for statistical comparisons, including Pearson’s t-test, Wilcoxon-Mann-Whitney rank sum test (for two groups), ANOVA/Tukey’s test, or the Kruskal‒Wallis/Dunn test (for three groups). This illustration facilitates the exploration of the molecular differences underlying gene expression variations between conditions.

Discussion

RT-qPCR is the gold standard for gene expression studies in molecular research and clinical practice. This method is prized for its rapidity, reproducibility, high sensitivity, and specificity, enabling the detection of gene expression even in low-yield samples. Additionally, RNA-seq is recognized as an increasingly utilized and potent tool for evaluating gene expression [24]. In real-time experiments, reference genes provide a comparative basis for assessing relative variations in the expression of target genes [18]. Therefore, the identification and validation of these genes are mandatory steps.

Some studies have questioned the use of conventionally established reference genes that may fail to enable the detection of subtle differences in target gene expression under certain conditions due to their high variability [2527] and may lead to misinterpretation of results depending on the experimental context [28].

Outlier removal is critical for the statistical analysis of qPCR data [29], as these values can bias descriptive statistics, such the as mean and standard deviation, that are used to describe gene expression. Thus, removing outliers and identifying genes that vary with the studied conditions helps to ensure accuracy and reliability in interpreting results.

EndoGeneAnalyzer is an invaluable tool for researchers conducting RT-qPCR experiments. This tool offers several features that significantly improve the accuracy and reliability of gene expression studies. It enables precise data normalization, a critical step in gene expression analysis, ensuring that the results are appropriately adjusted and comparable across samples.

Conclusion

In summary, EndoGeneAnalyzer is a new instrument used for RT-qPCR that has produced remarkable results in terms of gene expression analysis. To address the challenge of reference genes selection, our platform, which uses simulative data, was used to demonstrate the identification and assessment suitable genes in unique study settings. All the statistical analyses were combined with stability measures and outlier filtering to allow for informed selection of reference genes to provide a stable basis for data normalization.

EndoGeneAnalyzer provides a comprehensive set of features tailored to genetic research needs, including group and variation analysis, outlier removal, user-friendly data exploration, and differential expression analysis.

EndoGeneAnalyzer promises to be a useful tool for improving the quality of gene expression experiments. Such integrations are anticipated to make the research stronger and reproducible.

Data Availability

This study presents the EndoGeneAnalyzer tool, available at https://npobioinfo.shinyapps.io/endogeneanalyzer/ and the open-source code can be found at https://github.com/MoreiraFC/EndoGeneAnalyzer.

Funding Statement

This study was supported by Brazilian funding agencies: Coordination for the Improvement of Higher Education Personnel (CAPES; to E.B.T), the National Council of Technological and Scientific Development (CNPq grant number [404213/2021-9 to CAM-N; Productivity in Research PQ scholarships to P.P.A, A.S.K., and CAM-N]), and the Cearense Foundation of Scientific and Technological Support (FUNCAP grant number [P20-0171-00078.01.00/20 to CAM-N]); we also thank PROPESP/UFPA for the publication payment. There was no additional external funding received for this study.

References

  • 1.Gutala R V, Reddy PH. The use of real-time PCR analysis in a gene expression study of Alzheimer’s disease post-mortem brains. J Neurosci Methods. 2004;132: 101–107. doi: 10.1016/j.jneumeth.2003.09.005 [DOI] [PubMed] [Google Scholar]
  • 2.Nogueira BMD, da Costa Pantoja L, da Silva EL, Mello Júnior FAR, Teixeira EB, Wanderley AV, et al. Telomerase (hTERT) Overexpression Reveals a Promising Prognostic Biomarker and Therapeutical Target in Different Clinical Subtypes of Pediatric Acute Lymphoblastic Leukaemia. Genes (Basel). 2021;12: 1632. doi: 10.3390/genes12101632 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Chaves JR, de Souza CRT, Modesto AAC, Moreira FC, Teixeira EB, Sarraf JS, et al. Effects of alkaline water intake on gastritis and miRNA expression (miR-7, miR-155, miR-135b and miR-29c). Am J Transl Res. 2020;12: 4043–4050. [PMC free article] [PubMed] [Google Scholar]
  • 4.Ho-Pun-Cheung A, Bascoul-Mollevi C, Assenat E, Boissière-Michot F, Bibeau F, Cellier D, et al. Reverse transcription-quantitative polymerase chain reaction: description of a RIN-based algorithm for accurate data normalization. BMC Mol Biol. 2009;10: 31. doi: 10.1186/1471-2199-10-31 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Jin P, Zhao Y, Ngalame Y, Panelli MC, Nagorsen D, Monsurró V, et al. Selection and validation of endogenous reference genes using a high throughput approach. BMC Genomics. 2004;5: 55. doi: 10.1186/1471-2164-5-55 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chervoneva I, Li Y, Schulz S, Croker S, Wilson C, Waldman SA, et al. Selection of optimal reference genes for normalization in quantitative RT-PCR. BMC Bioinformatics. 2010;11: 253. doi: 10.1186/1471-2105-11-253 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Ruiz-Villalba A, Ruijter JA, Van den Hoff MJB. Use and Misuse of Cq in qPCR Data Analysis and Reporting. Life. 2021; 6: 496. doi: 10.3390/life11060496 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Bustin SA, Benes V, Garson JA, Hellemans J, Huggett J, Kubista M, et al. The MIQE Guidelines: Minimum Information for Publication of Quantitative Real-Time PCR Experiments. Clin Chem. 2009;55: 611–622. doi: 10.1373/clinchem.2008.112797 [DOI] [PubMed] [Google Scholar]
  • 9.Hruz T, Wyss M, Docquier M, Pfaffl MW, Masanetz S, Borghi L, et al. RefGenes: identification of reliable and condition specific reference genes for RT-qPCR data normalization. BMC Genomics. 2011;12: 156. doi: 10.1186/1471-2164-12-156 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.González-Bermúdez L, Anglada T, Genescà A, Martín M, Terradas M. Identification of reference genes for RT-qPCR data normalisation in aging studies. Sci Rep. 2019;9: 13970. doi: 10.1038/s41598-019-50035-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Chapman JR, Waldenström J. With Reference to Reference Genes: A Systematic Review of Endogenous Controls in Gene Expression Studies. PLoS One. 2015;10: e0141853. doi: 10.1371/journal.pone.0141853 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Touchberry CD, Wacker MJ, Richmond SR, Whitman SA, Godard MP. Age-related changes in relative expression of real-time PCR housekeeping genes in human skeletal muscle. J Biomol Tech. 2006;17: 157–62. [PMC free article] [PubMed] [Google Scholar]
  • 13.Zampieri M, Ciccarone F, Guastafierro T, Bacalini MG, Calabrese R, Moreno-Villanueva M, et al. Validation of suitable internal control genes for expression studies in aging. Mech Ageing Dev. 2010;131: 89–95. doi: 10.1016/j.mad.2009.12.005 [DOI] [PubMed] [Google Scholar]
  • 14.Kozera B, Rapacz M. Reference genes in real-time PCR. J Appl Genet. 2013;54: 391–406. doi: 10.1007/s13353-013-0173-x [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Butterfield DA, Hardas SS, Lange MLB. Oxidatively Modified Glyceraldehyde-3-Phosphate Dehydrogenase (GAPDH) and Alzheimer’s Disease: Many Pathways to Neurodegeneration. Journal of Alzheimer’s Disease. 2010;20: 369–393. doi: 10.3233/JAD-2010-1375 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Radonić A, Thulke S, Mackay IM, Landt O, Siegert W, Nitsche A. Guideline to reference gene selection for quantitative real-time PCR. Biochem Biophys Res Commun. 2004;313: 856–862. doi: 10.1016/j.bbrc.2003.11.177 [DOI] [PubMed] [Google Scholar]
  • 17.Andersen CL, Jensen JL, Ørntoft TF. Normalization of Real-Time Quantitative Reverse Transcription-PCR Data: A Model-Based Variance Estimation Approach to Identify Genes Suited for Normalization, Applied to Bladder and Colon Cancer Data Sets. Cancer Res. 2004;64: 5245–5250. doi: 10.1158/0008-5472.CAN-04-0496 [DOI] [PubMed] [Google Scholar]
  • 18.Vandesompele J, De Preter K, Pattyn F, Poppe B, Van Roy N, De Paepe A, et al. Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes. Genome Biol. 2002;3: RESEARCH0034. doi: 10.1186/gb-2002-3-7-research0034 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Pfaffl MW, Tichopad A, Prgomet C, Neuvians TP. Determination of stable housekeeping genes, differentially regulated target genes and sample integrity: BestKeeper–Excel-based tool using pair-wise correlations. Biotechnol Lett. 2004;26: 509–515. doi: 10.1023/b:bile.0000019559.84305.47 [DOI] [PubMed] [Google Scholar]
  • 20.Xie F, Wang J, Zhang B. RefFinder: a web-based tool for comprehensively analyzing and identifying reference genes. Funct Integr Genomics. 2023;23: 125. doi: 10.1007/s10142-023-01055-7 [DOI] [PubMed] [Google Scholar]
  • 21.Burns MJ, Nixon GJ, Foy CA, Harris N. Standardisation of data from real-time quantitative PCR methods–evaluation of outliers and comparison of calibration curves. BMC Biotechnol. 2005;5: 31. doi: 10.1186/1472-6750-5-31 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Livak KJ, Schmittgen TD. Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2−ΔΔCT Method. Methods. 2001;25: 402–408. doi: 10.1006/meth.2001.1262 [DOI] [PubMed] [Google Scholar]
  • 23.R Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing. Vienna, Austria: URL https://www.R-project.org/. [Google Scholar]
  • 24.Bustin S. Quantification of mRNA using real-time reverse transcription PCR (RT-PCR): trends and problems. J Mol Endocrinol. 2002;29: 23–39. doi: 10.1677/jme.0.0290023 [DOI] [PubMed] [Google Scholar]
  • 25.Solanas M, Moral R, Escrich E. Unsuitability of Using Ribosomal RNA as Loading Control for Northern Blot Analyses Related to the Imbalance between Messenger and Ribosomal RNA Content in Rat Mammary Tumors. Anal Biochem. 2001;288: 99–102. doi: 10.1006/abio.2000.4889 [DOI] [PubMed] [Google Scholar]
  • 26.Röhn G, Koch A, Krischek B, Stavrinou P, Goldbrunner R, Timmer M. ACTB and SDHA Are Suitable Endogenous Reference Genes for Gene Expression Studies in Human Astrocytomas Using Quantitative RT-PCR. Technol Cancer Res Treat. 2018;17: 153303381880231. doi: 10.1177/1533033818802318 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ayakannu T, Taylor AH, Konje JC. Selection of Endogenous Control Reference Genes for Studies on Type 1 or Type 2 Endometrial Cancer. Sci Rep. 2020;10: 8468. doi: 10.1038/s41598-020-64663-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Köhsler M, Leitsch D, Müller N, Walochnik J. Validation of reference genes for the normalization of RT-qPCR gene expression in Acanthamoeba spp. Sci Rep. 2020;10: 10362. doi: 10.1038/s41598-020-67035-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Goñi R GPFS. The qPCR data statistical analysis. Integromics White Paper. 2009;1: 1–9. [Google Scholar]

Decision Letter 0

Karel Sedlar

27 Oct 2023

PONE-D-23-21180EndoGeneAnalyzer: a tool for selection and validation of reference genesPLOS ONE

Dear Dr. TEIXEIRA,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

First of all, I would like to apologize for really long review process but I was struggling to find reviewers for your manuscript. I somehow expected that as your topic is neglected by many researchers. However, in my opinion, searching for HKGs is very important topic and tools as EndoGeneAnalyzer are of great uses for the community. Nevertheless, both reviewers raised numerous questions regarding your tool. In my eyes, they are meant to improve your manuscript and tool which could bring you more users and citations in the future. Both reviewers in many cases pointed to the same issues. Thus, please try to address all of their concerns.

Please submit your revised manuscript by Dec 11 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Karel Sedlar, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Did you know that depositing data in a repository is associated with up to a 25% citation advantage (https://doi.org/10.1371/journal.pone.0230416)? If you’ve not already done so, consider depositing your raw data in a repository to ensure your work is read, appreciated and cited by the largest possible audience. You’ll also earn an Accessible Data icon on your published paper if you deposit your data in any participating repository (https://plos.org/open-science/open-data/#accessible-data).

3. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, all author-generated code must be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

4. Thank you for stating in your Funding Statement:

“This study was supported by Brazilian funding agencies: Coordination for the Improvement of Higher Education Personnel (CAPES; to E.B.T), the National Council of Technological and Scientific Development (CNPq grant number [404213/2021-9 to CAM-N; Productivity in Research PQ scholarships to P.P.A, A.S.K., and CAM-N]), and the Cearense Foundation of Scientific and Technological Support (FUNCAP grant number [P20-0171-00078.01.00/20 to CAM-N]); we also thank PROPESP/UFPA for the publication payment.”

Please provide an amended statement that declares *all* the funding or sources of support (whether external or internal to your organization) received during this study, as detailed online in our guide for authors at http://journals.plos.org/plosone/s/submit-now.  Please also include the statement “There was no additional external funding received for this study.” in your updated Funding Statement.

Please include your amended Funding Statement within your cover letter. We will change the online submission form on your behalf.

5. We note that you have stated that you will provide repository information for your data at acceptance. Should your manuscript be accepted for publication, we will hold it until you provide the relevant accession numbers or DOIs necessary to access your data. If you wish to make changes to your Data Availability statement, please describe these changes in your cover letter and we will update your Data Availability statement to reflect the information you provide.

6. Thank you for stating the following in the Acknowledgments Section of your manuscript:

We note that you have provided additional information within the Acknowledgements Section that is not currently declared in your Funding Statement. Please note that funding information should not appear in the Acknowledgments section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form.

Please remove any funding-related text from the manuscript and let us know how you would like to update your Funding Statement. Currently, your Funding Statement reads as follows:

“This study was supported by Brazilian funding agencies: Coordination for the Improvement of Higher Education Personnel (CAPES; to E.B.T), the National Council of Technological and Scientific Development (CNPq grant number [404213/2021-9 to CAM-N; Productivity in Research PQ scholarships to P.P.A, A.S.K., and CAM-N]), and the Cearense Foundation of Scientific and Technological Support (FUNCAP grant number [P20-0171-00078.01.00/20 to CAM-N]); we also thank PROPESP/UFPA for the publication payment.”

Please include your amended statements within your cover letter; we will change the online submission form on your behalf.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: N/A

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: No

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Comments on the app

• In the Shiny app, on the “Gene Reference Analysis” page – please provide some extra written context to help with interpretation. For example, that the numbers in the topmost table are p-values, with significant p-vals in red. For the Normfinder Analysis, a brief description of how to interpret results to select the optimal combination of reference genes would be helpful (either on this page or the tutorial page). For example, what range of stability is acceptable for a gene to be retained as a reference, and in what range should one consider removing that gene before proceeding with gene expression analysis. Indeed, the ReadMe/Tutorial should provide a tutorial using the sample data for how to correctly analyse this dataset – what parameters provide evidence that a reference gene should be included/excluded, how this affects the interpretation of gene expression results – e.g. the tutorial could show how not excluding a reference gene with differential expression results in an incorrect conclusion on the differential analysis page. While the MS does go some way to assist in interpretation of results, the authors should assume that some users may access the app without reading the paper.

• The same comment applies to the Differential Analysis page. Please provide some context to assist in interpreting results. The MS (line 98) states that fold-change results are provided on this page, but the word “fold” is not used on this page at all, making it potentially difficult to interpret/understand the provided results (especially to researchers new to the field). Also suggest renaming “Differential Analysis” to “Differential Expression Analysis” or similar (to make clear that this page is about analysing the actual results of a gene expression experiment, after prior reference gene selection.

• On the Gene Reference samples page – it would be useful to have functionality to add outliers back into the analysis if the user chooses, so that the impact of removal and retention of these samples can more easily be explored. Also suggest adding some additional wording about what the two options for the radio buttons mean – this is clearer in the MS (lines 126-127) then in the software itself.

• On the Gene Reference Analysis page – Have you considered adding an analysis to determine the optimal number and/or combination of reference genes? Many of the algorithms do this. Using the intragroup variance would be a good place to start. As stated in the original Normfinder paper: “Intragroup variance estimates provide a natural way of identifying the number of genes to include. The optimal number of genes is reached when addition of a further gene leads to a negligible reduction in the average of the gene variance estimates.”

• On the Gene Reference Analysis page – please provide a key for abbreviations (IGroup = intergroup, SD = standard deviation etc)

• On the “Differential Analysis” page – it might be useful to offer an option to download the ggplot script as well as the PDF of the figure. This would allow users to edit the figure for publication. Also an option to download the tables. Figure and table download options on the Data Summary, Gene Reference Samples and Gene Reference Analysis pages would also be useful, for example so outputs can be saved in electronic notebooks. These are just suggestions though!

• I found the timing out somewhat annoying. Would it be possible to, for e.g., add a popup warning box along the lines of “you are about to be disconnected, do you wish to continue this session?". Or at least note on the tutorial page the amount of inactivity time a user has before they are disconnected from the server. Continually having to reupload the data was quite tedious.

• It would be helpful if the sample data file had a larger set of candidate reference genes, with some that obviously show differential expression and should therefore be excluded, and others that are very stable and should be included. In the currently provided dataset, interpretation is a little ambiguous (each ref gene is stable for some comparisons but not others. While this likely mimics real-world data much of the time, for the purposes of a demo dataset, more clear-cut interpretation would be useful). The file should also include more (and some very obvious) sample outliers.

Comments on the MS

• Introduction – please add at least one more sentence to connect between generic description of RT-qPCR (lines 41-42) and selection of reference genes (42-44). In what way is RT-qPCR useful for quantifying gene expression, what are reference genes, and why are they important – how do they normalise data? Apart from minimising differences in CT between treatment groups, what other properties are important for ideal reference genes (ubiquitous expression etc). Although this is a brief communication, I think a little more context in these opening sentences of the introduction would be helpful to provide context for why your software can play an important role in the field.

• Line 66 – none of the authors of Normfinder are authors of the current publication, and they are also not mentioned in the acknowledgements section. Please confirm that they have provided permission to recreate their algorithm in your app, and acknowledge them accordingly.

• Lines 187-217 – in the MS the ‘Gene Reference Analysis’ is discussed before the ‘Gene Reference Samples’ tab. But in the app, the gene ref samples tab is before the gene ref analysis tab. Please reorder either the MS or the app to be consistent. The order should be whichever process should occur first (ref gene ID or outlier removal) is discussed first in the MS and is the first of these two tabs in the app..

• The figures are based on a more comprehensive example dataset than provided as the example data file in the app. For example, RPLPO is not in the sample data, neither is TBP. HPRT is in the sample data but not in Fig4. There are more outliers in the Figure 4 than the provided dataset. Line 205 refers to condition 3, but provided file only has conditions 1 and 2. It would be useful if the figures in the MS where based on the same dataset as the provided sample dataset.

Minor comments

Line 30 – change ‘instrument’ to ‘platform’

Line 31 – amend to ‘an analysis tool to assist in decision making’

Line 34 – break into two sentences ‘under investigation. This interactive…’

Lines 32 and 41 – spell out RT-qPCR at first use (both in abstract and main body of MS)

Lines 55-57 – I would reword slightly. Nothing wrong with those reference genes per se, but need to be validated as stable in the particular context of the study.

Lines 59-60 – suggest citing reference directly after each algorithm, rather than grouping all refs at end of sentence

Line 135 – add fullstop after ‘genes’

Lines 148-149 – please add some further information about how to interpret NormFinder results and use these to select an optimal set of reference genes.

Line 157 – please provide reference for the 2-ΔΔCT method.

Line 183 – remove ‘meticulously constructed’

Lines 197-203 – change ‘screen’ to ‘table’ or ‘analysis’

Line 231 – change ‘the’ to ‘a’. RNAseq is also a very standard and reliable method these days.

Line 246 – remove ‘firstly’ or add other points to the sentence.

Lines 276 & 312 – same reference

Reviewer #2: Dear Authors,

I have carefully reviewed your manuscript on the "EndoGeneAnalyzer," a tool for the analysis of candidate reference genes for RT-qPCR studies. Unfortunately, I must express my concerns regarding the quality and completeness of the manuscript. While the idea behind the tool is promising, several critical issues need to be addressed before the paper can be considered for publication. Below, I've outlined my specific concerns and suggestions for improvement:

1. Data Origin and Description:

- The manuscript lacks a clear explanation of the origin of the test data used in the analysis. The phrase "To illustrate the usefulness of the EndoGeneAnalyzer, we used the RT-qPCR data found at " ext-link-type="uri" xlink:type="simple">https://npobioinfo.shinyapps.io/endogeneanalyzer/" is insufficient. You need to provide a comprehensive description of how you obtained the RT-qPCR data, including data sources, collection methods, and any relevant details.

2. Methodological Issues:

- There are issues with the method for identifying outliers that are unclear. It's vital to provide a detailed and transparent explanation of the outlier detection process.

- The use of boxplots combined with dots for data visualization is questioned. Authors should justify this choice or consider alternative visualization methods.

- The origin of values in the table "Reference of genes per group" / "Gene Reference by Group" and the color-coding is not adequately explained. Clarify what the values represent and the rules for color-coding.

- The paper lacks a clear rationale for the choice of statistical tests for assessing the stability of reference genes. Additionally, it does not compare the proposed tool with existing reference gene selection methods, such as NormFinder, geNorm, BestKeeper, RefGenes, or RefFinder, to establish its superiority.

- The reasons for conducting differential analysis and normality tests are unclear. Authors need to provide a more comprehensive description of how these steps aid in the selection of reference genes.

3. Conclusion:

- The conclusion is notably brief and fails to adequately summarize the findings. It lacks an overall assessment of the tool's effectiveness and significance.

4. Inconsistencies and Confusing Statements:

- The reference to "Postgraduate program in Biotechnology, Federal University of Pará, Belém, Pará, Brazil" as the institution appears unusual and should be clarified.

- Several sentences throughout the manuscript are confusing and should be rephrased for clarity such as lines 31-34 and 54-57.

- The use of "Ct (threshold cycle)" is inconsistent with MIQE guidelines, which recommend "Cq (quantification cycle)" values.

- Address the ambiguity in the statement "ΔCt greater or less than 2" (Line 124) for better clarity.

- Correct the statement "(p-value in reference genes between the studied groups or conditions" (Line 137) for clarity.

5. References:

- Ensure that references are correctly linked to their respective tools (line 60), and provide an accurate reference list.

- Correct punctuation and grammar in the text, including missing commas.

6. Figures:

- Figures 2, 3, 4, and 5 do not correspond with the current state of the tool, leading to confusion. Update these figures to accurately represent the tool.

- Figure 1 is misleading due to the use of dotted lines. Reconfigure the workflow diagram to eliminate confusion and better summarize the steps for users.

7. Tool Inconsistencies:

- Address inconsistencies in the tool's interface and the manuscript regarding the naming of sections and options.

- Resolve issues with the "Update Target Gene" button not functioning correctly.

- Clarify the purpose and effects of the "Select one or more Genes" option in the tool.

- Make sure that the results in the NormFinder analysis are presented in a clear and understandable manner.

- Add titles to all sections in the Differential Analysis tab for clarity.

8. Data Saving and Reporting:

- Consider implementing a feature to save the results of all analyses.

- Explore the possibility of generating downloadable reports for users, which can enhance the utility of the tool.

Overall, I believe your manuscript and tool require substantial revisions and clarifications to meet the standards for publication. Addressing the issues, I raised could substantially improve the quality and user-friendliness of your tool and its associated documentation a thus, bring you more users in the future.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2024 Apr 3;19(4):e0299993. doi: 10.1371/journal.pone.0299993.r002

Author response to Decision Letter 0


30 Nov 2023

RESPONSE LETTER TO THE REVIEWERS

Dear reviewer, my co-authors and I would like to thank you for the suggestions made during this high-quality review and then we present the answer to the questions.

We inform that with the reviews and suggestions, we were able to improve the idea presented by our work and we appreciate the opportunity. We hope this review has left the article suitable for publication in this high-impact and prestigious journal.

Kind Regards.

Reviewer #1: Comments on the MS

• Introduction – please add at least one more sentence to connect between generic description of RT-qPCR (lines 41-42) and selection of reference genes (42-44). In what way is RT-qPCR useful for quantifying gene expression, what are reference genes, and why are they important – how do they normalise data? Apart from minimising differences in CT between treatment groups, what other properties are important for ideal reference genes (ubiquitous expression etc). Although this is a brief communication, I think a little more context in these opening sentences of the introduction would be helpful to provide context for why your software can play an important role in the field.

R = Thank you for the suggestion. We have carefully considered your suggestions and addressed each point in the revised manuscript. Specifically, in response to your comment on the Introduction section.

• Line 66 – none of the authors of Normfinder are authors of the current publication, and they are also not mentioned in the acknowledgements section. Please confirm that they have provided permission to recreate their algorithm in your app, and acknowledge them accordingly.

R= Thank you for the suggestion. We have reached out to the authors; however, we have not received any response thus far. It is important to note that we are not modifying or recreating the code of the NormFinder tool; rather, we are simply incorporating it into our platform, making authorization unnecessary. Furthermore, the article is being cited in our manuscript.

• Lines 187-217 – in the MS the ‘Gene Reference Analysis’ is discussed before the ‘Gene Reference Samples’ tab. But in the app, the gene ref samples tab is before the gene ref analysis tab. Please reorder either the MS or the app to be consistent. The order should be whichever process should occur first (ref gene ID or outlier removal) is discussed first in the MS and is the first of these two tabs in the app..

R= The tabs 'Gene Reference Samples' and 'Gene Reference Analysis,' respectively, have been reorganized in the manuscript to be consistent with the platform.

• The figures are based on a more comprehensive example dataset than provided as the example data file in the app. For example, RPLPO is not in the sample data, neither is TBP. HPRT is in the sample data but not in Fig4. There are more outliers in the Figure 4 than the provided dataset. Line 205 refers to condition 3, but provided file only has conditions 1 and 2. It would be useful if the figures in the MS where based on the same dataset as the provided sample dataset.

R= We appreciate your valuable feedback. We value your observation regarding the inconsistencies in the figures and dataset. We have carefully reviewed your comments and are committed to addressing these concerns. The figures have been revised to ensure alignment with the provided sample dataset to enhance clarity and accuracy.

Minor comments

Line 30 – change ‘instrument’ to ‘platform’

R=Thank you for your suggestion. We have incorporated the change and replaced the term 'instrument' with 'platform' on line 30.

Line 31 – amend to ‘an analysis tool to assist in decision making’

R=Thank you for your suggestion. We have revised line 31 to now read 'an analysis tool to assist in decision making' for improved clarity and precision.

Line 34 – break into two sentences ‘under investigation. This interactive…’

R: Thank you for your suggestion. We have revised line 34 by breaking it into two sentences: 'under investigation.' and 'This interactive...' for improved readability.

Lines 32 and 41 – spell out RT-qPCR at first use (both in abstract and main body of MS)

R=Thank you for your suggestion. We have revised both the abstract and the main body of the manuscript to spell out RT-qPCR at its first use in lines 32 and 41 for improved clarity.

Lines 55-57 – I would reword slightly. Nothing wrong with those reference genes per se, but need to be validated as stable in the particular context of the study.

R: Thank you for your comment. We have thoroughly analyzed the genes and taken measures to ensure their stability. In fact, we have even inserted additional genes to further enhance their stability.

Lines 59-60 – suggest citing reference directly after each algorithm, rather than grouping all refs at end of sentence

R=Thank you for your feedback. We have made the suggested modification by citing the reference directly after each algorithm, as opposed to grouping all references at the end of the sentence.

Line 135 – add fullstop after ‘genes’

R= Thank you for your suggestion. We have added a full stop after 'genes' in line 135 for improved punctuation.

Lines 148-149 – please add some further information about how to interpret NormFinder

results and use these to select an optimal set of reference genes.

R=Thank you for your insightful suggestion. We have included additional information in lines 148-149 to provide a more comprehensive understanding of how to interpret the NormFinder output.

Line 157 – please provide reference for the 2-ΔΔCT method.

R= We have provide reference for the 2-ΔΔCT method

LIVAK, Kenneth J.; SCHMITTGEN, Thomas D. Analysis of relative gene expression data using real-time quantitative PCR and the 2− ΔΔCT method. methods, v. 25, n. 4, p. 402-408, 2001.

Line 183 – remove ‘meticulously constructed’

R=Thank you for your suggestion. We have removed the phrase 'meticulously constructed' from line 183 to align with your feedback.

Lines 197-203 – change ‘screen’ to ‘table’ or ‘analysis’

R=Thank you for your suggestion. We have made the change from 'screen' to 'table' in lines 197-203, as per your recommendation.

Line 231 – change ‘the’ to ‘a’. RNAseq is also a very standard and reliable method these days.

R= Thank you for your suggestion. We appreciate your keen observation. In response to your suggestion, we have revised line 231 by replacing 'the' with 'a.' and we added about RNA-Seq in MS

Line 246 – remove ‘firstly’ or add other points to the sentence.

R=Thank you for your suggestion. We have removed 'firstly' from line 246 to enhance the clarity of the sentence.

Lines 276 312 – same reference

R=Thank you for your suggestion. We have made the modification in line 312 as suggested, using the reference from RÖHN, Gabriele et al. "ACTB and SDHA are suitable endogenous reference genes for gene expression studies in human astrocytomas using quantitative RT-PCR. Technology in cancer research treatment, v. 17, p. 1533033818802318, 2018."

Reviewer #1: Comments on the app

• In the Shiny app, on the “Gene Reference Analysis” page – please provide some extra written context to help with interpretation. For example, that the numbers in the topmost table are p-values, with significant p-vals in red.

For the Normfinder Analysis, a brief description of how to interpret results to select the optimal combination of reference genes would be helpful (either on this page or the tutorial page). For example, what range of stability is acceptable for a gene to be retained as a reference, and in what range should one consider removing that gene before proceeding with gene expression analysis.

R=Thank you. We added the following text at the Gene Reference Analysis page:

Gene Reference by group

The first line of the table is the p-value of the Kruskall-Wallis test for each reference gene.

The other lines are p-values of the Wilcoxon-Mann-Whitney test between each condition/group.

The last column is the test p-value of the reference genes mean.

Red values indicate p-values 0.05. It is desirable that reference genes and/or the reference genes mean do not vary significantly among groups/conditions.

Gene Reference Descriptive Statistics

It is preferable for reference genes not to vary across different conditions or groups.

Additionally, it is desirable for reference genes to exhibit low variances,

enhancing the ability to discern small yet significant differences in the target.

If a reference gene is varying across group/condition and/or exhibits high variances and/or has low stability, consider removing it from the analysis.

Normfinder Analysis

Normfinder is a method to evaluate reference gene stability using the variation of the gene expression.

The lower the stability value, the more stable the gene is considered.

Indeed, the ReadMe/Tutorial should provide a tutorial using the sample data for how to correctly analyse this dataset – what parameters provide evidence that a reference gene should be included/excluded, how this affects the interpretation of gene expression results –

e.g. the tutorial could show how not excluding a reference gene with differential expression results in an incorrect conclusion on the differential analysis page. While the MS does go some way to assist in interpretation of results, the authors should assume that some users may access the app without reading the paper.

R= Thank you for your suggestion. At the tutorial section we added an extensive explanation of the tool, using the new example data set. We included data of a reference gene that should clearly be removed from the analysis, and showed its impact at the differential analysis.

• The same comment applies to the Differential Analysis page. Please provide some context to assist in interpreting results. The MS (line 98) states that fold-change results are provided on this page, but the word “fold” is not used on this page at all, making it potentially difficult to interpret/understand the provided results (especially to researchers new to the field). Also suggest renaming “Differential Analysis” to “Differential Expression Analysis” or similar (to make clear that this page is about analysing the actual results of a gene expression experiment, after prior reference gene selection.

R=Thank you. We changed the section name, as suggested. Included the topic "fold-change" with succinct explanation of its meaning at the tutorial section.

• On the Gene Reference samples page – it would be useful to have functionality to add outliers back into the analysis if the user chooses, so that the impact of removal and retention of these samples can more easily be explored. Also suggest adding some additional wording about what the two options for the radio buttons mean – this is clearer in the MS (lines 126-127) then in the software itself.

R=Thank you. We added a button to restore removed outliers, as suggested, and explain the impact of the two radio button options at the tutorial section.

• On the Gene Reference Analysis page – Have you considered adding an analysis to determine the optimal number and/or combination of reference genes? Many of the algorithms do this. Using the intragroup variance would be a good place to start. As stated in the original Normfinder paper: “Intragroup variance estimates provide a natural way of identifying the number of genes to include. The optimal number of genes is reached when addition of a further gene leads to a negligible reduction in the average of the gene variance estimates.”

Thank you for the suggestion. While this analysis wasn't initially included in the design, we can incorporate it in future versions, along with other features based on user feedback.

• On the Gene Reference Analysis page – please provide a key for abbreviations (IGroup = intergroup, SD = standard deviation etc)

Thank you for your suggestion. We added in this section the respectives key words for abbreviations.

• On the “Differential Analysis” page – it might be useful to offer an option to download the ggplot script as well as the PDF of the figure. This would allow users to edit the figure for publication. Also an option to download the tables. Figure and table download options on the Data Summary, Gene Reference Samples and Gene Reference Analysis pages would also be useful, for example so outputs can be saved in electronic notebooks. These are just suggestions though!

R= In the Differential Expression Analysis section, we've introduced a button that allows users to download all analysis results. This report includes the script employed to generate the Expression Boxplot. We appreciate your suggestion; however, we believe the steps involved in analyzing the reference genes are intermediate to obtaining the final result. Therefore, we have decided, for the time being, not to include downloads for these intermediate reports.

• I found the timing out somewhat annoying. Would it be possible to, for e.g., add a popup warning box along the lines of “you are about to be disconnected, do you wish to continue this session?". Or at least note on the tutorial page the amount of inactivity time a user has before they are disconnected from the server. Continually having to reupload the data was quite tedious.

R=Thank you for your suggestion. As far as we know, the time out is controlled by the shiny server. We may give more time to each user, however we don't have any more management over it.

• It would be helpful if the sample data file had a larger set of candidate reference genes, with some that obviously show differential expression and should therefore be excluded, and others that are very stable and should be included. In the currently provided dataset, interpretation is a little ambiguous (each ref gene is stable for some comparisons but not others. While this likely mimics real-world data much of the time, for the purposes of a demo dataset, more clear-cut interpretation would be useful). The file should also include more (and some very obvious) sample outliers.

R=Thank you for your suggestion. We update the sample dataset, with clearer examples. And used it at the tutorial section for users to follow it step-by-step, explaining each step of the analysis.

Reviewer #2: Comments on the MS

1. Data Origin and Description:

- The manuscript lacks a clear explanation of the origin of the test data used in the analysis. The phrase "To illustrate the usefulness of the EndoGeneAnalyzer, we used the RT-qPCR data found at https://npobioinfo.shinyapps.io/endogeneanalyzer/" is insufficient. You need to provide a comprehensive description of how you obtained the RT-qPCR data, including data sources, collection methods, and any relevant details.

R= We appreciate your detailed observations. In response to your concern regarding the origin of the test data used in the analysis, we would like to clarify that the data were obtained from an unpublished experiment conducted in our laboratory. We aimed to ensure the integrity and authenticity of the data, reflecting a real-world context. It is pertinent to note that, while the data originate from a genuine experiment, we made alterations to the names of the target genes and the groups/conditions to preserve confidentiality and avoid any potential conflicts of interest.

2. Methodological Issues:

- There are issues with the method for identifying outliers that are unclear. It's vital to provide a detailed and transparent explanation of the outlier detection process.

Outliers were identified as samples with a mean Cq two standard deviations away from their respective group in a specific gene, or from the mean of all reference genes (based on the user's selection of 'All outliers' or 'Reference Mean'). This information has been incorporated into the tool tutorial.

- The use of boxplots combined with dots for data visualization is questioned. Authors should justify this choice or consider alternative visualization methods.

A combination of boxplot with scatter plots is intriguing for observing the distribution of samples, which may be obscured by outliers and small sample-sized groups. However, we acknowledge that there are various ways to visualize this data. Therefore, we have included a functionality that allows users to choose the best plot type, including boxplot + dispersion, boxplot, scatter plot, and violin plot. Additionally, the tool generates a report, providing users access to tables and scripts for the plots, enabling them to create their own visualizations

- The origin of values in the table "Reference of genes per group" / "Gene Reference by Group" and the color-coding is not adequately explained. Clarify what the values represent and the rules for color-coding.

R= Red values indicate p-values 0.05. We added a text in this section informing it. We also explained it at the tutorial section.

- The paper lacks a clear rationale for the choice of statistical tests for assessing the stability of reference genes. Additionally, it does not compare the proposed tool with existing reference gene selection methods, such as NormFinder, geNorm, BestKeeper, RefGenes, or RefFinder, to establish its superiority.

R=Statistical tests were conducted to assess the variation of reference genes across different groups/conditions. This is crucial as, when calculating Delta Cq, any variability in reference genes can be transferred to target genes, potentially leading to misinterpretations in the differential analysis. This serves as an extra measure of stability, complementing the stability assessment provided by NormFinder. Furthermore, we integrated an outlier analysis, distinguishing our tool from others that evaluate reference genes but do not include this additional analytical step.

- The reasons for conducting differential analysis and normality tests are unclear. Authors need to provide a more comprehensive description of how these steps aid in the selection of reference genes.

R=The differential expression analysis is an additional step that is not directly related to the evaluation of reference genes. Nevertheless, considering that all reference genes are already integrated into the tool, we believe it would enhance the tool's appeal and utility if the analysis of differential expression for target genes were also included.

3. Conclusion:

- The conclusion is notably brief and fails to adequately summarize the findings. It lacks an overall assessment of the tool's effectiveness and significance.

R=Thank you for your suggestion. In our revised manuscript, we will ensure to provide a more comprehensive summary of the findings, including a thorough assessment of the tool's effectiveness and significance.

4. Inconsistencies and Confusing Statements:

- The reference to "Postgraduate program in Biotechnology, Federal University of Pará, Belém, Pará, Brazil" as the institution appears unusual and should be clarified.

R= Thank you for your observation, and we understand the need to clarify the reference to the "Postgraduate program in Biotechnology, Federal University of Pará, Belém, Pará, Brazil" in our manuscript. We recognize that this reference may appear unusual, and we agree to remove the citation to avoid confusion.

- Several sentences throughout the manuscript are confusing and should be rephrased for clarity such as lines 31-34 and 54-57.

R= We appreciate your feedback, and we have made several improvements to the text based on your suggestions to enhance clarity.

- The use of "Ct (threshold cycle)" is inconsistent with MIQE guidelines, which recommend "Cq (quantification cycle)" values.

R= We have changed Ct to Cq as proposed to MIQE guidelines.

Bustin SA, Benes V, Garson JA, Hellemans J, Huggett J, Kubista M, Mueller R, Nolan T, Pfaffl MW, Shipley GL, Vandesompele J, Wittwer CT. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments. Clin Chem. 2009 Apr;55(4):611-22

- Address the ambiguity in the statement "ΔCt greater or less than 2" (Line 124) for better clarity.

R=Thank you for your observation. We would like to inform you that the statement "(ΔCt greater or less than 2)" on Line 124 has been corrected to "(∆Cq |2| standard deviations)" for improved clarity.

- Correct the statement "(p-value in reference genes between the studied groups or conditions" (Line 137) for clarity.

R=Thank you for your observation. We would like to inform you that the statement "(p-value in reference genes between the studied groups or conditions)" on Line 137 has been corrected to "(p-value in reference genes between the studied groups or conditions)" for improved clarity. The necessary modifications have been made in the revised version of the manuscript.

5. References:

- Ensure that references are correctly linked to their respective tools (line 60), and provide an accurate reference list.

R=Thank you for your feedback. We have made the suggested modification by citing the reference directly after each algorithm, as opposed to grouping all references at the end of the sentence.

- Correct punctuation and grammar in the text, including missing commas.

6. Figures:

- Figures 2, 3, 4, and 5 do not correspond with the current state of the tool, leading to confusion. Update these figures to accurately represent the tool.

R= Thank you for your suggestion. The figures 2, 3, 4 and 5 have been updated to reflect the current status of the tool.

- Figure 1 is misleading due to the use of dotted lines. Reconfigure the workflow diagram to eliminate confusion and better summarize the steps for users.

R= Thank you for your suggestion. The figure 1 have been updated to reflect the current status of the tool.

Reviewer #2: Comments on the app

7. Tool Inconsistencies:

- Address inconsistencies in the tool's interface and the manuscript regarding the naming of sections and options.

R= Thank you for your observation. We have carefully reviewed and corrected the inconsistencies pointed out between the tool's interface and the manuscript, specifically regarding the naming of sections and options.

- Resolve issues with the "Update Target Gene" button not functioning correctly.

R= If feasible, we kindly request the reviewer to provide more specific feedback. We thoroughly examined the "Update Target Gene" button and did not identify any malfunctions. However, its functionality may lack interaction as it does not provide feedback to the user regarding the selection of target genes (although this can be observed in the following sample data in "target" column). To address this, we have developed a comprehensive tutorial, which has been added to the Tutorial Section. This tutorial explains and illustrates the usage of the "Update Target Gene" button.

- Clarify the purpose and effects of the "Select one or more Genes" option in the tool.

R= This checkbox is used to remove Reference Genes from the analysis. To address this, we have developed a comprehensive tutorial, which has been added to the Tutorial Section. This tutorial explains and illustrates the usage of the "Select one or more Genes" checkbox.

- Address inconsistencies in the tool's interface and the manuscript regarding the naming of sections and options.

R= Thank you for your observation. We have carefully reviewed and corrected the inconsistencies pointed out between the tool's interface and the manuscript, specifically regarding the naming of sections and options.

- Resolve issues with the "Update Target Gene" button not functioning correctly.

R= If feasible, we kindly request the reviewer to provide more specific feedback. We thoroughly examined the "Update Target Gene" button and did not identify any malfunctions. However, its functionality may lack interaction as it does not provide feedback to the user regarding the selection of target genes (although this can be observed in the following sample data in "target" column). To address this, we have developed a comprehensive tutorial, which has been added to the Tutorial Section. This tutorial explains and illustrates the usage of the "Update Target Gene" button.

- Clarify the purpose and effects of the "Select one or more Genes" option in the tool.

R=We added the following text to the tool: "Gene Reference Descriptive Statistics

It is preferable for reference genes not to vary across different conditions or groups.

Additionally, it is desirable for reference genes to exhibit low variances,

enhancing the ability to discern small yet significant differences in the target.

If a reference gene is varying across group/condition and/or exhibits high variances and/or has low stability, consider removing it from the analysis.

- Make sure that the results in the NormFinder analysis are presented in a clear and understandable manner.

R= In the NormFinder section, we have included a brief introduction to the method: "NormFinder is a reliable method for evaluating the stability of reference genes by quantifying the variation in gene expression. The lower the stability value, the more stable the gene is considered."

Furthermore, we have developed a comprehensive tutorial that guides users on how to effectively utilize NormFinder and other analysis techniques to select reference genes more accurately. This tutorial will assist users in making informed decisions and improving the reliability of their gene expression studies.

- Add titles to all sections in the Differential Analysis tab for clarity.

R= We added the respective titles, as suggested. .

8. Data Saving and Reporting:

- Consider implementing a feature to save the results of all analyses.

R=Thank you for the suggestion. While we acknowledge that this could be a valuable feature, the current development and hosting structure of EndoGeneAnalyzer does not support the creation and management of user accounts. It is possible that we may explore implementing this feature in the future.

- Explore the possibility of generating downloadable reports for users, which can enhance the utility of the tool.

R= In the Differential Expression Analysis section, we've introduced a button that allows users to download a report with the analysis results. Thank you for the suggestion.

Attachment

Submitted filename: Response to Reviewers.docx

pone.0299993.s001.docx (19KB, docx)

Decision Letter 1

Karel Sedlar

21 Dec 2023

PONE-D-23-21180R1EndoGeneAnalyzer: a tool for selection and validation of reference genesPLOS ONE

Dear Dr. TEIXEIRA,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Feb 04 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-emailutm_source=authorlettersutm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Karel Sedlar, Ph.D.

Academic Editor

PLOS ONE

Journal Requirements:

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

Additional Editor Comments:

While you have already done substantial improvements to both, manuscript and your tool, there is still need for additional improvement as both reviewers raised numerous, yet only minor, concerns. Please do no prepare your revision in hurry. I am aware that the reviewing process is quite lengthy for your paper but the manuscript must meet publication criteria before it can be accepted. Please consider detailed proofreading of your manuscript and try to eliminate any other mistakes. The concerns that reviewers raised are mostly formal, so I believe you should be able to correct it quite easily.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: (No Response)

Reviewer #2: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: No

Reviewer #2: No

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: Well done to the authors for addressing my comments, I think the improvements to the app will make it much easier to use. My remaining comments are largely minor and can be easily addressed.

Comments (all line numbers relate to the tracked changes version of the MS)

Line 31 – change to ‘tool to assist’

Line 32 – lower case Q

Line 46 – remove The

Line 63 – only as it pertains to gene expression studies. RT-qPCR is used for many other purposes as well.

Lines 65-67 – This sentence should not be a paragraph, instead, make it the first sentence of the previous paragraph.

Lines 72-74 – cite Chapman and Waldenström for this point (stability of ref genes within study context)

Line 206 – remove second ‘the’ on this line

Some of the wording in the tutorial section of the app needs proof-reading (perhaps by a native English speaker) for clarity.

Have you considered having an option for Bonferroni (or other method) correction for multiple comparisons in the Differential Expression Analysis?

Reviewer #2: The revised manuscript and corresponding application have undergone significant improvements; however, some of my previous concerns were not fully addressed.

Comments on the manuscript:

I appreciate the creation of a comprehensive tutorial page for your tool. Nevertheless, some information is missing in the manuscript. The method for outliers removal, a previous concern of mine, is now clearly described in the tutorial page but not adequately addressed in the manuscript (lines 143-144). Additionally, the new functionality of restoring outliers is not mentioned in the manuscript.

While the authors provided a conclusion for their manuscript, the sentences are overly complicated, hindering understanding, and contain grammar mistakes. The conclusion should also mention the differential expression analysis as one of the additional features of this tool. Please delete the last sentence.

Regarding the grammar and flow of the text, I recommend having a native speaker review and edit the whole manuscript.

Comments on the figures:

Fig 1: Correct the name of third step from: "Gene Reference Sample" to "Gene Reference Samples".

Fig 2: Check the description of this figure in the manuscript. Change "Ct" to "Cq".

Fig 4: The values in the figure are currently presented only in black and white, despite the description referencing values highlighted in red. To ensure accuracy, update the figure to include the specified coloration. Additionally, align the numbering format consistently between the figure and its description (e.g., use 1, 2, 3 in both). Furthermore, incorporate the respective p-values associated with the color change into the figure description. This ensures that the significance of the color variation is appropriately communicated in the context of the data presented.

Other minor comments:

Lines 32 + 41: According to the MIQE guidelines (Nomenclature 1.1) RT-qPCR stands for "reverse transcription quantitative real-time polymerase chain reaction".

Line 55: Explain the meaning of Cq here.

Lines 58 – 61: Check the grammar and clarity of this sentence. Firstly, correct "data normalization process to expression comparison" to "data normalization process for expression comparison". Secondly, consider making the last addition a standalone sentence: "The normalization of data using reference genes involves correcting errors that arise from the initial concentration of RNA/cDNA. "

Line 67: Check references on this line. Either put respective reference after mentioned gene or put all references at the end of the sentence into single brackets.

Line 128: Delete comma in this sentence.

Line 150: Correct "Reference of genes by group" to "Gene Reference by group" to be consistent with the tool.

Lines 163 + 165: Check equations, especially the summation symbols and remove word "group" from the summation symbol, it is redundant.

Line 164: Remove this line.

Line 200: Change "tutorial" to "Tutorial".

Line 226: Change "meanRef" to "MeanRef".

Line 269: Check references on this line. Put all references at the end of the sentence into single brackets.

Comments on the App:

Regarding the "Update Target Gene": I was testing your tool and I have selected several genes and pressed the "Update Target Gene" button. After that I tried to unselect any targeted genes and press this button again, but nothing has changed and targets stayed the same. I do not think this a desirable feature.

Regarding this concern "Clarify the purpose and effects of the "Select one or more Genes" option in the tool.": Although the tutorial and the tool provide an explanation now, I would recommend to change the title "Select one or more Genes" to "Selected Reference Genes" to be more intuitive.

On the "Gene Reference Analysis" tab change all occurrences of "Normfinder" to "NormFinder".

On the "Tutorial" tab:

Point 1: Change "meanCT" to "mean Cq".

Point 2,3 and 5: Corresponding figures have on y-axis "meanCT", change it to "meanCq".

Point 5.2: Here, you mention "2^-ΔCq" and in the manuscript "2-ΔΔCq" formula.

Please correct the Tutorial page for consistency and check grammar.

Finally, let me make two more suggestions for the tool.

Regarding the "Gene Reference Samples" tab: I would recommend to add a short description to outlier removal methods, so user can read the description here and decide which method to use. I really like the added description for the "Gene Reference Analysis" tab, which made the analysis easier. I believe that this tab should be treated in the same way and this addition would make it more user-friendly.

Create a new tab in the tool called "About". There, provide your contact information for the users as you have also added a warning to the upload page saying "An error has occurred. Check your logs or contact the app author for clarification." but contact information is missing. I would also recommend to provide a link to GitHub page with your code and keep your GitHub page with codes updated. Users can leave feedback there, if they run into some problems, or if they request a new feature to be added.

In summary, commendable progress has been made in refining the manuscript and its corresponding application; however, there are still areas that require attention. Focusing on specific concerns, particularly in enhancing clarity, consistency, and functionality, will undoubtedly contribute to the continuous improvement and effectiveness of the tool. The commitment to elevate the overall quality and user experience of both the manuscript and the application is evident, and I look forward to seeing the implementation of these valuable improvements for the benefit of users and the advancement of your tool within the scientific community.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachment

Submitted filename: Review of EndoGeneAnalyzer R1.docx

pone.0299993.s002.docx (12.3KB, docx)
PLoS One. 2024 Apr 3;19(4):e0299993. doi: 10.1371/journal.pone.0299993.r004

Author response to Decision Letter 1


15 Feb 2024

RESPONSE LETTER TO THE REVIEWERS

Dear reviewer, my co-authors and I would like to thank you for the suggestions made during this high-quality review and then we present the answer to the questions.

We inform that with the reviews and suggestions, we were able to improve the idea presented by our work and we appreciate the opportunity. We hope this review has left the article suitable for publication in this high-impact and prestigious journal.

Kind Regards.

Reviewer #1: Well done to the authors for addressing my comments, I think the improvements to the app will make it much easier to use. My remaining comments are largely minor and can be easily addressed.

Comments (all line numbers relate to the tracked changes version of the MS)

Line 31 – change to ‘tool to assist’

R=Thank you for your suggestion! The change to 'tool to assist' has been implemented.

Line 32 – lower case Q

R=Thank you for your suggestion! The change to lower case 'Q' on Line 32 has been implemented

Line 46 – remove The

R= Thank you for your suggestion! We remove ‘The’ on Line 46

Line 63 – only as it pertains to gene expression studies. RT-qPCR is used for many other purposes as well.

R= Thanks for your suggestion! We have clarified the sentence in Line 63 to explicitly state that in the study RT-qPCR refers only to gene expression studies.

Lines 65-67 – This sentence should not be a paragraph, instead, make it the first sentence of the previous paragraph.

R=Thank you for your input! We have reorganized the text as per your suggestion. The content from Lines 65-67 is now the first sentence of the preceding paragraph.

Lines 72-74 – cite Chapman and Waldenström for this point (stability of ref genes within study context)

R=Thank you for your suggestion! We cite Chapman and Waldenström on lines 72-74

Line 206 – remove second ‘the’ on this line

R= Thank you for your suggestion! We remove second ‘the’ on Line 206

Some of the wording in the tutorial section of the app needs proof-reading (perhaps by a native English speaker) for clarity.

R= Thank you for pointing it out, we reviewed the tutorial section wording.

Have you considered having an option for Bonferroni (or other method) correction for multiple comparisons in the Differential Expression Analysis?

R=After ANOVA, Tukey's post hoc test is used for multiple comparisons and already provides the adjusted p-value. Similarly, following Kruskal-Wallis, the Dunn post hoc test is applied for multiple comparisons, and the p-values are adjusted using the Benjamini-Hochberg method. We added this information at the tutorial section.

Reviewer #2: The revised manuscript and corresponding application have undergone significant improvements; however, some of my previous concerns were not fully addressed.

Comments on the manuscript:

I appreciate the creation of a comprehensive tutorial page for your tool. Nevertheless, some information is missing in the manuscript. The method for outliers removal, a previous concern of mine, is now clearly described in the tutorial page but not adequately addressed in the manuscript (lines 143-144). Additionally, the new functionality of restoring outliers is not mentioned in the manuscript.

R= Thank you for your feedback. We appreciate your acknowledgment of the comprehensive tutorial and note your concerns. We will enhance the manuscript by providing a more detailed explanation of the outliers removal method (lines 143-144) and include information about the new functionality for restoring outliers.

While the authors provided a conclusion for their manuscript, the sentences are overly complicated, hindering understanding, and contain grammar mistakes. The conclusion should also mention the differential expression analysis as one of the additional features of this tool. Please delete the last sentence.

Regarding the grammar and flow of the text, I recommend having a native speaker review and edit the whole manuscript.

R= We've made the changes as requested, and the manuscript has been sent for translation by native speakers. Regarding the conclusion, the sentences were overly complex, making it difficult to understand, and there were grammar mistakes. We've also added a mention of the differential expression analysis as one of the additional features of the tool. However, we've removed the last sentence.

Comments on the figures:

Fig 1: Correct the name of third step from: "Gene Reference Sample" to "Gene Reference Samples".

R= Thank you for your feedback. We have made the correction in Fig 1 by updating the name of the third step from 'Gene Reference Sample' to 'Gene Reference Samples.

Fig 2: Check the description of this figure in the manuscript. Change "Ct" to "Cq".

R= Thank you for your observation. We have revised the description of Fig 2 in the manuscript, changing 'Ct' to 'Cq' as per your suggestion."

Fig 4: The values in the figure are currently presented only in black and white, despite the description referencing values highlighted in red. To ensure accuracy, update the figure to include the specified coloration. Additionally, align the numbering format consistently between the figure and its description (e.g., use 1, 2, 3 in both). Furthermore, incorporate the respective p-values associated with the color change into the figure description. This ensures that the significance of the color variation is appropriately communicated in the context of the data presented.

R= Thank you for your observation. For Figure 4, we've updated the coloration to include the specified red highlighting to match the description accurately.

Other minor comments:

Lines 32 + 41: According to the MIQE guidelines (Nomenclature 1.1) RT-qPCR stands for "reverse transcription quantitative real-time polymerase chain reaction".

R= Thank you for your suggestion. We have implemented the change, and the text now correctly states: 'RT-qPCR stands for reverse transcription quantitative real-time polymerase chain reaction.

Line 55: Explain the meaning of Cq here.

R=Thank you for your feedback. We have addressed the request on Line 55 and provided an explanation for the meaning of Cq.

Lines 58 – 61: Check the grammar and clarity of this sentence. Firstly, correct "data normalization process to expression comparison" to "data normalization process for expression comparison". Secondly, consider making the last addition a standalone sentence: "The normalization of data using reference genes involves correcting errors that arise from the initial concentration of RNA/cDNA.

R= Thank you for your suggestions. We have addressed the concerns in Lines 58-61. Firstly, we corrected "data normalization process to expression comparison" to "data normalization process for expression comparison." Additionally, we separated the last addition into a standalone sentence, which now reads: 'The normalization of data using reference genes involves correcting errors that arise from the initial concentration of RNA/cDNA.'

Line 67: Check references on this line. Either put respective reference after mentioned gene or put all references at the end of the sentence into single brackets.

R= Thank you for your guidance. We have revised Line 67 by placing all references at the end of the sentence within single brackets.

Line 128: Delete comma in this sentence.

R= Thank you for your suggestion. The comma has been deleted in this sentence.

Line 150: Correct "Reference of genes by group" to "Gene Reference by group" to be consistent with the tool.

R=Thank you for your feedback. The change "Reference of genes by group" has been corrected to "Gene Reference by group" for consistency with the tool.

Lines 163 + 165: Check equations, especially the summation symbols and remove word "group" from the summation symbol, it is redundant.

R=Thank you for your feedback. Equations have been checked, and the word "group" has been removed from the summation symbol as it was redundant.

Line 164: Remove this line.

R=Thank you for your suggestion. This line has been removed.

Line 200: Change "tutorial" to "Tutorial".

R=Thank you for your suggestion. We change "tutorial" to "Tutorial."

Line 226: Change "meanRef" to "MeanRef".

R=Thank you for your suggestion. We change "meanRef" to "MeanRef."

Line 269: Check references on this line. Put all references at the end of the sentence into single brackets.

R= Thank you for your suggestion. References on this line have been checked, and all references at the end of the sentence are now in single brackets.

Comments on the App:

Regarding the "Update Target Gene": I was testing your tool and I have selected several genes and pressed the "Update Target Gene" button. After that I tried to unselect any targeted genes and press this button again, but nothing has changed and targets stayed the same. I do not think this a desirable feature.

R= Thank you for your careful review. We hadn't noticed this misbehavior. The issue has been fixed.

Regarding this concern "Clarify the purpose and effects of the "Select one or more Genes" option in the tool.": Although the tutorial and the tool provide an explanation now, I would recommend to change the title "Select one or more Genes" to "Selected Reference Genes" to be more intuitive.

R= Suggestion accepted. Thank you.

On the "Gene Reference Analysis" tab change all occurrences of "Normfinder" to "NormFinder".

R= Suggestion accepted. Thank you.

On the "Tutorial" tab:

Point 1: Change "meanCT" to "mean Cq".

R= Thank you for your careful review.Suggestion accepted.

Point 2,3 and 5: Corresponding figures have on y-axis "meanCT", change it to "meanCq".

R= Thank you for your careful review.Suggestion accepted.

Point 5.2: Here, you mention "2^-ΔCq" and in the manuscript "2-ΔΔCq" formula.

R= The term 2^-ΔCq is employed for estimating expression levels, while 2^-ΔΔCq is utilized for estimating fold-change. In this section, where we elucidate the estimation of expression levels, the formula was applied within the appropriate context.

Please correct the Tutorial page for consistency and check grammar.

R=Thank you for your careful evaluation. Suggestion accepted. The tool page and manuscript have been revised.

Finally, let me make two more suggestions for the tool.

Regarding the "Gene Reference Samples" tab: I would recommend to add a short description to outlier removal methods, so user can read the description here and decide which method to use. I really like the added description for the "Gene Reference Analysis" tab, which made the analysis easier. I believe that this tab should be treated in the same way and this addition would make it more user-friendly.

R= Thank you for your careful review.Suggestion accepted.

Create a new tab in the tool called "About". There, provide your contact information for the users as you have also added a warning to the upload page saying "An error has occurred. Check your logs or contact the app author for clarification." but contact information is missing. I would also recommend to provide a link to GitHub page with your code and keep your GitHub page with codes updated. Users can leave feedback there, if they run into some problems, or if they request a new feature to be added.

R=We have implemented the "About" tab as recommended, incorporating a concise tool description, contact and support details, GitHub link, information about participants and funding partners. Additionally, upon acceptance, we will include the publication link.

In summary, commendable progress has been made in refining the manuscript and its corresponding application; however, there are still areas that require attention. Focusing on specific concerns, particularly in enhancing clarity, consistency, and functionality, will undoubtedly contribute to the continuous improvement and effectiveness of the tool. The commitment to elevate the overall quality and user experience of both the manuscript and the application is evident, and I look forward to seeing the implementation of these valuable improvements for the benefit of users and the advancement of your tool within the scientific community.

Attachment

Submitted filename: Response to Reviewers.docx

pone.0299993.s003.docx (14.1KB, docx)

Decision Letter 2

Karel Sedlar

20 Feb 2024

EndoGeneAnalyzer: a tool for selection and validation of reference genes

PONE-D-23-21180R2

Dear Dr. TEIXEIRA,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Karel Sedlar, Ph.D.

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Congratulations! I believe your contribution will be of great use for the community. Special thanks also belongs to reviewers. I am once againg sorry it took so long to get them but it was not easy to get specialists in your specific field. Nevertheless, their thorough reviews helped to improve your manuscript and tool a lot.

Reviewers' comments:

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    Attachment

    Submitted filename: Response to Reviewers.docx

    pone.0299993.s001.docx (19KB, docx)
    Attachment

    Submitted filename: Review of EndoGeneAnalyzer R1.docx

    pone.0299993.s002.docx (12.3KB, docx)
    Attachment

    Submitted filename: Response to Reviewers.docx

    pone.0299993.s003.docx (14.1KB, docx)

    Data Availability Statement

    This study presents the EndoGeneAnalyzer tool, available at https://npobioinfo.shinyapps.io/endogeneanalyzer/ and the open-source code can be found at https://github.com/MoreiraFC/EndoGeneAnalyzer.


    Articles from PLOS ONE are provided here courtesy of PLOS

    RESOURCES