Skip to main content
. Author manuscript; available in PMC: 2012 Apr 1.
Published in final edited form as: Genet Epidemiol. 2011 Jan 31;35(3):159–173. doi: 10.1002/gepi.20564

Table 1. Phenotype harmonization and analysis plan check list.

Items may be addressed concurrently and not necessarily in this order.

Establish a group and group leader
  • Identify a leader interested and willing to drive the process forward

  • Include representatives from all studies and groups contributing data and any other groups (such as a coordinating center) contributing to the work

Identify common phenotypes of interest (including covariates)
  • Review available data—what was collected and what can be shared

    • Review data collection forms and questionnaires

    • Create a spreadsheet of phenotypes showing which studies collect which variables

    • Create a web-based or other survey, if appropriate

Determine feasibility of cross-study analyses
  • Identify participating studies

    • Are the subjects representative of the cohort

  • Identify approximate numbers of subjects to be included in cross-study analyses

    • Is the combined total number of subjects sufficient to detect a significant association

  • Review consent status of each contributing study

    • Do planned analyses fall within the scope of the data use statements of the consent forms used by each study

  • Review each study's inclusion and exclusion criteria

    • Are there studies for which the inclusion/exclusion criteria make their inclusion in the cross-study analysis not useful

  • Review data definitions from each contributing study

    • What exactly was asked (or measured)—compare the wording of the actual questions

    • Are response options comparable

  • Compare data distributions

    • If data distributions are not similar, i.e. if there is little overlap, the data may not be comparable

  • Confirm that the data are not being used for similar analyses by another collaboration or consortium

Prepare common definitions
  • Prepare phenotype definition

    • Define the outcome of interest and values

    • Define the outcome type—whether discrete, ordinal, continuous

Code variables
  • Create algorithms for converting each study's raw data to conform to the agreed upon definitions

Draft an analysis plan
  • Describe the phenotype, covariates, and inclusion/exclusion criteria

  • List participating studies and anticipated number of subjects contributing to the analyses

  • Describe the planned subgroup analyses

  • Describe the analysis approach

    • How will outliers be handled

      • Will outliers be Windsorized (i.e. equated to the next highest/lowest value)

      • Will extreme values be truncated (e.g. to 4 standard deviations of sample mean)

    • How will missing values be handled

    • How will covariates be adjusted for

    • How will data from longitudinal studies (studies with data from multiple time points) be combined with data from cross-sectional studies

      • Will a date range be chosen for which data values to include

    • Determine which statistical model fits best; often several models are needed

    • Identify type of analysis, i.e. meta-analyses of summary statistics or analyses of pooled individual data

    • Describe the statistical support required and who will provide it

    • Describe the imputation plan

      • Who will be responsible for imputation, i.e. will this be done by individual studies or will it be done centrally

    • Identify any candidate genes that can be examined

    • Identify key personnel and their roles

      • Who at each site and/or the CC will be doing analyses

      • Who will be the primary author

      • Will any assistance be required from the CC

      • What will be the role of the CC

    • Develop a timeline—genotype and phenotype data for different studies may not be ready at the same time

      • When can analyses start and what studies will be included

      • How will studies whose data will be ready later be incorporated into the analysis

    • Relationship to work being done in other consortia/collaborations

      • How do these analyses fit in with what other consortia are doing

      • Can collaborations be established with another consortium to combine data for a primary analysis

      • Can the other consortium's data serve as a replication study

      • What specific associations and G*E analyses can be performed that have not been already examined