Skip to main content
BMJ Open Access logoLink to BMJ Open Access
. 2025 Oct 23;80(2):e224561. doi: 10.1136/jech-2025-224561

oriGen cohort: a Mexican population-based epidemiological and genomic research platform

Pablo Kuri-Morales 1, Rocio Ortiz-Lopez 1,2,3, Elena-Cristina Gonzalez-Castillo 1,3, Nestor Rubio-Infante 1, Jose Ramírez-Vega 1, Rocio-Alejandra Chavez-Santoscoy 1,4, Cuitláhuac Ruiz-Matus 1, Martin De-La-Cruz 5, Israel Aguilar-Ordoñez 1, Gerardo Garcia-Rivas 3,6, Servando Cardona 2, Miguel Betancourt-Cravioto 7, Victor Trevino 1,2,3,*, Guillermo Torre-Amione 1,6,
PMCID: PMC12911581  PMID: 41136196

Abstract

Background

Genomic information is transforming public health and personalised medicine by identifying genetic and environmental contributors to disease. However, Hispanic populations remain under-represented in global genomic datasets, limiting the relevance of findings for groups such as the Mexican population. The oriGen project, launched by Tecnológico de Monterrey, aims to address this disparity by establishing a nationally representative cohort of 100 000 Mexican adults. The project integrates clinical, lifestyle and genomic data to enable the study of gene-environment interactions, disease susceptibility and health disparities in Mexico and Latin America.

Methods

oriGen is a prospective, population-based biobank recruiting participants from 18 metropolitan areas across 19 Mexican states using probabilistic, stratified, multistage sampling based on national statistical frameworks. Data collection includes electronic questionnaires, anthropometric measurements, vital signs and biological samples for biochemical analysis and genomic sequencing. Whole-genome and whole-exome sequencing are conducted in collaboration with national and international partners.

Results

As of April 2025, 83 764 individuals (61% female) have been enrolled. The cohort slightly over-represents older adults and women. Baseline data show a high prevalence of obesity (40%), elevated blood pressure (mean 127.4/81.7 mm Hg) and elevated blood glucose (mean 133.3 mg/dL). Initial genomic analyses (n=1318) indicate an average admixture of 61.8% Native American, 32.4% European and 5.1% African ancestry, with regional variation.

Conclusion

oriGen represents one of the most comprehensive genomic epidemiology efforts in Mexico, offering a valuable resource for advancing equitable precision medicine and public health research in under-represented populations.

Keywords: population genomics, biobank, Mexico, genomic epidemiology, personalized medicine, health disparities, gene-environment interaction, admixture, prospective cohort, Latin America


WHAT IS ALREADY KNOWN ON THIS TOPIC

  • Global genomic datasets have historically lacked ancestral and demographic diversity, with less than 1% of participants classified as Hispanic.

  • This gap compromises the generalisability of findings and perpetuates inequities in biomedical research and precision health, particularly for Latin American populations.

WHAT THIS STUDY ADDS

  • This study demonstrates the feasibility and design of a large-scale, nationally representative genomic and epidemiological cohort in Mexico, integrating rigorous sampling, deep phenotyping and state-of-the-art genomic analysis.

  • It also highlights the methodological strategies for inclusive recruitment and initial evidence of regional genetic diversity within an admixed population.

HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY

  • By generating high-quality, population-specific data, oriGen can support more accurate disease risk prediction and inform health interventions tailored to Mexico’s unique genetic and socio-environmental landscape.

  • The cohort also sets a precedent for ethically grounded, collaborative data sharing in low- and middle-income country settings, which may influence national research infrastructure and health equity policy.

Introduction

Genomic information plays a fundamental role in the development of strategies for disease prevention and control at the population level, as it allows for the identification of individuals more susceptible to developing specific conditions by designing effective and personalised public health interventions. It also provides the necessary data to develop more effective personalised treatments with fewer side effects.

To create these personalised medicine tools and improve the impact of public health programmes, it is essential to have specific phenotypic and genotypic information for individuals in different populations. With this information, it is possible to identify the most effective preventive and therapeutic strategies, which can significantly improve health outcomes.

Since the completion of the Human Genome Project in 2003, a number of important international efforts to help understand the relationship between genetic variations and health-disease processes have been carried out. These projects have allowed for a greater understanding of the underpinnings of a wide range of chronic diseases such as cancer and heart disease, leading to relevant innovations in the prevention, diagnosis and treatment of these health issues.1

Nevertheless, most large-scale population-based genomic studies to date have focused on European populations with little participation of other ethnicities such as African, Asian and Hispanic populations. It has been estimated that of the total genomes included in large-scale genomic studies, approximately only 1% corresponds to individuals defined as Hispanic.2 This has led to the use of extrapolations from models based on the genomics of European populations to predict diseases and phenotypes in Hispanic peoples such as the Mexican population, which is inadequate and inaccurate given the ancestral diversity and admixture of the country’s population.3 Additionally, these population biases in large genomic studies hinder the adequate identification of the relationship between environmental risk factors and their interaction with the population’s genetics, as well as their influence on health-disease states.4

To contribute to reducing the imbalance in the ancestry of large genomic projects, the Instituto Tecnológico y de Estudios Superiores de Monterrey (Tecnológico de Monterrey) launched oriGen, a prospective cohort of 100 000 Mexicans, men and women, 18 years and older, with national-scale representativeness.

Tecnológico de Monterrey began recruiting participants in January 2023 with the aim of generating information to better understand how genetic variants, lifestyle, and nutritional and familial factors contribute to the health and disease of the Mexican population. oriGen’s clinical and genomic information will be made available to researchers worldwide, in order to promote research and knowledge for the benefit of the population. The purpose of this paper is to present a general description of the oriGen project and cohort profile and to describe the most relevant findings of the cohort at baseline.

Cohort description

oriGen is designed as a nationally representative longitudinal survey and biobank of the adult Mexican population including assessments of health status and frequency of relevant risk factors and socioeconomic determinants of health, as well as anthropometry, biochemical markers and genomic data. The cohort is being established across 18 metropolitan areas in 19 states. The metropolitan areas were selected by convenience, focusing on those with Tecnológico de Monterrey campuses and the necessary facilities to support laboratory operations. Nevertheless, selected metropolitan areas are representative of the northern, central and southern regions of Mexico.

To mitigate the selection bias derived from the convenience sampling, enrolment of individuals in each metropolitan area is done using a probabilistic, multi-staged, conglomerated, stratified design based on the National Institute for Statistics, Geography and Informatics of Mexico’s (INEGI)5 basic geo-statistical areas (AGEBs—áreas geo-estadísticas básicas). AGEBs are the smallest statistical-geographical unit used by INEGI for organising censuses in Mexico. Each AGEB has at least 2500 inhabitants in urban settings and represents relative homogeneous areas in terms of land use, population density and morphology. AGEBs are the basis of the sampling framework used for the Mexican National Health and Nutrition Surveys (ENSANUT).6

To date, recruiting in 14 metropolitan areas has been completed (78% of all designated areas) having enrolled 83 764 individuals into the cohort. This includes the selected boroughs and municipalities of the Mexico City Metropolitan Area which is ongoing. Recruiting in the metropolitan areas of Veracruz, Merida and Tuxtla Gutierrez will be completed by mid-2025 when the cohort goal of 100 000 participants will be achieved (figure 1).

Figure 1. Selected metropolitan areas and states represented as of April 2025.

Figure 1

Recruitment begins with a local media awareness campaign, after which all households within a selected AGEB are visited by the data and sample collection teams. Each of the three teams is made up of a field supervisor, four interviewers and two nurses for somatometry and blood sample collection. A coordinator harmonises three supervisors, two transporters and two data managers. To ensure homogeneity in data collection procedures, all members of the Data Sample Collection Teams are continuously trained on an overview of the project, data-flow procedures, quality control steps, use of hardware and software, ethics in human research and informed consent. Also, interviewers are trained on household selection methods and on the specifics of the electronic questionnaire. Finally, nurses are trained in standardisation of vital signs measurement, somatometry and blood sampling and handling.

All members aged 18 or older in each household are invited to participate in the study and given a short description of the project, its objectives and the expected participation from each participant. Individuals who accept to be enrolled are then presented with an information sheet and the Informed Consent Form for signing.

To avoid high levels of relatedness in the sample, no more than 10% of related individuals within each household were enrolled. The overall household response rate was 30%, with an average recruitment of slightly more than one participant per household (mean=1.13).

Data collection procedure

Individuals who accept to participate in oriGen are enrolled through face-to-face interviews. Each interview includes an electronic questionnaire (performed by an interviewer), anthropometry and blood sampling (performed by nurses) and lasts around 75 min.

Electronic questionnaire

The questionnaire was specifically designed for oriGen and included validated items from the UK Biobank7 and ENSANUT8 besides items developed specifically for the project. Each participant responds to a minimum of 470 questions for male participants or 550 for females, and a maximum of 750 questions depending on family and personal history of illnesses (box 1, Section 1.1). Answers are collected by trained interviewers on electronic tablets using the RedCap (Research Electronic Data Capture) application. The questionnaire is included as online supplemental file 1.

Box 1. Sections included in the electronic questionnaire and physical evaluation measurements.

1.1 Questionnaire sections
  1. General information and identification

  2. Work, occupation and education

  3. Medical history—Common diseases (type 2 diabetes mellitus, hypertension, cardiovascular disease, chronic kidney disease, dyslipidaemia, overweight and obesity)

  4. Medical history—Other diseases (cognitive and neurological, psychiatric, pulmonary and respiratory, gastrointestinal, infectious, haematologic, rheumatologic, immunologic, endocrine, genitourinary, skin diseases and cancer)

  5. Mental health (mood perception, sleep, physical and personal self-perception)

  6. Hearing and motor health

  7. Dental health

  8. Risk factors: Blood transfusions, environmental risks, smoking and alcohol consumption

  9. Reproductive health and sexuality

  10. Fractures and pain

  11. Nutrition

  12. Physical activity

  13. Immunisations

  14. Common medications and supplements

  15. Hospitalisations

  16. Family medical history

  17. Household characteristics

1.2 Physical evaluation measurements
  1. Standing height

  2. Waist circumference

  3. Weight

  4. Resting pulse rate

  5. Sitting blood pressure

  6. Body fat mass

  7. Skeletal muscle mass

  8. Body mass index

  9. Percent body fat

  10. Basal metabolic rate

  11. Waist-hip ratio

  12. Visceral fat level

  13. Obesity

  14. Impedances

  15. Fat-free mass

  16. Body fat mass

Vital sign measurement and anthropometric data

In addition to completing the questionnaire, each participant undergoes measurement of resting heart rate and seated blood pressure, anthropometric evaluation and body composition analysis using bioelectrical impedance analysis devices (InBody 120, InBody, Seoul, South Korea) (box 1, Section 1.2). All measurements are carried out by trained health professionals.

Biological samples

All participants undergo capillary blood sampling for the measurement of glucose, cholesterol and triglyceride levels. Additionally, two venous blood samples are obtained from each subject. The first is a 4.0 mL tube with EDTA (BD Vacutainer EDTA tubes, 366643) and the second is an 8.5 mL tube without an anticoagulant (BD Vacutainer blood collection red tubes with clot activator, 367988). Both tubes are sent to the local laboratory for storage or processing.

The tube containing the EDTA anticoagulant is stored at 4°C. The tube without anticoagulant is centrifuged to obtain serum, and two 1.5 mL aliquots are prepared and stored at −20°C. The blood tubes and serum aliquots are stored at their respective temperatures and sent weekly to the central laboratory.

Once in the central laboratory, the serum samples are stored at −20°C and the sample with EDTA is also centrifuged to separate plasma (1.5 mL cryovial) and buffy coat (200 µL for DNA extraction and 500 µL buffy coat). All samples are labelled with a unique identifier generated at enrolment and stored at −80°C at the oriGen Biorepository at the Tec-Salud Zambrano-Hellion Hospital in Monterrey, Mexico.

Genomic tests

As a pilot test, whole-genome sequencing (WGS) was performed on 1481 DNA samples using a NovaSeq 6000 platform at the Tecnológico de Monterrey Genomics Core Laboratory in Monterrey, Mexico. Raw sequencing data were processed with a computational pipeline analogous to that employed by the UK Biobank,9 which is shown in online supplemental figure 1 including specific programmes and critical parameters. Briefly, from raw sequencing data, we filter, align and merge reads. Then, verify quality, perform variant calling and finally estimate kinship and admixtures.

An agreement was established with the Regeneron Genetics Centre (RGC) to sequence the entire cohort, primarily through whole-exome sequencing, with approximately 10% undergoing WGS to support genotype imputation.10

Quality control

To ensure the quality of both data and biological samples, several quality control checkpoints have been established throughout the entire process, including field, sample collection, field laboratory processing, transport, storage, DNA extraction, sequencing, data verification and data completeness, among others. For more details, see online supplemental table 1.

Data storage

All data items and records from each participant (electronic questionnaire, vital signs, anthropometry, results of biochemical and genomic tests) are stored in a specific data repository setup for the project based on national and international best practices and regulations regarding storage of personal data and clinical samples. These practices include multiple copies of databases, encryption, controlled access, physical and cloud-based repositories, and capacity to expand according to need.

oriGen is planned as a long-term cohort and therefore it is expected that data will be maintained and updated as necessary for the foreseeable future. However, according to the informed consent, participants can opt out of the study at any time.

Access to data by collaborators and external researchers will be subject to limits established in specific collaboration agreements as per oriGen’s Data Sharing Policy included in online supplemental file 2.

Baseline findings

Of the 83 764 individuals enrolled by April 2025 in oriGen, 39% (32 769) are male, and 61% (50 995) are female.

Table 1 presents the baseline socioeconomical, physical, lifestyle and biomedical characteristics of the participants currently enrolled in the cohort. All collected measurements are reported.

Table 1. Baseline characteristics of the participants in the oriGen cohort.

Total Female Male
Participants 83 764 50 995 32 769
Age, years, mean 50.61 50.39 50.96
Distribution by age group, n (%)
 <35 14 940 (17.84) 8745 (17.15) 6195 (18.91)
 35–54 30 436 (36.34) 19 364 (37.97) 11 072 (33.79)
 55–74 32 103 (38.33) 19 465 (38.17) 12 638 (38.57)
 >=75 6299 (7.52) 3428 (6.72) 2871 (8.76)
Education level, n (%)
 Without formal education 3671 (4.38) 2376 (4.66) 1295 (3.95)
 Primary school 20 745 (24.77) 13 277 (26.04) 7468 (22.79)
 Secondary school 41 790 (49.89) 24 687 (48.41) 17 103 (52.19)
 Technical education 5920 (7.07) 4393 (8.61) 1527 (4.66)
 University and above 11 637 (13.89) 6262 (12.28) 5375 (16.4)
Occupation, n (%)
 Unemployed/not economically active 45 747 (54.61) 33 425 (65.55) 12 322 (38.6)
 Managers 1968 (2.35) 1193 (2.34) 775 (2.37)
 Skilled agricultural, forestry and fishery workers 322 (0.38) 50 (0.1) 272 (0.83)
 Craft and related trades workers 3064 (3.66) 593 (1.16) 2471 (7.54)
 Sales workers 13 082 (15.62) 7665 (15.03) 5417 (16.53)
 Public administration and armed forces personnel 289 (0.35) 133 (0.26) 156 (0.48)
 Plant and machine operators, and assemblers/construction workers 4966 (5.93) 884 (1.73) 4082 (12.46)
 Professionals 2751 (3.28) 1577 (3.09) 1174 (3.58)
 Technicians and associate professionals 2723 (3.25) 1508 (2.96) 1215 (3.71)
 Service and sales workers—personal and domestic services 4240 (5.06) 3000 (5.88) 1240 (3.78)
 Technicians and associate professionals 840 (1) 367 (0.72) 473 (1.44)
 Drivers and mobile plant operators/logistics workers 1736 (2.07) 129 (0.25) 1607 (4.9)
 Occupations not elsewhere classified 2035 (2.43) 764 (1.5) 1271 (3.88)
Physical examination
 Systolic blood pressure, mm Hg, mean (SD) 127.37 (21.99) 125.65 (21.21) 131.41 (19.68)
 Diastolic blood pressure, mm Hg, mean (SD) 81.73 (12.69) 79.44 (12.06) 82.35 (12.21)
 Body mass index, kg/m2, mean (SD) 29.61 (9.95) 29.54 (9.13) 28.22 (8.64)
 Obesity (BMI ≥30 kg/m2), % 39.95 40.20 30.07
 Waist circumference, cm, mean (SD) 97.06 (16.91) 95.25 (17.12) 97.71 (16.84)
 Waist:hip* ratio, mean (SD) 0.96 (0.09) 0.96 (0.10) 0.94 (0.08)
Biochemical assessment
 Blood glucose, mg/dL, mean (SD) 126.38 (56.41) 124.58 (55.96) 129.33 (57.02)
 Blood triglycerides, mg/dL, mean (SD) 183.64 (79.26) 180.67 (77.56) 188.49 (81.73)
 Blood cholesterol, mg/dL, mean (SD) 182.57 (49.30) 187.64 (50.35) 174.30 (46.35)
 Blood HDL, mg/dL, mean (SD) 51.31 (14.40) 53.83 (14.59) 47.20 (13.09)
 Blood LDL, mg/dL, mean (SD) 95.10 (40.32) 98.20 (40.95) 90.03 (38.72)
Lifestyle
Smoking
 Current smoker (%) 13.92 11.74 16.09
 Ex-smoker (%) 9.33 7.23 12.26
 Never smoker (%) 70.47 80.58 56.50
 Vaping (%) 1.2 0.97 1.85
Alcohol drinking
 Never drank (%) 80.21 87.79 71.11
 Up to once a month (%) 10.64 8.02 13.31
 Two to four a month (%) 5.7 3.03 9.12
 Two or three a week (%) 1.94 0.74 3.55
 Four or more a week (%) 1.28 0.28 2.56
Dietary habits
 Meat, days/week, mean (SD) 2.31 (1.59) 2.10 (1.47) 2.47 (1.64)
 Vegetables, days/week, mean (SD) 4.37 (2.27) 4.51 (2.26) 4.07 (2.27)
 Fruits, days/week, mean (SD) 4.35 (2.35) 4.53 (2.33) 4.05 (2.33)
 Legumes, days/week, mean (SD) 3.88 (2.33) 3.82 (2.34) 4.01 (2.28)
 Corn cereals, days/week, mean (SD) 6.12 (1.77) 6.06 (1.81) 6.15 (1.71)
 Wheat cereals, days/week, mean (SD) 3.53 (2.54) 3.37 (2.53) 3.71 (2.5)
 Other cereals, days/week, mean (SD) 1.70 (1.99) 1.79 (2.03) 1.55 (1.91)
Physical activity
 Vigorous, days, mean (SD) 4.73 (1.97) 4.73 (2.04) 4.56 (1.97)
 Moderated, days, mean (SD) 5.28 (1.86) 5.02 (1.91) 5.34 (1.88)
 Walking, days, mean (SD) 5.63 (1.75) 5.59 (1.76) 5.54 (1.77)
 TV, hours/day, mean (SD) 1.97 (1.76) 2.05 (1.72) 2.03 (1.77)

Data are presented as mean±SD for continuous variables or n (%) for categorical variables.

Only biochemical measurements within the CardioChek Plus device's detection limits were used to estimate statistics: glucose (20–600 mg/dL), triglycerides (50–500 mg/dL), cholesterol (100–400 mg/dL), and HDL (15–100 mg/dL). Estimated LDL values (LDL = cholesterol – HDL – (triglycerides/5)) were only considered when triglyceride levels were 400 mg/dL or less, consistent with manufacturer guidelines.

*

Waist:hip ratio is estimated by the bioimpedance equipment (INBODY).

BMI, body mass index; HDL, High-density lipoproteins; LDL, Low-density Lipoproteins; TV, Television.

There are notable differences in the sociodemographic composition of the oriGen cohort compared with the general Mexican adult population, as reflected in national statistics. Women account for 60.9% of the cohort, compared with approximately 51% in the general population. Adults aged 55 years and older represent 45.8% of participants, exceeding the national estimate of 29% reported in the 2020 census. Conversely, individuals under 35 years are less represented in the cohort (17.8%) compared with around 34% nationally.11 These variations are consistent with patterns observed in household-based surveys, where women and older adults are more likely to be present and willing to participate. As such, these demographic characteristics should be taken into account when analysing and interpreting age-dependent and sex-dependent health outcomes, using bias-mitigation strategies such as post-stratification weighting, rate adjustment and sensitivity analyses.

The educational and occupational profiles of the oriGen cohort reflect a population with relatively higher formal education and a distinct employment pattern compared with the country’s general adult population. 49.9% of participants reported having completed secondary education, and 13.9% reported completing university-level education or higher, figures that exceed national averages, particularly among older adults. The number of participants who reported no formal education (4.4%) or complete primary school (15.8%) is lower than expected based on national census data. Occupationally, the cohort presents a high proportion of individuals who are unemployed or not economically active (54.9%), especially among women (65.6%), which is notably above national estimates. compared with national labour statistics, managerial and professional roles are slightly underrepresented in the cohort.11

Participants in the oriGen cohort have an average body mass index (BMI) of 29.6 kg/m², with 40% of participants classified as obese (BMI ≥30), a figure slightly higher than the 2022 ENSANUT national obesity prevalence of ~36% among adults. Mean systolic and diastolic blood pressures (127.4/81.7 mm Hg) are also slightly higher than national estimates. The mean random (non-fasting) blood glucose levels in the oriGen cohort were 126.4 mg/dL (124.6 mg/dL in women and 129.3 mg/dL in men), a value that falls below the diagnostic threshold for diabetes based on casual glucose testing (≥200 mg/dL) but may still reflect underlying metabolic risk.12

Table 2 shows the prevalence of self-reported history of chronic diseases by cohort participants.

Table 2. Self-reported history of chronic diseases.

Total Female Male
Diabetes (%) 20.8 (40.6) 21.8 19.17
Ischaemic heart disease (%) 0.1 (3.2) 0.1 0.1
Stroke (%) 2.0 (13.9) 1.8 2.9
Cancer (any site) (%) 2.5 (15.7) 3.27 1.4

The self-reported prevalence of diabetes among cohort participants is higher than the national prevalence of 12.8% reported in ENSANUT 2022 for adults aged ≥20 years. The same is observed for self-reported history of stroke and cancer. On the other hand, the self-reported history of ischaemic heart disease is notably lower than national estimates.12

Genomic ancestry

Of the 1481 samples sequenced, 1318 unrelated individuals were retained after quality control for ancestry inference. Briefly, after removing low quality samples, sex mismatches and up to third-degree relationships excluded based on an estimated genetic relatedness threshold of ~7% (see online supplemental figure 1 for details). Ancestry estimation was performed using ADMIXTURE (V.1.3.0) with the number of ancestral populations (k) set to 5, incorporating reference samples from the 1000 Genomes Project. All individuals were recruited from the Monterrey metropolitan area in northeastern Mexico. On average, we observed 61.8% Native American, 32.4% European and 5.1% African ancestry (figure 2A), consistent with previous reports of population admixture patterns in the region.13 14

Figure 2. Estimated ancestry from whole genomes. (A) The overall ancestry of the 1481 samples of all samples taken from Nuevo León (delimited by dashed lines). (B) The ancestry estimated from self-reported born sites stratified by regions in colours. (C) Average regional ancestry estimation of the Native American (AMR) component. The asterisk (*) marks regions poorly estimated due to a low number of samples.

Figure 2

Although all samples were collected in the Monterrey metropolitan area, we stratified ancestry estimates based on participants’ self-reported responses to the question ‘Where were you born?’. Geographical regions were defined according to the seven-region framework established by the Instituto Nacional de Salud Pública (INSP),6 enabling comparability with prior studies. Overall, individuals originating from southern regions exhibited higher proportions of Native American ancestry, while those from northern regions showed greater European ancestry (figure 2). These patterns are consistent with previous genetic studies conducted in non-metropolitan populations across Mexico.14

Perspectives

Follow-up and data updating

Participants in the oriGen cohort will be followed up longitudinally to assess changes in their health status, prevalence of risk factors and incidence of chronic diseases over time. Once the recruitment phase is complete (Q3 2025), follow-up assessments will be carried out at regular intervals yet to be defined, to incorporate updated clinical, biochemical and lifestyle data, as well as linkages to mortality records when appropriate. Follow-ups will be done through electronic and mobile messaging (eg, WhatsApp).

Future waves of data collection will allow for the examination of temporal patterns, causal inferences and the validation of predictive models developed from baseline findings.

Data availability and collaboration

The oriGen study has established a comprehensive data and sample-sharing policy aimed at maximising the scientific and public health value of its extensive biomedical repository. oriGen’s data sharing and collaboration policy is grounded in the project’s commitment to advance global knowledge in disease prevention, diagnosis and treatment, particularly for health conditions with high prevalence in Mexico. The data repository is being designed as a national and international resource for qualified researchers, and the policy emphasises collaboration with both national and international institutions.

Data sharing will operate under clear ethical, legal and quality safeguards, adhering to international best practices such as those outlined by UK Research and Innovation (UKRI)15 and the University of Oxford’s Nuffield Department of Population Health policy for data sharing.16 Access to data will have two modalities: formal collaborative agreements between research organisations and oriGen, and via open-access data requests evaluated by a dedicated review committee. English versions of the data and sample sharing policy and forms are included as online supplemental file 2.

Strengths and limitations

The oriGen cohort represents a major advancement in population-based biomedical research in Mexico due to its size, national representativity and integration of genomic and epidemiological data. With a target enrolment of 100 000 individuals across 18 metropolitan areas, oriGen offers a broader geographical and sociodemographic representation than earlier efforts such as the Mexico City Prospective Study (MCPS),13 which was limited to two districts within a single metropolitan area. In contrast to MCPS’s focus on mortality and cardiometabolic risk factors, oriGen aims to integrate multi-omic platforms including WES and WGS besides clinical and lifestyle variables. compared with the Mexican Biobank, another large-scale project highlighting the peculiarities of Mexican genomics which genotyped over 6000 individuals from all 32 states with an emphasis on capturing Mexico’s genetic diversity and population structure, oriGen focuses on urban populations and longitudinal health outcomes at a larger scale, with deeper clinical phenotyping and multi-omic integration.

The cohort’s centralised biobank will ensure high-quality biospecimen preservation, and its data sharing policy, aligned with international standards gives oriGen the potential to drive discoveries in precision health, chronic disease prevention and gene-environment interaction with emphasis on the Mexican and Latin American population.

The demographic profile of the cohort shows an over-representation of women and individuals aged 55 years or older, suggesting possible selection bias related to recruitment methods or differential willingness to participate; this limits generalisability to the national population structure and could explain the differences in chronic disease prevalence observed between the cohort and the general population. Additionally, while oriGen is geographically diverse, its sampling is largely urban-centric, potentially underrepresenting rural and indigenous populations whose health burdens and exposures may differ substantially.

Ethics approval

The project was approved by the Ethics Committees of Hospital La Misión with records CMZZGA-ORNI.001, CMZZGA-ORNI.002 and CMZMMA-ORGNI.003 in 2021; then Hospital Zambrano-Hellion registered with the Nacional Bioethics Commission (CONBIOETICA—Comisión Nacional de Bioética), IDs CMZMMA-ORNI-003 (2021), 121-2022-CEI-R and 121-2022 CI-R (2022), 060-2022-CEI-R and 060-2022 CI-R (2023); and the Federal Commission for the Protection of Health Risks (COFEPRIS—Comisión Federal de Protección contra Riesgos Sanitarios), ID 213301410D0020/2022. The protocol and other related files are included as an online supplemental file 2.

Supplementary material

online supplemental file 1
jech-80-2-s001.xlsx (120.7KB, xlsx)
DOI: 10.1136/jech-2025-224561
online supplemental file 2
jech-80-2-s002.pdf (6.7MB, pdf)
DOI: 10.1136/jech-2025-224561

Acknowledgements

The authors wish to thank the oriGen laboratory team (Erika Leticia Castillo-González, Belem Torres-Longoria; Francisco Rodríguez-Recio, Ana Victoria Camero-Maldonado, Rocío Marisol Martínez-Rentería, Ana Karen García-García, Susana Cristina Álvarez-Carrasco, Genaro Martínez-García, Sara González-Ayala, Gabriela Martínez-Soto, Gabriel Milán-Sandoval, Alonso Aguilar-Valencia, Jessica Álvarez-Salgado, David Alberto Vazquez and Patrick Alonso del Real Villa), the operations team (Ariel Gutierrez-Buendía, Antonio Arteaga-Rodríguez, Milton De León-Ramos, Bryseida De la Torre-Castro, Cielo Delgado-García, Maria Palcastre-Hernandez and Daniel Dibene de la Fuente) and the bioinformatics team (Eugenio Guzmán-Cerezo, Victor Aguayo López and Álvaro Colin), facility adaptations (Ricardo Marroquin-Novelo), legal support (Monica Ortiz-Garza), communications and media services (Yebel Duron-Villaseñor), laboratory support (Rosa Icela Gonzalez-López and Jade Estefany Pérez-Ong), administrative support (Joanna Angélica Nuncio-Castillo, Raúl Aguirre-Calvillo and Maria Alejandra Venegas-Pinto), TEC MED facilities (Martín de la Cruz-González) and Suasor for field work and map construction.

Footnotes

Funding: The project is fully funded by private capital from FEMSA (Fomento Económico Mexicano, S.A.B. de C.V.) and Tecnológico de Monterrey. FEMSA is one of the main funders of academic and research activities of Tecnológico de Monterrey. FEMSA has no control or decision-making on the oriGen project and will not use the data generated nor the results out of the public data and sample sharing policy.

Provenance and peer review: Not commissioned; externally peer reviewed.

Patient consent for publication: Consent obtained directly from patient(s).

Ethics approval: The project was approved by the Ethics Committees of Hospital La Misión with records CMZZGA-ORNI.001, CMZZGA-ORNI.002 and CMZMMA-ORGNI.003 in 2021; then Hospital Zambrano-Hellion registered with the Nacional Bioethics Commission (CONBIOETICA—Comisión Nacional de Bioética), IDs CMZMMA-ORNI-003 (2021), 121-2022-CEI-R and 121-2022-CI-R (2022), 060-2022-CEI-R and 060-2022-CI-R (2023); and the Federal Commission for the Protection of Health Risks (COFEPRIS—Comisión Federal de Protección contra Riesgos Sanitarios), ID 213301410D0020/2022. Participants gave informed consent to participate in the study before taking part.

Map disclaimer: The inclusion of any map (including the depiction of any boundaries therein), or of any geographic or locational reference, does not imply the expression of any opinion whatsoever on the part of BMJ concerning the legal status of any country, territory, jurisdiction or area or of its authorities. Any such expression remains solely that of the relevant source and is not endorsed by BMJ. Maps are provided without any warranty of any kind, either express or implied.

Data availability free text: Genomic data is available upon request to corresponding author (VT) or following the procedure declared in the section data availability and collaboration.

Correction notice: This article has been corrected since it first published. A typographical error in Dr Garcia-Rivas's first name has been corrected.

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

online supplemental file 1
jech-80-2-s001.xlsx (120.7KB, xlsx)
DOI: 10.1136/jech-2025-224561
online supplemental file 2
jech-80-2-s002.pdf (6.7MB, pdf)
DOI: 10.1136/jech-2025-224561

Data Availability Statement

The oriGen study has established a comprehensive data and sample-sharing policy aimed at maximising the scientific and public health value of its extensive biomedical repository. oriGen’s data sharing and collaboration policy is grounded in the project’s commitment to advance global knowledge in disease prevention, diagnosis and treatment, particularly for health conditions with high prevalence in Mexico. The data repository is being designed as a national and international resource for qualified researchers, and the policy emphasises collaboration with both national and international institutions.

Data sharing will operate under clear ethical, legal and quality safeguards, adhering to international best practices such as those outlined by UK Research and Innovation (UKRI)15 and the University of Oxford’s Nuffield Department of Population Health policy for data sharing.16 Access to data will have two modalities: formal collaborative agreements between research organisations and oriGen, and via open-access data requests evaluated by a dedicated review committee. English versions of the data and sample sharing policy and forms are included as online supplemental file 2.


Articles from Journal of Epidemiology and Community Health are provided here courtesy of BMJ Publishing Group

RESOURCES