Skip to main content
American Journal of Epidemiology logoLink to American Journal of Epidemiology
. 2024 Aug 9;194(5):1436–1447. doi: 10.1093/aje/kwae272

Hierarchical clustering analysis to inform classification of congenital malformations for surveillance of medication safety in pregnancy

Loreen Straub 1,, Shirley V Wang 2, Sonia Hernandez-Diaz 3, Kathryn J Gray 4, Seanna M Vine 5, Massimiliano Russo 6, Leena Mittal 7, Brian T Bateman 8, Yanmin Zhu 9, Krista F Huybrechts 10
PMCID: PMC13070530  PMID: 39123096

Abstract

There is growing interest in the secondary use of health care data to evaluate medication safety in pregnancy. Tree-based scan statistics (TBSS) offer an innovative approach to help identify potential safety signals; they use hierarchically organized outcomes, generally based on existing clinical coding systems that group outcomes by organ system. When assessing teratogenicity, such groupings often lack a sound embryologic basis, given the etiologic heterogeneity of congenital malformations. The study objective was to enhance the grouping of congenital malformations to be used in scanning approaches through implementation of hierarchical clustering analysis (HCA) and to pilot test an HCA-enhanced TBSS approach for medication safety surveillance in pregnancy in 2 test cases using > 4.2 million mother-child dyads from 2 US-nationwide databases. Hierarchical clustering analysis identified (1) malformation combinations belonging to the same organ system already grouped in existing classifications, (2) known combinations across different organ systems not previously grouped, (3) unknown combinations not previously grouped, and (4) malformations seemingly standing on their own. Testing the approach with valproate and topiramate identified expected signals and a signal for an HCA-cluster missed by traditional classification. Augmenting existing classifications with clusters identified through large data exploration may be promising when defining phenotypes for surveillance and causal inference studies.

Keywords: clustering, malformations, pregnancy, surveillance, teratogen, topiramate, valproate

Introduction

When evaluating medication safety in pregnancy, teratogenicity tends to be a primary concern, given the morbidity, and, in the most severe cases, infant death associated with major structural congenital malformations. Traditionally, studies have focused on malformations overall or on broad categories of organ-specific malformations (eg, cardiac malformations) for several reasons. The limited number of exposed pregnancies available for study often does not support the evaluation of individual malformation types, which tend to be rare. Additionally, because drug structure and function do not predict teratogenicity and because animal studies are limited in their ability to predict human teratogenesis, clear hypotheses in terms of which specific malformations might be associated with the drug and, therefore, are worth evaluating are usually lacking.

The increasing availability of large health care utilization databases (eg, insurance claims databases, electronic health records, nationwide or provincial registries) with potential for linkage across family members and methodological advancements to mitigate confounding have changed the landscape of perinatal pharmacoepidemiology in the past decade.1 These secondary data sources give the research community access to data on many more exposed pregnancies than used to be the case when primary data collection was the main source of information. Outside of pregnancy, the ready access to such large databases has spurred the development of scanning approaches, such as tree-based scan statistics (TBSS), which use a hierarchical outcome classification structure and systematically screen for a broad range of potential adverse effects associated with the exposure of interest.2-4 Tree-based scan statistics have been used by the US Food and Drug Administration and the Centers for Disease Control and Prevention for several vaccination and medication safety surveillance activities.5-8 For the mentioned reasons, such a scanning approach is very appealing to evaluate teratogenicity, and early assessments have shown promise for its application in pregnancy.9-11

These prior assessments have relied on a hierarchical outcome classification structure embedded in existing clinical coding systems, such as the International Classification of Diseases (ICD) or the Multi-Level Clinical Classification.12 Inherent in these coding systems is their grouping of codes based on organ systems. Given the etiologic heterogeneity of malformations, such groupings often lack a sound embryologic basis.13-15 More meaningful groupings would be based on knowledge of the underlying embryology and pathogenesis. For example, it is important to group malformations that are anatomically different but pathogenetically similar, such as defects of presumed vascular etiology.15 Sometimes the pattern of structural defects can be attributed to a primary problem in morphogenesis that leads to a cascade of consequent defects, known as a “sequence” (eg, congenital diaphragmatic hernia leading to pulmonary hypoplasia). Not considering potential common pathological pathways when screening for associations could result in missed safety signals for uncommon clinically related outcomes.

The objective of this study was 2-fold: (1) to enhance the grouping of congenital malformations to be used in scanning approaches through the implementation of agglomerative hierarchical clustering analysis (HCA), and (2) to pilot test an HCA-enhanced TBSS approach for medication safety surveillance in pregnancy using 2 test cases of medications (valproate and topiramate) known to cause specific congenital malformations. Although we focus on congenital malformations, the concepts and approach extend readily to other maternal or child outcomes.

Methods

Data sources and study cohorts

We used mother-child linked pregnancy cohorts nested within the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files (MAX/TAF; 2000-2018) and the Merative MarketScan Commercial Claims Database (MarketScan; 2003-2020). The creation of these cohorts has been described previously.16,17 Mothers aged 12-55 years were required to have continuous insurance coverage from ≥3 months before pregnancy to ≥1 month thereafter; children were required to have continuous coverage from birth until ≥3 months thereafter, unless they died sooner. Pregnancies with fetal chromosomal anomalies or with exposure to known teratogenic medications were excluded (Figure 1, Figure S1).

Figure 1.

Figure 1

Cohort selection for pregnancies in the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files (MAX/TAF; 2000-2018) and the Merative MarketScan commercial claims database (MarketScan; 2003-2020). Abbreviation: LMP, last menstrual period.

The subcohort for the HCA approach consisted of all children with ≥1 recorded congenital malformation based on the International Classification of Diseases, 9th Revision (ICD-9; or 10th Revision, ICD-10) coding system (ICD-9: 740.x-759.x; ICD-10: Q00.x-Q99.x) either in the child records within the first 3 months after birth or in the maternal records within the first month after delivery. Maternal records were considered because insurance claims are sometimes recorded under the mother’s information while the child’s eligibility is being processed.18 This research was approved by the institutional review board of Brigham and Women’s Hospital, which granted a waiver of informed consent.

Tree-based scan statistics

The implementation of TBSS for medication safety surveillance in the general population and during pregnancy has been described before.2,3,9 Briefly, this approach scans a hierarchical outcome tree for associations with the exposure of interest. For each outcome (node) in the tree, a log-likelihood ratio is calculated, with the maximum log-likelihood ratio being the test statistic T. Because the distribution of T is unknown, it is derived by generating distributions under the composite null hypothesis of no difference in risk between the exposed and comparator groups for any node in the tree via Monte Carlo simulation. The alternative hypothesis is that there is at least 1 node for which there is elevated risk for the exposure group. The test statistics from randomly generated data sets under the composite null hypothesis, the observed data set are then ranked, and the P value is determined by the rank of the observed test statistic divided by the total number of data sets (randomly generated and observed).

One example of a hierarchical tree used in TBSS is the ICD classification, which groups clinical concepts into increasingly aggregated, higher-level, clinically related categories. By scanning each node (individual code or code category) at each hierarchical level of the tree, TBSS allow for the evaluation of a broad range of potential adverse events, as well as of clinically related conditions.

Tree-enhancement through hierarchical clustering analysis to identify malformation clusters

We used the ICD-based hierarchy as the starting point for the outcome tree. Because most of the data used in perinatal epidemiology studies include many years from a period when ICD-9 codes were in use (ie, any time before October 2015 in the United States), we initially developed a tree based on the ICD-9 structure. To accommodate inclusion of data from the ICD-10 era, we subsequently linked ICD-10 codes to their corresponding ICD-9 codes using the forward-backward mapping method.19,20 When ICD-10 codes linked to multiple ICD-9 codes or did not link to any ICD-9 code, we manually assigned ICD-10 codes to the most appropriate ICD-9 code (Appendix S1).

To maximize the power to detect adverse effects on developmental processes (eg, malformations that share common causes or are part of a sequence), we enhanced the ICD-based tree, which is organized strictly by organ system, through the implementation of HCA to identify clusters of pathogenetically similar malformations. Hierarchical clustering analysis is a method that groups similar objects into clusters. In this implementation, we grouped individual ICD codes that co-occurred, as assessed based on the Pearson correlation coefficient.21,22 Each individual code was treated as a cluster and pairs of clusters were sequentially agglomerated using the average Euclidean distance between their members until all clusters were merged into a single cluster comprising all malformation codes. A cluster dendrogram was created to visualize the hierarchical clustering structure between all data points in the cohort (Figure 2, Appendix S2).

Figure 2.

Figure 2

Example of the hierarchical clustering analysis approach illustrated based on malformations grouped at the 3-digit International Classification of Diseases, 9th Revision (ICD-9) level, restricted to ICD-9 codes 740.X-749.X. This figure illustrates the application of the hierarchical clustering analysis and the creation of a cluster dendrogram in an example that focuses on a sample of children with ≥1 ICD-9 malformation diagnosis starting with code 740-749. Because there are several hundred individual ICD-9 malformation codes, for illustration purposes, this example focuses only on malformations already grouped at the 3-digit ICD-9 level. The first table shows the respective malformations (noted with value 1) recorded for each child. The second table shows the Euclidean distance between each 2-code combination respectively. For details on the calculation, see Appendix S2. In this table, the 2-code combination with the smallest Euclidean distance is highlighted in red (combination of ICD-9 code 745.x and 747.x). This combination builds the first cluster in the cluster dendrogram. The third table shows the Euclidean distance between each 2-code combination respectively, after considering ICD-9 codes 745.x and 747.x as 1 cluster. In this table, the next 2-code combination with the smallest Euclidean distance is highlighted in green (combination of clusters 745.x and 747.x and ICD-9 code 746.x). This combination builds the second cluster in the cluster dendrogram. The cluster dendrogram is a visualization of the hierarchical clustering structure among all 3-digit ICD-9 codes 740.x-749.x in the cohort, with the y-axis representing the Euclidean distance between ICD-9 codes within the respective cluster. More details on the implementation of the hierarchical clustering analysis are shown in Appendix S2.

We initially implemented this approach separately for the MAX/TAF and MarketScan cohort. However, because the identified clusters and cluster hierarchies were consistent across data sources, we implemented the approach in both cohorts combined. We subsequently reviewed the clusters to screen out groupings of codes that appeared to be used interchangeably rather than representing shared disease processes (eg, codes referring to “other ear anomaly” or “unspecified ear anomaly”) and clusters based on extremely sparse data (n < 10 children with ≥2 recorded codes contributing to the cluster). We selected the top 50% clusters, which reflect the strongest clusters based on the smallest Euclidian distance (Appendix S2). The ICD-based congenital malformation tree was enhanced with these clusters by adding new “branches” to the tree (Figure 3).

Figure 3.

Figure 3

Example of an International Classification of Diseases, 9th Revision (ICD-9) digit-based tree structure (blue boxes), enhanced with malformation clusters identified through hierarchical clustering analysis (HCA; red boxes) for 1 child’s unique combination of congenital malformations. The level 5 node (the leaf node) comprises unique combinations of ICD-9 diagnostic codes for malformations observed during the outcome assessment period. Each code within this leaf node has a unique level 4 node, a 5-digit ICD-9 code. A placeholder code, marked with “P” and a dashed gray line, was used if there was no 5-digit code available. The respective level 4 code is then grouped at increasingly aggregated higher levels based on 4-digit codes (level 3), 3-digit codes (level 2), and, finally, any congenital malformation code at the highest level (level 1) of the hierarchical ICD-9 tree. All blue boxes and blue lines correspond to the existing ICD-9 tree-based structure. All red boxes and red lines correspond to the malformation clusters that were identified through HCA and added to the tree. We evaluated each node at every level of the HCA-enhanced tree above the leaf level (level 5). Nodes in level 5 were not tested.

Test cases

We evaluated the HCA-enhanced TBSS approach using 2 test cases: the association between congenital malformations and first trimester exposure to (1) valproate and (2) topiramate. For valproate, teratogenicity has been well documented, with risk increases previously observed for multiple malformations including oral cleft, neural tube, cardiovascular, genitourinary, and limb defects.23-25 We previously showed that TBSS successfully detected such known risks associated with prenatal valproate.9 Here, we used the same test case, with the inclusion of more years of data (2015-2018 for MAX/TAF and 2016-2020 for MarketScan) and thus more statistical power to detect potential alerts, to assess whether the HCA-based tree enhancement can help identify additional, previously unknown, alerts.

Although multiple studies have shown an increase in risk of oral clefts after early-pregnancy topiramate exposure,26-30 the risk of other malformation types is less well understood. Therefore, we selected topiramate as the second test case.

Pregnancies were considered exposed if ≥1 medication prescription of interest was dispensed during the first trimester. Unexposed pregnancies were defined based on the absence of any anticonvulsant dispensing between 3 months prepregnancy until the end of the first trimester. A broad range of potential confounding variables was considered, including maternal demographic characteristics (race/ethnicity [available for MAX/TAF only], US state, delivery year, age), indications for anticonvulsant use (eg, epilepsy/seizure, bipolar disorder, migraine/headache), other chronic comorbidities (eg, diabetes, hypertension, obstetric comorbidity index31), substance use (eg, alcohol use disorder, smoking), concomitant medication exposure (eg, antipsychotics, antidepressants), and proxies for health care utilization (eg, number of outpatient visits, hospitalizations and emergency department visits) (Figure S1, Table S1). Propensity scores (PSs) were estimated separately for MAX/TAF and MarketScan based on logistic regression models that included all prespecified covariates. Observations from the nonoverlapping regions of the PS distribution were trimmed, and PS fine stratification (with 50 equally sized PS strata based on the distribution among the exposed pregnancies) was used to account for confounding.32,33 Unexposed pregnancies were then weighted using the distribution of the exposed among PS strata. After pooling weighted PS strata from both cohorts while accounting for cohort heterogeneity (see Tables 1 and 2 for details), an unconditional Poisson scan statistic was used to calculate weighted relative risks (RRs).2,3,10,34,35 In secondary analyses, to further mitigate confounding by indication and associated factors, we restricted the cohorts to those with an epilepsy/seizure indication.

Table 1.

Safety alerts for valproate exposure during the first trimester of pregnancy, based on pregnancy cohorts nested within the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files (2000-2018) and the Merative MarketScan commercial claims database (2003-2020).

ICD-9 code (cluster node identifier) a , b Node description No. of observed cases c No. of expected cases d , e Relative risk P value
Full cohort
P < .05
 759.7 Multiple congenital anomalies, so described 16 4.19 3.82 .0019
 741.93 Spina bifida without mention of hydrocephalus, lumbar region <11 e 9.22 .0027
 Cluster (759.7, 759.89) Multiple congenital anomalies, so described; other specified congenital anomalies 24 8.89 2.7 .0043
 741.0x Spina bifida with hydrocephalus <11 e 6.25 .0091
 741 Spina bifida with hydrocephalus, unspecified region <11 e 7.17 .0118
 741.9x Spina bifida without mention of hydrocephalus <11 e 4.76 .0123
 741.03 Spina bifida with hydrocephalus, lumbar region <11 e 8.73 .0124
 752.61 Hypospadias 37 18.24 2.03 .0169
 Cluster (745.5, 747.0,  747.83) Ostium secundum type atrial septal defect; patent ductus arteriosus; persistent fetal circulation 229 175.68 1.3 .017
 Cluster (745.5, 747.0) Ostium secundum type atrial septal defect; patent ductus arteriosus 225 172.8 1.3 .0209
 745.x Bulbus cordis anomalies and anomalies of cardiac septal closure 210 160.32 1.31 .025
 759.x Other and unspecified congenital anomalies 35 17.38 2.01 .0288
P ≥ .05, RR  4
 Cluster (743.47, 756.51) Specified congenital anomaly sclera; osteogenesis imperfecta <11 e 6.58 .0531
 752.7 Indeterminate sex and pseudohermaphroditism <11 e 6.16 .0748
 741.9 Spina bifida without hydrocephalus, unspecified region <11 e 4.29 .1112
 756.51 Osteogenesis imperfecta <11 e 6.23 .4196
 755.1 Syndactyly of multiple and unspecified sites <11 e 4.04 .8753
Restricted to patients with epilepsy/seizure
 P < .05
 752.x Congenital anomalies of genital organs 40 15.41 2.6 .0001
 740.x-759.x Congenital anomalies (overall) 232 162.13 1.43 .0001
 745.x Bulbus cordis anomalies and anomalies of cardiac septal closure 75 39.25 1.91 .0001
 745.5 Ostium secundum type atrial septal defect 65 33.8 1.92 .0002
 Cluster (745.5, 747.0,  747.83) Ostium secundum type atrial septal defect; patent ductus arteriosus; persistent fetal circulation 77 42.57 1.81 .0002
 752.7 Indeterminate sex and pseudohermaphroditism <11 e 39.94 .0004
 Cluster (745.5, 747.0) Ostium secundum type atrial septal defect; patent ductus arteriosus 76 42.33 1.8 .0004
 Cluster (752, 752.7, 752.8,  752.81, 752.89, 752.9) Congenital anomalies of genital organs; indeterminate sex and pseudohermaphroditism; other specific anomalies of genital organs <11 e 20.55 .0004
 752.5x Undescended and retractile testicle 17 4.61 3.69 .0008
 759.x Other and unspecified congenital anomalies 17 4.69 3.63 .001
 752.61 Hypospadias 16 4.44 3.6 .0028
 752.51 Undescended testis 16 4.53 3.53 .0037
 747.x Other congenital anomalies of circulatory system 56 30.53 1.83 .0043
 759.7 Multiple congenital anomalies, so described <11 e 6.39 .0058
 747 Patent ductus arteriosus 41 20.34 2.02 .0058
 Cluster (752.61, 752.63) Hypospadias; congenital chordee 17 5.38 3.16 .0064
 754.6x Congenital valgus deformities feet <11 e 11.31 .0086
 Cluster (746, 746.8,  746.89, 746.9) Other congenital anomalies of heart 16 5.47 2.93 .0184
 Cluster (759.7, 759.89) Multiple congenital anomalies, so described; other specified congenital anomalies 11 2.91 3.78 .0209
 Cluster (745.3, 746.7,  745.1, 745.10, 745.11,  745.12, 745.19, 746,  746.1, 746.8, 746.89,  746.9) Transposition of great vessels; common ventricle anomalies; other congenital anomalies of heart; congenital tricuspid atresia/stenosis; hypoplastic left heart syndrome 16 5.57 2.87 .0232
 749.0x Cleft palate <11 e 6.09 .0433
 749 Cleft palate, unspecified <11 e 6.09 .0433
P ≥ .05, RR ≥ 4
 759.9 Unspecified congenital anomaly <11 e 5.42 .0737
 742.4 Other specific congenital brain anomalies <11 e 5.98 .1134
 749.x Cleft palate and cleft lip <11 e 4.02 .1499
 741.9 Spina bifida without hydrocephalus, unspecified region <11 e 5.44 .1609
 741.03 Spina bifida with hydrocephalus, lumbar region <11 e 6.96 .1684
 741.9x Spina bifida without mention of hydrocephalus <11 e 5.12 .1828
 741.0x Spina bifida with hydrocephalus <11 e 4.82 .2311
 Cluster (756.7, 756.70,  756.79) Congenital anomaly of abdominal wall (4-digit code); congenital anomaly of abdominal wall, unspecified; congenital anomaly of abdominal wall, other <11 e 5.21 .3487

Abbreviation: ICD-9, International Classification of Diseases, Ninth Revision.

aAll nodes with P < .05 or relative risk ≥4 are included.

bChildren contributed to a certain cluster if they had at least 1 of the specific malformations included in the cluster definition recorded.

cObserved cases represent the number of children with the respective ICD code or included in the respective cluster with prenatal exposure to valproate. Cell sizes <11 for observed cases are suppressed in accordance with the Centers for Medicare and Medicaid Services cell size suppression policy.

dExpected cases represent the number of children with the respective ICD code or included in the respective cluster in the reference group after propensity score (PS) fine stratification and weighting of the reference group using the distribution of the exposed among PS strata. Specifically, expected counts were calculated separately for the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files and Merative MarketScan commercial claims database cohorts based on (number of PS-weighted unexposed children contributing the respective code or cluster) × (total no. of exposed children) / (total no. of unexposed children). Both cohort-specific expected counts were then summed.

eNumbers for expected cases were suppressed whenever cell sizes for observed cases were <11, to avoid back calculation.

Table 2.

Safety alerts for topiramate exposure during the first trimester of pregnancy, based on pregnancy cohorts nested within the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files (2000-2018) and the Merative MarketScan commercial claims database (2003-2020).

ICD-9/Cluster node identifier a , b Node description No. of observed cases c No. of expected cases d , e Relative risk P value
Full cohort
P < .05
 No alerts identified
P ≥ .05, RR ≥ 4
 744.0x Congenital anomalies of ear causing impairment of hearing <11 e 4.19 .152
Restricted to patients with epilepsy/seizure
P < .05
 749.2x Cleft palate with cleft lip <11 e 10.92 .0006
 749.2 Cleft palate with cleft lip, unspecified <11 e 10.92 .0006
 749.x Cleft palate and cleft lip <11 e 6.59 .0013
 755.63 Other congenital deformity of hip (joint) 21 7.55 2.78 .006
 749.0x Cleft palate <11 e 7.05 .0084
 749 Cleft palate, unspecified <11 e 7.05 .0084
 748.3 Other anomalies of larynx, trachea, and bronchus 19 7.11 2.67 .0191
 755.6x Other congenital anomalies of lower limb including pelvic girdle 22 8.92 2.47 .0191
 Cluster (754.3, 754.30,  754.31, 754.32,  754.33, 754.35, 755.63) Congenital dislocation of hip; other congenital deformity of hip 23 9.7 2.37 .0226
 748.x Congenital anomalies respiratory system 22 9.47 2.32 .0409
P ≥ .05, RR ≥ 4
 No alerts identified

Abbreviation: ICD-9, International Classification of Diseases, Ninth Revision.

aAll nodes with P < .05 or relative risk ≥ 4 are included.

bChildren contributed to a certain cluster if they had at least 1 of the specific malformations included in the cluster definition recorded.

cObserved cases represent the number of children with the respective ICD code or included in the respective cluster with prenatal exposure to topiramate. Cell sizes <11 for observed cases are suppressed in accordance with the Centers for Medicare and Medicaid Services cell-size-suppression policy.

dExpected cases represent the number of children with the respective ICD code or included in the respective cluster in the reference group after propensity score (PS) fine stratification and weighting of the reference group using the distribution of the exposed among PS strata. Specifically, expected counts were calculated separately for the Medicaid Analytic eXtract/Transformed Medicaid Statistical Information System Analytic Files and Merative MarketScan commercial claims database cohorts based on (number of PS-weighted unexposed children contributing the respective code or cluster) × (total number of exposed children) / (total number of unexposed children). Both cohort-specific expected counts were then summed.

eNumbers for expected cases were suppressed whenever cell sizes for observed cases were <11 to avoid back calculation.

We evaluated each node at each level of the tree above the leaf level, which comprised unique combinations of malformation codes observed for ≥1 child, without double-counting outcomes when children had multiple diagnoses grouped in the same higher-level aggregated outcome nodes of the tree. Outcome nodes with > 3 exposed cases and P < .05 (1-sided) were considered statistical alerts. We used P values to rank and prioritize alerts for further investigation rather than to decide on the presence of a causal association. However, because most malformations tend to be rare and adjustment for multiple outcomes reduces power to detect true effects, we also assessed statistical nonalerts with high RRs that might merit further assessment. Although we report all nodes with P < 1, we focus our discussion specifically on all statistical alerts as well as on strong associations (RR ≥ 4) independently from statistical significance. Lastly, to evaluate the trade-off between increasing the likelihood of identifying signals through inclusion of potentially more meaningful groupings through HCA and decreasing power by adding more nodes to the tree, we further compared the results from the HCA-enhanced TBSS to those from the TBSS prior to cluster inclusion.

Results

Hierarchical clustering analyses

The MAX/TAF and MarketScan cohort included, respectively, 358 644 and 341 903 children with ≥1 malformation code. When implementing HCA for these children, we identified different types of malformation clusters: (1) malformation combinations belonging to the same organ or organ system (eg, various urinary tract anomalies, cleft lip and cleft palate, components within cardiac anomalies), which are already grouped similarly in the ICD coding system; (2) known combinations and sequences across different organ systems (eg, diaphragmatic anomaly and pulmonary hypoplasia) not grouped in existing coding systems; and (3) previously not recognized combinations of malformations with strong correlation (eg, Meckel’s diverticulum and urachus anomaly). In addition, we found several malformations that seemed to stand on their own, showing only a very weak correlation with other malformations (eg, pyloric stenosis, abdominal wall defects, transverse limb defects). The final 62 clusters that were added to the malformation tree are listed in Table S2.

Test cases

The MAX/TAF cohort included 2 448 605 unexposed pregnancies (n = 9080 with a recorded epilepsy/seizure diagnosis), 3935 valproate-exposed pregnancies (n = 907 with epilepsy/seizure), and 6414 topiramate-exposed pregnancies (n = 990 with epilepsy/seizure). Respective numbers in MarketScan were 1 766 299 unexposed pregnancies (n = 2625 with epilepsy/seizure), 328 valproate-exposed pregnancies (n = 114), and 2718 topiramate-exposed pregnancies (n = 322). Congenital malformations were recorded in 347 338 (MAX/TAF) and 337 264 (MarketScan) prenatally unexposed, 683 and 67 valproate-exposed, and 1224 and 600 topiramate-exposed children, respectively (Figure 1). Mothers exposed to valproate or topiramate were more likely to be White (MAX/TAF), to have comorbid conditions, concomitant medication exposure and other substance use. The cohort characteristics among exposed and unexposed before and after PS weighting, stratified by cohort, are shown in Table S3.

Test case 1: Valproate exposure during the first trimester.

Consistent with the findings from our previous study,9 we found statistical alerts at the P < .05 threshold for several known associations including spina bifida (RR range = 4.76-9.22), bulbus cordis and cardiac septal closure anomalies (RR = 1.31), hypospadias (RR = 2.03), and other/unspecified malformations (RR = 2.01), driven by the ICD code for multiple anomalies (RR = 3.82).

With the consideration of the additional HCA-identified clusters, we further identified a potential alert for ostium secundum atrial septal defect (ASD) and/or patent ductus arteriosus (PDA) (RR = 1.3). Although we also observed similar increased risks for ASD and PDA individually (RR = 1.29 and 1.3), they would not have been detected with the prespecified threshold (P = .13 and .67, respectively) due to insufficient power.

When focusing on nodes with high RRs that did not meet the alert threshold, additional potential associations (all driven by a small number [<11] of exposed cases; RR range = 4.04-6.58) were observed for syndactyly—which, together with other limb defects, has previously been suggested to be associated with prenatal valproate36—for another spina bifida node, as well as for unanticipated signals of indeterminate sex and pseudohermaphroditism, osteogenesis imperfecta, and the cluster of osteogenesis imperfecta and/or anomalies of the sclera.

When restricting the cohort to those with recorded epilepsy/seizure, statistical alerts were observed for bulbus cordis and cardiac septal closure anomalies (RR = 1.91), ASD (RR = 1.92), several genital malformation codes (including indeterminate sex and pseudohermaphroditism, undescended and retractile testicle, and hypospadias; RR range = 2.6 to 39.94), malformations overall (RR = 1.43), other/unspecified malformations (RR = 3.63, driven by the ICD-code for multiple anomalies), other circulatory system anomalies (RR = 1.83, driven by PDA), cleft palate (RR = 6.09), and valgus deformities of feet (RR = 11.31). Although we further identified alerts for several clusters, these were driven by individual malformations for which statistical alerts were already identified. For associations with RR ≥ 4 that did not reach the alerting threshold (all with < 11 exposed cases; RR range = 4.02-6.96), we identified several nodes for spina bifida, unspecified congenital anomaly, cleft palate and cleft lip, and other specific brain anomalies, as well as a for a cluster of anomalies of the abdominal wall (Table 1, Table S4).

Test case 2: Topiramate exposure during the first trimester.

When assessing the association with topiramate in the full cohort, no alerts were observed at the P < .05 threshold, and inclusion of HCA-identified clusters also did not yield any alerts. Although a 1.67- to 2.59-fold increase in risk was observed for several oral-cleft–related codes, the alert threshold was not reached. Only 1 association with RR ≥ 4 (with < 11 exposed cases) that did not reach the alerting threshold was observed; for ear anomaly causing impaired hearing (RR = 4.19).

When restricting to those with epilepsy/seizure, alerts were observed for multiple cleft palate (± cleft lip) codes (RR range = 6.59-10.92), other congenital deformity of hip (RR = 2.78), other congenital anomalies of lower limb including pelvic girdle (RR = 2.47), and congenital anomalies of the respiratory system (RR = 2.32, driven by other anomalies of larynx, trachea, and bronchus [RR = 2.67]). Although we observed an alert for the cluster of congenital dislocation of hip and/or other congenital deformity of hip (RR = 2.37), this was driven by the ICD-code for deformity of hip for which the alert threshold was met. No additional associations with high RRs beyond the alert threshold were observed (Table 2, Table S5). When evaluating the TBSS approach without enhancement of the tree based on the HCA-identified clusters, the P values changed only marginally (Table S6).

Discussion

Our study aim was to improve surveillance of medication safety in pregnancy through augmentation of existing hierarchical malformation classifications organized by body system used in TBSS with new clusters that reflect the embryologic tissue of origin or teratogenic mechanism. These clusters were identified using HCA, a form of unsupervised machine learning, followed by human adjudication. The approach largely coincided with the ICD classification hierarchy but also suggested splitting and lumping certain malformations that are not considered in classification systems organized by body system.

Pilot testing of the approach for valproate, a medication with a well-established safety profile, resulted in the expected signals. Through the inclusion of HCA-identified clusters, we detected a signal for ASD and PDA (RR = 1.31; P = .021), which was not detected through codes or code groupings in the ICD hierarchy based on the statistical alert threshold or high RRs and, therefore, would have been missed. The ICD codes do not provide sufficient information to distinguish secundum ASD from patent foramen ovale, yet many of the recorded ASD diagnoses might represent patent foramen ovale.37,38 Thus, grouping the 2 persistent cardiac shunts—patent foramen ovale and PDA—which allow the oxygenated blood from the placenta to bypass the nonfunctioning lungs during the fetal period—is pathogenetically meaningful and, as shown here, can increase the power to detect potential signals. These findings, therefore, further support lumping of embryologically similar phenotypes for malformation risk evaluation. The other signals detected for valproate were identified based on the ICD hierarchy and, in several instances, were confirmed on the basis of an HCA-identified cluster.

In addition to the expected signals for valproate, we also found associations with a high RR that did not reach the alerting threshold for indeterminate sex and pseudohermaphroditism, osteogenesis imperfecta, and the cluster of osteogenesis imperfecta and/or sclera anomalies in the full cohort, as well as for abdominal wall anomalies in the epilepsy/seizure-restricted cohort. The alert for indeterminate sex and pseudohermaphroditism might be due to the estrogenic effect of phthalates, which have been inactive ingredients in some valproate formulations (as our team previously hypothesized9) or could represent cases of severe hypospadias initially recorded as indeterminate sex.

The association with abdominal wall defects (ie, omphalocele) has been suggested before39; however, to our knowledge, the associations with osteogenesis imperfecta and sclera anomalies have not been described previously. Osteogenesis imperfecta is a genetic disorder typically caused by defects in or related to the collagen type 1 protein. Although genetic variants associated with medical conditions are often inherited, some genetic variants causing osteogenesis imperfecta might occur after conception.40 Furthermore, medication exposure could potentially worsen the phenotype of an underlying genetic condition. Blue sclera is commonly seen in patients with osteogenesis imperfecta and is caused by thinning and increased transparency of the scleral collagen fibers. Given that valproate is known to downregulate synthesis of bone proteins, including collagen type 1,41,42 an association with these 2 conditions seems biologically plausible.

For topiramate, the association with individual malformations is less well understood. Although multiple studies have shown an increase in risk of oral clefts after early-pregnancy topiramate exposure, most of the evidence is based on topiramate treatment for epilepsy.26-30 Our group previously found a substantially weaker association between oral clefts and lower doses of topiramate, typically used to treat nonepilepsy conditions, but with imprecise estimates (RR = 1.45; 95% CI, 0.54-3.86).26 Using TBSS, oral cleft signals were only detected at the alert threshold in patients with epilepsy/seizure, rather than in the entire cohort. Restriction to the underlying indication can alter the RR estimate for several reasons: it might represent a population exposed to higher doses, but also better treatment compliance and longer-term exposure. Such restriction can further mitigate potential residual confounding by indication by making the exposure and reference groups more comparable to each other. Thus, depending on the exposure of interest, sensitivity analyses restricting to specific indications, imposing dose thresholds, or requiring multiple dispensings to reduce exposure misclassification could be considered, with the acknowledgment that this can substantially reduce sample size.

To our knowledge, the potential alerts observed for ear, respiratory, and lower limb/hip anomalies after topiramate exposure have not been previously reported. Although we cannot rule out the possibility of chance findings, it is advisable to conduct adequately powered follow-up studies focusing on these malformations that incorporate robust confounding control.

An important limitation of the HCA-based approach for cluster identification is that it relies on malformations co-occurring in the same child. Consequently, malformations that are embryologically or pathogenetically related but only occurred in different children did not contribute to the cluster identification. Also, HCA proceeds in a pairwise fashion; that is, pairs of individual codes and pairs of clusters are sequentially agglomerated. Although we do not anticipate this to have adversely affected the identification of meaningful clusters, other approaches that consider larger clusters could be explored (eg, latent semantic analysis43). It is also important to note that we did not include all malformation groupings with potentially shared pathways but prioritized those with the strongest correlations. For example, although we observed pairwise correlations between several characteristic features affected by the vertebral, anal, cardiac, trachea-esophageal, renal, and limb anomalies association, these were not selected, because each feature had a much stronger correlation with other malformations (eg, the diagnostic code for anal atresia/stenosis was most strongly correlated with other intestinal malformations). To balance the benefit of including meaningful groupings against potential power loss when scanning many nodes, we focused on the strongest correlations. As demonstrated in the 2 test cases, when compared with the TBSS results without HCA-based enhancement of the tree, adding the 62 clusters to the tree hardly decreased statistical power. Future studies could explore the value of including known teratogenic syndromes that are not already coded as such in existing classification systems and are not known to be caused by inherited genetic mutations or chromosomal anomalies unrelated to the exposure of interest. Future refinements of this approach could further focus on excluding groupings of malformations coming from the same organ system, which are already grouped similarly in existing classification systems, requiring repeated documentation of the respective malformation to increase outcome specificity (as we describe later in this section), or excluding malformations generally considered minor.

Despite the use of large data sets, limited power remains a challenge with scanning-based approaches. Based on simulation work, Suarez et al.10 found that with the Poisson model, approximately 4000 exposed pregnancies are needed to detect a 2-fold increase in risk for more common outcomes, which they defined as a prevalence of around 8 per 1000 births. This finding clearly highlights the importance of incorporating clinically relevant groupings of malformations in the hierarchical outcome tree, particularly for uncommon exposures, because they will have a higher prevalence than the individual malformations contributing to the cluster. Regardless, for less common exposures, relaxing the P value threshold >.05 may be advisable to reduce the likelihood of false-negative results. Given that our test cases comprised almost 4000 valproate-exposed and almost 6500 topiramate-exposed pregnancies, we set the alerting threshold to P < .05 but provided all alerts for P < 1, and we also focused on high relative risks irrespective of the P value. Indeed, it may be counterproductive to rely exclusively on an arbitrary significance threshold for alerting. Outcomes that do not alert at the prespecified threshold may still have a higher-than-expected incidence and may be worthy of additional follow-up through continued targeted screening or causal inference studies.9

In addition to considering HCA-identified clusters, our study is distinguished by including data from both the ICD-9 and ICD-10 eras in the outcome tree. This required a careful mapping to align codes between the 2 systems prior to aggregation in the hierarchical tree. The ability to consider all available data before October 2015 (when ICD-10 was introduced in the United States) further benefits study power, reducing the likelihood of false-negative findings, and is a major strength of our approach.

In causal inference studies, it is generally recommended to use a highly specific outcome definition, at the cost of sensitivity, to ensure unbiased relative risks.44 In the context of malformation outcomes, this often translates to requiring more than 1 date with a code indicating a particular malformation or at least 1 date plus a corrective surgery or procedure for the malformation (eg, to reduce the risk of including rule-out diagnoses or less severe conditions that might resolve on their own).1 In scanning-based studies, a single diagnostic code is often considered sufficient evidence of the outcome. This is the approach we implemented here for the test cases. Suarez et al.10 demonstrated that a highly sensitive outcome definition optimizes power, despite the potential for outcome misclassification, which will bias toward the null. Nevertheless, use of more specific outcome definitions may be advisable as a sensitivity analysis, or even as the primary analysis when power is less of a concern.

The current implementation of TBSS uses a frequentist approach. Our focus was on enhancing the tree structure. A promising next step to further optimize the method would be to explore Bayesian approaches that borrow information from hierarchically related outcomes in the enhanced tree. Bayesian approaches to testing are particularly promising for analyzing complex linked data.45,46 In the context of using TBSS for assessing medication safety in pregnancy, the ability of the Bayesian approach to use prior distributions that can leverage the dependence induced by data may provide more precise estimates of the associations of interest.

In summary, there has been a growing interest in the secondary use of health care utilization data to detect safety signals for medications used in pregnancy. In the US regulatory context, the Prescription Drug User Fee Act VII includes several commitments related to this objective.47 Tree-based scan statistics offer an innovative approach to exploit the large number of exposed pregnancies and rich outcome information included in these data to help identify potential safety signals. This study suggests that augmenting existing clinical classifications typically used in TBSS, with clusters identified through large data exploration, may be a promising approach when defining phenotypes for surveillance activities and causal inference studies of congenital malformations.

Supplementary Material

Web_Material_kwae272

Acknowledgments

This research was presented at the Annual Meeting of the Society for Birth Defects Research and Prevention, Vancouver, BC, Canada, June 25-29, 2022; and at the Annual Meeting of the International Society for Pharmacoepidemiology, Copenhagen, Denmark, August 24-28, 2022.

Contributor Information

Loreen Straub, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Shirley V Wang, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Sonia Hernandez-Diaz, Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, United States.

Kathryn J Gray, Division of Maternal-Fetal Medicine, Department of Obstetrics and Gynecology, Brigham and Women's Hospital, Boston, MA, United States.

Seanna M Vine, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Massimiliano Russo, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Leena Mittal, Department of Psychiatry, Brigham and Women’s Hospital, Boston, MA, United States.

Brian T Bateman, Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University School of Medicine, Stanford, CA, United States.

Yanmin Zhu, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Krista F Huybrechts, Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women’s Hospital and Harvard Medical School, Boston, MA, United States.

Supplementary material

Supplementary material is available at the American Journal of Epidemiology online.

Funding

This study was supported by grant R01-HD104646 from the Eunice Kennedy Shriver National Institute of Child Health and Human Development.

Conflict of interest

L.S. is supported by training grant T32HD040128 from the National Institute of Child Health and Human Development. S.V.W. is an ad hoc consultant to Veracity Healthcare Analytics, Exponent Inc, and MITER Corporation’s Federally Funded Research and Development Center for the Centers for Medicare and Medicaid Services on work unrelated to this article. S.H.-D. is an investigator on research grants to her institution from Takeda and UCB and provided methodologic consulting to Moderna and Johnson & Johnson outside the presented work. K.J.G. reports having served as a consultant for BillionToOne, Roche, and Aetion outside the scope of the submitted work. Y.Z. is an investigator on research grants to her institution from Takeda. K.F.H. is an investigator on research grants to her institution from Takeda and UCB outside the presented work. The other authors declare no conflicts.

Disclaimer

The funding source had no role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; and preparation, review, or approval of the manuscript; and decision to submit the manuscript for publication.

Data availability

Patient-level data cannot be made publicly available.

References

  • 1. Huybrechts  KF, Bateman  BT, Hernández-Díaz  S. Use of real-world evidence from healthcare utilization data to evaluate drug safety during pregnancy. Pharmacoepidemiol Drug Saf.  2019;28(7):906-922. 10.1002/pds.4789 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Kulldorff  M, Dashevsky  I, Avery  TR, et al.  Drug safety data mining with a tree-based scan statistic. Pharmacoepidemiol Drug Saf.  2013;22(5):517-523. 10.1002/pds.3423 [DOI] [PubMed] [Google Scholar]
  • 3. Kulldorff  M, Fang  Z, Walsh  SJ. A tree-based scan statistic for database disease surveillance. Biometrics. 2003;59(2):323-331. 10.1111/1541-0420.00039 [DOI] [PubMed] [Google Scholar]
  • 4. Kulldorff  M. TreeScan User Guide. Vol. v2.1. 2022. https://www.treescan.org/cgi-bin/treescan/register.pl/treescan.v2.1.userguide.pdf?todo=process_userguide_download [Google Scholar]
  • 5. Lee  H, Hong  B, Kim  S, et al.  Post-marketing surveillance study on influenza vaccine in South Korea using a nationwide spontaneous reporting database with multiple data mining methods. Sci Rep.  2022;12(1):20256. 10.1038/s41598-022-21986-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Yih  WK, Kulldorff  M, Dashevsky  I, et al.  Sequential data-Mining for adverse events after recombinant herpes zoster vaccination using the tree-based scan statistic. Am J Epidemiol.  2023;192(2):276-282. 10.1093/aje/kwac176 [DOI] [PubMed] [Google Scholar]
  • 7. Brown  JS, Petronis  KR, Bate  A, et al.  Drug adverse event detection in health plan data using the gamma Poisson shrinker and comparison to the tree-based scan statistic. Pharmaceutics.  2013;5(1):179-200. 10.3390/pharmaceutics5010179 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Liu  CH, Huang  WT, Chie  WC, et al.  Safety surveillance of varicella vaccine using tree-temporal scan analysis. Vaccine.  2021;39(43):6378-6384. 10.1016/j.vaccine.2021.09.035 [DOI] [PubMed] [Google Scholar]
  • 9. Huybrechts  KF, Kulldorff  M, Hernández-Díaz  S, et al.  Active surveillance of the safety of medications used during pregnancy. Am J Epidemiol.  2021;190(6):1159-1168. 10.1093/aje/kwaa288 [DOI] [PubMed] [Google Scholar]
  • 10. Suarez  EA, Nguyen  M, Zhang  D, et al.  Monitoring drug safety in pregnancy with scan statistics: a comparison of two study designs. Epidemiology.  2023;34(1):90-98. 10.1097/EDE.0000000000001561 [DOI] [PubMed] [Google Scholar]
  • 11. Suarez  EA, Nguyen  M, Zhang  D, et al.  Novel methods for pregnancy drug safety surveillance in the FDA Sentinel System. Pharmacoepidemiol Drug Saf. 2023;32(2):126-136. 10.1002/pds.5512 [DOI] [PubMed] [Google Scholar]
  • 12. Agency for Healthcare Research and Quality Healthcare Cost and Utilization Project. Clinical Classifications Software (CCS) 2015 . Accessed March 20, 2023. https://www.hcup-us.ahrq.gov/toolssoftware/ccs/CCSUsersGuide.pdf
  • 13. Rasmussen  SA, Olney  RS, Holmes  LB, et al.  Guidelines for case classification for the National Birth Defects Prevention Study. Guideline. Birth Defects Res A Clin Mol Teratol.  2003;67(3):193-201. 10.1002/bdra.10012 [DOI] [PubMed] [Google Scholar]
  • 14. Khoury  MJ, Moore  CA, James  LM, et al.  The interaction between dysmorphology and epidemiology: methodologic issues of lumping and splitting. Teratology.  1992;45(2):133-138. 10.1002/tera.1420450206 [DOI] [PubMed] [Google Scholar]
  • 15. Martin  ML, Khoury  MJ, Cordero  JF, et al.  Trends in rates of multiple vascular disruption defects, Atlanta, 1968-1989: is there evidence of a cocaine teratogenic epidemic?  Teratology.  1992;45(6):647-653. 10.1002/tera.1420450609 [DOI] [PubMed] [Google Scholar]
  • 16. Palmsten  K, Huybrechts  KF, Mogun  H, et al.  Harnessing the Medicaid Analytic eXtract (MAX) to evaluate medications in pregnancy: design considerations. PLoS One.  2013;8(6):e67405. 10.1371/journal.pone.0067405 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. MacDonald  SC, Cohen  JM, Panchaud  A, et al.  Identifying pregnancies in insurance claims data: methods and application to retinoid teratogenic surveillance. Pharmacoepidemiol Drug Saf.  2019;28(9):1211-1221. 10.1002/pds.4794 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Centers for Medicare & Medicaid Services . Medicaid Analytic eXtract (MAX) General Information. Accessed November 2017. https://www.cms.gov/Research-Statistics-Data-and-Systems/Computer-Data-and-Systems/MedicaidDataSourcesGenInfo/MAXGeneralInformation.html
  • 19. Centers for Medicare and Medicaid Services. ICD-10 . January 15, 2021. Accessed March 20, 2023. http://www.cms.gov/Medicare/Coding/ICD10/
  • 20. Fung  KW, Richesson  R, Smerek  M, et al.  Preparing for the ICD-10-CM transition: automated methods for translating ICD codes in clinical phenotype definitions. EGEMS (Wash DC). 2016;4(1):1211. 10.13063/2327-9214.1211 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Zhang  Z, Murtagh  F, Van Poucke  S, et al.  Hierarchical cluster analysis in clinical research with heterogeneous study population: highlighting its visualization with R. Ann Transl Med.  2017;5(4):75. 10.21037/atm.2017.02.05 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Petushkova  NA, Pyatnitskiy  MA, Rudenko  VA, et al.  Applying of hierarchical clustering to analysis of protein patterns in the human cancer-associated liver. PLoS One.  2014;9(8):e103950. 10.1371/journal.pone.0103950 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Weston  J, Bromley  R, Jackson  CF, et al.  Monotherapy treatment of epilepsy in pregnancy: congenital malformation outcomes in the child. Cochrane Database Syst Rev.  2016;11(11):CD010224. 10.1002/14651858.CD010224.pub2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Mutlu-Albayrak  H, Bulut  C, Çaksen  H. Fetal valproate syndrome. Pediatr Neonatol.  2017;58(2):158-164. 10.1016/j.pedneo.2016.01.009 [DOI] [PubMed] [Google Scholar]
  • 25. Jentink  J, Loane  MA, Dolk  H, et al.  Valproic acid monotherapy in pregnancy and major congenital malformations. N Engl J Med.  2010;362(23):2185-2193. 10.1056/NEJMoa0907328 [DOI] [PubMed] [Google Scholar]
  • 26. Hernandez-Diaz  S, Huybrechts  KF, Desai  RJ, et al.  Topiramate use early in pregnancy and the risk of oral clefts: a pregnancy cohort study. Neurology.  2018;90(4):e342-e351. 10.1212/WNL.0000000000004857 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Margulis  AV, Mitchell  AA, Gilboa  SM, et al.  Use of topiramate in pregnancy and risk of oral clefts. Am J Obstet Gynecol. 2012;207(5):405.e1-405.e7. 10.1016/j.ajog.2012.07.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Hunt  S, Russell  A, Smithson  WH, et al.  Topiramate in pregnancy: preliminary experience from the UK Epilepsy and Pregnancy Register. Neurology.  2008;71(4):272-276. 10.1212/01.wnl.0000318293.28278.33 [DOI] [PubMed] [Google Scholar]
  • 29. Hernández-Díaz  S, Smith  CR, Shen  A, et al.  Comparative safety of antiepileptic drugs during pregnancy. Neurology.  2012;78(21):1692-1699. 10.1212/WNL.0b013e3182574f39 [DOI] [PubMed] [Google Scholar]
  • 30. Mines  D, Tennis  P, Curkendall  SM, et al.  Topiramate use in pregnancy and the birth prevalence of oral clefts. Pharmacoepidemiol Drug Saf.  2014;23(10):1017-1025. 10.1002/pds.3612 [DOI] [PubMed] [Google Scholar]
  • 31. Bateman  BT, Mhyre  JM, Hernandez-Diaz  S, et al.  Development of a comorbidity index for use in obstetric patients. Obstet Gynecol.  2013;122(5):957-965. 10.1097/AOG.0b013e3182a603bb [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Desai  RJ, Rothman  KJ, Bateman  BT, et al.  A propensity-score-based fine stratification approach for confounding adjustment when exposure is infrequent. Epidemiology.  2017;28(2):249-257. 10.1097/EDE.0000000000000595 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Huybrechts  KF, Straub  L, Karlsson  P, et al.  Association of in utero antipsychotic medication exposure with risk of congenital malformations in Nordic countries and the US. JAMA Psychiatry.  2023;80(2):156-166. 10.1001/jamapsychiatry.2022.4109 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Kulldorff  M.  TreeScan User Guide for version 2.1 . 2014. Updated July 2022. Accessed March 20, 2023. https://www.treescan.org/cgi-bin/treescan/register.pl/treescan.v2.1.userguide.pdf?todo=process_userguide_download
  • 35. Wang  SV, Maro  JC, Baro  E, et al.  Data mining for adverse drug events with a propensity score-matched tree-based scan statistic. Epidemiology. 2018;29(6):895-903. 10.1097/EDE.0000000000000907 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Rodríguez-Pinilla  E, Arroyo  I, Fondevilla  J, et al.  Prenatal exposure to valproic acid during pregnancy and limb deficiencies: a case-control study. Am J Med Genet.  2000;90(5):376-381. 10.1002/(SICI)1096-8628(20000228)90:5<376::AID-AJMG6>3.0.CO;2-V [DOI] [PubMed] [Google Scholar]
  • 37. Straub  L, Huybrechts  KF, Bateman  BT, et al.  The impact of technology on the diagnosis of congenital malformations. Am J Epidemiol.  2019;188(11):1892-1901. 10.1093/aje/kwz153 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Bateman  BT, Hernandez-Diaz  S, Straub  L, et al.  Association of first trimester prescription opioid use with congenital malformations in the offspring: population based cohort study. BMJ.  2021;372:n102. 10.1136/bmj.n102 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Clayton-Smith  J, Bromley  R, Dean  J, et al.  Diagnosis and management of individuals with fetal valproate spectrum disorder; a consensus statement from the European Reference Network for Congenital Malformations and Intellectual Disability. Orphanet J Rare Dis.  2019;14(1):180. 10.1186/s13023-019-1064-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Eunice Kennedy Shriver National Institute of Child Health and Human Development. What causes osteogenesis imperfecta (OI)?  Accessed July 7, 2023, https://www.nichd.nih.gov/health/topics/osteogenesisimp/conditioninfo/causes
  • 41. Humphrey  EL, Morris  GE, Fuller  HR. Valproate reduces collagen and osteonectin in cultured bone cells. Epilepsy Res. 2013;106(3):446-450. 10.1016/j.eplepsyres.2013.06.011 [DOI] [PubMed] [Google Scholar]
  • 42. Pitetzis  DA, Spilioti  MG, Yovos  JG, et al.  The effect of VPA on bone: from clinical studies to cell cultures-the molecular mechanisms revisited. Seizure.  2017;48:36-43. 10.1016/j.seizure.2017.03.013 [DOI] [PubMed] [Google Scholar]
  • 43. Deerwester  S, Dumais  ST, Furnas  GW, et al.  Indexing by latent semantic analysis. J Am Soc Inf Sci.  1990;41(6):391-407. 10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9 [DOI] [Google Scholar]
  • 44. Ahlbom  A. Modern Epidemiology, 4th edition. TL Lash, TJ VanderWeele, S Haneuse, KJ Rothman. Wolters Kluwer, 2021. Eur J Epidemiol.  2021;36(8):767-768. 10.1007/s10654-021-00778-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Russo  M, Durante  D, Scarpa  B. Bayesian inference on group differences in multivariate categorical data. Comput Stat Data Anal. 2018;126:136-149. 10.1016/j.csda [DOI]
  • 46. Durante  D, Dunson  DB. Bayesian inference and testing of group differences in brain networks. Bayesian Analysis. 2018;13(1):29-58. 10.1214/16-BA1030 [DOI] [Google Scholar]
  • 47. US Food and Drug Administration . PDUFA reauthorization performance goals and prodecures fiscal years 2023 through 2027.  Accessed March 7, 2023. https://www.fda.gov/media/151712/download

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Web_Material_kwae272

Data Availability Statement

Patient-level data cannot be made publicly available.


Articles from American Journal of Epidemiology are provided here courtesy of Oxford University Press

RESOURCES