Abstract
As health care embraces learning health systems, data‐oriented approaches provide a path for genetic counselors to actively contribute to improving clinical care through evidence‐driven insights. The informatics‐driven querying of electronic health record (EHR) data in genetic counseling has the potential to advance clinical practice, quality improvement, and research. This paper examines the opportunities and challenges associated with leveraging EHR data in genetic counseling, with a focus on practical, collaborative, and future‐oriented applications. Quality improvement initiatives focused on identifying eligible patients for genetic services, assessing access and uptake patterns, and evaluating genetic testing outcomes serve as a natural gateway to research. Here, we demonstrate how genetic counseling professionals can use EHR data to conduct research that drives impactful changes in patient care and service delivery. By highlighting key applications and identifying areas for future exploration, this paper argues that EHR‐based research represents not only a practical solution to current challenges but also the future of genetic counseling inquiry. This approach promises to unlock new opportunities to measure and enhance the effectiveness, equity, and accessibility of genetic counseling services.
Keywords: electronic health record, genetic counseling research, health informatics
What is known about this topic
There are a growing number of examples of genetic counseling research that use informatics‐based approaches.
What this paper adds to the topic
This paper provides a review of informatics‐based approaches that can be incorporated into genetic counseling practice to improve outcomes and reduce the resources needed to conduct genetic counseling research.
1. INTRODUCTION
Genetic counseling practice is full of questions: Are we reaching the right population? What proportion of eligible patients are accessing genetic services, comprised of both genetic counseling and genomic testing? Who currently receives genomic testing, and how is that impacting the overall health system? While manual electronic health record (EHR) abstraction could answer these questions, the time and resource requirements prohibit many clinicians from exploring these concepts. Informatics‐driven approaches to EHR‐based research can expedite these processes, facilitate trans‐disciplinary collaborations, and broadly impact clinical practice. Consider the following examples:
An OB/GYN department at an academic center employs a reproductive genetic counselor for the first time in 3 years. While the population suggests that multiple genetic counselors are needed, there is a reluctance to hire additional genetic counselors without data. They would consider adding new positions if data could show patient volume, referral sources, growth projections, and financial aspects of the current genetic counseling practice suggest the need for additional capacity.
A cardiology department has provided genetic testing in their provider clinics for decades without genetic counseling services. They believe they capture all eligible patients but have never evaluated the data. Specifically, they do not know how many patients qualify for genetic services and the impact of that volume on the nurses who currently order testing.
Leveraging data to inform these clinical practices is the foundation of learning health systems (Anderson et al., 2022). In these systems, teams work toward a continuous cycle of identifying clinical gaps, collecting data, and analyzing the data to improve practice (Institute of Medicine, 2007). Using EHR‐based approaches to conduct gap analyses, guideline acceptance studies, and time‐series research is methods to improve clinical care and advance generalizable knowledge simultaneously. This paper aims to describe how the EHR can be used to conduct access, uptake, service delivery, financial, and other types of research in a clinical practice.
2. THE UNDERLYING EHR INFRASTRUCTURE
To consider how data in the EHR can be used for genetic counseling research, we must first consider how the EHR works at both a system and informatics level. EHRs serve as repositories for a depth of healthcare data beyond clinical information. The EHR includes logistical details about patient encounters, sample collection and transportation, and billing and claims data, all of which can enrich health research. The format of recorded data influences its future accessibility. Data may be collected in discrete (structured) fields/boxes; such as a specific entry for surgery dates, height and weight, or age at menarche; or as narrative (unstructured) text, like clinical notes, pathology notes, and imaging impressions. While narrative data provide rich contextual detail, it can be challenging to analyze systematically, whereas structured data ensure consistency and facilitate easier reporting. For example, in genetic counseling, family history (fhx) data may be entered narratively by a genetic counselor recording notes from a pedigree or through structured fields in the EHR. The discrete fhx fields will utilize drop‐down selections for health conditions and dedicated fields for diagnosis dates.
In addition, EHR infrastructure uses integrated databases to manage both transactional data, which records time‐sensitive events like the steps in placing and fulfilling laboratory orders (e.g., sample collected, sample received, result processing, result reported), and cross‐sectional data, which captures snapshots of data that may change over time such as demographic profiles. For genetic counselors looking to explore, there is a wealth of data nestled within the EHR.
2.1. Defining interoperable discrete variables in the EHR
Multiple clinical terminologies, or “electronic vocabularies,” underlie the EHR infrastructure and provide standardization across health informatics and the EHR. Four terminologies are primarily used for defining the conditions, problems, and processes occurring in health care:
The International Classification of Diseases (ICD) database defines medical problems and diagnoses, consisting of a letter and number (e.g., Z80.0 defines a family history of cancer). An ICD code may also describe an event related to a diagnosis, such as “encounter for ultrasound screening.”
Current Procedural Terminology (CPT®) is a registered trademark of the American Medical Association and is a billing code set. It consists of descriptors and numerical codes for medical procedures, including consults, invasive and noninvasive procedures, and laboratory tests (e.g., 96,041 is a code that can be used to report for each 30 min of genetic counseling service (American Medical Association, 2024)).
SNOMED Clinical Terms (SNOMED CT) provides numerical codes for qualitative outcomes, like the ICD. However, SNOMED CT provides additional definition and granularity, using a hierarchical structure (e.g., Hypertrophic Cardiomyopathy (HCM) is SCTID: 233873004, with child codes for obstructive HCM, fetal HCM, etc.).
Logical Observation Identifiers, Names, and Codes (LOINC) numerically defines observations or processes that evaluate clinical features. For instance, Huntington's disease testing is defined as an HTT gene mutation panel (code 53783‐7) with subparts to analyze the alleles and repeat number. Generally, if LOINC is the question, SNOMED CT is its answer.
The discrete nature of these clinical terminologies provides a starting point for EHR‐based research. For example, let us say an institution bills for genetic counseling services and recommends all pancreatic cancer diagnoses undergo genetic counseling. The ICD version 10 codes C25, C25.1, C25.4, C25.7, C25.8, and/or C25.9 and the CPT® code 96041 (American Medical Association, 2024), for genetic counseling services, can be queried across a specific period. Whether individuals with the specific ICD10 code have a 96,041 CPT® code associated with their management can provide a starting point to determine access and uptake of pancreatic cancer patients undergoing genetic counseling.
The development of computable phenotypes, or conditions and/or clinical features determined solely from the EHR, uses existing terminologies to define clinical concepts. Resources such as the Value Set Authority Center (vsac.nlm.nih.gov) help advance our definitions of conditions using multiple clinical terminology vocabularies. For instance, Type 2 Diabetes Mellitus can be identified through either (A) ICD‐10‐CM codes; (B) the use of Diabetes‐related medication; and/or (C) Hemoglobin A1c values suggestive of uncontrolled Diabetes (Richesson et al., 2022). The combination of multiple terminologies allows us to refine and clarify definitions of medical conditions, including conditions indicated for genetic counseling and/or testing services.
While these four terminologies provide discrete codes that can be queried for EHR‐based research, these codes are limited by institutional localization and nuances within the clinical terminologies. Specific consult codes, such as the referral code used for genetic counseling at an institution, are not defined by the terminologies and would require an institution‐specific query. Additionally, lab and biomarker values can be derived from LOINC codes. However, suppose the diagnosis of a disease requires a specific lab value cut‐off (e.g., individuals with cystic fibrosis must have a specific IRT level on newborn screening). In that case, there must be additional cut‐offs and alerts added to generate useful EHR‐derived data for analysis. Institutional variation makes EHRs functional for clinical use and queryable for institutional analyses. However, beyond nationally standardized terminologies, each EHR has a nuanced structure that must be evaluated before embarking on EHR‐based research.
2.2. Examples of advancing EHR infrastructure for genetic counseling benefit
In medicine, we have a responsibility to manage and design EHR‐based clinical data. Imagine the following scenario: A genetic counselor uses an EHR with a discrete Family History section. The section includes addable/subtractable rows for each family member and drop‐downs for health conditions family members may have. However, the list of family members lacks sensible options such as a distinction between a maternal versus paternal aunt. The health condition drop‐downs moreover only allow “cancer” as a diagnosis, rather than specific cancer types. The genetic counselor could (A) proceed with taking pedigrees on paper and upload them via PDF to meet clinical needs; or (B) reach out to their EHR stewardship committee to discuss the need for adjusted options.
In the latter option, the genetic counselor argues their case and the stewardship group agrees, updating the family history options and adding over 40 distinct cancer types (breast cancer, colon cancer, etc.). The genetic counselor, recognizing limitations of EHR discrete data for family history, (Polubriaginof et al., 2015) still scans their paper pedigrees into Epic but also updates the discrete fields during the visit to record distinct cancer types in the EHR. In 3 years of seeing a caseload of 500 patients a year, the genetic counselor who only records family history data via PDF could perform a time‐intensive chart review of the now greater than 1500 cases. At an efficient 1–5 min per chart, that would take 25–125 h. However, the genetic counselor who successfully navigates updating their EHR fields could pull the same number of cases worth of data within a few hours of setting up an EHR report with their informatics team.
3. USING EHR ABSTRACTIONS FOR GENETIC COUNSELING RESEARCH
Multiple broad mechanisms lend themselves to EHR‐derived data. Forecasting, quality improvement research, assessing patient volume, and analyzing current patterns of referral and uptake can be conducted across multiple conditions and settings (Chishtie et al., 2023). For genetic‐specific analyses, three core functions occur with EHR‐derived data: (A) identifying eligible patients; (B) evaluating referral and uptake patterns; and (C) measuring germline testing and hereditary condition outcomes. Figure 1 describes the three core functions and possible approaches to measure and analyze each core function of EHR‐derived data.
FIGURE 1.

Visual example of the types of analyses and outcome measures that can be assessed during the process of assessing eligible patients, referral rates, scheduling and visit completion outcomes, and long‐term medical outcomes. Notation in parentheses indicates a strategy for the desired analysis. While loosely based on real‐world clinical experience, the example metrics are theoretical. BPA, best practice advisory; CDS, clinical decision support; CPT®, current procedural terminology; EHR, electronic health record; ICD, International Classification of Disease.
3.1. Identifying eligible patients
A fundamental strength of EHR‐based research is its ability to assess the volume of patients who are indicated to receive genetic counseling, meet criteria for genomic testing, and/or are eligible for a clinical trial. Through terminology codes like ICD and SNOMED, you can search for individuals with premature ovarian failure, prolonged QT syndrome, paragangliomas, Down Syndrome, Friedreich ataxia, etc. Some of these indications, such as Down Syndrome, provide genetics teams estimated patient volume for annual management of patients. In contrast, quantifying individuals who need pretest genetic counseling to identify whether they have hereditary risk, such as premature ovarian failure, can help assess approximate new patient volume and genetic counseling need annually. Furthermore, EHR queries can assess eligible patients for gene therapy and other clinical trials for individuals with hereditary conditions. Identifying eligible patients can be done at a national or institutional level.
3.1.1. National and global identification of eligible patients
There are multiple cross‐institutional mechanisms to assess cohort volume using ICD, CPT®, and other clinical terminology codes. Corporate sources like TriNetX, commonly available through academic institutions, use real‐world de‐identified data to incorporate diagnoses, genetic variants, and other demographics (Talviste et al., 2024; Wang et al., 2022). In the genomic space, TriNetX has been used for studies on uveitis risk after Down Syndrome diagnosis (Hsu et al., 2024) and the association of Marfan Syndrome with Peyronie's disease thus far (Kolanukuduru et al., 2024). Similarly, an EHR‐affiliated tool like Epic Cosmos provides a similar dataset across a community of health systems that use the Epic EHR (Tarabichi et al., 2021). Through the integration of inpatient and outpatient charts into a single patient record across multiple health systems, Cosmos can generate higher patient numbers than any one institutions' rare disease volume for research (Kranyak et al., 2023; Wehrli et al., 2023). Lastly, federally supported organizations also sponsor rich data repositories for research. The Patient‐Centered Outcomes Research Institute (PCORI) funds a national health data resource, PCORnet (Fleurence et al., 2014). PCORnet focuses on comparative effectiveness studies and enhancing overall public health (Jackson et al., 2024; McTigue et al., 2020). While not directly EHR‐derived, the same concepts can be applied to PCORnet to identify eligible patients and overall disease burden and genetic counseling need in specific communities. The availability of the aforementioned databases may vary in the future; however, the growing priority on interoperable sharing of health data suggests that some form of these data will be accessible.
3.1.2. Institutional identification of eligible patients
EHRs typically come with a suite of tools designed to query data. These tools are used by administrators to measure all aspects of healthcare delivery and can similarly be leveraged to enhance research. Basic tools such as reporting workbenches, dashboards, and registry overviews can serve as a quick summary or data‐rich environment for large datasets. Increasingly, health systems are integrating more sophisticated versions of these tools, which use optimized data warehouses to allow immediate customizable data abstraction queries. While focused expertise or training‐intensive programs exist to allow EHR superusers to navigate complex data, EHRs have recently focused development on making user‐friendly tools (e.g. Epic's SlicerDicer) accessible to support everyday user inquiries. (Hae et al., 2024; Saini et al., 2021).
Among new‐age EHR technology, Clinical Decision Support (CDS) tools and Best Practice Advisories (BPAs) are additional mechanisms for evaluation. CDS refers to any EHR‐embedded software designed to support clinician, researcher, and/or patient decision‐making (Bright et al., 2012; Scalia et al., 2021; Sutton et al., 2020). Examples include an automated flag for abnormal results, diagnostic assistance (i.e., an EHR suggesting a differential based on the problem list), or a notification that an order would be a duplicate. BPAs are a subtype of CDS best known as “pop‐ups” or in‐line callouts to direct a user to complete an action (Health Journalism Glossary, n.d.). Notably, BPAs can be successful but run a risk of alert fatigue due to repetition and redundant alerts (Ng et al., 2023). BPAs have been used in genetics to recommend providers refer patients to genetics based on certain diagnoses of a problem list (Ramirez et al., 2023; Reddy et al., 2023). Once designed, CDS tools and the resulting user actions can be queried to measure outcomes.
CDS tools are a promising option in facilitating the identification of eligible patients for genetic services (Figure 1). While EHR data queries can be used to identify patients, connecting families with the respective service has traditionally presented a major obstacle. Many clinics have reported leveraging clinician education, patient resources, and navigation (Bednar et al., 2022), only to see a significant drop‐off in patients showing up to genetics appointments after referral (Greenberg et al., 2022). CDS tools can help with patient identification and assist in closing the gaps in patient navigation. A CDS can be designed to identify patients for genetics referral based on family history, notify the clinician at their next visit, and then send the reminder to a patient regarding their appointment. The outcomes of these CDS can then be quantified and used for genetic counseling research. These tools are flexible opportunities to leverage EHR software that can identify and improve follow‐up for patient care.
3.2. Evaluating referral and uptake patterns
A robust area for research in genetic counseling includes recognizing health disparities through gap analysis. Gap analysis refers to the research practice of identifying discrepancies between current practices and recommended standards (Golden et al., 2017). This work highlights areas where care delivery does not align with standards or expectations. Guideline acceptance research is a specific example of gap analysis measuring differences in guideline adherence. This work can focus on different populations such as clinicians, patients, or health systems. Gap analyses and guideline acceptance research include studying recommended services, like genetic counseling referrals. By analyzing referral and uptake patterns (Figure 1), we can evaluate both adherence to guidelines (e.g., what percentage of eligible patients receive a referral) and practice gaps (e.g., how many referred patients attend their genetic counseling consultation).
Within the EHR, discrete data fields can support this evaluation, including:
Referral details—who placed the referral, the referring provider, and referral date.
Referral codes—ICD10, CPT®, and procedural codes.
Referral status—such as “scheduled, declined, or unable to contact”.
Referral outcomes—seen, no show, rescheduled.
Querying these discrete fields can track both whether referrals were made (indicating guideline acceptance) and whether patients completed their referrals (helping assess practice gaps).
Another important use case is time‐based evaluation, tracking time processes, such as the time from diagnosis to referral and from referral to counseling uptake. This approach enables a detailed examination of patient flow and potential bottlenecks. Since EHRs record all relevant information, they offer the flexibility to customize data retrieval for specific metrics. However, consistency is crucial: Clearly defining data points and measurement criteria ensures accuracy and reliability in such analyses.
3.3. Measuring germline testing outcomes
As informatics capacity expands in the wake of modernized EHRs, the ability to record and query genetic data is vastly improving. Increasingly, health systems are investing in bioinformatic data collection and retention pathways for genetic testing results. However, these data remain highly variable depending on the setup and location of the genetic testing:
Internal Reference Laboratory: A laboratory connected directly with a health system may store discrete data directly into the EHR or may provide narrative data, which is highly variable. A common example is pathology. Many health systems have an internal pathology department; however, pathology results are frequently unstructured, opting for PDFs or narrative results letters (Kim et al., 2024). In contrast, nearly all EHRs have automated discrete fields for straightforward internal tests such as complete blood counts (CBCs).
Commercial laboratory with return of unstructured results: The frequent use of commercial laboratories in genetics is unsurprising given the history of development and complexity of these tests. Return of results has however been a challenge since the era of paper records. Genetics reports, frequently summarized for provider convenience, may be relayed back via narrative documents such as PDF files without structured data. While it is possible to manually extract information from PDFs for entry into discrete EHR fields, the time, fidelity error rate, and labor‐intensive nature of this work are often a barrier.
Commercial laboratory with EHR integrations: Increasingly, institutions and laboratories are moving to EHR‐integrated results. These integrations may be direct custom builds or facilitated through an interface developed by the EHR vendor (e.g., Epic's AURA) where the laboratory pays the EHR vendor to mediate connections to their associated health systems. This model can provide discrete genetic data which flow into the EHR without need for manual entry. A direct integration offers customizability, lower costs to the laboratory, and flexibility in queries. Conversely, vendor‐mediated integrations pass costs onto laboratories and promise reduced upkeep from a health system's EHR technical support team.
When contemplating an EHR: laboratory integration, start early by developing a multidisciplinary roadmap of data needs. With structured genetic testing data in the EHR, clinicians and researchers can now determine who received testing, the results of the testing, and the downstream impact of the results on clinical next steps. Consider questions such as “do we need CPT® information for send out tests?” and “how do we want results to be distributed to the patient?” With integrations, the data won't flow if it isn't built. A reminder that the development of these tools warrants careful planning as once implemented they may be burdensome to update or add new data categories to. As genetic counselors work with their health system IT and bio informatics teams moving forward, there is a clear takeaway: The more structured and integrated your family history and genetic testing is, the more data you'll be able to abstract.
3.4. Additional measurable mechanisms in the EHR
Beyond identifying eligible patients, assessing referral patterns and uptake, and quantifying genetic testing outcomes, multiple administrative, medical, and financial measures can be completed within the EHR.
3.4.1. Billing outcomes
Clinics interested in measuring financial outcomes will find a wealth of EHR data directly tied to billing workflows. Given the operational importance of cash flow information, many hospitals regularly track detailed financial data. These data include, but are not limited to: ICD‐10 and CPT® codes, visit dates and indications, insurance type, payer and insurance payments, denial rates and reasons, and latency periods between bill dates and payments. Administrators within the health system have reports on this data readily available. As the genetic counseling field considers sustainability, assessing the financial impact and coverage for services at their institutions is a practical research approach.
3.4.2. Clinical dashboards
Clinical dashboards serve as dynamic data snapshots, often updating in real‐time and providing an organized visual depiction of the current processes occurring in a clinical setting (Campbell et al., 2023). If you are looking for clinic metrics to understand volume, service delivery mechanisms, and referral patterns, a dashboard is a great “first step” en route to EHR‐based research. An example of this, highlighting the volume of eligible patients, the number of referrals and completed genetic counseling visits, and gaps in the process, is shown in Figure 2. The first example in the introduction, where a reproductive genetics clinic needed additional FTE but lacked meaningful metrics to advocate for the position, benefitted from developing a dashboard. Over 2 years, collaborations between the department administration, information technology team, and the genetic counselors led to the creation of a “reproductive genetic counseling dashboard.” Here, they can show patterns in referral and encounter volume month‐over‐month, categorize the mechanism of delivery (e.g., in person versus telehealth), and analyze provider referral patterns. After their initial dashboard generated discrete data proving the need for a second genetic counselor, they successfully advocated for a new position (1.0 FTE). Now, their dashboard prioritizes by location, provider, referral type, etc. resulting in data that tripled the team's FTE in less than 12 months.
FIGURE 2.

This example dashboard highlights the potential for electronic health record‐based research to derive real‐time data that can result in quality improvement initiatives and/or outcomes research.
3.4.3. Downstream revenue
Downstream revenue (DSR) analyses can measure the impact of genetic counseling on subsequent and related services performed at an institution. As genetic counselors look for options to justify value outside of direct billing, DSR provides an important path toward quantifying genetic counselor impact. DSR can differ by specialty, from assessing colonoscopy uptake after a Lynch syndrome diagnosis to evaluating medication adherence with lysosomal storage disorders. EHR data abstraction is perfectly suited for DSR projects as the data required often span multiple hospital programs and departments. To date, genetic counselors have presented examples of DSR in cancer, prenatal, and cardiology; providing researchers interested in DSR studies with multiple examples to model from (Mauer et al., 2021; Mauer Hall et al., 2024; Olson et al., 2024).
3.4.4. Medical insights
The potential to track large‐scale health information is an overt benefit of EHR data. Comorbidity analyses, identification of patients for eligibility in research trials, and longitudinal follow‐up studies all serve as hallmarks of research benefiting from EHR‐derived analyses. In the genetic counseling space, we can consider leveraging these data to think forward. Work in the spaces of health equity and the impact of genetic counselor intervention in downstream health benefit greatly from access to data not contained solely within a genetic counselor's note.
3.4.5. Service delivery mechanisms
The COVID‐19 pandemic ushered in the use of telehealth unexpectedly and rapidly (Ma et al., 2021). However, the continued use of telehealth warrants evaluation for both patient and provider impact. The EHR discretely measures the mechanism of care delivery (e.g., in person, video, phone, etc.) for billing and other purposes and inherently includes visit status and progress (e.g., person showed up, person checked out, etc.). EHR‐derived data can be used to analyze the impact and outcomes of telehealth services at an institution (Muppavarapu et al., 2022). Through genetic counseling referral and visit codes, one could determine whether rates of no‐shows, ordering of genetic testing, or even billable units vary by service delivery (Lin et al., 2020). Overall, evaluating service delivery mechanism outcomes in the EHR provides not only a research measure but also clinical quality improvement measures that can translate to operational process improvement.
4. CONSIDERATIONS WHEN USING INFORMATICS APPROACHES TO EHR ABSTRACTION
While innovative and potentially practice‐changing, aspects of data quality and ethics must be incorporated into EHR‐derived research. Specifically, there is a trade‐off between efficiency and accuracy between EHR‐derived and manually curated data (Early et al., 2022).
4.1. Data quality
Research using the EHR is inherently a retrospective and secondary analysis. The clinical nature of the EHR means that compared to robust research registries, the abstracted EHR data are at risk of “copy paste syndrome” (Amirav & Borycki, 2021), demographic and/or clinical discrepancies (Amirav & Borycki, 2021; Cai et al., 2024), and overall bad data. EHR‐based research must evaluate “fitness for us,” acknowledging this potential and incorporating data quality assessments in their methodology. A variety of frameworks exist, including Weiskopf's measures of data quality (correctness, completeness, concordance, plausibility, and currency) and Wang & Strong's framework of data quality (intrinsic, contextual, representational, and accessible) (Wang & Strong, 1996; Weiskopf et al., 2013; Weiskopf & Weng, 2013). Before data analysis, the study team should consider how they will assess data quality and reconcile conflicting EHR variables.
Consider EHR financial data. Collection rates, reimbursement, and denial statistics represent foundationally important data for hospital financial stewardship. In contrast, a genetic counselor may want to study patient financial toxicity, copay amounts, and denial indications. If the financial data have been designed for administrative use, careful wading into this EHR data is necessary. Does the field “reimbursement” distinguish between payments made by a patient versus Insurance? When the system indicates the patient paid $500 was that due to a denial or a deductible? Nuances that might seem intuitive to a biller may not be readily translatable for a clinician. One claim may be billed with telehealth modifier −95 and the next modifier –GT. The medical coder may know that this is a preference of the insurer the claim is being submitted to, but this information would not be readily available on a spreadsheet being reviewed by a clinical researcher. As genetic counselors translate data derived for one purpose to enrich understanding of another, careful consideration and consultation with expert colleagues are key to ensuring data quality.
Completeness and currency are especially key in time‐based studies and accurate data reporting (Weiskopf & Weng, 2013). If only 1/3 of a patient population can be evaluated, you must determine whether there are patterns of missingness. Does fragmentation of care across institutions change the amount of missing data among cases? Do those with less access to specific health services have more missing data? Can you accurately describe the entire patient population if only 1/3 of the population is queriable? Similarly, progressive diseases are expected to change over time, whether improved from or unresponsive to treatment. Time‐based studies must capture both the date of the event and the clinical state of a condition. It is also possible that an individual qualified for genetic services or specific intervention prior to establishing care in our clinics. Therefore, time studies are more complex due to the event at which they qualified compared to establishing care. In the genetic counseling field, we must move to defining event‐based measures (e.g., time from initial presentation at institution to genetic counseling visit) using the time‐stamped events in the EHR. This will require cross‐institutional collaborations and/or examples at single institutions that advance our shared definition of event‐based measures. Once defined, the opportunity to measure temporal dimensions of care is infinite. To reach these end goals, consideration of the required variables and issues with completeness and/or currency (e.g., the time‐based nature of variables and measures) must occur before study initiation and data abstraction.
Beyond general data quality concerns, two additional biases may exist. First, the EHR cannot speak; therefore, it cannot add social or contextual information that further shapes the findings. If there are cognitive biases in clinical practice relative to the primary outcome (Dobler et al., 2019; Stiegler et al., 2012), interpreting the data runs the risk of inaccurate conclusions. Second, there are multiple considerations for abstracting race, ethnicity, and ancestry data. Recent studies found discordant reporting between the EHR and patient report, misclassification of non‐Hispanic White individuals, and higher rates of inaccurate recording for patients who self‐identify as minorities (Samalik et al., 2023). Conducting rigorous research requires considering the accuracy of self‐reported data and the usability of the EHR to extrapolate findings from the analyses.
4.2. Ethics of data extraction
The EHR is primarily used for entering clinical data and documentation that allows for high‐quality continuity of care. However, when used for research, its secondary use must be strategic and ensure the protection of patients, who become human subjects. Institutional requirements for Institutional Review Board (IRB) permissions must be incorporated. Considerations include patient consent at the time of the medical visit for additional EHR use, limiting research that identifies patients due to the rarity of a condition and/or demographic, and/or waivers of consent when reasonable. Additionally, discordance is common between EHR‐reported demographics and patient‐reported identities (Samalik et al., 2023). Therefore, transparency in reporting the processes and outcomes of demographic reconciliation is crucial.
4.3. Visualizing information gaps
Multiple tools exist to visualize the gaps generated by data analysis to end users. While some EHRs have their own visualization tool (e.g., Epic and Slicer Dicer), other institutions may rely on Power BI (or equivalent softwares) to build a visual of the data. Direct queries using Structured Query Language (SQL) can also generate a summative visual of structured data and its relations among variables. While these tools are optimally used for different audiences or outcomes, collaboration with data visualization and other informatics experts can identify the best tool for you.
4.4. Collaboration with IR and/or other informatics specialists
Advancing healthcare systems and outcomes for patients and the genetic counseling profession requires teamwork. Particularly, using EHR‐derived data requires a multidisciplinary collaboration across researchers, data scientists, informatics specialists, and clinicians. Each team member contributes a specific lens to ensure the appropriate data are abstracted, assessed, and analyzed. There can be significant technical challenges with EHR‐based analyses, and simultaneously, the bar to conduct EHR‐based analyses is lower than one might anticipate. Therefore, genetic counselors with less data experience can still successfully lead EHR‐based analyses, within a larger team or using clinician‐facing software, such as Epic's Slicer Dicer.
At a broader level, we must work to increase the amount of discrete genetic data in the EHR. Currently, aspects of clinical genetic counseling such as the pedigree, genetic test order, and narrative discussion are not easy to query or abstract. However, the more discrete the data, the easier it is to evaluate clinical practice and outcomes at a higher level. EHR data could even be abstracted across specialties to account for the wide reach of genomic medicine to several areas of health care. Overall, improving discrete data abstraction from the EHR enhances the overall health system, contributing to a learning health system that can generatively evaluate and improve clinical practice.
5. THE FUTURE OF DATA ANALYSIS
Simultaneously, the rise of large language models (LLM) ushers in more need for informatics‐focused collaborations between genetic counselors and information scientists. LLM strategies range from chatbots and conversational artificial intelligence (Sato et al., 2021; Siglen et al., 2022) to analyzing EHR data for sentiment and content evaluation (Miah et al., 2024; Tang et al., 2023). In this growth era, the possibilities of EHR‐based research that push the envelope in genetic counseling are endless. Currently, disease‐specific synoptic reports bring data together to ease clinical understanding and interpretation of data, while LLM patient summaries help prepare for initial visits. The use of LLMs can also advance our current EHRs by defining discrete fields for genetic variants and streamlining genetic testing outcomes into measurable variables in the EHR. Additional training of LLMs may soon bring the ability to create pedigrees and family histories from text‐based entries. While the future of LLMs and informatics‐based EHR research continues to blossom, the need for clinicians and other genetic counselor researchers will grow to ensure that models are trained with correct information and genetic counselor expertise.
6. CONCLUSIONS
Integrating informatics‐driven approaches into genetic counseling provides clear opportunities to improve patient care, streamline clinical processes, and drive impactful research. By leveraging EHR data, genetic counselors can identify eligible patients, assess access and uptake patterns, evaluate the outcomes of germline testing, or analyze downstream revenue and billing practices. One example of this could be seen in a cancer genetic setting, where all ovarian cancer diagnoses are recommended to undergo germline testing, often with pretest genetic counseling:
A genetic counselor works with a data specialist to generate all ICD‐10 codes for ovarian cancer (e.g., C56, or more location‐specific codes like C56.1). They would also decide whether to include individuals with an ICD‐10 code for a personal history of ovarian cancer.
The data specialist would use the genetic counseling referral information, alongside the above ICD‐10 codes, to create a list of all patients with ovarian cancer and whether they were referred for genetic counseling.
The genetic counselor could then use the list to evaluate referral patterns.
Through collaboration with the data team, the genetic counselor subsequently evaluates germline testing uptake and outcomes. This could happen through an integrated EHR process for test orders or using the code for a germline test lab order.
The clinical team learns that while most patients undergo germline testing, they may have gaps in genetic counseling uptake. They can now reach out to clinicians about patients who need germline testing, implement posttest counseling follow‐up initiatives, or analyze patterns of referral and testing to see whether there are inequities in access. Through partnerships with data specialists and other informatics experts, genetic counselors are well positioned to advance genetic counseling practice and the broader field using EHR‐centered research approaches.
AUTHOR CONTRIBUTIONS
Each author contributed to the ideation, writing, and editing of the manuscript.
CONFLICT OF INTEREST STATEMENT
The authors do not have any conflicts of interest related to the work and did not require permission to reproduce materials from other sources.
ETHICS STATEMENT
Ethics and integrity policies: Due to the nature of this methods review, a data availability statement, funding statement, ethics approval statement, patient consent statement, and clinical trial registration statement are not applicable.
Permissions: CPT® is a registered trademark of the American Medical Association. However, the material included in this article does not reproduce or directly quote the manual; therefore, it is referenced but does not require separate permissions.
Human studies and Animal studies: This methods review did not utilize humans or animals in its research review, and was not categorized as human subject research.
Greenberg, S. E. , Reys, B. , Fisher, H. , & Basit, M. (2025). The future of electronic health record‐based research: Leveraging informatics in genetic counseling research. Journal of Genetic Counseling, 34, 1–11. 10.1002/jgc4.70058
REFERENCES
- American Medical Association . (2024). CPT 2025 professional edition. American Medical Association. [Google Scholar]
- Amirav, D. , & Borycki, E. M. (2021). Copy and paste in the electronic medical record: A scoping review. Knowledge Management & E‐Learning, 13(4), 522–535. [Google Scholar]
- Anderson, J. L. , Mugavero, M. J. , Ivankova, N. V. , Reamey, R. A. , Varley, A. L. , Samuel, S. E. , & Cherrington, A. L. (2022). Adapting an interdisciplinary learning health system framework for academic health centers: A scoping review. Academic Medicine, 97(10), 1564–1572. [DOI] [PubMed] [Google Scholar]
- Bednar, E. M. , Nitecki, R. , Krause, K. J. , & Rauh‐Hain, J. A. (2022). Interventions to improve delivery of cancer genetics services in the United States: A scoping review. Genetics in Medicine, 24(6), 1176–1186. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bright, T. J. , Wong, A. , Dhurjati, R. , Bristow, E. , Bastian, L. , Coeytaux, R. R. , Samsa, G. , Hasselblad, V. , Williams, J. W. , Musty, M. D. , Wing, L. , Kendrick, A. S. , Sanders, G. D. , & Lobach, D. (2012). Effect of clinical decision‐support systems. Annals of Internal Medicine, 157(1), 29–43. [DOI] [PubMed] [Google Scholar]
- Cai, L. , DeBerardinis, R. J. , Zhan, X. , Xiao, G. , & Xie, Y. (2024). Navigating electronic health record accuracy by examination of sex incongruent conditions. Journal of the American Medical Informatics Association, 31(12), ocae236. 10.1093/jamia/ocae236 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Campbell, I. M. , Karavite, D. J. , McManus, M. L. , Cusick, F. C. , Junod, D. C. , Sheppard, S. E. , Lourie, E. M. , Shelov, E. D. , Hakonarson, H. , Luberti, A. A. , Muthu, N. , & Grundmeier, R. W. (2023). Clinical decision support with a comprehensive in‐EHR patient tracking system improves genetic testing follow up. Journal of the American Medical Informatics Association, 30(7), 1274–1283. 10.1093/jamia/ocad070 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chishtie, J. , Sapiro, N. , Wiebe, N. , Rabatach, L. , Lorenzetti, D. , Leung, A. A. , Rabi, D. , Quan, H. , & Eastwood, C. A. (2023). Use of epic electronic health record system for health care research: Scoping review. Journal of Medical Internet Research, 25, e51003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dobler, C. C. , Morrow, A. S. , & Kamath, C. C. (2019). Clinicians' cognitive biases: A potential barrier to implementation of evidence‐based clinical practice. BMJ Evidence‐Based Medicine, 24(4), 137–140. [DOI] [PubMed] [Google Scholar]
- Early, M. , Gatti, J. , Slobogean, B. , Woods, B. , Ackerman, R. , Mathur, A. , Roberts, J. , & Blakeley, J. (2022). Comparing methods used to identify people with rare tumor predisposition syndrome in the electronic health record (P4‐9.004). Neurology, 98(18_supplement), 1719. 10.1212/WNL.98.18_supplement.1719 [DOI] [Google Scholar]
- Fleurence, R. L. , Curtis, L. H. , Califf, R. M. , Platt, R. , Selby, J. V. , & Brown, J. S. (2014). Launching PCORnet, a national patient‐centered clinical research network. Journal of the American Medical Informatics Association, 21(4), 578–582. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Golden, S. H. , Hager, D. , Gould, L. J. , Mathioudakis, N. , & Pronovost, P. J. (2017). A gap analysis needs assessment tool to drive a care delivery and research agenda for integration of care and sharing of best practices across a health system. Joint Commission Journal on Quality and Patient Safety, 43(1), 18–28. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Greenberg, S. , Orlando, E. , Devlin, M. , Low, S. , Anson, A. , Neil, B. O. , Kohlmann, W. , Wong, B. , Dechet, C. B. , Sanchez, A. , Tward, J. D. , Johnson, S. B. , Agarwal, N. , Kohli, M. , Gupta, S. , Swami, U. , & Maughan, B. L. (2022). Comparing pretest video genetic education for prostate cancer patients: Do patients need assistance? Journal of Clinical Oncology, 40(16_suppl), 5061. 10.1200/JCO.2022.40.16_suppl.5061 [DOI] [Google Scholar]
- Hae, R. , Madan, S. , Acai, A. , Wong, S. , & Gangji, A. S. (2024). Slicer dicer as a potential tool for self‐assessment: SA‐PO1118. Journal of the American Society of Nephrology, 35(10S), 10–1681. [Google Scholar]
- Health Journalism Glossary . (n.d.) Accessed December 15, 2024. https://healthjournalism.org/glossary/
- Hsu, A. Y. , Wang, Y.‐H. , Lin, C.‐J. , Li, Y.‐L. , Hsia, N.‐Y. , Lai, C.‐T. , Kuo, H. T. , Chen, H. S. , Tsai, Y. Y. , & Wei, J. C. C. (2024). Assessing uveitis risk following pediatric down syndrome diagnosis: A TriNetX database study. Medicina, 60(5), 710. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Institute of Medicine . (2007). The learning healthcare system: Workshop summary. National Academies Press. [PubMed] [Google Scholar]
- Jackson, S. L. , Lekiachvili, A. , Block, J. P. , Richards, T. B. , Nagavedu, K. , Draper, C. C. , Koyama, A. K. , Womack, L. S. , Carton, T. W. , Mayer, K. H. , Rasmussen, S. A. , Trick, W. E. , Chrischilles, E. A. , Weiner, M. G. , Podila, P. S. B. , Boehmer, T. K. , Wiltz, J. L. , & PCORnet Network Partners . (2024). Preventive service usage and new chronic disease diagnoses: Using PCORnet data to identify emerging trends, United States, 2018–2022. Preventing Chronic Disease, 21, E49. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim, M. K. , Rouphael, C. , McMichael, J. , Welch, N. , & Dasarathy, S. (2024). Challenges in and opportunities for electronic health record‐based data analysis and interpretation. Gut Liver, 18(2), 201–208. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kolanukuduru, K. P. , Mandel, A. L. , Simhal, R. K. , Sholklapper, T. N. , Sun, K. , Poluch, M. , Wang, K. R. , Shah, Y. B. , & Chung, P. H. (2024). Marfan's syndrome is associated with a greater risk of Peyronie's disease: A case‐control study of the TriNetX database. International Journal of Impotence Research. 10.1038/s41443-024-00923-5 [DOI] [PubMed] [Google Scholar]
- Kranyak, A. , Rork, J. , Levy, J. , & Burdick, T. E. (2023). Alopecia areata and thyroid screening in down syndrome: Leveraging epic cosmos data set. Journal of the American Academy of Dermatology, 89(2), 360–361. [DOI] [PubMed] [Google Scholar]
- Lin, J. C. , Kavousi, Y. , Sullivan, B. , & Stevens, C. (2020). Analysis of outpatient telemedicine reimbursement in an integrated healthcare system. Annals of Vascular Surgery, 65, 100–106. [DOI] [PubMed] [Google Scholar]
- Ma, D. , Ahimaz, P. R. , Mirocha, J. M. , Cook, L. , Giordano, J. L. , Mohan, P. , & Cohen, S. A. (2021). Clinical genetic counselor experience in the adoption of telehealth in the United States and Canada during the COVID‐19 pandemic. Journal of Genetic Counseling, 30(5), 1214–1223. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mauer, C. B. , Reys, B. D. , Hall, R. E. , Campbell, C. L. , & Pirzadeh‐Miller, S. M. (2021). Downstream revenue generated by a cancer genetic counselor. JCO Oncology Practice, 17(9), e1394–e1402. 10.1200/OP.20.00464 [DOI] [PubMed] [Google Scholar]
- Mauer Hall, C. B. , Reys, B. D. , Gemmell, A. P. , Campbell, C. L. , & Pirzadeh‐Miller, S. M. (2024). Downstream revenue generated by patients with hereditary cancer in the multigene panel testing era. JCO Oncology Practice, 20(12), 1695–1704. [DOI] [PubMed] [Google Scholar]
- McTigue, K. M. , Wellman, R. , Nauman, E. , Anau, J. , Coley, R. Y. , Odor, A. , Tice, J. , Coleman, K. J. , Courcoulas, A. , Pardee, R. E. , Toh, S. , Janning, C. D. , Williams, N. , Cook, A. , Sturtevant, J. L. , Horgan, C. , Arterburn, D. , & PCORnet Bariatric Study Collaborative . (2020). Comparing the 5‐year diabetes outcomes of sleeve gastrectomy and gastric bypass: The National Patient‐Centered Clinical Research Network (PCORNet) bariatric study. JAMA Surgery, 155(5), e200087. 10.1001/jamasurg.2020.0087 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Miah, M. S. U. , Kabir, M. M. , Sarwar, T. B. , Safran, M. , Alfarhood, S. , & Mridha, M. F. (2024). A multimodal approach to cross‐lingual sentiment analysis with ensemble of transformer and LLM. Scientific Reports, 14(1), 9603. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Muppavarapu, K. , Saeed, S. A. , Jones, K. , Hurd, O. , & Haley, V. (2022). Study of impact of telehealth use on clinic “No show” rates at an academic practice. The Psychiatric Quarterly, 93(2), 689–699. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ng, H. J. H. , Kansal, A. , Abdul Naseer, J. F. , Hing, W. C. , Goh, C. J. M. , Poh, H. , D'souza, J. L. A. , Lim, E. L. , & Tan, G. (2023). Optimizing best practice advisory alerts in electronic medical records with a multi‐pronged strategy at a tertiary care hospital in Singapore. JAMIA Open, 6(3), ooad056. 10.1093/jamiaopen/ooad056 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Olson, M. , Anderson, J. , Knapke, S. , Kushner, A. , Martin, L. , Statile, C. , Shikany, A. , & Miller, E. M. (2024). Cardiac genetic counseling services: Exploring downstream revenue in a pediatric medical center. Journal of Genetic Counseling, 34(2), e1984. 10.1002/jgc4.1984 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Polubriaginof, F. , Tatonetti, N. P. , & Vawdrey, D. K. (2015). An assessment of family history information captured in an electronic health record. American Medical Informatics Association Annual Symposium Proceedings, 2015, 2035–2042. [PMC free article] [PubMed] [Google Scholar]
- Ramirez, D. L. , Corredor, J. , Cordero‐Hernandez, I. , & Arun, B. (2023). Development of provider ordered genetic testing pathway and electronic medical record best practice advisory (BPA) to aid in germline genetic testing for patients with metastatic breast cancer. JCO Oncology Practice, 19(11_suppl), 443. 10.1200/OP.2023.19.11_suppl.443 [DOI] [Google Scholar]
- Reddy, P. , Mersch, J. , Howell, M. , Bullock, J. , Reys, B. , Read, P. , Rajpurohit, N. , Ghabach, B. , & Narra, K. (2023). Increasing cancer genetic referrals via best practice alerts (BPA). JCO Oncology Practice, 19(11_suppl), 565. 10.1200/OP.2023.19.11_suppl.565 [DOI] [Google Scholar]
- Richesson, R. W. L. , Gold, S. , Rasmussen, L. , & NIH Health Care Systems Research Collaboratory Electronic Health Records Core Working Group . (2022). Electronic health records–based phenotyping: Definitions. Rethinking clinical trials: A living textbook of pragmatic clinical trials. NIH Pragmatic Trials Collaboratory. [Google Scholar]
- Saini, V. , Jaber, T. , Como, J. D. , Lejeune, K. , & Bhanot, N. (2021). 623. Exploring ‘slicer dicer’, an extraction tool in EPIC, for clinical and epidemiological analysis. Open Forum Infectious Diseases, 8(Supplement_1), S414–S415. [Google Scholar]
- Samalik, J. M. , Goldberg, C. S. , Modi, Z. J. , Fredericks, E. M. , Gadepalli, S. K. , Eder, S. J. , & Adler, J. (2023). Discrepancies in race and ethnicity in the electronic health record compared to self‐report. Journal of Racial and Ethnic Health Disparities, 10(6), 2670–2675. [DOI] [PubMed] [Google Scholar]
- Sato, A. , Haneda, E. , Suganuma, N. , & Narimatsu, H. (2021). Preliminary screening for hereditary breast and ovarian cancer using a Chatbot augmented intelligence genetic counselor: Development and feasibility study. JMIR Formative Research, 5(2), e25184. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Scalia, P. , Ahmad, F. , Schubbe, D. , Forcino, R. , Durand, M.‐A. , Barr, P. J. , & Elwyn, G. (2021). Integrating option grid patient decision aids in the epic electronic health record: Case study at 5 health systems. Journal of Medical Internet Research, 23(5), e22766. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Siglen, E. , Vetti, H. H. , Lunde, A. B. F. , Hatlebrekke, T. A. , Stromsvik, N. , Hamang, A. , Hovland, S. T. , Rettberg, J. W. , Steen, V. M. , & Bjorvatn, C. (2022). Ask Rosa—The making of a digital genetic conversation tool, a chatbot, about hereditary breast and ovarian cancer. Patient Education and Counseling, 105(6), 1488–1494. [DOI] [PubMed] [Google Scholar]
- Stiegler, M. P. , Neelankavil, J. P. , Canales, C. , & Dhillon, A. (2012). Cognitive errors detected in anaesthesiology: A literature review and pilot study. British Journal of Anaesthesia, 108(2), 229–235. [DOI] [PubMed] [Google Scholar]
- Sutton, R. T. , Pincock, D. , Baumgart, D. C. , Sadowski, D. C. , Fedorak, R. N. , & Kroeker, K. I. (2020). An overview of clinical decision support systems: Benefits, risks, and strategies for success. npj Digital Medicine, 3(1), 17. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Talviste, G. , Leinsalu, M. , Ross, P. , & Viigimaa, M. (2024). Lipid‐lowering treatment gaps in patients after acute myocardial infarction: Using global database TriNetX. Medicina, 60(9), 1433. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tang, L. , Sun, Z. , Idnay, B. , Nestor, J. G. , Soroush, A. , Elias, P. A. , Xu, Z. , Ding, Y. , Durrett, G. , Rousseau, J. F. , Weng, C. , & Peng, Y. (2023). Evaluating large language models on medical evidence summarization. npj Digital Medicine, 6(1), 158. 10.1038/s41746-023-00896-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tarabichi, Y. , Frees, A. , Honeywell, S. , Huang, C. , Naidech, A. M. , Moore, J. H. , & Kaelber, D. C. (2021). The cosmos collaborative: A vendor‐facilitated electronic health record data aggregation platform. ACI Open, 5(1), e36–e46. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wang, R. Y. , & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33. [Google Scholar]
- Wang, W. , Wang, C.‐Y. , Wang, S.‐I. , & Wei, J. C.‐C. (2022). Long‐term cardiovascular outcomes in COVID‐19 survivors among non‐vaccinated population: A retrospective cohort study from the TriNetX US collaborative networks. eClinicalMedicine, 53, 101619. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wehrli, L. A. , Reppucci, M. L. , Ketzer, J. , Dominguez‐Muñoz, A. , Cooper, E. H. , Peña, A. , Bischoff, A. , & de la Torre, L. (2023). Incidence of medullary thyroid carcinoma and Hirschsprung disease based on the cosmos database. Pediatric Surgery International, 39(1), 227. [DOI] [PubMed] [Google Scholar]
- Weiskopf, N. G. , Hripcsak, G. , Swaminathan, S. , & Weng, C. (2013). Defining and measuring completeness of electronic health records for secondary use. Journal of Biomedical Informatics, 46(5), 830–836. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Weiskopf, N. G. , & Weng, C. (2013). Methods and dimensions of electronic health record data quality assessment: Enabling reuse for clinical research. Journal of the American Medical Informatics Association, 20(1), 144–151. [DOI] [PMC free article] [PubMed] [Google Scholar]
