Abstract
Abstract
Objectives
Automated surveillance of endoscopy-related adverse events (AEs) from non-English free-text reports remains insufficiently evaluated. We compared rule-based and transformer-based natural language processing (NLP) for classifying index procedure reports according to 30-day clinician-adjudicated AE status.
Design
Retrospective, single-centre diagnostic accuracy study.
Setting
Tertiary endoscopy unit in Türkiye, August 2015 to August 2025.
Participants
The source cohort comprised 140 385 reports. All 1512 lexicon-positive and 2000 randomly sampled lexicon-negative reports underwent clinician adjudication, identifying 1208 AE-positive reports. A 13 333-report corpus was allocated at patient level to training (n=9333), validation (n=2000) and a locked, fully adjudicated test set (n=2000; 180 AE-positive).
Primary and secondary outcome measures
The primary outcome was confirmation of an attributable AE within 30 days. Models analysed the index report only, whereas adjudication included subsequent documentation. Secondary outcomes were AE category and documentation timing. Diagnostic accuracy measures included sensitivity, specificity, predictive values, F1 score and transformer precision–recall area under the curve (PR-AUC).
Results
The rule-based model achieved 84.4% sensitivity (95% CI 78.4% to 89.0%) and 99.5% specificity (95% CI 99.1% to 99.7%). The transformer achieved higher sensitivity (92.2%, 95% CI 87.4% to 95.3%; McNemar’s test, p=0.01) and 98.4% specificity (95% CI 97.7% to 98.9%). Positive predictive values were 94.4% and 85.1%, and negative predictive values were 98.5% and 99.2%, for the rule-based and transformer models, respectively; both had F1 scores of 0.89. Transformer PR-AUC was 0.91 (95% CI 0.87 to 0.94). The transformer reduced false negatives from 28 to 14 but increased false positives from 9 to 29. Cohort-wide AE incidence was not estimated because lexicon-negative reports were only partially verified.
Conclusions
Both approaches had high specificity. The transformer reduced false-negative classifications but increased false-positive classifications. These findings support clinician-supervised retrospective case identification and quality assurance, but not autonomous diagnosis, prospective prediction or point-of-care use. External validation is required.
Keywords: Natural Language Processing, Endoscopy, Adverse events, Artificial Intelligence
STRENGTHS AND LIMITATIONS OF THIS STUDY.
The study used a large source cohort of consecutive, routinely collected endoscopy reports spanning 10 years.
Independent dual-clinician adjudication, with disagreements resolved by consensus and, when required, review by a third senior endoscopist, provided the reference standard.
Patient-level dataset allocation and evaluation in a locked, fully clinician-adjudicated test set reduced the risk of information leakage and excluded provisionally labelled reports from the diagnostic-accuracy assessment.
Lexicon-based candidate enrichment and partial verification of lexicon-negative reports may have introduced selection and verification bias and precluded direct estimation of unweighted cohort-wide adverse event (AE) incidence.
The single-centre design limits generalisability, and AEs documented exclusively after the index procedure were not directly observable from the models’ input text.
Introduction
Endoscopic procedures span a broad clinical spectrum, from high-volume diagnostic examinations such as upper gastrointestinal endoscopy and colonoscopy to advanced interventions including endoscopic retrograde cholangiopancreatography (ERCP) and endoscopic ultrasound (EUS)-guided procedures. Although endoscopy is generally safe, adverse events (AEs), including perforation, clinically significant bleeding, cardiopulmonary or sedation-related events, post-ERCP pancreatitis and infection or cholangitis, may result in unplanned admission, repeat intervention, surgery or intensive care. Reliable identification of these events is therefore central to clinical audit and endoscopy quality assurance.1 2
In routine practice, however, systematic AE surveillance is difficult to sustain. Procedure reports are commonly stored as unstructured free text, with considerable variation in terminology, structure and completeness. Some AEs are documented in the index procedure report, whereas delayed events may appear only in an addendum, emergency department record, inpatient note or discharge summary. Passive reporting may therefore miss cases, while comprehensive chart review requires substantial specialist time and scales poorly in high-volume services. This creates a persistent gap between the need for systematic surveillance and the resources available to perform it.
Natural language processing (NLP) offers a means of converting narrative clinical documentation into structured data suitable for analysis at scale. Previous studies have used NLP to calculate adenoma detection rates, extract quality indicators from electronic and scanned reports, link colonoscopy findings with pathology records, identify gastrointestinal bleeding in clinical documentation and generate guideline-based follow-up recommendations.3–7 These applications support the feasibility of analysing heterogeneous endoscopy records at scale. By contrast, the use of NLP for broader surveillance of endoscopy-related AEs across different procedure types has received comparatively little attention.
Classifying procedure reports according to AE status presents a more complex problem than extracting a discrete clinical finding. Events are uncommon, documentation is heterogeneous, and the definitive diagnosis may not be recorded until a later clinical encounter. Interpretation of the index procedure report may therefore depend on negation, uncertainty, temporality and indirect procedural or contextual cues. Rule-based systems are transparent and readily auditable, but fixed lexicons may fail when clinicians use unexpected or context-dependent expressions. Transformer encoders can represent contextual language more flexibly than fixed rule sets.8–10 Nevertheless, evidence supporting their use for endoscopy-related AE surveillance remains limited, particularly in non-English clinical settings, and few studies have evaluated performance against a clinician-adjudicated reference standard.11 12
We therefore conducted a retrospective, single-centre diagnostic accuracy study at a tertiary referral centre in Türkiye to develop and compare two NLP approaches—a rule-based algorithm and a fine-tuned Turkish transformer encoder—for classifying free-text index procedure reports against a 30-day clinician-adjudicated reference standard for endoscopy-related AEs. We assessed overall diagnostic accuracy and examined model performance according to AE category, documentation timing and source of classification error. Because both models received only the index procedure report, their intended use was to support clinician-supervised retrospective case identification for research and quality assurance. The study did not evaluate autonomous diagnosis, prospective AE prediction, point-of-care decision support, changes in clinical management or effects on patient outcomes.
Methods
Study design, setting and ethics
We conducted a retrospective, single-centre diagnostic accuracy study using routinely collected endoscopy reports from Kayseri City Hospital, a high-volume tertiary referral centre in Türkiye. The source cohort comprised eligible procedures performed between 1 August 2015 and 31 August 2025. This 10-year interval was selected to accrue sufficient numbers of uncommon AEs across different procedure types and to capture variation in routine documentation practices over time.
The study was reported in accordance with the Strengthening the Reporting of Observational Studies in Epidemiology statement and the Standards for Reporting Diagnostic Accuracy Studies 2015 guideline.13 14 Risk of bias and applicability concerns were assessed using an adapted framework informed by the Quality Assessment of Diagnostic Accuracy Studies–Artificial Intelligence (QUADAS-AI) initiative and structured around the patient-selection, index-test, reference-standard and flow-and-timing domains.15 Lexicon-based candidate enrichment and partial verification of lexicon-negative reports were prespecified as potential sources of selection and verification bias. Because a prespecified subset of the candidate-retrieval lexicon also formed the lexical basis of the rule-based classifier, the potential influence of the enriched sampling strategy on the apparent performance of that model was considered in the risk-of-bias assessment. The rule-based and transformer-based NLP models constituted the index tests, and independent clinician adjudication served as the reference standard. The unit of analysis was the individual procedure report.
Direct patient identifiers were removed before analysis, and coded identifiers were used for patient-level linkage and dataset allocation. Data extraction, clinical annotation, model development and statistical analysis were conducted on an access-restricted, password-protected institutional server. No identifiable patient information was transferred outside the institution, and results were reported only in aggregate form.
Patient and public involvement
Patients and members of the public were not involved in the development of the research question, study design, conduct, analysis, reporting or dissemination plans. The study used retrospectively collected, de-identified clinical records and involved no direct participant recruitment or contact.
Data sources, eligibility and target condition
The source cohort included all consecutive endoscopic procedures performed during the study period with a retrievable free-text report. Reports were obtained from the institutional endoscopy reporting system and included the available ‘Findings’, ‘Procedure Note’ and ‘Conclusion’ fields. Structured metadata extracted alongside the free text included patient age, sex, procedure date, procedure type, care setting and a coded patient identifier. All reports included in the source cohort had retrievable free-text documentation. No irreconcilable duplicate records were identified after reconciliation using procedure identifiers, procedure dates and coded patient identifiers. Multiple procedures from the same patient were retained because the procedure report was the unit of analysis; however, all reports belonging to the same patient were assigned to the same dataset partition to prevent information leakage. The transition from the source cohort through candidate screening, adjudication and dataset allocation is shown in figure 1.
Figure 1. Study flow diagram. All 140 385 eligible endoscopic procedure reports had retrievable free-text documentation, and no irreconcilable duplicate records were identified. Candidate-retrieval lexicon screening identified 1512 lexicon-positive reports, all of which underwent independent dual-clinician adjudication, confirming 1205 AE-positive and 307 AE-negative reports. From the remaining 138 873 lexicon-negative reports, 2000 were randomly sampled without replacement for the same adjudication process, identifying three additional AE-positive and 1997 AE-negative reports. The fully clinician-adjudicated subset therefore comprised 3512 reports. Of the 136 873 lexicon-negative reports remaining after verification sampling, 9821 were selected for model development and assigned provisional negative labels. The other 127 052 reports were not selected, were not assigned provisional labels and were not included in the analytic corpus. Combining the 3512 fully clinician-adjudicated reports with the 9821 provisionally negative reports yielded a final analytic corpus of 13 333 reports. Reports were allocated at patient level to the training (n=9333), validation (n=2000) and locked test (n=2000) datasets. Provisionally labelled reports were restricted to the training and validation datasets. The locked test dataset comprised only fully clinician-adjudicated reports, including 180 AE-positive and 1820 AE-negative reports. The full source cohort of 140 385 reports was not partitioned into the training, validation and test datasets. AE, adverse event.

The primary target condition was the presence of at least one endoscopy-related AE attributable to the index procedure within 30 days. Both NLP models received only the free-text index procedure report as input. For reference-standard assessment, clinicians additionally reviewed available addenda and relevant emergency department, inpatient and follow-up documentation recorded within 30 days. Events documented exclusively after the index procedure were retained in the reference standard and analysed separately as delayed-only events. Accordingly, the task was interpreted as classification of index procedure reports according to 30-day clinician-adjudicated AE status rather than direct extraction of every reference-standard event.
The predefined AE categories were perforation, clinically significant bleeding, post-ERCP pancreatitis, cardiopulmonary or sedation-related events, infection or cholangitis and other unplanned events. Multilabel classification was permitted when more than one category applied. One primary category was selected by adjudicator consensus for descriptive reporting, while additional applicable categories were retained as secondary labels. Events were not dichotomised as major or minor. Instead, severity-related consequences were represented through prespecified indicators of clinical escalation, including unplanned admission, blood transfusion, repeat endoscopy, interventional radiology, surgery, intensive care unit admission and death. Operational definitions and cross-cutting adjudication rules are provided in supplementary table s1.
Candidate screening and reference-standard adjudication
Two gastroenterologists developed a deliberately sensitive candidate-retrieval lexicon using the American Society for Gastrointestinal Endoscopy AE lexicon and iterative review of 500 representative procedure reports.1 The candidate-retrieval lexicon included Turkish and English terms, mixed-language expressions, synonyms, abbreviations, common spelling variants and multiword expressions encountered in routine Turkish endoscopy reports. Refinement continued until review of 100 consecutive development reports yielded no additional candidate-retrieval terms. The final candidate-retrieval lexicon comprised 214 terms. A further 126 contextual cues were developed for use by the rule-based classifier to interpret negation, historical or temporal context, uncertainty, hypothetical or counselling statements and clinical escalation, yielding a combined lexical inventory of 340 entries. Reports used during lexical development were treated as development data and excluded from the locked test set.
The complete inventory of 214 candidate-retrieval terms and 126 contextual rule cues is provided in online supplemental data file 1. The lexical development process, contextual classification framework, rule-family counts and representative examples are presented in supplementary table S2.
The completed candidate-retrieval lexicon was applied to all 140 385 reports in the source cohort and identified 1512 lexicon-positive reports. A lexical match was used only to identify a report for further review and did not itself establish the presence of a confirmed AE. Two board-certified gastroenterologists independently reviewed all 1512 reports, including the index procedure report, available addenda and relevant clinical documentation recorded within 30 days. They independently assigned AE status, primary and secondary AE categories, documentation timing and escalation indicators while blinded to each other’s assessments and to the outputs of both NLP models. Disagreements were resolved by consensus; cases that remained unresolved were reviewed by a third senior endoscopist. This process identified 1205 AE-positive reports.
To identify events potentially missed during candidate screening, a computer-generated simple random sample of 2000 reports was drawn without replacement from the remaining 138 873 lexicon-negative reports and underwent the same independent dual-clinician adjudication process. Three additional AE-positive reports were identified. The fully clinician-adjudicated subset therefore comprised 3512 reports: 1208 AE-positive and 2304 AE-negative reports. Inter-rater agreement before consensus resolution was quantified using Cohen’s κ for binary AE status, with 95% CIs estimated using asymptotic SEs.
Because only a sample of the lexicon-negative stratum underwent clinician adjudication, the identified AE-positive reports were not considered a complete census of AEs in the source cohort. The verification sample was used to assess potentially missed events and the implications of partial verification rather than to support an unweighted cohort-wide incidence calculation.
Dataset construction and NLP models
A further 9821 lexicon-negative reports were selected from the remaining lexicon-negative pool to provide sufficient negative examples for model development. These reports were assigned provisional negative labels based on the absence of a candidate-retrieval lexicon match and were restricted to the training and validation datasets. They were not included in the locked test set or in the calculation of diagnostic accuracy measures. The final analytical corpus comprised 13 333 reports.
Before model fitting, hyperparameter selection or decision-threshold optimisation, a locked test set of 2000 fully clinician-adjudicated reports was reserved using patient-level allocation. The test set contained 180 AE-positive and 1820 AE-negative reports and included no provisionally labelled records. The remaining 11 333 reports in the analytic corpus were allocated at patient level to the training dataset (n=9333) and validation dataset (n=2000). The full source cohort of 140 385 reports was not divided into training, validation and test datasets. All reports belonging to the same patient were retained within a single dataset partition. The locked test data were not accessed during lexical refinement, model fitting, hyperparameter selection, model selection or decision-threshold optimisation. Dataset composition, sampling details and leakage-control procedures are provided in supplementary table s3, panel a.
Report texts were standardised for whitespace and non-informative punctuation while preserving clinically relevant symbols, numerical expressions and mixed Turkish–English or Latin terminology. Texts were segmented into sentences and tokenised. All preprocessing decisions were established using the training and validation datasets and applied unchanged to the locked test dataset.
The rule-based classifier used the candidate-retrieval terms designated for rule-based classification in online supplemental data file 1, together with 126 contextual cues for negation, historical or temporal context, uncertainty, hypothetical or counselling statements and clinical escalation. Negation was handled using a NegEx-style approach adapted to Turkish clinical language within a five-word bidirectional context window.16 Contextual cues modified or supported the interpretation of an existing candidate AE mention and did not independently generate an AE-positive classification. All error-driven rule refinement was restricted to the training and validation datasets, and the final rules were frozen before application to the locked test set. The rule-based model produced a binary AE classification and multilabel predictions for the predefined AE categories. Because the model generated binary rather than continuously ranked outputs, curve-based discrimination measures were not calculated for this model.
BERTurk, a pretrained Turkish transformer encoder, was fine-tuned for binary AE classification and multilabel AE categorisation.17 Training was performed for five epochs using the AdamW optimiser, a learning rate of 2×10⁻⁵, a batch size of 16 and weighted binary cross-entropy loss. Decision thresholds were selected using the validation dataset to maximise the F1 score and were fixed before evaluation in the locked test dataset. The exact pretrained checkpoint, tokenizer, sequence length, truncation strategy, class-weight calculation, decision thresholds, software versions, hardware configuration and random seed are reported in supplementary table s3, panel b. No online learning, post-test refinement, recalibration or periodic model updating was performed. Any future update prompted by changes in terminology, reporting templates or clinical practice would require retraining using newly adjudicated reports and independent revalidation before implementation.
Statistical analysis
Continuous variables are summarised as median and IQR, and categorical variables as number and percentage. No formal a priori sample size calculation was performed. The source cohort included all eligible procedures during the study period, while the locked test-set size was determined by the prespecified dataset-construction strategy and the availability of fully clinician-adjudicated reports.
Diagnostic accuracy was evaluated exclusively in the locked, fully clinician-adjudicated test dataset. Performance measures for both models included sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and F1 score. For the transformer model, which generated continuously ranked probability estimates, the precision–recall area under the curve (PR-AUC) and receiver operating characteristic area under the curve (ROC-AUC) were also calculated. PR-AUC was the primary curve-based measure because of class imbalance.
95% CIs for sensitivity, specificity, PPV and NPV were calculated using the Wilson method. CIs for transformer PR-AUC and ROC-AUC were estimated using 2000 report-level bootstrap resamples. Because multiple reports from the same patient could be included in the test dataset, these report-level intervals may not fully account for within-patient correlation; this limitation was considered when interpreting their precision. Because both models were evaluated on the same reports, the paired difference in sensitivity was assessed using McNemar’s test applied to paired binary detection outcomes among the 180 AE-positive reports in the locked test dataset.
The test dataset was deliberately enriched for AE-positive reports and had an AE prevalence of 9.0%. PPV, NPV and F1 score were therefore interpreted as conditional on the test-set composition and were not treated as population-level estimates. Transformer PR-AUC and ROC-AUC, together with the sensitivity and specificity of both models, were interpreted in light of the enriched case spectrum, lexicon-based candidate selection and partial verification of the lexicon-negative stratum.
Subgroup analyses according to AE category and documentation timing were prespecified. Estimates were reported with numerators, denominators and 95% CIs when the number of reference-positive reports permitted meaningful calculation and were interpreted descriptively. No inferential comparisons were performed within individual subgroups, and no adjustment was made for multiple comparisons.
All eligible reports had retrievable free-text documentation; no imputation was performed. Statistical analyses were conducted using R V.4.3.2 and Python V.3.10. All statistical tests were two-sided, and p values below 0.05 were considered statistically significant. Domain-level judgements from the adapted QUADAS-AI-informed assessment are presented in supplementary table s3, panel c.
Results
Source cohort and reference-standard assembly
During the 10-year study period from 1 August 2015 to 31 August 2025, the source cohort comprised 140 385 endoscopic procedure reports. Baseline patient and procedure characteristics are summarised in table 1. The median patient age was 56 years (IQR 47–64), 44.8% of procedures were performed in female patients and 81.5% were conducted in the outpatient setting. The cohort included 73 400 upper gastrointestinal endoscopies (52.3%), 47 800 colonoscopies (34.1%) and 19 185 ERCP/EUS procedures (13.7%).
Table 1. Characteristics of the source cohort and reference-standard assembly (August 2015 to August 2025).
| Characteristic | Value |
|---|---|
| Source cohort | |
| Total procedures, n | 140 385 |
| Upper gastrointestinal endoscopy, n (%) | 73 400 (52.3) |
| Colonoscopy, n (%) | 47 800 (34.1) |
| ERCP/EUS, n (%) | 19 185 (13.7) |
| Age, median (IQR), years | 56 (47–64) |
| Female sex, % | 44.8 |
| Outpatient setting, % | 81.5 |
| Reference-standard assembly | |
| Lexicon-positive reports, n | 1512 |
| Confirmed AE-positive reports among lexicon-positive reports, n/N (%) | 1205/1512 (79.7) |
| Randomly adjudicated lexicon-negative reports, n | 2000 |
| Additional AE-positive reports identified, n/N (%) | 3/2000 (0.15) |
| Fully clinician-adjudicated reports, n | 3512 |
| Total adjudicated AE-positive reports, n | 1208 |
| Total adjudicated AE-negative reports, n | 2304 |
Demographic and procedural-setting characteristics were not stratified by procedure type. The full source cohort was not clinician-adjudicated, and the observed AE counts should not be interpreted as unweighted cohort-wide or procedure-specific incidence estimates.
AE, adverse event; ERCP, endoscopic retrograde cholangiopancreatography; EUS, endoscopic ultrasound.
Candidate-retrieval lexicon screening identified 1512 reports containing at least one AE-related term. Independent clinician adjudication confirmed at least one endoscopy-related AE in 1205 of these reports, corresponding to a confirmation yield of 79.7% (95% CI 77.6% to 81.6%). Among the 2000 randomly sampled lexicon-negative reports that underwent the same independent dual-clinician adjudication process, three additional AE-positive reports were identified (0.15%, 95% CI 0.05% to 0.44%).
The fully clinician-adjudicated subset therefore comprised 3512 reports, of which 1208 were AE-positive and 2304 were AE-negative. Before consensus resolution, inter-rater agreement for binary AE status was κ=0.91 (95% CI 0.88 to 0.94). Because the remaining lexicon-negative reports did not undergo clinician adjudication, these observed counts were not used to calculate an unweighted AE incidence for the full source cohort. No endoscopy-attributable death was identified within 30 days among the clinician-adjudicated reports.
Model development and test datasets
The analytical corpus comprised 13 333 reports. Patient-level allocation assigned 9333 reports to the training dataset, 2000 to the validation dataset and 2000 to the locked test dataset. The full source cohort of 140 385 reports was not divided into training, validation and test datasets.
The locked test dataset was drawn exclusively from fully clinician-adjudicated reports and contained 180 AE-positive and 1820 AE-negative reports, corresponding to an intentionally enriched AE prevalence of 9.0%. No provisionally labelled report was included in the locked test dataset. The complete transition from the source cohort through candidate screening and clinician adjudication to dataset allocation is shown in figure 1.
Spectrum and documentation timing of adverse events
Among the 1205 AE-positive reports confirmed after lexicon-based candidate screening, the primary AE category was post-ERCP pancreatitis in 540 reports (44.8%), followed by clinically significant bleeding in 301 (25.0%), perforation in 120 (10.0%), infection or cholangitis in 48 (4.0%) and cardiopulmonary or sedation-related events in 48 (4.0%). The remaining 148 reports (12.3%) were classified as other unplanned events figure 2. The three AE-positive reports identified in the lexicon-negative verification sample were included in the overall adjudicated count but were not included in this category distribution, which was restricted to the lexicon-positive set.
Figure 2. Distribution of primary adverse-event categories among the lexicon-positive AE cases. Distribution of primary adverse-event categories among the 1205 AE-positive reports identified through candidate-retrieval lexicon screening. Post-ERCP pancreatitis accounted for 44.8% of events, followed by clinically significant bleeding (25.0%), other unplanned events (12.3%), perforation (10.0%), infection or cholangitis (4.0%) and cardiopulmonary or sedation-related events (4.0%). The three AE-positive reports identified in the lexicon-negative verification sample were included in the total adjudicated AE count but were not included in this category distribution. These unweighted counts should not be interpreted as overall or procedure-specific incidence estimates. AE, adverse event; ERCP, endoscopic retrograde cholangiopancreatography.

Of the 1205 AE-positive reports identified in the lexicon-positive set, 94 followed upper gastrointestinal endoscopy, 150 followed colonoscopy and 961 followed ERCP/EUS. These figures describe the distribution of confirmed AEs identified through candidate screening and should not be interpreted as procedure-specific incidence estimates.
Documentation of the definitive AE was present in the index procedure report in 1024 cases (85.0%). The remaining 181 events (15.0%) were identified only in subsequent clinical records. Because these delayed-only events fell outside the models’ input window, they represented a prespecified source of false-negative classification.
Overall model performance
The rule-based model correctly identified 152 of the 180 AE-positive reports and 1811 of the 1820 AE-negative reports in the locked test dataset. Sensitivity was 84.4% (95% CI 78.4% to 89.0%) and specificity was 99.5% (95% CI 99.1% to 99.7%). The PPV was 94.4% (95% CI 89.7% to 97.0%), the NPV was 98.5% (95% CI 97.8% to 98.9%) and the F1 score was 0.89.
The transformer model correctly identified 166 of the 180 AE-positive reports and 1791 of the 1820 AE-negative reports. Sensitivity was 92.2% (95% CI 87.4% to 95.3%) and specificity was 98.4% (95% CI 97.7% to 98.9%). PPV was 85.1% (95% CI 79.5% to 89.4%), NPV was 99.2% (95% CI 98.7% to 99.5%) and the F1 score was 0.89.
The transformer model achieved higher sensitivity than the rule-based model (McNemar’s test, p=0.01), at the cost of lower specificity and PPV, while the F1 scores were identical. For the transformer model, the PR-AUC was 0.91 (95% CI 0.87 to 0.94) and the ROC-AUC was 0.98 (95% CI 0.96 to 0.99). Curve-based discrimination measures were not calculated for the rule-based model because it generated binary classifications rather than continuously ranked scores. Performance metrics are summarised in table 2, and the selected operating point of the transformer model at the fixed decision threshold of 0.48 is shown in figure 3.
Table 2. Overall diagnostic performance of the NLP models in the locked, fully clinician-adjudicated test set (n=2000).
| Metric | Rule-based NLP | Transformer-based NLP | P value |
|---|---|---|---|
| Sensitivity, % (95% CI) | 84.4 (78.4 to 89.0) | 92.2 (87.4 to 95.3) | 0.01 |
| Specificity, % (95% CI) | 99.5 (99.1 to 99.7) | 98.4 (97.7 to 98.9) | — |
| PPV, % (95% CI) | 94.4 (89.7 to 97.0) | 85.1 (79.5 to 89.4) | — |
| NPV, % (95% CI) | 98.5 (97.8 to 98.9) | 99.2 (98.7 to 99.5) | — |
| F1 score | 0.89 | 0.89 | — |
| PR-AUC (95% CI) | — | 0.91 (0.87 to 0.94) | — |
| ROC-AUC (95% CI) | — | 0.98 (0.96 to 0.99) | — |
McNemar’s test was applied to paired binary detection outcomes among the 180 AE-positive reports in the locked test set; the reported p value therefore corresponds to the paired comparison of sensitivity. CIs for sensitivity, specificity, PPV and NPV were calculated using the Wilson method. CIs for transformer PR-AUC and ROC-AUC were estimated using 2000 report-level bootstrap resamples. The rule-based model produced 152 true-positive, 28 false-negative, 1811 true-negative and 9 false-positive classifications. The corresponding counts for the transformer model were 166, 14, 1791 and 29. Curve-based discrimination measures were not calculated for the rule-based model because it produced binary classifications rather than continuously ranked scores. PPV, NPV, F1 score and transformer PR-AUC are conditional on the enriched test-set composition.
AE, adverse event; NLP, natural language processing; NPV, negative predictive value; PPV, positive predictive value; PR-AUC, precision–recall area under the curve; ROC-AUC, receiver operating characteristic area under the curve.
Figure 3. Selected operating point of the transformer-based NLP model in the locked test set. At the fixed decision threshold of 0.48, selected in the validation dataset before locked-test evaluation, the transformer model achieved a recall of 0.922 (166/180) and a precision of 0.851 (166/195) in the locked, fully clinician-adjudicated test set (n=2000; 180 AE-positive reports). The horizontal dashed line represents the no-skill precision corresponding to the test-set AE prevalence of 0.09. AE, adverse event; NLP, natural language processing.

Because the test dataset was deliberately enriched for AE-positive reports, PPV, NPV and F1 score estimates for both models, together with the transformer PR-AUC, are conditional on the test-set composition and should not be interpreted as population-level values for the full source cohort. Sensitivity and specificity for both models and the transformer ROC-AUC should likewise be interpreted in light of the enriched case spectrum, lexicon-based candidate selection and partial verification of the lexicon-negative stratum.
Subgroup analyses
Table 3 presents prespecified subgroup sensitivity estimates according to AE category and documentation timing, with numerators, denominators and 95% CIs. Estimates based on small numbers were interpreted cautiously because of their wide CIs.
Table 3. Subgroup sensitivity of NLP models by adverse-event category and documentation timing in the locked, fully clinician-adjudicated test set (n=2000).
| Subgroup | Rule-based NLP, n/N, % (95% CI) | Transformer-based NLP, n/N, % (95% CI) |
|---|---|---|
| By AE category | ||
| Perforation | 18/18, 100 (82 to 100) | 18/18, 100 (82 to 100) |
| Clinically significant bleeding | 44/45, 98 (88 to 100) | 45/45, 100 (92 to 100) |
| Post-ERCP pancreatitis | 56/80, 70 (59 to 79) | 66/80, 82 (73 to 89) |
| Cardiopulmonary or sedation-related events | 4/7, 57 (25 to 84) | 6/7, 86 (49 to 97) |
| By documentation timing | ||
| AE documented in the index procedure report | 144/153, 94 (89 to 97) | 150/153, 98 (94 to 99) |
| Delayed-only AE documented exclusively in subsequent clinical records | 8/27, 30 (16 to 49) | 16/27, 59 (41 to 76) |
Wilson 95% CIs are reported. Percentages are rounded to the nearest whole number. Subgroup analyses were prespecified and interpreted descriptively; no inferential comparisons were performed within individual subgroups, and no adjustment was made for multiple comparisons. Estimates based on small numbers should be interpreted cautiously because of their wide CIs. Sensitivity estimates for infection or cholangitis and other unplanned events were not reported separately because of the limited number of reference-positive reports in these categories.
AE, adverse event; ERCP, endoscopic retrograde cholangiopancreatography; NLP, natural language processing.
In descriptive analyses, sensitivity was numerically higher with the transformer model for post-ERCP pancreatitis (82% vs 70%) and cardiopulmonary or sedation-related events (86% vs 57%). For AEs documented in the index procedure report, sensitivity was 94.1% with the rule-based model and 98.0% with the transformer model. For delayed-only events, sensitivity decreased to 29.6% and 59.3%, respectively. No inferential comparisons were performed within individual subgroups.
Error analysis
The rule-based model produced 28 false-negative and 9 false-positive classifications. The transformer model produced 14 false-negative and 29 false-positive classifications.
Of the 180 AE-positive reports in the locked test dataset, 153 contained direct AE documentation in the index procedure report, whereas 27 were delayed-only events identified through subsequent clinical records. For reports with index-report documentation, sensitivity was 94.1% (144/153) for the rule-based model and 98.0% (150/153) for the transformer model. For delayed-only events, the corresponding estimates were 29.6% (8/27) and 59.3% (16/27), respectively.
Although the definitive AE diagnosis was documented only in subsequent records for delayed-only events, some index reports contained indirect procedural or contextual cues—such as difficult cannulation, repeated pancreatic duct instrumentation or anatomical complexity—that resulted in a positive model classification.
False-positive classifications included negated, hypothetical and historical statements that resembled current AE documentation, including statements describing the absence of bleeding or discussion of potential perforation risk. False-negative classifications additionally included reports containing atypical terminology, abbreviations or mixed Turkish–English expressions.
Risk-of-bias assessment
The adapted QUADAS-AI-informed assessment identified some concerns in the patient-selection and flow-and-timing domains, primarily because the diagnostic-accuracy sample was enriched through lexicon-based candidate retrieval, only a random subset of lexicon-negative reports underwent clinician adjudication and the index tests received a narrower information window than the reference standard.
Because a prespecified subset of the candidate-retrieval lexicon also formed the lexical basis of the rule-based classifier, the enriched sampling strategy may have influenced the apparent performance of that model. Risk of bias was judged to be low in the index-test and reference-standard domains. Applicability concerns were judged to be low in the patient-selection, index-test and reference-standard domains; applicability was not assessed for flow and timing. Domain-level judgements and their supporting rationales are presented in supplementary table s3, panel c.
Discussion
In this retrospective, single-centre diagnostic accuracy study, both NLP approaches classified free-text index procedure reports against a 30-day clinician-adjudicated reference standard for endoscopy-related AEs with high specificity, although their error profiles differed. In the locked, fully clinician-adjudicated test set, the transformer model increased sensitivity from 84.4% to 92.2% and reduced false-negative classifications from 28 to 14. This gain was accompanied by lower specificity and PPV, while both models had an identical F1 score of 0.89. Sensitivity was numerically higher with the transformer for post-ERCP pancreatitis and cardiopulmonary or sedation-related events; however, these subgroup findings were descriptive and based on limited numbers. Events documented only in subsequent clinical records remained more difficult to classify because the models received only the index procedure report. Accordingly, the reported performance represents retrospective classification of index procedure reports according to 30-day clinician-adjudicated AE status, rather than prospective identification of previously unrecognised events or demonstrated clinical utility at the point of care.
Several features strengthen the study, including the large decade-long source cohort, independent dual-clinician adjudication, patient-level dataset allocation, use of a locked test set without provisionally labelled reports and direct comparison of a transparent rule-based system with a Turkish transformer model. The sampling strategy nevertheless imposes important limitations on interpretation. Lexicon-based candidate enrichment and partial verification of the lexicon-negative stratum introduced risks of selection and verification bias. Because a prespecified subset of the candidate-retrieval lexicon also formed the lexical basis of the rule-based classifier, the enriched sampling strategy may have influenced the apparent performance of that model. The identification of three AE-positive reports among 2000 randomly sampled lexicon-negative reports demonstrates that candidate screening was not perfectly sensitive. Although the low yield in the verification sample indicates effective candidate enrichment, the large lexicon-negative stratum means that even a small residual event rate could correspond to a meaningful number of unascertained AEs. Accordingly, the 1205 AEs confirmed among lexicon-positive reports should not be regarded as a complete census of events in the source cohort, and neither overall nor procedure-specific incidence can be estimated directly from these unweighted counts.
Provisionally negative reports used during training and validation may also have contained unrecognised events, introducing label noise and potentially biasing model development and threshold selection towards the negative class. The deliberately enriched AE prevalence in the test set means that PPV, NPV and F1 scores for both models, together with the transformer PR-AUC, are conditional on the test-set composition. Sensitivity, specificity and the transformer ROC-AUC should likewise be interpreted in light of the enriched case spectrum, lexicon-based candidate selection and partial verification process. Small numbers in several subgroups resulted in wide CIs, while report-level bootstrap intervals may not have fully accounted for within-patient correlation. Finally, the single-centre design limits generalisability, and the decade-long study period does not itself establish temporal transportability because dataset allocation was performed at the patient level rather than by calendar period.
Most previous NLP applications in gastroenterology have focused on extracting structured findings and quality indicators from endoscopy narratives rather than identifying complications. NLP approaches have been validated for adenoma detection rate calculation,3 quality-indicator extraction from scanned or non-native electronic reports,4 linkage of colonoscopy and pathology narratives5 and automated generation of guideline-based surveillance recommendations.7 Transparent frameworks for endoscopy data extraction have also highlighted the importance of reproducibility and explainability.18 AE classification presents additional challenges because events are uncommon, context-dependent and vulnerable to negation, temporality and hypothetical-language artefacts. Transformer-based approaches are increasingly being studied for clinically relevant event identification, including gastrointestinal bleeding detection.6 Multimodal systems combining deep learning and NLP have also been developed to identify high-risk patients undergoing upper endoscopy and assign surveillance intervals across multiple centres.19 However, these applications address risk-stratified surveillance rather than the detection of procedure-related AEs. Systematic reviews continue to identify heterogeneous definitions, incomplete methodological reporting and limited external validation, particularly in non-English settings.11 12
The present study extends this literature by comparing rule-based and transformer-based classification of Turkish-language index procedure reports against a clinician-adjudicated 30-day reference standard covering multiple endoscopy-related AEs. The transformer’s higher sensitivity may reflect greater flexibility in representing variable and context-dependent clinical phrasing, but this interpretation is inferential and was not directly tested. Its lower specificity and PPV, together with the identical F1 scores, indicate that it was not uniformly superior to the rule-based approach. A rule-based approach may remain preferable when interpretability and a low false-positive burden are priorities, whereas the transformer may be preferable when minimising false-negative classifications is considered more important.
Systematic identification and review of AEs is central to contemporary endoscopy quality assurance.1 2 The most defensible application of the present pipeline is clinician-supervised retrospective case identification for research and quality assurance, rather than autonomous diagnosis, prospective prediction of postprocedural complications or point-of-care decision support. Such a system could be evaluated as a method for prioritising reports for manual review, but the present study did not measure review time, workload reduction, changes in clinical management or patient outcomes. The models cannot establish procedural causality, identify events that leave no signal in the index procedure report or replace clinician adjudication. Consequently, no conclusions can be drawn regarding clinical effectiveness, workflow efficiency, safety or patient benefit.
Prospective implementation would require evaluation of the number of reports flagged, manual review workload, false-positive burden, operating thresholds and clinician responses, together with ongoing monitoring for changes in terminology, reporting templates and case mix. Performance gains must also be balanced against transparency, governance and clinical accountability.20 21
Future research should prioritise multicentre external validation and prospective temporal validation using later reports that played no role in model development. Linking endoscopy reports with emergency department, inpatient, laboratory, imaging and readmission data may improve ascertainment of delayed complications and help distinguish extraction of documented events from prospective prediction of postprocedural AEs. Larger independently adjudicated test sets will be required to evaluate uncommon AE categories and lower-risk procedure groups with adequate precision.
Prospective implementation studies should report the number of records flagged, manual review time, PPV at routine clinical prevalence, missed-event frequency, clinician acceptance and resulting changes in quality-assurance processes. Standardised AE definitions and severity criteria would improve comparison across centres, and any future model update should use newly adjudicated data followed by independent revalidation before implementation. Similar principles apply to clinical NLP in other procedural fields and to the broader integration of artificial intelligence in endoscopy.21–23
Conclusion
In this single-centre diagnostic accuracy study, both NLP approaches classified free-text index procedure reports against a 30-day clinician-adjudicated reference standard for endoscopy-related AEs with high specificity. The transformer model improved sensitivity over a transparent rule-based baseline but generated more false-positive classifications and a lower PPV, while both models had the same F1 score.
These findings support clinician-supervised retrospective case identification for research and quality assurance but do not establish the safety or effectiveness of autonomous use, prospective AE prediction, point-of-care clinical application or improvement in patient outcomes. Multicentre external validation, prospective temporal validation and formal implementation studies are required before routine use.
Supplementary material
Acknowledgements
The authors thank the endoscopy unit staff for their contributions to routine clinical documentation and the institutional information technology team for assistance with secure data extraction.
Footnotes
Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial or not-for-profit sectors.
Prepublication history and additional supplemental material for this paper are available online. To view these files, please visit the journal online (https://doi.org/10.1136/bmjopen-2026-118459).
Provenance and peer review: Not commissioned; externally peer reviewed.
Patient consent for publication: Not applicable.
Ethics approval: Ethics approval was obtained from the Kayseri City Hospital Non-Interventional Clinical Research Ethics Committee (Decision No: 586; 26 September 2025). Given the retrospective design and use of de-identified data, the requirement for individual informed consent was waived.
Data availability free text: De-identified report-level data are not publicly available because free-text clinical records may retain a residual risk of patient re-identification. The complete candidate-retrieval lexicon and contextual rule-cue inventory are provided in Supplementary Data File 1. Requests for additional de-identified data will be considered by the corresponding author, subject to ethics committee and institutional approval and execution of an appropriate data-use agreement.
Patient and public involvement: Patients and/or the public were not involved in the design, or conduct, or reporting, or dissemination plans of this research.
Data availability statement
Data are available upon reasonable request.
References
- 1.Cotton PB, Eisen GM, Aabakken L, et al. A lexicon for endoscopic adverse events: report of an ASGE workshop. Gastrointest Endosc. 2010;71:446–54. doi: 10.1016/j.gie.2009.10.027. [DOI] [PubMed] [Google Scholar]
- 2.Rex DK, Anderson JC, Butterly LF, et al. Quality indicators for colonoscopy. Gastrointest Endosc. 2024;100:352–81. doi: 10.1016/j.gie.2024.04.2905. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Tinmouth J, Swain D, Chorneyko K, et al. Validation of a natural language processing algorithm to identify adenomas and measure adenoma detection rates across a health system: a population-level study. Gastrointest Endosc. 2023;97:121–9. doi: 10.1016/j.gie.2022.07.009. [DOI] [PubMed] [Google Scholar]
- 4.Laique SN, Hayat U, Sarvepalli S, et al. Application of optical character recognition with natural language processing for large-scale quality metric data extraction in colonoscopy reports. Gastrointest Endosc. 2021;93:750–7. doi: 10.1016/j.gie.2020.08.038. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Vithayathil M, Smith S, Goryachev S, et al. Development of a Large Colonoscopy-Based Longitudinal Cohort for Integrated Research of Colorectal Cancer: Partners Colonoscopy Cohort. Dig Dis Sci. 2022;67:473–80. doi: 10.1007/s10620-021-06882-x. [DOI] [PubMed] [Google Scholar]
- 6.Zheng NS, Keloth VK, You K, et al. Detection of Gastrointestinal Bleeding With Large Language Models to Aid Quality Improvement and Appropriate Reimbursement. Gastroenterology. 2025;168:111–20. doi: 10.1053/j.gastro.2024.09.014. [DOI] [PubMed] [Google Scholar]
- 7.Karwa A, Patell R, Parthasarathy G, et al. Development of an Automated Algorithm to Generate Guideline-based Recommendations for Follow-up Colonoscopy. Clin Gastroenterol Hepatol. 2020;18:2038–45. doi: 10.1016/j.cgh.2019.10.013. [DOI] [PubMed] [Google Scholar]
- 8.Yang X, Chen A, PourNejatian N, et al. A large language model for electronic health records. NPJ Digit Med. 2022;5:194. doi: 10.1038/s41746-022-00742-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Gu Y, Tinn R, Cheng H, et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans Comput Healthcare. 2022;3:1–23. doi: 10.1145/3458754. [DOI] [Google Scholar]
- 10.Chen Q, Hu Y, Peng X, et al. Benchmarking large language models for biomedical natural language processing applications and recommendations. Nat Commun. 2025;16:3280. doi: 10.1038/s41467-025-56989-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Omar M, Nassar S, SharIf K, et al. Emerging applications of NLP and large language models in gastroenterology and hepatology: a systematic review. Front Med (Lausanne) 2024;11:1512824. doi: 10.3389/fmed.2024.1512824. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Stammers M, Ramgopal B, Owusu Nimako A, et al. A foundation systematic review of natural language processing applied to gastroenterology & hepatology. BMC Gastroenterol. 2025;25:58. doi: 10.1186/s12876-025-03608-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.von Elm E, Altman DG, Egger M, et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Ann Intern Med. 2007;147:573–7. doi: 10.7326/0003-4819-147-8-200710160-00010. [DOI] [PubMed] [Google Scholar]
- 14.Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. doi: 10.1136/bmj.h5527. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Sounderajah V, Ashrafian H, Rose S, et al. A quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies: QUADAS-AI. Nat Med. 2021;27:1663–5. doi: 10.1038/s41591-021-01517-0. [DOI] [PubMed] [Google Scholar]
- 16.Chapman WW, Bridewell W, Hanbury P, et al. A simple algorithm for identifying negated findings and diseases in discharge summaries. J Biomed Inform. 2001;34:301–10. doi: 10.1006/jbin.2001.1029. [DOI] [PubMed] [Google Scholar]
- 17.Schweter S. Zenodo; [DOI] [Google Scholar]
- 18.Fevrier HB, Liu L, Herrinton LJ, et al. A Transparent and Adaptable Method to Extract Colonoscopy and Pathology Data Using Natural Language Processing. J Med Syst. 2020;44:151. doi: 10.1007/s10916-020-01604-8. [DOI] [PubMed] [Google Scholar]
- 19.Li J, Hu S, Shi C, et al. A deep learning and natural language processing-based system for automatic identification and surveillance of high-risk patients undergoing upper endoscopy: A multicenter study. EClinicalMedicine. 2022;53:101704. doi: 10.1016/j.eclinm.2022.101704. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Cascella M, Semeraro F, Montomoli J, et al. The Breakthrough of Large Language Models Release for Medical Applications: 1-Year Timeline and Perspectives. J Med Syst. 2024;48:22. doi: 10.1007/s10916-024-02045-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Parasa S, Berzin T, Leggett C, et al. Consensus statements on the current landscape of artificial intelligence applications in endoscopy, addressing roadblocks, and advancing artificial intelligence in gastroenterology. Gastrointest Endosc. 2025;101:2–9. doi: 10.1016/j.gie.2023.12.003. [DOI] [PubMed] [Google Scholar]
- 22.Soffer S, Klang E, Shimon O, et al. Deep learning for wireless capsule endoscopy: a systematic review and meta-analysis. Gastrointest Endosc. 2020;92:831–9. doi: 10.1016/j.gie.2020.04.039. [DOI] [PubMed] [Google Scholar]
- 23.Le KDR, Tay SBP, Choy KT, et al. Applications of natural language processing tools in the surgical journey. Front Surg. 2024;11:1403540. doi: 10.3389/fsurg.2024.1403540. [DOI] [PMC free article] [PubMed] [Google Scholar]
