Siler (1) estimates that by 2025 roughly 57% of published articles showed large language model (LLM) influence, measured by each article’s rate of 228 “focal words” (e.g., delve, intricate) selected for their post-2022 growth. The study also reports systematic variation across disciplines, publishers, regions, institutions, and gender. The underlying lexical shift is real and worth studying. But three features of the analysis keep it from supporting its core claims. The headline statistic, the detector beneath it, and the subgroup regression each support something weaker than reported.
First, the “57%” is not an article-level prevalence. For each year the figure is the integral of the positive part of the difference between that year’s focal-word-rate density and the pooled 2020–2022 baseline, ∫ max{ft(r)−f0(r), 0} dr. For normalized densities this equals the total-variation distance between the two distributions: It measures how far the distribution of focal-word rates has moved, and it equals a fraction of articles only under a two-component mixture in which an unaffected component is unchanged from the baseline and an influenced component is separable from it. The paper supplies no such structure; under its own “osmosis” channel, diffuse ambient exposure can shift the whole distribution rightward without partitioning articles into a discrete influenced class to count. Applied to the baseline years 2020–2022, the same estimator already returns 1.5 to 2.6% “AI-assisted,” an uncorrected baseline the statistic never benchmarks against.
Second, the focal-word list is not a validated instrument. The words were selected because their post-2022 rise was large, and articles’ use of that selected high-growth vocabulary is then used to quantify LLM influence; the list is never checked against text of known origin—human-written, LLM-assisted, LLM-generated, or human-edited nonnative English. Figure 1’s comparison, showing that focal words grew much more rapidly than other words, is circular, since the focal set is, by construction, the high-growth tail. The procedure identifies a lexical shift; it does not calibrate that shift to article-level provenance.
Third, the difference-in-differences coefficients for field and publisher (figure 5 and SI Appendix, Table S2) are not identified. The model includes field-by-year and publisher-by-year fixed effects together with post-by-field and post-by-publisher interactions. But Postt = 1{t ≥ 2023} = Σs ∈{2023,2024,2025} 1{t = s}, so for any field f, 1{F = f} · Postt = Σs 1{F = f}1{t = s}—exactly the field-by-year columns already in the model (likewise for the publisher). These coefficients are aliased; the published values are a minimum-norm solution to a rank-deficient least-squares problem, not estimands the data determine. Separately, even the estimable interactions are not causal. LLM availability is a common post-2022 shock and the reference categories are themselves exposed to it, so the coefficients are differential changes in focal-word rate, not causal effects of region, rank, publisher, or gender on adoption (2).
The paper convincingly documents a large post-2022 change in academic vocabulary. It does not establish the prevalence of LLM use, a validated detector of it, or its institutional determinants. Code reproducing each point above is available (3).
Acknowledgments
Author contributions
C.M.T. and U.B. designed research; performed research; and wrote the paper.
Competing interests
The authors declare no competing interest.
Data, Materials, and Software Availability
Code data have been deposited in a permanent repository (10.5281/zenodo.20645197 (3)).
References
- 1.Siler K., The diffusion of large language models in published academic articles. Proc. Natl. Acad. Sci. U.S.A. (2026), 10.1073/pnas.2605754123. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Roth J., Sant’Anna P. H. C., Bilinski A., Poe J., What’s trending in difference-in-differences? A synthesis of the recent econometrics literature J. Econom. 235, 2218–2244 (2023). [Google Scholar]
- 3.Topaz C. M., Bahl U., Code and data to reproduce ‘Lexical change is not a calibrated measure of LLM prevalence or its determinants.’ Zenodo. 10.5281/zenodo.20645197. Deposited 11 June 2026. [DOI] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Code data have been deposited in a permanent repository (10.5281/zenodo.20645197 (3)).
