Language models shape 89 percent of biomedical papers
Friday 21st August 2026 on 10:45 in
Estonia
Language model-written or edited text appears in 89 percent of new open-access biomedical research articles, according to a study by Belgian and German researchers reported by ERR. The finding has not yet been assessed by independent researchers.
The analysis also found that the influence of artificial intelligence is uneven across research papers. Researchers rely on language models most heavily when interpreting results, while they use them least often to describe research methods.
Language models began spreading widely in late 2022, giving researchers an efficient tool for editing and translating drafts. Researchers whose first language is not English are particularly likely to use the technology for these tasks.
The growing use of artificial intelligence has raised concerns in the academic community, including the risk of fabricated references and more uniform scientific writing. Policymakers therefore need accurate data on how widely language models are being used.
To avoid the conflicting results produced by various text-detection tools, computer scientist Dmitry Kobak of Ghent University and his colleagues analysed the full texts of nearly 1.2 million biomedical articles in the PubMed Central database. The articles were published between 2017 and 2025.
The share of articles produced with artificial intelligence assistance reached 89 percent by December 2025, far exceeding earlier estimates that often put machine involvement at below 20 percent.
Language model use was detected in nearly 70 percent of short passages in discussion sections, where researchers interpret their findings. The corresponding figures were nearly 60 percent in introductions and 46 percent in results sections.
Researchers relied least on language models in descriptions of research methods, where signs of their use appeared in 32 percent of short passages. This difference reflects the need to present laboratory protocols and numerical data with great precision.
At the same time, more than half of the text in methods sections as a whole had received some linguistic polishing from a language model. The researchers said this suggests that models are used in those sections to revise individual passages rather than to generate the text sentence by sentence.
To reach their conclusions, the researchers tracked the frequency of hundreds of individual words across the scientific corpus.