No one needs to be told that AI is moving fast and disrupting working practices in all manner of ways. We are not opposed to the responsible use of AI and we are all using it in our daily lives.
It has become clear that we need to develop our policy for the use of large language model (LLM) AI-generated text in articles for BJGP.
We have seen a significant uptick in submissions of short article types in recent months (Editorials, Letters, Analysis, Clinical Practice and Viewpoint/Opinion), often with single authors, and we believe this is a direct consequence of the widening availability of LLMs. One might argue this is a positive — it means that people with interesting ideas who are less certain of how to write and construct an article are able to produce publishable material. There is potential to hear from voices who are unfamiliar and less confident navigating the processes of academic publishing. There is still scope to use AI to help develop ideas and shape articles. However, we are under no illusions that we are receiving articles with high percentages of AI-generated text, often described as AI ‘slop’, where the LLMs generate all the text with insufficient or no human input.
If we are to maintain trust in the integrity of BJGP content we need to take immediate action.
This requirement and other policy changes below are coming in with immediate effect — this will also apply to papers in the pipeline at various stages. We are implementing changes in several areas related to AI: AI detection; research articles; AI and peer review: and fabricated and hallucinated references.
AI-detection
We will be using AI-text detection as part of the editorial decision-making process across BJGP.org and BJGP Life. This may, in the future, be integrated into the ScholarOne submission platform but initially we will make use of a third-party commercial service. This service undertakes not to use materials for training purposes and is compliant with GDPR.
We will be asking for more detailed declarations about the use of AI and we will facilitate this by providing a tiered set of options to aid authors.
We understand that AI detectors are not perfect and these will be used as tools alongside editorial judgement. However, they are improving — for instance, Pangram is a well-known platform and they claim a false positive rate as low as 0.0041% across a million English-language texts. This increases to around 0.02% for academic writing. Minimising a type 1 error of this nature is exactly where we want to prioritise diagnostic accuracy in these circumstances. These figures are based on the company’s own data but there have been some small independent evaluations of Pangram and they have been encouraging1,2. In the future, AI detection may get easier as the EU has required AI companies to ‘watermark’ their generated text and Anthropic, the company behind Claude, announced in August 2026 that they would be introducing this3.
Research articles
The pipeline for research is slower than that for opinion articles. It is a certainty that AI will be used as part of the process of research and, as per our current policy of transparency, that use will need to be detailed in the methods section of paper. It is inevitable that the writing of the prose in research papers will also be affected and the processes outlined in this article will also be applicable to research papers. It matters that sections such as abstracts, introductions, and the discussion are crafted with care. It is not acceptable that these sections are written by AI and we expect the prose to be written by humans in the main body of research articles. (This will not necessarily apply to supplementary files.)
AI and peer review
Peer reviewers are using AI. Researchers know this and are already having the experience of receiving detailed peer reviews they suspect are AI-generated. Evidence from a Frontiers survey of 1600 academics (across 111 countries) found that more than half of researchers were using AI for peer review. At BJGP we have had a policy that has said peer reviewers should not use AI. However, in effect, this policy can’t be enforced and the evidence suggests is being widely ignored.
We understand that peer reviewers may use AI as tools, for example to check references, and we will advise them that they must be mindful of the need not to breach the confidentiality of manuscripts. Peer reviewers will be asked to make full disclosures of AI use and we will ask them to write the reviews themselves and not use generative AI. From now on, we will be screening peer reviews that raise concerns for AI-generated text and we reserve the right to send peer reviews back or even to strike them out. We understand this may result in delays for papers but we will work to reduce this impact.
Fabricated and hallucinated references
This is a serious problem in the academic literature. Fiorillo described a typology with three forms of problematic references including fully fabricated reference, authentic references with corrupted data, and ‘chimeric’ references which are a mash-up of different authentic references4. We conducted an internal audit and searched for fabricated and hallucinated references in the BJGP going back to Jan 2025, using the same methodology as Topaz5. We found none. Opinion is divided in the scientific community on how to handle these references: is it always serious academic conduct or can these be handled with corrections? 6 Some argue that once a fabricated reference has been identified then all trust in the content of the article drains away. Clearly though, some context is needed in these decisions.
At the BJGP, accepted papers are published immediately as ‘author accepted manuscripts’ and we have identified a small window where fabricated references can be published in the literature with a DOI. Many reviewers do check references but, as per Fiorillo’s recommendations, we will be using AI to search reference lists to flag any potential concerns that need immediate attention. In the future, this is likely to become part of the automatic checks at submission points.
Summary
We have always been clear that authors and reviewers are responsible and accountable for the work they submit. The nature of the changes as a consequence of increasing use of AI means we need, at this time, to introduce additional safeguards. The BJGP is a non-profit scholarly journal publishing high-quality research related to primary care. This current shift in editorial policy is being implemented, with immediate effect, to maintain the integrity of the research and commentary that has given the BJGP its reputation as a trusted source for primary care clinicians, researchers, and policymakers.
We will be updating our AI policies on BJGP.org and on ScholarOne in the coming weeks. We understand that the rapidly evolving use of AI will, inevitably, require future policy changes and we will keep you informed.
References
- Van Vlasselaer M, Van Droogenbroeck F, Spruyt B. Who wrote this? Evaluating the reliability of AI detection tools in higher education. Int J Educ Integr. 2026 Jun 29;22(1):16. doi:10.1007/s40979-026-00226-w
- Third-Party Pangram Evaluations [Internet]. [cited 2026 Sep 30]. Available from: https://www.pangram.com/blog/third-party-pangram-evals
- How Claude’s text watermarking works [Internet]. 2026 [cited 2026 Sep 30]. Available from: https://www.anthropic.com/news/claude-text-watermark
- Fiorillo L. Confabulated references in the age of AI: contamination of the biomedical scientific literature. Explor Med. 2026 Mar 4;7:1001385. doi:10.37349/emed.2026.1001385
- Topaz M, Roguin N, Gupta P, Zhang Z, Peltonen LM. Fabricated citations: an audit across 2·5 million biomedical papers. The Lancet. 2026 May;407(10541):1779–81. doi:10.1016/S0140-6736(26)00603-3
- Bauchner H, Frederick P R. Fabricated references: a new threat to editorial integrity. The Lancet. 2026 May 9;407(10541):1765–6.
Featured photo by BoliviaInteligente on Unsplash
