The question “Was this research paper written by AI?” sounds simple, but current evidence shows that the answer is rarely reliable when it is reduced to a single detector score. Generative AI can contribute to academic writing in many ways: drafting, rewriting, translation, summarisation, language polishing, idea organisation and partial paragraph generation. That means real manuscripts increasingly sit on a spectrum between fully human-written and fully machine-generated text.
For authors, editors and publishers, the more useful question is often broader: does the manuscript contain signals that warrant closer review, and does the scholarship behind the text hold up when its references, claims, scientific reasoning and evidence are examined?
AI detection should be treated as a signal, not a verdict
Even major commercial systems explicitly caution against using AI detection as the sole basis for adverse decisions. A responsible scholarly workflow combines automated indicators with evidence that can be independently reviewed.
How AI detectors typically work
Most AI-writing detectors do not identify a hidden marker that proves who authored a manuscript. Instead, they analyse statistical and linguistic patterns that may be more common in machine-generated writing than in human writing. Different products use different models, training data and thresholds, which is one reason their outputs can differ for the same document.
Turnitin describes its own AI-writing result as an assessment of text that may have been prepared using a generative AI tool. Importantly, Turnitin also states that its model can misidentify human-written, AI-generated and AI-paraphrased text and should not be used as the sole basis for adverse action. Its guidance calls for further scrutiny and human judgment.
Why research papers are especially difficult
Scientific writing has characteristics that can complicate AI detection. Methods sections often use repeated structures. Technical terminology constrains word choice. Academic conventions encourage formal, impersonal phrasing. Review articles may summarise many studies using similar sentence patterns. Authors writing in a second language may also use more regular syntax or formulaic academic expressions.
A 2026 study in the International Journal for Educational Integrity evaluated Turnitin and Originality on 192 texts and reported overall accuracy of 0.61 and 0.69 respectively. Both systems struggled with hybrid human-AI texts, and detector performance was weaker on scientific writing than on humanities writing. The authors concluded that detector outputs are better treated as supplementary indicators than as sole evidence for misconduct decisions.
Another 2026 comparison examined GPTZero, Pangram, Copyleaks and Turnitin across fully human, fully AI-generated, hybrid and humanised-AI academic papers. Performance varied substantially across systems and document types. Pangram performed best in that particular benchmark, but the authors still recommended that AI detection be implemented within a broader evaluation strategy rather than used as sole evidence in high-stakes decisions.
Hybrid authorship changes the problem
The binary labels “human” and “AI” are becoming less representative of real scholarly workflows. A researcher may write the scientific content, use an AI system to improve grammar, manually rewrite the output, add new references, and later ask another tool to shorten the discussion. A detector may then be asked to assign one percentage to a document created through multiple stages of human and machine intervention.
Hybrid authorship is precisely where multiple recent studies report difficulty. This matters for publishers because many institutional and journal policies distinguish between acceptable disclosed uses of AI and unacceptable uses. Detecting that some text statistically resembles AI output is therefore not the same as establishing whether a policy has been breached.
False positives and multilingual writing
False positives deserve particular attention in scholarly publishing because an incorrect AI flag can affect an author’s reputation and editorial treatment. A July 2026 article in Assessing Writing examined Turnitin in multilingual writing contexts and highlighted concerns that syntactic regularity and formulaic phrasing in English as an additional language can increase the risk of misclassification. The authors argued for caution and for treating automated scores as a cue for further examination rather than as proof.
This does not mean AI detection has no value. It means the evidentiary standard should match the consequence of the decision. Screening a document for closer editorial attention is different from accusing an author of misconduct.
What should an editor examine after an AI flag?
For scholarly manuscripts, the text itself is only one layer. A more informative review asks whether the manuscript’s scholarly infrastructure is reliable. Four questions are particularly useful:
Are the references real?
Verify that cited publications exist and that titles, authors, journals, years and DOIs correspond to identifiable records.
Do citations support the claims?
A genuine paper can still be cited inaccurately. Check whether the source supports the specific statement made around it.
Is the science coherent?
Examine mechanisms, terminology, classifications, causal claims and technical interpretations for factual consistency.
Is the evidence sufficiently developed?
Look for quantitative results, comparisons, limitations and distinctions between preliminary and established evidence.
How EAIPAM approaches AI detection for scholarly manuscripts
EAIPAM uses potential AI-mediated drafting as a reason for structured scholarly review rather than as a standalone conclusion. Its assessment framework examines Reference Integrity, Citation-to-Claim Alignment, Scientific Accuracy and Evidence Depth alongside AI-involvement indicators. The aim is to provide evidence that an author, editor or publisher can inspect and interpret rather than relying on an AI percentage alone.
This distinction is important. A manuscript can be human-written and still contain unreliable references or unsupported scientific claims. Conversely, a manuscript may have involved permitted AI-assisted language editing while retaining sound scholarship. A useful integrity review therefore examines both the writing signals and the evidence behind the writing.
Explore evidence-backed AI detection for scholarly manuscripts
Practical guidance for authors
Authors should retain drafts and source materials, verify every reference, check that citations genuinely support the associated claims, and follow the AI-use disclosure policy of the target journal or publisher. If AI has been used for language support, translation, summarisation or drafting, the relevant policy—not a detector percentage—should determine whether disclosure is required.
Practical guidance for editors and publishers
Editorial teams should define in advance what an AI-detection result can and cannot trigger. A flag may justify additional review, reference verification or a request for clarification. It should not automatically be treated as proof of authorship, intent or research misconduct. Decisions should be documented, proportionate and based on evidence that can be reviewed independently.
Sources and further reading
Turnitin — Using the AI Writing Report. Turnitin explains the interpretation and limitations of its AI Writing Report and states that it should not be the sole basis for adverse action.
Hadra, Cambridge & Mesbah (2026) — Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity.
Who wrote this? Evaluating the reliability of AI detection tools in higher education (2026). International Journal for Educational Integrity.
Automated integrity or automated injustice? Turnitin’s AI detection in multilingual writing contexts (2026). Assessing Writing.
Responsible-use note: EAIPAM findings are intended to support responsible scholarly and editorial review. They do not independently establish AI authorship, fabrication, plagiarism, misconduct or other wrongdoing. Final interpretation remains with the responsible author, editor, publisher, institution or authorised decision-maker.