ArXiv submissions show over 30% now read as AI-written

The accelerating use of AI writing tools in academic preprint repositories raises questions about research authenticity and peer review effectiveness.

Abstract illustration showing manuscript pages transitioning to algorithmic patterns
AI-generated illustration · Sylvaris

Sharp Increase in AI-Generated Text

Analysis of new submissions to ArXiv, the widely-used scientific preprint repository, indicates that over 30% of recent papers exhibit characteristics consistent with AI-generated writing. The measurement comes from Unslop, a service focused on detecting AI-generated content in academic contexts.

The shift represents a significant change in how researchers are preparing and submitting scientific manuscripts. ArXiv serves as a primary distribution channel for early-stage research across physics, mathematics, computer science, and related fields before formal peer review.

Implications for Research Quality

The prevalence of AI-assisted writing in academic submissions raises questions about the authenticity of research communication and whether the substance of scientific work is being adequately conveyed. While AI tools can assist with language and structure, their use in scientific writing remains controversial.

Peer reviewers and journal editors now face the additional challenge of distinguishing between AI-enhanced presentation and potential issues with research originality or methodological rigor. The trend may require updated policies on disclosure of AI assistance in manuscript preparation.

Detection Methods and Limitations

Identifying AI-generated academic text relies on pattern recognition and statistical analysis of writing characteristics. However, as language models improve and researchers become more sophisticated in their use, distinguishing human from machine-assisted writing becomes increasingly difficult.

The 30% figure represents submissions that read as AI-written based on current detection methods, but does not definitively prove AI authorship or necessarily indicate problems with the underlying research quality. The boundary between legitimate writing assistance and problematic over-reliance on automated text generation remains contested.

sources
more in Artificial Intelligence
Text-to-SQL benchmarks fail to address real-world data store complexities AI code generation tools struggle with messy production databases that lack the clean schemas found in test environments. Meta launches Content Seal watermarking system for AI-generated content detection Meta's new invisible watermarking technology addresses platform accountability for AI-generated content, though it remains less accessible than Google's existing SynthID solution. MCP servers fail agent usability testing, one-third score D or F grades Poor server design undermines the Model Context Protocol's promise to standardize AI agent tool access, creating friction in production deployments.