AI Writing Research Data 2024-2025: Word Surge, Academic Fields & Detection Bias - Geostar
Signature Word Frequency Surges
Documented increases in specific word usage correlated with ChatGPT adoption
Detailed Surge Analysis
| Word | Surge | Before (Year) | Before (%) | After (Year) | After (%) | Context | Source |
|---|---|---|---|---|---|---|---|
| delve | +654% | 2022 | 0.15% | 2025 | 46% | Academic papers on PubMed | Nature Human Behaviour 2025 |
| underscore | +900% | 2022 | 3% | 2025 | 30% | Academic papers | Weizmann Institute + APA Journal |
| meticulous | +200% | 2020 | Baseline | 2023 | 2x frequency | Scopus abstracts | Scopus Database Analysis |
| tapestry | +800% | 2022 | <1% | 2024 | ~8% | Creative writing outputs | Forbes AI Content Study |
| pivotal | +450% | 2022 | Baseline | 2024 | 4.5x increase | Academic abstracts | Correlation analysis r=0.449 with 'underscore' |
| intricate | +335% | 2022 | r=0.03 | 2024 | r=0.335 | Co-occurrence with 'delve' | Weizmann Institute Study |
Key Findings
- 46% of all historical uses of "delve" in academic papers occurred in just 15 months (2023-2025)
- 98.8% co-occurrence rate: Papers with "delve" also contain "underscore" — nearly perfect correlation
- Many AI signature words trace to Nigerian English-speaking annotators who performed RLHF training.
Model Readability Comparison
Flesch-Kincaid grade level required to understand AI-generated text
| Model | Grade Level | Complexity |
|---|---|---|
| ChatGPT | Grade 12 | High |
| Claude | Grade 11.5 | High |
| Gemini | Grade 10.8 | Medium-High |
| Grok | Grade 9.2 | Low-Medium |
| DeepSeek | Grade 11.8 | High |
| Human Avg | Grade 8.5 | Low |
Critical Finding: All major AI models write at a college-level reading grade (10-12), while average human writing registers at 8th grade. This complexity gap is a reliable detection signal — AI systematically overcomplicates language.
AI Prevalence by Academic Field
Measured AI involvement across scientific disciplines (Nature Human Behaviour 2025)
| Field | AI Involvement | Total Papers (2023) | Estimated LLM-assisted papers | Co-occurrence |
|---|---|---|---|---|
| Computer Science | 22.5% | 60,000-85,000 | 1-2% of total | 98.8% "delve" + "underscore" appear together |
Detection Bias: The False Positive Problem
Certain human populations are systematically misidentified as AI-generated text.
| Population | Detection Rate | Description |
|---|---|---|
| Non-native English | 61.3% | Low perplexity mistaken for AI |
| ESL Students | 97.8% | Low burstiness patterns |
| Formal Writing | 35% | Register matching AI |
| Technical Writing | 28% | Structured patterns similar to AI |
Critical Finding: Non-native English speakers and ESL students are dramatically over-flagged as AI (61.3% and 97.8% respectively). Both groups write with lower perplexity and burstiness than native speakers, putting them squarely in the AI detection range. These tools are not reliable for evaluating their work.
Domain-Specific AI Content Analysis
How AI-generated content manifests differently across professional domains
Medical/Academic
- Prevalence: 20-30%
- Top Detection Signals:
- delve (654% surge)
- underscore (900% surge)
- meticulous (2x)
- uniform sentence length
- Quality Concern: Systematic literature reviews affected; accuracy risks
- Source: PubMed + Scopus analysis
Journalism
- Prevalence: 9.1%
- Top Detection Signals:
- balanced paragraphs
- no distinctive voice
- generic transitions
- lacks narrative arc
- Quality Concern: Lower comprehension; ethical issues with source confidentiality
- Source: 1,500 US newspapers (2024)
- Prevalence: 54%
- Top Detection Signals:
- 8-step viral formula
- single-sentence paragraphs
- question CTA
- perfect emotional arc
- Quality Concern: Algorithm-optimized engagement farming; authenticity erosion
- Source: LinkedIn long-form post analysis
CS/Tech Papers
- Prevalence: 22.5%
- Top Detection Signals:
- AI signature word clusters
- excessive citations
- formal register
- RLHF patterns
- Quality Concern: 16% of peer reviews AI-generated; quality control breakdown
- Source: Nature Human Behaviour 2025
Journalism: Ethical Crisis & Quality Breakdown
- Overall Prevalence: 9.1%
- Disclosure Rate: <1%
- ROUGE-L Score: 0.62
- Quality Gap: -22%
Critical Ethical Issues
- Confidential Source Violation
- Non-Disclosure Standard
- Limited Oversight
Detection Strategy
- Structural Tells:
- Perfectly balanced paragraphs
- Uniform sentence length throughout
- Generic transitions between sections
- Content Tells:
- No distinctive voice or personality
- Lacks narrative arc
- No unique angle or insight
Professional Standards Update
- Academic Publishing (ICMJE):
- AI cannot be listed as author (January 2024)
- Mandatory disclosure of all AI use
- Authors fully responsible for accuracy
- Journalism Ethics:
- Transparency required for reader trust
- Fact-checking and verification mandatory
- Source confidentiality must be protected
- Editorial review cannot be automated
RLHF: The Root Cause of AI Writing Patterns
- Fancy vocabulary preference: System learns elaborate words
- Nigerian RLHF annotators' vocabulary amplified
- Formal response rewards: Mid-formal register everywhere
Critical Finding: Human evaluators prefer elaborate language → system amplifies ornate vocabulary → cycle repeats.