AI Writing Research Data 2024-2025: Word Surge, Academic Fields & Detection Bias - Geostar

Signature Word Frequency Surges

Documented increases in specific word usage correlated with ChatGPT adoption

Detailed Surge Analysis

Word Surge Before (Year) Before (%) After (Year) After (%) Context Source
delve +654% 2022 0.15% 2025 46% Academic papers on PubMed Nature Human Behaviour 2025
underscore +900% 2022 3% 2025 30% Academic papers Weizmann Institute + APA Journal
meticulous +200% 2020 Baseline 2023 2x frequency Scopus abstracts Scopus Database Analysis
tapestry +800% 2022 <1% 2024 ~8% Creative writing outputs Forbes AI Content Study
pivotal +450% 2022 Baseline 2024 4.5x increase Academic abstracts Correlation analysis r=0.449 with 'underscore'
intricate +335% 2022 r=0.03 2024 r=0.335 Co-occurrence with 'delve' Weizmann Institute Study

Key Findings

Model Readability Comparison

Flesch-Kincaid grade level required to understand AI-generated text

Model Grade Level Complexity
ChatGPT Grade 12 High
Claude Grade 11.5 High
Gemini Grade 10.8 Medium-High
Grok Grade 9.2 Low-Medium
DeepSeek Grade 11.8 High
Human Avg Grade 8.5 Low

Critical Finding: All major AI models write at a college-level reading grade (10-12), while average human writing registers at 8th grade. This complexity gap is a reliable detection signal — AI systematically overcomplicates language.

AI Prevalence by Academic Field

Measured AI involvement across scientific disciplines (Nature Human Behaviour 2025)

Field AI Involvement Total Papers (2023) Estimated LLM-assisted papers Co-occurrence
Computer Science 22.5% 60,000-85,000 1-2% of total 98.8% "delve" + "underscore" appear together

Detection Bias: The False Positive Problem

Certain human populations are systematically misidentified as AI-generated text.

Population Detection Rate Description
Non-native English 61.3% Low perplexity mistaken for AI
ESL Students 97.8% Low burstiness patterns
Formal Writing 35% Register matching AI
Technical Writing 28% Structured patterns similar to AI

Critical Finding: Non-native English speakers and ESL students are dramatically over-flagged as AI (61.3% and 97.8% respectively). Both groups write with lower perplexity and burstiness than native speakers, putting them squarely in the AI detection range. These tools are not reliable for evaluating their work.

Domain-Specific AI Content Analysis

How AI-generated content manifests differently across professional domains

Medical/Academic

Journalism

LinkedIn

CS/Tech Papers

Journalism: Ethical Crisis & Quality Breakdown

Critical Ethical Issues

Detection Strategy

Professional Standards Update

RLHF: The Root Cause of AI Writing Patterns

Critical Finding: Human evaluators prefer elaborate language → system amplifies ornate vocabulary → cycle repeats.