Garp Independent AI & technology journalism
Saturday, August 8, 2026 Sign In · Join Subscribe
Latest Naïve raises $28.5M to automate the grunt work of setting up and running a company

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  AI text detectors struggle when language models mimic an author’s style

AI News

AI text detectors struggle when language models mimic an author’s style

AI text detectors struggle when language models mimic an author’s style

Popular AI text detectors catch plain AI-generated text with near-perfect accuracy. But when language models deliberately copy a specific author’s writing style, up to one in five AI texts slips through undetected.

Scientific writing is where the detectors fail the hardest. A research team from Epoch AI tested three of the most widely used AI text detectors: Pangram (version 3.3.2), GPTZero (model 2026-05-11-base), and Originality.ai (Turbo 3.0.2). The test covered three categories: genuine human writing, AI text generated from simple prompts, and AI text that deliberately mimicked a specific author’s style. The team built a corpus of 495 human passages from 99 authors, evenly split across blogging, fiction, and scientific writing. All texts were written before ChatGPT’s release in November 2022, which effectively rules out contamination by language models.Ad When dealing with plain AI-generated text, all three detectors performed almost flawlessly, with the false-negative rate topping out at 0.7 percent. Human texts were also classified correctly for the most part. Pangram and GPTZero didn’t produce a single false alarm. Originality.ai, however, flagged 19 out of 495 human passages as AI-generated, a troublingly high false-positive rate of 3.8 percent.AdDEC_D_Incontent-1 That result changes when language models receive writing samples from an author as reference material. For this test, three frontier models (Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro) each received five real text passages from an author and were asked to write new text in the same style. Of the 297 passages generated this way, an average of 38 went undetected, according to Epoch AI, which works out to a false-negative rate of about 13 percent. Pangram missed 10 percent of style-imitated texts, GPTZero missed 11 percent, and Originality.ai missed 18 percent.Ad For fiction, the false-negative rate across all detectors sat at just 1 to 5 percent. scientific writing told a very different story. Pangram failed to catch 25 percent of style-imitated academic AI texts, GPTZero missed 24 percent, and Originality.ai missed 29 percent. The worst individual results showed up in specific model-genre combinations within scientific writing. Pangram missed 48 percent of Gemini-generated academic passages, according to the published data.