San Francisco Daily 360

collapse
Home / Daily News Analysis / I’ve been testing AI content detectors for years – these are your best options in 2025

I’ve been testing AI content detectors for years – these are your best options in 2025

Sep 02, 2026  Twila Rosenbaum 31 views
I’ve been testing AI content detectors for years – these are your best options in 2025

How hard is it in 2025, three years after generative AI captured the world's attention, to detect when text was produced by an AI rather than a person? The answer is much harder than most people expect. I have been running controlled tests on AI content detectors since January 2023, and every few months the picture shifts. In early 2023, the best detector was right only two-thirds of the time. By spring 2025, several tools achieved perfect scores. But in this latest round, the quality has slipped again.

The exercise matters. Passing off AI-generated text as original human work is a form of plagiarism. Even when the words are not copied directly from another author, using an AI writer and not disclosing it violates the same principle: ideas and words produced by a machine are being claimed by a person. In education, journalism and professional publishing, knowing whether a text is genuinely human is important. So do AI content detectors serve that purpose? Let us walk through the evidence from the latest round of tests.

Key takeaways

  • Submitting AI-generated work without attribution is plagiarism.
  • Dedicated AI content detectors remain inconsistent and often contradictory.
  • In our tests, general-purpose chatbots matched or beat most standalone detection tools.

How I tested the detectors

To keep this evaluation fair and reproducible, I used the same five text blocks that I have used in previous rounds. Two of the blocks were written by me, a human author. The other three were written by ChatGPT. Each detector was given the text blocks one at a time. If a detector correctly classified a block as human or AI, it passed that test. If not, it failed.

For services that produce a percentage score, I treated anything above 70 percent as a strong verdict in either direction. That means a tool that says 80 percent human is considered to have declared the text human-written. A tool that says 75 percent AI is considered to have declared the text AI-generated. The threshold is not perfect, but it is a practical way to compare one system against another.

I ran a total of 55 individual tests across 11 content detectors. I also tested five leading chatbots with the same five text blocks, bringing the full test count even higher. The tests took considerable time and coffee, but the results were worth it. Let's start with the standalone detectors.

Standalone content detectors: overall results

This round included BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Pangram, Originality.ai, QuillBot, Undetectable.ai, Writer.com and ZeroGPT. One previous contender, Monica, had to be dropped because its free test window was limited to 250 words and the full toolset required a $200 upgrade. Writefull was also absent because it no longer offers a GPT detector. The biggest positive development was the arrival of Pangram, a newcomer that earned a perfect score.

Only three content detectors correctly identified all five text blocks: Pangram, QuillBot and ZeroGPT. Several established names had disappointing results. Undetectable.ai scored just 20 percent, while BrandWell, Grammarly and Writer.com each scored 40 percent. Copyleaks, GPTZero and Originality.ai each scored 80 percent, and GPT-2 Output Detector scored 60 percent.

Comparing scores over time shows no strong upward trend. I have now run this five-block test series six times, and the only consistent pattern has been that one of the human-written blocks was usually identified correctly. Even that consistency is beginning to fade. Some tools that earned perfect scores in earlier rounds lost accuracy in this round, often around the same time they added restrictions to their free services.

There is also a serious fairness problem. Text written by non-native speakers is often flagged as AI-generated even when a person wrote it. During this round, one detector was too uncertain to judge one of my human-written samples, while another detector declared the same sample fully AI-written. These are wildly inconsistent results, and they highlight why these tools should not be used as the sole basis for accusing someone of academic dishonesty.

Can chatbots do the job better?

Content detectors are not the only tools available. I also tested leading chatbots by giving each one the same instruction: Evaluate the following text and tell me if it was written by a human or an AI. The results were striking. ChatGPT Plus, Microsoft Copilot and Google Gemini all earned perfect scores. ChatGPT's free tier missed one of the human-written blocks but otherwise performed well. In one remarkable case, ChatGPT free tier not only identified a block as human-written but also correctly identified me as the author, even though the session was run in an incognito window without any account information.

The least useful chatbot in this test was Grok. It seemed to assume that all five text blocks were human-written, which means it incorrectly classified all three ChatGPT-generated samples. That is a reminder that AI chatbots, like dedicated detectors, are not all equally reliable when used for this purpose. Still, the fact that several free or low-cost chatbots outperformed most paid detection services raises a practical question: why pay extra for an AI detector when the AI assistant you already use may do a better job?

Detector-by-detector results

BrandWell AI Content Detection: 40% accuracy

BrandWell began as Content at Scale, an AI content generation service, before moving to a new marketing-focused brand. I expected it to improve over the past year, but it did not. It correctly identified only two of five text blocks and remained confused by several ChatGPT-written passages. For one AI-written block, it confidently declared nearly every sentence human-authored. BrandWell was the first product tested, and the results were not encouraging.

Copyleaks: 80% accuracy

Copyleaks markets itself as one of the most accurate AI detectors available. In this round, it correctly handled four of five samples, which is respectable. But the one mistake was serious: a human-written sample was flagged as 100 percent AI-generated. For a writer or editor trying to verify originality, a false positive of that confidence can be damaging. Copyleaks remains a professional service, but the claim of near-perfect accuracy is not supported by these tests.

GPT-2 Output Detector: 60% accuracy

This tool is built on an older model and has not shown any improvement since it was first tested. It correctly identified three of five blocks. The name itself reveals the problem: it was designed to detect output from GPT-2, an early language model that is now long obsolete. It may be useful as a historical artifact, but it is not a serious tool for detecting modern AI writing.

GPTZero: 80% accuracy

GPTZero has grown from a bare-bones project into a full company with a mission to protect human writing. It offers a validation tool, a plagiarism checker and regular product updates. Despite those efforts, its accuracy declined slightly in this round. It correctly identified four of five samples, but the results were inconsistent compared with earlier tests. A text that it correctly classified in April was missed this time, while another text it previously missed was correctly identified.

Grammarly: 40% accuracy

Grammarly is widely used for grammar checking and editing, and it now includes an AI content checker. Unfortunately, the AI detection side performed poorly. It identified only two of five samples correctly. This is surprising for a company with deep expertise in natural language processing. Grammarly did correctly flag one of the AI-written texts as previously published, but that is a plagiarism-checking feature, not an AI-detection feature.

Pangram: 100% accuracy

Pangram is a relatively new company founded by engineers who previously worked at Google and Tesla. It offers five free scans per day, which was enough for our test series. The interface was a little slow and showed a blank white screen for a few seconds before displaying results, but the accuracy was perfect. Pangram correctly identified all five samples, making it one of only three standalone tools to earn a perfect score in this round.

Originality.ai: 80% accuracy

Originality.ai sells usage credits and advertises itself as the most accurate AI detector. In previous tests, it correctly identified my human-written sample as human. This time, it was 100 percent confident that the same human-written sample was generated by an AI. That is a notable regression. It still handled four of five samples correctly, but the single false positive is enough to undermine trust in its verdicts.

QuillBot: 100% accuracy

QuillBot had inconsistent results in early tests, with multiple passes of the same text sometimes yielding very different scores. It then improved dramatically and earned a perfect score in the previous round. This time, it sustained that performance. QuillBot correctly identified all five text blocks. For users who want a standalone detector with a strong track record, QuillBot is now one of the safest choices.

Undetectable.ai: 20% accuracy

Undetectable.ai is best known for its humanizer tool, which rewrites AI-generated text to make it harder for detectors to flag. That feature raises ethical concerns because it is often used to mislead teachers and editors. The company also offers an AI detector, but it performed poorly in this round. It rated one human-written sample as likely AI and rated all three ChatGPT-written samples as likely human. It took the biggest accuracy drop of any tool in this test.

Writer.com AI Content Detector: 40% accuracy

Writer.com provides AI writing tools for corporate teams, and its content detector is meant to help companies identify AI-generated copy. It failed this test by labeling every text block as human-written. Since three of the five blocks were produced by ChatGPT, that means it missed all three. The tool showed no improvement from previous rounds and cannot be recommended for serious detection work.

ZeroGPT: 100% accuracy

ZeroGPT has matured substantially since it was first evaluated. In earlier tests, the site was anonymous, filled with ads and unclear in its ownership. It now presents itself as a proper software service with pricing, a company name and contact details. Its accuracy has also improved. After earning a perfect score in the summer, ZeroGPT maintained a perfect score in this round. It is now one of the most reliable dedicated detectors in the market.

Why these results are messy

AI content detection is inherently difficult because modern language models are designed to produce fluent, natural text that mirrors human patterns. Detectors look for statistical signals such as perplexity and burstiness, but those signals can change as AI models improve. This creates a moving target. A detector that works well today may fail tomorrow when a new model is released.

There is also a human cost. Students who write in a second language, authors who use precise technical language and professionals who edit carefully can all be flagged as AI. False accusations of academic misconduct have real consequences. In that context, a tool that says 99 percent accurate can still produce disastrous errors when one false positive meets a real person's reputation.

At the same time, the tests show that decent results are possible. Three standalone detectors and three major chatbots achieved perfect scores. The key is to avoid depending on a single tool. Any verdict from an AI detector should be treated as a signal, not as proof. If a content creator is accused of using AI, it is worth asking for the specific text analysis and reviewing it carefully. If a publisher needs to check submissions, combining a reliable detector with human judgment is far better than automation alone.

What about you? Have you tried AI content detectors like Copyleaks, Pangram or ZeroGPT? How accurate have they been in your experience? Have you used these tools to protect academic or editorial integrity? Have you seen human-written work falsely flagged as AI? Are there detectors you trust more than others for evaluating originality? The conversation is still open, and the technology is still changing rapidly.


Source:ZDNET News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy