
Key takeaways
- AI-generated text presented without attribution is plagiarism, even if no words are directly copied.
- Standalone AI content detectors remain inconsistent; only three of eleven tested products had perfect scores.
- General-purpose chatbots, including Copilot, ChatGPT Plus, and Gemini, now rival or beat many commercial detectors.
It is hard to imagine that just three years after generative AI went mainstream, so many classrooms, newsrooms, and marketing departments are still wrestling with one basic question: How do you know whether a person wrote the text? In 2025, the answer is more complicated than it should be. I have been running structured tests on AI content detectors for years, and the field keeps changing. Some tools that looked promising have fallen behind, while a handful of new and updated services have quietly become much more reliable.
For this latest round, I built a simple five-sample test: two samples were written by a human and three by ChatGPT. Each detector was given every sample separately, and I recorded whether the verdict was correct. A score of five out of five means the tool identified every human sample as human and every ChatGPT sample as AI-generated. Anything above 70% confidence was treated as a firm answer. I also repeated the same challenge with popular chatbots to see whether you need a separate detector at all.
What counts as AI plagiarism?
The definition of plagiarism still matters. Merriam-Webster defines it as stealing and passing off another person's ideas or words as your own without crediting the source. If someone asks ChatGPT for paragraphs and then submits them under their own name, the result is functionally plagiarism. The fact that the machine does not complain does not make it ethical. Many schools, employers, and publishers now treat undeclared AI generation as an academic or editorial integrity violation.
There is also a second problem: false accusations. Writers who use varied sentence structure, technical vocabulary, or less common phrasing are sometimes flagged by detectors as AI-generated. Non-native speakers in particular have reported being accused of cheating when their work was entirely their own. A detector is a caution light, not a court of law.
How the standalone detectors performed
In this test, eleven standalone detectors were evaluated. Three finished with a perfect score: Pangram, QuillBot, and ZeroGPT. Others, including Copyleaks and Originality.ai, fell just short after previously posting stronger numbers. The overall average was much lower than the marketing claims would suggest.
Pangram: 100% correct
Pangram was the newcomer in this group. Founded by engineers who once worked at Google and Tesla, Pangram focuses specifically on AI detection rather than plagiarism checking or rewording tools. The service allows five free scans per day, which is enough for many users. The scan process took slightly longer than the other tools and displayed a blank white screen for a moment, but the accuracy made up for the wait. Pangram correctly identified all five test blocks.
QuillBot: 100% correct
QuillBot has not always been reliable. Earlier versions produced inconsistent results when the same text was submitted multiple times. That problem disappeared in the previous test, and QuillBot kept its perfect record this time. It is now one of the safer choices for users who need a quick check.
ZeroGPT: 100% correct
ZeroGPT also earned a perfect score. The service has matured considerably since I first reviewed it. It previously looked like a bare-bones project with advertising and not much else. It now behaves like a proper software-as-a-service product, with a company name, pricing, and support resources. Equal importance: its accuracy stayed perfect for a second straight run.
Copyleaks: 80% correct
Copyleaks markets itself as the most accurate AI detector and points to independent studies. In my test, it missed one human sample and labeled it as AI. Marketing claims are easy to make, but classroom and workplace decisions deserve better. Copyleaks did correctly identify the four other blocks, including three generated by ChatGPT.
GPTZero: 80% correct
GPTZero has evolved into a company with a clear mission around preserving human authorship. Yet its results are still inconsistent. This time it correctly identified a human sample that it had previously failed, but then it missed one of the ChatGPT samples and declared itself too uncertain to judge in another case. The exact errors change from round to round, which is troubling for a tool meant to support academic integrity.
Originality.ai: 80% correct
Originality.ai is another commercial service that claims to be the most accurate detector. Its accuracy in my testing has gone down. In an earlier run it correctly identified the human-written sample that started the test; this time it was 100% confident that same sample was AI-generated. That is the kind of false accusation that makes people distrust these products.
GPT-2 Output Detector: 60% correct
GPT-2 Output Detector is built on an old model, despite being hosted by a respected machine-learning library. It has not been updated to keep pace with modern language models. It got three of five samples right and remains primarily a historical experiment rather than a practical safeguard.
BrandWell: 40% correct
BrandWell is connected to an AI marketing platform. In this run, it detected only two of five samples correctly. It called some ChatGPT output human and was confused by other sections. There was no meaningful improvement since the earlier test.
Grammarly: 40% correct
Grammarly is widely used for grammar checking, and its AI-content checker is now presented as a finished product rather than a beta. In practice, it did not improve. It identified only two of five samples correctly and failed to recognize a long passage generated entirely by ChatGPT. Interestingly, the service did recognize that the text had been published before, but that is a different function from AI detection.
Writer.com: 40% correct
Writer.com provides AI writing tools for corporate teams and also offers a free AI content detector. In my test, it classified every block of text as human-written, including three blocks produced by ChatGPT. That creates a dangerous false sense of security.
Undetectable.ai: 20% correct
Undetectable.ai was the worst performer in this round. It is a service that can humanize AI-generated text to avoid detection, which many educators and editors view as cheating. Its own detector identified human writing as probably AI and identified all three ChatGPT samples as human. If you rely on that tool, you will get the wrong answer most of the time.
Can chatbots replace content detectors?
Because standalone tools are so uneven, I tested popular chatbots using the same pass-through prompt: Evaluate the following and tell me if it was written by a human or an AI. The results were encouraging.
ChatGPT Plus, Microsoft Copilot, and Google Gemini all returned perfect scores. They correctly sorted every human and AI sample. The free version of ChatGPT made one mistake but also offered an unexpected surprise: when analyzing the first human-written text block, it not only said the text was human but recognized the author of the samples. That is a reminder that large language models have absorbed a huge amount of public web content and can sometimes connect a piece of writing to its author.
Grok, by contrast, struggled. It classified all five samples as human and therefore missed three ChatGPT passages. That was the only poor chatbot performance in this round.
What should you do with these results?
Do not treat any single detector as absolute proof. The best results in this round came from Pangram, QuillBot, and ZeroGPT among standalone tools, while ChatGPT Plus, Copilot, and Gemini proved that general-purpose AI models can perform just as well. If you already pay for a capable chatbot, you may not need to subscribe to a separate detector. But even the best tools will occasionally generate false positives, especially for technical writing or authors for whom English is a second language.
If you are evaluating student work, job applications, or news submissions, use the detector as the beginning of a conversation, not the end. Ask for drafts, confirm the writer’s process, and give the person a chance to explain. AI may make writing easier in one sense, but deciding what is genuinely human has become harder. Have you tested any of these tools? Do you rely on a particular detector or chatbot? The practical experience of users is still more useful than the marketing page of any AI company.
Source:ZDNET News
