Is Claude harder to detect than ChatGPT? There is no dependable, model-wide winner. A score can change when the model version, prompt, text length, subject, editing history, or detector changes.
The Short Answer
Most commercial AI detectors do not identify a specific assistant. They estimate whether a passage resembles text in their human and machine-written training data. A result therefore says more about the passage and the detector's current model than it does about the logo on the chatbot that produced the first draft.
Claude may score as more human on one prompt and ChatGPT may score as more human on another. Even the same text can receive different results from different detectors. Treat any percentage as a signal for review, not proof of authorship.
Why the Result Changes
Model and Product Versions Move Quickly
"ChatGPT" and "Claude" each refer to a changing family of models. Providers update model behavior, and detector vendors update their classifiers. A comparison that does not name the exact models, detector versions, prompts, dates, and scoring rules cannot support a durable ranking.
Prompt and Genre Matter
A generic school-style essay is structurally different from a customer email, technical explanation, or story. Prompt constraints also affect repetition, vocabulary, and sentence rhythm. Those features can change a detector score without establishing who wrote the text.
Length and Editing Matter
Short passages provide less evidence for a classifier. Human edits can alter the patterns a detector uses, while highly templated human writing can look statistically predictable. This is one reason a score should not be interpreted as a probability that a person used AI.
What Independent Evidence Supports
Research has repeatedly found that text detection is sensitive to the detector and evaluation setup. The paper Can AI-Generated Text be Reliably Detected? describes practical limits across several detection approaches. A separate peer-reviewed study, GPT detectors are biased against non-native English writers, found important fairness concerns for evaluative use.
OpenAI withdrew its own text classifier because of its low rate of accuracy. Turnitin's current educator guidance similarly says its AI writing report is one data point rather than a definitive answer.
Can a Detector Tell Claude From ChatGPT?
Usually not from a standard AI-likelihood result. A detector may label a passage as likely human or machine generated, but that is different from attributing it to Claude, ChatGPT, or a particular model release. Reliable model attribution would require a purpose-built system and a validated benchmark for the exact models being compared.
How to Compare Them Responsibly
If you want to run your own comparison, make the test reproducible:
- Record the exact model names, dates, and settings.
- Use the same prompts and comparable output lengths.
- Include several genres instead of one essay template.
- Save raw outputs before making edits.
- Report every result, including false positives on human control samples.
- Avoid presenting a detector score as proof of misconduct or authorship.
Choose the Better Writing Tool, Not the Lower Score
Pick ChatGPT or Claude based on the work you need to do: instruction following, context handling, research workflow, privacy requirements, integrations, and the quality of the draft. Then verify facts, add source-backed detail, revise for the real reader, and disclose AI assistance when a school, employer, publisher, or client requires it.
That workflow remains useful even when detectors disagree. Better evidence, clearer reasoning, and accountable editing matter more than chasing a single score.
Compare, Edit, and Review Your Draft
Whether you use ChatGPT, Claude, or another writing assistant, AI Undetectable helps you revise the output and compare it with the original before you publish.
Open AI Humanizer