I detect Claude.
Like, if I had an AI detector that actually worked—and all the other ones didn't, which they don't—I'd put it exactly the same way.
https://www.pangram.com/blog/pangram-4-technical
https://pangram-public.s3.us-east-1.amazonaws.com/pdf/pangra...
https://www.pangram.com/blog/introducing-pangram-image-detec...
It depends what you mean by "reliably." If you mean, "we should be comfortable relying on this kind of tool at scale to identify and punish students, professionals, and writers who may have used AI," absolutely not.
If you take "reliably" to mean "1 in 200 false positive rate" as they disclose on their front page, absolutely that is possible (they are doing it today!). If you think there are more than 200 assignments turned in over a given year at university, you probably do not consider a tool like this fit for purpose. It's an open question whether those procuring said tool are aware of this
Unfortunately their marketing is really insisting on the former, and trying to push it into the zeitgeist that detection of AI-generated or edited text is reliable-type-1 now and long-term. They fail to make it clear that this is merely a tool that strongly suggests text follows patterns known to us at the present time of known LLMs. However, that fingerprint will drift over time, as LLMs get better, human writing style evolves, and the line between human and "smart autocorrect" becomes even blurrier (does speech-to-text push the model into "AI assisted" mode, because it tidied up your punctuation, for example?)
"What color are your bits" is good reading today as it was 20 years ago: https://ansuz.sooke.bc.ca/entry/23
Pangram claims a 1 in 10,000 false positive rate (rate at which human-authored texts are incorrectly classified as AI-generated). 1 in 200 sounds like the false negative rate (rate at which AI-generated texts are classified as human-authored), or perhaps a rate for a specific category of text.
In my experience, Pangram is great for detecting an author who is trying to pass someone else's work as theirs, or if they are tackling a subject they have little to no knowledge in.
Or, and—hear me out—the writer typed two successive hyphens on their iDevice.
Outlook used to do this. I'm using a Mac, and Outlook on the Mac is a different animal, so Outlook on Windows may still do this conversion.
To get an en dash in Word, type "<word> - <word>" (word space hyphen space word) followed by spacebar.
To get an em dash, type "<word>--<word>" (word hyphen hyphen word) followed by spacebar.
If you've tried AI detectors a couple years ago, they're basically in the same position that coding agents were a few years ago, where everyone was skeptical at first, but the tech has gotten a lot better. Give it a shot, it's quite good. They are slightly tuned a bit towards classifying things as AI, but I imagine that's deliberate.
The only thing is that their models are pricey, but, very useful.
It's easy to get high accuracy numbers if you're testing on long (50+ words) texts. Much harder when the documens are short, as code comments tend to be.[2]
[1]: https://xkqr.org/aicomment
[2]: https://entropicthoughts.com/better-ai-comment-classifier