Pangram and the problem of answer shopping · ↗ freddiedeboer.substack.com
Real users do not behave like benchmark evaluators
Freddie deBoer has a good piece on Pangram, the AI-detection tool, that closely mirrors my experience with it. Pangram is still far too easy to manipulate into contradicting itself, even while reporting “high confidence” in both results.
To level set, I no longer think formal AI detection is quite as hopeless an endeavour as I did in the early days. OpenAI famously abandoned its own classifier in 2023 because of its low accuracy. The field was then a Wild West of slapdash tools being granted far more authority than they deserved. Most were probably worse than a vibe check from someone familiar with AI-generated text.
Pangram changed my mind somewhat. It seemed to take the problem more seriously than its competitors and to have made genuine advances. Perhaps this is just successful marketing, but Pangram is the best detector I have tried, by which I mean it is the one most likely to confirm my own priors.
I am still not comfortable with the “99.98% accuracy” claim that appears on Pangram’s website. As deBoer demonstrates, it is too easy to produce contradictory results by breaking a text apart or combining different pieces. Pangram declared his complete essay 100% human-written with high confidence, while declaring a passage within that essay 100% AI-written with high confidence. He could then divide that passage into smaller pieces that Pangram again called human-written. It makes no sense to produce directly contradictory results, all with high confidence. I have seen this phenomenon in my own experiments.
This matters because users are not performing clean, one-shot evaluations of random documents and accepting the results. They are answer shopping, trying to manipulate the tool into giving them the answer they want. One group is content creators trying to conceal AI-generated text by surrounding it with human-written text or otherwise cutting it up and reassembling it until Pangram gives them a clean bill of health. The other group is social media addicts picking apart a piece of writing until Pangram flags something, so they can smear its author as an AI slop farmer. These are two prominent real-world uses of the tool, and neither process remotely resembles the one that produced Pangram’s claimed accuracy figure. Certainly, no one should see a Twitter screenshot from Pangram declaring a text “100% AI-generated” and expect it to conform to the tool’s claimed “1 in 10,000” false-positive rate.
Pangram’s interface adds to the confusion. In Twitter dunks, people routinely interpret “100% of this text is AI generated” (meaning Pangram believes the entire passage was generated by AI) as “Pangram is 100% certain that this text was generated by AI.” The confidence measure is separate, although Pangram appears to have recently removed it from the overall summary and now displays it only for individual sections. They also seem to have added a warning that confidence is limited for short passages, which is also good.
Pangram may still be the best AI detector available, but its verdicts are treated with more confidence than they deserve.
