Writer and technologist Jon Stokes has a theory about why Anthropic’s text watermarking sparked such a disproportionate backlash: the argument was never really about the watermark.
Stokes, posting on X, contends that critics defending unwatermarked output are reacting to something deeper than a labeling policy. They sense, correctly in his view, that swapping one word for a near-synonym is not a neutral operation. A model choosing “overcast” instead of “gray” is not flipping a coin between interchangeable options, even though defenders of watermarking often describe it that way to argue the technique costs nothing.
His case rests on a distinction between verifiable and unverifiable qualities. Whether a given block of text came from a model is something you can check with math. Whether a given word was the right one for that sentence, in that context, for that author’s intent, is not. There is no scoreboard for word choice the way there is for a chess move or a benchmark score, because rightness here depends on what the author was trying to do to the reader, and that shifts by speaker, audience, and moment.
Stokes illustrates the point with a dodgeball scenario: even a perfectly accurate throw cannot be judged “best” without knowing the thrower’s intentions and the consequences that unfold afterward, consequences nobody can fully specify in advance. He applies the same logic to language. A training run can be pushed toward provenance, which is checkable, but it has no equivalent target for “best,” because nobody can quantify how many readers got the second choice of word or what that substitution cost them.
That asymmetry, in his framing, is the actual fight. When one property can be measured and a competing one cannot, the systems that companies build, and the incentives that shape them, will tend to optimize the measurable one. Provenance gets engineered for. Craft, not being scoreable, gets treated as if it barely matters, even by people who privately know better.
Stokes is careful to frame this as his own read of the controversy, not a settled account of Anthropic’s motives or of how the company describes its own tradeoffs. He does not claim the company set out to sacrifice quality for verification; he argues the effect follows regardless of intent, because that is what optimizable metrics tend to do to properties that resist measurement.
The argument holds up best as a diagnosis of incentive design: any system will chase what it can score, and that dynamic long predates AI. It gets shakier as a claim about magnitude. Stokes offers no estimate, and by his own logic could not, of how often watermarking actually forces a worse word choice versus a functionally identical one, which leaves the practical stakes of his argument as asserted rather than demonstrated.
There is a broader pattern worth watching here. Capability tends to compound fastest in domains where success is checkable, coding, math, benchmark-scored tasks, because checkable outcomes are exactly what optimization can grab onto. Stokes’s framing suggests the inverse holds too: wherever quality resists measurement, whether that is prose style, editorial judgment, or design taste, model improvement and product tradeoffs will lag or get deprioritized, not because the work matters less but because nobody has found a way to score it. For anyone planning where AI tools will and won’t reliably help over the next year, that gap between checkable and uncheckable work is a better predictor than the marketing copy.
Posted by Jon Stokes on X, on August 23, 2026.