Ed Newton-Rex, a vocal critic of how AI labs source training data, posted a thread on X describing what he calls “derived data”: a model rewrites a book or article, and a company trains on that rewritten version instead of the original text.
He argues this causes the same damage as scraping the original outright. A system trained this way can still compete with the work it drew from, even though it rarely reproduces the source sentences, which makes the practice harder to catch than straight copying.
The technique also complicates opt-outs, in his telling. A creator might block a platform from training on their writing directly, yet never have agreed to let that platform rewrite the same material and train on the result, a gap he says companies could exploit while insisting they used nothing original.
Newton-Rex’s proposed fix is training data transparency laws that would force disclosure of what models were trained on. Publishers negotiating opt-out terms should now confirm whether the clause covers rewritten derivatives, not only verbatim copies of their work.
Ed Newton-Rex outlined the argument in a thread he posted on X.