Oxford University allowed OpenAI to use scans of texts from its Bodleian Library as training data for its AI models, according to internal documents reported by the Guardian on Saturday. The Guardian’s Ethan Penny and Dan Milmo reviewed the papers, which show that a scanning project university officials described publicly as a preservation effort also functioned as a data pipeline for OpenAI.

The distinction matters for every research library weighing a similar offer. Oxford announced its OpenAI partnership in March 2025 as a way to digitize rare texts for students and scholars. Nothing in that announcement told the public the scans would end up in a model’s training set.

By June 2025, the Bodleian had sent OpenAI 125,000 scans of historic PhD theses written at European and American universities across the 19th and 20th centuries, the Guardian reported. Librarians raised concerns at the time about reputational damage to Oxford and about the energy costs of AI training, according to internal staff meeting notes that only came to light after the university disclosed them under a public-records request. Those notes did not stop the transfers.

Oxford’s position, relayed through a spokesperson to the Guardian, is that the material was out of copyright, modest in volume and not licensed to OpenAI exclusively. The library says it retains the underlying rights and plans to publish the scans itself within months. The spokesperson also said the training use was never concealed: scanning access for the public was the primary goal, and staff had been told the texts might also train models.

That defense sidesteps the real question, which is whether the public announcement gave anyone outside the university a way to learn the same thing. It did not.

Oxford is the sole UK member of OpenAI’s NextGenAI program, a group of research institutions that also includes Boston Public Library, Caltech, MIT and the University of Michigan (OpenAI has not disclosed whether those partners received the same disclosure treatment). An OpenAI spokesperson told the Guardian the company wants its models to “reflect different cultures, histories and perspectives,” given that more than a billion people now use the technology daily. That is OpenAI’s own framing of the deal’s value, not an independent assessment of how the training data was sourced.

The episode lands inside a broader scramble for training material untouched by AI generated text. Model developers have turned to physical books precisely because the open web is now saturated with machine written prose, and 404 Media reported in August that it traced a shipment of rare books to an Amazon facility that scans and destroys them for training data. The Bodleian’s copies survive the process intact, a distinction Oxford is emphasizing as its main point of difference.

It is a low bar. A library whose primary asset is centuries of institutional trust just showed that “we did not conceal it” and “we told the public” are not the same disclosure. Any university now negotiating a digitization deal with an AI lab should put the training data use in the same press release as the preservation pitch, not in a staff meeting that later required a records request to surface.

Reported by The Next Web, citing the Guardian’s investigation by Ethan Penny and Dan Milmo, on 26 September 2026.