OpenAI Used Oxford’s Bodleian Texts for AI Training, Documents Show
Oxford’s 2025 announcement focused on making historical texts easier to access. It did not mention model training, though the university says staff were open about that use.
Listen to this story
The audio brief
Story brief
3 key pointsThe newly reported training use adds a second purpose to Oxford’s digitisation partnership with OpenAI: scanned Bodleian material was used to populate OpenAI’s training set, though the documents do not identify specific works or models. By June 2025, Oxford had shared 125,000 dissertation images; that does not mean every image was used for training. Oxford says it will publish the scans openly within months and...
- 01
The images came from 19th- and 20th-century European and American theses; the project also scanned 10,000 16th-century broadside ballads.
- 02
Oxford says the material is out of copyright, modest in scale, non-exclusive to OpenAI, and that the Bodleian retains rights to the scans.
- 03
Freedom-of-information meeting minutes record staff concerns about reputational risk and energy use, but provide no measured environmental cost.
Historical texts from Oxford’s Bodleian Library have entered OpenAI’s model-training set, according to internal documents reported by the Guardian. That use was absent from Oxford’s public announcement of the partnership, which focused on making the texts more accessible. Oxford disputes the suggestion that it hid the training element.
How the texts reached OpenAI
Oxford announced the partnership in March 2025. OpenAI would help digitise Bodleian texts, and Oxford said the work would make them more widely available to students and researchers. The newly reported documents establish a second use: material digitised through the project was used to populate OpenAI’s training set.
Digitising a library collection makes its contents available in a form that can be shared beyond the reading room. Model training serves a different purpose: feeding material into systems that learn patterns in words. The documents connect the Bodleian project to that training process, but they do not identify which particular scanned works entered which models.
What was scanned
The shared dissertation images came from theses written at European and American universities in the 19th and 20th centuries. Texts scanned through the project also include a collection of 10,000 16th-century broadside ballads, with song lyrics and musical notes. These details show the range of historical material involved; they do not show that each work was used to train a model.
Oxford is the only UK member of OpenAI’s NextGenAI project, which also includes US research libraries such as Boston Public Library, MIT and the University of Michigan. The Bodleian arrangement is therefore part of a wider effort involving library collections, though the newly reported documents concern Oxford’s own material.
The disclosure Oxford disputes
Oxford’s March 2025 announcement did not say the digitised material would be used to train OpenAI models. A university spokesperson rejected the suggestion that this use had been hidden from students or the public. Digitisation was Oxford’s primary interest, the spokesperson said, but staff had been open that the project would also provide training data.
Meeting minutes obtained through a freedom of information request record staff concerns, including from members of the Bodleian governance committee. They raised the reputational risk of partnering with OpenAI and questioned how a deal involving energy-intensive technology sat with Oxford’s environmental commitments. The minutes record concerns, not a measured environmental cost for this project.
The access Oxford promises
Oxford says the material being digitised is modest in scale, out of copyright, and not reserved for OpenAI’s exclusive use. The Bodleian retains rights to the scans, and Oxford says it will begin publishing them openly online within months. OpenAI used the material for training; the promised public release of the scans has yet to happen.
OpenAI says including historical knowledge in today’s models helps them reflect different cultures, histories and perspectives. Oxford’s case puts that aim beside a more immediate question for libraries: how clearly should they explain training access when they present a partnership chiefly as a way to open their collections?
Sources
- theguardian.comOxford lets OpenAI train its AI models on Bodleian Library
Reader comments
Newest comments first. Replies stay oldest first.