OpenAI Used Oxford’s Bodleian Texts for AI Training, Documents Show

Oxford’s 2025 announcement focused on making historical texts easier to access. It did not mention model training, though the university says staff were open about that use.

By 3 min read
OpenAI Used Oxford’s Bodleian Texts for AI Training, Documents Show
OpenAI Used Oxford’s Bodleian Texts for AI Training, Documents Show

Listen to this story

The audio brief

About 0:43
0:000:43
Read transcript
Scanned texts from Oxford’s Bodleian Library were used to populate OpenAI’s model-training set, according to internal documents reported by the Guardian. Oxford’s public announcement described a digitisation partnership aimed at making historical material easier for students and researchers to access; it did not mention model training. The documents establish that training was a second use, but they don’t say which works went into which models. By June 2025, Oxford had shared 125,000 images of historical dissertations with OpenAI. That figure is not evidence that every image was used for training. The dissertations date from the nineteenth and twentieth centuries and come from European and American universities. The project also scanned 10,000 sixteenth-century broadside ballads, including song lyrics and musical notation. Oxford says staff were open about the training use, and rejects the suggestion that it was concealed. It describes the material as out of copyright, modest in scale, and not exclusive to OpenAI; the Bodleian retains rights to the scans. Meeting minutes obtained through a freedom-of-information request record staff concerns about reputational risk and energy use. They do not quantify this project’s environmental cost. OpenAI says historical material can help models reflect a wider range of cultures and perspectives. Oxford’s stated goal was access, too, but the promised online release of the scans has not happened yet. The key question now is whether the scans become publicly available within the months Oxford promised—and how clearly future library partnerships disclose training use.

Story brief

3 key points

The newly reported training use adds a second purpose to Oxford’s digitisation partnership with OpenAI: scanned Bodleian material was used to populate OpenAI’s training set, though the documents do not identify specific works or models. By June 2025, Oxford had shared 125,000 dissertation images; that does not mean every image was used for training. Oxford says it will publish the scans openly within months and...

  1. 01

    The images came from 19th- and 20th-century European and American theses; the project also scanned 10,000 16th-century broadside ballads.

  2. 02

    Oxford says the material is out of copyright, modest in scale, non-exclusive to OpenAI, and that the Bodleian retains rights to the scans.

  3. 03

    Freedom-of-information meeting minutes record staff concerns about reputational risk and energy use, but provide no measured environmental cost.

Historical texts from Oxford’s Bodleian Library have entered OpenAI’s model-training set, according to internal documents reported by the Guardian. That use was absent from Oxford’s public announcement of the partnership, which focused on making the texts more accessible. Oxford disputes the suggestion that it hid the training element.

How the texts reached OpenAI

Oxford announced the partnership in March 2025. OpenAI would help digitise Bodleian texts, and Oxford said the work would make them more widely available to students and researchers. The newly reported documents establish a second use: material digitised through the project was used to populate OpenAI’s training set.

Digitising a library collection makes its contents available in a form that can be shared beyond the reading room. Model training serves a different purpose: feeding material into systems that learn patterns in words. The documents connect the Bodleian project to that training process, but they do not identify which particular scanned works entered which models.

What was scanned

The shared dissertation images came from theses written at European and American universities in the 19th and 20th centuries. Texts scanned through the project also include a collection of 10,000 16th-century broadside ballads, with song lyrics and musical notes. These details show the range of historical material involved; they do not show that each work was used to train a model.

Oxford is the only UK member of OpenAI’s NextGenAI project, which also includes US research libraries such as Boston Public Library, MIT and the University of Michigan. The Bodleian arrangement is therefore part of a wider effort involving library collections, though the newly reported documents concern Oxford’s own material.

The disclosure Oxford disputes

Oxford’s March 2025 announcement did not say the digitised material would be used to train OpenAI models. A university spokesperson rejected the suggestion that this use had been hidden from students or the public. Digitisation was Oxford’s primary interest, the spokesperson said, but staff had been open that the project would also provide training data.

Meeting minutes obtained through a freedom of information request record staff concerns, including from members of the Bodleian governance committee. They raised the reputational risk of partnering with OpenAI and questioned how a deal involving energy-intensive technology sat with Oxford’s environmental commitments. The minutes record concerns, not a measured environmental cost for this project.

The access Oxford promises

Oxford says the material being digitised is modest in scale, out of copyright, and not reserved for OpenAI’s exclusive use. The Bodleian retains rights to the scans, and Oxford says it will begin publishing them openly online within months. OpenAI used the material for training; the promised public release of the scans has yet to happen.

OpenAI says including historical knowledge in today’s models helps them reflect different cultures, histories and perspectives. Oxford’s case puts that aim beside a more immediate question for libraries: how clearly should they explain training access when they present a partnership chiefly as a way to open their collections?

Sources

  1. theguardian.comOxford lets OpenAI train its AI models on Bodleian Library

Loading discussion...

YOUR READING SPACE

Notifications

OpenAI Used Oxford’s Bodleian Texts for AI Training, Documents Show | Superpower Daily