Oxford Allowed OpenAI to Train Models on 125,000 Bodleian Library Scans
Key point
Internal papers reveal Oxford University allowed OpenAI to use 125,000 scans of Bodleian Library texts for AI training, a use case not explicitly stated in the original public announcement.
Details
Training Data Disclosure
Internal documents reviewed by the Guardian reveal that Oxford University allowed OpenAI to use texts from its Bodleian Library to train AI models. While the initial deal announced in March 2025 was framed as a digitization effort to help students and scholars access materials, it did not explicitly state that the texts would be used for AI training.
By June 2025, the Bodleian had transferred 125,000 scans of old PhD theses to OpenAI. These documents include theses written at European and US universities during the 19th and 20th centuries. An OpenAI spokesperson stated that with over a billion people using their technology, it is important for models to reflect diverse cultures, histories, and perspectives.
Institutional Response and Context
Oxford maintains that the AI training aspect was not hidden from staff, noting that while scanning was the primary goal, training was a known secondary use. The university emphasized that the scans are:
- Out of copyright
- Small in scale relative to the library's holdings
- Non-exclusive, as the library retains rights and plans to post the scans online soon
Internal staff concerns regarding reputational harm and energy consumption were noted in papers obtained via a freedom of information request. Oxford is the only UK member of OpenAI’s NextGenAI group, which also includes MIT, Caltech, and the University of Michigan. The agreement occurs as AI firms increasingly seek high-quality, non-AI-generated training data, though unlike other reported incidents involving rare books, the Bodleian’s books remain intact.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.