Security & AI news

Oxford, OpenAI and the Bodleian Library · By FireAI Security & Research Team · Published

Oxford let OpenAI train on the Bodleian Library. AI’s hunger for fresh data doesn’t stop at old books

Scans from Oxford’s Bodleian Library went into OpenAI’s training set. The story is about public-domain books, but it shows how hard AI labs now hunt for fresh human writing.

The University of Oxford has let OpenAI, the company behind ChatGPT, train its AI models on historical texts digitised from the Bodleian Library, The Guardian reported on 26 September 2026. Internal documents seen by the paper say the scanned Bodleian material was used to “populate the OpenAI training set”.

Before anything else, the part that matters for your own privacy: this story is about old, out-of-copyright books and theses, not about anyone’s personal data. Nobody’s messages, photos or files were involved. It is still worth reading closely, because it shows where the AI industry is right now: short of fresh, human-written material, and looking for it everywhere.

What happened

Oxford announced a partnership with OpenAI in March 2025. The stated goal was to use OpenAI software to digitise texts from the Bodleian, one of the world’s most famous libraries, so that students and researchers could reach them more easily. According to The Guardian, that announcement did not say the material would also be used to train OpenAI’s models.

Oxford rejects the idea that anything was hidden. A university spokesperson told the paper that digitisation was its primary interest, and that staff had been open that the project would also contribute training data.

The Guardian lists what has been scanned or discussed so far:

  • By June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI, including PhD theses from European and American universities written in the 19th and 20th centuries.
  • A rare collection of 10,000 16th-century “broadside ballads”: song lyrics and musical notes once circulated on Tudor street corners.
  • Staff have also discussed digitising 18th-century Irish state papers, the private letters of the Irish novelist Maria Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.
  • The contract raises the prospect of mass digitisation of the Bodleian’s whole collection of 23 million items, and meeting minutes discussed an “Ask the Bod” chatbot.

Meeting minutes obtained by freedom-of-information request record concerns from university staff, including members of the Bodleian’s governance committee, about the reputational risk of partnering with OpenAI, and about what a deal built on an energy-intensive technology means for the university’s environmental commitments.

What Oxford and OpenAI say

The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.

University of Oxford spokesperson, via The Guardian

The university says the Bodleian keeps the rights to the scans and will begin publishing them openly online within months, as it does with other digitisation partnerships. Unlike some book-scanning projects elsewhere, the collections themselves stay intact.

An OpenAI spokesperson told the paper the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”. Oxford is the only UK member of OpenAI’s NextGenAI project, which also includes Boston Public Library, Caltech, MIT and the University of Michigan.

The bigger story: AI is running out of fresh human writing

The Guardian puts the deal in context. Text scraped from the open web is increasingly saturated with AI-generated material, which makes it less useful for training new models, so developers have turned to physical, often historical, book collections that were never online in the first place.

  • Secondhand booksellers report a spate of orders for obscure titles, such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Shop owners suspect these are bought precisely because they don’t exist online in digitised form.
  • Anthropic, OpenAI’s close rival, has spent tens of millions of dollars buying books, cutting off their spines so the pages can be scanned, and then pulping them. Anthropic says it does not buy and destroy rare or antiquarian books.
  • The tech news site 404 Media placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where books were also dismantled and scanned.

Public-domain books are a fair and useful thing to learn from. The point here is the appetite: when a lab will buy a forgotten biography just to scan it, it is fair to ask what else counts as “fresh data” to the same industry.

Where your own data fits in

Some of the freshest human writing isn’t in a library at all. It is what people type, paste and upload into AI assistants every day. According to The Guardian’s reporting, ChatGPT consumer accounts have to opt out if they don’t want their data used for training (enterprise data isn’t eligible). Whichever assistant you use, look for that setting in its data or privacy controls today. As a rule of thumb, don’t upload photos of other people, contracts, medical records or client files to any cloud AI you wouldn’t be comfortable seeing in a training set.

This isn’t a hypothetical risk. The same week, OpenAI said its own AI agents had leaked 53 images belonging to ChatGPT users, images the agents could reach because consumer data is used for part of the training process. We covered it in OpenAI’s agents leaked 53 ChatGPT users’ images.

The other half of the picture is quieter: the apps on your Mac that send data you never typed at all, such as analytics, crash reports, usage pings and advertising identifiers. That flow feeds the data-broker economy described in our article on data brokers and your Mac.

What FireAI does about the part you control

FireAI has nothing to do with the Oxford–OpenAI deal, and no firewall can take back something that has already been uploaded. What FireAI changes is what happens on your own Mac from now on:

  • FireAI’s own AI runs entirely on your Mac. The connections it explains and the rules it suggests are analysed locally, never sent to HisnLabs, to a cloud AI or into anyone’s training set.
  • The world map shows every connection your apps make and where it goes, including the ChatGPT app or any other AI app, so you can see which apps are sending data at all.
  • Per-app rules let you block an app, or just one of the domains it talks to, without breaking the rest of your Mac.
  • Paranoid mode blocks known telemetry and analytics destinations for every app you haven’t explicitly allowed, so background data you never chose to share stays on your Mac.

Libraries can decide what to do with their books. You get to decide what leaves your Mac. Download FireAI and try it free for 17 days, or read the Docs first. FireAI is made by HisnLabs.

Sources