Libraries & Archives
9 min read

AI for UK Libraries and Archives: Preserve Before You Predict

A 2026 guide to AI transcription, cataloguing and discovery that preserves masters, provenance, rights, sensitive records, accessibility and recovery evidence.

AI for UK Libraries and Archives: Preserve Before You Predict
Libraries & Archives / 9 min read
AIENGINE

9 min read

Share

AI for UK Libraries and Archives: Preserve Before You Predict

AI can propose a transcription for a faded page, suggest names in a photograph and expand a catalogue search beyond exact keywords. None of those outputs preserves the record. Preservation begins with custody, authenticity, fixity, format knowledge, rights, context and a plan to keep the object usable after today’s model and vendor disappear.

That distinction prevents a common failure: building an impressive semantic-search demo on derivatives that cannot be traced to stable masters. Use AI to add reversible layers of access. Never let it become the only surviving description or representation.

This guide is current to 31 July 2026. The National Archives is the official archive for the UK Government and for England and Wales; National Records of Scotland and the Public Record Office of Northern Ireland have separate statutory roles. Public-records, freedom-of-information and access regimes vary. Copyright is UK-wide, while contracts, donor terms and collection policies add constraints. Confirm the institution and jurisdiction for each collection; this is not legal advice.

Define three layers and do not mix them

LayerExamplesGoverning rule
Preservationoriginal bitstream, preservation master, checksums, format and event metadatachanged only through controlled preservation action
Curatorial descriptionaccession record, arrangement, authority file, rights and sensitivity notesapproved by accountable archival staff
AI enrichmentOCR/HTR text, entities, translations, image labels, embeddings, summariesreplaceable derivative, labelled and traceable

Users should be able to distinguish what came from the record, an archivist, a contributor and a model. A search result that merges these voices without provenance creates false authority.

The National Archives defines digital preservation as management and protection that maintains authenticity, integrity, reliability and long-term accessibility. Its practical advice includes inventory, appraisal, consistent organisation, geographically separate copies and fixity checking.

Start with a collection question, not a foundation model

Write a bounded objective:

Then compare non-AI options: improved metadata, collection-level descriptions, digitisation, a controlled vocabulary, better filters or additional staff indexing may solve the need more sustainably.

Good first uses are narrow and reviewable:

Use caseBounded AI roleMain riskUseful measure
Printed OCRDraft searchable text linked to page coordinatesnames and numerals silently wrongcharacter/word error by print condition
Handwritten-text recognitionPropose lines and confidencemodern spelling imposed on historical textline accuracy by hand, date and language
Metadata suggestionSuggest terms from a controlled schemeoffensive, anachronistic or unsupported labelsprecision, reviewer acceptance and correction
Entity linkingOffer candidate authority recordstwo people or places mergedfalse-link rate and unresolved ambiguity
Semantic retrievalRank source items for a queryplausible but irrelevant associationsjudged relevance, source-open rate and misses
Redaction supportFlag likely sensitive patternsunflagged personal or security informationrecall on a risk-weighted test set

Do not begin with autonomous appraisal, disposal, access closure or removal of historic description. Those decisions require evidential, legal and curatorial judgement.

For public-facing interpretation, see AI curation in UK museums. For security operations, use our UK AI cybersecurity guide.

Build the preservation package first

At ingest:

  • assign a persistent collection and object identifier;
  • retain the received object and document its source and custody;
  • scan for malware in a controlled environment;
  • identify the file format and record the tool and result;
  • generate and independently verify cryptographic checksums;
  • capture technical, structural, rights and preservation-event metadata;
  • create separate preservation and access copies where the policy requires;
  • store redundant copies with independent failure domains; and
  • test restore, not only backup completion.

The National Archives’ digital-preservation role describes an original-file-focused approach using virus checks, fixity checks and file-format identification. Its PRONOM registry records formats and software capable of reading or writing them; DROID uses PRONOM signatures for automated identification.

Record every preservation action as an event: input, output, tool and version, parameters, date, agent, outcome and checksum. PREMIS, maintained by the Library of Congress, provides an international preservation-metadata standard covering objects, events, rights and agents.

An AI-derived TIFF, transcript or summary is never a new “original”. Keep it linked to the exact source version and regeneration recipe.

Treat transcription as evidence with uncertainty

Evaluate OCR and handwriting recognition on a sample designed by the collection, not a vendor. Include:

  • clean and damaged pages;
  • handwriting styles, scripts and languages;
  • tables, marginalia, stamps and overwritten text;
  • historic spelling and abbreviations;
  • names, dates, amounts and catalogue identifiers; and
  • material from communities likely to be underrepresented in training data.

Report character or word error rates by relevant slice, plus exact accuracy for critical fields. A single average can hide systematic failure on Welsh, Gaelic, historic type, non-Latin scripts or a particular hand. Set a “cannot confidently transcribe” route.

Keep page coordinates and display the image beside the text. Let researchers report corrections and record whether a change came from staff, a community contributor or a model. Do not silently modernise spelling. Search may optionally expand variants, but the diplomatic transcription should preserve what is visible.

For charred, faded or sealed material, machine enhancement can propose readings. Call them hypotheses until appropriately qualified experts confirm them. “The AI decoded the manuscript” is not an archival description.

Make catalogue enrichment reversible and accountable

Historical catalogues can contain outdated, harmful or discriminatory language; collections can also preserve the evidence of those systems. A model trained on that text may repeat the terms, remove context or apply modern categories anachronistically.

Adopt four fields:

  • original description: retained as historical evidence where policy allows;
  • current preferred access term: controlled and dated;
  • context note: why wording differs and who reviewed it;
  • machine suggestion: separate, provisional and never silently published.

Consult affected communities for collections that describe them, but do not ask one contributor to speak for an entire group. Record disagreements and uncertainty. Authority-control changes should be versioned and reversible.

Semantic search needs similar restraint. Return actual items, catalogue scope and matched passages before a generated answer. When the evidence is incomplete, say so. Never claim that a ranked set is comprehensive or that an embedding proves a historical connection.

Rights clearance is not solved by digitisation

Owning a physical item does not necessarily confer copyright. Preservation copying, researcher supply, exhibition and worldwide online publication are different uses.

The Intellectual Property Office’s copyright-exceptions guidance describes limited exceptions, including preservation and non-commercial text and data mining under conditions. It also notes that UK cultural institutions can no longer rely on the former EU orphan-works exception. The current UK orphan-works scheme requires a diligent search and, where applicable, a licence that is limited to UK use.

Maintain a rights record for each item or series:

  • work type, creator and creation/publication date;
  • copyright owner and evidence;
  • donor, depositor and contractual restrictions;
  • permitted preservation, research and publication uses;
  • territory, duration, attribution and takedown terms;
  • orphan-work search and licence status; and
  • rights-review date.

Do not upload restricted collection material to a model whose terms allow training or cross-customer retention. A non-commercial institutional purpose does not make every vendor or downstream use non-commercial.

Public-interest archiving is a managed purpose

Archives often contain personal data about living people. The ICO’s research-related processing guidance says public-interest archiving preserves records of enduring value and requires active, ongoing management; it must not be used to justify indefinite retention of records with no potential public value.

Identify an Article 6 lawful basis and, where needed, an Article 9 condition or DPA 2018 condition. Apply data minimisation, access controls, security and the safeguards required for the research provisions. Assess whether AI creates a new purpose, increases discoverability of sensitive facts or enables people to be re-identified across collections.

Sensitivity review should cover living people, children, health, criminal allegations, addresses, security-sensitive sites, culturally restricted knowledge and donor promises. Redaction AI may prioritise review, but a trained person must release the object. Test recall on the harms that matter; a system that finds 99% of email addresses but misses abuse-survivor identities is not “99% safe”.

For a broader framework, see AI and UK data-privacy compliance.

Access copies need accessible design

An image viewer alone does not make a handwritten letter accessible. Provide corrected text alternatives where feasible, structured headings, keyboard-operable viewers, visible focus, sufficient contrast, meaningful labels and download choices. Preserve reading order for multi-page objects and associate captions with the right image.

For public-sector bodies, apply the current accessibility regulations and publish an accurate accessibility statement. WCAG 2.2 is a useful technical baseline, but legal scope and exceptions require institution-specific assessment. Test with disabled users and assistive technology; automated checks do not assess historical meaning or usable reading order.

Security and recovery are preservation requirements

The British Library’s review of its 2023 cyber-attack describes extensive service disruption, data exfiltration and the recovery difficulty created by complex legacy infrastructure. Cultural institutions should treat that as an operational preservation lesson, not a one-off IT story.

Separate preservation storage from public delivery and model-processing environments. Use least privilege, multifactor authentication, segmented administration, logged exports, patched systems and independently recoverable backups. Do not let an AI connector index every repository merely because it can.

Test restoration to a clean environment. Measure time to recover catalogue, access service, identity system and preservation metadata separately. Maintain an offline contact and decision plan.

A 90-day pilot with release gates

Days 1–30: select one rights-understood series; establish purpose, baseline, data and rights assessments, preservation package, representative evaluation sample, user needs and non-AI comparator.

Days 31–60: generate enrichment in an isolated workspace; measure transcription and retrieval by material type; conduct sensitivity, security, accessibility and community review; prove deletion and regeneration.

Days 61–90: release to a limited research group with clear machine labels, image/text comparison, feedback and correction tools. Sample outputs daily and review harms weekly.

Release only when:

  • 100% of preservation masters retain verified checksums and untouched originals;
  • every derivative links to source object, tool/model version, parameters and creation event;
  • critical names, dates, amounts and identifiers meet the defined exact-accuracy gate;
  • transcription and retrieval thresholds pass for every agreed language/material slice, not only overall;
  • zero AI-generated catalogue changes publish without named review;
  • 100% of released items have rights, sensitivity and access status;
  • no unresolved high-severity privacy, security, accessibility or cultural-harm issue remains;
  • a clean restore meets the recovery-time and recovery-point targets; and
  • staff can remove the AI layer without losing the catalogue or preserved objects.

Pause on fixity failure, unexplained source mismatch, rights claim, sensitive-data disclosure, systematic group failure, compromised credentials or vendor change. Re-run evaluation after model, collection, description policy or access purpose changes.

The practical verdict

AI can help a researcher read, search and connect—but only preservation lets the next researcher check the evidence. The durable architecture keeps authentic masters and professional context independent from every generated layer.

Preserve first. Label uncertainty. Keep rights and sensitivity visible. Evaluate by language, material and harm. Make the AI enrichment useful today and disposable tomorrow; the archive must survive both.

TaggedDigital PreservationArchivesLibrariesCultural HeritageResponsible AI
Work With Us

Interested in implementing this for your business?

We help UK businesses put these ideas into practice. Book a call to discuss your specific situation.