A useful library AI system helps a reader find a book, makes a fragile collection searchable or shows staff where a service is difficult to use. It does not turn borrowing history into a judgement about a person, replace consultation with a demographic score or publish machine-written catalogue facts as if a librarian verified them.
People use libraries to explore health, religion, politics, sexuality, employment and family problems. A recommendation can therefore create surveillance or self-censorship if its data and purpose are not tightly controlled.
This guide is current to 31 July 2026. Public-library law and administration are devolved: the Public Libraries and Museums Act framework described below concerns England and Wales, while delivery and oversight differ in Scotland and Northern Ireland. Data protection, copyright, equality and accessibility also depend on the organisation, material and processing. Confirm the applicable national and local rules before deployment; this is not legal advice.
Start with the public purpose
In England, library authorities have a statutory duty to provide a comprehensive and efficient service for people who live, work or study locally. The Department for Culture, Media and Sport’s current statutory-service guidance says provision should be based on local need in the context of available resources. It also notes Equality Act and Public Sector Equality Duty obligations. AI is a means of improving that service, not evidence that the duty has been met.
Write a service statement before a model specification:
- the public need and groups the service must reach;
- the data required and a less intrusive alternative;
- the decision a librarian or user retains;
- the harm caused by a wrong, missing or overly narrow result;
- the non-digital route that remains available; and
- the evidence that would justify continuing the feature.
DCMS research on engaging library non-users describes computer access, employment advice, literacy support, activities and training as well as lending. Optimising digital loans for existing members may improve a dashboard while leaving the service less inclusive.
Choose a bounded use case
Different library applications need different controls:
| Use case | Defensible output | Human boundary | Useful measure |
|---|---|---|---|
| Search assistance | ranked records with matched terms and filters | user can inspect, change and clear the query | successful discovery and reformulation rate |
| Recommendations | optional suggestions from declared interests or current session | no sensitive inference or hidden commercial ranking | useful saves, diversity and opt-out rate |
| OCR and transcription | draft text linked to page image and confidence | qualified review for public or scholarly records | character or word error by material type |
| Metadata enrichment | proposed subjects, entities or links with provenance | cataloguer approves authoritative fields | accepted edits and false-entity rate |
| Translation or summary | labelled access aid beside the source | source remains available; material facts checked | task success with target-language users |
| Service analysis | aggregated evidence about demand and barriers | staff and communities decide provision | participation, reach and access outcomes |
| Generative help | answers grounded in approved library information | referral for legal, health or safeguarding needs | supported-answer rate and safe escalation |
Start with one collection, one user task and an answer the system can decline. Separate experimental enrichment from authoritative catalogue fields and public-service decisions.
Minimise reader data
Borrowing, searches, computer sessions, event attendance and help-desk questions can reveal intimate interests even when they are not labelled “sensitive”. Define an Article 6 lawful basis, purpose, retention period, access rules and processor roles for every field. A library card should not quietly become a cross-service behavioural profile.
The ICO’s DPIA guidance identifies large-scale profiling, data matching, innovative technology, vulnerable people and decisions affecting access to a service as risk indicators. Complete and maintain a data-protection impact assessment where required, consult the data protection officer and users, and do not start processing if a high residual risk cannot be reduced.
Prefer:
- session-level suggestions over a permanent reading profile;
- explicit topic choices over inferred traits;
- coarse, thresholded service counts over individual histories;
- separate operational and research datasets;
- short retention with tested deletion; and
- local or controlled processing where it materially reduces exposure.
Children need additional care. The ICO says organisations should avoid profiling children where possible and make any human involvement meaningful in its current children and automated decisions guidance. Do not infer a child’s reading level, vulnerability or family circumstances from borrowing. Make personalisation off by default when appropriate, explain it in age-appropriate language and provide a neutral catalogue path.
Build discovery that broadens rather than narrows
A recommender trained only on historic borrowing will reproduce historic availability, promotion and membership patterns. Popularity can crowd out local authors, minority languages, accessible formats and unfamiliar subjects. “People like you” can also expose sensitive associations on a shared screen.
Use a controlled candidate set from current catalogue records. Apply deterministic availability, age and format rules before ranking. Expose why an item appears—such as a topic selected for this session—and let the user remove the signal. Do not infer protected characteristics or sensitive interests to personalise discovery.
Evaluate more than clicks:
- whether the selected item answers the stated need;
- catalogue and format coverage, including large print, audio and easy read;
- exposure across authors, subjects, languages and publication eras;
- performance for sparse, new and multilingual queries;
- inappropriate or revealing suggestions on shared devices;
- corrections, hides, opt-outs and “show me something different”; and
- outcomes for users of screen readers, keyboard navigation and low-bandwidth devices.
Keep direct search primary: a reader must be able to browse without being scored. For wider controls, see AI and UK data-privacy compliance.
Treat digitisation output as a draft with provenance
OCR can unlock newspapers, manuscripts and local-history collections, but error varies with typeface, language, layout, damage, handwriting and scan quality. A fluent summary can silently “repair” a name, date or quotation that the source never contained.
For every generated text or metadata record, preserve:
- stable source and image identifiers;
- page or region coordinates;
- capture and preprocessing history;
- model and prompt version;
- field-level confidence where meaningful;
- reviewer, correction and publication status; and
- a visible path back to the source image.
Sample errors by collection, script, decade and document condition rather than reporting one average. Names, dates, addresses and rights statements deserve targeted review. Search may use unverified OCR, but the interface should label it and let users see the scan. Authoritative descriptions require cataloguer approval.
Owning a scan does not solve copyright. The Intellectual Property Office’s copyright-exceptions guidance says the UK text-and-data-mining exception is limited to non-commercial research with lawful access; other uses need separate analysis. Record rights holder, term, licence, permitted transformations and takedown route.
Our guide to AI, archives and digital preservation covers preservation-specific controls. A public-library discovery project should reuse that provenance discipline without confusing access copies with preservation masters.
Use community evidence without manufacturing consent
Search logs and event bookings show interactions with the existing service, not the complete needs of a community. Low use may mean poor opening hours, inaccessible transport, language barriers, lack of awareness or distrust. Demographic correlations do not establish what individuals want.
Combine bounded analytics with resident panels, staff observations, accessible surveys, partner organisations and non-user research. Publish the question, data limitations and decision process. Use minimum cell sizes and suppress small or revealing combinations. Never generate a neighbourhood “risk”, literacy or political-interest score.
Staff should be able to challenge the analysis and document a different decision. If a model suggests cancelling a low-attendance session, decision-makers must inspect seasonality, capacity, promotion, accessibility, qualitative benefit and statutory equality considerations. Community programming remains a public judgement for which the authority is accountable.
Track reach rather than raw volume: new users who complete a task, participation from underserved areas, accessible-format fulfilment, repeat attendance where desired, wait time and reasons people could not use the service. Report uncertainty and missing groups.
Make the AI service itself accessible
Public-sector websites and apps in scope must meet the accessibility regulations. Current GOV.UK accessibility guidance points to WCAG 2.2 AA and an accessibility statement, while also making clear that public bodies retain wider reasonable-adjustment duties.
Test the complete journey with disabled users: discovery, consent, input, results, correction, reservation and help. Support keyboard and switch access, screen readers, zoom and reflow, clear focus, captions or transcripts, sufficient target sizes, plain language, adjustable timeouts and error recovery. AI summaries need meaningful headings and source links; dynamic suggestions must not steal focus or constantly re-announce.
Maintain staffed help, telephone or in-person routes, accessible terminals and downloadable formats. An automated translation or easy-read draft can assist production, but native speakers and intended users should review high-impact public information. Do not claim that a model “makes everything accessible”.
Secure the collection and the service
Library catalogues are public, but member data, unpublished collections, licences, API credentials and administrative systems are not. A retrieval system can be attacked through malicious document text, poisoned metadata or a compromised vendor update.
Follow the NCSC’s secure AI development guidelines: threat-model the service, document dependencies, protect models and infrastructure, log safely, manage updates and prepare incident response. Give retrieval services read-only access to an approved index, not the library-management database. Strip active content, scan uploads, separate instructions from collection text and prevent generated answers from triggering tools without explicit controls.
Contracts should state data locations, subprocessors, training terms, deletion, breach notice, audit evidence, export formats and exit support. Rehearse operation without the model and recovery from a corrupted index.
A measurable 90-day pilot
Days 1–30: choose one low-risk task and collection; document public purpose, jurisdiction, rights, lawful basis and DPIA screening; appoint library, accessibility, privacy, security and community owners; build a verified source set and non-AI baseline.
Days 31–60: test offline with librarians and representative users. Challenge obscure queries, minority languages, accessible formats, children’s use, harmful or revealing suggestions, bad OCR, prompt injection, stale holdings, deletion and vendor failure. Publish limitations for the pilot.
Days 61–90: offer the feature to an opt-in group while preserving normal search and staffed help. Sample outputs daily, review user corrections weekly and hold a community and equality review before expansion.
Release only when:
- 100% of public answers cite an approved library source or clearly decline;
- zero authoritative metadata field is published without the defined review;
- every digitised result links to its source image, rights record and processing provenance;
- no test case reveals another reader’s activity or generates a sensitive profile;
- discovery task success meets the pre-agreed improvement over standard search;
- agreed quality floors pass for every language, format and accessibility slice;
- users can browse, opt out, correct and delete without losing the core service;
- all critical journeys pass keyboard, screen-reader and disabled-user testing;
- retention, deletion, backup restoration and supplier exit are demonstrated; and
- no unresolved critical privacy, copyright, security, equality or safeguarding issue remains.
Pause after a privacy disclosure, rights violation, fabricated catalogue fact, harmful recommendation, material group disparity, inaccessible update or compromised source. Revalidate after a model, index, catalogue schema, vendor, rights status or service-purpose change.
The practical verdict
Libraries do not need an artificial librarian. They need carefully bounded tools that improve discovery and access while preserving confidentiality, provenance, intellectual freedom and accountable human service.
The strongest deployment is often quiet: optional suggestions, reversible metadata proposals, source-linked OCR and aggregated evidence that librarians and communities can question. Success is not how much activity a model predicts. It is whether more people can use a trusted library service without surrendering the freedom to read privately.



