Heritage
9 min read

AI Genealogy in the UK: Evidence Before Family-Tree Certainty

A 2026 guide to handwriting AI, record linkage and DNA matching for UK family history, with archival provenance, privacy and verification controls.

AI Genealogy in the UK: Evidence Before Family-Tree Certainty
Heritage / 9 min read
AIENGINE

9 min read

Share

AI can propose a transcription of a faded register, rank records that might describe the same person or help organise a DNA match list. It cannot prove a parent-child relationship merely because names align, turn an ethnicity estimate into a historical fact or build a reliable family tree without cited evidence.

Family history is unusually vulnerable to plausible error. Several people can share a name and year; ages drift between censuses; boundaries and spellings change; an online tree can copy another tree’s mistake. One confident automated link can then create hundreds of invented descendants.

This guide is current to 31 July 2026. Civil registration, census access, archives and family law differ across England and Wales, Scotland and Northern Ireland; the Republic of Ireland is a separate jurisdiction. Record-access rules also vary by series and repository. Check the catalogue, licence and living-person restrictions for each source. This is research and data-governance guidance, not legal or genetic counselling.

Begin with a research question and evidence table

Do not ask a model to “find my ancestors.” Start with a bounded question: for example, whether the Mary Jones who married in Cardiff in 1888 is the same person recorded with particular parents in a later census.

Create an evidence table before generating a tree:

FieldWhat to preserve
Sourcerepository, collection, series, piece, folio/page or certificate reference
Imagestable link or permitted local copy, with rights and access note
Transcriptionexact text, uncertain characters and transcriber
Eventtype, date as recorded, place as recorded and registration date if different
Personnames and relationships exactly as the record states them
Interpretationproposed identity or relationship, confidence and competing explanation
Verificationsecond source, original/certificate check and reviewer

Keep observation separate from inference. “Household member described as daughter” is an observation; “biological daughter of both adults” may be an unsupported inference. Historical terms, household structures and recording practices require context.

Every relationship edge needs its own evidence and status: confirmed, probable, possible, contradicted or rejected. Do not let a tree-builder turn “possible” into a permanent fact because it needs a single parent field.

Use handwriting recognition as a draft

Handwritten-text recognition can make page images searchable and accelerate a first pass. Performance changes with script, clerk, language, ink, bleed-through, image quality, abbreviations, columns and page layout. A polished sentence may conceal a wrong surname, date or relationship.

The National Archives warns on its research-enquiry page that generative AI cannot accurately search its catalogue and may suggest misleading or incorrect references. A model-generated reference must never be treated as a source until it opens in the repository and matches the document.

For each transcription:

  • preserve the original image and repository reference;
  • store model and processing version;
  • mark unreadable text instead of inventing a completion;
  • retain alternatives for ambiguous characters;
  • show line or bounding-box alignment where possible;
  • have a person inspect names, dates, places and relationships;
  • compare recurring handwriting elsewhere on the page or volume; and
  • record corrections without overwriting the original output.

Measure character and word error, but create a separate critical-field error rate. A transcription that gets 98% of words right but turns “widow” into “wife” or 1841 into 1871 can break the research conclusion.

Do not modernise spelling silently. Search can use aliases and normalised forms, while the transcript preserves the original. For broader digitisation practice, see AI for archives and digital preservation.

Respect what each record system actually covers

The National Archives’ birth, marriage and death guide for England and Wales explains that civil registration began on 1 July 1837 and that certificates come from the General Register Office or local register offices, not The National Archives. Before that date, parish and other local records are central.

Its census guide covers historical censuses from 1841 to 1921 and explains routes for Scotland and Ireland. A census is a snapshot created for administrative purposes, not a certified genealogy. Age, birthplace and relationship may be approximate, mistranscribed or supplied by another household member.

Coverage narrows further back. The National Archives’ medieval and early modern family-history guide notes that most ordinary people are less well documented and that a catalogue keyword search is not comprehensive. Its pre-1858 wills guide explains the fragmented probate-court landscape and that not everyone left or proved a will.

Scotland has its own record system. National Records of Scotland’s 2026 family-history guide covers statutory registers from 1855, earlier parish records and Scottish censuses, with privacy cut-offs for images. Northern Ireland researchers should use the official GRONI family-history service and relevant Public Record Office collections.

Encode these boundaries. The model should not search an English GRO index for an 1820 Scottish birth or claim a missing result means the event did not occur. Display repository coverage, dates, geography and known gaps beside every query.

Record linkage can rank candidates using name, age, occupation, address, relatives and place. It should not collapse them automatically. Common names, reused family names, remarriage, informal adoption, migration and transcription error create convincing false matches.

Use a comparison panel that shows:

  • exact agreements and disagreements;
  • whether each field came from source text or normalisation;
  • geographic and temporal plausibility;
  • household or witness connections;
  • records that should exist but have not been found;
  • alternative candidates searched; and
  • the consequence of accepting the link.

Set higher evidence thresholds for relationships involving living people, inheritance, citizenship, identity, adoption, donor conception or allegations. Do not infer paternity from surname or household position. Keep sensitive family stories as attributed claims unless corroborated.

Never use public online trees as independent confirmation when they copy the same uncited record. Track source lineage so ten derivative trees count as one unsupported assertion, not ten votes.

The ICO’s accuracy principle guidance requires organisations processing personal data to record sources, take reasonable steps on accuracy and distinguish fact from opinion. Those habits are valuable even in a personal research project and essential for a commercial service.

Explain what DNA matching can and cannot say

Consumer DNA matching estimates shared segments and compares a customer with reference populations or other customers in the provider’s database. Results depend on the tested markers, reference panel, algorithm, thresholds and who else has tested. A regional percentage is an estimate that can change after an update; it is not a nationality or a complete migration history.

The Government Office for Science’s Genomics Beyond Health report describes ancestry uses alongside limits, identifiability of relatives, international storage and cybersecurity concerns. It also distinguishes non-medical ancestry analysis from regulated medical-purpose testing.

For a relative match, report:

  • total shared DNA and segment method;
  • plausible relationship ranges, not one definitive label;
  • known pedigree and age constraints;
  • endogamy, pedigree collapse and population limits;
  • provider and algorithm version;
  • whether both people opted into matching; and
  • documentary evidence needed to test the hypothesis.

Do not convert a DNA match into an automatic family-tree edge. A close match can reveal unexpected parentage, donor conception or adoption. Provide private controls, neutral language and a route to counselling or specialist support where appropriate.

The Human Fertilisation and Embryology Authority’s April 2026 guidance on direct-to-consumer DNA matching explains that donors, donor-conceived people and close genetic relatives may become identifiable by inference even when they did not sign up themselves. Treat this as a foreseeable product effect, not an edge case.

Protect living people and their relatives

Genetic data about an identifiable person is special-category data. The ICO’s special-category guidance explains that sufficiently identifying genetic analysis remains special-category personal data even when names have been removed.

A family-history platform must identify controller and processor roles, lawful basis, Article 9 condition, purpose, retention, international transfer, recipient and deletion. Complete a DPIA for likely high-risk processing. Separate consent or choices for:

  • producing the customer’s result;
  • relative matching;
  • public tree visibility;
  • research or product improvement;
  • health or trait reports;
  • third-party uploads; and
  • law-enforcement disclosure where legally required.

One person cannot meaningfully waive every relative’s interests. Default living profiles to private, minimise dates and locations, prevent search-engine indexing and make invitations revocable. Do not encourage users to upload another person’s raw DNA or records without authority.

Children require particular care. Avoid testing or public profiling merely to make an adult’s tree more complete. Plan how choices and access change when the child becomes able to decide.

Explain deletion honestly: removing a profile may not recall matches already seen, downloaded trees or biological inferences made by relatives. For the wider framework, see AI and UK data-privacy compliance.

Secure irreplaceable genetic and family data

Threat-model account takeover, credential stuffing, raw-DNA theft, malicious GEDCOM or image uploads, prompt injection in transcribed records, cross-tree disclosure, insider access, vendor acquisition and model-training reuse.

Use multifactor authentication, encryption, matter or tree separation, least privilege, download alerts and time-limited sharing. Scan imports and treat document text as data, not instructions. Keep an audit trail for relationship changes and public visibility.

Let users export sources, citations, uncertainty and media in interoperable formats. A tree without provenance is not a meaningful backup. Test account recovery and deletion across raw files, derived matches, embeddings, caches and support systems.

Follow the NCSC’s secure AI system development guidelines and maintain an incident plan that recognises genomic data cannot be reset like a password.

A measurable 90-day pilot

Days 1–30: choose one collection and one task, such as draft transcription or candidate ranking. Document repository rights and coverage, create a gold-standard sample, define evidence statuses, living-person rules, data flows and non-AI research baseline.

Days 31–60: test across handwriting styles, image quality, common names, boundary changes and contradictory records. Include empty results, duplicate people, malicious files, inaccessible source links, vendor outage and a deliberately persuasive false tree.

Days 61–90: release to an opt-in cohort with the original image and citations always visible. Review high-impact relationship proposals before display, privacy or unexpected-family incidents immediately and correction patterns weekly.

Release only when:

  • critical-field transcription error stays below the agreed threshold;
  • every record opens at the stated repository and reference;
  • every relationship displays evidence, contradictions and confidence status;
  • no missing search result is presented as proof of absence;
  • DNA relationships remain ranges until documentary and human review;
  • living profiles are private by default and resist search indexing;
  • matching, research, training and public sharing have separate choices;
  • correction, export and deletion work for source and derived data;
  • cross-tree and cross-user security tests pass; and
  • users can continue research without the model or vendor.

Pause after an invented archive reference, wrongly merged living people, unexpected parentage disclosure without the designed safeguards, raw-DNA leak, unauthorised public profile, material population disparity or loss of provenance. Revalidate after collection, OCR, matching model, reference panel, privacy setting or vendor change.

The practical verdict

AI can reduce the labour of reading and comparing records. It cannot remove ambiguity from history.

Preserve the image, cite the repository and keep every relationship contestable. A valuable family tree is not the largest tree an algorithm can generate; it is the smallest set of claims that the surviving evidence can honestly support.

TaggedAI Genealogy UKFamily History ResearchDNA Ancestry PrivacyHandwriting RecognitionArchive RecordsGenealogy Evidence
Work With Us

Interested in implementing this for your business?

We help UK businesses put these ideas into practice. Book a call to discuss your specific situation.