Predictive maintenance and air-traffic management are often grouped under “aviation AI”, but they are not one assurance case. A maintenance model recommends attention to an aircraft or component. An ATM tool can alter the information, sequence or workload through which controllers manage live traffic. Their hazards, users, evidence and safe fallbacks differ.
The common principle is that a model never owns airworthiness or separation responsibility. It supplies evidence inside an approved operational system. The organisation must show what decision the output supports, who remains accountable, how the system fails and what happens when its performance moves outside the approved envelope.
This article reflects UK Civil Aviation Authority material available on 31 July 2026. Applicability depends on the aircraft, operation, approval and service provider. UK CAA, EASA, military, Crown Dependency and overseas-territory regimes are not interchangeable; an organisation must identify every competent authority and approval in scope.
Build Two Separate Safety Cases
Do not use a platform-level claim such as “validated for aviation”. Write an intended-use statement for each function.
| Dimension | Predictive maintenance | ATM decision support |
|---|---|---|
| Operational object | Aircraft, engine, component or system | Functional ATM/ANS system and live traffic context |
| Typical output | Anomaly, remaining-life estimate or inspection priority | Conflict, demand, sequence or sector-load advisory |
| Accountable user | CAMO, operator, engineer and approved maintenance organisation as applicable | ANSP, controller, supervisor and safety-management roles |
| Primary harm | Missed degradation, unnecessary removal or invalid maintenance action | Missed conflict, distraction, mode confusion or unsafe workload |
| Evidence focus | Configuration, sensor provenance, failure labels and maintenance outcome | Scenario coverage, timing, human performance and system interaction |
| Fallback | Approved maintenance programme and existing reliability process | Approved operational procedure and contingency mode |
The CAA’s current AI strategy page frames its role around safety, security and consumer outcomes while enabling responsible adoption. Its 22 July 2026 strategy announcement is a roadmap, not a blanket approval for AI products. Existing airworthiness, ATM/ANS, safety-management and occurrence-reporting obligations still govern the operational change.
For the cross-cutting dossier structure, use an AI assurance evidence pack, but keep the two hazard logs and approvals distinct.
Anchor Maintenance AI to the Approved Programme
A useful maintenance model can identify a changing vibration pattern, rising temperature, recurring defect combination or unusual degradation. Its recommendation does not amend an approved aircraft maintenance programme, release an aircraft to service or certify a repair.
CAA rule M.A.708 requires an approved continuing airworthiness management organisation, where applicable, to control the aircraft maintenance programme, coordinate maintenance, ensure applicable directives are implemented, manage defects and archive continuing-airworthiness records (M.A.708 regulatory text). Contracting a model or analytics provider does not move those responsibilities.
Start with a decision statement such as:
Then map every input:
- aircraft, engine and component configuration;
- sensor identifier, calibration and sampling behaviour;
- flight phase and environmental context;
- maintenance and defect coding;
- part removal and inspection finding;
- deferred-defect and minimum-equipment context;
- data gaps, replacements and software changes; and
- time at which each field became available.
Prevent label leakage. If a model sees a maintenance code or measurement created after the engineering decision, a retrospective score can look excellent while being impossible in live use. Use rolling temporal evaluation and separate aircraft or fleets where required to test transfer.
Validate the Decision Cost, Not Generic Accuracy
Define false-negative and false-positive consequences for the particular component and task. A missed advisory can permit degradation to continue until the approved programme or another warning catches it. A false advisory can cause unnecessary inspection, removal, spares use and maintenance exposure.
Measure:
- sensitivity for the defined failure or finding;
- positive predictive value at the proposed alert threshold;
- warning time before the actionable maintenance point;
- alerts per aircraft or operating hour;
- no-fault-found removal and inspection rate;
- abstention and missing-data frequency;
- performance by configuration, age and operating pattern; and
- engineering overrides and their later outcomes.
Calibration matters when a score is interpreted as risk. Test whether similarly scored cases show similar outcome frequency in the target operation. Do not translate an uncalibrated anomaly score into “days remaining”.
An illustrative shadow trial ranks candidates while engineers continue the approved reliability and maintenance process. Reviewers record whether the advisory was new, already known, plausible, actionable and supported by source evidence. No interval or task changes during shadow mode. Any proposal to change the programme follows the organisation’s established approval and reliability process.
The broader operational pattern is similar to predictive fleet maintenance, but aviation configuration and release controls make direct transfer of thresholds inappropriate.
Preserve a Reconstructable Maintenance Record
An alert must be reproducible after the component, model or software has changed. Retain:
- aircraft and component configuration at inference time;
- input-data window and quality flags;
- preprocessing and feature version;
- model identifier and threshold;
- output, explanation and uncertainty;
- reviewer, decision and reason;
- work order or inspection linkage;
- physical finding and later outcome; and
- corrections without deleting the original entry.
Separate the model’s inference from confirmed engineering facts. Generated summaries should never silently convert “possible bearing anomaly” into “bearing defect”. The signed technical and continuing-airworthiness records remain controlled under the applicable system.
Supplier updates need formal change classification. Retraining, a new sensor, a different feature window or threshold can alter the safety behaviour even if the user interface is unchanged. Keep an approved version available for rollback and prohibit unreviewed cloud updates in the operational path.
Treat ATM AI as a Functional-System Change
ATM safety is not proved by testing the model alone. The functional system includes procedures, people, hardware and software. The CAA’s ATM change-management guidance requires ANSPs to use an approved change-management procedure and assess change risk. Changes to equipment, software, roles, procedures, configuration and operating context can require notification or prior approval.
The current consolidated UK ATM/ANS provision-of-services material includes UK Regulation 2017/373 and associated acceptable means and guidance. An AI advisory entering a controller display or influencing flow management must be assessed in that wider context.
Define the operational envelope:
- airspace, traffic mix and demand range;
- surveillance and flight-data sources;
- latency, update and stale-data limits;
- weather and degraded modes;
- controller role and phase of operation;
- display position, salience and acknowledgement;
- recommendation and prohibited action types;
- coordination with adjacent units or services; and
- fallback and recovery procedure.
Do not train on an idealised world and validate only nominal traffic. Include rare but credible combinations: rapid runway change, surveillance degradation, communications failure, emergency aircraft, unusual military or general-aviation activity, severe weather, staffing constraints and interacting system alerts.
Make Human Factors a Primary Evidence Stream
An advisory can be factually correct and still unsafe if it arrives late, masks another cue or creates mode confusion. The reviewer’s role cannot be described merely as “human in the loop”.
Test whether controllers and engineers can:
- understand what the tool is and is not considering;
- recognise stale, incomplete or contradictory inputs;
- identify when the system is outside its envelope;
- disagree without procedural or interface friction;
- maintain skills when the tool is available for long periods;
- recover when it disappears suddenly; and
- coordinate the changed information with other roles.
Measure response time, head-down time, rejected recommendations, automation surprise, workload, missed concurrent tasks and recovery performance. Compare trained crews with different experience levels and include fatigue-relevant operating periods where appropriate.
The CAA’s Human Factors Action Plan 2024–2026 treats human performance as part of aviation-system safety. Evidence from simulation and shadow use should therefore feed the safety assessment, training design and operating procedures rather than sit in a separate usability appendix.
Control Data, Cyber and Timing Failures
ATM and maintenance tools depend on data that can be delayed, duplicated, misconfigured or maliciously altered. Record authoritative sources and reject data outside defined quality bounds. A model should expose missing or stale inputs instead of filling them with plausible values.
For ATM/ANS providers, CAA cyber-security regulation guidance points to duties to protect systems, constituents and data against information and cyber threats. Segment experimental analytics from operational control, restrict supplier access, use signed releases and log privileged changes.
Design these failure behaviours:
- stale inputs visibly disable the advisory;
- loss of model service leaves the approved process available;
- conflicting sources produce an abstention, not a guessed resolution;
- a display fault cannot issue an operational command;
- time synchronisation failure is detected;
- rollback does not lose the audit trail; and
- recovery is exercised with the people who will use it.
Run tabletop and simulator exercises for both silent error and obvious outage. Silent plausibility is often more dangerous than a visible system failure.
Feed Incidents Into Just Culture
AI-related hazards must enter the normal reporting and learning system. Staff should be able to report an implausible recommendation, misleading display, data anomaly or automation surprise without first proving that the model caused an occurrence.
CAA occurrence-reporting guidance explains the mandatory and voluntary reporting routes under UK Regulation 376/2014. The CAA’s Just Culture guidance emphasises learning and fair treatment while retaining accountability for gross negligence or wilful violations.
Update internal taxonomies so reports can identify the model version, data condition, interface state and human response. Do not use monitoring of overrides as a performance tool that discourages disagreement. Review recurring workarounds as design evidence.
Use an Evidence Ladder Before Operational Release
| Gate | Required evidence |
|---|---|
| Intended use | Decision, user, configuration, jurisdiction, prohibited actions and accountable approval holder |
| Offline validation | Representative temporal and configuration coverage; leakage checks; consequence-based metrics |
| Shadow operation | No operational effect; every advisory reviewed; workload and false-alert capacity measured |
| Simulation | Nominal, adverse and degraded scenarios; trained-user recovery; cyber and timing failures |
| Bounded live use | Narrow envelope, independent monitoring, immediate fallback and version lock |
| Scale | Safety objectives remain met across sites, shifts and configurations; change control and reporting work |
Production release should require zero unreviewed changes to maintenance programmes, releases to service or operational ATM actions. Every safety-relevant advisory must be attributable to its inputs and version. Every fallback must be demonstrated under realistic workload, not only described in a document.
Pause if a configuration falls outside validation, advisory load exceeds staffed review, subgroup performance breaches a safety floor, source data changes or a supplier update cannot be independently assessed. Benefits such as delay reduction or maintenance efficiency are considered only after safety gates pass.
Aviation AI becomes credible when it strengthens the approved safety system without obscuring responsibility. Predictive maintenance should create better engineering evidence; ATM support should improve awareness without weakening controller authority, human performance or recoverability.



