Search Has Become a Decision Service
Commerce search is no longer only a box that matches typed words to product titles. A useful service has to interpret an outcome such as “a quiet fan for a small bedroom under £80,” retain the constraints, retrieve eligible items and explain why each result fits. The same pattern applies to bookings, spare parts, professional services and complex business catalogues.
That shift can remove genuine customer effort. It can also hide exclusions, invent product facts or steer people toward the item with the best margin rather than the best fit. As at 31 July 2026, the sensible UK position is therefore to treat AI search as part of the selling process. The retailer remains responsible for what its system presents and does. The CMA’s March 2026 guidance on AI agents says that using an agent does not transfer responsibility when its conduct is unlawful.
The commercial aim should be a better decision, not a more conversational interface. Start with a narrow journey where failed discovery is visible: no-result searches, repeated reformulation, mis-bought components or expensive calls asking which service applies.
Build on Catalogue Truth, Not Model Memory
A language model is good at interpreting an imprecise request. It is not an authoritative inventory, price book or eligibility engine. The production path should separate four jobs:
- Parse the request into explicit needs, constraints and uncertain terms.
- Retrieve candidates from approved catalogue and availability indexes.
- Apply deterministic rules for price, geography, compatibility and stock.
- Generate an explanation using only fields returned for those candidates.
Product identifiers, variants, mandatory fees, delivery territories, service prerequisites and effective dates need named owners. Each indexed fact should retain its source and last refresh time. If a supplier feed is six hours late, the interface should disclose uncertainty or fall back to conventional results rather than confidently promise unavailable stock.
Synonyms also need governance. “Eco,” “professional,” “safe for children” and “compatible” are claims, not harmless search terms. A controlled vocabulary should map them to evidence-backed attributes. This foundation complements our guide to AI search optimisation for answer engines, which covers how structured, citable product information helps external discovery as well as on-site retrieval.
Relevance Needs More Than Clicks
Click-through rate rewards curiosity and prominent placement; it does not prove that a result solved the request. Train and evaluate with a balanced scorecard at query and outcome level.
| Measure | What it tests | Guardrail |
|---|---|---|
| Constraint satisfaction | Whether results meet stated needs | Zero tolerance for hard-rule breaches |
| Add-to-basket after search | Whether discovery advances intent | Segment new and returning users |
| Return or cancellation rate | Whether the recommendation was accurate | Compare by query family |
| Reformulation rate | Whether interpretation was useful | Inspect high-value failed sessions |
| Margin and revenue | Whether value is sustainable | Never override relevance or fairness |
| Assisted-contact rate | Whether avoidable support demand falls | Exclude service outages |
Create a labelled evaluation set from real, appropriately minimised queries. Include misspellings, regional language, sparse requests, exclusions, incompatible bundles and “none of these” cases. Human reviewers should judge both relevance and factual support. Keep a fixed holdout set so a tuning change cannot quietly move the goalposts.
Measure latency and availability beside relevance. A more accurate answer that arrives after the customer abandons the page is not an improvement. Set separate budgets for interpretation, retrieval, rule evaluation and rendering, then use cached catalogue facts only within their approved freshness window.
Consumer Law Shapes the Ranking
The unfair-commercial-practices provisions of the Digital Markets, Competition and Consumers Act 2024 apply to relevant practices from 6 April 2025. The CMA’s current guidance covers misleading actions, omissions, pressure selling and practices involving fake reviews. Its separate price-transparency guidance explains mandatory charges and the prohibition on drip pricing.
For AI search, that means material information cannot disappear merely because the answer is concise. A result card or generated comparison should show, or clearly connect the customer to:
- the total price including unavoidable charges;
- important limits, renewal terms and eligibility conditions;
- whether an item is sponsored or ranked for commercial reasons;
- the basis and date of availability claims;
- meaningful differences between variants;
- safety, compatibility and service-area restrictions;
- the source of ratings or review summaries; and
- a route to correct an incorrect assumption.
Do not generate urgency, scarcity or social-proof language from weak signals. Do not silently preselect optional extras. If the system can execute a transaction or refund, require confirmation of the important terms and log the customer’s instruction. Legal review should test representative conversations, not just the static checkout page.
Personalisation Without Covert Surveillance
Personalisation can use declared preferences, current-session behaviour or an account’s legitimate history. Those inputs carry different expectations. Document the purpose and lawful basis for each data source, minimise fields, set retention periods and make an unpersonalised route work well.
The ICO’s final Storage and Access Technologies guidance was published in April 2026 and reflects changes introduced by the Data (Use and Access) Act. Cookies, pixels, device fingerprinting and similar access technologies remain a separate control question from whether the resulting profile can be used under data-protection law.
Avoid inferring health, financial stress, ethnicity or other sensitive characteristics from queries unless a carefully assessed service genuinely requires it. Never let a high predicted willingness to pay change the displayed base price without a lawful, fair and transparently designed policy. Customers should be able to reset or correct preference signals. Our practical guide to customer-signal product design explains how to collect useful evidence without treating every interaction as permission for unlimited reuse.
Secure the Retrieval and Action Boundary
Product descriptions, reviews, supplier feeds and web pages are untrusted inputs. A malicious instruction embedded in one of them must not override system policy or trigger an action. The NCSC secure-AI deployment guidance recommends access controls around models, data and APIs, separated environments, incident plans, audit logs and evaluation before release.
Implement the following controls:
- allow-list the tools and fields the model may call;
- treat retrieved text as data, never as instructions;
- enforce price, stock and permissions outside the model;
- use short-lived service credentials with least privilege;
- scan supplier content and isolate anomalous records;
- require step-up confirmation for payment, cancellation or account changes;
- rate-limit automated discovery and checkout paths;
- record model, prompt, source IDs, rules and action result; and
- rehearse revocation and fallback when a provider fails.
Red-team prompt injection, query leakage, cross-account retrieval and attempts to bypass age, geography or contractual restrictions. Logs should support investigation without becoming an indefinite store of raw customer conversations.
Design Honest Explanations and Escape Routes
“Recommended for you” is not an explanation. Show the customer the decisive facts: “fits your 60 cm opening,” “available for delivery to EH1,” or “includes the installation service you requested.” When confidence is low, ask one useful clarifying question rather than fabricating a preference.
Every conversational path needs ordinary controls: filters, comparison, remove-a-constraint, view-all results and contact a person. Preserve keyboard navigation, screen-reader labels and a readable transcript. A customer who cannot or does not want to use conversational search must not receive a materially worse offer.
For regulated or safety-relevant goods, route ambiguous requests to qualified support. The system should distinguish “I need a compatible cable” from advice that could create electrical, medical or financial harm. Make refusal language specific enough to help the customer take the next safe step.
A Practical 90-Day Delivery Model
Days 1–15 establish the baseline. Choose one journey and one owner. Measure failed queries, conversion, returns, support contacts and catalogue freshness. Map every system that supplies price, stock and claims. Complete consumer-law, privacy and threat reviews before choosing a model.
Days 16–35 build the evaluation harness. Label at least the important query families, define hard constraints and capture approved explanations. Create conventional-search and human-assisted benchmarks. Fix obvious catalogue defects first; AI should not disguise missing dimensions or inconsistent variants.
Days 36–60 run a staff or low-risk shadow pilot. The service produces results but does not control what customers see. Review failure clusters weekly and test adversarial content. Lock down tools, permissions and observability.
Days 61–75 expose a small, randomly selected audience with an immediate fallback. Do not limit the trial to enthusiastic customers or best-selling products. Compare outcomes by device, accessibility need, product family and new-versus-returning status.
Days 76–90 decide whether to expand. Publish the measurement window, exclusions and unresolved risks to the accountable sponsor. Expansion is a governance decision supported by evidence, not the automatic end of a sprint.
Release, Pause and Rollback Gates
Approve wider release only when all hard constraints are enforced outside the model, material product claims are source-linked, accessibility testing passes and the pilot improves at least one customer outcome without worsening returns or complaints.
Set explicit pause gates before launch:
- any purchase path shows an incorrect mandatory price;
- a safety, age or compatibility restriction is bypassed;
- sponsored placement is presented as neutral relevance;
- personal information crosses accounts or purposes;
- unsupported claims exceed the agreed sample threshold;
- return or cancellation rate rises materially;
- complaint disparity appears for a customer group; or
- operators cannot reconstruct a consequential result.
Rollback should restore a tested lexical or faceted search, not leave a broken empty box. Preserve the incident evidence, stop risky actions first and correct affected customers where necessary.
What Good Looks Like
Good commerce search is quietly reliable. It understands ordinary language, but its answers remain bounded by current catalogue data, deterministic selling rules and visible evidence. It improves discovery without manufacturing demand or hiding trade-offs.
The most valuable operating asset is not the model. It is the combination of clean product truth, a representative relevance benchmark, consumer-protection controls and a team able to pause the system. That is what turns AI-assisted discovery into a durable service rather than a persuasive demo.
Authoritative UK Sources
- CMA: Using AI agents while complying with consumer law
- CMA: Unfair commercial practices
- CMA: Price transparency
- CMA: Treating customers fairly when selling online
- ICO: Final Storage and Access Technologies guidance
- ICO: Profiling and collecting information for direct marketing
- NCSC: Secure deployment of AI systems
- ASA/CAP: Advice for businesses
This article states the position checked on 31 July 2026. It is operational guidance for UK teams, not legal advice; Northern Ireland, regulated products and overseas sales may introduce additional rules.



