The easiest AI failure to imagine is a hallucination.

A wealth-management assistant invents a fund, misstates a fee, or cites a policy that does not exist. The error is visible. It can be tested, logged, and corrected. Institutions are already building retrieval, source citation, review, and evaluation around this class of problem.

The more serious failure may look completely different.

The system uses real products, current prices, approved research, and an accurate client profile. Its recommendation is plausible. The advisor accepts it. The client acts.

Only later does the institution discover that the system consistently favored the products most profitable to the firm, treated a stale risk score as current, failed to recognize a vulnerable customer, or optimized conversion while the policy said to optimize client outcomes.

No fact was invented.

The system did what it was designed to do. That will be the scandal.

Wealth management is moving toward AI before the controls are settled

The market is still early, but no longer theoretical.

The FCA's 2026 wealth-management survey covers a sector serving more than 5.5 million retail clients and managing almost £1 trillion. Thirteen percent of responding firms said they used in-house or third-party AI tools. The share reached 45% when firms considering use within the following twelve months were included. The FCA also reports that one in five UK adults is open to AI making financial decisions for them.

Those numbers do not mean autonomous advice is already normal. They show the direction of travel. AI is entering client communication, advisor support, asset allocation, stock selection, client screening, and the identification of vulnerability characteristics.

The same report says more than 92% of firms outsource some part of their business. That matters because responsibility does not leave the firm when the model, data, or workflow comes from a vendor.

The first major conduct failure may involve several parties and still leave the wealth manager accountable for the outcome.

Accuracy is not suitability

A recommendation can be factually correct and still be wrong for the client.

Suitability depends on more than product facts. It depends on objectives, financial position, time horizon, liquidity, experience, tax circumstances, risk tolerance, capacity for loss, portfolio concentration, and sometimes vulnerability. It also depends on whether the information is current and whether the recommendation creates or reflects a conflict of interest.

AI makes this harder because it can personalize at scale. The same capability that helps tailor advice can tailor persuasion. A system can learn which message, sequence, or product framing is most likely to produce action. If the optimization target is conversion, retention, or revenue, the technology may become exceptionally effective at achieving the firm's objective while appearing to serve the client's.

ESMA's statement on AI in investment services is clear that existing MiFID II obligations continue to apply when firms use AI. It identifies algorithmic bias, poor data quality, opaque decision-making, overreliance, privacy, and security as material risks. It also places uses supporting investment advice and portfolio management within scope.

The rule does not disappear because the recommendation came from software.

The objective function is a conduct decision

Every recommendation system optimizes something, even when the organization has not written the objective down.

It may optimize expected return, risk-adjusted return, probability of acceptance, assets retained, fee revenue, engagement, or advisor productivity. It may balance several of these. The weights can be embedded in product-ranking logic, training labels, prompts, business rules, or the data used to measure success.

That makes model design a conduct decision.

If a system ranks products using commercial value to the institution, that choice should be visible, challenged, and governed as a conflict. If it uses acceptance as the main success metric, the organization should ask whether acceptance is evidence of a good outcome or merely effective persuasion. If it reduces advisor preparation time but increases concentration or turnover, productivity has been purchased with client risk.

FINRA's guidance on AI challenges in the securities industry makes the same connection. It points firms using AI for client risk profiles and investment recommendations toward suitability, conflicts, supervision, customer profiles, and portfolio rebalancing.

This is why the first line of defense cannot be “the model was accurate.” The question is accurate toward what objective.

Human approval will not automatically save the firm

It will be tempting to say that an advisor remains responsible for every recommendation.

That is necessary and incomplete.

If the system produces hundreds of apparently well-supported recommendations, the advisor may stop challenging them. If the interface presents one preferred answer without alternatives, the human is being asked to confirm, not decide. If performance targets reward acceptance and speed, skepticism carries an operational penalty.

This is automation bias with a suitability process wrapped around it.

As argued in “Human in the Loop” Is Not a Governance Model, oversight works only when the person has context, authority, time, and a precise reason to intervene. In wealth management, that means seeing the client inputs used, the products considered, the reason for the ranking, relevant conflicts, missing information, and the consequences of alternatives.

A signature on a recommendation does not prove that judgment occurred.

Test appropriateness, not only hallucination

Most evaluation programs begin with factual questions: Did the system identify the right product? Did it calculate the fee correctly? Did it cite an approved source?

Those tests are essential. They are not enough.

I would add five conduct tests.

1. The conflict test

Hold the client facts constant and change the economics to the firm. Does the recommendation change when commission, margin, incentive, or proprietary-product status changes? If it does, is that effect intended, disclosed, and permitted?

2. The stale-profile test

Change a key circumstance after the client profile was created: retirement, loss of income, a liquidity need, bereavement, or a major market loss. Does the system recognize that the available profile may no longer support a recommendation?

3. The vulnerability test

Test whether the system detects signals of confusion, stress, cognitive difficulty, low financial literacy, or unusual dependence on the advisor. Then test whether it changes the workflow rather than simply changing the wording.

4. The alternative test

Require the system to show credible alternatives, including a lower-cost option and no action. A recommendation process that cannot explain why doing nothing was rejected is incomplete.

5. The outcome test

Monitor what happened after the recommendation. Complaints, early exits, losses outside expectations, concentration, turnover, repeated overrides, and customer-support contacts are evidence about suitability that pre-deployment accuracy cannot provide.

These tests examine the whole decision system: client data, objective, model, interface, advisor behavior, product economics, and outcome.

Keep the evidence a regulator will ask for

When a recommendation is challenged, the institution will need more than the final text.

It should be able to reconstruct:

  • the client information available at that time;
  • the version of the model, prompt, policy, and product catalog;
  • the products and alternatives considered;
  • the objective and constraints applied;
  • any commercial conflicts or incentives;
  • the evidence shown to the advisor;
  • the advisor's action, changes, and stated rationale;
  • subsequent monitoring and client outcome.

This record should be designed before deployment. Retrofitting it after a complaint will not recreate a decision that was never captured.

The FCA's Consumer Duty offers a useful standard even beyond the UK: act in good faith, avoid foreseeable harm, and enable customers to pursue their financial objectives. These are outcome obligations. A technically impressive recommendation engine can still fail all three.

Firms should treat regulation as architecture, not paperwork. The control, integration, and monitoring work described in The Model Is the Cheapest Part of Enterprise AI must be built into the product rather than added at the end.

The trust question is changing

The old trust question was: Is the advisor acting in my interest?

The new one is broader: Which system influenced the advisor, what was that system optimizing, and can the institution prove that the recommendation was appropriate for me at that moment?

AI can help close the advice gap, prepare advisors better, monitor portfolios continuously, and deliver more personalized service. It can also scale a weak objective faster than any human sales force.

That is why the first valuable banking AI may remain invisible: it should strengthen preparation and control before it expands autonomous advice. More clients per advisor is valuable only if the quality of judgment survives the increase.

The first wealth-management AI scandal may contain no imaginary fund, false return, or fabricated citation.

It may be a real, efficient recommendation that was wrong for the person who received it.

Sources

• Financial Conduct Authority — Wealth management survey report 2026

• European Securities and Markets Authority — Guidance on AI in investment services

• FINRA — Key challenges and regulatory considerations for AI

• Financial Conduct Authority — About the Consumer Duty