"There will always be a human in the loop."

It may be the most reassuring sentence in enterprise AI. It is also one of the least specific.

Which human? In which loop? Looking at what information? With how much time? What decision can that person change? What happens when the person and the model disagree?

If those questions do not have clear answers, human oversight is not a control. It is a phrase in a governance document.

Financial institutions are moving from AI that drafts and recommends to AI that can route cases, call tools, and initiate actions. The more capable the system becomes, the less useful vague oversight becomes. A bank cannot govern an agent by adding an approval button at the end of a workflow and hoping someone clicks it thoughtfully.

Human involvement only reduces risk when the human has a defined job to do.

A person can be present and still have no control

The easiest way to misunderstand human oversight is to treat presence as effectiveness.

Imagine a reviewer receiving hundreds of AI-generated alerts. The interface shows a recommendation but not the evidence behind it. The reviewer has thirty seconds per case, is measured on clearing the queue, and cannot change the system's underlying rule. Technically, a human approves every decision. Operationally, the human is a rubber stamp.

This is automation bias with a workflow around it.

The opposite can fail too. If every low-risk output requires manual approval, the review queue becomes the bottleneck. People stop checking carefully because most cases are routine. The system creates work instead of removing it, and the institution concludes that AI does not scale.

The answer is not always more human review. It is better allocation of human judgment.

The Bank of England and FCA's 2024 survey shows why proportionality matters. Fifty-five percent of reported AI use cases had some degree of automated decision-making, but only 2% were described as fully autonomous. At the same time, 62% of all use cases were rated low materiality and 16% high materiality. Those categories should not have the same review design.

A system drafting an internal summary may need sampling and monitoring. A system affecting creditworthiness, fraud intervention, or a client's financial position needs tighter boundaries, stronger evidence, and a clear path for a person to stop or override it.

Regulation asks for more than a checkbox

The legal direction in Europe is more practical than the slogan.

Article 14 of the EU AI Act requires high-risk systems to be designed so natural persons can oversee them effectively. The oversight has to be proportionate to risk, autonomy, and context. The people assigned to it need the competence, training, authority, and support to understand limitations, detect problems, avoid over-reliance, interpret outputs, disregard or override them, and stop the system when necessary.

That is a job description, not a disclaimer.

It also exposes a common governance gap. Many institutions assign accountability to a committee or a senior executive, while the actual intervention sits with an operations team that has limited information and no authority to change the process. Accountability travels upward on the organization chart. Control does not.

The Financial Stability Board's 2026 consultation on responsible AI adoption proposes twelve practices across the AI lifecycle and addresses boards and senior management directly. The message is important: governance is not one checkpoint before deployment. It follows the system through selection, development, use, monitoring, change, and retirement.

A human approval step cannot compensate for weak governance everywhere else.

Five questions every oversight design must answer

I would test any "human in the loop" claim with five questions.

1. What triggers intervention?

Do not send everything to a person. Define conditions: low confidence, missing evidence, unusual value, policy conflict, vulnerable customer, irreversible action, or a material deviation from normal behavior.

The trigger should be tied to the risk of the decision, not merely to model uncertainty. A confident model can still be wrong in a consequential way.

2. What does the person see?

A reviewer needs more than the recommendation. They need the relevant inputs, source provenance, confidence or uncertainty where meaningful, applicable policy, previous actions, and the reason the case was escalated.

Oversight without context creates ceremonial approval.

3. What authority does the person have?

Can the reviewer pause the workflow, change the decision, request more information, route the case to a specialist, or shut down the system? If the only available action is "approve," the institution has not created oversight. It has created liability with a name attached.

4. How quickly must the person respond?

Time changes the control. A fraud alert that requires action in seconds cannot rely on the same process as a quarterly portfolio review. The operating design needs response-time targets, coverage, escalation paths, and a safe default if nobody acts.

5. What happens after intervention?

Overrides, disagreements, near misses, and repeated escalation patterns should become evidence. They should feed monitoring, policy changes, retraining where appropriate, and decisions about whether the system's scope should expand or contract.

If the organization records only approvals, it learns nothing from the moments when human judgment mattered.

In the loop, on the loop, or outside the boundary

Not every workflow needs the same human-machine relationship.

For low-risk, reversible work, a person can be on the loop: the system operates, performance is monitored, samples are reviewed, and anomalies trigger intervention.

For material or ambiguous decisions, a person may need to be in the loop: the system prepares evidence or a recommendation, but a qualified person decides before action.

For prohibited, unsuitable, or excessively risky actions, the answer is neither. The action should sit outside the system's authority. Governance sometimes means deciding that the AI cannot do something at all.

This is especially important for agentic AI. An agent's risk is not only the quality of its text. It is the combination of model behavior, tools, credentials, permissions, transaction limits, and the reversibility of its actions.

A system that can read a policy document and draft a response is one thing. A system that can access client records, alter a case, send a communication, or move money is another. The governance model must follow the action, not the marketing label.

What current banking practice tells us

Public supervisory evidence suggests that banks already vary human validation according to risk, but the operating details remain critical.

In workshops with thirteen banks, ECB Banking Supervision reported that higher-risk models were subject to more human validation. None of the banks in that small sample allowed models to keep learning after deployment, and banks described human oversight for high-risk decisions and real-time fraud alerts. The ECB also noted differences in how banks understood explainability and gaps in the practical application of data standards.

That combination is revealing. A person cannot provide meaningful oversight if the institution cannot explain the output or trust the inputs. Human review is the last layer of a system of controls, not a substitute for that system.

The Bank for International Settlements reaches a similar conclusion when discussing the move from copilots to agents: human oversight remains essential, but institutions also need new skills, retraining, and redesigned workflows. The human role changes as the machine's role changes.

Replace the slogan with an operating contract

Every material AI workflow should have a short oversight contract:

  • the decisions the AI can make;
  • the decisions reserved for people;
  • the events that force escalation;
  • the evidence the reviewer receives;
  • the reviewer's authority and response time;
  • the safe fallback when review is unavailable;
  • the logs and outcomes used for continuous monitoring;
  • the executive who remains accountable.

That document should be readable by the business owner, risk, compliance, operations, and the people doing the review. If it exists only in technical documentation, it will fail at the moment it is needed.

The phrase "human in the loop" survives because it makes everyone feel safer without forcing difficult decisions about accountability. But the next phase of AI in financial services will not be governed by reassuring phrases.

It will be governed by boundaries, authority, evidence, and response time.

The human is not the governance model. The design of the human's role is.

Sources

• Bank of England/FCA — Artificial intelligence in UK financial services 2024

• Regulation (EU) 2024/1689 — Artificial Intelligence Act

• Financial Stability Board — Sound Practices for Responsible Adoption of AI, 2026 consultation

• ECB Banking Supervision — AI's impact on banking

• Bank for International Settlements — AI and human capital