No responsible bank would hire someone without deciding what that person is allowed to do.

The role would have a title. A manager. A set of responsibilities. Access to specific systems. Spending or approval limits. Performance measures. Rules for escalation. A process for removing access when the role ends.

Yet many organizations are introducing AI agents in the opposite order. They begin with a model, connect tools, add credentials, write a broad instruction, and only then ask what authority the system should have.

That is not an agent strategy. It is an unstructured delegation of power.

Before an AI agent gets a prompt, it needs a job description.

An agent is not another chatbot

A chatbot produces an answer. An agent can pursue a goal across multiple steps. It may search systems, call APIs, update records, draft communications, open cases, execute code, or trigger a transaction.

That difference changes the risk.

A hallucinated paragraph is visible and reviewable. An incorrect action may propagate through several systems before anyone notices. The quality of the model still matters, but so do the agent's identity, permissions, tools, memory, approval rules, and the reversibility of what it can do.

This is where the invisible AI operating inside banks becomes operational. The future will not be created by calling every workflow “agentic.” It will be created by giving narrowly defined agents specific work and observing whether they perform it safely.

Current adoption supports a cautious progression. The Bank of England and FCA's survey found that 55% of reported AI use cases involved some degree of automated decision-making, but only 2% were fully autonomous. Most institutions are not handing over entire processes. They are testing bounded delegation.

That is the right direction.

The prompt is not the role

Prompts describe desired behavior. Job descriptions define organizational authority.

“Investigate this alert” sounds clear until the agent encounters missing customer data, a conflicting policy, an external website, or a request to contact the client. Can it access another account? Can it close the case? Can it send a message? Can it decide that a transaction is suspicious? Which country’s policy applies? What happens if the source material contains an instruction aimed at manipulating the agent?

A longer system prompt will not answer those questions reliably. They need to be enforced outside the model through identity, policy, application logic, and workflow design.

NIST has made agent identity and authorization a dedicated area of work. Its 2026 AI Agent Standards Initiative focuses on agents that act autonomously across external systems and internal data. NIST's identity guidance argues that agents should be treated as first-class entities with their own identifiers, credentials, and entitlements rather than borrowing a person's password or operating through a shared high-privilege account.

That is the technical version of giving the agent an employee badge instead of the master key.

The nine fields in an agent job description

I would require nine fields before any financial institution allows an agent into production.

1. Mission

Describe one outcome, not a vague aspiration.

“Improve compliance” is not a mission. “Assemble the evidence required for an analyst to review a transaction alert” is. The narrower definition makes evaluation and control possible.

2. Work accepted

Specify what can enter the agent's queue: which case types, jurisdictions, products, customer segments, and data formats. An agent trained on retail banking policy should not quietly expand into corporate lending because the documents look similar.

3. Data scope

List what the agent may read at the document, record, field, and customer level. Access should follow the user and the purpose of the task. The agent should not gain broader visibility merely because it can search faster than a person.

4. Tools and actions

Separate reading, drafting, recommending, updating, communicating, and executing. Each verb represents a different level of authority. A tool that retrieves a balance does not need permission to transfer funds. An agent that drafts an email does not need permission to send it.

5. Decision rights

State what the agent can decide independently, what requires approval, and what sits outside its authority. This is where materiality, customer impact, regulatory consequences, and reversibility become concrete.

6. Escalation triggers

Define the conditions that stop autonomous work: missing evidence, policy conflict, vulnerable customer, unusual value, low confidence, suspected manipulation, or a request outside scope. Escalation is part of the job, not evidence that the agent failed.

7. Supervisor

Name the role accountable for performance, access review, incidents, and scope changes. “The business” is not a supervisor. Neither is a committee that meets once a quarter.

8. Performance measures

Measure completed work, quality, exceptions, overrides, losses avoided, cycle time, and customer outcomes. Prompts executed and tasks attempted are activity, not performance.

9. Termination conditions

Define when the agent is paused, rolled back, or retired. Repeated policy breaches, unusual access patterns, deteriorating quality, vendor changes, or an inability to reconstruct actions should trigger a response automatically.

This document should be short enough for an operations leader, risk officer, engineer, and auditor to understand in the same way.

Start with work that already has a shape

The strongest early agent use cases are not open-ended executive assistants. They are processes with a recognizable queue, documented steps, and an existing exception path.

The Bank for International Settlements gives a good example in its 2025 Annual Economic Report. It describes AI agents assisting anti-money-laundering investigations by operating existing computer systems and preparing suspicious activity reports. The report suggests that these agents can begin as copilots, handling routine work and identifying where human involvement is required.

The pattern matters. The agent does not receive a general mandate to “fight financial crime.” It performs a defined part of an investigation, inside an existing control environment, with escalation.

That is also why human in the loop is not a governance model. A person cannot supervise an agent effectively unless the person knows why a case was escalated, sees the relevant evidence, has time to review it, and has authority to change or stop the outcome.

Give agents a promotion path

Autonomy should be earned through evidence.

I would use a five-stage progression:

  1. Observe: the agent watches the workflow and produces no operational output.
  2. Draft: it prepares work that a person reviews before use.
  3. Recommend: it proposes a decision and records supporting evidence.
  4. Act reversibly: it completes low-impact actions that can be undone, with monitoring.
  5. Act materially: it performs narrowly authorized, higher-impact actions with explicit approval or policy controls.

Promotion between stages should depend on measured performance, not a launch date. Scope should be able to move backward when conditions change.

The Financial Stability Board's 2026 consultation on responsible AI adoption reinforces this lifecycle view. Its proposed practices address strategy, governance, development, deployment, monitoring, and change rather than treating approval as a one-time event.

The control plane should be outside the agent

The agent can propose what to do. It should not decide the limits of its own authority.

Identity, entitlements, transaction thresholds, prohibited actions, approval requirements, and audit logging should sit in systems that the model cannot rewrite. This protects the institution not only from hallucination, but from prompt injection, compromised tools, ambiguous instructions, and perfectly rational actions taken toward the wrong objective.

The OWASP guidance on excessive agency identifies three recurring causes of harm: too much functionality, too many permissions, and too much autonomy. Its recommended response is simple in principle: minimize tools, minimize permissions, run actions in the user's context, and require approval for high-impact steps.

Simple does not mean easy. It requires the integration and control investment that is often hidden when teams price AI as a model subscription. As explained in The Model Is the Cheapest Part of Enterprise AI, the operating model is the real investment.

Hire for a job, not for intelligence

The excitement around agents comes from their generality. Production value will come from specificity.

A financial institution does not need a digital colleague that can theoretically do anything. It needs a reliable investigator, document preparer, service router, coding assistant, or advisor-support agent that knows its job and its limits.

The question is not whether the agent is intelligent enough.

The question is whether the institution has been precise enough about the work.

Write the job description first. Issue the credentials second. Grant autonomy last.

Sources

• Bank of England and FCA — Artificial intelligence in UK financial services 2024

• NIST — AI Agent Standards Initiative

• NIST — Why Agentic AI Needs a Strong Identity Foundation

• Bank for International Settlements — The next-generation monetary and financial system

• Financial Stability Board — Sound Practices for Responsible Adoption of AI, 2026 consultation

• OWASP — Excessive Agency