Skip to content

Financial Regulators Will Judge AI Agents by What They Can Access

Rules written specifically for AI agents remain unfinished in both the US and the EU. But the access-control rules examiners already enforce cover agents today, and engineering teams can meet them now.

Financial Regulators Will Judge AI Agents by What They Can Access
Photo by Sufyan / Unsplash
Published:

Engineering teams in financial services keep asking when regulators will publish clear rules for AI agents, and many of them put off governance work until those rules appear. In my experience, that wait misreads where supervision already stands, because examiners do not need an agent-specific rulebook to ask what an agent can reach and who approved that access.

The failures that worry me most in my own work do not require a malicious agent. An agent can misread an instruction or act on bad or manipulated context and still stay inside the permissions my team gave it. So the permissions themselves become the control that matters most, and financial regulation already has a great deal to say about permissions.

Agent-specific rules are still catching up

US banking supervisors made the gap explicit on April 17, 2026, when the Federal Reserve, OCC and FDIC issued SR 26-2 as their revised model risk guidance. The agencies write that "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." Banks running agents today cannot lean on model risk guidance as their governance anchor, since the agencies deliberately left agents out of it.

Europe gave itself more time in July 2026 through the Digital Omnibus on AI, published as Regulation (EU) 2026/1744. The omnibus moves the deadline for stand-alone high-risk systems under Annex III [the AI Act's list of high-risk uses] to December 2, 2027. Annex III covers creditworthiness scoring and life and health insurance pricing, so many financial AI use cases now have more than a year before the AI Act's human oversight rules apply to them.

Standards work is moving at a similar pace, with NIST still collecting views on how identity standards should apply to agents. Its concept paper on software and AI agent identity and authorization came out on February 5, 2026, and comments on it closed on April 2. None of this leaves agents unsupervised in the meantime, because the older access-control layer already applies to them.

Existing access rules already cover agents

New York's cybersecurity regulation offers the clearest example for US firms, and the New York State Department of Financial Services spelled it out in an October 16, 2024 industry letter on AI cybersecurity risks. The letter reminds covered entities that 23 NYCRR 500.7 requires them to limit each authorized user's access privileges to what that user's job functions need. It also tells firms to review access privileges at least once a year and remove any they no longer need.

Across the Atlantic, the Digital Operational Resilience Act has applied to EU financial entities since January 17, 2025. Its Article 9(4)(c) requires policies that limit physical or logical access to information and ICT assets to what legitimate and approved functions require. I read an agent's service account as one more identity under both regimes, and nothing in either text suggests an exemption for it.

Supervisors writing about AI directly keep returning to the same pressure points. FINRA's 2026 Annual Regulatory Oversight Report, published in December 2025, warns that agents "may act beyond the user's actual or intended scope and authority" and suggests that firms monitor agent system access and track agent actions. The Financial Stability Board's Sound Practices for Responsible Adoption of AI, released on June 10, 2026, asks for human oversight that matches the materiality and autonomy of each use case. It also warns that real-time human monitoring of agent decisions becomes impractical as agent use scales.

For a firm operating in both markets, the AI-specific rules differ in timing and scope while the access-control expectations line up closely. That overlap gives engineering teams one design target, which is an agent that can reach only what its workflow needs and leaves a record of what it did.

Tokens only answer half the question

My team stopped letting agents rely on shared or long-lived credentials as one of the first steps in our rollout. When an agent needs to perform a task, it requests a short-lived token [a credential with a built-in expiry] scoped to the systems and actions that task requires. That access expires once the task finishes, so even a leaked credential stops working soon after the task that needed it ends.

But a token tells you who is calling and very little about how much that caller should be able to change, so we review permissions as a separate step from authentication. If a workflow only needs to read information, the same identity cannot modify records or kick off downstream actions. And that separation maps directly onto the job-function test in 500.7, because an examiner wants to know what the identity can do once it gets in.

Provider guardrails can't decide what an agent may do

We also learned to stop treating the LLM provider's built-in controls as a permission system. In one workflow, an agent had to process a legitimate structured JSON payload containing descriptions and labels, and the provider's gateway blocked the request as a possible prompt injection attempt. The input was expected and came from our own workflow, and we resolved the problem by restructuring the request and using structured output formatting so the gateway could process it.

The incident showed me how little a provider guardrail knows about the application around it. A guardrail has no view of the internal systems the agent will call or the permissions and business rules behind that call. We now keep provider guardrails as one detection layer and let our own authorization model decide what an agent may do. OWASP draws the same line in its guidance on excessive agency, which tells developers to limit the permissions LLM extensions have in other systems to the minimum necessary.

Save human approval for high-impact actions

Human approval is where regulation and engineering meet most directly, and the EU has written it into law for high-risk systems. Article 14 of the EU AI Act requires those systems to let people override or reverse an output and bring the system to a safe halt. FINRA also lists autonomy without human validation as a core agent risk for its member firms. But the FSB's warning about scale applies here too, because a reviewer buried in routine approvals soon stops looking at any of them closely.

My team handles that tension by reserving approval for the actions that need it. Before an agent acts on its own, we look at the privileges the action requires and the damage a wrong decision could cause. We also weigh how easily we could reverse the action and how much context the agent has to check its own call. Actions with limited impact and a clear rollback path can run with more automation, while anything that touches production systems, sensitive data or money movement gets narrower permissions and an approval step before execution.

Decisions that depend on business context the agent cannot verify stay with a person. A pull request review agent showed us how easily a low-risk label can hide real authority, because its service account could read the codebase and also merge approved changes. We removed the merge permission and kept final approval with a human reviewer. We also had the agent cross-reference business documentation from domain owners, which improved its analysis without giving it a path to push that analysis into production.

Walk compliance teams through the whole workflow

Our risk and compliance teams review the full workflow before any AI feature reaches production, and they ask us to demonstrate it end to end. They want to see what data reaches the model, which systems the agent can access, what actions it can take, and where the approval gates are. They also check LLM requests for personally identifiable information, review database access for risks such as SQL injection, and test how the application handles prompt injection attempts.

Our engineers prepare for that review with one check that has become especially useful before release. We map every system the agent can reach and every action its identity can perform, then compare that list with what the workflow actually needs and remove anything without a clear purpose. For sensitive actions, we also feed the model malformed, misleading and hostile input to confirm that the permission boundaries hold when the model behaves unexpectedly.

Engineers use the access map as a release gate, and compliance teams get the evidence FINRA has in mind when it asks firms to monitor agent system access and track what agents do.

Start the access record before examiners ask

The safest design I know gives an agent only the access its task needs and makes the engineering team justify every additional capability. Adopting that rule costs little today, and it leaves a written record of why each permission exists. The EU's high-risk deadline takes effect in December 2027, and US banking agencies still have to decide how they will supervise the agentic AI that SR 26-2 left out. Teams with that record will adjust their controls when those rules take shape instead of rebuilding them from scratch.

Regulators on both sides of the Atlantic are still writing their AI rulebooks, but they already agree that an identity should reach only what its job requires. Engineering teams that treat each AI agent as an identity with a specific job will have far less to explain once an examiner starts asking questions.

More in Agentic AI

See all

More from S Pattnaik

See all