Direct Answer
Organizations should control agentic AI risk through a control plane that constrains what an agent may do, observes what it actually does, records who authorized each action, and stops or reverses harmful behavior. This is stronger than relying only on written AI policies, because an agent can interpret an objective, select tools, modify data, and take external actions with limited step-by-step human involvement. The appropriate starting point depends on autonomy: a system that drafts a document needs different controls from one that can issue payments, change production infrastructure, access confidential records, or send communications to customers.
Also worth reading: How Should Organizations Build AI Governance Evidence Architecture for Auditable Agentic Systems? · How Can Structural Engineers Implement Rigorous Agentic AI Control Testing to Prevent Systemic Failure? · How Should Engineering Organizations Govern AI Used in Structural Decisions?
A defensible design uses four linked functions: preventive limits, detective monitoring, responsive intervention, and governance evidence. Preventive limits include least-privilege credentials, tool allowlists, spending ceilings, data-access boundaries, timeouts, and approval requirements. Detective controls capture prompts, tool calls, retrieved data, decisions, outputs, and policy decisions in tamper-evident logs. Responsive controls include a kill switch, session termination, credential revocation, transaction rollback, and escalation to a named human owner. Governance evidence shows which model and prompt were used, which policy applied, who approved the deployment, and what residual risk was accepted.
There is no single product category that makes this problem solved. Governance platforms, agent observability tools, identity and access management systems, AI gateways, security information and event management platforms, and conventional application-security controls each cover part of the problem. By 2026, banks and other regulated organizations are moving beyond static policy documents toward continuous control, but that transition is incomplete. The most important question is not whether an organization has an “AI governance framework”; it is whether it can prevent an unauthorized action, detect a harmful sequence of actions, and explain the outcome with reliable evidence.
Why Conventional AI Policies Are Not Enough
An agentic system differs from a conventional chatbot because it can act through tools. It may read a ticket, retrieve a customer file, generate code, run a test, deploy a release, or call another software system. Each step can be individually plausible while their combination creates unacceptable risk. A policy saying “do not make unauthorized changes” does not define how the system recognizes authorization, isolate production credentials, or prevent a chain of individually valid actions from exceeding the intended mandate.
The core problem is indirect authority. A user delegates an objective, and the agent translates that objective into multiple technical actions. Natural-language instructions are variable, model behavior is probabilistic, and retrieved context can change during a session. Consequently, control cannot depend solely on the agent following its prompt. External enforcement must operate at the tool and data boundaries, where decisions can be tested against explicit rules before execution. Research and product development around frameworks such as MAESTRO and STRIDE reflect the need to model agent behavior as a system of components, data flows, trust boundaries, and attack paths rather than treating the model as an isolated component.
Continuous authorization should therefore replace blanket trust. An agent acting within a low-risk drafting task might run unattended for hours, while a production deployment agent should require approval for every change. A customer-service agent may be able to read account data but not alter balances; a coding agent may modify a branch but not merge it. The control boundary should be stated in machine-readable policy, tested before deployment, and monitored during execution. Human review of every token would be impractical, but human review at irreversible or unusually consequential actions is feasible and often necessary.
A Practical Control Architecture
The first layer is identity. Every agent, service account, human operator, and delegated tool should have a distinct identity. The agent should not share an employee’s general-purpose login, and separate development, testing, and production environments should use different credentials. Access should be limited by role, resource, environment, time, and transaction value. Temporary credentials are preferable for short jobs, while non-human identities should not retain dormant privileges after a task ends.
The second layer is an action broker positioned between the model and external systems. Instead of allowing the model to call a database, cloud console, payment system, or repository directly, the broker evaluates each proposed action. It can block prohibited operations, redact sensitive context, require approval, enforce a monetary or data-volume ceiling, and return only the minimum information needed. This architecture also creates a reliable audit record because all consequential tool calls pass through a consistent control point.
The third layer is continuous observation. Logs should include the agent’s objective, system and user instructions, model version, tool requests, authorization results, data sources, action outputs, latency, cost, and final outcome. Security teams need behavioral alerts for repeated failed access, sudden data retrieval, privilege changes, new destinations, code execution, or attempts to bypass approval. Thresholds should reflect context: 20 database queries may be normal for a migration agent but suspicious for a short customer-support task. Alerts should be based on deviations from an expected action envelope, not arbitrary volume limits applied without a workload baseline.
The fourth layer is response. A kill switch should stop new sessions, revoke active tokens, interrupt running tool calls, and prevent the agent from restarting automatically. Reversible actions, such as a draft email or branch commit, should have an undo path; irreversible actions, such as a wire transfer or production deletion, need stronger preventive controls. Incident playbooks should identify who can pause the system, who investigates, who communicates externally, and who authorizes recovery. Testing matters because a control that exists only in a slide deck provides little assurance during an incident.
Implementing the Controls in Practice
A 90-day initial program can produce useful evidence without pretending that the risk has been eliminated. During days 1–30, organizations should inventory every agent and the tools it can invoke, assign business and technical owners, classify the data involved, and identify autonomous versus approval-gated actions. They should also suspend unused credentials and remove direct production access from experimental agents. By day 30, each priority use case should have an explicit risk tier, an accountable owner, and a documented action boundary.
During days 31–60, teams should implement a tool gateway or equivalent enforcement layer, least-privilege identities, session-level budgets, log collection, and approval routes. For a high-risk action, the broker can present the intended action, target, estimated cost, data access, and reason to an authorized person. Approval should be short-lived and bound to that exact action; it should not become a reusable permission for the rest of the session. A second reviewer can be required for transactions above a defined threshold, such as $10,000, although the correct number depends on the organization’s risk appetite and control environment.
During days 61–90, the organization should run adversarial tests and controlled failure exercises. Scenarios might include prompt injection in retrieved documents, credential theft, tool-name spoofing, excessive retries, data exfiltration, conflicting objectives, and attempts to escalate privileges. Teams should verify that a blocked action does not execute, the event is logged, the alert reaches the correct owner, and the kill switch works. Results should determine whether the system is approved, restricted, or rejected. An initial control target of 100% of production agents having an owner and inventory entry is more meaningful than claiming universal effectiveness from a small pilot.
Policies and procedures remain necessary, but they should be generated from observed system behavior and incorporated into engineering requirements. A control such as “minimize data access” must become a role-based permission, query row limit, field masking rule, or approved retrieval service. A policy requiring human supervision must become an approval node with a timeout and default-deny behavior. This translation from principle to enforcement is often the difference between an AI policy program and operational risk control.
Comparing the Main Control Alternatives
Organizations can combine control categories, but they should understand what each one does and does not provide. A governance platform is useful for documenting policies, owners, approvals, and evidence, while an agent observability platform records traces and detects suspicious behavior. Neither necessarily prevents a tool call unless it is integrated with enforcement. Identity systems supply strong credentials and authorization, and gateway or broker patterns provide the most direct point for action gating. No option is sufficient alone.
| Feature | Governance and policy platform | Agent observability platform | Identity and action-control platform |
|---|---|---|---|
| Primary purpose | Define ownership, obligations, approvals, and evidence | Trace model behavior, tool calls, latency, cost, and anomalies | Enforce identity, permissions, action limits, and approval |
| Preventive control | Usually limited unless connected to enforcement | Usually weak for direct action blocking | Strong at tool and resource boundaries |
| Detective control | Moderate; depends on integrations | Strong for traces and behavioral alerts | Strong for authorization and access events |
| Best deployment stage | Portfolio inventory and risk acceptance | Testing, production monitoring, and investigation | Runtime control of high-consequence tools |
| Common weakness | Policies can remain disconnected from actual actions | High visibility does not necessarily create reversibility | Poor configuration can create a single bottleneck or unsafe policy |
| Typical cost model | Per user, workflow, or enterprise agreement | Per trace, volume tier, or agent/workload | Per identity, protected tool, request, or platform agreement |
| Appropriate expectation | Accountable decision record | Evidence of what the agent did | Demonstrable prevention and intervention |
Free and open-source tracing or gateway components can support experiments, but they do not remove enterprise responsibilities. A team may begin with open telemetry formats, policy-as-code, proxy logs, and role-based access controls, then add commercial products where support, scale, or compliance evidence is required. The decisive criterion is whether the approach produces enforceable and testable controls, not whether the tool is marketed as “agentic” or uses a fashionable architecture label.
Common Mistakes and Trade-Offs
The first common mistake is treating all agents as though they pose the same risk. A read-only research assistant and an infrastructure-management agent should not receive identical controls. Risk classification should consider autonomy, reversibility, data sensitivity, affected population, tool reach, financial exposure, and the consequences of a wrong decision. The second mistake is assuming that a more capable model is safer because it follows instructions better; greater capability can increase the speed and scale of harm when controls fail.
Another mistake is creating an “agent sandbox” that contains credentials and tools but lacks a real security boundary. A sandbox is useful only if network access, filesystem permissions, secrets, and external tools are technically isolated. Cosmetic restrictions embedded in a prompt are not equivalent. Teams also make the mistake of logging everything without defining protected retention, access controls, and redaction. Full traces may contain personal data, authentication material, source code, or confidential customer information, so observability creates its own privacy and security obligations.
Human approval can become ineffective if people receive 100 alerts a day, lack context, or approve mechanically. The approval interface should show the intended action, target, evidence, expected cost, and reason, while rejecting vague requests to “approve the agent.” Conversely, requiring approval for every harmless read may make the system unusable and train reviewers to click automatically. Controls should be proportional to consequence, with low-risk actions bounded by strict limits and high-risk actions requiring stronger verification.
No control is perfect. Blocking a known tool call can miss a novel abuse path; a monitoring system can generate false positives; a human reviewer can be deceived; and an identity platform can be misconfigured. The correct goal is to reduce likelihood, constrain impact, shorten detection time, and support recovery. Organizations should document residual risk rather than describe controls as guarantees. Claims that a system is “completely secure” should be treated as a governance defect because no current evidence supports that conclusion for a probabilistic, changing system.
When to Act and What It May Cost
Action should begin before an agent receives production credentials, customer data, or authority to change external systems. Waiting for a public incident increases urgency but leaves the organization without a tested inventory or kill switch. Immediate priorities should be systems that can access regulated information, execute code, move money, change security settings, or communicate externally at scale. A smaller organization with one internal drafting agent may begin with identity separation, logging, and a documented owner, while a bank handling payments should require action brokering, transaction limits, dual approval above set thresholds, and continuous session monitoring.
There is no universal percentage of agents that must receive human approval. A reasonable design uses a matrix with at least three practical tiers: low-risk actions can run automatically within narrow limits; medium-risk actions require sampled review or confirmation; high-risk or irreversible actions require explicit approval or are prohibited. Organizations should review the thresholds quarterly and after model, tool, data, or regulatory changes. A 2026 deployment review should also test whether new agent frameworks or protocols have changed tool behavior since the last assessment.
Budgeting should cover more than software licenses. For an enterprise deployment, costs can include model inference, retrieval and storage, identity services, policy enforcement, telemetry, security testing, legal review, employee training, and 24/7 operations. Exact figures vary greatly by scale, so vendors should be required to disclose assumptions about traces, users, tool calls, retention, environments, and support. Cheaper does not mean adequate if the product cannot export evidence or integrate with incident response. More expensive does not mean adequate either, because configuration and governance remain the customer’s responsibility.
By September 2026, the practical standard is continuous control rather than a one-time certification. Organizations should be able to answer specific questions: Which agents are running? What can each one access? Which actions are automatic? Who approved the current configuration? Can a session be stopped in minutes? What happened at 03:14 during an incident? If those answers cannot be produced reliably, the organization is not ready to grant broader autonomy. The right objective is not unlimited agent capability; it is capability whose authority, evidence, and failure modes are bounded by the organization.
AI Structural Engineering Relevance
For AI structural engineering and architecture, publishing, design, and asset-management systems, the same control logic applies with domain-specific consequences. An agent connected to BIM models, structural calculations, specifications, procurement records, or a document-control system may create or approve safety-relevant information even if it never directly constructs a structure. The relevant boundary includes drawing revisions, calculation parameters, material specifications, inspection records, and release status. A useful architecture separates proposal generation from authoritative sign-off, preserves revision history, and makes the source of each value traceable.
Controls should be tailored to engineering workflows. A design agent might propose reinforcement changes, but a licensed engineer may need to review load combinations, units, code references, and design assumptions. A document agent may summarize a clause, but it should identify the exact source and uncertainty rather than silently treating generated text as approved specification text. For manufacturing or construction operations, an agent may be allowed to create a work instruction draft but not release it to a field crew without an accountable approval path. These controls reduce both cyber risk and the risk of propagating an incorrect assumption across many downstream artifacts.
The architectural principle is “least authority with inspectable handoffs.” Tools should expose only approved interfaces; data should be partitioned by project and role; and generated changes should pass through a review boundary before becoming authoritative. Engineers should log the model version, input documents, retrieved clauses, calculations performed, and the reviewer’s decision. That record is valuable not only for cybersecurity but also for professional accountability, quality assurance, and future maintenance. The system is safer when an agent can be interrupted, compared with a source, and corrected without leaving hidden changes across a project environment.