Direct answer: what runtime controls mean for AI agents

Runtime controls for AI agents are policies and technical mechanisms applied while an agent is running, rather than only before deployment or after an incident. They can restrict which tools an agent may call, limit spending and execution time, inspect prompts and tool arguments, require human approval for sensitive actions, and terminate or quarantine a session when behavior exceeds policy. This differs from ordinary application permissions: a service account may be allowed to read a database, while an individual agent should not be able to export every row from it. As of 29 September 2026, the control problem has moved beyond basic IAM because agents can plan, generate code, call APIs, and make decisions across multiple systems. NVIDIA’s OpenShell direction, Okta’s agent-runtime security architecture, and newer products such as Agno, G0, Traccia, and Alterion Draco all reflect the same market shift toward a runtime control plane. The right answer is not a single product or a universal “kill switch”; it is a layered control system that preserves useful autonomy while reducing the blast radius of mistakes, prompt injection, excessive tool use, and compromised dependencies.

Also worth reading: How Can eBPF Security Controls Protect AI Agents Without Slowing Down Inference? · How Should Structural AI Risk Controls Be Applied in Engineering and Infrastructure Projects? · How Should Organizations Control Runtime Agent Access in 2026?

How runtime controls actually work

A typical agent runtime receives a user request, selects a model, generates a plan, invokes tools, receives results, and repeats until it reaches an answer or an execution limit. Controls are inserted at each transition. Input controls can detect secrets, malicious instructions, excessive context, or requests outside the agent’s assigned role. Planning controls can cap the number of steps, estimated cost, token volume, and allowed destinations. Tool controls can inspect arguments, validate schemas, enforce least-privilege credentials, and prevent irreversible operations without approval. Output controls can scan for sensitive data, validate structured responses, and block policy violations before the result reaches a user or another agent. Finally, session controls can pause, revoke, roll back, or terminate execution. These mechanisms work best when they are enforced outside the model itself, in a proxy, gateway, sandbox, policy engine, or tool broker. Asking an LLM to “follow these rules” is useful as a behavioral instruction, but it is not a reliable security boundary because the same model may be manipulated by untrusted content.

The main control categories and measurable thresholds

The first category is identity and authorization. Each agent should have a distinct identity, short-lived credentials, and permissions tied to a specific task, environment, and time window. If an agent is responsible for creating support tickets, it should not automatically receive production deployment rights. The second category is tool governance: maintain an allowlist of tools, constrain parameters, require typed outputs, and separate read operations from write, delete, payment, and privilege-changing operations. The third is resource governance, with explicit limits such as a maximum of 10 tool calls for a simple classification task, 50 for a research workflow, or a fixed 15-minute execution window. Organizations should also set budgets, for example stopping a run after 2,000 model tokens or a configured dollar amount. The fourth category is behavioral monitoring, which records prompts, tool calls, latency, errors, approvals, and state changes. The fifth is emergency response, including a session kill switch, credential revocation, quarantine, and a clean incident record. Thresholds must be workload-specific. A threshold that is appropriate for a coding agent may be too restrictive for a research agent, while an unrestricted research agent may create unacceptable cost and data exposure.

A practical implementation sequence

Start by inventorying what the agent can do rather than beginning with a large policy platform. Map every model, tool, API, data source, identity, and human approval point, then classify actions by reversibility and sensitivity. Reading a public document and sending an email are different risk classes; deleting a database and changing an IAM policy are higher still. Next, place a policy enforcement point between the agent and each tool. Replace broad credentials with scoped, short-lived tokens, and route calls through a broker that can validate arguments and enforce rate, cost, and time limits. Add a sandbox for code execution, network egress controls, and filesystem isolation. Then test with adversarial cases: indirect prompt injection in retrieved documents, malicious tool output, credential requests, infinite loops, unexpected data volume, and attempts to bypass an approval requirement. A sensible initial target is to block 100% of unapproved production writes, sensitive-data exports, and use of credentials outside the agent’s assigned role. These are stronger operational objectives than claiming that the agent is “safe.”

After basic enforcement is working, add observability and approval workflows. Every tool call should produce an event containing the agent identity, task identifier, policy decision, input hash, destination, cost, result status, and correlation ID. Sensitive values should be redacted or encrypted rather than copied indiscriminately into logs. Human approval should be reserved for defined transitions, such as external publication, financial transactions, production changes, or access to regulated records. Approvals should be short-lived and bound to the exact action, not a general approval for an entire open-ended session. Finally, rehearse failure. On 29 September 2026, teams should know who can stop an agent within minutes, how credentials are revoked, which systems are quarantined, how customers are notified, and how they can determine what the agent did. Runtime controls are operational infrastructure, so they need runbooks and exercises just as conventional security systems do.

Comparing the available approaches

There is no single category that covers identity, execution, observability, and emergency response equally well. The following comparison is directional rather than a vendor scorecard; actual capabilities change quickly, and deployments must be verified against the product documentation and a proof of concept.

FeatureModel or prompt controlsGateway or policy engineSandbox or execution runtimeFull agent control plane
Main purposeGuides model behaviorEnforces routes, tools, and policiesIsolates code and limits actionsCombines governance, identity, monitoring, and response
Defense against prompt injectionWeak to moderateModerate when policies inspect contextModerate for contained executionModerate to strong across layers
Tool-level authorizationUsually inconsistentStrongStrong inside the sandboxStrong and centrally governed
Cost and step limitsOften advisoryStrongPossibleStrong
Human approval workflowsBasic or customGoodDepends on integrationUsually integrated
Audit and incident responseLimitedGood at gateway eventsGood for execution tracesBroad, session-level records
Best useLow-risk assistantsProduction API-mediated agentsCode and data processing agentsEnterprise or multi-agent operations
Typical trade-offCheap but not a security boundaryRequires careful policy designAdded infrastructure and latencyHigher cost and implementation effort
Open-source frameworks can provide scheduling, tool abstractions, and tracing, but governance is often assembled separately. A commercial control plane may reduce integration work, yet it introduces vendor dependency, pricing, and questions about data residency. NVIDIA OpenShell and related runtime-security efforts are more relevant where agents execute code, interact with infrastructure, or operate in environments requiring hardware- and system-level isolation. Okta and other IAM providers are more relevant where identity, lifecycle, and credential policy dominate. A practical architecture commonly combines these rather than selecting one winner.

Alternatives, limitations, and what not to buy

Runtime controls are not equivalent to model alignment, red teaming, or conventional observability. Alignment can improve behavior, but it does not guarantee that a compromised retrieval document cannot cause an unintended tool call. Observability tells you what happened after the system emits enough signals; it does not necessarily stop the action. Secrets management protects credentials, but it does not decide whether the agent should use a credential for a particular destination. A firewall or egress proxy can block network destinations, but it cannot decide whether a permitted API operation is sensible for the user’s request. These controls are complementary. Buyers should be skeptical of claims that a product provides “complete agent safety” without defining the threat model, enforcement location, policy language, failure mode, and audit evidence. They should also ask whether a prompt-injection test was performed against retrieved content, whether an agent can call a tool directly, whether a human approval can be bypassed, and whether the control remains active when a model endpoint is changed.

Pricing is usually not a single industry-wide figure. Open-source components may be free to install, while hosted tracing, policy, evaluation, and security services commonly use combinations of per-agent, per-session, per-event, per-tool-call, or monthly platform fees. Infrastructure costs also matter: sandboxes, proxies, vector databases, logging storage, and human review can dominate the nominal software subscription. A small team might begin with an open-source gateway, structured logs, a secrets broker, and a simple approval service. A regulated enterprise should budget for integration, policy engineering, security testing, retention controls, and 24/7 operations in addition to licenses. The presence of funding announcements, such as Arrakis’s reported $8 million round, signals investor interest but does not validate product effectiveness or establish a standard price. No percentage of the market can responsibly be inferred from these announcements, and no vendor should be described as essential without a deployment-specific evaluation.

Common mistakes and when to act immediately

The most common mistake is allowing the agent to inherit the developer’s broad permissions. Another is treating prompt instructions as enforcement. Teams also over-log raw prompts and tool results, creating a new data-leakage problem, while failing to log policy decisions and denials. Other errors include approving an entire workflow instead of one irreversible action, setting no maximum runtime, permitting unrestricted network egress, and testing only known attacks. A further mistake is deploying a control plane without an owner: when a session is stuck or a credential is exposed, nobody knows whether to pause the task, revoke access, or preserve evidence. Do not wait for a major incident if an agent can write to production, execute generated code, access regulated information, spend meaningful money, or act on behalf of external parties. These conditions justify controls before broad release. For a read-only internal search assistant with no sensitive data and a fixed corpus, a lightweight gateway and logging policy may be enough initially. The risk changes as autonomy, tool count, data sensitivity, and number of connected users increase.

The recommended architectural pattern

The strongest pattern is defense in depth: an identity layer, a policy decision point, a tool broker, an isolated execution environment, an event stream, and an independent emergency control. The identity layer issues a unique credential for each agent and task. The policy engine evaluates the requested action, resource, context, and risk level. The tool broker supplies only approved capabilities and transforms untrusted inputs into validated parameters. The execution environment runs generated code with restricted files, memory, CPU, and network access. The event stream records decisions and actions with sensitive fields redacted. The emergency control remains usable even if the model provider, agent framework, or policy service is unavailable. This architecture also supports multi-agent systems: one agent should not silently inherit another agent’s authority merely because they share a message. A control plane should make delegation explicit, including the source identity, permitted scope, expiration, and audit record. In practice, a lightweight implementation can follow this pattern in weeks, but production readiness depends on integrations, data classification, and the organization’s tolerance for residual risk. The correct target is measurable containment, not perfect prevention.

How to evaluate a runtime-control product

Evaluation should be scenario-based and reproducible. Build a test suite containing direct prompt injection, indirect injection through a web page, malicious tool output, unauthorized data access, excessive looping, budget exhaustion, and attempts to perform high-impact actions. Measure prevention rate, false-positive rate, approval latency, time to revoke access, completeness of audit logs, and total cost per successful task. For example, require zero unauthorized production writes, zero unapproved access to restricted records, and 100% identity attribution for tool calls. Also test bypass paths: direct API access, alternate credentials, alternate tools, nonstandard endpoints, and changes to the model or prompt. Compare a minimal policy layer with a more complete control plane rather than treating feature counts as proof. Ask whether policies are deny-by-default, whether failures fail closed for sensitive operations, whether decisions are explainable, and whether the system can export evidence to an existing SIEM. A product that adds a dashboard but leaves enforcement to the model should not receive a high security rating. The final decision should include total cost, operational burden, portability, and the difficulty of removing the vendor later.

The bottom-line recommendation

For 2026, the best runtime controls for AI agents are enforced outside the model and applied continuously across identity, tools, execution, data, and response. Organizations should begin with an inventory and threat model, then implement short-lived scoped identities, an allowlisted tool broker, step and cost ceilings, sandboxing, redaction, immutable audit events, approval gates, and a tested shutdown procedure. They should add a commercial or open-source control plane only where it closes a measured gap; no framework or vendor can replace policy design or incident response. The maturity milestone is not that an agent never fails, but that failures remain bounded, attributable, reversible where possible, and visible to operators. As of 29 September 2026, runtime governance is becoming a normal part of agent architecture rather than a specialized research topic. Teams should act before exposing agents to production systems, especially when the agent can execute code, access sensitive information, spend money, or make irreversible changes. The correct level of control follows the level of authority granted.