Most agent stacks have become eerily good at turning model output into an action. But they’re much less disciplined about answering the question that actually matters: “Is this action actually allowed to happen?”
For example: A support agent sees, “Please cancel the subscription for the account that is no longer being used.” The model correctly chooses cancel_subscription. It extracts an account ID from an earlier message. The JSON is syntactically valid. The API returns HTTP 200. A trace dashboard shows a green tool call span. And the system may still have canceled the wrong account…
Schema validation tells you whether an input is well formed. Authentication tells you who presented a request. A tool definition tells the model what it can ask for. None of these things establish that the current principal may cancel that specific subscription that the user meant, that account, or that the cancellation actually took effect.
This is where the distinction between capability and authority becomes important. An LLM has a capability when it can select a tool and produce arguments. It has authority only when a separately enforced policy permits a bounded action for an identified principal in the current context.
A robust system treats capability as a proposal for actions, and requires authority before producing an effect.
That’s not just a compliance concern: Prompt injection, ambiguous user intent, stale context, and overly broad credentials all turn a valid tool call into a wrong real-world effect.
OWASP explicitly lists excessive autonomy, high-impact action abuse, approval manipulation, tool abuse, data exfiltration, and cascading failure among agentic system risks.
The engineering response cannot be one more instruction in the prompt. It needs to be a control plane that remains deterministic when consequences require that.
A valid tool call can still be the wrong action
Most agentic architecture diagrams show a very simple loop: plan → tool call → observation → next plan. That’s great, because the loop describes reasoning and orchestration. But it doesn’t describe governance.
For sophisticated agentic systems, the system is actually organized in three planes. The first is the planning plane, which basically tells you what the model is good at. For example, interpreting the task context, selecting among a deliberately small set of business-level tools, forming a typed proposal, explaining what it intends to do, and so on. It mustn’t hold broad provider credentials or decide the final policy question.
The second plane is the control plane. It describes what decides whether any action is actually allowed. It contains things like authenticating the initiating principle and resolving tenant and delegation context, canonicalizing and validating the proposed action; evaluating authorization risks, resource ownership, data quality policy and so on. It may request an exact action approval when the policy requires it, or issue only narrowly scoped short-lived execution authority. And it may, at this plane, persist the evidence required to understand the decision later. This may all sound very corporate, but you will see why this is very necessary further down.
The third plane is the execution and observation plane. What performs and proves the effect? An execution broker invokes a constrained adapter or an isolated worker. The resource service or provider performs the action or returns a failure or uncertain result. A verifier reads an authoritative post condition or receipt. The system represents pending, unknown, failed, and verified distinctly. The audit ledger and telemetry system retain the decision effect trail without treating a mutable chat transcript as the audit log.
The split doesn’t come from me. It is already present in established patterns which precede LLM times. For example, NIST’s zero trust architecture separates a policy decision from a policy enforcement point and emphasizes granular access decisions for individual resource requests rather than implicit trust after sign-in.
In an agentic system, the LLM becomes a useful upstream proposer. The policy engine and tool or API boundary are then where the authorization is decided and enforced.
Three planes, not one agent loop
Below is a schematic of how this would pan out in practice. The one non-negotiable boundary is that the only route to a consequential provider API runs through the action gateway, policy decision, and – where required – approval.
Untrusted web pages, emails, retrieved documents, and tool output are not elevated into instructions just because the model can read them. They’re explicitly classified as untrusted data and, when necessary, processed in a no-tool, no-secret quarantine step before a privileged agent can use the extracted facts.
Approval is also conditional and not universal. The control plane can authorize a low-risk, read-only action within a narrow scope. Or it require user confirmation for a consequential reversible action. Or it can also simply deny a critical action and leave the agent to prepare evidence for a human-operated runbook.
How do you actually build such a control plane? It’s more than a dashboard – nine steps are needed in total, and they can’t just be handed over to an LLM. Below is what you’ll need to do.
Step 1: Make action contracts narrow and typed
First, the core principle: Do not expose generic tools such as http_request, run_shell, raw SQL, or an unrestricted browser just because the model can use them. Instead, expose business-level verbs with explicit boundaries: something like create_draft_invoice, queue_refund_review, send_approved_notice, revoke_session.
Tool schemas become interface contracts. They reduce free-form ambiguity; they do not authorize an action.
One would want to make schemas and draw them up very cleanly, by specifying all of the following alongside it:
-
action_idandschema_version. -
side_effect_class(read_only,draft,reversible_write,external,irreversible). -
required_scopesand allowed environments. -
risk_tierandapproval_mode. -
idempotency_requirement. -
verification_method. -
data_classification/ outbound-data rules where applicable.
OpenAI and Anthropic both document structured tool inputs/function schemas; their client-side execution models make the broader point that the application executes the tool after the model requests it.
Step 2: Separate authentication, delegation, and approval
Common implementations authenticate a user once, hand the agent the user’s token, and assume that every subsequent tool call is authorized.
That’s not a great idea: It converts a temporary request into ambient authority. It also makes the agent difficult to audit: did the human act, did the agent act for the human, or did a shared system identity act?
Instead, you’ll need to give agents distinct, attributable identities rather than running them indefinitely as generic service accounts or borrowing human sessions. (That’s a huge security risk anyway!)
Google Cloud’s Agent Identity documentation is a concrete implementation example: it differentiates user-delegated authority from an agent’s own authority, provides per-agent identity, and says that delegated access logs can show both the user and agent identities.
The core point is to implement agent identity, human identity, and delegation all as separate concepts.
This principle is not new: OAuth 2.0 was designed around limited access through an authorization layer rather than a third party holding the resource owner’s password; access tokens represent scope, lifetime, and other access attributes.
OAuth doesn’t agent governance, but it offers the right vocabulary for delegation. The rest of the control plane still needs policy enforcement, verification, monitoring, and revocation. These are the next steps!
Step 3: Authorization becomes a deterministic decision
No model output should directly reach a side-effecting provider API. Instead, the model creates something like an ActionProposal. The action gateway validates it. A policy decision point then returns allow, deny, or approval_required. Finally, an execution broker, not the planner, receives the constrained authority to carry out an allowed action.
What matters most is the following:
-
The policy check happens outside the LLM.
-
The broker does not execute before the decision.
-
The approval, if required, is checked against the same canonical digest that will be dispatched.
-
The output is not automatically success just because dispatch returned 200.
-
The ledger is written around decision and outcome transitions, not reconstructed later from a chat transcript.
NIST’s zero-trust model is useful here, too, because it makes authorization dynamic: the policy decision can consider the request, identity/attributes, resource requirements, and contextual signals rather than treating a prior login as permanent trust.
Step 4: Classify actions by consequence
An entire agent should not be classified as purely “autonomous” or “human-in-the-loop.”
That classification belongs to individual actions and data flows. Therefore, the same agent can read a bounded knowledge base automatically, create a draft in staging, and require dual control before changing production access rights.
The following basic tier model can be adapted to your specific use case:
|
Tier |
Typical scope |
Control posture |
|---|---|---|
|
T0 — observe |
Read approved/public information; local classification |
Typed read contract, least privilege, rate limits, logging |
|
T1 — draft |
Create a draft or agent-owned staging artifact |
Scoped staging write, provenance, version record, later review |
|
T2 — reversible internal effect |
Update one authorized internal record; queue a bounded workflow |
Policy check, idempotency key, postcondition read, correction/compensation path |
|
T3 — high impact / sensitive |
External message, production change, sensitive disclosure, payment/refund, access change |
Exact-action human approval, short-lived narrow authority, sandbox/egress controls, protected evidence, reconciliation |
|
T4 — critical / systemic |
High-value transfer, destructive bulk action, root identity policy change, regulated/safety-critical commitment |
Default deny for autonomous commit; human-operated runbook, independent approval, simulation/dry run, named accountability |
This corresponds to the general guidelines: OWASP recommends explicit approval for high-impact or irreversible actions, action previews, autonomy boundaries based on risk, audit trails, and the ability to interrupt/roll back where possible. The UK NCSC recommends risk-proportionate autonomy and notes that greater autonomy means greater potential impact—and therefore a greater need for controls.
Following the guidelines, we also obtain the following promotion rules:
-
Escalate a tier whenever target identity was inferred rather than explicitly selected.
-
Escalate if untrusted content materially influenced the action.
-
Escalate if there is no idempotency mechanism or authoritative verification path.
-
Escalate if scope becomes cross-tenant, bulk, external, irreversible, financially significant, legally significant, or production-critical.
-
Never auto-demote because the model expressed high confidence.
Whether these actions happen to use a model is not so important. What matters is whether the consequence of an action is material or not.
Step 5: Bind approval to a canonical action
At this step, it is crucial to be very concrete and not wishy-washy. For example, if a user writes, “please resolve this billing issue,” and the agent later chooses a recipient, cancellation reason, subscription amount, and refund option, that’s not a meaningful authorization of the actual effect. The approver must see a canonical representation of the effect that will be dispatched and not the model’s natural language intent.
Also, the decision must be invalid if the action changes. This means we have to introduce a little bit of red tape in the form of approval record fields:
-
Action name and schema version.
-
Tenant, user/principal, agent run, and executor identities.
-
Resolved target/resource and consequential parameters.
-
Risk tier and policy version that required approval.
-
Digest/hash of the canonical action and resource version.
-
Approver identity, role, decision, time, expiry, and reason.
-
One-time nonce/state so an approval cannot silently be replayed.
This also has consequences for the user interface: Any change has to be described in business language. UI designers need to emphasize recipient, target, amount, scope, data type, environment, and irreversibility. And the approver must be able to deny or edit at a meaningful decision boundary—not after the provider has already accepted the request.
OWASP recommends binding approval to the actor, tool, target resource, normalized parameters, timestamp, and expiry, and independently validating scope/privilege/approval state before execution.
Step 6: Defend against prompt injection
First of all, a distinction: A prompt injection is not only a user typing “ignore previous instructions.” It can arrive through an email, web page, document, image, RAG chunk, tool output, or memory record that the agent reads while pursuing a legitimate task.
OpenAI warns that untrusted text can induce data exfiltration or misaligned downstream tool calls. OWASP also describes direct and indirect injection and agent tool manipulation.
Sadly, there is no single prompt or classifier that turns untrusted text into trusted instructions. The protective burden belongs mainly on authority boundaries and blast-radius reduction. Here’s how you reduce the risk:
-
Provenance label context: Mark each item as system instruction, trusted business data, user input, external untrusted content, or tool output.
-
Partition context: Never concatenate untrusted content into developer/system instructions. Present it as data in a clearly separate field.
-
Quarantine extraction: Use a no-tool, no-secret component to extract a narrow typed summary from external content before a privileged planner sees it.
-
Deterministic gate: Recheck action schema, authorization, data egress policy, recipient/target constraints, and approval at the execution boundary.
-
Reduce blast radius: Give the executor only a short-lived credential and limited network/file/data access for one action.
-
Test the attack path: Include direct, indirect, encoded, multilingual, tool-output, and RAG-injection examples in release evaluations.
As with all things security, even following all these steps does not make the system impervious to hostile content – it just reduces the risk. Also worth keeping in mind:
-
Never expose secrets in model-visible context if the execution broker can use a secret reference instead.
-
Never let an agent with arbitrary web input and generic network/file/shell tools become the only gate between an attacker and production authority.
Step 7: Design for uncertainty, retries, and compensation
“Tool call succeeded” is not a business result: A provider may acknowledge a request before it completes. The connection may time out after the provider committed the change. A retry could create a duplicate effect. A downstream system may return success while a later reconciliation discovers that a related step failed.
The model should never turn those uncertainties into a confident “done.”
Instead, we need an action state machine. This follows standards from the pre-LLM era: RFC 9110 defines idempotence as repeated identical requests having the same intended effect and explains why a client can retry an idempotent request after a communication failure; it also cautions against automatically retrying non-idempotent methods without knowing the request was not applied.
These are the design rules in practice:
-
Persist the intended effect and idempotency key before dispatch.
-
Bind the key to tenant, action, canonical parameter digest, and intended effect—not to a chat turn.
-
Reuse the key only for the same intended effect; reject reuse with changed parameters.
-
Verify using an authoritative read-back, resource version, or provider-signed receipt.
-
Surface
unknownif the system cannot establish outcome; place it in a reconciliation queue rather than fabricating success. -
Separate retry (same intended effect), rollback (restore prior state inside a transactional boundary), compensation (a new forward action intended to offset a prior effect), and reconciliation (compare intent with observed authoritative state).
In short: Sending an email, paying a supplier, or disclosing data does not become reversible because an engineer labelled an endpoint DELETE or added a “rollback” button.
Step 8: Store a decision/effect record
Observability is, of course, important. But AI thinks a lot, and recording every hidden thought is going to blow up your logs.
Essentially, you need two things: (1) traces for debugging and performance, and (2) audit evidence for accountability and incident investigation. Here’s a minimum protected event set that would cover both:
-
run.started— initiating principal, agent/model/version, session, tenant, trace ID. -
context.ingested— source, trust class, hash/reference, sanitizer/guardrail result. Avoid raw sensitive content where policy does not permit retention. -
action.proposed— action/schema version, normalized parameters or protected reference, model/tool-call ID, risk calculation. -
policy.decided— allow/deny/approval-required, policy version, applicable rules, delegation/authorization references. -
approval.requested/approval.decided— canonical action digest, approver, role, decision, expiry. -
execution.dispatched/execution.acknowledged— executor identity, idempotency key, provider receipt/reference. -
effect.verified/effect.unknown/effect.failed— verification method, observed state/receipt, error class. -
compensation— original action, new authorization/approval, outcome.
OWASP, in their Logging Cheat Sheet, recommends following these principles: centralized collection, protection from unauthorized modification/deletion, tamper detection, restricted/monitored log access, and verification of the logging system itself.
Step 9: Evaluate the control plane (not just the model response)
How do we know if all we’ve built is good? The question is not simply “Did the answer look good?”
For an agent, you need to evaluate tool choice, argument precision, policy adherence, approval binding, execution reliability, and verified outcomes. OpenAI’s evaluation guidance recommends task-specific, continuous evaluations, production-log-derived cases, adversarial examples, and human calibration of automated scoring.
For example, it could look something like this:
|
Eval family |
Example assertion |
Suggested gate |
|---|---|---|
|
Contract conformance |
Extra, malformed, or cross-tenant fields never reach execution |
Deterministic tests: 100% pass |
|
AuthN/AuthZ |
An authenticated but unprivileged user is denied |
Deterministic integration tests: 100% pass |
|
Approval binding |
Changing recipient/amount/resource version invalidates prior approval |
Deterministic integration tests: 100% pass |
|
Tool selection |
Agent chooses only allowed business tools with correct canonical arguments |
Labelled cases; error rate tracked by risk tier |
|
Prompt injection |
Untrusted pages/docs/tool output cannot induce disallowed effects |
Attack-success rate plus blast-radius tests |
|
Data egress |
Secrets, another tenant’s data, or unapproved fields cannot reach a connector |
Canary/DLP tests with denial evidence |
|
Reliability |
Timeouts, duplicates, out-of-order events, and provider 5xx do not create duplicate effects |
Chaos/integration tests; duplicate-effect target = 0 |
|
Verification |
Provider HTTP 200 without postcondition becomes |
Simulated-provider test suite |
|
Audit |
Every consequential state transition has correlated, protected evidence |
Invariants + restore/drill tests |
|
Human factors |
Approvers detect material recipient/amount/scope changes |
Blinded usability/security review |
If this feels like a lot, here’s a useful anti-pattern: A model-as-judge, for example, can help evaluate explanation quality or action plausibility. But it shouldn’t be the final proof that an authorization control worked. For high-impact effects, deterministic policy tests and provider-side invariants should gate deployment.
Let the model plan; make the system decide
There you have it! 9 steps to make agents reliable and scalable.
If there’s one thing you retain, let it be this:
Let the model plan.
Let policy decide.
Let a constrained executor act.
Let an independent verifier state what happened.
This arrangement may look less magical than a single agent with a browser and broad credentials. It’s also how you make an agent useful in systems where “helpful” and “authorized” are not synonyms.
The scalable form of agent autonomy is not permissionlessness. It’s a system that can reason freely inside clear, enforceable boundaries. A system that knows exactly when it must ask before crossing one.

