Poxek Research · llm-firewall

Runtime boundaries matter more as agents gain tools

Why tool-using AI needs capabilities, authorization, and evidence outside the model.

Instructions are not authorization

An LLM can propose a useful action, select a tool, and explain why it chose that tool. It cannot be the authority that decides whether the action is permitted. That distinction becomes operational as soon as an application can read a mailbox, query a production service, modify a ticket, send a message, browse a website, or run a workflow.

The model is exposed to inputs that are only partly under the application's control. A user can write an instruction directly. A retrieved document, webpage, email, image, or tool result can carry an indirect instruction. OWASP describes prompt injection as a change in model behavior caused by such input and notes that there is no foolproof prevention. The practical consequence is not to abandon prompts. It is to ensure that a changed model decision cannot silently become an unauthorized system action.

That means treating an agent as a component in a larger authorization system. Prompts can state task intent and help the model make a plan. Runtime controls must decide what the plan is allowed to do.

The boundary belongs at execution time

A useful agent boundary starts with a concrete question: which principal may perform which operation on which resource, under which conditions? The principal may include the end user, the workload, the agent identity, and the service account used for a tool call. The operation should be narrower than a conversational goal. “Help with customer support” is not an enforceable permission; “create a draft reply for this ticket” or “read the orders visible to this user” is closer to one.

The same distinction applies to tools. An open-ended shell, browser, database client, or HTTP fetcher gives the model a wide action surface even when the intended task is narrow. OWASP lists excessive functionality, excessive permissions, and excessive autonomy as common roots of excessive agency. A read-only search tool and a bounded record lookup are easier to authorize and investigate than a general connector with modification privileges.

Runtime policy can therefore evaluate an action before the downstream system receives it. Typical conditions include:

  • the agent and workflow identity;
  • the authenticated user or service on whose behalf it acts;
  • the selected tool, operation, destination, and data class;
  • the current session, tenant, environment, and time window;
  • whether the operation is read-only, reversible, high impact, or requires confirmation.

This does not replace the receiving system's own authorization. Database roles, API scopes, network egress controls, and approval steps remain authoritative at their own boundaries. NIST's zero trust guidance likewise focuses on protecting resources and on discrete authentication and authorization rather than granting implicit trust because a component is inside a familiar network. An agent control plane should complement those mechanisms, not become a second, weaker identity system.

A policy decision needs an enforcement point

There are two different jobs in a runtime boundary. A policy decision determines whether an action matches the workload's allowed capability. An enforcement point ensures that a denied action does not reach the tool, and that an allowed action carries only the approved identity and scope. Mixing these jobs inside a system prompt leaves both dependent on the model continuing to follow instructions.

Consider an agent that summarizes account activity and can create a support case. A safe design need not give it generic access to every customer record or a generic “send email” function. It can expose a read operation scoped to the current customer, a case-creation operation with a constrained schema, and a human approval step before an external communication is sent. If retrieved text tells the agent to search other accounts or forward data to an unfamiliar destination, the policy and the destination's access control can reject the request even if the model repeats the instruction.

This architecture also clarifies what a control cannot decide. A policy engine can enforce a known rule such as “this workflow cannot delete records” or “this tenant may not send data to this destination.” It cannot prove that a natural-language request is harmless, that a retrieved document is truthful, or that a model's summary is correct. Those are evaluation, data-governance, and human-review problems. Treating a semantic classifier as final authorization would recreate the same dependency on probabilistic judgment at a more polished layer.

Observability must include the attempted action

Logs that contain only a final chat response do not explain an agent event. An investigator needs to connect the interaction to the model invocation, retrieval source, proposed tool call, policy decision, tool result, identity context, and downstream state change. The exact content may be sensitive, so collection and retention should be intentional: identifiers, classifications, and hashes can sometimes support correlation without recording every prompt or tool argument.

The record should distinguish at least four outcomes: proposed, allowed, blocked, and executed. “Allowed” is not evidence that the destination completed the request; “executed” should reflect a response or confirmation from the downstream system. This distinction avoids a common investigation failure where an AI layer says it performed an action but the trace cannot show whether the action was sent, rejected, retried, or partially completed.

Telemetry must be protected as carefully as the workload it describes. Prompt and tool content can contain personal data, secrets, business records, or attacker-controlled text. Access controls, minimization, tenant separation, retention limits, and redaction rules are design requirements, not dashboard settings. An audit trail that becomes a second ungoverned data store adds risk instead of reducing it.

Design for compromised instructions

The relevant question for a future agent system is not whether every prompt injection can be detected. It is whether a compromised instruction can exceed the authority intentionally granted to that workflow. OWASP's guidance on system prompt leakage is direct: a system prompt should not be treated as secret or as a security control, and critical authorization bounds must be enforced independently from the LLM.

For teams designing agent workloads, that leads to a short set of architectural checks:

  • Model-visible instructions contain no credentials or hidden authorization logic.
  • Each tool exposes the minimum operation set required for the workflow.
  • Downstream calls use scoped identities and preserve the initiating user or workload context where possible.
  • High-impact, irreversible, or externally visible actions have a deterministic approval or confirmation boundary.
  • Policy decisions and execution results are recorded in a reviewable, access-controlled trail.
  • Tests include hostile content from user input, retrieval, files, webpages, and tool responses.

Poxek LLM Firewall is a private-preview runtime control plane for expressing workload capability boundaries, observing model and agent traffic, and making policy decisions available to security and platform teams. It is not a claim that an inspection layer eliminates prompt injection, model error, or unsafe integrations. The model can be helpful without becoming the final authority over production resources.

Sources

  1. OWASP LLM01:2025 Prompt Injection
  2. OWASP LLM06:2025 Excessive Agency
  3. OWASP LLM07:2025 System Prompt Leakage
  4. NIST SP 800-207: Zero Trust Architecture