Poxek Research · llm-firewall

An LLM firewall is a runtime control plane, not a prompt filter

A product view of capability boundaries, traffic monitoring, and governed agent execution.

The product question is where a model's suggestion becomes an action

Most LLM applications begin with a prompt boundary: instructions define the assistant's role, expected behavior, and prohibited outputs. That work is useful, but it is not sufficient when the same application can call tools or access production data. A prompt is input to a probabilistic model. Authorization, approval, and downstream validation need to remain deterministic controls owned by the systems that hold the resource.

Poxek LLM Firewall is a private-preview product for that runtime layer. Its role is not to pronounce a conversation safe or unsafe in isolation. It helps teams describe what an AI workload may do, observe the traffic that leads to a decision, evaluate actions against explicit policy, and preserve evidence for review. It operates in the space between model output and consequential execution.

That scope is deliberately narrower than a promise to “secure AI.” A control plane cannot make a model truthful, prevent every prompt injection, or repair an application that gives a broad service account to an unsafe connector. It can provide a place to make authority visible and to enforce defined boundaries before a governed integration is invoked.

Capability boundaries turn intent into enforceable rules

An agent's goal is normally expressed in natural language: investigate an alert, summarize a case, prepare a change, or answer a customer. A runtime policy needs concrete objects instead. It should identify the workload, its permitted tools and operations, the target resources, the acting identity, the data sensitivity, and any approval requirement.

For example, an investigation assistant might be able to read a defined alert index and create a draft case note. It should not inherit delete permissions, unrestricted HTTP egress, or the right to apply a configuration change merely because another workflow in the organization needs those functions. The policy model should make the difference between read alert, create draft, and apply change explicit.

OWASP's excessive-agency guidance makes this design principle practical: reduce unnecessary extensions, reduce the functions inside each extension, use granular operations rather than open-ended tools, and minimize downstream permissions. Those decisions reduce impact even if the model follows a malicious instruction or simply makes a poor choice. They also leave less ambiguity for a reviewer who must explain why an action was available.

The Poxek approach expresses these workload-specific boundaries at runtime. A policy decision can consider an agent or workflow identity, user context, tenant, tool and operation, resource or destination, request classification, and action risk. A deny result should stop the request before it reaches the governed tool. An allow result should not bypass authorization in the target system; the target still needs its own scoped identity checks and validation.

Traffic monitoring should connect decisions to execution

AI traffic is more than prompt text. In a tool-using workflow, a single user request may lead to retrieval, several model calls, tool proposals, policy evaluations, approval steps, retries, and a downstream state change. Looking only at a chat transcript loses the chain that matters during a security review.

The intended Poxek LLM Firewall record connects relevant stages of that chain: workload and session identifiers, model and operation metadata, tool selection, policy outcome and reason, approval status, downstream result, and timestamps. Where full content collection is justified, it must be treated as sensitive telemetry, not as harmless debugging material. Where it is not justified, a team may prefer structured metadata, classifications, controlled excerpts, or correlation identifiers.

This distinction supports a practical question: what happened after the model proposed an action? A policy event marked allowed is different from an execution event marked completed, and both differ from a tool failure or a downstream authorization denial. Retaining those states makes it possible to investigate failure paths without assuming that a natural-language answer is evidence of an external effect.

Govern the model-facing and non-model-facing layers separately

The LLM should never become the sole gatekeeper for security-critical behavior. OWASP recommends treating model output as untrusted before it is passed to a backend function, with validation and context-aware handling at the receiving boundary. This applies even if a separate prompt filter marks the output as acceptable.

The resulting architecture has several layers with distinct jobs:

  • Model layer: prompt instructions, model configuration, and output shaping to support the task.
  • Runtime policy layer: capability allowlists, constraints, policy decisions, and approval requirements for known actions.
  • Integration layer: typed tool interfaces, schema validation, scoped credentials, rate limits, and destination-specific guards.
  • Resource layer: the target service's authorization, validation, transaction controls, and audit trail.
  • Governance layer: ownership, review cadence, exceptions, incident handling, and evidence retention.

These layers overlap by design. A policy deny can prevent an unnecessary call; a tool schema can reject an invalid argument; a destination can refuse a token that lacks a required scope. A missing control at one layer should not turn the model's text into a privileged command at the next.

NIST's Generative AI Profile treats risk management as a lifecycle activity and includes suggested actions around incident monitoring, after-action review, and system inventory information. A product control plane can support that work only if its operational records are connected to named owners and a defined response process. More telemetry without ownership is merely a larger investigation queue.

What private preview means here

Poxek LLM Firewall is in private preview. Its scope centers on three capabilities: expressing capability boundaries for AI workloads, monitoring model and agent traffic, and producing policy and execution context that security and platform teams can review.

It should not be treated as a substitute for secure application design, identity and access management, data classification, output validation, sandboxing, red-team testing, or human ownership of high-impact decisions. It also does not guarantee detection of malicious instructions, accurate model reasoning, regulatory compliance, or prevention of all data disclosure. Those claims would confuse a control surface with a complete security program.

The useful evaluation question is concrete: can a team state what a given workflow may do, prevent it from doing what it may not, and reconstruct the decision when something goes wrong? If the answer is no, the next step is usually to narrow the tool or credential first, then decide what runtime policy and telemetry are needed around it.

Sources

  1. OWASP LLM05:2025 Improper Output Handling
  2. OWASP LLM06:2025 Excessive Agency
  3. OWASP Agentic AI: Threats and Mitigations
  4. NIST AI 600-1: Generative AI Profile