Poxek Research · llm-firewall
AI traffic monitoring needs execution context
Why LLM and agent observability must connect model activity to policy, tools, identities, and outcomes.
A transcript is not an operational record
Teams often discover an AI incident through a chat transcript. The transcript may show a surprising answer, a model refusal, or a sentence claiming that an agent completed an action. It rarely answers the questions that determine impact: Which workload was running? Which identity was used? Which retrieved content shaped the decision? Which tool call was proposed? Was it blocked, approved, sent, retried, or completed? What did the destination system actually do?
Those questions define the difference between conversation logging and AI traffic monitoring. The first captures language. The second connects language-driven decisions to execution. As LLM applications add retrieval, tools, memory, delegation, browser steps, and background workflows, that connection becomes necessary for reliability, security operations, and governance.
Poxek LLM Firewall applies this runtime view in private preview: observe model and agent traffic in the context of stated workload capabilities and policy decisions. The objective is not a universal risk score for prompts. It is evidence that helps a team reconstruct the route from an input to a system effect, while respecting the sensitivity of the data collected.
Start with the unit of work
The most useful observability unit is not always an individual model call. For a simple assistant, it may be enough. For an agent, the useful unit is usually a user request, scheduled job, or workflow run with child operations underneath it.
A trace for a customer-service agent, for example, may contain a request intake, a retrieval operation, a model invocation, a proposed lookup, a policy check, a tool execution, and a final answer. The trace needs a durable correlation identifier plus workload, environment, tenant, and identity context. It should make it possible to ask whether two events belong to the same agent run without using a prompt body as the primary key.
OpenTelemetry's GenAI semantic conventions illustrate the direction of travel. They define attributes for agent and conversation identifiers and recognize operations such as chat, retrieval, execute_tool, invoke_agent, and invoke_workflow. Shared vocabulary does not solve the policy problem, but it gives distributed systems a better chance of correlating their records across languages and services.
Not every operation requires the same level of detail. A model-call span can carry model name, token use, duration, and termination reason. A tool span needs the tool identity, operation, decision context, request status, error type, and downstream result. A policy event needs the rule or policy version, decision, reason category, and approval state. The goal is a legible causal chain, not a maximal archive of raw text.
Record decisions and outcomes as separate events
An agent's natural-language output is a statement of intent, not proof of execution. Even a structured tool call is only a proposed action until an enforcement point admits it and a target system returns a result. Monitoring must preserve the intermediate states.
A minimal state sequence may look like this:
- An agent proposes a tool call with a typed operation and target.
- Runtime policy evaluates the call against the workload capability, identity context, and risk conditions.
- The system records an allow, deny, or approval-required decision with a reason that a reviewer can understand.
- If allowed or approved, the integration invokes the downstream service using an appropriately scoped identity.
- The destination reports completion, rejection, partial completion, or failure; that result is linked back to the policy decision.
This separation is useful for more than security. It identifies a tool schema problem, a downstream permission problem, an intermittent dependency, an unnecessary retry loop, or an approval queue that is too vague to support decisions. It also prevents the misleading conclusion that a model “did” something merely because it said it did.
OWASP's guidance on improper output handling is relevant here. Model output can be influenced by prompt input and must be validated before it reaches backend functions. Logging policy evaluation and tool execution does not replace that validation, but it establishes whether the validation and authorization boundaries were actually crossed during a particular run.
Content capture is a data-governance decision
Full prompt, response, tool-argument, and tool-result capture can be valuable in debugging and incident response. It can also reproduce exactly the data a team is trying to protect: personal information, credentials, customer records, proprietary code, or untrusted content supplied by an attacker. OpenTelemetry's GenAI guidance explicitly warns that input and output message attributes can contain sensitive information; its observability examples also note that content capture is not enabled by default.
That is a sound default for an operational design. A monitoring policy should specify what is collected, why it is collected, who can read it, how long it is retained, and how it is segregated by tenant or environment. It should distinguish low-risk technical metadata from content that requires access control, redaction, truncation, encryption, or explicit opt-in.
There is no single correct level of capture. A production workflow processing sensitive data may retain event metadata and controlled evidence references while a pre-production test environment captures more content for debugging. The important point is that the choice is owned and reviewable. A raw transcript store with broad access is not observability maturity.
Observability supports governance only when someone owns the response
NIST's Generative AI Profile frames risk management across the system lifecycle and includes suggested actions for incident monitoring, after-action review, and maintaining system inventory information. AI telemetry should support those operational responsibilities: named workload owners, an approved risk posture, data-handling rules, an escalation path, and a way to revise a policy after an event.
For security teams, useful queries may include: Which workflows tried to access a prohibited destination? Which tool operations were denied after a policy update? Which agents repeatedly request approvals they cannot use? For platform teams: Which workflow versions increased tool failures or token consumption? For governance teams: Which applications process a sensitive data class, and where is the evidence for the stated control boundary?
These are questions about behavior in context. They cannot be answered well with model latency and token counts alone, although those measurements remain useful. They also cannot be answered solely with a prompt classifier, because a classifier cannot establish which scoped identity reached which destination or whether the target service committed a state change.
Build a reviewable telemetry contract
Before deploying an agent, teams can define a small telemetry contract alongside the capability policy. It should cover the identifiers that join events; expected operations and tools; policy decision states; approval and execution outcomes; protected fields; retention and access rules; and the owner who reviews exceptions. Versioning the contract matters: when a workflow, tool, policy, or model configuration changes, investigators need to know which version produced the event.
The Poxek LLM Firewall direction is to make this runtime context available across model and agent interactions: traffic monitoring connected to capability boundaries, policy outcomes, and execution results. It does not promise complete visibility, perfect content classification, or prevention of every unsafe action. Systems still need sound tool design, least-privilege identities, target-side controls, and human decisions for high-impact work.
The market question is therefore not whether an organization has “LLM logs.” It is whether it can reconstruct an important AI-driven action without guessing. If the answer depends on a transcript and a series of disconnected service logs, the next architectural improvement is to connect the decision path before adding more autonomous capability.