Poxek Research · ai-soc

An evidence-aware handoff is the unit of work for an AI SOC

A product approach for using bounded automation to prepare reviewable L3/L4 SOC cases.

A case should make review faster, not create a second investigation

An escalation from L1 or L2 is often treated as a container for notes. That is too weak for an AI-assisted workflow. A useful handoff is a reviewable unit of work: it tells the next analyst what was observed, which questions were checked, what evidence supports the current assessment, and which uncertainties remain. If the reviewer must search again for the original alert, reconstruct the timeline, or guess why an enrichment was included, the handoff has added prose without reducing investigation work.

Poxek AI SOC operates in Private Preview around that narrower unit of work. Bounded ML, LLM, and agent workflows can support repetitive context assembly and triage, but L3/L4 analysts retain the authority to assess impact, direct response, and close or escalate a case. The product is not a system that makes security decisions in isolation. It prepares evidence-aware cases that make human decisions more informed and auditable.

Define a case contract before choosing a model

The first design choice is the case contract, not the model or prompt. A contract should specify the output that every automated workflow must produce and the evidence it must preserve. At a minimum, it should contain:

  • the alert or initiating event, including source and time;
  • entities observed in the case, such as account, host, workload, IP address, or application;
  • a chronological timeline that links back to source records;
  • enrichment results with their origin and collection time;
  • checks that ran, their status, and any tool or data-access failures;
  • observations, hypotheses, and recommended follow-up separated from each other; and
  • the escalation reason, proposed owner, and decision that still requires human review.

These fields are not bureaucracy. They give the automation a bounded job and make gaps inspectable. An LLM may help express the relationship between events, but it should not be the sole record of the relationship. When a case says a host contacted an unusual destination after a suspicious process started, the reviewer needs links or identifiers for both observations and their timestamps. A generated sentence is a convenience layer over evidence, never the evidence itself.

NIST SP 800-61 Rev. 3 positions incident response within cybersecurity risk management. That supports treating a case as an operational record with a clear path to detection, analysis, response, recovery, and improvement. The useful question is therefore not “Did the model identify the incident?” It is “Did the workflow produce enough traceable material for the accountable analyst to make the next decision?”

Use bounded tools for bounded questions

An AI SOC workflow should call approved tools for specific purposes, rather than receive unrestricted access to every data source. For example, the workflow may be allowed to retrieve the initiating detection, resolve an account to an identity record, find recent activity for the same entity, check an asset's owner and criticality, and retrieve related endpoint events for a defined time window. Each operation has a request, response, timestamp, and failure mode that can be retained with the case.

This pattern has two benefits. First, it limits what an automated process can access and do. Second, it gives the analyst a compact audit trail. A reviewer can see that a host lookup returned no owner, or that a cloud query used a fixed window rather than an open-ended search. Those facts are more useful than a summary that quietly treats absent data as benign.

NIST SP 800-92 describes log management as a set of organizational processes, including the infrastructure and practices needed to make log data useful. That is why data-source quality belongs in the case contract. A workflow should record when source telemetry is unavailable, delayed, incomplete, or inconsistent. It should not use natural-language confidence to disguise a collection problem.

The same rule applies to identity and asset context. An ownership field may be stale. An IP address may be shared. A device name may be recycled. Automation can collect these associations and flag contradictions, but it should not turn an ambiguous match into an asserted fact merely because the wording reads smoothly.

Keep analytic hypotheses testable

MITRE ATT&CK is valuable because it offers a shared knowledge base of tactics and techniques grounded in real-world adversary observations. In a case workflow, it can structure a question: whether the available evidence is consistent with a technique, what adjacent telemetry may be relevant, and which behaviors deserve a follow-up check. It should not be used to convert a detection label into a definitive attribution or incident classification.

A testable case statement has a different shape from a model conclusion. Instead of “this is lateral movement,” it might say: “Remote-service activity from the source host to two peer systems was observed after the initiating event; no interactive user context was available in the collected records. Review authentication and process telemetry for those connections.” The statement identifies the observation, its limitation, and the next check. It leaves room for the reviewer to refute it.

This is also where structured output is preferable to a single narrative. A case can contain an evidence table, a separate hypotheses section, and an explicit list of unanswered questions. The system may write a concise summary for speed, but the summary should be derived from those structured components and should link back to them.

Set human decision gates where consequences begin

Some actions have a different risk profile from evidence collection. Disabling an account, isolating a host, blocking a domain, notifying a business owner, or declaring an incident can affect operations, legal obligations, and recovery. These should be human decision gates in a Private Preview workflow unless a customer has separately defined, authorized, and tested a narrowly scoped response path. A confidence score, a threat label, or a generated recommendation is not authorization.

NIST CSF 2.0 is useful here because it makes governance a first-class function. SOC automation must fit named decision rights, risk appetite, and recovery requirements. The case should say what it recommends and why, but the accountable role should decide whether to act and record that decision.

For Poxek AI SOC, this boundary makes the product easier to evaluate honestly. The system can assist repetitive L1/L2 gathering, normalization, and handoff preparation. It does not claim to replace experienced analysts or incident leadership. Its usefulness depends on the customer's telemetry, approved data paths, operating procedures, and human review.

Evaluate the workflow as a case-making system

Evaluation should use representative, permissioned historical or synthetic cases and sampled analyst review. Teams can inspect whether the workflow included the correct initiating evidence, preserved source links, separated facts from hypotheses, surfaced missing data, and routed the case according to the documented policy. They can also look for unsafe patterns: invented evidence, overconfident conclusions, incorrect entity joins, or recommendations that bypass a decision gate.

This evaluation is more informative than asking whether a summary sounds credible. A credible-sounding paragraph can still omit the event that disproves it. An evidence-aware handoff gives reviewers a way to detect that failure and improve the workflow, playbook, or data source.

The intended outcome is modest and operational: fewer repetitive steps between an alert and a reviewable case, while the analyst retains access to the facts needed to challenge the case. That is a better foundation for AI SOC work than a promise that the SOC can run without the people accountable for security decisions.

Sources