Poxek Research · ai-soc

The future of SOC automation is a review system, not an autonomous authority

A future-state view of AI-assisted investigations with constrained tools, recorded evidence, and human authority.

The useful future is less theatrical than autonomy

The future of security operations will include more automation, but the important design question is not whether a model can produce an answer. It is whether a SOC can review, reproduce, and govern the work behind that answer. Investigations cross systems with different levels of reliability, involve decisions with unequal consequences, and often continue while the available evidence changes. An operating model that hides those realities behind an autonomous-agent label creates an attractive demo and a difficult incident process.

A more useful future state is a review system. It uses automation to collect and organize routine context, checks defined conditions with constrained tools, creates an evidence trail, and presents decisions to the people authorized to make them. L3/L4 analysts, incident commanders, and service owners remain part of the system because they hold context that telemetry and models do not: operational priorities, safety constraints, regulatory obligations, and the authority to accept or change risk.

Poxek AI SOC assists repetitive L1/L2 triage and context preparation in Private Preview, then produces evidence-aware handoffs for L3/L4 review. It does not claim autonomous incident response, universal detection coverage, or replacement of senior analyst judgment.

Treat AI risk as an operational design input

NIST's AI Risk Management Framework describes a voluntary approach for incorporating trustworthiness considerations into the design, development, use, and evaluation of AI systems. Its functions — Govern, Map, Measure, and Manage — give a practical way to think about an AI SOC without assuming that the model is the system boundary.

Govern means assigning responsibility for the workflow: who owns the playbook, approves data access, reviews failures, and decides which actions can ever be automated. Map means documenting the intended use, affected users, data sources, potential harms, and operating conditions. Measure means testing the workflow for accuracy and failure modes that matter to the SOC. Manage means deciding whether to deploy, constrain, change, or withdraw a capability as evidence from those tests and operations accumulates.

For a SOC workflow, this turns broad AI principles into concrete questions. Can it distinguish observed facts from derived text? Does it preserve the source for each material assertion? Does it accidentally merge two identities? Can it be induced to request information outside the case? Does it surface a data-quality problem rather than write around it? These questions matter before the system is allowed to influence an analyst's priority queue.

NIST's Generative AI Profile adds considerations specific to generative systems. In an investigation setting, the relevant lesson is not that every output is untrustworthy; it is that teams must manage risks from generated content in the context of the intended use. A polished narrative is not a verification mechanism. The workflow needs independent evidence links, constrained retrieval, and human review appropriate to the decision at stake.

Build from observable components with clear limits

The next generation of AI SOC tooling should be composed from small, observable components. One component might normalize identifiers. Another might retrieve events for a fixed entity and time window. A third might compare collected facts against a documented investigation checklist. A fourth might format the resulting material as a handoff. Each should have defined inputs, allowed data paths, output schema, and failure behavior.

This is not merely an engineering preference. It changes how a security team can investigate the automation itself. If a case contains an incorrect association, reviewers can trace whether the issue arose during identity resolution, data retrieval, timeline assembly, or language generation. If the whole process is an unrestricted agent with broad access, the team has fewer practical ways to reproduce the path that produced the error.

Bounded components also support proportionate authority. A workflow that retrieves evidence is different from one that changes production state. A workflow that drafts a recommendation is different from one that sends an external notification. The safest deployment path adds automation where the consequences of error are understood and reversible, then keeps higher-impact actions behind explicit human approval.

Preserve the distinction between behavior and interpretation

MITRE ATT&CK is a useful knowledge base for structuring interpretation, because it captures adversary tactics and techniques from real-world observations. It can help a system propose investigation questions or explain why a sequence of events merits attention. It should not turn a sparse event sequence into a fact about intent, attribution, or impact.

The future AI SOC should record this distinction directly in the case. Behavior is what happened in the available telemetry: a process executed, a connection occurred, a credential was used, a configuration changed. Interpretation is what those facts may mean, the confidence and limitations of that assessment, and the evidence that would change it. A senior analyst should be able to accept, reject, or refine the interpretation without discarding the collected facts.

This matters particularly when telemetry is incomplete. An absent log may mean no activity occurred, or it may mean the source was not collected, retained, or queried correctly. The case should expose that uncertainty. A system that writes “no evidence of malicious activity” when it only means “no result returned” makes an analytic and operational error.

Connect automation to incident response and learning

NIST SP 800-61 Rev. 3 places incident response within the broader activities of cybersecurity risk management and emphasizes continuous improvement. AI-assisted cases should participate in that same loop. Analysts should be able to mark an association wrong, an enrichment unhelpful, an escalation late, or an evidence gap material. Those corrections should improve the relevant playbook, data source, evaluation set, or policy — not become invisible comments under a generated summary.

That feedback loop also keeps evaluation grounded. Rather than relying on generic demonstrations, teams can use permitted historical cases, synthetic exercises, and carefully reviewed live samples. They can test whether the workflow preserved evidence, used only approved tools, recognized missing inputs, and routed decisions to the correct role. When it fails, the team can narrow the capability, update the contract, or remove it from a workflow. An AI SOC should be allowed to become more constrained when evidence justifies caution.

Human supervision is a capability, not a disclaimer

Human supervision is sometimes described as a final approval button. That is insufficient. Effective supervision requires enough context to make a meaningful decision, time to challenge a recommendation, the ability to see source evidence, and clear ownership of the consequence. If the reviewer receives only a confidence score and a paragraph, approval becomes a rubber stamp.

In a mature AI SOC, the handoff itself supports supervision. It shows what triggered the case, what the automation did, what it found, what it could not find, and why the case needs a decision. It keeps response authority with the people and processes that hold it. That design can make L1/L2 work less repetitive without pretending that escalation, containment, recovery, or risk acceptance are language-model tasks.

The future worth building is therefore not an autonomous authority sitting above the SOC. It is a well-instrumented system of assistance inside the SOC: constrained enough to inspect, useful enough to reduce repetitive reconstruction, and accountable enough to support the people who must make difficult decisions under uncertainty.

Sources