On-Demand Webinar

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

Agentic SecOps
August 3, 2026 11:06 AM
CST
Online
On-Demand Webinar

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

Detection Strategies

Part two of our Agentic SecOps series ended on a line that sounds simple, but isn't: an agent is only as good as the data it can reach. Put an investigation agent on an alert and starve it of the log that would have explained it, and you don't get a wrong answer. You get a confident one built on half the picture. So before any of the four jobs of the SOC can run well, one question has to have an answer: can the work actually see everything it needs to see?

For most SOCs, the honest answer is no. And the strange part is how they got there. Every team we talk to spent the last decade buying more data to see more, but most of them actually see less.

The data paradox

The instinct was right. More telemetry should mean more visibility. Every new cloud account, SaaS app, identity provider, and endpoint produces a stream worth watching, and teams rightfully turned them on. But visibility didn't compound with the data; it fragmented, because each stream landed in a different place, under a different query language, on a different bill. The result is a SOC that owns more security data than ever, but can answer fewer questions across it. The data grew, but the ability to see across it didn't.

If the data is scattered by design, one question can't cross it.

Two things that break "one question across everything"

If you ask an analyst to run a single question across their environment, like "has this indicator touched anything we own in the last thirty days," it will stall in two places. The first is the pivot tax. Security data lives in a dozen consoles: a SIEM, a data lake, cloud storage, an EDR console, each with its own syntax and its own login. One question becomes a dozen queries, asked a dozen ways, stitched back together manually. The analyst spends more time reconciling four dashboards than reasoning about the answer, and the reasoning is the part they actually get paid for.

The second is the cost blind spot. To control what a SIEM charges per gigabyte, teams route the noisier, high-volume sources to cheap cloud storage instead. Those logs are retained, but they can't be searched, so in practice they're invisible. The data dropped to afford the SIEM is the data you now can't see when you need it. And it's rarely the unimportant data. It's the firewall logs, cloud audit trails, and flow data, the exact sources an investigation reaches for. So the SOC ends up with two failure modes at once: the data it can search is scattered across tools that don't talk to each other, and the data it can't afford to centralize is stored somewhere it can't query at all.

The old fix is what created the visibility gap

Part one traced the SIEM cost curve, around $150,000 a year at 100 GB a day, into the millions at terabytes, increasing 30 to 100% year over year as data grows. For a decade the answer to fragmented visibility was to centralize, pull everything into one place and search it there. That worked until the bill made it impossible.

Once centralizing everything got too expensive, teams started dropping sources to control spend, and "centralize everything" stopped being true. The fix became the cause. The model that promised one place to see everything is the same model that forced teams to leave half their data outside it. That's why the answer can't be one more place to centralize data. Another destination to ingest into is another bill, and the same pressure that emptied the last one empties this one, too. The gap doesn't close by moving the data one more time, but by searching the data where it already sits.

What visibility actually takes

Searching data in place, rather than moving it first, changes what's possible. Three things have to be true for it to work:

  • One search across every source. An analyst should ask one question, in one interface, and have it run across every connected platform at once, the SIEM, the data lake, the cloud storage bucket, with each result labeled by where it came from. One question, one result set, and no pivoting between consoles. This is what Anvilogic's Federated Search does, and it turns "which tool do I check first" into a non-question.
  • Reach the data you couldn't search before. The logs stored in cloud storage to control cost stay parked. Anvilogic Compute indexes that data in place, in S3, Azure Blob, or Google Cloud Storage, and makes it searchable without ingesting it into a SIEM or standing up a pipeline, so the blind spot teams accepted as the price of affordability stops being a blind spot. Nothing moves; it just becomes visible.
  • Ask in plain language, under your own credentials. Queries run where the data lives, under each analyst's own access, so nothing is centralized, and existing access controls carry through. And because a Search Agent turns a natural-language question into the optimized query each platform expects, hunting stops being gated on whether an analyst happens to know that platform's syntax. The question is the skill, not the translation.

And none of this requires ripping anything out. A team keeps its SIEM and searches alongside it, or runs standalone on a data lake, or modernizes toward one gradually while the SIEM stays live. The point isn't a new place to put data. It's the ability to see all of it, wherever the team already decided to keep it.

The gap closes by searching the data where it sits, not moving it again.

Why search feeds everything else

Dubbing search as "nice to have" can't be the case anymore. In an Agentic SecOps model, search is the foundation, and it feeds the rest of the lifecycle. Onboarding makes a source usable, search makes it reachable, detection runs on what's reachable, and investigation pulls its context from the same well. An agent that can't reach a source is blind to whatever that source would have shown, and it won't tell you it's blind, so it'll just work with less.

AI agents are only as strong as the data underneath them, so if you widen what the SOC can see, every job on top of it gets better. More data sources can be onboarded, more questions can be answered, more detections can be grounded in real data, and more investigations can close on full context instead of partial.

Every job runs on what search can reach.

What changes for the SOC

Right now teams pay for this problem twice, in what the SIEM charges and in the data they drop to afford it. Fixing it ends both bills. The question that used to take four consoles and hours takes one search. The data that used to be in the dark becomes searchable, and the agents doing the everyday work of the SOC finally run on all of it, not just the affordable slices. But reaching the data is the first move, not the last one.

Once an agent can see everything and act on what it finds, the harder question becomes: can you trust it? That's what we'll cover in part 4 of this blog series, why most AI for the SOC can't be trusted to act, and what it takes to trust an agent that does. Stay tuned.

Get the Latest Resources

Leave Your Data Where You Want: Detect Across Snowflake

Demo Series
Leave Your Data Where You Want: Detect Across Snowflake
Watch

MonteAI: Your Detection Engineering & Threat Hunting Co-Pilot

Demo Series
MonteAI: Your Detection Engineering & Threat Hunting Co-Pilot
Watch
White Paper

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

Agentic SecOps
August 3, 2026

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

Agentic SecOps
No items found.

Part two of our Agentic SecOps series ended on a line that sounds simple, but isn't: an agent is only as good as the data it can reach. Put an investigation agent on an alert and starve it of the log that would have explained it, and you don't get a wrong answer. You get a confident one built on half the picture. So before any of the four jobs of the SOC can run well, one question has to have an answer: can the work actually see everything it needs to see?

For most SOCs, the honest answer is no. And the strange part is how they got there. Every team we talk to spent the last decade buying more data to see more, but most of them actually see less.

The data paradox

The instinct was right. More telemetry should mean more visibility. Every new cloud account, SaaS app, identity provider, and endpoint produces a stream worth watching, and teams rightfully turned them on. But visibility didn't compound with the data; it fragmented, because each stream landed in a different place, under a different query language, on a different bill. The result is a SOC that owns more security data than ever, but can answer fewer questions across it. The data grew, but the ability to see across it didn't.

If the data is scattered by design, one question can't cross it.

Two things that break "one question across everything"

If you ask an analyst to run a single question across their environment, like "has this indicator touched anything we own in the last thirty days," it will stall in two places. The first is the pivot tax. Security data lives in a dozen consoles: a SIEM, a data lake, cloud storage, an EDR console, each with its own syntax and its own login. One question becomes a dozen queries, asked a dozen ways, stitched back together manually. The analyst spends more time reconciling four dashboards than reasoning about the answer, and the reasoning is the part they actually get paid for.

The second is the cost blind spot. To control what a SIEM charges per gigabyte, teams route the noisier, high-volume sources to cheap cloud storage instead. Those logs are retained, but they can't be searched, so in practice they're invisible. The data dropped to afford the SIEM is the data you now can't see when you need it. And it's rarely the unimportant data. It's the firewall logs, cloud audit trails, and flow data, the exact sources an investigation reaches for. So the SOC ends up with two failure modes at once: the data it can search is scattered across tools that don't talk to each other, and the data it can't afford to centralize is stored somewhere it can't query at all.

The old fix is what created the visibility gap

Part one traced the SIEM cost curve, around $150,000 a year at 100 GB a day, into the millions at terabytes, increasing 30 to 100% year over year as data grows. For a decade the answer to fragmented visibility was to centralize, pull everything into one place and search it there. That worked until the bill made it impossible.

Once centralizing everything got too expensive, teams started dropping sources to control spend, and "centralize everything" stopped being true. The fix became the cause. The model that promised one place to see everything is the same model that forced teams to leave half their data outside it. That's why the answer can't be one more place to centralize data. Another destination to ingest into is another bill, and the same pressure that emptied the last one empties this one, too. The gap doesn't close by moving the data one more time, but by searching the data where it already sits.

What visibility actually takes

Searching data in place, rather than moving it first, changes what's possible. Three things have to be true for it to work:

  • One search across every source. An analyst should ask one question, in one interface, and have it run across every connected platform at once, the SIEM, the data lake, the cloud storage bucket, with each result labeled by where it came from. One question, one result set, and no pivoting between consoles. This is what Anvilogic's Federated Search does, and it turns "which tool do I check first" into a non-question.
  • Reach the data you couldn't search before. The logs stored in cloud storage to control cost stay parked. Anvilogic Compute indexes that data in place, in S3, Azure Blob, or Google Cloud Storage, and makes it searchable without ingesting it into a SIEM or standing up a pipeline, so the blind spot teams accepted as the price of affordability stops being a blind spot. Nothing moves; it just becomes visible.
  • Ask in plain language, under your own credentials. Queries run where the data lives, under each analyst's own access, so nothing is centralized, and existing access controls carry through. And because a Search Agent turns a natural-language question into the optimized query each platform expects, hunting stops being gated on whether an analyst happens to know that platform's syntax. The question is the skill, not the translation.

And none of this requires ripping anything out. A team keeps its SIEM and searches alongside it, or runs standalone on a data lake, or modernizes toward one gradually while the SIEM stays live. The point isn't a new place to put data. It's the ability to see all of it, wherever the team already decided to keep it.

The gap closes by searching the data where it sits, not moving it again.

Why search feeds everything else

Dubbing search as "nice to have" can't be the case anymore. In an Agentic SecOps model, search is the foundation, and it feeds the rest of the lifecycle. Onboarding makes a source usable, search makes it reachable, detection runs on what's reachable, and investigation pulls its context from the same well. An agent that can't reach a source is blind to whatever that source would have shown, and it won't tell you it's blind, so it'll just work with less.

AI agents are only as strong as the data underneath them, so if you widen what the SOC can see, every job on top of it gets better. More data sources can be onboarded, more questions can be answered, more detections can be grounded in real data, and more investigations can close on full context instead of partial.

Every job runs on what search can reach.

What changes for the SOC

Right now teams pay for this problem twice, in what the SIEM charges and in the data they drop to afford it. Fixing it ends both bills. The question that used to take four consoles and hours takes one search. The data that used to be in the dark becomes searchable, and the agents doing the everyday work of the SOC finally run on all of it, not just the affordable slices. But reaching the data is the first move, not the last one.

Once an agent can see everything and act on what it finds, the harder question becomes: can you trust it? That's what we'll cover in part 4 of this blog series, why most AI for the SOC can't be trusted to act, and what it takes to trust an agent that does. Stay tuned.

Resources

Blog

Agentic SecOps Part 1: The SOC Math Problem Hiring Can't Solve

Security data compounds 35 to 50% a year while SOC headcount stays flat. Part 1 of our Agentic SecOps series on why hiring can't close the gap.
Blog

Agentic SecOps Part 2: What It Actually Is, and Why Most "AI SOC" Tools Aren't It

Agentic SecOps is a model where AI agents do the SOC's everyday work, onboarding, search, detection, and investigation, with a human approving what matters.

See what Anvilogic can do for your SOC.

Talk to a practitioner who has been on your side of the problem.

August 3, 2026

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

Agentic SecOps

Resources

Blog

Agentic SecOps Part 1: The SOC Math Problem Hiring Can't Solve

Security data compounds 35 to 50% a year while SOC headcount stays flat. Part 1 of our Agentic SecOps series on why hiring can't close the gap.
Blog

Agentic SecOps Part 2: What It Actually Is, and Why Most "AI SOC" Tools Aren't It

Agentic SecOps is a model where AI agents do the SOC's everyday work, onboarding, search, detection, and investigation, with a human approving what matters.

See what Anvilogic can do for your SOC.

Talk to a practitioner who has been on your side of the problem.

Product Vision
|
August 3, 2026
|
4 min read

Agentic SecOps Part 3: How Buying More Data Made the SOC See Less

This is some text inside of a div block.

| Author

More data was supposed to mean more visibility. It fragmented instead. Why centralizing backfired, and what it takes to search across everything a team already has. Part 3 of the series.

Part two of our Agentic SecOps series ended on a line that sounds simple, but isn't: an agent is only as good as the data it can reach. Put an investigation agent on an alert and starve it of the log that would have explained it, and you don't get a wrong answer. You get a confident one built on half the picture. So before any of the four jobs of the SOC can run well, one question has to have an answer: can the work actually see everything it needs to see?

For most SOCs, the honest answer is no. And the strange part is how they got there. Every team we talk to spent the last decade buying more data to see more, but most of them actually see less.

The data paradox

The instinct was right. More telemetry should mean more visibility. Every new cloud account, SaaS app, identity provider, and endpoint produces a stream worth watching, and teams rightfully turned them on. But visibility didn't compound with the data; it fragmented, because each stream landed in a different place, under a different query language, on a different bill. The result is a SOC that owns more security data than ever, but can answer fewer questions across it. The data grew, but the ability to see across it didn't.

If the data is scattered by design, one question can't cross it.

Two things that break "one question across everything"

If you ask an analyst to run a single question across their environment, like "has this indicator touched anything we own in the last thirty days," it will stall in two places. The first is the pivot tax. Security data lives in a dozen consoles: a SIEM, a data lake, cloud storage, an EDR console, each with its own syntax and its own login. One question becomes a dozen queries, asked a dozen ways, stitched back together manually. The analyst spends more time reconciling four dashboards than reasoning about the answer, and the reasoning is the part they actually get paid for.

The second is the cost blind spot. To control what a SIEM charges per gigabyte, teams route the noisier, high-volume sources to cheap cloud storage instead. Those logs are retained, but they can't be searched, so in practice they're invisible. The data dropped to afford the SIEM is the data you now can't see when you need it. And it's rarely the unimportant data. It's the firewall logs, cloud audit trails, and flow data, the exact sources an investigation reaches for. So the SOC ends up with two failure modes at once: the data it can search is scattered across tools that don't talk to each other, and the data it can't afford to centralize is stored somewhere it can't query at all.

The old fix is what created the visibility gap

Part one traced the SIEM cost curve, around $150,000 a year at 100 GB a day, into the millions at terabytes, increasing 30 to 100% year over year as data grows. For a decade the answer to fragmented visibility was to centralize, pull everything into one place and search it there. That worked until the bill made it impossible.

Once centralizing everything got too expensive, teams started dropping sources to control spend, and "centralize everything" stopped being true. The fix became the cause. The model that promised one place to see everything is the same model that forced teams to leave half their data outside it. That's why the answer can't be one more place to centralize data. Another destination to ingest into is another bill, and the same pressure that emptied the last one empties this one, too. The gap doesn't close by moving the data one more time, but by searching the data where it already sits.

What visibility actually takes

Searching data in place, rather than moving it first, changes what's possible. Three things have to be true for it to work:

  • One search across every source. An analyst should ask one question, in one interface, and have it run across every connected platform at once, the SIEM, the data lake, the cloud storage bucket, with each result labeled by where it came from. One question, one result set, and no pivoting between consoles. This is what Anvilogic's Federated Search does, and it turns "which tool do I check first" into a non-question.
  • Reach the data you couldn't search before. The logs stored in cloud storage to control cost stay parked. Anvilogic Compute indexes that data in place, in S3, Azure Blob, or Google Cloud Storage, and makes it searchable without ingesting it into a SIEM or standing up a pipeline, so the blind spot teams accepted as the price of affordability stops being a blind spot. Nothing moves; it just becomes visible.
  • Ask in plain language, under your own credentials. Queries run where the data lives, under each analyst's own access, so nothing is centralized, and existing access controls carry through. And because a Search Agent turns a natural-language question into the optimized query each platform expects, hunting stops being gated on whether an analyst happens to know that platform's syntax. The question is the skill, not the translation.

And none of this requires ripping anything out. A team keeps its SIEM and searches alongside it, or runs standalone on a data lake, or modernizes toward one gradually while the SIEM stays live. The point isn't a new place to put data. It's the ability to see all of it, wherever the team already decided to keep it.

The gap closes by searching the data where it sits, not moving it again.

Why search feeds everything else

Dubbing search as "nice to have" can't be the case anymore. In an Agentic SecOps model, search is the foundation, and it feeds the rest of the lifecycle. Onboarding makes a source usable, search makes it reachable, detection runs on what's reachable, and investigation pulls its context from the same well. An agent that can't reach a source is blind to whatever that source would have shown, and it won't tell you it's blind, so it'll just work with less.

AI agents are only as strong as the data underneath them, so if you widen what the SOC can see, every job on top of it gets better. More data sources can be onboarded, more questions can be answered, more detections can be grounded in real data, and more investigations can close on full context instead of partial.

Every job runs on what search can reach.

What changes for the SOC

Right now teams pay for this problem twice, in what the SIEM charges and in the data they drop to afford it. Fixing it ends both bills. The question that used to take four consoles and hours takes one search. The data that used to be in the dark becomes searchable, and the agents doing the everyday work of the SOC finally run on all of it, not just the affordable slices. But reaching the data is the first move, not the last one.

Once an agent can see everything and act on what it finds, the harder question becomes: can you trust it? That's what we'll cover in part 4 of this blog series, why most AI for the SOC can't be trusted to act, and what it takes to trust an agent that does. Stay tuned.

Resources

Blog

Agentic SecOps Part 1: The SOC Math Problem Hiring Can't Solve

Security data compounds 35 to 50% a year while SOC headcount stays flat. Part 1 of our Agentic SecOps series on why hiring can't close the gap.
Blog

Agentic SecOps Part 2: What It Actually Is, and Why Most "AI SOC" Tools Aren't It

Agentic SecOps is a model where AI agents do the SOC's everyday work, onboarding, search, detection, and investigation, with a human approving what matters.