When Every Security Test Passes and the Breach Still Happens

A security team runs its usual checks. The phishing simulation gets flagged. The endpoint detection tool catches the test payload. The SIEM rule fires exactly as designed. Every box is green. And yet, months later, that same organization discovers attackers walking out the door with sensitive data. How is that possible if every individual control did its job?

Security analyst reviewing an attack chain validation dashboard and network incident timeline

The answer is uncomfortable but simple: attackers rarely defeat one control. They walk through the seams between several controls that were each tested — and passed — on their own.

Why testing one thing at a time misses the point

Most security validation programs are built around a library of discrete techniques, often mapped to a framework like MITRE ATT&CK — a catalog of observed adversary behaviors, not an exhaustive list of everything an attacker could do. A team runs technique 1234, checks whether it’s detected. Runs technique 5678, checks whether it’s blocked. Score it, move to the next one.

That kind of testing tells you something real. It just doesn’t tell you the thing that actually matters: whether an adversary who links ten of those techniques together — adapting at each step based on what worked — could walk straight through the environment while every individual control reports "no issue detected".

This is the essence of what one security-research analysis calls the "validation gap": the distance between what you assume your controls do and what you can actually demonstrate they do against the way real intrusions unfold. It’s a gap of both timing (environments change faster than quarterly or annual tests can track) and fidelity — confusing control presence, like "MFA is enabled," with control efficacy, like "MFA actually stops account takeover in this specific workflow".

In plain terms: in cybersecurity, as in any chain, strength is measured by the weakest link, not the average one. Fixing 90% of your exposures does not help if the missing 10% is exactly where an attacker’s path happens to run.

What a real intrusion actually looks like

Consider the 2025 breach of the French tax authority, the DGFiP, which affected roughly 678,000 individuals and professionals. Investigators found that attackers used compromised credentials to access internal systems over several weeks before the intrusion was detected, and by the time investigators fully understood the scope, sensitive tax and property data had already been extracted. None of the individual steps in that sequence was exotic. What mattered was that they connected — access, credential use, movement through the environment, and extraction — into one uninterrupted path.

That pattern shows up again and again: initial access leads to some form of credential exposure, which leads to a foothold, which enables movement to another system, which enables staging and removal of data. Each step, tested alone, might look survivable. Strung together, they add up to a breach.

flowchart TD
 A[Initial access: phishing or stolen credential] --> B[Foothold established]
 B --> C[Privilege escalation]
 C --> D[Lateral movement]
 D --> E[Data staged]
 E --> F[Exfiltration]

The point of attack-chain validation isn’t to catalog every possible sequence — that’s practically infinite. It’s to check, deliberately and repeatedly, whether a realistic path like this one can be interrupted at any single point, and to identify which point that is.

From point-in-time checks to continuous validation

This is also why the idea connects to Continuous Threat Exposure Management, or CTEM: a philosophy that treats exposure management as an ongoing cycle of discovery and prioritization rather than a once-a-year audit. A survey of 550 security decision-makers found that 93% agreed delaying improvements to exposure management increases the likelihood of serious incidents, yet only 38% had implemented continuous, automated validation of their defenses, and teams reported spending an average of 42% of their time chasing risks that turned out to be low priority. That gap between knowing about exposure and being able to act on it quickly is exactly where chain-based thinking becomes useful — because it forces the question of exploitability and interruption, not just detection.

This is also where red teaming fits in. A red team exercise, in which specialists simulate a real adversary rather than testing one technique in isolation, has long been the closest approximation of "realistic" testing available — but it’s typically periodic and resource-intensive, more like a snapshot than a continuous signal. The newer generation of automated chain-testing tools tries to bring some of that realism to a faster, more repeatable cadence, though how far these tools have actually spread across the industry, and whether they measurably outperform technique-based testing over time, is not yet independently established. Vendor accounts of "autonomous" attack simulation should be read as product positioning until confirmed by outside evidence.

Comparing the two approaches side by side

Dimension Single-technique validation Attack-chain validation
What it tests Whether one control catches or blocks one specific action Whether a connected sequence of actions can reach its objective
What it misses Interactions between controls, teams, and tools Nothing is guaranteed to be caught — no method eliminates that risk
Typical output A pass/fail score per technique A mapped path showing where the sequence could be stopped, and where it wasn’t
Realism Useful but narrow; controls tested in a vacuum Closer to how intrusions actually unfold, though still a simulation, not proof
Best use Baseline coverage checks, tool tuning Understanding chokepoints and prioritizing fixes with the most leverage

Neither column is "wrong." Technique-level testing remains useful for tuning individual tools and catching regressions. What changes with chain validation is the question being asked — not "does this control work?" but "does the whole path survive?"

What this means for boards and budgets

For a security leader in front of a board, this reframing matters because it changes what "we’re covered" is allowed to mean. A stack of green checkmarks from isolated tests can create false confidence — the appearance of assurance without the evidence that a realistic attacker path would actually be stopped. That doesn’t mean every organization needs full autonomous offensive tooling; for many, the more urgent step is simply asking whether their existing tests connect to each other at all, or whether tools, teams, and alerts operate in separate silos that a real attacker would simply walk between.

It’s also worth resisting a tempting but unproven shortcut: pointing to AI as the reason all of this suddenly matters. AI-assisted tools may be changing how quickly some attackers move once inside a network, but the deeper issue — testing techniques instead of paths — predates any of that, and isolated human-driven intrusions still account for the overwhelming majority of real-world breaches.

The takeaway

The organizations that get breached despite having "passed" their security testing usually didn’t fail a control test — they never ran the chain test in the first place. A passing grade on ten individual checks says nothing about whether those ten things, connected in the right order, still add up to a door left open. The more useful question isn’t whether a tool works. It’s whether, somewhere along the realistic path from a first click to a stolen file, something actually stops it.

Sources

  1. What Is MITRE ATT&CK Framework?
  2. Attack Chains, Not Just Attack Surfaces: Why Testing Individual Techniques Misses the Point
Scroll to Top