The Cybersecurity risk you couldn’t see

insight
August 09, 2026
9 min read

Author

Thomas-Aardal

 

Thomas Aardal is Chief Technology Officer (CTO) at Nagarro. He is an architect and consultant for cross-functional technology practices, focusing on cross-cutting technologies such as enterprise integration, open-source platforms, and process automation.

 

Frontier AI didn’t make your systems more vulnerable. It made your assumptions more visible.

Over the past several months, a familiar pattern has emerged across boardrooms and security teams alike. Large, capable, and increasingly autonomous frontier AI models are demonstrating a remarkable ability to identify vulnerabilities in software systems at a scale and speed that previously required entire red teams and months of coordinated effort. The reaction has been a mixture of concern and genuine uncertainty: Should organizations be running more testing? Do existing threat models need to be reconsidered entirely?

The answer to both is yes. But something more fundamental is at stake here, because this moment exposes a structural flaw in how the industry has understood cybersecurity risk for a long time, one that has been quietly tolerated rather than honestly examined.

Visibility was never the same as safety

The core issue, stated plainly: cybersecurity has long treated the visibility of vulnerabilities as a proxy for the existence of risk. When a system is tested and ten vulnerabilities are found, they get scored, prioritized, partially remediated, and the remainder accepted as residual risk. The audit closes. The paperwork is complete.

What frontier AI is doing, with uncomfortable clarity, is demonstrating an idea of how much was always present beneath the surface.

1

The vulnerabilities these models are now surfacing did not appear the moment a model went looking. They were present yesterday. They were present when the last compliance audit passed. They existed in the same codebase, the same infrastructure, the same product customers are using today. The only thing that has changed is the speed and volume at which they can be found.

2

This distinction matters and is often lost in the current conversation: systems are not more vulnerable because of AI. They are, however, more susceptible to attack. Those are not the same thing, and conflating them leads to the wrong response.

3

A system’s vulnerability is a property of the system itself; it exists independently of whether anyone has observed it. Susceptibility to attack is a function of how accessible those vulnerabilities are to adversaries. What frontier AI has done is dramatically shift that accessibility, and in doing so, it has exposed a premise that has always had limits: that what had not yet been found was not yet a meaningful problem.

The structural problem with the current model

The model most organizations follow today runs as follows: test the application to discover vulnerabilities, score and prioritize what is found, fix or mitigate the highest-risk items, and declare acceptable risk on what remains. Compliance frameworks were largely built around variations of this loop. Pass the audit, demonstrate due diligence, maintain documentation, repeat on the next cycle.

The problem is not that this model reflects careless intent.

Many serious, skilled professionals have built their practice around executing it well. The problem is that it is structurally non-exhaustive. No practical testing effort will ever discover all vulnerabilities in a complex system, not because of any failure of the practitioners, but because vulnerability discovery is an open-ended problem. The more you look, the more you find, and the decision to stop looking is invariably driven by cost and schedule rather than by any objective measure of completeness.


There is a revealing pattern in how vulnerability discovery actually unfolds over the course of a testing engagement.

Historically, findings tend to follow a Bell curve: as testing effort accumulates, the rate of new discoveries rises, peaks, and then diminishes. On a cost-and-schedule basis, that diminishing return can appear to be a rational signal to stop. In most disciplines, it is. Security is not most disciplines. The vulnerabilities found on the tail of that curve — after the peak, when effort yields fewer results — are not necessarily minor. They can be equally critical, or more so, than those uncovered in the earlier, more productive phase, simply because they are harder to surface. Difficulty of discovery does not correlate with severity. This is precisely why the principle that the more you look, the more you find retains its force in security contexts in a way it does not elsewhere. A model that treats the tailing-off of findings as a signal of completeness is measuring the limits of the testing effort, not the limits of actual exposure.


When the benchmark for “secure enough” is calibrated against what a bounded testing effort found, the result is not a measurement of system security.
It is a measurement of one testing effort’s output at one point in time. If a frontier AI model can surface in hours vulnerabilities that a prior audit missed entirely, the honest conclusion is that the audit measured testing scope, not system security. The risk did not change. The visibility did.

 

This is not an indictment of the teams doing the work. It is a structural observation about the framework within which their work is measured.

A different starting point

There is a compelling case for inverting the model.

Instead of beginning with testing to discover what might be wrong, the starting point could be a predictive, statistical analysis that estimates the vulnerability profile of a system before a single test is run. The characteristics of the system itself (its architecture, dependencies, technology stack, attack surface, and age) can generate a probabilistic risk profile that reflects a more complete picture of its security posture, not merely the results of the last audit cycle.

The data to support this already exists. Systems built on certain frameworks carry statistically consistent vulnerability distributions. Systems with high dependency counts, legacy components, or particular architectural patterns have measurable risk profiles. Years of CVE data, breach records, and incident histories form a credible foundation for genuinely predictive models.

Fog-is-lifting

In this inverted model, statistical analysis of a given system produces a risk score that becomes the primary input for cybersecurity decision-making. Testing then becomes a targeted instrument, used not to discover risk from scratch but to identify and confirm what specifically needs to be addressed within an already-estimated risk profile. Fix, mitigate, reassess. Repeat.

This is a foundational shift. The advantage, however, is significant: organizations are no longer relying solely on what was found during the last audit cycle. They are operating with a continuously updated, statistically grounded view of actual exposure, one that remains valid even as more capable testing tools change what can be discovered.

This shift in primacy does not mean abandoning testing. Organizations should continue conducting security testing as a matter of ongoing practice — including, where appropriate, deploying frontier AI models as part of that testing effort. The change is one of role, not of volume. Testing remains a valuable instrument. What changes is its position in the decision-making hierarchy. In the inverted model, testing no longer sets the baseline from which risk is measured; that baseline is established by the predictive framework. Testing results are interpreted within it, used to sharpen, confirm, or update the probabilistic picture rather than to generate the primary risk assessment from scratch. The two operate in parallel — testing as an active investigative tool, the predictive model as the persistent framework that contextualizes what testing finds.

What this means for compliance

Current compliance and regulatory frameworks would face real friction with this proposal in its present form. Compliance, by design, operates on verifiability: documented controls, evidence of testing, records of remediation. A statistical risk score, however well-constructed, does not map cleanly onto that paradigm today.

Fog-is-lifting-2 (1)-1

 

That tension is worth holding rather than resolving too quickly in compliance’s favor. The assumption that “compliant equals secure” already has well-documented limits; organizations that maintained full compliance and still experienced significant breaches understand this directly. Compliance frameworks are, at their best, a floor (a minimum standard), not a complete measure of resilience.

If the genuine goal is to build systems that are more resilient, rather than systems that are more thoroughly documented as secure, waiting for regulatory frameworks to lead the way is not a viable posture. Regulation tends to follow practice. The organizations willing to develop more rigorous approaches now will be better positioned as the regulatory environment eventually catches up.

The leadership dimension

This question belongs at the leadership level because the decisions that sustain the current model are not primarily technical. They are business decisions.

The acceptable risk threshold has always been set, implicitly, around what was visible. “Testing was conducted, critical findings were addressed, residual risk was accepted.” That approach made reasonable sense when visibility was the binding constraint. It holds up less well in an environment where frontier AI models are systematically expanding what is visible to anyone, including adversaries operating with the same tools.

The question for leaders is what risk acceptance actually means in that context. Whether a passed audit is sufficient, or whether it is worth asking what the audit did not find. Whether investment in predictive approaches is warranted, giving security teams a more complete picture between audit cycles and before vulnerabilities become incidents.

The fog lifting is not, in itself, the problem. The landscape it reveals was always there. The more constructive response is to ask whether the current framework is equipped to deal honestly with that landscape, and if not, what a more honest framework would look like.

That question deserves serious attention now, before circumstances force the answer.

Nagarro’s cybersecurity practice works with organizations navigating the strategic and operational dimensions of security in complex, evolving technology environments.

Frontier AI and cybersecurity risk: What leaders need to know

Get in touch

Have you been measuring Frontier AI wrong?