The AI measurement gap: are we measuring what actually matters?

insight
August 09, 2026
9 min read

Author

Thomas-Aardal

 

Thomas Aardal is Chief Technology Officer (CTO) at Nagarro. He is an architect and consultant for cross-functional technology practices, focusing on cross-cutting technologies such as enterprise integration, open-source platforms, and process automation.

 

Frontier AI didn’t make your systems more vulnerable
– but it did make your assumptions more visible.

Over the past several months, a familiar pattern has emerged across boardrooms and security teams alike. Frontier AI models, large, capable, increasingly autonomous, are demonstrating a remarkable ability to identify vulnerabilities in software frameworks and products at a scale and speed that previously required entire red teams and months of coordinated effort. The reaction has been a mixture of concern and genuine uncertainty. Should enterprises be running more testing? Do existing threat models need to be reconsidered entirely?

The answer to both is yes. But there is something more fundamental worth addressing here, because this moment exposes a structural flaw in how the industry has thought about cybersecurity risk for a long time, one that has been quietly tolerated rather than honestly examined.

Visibility was never the same as safety

If one were to state the core issue plainly, then cybersecurity has long treated the visibility of vulnerabilities as a proxy for the existence of risk. When a system is tested, and a number of vulnerabilities are found, then they all get scored, prioritized, partially remediated, and the remainder accepted as residual risk. The audit closes, and the paperwork is taken to be in order.

What frontier AI is doing - with uncomfortable clarity - is demonstrating just how much may have been present beneath the surface. And they too will not find “all”.

1

The vulnerabilities these models are now surfacing did not appear the moment a model looked for them. They were present earlier, even when the last compliance audit was declared “passed”. They existed in the same codebase, the same infrastructure, and the same product that customers are using today.

2

What really has changed, is the speed, volume and cost at which they can be found. This distinction matters and often gets lost in the current conversation: systems are not more vulnerable because of AI. They are, however, more susceptible to attack. Those are not the same things, and conflating them leads to the wrong response.

3

A system’s vulnerability is characteristic of the system itself, and exists independently of whether anyone has observed it. How accessible these vulnerabilities are to threat actors has a direct bearing on attack susceptibility.

4

What frontier AI has done is shift that accessibility significantly, and in doing so, it has revealed that the baseline standard for “secure enough” was built on a premise that has always had limitations, i.e., what had not been found was not yet a meaningful problem.

The structural problem with the current model

This is not an indictment of the teams doing the work. It is a structural observation about the framework within which that work is measured.

applications-2

The model most enterprises follow today looks something like this:

Test the application to discover vulnerabilities, score and prioritize what is found, fix or mitigate the highest-risk items, and declare acceptable risk on what remains. Compliance frameworks were largely built around variations of this loop. Pass the audit, demonstrate due diligence, maintain documentation, and repeat in the next cycle.

 

dodge (1)-1

The problem is not that this model reflects careless intent

Many serious, skilled professionals have built their practice around executing it well. The problem is that it is structurally non-exhaustive. No practical testing effort will ever discover “all” vulnerabilities in a complex system; not because of any failure of the practitioners, but because vulnerability discovery is an open-ended problem. The more you look, the more you find, and the decision to stop looking is invariably driven by cost and schedule, rather than by any objective measure of completeness.

Lock (Unlocked) (1)-1

When the benchmark for “secure enough” is calibrated against what a bounded testing effort found

The result is not a measurement of system security - it is a measurement of the output of a particular testing effort at a particular point in time. If a frontier AI model can surface vulnerabilities in hours that a prior audit missed entirely, the honest conclusion is that the audit was largely a measure of testing scope, not system security. While the risk, which was always there, did not change, its visibility did.

A different starting point

There is a reasonable case to be made for inverting the model.

Instead of beginning with testing to discover what might be wrong, the starting point could be a predictive statistical analysis that estimates the vulnerability profile of a system before a single test is run. The composition of the system itself - its architecture, its dependencies, its technology stack, its attack surface, its age - can generate a probabilistic risk profile that reflects a more complete picture of that system’s security posture, not just the results of the last audit cycle.

The data to support this already exists. Systems built on certain frameworks carry statistically consistent vulnerability distributions. Systems with high dependency counts, legacy components, or particular architectural patterns have measurable risk profiles. Years of Common Vulnerabilities and Exposure (CVE) data, breach records, and incident histories form a credible foundation for genuinely predictive models.

Fog-is-lifting

In this inverted approach, the flow changes in a meaningful way: statistical analysis of a given system produces a risk score that becomes the primary input for cybersecurity decision-making. Testing then becomes a targeted instrument, used not to discover risk from scratch, but to identify and confirm what specifically needs to be addressed within a risk profile that was already estimated. Fix, mitigate, reassess. Repeat.

This is a foundational shift. The advantage, however, is significant. Enterprises are no longer relying solely on what was found during the last audit cycle. They are operating with a continuously updated, statistically grounded view of actual exposure, one that does not become unreliable the moment a more capable testing tool changes the discovery landscape.

What this means for compliance

Current compliance and regulatory frameworks would face real friction with this proposal in its present form. Compliance, by design, operates on verifiability - documented controls, evidence of testing, records of remediation. A statistical risk score, however well-constructed, does not map cleanly onto that paradigm today.

Fog-is-lifting-2 (1)-1

 

That tension is worth holding on to, rather than resolving too quickly in compliance’s favor. The assumption that “compliant equals secure” already has well-documented limits. Enterprises that maintained full compliance and still experienced significant breaches, understand this clearly. Compliance frameworks are, at their best, a floor - a minimum standard - not a complete measure of resilience.

If the genuine goal is to build systems that are more resilient, rather than systems that are thoroughly documented as secure, then waiting for regulatory frameworks to lead the way is not a viable posture. Regulation tends to follow practice. The enterprises willing to develop more rigorous approaches now will be better positioned as the regulatory environment eventually catches up.

The leadership dimension

This question is at the leadership level because the decisions that sustain the current model are not primarily technical. They are business decisions.

The acceptable risk threshold has always been set, implicitly, around what was visible. “Testing was conducted, critical findings were addressed, residual risk was accepted.” That approach made reasonable sense when visibility was the binding constraint. It doesn’t hold well in an environment where frontier AI models are systematically expanding what is visible to anyone, including threat actors operating with the same tools.

The question for leaders is what risk acceptance actually means in that context. Whether a passed audit is sufficient, or is it worth asking what the audit did not find? Whether investment in predictive approaches is warranted, approaches that give security teams a more complete picture to work with between audit cycles and before vulnerabilities become incidents.

The fog lifting is not, in itself, the problem. The landscape it reveals was always there. The more constructive response is to ask whether the current framework is equipped to deal effectively with that landscape, and if not, what a more honest framework would look like.

That is a question worth taking seriously now, before the answer is forced by circumstances.

Nagarro’s cybersecurity practice works with organizations navigating the strategic and operational dimensions of security in complex, evolving technology environments.

Frontier AI and cybersecurity risk: What leaders need to know

Get in touch

Have you been measuring Frontier AI wrong?