The “Best” AI Model Is Getting Harder to Define

The best AI model for application security depends on vulnerability detection, accuracy, cost, human review, data handling, and deployment needs.

Oct 7, 2026
5 minute read
two people monitor computer dashboards beneath four glowing network spheres and digital panels with status indicators and warning icons.

AI model benchmarks reveal major tradeoffs in security performance and cost. Image: ChatGPT

eSecurity Planet content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

For the past few years, choosing an AI model often meant looking for the strongest benchmark scores or selecting a familiar provider. 

Debates over which AI model was “best” became common among AI security researchers on social media. For security teams evaluating AI for vulnerability detection, those measures are no longer enough.

Raw model capability is only part of the decision. Teams need to know whether a model can find real vulnerabilities without overwhelming analysts with noise. 

Cost also matters, especially when security testing must scale across large codebases. 

Organizations must consider where sensitive code is processed and whether changes in provider policies could disrupt continued use.

The more useful question, then, is not which model leads a general-purpose leaderboard. It is whether the entire approach produces reliable results for the organization. 

That means evaluating the model alongside the surrounding workflow and determining whether the resulting system fits the organization's security requirements.

A lower model bill does not equal lower cost

Take insecure direct object reference (IDOR) vulnerabilities as an example. These access control flaws allow users to access or modify resources they should not be able to. IDORs are often among the first vulnerabilities new penetration testers learn to identify, but their apparent simplicity can be misleading.

Detecting an IDOR requires contextual reasoning rather than finding an obviously dangerous function. A model must understand who should have access to a resource and determine whether the necessary authorization check is missing.

Semgrep’s IDOR vulnerability-detection benchmark tests models against real vulnerabilities in open-source codebases. Detection performance is measured using precision and recall, with the F1 score balancing the two to provide an overall measure of quality. Semgrep also tracks the cost of each confirmed vulnerability.

In one recent run, China-based Z.ai's GLM-5.3 achieved a 23.8% F1 score at $0.15 per true positive. Claude Opus 4.8 produced a nearly identical F1 score of 23.6%, but each true positive cost $1.04. For this particular test, the models therefore delivered similar overall detection quality at substantially different costs.

That difference can become significant when application security teams are reviewing large codebases or frequently testing new changes. However, price alone does not make GLM-5.3 the better choice. Its recall in the same test was just 13.9%, meaning it failed to identify most of the known IDORs in the dataset.

Advertisement

GLM-5.3 also performed worse than its predecessor, GLM-5.2. Semgrep researchers caution that additional runs were necessary to determine whether that difference represented an actual regression or normal testing variance.

Ultimately, cost per true positive matters only in the context of the detection performance an organization requires. A lower inference cost provides little benefit if too many vulnerabilities go undetected or analysts must spend substantially more time compensating for the model's limitations.

The full cost includes human review

False positives illustrate another problem. Two models can produce similar aggregate scores while creating very different workloads for the engineers responsible for validating their findings.

A separate Semgrep benchmark of Kimi K3 demonstrated that tradeoff. Using the same guided-prompt configuration, Kimi achieved a 34.0% F1 score, close to GLM-5.2 at 34.5% and Claude Opus 4.8 at 34.3%. However, Kimi's precision was 68.4%, compared with 86.3% for GLM-5.2 and 91.0% for Claude Opus 4.8. As a result, engineers had to investigate and dismiss a larger share of its findings as false positives.

The repository-level results raised another concern. Kimi averaged roughly 6% F1 on the largest enterprise-style repository in the benchmark, while GLM and the frontier models averaged around 20%. One repository cannot establish that codebase size caused the decline, but the result shows why teams should test models against environments that resemble their own instead of relying solely on aggregate scores.

The true cost of AI-assisted security also extends beyond model usage. Organizations must account for the infrastructure required to operate the system and the engineering effort needed to support it. 

Missed vulnerabilities introduce another potential cost that will not appear on a provider's pricing page. Together, these factors determine whether a system is viable in practice.

The model itself is only one component. The surrounding harness can materially affect security performance by shaping how the model receives information and performs its analysis. 

In the same benchmark set, GPT-5.6 Sol's recall increased from 25% with a guided prompt to 73% in Semgrep's multimodal harness. The underlying model remained the same, but the system around it produced a substantially different result.

For security leaders, that difference changes the evaluation process. Knowing which model a product uses is less informative than understanding how the complete system performs against representative code and how much work its output creates for security teams.

Advertisement

Geography changes the calculation

There's another wrinkle here. Some compelling alternatives aren't coming exclusively from U.S. frontier labs.

GLM-5.2, for example, is open weight, allowing organizations to download and operate it on their own infrastructure. In Semgrep testing, it outperformed Claude on an IDOR benchmark while costing roughly $0.17 per vulnerability found.

For organizations handling sensitive intellectual property or operating under strict data requirements, where code is processed may matter as much as benchmark performance. An open-weight model that runs within an organization's own environment presents a different security and governance proposition from a closed model accessed through an external provider.

Resilience matters as well. Providers can change model availability or pricing with little control from customers. Policies governing access can also shift over time. Building a critical security workflow around a single provider means accepting that external dependency.

None of this means foreign or open-weight models deserve a free pass. Organizations still need to evaluate where a model comes from and whether its behavior is reliable enough for security work. They also need to determine whether the deployment approach meets their security requirements. Country of origin isn't a useful shorthand for capability, just as a recognizable U.S. brand doesn't guarantee that a model is the right fit.

The growing quality of models like GLM and Kimi gives security teams something valuable: leverage and choice.

Benchmark the job, not the brand

There is unlikely to be one best AI model for every application security program. 

A team prioritizing broad vulnerability discovery may place more weight on recall. Another struggling with alert volume may favor precision. 

An organization with strict data residency requirements may prioritize self-hosting, while a team reviewing large amounts of code may place greater weight on cost per confirmed finding.

A useful evaluation starts with representative codebases, including the largest repositories a team expects the system to handle. Teams should measure precision and recall while examining which vulnerabilities go undetected. Cost calculations should account for the resources required to operate the system as well as the effort needed to review its findings. 

Advertisement

Testing should also reflect the actual environment surrounding the model rather than evaluating the model in isolation. These assessments need to be repeated as the technology and provider offerings change.

For application security leaders, model selection should look less like a leaderboard decision and more like an engineering evaluation. 

The most recognizable model may still be the right choice, but it should win because it delivers the best security outcome under the organization's constraints, not because its name appears at the top of someone else's benchmark.

AI models are only one option for finding security weaknesses in code and infrastructure. See our guide to the best vulnerability scanning tools to compare additional approaches to vulnerability detection.


Katie Paxton-Fear

Staff Security Advocate at Semgrep

eSecurity Planet Logo

eSecurity Planet is a leading resource for IT professionals at large enterprises who are actively researching cybersecurity vendors and latest trends. eSecurity Planet focuses on providing instruction for how to approach common security challenges, as well as informational deep-dives about advanced cybersecurity topics.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.