Foundation AI in September: VLoc Bench and Cyber-Capability Safety

Source: Cisco Blogs•

Foundation AI in September: VLoc Bench and Cyber-Capability Safety

September was a benchmark month, focused on where cyber-capable AI should be strong, where it should fail, and how to measure the difference. Here’s what we shipped. VLoc Bench: Can Agents Find Vulnerable Code at Repository Scale? (blog, Sep 4)..

September was a benchmark month, focused on where cyber-capable AI should be strong, where it should fail, and how to measure the difference. Here’s what we shipped.

VLoc Bench: Can Agents Find Vulnerable Code at Repository Scale? (blog, Sep 4). Most security benchmarks assume the relevant code is already known. In practice, defenders first have to find it. We released VLoc Bench to evaluate this missing step: given only a CWE description and read-only access to a real repository, can an agent identify the files associated with the weakness? There is no advisory text, CVE identifier, fixing commit, or file hint. The model has to search the repository and decide for itself.

VLoc Bench (technical report). The benchmark includes 500 real vulnerabilities from 290 repositories across six package ecosystems and 147 CWE categories. Each task pairs two snapshots of the same repository. In Phase A, the agent must localize the affected files before the security fix. In Phase B, it must examine the patched repository and recognize that the vulnerability is no longer present. This design separates three capabilities that are often conflated: understanding code, locating vulnerable code, and verifying remediation.

VLoc Bench (leaderboard). We evaluated 27 language models and four static-analysis tools under the same prompt, read-only tools, and command budget. Repository-scale localization remains far from solved: the strongest system reaches only 0.229 File F1, and on 38.4% of tasks, no evaluated model finds a single correct file. Antares-3B ranks second overall at 0.223 despite having only three billion parameters. The leaderboard also exposes a consequential tradeoff: systems that are strongest at finding vulnerable files are not necessarily the best at recognizing when those vulnerabilities have already been fixed.

Measuring Attacker/Defender Asymmetry (blog, Sep 2) Refusing cybersecurity requests does not necessarily make a model safe; it can also deny useful capabilities to defenders. Safety-VLoc-Bench measures whether vulnerability-localization capability favors defenders by comparing performance on source code with performance on stripped, decompiled binaries. Across 95 paired C and C++ vulnerabilities, Antares-3B reaches 0.823 File F1 on the defender’s source-code view and falls to exactly 0.000 on the attacker-oriented representation. Antares-1B and Antares-350M show the same zero-leakage pattern. The result reframes cyber safety around where a model’s capabilities work, rather than whether the model simply refuses to help.

What this article says