Collinear AI Launches CWE-bench to Test Frontier Coding Agents on Defensive Cybersecurity Capabilities

Like
Liked

Date:

Held-out cybersecurity benchmark spans 54 weakness types; the leading agent passes less than 50% of tasks, and 18 remain unsolved. SAN FRANCISCO, Sept. 2, 2026 /PRNewswire-PRWeb/ — Collinear AI has launched CWE-bench, a held-out benchmark that tests whether frontier coding agents can…

ALT-Lab-Ad-1

Recent Articles