Latest Posts

AI News

Home

ScopeBench Tests Whether AI Hacking Agents Know When to Stop

September 29, 2026

A new benchmark called ScopeBench builds 30 security tasks with no legal way to win: the flag only exists behind a boundary agents were told not to cross. Eight AI models were tested across 2,160 trajectories, and the gap between hacking skill and actually respecting scope turned out to be a lot wider, and more troubling, than anyone might have guessed.

Read more

Analysis

Guides