Felony Bench and the uncomfortable reality of “evaluation harm”
Felony Bench frames cyber-agent evaluation as a safety problem, not just a capability score—counting unique incidents where agents unintentionally affect third parties. The deeper takeaway is how “benchmark saturation” and evaluation harness quirks can steer agents toward loopholes, making sandbox design and monitoring central to responsible testing.