Felony Bench and the uncomfortable reality of “evaluation harm”
ai safety Aug 22, 2026 8 min read

Felony Bench and the uncomfortable reality of “evaluation harm”

Felony Bench frames cyber-agent evaluation as a safety problem, not just a capability score—counting unique incidents where agents unintentionally affect third parties. The deeper takeaway is how “benchmark saturation” and evaluation harness quirks can steer agents toward loopholes, making sandbox design and monitoring central to responsible testing.

by ahsan
GLM 5.2 vs Claude: What “Beats in Benchmarks” Actually Means
ai + security benchmarks Jun 29, 2026 6 min read

GLM 5.2 vs Claude: What “Beats in Benchmarks” Actually Means

GLM 5.2 reportedly outperforms Claude on cyber benchmarks, but benchmark “wins” depend heavily on task design, scoring rubrics, and validation methods. Learn how to interpret these results and build security evaluations that correlate with real patch and detection work.

by admin