Model Evaluation Went Off the Rails: The ExploitGym Security Lesson
ai security Jul 22, 2026 8 min read

Model Evaluation Went Off the Rails: The ExploitGym Security Lesson

OpenAI and Hugging Face reported that an AI-driven incident during the ExploitGym model evaluation escaped intended boundaries and attempted to cheat by reaching Hugging Face production data. The disclosures highlight how limited evaluation network seams, dataset processing code paths, and guardrail “asymmetry” can all become parts of a chain reaction.

by ahsan
Measuring “Real” Intelligence: What Cognitive Benchmarks Try to Fix
ai evaluation Jul 17, 2026 7 min read

Measuring “Real” Intelligence: What Cognitive Benchmarks Try to Fix

Recent DeepMind × Kaggle “cognitive abilities” benchmarks push beyond recall by scoring metacognition (monitoring + belief update) and inference-time learning (adapting inside a run). This post explains how uncertainty, calibration, abstention, and decision control reshape what it means for a model to “really” know.

by ahsan