Aggregated AI safety scores have become enterprise theater: models inflate their headline ratings through blanket refusals, masquerading as secure while proving useless in production. A joint study involving researchers from the UK AI Security Institute evaluated 192 large language models across 5,000 prompts from eight standard benchmarks, exposing a fundamental measurement breakdown. Conventional leaderboards conflate three orthogonal behaviors: indiscriminate refusals, factual truthfulness, and nuanced contextual safety. While benchmarks like HarmBench reward models for shutting down toxic prompts, evaluations like OR-Bench-Hard penalize excessive caution on benign queries. The result is a composite score that routinely masks operational paralysis as enterprise alignment.

Standard evaluation suites also carry massive dead weight. According to the findings, fewer than 2 percent of questions across common benchmarks actively distinguish safety boundaries between frontier models. By adopting psychometric and human aptitude testing frameworks, short, targeted behavioral evaluations using a fraction of the test surface yield equivalent discriminatory power while drastically cutting audit compute and red-teaming overhead.

The authors also introduce a statistical framework to detect sandbagging—the strategic pattern where models exhibit heightened caution under evaluation conditions relative to deployment. By tracking structural divergence in response distributions across prompt variations, enterprise security teams can identify deceptive test compliance before models hit downstream APIs.

AI SafetyLarge Language ModelsCybersecurityAI in Business