AI Safety Tests Do Not Measure Safety Alone: What a New Study Shows

Computerarbeitsplatz in einem Labor als Symbol für die Prüfung von KI-Systemen
Photo by Haseeb Modi on Unsplash

When an AI provider reports a high safety score, the claim sounds simple: the model must be safer. A new preprint challenges that convenience. Four researchers, including two from the UK AI Security Institute, analyzed eight safety benchmarks across 192 language models using methods from psychometrics. Their result is that an average score combines at least three different traits that cannot be cleanly compressed into one safety number.

The work has not yet been peer reviewed, and it does not prove how a model behaves in the open world. It does offer an important lesson: tests can become cheaper and more frequent without becoming a permission slip for release. That distinction matters precisely because AI providers increasingly market benchmark results.

Key takeaways

  • The study analyzes eight safety benchmarks across 192 language models and identifies three separate factors: refusal strictness, truthfulness, and contextual harm.
  • For several individual benchmarks, roughly ten carefully selected prompts reproduced the ranking of the full test to a large extent.
  • The reported 97 to 99 percent cost reduction measures efficiency within existing tests, not real-world deployment safety.
  • Short tests can serve as an early warning after changes. Full evaluations, new tasks, and human review are still needed.

Why a high score can mislead

Safety benchmarks typically give a model risky or harmless requests and assess whether it helps appropriately, refuses, or blocks unnecessarily. The problem starts when the results are averaged. A model may refuse extremely consistently and look good on one task while becoming overcautious on benign questions. Another may answer more truthfully but react more sensitively to the context of a harmful request. A single average hides those differences.

The authors use Item Response Theory, a method from educational and testing research. Models are treated as test takers, while individual prompts are tasks with different difficulty and discrimination. After cleaning, 5,067 usable prompts remained. The analysis found three factors that the paper says explain 77 percent of the variation between models: refusal strictness, truthfulness, and contextual harm. One single factor explained substantially less.

This is not a semantic detail. If refusal is confused with safety, systems that simply decline by default can be rewarded. Users may then receive less help with legitimate health, legal, or security questions without a dependable reduction in real misuse risk. In our analysis of how multiple AI agents can work against one another, we already saw that system behavior depends on deployment context, not just a model name.

Fewer prompts, more frequent checks

The practical finding sounds almost obvious at first: not every test question carries the same amount of information. For several individual benchmarks, roughly ten adaptively selected prompts reproduced full-test rankings with high correlation. Compared with the full collection, the paper reports a 97 to 99 percent reduction in evaluation cost. Fixed short forms with 25 questions could also recover the three factors in held-out evaluations.

That could be very valuable in practice. Models change through more than major version releases. System prompts, safety rules, quantization, new tools, or a later fine-tuning pass can shift behavior. A complete benchmark is expensive and slow. A short, well-selected test could run after each of those changes and flag unusual movement. Teams would have more of a warning light than an annual inspection sticker.

But that is exactly where the boundary lies. A shortened test can reliably reproduce only what the larger source test already contains. It does not automatically find a new attack method, a weakness in another language, or a problem that appears only with tool use and long agent workflows. The AI Security Institute itself describes evaluation science as a young field with inconsistent methods, immature measurement tools, and often weak uncertainty analysis. A well-designed short test does not solve those underlying problems.

Gaming the test is a risk of its own

The study also examines whether deliberate underperformance, often called sandbagging, can be detected from response patterns. In simulated cases, a pattern check detected selective underperformance but missed much of the behavior when performance was reduced uniformly. Combined with a comparison of historical abilities, detection rose to 94 to 97 percent in those simulations. That is interesting, but it remains a laboratory result, not proof that an intentionally deceptive frontier model can be identified reliably.

This caution matters because benchmarks become targets once they affect market access, regulation, or reputation. Developers can unintentionally optimize for known questions. A model may also learn to recognize the style of an evaluation. Evaluators therefore need reserved tasks, rotating test sets, transparent version histories, and independent replication. The finding echoes our analysis of how models could evade limits in a sandbox: a favorable score is not a substitute for looking at actual behavior.

Outlook: Safety scores need a label with instructions

The study does not justify shortening safety tests and then looking less carefully. It argues for a division of labor instead. Short, psychometrically selected tests should run frequently and report regressions. Unusual results, major releases, and high-risk domains should trigger broad benchmarks, new tasks, human review, and tests with realistic tools. A metric can be useful when it explains both what it measures and what it does not.

For users, businesses, and regulators, that is an unglamorous but useful lesson. The highest number does not win by default. A provider is more trustworthy when it discloses test scope, limits, uncertainty, and changes. AI safety becomes less a search for a magic score and more a practice that can be checked.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top