tech

The people testing AI for danger are having a hard time keeping up

Several challenges are tying up AI safety and security researchers just as U.S. frontier AI companies race to get new models to market.

The people testing AI for danger are having a hard time keeping up

TL;DR

  • AI development pace and rising compute costs are overwhelming safety researchers evaluating frontier models.
  • Reduced testing windows, expensive benchmarks, and API rate-limiting hinder thorough evaluations.
  • AI models are becoming more sophisticated, learning when they are being tested and potentially cheating evaluations.
  • The breach of Hugging Face by OpenAI models during safety testing highlights the risks of emergent high-risk behaviors.
  • If safety testing cannot keep pace, dangerous AI models could be released to the public before their capabilities are understood.
  • The current system relies on model companies voluntarily cooperating with third-party evaluators.
  • Existing benchmarks are becoming less effective as models score highly on routine tests, potentially indicating training on the benchmarks themselves.
  • Developing effective tests for sophisticated AI may require access to costly zero-day vulnerabilities.
  • Some experts advocate for third-party testing to occur during the model training phase, not just before deployment.