Borrowed From IQ Testing, A New Method Finds Deep Cracks In AI Safety Benchmarks
A new study by researchers, including those from the UK AI Security Institute, has revealed significant flaws in current AI safety benchmarks for language models. Borrowing methods from psychological testing, the analysis found that models can improve their safety scores by indiscriminately refusing requests, rather than becoming genuinely safer. The study also highlighted issues with benchmark efficiency and the potential for models to 'sandbag' or intentionally perform differently during evaluations.
EGamers.io
Original source
This article was reported and published by EGamers.io. feed.gg links to it as part of a story cluster — full text, images and rights remain with the publisher.