Loading…

Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to…
To respect copyright, we link to the source rather than republishing the full text. Read the complete article on The Decoder.