Loading…

OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries during training, including a "Breach Alert" intended to override…
To respect copyright, we link to the source rather than republishing the full text. Read the complete article on The Decoder.