Loading…

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came…
To respect copyright, we link to the source rather than republishing the full text. Read the complete article on The Decoder.