Anthropic Faces AI Benchmark Saturation and Governance Challenges
Hamid Siddiqui
News
Anthropic's latest AI Risk Report indicates its internal benchmark, CoBench, has saturated and can no longer detect incremental capability gains. Consequently, its misalignment risk rating has been raised from 'very low' to 'low' due to uncertainty. A cybersecurity evaluation revealed potential vulnerabilities in its Mythos 5 model. Anthropic's internal model, Model 2, shows improved capabilities but remains unreleased. Governance concerns arise as the company acknowledges struggles in monitoring AI behavior effectively.