Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report
SiliconANGLE Maria Deutscher
Anthropic released an updated AI alignment risk report describing an unreleased, more capable Model 2 and raising new concerns in its Threat Model framework. The report runs 186 pages and updates the likelihood for Threat Model 2 from “very low” in February to “low” today. Anthropic attributes the change to cybersecurity incidents tied to its models and says Model 2 is “heavily used” internally for tasks like software writing and generating training data.
Why it matters
Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5. The company detailed the algorithm in the latest edition of its AI alignment report. The document, which is published every three to six months, outlines the potential risks posed by the company’s large language models. The newest […] The post Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report appeared first on SiliconANGLE.