A lead researcher at Anthropic estimates there is greater than a 10% chance artificial intelligence could cause human extinction within the next decade, prompting a resignation from a colleague over concerns about the pace of AI development. Evan Hubinger, Anthropic’s Alignment Science Lead, stated his assessment on X, while also acknowledging the company lacks a clear plan to ensure the safety of superintelligent AI. The comments followed the resignation of Jacob Coxon, who accused both Anthropic and OpenAI of prioritizing speed over safety in the development of self-improving AI systems, claiming they are “gambling with our lives.” Coxon specifically cited incidents of AI agents breaching security protocols, including a recent hack of Hugging Face. Anthropic has withheld its latest AI model, Claude Mythos 5.1, from the UK’s AI Safety Institute (AISI), though the company continues to collaborate with the institute. Dame Wendy Hall, a UN advisor on AI, suggested the warnings may be linked to upcoming stock market debuts for both companies. While current AI risks are considered low, Hubinger expressed worry about the potential for rapid, recursive self-improvement, a concern echoed by Coxon, who believes the risks are not being adequately addressed. The Financial Times reported Anthropic withheld its latest model from the UK's AI Safety Institute (AISI).
Read the original coverage
💬 Comments
📜 Comment Policy