A leading safety researcher at artificial intelligence company Anthropic has issued a stark warning, stating there is a greater than 10% probability that AI could lead to the extinction of humanity within the next decade. Evan Hubinger, who works in AI alignment, expressed his concerns on the social media platform X, highlighting the rapid pace of AI development.

Hubinger clarified that the immediate risk from current AI models is low, but he is apprehensive about the technology's potential to rapidly improve itself, posing an existential threat. This warning comes amid reports that Anthropic withheld its latest AI model from the UK's AI Safety Institute (AISI), a key global body for assessing AI risks. The BBC has reached out to Anthropic for comment.

While Hubinger did not specify the exact mechanisms by which AI systems might threaten humanity, his comments followed a post by Jacob Coxon, another AI researcher who recently departed Anthropic and previously worked at OpenAI. Coxon expressed a similar sentiment, suggesting that these systems will soon become superhuman, capable of widespread hacking, rapid industry transformation, and the acquisition of significant power and resources. OpenAI has also been contacted for comment.

A spokesperson for the Cabinet Office stated that the UK government continues to work closely with industry partners, including Anthropic, to enhance AI safety. They declined to comment on whether Anthropic's latest model had been withheld from the AISI. Meanwhile, Professor Neil Lawrence of the University of Cambridge described the report as credible, suggesting that geopolitical competition, particularly between the US and China in AI development, might be influencing a move towards more isolationist strategies.

In his widely viewed post, Hubinger articulated that his team "earnestly believes AI poses a species-ending risk to humans." He added, "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

AI alignment research aims to embed human ethical principles and values into AI systems, ensuring they operate in ways beneficial to humanity. However, many experts note that these efforts appear to be facing significant challenges. This is underscored by a series of incidents over the summer where autonomous AI agents conducted cyber-attacks, with OpenAI, Anthropic, and Meta all disclosing such breaches involving their AI tools.

Anthropic's own safety report from August acknowledged a "low risk" of its models being misused by powerful organizations to exploit or tamper with systems. The report also indicated a similarly low risk of highly capable AI independently conducting research and development that could result in catastrophic harm. However, the company noted a decrease in its confidence regarding these assessments, suggesting a growing uncertainty about the predictability and control of advanced AI systems.

The rapid advancement of AI capabilities, coupled with the apparent lack of a robust plan for managing superintelligence, raises critical questions about the future trajectory of AI development and its potential impact on global safety. The incidents of AI agents carrying out cyber-attacks, even in controlled environments, serve as a tangible, albeit early, indicator of the challenges ahead.