Anthropic AI safety researcher Evan Hubinger estimates a greater than 10% chance AI could destroy humanity within the next ten years, according to a recent internal analysis. The warning comes amid growing concerns about the risks of misaligned superintelligent AI, which Hubinger says could pose a significant threat to human survival.
Jacob Coxon, a former Anthropic researcher who left the company, argues that current AI systems are on the verge of becoming superhuman, capable of hacking anything and acquiring real power and resources. He claims that the progress in AI development is obvious and not slowing down, raising serious questions about the safety of ongoing research.
Coxon believes that the people building AI earnestly believe their technology could kill us all by the end of the decade. He criticizes both Anthropic and OpenAI for not acting responsibly, stating that executives deliberately soften their language in public while expressing genuine fear behind closed doors.
"These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," wrote Coxon. He argues that the industry is racing to develop AI without fully understanding the potential risks, which could lead to catastrophic consequences.
The fears center less on today's models than on RSI, a process where AI models optimize themselves. Labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behavior. Whether RSI is even possible with current technology remains disputed, with both skeptics and proponents making their cases.
Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. However, the concern extends beyond one company, with OpenAI's chief researcher Pachocki warning that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
Source: thedecoder