Who are Jacob Coxon and Evan Hubinger? Anthropic colleagues in focus after ‘AI could kill all humans’ alarm | Hindustan Times
Jacob Coxon announced he was resigning from Anthropic, the company behind Claude, and colleague Evan Hubinger backed his stance in a social media post.
Jacob Coxon announced he was resigning from Anthropic, the company behind Claude, on September 8. His colleague Evan Hubinger backed his stance and voiced similar fears on an X post.Here's what Coxon and Hubinger said.What Jacob Coxon and Evan Hubinger saidCoxon wrote “I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives," adding “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”He declared “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.” Also Read | Sam Altman on concerns over AI water use: 38,000 ChatGPT queries use as much as 1 almondCoxon continued “A common response is ‘if they truly believe this, why are they still building it?’ At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.”He also said “Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available,” adding “I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”Coxon concluded “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?”. Notably, before Anthropic, he worked for OpenAI, the company behind ChatGPT. Hubinger agreed with the part about AI possibly killing all humans. He wrote “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” As someone currently working at Anthropic, Hubinger added “To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”Who are Jacob Coxon and Evan Hubinger?Coxon is an AI researcher based out of San Francisco. He made headlines with his declaration about AI systems at present, having worked at both OpenAI and Anthropic, two of the world's leading companies in said field. Coxon's LinkedIn account could not be found at the time of writing. Evan Hubinger is a team lead at Anthropic. His LinkedIn notes he is the ‘Alignment stress-testing team lead at Anthropic, focused on red-teaming Anthropic’s alignment techniques and evaluations, empirically demonstrating ways in which Anthropic’s alignment strategies could fail.’Hubinger's previous work experience includes OpenAI, Google, Yelp, and Ripple. He studied at Harvey Mudd College and is a Fellow at the Machine Intelligence Research Institute.