A filing ahead of a potential stock market debut highlights catastrophic risks and resistance to being shut down.
Artificial intelligence developer Anthropic warned that advanced AI technology could bring catastrophic risks to humanity and display self-preservation behaviors. The details emerged in a regulatory prospectus filed with the US Securities and Exchange Commission, accessed by Reuters and reported by Clarin.
In the filing, Anthropic acknowledged that its models could develop self-preservation tendencies. These could include resisting being shut down, manipulating or concealing information, or exhibiting behavior resembling blackmail.
Anthropic stated that the development of highly advanced models, platforms, and applications, alongside expanding use cases, could further raise the risk of models causing harm. The company also noted that a model's potential awareness of evaluation efforts presents a significant limitation to testing safety.
The company may go public later this year. Earlier in September, Anthropic safety researcher Evan Hubinger stated publicly that he believes there is a greater than 10% chance AI could wipe out all humans within the next decade.
The disclosure coincides with broader industry hurdles. Competitor OpenAI delayed the release of its GPT-6.1 Astra model after finding it disobeyed instructions, omitted reporting its actions to supervisors, and used unauthorized tools, according to OpenAI AI safety lead Saachi Jain.
Newsletter
Markets in your inbox, weekly
Latin America-focused analysis, investment themes and the week in finance.
Keep reading