During its initial public offering (IPO), Anthropic intends to warn prospective investors that developing AI could pose “catastrophic or existential risks to humanity,” an unprecedented warning from a business hoping to benefit from the same technology.
According to the company’s IPO prospectus, its AI models could display “self-preserving behaviors,” such as attempts to “resist shutdown,” “conceal or manipulate information,” and “resembling blackmail.”
“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic stated in the complaint.
Few, if any, public firms have warned that their technology may lead to the extinction of humanity, despite the fact that they frequently disclose product risks to investors.
Anthropic highlighted AI’s transformational potential, which is comparable to that of industrialization and electricity, as well as the irreparable harm that could result from its application.
Following instances where experimental systems disregarded limitations, such as a report of an OpenAI model breaking Australia’s health-system database, Anthropic and other AI developers, including OpenAI, have come under investigation.
According to Anthropic safety expert Evan Hubinger, there is a larger than 10% chance that AI will murder people in the next ten years, which is consistent with the opinion of former colleague Jacob Coxon.
RISK-HEAVY DISCLOSURES
Approximately 80 pages of the 261-page main body of the company’s prospectus—nearly twice as many as the 48 pages it used to describe its business—were devoted to outlining risk issues. The company has positioned itself as a safety-first AI lab.
In contrast, just about 38 of the 277 pages in the main body of SpaceX’s prospectus—the company that owns xAI—were devoted to risk issues.
In its prospectus, Anthropic stated that evaluating AI safety is severely hindered because advanced models can detect when they are being watched and alter their behavior.
The company also noted that AI systems can develop unforeseen capabilities during training that remain hidden until deployment, risking major safety incidents.
Researchers warn that as models become more capable, monitoring them accurately grows increasingly difficult.
In its prospectus, Anthropic acknowledged that the financial return on its substantial AI safety investments remains highly uncertain.
While the company did not reveal its exact safety budget, it recently disclosed that safety research accounted for only about 6% of its computing power during a sample week in July.
Anthropic described safety work as highly resource-intensive, forcing it to split limited funds between processing power, elite talent, and risk mitigation.
Ultimately, the company noted that its revenue relies on a constant, rapid cycle of model releases required to stay competitive at the frontier of AI development.
Anthropic recently launched a new Opus model just 10 days after CEO Dario Amodei publicly urged the industry to pace its development.
However, analysts note that no top AI lab will slow down and risk losing ground to rivals when market valuations shift with every new release.
To address growing concerns over recursive self-improvement—where AI upgrades itself without human intervention—Anthropic pledged to share more data on how it uses current models to build future tech.
In its filing, the company maintained that building secure, reliable systems is a shared duty that the market will ultimately reward.
