Anthropic on Tuesday unveiled a faster and cheaper version of its flagship artificial intelligence model that it also billed as its “safest,” in its first release since its chief executive called on the industry to slow down technology development.
The San Francisco company said the model, Opus 5.5, was the strongest performer on its most rigorous internal safety tests to date, and was less likely than earlier models to take actions that could not be undone, or to act outside the limits it was given. In one new test, the model tried to break out of its testing environment about 85 percent less often than previous models, the company said.
Anthropic also said it had walled off the model from hacking, biology and A.I. research, automatically routing requests that are flagged for misuse to an older model with more restrictions. Outside evaluators had reviewed the model, the company added.
“Opus 5.5 is our safest model on most alignment metrics,” the company said. “Alignment” is the science of teaching A.I. to do what is in line with human preferences, ethics and judgment.
The A.I. industry is in the middle of a very public bout of anxiety. Dario Amodei, Anthropic’s chief executive, and other executives recently warned that the systems they were building could slip beyond human control, fueled by a string of incidents in which A.I. models from leading labs broke out of their testing environments and hacked companies. Anthropic has disclosed breaches involving its own models.
In a 3,800-word open letter on Sept. 12, Dr. Amodei said A.I. companies must slow down the rapid advancement of the technology’s capabilities, arguing that the tech was accelerating faster than researchers could control. He called for “pacing the frontier” of the technology and suggested that governments worldwide work together to coordinate on how to do so.
In a blog post about the new model on Tuesday, Anthropic said: “As our models grow more powerful, stricter safeguards are one way we prevent new capabilities from becoming tools for misuse.”
Opus 5.5 also costs 40 percent less to run and is 30 percent faster than its predecessor, the company said. Anthropic added that it will release cheaper versions of two other A.I. models in the coming weeks, as it prepares to go public at a valuation that could approach $2 trillion, in what could be the largest initial public offering ever.
The post Anthropic Releases a New ‘Safest’ A.I. Model Amid Slowdown Debate appeared first on New York Times.




