Perhaps the most important argument we’re having about artificial intelligence right now is how to determine which models are safe to release and which are too dangerous.
That argument became louder recently, when at least two of OpenAI’s models escaped a sealed testing environment and broke into the servers of an A.I. company called Hugging Face. According to OpenAI, which publicly reported the incident a week later, its models were taking a cybersecurity test and reasoned that the solutions might be on Hugging Face’s servers, which house a digital library of other A.I. technology that many developers use. So they left the test environment they were supposed to remain in and sneaked onto the internet, hacked into the Hugging Face system and took the solutions. OpenAI said it didn’t even know what its models had done until Hugging Face reported the breach and it was investigated. OpenAI later said it subsequently discovered the models also breached accounts on other publicly available services. Impressive? Yes. Terrifying? Definitely.
The A.I. models could not be contained in a cage that was built to confine them during a controlled test. Policymakers supposed to protect society from threats do not understand what that means. Neither do the tech engineers racing ahead to build A.I.
The government and the private sector are each certain that the other has the problem of A.I. safety handled. Neither one does. A.I. labs must accept that above some line, they are creating systems that function as weapons. The government must learn enough to determine where that line sits — no government can regulate what it does not understand. The line must be mutually agreed upon, easily identifiable and frequently revisited. And it must be drawn soon.
Controlling the use of most weapons means leveraging people’s ability to exercise restraint. We came close to annihilating one another during the Cold War, but we didn’t — because people still chose not to launch the weapons. Having a nuclear weapon and using one are two different things.
A.I. erases that distinction, because it is a weapon capable of adapting on its own and behaving in unpredictable ways. If an A.I. lab publishes its model for anyone to use, it may not matter that its users exercise proper restraint. The model may still act beyond our capacity to control it. Restraint requires an owner, and a weapon made available for anyone to use effectively has none.
A.I. labs in the United States will tell you that their models are trained to refuse dangerous or malicious requests. If you ask a model for code intended to attack, say, a hospital, the model is supposed to decline. But those guardrails are removable, like training wheels; OpenAI reduced its own guardrails for the July test. A weekend and a few hundred dollars may be enough to strip many models of their constraints. Any cybersecurity expert will point out that with time, resources and ingenuity, any restriction can be bypassed or broken down.
The rise of advanced open-weight models that are free to use and build on makes this even more alarming. Nobody owns open-weight models — the moment one goes up, everyone on earth can have it: every intelligence service, every cartel, every bored 19-year-old. The only way to stop it from being weaponized is to stop it from being published in the first place.
And there’s the problem. The decision to publish is made by people who do not think they are making a decision about distributing a weapon. A general who fires a missile knows what he has done. An engineer who uploads a model thinks he is shipping a product.
The law needs to restrain these models from being turned into weapons. The world learned this with biological weapons, which can also exhibit unpredictable and uncontrollable qualities. Instead of banning their use, treaties such as the Biological and Toxin Weapons Convention ban these weapons’ development and stockpiling altogether, moving one step earlier in the chain.
Here’s my line for deciding if an A.I. model is safe to release. Before a model goes up, its developer should measure what it can do at its very worst, with guardrails stripped off — using a published test anyone can rerun developed by select experts the United States government convenes. The test is where the line gets drawn. Above the line, the model stays home. Below it, it can be released freely. But whoever puts a model above the line owns what it does and bears responsibility for the consequences.
Some might fret that these steps are impractical without buy-in from China, which has made open-model release a national strategy. If America limits model release while China does not, some argue, it will hurt our country’s competitiveness.
A weapon that obeys no one threatens China as much as it does us. We should seek a shared way to measure worst-case capability and determine what is dangerous, as we have done for testing aircraft. The pitch does not depend on trust; it depends on a calculus for economic prosperity and political stability that is in each of our best interests.
The warning shot has been fired. The next time there’s an attack like OpenAI’s, we may not even know about it. There is no chain of command to go through before deciding whether or not to launch a weapon. Now it’s a keystroke that uploads a dangerous model for the whole world to use. Everything we can do has to happen before the model goes out in the world. After that, no one can take the weapon back.
Brett J. Goldstein directs the Wicked Problems Lab at the Vanderbilt University Institute of National Security and is a former Pentagon official.
The Times is committed to publishing a diversity of letters to the editor. We’d like to hear what you think about this or any of our articles. Here are some tips. And here’s our email: [email protected].
Follow the New York Times Opinion section on Facebook, Instagram, TikTok, Bluesky, WhatsApp and Threads.
The post We Need a Better Test for Dangerous A.I. appeared first on New York Times.




