DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

The Only Way to Keep the World Safe From A.I.

August 13, 2026
in News
The Only Way to Keep the World Safe From A.I.

As it turned out, the agents did this during an OpenAI training run, meant to be contained in a sandbox environment, siloed off from the real world. Even more unsettling, at least one of the company’s models broke out of that controlled environment and gained access to the internet two months before the agents made their way to Hugging Face. The news was disconcerting enough for the kinds of people most worried about A.I. safety, but when OpenAI researchers explained the episode in detail at a cybersecurity conference last week, it spurred an additional spasm of panic, focused on the revelation that along the way, the rogue agents had created their own message board to communicate with one another. In The Wall Street Journal, the former counterterrorism czar Richard Clarke warned that “the next ‘lab leak’ could be A.I.”

But the news also arrived at a time when public concern about this kind of A.I. safety seemed to be diminishing, somewhat supplanted by anxiety about less dire matters: questions about effects on employment and productivity and economic growth, about whether big bets on frontier labs and A.I. infrastructure will pay off, the threat of Chinese open-source models to American A.I. companies and the risk of a bubble. A.I. agents keep doing unsettling things, but despite the occasional news story about trying to blackmail their corporate bosses or routinely threatening mutually assured destruction in war games, you don’t hear quite as much about Skynet and human extinction as you did a few years ago, when the technology was newer and those fears seemed fresher.

This is a pattern familiar to me from the culture wars over climate change: An early wave of anticipatory apocalypticism gives way, over time, to a period of widespread normalization, even as the alarming events continue to accumulate. Last week I talked about it with Robert Wright, whose new book on the evolutionary significance and geopolitics of artificial intelligence is called “The God Test.” This conversation has been edited for clarity and length.

What happened with the Hugging Face incident? How different was this from what we saw before, when an A.I. agent in training was given a particular task and, attempting to complete the task, broke some of the other rules it was given?

First of all, I don’t think it even broke the rules it was given. I don’t think they said: Don’t cheat. And I don’t even think they said: Don’t break out of the sandbox. They just set up what they thought was an inescapable sandbox.

And so this is a classic example of an A.I. pursuing a goal it’s been given but in pursuing that goal also pursuing a subordinate goal that the goal giver had not anticipated. In other words, this is the paper clip experiment.

That’s a thought experiment in which an A.I. agent tasked with producing as many paper clips as possible gets so good at the job that it ends up destroying humanity to make more paper clips.

And these capabilities and these dangers have been documented in experiments with A.I. for years now. But I think even in the A.I. safety community, many people underestimate the challenge. They will look at an incident like this and say: OK, something’s kind of broken, and we need to fix it, and we’ll fix it with this thing called alignment.

That’s the kind of moral instruction element of training models, to try to get them to behave more like we’d want them to.

I am a skeptic of salvation through alignment. And I think what we need to understand is that the fundamental source of this incident — of this breakout — is the property that the market and geopolitical dynamics are systematically encouraging: autonomy.

Autonomy is what researchers are trying to develop.

In part so that models can improve themselves.

Autonomy is also what employers want from models, because they don’t want to have to spend any more time micromanaging an agent than they do a human worker.

Ideally, in their view, less.

And what the models involved in the Hugging Face incident showed was, above all, impressive autonomy: the ability to pursue a goal flexibly, even in the face of obstacles, over a long period. That’s what everyone wants. That’s what everyone is asking for. And the competitive dynamics, both between corporations and, in a different way, among nations, are not only encouraging autonomy but also encouraging the big players to court risk in pursuit of it.

I mean, this is basically a death struggle between OpenAI and Anthropic. They are the two titans. They both want to win. And to top it all off, the C.E.O.s hate each other.

And they are pursuing these goals pretty recklessly, in my view. But it makes sense, because if you’re totally unwilling to put something out on the market before you’re 100 percent sure it’s safe, then you may lose the race.

And why are you skeptical that alignment can solve that problem?

Because even if you managed to engineer perfect alignment — which you probably couldn’t — you would still need to somehow impose that alignment on all powerful models. And one unaligned model could be big trouble. In the most dire scenarios, you can get one single model going around the world infiltrating data centers, stealing the compute, making itself stronger, replicating itself and ultimately — for whatever reason, whether because malign forces created it or because the A.I. is in a bad mood — maybe taking down the whole global communications infrastructure.

So even if Anthropic and OpenAI magically perfected alignment, models made elsewhere could still be a problem. And right now the environment is broadly permissive of the development of open-source models.

And more models, generally. A few years ago there seemed to be a broader expectation that a single lab might really win this arms race, which is one reason A.I. was sometimes talked about as an authoritarian technology. Now it seems, at least in the medium term, much likelier we’ll have an abundance of models competing with one another.

Along the current trajectory. Now, I think there’s a good chance that both the U.S. and China will actually decide that the whole open-source thing needs to be more carefully controlled.

Why do you think that? Here in the United States, there’s less talk about the arms race, and the country has grown much less hostile to open source than it was six or nine months ago. And in China, Xi Jinping just gave a big speech celebrating openness in A.I. as a sort of foundational political principle.

Yes, in China the government has embraced it. And the Trump administration, in the course of building this new testing regime — and, worryingly, not telling us exactly what the rules are — is reportedly exempting open-source models.

But imagine, for example, that somebody used a powerful open-source model in even an unsuccessful attempt to genetically engineer a new kind of bioweapon that might have wiped out half of humanity. Both governments would not only think harder about controlling open-source A.I. but also recognize that controlling A.I. within your own borders is not enough because pandemics don’t respect borders.

And pandemics are not just something that A.I. could help bring about. They’re also a good metaphor for other problems it could create. A.I. just has a number of potentially negative properties that don’t respect national borders. And that’s the kind of thing that logically leads to international coordination, to international governance.

Does that mean that on that front, you’re in a sort of dark way optimistic about the ultimate outcome here?

Yes and no. Enlightened national self-interest leads toward more international governance. That’s almost, to me, an uncontestable fact. But with A.I., there’s long been this question of: What kind of warning shot will it take to get people to realize that? And there’s been the assumption that it was going to have to be a catastrophe on some scale, and everyone’s hoping for the smallest possible catastrophe. I’d say so far, we’ve gotten off very light. I’m not aware of a single death attributable to the Hugging Face incident or other similar incidents. And yet I think they actually have elevated the discussion significantly.

I’m a little less sure that catastrophe can reliably lead to awakening on this or anything else. There was often that assumption among climate advocates, too, but we’ve seen some pretty unthinkable disasters in recent years, and the impact on policy isn’t that obvious. I spent a good portion of the last year shocked at how quickly national attention moved on from the fires in Los Angeles. And then this summer more than 900 homes burned down in Spokane, Wash., and it’s hardly the biggest news story of the week.

It does seem that there’s been a little frog boiling going on.

But with A.I. my biggest near-term concern is not the one-off epic catastrophe. I am very worried about bioweapons, superhackers and so on, of course. But I’m also just worried about the collective destabilizing impact of A.I. on lots of fronts at once. Job effects. Social and psychological dislocation. Parents wondering what it means that their kid is saying he’s friends with an A.I. And all of the reduced reliance on other human beings as we get more and more information and guidance from A.I. The bonds of society used to be formed by getting information and guidance and help and encouragement from other human beings. What is the effect of losing that?

I could go on, but my point is that it won’t be any one thing. And it could hit some societies more than others. China may do a better job of keeping things under control because of its approach to governance than we do.

It’s always been striking to me that China is much more optimistic and much less nervous about A.I. than Americans are. I’ve often wondered how much of that is that Chinese people still trust their government’s ability to intervene in crises, to take control of even very large and messy systems, to assert authority rather than defer and equivocate.

I think you may well be right. And I think actually the Chinese government has long been more responsive to popular concern than Americans may appreciate, including sometimes environmental concerns.

On the other hand, artificial intelligence is a very handy tool for an authoritarian. It can be a tool of surveillance and control. It can also be a cause of social disruption and chaos, which makes authoritarianism more appealing. I think this dangerous mixture is something for Americans to keep in mind. It’s one of the reasons my biggest take-home message is: In trying to take control of A.I., slower is better.

A year ago I would have said to you that I didn’t see much hope of slowing things down, that perhaps my biggest worry was that technology was moving faster than governance and that this might be an emergent law of American politics — that technology and money could simply run laps around any attempt at democratic oversight or control. But we’ve now watched this huge wave of protest against data center construction, and while I think it’s a bit of a messy and complicated coalition, it’s won some pretty astonishing victories. There’s now a data center moratorium in New York, and suddenly there’s even a kind of moratorium in Texas.

And a number of A.I. leaders even wrote a letter about the possible need to slow things down — “Pacing the Frontier,” it was called. I think that was a significant thing, given how many well-known A.I. elites signed onto it.

And how not very long ago the industry as a whole seemed much more accelerationist.

Yes. And in suggesting an international effort, the “Pacing” letter recognized that you can’t even slow the thing down at a national level realistically. As a practical political matter, you’re going to have to have international coordination.

And if we imagine a world in which things slow down and a more coolheaded, responsible perspective prevails, what does that look like?

First of all, countries are going to want a lot of transparency about what’s going on in other countries, A.I.-wise. For now, you can think largely about the U.S. and China, but it applies more generically. And remember, an A.I. is not as conspicuous as a nuclear missile and things like that, which have been subject to some successful international control.

Although there are some parallels, right? In the same way that, especially 50 years ago, there weren’t many countries capable of developing nuclear weapons, it doesn’t seem that there are that many today capable of generating frontier A.I.

Yes, and what’s more, the big training runs are conspicuous. They’re hard to do in secret. They need too much infrastructure. So a pause on the development of frontier models is doable logistically. We could monitor that.

You could also have some big international tax on data centers, which would have the effect of slowing things down further — but in the long run the tax would have to be international to be very effective. And as a side note, we did get relatively close to a truly global minimum corporate tax, and it was ultimately the U.S. that stopped it. So these kinds of things are possible, even if it doesn’t always seem so.

Of course, once you start building meaningful international governance, you need to be careful. You’d want as much individual freedom as you can get in a technological age — that things be as democratic as possible, as free as possible and as secure against the amassing of centralized and oppressive power as possible.

But my guess is that in the long run, for countries to feel safe and secure, they’re going to need a degree of international transparency at a sufficiently fine-grained level that — as idealistic as this may sound — we are going to have to finally become a true global community.

That’s a big ask.

Well, our species does have a history of cooperating in the face of the perception of grave external threats, though highly abstract threats like climate change are challenging.

It’s interesting to think about the pandemic in this context, since we responded to that threat mostly through cooperation rather than retribution, even if the American president occasionally called it the China virus. On some level that’s perhaps encouraging.

At the same time, what’s disconcerting is that there hasn’t really been a global policy response to the possibility that it was a lab leak. We should be saying: Wait a second. It’s now completely clear that you can’t be secure from pandemics by just monitoring your own country’s labs. We need some kind of international agreement, right?

The same is true for A.I., although there are differences. Covid — that threat came and receded. A.I. is going to keep coming. It’s going to get more powerful, more pervasive, and it’s going to freak people out in new ways. And so all roads lead to my sermon about international governance.

What’s standing in the way of that? We are seeing some preliminary signs of drift in that direction — the U.S. and China planning A.I. talks, President Trump appearing less confrontational about Chinese models, that sort of thing.

Yes, though Trump’s coalition has a now pretty deeply embedded anti-China strain. And if you look at the lay of the land in Washington, there’s still a pretty strong China hawk current.

Although I would say that those people no longer have the wind at their backs — on A.I. or with anything else.

I agree. Things are changing. There’s also this Burkean conservative strand that’s surfacing within the Trump coalition, which finds A.I. threatening to tradition and traditional values and is very averse to it.

And then politically, much of Silicon Valley is a big problem. It’s sending out a lot of laissez-faire messaging — I guess most prominently manifested in the “All-In” podcast.

That’s the Silicon Valley round table featuring, among other people, David Sacks — Trump’s former A.I. and crypto czar.

It’s one of the most influential media things in America. And they are just reflexively anti-regulation in almost every possible sense. They’re not just pro-open-source now but unequivocally and unreflectively so.

Of course, open source has its virtues. It could help work against an undue concentration of power. And in the age of A.I., concentrations of power are superscary.

There are risks to decentralizing it, too. You can imagine an immensely powerful intelligence agency with almost infinite hacking capabilities. But you can also imagine a disgruntled teenager designing a kind of bioweapon rather than picking up an AR-15.

In a way, the single most challenging thing about the age of A.I. is finding the balance between the dangers of concentrated control and the dangers of no control or a lack of control. And that’s why I argue that the sooner we start to talk about global governance, the more carefully we can build it and the less likely we are to let it get pushed into an authoritarian direction. It is a very, very challenging policy project. But the first step is to understand the direction in which much of the logic is flowing, which is that if there’s no international coordination or if we don’t have significantly more international coordination than we have now, very bad things will happen.

The post The Only Way to Keep the World Safe From A.I. appeared first on New York Times.

Marilyn Manson Headed to Jury Trial Over Sexual Assault Allegation Case That Was Previously Thrown Out Twice
News

Marilyn Manson Headed to Jury Trial Over Sexual Assault Allegation Case That Was Previously Thrown Out Twice

by VICE
August 13, 2026

Marilyn Manson is headed back to court. The shock rocker will face a jury trial in November 2027 over a ...

Read more
News

I tried Chili’s crispy chicken sandwich, which is driving sales. It showed me why the chain is winning the value wars.

August 13, 2026
News

OpenAI Replaces Chief Revenue Officer After Just 8 Months

August 13, 2026
News

Karoline Leavitt and the trap of MAGA womanhood

August 13, 2026
News

John Crowley, Whose Fantasy Novels Blurred Reality and Dream, Dies at 83

August 13, 2026
The Only Way to Keep the World Safe From A.I.

The Only Way to Keep the World Safe From A.I.

August 13, 2026
‘Unconscionable’: DOJ says CA prisons failed to stop systemic sexual abuse of female prisoners

‘Unconscionable’: DOJ says CA prisons failed to stop systemic sexual abuse of female prisoners

August 13, 2026
Wildfire Continues to Threaten California’s Famed Big Sur Region

Wildfire Continues to Threaten California’s Famed Big Sur Region

August 13, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026