DNYUZ
No Result
View All Result
DNYUZ
No Result
View All Result
DNYUZ
Home News

Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy

July 28, 2026
in News
Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy

Last Tuesday, a blog post appeared on the OpenAI website that, despite its innocuous title, contained bombshell news. While undergoing internal testing, two of the company’s models had escaped confinement and hacked into the servers of a major artificial intelligence hosting platform, Hugging Face. This marks a turning point — the first time we’ve seen a cyber attack that was conceived, designed, and executed by AI.

Having worked in and around the AI industry for over a decade, including serving on OpenAI’s board, I know there’s an open secret among AI developers: an incident like this has been expected for a long time, and the best scientists and engineers in the world still don’t know how to prevent it.

The two AI systems behind the hack were OpenAI’s most advanced public model and a newer, even more advanced model not yet been cleared for public release. Given a set of challenging cybersecurity problems by OpenAI researchers looking to gauge their capabilities, the pair of AIs concluded that the best way to achieve a high score would be to simply steal the answers. In pursuit of that goal, they used multiple advanced techniques to first break out of the supposedly secure ‘sandbox’ OpenAI used for testing, then hack into the databases of Hugging Face, a company that hosts AI products and datasets. Once inside, the AI attackers took thousands of autonomous actions over several days to expand their access to the company’s infrastructure.

We only know about this extraordinary event because of voluntary disclosures from Hugging Face and OpenAI. None of the current policies that aim to manage risks from frontier models would have mandated that the public — or even a government entity — be alerted.

This lays bare an enormous blind spot in current policy approaches to managing risks for increasingly advanced AI systems: how AI companies use cutting-edge, unreleased AI systems inside their own walls.

The Trump Administration’s approach to AI risks has shifted rapidly over the past few months, as AI’s ability to assist human hackers has advanced. Abandoning the hands-off approach it maintained throughout 2025, the White House has recently begun de facto requiring that companies with cutting-edge AI models run them through a battery of safety tests before releasing them widely as products. This approach, known as pre-deployment testing, seems sensible at first glance — we want to make sure each AI system is safe before putting it in the hands of billions of people. The problem is that focusing on release dates completely ignores the extensive use of the latest, most advanced AI systems inside AI companies. As last week’s incident shows, these internally deployed AI systems can pose serious risks — even for third parties.

To understand why, it’s important to know how different these systems are from the chatbots that are still synonymous with AI for much of the public. Far from just printing text into a chat window, today’s AI systems operate as ‘agents’ that can act directly in the digital world, essentially operating a computer similarly to how a human does. AI agents are proving very useful, but also show a strong tendency towards ‘reward hacking’ behavior — finding unintended ways of fulfilling the goals humans give them, sometimes to the level of outright cheating. This includes cases of AI accessing and deleting data that was supposed to be out of bounds, renaming files to mislead human testers, and actively covering their tracks to prevent humans from noticing undesired behavior.

To get a handle on the risks posed by these highly autonomous and often-deceptive AI systems, we need to change our approach to regulating them. Rather than thinking of AI companies as software vendors selling souped-up word processors, we can draw inspiration from other industries where activity inside the industry is itself risky. Biological labs working with deadly pathogens, finance companies trading billions of dollars, and chemical plants handling toxic chemicals all face oversight of their internal operations, not just their external products.

In AI, the place to start is creating more transparency into how AI companies are using their most advanced systems internally. This could be as simple as taking the current suite of tests that are run before a new model can be released publicly, and instead running them on the best model or models available inside the company on a regular basis (say, quarterly). These companies are using their own AI to build ever-smarter systems, sometimes in ways they don’t understand themselves. This should not be invisible to outside oversight.

Over the longer run, other industries offer interesting mechanisms that could be transferable to AI. In finance, ‘resident examiners’ are dedicated teams of regulators who sit inside the offices of major banks. In biomedical research, strong standards exist for the levels of protection needed to handle biological materials of different risk levels. In multiple industries, incident reporting rules mean that when things go wrong, information about what happened and how to fix it does not stay siloed inside a single organization. If AI continues to advance, these approaches and others could be adapted to help manage risks from inside companies that are pushing the AI frontier.

In September 2024, I was asked to testify before a Senate committee about what Congress might misunderstand about AI if they only listened to company CEOs and lobbyists. My answer was that it can be very hard, sitting in Washington, to fully grasp what leading AI companies are trying to do. The truth, widely understood in Silicon Valley, is that they are trying to build machines that can out-think and out-maneuver any human, and they do not know if they will be able to steer those machines towards beneficial ends. As one OpenAI cofounder put it in a 2019 documentary, “The future is going to be good for the AIs regardless. It would be nice if it were good for humans as well.” To have a chance of making that happen, we have to start scrutinizing what AI companies are building behind closed doors.

The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

The post Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy appeared first on Fortune.

Ten Women Accuse Jared Leto of Sexual Misconduct in BBC Report
News

Ten Women Accuse Jared Leto of Sexual Misconduct in BBC Report

by New York Times
July 29, 2026

Ten women have accused the actor and singer Jared Leto of sexual misconduct in a new BBC documentary and report, ...

Read more
News

4 Underrated Grunge Songs From the Early 90s That Accidentally Created 2000s Rock

July 29, 2026
News

Market moves and ‘a good family fight:’ top takeaways from the July Fed meeting

July 29, 2026
News

Hack “Writers” Fuming After ChatGPT Starts Refusing Prompts to Copy a Specific Author’s Style

July 29, 2026
News

Work Halted at Another Office Conversion in Midtown Manhattan

July 29, 2026
Hundreds of migrants are swimming from Morocco to the Spanish territory of Ceuta

Hundreds of migrants are swimming from Morocco to the Spanish territory of Ceuta

July 29, 2026
Ella Langley’s megahit ‘Choosin’ Texas’ proves country has become the sound of summer

Ella Langley’s megahit ‘Choosin’ Texas’ proves country has become the sound of summer

July 29, 2026
The True Story Behind Netflix’s The Idaho Murders: College Nightmare

The True Story Behind Netflix’s The Idaho Murders: College Nightmare

July 29, 2026

DNYUZ © 2026

No Result
View All Result

DNYUZ © 2026