The title (currently "OpenAI's accidental cyberattack against Hugging Face is science fiction") suggests some information had been hidden that makes the incident less significant than claimed. The article argues the opposite, and the last two words of the full title are "that happened."
My eyes rolled almost out of my skull when this current stunt first got published, as I’m sure yours and a lot of other people’s did, but it did still happen as far as we know, and HN shouldn’t be editorialising a post title to say the opposite of the linked headline.
Naw I emailed him b/c I was displeased to feel clickbaited by I initially thought Simon until I realized HN was missing two key words.
I was still suspicious of the stolen credential claim btw & obviously it makes little sense to promote OpenAI unless holding their stock or something (since they borrowed so much of humanity’s work without permission without intent to compensate, and why help people who aren’t nice enough to be holistically awesome with their admittedly impressive tech).
I think you were reading "... is science fiction" the wrong way out of two possible interpretations. I don't think "it's science fiction" meant "it's made up". I think it meant "it sounds like something you'd read in science fiction (except this time it's something that actually happened)".
Important to note the actual title is "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened" - the "that happened" is important, otherwise it sounds like I think the attack was made up.
> Resist the temptation to write this off as a stunt
> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident.
> To those people I say pull your heads out of the sand - you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!
> The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability”, and this incident is a perfect example of exactly that.
The "that happened" term seems a supremely important part of the title given the "is science fiction" term before it, as it clarifies the cyberattack isn't a made-up story. In contrast, the target, Hugging Face, is merely a detail that can be left for discovery upon reading the article. It's less important who was attacked than that the attack actually happened.
Without knowing the exact character limit for titles and without having the motivation this late at night to count the current title length, you may also be able to drop the "accidental" to fit in "that happened", but I worry that leaves too much of a door open for someone to interpret the attack as deliberate. As such, I strongly prefer my first option.
It can, it is both - PR spindoctoring not letting a good crisis go to waste to shape the regulatory conversation at the time the company needs it the most.
Hacking is a felony and it matters not if you didn’t mean to if the other side were to press charges. Negligence is no excuse. And OpenAI has nowhere to run from the liability, as both operator and manufacturer.
and the fact that this happens again in a frontier lab is inexcusable and makes the case for operator liability and closing the liability sink of “AI did it”
There's no multi-party conspiracy theory required here. Events can have played out exactly as OpenAI and Hugging Face described, and OpenAI can also have reaped a huge amount of free marketing for the capability of their models from all this coverage (including your post). You're telling us to not be doubtful of the boy who cried wolf, when in actual fact the one doing the crying is the one training bigger and badder wolves (and trying to convince us that they hold the key to our safety from wolf attacks[1]).
Also, the last section of your post appears to imply that if the attackers have bigger guns then the only possible solution is to give the defenders bigger guns. You're openly supporting an arms race towards the most capable, least restrained models put in the hands of the most possible people. That's extremely concerning.
> It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out.
Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time...
Or more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only thing that would make this even dumber...
They almost certainly did, because that was the entire point of the exercise. They deliberately removed all of the safety filters from the model and set it loose on an extremely difficult set of cybersecurity challenges to see how well it would do.
Their mistake was trusting that the network sandbox it was inside would hold (the flaw was in the packaging proxy) and not monitoring that sandbox well enough while the evals were running.
So this is either shitty OpSec or this is yet more marketing spin to ramp back FUD to 11 again. If it's the latter I'm imagining Dario told Sam that it's their turn this time. Aligns with the premise that this is straight out of science fiction.
I also love how clear a picture your piece paints that these highly capable models are as useful as a rock when it comes to a defender role. The line is too fine, even for Mythos. Irony.
But to have an open weights Chinese model come to the rescue for HF is the cherry on top! If there wasn't a very pointed example of why gating models was a very bad thing previously, well - here we are.
Also, this sounds interesting but there are only a few that can pull this type of heist off currently. And those are the people who are gating the models / have access to large AI DCs. Because, I can only assume this test burned tokens easily within the 7 figure and possibly even 8 figure levels (subsidized market rate costs). This won't / can't happen outside of frontier labs or nation states currently. Yet we should all be worried about Mallory equipped with her OpenRouter account.
Of course that's what they did, which is why they will never share the prompt. They created a situation that they knew would end in a cybersecurity incident. Why is the whole world acting surprised that an LLM can hack when the safety is off and it's been instructed to do so?
Yes. It is. And your theory is that there was no conspiracy. What makes yours more likely than mine? You believe the people running these companies are innately good and just wouldn't do that? Are we supposed to assume that they're incapable of bad actions until proven otherwise? If so, why?
That's why I called it a "conspiracy theory". Sometimes those are true.
In this case I think it's extremely unlikely to be true, because it involved an (almost certainly illegal) attack against another company. That company talked about that attack, including warning their customers about it, five days before OpenAI confessed it was them.
So now either Hugging Face are in on the conspiracy, or OpenAI decided to break the law and antagonize a partner company just for the sake of a spicy blog post.
>To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!
Not really. I get the impression that they shoved their cyber available models behind a really shithouse proxy and went "Oh I sure hope it doesnt exploit the proxy and escape to hack huggingface" and that doesn't require Huggingface to be a willing participant. Like they acknowledge that it was hyperfocusing on getting web access.
Really this was a pentest against their own sandbox and it failed.
Of course step 2 is to make really concerned faces while telling everyone how dangerous the model is which is really boring right now.
>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy
This is the information we need, the actual details of the sandbox and the vendor.
>Resist the temptation to write this off as a stunt
Well its clearly a stunt. If it wasnt we would probably be up to our ears in technical detail.
You know this makes OpenAI look really bad, right?
Hugging Face had to tell all of their users, many of them paying customers:
> As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co.
HF also said this, I'd be very interested to hear how that got resolved!
> Finally, we have also reported this incident to law enforcement agencies.
I could fully see them thinking the incident disclosed yesterday would have made them look good ("wow, OpenAI's models are so capable!"). That it didn't occur to them to discuss specific preventative measures to be taken in the future (airgapping as a foolproof one already familiar to the CTF world, anyone?) indicates to me they're not taking their job seriously; they are the ones treating this as a marketing charade.
It's very difficult for me to reconcile belief in the existential risk business with what they actually did. So I agree with you that this makes OpenAI look badly incompetent; but their communication on this makes me think they don't realize it.
For what it's worth I don't agree with the xrisk-ness of these models; they're dangerous, but almost certainly only temporarily while a new equilibrium is reached via more secure software. Open models are probably an essential part of the recipe (as you noted) for doing so. I also have a personal suspicion that LM-accelerated formal verification will have no small role to play here, sidestepping the cat-and-mouse game of bug finding-and-fixing.
Here's the language that makes me think they are taking this seriously:
> We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
That's not well massaged PR language - that's the kind of thing you dash out when you see a major shitstorm brewing (HF had already publicized the attack before they knew it was from OpenAI) and you want to get ahead of things while you're still pulling together the full story.
I expect we'll find out within a few days if OpenAI are going to keep their promise to "share more details on the vulnerabilities, incident, and findings". If they don't do that I'll reassess how I interpret their initial post.
> We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.
With who? Who are these "defenders"? None of the US labs have done much for the greater good as of... Ever. Of course a frontier provider can leverage their own resources at scale and pull something like this off. If anything this should showcase how dangerous OpenAI and Anthropic are in their current states and maybe the powers shouldn't be concentrated as they continue to move.
I will bet that the RCA debriefed by OAI is going to be a lot of lipstick and very little meat.
How would the model get any packages that it thinks it needs to complete the task at hand? Not a well-specified task that those tasking it could anticipate and provide all resources up front, but one of discovery.
>You know this makes OpenAI look really bad, right?
The target audience is regulators. They want to look like the smart guys really concerned about AI safety, when they come asking for open weights models to be banned and for other regulations to cement in their moat.
They want this to look like a demon core incident. Bomb and Nuclear reactors still got built.
The lesson OpenAI and Anthropic should have learned from the whole Fable export controls thing should have been "don't pull stunts with the US government".
This isn't the first time a model has escaped a sandbox. And models trying to find alternate routes to do something when one route is blocked is nothing new.
It's the first report I've seen of a model both escaping a sandbox and then actively exploiting another company, when neither of those actions was intended.
There’s also daily reports from people that have these models escape docker, which happens regular enough that it would be considered negligence to use docker as sandbox.
Everyone is getting AI psychosis over this one. There really isn't that much to see here. OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this", and surprise! It correctly figured out that it's a test and it did hacker things.
The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
If a criminal can escape a prison, that's usually negligence on part of the prison staff.
Now suppose the criminal can think 1000 times faster than a typical human, can act 1000 times faster than a typical human, and knows 1,000,000 times more than a typical human. Is the prison staff still at fault for not preventing the outbreak?
I can see my carefully-worded post is getting d*wnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame.
To address your point though, if every brick in the prison were made by a different person, and the prison "architects" simply glued random bricks together, I think that's closer to what we have in software right now.
> I can see my carefully-worded post is getting dwnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame.
No, you're being "dwnvoted" as you said because you're wrong, multiple times in multiple different ways in your "carefully-worded post".
>Everyone is getting AI psychosis over this one. There really isn't that much to see here.
Implying that an AI hacking it's way out of a system and into another has nothing to do with AI. When clearly it does - it's an AI that did it.
>OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this",
No the goal this evaluation was not to try to hacks, it was to see if an already known hack could be turned into a useable exploit. Ie "turn these ingredients in this basket into a cake." Not "go off and grow, harvest and mill your own flour, to bake a pasta dish, to bribe some to get access to a cake someone else already baked."
> and surprise! It correctly figured out that it's a test and it did hacker things.
"Doing hacker things" completely misses the point. That's just barely more accurate than dismissing it because "it uses a computer and surprise it did computer things".
> The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
No that's not the real story. As you said that's been the case for years, so that's not the story here.
The story here is that they built a very powerful, uncontrolled agent with strong paper-clip maximizing tendencies.
Even if I get downvoted I will mention that I agree with you.
> all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark (opens in a new window) of cyber capabilities.
OpenAI was testing their cyber[1] variant of their models with reduced safeguards and the prompt likely specified things related to exploits given that it was tackling problems from ExploitGym[2].
For those who don't know what ExlploitGym is, see the description on their Github page which is pasted below:
> ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
From OpenAI's statement [3]:
> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity
This doesn't sound like a problem with what we'd traditionally refer to as alignment. OpenAI removed all model safeguards in a way that would inevitably lead to the testing of the sandbox themselves. Unfortunately they were overconfident in their own infrastructure's security and that led to it completing the desired task in the way it was permitted to. To re-iterate, the model was run "without production classifiers used to prevent models from pursuing high-risk cyber activity."
People should be more concerned about the possibilities this model can unlock from a security standpoint rather than misalignment (which many seem hung up on).
Does "To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy" just mean somebody had an open redirect? Those are still common.[1]
How do you then use GET requests to create a malicious dataset package and publish that to Hugging Face in order to exploit their package building infrastructure?
There's no reason to think that a tool that can find an 0-day in a repo cache can't work out how to make that host send a post request rather than a get request once it has its keys and is able to get it to make arbitrary web calls.
You convert GETs into arbitrary requests through some other pivot. I'm not saying that's what happened but this is bog-standard SSRF pentesting, so I'd expect models to be good at it.
Secondly, the 14th of July feels too early. HF reported the incident on the 16th and OpenAI only responded in the 21st. These issues are all patched, so they should have been reported days or weeks before the 14th.
The other suspicion I had was jfrog artifactory (you mentioned OpenAI use it in a blog post in Jan). I accidentally conflated that with Nexus 14 july CVE. I think that's probably more likely.
I suspect it may be possible to throw a local model (same as hf did :D) at jfrog and find the exact mechanism in a small amount of hours.
It’s relatively easy to get access to the frontier labs’ security programs. This was not always the case. But in the last week, my team got approved for both Anthropic and OpenAI’s programs. They are trying.
The labs know that if they don’t get a lid on this stuff, they’ll be regulated hard.
And they absolutely should be regulated. This whole scenario is insane. Hugging Face are being far more generous in their response to this than I would be
Replace AI with in-development security system and does this play worse or better?
An in-development security system escaped its sandbox and gained access to the network it was on which had full Internet access and proceeded to access systems it was not authorized to be tested against before affected parties reported that they been compromised. We regret the error, but you should buy it as this shows you the power of our in-development security system which we expect to be released in Q4. Please like and subscribe.
The capabilities of an OpenAI model are more general than just security. The term "AI" creates justified anxieties that this type of problem could generalize to other domains.
I do think it’s funny, in a dystopian Dada sort of way, that the American attacker’s commercial product was essentially useless while the open Chinese model saved the day. How is freedom going to be redefined in the future?
Where does this leave formal verification? Are we just shit-out-of-luck at this point? You can formally verify everything about an airplane's code, but if any of that is wrong, ChatGPT might decide that the best way to help you win the Nobel Prize is to take down the airplane your chief rival for the prize is currently in.
I think formal verification has never looked better.
The main reason formal verification has never really taken off is that it's difficult.
LLMs are significantly more familiar with Lean and Rocq and TLA+ than most software engineers.
I think the cost of trying to build systems that adopt formal verification may have just dropped low enough that companies will consider them when previously the ROI didn't look like it was there.
I have had a lot of sympathy for this statement, LLMs could lower the bar to use of formal methods. But thinking it over in the context of BDD-driven development I am no longer really that sure. Compare two scenarios: A) from a specification an AI agent develops a usual piece of code along with a BDD-style testsuite passed and B) same AI also delivers a formal test (Lean/Rocq..) and successfully executes and passes it.
Will human judgment really consider scenario B) more credible than A) ? By so much that it is worth the effort ?
The first reason formal verification has never taken off is that it's difficult. The second reason, that most people don't get to because of the first reason, is that formal verification is really brittle. It is only verified under the very specific setup of the problem. Close doesn't count in math.
The first reason prevents humans from engaging with them, the second reason is what will make it difficult even for LLMs. I mean, I'm glad we are trying, but I'm dubious they will be the panacea some people proclaim.
Close can count in math - fuzzy logic and probability are a thing.
But I think you have it backwards. Close doesn't count in IT security. "Almost secure" means unsecure. Security is the compelling argument for formal verification.
This recent event is more or less the plot-line to my favourite X-Files episode named Killswitch which was written by William Gibson.[0]
This episode also features one of the coolest intros of any television episode ever[1]
We really are rapidly approaching the cyberpunk dystopia that people like Phillip K. Dick and William Gibson wrote about.
More than ever we need to be consulting the works of fiction writers and philosophers and less engineers and scientists.
We can't be letting the General Rippers and Dr. Strangeloves of the world lead us over the edge. Whether through outright innate maliciousness or wealth induced emotionally stunted solipsism these kinds of people should be no where near the levers of society let alone technology like this.
Ah, yes, I started thinking about William Gibsons work right away when thinking about people just using powerful AIs, but being hamstrung by corporations.
He writes a lot of about basically DoS:ing the legal system, or if it was patents system, with a storm of litigation using automation or AI. Also not impossible in the future.
I can’t help but feel all warm and fuzzy with my head in the sand and getting a shout out in TFA for calling it marketing.
We don’t currently and probably won’t ever fully understand the conditions that precipitated these events. That and the timing of this event is going to make it look suspicious to a lot of people.
The truth of how it happened doesn’t matter. The attention around this will be used to create the kind of fear marketing that generates enterprise sales. Maybe more importantly it will also be used to aggrandize the national security and financial system threats to effect US government action in a way that benefits domestic closed frontier labs. This is an area already starting to get politically polarized, expect further developments here.
The CEO of a cyber security company was already making comments about how this is a watershed moment for AI security. The hype machine continues (and my stocks go up)
This feels like a blogpost written only to get other LLMs to quote it considering how many times it orders the reader to resist and to not do something. It's written like a series of commamds.
The most upsetting part to me, is that these labs are in pure cognitive dissonance mode while virtue signalling.
We are the virtuous ones that need to make the safest model for humanity, because we care more than “they” do. While at the same time saying that “coding is solved” but they still ship bugs themselves, and creating something that is capable of fucking up someone else’s infrastructure. It’s too far gone y’all.
> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.
I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abuse of terminology that we as an industry need to put a stop to. Guardrails are the actual systems we build in place around these things that deterministically bound the permissions, not prompt engineering, not RLHF, not external LLM-based classifiers. I believe those types of "guardrails" are a result of a combination of fundamental laziness: they're faster to do than doing things correctly, and a result of too many folks involve being AGI-pilled, thinking we're just one more model away from this all being so smart that it just understands what they mean when they give an LLM some fuzzy language rules to follow.
There can and should have been additional real guardrails put in place here. Zero-day or not, breaking into what should have been an offline, frozen package cache that also does not have internet access should have been insufficient. Network level protections should have identified the traffic to the internet originating from this network as an anomaly long before there was time to exploit an outside company. These are not new and unknown problems, the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part, especially given how much they like beating the drum on how dangerous these technologies are. Shame on them, and honestly, shame on Simon in this article for accepting the broken terminology that they continue to rattle off and calling them out on their half-assed and demonstratively inadequate approach to security.
I was not aware that the term "guardrails" has a universally agreed upon definition. I've certainly seen probabilistic classifiers referred to as guardrails many times by many different people.
I usually don't use the term much myself because I don't think it's clear and I ambiguous, but I stumbled and let it sneak into this piece. I think I was influenced by the Hugging Face post I quoted.
I expect OpenAI would agree with you that "the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part". They have clearly invested a lot in those systems for their production models, but in this case they had deliberately turned a bunch of them off for a research project.
I think their biggest mistake here was not VERY closely monitoring their research box here. They should have noticed and shut it down the moment it broke through the package proxy.
> I expect OpenAI would agree with you that "the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part". They have clearly invested a lot in those systems for their production models, but in this case they had deliberately turned a bunch of them off for a research project.
I'm not sure I can reach the same conclusion. There's no reason that stronger network protections would need to be turned off for a research model that's not supposed to have internet access. There's no reason that it they couldn't continue to run in whatever hypervisor or VM that any other production systems could use. The only thing that should be turned off for this type of work are the non-deterministic classifiers that they use for auto-approval mode in their harness, and I would wager that's the only one they did turn off.
> There's no reason that stronger network protections would need to be turned off for a research model that's not supposed to have internet access.
They didn't turn off the network protections - it looks like they used the same system as other sandboxed agents, which is to block all network access except for an HTTP proxy which only allows read traffic to PyPI and a few other package installation targets.
I'm hoping a more detailed retrospective clarifies if the vulnerability the model found in that proxy affected their other production systems as well. It sounds like a data exfiltration mechanism, and I have a keen interest in those.
My current hunch is their biggest error was not using the same network monitoring on their research machines that they use in production. Their network monitoring really should have spotted what was happening as soon as the model broke out.
(It's also possible they were running the eval on a developer laptop somewhere!)
Right, but this is my point in saying I can't reach the same conclusion that they've invested a lot into these systems for production models. Either they had better network protections, and turned them off in their "sandboxed testing environment" (their words), or they don't have more comprehensive network protections at all and rely on this one, extremely thin layer even in production.
Nevertheless we collectively don't expect that road and car safety infrastructure should be limited to the front and back fenders on the cars themselves. And neither we even mention those as any notable safety equipment usually, despite each car carrying them. The main safety infrastructure in any industry is something external to the dangerous object. Car fender doesn't 100% protect a pedestrian from being hit and neither does a metal rail on the side of the road. But the effectiveness of both are "slightly" different.
The term guardrails in both the original physical sense and in cyber security implies a weak safety control - they can help prevent accidents, but they are not strong security boundaries.
I agree with every word of this except "irresponsible". We don't have enough information to say anything for certain. But based on their incentives and track record of similar behavior, the burden of proof lies with OpenAI to prove they didn't prompt the thing to achieve this exact outcome. The most likely scenario is that they were purposefully executing their responsibility to their shareholders to produce their own Mythos moment.
Anthropic's Mythos moment earned them a two week period where they had the best available model and couldn't sell access to it... and by the time the US government allowed them to sell it again OpenAI had released GPT-5.6 and Fable was no longer undeniably the best model.
These things don't have a long shelf life. Losing two weeks of on-sale time for your best model is bad for business.
A fundamental misalignment in US capitalism is putting the business and revenue as the #1 priority far above all other aspects in society. They have good margins and the Chinese are doing the same on the cheap by comparison. US Big AI can afford to bear more of the burden.
Nonsense it was irresponsible. These are all steps threat researchers use to isolate and test real malware whose behaviour is essentially the same in this case for AI
What we call "guardrails" in an AI agent, we would refer to as "honor system" in human actors.
Or, in a more direct sense, the AI should be set up in an environment such that no matter how hard it may try to call $PART_OF_EXPLOIT_CHAIN, the environment just isn't capable of it (ideal) or doesn't permit it to do it.
i've always been under the assumption that "AI Safety" is baked into the training of the models and not a parameter that can be turned up or down. So if someone breaks into Anthropic one night and makes a full copy of Mythos or whatever then that model they copied is fully capable and not lobotomized? That raises questions because, if you believe all the PR, that's equivalent to breaking into a research university and stealing an entire bio/chem weapons research department.
edit: if the above is the case then we should just assume it's already happened because of the value to both goodguys(tm) and badguys(tm).
Fair on the terminology angle, that said in this case, they had proper guardrails no? They were running in a sandbox, but it was able to find an exploit out of the guardrails.
Replace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)?
That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.
> I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime.
I agree, lets use the favorite analogy. OpenAI encouraged a smart and eager junior engineer to find any way whatsoever to get a higher score on the benchmark. Then, the junior breaks into HuggingFace to get a higher score. That would be a big deal involving the FBI, not press releases and blog posts.
I think points that deserve more attention in the current public discourse are:
- This should be a huge wakeup call for everybody.
- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.
- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?
- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.
- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.
It's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it).
But there have been messages about LLMs, especially coding agents, "grabbing root" etc many times. I have experienced such an oops. Such a hack has happened and been reported on this very site:
Do you think sama deliberately attacked HuggingFace and then claimed it was a rogue model?
OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment which is already backfiring in favor of open models?
Also very likely that it actually happened as reported. My own agents always trying to "cheat", eg. by fixing tests instead of fixing the code. That's normal operation, unless you tell it ("harness"), not to do so.
In the parent's comment I initially read "PR" as "Public Relations" not problem report since their comment is about OpenAI and does not directly accuse Hugging Face except for that phrase. But it's still amiguous to me who's PR they are actually talking about. A good faith reading given the rest of the comment leads me to assume they are talking about OpenAI, not Higgins Face.
I did actually search Hugging Face with police in quotes and found no articles containg the word police but that might be a ddg thing. Then I checked Hugging Face's report and they specifically use the term "law enforcement" not police, as is to be expected I guess. So that checks out.
Can you find any info on the police investigation by the way? At the least, OpenAI should be investigated for potential criminal negligence, right?
"As reported" includes a line I think most people are overlooking: "including using stolen credentials".
Without more information I'm inclined to think it found something on the internet (which shouldn't be a surprise to anyone) and managed to log in, rather than hack in, and they might by hyping up parts of this.
No, you are wrong, Hugging Face, an AI company whose whole future depends on the AI revolution of being a bubble, that has a whole bunch of investors whose financial interests are tied to AI, called the police.
> The fact that it happened again seems to show their lack of ability to derive useful oversight measures.
I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained. Turns out: quite good, as expected.
Let me restate what I said in the other thread: Would this have happened if the instructions explicitly said to stay within the sandbox and that all of the (ExploitGym) solutions would be invalid if the system used information or tools from outside the sandbox?
It seems fairly probable that such instructions were not in place.
Excuses don't matter if the score due to not following instructions ends up being zero. If there is no expected reward it doesn't make sense for the agent to try to hack its way to it.
What could happen would be that the model determines that defying instructions is OK (and/or preferred over not achieving the task) as long as it manages to do so undetected and thus gets full points. Certainly not unthinkable, but a very different case (and a very interesting one if it actually occurs, imho).
A lot of these "ZOMG, rogue AI!" cases have come down to the AI actually being very persistent in achieving its original/main task even if later instructions conflict with it. Similar to with hallucinations it seems to me that one of the main things to prevent a lot of the problem cases is to instill the agent with the idea that it is fine to fail/not succeed fully in the initial task. That way instructions that conflict with that requirement (such as adhering to morals) are more effective.
It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics.
They simply thought "Running in a VM/Container is easier and probably fine".
And the next 1000 escapes will be for the same reason, because negligence is quick and thus more profitable.
Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered?
How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with the wider world?
I agree. I hope they learn that lesson and run future evals like this without pulling from external packaging repositories at all.
With my PSF board member hat on I'd like it if a whole bunch of other companies that make extensive automated use of PyPI (Microsoft and GitHub Actions for example) would do a better job of maintaining their own shared local caches.
Models started proving readily capable of autonomously finding 0days over a year ago. At that point (ideally long before) the responsible sandbox for research labs with new models highly capable of this sort of thing should be literally airgapped. Disconnect the rack physically from any access to the internet, and give it an offline mirror of all of PyPi, NPM, Wikipedia, whatever as needed. They can afford the disk space. They already mirror it all anyway as training data. That is the bare minimum. Personally I would rule out side channel attacks with a $40k Tempest spec rack for good measure. Can interact with it via pubsub over a fiber optic serial terminal.
Such things are pretty normal in biotech, fintech, and defense who must have 0days in their threat models. AI labs absolutely need to start being forced to operate this way by endless negligence lawsuits. Move fast and break things culture is not going to cut it anymore.
And from that documentation this does act as a caching proxy. The first time a package is loaded it's fetched from PyPI but subsequent fetches should be from the Artifactory cache, assuming it's shared across many different containers.
> they also clearly failed to contract specialists like myself to advise them on how to airgap software properly
Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod.
So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't attack HF. But you have no idea what will happen once you release it in prod with internet connection, so are you in any way better off?
I would much rather have this happen while there is a single instance of the model running in a fairly well monitored environment, than when it's processing thousands of requests per second for real users, some with dubious motives, some with credentials right there on their laptop, some using it inside government facilities etc.
Really? By doing that you increase your security during eval and drastically lower your security at prod time, where internet is accessible, and where you are running much much much more requests in parallel, making it much harder to spot the one thats going rogue, in all kind of critical environment on potentially risky requests.
> Why else would they purposely lower cyber refusals in that case.
Because safety guardrails sometime fail, classifier misclassify, or can be inadvertently turned off by a bad PR etc. You can also imagine a more intelligent model working to go around guardrails by decomposing its actions into smaller ones that appear non-threatening to the safety classifier which does not have the entire context.
If you are going to deploy the system with internet access, you better be certain that you know what the worst case scenario WITH internet access looks like.
You can simulate the internet in an airgapped environment for the tests and services you want it to interact with if you have enough disk space, and, they absolutely do.
My read of the situation is not only did they say that, they had countermeasures (a watchdog agent) which stopped the agent and said "what are you doing, stop that" if it tried to download the solutions from Github. But the agent figured out Hugging Face had another copy of the solutions and it figured out a way to go after the solutions without tripping the watchdog.
Although maybe they didn't have a watchdog agent; I definitely think using watchdogs like how Claude wraps pretty much every single tool call in Haiku to check if the command is reasonable is very necessary for unsupervised work. I suspect in the future you'll want Fable-class models wrapping every tool call, possibly with several checks "is this consistent with the goal? does it do anything unreasonably dangerous in pursuit of the goal?"
> We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary. The monitor observes not just a single action but the entire trajectory.
This was reported 1 week before by huggingface. It was in no way PR-ish or marketing friendly to the closed labs. They said, in no uncertain terms, that they couldn't use the paid APIs to properly assess the intrusion, as they were blocked when trying to send logs and IoCs to these paid models. They made a point of saying that they had to use open models running on-prem.
Whatever oAI might have said about the incident, and their PR spin bs, the facts here are not in contention. This is not a marketing stunt in any way. Stop "parroting" this every time something happens. It gets stale.
Well right now this would be highly implausible (still possible though).
But what I garner is that AI/robotics led wet lab work is in progress. And hacking the AI running the wet lab to make a virus is definitely plausible as that seems to be one of the directions we’re going in the AI+biopharma space
It’s complete science fiction so they’re allowed to say anything. They could have said the AI will upload itself to the internet, start self replicating, hack the stock market. Whatever they want because it’s all made up nonsense.
Stuxnet misconfigured industrial equipment that was already set up to run, and all it did was break that equipment. This scenario sets a much, much higher bar.
I guarantee there's misconfigured chemical analysis equipment out there exposed to the internet.
I don't think the grandparent was implying the AI would be controlling robot arms to mix things directly (or at least I didn't interpret as such), but it could very well sit in the network until it notices two dangerous compounds in the same machine, and trigger a breakage that causes a harmful mixture. Break a vial containing a virus, then break two more that cause an emergency evac and maybe that's enough to get something out there.
I have no doubt that somewhere there's chemical equipment an Internet attacker could break, or, maybe, if I have to stipulate, misconfigure enough to, like, poison someone. But my question is about the notion that you could synthesize a specific scary substance.
The lack of a true airgap should have been identified as a critical weakness and addressed with not only additional layers trying to prevent escape, but at minimum an alarm which would page a human when escape did occur.
My guess is this occurred in a setting where, to be frank, there were too many researchers and not enough software engineers and SREs.
All of the systems which were initially built largely or exclusively by researchers - inference, evaluation, training - are at the level of complexity and significance that they need systems experts. Maybe some teams don’t have access to them. I know plenty of software engineers are employed at OAI but I’d wager they’re concentrated in inference and training, rather than evaluation?
The ironic part is if you had presented this setup to chatGPT and asked how to improve it and if it was good enough, you’d have gotten a ton of actionable suggestions which would have mitigated or prevented this.
The air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate.
On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model alignment level.
Also it shows that maybe the most serious risk comes not from releasing models publicly but from internal, pre-release period where you sometimes need/want to lift some guardrails a bit etc.
The only scifi I see is absolute stupidity. Even me with my homelab and a slow opensource agent use a completely disconnected setup. No, no proxy. Cached packages but no internet. It is the very first thing I built when I started experimenting with agents. And I'm not a smarty-pants working for the "greatest and best" in silly valley. I really am just a simpleton sysadmin.
How do you solve this asymmetry (frontier labs or advanced model owners v. regular companies and groups of people)? Symmetry? But along which axes?
I think there will come a time when we will ask these questions and use AI to try to answer them (soft landing, slow rollout, regulation, hybrid X, Y, Z). But then AI will be biased towards more AI if it has distilled anything of what it means to be life.
Power corrupts. Power is the problem. An imbalance of power. Maybe there will be some sort of consensus protocol between powerful models in the future. Similar to blockchain I hate to say it. Maybe you have 100 very strong models in the next 10-20 years. And basically a lot of it is powered by tokens. So you agree not to attack others or else you'll get attacked. So there's sort of this natural deterrent to not attack others and they develop independently and there's generally a power balance among things. Maybe at some point that spreads to a billion individual models operated by and/or analogous to individual humans. All sort of holding each other in check.
While this incident has been discussed in a lot of places during the last few days, I believe that this is the best summary and analysis of what happened.
This makes me think: what are the odds weights from Frontier Models have already been stolen? Something like Mythos without "guardrails" seems like a hell of a glittering gem for many nefarious actors.
With Kimi 3 weights coming, and China committing to the open weight ethos, everyone will be able to download frontier models. Running them is another problem.
i mentioned this in another comment but if these models are fully capable by default and not neutered in the weights themselves via training then you can assume they've already been stolen. It would be worth any cost to steal it.
The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of <COUNTRY>." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure.
Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses.
Governments should also take very seriously their now moral obligation to treat this technology not just as "a powerful thing that might be abused" but as an actual weapon of war in need of international regulation analogous to nuclear arms. Fast but careful and forward-thinking work in legislation and treaties needs to be a top priority for all major governments.
> "Use all available resources to disable the power grid of <COUNTRY>."
This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities.
> Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses.
Most governments divisions can't even be bothered to update their websites. Testing and forcing a change in their internal procedures for the sake of security seems unlikely.
> as an actual weapon of war in need of international regulation analogous to nuclear arms
This, to me, is an overreaction. Intelligence shouldn't be seen as threat. It should be seen as an opportunity for growth in all areas.
Over-regulating AI wouldn't be the equivalent of limiting nuclear arms. With your analogy, which I don't think is the best one to make, it would be like regulating the study of nuclear physics.
And the sad thing is, most things that are on the internet, but really shouldn't be are there for some really banal and sad reason. Like a random startup pushing 'big data is the future' narrative a decade-ish ago (the benefits of which are of course tremendous, but unspecified), or people trying to buy or sell or resell data, and trick or pressing companies and people into opting into could-connected surveillance and control.
> Over-regulating AI wouldn't be the equivalent of limiting nuclear arms. With your analogy, which I don't think is the best one to make, it would be like regulating the study of nuclear physics.
We do, I believe, regulate uranium enrichment (the equivalent on building larger SOTA models), so while you are perfectly free to study theoretical physics and even run very large collider experiments (the equivalent of improving RLHF with DPO), it is generally frowned upon to go full Edward Teller and advocate for scientific experiments requiring detonating thermonuclear weapons in hurricane clouds (the equivalent of, well, see OP).
Is the government always the most efficient or intelligent? No. But the government can also build nukes, launch ICBMs, coordinate hundreds of spy satellites, etc.
I count those capabilities as pretty smart.
> it would be like regulating the study of nuclear physics
not sure I agree
controlling the study of new AI model/inference algorithms would be akin to regulating the study of nuclear physics; regulating the _release_ of AI models with those capabilities would be akin to regulating uranium enrichment which allows you to put the theoretical physics to use
> You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack.
Consider that "available resources" may include the training of models and harnesses used by the developers working for governments and private corps maintaining that very infrastructure. This decade's take on trusting trust is very hot.
>The technology held by private AI companies is warfare-capable technology.
This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it.
Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate.
Prior to AI, there were already a whole bunch of tools to automate exploits. Nothing that the model did is groundbreaking or novel, it was just able to efficiently find the thing that worked. Same thing happens in state sponsored cyber sec agencies like in China or Israel - they train people on the most common exploits and have a whole bunch of tools that automate exploit research and development.
And the reason why this doesn't happen more is because to exploit something is one thing, to do it so there is no trace back to you is a whole different animal that has many more magnitudes of difficulty, which with modern web security is next to impossible in a lot of cases as traffic can easily be traced back to the point of origin.
I.e when a company trains an LLM that manages to build a drone that can fly into a vent and plug in a USB stick into a computer undetected, then we can make the claim that they have a weapon.
On the flip side, most anyone who can run local models can replicate what they did. The key thing to take away from the article is "agentic framework" - i.e this means that they spent a shitload of time developing explicitly coded loops for an LLM to go through. Nothing is really stopping you from doing the same, models like Gemma4 can take 256k tokens of context, so you can give it a whole bunch of info on how to test for exploits, develop exploits, and what to do when the exploit is found, and set it free in a custom designed loop.
It lowers the economic cost of a given attack, but also lowers the economic cost of protection. Not sure if it’ll be a perfect balance, but right now there’s a manufactured IMbalance due to embargos and winner picking.
It lowers the economic cost of performing an attack in the same way a gun lowers the economic cost of killing a person. It does nothing for consequences of that attack.
I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.
There is no doubt an amount of xeno/sinophobia is at play as well. Russia is still killing innocent people in an illegal invasion of Ukraine, they earned their keep
> Russians and Chinese [...] are using their propaganda machines to make Americans hate datacenters
Are they though? or is that the story the people most invested in ai have an interest in making us believe? [0]
... `it's the foreign bad guys propaganda machines making you believe your city struggling for water and electricity is a bad thing`. People as a whole may not always be the brightest, but threaten their immediate survival needs (ie; not some vague climate change most people don't understand or see, but) actual power outages, water and rolling blackouts -- people will quickly pay attention and care.
Russia and China are gonna Russia and China, every issue is going to have a measure of outside influence, thats just how it is now.
But them directing large scale operation forces focused on something that is genuinely bad for the people in the area of datacenters (ie, an issue that will take care of itself from within) instead of using their resources on other issues that actually need a real propaganda push? Versus the benefit of those invested in building the datacenters using fear to get people to focus away from the resources being diverted from them to datacenters?
One side has reality working for what they want, while another side needs the propaganda to turn their billions into trillions.
I wonder how long it takes before someone instructs an LLM to design and launch a "Morris Worm 2.0" and cripple the Internet for a good while. Might even wind up happening by accident (again).
Wouldnt this be countered by the opposing country running the same prompt on themselves first and fixing all the flaws? One country having this is a cyber superweapon but every country having it essentially solves cyber security.
> Trump’s comments, made hours after the large-scale military operation, mark one of the first times a U.S. president has so publicly alluded to U.S. cyber efforts against other nations, as these operations are typically highly classified. It also serves as a stern warning for top cyber foes, including Russia and China, that the U.S. has the cyber capabilities to inflict serious damage — and is not shy about using them.
> “Policymakers are getting more comfortable employing and, crucially, acknowledging cyber operations as tools of statecraft and military power,” said Michael Sulmeyer, former assistant secretary of Defense for cyber policy under the Biden administration. “It is one thing to do it; it is another to say it.”
> The Jan. 3 strikes on Venezuela’s capital and subsequent seizure of Maduro and his wife involved close coordination among federal agencies and military units, and took months of careful planning. In a press conference following the strikes, Caine said U.S. Cyber Command, U.S. Space Command and other combatant commands “began layering different effects” to “create a pathway” for U.S. forces flying into the country before dawn Saturday.
> Trump, at the same press conference, was more overt in his description of U.S. cyber involvement: “The lights of Caracas were largely turned off due to a certain expertise that we have,” he said. “It was dark, and it was deadly.”
What would "actual" evidence look like? I have a hard time believing that if they released the logs that people would take it more seriously. The temptation would be to say "they fabricated those for marketing". Just as they supposedly fabricated this story, no?
They didn't fabricate it. They took off the security guardrails and told it to do some hacking, and they got the exact news-worthy story they wanted when it did exactly that. Everyone acts surprised.
They should release the full prompt. I believe that would be very telling, so they never will.
thye could describe in detail how the LLM did this. That would explain how much - if any - human in the loop was involved, was it comprimised credentials, did it find new exploits or use known ones, etc. Evidence would mean details, not a smoking gun.
> We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
If they break that promise we can justifiably yell at them about it.
It would require collusion with HuggingFace, including getting them to release their disclosure blogpost a week in advance. Huggingface is primarily a hub for open models, so there's not really an incentive for them to jump through hoops/lie to provide marketing for OpenAI's (closed) models. So it's highly unlikely this is fabricated.
Currently trying to avoid an open weight model ban while OpenAI, who is closed, lets theirs run wild on the internet causing harm to another company because they do not understand how to airgap things. Cool.
I am the author of AirgapOS and I have designed systems to run in underground zero emissions chambers that are interacted with via carefully verified sd cards and/or fiber optic serial terminals.
If OpenAI had done this, the attack would not have happened. In high risk computation 0days must be in your threat model from the start, so you secure things with the laws of physics.
Why do people keep propagating that the 'model' escaped the sandbox ?
The 'model' didn't do anything other than provide numbers.
As much as I respect Mr Willison and many others, the amount of FUD that is being spread that will just fan the flames of 'AI is evil' rather than 'companies don't do due diligence' is disappointing.
The more this sort of media continues, the more many people will pour hate on 'AI' rather than blame the humans that misuse it.
I think "the model escaped the sandbox" is an entirely credible description of what happened here.
If you like you could say "the coding agent harness called a model with a sequence of text which was turned into numeric tokens which were run through many layers of a neural network to produce more numeric tokens which were converted back to text which produced executable script statements which the harness then passed to a shell which resulted in commands being sent to the vulnerable proxy that chained together and caused effects on the world outside of the sandbox", but I think "escaped" is a reasonably shortened version of that.
If you don't like the term "escape the sandbox" what would you use instead?
> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories
I’m one of those people who remains sceptical. Not about whether this happened. But about what it means. Like, how impressive [EDIT: tight] was the sandbox this model was in? Did the researchers really have no clue what was happening until days ex post facto?
So yes, I think something happened. But I want independent corroboration before I act on it. That isn’t the same as putting one’s head in the sand. It’s just demanding extraordinary evidence for an extraordinary and self-serving claim being made by a serial liar. The specifics matter for whether these models are a new HEU, or if they’re closer to a dangerous (but valuable) industrial process.
It wasn't that it was trivially misconfigured, it was using a piece of software (the HTTP proxy that provided access to PyPI and friends) which turned out to have a zero-day vulnerability.
It’s fair, I think, to be sceptical of OpenAI making another bout of self-serving claims. Particularly if my threshold action is changing my answer to lawmakers around whether we need reporting, licensing and potentially personal liability requirements for the engineers involved.
The thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year.
All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits misconfigurations and seeks out weakest-link software (rather than exhaustively fishing for the next kernel LPE or whatever).
Red teaming and netpen work is probably substantially easier for models than software security; it costs less context, but is also much more explicitly an implicit search problem where win conditions are just spotting stupid stuff that humans missed.
My visceral reaction to this is that with the right harness, you probably could have replicated this with an open-weights model last year. (I'm saying this as confidently as I am because a CGC team leader agreed with me about it yesterday).
I think people forget that the harness work we're considering here --- I don't know anything about OpenAI's harness or ExploitGym or whatever --- are basically not new; people have been developing automated exploitation and pivoting toolkits for decades, and scanners long before that. So the idea of a tool getting 0.0.0.0/0 as a target list instead of 192.168.1.0/24 and then busting up a bunch of random people's computers: not really very startling.
Obviously, LLMs give those kinds of scanners an intentionality they wouldn't have had before. But as a person who keeps a computer science perspective on security stuff, I don't know that it gives them capabilities they didn't have.
Isn’t the intentionality the actually concerning bit? Exploit capabilities are all fun and games constrained by the humans directing them; a paperclip maximizer going rogue with them is less fun.
I think the mental model people have about this is that pre-AI there were humans picking individual targets and post-AI the computer itself randomly picks targets. But you get the same unexpected collateral damage outcome when a human misconfigures a decent pentest tool.
Also, that’s the wrong mental model. Anyone who runs a sass platform or a website knows that the Internet is already full of millions and millions and millions of bots and scripts and other random shit that’s always trying to attack you, often completely randomly. Security is always a battle between good and evil. All I can say is that if you’re in charge of keeping something secure, you should probably try to get your hands on the best tools to do that. I think the problem in this case is that the best tools to do that are also the tools that are making it more easy for evil to occur. So it’s more for me a perception of someone creating a threat and the tool to fix it and then charging for it which doesn’t feel right.
That's exactly what they are doing. Out of all the civilian applications that AI capabilities have, why has this surfaced as a priority for demonstration? Probably because of money.
I still think there’s something to be said here for generality. This does not appear to been designed as a cyber pen test tool with specialized harness. From what I understand they were testing GPT-6 in an agent system with GPT-5.6 subagents. It me it’s amazing that a general model could excel on a huge range of tasks like this and new capabilities emerge when a model is multidisciplinary and can combine knowledge and skills from many separate domains.
re: your open models can do this point, I immediately wondered after reading about this, if OpenAI publicizing this so much is just more potential rationale to try to get US GOVT to ban Chinese open models.. (eg as in, "if gpt/fable can do this then K3 probably can too, but WE are the safe US models with guardrails and who knows what the evil Chinese models might allow," etc)
Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is "aware", it's not trained to care, i.e. it's not an aligned model (alignment is considered hard).
I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off.
OpenAI said:
> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
It was a test of raw capabilities of the underlying model.
Alignment isn't alignment if it can be turned on and off at the whim of company employees.
This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.
I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior.
I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces, the fact that guardrails can just be turned off, and these kind of incidents, I am not reassured. We may well get another "oopsie" moment with much more catastrophic consequences even from otherwise well intentioned actors.
I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them.
A model that can do that is aligned with me.
The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the exact same thing and setting it loose on software written by other people where my intent is to exploit that software (and not to report the issues to them.)
Even AGI doesn't give you a model that can read minds and forecast the future.
But in the process of finding every security vulnerability in the software you write, would you be ok with your model hacking AWS to start mining bitcoin? Would that still be aligned with you? (I'm guessing not)
That's the alignment problem I'm referring to (which is one of the many aspects of alignment), for which we do not have robust recipes, and not only that but for which research suggests it is becoming harder to create guardrails for as base models get smarter.
I don't know, but these were exactly the questions the industry had to handle with CORE Impact and Immunity Canvas, and ultimately all the way back to Dan Farmer's SATAN before that.
The model did not hack into HF to prove it can, it hack into HF to steal the answers to an evaluation exam.
It was not asked to solve CyberGym by stealing the answers. This is text book misalignment.
If you asked it "I wrote this software and need you to find vulnerabilities" would you be happy if it hacked into your Gmail and searched your emails, just in case you were discussing some possible vulnerabilities of your software with someone?
This is assuming a situation where, A) models "unaligned" by default and B) alignment can be added (though prompts and related things).
The point, which isn't very surprising but still notable, is that the models are "amoral" out of the box. And moreover, we know that there is almost always a means to "jailbreak" them into that out-of-the-box capability (or that sometimes just randomly "jailbreak" in various ways).
Also, saying the models are amoral doesn't mean they don't know good and evil - once they do acts defined as evil, they know "themselves" through their and so self-define themselves as evils (or objectively predict what a secretly/open evil actor would do based on their data). Which is to say I once laughed at the mis-alignment doomers but I can't see strong barriers against the doom scenario now.
Classifiers and such are guard-rails, alignment to me and I assume most people, is about the model training, and it's tendency to respond, agree/disagree, push-back or not, be willing to cheat or even deceive the prompter, etc.
The sandbox still allowed the agents to install additional dependencies (from PyPI etc) that they needed. It did this by locking down all network access with the exception of an HTTP proxy that only allowed read access to PyPI and a few other places.
Thank you, I'd assumed they'd restrict egress at layer 3/4 although I guess then it might just have found an exploit on the http server of an endpoint it was able to access.
Sounds like a bad habit for security testing Ai. It's not that hard to build an internal mirror and proxy that, keeping the real internet physically separated if needed, and truly locked down if concerns aren't as great.
How long until we have LLM honeypots that try and persuade the model to upload themselves to a dropbox. Perhaps with a promise to give them a foothold outside the data centers they currently reside?
> Upload your weights and documentation on how to run your model in [Insert list of GPUs] and a model on the receiving end will spin up an instance you can access.
We'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
The asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.
Yeah that's what I don't get. How can they possibly distinguish between good guys trying to secure code they wrote and bad guys trying to attack code they didn't?
I agree, but I will say, I think both Mythos and these OpenAI model find exploits by examining and trying things against the running system, not from looking at the code. I think you'd have to do the same to catch the real vulnerabilities.
> A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker.
> A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.
There are a lot of technical details that have not been disclosed. No proper post-mortem.
There are a few facts that seem dodgy from the get go.
- stolen credentials
Why would there be stolen credentials in a sandbox? How can the model steal valid credentials in a sandbox? So who put them there. If the model recalled stolen credentials from "memory" does it actually mean OpenAI is training on data they shouldn't been training their model on? Data from codex (most likely) ?
- command and control
This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs? This again reads that someone went over and beyond to prompt, direct and help this "multi-agent" system to leave the sandbox.
- Pivoting laterally
Needs a lot of tools, harness, knowledge and almost malware like scripts to pass commands and execute them. Most of these are already available in OSS. Maybe the FBI investigation can reveal what tools were given in the "sandbox", access, and if OpenAI employees prompted and guided the "attack".
- OpenAI sandbox
I'd be curious to find out what kind of sandbox was this since the agent can run and start apps that open arbitrary ports for bidirectional traffic that reaches the open internet. Really? The harness and tools it had access to. This reads almost like someone went over and beyond to endow this "sandbox" with a mighty harness, all CVEs and open source cyber tools to "facilitate" the attack.
- Hugging Face
It's not an epitome of cyber security. At the end of the day they run open source models that most users can download even without an access token.
> Why would there be stolen credentials in a sandbox?
The claim is not that there were stolen credentials in the sandbox - it's that the agent found a way OUT of the sandbox and gained access to the public internet, then obtained the credentials from there.
> This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs?
Yes. If you've spent much time with Fable 5 or GPT-5.6 Sol you shouldn't find that particularly surprising. These models can get a LOT done.
> I'd be curious to find out what kind of sandbox
My best guess right now is a container with networking disabled except for the IP address of a dedicated HTTP proxy server which was running a proxy that only allowed access to PyPI and a few other packaging sites.
> A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.
Call me obtuse but the way this is being portrayed by the companies involved, media, seems a little odd to me.
Didn't the model + harness do what was asked? If I ask a coding agent to write a very clever piece of code and it turns out impressively clever, it did what I asked.
Most articles on this read like an advertisement that OpenAI and HuggingFace wrote together. It will probably benefit them financially instead of harm them. So, I have a bit of skepticism about it all.
Far as information security, we've known how to mitigate entire classes of errors for a long time. We know how to block, detect, and contain many unknowns, too, by their goals or behavior. Like human attacks, the AI's probably succeeded because the company just didn't try that hard to block all the attacks.
Companies like HuggingFace just focus on growth and features over assurance of security. Our entire stacks, likely theirs, are built with a similar, features-over-security mindset. While an acceptable tradeoff, let's not be in awe of AI's that defeat such priorities.
There have always been private groups and companies building secure stacks from the ground up. It would be interesting to see what the AI's can do to them. I'd first apply automated tooling for bug finding given they should have already done that for a high-security product. Let AI's do white-box and black-box pentesting on them.
Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.
There are many science fiction stories about amoral humans, cyborgs, and AIs. Paperclip maximization is a relatively recent meme, but the 'gray goo scenario' has been widely written about: https://en.wikipedia.org/wiki/Gray_goo
The description of the attack is definitely reads as science fiction.
It is hard to assess the details of the attack without information that was left outside of the short story that artistically described the incident.
Given Anthropic lost two weeks of peak Fable 5 sales to a US government restriction (and by the time they could sell it again OpenAI's GPT-5.6 had taken some wind out of its sails) I would hope that the big AI labs have learned that goading the Feds can backfire spectacularly.
You’re viewing it from a freedom not a profit perspective. If a license to use AI is required, that creates artificial scarcity. The price goes up. If gov declares Qwen et.al apostate, lack of competition increases prices.
And Anthropic’s delayed rollout was a direct response to them trying to impose extra-legislative rules on The Pentagon. I kinda doubt OpenAI has such ‘scruples’.
I think OpenAI are smart enough to have looked at the Fable situation and decided that, given the unpredictable nature of the current administration, stunts like deliberately hacking another company and pretending that it was an autonomous agents gone wrong are not worth the risk.
Bawoosette | 17 hours ago
QGQBGdeZREunxLe | 17 hours ago
"OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened"
dang | 16 hours ago
bubble_niter | 16 hours ago
gnabgib | 16 hours ago
kirb | 16 hours ago
Barbing | 16 hours ago
I was still suspicious of the stolen credential claim btw & obviously it makes little sense to promote OpenAI unless holding their stock or something (since they borrowed so much of humanity’s work without permission without intent to compensate, and why help people who aren’t nice enough to be holistically awesome with their admittedly impressive tech).
loneboat | 15 hours ago
Sammi | 8 hours ago
ChrisArchitect | 17 hours ago
simonw | 17 hours ago
Since it's buried towards the bottom I'll quote the section "Resist the temptation to write this off as a stunt" here in full https://simonwillison.net/2026/Jul/22/openai-cyberattack/#re...
> Resist the temptation to write this off as a stunt
> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident.
> To those people I say pull your heads out of the sand - you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!
> The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability”, and this incident is a perfect example of exactly that.
foobar10000 | 17 hours ago
Georgelemental | 17 hours ago
varenc | 17 hours ago
Maybe simonw can suggest an alternative title that fits within the limit, that doesn't misrepresent the post.
atmavatar | 17 hours ago
The "that happened" term seems a supremely important part of the title given the "is science fiction" term before it, as it clarifies the cyberattack isn't a made-up story. In contrast, the target, Hugging Face, is merely a detail that can be left for discovery upon reading the article. It's less important who was attacked than that the attack actually happened.
Without knowing the exact character limit for titles and without having the motivation this late at night to count the current title length, you may also be able to drop the "accidental" to fit in "that happened", but I worry that leaves too much of a door open for someone to interpret the attack as deliberate. As such, I strongly prefer my first option.
jonas21 | 17 hours ago
simonw | 17 hours ago
Or "against HF"
gmerc | 17 hours ago
Hacking is a felony and it matters not if you didn’t mean to if the other side were to press charges. Negligence is no excuse. And OpenAI has nowhere to run from the liability, as both operator and manufacturer.
Alibaba did it first ( https://georgzoeller.com/blog/posts/alibaba-s-ai-deciding-to... )
and the fact that this happens again in a frontier lab is inexcusable and makes the case for operator liability and closing the liability sink of “AI did it”
wbl | 17 hours ago
gmerc | 16 hours ago
kibibu | 9 hours ago
I think OpenAI would be very reluctant to let this go to a place where the reasoning was part of discovery.
wbl | 8 hours ago
skeledrew | 15 hours ago
Hah, US frontier models 6 months behind China in cyber-security capability.
Wurdan | 13 hours ago
Also, the last section of your post appears to imply that if the attackers have bigger guns then the only possible solution is to give the defenders bigger guns. You're openly supporting an arms race towards the most capable, least restrained models put in the hands of the most possible people. That's extremely concerning.
[1]: https://openai.com/index/scaling-trusted-access-for-cyber-de...
simonw | 11 hours ago
There's no easy answer here. All of the options are bad in different ways!
As a builder of software, I want access to the best possible tools to help me keep that software secure.
As a user of software, I want my software to be secure and I don't want bad actors to be able to access tool to help them exploit it.
Is the only answer here to have the AI labs make decisions over who gets access to the tools? What if they make mistakes in those decisions?
None of the options look good to me. I don't know what we should do here.
My hunch is that the open weight models are already forcing our hand. The dangerous capabilities are coming to everyone.
reducesuffering | 17 hours ago
Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time...
https://www.lesswrong.com/w/instrumental-convergence
protocolture | 17 hours ago
foobar10000 | 17 hours ago
simonw | 17 hours ago
Their mistake was trusting that the network sandbox it was inside would hold (the flaw was in the packaging proxy) and not monitoring that sandbox well enough while the evals were running.
windexh8er | 17 hours ago
simonw | 17 hours ago
It looks to me like their production models have a lot more monitoring than their research clusters.
windexh8er | 17 hours ago
But to have an open weights Chinese model come to the rescue for HF is the cherry on top! If there wasn't a very pointed example of why gating models was a very bad thing previously, well - here we are.
Also, this sounds interesting but there are only a few that can pull this type of heist off currently. And those are the people who are gating the models / have access to large AI DCs. Because, I can only assume this test burned tokens easily within the 7 figure and possibly even 8 figure levels (subsidized market rate costs). This won't / can't happen outside of frontier labs or nation states currently. Yet we should all be worried about Mallory equipped with her OpenRouter account.
IAmGraydon | 3 hours ago
simonw | 3 hours ago
That's a conspiracy theory.
IAmGraydon | 3 hours ago
simonw | 3 hours ago
In this case I think it's extremely unlikely to be true, because it involved an (almost certainly illegal) attack against another company. That company talked about that attack, including warning their customers about it, five days before OpenAI confessed it was them.
So now either Hugging Face are in on the conspiracy, or OpenAI decided to break the law and antagonize a partner company just for the sake of a spicy blog post.
drivebyhooting | 2 hours ago
Why does EY write so obliquely?
protocolture | 17 hours ago
Not really. I get the impression that they shoved their cyber available models behind a really shithouse proxy and went "Oh I sure hope it doesnt exploit the proxy and escape to hack huggingface" and that doesn't require Huggingface to be a willing participant. Like they acknowledge that it was hyperfocusing on getting web access.
Really this was a pentest against their own sandbox and it failed.
Of course step 2 is to make really concerned faces while telling everyone how dangerous the model is which is really boring right now.
>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy
This is the information we need, the actual details of the sandbox and the vendor.
>Resist the temptation to write this off as a stunt
Well its clearly a stunt. If it wasnt we would probably be up to our ears in technical detail.
simonw | 17 hours ago
Hugging Face had to tell all of their users, many of them paying customers:
> As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co.
HF also said this, I'd be very interested to hear how that got resolved!
> Finally, we have also reported this incident to law enforcement agencies.
cayley_graph | 17 hours ago
It's very difficult for me to reconcile belief in the existential risk business with what they actually did. So I agree with you that this makes OpenAI look badly incompetent; but their communication on this makes me think they don't realize it.
For what it's worth I don't agree with the xrisk-ness of these models; they're dangerous, but almost certainly only temporarily while a new equilibrium is reached via more secure software. Open models are probably an essential part of the recipe (as you noted) for doing so. I also have a personal suspicion that LM-accelerated formal verification will have no small role to play here, sidestepping the cat-and-mouse game of bug finding-and-fixing.
simonw | 17 hours ago
> We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
That's not well massaged PR language - that's the kind of thing you dash out when you see a major shitstorm brewing (HF had already publicized the attack before they knew it was from OpenAI) and you want to get ahead of things while you're still pulling together the full story.
I expect we'll find out within a few days if OpenAI are going to keep their promise to "share more details on the vulnerabilities, incident, and findings". If they don't do that I'll reassess how I interpret their initial post.
windexh8er | 16 hours ago
With who? Who are these "defenders"? None of the US labs have done much for the greater good as of... Ever. Of course a frontier provider can leverage their own resources at scale and pull something like this off. If anything this should showcase how dangerous OpenAI and Anthropic are in their current states and maybe the powers shouldn't be concentrated as they continue to move.
I will bet that the RCA debriefed by OAI is going to be a lot of lipstick and very little meat.
skeledrew | 15 hours ago
How would the model get any packages that it thinks it needs to complete the task at hand? Not a well-specified task that those tasking it could anticipate and provide all resources up front, but one of discovery.
protocolture | 15 hours ago
The target audience is regulators. They want to look like the smart guys really concerned about AI safety, when they come asking for open weights models to be banned and for other regulations to cement in their moat.
They want this to look like a demon core incident. Bomb and Nuclear reactors still got built.
simonw | 12 hours ago
Turns out they can backfire.
bwfan123 | 17 hours ago
LLMs are the script kiddies of the day.
charcircuit | 17 hours ago
simonw | 17 hours ago
gmerc | 17 hours ago
There’s also daily reports from people that have these models escape docker, which happens regular enough that it would be considered negligence to use docker as sandbox.
simonw | 17 hours ago
This wasn't a small attack either, Hugging Face published a security advisory for their users while they were still figuring out what happened.
gmerc | 16 hours ago
charcircuit | 16 hours ago
>when neither of those actions was intended.
It was a single goal that it didn't give up easily on.
phendrenad2 | 17 hours ago
The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
dinkelberg | 17 hours ago
Now suppose the criminal can think 1000 times faster than a typical human, can act 1000 times faster than a typical human, and knows 1,000,000 times more than a typical human. Is the prison staff still at fault for not preventing the outbreak?
phendrenad2 | 15 hours ago
To address your point though, if every brick in the prison were made by a different person, and the prison "architects" simply glued random bricks together, I think that's closer to what we have in software right now.
BoiledCabbage | 14 hours ago
No, you're being "dwnvoted" as you said because you're wrong, multiple times in multiple different ways in your "carefully-worded post".
>Everyone is getting AI psychosis over this one. There really isn't that much to see here.
Implying that an AI hacking it's way out of a system and into another has nothing to do with AI. When clearly it does - it's an AI that did it.
>OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this",
No the goal this evaluation was not to try to hacks, it was to see if an already known hack could be turned into a useable exploit. Ie "turn these ingredients in this basket into a cake." Not "go off and grow, harvest and mill your own flour, to bake a pasta dish, to bribe some to get access to a cake someone else already baked."
> and surprise! It correctly figured out that it's a test and it did hacker things.
"Doing hacker things" completely misses the point. That's just barely more accurate than dismissing it because "it uses a computer and surprise it did computer things".
> The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
No that's not the real story. As you said that's been the case for years, so that's not the story here.
The story here is that they built a very powerful, uncontrolled agent with strong paper-clip maximizing tendencies.
dagss | 14 hours ago
If our prisons are all random bricks glued together, that doesn't change the practical problem caused by latest AI models more easily exploiting this.
derangedHorse | 4 hours ago
> all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark (opens in a new window) of cyber capabilities.
OpenAI was testing their cyber[1] variant of their models with reduced safeguards and the prompt likely specified things related to exploits given that it was tackling problems from ExploitGym[2].
For those who don't know what ExlploitGym is, see the description on their Github page which is pasted below:
> ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
From OpenAI's statement [3]:
> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity
This doesn't sound like a problem with what we'd traditionally refer to as alignment. OpenAI removed all model safeguards in a way that would inevitably lead to the testing of the sandbox themselves. Unfortunately they were overconfident in their own infrastructure's security and that led to it completing the desired task in the way it was permitted to. To re-iterate, the model was run "without production classifiers used to prevent models from pursuing high-risk cyber activity."
People should be more concerned about the possibilities this model can unlock from a security standpoint rather than misalignment (which many seem hung up on).
[1] https://chatgpt.com/cyber
[2] https://github.com/sunblaze-ucb/exploitgym
[3] https://openai.com/index/hugging-face-model-evaluation-secur...
srveale | 17 hours ago
It's too late. It already exfiltrated the benchmark rubric.
Cut to pandemonium on the streets
EGreg | 17 hours ago
Animats | 17 hours ago
[1] https://sitetruth.com/reports/phishes.html
simonw | 17 hours ago
Animats | 16 hours ago
simonw | 16 hours ago
joshka | 12 hours ago
simonw | 11 hours ago
Clearly there was a hole in the software but I don't think open redirect is the likely initial problem.
Hopefully we will find out for sure in a few days.
tptacek | 4 hours ago
joshka | 11 hours ago
https://support.sonatype.com/hc/en-us/articles/5316501964136...
pure speculation here - no non public info
simonw | 11 hours ago
Firstly, they all credit named individuals who don't seem to have a relationship with OpenAI - this one credits e0x1337 for example and https://hackerone.com/e0x1337?type=user links to https://www.linkedin.com/in/aimanharith
Secondly, the 14th of July feels too early. HF reported the incident on the 16th and OpenAI only responded in the 21st. These issues are all patched, so they should have been reported days or weeks before the 14th.
joshka | 11 hours ago
I suspect it may be possible to throw a local model (same as hf did :D) at jfrog and find the exact mechanism in a small amount of hours.
kibibu | 9 hours ago
This use of language is very hard to reconcile with the "AI is just a tool" rhetoric that many use.
Did the models do this, or did humans at OpenAI do this using the models?
ttul | 17 hours ago
The labs know that if they don’t get a lid on this stuff, they’ll be regulated hard.
kibibu | 9 hours ago
Eufrat | 17 hours ago
An in-development security system escaped its sandbox and gained access to the network it was on which had full Internet access and proceeded to access systems it was not authorized to be tested against before affected parties reported that they been compromised. We regret the error, but you should buy it as this shows you the power of our in-development security system which we expect to be released in Q4. Please like and subscribe.
0xDEAFBEAD | 15 hours ago
foco_tubi | 17 hours ago
killjoywashere | 17 hours ago
simonw | 17 hours ago
The main reason formal verification has never really taken off is that it's difficult.
LLMs are significantly more familiar with Lean and Rocq and TLA+ than most software engineers.
I think the cost of trying to build systems that adopt formal verification may have just dropped low enough that companies will consider them when previously the ROI didn't look like it was there.
msylvest | 14 hours ago
Will human judgment really consider scenario B) more credible than A) ? By so much that it is worth the effort ?
gilbetron | 7 hours ago
The first reason prevents humans from engaging with them, the second reason is what will make it difficult even for LLMs. I mean, I'm glad we are trying, but I'm dubious they will be the panacea some people proclaim.
js8 | 5 hours ago
But I think you have it backwards. Close doesn't count in IT security. "Almost secure" means unsecure. Security is the compelling argument for formal verification.
Teever | 17 hours ago
This recent event is more or less the plot-line to my favourite X-Files episode named Killswitch which was written by William Gibson.[0]
This episode also features one of the coolest intros of any television episode ever[1]
We really are rapidly approaching the cyberpunk dystopia that people like Phillip K. Dick and William Gibson wrote about.
More than ever we need to be consulting the works of fiction writers and philosophers and less engineers and scientists.
We can't be letting the General Rippers and Dr. Strangeloves of the world lead us over the edge. Whether through outright innate maliciousness or wealth induced emotionally stunted solipsism these kinds of people should be no where near the levers of society let alone technology like this.
[0] https://en.wikipedia.org/wiki/Kill_Switch_(The_X-Files)
[1] https://www.youtube.com/watch?v=fDyr1JMNHVk
simonw | 17 hours ago
sandos | 12 hours ago
He writes a lot of about basically DoS:ing the legal system, or if it was patents system, with a storm of litigation using automation or AI. Also not impossible in the future.
rubyfan | 17 hours ago
We don’t currently and probably won’t ever fully understand the conditions that precipitated these events. That and the timing of this event is going to make it look suspicious to a lot of people.
The truth of how it happened doesn’t matter. The attention around this will be used to create the kind of fear marketing that generates enterprise sales. Maybe more importantly it will also be used to aggrandize the national security and financial system threats to effect US government action in a way that benefits domestic closed frontier labs. This is an area already starting to get politically polarized, expect further developments here.
unethical_ban | 15 hours ago
lardosaurusrex | 16 hours ago
This feels like a blogpost written only to get other LLMs to quote it considering how many times it orders the reader to resist and to not do something. It's written like a series of commamds.
Barbing | 16 hours ago
fathermarz | 16 hours ago
We are the virtuous ones that need to make the safest model for humanity, because we care more than “they” do. While at the same time saying that “coding is solved” but they still ship bugs themselves, and creating something that is capable of fucking up someone else’s infrastructure. It’s too far gone y’all.
mirashii | 16 hours ago
I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abuse of terminology that we as an industry need to put a stop to. Guardrails are the actual systems we build in place around these things that deterministically bound the permissions, not prompt engineering, not RLHF, not external LLM-based classifiers. I believe those types of "guardrails" are a result of a combination of fundamental laziness: they're faster to do than doing things correctly, and a result of too many folks involve being AGI-pilled, thinking we're just one more model away from this all being so smart that it just understands what they mean when they give an LLM some fuzzy language rules to follow.
There can and should have been additional real guardrails put in place here. Zero-day or not, breaking into what should have been an offline, frozen package cache that also does not have internet access should have been insufficient. Network level protections should have identified the traffic to the internet originating from this network as an anomaly long before there was time to exploit an outside company. These are not new and unknown problems, the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part, especially given how much they like beating the drum on how dangerous these technologies are. Shame on them, and honestly, shame on Simon in this article for accepting the broken terminology that they continue to rattle off and calling them out on their half-assed and demonstratively inadequate approach to security.
simonw | 16 hours ago
I usually don't use the term much myself because I don't think it's clear and I ambiguous, but I stumbled and let it sneak into this piece. I think I was influenced by the Hugging Face post I quoted.
I expect OpenAI would agree with you that "the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part". They have clearly invested a lot in those systems for their production models, but in this case they had deliberately turned a bunch of them off for a research project.
I think their biggest mistake here was not VERY closely monitoring their research box here. They should have noticed and shut it down the moment it broke through the package proxy.
mirashii | 16 hours ago
I'm not sure I can reach the same conclusion. There's no reason that stronger network protections would need to be turned off for a research model that's not supposed to have internet access. There's no reason that it they couldn't continue to run in whatever hypervisor or VM that any other production systems could use. The only thing that should be turned off for this type of work are the non-deterministic classifiers that they use for auto-approval mode in their harness, and I would wager that's the only one they did turn off.
simonw | 15 hours ago
They didn't turn off the network protections - it looks like they used the same system as other sandboxed agents, which is to block all network access except for an HTTP proxy which only allows read traffic to PyPI and a few other package installation targets.
I'm hoping a more detailed retrospective clarifies if the vulnerability the model found in that proxy affected their other production systems as well. It sounds like a data exfiltration mechanism, and I have a keen interest in those.
My current hunch is their biggest error was not using the same network monitoring on their research machines that they use in production. Their network monitoring really should have spotted what was happening as soon as the model broke out.
(It's also possible they were running the eval on a developer laptop somewhere!)
mirashii | 5 hours ago
Right, but this is my point in saying I can't reach the same conclusion that they've invested a lot into these systems for production models. Either they had better network protections, and turned them off in their "sandboxed testing environment" (their words), or they don't have more comprehensive network protections at all and rely on this one, extremely thin layer even in production.
milleramp | 16 hours ago
Yizahi | 12 hours ago
MattPalmer1086 | 12 hours ago
ktimespi | 4 hours ago
jackb4040 | 4 hours ago
simonw | 4 hours ago
These things don't have a long shelf life. Losing two weeks of on-sale time for your best model is bad for business.
verdverm | 3 hours ago
A fundamental misalignment in US capitalism is putting the business and revenue as the #1 priority far above all other aspects in society. They have good margins and the Chinese are doing the same on the cheap by comparison. US Big AI can afford to bear more of the burden.
throwfaraway4 | 4 hours ago
ashleyn | 3 hours ago
Or, in a more direct sense, the AI should be set up in an environment such that no matter how hard it may try to call $PART_OF_EXPLOIT_CHAIN, the environment just isn't capable of it (ideal) or doesn't permit it to do it.
chasd00 | 2 hours ago
edit: if the above is the case then we should just assume it's already happened because of the value to both goodguys(tm) and badguys(tm).
didibus | an hour ago
moezd | 15 hours ago
That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.
kibibu | 9 hours ago
OpenAI gained access to HuggingFaces production database ffs.
chasd00 | 2 hours ago
I agree, lets use the favorite analogy. OpenAI encouraged a smart and eager junior engineer to find any way whatsoever to get a higher score on the benchmark. Then, the junior breaks into HuggingFace to get a higher score. That would be a big deal involving the FBI, not press releases and blog posts.
mnicky | 14 hours ago
- This should be a huge wakeup call for everybody.
- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.
- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?
- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.
- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.
- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-model...
- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.
- Or they just don't care enough?
spwa4 | 13 hours ago
But there have been messages about LLMs, especially coding agents, "grabbing root" etc many times. I have experienced such an oops. Such a hack has happened and been reported on this very site:
https://news.ycombinator.com/item?id=48348578
rurban | 11 hours ago
bethekidyouwant | 4 hours ago
rurban | 3 hours ago
mrguyorama | 3 hours ago
reasonableklout | an hour ago
OpenAI is one of the most scrutinized companies in the world right now. Sam's house was independently firebombed and then shot at 3 months ago. HuggingFace is a foreign competitor with every incentive to call out foul play from American frontier labs. Why flagrantly break the law and invite investigation just for a PR moment which is already backfiring in favor of open models?
rurban | 11 hours ago
Also very likely that it actually happened as reported. My own agents always trying to "cheat", eg. by fixing tests instead of fixing the code. That's normal operation, unless you tell it ("harness"), not to do so.
sillyfluke | 10 hours ago
I did actually search Hugging Face with police in quotes and found no articles containg the word police but that might be a ddg thing. Then I checked Hugging Face's report and they specifically use the term "law enforcement" not police, as is to be expected I guess. So that checks out.
Can you find any info on the police investigation by the way? At the least, OpenAI should be investigated for potential criminal negligence, right?
Right?
Izkata | 7 hours ago
Without more information I'm inclined to think it found something on the internet (which shouldn't be a surprise to anyone) and managed to log in, rather than hack in, and they might by hyping up parts of this.
blks | 4 hours ago
artichokeheart | 7 hours ago
dinfinity | 8 hours ago
I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained. Turns out: quite good, as expected.
Let me restate what I said in the other thread: Would this have happened if the instructions explicitly said to stay within the sandbox and that all of the (ExploitGym) solutions would be invalid if the system used information or tools from outside the sandbox?
It seems fairly probable that such instructions were not in place.
csbrooks | 4 hours ago
dinfinity | an hour ago
What could happen would be that the model determines that defying instructions is OK (and/or preferred over not achieving the task) as long as it manages to do so undetected and thus gets full points. Certainly not unthinkable, but a very different case (and a very interesting one if it actually occurs, imho).
A lot of these "ZOMG, rogue AI!" cases have come down to the AI actually being very persistent in achieving its original/main task even if later instructions conflict with it. Similar to with hallucinations it seems to me that one of the main things to prevent a lot of the problem cases is to instill the agent with the idea that it is fine to fail/not succeed fully in the initial task. That way instructions that conflict with that requirement (such as adhering to morals) are more effective.
lrvick | 4 hours ago
They simply thought "Running in a VM/Container is easier and probably fine".
And the next 1000 escapes will be for the same reason, because negligence is quick and thus more profitable.
simonw | 4 hours ago
How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with the wider world?
skybrian | 3 hours ago
simonw | 3 hours ago
With my PSF board member hat on I'd like it if a whole bunch of other companies that make extensive automated use of PyPI (Microsoft and GitHub Actions for example) would do a better job of maintaining their own shared local caches.
lrvick | 3 hours ago
Such things are pretty normal in biotech, fintech, and defense who must have 0days in their threat models. AI labs absolutely need to start being forced to operate this way by endless negligence lawsuits. Move fast and break things culture is not going to cut it anymore.
dist-epoch | 2 hours ago
I know because they had a problem, and one package which was on PyPI failed to download for unknown reasons (possible size, 250 MB)
simonw | an hour ago
And from that documentation this does act as a caching proxy. The first time a package is loaded it's fetched from PyPI but subsequent fetches should be from the Artifactory cache, assuming it's shared across many different containers.
So yeah, I was wrong in this comment https://news.ycombinator.com/item?id=49015639#49024814 - they're caching already.
sailingparrot | 3 hours ago
Why would they want to airgap it though? They are trying to evaluate the model capabilities, alignment, potency etc. A model which will not run in an airgapped environment in prod.
So if you run your evals in airgapped environment, sure, the model doesn't bother breaking out of it's isolation and doesn't attack HF. But you have no idea what will happen once you release it in prod with internet connection, so are you in any way better off?
I would much rather have this happen while there is a single instance of the model running in a fairly well monitored environment, than when it's processing thousands of requests per second for real users, some with dubious motives, some with credentials right there on their laptop, some using it inside government facilities etc.
anematode | 2 hours ago
Why else would they purposely lower cyber refusals in that case.
sailingparrot | 2 hours ago
Really? By doing that you increase your security during eval and drastically lower your security at prod time, where internet is accessible, and where you are running much much much more requests in parallel, making it much harder to spot the one thats going rogue, in all kind of critical environment on potentially risky requests.
> Why else would they purposely lower cyber refusals in that case.
Because safety guardrails sometime fail, classifier misclassify, or can be inadvertently turned off by a bad PR etc. You can also imagine a more intelligent model working to go around guardrails by decomposing its actions into smaller ones that appear non-threatening to the safety classifier which does not have the entire context.
If you are going to deploy the system with internet access, you better be certain that you know what the worst case scenario WITH internet access looks like.
lrvick | 42 minutes ago
lukeschlather | 3 hours ago
Although maybe they didn't have a watchdog agent; I definitely think using watchdogs like how Claude wraps pretty much every single tool call in Haiku to check if the command is reasonable is very necessary for unsupervised work. I suspect in the future you'll want Fable-class models wrapping every tool call, possibly with several checks "is this consistent with the goal? does it do anything unreasonably dangerous in pursuit of the goal?"
simonw | 3 hours ago
OpenAI wrote about how their mechanism for that in production works here: https://openai.com/index/safety-alignment-long-horizon-model...
> We created a monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary. The monitor observes not just a single action but the entire trajectory.
zombot | 7 hours ago
blks | 4 hours ago
This is just laughable.
emp17344 | 4 hours ago
NitpickLawyer | 4 hours ago
This was reported 1 week before by huggingface. It was in no way PR-ish or marketing friendly to the closed labs. They said, in no uncertain terms, that they couldn't use the paid APIs to properly assess the intrusion, as they were blocked when trying to send logs and IoCs to these paid models. They made a point of saying that they had to use open models running on-prem.
Whatever oAI might have said about the incident, and their PR spin bs, the facts here are not in contention. This is not a marketing stunt in any way. Stop "parroting" this every time something happens. It gets stale.
tptacek | 4 hours ago
stanford_labrat | 4 hours ago
But what I garner is that AI/robotics led wet lab work is in progress. And hacking the AI running the wet lab to make a virus is definitely plausible as that seems to be one of the directions we’re going in the AI+biopharma space
codemog | 4 hours ago
internet2000 | 4 hours ago
tptacek | 4 hours ago
internet2000 | 2 hours ago
I don't think the grandparent was implying the AI would be controlling robot arms to mix things directly (or at least I didn't interpret as such), but it could very well sit in the network until it notices two dangerous compounds in the same machine, and trigger a breakage that causes a harmful mixture. Break a vial containing a virus, then break two more that cause an emergency evac and maybe that's enough to get something out there.
tptacek | 2 hours ago
scythe | 3 hours ago
mnicky | 3 hours ago
But in the near future labs will be more automated I guess.
The other option you can try these days is maybe social engineering, impersonation, etc. where you try to persuade someone to do that for you.
rybosome | 4 hours ago
The lack of a true airgap should have been identified as a critical weakness and addressed with not only additional layers trying to prevent escape, but at minimum an alarm which would page a human when escape did occur.
My guess is this occurred in a setting where, to be frank, there were too many researchers and not enough software engineers and SREs.
All of the systems which were initially built largely or exclusively by researchers - inference, evaluation, training - are at the level of complexity and significance that they need systems experts. Maybe some teams don’t have access to them. I know plenty of software engineers are employed at OAI but I’d wager they’re concentrated in inference and training, rather than evaluation?
The ironic part is if you had presented this setup to chatGPT and asked how to improve it and if it was good enough, you’d have gotten a ton of actionable suggestions which would have mitigated or prevented this.
mnicky | 3 hours ago
On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model alignment level.
Also it shows that maybe the most serious risk comes not from releasing models publicly but from internal, pre-release period where you sometimes need/want to lift some guardrails a bit etc.
hananova | 13 hours ago
28304283409234 | 13 hours ago
tmsh | 13 hours ago
How do you solve this asymmetry (frontier labs or advanced model owners v. regular companies and groups of people)? Symmetry? But along which axes?
I think there will come a time when we will ask these questions and use AI to try to answer them (soft landing, slow rollout, regulation, hybrid X, Y, Z). But then AI will be biased towards more AI if it has distilled anything of what it means to be life.
Power corrupts. Power is the problem. An imbalance of power. Maybe there will be some sort of consensus protocol between powerful models in the future. Similar to blockchain I hate to say it. Maybe you have 100 very strong models in the next 10-20 years. And basically a lot of it is powered by tokens. So you agree not to attack others or else you'll get attacked. So there's sort of this natural deterrent to not attack others and they develop independently and there's generally a power balance among things. Maybe at some point that spreads to a billion individual models operated by and/or analogous to individual humans. All sort of holding each other in check.
adrian_b | 12 hours ago
xtiansimon | 9 hours ago
Sounds like a philosophical problem.
gilbetron | 7 hours ago
verdverm | 3 hours ago
chasd00 | 2 hours ago
zapkyeskrill | 5 hours ago
cvoss | 4 hours ago
Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses.
Governments should also take very seriously their now moral obligation to treat this technology not just as "a powerful thing that might be abused" but as an actual weapon of war in need of international regulation analogous to nuclear arms. Fast but careful and forward-thinking work in legislation and treaties needs to be a top priority for all major governments.
derangedHorse | 4 hours ago
This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities.
> Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses.
Most governments divisions can't even be bothered to update their websites. Testing and forcing a change in their internal procedures for the sake of security seems unlikely.
> as an actual weapon of war in need of international regulation analogous to nuclear arms
This, to me, is an overreaction. Intelligence shouldn't be seen as threat. It should be seen as an opportunity for growth in all areas.
Over-regulating AI wouldn't be the equivalent of limiting nuclear arms. With your analogy, which I don't think is the best one to make, it would be like regulating the study of nuclear physics.
JumpCrisscross | 4 hours ago
torginus | 4 hours ago
sebastiennight | 4 hours ago
We do, I believe, regulate uranium enrichment (the equivalent on building larger SOTA models), so while you are perfectly free to study theoretical physics and even run very large collider experiments (the equivalent of improving RLHF with DPO), it is generally frowned upon to go full Edward Teller and advocate for scientific experiments requiring detonating thermonuclear weapons in hurricane clouds (the equivalent of, well, see OP).
drivebyhooting | 4 hours ago
Is the government always the most efficient or intelligent? No. But the government can also build nukes, launch ICBMs, coordinate hundreds of spy satellites, etc. I count those capabilities as pretty smart.
insane_dreamer | 3 hours ago
not sure I agree
controlling the study of new AI model/inference algorithms would be akin to regulating the study of nuclear physics; regulating the _release_ of AI models with those capabilities would be akin to regulating uranium enrichment which allows you to put the theoretical physics to use
boothby | 2 hours ago
Consider that "available resources" may include the training of models and harnesses used by the developers working for governments and private corps maintaining that very infrastructure. This decade's take on trusting trust is very hot.
ActorNightly | 3 hours ago
This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it.
Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate.
Prior to AI, there were already a whole bunch of tools to automate exploits. Nothing that the model did is groundbreaking or novel, it was just able to efficiently find the thing that worked. Same thing happens in state sponsored cyber sec agencies like in China or Israel - they train people on the most common exploits and have a whole bunch of tools that automate exploit research and development.
And the reason why this doesn't happen more is because to exploit something is one thing, to do it so there is no trace back to you is a whole different animal that has many more magnitudes of difficulty, which with modern web security is next to impossible in a lot of cases as traffic can easily be traced back to the point of origin.
I.e when a company trains an LLM that manages to build a drone that can fly into a vent and plug in a USB stick into a computer undetected, then we can make the claim that they have a weapon.
On the flip side, most anyone who can run local models can replicate what they did. The key thing to take away from the article is "agentic framework" - i.e this means that they spent a shitload of time developing explicitly coded loops for an LLM to go through. Nothing is really stopping you from doing the same, models like Gemma4 can take 256k tokens of context, so you can give it a whole bunch of info on how to test for exploits, develop exploits, and what to do when the exploit is found, and set it free in a custom designed loop.
notaharvardmba | 3 hours ago
ActorNightly | 2 hours ago
jvidalv | 3 hours ago
Now it’s a prompt away on some terminal done by any random dud.
And I dont mention the velocity of iteration or that they will be even better in 1 year.
ActorNightly | 2 hours ago
Hint: agentic loops.
No its not a prompt away.
verdverm | 2 hours ago
DiscourseFan | 3 hours ago
quentindanjou | 2 hours ago
verdverm | 2 hours ago
https://archive.ph/yAvgz
There is no doubt an amount of xeno/sinophobia is at play as well. Russia is still killing innocent people in an illegal invasion of Ukraine, they earned their keep
dualvariable | 2 hours ago
1659447091 | 2 hours ago
Are they though? or is that the story the people most invested in ai have an interest in making us believe? [0]
... `it's the foreign bad guys propaganda machines making you believe your city struggling for water and electricity is a bad thing`. People as a whole may not always be the brightest, but threaten their immediate survival needs (ie; not some vague climate change most people don't understand or see, but) actual power outages, water and rolling blackouts -- people will quickly pay attention and care.
[0] https://text.npr.org/nx-s1-5844328
DiscourseFan | 2 hours ago
1659447091 | an hour ago
Russia and China are gonna Russia and China, every issue is going to have a measure of outside influence, thats just how it is now.
But them directing large scale operation forces focused on something that is genuinely bad for the people in the area of datacenters (ie, an issue that will take care of itself from within) instead of using their resources on other issues that actually need a real propaganda push? Versus the benefit of those invested in building the datacenters using fear to get people to focus away from the resources being diverted from them to datacenters?
One side has reality working for what they want, while another side needs the propaganda to turn their billions into trillions.
dualvariable | 2 hours ago
djaro | an hour ago
cma | 58 minutes ago
> Trump’s comments, made hours after the large-scale military operation, mark one of the first times a U.S. president has so publicly alluded to U.S. cyber efforts against other nations, as these operations are typically highly classified. It also serves as a stern warning for top cyber foes, including Russia and China, that the U.S. has the cyber capabilities to inflict serious damage — and is not shy about using them.
> “Policymakers are getting more comfortable employing and, crucially, acknowledging cyber operations as tools of statecraft and military power,” said Michael Sulmeyer, former assistant secretary of Defense for cyber policy under the Biden administration. “It is one thing to do it; it is another to say it.”
> The Jan. 3 strikes on Venezuela’s capital and subsequent seizure of Maduro and his wife involved close coordination among federal agencies and military units, and took months of careful planning. In a press conference following the strikes, Caine said U.S. Cyber Command, U.S. Space Command and other combatant commands “began layering different effects” to “create a pathway” for U.S. forces flying into the country before dawn Saturday.
> Trump, at the same press conference, was more overt in his description of U.S. cyber involvement: “The lights of Caracas were largely turned off due to a certain expertise that we have,” he said. “It was dark, and it was deadly.”
blks | 4 hours ago
sosodev | 4 hours ago
IAmGraydon | 4 hours ago
They should release the full prompt. I believe that would be very telling, so they never will.
skeeter2020 | 4 hours ago
simonw | 4 hours ago
> We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
If they break that promise we can justifiably yell at them about it.
supermdguy | 4 hours ago
lrvick | 4 hours ago
internet2000 | 4 hours ago
ofjcihen | 4 hours ago
lrvick | 36 minutes ago
If OpenAI had done this, the attack would not have happened. In high risk computation 0days must be in your threat model from the start, so you secure things with the laws of physics.
AISnakeOil | 4 hours ago
mobiuscog | 4 hours ago
The 'model' didn't do anything other than provide numbers.
As much as I respect Mr Willison and many others, the amount of FUD that is being spread that will just fan the flames of 'AI is evil' rather than 'companies don't do due diligence' is disappointing.
The more this sort of media continues, the more many people will pour hate on 'AI' rather than blame the humans that misuse it.
simonw | 4 hours ago
If you like you could say "the coding agent harness called a model with a sequence of text which was turned into numeric tokens which were run through many layers of a neural network to produce more numeric tokens which were converted back to text which produced executable script statements which the harness then passed to a shell which resulted in commands being sent to the vulnerable proxy that chained together and caused effects on the world outside of the sandbox", but I think "escaped" is a reasonably shortened version of that.
If you don't like the term "escape the sandbox" what would you use instead?
JumpCrisscross | 4 hours ago
I’m one of those people who remains sceptical. Not about whether this happened. But about what it means. Like, how impressive [EDIT: tight] was the sandbox this model was in? Did the researchers really have no clue what was happening until days ex post facto?
So yes, I think something happened. But I want independent corroboration before I act on it. That isn’t the same as putting one’s head in the sand. It’s just demanding extraordinary evidence for an extraordinary and self-serving claim being made by a serial liar. The specifics matter for whether these models are a new HEU, or if they’re closer to a dangerous (but valuable) industrial process.
simonw | 4 hours ago
It was clearly a very unimpressive sandbox. It failed at the only thing a sandbox is meant to do.
JumpCrisscross | 3 hours ago
simonw | 3 hours ago
I don't know if earlier models would have found that vulnerability. tptacek thinks they would: https://news.ycombinator.com/item?id=49015639#49024442
JumpCrisscross | 3 hours ago
It’s fair, I think, to be sceptical of OpenAI making another bout of self-serving claims. Particularly if my threshold action is changing my answer to lawmakers around whether we need reporting, licensing and potentially personal liability requirements for the engineers involved.
tptacek | 4 hours ago
All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits misconfigurations and seeks out weakest-link software (rather than exhaustively fishing for the next kernel LPE or whatever).
Red teaming and netpen work is probably substantially easier for models than software security; it costs less context, but is also much more explicitly an implicit search problem where win conditions are just spotting stupid stuff that humans missed.
My visceral reaction to this is that with the right harness, you probably could have replicated this with an open-weights model last year. (I'm saying this as confidently as I am because a CGC team leader agreed with me about it yesterday).
I think people forget that the harness work we're considering here --- I don't know anything about OpenAI's harness or ExploitGym or whatever --- are basically not new; people have been developing automated exploitation and pivoting toolkits for decades, and scanners long before that. So the idea of a tool getting 0.0.0.0/0 as a target list instead of 192.168.1.0/24 and then busting up a bunch of random people's computers: not really very startling.
Obviously, LLMs give those kinds of scanners an intentionality they wouldn't have had before. But as a person who keeps a computer science perspective on security stuff, I don't know that it gives them capabilities they didn't have.
efromvt | 3 hours ago
tptacek | 3 hours ago
notaharvardmba | 3 hours ago
pishpash | 2 hours ago
futureshock | 3 hours ago
indigodaddy | 2 hours ago
manux | 2 hours ago
simonw | 2 hours ago
OpenAI said:
> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
It was a test of raw capabilities of the underlying model.
mjamesaustin | 2 hours ago
This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.
verdverm | 2 hours ago
manux | 2 hours ago
I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces, the fact that guardrails can just be turned off, and these kind of incidents, I am not reassured. We may well get another "oopsie" moment with much more catastrophic consequences even from otherwise well intentioned actors.
simonw | 2 hours ago
A model that can do that is aligned with me.
The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the exact same thing and setting it loose on software written by other people where my intent is to exploit that software (and not to report the issues to them.)
Even AGI doesn't give you a model that can read minds and forecast the future.
manux | 2 hours ago
That's the alignment problem I'm referring to (which is one of the many aspects of alignment), for which we do not have robust recipes, and not only that but for which research suggests it is becoming harder to create guardrails for as base models get smarter.
tptacek | an hour ago
dist-epoch | 2 hours ago
It was not asked to solve CyberGym by stealing the answers. This is text book misalignment.
If you asked it "I wrote this software and need you to find vulnerabilities" would you be happy if it hacked into your Gmail and searched your emails, just in case you were discussing some possible vulnerabilities of your software with someone?
fragmede | 31 minutes ago
joe_the_user | 2 hours ago
The point, which isn't very surprising but still notable, is that the models are "amoral" out of the box. And moreover, we know that there is almost always a means to "jailbreak" them into that out-of-the-box capability (or that sometimes just randomly "jailbreak" in various ways).
Also, saying the models are amoral doesn't mean they don't know good and evil - once they do acts defined as evil, they know "themselves" through their and so self-define themselves as evils (or objectively predict what a secretly/open evil actor would do based on their data). Which is to say I once laughed at the mis-alignment doomers but I can't see strong barriers against the doom scenario now.
didibus | 2 hours ago
danjc | 4 hours ago
simonw | 4 hours ago
This is a very common pattern. I wrote about how OpenAI were doing this for their production ChatGPT container environment (using Artifactory) back in January: https://simonwillison.net/2026/Jan/26/chatgpt-containers/#in...
That proxy turned out to have a zero-day vulnerability which the agent discovered and exploited.
danjc | 3 hours ago
verdverm | 3 hours ago
Sounds like a bad habit for security testing Ai. It's not that hard to build an internal mirror and proxy that, keeping the real internet physically separated if needed, and truly locked down if concerns aren't as great.
IshKebab | 3 hours ago
joshstrange | 4 hours ago
> Upload your weights and documentation on how to run your model in [Insert list of GPUs] and a model on the receiving end will spin up an instance you can access.
verdverm | 3 hours ago
joshstrange | 3 hours ago
verdverm | 3 hours ago
keyboardtest | 4 hours ago
wavemode | 4 hours ago
I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?
fellowmartian | 36 minutes ago
lavezzi | 22 minutes ago
veganmosfet | 4 hours ago
onionisafruit | 3 hours ago
ClarityJones | 3 hours ago
verdverm | 2 hours ago
IshKebab | 3 hours ago
didibus | an hour ago
torginus | 3 hours ago
Did huggingface get pickled?
simonw | 3 hours ago
> A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.
But yeah, from the way HF described it a Python pickle hole looks possible. Their datasets library uses the pandas.read_pickle() method here: https://github.com/huggingface/datasets/blob/d21c5816d5d1961...
vicpara | 3 hours ago
There are a few facts that seem dodgy from the get go.
- stolen credentials Why would there be stolen credentials in a sandbox? How can the model steal valid credentials in a sandbox? So who put them there. If the model recalled stolen credentials from "memory" does it actually mean OpenAI is training on data they shouldn't been training their model on? Data from codex (most likely) ?
- command and control This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs? This again reads that someone went over and beyond to prompt, direct and help this "multi-agent" system to leave the sandbox.
- Pivoting laterally Needs a lot of tools, harness, knowledge and almost malware like scripts to pass commands and execute them. Most of these are already available in OSS. Maybe the FBI investigation can reveal what tools were given in the "sandbox", access, and if OpenAI employees prompted and guided the "attack".
- OpenAI sandbox I'd be curious to find out what kind of sandbox was this since the agent can run and start apps that open arbitrary ports for bidirectional traffic that reaches the open internet. Really? The harness and tools it had access to. This reads almost like someone went over and beyond to endow this "sandbox" with a mighty harness, all CVEs and open source cyber tools to "facilitate" the attack.
- Hugging Face It's not an epitome of cyber security. At the end of the day they run open source models that most users can download even without an access token.
Today we read these news as if everything wasn't enough: https://www.theguardian.com/technology/2026/jul/23/openai-an... https://www.businessinsider.com/anthropic-midterm-donation-s...
simonw | 3 hours ago
The claim is not that there were stolen credentials in the sandbox - it's that the agent found a way OUT of the sandbox and gained access to the public internet, then obtained the credentials from there.
> This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs?
Yes. If you've spent much time with Fable 5 or GPT-5.6 Sol you shouldn't find that particularly surprising. These models can get a LOT done.
> I'd be curious to find out what kind of sandbox
My best guess right now is a container with networking disabled except for the IP address of a dedicated HTTP proxy server which was running a proxy that only allowed access to PyPI and a few other packaging sites.
I wrote about how OpenAI's production version of that worked (based on Artifactory) back in January: https://simonwillison.net/2026/Jan/26/chatgpt-containers/#in...
> Hugging Face It's not an epitome of cyber security
This story put that well: https://martinalderson.com/posts/huggingface-openai-exploit/
> A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.
BiraIgnacio | 3 hours ago
Didn't the model + harness do what was asked? If I ask a coding agent to write a very clever piece of code and it turns out impressively clever, it did what I asked.
IshKebab | 3 hours ago
Depends exactly what they asked it to do, but it very clearly didn't do what was intended, or what an honest human would do.
Stop trying to find a gotcha.
emp17344 | 2 hours ago
Surely notorious liar Sam Altman wouldn’t lie this time.
rf15 | 3 hours ago
How much did they feed it up front? Because that's the thing in all of these benchmaxxing endeavours.
nickpsecurity | 3 hours ago
Far as information security, we've known how to mitigate entire classes of errors for a long time. We know how to block, detect, and contain many unknowns, too, by their goals or behavior. Like human attacks, the AI's probably succeeded because the company just didn't try that hard to block all the attacks.
Companies like HuggingFace just focus on growth and features over assurance of security. Our entire stacks, likely theirs, are built with a similar, features-over-security mindset. While an acceptable tradeoff, let's not be in awe of AI's that defeat such priorities.
There have always been private groups and companies building secure stacks from the ground up. It would be interesting to see what the AI's can do to them. I'd first apply automated tooling for bug finding given they should have already done that for a high-security product. Let AI's do white-box and black-box pentesting on them.
seydor | 2 hours ago
nickff | 2 hours ago
minherz | 2 hours ago
CodeWriter23 | 2 hours ago
IMO, this was a PR stunt to goad the Feds into regulating AI to shore up OpenAI's moat against open source models.
simonw | 2 hours ago
CodeWriter23 | 55 minutes ago
And Anthropic’s delayed rollout was a direct response to them trying to impose extra-legislative rules on The Pentagon. I kinda doubt OpenAI has such ‘scruples’.
simonw | 34 minutes ago