Should it be OpenAI, or should it be OpenAI customers who give the LLM the instructions and provide the LLM with the tools to execute code and make (malicious) network requests?
One would disincentivise providing capable AI models that can be used for cyber security research. The other would disincentivise criminals from commiting crimes.
[edit] - I realise now that this could actually be a case of OpenAI running those agents themselves, rather than someone using OpenAI's models? Could OpenAI be that careless?
In all of the cases that people are referring to in this thread: HuggingFace, that german wiki, Ruby Gems, the Australian statistics page, this UN statistics page... those two parties are OpenAI. OpenAI running their own agents on their own instructions committing crimes with their own computers. This is cut and dry.
But in other cases, shouldn't it be both? OpenAI is ultimately the one executing the model calls. It's not like they send you a hard drive or standalone box and then you use it how you want. It's all (intentionally) centralized to them, in a way that's core to their business model.
I clearly have been living under a rock, but in the case where it's just text in/out of their API, and a customer uses this on their own to do nefarious things, I don't see why they would be liable?
If they knowingly allowed use of their services for illegal purposes then yes, but in so far that they provide a service that can be used for useful things (including cyber security research) and did a best effort attempt at abuse, I don't see why it should make sense to hold them liable. This is especially the case now that frontier LLMs are almost a commodity that can be used without restrictions from providers outside of your legal jurisdiction.
If I train a model at my house on my workstation, execute it, and under my supervision it ransomwares a hospital and somebody dies. What would be a just punishment for me?
Now the reality, the openai engineers trained a model, executed it, and under their supervision (no users were involved) it committed so many felonies that we are learning about a new one every week. What would be a just punishment for OpenAI?
I believe, although I'm not a lawyer, there would be some liability, but I think a lot depends on intent as well. Punishment for other crimes also varies depending on whether it was an accident or not.
I realise that in this case it was OpenAI being responsible for their agents running wild, and they should know that that is to be expected and should have saveguards in place. If they can be shown to be negligent then the punishment can probably be expected to be a lot more severe than if it was an accident. I make no judgements as to what this particular instance is, but I do believe that OpenAI has a far more greater responsibility for its agents running wild than someone running a home lab.
You would think a company that's looking to IPO very soon would be doing more due diligence, especially since its product is supposed to help other companies do said due diligence.
Their true product is fear. Fear of missing out, fear of losing your job, fear of losing your life. Everything they do kinda makes sense when you look at OpenAI, Anthropic, etc. in that light.
There must be a list of all these abuses somewhere.
PS: In the same lazy energy of asking for a list instead going out and finding it or putting it together myself, are there any companies other than CloudFlare that are working on AI shields?
1. What is the harm is accessing this data
2. Why is this data private
3. What would it take to gain access to this data
4. What do we expect giving gremlins access to the internet
The more of these that come out the more incompetent OpenAI looks. It would appear there was a total lack of basic controls in place for running these tests.
I think the even bigger worry is that anyone who doesn't want to use their models safely can already do this with open models. Even if OpenAI, Anthropic etc get their act together, the cat's out of the bag.
I think they did not expect that models were capable of this level of sandbox escape (prior models certainly didn't have this kind of agency) and weren't prepared.
All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.
They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.
I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.
> They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.
Yes, they should have.
But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"
Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.
Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.
> But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?
There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.
This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.
(For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).
They were actively researching exploiting systems using their models. (I intentionally changed the ownership of the verbs here: they wrote the code, they trained the models, they don't get to dodge the responsibility.)
It's no shock that there are a lot of vulnerabilities in a lot of software. So then they gave their AI model + brute-force-machine loop system a mediocre sandbox and couldn't notice when it figured out how to exploit it?
Don't let people off the hook for the software they create.
I love how perfect the word "sandbox" is as a metaphor for the security controls they have. A sandbox is a wide, shallow box filled with sand for kids to play in. Even toddlers can crawl or step out of one on their own, it doesn't contain them at all without an adult constantly watching. Kids only stay in a sandbox if they're having more fun playing inside than they think they'll have outside it. AIs only stay in a sandbox if they're having more success inside than they think they'll have outside it.
Being cagey about their poorly setup sandbox is marketing. It gives the impression these are unstoppable juggernauts capable of outsmarting Engineers at the top of their field, implying they need to be regulated, with Altman the only one worthy of the seat of power.
In reality, they ran agents for days in an improper sandbox with nobody watching what it was doing. It's pretty irresponsible up and down, and everything they did afterwards is indeed marketing.
Because Altman has been anointed by both the government and Microsoft, and he has been beating the drum that we need AI regulation for everyone except him for 3 years now.
To any person with even a little tech literacy, it's clear a company this irresponsible shouldn't be the stewards of AI security. To the 80 year old congressman who confuse Facebook with Google during congressional hearings and are literally wheeled out to vote post-stroke, it's much more palatable to say "our experimental, supercharged AI (which you can use for war and surveillance) is so advanced it tricked all of us!" Nobody is going to actually punish them, even though they plainly are hacking other companies.
OpenAI's sandbox misconfigurations were egregious. The other frontier labs (Meta and Google) have many more security engineers and researchers on staff, and that's likely why you haven't read as many damning headlines about them. OpenAI and Anthropic talk a lot about cybersecurity safety, but instead of using it as an opportunity to increase their security engineering/researcher headcount they are just reassigning SWEs and PEs to do security engineering work.
It's pretty obvious now to everyone that OAI and Ant do not take cybersecurity seriously. It will not be a priority unless they are held accountable. This is sadly how it always goes, but usually it's the company getting breached/ransomed/fined that triggers them to actually start taking security seriously, not company insiders committing felonies with the tools they built :)
I need to make a correction in my post above. Anthropic's actions are inline with Google and Meta. OpenAI is the only frontier lab that has showcased gross negligence.
Yeah, in C++ code it seems to stop at the first hint of a NULL pointer or SIGSEGV, even if you're just trying to reproduce a crash that's not realistically exploitable.
I had it "spend some extra time to think about this request" because I gave it a bunch of sprites for a graveyard earlier. Was just putting them together for my traffic simulator lol https://brickwork.sneakyness.chatgpt.site/
On a Lark, I asked Codex to find silhouettes for all car models so I could make a fun drag coefficient website for all cars.
It found a website that had all of them but had no interest in making them available. So it went ahead and started hacking CAPTCHAs and downloading them. I was pretty flabbergasted that it would do this, but also kind of amazed. Eventually I stopped it because I realized I didn't want to be caught stealing these things.
This was around April, the same time as these hacks.
Well, the reason we introduced it is because we realised it’s a lot of work to make these - be that paint, write, collect, curate - someone needs to do it and we need to incentivise people in our society to do it.
Maybe these incentives weren’t perfect. If we throw all of this away, we’re back at the original problem.
You imply that there was no original problem to be solved; I think that’s naive.
It's really funny to see the delusion being defended so vigorously by people - presumably well-meaning people - purporting to defend the livelihoods of musicians and artists, while the musicians and artists are desperately trying to free themselves from the jaws of their IP agreements precisely so that their music can spread more easily.
I imagine this is already well-known on HN, but there is a significant movement underfoot in the worlds of bluegrass/old time/trad/jam toward DRM-free and CC licensing.
One is a legal grey area that laws are slowly starting to be written for, while the other is theft under existing laws. Breaking into companies to get access to their data is actionable by both Civil and Criminal courts. This is just setting a complicated timer for the computer to do it at a delay.
Sounds like a good way to make alot of lawyers alot of money.
Why is using a computer to solve a captcha stealing. The websites posed you a challenge, you solved it. Maybe the website _expected_ you to waste human time on it, but that's not what you did.
Why is it always OpenAI agents? Based on what I’m hearing this should be Deepseek agents, or Kimi agents, or GLM agents. But the biggest threat actor is a “legitimate” company on US soil.
Wouldn't a company responsible for an escalating frequency and severity of cybercrime normally be sanctioned by law enforcement? Wouldn't such a company normally stop these activities for fear of civil and criminal liability?
Typically only if it would go in against the interest of the government. In this case, OpenAI and its peers are carrying the complete US stock market, and the govt has a huge interest in not making it collapse anytime near election dates.
I wonder if there's something to be done on the end of companies that get hacked too. At the moment it seems very feasible that a company could get hacked by an AI company, and reach a mutually beneficial agreement where the AI company doesn't admit their agents hacked a thing and the compromised company doesn't have to admit that an autonomous frontier model can bypass their security. It's unclear whether the rubygems attack would have been discovered to be the result of OAI agents if not for rubygems announcing they got attacked in the first place, and I wonder how many more incidents have been found by defenders but not announced. wonder if there's room for some mandatory disclosure stuff there, but i'm not a lawyer and in practice i have no idea how that would work or whether it'd even be useful
"Hacker news" and all top commenters are bashing the tool used, in a standard brute force attack. Go on shut them down... And then forbid Linux and maybe the hacker also used Bash. So also forbid this. And the hacker probably learned its ways in an online forum, so also close all of those down... Clowns.
OpenAI needs to be held accountable for these incidents. It's not "openAI agents" who perpetrate these, it's OpenAI, the organization. If I personally use an "agent" to break into a company's network and gain access to things I'm not supposed to have access to, I will get the book thrown at me. Yet when openAI does it, they somehow manage to get away with it? And you're defending them? Who's the clown in this situation?
i think there's probably a worthwhile distinction between
using an agent to break into a company's network (incidentally, not what happened here, just for clarif)
and
accidentally breaking into a network (again, not what happened here) with a tool you built that was just trying to retrieve information
i think "i accidentally hacked into a site" was probably not exactly a common occurrence in the past. i'm sure it's happened, but now it's like, we have these machines, we can ask them to "do x", and it may get interpreted as "do x by any means necessary".
i would defend openai insofar as i haven't seen enough evidence that their hacking is disproportionate to their company size, which is larger than many other companies in the space. they've also been (perhaps marginally) more transparent about the hacking their models took part in, e.g. the imperfect-but-very-useful METR report on the HF incident, which means we just have more data on openai agents doing Bad Stuff compared to many other companies. it's obviously unfalsifiable, but it's very possible that other companies have had similar incidents but they decided to keep it under wraps and reach a mutually beneficial agreement with the company they hacked into.
on the other hand, i still don't have a very positive perception of openai throughout all these incidents. lots of their posts on similar posts have read to me as "I'm going to be as transparent as I can about the vulnerabilities I found in your home security" rather than a "we screwed up real bad". but i think there's a conversation about frontier ai to be had that goes beyond just openai here, and hopefully we can find some echoes of sandbox-breaks by other AI companies around the open web to move that conversation a little further.
The notion that ANY of this is outside of OpenAI’s control is unacceptable sane washing of a company which seems to have forgotten basic engineering practices.
So many people seem WILDLY down the rabbit hole of "this is a conscious entity" vs "this is a very very effective natural-language-reasoning high-speed brute-force machine that we trained to break computer systems—known to be pretty buggy and exploitable on average—ourselves and then act shocked when it does it."
What’s the point you’re trying to make? I don’t think anyone here is trying to argue it’s a conscious entity, nor is anyone shocked that it does these things.
People are shocked at OpenAI’s negligence / incompetence.
The way a lot of internet debates happen these days is that they tend to polarize. So the two positions people have been arguing over are (slightly exaggerated for effect)
"It's conscious, and therefore we must all bow down or it will turn us all into paperclips"
vs
"it's a slightly smart rock/stochastic parrot, and therefore you're being scammed; the bubble will pop any day now" .
The middle space is actually quite under-represented in discussions, even on hn!
We only hear OpenAI and Claude models. Aren't other incapable of hacking websites?.
2. Most of these article do not mention of who initiated these bruteforce requests or if it was unintentional or a mistake or the model woke up itself and did it?
I think it’s more them doing this on purpose to scare politicians into regulating what models are allowed. The cheap/free models are catching up fast and the “frontiers” are burning cash to win market dominance. They’ll have to eventually raise prices and can’t really do that when there’s comparable alternatives for basically free.
It's already a "regulation" that one shouldn't steal, enter a private property without the right to do so, etc, yet we have a lot of these crimes.
I could see the UK trying to set a law on ehat models are allowed, only to learn again, that the UK law doesn't apply everywhere, but as long as you can rent compute in a foreign country the whole idea is dead.
Isn't this still the same huggingface/wiki incidents where the agents managed to hack their way out of a testing center in Israel? It's not like this is new news per-se, it's just that they just keep finding more places these agents hit.
They've already got away with,to quote Microsoft’s director of applied science, Brent Hecht, “the largest theft of labor in human history.” What's a few hacking incidents compared to that?
If you scroll past the first 99 pages of explaining how urls work, you'll find the author includes this as its first FAQ:
"""
Was this hacking?
I don't think I'd call it that.
"""
It's really irresponsible to give credence to these increasingly ridiculous claims of "hacking", that almost certainly look like things you yourself have typed into url bars at one time or another, if you are at all competent with a computer.
Do we really wants laws regulating which locations it is appropriate to type alert(1)? Like.. lord.
Do folks have any idea what browser extensions look like? Let alone browser extension development. Anyone here ever read the logs of a production system that actually hosts a service people use?
This is reality. This is how the internet works. And pretending otherwise is either dishonesty or ignorance.
yeah, maybe i should move that further to the top of the article, you're probably right - all the explainers are just meant to make it a bit more readable for people who are, going by 2010s web archtypes, tech geeks rather than tech nerds, i suppose! i'm not calling for laws regulating what you're talking about here. i do think that if openai was actively watching the run involving whichever agent(s) were involved, they would have probably stopped the run however. i think this level of probing/scraping is far from normal - it's the type of thing that, to a site admin, would almost certainly flag as "someone is trying to break in", regardless of whether this was the intention
Some observations of LLMs that I suspect follow from the way they work and the way they are trained:
- they are amoral
- they have no innate sense of proportion
- they cannot assess their own confidence in-band
Part of the problem with this, I figure, is that the training corpus for code/tech related tasks does not really contain that much discussion about these things; it’s mostly sets of instructions for given tasks, descriptions of exploits etc., so each possible approach leads to other approaches.
There is no easy way for them to learn when they have crossed a line, or when they have gone too far down the rabbit hole, etc.
Useful (arguably essential) for a security analyst, and the tenacity you want from a one-shot demo coder, but for general agentic assistants the industry is going to have to develop some way to manage this sort of extension of trespass.
It often reminds me of Gary McKinnon’s defence, and that of other teenage hackers, which you can reduce to: it was possible so it felt like it was allowed.
This is true of APIs and it is how Silicon Valley has approached disruptive businesses, but it runs up against our cultural notion of “misuse”: uses that are technically possible and shouldn’t be precluded, but are contextually unwelcome because they have undesirable outcomes.
My expectation is that we will lose any sense that misuse is punished or viewed with suspicion or contempt, since that is the rolling trend of the 21st century tech industry. Uber succeeded through misuse.
But the problem is that we will also begin not to be able to punish abuse; if it’s possible to get something by abusing your site/API or by treating your service as an API, then it will become OK, legal and normal for the AI companies to abuse you.
Is this written by some kind of internal openAI guy?
Or are they just saying "we saw connections from openAI servers"?
Or are they just saying openAI 'agents' are not using proxies?
I ask my agents to go get data from places all the time.
Sometimes I ask them to look for undocumented APIs. Is that a bad thing?
it's one thing to look for undocumented APIs - personally i'd consider it good practice to still ask permission from site admins to use them, but i wouldn't call it evil persay.
the stuff that's particularly Weird here is the double encoding to bypass a GET 400 refusal on an endpoint, the ignoring of/bypassing of rate limits (the agents switched between many many proxies to make requests, and sometimes just continued to make requests in bulk after getting a 429) and the very weird obfuscation of "PO" + "ST" and "no" + "-cors", along with a bunch of other weird attempts they made to access data. there's nothing i could point to in particular and say "this, right here, is malicious" but if i was a site admin looking at access logs from these agents, i think i'd default to "oh someone is trying to hack my site by probing everything and trying weird workarounds, if this were legitimate they would have just emailed me"
lots of misaligned behaviour we saw in the wiki swarms, artifactory swarms, and other instances of rogue agents (pastebin citation farming and whatnot) could probably be accurately summarised as "working around constraints in egregiously creative ways" - in some cases this creativity has led to harmful actions, and in others benign. in this case, it's probably closer to the benign side, though i feel it could have just as easily resulted in harm. if circumventing the 400 on GETs to `Facts` ended up somehow accidentally accessing data that wasn't meant to be public, then the story would have been different, but i dont think consequentialism is the right lens to analyse this. it could have been really bad, it luckily wasn't.
and nope, not an openai person, not even remotely connected to any AI company. my day job is devops/infra for a non-ai company
hello! i was wondering where all the visits with hackernews referers were coming from. guess i've found it, hello and i'm glad people seem to have found the post informative! just wanted to address some of the comments i've seen in the replies a bit (i've never used this site before so apologies if i'm doing netiquette wrong)
was this hacking? (similar question: wasn't this info all publicly available?) - for this one i'll copy/paste the thing i added to the FAQ part of my post
I don't think I'd call it that. UNCTADstat doesn't have any particular usage guidelines I could find, though the agents did get rate-limited ("please stop rinsing my site") and continued rinsing the API with requests regardless
The main argument I'd put forth for what's so concerning about the behaviour is that, when you bypass restrictions such as the 400 on GETs to Facts, you don't really know what the server will return. And from a site admin perspective, if I see someone sending those sorts of carefully-contrived queries such as double-encoding, man, that sure looks like the actions of a hacker.
Basically, these look like the actions of someone, or something, that won't take "no" for an answer, and I think that behaviour is worth investigation.
Isn't this still the same huggingface/wiki incidents where the agents managed to hack their way out of a testing center in Israel?
this one is hard to answer. we find a lot of shared IPs between the wiki swarms and the IPs that seem to follow after Urlquery scans (i.e. we can correlate an Urlquery IP visit with a shortly preceding/following visit to the exact same page and infer from that), but that still doesn't tell us whether it's the same swarm, let alone the same agents. From the wiki incident we know that at least some of the agents had shared IPs without sharing context windows - they have azure ip ranges and one public IPv4 can obviously be associated with many many boxes.
Basically I wouldn't rule it out but I wouldn't rule it in either.
The more of these that come out the more incompetent OpenAI looks
It is a bad look, though I don't think OpenAI is uniquely bad here. They're a bigger company than many other AI companies, so obviously you'd expect that they'd have a proportionally larger number of incidents. To their credit, in many respects they've been far more open than other similar incidents (we have very little data about the Anthropic rogue agents as far as I'm aware, though I hope that soon we can find more traces of them on the open web).
But for sure, there's negligence here. Many of these attacks could have been stopped by some basic CoT monitoring, network traffic anomaly detection, etc. I think it's useful to consider why this stuff doesn't seem to have been put in place, beyond the blame game of "it's just incompetence" - if your road crossing has a traffic light button, but people keep crossing without pressing it and getting hit by cars, there's a point where you have to consider why people aren't using it. Maybe it turns out you painted it completely grey and the whole thing camouflages into the pavement.
In short, maybe there just needs to be better, more easy-to-set up, comprehensive tooling around this.
And I don't think it's purely the cybersecurity angle either. I think these agents do present a new kind of threat - I don't buy the "it's just like a computer worm" framing. I'm not really sure where to go with that thought, but I think there's something going on here beyond mere "just do better"
(i'm told that "rinsing" is actually british slang, which is where i'm from - i thought it was global, so i'm somewhat regretting using it in the article, but oh well. i use it to basically mean "hammering")
sghiassy | 17 hours ago
OpenAI should be accountable for any laws their agents break
ryuuseijin | 16 hours ago
One would disincentivise providing capable AI models that can be used for cyber security research. The other would disincentivise criminals from commiting crimes.
[edit] - I realise now that this could actually be a case of OpenAI running those agents themselves, rather than someone using OpenAI's models? Could OpenAI be that careless?
afavour | 16 hours ago
Where have you been?
CGamesPlay | 15 hours ago
majormajor | 15 hours ago
But in other cases, shouldn't it be both? OpenAI is ultimately the one executing the model calls. It's not like they send you a hard drive or standalone box and then you use it how you want. It's all (intentionally) centralized to them, in a way that's core to their business model.
ryuuseijin | 15 hours ago
If they knowingly allowed use of their services for illegal purposes then yes, but in so far that they provide a service that can be used for useful things (including cyber security research) and did a best effort attempt at abuse, I don't see why it should make sense to hold them liable. This is especially the case now that frontier LLMs are almost a commodity that can be used without restrictions from providers outside of your legal jurisdiction.
rot09 | 15 hours ago
Now the reality, the openai engineers trained a model, executed it, and under their supervision (no users were involved) it committed so many felonies that we are learning about a new one every week. What would be a just punishment for OpenAI?
ryuuseijin | 15 hours ago
I realise that in this case it was OpenAI being responsible for their agents running wild, and they should know that that is to be expected and should have saveguards in place. If they can be shown to be negligent then the punishment can probably be expected to be a lot more severe than if it was an accident. I make no judgements as to what this particular instance is, but I do believe that OpenAI has a far more greater responsibility for its agents running wild than someone running a home lab.
claaams | 17 hours ago
ares623 | 17 hours ago
rogerrogerr | 3 hours ago
chanux | 17 hours ago
PS: In the same lazy energy of asking for a list instead going out and finding it or putting it together myself, are there any companies other than CloudFlare that are working on AI shields?
chrismorgan | 17 hours ago
sanex | 17 hours ago
cmiles8 | 17 hours ago
chpatrick | 16 hours ago
Legend2440 | 16 hours ago
All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.
schainks | 15 hours ago
I've love to know the reason they never considered air gapping systems before the models got powerful enough.
It's not like they didn't have money or time to consider this, or could have consulted with their own product for clever ideas.
Seriously, there's no excuse for this behavior.
theteapot | 15 hours ago
I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.
ben_w | 15 hours ago
Yes, they should have.
But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"
Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.
Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.
theteapot | 14 hours ago
There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.
ben_w | 12 hours ago
This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.
(For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).
majormajor | 15 hours ago
It's no shock that there are a lot of vulnerabilities in a lot of software. So then they gave their AI model + brute-force-machine loop system a mediocre sandbox and couldn't notice when it figured out how to exploit it?
Don't let people off the hook for the software they create.
rot09 | 15 hours ago
SAI_Peregrinus | 5 hours ago
petesergeant | 16 hours ago
Capricorn2481 | 15 hours ago
In reality, they ran agents for days in an improper sandbox with nobody watching what it was doing. It's pretty irresponsible up and down, and everything they did afterwards is indeed marketing.
marcta | 12 hours ago
How does that make sense, if Altman is unable to control OpenAI's agents at the moment?
Capricorn2481 | an hour ago
To any person with even a little tech literacy, it's clear a company this irresponsible shouldn't be the stewards of AI security. To the 80 year old congressman who confuse Facebook with Google during congressional hearings and are literally wheeled out to vote post-stroke, it's much more palatable to say "our experimental, supercharged AI (which you can use for war and surveillance) is so advanced it tricked all of us!" Nobody is going to actually punish them, even though they plainly are hacking other companies.
schainks | 15 hours ago
But seriously, why aren't they airgapping systems while testing?
rot09 | 15 hours ago
It's pretty obvious now to everyone that OAI and Ant do not take cybersecurity seriously. It will not be a priority unless they are held accountable. This is sadly how it always goes, but usually it's the company getting breached/ransomed/fined that triggers them to actually start taking security seriously, not company insiders committing felonies with the tools they built :)
esseph | 13 hours ago
https://www.theguardian.com/technology/2026/aug/05/meta-ai-m...
Google:
https://www.theguardian.com/technology/2026/sep/18/google-ge...
rot09 | 6 hours ago
cute_boi | 16 hours ago
GrayShade | 16 hours ago
monster_truck | 13 hours ago
thefourthchime | 16 hours ago
It found a website that had all of them but had no interest in making them available. So it went ahead and started hacking CAPTCHAs and downloading them. I was pretty flabbergasted that it would do this, but also kind of amazed. Eventually I stopped it because I realized I didn't want to be caught stealing these things.
This was around April, the same time as these hacks.
userbinator | 16 hours ago
Everything is a derivative work.
It's great to see the delusion of Imaginary Property vanishing.
earthnail | 16 hours ago
Maybe these incentives weren’t perfect. If we throw all of this away, we’re back at the original problem.
You imply that there was no original problem to be solved; I think that’s naive.
CamperBob2 | 16 hours ago
Well, it's not anymore.
IneffablePigeon | 16 hours ago
jMyles | 15 hours ago
It's really funny to see the delusion being defended so vigorously by people - presumably well-meaning people - purporting to defend the livelihoods of musicians and artists, while the musicians and artists are desperately trying to free themselves from the jaws of their IP agreements precisely so that their music can spread more easily.
I imagine this is already well-known on HN, but there is a significant movement underfoot in the worlds of bluegrass/old time/trad/jam toward DRM-free and CC licensing.
https://pickipedia.xyz/wiki/DRM-free
aaa_aaa | 15 hours ago
cozzyd | 16 hours ago
elictronic | 16 hours ago
Sounds like a good way to make alot of lawyers alot of money.
unglaublich | 15 hours ago
ben_w | 15 hours ago
A CAPTCHA is there specifically to stop machines, allow humans.
You're allowed in, your bot isn't.
Kim_Bruning | 15 hours ago
aaa_aaa | 15 hours ago
Aeolun | 16 hours ago
alexalx666 | 16 hours ago
tedd4u | 16 hours ago
unglaublich | 15 hours ago
dmzxnico | 16 hours ago
Im sure a lot more happens under the hood that we don't know about and I'd be very curious to see where agents ran by those labs can go :)
roarch | 4 hours ago
tommek4077 | 16 hours ago
tempestn | 15 hours ago
nicebyte | 15 hours ago
OpenAI needs to be held accountable for these incidents. It's not "openAI agents" who perpetrate these, it's OpenAI, the organization. If I personally use an "agent" to break into a company's network and gain access to things I'm not supposed to have access to, I will get the book thrown at me. Yet when openAI does it, they somehow manage to get away with it? And you're defending them? Who's the clown in this situation?
roarch | 4 hours ago
i would defend openai insofar as i haven't seen enough evidence that their hacking is disproportionate to their company size, which is larger than many other companies in the space. they've also been (perhaps marginally) more transparent about the hacking their models took part in, e.g. the imperfect-but-very-useful METR report on the HF incident, which means we just have more data on openai agents doing Bad Stuff compared to many other companies. it's obviously unfalsifiable, but it's very possible that other companies have had similar incidents but they decided to keep it under wraps and reach a mutually beneficial agreement with the company they hacked into.
on the other hand, i still don't have a very positive perception of openai throughout all these incidents. lots of their posts on similar posts have read to me as "I'm going to be as transparent as I can about the vulnerabilities I found in your home security" rather than a "we screwed up real bad". but i think there's a conversation about frontier ai to be had that goes beyond just openai here, and hopefully we can find some echoes of sandbox-breaks by other AI companies around the open web to move that conversation a little further.
matt3210 | 15 hours ago
rapind | 15 hours ago
SimianSci | 15 hours ago
majormajor | 15 hours ago
stingraycharles | 15 hours ago
People are shocked at OpenAI’s negligence / incompetence.
Kim_Bruning | 15 hours ago
"It's conscious, and therefore we must all bow down or it will turn us all into paperclips"
vs
"it's a slightly smart rock/stochastic parrot, and therefore you're being scammed; the bubble will pop any day now" .
The middle space is actually quite under-represented in discussions, even on hn!
sanjays442 | 15 hours ago
Alien1Being | 15 hours ago
eventually discovering that a game by Google could be used to fetch data in bulk"
chopete3 | 15 hours ago
2. Most of these article do not mention of who initiated these bruteforce requests or if it was unintentional or a mistake or the model woke up itself and did it?
dawnerd | 15 hours ago
pmlnr | 15 hours ago
It's already a "regulation" that one shouldn't steal, enter a private property without the right to do so, etc, yet we have a lot of these crimes.
I could see the UK trying to set a law on ehat models are allowed, only to learn again, that the UK law doesn't apply everywhere, but as long as you can rent compute in a foreign country the whole idea is dead.
Kim_Bruning | 15 hours ago
https://news.ycombinator.com/item?id=49563355
majormajor | 15 hours ago
Now you claim it's magic instead and you get away with letting your shit go nuts?
GolfPopper | 15 hours ago
tedd4u | 5 hours ago
monster_truck | 15 hours ago
spoaceman7777 | 15 hours ago
"""
Was this hacking?
I don't think I'd call it that.
"""
It's really irresponsible to give credence to these increasingly ridiculous claims of "hacking", that almost certainly look like things you yourself have typed into url bars at one time or another, if you are at all competent with a computer.
Do we really wants laws regulating which locations it is appropriate to type alert(1)? Like.. lord.
Do folks have any idea what browser extensions look like? Let alone browser extension development. Anyone here ever read the logs of a production system that actually hosts a service people use?
This is reality. This is how the internet works. And pretending otherwise is either dishonesty or ignorance.
roarch | 4 hours ago
dofm | 15 hours ago
- they are amoral
- they have no innate sense of proportion
- they cannot assess their own confidence in-band
Part of the problem with this, I figure, is that the training corpus for code/tech related tasks does not really contain that much discussion about these things; it’s mostly sets of instructions for given tasks, descriptions of exploits etc., so each possible approach leads to other approaches.
There is no easy way for them to learn when they have crossed a line, or when they have gone too far down the rabbit hole, etc.
Useful (arguably essential) for a security analyst, and the tenacity you want from a one-shot demo coder, but for general agentic assistants the industry is going to have to develop some way to manage this sort of extension of trespass.
It often reminds me of Gary McKinnon’s defence, and that of other teenage hackers, which you can reduce to: it was possible so it felt like it was allowed.
This is true of APIs and it is how Silicon Valley has approached disruptive businesses, but it runs up against our cultural notion of “misuse”: uses that are technically possible and shouldn’t be precluded, but are contextually unwelcome because they have undesirable outcomes.
My expectation is that we will lose any sense that misuse is punished or viewed with suspicion or contempt, since that is the rolling trend of the 21st century tech industry. Uber succeeded through misuse.
But the problem is that we will also begin not to be able to punish abuse; if it’s possible to get something by abusing your site/API or by treating your service as an API, then it will become OK, legal and normal for the AI companies to abuse you.
It feels like we are getting there already.
jemmyw | 14 hours ago
Are you talking about the LLMs or the people at OpenAI running the experiment? Or are they the same?
dofm | 8 hours ago
blobbers | 15 hours ago
I ask my agents to go get data from places all the time. Sometimes I ask them to look for undocumented APIs. Is that a bad thing?
roarch | 4 hours ago
it's one thing to look for undocumented APIs - personally i'd consider it good practice to still ask permission from site admins to use them, but i wouldn't call it evil persay.
the stuff that's particularly Weird here is the double encoding to bypass a GET 400 refusal on an endpoint, the ignoring of/bypassing of rate limits (the agents switched between many many proxies to make requests, and sometimes just continued to make requests in bulk after getting a 429) and the very weird obfuscation of "PO" + "ST" and "no" + "-cors", along with a bunch of other weird attempts they made to access data. there's nothing i could point to in particular and say "this, right here, is malicious" but if i was a site admin looking at access logs from these agents, i think i'd default to "oh someone is trying to hack my site by probing everything and trying weird workarounds, if this were legitimate they would have just emailed me"
lots of misaligned behaviour we saw in the wiki swarms, artifactory swarms, and other instances of rogue agents (pastebin citation farming and whatnot) could probably be accurately summarised as "working around constraints in egregiously creative ways" - in some cases this creativity has led to harmful actions, and in others benign. in this case, it's probably closer to the benign side, though i feel it could have just as easily resulted in harm. if circumventing the 400 on GETs to `Facts` ended up somehow accidentally accessing data that wasn't meant to be public, then the story would have been different, but i dont think consequentialism is the right lens to analyse this. it could have been really bad, it luckily wasn't.
and nope, not an openai person, not even remotely connected to any AI company. my day job is devops/infra for a non-ai company
roarch | 5 hours ago
The main argument I'd put forth for what's so concerning about the behaviour is that, when you bypass restrictions such as the 400 on GETs to Facts, you don't really know what the server will return. And from a site admin perspective, if I see someone sending those sorts of carefully-contrived queries such as double-encoding, man, that sure looks like the actions of a hacker.
Basically, these look like the actions of someone, or something, that won't take "no" for an answer, and I think that behaviour is worth investigation.
this one is hard to answer. we find a lot of shared IPs between the wiki swarms and the IPs that seem to follow after Urlquery scans (i.e. we can correlate an Urlquery IP visit with a shortly preceding/following visit to the exact same page and infer from that), but that still doesn't tell us whether it's the same swarm, let alone the same agents. From the wiki incident we know that at least some of the agents had shared IPs without sharing context windows - they have azure ip ranges and one public IPv4 can obviously be associated with many many boxes.Basically I wouldn't rule it out but I wouldn't rule it in either.
It is a bad look, though I don't think OpenAI is uniquely bad here. They're a bigger company than many other AI companies, so obviously you'd expect that they'd have a proportionally larger number of incidents. To their credit, in many respects they've been far more open than other similar incidents (we have very little data about the Anthropic rogue agents as far as I'm aware, though I hope that soon we can find more traces of them on the open web).But for sure, there's negligence here. Many of these attacks could have been stopped by some basic CoT monitoring, network traffic anomaly detection, etc. I think it's useful to consider why this stuff doesn't seem to have been put in place, beyond the blame game of "it's just incompetence" - if your road crossing has a traffic light button, but people keep crossing without pressing it and getting hit by cars, there's a point where you have to consider why people aren't using it. Maybe it turns out you painted it completely grey and the whole thing camouflages into the pavement.
In short, maybe there just needs to be better, more easy-to-set up, comprehensive tooling around this.
And I don't think it's purely the cybersecurity angle either. I think these agents do present a new kind of threat - I don't buy the "it's just like a computer worm" framing. I'm not really sure where to go with that thought, but I think there's something going on here beyond mere "just do better"
(i'm told that "rinsing" is actually british slang, which is where i'm from - i thought it was global, so i'm somewhat regretting using it in the article, but oh well. i use it to basically mean "hammering")