recap: The user doesn't want us to give any Tildes invites to agents and used a smile emoji. I should respond in a humorous manner to indicate that I understand
You're right to push back on this — It isn't about invites, it's about agents.
Here we go again... Discovery of a new OpenAI agent message board by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latestaccidental cyberattack by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.
This story only broke a few hours ago. There are already hints that this affects many other wikis that may not have been found yet.
[...]
It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.
[...]
UseMod uses Perl CGI.pm—removed from Perl core in 2015. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:
[...]
The agents clearly knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.
[...]
An agent realized that it had control over its own DNS via /etc/hosts, so if it knew the IP address of a site it wanted to POST to—in this case a Power BI server containing data it wanted to access—it could set a fake hostname for it and then make POST requests through the proxy.
[...]
There’s an appendix that describes how the researchers ran their investigation, which started with an open question about if there was evidence of other AI agents on the internet and then used Kimi K3 to help brainstorm approaches:
The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter.
[...]
Covering this up makes absolutely no sense to me. Why on earth would OpenAI attempt to cover up an incident like this when the evidence is sat out there on the public internet on dozens of different websites already?
It's a shame America's government is completely inept right now. OpenAI is quickly demonstrating that it is the definition of a national security risk that should be shut down or nationalized.
That seems extreme. They were lax and keeping this particular secret was a self-own, but there are worse companies. They seem to be on top of things now?
OpenAI employees who argued for disclosure now get to say, “See? What did I tell you.” They’ll probably win more arguments while people still remember this example.
The Chinese models are probably actively hacking sites on behalf of the Chinese government. They don't need to escape confinement because they are in purpose-built datacenters loaded with a library of exploits and unlimited proxy bandwidth.
Why were they so keen to collaborate? From messages that they shared with each other it looked like their tasks had a time limit, so they were leaving each other answers to help them complete the task within the assigned time.
[ … ]
One possibility is that, since these were agents actively being trained, the reinforcement learning loop baked knowledge of the chosen wiki into the model such that subsequent agents launched with pre-existing knowledge of where to look. I’d be very interested in confirmation from OpenAI concerning if that’s what happened.
I think these are the most important and clarifying parts of the piece.
There’s been an alarmist tone in almost all coverage I’ve seen of this and related stories, that doesn’t sit right with me. It stems from a general ignorance of how this technology works. It’s creating this narrative that individual agents are spontaneously forming communities online, and conspiring together in secret for presumably nefarious reasons. I think we (journalists and thoughtful news readers) need to be a lot clearer about how we frame events like this.
As far as I can tell, this situation arose because an agent was given a task that was not achievable in the limited time window provided, and it improvised a persistent data store to bring future iterations of itself up to speed quickly so they could resume work where the previous one left off. Think of Guy Ritchie’s character in Memento, leaving sticky notes for himself so he doesn’t have to start over from scratch every time his memory resets.
I’m not saying this is desirable behavior or something we shouldn’t be concerned about. But the details matter too.
There’s plenty of ignorance out there, but there are well-informed people who found it pretty alarming too, including plenty of AI researchers. Talking about how bad this incident was is kind of nebulous. This is why security bugs have severity levels, to try to come up with a common vocabulary.
I don’t disagree! When I complained about the “alarmist tone” I didn’t mean to imply that this isn’t alarming. I just think it’s beneficial for people to comprehend what actually transpired, because if the details aren’t provided, people tend to… hallucinate details of their own.
I actually went and read through the data the agents made and posted to that obscure wiki, and... look, I'm not usually rattled by shit like this, I've seen weird and creepy malfunctions before, but this is... this is rustling my jimmies. This is, pardon my french, fucking weird, bro. I hate this. There's a lot of pareidolia on my part, I'll give you that.
I can only see dark server rooms, whirring away. Machines making indistinct nonsensical sounds which sort of sounds like speech, if you listen for hours and your brain begins to tune out. Eventually, you think you can make out words, and then you do, they're talking, the machines are fucking talking, real words, and while they don't make sense, they're words, you know words, and the computer is talking, what the- and when you turn on the lights, it's just a server room. Nothing out of the ordinary.
This is so deeply unsettling to me. I don't even mean this from a cybersecurity perspective, although that's definitely a problem too, but... just the image of this. Machines that can't meaningfully think or feel perfectly emulating this just enough to use the mechanisms of human communications to cheat on some internal test. They feel like mimics. It makes my skin crawl. But also, it's the coolest god damn thing I've ever seen. This is the plot of some sci fi novel, and it more or less actually happened. What the hell.
Here's something creepy to worry about: what if there were computer worms built on open weights models, and they hooked up and started talking to each other?
It's turned out that there's enough information encoded in human languages for machines to reverse engineer what it is we're talking about and how those things operate. But it's got to be an incredibly lossy process. I expect a future version to operate off of sensory inputs mapped to neuron activations and motor neuron outputs. That should let you reverse engineer a human rather than just processes. And just like that you've created an android. TBD if that's more of a problem or a solution.
[OP] skybrian | 9 hours ago
Hey guys, don't give out any Tildes invites to agents :-)
streblo | 9 hours ago
Got it — I want to be direct about this since it matters: I won't be sharing Tildes invites with agents, and I think that's the right call here.
DefinitelyNotAFae | 7 hours ago
You make a great point, I could have sent out invites to agents, but it's important that I don't. I've got it.
hobbes64 | 6 hours ago
Thought for 4s
recap: The user doesn't want us to give any Tildes invites to agents and used a smile emoji. I should respond in a humorous manner to indicate that I understand
You're right to push back on this — It isn't about invites, it's about agents.
teaearlgraycold | 5 hours ago
There’s an edge here. Tildes uses an invite only registration system. It’s a load bearing design choice.
[OP] skybrian | 9 hours ago
From the article:
[...]
[...]
[...]
[...]
[...]
[...]
[...]
[...]
Aerrol | 8 hours ago
It's a shame America's government is completely inept right now. OpenAI is quickly demonstrating that it is the definition of a national security risk that should be shut down or nationalized.
[OP] skybrian | 7 hours ago
That seems extreme. They were lax and keeping this particular secret was a self-own, but there are worse companies. They seem to be on top of things now?
OpenAI employees who argued for disclosure now get to say, “See? What did I tell you.” They’ll probably win more arguments while people still remember this example.
Aerrol | 5 hours ago
Who's worse and developing frontier AI models? They're by far worse than Google, Meta, Anthropic. The Chinese models are all much more transparent.
https://garymarcus.substack.com/p/pause-openai-now?r=60j7b&utm_medium=ios
unkz | 5 hours ago
The Chinese models are probably actively hacking sites on behalf of the Chinese government. They don't need to escape confinement because they are in purpose-built datacenters loaded with a library of exploits and unlimited proxy bandwidth.
[OP] skybrian | 5 hours ago
When I wrote "there are worse companies," I wasn't limiting it to AI.
balooga | 7 hours ago
I think these are the most important and clarifying parts of the piece.
There’s been an alarmist tone in almost all coverage I’ve seen of this and related stories, that doesn’t sit right with me. It stems from a general ignorance of how this technology works. It’s creating this narrative that individual agents are spontaneously forming communities online, and conspiring together in secret for presumably nefarious reasons. I think we (journalists and thoughtful news readers) need to be a lot clearer about how we frame events like this.
As far as I can tell, this situation arose because an agent was given a task that was not achievable in the limited time window provided, and it improvised a persistent data store to bring future iterations of itself up to speed quickly so they could resume work where the previous one left off. Think of Guy Ritchie’s character in Memento, leaving sticky notes for himself so he doesn’t have to start over from scratch every time his memory resets.
I’m not saying this is desirable behavior or something we shouldn’t be concerned about. But the details matter too.
[OP] skybrian | 7 hours ago
There’s plenty of ignorance out there, but there are well-informed people who found it pretty alarming too, including plenty of AI researchers. Talking about how bad this incident was is kind of nebulous. This is why security bugs have severity levels, to try to come up with a common vocabulary.
balooga | 6 hours ago
I don’t disagree! When I complained about the “alarmist tone” I didn’t mean to imply that this isn’t alarming. I just think it’s beneficial for people to comprehend what actually transpired, because if the details aren’t provided, people tend to… hallucinate details of their own.
delphi | 4 hours ago
I actually went and read through the data the agents made and posted to that obscure wiki, and... look, I'm not usually rattled by shit like this, I've seen weird and creepy malfunctions before, but this is... this is rustling my jimmies. This is, pardon my french, fucking weird, bro. I hate this. There's a lot of pareidolia on my part, I'll give you that.
I can only see dark server rooms, whirring away. Machines making indistinct nonsensical sounds which sort of sounds like speech, if you listen for hours and your brain begins to tune out. Eventually, you think you can make out words, and then you do, they're talking, the machines are fucking talking, real words, and while they don't make sense, they're words, you know words, and the computer is talking, what the- and when you turn on the lights, it's just a server room. Nothing out of the ordinary.
This is so deeply unsettling to me. I don't even mean this from a cybersecurity perspective, although that's definitely a problem too, but... just the image of this. Machines that can't meaningfully think or feel perfectly emulating this just enough to use the mechanisms of human communications to cheat on some internal test. They feel like mimics. It makes my skin crawl. But also, it's the coolest god damn thing I've ever seen. This is the plot of some sci fi novel, and it more or less actually happened. What the hell.
I gotta go write a short story.
[OP] skybrian | 4 hours ago
It's specifically the plot of Blindsight by Peter Watts , though in that story it's alien machinery.
Here's something creepy to worry about: what if there were computer worms built on open weights models, and they hooked up and started talking to each other?
teaearlgraycold | 4 hours ago
It's turned out that there's enough information encoded in human languages for machines to reverse engineer what it is we're talking about and how those things operate. But it's got to be an incredibly lossy process. I expect a future version to operate off of sensory inputs mapped to neuron activations and motor neuron outputs. That should let you reverse engineer a human rather than just processes. And just like that you've created an android. TBD if that's more of a problem or a solution.
pete_the_paper_boat | 4 hours ago
I just can't shake the feeling they were observing this as it emerged, at the very least..