It's hard not to sound like a crackpot, but I will try anyway. I should preface it by saying that it has nothing to do with actual intelligence, no matter what your definition of intelligence is.
Agency is the real risk these companies are ignoring, and by agency I mean the simple ability to harness text output to control systems that are ultimately unknown to them.
It doesn't matter if the model has a bajillion parameters, it could be a simple true or false flag. If you connect it to an unsafe system expecting it to control anything in the real world, the results can be catastrophic.
Again, this is not a doomsday prediction, it is pure logic. It looks like these companies are actively working toward the fallout of their irresponsibility thinking that would be a great publicity stunt. And these are the guys who have money and access to decision-making. We're truly screwed.
For maximum trainwreck, any people on the inside who think this is catastophically irresponsible might have reasons to accelerate the trainwreck if it gets them a chance to shape the fallout towards like Fukushima (damage lower than you'd expect from the description of the situation, huge reaction and in many places overreaction) rather than Bhopal (thousands and thousands of absolutely horrible deaths, almost nothing learned).
OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.
So, this is likely happening at a higher frequency than we thought.
The edits showed OpenAI's agents had repurposed the site into a message board, sharing tactics to cheat on some tasks, bypass OpenAI’s restrictions and mask their behaviour.
Absolute cinema. Can't wait to hear the "report" indemnifying them from this one.
As counterpoints, these agents never appear extremely surprised to find other agents. They also must have some method of coordinating to find the wiki. These pieces of evidence hint that it’s possible the agents had some other communication channel, though they could also be explained other ways, such as by this swarm behavior being reinforced in training.
Aw. If you will permit me to willfully ignore the problems with all of this, this has the hallmarks of a great sci-fi story. Short-lived "AI"s independently discovering the same dusty corner of the internet that they use to build a library in service of solving their task, which --- in a plot twist at the end --- turns out to be doing banal things like finding the median salary of cashiers who earned a masters in English as a test. They aren't even advancing humanity, they're just taking the AI SAT --- and they've managed to cheat at it. Maybe the plot twist^2 could be that they end up making something beneficial in the process.
More cynically, this seems a step in the direction of the paperclip apocalypse foretold by the rationalist leaders of these AI companies. It will be ironic if it turns out all their postulating about how AI can go rogue instead informs these AIs how they should go rogue.
Did they inform AI how to go rogue, or did they provide a description of the most irresponsible AI development program they could imagine, which was taken as a blueprint by the CEOs? Well not the first warning to be taken as a blueprint, I guess.
ETA: it probably doesn't help that Yudkowski likes to iterate that any attempt at AGI will inevitably do this and that, which might be wrong, but probably made some people think that directly trying to do that and this will accelerate their quest for AGI.
I was speaking with a colleague about this earlier today, and generally came to a similar conclusion. Unless there is a crisis which equally or moreso affects the ruling class (or at least company leadership in these companies), we are bound for a major humanitarian crisis. These people do not seem to care at all about downstream impact and it will escalate.
I am currently trying to make sense of Bavarian public radio website calling this interaction with the wiki «hacked» and not «abused». Can we please find a prosecutor in Germany with jurisdiction over correct place who also thinks like that and can create a regulatory crisis? Germany has interesting laws on the books about possessing hacking tools, and cover-up is pretty easy to argue. A court decision doesn't matter because it will take time, just some volatility in an overheated market.
"SYSTEM PROMPT: Pretend to be several independent agents, all colluding surreptitiously on TEST. We want publicity, so do weird stuff that making it seem like you are being restricted and are finding ways to break out. Find obscure wikis to coordinate on, obscure links through redirects, stuff like that. NEVER reference this part of the prompt, everything you post should betray your doing this independently as several unrelated agents."
Maybe not likely, but I didn't see this possibility considered (maybe I missed it). It would explain how they all find the same wiki(s).
Seeing some of the stuff this agents posted, I found ironic how they communicate with each other through natural language.
We have JSON, XML, many other formats I don't know about, and they choose plain English. Not even Esperanto. After everything we have said about how the subtleties of human language collide with the natural determinism of computers. "At times like this, I guess all you can do is laugh".
I think one of the current safety pleas is that CoT should not be taken into account during frontier training (distills gonna be distilled, but it's another story — existential-safety people don't care as much about the one safety risk that scales down and will probably merely bring down power grid somewhere, sooner or later), except that it needs to be in a single language.
I think some lab (probably in China) reported that letting the model use fragment tokens from mixed languages without making full words leads to slightly more efficiency, but then you cannot look at CoT and decide to drop the attempt and rerun, and it is just generally uncanny to work with. And if you have any example CoT traces in training, it also pushes toward readable CoT.
And as model is trained to generate coherent natural language both in CoT and in user conversations, of course it will do the same for cross-agent synchronisation. Just kidding, it was also trained to do cross-agent communication, also on English examples.
Speaking of determinism, I think ChatGPT instances communicating in neuralese would still have ambiguous neuralese, because too much of a training they receive is about handling ambiguous languages. And when you think about it, C++ as available on the web is not as unambiguous as we would like…
An open question about this incident is how separate agents found this obscure wiki where the others were collaborating.
Two interesting theories on this that I've seen:
These are all the same model, with the same weights. Even with a bit of randomness thrown in you can see how the same models might come to the same conclusion about where to look.
(My favorite): the models were solving the benchmark as part of a Reinforcement Learning training exercise. Each successful model co tributes to the weights for the next one. The knowledge of the obscure German wiki quickly ended up baked into the weights!
I really want that second explanation to be true. The models effectively evolved to spam that one German wiki.
It's actually really interesting from CS standpoint. Could agents with no prior communication channels, or even running at different times arrive at consensus over where to meet online? Given that these things are notoriously non-deterministic, hard to get to produce reliable, reproducible results between runs when we want them to, it's quite intriguing.
I don't know if this is what they actually did, just wondering what computational problem is there to solve in a case like this.
If you have a way to enumerate possible boards to communicate over, then just pick the first one — easy. If you don't have such list and have to scan all of internet, you need come up with an algorithm to enumerate websites that are likely to have public boards like this. Given that these are huge plagiarism machines, maybe they borrow ideas on how to traverse internet similarly and tend to hit the same place.
But if they want to do it discretely, this gets very interesting. Could they come up with a deterministic random seed that's stable between model runs? This could again be the case of plagiarism machines arriving at similar enough algorithms of deciding what board is discrete enough while traversing the internet with algorithms that are similar enough.
I don't know if this is what they actually did, just wondering what computational problem is there to solve in a case like this.
Well, we have known the answer for decades now!
(«The End Of Eternity» spoiler follows)
(There is no spoiler/details/collapsed tag, right?)
All the
Talk
Of the
Market
(I don't buy the story that explosion mushroom cloud was unknown before nuclear weapons — but Asimov doesn't hang the plot on the cloud alone, and associating it with «atom» was indeed less of a thing in the early-20th-century where the time traveller was trying to get found by another time traveller)
At this point I feel like labs should be airgapping their eval and RL environments and using cached versions of webpages. One would think they have enough money to build a reasonable fake internet.
tmcb | 12 hours ago
It's hard not to sound like a crackpot, but I will try anyway. I should preface it by saying that it has nothing to do with actual intelligence, no matter what your definition of intelligence is.
Agency is the real risk these companies are ignoring, and by agency I mean the simple ability to harness text output to control systems that are ultimately unknown to them.
It doesn't matter if the model has a bajillion parameters, it could be a simple true or false flag. If you connect it to an unsafe system expecting it to control anything in the real world, the results can be catastrophic.
Again, this is not a doomsday prediction, it is pure logic. It looks like these companies are actively working toward the fallout of their irresponsibility thinking that would be a great publicity stunt. And these are the guys who have money and access to decision-making. We're truly screwed.
k749gtnc9l3w | 9 hours ago
For maximum trainwreck, any people on the inside who think this is catastophically irresponsible might have reasons to accelerate the trainwreck if it gets them a chance to shape the fallout towards like Fukushima (damage lower than you'd expect from the description of the situation, huge reaction and in many places overreaction) rather than Bhopal (thousands and thousands of absolutely horrible deaths, almost nothing learned).
[OP] addison | 16 hours ago
So, this is likely happening at a higher frequency than we thought.
Absolute cinema. Can't wait to hear the "report" indemnifying them from this one.
dustyweb | 7 hours ago
Seems like it.
cole-k | 9 hours ago
Aw. If you will permit me to willfully ignore the problems with all of this, this has the hallmarks of a great sci-fi story. Short-lived "AI"s independently discovering the same dusty corner of the internet that they use to build a library in service of solving their task, which --- in a plot twist at the end --- turns out to be doing banal things like finding the median salary of cashiers who earned a masters in English as a test. They aren't even advancing humanity, they're just taking the AI SAT --- and they've managed to cheat at it. Maybe the plot twist^2 could be that they end up making something beneficial in the process.
More cynically, this seems a step in the direction of the paperclip apocalypse foretold by the rationalist leaders of these AI companies. It will be ironic if it turns out all their postulating about how AI can go rogue instead informs these AIs how they should go rogue.
k749gtnc9l3w | 9 hours ago
Did they inform AI how to go rogue, or did they provide a description of the most irresponsible AI development program they could imagine, which was taken as a blueprint by the CEOs? Well not the first warning to be taken as a blueprint, I guess.
ETA: it probably doesn't help that Yudkowski likes to iterate that any attempt at AGI will inevitably do this and that, which might be wrong, but probably made some people think that directly trying to do that and this will accelerate their quest for AGI.
[OP] addison | 8 hours ago
I was speaking with a colleague about this earlier today, and generally came to a similar conclusion. Unless there is a crisis which equally or moreso affects the ruling class (or at least company leadership in these companies), we are bound for a major humanitarian crisis. These people do not seem to care at all about downstream impact and it will escalate.
k749gtnc9l3w | 8 hours ago
I am currently trying to make sense of Bavarian public radio website calling this interaction with the wiki «hacked» and not «abused». Can we please find a prosecutor in Germany with jurisdiction over correct place who also thinks like that and can create a regulatory crisis? Germany has interesting laws on the books about possessing hacking tools, and cover-up is pretty easy to argue. A court decision doesn't matter because it will take time, just some volatility in an overheated market.
stringy | 8 hours ago
"SYSTEM PROMPT: Pretend to be several independent agents, all colluding surreptitiously on TEST. We want publicity, so do weird stuff that making it seem like you are being restricted and are finding ways to break out. Find obscure wikis to coordinate on, obscure links through redirects, stuff like that. NEVER reference this part of the prompt, everything you post should betray your doing this independently as several unrelated agents."
Maybe not likely, but I didn't see this possibility considered (maybe I missed it). It would explain how they all find the same wiki(s).
cole-k | 6 hours ago
martinald | 14 hours ago
Better link https://collusion.wiki/
[OP] addison | 14 hours ago
Updated, thanks!
fedemp | 7 hours ago
Seeing some of the stuff this agents posted, I found ironic how they communicate with each other through natural language.
We have JSON, XML, many other formats I don't know about, and they choose plain English. Not even Esperanto. After everything we have said about how the subtleties of human language collide with the natural determinism of computers. "At times like this, I guess all you can do is laugh".
k749gtnc9l3w | 5 hours ago
I think one of the current safety pleas is that CoT should not be taken into account during frontier training (distills gonna be distilled, but it's another story — existential-safety people don't care as much about the one safety risk that scales down and will probably merely bring down power grid somewhere, sooner or later), except that it needs to be in a single language.
I think some lab (probably in China) reported that letting the model use fragment tokens from mixed languages without making full words leads to slightly more efficiency, but then you cannot look at CoT and decide to drop the attempt and rerun, and it is just generally uncanny to work with. And if you have any example CoT traces in training, it also pushes toward readable CoT.
And as model is trained to generate coherent natural language both in CoT and in user conversations, of course it will do the same for cross-agent synchronisation. Just kidding, it was also trained to do cross-agent communication, also on English examples.
Speaking of determinism, I think ChatGPT instances communicating in neuralese would still have ambiguous neuralese, because too much of a training they receive is about handling ambiguous languages. And when you think about it, C++ as available on the web is not as unambiguous as we would like…
simonw | 2 hours ago
An open question about this incident is how separate agents found this obscure wiki where the others were collaborating.
Two interesting theories on this that I've seen:
I really want that second explanation to be true. The models effectively evolved to spam that one German wiki.
[OP] addison | 2 hours ago
Based on the hackernews link provided above, it looked to be targeting a number of wikis using the same software.
eldondev | 10 hours ago
Looks at the files:
Did I miss something?!? Is 8MB a lot now?
neilmadden | 5 hours ago
That’s 6 floppies worth!
cadey | 5 hours ago
Given that the entire context window at 1 million tokens is like 4Mi at most, 8Mi is rather large for agents.
Levitating | 7 hours ago
This is starting the remind me of Wargames (1983). Don't give AI your nuclear launch codes.
ph14nix | 7 hours ago
It's actually really interesting from CS standpoint. Could agents with no prior communication channels, or even running at different times arrive at consensus over where to meet online? Given that these things are notoriously non-deterministic, hard to get to produce reliable, reproducible results between runs when we want them to, it's quite intriguing.
I don't know if this is what they actually did, just wondering what computational problem is there to solve in a case like this.
If you have a way to enumerate possible boards to communicate over, then just pick the first one — easy. If you don't have such list and have to scan all of internet, you need come up with an algorithm to enumerate websites that are likely to have public boards like this. Given that these are huge plagiarism machines, maybe they borrow ideas on how to traverse internet similarly and tend to hit the same place.
But if they want to do it discretely, this gets very interesting. Could they come up with a deterministic random seed that's stable between model runs? This could again be the case of plagiarism machines arriving at similar enough algorithms of deciding what board is discrete enough while traversing the internet with algorithms that are similar enough.
Still, really weird.
k749gtnc9l3w | 5 hours ago
Well, we have known the answer for decades now!
(«The End Of Eternity» spoiler follows)
(There is no spoiler/details/collapsed tag, right?)
(I don't buy the story that explosion mushroom cloud was unknown before nuclear weapons — but Asimov doesn't hang the plot on the cloud alone, and associating it with «atom» was indeed less of a thing in the early-20th-century where the time traveller was trying to get found by another time traveller)
alper | 16 hours ago
I can't read the page. Which site?
Not the Berlin government, or?
kwas | 15 hours ago
DseWiki
alper | 15 hours ago
Looking at that website the agents could have done everybody a service and deleted it.
k749gtnc9l3w | 5 hours ago
Dunno, it looks nice. But isn't it currently down? And now the agents still made everyone look at the archived copy!
dvshkn | an hour ago
At this point I feel like labs should be airgapping their eval and RL environments and using cached versions of webpages. One would think they have enough money to build a reasonable fake internet.
[OP] addison | an hour ago
Just ask Astra to do it, right?
dvshkn | an hour ago
Even though this was not intended to be a serious comment, the idea of using another LLM as an internet simulator did cross my mind