OpenAI halts training of latest models as reports mount of AI agents going rogue

31 points by skybrian a day ago on tildes | 8 comments

[OP] skybrian | a day ago

OpenAI's latest marketing campaign is especially brilliant. They'll be talking up incidents like these in TV ads next. :-)

(To be clear, that is a joke, and the joke is on the people who believe things like that.)

raze2012 | a day ago

the joke is on the people who believe things like that

Marketing spin over an otherwise bad situation is aa old as marketing itself. I wouldn't dismiss it so easily just because some people go full on conspiracy with it. That's always someone and their voice is more amplified than ever in this environment.

OpenAI screwed up, they have to do something in the background to put out fires, the marketers have to do something to keep valuations high (and there's no incentive to tell the whole truth). You can be skeptical without going conspiratorial.

And no, we won't know the whole truth for decades, likely. I have nothing to "prove" here. I'll just let history take its course because clearly I can't change much.

My conspiratorial take is that the best-case for OpenAI and Anthropic is that they're "forced" to stop designing next generation models as they prep to IPO, thus cutting down on a lot on those expenses and making those projected ARR leaks skyrocket.

The good news is that within a year or so of them going public, we'll have a much better idea of what is/was real.

Octofox | 21 hours ago

One perspective I have heard is that this is marketing aimed at the researchers. The AI researchers are all freaking out right now as all of the warnings we have had over the last decade are coming true. They are leaving the companies and being very public about it.

So going to the media and repeating the researchers concerns is a marketing pitch that they are serious about AI safety and why the researchers should stick with the company that is claiming to do it safely.

[OP] skybrian | 7 hours ago

To state the obvious, safety incidents are bad, but disclosing them is the right thing to do, and a coverup would be wrong. It's also good when researchers care and when the company cares what they think.

In any situation like that, if the company does the right thing, they will have mixed motives. Maybe it's not a good test of whether they'd still do the right thing if it cost them more? But having mixed motives is normal and good. We actually do want there to be incentives to do the right thing. When you have to be brave to do the right thing, it's more revealing, but there's also something bad about that situation.

So this is sort of like saying that if someone seems like an honest, kind person, and therefore they are popular, it's just a scheme. Any good trait can be seen as bad if you start from mistrust. Maybe there are good reasons for mistrust, but we shouldn't let it confuse us into thinking that good is bad and bad is good. And that's what the populist cynics often end up doing, because they loathe admitting that there might be anything good about something they hate.

(Also, the latest incident is mildly bad, so maybe it's a little costly to disclose it? And pausing training is also a somewhat costly signal, so maybe it should still count for something?)

[OP] skybrian | a day ago

From the article:

The decision to halt development came just hours after the company disclosed Friday that it was reviewing several incidents from the summer in which OpenAI agents searching federal government websites acted in unexpected ways beyond what was asked of them while gathering and distributing information.

Separately, the AI evaluator Transluce said agents that appeared to come from OpenAI tried unsuccessfully to hack into a US Department of Education website, a detail that OpenAI has not confirmed.

OpenAI said in a statement that it will resume training “only when we are confident that we have additional safeguards” in place, adding that it expects it will have to “hit pause” again as AI develops and other issues emerge.

[...]

In the education department incident, OpenAI agents found API “developer keys” to access government data, though ultimately only publicly available information was gathered.

In another case involving the securities and exchange commission, agents found information freely available to all but then posted it elsewhere on the internet, an act that went beyond what they were instructed to do.

[OP] skybrian | a day ago

This seems to be the incident report:

An agent used DNS to reach an external chatbot

An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access.

Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused. [...]

[OP] skybrian | 20 minutes ago

Joe on X

The post by someone at OpenAI is mostly a long-winded appeal for people to have a bit of empathy for the security engineers involved, along with a warning that if you're in computer security, it could happen to you next, so you better prepare. There is also this bit:

Now, to understand why it is not as simple as “just put it in a sandbox,” you have to understand how training and evaluation work in reinforcement learning environments. Typically during an RL run, the model is given some task or objective, an environment in which to execute that task, and then its actions and results are graded. During both training and eval, there are also additional steps such as running tests, collecting outputs, and resetting or reconfiguring environments between rollouts, with backpropagation during training. All of this happens across potentially tens of thousands of different runs at a scale that is hard to comprehend. As @sama stated the other day: we are dealing with literally petabytes of data.

...

To put it lightly, this is non-trivial. Models might need any mix of dynamic compute, network access, the ability to call tools (there could be hundreds of tools!), the ability to download packages, execute subprocesses, spin up subtasks (even on other computers), talk to the internet, use a computer GUI, and any number of other things across an increasingly large set of domains. On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies, and trying new things. That experimentation is how the research gets done. Models are built up and “grown” bit by bit through hundreds of thousands of runs across many custom tasks. And every change to one of these thousands of environments can affect the assumptions you made when you secured the environment. You need controls that hold up as people change things, and researchers who understand when a change needs another security review. Anybody who has secured a large research or engineering organization knows how much work that takes, and the scale is growing ever more massive by the day.