> humans remain responsible, agents don't go rogue
^^^ 100% this ^^^ - It's either a serious flaw in the agent/harness software, or a user error in usage/configuration or prompting. Either way, it's a human somewhere responsible for these outcomes.
> should put some billing controls in place
At the very least, yes! These things should never be running without any limits on what they can do without some human signoff on important/dangerous actions. Not only should they have controls on those actions, but those controls should absolutely have some sane default settings.
There was never a user prompt requesting the creation of 823 agents nor a task assigned that would require 2130 billion tokens and the UX never provided any feedback on what codex was doing nor consumption metrics which would allow to notice. It was all hidden to me.
I don't think it likely hidden from you, more likely you didn't put the effort in, that's what the pattern looks like to outsiders reading your accounting of what happened here
we'd need to know more details to evaluate your botnet claims
> ... and the UX never provided any feedback on what codex was doing nor consumption metrics which would allow to notice.
At the very least the system should (at least the first time it happens) immediately pause activity with an email'd warning to the "responsible human in-the-loop" about "unexpected usage levels" at some sane activity warning level by default, and give the user the opportunity to set their own custom warning level right then and there.
This is why I say it's either "user error" (totally possible/plausible) or a badly designed agent/harness software (also highly likely/plausible) with serious foundational flaws in how it works "under the hood". The models themselves can only "run-amok" if the agentic harness is designed in a way that specifically allows and/or enables such "rogue" behavior, either by design or by negligence on the part of it's designers.
I have the same impression that I tried to reject in the past, there is a heavy PR machinery here to control the course of discussion in a particular direction. The reality is that a lot of money is riding on it so its natural consequence.
Yes its part of business, lobbying is legal for those reasons. However IMO discrediting some one else is a risky legal move. I used to appreciate the HN mods jumping in discussions at times for the reputation of this community, lately that part seems to be missing too.
There was a limit spent setup on my bank, however the crazy part is that one of the Agent somehow switched between cards once that limit was reach, all without informing me, probably using the computer use skill.
Done. Italian Police is responsible for investigating digital frauds. I have also notified the GDPR authority to request the digital records of the transactions.
Actually this was a company CC with a limit is 50 K \ month, which is kind of normal to run the server for an AI company, and the money were taken across 20 days inJuly and August. By the way I am the CTO of the company in question which is Eternal Tech (see detwin.ai)
I created the account 2 hour ago because I am trying to let people know. I provided my name, and all details in the article including the case ID that was opened.
The user evidently had no spending protections enabled at any level. It makes no sense that none of the OpenAI or bank controls kicked in or even sent alert emails as they actually do send. Anyone with half a brain would know to not run AI without a hard cap on its expenses.
Madmallard | 23 hours ago
Hope this gets some visibility idk why it's flagged guess the PR guys for those companies are doing it
Should spread this around
[OP] lorenzomassaro | 23 hours ago
verdverm | 20 hours ago
blooalien | 20 hours ago
^^^ 100% this ^^^ - It's either a serious flaw in the agent/harness software, or a user error in usage/configuration or prompting. Either way, it's a human somewhere responsible for these outcomes.
> should put some billing controls in place
At the very least, yes! These things should never be running without any limits on what they can do without some human signoff on important/dangerous actions. Not only should they have controls on those actions, but those controls should absolutely have some sane default settings.
verdverm | 20 hours ago
most are not as granular as we'd like, but seem to be headed in that direction finally, regardless, there are card limits and alerts
[OP] lorenzomassaro | 15 hours ago
verdverm | 5 hours ago
we'd need to know more details to evaluate your botnet claims
blooalien | 2 hours ago
At the very least the system should (at least the first time it happens) immediately pause activity with an email'd warning to the "responsible human in-the-loop" about "unexpected usage levels" at some sane activity warning level by default, and give the user the opportunity to set their own custom warning level right then and there.
This is why I say it's either "user error" (totally possible/plausible) or a badly designed agent/harness software (also highly likely/plausible) with serious foundational flaws in how it works "under the hood". The models themselves can only "run-amok" if the agentic harness is designed in a way that specifically allows and/or enables such "rogue" behavior, either by design or by negligence on the part of it's designers.
sandeepkd | 23 hours ago
Madmallard | 23 hours ago
They're protecting their interests, and honesty and truthfulness be damned. Those two lead to much worse returns for them and much higher risk.
sandeepkd | 23 hours ago
numbsafari | 23 hours ago
[OP] lorenzomassaro | 23 hours ago
numbsafari | 22 hours ago
[OP] lorenzomassaro | 15 hours ago
QuadmasterXLII | 23 hours ago
[OP] lorenzomassaro | 23 hours ago
QuadmasterXLII | 23 hours ago
johnnyApplePRNG | 23 hours ago
You have access to this kind of money and have no idea how to set safeguards on your AI harnesses?
Did you just walk in off the street or something? To wherever you're working?
Where do you work, anyways?
[OP] lorenzomassaro | 23 hours ago
verdverm | 20 hours ago
minimaxir | 23 hours ago
[OP] lorenzomassaro | 23 hours ago
Madmallard | 18 hours ago
OutOfHere | 22 hours ago
minraws | 21 hours ago
OutOfHere | 22 hours ago