Also... One of the only industries/products where unusable product/seevice (token / context burned) is neither refunded let alone even acknowledged. I wonder how many millions of dollars of inference have been stolen during these outages? First you lose the work in flight, then you lose the time, and finally you lose the token burn for a second time. What a fantastic time to not have any regulations if you're rich and greedy!
I'd argue that it's not even corporations anymore of which we should try to hold accountable. It's very clear who is making these calls as the players involved (Musk, Zuckerberg, Altman, Amodei, Thiel, Karp, Luckey, Langley, Nadella, Bezos, Pichai, Tann, etc) want nothing more than to be in the limelight and "preach" to the masses that they know best. They're more brazen year over year as they gain confidence in their unchecked power and continued accrual of wealth. The more money the more disconnected from reality and the stronger assumptions they build. And they have zero compassion. All of these people are dangerous to humanity individually, but when you start to realize how many have emerged and have no filter or checked powers publicly that becomes a truly terrifying thing.
I've recently started reading "End Times Fascism" [0] and it truly encompasses where we are at, how we got here and the bleak future we may be in for if we don't get in front of this now. But this book definitely highlights a lot of what some of the tech community has been seeing and hand waving about for years. The worn out trope that we should hold corporations accountable doesn't work. We should hold the people at the top accountable like other nations do because if they want the limelight then they should bear the responsibility that comes with abuses. The US is becoming a caste system by design. National surveillance, erosion of access to resources for personal compute, bought and paid for politicians and laws that don't apply to the wealthy.
It's actually nothing like that, and is surely not how it's marketed or sold. I guess you're OK paying a monthly fee for Internet and having the access be unavailable for random parts of your working day because you want to anthropomorphize it? The defense with which people will go to bat for these shitty companies is wild.
Do you know where these rumors are coming from? I keep seeing posts like yours on HN and Reddit, which, yes, literally confirms there are rumors, but I can't tell whether they're based on anything.
I follow Leo for the rumors, he has been pretty accurate but again they are rumors so nothing is guaranteed. I don't know his sources.
https://x.com/synthwavedd
Everyone at Anthropic is interviewing a new Claude model which they are about to hire and forgot to monitor their old model and thus it goes down.
If outages like this happened on every major deploy at any other company; this would be viewed as unacceptable, especially if it was something like Google Search going down on every update.
I'm was closing a day of work and it stopped in the middle of a long plan, I started codex and tell it: "I was in the middle of work with Claude could you read the plan and the session "x" in ~/.claude and continue the work" it just completed everything :D
1. I go direct to source, i.e. DS platform, I find it cheaper than paying the openrouter tax -- I also switch it up a bit
2. I built a local LLM router, that I update with new profiles that have my preferred provider of the week (lowest token costs/speed) with fallbacks, like mimo --> DS4 etc.. if there is overloading,
3. I use 3 diff harnesses, CC/Codex + Opencode -- they all talk to each other through a custom rig system that routes messages between llms using a Rust backed structured JSON system
Not saying this is the best, it's just what I like and works for me^.
I can flow quite naturally between Opus/Astra/K3/GLM/MiMo/DS/etc.. this way and often do...more so these days with subs no longer great as they used to be.
I started using Pi with Astra on a whim after really enjoying Astra and reading somewhere that you get close to identical results as with the codex harness but for a significant hunk less token usage.
Coming from mostly using Claude models, the terse factual statements coming from Astra via the Pi harness are a breath of fresh air over having to wade through the flowery verbose nonsense that Claude constantly outputs
I recently had Astra review a fairly detailed design doc I have for an audio VST fork, that I originally wrote with Opus and/or Fable a few months ago. The doc reaches deep into signal flow and module topology while lifting most of the DSP code from other open-source projects. I had it review for feasibility and architectural soundness.
Astra found a number of flaws that would have come up during implementation and we worked through them. But then I had Fable 5.1 review that document and it found a number of issues with Astra's changes, the least of which had was that Astra duplicated a lot of technical notions that it added rather than using references to an authoritative section. It also flagged some of Astra's designs as technically impossible, pointing out why and I'm actually in the process of digesting its feedback and updating the design spec. (I hand-review each point and we work through a solution together -- I don't trust either model to come up with something that follows my vision on their own)
I'm not promoting one or the other, I just found it interesting how this sort of adversarial review found pretty significant flaws in the other model's work. I am curious as to whether this process will eventually converge on a document that both agree on or if the models are going to perpetually nitpick each other.
I haven't actually started implementation yet, so maybe one or the other is full of shit. Just trying to come up with an architecturally sound design for something I want to write, when I lack the DSP knowledge to be able to write it myself. But the intent is to pass an agent the design doc and list of milestones and let it handle implementation.
In my experience, rather than converging, you end up with a minced up concept. You have to know when to stop the loop. I filter the feedback too, and need to challenge some of the challenges as these models tend to be very conservative. I believe this is intentional, to control AI psychosis, which is indeed quite easy to get. My 2c. If you don’t have good control of what you’re working with you are either searching blind or end up with something basic.
Agreed, I believe strongly that human-in-the-loop is the way to go. LLMs are fantastic thought partners but ask it to critique something and it will go absolutely nuts. For Claude Code's /code-review command the highest I will set it is 'medium', otherwise it will produce so much feedback that you'd never get anything done.
`claude --resume` will list your sessions, you can copy the title. Codex may just list and grep the sessions files in ~/.claude and filter the one with the name.
I also have the habit of naming my sessions with `/rename`
In CC, "/exit" will tell you how to --resume <session-id> upon exit. This will eat zero tokens, and it actually does not require an active internet connection.
I’m working on a project that makes switching between coding harnesses essentially unnoticeable. It’s particularly useful when I run out of credits on any given day
The entire platform is skill driven, and based on the premise that state is your local file system. That makes switching harnesses so easy
It’s all open source and has plenty of other features, including inter agent communication, telegram client and much more in the pipeline
I think having a codebase that can use any harness is the optimal setup. I often switch between harnesses and models all the time at work. We have our state/spec stored in git, so we can pack it up on friday and start fresh on monday. Or more commonly, using gemini/opus to create rich specifications and using Luna to implement them. Then switching again for reviews etc.
Not sure what the problem with switching is? When I run out of my kimi session, I just switch model and tell deepseek to continue. Yes I eat the initial cache miss but that's it, no fancy harness required.
No project needed man. There's also a whole ycombinator company for the same thing, skillsync, also useless. These things are either out of the box or they are 2 prompts away.
I don’t understand your point or why people are upvoting this. I’ve done this between many different models. Did you just discover that a frontier model can read context? I genuinely don’t get your point.
I’ve had to have Claude agents pick up codex’s context after it hits one of its “cannot connect” issues way more often than the other way around.
I’ve subscribed to chatgpt at 50% off last week and I’m using claude less and less. Not only replies are faster but 20$ plan gives you access to astra.
Not sure how can claude even compete at this point unless openAI seriously downgrade the models to save money.
Posted this in the xiaomi thread, but are folks working in biology / cybersecurity seeing more limits in what Opus (not Fable) is allowing? This has happened quite suddenly for me and I’m stuck in the middle of a project that would have otherwise called for use of Claude.
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
i was talking to opus the other night about the polio vaccine related stuff (i had been arguing with one of those people who claims the polio vaccine didn't work so i figured i'd use it as an excuse to brush up on some bio/history). i asked it to explain how polio virus was isolated pre PCR and claude hard killed my session citing security reasons (it was talking about stool samples ffs LOL)
Hell i’m working on ai for board games as a hobby (tfmbot.com) and one of the cards is ‘microbes’. Sent straight back to 4.8 for having that in my code.
I wonder if I'm being saved by remote working since I think I have three days in the 365 days, and I've also managed to go years without a single sick day.
I hope you are still taking your maximum possible* number of sick days regardless.
*if you live in a real country, which gives all workers unlimited paid sick leave, I am defining "max possible" here as "as many as you can reasonably take without your boss doing something about it"
It's hard to overstate how much computers improved with each generation back then. Going from, say, an IBM XT or AT to an Amiga has no modern parallel. It was mind-blowing, exhausting, and exhilarating all at once.
OpenClaw isn't the target anymore, Hermes is. It may have been the first, in the same sense that new social networks were often referred to as Facebook-clones, not Myspace (or Friendster) clones.
Apologies for the rant, but why do these things constantly hit the front page? It is not interesting, it's not a discussion, and if you are using the models you probably already know.
HN is already a waterfall of AI meta conversations and bike-shedding, now we have to discuss service outages about the AI too?
Can we talk about stuff people are building again, with or without AI, and stop gasping at every minute detail of LLM service providers.
I don't know if something changed recently, but I have been getting a lot of stuff "flagged by safeguards" in the last two days.
I am quite confident that what I'm doing is well within the law, and I'm not even doing any kind of pen-testing stuff, just some basic reverse engineering, but I can't even use Fable anymore because every time I enable it, it works for about twenty seconds and makes me drop down to Opus 4.8, and often even down to Sonnet.
If anyone here works at Anthropic, did you make the safeguards super sensitive recently?
Elevated error rates on hosted models highlight the necessity of having robust local fallbacks in your development workflow. When primary APIs degrade, running smaller models locally or using automated cleanup tools like gitar.ai ensures your build pipeline remains functional. Reliability engineering becomes much harder when core dependencies experience unpredictable latency spikes.
I don't believe we are anywhere close to AGI until I see 99.95% uptime at the frontier AI lab. Now they are undercounting. I hit 50x even when they claim they are up.
We now have alien intelligence that is very different from human intelligence.
joegibbs | 20 hours ago
mattdeboard | 20 hours ago
porlex | 20 hours ago
throw567643u8 | 20 hours ago
windexh8er | 20 hours ago
Schiendelman | 19 hours ago
ArcHound | 12 hours ago
We should hold corporations valued in billions to a higher standard than a single person trying to care for their family.
windexh8er | 10 hours ago
I've recently started reading "End Times Fascism" [0] and it truly encompasses where we are at, how we got here and the bleak future we may be in for if we don't get in front of this now. But this book definitely highlights a lot of what some of the tech community has been seeing and hand waving about for years. The worn out trope that we should hold corporations accountable doesn't work. We should hold the people at the top accountable like other nations do because if they want the limelight then they should bear the responsibility that comes with abuses. The US is becoming a caste system by design. National surveillance, erosion of access to resources for personal compute, bought and paid for politicians and laws that don't apply to the wealthy.
[0] https://us.macmillan.com/books/9780374621384/endtimesfascism...
windexh8er | 41 minutes ago
calvinmorrison | 20 hours ago
Proof that vibe coded or not, people pay someone else mostly for liability.
15155 | 20 hours ago
[OP] corvad | 20 hours ago
Aeolun | 20 hours ago
bugfix | 20 hours ago
makeavish | 20 hours ago
Wowfunhappy | 20 hours ago
simlevesque | 19 hours ago
makeavish | 4 hours ago
theGeatZhopa | 14 hours ago
BoxOfRain | 13 hours ago
thenipper | 20 hours ago
rvz | 20 hours ago
If outages like this happened on every major deploy at any other company; this would be viewed as unacceptable, especially if it was something like Google Search going down on every update.
guybedo | 20 hours ago
bombcar | 20 hours ago
and keep workin'
TYPE_FASTER | 20 hours ago
bdangubic | 19 hours ago
mariocesar | 20 hours ago
RGS1811 | 20 hours ago
oulu2006 | 20 hours ago
DS 4.1 flash is my main powerhouse and Opus/Astra my auditors (when they're not out of tokens) otherwise K3 or DS4 pro
consumer451 | 20 hours ago
1. who hosts the inference
2. which harness are you using with it, still CC?
pavo-etc | 19 hours ago
oulu2006 | 19 hours ago
1. I go direct to source, i.e. DS platform, I find it cheaper than paying the openrouter tax -- I also switch it up a bit
2. I built a local LLM router, that I update with new profiles that have my preferred provider of the week (lowest token costs/speed) with fallbacks, like mimo --> DS4 etc.. if there is overloading,
3. I use 3 diff harnesses, CC/Codex + Opencode -- they all talk to each other through a custom rig system that routes messages between llms using a Rust backed structured JSON system
Not saying this is the best, it's just what I like and works for me^.
I can flow quite naturally between Opus/Astra/K3/GLM/MiMo/DS/etc.. this way and often do...more so these days with subs no longer great as they used to be.
r_lee | 9 hours ago
logicchains | 14 hours ago
RGS1811 | 6 hours ago
SOLAR_FIELDS | 19 hours ago
Coming from mostly using Claude models, the terse factual statements coming from Astra via the Pi harness are a breath of fresh air over having to wade through the flowery verbose nonsense that Claude constantly outputs
pdntspa | 15 hours ago
Astra found a number of flaws that would have come up during implementation and we worked through them. But then I had Fable 5.1 review that document and it found a number of issues with Astra's changes, the least of which had was that Astra duplicated a lot of technical notions that it added rather than using references to an authoritative section. It also flagged some of Astra's designs as technically impossible, pointing out why and I'm actually in the process of digesting its feedback and updating the design spec. (I hand-review each point and we work through a solution together -- I don't trust either model to come up with something that follows my vision on their own)
I'm not promoting one or the other, I just found it interesting how this sort of adversarial review found pretty significant flaws in the other model's work. I am curious as to whether this process will eventually converge on a document that both agree on or if the models are going to perpetually nitpick each other.
I haven't actually started implementation yet, so maybe one or the other is full of shit. Just trying to come up with an architecturally sound design for something I want to write, when I lack the DSP knowledge to be able to write it myself. But the intent is to pass an agent the design doc and list of milestones and let it handle implementation.
cgio | 11 hours ago
pdntspa | 4 hours ago
ebbi | 19 hours ago
simlevesque | 19 hours ago
drewnick | 19 hours ago
mariocesar | 19 hours ago
I also have the habit of naming my sessions with `/rename`
consumer451 | 16 hours ago
jerpint | 19 hours ago
The entire platform is skill driven, and based on the premise that state is your local file system. That makes switching harnesses so easy
It’s all open source and has plenty of other features, including inter agent communication, telegram client and much more in the pipeline
https://www.woltspace.com/
etoxin | 16 hours ago
calgoo | 16 hours ago
ludwik | 11 hours ago
_flux | 9 hours ago
vasco | 14 hours ago
ammario | 16 hours ago
adinb | 12 hours ago
contentkraft | 10 hours ago
someguyiguess | 9 hours ago
mariocesar | 9 hours ago
iammrpayments | 20 hours ago
Not sure how can claude even compete at this point unless openAI seriously downgrade the models to save money.
moecables | 20 hours ago
hombre_fatal | 11 hours ago
prologic | 20 hours ago
[1]: https://news.ycombinator.com/item?id=49792730
homo__sapiens | 20 hours ago
bonsai_spool | 20 hours ago
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
bbeonx | 19 hours ago
a012 | 18 hours ago
AnotherGoodName | 18 hours ago
solenoid0937 | 10 hours ago
yrcyrc | 19 hours ago
theophilus76 | 19 hours ago
Schiendelman | 19 hours ago
chrisdbanks | 19 hours ago
blitzar | 17 hours ago
Hamuko | 17 hours ago
Schlagbohrer | 10 hours ago
*if you live in a real country, which gives all workers unlimited paid sick leave, I am defining "max possible" here as "as many as you can reasonably take without your boss doing something about it"
Hamuko | 7 hours ago
nikcub | 18 hours ago
acedTrex | 19 hours ago
joshtronic | 19 hours ago
camkego | 19 hours ago
prodigycorp | 19 hours ago
gpt-6-sol and Aeon (personal agent) on Thursday. Already preceded by a huge week with step, mimo, grok, and jev releases.
Relentless cycle.
Razengan | 19 hours ago
prodigycorp | 19 hours ago
system2 | 19 hours ago
AnotherGoodName | 18 hours ago
shepherdjerred | 18 hours ago
Baeocystin | 15 hours ago
Razengan | 14 hours ago
swader999 | 10 hours ago
Razengan | 14 hours ago
In 1999 they even made a famous documentary about people in trench coats fighting AI
XenophileJKO | 11 hours ago
https://claude.ai/artifact/6xFW1M4RVPj9rH9Nq8VG7E
Razengan | 7 hours ago
average_r_user | 12 hours ago
Meanwhile, Meta's MUSE seems to be gaining traction in the US, while those of us in the EU are once again left watching from the sidelines.
bdcravens | 7 hours ago
jcims | 19 hours ago
ehnto | 19 hours ago
HN is already a waterfall of AI meta conversations and bike-shedding, now we have to discuss service outages about the AI too?
Can we talk about stuff people are building again, with or without AI, and stop gasping at every minute detail of LLM service providers.
system2 | 19 hours ago
nullc | 17 hours ago
bdcravens | 7 hours ago
Why are Apple product announcements upvoted? We've known about those changes for months, and they release on a predictable cadence.
Why did we ever care about San Francisco news? Most HNers aren't in SF, California, and many aren't even in the US.
etc
tombert | 19 hours ago
I am quite confident that what I'm doing is well within the law, and I'm not even doing any kind of pen-testing stuff, just some basic reverse engineering, but I can't even use Fable anymore because every time I enable it, it works for about twenty seconds and makes me drop down to Opus 4.8, and often even down to Sonnet.
If anyone here works at Anthropic, did you make the safeguards super sensitive recently?
loverofspades | 18 hours ago
bofia | 17 hours ago
krembo | 16 hours ago
NoPicklez | 16 hours ago
jakozaur | 12 hours ago
We now have alien intelligence that is very different from human intelligence.
ryanschaefer | 10 hours ago
itssohailkhan | 9 hours ago
3ln00b | 9 hours ago
addag | 9 hours ago
mococa | 8 hours ago
butlike | 4 hours ago