Haha, I was more like: I have to think alone now. But I get the feeling, as someone who has always made a point of using only foss on my own machines (own the means of production so to say), this is a weird time. I don't like it, I hope soon local models get good. But in the mean time... We have this... dependency.
I use a lot of spec driven development today. It does allow me to make workflows that were on my Life's backlog for months or years in 10~20 minutes. It's amazing, but also I lack ownership over the thing made. It's like I paid a contractor.
I can feel that my actual cognitive engineering skills are in decline. Does anyone else see this? To those of you who haven't hand written a PR in months, can you still write your own personal software as easily as you once could?
There are little tools/optimisations that I have been wanting to do for years that I can just throw at AI now. I am almost annoyed sometimes because the little easy things I could do to 'relax' are now gone: if I am working on it is something AI can't do, normally because there are moving parts outside of the repository it can't see (i.e. integration problems, and I hate those).
It feels like I’ve got TikTok brain for coding and have trouble staying focused on the boring grunt work of coding (like threading an API change across the required stack).
I can still do it but it requires more willpower than it used to in part because I’m getting used to telling some LLM to do it for me.
Kind of cool how fable 5.1 just worked around these by reassigning subagents to non-failing models with more supervision and now with a hint to Sol and Flash 3.8.
Officially signing up for the $200 Codex plan now. I do like CC, but these errors are happening every week now, and if they (Anthropic) gave regular resets that's one thing, but somehow I'm already at 62% weekly usage after the random reset this Tuesday using Fable 5.1 medium. Just not worth it any longer compared to Codex IMO. I already set up my skills and everything to be generic and not CC specific, so it shouldn't be too hard to transition over. I'll run both for a month and decide.
Couldn't recommend it more. I've been on their $100/mo. plan and they keep randomly resetting my weekly usage limit. It's been wild. Don't know how long it will last, but I couldn't go through all my usage if I tried.
Their Sol model has been doing great for most of my tasks, even at its "Medium" setting.
Much different picture then back when I was using Claude. I would sneeze and half my usage would be gone and then it go down several times in between.
It's been a while since I messed with serious hardware, what's the cooling noise level? I guess they have to be water cooled or the rack format 8u ones would be about 100db
My 4x 15k rpm 1U fans I use to cool my twin AMD V620s make a horrible racket, especially how they resonate with the rest of the workstation's case fans. Similar in volume and annoyance to running an 8" circular saw. I can only imagine it would probably be as bad or worse than that.
Yeah my experience in the past was that datacenter hardware (air cooled, at least, I've never dealt with water cooled) uses tiny fans at or about 15000RPM, so it matches basically what you're saying or worse.
That was one of those fantastic early google videos/youtube things I probably watched 100 times. This was another from (roughly) that era: https://www.youtube.com/watch?v=ULSA5wktywI
Setting aside annoyance at the downtime, I'm really curious about the reasons for the failures, because I have to imagine there are some novel failure modes when serving these giant models that I haven't experienced with the kind of work I've done.
Anyone out there working in this space who can elucidate us on interesting failure scenarios unique to the space?
A nice chance to try sonnet again, must say, I'm not missing the load bearing assumptions, honest read of the permission matrix, what's genuinely a Django convention and what's a design choice, what's worth my thinking, and what's worth being precise about.
Long story short, it seems to be faster and less vocal but not much dumber (it's just my thinking partner, so I read a lot of the output as I create a large data model).
I've almost entirely stopped using Opus 5, at least with Sonnet I know what i'm getting -- Fable for the complex stuff and Sonnet for the precision changes.
The verbosity of Opus 5 isn't even my issue, it's consistency. For every 10 tasks Opus 5 accomplishes, there's at least one task that Opus 5 does just an atrocious job of, or a debugging investigation that it just completely goes off the rails on.
My engineering has changed so much that I'm now using this time to catch up on emails and think about design and architecture for the next leg of work. (And post on HN, of course.)
I would never have predicted this a year ago.
I always scoffed about engineers not doing work during Github outages - I'd always find some kind of other engineering to do if PRs or builds were piling up.
But during model outages, and since I'm not actively set up to use OpenAI/Codex at the moment, I'll just find something else productive to do. I don't think I'm ever going to write code by hand again unless I'm fixing something the AI can't manage or the change is small enough. Why even try when the machine is 100x faster?
Should I give Codex another try? It's been a month or so since I last used it, and it always felt inferior at Rust and TypeScript.
I think if you've got the money, you should set up a coding harness that you can plug to openrouter and then be able to switch quickly, as these new models seem to be coming out pretty regularly now
switching subs is kind of a PITA since you need to commit to a month at least
This is a term fellow devs have been using for years and years and my brain just cannot come to re-index its meaning. So I am immediately distracted by the context switch to sociopathic liars and related psychology.
In old school ML a pathological solution is often a collapse of some kind due to an underspecified objective. Maybe one parameter explodes or the model predicts "0" for everything
I read this as "the code is working as intended but the bug is an artifact of the original design"
I genuinely have no idea what it is meant to mean in the context it uses it. Load-bearing etc is annoying but roughly understandable, some of the opus5 classics are completely out there.
I'm a native English speaker as well, I shudder to think what a second language speaker would do (even if they were very confident with the language!).
Has anyone else had luck having a system level prompt for “express all responses in bullet point form unless full sentences are specifically necessary”? Because I’ve been having tremendous luck with that.
Just over the last couple of days I've found that forcing the models to use bullet points has helped a ton. I was even tempted to post something/ask around if others had any similar experiences with this.
I remember Jetbrains Junie which read like that all the time. Never used first-person either. I've not used Jetbrains AI for a while though, not sure if it still reads that way.
Demand > supply. It's impressive that customers have not migrated en masse to other providers, given the frequency of these outages. Perhaps switching costs are greater than some would believe. Or, qualitative differences between models continue to exist, despite matching on public benchmarks.
Do you really think it’s a coincidence claude outages always happen when US and EU workdays overlap?
Anthropic is a trillion dollar company and employs way more skilled high-paid engineers than you are btw, you really think it’s a systems issue not capacity
That's quite a reach. More likely someone merged and deployed their vibe-coded PR and is now figuring out how to bring the service back up.
If there's a lot of demand for rollercoaster rides, the rollercoaster will not stop operating; instead the queue of people in front of it will increase.
It's not like a bridge or elevator where we have a certain number of people that can use it, and if one more person joins, the whole structure breaks apart and everybody perishes.
Those guys are running a website that provides an interface for some specific hardware. Just like file hosting providers back in the day selling their terabyte-sized hard disks in 100MB-increments.
> instead the queue of people in front of it will increase.
This can ultimately result in the system breaking. I don't think Anthropic engineers are that much worse than their peers, such that they are 10x more prone to causing outages due to bad deployments.
>> Perhaps switching costs are greater than some would believe.
I haven't switched because there's nothing to switch to that is anywhere as good. I've been making dedicated attempts at using Sol but it falls short, despite what some people claim.
I did migrate 80% of my tokens. But for some tasks claude models are still the best.
Easy workaround is to work outside US peak hours (europe morning). I love this outages, I am hardly affected, and weekly reset usually promptly follows!
If I didn't know any better, I would have said Grok is using Claude behind the hood. But definitely curious now why it’s happening with both these LLM providers around the same time
During each forced workflow interruption I look for alternatives. No matter if I can't use the product due to a server outage or due to their weird quota limitations.
This can't be good for retention numbers. Old-school VCs would've ripped them apart in the air. Where did all the expertise go?
Yeah, I really like GLM 5.3 Flash with vision. Reminds me of when Claude was lots of fun to work with. It does feel slow, thinking away on the Chinese servers, but it's so incredibly cheap that for a lot of tasks I don't mind it taking a bit longer. I'm finding it mindblowing having entire software features built out for 5 cents, 20 cents.
I suspect the only reason those alternative providers have better up-time and more generous quotas, is because they don't have nearly the same amount of demand. Notice that Deepseek recently had to increase their pricing, once it gained it popularity.
Bit of a misleading status page, if you click in you can see that grok 4.5 and 4.6 etc are totally down, with the rest of the models showing as "up". I_strongly_ suspect they are not weighting it to actual number of requests!
Isn't this the classic blackout scenario? Claude goes down, then people move over to Codex, which is overwhelmed, and crashes, so people move to Grok...
When you're dealing with millions of users and response times go up above timeouts, there isn't much difference between "at capacity" and "down" if most users can't reliably use the service.
Every time I see Anthropic ship an issue like this, I'm reminded of Boris Cherny's glorious quote: "At this point it's safe to say that coding is largely solved."
Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa depending on use case).
Personally I like Claude and have the 20x Max plan but even there I burned through the whole weekly quota with 3 prompts in less than a day using the new Fable 5.1 which is crazy. Now Opus 5 is down. These two issues are really testing my patience.
The interesting thing is they changed default mode for Claude code to auto mode and auto mode uses sonnet to decide whether the command is safe or not. With their Sonnet model outage, the entire thing stopped working.
here is the error it was throwing:
> Error: claude-sonnet-5[1m] is temporarily unavailable (overloaded), so auto mode cannot determine the safety of Edit right now. Wait a moment and then try this action again. If it keeps failing, continue with other tasks that don't require this action and come back to it later. Note: reading files, searching code, and other read-only operations do not require the classifier and can still be used.
I just experienced this through OpenRouter and filed a refund request. I'm curious how OpenRouter handles incidents like this across such a large number of models and providers.
chat.com is down for me but the models seem to still be working. But actually i noticed instability last night at 3am eastern when thinking kept repeatedly failing.
This of course could be a coincidence. But I suppose also that if one goes down the others get slammed, so there could be a cascade. If I were another provider I would possibly monitor Claude outages (or vice versa) as a load predictor.
We've all noticed that when one company announces a new model, the other one announces a new model usually within hours. I guess now we're syncing downtime too.
As an aside, reminder to consider keeping your status page on a separate domain lest the problem be e.g. with DNS. You can see cloudflare does this (and we all know github does too as they give us frequent reasons to check it)
What I find fascinating is that this thread was chock full of people saying “Well this is the last straw: I’ve just signed up for Codex.” And now every single one of them is gone. They felt like marketing shitposts when I read them and their sudden absence now that Codex is down too only confirms it in my mind.
It lols like the timeline is Claude went down, some % migrated to Codex, then Codex went down. So much for the “multiple providers” theory. Wonder how token-based providers are doing.
Is anyone else noticing degradation of quality many evenings from 9pm-2am-ish ET? I wonder if anthropic has some sort of usage pattern that incentivizes them to route to weaker models.
I also fully acknowledge that I might be making spurious associations
codex was working fine until I installed the update and now getting the same error. Gemini and claude are fine but there is an update for both and I'm not taking a chance.
honestly boggles my mind that actual companies put load-bearing (lol) work behind 3rd party api's like claude or gpt.. also its really funny because some of these companies fired actual humans to put said 3rd party api to work in place of humans... like what happens when ur boy claude stops showing up for work huh!?
_joel | 14 hours ago
Linello | 14 hours ago
teekert | 14 hours ago
airhangerf15 | 13 hours ago
I can feel that my actual cognitive engineering skills are in decline. Does anyone else see this? To those of you who haven't hand written a PR in months, can you still write your own personal software as easily as you once could?
sire-vc | 13 hours ago
whattheheckheck | 13 hours ago
Espressosaurus | 13 hours ago
I can still do it but it requires more willpower than it used to in part because I’m getting used to telling some LLM to do it for me.
Maybe those no AI Friday ideas have some merit.
nikhizzle | 14 hours ago
hmokiguess | 14 hours ago
nikhizzle | 13 hours ago
giancarlostoro | 14 hours ago
tartoran | 13 hours ago
giancarlostoro | 12 hours ago
Bluestein | 12 hours ago
rob | 14 hours ago
SunshineTheCat | 13 hours ago
Their Sol model has been doing great for most of my tasks, even at its "Medium" setting.
Much different picture then back when I was using Claude. I would sneeze and half my usage would be gone and then it go down several times in between.
Hope it goes well for you too!
dgellow | 14 hours ago
throwup238 | 14 hours ago
jaggederest | 13 hours ago
27183 | 12 hours ago
jaggederest | 9 hours ago
I always loved this video: https://www.youtube.com/watch?v=tDacjrSCeq4
27183 | 2 hours ago
aqme28 | 14 hours ago
aurareturn | 14 hours ago
dofm | 13 hours ago
aurareturn | an hour ago
Arguably, my brain usage is higher when using Claude Code. It's just used in a different way.
jodacola | 14 hours ago
Anyone out there working in this space who can elucidate us on interesting failure scenarios unique to the space?
martinald | 14 hours ago
D13Fd | 13 hours ago
Pungsnigel | 13 hours ago
whh | 13 hours ago
fimdomeio | 13 hours ago
teekert | 14 hours ago
Long story short, it seems to be faster and less vocal but not much dumber (it's just my thinking partner, so I read a lot of the output as I create a large data model).
nonethewiser | 14 hours ago
Good for more precise changes.
devin | 13 hours ago
teekert | 12 hours ago
Syntaf | 13 hours ago
The verbosity of Opus 5 isn't even my issue, it's consistency. For every 10 tasks Opus 5 accomplishes, there's at least one task that Opus 5 does just an atrocious job of, or a debugging investigation that it just completely goes off the rails on.
echelon | 13 hours ago
I would never have predicted this a year ago.
I always scoffed about engineers not doing work during Github outages - I'd always find some kind of other engineering to do if PRs or builds were piling up.
But during model outages, and since I'm not actively set up to use OpenAI/Codex at the moment, I'll just find something else productive to do. I don't think I'm ever going to write code by hand again unless I'm fixing something the AI can't manage or the change is small enough. Why even try when the machine is 100x faster?
Should I give Codex another try? It's been a month or so since I last used it, and it always felt inferior at Rust and TypeScript.
r_lee | 13 hours ago
switching subs is kind of a PITA since you need to commit to a month at least
martinald | 13 hours ago
evanmoran | 13 hours ago
Waterluvian | 12 hours ago
devin | 12 hours ago
lachlan_gray | 12 hours ago
I read this as "the code is working as intended but the bug is an artifact of the original design"
martinald | 2 hours ago
I'm a native English speaker as well, I shudder to think what a second language speaker would do (even if they were very confident with the language!).
Waterluvian | 13 hours ago
nater5000 | 12 hours ago
visarga | 12 hours ago
chuckadams | 12 hours ago
jason-phillips | 12 hours ago
doawoo | 12 hours ago
estetlinus | 11 hours ago
stri8ted | 14 hours ago
testfrequency | 14 hours ago
stri8ted | 14 hours ago
ieie3366 | 14 hours ago
Anthropic is a trillion dollar company and employs way more skilled high-paid engineers than you are btw, you really think it’s a systems issue not capacity
bflesch | 14 hours ago
If there's a lot of demand for rollercoaster rides, the rollercoaster will not stop operating; instead the queue of people in front of it will increase.
It's not like a bridge or elevator where we have a certain number of people that can use it, and if one more person joins, the whole structure breaks apart and everybody perishes.
Those guys are running a website that provides an interface for some specific hardware. Just like file hosting providers back in the day selling their terabyte-sized hard disks in 100MB-increments.
stri8ted | 14 hours ago
This can ultimately result in the system breaking. I don't think Anthropic engineers are that much worse than their peers, such that they are 10x more prone to causing outages due to bad deployments.
kilroy123 | 14 hours ago
khalic | 12 hours ago
enraged_camel | 14 hours ago
I haven't switched because there's nothing to switch to that is anywhere as good. I've been making dedicated attempts at using Sol but it falls short, despite what some people claim.
hmokiguess | 13 hours ago
throw83930489 | 13 hours ago
Easy workaround is to work outside US peak hours (europe morning). I love this outages, I am hardly affected, and weekly reset usually promptly follows!
jephs | 12 hours ago
stacktrace | 14 hours ago
If I didn't know any better, I would have said Grok is using Claude behind the hood. But definitely curious now why it’s happening with both these LLM providers around the same time
Phemist | 14 hours ago
zipy124 | 14 hours ago
aurareturn | 14 hours ago
stacktrace | 14 hours ago
bflesch | 14 hours ago
This can't be good for retention numbers. Old-school VCs would've ripped them apart in the air. Where did all the expertise go?
derwiki | 14 hours ago
SyneRyder | 13 hours ago
stri8ted | 14 hours ago
trjordan | 14 hours ago
Looks like trouble in the SpaceX datacenters.
alansaber | 14 hours ago
clickety_clack | 14 hours ago
genidoi | 13 hours ago
clickety_clack | 10 hours ago
empath75 | 14 hours ago
matt-p | 13 hours ago
r_lee | 13 hours ago
those were only relieved once they did the Colossus deal with SpaceX
are you only saying that because of Elon?
pantalaimon | 12 hours ago
matt-p | 13 hours ago
Aldipower | 13 hours ago
foresterre | 13 hours ago
Aldipower | 13 hours ago
rmujica | 12 hours ago
Aboutplants | 13 hours ago
martinald | 13 hours ago
torginus | 13 hours ago
kristofferR | 13 hours ago
reinhash | 13 hours ago
embedding-shape | 12 hours ago
esskay | 12 hours ago
Lalabadie | 12 hours ago
mtgh2s | 14 hours ago
lol
eis | 14 hours ago
bakies | 13 hours ago
eis | 12 hours ago
d1ss0nanz | 14 hours ago
rrrx3 | 13 hours ago
jaggederest | 13 hours ago
rrrx3 | 12 hours ago
scottydelta | 14 hours ago
here is the error it was throwing:
> Error: claude-sonnet-5[1m] is temporarily unavailable (overloaded), so auto mode cannot determine the safety of Edit right now. Wait a moment and then try this action again. If it keeps failing, continue with other tasks that don't require this action and come back to it later. Note: reading files, searching code, and other read-only operations do not require the classifier and can still be used.
throw83930489 | 13 hours ago
rrrx3 | 13 hours ago
chrisjj | 14 hours ago
Anthropic, please hire a literate human.
jh54 | 14 hours ago
kingkawn | 13 hours ago
Khaine | 13 hours ago
nr378 | 13 hours ago
cute_boi | 13 hours ago
Aboutplants | 13 hours ago
wslh | 13 hours ago
gtsnexp | 13 hours ago
ModernMech | 13 hours ago
gtsnexp | 13 hours ago
ModernMech | 13 hours ago
gtsnexp | 13 hours ago
TomGarden | 13 hours ago
tekacs | 13 hours ago
https://www.reddit.com/r/codex/comments/1w69ff6/whats_this_e...
graemep | 13 hours ago
jsw97 | 13 hours ago
tekacs | 13 hours ago
jotaefea | 13 hours ago
Their status page is hard to believe given how many folks are chiming in to report on Reddit https://status.openai.com/
anonbuddy | 13 hours ago
RinTohsaka | 13 hours ago
_Tev | 13 hours ago
Status page is very funny - claiming "We’re fully operational" while https://chatgpt.com gives 404, lol.
continuum | 13 hours ago
thinkingtoilet | 13 hours ago
postalcoder | 13 hours ago
Seeing as how xai, anthropic, and openai are all having issues at the same time, it has to be a common provider, right? Maybe cloudflare?
edit: ChatGPT and Grok are back up. Claude is not. Typical.
https://status.x.ai
https://status.claude.com
https://status.openai.com
https://www.cloudflarestatus.com
anonymars | 13 hours ago
https://news.ycombinator.com/item?id=27102020
small_model | 13 hours ago
nprateem | 13 hours ago
_Tev | 13 hours ago
EDIT: it was just 3.6 Flash, 3.1 Pro works it seems.
alberth | 13 hours ago
They are experiencing increased errors themselves.
https://www.cloudflarestatus.com/?t=1
ksimukka | 13 hours ago
Rover222 | 13 hours ago
reinhash | 13 hours ago
sire-vc | 13 hours ago
goonersallofyou | 13 hours ago
damsta | 13 hours ago
messh | 13 hours ago
khalic | 12 hours ago
lowbloodsugar | 12 hours ago
It lols like the timeline is Claude went down, some % migrated to Codex, then Codex went down. So much for the “multiple providers” theory. Wonder how token-based providers are doing.
mandadtabrizi | 12 hours ago
chb | 12 hours ago
dnautics | 12 hours ago
I also fully acknowledge that I might be making spurious associations
mandadtabrizi | 12 hours ago
macwhisperer | 12 hours ago