Models do not get nerfed. There has never been evidence of this. This would be trivial to prove if it were true, and such a proof would be a huge story and scandal to a news market hungry for a shred of a signal on AI's downfall.
This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s.
One theory is that they start serving at full precision, then quantize to save on costs as adoption grows. It's kind of a conspiracy theory atm, but I have definitely felt it - I had a large project done on Opus 4.8 release day, a week later it was struggling to complete partial tasks in the same area.
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
It is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops.
Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.
Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.
my personal experience with swe has been…suboptimal. I am not sure how much I buy these benchmarks and it has a very “just blurt it out even if it’s probably not right” style but hey it’s free.
varispeed | 15 hours ago
cbg0 | 15 hours ago
varispeed | 14 hours ago
nba456_ | 14 hours ago
d_tr | 15 hours ago
varispeed | 14 hours ago
cliche | 11 hours ago
johnfn | 14 hours ago
This is the "your iPhone is listening to you and serving ads based on what you say" of the 2020s.
esafak | 14 hours ago
4chandaily | 14 hours ago
Of course, it turned out that this wasn't actually completely BS. We just were accusing the wrong vendor. Not disagreeing with you on models.
johnfn | 12 hours ago
4chandaily | 11 hours ago
johnfn | 6 hours ago
ricardobeat | 13 hours ago
kzrdude | 15 hours ago
varispeed | 14 hours ago
Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.
EPWN3D | 11 hours ago
singingtoday | 6 hours ago
forgot-my-pw | 15 hours ago
Still quite impressive though.
walrus01 | 14 hours ago
ricardobeat | 13 hours ago
walrus01 | 13 hours ago
p1esk | 12 hours ago
llm_nerd | 10 hours ago
Eh, it isn't news. I mean, it's just an echo bit of news to the great K3 release.
ChrisArchitect | 14 hours ago
Cognition launches new SWE-2 model
https://news.ycombinator.com/item?id=49645443
captainregex | 12 hours ago
samusiam | 11 hours ago