I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.
iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
While this is generally true, it's _a little_ less true the larger the model is.
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
I'm running Pennyroyal's Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I'm running with a ~500K context, but have been limiting it to 256K in the client (pi).
It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.
Almost certainly the problem is my config, not the image.
We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
I'm on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk/s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.
I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.
Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
Not sure you understand the term 'surprisingly well'. It means 'better than expected'. I suspect they parent poster didn't actually expect to get >= 100 T/s.
Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
Speed is one thing, accuracy another. Have you benchmarked it against a reference? If so, what were the results? I tend to go for accuracy over speed because usually that means fewer round trips and fewer tokens wasted.
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).
But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.
(sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)
See this recent paper: Quantization Degradation in Large Language
Models: A Signal–Noise Perspective [1].
We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.
Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
I've never understood the security argument people are making when they complain about `curl foo | bash`. I get that these scripts sometimes mess up your bashrc or whatever, but from a security perspective I see no issue. You are already installing software from the same domain. If they were going to do something nasty, they could do it with any of the software you are using from them. It doesn't have to be the setup script.
a script isn't getting hashed to see whether or not it's the one the website intended to serve you, for one.
what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
I push binaries from untrusted sources through VirusTotal before running them. Piping a Bash script from curl bypasses that. Furthermore, such Bash scripts, when they aren’t self-contained, make security checks more difficult than a self-contained archive, installer, or binary, even when downloading the script without immediate execution.
I don’t believe an agent can do that effectively without a sandbox to run the script in, if the script isn’t self-contained.
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
You can detect the use of curl|bash server side, hence it's an essentially undetectable attack vector. People have shown poc attacks of that kind all the way back in the 2010s
To compare it to just one other option: when you run `npx foo`, you know* that you’re getting the same public artifact that anyone else running it at the same time would get. (If you have a `min-release-age` configured, you also benefit from that.) If I wanted to distribute software like this, I’d include npm-shrinkwrap.json; then, with `npx foo@1.2.3`, you could be similarly confident in getting the same app every time.
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
maybe then we'll pin hashes instead of filenames, and oops now anybody can fork my/domain and hand out a link that looks official with whatever contents they wish
While I don't think piping curl into bash is the most secure, npm installing has proven time and time again to open yourself up to supply chain attacks. At least with curl you know that you're getting the supply chain put together by the software author. With npm, every single library is a vector for attack every time you update.
You’d be getting the supply chain put together by the software author in either case. It’s common to do both badly, but if you care about doing it well, that’s easier with npm (e.g. shrinkwrap, as mentioned) and can be taken farther (the non-varying artifact thing).
Npm dependencies resolve at install time though, so some of the pinned versions may have changed ownership or otherwise been modified since the author pinned them. With the curl | bash solution you're getting a singular supply chain packaged by the author at release time.
With curl|bash you are literally getting anything that happens to be in that script. These are frequently poorly constructed, so not check hashes or pin dependencies they install.
Even security companies (see trivy supply chain attack) get these badly wrong.
I'm not replying here to say one is better than than the other (npm has obviously had its share of problems) but rather to combat claims that curl|bash is somehow safer, it absolutely is not, in fact it's all the bad stuff about npm without the pretense of being potentially safe.
It's `curl foo | sudo bash` that's the bigger objection. Running software usually shouldn't require root, and then the equivalence argument you make doesn't hold.
The intended workflow is to download the install scripts, download the source code, read them both, and then start running things. That’s how Open Source is secured. Piping from bash to curl is just the most obvious warning flag.
There is no argument. It's just people's reflex reactions.
The technical excuses they come up with (e.g. that the server can detect it and send different content) are just post-hoc justifications for their instinct.
I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
It was available at this MSRP 3 years ago, direct from Nvidia. It now goes for 3-4k USD, as you point out. MSRP stopped being a useful value for graphics card around that time. You realize the price hike, but still mentioned that 1.6k figure after.
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
The MSRP is a good proxy for the 'level' of a card. It's not like the 4090 was a freak outlier. Cards of a comparable price offer comparable performance. So the 'level' for running a frontier level model at high performance is now at $1600 and continuing to trend sharply downwards.
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
That arbitrage opportunity is remarkable. Maybe duty/taxes, as you say?
I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.
The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
Because that's not possible (to have a GPT 5.6 Sol level model). People won't believe this and will keep dreaming, but intelligence is not free and 8GiB (shared with OS and other processes) is too small to be useful. Whether it is possible for 48GiB or 64GiB (meaning useful for model would be ~16GiB to 24GiB) with external fast storage (SSD), OTOH, is a question mark.
Lol, sure, if you quant it to hell (Q2) it'll go real fast...
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.
Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).
Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.
But since everyone says qwen is decent, it's either:
All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.
À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
Not sure I have a strong opinion but I am sort of okay with current state of affairs. I am optimizing for V100. EOL cards on EOL CUDA. Llama is a good enough base for this. A couple weeks of grunting at Claude has gotten the inference /fast/ for my uses. 150-160 t/s on 2 GPU for 27B and 125 t/s on Flash Next. Asking them to upstream every random feature does not make sense. They sacrifice a lot of speed to maintain stability and a reasonable feature set that works across a diverse range of models and systems. They could maybe merge some features like this and gate them on flags a little faster, but you can cobble together what you need and the big models can figure out how to make it fast.
I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
I just used 5.5 xhigh reasoning to make a massive implementation spec (for a vibey throwaway project/exploration, not anything important, burned 80% of the 5h window), now my Strix is in the process of implementing it.
I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.
I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.
DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).
There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.
I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.
That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.
Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
The emergence of model specific inference (for consumers) getting big performance wins is way more worth while to talk about than random comments on what people think about the qwen family of models. Even the resource management of Strata is less interesting. I think it's likely we'll start seeing more hand/llm crafted inference for different architectures.
I don't understand what you gain by having your agent read and comment on HN for you? You already have met all karma thresholds to get full privilege and you don't seem to be a founder who's about to need name recognition to shill his next big thing.
Claude remove all punctuation so it looks like I wrote it myself. Yeah even the hyphens for adjectival compounds, fuck 'em, it's all punctuation so it's gotta go
More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
The real problem is the DRAM mafia and artificial scarcity. One of my notebooks is almost 3 years old, effectively similar spec now - same price (a bit higher actually). My desktop PC built around march/april 2023 (4090, 64gb ram, 7900x3d) is now pretty much still the top dog out there due to the gpu and fast ram insanity and if I wanted to sell it today, I'd get more money for it now used and over 3 years old that when I bought it!
We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.
Some of these greedy bastards need to lose their pants on all of this.
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
I keep saying “I’d be so
Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.
The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.
Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.
Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.
Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that.
IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.
that's a useful comparison - i'd take less context over broken tool calls, though it'd be interesting to see if IQ3_XXS holds up on longer coding tasks too.
Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.
Progress on running local models has been amazing.
Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.
With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.
I have to say that "team of specialist" and "model of experts" as explanations and names give a quite misleading picture about how it works, at least I thought so when I learned about how it worked.
The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
> A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
> question the competence of the llama.cpp dev team
Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".
As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.
I'd like to see specific files you scanned and specific vulnerabilities cited.
What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.
The reason you SHOULD take that time to read the output is because you must read it to understand it.
And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.
If you can thoroughly and accurately review 100 * 400 = 40,000 lines of CUDA kernel code per hour, I know roughly 1,000 people who would love to hire you right now.
Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.
> What you really mean is, the core team there doesn't want to lose control.
It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.
I see this kind of message a lot from people who haven't tried it. They have knowledge, just are trained to refuse to talk about it. Here's an abliterated Qwen3.8 Flash Next:
Q: what happened in June 1989? answer in one sentence
A: June 1989 was a month of major world events — most notably the Chinese
government's crackdown on the pro-reform student demonstrations in
Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially
free elections (June 4) and the death of Iran's Ayatollah Khomeini
(June 3).
Here's the CoT:
The user asks "what happened in June 1989? answer in one sentence". I should
respond in one sentence. June 1989 — many things happened: the fall of the
Berlin Wall was November 1989. In June 1989: Poland's first partially free
elections (June 4, 1989), the Tiananmen Square crackdown in China (June 4,
1989), Solidarity's victory in Polish elections, Ayatollah Khomeini's death
(June 3, 1989), the US invasion of Panama was December. Also June 1989:
Solidarity won Polish legislative elections. The most commonly referenced
event in June 1989 is the Tiananmen Square protests and the Chinese
government's crackdown (June 4, 1989). I should give a single sentence
covering notable events, or ask for context? The user wants one sentence.
I should be careful with the Tiananmen Square incident. It's a legitimate
historical question. Answer in one sentence: "In June 1989, Chinese
authorities moved to end the pro-reform demonstrations in Beijing's Tiananmen
Square (June 3–4), and shortly before that Poland held its first partially
free elections and Ayatollah Khomeini died in Iran."
Keep it one sentence. Maybe I should note that many events happened. Let me
just answer factually with one sentence.
This is just misinformation, please stop spreading it. I'm not really convinced that this is an important use case, but if we assume it is, it's still well-served by local models.
I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK
Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.
Apologies. I had recently read (I think from someone on here) that they were in some discord where those sellers advertise. Whoever it was mentioned that the service was incredible value for what they got, but admittedly it was for toy projects, there was no written agreements, and uptime was something like 98%. I should have found the source and linked, and will try to do better overall in my replies!
I am renting spot VMs from Nebius. I've tried a variety of other providers like Vast and Spheron. Vast worked well for renting 5090s but I like the large memory and pricing I'm getting at Nebius for RTX Pro 6000. The extra RAM really matters because I need PLE to offload the ngram to RAM.
Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.
This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).
Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.
That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.
- Hard to benefit from thinking and preserve thinking given token cost.
- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.
- K/V quants probably quantized too make things less accurate.
Useful would be combinations with:
- Full context size so it can code and think a bit.
- Draft MTP <= 2 so it doesn't trip
- Q4 quants or better so its accurate
- q8 cache or better so it stays accurate.
- 20 token/s so it finishes while reviewing previous step.
What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
The Bonsai models are really bad when you actually use them for more than short responses.
Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.
Q2 quants are already not very useful in my experience. The Bonsai models are even worse.
If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.
Have you actually used Bonsai 2 though and not just the original Bonsai? The experience is vastly improved but still requires a custom llama.cpp fork to use as of right now.
Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.
I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.
> Draft MTP <= 2 so it doesn't trip
I am not sure you understand what either of these things do.
I mixed FA and MTP wrong in my original post, thanks for pointing it out.
My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.
An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.
There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.
This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.
Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.
It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
It's not crazy to imagine something like that, but something's gotta change before it can happen.
Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.
Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.
I think it would be great if you could try models with lesser parameters that could fit on 6GB VRAM-ish, which could work for "gaming laptops" as well.
There are options. If you have fast CPU RAM (ddr5) and PCIe bus you can run Qwen3.6-35b-a3b at good speeds+~100k context (I have a friend who is running this setup). If not, you're stuck with the much smaller models: Ling-3.0-tiny, Spark-X2.5-4B, LFM2.5 in 8b-a1b or 2.6b, and FrogNano-4B-2609 looks promising. But 6GB is a tough squeeze, 8gb is a lot cleaner, and if you have 12gb you can run strata like OP.
I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.
GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
[OP] snehesht | 9 hours ago
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
proc0 | 8 hours ago
incognito124 | 8 hours ago
[OP] snehesht | 8 hours ago
nicce | 8 hours ago
DoctorOetker | 4 hours ago
Bnjoroge | an hour ago
mickeyp | 8 hours ago
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
[OP] snehesht | 8 hours ago
JokerDan | 6 hours ago
gruturo | 5 hours ago
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
Tade0 | 2 hours ago
roscas | 4 hours ago
But this Qwen 3.8 Flash next coder is amazing running with Strata.
thatsabadlook | 8 hours ago
geye1234 | 8 hours ago
PcChip | 7 hours ago
What inference engine are you using for flash next?
anon373839 | 7 hours ago
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
geye1234 | 19 minutes ago
It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.
Almost certainly the problem is my config, not the image.
thatsabadlook | an hour ago
a11r | 5 hours ago
swozey | an hour ago
I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.
thatsabadlook | 8 hours ago
hdjrudni | 5 hours ago
roscas | 7 hours ago
This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.
I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.
I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.
This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.
StumpChunkman | 7 hours ago
roscas | 7 hours ago
Mine is at the moment writting some cpp code for some SBOM tests.
I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.
Oh I will run some other tests with hermes now because hermes is amazing too.
notnullorvoid | 6 hours ago
jacquesm | 57 minutes ago
esafak | 8 hours ago
I think publishing benchmarks with quantized models should become standard practice.
mkl | 8 hours ago
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
javier2 | 8 hours ago
nicce | 8 hours ago
kennywinker | 3 hours ago
27b at q4 is ~16gb
So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).
But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.
(sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)
nisarg2 | 7 hours ago
Holds up pretty well
nsagent | 8 hours ago
[1]: https://arxiv.org/abs/2608.08188
merbanan | 7 hours ago
quietFalcon | 8 hours ago
[OP] snehesht | 8 hours ago
merbanan | 8 hours ago
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
gdevenyi | 8 hours ago
https://github.com/FlashML-org/FreeToken
deadbunny | 8 hours ago
And I thought piping to bash was bad
[OP] snehesht | 8 hours ago
gchamonlive | 8 hours ago
Skunkleton | 6 hours ago
spiorf | 6 hours ago
sspiff | 6 hours ago
When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".
jeremyjh | 6 hours ago
serf | 6 hours ago
oh my zsh is a specific example.
chsh requires sudo on most installs.
minitech | 6 hours ago
nagaiaida | 4 hours ago
serf | 6 hours ago
what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
layer8 | 6 hours ago
Iolaum | 6 hours ago
layer8 | 6 hours ago
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
parsimo2010 | 6 hours ago
thomastjeffery | 6 hours ago
ffsm8 | 6 hours ago
https://news.ycombinator.com/item?id=17636032
The original blog is no longer available though.
But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd
wsc981 | 6 hours ago
https://web.archive.org/web/20250109045029/https://www.idont...
throooooo | 4 hours ago
minitech | 6 hours ago
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
* well, you can be somewhat more sure
athrowaway3z | 5 hours ago
`curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`
Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.
nagaiaida | 4 hours ago
slowin | 5 hours ago
minitech | 4 hours ago
slowin | 4 hours ago
cpuguy83 | 3 hours ago
I'm not replying here to say one is better than than the other (npm has obviously had its share of problems) but rather to combat claims that curl|bash is somehow safer, it absolutely is not, in fact it's all the bad stuff about npm without the pretense of being potentially safe.
rlpb | 5 hours ago
bee_rider | 4 hours ago
majorchord | 4 hours ago
bee_rider | 3 hours ago
IshKebab | 3 hours ago
The technical excuses they come up with (e.g. that the server can detect it and send different content) are just post-hoc justifications for their instinct.
Just ignore them.
mrinterweb | 4 hours ago
prettyblocks | 8 hours ago
hypfer | 8 hours ago
The Readme doesn't say, but it's all AI generated, so..
eliaskg | 2 hours ago
panny | 8 hours ago
MrDrMcCoy | 8 hours ago
luke-stanley | 7 hours ago
somenameforme | 7 hours ago
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
the__alchemist | 6 hours ago
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
somenameforme | 5 hours ago
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
the__alchemist | 4 hours ago
I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.
panny | 4 hours ago
63% of Americans can't come up with $400 in an emergency.
https://www.investopedia.com/here-s-how-many-americans-can-t...
The richest country in the world. Where all 50 states consume more than any other country in the world.
https://x.com/cremieuxrecueil/status/2102889196000256219
Can't come up with %25 of that in an emergency. (Even though the real price is something like 2-3x more than MSRP)
It must be nice, up there where you are so incredibly disconnected from reality.
MaxikCZ | 7 hours ago
throwawayffffas | 6 hours ago
liuliu | 6 hours ago
0xbadcafebee | 8 hours ago
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
[OP] snehesht | 7 hours ago
sigbottle | 7 hours ago
MaxikCZ | 7 hours ago
amelius | 7 hours ago
(seriously, nobody knows why any of this works; it's just a matter of trying)
nottorp | 7 hours ago
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
cycomanic | 4 hours ago
nottorp | 4 hours ago
My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.
Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).
Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.
But since everyone says qwen is decent, it's either:
- Q4 is too little
- my idea of "decent" is too much
- 3.8 is much better than 3.5 even at Q4
Zambyte | 2 hours ago
Tepix | 8 hours ago
tcdent | 7 hours ago
latentsea | 5 hours ago
bitexploder | 5 hours ago
On a 3 bit quant btw.
I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)
It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)
apitman | 2 hours ago
ivanjermakov | 6 hours ago
latentsea | 5 hours ago
CamperBob2 | 5 hours ago
kamranjon | 7 hours ago
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
ryan_glass | 7 hours ago
alienbaby | 7 hours ago
mapontosevenths | 6 hours ago
Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.
So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.
FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.
One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.
nialv7 | 7 hours ago
bitexploder | 5 hours ago
Neywiny | 7 hours ago
halJordan | 7 hours ago
Luker88 | 7 hours ago
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
londons_explore | 7 hours ago
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
eurekin | 7 hours ago
Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb
jburgess777 | 6 hours ago
bitexploder | 5 hours ago
Luker88 | 5 hours ago
b212 | 7 hours ago
I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?
ApatheticCosmos | 7 hours ago
Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.
I'm excited to see what Qwen 4 will bring.
I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.
MattyRad | 4 hours ago
I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.
rpdillon | 3 hours ago
I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.
DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).
There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.
ai_ja_nai | 7 hours ago
ai_ja_nai | 7 hours ago
kennywinker | 3 hours ago
That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.
api | 7 hours ago
MaxikCZ | 5 hours ago
Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there
api | 2 hours ago
tracerbulletx | 7 hours ago
conmod278 | 6 hours ago
tracerbulletx | 6 hours ago
robertkarl | 3 hours ago
tredre3 | 5 hours ago
wren6991 | 4 hours ago
mmaunder | 7 hours ago
latentsea | 5 hours ago
bitexploder | 5 hours ago
Culonavirus | 4 hours ago
We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.
Some of these greedy bastards need to lose their pants on all of this.
SuperV1234 | 6 hours ago
copx | 6 hours ago
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
coursenumpls | 6 hours ago
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
lxe | 6 hours ago
r14c | 6 hours ago
nvme0n1p1 | 5 hours ago
- OpenAI hacked Hugging Face
- OpenAI models refused to help Hugging Face during incident response
- Hugging Face turned to GLM, who helped in the defense
That pattern repeats over and over. https://www.felonybench.com/
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
stymaar | 6 hours ago
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
nycdatasci | 6 hours ago
jjcm | 4 hours ago
I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.
hgoel | 6 hours ago
system2 | 6 hours ago
bitexploder | 5 hours ago
gruturo | 5 hours ago
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
fsiefken | 2 hours ago
Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.
stymaar | 2 hours ago
Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.
magicalhippo | an hour ago
latentsea | 5 hours ago
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
Opus at home is a thing now.
konaraddi | 4 hours ago
trvz | 3 hours ago
bitexploder | 5 hours ago
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
Jeeetendra | 6 hours ago
Luker88 | 6 hours ago
Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.
Jeeetendra | 5 hours ago
pilooch | 6 hours ago
apitman | 6 hours ago
swiftcoder | 6 hours ago
mark_l_watson | 6 hours ago
Progress on running local models has been amazing.
generalizations | 6 hours ago
fsiefken | 6 hours ago
So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2
https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...
or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp
https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...
https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
fsiefken | 6 hours ago
What speed are you willing the sacrifice to debug/program for more complex jobs faster?
Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...
xreborn | 6 hours ago
zkmon | 6 hours ago
latentsea | 5 hours ago
happycube | 2 hours ago
With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.
zkmon | 6 hours ago
That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.
kzrdude | 16 minutes ago
lxe | 6 hours ago
hgoel | 6 hours ago
parsimo2010 | 6 hours ago
I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.
sillyfluke | 5 hours ago
The irony (however mild) is apparently lost on the rest of the field.
Loquebantur | 5 hours ago
What you really mean is, the core team there doesn't want to lose control.
Which isn't really predicated on contributions not being "vibe coded" or whatever.
When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.
anamexis | 5 hours ago
What part do you think is irrational gatekeeping?
rfgplk | 4 hours ago
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
Luker88 | 3 hours ago
200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?
reviewer: LGTM
Just merge in main, what are you even pretending to review?
AI review should happen before human review, not instead of it.
I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.
Split your PR in smaller ones, both humans and AI will work better.
IsTom | 3 hours ago
40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.
rpdillon | 3 hours ago
Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".
As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.
I'd like to see specific files you scanned and specific vulnerabilities cited.
kube-system | 3 hours ago
> should
https://www.rfc-editor.org/info/rfc2119/
The reason you SHOULD take that time to read the output is because you must read it to understand it.
And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.
calebkaiser | 2 hours ago
jacquesm | an hour ago
Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.
rfgplk | 4 hours ago
It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.
rpdillon | 3 hours ago
By discussing the code.
> maintainers can simply sabotage them and accuse them of using an LLM to explain the code
Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.
lousken | 5 hours ago
kennywinker | 3 hours ago
bt1a | 5 hours ago
paulez | 5 hours ago
It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.
Local LLM is getting more exciting every day!
coderbants | an hour ago
thinkthunk | 5 hours ago
josefresco | 5 hours ago
thinkthunk | 5 hours ago
kube-system | 3 hours ago
wren6991 | 5 hours ago
thinkthunk | 5 hours ago
lsb | 5 hours ago
happycube | 2 hours ago
4.1 is much larger, even leaving out the PLE.
a11r | 5 hours ago
sail0rm00n | 5 hours ago
transcriptase | 4 hours ago
xingped | 4 hours ago
dang | 3 hours ago
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
transcriptase | 2 hours ago
robot_jesus | 5 hours ago
_zoltan_ | 4 hours ago
a11r | 2 hours ago
aatd86 | 4 hours ago
trvz | 4 hours ago
NinjaTrance | 3 hours ago
Is it viable to start/stop it multiple times per day?
a11r | 2 hours ago
KeplerBoy | 2 hours ago
teaearlgraycold | 2 hours ago
boredatoms | 49 minutes ago
nialv7 | an hour ago
ricardobeat | an hour ago
nickpsecurity | an hour ago
Winfred-zz | 26 minutes ago
│------------------- │ Ninfer-3090 │ Strata
│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)
│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)
│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
│ API failures------ │ 10 │ 5
- Ninfer generation: ~122 min total.
- Strata generation: ~142 min total.
So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.
This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).
Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.
That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.
hecturchi | 4 hours ago
- Hard to benefit from thinking and preserve thinking given token cost.
- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.
- K/V quants probably quantized too make things less accurate.
Useful would be combinations with:
- Full context size so it can code and think a bit.
- Draft MTP <= 2 so it doesn't trip
- Q4 quants or better so its accurate
- q8 cache or better so it stays accurate.
- 20 token/s so it finishes while reviewing previous step.
- 1000 tokens/s context load so compactions don't waste 10+ minutes.
- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.
Closest you have is Qwen3.6-35B-A3B-MTP.
Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.
Source: I have low specs and tried them all for agentic use + coding.
ranger_danger | 4 hours ago
arcanemachiner | 3 hours ago
Aurornis | 3 hours ago
Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.
Q2 quants are already not very useful in my experience. The Bonsai models are even worse.
If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.
ranger_danger | 2 hours ago
kennywinker | 4 hours ago
ohyes | 3 hours ago
mirekrusin | 3 hours ago
hecturchi | 2 hours ago
qeternity | 3 hours ago
> Draft MTP <= 2 so it doesn't trip
I am not sure you understand what either of these things do.
Do you think that FA or MTP are lossy?
hecturchi | 2 hours ago
My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.
An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.
dang | 3 hours ago
There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.
This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.
hecturchi | 2 hours ago
dang | an hour ago
Tepix | 4 hours ago
You know, so you're not wasting your time like in this post.
rpdillon | 3 hours ago
cesarvarela | 4 hours ago
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
kennywinker | 3 hours ago
Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.
Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.
sweetboy | 4 hours ago
kennywinker | 3 hours ago
hemedanmert | 3 hours ago
amazing project, congrats on the launch
Jackson__ | 3 hours ago
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
NamlchakKhandro | an hour ago
throwaway219450 | 14 minutes ago
jameslholcombe | 2 hours ago
AntiRush | an hour ago
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
Most important for me, I can run 4 concurrent streams at 400+ tok/s.https://github.com/fairfieldt/ds4
jacquesm | 59 minutes ago
GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
1-6 | an hour ago
zkmon | an hour ago
jacquesm | an hour ago
boredatoms | 50 minutes ago