Nah, we're long past the point where that would make a difference - if they could have done so effectively, they would have before K3 and GLM 5.3 were nipping at their heels...
Anti distillation is just American nationalistic marketing, the amount of data they have "distilled" makes no difference. Deepseek "distilled" like 10000 messages? That's clearly not enough. You gain a lot more doing RL in house.
Hopefully it will be open weights and have the same architecture and size as the current v4 flash vision, which is probably the best LLM that can be run on 128G devices.
IQ3_XXS (~3.2 BPW). For me this is an option because my Mac studio is only used for serving LLMs, so I can afford to dedicate most of its RAM to this. I can run with 256k context and only uses ~117G, with the remaining (up to 125G which I can allocate to VRAM) being used for prompt caching and context checkpoints.
I'm making my own quants, though the Vision-Exp version is outdated and won't work on llama.cpp master branch (I built it before llama added support):
Don't use my Vision-Exp GGUF though. As I said I built those GGUFs before llama.cpp supported, and they can't be loaded on current master (require my own branch).
I already have new GGUFs but haven't uploaded yet. If you want Vision-Exp, maybe use bartowski or unsloth's GGUFs.
Side note:
As an alternative to deepseek v4, you might want to give it a shot at qwen 3.8 flash next. I have IQ4_NL GGUFs that can be loaded fully into 128G, or Q5_K GGUFs that can offload the PLE to disk (use --load-mode none --lazy-mode on for that): https://huggingface.co/tarruda/Qwen3.8-Flash-Next-GGUF.
llama.cpp master is still somewhat bad in Qwen 3.8 next performance, but I was able to achieve 40tps tg and 600 tps pp on my private branch.
I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that route instead? I've been using ds4 for months now and it's already got support for the new vision model, haven't tried it yet, still on 0731 but it's been very solid for me.
I tried dwarfstar when llama.cpp DSV4 support was still very weak, and while it worked, I didn't see anything that would make me want to stick with it vs llama.cpp. llama.cpp is simply better with its awesome built-in webui, router and server APIs and certainly support much more models and quantizations than dwarfstar.
Since then, I started maintaining my own vibe coded dsv4 branch with metal optimizations, so I actually get much better metal performance on my llama.cpp branch than on dwarfstar (plus all the extra llama.cpp features). Here it is in case you want to give it a shot: https://github.com/tarruda/llama.cpp/tree/qwen4exp-dsv4-opti...
v4 flash has been working quite well for the majority of my personal projects, with occasional v4 pro or Kimi 3 for the most complicated tasks or to check the overall project progress (when vibe coding).
I must be doing something wrong. I gave v4 pro a try a couple of days ago, gave it a simple prompt like "clean up functions x and y in file z" and it would always start off promising, just to quickly get sidetracked, start hallucinating problems in the code, and just get stuck for hours until I interrupt it:
— hmm — 0x2D696370 — little-endian bytes: 70 63 69 2D = 'p','c','i','-' — hmm — WAIT — WAIT — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *HOLD ON — HOLD ON — HOLD ON — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — !!!!!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *OK — WAIT — I THINK I FINALLY SEE THE WHOLE PICTURE — I NEVER READ IT — AND — THE LAYOUT — hmm — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — HOLD ON — HOLD ON — HOLD ON — HOLD ON
Then gave the same to Sonnet 5 and it was done 15 - 30 minutes later.
I tried v4 pro both in claude code and codewhale with similar results. Haven't tried the new deepseek harness.
I've been using the 0731 Flash V4 model, via Opencode, and I've had no major issues. It feels very comparable to Opus 4.6/4.7, that I use at work. I haven't ran into any of the problem you mention, so that might be a quirk of V4 pro, the specific harness, or maybe the host you're using?
I guess this was a heavily quantized version from openrouter? I've never had that experience in the last months of quite intensive use of the official deepseek api.
No, that was directly from deepseek. And it happened very consistently (3 or 4 times in a row, clean session every time, and across 2 harnesses). Guess I'll give it another try when v4.1 flash is available.
> In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price. If you encounter any issues during your comparative testing between V4 Pro and V4.1 Flash, please do not hesitate to reach out to us with your feedback. Thank you for your support!
I will continue to be amazed by how much power you get from DeepSeek Flash for the cost. I have let that puppy lose on so many projects and it is has never let me down. It can build and entire Rails app in no time and even do the tests. For most things, I don't get why people pay the money for Claude. DeepSeek Flash is my default agent in Omarchy.
True, make them cheap and fast enough and you can scope and stack agents sufficiently that the error rate tends close enough to zero to be meaningfully useful.
Same for me. DeepSeek models are incredibly good at implementation and light planning. I still default to Opus models for feature planning, but for most simple features the Pro models suffice.
Incredible good value and product they have built.
My mental bias always kept me away from Chinese models. Because i know that china is a surveillance state and all the things we know about CCP. But after what we learned about OpenAI and how they most likely used user data to basically cheat in an open competition i think it does not matter which AI provider you use all of them will own your data and all of them can spy on you. So I am willing to switch to Chinese models. This way we help them develop and improve models some day we can run them locally.
These models are open-weights. Anyone can host them, you don’t have to use chinese servers even though most of them offer zero data-retention policies.
You're naive if you're thinking the scumbags running the US companies aren't using your data.
In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc.
Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol.
The new meta model is fast and very cheap as well, and when used through OpenCode you get quite a lot of free tokens. But meta is also THE surveillance company, so probably also not a good choice in your case.
Is it? You're still putting a lot of thought and guidance into the agent's harness, the final code is just a tiny bit of that. It's like giving a junior developer final code vs explaining the whys and nuance. Which I'm not sure I want to give Meta
The Chinese labs have released interesting papers to accompany their releases too, especially DeepSeek and Kimi. This improves their standing among a few of us, who really like to see and read the papers with details about what they have changed and how their models work.
Yes. I'm working in the agent industry and my god are we excited on new versions of Chinese flash models. The direct competition is Gemini Flash, and these models are much better on agentic tasks with fraction of the task price compared to Gemini. Things like oh here's a set of simple instructions for you to follow, call these tools, return this report. 20-30% of the price per task. And especially Deepseek Flash produces better quality than Gemini does.
Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.
If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.
Well, it's much more than that. In general everybody's building agents now. You see these things that can help you to do things like adding things like OCR an appointment from a picture of a hand-written paper and add it to your calendar, search things from the internet, find that email with a PDF and add it to your local paperless instance.
Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?
What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.
Which versions of flash and at what thinking levels? Which chinese flash models and at what thinking levels? What tasks? What completion rates? How was quality evaluated?
- Which versions: 3.6 vs 3.7 vs. 3.8 for Gemini Flash, and v4 0731 for Deepseek v4 Flash, and GLM 5.3 Flash
- Medium for Gemini, high for Deepseek.
- Things like find information, then understand something about it, then send a slack message or email etc.
- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini
- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.
Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.
Thanks... I have been trying to figure out some things. Been doing my own evals. Flash 3.8 does burn a lot more tokens on high. Interesting how smart and not smart it is. For personal use almost impossible to justify the cost of 3.8 Flash cost.
Deepseek also burns a lot of tokens, its output on high is 2x of Gemini on medium. But it's dirt-cheap so it still can be 60-70% cheaper.
From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.
All this really needs evals, the token prices tell nothing.
I am surprised at how well DSv4 flash does in the real world vs many benchmarks. You look at Flash 3.8 and it supposedly beats opus 5 and deepseek is far below.. but they were measuring efficiency, whatever that is…
Something doesn’t add up for me on the published benches
I recently had to config my harness to watch for cybersecurity flags from astra and funnel requests to flash when they occur because Astra gets queezy when you talk to it about UDP packets in games.
Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive.
Yeah, the latest batch of <256GB Chinese models are really nice. They're far less cryptic than Claude, and competent enough to feel almost near Opus. I canceled all my subscriptions and switched to running the Chinese models locally (not as a cost saving measure).
Do you use them for coding with your harness or do you use them in production? I found the latency distribution on OpenRouter to be unusable for DeepSeek v4 Flash.
I hope DeepSeek takes some time to improve their tuning for reasoning effort. Right now, there are only three reasoning efforts: low, high, and max.
For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting.
I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.
But, the web ui chat version of flash has very poor language following abilities in my experience:
You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results.
Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.
I finally uninstalled the app yesterday after giving it plenty of chances over several months. Yesterday, I asked it whether «DeepSeek has fixed the issue where it erroneously answers in Chinese?» and it answered in Chinese.
This shouldn’t be a user-facing issue. The web UI should inject the account’s language setting or solve it like competitors. They’ve mentioned giving it multiple chances but it’s still not fixed.
Oh, I’ve tried that too. It will promise to keep it in English from here on out, then switch back to Chinese after two or three exchanges. When ever it needs to do a web search, it seems to load so much Chinese text that it forgets any language instructions. Just thought my experience yesterday was more to the point. Right now the chat is absolutely hopeless.
I'm also totally not sure why it do that, but I guess because they're searching from China and web results comeback in Chinese so the model start using that.
I don’t know what model codex uses for session summarization (I use Pro subscription, no third party models), but I get Chinese summaries from time to time, when the only Chinese that could have appeared in the session would be an i18n strings file that it may or may not have loaded. Very puzzling. Last happened yesterday.
Yep, the same issue. I even defined a dictionary shortcut on my phone to expand aie to "Answer in English!", but every so often it takes 5 times to force it to switch to English.
Interesting though, when I ask questions in German or my native language, I rarely get Chinese answers. Looks like English is most affected.
It's not just web chat, V4 Flash 7/31 suffers from a lot of pathological behavior in coding harnesses as well, e.g. infinite loops, hallucinations, premature termination, and invalid tool calls.
Same. My side projects are coded almost exclusively with the Deepseek V4 Flash 07/31 in omp, and it recovers beautifully in every case. I'm using OpenCode Zen.
OpenCode but I variously use DeepSeek API/OpenRouter/Vercel AI gateway. I'm sure it's the combo of model + inference provider that is the issue and not the model alone. DeepSeek API also has far better inference speed and reliability than the cheapest providers. That said I never seem to have these issues when using GLM 5.3 flash served by OpenRouter/Vercel.
All of these flash models have this. You have to build your harness so that it deals with it. Infinite loops are solved by having an error message that says what to do differently on failure, invalid tool calls are solved by making the tool schema less strict and detect things in the runtime etc.
Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.
With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.
Can you share what tools and processes you're using to do this?
I've been using Pi to build custom extensions and wrapping workflows in shell processes to make it more deterministic and enforce certain validations, all guided by Fable. This isn't production work, though, just playing llm factorio at home.
What you want is a bunch of sessions to replay. Something anonymized if it's not yours, and something that's not depending on state.
You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.
Finally use a blind judge to check everything, and score the output.
Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.
I have used the flash model for over 3b tokens and ofc. I saw some hallucinations and premature termination (I also get this on Astra - way more often than with deepseek v4 flash), but I never had a infinite loop (using the copilot as harness).
I used a lot V4 flash to implement plans built by other models, and it was honestly top notch. The thing was a workhorse, and I got none of the isuses you describe.
I was mostly using DeepSeek on Pi, connecting to their API directly (not some third party provider).
I honestly have more issues steering Sonnet properly.
disagree; been using flash as my exclusive model (other contributors have used other models) to build a complicated software project, a web engine. See https://github.com/gterzian/formal-web, which as you can see comes with very specific guidance explaining how to implement features.
I'm using headless Pi with my own UI and sandbox client, https://github.com/gterzian/uni03C0, as well as a bunch of Pi extensions for things like accessing Web standards and browser use via CDP for testing.
Switching to 4.1 today...
Edit: it seems they pushed the date at which they route the Pro calls to new Flash, so today I ended up paying regular Pro rates thinking I was using the new Flash; an example of how their offering is not quite as predictable as I would like it to be (the other is cache performance being unpredictable).
You see this on Reddit where the bot accounts will just comment in German, French or Italian randomly (and other bot accounts responding to it won't even bat an eye, responding in English as if it's the most natural thing in the world)
Very odd. I've used DeepSeek heavily for some weeks, and haven't seen a single Chinese character either in its replies or its thinking. Are you using a quantified model or a different provider by any chance?
There is a chrome extension that injects “respond in English” and “English [checkbox emoji]” to every query. This helps a lot but I still sometimes get Chinese responses. I have not had this issue via api on openrouter.
This is the notice from DeepSeek regarding their API:
We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
-------------------------------------------------------
Hoje em sites como openrouter o valor é de $0.16 output .
Direct 1:1:1 comparison, for V4.1 Flash - V4 Pro 0813 - V4 Flash 0731
Input cache hits (per 1m tokens) - $0.003 Vs. $0.022 Vs. $0.007
Input cache miss (per 1m tokens) - $0.15 Vs. $0.66 Vs. $0.22
Output (per 1m tokens) - $0.6 Vs. $1.98 Vs. $0.66
This is taken from https://api-docs.deepseek.com/quick_start/pricing, and it's comparing only off-peak hours pricing. It looks like V4.1 Flash is cheaper than the current 0731 flash model, and much cheaper than the current V4 Pro model.
Like _aavaa_ said, make sure you're comparing the right vals 1:1. There's different costs for cache hit, cache misses, output tokens, etc. This one seems, during non-peak hours, cheaper. Peak hours are obviously more expensive, if they're gonna be 2x non-peak pricing. But, that might end up decreasing in the future.
In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price.
just in terms of user perception when selling this sort of service, this is what they call a "good look"
I've been using deepseek-v4-flash as a "worker" model with Claude Code to implement a tool using Rust/Iroh for my personal use, and it works fairly nicely when I use Opus as the planner/reviewer model. It seems to follow the plan generated by Opus, albeit with a few misses here and there that it cleans up later after being reviewed by Opus.
Fairly excited for the v4.1 launch. Input cache hit prices have been halved, which looks nice.
If you are okay with waiting use GLM 5.3 max. It costs more but still cheap. It is slow, but a very strong worker. Still dollars per day (at most) with heavy concurrent agent running. I load up planning and tasks in Opus or Sol, and just have glm flash workers go to town every night. My project has never advanced more smoothly.
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. There are some rumours on chinese forums talking about problems with the pretraining phase, so this is not mid/post training related.
I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...
Just my intuition about it but it does seem like a data issue.
V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus.
All that would suggest to me is that V4 Flash is capable of absorbing the data they’re throwing at it, and we’re still nowhere near the data limits of their larger 1.6T model
> all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price
If I'd carefully tested and optimized prompts against Pro I wouldn't be keen on this particular news. I feel like API model providers should lean towards not swapping out models on their paying customers, no matter how much "better" the new model is meant to be.
Sure, but if a company decided to place a remote chinese hedge fund's API at the center of a critical business workflow, this is a lesson better learned sooner rather than later.
I imagine they're doing this due to capacity issues or somesuch. They can always relaunch Pro later, meanwhile a little ricered benchmaxxing of their existing flash model provides a temporary cover story. They certainly aren't silly enough to think this won't impact existing Pro users </paranoia>
They can't increase capacity by raising prices. Right now they calculate that they're at an optimal revenue curve. Simply reducing demand by increasing prices doesn't mean they get more money. The limit is in the ability to buy hardware.
Absolutely nobody commenting that has done that, it's just roleplay. Flash-0731 and Pro-0813 replaced previous models too, and the V4-preview models replaced V3.2 before that. You have your official API from a tiny cutting edge research lab that can only realistically host one model at a time, which you know from every model they released before, but they fully openly provide every model so if you want that infinite stability you can easily host it yourself for eternity. So those commenters want to pretend they require RHEL-like stability for their prompts with enterprise budgets, but somehow can't host those models, yet at the same time offload their entire RHEL-stability requirements to the research lab. Even Google and OpenAI retire models that were still relevant a year ago, but open models actually last forever.
I've been watching a bunch of bycloud on YouTube recently, and although he's done a great job reviewing papers from the big AI labs, I feel like I'm missing something - how have all the labs seemingly made a model that's cheaper, faster AND has better performance? Historically `flash` variants (like codex spark as well) have been faster but perform worse
My problem with V4 flash is output limit. When I need to write or rewrite a larger file (~1000 lines of code) it will fail with message like output limit reached.
>In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price
Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.
At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
Usually I would very much agree with you, but those things are not deterministic so if that's an issue for you you're probably not making the right choices.
Nope, strong disagree. The model is one small part of the process harness; behaviors are usually routable with expected propensities. Unexpected model changes avoiding change management messes with monitoring and observability thresholds. Stochastic controls are a real thing when you have your distributions defined; your workflows on a new model will throw that expected prior out the window.
I think it depends on the workload, which put you in the right (in some case, it's dangerous to do that), but for me even for very clean always correct path I always assume a model can, at time, have a vector brain fart (because it keeps happening).
Also in principle it's similar to Anthropic downgrading.
Personally I use the basis that if I don't self host (I include remote host, but that I pay per hosting nor per model or api), it can change behavior without me asking. But they shouldn't, but it doesn't matter that's what they do.
They’re nondeterministic at a fine level, but can be “deterministic” at a more general level: e.g. you might know that one model will always return properly formatted json when asked. That might not be true of the replacement, even if it is in general “better” and cheaper.
Just the risk of such a thing means regression testing every time you update the model, and you want to be able to run that testing on your schedule rather than having it forced on you.
> They’re nondeterministic at a fine level, but can be “deterministic” at a more general level: e.g. you might know that one model will always return properly formatted json when asked. That might not be true of the replacement, even if it is in general “better” and cheaper.
This isn't true. Even Sol messes up JSON formatting for me on occasion.
Do not delude yourself into thinking these things are reliable. They are not.
Is nobody using structured outputs? They use constrained decoding at the generation stage to ensure the probability of tokens that would break the format are set to 0. I kinda figured everyone was doing this at this point.
If you use a large enough volume, you will know this isn't fully reliable. You might get json, and it might not match what the model actually sent because the last layer cut it up to match what you want. At the end, not json, or json but not really matching what the model wanted, it's sort of the same issue: when you use them you HAVE to assume they can have a brain fart. That's fine, just code around it.
Most serious providers are now supporting structured outputs in a reasonable way for all model configs. But for example on ollama structured outputs are still incompatible with tool calling and with reasoning
We use structured output and To my knowledge it has never failed (millions of data points). There seem to be two classes of people: those doing productive work with LLMs, and those who only get replies insulting their mothers…
Models are not deterministic, but they do have a flavor. When that flavor changes it can change the nature of output in a way that is undesirable.
Sort of like shooting a rifle - where the bullets hit is (to some order of magnitude, no philosophizing please) non-deterministic, but different very similar rifles will group differently and need to be appropriately adjusted to hit anything.
That flavor profile is known -- it's typical behavioral distribution is somewhat understood (and, often, common failure modes addressed). If JSON breaks about 20% of the time, and that drops for 2% or blows up to 90%, it can drive all sorts of issues (not the least, costs for retries).
Yeah but then it doesn't matter if it's "the model changed to another one" or "the model changed but it's the same name". Point is, you're not hosting it, as far as you know it can change at any moment, build around that idea. Is that great no, is that ideal no, that's why I self host (I include actually renting online the capacity and hosting the model myself on it).
You've added a similar response like 5 times, but I don't see you adding more information.
Yes - model hosts can do nasty things to you aside from changing the underlying model. That doesn't mean it's cool to have them change the model automatically.
Yes, it would be preferable to have complete control over your model serving, and no - not everyone is in a position to do that themselves.
Often via api you can pin to the specific snapshot. Yes the model provider can screw you over - it's physically possible for many vendors to screw you over, that doesn't mean that's not a dick move and that you can't expect/push for better behavior
That's a narrow take. Non-deterministic doesn't mean random; workflows can be reasonably validated and consistent to some known degree.
I work for an education department that serves a chatbot for students, and model changes go through painstaking content safety reviews. I initially assumed it's just a bunch of bureaucratic paranoia. But every other model upgrade has a measurably different adherence to the existing system prompts about not talking to the kids about sex and drugs and mental health issues.
I work in the same field for one of my company, in europe, and if you're not self hosting sorry but your worries are not something I can accept because models are very much not reliable on that front, let alone when you let the host decide HOW to serve a model (ressources allocated, different version of the same model, etc ...).
I'm not being a d**, just saying, the problem you have is something that I have faced EXACTLY, and at least here it's not working until you host in house or remote but on raw hardware. Otherwise it keeps having subtle changes, and you will notice no LLM API providers has guarantees about these.
It's not that you're a d*, it's just that you lack any kind of nuance
There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you
Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.
Sure, but if you're working in a field where that limitation is not "because you like it" but "because you have to" it doesn't matter. If you cannot assume it to be true, then you have to assume it isn't.
If self-hosting is the only solution you can think of to maintain model stability, it's no wonder you missed the point. We already do all those things you mentioned, even with cloud providers. Model versioning exists for exactly this purpose, just like it does with any software dependency. We can specify sonnet:v1.23 and decide if and when to upgrade.
But even if we only asked for sonnet:latest, the last thing we'd expect is opus. Model names should be indicative of breaking changes, and change management doesn't just go in the bin because of non-deterministic tools.
Crossing the street and Russian roulette both have non-deterministic risks of injury. And yet I would be bothered to find out that that on my way to work, I was playing Russian roulette by surprise.
Yes and if you decide to play not using a game rulebook but a website that call it "game A" you can't be shocked if "game A" switched from one to the other, even though the doc said opposite yesterday.
That would absolutely be a dick move on the part of the website and would confuse users.
Also, in this case, the game name is not “Game A” but something like “Deep Seek v4 Pro”, which they have previously chosen to use to describe Deep Seek v4 Pro, not Deep Seek v4.1 Flash.
The user cares about the distribution of outputs. That distribution is structurally determined by the distribution of inputs (i.e. prompts), the weights of the model, and (these days) the dynamics of the harness guiding successive generations.
The only way to characterize whether a choice is 'right' is to characterize the output distribution (i.e. evals)! Changing the underlying weights necessarily invalidates whatever characterization may have been done. One may assert that one's harness regularizes outputs back toward the desirable distribution, or one may hope the different weights induce a sufficiently similar output distribution.
But no, one should not be completely agnostic to the choice of weights just because there's some nondeterminism.
These models have specific behavioral characteristics trained into them from reinforcement learning and prompts optimized for one aren't guaranteed to transfer to the new generation. Think if the difference between gpt 5.4 and 5.5 and then 5.5 to 5.6 for example. 5.5 was "better" than 5.4 for struggled more across compaction boundaries and needed much more precise instructions before 5.6 sol recovered some of 5.4's ergonomics. All from the same lab but each model was trained with specific behavioral patterns that were basically product decisions. I would be quite annoyed to find that a model provider was routing a promt optimized for one model to a different one, especially for a dumber/cheaper non frontier model that's not going to be as good at just figuring out what you meant
It’s like being told one day that one of your teammates will be replaced tomorrow by another teammate who’s more capable (has a higher test score), regardless of how long you’ve already worked together and gotten used to each other.
In deterministic scenarios, this kind of switch can sometimes actually be safe, because we’ve used various engineering techniques to converge from non-deterministic behavior to deterministic decisions. On the other hand, in scenarios that are inherently non-deterministic, the impact of such a change is much harder to predict, so we need comprehensive evaluations to assess the extent of its impact.
You absolutely cannot consider an LLM production build number something to be pinned against as a static dependency in a product chain, so it's a non-issue.
In the openai api and many other rapid you can pretty trivially pin not only the model but also the specific snapshot you want to use by using the model id for that snapshot. It's not that big a deal and has been available for years
I'm the founder of a SaaS in the marketing industry with tens of thousands of paying users and we have lots of AI powered workflows where small changes between model versions can make some big differences so whenever we update models we have to thoroughly test them which is why we pin them at specific versions which has been a best practice (and common sense) for the last 3+ years
In this case, Deepseek organization is under a lot of pressure due to compute constraints. It would be better if they just throw a 404 instead of rerouting though so customers are not surprised by subtle changes in behavior.
Yes exactly. Automatic model downgrade seems horrible for a lot of production workloads, even if you are deterministically constraining the behavior of your agents.
Strict determinism is a different, but related, issue.
E.g. if I've written a role playing character using a specific model I may want to pin the character to that model until I've been able to test the model being "better" doesn't affect the feel of the character before switching. That doesn't mean I need the character's responses to be completely deterministic, but that doesn't imply I'm fine with the character having a different quality or feel of response just because the new model is out.
It'd be nice if there was a more explicit way to signal in the request "I want what you think is best per dollar for this class of answer" vs "I want this model to answer".
its probably more expensive to run, and I'm thinking 4.1 flash is a smaller more efficient model. You can always host your own. This is what they need to do to stay competitive.
I wonder if this is a sign of things to come for dirt-cheap model hosting: no servers running old versions, only new versions. Just to keep costs down.
I think keeping models around for a defined period of time is fine, but fracturing your model offerings like that (keeping around multiple versions of the same model) is very hard to do economically. The economics of the AI ecosystem are dominated by queueing theory constraints that make it extremely cheap to serve predictable traffic loads, and extremely expensive to serve unpredictable loads, and any time you split your offerings like that, you make both less predictable and therefore more expensive to serve both versions.
If I were paying anthropic prices, I'd expect it, but Deepseek is a super scrappy upstart in comparison and intentionally arbitraging on price. I would never expect them to do that.
I'm not sure that anyone is running production workloads against an API that bills twice as much for a chunk of the day. One of the best things about Deepseek is that you can host it yourself and get a ridiculous multiple of usage for what the same dollar amount would yield from their API
I don't really agree, but we shouldn't have to debate it. An Auto option at each level would preclude this kind of decisioning. Pick a discrete model, that's what you get.
Pick Auto (Deepseek v4 Flash Auto vs Deepseek v4.x Flash), and let the vendor decide. I think OpenRouter uses this method.
Crazy that they still keep the price -- or actually decrease it, even -- despite it is a big improvement. I hope it retains some of the tps speed of the preview release though, 300 tps means gemini flash is no longer "the fastest option" any more.
DeepSeek v4 Flash with high is already a really great work horse. Reliable. But this time, not only that it is better but they are reducing the price by 50% so that's great.
I also find the DeepSeek models to be more precise than Claude models (last I used 4.7) in that I yet had not the occasion where model did something unintentional that I did not direct it to.
What I mean is that price has been effectively halved.
During off-peak hours, the unit price is reduced from $0.007 for input cache hits to $0.003, $0.22 for input cache misses to $0.15, and $0.12 for output to $0.6
I've been trying out the 4.1 flash preview for some bulk tasks: it did a pretty good job refactoring a bunch of .metal kernels to .cu. It needed less steering than Opus on refactoring, IMO, and writes better comments. It failed to port a root exploit from modern Android to an older Pixel 3, but I suspect part of that might have been harness configuration (it asked me to give it a longer timeout for tasks at some point, but I didn't have a chance to finish that).
I was getting something like 300-400 tok/s which was just insanity. It was running so much faster than the toolcalls themselves. Honestly, even if it's not quite as strong in reasoning, it just throws so much so fast that it can do a lot more than you might expect.
What are people using in place of the cowork web/chrome integration? I find that to still be compelling reason to use Claude. Sometimes I have a menial task that involves a lot of web browsing/clicking/searching and it's much easier to let claude using my existing browser/login. I haven't find a replacement for it.
The price drop is probably more interesting than the benchmark improvement.
At these prices, you can start throwing Flash at a lot of small, repetitive tasks where you wouldn't even consider using a bigger model before. It feels like the interesting shift is not “Flash replaces Pro”, but “there are now a lot more things worth automating.”
Their free web interface has been upgraded to the new model. No more "expert mode" just a single interface; but it lost its ability to read PDFs that it could process normally yesterday. Probably a temporary thing.
neugls | 22 hours ago
oefrha | 22 hours ago
[OP] nickweb | 21 hours ago
NooneAtAll3 | 19 hours ago
[OP] nickweb | 18 hours ago
GreenWatermelon | 17 hours ago
swiftcoder | 21 hours ago
dude250711 | 20 hours ago
_aavaa_ | 19 hours ago
swiftcoder | 19 hours ago
alightsoul | 18 hours ago
tarruda | 21 hours ago
fluoridation | 21 hours ago
tarruda | 21 hours ago
I'm making my own quants, though the Vision-Exp version is outdated and won't work on llama.cpp master branch (I built it before llama added support):
- https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF
- https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-...
For the Vision-exp version, I also ran perplexity + KLD against the original MXFP4. Seems quite OK: https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-...
fluoridation | 20 hours ago
tarruda | 19 hours ago
I already have new GGUFs but haven't uploaded yet. If you want Vision-Exp, maybe use bartowski or unsloth's GGUFs.
Side note:
As an alternative to deepseek v4, you might want to give it a shot at qwen 3.8 flash next. I have IQ4_NL GGUFs that can be loaded fully into 128G, or Q5_K GGUFs that can offload the PLE to disk (use --load-mode none --lazy-mode on for that): https://huggingface.co/tarruda/Qwen3.8-Flash-Next-GGUF.
llama.cpp master is still somewhat bad in Qwen 3.8 next performance, but I was able to achieve 40tps tg and 600 tps pp on my private branch.
kamranjon | 19 hours ago
I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that route instead? I've been using ds4 for months now and it's already got support for the new vision model, haven't tried it yet, still on 0731 but it's been very solid for me.
tarruda | 18 hours ago
Since then, I started maintaining my own vibe coded dsv4 branch with metal optimizations, so I actually get much better metal performance on my llama.cpp branch than on dwarfstar (plus all the extra llama.cpp features). Here it is in case you want to give it a shot: https://github.com/tarruda/llama.cpp/tree/qwen4exp-dsv4-opti...
[OP] nickweb | 21 hours ago
Looks like the new model can be used if summoned via the API but the API won't list it.
igleria | 21 hours ago
As a consumer I feel like hansel and gretel combined, deepseek could be the witch.
throwaway473825 | 21 hours ago
calgoo | 21 hours ago
HansHamster | 20 hours ago
— hmm — 0x2D696370 — little-endian bytes: 70 63 69 2D = 'p','c','i','-' — hmm — WAIT — WAIT — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *HOLD ON — HOLD ON — HOLD ON — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — !!!!!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *OK — WAIT — I THINK I FINALLY SEE THE WHOLE PICTURE — I NEVER READ IT — AND — THE LAYOUT — hmm — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — HOLD ON — HOLD ON — HOLD ON — HOLD ON
Then gave the same to Sonnet 5 and it was done 15 - 30 minutes later. I tried v4 pro both in claude code and codewhale with similar results. Haven't tried the new deepseek harness.
k__ | 20 hours ago
It built this whole IaC plugin from scratch: https://github.com/fllstck/nebius-alchemy
KyleTheDev | 20 hours ago
atwrk | 18 hours ago
HansHamster | 16 hours ago
nicce | 21 hours ago
Wow. Imagine OpenAI/Google/Anthropic doing this! Nope.
thrownaway561 | 21 hours ago
bwfan123 | 19 hours ago
myaccountonhn | 2 hours ago
EbNar | 21 hours ago
ActionHank | 21 hours ago
I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale.
Oras | 18 hours ago
ActionHank | 17 hours ago
XzAeRosho | 21 hours ago
Incredible good value and product they have built.
darkoob12 | 21 hours ago
ricardobeat | 21 hours ago
m00dy | 19 hours ago
Yeah, that’s basically an industry-wide scam.
Xiol32 | 18 hours ago
epolanski | 20 hours ago
In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc.
Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol.
el_io | 20 hours ago
Mashimo | 20 hours ago
miroljub | 20 hours ago
hn8726 | 19 hours ago
jsw97 | 19 hours ago
kzrdude | 17 hours ago
efficax | 17 hours ago
pimeys | 20 hours ago
Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.
If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.
urieiejr | 19 hours ago
I make vaporware that doesnt do shit reliably and this chinese crap spouts plausible demos and spam calls more cheaply than the competition saaar
pimeys | 19 hours ago
Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?
What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.
bitexploder | 18 hours ago
pimeys | 18 hours ago
- Medium for Gemini, high for Deepseek.
- Things like find information, then understand something about it, then send a slack message or email etc.
- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini
- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.
Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.
bitexploder | 15 hours ago
pimeys | 15 hours ago
From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.
All this really needs evals, the token prices tell nothing.
bitexploder | 13 hours ago
serf | 20 hours ago
Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive.
nicce | 20 hours ago
hgoel | 17 hours ago
iamjs | 13 hours ago
postalcoder | 21 hours ago
For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting.
I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.
iamniels | 20 hours ago
alfiedotwtf | 15 hours ago
I just wish they kept parameter count down in order to fit entirely within commonly used RAM sizes
tarruda | 20 hours ago
jiehong | 21 hours ago
But, the web ui chat version of flash has very poor language following abilities in my experience:
You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results.
Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.
swiftcoder | 21 hours ago
gentlewater | 19 hours ago
elaus | 19 hours ago
michimagdesign | 18 hours ago
alightsoul | 18 hours ago
gentlewater | 18 hours ago
swiftcoder | 18 hours ago
jtbayly | 15 hours ago
fdsjgfklsfd | 14 hours ago
swiftcoder | 47 minutes ago
el_io | 20 hours ago
kgeist | 20 hours ago
epolanski | 20 hours ago
Hasn't happened in a while, last time was when I was testing fable 5 in june.
oefrha | 20 hours ago
apexalpha | 20 hours ago
I initially thought it was a trick, that using Chinese chars is somehow more info dense and it saves tokens to 'think' in Chinese.
But later on it became more erratic. I still wonder if token reduction would work that way.
miroljub | 20 hours ago
Interesting though, when I ask questions in German or my native language, I rarely get Chinese answers. Looks like English is most affected.
API never answers in Chinese.
tinyhouse | 20 hours ago
mattmcal | 20 hours ago
CharlesW | 19 hours ago
rpdillon | 18 hours ago
mattmcal | 16 hours ago
bendangelo | 19 hours ago
pimeys | 19 hours ago
Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.
With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.
tempoponet | 16 hours ago
I've been using Pi to build custom extensions and wrapping workflows in shell processes to make it more deterministic and enforce certain validations, all guided by Fable. This isn't production work, though, just playing llm factorio at home.
pimeys | 15 hours ago
You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.
Finally use a blind judge to check everything, and score the output.
Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.
K0IN | 18 hours ago
surgical_fire | 14 hours ago
I was mostly using DeepSeek on Pi, connecting to their API directly (not some third party provider).
I honestly have more issues steering Sonnet properly.
anramon | 14 hours ago
polyglotfacto | 5 hours ago
I'm using headless Pi with my own UI and sandbox client, https://github.com/gterzian/uni03C0, as well as a bunch of Pi extensions for things like accessing Web standards and browser use via CDP for testing.
Switching to 4.1 today...
Edit: it seems they pushed the date at which they route the Pro calls to new Flash, so today I ended up paying regular Pro rates thinking I was using the new Flash; an example of how their offering is not quite as predictable as I would like it to be (the other is cache performance being unpredictable).
lampe3 | 18 hours ago
I take the free chat gpt one writ with it in polish suddenly english.
djeastm | 18 hours ago
prussia | 18 hours ago
cheema33 | 18 hours ago
vintermann | 2 hours ago
red_green_yell | 15 hours ago
hinow | 20 hours ago
_aavaa_ | 20 hours ago
hinow | 19 hours ago
We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
------------------------------------------------------- Hoje em sites como openrouter o valor é de $0.16 output .
_aavaa_ | 19 hours ago
0.007 -> 0.003
0.22 -> 0.15
0.66 -> 0.60
Each one is now cheaper.
[0]: https://api-docs.deepseek.com/quick_start/pricing/
KyleTheDev | 19 hours ago
Input cache hits (per 1m tokens) - $0.003 Vs. $0.022 Vs. $0.007
Input cache miss (per 1m tokens) - $0.15 Vs. $0.66 Vs. $0.22
Output (per 1m tokens) - $0.6 Vs. $1.98 Vs. $0.66
This is taken from https://api-docs.deepseek.com/quick_start/pricing, and it's comparing only off-peak hours pricing. It looks like V4.1 Flash is cheaper than the current 0731 flash model, and much cheaper than the current V4 Pro model.
petu | 18 hours ago
You're comparing different providers then. DeepSeek price on OpenRouter is $0.66 output.
KyleTheDev | 20 hours ago
VulgarExigency | 19 hours ago
KyleTheDev | 19 hours ago
hinow | 19 hours ago
tensegrist | 20 hours ago
indigodaddy | 20 hours ago
ComputerGuru | 19 hours ago
a-ve | 20 hours ago
Fairly excited for the v4.1 launch. Input cache hit prices have been halved, which looks nice.
bitexploder | 18 hours ago
ThouYS | 20 hours ago
vib08 | 19 hours ago
NitpickLawyer | 20 hours ago
I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...
pixelesque | 19 hours ago
The announcement specifically says 4.1 Pro will be released in the future.
petu | 18 hours ago
Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same situation, but Pro is taken offline this time.
wolttam | 19 hours ago
V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus.
All that would suggest to me is that V4 Flash is capable of absorbing the data they’re throwing at it, and we’re still nowhere near the data limits of their larger 1.6T model
simonw | 20 hours ago
If I'd carefully tested and optimized prompts against Pro I wouldn't be keen on this particular news. I feel like API model providers should lean towards not swapping out models on their paying customers, no matter how much "better" the new model is meant to be.
tjwebbnorfolk | 20 hours ago
k__ | 19 hours ago
badatnames | 18 hours ago
dandaka | 17 hours ago
gpugreg | 16 hours ago
chronogram | 14 hours ago
surgical_fire | 14 hours ago
I very much prefer they don't raise prices.
vintermann | 2 hours ago
notatoad | 16 hours ago
Anybody actually using deepseek in a production system affected by this want to share their experience?
chronogram | 15 hours ago
edude03 | 20 hours ago
_3u10 | 20 hours ago
Its cost is now 1/10th per token, and 1/5th per task.
Basically they have shitty hardware so they have to do a lot of optimization. Think of it like replacing an O(n) algorithm with O(log n).
Anthropic / Open AI think the best path is the most intelligent models deepseek is more focused on tok/$
stanac | 19 hours ago
surgical_fire | 14 hours ago
All DeepSeek models have 384k maximum output tokens:
https://api-docs.deepseek.com/quick_start/pricing
k__ | 19 hours ago
https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/
I hope some of those speed increases will make it to production.
esafak | 19 hours ago
aftbit | 19 hours ago
Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.
At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
m3kw9 | 19 hours ago
KoolKat23 | 19 hours ago
But I agree with you. I have a dumb workflow that worked well with v4-flash-0731 and I suspect is directing to a newer model that now breaks it.
petu | 19 hours ago
nolok | 19 hours ago
tomrod | 19 hours ago
nolok | 18 hours ago
Also in principle it's similar to Anthropic downgrading.
Personally I use the basis that if I don't self host (I include remote host, but that I pay per hosting nor per model or api), it can change behavior without me asking. But they shouldn't, but it doesn't matter that's what they do.
gcanyon | 19 hours ago
Just the risk of such a thing means regression testing every time you update the model, and you want to be able to run that testing on your schedule rather than having it forced on you.
packetlost | 18 hours ago
This isn't true. Even Sol messes up JSON formatting for me on occasion.
Do not delude yourself into thinking these things are reliable. They are not.
kamranjon | 18 hours ago
nolok | 18 hours ago
wongarsu | 18 hours ago
Zopieux | 14 hours ago
gcanyon | 17 hours ago
idiotsecant | 19 hours ago
Sort of like shooting a rifle - where the bullets hit is (to some order of magnitude, no philosophizing please) non-deterministic, but different very similar rifles will group differently and need to be appropriately adjusted to hit anything.
tomrod | 18 hours ago
That flavor profile is known -- it's typical behavioral distribution is somewhat understood (and, often, common failure modes addressed). If JSON breaks about 20% of the time, and that drops for 2% or blows up to 90%, it can drive all sorts of issues (not the least, costs for retries).
nolok | 18 hours ago
switchbak | 17 hours ago
Yes - model hosts can do nasty things to you aside from changing the underlying model. That doesn't mean it's cool to have them change the model automatically.
Yes, it would be preferable to have complete control over your model serving, and no - not everyone is in a position to do that themselves.
vikramkr | 15 hours ago
lkois | 18 hours ago
I work for an education department that serves a chatbot for students, and model changes go through painstaking content safety reviews. I initially assumed it's just a bunch of bureaucratic paranoia. But every other model upgrade has a measurably different adherence to the existing system prompts about not talking to the kids about sex and drugs and mental health issues.
nolok | 17 hours ago
I'm not being a d**, just saying, the problem you have is something that I have faced EXACTLY, and at least here it's not working until you host in house or remote but on raw hardware. Otherwise it keeps having subtle changes, and you will notice no LLM API providers has guarantees about these.
frde_me | 17 hours ago
There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you
Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.
nolok | 17 hours ago
lkois | 7 hours ago
But even if we only asked for sonnet:latest, the last thing we'd expect is opus. Model names should be indicative of breaking changes, and change management doesn't just go in the bin because of non-deterministic tools.
WhyNotHugo | 18 hours ago
Replacing a six-sided die for an eight-sided die also keeps rolls non-deterministic.
That doesn't mean it's fine to just replace the dice mid-game.
disiplus | 17 hours ago
neodymiumphish | 17 hours ago
hyperpape | 17 hours ago
nolok | 17 hours ago
hyperpape | 16 hours ago
Also, in this case, the game name is not “Game A” but something like “Deep Seek v4 Pro”, which they have previously chosen to use to describe Deep Seek v4 Pro, not Deep Seek v4.1 Flash.
genxy | 12 hours ago
gpugreg | 17 hours ago
qeternity | 14 hours ago
They think that sampling is an inherent part of Transformers.
Even on this site, it is regurgitated with confidence.
kristjansson | 17 hours ago
The only way to characterize whether a choice is 'right' is to characterize the output distribution (i.e. evals)! Changing the underlying weights necessarily invalidates whatever characterization may have been done. One may assert that one's harness regularizes outputs back toward the desirable distribution, or one may hope the different weights induce a sufficiently similar output distribution.
But no, one should not be completely agnostic to the choice of weights just because there's some nondeterminism.
vikramkr | 15 hours ago
allenxu | 5 hours ago
allenxu | 5 hours ago
weego | 19 hours ago
weird-eye-issue | 18 hours ago
Xunjin | 18 hours ago
vikramkr | 15 hours ago
weird-eye-issue | 5 hours ago
samuelknight | 18 hours ago
samuelknight | 18 hours ago
DetroitThrow | 17 hours ago
jmathai | 18 hours ago
cyanydeez | 18 hours ago
jmathai | 18 hours ago
ycui7 | 18 hours ago
if you want deterministic returns, you should set the temperature to 0 to get the best possibility of deterministic.
zamadatix | 18 hours ago
E.g. if I've written a role playing character using a specific model I may want to pin the character to that model until I've been able to test the model being "better" doesn't affect the feel of the character before switching. That doesn't mean I need the character's responses to be completely deterministic, but that doesn't imply I'm fine with the character having a different quality or feel of response just because the new model is out.
It'd be nice if there was a more explicit way to signal in the request "I want what you think is best per dollar for this class of answer" vs "I want this model to answer".
g023 | 18 hours ago
bicx | 18 hours ago
darksaints | 17 hours ago
If I were paying anthropic prices, I'd expect it, but Deepseek is a super scrappy upstart in comparison and intentionally arbitraging on price. I would never expect them to do that.
monster_truck | 17 hours ago
happycube | 16 hours ago
If they were still at original price I'd get a couple of DGX Sparks myself to run Flash models at a decent quant/context combo.
Sha1rholder | 14 hours ago
I'm not sure that anyone will mind running production workloads against an API that bills half as much for a chunk of the day.
teamv02 | 9 hours ago
Pick Auto (Deepseek v4 Flash Auto vs Deepseek v4.x Flash), and let the vendor decide. I think OpenRouter uses this method.
aftbit | 19 hours ago
nicman23 | 19 hours ago
hope deepseek makes me change my setup again
damsta | 19 hours ago
While V4.1 Flash performance and cost looks promising this auto re-routing sounds concerning
throwa356262 | 3 hours ago
coopykins | 19 hours ago
npn | 19 hours ago
wg0 | 19 hours ago
I also find the DeepSeek models to be more precise than Claude models (last I used 4.7) in that I yet had not the occasion where model did something unintentional that I did not direct it to.
EDIT: Updated percentage reduction.
kennywinker | 19 hours ago
wg0 | 19 hours ago
During off-peak hours, the unit price is reduced from $0.007 for input cache hits to $0.003, $0.22 for input cache misses to $0.15, and $0.12 for output to $0.6
riknos314 | 18 hours ago
0.15/0.22 ≈ 0.68, meaning a roughly 32% reduction on inputs. The 50% reduction is only outputs and cached inputs.
kakacik | 18 hours ago
c0rruptbytes | 18 hours ago
eli | 18 hours ago
It’s good and very fast.
(Note that the deepseek API trains on your data)
mmastrac | 18 hours ago
I was getting something like 300-400 tok/s which was just insanity. It was running so much faster than the toolcalls themselves. Honestly, even if it's not quite as strong in reasoning, it just throws so much so fast that it can do a lot more than you might expect.
I'd say it was comparable with GLM5.3 Flash.
mrbonner | 18 hours ago
nullbio | 17 hours ago
You'd think it would have been something they did a year ago, but here we are. Still.
declan_roberts | 16 hours ago
mermadicsolutio | 15 hours ago
At these prices, you can start throwing Flash at a lot of small, repetitive tasks where you wouldn't even consider using a bigger model before. It feels like the interesting shift is not “Flash replaces Pro”, but “there are now a lot more things worth automating.”
Axonis | 2 hours ago