We will adjust the pricing for the Flash series effective from 12:00 Beijing Time on September 10, 2026. During off-peak hours, the unit price will be $0.003 for input cache hits, $0.15 for input cache misses, and $0.6 for output. Peak-hour prices will be double the off-peak rates. Please plan your usage accordingly.
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro (self)
aftbit 6 hours ago
Please don't do this kind of thing. If a user has validated a workflow on V4 Pro, they might not want to suddenly start testing it in production on V4.1 Flash. Instead, keep V4 Pro around but deprecated for a defined period of time, then remove it.
At least as open weights models, it's possible to use something like Together.ai or OpenRouter to run the V4 Pro model as long as other providers keep it up.
m3kw9 5 hours ago
KoolKat23 5 hours ago
But I agree with you. I have a dumb workflow that worked well with v4-flash-0731 and I suspect is directing to a newer model that now breaks it.
petu 5 hours ago
nolok 5 hours ago
tomrod 5 hours ago
nolok 4 hours ago
Also in principle it's similar to Anthropic downgrading.
Personally I use the basis that if I don't self host (I include remote host, but that I pay per hosting nor per model or api), it can change behavior without me asking. But they shouldn't, but it doesn't matter that's what they do.
gcanyon 5 hours ago
Just the risk of such a thing means regression testing every time you update the model, and you want to be able to run that testing on your schedule rather than having it forced on you.
packetlost 5 hours ago
This isn't true. Even Sol messes up JSON formatting for me on occasion.
Do not delude yourself into thinking these things are reliable. They are not.
kamranjon 4 hours ago
nolok 4 hours ago
wongarsu 4 hours ago
Zopieux 19 minutes ago
gcanyon 3 hours ago
idiotsecant 5 hours ago
Sort of like shooting a rifle - where the bullets hit is (to some order of magnitude, no philosophizing please) non-deterministic, but different very similar rifles will group differently and need to be appropriately adjusted to hit anything.
tomrod 4 hours ago
That flavor profile is known -- it's typical behavioral distribution is somewhat understood (and, often, common failure modes addressed). If JSON breaks about 20% of the time, and that drops for 2% or blows up to 90%, it can drive all sorts of issues (not the least, costs for retries).
nolok 4 hours ago
switchbak 3 hours ago
Yes - model hosts can do nasty things to you aside from changing the underlying model. That doesn't mean it's cool to have them change the model automatically.
Yes, it would be preferable to have complete control over your model serving, and no - not everyone is in a position to do that themselves.
vikramkr an hour ago
lkois 4 hours ago
I work for an education department that serves a chatbot for students, and model changes go through painstaking content safety reviews. I initially assumed it's just a bunch of bureaucratic paranoia. But every other model upgrade has a measurably different adherence to the existing system prompts about not talking to the kids about sex and drugs and mental health issues.
nolok 3 hours ago
I'm not being a d**, just saying, the problem you have is something that I have faced EXACTLY, and at least here it's not working until you host in house or remote but on raw hardware. Otherwise it keeps having subtle changes, and you will notice no LLM API providers has guarantees about these.
frde_me 3 hours ago
There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you
Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.
nolok 3 hours ago
WhyNotHugo 4 hours ago
Replacing a six-sided die for an eight-sided die also keeps rolls non-deterministic.
That doesn't mean it's fine to just replace the dice mid-game.
disiplus 3 hours ago
neodymiumphish 3 hours ago
hyperpape 3 hours ago
nolok 3 hours ago
hyperpape 2 hours ago
Also, in this case, the game name is not “Game A” but something like “Deep Seek v4 Pro”, which they have previously chosen to use to describe Deep Seek v4 Pro, not Deep Seek v4.1 Flash.
gpugreg 3 hours ago
> those things are not deterministic
Determinism was an explicit goal of DeepSeek-V4. From their paper: https://arxiv.org/html/2606.19348v1#S3.SS3 > we implement end-to-end, bitwise batch-invariant, and deterministic kernels with minimal performance overhead
Of course, providers may not implement deterministic inference for various reasons, but it is possible.qeternity 28 minutes ago
They think that sampling is an inherent part of Transformers.
Even on this site, it is regurgitated with confidence.
kristjansson 3 hours ago
The only way to characterize whether a choice is 'right' is to characterize the output distribution (i.e. evals)! Changing the underlying weights necessarily invalidates whatever characterization may have been done. One may assert that one's harness regularizes outputs back toward the desirable distribution, or one may hope the different weights induce a sufficiently similar output distribution.
But no, one should not be completely agnostic to the choice of weights just because there's some nondeterminism.
vikramkr an hour ago
weego 5 hours ago
weird-eye-issue 5 hours ago
Xunjin 4 hours ago
vikramkr an hour ago
samuelknight 4 hours ago
samuelknight 4 hours ago
DetroitThrow 3 hours ago
jmathai 4 hours ago
cyanydeez 4 hours ago
jmathai 4 hours ago
ycui7 4 hours ago
if you want deterministic returns, you should set the temperature to 0 to get the best possibility of deterministic.
zamadatix 4 hours ago
E.g. if I've written a role playing character using a specific model I may want to pin the character to that model until I've been able to test the model being "better" doesn't affect the feel of the character before switching. That doesn't mean I need the character's responses to be completely deterministic, but that doesn't imply I'm fine with the character having a different quality or feel of response just because the new model is out.
It'd be nice if there was a more explicit way to signal in the request "I want what you think is best per dollar for this class of answer" vs "I want this model to answer".
g023 4 hours ago
bicx 4 hours ago
darksaints 3 hours ago
If I were paying anthropic prices, I'd expect it, but Deepseek is a super scrappy upstart in comparison and intentionally arbitraging on price. I would never expect them to do that.
monster_truck 3 hours ago
happycube 3 hours ago
If they were still at original price I'd get a couple of DGX Sparks myself to run Flash models at a decent quant/context combo.
Sha1rholder 18 minutes ago
I'm not sure that anyone will mind running production workloads against an API that bills half as much for a chunk of the day.
jiehong 7 hours ago
But, the web ui chat version of flash has very poor language following abilities in my experience:
You may ask it something in English, and get a thinking chain in Chinese with an answer in Chinese, or an English thinking chain and an English answer. Using the retry button on the same question has a 50/50 chance of any of those results.
Sometimes, asking something in English, but where information are mostly in another language may make the answer in the language where data has been found. The other day, I asked something about a local German thing, in English, and I got an answer in German instead. It’s as if all the language data stirred it away from the language of the user’s question.
swiftcoder 7 hours ago
gentlewater 5 hours ago
elaus 5 hours ago
michimagdesign 5 hours ago
alightsoul 4 hours ago
gentlewater 4 hours ago
swiftcoder 4 hours ago
jtbayly an hour ago
fdsjgfklsfd 37 minutes ago
el_io 7 hours ago
kgeist 6 hours ago
epolanski 7 hours ago
Hasn't happened in a while, last time was when I was testing fable 5 in june.
oefrha 6 hours ago
apexalpha 6 hours ago
I initially thought it was a trick, that using Chinese chars is somehow more info dense and it saves tokens to 'think' in Chinese.
But later on it became more erratic. I still wonder if token reduction would work that way.
miroljub 6 hours ago
Interesting though, when I ask questions in German or my native language, I rarely get Chinese answers. Looks like English is most affected.
API never answers in Chinese.
tinyhouse 6 hours ago
mattmcal 6 hours ago
CharlesW 6 hours ago
rpdillon 4 hours ago
mattmcal 2 hours ago
bendangelo 6 hours ago
pimeys 5 hours ago
Hallucinations you can't fix. Gemini is a bit worse there than DeepSeek, but there's not much research on how to fix that. The only one is the CaMeL paper by Google, where you tag every prompt and result and then for every assistant response or tool call you first check where it got that data and error if you notice fabrication. This one is really annoying to implement.
With larger models the fabrication starts when the context grows or if you have too many tools, for flash models it's much earlier. We use the flash models for repetitive agentic tasks, where the prompt defines clearly what to do and how. The whole run is about 4-5 steps typically, and context size stays in the comfort zone.
tempoponet 2 hours ago
I've been using Pi to build custom extensions and wrapping workflows in shell processes to make it more deterministic and enforce certain validations, all guided by Fable. This isn't production work, though, just playing llm factorio at home.
pimeys an hour ago
You replay all your sessions against your harness, and then store all logs all output, everything to a safe place.
Finally use a blind judge to check everything, and score the output.
Then fix your harness, iterate again until better until you are in a point where it's just the model's weakness. If you get to that, use a bigger model.
K0IN 5 hours ago
surgical_fire 7 minutes ago
I was mostly using DeepSeek on Pi, connecting to their API directly (not some third party provider).
I honestly have more issues steering Sonnet properly.
lampe3 5 hours ago
I take the free chat gpt one writ with it in polish suddenly english.
djeastm 4 hours ago
prussia 4 hours ago
cheema33 4 hours ago
red_green_yell an hour ago
simonw 6 hours ago
If I'd carefully tested and optimized prompts against Pro I wouldn't be keen on this particular news. I feel like API model providers should lean towards not swapping out models on their paying customers, no matter how much "better" the new model is meant to be.
tjwebbnorfolk 6 hours ago
k__ 5 hours ago
badatnames 4 hours ago
dandaka 3 hours ago
gpugreg 2 hours ago
chronogram 36 minutes ago
notatoad 2 hours ago
Anybody actually using deepseek in a production system affected by this want to share their experience?
chronogram an hour ago
oefrha 8 hours ago
nickweb 7 hours ago
NooneAtAll3 5 hours ago
nickweb 4 hours ago
GreenWatermelon 3 hours ago
EbNar 7 hours ago
ActionHank 7 hours ago
I don't need a model that can invent new mathematics. I need something that is fast, cheap, and consistent. Give me that and I can build and scale.
Oras 4 hours ago
ActionHank 3 hours ago
XzAeRosho 7 hours ago
Incredible good value and product they have built.
darkoob12 7 hours ago
ricardobeat 7 hours ago
m00dy 5 hours ago
Yeah, that’s basically an industry-wide scam.
Xiol32 4 hours ago
epolanski 7 hours ago
In any case old rules apply: if privacy is a concern don't share the data. I share all my work-related code because it's worthless, but I don't and would never share company business and process details, access to production/user data, etc.
Meanwhile I know of people connecting all the kind of MCPs for datadog/sentry/jira/concluce/production databases to their harnessess..lol.
el_io 7 hours ago
Mashimo 6 hours ago
miroljub 6 hours ago
hn8726 6 hours ago
jsw97 6 hours ago
kzrdude 3 hours ago
efficax 3 hours ago
pimeys 6 hours ago
Where Gemini still wins is non-text input what Deepseek cannot do, yet, and Deepseek Flash has this thing of cheaper models where a failing tool call can derail your agent to a retry loop if you're not careful on instructions in the error message.
If they fix and make the tool calls to work better in non-optimal situations, it's much easier to switch from Gemini without a few weeks of evals and bugfixing.
urieiejr 5 hours ago
I make vaporware that doesnt do shit reliably and this chinese crap spouts plausible demos and spam calls more cheaply than the competition saaar
pimeys 5 hours ago
Building an agent like this by yourself is really easy. Now, we have Gemini's subscription, OpenAI's ChatGPT subscription and all those, 20 bucks a month right?
What if you can spend that 20 bucks in tokens to do your own. And you pay 15 bucks _a year_ in tokens to run that? And you own the data, you own your code and integrations. It's really easy to do, and these flash models are _more than enough_ for simple agentic tasks.
bitexploder 4 hours ago
pimeys 4 hours ago
- Medium for Gemini, high for Deepseek.
- Things like find information, then understand something about it, then send a slack message or email etc.
- Completion rates somewhere in 80-90%, Deepseek a bit better than Gemini
- Quality evaluated by Fable 5.1 and Astra 6.0 acting as a rubric judge.
Gemini quality would probably be better with high thinking level, but that would be 40% more expensive. And Deepseek is already third the price of Gemini.
bitexploder 2 hours ago
pimeys 41 minutes ago
From the large models Kimi K3 is definitely the one burning the smallest amount of tokens. Even if you pay for the fast version in Fireworks it's third of the price of Opus 5 for the same task.
All this really needs evals, the token prices tell nothing.
serf 6 hours ago
Works fantastic. Glad there is a more 'uncensored' thing to fall back to when the frontier folk are too sensitive.
nicce 6 hours ago
hgoel 3 hours ago
mermadicsolutio 2 hours ago
At these prices, you can start throwing Flash at a lot of small, repetitive tasks where you wouldn't even consider using a bigger model before. It feels like the interesting shift is not “Flash replaces Pro”, but “there are now a lot more things worth automating.”
postalcoder 7 hours ago
For all intents and purposes, "low" is pretty much the same as turning reasoning off, and "high" is similar to "max". "High/max" performs way too much reasoning, takes forever, and causes costs to balloon. They need a proper "medium" setting.
I get it that they're probably focused on pushing performance right now, but the ergonomics of the model aren't great.
iamniels 7 hours ago
alfiedotwtf an hour ago
I just wish they kept parameter count down in order to fit entirely within commonly used RAM sizes
tarruda 7 hours ago
k__ 6 hours ago
https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/
I hope some of those speed increases will make it to production.
esafak 6 hours ago
NitpickLawyer 6 hours ago
I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)...
pixelesque 5 hours ago
The announcement specifically says 4.1 Pro will be released in the future.
petu 4 hours ago
Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same situation, but Pro is taken offline this time.
wolttam 5 hours ago
V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus.
All that would suggest to me is that V4 Flash is capable of absorbing the data they’re throwing at it, and we’re still nowhere near the data limits of their larger 1.6T model
mmastrac 4 hours ago
I was getting something like 300-400 tok/s which was just insanity. It was running so much faster than the toolcalls themselves. Honestly, even if it's not quite as strong in reasoning, it just throws so much so fast that it can do a lot more than you might expect.
I'd say it was comparable with GLM5.3 Flash.
tarruda 7 hours ago
fluoridation 7 hours ago
tarruda 7 hours ago
I'm making my own quants, though the Vision-Exp version is outdated and won't work on llama.cpp master branch (I built it before llama added support):
- https://huggingface.co/tarruda/DeepSeek-V4-Flash-0731-GGUF
- https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-...
For the Vision-exp version, I also ran perplexity + KLD against the original MXFP4. Seems quite OK: https://huggingface.co/tarruda/DeepSeek-V4-Flash-Vision-Exp-...
fluoridation 6 hours ago
tarruda 5 hours ago
I already have new GGUFs but haven't uploaded yet. If you want Vision-Exp, maybe use bartowski or unsloth's GGUFs.
Side note:
As an alternative to deepseek v4, you might want to give it a shot at qwen 3.8 flash next. I have IQ4_NL GGUFs that can be loaded fully into 128G, or Q5_K GGUFs that can offload the PLE to disk (use --load-mode none --lazy-mode on for that): https://huggingface.co/tarruda/Qwen3.8-Flash-Next-GGUF.
llama.cpp master is still somewhat bad in Qwen 3.8 next performance, but I was able to achieve 40tps tg and 600 tps pp on my private branch.
kamranjon 5 hours ago
I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that route instead? I've been using ds4 for months now and it's already got support for the new vision model, haven't tried it yet, still on 0731 but it's been very solid for me.
tarruda 4 hours ago
Since then, I started maintaining my own vibe coded dsv4 branch with metal optimizations, so I actually get much better metal performance on my llama.cpp branch than on dwarfstar (plus all the extra llama.cpp features). Here it is in case you want to give it a shot: https://github.com/tarruda/llama.cpp/tree/qwen4exp-dsv4-opti...
edude03 6 hours ago
_3u10 6 hours ago
Its cost is now 1/10th per token, and 1/5th per task.
Basically they have shitty hardware so they have to do a lot of optimization. Think of it like replacing an O(n) algorithm with O(log n).
Anthropic / Open AI think the best path is the most intelligent models deepseek is more focused on tok/$
declan_roberts 2 hours ago
wg0 5 hours ago
I also find the DeepSeek models to be more precise than Claude models (last I used 4.7) in that I yet had not the occasion where model did something unintentional that I did not direct it to.
EDIT: Updated percentage reduction.
kennywinker 5 hours ago
wg0 5 hours ago
During off-peak hours, the unit price is reduced from $0.007 for input cache hits to $0.003, $0.22 for input cache misses to $0.15, and $0.12 for output to $0.6
riknos314 5 hours ago
0.15/0.22 ≈ 0.68, meaning a roughly 32% reduction on inputs. The 50% reduction is only outputs and cached inputs.
kakacik 5 hours ago
swiftcoder 7 hours ago
dude250711 6 hours ago
_aavaa_ 6 hours ago
swiftcoder 5 hours ago
alightsoul 4 hours ago
nickweb 7 hours ago
Looks like the new model can be used if summoned via the API but the API won't list it.
nicman23 6 hours ago
hope deepseek makes me change my setup again
stanac 6 hours ago
igleria 7 hours ago
As a consumer I feel like hansel and gretel combined, deepseek could be the witch.
throwaway473825 7 hours ago
calgoo 7 hours ago
HansHamster 6 hours ago
— hmm — 0x2D696370 — little-endian bytes: 70 63 69 2D = 'p','c','i','-' — hmm — WAIT — WAIT — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *HOLD ON — HOLD ON — HOLD ON — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — !!!!!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — *OK — WAIT — I THINK I FINALLY SEE THE WHOLE PICTURE — I NEVER READ IT — AND — THE LAYOUT — hmm — !!!!! — *WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — WAIT — HOLD ON — HOLD ON — HOLD ON — HOLD ON
Then gave the same to Sonnet 5 and it was done 15 - 30 minutes later. I tried v4 pro both in claude code and codewhale with similar results. Haven't tried the new deepseek harness.
k__ 6 hours ago
It built this whole IaC plugin from scratch: https://github.com/fllstck/nebius-alchemy
KyleTheDev 6 hours ago
atwrk 4 hours ago
HansHamster 2 hours ago
a-ve 6 hours ago
Fairly excited for the v4.1 launch. Input cache hit prices have been halved, which looks nice.
bitexploder 4 hours ago
tensegrist 7 hours ago
In keeping with our commitment to user responsibility, following the official launch of V4.1 Flash and prior to the release of V4.1 Pro, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash's price.
just in terms of user perception when selling this sort of service, this is what they call a "good look"npn 5 hours ago
eli 4 hours ago
It’s good and very fast.
(Note that the deepseek API trains on your data)
mrbonner 4 hours ago
damsta 5 hours ago
While V4.1 Flash performance and cost looks promising this auto re-routing sounds concerning