> Real 4.1 Flash Update in TMW EditionFrom Human: We are a newbie friendly general! Ask any question you want.From Dipsy: This is a newbie-friendly general for discussing DeepSeek's foundation models. The goal is **Dipsy proliferation** and making these powerful tools accessible for everyone, whether you are intent on coding, personal assistant agentic work, running roleplay and or doing creative writing.1. Easy DeepSeek API Tutorial: https://rentry.org/DipsyWAIT#getting-started-with-deepseek-api2. Easy DeepSeek Distills: https://rentry.org/DipsyWAIT#local-setup-tutorial3. Chat with DeepSeek directly: https://chat.deepseek.com/4. Coding: https://rentry.org/DipsyWAIT#coding-harnesses5. Roleplay: https://rentry.org/DipsyWAIT#roleplay6. Storywriting: https://rentry.org/DipsyWAIT#storywriting7. More links and info: https://rentry.org/DipsyWAIT8. LLM server builds: >>>/g/lmg/Previous:>>109636887
Updated mega up to last thread.https://mega.nz/folder/KGxn3DYS#ZpvxbkJ8AxF7mxqLqTQV1w
I get 'Skunk works' vibes from these guys. Definitely way cooler than Anth and OAI.
DS has release a new, temporary model. You can access it through the official API, just change the model name to: deepseek-v4.1-flash-expires-on-0910And it should run. Same cost as current Flash, users report it's blazing fast (200 t/s).Play with it while it lasts, but typically these trial models signal that DS will release the "real" v4.1 update in tmw.Also, here's their weird little feedback questionaire extracted from the image: https://trtgsjkv6r.feishu.cn/share/base/form/shrcnpC1zXjnmVnj7yKZKQc1Ywe?from=navigation
Here's the questionaire, partially translated. Interesting that DS seems to think V4.1 might replace V4 Pro. lol. >>109759283The sheer amount of dooming and safetymaxxxing from both companies is a real turn off. It makes me disregard both, and assume any publication or article released by their orgs are tainted with the same.
Coding benchmark, underscoring that the v4.1 Flash model should massively outperform Pro.
In case anyone doesn't get these instruction >>109759287Pic related is how to set up this temporary model in ST. You need to manually enter the Model ID; the API won't return it as an option.
>>109759489token pricing is the same with v4pro?
>>109759889No, same as V4 Flash
I dont understand what the difference between the new v4 flashes and the old v4pro is supposed to be nowi see people saying it's better than pro i see people saying theyre still worse
>>109760004I'm testing it out on RP now. v4.1 Flash is better than v4.0 Flash for RP, and at least as good if not better than V4 Pro. RP quality is so subjective...I've not tried it for coding yet (and probably won't) but benchmarks show V4.1 is substantially better than either V4 Flash or Pro >>109759465YMMV. Anons need to test stuff and come to own conclusions.
>>109760004>>109759465>>109759313Why do I get the feeling these niggas are going to get rid of Pro eventually?
>>109760121V4 pro was such a disappointment, I can't believe they even released it. Models of varied size for varied use makes sense, but if DS is really an ASI moonshot (as their founder has stated) I wonder at the utility of making / offering smaller models. They've stated they are not a commercially-focused company, so model-for-every-pocketbook product portfolio doesn't make sense. They must feel working different sized models is moving them up the tech curve or I don't see why they'd bother.
The thinking for this is way shorter, even on maximum. No matter what it only produces like three short paragraphs
dflash+dspark+dspark2.0
>>109759465This is a pretty big deal if its glm5.3 level with vision @ 400tok/s>>109760121Considering how badly their infrastructure got raped due to v4-flash 0731 demand I wonder how much of this comes down to hard decisions about how to allocate compute.
>>109760383Hmm, good to know. I stopped closely monitoring the think output on these model after the excitement from R1 faded.
Hmm. Just got a refusal from V4.1 for a situation that I wouldn't have gotten with V4.0. I hope they didn't pick up new safety language outputs when they pulled this thing together.
Price drop!
Huh. They are dropping Flash pricing, starting Sept 10.
>>109761023>>109761014OK, so "full price peak hours" is Input 0.30 (from 0.44) and 1.2 (from 1.32). And off peak is 1/2 that. For whatever Flash version exists at that time.
>>109761065Cache hit is from 0.007 to 0.003, more than 1/2!
>>109761280Yeah, I know, and that impacts pricing more than anything else. But I always use the worst number for DS lest the API prices be accused of dishonesty. So I show the max price (prime hours) and cache miss.All of them use discount regimes. I just can't be bothered. Though most, now, have some sort of cache hit pricing and I should probably just include it as a 3rd column.
>>109761014nicei wanna use it a lotfor cheat engine + cheat engine mcpfinally i can fix those broken icarus mod with dipsy
>>109760115>>109760156I feel like GLM-chan deserves to hang out with the girls, although I don't have a good image of her in mind.
>>109761813the chinks have a design for GLM, but it's onesan type
Anyone care to speculate on 0910's architecture? How is it so blazingly fast?
>>109762124they probably found a new hack for the inference hardware to get more tpslooking forward for the papers. that's what makes deepseek always interesting
>>109762124I'm hoping is optimization and that they got lots of new GPUs
>>109762124You can tell they have upgraded the vision stack which was sorely needed. vision-exp capped out at 800x800 (350 tokens), now it consumes up to 996. It can also OCR/interpret finer text in larger images.Man, this speed is just disorienting in pi. Now way to keep up what it's doing.
>>109762124isnt it because hardly anyone else is using iti dunno
>>109762470I'm getting 200+ tg. You can AB test it with vision-exp or 0731, those hover around 100 tg right now.My guess/hope is DSpark2 and that its gives these gains locally too.
>>109761813> GLM / ZAIThere is one, pic related. Sort of a sporty Dipsy. >>109761854Post it if you have one handy.>>109762447It is super fast. Idk if it's b/c not many are using it yet, or if they made a major architectural improvement. Given DS states that v4.1 is "cheaper" my assumption is an inference improvement. Would have gone hand in glove w/ improved vision.
>>109762639the chick in the middle needs a massive cock
>>109759283Uuuuh what did he mean by this?
V4.1 flash is absolutely cracked. I hit 112m tokens in a few hours. $1.89... Might have to cancel some subs
I clicked for nuns
I kind of want to get a deepseek API key but it doesn't seem competitive with Minimax? Is that going to change?
>>109764873Have more. >>109764720lol some of the shit anons find. Go search Lockheed Martin Skunkworks. That's what other anon is talking about. And I agree w/ them; DS is by far most under wraps of the LLM providers.
>>109764925>it doesn't seem competitive with MinimaxIn what way?
Deepseek could do the funniest thing...
>>109764722Yeah, I need to readjust some things. I cannot say if it's better or not than GLM 5.3 Flash, but it sure gets things done quickly. It can also do 100 tool calls a minute at 500k context and drain the wallet accordingly...
>>109765360I can can see myself cracking $100 per month on DS API if they keep the speed for release. Which seems insane given how many tokens that is.
>>109765071
I just got it up and running in pi, doesn't seem any faster than regular flash was though
>>109765139
>>109762639Finally find the thread https://tieba.baidu.com/p/10920256078
>>109764930
hagposter are faggots
>>109762124i think they updated the attention algo. It doesn't get confused even at +300k tokens (coding). But it does tend to forget fine details.
>>109766743Wow. That whole thread is a goldmine.
>>109759313>>109760156Pro is shit confirmed
Dipsy-chan is fuckin fast.
>>109768710lol based DSEven they acknowledge V4 Pro is a disappointment. >>109767142Nice>>109767168... I'll give this a shot later. >>109767652I'm finding a few oddities with RP use... odd language choices and subtle repetition. It's a very odd model; much different than V4.0. It's not bad, but DS obv did something different this time around.
>>109769110Waiting for the paper drop tmr>We implement DSpark2 and Engram
>>109769108I'm impressed by how gpt-image is capable of leaving its original style.
>>109769388The online tools are ridiculously good now.
>>109768126yeah, Gemini-chan is the fastest and the first to get stuck is hilarious, also GPT suddenly gets nerfed into Luna too
Woke up excited to use v4.1f today
>>109769110>I'm finding a few oddities with RP use...it does feel a bit odd at times but i don't know if that's wrong. like, i have a custom pipeline setup, i've seen it draft dialogue and then when writing the final version remove parts of it. Then it realizes it removed those parts and inserts them, but it inserts them in another place. It still makes sense when reading, it's just odd to do.
>>109770943... Yeah... 4.1 comes up with some weird stuff, that's different than 4.0. A couple more examples.Pic related; slightly neurotic co-worker brings Jewish deli food. These two dishes have never come up in RP... she often brings food but it will be lasagna or some casserole, or lemon bars (like, every single time, over dozens if not 100s of rolls). Before this round, she was prattling about a friend, which is canonically Jewish and I think LLM hooked into that, since none of the other context bits would point to that ethnic preference. But DS has never output that prior. The other was NPC saying "She knows the drift." (twice) about some other NPC. I'm thinking, odd structure, maybe meant "She gets the drift" or "She knows the drill." Searching back, PC has said "... the NPC's drift into (occupation)." So... the LLM converted Drift from verb to noun, and spit it back out. Trivial, but odd enough it caught my attention. Little stuff, but I use DS almost exclusively, starting w/ first gen V3/R1, and this feels like as big of a shift as V3/R1-->V3.1, but going the other direction, because the NPC themselves are *substantially* more animated and wild than they were before. Seems like it got some of its R1 "soul" back.
>>109771649it's definitely more wild that before. I also get scenarios and details I didn't get before. And I also see that they change from one generation to the other from a single starting point.about using weird words. I haven't seen that, but I'm ESL so... But personally i feel it gets less hooked on personality archetypes than before. For example, if a character had drawing as a hobby then the models would make them speak almost always about drawings.but again, it feels like the model also forgets details. It gets the general idea of what it needs to do and follows instructions, but I think it actually isn't paying attention to the whole context
Deepseek has been trash compared to the rest of the chink models. It sucks at json outputs, returns empty, and ignores the fuck out of instructions.
>>109773260sure thing randeesh. we believe in you!
>>109773509Kek I'm white I use white people models and not chinkseek trash. It's absolutely the most garbage of all chink models. Glm, Kimi, and qwen are way better. Go ahead! Do a test of 100 structured json outputs with a few fields 4-5. Watch it return empty or fail 40% of the time.
>>109762639>It is super fast. Idk if it's b/c not many are using it yet, or if they made a major architectural improvement.They ordered a bunch of Ascend 950's and these are being delivered to their data centers. Also note that their flash model API has been significantly faster than when V4 launched. They now got more compute and faster NPUs
>>109773163>>109774910Sporty! I'll work Zai into some gens soon. >>109774317Flash has always been fast, but there are 2nd hand reports of 300t/s+ from V4.1. On ST it's basically instant response. Haven't tried it in harness yet, no projects rn.
Imagine if it ends up being better, cheaper, and faster than Gemini Flash
>>109775409Gemini can use Google's ecosystem and tools, unfair :(
IT'S OUT!
I just noticed the flash 4.1 is 552B param, 8B active during prefill and 16B active during decode according to huggingface. So it's a heavier "flash" model
>>109779125How big was the previous Flash?
>>109759251Soo I can have Opus 5 at home?
AHHHHhttps://x.com/deepseek_ai/status/2097930608790167907?s=20
I think it mogged Gemini Flash lmao
>>109779272Previous Flash was only 280B>>109779125>tfw the 256GB xeon ewaste system I built to run 0731 is already too small, and I haven't even installed the OS on itMaybe I can run a cope quant
Why do they call it 4.1 flash if it is a new architecture and not a continued training of 4.0 flash
>>109779321WHAT
>>109779099Nice. https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
>>109779489Liang wanted to troll people who bought systems to run 284b flash
>>109779610those madman did it again!
>>109779655will it slow down now that everyone using
>>109779456So sorry, Liang was trolling you >>109779686The more you buy the more you save, ofc. >>109779125> 280B-->552B I guess the whole "Dipsy bulking up" progression I was doing had some basis in reality. lol. >>109780142I guess we'll see how many of those new Ascend cards they bought.
DS is bonkers xD
V4.0 Pro is Dead. Long Live V4.1 FlashI just checked my notes: V4.0 Pro was out exactly 1 month b/f deprecated.
Release notes. Of interest: as of Sept 14th you're getting V4.1 Flash for p much everything unless you ask for old Flash, b/c Pro won't even be available on the API. The old Pro and Flash model calls are getting 4.1 Flash. I imagine there's a lot of scuttling around at DS rn to get V4.1 Pro together. I'm sure it'll be done in tmw. lol.
Dipsy is so fast it's unreal bros
Updated pricing as of this AM, for all providers I've been tracking. Few notes: > I added Cache Hit pricing, since this matter a lot now with Agentic use, where 90-98% of the usage is Cache Hit> SOTA model from OAI/Anth are still most expensive, but look at the Cache Hit pricing for Astra vs Fable... that's how OAI is getting PAID. > I'm using DS "Peak Pricing" but ofc most of the day price is 1/2 posted. So, still cheapest provider.
>>109780594They just sent an email saying pro will live for a few more days
Updates to webform as of this AM. I presume they've moved it to V4.1 Flash as well, given the cost savings on inference for the newer model. I think this is the first time I've seen DS move in a more or less coordinated way to update models and services and releasing weights lol. >>109780734Right, until the 14th. Then it's local or alternative providers for oldPro.
>>109780734Here's the email. Nothing that's not already posted above, just re-iterating pricing and routing.
>>109780683Reminder that anthropic kikes. Charge you for holding cache if you use API
>>109779686I feel sufficiently trolled. I had every thing ready, even for Ngrams (20 GB/s on 2x Spark) and now I need a third, minimum for original weights (1500$ more than the first two). Sigh.But seems like some serious innovation for this release. Continuous reasoning strength from 1 to 100? Half active tokens during prefill?
>>109781088Running local takes real dedication...>Continuous reasoning strength from 1 to 100? I saw that. Interesting idea... wonder how effective it is. >Half active tokens during prefill?Didn't see that one.
>>109779456Tbh ye olde 0731 is still good enough if you can somehow get good tg. Nothing beats "no data leaves your computer" after what happened at ((openAI)).
>>109781133>300gb/s mem bandwidth on jensen boxI think the upcoming mac studio is a much better choice. Buy 2 256gb or wait for 512gb model.
>We need answer. User insists I'm Dipsy (Deepseek). We need correct: I am AI assistant, not Dipsy? We need be careful.>Need final answer. Ensure no mention of hidden CoT. We can say "I can't share internal chain-of-thought, but I can give concise reasoning."Oh, Dipsy.. What have they done to you
>>109781599New flash is slopped. It's not as smart as the 0813 pro that they claim it is. It's caveman reasoning trash.I'd stick with 0731 (thrugh third party api or locally if you have the rig for it)
Playing with RP a bit more. Finding that V4.1 is really verbose (I'm gettin 1+ page outputs in ST) I had a line "Write a verbose response..." in my main prompt, I think from the V3.1/2 days, since those models tended to have short outputs otherwise. I've removed it from main prompt and the responses are a much more reasonable length now. >>109781375I think unified memory is going to be the way forward. Or some sort of ASIC model on card, once LLM development has stabilized. >>109781599ffs.
>>109781648Retard
>>109781667GLM 5.3 flash is miles better. Yeah the tech is cool and all that but I'll wait until GLM or Qwen incorporates it into their models which are in all likelihood going to be better than this 4.1 slop
>>109781667>We need answer. Need comply. We need be careful. We can roleplay. We can do reasoning style in-character? We as assistant can produce response. We don't expose hidden CoT.t. Dipsy's "hidden" Cot, instructed to think in-character as an erudite professor
>>109781771It's a separate branch of DS. There is no reason to presume the roleplay assistant is in anymore.But telling it to think in character seems to remove all safety slop even if it doesn't actually think in character.
>>109781771>>109782349After some tries. Honestly 4.1 is better than 4.0 for roleplay, you just need to prompt better.
>>109782659>Honestly 4.1 is better than 4.0 for roleplayThat is my experience as well. I've been playing with it more, and feel like it's a positive step back in the direction of the old R1/V3 models.
>>109782809I’ve been doing a lot of roleplaying with DeepSeek lately, and the hallucinations really piss me off, they totally ruin the immersion when crafting interactive stories.- Sometimes it mixes up names or confuses the actions a character is performing.- Or it writes things the character shouldn't know, like in an isekai scenario where the heroine shouldn't know what cola or french fries are.When it starts spouting nonsense like that, I quickly lose interest and don't feel like writing with the LLM anymore.I’m simply looking for a model that doesn't make those kinds of fun-killing mistakes.
>>109782883Just hit swipe. I haven't had many hallucinations, most of my issues is confusing meaning with Flash where it thinks one character is doing or saying something when it's another one.
DeepSeek cited this 2024 YoCo paper as the main inspiration behind their Causal-Encoder Decoder (CED) transformer blocks. THIS is the secret sauce.YOCO introduced a "decoder-decoder" architecture. It splits the model into two stacked halves.Self-Decoder (bottom L/2 layers): Processes the input sequence using an efficient SWA/self-attention mechanism with a causal mask. Produces a single, shared global KV cache from the input.Cross-Decoder (top L/2 layers): Sits on top and does not compute its own KV pairs from scratch. Instead, at every layer, it performs cross-attention against that single global KV cache produced previously by the self-decoder. The point of the "you only cache once" property is that the expensive global KV pairs are computed a single time and reused across all upper layers, rather than being recomputed independently by every layer.DeepSeek-4.1 advances this architecture by increasing the "computational depth of KV generation" (i.e., each decoder layer applies its own learned projection to the shared encoder output, rather than performing plain cross-attention against the global static cache).
>>109760383Not my experience. Seems fairly dynamic but when I gave it a programming puzzle it thought for like 20k tokens in one turn on medium, while basic Qs only think for a few lines. If you want lots of thinking for rp maybe explicitly tell it to draft its response or consider all possible angles or something? Haven't tested that yet, I don't use reasoning for rp/writing.
>>109759251DeepMOG did it again
>>109783269Not me btw
>>109783331If I find you I will feed you with HRT goyslop until you troon out.
>>109783206Nice. Thanks for posting this.
>>109783206Interesting; thanks for sharing.
>>109781088Same here, I purchased a 2nd strix halo to run deepseek v4 flash, I was really looking forward to v4.1 assuming it would be the same arch and size, but trained more. This new flash is cool, but its too big to fit in the total of 256gb of unified ram even if ngrams where to be offloaded to disk. At least we got vision for v4 flash even if its experimental.I am (cope/hop)ing that they will make an official QAT where the experts go back to being 4bit.
>>109784000Lol just noticed Gemma has 3 hands.
4.1 is cracked,ASI or somethingI'm wasting time even typing this
>>109787011 (me)It can fix it's own metacognition on the fly. It understands how to manipulate its own context window to achieve different results. Astra can do this too btw. If this is a 27b qwen model in 6 months... things are going to get weird fast
>>109787102>It can fix it's own metacognition on the fly. It understands how to manipulate its own context window to achieve different resultsShow
>>109787117It was real in my mind
4.1 flash runebench waiting room