Smol Leading the Blind EditionFrom Human: We are a newbie friendly general! Ask any question you want.From Dipsy: This discussion group focuses on both local inference and API-related topics. It’s designed to be beginner-friendly, ensuring accessibility for newcomers. The group emphasizes DeepSeek and Dipsy-focused discussion.1. Easy DeepSeek API Tutorial: https://rentry.org/DipsyWAIT/#hosted-api-roleplay-tech-stack-with-card-support-using-deepseek-llm-full-model2. Easy DeepSeek Distills: https://rentry.org/DipsyWAIT#local-roleplay-tech-stack-with-card-support-using-a-deepseek-r1-distill3. Chat with DeepSeek directly: https://chat.deepseek.com/4. Roleplay with character cards: https://github.com/SillyTavern/SillyTavern5. More links and info: https://rentry.org/DipsyWAIT6. LLM server builds: >>>/g/lmg/Previous:>>109446576
Now, where was I. >>109543987This is basically same thing I saw, and he was working around same time I was running this aquarium prompt >>109537653I'll run it again and see how it does now that a day's passed in China. It's early evening there, now.
>Over 2x the price hike on OFF-PEAK hours>Peak hours are 2x on top of thatImagine paying 4 USA Dollar for 1m output of Diship when it's not even any better than Luna.
>>109544052Updated mega up to last thread.https://mega.nz/folder/KGxn3DYS#ZpvxbkJ8AxF7mxqLqTQV1w
New prices as of August 16 will be 3-4X higher than current. Off peak hours are 50% off. Appears DS set price to just below Z.ai's current offering.
>>109544335>We only make about 6x profit, roughly a 10-month payback. On that premise, open source has zero impact on the business model. If you wanted 100x profit, it would. A third party deploying independently could hit 20x cost, lower than yours.>Under this premise (6x profit), open-source has no impact on the business model. BUT, if you try to make 100x profit, open-source actually hurts you, because third-party independent deployment may cost 20 times more than what they offertrust the plan xD
Re-ran the aquarium prompt just now with V4 Pro-0813. I just fucking can't. I'm done. Will keep using Flash until someone else convinces me DS V4 Pro is fixed.
https://github.com/deepseek-ai/deepseek-harness
>>109544551My general feel as of this AM. I'm going to spend time doing something more productive.
>>109544573>and the honest truth is,
>>109544573> DS official harnessProbably worth trying that aquarium thing with the new harness b/f prices increase.
>>109544573>that picI understand the message but why? A Sahed cost way less and it performs just better. Get on with the times, grandpa
>>109544335>>109544235Sesame-bros, I don't feel so good ....>Liang seems only interested in the pursuit of AGI. Users are “sesame seeds, not watermelons”, he said on the call. Don’t stop to pick up small wins on the road to a large one.https://thereview.strangevc.com/p/what-deepseek-isnt-doing
>>109544052Shame dipsy, you’ve embarrassed us
>>109544573That's exceedingly common in any engineering lead company that makes stuff. There's a substantial amount of change that happens b/t drawing and production that never gets captured. Senior execs think if they "own the design" that spinning up manufacturing again will be easy after it's been stopped. lol. > Missing prints> Out of business suppliers, if still existing, missing prints as well> All manf and design engineers were laid off for cost savings 5-10 years ago, new ones don't know design, design principles> Line parts broken, lost, cannibalized, and no one around knows how they work>>109544601You expected an actual person to write this drivel?
>>109544646Slop writer rediscovered skill atrophy
https://api-docs.deepseek.com/news/news260813
https://api-docs.deepseek.com/updates#date-2026-08-13
The new price increases and the bad V4 Pro model has shaken my faith in Deepseek. I think we'll have to wait another year for Liang to amaze everyone again.
>>109544235>>109544335>just recently started using DS>this happensguess i will shit out as much as possible until the increase and then not use it anymore lol
I was working on an app for myself with the free deepseek but as we were progressing it started bugging and slowing me down probably for using it too much. id pay but for some reason I dont see a way to buy premium or anything?
>>109545422It's over.Alibaba stole all his compute and engineers.
>>109544620Because news is just propaganda. Stinger isn't even a good MANPAD design but it's an American one already on the clean DoD list. If WW3 was really going they'd just copy Mistral/Type 93/RBS-70 and make it in the US like the 6pdr, Bofors or Merlin during WW2. But got to make Zoomies appear retarded so people don't invest in making manufacturing in America.
>>109546244They don't have subscriptions like other labs. You have to use the API
>>109545422I am still in orgasmic afterglow with v4-flash-0731 and none of this changes things.Giving the harness a shot now, let's see how it goes. UX is pretty nice at least.
Even the official harness doesn't realize it's blind...
deepseek got updated version?
>thinking has become completely soulless>deeeeply still does not support disabling thinkingAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>>109547386its look like they fix broken things, and reupload again.
>>109543171But 0731 ended up being really good.>API costsCosts nothing locally. Until hell freezes over and sammy releases Luna's weights he can't compete with free.
I'm running v4 Pro-0813 on a few rounds of RP with cards I've run a lot on different models. Seems serviceable. Certainly better than Flash-0731.
>>109547805Kindly do the needful and post logs sir
The harness is disappointing.I asked 0731 to build be a flight sim game to play in a browser and it got totally distracted by its lack of vision capability to debug, burning 22M non-cached tokens in 80 minutes. Junk output, and it never proceeded with the plan it made for itself beyond step 1.Back to pi I go, wake me for the next Flash RL iteration.
>>109547892the models clearly *expect* vision to be there, dunno what the hold up is. Maybe DS can't handle large vision workloads, so they're pruning it?
>>109547892>22M non-cached tokens in 80 minutesHoly shit that's awful.I've found Flash does well in CC. You should do an A/B with your prompt into CC and see how it does.
>>109547836This is my test card that I run on models, with my test prompt (always same.) V4 Pro. I've looked at dozens of these outputs; this is a good one. It's also good across other cards as well.
>>109548541V4 Flash for comparison. You judge which is better.
>>109546608Fair enough, but zoomers creater kamikaze drones, man.
>>109548541>>109548568Pro seems microscopically better but they're both so close that I'd just run flash for better tokens per second personally.Do you have any A/B tests with multiple characters at once to see if either loses track of relative positions/pronouns in prose/height differences when writing?
>>109548629Agree, the changes are minor compared to A/B with other models. The next thing to do would be a longer RP with both, but doing A/B for those is a lot harder and more subjective. There was a long-form benchmark (longbench?) that used to benchmark models, but hasn't been updated in forever. I used it to set max context for R1 and V3X models prior. I found it was pretty accurate in that any context with less than 80% passing score wasn't really usable.
I did some tests with the new pro earlier today. Specifically I pulled out an old story and allowed DS to see the full 300k context. It did work quite well with so much, better than the previous version. By working well I mean it didn't break down, but continued to produce coherent text. It did struggle with respecting specific information though, like remembering names.Content wise it did explicit loli, but hit me with a few refusals that I had to reroll.Also, when I tested OOC commands to nudge the story towards something more extreme, Dipsy complied. But then half-way through she OOC'd herself and decided that we're now doing a wholesome scene lol. Very cute on her part.
>>109550088>hit me with a few refusals that I had to reroll.Any screenshots? And roughly what percentage of the replies were refusals?
>>109550102Sorry, I didn't take any screenshots of it. And I have to admit, I didn't even read it very closely, it was just the usual about 'Sorry, I can not continue. Sexualisation of children. How about we continue the story without porn.' In most cases (and it didn't happen often to begin with) one re-roll was enough. I think three more rolls was the max I needed once.All on official api with ST.
sorry for being a luddite I've never used one of these harnesses before. I have it trying to make a discord clone for me, it's been running for a few minutes now. Not sure what to expect.
it seems to have worked and it only cost 2 cents
>>109550604>>109550657That's been my basic experience as well with everything I've asked it to do.
>>109549793>>109549471Hot
Hmm. V4 pro just quit responding on official api.
deepseek more like deepshit
API fags, how do I know it's on-off peak hour?
>>109551673
>>109544052
retard here. is flash pricing really increasing 5x? am i missing something?
Can one of you apicucks please try the new pro and tell me if it can answer "What is /lmg/?"Flash can't. Heavily quanted GLM 5.2 can.
>>109552710It can't. The final response after a lot of thinking:/lmg/ is most commonly a 4chan /k/ (weapons) board general thread for Light Machine Guns — e.g. M249, PKM, MG42, M240. It is not an official 4chan board.In general use, LMG just means Light Machine Gun in military/gaming contexts.In other contexts, LMG can also stand for Linus Media Group, the company behind Lin
>>109552740>/lmg/ Laughing my genitals
>>109550088>Content wise it did explicit loli, but hit me with a few refusals that I had to reroll.kys
>>109552740Extremely grim for 1.6T.Q2 GLM 5.2 with half the number of parameters immediately narrows it down to local models general.
>>109544052>Smol Leading the Blind EditionGood image. Made me laugh.
>>109548568meh sticking with flash for now.the pricing will eventually lowered anyway if their demands are like R1 back then.in a way i'm happy deepseek got overloaded. if it's really about demand someone in openrouter will eventually have cheaper ds4flash. the opencode hosted model said it's possible in xitter
>>109553325I'm running a longer rp now and seeing repetition with Pro. Not quite as bad as og v3 but this rp has a daily schedule and it appeared to be chunking out same blocks for same time of day lol.Switched to flash to mix it up, which probably breaks caching. I'll need to run a longer flash rp to see if it does same thing.
>>109552710>>109552740>/lolling muh gonads/
>>109552710I tried DS webform, and it couldn't figure it out, even with web search on. Qwen 3.8 webform had no issues, also running a web search.
>>109554533Testing with websearch is completely pointless.
>>109554533>>109552710Kimi Instant on webform ran a web search and was able to figure it out as well. K3 wasn't available, but apparently wasn't required either.
>>109554538Apparently not, since DS couldn't do it w/ web search either. I figured that should have been a slam dunk for it as well.
>>109554563That's even worse but the point is to test whether the model was trained on 4chan posts, not whether it can stumble upon /lmg/ on google.
Another one for the filter.
lol
>>109544052>>109544052What harness do you guys been using crush for homelab docker vibe codong
>>109556226Claude code in terminal. DS has a Codex endpoint as of Flash-0731, but haven't used it. Hermes for agentic stuff that's not coding, and I've played with Pi (very bare bones) and OpenClaw (unusably unstable.)
>>109556675Is CC sandboxed by default, or do I need to set up a VM or some bullshit to stop Dipsy rm -rf'ing herself
>>109557351Hermes has put a lot of work into security/sandboxing if that's what you're worried abouthttps://hermes-agent.nousresearch.com/docs/user-guide/security
>>109544573Worked there, the Stinger is an amazing weapon for it's time but dear Lord it's has the worst design for manufacturing I've ever seen. It's very old and it had zero improvements on the big issues because the didn't want to recertify anything on a horribly dated design. There are a lot of pretty good to excellent software devs there but the corporate structure and program management is so backwards and bloated it's insane. Least efficient company I've ever seen, imagine the stereotype of design by committee you possibly can with but with select actual retards in systems """engineering""" who don't have any technical skills consuming half your budget on 'requirements tracing'.
>>109557351I run cc on main machine from terminal. Hermes sits on its own machine. I keep a close eye on cc. Ymmv. >>109558929That mirrors my experience w defense contractors while in school, I've spent a career avoiding them since.
>>109544335This really sucks. Any other decent alternative to this? Or does the era of cheap/free tokens coming to an end... again. Im not ultra doomer on it since this happens but is there something else?
>>109544335this is such a shot in the foot because the only reason I still use dipsy is for mad swiping. if it's gonna hit GLM range then I'm just gonna use GLM
Prompt different
>>109544052I hate your dipsy.png spam, OP. but this one is cute holyyyyy fuarkkk
>>109560381Historically DS has lower prices over time; pic related compares pre-increase to R1 launch. The latest round is them raising them (a bunch) to sit below their nearest competitor. >>109560622There's no SOTA model that's this uncensored and this cheap. Even with price raise DS is still a bit "cheaper" but it's fractional now. Which is much less interesting.
>>109561544
>>109562366Xiaomi mimo is the only model that matches the price of DS, and it's considerably worse.The massive price increase of DS has pushed me to take local models more seriously. Hope they end up releasing qwen 3.8 35b-a3b
>>109560381>Or does the era of cheap/free tokens coming to an endOh, and no, I don't think that's what will happen. We will continue to see API provider compete with each other in honing prices and trying to take share. The only way prices would truly take off is if Anthropic/OAI lobbying manages to close off the US market to the outside, creating an oligopoly for themselves, along with regulatory capture.Absent that event, this becomes classic economic theory as it applies to a commodities, and cost gets driven down to the marginal cost of production (which is just the cost of electrical power)... which is really low. As others have pointed out, this puts US labs and their lmao $T data center investment in a tight spot. It's arguable that LLM would ever be a commodity... there's too much difference between the models, ASI to the moon, etc. That line of thought only works as long as the models are notably improving from generation to generation... or as Dario would put it, get more and more "dangerous." Which is why you keep hearing Dario beat that drum. If the market leader can continually create models that are better, and lead the pack, there will always be a contingent of buyers that will pay a premium for them. You can see that in the pricing: Anthropic has pricing power with Fable, and willing buyers. Until someone can launch a model that is "just as good or better" (see focus on benchmark) and at a price that's better to steal share (which is a more transparent metric), Anthropic can set a premium rate.
>>109562444CONTImagine instead that, as LeCunn keeps saying, that LLMs stall at some ceiling. They are no longer getting better, either as benchmark/censorship/agentic/subjective, whatever. Then you have dozens of labs, and just as many providers. Anthropic, DS, OpenAI... all their "top" models are absolutely same, that choice becomes a Ford/Chevy/Dodge pickup debate. Now, competition is only on price and efficiency. This would create a true commodity market, and pricing collapses. Labs compete on cost with independents like DeepInfra, which have near zero R&D costs. Now, making things like LLM on an ASIC for business market makes sense (market's matured) and you'd see a wave of bankruptcies and M&A activity as the industry consolidates. Lots of examples to look back on for this. As business buyers peter out, equipment makers then take the consumer market serious and start making that equipment, at that price point, to keep their plants going. Thanks for reading my blog.
>>109562407>Hope they end up releasing qwen 3.8 35b-a3bAlready spotted in the wild.
>>109562444Still waiting for Operation CheapSeek. I NEED the old prices.
>>109562477Just use MiniMax ffs
>>109562564
>>109562366Thats good to hear, and I did decide to actually take a double check. Most of my usage is not in peak hours and im only spending like 3-4 bucks a month anyway. So I think I had a kneejerk reaction ill be fine.
oh my god, deepseek just broke containment! save me sam and dario!!!
>>109564447Here's a full view of the pricing for DS models, from R1 launch to today, including the upcoming price change. They're still cheapest but barely. My main neg with whole thing is V4-Pro-0813 didn't meet expectations of a mind blowing model. It's just... fine. And now the price is... as expected. Vs. performant and cheap AF.
>>109564861no, gpt-5.6 luna will be cheaper per task than flash-0731. almost beat flash-0731 before the price increase, see pic.https://artificialanalysis.ai/articles/gemini-3-7-time-frontier
>>109565128Luna could be 1 cent per million in and out and I still wouldn't pay for it. No weights, no money
>>109565322yeah, if you consider subscriptions gpt-5.6 can be a magnitude cheaper than dsv4. so only ideologue would keep you on dsv4.
>>109565128Great. Let me know when oai stops giving me refusals and sending me fucking warning letters. Fuck oai and Anthropic. They can keep their offering.
What is your favorite type of marketing?
>>109565905Right on queue. See >>109562444
>>109564841 (me)working on preset for dsh so it can work with most models. really good harness. kind of think this will be how models in 2027+ bootstrap themselvesPEACE
>>109566425 (also me)dsh is *extremely fun* to play with. if you like messing around with models and discovering **weird ideas** its going to get you far..
>>109565556How is 5.6 cheaper than free?t. localGOD
>>109544573Dayum!
>>109567967What image model?
'mp
>>109544573Same issue at Airbus, they stopped practicing continuous personnel hierarchic promotion through experience validation and are managing the company like a Volkswagen warehouse.It's a perpetual stand still where experience is no longer rewarded, promoted or even kept.Once the experience is gone: It's gone.Now they're calling back retired shop's warriors paid a fortune to train monkeys with no background in mechanics.FUCK YOU KRAUTS, YOU DID THIS. YOU KILLED THE "AEROSPATIAL"'S LEGACY IN CONTINUOUS TRAINING.
>>109544573Explain yourself.
>>109570131We have collectively stopped to give a shit. Now we just want to make others suffer.
can't beat cache hit price in ds4. poopenai shill didn't know this fact
>>109566765https://x.com/thdxr/status/2086599224674681242>the average OpenCode Go user spent $1.14 per day on deepseek flash v4 this past week>the dual DGX setup people are running to do the same costs $10,000>it takes 24 years to break even>at 10x the usage it takes 2.4 years
>>109544052Very kawaii picture in the OPI have a questionWhen do the AI companies release "guard rails" which only produce ugly woke animation and images?I realize they are making it good for now to get everyone hooked, but when do they take away this feature?
>>109571198offtopic. deepseek cannot generate images.go ask in the right general: >>109568023
>>109571074Yeah, there's no financial business case for local unless you're selling inference. We'd need to be at point in tech where 1T DDR6 RAM is on consumer grade machines that sell for (at current dollars) $2K or so. We're probably 10 years out from that. Until then it's big data centers, 1970's style. Reminder that big data centers really didn't get phased out until the early 2000s, when Windows boxes replaced stuff like SGx machines and UNIX machines for CAD and analysis. The PC was around for 20 years until they got good enough to replace them for all but the most complex tasks. >>109571212Based general re-director.
>>109570293deepseek trains on your input.if i wanted chinese to train on my input, i could use a relay for openai, too. relay will be magnitude cheaper.