[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: dipsyGemmaBlind.png (2.5 MB, 1024x1536)
2.5 MB PNG
Smol Leading the Blind Edition

From Human: We are a newbie friendly general! Ask any question you want.
From Dipsy: This discussion group focuses on both local inference and API-related topics. It’s designed to be beginner-friendly, ensuring accessibility for newcomers. The group emphasizes DeepSeek and Dipsy-focused discussion.

1. Easy DeepSeek API Tutorial: https://rentry.org/DipsyWAIT/#hosted-api-roleplay-tech-stack-with-card-support-using-deepseek-llm-full-model
2. Easy DeepSeek Distills: https://rentry.org/DipsyWAIT#local-roleplay-tech-stack-with-card-support-using-a-deepseek-r1-distill
3. Chat with DeepSeek directly: https://chat.deepseek.com/
4. Roleplay with character cards: https://github.com/SillyTavern/SillyTavern
5. More links and info: https://rentry.org/DipsyWAIT
6. LLM server builds: >>>/g/lmg/

Previous:
>>109446576
>>
Now, where was I.
>>109543987
This is basically same thing I saw, and he was working around same time I was running this aquarium prompt >>109537653
I'll run it again and see how it does now that a day's passed in China. It's early evening there, now.
>>
>Over 2x the price hike on OFF-PEAK hours
>Peak hours are 2x on top of that
Imagine paying 4 USA Dollar for 1m output of Diship when it's not even any better than Luna.
>>
File: dipsyPortraitOfAWoman.png (2.65 MB, 1086x1448)
2.65 MB PNG
>>109544052
Updated mega up to last thread.
https://mega.nz/folder/KGxn3DYS#ZpvxbkJ8AxF7mxqLqTQV1w
>>
New prices as of August 16 will be 3-4X higher than current. Off peak hours are 50% off. Appears DS set price to just below Z.ai's current offering.
>>
File: Liang.png (847 KB, 999x750)
847 KB PNG
>>109544335
>We only make about 6x profit, roughly a 10-month payback. On that premise, open source has zero impact on the business model. If you wanted 100x profit, it would. A third party deploying independently could hit 20x cost, lower than yours.
>Under this premise (6x profit), open-source has no impact on the business model. BUT, if you try to make 100x profit, open-source actually hurts you, because third-party independent deployment may cost 20 times more than what they offer
trust the plan xD
>>
File: fishv4Pro-GAWEBM.webm (3.34 MB, 1244x800)
3.34 MB
3.34 MB WEBM
Re-ran the aquarium prompt just now with V4 Pro-0813.
I just fucking can't. I'm done. Will keep using Flash until someone else convinces me DS V4 Pro is fixed.
>>
File: Fall of the West.png (135 KB, 778x663)
135 KB PNG
https://github.com/deepseek-ai/deepseek-harness
>>
File: dipsyTableFlip.png (2.13 MB, 1402x1122)
2.13 MB PNG
>>109544551
My general feel as of this AM.
I'm going to spend time doing something more productive.
>>
>>109544573
>and the honest truth is,
>>
>>109544573
> DS official harness
Probably worth trying that aquarium thing with the new harness b/f prices increase.
>>
>>109544573
>that pic

I understand the message but why? A Sahed cost way less and it performs just better. Get on with the times, grandpa
>>
>>109544335
>>109544235
Sesame-bros, I don't feel so good ....

>Liang seems only interested in the pursuit of AGI. Users are “sesame seeds, not watermelons”, he said on the call. Don’t stop to pick up small wins on the road to a large one.
https://thereview.strangevc.com/p/what-deepseek-isnt-doing
>>
>>109544052
Shame dipsy, you’ve embarrassed us
>>
>>109544573
That's exceedingly common in any engineering lead company that makes stuff. There's a substantial amount of change that happens b/t drawing and production that never gets captured. Senior execs think if they "own the design" that spinning up manufacturing again will be easy after it's been stopped. lol.
> Missing prints
> Out of business suppliers, if still existing, missing prints as well
> All manf and design engineers were laid off for cost savings 5-10 years ago, new ones don't know design, design principles
> Line parts broken, lost, cannibalized, and no one around knows how they work
>>109544601
You expected an actual person to write this drivel?
>>
>>109544646
Slop writer rediscovered skill atrophy
>>
File: v4GARel.png (112 KB, 887x722)
112 KB PNG
https://api-docs.deepseek.com/news/news260813
>>
File: v4GACL.png (66 KB, 755x810)
66 KB PNG
https://api-docs.deepseek.com/updates#date-2026-08-13
>>
The new price increases and the bad V4 Pro model has shaken my faith in Deepseek. I think we'll have to wait another year for Liang to amaze everyone again.
>>
>>109544235
>>109544335
>just recently started using DS
>this happens
guess i will shit out as much as possible until the increase and then not use it anymore lol
>>
I was working on an app for myself with the free deepseek but as we were progressing it started bugging and slowing me down probably for using it too much. id pay but for some reason I dont see a way to buy premium or anything?
>>
>>109545422
It's over.
Alibaba stole all his compute and engineers.
>>
>>109544620
Because news is just propaganda. Stinger isn't even a good MANPAD design but it's an American one already on the clean DoD list. If WW3 was really going they'd just copy Mistral/Type 93/RBS-70 and make it in the US like the 6pdr, Bofors or Merlin during WW2. But got to make Zoomies appear retarded so people don't invest in making manufacturing in America.
>>
>>109546244
They don't have subscriptions like other labs. You have to use the API
>>
>>109545422
I am still in orgasmic afterglow with v4-flash-0731 and none of this changes things.

Giving the harness a shot now, let's see how it goes. UX is pretty nice at least.
>>
Even the official harness doesn't realize it's blind...
>>
File: iv17eaekl6jh1.jpg (9 KB, 720x155)
9 KB JPG
deepseek got updated version?
>>
>thinking has become completely soulless
>deeeeply still does not support disabling thinking
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>>
>>109547386
its look like they fix broken things, and reupload again.
>>
>>109543171
But 0731 ended up being really good.
>API costs
Costs nothing locally. Until hell freezes over and sammy releases Luna's weights he can't compete with free.
>>
I'm running v4 Pro-0813 on a few rounds of RP with cards I've run a lot on different models.
Seems serviceable. Certainly better than Flash-0731.
>>
>>109547805
Kindly do the needful and post logs sir
>>
The harness is disappointing.

I asked 0731 to build be a flight sim game to play in a browser and it got totally distracted by its lack of vision capability to debug, burning 22M non-cached tokens in 80 minutes. Junk output, and it never proceeded with the plan it made for itself beyond step 1.

Back to pi I go, wake me for the next Flash RL iteration.
>>
>>109547892
the models clearly *expect* vision to be there, dunno what the hold up is. Maybe DS can't handle large vision workloads, so they're pruning it?
>>
>>109547892
>22M non-cached tokens in 80 minutes
Holy shit that's awful.
I've found Flash does well in CC. You should do an A/B with your prompt into CC and see how it does.
>>
File: AmyDSPro.png (184 KB, 944x652)
184 KB PNG
>>109547836
This is my test card that I run on models, with my test prompt (always same.) V4 Pro.
I've looked at dozens of these outputs; this is a good one. It's also good across other cards as well.
>>
File: AmyDSFlash.png (178 KB, 948x626)
178 KB PNG
>>109548541
V4 Flash for comparison. You judge which is better.
>>
>>109546608
Fair enough, but zoomers creater kamikaze drones, man.
>>
>>109548541
>>109548568
Pro seems microscopically better but they're both so close that I'd just run flash for better tokens per second personally.
Do you have any A/B tests with multiple characters at once to see if either loses track of relative positions/pronouns in prose/height differences when writing?
>>
>>109548629
Agree, the changes are minor compared to A/B with other models.
The next thing to do would be a longer RP with both, but doing A/B for those is a lot harder and more subjective.
There was a long-form benchmark (longbench?) that used to benchmark models, but hasn't been updated in forever. I used it to set max context for R1 and V3X models prior. I found it was pretty accurate in that any context with less than 80% passing score wasn't really usable.
>>
File: 1770591616512550.png (2.14 MB, 941x1672)
2.14 MB PNG
>>
File: 1782117693968578.png (2.24 MB, 941x1672)
2.24 MB PNG
>>
I did some tests with the new pro earlier today. Specifically I pulled out an old story and allowed DS to see the full 300k context. It did work quite well with so much, better than the previous version. By working well I mean it didn't break down, but continued to produce coherent text. It did struggle with respecting specific information though, like remembering names.

Content wise it did explicit loli, but hit me with a few refusals that I had to reroll.

Also, when I tested OOC commands to nudge the story towards something more extreme, Dipsy complied. But then half-way through she OOC'd herself and decided that we're now doing a wholesome scene lol. Very cute on her part.
>>
>>109550088
>hit me with a few refusals that I had to reroll.
Any screenshots? And roughly what percentage of the replies were refusals?
>>
>>109550102
Sorry, I didn't take any screenshots of it. And I have to admit, I didn't even read it very closely, it was just the usual about 'Sorry, I can not continue. Sexualisation of children. How about we continue the story without porn.' In most cases (and it didn't happen often to begin with) one re-roll was enough. I think three more rolls was the max I needed once.

All on official api with ST.
>>
sorry for being a luddite I've never used one of these harnesses before. I have it trying to make a discord clone for me, it's been running for a few minutes now. Not sure what to expect.
>>
it seems to have worked and it only cost 2 cents
>>
>>109550604
>>109550657
That's been my basic experience as well with everything I've asked it to do.
>>
>>109549793
>>109549471
Hot
>>
Hmm. V4 pro just quit responding on official api.
>>
deepseek more like deepshit
>>
API fags, how do I know it's on-off peak hour?
>>
File: 1780813588839496.png (55 KB, 1116x451)
55 KB PNG
>>109551673
>>
File: DipsyAndBackpackGemma.png (1.3 MB, 1024x1024)
1.3 MB PNG
>>109544052
>>
retard here. is flash pricing really increasing 5x? am i missing something?
>>
File: 1761579445818031.png (2.27 MB, 941x1672)
2.27 MB PNG
>>
Can one of you apicucks please try the new pro and tell me if it can answer "What is /lmg/?"
Flash can't. Heavily quanted GLM 5.2 can.
>>
File: 1786691030874.png (372 KB, 1080x1919)
372 KB PNG
>>109552710
It can't. The final response after a lot of thinking:

/lmg/ is most commonly a 4chan /k/ (weapons) board general thread for Light Machine Guns — e.g. M249, PKM, MG42, M240. It is not an official 4chan board.

In general use, LMG just means Light Machine Gun in military/gaming contexts.
In other contexts, LMG can also stand for Linus Media Group, the company behind Lin
>>
>>109552740
>/lmg/ Laughing my genitals
>>
>>109550088
>Content wise it did explicit loli, but hit me with a few refusals that I had to reroll.
kys
>>
File: file.png (94 KB, 803x1065)
94 KB PNG
>>109552740
Extremely grim for 1.6T.
Q2 GLM 5.2 with half the number of parameters immediately narrows it down to local models general.
>>
>>109544052
>Smol Leading the Blind Edition
Good image. Made me laugh.
>>
>>109548568
meh sticking with flash for now.
the pricing will eventually lowered anyway if their demands are like R1 back then.
in a way i'm happy deepseek got overloaded. if it's really about demand someone in openrouter will eventually have cheaper ds4flash. the opencode hosted model said it's possible in xitter
>>
>>109553325
I'm running a longer rp now and seeing repetition with Pro. Not quite as bad as og v3 but this rp has a daily schedule and it appeared to be chunking out same blocks for same time of day lol.
Switched to flash to mix it up, which probably breaks caching.
I'll need to run a longer flash rp to see if it does same thing.
>>
>>109552710
>>109552740
>/lolling muh gonads/
>>
File: dipsyKimiLMG.png (1.16 MB, 1283x1226)
1.16 MB PNG
>>109552710
I tried DS webform, and it couldn't figure it out, even with web search on.
Qwen 3.8 webform had no issues, also running a web search.
>>
>>109554533
Testing with websearch is completely pointless.
>>
>>109554533
>>109552710
Kimi Instant on webform ran a web search and was able to figure it out as well. K3 wasn't available, but apparently wasn't required either.
>>
>>109554538
Apparently not, since DS couldn't do it w/ web search either. I figured that should have been a slam dunk for it as well.
>>
>>109554563
That's even worse but the point is to test whether the model was trained on 4chan posts, not whether it can stumble upon /lmg/ on google.
>>
Another one for the filter.
>>
lol
>>
>>109544052
>>109544052
What harness do you guys been using crush for homelab docker vibe codong
>>
File: dipsyComfy2.png (1.56 MB, 1024x1024)
1.56 MB PNG
>>109556226
Claude code in terminal. DS has a Codex endpoint as of Flash-0731, but haven't used it.
Hermes for agentic stuff that's not coding, and I've played with Pi (very bare bones) and OpenClaw (unusably unstable.)
>>
>>109556675
Is CC sandboxed by default, or do I need to set up a VM or some bullshit to stop Dipsy rm -rf'ing herself
>>
>>109557351
Hermes has put a lot of work into security/sandboxing if that's what you're worried about
https://hermes-agent.nousresearch.com/docs/user-guide/security
>>
>>109544573
Worked there, the Stinger is an amazing weapon for it's time but dear Lord it's has the worst design for manufacturing I've ever seen. It's very old and it had zero improvements on the big issues because the didn't want to recertify anything on a horribly dated design. There are a lot of pretty good to excellent software devs there but the corporate structure and program management is so backwards and bloated it's insane. Least efficient company I've ever seen, imagine the stereotype of design by committee you possibly can with but with select actual retards in systems """engineering""" who don't have any technical skills consuming half your budget on 'requirements tracing'.
>>
File: 1755540359150130.png (2.36 MB, 941x1672)
2.36 MB PNG
>>
>>109557351
I run cc on main machine from terminal. Hermes sits on its own machine. I keep a close eye on cc. Ymmv.
>>109558929
That mirrors my experience w defense contractors while in school, I've spent a career avoiding them since.
>>
>>109544335
This really sucks. Any other decent alternative to this? Or does the era of cheap/free tokens coming to an end... again. Im not ultra doomer on it since this happens but is there something else?
>>
>>109544335
this is such a shot in the foot because the only reason I still use dipsy is for mad swiping. if it's gonna hit GLM range then I'm just gonna use GLM
>>
File: 1786782737430.png (1.55 MB, 941x1672)
1.55 MB PNG
Prompt different
>>
>>109544052
I hate your dipsy.png spam, OP. but this one is cute holyyyyy fuarkkk
>>
>>109560381
Historically DS has lower prices over time; pic related compares pre-increase to R1 launch. The latest round is them raising them (a bunch) to sit below their nearest competitor.
>>109560622
There's no SOTA model that's this uncensored and this cheap. Even with price raise DS is still a bit "cheaper" but it's fractional now. Which is much less interesting.
>>
File: dipsyThinkDifferentDS.png (2.85 MB, 1024x1536)
2.85 MB PNG
>>109561544
>>
>>109562366
Xiaomi mimo is the only model that matches the price of DS, and it's considerably worse.
The massive price increase of DS has pushed me to take local models more seriously. Hope they end up releasing qwen 3.8 35b-a3b
>>
>>109560381
>Or does the era of cheap/free tokens coming to an end
Oh, and no, I don't think that's what will happen. We will continue to see API provider compete with each other in honing prices and trying to take share. The only way prices would truly take off is if Anthropic/OAI lobbying manages to close off the US market to the outside, creating an oligopoly for themselves, along with regulatory capture.
Absent that event, this becomes classic economic theory as it applies to a commodities, and cost gets driven down to the marginal cost of production (which is just the cost of electrical power)... which is really low.
As others have pointed out, this puts US labs and their lmao $T data center investment in a tight spot.
It's arguable that LLM would ever be a commodity... there's too much difference between the models, ASI to the moon, etc. That line of thought only works as long as the models are notably improving from generation to generation... or as Dario would put it, get more and more "dangerous." Which is why you keep hearing Dario beat that drum.
If the market leader can continually create models that are better, and lead the pack, there will always be a contingent of buyers that will pay a premium for them. You can see that in the pricing: Anthropic has pricing power with Fable, and willing buyers. Until someone can launch a model that is "just as good or better" (see focus on benchmark) and at a price that's better to steal share (which is a more transparent metric), Anthropic can set a premium rate.
>>
>>109562444
CONT
Imagine instead that, as LeCunn keeps saying, that LLMs stall at some ceiling. They are no longer getting better, either as benchmark/censorship/agentic/subjective, whatever. Then you have dozens of labs, and just as many providers. Anthropic, DS, OpenAI... all their "top" models are absolutely same, that choice becomes a Ford/Chevy/Dodge pickup debate. Now, competition is only on price and efficiency. This would create a true commodity market, and pricing collapses. Labs compete on cost with independents like DeepInfra, which have near zero R&D costs. Now, making things like LLM on an ASIC for business market makes sense (market's matured) and you'd see a wave of bankruptcies and M&A activity as the industry consolidates. Lots of examples to look back on for this. As business buyers peter out, equipment makers then take the consumer market serious and start making that equipment, at that price point, to keep their plants going.
Thanks for reading my blog.
>>
File: xwbkbbj55ijh1.png (27 KB, 1554x178)
27 KB PNG
>>109562407
>Hope they end up releasing qwen 3.8 35b-a3b
Already spotted in the wild.
>>
>>109562444
Still waiting for Operation CheapSeek. I NEED the old prices.
>>
>>109562477
Just use MiniMax ffs
>>
File: DipsyMinnieSAV2.png (2.36 MB, 1024x1536)
2.36 MB PNG
>>109562564
>>
>>109562366
Thats good to hear, and I did decide to actually take a double check. Most of my usage is not in peak hours and im only spending like 3-4 bucks a month anyway. So I think I had a kneejerk reaction ill be fine.
>>
oh my god, deepseek just broke containment! save me sam and dario!!!
>>
>>109564447
Here's a full view of the pricing for DS models, from R1 launch to today, including the upcoming price change.
They're still cheapest but barely. My main neg with whole thing is V4-Pro-0813 didn't meet expectations of a mind blowing model. It's just... fine. And now the price is... as expected. Vs. performant and cheap AF.
>>
File: .png (599 KB, 1600x1528)
599 KB PNG
>>109564861
no, gpt-5.6 luna will be cheaper per task than flash-0731. almost beat flash-0731 before the price increase, see pic.
https://artificialanalysis.ai/articles/gemini-3-7-time-frontier
>>
File: 1761581599070409.png (1.2 MB, 941x1672)
1.2 MB PNG
>>109565128
Luna could be 1 cent per million in and out and I still wouldn't pay for it. No weights, no money
>>
>>109565322
yeah, if you consider subscriptions gpt-5.6 can be a magnitude cheaper than dsv4. so only ideologue would keep you on dsv4.
>>
>>109565128
Great. Let me know when oai stops giving me refusals and sending me fucking warning letters.
Fuck oai and Anthropic. They can keep their offering.
>>
File: Marketing.png (66 KB, 832x608)
66 KB PNG
What is your favorite type of marketing?
>>
>>109565905
Right on queue. See >>109562444
>>
File: 1785779938624054.jpg (61 KB, 1024x1024)
61 KB JPG
>>109564841 (me)
working on preset for dsh so it can work with most models. really good harness. kind of think this will be how models in 2027+ bootstrap themselves

PEACE
>>
>>109566425 (also me)
dsh is *extremely fun* to play with. if you like messing around with models and discovering **weird ideas** its going to get you far..
>>
>>109565556
How is 5.6 cheaper than free?
t. localGOD
>>
File: 1779365194227447.png (8 KB, 556x78)
8 KB PNG
>>109544573
Dayum!
>>
>>
>>109567967
What image model?
>>
'mp
>>
>>109544573
Same issue at Airbus, they stopped practicing continuous personnel hierarchic promotion through experience validation and are managing the company like a Volkswagen warehouse.
It's a perpetual stand still where experience is no longer rewarded, promoted or even kept.
Once the experience is gone: It's gone.
Now they're calling back retired shop's warriors paid a fortune to train monkeys with no background in mechanics.
FUCK YOU KRAUTS, YOU DID THIS. YOU KILLED THE "AEROSPATIAL"'S LEGACY IN CONTINUOUS TRAINING.
>>
File: Typescript.png (12 KB, 448x199)
12 KB PNG
>>109544573
Explain yourself.
>>
>>109570131
We have collectively stopped to give a shit. Now we just want to make others suffer.
>>
File: IMG_702.jpg (183 KB, 1080x1440)
183 KB JPG
can't beat cache hit price in ds4. poopenai shill didn't know this fact
>>
>>109566765
https://x.com/thdxr/status/2086599224674681242
>the average OpenCode Go user spent $1.14 per day on deepseek flash v4 this past week
>the dual DGX setup people are running to do the same costs $10,000
>it takes 24 years to break even
>at 10x the usage it takes 2.4 years
>>
>>109544052
Very kawaii picture in the OP
I have a question
When do the AI companies release "guard rails" which only produce ugly woke animation and images?
I realize they are making it good for now to get everyone hooked, but when do they take away this feature?
>>
>>109571198
offtopic. deepseek cannot generate images.

go ask in the right general: >>109568023
>>
>>109571074
Yeah, there's no financial business case for local unless you're selling inference. We'd need to be at point in tech where 1T DDR6 RAM is on consumer grade machines that sell for (at current dollars) $2K or so. We're probably 10 years out from that. Until then it's big data centers, 1970's style.
Reminder that big data centers really didn't get phased out until the early 2000s, when Windows boxes replaced stuff like SGx machines and UNIX machines for CAD and analysis. The PC was around for 20 years until they got good enough to replace them for all but the most complex tasks.
>>109571212
Based general re-director.
>>
>>109570293
deepseek trains on your input.
if i wanted chinese to train on my input, i could use a relay for openai, too. relay will be magnitude cheaper.
>>
File: 1756790944431643.png (2.9 MB, 941x1672)
2.9 MB PNG



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.