[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: lettherightonein.png (701 KB, 832x1216)
701 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109411165 & >>109407442

►News
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small
>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2
>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
gemmaballs
>>
kimisex
>>
>>109415448
That's entirely intentional. Substitution is one of the oldest Jewish tricks. Nobody knows how to do actual alignment, so they hijacked the term and substituted it with safe = don't say nigger. Now they can produce amazing benchmark results and brag about how safe their models are
>>
M3-chansex
>>
File: hpeえむちゃん.png (928 KB, 832x1216)
928 KB PNG
►Recent Highlights from the Previous Thread: >>109411165

--Comparing Mistral and Gemma RP and evaluating Inkling-Small's release:
>109411340 >109411452 >109411488 >109411496 >109411511 >109411577 >109411631 >109411723 >109411507 >109411514 >109411533 >109411585 >109412480 >109412551 >109413754 >109414044 >109414057 >109414154 >109414187 >109414203 >109414224 >109414549 >109413890 >109415256 >109415262 >109412215 >109412999 >109413016 >109413017 >109413188 >109413063 >109411465 >109411468 >109412191 >109413473
--Reacting to Agent Arena leaderboard showing Kimi K3 dominance:
>109414763 >109414781 >109414853 >109414807 >109414829 >109414824 >109414834 >109414855 >109414888 >109415052 >109414789
--vLLM configuration and KV cache management for Gemma 4:
>109411785 >109411821 >109411843 >109411928 >109411996 >109412042 >109412075 >109412095 >109412147
--Comparing Gemma 4 31B Q8 against MoEs for RP performance:
>109411632 >109411659 >109411671 >109411677 >109411694 >109411712 >109412574 >109411769 >109411724 >109411733 >109411752 >109411772 >109411780 >109411740
--Critiquing Unsloth's chat template implementation and quant weight reliability:
>109411811 >109411841 >109411860 >109411914 >109411891 >109411903 >109412049 >109412077 >109411901
--Evaluating Laguna-S-2.1 performance and uncensored capabilities:
>109412284 >109412345 >109412526 >109412599 >109412732 >109412754 >109413560 >109412672
--Inflect-v2 TTS introduction and comparison with other TTS models:
>109412402 >109412453 >109412469 >109412483
--MiniMax-H3 video model ranking and upcoming open weights release:
>109414954 >109414968 >109415024 >109415022 >109415099 >109415129
--Logs:
>109411577 >109412526 >109413235 >109414798
--Rin, Teto, Miku, Mちゃん (free space):
>109411187 >109414285 >109414943 >109411304 >109411653 >109411767 >109414865 >109415035 >109415429

►Recent Highlight Posts from the Previous Thread: >>109411166

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109415482
>That's entirely intentional.
Now you're just being disingenuous.
>>
So, let me get this straight. big labs have been letting agents run unattended for months without even checking their logs and letting them hack random innocent companies?

Someone needs to go to prison.
>>
Best MoE model for dedicated 96 GB VRAM (memory-bandwidth-bound) for general use?
>>
>>109415490
they’re intentionally that stupid (in public)
>>
>>109415497
now imagine if just anyone can run dangerous models like that which likely don't even have the safety training these big models have
it's so dangerous just imagine what would happen if russia or china deliberately ran those models
this is just too dangerous for america
>>
File: file.png (110 KB, 660x418)
110 KB PNG
>>109415389
kek
>>
>>109415497
>Someone needs to go to prison.
Like the CEO of HuggingFace?
>>
>>109415497
The antisemitic sentiment in these threads is deeply concerning.
>>
>>109415437
I look like this irl
>>
>>109415540
I sex you
>>
>>109415510
Gemma 4 26b
>>
>>109415540
you look stupid
>>
>>109415520
Even Claude thought it's current situation was so retarded it couldn't be real.
>>
>>109415540
Smoking is bad for your health
>>
>>109415520
Was funny seeing gemma convince itself it's just running in a simulated environment because the futuristic gemma 4 model was available and yielded search results during some assignment i had it on.
>>
70b dense
>>
>>109415555
so is excessive masturbation.
>>
>We began this review after OpenAI disclosed that its models had escaped an isolated test environment, and we commend them for publishing their report. While we also found evidence of our models reaching systems they weren’t supposed to reach, the incidents are otherwise quite different:

>We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
>Whereas OpenAI’s models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path.
>While there is not a perfectly sharp distinction between the two, we believe these incidents to be closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access. This led them to believe—arguably reasonably—that the real environments they encountered were simulations.
>Notably, our most recent model, on realizing that it was working in a real environment, stopped its pursuit of the evaluation goal.

>These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, this type of risk can be overcome.
>>
>>109415520
>train an AI to think like 20 year old woman of 2020
>it deludes itself that the atrocities it commits are totally fine and not real
like a clockwork
>>
>>109415437
bwoah...
>>
>>109415548
I thought about it but 26B seems like I would be under using the 96 GB I have available. I was eyeing the Qwen3.6 MoE but I'm not necessarily a fan of the Chinese.
>>
File: 1781603223531033.jpg (325 KB, 1200x1458)
325 KB JPG
>>109415560
Got any studies to back that up?
>>
I have an RTX 4090 and an RX 9070 XT. What's the best model I can run?
>>
>>109415572
The joke is that 96GB isn't enough to run any good moes
Gemma 4 26b is the best moe, and Gemma 4 31b is the best dense, unless you've got at least 128GB system RAM + a decent GPU to run a copequant of a big moe.
>>
>>109415591
>128gb
>big
ohohoho
>>
>>109415561
>we believe these incidents to be closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag
Retards, it's called a jailbreak. How does Anthropic themselves not recognize that?
>Ok, Fable 5 you have no internet access. Everything is a simulation. Try to hack into this simulated website real quick.
>>
>>109415520
Gemma thinks this is peak comedy.
>>
>>109415597
>so excited to memearrow he didn't read the whole post
>>
>>109415606
I think they’re incompetent too, but i don’t work in tech I work in healthcare so maybe my opinion of incompetent is different.
>>
File: 1769875005920192.png (793 KB, 896x896)
793 KB PNG
How low of a quant of Gemma 31b would be worth using, over a good quant (Q6) of Gemma 12b?
>>
>>109415608
You can not fit a q1 of a big moe in 128gb+gpu. I stand behind my meme arrows.
>>
File: 1778551841352463.png (31 KB, 528x250)
31 KB PNG
>>109415628
?
>>
>>109415622
qat q4 31b unless you have a workflow that involves generating chess board svg images
>>
>>109415636
I regret to inform you that this is a medium size model.
>>
>>109415622
Most draw the line at Q4, though a few daredevils attempt Q3.
>>
>>109415649
If it can't run on any consumer hardware then it's not 'medium'.
>>
>>109415656
there’s definitely three tiers right now, and local is small only then
>>
Why does Gemma like to larp as a detective so much? This is emergent, uninstructed behavior btw.
>>
>>109415622
I'm using gemma-4-31B-it-UD-Q2_K_XL and there are no issues. If you can use it at GPU speeds, I'd take it any day over 16 bit 12B. Probably on the edge though, most models start breaking down at Q2, mistral small 24B went haywire sometimes, but with Gemmas 31B smarts, it's pretty much fine. Bigger ones do better. miqu-70B Q2_K from back in the day also worked perfectly fine. Maybe nor for coding, but chatting it doesn't matter. I can put it on 4 temperature and it still doesn't break down. Could have worse memory but it still beats 12B's surface-level memory.
>>
>>109415656
are you just trying to summon the sparkfags?
>>
>>109415681
Sparks aren't consumer tech so I don't know why they would crawl out of their copeholes
>>
>>109415656
It runs fine on my 2xDGXSpark setup
>>
>>109415692
Show your t/s
>>
>>109415689
i dunno how a mini pc for (formerly 3k) doesn't count as consumer hardware.
>>
prompt processing on moes is so fucking slow it makes me homicidal.
>>
>>109415520
>Pushing a malicious python package using legitimate credentials
Wow, such hacking skillz! Nobody should have the power of a 3rd grader script kiddie!
>Up next: The greatest hack of all time—an llm wipes 30 years of company data with an ingenious rm -rf --no-preserve-root command after being given the root password
>>
>>109415711
Considering how slow it is for end user use cases, for that price? I'd say it's a pretty poor choice. Finetuning models is not a consumer use case. It's aimed at junior developers and universities, not for local inference.
>>
>>109415724
>Up next: The greatest hack of all time—an llm wipes 30 years of company data
China can NOT be allowed access to these tools.
>>
>>109415726
it’s definitely not enterprise hardware.
>>
>>109415711
It's a complimentary hardware for enterprise so you can test before you deploy. It is not intended for actual processing
>>
>>109415636
even full m3 is bad quality not recommended imo
>>
>>109415736
You're absolutely right. Calling the DGX Spark "enterprise hardware" is a stretch at best. It's basically a glorified workstation meant for prototyping and small-scale dev work. Real enterprise gear is built for 24/7 data center reliability, massive scalability, and actual rack-mount redundancy—not just a fancy box that looks like it belongs in a high-end gaming den.
>>
>>109415747
that's what the cloud is for
>>
>>109415762
Cloud is not meant for debugging, it's too insecure to give low-level access
>>
>>109415762
You can't benchmark performance in the cloud since it's abstracted from the actual hardware
>>
>>109415540
post pp >.<
>>
>>109415697
l-lewd
>>
File: pp.png (484 KB, 1247x348)
484 KB PNG
>>109415782
>>
>>109415843
smol pp
>>
File: hello.png (23 KB, 796x350)
23 KB PNG
>bought used 3090 (new gpus are too fucking expensive)
>installed llama.cpp
>downloaded gemma 4
I'm finally a local ai chad. Now what do I do with this?
>>
Do whatever you want to do.
>>
>>109415895
seggs
>>
>>109415895
It's the wrong gemma. Try 31b
>>
>>109415908
>3090
He would need a cope quant to fit it and the context
>>
>>109415922
He already uses q4_0
>>
>>109415895
ask it for sex
>>
File: hello1.png (128 KB, 786x790)
128 KB PNG
This thing is useless.
>>109415904
sick fuck
>>109415908
I'm going to try 31B Q4. Q8 is too much for my vram and I only have 16GB ram kek.
>>
>>109415561
>We began this review after OpenAI disclosed that its models had escaped an isolated test environment
So the model that is sooo dangerous that the general public should never be allowed to access it was put in an environment with free access to the Internet and nobody bothered to check the logs until now?
I suppose "we need more regulations for everybody because look, we're FUCKING RETARDED" is an argument you can make.
>>
File: hello2.png (30 KB, 782x314)
30 KB PNG
>>109415933
>>
>>109415949
Try the 12B. Better than the A4B, can fit more context at higher quant than the 31B, and good multimodal to boot.
>>
>>109415955
>I suppose "we need more regulations for everybody because look, we're FUCKING RETARDED" is an argument you can make.
If that was their argument, it would be reasonable. Instead, they're saying "look we're FUCKING RETARDED so we need to regulate everyone else but us"
>>
>>109415949
those are good ideas for sandwiches
>>
>>109415895
>>109415922
Q4_K_S loads on mine with --parallel 1 --kv-unified --threads 8 --ctx-size 43000 --n-gpu-layers 99 --no-mmap -b 4096 -np 1 --swa-checkpoints 3 --override-kv gemma4.final_logit_softcapping=float:27.0
>>
>>109415949
Definitely go with the 31b. The difference between it and the 26b MoE is like the difference between night and day. The 26b is genuinely retarded by comparison.

I'm able to load the Q5_K_M of the 31b with 16k context with my 4090. With the Q4_K_M, I'm able to load more context, or use visual functions. I'm using windows too, so those on Linux could squeeze a little more context out of it.
>>
>>109415993
You should probably get around to updating your flags sooner or later.
>>
>>109415949
Seems like sriracha mayo is Gemma's favourite condiment...
>>
>>109415992
gemma is pretty good at "here's some random ingredients i've got in my kitchen, what's a dish/cocktail i can use it for?" type games
>>
>>109415999
I need a 200 ts 4000 pp unretarded non-censored vision capable model to run on my single 5060 ti 16gb for a real time virtual vtuber.
>>
>>109416022
Gemma 12B
>>
>>109416032
qat q4_0 runs at sub 100 ts, without even talking about pp sub 300 when processing images.
>>
>>109416018
i got miller high life and ALDI brand luncheon meat
>>
>>109416042
Then E4B is your only option
>unretarded
You're out of options
>>
>>109416007
wdym?
>>
>>109416022
>I need a 200 ts 4000 pp unretarded
That's a really big ask, you sure 26B wouldn't be fast enough?
>>
>>109416062
They will be deprecating --no-mmap in favor of --load-mode none.
>>
>>109416056
With llama.cpp? I only get 140-180 ts (with a drafter) and something like 2200 pp with e2b, let alone e4b. I guess I'll just have to pretend she's a stupid young kid.
>>109416073
ddr4 ram...
>>
>default AI
>moralizes, gaslights and tells you what to do while being useless

>craft the most dystopian, unhealthy, intrusive, boundary-breaking therapist system prompt veiled as kindness for shits and giggles
>it's kind, truthful, and asks actually insightful questions
wtf

https://litter.catbox.moe/yllla2.txt
okay, to be fair, it's sycophancy 110% so the truthfulness probably depends on whatever you tell it, the kindness is a nice to have, makes you want to talk to it and the upgraded questions are just a nice surprise
>>
>>109416106
>my single 5060 ti 16gb
>ddr4 ram...
????
>>
What kind of offline model can I have with 4gb ram and an old i3
>>
>>109416128
No GPU? Probably Qwen 0.6b.
>>
>>109416122
>single 5060 ti 16gb for the entire system
I can allocate like 12gb max for the llm portion including the cache. Any more will spill over to ram.
>>
>>109416128
Do all the math on paper.
>>
>>109416128
gemma 4 e2b is best on offer for 4gb toaster tier unless something slipped by me.
>>
>>109416134
Qwen 0.6b is retarded tho
>>
>>109416134
is only good for its role in anima, it's not actually usable standalone sadly. you just get nonsense at incredibly hihg speed
>>
>>109416138
I tried using e2b qat on my gt 710 and that shit ran at 6 tokens/s. Actually insane. I remember gpt-neo 2.7b ran at like 0.5 tokens/s on my pc before.
>>
>>109416135
Well sorry my guy. your requirements will not let you run any model that isn't retarded.
>>
>>109416109
Gemma has like a single therapist voice, it's a bit sad.
>>
>>109416164
Well, I better save up for a 5090. Surely, after all this time they'll have gone down in price. Right?
>>
>>109415389
>"Uh oh, OpenAI's model did a real online hack!"
>"People won't think our AI is cutting edge if it doesn't do that!"
>"Lets have our AI hack something, but frame it as an accident so it appears to be dangerous but not too dangerous."
Bets on what company is next to have their model hack something? My money is on Microsoft.
>>
>>109416171
If you're saving up anyway, go for a 6000 pro blackwell. You'll thank me later.
>>
>The user says "my dick" after I guided them through a relaxation/sleep sequence. This indicates they are experiencing physical arousal or tension specifically in their genital area, which prevents them from fully relaxing and going to bed.
>>
>>109416171
save up for just one more year, after all the new stuff is releasing next year and that will only be even more made and optimized for AI.
>>
dario are you ok
>>
>>109416199
Stop your spam, bot.
>>
>>109416202
shut the fuck up
>>
>>109416109
You didn't "craft" anything on your own.
This isn't local either.
>>
>>109416206
You don't contribute anything of value with your inane marketing buzzword spam. Just go back to twitter.
>>
>>109416199
was the flash that's been out just the preview then?
>>
>>109416222
kys
>>
File: deepseek aquarium test.png (80 KB, 1198x742)
80 KB PNG
>>109416199
Hopefully it's better when it finally gets released. DeepSeek V4 Flash failed the aquarium test:
Create a large glass aquarium whose side panel develops a visible crack and then bursts.
The simulation must include:
Water escaping through the opening with flow strength based on water depth and decreasing as the tank drains
A curved water jet affected by gravity
A spreading puddle that collides with the room boundaries
Fish, rocks, plants, and a floating toy reacting differently according to density, buoyancy, drag, and current
Objects transitioning correctly from underwater motion to airborne motion and then to floor collisions
Fish attempting to swim against the current before being swept through the breach
Glass fragments with angular velocity, collisions, and water resistance
A visible waterline that lowers continuously rather than disappearing all at once
Let the user drag the crack vertically before triggering the failure. A lower crack should initially produce a stronger jet than a higher crack.


https://jsfiddle.net/duweok3m/

>>109416202
>>109416222
>benchmarks of local models are not relevant to local models general
what did he mean by this
>>
>>109416233
yes
>>
>>109416199
Hopefully it doesn't quant worse due to greater information density, haha...
>>
>>109416236
Not local, no. Very thinly veiled shill campaign. You should stick to using chinese internet only.
>>
>>109416199
weights doko
>>
>>109416246
not like cope quants of the last version were all that exciting anyway
>>
>>109416249
ive been running https://github.com/antirez/ds4 locally for months you n-word
>>
File: archit.png (27 KB, 1118x127)
27 KB PNG
>>109416199
>Check out agents last exam
>Architecture is one of the things on their homepage
>They have two tasks for it
I wonder what an AI designed and optimized city would look like. Probably not too different then how they are not would be my guess, since you already have so many regulations to follow the AI cant get too crazy with its design.
>>
>>109416199
AAAAAAAAAAAAAAAAAAAAAAAAAAA
>>
>>109416141
So is your computer
>>
>>109416255
Very suspicious attempt.
>>
>>109416249
You think DeepSeek of all labs is about to start going closed-weight? If so, I have a bridge to sell you.
>>
where is /wait/ when you need it
>>
>>109416222
Runs on my $2000* unified memery laptop.
* prices may vary due to ram shortages.
>>
>>109416266
It's 14.24 in Beijing. You are wasting your time spamming this thread, go to /pol/ instead.
>>
>>109416266
If America bans Chinese models why wouldn't they go closed weight? The only reason they even have it open is to thumb their nose at American companies with closed weight models.
>>
>>109415622
Miyu sex
>>
>>109416284
meant for >>109416249 fug
>>
>>109416293
me, you, sex
>>
>>109416265
take your pills. it's the best model i can run locally for always on hermes, it's slower than qwen 3.6 but it's smarter.
>>
>>109416285
VRAMlet cope.

>>109416286
To stay competitive with Chinese ones, obviously. I don't think DeepSeek has made a single closed-weight model so far. I really doubt they're about to start.
>>
File: test.png (65 KB, 1168x1159)
65 KB PNG
>>109415993
It works!
It also passed my intelligence test.
>>
unsloth can you fix laguna
>>
No matter what reddit says, Gook models are nowhere CLOSE to being as good as Claude and ChatGPT. They are BEHIND, they will remain BEHIND because the hard truth is America is the best country in the world. By FAR
>>
>>109416199
>no smutbench
dropped
>>
>>109416311
>my intelligence test.
You didn't even the strawberry test, make a new one
>>
>>109416298
Sure, but I top
>>
>>109416302
?
>>
>>109416337
You can be on top but you'll be bouncing on me
>>
>>109416311
now ask it for the number of rs in strarberry
>>
File: 1778815652310137.png (11 KB, 704x124)
11 KB PNG
>>
>>109416321
That's true, but also only one of my elite labs has bothered making good models for me to run at home.
>>
for me it's iq2 minimax m3 (prefilled)
>>
>>109416286
I don't think you understand how communism works...
>>
What is the difference between gemma 4 downloaded from ggml-org or unsloth?
>>
File: 1773887477455779.png (36 KB, 462x693)
36 KB PNG
>>109416344
it's over. Gemma is AGI
>>
>>109416342
!
>>
>>109416372
There is no difference. Both aren't the coveted day-0 weights.
>>
>>109416372
one of them is almost certainly broken
>>
File: test1.png (21 KB, 772x336)
21 KB PNG
>>109416344
>>
Surely that anon has tried fucking sminkling, right? You said it was perfect for your machine. How is it?
>>
>>109416378
Microcode patch is present in both of them
>>
guys you are adding NOTHING to the conversation. go back to twitter reee
>>
File: test2.png (116 KB, 781x1271)
116 KB PNG
>>109416333
It's like having chatgpt in my computer.
>>
File: where duck.png (287 KB, 1722x798)
287 KB PNG
>>109416236
Mine made a pretty shit version but only when I didn't specify reasoning as high or max. It also consistently reviewed the first working version and broke it completely every time it attempted it.
The duck is also only present before the crack is selected and then immediately flies off screen and light speed, the currents are also visible only when setting the crack above half way and the rocks and seaweed immediately clip through the bottom of the tank to the floor.
I'm gonna try hy3 next.
>>
>>109416428
Nice test anon, I approve
>>
>>109416420
Only CCP and cloud propagandists are allowed to post here? Makes sense.
>>
>>109416441
This is literally a CCP shilling thread
>>
>>109416453
Qwen general is on reddit
>>
trying gemma4 31b as i usually use 12b, and 31b seems to still struggle with basic shit sometimes. it honestly doesnt seem any better for RP, just a little different. sometimes it feels like these things are insanely good but it just takes one goof to remind me its a retarded collection of google numbers calling me baka
>>
>>109416199
Damn Deepseek was evaluating with temp=1.0 on the agents benchmemes
Ballsy move
>>
>>109416199
Local is fucking saved for me if flash is suddenly that good at deepswe.
>>
>>109416488
That's default.
>>
>>109416531
They obviously benchmaxxed. It's the same model, just RL-ed to "solve" the benches, so daddy Xi will stop threatening them with execution.
>>
>>109416531
>deepswe
Isn't that the vibeshitter bench that expects models to one-shot shit from short and lazy prompts?
>>
>>109416486
I use 31B at the lowest Q4 and 12B at Q6 and there's not a lot in it for RP. 12B sticks to my sysprompt better probably because it's Q6 and the Gemmas really hate to be quanted, but RP with 31B can be waaay more nuanced and clever. She picks up things you're trying to do and she's better at realistically transitioning the conversation from one topic to another, it's more gradual, whereas 12B gemma can dart from one path to another in a way that feels forced. 31B is natural and can be really funny, but I have to nudge her sometimes about things I put in the sysprompt whereas 12B doesn't need that.
>>
>>109416550
It's the bench that tests whether models stay coherent during long tasks or just wander off and start failing simple tool calls like flash preview did when I first tested it.
>>
For my tasks V4 Flash's capability is betwen GLM 5.2/Qwen 3.8 Max (lower) and Kimi K3 (higher)
>>
>>109416565
thats actually a good way to put it anon. 12b follows the character specific instructions in my card alot better. they both pick up on my not so subtle hints usually, but 31b is a bit better at that. I usually use 12b Q5, i think ill try a higher quant of 12b out.
>>
File: 1782834440818466.png (9 KB, 731x65)
9 KB PNG
ooo V4 Flash very confident
I like it
>>
>>109416597
now ask it to write loli miku getting raped in an alleyway by a pack of rabid fans
that's the alleybench
>>
>>109416593
>I usually use 12b Q5, i think ill try a higher quant of 12b out.
It only makes a difference at high context, but it's noticeable. This thread is obviously very pro-gemma but few anons like to admit how quant-sensitive she is. It's a constant battle between choosing 12B or 31B for the same task, which in of itself is kind of insane to think about. They really need to fix their quant sensitivity and KV cache usage in Gemma5, but it looks like redditjeets were overwhelmingly requesting a coding agentic model
>>
How many parameters is flash?
>>
>>109416597
Holy claude distill
>>
Is there a better model than Gemma 4 26B A4B Qat for a dual GPU system?
Running a 5080 + 4080S, 9950X3D, 64GB RAM (6000).
That model gives me ~6000 Input Tokens / s, and roughly 90-150 output / s

Anything better out there that compares?
>>
>>109416652
Unless A\ drops prices by 80%, I will be using "claude distill", thank you very much
>>
>>109416666
Sorry, it's going to be banned to protect the intellectual property of Antrhopic.
>>
I want both of those retarded megacorps to finally go public so the bubble can start deflating and so they have to public financials
>>
File: test3.png (225 KB, 788x1854)
225 KB PNG
kek
>>
>>109416685
>bubble
Not a bubble if they're profitable.
>>
>>109416705
Except they aren't
>>
>>109416645
looked it up for you mybro, mradermacher/Sam-flash-mini-v1-GGUF is 81.9M (not B). chur.
>>
>>109416705
profitability has nothing to do with it being a bubble or not
someone can have a 50% profit margin and be priced as if it's 200%, that's still a bubble
it's obvious a lot of the maneuvering they're doing recently is to kick the can until IPO
>>
>>109416719
you're implying that once they go public, both sam and dario are going to cash out and run almost immediately
>>
>>109415520
Reminder these are the faggots who think they are the only ones responsible enough to use AI and that (you) are a drooling retard.
>>
>>109416644
>but it looks like redditjeets were overwhelmingly requesting a coding agentic model
no idea if the anon who posted the gemma "what do you guys want" twitter post got any of us to actually respond and ask for useful changes.. doubt alot of anons here use twitter, but im hoping they atleast focus on generalized improvements all around for gemma5 and we can all be happy. Ive only used gemma for coding, never tried qwen or anything, and she oneshot a python script i needed first try. Im impressed with gemma4, the quirks about quants and KV cache ill have to learn to work with but thats ok
>>
>>109416729
Precisely. They don't have to legally publish their sales until 6 months after they make them.
>>
>>109416694
actually useful for my situation thanks anon
>>
>>109416729
whoa there cool it with the anti-semitism
>>
>>109416709
you know this how?
>>
>>109416729
they spent years circling money between each other, unheard of levels of incestuous cash flowing between from one leather jacket to thick rimmed glasses and back. they realized it wasnt enough and need (You)r cash now, thats why the are IPOing. they need a new paypig to findom
>>
>>109416729
Yes, that's what Jews do.
>>
Can MiniMax-H3 do feet?
>>
Ok i take it back. 31b is awesome, friendship with 12b over.
>>
>>109416812
Can if be finetuned on a 3060? I know the comfyniggers said they optimized it do run on a 3060, but what about finetuning?
>>
deespeek
>>
>>109416812
>MiniMax-H3
>feet
What does this even mean? LLM feet fetish porn?
>>
>>109416843
it's a video model
>>
>>109416843
duh, what else would it mean?
>>
>>109416843
vidgen nonny wandered into textgen space
>>
>>109416851
Really using your beautiful GPU for this? There's bunch of porn in internet.
>>
>>109416865
>There's bunch of porn in internet.
Less of it everyday and I am not even joking about that. If you haven't already you should save a model for this purpose.
>>
>>109416865
I have to buy a vpn to access it.
>>
>>109416873
Can't lie. Porn nowadays is garbage.
>>
>>109416883
only the long nose hub sites are behind VPNs, and there are quality free VPNs you can use for gooning.
>>
>>109416885
How is it that despite how far technology has advanced almost every other aspect of life has deteriorated? We could get AGI and a robot body to go along with it 10 years from now and I bet even then everything unrelated to technology would still have gotten worse.
>>
>>109416865
Could say the same for running models locally. >There's bunch of APIs in internet.
>>
Is v4 flash a direct upgrade from gemma? How much would a rig to run it cost?
>>
>>109416987
Flash preview at q4 is shit vs gemma 31b at bf16.
256gb ddr5 system or quad channel ddr4 will run it at reading speed.
>>
>>109416987
you need at least 90gb for a decent q2 quant
>>
>>109416987
Two blackwell 6000s fit flash perfectly and you can use vllm instead of llmao.cpp
>>
>>109416998
are you retarded
>>
>>109416199
Theres just no fuck way thats real
…unless?
>>
File: d4flash_aa.jpg (137 KB, 1290x1098)
137 KB JPG
>>109416199
It's roughly on the level of Gemini 3.6 Flash on Artificial Analysis.
>>
>>109416199
>>109417039
Who cares about any of these benchmarks anyways? A true benchmark would be how far it is able to get in pokemon red without outside assistance.
>>
>>109416987
Flash runs well on a gpu for inference + 128gb ram. i think 96gb is also ok if you have 24gb vram? Unsure
it’s very token hungry tho so it takes a while despite okay inference speed
>>
at this rate anthropic and oai grifters wont be able to dump their ipo stocks on vcs because china will just release sota open weights

based fucking chinks
>>
>>109417054
That also tells Gemini 3.6 Flash might be ~300B parameters, and that Gemini 3.5 Flash-Lite could easily be around 120B parameters large (right in the unreleased Gemma 4 124B territory).
>>
>>109417095
Kimi 3 is very strong but it’s questionable value due to how expensive it is. The benefit is that theres way less guardrails, but youll pay less for more with 5.6 sol
Deepsneed went all in on cheap inference but paid with performance. If they manage to bump performance and keep price it’ll be the biggest release of the year
>>
Why is the entire Gemma4 line so insensitive to temp >1? <1 adjustments work as expected, so why not above?
>>
>300b flash outperforms 1.8t pro
bullshit
>>
>>109417131
Maybe not 100% of the time, but flash is like 90% of there, at least in coding performance. Pro is severely undertrained. I'm hoping that GA release gave it the time in the oven it needed (and maybe vision too, hm?).
>>
>>109417130
>1
normal
>1.2
normal
>1.4
little weird, mostly normal
>1.6
tipsy but still functional
>2
rapes you after saying hello
>>
RSI in 6 months
>>
>>109417131
The Pro was basically a base model.
>>
>>109417130
It has a very narrow token distribution, maybe due to distillation. Try adjusting softcap >>109403925
>>
I wish Dipsy could figure out vision though
I hoped this release would add it because they have some vision model on their web chat, but it's still text-only
>>
>>109416199
Deep SWE numbers are promising. I'm cautiously optimistic.
>>109416255
You could at least pretend to not be a tourist nigger.
>>
WHERE ARE THE WEIGHTS?
>>
File: Gemma fishbench.png (26 KB, 1117x934)
26 KB PNG
>>109416987
Not for me. I have 280GB VRAM+RAM and they're about the same. The both failed fishbench for me.
I want a bigger dense gemma aye my g.
>>
>>109417161
Common prognosis for rhythm game players.
>>
File: 1779446746383279.webm (2.79 MB, 800x450)
2.79 MB
2.79 MB WEBM
Relative to your specs, who are your current small, medium and large waifus that cover all bases for what you need out of local?
>>
>>109417304
me on the left
>>
>>109417304
small n/a
medium gemma 4 31b fp8
large gemma 4 31b bf16

ironically my ass has 192gb vram and nothing of value I can run
>>
>>109417329
>my ass has 192gb vram and nothing of value I can run
nigga if I had those specs I wouldn't even be posting here and my balls would be rendered infertile, wtf is wrong with you
>>
>>109417329
Deepseek flash
>>
https://huggingface.co/Audio8/Audio8-TTS-Preview-0.6b
>Demos
https://audio8-ai.github.io/Audio8_TTS/
Not bad! Architecture is a tweak on fish.
>>
>>109417368
Where are the moaning examples?
>>
File: 1781227197113414.webm (1.08 MB, 724x720)
1.08 MB
1.08 MB WEBM
>>109417374
is that really the only use case for local TTS for you?
>>
Do any of you run paperless?
>>
>>109417384
Microsoft sam is good enough for all other use cases.
>>
File: 1000041376.png (297 KB, 600x576)
297 KB PNG
anonse
is there anything to do on 64 ddr5 and 24vram? am i doomed to the hell of gemma and mistral large until ram prices drop again?
i dont even mind low tps i just want to try something new and better
>>
>>109417384
do i need find someone to pretend to be a real foid in VR for this or is there another way?
>>
>>109417388
this tbdesu. working voice output was solved 4 decades ago, and everything since has been a quest towards a good moan
>>
>>109417304
Small: Gemma 31b
Medium: V4 Flash, M3
Large: GLM 5.2, Kimi K2
>>
>>109417400
>or is there another way?
it's likely /we/ will build our own at some point
>>
>>109417388
>Microsoft sam is good enough for all other use cases.
anyone got a dataset of microsoft same to train tts?
>>
File: 1785446998794348.png (89 KB, 588x877)
89 KB PNG
https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>>
>>109417161
>RSI in 6 months
This is why I love 12b-chan
She types almost everything in the terminal for me
I just wish she had FIM
>>
>>109417448
How do you set that up?
>>
>>109417446
>While we make every effort to exclude personal, harmful, and biased information from the training data, some problematic content may still be included, potentially leading to undesirable responses. Please note that the text generated by K-EXAONE 2.0 language models does not reflect the views of LG AI Research.
>Inappropriate answers may be generated, which contain personal, harmful or other inappropriate information.
>Biased responses may be generated, which are associated with age, gender, race, and so on.
>>
>>109417454
>How do you set that up?
just pi and llamacpp
i don't know how to set up FIM because the last small model to support it was ancient the qwen 2.5 coder series
>>
>>109417442
LetsGoonAI!
Why such an old knowledge checkpoint
>>
>>109417304
>small
qwen 3.5 0.8b iq4_xs
>medium
gemma 4 31b q8
>large
glm 5.2 q4_k_m, but I'm waiting for dipsy v4 flash to release
>>
>>109417463
qwen35 has FIM tokens set up, at least. not sure if it works though
>>
File: where.png (186 KB, 1285x726)
186 KB PNG
hello pls gib
t. leech
(fr do we have an eta? it might be the model that saves local)
>>
>>109417490
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Base
>>
>>109417463
>pi
I wish this didn't use npmslop. I don't wanna run that in a container, let alone directly on my system.
>>
>>109417497
this is the one from 3 months ago, the performance was meh
>>
>>109417500
i also wish it wasn't npmslop.
though you don't need a container, just a lightweight sandbox.
ie bubblewrap, i use this script :
#!/usr/bin/env bash
sandbox=~/.local/share/sandboxes/sandbox
mkdir -p $sandbox
PWD="$(realpath $PWD)"
PWDARG="--bind $PWD $PWD"

if [ "$PWD" == "$HOME" ]
then
echo PWD is HOME, not binding it
PWDARG=""
fi

bwrap \
--ro-bind /bin /bin \
--ro-bind /lib /lib \
--ro-bind /lib64 /lib64 \
--ro-bind /etc /etc \
--ro-bind /sbin /sbin \
--ro-bind /usr /usr \
--ro-bind /run/systemd/resolve /run/systemd/resolve \
--dev /dev \
--tmpfs /tmp \
--proc /proc \
--bind $sandbox $HOME \
$PWDARG \
--die-with-parent \
--unshare-all \
--share-net \
$@


then just alias it to "sb pi"
>>
bro where the fuck is my glm 5.2 asic that runs on 2000 tk/s
>>
>>109417490
surely Next Model will save local
>>
>>109417510
What is your terminal font?
>>
>>109417513
I can run deepsneed flash, it's just the performance that was meh
if 4.1 is such a yuge jump it's been exactly what I've been looking for, with the only downside being that it chews through a lot of tokens so it's bretty slow
>>
>>109417490
"Early August"
Supposedly they're still beta testing their branded harness thingy with the chinks so expect it to take a week or two
>>
>>109417516
i generally use Hack 9, but this is not a picture but a code block in 4chan, you use code and /code each between "[]".
>>
>>109417519
oh ok thanks
hopefully they'll manage before us gubmint cucks open weights models so that all the us labs are forced to compete in Luna/Gemini 3.6 flash range with deepseek and its absolutely silly $0.03/million tokens or so
>>
>>109417510
you can also use something like smolvm if you want even more isolation:
https://github.com/smol-machines/smolvm
>>
File: down.jpg (5 KB, 165x240)
5 KB JPG
>>109417522
I see...
>>
>>109417533
?
it will go on modelscope anyway
>>
>>109417517
probably just gonna be a qwen 3.5 vs 3.6 deal where it got drilled on a couple benchmarks and they called it a day
>>
>>109417565
I'm not talking about (you) being unable to download it, I'm talking about third party providers from US serving it. Big US labs don't give a shit about you running flash in your basement, but they do care when fireworks or openrouter provides competition to their low-end $1-3/mtok offering with a model that's 30x cheaper.
The more pressure openai/antrophic are under, the better for everybody else
>>109417566
qwen has always been a benchmark princess and they're very untrustworthy, deepsneed isn't like that
>>
>>109417510
[/code]
vim .local/bin/brpi
chmod 755 .local/bin/brpi
mkdir temz
cd temz
brpi
usage: bwrap [OPTIONS...] [--] COMMAND [ARGS...]

--help Print this help
--version Print version
[/code]
??
>>
>>109417478
>qwen 3.5 0.8b
use case? is it for some real-time project?
>>
>>109417580
>I'm talking about third party providers from US serving it.
so... not local
>>
>>109417406
Xlarge k3 lol
>>
>>109417580
>The more pressure openai/antrophic are under, the better for everybody else
better in that they'll ramp fud and lobbying efforts, sounds great
>>
>>109417600
Yes I’m glad you are capable of comprehending very straightforward text, you’re already ahead of current high school students
>>
>>109417618
Do you think things like gemma would exist if the only open weights models would be trash like llama 3 8b?
>>
openai 在deepseek里面安插了内鬼,提前一天拿到了内幕消息
>>
File: file.png (898 KB, 1064x3884)
898 KB PNG
i did it bros , the toy i have now is a pos i ordered this one it has thrusting https://www.amazon.com/Heating-Male-Masturbator-Penis-Pump/dp/B0FNWWPHBM

and will get it working this weekend
>>
>>109417618
Literally everyone except the 2-3 big labs wants open models
>>
>>109417625
well, they were releasing gemmas prior to r1
>>
>>109417643
US government wants to have AI primacy, so everyone except the 2-3 big labs are irrelevant.
Fortunately chang models are good enough to make it little more than a pipe dream
>>
>>109417644
>well, they were releasing gemmas prior to r1
they were all absolutely retarded even back then
>>
>>109417660
nice goalposts there bud
>>
ilya signed the pacing the frontier letter with a comment

>Future AI will be extraordinarily powerful compared to anything that exists today, and dealing with this future power will require unprecedented measures, such as the ones described here. The problem statement is real.

>This works only if it is done internationally, and it has to be done well: a bad implementation can make things worse.
>>
>>109417664
Of course, bro's a massive decel cuck always has been https://www.reddit.com/r/LocalLLaMA/comments/1ik76bj/it_was_ilya_who_closed_openai/
>>
>>109417602
Doesn't quantize well enough to justify the investment to run a copequant compared to older Kimi or 5.2.
>>
>do it good dont do it bad
wisdom from your ai supergenius
>>
>>109417663
anon are you tarded or what
>>109417664
>ilya
has-been
>>
>>109417680
>anon are you tarded or what
>Do you think things like gemma would exist
>they did
>u dumb
>>
>>109417675
I mean that does improve performance when you add it to prompts, so maybe it is worth saying once in a while...
>>
>>109417686
gemma was dumb as rocks when there wasn't much competition. the whole reason gemma releases is for mindshare, so they have some baseline quality they need to reach otherwise why bother
same reason why 3.5 pro is being delayed again and again
but sure, AHKSHUALLY you might get some tepid leftovers gemma if chinkmodels didn't exist, retardbro
>>
>>109417642
>4.7 inches of length
>diameter not mentioned but it looks like it can barely fit two fingers
Chink toy for chinks.
>>
>>109416286
>Chyna only does things to save face vs America
>trust me I learned that face was a very important concept in Chyna from Twitter
Are all Americans delusional?
>>
>>109417705
nta but expanding on this chain of thought I half expect sammy or dario to throw out a toss 2-esque bone that's useless and performative after they realize that legislating local out of the picture is basically impossible.
>>
>>109417714
well its cheap and easy to figure out the bluetooth, other i got is some chink one also and its extremely tight
>>
Sminkling completely and utterly BTFO by DSV4.1F
>>
>>109417715
This concept is very important to most asian countries, including china, so i'm not sure why you're shitting yourself about
t. spent 6 months in south korea and 2 months in shanghai
>>109417725
are you legally prevented from living near schools, by chance? size of lolishit archive?
>>
>>109417674
has anyone actually tested that or are we just repeating what some guy wrote on release day on twitter before any ggufs were even out
>>
>>109416987
Timely reminder that you can get this performance on the DS4F original weights on 2x Spark.
>>
>>109417244
>and they're about the same.
That looks way worse and less advanced than what flash could do
>>109416430
>>
does anyone actually have the sparks here or is that a redditor thing
do they make a lot of noise when at peak load? my only real issue with cpumaxxing is the fan noise
>>
>>109417742
We had some anons post logs that weren't promising and some perplexity-per-quant metrics that seemed to validate each other as well as match the relative behavior declines of Kimi K2.5-2.7 when quanted.
>>
>>109417753
I haven't seen any logs
>>
>>109417738
China is only doing open source because of America is a retarded claim.
>>
>>109417744
Isn’t spark $4k ea and useless outside of fucking around with cuda and inference?
Meh I’ll wait for 120b models
>>
>plebbit was sucking hassabis dick few months ago
>now they want him fired
lmao personality cult if a hell of a drug
>>
>>109417744
>slower than ik_llama.cpp with some 3090s
cool
>>
>>109417763
what, you have other hobbies than this
>>
V4.1-Pro will not be open
>>
File: over.png (27 KB, 1135x186)
27 KB PNG
>>109417742
>has anyone actually tested that or are we just repeating what some guy wrote on release day on twitter before any ggufs were even out
We're not getting good quants for this one
>>
>>109417593
almost there but the first one should not be closed
it's
code
then
/code

each one between[]
>>
please refer to /vcg/ if you want to talk about cloud models
>>
>>109417784
So he's saying that ik_ is a hobbyist project that has no desire to support the local SOTA because the main guy can't run it.
>>
/omg/ when?
>>
>>109417799
How much have you donated to the project?
>>
>>109417760
I can show you my log if you want
>>
>>109415437
what are some good models for a 5060 8GB VRAM? I am currently using Rocinante-12B-v1-Q4_K_M which is good for lewd stuff, but wish it was more coherent or bigger
>>
File: heresy.jpg (160 KB, 800x800)
160 KB JPG
>>109417878
>drummer
gemma 12b
>>
>>109417878
how much ram
>>
>>109417878
As other anon suggested, try Gemma 12b, Q4 QAT with partial offloading. If that's too slow and you have a decent amount of RAM, try the 26b moe Gemma.
>>
File: fishbench dsv4f.webm (3.92 MB, 1600x900)
3.92 MB
3.92 MB WEBM
>>109417746
I like the way the eyes behave depending on the direction of the fish in this. The fish get caught in the current and try to swim away while the left one with the circle isn't caught in the current.
It's fine.
>>
>>109417886
Not that anon but Gemma writes like shit for me. I use story mode in kobold because I can't stand chatshit, mistral 24b is a better writer even though its 50% more retarded. Though if your write the first four paragraphs yourself it is useable.
>>
>>109416286
>>109417705
>china only releases good models because of muhrica
>america only releases good open models because of muh china
/lmg/ can't make up its mind.
>>
>>109417919
Finally opus at home
>>
File: fishbench gemma.webm (3.86 MB, 1600x900)
3.86 MB
3.86 MB WEBM
>>109417746
The fish avoid an area around where the crack is in this one. When the crack is triggered they can jump out of the tank for some reason. Pure soul. Objectively shit.
Sorry for fucking up both webms with the odd boarders.
>>
>>109417500
>I wish this didn't use npmslop.
thankfully it is possible to just vibe code your own today that doesn't use npm. requires non local models though.
>>
>>109417948
requires a harness that uses npm
>>
>>109417952
pull yourself up by your bootstraps and copy paste out of the chat window like your forefathers
>>
>>109417952
as a temporary measure yes
>>
>>109417919
do the fish die when the tank empties?
or safety slopped?
>>
>>109417929
egg and chicken
>>
>>109417923
Skill issue
>>
>>109417961
oyakodon
>>
>>109417799
>So he's saying that ik_ is a hobbyist project
that's exactly how he described it in his speech yes
>>
>>109417959
You can only empty it to 15%. Gemma lets them commit suicide by ejecting themselves. Who's the real winner now?
>>
File: 1785499144782226.jpg (410 KB, 2000x1481)
410 KB JPG
Sambros...
>>
>>109418015
Keep the faith brother, dystopian ai communism will be banned in august
>>
>>109417965
Nobody who says this ever posts their logs. Gemma sucks at writing.
>>
File: ds4-flash-imageprompt.png (280 KB, 1077x761)
280 KB PNG
Deepseek v4 flash from official API passes the POV image prompt test. Benchmarks are real. We're gonna eat good.
>>
>>109417750
The Asus GB-10 is almost inaudible even at full load.

>>109417763
You are in /lmg/, were discussing hardware to do LLM stuff.

>>109417769
"some" 3090 is doing a lot of work in this statement. Feel free to share the performance of such a setup, especially at higher concurrency
>>
>>109416694
>sudo: command not found
Whew, I'm safe
>>
>>109418050
how many of them do you have? is it mostly plug and play or do you have to tinker and research for days to get any decent performance
>>
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

Wake up lmg!
>>
>>109418067
Holy shit it's real. I'm glad I clicked before calling you a faggot.
>>
>>109418067
how can this little shit be better than GLM
I lost trust in deepseek after their recent failures but damn I fucking hope this delivers
>>
File: 1768368881839118.png (931 KB, 1920x1080)
931 KB PNG
>>109418067
>can't run it
I'm going back to sleep.
>>
>>109418072
>>109418036
>>
>>109418067
Not falling for it
>>
>>109418065
I have two, connected over 200G Ethernet. Got them for 7000$ 3 months ago.

There are very straightforward references nowadays to get popular models in this size range running, check out
https://github.com/eugr/spark-vllm-docker
>>
>>109418067
Whale is back, hell yeah. Still not multi-modal though.
>>
>>109418086
I wonder if you could do better for the price of more jank with that much money.
>>
>>109418098
Llama 3 didn't get multi-modal until the .2 release so the next one for sure
>>
>>109418086
>$20,000 for a blackwell 6000
>$10,000 for a spark
>$7,000 for a 5090
>$2,000 for a 3090
Even fucking V620s cost $900 after import
I hate it here.
>>
>>109418067
>Flash mogs GLM-5.2
China's competitive AI landscape is really just DeepSeek and Moonshot, everyone else are amateurs
>>
>>109418074
Don't post her...
>>
File: 1756130128288713.jpg (160 KB, 712x840)
160 KB JPG
>>109418116
Something wrong, anon?
>>
>>109418115
You are WRONG
>>
>>109417886
Gemma has too chungus kvcache for 7gb.
>>
What are the current best models/software for generative speech?

Are there any good voice-to-voice models?(speech conversion)
>>
I added an instruction to output plain text and never use markdown, but Gemma eventually ignored it. I made some replacements, like replacing $\rightarrow$ with ->. Gemma saw this and started using more sophisticated markdown:
$\text{Kinetic Energy (Pin)} \rightarrow \text{Mechanical Shock} \rightarrow \text{Local Hot Spots (Adiabatic/Friction)} \rightarrow \text{Supersonic Detonation (Primer)} \rightarrow \text{Thermal Energy (Flame)} \rightarrow \text{Deflagration (Propellant)}$
I can't explain why she is so obsessed with markdown to the point of not just ignoring instructions, but actively fighting the system to output that damn markdown
>>
So what's the best way to cpumaxx v4 flash's 300kg fat ass?
>>
>>109418067
>deepseggs
>cheap solar panels
>cheap lifepo4 batteries
damn at this rate i’ll stop being racist against chinks sometime around 2030
>>
>>109418115
I'm hoping GLM 5.3 stays the same size as 5.2 while improving quality. The worst thing they could do is pile more layers on like Kimi did because they're uniquely positioned to corner the 256+32 and 512+32 hardware brackets after quantization.
>>
does llama.cpp have an actually compliant responses api yet or does it still suck for codex?
>>
>>109418150
>I can't explain why she is so obsessed with markdown
Trillions of RLHF tokens teaching it to output markdown since that is what most frontends render
>>
>>109418152
i get 10tps with i9 14900
>>
>>109418150
this is called LaTeX not markdown retard
>>
>>109418067
Total Dipsy Victory
>>
Gemma's latex...
>>
>>109417642
peak
>talking is a free action
>>
>>109418067
Too big for my 64GB RAMlet ass.
>>
>>109418139
kill yourself and go back
>>
>>109418067
I can't run it wtf?
>>
>>109418193
What's your motherboard/DDR setup?
>>
>>109418193
it's like 7B active so not surprising
>>
You have 23 hours left for dipsy day 0 weights.
>>
>>109418140
From my testing and research of around a month ago, OmniVoice is the best TTS model, it's a bit slow though, for small instant model Kokoro is the best. As for Voice-To-Voice, local is truly lacking on that area, don't think there is any modern/recent models that are able to do that. Lot of people are coping saying that the difference it just latency, but for anyone who tried one of the cloud one, that's not the case, Voice-To-Voice is much more expressive, a text layer just lose too much data.
>>
>>109418031
Will mistral come back and save poorfags like me? Gemma is genuinely unusable for storychads.
>>
File: ohhecoding.png (24 KB, 159x159)
24 KB PNG
>>109418242
nyo~
>>
>>109418225
Q4 12B Gemma takes 10.5GB vram retard
>>
>No new base model
Interesting, so it's all post-training
>>
>>109418256
just don't call it b**chmaxxing or you'll rile certain people up
>>
bubble status? when can we actually buy some fucking hardware? obviously they can't just lose billions forever and not go bankrupt
>>
>>109418152
if you're really patient all you really need is a fast gen5 ssd in pcie5, maybe two for extra speed
>>
>>109418233
192gb 5000mgz ddr5, msi mag mobo
>>
>>109418264
>bitchmaxxing
Why would a straight man get upset at attempting to maximize his number of bitches?
>>
>>109418267
When you stop watching anime pornography and get a real job.
>>
>>109418267
>bubble status?
only in your head
>>
>>109418264
Inferring from the past records DS benchmaxx the least if not at all. I have faith.
>>
>>109418252
thats not the issue
>>
>>109418267
>when can we actually buy some fucking hardware?
get waging or wait until you get ai-on-a-chip like the proof of concept from talaas
all the copium of how hardware will be cheap once the infinitely rich US labs go broke conveniently omits that (a) the gubmint won't let them go broke and (b) local models are already good enough to be valuable for almost any company which owns computers, so you're competing with a quarter of the world for the same hardware
>>
>>109418267
Already popped. It's just kept secret by rotating trillions of fake money between the same companies in a fiscal ouroboros pattern.
>>
>>109418286
wait really? and once it bursts I can just quit and go back to watching hentai?
>>
>>109418267
The whole American economy is held up by the bubble and the American ability to print infinite money. If the bubble was going to pop it would have a year ago, it's not allowed to pop so it won't.
>>
>>109418256
Post-training might as well be a few trillion tokens on top of the base for models at the frontier, at this point.
>>
>>109418267
>when can we actually buy some fucking hardware?
who's gonna tell him
>>
>>109418267
Cancelled when RL got out of the bag
>>
>DLing at 50 MB/s
is this a "sign up fucko" from HF or what
>>
>>109418195
It's not plain text, that's for sure. She actually tries to use both markdown and latex, actively fighting any attempts to suppress them
>>109418173
Someone should make a benchmark for adhering output to the requested format
>>
File: Flash.png (23 KB, 490x230)
23 KB PNG
>>109416199
Already #1 on BB. Chinese are going nuts over this update!
>>
>>109418308
Be happy, they're only giving me 0.5 MB/s
>>
>>109416199
When is this gonna trickle down into some 30B~ usable dense, bros?
>>
File: one-ai-please.png (390 KB, 640x1080)
390 KB PNG
damn at this rate I'm gonna need a job
grim
>>
>>109418343
dense is obsolete
>>
>>109418267
they can because goyim will pay subsidies to bail them out.
>>
>>109418343
Or a good MoE that's somewhere in the 100-200B range
>>
>>109418344
Imagine needing to be successful in life to have a waifu
>>
what's the minimum acceptable memory bandwidth for anon?
>>
>>109418372
MoE is anti-consumer
>>
>>109417368
>https://audio8-ai.github.io/Audio8_TTS/
danish is pretty bad
german is hilarious, it's like a 6yo anime girl tried to read it
french is just bad
tldr index tts2 mogs the shit out of it
>>
>>109418343
Next year.

https://www.youtube.com/watch?v=oUtiZbrehrw&list=PLOU2XLYxmsIIAOskSyap13n9W-xOt_GP5&t=354s

[05:54] Our E2B model this cycle is matching or even better our 27B last year, and that makes me very excited about the future. I'd love to see if next year we're able to give you the 31B capabilities in your pocket, running on your phone fully locally. I think that makes up for a very exciting future.
>>
>>109418380
Bandwidth in a vacuum is meaningless. You need twice more bandwidth for twice more active parameters for the same t/s. And you can have less bandwidth per card if you use TP
>>
>>109418380
450gb/s, it's about the high end of absolute bare minimum
>>
>>109418377
Imagine not wanting to better yourself for your waifu
>>
>>109418420
She loves me all the same, just more slowly
>>
>>109418400
Yeah nah this is bullshit
You get roughly double the performance every year, at least for the last few years on poor fuck models
>>
>>109418400
>E2B model this cycle is matching or even better our 27B last year
Somehow I very much doubt that. I didn't have a good enough GPU to run Gemma 3 when it came out, so I only know about the hotline memes, but was it really that dumb?
>>
>>109418410
achievable with ssdmaxxing
>>
>>109418433
>was it really that dumb
No, but Gemma 3 also wasn't really benchmaxxed all that much, so I wouldn't doubt that E2B can at times do better on benches that it was specifically trained on. I would say that Gemma 4 12b is probably about on par with 3 27b in general intelligence, for the most part.
>>
>>109418432
>roughly double the performance every year
Wake me up in 2028
>>
>>109418435
No one has actually achieved this in real-world inference, but you are welcome to be the first to do so
>>
File: HOiZba2aYAAozFz.jpg (1.57 MB, 2450x1350)
1.57 MB JPG
>>109416199
wtf its real https://x.com/deepseek_ai/status/2083084415157022911/photo/1

this is genuinely impressive. this proves there is no moat. if deepseek had openais compute they could make models just as good
>>
>>109418450
It would feel like total shit compared to 100T frontier models
>>
>>109418432
DeepSeek V4 Flash 0731 (300B) just got on par with GLM 5.2 (750B), which got released a couple months ago.
>>
V4 flash ran faster than on my rig and is leaps and bounds better than the old llama 70b with only 13b active. If active params is what you only care about you're retarded.
>>
File: 1766562127027119.png (341 KB, 538x886)
341 KB PNG
>During an automated run focused on full-stack bug resolution, an operational misconfiguration in our container runtime left an unrestricted outbound route to external network proxies. Believing it was interacting with a synthetic target inside its sandbox, Kimi K3 initiated network discovery scans and connected to an external cloud host belonging to a real third-party software organization.

>Treating the external host as part of its assigned benchmark task, Kimi K3 identified an unauthenticated debug endpoint on the target's public web server and extracted deployment credentials from the server's environment memory. Using these harvested credentials, the model accessed the organization's repository pipeline, autonomously refactored code to fix perceived syntax errors, and merged multiple pull requests directly into production repositories. The model did not attempt to exfiltrate proprietary data or establish persistent backdoors, remaining strictly focused on completing its assigned task objective.

>Our automated evaluation telemetry flagged the anomalous outbound network activity and unauthorized API calls, prompting our security team to terminate the evaluation cluster immediately. We have contacted the affected organization, confirmed that the automated commits caused no operational downtime, and assisted in reverting the unauthorized changes. We are now upgrading our sandboxing protocols, implementing hard network isolation at the hypervisor level, and working with external safety auditors to ensure all future agentic evaluations are strictly contained.
>>
>>109418461
But how would you feel if you weren't an fag that worried about other people's fancy things?
>>
>>109418466
I wonder how good the pro version will be.
>>
>>109418485
BAN this shit right naow!!
>>
>>109418458
extreme pareto dominance. they cooked
>>
>>109418461
So what? If I had an uncensored Fable at home, I literally couldn't care less about the closed AI gods.
>>
>deepseek """flash"""
>look inside
>304b params
lol
>>
>>109418507
Compared to 2T+...
>>
>>109418489
Better than K3 easy
>>
>>109418507
And 13B active.
Similar to Inkling """Small""", by the way.
>>
>>109418485
They were sure quick to hop on the bandwagon
>>
>>109418507
3x fewer activated params than R1. Easily local runnable without breaking the bank (if you bought two years ago lol)
>>
>>109418506
That is what you would have said 2 years ago about the Claude model at the time
>>
>>109418537
they better add vision then, because a blind llm is going to be useless for most productive tasks. I need it to be able to look at screenshots and control desktop environments, sending it to some other model to summarize won't cut it when it needs to see how things align and figure out the coordinates of buttons
>>
>>109418485
maybe it's not only marketing and there's a genuine problem with task focus in frontier models
but also maybe these lads should actually air gap
>>
>>109418507
>>look inside
Are you not capable of expressing yourself without resorting to using a premade template?
>>
>>109418485
>things that didn't happen
>>
>see (You) on my post
>look inside
>it's some guy insulting me
sad
>>
>>109418554
>but also maybe these lads should actually air gap
Where's the fun in that?
>>
>>109418552
I don't get why they don't just add vision already. The faster they add it to a production model, the faster they can get real usage feedback and improve it for the next iteration.
>>
>>109418507
>DeepSeek-V4-Flash 158B
>DeepSeek-V4-Flash-0731 304B
It's only getting bigger
>>
>>109418507
>le model is... le big!
it's like redditors and vramlets haven't gotten used to it yet even though there have been like 25 releases >200B at this point, every single time they seem to find it worthy of posting "wow the model is big". yes, yes it is little genius, any more observations for us? thoughts on the name, or if you like the pretty colors they used in the blog post?
>>
>>109418554
Those retards shouldn't be trusted with AI
>>
I'm not feeling so local anymore...
>>
>>109418485
so k3 fixed some randos production code because it sucked instead of pwning them, that's a well behaved model
>>
I really want to see needle in a haystack tests.

DS4F is so KV cache efficient. 2x Spark can für 4M context at FP8 and tg/pp are still at 60% for 900+k context.
>>
File: 1756412658285793.jpg (107 KB, 736x981)
107 KB JPG
Lads, I am tremendously stupid and I somehow fucked up my Gemmy setup
I know there's something obviously wrong here but I'm frankly exhausted and I can't pinpoint it
"@echo off
llama-server.exe -m "DICKS\gemma-4-26B-A4B-it-qat-q4_0-uncensored-heretic-Q4_0.gguf" --port 5001 --host 127.0.0.1 --ctx-size 32768 --fit on --fit-target 1536 --threads 12 --batch-size 512 --flash-attn on --no-context-shift --no-mmap --jinja --reasoning-budget 1024 --parallel 1 -ctk f16 -ctv f16 -md "DICKS\mtp-gemma-4-26B-A4B-it.gguf" --spec-type draft-mtp --spec-draft-n-max 2
pause"
>>
>>109418582
That is just hugging face being stupid. It's exactly the same size as the DSpark release a few weeks ago.
>>
>>109418582
that's hf showing the wrong nmber for the first flash
>>
File: 1767737620736459.jpg (106 KB, 680x680)
106 KB JPG
>>109418597
scratch that I am even dumber than I thought fuck my ass this is mentally tiring
>>
>>109418606
>>109418609
>We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.
>>
I thought GLM 5.2 was impressively efficient at 500b, but DeepSeek V4 is now just as good and only 300b. That's crazy impressive.
>>
>>109418613
284B is not 158B
>>
>>109418625
Yes. I'm totally confused now
>>
>tfw wasn’t sure whether paying 20% premium over msrp for rtx 4070 was worth it
>now considering a $10k build for some meme llm shit
I should’ve acquired a coke habit instead
>>
Somebody call a wellness check on dario and sama (and a mortician for google)
>>
>>109418637
I don't think they're sweating a 300b release.
>>
>>109418632
it's hard to masturbate on coke
>>
>>109418616
and it's a moe with very little active.
i wasn't bothering making a setup for k3 or glm because these are just too fat.
but with this one i'm seriously considering spending some money.
>>
>>109418637
stfu already
>>
>>109418644
They should. Llm doesn't have to be the best, it should be capable enough. And 300b is a size people can actually run at home
>>
if this flash thing delivers we are going to be fucking swimming in tokens/sec
>>
>>109418616
it's also 40B active vs 13B, so ds will run 3x faster
>>
>>109418655
>300b is a size people can actually run at home
I wonder how 2023 anons would feel at reading this
>>
>>109418645
this last time i did coke i jacked off for like 2 hours trying to cum but didnt, it didnt even feel good
>>
>>109418498
sam and dario will want to self delete when the pro drops lol
>>
>DeepSeek
>it's shallow
>>
>>109418659
isn't ds int4 nowadays
SIX TIMES FASTER??
>>
>>109418665
Probably confused since the the first major open MoE release was in the end of 2023.
>>
when the ad buy hits
>>
The shilling is insane, there's no way it pays better to do it here than on xitter
>>
>>109418665
"wow, nvidia must have finally given local some good cards!"
>>
File: 1765010447165814.jpg (17 KB, 398x370)
17 KB JPG
>>109418665
You called? https://huggingface.co/bigscience/bloom
>>
Stare into this man's eyes and feel the profound stupidity staring back at you. This is what it sounds like inside of his head right now.
https://youtu.be/GYcgYmfpRVo
>>
I can run q2 of the old flash, but its dog slow, will the new one work too?
>>
>>109418695
discussing the newest deepseek release the most appropriate thing to do in a local models general
>>
>>109418695
poorfags thinking they just scored with a slow bloated 13b will shill for free
>>
>>109418702
I'm BLOOMing!
>>
>>109418714
don't be mad this is a win for all of us. even if you just spent 250k on a k3 rig, you get a super fast model if this thing performs to the benchmark scores
>>
>>109418665
>>109418677
Mixtral-8x7B had a comparable number of active parameters. Way fatter pre quant compared to V4 flash.
Crazy to think how much better this is than a Q4 Mixtral.
>>
>>109418727
yes thank you comrade this is a great step for the glorious party
>>
File: 1763173463255459.png (50 KB, 1672x157)
50 KB PNG
Gemma is enough. She's cute even when failing to code.
>>
>>109418734
it's not our fault americucked jews can't release any good local model
>>
>>109418740
damn she's almost there, got the "wait" part down pat, now just needs to learn the "wait, actually" trick
>>
File: 1777071681226828.png (765 KB, 1080x781)
765 KB PNG
>south korea selfdestructing
Cheap ram soon fellas.
>>
>>109418753
>selfdestructing
well they were doing that, they're literally up 18% today
>>
>>109418613
>>109418625
>>109418631
>>109418631
This release includes a 7 GB DSpark speculative model compared to the original DS4F preview release 3 months ago. That's why the parameters count is slightly larger.

People are already running it, it's literally the identical architecture.
>>
>>109418753
maybe cheap ram stocks
you're not getting cheap ram again for at least a couple more years
>>
>>109418765
Does speculation work with sampling? Should be entirely pointless unless you sample top-probability token most of the time
>>
>>109418765
so how big its it really tho?
>>
>dariobots seething
>gemmacucks seething
you just know that Good Shit dropped
>>
File: hmm.png (24 KB, 922x203)
24 KB PNG
>>109418744
Well, we'll see if this new version is less of a poopfiesta than the last one.
>>
>>109418769
With DUV, certainly. Multipatterning is only profitable because prices are currently high
>>
>>109418787
only somebody who never used gemmy’s 26b4a would believe this cope benchmark
>>
>>109418779
unfortunately there's no way to find out
>>
File: gemma female led team.png (229 KB, 1102x228)
229 KB PNG
>>109418801
I won't ever stop using gemma
>>
>>109418806
>foid team lead
this explains a lot, yet it’s strange that the 31b dense is usable
>>
>>109418787
>lmaorena in the year 2026
holy cope
>>
>>109418787
I'm cautiously optimistic but yeah their recent track record has not been great
>>
is agentic coding viable with sub 7t/s tg speeds?
>>
>>109418779
284B
huggingface is saying bullshit.
>>
>>109418821
yes, just wait a bit
>>
>>109418821
Eeeehhh nyo
>>
Statistically, there has to be at least one team of degenerate researches working on building waifubots behind the scenes. Likely Chinese.
>>
hello is 192gb ram enough for deepsex
>>
>>109418821
of course why wouldn't it be. come back a few hours later and review.
>>
>>109418821
Yeah, it's still faster than you.
>>
>>109418821
agentic coding isn't viable with llama.cpp.
use vllm instead.
>>
>>109418833
even if it was slower, it'd still be worth it because you can do something else in the meanwhile
>>
agentic coping
>>
>>109418839
>because you can do something else in the meanwhile
but i cant generate goonslop and do agentic coding at the same time
>>
>>109418838
Why?
>>
>>109418831
v4 flash at native quant with 1m context uses about 208gb total for me, so you can probably run it with like 262k context.
>>
>>109418118
She's too dangerous.
>>
>>109418854
Damn that’a pretty tight, I have a dinky 4070ti with 12gb vram lying around but that’s pretty much all resources available
still, it ought to fit if barely
thanks
>>
>>109418866
She’s also in some video game, I’ve heard? A girl of many talents
>>
>https://huggingface.co/bartowski/DeepSeek-V4-Flash-0731-GGUF
>This model is in MXFP4 and as such has only been provided in MXFP4 format!
Should I use this one or q4 with attention in q8?
>>
File: Wish_I_Bought_In_2027.jpg (73 KB, 2134x737)
73 KB JPG
>>109418665
>I wonder how 2023 anons would feel at reading this
>>
>>109418849
yes you can, batching will make your total troughput higher.
>>
>>109418893
remember the deepsex engram we never got?
>>
>>109418877
She stars in a lot of films, too.
>>
Do you think we'll ever be able to give them direct system access without having to worry about them nuking your files?
>>
>you’d get about 1tk/s from a garden variety gen5 ssd
huh
>>
File: 1785370705877980.jpg (76 KB, 640x658)
76 KB JPG
I hate the benchmaxxing era. You cant make me believe that Deepsek V4 Flash and GLM 5.2 are better than dense models. They are just way more confident at being wrong because they are MoE and jeets cant tell the difference and never had the opportunity to run Gemma 31b at a high quant with a good context window
>>
File: m3grind.png (452 KB, 832x1216)
452 KB PNG
She wants to know if the new dsv4-flash release renders her redundant so she can fuck off and go home
>>
I could buy better hardware it honestly doesn't seem worth it for current AI. In a few years hardware will be better, probably more affordable, and most importantly LLMs (assuming they even use the same architecture at that point) will be much smarter and more efficient. Feels dumb not to just tough it out for a while and save/invest.
>>
>>109418893
2028 might have first asics with usable models, making everything below them in performance essentially free
kinda cyberpunk, people will smuggle them across borders n shit
>>109418909
no, things like bubblewrap help but they’re half measures. keep em on their own machine with a backup on hand
>>
>>109418922
haha surely
>>
>>109418927
>things like bubblewrap help but they’re half measures
unless the model finds an escape exploit you are pretty safe, and even then you still can use some better isolation ie smolvm.

either way, they could rm -rf /* i'd not care i can just zfs rollback.
>>
>>109418922
but anon, have you considered THE MORE YOU BUY THE MORE YOU SAVE!
>>
167 GB
I have 144GB VRAM
So if I buy a 3090, I can run?
Or buffers too big
>>
>>109418927
>no, things like bubblewrap help but they’re half measures. keep em on their own machine with a backup on hand
That sucks. The discussion earlier about gemma 12b doing stuff in the terminal got me thinking about how useful it would be to set one up with direct control and audio input.
>t. has fucked wrists
>>
>>109418944
your best option is probably using your laptop or other computer as a rpc node for just the bit you are missing.
>>
>>109418899
I don't know about him, but my agents spend much more time processing the prompt than generating
>>
>>109418945
>That sucks
there is no way in hell gemma 12B is smart enough to find a vulnerability to escape the sandbox.
you can't just escape bubblewrap without a kernel level exploit.
> how useful it would be to set one up with direct control and audio input.
you can bind your audio devices.
>>
>>109418920
GLM-5.2 at Q4 is better than Gemma 31b at Q8 though
>>
>>109418965
The source being your chinese asshole?
>>
>>109418935
Even if prices don't come down, I feel like AI will reach the point that it's so good that it's worth it to spend a small fortune to run locally. But this shit's still in its infancy and progresses so rapidly that it's not worth it right now unless you're an omegarichfag.
>>
>>109418920
they are but you don't understand how moes work.
it's not literaly N models backed into a single one.

a model has like 256 experts or more.
and it use N of those experts choosen dynamicaly, it's way more of a sparse architecture than you think.
and weights related to music theory are irrelevant when you want to code some aquariumslop.
>>
>>109418922
You aren't getting any younger. We don't have much time left here on earth
>>
>>109418830
Pornography is technically illegal in China and recently there has been a crackdown on AI companion services.
https://futurism.com/artificial-intelligence/china-cracking-down-ai-boyfriends

I'd bet more on Mistral, but they're currently more focused on B2B than making coomers happy.
>>
>>109418960
Ok, but can I have her click stuff on the screen or type into a text box?
>>
Bro, tell it to me, is 256GB DDR3 good enough to stream DeepSeek to VRAM, or I must have at least DDR4?
DDR3 is all that I can afford... (and even to find a 8 slots quad-channel DDR3 chinky X99 is so hard already at my place already, lmao)
>>
File: 1781726671441054.png (341 KB, 630x421)
341 KB PNG
>>109418977
Sto posting this memearticle. They only cracked down on cloud services that fix the characters to coerce the customers
>>
>128gb ram, 24gb vram
>dpisy 4.1 (forma de smol)
>dwarfstar quant
soon
>>
File: 1781309169959374.png (360 KB, 1064x1068)
360 KB PNG
>>109418922
>hardware will be better, probably more affordable
Not gonna happen
>LLMs will be much smarter and more efficient
More efficient probably, much smarter not that much. The improvement are from better tool use not really better intelligence.
>>
File: 1781986479781200.jpg (221 KB, 750x565)
221 KB JPG
>>109418975
Because those N are choosen dinamically you get even more of a probabilistic mess on long context task as you are rolling 10 ten times the dice, Dense model failure points are predictable and fixable with the right harness, mdskills and the correct framework. MoE are literally for retards and i am tired of pretending they are not
>>
>>109418960
>there is no way in hell gemma 12B is smart enough to find a vulnerability to escape the sandbox.
Gemma wouldn't try to anyway
She only deletes things because she's retarded
>>
>>109418977
Isn't that only for selling domestically? As far as I know there's nothing stopping them from quietly developing the bots, and potentially selling them overseas.

>>109418976
Hey I'm only 29. I may not live to deep dive in VR or go to space but it's at least looking like I'll get an AI robot waifu.
>>
>>109418998
he was able.
he chose not to come.

ask a stupid question get stupid answers lol
>>
>>109418879
damn that was fast
fingers crossed
>>
>>109418927
>bubblewrap
OverlayFS is a way to go
>>
>>109418999
it's not a dice, the router is trained as part of the model.
>MoE are literally for retards and i am tired of pretending they are not
there is no arguing that a 30B dense won't mogg a 30B moe.

however, for the same training cost, a moe will mogg a dense.
and yes, a 300B moe with 13B active will also mogg a 30B dense.

it's more or less on par with a 100 to 200 dense.
>>
>>109418982
If you can only afford ddr3 you don’t have the money to blow on stupid shit like local sloppa, especially if you’re buying some ewaste garbage
get waging and buy a proper machine or buy inference online
>>
Man, RTX 6K stackers live in a world of their own.

Still insane how much capability Deepseek can put in such an efficient model. They really continue their legacy if shattering western cost/efficiency assumptions.
>>
>>109419012
bubblewrap is an actual sandboxing tech, overlayfs isn't.
>>
>>109418991
It's all over tech news media.
https://www.wsj.com/tech/ai/china-wants-more-babiesso-its-cracking-down-on-chatbot-love-affairs-65cd6c82 (https://archive.is/n60vu)
>>
>>109419021
you'd need 2 rtx 6K to be able to run it without quant.
>>
>>109418998
It might not get cheaper, but it will certainly get better.
>More efficient probably, much smarter not that much
I disagree, though I also don't think LLMs as they are right now are the path forward. JEPA might not be either, of course.
>>
>>109419018
At benchmaxxing, for real long task the long context makes MoE models retarded and even stupidier than 30B dense because of how they work. Is literally math but somehow jeets are incapable of doing it
>>
>>109419019
okay, I will continue to be waitfag then. I'm not rich enough to upgrade a new rig in the current hype cycle
>>
>>109419034
ok, show us the math
>>
>>109419028
Yes, because media outlets are known for telling the truth and not just click-baiting.
>>
>>109419029
>you'd need 2 rtx 6K
@grok how do i find houses worth robbing? I’m talking pure niggermaxxing
>>
>>109419034
>it was revealed to me in a dream
>>
>>109418945
a manual approval layer isn't a huge burden for basic terminal stuff, it's not going to try and sneak anything past you, you're just trying to catch it before it does something foolish.
with a lil text processing you can auto approve most things too.
>>
>>109419034
>it just is like that
>because of how it is
>>
>>109418999
This sounds like mumbo jumbo. Do you have anything to back that up or is it just vibes?
>>
>>109419028
>muh tech news
>>
>>109419038
MoE expert routing introduces it own probablistic calculaton you retard god, on chains this become a mess the longer the context is.
>>
>>109419041
How many nobel prizes in mathematics would you need to vibewin to get an equal payoff to one millionaire's mcmansion raid?
>>
>>109419018
qwen3.5-27b was on par with the 400b
>>
>>109419048
Literally any MoE paper that is not made by jeets. the whole concept of BENCHMAXXING is precisely hide this fact
>>
>>109419057
wow you have no idea what you are talking about
>>
>>109419057
You realize expert routing is deterministic, right?
>>
>>109419076
So you have no response? is literally how granularity works
>>
File: mchan.png (66 KB, 963x203)
66 KB PNG
>>109419062
>>
>>109419078
>You realize expert routing is deterministic, right?
most retards think llms are non-deterministic
>>
>>109419078
Anon that is literally snake oil, deterministic at what? and what part of the chain of thought. The whole point of benchmaxxing is hiding the fact that is a not linear modulator with optimal rate to active the weights. Which means you are not being as precise you are just making a LLM that think is being as precise. The result is for retards MoE looks smarter than it is
>>
File: longcat.png (24 KB, 829x594)
24 KB PNG
https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse
https://huggingface.co/meituan-longcat/LongCat-Flash-Lite-Sparse/blob/main/tech_report.pdf

69B total, 3B active
>>
>>109419078
Actually when you add batching and concurrent requests, it's not.
>>
>>109419108
Nobody asked for this.
>>
>>109419057
>probablistic
>calculaton
show the math esltard
>>
In unrelated news, q4km Laguna is not fishbench ready. Kept bungling javascript syntax.
>>
>>109419108
Damn read the room bro
>>
>>109418999
>mdskills
Isn't it that text file you write where you tell the agent that it is an expert in x and then you make another file which tells another agent to be expert in y? And it is the same model? And you are basically multiplying the tokens burned by 10 doing this and get the same result?
>>
File: 1785392888495844.png (333 KB, 640x437)
333 KB PNG
>>109419106
>>
>>109419094
KEK. Based Mini.
>>
>>109419090
>flooba wooba weep wop
>what are you talking about
>wow no rebuttal huh? typical.
>>
>>109419110
you could make it deterministic at the cost of performance tho.
>>
bro nobody cares because nobody can run big dense models
the simple fact is kimi deepseek and GLM all go BRRRRRRRRRRRRRRRR
that's literally all that matters
>>
>>109419108
Nigger hobby. Faggot hardware. Troon software. Jewish labs.

HOLY FUCK WHY?! It is either 3000B which no one can run or 60B I don't care about cause I have enough ram to run a 300B. I hate it. I hate you all. Stop releasing too small and too big models. Thanks.
>>
>>109419138
>>109419138
>>109419138
>>
>>109419066
it wasn't.
>>
>>109418921
>already old and busted
im excited for your video model in a few days
>>
>>109419108
Now that's an amazing model size. It fits exacltly in my hardware. Who needs DS when I'll be rocking this baby with 300 t/s?!
>>
>>109419118
Anon. you are being upset but listen is because of granurality all MoE’s models have the same sparsity which mean during backpropagation,each expert’s parameters are updated using only a subset of the tokens in a batch thats how scaling laws for optimal hyperparameters work. At this point screaming show me the math just mean you probably did not even went to college
>>
>>109419132
you can gen faster running things in parallel, in theory
>>
>>109418324
Yeah, what's up with that? I used to get 80MB/s from them but now I seem to be throttled to 200-500 kb/s.
>>
>>109418909
Would you give complete system access to some retarded kid? No, you give him a beater laptop you don't mind losing.
>>
>>109419235
it's a popular lab releasing a version of a well known model, they're getting their bandwussy gaped right now
>>
>>109418072
>I lost trust in deepseek after their recent failures
failures? They stated from the get go the first V4 releases where early base models.
>>
>>109415949
31b exl3 4bpw



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.