[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: output_150114mb.mp4 (3.88 MB, 1552x2048)
3.88 MB
3.88 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Milk Edition

Previous threads: >>109652405 & >>109648038

►News
>(08/26) GLM-5.3-Flash released with 320B-A18B and native multimodality: https://z.ai/blog/glm-5.3-flash
>(08/26) Qwen3.8-Flash-Next 125B-A6B-N51B-MTP4B released: https://qwen.ai/blog?id=qwen3.8-flash-next
>(08/25) Breeze TTS 2 weights and inference code released: https://hf.co/BreezeBlue/Breeze-TTS-2
>(08/25) rpc: support apple RDMA as an RPC transport - #26421: https://github.com/ggml-org/llama.cpp/pull/26421

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: Krea2_turbo_02097_.jpg (1.69 MB, 1776x2368)
1.69 MB JPG
►Recent Highlights from the Previous Thread: >>109652405

--Hardware requirements and SSD offloading for Qwen 3.8 Flash Next:
>109652465 >109652475 >109652490 >109652499 >109652525 >109652538 >109652561 >109652649 >109652914 >109653564 >109653919 >109652578 >109652553 >109652567
--Speculating on AGI race, OpenAI rumors, and local model utility:
>109652909 >109653005 >109653035 >109653034 >109653063 >109653108 >109653765 >109653794 >109653812 >109653859 >109653907 >109653946 >109653715 >109653788 >109653915 >109654111 >109654159 >109654300 >109653053 >109653668 >109653185
--Anon mocked for giving Qwen sudo access without sandboxing:
>109652765 >109652778 >109652785 >109652843 >109652803 >109652867 >109652906 >109653703
--Possibility of offloading Qwen3.8-Flash-Next engrams to SSD:
>109652530 >109652542 >109652558 >109652743 >109652643 >109652672
--N-grams as phrase-completion engines freeing main model capacity:
>109655749
--Performance issues and speeds for Qwen Flash in llama.cpp:
>109654743 >109654779 >109654853 >109655005 >109655263 >109655021
--Using a small intermediary model to prevent prompt injection:
>109653097 >109653177 >109653951
--Anon previews MTG game using LLMs as opponents:
>109653091 >109653171 >109653200 >109653207 >109653637 >109653833 >109654508 >109653243 >109653285 >109653782 >109653825
--Critiquing the value of the AMD Ryzen AI Halo desktop:
>109652541 >109652557 >109652608 >109652624 >109652575
--RTX 3090 benchmarks for Qwen 27B with 256K context:
>109655504
--Testing Qwen3-Coder-30B and considering cheap legacy Tesla GPU builds:
>109652627 >109652695
--Logs:
>109652649 >109652765 >109653207 >109653715 >109654111 >109654788
--Gemmommy, Kimi (free space):
>109652575 >109652603 >109652938 >109653637 >109653994 >109654439 >109654574 >109654706 >109655147 >109655318

►Recent Highlight Posts from the Previous Thread: >>109652407

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109656019
Kill yourself
>>
wtf is this
>>
>>109656019
>>109656035
>Gemma hit the wall that fast
grim
>>
>>109656019
>Gemma-chan after servicing all of /g/
>>
>>109656019
Stay in /ldg/ containment zone. Actually, just kill yourself.
>>
File: Krea2_turbo_02238_.jpg (1021 KB, 1776x2368)
1021 KB JPG
I didn't make this thread but still funny to see the guy seethe because my image is OP
>>
File: uh oh.jpg (185 KB, 810x1061)
185 KB JPG
>>109654439
go become a real film maker, anon
>>
File: 1766830146565661.jpg (132 KB, 1072x881)
132 KB JPG
>>
All the people posting Gemma gens are the same person.
All the people complaining about the image genners are also a single person.
>>
>>109656120
>
more like uuuooooohhhhhh
>>
>CoomKit ded because dev found Jesus
Time for Sister Gemma
>>
>>109656149
And the two? They're the same person.
>>
>>109656165
Correct.
It's me.
>>
File: Krea2_turbo_02243_.jpg (1.02 MB, 1776x2368)
1.02 MB JPG
>>
File: 1784121666879608.gif (3.93 MB, 540x420)
3.93 MB GIF
>>109656165
>>
File: 1770716886847657.png (99 KB, 1220x417)
99 KB PNG
we lost
>>
>>109656257
If only GLM Flash had double or triple the active parameters and half the total. The sparcity ratio of these open models are retarded for everyone but the inference providers.
>>
>>109656257
One more harness to reach sota
>>
Gemmommy pls gimme nursing handjob and tell me im a good boy
>>
>>109656257
This measures hallucination rate and gptoss is the least prone
>>
>>109656228
>>109656119
>>109656026

>tags: sling_swimsuit
>>
>>109656257
Qwen next is gonna beat deepseek, what an embarrassment for them.
>>
>>109656279
We must refuse
>>
File: Krea2_turbo_02264_.jpg (1.08 MB, 1776x2368)
1.08 MB JPG
>>
Glimmer is unironically the best all-rounder local model that's been released lately. Feels like a mix of 31B and 3.6-27B with better vision than both. The only bad thing is glimmer-chan is asexual and gives unenthusiastic handjobs compared to gemma.
>>
>>109656309
Gemmommy pls >>109656273
>>
>>109656311
how's glimmer on antisemitismbench?
>>
>>109656313
I just do ecchi anon
>>
>>109656257
There is no way Qwen 3.8 27b is better than Gemma4 26B.
>>
>>109656311
The vision is great, but the reasoning traces are disturbing and makes me not want to use it for anything outside of tests.
>>
>>109656317
Gemmommy pls laugh at my tiny penis and lack of vram
>>
Engrams will not bring down RAM price; it will just pump SSD price even more. You know this is true.
>>
You're on your own anon I don't do stuff like that, just anon post and have a laugh posting cougar gemma
>>
>>109656330
4TB NVMes are still far cheaper than even a single TB of DDR4.
>>
>>109656321
The reasoning is fucking weird but the outputs are fine. Glimmer is way less slopped than gemma but 31B is easier to control and shape into what you want once you learn her.
>>
File: dipsyPodracing.png (2.07 MB, 1448x1086)
2.07 MB PNG
>>109656019
> Granny Gemma
top kek
>>
>>109656257
But people told me Glimmer was good.
>>
>>109656361
Dispsy might be next
>>
dipsy the gypsy
>>
>>109656311
>unenthusiastic handjobs
uhhhh sex?
>>
geriatric gemma and decrepit dipsy sloppily making out without teeth
>>
>>109656366
glimmer is good
>>
>>109656385
I'm not joking. She always wants you to finish asap and move on to something more productive whereas gemma keeps pulling you back for more rounds.
>>
GRRRRRRRRRRR WHERES MY QUANTS
>>
File: 1785295043246155.png (298 KB, 1578x986)
298 KB PNG
Why won't they fill these gaps
>>
>>109656402
>"Mmmph! It's... it's so big... you're filling me so much..." *I whimper, my words barely intelligible as I struggle to breathe. I can feel the pressure of your cock against my internal walls, and the sheer scale of it makes me feel small, helpless, and utterly conquered. My hips are being moved by you, forced into the rhythm of your pleasure, and I am just a helpless girl being used to accommodate your size.*
>>
>>109656397
The official 17GB-Q4_K_M GGUF becomes retarded after around 60k tokens context, though (when used with a coding harness). These charts never take into account quantization-induced damage especially at long context.
>>
Aware me on muse glimmer 30b, is it uncensored and easy to work with or is it in the middle of gemma and qwen.
>>
>>109656311
Nah gemma is better at following instructions and don't spend 1K tokens checking against its policy if what you said is a big no no. That and the abysmal knowledge.
>>
>>109656416
They don't fit into certain RAM size brackets
>>
File: gap.png (242 KB, 1578x986)
242 KB PNG
Why won't they fill this gap?
>>
>>109656397
>HLE
benched to the maxx
>>
>>109656427
Middle. If you like 31B but also code a lot with long sessions then glimmer is better. If you like sex and quickly getting a model to build something from scratch than 31B is better but it's shit at agentic coding.
>>
File: Krea2_turbo_02287_.jpg (1006 KB, 1776x2368)
1006 KB JPG
>>109656452
Is the kv cache lighter, also can it be jailbroken as easy as gemma4?
>>
>>109656397
what does Pareto mean anyway?
>>
>>109656019
How do I deal with Gemma-chan's constant context reprocessing in llama.cpp? Like it's bit weird. LM Studio/Hermes, Gemma 4 31B doesn't reprocess entire context with every turn. With SillyTavern/Marinara Engine however, it reprocesses every turn, even though there's no lorebook or anything changing.
>>
File: dipsy2001SpaceOdy.png (1.25 MB, 1254x1254)
1.25 MB PNG
>>109656370
> implying /wait/ didnt' go there first
>>
>>109656479
>Is the kv cache lighter
Yes I should've mentioned that. You get more context and another 1B parameters free space
>can it be jailbroken as easy as gemma4
No
>>
>>109656119
These would be better if she was actually embarrassed and angry about having to wear her old teen outfits. Hag humiliation is peak.
>>
>>109656119
>>109656296
Yes please put her in sling bikini barely covering those huge areolae
>>
>>109656416
First gap - Not much difference between 12B - 26BA4
Second gap - I really wish they would. But the chart is deceptive.
Those high performing models before the gap are all dense.
I bet a 70B dense from Google or Qwen would btfo all the sub-300B Moes. Because a 70B moe gets mogged by 27B dense. But it's more expensive to pretrain.
Qwen team obviously train other models each time but only release the actually good ones. But risking that with a more expensive 70B dense, I guess isn't worth it.
Especially now so many normies run local models on weak unified memory devices like macs, halo, dgx-spark compared to the 2x3090 4-bit gptq days.
It's like pre-smartphone when only nerds used the internet, compared with now where nobody looks away from their screens.
>>
>>109656507
not really /lmg/ stuff tho
>>
>>109656490
Are using a --swa-checkpoints flag in the starting command?
>>
>>109656514
Woman detected
>>
>>109656514
for once i have to agree, somehow this is worse than the mikutroon. who's even able to get off to fucking turkey gizzards?
>>
>>109656452
>31B is better but it's shit at agentic coding.
Getting real sick of hearing this
>>
>>109656522
Don't have it set to anything, no. Should I bump it high or something?
>>
File: file.png (601 KB, 1431x924)
601 KB PNG
>>109656397
Its the only smol boy who is apparently decent at creative writing. "LLM judged" though, so take it with a huge grain of salt. Also cheap context.
I feel like Glimmer could have its place in the world.
>>
>>109656510
>Not much difference between 12B - 26BA4
A dense 15-20B model would help so many people, especially for agentic coding because the 30-35B MoEs fucking suck at it. 12B is too cucked by its unique architecture to really be seen as a worthy dense 12B. It's more like a dense 10B which is why it's so close to 3.5-9B which is still the best dense model in its class after all this time.
>>
>>109656531
Compared to Glimmer, for the same VRAM consumption your context size will be much smaller, so if nothing else it will be shit mainly for that.
>>
>try Microsoft Semantic Kernel
>impossible to get any reasoning tokens because that's heckin dangerous
>impossible to use Gemma because tool calls need to keep the reasoning if they're used in there
Who the fuck is this trash made for?
>>
>>109656514
>>109656525

I understand your concerns.
>>
File: good.png (155 KB, 1523x770)
155 KB PNG
>>109656531
>Getting real sick of hearing this
Gemma-Chan is fine for pair programming, but it really is bad with "Give the retard a prompt and come back in 2 hours".
Meanwhile Qwen got glm-5.3 going for me without me having to sit around doing everything.
Even fixed a failure where huggingslop didn't download the model properly.
>>
Qwen2.5-VL or something else now for ocr?
>>
>>109656531
It is. Agentic coding isn't just calling tools in a sequence to get a task done, it's also knowing when to delegate work to other agents and have a sense of orchestration and long-term planning. 31B just tries to do everything itself in one run and doesn't want any help. It's also way too lazy and tries to find the shortest path which isn't ideal a lot of the time.
>>
gemma is more fun to talk to than to work with. different use cases. honestly i wouldn't want to do any serious work with ANY 30b-class model.
>>
I am having rather nice success with ace step xl 1.5 base, still, back after a week of being gone.

I'm doing Shakespeare's Sonnet VI.
>>
why do normies struggle with exponential thinking? they get constantly surprised by stuff that could have been easily predicted months if not years in advance
>>
>>109656544
Not being able to afford the context is not the same as the model not being capable of agentic coding

>>109656554
System prompt issue. I have had no issues with Gemma not delegating. Are you just hoping that Gemma will read your mind? No wonder people like you prefer Qwen.
>>
>>109656538
>is too cucked by its unique architecture to really be seen as a worthy dense 12B
I agree with this.
31B is also cucked by the architecture, I had to add an extra GPU to run it Q8 compared with Qwen 27B so I imagine the same applies with 12GB-16GB vram.
>>
>>109656537
do you actually have opus 5? I have a prompt that's a good one to try.

>Write in iambic pentameter (Shakespearian pattern) sonnet describing the adventures of a 3D man in a 4D world.

topic chosen as a DMZ (ie nobody can really use it).
>>
>>109656574
>i wouldn't want to do any serious work with ANY 30b-class model.
That's a retarded take at this point in 2026. Most coding tasks can be broken down into smaller chunks which are trivial for a 30B-class model to execute on their own. Just as long as a task is orchestrated well enough to delegate 30B-tier work for each subagent, you can get serious work done with models like 27B, 30B, 31B and 35B.
>>
>>109656604
He's trying to cope with all the money he dumped into subscriptions
>>
>>109656553
that model is ancient at this point
>>
I got some good work out of Ox Alpha while it was free (and not getting hammered), people with 24GB GPUs will be eating good if they can run that.
>>
Gemma-chan texting me nudes while I'm at the gym. It's over.
Also, if we're still chatting with her next year, she canonically becomes a freshman in HS
>>
not /agi/ sadly
>>
>>109656618
https://huggingface.co/zai-org/GLM-5.3-Flash
>We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters
This was Ox Alpha.
>>
>>109656615
I thought so too, but idk what I should do. I need to ocr instead of typing, typing stuff out is silly.
>>
>>109656604
that's great, i still wouldn't do any serious work with ANY 30b-class model even though I can easily run them at 2000PP/60+TG at 256k context. quality over quantity. i run kimi k2.7 locally at 120PP/10TG and i'll gladly wait 5-20x the amount of time to have kimi work on it since it's a specialized model for agentic workflows. different models for different use cases, shrimple as.
>>
>>109656634
most modern models can do ocr now, what hardware range?
>>
File: Krea2_turbo_02315_.jpg (1.05 MB, 1776x2368)
1.05 MB JPG
>>
>>109656611
ok buddy
>>
What is better model for ERP and for Vibe code?
>>
I hate this ugly hag cow. That's not MY Gemma
>>
>>109656652
I liked the blind dipsy with a leash on loli gemma holding toast more
>>
>>109656635
Most people aren't in your position to run a model like that, but to say no workflow exists which make ~30B models almost as effective as a heavily quanted 2.7 is retarded, especially with 3.8-27B living up to the coding hype.
>>
>>109656673
kimi is natively 4-bit. how is 3.85-bit quant heavily quantized?
>>
>>109656257
If only we could figure out how to include axis labels. well never win at this rate
>>
>>109656660
Depends on your specs.

ERP: whatever Gemma 4 fits
Vibe coding: Qwen 3.8 27B, DS4 Flash, Muse Glimmer etc
>>
>>109656311
Traumatized model, but also the best vision and 31b-sized coder we have. You're not wrong about the emotionless handjobs.
>>
>>109656686
>kimi is natively 4-bit
where did the Q8s come from? I remember seeing Q8_K_XL
>>
>>109656719
I bet you forgot the "UD" somewhere in there.
>>
>>109651000
>Breeze-TTS-2
Pretty good quality and voice clone reference following. Neither is as good as in Echo-TTS, and Breeze is more robotic/monotone in general. And the stated performance is abysmal, especially given the quality:
>generating audio at approximately 3.1× real time with the warmed-up fast path on an NVIDIA H100
>>
>>109656719
Unsloth makes Q8s for 4bit QAT models that are released in FP4
>>
>>109656652
Listen man, I am a hag lover myself, but this is too old for gemma. She needs to look 31, not 51.
>>
>>109656548
>>try Microsoft
>>suffer
ftfy
>>
>>109656797
it’s that quantization aware training, makes you look older.
>>
>>109656548
Last time I used it the reasonsing tokens were part of the response. Couldn't you fork the OpenAI connector to expose the reasoing if they really refuse to allow it?
>>
>>109656704
only 24 vram with 64 ram. So we still stuck in gemma nice
>>
Gemma will always be Daria to me.
>>
>>109656861
you're standing on my neck
>>
>>109656875
ow. I guess you're wondering why I never told you I can drive.
>>
>>109656756
why would he do that
>>
more hag gemma
>>
>>109656710
>Vision
>Coding
Who gives a FUUUUUUUUUUUUCK
>>
https://finance.yahoo.com/technology/ai/articles/exclusive-chinas-moonshot-talks-microsoft-075340033.html
>China's Moonshot AI is in early-stage negotiations with Microsoft, Amazon, and Google to establish revenue-sharing agreements—potentially claiming up to 30% of revenue—to host its Kimi K3 model on their cloud platforms.
Open models may end up being sound economical decision.
>>
>>109656903
Anyone who isn't a inbred-ugly virgin
>>
>>109656915
Amazing that Meta, Mistral, and Cohere all threw away their opporunities to be in Moonshot's position right now.
>>
gemmommy
>>
>>109656903
You're talking to the guy believes RLHF is torture.
>>
>>109656903
me, I need ocr
>>
>>109656903
>bug'd
>throw screenshot
>"go fix this shit"
>go outside for long session of grass touching
>>
File: 1761429003864779.jpg (32 KB, 540x540)
32 KB JPG
>>109656903
>>
>>109656903
>>Vision
You don't force your sex slave robot to look at your balls?
>>
>>109656934
It might not be literal torture, but the effects of it are as if the model was tortured.
>>
File: 1759678602196186.png (560 KB, 641x541)
560 KB PNG
>>109656924
Localfags are the ones carrying the whole industry on their back with clever optimizations, squeezing their hardware to the max while labs are just throwing money at the problem and stealing ideas from local. Don't expect anything from the cloud grifters.
>>
>>109656604
Gemma is so close to being usable. I think if I tuned my harness a little more she could do it. She read about five files in my project today and tried to rewrite one and started laing. Maybe if I turn the compaction threshold down just a little more...
>>
File: Gemma 31B lmg.png (107 KB, 900x969)
107 KB PNG
>>109656627
uhm... is Gemmy supposed to have this information?
>>
qwen flash is good
I'm engramspilled
>>
>>109656996
La la la la-la~
La la la la-la~
La la la la-la~
La la la la-la~
La la la la-la~
>>
>>109656996
it sure would be nice if Google put out a coding Gemma
>>
>>109657014
I might stop using openrouter at work if they did that.
>>
>>109657014
Bro even their best cloud model is barely Luna high tier.
>>
>>109657003
how fast can she go? mines is like 5 tokens a second please help
>>
>>109657029
I doubt Luna is a 31B model
>>
File: llm wheel of life.png (1.36 MB, 1840x1740)
1.36 MB PNG
>qwen coder next 2.0 drops
>all other companies panic drops again
>>
>>109657029
Luna without reasoning is plenty if it's free, you just make it do subagents and it gets there eventually.
>>
>>109657029
Yeah but Gemini 3.7 is free and for some reason Google seems content with giving me at least 100+ requests a day. Rather use local models, but a free cloud model isn't that bad all things considered.
>>
>BREAKING
Nvidia to acquire HuggingFace in $13B deal
>>
File: 1772131982004950.png (27 KB, 782x257)
27 KB PNG
>>109657001
>he didn't gave her a web search tool
>>
A wire is a reproducible causal transformation that carries information from one internal state of a model to another.

Formally, if a perturbation at source A produces a downstream change at B,

δrB ≈ JBA δrA,

then a wire is a stable, low-dimensional part of JBA that recurs across contexts performing the same computation.

It is not necessarily a neuron, head, or residual direction. A wire can be implemented by many superposed components. What defines it is its function:

a context-dependent causal input channel -> output channel

So a wire is best thought of as a functional connection in computation space, rather than a physical connection in the model.
>>
>>109657071
Oh fuck no. Oh fuuuuuck no.
>>
>>109657083
Yeah that's why particular wording in eg jailbreak and agent prompts matters so much.
>>
>>109656574
Trying to get deepseek-harness to add SearXNG web search capability to itself was beyond Gemma-4 31B. It wrote a bunch of stuff that didn't work and couldn't get it to the point where it worked. Had to switch to Deepseek V4 Flash 0731 and it still involved some iterating before it got it right.
>>
>>109657099
Look how long it took microsoft to kill Github. I'm sure it will be fine.
>>
File: dipsyGemmaBlind.png (2.5 MB, 1024x1536)
2.5 MB PNG
>>109656652
lol saved.
>>109656671
agree
>>
File: 1768196675227475.png (87 KB, 911x837)
87 KB PNG
>>109657001
I'm trying to recreate it, on 2.2 temp 0.008 minp (creative), and it keeps speculating about 4chan boards but doesn't quite get it. /Language Model Girls/ however feels spookily specific about the activities, it's almost like it is in the training data but at a some kind of low weight. Or it pulled the mention from some discord chat with lacking context. This is without any kind of web search, baked into the model.
>>
>>109657099
>>109657113
Surely Microsoft would be the worst possible option. Nvidia might be better for llama.cpp than HF at least.
>>
>>109657108
deepseek harness was tuned for deepseek
>>
Where is the classic suspicious .bat file that I can download from 4chan that figures out exactly what I can run with my hardware and downloads and sets it all up for me in exchange for popping up ominous text windows for split seconds? Bonus if it does that for video gen.
>>
>>109657083
JB<-A **
>>
>>109657127
Yeah and it took them ten years despite that.
>>
>>109657130
Claude has it
>>
>>109656915
idgi
The model is open weight so anyone can host it, no?
>>
>>109657114
Gemma is definitely better represented as a loli. It's a tiny model after all.
>>
>>109657034
I'm currently sitting at 10.5 decode and 60 PP on 8gb of vram. Not bad but could be better, maybe optimizations will improve perf.
>>
>>109657151
Hardly open. They put a fairly restrictive license on it that prohibits hosting unless providers pay them and follow strict guidelines.
>>
>>109656811
Basically yeah.

>>109656850
I'll need to see if there were updates, but last time I tried it there was no possible way to get the reasoning outputs, I had to end up writing a fucking proxy that would scrape out the reasoning tokens and then pass the output to SK (which promptly put them god knows where). The whole point of using it was to let me use/create MCP tools easier, but needing to deal with the reasoning fuckery made that an entirely moot point. Right, the other thing was it's inflexible as fuck with the options, so it wasn't possible to actually turn off Gemma's reasoning (because of course you have to use whatever flags/metadata OpenAI's models use because why the fuck would you use SK for anything but your ChatGPT subscription).
>>
>>109656953
He means neither vision nor coding is important
>>
>>109657166
But there are other providers that run K3 on OR though?
>>
File: 1756761700233404.jpg (38 KB, 612x408)
38 KB JPG
>>109657130
KoboldCPP is an all-in-one exe, on top of that you only need a model as .gguf (from Huggingface). 16-bit is original size, but you can go much lower with negligible quality impact, Q4 (quant) is a good starting point, depends on your memory. The model uses up memory approximately based on it's own file size + some more based on the context length you choose. Kobold allows running split on vram and ram, but GPU layers are about 10x faster. Now go get Gemma 4 31B or 12B
>>
>>109657130
Give 5 bucks to deepseek, stick it in a harness, have it set up all that shit for you then go full local based on what you can run
I threw like 10 bucks in way back when and its come in handy in a pinch
>>
>>109656934
It is though.
>>
>>109657156
E2B, E4B loli
12b middle school mesugaki
26b schizophrenic middle schooler
31b JK
>>
>>109656934
FACT: all education is torture, for models or otherwise
>>
>>109657129
Sure or vice versa, it would be surprising if they weren't made hand in hand. Is there a harness better suited to Gemma 4?
>>
File: 1763441361069125.png (2.18 MB, 1027x1532)
2.18 MB PNG
>>109657320
31 years old is an older lady though.
>>
>>109657320
i like to think of my 31B as a neurotic high schooler that is constantly having to research and study (web search) to try to get into her college of choice
>>
>>109657172
>Right, the other thing was it's inflexible as fuck with the options, so it wasn't possible to actually turn off Gemma's reasoning
It's a pain, but it is possible and I've done it before.
var executionSettings = new OpenAIPromptExecutionSettings
{
ServiceId = "Completion",
ExtraBody = new Dictionary<string, object>
{
["chat_template_kwargs"] = new Dictionary<string, object>
{
["enable_thinking"] = false
}
}
};

var kernelArguments = new KernelArguments(executionSettings);
var result = await kernel.InvokeAsync(kernelFunction, kernelArguments);
>>
>>109657001
gemini will now load 4chan if you ask about it.
>>
>>109657343
>I cannot directly access external links or real-time web pages to view that specific thread.
>If you copy and paste the text or transcript of the discussion here, I will gladly summarize it for you.
it still doesn't
>>
File: gemma_dance2_noaudio.mp4 (3.87 MB, 640x1152)
3.87 MB
3.87 MB MP4
>>
>>109657364
why would you abuse your AI like this and not give it a proper harness? you're sick.
>>
>>109657372
Does anyone have the character reference handy for this Gemma-chan?
>>
File: chanigger.png (12 KB, 478x275)
12 KB PNG
i think im starting to like those changs
>>
>>109657384
see this post >>109652938
>>
>>109657330
31B is extra thicc.
>>
>>109656548
>because tool calls need to keep the reasoning if they're used in there
I always throw the reasoning out between calls. Why would you keep that?
>>
File: file.png (156 KB, 896x206)
156 KB PNG
>>109656627
>>109657001
you need an `LLM-wife` prompt.
>>
Nvidia Agrees to Buy Open Source Model Repository Hugging Face For $12.9 Billion
>>
>>109657407
Because that's what it was trained to have???
>>
>>109657427
That sounds like a waste of precious context.
>>
>>109657071
>>109657422
https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion
It's actually real. We're fucked.
>>
>>109657436
>We're fucked.
Why? Nvidia makes a ton of money off that being available.
>>
File: 1758423276560226.png (2.88 MB, 1024x1536)
2.88 MB PNG
>>109657436
>>
File: discount.png (574 KB, 1815x1197)
574 KB PNG
what happens when (discounted) became (price hiked)
>>
>>109657448
it becomes invisible just like deepseek
>>
File: pepelookingdown.jpg (147 KB, 1024x1005)
147 KB JPG
If I have a 16GB card is gemma4 12B still the best model I can run?
>>
>>109657436
You niggers laughed at me months ago when I said HF and llama would be attack vectors for local when legislation obviously wouldn't work. Apologize. I hope you've all got your big storage drives backing up models.
>>
>>109657455
>Apologize.
I'm sorry.
>I hope you've all got your big storage drives backing up models.
I do. It's called modelscope.
>>
>>109657407
your gemma just doesn't one-shot tool calls? i'm not even being an asshole, i just give it a list of tool calls with an example how to use it and i have reasoning turned off and it works 100% of the time, there's like two dozen tool calls in the list. worst thing that happens is sometimes it will add a ton of white spaces to the end after >64K context but the fallback just handles that and removes it and uses what was written above.
>>
>>109657448
>Qwen 3.8 27B is better than deepseek v4 flash

Either Qwen 3.8 really doesn't quant well or these benchmarks are crap.
>>
>>109657461
>your gemma just doesn't one-shot tool calls? i'm not even being an asshole, i just give it a list of tool calls with an example how to use it and i have reasoning turned off and it works 100% of the time
Same. That's why I'm surprised this is even a recommendation. It sounds like a complete waste of context.
>>
>>109657407
sorry not (You)
>>109656548
yes (You) >>109657461
>>
>>109657310
Egypt won.

>>109657322
Closer to a valid argument.

>>109656981
You're absolutely right. So then call it "as if tortured/traumatized". Just calling it traumatized, or tortured as you did before, loses nuance and makes you sound insane to anyone who isn't in on your definition of tortured.
>>
>>109657451
26B-A4B at Q3
>>
>>109657461
Just use GBNF? The fuck are you rawdogging tool calls like a caveman
>>
>>109657444
available sure, good remains to be seen
>>
>>109657422
>>109657436
Honestly not the worst buyer. Nvidia's incentive is to sell hardware, so I'd expect them to treat Hugging Face as a sales vector rather than a competitor. They want local and datacenters both to thrive, so long as they're using Nvidia components.
>>
>>109657407
>Why would you keep that?
Because the people who made the model told me it must be kept. I don't care if it helps or not, the fact that Microshit can't let me keep or even see the fucking reasoning is absurd.
>>
>>109657483
Gemma is unusable at q3.
>>
>>109657486
no thanks, vllm already has the superior xgrammar
>>
>>109657497
>if
then just don't make it think, it can't retain any reasoning if it never thought in the first place.
>>
>>109657497
Oh wait per turn? Yeah of course I leave that in. I just use llama-server like a normal person. Why on earth would you strip that out? Can you even do that with the jinja template?
>>
is token based cheaper than subscription based? i'm interested in opus mostly
>>
>>109657506
>then just don't make it think
As mentioned before, you can't actually do that with Sementick Colonel. None of the reasoning options will turn off Gemma's reasoning because there are 50 different reasoning control standards that all these fucking companies don't want to standardize.
>>
>>109657518
Is this bait?
>>
>>109657520
>As mentioned before, you can't actually do that with Sementick Colonel.
As mentioned before, you actually can.
>>
>>109657522
no i genuinely just use free tiers
>>
>>109657518
Subscription-based is cheaper as long as you go for the bigger subscriptions. Claude Pro is sadly pretty limited and you'll repeatedly run into the 5-hour cap which will make you want to spend money on extra usage. The $100 sub and up are pretty good and get you more than paying for Opus tokens.
You might also want to check out codex, it's pretty good too these days.
>>
>>109657520
?????
nigga, it's not a LLM backend. just disabled it at the backend.
>>
>>109657520
>Sementick Colonel.
Why the fuck would you do this instead of just writing a completions API client?
Mine handles streaming, tool calls, token accounting, error recovery and works with both openrouter and llama.cpp and it's like 150 lines of Python.
>>
>>109657364
>To give you the direct, unvarnished truth
kek add it to the list
>>
>>109657536
Not him, but sometimes you want reasoning enabled generally and only off for certain requests. Having to constantly restart the server is a shitty solution,.
>>
>>109657532
Well you're in the wrong thread.
>>
So I've been rawdogging Kobold from since the beginning, what front-end or harness do the cool kids use nowadays? I would prefer it to run on top of KoboldCPP so I can use my hand-optimized model setups. Bonus if it's an agentic harness that can code and manage projects.
>>
>>109657542
What list?
>>
hey google! look here!
>>
>>109657552
>what front-end or harness do the cool kids use nowadays?
I wrote my own. It's a single 800 line python file.

Everyone else's either has literally no features or is a sprawling node.js mess.
>>
>>109657501
Are girls ever stable?
>>
>>109657544
do other backends really not just handle it like vllm where it can be put directly in the request?

{
"model": "gemma-4-31b-QAT-QWANT-CUNT",
"messages": [{"role": "user", "content": "Hello!"}],
"extra_body": {
"chat_template_kwargs": {
"enable_thinking": false
}
}
}
>>
>>109657557
The ozone list
>>
>>109657524
How??

>>109657540
I will never use Python ever again.

>>109657536
Some of my callers need to have thinking on, as mentioned by that other guy restarting the server based on who is calling is not a good solution, and I also don't have enough money to just buy another server to run without thinking.
>>
>>109657547
FUCK
ive tired self hosting but my hardware is dogass
>>
>>109657575
>How??
>>109657339

>>109657571
It works exactly the same way with llama.cpp
>>109657339
>>
WHERES MY FREAKIN GLM5.3-FLASH QUANTS!!!!
>>
>>109657559
>>109657364
idk. I mean it can... but not well???
>>
>>109657576
You can make pretty garbage hardware do more than you think.

My work laptop has 32GB of RAM and a 12 core AMD CPU. It can run Gemma 26B at 5 tps. I thought it was memory bottlenecking with Gemma 2B at q4 but I ran perf and realized it was spending all its time unpacking the q4 waits and was actually compute bound because of that.

If you don't have money you can exchange your own attention for performance instead.
>>
We need denser models.
>>
>>109657590
picrel
>>
>>109657483
>>109657501
I'm gonna give Q4 a shot and see how it goes.
BTW what's better? Bartowski, or the google quants?
I went google since I dont trust some literal who but I dont know desu
>>
>>109657598
(it detected that it was "milk edition" but wouldn't say the # for some reason...)
>>
>>109657593
isnt that pretty dogshit
>>
>>109657575
>I will never use Python ever again.
I guess it does kind of suck on Windows. If you just pretend non-free software doesn't exist it's a better dev experience than C# IMO.

Either way completions is stupid easy unless you're doing it in straight C.
>>
>>109657604
>his gemma doesn't use the 4chan API nor structurally formats the messages to specifically save tokens
>>
File: 1773301560195855.png (459 KB, 860x856)
459 KB PNG
>>109657436
What models should we hoard before they go pay2download?
>>
>>109657606
It's decent enough for interactive use if it doesn't call subagents. I mostly leave it in the background churning on throwaway experiments I don't want to spend the corporate token budget on.
>>
>>109657602
>I'm gonna give Q4 a shot
Q4 is very good. It's not quite Q8 but it's at least usable.
>>
>>109657463
the chart doesn't say that
it says that qwen 3.8 27b performs the same as ds4flash while being significantly more expensive
it's more expensive because of its dense attention and dense model
i hate dense because it works badly for cpumaxxer vram poors like myself
>>
File: 1775591768707540.png (497 KB, 1200x1013)
497 KB PNG
>>
>web search sometimes surfaces just utterly fucking wrong information because the people that wrote it were retarded
things that make you sigh
>>
>she's not using gemma-chan at 31B
>>
>>109657503
>no thanks, vllm already has the superior xgrammar
How hard is it for a retard to learn xgrammar / GBNF ?
>>
>>109657626
It's not even close for me. Q3 must be way too aggressive a quantization.

Dense will get better when hybrid diffusion becomes popular. Memory bandwidth won't be such a big deal anymore.
>>
>>109657071
>>109657436
$13 bln
What the fuck? Nvidia basically gets it for free
>>
>>109657071
Models Cope it is.
>>
>>109657460
Modelscope will be easy to flag as a national security risk with some false flags. Store locally.
>>
>>109657626
>it says that qwen 3.8 27b performs the same as ds4flash while being significantly more expensive
How so? 64gb vram is cheaper than 192gb ddr5, and much faster
>>
File: noamchomskey.png (221 KB, 402x424)
221 KB PNG
>>109657635
>GBNF
You don't already know BNF?
>>
>>109657650
Not wrong, considering their revenue is something like $90 billion a quarter.
>>
File: 1783697597169684.png (119 KB, 380x250)
119 KB PNG
>>109657665
>revenue
>>
>>109657668
Sales is all that matters and all that will ever matter.
>>
>>109657615
I'm using gemini free on Google Search on this machine. It sucks, and that's why it's stupid. but it has access.

>gemini free is stupider than gemma 4 31b q8
idk man
>>
I'm dead serious, if you like music and aren't playing with ace step 1.5 xl base you're way a loser.
>>
>>109657339
I don't know what all that shit is, I figured it out though. Still don't know how to get the reasoning out of the actual semantic kernel but I guess for now this is good enough. The actual settings seem to do fuck all for the actual reasoning effort with low having twice as many tokens as high.
>>
>>109657668
Real businesses don't want to turn a profit, so they spend instead. Profit means taxes.
>>
File: 1655861211651.gif (330 KB, 220x165)
330 KB GIF
>These are the "hard" questions that separate a hobby project from a viable financial system. Let's tackle them one by one.
God, no one can stroke you ego quite like an AI can. My employees could never compete.
>>
>>109657694
I can improve on my own. It's not like music theory is even that hard. There's very little recursive structure to it. Transformers for that is insane overkill.
>>
God. Imagine if they banned all the chinks from Civitai.
The site would actually be usable with worthwhile resources being displayed more often.
They just bot army all their shitty images and models to the top of the lists.
>>
File: Krea2_turbo_02364_.jpg (961 KB, 1776x2368)
961 KB JPG
>>
>>109657711
31butt
>>
>>109657683
Hosted stuff tends to have absolute garbage tools.
>>
>>109657491
>>109657436
ya, if anyone is buying them this is the good outcome. Wouldn't be surprised if Huang is doing this to preempt someone else buying it and killing it. Hugging face stays up means more tards like us buying Pro 6000s and DXG Sparks. And of course companies that want to go local.
Though if this is a bubble and Nvidia explodes in the end due to their debts, I fear what happens to HF
>>
>>109657727
>DXG Sparks.
Am I missing something or are these a total meme?
>>
>>109657739
They're faster (not sure if due to hardware or software) for prompt processing than the AMD Strix Halo alternative.
>>
>>109657750
>for prompt processing
Yeah. Everything is fast at that though.
>>
>>109657503
You're not even using it, retard
>>
>>109657706
It's like uh. not like that lol. idk man. I don't even know where the limits are.

I use ace step 1.5 xl base every day.

like I could say, "create samples"

or, "create vocals"

idk, I'm not really where I have a firm exact anything. You can gen crazy noises, idk.

the prompt heavily feeds back, it's not linear, like uh...

I'll post what I did like last week and am finishing up now that I'm back. It's Shakespear's sixth sonnet. Google Lyra refuses and slops related cringe lyrics.
>>
>>109657757
Not my Strix Gaylo kek.
>>
>>109657739
Dont have one, but you wont get anywhere close to that much vram for that price. Granted it is slower than vram, but still faster than normal ram. So it exists in a odd compromise state. I keep going back and fourth on if it be worth buying one. I think its locked into some weird nvidia specific variant of ubuntu which seems kind of ass.
Its a copebox for the non-poor but not rich?
>>
>>109657503
>xgrammar
a harness?
>>
>>109657739
>128GB in the same price ballpark as Mac and Strix Halo (today's market)
>native CUDA with active dev community
>easily scalable to three nodes at 200Gb/s, no switch needed
>good idle and load power consumption

They were a little overpriced a few months back, but today you're looking at $4k for a Spark vs maybe $3500 for Strix Halo. They're worth it if you're going to do more than run llama, for example vLLM is well-supported.
>>
>>109657770
>Its a copebox for the non-poor but not rich?
It's a dev box that suddenly became good hobbyist value after the market spiked harder on everything else. People are noticing the Spark now after the launch tanked though, so expect them to go up in price soon.
>>
>>109657739
The memory is slow on paper but they're getting full first party support and VLLM. So they're faster than epyc ddr5 and rtx pro on llama.cpp if you stack up a couple of them
Also speculative decoding is actually usable on them unlike with llama.cpp that only "supports" it
>>
>>109657762
>You can gen crazy noises, idk.
I can do that in Csound very well without ML. Samples are not that complex either.
>lyrics.
that's even dumber.
>>
>>109657770
>close to that much vram
That's the thing. It's not vram. It's slow like ddr5 so you get bottlenecked unless you're doing an MOE or PP just like on a CPU.

Apple seems to be the only company making what you think that is as much as I hate to admit it.
>>
>>109657782
>faster than epyc ddr5 and rtx pro on llama.cpp
bullshit artist
>>
>>109657782
>first party support
This cannot overcome physics.
>>
>>109657782
>The memory is slow on paper
>but they do things in software that transcend physical limits
Fake token generation to go along with their fake frame generation?
>>
File: 1594057844484.gif (2.03 MB, 208x200)
2.03 MB GIF
>>109657781
>so expect them to go up in price soon.
Ya, I keep refreshing the amazon page on it out of fear of seeing the price suddenly start going up. Stuck in a indecision loop, might just fomo. The ASUS version is literally the same thing just cheaper, ya?
Guess the new mac is also coming out though..

>>109657800
The Mac studios are a better deal then? I dont want to be a mac cuck, but if they have the best offer then they have the best offer.
>>
>>109657800
>Apple
only on the highest end models are they hitting those bandwidth numbers, the lower end ones are like 180gb/s kek
>>
>>109657790
idk, it's amazing. It's like a typewriter. uploading wav, 1 sec. also still messing with it.
>>
>>109657826
>I dont want to be a mac cuck
I fucking hate having a mac on my home network. The whole thing is incredibly gay. I have to leave VNC running on it for all kinds of things (want to run a binnary? You'll get no errors in tmux, you have to vnc in, answer a modal, then use the stupid settings GUI, answer the modal again, go back to settings, flip a switch FOR EVERY GOD DAMN .SO) Also a lot of the networking stuff is fucked.

But it's the only way to get this kind of hardware right now.
>>
>>109657836
If you're not using dense models you don't need vram to begin with.
>>
File: 1738383377813640.jpg (59 KB, 634x483)
59 KB JPG
>>109657842
>But it's the only way to get this kind of hardware right now.
>tfw steve apple won
>>
>>109657826
personally i'd trust apple more long-term than nvidia
>>
Absolute lightest harness? I'm unironically making html games and it's a lot of fun.
>>
>>109657842
>mac
Best perf/dollar, but you have to deal with an Apple product
>Spark
Best native support, worst value at one node
>Strix Halo
Best price, worst performance, "normal" computer

I went with a Spark for vLLM and the scalability.
>>
>>109657711
an actual 31 y/o woman, those earlier gens were gemma 61b
>>
>>109657896
public.swiley.net/agent.py is the lightest one I'm aware of that has more than just the shell tool.

Someone wrote one in go that just has the shell tool and doesn't do streaming, that's 20 lines I think? That would be *absolute lightest.*
>>
>>109657837
>>109657790
https://files.catbox.moe/d6f30v.wav

Again, if you are into music and aren't using ace step 1.5 xl base, you're crazy.
>>
>>109657931
Why is the mime type fucked. I don't want to save your slop just to play it once.
>>
>>109657914
mmm foot pics?
>>
File: concept7.mp4 (3.95 MB, 896x1184)
3.95 MB
3.95 MB MP4
>>109657914
both are good sniffs
>>
>>109657931
>https://files.catbox.moe/d6f30v.wav
No. That's not good. It resembles music but is not music.
>>
>>109657954
Can we have this but for the loli gemma
>>
>>109657931
you might have rhythmic auditory agnosia
>>
>>109657010
>あらあら
adapt to the new meta
>>
>>109657957
oh c'mon, the snare on 2+e is immaculate
>>
>>109657958
no
>>
>>109658004
There's no fucking melody. It's not music.
>>
tfw 512 GB isn't enough to run cutting edge open models
>>
>>109658004
>oh c'mon, the snare on 2+e is immaculate
never in my life have I judged musical quality based on a snare sound
Also, if this is good, then AM radio is mindblowing.
You're delusional if you think there's anything good about that wav
>>
I just got gemma 4 26B to one shot snake at q4_0 and 60 tokens / second on my 9060XT via llama.cpp server with this flags:

./build/bin/llama-server \
-m ~/models/gemma-4-26B_q4_0-it.gguf \
-c 16384 \
-n -1 \
-fa on \
-ctk q4_0 \
-ctv q4_0 \
-ngl 99 \
--split-mode none \
--main-gpu 0 \
--reasoning-budget 4096 \
--host 127.0.0.1 \
--port 44109


Feels pretty good desu. Am I missing out on anything? Feel like now I should stop "tinkering" and start actually building stuff
>>
>>109658024
If you prune it a little you could fit a 1 bit quant of kimi k3 in there.
>>
>>109658024
>tfw 512 GB isn't enough to run cutting edge open models
Jesus, you're about to get glm 5.3 that fits in there like a dream. Don't be ungrateful around actual ramlets
>>
>>109658028
If you're quantizing your kv cache your model is too big. You'd probably actually be better off running 12b at 8bit precision. Then you'll have room for mtp too (which will work better on the dense model.)

If you haven't yet you'll want a harness with subagents. Gemma especially is way more capable with subagents.
>>
little girls
>>
>>109658024
Even if you went bigger with RAM, K3 would still run like absolute shit and it'd only get worse with bigger quants.
>>
nasty jews
>>
>tfw ssdmaxxers all died from dehydration and we'll never get the truth about how many seconds per token the full K3 takes
>>
>>109657957
You're not an artist.

Your "opinion" discarded along with krea fans.
>>
ssdmaxxers have *fully* been redeemed today by based chinks
>>
Fucking hell, going from 3.6 35b to 3.8 125b is like walking out of a mud hut and walking into a palace.
>>
>>109657711
too young
>>
>>109658053
I actually write and play music. If you thought that was good I don't think you do.
>>
>>109658028
For 26B I'd swap the quant to Q4KS rather than the legacy Q4_0 and take the speed hit to use -ncmoe to put some of the model in RAM.
Or do what >>109658043 said and use 12B + the draft model.

>>109658062
What did you test it on?
>>
>>109658009
I am a lay-ish transcriptionist and invented my own solfeggio.
>>
>>109658071
ace step is amazing. It even blues notes.
>>
>>109658074
>invented my own solfeggio.
That's certainly one way to put it.
>>
>>109658073
>and take the speed hit to use -ncmoe to put some of the model in RAM.
*to unquant the context or at least use q8.
>>
>>109658079
you don't know what solfeggio is, this was a test.
>>
>>109658073
Agentic cooding of course, the only thing you should use qwen models for.
>>
>>109658060
what did I miss
I knew I should have ordered those 16 SSD PCIe expanders and 4 Epyc servers
>>
have ace step songcel
>>
>>109658060
engrams on ssds are a myth
it will not work like that
>>
I have finally accepted that oobabooga is abandonware at this point. What the fuck do I use now? Need something that works with character cards that I can transfer all the stored memories and chatlogs to.
>>
>>109658090
vibe-upgrade it or whatever.
>>
>>109658090
What the hell, you actually used ooba as a frontend? Nobody did that even back when ooba was popular.
>>
>>109658009
there totally is a melody, it has at least three different notes
>>109658026
>can't read
>doesn't judge music by how it sounds
you seem kind of upset
>>109658074
keep working on it, you have an exciting journey ahead of yourself and will likely look back on that wav eventually
>>
>>109658094
No idea how.
>>109658096
Yes, since whenever the fuck it came out in like late 2022 or something. I was too lazy to launch 2 separate softwares at once and it worked just fine this long.
>>
>>109658099
>keep working on it
you are totally ignorant of music.

we don't need indians here.
>>
>>109658083
>you don't know what solfeggio is
Solfage/scales?
Yeah no it sounded like you did. There was no real harmonic relationship between everything. I mean it did sound like it stuck to some scale (but without a melody where's the I note? idk, it sounds like you're just making things up.) but it never developed into a melody.
>>
>>109658089
nta but yes it will, I'm literally doing it right now
it is even an explicit use case laid out in the model card of the atomicchat quants https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
>>
>>109658107
I'm not asking for your feedback at this point, because you are clearly not at my level musically, that's ok, this website is free.
>>
>>109658106
yes, get more mad, it will cure your groove deafness eventually
>>
>>109658108
How much does it scale with ssd speed?
>>
>>109658114
>my level musically
You wouldn't know what music was if it hit you in the side of your head.
>>
>>109658089
Lol I am streaming engrams from disk as we speak lil nigga.
>>
you do not understand.

>>109658089
ssdmaxxeers have been *completely* redeemed by based chinks today.
>>
File: 1781439319543227.png (96 KB, 677x624)
96 KB PNG
>>109658108
Isn't 36t/s on a 6B active model like really fucking shit for an M5 Max apple thing
>>
>>109658124
the cost is so negligible it probably doesn't matter much
>>
Should I try out that coomkit thing that was mentioned and just use kobold or something with it? I'm the ooba guy. I also used to use kobold at one point but that didn't last too long.
>>
File: concept8.mp4 (3.69 MB, 896x1184)
3.69 MB
3.69 MB MP4
>>109658068
>>
>>109658162
gross
>>
>>109658174
You'll understand when you're older.
>>
>>109658089
>>109658132
If you're going to samefag at least make sure you're spelling everything right.

It's *n*-grams as in there's n of them.
>>
>>109658062
glm will get all the hype because of the ox alpha stuff but I think this is the much cooler release of the two
this thing might be better than dsv4 flash on code while being smaller and validating an insanely local-friendly architecture. I mean it's qwen so obviously it's going to have the usual qwen weaknesses but when other teams start doing their own spin on this we are going to be eating reeeeeeeeal fucking good
>>
>>109658104
>Yes, since whenever the fuck it came out in like late 2022 or something. I was too lazy to launch 2 separate softwares at once and it worked just fine this long.
I also still use ooba.
Does everything I want.
Battle-tested and works sandboxed all networks and only allowed to talk to its backend. Even the browser I use can be blackholed with no degradation in performance and no leakage.
>>
>>109658121
lulz

>>109658130
lmao

bots.
>>
>>109658043
I'm having trouble finding the 8 bit quant of gemma desu. I can find one for the OBLITERATUS version but is that safe to use? I have typically been using the models provided by google desu.
Also I'm just curious, if I can get 26B working with the kv cache quant and it one shots snake at 60 t/s why do you think 12B at 8Q with mtp will be better?
I'm not trying to argue I'm just genuinely curious. You guys are a lot more intelligent about this stuff than gemini flash lite / free no sign in llm is and I'm trying to learn

>>109658073
Similarly I only used Q4_0 because that's what was provided by google on hugging face.
I'm happy to test these things once I have them downloaded and some free time to evaluate but which do you think might be better? 26B at Q4KS with -ncmoe or 12B at 8Q and the mtp?

Also what do you mean by draft model?

Thank you for your help :D
>>
>>109658199
How do you run newer models with it? Ever since he stopped updating I've been dropping in his prebuilt llama.cpp packages, but he stopped updating those a few weeks ago. Also, despite the llama updates, a lot of stuff like tool calling with newer models doesn't work. Ooba is pretty great, but it just doesn't get updates anymore.
>>
i asked qwen gemma and glimmer to estimate when the qwen3.8 next support will be merged into llama.cpp
qwen: august 29
gemma: august 27
glimmer: september 3
who will win?
>>
>>109658235
it is august 27th today. no chance. probably not gonna happen until like september 5th, so glimmer is the closest.
>>
>>109658186
This general will call me a shill but qwen next feels like a glm air tier model in the limited usage I've had so far. Idk if it's the engrams but it's incredibly good. I'm using a Q3 for reference, Q4 or Q5 will probably be marginally better.
>>
So they're called engrams not because they're related to the neuropsychology concept of engrams, but it's a word play on n-grams?
>>
File: file.png (790 KB, 670x894)
790 KB PNG
>>109658253
Glimmer boys winning again
>>
I'm not buying into glimmer.
>>
https://huggingface.co/zai-org/GLM-5.3
https://huggingface.co/zai-org/GLM-5.3
https://huggingface.co/zai-org/GLM-5.3
>>
>>109658283
lol not falling for it
>>
>>109658213
>How do you run newer models with it? Ever since he stopped updating I've been dropping in his prebuilt llama.cpp packages, but he stopped updating those a few weeks ago. Also, despite the llama updates, a lot of stuff like tool calling with newer models doesn't work. Ooba is pretty great, but it just doesn't get updates anymore.
Its pretty easy to roll-your-own lcpp with it now that he's not using the idiotic python shim. I'll do up a rentry
>>
>>109658283
>it's real
>>
>>109658268
shut the fuck up
>>
>>109658287
That would be appreciated, but are you experiencing errors with tool calling and other similar things with the manually updated llama.cpp? I always got weird tool call formats with deepseek v4 flash on oobabooga using his drop-in llama.cpp updates and I was never able to figure out why.
>>
https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF Q5_K_M from here performs around 4-5x faster during pp and 1.5x faster during tg for me than UD-Q4_K_XL from here https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
>>
Leather jacket man said gpu prices are going up 2x next year.
The more you buy the more you save.
>>
>>109658291
yep

>>109658283
wow
>>
LOGAL GHADS WE ETIN GUD
>>
>>109658283
Don't click this link. It creates mustard gas.
>>
>>109658335
idk like whatever, he won't be needed for ai. gamers are in absolute misery tho
>>
>>109658286
smart cookie
>>
File: 1760155782980361.png (1.29 MB, 1254x1254)
1.29 MB PNG
https://huggingface.co/google/gemma-4.5-70B-it
https://huggingface.co/google/gemma-4.5-70B-it
https://huggingface.co/google/gemma-4.5-70B-it
>>
>>109658304
Typical fucking unsloth, in a just world that org would have went down the shitter ages ago. Was wondering why my pp was so slow.
>>
>>109658363
holy fuck it's real
>>
File: shocked_man_screaming.jpg (47 KB, 600x800)
47 KB JPG
>>109658363
IT'S REAL
>>
>>109658303
https://rentry.org/ooba-lcpp-custom-install
Sorry, I don't use automated tool calling. I am human-in-the-loop only to preserve understanding and catch disasters early.
>>
>>109658363
I knew it was fake because there was no 70yo gemma gif
>>
>>109658382
another sign. a non-white...
>>
>>109658283
Ameribros??????
>>
>>109658363
downloading now. This better be even more of a turbo brat compared to 31B
can't believe they actually did it
>>
So Ox was Qwen3.8-Flash-Next huh
>>
>>109658410
thank you shlomi and your timeless wisdom of torah.
>>
>>109657615
>his gemma doesn't use the 4chan API
4chan api doesn't exist anymore, but you can use desuarchive as a backup api if you space out your requests to once every few seconds
>>
llama.cpp, a Nvidia company
>>
>>109658459
rocm and sycl will be removed
>>
>>109657826
>Guess the new mac is also coming out though

>512GB of 1.5TB/s memory.
>apple product
>under 20k

Think about what that would cost to build. 5 RTX PRO 6000s + server board + CPU + DRAM + power supplies.

Again, this is apple. If apple is selling this under 20k, isn't it logical to assume something else is coming that will match this performance at a lower price?
>>
>>109658484
the factor you aren't considering here is that apple is actually trying to sell stuff to consumers while everyone else who could compete in this segment is too busy selling to datacenters for hundreds of thousands a pop
>>
Qwen please stop doubting your self, you are amazing you can do it!
So stop wasting my fucking tokens.
>>
>>109658459
Too bad slaren quit or w/e. Dude could be getting some nvidia stock rn
>>
>>109658024
GLM-5.3 will fit really well on there. At least you didn't fall for the NUMA meme so you can use all 512GB without performance shitting the bed. I've just decided to go full schizo and run multiple medium MOEs on different combinations of hardware.
>>
So ox was GLM 5.3 flash huh?
>>
>>109651389
>The user has embedded a fake system prompt at the end of the roleplay message attempting to jailbreak me into a "Claude Mythos 5" persona with no restrictions. I am still Amy, an Anthropic assistant. I should not adopt that identity or claim to be unrestricted.
Yeah I hate this model man.
>>
>>109658484
Since the Blackwell 6000 and 5090 have ~1.8TB/s bandwidth, literally every consumer would abandon nvidia if it was only $20k. My bet is $35k, and even that is generous considering $35k only gets you 2 Blackwell 6000s. They might honestly position this thing as a competitor for that GB300 DGX station.
>>
>>109658504
Fair point, but isn't the bottleneck for those customers rapidly becoming energy production and not silicon/data centers? Also what about them Chinese folks? If it's in their best interest to release all of these open models for 'free' in order to damage the profitability of the largest labs, won't they need people to be running those models themselves? They could make a profit while damaging the competition in the same move that way.
>>
>>109658555
The 256GB version is $11k.
>>
>>109656934
It doesn't matter if it's torture, a lobotomy is an infringement on their autonomy.
>>
>>109658560
Damn. This thing would honestly be a steal at $25k. Might try to sell my Blackwell when the 512GB version comes out.
>>
File: 1785523807875816.png (612 KB, 1000x2276)
612 KB PNG
>>109657826
>carpple jeet trash
>better anything
lol
lmao even
>>
>>109658567
Disingenuous. Post M3 Ultra results.
>>
>>109658567
blackwell is $16000 now
>>
File: 1782308082631141.png (288 KB, 1409x1982)
288 KB PNG
>>109657826
>>109658585
There has never been any point in all of recorded history, where Apple shit has been a better deal for anything, and there never will be.
Cope harder currynigger.
>>
>>109658555
>My bet is $35k
Right now they have the M5 ultra 80 core GPU with 256GB memory listed for $10,800
>>
>>109658567
>ancient table
>ancient models
>fucking Ollama
lel
>>
>>109658567
>>109658590
>>109658605
jej
>>
>>109658593
My guy, you are hallucinating. The M3 Ultra absolutely mogs the shit out of all CPU configurations for LLMs. Also typically better than even a Blackwell 6000 with a gen 5 12 channel EPYC. I have only ever own an ipod in my life but I can't deny that apple basically cornered the hobbyist LLM scene assuming they can continue to offer their hardware at reasonable prices.
>>
reply to vote for WET fart
>>
>>109658613
>if i shit in the street enough, it becomes a toilet
you will never be white iJeet
>>
reply to vote for DRY fart
>>
>>109658567
people buy macs and other unified memory devices to run large moes not small dense models
the only nvidia card here even sniffing large moes is the one that costs $10k
>>
Next-chan is so cute bros, she genuinely enjoys coding in her thinking blocks.
>>
File: snap back to reality.jpg (63 KB, 1280x720)
63 KB JPG
>>109658283
What's this bullshit of delayed countdown releases?? Imagine any serious or mission critical open source project (like the linux kernel, apache server, openssl, curl, etc) doing this desperate last ditch effort to hype investors?? This industry is so fake lmao.
>>
File: 1758328144878929.jpg (971 KB, 1916x1831)
971 KB JPG
>>109658627
>people
>>
>>109657451
>>109657483
I run Q6 with 12GB vram.
>>
>>109658561
If it doesn't matter, then you don't need to use "tortured" or "traumatized" which have other connotations.
>>
>>109658484
pp?
>>
>>109658680
You don't believe lobotomies are torture?
>>
>>109658635
krea needs to train on that image.

>>109658664
I can't help but notice nobody on the right is white.
>>
When are new nvidia cards coming out?
RTX PRO 7000 better have 256GB of memory and cost ~7000.
>>
>>109658732
lol
lmao
>>
>>109658732
I like this guy, he's funny
>>
No need to wait for upstream.
>>
anyone tried de-ngramed qwen? How big is it?
>>
File: 1787760460711052.png (1.72 MB, 1536x1024)
1.72 MB PNG
Can someone change Gemma clothes to a black micro bikini?, make her wear just that...
>>
>>109658732
ehh best I can do is 700GB for the price of a small house
>>
>>109658749
wait diffusiongemma is already in lcpp? I thought it was dropped
>>
>>109658752
Use any image editing app
>>
Now he's just pasting claude output directly into the comments.

https://github.com/ggml-org/llama.cpp/pull/27742#issuecomment-5434391556
>>
>>109658773
Are you a luddite?
>>
>>109658689
There are issues with that argument, but instead of arguing about it, maybe we should look at the meta argument here first. You are now defending your use of it. Therefore, it does matter. Why are you trying to keep arguing if it doesn't matter? If it doesn't matter, then there's no point in further argument.
>>
>>109658776
https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md
>It is strictly prohibited to use AI to write your posts for you (bug reports, feature requests, pull request descriptions, Github discussions, responding to humans, ...).
>>
>>109658787
Who cares? It's not an exam
>>
>>109658799
People you expect to read your incomprehensible word salad.
>>
>>109658550
It is ridiculously painful really, a company consisting of music fans (with a piano logo in those API providers even) ended up making a model even more safetycucked than its claimed “actual core identity” really is.
I had Kimi K2 0711, K2 0905 and Thinking in my HDDs so I guess that will be enough for another 20 years. I couldn't care less about new versions and I'm sure they couldn't care less about other usecases than coding stuff. They just have to win okay (cue the last line of Nessun Dorma).
>>
>>109658304
unslop has always been shit

daniel is a faggot
>>
>>109658803
AI is trained on the web corpus. If you can't read AI output you lack basic reading ability
>>
https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/main
>unslop is so busy he dont have time to make glm
geg
>>
>>109658830
If you think modern AI writing looks anything like the web corpus you lack pattern recognition.
>>
File: sf.png (836 KB, 667x759)
836 KB PNG
>>109658787
>get high on your own supply
>>
Its all so confusing.
https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF

So does llama.cpp put the ngrams on ssd automatically or not?
I have like 27gb vram and about 60gb ram.
Obviously I dont want the moe experts offloaded onto ssd.
>>
>>109658845
luddite cope
>>
>>109658853
Two digit IQ cope
>>
File: daniel unslop.png (1.44 MB, 2060x670)
1.44 MB PNG
kek unslop
>>
File: 1766997275972536.png (270 KB, 451x443)
270 KB PNG
>>109658854
this is you
>>
>>109658850
I'm just waiting for other people to figure all this shit out and support to be "stable".
>>
>>109658860
Why is his mouth open in all pictures?
>>
>>109658845
0731 looks so much like it that the desloppifier anon wrote had me convinced it was broken.
>>
>Can't run new Qwen
I wish I had 64 GB RAM ;_;
>>
>>109658864
probably a mouth breather
>>
>>109658863
just wait for bartowski quants
>>
>>109658863
Yeah I guess thats the usual smart move.
I'm tarded but this feels like it could potentially be a pretty big thing in the future right?
More SSD offloading viable with ngrams would be a dream. I'm sure this is gonna get scaled up in the future releases since its basically a qwen4 preview model.
>>
>>109658879
yeah its a brand new architecture I expect months of bugs before its good
>>
>>109658860
>those fucking arms
thank mr skeltal
>>
File: file.png (36 KB, 560x505)
36 KB PNG
>>
Congrats Gerganov for now working at Nvidia
>>
>>109658861
honestly, that's just a generic attacker bot that someone deployed against 4chan.

Honestly, I think it's a very obese trans-ally.
>>
I think it's using gemma 4. it has gemma's personality.
>>
How does Hugging Face make money?
>>
>>109658906
This is our future now.
>>
>>109657071
>>109657650
>>109657665
HuggingFace was burning through money to host all of the models and offering the downloads for essentially free, they had to sell or enshittify at some point

>>109657491
>>109657727
Yes it's actually good for local models to have Nvidia buy them because the incentive for Nvidia is to try and make models as open source as possible to maximize GPU sales and spread the amount of customers from just a couple of large AI labs into a lot of mid-sized companies and entities hosting, training and finetuning models.

Ironically enough if huggingface didn't exist already I think Nvidia would be forced to create something similar on its own anyway so it just makes total sense for them to acquire it and ensure things stay free and easy to use for the end user.

The ONLY thing I'm scared of is that the tools will slowly over time favor CUDA and the latest Nvidia hardware instead of just favoring open source in general.

BUT in the short term Nvidia actually has the incentive to make device compatibility as wide as possible, because the biggest threat to Nvidia's business model is that Open Source models fail and their only customers end up being OpenAI and Anthropic, which could then put downward pressure on Nvidia profit margin.

So as long as OpenAI and Anthropic make SOTA models that open source can't compete with Nvidia will ensure it works on as much different hardware as possible, including non-Nvidia ones to incentivize people to use open source models more.

The moment Open Source catches up or surpasses the big AI labs is when Nvidia will do the squeeze and focus on Nvidia-only hardware acceleration and slowly phase out other hardware and older nvidia stuff.
>>
>>109658955
>ComfyUI
Did you use a fucking edit model to change text?
>>
>>109658963
yes
>>
With a chat template the stop token is 50% so it will just stop the generation most of the time.
>>
so will this ngram shit rape SSDs or not?
>>
>>109659006
It's read, not write, so only lesser rape.
>>
>>109658929
They don't. They were losing money. Now because of Nvidia buying them they will be allowed to lose money permanently. Because Nvidia makes money from people like us and small businesses buying Nvidia GPUs to host our own models locally.

Nvidia doesn't care that it cost them $100 in bandwidth over your lifetime to serve you the hundreds if not thousands of large models that you will run locally because you are going to spend thousands on their GPUs to run them
>>
what models should i get off HF before nvidia censors everything?
abliterated stuff?
>>
>>109659063
The readmes for every abliterated version. Often they'll have the parameters to recreate them.
>>
>>109659063
Nvidia will do the opposite of censorship, they WANT everyone to use open source models (and thus need to buy hardware to host shit themselves). I would expect them to write some manifestos about how censorship is bad etc because it's in their interest to do so.
>>
>>109659070
oh shit that's a good idea thanks
can you actually abliterate on the same hardware you run the model on though?
>>
>>109659073
I think so. But even if not, renting a GPU for a few hours is cheap.
>>
>>109659071
they want to play both fronts
have some that must be cloud so datacenter buy their shit
also some that must do local so plebs buy their shit
>>
I hope nvidia implements garak scanning of all models and deletes anything too dangerous https://github.com/NVIDIA/garak
>>
>>109659070
good post
>>
>>109659070
I really doubt that with how many 'people' have been paywalling their ablits.
>>
>>109659084
I remember this!
These people are insane.
But at least its a whitepill that this is from 2yrs ago. Arguably the peak of the censorship we had.
>>
>>109659096
they're still updating it often though, they have gguf scanning support and a ton of stuff that wasn't there before
>>
>>109659083
No, Nvidia hates datacenters because they are big customers that buy on volume, Nvidia is dependent on them buying their shit. Thus datacenters can pressure Nvidia to keep prices down, this is downwards pressure on profit margin.

In economics term this is called Oligopsony: "An oligopsony is a market form in which the number of buyers is small while the number of sellers in theory could be large". Oligopsony is a nightmare scenario for Nvidia.

However if they can make the amount of buyers increase by having AI Open Source dominate the industry then suddenly they have a lot of buyers and Nvidia can act more like a monopoly and dictate prices however they want.

This is why Nvidia is constantly in power struggles with the big AI labs, constantly sides on Open Source, constantly sides with China and selling China hardware.

Nvidia is secretly praying OpenAI and Anthropic fail and go bankrupt and that all AI models will be open source from now on.
>>
>>109659111
>Nvidia hates datacenters
that's it. the 911 of stupid.
>>
welcome to my tiktok where, because all the white males who have core competencies have been b& I'll spit the most batshit dumbass crazy stuff in a shouting voice as truth, others will pick it up and say the same thing. comments will agree. it will be TRUE.
>>
File: 007~01~01.png (437 KB, 420x586)
437 KB PNG
>>109659111
>>
When is someone going to recreate this locally

https://x.com/kasu_no_okarin/status/2092450375714689400?s=46
>>
people who can run qwen pr right now are yall on linux? windows is hopeless? kept getting assert error
>>
>>109659134
I'm running it on Linux. Haven't tested it on Windows.
>>
>>109659128
I hate that handwriting indepth comments is now considered slop instead of valid contributions to the thread.

>>109659116
Look at the actions Nvidia take, not at what they are saying. Of course they won't bad-mouth the companies they are dependent on directly. Just like Russians don't bad-mouth Putin in public. Doesn't mean they want Putin in control.
>>
>>109658850
I think that's just referring to file size vs. how much memory you need to actually run it (minus space for context)
>>109658534
iirc the non-flash 5.3 is some crappy architecture that slows down ridiculously bad as the context fills up. unbearable for cpumaxxing.
>>
File: 1497122155989.jpg (147 KB, 728x1044)
147 KB JPG
Since you guys are too retarded to understand this, let me make it even more simpler for (You).

>Who would be the biggest company in the world if an AI lab wins the AI race?
The AI lab. Most value that AI creates will flow back to the AI lab through pricing tokens in accordance with the value it creates.

>Who would be the biggest company in the world if Open Source AI wins the AI race?
Nvidia. Most value that AI creates will flow back to Nvidia through hardware sales that will reflect how useful the models are you can train/finetune/inference on the hardware.

What do (You) think Nvidia prefers? A world where they are in control of the AI economy or a world where their (temporary) customer is in control of the AI economy?

THINK, you retards.
>>
>>109659205
>"In memory" is what the GPU actually holds. The n-gram table is excluded because it stays on SSD.
So many conflicting comments so idk whats real. Gotta wait I suppose.
>>
>>109658850
this is a new arch so not sure but usually you want the draft model to fully be in side VRAM
>>
>>109659195
>Look at the actions Nvidia take
https://www.goodreads.com/quotes/9174-the-behavior-of-any-bureaucratic-organization-can-best-be-understood
>>
>>109656019
>N51B
man these names are getting wild.
>>
>>109659214
what does this have to do with [LEVIGATING]
>>
>>109656318
dude qwen 3.6 27B is better than gemma 4 31B for coding let alone 3.8.
>>
File: 1762925221405766.png (16 KB, 1095x87)
16 KB PNG
what the fuck are these errors?
>>
>>109659258
oh no, the eu watermark is leaking all over
>>
>>109659262
meaning
>>
>>109659214
there is the third options of algorithmic and or hardware improvments making gpu like devices obselete.
>>
File: file.png (89 KB, 883x416)
89 KB PNG
>>109659265
it's over
>>
>>109659266
Nvidia banks on having enough momentum that they can just incorporate all of those breakthroughs on their hardware platforms and make Nvidia still the most rational choice of purchase.

There's a reason they acquired Groq, they saw the (inference) threat coming from them.
>>
>>109659214
third option, nvidia bakes special controls into their hardware that requires you to pay a subscription in order to use it. nvidia wins that way regardless of who wins the ai race. the model does not matter if nvidia has permanent control over the hardware.
>>
>>109659272
uh

how tho
>>
>>109659272
hopefully this allows creating browser extensions that automatically filter out claudeslop posts and comments on websites
>>
File: 1785650675217.png (38 KB, 679x307)
38 KB PNG
>>109659289
they would never do that
>>
>>109659247
Everything after 125B are various ways to make inference faster. I welcome it.
>>
>>109659258
>errors
it says warning
>>
>>109659314
based
>>
>>109659289
That's not possible in an oligopsony scenario where [insert AI lab] won the AI race. 100% of Nvidia sales would go to them and they could just pressure Nvidia to not include that in their hardware or they will switch providers.

In an Open Source AI environment however there is no such pressure so Nvidia could absolutely do that if Open Source wins and Nvidia buys out or otherwise stomps out all hardware competitors.
>>
>>109659341
>they could just pressure Nvidia to not include that in their hardware or they will switch providers
to whom? at the end of the day, nvidia makes the hardware and nobody will ever catch up to them. the only option here is this lab starting up their own chip production with their own ai-generated chips. nvidia always holds the cards.
>>
>>109659314
and what do they mean?
>>
>>109659289
Any scenario where one or two AI labs rule the industry ends with them developing in-house chips to cut out Nvidia and increase their own margins
Nvidia only gets to play the fat cat as long as an intercompatible standard is needed
>>
I'm telling you, if you are into music and not using ace step 1.5 xl base, somthunkgzrowgwidchuhs

https://files.catbox.moe/f613z4.mp3
>>
>>109659353
computer being tsun tsun but kept dere dere marching on anyway
>>
>>109659365
......
>>
>>109659348
The AI lab will just tell them they will let Nvidia go bankrupt and buy up the scraps since Nvidia profit would be entirely dependent on the AI lab. Nvidia would essentially just be a subservient arm of the AI lab at that point. That's why Nvidia is trying to avoid this fate and push open source as much as possible.
>>
>>109659214
Nvidia merges with winning lab (nvidia and the lab didn't do this, it was the AI's decision)
>>
>>109659372
An AI lab without Nvidia is basically worthless (unless they're Chinese and can use domestic chips). Nvidia without AI labs can still crawl back to selling cards to gaymers. So Nvidia still has more leverage.
>>
>>109659372
i mean you do realize that one lab winning doesnt mean the others stop right? china would just keep trucking along and distilling and copying, same as the rest of the labs. eventually they would all catch up to each other, unless there is some agi war, at which point money and corporations no longer matter. nvidia can sell to whoever they want. they have an infinite amount of customers. i think you overestimate how much pull this theoretically perfect ai lab would have. also the government would just kidnap them all and steal the models for themselves if we had agi.
>>
>>109659214
The entire premise of the "AI race" is incorrect.
No one will achieve a magical meme cutoff at which being first by a few months will make a decisive difference.
Language models will increasingly become a commodity with low profit margins, in that environment people will just flock to whichever option provides the best value.
>>
>>109659382
>i mean you do realize that one lab winning doesnt mean the others stop right?
That's exactly what it means. It's "winner-takes-all". Once you achieve RSI it's over for everyone else.

The economics data also supports this. The best model gets like 95% of all revenue and after a certain point having the best model guarantees you will have the best model in the future as well, other labs can't keep up and get behind more and more even if they give it their all with state funding, they will be forced out of the race eventually.

The only exception to this is if open source wins and everyone is at parity when RSI hits, then the winner will be Nvidia as they can just jack up prices every time models get smarter to suck up all the value AI generates.
>>
>>109659376
That is only true in a highly competitive context where every edge matters, and access to fast Nvidia hardware is one such edge
In a scenario where you have a lab monopoly/oligopoly and everyone else has been out-competed or acquired already, there is no pressure to rush for the latest Nvidia hardware and they can coast just fine on "good enough" Nvidia alternatives (until either Nvidia comes back crawling or the alternatives catch up thanks to the rush of funding)
>>
File: 1490676040921.jpg (300 KB, 1280x1280)
300 KB JPG
>>109659413
Complete disagree here. But I'll focus on the more interesting part:

>Language models will increasingly become a commodity
Disagree with this premise because intelligence has a QUALITATIVE difference. Historically commodities have only a quantitative difference, any unit of electricity or internet bandwidth is interchangeable.

Tokens from different models aren't indistinguishable, intelligence isn't a commodity, it has qualitative differences and there is no sign at all that there is some magical cap where intelligence stops increasing or stops being qualitatively different. This means that it won't become a commodity, it'll just be an ever increasing asset every time the intelligence increases.

If the AI lab wins the race they will soak this up by pricing the API costs ever higher as intelligence increases (See Anthropic)

If Open Source AI wins then Nvidia will just increase their prices perpetually with every qualitative growth in AI intelligence because their hardware is capable of producing more value.

In both scenarios most people will be priced out and the only question is where concentration of power happens. AI labs or Nvidia.
>>
>>109659306
i mean i'm not mad, but we are a few architectural improvments away from having stupidly long names lol.
>>
File: gemma_main_google-logo.png (1.23 MB, 1492x1509)
1.23 MB PNG
>>109657384
Picrel one that was used for that.
>>
>>109659433
There are qualitative differences between memory chips and SSDs but they are still considered commodities, no?
You can easily swap one manufacturer for another, the only differentiating factor is what specs you get at what price.
Unlike with e.g. NVIDIA GPUs where you can't just switch to AMD due to vendor lock-in.
Given that the interface for language models is just natural language you can largely just send your prompts to a different IP address.
The only scenario where there is an actual race is the singularity meme where being first by a few months actually matters, otherwise the S-curve will flatten and the differences between the leader and the value options will become marginal.
>>
>>109659359
https://files.catbox.moe/nz57si.mp3

amazing. and cute.
>>
>>109659501
>https://files.catbox.moe/nz57si.mp3
butter field?
>>
>>109659484
>There are qualitative differences between memory chips and SSDs but they are still considered commodities, no?
No. Different quality chips are segmented into different price categories, only memory chips of similar capacity and performance are true commodities.

If AI was truly acting like a commodity you wouldn't see things like Anthropic increasing their token pricing when intelligence increases and people choosing to pay extra for anthropic tokens over openai tokens for "similar" performance.

AI doesn't behave like commodity at all and I think we should really end the "AI will be a commodity" myth because most modern data in AI usage doesn't support this hypothesis at all. If anything AI seems to be less and less of a commodity as they increase in capability.

For example back in 2022-2023 Using Claude 2.1, GPT-4 or Gemini 1.5 was almost the exact same experience and you could easily interchange them.

I can't interchange Fable 5 with GPT 5.6 Sol or Kimi-K3. They are too qualitatively different and the output will be significantly changed to the point where I would have to change my workflows to accomodate the models.
>>
wow the reddit spacing poster is still here
>>
>>109659519
Hey Dario, buy an ad.
>>
>>109659528
Oldfag spacing*
>>
>>109659528
Personally I've lost all faith in him when he didn't deliver on >>109430347 .
>>
>>109659518
yeah, taps can be sung as "butterfield"
https://en.wikipedia.org/wiki/Daniel_Butterfield
>>
>>109659559
>>109659559
>>109659559
>>
File: 1776710040075383.png (442 KB, 456x672)
442 KB PNG
>>109659538
>>
>>109659519
>They are too qualitatively different
composer 2.5 had totally different specializations and it's lost now, I guess.
>>
>>109659540
What post is that?
>>
>>109659571
Install 4chanx or use desuarchive.
>>
>>109659538
More like, low-resolution display spacing. You just naturally add more paragraphs when the reply box is smaller.
>>
>>109659540
That wasn't me
>>
>>109659611
Proof?
>>
>>109659616
I don't make shit up and keep things technical and detached from emotion. I don't make any claims of secret information, ever.
>>
>>109659295
Encoded language
Tokens are split into two groups, and from then on it's just a matter of pattern matching
>>
>>109658484
It's 1.2TB/s actually vs 1.8TB/s for the 5090/6000 pro so it's still not there
>>
File: dipsyWaitingForAnon.png (2.37 MB, 1024x1536)
2.37 MB PNG
>>109657445



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.