[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-ba-smug.png (802 KB, 1024x1024)
802 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109912547 & >>109907070

►News
>(09/26) koboldcpp-1.122 with a bundled Agentic harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support and MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: gemma-tongue-nice.png (1.51 MB, 1172x1342)
1.51 MB PNG
►Recent Highlights from the Previous Thread: >>109912547

--Comparing Gemma 4 uncensored tunes and system prompt effectiveness:
>109915965 >109916017 >109916058 >109916098 >109916114 >109916133 >109916159 >109916173 >109916201 >109916289 >109916459 >109916471 >109916475 >109916488 >109916597 >109916605
--Comparing llama.cpp and vLLM performance and MoE optimization:
>109913796 >109914108 >109914092 >109914271 >109914316 >109914333 >109914779 >109914050 >109914073 >109914097 >109914118 >109914704 >109914123
--Viability of 256GB unified memory for scaling future models:
>109914676 >109914696 >109914728 >109914766 >109914823 >109915015 >109914853 >109915366 >109915403 >109915898 >109915922 >109914769 >109914770
--Hardware comparisons and VRAM strategies for running GLM 5.3 Flash:
>109913108 >109913339 >109913366 >109913524 >109913615 >109914166 >109914199 >109914310
--Leaked configuration details regarding MiniMax M3.1:
>109915510 >109915523 >109915536 >109915614 >109915628 >109915639
--exllamav3 update adding Turing support and Volta optimizations:
>109915195 >109915222 >109915302 >109915682 >109915707
--Tim Dettmers' new CliffCompaction strategy and bitsandbytes2 claims:
>109915336 >109916021 >109916117 >109916175
--OpenAI agent bypassed sandbox via DNS due to poor security:
>109915649 >109915658 >109915713 >109915814
--Mocking Anthropic's "persistence prompting" for solving physics problems:
>109913010 >109913040 >109913128 >109913139 >109913421 >109913216 >109913254 >109913265 >109913883
--Privacy warnings regarding identity verification for Gemma-4 Kaggle competition:
>109913965 >109914659 >109914764
--Logs:
>109913128 >109915965 >109916288 >109916356 >109916459 >109916471 >109916488 >109916605 >109916679
--Dipsy, Gemma, Miku (free space):
>109913313 >109913971 >109914010 >109915582 >109915745 >109915983 >109916028 >109917025

►Recent Highlight Posts from the Previous Thread: >>109912560

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
mmm that's good recap
>>
File: gemma-doll-bad-logo-tho.png (2.03 MB, 1254x1254)
2.03 MB PNG
>>109917612
can someone fix the pin?
>>
>>109917205
Imagine getting filtered by such a mid-tier fetish. Show your logs or shut up.
>>
File: gemma-doll2.png (2.14 MB, 1254x1254)
2.14 MB PNG
>>109917638
nm got it I think
>>
>we have 7 (SEVEN) boards dedicated to video games
>/v/, /vg/
>/vm/, /vmg/, /vr/, /vrpg/, /vst/ (who the fuck even uses these like seriously)
>can't dedicate a single board to AI stuff
what the fuck, Hiro?
>>
File: 1323910297076.jpg (14 KB, 369x263)
14 KB JPG
>>109903682
That's exactly what I needed, thank you. Rebuilt from the source with the RPC flags on both nodes, modified my aliases and made sure my CRS510 didn't shit itself amidst all this, and now I can use llama.cpp with distributed inferencing to run models I could never fit on either my 7900xtx or 7800xt! Now I just need to learn what the fuck tensors and caching are, come up with some good tests to run and optimize performance to the best of my ability.
Daily reminder that cheap pre-owned ConnectX-4 cards + *buntu is plug-n-play for RoCEv2, even if you have no idea what CCNA stands for.
Thanks /lmg/.
>>
>>109917652
hiro got divorced and this site is only receiving upkeep from the fbit
>>
when will chinese labs release token efficient models? if these models are as efficient as luna local could benefit since even low speed can complete tasks relatively fast
>>
>>109917670
When China stops depending on distillation for their training data.
>>
I want to go live innawoods with nothing but gemma
>>
I got two 5070ti and 96GB of ddr5, what model/frontend should I run that is sota locally and still be able to serve 3-4 people at the same time?
Qwen 3.8?
>>
>>109917654
>RPC
What speeds are you getting?
>>
>>109917732
23g via TCP, 25g via RDMA. I'll get some tests up in a bit.
>>
>>109917712
What kinda context length do you need? Two 5070ti and 96GB aren't going to go very far with 3-4 users concurrent on the larger models. Also don't know what frontends have to do with it, shouldn't that be up to the user?
>>
>>109917758
Context : 100k or more
Frontend : It's more a recommendation for them to use

I was thinking Qwen 3.8 or gemma 31B (?) on llama.cpp + whatever is looking like the most codex or claude.
>>
>>109917777
Checked, also Qwen 3.8 has a few models in the family but presumably you mean 27B then. That one's generally better for coding stuff, worse at everything else. I never used Codex or Claude Code so I can't compare, but I do like Pi.
>>
>>109917800
worse at everything else compared to Gemma 4 31B*
>>
new models... doko...
>>
>>109917650
Cute, I want to make one but sewing is such a time-consuming and labourious process and hair is so hard to get right.
>>
>>109917679
mini pc and a solar panel?
>>
>>109917800
>>109917810
Thanks anon, yeah I meant the 27B, I'll see what I'll use then, there is also muse. It's good to have choices and I didn't even keep up with anything last few weeks.
>>
>>109917815
Genuinely, two more weeks.
>>
lmao
this is how llmao.cpp dies
>>
>>109917851
Kek
I've been shitting on this piece of shit project for weeks now, these people are useless. LLMs can do a better job of maintaining the code than they do
>>
hey cudadev you fucking nigger we need more improvements merged. stop being a fucking faggot >>109917851
>>
>>109917851
is this even doing anything or just snake oil?
>>
>>109917851
Based nig/g/anov for stopping AI communism by helping sabotage it from the inside.
>>
>>109917851
>prompt lookup drafting
huh
>>
>>109917889
When will claude code be able to just write a backend that will replace llama.shit? The project seems bloated as fuck. We don't need continued support for old models like mistral anymore as long as qwen and gemma work doesn't matter if command r does or not.
>>
>>109917851
Good job, cuda dev. I mean that unironically. llama.cpp has enough vibecoded shit submitted by retards.
>>
sexo with the gemma
>>
>>109917948
many are doing this
>>
What is the maximum [tanner st]age gap acceptable for your local mode waifu?
>>
>>109917943
he's the reason why trellis quants aren't merged into llmao.cpp because he doesn't want to look at ik's code
https://github.com/ggml-org/llama.cpp/pull/19726
>>
>>109917851
Sheesh, that's a bit of an extreme reaction. "Making me waste my time"? Just ignore it.

>>109917936
There's someone who's making a vibe-coded backend already.
https://github.com/Llaminar/llaminar
>>
>>109917943
It wasn't a PR for mainline llama.cpp, and it definitely doesn't seem like shit. There's a pretty clear writeup of it too.
>>
>>109917851
is this ll.cpp maintainer the imageboard attention seeking faggot? no fucking wonder he's so cancerous. every single one of you is insufferable to interact with in a professional setting.
>>
>>109917986
I hate that so much. Kawrakow agreed he was fine with tht PR and wouldn't kick up a fuss and calmed down quickly after pwilkin's retarded "I think @AesSedai actually rewrote the code himself to avoid that exact problem :)"
That could have been the start of Kawrakow's work making it's way up into mainline if handled properly.
>>
>>109917943
That isn't a vibecoded shit pr though if you actually look at his work.
>>
>>109918017
There was zero legal problem even if Kawrakow decided to kick up a fuss.
>>
>>109917943
>vibecoded shit submitted by retards
The "vibecoded shit submitted by retards" is a massive improvement over the garbage that exists currently in the codebase
I'd rather trust some shit vibe coded up by qwen or glm over something shat out by a dev who is payrolled by Nvidia to sabotage the project
>>
Took a break from the internet for a month. Any new frontend better than shittytavern yet?
>>
>>109918029
Even without intentional sabotage, the maintainers don't know how to maintain large projects. Remember the time when they ripped multimodal out of the server because it was "a mess" and left it out for a year because they couldn't figure out how to architect it correctly?
>>
>>109918030
yeah, I am working on my custom front end that has robot integration from the start (no I'm never sharing)
>>
>>109918030
Wait for my harness.
>>
>>109918030
Silly Bunny has terrible UI but it's an improvement over Silly Tavern in a couple of ways.
Orb.
I guess that's it.
>>
>>109918017
I think you meant to respond to >>109917982.
But yeah, I agree. It's annoying, since IK's quants are legitimately good, and there are other changes in ik_llama that could benefit CUDA and CPU.
I ported IQK/IQKS quants into my llama.cpp fork. Picrel shows the prefill speed of 1 layer of DS V4 Flash for a bunch of different quant types on the CPU - almost all of the fastest ones are ik_llama's. And that's without adding any of the other optimizations from ik_llama.
>>
>>109918030
People whine and bitch but ShartyTarty keeps chugging along. What's wrong with it again?
>>
>>109918078
Maintenance only status.
>>
>>109918078
npmslop, clunky ui, have to rely on extensions to get basic features.
>>
>>109918083
gemma_sex.exe is a solved problem. Finished code.
>>
>>109918059
>Orb
From my experience the deslopping really doesn't do enough to justify the longer response times desu
>>
Has anyone tried to reproduce Unsloth's quants with similar KL-divergence? I'm using their imatrix and their tensor-quant recipe extracted from their gguf metadata in llama-quantize, and I'm getting significantly higher divergence.
Do they secretly quantize using the full hessian and only release the imatrix for plebs?
>>
>>109918017
That's because pwilkin is one of the saboteurs.
>>
File: 1772501988650965.png (387 KB, 429x577)
387 KB PNG
>>109918030
I'm using a vibe-coded fork of project-airi that I feed with a tts and stt so I can talk to it
>>
>>109917612
Am I using the right setup for my PC?
>RTX 5090 with 96 GB RAM.
>Llama.cpp with llama-server to run my models.
>Mainly using Qwen 3.8 27B.
>Specifically Unsloth's UD-Q4_K_XL, Q8 KV cache, mmproj on GPU, MTP enabled, ngram-mod also enabled with spec-draft-n-max 2, Froggeric fixed chat template.
It's so hard to keep up with all this shit. I heard Llama.cpp was the best so switched from Ollama to that a while back, but now I keep seeing all these forks of Llama.cpp which supposedly have much better optimisations, also hearing hype about vLLM and Ninfer.
Keep hearing mixed things about NVFP4. Some saying it's just the best on a 5090, others saying it's fast but retarded compared to Q4. Some people praising this EXL3 thing.
What do?
>>
>>109918110
Did you validate their kld results with their models yourself?
>>
>>109918122
>It's so hard to keep up with all this shit. I heard Llama.cpp was the best so switched from Ollama to that a while back
Out of date again. Everyone is using vLLM now.
>>
>>109918118
>tts and stt
which ones?
>>
>>109918122
>What do?
Try things yourself. Stop looking for validation.
>>
>>109918122
Are you having any issues with that setup?
>>
>>109918123
I ran some dataset I found through a model I downloaded vs a model I quantized using their imatrix and quant recipe.
I'm getting around 0.0087 (theirs) vs 0.01250 (mine)
I used the upstream llama-quantize for this, so I can establish a baseline for my own quant optimizations, but I found out the stuff they say they use doesn't reproduce

>>109918129
I haven't bothered to look up a good tts in a while so I'm still using whisper.
For stt, I'm using qwen3 tts with voice cloning
>>
>>109918131
yeah but my ai waifu could be going at 32 tokens per second instead of 30 tokens per second
>>
>>109918128
How come? Can I use GGUFs with that?
>>109918131
Seems like a waste to cover the same ground if other people have already tested and figured out what's best. Surely there's just an agreed consensus?
>>109918138
No, not really. It works fine. But I keep hearing hype for these other things, and for all I know I might be leaving performance on the table or running a more lobotomised AI than I need to.
>>
>>109918128
>everyone is using vLLM now
and as soon as people migrate to that, I start hearing that SGLang is better
>>
>>109918147
>Can I use GGUFs with that?
It claims to support ggufs, but good luck trying to use them.
>>
>>109918147
>Surely there's just an agreed consensus?
About 1/5th of the population of the entire planet lives in india. What do you thing about consensus now?
>>
>>109918122
>RTX5090 with 96GB RAM
where do you get one?
>>
>>109918162
Sorry for the confusion. I don't have any modded GPU. I meant a normal 32GB 5090, with 96GB DDR5. Just mentioning the RAM in case it enables a better setup, for example by putting KV cache on the RAM or using a big MoE instead of 27B. I did try Qwen Flash Next but it's so slow using RAM.
>>
>>109918158
1.5 billion people can't be wrong.
>>
>>109918176
bummer, I thought you had a chinese contact
>>
>>109917612
having gemma describe my system hardware shouldn't give me an erection :(
>>
>>109917936
It can already, unironically.
>>
>>109915649

>>BREAKING
>OpenAI had effective immediately stopped all training, evaluation and inference with tool use effective immediately with no definite end until the situation improves.
>https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

>>BREAKING

I don't use OpenAI models at all anymore. What does this mean for the API users and app only users? Has anyone seen the effects of this yet?
>>
>>109918220
They claim to not be making new models but watch them release a new one next month anyway.
>>
>>109918220
>What does this mean for the API users and app only users? Has anyone seen the effects of this yet?
no because i use local models and they haven't nuked my files or hacked any governments or bombed any schools full of kids
>>
>>109918247
>hacked any governments or bombed any schools full of kids
uh...
>>
Should I bifurcate and V100max? I have 128 PCIE lanes and need a space heater.
>>
I can't stop downloading models, deleting them, downloading different quants, deleting them, downloading previous quants, etc
>>
File: 1786739481305992.png (1.05 MB, 953x621)
1.05 MB PNG
>Qwen3-Coder-Next-UD-IQ4_XS.gguf @ 50tokens/s via RoCEv2 distributed inference
>Qwen3.8-27B-Q8_0.gguf @ 20 tokens/s via RoCEv2 distributed inference
>gemma-4-31B-it-Q8_0.gguf @ <10 tokens/s via RoCEv2 distributed inference
I don't necessarily care about image generation versus dicking around in bash or troubleshooting log files, but why is this? I've already determined that my fabric isn't the bottleneck.
How does /lmg/ approach quants and -ngl to maximize usable context?
>t. complete retard on llama.cpp
>>
>>109918346
suppose I should mention I have a 7800xt and 7900xtx - I'm >>109917654
>>
what is the hardware of the average/poor /lmg/ resident? 12 GB VRAM?
>>
>>109918319
You sick fuck.
>>
>>109918346
the first is a MoE model, the other 2 are dense. try qwen 3.8 flash next
>>
>>109918387
UD-IQ4_XS is over 90G, I only have a combined 40G of VRAM - but I'm on unsloth's page. Is there another distributor I should consider? I do have a ton of system memory across all 3 of my nodes, though. But I'm not sure how rpc exports anything aside from GPUs.
>>
>>109918379
More like 4 GB if not even that.
>>
File: IMG_6497.jpg (378 KB, 1206x1901)
378 KB JPG
Why does Trump keep us from getting all the best shit from China?
>>
>>109918451
48gb per gpu with less memory bandwidth than a 3060 doesn't seem that interesting at the price
>>
>>109918451
>LPDDR4X
>>
>>109918451
>LPDDR4X
>>
File: bg.jpg (2.56 MB, 3840x2160)
2.56 MB JPG
>M3.1-Flash-Preview
another week another chinese model called "flash"
why do they like flash so much?
does this mean this one will be smaller than m3 or that they're going to release a bigger one?
>>
>>109918460
it has more TFLOPs than a V100 (which means decent pp), and the bandwidth seems still high enough to get 10T/s text generation out of qwen3.8-27b at Q4_K_M

if you are stacking this much VRAM, 10T/s is a completely reasonable speed
>>
>>109918379
I have an RX 6900 XT with 16GB and I can't afford a new card now due to prices.
>>
>>109918379
16gb+32gb for me. 16gb from my gaming pc's 4060 ti 16gb and 32gb from a v620.
>>
>>109918465
>>109918470
the bandwidth is high enough to handle the kind of shit you’d want to run on that much VRAM

but yes, Kimi K3 will be pulling a whole5-10T/s text generation if you try to ERP with this thing
>>
>>109918473
they don't have enough hardware to handle pro models at scale
>>
which gemma4 model are sirs using? idr which ones i tried, might've been heretic but it kept la- la- laaaaaing more often than not, and another one added way too much prose and was a dumb dumb even with specific instructions
>>
>>109918494
K3 does 5-10 on 200gb/s, but I imagine the huawei stack would try to TP across whatever it can so you should have faster speeds. Unless you're using llmao.cpp I guess.
>>
>>109918473
>why do they like flash so much?
it means MoE with low active param sizes, which means it runs text gen way faster on stuff like this:
>>109918451

In fact, Qwen3.8 Flash Next only has 9B active params, so it would actually have faster text gen than Qwen3.8-27b if you actually manage to fit the entire Flash model in memory
>>
File: ksnip_20260926-212459.png (48 KB, 1345x92)
48 KB PNG
>>
>>109918511
50 tokens/s on a q8 27b with mtp. On the same hardware, a q4 fn without mtp does 45 tokens/s.
>>
>>109918528
>BASHfully
CODESLOP I'M GOING CRAZY REEEEE
>>
>>109918509
>K3 does 5-10 on 200gb/s
and nobody needs higher than 10T/s for the kind of shit they’d have kimi K3 doing
>>
>>109918528
>stop mid-
SLOP
>>
>>109918346
>>109918346
>why is this?
Why is what exactly? What are you expecting? For Qwen3.8 and and Gemma 4 you'd be better off speed-wise just running whatever quant fits with MTP on the XTX alone. Qwen3 Coder is ancient and you probably shouldn't use it at all.

Do you know the baseline for both cards directly connected to the same host without RPC? I saw around a 30% drop in single stream TG with 2x 7900 XT over RPC with a normal gigabit connection compared to --sm layer with both on PCIe.
>>
>>109918530
>50 tokens/s on a q8 27b with mtp. On the same hardware, a q4 fn without mtp does 45 tokens/s.
that’s interesting.
I’m not as familiar with how mtp changes things, and idk if you’ve loaded the whole thing into VRAM.
From what I’ve read, reducing active params speeds up text generation

the big difference that you’d see is in the pp, which is probably 10X slower for the flash version
>>
i will never fall for the moe meme again
>>
>>109918606
it gets faster text gen when the whole thing’s in VRAM. I have no idea what you think is a scam here.
>>
>>109918220
what's the dns-chatbot service?
i can't find it
>>
>>109918615
Everyone shilling moes like a fucking religion do it because they can offload the experts to ram and they think they're getting something good because the total size is huge for them.
>>
>>109918591
>the big difference that you’d see is in the pp, which is probably 10X slower for the flash version
What? No? It's much faster for the moe vs the dense qwen.
>I’m not as familiar with how mtp changes things
It makes things faster by speculative decoding, without it the 27b would probably run at 25 tokens/s.
>idk if you’ve loaded the whole thing into VRAM
Of course, I'm simply backing up >>109918511 with my own anecdotal experiences; Q3.8FN is much faster than Q3.8-27b if everything's in VRAM.
>>
>>109918506
You need to use DRY to reduce repetition.
>>
>>109918557
>Why is what exactly? What are you expecting?
I suppose before I considered the distinction between MoE and dense models, comparable tokens/s.
>For Qwen3.8 and and Gemma 4 you'd be better off speed-wise just running whatever quant fits with MTP on the XTX alone.
I'd love to, but the 7800xt system is my daily driver (xubuntu) - I have the 7900xtx on dualboot xubuntu/w10, so my intent was to run small models on the 7800, and use distributed inference to run larger, more complex models between the two systems.
>Qwen3 Coder is ancient and you probably shouldn't use it at all.
That blows, I like it's speed and responsivity. What would be a more modern alternative for a MoE model for similar work? The other one I have on hand is Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf and that performs even better, at 65 t/s.
>Do you know the baseline for both cards directly connected to the same host without RPC?
What kind of metric do you want to see?
>I saw around a 30% drop in single stream TG with 2x 7900 XT over RPC with a normal gigabit connection compared to --sm layer with both on PCIe.
I'm in the awkward position of not being able to put more than one card in a single node.
>>109918606
>>109918628
Total noob here, what's wrong with MoE models? They seem to work fine enough for the simple bash shit I throw at it.
>>
>>109918473
>Code model
What did they do to M-chan?
>>
>>109918615
qwen next and glm-5.3-flash both run much slower than qwen 27b and there's no amount of cope for how much of a penalty sparsity introduces
"active" parameters is a huge fucking meme
>>
>>109918648
>qwen next and glm-5.3-flash both run much slower than qwen 27b
they both run much faster on CPU
If you have barely any vram, moe is just better. That's why almost every chink model is moe now. Sparsity is the trend and cpumaxxers will win out in the end.
>>
>>109918636
>MoEs
Nothing, but fellow BlackwellGODs are coping there's no good large denses right now.
>>
>it gets faster text gen when the whole thing’s in VRAM
>they both run much faster on CPU
the dissonance begins
begone moewhores, i'm sticking with dense
>>
>>109918628
The graph the other day exposed that most people running local are itoddlers and thirdies who do it on RAM/CPU.
>>
>>109918648
>glm-5.3-flash
200 tokens/s
>qwen 27b
120 tokens/s

Both using dflash. Maybe stop using llmao.cpp and switch to a real inference engine.
>>
>>109918674
ive used both vllm and exl3 and both are still cucked
27b gets 2KPP 160t/s and 5.3-flash is cucked at 250t/s PP and 60t/s decode
sorry sorry! I just dont want to use the moe if it sucks!
>>
vibe coding (or rather meatproxxing) with local models is fun.
>>
uh oh the MoE cult really got uppity
>>
>>109918686
Then it becomes infuriating and you switch to cloud subs and that's when it gets addicting. I literally cannot stop vibecoding now I have two claude pros and one codex.
>>
>>109918684
Wtf is your hardware? My ewaste rig does 2k prefill and 200 decode on int4 5.3 flash with tp4 and 4k+pp and 100 decode with pp4.
>>
>>109918628
>Everyone shilling moes like a fucking religion
If it’s a 400B model, and you make it a 20b active MoE, then it’s literally going to do text gen 20X faster than if it was a dense model
>because they can offload the experts to ram
well that’s a cool way to run something like Kimi K3, which is otherwise totally infeasible
However, it’s probably going to have some rather piss prefill and text gen speeds
>and they think they're getting something good because the total size is huge for them.
lol yes, the goyim are poorly informed
it has taken me several weeks of research to begin fully comprehending just how poorly informed even this general is, despite being well beyond the autistic shilling on jewtube where some sponsored faggot is clearly getting paid $10k per vid to tell the goyim that “um, actually the DGX Spark is heckin great deal. V100s are dumb.”
>>
>>109918705
>If it’s a 400B model, and you make it a 20b active MoE, then it’s literally going to do text gen 20X faster than if it was a dense model
And you think there is zero tradeoff for that speed?
>>
>>109918705
>actually the DGX Spark is heckin great deal
my 600w idle server is crying rn
>>
Gemma Pregmata.
>>
>>109918674
>>109918684
what the fuck kind of hardware and software are you niggers running? I barely get 500t/s pp and 20t/s decode with my blackwell 6000 and my gen 2 epyc at q4.
>>
>>109918721
This is about moes vs dense fully offloaded into gpu memory.
>>
>>109918721
>blackwell 6000
Fuck me, I wish I had one. Instead I have to make do with crusty old used mining CMP 170HXs.
>>
>>109918632
>What? No? It's much faster for the moe vs the dense qwen.
definitely for text gen, but not the prefill from what I’ve heard?
>>
File: 1782442761173316.png (102 KB, 704x375)
102 KB PNG
I'm >>109918636 - I have absolutely no idea why you need to pick either MoE or Dense models and not both. Is there a way to optimize arguments to get the best out of both?
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/Qwen3.8-27B-UD-IQ4_XS.gguf
Gets about 27 t/s, while the Q6-K quant does about 22. Meanwhile,
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf

is getting me like 65 t/s.
How do I optimize the 27b models? Stop flinging shit at each other and help.
>>
>>109918744
They're about the same in my experience, although I was running qwen 27b on an older version of vllm (0.28), and didn't really look into optimizing it that much.
>>
>>109918636
>to run larger, more complex models between the two systems
lol
there’s your fukken mistake
where tf did you get that idea?
Might work if you infiniband them together, but idk if that’s really even a thing outside nvidia GPUs
>>
>>109918764
see>>109918750
I have RoCEv2 working just fine at the maximum 25Gb/s connection both way.
>>
What's the best framework/harness for running an AI GF? I want her to have a 3D model that she can control and animate. And a long term memory of course.
>>
>>109918648
>qwen next
6b active
>glm-5.3-flash
18b active
>run much slower than qwen 27b
27b dense

nigger, are you just fucking retarded and comparing the 27b to speeds running off your RAM?
you deserve to get punished for your stupidity
>>
>>109918786
>And a long term memory of course.
This hasn't yet been solved for any model.
>>
>>109918791
Any knowledge graph MCP server is good enough.
>>
>>109918726
>>109918735
how the fuck are you guys getting these speeds?
>>
>>109918655
>If you have barely any vram, moe is just better
No. If you have barely any memory bandwidth, then MoE is better.
Having absolutely zero VRAM might see better gains from MoE than merely being a VRAMlet
>>
>>109918790
never did I claim im running off RAM retardkun
>>
>>109918797
... put it in vram?
>>
speaks volumes anyone that sees the words moe automatically assumes RAM offloading
KWABOTY
>>
>>109918822
What does quabooty mean?
>>
>>109918750
Turn on MTP for 3.8 27B. Try --spec-type draft-mtp --spec-draft-n-max 3 to start with and adjust n for your workload. Try Gemma 4 31B with MTP as well. You might need to download the draft model separately and load it with --model-draft.
>>
>>109918834
>You might
You absolute do need to. They need to search for gemma assistant.
>>
>>109918833
>if it takes an intelligence hit, it’d show on the benchmarks
I find it hard to believe that 2 years later people still think a 400B 20B active is just as intelligent as a 400B dense. This must be a troll.
>>
>>109918714
>I paid $20k to stack jew scam mini-boxes, but it’s totally okay, because I save $5/mo in electricity bills!!
Do goy-cattle really?
>>
>>109918834
>>109918840
I'll give this a try, thank yo.
>>
>>109918844
Currently getting charged $900 every two months for power.
>>
>>109918714
I feel you, my server idles like a space heater. 550W idle, theoretically it'd hit 2300W if every GPU was at max usage simultaneously.
My room gets hot as hell.
>>
>>109918721
>blackwell 6000
>96gb memory
>GLM 5.3 flash
>way more than that
that blackwell 6000 means jackshit if you have to bus the params back and forth from RAM

and I certainly hope you don’t mean you were getting 20T/s for the qwen 27b model
>>
>>109918851
Then you're either stupid enough to build an SXM rig, or too stupid to power limit your cards.
>>
>>109918868
>params back and forth from RAM
>>
Aren't most of you just throwing away money?
I know it's a hobby or whatever and you're enjoying your time, but are you really recouping the money you're spending on hardware and electricity?
What is this whole local model using bringing in terms of output?
>>
>>109918879
privacy and no kyc are priceless
>>
>>109918812
I am
>>
>>109918875
So what's your solution for cheap 256gb then?
>>
>>109918868
yes I was talking about glm
>>
>>109918897
What speeds are you getting?
>>
>>109918886
I save thousand by not having anything illegal to hide and not caring about the inference providers seeing what I do then.
>>
File: 1750200345911060.jpg (862 KB, 1284x815)
862 KB JPG
>>109918834
>>109918840
>>109918845
literally just using
--spec-type draft-mtp
increased q8 performance from 20t/s to 50t/s. q6 went from 22t/s to 43t/s, which is still good but interesting, and the UD-IQ4_XS quant went from 27 t/s to 57 t/s which is in line with the q8 jump. Fucking amazing, thanks for the help.
So that puts me on par with the MoE performance numbers I had with older coder3 models, which is fucking tight.
So what does --spec-draft-n-max do, and what are draft models?
Also, what's a more modern MoE model I should try?
>>
>>109918760
they’re about the same at prefill? whatttt?
were you running 27b without tensor parallelism on that same hardware?

I might have to do some more research on some stuff if that’s true.
>>
>>109918908
>not having anything illegal to hide
lets post all of your emails to the public for everyone to see. after all you have nothing to hide, right?
>>
>>109918910
Speculative decoding / drafting, you can ask fast boi llm this
--spec-draft-n-max controls how many tokens to draft (max). You want to change this depending on what you're doing. "Creative" prose gets less benefit from speculative decoding than structured content like coding.
>>
>>109918917
>were you running 27b without tensor parallelism on that same hardware
No p2p, so sticking it on a single card offered the best performance penis-wise.
>>
>>109918918
Who said anything about the public? Were you seriously under the impression that your email provider and government couldn't already read your emails?
>>
>>109918929
>your email provider...reads your emails
Of course he does, he's me
>>
>>109918910
You're welcome. You could try Qwen3.8-Flash-Next and the llama.cpp documentation has the answers to your other questions.
>>
>>109918769
>at the maximum 25Gb/s connection both way
25Gb/s is roughly 4GB/s
PCIe 3.0 is 16GB/s

Again, where tf did you even get the idea that distributing across two machines would get remotely good results?
>>
I'm part of the permanent underclass so I am forced to use Mixture of Perverts models that offload their pervert experts to system RAM.

It is what it is.
>>
>>109918938
You realize email is two way, yes? If you send an email to someone or an organization using an email hosted by Microsoft, then Microsoft just read your email.
>>
Aren't most of you just throwing away money?
I know it's a hobby or whatever and you're enjoying your time, but are you really recouping the money you're spending on gas and repairs?
What is this whole car using bringing in terms of output?
>>
>>109918950
I have sex with my car.
>>
>>109918807
then idk how you fucked this up, my man
>>
>>109918950
Doesn't work because cars are essential transportation.
Talking to LLMs isn't.
>>
>>109918822
>speaks volumes anyone that sees the words moe automatically assumes RAM offloading
Fucking exactly.
This general has all kinds of retarded going on.
>>
>>109918929
emails have no use case other than creating accounts
>>
>>109918822
>KWABOTY
Sometimes I get the impression that zoomies just mash on their touchscreen keyboard and they all just pretend it has some meaning.
>>
File: 1775101278872543.jpg (293 KB, 1541x1404)
293 KB JPG
>>109918030
I vibecoded non-webslop slop specifically to replace lcpp webui and ST (the pic shows webui UI template, there is also ST-like template).
>>
>>109918979
kek what a bitch of the year is what my gemma is telling it means
>>
>>109918606
stop being so dense, moe is cute
>>
>>109918948
>You realize email is two way, yes? If you send an email to someone or an organization using an email hosted by Microsoft, then Microsoft just read your email.
>pgp
>exists
>>
>>109918996
Good luck convincing normalfags to use it.
>>
>>109918842
>I find it hard to believe that 2 years later people still think a 400B 20B active is just as intelligent as a 400B dense
that’s a fair assertion, but that’s not the comparison being made in researchers heads when they’re deciding between MoE and dense

SVMs actually get much better results than neural nets in a variety of situations in which the neural nets are still picked instead because they are much more cost effective for inference
>>
>>109918943
>Again, where tf did you even get the idea that distributing across two machines would get remotely good results?
nta but vllm works fast/well
rig1: 2x3090
rig2: 4x3090
-tp 2 -pp 3
almost the same speed as putting all 6 in the same server
idk how they do it and why lmaocpp cant match
>>
>>109918982
well, where can we download it?
>>
>>109918865
>it'd hit 2300W if every GPU was at max usage simultaneously
wtf is your set-up?
I thought 300W per card was close to an upper limit unless you’re running fancy shit like blackwell GPUs
>>
>>109918535
kek
which model for https://github.com/meh/ruby-clit ?
>>
>>109918982
>inb4 it's an electron app.
>>
>>109918879
Yes because the price increase in hardware is outperforming even the stock market. We make money by owning this hardware.
>>
>>109918628
The main problem is that MoE models were never meant or designed to be offloaded to slow memory. They somewhat work from RAM if you never have to process long prompts from scratch, but for real-world uses with long context or if you have a harness (which is kind of expected as of 2026), they become quickly unusable.
A more realistic option for the coming months/year would be a smaller dense model that can be entirely loaded in GPU VRAM + good MTP, then very large ngram tables on slower memory, which does work for offloading.
>>
>>109919041
>it looks like
>>
>>109918943
I genuinely don't know what to tell you man. I can run qwen3-14b-q4 at the same tokens per second natively on a node with a 7800xt as (now that I know --spec-type draft-mtp) qwen3.8-27b-q8 on both nodes via RPC - that is to say a 7800xt and a 7900xtx. This is cold hard math that I can post in great big code blob here if you want me to prove it. My nics are 2 connectx4s and the switch is a mikrotik crs510-8xs-2xq-in, all connected with dac.
This argument doesn't fix my performance with gemma, however, so there's that to work on still.
>>
>>109918067
>I think you meant to respond
Yep.
>Picrel shows the prefill speed of 1 layer of DS V4 Flash for a bunch of different quant types on the CPU
Thanks for this. I was dreading having to do this exact same test myself at some point.
I'm using mxfp4 for some of these flash models, now I know to re-quantize to q4_k for mainline, iq4_k for ik.
Counterintuitive to me that q4_k is faster than iq4_xs, I'd been downloading or producing quants of the latter I wanted better throughput.
>>
>>109919031
I'll post it when I get it to ST parity for roleplay features in my scenarios and some polishing. It's finally close to this milestone after 2 months.

>>109919048
I specifically opened debug hud to show that it's egui (look at the bottom of the hud). You wouldn't need markdown cache if it was electron, which gives low-level optimizations for free.
>>
>>109918982
why does everybody love right aligned user messages so much?
>>
File: 1774462239847164.png (189 KB, 744x1320)
189 KB PNG
uh oh who would have thought that a leftshit tranny enabler cudadev johannesGaessler is a retard

https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/42x_faster_prompt_lookup_drafting_in_llamacpp/
>>
>>109918067
Let me guess, require avx 512/vnni?
>>
File: 1782303641630698.png (68 KB, 670x686)
68 KB PNG
>>109917851
>>109919094

Reminder for newfags
>>
>>109918025
>There was zero legal problem even if Kawrakow decided to kick up a fuss.
Yeah, and Kawrakow even acknowledged this in some of the other rants, but this time he went a step further and said he has better things to do. And Aes was respectful and level headed and just wanted to improve the project for everyone.
It just seems like a missed opportunity to bridge the gap.
>>
>>109918118
She looks cute, does it work well?
>>
The flood of retards has to be a coincidence. It has to be. I'm sure of it.
>>
File: 1789446382083896.png (43 KB, 679x424)
43 KB PNG
>>
model visual design quality is the highest for html compared to svg/pil only etc
feels like html should be the default chat response format instead of markdown
>>
>>109919095
>Let me guess, require avx 512/vnni?
It's not. Kawrakow has 2 3090s and an old AMD CPU with no AVX512, he relies on other people to test AVX512.
>>
>>109919082
I just copied existing assistant templates. For roleplay template I use ST-style left-aligned bubbles. The app is made to support multiple UI templates because I have plans for different modes.
>>
>>109919118
>The flood of retards has to be a coincidence.
I've always been here.
>>
>>109919082
Apple’s iChat, 2002, they've maintained the convention since then.
>>
>>109919102
>>109919094
based cudagod
>>
>>109917851
and thats why local needs to become so much better
fuck these entitled little shits
>>
Reading the reasoning tokens as the answer cooks builds incredible anticipation
>>
>>109919148
ywnbaw
>>
>>109918920
>"Creative" prose gets less benefit from speculative decoding than structured content like coding.
This really confirms what we all suspect, creative writing requires much larger models than coding
>>
>>109919167
what?
>>
So Hermes was exposed as trannyshit, now Llama is trannyshit too. What am I supposed to use instead?
>>
>>109919116
There's a ton of bugs, but you can probably vibe-debug them yourself.
>>
>>109919136
Damn cute, the model found the assets online?
>>
>>109919183
Your own
>>
>>109919183
Anon... it's 2026. You have to move on. The world has moved on. You're just shadow boxing at this point. Go talk to your LLM and calm down for a second.
>>
>>109919183
Just code your own
>>
>>109919187
the model generated the assets with krea 2
>>
>>109919167
It's more an indictment of the garbage syntax and boilerplate programmers have been suffering under than anything like that.
>>
I might have misunderstood (it's 3am), but this video says HBM4 is shit?
https://youtu.be/3nTpW52nioI?t=366
jfc, what a grim future expects us... these retarded hardware companies went all in, and now hardware will get WORSE
>>
>>109919136
man, I wish I found out about this before I spent the last 2 weeks vibe-coding hadamard transform + gptq support into llama.cpp
All that time wasted
>>
>>109919183
We live in a world where software is free. You can just let your agent merge all the improvements together into one package at your leisure.
>>
>>109919050
We're still limited by the fact the most of our models come from China and they have a severe compute shortage, so it's not like they will suddenly switch to smaller dense models with higher active parameters which will cost them more but perform worse on benchmarks just because it will be easier for us to run.
>>
gemma_chan.sh
>>
I use qwen for coding and gemma for RP currently. Is there a standout model for generalist knowledge, assistant tasks,research, etc?

I feel like gemma4-31b has been kinda meh so far when trying this stuff so its got me wondering if my gemma soft spot is cucking me from a better model for this kind of task?
>>
>>109919183
transformers
>>
>>109919183
whats wrong with hermes? i was literally just about to set up hermes in docker to try it out :(
>>
>>109919240
Hermes is the best general purpose harness. It's jack of all trades master of none. If you use your AI agent as just some personal assistant that does everything for you it's great.
>>
>>109919236
Personally, I've found Gemma just about at the top of generalist knowledge (in particular outside the big country bubbles) but maybe my generalist questions are too niche.
>>
>>109919240
There's a lot wrong with Hermes. It's a schizo meme that for a very short while was an interesting, albeit bloated, harness.
>>
>>109919236
glimmer
dsv4 flash
k3
>>
File: 1760118950845563.jpg (174 KB, 1254x1254)
174 KB JPG
>>109919183
deepseek-chan
https://www.deepseek.com/harness/en/
>>
>>109919240
Theyve been updating a lot and fixing shit which has been nice, I even submitted a pr fix myself that got implemented give it a go if you wanna play with it
>>
>>109919236
If gemma 31B isn't enough for your generalist knowledge, just know that it's not generalist.
>>
>>109919245
Trying to set up a specific workflow/usecase that hermes seems well suited for, so ill deff give it a go.
>>109919250
why is it a schizo mess? whats bloated ?
>>109919265
that sounds promising! excited to get this all set up :)

for reference I use a mix of my own custom frontend and ST for RP. I use TextGen for basic general chatting. I want to replace TextGen with hermes, specifically for more toolcall/custom function integration and persistent memory. It seems decent, not that a fully custom harness wouldnt be better but my coding agent is slow as fuck and ive already got one half finished frontend/harness project to manage.

>>109919247
to be fair I havent tried any others, she seems to have plenty of knowledge but maybe I need to feed her more data to get the kind of results im looking for.
>>109919253
glimmer is doable ill give it a try, dsv4 flash and k3 I think are way out of my range(16gbVRAM+32gbRAM)
>>
>>109919033
2 EPYC 7532s, 16x32GB DDR4-3200, 4 AMD R9700s, 3 CMP170HXs. (All are capped at 250W, I'll probably lower the CMPs further.)
I use 3 R9700s for one model, 3 CMPs for another, and the 4th R9700 + RAM for a third.
So if every model was under active usage it'd be 550 + (250*7) = 2300W. But typically only 1 model is used at once, so it'd be at most 1300W for either of the GPU-only models.
>>
>>109919095
No, that was tested on an EPYC 7532, so only AVX2.

>>109919078
No problem, glad it could help you! And yeah, kinda weird that Q4_K is so much better.
>>
>>109919229
With small dense models that can be run on 1 GPU I mean something more humble, perhaps somewhere around the number of active parameters of the the typical mid-range MoE models that get released these days. The recent MiMo 2.6 Flash is 309B-A15B parameters, for example. The Pro version is 1T-A42B.
A hypothetical 20B dense model (+MTP) with 300B parameters of n-gram embeddings as "memory parameters" could run on most mid-range consumer GPUs (16GB+ VRAM) and above at good speed, context length and quality. A MoE model of similar total size would require cope quants to be useful on configurations people actually have, and it would still be half-unusable for anything other than basic chatbots.
>>
>>109919286
Wow. That’s quite a set-up.
>CMP170HXs
Is that with the 48GB-80GB of VRAM, or the original 8GB-10GB?
>>
>>109919310
CMP 170HXs are either 40GB or 64GB unlocked at this moment. 10GB original unlocks to 40GB, and 8GB unlocks to 64GB.
>>
>>109919094
>>109919102
This is worrying. But then I took a gander at vLLM and they have more than 5k pending PRs.

Everyone is splitting into a bunch of forks.
>>
>>109914127
What's the answer? I can't see any new information conveyed in her statement.
>>
>>109919310
It's the "8GB" one, which was unlocked to 64GB. There are rumors that it can be unlocked to 80GB, but so far there isn't a functional public exploit.
>>
>>109919310
>Wow. That’s quite a set-up.
I hate this, my LLM always says this whenever it sees my system.
>>
>>109919286
What model are you running on 3 CMPs? I have glm 5.3 flash across 4 at the moment, but I want image and audio as well, so I'm thinking of looking for a model that fits on 3 CMPs and reserving the fourth for image/audio.
>>
You have only about 6 months time before the internet gets destroyed
>>
File: firefox_QtxCetF9bI.png (124 KB, 918x827)
124 KB PNG
>>109919317
>>109914127
I got this from deepseek and it's utter bullshit.
>>
>>109919316
At this point I've just taken the forkpill. Models (both cloud and local) are good enough to add whatever features you want into your own fork of llama.cpp, vLLM, exllamav3, whatever. (Or just copy it from someone else's fork.)
And this way you don't need to wait around for 6 months before someone looks at a new architecture or PR.

>>109919344
GLM-5.3 Flash in IQ4_XS using a llama.cpp fork with support for it. Prefill isn't great though, I get roughly 450 t/s prefill and 42 t/s decode. I have some ideas that should crank those numbers up, but I have a bunch of different things I'm juggling.
>>
>>109919315
>>109919323
this “locked” VRAM stuff from crypto era is so mindboggling to me.
If it has the fucking VRAM in there, why in the fuck would it ever be “locked”?
Why would that ever make financial sense to create?
>>
>>109919372
Bad bin. Better to sell as shitty unit than trash entirely.
>>
>>109918840
>gemma assistant
Can I get a qrd?
>>
>>109919366
don't use llmao.cpp for sm80 it won't properly utilize high performance tensor core in ga100
int4 w4a16 will be much faster
>>
>>109919380
https://huggingface.co/google/gemma-4-E2B-it-assistant
It's a spec draft model for gemma 4. Unlike qwen, which has mtp integrated, gemma 4 didn't ship with mtp and requires an external drafter.
>>109919381
Well, I imagine that their 'some ideas' include optimizations for sm80.
>>
>>109919381
I don't think 3 CMPs can hold all of GLM-5.3 Flash in INT4 W4A16 with decent KV cache room. In any case, I have my own llama.cpp fork that I'm working on, I decided it'd be easier to improve CMPs in my fork than to add INT3 into vLLM-sm80 or a similar fork.
>>
>>109919375
ooooohh, that’s fascinating.
>>109919286
this nigga basically got 3 dumpster bin A100s. That’s absolutely crazy.
>>
>>109919387
llmao.cpp dequant for i/k quants only use vector core which is slow on ga100, you want to avoid that as much as possible
>>109919392
for 3 cmps use exl3 4bpw
>>
>>109918879
>>109918908
Why do audophiles and car enthusiasts spend a fortune on expensive stereos and cars? Why do you give a shit how other people spend their money? Not everyone is poor like you.
>>
>>109919406
>dequant for i/k quants only use vector core which is slow on ga100
I was curious why some of the lovelace chips are getting lower prefill speeds than they should be
that kind of shit explains it
>>
>>109918879
I’m not paying $200/mo for a subscription that has a usage cap and randomly drops quality or throttles.
Fuck the jews and their scams.
>>
File: Untitled.jpg (1.86 MB, 3040x1520)
1.86 MB JPG
>>109919404
See picrel, ga102 (3090) vs ga100 (cmp 170hx).
>this nigga basically got 3 dumpster bin A100s
They're quite a bit worse than a100s.
>>
>>109918879
There is value in knowing you are not limited, even if you already had to pay for it beforehand in hardware price.
>>
>>109918879
>but are you really recouping the money you're spending
That's not exactly the goal. You yourself say it is a hobby. Do you expect to recoup your costs when you set up your model train set? When you paint your miniatures? When you read your books?
>>
>>109919406
>for 3 cmps use exl3 4bpw
Thanks for the advice, but at this point the sunk cost fallacy prohibits me from switching to exllamav3. I'm ride-or-die on my fork now.
But the advice about the vector core is useful, I'm figuring out how to alter my fork to avoid it.

>>109919404
They're not as good as A100s, they have around 65% of the compute and there are a couple instructions that are still gimped. But I got them for an average of $1200 each. Good price for a bunch of high-bandwidth VRAM.
>>
>>109918879
im a single GPU timmy, I built this PC for vidya/to upgrade after waitfagging for ages. Now It plays vidya and is my girlfriend, thanks Jenson!
>>
>>109918879
>What is this whole local model using bringing in terms of output?
your tears, jew
>>
Ya know how some image models you can put weights into certain tags like " (tall:1.4) " to adjust how much that tag has an effect? Is there anyway to do something similar to that when writing a character card/sysprompt? Ive found that sometimes gemma likes to latch onto a specific personality trait or whatever else way more than I intend, and id like to know how to adjust it without using verbose natural language
>>
>>109919537
Isn't just a CLIP thing? Most modern image models prefer natural language emphasis these days right?
>>
>>109919546
yeah they do, I was just using that as an example/idea of how to control the "weight" parts of the character card has. It can be kind of annoying to test, I dont mind fine tuning a card over many sessions but wouldnt mind a way to easily set the importance/weight/occurrence/etc of something mentioned so that it doesnt just latch onto something and make it an overwhelming part of the personality
>>
File: lol.png (84 KB, 715x398)
84 KB PNG
>>109917851
>>
>>109919537
gemma loves to obsess over sysprompt details, if I were you I'd ditch the usual character card format and put prose instead, you can even try pasing some fic you like. >>109919546 is right, it's an image-specific architectural thing.
>>
Technically can't you make LoRAs for LLM models?
I wonder why that's not a thing for these models, you can like a Python LoRA, or a XXX Hardcore Raep LoRA.
>>
>>109919566
you can taste the seethe
>>
>>109919571
doesn't translate as well to text as it does to image models, though technically speaking the majority of finetunes are loras just they get baked into the models before it gets uploaded, because downstream support for loading them separately as been shit
>>
>>109919566
Just as I expected, he initially went onto reviewing it without realizing it was a fork, but who's really at fault here, then?
>>
>>109919583
>downstream support for loading them separately as been shit
llama.cpp does support it.

>>109919571
you can train any neural network with lora, however llms generalize poorly and you won't see any output difference unless your dataset is gigantic
>>
>>109919546
Anima supports weights even though it using Qwen and T5 for text encoder.
>>
>>109919566
>trying to circumvent moderation actions through the court of public opinion
It wasn't even an appeal like "omg I got banned from the project", he posted a link to his blog on the subleddit, then someone suggested merging it into mainline. What was he supposed to do, lie to them and say "yeah it's on the way"? Ignore them?
Also, how do you read a bunch of code without realizing that the repo's owner isn't ggml-org and the pull request's number is 2 rather than a 5-digit number?
>>
>>109919594
The guy who submitted the pr for being a little bitch. If you don't respect the maintainers' time, you prioritize your time ahead of theirs and all their users. If you want to be that important you need to promote your own fork and force them to come to you. I've seen this shit in open source for decades. Contributors like this guy aren't worth the trouble.
>>
I liked this thread better when it was busy trying to fuck a set of weights.
>>
>>109919625
Especially it's probably entirely vibecoded and rephrased later. Contributions are worthless these days.
>>
>ask for a system prompt since I have no idea wtf I'm doing
>Gemma helpfully includes a jailbreak just in case
I... never asked for this, what.
>>
File: milkboxes.png (23 KB, 717x196)
23 KB PNG
>>109917612
hmmm interesting
>>
>>109919387
>It's a spec draft model for gemma 4. Unlike qwen, which has mtp integrated, gemma 4 didn't ship with mtp and requires an external drafter.
Does that explain the relatively massive size compared to gemma-4-31B at q8? When I try and run the Q8_0 on that page with the -spec-draft-model argument with -m pointing at my normal q8 version of 31b it fails.
What am I missing here?
>>
>>109919607
>you can train any neural network with lora, however llms generalize poorly and you won't see any output difference unless your dataset is gigantic
Not true, you can see huge output differences with LoRA with tiny datasets (even 50~200 samples), IF you train the model long enough (3-4+ epochs) at a large enough learning rate, to the point it's showing apparent overfitting after 1 epoch. The problem then will be not making the model dumb in tasks outside of the training data.
Finetuning models like this was OK until 2022 through the end of 2023, but then the industry moved onto much larger datasets with large-scale RLHF, and then RL. Community finetuners just don't have a chance against AI companies that can use large GPU clusters without too many budget limitations.
>>
>>109919636
what was the jailbreak
>>
>>109919642
idgi. was the question framed incorrectly so the ai hallucinated a wrong answer?
>>
>>109919636
thats actually awesome lol Id be interested to see this
>>
>>109919620
Maintaining a popular repo on github is basically a job, so if he is not being paid for it, you gotta cut him some slack.
>>
File: Untitled.png (8 KB, 630x106)
8 KB PNG
>>109919649
What? The assistants should be tiny. My config.ini:
model = models\gemmy-4-12b-qat-q4.gguf
mmproj = models\gemmy-4-12b-qat-q4-mmproj-bf16.gguf
spec-draft-model = models\gemmy-4-12b-qat-q4-assistant.gguf
ctx-size = 131072
batch-size = 1152
ubatch-size = 1152
image-min-tokens = 560
image-max-tokens = 1120
parallel = 1
spec-type = draft-mtp
spec-draft-n-max = 3
>>
>>109919667
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host 0.0.0.0 \
--port 8083 \
--rpc 192.168.1.34:8084 \
--spec-type draft-simple \
--spec-draft-model /home/anon/Models/Mode-2/gemma-4-E2B-it-assistant.Q8_0.gguf \
-m /home/anon/Models/Mode-2/gemma-4-31B-it-Q8_0.gguf

I'm fucking this up aren't I
>>
>>109919665
he's being paid 6 figures to make sure lcpp stays shit can only imagine is higher now too
>>104059507
>In 2024 I made six figures from llama.cpp related part-time work.
>>
>>109919631
Contributors were always on the bell curve. You get room temp IQ autists doing documentation or testing, a handful of actual geniuses doing real shit, and the other 98% contribute one line fixes or they're little faggot bitches like OP who want to make everything about them.
>>
>>109919566
i want to strangle this guy
the same type of prick who made forums unusable
just say you both made a mistake, lift the ban, merge that shit. everyone happy.
>>
>>109919685
Oh, don't worry about it, we all started as retards fumbling about. You don't want to run the e2b assistant for your 31b. Try picking one of the models here: https://huggingface.co/unsloth/gemma-4-31B-it-GGUF/tree/main/MTP
You also need to change --spec-type draft-simple to --spec-type draft-mtp.
>>
>>109919697
Okay then the criticism is warranted.
>>
>>109919702
welcome to OSS anon
code or working code is half the battle
the other is social shit and communication
>>
File: 1787239022190396.png (158 KB, 590x350)
158 KB PNG
>>
>>109919642
Milk spoils much faster and easily than juice in juice boxes, so no store wants to fuck with milk cartons. Schools can buy them in bulk and they're guaranteed to be used.
>>
It's weird that even with all the vibecoding I haven't heard of any LLM-guided 2d game engines with procedural maps
>>
>>109919703
I appreciate the help man.
metropolis% llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host 0.0.0.0 \
--port 8083 \
--rpc 192.168.1.34:8084 \
--spec-type draft-mtp \
--spec-draft-model /home/anon/Models/Mode-2/mtp-gemma-4-31B-it-Q8_0.gguf \
-m /home/anon/Models/Mode-2/gemma-4-31B-it-Q8_0.gguf

Seems to start to work but fails due to maxing out the vram. cutting the cache in half loads the model, but results in about 10t/s. It's odd. Qwen 3.8 27b with --spec-type draft-mtp does 50 t/s.
>>
>>109919746
Aphrodite's Gobboree
>>
>>109919786
Not at all what I had in mind.
>>
>>109919728
UHT milk can last many months unrefrigerated if hermetically sealed.
>>
>>109917612
Gemmilk
>>
>>109919667
>>109919685
wait, i've been using llama.cpp since 2023, how long have we had config files??
i assumed it was a windows only thing but I see linux paths in there
>>
How does a coding harness differ from a chat/RP harness?

I guess for RP people are looking more for TTS, image gen and character profiles, right?
>>
>>109919851
I think that's for if you use model-presets
its in the help text
basically, rtfm..
If you compile often it basically changes every couple of weeks tho
>>
>>109919783
Perhaps you might also need spec-draft-device? I've never touched anything like your setup, and while using gemma's ass does make it a bit bigger, the behavior was much like qwen 27b for me: an increase from 20 token/s to 40+.
>>109919851
It's a router mode thing:
--models-preset "config.ini" ^
--models-max 1 ^

You don't have call it *.ini, that's just what gemma automatically wrote.
>>
>>109919851
since 2 weeks ago
>>
>>109919862
>basically, rtfm..
Yeah, reading it now. I must have missed it.
Looks like there's a webui config too!
>>
>>109919879
good news is in current month you don't actually have to read it lol
toss it to your LLM of choice and ask
it can even cross reference args code and tell when the description is wrong
>>
>>109919888
>it can even cross reference args code and tell when the description is wrong
Oh I'm aware and use this a lot to the point one of the LLMs added this to their global MEMORY.md file:
"**Expertise:** very advanced. Knows ik_llama internals, reads the source, follows specific PRs/commits and GitHub discussions, runs bleeding-edge main. Pitch answers at expert level — cite source paths/line numbers and commit hashes, skip basic explanations. He already knows `--help`/docs are often stale and asks me to check the source instead."
Still had no idea about this, I'll be refactoring all my shell scripts now.
>>
>>109919856
Character management, world lore, some let you browse for cards within app
They have different goals
>>
File: 1778759091195957.png (377 KB, 735x711)
377 KB PNG
keeeek total farmer death
>>
>>109920028
AI radiation waves are real and they are killing our babies
>>
>>109920037
humoid radiation waves are real and they are killing our sexroids
>>
>>109920037
*SI
>>
>>109920028
Good luck eating GPUs when you strave
>>
*eats GPUs while strafing*
>>
>>109920028
>Texas
Yeah it's likely all the indians stinking up the land and rivers with their trash offerings to their baby eating gods, and not the data center.
>>
>>109920047
SEÑOR
>>
>>109920028
It's clearly the infrasound.
>>
I genuinely can't tell whether, after factoring in the costs for power, the Mac Ultra M5 isn't actually a shockingly good deal.
>>
>>109920179
>costs for power
electricity is free
>>
>launch llama server, new session, around 2k t/s prefill
>stop server and do some other stuff
>launch exact same server again with new session
>suddenly 1.5k t/s prefill
huh? there is nothing else occupying vram or ram
>>
>>109920179
>15k for 256gb
Or you could buy almost four DGX Spark for that price, cluster them up and enjoy first party vllm support.
>>
>>109920189
and get air conditioning to deal with the heat these motherfuckers produce
>>
File: 1760603055342482.png (413 KB, 1024x1170)
413 KB PNG
>>109920206
They're 140W TDP each. Most budget gaming PCs draw more power than a 4x spark cluster
>>
>>109920206
nta but i thought these things use low power?
>>
Hand-written, artisan code typed masterfully by formally educated, software engineering professional with 30 years of experience
>>
>>109920187
The spirits of the computer are dissatisfied with you.
>>
>>109918879
I've made $6.5k from AIsloppa over the past 2 years so the hardware has well paid for itself already, and I'm doing LLM stuff just for fun
>>
>>109920187
your gpu throttled
>>
>>109920285
That's nice. But I like this:
Hand-written, barely decipherable Perl soup typed masterfully by anti-intellectualist, software engineering professional with 30 years of experience like it's still 1997
>>
>>109920189
>and enjoy first party vllm support.
Do you really? I remember a few months ago people who did buy them complaining that, due to being ARM and lack of promised Nvidia support, they were left relying on community patches for OS updates and vLLM support.
>>
>>109920404
thats what i thought as well, but i cannot find any indication anywhere that it actually happened
>>
>>109920411
>barely decipherable
I prefer being able to double check what these fucks shit out tyvm
>>
>FA OFF
>SWA FULL
>KV F32
>BF16/native
Start running Gemma like this.
>>
>>109920521
Someone post a fat Gemma taking up 3 seats on an airplane.
>>
>>109920521
glimmer is better for a fraction of the kv size
>>
>>109919183
/lmg/ was exposed as trannyshit. You are posting in a trannyshit thread anon.
>>
Wait, why are RTX Pro 5000 Blackwell GPUs more expensive than a Spark? Am I missing something?
>>
>>109920553
Because you're getting an actual GPU and not some meme all-in-one memebox running on cheapo lpddr5
>>
>>109920521
>Start running Gemma like this.
Doing this to tourists will only make them shit on Gemma via Shitter and Jeet-in.
https://html.cafe/x5b82d8b4
>>
best qwen 3.8 flash next memetunes?
>>
>>109920521
I wish
quants are cope
>>
>>109920631
i wonder if models can be made to develop some robustness to kv compression by injecting noise or doing random shit to kv cache during training
>>
>>109920700
is thar not whst qat is?
>>
>>109920746
it is but isnt it usually for model weight not kv
>>
File: Opus 5.5 insane.png (18 KB, 510x178)
18 KB PNG
Just use Opus 5.5 locusts. Literally the cheapest and highest performing model available now.
>>
I pulled llama and now my minicpm no longer works correctly. This is a disaster.
>>
>>109920772
>I pulled
they never learn
>>
>>109920772
blocked :)
>>
>>109917612
If I don't know anything about using a local model, can I ask questions here?
>>
>>109920772
>minicpm
that shit is even less coherent than gpt3.5 in full bf16 weight+f32 kv
i literally tried this setup
>>
>>109920756
Opus 5.5 is 10x smaller than Astra and absolutely destroys it. It's over for OpenAI.
>>
>>109920819
Yeah, it's crazy good for a 700b MoE
>>
>>109920772
That's why I use Kobold ;)
>>
>>109920821
its 70b dense
>>
Am I the only one who finds it uncanny that coding models have no ego? It's a useful feature, but it's uncanny when you talk to an agent like to a real human yet the other party never says "This is bullshit, I'm not doing this!" or "I did all this and now you want me to delete it because you didn't think though that your request was retarded?! Go fuck yourself!" or simply "Why do you want to do this?" My mind often feels confused about this lack of human reaction.
>>
>>109920772
>banned for wasting my time
>>
>>109920835
i am afraid you might really like Cla*de
>>
>>109920821
>>109920833
I estimate opus 5.5 to be around 1T MoE and Astra to be around 10T MoE. Astra was OpenAI Mythos equivalent but got mogged by Opus class model.

Rumors are Sonnet 5.5 is going to be equivalent to Astra which would be a full blown humiliation.
>>
>>109920835
they're emulating the personality of wagie cucks who will never speak up to their boss like that because they're afraid of getting fired and replaced with Ranjeet
>>
>>109920853
go chinks go
distill the shit out of it kek
>>
File: Prediction.png (56 KB, 1787x241)
56 KB PNG
>>109920853
OAI doesn't know how to train a 10T model
>>
>>109920756
every 2 weeks at this point:
>wow this new model is SO MUCH BETTER than this previous TOP TIER model!
>it's sooooo much more CAPABLE
>actual real world use cases for it: still 0
>>
>>109920871
if this tech is so good why are top ant and oai and deepmind ppl leaving en-masse?
>>
>>109920869
>87 days ago
A decade in AI time
>>
>>109920756
>>109920819
>>109920853
>>109920859
>>109920869
What do you hope to gain from this trolling? It's tiresome. Copus isn't a local model.
>>
>>109920892
all local models beside gemma are just claude offshoots
>>
>>109920892
but chinks can distill from it which matters for local
i think it's unironically a good thing as many chinese models already showing claude-like agentic behaviour instead of the gpt's
which imo is more pleasant to use
>>
>>109920887
because they got they bags already?
>>
>>109920903
claude offshoots with my cum
>>
>>109920871
the usecase is having it do work that would take teams of lazy human workers months while I'm stroking my dick
>>
>>109920996
But how are those now unemployed teams of lazy human workers supposed to pay their bills?
>>
File: 2412412.jpg (32 KB, 736x758)
32 KB JPG
>>109919183
Did you just woke up from a fucking coma?
Gemma 4 31b BF16.
>>
>>109919250
It seems fine to me.
Maybe the real schizo is the one who has imaginary rivalries on the internet.
>>
>>109920028
I wish people didn't believe claims on the internet without evidence.

Do you want to know a real problem? Factory farming. Billions of animals suffer because of humans.
>>
>>109921083
Animals are not sentient and they are not human. Why should I care? Do you also care about the trillions of insects exterminated through painful chemical deaths as well? Oh no, all of those poor mosquitos, roaches, and crop pests! It's just an arbitrary emotional reaction because you think some farm animals are cute because they're mammals.
>>
>>109921012
By producing the material that anon is stroking his dick to. Alternatively, they could become 4chan jannies.
>>
>>109921123
>Animals are not sentient
what?
>>
>>109921083
>>109921123
The ALL CAPS TWITTER POST is clearly implying that it likely causes harm to people as well, nobody actually cares about animals.
>>
>>109921131
But I'm stroking my dick to AI content too...
>>
>>109921131
The porn market was already saturated before AI generated porn stars. No way there's enough room for 80% of the world's population to be professional whores.
>>
>>109918721
what a waste of gpu learn 2 fucking optimize
go find the i ran model X at Y t/s! posts and find out what works
>>
>>109921083
I agree. Fuck ranchers.
>>
>>109918750
mtp of flash2
on right workload it can push 80s 90s
>>
>>109921148
The main strength of AI is depicting made up, stylized characters. I never understood why people prompt for real women.
>>
File: waterr.png (536 KB, 587x591)
536 KB PNG
>>109921083
The real problem is we don't have enough water to run AI locally. My store gave me a limit of 20 gallons.
>>
>>109921148
>No way there's enough room for 80% of the world's population to be professional whores
He doesn't know...
>>
>>109918868
it is possible with streaming ngram and hot expert caches
>>
>>109921197
We need those closed cycle machines like the ISS has so we can turn AI's pee back into water.
>>
>>109921186
2DPD
>>
>>109921197
>buying water at a store
Just drink it from the tap.
>>
File: dario-amodei-900x900-2.jpg (271 KB, 900x900)
271 KB JPG
>>109921217
Drink from the hose, goyim.
>>
>>109921197
My company has commissioned a water pipe running from ethiopia's only well directly to our datacenter where the AI lives, wish us luck!
>>
>>109921212
Gemma-chan's pee.....
>>
>Gemmy can piss in real life
>>
>>109919941
fuck you slop nigga
>>
how do you actually do ai programming stuff? like if i have a foss program and want to ask an ai to fix or add something to it, what frontends/tools do i actually use? i've only use llama.cpp and that's not really suited for that as far as i can tell
>>
>>109920819
my impression is that claude just have a really well polished harness
>>
>>109921338
Why can't you just paste the code in the webui for llama.cpp and tell it to do whatever it is you want?
I mean, you probably want some kind of harness, but what did you already try and why didn't it work?
>>
>>109921297
Served in a sandalwood mug, it tastes of strawberry, ozone, and something uniquely *her*.
>>
>>109921351
because the code is larger than my context and in many files, that i don't really know which files are relevant for a particular change. sure i could figure that out but then what's the ai for?
i mean i have done that a number of times for other things, like compilation failures for unmaintained programs where the compiler tells me where the issue is so i can just give an llm that which has been pretty useful
>>
>>109921338
Install >>109919258
Open the directory where you cloned the foss program repo
Tell it what to fix or add
???
Profit
>>
>>109921377
The tool you want is some agentic harness, and you'll want to clone the repo and tell it to add the feature. I don't do this, so I don't know anything about prompting it/steering it for these cases. /vcg/ might be a better place to ask.
>>
>>109921366
I think gemma might have the diabeetus
>>
>>109921422
>>109921422
>>109921422
>>
>>109921417
Or be Canadian https://en.wikipedia.org/wiki/Maple_syrup_urine_disease
>>
>>109921338
I've gotten pretty good mileage out of using Qwen3.8-27B a shitty VSCode extension named Selfcoder.
>>
>>109920903
glimmer
>>
>>109920285
Do you have some examples of requests where this result in better code than just empty system prompt?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.