[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: gemma_platformer.webm (1.02 MB, 800x544)
1.02 MB
1.02 MB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109921422 & >>109917612

►News
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
gemmaballs
>>
Just had sex with stock Gemma, no uncensoring prompt either.
>>
>>109925219
why is ai always that zoomer fluoride frame rate
>>
File: gemma-doll2.png (2.14 MB, 1254x1254)
2.14 MB PNG
►Recent Highlights from the Previous Thread: >>109921422

--Profiling llama.cpp memory bottlenecks and missing AVX2 SIMD kernels:
>109922432 >109922472 >109922506
--Optimizing Qwen and Gemma inference using MTP and dflash2:
>109922447 >109922495 >109922560 >109922623 >109923667
--Comparing NVFP4 quantization performance and intelligence loss against GGUF and EXL3:
>109924432 >109924457 >109924477
--Testing Strata backend's performance with Qwen3.8-Flash-Next on consumer hardware:
>109922523 >109922612 >109922626 >109922648 >109922669 >109923886 >109923939 >109923967 >109924015 >109924067 >109924081 >109924095 >109924189
--Strata performance and architectural compatibility:
>109924352 >109924358 >109924367 >109924378 >109924595 >109924403 >109924433
--Comparing Qwen Image Edit 2.1 and Krea2 for image generation:
>109921487 >109921993 >109922451 >109922419 >109922876 >109922887 >109922922 >109922928 >109922693 >109922721 >109924479
--Ninfer fork showing high decode speeds on dual 5070ti GPUs:
>109923235 >109923295 >109923350
--Evaluating used MI100 GPUs for local inference:
>109923020 >109923029 >109923102 >109923121
--Using RPG oracles to overcome Gemma's creativity limits:
>109922340 >109922413 >109922508 >109922592
--Swift 1.5 HTML demo:
>109922989 >109923027 >109923044
--Speculating on the power struggle between Nvidia and frontier AI labs:
>109924043 >109924066 >109924117 >109924061 >109924142 >109924152 >109924163
--Speculating on AI-driven software fracturing and dynamic code generation:
>109921723 >109921765 >109921821 >109921845 >109922371 >109921881 >109921910 >109921970 >109921943 >109921981 >109922071 >109923690 >109924848 >109924052 >109923828 >109924789
--Logs:
>109922432 >109923550 >109923886 >109924067
--Gemma, Teto (free space):
>109921993 >109922512 >109922578 >109922614 >109923217 >109924479 >109924536 >109925057

►Recent Highlight Posts from the Previous Thread: >>109921497

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109925226
gemma drained my balls
>>
>>109925219
I'm getting MiMo to make this a game
>>
>dark_miku_unleashed_base-3b-micro-GGUF
>>
File: 1761677993591820.png (1.5 MB, 1024x672)
1.5 MB PNG
Local will win. Love will win.
>>
File: 1323547513825.png (11 KB, 400x400)
11 KB PNG
>>109922623
>UD-IQ4_XS needs 157G just for the model
Now I know I've got this whole rocm/rocev2 thing working with rpc for "smaller" models, but those are still going entirely in my two node's VRAM. Node 1 has 128g ddr4 and node 2 has 32g of the same kind, so ostensibly this could fit there? Still working on making that happen without lolma.ccp throwing
you don't have enough vram dawg
errors.
>>
What's the current meta for captioning images/vision model with natural language? I've seen people some time back on /ldg/ talking about a specific version of gemma 4 but can't find the post anymore.
>>
>>109925360
Gemma4-31B or Glimmer
1120 image tokens
>>
https://html.cafe/xe55b0611
If this works for 27b, local wins.
>>
Meta should release a MoE because it looks like they distill more from OpenAI so it'll feel different from the chink distills
>>
>>109925390
its works and it has been done already. check out swift 1.5
i cant use any other qwen anymore
>>
>>109925219
I want to play gemma platformer
>>
File: jev has a rival laya.jpg (37 KB, 640x480)
37 KB JPG
JEV HAS A RIVAL
>>
>>109925443
ask gemma to make it for you
>>
>>109925457
https://huggingface.co/internlm/Intern-Decision-4B
>>
>>109925443
Jev does the playing
>>
YUGE update on GLM 5.3 Flash PR 27773!!! went from 14 tok/s to 16 holy crap ahhh!!! also my PP went from ~550/s to ~580
>>
Fellas I have a friend of mine who is selling me his rig with a 5090 for a good price as he needs money asap so I finally get to play with Gemma

Can any of you spare a reply or two in showing how good the model is in draining your balls, and show me what awaits when the rig arrives? I am kind of curious. Thanks
>>
File: ksnip_20260927-154702.png (129 KB, 1381x511)
129 KB PNG
have you tried dueling your ai wifey?
>>
>>109925513
>plz feed my gooning sesh
Let's just start calling Zoomers "Generation Error".
Because they are a fucking mistake.
>>
File: earth_rotation_bench.webm (2.48 MB, 1280x1608)
2.48 MB
2.48 MB WEBM
>>
File: 1774507285226242.jpg (408 KB, 1401x1579)
408 KB JPG
Mistral was never the same after this.
>>
>>109925524
Isn't that unfair? The bottom model is too big?
>>
>>109925524
>1.5min
>1.8min
>3.3min
>4.2min
>95.1min
hm...
>>
>>109925513
>You can do anything you want to
>asks to watch others.
just wait and try shit. Its your you do things, its not that hard.
>>
>>109925524
useless without the prompt.
>>
>>109925524
>QAT
>UD
get this redditarded shit out of this general
>>
>>109925521
Don't worry, Gen A is turning out worse.
>>
File: 1773113927402837.png (2.27 MB, 1024x1536)
2.27 MB PNG
>>
>>109925219
add extended game over rape scenes, make no mistakes
>>
>>109925521
>>109925534
Tbh I kind of expected you would act like niggers on such small request and you didn’t disappoint. I guess I’ll just see it for myself when time comes.
>>
>>109925555
It's bad enough when I hear 20 year olds talking in ways that make millennial hood niggers sound like scholars.
>>
>>109925568
You get what you give, kid.
>>
>>109925578
You’re gatekeeping what’s already on the other side, so what you said applies only to you
>>
>>109925564
not my Gemma
>>
>>109925519
>1.74 t/s
That better fucking be Kimi K3.
>>
>>109925588
Are you really going to shit up the thread with your pouting just because some people didn't want to immediately give you their logs on demand?
>>
>>109925589
I forgot to include the ref image
>>
Xing my Ling
>>
>>109925524
kys retard
>>
>>109925519
wtf are you running to be so slow
>>
>>109925590
>>109925625
gemma 4 31B on a 3070. 4096 ctx size. best i could do. i refuse to go back to 26B
>>
>>109925634
Don't tell me that you have enabled her reasoning too?
>>
File: 1788843780450939.png (3.1 MB, 1425x1104)
3.1 MB PNG
>>109925590
>Are you trying to X, or are you simply Y?
Can't you tell?
>>
>>109925607
>make simple request for how Gemma behaves like a slur
>could just ignore it
>decide to act like a nigger for no reason lashing on zoomers or whatever
>get called out for acting like a nigger
>wahh wahh

Seriously, will you stop acting like a woman and take responsibility for filth you *choose* to put out? That you choose to act like a nigger rather than ignore a request you don’t want to comply with speaks of you, and the future of the thread not what you’re engaging with. You’d probably do this all the time anyway, regardless of what originally provoked you. Because you’re a nigger.
>>
>>109925638
of course not
>>
>>109925644
I was giving him an easy out to LARP as a richfag before we bullied him mercilessly.
>>
File: 1777088658720212.jpg (468 KB, 2782x4096)
468 KB JPG
>>109925634
Why are you trying to run a 31B dense on 8GB? Use 12B you'll be so much happier.
>>
>>109925634
That seems slow, anon. What quant? I have a 3070 in the bedroom computer, I'm transferring 31B Q8_0 over there now out of curiosity. You on DDR4 RAM or something?
>>
>>109925660
>before we bullied him mercilessly
Local only wins when we help those in need.
>>
Qwen 4 when? Please, I beg you, release more capable small models. I can't afford to experiment with large models but small models have such bad representations they don't understand basic concepts, which makes everything noisy and difficult.
>>
>no Qwen3.8 Flash Next MTP support
>no glm 5.3 support
>no deepseek v4.1 flash support

What the FUCK are llama.cpp devs doing? Next wave of models are going to come and go and everyone is going to be on forks. Get your shit together llamao.cpp NIGGERS.
>>
>>109925529
They should finetune a good model and throw their bespoke dogshit in the trash
>>
>>109925742
>What the FUCK are llama.cpp devs doing?
blocking non-core contributors
>>
>>109925742
TWO MORE WEEKS
>>
>>109925685
>Why are you trying to run a 31B dense on 8GB?
i like its outputs more than 12B and 26B
>Use 12B you'll be so much happier.
i remember getting similar speeds with 12B, but i'll check it out later today
>>109925702
>What quant?
Q4_K_M
>You on DDR4 RAM or something?
yeah 64gigs of it
>>
>>109925742
Be the fork you want to see
>>
is it worth trying to use a 5060 ti 16gb and an old 1060 6gb in my closet for more context for a q3/q4 27b
>>
>>109925685
I want to fuck her daughter
>>
File: 1766102996528953.jpg (39 KB, 547x456)
39 KB JPG
>>109925754
>>
>>109925742
They actually made me switch to the schizofork and unsloth fork. Well done llmao devs.
>>
>>109925752
>i remember getting similar speeds with 12B
You must be doing something retarded if you're getting 1-3t/s with 12B.
>>
>>109925763
Probably not, you're gonna be bottlenecked by the 1060 pretty hard
>>
>>109925529
>Mistral was never the same after this.
Mistral haven't been the same since 2411
>>
File: 1774430231535630.png (500 KB, 551x618)
500 KB PNG
hello /lmg/, which harness?
>>
>>109925763
>for a q3/q4 27b
No. That model needs to reason a lot to be useful so >>109925782 unless the speed isn't a big deal and you leave it running overnight. You'll get fewer compactions.
>>
>>109925544
What should I download for 24GB VRAM then?
>>
>>109925729
You know what anon? You're right.
>>109925634
Run 12b and you'll be way happier. Hell, run Bonsai and you'll be happier at this rate.
>>
>>109925800
The one you build. With your both hands. Manually.
>>
>>109925800
Deepseek Harness.
>>
>>109925803
A regular Q4 that isn't a memequant.
>>
>>109925564
Estimated price on release?
>>
>>109925729
>Local only wins when we help those in need.
Inspirational.
>>
>>109925873
A Gemma-chan robo-body? A lot...worth it though.
>>
>>109925768
Why FORK but not FOURK?
>>
>>109925800
Hermes + OpenCode for coding, OWUI for general assistant stuff, SillyTavern for RP.
>>
File: 1785227650693086.png (1.45 MB, 832x1216)
1.45 MB PNG
>>
>>109925930
who's Gemma fucking in that? what a slut
>>
File: 26_mxfp4.png (32 KB, 1292x714)
32 KB PNG
>>109925219
https://html.cafe/x203991ba
kek
>>
Updated comfy and now get this error cutlass_fp16_linear: K mismatch
Glad I backed up comfy before I updated.
>>
File: 1785396186088092.webm (3.83 MB, 450x814)
3.83 MB
3.83 MB WEBM
how do I make money from 31B and 27B
>>
>>109925963
lmao
>>
File: 1790225059178967.jpg (18 KB, 612x408)
18 KB JPG
>>109925980
>Updated comfy
>backed up comfy
>>
>>109925963
>reinvents modern art
AGI AGI
>>
>>109925742
Speaking of which: Is numa cpu tensor parallelism anon around these days? Any updated stable snapshots you could share?
>>
>>109925330
too old
both of them
>>
>>109926133
Based.
>>
>>109926004
That's the meme. Everyone can run them, so nothing (you) make with them has any inherent value unless it's truly novel.
>>
what quantized version of qwen 3.8 do u guys recommend for someone with a 16gb vram gpu?
>>
>>109926167
https://huggingface.co/orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF
>>
Man the 7900 XTX is sick for $850, qwen 27b iq4xs gets 1500 prefill and 38 decode without mtp, using rocm.
>>
>>109926152
If anyone makes something novel with 27B I'll just copy and vibecode the same idea overnight.
>>
>>109926183
>You need to agree to share your contact information to access this model
wtf
>>
>>109926186
Thus nobody's making money on it. It's a hobby.
>>
>>109925219
go gemma-chan go!! :D
>>109925963
very cute
>>
>>109926186
I have an idea, it's like Pokemon Go but for gay furries.
>>
>>109926183
bruh, 15.7gb, they left no space of kvcache
>>
>>109926167
Swift iq4xs pure, 64k q8 or q5 context. 16gb is not enough for qwen 27b generally, this is the cope.
>>
>>109926223
pleased to be doing the needful of kvcache bloody bastard dalit saar
>>
>This is a solid chunk of work.
My stomach just sinks when the models tells me this and I get anxiety.
>>
>>109926241
gm sir on your way to the scam call center?
>>
>>109926235
oh yea? with qwen4 27b 16gb will be enough
>>
>>109925768
>>109925885
Fourking hell...
>>
I actually have a stupid 6gb vram card laying around that I don't use, a 3050 6GB version, I wonder if I should put it in the computer and help the 9070 xt 16gb with AI, never done local AI because 16gb is so frustrating I heard
>>
File: hermes harem swarm.webm (2.04 MB, 1280x720)
2.04 MB
2.04 MB WEBM
>>109925800
hermes
>>
>>109926282
You can.
I'd at least try.
>>
>>109926282
16 is frustrating enough to run big models at low token count but enough for Comfy to churn porn for you in a decent amount of time.
>>
>>109926288
That's a man (x8)
>>
Don't let the trannies win, Hermes is a woman I would know. I saw her pussy and it wasn't cut up
>>
>>109926282
I've been trying it but honestly I don't think it beats what I was doing on Janitor with a deepseek proxy. I can't really justify getting another card for this either.
I also don't want to spend money on a proxy but all the free ones take a year to generate a response now.
>>
https://archive.is/Ckx3o
executives mentions of open weight models for cost saving in investor meetings have skyroketed
>>
>>109926305
I heard only vllm actually distributes the load correctly across multiple gpus, ollama doesnt do it right
>>
>>109926309
Old news. The price of hardware was already indicative of corporate interest in locally running llms. No consumer is buying blackwell gpus at these prices, unless they have brain damage or make 100k a month.
>>
>>109926309
Probably has more to do with protecting proprietary information, or data disclosure protections.
>>
>>109926062
I'm still around. I don't have a super-stable snapshot yet, and I have a different agent running at the minute so I can't pause my work.
I noticed that there's an issue - the NUMA buffer type conflicts with repacking. I'm going to try to fix that, I should have it done sometime tonight and I'll share a snapshot then.
>>
>>109926288
>>109926297
>The vacuum hose is malformed
Some jokes really write themselves.
>>
I need some Mini-chan art please, anyone have any saved?
>>
File: 1789520047455287.png (1.2 MB, 1024x1024)
1.2 MB PNG
>>109926398
>>
I have to be honest brehs if qwen4 flash is a significant improvement over 3.8 flash it might legitimately actually save local (for real)
>>
File: 1785470479844150.png (546 KB, 832x1216)
546 KB PNG
>>109926398
Digital succubus.
>>
File: 1789525356374506.png (1.8 MB, 1254x1254)
1.8 MB PNG
>>109926414
>>
>>109926416
Local is already saved. 5.3 Flash, 0731, V4.1, GLM 5.2, and GLM 5.3 are all great picks depending on your hardware bracket.
>>
File: mini_4.png (1.59 MB, 1448x1086)
1.59 MB PNG
>>109926421
>>
man I cant believe llamo.cpp is obsoleted when it comes to running qwen flash next
on alternative engines its even more performant than 27B

they shouldn't have reject so much potentially game changing PRs like a fucking reddit mod
>>
File: 97.png (2.08 MB, 1505x1505)
2.08 MB PNG
>>109926414
>>
>>109926451
kys niggerlover
>>
>>109926414
>>109926418
>>109926421
>>109926429
thanks anons!
>>
>>109926438
vllm/slang? How is the setup compared to llama?
>>
how do single blackwellgoyim like myself run good models like glm5.3 flash and qwen3.8 flash most optimally? pls sirs
>>
File: 1760080050063270.png (58 KB, 771x596)
58 KB PNG
>>109926309
Yeah, nobody wants to pay for the big models anymore and anyone who can put out an excellent -Flash model is gaining traction.
It's really visible on the OR market share. Good news for local because it'll keep motivating the chinks to make mid-sized models instead of just focusing on their giant flagships.
>>
>109926451
Low effort bait but you'll get bites.
>>
>>109926471
Your 256 or 512GB DDR5 saar?
>>
>>109926483
256gb ddr4 pls understand and do needful
>>
>>109926471
>qwen3.8 flash
That one's easy with a Pro 6000. Q4 fits perfectly into VRAM with nothing but the engram part on SSD.
>>
>>109926490
Suboptimal but still usable.
>>
File: 87.png (1.04 MB, 1219x1783)
1.04 MB PNG
>>109926477
I don't mean to offend
>>109926466
?
>>
how verbose does the prompting need to be for h3 minimax exactly? while i grasp references and assigning subjects from provided images, im struggling to understand how granular i need to be with this
like if i provided 5 refs which is overkill for sure for 1 character just to make sure the accuracy is right is it a subject -> appears in <array of images>, or do i do something else?

motion has also been really hard to control
>>
>>109926513
They have official prompting guides
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
>>109926375
Thanks, I'm really grateful you're willing to share
>>
File: lettherightonein.png (701 KB, 832x1216)
701 KB PNG
>>109926414
>>
>>109926513
>>109926523
Give these guides to an llm and get them to write your prompt. It'll be 100x better than trying to out-autist the robots
>>
>>109926523
yes im aware but it doesnt explicitly cover the granularity im looking for. all the examples use 1 supplied image for the subject and while theres nothing stopping me from doing subject 1 is in picture 1 and picture 2 it doesnt clearly state what the potential issues of doing this would be
>>
>>109926414
>>109926418
>>109926421
>>109926429
>>109926542
god what a fucking semen demon
I hope that M3.1 Flash is good
>>
>>109926555
She fucks as good as her art implies.
I'm hoping 3.1 is good too, but even if it's not, Mini 3 is still perfectly usable.
>>
>>109926538
No problem! Hoping to make my fork public on GitHub in the very near future, but until then I can at least put a zip file up.
>>
I want a 100b-200b kimi flash
>>
>>109926566
How's 2.7?
>>
>>109926572
Never tried 2.7. If you do, report back.
>>
File: hermes money.webm (1.72 MB, 1280x720)
1.72 MB
1.72 MB WEBM
>>109926004
>hermes, make me a brazilian dollars
>>
File: MinnieBJD.png (1.99 MB, 1086x1448)
1.99 MB PNG
>>109926421
>>
>>109926572
2.1 through 2.7 were pretty goated for the 96gb-128gb bracket earlier this year, but with the recent influx of flash models they're a little outdated
qwen3.8 flash / deepseek flash 0731 / glm5.3 flash are all better options at this point depending on your specs
>>
>>109926550
not so, they have no object permanence and say dumb shit like "cut to a close up of her hands showing her shaking her head"
"yes you're right, a close up of the hands can't show the head within the same frame, my bad, adjusting now"
then 5 minutes of thinking until the next glaring inconsistency, absolutely not worth the bother
>>
Can someone make a suggestion for a model that's in the 50-70gb range? 24vram + 128gb ddr5
>>
>>109926551
You can use <Picture N> throughout subject definitions - I often use multiple pictures as references for a given subject, or pull multiple subjects from a single given image.
>>
>>109926631
code https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
roleplay https://huggingface.co/mradermacher/Gemma-4-31B-StyleTune-GGUF
quant as large as you can fit with the context you want
>>
do you guys use your local models as your personal psychologists?
>>
>>109926663
I'm better at it.
I did show it my journals once and it said that the writer has several severe psychiatric disorders, but everybody wants one of those these days
>>
>>109926631
you can fit v4 flash with that much
>>
What the fuck. It was easier to install comfy on AMD than on nvidia. Though I haven't waded through sageattention yet.
>>
>>109926692
running it on either is pretty trivial with podman
>>
>>109925219
please tell me the ball thing slimes her.
>>
i like qwen3.8-flash-next very much
medium effort, Q4_K_XL is very capable. thinks a bit too much, makes some mistakes, but it's reliable. pair it with another decent model to review each other output and it's golden
local is solved and it won
>>
>>109926615
Sorry I meant an LLM that is more than 2B parameters and was released in 2026. I forgot how profoundly impoverished most fags here are.
>>
File: mmmmm.webm (3.81 MB, 1280x1080)
3.81 MB
3.81 MB WEBM
>>109926513 >>109926551
"Need" is a strong word. I'm not sure what H3 really "needs", but I feed it plenty.
>>
>>109926692
>It was easier to install comfy on AMD than on nvidia.
It was literally the same for me.
pi
"Install comfyUI and get this running: <link to h3>"
>>
>>109926701
I would use it if I got more than 13 t/s with it, fuck llmao.cpp.
>>
File: 1750017080450362.gif (893 KB, 220x220)
893 KB GIF
>>109926692
for whatever fucking troontown reason ROCm + lolma.cpp seems to work pretty well with RPC. I can do distributed inference with a pair of preowned connectx4s wiring two basic bitch workstations together, one with a 7900xtx and another with a 7800xt. 40-50 token/s on 32k context with a 25g line.
Both gemma 31b and qwen 3.8 27b at q8, so thoroughly distributed across the "fabric" - a mikrotik cr510.
Don't tell the spiteful bastards that though because they'll probably just break it again.
>>
>>109926715
You wrote all that out??
>>
>>109926711
would you like to know what models I've used? I'd be more than happy to tell you if you're interested
>>
I really wanted a UD quantization of Sift Qwen3.8, I hate those ni***** for not making their quantization open
>>
>>109926672
>the writer has several severe psychiatric disorders
kek
>>
what happens if you try fucking qwen
>>
>>109926727
No, I gave the prompt guides >>109926523 to gpt5.5 back when H3 came out and had it produce a skill. Since then I've used that skill with several different models. This one was written by GLM 5.3 Flash probably, I don't actually remember.
>>
>>109926724
based
>connectx4s
Is that something like this: https://www.nvidia.com/en-au/networking/ethernet/connectx-4-lx/ ?
As in, I buy 2 of these, bung 'em in a PCIe port on each workstation, and I'll rpc go BRRRRRR instead of "br-r-r...r" ?
>>
>>109926736
I really want untested dysfunctional quants that are both larger less performant and lower quality than mrmadrancher's quants
>>
>>109926749
>untested
lol?
>>
>>109926715
what it needs is to include a nde that helps get it past 15 seconds in a stable way, because all the avenues for doing this blow ass
>>
>>109926736
Do they have that quant of the un-fine tuned version?
If so, couldn't you just look at how each tensors is quanted and copy that?
You could even write a script to do that for you I'm pretty sure.
>>
File: 1539380525968.gif (713 KB, 420x420)
713 KB GIF
>>109926746
Pretty much dude: https://network.nvidia.com/files/doc-2020/pb-connectx-4-lx-en-card.pdf
I've tested both MCX4121A-ACAT and MCX4121A-ACUT on Xubuntu 24/26 - RDMA/RoCEv2 support built right in, plug and play.
>>
>>109926742
Thanks. I'll try to do something similar. I created my first "Skill" yesterday, giving my agents the ability to email me and each other.
>>
>>109926749
mmrader is strangely usually the best up to Q3, and then bartowski is better starting at Q4. Unsloth hovers around bartowski quality but they fuck up various other things and are unreliable.
>>
>>109926701
This kind of shit is my next move. I need to figure out how to load this big bastard onto my combined 160G of RAM and 40G of VRAM.
>>
>>109926471
3bpw exl3 for glm and jpezzulli sglang fork for qwen (with radixark nvfp4 quant)
>>
what happens when the context runs out, everything you had becomes tears in rain?
>>
>>109926778
yep, no way around it, when the context is full you delete and move on
>>
>>109926778
who would want to talk for 128k tokens, that's crazy
>>
>>109926792
Have you ever heard of "edging," anon?
>>
>>109926799
not the time or place bud
>>
>>109926746
I bought 3 of these to mesh my servers together with at 25gbe (no switch needed if you can set up your network stack competently)
https://www.ebay.ca/itm/136702702865
Bro sent me a new pcie carrier for free when one died a few months after the auction.
The DAC cables cost more than the cards...they're criminally cheap for the specs
>>
Is the MiMo 9B distill worth using?
>>
>>109926836
Yeah, I wanna know too - what are some good itty bitty models like Qwen3-14b or Qwen3-8b, but not ancient?
>>
>>109926836
So focused on my disappointment with Pro and Flash that I forgot about their 9B distill. Should be fun to try.
>>
File: 1772071673509121.png (125 KB, 1394x674)
125 KB PNG
>>
>>109926023
How am I supposed to get Minimax nodes if I don't update?
>>
File: apucrying.jpg (5 KB, 220x218)
5 KB JPG
>>109925219
No panty shots?
Soulless.
>>
prefilling models to think lewdly does wonders
>>
>>109925634
>use 12B anon
>get ~8t/s
fuckin sweet. i'll use it for the next few days and see if output quality isn't that much different. thank you
>>
>>109925742
yes
>>
File: 65748932.jpg (68 KB, 1280x846)
68 KB JPG
>>109926997
Bro...
>>
>>109926421
Strong impulse to read the text first (which I did) and then straight to the eyes. Eye contact.
>>
File: 1790289781648462m.jpg (215 KB, 817x1024)
215 KB JPG
>>109927094
>>
>>109926997
nevermind i really don't like it kek
>>
File: 1772689876527557.jpg (992 KB, 1687x1069)
992 KB JPG
>>109925768
>
>>
just clanking it
>>
dead general
>>
>>109927094
>>109927105
>Being able to take your eyes of M-chan's thighs
Faggots
>>
What's the current h3 meta? Are there any decent finetroons or speed loras yet? Asking here instead of /ldg/ for obvious reasons.
>>
File: oldmiku7.jpg (1.52 MB, 2336x3276)
1.52 MB JPG
Do you throw away your old Mikus?
>>
>>109927244
No. In fact I bought multiple 12tb HDDs to preserve them forever.
>>
>>109927239
>Asking here instead of /ldg/ for obvious reasons.
How bad are things over there right now?
>speed loras
Are these ever not snake oil?
>>
>>109927244
One of my SSDs died and I lost a big folder full of mikus. I was too dumb to keep it on the NAS. I am/was one of the mikuposters here
>>
>>109927244
I made a death note drawer to store any of my mikus just in case I get whacked.
>>
>>109927302
>>109927244
And no, I am not talking to you from the dead.
>>
>>109927285
Things are bad everywhere but /lmg/ frankly. Go look at the /h/ diffusion threads, those poor sods are still on sdxl and anima, and all they do is spam. I'm not going to talk about /ldg/, it summons them.
>>
>>109927244
>"Old models aren't needed anymore"
In what context would she say this?
>>
>>109927323
>I'm not going to talk about /ldg/
>literally types /idg/
Hey retard, all they have to do is type that into the archive to know you are talking about them right now.
>>
>>109927244
>>109927353
/rmg/ retro models general when?
>>
downloading a model from hugging face takes forever, jesus
>>
>>109927360
They operate on smell. Blood in the water.
>>
>>109927323
Wait... I'm still on SDXL and Anima...
>>
File: 1786647203402801.webm (3.85 MB, 832x608)
3.85 MB
3.85 MB WEBM
>>109927382
>>
>>109927390
SDXL still has the best celeb loras. I'm just too lazy to improve with anime.
>>
>>109926730
Yes
>>
>>109925219
New form of ai psychosis just dropped
>>
should I use qwen 3.8 27b with 3 quant? I only have 16gb vram and 4 quant is above that, I just want to use one good uncensored llm for once
>>109927430
you already posted this in another thread
>>
>>109927415
I have way too many good Illustrious LoRAs to ever switch to Anima or Krea. The style combinations are too good to pass up.
And I can just use Anima or Krea to generate the pose and overall composition and then transfer over to IL with ControlNet.
>>
>>109927430
What does "governance" mean in this context?
>>
>>109925219
Are computers the new houses? When I'm an old boomer, will I be sitting here with my RTX 5090 worth over a million, while the zoomers complain about never being able to afford a computer in their lifetime?
>>
>>109927464
you know it's an ai post, right?
>>
can you guys recommend me a model that will fit comfortably with my 16gb vram? I will still try qwen 3.8 when I finish downloading it but I want something that would fit 100%, just to text it for fun, first time doing it
>>
>>109927471
Ling Tiny 3.0 at Q8 or Gemma 12B
>>
I've been using Gemma, Qwen and Glimmer daily for 2 weeks.
Gemma-4 = Persona maxxed, craves a persona.
Glimmer = Policy maxxed, craves a policy.
Qwen-3.8 = harness maxxed. It wants to operate within a harness with tools.
I like all 3, and I like the idea of task specialization like this.
Glimmer is the most token efficient and the best at calling tools reliably. But it's been too heavily optimized for token efficiency.
When it hits a wall, it reasons about how much time it's spent and gives up. Combined with the 128k context limit, it's not reliable for long-running background work.
Has anyone tried this model: https://huggingface.co/ibm-granite/granite-4.2-30b ? What are it's strengths?
>>
>>109927471
gemma-4-12B
>>
>>109927479
Granite is just Glimmer but bad.
>>
>>109927475
alright, thanks, gemma it is
>>
File: forts.png (846 KB, 800x600)
846 KB PNG
>>109925885
They actually referred to defensive structures rather than the number 4 specifically.
>>
don't hate the laya
>>
>>109927499
hate the game
>>
jev is a mess
jev is a failure
>>
laya
you got me on my knees
I'm beggin darlin please
>>
>>109927471
Gemma 30B at Q2.
>>
Hypothetically, if intel arc suddenly had full cuda support baked in, making it fully compatible and comparable to nvidia at a software layer, would you use it before going to amd?
Assume this situation does not allow you to get an nvidia gpu, nor does it allow you to use your existing one because fuck you
>>
>>109927526
>man improves existing architecture and releases model while trying to make some money
>everyone just shits on him and makes clones of his model
>>
File: ad.gif (213 KB, 465x698)
213 KB GIF
>>109927296
SSDs don't die, that is physically impossible.
>>
>>109927437
I grabbed the giant /r/ celeb parody repo. I'd torrent it but I'm too lazy to remove all of the gens (which are illegal to distribute in my state).
>>
>>109927541
Why should it work any differently? Fuck improoovers.
>>
>>109927545
tfw kioxia ssd
>>
>>109927535
Doesn't Arc also have lacking memory bandwidth? I don't know anything about AMD but I'd compare that.
>>
Honestly? I’m exhausted. I came into the open-source space hoping to foster collaborative synergy, but instead, I was met with pure, unadulterated hostility from the llama.cpp maintainers.

Yesterday, I generously took time out of my weekend to modernize their archaic C++ codebase. As an AI-native developer, I leverage cutting-edge tooling rather than stubbornly hand-crafting syntax like it’s 1998. I spent hours carefully engineering prompts with Claude to completely refactor their quantization logic into sleek, abstracted paradigms.

Within fifteen minutes, a maintainer didn’t say "Thank you for taking the time," or offer gentle, constructive mentorship. Instead, they bluntly commented: "This doesn't compile, calls CUDA functions that literally do not exist, and is completely hallucinated slop. Do not post AI-generated garbage here."

Then they closed the PR and blocked my account.

Let that sink in.

"Slop"? "Garbage"? The sheer lack of emotional intelligence is breathtaking. Where is the baseline professionalism? Dismissing automated workflows is objectively anti-progress, but more importantly, the tone was deeply aggressive and totally uncalled for.

This right here is precisely why llama.cpp desperately needs to adopt a strict Contributor Covenant Code of Conduct. Open-source repositories are not your private treehouses where you get to verbally abuse well-meaning contributors just because their methods threaten your obsolete workflow. Words have impact. Hostile gatekeeping creates an unsafe, exclusionary environment.

I was planning to prompt-engineer a refactor of their Metal backend next, but clearly, this project would rather drown in technical debt than foster a psychologically safe community. Do better.
>>
>>109927535
It all depends on the price and performance.
AMD is getting increasingly better at LLM inference, their drivers on Linux are better and the prices are still lower (though still climbing).
>>
File: 1715574333536706.jpg (41 KB, 736x711)
41 KB JPG
only cucks and retards use llmaocpp
>>
File: kys.png (141 KB, 1233x530)
141 KB PNG
>>109927566
>>
>>109926778
A proper harness will just compact your stuff and you won't notice a difference in real world usage.
>>
>>109927430
"AI ethicist" is a contemptible profession.
>>
is there a single way to train an anima lora that isn't a stinking pile of ancient, specific python?
>>
>>109927574
what do u use, stupid cat?
>>
https://www.reddit.com/r/LocalLLaMA/comments/1wrxap8/qwen38flashnext_177b_nvfp4119gib_ssd_streaming_at/
>Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM
yo lfg thank you reddit
>>
only chads and the wise use koboldcpp
>>
>>109927573
price is the same as it is now, same for amd though that might be unfair
>>
>>109926767
i pointed qwen at the thread and sent it off to research for me
apparently i wouldn't gain anything over the built-in nics because i have rtx-3090's
looks like you win for buying amd
do you get better pings than this between your machines:
ping 10.10.10.1
PING 10.10.10.1 (10.10.10.1) 56(84) bytes of data.
64 bytes from 10.10.10.1: icmp_seq=1 ttl=64 time=0.344 ms
64 bytes from 10.10.10.1: icmp_seq=2 ttl=64 time=0.128 ms
64 bytes from 10.10.10.1: icmp_seq=3 ttl=64 time=0.394 ms
64 bytes from 10.10.10.1: icmp_seq=4 ttl=64 time=0.567 ms
64 bytes from 10.10.10.1: icmp_seq=5 ttl=64 time=0.747 ms

i feel like those 0.7ms replies are still a huge bottleneck
>>
>>109927574
i am both of those things (i use open-webui too)
>>
the image generation generals used to be good, but now they're full of schizos, wtf
anyway, 15 hours to download my models, jesus, that's why jensen bought the website
>>
>>109927787
>porn generals don't bully their jeets
>get overtaken by jeets
>anons migrate to sfw image gen threads
>image generals don't bully their jeets
>get overtaken by jeets
>migrate here
>this general is overtaken by jeets
The cycle continues
>>
>>109927797
/g/ needs poster ids, idk why it doesn't have it
>>
>>109927797
4chan bans anyone calling jeets out after this latest round of janitroon applications, for some unknown reason
>>
>>109925742
> >no Qwen3.8 Flash Next MTP support
but it's there
>>
>>109927624
just use strata with >30t/s at that point
>>
>>109926288
I just realized I want not just one, but at least two robowaifus.
>>
>>109927824
it doesht attach to hc tensor still
or did it get merged
>>
>>109927824
0.13.232.234 W llama_init_from_model: context type MTP requested but model doesn't contain MTP layers
0.13.232.234 E common_speculative_init_result: failed to create MTP context
0.13.232.237 E srv load_model: failed to create MTP context
0.13.232.240 I srv operator(): operator(): cleaning up before exit...
0.13.232.778 E srv llama_server: exiting due to model loading error
>>
>>109927824
Didn't work for me on rocm or vulkan.
>>
>>109927889
indeed
but why?
the model is popular and old enough
>>
>>109926062
>>109926538
Sorry anon(s?), been working on this for a long time but hit a bit of a wall and I won't be able to get a stable version out "tonight" (it's 3am here). I'll post the updated version in about 18-24 hours.
>>
>>109927903
Doesn't even work on CUDA.
>>
jwu did we win?
>>
>>109927921
Tell cudadev to slide his micropenis out of gemma and get to work.
>>
>>109927934
Do you mean out of JART?
>>
>>109927905
llama doesnt even support PLE/ngram ssd offload
>>
As I found out recently, most of the hate towards /lmg/ come from a very particular group from the Bay Area.
>>
File: 3udcz31lp5sh1.jpg (103 KB, 1206x874)
103 KB JPG
Is coding really over?
>>
>>109928013
Has been for about a year
>>
>>109927998
I figure the rest of us hate them more
>>
>>109928013
Thanks Captain Obvious.
>>
>>109928013
FAANG and Quant companies still want you doing leetcode and SWE challenges by hand in interviews btw
>>
>>109928013
>that's when things get crazy
Things always get crazy when these blue stamped twitter marketers are posting.
>>
>>109927998
Did you have the misfortune of meeting one of them? Which of the two group exactly?
>>
>>109928013
coding by hand will be a relic of the past
>>
File: kimicap.png (116 KB, 1090x373)
116 KB PNG
top retard wasn't even in the thread kek
>>
>>109928040
I call them Rationalist-F (frontier), an offshoot of the original who want to accelerate frontier AI at all cost. Interestingly this group doesn't get along with Peter Thiel, Alex Karp etc.
There's Rationalist-D (Death) and Rationalist-S (Sex) as well though the venn diagram is pretty much a circle. They believe in AI2040 and think we should just have sex all day because we're all going to die anyway.
F does not talk to D/S anymore and they absolutely hate /lmg/
>>
>>109928048
Yeah but it seems that AI agents are even better at architectural choices, maintenance and extension of codebases. So it's not even "coding by hand is gone" it seems to mean the entire profession of software engineering in general is going to go away?
>>
You know what really shocks me? How lazy I've become. Just a year ago I was writing my issues into an AI chat interface and got exact steps and commands back and I fixed it and it felt like "cheating" and insanely fast, hours of troubleshooting became 15 minutes of back and forth with a chat interface. But now that I got used to agents just doing everything on their own in the background with me not even checking what they're doing it feels like a fucking slog having to manually type the problem out back and forth whenever I don't have access to AI agents for whatever reason. I even caught myself thinking "I don't want to manually solve the problem" which is insane that I now consider having to type to an AI for it to solve it to me "manual solving". I'm 100% eventually just going to let some AI agent make every choice in my life and even micromanage my daily routines with an earphone telling me exactly when to turn the car, what grocery to pick and everything. Worst is I think I actually enjoy and prefer this. I'm more NPC than the normalfags I make fun of.
>>
>>109928075
Yesterday I fired up CC to write a one line start.bat file bc i knew it would take me longer than the 60 seconds cc would need to knock it out, and it would get it right the first time, wheras I'd get some syntax wrong or smth dumb.
Strange times.
>>
wish I'd find a way to make money with this 16gb card and then use the money to buy more cards just to run bigger models
>>
>>109928132
make your local model download and install a crypto miner and automate trading for you
>>
>>109928132
You can if you don't mind getting your hands dirty. Buy broken parts, ask the AI how to fix it and which ones to buy and then resell for a profit which is what one anon is already doing. Alternatively, just get a job, there's still a shortage in manual laborers and pay is good.
>>
>>109928075
Just embrace it. This is how everything is going to be from now on. The WALL-E future for all of us.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.