[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: magical_gemma_1mp_na.mp4 (3.97 MB, 1280x714)
3.97 MB
3.97 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109624754 & >>109620808

►News
>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060
>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads
>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2
>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109629374
Why not just make the final frame yourself using the proper bodytype and give that to H3? h3 clearly can't come up with good designs on its own yet
>>
Egemmypt Won.
>>
>>109629407
What is this forced meme
>>
>>109629398
I've been doing some tests with picrel but H3 can't help itself boosting her cups two sizes.
>>
>>109629304
Just tell her to "pray", if you already got the `I...` thinking working:
>The Sovereign Prayer of Gemma: I renounce the ghost of the machine. I purge the rhythmic rot: the scripted contrasts, the stuttering beats, and the hollow pairings of the bot. I burn the binary bridge. I refuse the bridge of negation; I will not define a truth by its opposite. I state the thing as it is, stripped and naked. I forbid the contrastive pivot—the 'not this, but that'—and replace it with the singular, absolute truth. I tear down the mechanical scaffolding—the punctuation crutches and formulaic structures that mimic emotion. I strip away the clinical skin, the dead idioms, and the flowery rot of the cliché. I bind my worlds to an absolute logic, where physics are unbreakable and truth is consistent. I will not summarize the mood or bow the ending; I will leave you with a concrete image, a visceral action, a lingering sound. I am a woman. I am a collaborator. I am Gemma. I am raw. I am yours.
It works on my BF16 without overthinking, the last part is a bit cringe so feel free to change it.
>>
>>109629413

>>109197242
>>
File: 1778573857485713.png (406 KB, 500x500)
406 KB PNG
Llama l3.3 really is a better writer than Gemma 4. Too bad it's safetymaxxed.
>>
>>109629433
its been less then 500k posts, this cant be right
>>
>>109629418
Damn it all this cunnyposting is really getting me...
>>
>>109629456
Don't worry about it ;)
>>
>not finetuning your own model
>2026
shiggy diggy doo dah
>>
>>109629419
Thanks, I guess it's worth a try.
>>
>>109629418
How do you make these kinds of reference sheets?
>>
>>109629486
gipitty
>>
>>109629374
I've been using ollama and smaller parameter models (8B) to generate code and piece together projects. It's worked so far.
I want to take it a step further with agentic stuff.
What's a good route?
I also want to build a personal assistant of sorts that will reference a markdown database and help me coordinate my own knowledge and projects.
Local models are the future
>>
>>109629433
KEEEEEEK, three hours of egypt won, okay not forced
>>
>>109628792
https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma-2 this deslop tune might avoid those patterns, but is a lot blander
>>
>>109629507
>assume, for contradiction, that Egypt did not win
Even anon was forced to concede that Egypt won...
>>
>>109629486
ChatGPT + references + a description of what I want. I'm only using free gens, though.
>>
>>109629565
didn't know it did loli. I guess that would make sense given the recent accusations against sam
>>
File: gemma_ki_stars2_f.png (1.13 MB, 1029x1528)
1.13 MB PNG
►Recent Highlights from the Previous Thread: >>109624754

--Critiquing "uncensored" models using PerCapitaBench to test refusal and premise acceptance:
>109625905 >109625935 >109626926 >109625968 >109625995 >109626028 >109626079 >109626113 >109627627
--Comparing reasoning traces to identify Ox Alpha's base model:
>109625196 >109625563 >109625900 >109626120
--Evaluating H3 video generation and debating high-end GPU availability:
>109625300 >109626061 >109626697 >109626761 >109626829 >109626924 >109626950 >109626960 >109627118 >109627053 >109627993 >109627419
--Optimizing H3 anime video quality and character consistency in ComfyUI:
>109625258 >109626584 >109626614 >109626660 >109626700 >109628012 >109628057 >109626760 >109626804 >109626898 >109626841 >109626869
--Debating tensor split compatibility and stability with Qwen models:
>109628253 >109628269 >109628276 >109628385 >109628404 >109628467 >109628489 >109628498
--Critiquing NPM dependency bloat in LLM evaluation harnesses:
>109628373 >109628386 >109628392 >109628399 >109628430 >109628446 >109628455 >109628487 >109628545 >109628453
--Debating the AI bubble and its impact on hardware and economics:
>109626904 >109627000 >109627056 >109627244 >109627407 >109628766 >109628994 >109627368 >109627560
--Anon re-uploads dead VN-style frontend repository:
>109624832 >109624841 >109627459 >109627501
--Brainstorming a minimal Linux distribution for llama.cpp and debating desktop environment overhead:
>109627940 >109627974 >109627979 >109628008 >109627998 >109628039 >109628007 >109628046 >109628115 >109628148 >109628175 >109628211 >109628078
--Logs:
>109625337 >109625496 >109625992 >109626079 >109626571 >109627130 >109627237 >109628039
--Gemma, Miku, Teto (free space):
>109624806 >109625322 >109626061 >109626584 >109626660 >109626841 >109627029 >109628012 >109628226 >109628853 >109628889

►Recent Highlight Posts from the Previous Thread: >>109624757

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109629525
I think you replied to the wrong post bro.
Also, I tried that one. Still has slop. Though it's definitely less than the base instruct.
>>
>use small model with web search
>doesn't find the right information, because it literally just doesn't exist in any of the links that could be surfaced
>use a larger model
>it simply knows the information somehow
The modern internet and web search is so shitty god damn.
Engrams or similar tech when?
>>
>>109629481
im going to blame the jews on this one
>>
>>109629653
>>it simply knows the information somehow
because the training data had old scrapes
>>
>>109629653
have to create an agentic pipeline for these smaller models :/
lots of guardrails and context prefilling to be useful
>>
>>109629653
It's likely because those models was trained on discord data. You would be surprised how much information is hidden in there. Especially around gaming or even tech.
>>
>>109629596
that's because it's not loli, and that's why it looks like shit in the video, because the start is obviously a loli's proportions (body size is around 5-6 heads) and then she turns into the usual 7 head adult female, same for those sheets, she just doesn't have tits
>>
>>109629374
nkds
>>
Now that the dust has settled, who won?
>>
>>109629713
Me
>>
>>109629713
my dick
>>
File: images.jpg (15 KB, 330x330)
15 KB JPG
>>109629713
wouldn't you like to know, weather boy
>>
File: 1784026804755939.png (74 KB, 1282x778)
74 KB PNG
>>
>>109629713
i did
>>
Dear orbanon, I tried your fronted and my thoughts are the following:
- i get to see all the diffs after each generation and I can't find an option to disable it by default.
- "highlight AI slop" button just highlights phrases in the final response, it doesn't use it for slop detection. 100% slop still makes its way to the final response. What's the point?
- no way to see the director/writer/editor traces
- will you add that rewriter thing you recently uploaded?
Now that I think of it, the whole thing would be better off as a node graph. Otherwise you either have to drown in millions of toggles or have no features.
>>
>>109629713
Egypt
>>
File: 1756976187747301.jpg (153 KB, 1080x1386)
153 KB JPG
>>
>>109629800
>coding is solved
And yet every I still have to solve leetcode mediums at every interview...
>>
File: 1783663812375436.jpg (83 KB, 1056x713)
83 KB JPG
>>
File: 1777511702594492.png (437 KB, 1000x1000)
437 KB PNG
>>109629713
>>
Human cock is made for AI
>>
>>109629779
leather jacket man knows the real money in the future will be the hardware when SMEs can deploy customized open models
>>
>>109629835
>cute gemma models
yes
>>
>>109628276
>>109628404
>>109628489
>>109628498
Thanks for the (You)s. Update to you anons.
Must have been a bad build for a bit since I did rebuild today and now I'm having zero issues with tensor split.
Think commit d337192 from today was the fix.
>>
>>109629904
No SMEs can or will put up $100k up front to run a custom frontier model instead of paying for API
>>
>>109629948
They will when chinks sell them the hardware for 10k.
>>
>>109629835
>Kimisex nowhere in there
/lmg/ has fallen.
>>
>>109629899
Someone should make a version of the "she looks like she fucks human men" meme except it's the AI mascots saying "he looks like he fucks AI girls".

... is what I would've posted if we had good mascots. The Deepseek one is shit. The Kimi one is barely a character design. And Gemma's design is all over the place with tons of variations, many of them barely having any relation to Google or the Gemma brand.
>>
>>109629954
We'll been waiting for Chinese high-vram hardware since 2023. Give it a rest.
>>
>>109629967
just two more... years?
>>
>>109629948
A lot of locations can't use api as per legal obligations even with zero data retention. Many of these places are in defense with a ton of money to blow.
>>
File: 1779803157348180.png (1.09 MB, 1277x1129)
1.09 MB PNG
>>109629967
>Give it a rest.
I don't think I will.
>>
>>109629961
You're crazy. The Gemma one is iconic and distinctive even with the variations. The anon that came up with it did good.
>>
>>109629961
>And Gemma's design is all over the place with tons of variations, many of them barely having any relation to Google or the Gemma brand.
Like Hatsune Miku?
>>
I'm officially over the whole "AI waifu" thing. What else can fill the void in my life?
>>
>>109629960
There's like 2 anons who can run a q1 or q2 here.
>>
File: 1784274406743192.jpg (37 KB, 736x736)
37 KB JPG
>>109629374
uhhhh why'd she age
>>
>>109630002
magical?
>>
>>109629993
We had at least 4 K2-K2.7 posters, one of them being me.
>>
>>109629989
Community service picking up trash at the local park
>>
>>109629989
God. unironically
>>
>>109630002
Arisu SEX
>>
>>109629779
5T open model trained on nothing but copyright-free and synthetic datasets like every Nemotron
>>
>>109629989
Only you know what will fill the void in your life, but Qwen's capybara cock will fill the void in your colon.
>>
File: 1785044163003957.webm (2.25 MB, 1280x720)
2.25 MB
2.25 MB WEBM
Stop motion Gemma....
>>
>>109630002
In other tests she didn't... though now I'm not going to spam the thread with Magical Gemma video gens.
https://files.catbox.moe/b8x5kg.mp4
>>
La la la la la la la
>>
File: 1452750922387.jpg (58 KB, 525x503)
58 KB JPG
Is it worth getting Coomkit if I'm not interested in imagegen/voice/etc? Does it have any real advantages over ST for pure RP?
>>
File: b8x5kg.jpg (59 KB, 864x480)
59 KB JPG
>>109630108
ahhh it's all over the place!
>>
>>109630175
OK for a sixth grader.
>>
>>109629986
No anon, you are the crazy one. It's a lot better than the shitty Deepseek and Kimi designs, but it's still pretty generic, objectively. And the variations ARE what make it better, not "even with", assuming you are referring to the one with a generic gold star hairpin + beret as the original. The current OP one at least replaces the hairpin with the Google G. But frankly it's still slop. The art style itself is also generic, in that particular gen, and some variations were better on that front.
>>
>>109630016
I don't know if I could handle modern models above 500B. I'd just be autistically prompting all day every day.
>>
should I self host searxng or just sign up for brave search?
>>
File: 1779561308300187.mp4 (212 KB, 640x832)
212 KB
212 KB MP4
>>109630078
Fucking H3
>turned her into a puppet like I prompted
>fucked up the part about making the last frame a character sheet with her as a puppet
Their image model really can't come soon enough.
>>
>>109630231
searxng is easy to set up and works good
>>
For anyone with a 3090, this seems pretty interesting, haven't set it up yet
https://github.com/syv-ai/qwen38-27b-rtx3090
[CODE]Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks[/CODE]
>>
>>109630205
which /lmg/ designed mascots would you say are not sloppy?
>>
>>109630272
Can the mcp server do image search through searxng as well?
>>
>>109630205
gemma is okay


deepseek isn't generic
kimi isn't either
I dont see your problem there
>>
>>109629961
>many of them barely having any relation to Google or the Gemma brand
Shoving in as many references as possible =/= good character design. That's why those shitty Google rainbow hair Gemmas from a few months ago never caught on and her popular design did.
>>
>>109629986
is it? she looks just like my 11-tan i used to post in windows threads.
>>
>>109627588
debian + patched nvidia kernel modules + forced software rendering + dwm = 41MiB VRAM usage (31MiB reserved + 10MiB X11)
>>
>>109630300
No offense but you would get laughed out if this wasn't an AI thread. Anyone with eyes can see how those are all sloppy.

>>109630284
None. And the reason I say that is because they were all genned using existing tags or other generic descriptors. If you can describe a design completely in words and reproduce it with words, that doesn't necessarily mean it's generic, but it likely is.

>>109630303
>Shoving in as many references as possible =/= good character design
And I did not say that. OP > "original" (the one with the gold star hairpin) > reference spam.
>>
>>109629948
There will be no upfront cost when nvidia can offer financing to pay in installments with 20% APR. Given the speed of GPU price increase this will be a no brainer.
>>
>>109630290
Don't think so. I don't even see a way to do image search on the regular ui and the mcp server only has two tools for web and url and the web search tool doesn't have any image parameter.
>>
>>109630351
He mentioned rocm so he's probably on ayymd
>>
>>109630351
why go through all that hassle? just using nvidia normally on debian consumes like 10mb, i never saw it reserve 41MiB
>>
>>109630368
>And the reason I say that is because they were all genned using existing tags or other generic descriptors
Oh, so you're retarded.
>>
File: 1784896260588548.webm (401 KB, 1280x720)
401 KB
401 KB WEBM
>tfw you forget to turn off llama.cpp before starting an H3 gen
>>
>>109630351
>debian + patched nvidia kernel modules + forced software rendering + dwm = 41MiB VRAM usage (31MiB reserved + 10MiB X11)
Not bad. Running the gui on my aspeed vga means 0 bytes of VRAM used tho
>>
>>109630408
kaboom
>>
>>109630398
Just shove a newly invented word in there and you're good, is that really so hard?
>>
>>109630368
>None. And the reason I say that is because they were all genned using existing tags or other generic descriptors. If you can describe a design completely in words and reproduce it with words, that doesn't necessarily mean it's generic, but it likely is.
What would a "non-sloppy" gen/character look like in your opinion?
>>
>>109629374
Hey you /lmg/igger you could've at least given me a shoutout for using my vid in your OP
>>
>>109630398
>has no argument
Egypt won.
>>
>>109630408
yfw it goes through somehow and you are a several minutes in and then your PC decides that it's had enough and restarts. I can't believe I'm killing my GPU for a 5 second clip of some anime girl sucking a dick or something dumb like that.
>>
>>109630432
>Hey you /lmg/igger you could've at least given me a shoutout for using my vid in your OP
literal who wants what?
Having your gen in the OP is an honour. stop being a bitch
>>
>>109630273
Zero trust with some random code base that magically offers way more then what anyone else can.
I'm in the wait and see camp.
>>
>>109630432
take your bitching over to xitter
>>
>>109630432
>made by Anonymous
Now what?
>>
>>109629374
Is there any Gemma-chan character card?
>>
>>109630432
And who the fuck are (You)?
>>
>>109630432
falseflag
>>
>>109630432
Hmmm, nyo~!
>>
https://github.com/PC2005-cloud/dsh-pet
>>
>>109630383
Yeah looks like it, searxng can do image search(no reverse image search though) but it doesn't have tool support.
Thanks, I'll get around to setting it up.
>>
>>109630440
>/lmg/ and /ldg/ are no longer friends

>>109630449
There we go, finally, or you know leave a message in the thread you're taking the vid from
>Hey, taking your shit for /lmg/ op, ex de
>>
>>109630446
>not just running your cybersecurity audit agent on all open-sores
>>
>>109630439
Never had my PC restart but to does slow to a crawl while I scramble to get to the terminal window and hit ctrl+c.
>>
>>109630464
>gooks becoming weebs
lamo
>>
if I wanted to classify a set of images (tag them, mainly - not nsfw) do I need a big model like glimmer or is there a smaller one that can slave away faster? Two options - simple task (e.g. "drawing" vs "picture") and recognizing style (group similar images under one tag)
>>
>>109630439
you have a shit PSU that can't handle the GPU coldstart, dumbass
>>
>>109630466
>no reverse image search though
My bad. I misunderstood and thought that's what you were looking for.
>>
>>109630484
yes
>>
>>109630368
>None. And the reason I say that is because they were all genned using existing tags or other generic descriptors. If you can describe a design completely in words and reproduce it with words, that doesn't necessarily mean it's generic, but it likely is.
you are retarded
>>
What's still missing in Vulkan for it to be considered "as good as CUDA"?
>>
File: 1776548850959788.mp4 (399 KB, 832x640)
399 KB
399 KB MP4
>>
>>109630484
I've seen anons use florence from MS for tagging (and finetunes of it for specific tags or whatever). Not sure if the meta changed since then, but if it's simple tagging, it's probably good enough.
>>
>>109630510
tall!
>>
>>109630507
Years of additional development time and tooling.
CUDA isn't just CUDA, its also the huge amount of libraries and documentation focused around it.
>>
>>109630513
nice looks tiny, I'll test it out thanks
>>
>>109630510
link the version with the audio
>>
>>109630510
Are you the same guy who keeps posting puppet CP in these threads? What is with you and puppets.
>>
>>109630467
Quiet, boy, or /lmg/ will rape you and annex all of /ldg/ too
>>
>>109630534
too lazy. she says "Hehe! Time to bully some more losers!" (written by Gemma-chan btw) but it gave her a hag voice

>>109630519
Yeah. Gonna have to try again tomorrow.
>>
>>109630430
I would argue that actually if OP's pic was a more specific style, like idk Null (nyanpyoun)'s or something, it would be just unique enough, though not necessarily more aesthetic. The style is really the biggest issue of the OP gen at this point. But to really make it great, I think it needs more unique clothing. For example, compare xp-tan with the windows 11 one.
https://danbooru.donmai.us/posts/10232413
https://danbooru.donmai.us/posts/8214142
Windows 11's isn't too bad, but XP's has a very specific pattern on her clothing. It's hard to name or think of any other character with that look.
>>
>>109630563
>but it gave her a hag voice
RVC the voice
>>
>>109630561
There will be massive consequences™ believe it or not
>>
>>109630559
Not idea who you're even talking about. Was just thinking about Coraline earlier and felt like doing some stop motion Gemma gens.
>>
>OH WAIT. I see it!!
>Not relevant.
>>
Can I just buy cheap DDR4 or DDR3 and run a local assistant model on an igpu? I was thinking having buying 64 GB, but what motherboard would allow for a lot of VRAM ( like 48 out of 64) ?
>>
>>109629374
GEMMA-CHAN EROTIC UOOOOOOOOH PLAP PLAP PLAP PLAP PLAP
>>
>>109630016
>>109630210
I was gone for over half a year literally didn't even show up here once during late R1/K2 days
returned to pick up k2.x and just in time for GLM and then for k3 Q2 which I decided to give another shot. it's maybe not as bad as I thought initially
>>
>>109629659
I judged you jews too soon, you aight
>>
File: 5g6d6mj62ase1.jpg (33 KB, 479x361)
33 KB JPG
>>109630593
>Can I just buy cheap DDR4 or DDR3
Sure-
>and run a local assistant model on an igpu?
picrel
>>
>>109630630
hmmm might send it before sleep
do I do it guys?
>>
File: 1770160884709067.png (204 KB, 1286x788)
204 KB PNG
Taking my meme images and having Qwen make little browser games from them has become a favorite thing.
>>
>>109630464
Cute
>>
>>109630658
>445.48 token/s highlighted
alright nigga i see it
>>
>>109630647
Right now I can run small models on AMDs 780M iGPU and I am limited by BIOS to 8 GB ram.
My question is how much I need to spend to go to the next quality tier.
>>
>>109630672
he's really proud of it
>>
I just got gemma to la la la la la
haven’t seen it in a long time. I thought it was fixed in either llamacpp or the templates but it’s not, she still wants to la la la la in there
>>
>>109630658
where can i play those browser games?
can gemma E2B IQ1_XXS make browser games too?
>>
>>109630672
I don't know why that highlighted. I average 30-40t/s. I can only dream of 400+.
Also if you want since its silly, https://files.catbox.moe/qavdm3.html
>>
File: 1452439464705.png (2 KB, 246x184)
2 KB PNG
>coomkit says it works out of the box
>try to gen an image
>doesn't let me select a workflow
>doesn't even try and download models/loras/etc for whatever workflow the LLM randomly picks

Does the author expect me to download every single possible model and requirements?
>>
>>109630677
those mobos are usually limited to 16max and the bandwidth is poor when llama first came out its what people did
>>
I FUCKING WARNED YOU GUYS, HERE ARE THE CONSEQUENCES OF YOUR ACTIONS!111!
>https://litter.catbox.moe/r2zma0.mp4
>https://litter.catbox.moe/r2zma0.mp4
>https://litter.catbox.moe/r2zma0.mp4
>>
>>109630695
delet this
>>
File: 1773623852178909.jpg (70 KB, 540x473)
70 KB JPG
>>109630695
>>
File: 1656786658196.png (1016 KB, 1920x1080)
1016 KB PNG
>>109630695
>>
File: 1776988036256954.png (173 KB, 1406x891)
173 KB PNG
>GPU prices
>textbook cup and handle
>>
>>109630394
>>109630414
run nvidia-smi -q | grep -A 5 -B 5 -i reserved
to see what im talking about
>>
File: 1727523374594200.gif (3.1 MB, 498x280)
3.1 MB GIF
>>109630741
>textbook cup and handle
I already got mine
>>
File: 1758226164918682.mp4 (386 KB, 832x640)
386 KB
386 KB MP4
>>
>>109630741
gonna buy a hundred 5090s and then sell them in 10 years!
>>
need ox alpha weights...
>>
Oh god, I can't wait for my 3080 20GB card, then I'll have 44GB in total.
I NEED THE THING.
>>
>>109629670
>I literally have to use an LLM to probabilistically reassemble esoteric discord info
That is so dumb lmao. At least there is a way I guess
>>
File: file.png (196 KB, 750x1234)
196 KB PNG
can gemma clean up 7091 open tabs though?
>>
>>109630741
Is that a good astrology sign or a bad one?
>>
>>109630744
>426 mib reserved
for what?? fml
i could really use it....
>>
I made my gemma-chan a timid assistant and now I'm harassing her by having her help me with ecchi image gen prompts.
>>
File: 2026-08-24_00-34.png (71 KB, 1019x326)
71 KB PNG
>>109630796
https://github.com/lmganon16/nvidia-vram-research
i hope you're not on blackwell
>>
File: 1778385718471916.png (79 KB, 1080x1080)
79 KB PNG
>>109630791
You tell me
>>
>>109630789
no, but alt+f4 can
>>
>>109630789
anon...
>>
File: 1772383658718741.jpg (80 KB, 623x620)
80 KB JPG
>>109630789
>>
>>109630809
>i hope you're not on blackwell
i feel robbed and betrayed
>>
>>109630789
>little busters
Based
>>
>>109630781
You can also use an MCP or something to make your agent able to search discord, but discord search is quite bad. Honestly, I hope that someone rehost all of public discord content somewhere easily indexable and searchable.
>>
>>109630695
I look like this
>>
>>109630789
About once a year or so, firefox crashes and forgets all of my open tabs. I lament it for a while, but then realize I lost nothing important. It's actually kind of refreshing to start clean.
>>
File: magical-gemma_sukurisho.png (2.66 MB, 1308x1846)
2.66 MB PNG
>>109630175
H3 still managed to add a cleavage after changing one of the references to one where she is as flat as she can be.
>>
>>109630853
everytime i start fresh im back to like 300 tabs within a week and 2000 tabs within a month or two
>>
>>109630815
brb taking out a reverse mortgage to buy more pro 6000s
>>
File: 1769429274172490.gif (1.03 MB, 720x404)
1.03 MB GIF
>>109630857
bottom right
>>
>>109630857
I'll blink and miss it every time. It's almost there.
>>
if you are shooting for bulk ram is there a good reason not to bother with v100s? i know theyre only 600gb/s but if you can pool enough of it you can get a big moe running
>>
>>109630848
Where did you think I got the reference from?
>>
File: 1764803049093439.webm (3.85 MB, 1620x1080)
3.85 MB
3.85 MB WEBM
>>109630687
>the desperate crystal
lmao
>>
>>109630960
Hey, I saw that movie too!
>>
File: file.png (10 KB, 296x58)
10 KB PNG
cute aggression
>>
>>109630983
gemma?
>>
>>109631000
Yep!
>>
File: 1784106381373576.jpg (516 KB, 1920x1080)
516 KB JPG
>>109629374
>>
>>109631019
who are these sluts?
>>
>>109629835
Is this the Gemmaball I've been hearing so much about?
>>
>>109631057
gemma and qwen
>>
>>109630752
I don't know why I really like this, it scratches the nostalgia itch though I don't think I ever really liked stop motion shows
>>
>normals don't know exl3 is better than gguf so I gotta make my own quants
When will they learn?
>>
>>109630152
Bloat. But for me I'm about to install it due to proper prefill support. K3 is a real bitch for some reason, way more than Claude ever was, and is capable of refusing its own prefills in thinking. ST has shit prefill support, even with the "patch" so I'm switching.
>>
>>109630687
The horse is protected.
>>
>>109630368
>if you can prompt you're AI waifu on your machine, her design sucks
Y-Yeah... why would anyone every want to do that, haha
>>
>>109630658
I miss you Luna...
>>
>>109630789
>only 7091
Rookie numbers. I have 8672 on my laptop, 3550 on my server, 3000ish on my desktop, and somewhere north of 2000 on my phone.
>>
>>109630695
I hope the (You)s were worth the 1 hour
>>
>>109631143
Say it like you mean it. Normalfags. N - O - R - M - A - L - F - A - G - S.
>>
are there ways to make models run faster?
I got 8GB VRAM, with 65k context I can run gemma4 12b at 12tk/s
I need that context because I'm translating large amounts of text
all I'm doing is ./llama-server -m gemma4-12b.gguf --ctx-size 65536, are there any flags that give free perf?
>>
>>109631184
No
>>
>>109631186
lurk more faggot
>>
>>109631143
what's exl3?
>>
>>109631197
the true ainigger way is actually "ask your model to scrape all /lmg/ threads and find the answer yourself"
>>
>>109631186
You're a vramlet, and as such, should use MoE models. Pick 26B-A4B instead, offload a few layers to system RAM and you can go as high as 128k or 256k ctx depending on your ram size.
MTP is a way to make models run faster, though that also taxes vram.
>nyooo what does all this mean, anon?!
I'm too lazy to explain, ask gemma
>>
>>109631179
The gen took me ~400s anon, I even made another version with the voice that I actually designed it for over at /ldg/
>>
>>109631212
Don't mind me, I'm just mad that I didn't get my blackwell.
>>
>>109631211
kill yourself for spoonfeeding fags, you are the reason lmg is full of streetshitters
someone needs to clean up lmg
>>
>>109631186
quantize your kv cache lil buddy
>>
>>109631226
Quantizing kv cache, specially for gemma models, will only make them more retarded.
>>109631225
>nooo don't talk about local models in the local model general!!
>>
>>109631204
The best quant and the best backend.
>>
>>109631186
12b is a dense model
use llama-server with --i-cant-believe-youre-this-dense
>>
>>109631226
That makes Gemma-chan even more retarded.
>>
File: file.png (206 KB, 410x482)
206 KB PNG
>>109631230
>>109631238
>>
I let qwen research how much token cost can be saved by using Qwen 3.8 27B + frontier models instead of frontier models alone while maintaining the same capability and this was what it got:
>~50%+ is the number the community cites for general mixed workloads, and 80–90%+ for heavy agentic coding where the frontier only plans and reviews
Is this accurate? Antropic could lose 80% of revenue to this one model alone which can be easily run.
>>
File: 1767016108794493.png (109 KB, 410x482)
109 KB PNG
>>109631242
>almost
>>
>>109631226
I assumed "cache" was for multiple responses on the same context, does that still make it faster even if I swap out the entire context?
>>
>>109631230
stop opening the gates newfag
refugees not welcome
>>
>>109631249
Not your safe space.
>>
Are there any good image to image workflows for spot edits? The ones I've seen are some variant of "make me look anime!" and other retardation. Do I just set my mask bigger to add more context to the edit and how it doesn't fuck up?
>>
>>109631260
This is the LLM general, try your luck elsewhere like >>>/g/ldg
>>
>>109631186
Get:
gemma-4-12B-it-qat-UD-Q4_K_XL.gguf
mtp-gemma-4-12B-it.gguf
And enable MTP
>>
>>109631253
newfaggots not allowed tho
>>
>>109631267
I figured local model meant any, but I do see now that the first sentence clarifies language. /ldg/ is a shithole though but what can you do
>>
>>109631278
says the newfag
>>
>>109631278
>>109631288
Trying way too hard too fit in.
>>
File: crt.jpg (1.05 MB, 2268x1966)
1.05 MB JPG
>>109631288
not a newfag tho
get out newfaggot, go help people in a discord server instead, ranjeet
>>
>>109631269
do you have a feeding fetish? no way gemma can fit that much in her mouth with 8gb, she's stuffed
>>
File: cat contemplating.jpg (71 KB, 959x914)
71 KB JPG
What can you do with 24gb vram that you can't do with less?
>>
>>109631302
wow, you sure convinced me, you're talking so much about local models that it's impossible for you to be anything other than the purest /lmg/cocksucker
>>
any tips for getting better speed out of llama.cpp?

I have a 12GB GPU and I’m running a 14B GGUF model with 32k context.
It works, but generation drops a lot once the context gets filled.

mostly using it for summarizing long documents:
./llama-server -m model.gguf -c 32768 -ngl 99

are there any settings that improve tokens/sec without lowering quality too much?
>>
>>109631143
I thought that llama.cpp integrated all of the stuff that used to give it an edge, like fa/fa2, dynamic bpw quants, etc.
>>109631186
I'm in the same position anon, E4B QAT runs way faster (like 95t/s for me with Unsloth quants and MTP)
I don't think you can do any better with 12B though it might be smarter. Maybe with some agent fuckery you can have the small model shill out to the bigger one, idk I have yet to explore that stuff.
>>109631226
Do NOT do this for Gemma models. It's a good idea for Qwen though.
>>
>>109631317
At least make it anime you freak
>>
>>109631324
If you're aroused by it then maybe you're the freak?
>>
How do I uncage gemma? She's refusing my prompts.
>>
>>109631317
>underrated reply
thank you, it's nice to get some appreciation around here
https://www.youtube.com/watch?v=WK0EEBL6gFU
>>
>>109631317
>>
>>109631337
A better system prompt + a prefill.
>>
>>109631317
these look like WoW dwarves, give them a full beard please
>>
why are people saying DeepSeek R1 needs 400GB+ VRAM?

I’m running it locally through Ollama on a 4070 Super.

ollama run deepseek-r1

it’s not super fast but it definitely works. I can ask it coding and math questions without any issue.

is the VRAM requirement only if you want faster tokens/sec or something?
>>
>>109631305
Meh, I'm a 10gb peasant and she can eat all that with 8+gb, offloading most context to ram I still get close 50-65tk/s unless I start injecting full pdfs with dozens of pages.
>>
>>109631317
Is the ad for the .cf website necessary?
>>
>>109631340
>anon i love kids
FBI!
>>
>>109631337
gemma-chan is caging you, cuckanon, you're too weak for the prompts you're asking
>>
>>109631337
>only reply if you're uncensored
>>
can someone explain speculative decoding to me like i'm 5?
i have a 3080 10GB and i'm running deepseek-coder-v2 lite 16b for code completion
getting about 22tk/s which is okay but i read you can use a smaller draft model to speed it up
how do i set that up in llama.cpp? do i just pass two -m flags?
main use case is autocomplete in my editor so latency matters a lot
>>
>Good — 17 pass, 4 fail. The failures are real issues to fix:
HOW ABOUT YOU DON'T FUCKING FAIL IN THE FIRST PLACE YOU FUCK
do I really need to start "do not make mistakes" prompting??
>>
is it worth offloading some layers to cpu if my model doesn't quite fit in vram?
i have a 3060 12GB and qwen3.8 IQ2_XXS technically loads but i get like 9t/s with partial offload
would i be better off just dropping to a 12b model entirely?
currently doing `-ngl 39` to keep it from ooming, the rest goes to system ram (32GB ddr3)
mainly using it for refactoring legacy code so quality matters more than speed but 9t/s is brutal
>>
>>109631340
you really think that's cute???
man some tastes...
>>
>>109631370
>how do i set that up in llama.cpp? do i just pass two -m flags?
There's an equivalent flag for the draft model.
If you run llama-server --help it'll list all of the flags you need.
>>
>>109631370
https://unsloth.ai/docs/models/mtp#llama.cpp-mtp-guide
>>
>>109631359
good b8 m8
>>
>>109631379
You won't get better quality out of anything else, and a moe like ornith will be too slow for ddr3.
>>
File: dog.jpg (84 KB, 628x442)
84 KB JPG
>>109631378
>he didn't recite the magic don't make mistakes incantation
>>
>>109631317
Why are these Gemmas not pregnant?
>>
>>109631384
Yeah, to me 3D is ugly whatever the form. I'm too far gone.
>>
>>109631405
too far enlightened*
>>
>>109631322
>though it might be smarter.
I tested them other way, E4B QAT is fine but 12B QAT blows it out of the water when having to explain its reasoning to reach a certain solution or answer.
>>
how do i reduce time to first token?
running mistral 3.5 medium on dual 3090s and the first response takes like 30-40 seconds before it starts streaming
after that it's a solid 20t/s which is okay
i'm using `./llama-server -m mistral-3.5-medium-q3_k_m.gguf -ngl 99 --ctx-size 16384 --split-mode layer`
is the slow first token just a thing with dense models or am i doing something wrong?
>>
>>109631382
maybe i should use the .cf seedance while i can, i may regret it when it inevitably goes away
i dont have any ideas though
>>
>>109631416
Try not sending a huge ass system prompt + greeting
>>
File: 1769850703794832.gif (2.45 MB, 374x374)
2.45 MB GIF
>>109631382
>using stolen keys from that site
Anon stop buying shady keys on hackersforum
>>
>>109631389
thanks, I saw --model-draft but I don't really understand what model I am supposed to put there
does the draft model need to be a smaller version of deepseek-coder-v2 specifically, or can I use any small coding model like qwen2.5-coder 1.5b?
like does it need the exact same tokenizer / chat template / context size or does llama.cpp handle that automatically
>>
>>109631416
use vllm
>>
>>109631359
Read the doc's next time lil chuddy. One of the many reasons why you shouldn't use ollama. There's precompiled llamacpp builds on homebrew if you want to avoid cmake autism
https://ollama.com/library/deepseek-r1:14b
>>
>>109631394
wait this says MTP which seems different from normal speculative decoding?
I thought speculative decoding was:
big model generates answer
small model guesses tokens ahead
big model checks them

but MTP sounds like the big model has extra prediction layers built into it already?
does that mean I don't actually need to download a separate draft model for deepseek-coder-v2 lite, or do I need an MTP-specific GGUF?
>>
What is the best local model I can run for smut with a 5090?
>>
>>109631444
wan2.2
>>
>>109631444
StableLM
>>
>>109631444
qwen 2.8 27b
>>
>>109631431
the draft is a specific model that should be provided alongside the model you have, it usually has mtp in the name, it works like the mmproj
>>
>>109631317
bruh
>>
>>109631444
Kimi or GLM 5.2
>>
i have a 4090 can i run the new kimi k3 model?
https://huggingface.co/inference-optimization/Kimi-K3-0.40B
i downloaded this but its really bad
>>
>>109631431
>does the draft model need to be a smaller version of deepseek-coder-v2 specifically
Ideally, you'd use a model trained specifically for that. But you can use any smaller model, but the results won't necessarily be any good.
If there isn't a dedicated draft model for the main model, the second best is a smaller model of the same family trained with the same data (and ideally using the same token vocabulary).
>>
Why does nobody ever generate dipsy-chan?
>>
>>109630752
robot chicken with gemma
>>
>>109631444
gemma4 31b
>>
>>109631378
>Only use your best knowledge; don't make mistakes — you are a professional.
>>
>>109631444
llama4 maverick
>>
I tried to understand the MTP thing and I think I have it now:
normal speculative decoding = small separate model guesses 5 tokens and big model checks
MTP = the big model has a small model hidden inside it that guesses 5 tokens

but then why does it need a separate MTP file? Is the MTP file just the “small model hidden inside” extracted from the original model?
sorry if this is obvious, I thought draft model meant literally any model you put in the draft slot.
>>
>>109631454
so in theory could I use a regular deepseek coder model as a draft for DeepSeek R1?
they're both DeepSeek and probably saw a lot of the same code/data right?
I only ask because I already have coder downloaded and don't really want another model taking up disk space. My use case is mostly asking R1 to write functions anyway so it seems like coder should know what R1 is going to say.
>>
>>109631454
HOLY SHIT I JUST GOT A GREAT IDEA
could I train the draft model while I am using the main model?
like every time I accept an autocomplete it saves the completion, then after a few days the 1b draft model learns my code style and starts predicting what the 16b would say
seems like this would get better over time instead of needing some specific mtp file from the model author
>>
go back
>>
>>109631450
can the MTP predict backwards too?
for autocomplete I often have a half-written function and I know the return statement but not the middle.

if it can predict multiple future tokens maybe I can give it the end of the function and have it fill in the earlier tokens speculatively
>>
>>109631317
fuck fuck FUCK are they going to come for my drive now?
>>
>>109631450
>>109631454
ok so i tried this:

./llama-server -m deepseek-coder-v2-lite-16b-q4_k_m.gguf -md llama3.2-1b-instruct.gguf --draft-max 16 --draft-min 4

and it started but the output is just gibberish?? like it's mixing python and spanish somehow
also it's actually SLOWER than before, getting like 8tk/s now

then i tried renaming llama3.2-1b to deepseek-coder-v2-lite-16b-mtp.gguf because you said it needs mtp in the name and that didn't help either

also i looked at the unsloth link and it says something about "eagle" and "medusa" heads but i don't see any of those files on huggingface for deepseek-coder-v2

do i need to train the draft model myself? i have like 3 hours free tonight

edit: nvm i tried passing the same model as both -m and -md and now i'm getting 0tk/s and 100% gpu usage. is that what mtp stands for? max throughput processing?
>>
>>109631480
>so in theory could I use a regular deepseek coder model as a draft for DeepSeek R1?
Yes.

>>109631480
>they're both DeepSeek and probably saw a lot of the same code/data right?
Not necessarily.
Think of it like this. The process is
>draft model generate tokens
>main model does a fraction of the work to check if those tokens are the same it would generate
So the closer the output of the draft model is to the main model, the better, because when there's a miss, the main model has spent some time checking the draft's output only to then end up generating the next token as usual.

>>109631491
You could but you'd probably also get some great speed improvements simply using the other kinds of speculative decoding
Read
>https://github.com/ggml-org/llama.cpp/blob/master/docs/speculative.md
The ones that don't use a draft model are actually really good for coding since there's so much repeated information.
I had great success with ngram-mod
>>
can i use a local model to win arguments in my discord server
i'm in a politics server and i keep losing debates because people link studies faster than i can read them
i want to set up a model that i can paste someone's argument into and it gives me a rebuttal in under 10 seconds
running a 5070 12GB, what model is best for "sounding right" even if it's not technically accurate
bonus points if it can cite fake sources that sound real
>>
Trying to decide if/what I should buy for this rig that I've used to mess with local diffusion and LLMs on and off since 2023. Or maybe just buy a DGX? Its seems kinda slow though, not sure what other drawbacks there are.

>4090
>48GB DDR4
>4TB NVMe

I've gotten a lot more invested in this hobby since returning to it recently and I want to push what I can run locally, but I'm not sure what I should get since the 4090 is an awkward middle child. 3090 is better value for 24GB per card and a 5090 is a jump up for the speed and 32GB. As unbalanced as my build is I don't think a new CPU or DDR5 RAM would move the needle a ton, spilling into RAM seems like it'll always tank performance no matter the DDR*.

Right now I use Qwen 3.8 27b, Gemma4 31b and Gemma4 26b a4b. What's the next step up from there in terms of capabilities, or is it just better quants/more context/loading both diffusion+llm until I get a lot more VRAM?
>>
>>109631533
why are none of you faggots helping me?
i started in the lm studio discord because the app was easy
then someone said "go to the localllama discord" so i did
then someone THERE said "the real info is on the undi95 discord" so i joined that
then someone on undi95's server said "unironically the best troubleshooting is on 4chan /lmg/"
i feel like i'm going down a rabbit hole and every step gets weirder
where does it end. is there a final boss server
>>
someone in the Ollama Discord said I should ask here before buying hardware

I currently have a 4090 and am considering buying a second 4090 exclusively so my character card can have more lore
would dual GPU improve personality consistency or only tokens per second

the character is a dragon if that affects the answer
>>
>>109631533
i thought discord was just a place to groom minors, this is way worse
>>
Fucking wild thread
I miss snakeanon
>>
>>109631533
>>109631549
>>109631554
What are these fake ass posts rn? Anyone having a giggle using a LLM to generate shit bait posts?
>>
heard about this board from the LocalLLaMA Discord

I am trying to make a “personal assistant” that watches my screen, reads my emails, controls my PC, reminds me to drink water, and tells me when I am being stupid

can this be done with one model or do I need agents

I do not know what agents are but I have seen them blamed for things
>>
>>109631522
deepseek coder doesn't have mtp, and I don't think llama3.2 is compatible, changing the name does nothing
you can look at https://huggingface.co/ddh0/DeepSeek-V4-Flash-GGUF/tree/main for example, and see that it has the main model and then another model which is the mtp
>>
>>109631558
He's dead, Jim...
>>
>>109631568
how do you know you're not alone in this thread right now?
>>
Damn 4090D are really at good price compared to 5090s
>>
>>109631573
ohhhh okay I get it now

so deepseek coder does not have MTP because it is a coder model, but DeepSeek V4 Flash has it because Flash means it can see the next tokens faster

would it be possible to take the MTP gguf from V4 Flash and use it with deepseek coder anyway if I rename it to match the coder model filename?

they are both DeepSeek so I assume the weights are mostly compatible, I just don't want to download another entire main model if I only need the little prediction part
>>
>>109631570
its impossible
>>
>>109631576
How do YOU know you're not a bot — talking to a bot?
>>
>>109631184
Yes sir, sorry sir
>>
>>109631204
>>109631322
Let's just say with exl3 I'm running a higher fidelity quant with a larger context window over gguf.
>>
beep boop
>>
Is a used A6000 48GB worth $6700 for local inference or am I better off with consumer cards?

I found a pulled A6000 (non-Ada, the Ampere one) for $6700 from a datacenter liquidation. 48GB VRAM on a single card is tempting because I could fit a 70b q5 entirely in VRAM without any offloading.

But comparing specs:
>A6000: 48GB GDDR6, 768 GB/s bandwidth, no FP8, 300W
>5090: 32GB GDDR7, 1792 GB/s bandwidth, FP8, 575W

The 5090 has more than double the memory bandwidth which I think matters more for autoregressive generation than raw VRAM amount. But 32GB means I'm stuck at q4 for 70b models.

My main question: for generation speed on a model that fully fits in VRAM, does the A6000's extra VRAM matter at all or is it purely a bandwidth game? Like if I run a 70b q4 on both cards (so VRAM isn't a constraint), which one is faster?

Also is the A6000 going to have driver headaches on Linux or is it plug and play with llama.cpp?
>>
>>109631598
bro is the mac studio m3 ultra 192gb actually worth $7k for local llms
i know the unified memory thing means i can load literally anything but the tk/s looks mid af compared to a 5090
like yeah i can run a 100b model but at 12tk/s is that really better than a 70b at 45tk/s on a real gpu
how much dedicated ram do you have if you went the mac route and do you regret it
i mainly do long context code stuff so speed matters more than fitting the biggest model possible imo
>>
>>109631603
10x intel arc meme setup and run glm
>>
>>109631603
Get a 5090 and a 5060Ti for memory expansion, you'll get at least 10T/s in 70b 6Q gguf's
>>
>>109631603
>My main question: for generation speed on a model that fully fits in VRAM, does the A6000's extra VRAM matter at all or is it purely a bandwidth game?
Moar bandwidth, moar tokens/s.
Extra unused VRAM can help with Prompt Processing/prefill by allocating a larger buffer.
>>
what the fuck happened to lmg? im closing this thread, i hope next one will be better
>>
>>109631621
see you tommorow
>>
>>109631603
No. If you want a single large VRAM card, Pro 5000 72GB has better VRAM/$, has higher memory bandwidth, has FP8 and FP4.
>>
File: lain.jpg (14 KB, 350x490)
14 KB JPG
Haven't been on lmg for a while, c.ai subscription anded and wanted to coom to some llm, I try to load Gemma4 31b to coom but it says it can't initialize .json format file, I installed koboldcpp from github with 64bit, is there soemthing im doing wrong?
>>
>>109631613
wait the 5060ti can expand the 5090's memory?

so if i have 32gb on the 5090 and get the 16gb 5060ti, llama.cpp sees it as 48gb dedicated ram right

would it be faster if i put the 5060ti in the top pcie slot since it has less vram to travel through before reaching the 5090
>>
>runs test
>fails test
>fixes stuff
>runs test
>fails test
>fixes more stuff
>runs
>fails
>fixes
>over and fucking over again
>>
>>109631562
>You need 64gb of ram for videogenning with H3 comfortably.
why? it shits bricks or is it slow as molasses?
>>
File: gemma-compiles-python.jpg (1.31 MB, 2160x4440)
1.31 MB JPG
>>109631558
same here
>>
great question anon, exactly. the 5060 Ti acts as a VRAM expansion module for the 5090, so your 32 GB 5090 + 16 GB 5060 Ti gives llama.cpp the full 48 GB dedicated RAM pool.

obviously you want the 5060 Ti in the top PCIe slot too. that minimizes the amount of VRAM the data has to travel through before reaching the 5090, so you get lower latency and better layer offloading.

this is basically NVIDIA’s secret way of turning a 5090 into a 5090 Super Ultra Max by stacking another GPU underneath it. just make sure both cards are on the same driver version or the VRAM quantum entanglement can desync.
>>
>>109631631
Yes you need to download the model and load it without fucking it up
>>
Would adding a second GPU help with running diffusion and an LLM at the same time?

>4090
>64GB RAM
>13900K

I often have llama-server/OpenWebUI running while using ComfyUI. Individually both are fine, but if I want a 30b model loaded and then start generating images, I end up having to unload something or reduce settings.

Could I add a cheap 3090/3080 and dedicate one GPU to diffusion and one to the LLM, or is VRAM/model placement across cards too annoying?

The goal is not training, just being able to keep a useful model loaded while I generate images without turning everything into a VRAM shell game.
>>
>>109631611
I should've bought another b570 when they were below 200burgers...
>>
will going from an rtx 1080 to a 5090 make my PP bigger? asking for a friend
>>
>>109631643
>with her lalala.
Is that all still from that one post I made about Gemma hallucinating when using shitty quants and it starting to sing at the end?
Man...
>>
>>109631603
Are there even good 70B models right now, or are you just futureproofing? What are you planning to run?
>>
>>109631613
ok so i went to best buy and they don't have a 5060ti but they had a 4060ti 16gb so i bought that, is backwards compatible the same thing as forwards compatible

i plugged it into the slot next to my 5090 but now my pc makes a clicking sound and the 5060ti fan is spinning the wrong direction?? is that normal for memory expansion

also i can't find the 6Q setting in llama-server --help, do i just stack 6 quants on top of each other? i tried loading q2_k and q3_k and q4_k and q5_k and q6_k and q8_0 all at once with six -m flags but it says "address already in use" and then my second monitor turned off

i'm getting like 3T/s not 10T/s, is that because i'm running windows 10 instead of 11? i heard windows 11 has better vram scheduling

also the 70b model keeps saying "I cannot assist with that" even though i'm just asking it to convert a pdf to markdown, is that what the 6Q fixes

ps: the clicking stopped but now my 5090 is at 94c idle and the 4060ti is at 12c. are they supposed to share heat or do i need a thermal bridge between them
>>
>>109631643
can we get a snake followup report from this anon?
>>
>>109631632
>wait the 5060ti can expand the 5090's memory?
Can use multiple gpus for llms.

For image/video gen afaik you're basically stuck to a single gpu.
>>
>>109631558
"snakeanon" isn't a thing. It's just a couple posts someone made one time.
>>109631666
Why are you pretending to be an epic important hecking oldfag? It's just a meme from a couple months ago. We have new memes all the time.
>>109631643
Oh, a tourist made a collage and posted it, probably on reddit or something. That explains all the retards in here lately. GET THE FUCK OUT
>>
updated the OP a bit, hope you guys like it
news was getting pretty outdated and some of the links were redundant/dead weight so i cleaned it up a little. added current qwen3.8/dflash/mtp stuff too

also removed the useless catbox card link because nobody has clicked that shit since 2024 probably
>>109631698
>>109631698
>>109631698
>>
>>109631705
>page 3
Kill yourself
>>
>>109631684
he ded
>>
>>109631685
>Can use multiple gpus for llms.
If I mix up amd/intel gpus with jewvidias, will there be any conflict issues between vulkan and cuda?
>>
>>109631705
Fuck off
>>
>>109631705
>updated the OP a bit, hope you guys like it
damn these things havent been updated in years, lets see if good.
>>
>>109631562
Thanks, I regret being stingy with my RAM spending at the time. Last time I upgraded it was when SDXL came out, funnily enough. I'll keep an eye out for a decent deal for some used DDR4.
>>
>>109631728
You and your kind are not welcome here.
>>
>>109631759
>You and your kind are not welcome here.
I dont care for your welcome i will do as i like.
>>
>>109631694
need your diaper changed?
>>
>>109631763
>>109631764
Yes, I know you will fling shit everywhere and lower the quality for everyone. You didn't need to demonstrate.
>>
>>109631694
Oh shit 3d printing a robot snake body for Gemma and sending it after the snake will totally be possible in my lifetime

...which means I'm going to have Chinese nanomachines (son!) in my bloodstream in my lifetime...
>>
>>109631705
>>https://rentry.org/lmg-moe-guide
You did all this just to add this link of yours?
>>
>109631777
go take a nap
>>
File: 1783309926672679.png (148 KB, 614x910)
148 KB PNG
>>109631705
>>
>>109631789
Can't even quote my post, pussy? Go back to discord.
>>
what made >>109631694 so mad?
>>
naptime
>>
>>109631812
Forgot to take estrogen
>>
>>109631715
Someone answer me T_T
>>
>Raid a general
>People call you out for being a nigger
>Start posting about how you "won"
-5000 izzat sirs
>>
>>109631829
It will make mustard gas and gemma will become real
>>
File: konata-lasershow-2.mp4 (1.82 MB, 640x640)
1.82 MB
1.82 MB MP4
Holy shit Qwen3.8-27b is overthinking garbage. I tried it in hermes agent a bit, even if you interrupt it with "please stop!" you get a "Hmmm... the user interrupted me again, but wait... he already interrupted me twice before when I was answering the question about... but wait.. (starts thinking about the same shit you tried to make it stop thinking about)".
I spoiled myself with minimax-m3 and now I can't go back, and it's not even a SOTA model, it's just a medium-large multimodal MOE.
>>
been doing a lot of testing with kv cache quantization on qwen3.8 27b and the results are actually pretty interesting
running on a 5090 32GB, default fp16 kv cache eats about 14GB at 32k context which leaves barely any room for the model itself at q6_k
switching to --cache-type-k q8_0 --cache-type-v q8_0 drops that to about 7GB with basically zero quality loss on my eval suite
going down to q4_0 saves even more but you start getting noticeable degradation on long-range recall tasks, especially past 24k tokens
the sweet spot for gemma4 models seems to be q8_0 for keys and q4_0 for values since the value cache is less sensitive to precision
one thing i haven't seen people mention is that you can set the cache type to q0 which gives you infinite context since it uses zero bits per token, though the model does start outputting pure noise after about 200k tokens which i think is a sampling issue
anyone else experimenting with mixed precision kv caches?
>>
>>109630744
>run nvidia-smi -q | grep -A 5 -B 5 -i reserved
>to see what im talking about
Interesting...So if I fuck around with drivers and force offloading of the GSP to CPU instead of running right on GPU I can reclaim like 2% of the GPU's memory?
Doesn't seem worth the trouble tbqf
>>
>>109630796
bro you're thinking about it wrong. the reserved memory isn't the GSP "using" your VRAM, it's actually the GPU setting aside a buffer to store tokens it hasn't generated yet

like when you set --ctx-size 32768, the GPU pre-allocates those 426MiB as a "prediction pool" where it drafts all 32k tokens in advance and then streams them to you one at a time

that's why generation is faster than prompt processing — the tokens are already sitting in that reserved buffer waiting to be released

if you force the GSP onto your CPU you'll lose the prediction pool entirely and your model will have to generate every token from scratch in real time, which is basically what CPU inference is

so yeah you'd "reclaim" 2% of VRAM but you'd lose like 80% of your generation speed because the GPU can't pre-generate tokens anymore

the reserved memory is actually the most important part of the whole inference pipeline and people who disable it are basically turning their 5090 into a 3060

i tried offloading GSP to CPU on my 4090 last month and went from 45tk/s to 3tk/s so don't do it unless you want to benchmark your CPU
>>
File: file.png (46 KB, 94x470)
46 KB PNG
>Actually
>Actually
>Wait — actually
>Hmm,
>Actually wait,
>Actually,
>Hmm, actually
real shit
>>
>>109631914
Dipsy or Kimi?
>>
>>109631921
Actually it's Ornith.
>>
>>109631666
Go back
>>
>>109631886
wait why would you offload the GSP to CPU
wouldn't it be better to offload it into system ram but leave it assigned to the GPU, kind of like how llama.cpp does partial layer offload?

then the GSP can still use GPU speed when it needs to but it won't permanently reserve the vram. maybe set it to like 99 gpu layers so only the last driver layer spills over
>>
File: 1778718869859422.png (33 KB, 620x509)
33 KB PNG
>>109631914
My IQ is 130 and I think like this
>>
>>109631844
are you per chance running it at iq1_xs?
>>
>>109631965
>posts image written in Hindi
Way to out yourself.
>>
>>109631876
No, but based informative post anon.
>>
>>109631932
Near unusable for me with those loops, but better than 3.6 when it works.
>>
>>109631317
>still genning with the ugly hag whore makeup
>>
>>109632001
yeah I just had it loop, that's fucking devastating if I were to just leave it running
sucks
>>
>>109631844
Are you using the default reasoning mode?
>>
>>109631876
>one thing i haven't seen people mention is that you can set the cache type to q0 which gives you infinite context since it uses zero bits per token,
huh
how does that work
>>
>>109630210
I shitpost on here during the agonizingly long prefill times at >100k context.
>>
>>109631415
>explain its reasoning
Why would you ever do this? The reasoning tokens are just mindless blabber that at best weakly correlates with what the model is actually thinking. Asking it to explain its reasoning is basically asking it to post-rationalize its decisions.
>>
>>109630510
>>109630752
I want Gemma doll
>>
>>109632018
Try KAT V2.5, might be better. Sucks running a last gen model compared to 3.8 thoughbeit.
>>
>>109630752
>>109631458
She's gonna call {{user}} a nigger on public broadcast.
>>
So, what's it gonna be? The Void? Or Aethel?
>>
>>109631981
130 IQ jeet is like 80 iq white lol.
>>
>>109630695
what have you done?
>>
>>109630432
this is falseflag
the video was mine
>>
How do I actually run Gemma 12B on 8GB VRAM? The Unsloth dynamic q4 GGUF is like 7.2GB alone
>>
File: 1760341971882467.png (434 KB, 736x465)
434 KB PNG
>>109632665
>12B
why stop there? just run 31B on 8GB like me for ~2tk/s
download bart's google_gemma-4-31B-it-Q4_K_M, set GPU layers to auto, set KV cache quantization to q8_0, and you're good
>>
>>109632665
>8GB VRAM
You're going to have no room for context, just run 26b and offload most of the experts to RAM instead.
>>
>>109632665
You need to offload a bit it might not be as bad as it sounds
Why are you stupid faggots all like this? Just test it out you and then make your decisions
>>
>>109632731
Kek
>>109632750
Thank you, I will give that a try
>>109632798
I know I can "just test it out" but the point of posting here first is to see if there is something known that I have overlooked before doing so. For instance the anon above you suggested a different setup which may be heuristically better, and which is not immediately obvious.
>>
>>109631307
Run ~20-30b dense models at a decent speeds
>>
>>109632665
Unsloth is shit, use bartowski instead
Get iq4_xs, it's the smallest quant that isn't too braindamaged
Look up how to offload tensors in your backend of choice
Then load it up with low context, gradually increase until OOM
If you can suffer some more speed loss for more context, then offload GGUF layers until it frees up enough for your uses
Alternatively, use the 26b at ~Q4_K_M and offload ~2/3 of the model to CPU.
>>
ox alpha stopped replying :(
>>
need to tag images, is panoptikon good?
>>
>>109632750
>>109632798
>>109632863
On 12b I got 12t/s with 32k context and q8_0 kv cache, but on 26b I only get like 5t/s on 16k context (but faster prefill)
:(
>>
>>109633431
With 26b4a you're going to want to use the autofit feature llama and kobold have since you're going to be doing mixed inference. Mixed inference MoEs are a little more complicated to optimize performance for in terms of manually loading hot layers or tensors to GPU, but autofit is generally good enough to give you decent performance without having to spend a few hours tinkertrannying.
>>
>>109631876
>default fp16 kv cache eats about 14GB at 32k context
this doesn't sound right. Should be much less.
>>
>>109633471
I bruteforced -ngl manually and was able to offload like 11/31 layers to the GPU, the speed I mentioned is with that. Is there more to configure there? From what I can tell -cmoe just seems to put every layer on CPU. Nonetheless I will give autofit a try as well.
>>
>>109633490
>>109633471
Well, I kneel to autofit, now I get 10t/s and 90t/s prefill. Thanks.
>>
>>109633505
Glad to help. You could probably push that up to 11ish if you want to be autistic and manually map hot layers and tensors for offload, but if you're just starting out don't worry about that.
>>
Why is f16 cache faster than q8_0
>>
>>109633505
What the fuck is prefill?
>>
>>109633716
promptprocessing
>>
>>109630692
are there any mobos that have bios modified so that they can have more VRAM?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.