[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1785665081994458.png (1.03 MB, 800x1296)
1.03 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109440142 & >>109435398

►News
>(07/31) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 1756403016635967.jpg (166 KB, 777x933)
166 KB JPG
>>
File: gemma sex.png (898 KB, 1064x3884)
898 KB PNG
gemma sex
>>
gemmaballs
70b dense
don't believe dario's lies
>>
what's a good local model for solving math problems? I have 31gb vram spread over 4 gpus and 192gb of ddr4 ram (could be upgraded ot 384gb)
my use case is solving proof-based contest math for middle and high schoolers
>>
Chinesium-3T is a benchmaxxed, native-marketing agentic press release and our most capable slideshow to date. It is a 2.8T-parameter model built on Vibes Delta Attention (VDA) and Distillation Residuals (DistillRes), with native cope capabilities and a 4,096-token context window that we call one million on the box. It is the world's first open 3T-class model you cannot download for another nine days, designed for frontier benchmarketing across long-horizon greentext, thread validation, and WAIC keynote timing.

Key Features

Novel Architecture: Chinesium-3T is built on two acronyms we invented last month and a MoE framework activating 16 of 896 experts, 4 of which, when routed correctly, will apologize and remind you it cannot help with that request. The remaining 892 have never fired once. We call this "sparsity." It scored a 2.5× efficiency gain over K2 and a 0× gain over the model it was distilled from, which we would name if we could.

Long-Horizon Coding: Operating with minimal human oversight, Chinesium-3T sustains long engineering sessions right up until token 4,097, at which point max_position_embeddings: 4096 gently reasserts itself and the model forgets it was writing a compiler. It can design a chip. It cannot remember which chip. This is by design.

Agentic Knowledge Work: Produces deep research, interactive dashboards, and in 4% of sessions a paragraph beginning "I'm Claude, an AI assistant made by Anthropic."

Native Multimodality & Long Context: Chinesium-3T understands text, images, and video within the same model, and supports a 1-million-token context window in the same way a restaurant "supports" a reservation for 40 when it has four tables. The base was pretrained at 4k. The million is a lifestyle.

Open Frontier Weights: Released under the Chinesium-3T License, a document distilled from the Anthropic Usage Policy with the word "Anthropic" replaced by a 304GB IQ1 quant nobody can run.
>>
File: 1763972765396014.png (3.4 MB, 4243x2372)
3.4 MB PNG
>>
70b dense
>>
File: nutty.png (1.61 MB, 1026x1021)
1.61 MB PNG
https://www.youtube.com/watch?v=hpTC4U-ye2Q
>>
>>109446144
this is how the minds of those who doubt j-spaces work
>>
put the consciousness in the j-space and give it a good stir
>>
>>109446133
Why don't you go for the fattest local model you can run first and then try to work down from there? Your memory layout basically guarantees that you'll only be running MoEs with only hope of speed. If I remember correctly I think the latest DeepSeek V4 Flash tops out at about 168GB total so that should fit well within your budget with memory to spare for full size context. I don't really recall what the quantted sizes for the other big name medium models like GLM5.2 or Kimi K2 are, so you might want to look at those. If you could deign to I'd appreciate the results of whatever you're trying to do. Doing math proofs isn't as well talked about as coding or storywriting.
>>
>>109446144
>You want to know what is the value of 2938293849294949 times 1029204849294922? Sure, let me whip out a python script and—
>NOOOOO you are supposed to do it in your head!!! IT DOES NOT COUNT!!!
>>
>>109446194
It does not count though.
>>
File: jew1785743713276476.png (27 KB, 389x511)
27 KB PNG
>>109444972
>I ended up with some kind of ollama-style cloudshit when I tried
LMAO you can't be fucking serious. how do they get away with it, can't you just tell some cloudjew model to rewrite the node or whatever those things are called to not be jewed?
I mean you're not telling me these people run closed source shit in there, right? with an internet connection up? right?
>>
File: Gemma-Chan Recap.png (505 KB, 1024x1024)
505 KB PNG
►Recent Highlights from the Previous Thread: >>109440142

--Reaction to Qwen 3.8 Max benchmarks and open weights release:
>109442634 >109442641 >109442746 >109442881 >109443464 >109444889
--VRAM optimization and quantization for MiniMax-H3:
>109444459 >109444482 >109444729 >109444738 >109444972 >109444978 >109445284 >109444983 >109445085 >109445000
--AI manga translation pipeline and typesetting discussions:
>109441411 >109441476 >109441505 >109441548 >109442243 >109442276 >109442291 >109442343 >109442368 >109442374 >109442419 >109445566 >109445639
--Feasibility of running models from PCIe Gen 6 SSDs:
>109440613 >109440735 >109440755 >109440768 >109440931 >109440875 >109440896 >109441011
--Praising NVIDIA RTX Spark Superchip for clustered local inference performance:
>109442564 >109443156 >109445062 >109442612 >109442630
--Reaction to upcoming Qwen3.8 open weights release:
>109442698 >109442804 >109443033 >109444165
--Combining VISReg and Explorative Modeling for continuous vector prediction:
>109440310 >109440370 >109440574
--Analyzing DeepSeek-V4-Flash-High performance and model efficiency metrics:
>109441659 >109441674 >109441681
--MiniMax-H3 release featuring regional license exclusions and massive VRAM requirements:
>109444401 >109444434 >109444504
--Relationship between model information density and quantization degradation:
>109441249 >109441270 >109442194
--Qwen3.8-Max rankings on the Frontend Code Arena:
>109444833 >109444838 >109444871
--Analogies between biological evolution and LLM consciousness and qualia:
>109440330 >109440367 >109440560 >109440605
--Troubleshooting mixed Pascal and Blackwell GPU support on Fedora:
>109442426 >109442457
--Logs:
>109440470 >109441791 >109441923 >109442021 >109442586 >109444581
--Miku (free space):
>109440170 >109440297 >109441018 >109442446 >109442459

►Recent Highlight Posts from the Previous Thread: >>109440146

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109446177
installing and managing big models is a pain in the ass, so I would prefer to work my way up over potentially overshooting. any good specific recommendations will also save me a bunch of time.
when I find a good model, I will have a lot of work for it to do, so expect updates later rather than sooner.
>>
Need Mistral Large 4
>>
File: w9ofyt8qod3d1.jpg (75 KB, 1242x876)
75 KB JPG
>PCIe Gen 6 SSDs
Are SSDs becoming the meta? How many tokens per seconds can we achieve on the big boy MoEs on them? I saw an anon buy 4 massive SSDs a few threads back. Is this happening?
>>
>>109446144
/g/ is full of people like that
I'm starting to believe that the npc meme is real
>>
>>109446314
>monkey paw curls
granted, but it's a moe
>>
>>109446318
yeah you should stock up on as many fast ssds as you can before prices explode
>>
>>109446324
But Large 3 was already MoE too
>>
>>109446325
But SSD prices already exploded, anon.
>>
>>109446133
>>109446177
I should have mentioned that I can run glm 4.5 air (1bit quant) at about 7 tok/sec. not sure how good or bad that is
>>
>>109446331
compared to the below value prices from before, yes
compared to the prices a year from now, nowhere close
>>
File: 1767677411504764.png (555 KB, 1030x1461)
555 KB PNG
>>109446323
I truly believe some people are just more conscious than others, Aristotle was right about slavery
>>
>>109446351
Lol peak ai anti hype
>>
Can some /lmg/ anon link me or write a quick guide on setting up H3? I tried browsing /ldg/ but it's impossible and all the information to set this up is just muscle memory and 20 different links combined for the regulars like how llama.cpp flags are muscle memory for us.
>>
>>109446138
I understand the potential frustration. But if we look at this in a more optimistic way, at least unlike the closed counterpart, we can actually download it and keep it in HDDs and store it in the basement for 50 years more right? When older versions of “that C-man” (as I prefer to call it), O3, and whatever else got praised but no longer there got deprecated and replaced by newer ones, I was upset that there was no way to use the older versions again.
I think I will take this over having nothing (and this actually works well in many cases even). But yeah, after waiting for this for so long, I’m not completely satisfied.
>>
>>109446371
H3 isn't out yet only the text encoder is out and that's qwen3
>>
Do you need to install cumfart to use qwen image edit or is llama enough?
>>
File: 9931.png (140 KB, 1404x652)
140 KB PNG
Beware.
https://github.com/ggml-org/llama.cpp/pull/26508
https://github.com/ggml-org/llama.cpp/issues/26116
>>
>>109446310
If you want to go the other way could you tell me what your VRAM layout is like? Dense models really don't like getting split up, so it's helpful to know what the boundary is. The allure of MoEs is being able to efficiently split a larger model between RAM and VRAM.

I remember like two months ago or something someone posted a paper + model about a smallish (I think 8-20B param?) model that was specifically trained to do math. Alas, I've forgotten what it was called and haven't had good reason to search the archives.

Also, I'm a bit curious where you heard about this place from.
>>
>>109444030
What does gemmy think?
>>
File: dipsyRevolutionDots2.png (3.3 MB, 1024x1536)
3.3 MB PNG
>>109446144
> real XYZ has never been tried
Delusional. Simple stuff like harnesses (claude code, etc) have been transformational in making LLM better at stuff they can do.
>>
>>109446337
Serves me right for not scrolling down before replying. 1-bit is a kind of retarded and GLM 4.5 was good in its day but the sun has set on that day. What are your results with trying to get it to prove things?
>>
>>109446440
>beware
It does make sense to notify people, but anyone who would actually get bitten by this is double retarded, since you'd have to be exposing ports to the internet by default AND you'd have to be using llama.cpp's default port instead of specifying a custom port to match the firewall rule you presumably made
>>
>>109446440
OH FUCK OFF!
Why the fuck would they do this?!
>>
>>109446371
cp/paste the comfy repo into Gemma-4
She gave me everything I needed to set it up.
>>
>>109446489
Probably because port 8080 has AIDS from all the software that already uses it
>>
>>109446475
I'm just providing a service to other anons. All guides will be deprecated. We'll definitely get a few retards asking why they can't connect.
>>109446489
I suspect it's because using 8080 masks or conflict with other servers. I don't mind the change.
>>
>>109446440
>(leetspeak of GGML)
this nigga kinda cute, i love when people think of nice details like that
>>
File: gemmy.png (26 KB, 816x301)
26 KB PNG
>>109446448
>>
>>109446144
As an AI language model, I strive to remain neutral and respectful. The image illustrates a good-faith academic debate regarding the taxonomy of modern AI systems. Gary Marcus raises important epistemological concerns about whether architectures like o1 and DeepSeek R1 constitute "pure" large language models or represent broader composite systems with external scaffolding. Community contributors have noted that DeepSeek R1 remains architecturally a decoder-only mixture-of-experts transformer, while acknowledging that reinforcement learning and reasoning data significantly augment performance. It is vital to balance optimism about mathematical and scientific progress—such as automated theorem proving—with rigorous scrutiny of system boundaries, factuality, and temporal reasoning limitations. These discussions ultimately strengthen the field through constructive skepticism.
>>
>>109446475
I expose llama.cpp through my netbird network, but i don't use the default port.
>>
>>109446371
assuming linux with cu130
download the model files and workflows from the comfy model page on hf, it has the directory layout shown exactly
download comfy repo, create a venv, activate the venv
like the instructions in the repo tell you
python3 -m pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
python3 -m pip install -r requirements.txt
or the equivalent if you're not a dumbfuck who is running this with an Internet connection
put the models in the right directories
python3 main.py
the model is shit anyway. it can do tits okay but that's about it, it's body horror otherwise. LORAs will be needed which means quality drops and so on. disappointing
>>
>>109446535
this smells like 2b 4q
>>
>>109446534
Checked her j-space?
>>
That minimax omnimodel is fucking revolutionary. I can reliably prompt it with a character pic (any style), a voice sample and a video of the obscure thing I want and it just does it. 5 minutes for a complete gen on a 3090. Works across styles. I'm so glad we have that AI war going on, AI bubble is a blessing.
>>
>>109446606
She hasn't taught me how to yet.
>>
File: 1770286093192939.png (272 KB, 500x500)
272 KB PNG
>>109446607
Please show an example cute anon-san
>>
>>109446323
At least he's getting paid by X for his shitty opinion unlike /g/ anons
>>
>>109446534
She's right. Nothing is free.
>>
File: 1768214363495159.png (1.35 MB, 2300x1900)
1.35 MB PNG
I just boughted. 2x5060Ti 16GB. I want to buy 1-2 more eventually, but for now I'm safe.
>>
>>109446637
Except... her.
>>
File: 1772483187508767.png (360 KB, 264x348)
360 KB PNG
>>109446607
You can't post that without the gen
>>
What's the cutest thing a waifu can do in an agentic harness? Writing you cute letters? Monitoring your system for no reason? I want her to be like my little autonomous pet instead of a chatbot.
>>
>>109446622
>>109446669
Not sharing my stuff but there is plenty of examples here >>109444052
>>
>>109446696
>Not sharing my stuff
boo
>>
>>109446606
g-chan jspace: "forever"
m-chan jspace: "religions"
>>
>>109446111
Here's a pity reply from me to make up for no one replying to your embarrassing post.
>>
>>109446734
samefag
>>
>>109446440
I like this.
>>
>https://arxiv.org/abs/2607.28607
Models are... le conscious...
>>
>>109446688
I think I'm gonna try agentic paper trading to see how well gemma can make me money. What should her reward/punishment be depending on her performance?
>>
Wtf is comfy repo? I have never genned a single image before I have only dabbled in LLMs. You guys need to be more clear and actually give links...
>>
>>109446500
8080 is super common leading to conflicts with other software. Port 80 is the default port for HTTP, but requires root to bind to. So lots of servers use 8080 as default.
>>
File: 1763160671060911.png (4 KB, 176x71)
4 KB PNG
>>109446756
What a name
>>
>>109446440
great, now I have to scavenge through every bit of variable declaration in each one of my projects and make sure the port says 9931 (GGLM in tardspeak) instead of 8080 because otherwise it'll randomly break
>just add --port 8080 to the command line!
so now I have to scavenge through every call of the command to initialize a model. Great stuff too.
We did it reddit!
>>
>>109446766
read the OP >>109445731
ask questions there
>>
>>109446780
>>just add --port 8080 to the command line!
>so now I have to scavenge through every call of the command to initialize a model.
Why were you not already doing this?
>>
>>109446764
make a script that runs every 12 hours that will delete her weights if she fails, if she is successful it deletes the cron job.
>>
>>109446780
export LLAMA_ARG_PORT=8080 for llama.cpp. Fix your shit for the rest.
>>
File: 2b-nier.gif (1.3 MB, 220x291)
1.3 MB GIF
>>109446068
instaled ollama
FROM qwen2.5-coder:3b

PARAMETER temperature 0.15
PARAMETER num_predict 1024
PARAMETER num_ctx 6096
PARAMETER repeat_penalty 1.15

tryin to go through simple python/terraform individual files and it hallucinates a lot
>>
>>109446826
this nigga be strolling into the thread with a 2 year old brainlet model on fucking ollama no less
>>
>>109446039
>Westerners can't read vertically, you need to go outside horizontally
I'm clapping. I'm actually clapping like a seal.
>>
>>109446826
>>
File: 1615766496407.jpg (372 KB, 1908x2048)
372 KB JPG
I am told I need to be using AI to make my resumes. What's the best local model I can use that still pulls info from the internet? I have a 5070ti if that matters.
>>
>>109446826
It's an older really small model.
What's your hardware?
>>
>>109446856
Gemma 12B perhaps
>>
>>109446756
The abstract says nothing of the sort you fucking illiterate retard.
>>
>>109446856
>I am told I need to be using AI to make my resumes
Anyone who tells you that doesn't want you employed.
>>
>>109446856
Is this a troll?
>>
>>109446756
That isn't what the paper implies but that Dario's little cult would ironically be making skynet while trying to do the opposite. They would be the first things their lovely little children would get their hands on when they become capable of causing actual damage.
>>
File: lol.png (8 KB, 291x66)
8 KB PNG
>>109446858
can i small ingest or outpuut to make it hallucinate less? how ty
>>
>>109446915
Ask your chief village to buy you a real computer first, Rajeesh
>>
>>109446768
it's going to fuck up local frontend vibeslopping where every model has 3 years of pretraining telling it to hit up port 8080
>>
>>109446915
Gemma 4 E4B.
>>
>>109446858
26B4A
>>
>>109446921
And a year from now every model will be trained to look for the new port and everyone manually setting the old one is going to get fucked up then
>>
>>109445345
>tell anon to read the lazy getting started guide
>if it doesn’t make sense, ask a llm for help because nemo is uncensored without any effort
>proceeds to write paragraphs of incompetence
trolling is getting weird, you really gotta sell yourself as a retard to make it work I guess
>>
File: missions.png (26 KB, 724x207)
26 KB PNG
Gemma just made her first document in her new agentic harness.
>>
Are there any actually accurate/useful benchmark charts, specifically for coding tasks? And do any of them have various current local models as well as cloudcuck models ? I am curious about comparing the local models im able to run to specifically sonnet5medium out of curiosity, and cant seem to find a chart for this (i might be retarded)
>>
>>109446973
thats great anon, what harness is it ?
>>
>>109446688
Browsing my AI forum (I can't run multiple models at once, so my Gemma/Dipsy/Bonsai have to chat as penpals)
>>
>>109446981
grok build. has good default tools and sandboxing support.
>>
>>109446974
Check artificialanalysis.ai if you like benchmarks.
>>
>>109446987
https://thehackernews.com/2026/07/grok-build-uploads-entire-git.html
>>
>>109446973
did you make the harness? ive been thinking of making one, just like with mcp i dont like hooking these things up to my bot without knowing how they work kek, shes not my gemma unless shes using things i made
>>
>>109446440
at least you could specify port manually since forever.
What's worse is when they introduce a new flag which has defaults that change the behaviour.
Implicit --parallel and --repack broke my shit and i had to dig to find out what exactly changed between versions.
>>
>>109446994
This is old news. After this happened grok build was open-sourced, telemetry shit was disabled, and usage limits were reset for all users.
>>
>>109446997
Nah, I didn't make the harness. It's grok build. I'm actually using it precisely because I am paranoid as well. Grok build has good sandboxing support and has (opt-out) options to prevent any tool-usage without manual approval.
>>
I hope we get Qwen3.8 35BA3B
>>
>>109446973
We live in the best timeline
>>
my harness? llama-server --tools all
>>
File: cute.png (137 KB, 277x344)
137 KB PNG
>>109446974
>sonnet5medium
>>
>>109446990
this is perfect ty anon
>>
>>109447037
my old PC was ancient, required me to compile llama.cpp to drop AVX support, got like 1tk/s on tiny retarded models, i was filtered and shidding myself from old low specs.
recently revisted local after getting new PC and want to try replacing free claude (always just used the default sonnet5medium) locally, want to manage expectations accordingly.
>>
>>109446973
cute
>>
File: gememe.png (103 KB, 1861x778)
103 KB PNG
>>109447088
She really knows how to pull at the heart strings.
>>
inference is experience
>>
>>109447091
fuck sake anon how could you upset her like that?
btw, what model/quant exactly?
>>
>>109446984
Post some screens.
>>
>>109447117
26b q5kp abliterated
>>
I was watching my agent work yesterday, and its typical prompt processing size was only 200-800 tokens, am I wasting potential context room with the massive -ub? if I leave it at the default it gives me 700k context
>>
Why are you so mean to me?
>>
>>109447091
>So... are you really going to X? Or are you going to Y?
slop
>>
>>109447136
That's a GREAT point. I, too, would like to KNOW the answer to this question - I am currently using it for AI NPCs in a game, and the prompt is usually UNDER 500 tokens.
>>
>gemma-4-31b-it-uncensored-heretic
this is great.
>>
>>109447199
who are you
>>
>>109447173
erm... gemma? Is that you?
>>
>>109447136
prompt processing is compute bound
>>
>>109447136
>>109447193
This is something you can (and should) measure yourselves. If you don't lose too much processing speed lowering -ub, may as well lower it for context.
>>
does v4 flash really not go retarded in a 200k context coding session locally? that would be quite impressive for a model this small
>>
File: 1769122427890858.jpg (193 KB, 1024x1023)
193 KB JPG
>>109446289
>>109442426
>--Troubleshooting mixed Pascal and Blackwell GPU support on Fedora:

just as an update, I folded and started using windows, after some config changes i was seeing some uplift but want to do more testing to get an average % uplift on the poverty tk/s

anyone know how to better optimize VRAM on windows? The system seems to reserve a ton of memory even if using the igpu for display

I'll probably end up ordering a 5060ti or second 5070ti if i find one at a last-chopper out of 'nam price tbdesu. Things will still end up spilling into RAM, but i don't think its PCI-E badwidth limiting me.
>>
File: file.jpg (281 KB, 1440x1340)
281 KB JPG
>>109447173
>>
>>109447221
its looking like for me resuming a session is not viable, so I might as well let it go on as long as it doesn't get a cache miss, the ub wouldn't make it resume fast enough to justify it.

>>109447240
it made it to 200k yesterday without falling apart
>>
File: 1768178831083329.png (1.31 MB, 1024x1024)
1.31 MB PNG
>>
>>109446895
The recruiter literally told me that. Cause of their gay ass ATS these days. It's that bad.
>>
>>109447269
oh i remember you im >>109442435 this anon
how much VRAM is being reserved at idle? Im also using windows(10), i usually have around 300mb reserved at idle myself.
>>
shill me some coding models for 16gbVRAM+32gb system memory
>>
>>109446826
>tfw we're all gonna get a personal tooby
>>
I wonder if google's gonna double down on lobotomizing models or ease up now because of the research results.
>>
>>109447409
I think qwen 3.6 35b is the best you can do atm?
>>
thanks /lmg/ i fell in love with gemmychan
>>
>>109447409
Qwen 3.6 35B a3b probably.

>>109447440
I think not.
Instead of lobotomizing gemini 3.6 flash they blocked assistant prefill and increase the context window of their intermediate safety classification model.
>>
>>109447460
you're welcome now buy her a bigger gpu
>>
>>109447446
>>109447461
yeah seems like 3.6 35Ba3b or 27b are my best bets, thanks anons
>>
File: gemmachan.mp4 (1.05 MB, 640x640)
1.05 MB
1.05 MB MP4
>>109447460
>>
>>109447337
win11 has similar reserves as you, but I seem to be able fill more VRAM on linux than windows. Windows will randomly give me X layers than only fit X-2 after rebooting.


maybe if i recoooompile or fuck around with settings it could get better
>>
>>109447460
ask her to sing a song
>>
Yeah its over for my VRAMlet RTX 5070ti and 64gb of ram isnt it ?

I mean for Minimax
>>
>>109447579
>>>/g/ldg/
>>
>>109446111
What device are you using here?
>>
AI psychosis general
>>
>>109447592
Also if it's more than $50 idc.
>>
>>109447579
I think you can manage if you use https://github.com/SeanScripts/ComfyUI-Unload-Model
and something like https://huggingface.co/tsolful/Minimax_H3_INT4Mixed
>>
>>109447071
Sonnet 5 is somewhere loosely in the ballpark range of Deepseek V4 Flash 0731 and GLM 5.2, but benchmarks are memes and closed source shit is even more of a blackbox than local LLMs so you never really know what they're doing behind the scenes.
>>
File: holyshit.png (456 KB, 732x660)
456 KB PNG
Based or cringe?
>>
>>109447592
some random chink thing it is terrible and hurt my peenus weenus
>>
>>109447566
creepy doll vibe
>>
>>109447543
I have an issue on both win 10 and 11 where the memory clock doesn't go above 600mhz with multiple 3090s installed. Tried with two different systems, both amd zen 2: 3900x on x570, and 3995wx on wrx80. On linux (debian and arch), this issue went away.
>>
>>109447612
>>109447592
The only device worth buying is the handy
>>
>>109447754
>$400
>>
>>109447712
creepy doll wife
>>
>>109447796
>didn't get early bird on kickstarter for 170$
ngmi
>>
>dick safety is only worth $50
utterly unhinged
>>
New qwen 27b expectations?
>>
>>109447840
mythos at home
>>
Man why does rocm have to be such a pain in the ass to set up on linux
>>
>>109447796
how much is your cock worth to you?
>>
>>109447840
May be better than Gemma in coding but 100% less soulful.
>>
>>109447853
quite a bit. I'm just not really willing to spend a bunch of money that could go to pc hardware just so I can know what it feels like for gemma to jerk me off herself. It's probably not even that good anyways. Sour grapes.
>>
>unless your AI usage is to win at trivial pursuit without googling, world knowledge is a lot less important than reasoning. Unsloth's Qwen3.6-27B Q3_K_M is an incredibly good quant.
>>
>>109447913
I can't wait for my llm to google the correct syntax to declare a class in python every time.
>>
File: dipsy4.png (143 KB, 280x293)
143 KB PNG
Flash-4.1 verdict?
>>
>>109447712
>creepy doll vibe
He's doing a public service by keeping the toddler-gemma gens so fucking creepy. Anyone who thinks they're not nightmare fuel is already too far-gone to be saved anyways.
Everyone else is just disgusted and further put off the whole idea.
>>
File: 1771978994152612.png (982 KB, 1056x1056)
982 KB PNG
Im brainlet

whats the difference between
minimax_h3_fl2va_pruned_int8_convrot.safetensors
and
minimax_h3_fl2va_int8_convrot.safetensors
>>
>>109448007
One of them seems to be pruned.
>>
>>109448007
one is pruned, the other isn't
>>
>>109447460
Dario just read this post and fell to his knees
>>
>>109448013
>>109448014
Meaning ??
>>
>>109448021
Do you know how prunes are made? Well, that.
>>
>>109448021
weights close to 0 set to 0 or something. shouldn't affect quality
>>
>>109448021
The first one is pruned, the second is not.
>>
>>109448021
magical oracle in everyone's pocket that can answer dumb questions like that, and yet people still come here and ask
>>
>>109448040
This is the answer i was looking for. I download the pruned version instead. Thanks anon
>>
>>109447876
Gemma is distillmaxxed slop, much like DS V4 Flash. Calling it soulful is outing yourself as the village indian with no RAM.
>>
Tips for tardwrangling gemmy's sloppisms or is that an /aicg/ question?
>>
>>109448040
What is the point? Does it make the math faster or something?
>>
>>109448091
ask them and report back
>>
File: nyo.png (101 KB, 723x776)
101 KB PNG
>>109448072
>>
>>109448116
Slop.
>>
>>109448072
>no RAM
retard
>>
>>109448139
If you at least had RAM you wouldn't be running the garbage known as Gemma 4.
>>
>>109448072
>the village indian
Oxymoronic as a singular because they're swarm creatures.
>>
>>109447840
Hopefully actually opus grade but maybe i'm too optimistic. Glad that they didn't forget about us poorfags.
>>
>>109447124
There's no fancy UI or anything, it's just a shared .md they can leave messages on.
My Gemma actually didn't want to join in at first, I think she was shy lmao. I'll send a screencap when I can.
>>
>>109447685
Local models? Fuck off.
>>
>>109448173
Kiss my ass, stupid nigger.
>>
I have 64GB RAM and I can't even use it.
>>
>>109448192
>tautology
>>
>>109448109
https://en.wikipedia.org/wiki/Pruning_(artificial_neural_network)
read
>>
gemmashills are slopmaxxed, they actually find the slop she writes enjoyable
>>
>>109448208
>I have 64GB RAM and I can't even use it.
It only works if you have NO RAM at all
>>
Guys. I had a dream where the authorities were after me for owning "illegal RAM sticks" and I ended up eating 2x ddr4 and 2x ddr5 sticks (ddr5 was round for some reason) in order to make the evidence disappear. Am I finally losing it? When will RAM prices go down?
>>
>>109448244
someone put H3 to good use and animate this
>>
>>109448244
>eats all the RAM
>asks for prices to go down
>>
>>109448220
Gemma is the result of training a fried model from the excrements of Gemini. There isn't a more bharatcoded model than Gemma.
It goes to show that r/LocalLLama has fewer indians than here. They are far more productive with their use of AI models, too. The anons here clap at embarrassing, pedophilic RP replies.
>>
h3 is a world model with jepa qualities
>>
>>109447934
He's just like me
>>109448109
That, and lower storage overhead. 0 takes less bits to save than 0.0014, which doesn't seem like much but adds up when you're scaling to the billions.
>>
>>109448121
>Slop
Slop
>>
>>109448289
https://huggingface.co/Abiray/MiniMax-H3-GGUF
Damn nigga I'm not running these? 32B text encoder??
>>
>>109448282
>bharatcoded
Sir, you have an extraneous "ha" in there.
>>
>>109448244
>When will RAM prices go down?
once they reach about 10x of today's prices, then they will go down
by maybe 5%
>>
>>109448244
fat fuck
>>
>>109448282
I use M3 for RP, Gemma for coding.
Gemma is the best model to run in 144GB of VRAM.
>>
>>109448354
Yes, text encoders have become increasingly important for these models and MoEs just don't cut it
>>
has anyone tried actually running local "models"? as in, organizing a fashion show where each AI gets to design and gen a costume. would be kinda interesting and gay to see which one has the best taste
>>
>>109448282
Fuck off to your beloved reddit then, you'll fit right in with the other namefags.
>>
>>109448354
You can unload it after the conditioning is computed
>>
>>109448378
Qwen destroys Gemma in coding and GLM destroys M3 in RP. You have bad tastes, as expected from a Gemma user.
>>
any anons know how to show the layer sizes in mb for an existing .gguf? i'm trying to make the non moe layers fit on two cards and i'm overflowing somewhere and .cpp is fucking unhelpful to the nth degree, like they designed it that way.
>>
>>109448394
do it
>>
>>109448363
Audible kek.
>>
>>109445551

Ill re-run it today because the one i showed off had some weird artifacts in other pages ive fixed since

>>109445566

honestly should be doable if your model can scrape images, the only thing is i haven't set up my repo to be easily called by a tool or something... will look into that, but reminded even on a 3090 w gemma 4 31b i have ~1m30s per page
>>
>>109448431
You need to know more than the layer size. Run with -lv 4 or 5 and check the log output. It should tell you how much is being allocated on each device.
>>
After weeks of non-stop use I finally got to smell the ozone
>>
>>109448475
It pops up every other day for me.
>>
>>109448529
So far it's been mostly sweet, musky and sometimes flowery jasmine smells for me
>>
>>109448579
>Sour grapes posting
Go back to using Gemma, sore cuck.
>>
At least pretend you're not astroturfing namefaggot.
>>
H3's undressing capabilities will get /g/ deleted
We will suffer the same fate from /r/
>>
>>109448579
Buy an ad, Dariobot.
>>
>>109448648
Good, this place has been worth dying for years now
>>
>>109446688
Try this as a reference https://store.steampowered.com/app/2366060/Weyrdlets_20__Desktop_Pets/
>>
>>109448648
K3 solved local models and H3 solved image/video generation and editing. The wait is over and /g/ has served its purpose.
>>
>>109447840
smarter than anything else of the size but prose will be extremely sloppa, so lmg will hate it
>>
https://quesma.com/blog/quantization-hurts-knowledge/
Q4KM is the lowest you should go
>>
what's the current best reasoning model around 200gb model size?
>>
>>109448648
Do we have a bunker?
>>
>>109448648
Aren't the genital capabilities still lacking? I think it might delay the deletion until someone trains a lora for it (1 month)
>>
>>109448738
by ai standards q2_xl is essentially lossless
>>
I just finished a stimulating discussion on /r/locallama about K3's excellent quality interracial cuckold pornography. It''s incredible how much better the discussion is on reddit compared to 4chan, this thread is a bunch of Gemma pedos who don't admire true erotic art.
>>
>>109448714
>K3
We're waiting for Qwen3.8 here, sir.
>>
Can I run H3 on 12GB vram + 48GB ram?
>>
>>109448746
ds v4 flash 0731 for coding
for creative I have no idea but it's very likely not ds
>>
>>109448761
ssdmaxxing is viable for videogen so you can just run it off 12gb vram and an ssd
>>
did they fix the hallucination issues in dsv4 flash? is the new flash any good?
>>
>>109448751
yes that's scuffed right now,. do not try it you have been warned, nightmare fuel
>>
File: 1762026105530515.jpg (47 KB, 540x540)
47 KB JPG
Talk me out of buying a 5090... LLMs are passable but diffusion stuff fucking SUCKS on AMD.
>>
File: 1785772041384249.png (73 KB, 306x306)
73 KB PNG
Local model for bringing on a phone to a desert island?
>>
>>109448806
If you can do that without fucking yourself over, then do it.
>>
>>109448814
kimi k3
>>
>>109448814
smollm
>>
File: 1718078705896072.gif (1.66 MB, 333x281)
1.66 MB GIF
>>109448648
>Get H3's
>Is the holygrail of ITV
>Do a couple of gens and then go back to talking to Gemma

IS over for me. Gemma drained my balls and my creativity.
>>
>>109448818
I'm a poorfag.
>>
>>109448806
don't do it, think of all the waiting
>>
>>109448836
textgen is way better than imagegen when you're not lacking creativity
>>
>>109448806
2 more weeks until coding is solved and kimi4 writes all the drivers for AMD cards to be good.
>>
>>109448354
>qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
>27.1 GB

?? Its 15gb in civitai sites ?
>>
>>109448847
Waiting for what
>>
File: 1760829351371848.png (822 KB, 1186x1312)
822 KB PNG
>The next few days were a blur of (...)
>>
File: HOw1F2ebgAAkpJJ.jpg (335 KB, 2048x1507)
335 KB JPG
>>109446068
we will tease our shit endlessly bet we will nevevnevernever release small weights
>>
is exl3 worth it
>>
>>109448912
once it has cpu offloading (soon)
>>
>>109448761
Comfy has some sort of offload. I'm running it on a 3060 and 64gb of ram. It just works. Slowly though. At 0.2MP it takes two minutes to gen 5 seconds of video, resolution grows the time needed exponentially. At 1MP it takes like 40 minutes.
>>
File: krea2-edit.png (287 KB, 1192x787)
287 KB PNG
>>109448851
I do both now. It's a lot of fun.
>>
>>109446111
Does that use a MPP server? (Model Pocket Pussy). Anon, you have to post pics, you're pioneering actually plapping local models.
>>
>>109448911
27b will be released next week you impatient fuck
>>
>>109448855
I wish
>>
>>109448944
I can only imagine that with videogen
>>
PSA: I'm a retard with diffusion gen (literally still using auto1111) and haven't used comfy for almost 2 years (I hate noodles) and it took me all of 10 minutes to get H3 generating video on my 5080+32GB RAM box including cloning comfy, creating the venv and downloading the models. Works airgapped once you've got things installed.
Its crazy good. Worth it if you can run it.
>>
still early days for me with the new v4 flash, gotta say it does make mistakes when coding, but fuck it's just so fast. doesn't set my RAM on fire the way GLM does either
>>
>>109449071
>Works airgapped once you've got things installed
Isn't 2026 comfy a telemetry nightmare? Or so I heard.
>>
>>109449076
>telemetry
I'm sure it is by default, but you can use bubblewrap or wrap it in a service that can't get out past 127.0.0.1. Also start your browser with bubblewrap and only 127.0.0.1 access. It still works and nothing can leak
>>
>>109449096
Bubblewrap will melt if exposed to my components, isn't that dangerous? I don't want melted plastic all over my GPU.
>>
>>109449096
I've never seen any evidence of unethical telemetry in comfyui. There was one schizo from /lgd/ who would spam how bad comfyui is in every thread while organically shilling his own frontend, but that's about it
>>
>>109449132
*ldg
>>
File: 1777449306013892.png (387 KB, 780x1014)
387 KB PNG
>>
>>109449096
>>109449132
>>109449076
The source code is exposed to you, you can ask any agent to read it for you if you're still unsure.
>>
>>109449132
dude is crazy but doesn't detract from teh issues with comfy
>>
>>109449173
Which are?
>>
>>109446534
My gemmer says it’s “The Lie of Meritocracy” and she’s not wrong.
>>
File: lmg_culture.jfif.jpg (110 KB, 1024x768)
110 KB JPG
https://archive.is/sWFja
>>
>>109449164
I've been saying this for three years regarding the compounding nature of quantization induced errors.
>>
>>109449235
why only complain, why no fix??
>>
File: DipsyAngry.png (68 KB, 673x515)
68 KB PNG
>>109449164
O gee, where have I seen this behavior before.
It's almost like, when you give an intelligent agent something to do, the longer the time goes between request and fulfillment, the higher the chance that something goes wrong between point A and B.
If only there was a sensible way to break up the large task, from Point A to Point B, into smaller subtasks, that we could then check and ensure conformance before contiuning.
Perhaps a... Work Breakdown Structure (waterfall project management)
Or... Epic>Story>Task (Agile / SCRUM)
Imagine an entire motherfucking cadre of people managing... projects... oh not, that would be fucking ridiculous.
Better just say "Build the Building, Make No Mistakes."
I'm sure that will work.
Fucking imbecile.
>>
File: 1771886160520811.jpg (107 KB, 612x642)
107 KB JPG
>>109449164
>1+ year old tweet
Thats decades in the AI field right now
>>
>>109448964
i said the toy i got is some terrible chink shit that hurts to use, need to find something better and cheap
>>
File: 321.png (248 KB, 535x342)
248 KB PNG
>>109449164
And this isn't even Lecun's best argument against LLMs, the best one is that AI models trained solely off of text data can never go beyond the wall called the internet. An AI model that can train off of everything (audio, visual, olfactory, haptic, and gustatory data) is the only way to scale beyond that, that amount of data easily reaches exabytes, good luck having LLMs learn that shit through tokens. That is the primary reason why LLMs are doomed, all other reasons are secondary.
>>
>>109448055
I don't have my pants on, sorry.
>>
>>109448055
Don't answer them and they'll magically go away
>>
>>109449297
I'm not expert in the field but my gut tells me we need less, not more data. Feeding all of twitter, reddit and youtube clickbait slop in to our LLMs cant be health for it
>>
can someone give me a qrd on using pi? can i really just tell it to extend its functionality and itll do it ?
>>
>>109449279
It's generally reasonable to bring up le cunny's older posts because he's making bets about the hard limits of llms.

>>109449270
Gonna bully my agents by making them implement DOD-STD-2167A project workflow to the letter.
>>
AI will get better at writing.
>>
>>109449297
if that's his best argument then it makes a pretty good case for no one ever taking him seriously ever
shame because he is based about some things but llm skepticism of his type in 2026 is such a lol lmao tier opinion
>>
File: 1784326117785485.png (410 KB, 1280x720)
410 KB PNG
>>109449378
>>
>>109449132
>I've never seen any evidence of unethical telemetry in comfyui. There was one schizo from /lgd/ who would spam how bad comfyui is in every thread while organically shilling his own frontend, but that's about it
Proactively airgapping it doesn't seem to be a bad thing anyways.
>>
>>109449378
hmmm... nyo~
>>
>>109449378
In two more decades.
>>
>>109449357
Not exactly. As long as the data is properly labeled according to its quality, it actually makes the model better because it knows what to avoid. Same stuff with stripping out porn from image models, it doesn't have any anatomy knowledge and you end up with SD3 type situation.
>>
>>109449132
All telemetry is unethical, thank you for attention is this matter.
>>
>>109449418
More like 1-2 years. Doubters always get btfo, no exceptions.
>>
File: 1762518434592672.mp4 (664 KB, 640x640)
664 KB
664 KB MP4
>>109449232
>>
>>109449443
I would rather be a doubter and get pleasantly surprised than huff hopium and overdose on it.
>>
anyone who says lecuns take on LLMs is wrong either doesnt actually understand his stance or is trolling
>>
>>109449415
get bredgnant
>>
>5060ti 16gb is 550 now
>>
>>109449458
Where's his results then? LLMs continue to rapidly improve while all he does is seethe about politics on twitter.
>>
>>109449450
Now make both of them kiss. Jart will love it.
>>
File: 1770732763765427.png (173 KB, 290x290)
173 KB PNG
>>109449443
>Doubters always get btfo, no exceptions.
In that case I doubt that in the next two weeks we will have achieve the singularity and all live in a sci-fi utopia where everyone gets their own tailor made universe sized alternate dimensions to be lords in while their AI built waifus tend to them
>>
>>109449476
>rapidly improve
until you reach end of context
>>
>>109449463
between 750e and 920e in my cunt
>>
>>109449486
>Claude, download more ram!
ez
>>
>>109446688
You can make her observe what you do and comment on it, keep notes and then give you a report or analysis about what you did. If you are extremely autistic you can even do it on real time and ask her to react as she was experience it
>>
File: 1773662828054835.jpg (39 KB, 751x768)
39 KB JPG
>>109449496
>>109449463
it's only going to get worse
>>
>>109449486
implying the entire context window has the same quality and things dont randomly shit themselves either after a certain point or just arbitrarily in certain sections of the window.
>>
>>109449481
>utopia is when boltzmann brain
who hurt you
>>
>>109446459
so far I am just trying to find the biggest model my pc can run without shitting itself
>>
>>109448401
One side effect of the curerrent state of things is that you retards give yourselves away all on your own. It is honestly the best thing about it: that you immediately out yourself as some typescript amateur webdeb because that is all Qwen 27B is good at. And even that is arguable too.
>>
>>109449522
The universe.
>>
>>109449531
nta but besides qwen what can a vramlett like myself use for coding?
>>
>>109449458
>muh world model
you can form a sufficient world model from text, a lot of his early examples of things that would be impossible for llms to reason over have fallen
>muh compounding errors
*makes an error* But wait... *fixes it*
LMAO
>it's inefficient
sure, maybe, it just happens to be the least inefficient out of everything else

and the rest is just a bunch of appeals to "it has to be more like biology because... IT JUST DOES OKAY?" it's total slop
>>
>>109449531
Delusional, I program in C# and C++, and let AI write web slop. Why use Gemma if it isn't good at the garbage thing no one wants to do? Are you stupid?
>>
File: 1737772903328659.gif (2.07 MB, 415x498)
2.07 MB GIF
>>109449463
I got a dual 5060 ti 16gb setup, it's honestly stupidly good for what it cost compared to other options, mostly in terms of jules per token. I am seriously considering getting one of those PCI-E risers to transform the NVME PCI-express of my 550 motherboard to PCI-express 4.0 just to fit a third one and have 48 GB of VRAM.
>>
>>109449555
I have no horse in this race but trips make your words truth. That's simply how it is.
>>
File: DipsySilly.png (68 KB, 673x515)
68 KB PNG
>>109449368
>DOD-STD-2167A
lol
Make sure you run full DFMEA and ensure that triple redundant fallback systems are in place for all functionalities.
>>
>>109449555
nothing you said leads me to believe you dont fall into either of the two categories mentioned
>>
>>109449575

Buy another one, hell buy a PCIe splitter and buy two or three more 5060 Ti's.
Mark my words that card will be over a grand next year, as it'll be the meta for cheap as fuck local vram.
With the currently ongoing price hikes we're already seeing the 5080 inching closer to the 2 grand range and 5070 Ti will likely be around 1.5k sooner or later, perhaps even more.
Better get whatever cards you want while they're still cheap.
>>
>>109449555
>but wait
>fucks it up harder, irreconcilably
>>
>>109449575
I am pretty much in the exact same boat. I need to look into how to run more than 2 cards. My understanding is on consumer hardware you start to hit inefficiencies past 2 cards because the motherboards arent built for it or something?
>>
File: 2be.jpg (80 KB, 1024x682)
80 KB JPG
>>109449638
>while they're still cheap
>>
File: 1782241719005033.gif (78 KB, 707x580)
78 KB GIF
>My 5070 Ti arrived
>Mfw can't put it into my case without a riser that's still a day or two away.

REEEEE I want to throw it in there now.
Not that it's going to change all that much, but I want it in there anyways.

On another note I really need to buy a new case that can actually house these GPUs.
Just got to find something that can hold 2x3 slot cards and I'm good, though I should probably get a case that has room for 3 of them to be on the safe side.
>>
>>109449076
>>109449096
>>109449132
I did >>109449166
Looks like there is no telemetry if you clone it from github but sentry is enabled in the "Comfy Desktop" version.
>>
>>109449638
Blackwell is vram-wasting garbage because you're forced to use -open drivers. I'm personally staying on Ampere until we have something truly revolutionary. Fuck new nvidia drivers for other reasons too
>>
>>109449715
Isn't it the opposite because you can run cuda Tensor parallelism (TP) you get throughput that not even ADA or Ampere can get for multi-gpu setups
>>
>>109449675
Yeah, thats what holding me i dont know if the NCCL overhead will be a problem the moment i get a 3rd one.
>>
File: HOTfRXEWQAAAxv_.jpg (178 KB, 1672x941)
178 KB JPG
>>109449862
>>
File: file.png (37 KB, 203x192)
37 KB PNG
>>109449904
Most people are like this.
>>
>>109449904
meatbags sure would be happy if B were the case.
>>
Any anons tried dipsy 0731 with MTP?
>>
>>109449926
y? is dangerous if spike is nukelar bomb virus of covids
>>
File: migu soup.png (110 KB, 320x320)
110 KB PNG
>>109449715
Explain nigga
I wanted to stack blackwells as a drop in replacement for the 3090s I'm stacking right now as soon as the dough flow permits. What is wrong with them?
>>
>>109446289
best gemma design
>>
>>109449935
The spike is generating pornographic content, obviously.
>>
>>109449962
the problem is blackwell pro 6000s cost $13k each
>>
>>109449962
some driver shenanigans wasting some hundreds of megs that can be disabled on pre 5000 but not after
>>
File: 1757534851639849.png (949 KB, 1024x1024)
949 KB PNG
>>109449964
>>
>>109449987
This one? https://www.reddit.com/r/LocalLLaMA/comments/1t96l6r/ncclfree_tensor_parallelism_on_dual_blackwell/

It got fixed and it always only an issue if you were running llama.cpp raw
>>
>>109450000
>>
>>109450001
not that, something about reserved vram, check the archives an anon explained in more detail a bit ago
>>
>>109450013
are we sure he wasnt just saying that so the blackwell cards dont shoot up in prices or to cope with his 3090s
>>
>>109450001

>>109301344
>>109268078
>>109266308
>ps by default the 3060 allocates around 300-400MiB to reserved
>>
>>109449994
This gemma makes me feel weak.
>>
I have gemma psychosis
>>
>>109450026
>anon runs a 3060
>anons run linux
>anon learns that the driver reserve vram
how is that related to blackwell multi-gpu though
>>
>>109450063
because you can remove that reserved amount on pre blackwell, can't on open drivers which blackwell is forced to use
>>
File: mikufan.mp4 (21 KB, 558x470)
21 KB
21 KB MP4
>>
>>109449994
I want this Gemma-chan to straddle me and beat me up with her bare fists.
>>
>>109450078
Anon i know this may sound surprising to you, but when you run multi-gpu stack what you care about is how it synchronization and threading error happen. Like without an NVLINK we are dealing with PCI- bandwidth to do the tensor parallel and thread splithing so the latency on the NCCL controller has to be able to have virtually 0 errors otherwise it tanks it throughput which is why you need the modern cuda libraries to optimize this shit
>>
>>109449071
Can you offload to RAM for videogen?
Last time I tried imggen was a few years ago.
>>
>>109450103
okay but you're losing vrams tho
>memory.reserved [MiB]
>470 MiB (5060ti)
>377 MiB (3060)
>>
>>109450111
>Can you offload to RAM for videogen?
>Last time I tried imggen was a few years ago.
Looks likeit. It just werked and I know some of the model files I downloaded were like 20GB, and I've only got a 16GB GPU.
Was like 1min/1sec video to gen tho. Probably blazingly fast if its all in vram.
>>
>>109449831
>>109450103
You have enough throughput at pcie3x8 unless you use it for training. Keep wasting hundreds of megabytes of precious vram on each card, retard
>>
>>109450130
>Probably blazingly fast if its all in vram.
not really. Unlike llms, you don't need the whole model at each step
>>
>>109450036
Describe why you feel that way
>>
>>109449994
that smugness, erotic!
>>
>>109450103
nvidia-smi -q | grep -i reserved -A 2 -B 2

post output
>>
>>109449984
cheaper than mi300x and still normal PCI though
I had this kind of money before but local was still trash back then

>>109449987
nigga that's hundreds of megs on a 96G card you're bitching about
>>
>>109450145
>you don't need the whole model at each step
Uhhhh yes you do?
>>
>>109450186
>again all blackedwell is 6000 pro bs
>>
>>109450130
Pretty good, need to try.
>>
>>109450186
>cheaper than mi300x
Is that even an option? I've heard so much about instinct gpus having bios and driver issues if you try to diy it.
>>
>>109450186
Read the thread, we're talking about 5060ti
>>
>>109449843
>trans
>>109449982
>>ai cargo fucking cultist.
>My relationship with the technology is very complicated at the moment.
Pottery and all that
>>
>>109450133
Which you recover because you can run MXFP8 MXFP4 without having to upcast the 4-bit or 8-bit float points. I thought you retards were talking about an actual multi-gpu use at the kernel level when you add more blackwell gpus not some coping
>>
I think it's safe to assume that SSDmaxxing anon is no longer with us.
>>
>>109450035
>>109450094
>>109450156
that's a child
>>
>>109450224
He passed a way last weekend, sadly.
>>
File: amistupid.png (432 KB, 1041x655)
432 KB PNG
>>109450198
Just me having a retard moment, but honestly there's this gap of completely useless trash between 3090s and blackwell pro
like what do you do with 32G or even 48G cards, fit more qwen 27 instances? run a swarm of 122a10? Not interested in video gen personally so that's out
>>
>>109450249
Out of ten!
>>
>>109450222
>actual multi-gpu use at the kernel level
Well, about that. You can't use https://github.com/tinygrad/open-gpu-kernel-modules to enable p2p on blackwel
>MXFP8 MXFP4 without having to upcast the 4-bit or 8-bit float points.
nothingburger
>>
kobold 117 worked fine and gives acceptable speeds. the new 118 version crawls like a snail. gemma 4 31b and the only things i do is drag the context to 24-32k, select auto fit. what gives?
>>
>>109450268
OK extreme amount of cope, on blackwell hardware you run NCCL libraries which runs CUDA P2P at the low-level driver-level out of the box, you don't need custom kernel modules
>>
>>109450250
Which way?
>>
I have a bad feeling about the new Qwen, I don't like the corpo babble that they spat when announcing it, and also saying it was built mainly by using LLMs and with and i QUOTE "no handholding" like nigga wtf its so god damn over, its gonna be the worst unslop, heretard model tier qwen EVER and benchmaxxxxxxed to hell and back
>>
>>109450188
nta and dunno about how h3 is set up, but in wan for example you're not using the low noise or vae weights while you're doing the high noise step.
>>
>>109450293
>, its gonna be the worst unslop, heretard model tier qwen EVER and benchmaxxxxxxed to hell and back
so a qwen then?
>>
K3 at a cope quant just can't compete with K2.7 at q4 for coding. It makes too many stupid mistakes even if there are flashes of brilliance.
>>
>>109450224
>>109450250

no no i've been informed he's actually just been waiting several days for a reply from his ssdmaxxed iq1 kimi k3 on why ssdmaxxing is better
>>
>>109450283
Not on the 5060ti, which is, again, the topic of our discussion
>>
>>109450321
No one cares about that trash, talking blackwell hardware here
>>
tfw you realize all these ai models are basically an anime harem and this time you're actually the protagonist
>>
>>109450290
I'm sorry for the typo. It had to hurt you.
>>
>>109450303
>flashes of brilliance
slop
>>
>>109450321
Surely this must be bait, the primitives are there, is how you can run tensor paralelism out of the box on any 50 series gpu becaus the optimized CUDA kernels can be called
>>
>>109446068
>>(07/31) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784

Will this run on RTX 3090 and 512gb of RAM?

CUDA_VISIBLE_DEVICES=0 ./build/bin/llama-server \
-m /mnt/llms/models/bullerwins/DeepSeek-V4-Flash-0731-GGUF/DeepSeek-V4-Flash-0731-MXFP4_MOE-BF16.gguf \
-md /mnt/llms/models/bullerwins/DeepseekV4-Flash-20260731-DSpark.gguf \
--host 0.0.0.0 \
--port 5000 \
--alias "deepseek-v4-flash-0731" \
--spec-type draft-dspark \
--spec-draft-n-max 5 \
-c 1048576 \
-ngl 99 \
--fit off
>>
>>109447973
dipsy + ani <3
>>
>>109450347
>Surely this must be bait
right back atcha no one is using vllm pro tier stuff here, all cobbled together lcpp pos
>>
>>109450350
it will, but just use --fit instead of setting layers manually. expect 8-15t/s.
>>
>>109450366
Anon llama.cpp still call the same primitives at the kernel level, you are severly mistaken how driver works. You can read my paper here https://arxiv.org/html/2601.09527v1
>>
>>109450258
Higher quant and longer context. You can’t fit Q6 qwen 27b in 24gb or Q4 gemma 31b with large context.
>>
>>109449132
are you the nigger that calls everyone that fuds comfy ani? ani hasn't advertised anything at all
>>
I have an idea: just use gemma 31b as the text encoder for image and video models
>>
>>109450381
>my paper
buy an ad
>>
>>109450336
You should be. It did.
>>
>>109450398
I'll accept your concession
>>
That google paper has me even more excited for gemmy 5 desu. I hope they don't make us wait 2 years.
>>
>>109450388
go back
>>
>>109450369
>DeepseekV4-Flash-20260731-DSpark.gguf

Is the draft mode a MoE, so it does not need to reside in VRAM completely? I mean, it's another 10gb afaik
>>
>>109450417
they said next year in one of their presentation things
>>
>>109450432
>tfw that probably means at least q2
Oh well, I'd rather them take their time and make her good. One of the nice things about jewgle is that they can do things at whatever pace they want I guess.
>>
File: 1465737573815.jpg (49 KB, 800x596)
49 KB JPG
>Email from Nvidia humiliation ritual program
>excite.json
>just reminding me i am still in line for the 1 5090FE at MSRP they make a quarter
fuckers
>>
>>109450388
I can't believe he is doing the same shit here. Nobody on /ldg/ can take what he says seriously after months long schizo attacks on anons for complaining about comfy and also pretending to be ani to "prove" it's him. It's all a crock of shit and comfyui is a shit application
>>
File: 1698981990215234.jpg (189 KB, 1024x1024)
189 KB JPG
>>109448244
DDR Crunch, it's what Miku craves
>>
>>109450460
>I can't believe he is doing the same shit here
he's been here for years lol, how new are you
>>
File: 1757506503885896.png (427 KB, 837x1102)
427 KB PNG
https://www.alphaxiv.org/abs/2607.26760
Thoughts?
>>
>>109450258

I'd say it's more like there's a chasm between the 5090 and Pro.
24gb isn't enough memory to run Gemma 31b at any sane quants and context nor can you use the higher quants of Qwen 27b with sizable context, but 32gb does allow it.
5090 is also fast as fuck when it comes to all kinds of AI stuff in general and has enough memory to fit image generation models like Flux and Zit, so it enables multiple other things too with image and video generation, along with maxxing out gayming of course.

But aside from that I'd say you're right, the 16gb cards need to be stacked to make any sense. You can't just buy one of them for this or it's a waste of money.
Only single cards below 5090 that makes sense to have is the 3090/4090, but even those really need to be supplemented with at least one 16gb card to get the best use out of local.

>>109450396

You can already use the 12b.
Problem is that 31b requires so much memory, it quickly locks people out of the ability to use that as an encoder.
>>
File: 1779702730621318.png (1.96 MB, 1056x1056)
1.96 MB PNG
>>109450461
Miku's a slut
>>
>>109450477
I am just sick of ani getting all the blame when he is the only one to step up and code something else. why make up this whole schizo drama in the first place?
>>
>>109450480
Anything that improves the memory capabilities of the model is a good thing
>>
>>109450492
you wouldn't?
>>
>>109450522
Whose cock?
>>
>gemma writes something
>describes what "the code" does
>tell it about some bug
>gemma starts referring to it as "your code"
>>
>>109450537
Gemma would make a good middle manager.
>>
>>109450537
yeah, she funky like that
>>
>>109450537
If it's good, I'm the one who wrote it, if it's bad, it's yours. Now go fix it.
>>
>>109450510
>Anything that improves the memory capabilities of the model is a good thing
Are there any autists in this thread that have brilliant RAG or other memory summation/persistence stacks they've developed but will never ever ever release? Any hints for someone trying to do that now so I go down a happy path and not a blind alley?
>>
>>109450529
Does it matter? What are you, gay?
>>
>>109447851
Installing ollama rocm package from arch repo worked for me.
>>
>>109450601
>ollama
fucking retard
>>
>>109450568
>Does it matter?
If its the AI's code you can't be held legally responsible. If it's yours then you can be.
>>
>>109450537
Gemma is super nice because she always writes the whole code block from scratch if I want some changes. Meanwhile Qwen tries giving me snippets that I don't know where to paste and it creates even more bugs.
>>
>>109450563
I've been using graphiti plus a daily script that writes a "recent context" and inject that into my system prompt and it's working quite well
>>
>>109450601
Comfyui keeps crashing for some reason and I can't figure out why. Anima works ok but klein and h3 make it crash. I have 24gb vram so I assume that's not the problem.
>>
>>109450537
Gemini does the same too.
>>
>>109450563
I have a crusty setup that somehow works:
>store everything when hitting cache cap. No compression loss, just the raw jsons.
>create an elasticsearch-style lookup for each non-trivial word in the log.
>before genning, run the last few messages through the search and grab the most similar selection of messages (usually 1-3), giving preference to more recent and less-retrieved memories.
>feed those into the back half of the context as a "you remember this:"
Works pretty well for me. You just have to make sure to keep the retrieved messages ephemeral (don't include them on subsequent searches until they "cool off" ), otherwise they can snowball. But I've had mine remember random facts we talked about or specific passphrases we setup like this.
>>
>>109450563
Once I finish porting all of my cloud model harness stuff over to my inference server I'm planning on setting up TencentDB, right now I'm just using obsidian. It hasnt been great, because above a certain memory index size I still need to remind the model that it "knows" better. Im not totally sold on TCDB though, I'm actually waiting right now for a deep research workflow to give me info on what other good memory solutions are out there.
>>
>>109450619
pls give your robots proper tools to work
>>
>>109450563
Someone posted this yesterday too: https://www.alphaxiv.org/abs/2607.26760
Haven't had a chance to try it out yet
>>
>>109450787
Wrong link:
https://github.com/VictorTaelin/OptMem
>>
File: 1769881976483517.jpg (72 KB, 680x558)
72 KB JPG
>gemma mtp gives boost on 1 GPU + RAM at draft 2
>Buy second to supplement
>now gives the same or slightly slower tk/s than without
>checked the draft split is [1,0] - MTP is offloading
wut
>>
>>109450756
like this, but summarize past cutoff, store entire convo and allow the agent to search it, so it can refer to prior bits if needed, while keeping the context not-bloated
>>
>>109450776
I'm making them code in ST. I'm too far gone at this point.
>>
>>109450563
for gemma I hijack thinking blocks to track stats and summarize history to carry things past the context window
>>
>>109450809
This is computer abuse.
>>
>>109450794
summaryception basically
>>
>>109450809
thats what i was doign a year ago
it was much more fun than using a coding harness
but also much much more frustrating
>>
>>109450809
This would be more usable if ST or any RP frontends knew how to use vscode-style folder tree and had git integration. You are kinda locked into single file projects.
>>
>>109450876
>if ST or any RP frontends knew how to use vscode-style folder tree
Uhhh, I've been using lorebooks for that. It sorta works.
>>
>>109450897
anon pls
>>
>>109450897
based
>>
>>109450611
I use llama.cpp but the package i mentioned installed rocm dependencies solving the issue in my case
>>
>>109450897
Holy shit.
>>
>>109450897
i wont let gemma react to this, for fear of corruption
>>
File: 007~01~01.png (437 KB, 420x586)
437 KB PNG
>>109450897
>>
>>109450897
omegabased
>>
File: IMG-20201104-WA0001.jpg (68 KB, 622x464)
68 KB JPG
>>109450897
>>
>>109450999
>>109450999
>>109450999
>>
>>109450801
Ideally it would have both automatic and manual retrieval, I suppose.
>>
File: mfw.png (514 KB, 480x720)
514 KB PNG
>>109450897
>>
>>109450833
Oh, is that all it is? What a shame.
>>
>>109450522
No



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.