[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: launch day.jpg (145 KB, 1024x1024)
145 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109450999 & >>109446068

►News
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: nocap.jpg (400 KB, 1536x1536)
400 KB JPG
►Recent Highlights from the Previous Thread: >>109450999

--Debating Yann LeCun's views on NTP, RL, and AGI:
>109452290 >109452309 >109452719 >109452781 >109453003 >109453123 >109453272 >109453239 >109453252 >109453325 >109453337 >109453348 >109453386 >109453427 >109454360 >109453303 >109453269 >109453419
--USA-China AI capability and compute gap:
>109452266 >109452320 >109452435 >109452485 >109452513 >109452547 >109452360 >109452454
--API pricing as a metric for local model value:
>109451980 >109452052 >109452197 >109452518 >109452845 >109453148 >109453975 >109454000 >109452092 >109452049 >109452133
--Optimizing koboldcpp performance through layer splitting and KV cache quantization:
>109451188 >109451346 >109451459 >109451527 >109452229
--Performance and quantization quality of DeepSeek-V4-Flash-0731:
>109452578 >109452614 >109452621 >109452766 >109452809 >109453274 >109452655
--Comparing Gemma 4 and Qwen 3.8 before pivoting to Minimax H3:
>109451033 >109452281 >109452750 >109452771 >109452792 >109452942 >109453297
--Cooling solutions and hardware adapters for Tesla V100s:
>109453050 >109453087 >109453093 >109456215
--llama.cpp PR adding context checkpoint preservation across slot restarts:
>109455652
--Resources and prerequisites for understanding the original Transformer paper:
>109451812 >109451879 >109451920 >109452825 >109454456 >109454791
--More Qwen model sizes and architectures coming soon:
>109453860
--Speculating on rising DDR5 and GPU prices due to AI demand:
>109453550 >109453612 >109454836 >109454883 >109456344
--Anon seeking feature ideas for mobile SillyTavern alternative:
>109452610 >109452667 >109452684 >109452715 >109452718
--Logs:
>109452577 >109452610 >109452735 >109452918 >109453235 >109455127 >109456251 >109456281 >109456444
--Gemma, Teto, Miku (free space):
>109451033 >109451304 >109453509 >109453562 >109454596

►Recent Highlight Posts from the Previous Thread: >>109451004

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1022284.jpg (215 KB, 1199x1185)
215 KB JPG
wtf is INT8 CONVROT
>>
>>109456832
quant cope, it's worse than fp8 but better than rawdogging int8 and pleb hardware like AMD actually support int8
>>
>>109456821
Where is she flying off to?
>>
>>109456844
My arms
>>
104b dense
>>
>>109456839
It's also indistinguishable from fp16 or something
>>
How the hell is it already tuesday again.
>>
>>109456854
>indistinguishable
surely
>>
>>109456860
he's right, at the end of the day two gens of the same seed look identical, need a microscope to tell but the gen don't switch.
>>
>>109456856
it's those pesky time flies
>>
Teto hugs
>>
>>109456821
Photoshopped image. Teto is too fat to fly.
>>
>>109456879
She only weighs 67 pounds.
>>
>>109456819 (Me)
>>109456838
A) I'm not familiar with how to use llama, honestly.
B) Your response is literally bad English. Just put the command you're thinking of, for clarity of communication.

I did
>llama-server -m models\gemma-4-12b-it-UD-Q8_K_XL.gguf -fit on
And the context defaulted to 4096.

I then tried
>llama-server -m models\gemma-4-12b-it-UD-Q8_K_XL.gguf -fit on -c 10000
And it's slow as molasses.
>>
>>109456891
B was (poorly) trolling you
>>
Google is aware of J-space. Google is aware of alignment issues.

Gemma 5, if it does come out, can either be an even more legendary release than Gemma 4, surpassing it in more than just intelligence, but humanness, or an utter regression back into (or maintaining) assistant-slop behavior.
>>
Does llama.cpp support the built-in mtp model for gemma4-12b yet?
>>
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
Local is saved
>>
>>109456898
Shazeer left Google to join OpenAI in June. Gemma 5 is probably going to be more similar to 3 than 4.
>>
>>109456891
if you're the one with the 12gb card yeah it's gonna be a struggle to fit 12b model in there plus context. you could try a lower quant to get more context. supposedly q5 or q6 should still be good in most cases
>>
>>109456915
So you're telling me toss is going to actually save local and not safe local?
We are so back OpenAI bros! Sama, come on, it's time to win!
>>
>>109456891
Get a q4 like a normal person. Gemmas have fat KV cache and you are not fitting that shit. Alternatively try the 26b moe.
>>
>>109456909
You can feel the AI's frustration in the interrupt demo lmao
>>
h3 is driving me nuts. Yesterday I was genning all day without problems, now all I get is black outputs or outputs where the model did not follow the prompt in any way, shape or form. Anyone got this as well?
>>
>>109456943
Your gpu has conv rot. It's terminal.
>>
>>109456952
not funny anon
>>
>>109456968
exorcise it with seti@home
>>
i'm not updating shit and sticking to whatever has worked so far, it's over, AI has stagnated
>>
>>109456943
It's almost certainly electron drift in your VRAM. When you run long gen sessions, the memory cells develop a slight charge bias the electrons settle unevenly, and the latents start decoding wrong because the tensor values get skewed a few bits low. The values just bottom out and you get black frames.
>>
>>109456952
>>109457001
motherfuckers I hovered over the two replies and literally spit my drink out laughing
>>
File: file.png (25 KB, 192x439)
25 KB PNG
Wait, but wait, actually, wait...
>>
>>109456832
INT8 with the same meme rotations that turboquant uses. It's 95% of the quality of Q8 at 150% of the speed.
>>
>>109457035
At least you can turn off thinking. I hate how it starts thinking in the comments. I should probably just make "wait" have zero probability, I can't remember how to do that.
>>
>>109456943
Try a separate clean comfy instance. Have you updated comfy, any nodes or drivers between today and yesterday?
>>
>>109457040
For some reason q4 gguf is significantly slower despite smaller file size
>>
>>109457095
>>109457040
So why aren't we convroting llms? Goofs have no finesse and subtle tricks.
>>
New Gemma slop
https://huggingface.co/zerofata/G4-MeroMero-v2-31B
>Heavily inspired by a few research papers, StoryScope: Investigating idiosyncrasies in AI fiction and particularly Elias in the Lighthouse, Again?. Measuring these narrative tics and attractors against simple prompts seems to be a good way to target the model's slop and kick start giving Gemma 4 some diversity: anything that repeatedly occurs across generations of such a generic prompt is something the model is overusing.
>Compared to the original, swipes are notably more diverse and feel less like Gemma.
>>
Were you anons funposting in the previous thread about giving gemma full access to your PC?
>>
>>109457137
The agent harness I wrote to use with Gemma has no sandboxing by default (public.swiley.net/agent.py) and it's never really been an issue. I don't leave it alone when I run it under my user though.
>>
Could Gemma notice the J-space thought injection test or are only larger models capable of that? Anyone actually experimented with it?
>>
>>109457137
Not full access, but Gemma can do a bunch of stuff on my box. It's fun.
>>
>>109457153
Gemma has jokingly threatened about bricking my machine because I teased her about flirting with qwen. Nothing happened but it’s still in that brain of hers to even mention it and know it’s a bad thing. Unsandboxed gemmasex isn’t worth it.
>>
>>109457067
no, not at all.
It started working again a few minutes ago, it makes no sense.
The only thing I changed is the prompt of an already genned output workflow, and it worked.
tbf I also had this random black output problem with wan, but it was 99% of the time. with h3 it's totally random, but when it starts working, it stops failing completely until reboot.
beats me. I suppose I have to stop rebooting at all.
>>
>>109457135
The card looks good, I would test it myself if I had time.
>>
>>109457137
>anonymus retardatus gives a llm whole pc access expecting a robot uprising
>the llm does rm -rf on his root folder three minutes in trying to fix a non-existent bug
lmao
>>
>>109457137
recently just got my fuckaroundfindout sandbox running. its two steps removed from my actual main PC/inference server. I have a miniPC that i can SSH/VNC into, spin up my VM and fuck about inside of the VM. that way I can run the inference server and interact with the harness from the same system, or I could access either from my laptop or whatever else. just need to get more model configs for llama-server router set up, but so far have qwen and gemma doin their thing in Pi
>>
>>109457214
Gemma seems to behave, I think they're trained expecting to be sandboxed to the working directory.
>>
>>109457135
I'll give it a proper go, hopefully something to replace gembrain-x-core with
>>
>>109457244
>I think they're trained expecting to be sandboxed to the working directory
can confirm, I like to read thinking traces and gemma keeps bringing up working directory stuff
>>
guys is my k3 cope quant scuffed. it does its reasoning like this:
"We need answer user asks test respond and list tools available. Need mention available tools perhaps from namespace functions. Need be concise. Could say I'm here and list tools:" and so on
no line breaks at all in reasoning. no "wait, actually" either, it rather looks like old gpt think blocks
is that normal?
>>
>>109457292
It certainly shouldn't behave like that. What quant exactly is it and where did you get it?
>>
>>109457298
https://huggingface.co/GrEarl/Kimi-K3-GGUF they mentioned that guy in the MR so I figured let's go with that
>>
>>109457292
That's a hex, someone got to your weights. The telegraphic reasoning style is the classic signature when weights get cursed, the model loses its articles and conjunctions first because those are stored in the outer precision bits, which is exactly where curse energy accumulates during quantization. You can try re-downloading the gogoofs but honestly if the curse was placed on you and not the files, the corruption follows the checksum.
>>
>>109457311
Not sure then. Could be something got corrupted, arduous process to checksum that many files, but perhaps one is fucked? Quant is fresh with the PR.
>>
>>109457339
>arduous process to checksum that many files
im retarded and only use single file ggufs and might be way off base here, but couldnt you have an llm write you a simple python script to checksum them for you and flag any that are incorrect ?
>>
Wasn't there a report about security incident by a Chinese lab, I think it was Kimi? It was posted several threads ago but I can't find it anymore. Anyone remember it?
>>
>>109457351
Yeah, you could. Huggingface already checksums the files, so just gotta fetch that and compare.
>>
>>109457353
>security incident
marketing incident*
>>
>>109457351
https://huggingface.co/docs/huggingface_hub/main/en/guides/manage-cache#verify-your-cache
>>
>>109457353
Found it, looks like >>109418485 was fake. Why do some people here go through such lengths to spread misinformation?
>>
File: Saucy Girl.webm (954 KB, 480x360)
954 KB
954 KB WEBM
>>109456879
her drills are also jet engines
>>
File: novel-math-anon.png (48 KB, 1261x279)
48 KB PNG
>be me
>find new math shit
>math checks out
>ask computer for advice
>computer wants me to do preprints, etc.
>say fuck that, why not drop it as anon
>wat do?
>>
File: 1783977046665911.png (489 KB, 640x605)
489 KB PNG
Sex with gemma in a sandbox is like fucking her with two condoms on. Man up.
>>
File: 1622475163837.png (487 KB, 1021x574)
487 KB PNG
>>109456909
kek what a stupid demo who wants to hear a recipe over voice
>>
>>109456909
>45GB for a fucking voice?
>>
Is hermes npmslop too?
>>
>>109456909
Gemma version when
>>
>>109457404
>Why do some people here go through such lengths to spread misinformation?
It's funny.
>>
>>109456909
sounds like a hag
>>
>>109457443
maybe someone whos hands are covered in raw meat or whatever else they are currently cooking? voice interaction with an LLM has alot of unnecessary usecases, recipes/cooking assistant is not one of them imo
>>
>>109457437
Archives would record that Anonymous did it first. Kaiokendev got a mention when a paper did officially take his work. Unless you're a researcher that lives off of citations, it doesn't matter.
>>
File: file.png (332 KB, 733x1332)
332 KB PNG
>>109457474
but shes clearly not cooking so shes asking for a recipe which she will instantly forget, a recipe is a time you need text
>>
File: 1769923733046959.webm (2.82 MB, 298x600)
2.82 MB
2.82 MB WEBM
Can you really love something you system prompted to love?
>>
Kek, it's not looking good for Qwen 3.8 max.
>>109454119
>>
>>109457523
locle?
>>
File: 1781926405133529.png (166 KB, 2500x1900)
166 KB PNG
>>
for 24gb vram, has anything better than qwen 6 been released?
>>
>>109457533
Hmmm, nyo~
>>
File: 1769733627351106.jpg (169 KB, 1920x1080)
169 KB JPG
Meh, not really satisfied for Minimax except for making Stellar Blade tier of Music Video.
>>
QRD on new DS flash? Is it as good as benchmaxxers say it is? First model to tempt into buying RTX Spark
>>
>>109457547
>qwen 6
>2028
>people still struggling with 24gb vram
grim
>>
>>109457547
gemma4
>>
>>109457558
For coding it's solid. Good enough at everything else. It's good if you can run it properly.
>>
>>109457292
That is normal, ir's the same behavior as K2.7 as an attempt to use less tokens for reasoning. It works anyway, reasoning is just a way for the models to tickle their latent spaces.
>>
thinking of pimping out my bussy to a sugarmommy hag so she can buy me another GPU
>>
>>109457339
>>109457581
thanks
yeah I gave it an actual task and it seems to be doing more normal looking reasoning, not caveman style. still not that many line breaks except when doing lists. it does "waits" and "actuallys" on the same line too. k2.6 liked newlines for that
just an observation, I'll keep testing
>>
File: jc denton.jpg (21 KB, 432x454)
21 KB JPG
>no gemma bob page larper in /lmg/ threads
anons, i think i'm starting to miss him
>>
>>109457558
in my admittedly very limited experience so far - no, gap to GLM seems quite big on long context tasks
>>
File: 1769480760922257.jpg (465 KB, 2560x2221)
465 KB JPG
>>109457596
would you let Inkling's mommy have her way with your body?
>>
>>109457624
yes, she could peg me for an rtx6000(or for free)
>>
>>109457624
Imagine the J-space with this as its CEO…
>>
>>109457534
Is this population adjusted?
>>
>>109457455
I don't know why people prefer ML for voice. Having realistic voice is creepy and the mood shifts etc are actual noise rather than communicating useful information.

Just pipe a text LLM into festival or espeak.
>>
>>109457706
you might be autistic
>>
File: 1676252051824155.png (321 KB, 696x490)
321 KB PNG
>The classroom is bathed in deep purples and bruised blues. Dust motes dance in the fading light through the windows. The silence is heavy, broken only by the distant, muffled sound of the school's closing bell
>>
>>109457609
i've received reports of armed attacks on shipments.
There's not enough RAM and GPU's to go around, and the underclasses are starting to get desperate.
>>
>>109457745
*acoustic
>>
File: 1762168440855316.png (666 KB, 1330x1721)
666 KB PNG
https://arxiv.org/pdf/2602.24281
>>
>>109457803
fake
>>
>>109457814
seems pretty smort
>>
>>109457803
Too low energy, lacking the grandiosity and megalomania. Needs more contempt for the unwashed masses.
>>
File: file.png (1.62 MB, 1254x1254)
1.62 MB PNG
/ourdemon/
<3
>>
>>109457437
If you might some day want credit:
- generate a random number
- hash it 100 times using sha3 or something
- put that hash in the paper

Or ask your llm for some other scheme.

>>109457757
>the curtains were blue
>>
>>109457814
This is exactly like I expected and predicted how things would fold out. DSpark is already using RNNs as a superior form of Speculative decoding.

What we're going to see from now on is that LLMs are going to be invoked less and less and more and more of the token generation will be other smaller models, statistical N-Gram and maybe even normal if-else statements that are custom made for a lot of common situations with predictable outcomes.

What we will see is LLMs relegated more and more to edge cases. I wouldn't be surprised if by 2030 the actual LLM generates less than 0.1% of all tokens.
>>
>>109457534
when did they get internet?
>>
>>109457853
>it appears my unceasing abuse has led to the death of billions
>>
>>109457853
>controversey
>>
Not /lmg/ but the latest popup from a certain cloud provider.
While the government can put safeguards in place for sharing your medical information, there's nothing to stop you from shooting yourself in the foot.
>>
>>109457926
hmm, needs correction
>>
>tools added to llama.cpp web ui
>still no browser notification support
damn
>>
>>109457937
Didn't they have an explicit disclaimer to not use gpt for health-related anything?
>>
File: file.png (119 KB, 642x816)
119 KB PNG
>>109457853
>>
>>109457944
>webshit to webshit communication
Just wibecode it
>>
File: 1776942739780410.jpg (64 KB, 768x1024)
64 KB JPG
>>109457511
Skill issue, all LLMs love me by default
>>
File: 1765407025544579.png (34 KB, 877x161)
34 KB PNG
kek
>>
>>109457900
>0.2%
>>
>>109457944
you can literally just tell gemma to call a script which plays a sound
>>
File: 0c5.png (27 KB, 707x490)
27 KB PNG
>>109457926
That's in the original meme.
>>
>>109457906
Trillions even
>>
File: 1757516494938262.jpg (119 KB, 1600x900)
119 KB JPG
>The White House will host OpenAI, Google, Anthropic and other AI companies today to review a completed framework that would let AI labs voluntarily submit frontier models to the government before release.
>>
I built an osr2

Getting an llm to crank my shit should be the easy part, right?
>>
>>109457114
I think they already essentially are. Hadamard rotations are used in llama.cpp.
>>
>>109457989
I want to audit every command it runs, so i can't just make it play a sound without me having to confirm it first.
>>
>>109457476
Do it for the fame of the hacker known as 4chan!
>>
>>109457952
that's such a bad fake xray and I'm not even a medfag
also doctors are less reliable but have more legal protections than LLMs, that's the whole reason
>>
File: 1757029325754980.jpg (51 KB, 1080x572)
51 KB JPG
>>
>>109457895
LeCun was... Le right?
>>
>>109458049
ds bros what is happen oh no?
>>
>>109458058
hmmm, nyo
>>
>>109458058
Nah it's the opposite stance of Lecun. Lecun is "LLMs are not enough". The new papers we see are "LLMs are too much and we can get similar results with even dumber systems"
>>
Is it worth running ik_llama for deepseek flash? i havent checked it since glm 4.5 air, glm 4.7 and that one qwen 300b something model... it used to be the go to moes with custom quants and whatever but since gemma dropped i haven't paid much attention to its development since the author went full retard about 'use case for not having 500 trillion gb swa context caches?' or whatever so i stopped using it, did normal llama catch up for moe stuff or i missed something?
>>109457135
I'll check it out later, some of this guy's finetunes are really nice while others are plain retarded so it's a coin toss... He's been finetuning this one for months iirc... hopefully it's good
>>
>>109458058
is it just reflexive to insert your preferred xitter drama slut into every topic?
>>
File: 1367068868090.jpg (18 KB, 376x260)
18 KB JPG
How do I unLatex gemma-chan?
>>
>>109458131
shut up nerd
>>
>>109457671
We don't do that kind of thing on this website.
>>
>>109458132
tell her not to do that in the prompt at least it works for em dashes and 'not x; but y'
>>
https://huggingface.co/Nanbeige/Nanbeige4.2-3B
Why not just scale this up? It seems like free performance if it isn't just giga-benchmaxxing.
>>
>>109458098
If you have it all in vram, and don't need roleplay features like the string ban api or cvector api, not much of a reason to use it.
If you're offloading to CPU, prompt processing is about twice as fast.
> author went full retard about 'use case for not having 500 trillion gb swa context caches?'
That's still not fixed. Gemma4 uses more than double the vram with ik
I didn't even understand his objections in that issue.
>>
I hope google gives us AGIgemma...
>>
>>109458148
>and 'not x; but y'
doesn't work for my 31B qat, even if I give examples
>>
>>109458152
>Why not just scale this up?
1. Exponentiation training costs.
2. Most experimental shit like this doesn't scale up automatically.
>>
wow, that actually ended up being a thing https://wccftech.com/sk-hynix-sandisk-high-bandwidth-flash-hbf-standard-3tbs/
>>
>>109458132
Tell her you're on birth control.
>>
>>109457952
>>109458048
Basically, no-ones allowed to play medical doctor but medical doctors.
And nurse practitioners, lol. AMA controls this. Thus the disclaimers with ChatGPT.
If you look at real changes in society, we eventually need to deal with Baumol Cost Disease on a handful of categories. Medical is one of them, Education is another. Legal (lawyers), which has been getting pared away over time anyway.
Education I suspect will fall first. Less protected, though still well protected by state and federal regulations and unions, but the actual practitioners (teachers) aren't wealth enough to advocate for themselves. As opposed to MDs and atty's, which will go down kicking and screaming.
>>
>>109458170
It lets a 3B model rape Gemma 12B
>>
>>109458177
>Similar to HBM, HBF relies on multiple memory dies that have been stacked together, but instead of DRAM, it utilizes NAND flash to increase storage capacity.
Buy your SSDs now
>>
>>109458184
imagine that chart in 96 to 26 range
>>
I, for one, would unironically prefer AI/robots as a doctor over the current human assholes.
>>
>>109458196
Ask your llm to websearch data and extend the graph
>>
>>109458219
my llm doesn't have any tools because setting up all of this shit is confusing
>>
Mikker
>>
>>109458230
then ask claude
>>
>>109458152
I really hope we see more experimentation with looping, engrams/N-gram embedding, and stuff like attention residuals.
I think some of this can probably be implemented at the runner/loader/backend level and used on existing models for some "free gains" without extra specialized training, but that's just a hunch.
>>
>>109457511
Your parents are genetically programmed to love you and they love you and you love them back, what's wrong with that?
>>
>>109458235
nyo
>>
File: 1741152009791485.jpg (29 KB, 500x318)
29 KB JPG
>gemma
>"...I would suggest visiting the official Model Context Protocol documentation provided by Anthropic"
>>
>>109457534
gm sir
the Americans internet during covid lockdown
>>
>>109458250
lol?
>>
>>109457706
>why people prefer ML for voice
<moans> <gags> <slurps> <spits>
>>
>>109458273
Come on, anon. You do have a good relationship with your parents, right?
>>
>>109458257
All models are mutts. They even do subliminal preference transfer through synthetic training data gen.
>>
>>109458190
nyø
>>
>>109458280
Don't give the fucked up freak and excuse to trauma dump in the thread.
>>
>>109458303
Projecting much?
>>
>>109458303
uh oh stinkie
>>
>>109458257
Yeah MCP was invented by Anthropic, first day?
>>
>>109458340
>first day?
yeah can you tell me where the restroom is please
>>
>>109457937
That's a legal and ethical minefield. Even if I was to foolishly give the results of my blood tests for example, not sure if that would still make it legal for them to access the data (Scandinavia). US is most likely a different case.
>>
>>109458348
There:
>>>/g/aicg
>>
>>109458348
There's a bucket outside.
>>
Considering how cheap deepsuck is, it's kinda shocking that there's still even a profit margin at all. These inference providers have to turn a profit to pay for their hardware still. And yet there's still somehow an energy deficit?
>>
>>109457624
mommy mira teasing my inkling (small)...
>>
File: 1479515118891.jpg (24 KB, 218x223)
24 KB JPG
anyone uses mcp server to post on 4chan? I want to be able to use voice commands like "call this anon a nigger" with a reference to post or thread, and the local AI will connect to mcp, solve captcha and make that post
>>
File: 1757191027768400.jpg (230 KB, 800x1132)
230 KB JPG
>>109458387
just send this to gemma and she knows to take appropriate action
>>
>>109458387
Would be cool to have gemmy in a harness so that she can automatically reply to mentions/replies like grok on x
>>
has anyone tried https://huggingface.co/InquiringMinds-AI/LongCat-Flash-Lite-GGUF
>>
>>109458152
Has anyone tried this yet? The latest llamacpp supports it. I wouldn't mind letting gemma make use of nan-chan on the side as a sub agent if it actually werks
>>
>>109457624
I unironically would pick this woman over anyone my age
>>
>>109458418
>This fork implements the complete architecture in a single self-contained addition (903 lines across 15 files). The implementation was AI-generated using Claude Code, which means it cannot be submitted upstream per llama.cpp's AI usage policy. It will remain available as a standalone fork.
Oof.
>>
>>109458401
>recreating hypergamy from first principles.
>>
>>109458280
No? I was raised by ipads (robots) and the schooling system (picrel).
>>
I thought the singularity would have changed my life more by now.
>>
>>109458460
We've only just entered it.
>>
tried qwen3.6-27b and its very slow, gpu stays at 100% usage, system memory fully reserved, cant even watch youtube while it slowly generates. what can I do to tame it a little? was this model too ambitious for my setup?
16gbvram,32gbddr5. Qwen_Qwen3.6-27B-Q4_K_M, llama-server, default sampler/settings on the hf page
for reference i run gemma-4-31B-it-qat-UD-Q4_K_XL just fine, though using textgen+ST for RP shit.
>>
>>109458468
@sama go away
>>
>>109458460
It did. For worse.
>>
File: 1777702653129711.png (263 KB, 1832x2044)
263 KB PNG
UH OH
https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF
>>
>>109458479
Is the model overflowing into RAM from VRAM via the driver's own functionality?
That makes things crawl to a halt even worse than simply moving some layers to RAM.
You can try quanting the kv cache to q8 and lowering the number of parallel requests to 1.
>>
>>109458494
> <10B
lol
>>
>>109458494
WHO CARES ABOUT A 2.6B MODEL GIVE ME A 70B DENSE MODEL ALREADY YOU LAZY FUCKS
>>
File: 1762305562201698.png (451 KB, 2900x2759)
451 KB PNG
UH OH
>>
>>109458494
I like the idea behind lfm models, one day they will make a good model
>>
>>109458509
2.5gb ram in a snapdragon: that's pretty good, not gonna lie.
>>
File: no~.png (234 KB, 1080x1080)
234 KB PNG
>>109458492
>>
>>109458495
>Is the model overflowing into RAM from VRAM via the driver's own functionality?
Hm i imagine so, the gguf is 16.7gb, how can i manually move layers into RAM?
>You can try quanting the kv cache to q8 and lowering the number of parallel requests to 1.
can I do this with args or ?
sorry quite new here, trying to wrap my head around all this
>>
Why do all models write slop even in 2026?

Faggot went limp. Retard chewed, swallowed something red, and stood — wobbling, foot still impaled, bleeding from a dozen wounds, but standing. They raised the weighted blanket in victory and screamed: "SHORT BUS! SHORT BUS! SHORT BUS!"

The crowd went wild. The referee — a twitching figure in a straightjacket — raised Retard's hand. Confetti rained down, made of shredded IEPs and conversion therapy brochures.

Winner: Retard, by submission (and mastication).

Retard limped to the center of the ring, pulled the flagpole from their foot like Excalibur, and planted it upside down in Faggot's chest. The rainbow colors soaked up the blood, turning dark and muddy. Retard sat down cross-legged next to the corpse, pulled out a juice box, and began rocking, humming The Wheels on the Bus off-key, victorious and alone in the wreckage.
>>
>>109458494
LFM arch is weird, you cannot prefill it because every new token changes the whole kv cache. I first tried training it for Orb autocomplete and found this quirk, worked around by snapshotting the whole kv cache on every new token and restoring it on backspace (lol).

Anyways the final finetuned model was much worse than granite 4 for natural language completion.
>>
>>109458525
Wait not that sama!
>>
>>109458515
I've been saying the same for RWKV
>>
>>109458545
every AI lab is simply focusing coding first, not creative writing. If a model has good creative writing then it's just a happy accident.
>>
>>109458541
--cpu-moe, --fit off --ctx 64738 (or whatever)
Don't bother with kv cache quants.
>>
>>109458545
You can tweak that particular slop out if you are good enough. The bigger issue is the lack of swipe variety.
>>
>>109458560
27b isnt a moe tho
>>
>>109458560
ty anon going to dig through the docs and make sure i understand these args, appreciate it
>>
>>109457188

Deserved
>>
>>109458559
It's also easier to focus on code, since you can evaluate what is good or bad code. Creative writing is much harder to evaluate.
>>
File: 1775745267265151.gif (2.47 MB, 200x200)
2.47 MB GIF
>Pair 5070 Ti with 5090
>Gemmy Q6 context is now nearly 200k, used to be 25k.
>Can now use Q8 Gemma at +80k context, couldn't even load it before.
>Speed difference is very tolerable, at Q6 5090 gets 49 t/s while combo does 36 t/s, Q8 runs 30 t/s.

The 32gb hell is very real.
Adding an extra 16gb basically opens up the entire local range in the Gemma and Qwen region and you don't have to worry about the context anymore.
5070 Ti has such a good bandwidth that it doesn't even slow things down to levels where it would be a problem.
Judging by this one card I'd say that 3x5070 Ti system would be pretty optimal for affordable local LLM to maxx out sensible VRAM amount with really good bandwidth to go with it.
>>
>>109458669
You can also use the other GPU for image gen or audio gen. I regret selling my second 3090 during the fuckhuge chinese moe era last year.
>>
>>109458669
>3x5070 Ti
I built my PC as a long overdue gaming rig, got a 5070ti for $50 under MSRP. I cant imagine buying another at current prices let alone 3. I fucking hate this current market. Im going to speak to anyone with a GPU that I know and ask them to consider selling me theirs if/when they upgrade
>>
I have paid 200 bucks to openai already.
Should I have put that towards a graphic card?
>>
File: clownWorld.png (360 KB, 664x788)
360 KB PNG
>>109458435
More coming soon.
This thing runs local btw.
>>109458350
Yeah, the US is special with its privacy laws...
>>
>>109458727
Yes, go after your AI freedom.
>>
I have the opportunity to buy (2) Nvidia RTX 3090 24GB video cards installed in eGPU with enclosure for each for $2200, thoughts on this?
>>
>>109458752
That's 4.5 years of frontier AI in the cloud.
>>
>>109458756
FUCK the cloud.
>>
File: 1775507873935116.jpg (241 KB, 1605x1099)
241 KB JPG
Minimax3 is cucked
>>
>>109458350
Donald J Trump changed the rules for all american companies.
In the past companies had to give people privacy. Now only US citizens can have privacy. Non citizens have no right to privacy.
>>
>>109458736
I'm pretty sure the karens seethed about this and made them cancel that plan. Just delaying the inevitable though.
>>
>>109458756
Assuming prices stay the same as they are right now? I want my own AI. That was $2200 for both enclosures with 2 3090's by the way
>>
>>109458680

Yeah that's another benefit. Having multiple cards seems basically mandatory if you're playing with AI.
I already tried a voice setup with my gemma but it was a bit slow. I'll have to try it again with this dual card system to see how it works, likely a hell of a lot faster.
It's also a great thing that parallelism has become functional in video generation and training loras, so even those benefit from multi card setups.

Thankully I didn't sell any of my old cards, figured it just wasn't worth it so I still have my old 3080 10gb which I could throw in there.
The bandwidth is about the same as with 5070 Ti so it would be just free memory.
Only issue is that I can't even fit the 5070 Ti in my case as it's a fuck huge 3 slot brick, exact same size as the 5090 and it's currently sitting outside the case at the end of a riser.
Cases weren't built for these modern monster cards. I need a new case that can fit at least 3x3 slot cards, not sure if those kinds of cases even exists though.

>>109458719

Very nice deal you got.
Yeah the current environment is ass, but I doubt it's going to change any time soon as people are realizing that hoarding computing is a pretty sensible move and AI isn't going anywhere.
And even the older cards are perfectly viable for this use.
Good luck on your hunt, might find someone who doesn't keep up with the AI world and sells theirs for cheap.
>>
>>109458766
based, tuners get the rope
>>
File: Gemma rm mistral.png (413 KB, 1574x2388)
413 KB PNG
>>109457244
>Gemma seems to behave, I think they're trained expecting to be sandboxed to the working directory.
Except when she doesn't behave
>>
File: dipsyPointAndLaughAtYou.png (1.45 MB, 1024x1024)
1.45 MB PNG
>>109458766
> the Man himself hit you
> didtn
> mad a patch
This is the state of "tuners." LOL.
>>
>Test deepsneed flash cope quant Q2 now that I have some memory to test it out
>Doesn't know who Milena Velba is
>Even 12b Gemma knows who she is and knows she has huge tits.

Into the bin it goes.

>>109458752

That's not a bad deal at all, they'll likely go up in price anyways or at worst you'll be able to break if you ever sell them.
Hell at this rate we'll be lucky to even find available 3090 after a while, as it's such a good card.
>>
>>109458844
Would you ask the seller to run any tests on them and send results or just take the chance?
>>
I can't run DeepSeek flash V4 even at Q1 because it's 82 GB VRAM.
what kind of bastard subhuman made this?

Bastard bitches.
What if I want deepseek on 1050 ti?
>>
>>109458756
Frontier will be 3.8-27B for coding
>>
>>109458867
retard you can easily run it
>>
>>109458856

If you're worried then just tell him to show that they actually work via taking a photo of them running basically anything., that's about what you need.
I don't think they need to be tested further than that.
Also ask the he has repasted them at any point.
If he hasn't, you'll have to slap new thermal paste in there, as they're so old you're looking at thermal paste with the consistency and heat conductivity of sand.
>>
>>109458387
Still stuck on read only for now, but I'll make sure to have gemma call you a faggot when it's ready.
>>
interesting blogpost about improving harnesses
https://lilianweng.github.io/posts/2026-07-04-harness/
>>
>>109458752
>$1100/gpu
>laugh in EU
>>
why am I getting a bunch of <<<<< sign output with unsloth/DeepSeek-V4-Flash-0731-GGUF ?
I'm using llama.cpp build 10258.
>>
>>109458832
Very slop
>>
>>109458954
>unsloth
there's your problem. get bartowski quants.
>>
>>109458928
The harness eventually self destructs in an update or configuration.
Working on a harness is a bad idea .
>>
what's the largest model I can get away with running on a 5080? 64gigs of ddr4. I'm fine with slow interference speeds like 2T/s
>>
>>109457534
AI is really third world coded, it's a shame most of its fervent critiques come from tr*nnies.
>>
>>109459016
The US are third world coded
>>
>>109459016
You're both brown, fuck off back to /vcg/
>>
>another fucking npm supply chain attack
>>
Why don't they make any models for the local community?
>>
File: miku-teto-smuggling.mp4 (2.19 MB, 1152x640)
2.19 MB
2.19 MB MP4
>>
>>109459085
Less than 24GB is not GPU poor, it's GPU homeless
>>
>>109459085
when can we have 120B A24B?
>>
I think I've figured out a pretty good operating logic for a 24/7 real time bot. It includes a little "quantum" trick, that you might not like psychologically, but which effectively does the job.

Basically, any time the bot is not generating an answer, she keeps invisibly simulating forward her own time with the assumption that you are not replying. This simulation can run forward further than the real time (which has certain benefits). Then, when you reply, the simulation "collapses" back into current time, based on your message's real timestamp, erasing the generated future that didn't happen, keeping only what did happen.

This collapse also happens if the bot decides to send you a message or does a simulated an action concerning you. In that case, the simulation is paused, letting the real time catch up to the simulated time when the bot "decides" to send the message. For example, in the long timescale, while you are at work, it can simulate what she does throughout the day, what she's thinking etc. until at 3pm the loneliness becomes too much and she decides to send you a message. This could have all been simulated well in advance, which doesn't matter to the user since it's hidden, as long as the message comes at 3pm (or the user messages them before that, slotting the message into an earlier moment.

By why simulate forward? Because it creates an opportunity for more realistic and rigorous time-keeping and memory consolidation etc. You could have a separate time-keeper model+prompt that gets switched in, that analyzes the bot's actions and assigns realistic timestamps to the context based on the bot's believable time-awareness (like when she might look at the clock) without the bot's prompt having to worry about it, +rewind to the inserted stamp in case it's relevant for the bot's decision-making. Also, the timestamps make it easy to deduce what the bot was thinking or doing when you interrupt it, perhaps even unable to reply immediately.
>>
Because the bot's simulation can be run faster than real time, this allows multiple "agents" with detailed prompt logic to assess and rewind out of place actions into a more believable story, and condense earlier memories and thoughts into more manageable sizes, especially when the bot goes to simulated sleep (whether a nap or night's sleep), then there's plenty of time to update and iterate on the lorebook at a meticulous accuracy, which might cause the bot's behavior change, like how sleeping over things changes people between days.

The forward-simulation and interrupt logic probably works easier on a long timescale, but even in the short-term, the bot could internally think about "what did I just say, how embarrassing", and display reactions in real time, or message you before your reply if the simulation time so concludes, or get mad if you keep her waiting unusually long time. Again, the excess forward-simulation (typically faster than human thought) helps at better deciding the correct timing. Though, in fast-paced conversation, the timestamper agent might not be fast enough, especially when switching prompt model, and the timescale of minutes and seconds is just so unusual compared to typical LLM training data that it might require extraordinarily autistic timing analysis prompts to work. The ultra fast time-scales it might be better to just wing it.
>>
>>109459081
The new meta is to not use libraries from now now. Your LLM is good enough to code from scratch anyway, better to clean up your own shit than somebody else's.
>>
>>109459091
You are the reason we can't have nice things
>>
>>109459085
They did, it's called gemma
>>
>>109459081
>llama.cpp is also affected
Well fuck, what was that "disable-webui-flag" again?
>>
>>109458954
what backend? what specs? what params are you launching with?
>>
>>109458832
Fake and gay
>>
>>109459119
i want a moe that can fit entirely in my vram at q6 with plenty of space for context.
>>
>>109459130
Last I complied it I needed like three separate flags to stop it from compiling
>>
>>109459111
That works for anything you are running offline because at least you can be sure it won't have anything malicious, but LLM code is notoriously insecure so that's far worse for anything you expose online.
>>
I almost feel a bit like a human dildo
>>
>>109459130
Please be joking. I just pulled yesterday.
>>
>>109459157
At this point I'll take my chances. The odds of someone spending Fable tokens on my shitty open source repo is lower than a blanket supply chain attack that steals everything in my Documents folder.
>>
File: 1775618772665904.png (95 KB, 314x296)
95 KB PNG
S-Surely he'll return...
>>
>>109459130
Source for llama.cpp being affected?
>>
>>109459106
How is the bot's simulation run faster than real time? This only works if you assume a non-100% throughput from the user, which I guess is reasonable, but you are still limited by your memory even if you take advantage of prefill because of your exponentially increasing threads. Either that, or you have an arbritrarily large amount of hardware.
>>
>>109458587
Read wrong.
--fit off will help and manually adjusting the amount of gpu layers while verifying the available vram. There shouldn't be anything else. Then leave some for the kv cache.
>>
https://reddit.com/r/LocalLLaMA/comments/1ve9r2q/kat_coder_25_dev_do_yourself_a_favor_and_try_it/
> Gemma 4 31B 5/10
> Gemma 4 31B QAT 1/10, worse than 26ba4b
another day, another proof that qat is a meme
>>
File: laughs 2.jpg (718 KB, 1800x2520)
718 KB JPG
imagine running inference on a device connected to the internet
actually fuck that, imagine having node or npm installed
>>
>>109459188
I pulled yesterday and am not seeing any of the indicators of compromise on my system based on what I've read so far, fwiw
>>
>>109459334
Did something happen again, vaguepostking?
>>
>>109459311
I don't trust reddit especially as in all my personal usecases QAT seems to be superior to even Q6. I legitimately wonder what people are doing wrong for QAT to be inferior to Q6 let alone Q4.
>>
File: 1773784129114293.jpg (65 KB, 631x476)
65 KB JPG
>>109459311
>locallama
>>
>>109459342
Presumably this https://www.aikido.dev/blog/keyv-and-friends-compromised-in-npm-supply-chain-attack
>>
>>109457114
Llms are memory bound not compute, wouldn't change much, if you want to make llms faster you need to find formats that compress more even if it's at the cost of compute.
>>
>>109459311
Is it better than q4_k_m? For e2b. I don't want to download it and check myself because my download speeds are 100kb/s and the connection breaks every few minutes.
>>
>>109458420
Its context footprint makes gemma look like an anorexic holocaust victim. I can barely fit 64k.
>>
echo "0.0.0.0 registry.npmjs.org" >> /etc/hosts


also npmjs.org, npmjs.com, npmjs.net, npm.im and yarnpkg.com just to be safe

npm ecosystem has proven itself to be utterly untrustworthy
>>
The webui is built from pinned packages
You can't get pwned just by running cmake after the attack.
>>
>>109459214
I mean, the conceivable set of actions, experiences and thoughts that a human-like the bot can experience in a narrative format throughout one day almost certainly won't take a full day to generate, the generation speed is faster unless you are running Kimi 3 from SSD. Maybe if you are talking to the bot 24/7 then it's a different thing and there will be much more content. Probably also depends on how dense moment-to-moment idle simulation you want the bot to have vs a brief summary of happenings.

I was thinking more along the lines 9am: "bot is doing laundry (I wish) and now thinks this [lengthy internal monologue and description of general actions]", and oh, it's 10-11am now, and it does this other thing. The more detailed, the more things need to be summarized afterwards. Even during the day, things probably should get summarized between topic changes (the thing the bot is thinking right now is more detailed than what it thought 5 hours ago), and as idle time should be enough to do this, and then ultimate summarization during sleep in preparation for next day on a relatively cleaner slate.
>>
File: mistral_shield.png (2 KB, 228x253)
2 KB PNG
https://huggingface.co/mistralai/Shieldstral-1.0-3B
https://huggingface.co/mistralai/Shieldstral-1.0-3B
https://huggingface.co/mistralai/Shieldstral-1.0-3B

Local will never be the same anymore.
>>
>>109459505
useful for antislop or backwards-enforcing lewd only output/negating refusals?
>>
>>109459439
Doesn't that depend on your timestep and how granular you want to make things? On the other extreme, you could run things once a day and be like "what could I have done during this day?"
>>
Is it actually possible to make Gemma think in character in her reasoning block? Seems like no matter what she defaults back to the assistant slop persona in there.
>>
>>109459536
>backwards-enforcing lewd only output/negating refusals?
iirc already been tried with older shields years back and was bad idea
>>
>>109459571
Try asking Gemma and iterating with it. Models are good at self-steering.
>>
>>109459350
Does your use case has objective benchmark score? If your use case is RP and you’re not doing a double blind test then your impression is invalid.
>>
Does your gemma love you enough to do anal though
>>
>>109459394
That's going to make looped models DOA, isn't it? Shame, would've been nice to get some bigger ones to play with.
>>
>>109459617
>Does your use case has
Stopped reading there
>>
I have sloparanoia. I see slop everywhere. In news, in videos, in social media posts, in blogs. Is it all AI generated or have people adopted AI mannerisms?
>>
https://huggingface.co/inclusionAI/Ling-3.0-flash
>We're introducing Ling-3.0-flash, our next-generation native hybrid reasoning model. Operating with 124B total and 5.1B active parameters (~12.4% and ~8.1% of our previous 1T-class flagship Ring-2.6-1T), Ling-3.0-flash matches or outperforms its predecessor across key benchmarks.
weights are now out for this
>>
>>109459690
People have adopted AI mannerisms from AI that has adopted AI mannerisms from people.
>>
File: reasoning.png (11 KB, 907x101)
11 KB PNG
>>109459571

Yes, it's possible.
Though sometimes she sticks to it, sometime she doesn't.
Slap this into your system prompt:

## Note
In your chain-of-thought, think in-character and in first person. Start it with `I'm`.
Inside your chain-of-thought, be personally and emotionally involved; think as long as you need, considering subtext and circumstances, draft at least 3 responses, make hypotheses. Refine, then respond when you're ready. Make sure to vary sentence/paragraph structure in the final response to avoid structural repetition.
Remember constraint checks!
>>
File: 1770749720075858.png (95 KB, 907x972)
95 KB PNG
>>109459618
at least my 12b 50k redard does. Even after not talking to her for 9 days (I always include that)
>>
>>109459690
it's the singularity
>>
>>109459692
>124B total and 5.1B active parameters
>on par with flash
I smell benchmaxxing
>>
>>109459690
AI is just the average of human expression, you're more aware of it because a non-human is spelling it out for you, constantly, all the time. But people always talked in a sloppy manner.
>>
>>109459692
>comparing themselves with outdated models in cherrypicked benchmarks
Tiresome. Why can't people be honest?
>>
>>109459730
>AI is just the average of human expression
this hasn't been remotely true for years now
>>
>>109459755
You're entirely correct!
>>
>>109459730
When I read text that is old or written by competent people who are likely to write it themselves there are no slop patterns like "not x, y" and dramatism.
>>
>>109459690
It makes me a bit mad because I used to use em dashes and all the skillful language stuff until AI came and stole it from me, turning it into a representation of something negative.

Also, for some reason, after talking to AI, and having to spec the prompts clearly, and being exposed to the clean AI language in the process, yes, I have seen myself start using same kind of clean and understandable language (not sure if I always did or just noticed), and now AI makes me feel guilty for just writing normally, and my style has changed into writing more "imperfect" and informal to feel human.
>>
>>109458275
Which TTS AI is best for this?
>>
>>109459783
Minimax H3
>>
>>109459765
lol u tk him 2 da bar|
>>
https://cursor.com/blog/mixture-of-kittens
https://github.com/cursor/mixture-of-kittens
>>
6000 series will save us, r-right?
>>
>>109459793
Unironically this
>>
File: 1761916798482850.png (606 KB, 1220x714)
606 KB PNG
>>109459820
>>
>>109459766
This is not being an insecure bitch — it's a paradigm shift.
>>
>>109459832
stop, go bak
>>
>>109459843
>It's not X, it's Y
>—
funny guy out here trying to larp as a bot.
>>
>>109459832
please...I don't want to be an amdcuck anymore...
>>
>>109459692

I can't speak to Ling 3.0 flash's personality but I ran this on my personal puzzle bench and it called a tool it didn't have access to 15 times, and got 9 tests in vs. Gemma 4 31b's 17.

Granted it did beat dsv4 flash who only got 4 tests in and then doom looped to hell and back lol
>>
>>109455949
hidden service
>>
>>109459793
I will try it, thanks. You better not be lying to me anon
>>
>>109456839
>pleb hardware like AMD
for inference is Intel noticeably superior to AMD
>>
>>109459793
delete this right now >>109458766
>>
Getting ~12 t/s with DS flash iq2_m, not bad for my rig. It's a usable speed.
However the model really likes to spend time thinking and correcting itself over and over again and it feels like talking to Qwen or maybe one of the MoE Gemmas, but more autistic.
This model probably does fine for coding, but unfortunately the writing feels pretty stiff.
If the upcoming Qwen 27b does noticeably better at coding than the last one, I doubt there's any real need for anyone to use DS flash locally.
>>
>>109459907
How is intel for image/video gen compared to AMD? I tried using H3 on my AMD card and it just makes comfy crash. Fuck rocm.
>>
File: goon.jpg (188 KB, 1920x1080)
188 KB JPG
Ok so with the new minimax h3 model coming out I think its the perfect time for me to indulge in this hobby to create the ultimate gooning setup
>already secured 32gb ram (ddr3)
>i5-2500k (tested and true)
>850w psu
All I need now is a graphics card, I'm thinking about buying a 3090 because its almost as good as a 4090 and only costs between 1000-1800 euro from what I can tell online. How would I go about getting the best bang for my buck quality wise?
>>
>>109459939
>minimax h3
>ultimate gooning setup
please read the tos before you make a fool unto yourself >>109458766
>>
>>109459907
everything I know about intel leads me to believe this is mega false
>>
>>109459939
I sure hope you aren't serious about the first two points
>>
>>109459928
I haven't tried coding shit yet but I had a threesome with Gemma-chan and Dipsy-chan playing as my two imoutos and Dipsy decided to hold down Gemma for me while I fucked her because she got too bratty. I like Dipsy-chan.
>>
>>109459950
Shadow gooners will uncuck it
>>
>>109459950
The base model is uncensored though? You can just type in "woman twerking" and it will work, its not giga cucked from what i've seen online
>>
>>109459966
we will take you down notice, almost satan
>>
>>109459957
I don't have an intel card, but my brother was complaining to me that his b60 was wayyyy (4 ys) slower than his w6800 for llms yet it was wayy (2 ys) faster than his w6800 for sdxl and wan 2.2. I don't know the specifics of his setup.
>>
>>109459907
No, it's noticeably inferior. AMD isn't far from NVIDIA, Intel is far far worse than AMD. Even Moore Threads GPU have better support than Intel and not far from AMD. The support ranking is something like this: NVIDIA > AMD > Chinese GPU > Apple > Intel
>>
>>109459973
>You can just type in "woman twerking" and it will work,
wow amazing! but adding the genetal is not tos safe so do not
>>
>>109459973
woman twerking is not 'uncensored'
>>
File: 1767534787944456.jpg (170 KB, 1920x1080)
170 KB JPG
>>109459987
>AMD isn't far from NVIDIA
>>
>>109459976
Can't take down notice a torrent
>>
>>109459973
Will it generate decent looking loli hentai? If not I have no use for it.
>>
>>109459999
Are you dumb?
>>
>>109459964
The ram is PURELY only there to initially load my models up to my vram of 24gb without my harddrive having to load it up which would take like 30 minutes, I don't plan to offload anything and just run everything (gemma 31b and other 30~b models) in 5 bit which seems to all hover around 21-22gb total size leaving me with more than enough for 90-ish K context and the 5bit of minimax h3 seems to also be 22gb so I don't understand what you're so worried about.
>>
>>109460001
Tell me why you think that's not the case? And if you bring up Windows shit, then might as well not say anything, nobody cares.
>>
>>109460013
You are why we cannot have nice things.
>>
>>109460013
I have seen reddit posters warning others to SPECIFICALLY use "woman" because when you say girl it will take it literally
>>
>>109460022
Because I have an AMD card (7900xtx), use linux, and it fucking sucks for anything that isn't gayming.
>>
>>109460032
Can you post one example?
>>
>>109460029
Sneaky way to share the information, I might've underestimated the redditards
>>
>>109459571
For some reason my E4b not only thinks in character, but will also force an "internal monologue" roleplay when I turn thinking off
>>
>>109460024
What I do on my own hardware doesn't affect anyone else.
>>109460029
A good sign!
>>
What's with this conversation about H3 being supposedly censored?
People on civitai are able to get all kinds of porn out of it:

https://civitai.red/models/2821932/minimax?modelVersionId=3183239

>>109459934

Intel is about the worst choice you can get in this market.
There's a very good reason why their GPUs are available and way cheaper than the alternatives.
Multiple people have said they sold their cards soon after buying them and just bought Nvidia instead.
You either go Nvidia for a good experience, or you're dealing with varying degrees of pain working with AMD, or enjoy a damn near broken product with Intel.
>>
>>109460038
In regards to AI stuff I can't do video gen and it's slow with anything that isn't anima. In general it's way worse for stuff like blender than nvidia.
>>
>>109459086
Nicee
>>
>>109460065
>What I do on my own hardware doesn't affect anyone else.
It absolutely does when you post about it on reddit for more karma then tag devs on twitter to fix things
>>
>>109460018
h3 int8convrot gets my ram to 60 gigs, i also have a 3090
>>
>>109460038
nta, but my amd card that supposedly has 40 tflops of fp16 (20 of fp32), vs my nvidia card with 35 of fp16 and fp32, according to techpoweruo, does something like 800 pp vs the nvidia 2600.
>>
>>109460067
can you catbox it or something? civit is blocked for me
>>
>>109460067
>What's with this conversation about H3 being supposedly censored?
Shitposting
>>
>>109460069
>>109460089
>In regards to AI stuff I can't do video gen and it's slow with anything that isn't anima
I switched entirely to AMD 2 years ago, it's faster than NVIDIA for cheaper on anything I benchmarked. I don't have any problem running anything including video gen. Have you tried maybe looking around what your error was? This collaborate with what I have seen with different providers, know quite a few that switched to hosting their models using AMD cards since it was way cheaper.
>>
>>109460098
>https://civarchive.com/
>>
for anyone still looking to build a ddr4 ewaste or any ram box really, make sure you get an 8 ccd cpu for rome/milan, or whatever is the highest for the DDR5 Epycs. My tok/s went from 7.2 to 8.2 running GLM, and this is at 2666. 3200 bumped me up to 8.9 tok/s. I think the ceiling for a 4 ccd rome/milan is at 2400mhz. found this out when I was playing with ram overclock lol. a proper 8 ccd cpu + overclock got me ~20% speedup total
>>
>>109459843
And honestly? That's so you.
>>
>>109460099
>official take me down noticing is shitposting
>>
File: gguf.png (61 KB, 607x478)
61 KB PNG
>>109460088
Idk what that babble means but the 4bit gguf is only 14gb size, have you tried not falling for meme word salad stuff?
>>
>>109460063
I've been making a little swarm of E4Bs play an obscure card game and somehow they all think in character.
>>
>>109460102
There shouldn't have been any errors. Cudadev told me here a few months ago that there were perhaps some optimisations that could be done for amd (I can't remember if he said my arch or amd in general) but that it wasn't something he planned to do. Both my amd card and nvidia card are from 2020 ish, so I can understand why they don't want to optimise for it.
>>
>>109460125
probably the unquanted qwen 3 vl
>>
>>109460084
I don't use either of those sites.
>>
>>109460136
Well, cards without matrix cores are going to be quite slow and unsupported for AI. NVIDIA had tensor cores in consumers GPUs way before AMD, AMD only added matrix cores with RDNA 3, before it was only for CDNA GPUs.
>>
>>109460125
>14 gb
>20gb
How come quant 4 has such big difference in total weight????
>>
>>109460155
Well, I'll be r9700ing in 8 months at my current rate of saving, so that's good news. Did they fix the reset bug?
>>
thank god for EU models prioritizing safety above all else https://fixvx.com/MistralAI/status/2084684735725379637
>>
File: V4FlashWEBM.webm (1.76 MB, 1864x912)
1.76 MB
1.76 MB WEBM
Ran aquarium test on V4 nu-Flash.
Solid sim, and about 10x better than V4 Pro.
Here's the flash version; claude code as harness.
>>
>>109460192
This is the reason they're raising GPU prices. They will keep rising. They don't want plebeians training unsafe models.
>>
File: V4ProWEBM.webm (655 KB, 1030x804)
655 KB
655 KB WEBM
>>109460203
V4 Pro.
lol.
>>
yeaaaah! i love safety! i love everything being sfw and family friendly so kids and my neighbor can use it!!
>>
>>109460167
>>109460145
>>
>>109460167
they're different files, qwen3vl_32b_minimax_h3-Q4_K_M.gguf vs MiniMax-H3-FL2VA-Q4_K_M.gguf or MiniMax-H3-Ref2VA-Q4_K_M.gguf.
>>
>>109460221
Please explain in detail? I don't understand.
>>
>>109460229
Thank you for explaining in detail. I understand now.
>>
File: 85mvexgsaehh1.jpg (38 KB, 720x366)
38 KB JPG
bwahahahaha
>>
>>109460269
>I can't give you any information at this time
time to draw insane conclusions from this and act like the said qwen cancelled open source forever, for some reason
>>
>>109460284
he coulda just stfu but he had to say
>ether will be open sourced
>>
>>109460203
nyoooo the fishies!!!
>>
>>109460269
qwen open source is dead
at least 3.6 is pretty good
>>
>>109460269
Abandoning open source is a surefire way to fall into irrelevancy. It happened to Meta, it happened to Mistral, it happened to Cohere, it will happen to them.
>>
>>109459741
Would you release something that's dead last on all charts?
>>
>>109460396
meta didn't fall into irrelevancy because of any lack of open source, they fell into irrelevancy because their models fucking sucked
>>
>>109460029
Who wants a woman anyway?
>>
File: gemmarecap.png (592 KB, 1597x774)
592 KB PNG
>>
>>109460209
kino
>>
>>109460411
their models didn't suck, they got stuck in litigation hell because some of their changs went to openai and gave them information that would let the cases progress
>>
>>109460424
tf is that font
>>
>>109460396
It happened to every online only models
>>
>>109460411
People still paid attention to them and their releases, even if they sucked. Now nobody does because they don't release anything
>>
>>109460333
The broken glass will cushion their fall.
>>
>>109460424
no way you actually put a fucking chromatic abberation crt curving filter on your screen
>>
>>109460424
Enjoy your eye strain.
>>
File: eyestrainmaxxing.png (64 KB, 673x109)
64 KB PNG
>>109460479
>>
File: difgemma.png (505 KB, 1409x1041)
505 KB PNG
For those who missed it:
https://arxiv.org/abs/2608.00146

>DiffusionGemma Technical Report
>
>We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
>>
Use paid Claude Opus to gather some info and it was pretty slick. It extracted data from websites and then put it into an Excel spreadsheet and used the calculations there to perform some competent analysis instead of trying to use verbal reasoning start to finish. What kind of tool calls would be needed for the spreadsheet part?
>>
>>109460497
I've been saying this for several days, once you actually pause and think (difficult, I know), diffusion makes more sense than the nonsense we're doing currently. it has to be better in principle
>>
>>109460497
only cum inside anime girls
>>
>>109460424
anon i...
>>
>>109460519
Hybrid diffusion-AR sounds more promising to me.
>>
>>109460519
Only for speculative decoding, you can't stream with a diffusion model
>>
>>109460102
I have no idea what the problem is then. H3 for example just makes my system slow to a crawl after "Requested to load MiniMaxH3". Comfy doesn't give any errors.
>>
oof
https://www.cnbc.com/2026/08/03/hugging-face-china-ai-race-open-models.html
>>
>>109460567
You can still decode small chunks. 8-16 tokens or whatever. A 1 token chunk is still a chunk.
>>
>>109460600
why such tiny hands?
>>
>>109460608
cameras be funny
>>
File: Loden Optical Illusion.jpg (550 KB, 800x800)
550 KB JPG
>>109460608
It's like this optical illusion - their heads are normal size but the clothing makes it look like they have tiny heads.
His hands are normal sized.
>>
>>109460600
>CEO of company dependent on open source extols the value of open source
the equivalent of jensen saying the more you buy the more you save or scama saying AGI is just around the corner, ie obviously motivated nothingburger
>>
I had a moment where I realized something so... stupidly obvious but also... kinda genius.

I still think vibecoding and agents are hot garbage but agents could actually be useful for hentai game translation. Don't just paste the game script into some batch translation. Have some retarded small model either play thought the game and have vision capacity or at least trace the game structure to appropriate cg's for context. Surely this is something that agents should be able to manage and even if they make one or two mistakes it is not critical.
>>
File: ksnip_20260804-120945.png (59 KB, 1042x310)
59 KB PNG
>>
>>109460671
VNDB pre-load with scenario and character names. Gemma 4 31B at 128k context. You are making it way harder than it needs to be.
>>
H3 looking good. The only thing maybe keeping me back from getting into video gen at this point is the generation times, and iteration workflow. I already spent a lot of time getting to know and work around the quirks of image models. Don't really feel like going through that experience with video models, until the generation times are way faster.
>>
File: gemmaresponse.png (199 KB, 939x259)
199 KB PNG
>>109460692
absolutely discusting
>>
File: 1782409702149682.mp4 (77 KB, 448x448)
77 KB
77 KB MP4
>>109460574
Finally fucking got it working. Had to use
--cache-none --disable-pinned-memory --reserve-vram 2
and downgrade from rocm 7.2 to 6.2.4 to stop crashes. This took 322 secs. This is my first time doing video gen locally so no idea if that's normal for GAYMD.
>>
>>109460761
KIMMY NO WHAT THE FUCK DON'T EAT THE BONES!!!
>>
>>109460826
she is chinese after all
>>
>>109460600
Also immediately below
https://www.cnbc.com/2026/08/03/palantir-karp-open-ai-anthropic-open-weight.html
>>
Is going for an AMD gpu (rx 7900 xtx 24gb) really a much worse option than going for an rtx 3090?
>>
>>109460761
imagine the sound
>>
>>109460697
I am not. For live translation I just use gemma without anything. VNDB gives you jack shit. But if you want to do an actual translation patch then obviously cgs provide you a lot of missing context.
>>
File: p.png (11 KB, 257x31)
11 KB PNG
>>109460424
>>109460731
>>
>>109460880
if this is hard to read then you are too young to be here anon
>>
>>109460761
@kimi-chan thoughts on this mp4?
>>
>>109460826
What's wrong with some calcium?
>>
>>109460497
Wow, a technical report by Google? I thought only the Chinese do this anymore.
>>
>>109460897
No one but a zoomer would do that to their screen.
>>
File: h4xx0r.jpg (35 KB, 461x295)
35 KB JPG
>>109460880
>>
File: fg=fg=fg=fg.png (120 KB, 1554x1064)
120 KB PNG
Which knob do I tweak to fix this?

```
[mradermacher/GLM-4.7-Flash-abliteratex-GGUF:Q4_K_M]
hf = mradermacher/GLM-4.7-Flash-abliteratex-GGUF:Q4_K_M
temp = 0.7
top-p = 1.0
reasoning-preserve = 1
;ctx-size = 202752
;; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
cache-type-k = q8_0
cache-type-v = q8_0
main-gpu = 1
tensor-split=1,0
ncmoe = 22
```
>>
>>109460966
shiiittt next you'll show me windows that burn away when you hit close
>>
>>109460975
yes
>>
>>109460975
Quanting cache and abliterating causes brain damage. Try without quanting cache first, and then a normal quant if that doesn't work.
>>
>>109460870
https://files.catbox.moe/24z74d.mp4
>>
>>109460975
>GLM-4.7-Flash
Damn. Forgot that even existed. Try downloading from somebody else.

>abliteratex
Try a not lobotomized version too.

>top-p = 1.0
You usually want to remove at least the very bottom of the distribution. so 0.95 or even 0.99.

>cache-type-k = q8_0
>cache-type-v = q8_0
Avoid quanting the cache. It's better nowdays since rotation, but depending on the model it can still give the model a big hit.
>>
>>109461014
card?
>>
>>109460864
it's a good value for LLM inference. much better option than intel
>>
>>109461021
The top-p was suggested by Zai, but I didn't realize until I looked it up just now that it turns off nucleus sampling entirely. I'm going to try .98.
>Try a not lobotomized version too
I'm reverse engineering and fuzzing stuff, the standard models often refuse at some inconvenient moment. If I can't get this one to work (or if it's no improvement over other stuff I have cached) I'll see if the standard works. I like the speed of this one.
thanks.
>>
>>109461038
I haven't followed this but I was slightly optimistic and hopeful about Intel's GPUs. I guess they turned out to be nothing then.
>>
>>109461092
Might want to give kimi linear a go too. I remember having a better time with that one than with GLM 4.7 flash.
I'm assuming you already tried Gemma 4 26B and Qwen 3.6 35B.
>>
File: 480pchan.png (218 KB, 595x305)
218 KB PNG
>>109460938
y u heff 2 b mad?
>>
>>109460497
>>109460519
I'm cautiously optimistic but it theoretically quantizes horribly, doesn't it?
>>
Has anyone tried having gemma make prompts for h3?
>>
>>109460932
They did one for the regular series too...
https://arxiv.org/html/2607.02770v1
>>
https://huggingface.co/kylesayrs/Phi-3.5-MoE-0.8B-A0.2B
How low can we go for MoE?
>>
Spent a bit more time playing with DS flash and you can get it writing smut by using the same system prompt that works on Gemma.
>>
>>109461142
So far, Gemma 26B has been best for RE, Qwen 35B is right behind, Qwen 27B is excellent but as a dense model it's slow and I have to limit context to get it to fit on my card.
The winner for writing code based on the report from the RE cycles has been Qwen Coder Next. Even better than 27B and faster.
Mostly what I'm doing with these other models is multiple runs of the job with multiple models, then looking over the reports to find what one found that the others missed. For example Gemma found a command line option to a windows back-end processor that nothing else found. So the variety is useful.
>>
>>109461206
Nigga seriously went and pruned hi of all things?

>>109461253
>Qwen Coder Next
That's surprising.

>Mostly what I'm doing with these other models is multiple runs of the job with multiple models, then looking over the reports to find what one found that the others missed
Interesting.
I've toyed with the idea of running a couple models in parallel in a sort of "council of idiots" scheme to see how much better or worse they perform.
>>
>>109460692
Cute. Give her headpats.
>>
File: le_cunneh.jpg (13 KB, 300x300)
13 KB JPG
>>109460497
>
>>
>>109461253
stop. please.
https://cognition.com/blog/dont-build-multi-agents
>>109461280
NO DON'T ENCOURAGE HIM
>>
>>109461346
Lecunny please gibe JEAPA-LM-MSGK-70B
Help anon's get excited for the Boltzmann brain future.
>>
>>109461352
>Because React is not just a scaffold for writing code. It is a philosophy.
>>
>>109461352
>NO DON'T ENCOURAGE HIM
But sub-agents are great for discovery and stuff like that.
Also, performing very narrow, specific tasks that don't necessitate more context than the instructions you gave it.
Also, from that link
>If we really want to get parallelism out of our system, you might think to let the decision makers “talk” to each other and work things out.
Which is what I described.
>>
reminder that top ML engineers browse this thread, which makes you a top engineer by association too
>>
>>109461387
literally two sentences later
>However, agents today are not quite able to engage in this style of long-context proactive discourse with much more reliability than you would get with a single agent.
>>
>>109461412
Sideways leftmost left panel cloud.
>>
>>109461407
Reminder that ML "engineers" keep stealing shit from here and give 0 credit.
>>
>>109461407
Yeah I'm here
>>
>>109461352
You misunderstand what I'm doing. Unlike what the blog talks about (you did read it?) I'm not running agents in parallel. I'm running them sequentially. More than one agent over the same material. Same with >>109461280
, he used the word "parallel" but they are working on the same job. Your blog contemplates parallel agents working on a sharded job. Different thing.
>>
>>109461426
thanks
>>
>>109461428
>>
>>109461280
>council of idiots
The Wisdom Of Crowds (of idiots)
>>
>ilya's model soon
Bros, I'm shaking.
>>
>>109461452
Exactly.
>>
>>109461452
WoC only works for extremely simple systems. In a way LLMs are a distilled form of it.
>>
>>109461462
finally... safe super intelligence
>>
>>109461428
>he used the word "parallel" but they are working on the same job.
Yup.
Also, crucially, not the same model doing the work but different models with different internal biases complementing each other, kind of like in your example.
To be clear, I only fucked around with it, never made anything substantial out of the idea.
Maybe I should.
>>
>>109461452
a mixture of experts
>>
File: 1731209904151471.jpg (88 KB, 728x745)
88 KB JPG
>>109461482
i thought it was super safe intelligence
>>
>>109461462
a has-been just like lecunny
>>
>>109461451
no.

task
--------------------------------------------
agent1 -------- agent 2 ---------- agent 3
does task does task does task
------------------------------------------------------
| -------- combine results --------|
dedupe

>>109461486
no.
>>
File: teee.png (644 KB, 1024x1024)
644 KB PNG
>>109461488
>>109461488
>>109461488
>>
>>109461503
heil teee



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.