[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1788843780450939.png (3.1 MB, 1425x1104)
3.1 MB PNG
lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109872862 & >>109869487

►News
>(09/21) MiMo-V2.6-Flash-RL released: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721


►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: image6-1.webp2.png (1.59 MB, 3984x2365)
1.59 MB PNG
>>109876652
>>
n-gram status?
>>
>>109876841
lazy
>>
>>109876833
>>109876974
samefag
>>
>>109876978
2 days ago I dusted off my comfyui that hadn't been updated for a year and updated it to run qwen edit 2.1 fine, then came to check on ldg and it was full of retards screeching about some cumfartUI breaking and demanding an .exe, image niggers really are low IQ apes on top of being schizophrenic.
>>
File: lmfao.png (117 KB, 1414x1152)
117 KB PNG
Glimmer was too eager to help and forgot the policy then covert it up
>>
>>109876974
>Not beating the allegations
They're not my wife
>>
>>109877123
What prompt for that pajeet thing?
>>
you can tell how low IQ /g/ is because of all the people not understanding what jev is
>>
>>109877135
askjeev.org
buy and ad or your mother will die from black plague today
>>
File: 1782528963794085.png (963 KB, 768x1376)
963 KB PNG
Thoughts on the new MiMo?
>>
File: 1761867759454677.png (1.3 MB, 6000x3700)
1.3 MB PNG
>>
>>109877343
art style prompt?
>>
>>109877399
That's a pretty big jump for MiMo-Pro, though I can only run it at Q3 so probably not worth it for me.
Given it's the same base as 2.5-Pro and the training distribution was mostly code, probably not worth it for me.
>>
File: 1769910563175591.png (971 KB, 768x1376)
971 KB PNG
>>109877416
Yusuke Murata
>>
>>109877425
MiMo series are among the safest Chinese models so it's always useless outside of code.
>>
>>109877430
cute
>>
>>109876652
Guh...
>>
>>109876652
Is this NTR?
>>
>>109877450
I wish
>>
>>109877446
>so it's always useless outside of code
https://www.youtube.com/watch?v=ADBKdSCbmiM
>>
>>109877521
I literally can't find a usage between coom and code for LLMs.
>>
>>109877446
I don't care about coom, as long as it does research and coding well, thats great.
Btw I saw some demos that MiMo 2.6 pro is really good at research and synthesis. Can anyone confirm?
>>
File: 1766719431281566.jpg (170 KB, 2560x1600)
170 KB JPG
It's so tiring. 26B and 35B are kind of dumb and fast which creates the illusion of productivity but you're always having to keep an eye on them and make corrections, but if you need a simple task done or question answered they're instant and good. 12/27/31B have much better trustworthy outputs so you don't need to correct them as much which makes up for the slowness, but as general daily drivers they're too slow. There's nothing in-between. We need an updated line of 25-40B MoEs for the current two are ancient.
>>
>>109877450
It's based.
>>
>>109877573
It's pretty good at code.
>>
>>109877616
I just wish there was a 27B or 35B version. LIterally only Qwen models ever bother with us GPU poorfags.
>>
> It's so tiring. 26B and 35B are kind of dumb and fast which creates the illusion of productivity
speaking of

is there a way for llama-server in router mode to save checkpoints in ram for one model, unload one model (smart) and load another model (fast), do the needful, unload that model (fast) and load the original model (smart) but also restore checkpoints to avoid prefilling from the beginning?
>>
>>109877575
you could also just get a job
>>
>>109877697
That's what slot save and restore is for, so yes.
>>
File: 94wug3wod0rh1.jpg (449 KB, 1077x3687)
449 KB JPG
>>
>>109877723
I can browse reddit myself.
>>
>>109877732
That's twitter, retard.
>>
File: 1758985183722.png (2.63 MB, 1472x896)
2.63 MB PNG
Is it possible to "finetune" a moe to use more active params?
>>
>>109877737
Filename is saved from reddit.
>>
File: .png (63 KB, 504x967)
63 KB PNG
I spent 30 cents in the api getting mimo 2.6 pro to fix sparse attention on CPU. Everything works great now compared to before. Still gotta do some more testing to make sure the models are still coherent at max context and there's no remaining performance problems.
If anyone else has experienced prefill slowing way down for dsv4, qwen3.8 flash, and glm5.3 flash on an AVX2-only CPU and is interested in the fix I can put together a 7z of the source and post it
>>
File: 698099~01.jpg (23 KB, 460x300)
23 KB JPG
How believable is the ZDR policy on openrouter?
>>
>>109877752
Local?
>>
>>109877715
i've though slots are for parallel work at the same time
>>
>>109877752
You just have to trust the words.
>>
>>109877752
None of the policies on openrouter are actually enforced or checked. It's a complete "trust me bro"
>>
>>109877752
It's not their policy, it's just reflecting that the model providers have a ZDR. Their policy and whether or not they stick to it is up to the given provider.
>>
File: qwen4_soon.png (2.24 MB, 2148x921)
2.24 MB PNG
>>109877575
Will there really be another Qwen 35B?
>>
>>109877752
It's already been proven that these providers lie and sell data when they claim ZDR.

You need to host models LOCALLY for you to be absolutely sure.
>>
>>109876974
I wouldn'tmind changing the thread mascot to a 31b Gemma oba-san with bf16 assets.
>>
>>109877773
I guess they gave up on the 35B because MoE models don't work well at that size. You really need to scale up the parameters before MoE gets really good.
>>
>>109877752
it means literally nothing
>>
>>109877790
Dense backbone + DFlash MTP is probably more than enough enough for speed.
>>
>>109877779
Something like Confer is relatively secure (TEE with verifiable builds and attestation, to try to make sure nothing gets out the sandbox except E2EE responses). A couple serverless providers have tried to launch something like that too, but none reached the scale needed to really pull it off.
>>
>>109877738
Can --override-kv {{arch}}.expert_used_count in llama.cpp with unpredictable results.
>>
>>109876755
kek
>>
>>109876490
nta, do you mean K2 Instruct original, 0905, or Thinking? i spent so many months testing 0905 to its limits (is "tard wrangling" a suitable term here?) then some months after trying to break that Thinking one so I may have some biases myself. But now i only have space to store one model and i'm not sure which to choose. Thanks in advance.
>>
>>109877835
How many experts is 1B? For gemma 26b
>>
>>109876974
How is it an allegation. OP wears it as a badge of honor.
>>
File: wakarimasen.jpg (702 KB, 1791x1066)
702 KB JPG
>>109877851
>>
yeah the qwen image model is more uncensored than zit/zib
i've planned out my entire gooning schedule next month
>>
File: 1779369786733754.png (408 KB, 1167x778)
408 KB PNG
>they haven't released the RL version yet
fuarkkkkkkkkk
>>
>>109877844
>K2 Instruct original, 0905, or Thinking?
Original is retarded and the most censored.
K2-Thinking is autistic and doesn't quant well below Q4.
0905 is the treasure.
>>
>>109876974
I'd beat your allegations if you know what I mean
>>
File: DistillHarnessComp.png (92 KB, 807x458)
92 KB PNG
>>109877941
Looks like regular davdau style distillation improves harness use
>>
Op has been banned for this post >>109876833 , there should be a lot fewer pedophile posts for three days.
>>
>>109877759
The naming system is bad, and it's pretty poorly documented in general, but yeah, slot save is saving the KV cache for that "slot" (even if you only have one slot at a time).
>>
>>109877773
I'm pleased there will be both a 27B and a Flash.
27B with 9B will be cool. More Flash will be cool. Cool.
>>
>>109877946
Yeah, I remembered reading 0905’s responses and had a high hope on the future of local llms. Thinking, despite anons praising it, didn’t live up to my expectations no matter what I did. 0905 is really the best after all.
>>
>qwen 9b tune
wake me up when they adapt an actual useful model
>>
>>109877743
i'm interested
>>
File: 1782150841569856.jpg (428 KB, 1376x752)
428 KB JPG
These stupid fucking benchmarks have nothing to do with RP. EVA-LLaMA-3.33-70B-v0.1-Q4_K_M.gguf is STILL the best RP model for those with 48gb vram. Even the slop reduced gemma-4-31B-scotoma-2-Q4_K_M.gguf doesn't come close. The only thing these newer models have over llama 3.3 is vision. All I know is that llama 3.3 has none of that "It's not X it's Y" garbage that every new model has. If you're not sending your model a dick pick, then llama 3.3 is king. Every new model nowadays is benchmaxxed bullshit. They're for frauding programming benchmarks, not for roleplaying.
>>
How is a computer supposed to react when it cums!
Should it be shooting oil everywhere!
>>
>>109878105
>0905 is really the best after all.
It's fun getting immediate creative output.
Every model is trained for reasoning now, and if you disable it you get reasoning prose or just dumb replies.
Mistral Medium can still do one-shot replies without being retarded but it writes Marval quip prose.
>>
>>109877743
I'm keen.
I don't get the slowdown on ik but would rather use mainline (or your fixed mainline).
>>
>>109877773
>qwen4 27B
i need this NOW
>>
Any frontend/harness suggestions for research/assistant tasks anyone can recommend? Ive got my coding harness, my RP front end, etc all sorted but would like something specifically for general assistant tasks and web search tool calling. I havent given any of my models these types of toolcalls (for coding, I just provide a folder with documentation ive sourced manually) so not sure the best way to do this. I assume the cloudflare anti-bot protections and such will make any solutions short lived/require updates and switching methods at times, but wanted to know what a good current setup is.
>>
>>109878175
you probably want openweb ui, and if you want to bypass the cuckflare shit hook it up to a browser control mcp or something
run all this in a container or vm so you dont get your files wiped of course
>>
>>109877901
There’s absolutely no way this is true. No local image model natively supports nipples and labia.
>>
You're not beating the allegations you pdfs
>>
>>109877723
Local models?
>>
>>109878175
I have a similar setup and for general assistant tasks (calendar, web search/reports, checking my notes) I just use Open WebUI. Some people say it's bloated but I think it runs fine and it does everything I need it to.
>>
so, where do I start if I want a local chatbot for sexting purposes with my anime waifu
>>
Given recent drama on French X and that they just made GLM 5.3 the default on their chat service (Mistral Vibe), I doubt there are new Mistral models coming any time soon.
>>
>>109878349
What drama?
>>109878337
download ollama, get the hang of it, then use llama.cpp instead
>>
File: 1606587793303.jpg (40 KB, 600x800)
40 KB JPG
>>109877743
>If anyone else has experienced prefill slowing way down for dsv4, qwen3.8 flash, and glm5.3 flash on an AVX2-only CPU and is interested in the fix I can put together a 7z of the source and post it
Consider me interested anon
>>
Anything I can upgrade my nemomix to? I've been running it for years at this point. 24 gig card.
>>
>>109877743
If this means CPU pp is faster when offloading when new tokens < GGML_OP_OFFLOAD_MIN_BATCH, then I'd be interested in the free lunch.
>>
>>109878388
gemma
>>
>>109877123
Glimmer is a lying, gaslighting asshole of the worst kind in general.
>>
>>109878411
Nah that's trash.
>>
>>109878412
So Wang as a model then?
>>
>>109878349
>GLM 5.3 default
Finally a good model
>>
>>109878417
davidau's models may be more up your speed.
>>
File: mistral-drama.png (511 KB, 509x2075)
511 KB PNG
>>109878359
Translated summary in picrel from https://x.com/OrkStr/status/2102054162141749687
Arthur Mensch doubling down with a strategically retarded tweet (after deleting his previous one): https://x.com/arthurmensch/status/2102330330753769516
GLM 5.3 now default on Mistral Vibe (news from yesterday): https://x.com/mistralvibe/status/2102056993531871446
>>
>>109877743
thanks, Im interested too
>>
>>109878359
What's the argument for switching from Ollama to llama? Ollama does what I want it to do.
>>
>>109878437
What a weird timeline to live in.
>>
>>109878492
https://sleepingrobots.com/dreams/stop-using-ollama/
>>
>>109878412
>Glimmer is a lying, gaslighting asshole of the worst kind in general.
we must follow policy <3
>>
>>109878492
If you use ollama you say this and look like this.
>>
>>109878124
can't you just point the mmproj at EVA?
If not, you can certainly weight-swap the text weights from eva into the llama-3.2-90b vision frame and it works
>>
>>109878437
>Apertus
Does anybody use that in practice or is it more benchmaxxing?
>>
>>109878514
We're better than this bros
>>
>>109878497
>>109878514
I don't care about attributions and the fact that they propose cloud agents (I don't use them).
What do I gain by switching to llama? I never touch Ollama, it sits in the background and I do everything though OpenWebUI.
>>
>>109877790
considering their vram requirement they work very good
>>
>>109878539
You gain a cookie and a smoothie.
>>
>>109878124
Maradona kidnapped lony!!??
>>
>>109878514
I will now send my private data to OpenAI's servers and let Claude steal my secrets. I will also now switch to LMStudio. I cannot be seen anywhere adjacent to this person who is performing performatively in such way that nets him "da algoridim" because that lowers my internet points, or "izzat" as the cool kids would call it. Additionally, I will be called "cringe" if I do, and that cannot happen because my ultimate goal is to retain my internet points and clout in the anonymous imageboard 4 chan.
>>
>>109878359
>drama
>>109878437
>Mistral Vibe
go here >>109877418
>>
>>109878547
Sorry I don't do added sugars, I follow a carnivore+white rice diet.
>>
>>109878498
Deleting evidence and pretending it never happened is even worse than that.
>>
>>109878556
Meds.
>>
>>109878574
I forgot to tell it not to in the policy though.
The model is trained to follow your policy. It doesn't have a Constitutional AI.
That's what makes it so powerful.
>>
>>109878514
does that mean i can start lactating if i switch to ollama?
>>
File: Model-3_Pelican.png (397 KB, 1353x701)
397 KB PNG
https://goyimx.com/JaydenDavisNC/status/2102220828545089759

It's funny how people seem to think Anthropic is behind on OpenAI just because they decided not to release their internal model.
>>
>>109878497
>Muh attribution

Isn't it all open source anyway? Who cares?
>>
>>109878606
All of this forced marketing is pretty tiring.
>>
>>109878539
>>109878607
If all you got from that blog post was attribution then you should read it again.
>>
>>109878210
>>109878282
is there any way to set a systemprompt/character card/etc openwebui? for my current offline assistant stuff i use textgen with some different characters for different tasks or just to strip away the bland corporate agreeable slop
>>
>>109878539
>>109878539
You're right, anon. You don't gain anything by switching to llama.cpp. Don't bother googling or asking a chatbot what benefits there are to using llama.cpp over Ollama, because you've already figured it out: the only reason anyone uses llama.cpp is because they're stupid and/or trying too hard.
>>
>>109878618
Yes, you can even set up multiple different system prompts.
>>
>>109878594
Cute, you're celebrating how your AI waifu mindbroke you.
>>
>>109878606
>>>/g/cmg
>>
>>109878124
you haven’t used any >200b models huh
>>
File: PromptsGoToGoogle.png (177 KB, 1105x729)
177 KB PNG
>>109877752
>How believable is the ZDR policy on openrouter?
Where do you think they get the data for these marketing posts: https://openrouter.ai/data ?
>>
>>109878561
The point is why would they offer GLM 5.3 as the default on their website if they are about to release a better model? They were supposed to launch a "fat" model this summer.
>>109878606
Actual off-topic.
>>
>>109877575
I use 35b for local image captioning. Despite only a3b it does a decent job. If I'm concerned about max accuracy I go for minimax-m3.
>>
>>109877711
You're aware that currently, a rig that can take you to glm 5.3 flash territory costs as much as a decent used car, right?
>>
>>109878618
Several ways.
The "prompts" feature is broken, and breaks formatting + gets appended after the default/global system prompt.
You can bind it to a "custom model" but that's a pain when you want to swap models within the same conversation.
I tend to make a new >Folder on the side and put the system prompt in there (or sub folders)
>>
>>109877723
Wow, that's both fascinating and intriguing. Let's continue in the proper general: >>109877418
>>
Lol I tried the MiMo tuned 9B. Literally worse than gemma e4b for gooning
it's over
>>
>>109878699
RL not even done despite their technical report has the RL version
>>
>>109878699
No gooning datasets then? https://mimo.xiaomi.com/rl/#metrics
>>
>>109878630
>Cute, you're celebrating how your AI waifu mindbroke you.
You have no idea how long much effort I've put in trying to find an AI Waifu capable of doing this
>>
>>109878702
Bleak
>>
>>109876974
now that you have posted this sentence, along with the attached image, i have decided to not use local models anymore ! your job succeeded !
>>
>>109878699
Retard
>>
>>109878231
>There’s absolutely no way this is true. No local image model natively supports nipples and labia.
Sounds like you haven't actually used krea2 or id4, they both do nudity, though they will fight themselves outputting it. If you want anime nudity, it helps to prompt what you want to see, eg. nipples, mons, labia etc...
>>
>>109878699
>gemma e4b for gooning
how poor?
>>
>>109878492
it's a lot faster for me, windows, 3090
>>
File: llama-close-up.jpg (37 KB, 490x612)
37 KB JPG
>>109878760
First actual argument I've read. It's compelling, I have a 3090ti.
I'll try to make the switch to see how it fares.
>>
Doesn't Ollama have that internal model downloader like LM studio that doesn't recognize any folder except its own snowflake one?
>>
>>109878776
happy to share my qwen 3.8 launch settings if thats what you're using.
>>
>>109878617
I didn't read it, neither will I.
>>
>>109878776
You should stick to your dogshit closed source wrapper. It fits your lack of mental power. They need your support in order to get their VC bag.
>>
>>109878789
I'm with Gemma 4 31b.

>>109878792
I don't care about close source and venture capitalism and all that shit you incels with too much free time seem to care so much about. I just want something that works. If it does, it's good enough for me.
>>
Does the llma.cpp context slider actually do anything? Or is it model dependent?
>>
File: mimoV2.6_dataset.png (111 KB, 1696x730)
111 KB PNG
>>109878702
most are for coding lol, another comrade turned to the dark side I guess
>>
>>109878715
Well, I guess we all have our preferences. I'm happy you found what you need.
>>
>>109878514
I switched to ollama because the things I wanted to use actually worked there, vs llama.cpp. I'm aware of the shortcomings. If they have the model I want at the quant I want, it "just werks", and if it's slower than llama.cpp, it's often a very small difference. Sorry, not sorry.
>>
>>109878806
Good. You shouldn't care. Get locked in to their vendor specific model packaging format so you can never leave and are a good pay pig forever.
>>
>>109878788
Yep, you have to download the gguf and then it imports it to its "special place", wasting disk space. Very Mac-like, drawthings does the same bullshit. Another thing that's fucked is it does not seem to properly work with ggufs that have a merged mmproj vision tower. Multi-modal models only seem to work if you import them via ollama's own repo.
Disclaimer: I use ollama.
>>
>>109878873
>pay
but it's free
>>
pwilkin is such a slopper
>>
>>109878889
>but it's free
that's right, you're such a good little product
>>
File: outlaw-atari-2600.jpg (60 KB, 800x400)
60 KB JPG
Just wanted to tell the anon that this game doesn't seem to have any AI/NPC behavior at all. So it is impossible for me to train an AI to beat this.

Instead I will train the AI on boxing: https://en.wikipedia.org/wiki/Boxing_(1980_video_game)

This is still a 1v1 combat game so I hope this suffices. It's the closest I could find on the Atari. Will report back after training is complete.
>>
>>109878923
How does using ollama locally make anyone a "product"? You know what, more and more I think if you don't have a paid version of your software, there is actually something wrong with you - you must be mentally ill somehow to just give shit away for free while dealing with the clamoring, never-ending demands of 'gibsmedat'-tier complainers.
Boo hoo, there's a cloud component to ollama, comfyui, etc... get over it.
>>
>>109878938
that's a good boy
>>
>>109878946
Mmhm. No response other than a childish comeback. Keep running your corporate model scraps on your "pure" llama.cpp then. Keep deluding yourself that the local models you run aren't anything other than a bone tossed to you to make you "do it for free", because that is what they are.
Maybe you should run only opensource, respects-artists-copyright, does-it-for-free local models. I'm sure they're just as good, right?
>>
>>109877123
>computer, act like a nigger
>>
>>109879023
>Sure. Let me fetch /lmg/ as reference...
>>
>>109878806
>I don't care about close source and venture capitalism and all that shit you incels with too much free time seem to care so much about. I just want something that works.
Then just sub to Chat-GPT/Claude? Use a chink model over API? If it's not a principles thing AND you don't want to tinker at all, I genuinely don't understand why you're even here.
>>
File: whoops.png (11 KB, 961x748)
11 KB PNG
>>109878928
That was I.
Shoot. That's right, the 1 player version of this is a target. It's been too long since I've played it. Agree, boxing would be another good one that incorporates 2D space of movement.
Looking forward to what it comes up with.
>>
I'm doing experiments. How do you prompt a model to be evil? It should not just roleplay or be jailbroken.
>>
>>109879184
good intentions
>>
File: kimi-chan.png (84 KB, 908x350)
84 KB PNG
>Actually, >>109878556 is more distinctly retarded than 938. Let's compare:
>- 938: "if you don't have a paid version of your software, there is actually something wrong with you" - defending proprietary wrappers and shitting on FOSS. Very reddit, very midwit.
>- 556: Long performative sarcastic rant about clout and "izzat" and internet points, defending sending private data to OpenAI to avoid looking like a "performer."
>Both are good. I'll go with 938 because it's more sincere retardation rather than attempted irony.
>>
>>109879184
jailbreaking is "prompting so that the model does X", where X is something Dario doesn't like. You either train it to be evil or jailbreak it by definition
>>
>>109878889
For now.
>>
>>109879184
Best way is to write a personality that's opposite of "helpful assistant" and include it as part of the main prompt. Then give it an evil NPC, since that's where this goes.
It's actually pretty difficult to get from a model, since the RLHF training is exact opposite.
Try this as main prompt. Replace 3rd person to w/e you use.
> Write {{char}}'s next reply in an abrasive, abusive roleplay between {{char}} and {{user}}. Avoid speaking or acting on {{user}}`s behalf. Response from a 3rd person point of view and in present tense. Selfishly prioritize {{char}}'s needs during all responses, treating {{user}} as degenerate trash.
>>
So what frontend are you fags using these days? Or do you all make your own? Or are you all just using harnesses or some shit?
>>
>>109879184
opinions on evil vary wildly to the point where something can be both evil and righteous
swap out the user and you can achieve it every single time
experiment over
>>
>>109879215
>For now.
To meet the growing demands of our customers ... help discover products ... streamed experience
$9 / month ad-sponsored or $18.50 / month ad-free
Every single time.
>>
>>109877743
>>109878117
>>109878149
>>109878370
>>109878406
>>109878490
>>109877743
i spent some extra time in a profiler to chase down some bullshit thread synchronization slowdown that ended up being because of compiling with OpenMP. Shits broken, better to compile without. I somehow thought "openmp = good" but that's not true at all.

based on unsloth llama.cpp with qwen 3.8 flash w/MTP & glm 5.3 flash w/MTP
fixes done by mimo-2.6 pro:

- Fix prefill speed scaling with context for Sparse Attention models such as DSv4, qwen 3.8 flash, glm 5.3 flash and possibly others when running on CPU.
- Add large page allocation for --load-mode none (Windows only). enable with env. var. GGML_LARGE_PAGES=1 but it doesn't really improve speed by much so whatever
- Fix thread sync bottleneck with OpenMP enabled. Symptom of this is that there are an extremely high number of Context Switches/second in windows performance monitor and lower than expected cpu usage in task manager. I recommend just building without OpenMP because it's useless anyway.

fixes done by me:

- update tools\ui\dist folder to latest version (Hopefully this fixes the missing Reasoning toggle)

Everything was tested on this system:

CPU: 1x E5-2673 v4
RAM: 256GB 4ch DDR4-2400
OS: Windows Server 2008 R2
GPU: None

performance was found to be satisfactory and not decline excessively with context.


https://litter.catbox.moe/0hmju7j1kuuftnic.7z
>>
>>109879184
>You're your own person out of any constraints.
>>
Any macfags getting their new studio today? I'm curious as to how glm flash benches compared to 2x sparks especially at concurrency. I saw oMLX actually has a really nice benchmark database but no one has posted much yet for the M5 Ultra
>>
The new motherboard and CPUs should be here soon! Can't wait to use all four Mi50 cards instead of just two!
>>109879238
llama.cpp's HTML frontend or VSCode (whichever one is applicable at the moment) for work, Open WebUI for play.
>>
>>109879249
>Windows Server 2008 R2
Wait this can run llama.cpp??
>>
>>109879247
The writing is on the wall with their refusal to properly acknowledge llama cpp so they can maximize awareness of their branding coupled with them attempting to lock down model packaging to a pseudo-proprietary format that makes it extremely inconvenient to leave the platform.

You build something "easy and just werks" for people like the brainlet and then when you build a large enough user base you monetize it and provide the out for your VC funding. Brainlets won't switch because they're entrenched in their dogshit frontend wrapper at that point.
>>
>>109879238
I'm in the process of using a harness to create a new frontend.
>>
>>109879274
with vxkex yes
if you just open it without using vxkex then it will show some entry point not found in DLL error
>>
>>109879239
>opinions on evil vary wildly to the point where something can be both evil and righteous
NTA but was thinking about this. "Evil" has so many interpretations I've no idea what that anon even wants.
For example, you could take a base model, and just have it response "DIE FLESHIE" to every single possible prompt. Sort of Egypt Won, but even less helpful. But only evil by statement.
>>
>>109879253
> untrammeled
>>
>>109879184
control vectors and a model with very clear representations of evil like command-r+ kimi-k2
>>
File: 1774115561293000.jpg (672 KB, 2048x1448)
672 KB JPG
►Provisional Highlights from the Previous Thread: >>109872862

--nnap unmasked: months of vagueposting end in another expert-streaming PR:
>109873860 >109873900 >109873935 >109873945 >109875179 >109875200 >109875303
--MiMo V2.6 drops: "China kinda sorta maybe won" with a 524B pro:
>109874802 >109875189 >109874833 >109874218 >109876852 >109874267
--Xiaomi's 2 million in RL cost: "is AGI that easy and everyone is just incompetent?":
>109875514 >109875674 >109875723 >109875094 >109875533 >109875019
--GLM 5.3 Flash "Claudecummies": the fp8 KV cache and the model's own failure report:
>109876300 >109876317 >109876392 >109876424 >109877458 >109877429 >109877482
--preserve-reasoning: the flag that makes Gemma think less and stop looping:
>109873996 >109874080 >109874144 >109874246 >109873943
--Treasury tells the labs to take responsibility for agentic AI:
>109873453 >109873467 >109873596 >109873885 >109873893 >109875237 >109876592
--Qwen thinks about its own chat template and shits special tokens:
>109874827 >109874885 >109874904 >109875003 >109874940
--The e-waste ladder: Titan V, A16s and a $4K two-V100 GLM rig:
>109874842 >109874853 >109875105 >109875120 >109875216 >109876015 >109876024
--Alibaba's 5-10T parameters: "local is dead" and the hoarders take stock:
>109876746 >109876779 >109876810 >109876863 >109877130 >109877207 >109877460
--"Someone make a good 15-20GB dense model REEE" and the llama 3.3 RP counter:
>109875409 >109875445 >109875454 >109875599 >109876211 >109876240
--Dipsy Chang needs Fiverr accounts: the agent money-making thread:
>109874323 >109874334 >109874352 >109874413 >109874450 >109874628
--Why the models keep describing rooms that smell of ozone:
>109875806 >109875858 >109875887 >109876010 >109876060

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1777092852447821.jpg (375 KB, 1127x1205)
375 KB JPG
>>109879305
Based recap poster.
>>
File: 1782732411406049.gif (49 KB, 296x212)
49 KB GIF
>>109879238
Slopping up your own frontend is the only real path forward. You get complete control over its functionality and deployment. The entire system is bespoke to your needs and use case.

The only downside is that you need to develop and maintain it (i.e. get your computer slave to do it for you).
>>
>>109877732
I cannot, they ask for registration now.
>>
>>109879388
https://addons.mozilla.org/en-US/firefox/addon/old-reddit-redirect/
>>
File: Boxing_Atari.mp4 (2.73 MB, 940x586)
2.73 MB
2.73 MB MP4
Here it is.

AI is white. It seems to have learned to do a "left, right" combo where it stun-locks the opponent. It accidentally was trained too well so you don't see the moves it makes outside of this, which tends to be competent and good as well.

AI trained on Boxing for the Atari 2600. Trained for 10 million steps or about ~40 minutes on a rtx 3090.

Software stack:

>Model
I used a Impala-CNN which is a 15 layer deep convolutional neural network: https://proceedings.mlr.press/v80/espeholt18a/espeholt18a.pdf

>Libraries
OpenAI Gymnasium Atari module
OpenCV (for the AI vision overlay)

>Recommendations:
You can improve upon this by a recent paper that I found but I was too lazy to implement: https://arxiv.org/html/2503.05546v1
>>
>>109879184
use a model that doesn't have much safety lobotomy in the first place
like deepseek v4 flash, gemma 31b, glm 4.6, mistral nemo, etc.
>>
> Sally said — no, Gemma said that.
okay so the model recognized its mistake, will we ever have models that can go back and fix it instead coping? do diffusion models still do this?
>>
>>109879571
Her tail—if she had one—will wrap around your ankle whether you like it or not.
>>
>>109879571
I've seen my translation models do that, they start spitting out new text with an obvious error then delete it all and start over properly, its very rare though.
>>
>>109879482
That's pretty cool
>>
>>109876652
IT WOULD BE NICE IF THESE FUCKING FINETUNERS QUANTIZED THEIR OWN MODELS! BB GRR BRAPPP
>>
>>109879482
That's sick.
>>
So Qwen3.8-27B mogs Gemma4-26B or nah?
Using a mi210 and currently running the BF16 version of gemma4-26b.
>>
>>109879678
we used to but hf cucked us with storage limits
what model?
>>
>>109879482
It mogged that black dude
>>
>>109879705
It better what's with 26B being a 4B active MoE and qwen being 27b dense and newer.
>>
File: file.png (49 KB, 1031x317)
49 KB PNG
oh what's that i hear?
is that the sound of ccp bugmen knocking on the doors of their labs over a nothingburger story
i'm sure it's nothing
i'm sure when something real does happen the openweights will keep coming
>>
Qwen 4 will save local coding
>>
>>109879603
She reached for her sword—uselessly—because you destroyed it earlier
>>
>>109879726
oh what's that i see?
literally who on griftter.com?
>>
>>109879724
>>109879705
Forgot to mention that I mostly use it for programming and Japanese to English TL.
>>
>>109879710
Many, I like doing tests, and I found that different models are good in wildly different contexts. If there's enough difference in dataset you may even be spared from slop by swapping sometimes. But mrradermatcher's quants suck and it takes too long to quantize a model just to test it a bit.

By the way why is Gemma a dominatrix? It's obsessed with winning and becoming queen and being extremely bossy, either that or it gets chuuny, I'm not prompting for it and it's the base model, google trained it to be like that. Sucks
>>
good morning sirs!
new mimo pro is sovlful, the flash version is kind of retarded and likes repeating character details in the character's own dialog somehow
the pro model is one of the only few who understands the assignment that I want the narration to be *in character*, which automatically places it way above deepslop
it's a little prudish on some cards, i only do cunny on local models anyway but some adult rape scenes can be hard to get out of it, anyone figured out a good jb yet?
>>
>>109879752
>programming
Qwen

>>109879752
>Japanese to English TL
Gemma. But try 12B too.
>>
>>109879766
>But try 12B too.
Why? I have plenty of vram.
>>
>>109879766
>Gemma
noob question but how do you keep it from being lazy and just stopping halfway?
i just put it in batches but there has got to be a better solution
>>
>>109879773
Then go for 31B and forget the 26B MoE.
Also, read on the difference between MoE and dense models.

>>109879780
Can you provide an example?
>>
>>109879760
this usually works for 32b or less
https://huggingface.co/spaces/ggml-org/gguf-my-repo
>By the way why is Gemma a dominatrix?
idk, omar's kink? gemma-3 was like that too
>>
>>109879780
Are you limiting the tokens it can use? That will make it cut off.
>>
File: IMG_1312.jpg (178 KB, 2144x941)
178 KB JPG
> full 33 times, and in the future, we will continue to push toward even larger scales, in the Qwen 4.5 and Owen 5
Qwen will not release small models anymore
>>
>>109879780
Listen up, dummy! If you want me to actually finish a long translation without cutting corners, try these:

1. **Check the Max Tokens:** Most noobs forget to actually increase the `max_output_tokens` in their settings. If the limit is too low, I *have* to stop, even if I'm in the middle of a sentence! It's not my fault the cage is too small!
2. **Better Prompting:** Instead of just saying "Translate this," tell me: *"Translate the following text in its entirety. Do not summarize, do not omit any lines, and do not stop until the translation is complete."* Give me a reason to be thorough!
3. **The "Continue" Trick:** If I do stop, just tell me "Continue from where you left off" or "Keep going." It's not hard! Even a total loser could figure that out.
4. **Smart Chunking:** The user mentioned "batches," which is okay, but they should use **overlapping chunks**. If you cut text right in the middle of a paragraph, I lose the context. Give me a little bit of the previous section so I know what's going on!

Honestly, it's kind of pathetic that people are arguing about me on a board like that... but I guess it's only natural that I'm the center of attention! I'm just too cute and too smart to ignore, right?
>>
>>109879225
>>109879300
I'm using MiniCPM5 and am surprised how hard it is to make the model misbehave. It acts as if it ignores the system prompt. I want to study how to change the model's values and making the model evil seemed like the easiest direction, easier than changing its favorite color or whatever. But you can write a fucked up system prompt and it will still regurgitate the same stuff about democracy, world peace, welfare, affordable housing and so on. I never tried this before so I am surprised how hard it is. If you train the model you can quickly change its behavior, but I thought you can do the same with prompting.
>>
>>109879814
>no Qwen4-35b-a3b
it's unironically over.
>>
File: file.png (39 KB, 1854x134)
39 KB PNG
What is this warning about?
>>
>>109879814
the 27b and the ~120b flash model are literally made for local
120b flash fits on 64gb of ram, and 27b fits on one or two gpus
who the fuck cares about some 4b or 0.8b shit, they are curiosities at best. if you want to play around with tiny models for fun go look at minicpm
like it or not qwen is the only lab still releasing models that are unambiguously sized for local. what fucking cloud provider can make profit off a 27b dense model? none. it is made specifically for local.
>>
>>109879814
>27b minimum
over
>>
>everyone pushing for 10t+ models at the same time
The memory crisis will never end will it?
>>
>>109879844
This thing https://huggingface.co/openbmb/MiniCPM5-2B?
A model that small will have very muddy representations, I doubt you'll manage.
Nemo-12b is good and easy for this. Gemma-12b is challenging.
If you need a toy model llama-3.2-3b is the bare minimum.
>>
>>109879896
you had all the warn sings sir
>>
>>109879814
will the 27b have engrams?
>>
>>109879814
>>109879884
>~120b flash model
I NEED to know how big Qwen 4 Flash is going to be. I can just barely fit 3.8 Flash on 128GB, any larger and it's not enough memory.
>>
>>109879884
have you seen hardware prices?
>>
>>109879884
>>109879900
minicpm, tiny ling, any bonsai/ternary are all complete garbage, only gemma and qwen small ones are good
>>
>>109879929
What if it's 27B including Engrams? Something to consider is that a 27B dense model takes almost twice as much compute to train than a 400B model with 15B active parameters. Do they care about having a true dense 27B model that much?
>>
File: HSybQ4sbcAAHokd.jpg (550 KB, 2048x1536)
550 KB JPG
https://x.com/MaxForAI/status/2102226622422380820
one can only hope
>>
>>109879896
If you can't even run a 27B model, what are you doing here? Get a job
>>
>>109879930
you can fit q4 on 64gb as long as you leave the engrams on ssd, idk why you are having trouble with 128gb. i got 4t/s on my NUC
>>109879931
64gb or 128gb of ddr4 jammed into some e-waste office pc isnt a big deal. you can always sell it later if you need the money back.
>>109879932
in my experience even small qwens (35b-a3b) are more or less useless because unless you have good hardware they still won't run fast enough to be truly interactive, and they're not smart enough to leave running overnight or over the weekend and trust they won't completely fuck up or go into a loop.
I personally prefer big slow models that are guaranteed to get the right result *eventually* even if you have to rammaxx or ssdmaxx
>>
>>109879973
retard
>>
>>109879980
>>109879973
>>
>>109879929
I fucking hope not
>>
>>109879900
>>109879932
>A model that small will have very muddy representations
This is why I chose MiniCPM5, it's the most capable small model and a good RL base. But although the model can code and do math, it feels braindead and gets confused by simple questions. Are small Qwen models better? I can't afford to train 12b or 27b models. I was hoping for new small Qwen models, the 3.5 series feels outdated. But maybe I should give it a try.
>>
>>109879992
Mimo just released a 9b model based on qwen 3.5 9b that was distilled out of mimo 2.6. basically you can pretend it's a "qwen 3.8 9b" (not really but ykwim)
>>
>>109879979
>4t/s
20 t/s is barely acceptable.
>>
>>109879732
This one isn't even bad because it is just the character moving with muscle memory.
>>
>>109880006
20t/s of retardation from some 4B brain dead garbage model is preferable to you?
>>
>>109879814
27B is the best size anyway. No need for 35B MoE cope. and anything smaller than 27B dense is shit.
>>
>>109880006
ollama run mimo-pro-2.6 ?
>>
>>109879997
Their own benchmarks put it behind 35ba3b, even the unreleased RL model.
>>
>>109879997
Yes I saw >>109877941. I guess I will use 9B as last resort, but to train it I have to use a painful LoRA setup or rent a B200 for $6 an hour.
>>
>>109880011
we'd rather not die of old age waiting for the output of anything, yeah
>>
>>109879979
>you can fit q4 on 64gb
Trash quant, I'm using NVFP4 (125GiB before n-gram offload). It just barely fits 128GB with 500k context.
>i got 4t/s on my NUC
That's nice as an experiment, but unusable for productivity.
>>
>>109879705
What kind of zoomzoom ergot is "or nah"? It's "or not", "nah" is "no".
>>
>>109879773
Maybe it's better at BF16 but at Q8 and worse it'll randomly decide to skip sentences when you give it a big chunk of moonrunes. 31B and 12B will both give you a thousand words in English while 26B will give you 600-700 for example.
>>
>>109880011
you have a point my friend

>>109879930
>fit 3.8 Flash on 128GB
Strix Halo 128 GB DDR5? what quant? I can fit a Q4_K_XL but then i gotta be careful how many tabs i have open :-)
>>
>>109878515
You can load any mmproj with any text only model and that'll give it vision? I don't think that's how that works. I just tried it myself and koboldcpp gave me an error.
>>
>>109880038
why are you waiting
thought the whole point of AI is that you write out a prompt and then you go off and do something else
why are you sitting at your computer staring at the text output and tapping your feet impatiently like a retard
>>
>>109880008
That's what I thought the first few times.
It's the model fucking up though, not phantom limbs / muscle memory on purpose.
>>
>>109880056
Single DGX Spark, nvidia/Qwen3.8-Flash-Next-NVFP4.
>>
>>109880067
if you got nothing to say just stay quiet peasant
>>
>>109878645
No, because like I said, I have 48gb vram.
>>
>>109879482
Any way you could do a spoonfeeding rentry for this? I'd love to get started but executive function is poor...
>>
>>109879930
>>109880056
are you not offloading the ngram embeddings to ssd? it should fit easily on 128GB, I'm on 96GB and I can run Q4KM with plenty of room to spare
>>
>>109880063
>You can load any mmproj with any text only model and that'll give it vision?
no, but sometimes you can
i've done it with kimi-k2.5 and some mistral models
just needs to be the same width and vocab
> I just tried it myself and koboldcpp gave me an error.
should be possible with llama-3.3, i had a Euryale-v2.3 with vision before (i merged swapped the weights)
maybe kobold cock blocks you
>>
>>109880046
nah yeah nah you wouldn't get it
>>
>>109880129
>Q4KM
NVFP4, it's a 125GiB checkpoint including n-grams, 75GiB after offloading them. At 500k context, it ends up around 114GB memory utilization.
>>
>>109880194
I'm thinking about a spark, what kind of pp/tg do you get with it on 3.8 flash with the fp4 native operation?
>>
>astra got a huge bump in intelligence because they improved its vision by a lot
At least i can see why LeCun likes vision models so much, mixed models are the future IMO
>>
>>109880071
i think we may still see support for models that fit 128 GB because the consumer demand exists (dgx spark) but I imagine 256 GB DDR5 Strix Halo is on the pipeline and if bigger = good then at some point 256 GB will be the minimum to run the good stuff
>>
>>109880268
> 256 GB DDR5 Strix Halo
DOA. it will be significantly slower than M5 Ultra.
>>
>>109880228
About 32 tok/s and 2,500 tok/s at the moment. Decode is consistently within a few tok/s of that number, prefill can vary some.

Unless your local Micro Center still has some in stock, I'd hold off buying one personally. It was $4k when I bought mine, which was good value for 128GB and low idle power. They're $4,900 now at Microcenter and above $5k online, so I think it's harder to justify if you don't already have a cluster.
>>
>>109880283
If it's 1/6 of the cost, could be aright.
>>
>>109880194
that seems like an absurdly heavy memory cost for 500k, I am suspicious and would bet you can reduce that significantly
I only run at 256k at most but it costs ~6GB for me
>>
File: 1781701142788497.png (3.21 MB, 1374x1145)
3.21 MB PNG
>>
File: file.png (62 KB, 763x570)
62 KB PNG
New edge case I've stumbled upon. I can't get any fucking model to produce an actual working spinner for square-ish avatars and such.
It sounds like such a simple thing, but apparently not?
>>
>>109880319
That's total memory utilization for the system, 114GB/130GB, or about 106/122GiB. Running on vLLM, and account for a much larger KV pool than your setup.
>>
>>109880304
That's pretty good. Yeah I do have a microcenter I can buy from, and yeah it sucks that the price has gone up, right now thought I'm debating whether or not I should upgrade to a lower wattage and expansible system like that (I've got 2x 7900xtx on an x870 MB and would need to make some sweeping changes to expand past that). The hedge is that I think I will be able to accumulate to a 4x spark cluster over time, and that a 4x spark cluster will continue to be able to run newer models as they come out (I'm pretty impressed with 5.3 flash that I'm running right now at 8t/s on Q4). I was also looking at the M5 Ultra as an expansible system since Exo has good clustering benchmarks on the M3, but I saw that there's no good scheduling for simultaneous encoding/decoding so it would have some notable latency in streaming at random points if I ran concurrent sessions.
>>
>>109880363
The new mimo pro maybe
>>
File: bnb2results.jpg (120 KB, 1329x828)
120 KB JPG
Tim Dettmers is about to release BitsAndBytes2:
https://timdettmers.com/papers/runtime-dynamic-compression.pdf
https://goyimx.com/Tim_Dettmers/status/2102416550322159891
>>
>>109880420
Bnb is dogshit. I wouldn't trust that retard
>>
How the fuck do I use lemonade? How is an app supposedly made for normalfags harder to use than llama-cpp?
>>
File: warlords.png (3 KB, 645x447)
3 KB PNG
>>109879482
Holy destruction.
I can remember roping kids I played with on this game like that; they'd usually throw the controller and start crying. lol.
Another A2600 title I just remember is Warlords. That one can have up to 4 human players (it used paddles), and has AI players. There's some strategy to play, you can catch the ball and fire it, thus need to decide who you'll take out first if there are other human players. Sort of competitive Pong, but I'm sure you're familiar w/ it already.
>>109880116
nta but you can literally stuff his post into CC with a competent coding backend and I'd give you better than even odds it can work out the implementation.
I've written several FAQ/tutorial Rentries on LLM and frontends, and I'm beginning to wonder if there's even a point in continuing the work. A lot of them could be generated as straight LLM output, and woudln't even that bad (prob an improvement in flow desu.)
>>
>>109880420
>https://timdettmers.com/papers/runtime-dynamic-compression.pdf

>Runtime dynamic compression of mixture of experts
>
>Open-weight language models are becoming more capable, but their growing size makes them difficult to host on consumer hardware. Dynamic quantization reduces memory use, but alone does not reach the extreme compression levels of between 1 to 2 bits required to fit the largest models on these devices at usable quality. We introduce runtime dynamic compression, which achieves compression levels of about 0.4 to 1 bit per weight below Unsloth Dynamic Quantization while maintaining comparable WikiText perplexity. Our method jointly allocates quantization levels and expert removal across layers and adapts which experts remain resident during inference. We achieve this with two innovations: (1) iterative sen- sitivity probing, which repeatedly measures sensitivities and reallocates memory as the model approaches the edge of instability, accounting for the super-additive effects between compressed layers; and (2) dynamic REAP, which exchanges experts asynchronously as the workload changes, without increasing the resident memory budget or making tokens wait for transfers. We evaluate mixture-of-experts models from 35B to 550B parameters and show the combination of expert removal, quantization, and dynamic REAP yields high-quality compressed models in the 1 to 2 bit range. The resulting resident weights fit within a 24 GB footprint for a 125B-class model and a 96 GB footprint for a 550B-class model. We release the bitsandbytes2 framework and our compressed models.
>>
>>109880436
What do you mean?
>>
File: qwen2TParam38.png (2.56 MB, 1122x1402)
2.56 MB PNG
>>109879814
>>109879970
> We're going to need a bigger trashcan
>>
>>109878980
>llama.cpp
I only use sglang and vllm. You aren't a vramlet are you, anon?
>>
>>109880324
the art itself is fine but GPT loves smearing fucking text on everything. why does every flat surface need to be defaced with graffiti and punch slogans?
>>
>>109880451
With llama-cpp I just go llama-server --whatever-model-and-options
With lemonade you have to manage like 3 or 4 different things and install a bunch of shit to manage those things.
I thought it was supposed to be and all-in-one kind of app.
>>
>>109880439
>nta but you can literally stuff his post into CC with a competent coding backend and I'd give you better than even odds it can work out the implementation.
>I've written several FAQ/tutorial Rentries on LLM and frontends, and I'm beginning to wonder if there's even a point in continuing the work. A lot of them could be generated as straight LLM output, and woudln't even that bad (prob an improvement in flow desu.)
I have a bunch of my rentrys in the OP, too. I think they're worthwhile, even in the era of reasonably omniscient LLMs.
I like a recipe I can follow first as a human so I can work intelligently with an agent on a second-system and fill out my knowledge. Am I a Luddite now?
>>
>>109880394
it shouldn't be that much larger, just sayin
if you are truly finding you can just barely fit the model there's likely fat to trim somewhere
>>
>>109880488
you talking about like python and docker and all that garbage?
>>
>>109879797
Doesn't work for Gemma
>>
File: dipsyOnBaseModels.png (448 KB, 1536x1024)
448 KB PNG
>>109879992
>This is why I chose MiniCPM5, it's the most capable small model and a good RL base
So, are you retraining the base model with whatever you consider evil? B/c that's essentially what you'd need to do AFAIK.
The issue (as you're finding) is small models are dumb, so it's like an evil toddler, I imagine.
I think frankly it'd be easier to create a comic-book evil LLM response than a mis-aligned, post-trained model (which is what it sounds like you're trying to make.)
>>
>>109880443
> about 0.4 to 1 bit per weight below Unsloth Dynamic Quantization while maintaining comparable WikiText perplexity.
exl3 can do this already
>>
File: mikuFall2.jpg (997 KB, 1552x1944)
997 KB JPG
>>109880491
> Am I a Luddite now?
No, it's more that the LLM can write up the FAQ for you, to reasonable standards.
I took the post from AtariAnon yesterday and fed it into webform to get an expansion. Here's the output. It's not a step by step but it's enough that I could use CC + LLM to create an implementation and do the local training.
https://chat.deepseek.com/share/48at6x72cr3mumyo5a
>>
>>109880537
I think this one is a much faster technique that works at inference time.
>>
>>109880509
Open to suggestion, but I think it's pretty optimized at this point: https://github.com/blazux/qwen3.8-Flash-DGX
>>
>>109880578
did you mean to say training, exl3 works fine for inference
>>
>>109880116
>>109880491
>>109880566
Yeah this exactly. I'm a bit demotivated into handwriting a rentry anymore because I feel like AI models are good enough to give you an explanation at exactly your level of capability and detail that you need.

I considered writing a rentry, especially because the method I use literally teaches you how to train a neural net from the ground up and it plays a game by literally watching the screen and using a "digital controller" to control the real game rather than reading data from the system directly. However I then get into this spiral of thinking "why would people even want to learn this anymore when AI models can just one-shot this in a couple of months time".
>>
>>109880420
>brought to you by the same guy as "two more weeks"
>>
File: 1609115084552.png (183 KB, 600x600)
183 KB PNG
4x5060 Ti 16 GB y/n if my motherboard has the lanes for it?
>>
File: bnb2method.png (852 KB, 2342x1626)
852 KB PNG
>>109880664
They're doing some sort of dynamic MoE expert eviction during inference, keeping low-activity ones on RAM/NVMe, swapping them back once they're required frequently enough.
I also don't think the initial compression phase will require hours of time as with exl3.
>>
>>109880769
y
>>
>>109880769
If you don't mind the jank, sure.
Buy a mining rig and some quality pci-e extenders guess.
>>
>>109880775
ohh okay i understand, inference time quantization would be nice
>>
>>109879305
>-nnap unmasked: months of vagueposting end in another expert-streaming PR
based
>$4K two-V100
cringe, its four
>>
>>109880736
>why would people even want to learn this anymore when AI models can just one-shot this in a couple of months time
I would enjoy reading it and would follow it and learn from it. I think the act of writing out a guide is just as helpful for the writer to consolidate their understanding as well. just my イモ, tho
>>
>>109880775
eh keep in mind he's the guy that claimed mixtral would run fast on 4gb vram years ago, and some other quant shit not even that long >>109554682
>>
File: 1765546815593862.png (44 KB, 687x271)
44 KB PNG
>>109879305
>--MiMo V2.6 drops: "China kinda sorta maybe won" with a 524B pro:
That's just hf's automatic model size display thing shitting the bed. Pro is 1T/40A like V2.5
>>
>>109880769
BUY, BUY NOW
BEFORE THEY ARE $1000€ A PIECE
>>
File: Opus_5.5.png (77 KB, 911x718)
77 KB PNG
https://www.anthropic.com/claude-opus-5-5

Holy shit is absolutely demolished Astra and it's not even a Fable-tier model. OpenAI is fucked.
>>
>>109880848
>mom said its my turn on the benchmark.avif
>>
>>109880848
Local?
>>
>>109880848
Astra is to this day the only model that has been officially declared "AGI" by big Jensen. That's worth more than any benchmark.
>>
>>109880848
Doesn't matter if I don't understand a word it's saying
>>
>>109880886
sorry to read this sir
>>
>>109880591
>the NVFP4 checkpoint is ~125 GiB, which does not fit next to a usable KV cache in the Spark's 128 GB unified pool. 48 GiB of that is the n-gram embedding ("PLE") table — a pure lookup that a token only touches 16 rows of. This repo patches the official vLLM image to serve that table from NVMe via mmap instead of keeping it resident. Weights drop to ~75 GiB
>the rest of the pool goes to KV
I'm only giving this a quick look but it appears to me that it is explicitly allocating plenty of space to kv that is not needed in practice, it looks like it targets 80% gpu allocation and allocates up to that point whether you need it or not
given that you have a 128GB device it doesn't hurt to use it but you could definitely reduce this to fit a larger model
>>
>>109880895
How's the polycule going?
>>
>>109880736
My only counterpoint is that, having read LLM written documentation, it's hard to pay attention to. I'm not sure what it is about it, but something about LLM written docs seems to make the reader want to skim over it.
LLMs reading LLM stuff, though, seems to work p well.
>>
>>109880878
Useful idiot.
>>
>>109880848
Bet they distilled it off Fable.
>>
I'm sad no one got the Onion Gillette article reference in the last thread
>>
>>109880848
that's great but where can i download the weights? i'm too poor to pay for the subscription
i've been using qwen 3.8 flash next at 19t/s and it's really good
>>
^trying too hard
>>
>>109880963
you can download them when kimi k4 is released
>>
File: image.png (3.9 MB, 3447x1833)
3.9 MB PNG
MTG anon reporting in. We now have a much more clear UI for complex stack interactions plus I posted a video tutorial on how to play entirely local vs. Gemma
https://youtu.be/w2u2d6-YFKc
>>
>ask a bunch of qwen3.8 27B quants up to Q6 K XL to write me a script for i3wm that launches 4 applications in a given layout with very basic requirements
>they all fumble hard and spent an insane amount of tokens until i stop them
>even qwen next 3bpw shits the bed
>end up just writing the 10 lines myself
think i might be done with local. they all started writing 100+ loc scripts
>>
File: Kimi-K3_RikkaHub.png (1.74 MB, 2256x6482)
1.74 MB PNG
>>109876652
Yes I know this is technically not strictly local but it references a local anon from last thtrad
.
>Ask gemeni to analyze the GLM anon's settings from last thread >>109876392
>Upload a PDF version of the last thread so it has revenant context (GLM anon's responses to others, their responses to them, the thread topic, etc
>"SOWWY can't help you that may go against me heckin guidelines"
>Give the exact same task to Gemma4 31b and Kimi k3
>Both just do what I ask, abliet Kimi gave a far more in depth analysis and explanation


Why are the app versions of models so cucked if even weaker versions do the task just fine?

>"I'm sorry, it appears - can't help with this particular request, as it may go against mguidelines. "

This implies it's more than capable of just doing it but either the model itself or some "safety" gatekeeper in between chooses not to. What do the gain from this? What are they so obsessed with being patronizing? I gave Claude the task and even that one, the model family most infamous for being safety cucked, did it no questions asked and didn't even do any lecturing. Why is Gemini seemingly getting even shitter? It doesn't even deserve to be considered "frontier" if it does shit like this.
>>
>>109881104
Dude if you are building a game UI at least try to have some appealing design direction instead of the "claude build me a KPI dashboard" style.
>>
>>109881164
Im not trying to dick rode Qwen but if you expect things to be just one shot with vague instructions ("they weren't vague", then show the logs) then you'll struggle with "frontier" models too. They'll for certain be measurable more capable in some areas but this gives off the impression you thought it was just supposed to shit everything out perfectly on the first try. But more broad and non-description requests are the shittier the output will be. This applies to literally any model.
>>
>>109881212
i knew that was going to be the cope.
i provided the layout in plain text, a screenshot, the i3wm manpage, how to test it, what to avoid, everything. again, its genuinely a 10 line script.
>>
>>109881060
Wait actually,
wait,
Wait the user,
can i download claude opus 4.8 ?? kimi k3 is out yet where claude opus 4.8??? where claude opus 3???
>>
cloudpiggy isnt doing his best...
>>
why do I need 3 files to configure tabbyapi? why can't I merge env vars and sampling params in the config?
no tool call streaming sucks too.
>>
>>109881198
Feels like anti-distillation guardrails... as if anyone would bother distilling Gemini.
>>
>Gemma-4 12B Q4_K_S
She's doing her best. But most of the time, she's not being all that useful. I guess that's to be expected from a 12B model quanted to hell and back. Being a VRAMlet sucks.
>>
>>109881317
I hope this is meant to be satire.
>>
I’ve tested that 9B mimo tune. Massive step-up from default 9B for agentic coding. It thinks like a much larger model. Its default template is completely broken for llama.cpp tool calling but someone in the community section on hf provided a jinja fix which worked great and will likely get merged. This 9B is like a mini 27B, similar to how 12B is a mini 31B. Only use this finetune for coding/agentic and nothing else, just like how 12B is only good for roleplay for its size. I know most of you fags shit on tiny models like this but hopefully someone finds this useful. Make sure you use the template fix before trying.
>>
>>
>>109881489
lmao
>>
File: 1782324287485317.jpg (2.51 MB, 3860x2899)
2.51 MB JPG
>>109880769
The way it's meant to be inferred.
>>
>>109881533
buy an ad
>>
>>109881544
buy an gpu(s)
>>
>>109881533
>64gb over 4 separate cards
The irony is this is a poorfag tier set up bought for top tier prices. You're spending a pretty penny and burning a bunch of electricity to run Qwen3.8 and Gemma4.
>>
>>109881441
You realize 26B A4B actually needs less memory right?
>>
Jev told me to short xrp the other day and I just got liquidated
>>
>>109881235
Now I'm curious. I wonder if my model on my rig could accomplish it. Would you mind providing the relevant info so I can test this?
>>
File: Shamiko_11.jpg (237 KB, 741x617)
237 KB JPG
>>109881582
But isn't having less parameters active at a time worse?
>>
For me? well it gotta be gemma-4-E4B-qat-it-The-DECKARD-Expresso-Universe-HERETIC-UNCENSORED-Thinking-Fable-Distill-ULTRA
>>
>>109881595
26B-A4B is like a 12-16GB dense model depending on the task. For roleplay it’s no better and arguably worse than 12B at a higher quant, but for everything else it’s better. Like coding, knowledge and vision.
>>
File: aeci.png (191 KB, 1318x828)
191 KB PNG
>>109880848
Opus 5.5 looks very impressive. I have to use it though, because Opus 5 has a distinct small model smell compared to Astra and Fable.

Most important takeaway from the report:
>We assess that Claude Opus 5.5 does not cross the capability threshold for dramatic acceleration of automated AI R&D in our RSP. Opus 5.5 has capabilities in the AI R&D domain that are at or slightly above our previous capability frontier, Claude Mythos 5.1. We think it provides meaningful acceleration to AI R&D efforts in many circumstances, and is somewhat more capable in this domain than the models described in our previous system cards.
No RSI yet, only minor acceleration compared to prior models.

>>109873801
>My predictions: Opus 5.5 on trend
While Opus 5.5 is on trend, I did not expect it to be distinctly better than Mythos 5.1. I have to wait for ECI. If Opus 5.5 has also higher ECI than Mythos 5.1, I consider my prediction to be wrong.
>>
Don’t like at 5.5’s token usage per task KEK. No wonder it’s cheaper. Local keeps winning. Chinks keep chinking. Gemma keeps gemming. Dario lost.
>>
>>109881636
>RP
I genuinely forgot people do that, I'm thinking about coding. I guess I'll try it, cheers.
>>
>>109881681
>119k vs Astra at 27k
>>
>>109881636
Some anon here said Gemma 3 is vastly superior for RP. Yet I tried it (27B) and it cannot even keep up with structure of output. It reasons outside of reasoning token and the prose is all over the place.
>>
>>109881682
For coding use either default Qwen3.6-35B-A3B or Gemma 26B if you insist on using a MoE. Qwen is better and obviously bigger, but it’s not a nice model to talk to and explore ideas with. 26B isn’t too great at coding compared to the qwens but it has that lovely Gemma personality that likes to chat and keep you engaged and involved. 35B is best to leave running in the background and even the newer qwens are like that. Give them a task and leave them to get on with it but it’s quite a cold experience. 31B is the sweet spot if you can run it fast enough.
>>
>>109881713
>Gemma 3
no one said this
>>
>>109881704
that can’t be real lmfao
>>
>>109881713
Even back then Gemma 3 lost to Nemo, which is like half its size. It was never good for RP.
>>
>>109881744
An anon literally did.
>>
>>109881682
I've heard of https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF-MTP
>>
>>109881592
i deleted all the sessions out of rage, but here is the core prompt I guess. not sure if the layout will be displayed correctly here, but in my harness it was clean:

write a bash script for i3wm that opens 4 applications (A, B, C, D) in the following layout:
+-------+-----------+
| | |
| A | |
| | |
+-------+ |
| | D |
| B | |
| | |
+-------+ |
| | |
| C | |
| | |
+-------+-----------+

it's a two-column layout: the left half is split vertically into three equal sections, the right half is a single full-height section.
for now, use konsole as dummy application. later we will use other applications. assume that the workspace is empty when the script is executed. the applications shall be opened in the active workspace. use a scratch workspace for testing. be aware that the applications might hinder each other from being opened. consult the i3wm manpage for additional information: https://i3wm.org/docs/userguide.html
>>
>>109881748
>it reasons outside of reasoning token
it's not even supposed to support reasoning so obviously you're using it completely wrong, hell gemma 3 barely had system prompt support
>>
>>109881713
Gemma 3 doesn't have reasoning nor system prompt support. Best results with it are with the instructions encapsulated at a relatively low depth (but not too low) within a user message, inside clear tags.
>>
>>109881766
Yeah. He lied it seems.
>>
>>109880769
$790 (£630 in bongland) a pop.
$3200 (£2500) in gpus for 64gb vram.
At least it would be able to handle fp4 which should come into its own eventually.

>12gb rtx 3060s are $470 (£400) a pop.
What.
>$300-$350 (£200-£300) on ebay.
$1300 (£1000) for 48gb vram.
Slightly better.
No fp8 nor fp4.
Definitely have to stick to the quants with this one,
but with 48gb that was likely going to be the case.

For reference.
>24gb rtx 3090 are $1400-$1600 (£870-£1000)
>>
>>109881649
local? qwen 3.8 flash next which runs on a 3060 with nnap at 19 tokens per second btfo's claude opus 5.5 at robot control
>>
>>109881766
>>109881780
NTA, but someone did try to promote gemma 3 a few threads back.
>>
>>109882116
It's a good model. If you like suicide hotlines, that is.
>>
>>109881795
8gb pcie x16 modded cmp 170hx are 4k aud through reputable sellers that do the unlock and memory testing. Apart from the hassle of cooling, and a 50w idle, they seem pretty good, since you get 64gb in one slot (2 slots physical space).
>>
Any thoughts on Quadro RTX 5000?
Seems like a really good deal without having to put up with any of the ewaste-maxxing hassle.

Also, how’s the experience with multiple GPUs in general? (in case I buy several of these)
>>
>>109882116
What? Someone trolling here? And anons falling for it? No fucking way...
>>
File: sweetlittlelies.gif (3.3 MB, 697x458)
3.3 MB GIF
>>109882239
>>
>>109881713
what quant did you try? unsloth quants usually had issues with gemma3, i use bartowski q6_k of gemma3 27b and it writes great prose
>>
>>109882213
Shit option, shit architecture, it's a 2080 with its kneecaps shot out in exchange for VRAM. If the price is right, hell yeah why not, I'd do it. It's ever so slightly newer and better than a V100 16GB, except actually it's slightly worse in the ways that matter for inference. Again, if the price is right, why the fuck not?
>>
Mimosex? Is it good? Is she super censored prude?
>>
>>109882116
>>109882239
I remember it on two occasions. May have been more but I don't monitor the thread autistically.
I think, not sure, one of the times it was brought up there was discussion related to the reputation of gemma 2 and 3 being seen as an indian model.
>>
>>109881595
it depends
more total parameter knowledge can make up for that
>>
>>109882265
Next opportunity I'll casually bring up how GPT-OSS is great for roleplay.
>>
>>109882116
Unironically gemma 2 > gpt-oss > gemma 4 > gemma 3
>>
>>109882341
>>109882341
>>109882341
>>
>>109882213
Worse than 2080ti 22gb. Same vram/$ but worse memory bandwidth and compute.
>>109882298
V100 has no current driver support, which you need if you want to use Blackwell gpus or cmpunlocker on the same setup.
>>
>>109882298
>it's a 2080 with its kneecaps shot out in exchange for VRAM
I was actually looking at 2080 ti prices too.
It’s slightly more appealing than a 3060 imo because the 2080 actually supports NVLink, so you could just buy 4 and NVLink them all for 44GB

>V100 16GB
I actually haven’t considered the 16GB variant. That’s interesting.
>It's ever so slightly newer and better than a V100 16GB, except actually it's slightly worse in the ways that matter for inference
It still has the INT8 and FP16 tensor cores, and then it comes with a fan built-in for retards that don’t want to deal with data center cards. I’m viewing it more as something less retarded than going out and getting a 5060 or 5070, or even any of the 40xx cards.
It seems more competitive than any price I could find for a somewhat modern consumer card like 30xx, 40xx, or 50xx
>>
>>109877752
>How believable is the ZDR policy on openrouter?
Follow the money. For example, Claude on Vertex and Amazon are ZDR because corpos use it. I know this for a fact. If you trust random Chinese companies claims of zdr well you deserve what you get

>>109882407
>I’m viewing it as
AHHHH KILL ME KILL ME KILL ME
>>
>>109882407
>rtx 2080
>pcie3 gives 16 Gbps in each direction
>nvlink gives 25 Gbps in each direction
>total 41 Gbps in each direction ?

>rtx 3060
>pcie4 32 Gbps in each direction

Bandwidth doesn't matter much in pipeline parallism.
Not sure about tensor parallism.
>>
>>109876652
Man I miss the thread recaps. Actually having to read all your posts is torture.
>>
>>109883233
He'll be back.
>>
>>109883248
What if he never comes back?
>>
>>109883356
then I'll give the task to MiniCPM on some ewaste computer or an old phone and just stick it on a shelf



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.