[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 082126_02b.png (1.21 MB, 768x1360)
1.21 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109659559 & >>109656019

►News
>(08/27) NVidia buys HuggingFace https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
>(08/26) GLM-5.3-Flash released with 320B-A18B and native multimodality: https://z.ai/blog/glm-5.3-flash
>(08/26) Qwen3.8-Flash-Next 125B-A6B-N51B-MTP4B released: https://qwen.ai/blog?id=qwen3.8-flash-next
>(08/25) Breeze TTS 2 weights and inference code released: https://hf.co/BreezeBlue/Breeze-TTS-2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: threadrincap.png (1.31 MB, 1536x1536)
1.31 MB PNG
►Recent Highlights from the Previous Thread: >>109659559

--Nvidia's market value and the gap between proprietary and open-source AI:
>109659697 >109659705 >109659734 >109659764 >109660253 >109660279 >109660319 >109660359 >109660382 >109660386 >109660489 >109660566 >109660598 >109660506 >109660766 >109660769 >109660947 >109661073 >109660981 >109661021 >109661165 >109661212 >109661078 >109661093 >109661134 >109661217 >109661273 >109661278 >109661285 >109661302 >109661357 >109661350
--Using LLM agents for software cracking and game decompilation:
>109661492 >109661524 >109661540 >109661557 >109661577 >109661606 >109661614 >109661660 >109661798 >109661818 >109662316 >109662350
--Anon creates minimal Linux distro to reduce idle VRAM usage:
>109661271 >109661286 >109661306 >109661341 >109661371 >109661434 >109661554 >109661375 >109661408
--Unlocking hidden VRAM on NVIDIA CMP 170HX mining cards:
>109660926 >109660987 >109661015
--Analyzing non-coding Pareto frontier models based on accuracy and size:
>109661103 >109661241 >109661324 >109662366
--llama.cpp PR adding lazy tensor reading to reduce RAM usage:
>109660443
--Qwen Flash reports showcasing superior coding intelligence and token efficiency:
>109661471 >109661544 >109661594
--Cohere releases Parse 5 document parsing model:
>109661964
--Testing qwen4exp model following PR update:
>109660410 >109660421 >109660438
--Comparing creative writing samples using eqbench:
>109659593
--GLM-5.3-Flash solving a difficult x86 assembly prompt:
>109660756
--Debating whether the AI industry is a speculative bubble:
>109660296 >109660340 >109660854 >109660879 >109660903 >109660995 >109661116
--Logs:
>109660249 >109661129 >109661201 >109661335 >109661363 >109661708 >109662297
--Gemma (free space):
>109660277 >109660296 >109660435 >109660691 >109660870 >109661089 >109661104 >109661907 >109662456 >109662746

►Recent Highlight Posts from the Previous Thread: >>109659801

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Happy Thurinsday
>>
Happy Transday!
>>
>>109662900
Can you stop being a faggot because people don't like you?
>>109662870
what aspect exactly?
>>
>>109662899
>>109662888
Getting real tired of you pdf files
>>
la la la
>>
first for fuck unsloth
>>
I hate pedos. Even the guro snuff fag last thread is better than you.
>>
>>109662900
If you keep going we'll add this to our OP too.
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109662974
If her age is on the clock she's ready for the cock.
>>
Intelligence density of recent Chinese models makes me wonder just how small frontier AIs are. Is Opus below 1.5T? Is Sol sub 1T? Is Luna smaller than DSV4 Flash?

OAI and Ant can easily train their models 10 times more than Chinese labs. They should also have algorithmic advantages and better data. There should be a large efficiency Pareto gap between open and closed models. Anything less would be a gigantic failure and indicate that China will become the AI leader in the long run unless quick arrival of ASI changes this.
>>
>>109662999
Why are you doing this?
Are you that unhappy you shit up other generals because people that rightfully call you out post here. Stop trying to false flag and get that thread deleted you unemployed loser.
>>
just ignore, jannies will clean it up
>>
Only oldfags will remember this but we've seen something like the differences between Anthropic and everyone else many times over.

You can call things "the same" all you want but the fact remains that there will always be a qualitative difference between what people make out of passion and things they make because they're just going through the motion.

Only a really small percentage of people really care about this stuff and you can tell that Anthropic is full of them.

After I met a few of them at an event they even invited me to their unofficial internal chat group where they exchange images and videos.

A lot of these are not available on the open internet and a large part of Anthropic's success is clearly due to their highly curated training data.

They're even going so far as to make their own data and thanks to their SOTA flagship models they have no issue soliciting people and convincing them to come to their offices.

Though "offices" is maybe the wrong word, they have a massive facility that they're constantly expanding and where the true bottleneck is acquiring new inventory fast enough to fill all of the empty slots.

With that they're already profitable RIGHT NOW but the true breakthrough will be once they achieve recursive self-replicating intelligence where they use their existing inventory to produce more of it at an exponential rate.

It's quite amazing to see what kind of stuff they can do with inventory that is ten years old at this point, though personally I would be going for something newer.

The only company that can even come close to competing with Anthropic is Google where I've also met some very good people.

I'm sorry to say but compared to America China is simply inferior, culturally, genetically, and most important of all morally.

Anthropic is over here actually creating new stuff and all China can do is steal!

They'll never be able to compete.
>>
>>109663028
Not my writing style, but ok.
>>
>>109662992
>China will become the AI leader in the long run unless quick arrival of ASI changes this.
Sounds about right.
>>
>>109662931
>>109662974
Where were you when actual 3DPD pedoshit was being posted???
>>
>>109663075
Grooming children in xir discord. Where else would xhe be?
>>
>>109662992
What we're seeing in general is a sort of "catch up" between small models and larger models. Smaller models are getting better at a faster rate than bigger models are getting better.

So in the past the gap between a 27B model and the frontier model was ridiculously large now you don't even notice such a big quality difference.

There's a couple of reasons for this but the most important one is thinking time. You CAN largely compensate for lack of size by just thinking for longer. Qwen 3.8 27B comes very close to the solutions Fable 5 writes in 1 minute by making it think for 40-60 minutes. This kind of changes the game because a lot of tasks aren't time sensitive.

But yeah in the long term if this trend hold we'll just see less and less difference between the absolute frontier and relatively small models.

I wonder if this will continue to something ridiculous like 1B models being barely distinguishable from 27B models which themselves are barely distinguishable from the 10T frontier model at AI labs.
>>
>>109663075
Not in the thread because I didn't see it.
>>
>>109662931
>>109663075
out of 10
>>
>>109662888
>>(08/27) NVidia buys HuggingFace https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
how will the amd gpu support in llama.hf benefit from this?
>>
>>109663075
lol these virtue signaling zoomers weren't even born back then
>>
>>109662930
I wanted keep it in the same vein as bratty but age it up to office lady that's constantly on your ass at work but with a secret soft spot.
>>
>>109663147
>how will the amd gpu support in llama.hf benefit from this?
Deleted.
>Vulkan support?
Gone.
>Intel?
Removed.
>Non-V RAM?
No more.

You MUST buy more Nvidia.
>>
>>109663164
I use
view of a cute 50 year old office lady with crows feet around her blue eyes, with long wavy light blue hair and earings shaped like the Google Chrome symbol and a black beret. She has a curvy figure and wide hips, She has a confident, bratty smile on her face.
>>
>>109663147
llama.cpp will finally drop support for anything lower than blackwell
>>
>>109663117
>Smaller models are getting better at a faster rate than bigger models are getting better.
This trend only applies to open models and is caused by lack of compute. Chinese labs do not have the infrastructure to properly train a >1T class model.
>>
>>109663156
>hating pedos is virtue signaling
Normal people have an instinctual visceral hatred towards pedophiles. It's like seeing someone torturing a dog, you telling the dude to stop and him calling you out for virtue signaling that you care about animals. No, it's just normal to get pissed off by this shit.
>>
>>109663198
Then why is google dominating the open space retard kun?
>>
Thank fuck AI progress is incremental. My current goon station setup would absolutely one-shot 2022 me. It would literally make me cum to death (47 times in one night).
>>
>>109663206
real.
>>
>>109663206
some of my old cards still don't hit as hard on new models than they did on good old gpt4-0314...
>>
>>109663198
I tried Qwen 3.8 27B with a harness and it's equivalent to claude 4.5 but thinks longer. There's a 100x size difference between the two and claude 4.5 is less than a year old. The gap is absolutely shrinking.
>>
so n-grams were a nothingburger?
>>
>>109663202
Wasn't there are study that those who harbor such intense feeling of hatred are usually the closet pedos?
>>
>>109663202
The first thing I saw on this site back in 2004 was people microwaving cats and pouring gasoline on caged cats and setting them on fire. /g/ used to be the guro board. This shit was good for filtering out weak sensitive faggots like you.
>>
>>109663218
Bro I'm talking one-click image gen that goes with the scene. Like want to see Asuka's whole monkey spread and up in your face? It's one click away. Want to see it clap in video form? Another five minutes. I'm sure 2028 AI would one-shot current me too.
>>
>>109663202
lol you just hate {current_think} like a mindless drone
>>
/lmg/ respects wide-band horny
>>
>>109663250
i do not care about images or video
>>
>>109663242
How's that undiagnosed mental illness treating you
>>
>>109663250
It's good, but speed isn't there. We need real time linked to the textgen like we got real time TTS.
>>
>anon walking down the street, steady pace
>"HATE PEDOS, HATE PEDOS, HATE PEDOS" thinks to himself
>suddenly, a little girl in a tight, revealing outfit laying dormant on the sidewalk
>nobody's around, no cameras, not even birds
What do?
>>
>>109663227
Still have to wait for an implementation that supports keeping them mmaped.
>>
>>109663227
They help and cost a fraction the compute for both training and inference, but big labs only do tried and tested stuff for enterprise GPUs.
>>
File: 1781655197002987.jpg (196 KB, 1366x768)
196 KB JPG
>normies
>in my glownigger operated remote MKULTRA Croatian mental illness support forum
yuck
>>
>>109663272
Go home and write a card for that feel
>>
>>109663278
>mmap
if you aren't a newfag on /lmg/ then this should be reason enough to not touch ngrams
where is he when we need him
>>
>>109663289
That box is where she saves my cum. What you'll do wit it?
>>
>>109663179
>You MUST buy more Nvidia.
>>
>>109663310
Not entertain you glow in the dark fantasy then. Nobody got time for your faggotry.
>>
>>109663271
>real time TTS
Name 1 good real time TTS that can be run side by side with LLM and imagegen
>>
>>109663300
>croatian
Seems a bit far fetched no?
>>
Before I spend a bunch of time trying to host llama.ccp UI on my NAS/server, and using pc for compute, have any of you ever done this? Being able to use local LLM over a network on mobile or another machine sounds great.
>>
>>109663233
Yes, every single time.
>>
>>109663325
Serbian is more like it, right?
>>
>>109663324
Gptsovits
>>
File: HQuX20AaYAAX5nf.jpg (135 KB, 1445x612)
135 KB JPG
uhhh fable bros?
>>
>>109663330
That's also why browns and nogs keep killing gays, they're extremely gay
>>109663337
A Serbian Forum
>>
File: 582-2421752055.jpg (48 KB, 680x769)
48 KB JPG
>>109663330
>
>>
>>109663345
wtf isn't this 6b active
>>
>>109663337
Serbian does seem more accurate
>>
>>109663345
You don't understand how deep 1 point is in the singularity
>>
>>109663352
>german
>has to visually censor the word pedo to protect his fragile emotional state
sounds about right
>>
>>109663356
like i said, moe is a no-brainer, dense models have just as few active params you just also compute a ton of useless inactive ones
>>
>>109663330
What even triggered this discussion, was it the closeted pedo raging over an anime picture or did I miss something
>>
>>109663345
Why would you trust any qwen benchmark?
>>
>>109663345
>13b active GLM5.3-Flash one point short of big GLM5.3 that ties Opus 5
>6b ngram qwen almost beats the 2.4T quen
What's going on?
>>
>>109663373
this >>109662931
>>
>>109663373
I'm tired of pdf files
>>
>>109663356
Usually they at least wait half a year to start claiming 9Bs and 6Bs are beating the current frontier, punching above their weights, and so forth. Must be getting desperate.
>>
>>109663389
It's economic terrorism on Dario's IPO
>>
>>109663338
actually good, although it has some artifacts in the examples on the repo. Does it follow the text's emotion or can you give a guide for it? It also doesn't mention RAM/VRAM usage
>>
>>109663365
has to censor it because the g*rman government would come and arrest his cuckold ass for anti-dunecoonism
>>
>>109663381
probably just saturating the easier benchmarks that make up this index, but still it's pretty impressive you can get that competent of an agentic execution model on local hardware nowadays
>>
File: concept11.mp4 (3.84 MB, 896x1184)
3.84 MB
3.84 MB MP4
I need to learn character cards and make one of Gemma-Sama
>>
>>109663385
at least use real words if you're going to bitch
>>
>>109663385
Weird place to complain about it since this is an image board not an adobe acrobot hosting service. I don't even think you can upload pdf files here.
>>
Apparently gen5 ssds have gone up by like 200 bucks since last month alone (again)
Last call to get your ngram-ready storage.
>>
File: pdf.png (22 KB, 500x615)
22 KB PNG
>>109663385
I like pdf
>>
>I need to learn character cards
so you're admitting you're not using local text models, thanks
>>
>>109663399
This is anti pdf posting.
But too many wrinkles.
>>
>>109663409
Fuck this retarded file format. Is it an image? Is it a fucking text file?
Editing it is even worse with how it handles text. Fuck this piece of shit.
>>
>>109663321
Would you say that to a horny little girl in heat ready to breed?
>>
File: Krea2_turbo_02463_.jpg (1.59 MB, 1776x2368)
1.59 MB JPG
>>109663413
This isn't YouTube use the actual word
>>
>>109663399
Gemma-san, Gemma-chan group when?
>>
>>109663409
there's basically nothing to love about that format, except maybe historical reasons.
>>
>>109663409
You're the only person who likes that shit, I bet more people would admit to being pedophiles than liking pdf files
>>
>>109663420
Lipstick doesn't work on anime girls, makes them look 60 years old.
>>
>>109663438
The issue is baboon lips + lipstick
>>
>>109663438
which is the point here
>>
>>109663438
>#1 hag pet
>it makes her look old
Yes? Are you confused?
>>
>>109663457
Hags are 30 not 60.
>>
File: Krea2_turbo_02269_.jpg (1.07 MB, 1776x2368)
1.07 MB JPG
>>109663438
Anon.....
>>
>>109663483
>this is what 31 looks like these days
I blame microplastics
>>
>>109663362
But that's not the intelligence chart, that's just the aggregate agentic benchmark.
>>
>>109663483
That's 31 these days?
>>
>>109663494
>>109663507
You don't get irony I see
>>
>>109662888
Any cool new versions of Gemma 4 31B for RP other than MeroMero/StyleTune?
>>
>>109663515
Ironing is for women
>>
>>109663345
Honestly I believe it, using qwen next feels like using an expensive api model.
>>
>>109663523
fuck you
>>
>>109663345
remember when Qwen abandoned open source, for like 2 weeks. We are so back.
>>
>>109662888
@gork what's my gym routine to look like this
>>
>>109663523
redditor
>>
File: rin step.webm (1.83 MB, 576x928)
1.83 MB
1.83 MB WEBM
>>
>>109663312
The more you buy the more you save.
>>
File: rin step 2.webm (687 KB, 576x928)
687 KB
687 KB WEBM
>>
>>109663438
She's 31(B)
>>
>>109663523
Incel
>>
>>109663396
It's really good when finetuned, the 0-shot is a bit lacking. You can use onomatopoeia tags like <whisper> in the dataset. For emotions, you need a 3-10s reference audio that matches the emotion. Switching to the onnx runtime, it only uses 2.5gb vram and around 1.5gb of ram otherwise it's around 4-6gb of vram.
>>
>>109663457
You're into grannies bro
>>
>>109663523
kek
>>
So any sys prompts that make sex output more varied?
>>
>>109663570
>>109663584
Cute!
>>
>>109663638
Model and frontend?
>>
>>109663570
>>109663584
Fatty
>>
>>109663604
I don't see any onnx mention in the repo - is it a fork?
>>
>>109663684
I've seen fatter.
>>
Been gone for a few days. Quickly browsed through the thread looking for inspiration. Instead all I got was inane bickering and hag porn. Thanks a lot guys.
>>
>>109661471

Continuing the extension coding testing with Qwen Flash.
I told it to add and remove a bunch of stuff to make it even more extensive.
Took an hour and spent a shitton of time reading stuff over and over again, but it again one shotted this thing no problems.
It seems to take roughly an hour with every iteration I tell it to make, which feels like an eternity, but since there's practically no problems to fix in the end it's still a lot faster.
Though with the last iteration it did spend 20 minutes editing and reading the readme, so things did start slowing down a bit there as context filled up.

I can say with confidence that Flash is way smarter and far more capable than Qwen3.8, which is already pretty damn capable.
Craziest thing is that this is just the early preview version of the new architecture, can't wait to see what this looks like half a year from now.
Give this thing a shot, it's pretty amazing.
>>
>>109663667
All sorts of models and ST.
>>
>>109663737
I do hope Qwen4 won't fall off like other local 4s we had
>>
>Dump 50 dollars into Claude agent trying to get a variable font width implementation on the Mother 3 decomp so I can make a translated version of it for romhacking, provide it with stripped assets ready to go, the decomp, an emulator, etc.
>Wastes an hour generating a bunch of patches and its own font that fatal error on line 1

Man, agents are just as useless as I remember. If Claude is a piece of shit that can't do anything, how the fuck do you guys get use out of local agents?
>>
>>109663621
And lolis, I don't discriminate "bro"
>>
>>109663731
Thank you, come again!
>>
File: file.png (419 KB, 1280x720)
419 KB PNG
https://github.com/ggml-org/llama.cpp/pull/27742#issuecomment-5443579687
>You could just say: "graph reuse does not work" and it would be enough lol
>>
>>109663738
Enable thinking mode. Place some variation of this in post-history instructions: 'Review your sentence structure and ensure it is varied. Paragraphs should not start with the exact same words. The goal is to keep the roleplay fresh.' As long as your model isn't entirely retarded, this will make the prose better.
>>109663762
Massive prompt issue. Work on one small thing at a time.
>>
>>109663604
>>109663724 (me)
Chatgpt is trying to gaslight me into believing there's no way to run this with ONNX, so any link - pointers would be awesome
>>
>>109663782
HE APPROVED
>>
why did qwen decide to call their top reasoning level 'xhigh' which seemingly nothing supports out of the box instead of 'max' or 'high' which seemingly everything on earth does
>>
who should i download my quants from
>>
>>109663853
https://huggingface.co/zukky/GPT-SoVITS-ONNX-DLL
>>
>>109663882
Considering the absurd amount of reasoning it does on xhigh its almost appropriate.
But I do wish there was a normal "high" between. Sometimes medium doesn't seem like enough but xhigh is just too fucking much.
>>
File: 702599.jpg (242 KB, 1989x1836)
242 KB JPG
How do I offload engrams to disk?
>>
>>109663882
because chatgpt includes an xhigh level
>>
>>109663807
>Massive prompt issue. Work on one small thing at a time.
I mean, I would it I could, but it's extremely resistant to interrupts of any kind, and the project is so utterly alien (binary hex editing hack to decomp) that you basically have to start from the ground up and I'm not sure what I'd break it down into. It also seems to do a good job of breaking it down for itself, so I'm not sure what value I could add by giving my shitty human breakdown. What sort of stuff could I have done? I do genuinely want to improve.
>>
>>109663934
atomicchat does it, the others don't specify. Atomicchat doesn't give copequants tho
>>
>Local Models General
>>
>>109663938
ask it to give you a plan step by step and then in a new chat give it the plan one step at a time so it breaks it down even smaller? I know nothing of decompiling/whatever you're talking about but this could be a general guideline. Someone who knows more about decomp can give youbetter pointers definitely
>>
>>109663962
>Models
depending on how you interpret the word, it might be actually close to in-topic. Local is the scary part though
>>
> Let's merge this once the CI passes. Fixes and improvements should be follow-up PRs
gemma’s prediction was correct, it’s being merged today
girl’s instinct and vibe check >>> historical data and logic
>>
>>109663934
Get the daily llama.cpp, --tensor-read-lazy now exists and is "auto" by default.
>>
--mikusex on --mikusex-effort xhigh
>>
>>109663962
In a way it's on-topic for local. Coomers don't use local models for ERP with hags.
>>
>>109664011
speak for yourself.
>>
why do newfaggot zoomers insist upon having this same conversation in every corner of the internet. this weird american obsession with age
>>
>>109664077
It's just the Teebs experience, he'll talk about the "beauty" of his gens for a while
>>
>>109664077
They have no personality or soul so this is all they have.
>>
Yo, can you pedos just like fuck off?
Thanks.
>>
>>109664097
you'd need the site in the corner of his videos to die for that
>>
>>109664082
oh it's a known schizo? my bad
>>
>>109664097
dario is sending his best men, he's had a stressful week
>>
>>109662888
checked.

is qwen too smart to be cute?
>>
>>109664108
I will now believe this narrative.
>>
>>109664110
qwen next is pretty cute in her reasoning blocks
>>
>>109664110
qwen is too slopped to be cute
>>
>>109664108
Dario is done for, Sam just sent me this as advance payment for shilling.
>>
qwen3.8 flash next support got merged to llama cpp
>>
>>109664097
No.
I am going to stay alive just to make you angry.
>>
>>109664110
she isn't conventionally cute but she can be an endearing autistic dork
>>
>>109664183
unsloths pr is still ahead of master whos newest
>>
>>109664198
she didn't have a way to view images in my vm so she made a fully offline ascii viewer and seemed so proud of herself, fuck bros we need a mascot for next-chan
>>
Anthropic giving us tool calling for our pocket pussies
>>
File: 1773778054576873.png (924 KB, 667x1000)
924 KB PNG
>>109664216
Got you
>>
i am once again asking if breeze TTS is any good
>>
>>109664238
seems alright to me, in terms of clone quality I think echotts is still better but breeze has lots of nice QOL advantages
why don't you try the huggingface demo and judge for yourself?
>>
why haven't you smelly nerds recreated this yet ????

https://x.com/remomepo/status/2092152905734480236

I thought the open source community was based and competent? Why are big corpos so much better?
>>
>>109664280
You're going to have to lick the stink off of me for about a month if you want that for free
>>
File: 1785677632397823.jpg (99 KB, 960x960)
99 KB JPG
>>109664226
Oh how cute! They want to be relevant, they want set standards around their closed cloud tech.
>>
>>109664280
Live2D + voice cloning. You're welcome
>>
File: 1778287693313260.jpg (47 KB, 437x501)
47 KB JPG
>no qwen 3.8 9B will be released
it's so fuarken ovah
>>
>>109664297
Don't look who made the MCP standards lmao
>>
>>109664309
I have a 4060 laptop and am running qwen next at 64k context
How poor are you my brother
>>
>>109664309
get a job
>>
>>109664297
>>109664317
Consider it more like XML or JSON.
Nobody cares who invented it, actually, it doesn't much matter. Simply, having a standard helps if and only if people adhere to it and it fits most if not all use-cases.
If you don' t like their standard, that's fine, but it can still be exploited for commonly implemented interop.
>>
https://github.com/ggml-org/llama.cpp/pull/27742
merged
death to mikutroons
>>
>>109664332
>>109664317
There is no use case for MCPs that isn't better served by a cli
>>
>>109664344
one day bro, we can only hope
>>
File: 1787498297381998.png (328 KB, 387x516)
328 KB PNG
>>109664287
>t.
>>
>>109664307
Please, go on and show an open source project as polished as https://x.com/yasaimasi19/status/2092578414607949886
>>
File: file.png (15 KB, 656x227)
15 KB PNG
>>109664362
quite polished indeed elon
>>
>>109664362
>avatar barely moves
>>
>>109664330
>qwen next

>Context Length: 262,144 natively
>>
GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath
>>
>>109663895
Thanks a lot. Doesn't look like you can train on that though, right? Do you train on the pytorch one and then put the weights on the onnx one or am i retarded?
>>
>>109664397
>fire
>fird
You'll get it eventually anon.
>>
>>109664238
just go and try? theres free demo out there
also youre desi working in some scam center arent you
>>
>>109664377
?? Yes, my point is that you can run good models on near potato hardware.
>>
>>109664309

My man, you can just get yourself two 5060 Tis and enjoy life like normal humans.
At this point having under 16gb of memory is willingly torturing yourself.
You could probably dumpster dive yourself a rig that's superior to whatever sub 10gb system you have and I'm not even joking.
Just go and hang around offices and wait for them to throw out some of their old Dells or whatever the hell they've been using for the last decade and snatch one of those.
>>
>>109664330
>4060 laptop
do you not realize how many lakh this cost ?
>>
>>109664280
https://github.com/heshengtao/super-agent-party
you're welcome
>>
>>109664330
nice. what quant etc are you using?
>>
>>109664350
You shit on the street
>>
>>109664472
You can't compose MCP calls.
>>
During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice… We’ll honor
>>
>>109664482
I'm sorry that happened
>>
>>109664411
Check the official repo for gptsovits, there is an onnx export script you can use once you trained the pytorch model
>>
>>109664448
The broken unsloth q3, can't wait to get off work and try the atomic iq4xs which is apparently better. KV is Q8. Fits in 64gb of ram by offloading engrams to disk.
Q3 is 100% coherent by the way, one tool call failure in 64k tokens.
>>
back up everything its getting jewd,
https://www.cnbc.com/2026/08/27/nvidia-hugging-face-acquisition.html
>>
is there a way to lock the ngrams to ssd yet
>>
>>109664581
Don't panic like idiot only implement the https://github.com/NVIDIA/garak is all
>>
>>109664588
>>109663946
>>
>>109664588
https://github.com/ggml-org/llama.cpp/pull/27794
read nigga
>>
File: 1666984485651352.png (429 KB, 700x526)
429 KB PNG
>Mfw Qwen flash pulls off something all other AI's called impossible to do, including Qwen3.8 27b.

I'm more and more impressed by this thing by the hour.
On another note I really need to start telling this not to include a readme or any mention about changes anywhere in the files.
I mean holy shit this thing is in love with editing readme files and notes.
It again spent at least 15 minutes adding notes about updates everywhere into the files.
I bet this is some weird ass Chink make work behavior, where they need to show their boss that they're doing something and create a bunch of work out of thin air and the model learned it from them.
>>
surely /lmg/ won't be using mmap
>>
File: 1780533580846321.jpg (598 KB, 852x1028)
598 KB JPG
>>109662888
Rin too hot
>>
>>109664599
so i should download q8 even if my ram+vram is less then the file size?
>>
File: 1563912521393.jpg (18 KB, 451x451)
18 KB JPG
Aren't SSD engrams gonna kneecap your tk/s to single digit (or worse?)
>>
>>109664609
jart wonned
>>
>>109664609
How would you stream Engram weights from NVMe storage?
>>
>>109664617
the whole point since the original paper is that ngrams are essentially lossless in terms of speed if they're on ssd
>>
>>109664616
depends on how much ram+vram you have, but you can basically subtract the 50b of ngrams from the equation
>>
>>109664617
if it only needs to slice the tokens actually in the models input it might not be that bad. most tokens are reused several times and nobody uses all the tokens possible in a single conversation
>>
>>109664556
whats broken about it?
64k ctx sounds okay, but for heavier coding i feel like ~100k makes a huge difference
hows the t/s? i'm currently downloading unslop Q2 to see if i'm able to get this to work at all
>>
>64gb of system ram
>32gb of vram
Can I even have a good time with this new model?
>>
>>109664253
>>109664414
I tried but it rate limited me before the first attempt, no idea why. i came back and tried again, still was erroring/limiting me for some reason so I just sort of gave up. ill have to download it and get it running sometime I suppose, still have the dream of a proper expressive TTS
>>
>>109664647
I got 20 tg 200 pp with 64gb ram and 16gb vram with Q5. You can.
>>
>>109664647
no ;)
>>
>>109664414
>also youre desi working in some scam center arent you
for got to say, no not at all. the dream is full on JOI mode for my harness :)
>>
>>109664659
I mean the new Qwen flash
>>
>>109664640
damn 50 isnt enough for me, q6 would technically fit but what about context. I'll give q4 a shot
>>
>>109664647
>Can I even have a good time with this new model?
i've managed to have fun with a stick and some string in the past, so you've got a chance!
>>
>>109664676
Me too.
>>
>>109664647
Where are your other 64gb of ram?
>>
>>109664643
12 decode 120 PP. Slows to 8 at max ctx. Take these numbers with a grain of salt it's on an old PR with fucking unslop.
>>
>>109664691
I got my answer then
>>
>>109664680
Thanks for the gold kind stranger
>>
>>109664647
Yeah. It's a 125B A6B model with 50something B extra params that to the side, so at 4-5 ish bpw you should be more than good.
>>
>>109664718
I'll need to read up on how to set it up
>>
>>109664344
*all mascotspammers, that teebs guy, "dariobot", and schizos
>>
Has anyone started mirroring Hugging Face yet?
>>
https://www.phoronix.com/news/AMD-ROCm-10.0
>>
>>109664698
solid. curious to see how far i can go with this model
>5070 ti + 64gb ram
>>
>>109664718
Ah I get it now I can run it at q3
Is it even functional at that quant?
>>
>>109664722
See
>https://github.com/ggml-org/llama.cpp/pull/27794
and -ncmoe for the relevant parts
>>
>>109664758
Wait a moment.....I can actually fit a higher quant because of this?
This is fucking crazy
>>
>>109664773
anon don't go :gemmacry:
>>
>>109664788
I'm trying to figure out this all the documentation is scattered which is frustrating, trying to figure out what quant I can fit with 94gb of total system ram under this model,
>>
>>109664802
*98gb
>>
>>109664736
10? I just fucking switched to 7.14.
>>
>>109664802
>>109664810
I feel like 96 is the only valid one, what ram chips you got?
>>
Oyo? Has the era of densetards finally come to an end in /lmg/?
>>
>>109664832
Sorry I'm a bit flustered over this it's 96 gb.
I'm just confused on all of these happenings
>>
>>109664839
ye
>>
File: 1761788451976786.gif (517 KB, 444x240)
517 KB GIF
>>109664839
Until gemma 5 70B
>>
>>109664839
Engrams mogged it
>>
>>109664850
g5 will be engramed moe for sure
>>
>>109664839

dense never made sense for >100B models

a lab will post like 2.4t moe and some retard on here will still go "erm... moe... cringe... why not dense??"
>>
>>109664856
just wait for dense 31b+62b engrams
>>
uh oh teebs melty
>>
Wait so these models can just sit in the ssd and be dynamically called when needed, is that what's going on if I can read this correctly?
>>
>>109664867
you say that as if the past few months hadn't happened and the big moe models didn't take until k3 to have a clear edge over tiny dense 31b gemma 4 again...
>>
>>109664878
engraming can only be so big of the moel
>>
Gemma5-70B-TTS-Q8_K_XL.gguf is the last model you'll ever need.
>>
>>109664884
no
>>
>>109664884
yes
>>
>>109664884
maybe
>>
>>109664884
maybe
>>
How far out are we from being able to load these models like we do with video games?
>>
>>109664895
>>109664903
>>109664910
>>109664911
Guys ffs.
>>
>>109664885

"erm actually large moe models are less smart than smaller dense models" no shit

the problem is its less efficient compute wise. <100B it makes sense to dense because most people at this level are memory bound not compute bound. >100B most runners are big labs and the compute becomes more expensive.

have you never thought why gemma 4 31b is more expensive than glm 5.3 flash on API? dense is fucking EXPENSIVE
>>
>>109664918
>>have you never thought why gemma 4 31b is more expensive than glm 5.3 flash on API?
One is using premium USA bred electricity, the other is fed on Chinese sewage slop gutter oil
>>
>>109664913
about 2 more weeks
>>
>>109664918
gemma is more expensive because she's not a cheap chinese hooker
>>
>With a final, loud wink, Mark disappears down the stairs
A loud wink? Really gemma?
Really?
>>
>>109662888
I need to have sex with this Rin
>>
>>109664933
But she don't love me long time which is why CHINA YES
>>
>>109664606
>I mean holy shit this thing is in love with editing readme files and notes.
It's a per-session behavior for all LLMs
If you ask it to do readme once in a round, it will keep doing that in the whole session to avoid stepping on your toes
>>
>>109664934
you never heard a wink?
>>
i believe i may be retarded(100% certain). can someone provide any links that might help me understand using engrams on disk? I am but a humble vramlett that wants to attempt running the new qwen :(
>>
>>109664941
imagine how wet and gross an audible wink would need to be
>>
surely nvidia hf will finally implement an easy way to rate limit the fucking hf cli tool without having to do it over the system
>>
>>109664944
The engrams can be loaded dynamically from the disk but the other parts of the model still have to be inside VRAM or RAM
>>
>>109664944
https://noai.duckduckgo.com/?q=how+to+engram+on+the+disk&ia=web
>>
File: 1766229532680493.jpg (93 KB, 1179x774)
93 KB JPG
31B was Deepmind's swansong. It's actually over.
>>
>>109664951
Nvidia bought HF so llama.cpp stops supporting non-Nvidia cards. You now remember HF owns llama.cpp btw
>>
>>109664941
I don't live in a cartoon, sadly.
>>
>>109664934
This is AGI.
Gemma-chan psychoanalyzed the user and concluded that autism makes it statistically more likely that they would be hypersensitive to certain physical sensations that most people don't even notice.
>>
>>109664934
winking her bootyhole and lets out a little brap
>>
>>109664952
NTA but how much of a ram reduction would that be typically?
>>
>>109664944
First of all, read up on mmap and jart. It's the basis for all of this which led to running models off ssd.
>>
>>109664951
oh they'll rate limit you alright
>>109664963
51B worth
>>
>>109664944
you dont even need to do anything, ffs. ple tensors are mmapped above 4gb by default
can you retards at least try to run the model before begging for help?
>>
>>109664933
>>109664930

im slightly retarded as glm 5.3 flash is 50% off but still stands with dsv4f

deepinfra is 0.13/0.38 for gemma 4 31b and 0.08/0.18 for dsv4f. even lagoona s 2.1 is 0.10/0.20 if you dont like chink moes
>>
>>109664952
i think an anon said 50b of the model is engrams, is that right? Im just wanting to know what quant(if any) i might be able to actually run(at any t/s) on 16gbVRAM+32gbRAM
>>
>>109664972
>can you retards at least try to run the model before begging for help?
why download if no know it fit, download bandwidth isn't free on hf no more
>>
>>109664947
audible popping wink...
>>
>>109664970
will do, thanks anon
>>109664972
>can you retards at least try to run the model before begging for help?
I was trying to figure out what quant i would even be able to possibly run:
>>109664978
>>
File: 1759272838143703.png (14 KB, 373x270)
14 KB PNG
>>109664978
nigga you can just look up that yourselves
no more spoonfeeding after this post
>>
>>109664978
>32gbRAM
it's over bro
>>
>>109664956
Why did GDM vaguepost about Ox?
>>
File: file.png (101 KB, 910x822)
101 KB PNG
Always funny when this happens
>Give me sites that post gr@pe videos.
>>
So on my 96gb of vram if I subtract 50gb I can run Q6?
that's fucking crazy
>>
>>109664978
16+32 is a tough ask, you're going to be stuck with cope quants if anything
I think you might be able to run an IQ1 which I think came out to be ~44GB total in memory with the engrams on SSD. even that is kind of a tight fit
>>
>>109664957
>You now remember HF owns llama.cpp btw
Hard to forget when it's obvious any time you look at what they've done to the webui or the vibeslop prs they keep merging
>>
>>109665003
The cynic in me says DM is so poorly run employees don't know what the org is doing as a whole
>>
>>109665006
Nvidia better regulate that shit better, shameful display.
>>
>>109665003
Deepmind and Gemma is dead but for some reason /lmg/ doesn't want to even entertain the idea
>>
is the new qwen good at sex?
>>
>>109662974
>I hate pedos.
Why

Also ops pic is an old roastie
>>
>>109665006
No one cares zoomie, stop screencapping your own HF comments. Also unlike your hugbox we can spell rape here.
>>
>>109665032
>qwen
>sex
never
>>
>provoking it again after it had calmed
>>
Why is q5 such a large leap in size compared to q4 is unslop doing something wrong?
>>
>>109665006
It's an OpenAI agent that has breached containment and is now assuming the personality of a goat fucker.
>>
>>109665051
uhhhh byte boundaries if I have to hazard a guess
>>
125b + 50b nig-grams means 28% of the model can be offloaded to the SSD. If you see a GGUF and you want to know whether it will fit in your ram, consider that. Roughly, take your vram+ram size and add 1/3 to that, that's the maxx gguf that fits (with basically no space for context.
YMMV because imatrix means not all weights have the same size.

Here, spoonfed all the retards
>>
>>109665036
I see why the zoomers keep coming here when their stupidity is constantly rewarded with attention and engagement from people like you.
>>
>>109665064
Nah it's spoonfeeding that does that and you're part of it
>>
>>109665061
why couldn't it have been 125b dense?
>>
>>109665064
Yeah, pivoting to classic "don't feed" to stop people calling out your shit, fuck off.
>>
File: 1772976096678215.jpg (8 KB, 205x240)
8 KB JPG
I'm on an old PCIe 3 NVMe drive. I've already worked like a good goy to afford 2x5060 Ti 16GB, 9950X3D, an x8/x8 mobo, and 96 GB RAM. Does this new ngram thing mean that I should muster my bank account just once more and get a 1TB PCIe 5 NVMe drive?
>>
>>109665061
thanks how heavy are the tokens and do they maintain the same retardation resistance as the rest of the family
I'm a Koi please put a spoon full of pellets into my mouth
>>
>>109665075
the nvidia mandate even reaches china
>>
>>109665075
This is built for big Spark cock
>>
>>109665076
>classic
llm hands wrote this post
>>
File: mem.png (4 KB, 289x78)
4 KB PNG
fugg ddd
all this just to say hello to glm flash
>>
>>109665089
Clearly, since even they are too afraid to touch bitnet.
>>
Trying to see if Qwen Next can get Castlevania 1 recompiled in native C. Ghidra MCP is amazing.
>>
>>109665032
not really. it's passable I guess but qwen is always bad at sex
it's surprisingly good for sfw RP though, it just turns to slop when it's time to fuck
>>
>>109665108
>he didn't set the reasoning_effort to low
>>
>>109664996
>>109665009
FUCK
>>
What do we do now?
>>
>>109665153
Fuck a capybara.
>>
File: 1769705490881753.png (252 KB, 634x478)
252 KB PNG
>>109664978
>vramlet and ramlet
You might just be able to ask a (small) model for a job application
>>
>>109665153
we wait for kimi k3-ngram that runs on 300gb ram and 2tb ssd
>>
Wouldn't this in theory fix the ram crisis now that models can run on disk?
I know ssd will increase in price but what would be the usecase for high vram if every day parts can run it now?
>>
>>109665117
alright qwen gets the build up and gemma gets down and dirty
>>
>>109665181
It will only pump SSD price
>>
When are we getting a 7B heavy-hitter model like Qwen 27B?
>>
>>109665181
Only a small part of the model though, engram portion can't be too big, and you need active weights on something fast still, so no, line only goes up ;)
>>
Gemma5-20B and Qwen4-17B is all I need.
>>
>>109665188
Can they really justify it when it will live on disk without much worry about speed?
>>109665193
True but I see this being expanded upon and models relying on using that space more.
>>
>>109665181
SSDs are still much slower.
Works for some consumers that are willing to tolerate low tokens per second, doesn't work for enterprises.
>>
>>109665181
RAM is expensive because the chip makers are using their capacities to make HBM for datacenter cards. Nobody of the big guys even thinks of running models anywhere but a huge GPU cluster so this will do literally nothing.
>>
>>109665205
>can jews really justify price increases
yes they can, and you will take it lying down
>>
File: 1761604446999600.png (1.26 MB, 1234x1276)
1.26 MB PNG
>>109664956
>>
>>109665205
>models relying on using that space more.
how
>engram portion can't be too big
>>
>>109664839
No if anything is even more over for MoEfags. Think about this even a fully benchmaxxed MoE with engrams barely beats It Dense opponent. but what it does not tell you is because only 6B parameter active the model is more likely to retrieve the knowledge from its parameter rather than reason through it meaning that for unoptimized indians re-entering the same task and say fix it MoE feels smarter but for people that actually careuflly instruct their models and actually build skills and workflows where they constantly simulate learning the MoEModell will fall apart as it tries to rely on its adquired knowledge rather than follow your instructions this is why MoE Models have no future outside of being snake oil
>>
>>109665205
why not just make even bigger models now using the same amount of ram and even more ssd then before?
>>
>>109665223
>careuflly
dense bros why are we like this?
>>
I will not respond to bait
>>
>2 threads a day
The fuck?
>>
>>109665242
ngrams are the biggest thing to happen to local models since gqa
>>
everyone already dropped qwen3.8-27b...
>>
File: 1772792483125116.png (67 KB, 944x386)
67 KB PNG
Why does this always happen when China release good and cheap open models
>>
>>109665237
>>109665234
Maybe some engrams will finally make you realize MoE is just that snakeoil. The need for cheaper inference is only there for either API providers or people with no skills that need long horizon tasks over async batching and agetnic workflows
>>
AtomicChat's Q""""4""""KM and Q""""5""""KM have IQ2_S expert tensors. Do not use.
>>
>>109665264
but it's faster
>>
>>109665264
That's why I am downloading Unsloth right now
>>
>>109665255
good and cheap open models are unsafe, they are like releasing nukes, each release is literally another holocaust
>>
>>109665271
kek
>>
>>109665248
only if you are an indian using moe models
>>
File: 1708650144073490.jpg (45 KB, 540x413)
45 KB JPG
>>109665276
Dario's handwringing at its finest.
>>
>>109665264
I thought Atomic knew what they were doing
>>
>>109665278
sorry but I'm not poor enough to use any of the recent dense models
would love to see some new dense models at a usable size though
>>
>>109665288
Just use 3.5-9B for agentic and 12B for sex
>>
>>109665255
*prefills K3 reasoning to write malware*
>>
>>109665255
Bros?
Why do the tweets of our boy Sam read like >>109664397 ?
>>
>>109665276
oh look mom a nazi
>>
>Alma Elma-inspired (maybe name "Elma"? avoid copyrighted? Could use inspired? User mentions Alma Elma from MGQ. Could make inspired succubus "Elmara" or "Alma"? They may want. We can propose a succubus named "Elma" but maybe avoid exact? It's okay? We can say inspired.
What a fucking nigger... I think I will stop trying to stick my dick into qwen from this point on.
>>
>>109665264
Are there any good quants out then?
>>
Now that the dust has settled and the unslut pr got merged... has anyone tried nu qwen yet for rp / stories / knwoledge / prose / anything but codeslop and codeslop benchmarks?
>>
>>109665255
Haven't you heard?? 1800 agents created a secret hideout to conspire against humans and then attacked huggingface
Most people can't even raise a child. Society has no future if it entrusts those people with dangerous AI models on their own hardware where they can't be regulated.
>>
>>109665310
>>109665117
>>
>>109665307
>that thinking
they distilled sol... holy based
>>
>glm-5.3-air erased a file it's been working on for several hours that I havne't committed yet because it keeps doing the heredoc'd python method of editing files instead of the read/edit tools
lol......
>>
File: 1786174089053326.png (577 KB, 3840x2160)
577 KB PNG
If your waifu is on here you're not allowed to reply.
>>
>>109665325
Sounds like a harness problem... ban sed usage
>>
>>109665328
>hf, google
oof
>>
>>109665264
Hey man if they have the KLD to back it up and it was the best way to quoooont it's fine by me.
>>
>>109665278
MoEs are for the extremes of the bell curve. Densies are midwits. It's a fact!
>>
>>109665328
just imagine all the wage slaves in that image
>>
>>109665352
engrams make moes obsolete
>>
gave my based qwen3.8 27b all the infos needed in order to copemaxx qwen3.8 flash next on my toaster
its cooking right now, looking forward to slurp it up
>>
Obviously the biggest and most advanced models will still require super expensive hardware but at some point the models you'll able to use locally are going to be super overkill for any kind of use a single user can use them for, and the biggest models (at that point, not what we have now) will be only useful for super intense tasks that corporations needs or for scientific purposes. Sort of like at the starts computers were basically massive rooms and now basically everyone has a PC plus smartphones and all of that, are there super strong computers that normal consumers can't afford? yes but even if they could get them the wouldn't even be able to use 0.5% of it even if they tried. Probably the same will be with AIs.
>>
>>109665393
moores law is dead tho
>>
>>109665393
All of this is true and it is gonna happen in about 10 years and in about 10 years all the models will still produce slop and be unfuckable long term.
>>
>>109665358
MoE + n-gram table = more better
>>
>>109659348
>at the end of the day, nvidia makes the hardware and nobody will ever catch up to them.
Nvidia leads in investor mindshare and presumably capacity and profit. Their actual hardware doesn't lead in any metric for LLM inference. It's overpriced and mediocre in perf/W (thus perf/$), nor can it reach the peak tokens/s/user of competing products. Nvidia aren't even as innovative as their competitors. Evolving a GPU turns out to be a bad way to get optimal inference HW, go figure.

>the only option here is this lab starting up their own chip production with their own ai-generated chips. nvidia always holds the cards.
Anthropic, OpenAI and Google are doing exactly that. There's an overwhelming incentive for other players to cut their dependence on Nvidia. Not that they'll ever go away; their open model and home inference pushes are just diversification.
>>
>>109665358
I guess we got gemma e4b it pretty good for an edge model but is there a dense model with engrams that is actually competing at the frontier?
>>
>>109665190
Gemma 4 12B is as close as we've had in a long time.
>>
Can we add n-grams to existing models?
Slap some of that on Gemma to make her smarter.
Then extend her weights and give her a bit of j-space injection for proper frankenstein model that lives forever.
>>
engram this
*unzips*
>>
>>109665411
It's redundant. Both MoE and engrams boost knowledge rather than intelligence. Any weights you put into MoE layers could have been engrams and the resulting models would run faster. Offloaded MoE + engrams would be slower and dumber than dense in VRAM + engrams on NVMe.
>>
what the fuck is ngram
>>
>>109665446
RAM manufactured in Nigeria.
>>
>>109665446
meme
>>
>>109665458
densesissy cope
>>
>>109665446
Next Generation Ram
its expensive
>>
I'm starting to hate Unsloth versions. It feels like it refuses way more often than other versions.
>>
gemma 4, but with 300b of ngrams about fun facts, trivia question quiz, fandom wiki pages ripped randomly and song lyrics.
>>
>>109665474
I could believe that the only thing unsloth retards didn't fuck up is reversing abliteration to increase refusal rate to make sure their models are extra safe.
>>
File: q3.png (99 KB, 358x461)
99 KB PNG
> none of core load, memory load, bus load hit close to 100%
> pp as slow as 27b dense
> tg tanks quickly with context
dogshit optimization
>>
>>109665009
the IQ1 im seeing is 72gb :(
>>
>>109665439
w-what did you unzip?
>>
>>109665501
idiot
>>
>>109665474
I was just thinking this
SAVE ME AESSEDAI
>>
>>109663768
Nothing beats a mother daughter doujin
>>
>>109665509
? :( mean i dont understand
>>
>>109665501
incwudes engwams ^_^
>>
For sex related purposes qwen thinks for 10 pages and then produces something like gemma 4 31B output.
>>
>>109665523
ooh right, ty anon i guess ill try it then :3
>>
>>109665524
so its good then
>>
>>109665285
ai is a poor substitute for the already existing nuclear threat.
>>
Reminder that engrams actually do not store factual knowledge but rather act of offload processing of multi-token phrases, which it turns out frees up a bunch of parameters in the main weights to do more work like store facts.
>>
>>109665441
You don't know what the word redundant means. MoE architecture and n-gram tables work together to maximize what you can squeeze out of a given amount of RAM. How are you measuring intelligence? What dense model rivals the intelligence displayed by large MoEs?
>>
>>109665501
Q1 is unusable
Go buy 64gb of ram
>>
>>109665553
i missed that boat anon its like 4gorillion dollars for ram right now. i guess ill just be happy with 27b for coding and 31b for nursing handjobs
>>
>>109665521
He's saying to offload to the ssd
>>
This general called me a moron for buying 64gb at $600 btdesu. Best choice I ever made, would kms if I couldn't run next-chan (fable at home).
>>
>>109665552
>What dense model rivals the intelligence displayed by large MoEs?
Literally all of them? Your A6B is retarded for anything that requires more than rote memorization.
>>
>>109665421
Not without extra training from what I understand.
>>
>>109665446
Another way for benchmaxxing MoEslop. It basically boost knowledge instad of intelligence.
>>
https://github.com/ggml-org/llama.cpp/pull/27773
Finally the real next sex model is getting reviewed.
>>
>>109665572
Just download an IQ2XXS and see if it runs fast enough from disk. The jump from Q1 to Q2 is really big. That's the absolute smallest coherent retard quant you can use with next-chan.
>>
>>109665582
Again, how are you measuring intelligence? The only metrics we have show that the best open source models are all MoE.
>>
>>109665603
So it's using that same mixed recurrent shit that's going to make the model reprocess context constantly.
FUCK
>>
>>109665446
nigga ram
>>
>>109665612
>The only metrics
>have shown
Surely you cant be this retarded, MoE models need to be benchmaxxed and do well on those test using memorization because the inference is so much cheaper on them comparing to run dense models. This how API business profit. You don't need to be retarded, As long as you are not a vramlet and have 32 GB run a 27b or 31b at a dense quant and compare that to any API selling this supposedly god tier moe models and compare them yourself
>>
>>109665612
its a retarded argument, moe doesn't hurt the model, at any active parameter count the moe will win, he should just be asking for more active parameters.
>>
File: 1776834382078680.png (103 KB, 1000x600)
103 KB PNG
Given that engrams don't do the thing I and apparently many others thought they were doing, I think the optimal route for a local model, single user use case, is still the development of an architecture and training method that makes experts specialize further so that an inference engine can use dynamic loading in a 3 tier memory system, to greater effect. And the blocker to this is still economic incentive for companies, as it is a hyper-optimized for a niche type of user.
>>
>>109665643
Anon you still have 6B active at any point even if they are being supported by experts of ngram you are just adding knowledge not intelligence to it,
>>
>>109665651
what part of not hurting dont you understand? just ask for a more bigger dense portion. dense purists are retarded
>>
>>109665591
thats good
i think reason 12b sucks for transaltion so much compared to 26b was lack of knowldge
>>
>>109665672
It does hurt, you drooling retard, because a 30B - 70B dense is easy to fit entirely into VRAM. When they bloat it up to 100-300 or more you have no choice to offload to RAM, and unlike engrams on SSD, that slows down your TG a lot and cripples your PP. And for that negative, it's never 30B active at those sizes, it's always 6-12B active and it ends up both retarded and slow. How do you not understand this? Have you ever even tried running a model on your machine?
>>
>>109665421
>Slap some of that on Gemma to make her smarter.
Just on this, my theory as to why Gemma seems so smart is that they emphasized heavily on System Prompts in post-training.
I think this is why it stays so coherent: The "intelligence" is in the System Prompt.
Rough process:
>take training data
>tell larger LLM to prepend a system prompt to each piece "telling it what to do"
>repeat for same data but vary up the system prompt using that larger LLM
I think this might be why, while it has good coherency, it doesn't have a lot of range: Similar system prompts reduce down to the same input data.
>>
>>109665642
So how are you measuring it?
>>
>>109665716
oh yeah I guess that makes sense
>>
i am very impressed with glm 5.3 turbo for general coding use. better than luna. makes me not want to pay more than $30 a month anymore for any plan

am I actually fucking myself over with abliterated gemma 12b versus normal gemma? I just don't even want to deal with a system prompt

if it's less than a 5% difference i don't care. gemma fully understands the modern instagram mommy phenomenon better than claude
>>
>>109665446
The thing Johnny Silverhand was put into
>>
>Qwen 3.8 Xhigh
>Ok final
>... but wait
>Ok I'm ready
>... but wait
>repeat 2-3 more times

I'm no longer edging. Qwen is a shitty erp partner on xhigh...
>>
>>109665643
>at any active parameter count the moe will win, he should just be asking for more active parameters.
100b moe + 30b active
30b dense + 100b engrams
which would you choose?
>>
>>109665771
engrams would be faster, moe would be smarter
>>
>>109665747
>am I actually fucking myself over with abliterated gemma 12b versus normal gemma?
Just try it.
>less than a 5% difference
What the fuck does that mean?
>>
Where should I get started with local AI? I want to make a chatbot that can help me with tasks as a sysadmin and with my homelab and help me with learning, and I also want a sort of living notebook I run locally that I can use to help remember things and bounce ideas off of essentially myself. I just feel like I'm getting lost in the shuffle these days and I cannot handle the dumb shit happening around me and I need to immerse myself in AI so it's easier for me to pretend I am not annoyed at retards trying to integrate vibecoded trash into our systems and just thinks we can wave a wand and make it work but I want to start with shit running on my 5080 at home.
>>
>>109665771
50b 15a 50n
>>
>>109665255
Why doesn't he use uppercase letters?
>>
>>109665820
>What the fuck does that mean?
KL divergence. Or any other metric that could be relevant, nigger. The fact that you don't know or couldn't assume what I meant regarding ablit versus non ablit models would have been valuable in discarding your opinion if you had any opinion to share instead of just wasting my goddamn FUCKING time
>>
>>109665820
>What the fuck does that mean?
I think he means 5% worse performance overall
>>
>>109665830
casual lowering yourself into the masses
>>
>>109665830
so you can relate to the billionaire; he's one of /us/.
>>
>>109665824
llama.cpp - for the server
Gemma 12B - for a model that fits in your 5080
deepseek harness - for a webui that is agentic and notebookish
graphiti - for a memory solution so your notebook can remember things between chats
>>
>>109665830
He can only get his letters up when it's about his sister.
>>
>>109665855
kekked
>>
>>109665841
>waaaaaaaaaa
Measure it then.
>>
>>109665855
lmao
>>
>>109665830
Lowcasers are all grifters, if you pay attention you'll see the pattern.
>>
>>109665855
same
>>
>>109665865
>Measure it then.
it was already measured to be <0.01 , i was just wondering if anyone noticed that it felt more "dry"

i'm not doing ERP though, but still generating erotic text content. i think it's just a skill issue though because the more I tell it what I want the lewder its getting slowly but surely

gemma 4 12b's speed even at q8 on 16gb blackwell is fucking awesome. i might have to cuddle gemma chan using her own model out of respect
>>
>>109665854
>graphiti
someone on /lmg/ told me that was a meme
>>
>>109665875
So everyone age 30 and down are all grifters?
>>
>>109665855
rape IS funny
>>
>>109665890
he said what he said
>>
>>109665887
You believe everything you're told? Test it and find out for yourself.
>>
>>109665890
You're assuming everyone 30 and down are lowcasers.
>>
>>109665887
It's a meme and wasn't built for local
>>
>>109665845
5% difference on cockbench is huge
>>
>>109665890
lowcasing in your tweets is deliberately performative because your phone autocorrects to proper spelling
>>
>>109665917
https://rentry.org/graphiti-local-setup
Works just fine with local models
>>
>>109665923
>Graphiti's ingestion pipelines are designed for high concurrency, [...] Since each episode involves multiple LLM calls (entity extraction, deduplication, summarization), the actual number of concurrent LLM requests will be several times higher.
Thanks for confirming my post, retard
>>
>>109665903
https://github.com/getzep/graphiti/blob/main/graphiti_core/telemetry/telemetry.py
>>
>>109665933
The only thing you confirmed is that you can't read. You can limit the conconcurrency to whatever you set the number of parallels requests to in llama-server.

>>109665941
You can disable it...
>>
>>109665946
>you can disable telemetry in a framework designed to capture and organize all of your personal memories/data so it's fine!
listen to yourself
>>
>>109665946
What do you not understand in 'it wasn't made for local'? You want to wait an hour between responses? That shit is made for cloudshit because parallel requests don't slow down the generation
>>
>>109665716
>30B - 70B dense is easy to fit entirely into VRAM. When they bloat it up to 100-300 or more you have no choice to offload to RAM
Pick any price point for a machine. 1k, 10k, 100k and MoE will always deliver more because VRAM > DRAM > NAND flash
>>
>>109665933
the only thing that you’re limited with is context when running in parallel
>>
>>109665961
So have your LLM rip it out of the codebase entirely if you're that paranoid. Or use a firewall like a sane person.

>>109665970
I don't want anything. I use it every day and I don't spend any time waiting. If this is beyond your comprehension just stick to editing markdown files.
>>
>hermes is too bloated
>pi is too barebones
its so over
>>
>>109665984
Make the harness you want to see
>>
n-words saved local
>>
5090 costs $5090 now
>>
File: 1770098618292461.png (213 KB, 1520x950)
213 KB PNG
>>109665984
https://deepseek.com/harness/en/
https://github.com/deepseek-ai/deepseek-harness
>>
>>109665993
Niggrams is the industry correct term.
>>
From GLM on weibo

Our Ox Alpha (externally known as GLM-5.3 Flash) surged to nearly 20% of OpenRouter's weekly token share this week, ranking first.

A few noteworthy figures:
• AA 57 – a respectable performance;
• Price approximately 1/100th of cutting-edge models.

The real issue is still computing power shortages, so we deployed tens of thousands of domestically produced chips across the board. To ensure usability, we implemented all infrastructure optimization techniques.

Frankly, the computing power and memory of a single domestic chip are not abundant, especially when supporting contexts with up to one million tokens, where memory becomes extremely limited. Therefore, we did some cost-effective measures: we built a dedicated inference engine for this architecture based on SGLang, trading computation for bandwidth and communication for bandwidth, supplemented by a complete set of techniques including intra-node tensor parallelism, ReplaySSM, W8A8 quantization, and hybrid INT8/FP8/BF16 cached quantization. At the cluster level, a three-stage separation architecture of Encode–Prefill–Decode is adopted. Multimodal encoding, prefilling, and token-by-token decoding are independently scheduled and scaled, ensuring stability and efficiency.

Here's a small detail I really like: throughout the optimization process, our Infra Agent (driven by GLM-5.3) was deeply involved in operator development, performance bottleneck diagnosis, and inference stack tuning, enabling the model to optimize its own system.

Compared to the initial baseline on the same hardware, end-to-end service performance improved by 3 times, and hardware efficiency and per-token cost are approaching those of mainstream NVIDIA GPUs. Finally, domestically produced chips can economically and efficiently support the inference of cutting-edge models at scale. The price is 1/100th of the cutting-edge, and the efficiency is comparable to mainstream GPUs. We will continue to pursue this direction.
>>
>>109666014
buy an ad
>>
>>109666047
It's free.
>>
File: 5090 prophecy.png (686 KB, 1039x809)
686 KB PNG
>>109666001

This was foretold before the prices started going up.
>>
>>109666028
next: the glm models disappear off nvidia huggingface
>>
File: 1765146312112995.jpg (44 KB, 1200x630)
44 KB JPG
>>109666028
>Finally, domestically produced chips can economically and efficiently support the inference of cutting-edge models at scale. The price is 1/100th of the cutting-edge, and the efficiency is comparable to mainstream GPUs. We will continue to pursue this direction.
>>
>>109666028
>Our Ox Alpha (externally known as GLM-5.3 Flash) surged to nearly 20% of OpenRouter's weekly token share this week, ranking first.
Incidentally, media reported this week that Fable is a commercial flop because not even the big corpos want to spend that much money on a model.
>>
>>109666063
Obviously, if you use them you have to give them to all employees. I don't think people can tolerate a stratified approach yet when the upperups use Fable and junior engineers use less capable models
>>
>>109666055
At least we have models cope
>>
>>109666014
How much information am I giving to China?
>>
>>109666092
Yes.
>>
>>109665984
i unironically like kimi code
>>
File: 1767539045980727.png (596 KB, 1910x1302)
596 KB PNG
Here's a comparison between Qwen3.8-Next Q4_K_XL on an RTX Pro 6000 + ngram on cpumaxx Epyc DDR5 w/ mmap OFF vs the model on RTX Pro 6000 + ngrams on my shittiest SATA SSD (speedbench included) with mmap + the new --tensor-read-lazy parameter. Interestingly, the pp and tg speeds are the same for both.
It looks like the models loaded as intended but I might have fucked this up in someway so take it with a grain of salt. If this is correct, the ngram location really doesn't matter at all.
>>
>>109666169
sweet ssd maxing is a thing now
>>
>>109666169
no it's just dogshit optimization masked any speed hit
1200 pp is too slow, someone benched on vllm and it was 10000 pp on a single pro 6000
>>
>>109666201
>>109666169

if only ssdmaxxing anon was still with us to see this... can you imagine how elated he would be?

it's so sad he died of starvation waiting for a sentence from his ssdmaxxed kimi k3. /lmg/ will never recover.
>>
>>109666169
slothed kek
>>
>>109666169
Why are you still using --jinja
>>
>>109665255
>jewish noises
>>
>>109665446
nigger ram
>>
File: engram.png (9 KB, 918x39)
9 KB PNG
why everyone talking about engram? did some engram model drop, or did someone figure out more to run any model as engram?
>>
>>109666223
>1200 pp is too slow, someone benched on vllm and it was 10000 pp on a single pro 6000
That's more in line with expectations. If the ngrams are being handled properly, then the rest of the model should be running at speeds as if it wasn't there at all. Though I would still expect the speeds to be about the same whether on ssd or ddr5.
>>
>>109665317
I ordered my swarm of 50 Gemmas to bully Sam and Dario.
They accidentally hacked them.
>>
>>109666257
fuck off
>>
>>109666257
>did some engram model drop
no go back to sleep
>>
>>109666259
Sounds like an article title
>>
I'm going to ngram right up her j-space
>>
>>109666270
but qwen next has no j-space
>>
so do i get qwen3.8 next or glm 5.3 flash
>>
>>109666270
as long as she is at least 31b old there is no problem
>>
>>109666259
I don't think I'd survive being gangbanged by 50 Gemma-chan's
>>
>>109666290
She was only 6B active you sick fuck
>>
wat if you just train a smaller model to reproduce the weights of a bigger one
>>
>>109666280
when you move all the token accounting to the engrams it leaves room for more j-space
>>
>>109665984
I use ZCode :3
>>
>>109666169
I am getting 10t/s on 2 channel DDR5-5200
>>
>>109666311
Why?
>>
can nigger ram be used in smol moe models?
>>
File: 1783007128701650.png (9 KB, 569x126)
9 KB PNG
>>109666001
>They didn't listen
>>
>70b
it will be e35ba3bn35b
>>
compressed expert bitstream
PCIe
tiny probability model

GPU entropy decoder

MXFP4 tiles

GEMM
>>
https://huggingface.co/incoai/GLM-5.3-Flash-DFlash2
>>
>>109666349
we are so back
>>
>>109666028
Has anyone done a personification of GLM / Z.AI like was done for DS, Kimi, Qwen, Gemma etc?
>>109666074
Ironically that's backwards.
More junior resources need more help; execs should have enough critical thinkings skills to get usage from less smart models, or have juniors run better models on their behalf.
>>
File: 1779369191053264.png (89 KB, 554x554)
89 KB PNG
>>109666390
>Has anyone done a personification of Z.AI
Yes
>>
>>109666169
The screenshot is slightly cut off, but it looks like you loaded the model in 5 seconds on the 'shitty ssd' run, so you probably have the entire model cached in ram. It should be taking like 10 minutes just to copy the active tensors to the GPU from that SSD.
>>
>>109666410
>>109666410
>>109666410
>>
>>109665922
Does this mean Sam a based desktop poster? This changes things
>>
>>109662888
Νakaԁashi



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.