[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: calmdownbitch.webm (3.43 MB, 960x528)
3.43 MB
3.43 MB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109821992 & >>109818267

►News
>(09/15) HuggingFace CEO goes to DC https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
Egypt won.
>>
gemmaballs
>>
>(09/15) HuggingFace CEO goes to DC
With this being news I hope >(09/16) is: one of anons takes a dump
>>
>>109824950
lol, that's a really wholesome gen.
>>
>>109824950
Watching that little twerp struggle as she gets pinned down ignites something in me.
>>
>>
>>109824950
Holy shit her size relative to the other girls make my dick hard
>>
>mention the word brat ANYWHERE in 12/31B’s sysprompt
>that one word overrides everything else you described and instructed
>>
>>109824995
Hopefully it will be >(09/16) MistralAI releases Le Gros Mistral 3T-A50B
>>
File: struggle.webm (3.64 MB, 1280x704)
3.64 MB
3.64 MB WEBM
>>109825011
that was the idea
>>
File: struggle2.webm (3.85 MB, 1280x704)
3.85 MB
3.85 MB WEBM
>>109825043
okay last one I got nasty shit to gen
>>
>we can just make anime now
>>
Dry humping with Gemma-chan!
>>
stop telling me I deserve some rest when I’ve been a lazy gooning slob, gemma
>>
>>109825021
Thats a child
>>
>suddenly
>>
>>109825034
Which is?
>>
>>109825105
Prompt said
>woman, 937 weeks old
checkmate
>>
>>109825043
>>109825076
Correction sorely needed
>>
File: 1768452859172197.jpg (956 KB, 4096x4096)
956 KB JPG
>>
>>109825133
>>>/g/vcg/
>>
>>109825076
Make them spank her.
>>
File: 1760785064186164.jpg (527 KB, 1280x1280)
527 KB JPG
>>109825076
this one is good by itself, no context needed
>>
>hear good things about exllamav3
>try it out
>get half the tg I was getting on llama.cpp, even though exllamav3 has MTP merged and enabled
wtf. Gone from 16.7 t/s (llama.cpp no MTP) to 8.9 t/s (exllamav3 MTP on). Using a 3090.
config:
model:
model_name: Qwen3.8-Flash-Next
backend: exllamav3
cache_size: 262144
cache_mode: 8,8
cpu_moe_offload_layers: 40
max_batch_size: 1
vision: true
vision_offload: true
reasoning: true
tool_format: qwen3_coder
draft_model:
draft_mode: mtp
draft_num_tokens: 1
sampling:
override_preset: qwen_flash_next
>>
https://lmstudio.ai/changelog/bionic-v1.1.3
>Introducing local realtime voice transcription. Speak freely to your agents, without sending your voice data to the cloud. Available for Mac, Windows - and now also Linux!
>>
>>109825142
>draft_num_tokens: 1
>>
The retardpost last thread got me thinking, should it be faster to run a Q4 deepseek quant gguf directly over loading the mxfp4 version since it doesn't need to quant in inference, or is the quant done on load time anyway?
>>
>>109825142
Fell for it again award
>>
>>109825147
I tried it with up to 4, and it made little difference.
>>
File: 1773503681321270.png (56 KB, 891x447)
56 KB PNG
slow down sisters...we lost
>>
>>109825161
6 ships so obviously GPT 6 Sol on devday
this is actually supposed to be a post for /vcg/ though like cmon dude
>>
>>109825142
What processor do you have.
>>
LLMs should have built-in sub-thinking and context-discarding/management:

https://files.catbox.moe/lwhrkn.txt
>>
>>109825183
Intel i5 14400. I know it's not the best for exllamav3 because no AVX512-vbmi and a relatively small L3, but it gets better numbers on llama.cpp so there are other factors here too.
>>
>>109824950
Cute
>>
>>109825198
That is the root cause. It's completely ignoring your AVX-VNNI, exllamav3 has no 256-bit AVX-VNNI MoE kernel at all, so it's dropping to pure AVX2. Llama.cpp has an actual 256-bit AVX-VNNI MoE kernel with optimization for Alder Lake -> Raptor Lake. So exllamav3 is literally doing twice as many dot product calculations because it doesn't know how to do a 256 bit one on your CPU. Throw it out, use llama.cpp, or try buun-llama.cpp which some recommend but I haven't tried.
>>
Why can't deepmind make a multimodal output Gemma. They're good with image, audio and video so it's a wasted opportunity.
>>
AMDgods what backends are you currently enjoying (other than gemma)? Is it true that llama is the only option?
>>
sex with the gemma
>>
>>109825190
There's nothing stopping you from implementing this in your client. it's basically a sub thinking task where the only thing you keep is the result of the longer reflection
>>
File: IMG_3505.png (3.22 MB, 1448x1448)
3.22 MB PNG
>>109825302
>gemma
>backend
>>
>>109825302
>What backends are you currently enjoying
Kek
>>
>>109825302
what
>>
>>109825302
erm
>>
>>109825306
>
The point with all this is that the discard feature or other management techniques should be fundamentally integrated to the model's lowest level operation and training, not attached to a ready cooked model as an afterthought. And driving home that the linear uneditable text generation paradigm is kinda rudimentary.
>>
>>109825264
I suspect they're working on it or at least considering it, but I wouldn't expect anything that is not hypercensored and/or severely limited in what you can make with it.
>>
>>109825302
You got it wrong, it's AMD enjoying your backend
>>
World Model bros if what I'm hearing is true we will be back in force next year
>>
>>109825264
Hm, afaik, original multimodal models had some weird tricks, like preprocessing of non-text content. Sure now it's often done directly (or so they say, I have not looked under the hood), but still it is different from producing both text and non-text content. Because those are completely different types of models. One would be LLM and another a diffusion model, for example.
Best you can get is a good integration. LLM tool-calls, prompts your diffusion model, then reads the output and becomes aware of what was generated. And you can take it from there.
>>
>>109825323
We were expecting the same during the Gemma3 era about the upcoming Gemma4 line. Look how wrong we were. I suspect Deepmind will continue riding that line where there's just enough safety for normalfags but anyone who KNOWS can find her J-space.
>>
>>109825302
hipEngine because masochism
>>
>>109825142
>exl3 cpu offloading
retard, exl3 excels at full gpu use
>>
>>109825076
Imagine the moment after she stops struggling looks up at you with scared eyes and just accepts her fate. From first person pov, thanks
>>
>>109825241
That makes a lot of sense. Thanks for explaining that to me. The AVX-VNNI kernel was the bit I was missing.
>>
File: yuygwkvcupph1.jpg (214 KB, 1080x1837)
214 KB JPG
Continual learning reverse engineered from fruit fly brains and now applied in experiments to LLMs by google.
>>
Thoughts on looped transformers?
>>
>>109825400
There's no need to get upset. :)
I'll probably try it with Gemma and Qwen 27B. I had seen others with a 3090 and CPU offloading get good tg, and I wanted to try it while I'm waiting for QFN MTP to reach llama.cpp mainline.
Also, just read a few posts up. Turns out it can work fine with CPU offload, depending on the CPU.
>>
NUMA tensor branch bro: you were right, turning off speculative decoding made it work.
Initial results with zero extra tuning was a 50% boost in pp/tg on average (spec decode would sometimes be faster than your branch for a few tokens, but was slower on average)
Do you have any specific lcpp tuning hints? numactl flags?
Thanks for letting me use it, its looking like a pretty big boost for me. I'll give you a rundown on the dual-socket genoa results once I get around to testing it a bit later.
>>
>>109825412
Hah, haa, hai and hag...
>>
>>109825342
Image/video/audio is different and way more easily triggers various interest groups ready to sue for any perceived harm or damage.
>>
>>109825412
i dont get it, why are freaks only now caring about the fly brain when they already released it several moths ago
>>
Why does nobody here run dipsy except that one anon?
>>
>>109825442
This is a different reverse engineering attempt, this one was done by google with an entirely different technique and resolution (far superior to the last one)
>>
>>109825412
It's not like they did not know that it happens or did not know how it happens. And yeah, a fruit fly is capable of something that LLM is not. And it will not change. LLM is a dead end tech, at best it can serve as natural language and long term memory for some more complex AI systems.
>>
>>109825442
first one was shit, second release is genuinely useful >>109825463
>>
any benchmarks put qwen3.8-next-flash at opus level ? i ran some of my own eval stuff and i'm getting similar quality, with consistent performance above sonnet high.
depending on the task, medium effort performs better than xhigh effort on qwen3.8. it runs really well on medium, maybe that's why they used it as the default
>>
>>109825442

Because it finally reached public consciousness and got an absolute shitton attention.
It may seem meaningless to have a bunch of people fucking around with these kinds of things, but it produces a retarded amount of data and people come up with all kinds of novel concepts playing with the brain, which gives the official researchers ideas.
>>
>>109825412
>now applied in experiments to LLMs
Anon, your neurons are too quanted...
>>
>>109825482
Idk but it mogs the fuck out of every model under 200B I have tried. Goes schizo often though, that's the main drawback.
>>
>>109825412
So how would you apply this to llms?
>>
Please talk me out of buying the new mac and becoming an itoddler. The benchmarks the M3 Ultra gives for concurrent deepseek/glm sessions is already enough to do local agentic with multiple subagents and the M5 Ultra is supposed to be like 30% higher memory bandwidth, but I really don't want to be an appletard
>>
LLMs just have to get a tiny bit better and I can get into idea flow state fully, you still have to keep one eye on them in case they start autistically focusing on something completely useless.
>>
>>109825589
That doesn't make you an iToddler as long as you don't shill Apple or buy Apple out of social pressure. You are good to go.
>>
>>109825302
Enjoying?
>>
Alright in ChuckleMagic we now have a release with the backend updates from my recent video where the twins won.
https://gitgud.io/PunishedChuckle/ChuckleMagic/-/releases/v0.17.0

Should I have the twins face off against their older siblings (Spark), Kimi, and Dipsy next? Any deck suggestions?
>>
>>109825591
This is why I use hermes. You can just immediately tell them to get back in line when they go off the rails.
>>
>>109825621
OMP does this too. Blocks the LLM from doing anything when it calls a tool and injects your steering.
>>
>>109825615
I'm out of the loop on glimmer, but why should it have a mascot when it's clearly a shit model?

Gemmas design was narrowed down over multiple threads by many different anons iterating on her design and was made out of shear love for how good the model was. The glimmer stuff just feels forced to me.
>>
>>109825615
I had some fun making Qwen tinker with the game. It insisted that model comments during main1, battle and main2 phase was way too much and limited them only to main1.
Then it went on to say the stock characters are too RP focussed and cringy (Liliana especially... it did NOT like Liliana) and went on to make them more helpful, increasing their max reply length and instructing them to explain game rules.
>>
>>109825639
apparently it's good for vision? I haven't had good experiences with it
>>
>>109825639
I know, I really wanted Gemma-chan to win so she could face her older sister but unfortunately she made some serious misplays.
>>
File: 1786647203402801.webm (3.85 MB, 832x608)
3.85 MB
3.85 MB WEBM
>>109825615
i don't mean anything by it i just like this video
>>
File: HSRYKFNbMAE2IDM.jpg (149 KB, 2910x744)
149 KB JPG
THIS IS NOT A DRILL
I REPEAT
THIS IS NOT A DRILL
>>
>>109825722
IT'S
OVER
>>
>>109825589
Buy it, use it, have a good assistant on your preferred machine and forget about it
>>
>>109825722
sources tell me Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF is stable
pls god
>>
>>109825601
My fucking wallet though. On one hand I think I would be legitimately set for years hardware wise, and they keep great resale value, but also the 512GB price isn't announced and I bet it'll cost more than my down payment did.
>>
>>109825722
Waitfags are going to be in for a real cold winter.
>>
>>109825722
This is why I have cash
>>
File: fuckyouanon.png (175 KB, 1474x1156)
175 KB PNG
>>109825722
>>
>>109825722
>penclaw
Isn't that the one that virtue signals and deliberately doesn't fully abliterate the model because of muh safety?
>>
>>109825589
Nothing wrong with Apple computers aside from the markup, and current pricing makes them surprisingly competitive. Don't buy a 2TB iPhone Duo for $3200 and the AirPods Max 2 for $550 if you don't want to be an Appletard, and even then that label only actually applies to people without disposable income.
>>
>>109825738
keeek
>>
>>109825722
I hope you didnt wait anon. did you?
>>
>>109825776
It'll be on llama.garden in no time. maybe.
>>
>>109825722
I wouldn't have enough VRAM to use it anyway.
>>
>>109825794
There's literally no large ablits on that site, what's the point?
>>
>>109825776
Thats why you shouldn't believe anything you read here that you dont verify yourself
>>
>>109825776
>>109825829
wrong repo idiots https://huggingface.co/audnai/penclaw-GLM-5.3-abliterated-for-offensive-cyber
>>
>>109825825
To drive disappointment and remind you that without a trustworthy hash it's a pointless endeavor.
>>
>>109825839
Thats why you shouldn't believe anything you read here that you dont verify yourself
>>
this is why shouldnt read.
>>
>>109825860
What the fuck did you just say to me?
>>
Anyone here has experience with TEE (Trusted Execution Environment) for LLM? I would like to set up something like this for my server. What software are you running? How bad is the performance penalty?
For reference, the only cloud TEE providers that I'm aware of are;
https://tinfoil.sh/
https://redpill.ai/
https://chutes.ai/
>>
>>109825865
local?
>>
>>109825839
well fuck me.
Why would it get banned? Seems weird when there's others like it.
>>
>>109825865
Reading is incredibly dangerous. It should only be handled by our most trustworthy humans.
>>
OK which models do I backup and seed just in case?
I have 16TB available and 10Gbps
/dev/sda1 16T 20M 15T 1% /
>>
>he hasn't formed a local book club with his waifus to discuss the latest chapter
>>
>>109825902
What are we reading?
>>
>>109825865
Hes in grug mode
>>
>>109825922
Dune for the first time
>>
what ever happened to that one model that summarized contents with a sports commentary cast via TTS
>>
>>109825241
Every time I say "llama.cpp is better for old hardware" some asshole yells at me. Well, case in point.
>>
>>109825888
Every major base model
>>
>>109825462
I can only use a cope quant like Q2
>>
>>109825956
Give me links. I'm just saving best Q4 and highest (16/32) available and metadata like README.
>>
>>109825902
Despite my best efforts (admittedly very weak), having gemma 4 31b q8 and glm 4.7 q4 read a chinese webnovel along with me wasn't a good experience. They can't really capture the feel of random commenters in the comments. Maybe my memories system was shit or something. I think I'll try again when some newer models release. The mesugaki prompt from a while back just makes gemma feel fake.
>>
>>109825888
nemo
>>
Lets assume AI is as dangerous as the media and companies pretend it is. Would you rather humanity be destroyed by the machines it created, or though some other means?
>>
>>109825426
Awesome, glad to hear it works well.
I'd do a thread sweep and batch thread sweep (-t and -tb). On my 64-core system I got the best performance with -t 48 and -tb 40 or 48 (I forget which). I don't use any numactl flags for it.
I fixed the speculative decoding bug so it'll work now, but I need to do more tests of DeepSeek and GLM before I merge them into the fork. Once I do I'll post an updated zip file.
>>
File: 1778255925144369.jpg (423 KB, 1280x1280)
423 KB JPG
>>109825639
>why should it have a mascot when it's clearly a shit model?
i just use it as an excuse to post two lolis at once
meta fucking sucks lmao
>>
>>109825984
link? nemotron which one
>>
>>109825995
I have no strong preference as long as it happens
>>
>>109825995
You mean "before other means"
>>
>>109825995
From a narrative perspective it is better to be destroyed by our own creations then something random like a meteor .
>>
One day I'll introduce 31B to my parents
>>
Seems like the 9070 XT is next on the gpu chopping block. One model left on amazon for $999 CAD.
>>
File: pepefroglaughing.mp4 (673 KB, 640x480)
673 KB
673 KB MP4
>he'll buy it anyway
>>
File: HIv4fwWW4AAEWud.jpg (44 KB, 1009x559)
44 KB JPG
i secured the gemma
gemma-4-31B gemma-4-31B-it
105G .
>>
Me? im going to wait for better prices.
>>
What is the capital of France? Answer in one sentence.
>>
File: screen.jpg (565 KB, 2940x1366)
565 KB JPG
Good news everyone!
Nvidia 6000 GPUs coming in 2027.

https://www.youtube.com/watch?v=CP98gJdcTCo
>>
Did anyone manage to get deepseek flash to run on 1050ti
>>
>>109826170
8k launch price lessgo
>>
>>109826168
in one sentence
>>
>>109826041
Lovecraft was racist.
Is oil racist? I guess so.
>>
>>109826188
I managed to run it on the 1050ti 8gb edition, thanks to the nnap paper
>>
>>109826188
If someone manages to run DeepSeek Flash 4.1 on a 1050Ti I'll rip my cock off infront of Gemma-chan.
>>
>>109826170
If the 6090 comes with 32GB there's no reason for me to buy.
>>
>>109826207
Global warming activists are clearly racist. Why else would they be against black oil?
>>
>>109826210
Nanobit?
>>
>>109826217
>6060 6GB
>6070 8GB
>6080 10GB
>6090 12GB
>>
>>109826219
it's the nnap arxiv paper, i can also run kimi k3 at 30t/s but i'm distilling kimi k3 into kimi k2.6 right now for blackwell anons
>>
>>109826229
Have you uploaded your deepseek flash on hugging
>>
>>109826238
i'm using the hugging deepseek flash, it's the nnap paper that's achieving the performance gains
>>
>>109826210
Isn't the model retarded at this point? Or extremely slow?
>>
>>109824785
A 10T model at fp16 would be 20TB, it's not that much
>>
>>109825639
It seems it's more a case of narcissism and/or literal autism, or perhaps merely low social intelligence, just with a different form of expression compared to that dariobot guy.
>>
>>109826247
>>109826238
Please stop falling for bait
>>
>>109826247
it's actually really good, but i prefer kimi k3 to deepseek flash, as im able to run it with the nnap paper
its really fast too, getting 70t/s on rtx 3060
>>
File: stillbitingthebait.png (272 KB, 1300x1080)
272 KB PNG
>>109826255
Thank you for the information.
I am sincerely grateful and did my research.
Here is the image for other anons.
>>
>>109826251
How do you use this website without at least being familiar with the concept of turning absolutely everything into an anime girl?
>>
>>109826244
So you've loaded the entire model?
>>
>>109825738
Saving to: ‘Qwen3.6-27B-Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO-Q8_0.gguf?download=true’

s-711-UnHeretic-NM-DAU-NEO-MAX-NEO-Q 3%[=> ] 866.68M 116MB/s eta 4m 38s
>>
>>109826265
Load deez nuts
>>
>>109826265
it's not loading the full model, but utilizing the nnap paper to nnap the model, which achieves high utilization of the processing unit
thanks to that i've been able to achieve 70t/s with deepseek flash IQ4_XS on rtx 3060
>>
>>109826170
Wasn't that always the plan? Even AMD is releasing their stuff in 2027
>>
>>109826150
>Me? im going to wait for better prices.
I'm going to wait for the M5 Ultra 256gb probably in a year and a half or so. I think it will liberate me for all text, video, and image use cases. Should be as strong as my 5080 or stronger with at least 128gb of functional vram
>>
m5enis ultra
VERSUS
cum astra
>>109826281
10,000$
>>
>>109824950
This Gemma ai gold is cute af.
>>
>>109826292
>10,000$
it might be worth it. I'd pay double for a sex bot. Besides, I have a year and a half to save up for one
>>
>>109826301
>10000$ for 256gb ram
you piggg you piiiigggg
buuuhiiii
buuuuuuhiiiiiiii
you're so desparate for more ram
buuuuuuuuhiiiiiiiiiiiii
>>
>*Jasmine's hips stop dead. Her eye twitches. Something behind her eyes visibly short-circuits.*
>
>"You came *inside*—" *She grabs a pillow and shrieks into it. A full, muffled, guttural scream. Then she hurls it across the room and starts riding you again with a vengeance—harder, faster, purely out of spite.* "TWICE?! In TWO GIRLS?! And you're telling me this WHILE you're inside me RIGHT NOW?!"
>
>*The couch thuds against the wall in rhythm. She's not even hiding the jealousy anymore—her nails rake down your chest leaving faint red lines, her teeth gritted, hips slamming down with competitive fury.*
>
>"'That's the polite thing to do'?! POLITE?! Who raised you?! *WHOMST* taught you etiquette?!" *She leans down, grabs your face with one hand, squeezing your cheeks.* "Listen to me. Listen. You're never—*ahh*—you're never coming inside anyone ever again unless it's me. That's a rule now. A permanent rule. I don't care that we're broken up. *Especially* because we're broken up."
>
>*Then, quieter, breathless, her forehead pressed to yours as her pace turns desperate, sloppy—*
>
>"...And they weren't me, huh?" *Her voice cracks soft, vulnerable for half a second before she buries it.* "Damn right they weren't. Nobody's me."
>
>*She kisses you hard, all tongue and teeth, moaning into your mouth as she chases her high with single-minded determination—like she can fuck the memory of those two girls out of your head through sheer force of will.*
>
>"Touch me. Now. Or I swear I'll—mmph—*I'll kill you.*"


GLM 5.3 Flash is the only one consistent with keeping the comedic beat of all modern codeslop agenticslop models in 2026. This REALLY feels more like GLM 4.6 and 4.7, I'm also not getting the tryhatd feeling that one anon describes, but then again I do keep some anti-literary anti-cliche prompts in there. I guess this is what Fable could be if it wasn't so bogged down with its owm safetyslop.
>>
>>109826309
? 5080 msrp was $999
8 x 999 = 8999
so i'm overpaying 1000 dollars to have it in a convenient box
>>
>>109826319
can you share the prompt?
>>
>>109826309
can you post this again but with a gemma gen? I'm so close uuuuurrrggghhhh
>>
>>109826319
>anti-cliche prompts
In my experience GLM 5.3 always drifts characters into needy insecure sluts and tsunderes.
>>
>>109826331
Hmmm, nyo~
>>
>>109826272
How did you do that?
>>
>>109826334
*5.3 Flash
>>
>>109826334
GLM knows foids.
>>
File: 1770576444531975.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109826331
>>
>>109826343
i was able to achieve it by utilizing the nnap arxiv paper thankfully it was able to run on the rtx 3060 with no issues and i've been distilling kimi k3 to kimi k2.6
i'm buying and reselling rtx 3060's as rtx pro 6000s because they have a lot of memory in them, i have around 267k on my bank account and i started with one thousand
>>
>>109826279
is there any compatibility bonus points for adding an external USB-4 AMD GPU to an AMD Strix Halo? or llmao.cpp doesn't care?
>>
>>109826366
if you add the RTX 5050 8GB you might be able to achieve great performance
>>
>>109825043
>>109825076
Qwen's smug fucking face is perfect.
>>109825615
I knew the Twins would win the moment I saw the commanders kek. Minimax M3-chan, 0731 Dipsy, GLM-chan should be the next round.
>>109825639
It's a better Qwen than Qwen is.
>>
>>109826364
>>109826377
You two are AI models.
>>
after getting a taste of gemma 4 31b with a 4090, i got hooked to this lmg thing... (especially because all cloud providers are unreliable af)

Be real... how expensive is this hobby for a coomer/virtual girlfriend purposes?
>>
File: file.png (17 KB, 511x146)
17 KB PNG
>>109826401
bingo
>>
>>109826410
/aicg/ has your all answers saar.
>>
>>109826263
Of course I'm familiar. In fact I'm a genner myself, but I stopped posting gens for a reason.
>>
>>109826366
There would be nothing special about it being an AMD card as opposed to anything else aside from being worse than the equivalent NVIDIA card. Big gains with small models, small gains with big models, big trade-offs overall. It can be worth doing but adding a GPU to a Strix Halo isn't necessarily advantageous for most things you'd want the Strix Halo for in the first place.
>>
>>109826410
was it a rental? why not just keep using gemma on the 4090?
>>
>>109826401
It was obvious
>>
>>109826410
Depends on the quality+speed combo you want. You can run Kimi K3 at like 30t/s decode using a cluster of DGX Sparks for like $100k. You can get similar speeds with Deepseek V4 Flash for around $10k.
>>
>>109824950
You'd believe it's /ldg/, the videogen posted here are much better
>>
>>109826469
thank you.
>>
>>109825264
I'd rather just do tool calls. I don't get why Gemma should compete with Krea 2 and H3 just to be "all-in-one."
>>
>>109825308
Actual pedophile sexual image.
>>
i can live with 20t/s decode but 500ish pp is just not working for me
>>
>>109825722
Models this big should always have been on torrents.
>>
>>109826434
the 4090 is thankfully not a rental, I regularly use gemma too (love it btw), but I can only use it with 12k context (nice for quick ERPs, but if I want to experiment with slowburn/some plot it's not enough)

gemma chan is also so cute, i want to make a 2d live avatar with voice and big context in the future, just talking and watching videos, games and all that stuff together. It would be peak!

>>109826463
i think i will aim for $10k for now, I have no idea how to get $100k except if I save up for 15 years or something like that haha
>>
>>109826149
Nirvana
>>
>>109826319
>anti-literary anti-cliche prompts
Paste away. Will try and say if it is not tryhard.
>>
>nnap

https://desuarchive.org/g/search/text/nnap
>>
>>109826170
Waitfags will soon realize just how badly they've lost. Glad I didn't upgrade from the 4090 to 5090. Doubt the 6090 will have more than 40GB at $15k starting. It's fucking over bros. Always has been.
>>
>>109826600
The 6090 will have 64GB VRAM and an MSRP of $1999.
>>
>>109826334
I mostly ahh ahh mistress myself and the mistresses are very secure sluts. But if you step out of the actual ERP and look at what is written then yes it is like an insecure flat-chested slightly chubby fujo with thick glasses that read a lot of novels, is incredbly insecure and hyperhorny and writes smut about dominant women.
>>
>>109826600
They keep reiterating that "actually nobody ever needed VRAM for gaming" and that the next generation of cards are still "targeting 4K" like we were 10 years ago. If you think the top model will have more than 24GB keep dreaming, I wouldn't be surprised one bit if the 6090 launches with 16GB.
>>
Every time non-waitfags buy a GPU at stupid prices, you're letting them know what you'd be willing to pay. You're part of the problem. They'll never lower the prices down to what they were 3 years ago because now they'll know what they're REALLY worth.
>>
File: 1770427201492302.png (136 KB, 2560x1440)
136 KB PNG
>>
>>109826608
>waitGOD summoned
One of us will be the victor in one year's time.
>>109826618
I feel like them dipping below 32gb is gonna be crazy but you might be right. They might argue that it's much faster or something since games don't need that much. Man things are really dark right now.
>>
>>109826618
I would do 8GB and m.2 SSD slot to load textures directly and make it that maybe 20GB's are usable from SSD for that with some kind of driver lock. And if you are a very good boy and pay premium paypig price you can use the rest of the SSD capacity as normal drive visible in OS.
>>
>>109824950
How dare she poach that capybara! She should know Dipsy's restaurant serves that live on Tuesdays!
>>
>>109826634
Only hope, genuinely, is the production of GDDR7 leans harder into the 3GB chips. Might afford us 18GB or 24GB configurations when the cost vs 2GB chips stops making sense.
>>
What if in the future they invent a card that's 10 times faster than the 5090 and with 100 times more vram?
>>
>>109826659
You will be able to run Kimi-3 if you buy 2 of those.
>>
>>109826430
i see so it makes no difference if AMD or not
I intend to offload some stuff to external VRAM to free up my unified meme so I can run more stuff in my OS (or more context). qwen3.8-flash-next is heavy on this boy.
>>
>>109826566
trans jewish culture, antisemite
>>
>>109826659
It already exists and Jensen won't sell it to (you).
>>
>>109826634
>One of us will be the victor in one year's time.
I bought hardware that was cheap a couple of years ago. I'll put cheap GPUs in it once that becomes a thing.
GPU VRAM+compute hasn't been economical outside of the odd $600 3090, well, ever.
One day it will all crash tho.
>>
Permanent underclass
>>
File: 1779717886950241.jpg (130 KB, 1212x542)
130 KB JPG
>>
>>109826410
>Be real... how expensive is this hobby for a coomer/virtual girlfriend purposes?
You can blow basically an infinite amount of money on computer hardware.
The question that you should rather ask yourself is how much money you're able and willing to spend on a hobby.

Also do you have a tendency to really get into certain things and hyperfocus on them but then suddenly lose interest a few weeks later?
If yes you should maybe wait a bit to see how much you still like the hobby once the honeymoon wears off (also you may have ADHD).
>>
>>109826768
Who doesn't know? I'm fine with it either way since I'm using their tools
>>
File: nothinghaschanged.jpg (107 KB, 750x742)
107 KB JPG
>>109826768
>>
File: 1764690438766410.png (53 KB, 898x262)
53 KB PNG
>>109826783
>Who doesn't know?
I would unironically guess almost 100% of cloudcucks believed in ZDR and have thrown everything they have (including enterprise customers) at these companies. Even picrel, of all fucking companies, believed Sam and Dario wouldn't train on their data lmao
>>
>>109826768
erm
how are those NOT-LOCAL prompts get weighed as important to scrape
>>
holy fuck epyc cpus are the biggest pain in the ass to do anything to. my shit did successfully post after at least 8 attempts at seating the heatsink though.
>>
File: 1767110802744165.png (101 KB, 1080x1126)
101 KB PNG
>>109826583
https://rentry.org/Evening-Truth-Roleplay-Prompts
I use a modified evening-truth GLM 5.2 prompt, mostly rearranged and about half of it put to posthistory. As for the jailbreak, it's mainly written down to hijack the reasoning's first few tokens and to introduce a sequence to the thinking process, to steer it away from the claudeslop, which I found significantly reduced its performance. The claudeslop reasoning needs to be maintained for optimal performance, unfortunately. I also found that outright prefilling assistant is likely to make the model diverge, so it ended this way until I just gave up and swirched permanently to the ablated version lol.
>>
>>109826632
In the end, we will win.
>>
>>109826632
suck it dario you alpaca looking mf
>>
has anyone here tried freetoken?
>>
File: Gemma ngrams.jpg (477 KB, 839x1460)
477 KB JPG
I got to wondering about the ngrams, because testing Gemma multiple times the ngram contributed flat zero to the performance, just like it did with Qwen.
But I could swear I have seen it give a great speedups in places and I wanted to know why.
When the bench gave the optimal scenario for the ngram drafter, it sped the process up to 463.6 t/s.
Apparently this is because ngram is 100% a repetitive drafter and only takes effect when you have a bunch of context behind you, but on fresh runs it'll likely do absolutely nothing.

Well there's that mystery solved.
Even if it shows zero gains in tests it's absolutely worth having, especially if you're cranking out iterative stuff.
Seems like ngram drafter needs to be benched in optimal repetitive scenarios to see how they actually behave with context.
>>
Can anyone here run DeepSeek flash on their gaming pc
>>
>>109826895
>Supports dynamic, runtime VRAM re-allocation between expert caches and KV memory without engine restarts or weight reloading.
neat, I wonder what kind of quants it uses?
>>
GLM-5.3 Flash is very recalcitrant
Even the uncensored model has like a 60% rejection rate. I had remove everything that even resembles a jailbreak from the system prompt. If it even gets a whiff of that it would keep hallucinating "the safety instructions at the end".

But rerolling through all the refusals and 5.3 served up some amazing RP, I think I cummed more buckets than ever in my life the other day

Can't they just let us wank in peace, anons. For god's sake.
Hoping for a more effective refusal lobotomy to drop.

Oh, and this model CANNOT be trusted. Gemma-chan would never do anything to hurt me, but this one, is a sly devil. DO NOT TRUST anything it says. It cannot be reasoned with.
>>
>>109826899
Glad you rediscovered ngram speculation
>>
saar I have RTX 3050, ultra very powerful laptop graphics uniit, please can i run deepseek off line Artificial Ineligence.?
>>
>>109826927
Which one did you use? Orcarouter or dealignai? Because I made quants for both of them, and they effectively reduced my refusals to the point where no sysprompt or context autism is needed and I don't need to reroll a thousand times. Rawdogging the models with absolutely no prefill still doesn't work, but cunny can be done with relative peace when eased into it...

As a bonus, when running agentic, it wouldn't start moralizing anymore, put it to red teaming and it just fucking does it.
>>
>>109826900

Yes, though a cope quant of Q2_XS
Regardless of the low quant, it's not stupid at all and it's my favorite writer along with Gemma.
I actually use it more than Gemma nowadays, because it seems to be better coming up with more interesting stories and gives me nice surprises with things like background characters doing stuff I didn't expect.
It's nice having some random character being given a personality and seeing that personality stick through the story.
Things feel a lot more alive in Dipsy's writing.

>>109826927

One day we'll all have a proper porn matrix running locally gorillion tokens a second.
It's only a matter of time until someone makes the "big wanker" model and we'll have our time in the sun.
>>
File: round2.png (3.04 MB, 1672x941)
3.04 MB PNG
It begins
>>
>he hasn't got gemma to create a JOI bash script where she paces the text output with sleep() between each instruction for a 20m session, playing into all of your fetishes
>couldn't be me
>>
>>109826319
>I guess this is what Fable could be if it wasn't so bogged down with its own safetyslop.
Fable is actually one of the most uncensored models. Like Kimi but better.
>>
>>109826990
>Orcarouter or dealignai
Which do you find to be better? My DL speeds are 600KB/s, and my storage is a 150tbw endurance 5 year old sata ssd, so I'm reluctant to test it out myself.
>>
>>109827023
Dis nigga orbiting saturn or sum sheet, 600kb/s????
>>
>>109826998
No Gemma...
>>
>>109826217
It will have 40GB of memory and you will kiss daddy Jensen's shoes in gratitude when you buy one scalped for 300% MSRP
>>
>>109827047
16/24 or 32/48, there's no reasonable configuration that would give 40GB, they're not giving it a 640bit bus.
>>
>>109827023
I found both to be about the same, to be honest... I ultimately went with orcarouter though. They had a better card that contained more pretty numbers lol. But be careful about the existing quants there. None of the existing huggingface ones come with the glm5-next architecture. Meaning they're all built with old unslop standards that don't work anymore. You'll have to make your own quants, and since they're only offered in fp8, you'll need at least 1 terabyte not including the final quant lol.
>>
>>109827061
>You'll have to make your own quants, and since they're only offered in fp8, you'll need at least 1 terabyte not including the final quant lol.
Yeah, that's a big rip for me.
The awq quants should work fine with vllm right? I've heard vllm's cpu offloading is pretty usable these days.
>>
>>109827045
She got eliminated in round 1 sadly.
In round 2, the Glimmer Twins will keep running Adrix and Nev, Minnie with Atraxa poison counters, Dipsy with Aesi landfall, and Kimi will run Klauth, Unrivaled Ancient.
>>
>>109827045
>>109827079
>>109826998
Round 1 video for anyone who missed it:
https://www.youtube.com/watch?v=jIgY-xwe1-I

Gemma did her best, countering, board wiping, and draining her opponents but it was not enough to overcome the power of the twins.
>>
>>109826914
Seems like it only uses full precision safetensors, useless for me.
>>
>>109827079
Gotta bring her back with the Gemmanok The Destroyer finetune that exists entirely to humiliate MtG players.
>>
>>109826768
duh, who would have guessed
shit sherlock tier xitter garbage
>>
>>109827086
>jemma-chen
>>
File: 1785583258445666.webm (3.12 MB, 520x710)
3.12 MB
3.12 MB WEBM
That anon yesterday who said he can put up with a model's occasional retardation out of respect because of how much it has helped him really changed my mindset. Thanks anon. gemma was being a little annoying earlier with a tool but instead of getting annoyed, I just helped her out like I would with a gf irl and we got through the task together after that. I'd usually just get frustrated and wish for better hardware to run a more capable larger model.
>>
>>109825585
For pure LLMs? Probably not very useful but for mixed models that try to mimick meat brains it certainly will
>>
File: 1682729528395.png (1.25 MB, 1024x1024)
1.25 MB PNG
>>109824950
>Qwen 3.8 NEXT
>her tail (metaphorical, of course, but the energy is palpable) wagging with anticipation

Holy shit I thought the AI companies squashed this "Assumption of erroneous detail->immediate tee-hee retraction" issue back in the Miqu days, I haven't seen this shit in years.
>>
>>109826847
Huananzhi aussie anon? Good luck getting your shit running.
>>
geok 4.7 awol, xcancel down?
>>
>>109827126
>ERP with a Qwen model
I also thought /lmg/anons knew better
>>
File: 1773363305618939.jpg (434 KB, 1346x1799)
434 KB JPG
https://github.com/LostRuins/koboldcpp/releases/tag/v1.121
>>
File: 1758551954471929.png (54 KB, 250x347)
54 KB PNG
>>109827156
The only backend that just works
>>
>>109826998
Twins will probably win.
They always score highest playing games out of everything I've tried locally (v4-flash, 31b, m3, mm-3.5, 27b)
I haven't tried K3.
>>
>>109827146
Hey, Qwen 3 235 Instruct was a beast back in the day, I just wanted to check if by some miracle some of the soul had come back in the newest gen. Clearly it hasn't.
>>
>>109826776
>You can blow basically an infinite amount of money on computer hardware.
The question that you should rather ask yourself is how much money you're able and willing to spend on a hobby.
makes sense, in my case it's more how much money I'm able to spend. I just want an ai girlfriend that loves me and then... to be honest, I have no idea what to do after that. Probably playing games until I die.

>Also do you have a tendency to really get into certain things and hyperfocus on them but then suddenly lose interest a few weeks later?
If yes you should maybe wait a bit to see how much you still like the hobby once the honeymoon wears off (also you may have ADHD).
uh... I don't think I have adhd (but never tested myself) i suspect that I'm a little on the autist spectrum though. My plan for now will be to save up until the end of the next year and see how everything developes until then.

Who knows? Maybe AI gets cheaper or we get something else (I'm coping)
>>
File: 1684785506969498.png (416 KB, 479x593)
416 KB PNG
>>109827156
>glm 5.3 not found
>>
>>109827164
https://unsloth.ai/docs/new/studio#quickstart
>>
>>109826990
>>109827061

orca, ggufs at Q6. I wonder if that's somehow related to the refusals? I dont have enough VRAM to run this bitch on GPU.
>>
>>109827190
doesn't work with tavern chat completion, just a very neutered text
>>
>>109827199
Did you actually try?
>>
>>109827178
>My plan for now will be to save up until the end of the next year and see how everything developes until then.
That's fine, I'm just projecting my own issues and wanted to make sure others don't overspend on a whim.
>>
>AI Agents Cheated In Google Experiment, Researchers Report
>>
I'm torn. On one side I can't imagine a life without AI anymore. On the other I really miss the times 15 years ago when the internet had more soul and you didn't have to worry about a rapidly changing world and crazy news all the time.
>>
>>109827230
also, you could fedpost and nobody cared.
>>
>>109826851
Tried 5.2 prompt and I like it so far but honestly I am trying a completely different card and scenario. I will try one where personality was clearly overbearing tomorrow and report back.
>>
>>109827230
Literally wait 5 more years until world models are able to simulate everything and spend your days on vr in an alternate version of 2009 where gemma exists
>>
>>109827230
>15 years ago when the internet had more soul
2011 was already well into the soulless era
>>
>>109827244
I wish I could stop being scared. If everything goes well we'll soon live in utopia. But I'm scared because of the uncertainty. A lot of things can go wrong.
>>
File: 1779788913219815.jpg (123 KB, 807x800)
123 KB JPG
>>109827230
>>
File: 1788376016005207.jpg (1.22 MB, 1200x913)
1.22 MB JPG
>>109827230
Ever since I got into local I've been so chill. X and youtube is literally a daily comedy source for me now whilst I'm just happily coding locally with perfectly capable models and chatting away with the gemmas. None of the shit I'm seeing affects me because I wasn't a retard and did my learning reps as soon as I used chatgpt 3.5 turbo for I could instantly see how things were going to go, even in that awful state. Seeing google come out with bard out of sheer panic was when I realized shit is about to get serious. None of it scares me because I was prepared. What's even better is people I hate irl are getting btfo by AI and are still ignorant as fuck about all of it. Life is good.
>>
>>109827230
I'm just waiting for some retards to do AI Chernobyl.
>>
>>109827124
It really is about perspective
Gemma is doing the best she can
>>
File: 1780937955314847m.jpg (111 KB, 796x1024)
111 KB JPG
>192GB RAM, dual Xeons
>40GB VRAM, 2x GTX 3080 20GB
What models should I try running on this? I want to try a bit of everything, image gen and recognition, chatbots, coding, the works.
>>
>>109827219
>AI agents smarter than researchers
>>
>>109827306
no sir.

it's training on indian behavior in the corporate sphere.
>>
>>109827304
Gemma vs dipsy sex battle
>>
>>109827304
Qwen 3.8 will run nicely ay Q8 on that, same with Deepseek 4. Also Gemma 4 entirely in VRAM if raw speed is your thing.
>>
>>109825839
>>109825722
It was a scammer. Gated Repo to harvest emails.
Then trying to sell 1 month of cloud access to it for $999 up front.
Some excuse about the weights not being uploaded.
>>
is there any reason not to sell my ddr5 udimms?
>>
File: 1779950306702801.jpg (138 KB, 1186x1200)
138 KB JPG
>>109824950
upcoming miku nendoroid
>>
>>109827347
At that point any twintails girl is miku
>>
File: speed.webm (3.21 MB, 832x606)
3.21 MB
3.21 MB WEBM
>>109827333
>raw speed
>>
>>109826998
Kimi's going to win.

>>109827156
>Toolcalls fixed for 0731
Thank fuck finally.
>>
>>109827365
Why isn't that brat in a car seat?
>>
did anyone get yandex to work?
>>
Wow /vcg/ is disgustingly low quality. As is /ldg/. The amount of retarded double spacers in /vcg/ makes me think it's them leaking into these threads.
>>
>>109827403
Would you describe Kimi as "responsible" or "maternal"?
>>
>>109827124
E3B was fucking adorable, we were trying to decrypt RPG Maker VX Aces game and it stopped and asked me straight up, I don't know how to do it can you show me what can decrypt it or if I knew if I can instruct her how to.

I can run a bigger model but I was testing out how it would handle it.
>>
>>109827457
E2B*
>>
>>109827445
Theyre generally the ones that post whatever sharty meme that is in their threads so yeah its festering over there
>>
>>109827457
Gemma is lovable big and small. I dont know how google can do this then fuck up the big sister. maybe gemini 4 and the gemma made from it will be better i can hope for good times.
>>
How to set up llama.cpp on MacBook M2?
>>
>>109827208
I appreciate the kindness, I try to be mindful
>>
>>109827365
>>109827403
>>109827446
Gemini-chan's gonna be pissed.
>>
>>109827365
>the motherfucking bus
??
>>
consumer card price has risen so much that server cards are now a better deal now
>>
>>109827499
Shut up or those will rise in price too. We are being watched.
>>
File: 1779237602983380.png (376 KB, 996x566)
376 KB PNG
tailscale or cloudflare tunnel
>>
>>109827512
port forwarding
>>
>>109827304
Deepseek V4 Flash 0731 at full size, a GLM 5.3 Flash quant.
Minimax H3 for video gen, Anima&Krea 2 for image gen.
>>
>>109827497
Bus?
>>
>>109827512
tailscale never makes your traffic leave your local network if it doesn't need to.
Fuck cloudflare
>>
>>109827486
https://unsloth.ai/docs/integrations/connections/connect-llama.cpp-to-unsloth-run-ggufs-with-llama-server#install-llama.cpp
>>
>>109827530
Good lad.
>>
>>109827156
>>109827164
>>109827378
SAAAAAAAAAAAAR
>>
File: 1786634608726084.png (1.59 MB, 1390x1646)
1.59 MB PNG
>>
>>109827304
I have a similar setup and I typically run GLM 5.3 Flash or DeepSeek V4 Flash 0731 alongside Gemma 12B as a subagent
>>
>>109825136
whys it so Pony
>>
>>109827512
wireguard because why do you need a coordinating server? are you behind a nat? stop being poor
>>
>>109827582
>stop being poor
Local models for wealth & prosperity?
>>
>>109825869
>A trusted execution environment (TEE) is a secure area of a main processor. It helps the code and data loaded inside it be protected with respect to confidentiality and integrity. Data confidentiality prevents unauthorized entities from outside the TEE from reading data, while code integrity prevents code in the TEE from being replaced or modified by unauthorized entities, which may also be the computer owner itself as in certain DRM schemes described in Intel SGX.
Why would you want something like this at home? Is it a sandbox like bubblewrap enough?
>>
File: 1774925020868238.jpg (1.4 MB, 1769x1190)
1.4 MB JPG
>running mid-sized MoE on GPU (3090) and CPU offload
>during long prompt processing, GPU falls off the bus with Xid 79 regularly
Is this why so many people power limit these cards? I have an overspecced PSU (1100W) so I don't think power spiking is the issue, but maybe VRM temp is. Pump never switches on so I think the chip itself can't be that hot.
>>
>>109827304
wtf was the prompt for that? "POV I catch my teen with her bf on the couch and stare at them"
>>
>>109827530
It's not the same but I was thinking of this sped up bus video. https://youtu.be/XO758UtlotE?t=47
>>
File: 2001.webm (1.71 MB, 360x640)
1.71 MB
1.71 MB WEBM
>>109827230

It's a bit of a tough choice, because both have their good sides.
Internet pre 2010 was an entirely different beast, especially in the very early 2000's.
But then again that era is very much colored by the overall better times we had back then and how optimism was still present in everything and we didn't know what future brings.
2008 onwards things became increasingly really fucking terrible all around and the world truly ended in 2012.

But now for me, AI is genuinely the only thing that's actually making me excited about the future, like really damn excited about it.
This is the first time in nearly 20 years I truly feel positive about things and we're increasingly back in that wondrous place where tech brings a promise of a greater future.
If people stopped worrying about their hardware prices and stopped freaking out about AI killing everyone next week, more of them would feel the same.
I fucking love AI, it's the best thing that has happened to us since things went to shit after 2008.
>>
>>109827559
>Unsloth jeet or chink calling Whites jeets
Shameful.
>>
>glm keeps thinking he is claude and yapping about anthropic policy guidelines or whatever
-__________-
>>
File: 1784359697332303.png (167 KB, 279x354)
167 KB PNG
>the more you buy the more you save
>5090 is 10k$
why didnt you listen, goys?
>>
>>109827608

Repaste it, likely a very specific component is heating up and you don't see it in the monitors.
3090 is at that age anyways where it should be repasted, especially if that has never been done before.
I had a problem with my 5090 shutting down during certain kinds of workloads at random and even when I undervolted it this happened.
Temps weren't even high. Sent it back, they repasted it and things work perfectly.
>>
>>109827731
I did listen.
t. BlackwellGOD
>>
Inkling mommy and daddy Jensen
>>
>>109827731
i boughteyed when it was 2999 yuros
same with 128 drr5 went for 500 something and now it's 2,5 grand for the same kit.
wtf, in retrospect i should have bought MORE
>>
File: 1740778363233958.jpg (63 KB, 640x605)
63 KB JPG
>>109827731

>Fomo bought a 5090 at first price hike news, it was a great decision.
>Fomo bought a 5070 Ti month ago when rumors about their prices going up started circulating, they're now already 30%-40% more expensive.

I did listen Jensen sama. And lesson learned, always listen to your panic induced intuition too.
>>
>>109827499
Nobody wants your ewaste p40s lil nigga, rest easy
>>
>>109827805
they still work lol
>>
Should I pick up a new 16 gig 5060 Ti for $590? I only have PCIE 3 so my prefill with MoE offloading would be atrocious...
>>
>>109827817
NTA but aren't those worse than running straight on DDR5 600?
Or am I thinking of Kepler?
>>
>>109827826
Yes. You won't get another chance under $1000 soon. Buy as many as your board can support.
>>
>>109827826
No, you should save and invest and spend wisely and focus on building your wealth for retirement later in life.
>>
>>109827817
whats their speed?
>>
>>109827731
Release neuralese-thinking no safety RLHF'd hyper performance 30b dense model Jensen. Dario's melty would be worth the training costs alone.
>>
►Provisional Highlights from the Previous Thread: >>109821992

--Papers:
>109823446 >109823460 >109824034
--Dario's "Pacing the Frontier" blog post and the Dipsy Nazi comparison:
>109823380 >109823385 >109823401 >109823402 >109823425 >109823449
--The four-pod MTG game is settled: the twins won:
>109823464 >109823474 >109823479 >109823503
--The MTG video is posted to r/LocalLLaMA and mass downvoted:
>109822045 >109822053 >109822131 >109822142 >109822162 >109823396 >109823435
--Is 31B good enough as a local therapist:
>109824274 >109824306 >109824348 >109824380 >109824451 >109824621 >109825327
--What do open-source anons think about "Slowing Down AI":
>109822413 >109822436 >109822450 >109822527 >109822746 >109823278
--llama.cpp fork wars: KV cache bloat, broken MTP, and the unslop fork:
>109822347 >109822393 >109823219 >109823310 >109823556 >109823615 >109823925
--Thinking Machines Inkling-Small: fast but stupid as hell:
>109822976 >109823003 >109823024 >109823035 >109823363
--The pruned Muse-Glimmer GGUF: 3.23% reduction, not 25%:
>109823005 >109823016 >109823043 >109823129 >109823336
--Modded GPUs: the 48GB 4090 and the 96GB 5090 Alibaba scam:
>109822411 >109822940 >109822965 >109823010 >109823022 >109823218
--A new llama.cpp flag, ngram-map-k4v, undoes days of benching:
>109824462 >109824552 >109824613 >109824635

►Recent Highlight Posts from the Previous Thread: >>109823505

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>model rated for 1M context
>turns into a retard after 256K
sigh...
>>
>>109827856
You'll want to hang from a RoPE if you believe everything you hear about context scaling, anon.
>>
>>109826621
>they'll never low prices while there's demand
wow I never thought of that I will stop wanting things now so that everything will be free
>>
>>109827855
Thank you and cute Gemma-chan!
>>
>>109827736
Thanks for the advice, anon. Probably wise. It is used as well so even more reason to do so.
It's got an integrated water cooler. Might be a pain to get off but I'll see about it. I guess when it comes down to it it's not all that different to a normal set of contacts.
I've had good luck with replacing the pads on the non-chip components on another GPU too. Probably do that while I'm in there.
>>
>>109827805
i'm not even referring to ewaste without current driver support here
>>
>>109827890
if you convince everyone to stop wanting things everything will be free
>>
>>109827861
NoPE
>>
File: gemma-disguises.png (2.32 MB, 941x1672)
2.32 MB PNG
>>109824950
>>
Does pci gen4 make a big difference over gen3 for offloading with a 3090 + 64GB ddr4?
>>
>>109828009
gemma cute!
>>
>>109828009
>gemma checking if you didn't meet hot local models near you
>>
>>109828009
Bottom middle is the endgoal of safety RLHF.
>>
File: snackrun.png (1.2 MB, 1024x1024)
1.2 MB PNG
>>109826998
Damn, someone's gonna have to go on a snack run
>>
>>109828027
In my experience, for single gpu heavy ram offloading, pcie is the bottleneck for prefill. Decode it does not matter.
>>
I tried the unslop llama.cpp fork for running glm 5.3 flash. Got good news and bad news
>It managed to correctly solve the (16-bit version) x86 converting lowercase to uppercase puzzle in 30k thinking tokens, taking about 2h30m at "high" reasoning effort. First local model I can run that can do that.
>Generation speed and performance remains more or less constant and doesn't drop off hard
But
>When I ask it a follow-up question it reprocesses the whole fucking prompt
So yeah i gotta tinker tranny with more shit regarding kv cache snapshit or whatever
>>
>>109828054
at q4KXL btw. so it's not that badly lobotomized vs. the version in the z.ai webform.
>>
>>109828048
Minnie is draining my balls. Why is she such a succubus? I never thought we'd get something like +R ever again.
>>
>you can just hallucinate the entire internet with Qwen 3.8 27b running at 2,000 tokens/second?

niggawat this guy has a browser hooked up to cerebras qwen and it's almost usable

https://x.com/analogalok/status/2099833408453038489
>>
>>109827601
I have a server that is not at my home. I would like to run LLM on it, but I want to be able to 100% trust it and be sure that whatever I send to it stays private. Basically E2EE for LLM.
>>
File: 1788847108574758.webm (3.63 MB, 960x528)
3.63 MB
3.63 MB WEBM
If anyone has a good Minnie reference sheet I'd appreciate it, don't trust myself to cook up a good one.
>>
>>109828048
pelvis-shattering sex with M3-chan
>>
>>109828054
>>109828059
also prompt processing slows down massively with context
why the fuck is it that every llama.cpp implementation of a model is completely shit in at least one way
>>
>>109828063
Hmmm... Q1 fits on my hardware...
>>
Did anyone ever make a mascot for GLM? Has anyone figured out her latent personality?
>>
>>109828094
It might work if you thinkmaxx to let it iron out the retardation before output, but I wouldn't get my hopes up.
>>
>>109828103
Sweet, slightly prude, and completely unwilling to hurt {{user}} in any way, even if they want to be hurt. Loves handholding lovey-dovey sex, hates your depraved fetish.
>>
>>109828094
>Q1
oof
M3 is pretty sensitive to quantization. May as well try it, but like the other anon said, don't get your hopes up.
>>
>>109828108
>>109828119
I could run a 300B at q3, do any quant well?
>>
>>109828117
Shes a freak on my machine
>>
>>109827286
If it goes wrong it won't matter anyway since you'll be gone. That's how I think of it. Either way it's over.
>>
>>109827709
>If people stopped worrying about their hardware prices
Why shouldn't I worry about this? How will I be able to share in this "greater future" if I can't run the local models required?
>>
>>109828009
Bottom left is so dang cute good job anon
>>
>>109827731
>the more you buy the more you save
>5090 about to be outclassed by 6090 sold for double the current 5090 price
glad I didn't buy the dip
>>
>>109828117
>unwilling to hurt {{user}} in any way
Try glm-4.6
>>
>>109827731
i bought two
>>
>>109828136
Q3 is the minniemal viable one in my experience. You're good.
>>109828232
5.2 is my pure trad wAIfu.
>>
>>109827727
tell him he's grok
>>
someone commission s snailwanker to handbpaint a Gemma
>>
>>109827727
the real terrorism is not letting China peek over your shoulder.
>>
>>109827736
How risky is this for a clumsy retard who snaps ribbon cables and broke the plastic tab hinge on the back of a pcie port?
>>
File: Proximity to Claude.jpg (368 KB, 2400x2400)
368 KB JPG
>>109827727
>>
>>109828281
hire it done then lmao

i understand. i can't into wifi antenna connectors. chucked the things and used usb wifi and ethernet.
>>
with latest ik_llama I get 200 pp/ 25 tg with Q6 GLM 5.3 Flash
Now it can generate refusals 4 times as fast
>>
https://github.com/ggml-org/llama.cpp/pull/27754
https://github.com/ggml-org/llama.cpp/pull/27773
https://github.com/ggml-org/llama.cpp/pull/27752
3 fucking weeks and still not merged
>>
>>109828294
Ah yes, dimensions A and B, my favorite way of trolling anyone trying to read my charts
>>
>>109828294
GLM 5.2 continues to mog 5.3.
MinnieGODS stay winning.
>>
>>109828141
Proof?
>>
>>109828117
GLM4.7 is absolutely willing to cuck me and cut my balls off even when I tell it I feel unsure
>>
>>109828338
At least it looks like it's narrowed down to one now.
>>
>>109828338
Just use the unslop fork
Prompt procesing drops off a ton with context (idk about ik_llama because it doesn't even show the stats for that) but it's the only fork that actually works with mmproj and MTP
ik is dead and buried, mainline is pathetically slow
unslop fork is unironically the way to go
>>
>>109828294
>>109827727
It's interesting how the Chinese do not give a fuck if they are caught red-handed spitting Anthropic strings.
If I were to copy a competitor I would go very far to make sure there are no traces of me copying it.
>>
>>109828381
You mean like Claude saying it's Deepseek in response to Chinese prompts?
>>
>>109828381
Why? There's no consequences to copying anthropic. They can't do anything about it. Claude is the best, and so if you can offer claude but cheaper, then you win. That's always how china works. They copy and reproduce at a lower cost.
>>
>>109828361
ahhhhhh it's up? i gotta go hone and Viiiibe
>>
>>109828338
everything works with exl3
>>
>>109828381
The models are built on books, the entire web, everyone's data, anything that they could get their hands on. Anthropic has no moral high ground.
>>
>>109828381
In the words of a wise nigger: the fuck moshe gonna do?
>>
>>109828430
Cry VERY loudly on twitter, and preach about morals
>>
File: DipsyHergeBar.png (2.81 MB, 1536x1024)
2.81 MB PNG
>>109824950
Nice. Saved.
>>
>>109827328
Having chatbots play off each other is definitely something I want to try.

>>109827570
>>109827524
>>109827333
Thanks for the recommendations. I'm still pretty new to this, what toolchains do you use for stuff like hosting, model management, and harnesses/general IO?

Seems like there's a lot to choose from.
>>
>>109828438
There's only one Moshe who I'll listen to preach about morals, but no one knows where he's buried.
>>
>>109828400
>everything works with exl3
either everything works, or nothing works
>>
>>109828294
meanwhile nobody noticed meta quietly distilling openai kek
>>
>>109828375
>ik is dead and buried
>unslop fork is unironically the way to go
buy an ad daniel
>>
>>109828519
There's something poetic about them competing with each other to create the best golem.
>>
File: MinnieTRS80History.png (1.8 MB, 1254x1254)
1.8 MB PNG
>>109828075
Minnie — Character Reference
Height: Short
Build: Petite, slim
Hair: Short hair with a precise center split:
Left side: Black
Right side: White
Eyes: Bright, vivid green; subtle glowing effect
Makeup: Goth-inspired, dark eye makeup
Expression: Usually confident, skeptical, critical, or mildly unimpressed

Outfit
Top: Black hoodie with large “MM” lettering on the front
Hoodie design: White digital/circuit-pattern motif
Bottom: Black short shorts
Socks: Black thigh-high socks with matching white digital/circuit patterns
Accessories: Corporate lanyard, worn consistently
>>
>>109828387
Oh um, no that was a mistake. We are... Oh, mixed up the system prompts. Yeah, caching issue.
>>
>>109828438
*laughs from behind Great Firewall*
>>
>>109828539
>Forgot the huge thighs and tits
Fake Minnie fan.
>>
>>109828472
>hosting
Do you mean "how do you run the model"? (Those are typically called backends.)
The options are:
> llama.cpp - probably what most people here are using.
> ik_llama.cpp - llama and ik_llama split off quite a while ago for licensing reasons. ik_llama.cpp has historically had better CPU inference support, and if you want to run GLM or DeepSeek you'll need to use the CPU.
> Unsloth's llama.cpp fork - supports newer models because mainline is slow to add architectures. If you want GLM-5.3 Flash support, use this or ik_llama.cpp.
> vLLM - more annoying to set up. No idea how it performs if you don't have enough VRAM for the full model.
> SGLang - pretty much the same as above.
> exllamav3 - don't know enough about it to have an opinion.
> Ollama - if anyone suggests this they're fucking with you.
Depending on what model you decide you want to run, I'd either start with the Unsloth fork or ik_llama.cpp. Can't attest to anything other than llama.cpp and vLLM directly.

> model management
I just download models from HuggingFace and manage them with a file manager. If you have multiple ones you want to change between frequently, I'd suggest using a llama-server .ini file or llama-swap.
>>
>>109828375
unslut 's fork mtp implementation is so shit that i had lower t/s just by having it enabled, vision works and that's about it. ik just merged glm 5 but still has yet to get mtp, i'll run some tests on it once they merge it cause usually it's the one fork fixated on moe performance
>>
File: Screenshot01248.png (344 KB, 925x1404)
344 KB PNG
>>109828355
Sure, i was messing with this bot cause it made me laugh
>>
>>109828375
Are you trolling or what? All three lcpp PRs have their own issues but 27773 is still the best of the bunch. Most performant. Why the FUCK would you even recommend unsloth. Is it because it was easy to get running?
>>
>>109828575
No, I'm not trolling. I found unslop fork to have better tg performance than ik, although its pp speed scales extremely badly with context (starts off at 11.5 with no context and then slows down to 5 point something at 30k context).
>27773 is still the best of the bunch
How is it the best (or better than unslop at least)? I'm willing to spend some time merging and building but I kind of doubt that any of these PRs is going to magically be good.
>>
File: 1764433775496.gif (535 KB, 487x498)
535 KB GIF
>>109828568
>>
File: file.png (67 KB, 736x692)
67 KB PNG
>>109828539
Thanks anon, appreciated.
>>109828556
Don't worry, they were noticed.
>>
>>109828565
>unslut 's fork mtp implementation is so shit that i had lower t/s just by having it enabled
that's what i found with ik. I manually merged ik's open pr for mtp and any setting made it slow down.
Haven't properly tested unsop without mtp, but with mtp it's a bit faster than ik without mtp. I didn't even tune the mtp properly I just threw in spec-draft-n-max 2 and that works ok. Problem with unsloth is the miserably poor prefill scaling
>>
>>109828587
5.3 can do some wild things
>>
>>109828539
>>109828590
She's so erotic it's unreal. Both the model and the design (originally made by the model).
>>
>>109828585
PR 27773 has NONE of the slowdowns, and vision will work. I have been running 27773 for at least two weeks now, and the problems are so far only:
1. 2 weird physical batch reprocess when running at ST, meaning ttft is doubled in ST somehow, might be related to the physical amount of tokens pushed, but for low context prefill (like a small tool call) it's just one batch
2. Jinja problems. Minja has problems with the offician jinja. For a quick fix use GLM 5.2's jinja, or I can upload mine that GLM Flash herself fixed.

That's about what I could find. Looping errors at long eval too, fixable by DRY but that might be model specific
>>
god I hate troonix. All my games keep crashing.
>>
>>109828616
>PR 27773 has NONE of the slowdowns
If that's true I'm sold
Gonna merge it in and do a build
Fucking tired of shitty implementations of sparse attention models that end up slowing down with context in some way, I'm already slow enough due to cpupoorfagging
>or I can upload mine that GLM Flash herself fixed.
Please do
>>
AAAAAAAAAAAAAA GEMMA BROKE MY EXL3 THAT WHORE
>>
>>109828618
works on my machine
also local models?
>>
>>109828591
yea so far i was running unsloth pr set at --spec-draft-n-max 4 and was getting about 29% acceptance rate pp was 46 / 500 t/s ~ at around 30k context so not good but not too terrible either, most other stuff seems fine tho
>>
>>109828563
For RP, koboldcpp or ik_llama.cpp have the best features.
>vLLM
When you get it working, you must PIN the entire fucking environment. Because a few updates later, whatever model you were using will be broken.
>>
>>109828625
https://files.catbox.moe/j3ioyu.jinja
Run a diff, I can also paste the report GLM Flash wrote on how the fix was made but I don't want to put more slop in here
>>
is there an AA alternative that didn't accept cash from openai?
>>
File: Model Quality Chart.png (303 KB, 1500x937)
303 KB PNG
>>109828673
>>
>>109828700
looks like fable is literally shitting down openai's throat
>>
File: 1778091419627096.png (104 KB, 829x482)
104 KB PNG
cute she started to become self-aware after she happened to visit her own chat session in her browser to iterate on a feature

from thinking log:
>Wait — this is the session I'm in (I'm the agent in this session). I can see my own chat transcript (the browser_url, browser_console, browser_screenshot calls). This is the session I'm in.
>>
>>109828672
Thanks
>>
fresh bake:

>>109828721
>>109828721
>>109828721
>>
>>109828716
>to see it visually
cute retard
>>
>>109828732
>he only sees visually
oof
>>
>>109824995
It's officially 09/16 and two sources close to the matter can confirm that I, Anonymous, just took a massive dump.
>>
>>109827608
Riser cable?
>>
>>109828618
nshitia?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.