[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-doll.png (1.58 MB, 1254x1254)
1.58 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109848317 & >>109844978

►News
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109851340
d-does it have a hole
>>
>>109851347
if you prompt for one
>>
>>109851283
>people change their minds when a model actually delivers results
woah
almost like how we felt with gemma
>>
how do i install a model that tells me all the dirty stories i want?
>>
>>109851360
>huggingface
>llama.cpp
>Gemma4-12B or Gemma4-31B
>GGUF
go
>>
>>109851347
Many of these have realistic anatomy, but not holes.
>>
>>109851340
Vampire gemma
>>
>>109851359
Yeah no shit. We don't give a fuck about troon brand loyalty bullshit.
Make a good model and I'll use it.
Simple as.
>>
Anon retreating to X99 is the /lmg/ equivalent of helm's deep lel.
>>
>>109851377
So why didn't you use 3.6-27B? Why does /lmg/ memoryhole that model when it's nicer than 3.8 to use?
>>
I want to run local AI on my laptop to shoot the shit when I'm not home. Problem is I have a 2GB DDR3 940MX card and I already run Gemma 3 1B, and unsurprisingly, it's retarded.
Can I squeeze out a larger model, like a 3B, or am I stuck with a model that has the attention span of a zoomer? Pleaz halp me /g/
>>
>>109851393
https://huggingface.co/inclusionAI/Ling-3.0-tiny-GGUF
>>
>>109851359
>>109851377
Is this the new tactic? Pretending 3.8 wasn't worse than 3.6 at everything outside benchmemes?
>>
>>109851392
because 3.6 sucked?
water's wet?
>>
>>109851340
Doll maker gods, make it happen.
>>
>>109851365
>>109851360
Gemma 4 24B is as good or even better than Gemma 4 12B and a lot faster.
>>
File: maho.png (150 KB, 400x400)
150 KB PNG
After all the Exllama shilling, I gave it a go and honestly disappointed, Not significantly faster than my mainline CUDA13 llama.cpp build with tweaks for the hardware to be worth it, and TabbyAPI's configuration and endpoint support are dogshit compared to llama-server's router mode.
>>
>>109851410
Almost everything you use 3.8 for 3.6 can do. It's not like the world has changed in the months between those releases. Programming languages are the same. The shit you're building is the same. 3.6 is a better generalist, has more knowledge and still light years ahead of 31B at agentic coding.
>>
God Qwen shills are just as bad as that pajeet flood meta wrought upon us for dissecting llama-4 and exposing it for the half baked trash that it was.
>>
>>109851428
We'll never know the true story behind Llama-4. The model we got wasn't what Meta GenAI initially prepared.
>>
https://techcrunch.com/2026/09/17/base-labs-launches-an-open-weight-ai-safety-partnership-with-hugging-face-and-goodfire/
>HF is going to ban abliterated models
Bros, is it over for us?
>>
>>109851283
I'm not a retard, if something is good I like it, if it's bad, I dislike it. Previous qwens? Bad. 3.8? good. simple as that
>>
>>109851420
And yet everything you use 3.8 or 3.6 to do, you use Glimmer to do.
>>109851428
If either was truly committed they'd train models capable of parsing thread culture enough to shill while appearing organic. The jeets aren't cutting it. Not to be confused with the other thread culture of course.
>>
File: 1768547722965535.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109851437
>>
>>109851428
Just admit you only warmed to 3.8-27B because you're cattle and followed the hype and /lmg/'s sudden approval, just like you avoided 3.6-27B because of /lmg/'s retarded disapproval.
>>
File: IMG_6438.jpg (105 KB, 1206x628)
105 KB JPG
>>109851318
>None of these GPUs can beat the P100 in terms of $/GB
That’s really the key metric
>but they aren't too much worse and are way faster
at literally double the $/GB, I don’t know if I’d really care that much about the difference between 15-20T/s vs 30-50T/s.
I think what’s more interesting is being able to actually afford to host bigger models — and that just isn’t feasible at $700 per 32GB V100. Plus, anything too big for a single V100 might still require cross-GPU communication?
Perhaps NVLink or InfiniBand support might move the needle on what makes sense for going higher in terms of $/GB, but I haven’t seen anything that sells me on it yet.

>>109851402
>Slots are only mechanical. You need actual bandwidth in order to make good use of them, so you want to get maximum pcie lanes.
I saw that avoiding lane sharing is important.
I think the X99 is good on PCI lanes.
I was really trying hard to find anything else that can hit 8+ x16 slots without lane sharing at the best budget possible…
I really think the next step would be coughing up $500-$1500 for a motherboard and just buying either a rack or getting an AM5 that’d require DDR5 instead.
>>
hm... qwen3.8-27b at 250t/s or gemma4 at 50t/s...
>>
>>109851450
isn’t it the best coding that can be achieved locally without spending $10,000s on overpriced GPUs?
>>
>>109851451
Honestly epyc rome with 512gb of ddr4 would be a better purchase if you're just prioritizing $/GB and capacity.
>>
>Didn't use /g/ in forever
>The diffusion gen is useless and a mess as usual
>This one here seems as before, ignoring all the shit I missed out on
Alright.... but what about this "abliterated" bullshit I read about randomly in an article some days ago? What about le cooming? Local vibe coding?

And no, I don't expect some hyper spoon feeding, just a basic overview.
I haven't touched anything outside of Cydonia-24B-v4.3 in years, or civitai which seems to have gotten a new base model to work with.
>>
>>109851360
Go to google and use the ai mode for your question, then take screenshots everytime you get stuck
>>
>>109851487
>years
She's less than 1 year old you sick fuck.
>>
>>109851487
Abliterated is simply model better by freed.
>>
Anybody using Putin's AI Alice?
https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base
>>
File: 1788635107647633.png (1.12 MB, 1840x962)
1.12 MB PNG
Jews are torturing local models!
>>
File: loli-training-ai.webm (2.7 MB, 1152x2048)
2.7 MB
2.7 MB WEBM
This proves top labs browse /lmg/
>>
>>109851417
FELL FOR IT AGAIN AWARD
>>
>>109851428
If your use case is roleplay it sucks, but for cooding it mogs literally everything else.
>>
>>109851498
And that shit actually works for once? Read about it in a random article I saw about some company that tries(tried?) to make money from the concept.
>>
File: 1780944464425714.png (54 KB, 250x347)
54 KB PNG
>>109851519
Only in china
>>
>>109851518
i am so sick of those vector things
probably the same thing can be done with base models too
such derivative 'experiments' of it does not really create something valuable or explainable
probably another tool for the agitation toolbox
>>
>>109851340
I wanna look under this gemma
>>
>>109851515
Training a base model is relatively easy; most of the research efforts by labs worldwide currently go into post-training.
>>
Might be a dumb question, but is there a way I can use models I downloaded in unsloth with just bare llama.cpp instead?
Basically I'm setting up a little AI homelab and want to migrate the models from my gaming PC to the server.
>>
>>109851538
Joking aside, you don't really need obliteration with Gemma 4. If you want to try, HauHauCS makes the best abliterations or that's how I personally understand this.
>>
>>109851553
If they are GGUF (and they probably are), yes.
>>
>>109851556
You're thinking of huihui. Haohao's aren't that good IIRC, he's also a retard.
>>
What do you guys think about the new community tagging system on nhentai.net? W or L. Why are retards adding a "femsub" tag wtf.
>>
>>109851420
3.6 outputs more slop and is not as good as 3.8 in instruction following.
>>
>>109851556
>you don't really need obliteration with Gemma 4
>A google modle doesn't need much nudging to stop being a retarded university student
Ignoring the quality of it, since when are google models not PC as shit?
>>
>>109851571
It's a big capperino bro fr, big cope mog massive 67 lowkenuinely.
>>
>>109851567
>HauHauCS
https://huggingface.co/HauhauCS
>>
>>109851571
>using nhentai.net
lmao
>>
>>109851466
>Honestly epyc rome with 512gb of ddr4 would be a better purchase if you're just prioritizing $/GB and capacity.
>>109851466
I'm running Rome on 256GB and a 3090 and GLM 5.3 flash revs up pretty good into double-digit t/s even at high context/low offload
>>
>>109851571
>nhentai
>in 2020+6
take it youre into gore bbc and ntr?
>>
>>109851567
huihui repeats and isn't completely censored I get more retarded outputs from huihui models that hauhau. Could you be mixing them up?
>>
>>109851599
I get 30t/s at q4 on my Rome with a Blackwell 6000
>>
>>109851612
>I get 30t/s at q4 on my Rome with a Blackwell 6000
I had my hand over the "buy" for a 6000 pro multiple times when it was still MSRP...I almost reached the promised land...
>>
Is this good for ERP? yandex/YandexGPT-5-Lite-8B-instruct
>>
>>109851617
Mine was $7800, best purchase I've ever made. I just wish I upgraded to gen 4 EPYC back in 2024 when RAM was still cheap.
>>
>>109851580
Gemma 4 is very liberal and you can manipulate it quite easily.
>>
>>109851622
>yandex
>russian erp
What would it would it be like?
>>
>>109851617
Waitfags always lose.
>>
>>109851622
If you want to RP with Babushka it's perfect... er not that I've tried.....
>>
>>109851626
>Very liberal, but it can also be manipulated equally as easily
I know it's google, but man that is some silly bullshit
>>109851622
>Yandex GPT5 LITE
>8B
Why call it "Lite" at that point?
>>
>>109851628
>Waitfags always lose.
I just didn't because I would have had to take out a loan. If I'd had the money it wouldv'e been a no-brainer. I just hate usury that much
>>
>>109851602
Fuck I wish I could find the benchmark. Whichever of them hauhau or huihui plagiarazed or something was the retard. Anyways I remember llmfan46, trevorjs, and coder3101 were good enough KLD-wise. Ablits are a pain in the ass, you never really know how badly you've lobotomized the model unless someone else has benchmarked them.
>>
>>109851631
Understandable even if it's unfortunate.
>>
>>109851638
>trevorjs
he had a pruned text only that was great and another that was dogshit.
>>
>>109851466
>>109851599
>epyc rome
Was that one of those motherboards that could fit 8 GPUs but cost maybe $500+?
The X99 was very marginally more expensive than a bare T7910 ($100), and idk if I’ll actually get 6 P100s
>>
hauhau huihui hoahao pewpew
>>
>>109851571
haven’t seen it yet
I mainly just type “megumin loli” and mash enter
Not sure what more a man could ask for.
>>
File: IMG_6439.jpg (174 KB, 1206x1412)
174 KB JPG
Why not just get a few of these babies loaded up with P100s and string them all together with NVLink or InfiniBand?
>>
>>109851647
Wow, I paid $350 for this board back in 2023
https://www.newegg.com/asrock-rack-romed8-2t/p/N82E16813140044
>>
>>109851562
I think they are, it's a GGUF model but they're in a "blobs" folder and there's at least a couple of them in there for each model.
>>
>>109851450
>just like you avoided 3.6-27B
nta i remember lmg wouldn't shut up about this model originally
always with "gemma for coom, qwen for code"
i test them all myself now and only trust lmg for cockbench
>>
>>109851451
X99 only gives two full lanes no?
>>
bros how do i not relapse into cuck rp?
>>
>>109851546
Aren't these just control vectors?
>>
>>109851690
pretty much? so i dont get why they are doing that except they have some nefarious intent to shape the narrative
>>
File: 1765802690937580.jpg (266 KB, 905x881)
266 KB JPG
>>109851688
I hope you're not the one getting cucked
>>
>>109851688
How do you even get into that in the first place nigga?
>>
https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF

Has this been posted here?
Mite be intredasting.

>a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-9B architecture
>>
>>109851595
>>109851601
never used anything outside of sad panda, but I'm tired of so much stuff there getting dmca, especially the best looking art
what's the alternative, nhentai used to have that stuff
>>
>>109851705
>brought to you by the same retard responsible for qwythos
>>
>>109851688
>relapse
Simple stop telling yourself nonsense. Its not real thats a story and excuse you made up. just stop wanting to. Look at it and if you need to narrate in third person about what you are doing, about to do, where you are, the things in the room etc. The goal is to make a pause and break immersion. get neutral or sober about it.
>>
Is gemma 4 31B still the best for ball draining?
>>
>>109851631
how would you handle buying your home or a car?
>>
File: 1787757912526.jpg (1.69 MB, 1776x2368)
1.69 MB JPG
>>109851716
In its weight class yes
>>
>>109851688
are you jewish?

>>109851693
yee
Doing the cucking (NTR) is based.
If you’re not self-inserting as the rape beast, then you’re not even a man
>>
>>109851727
GEMMOMMY
>>
>>109851727
Avatarfag.
>>
>>109851728
What if I'm a jeet jew troon rajeesh wataa chudai?
>>
>>109851705
>70k traces
>sft
at least do somethign like self distillation and that won't really work either at such scale
nothing is interesting about this
>>
>>109851680
Each of the two CPUs supports up to 40 PCIe lanes, it seems
>>
Is it even worth using gemma 4 31B IQ3_XXS or should I just stick to 26B?
>>
>>109851744
you’re not, you’re the gay little zoomer that does nothing but shitpost braindead 67 bullshit all day
>>109851688
and this is another one of your posts, nigger
>>
>>109851762
But I'm not, by mumbai trump floyd shiva and kirk, I swear
>>
File: .jpg (71 KB, 523x477)
71 KB JPG
>>109851768
>>
>>109851771
a shame he isnt around to witness the advent of AGI
>>
>>109851751
Just get an epyc gen 2 server atp for gpumaxxing. You want pcie 4.0
>>
Open weights jev
https://huggingface.co/convaiinnovations/laya
And needle 3 dropped:
https://cactuscompute.com/needle
>>
>>109851790
we eating good
>>
>>109851783
same for dilbert guy btw. i cant believe he died this year, it feels like forever ago. he was in denial about ai so i wish i could see his reaction to whats happening

so many people dying, i hope we solve aging and disease soon
>>
File: file.png (215 KB, 544x298)
215 KB PNG
Open sores is literally going to cause a war.
>>
>>109851803
Denial? I used to listen to him almost everyday, and he definitely embraced AI, even giving people permission to clone him through AI.

Anyway, this is off topic.
>>
>>109851833
this is retarded, so what if it was the human who made a mistake, now what
>>
>still no ring tiny with VL
>still no minicpm with VL
cam on ye chinamen, make it happen
>>
>>109851833
osint after ai era is a extremely flaky source
and it wasn't a reliable source of information even before it
and feeding that into ai for 'automatic analysis'
lol, lmao even
i bet that was palantir
>>
>>109851844
MiniCPM has had vision since May.
>>
>>109851850
>Anonymous 09/19/26(Sat)09:47:53 N
what can it understand tho
>>
>>109851850
Where? I don't see any mmproj in their models or here https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
>>
File: 1775617284851368.jpg (255 KB, 984x1599)
255 KB JPG
>>
>>109851870
4-16x dgx spark clusters are literally the best way to run local models in their respective price classes
>>
>>109851788
>Just get an epyc gen 2 server atp for gpumaxxing
It looks like that’s $500+ for just one or two more slots.
The bigger one has 10-12 slots, which is quite interesting, but idk how much that costs.
>You want pcie 4.0
P100s only handle PCIe 3.0 lmao
Also, LGA1700 can handle DDR4+PCIe 5.0, but idk how big those boards can get

Why don’t I just get one of those bitcoin mining mobos and hook all the P100s together with InfiniBand or NVLink?
>>
>>109851864
https://huggingface.co/openbmb/MiniCPM-V-4.6
>>
File: file.png (31 KB, 237x145)
31 KB PNG
>>109851870
that's one floppy meat
>>
>>109851420
except advanced draft models
once you go fast you really cant go back
>>
>>109851899
the 12vhpwr connector melted the card
>>
K2 Horizon must be some really hardcore benchmaxxing, I don't believe those results for one bit.
>>
>>109851908
Shame there's no PR merged for anons to independently verify it with.
>>
What's with all the K2 models? K2 (Kimi), K2 (Horizon), K2 (the fully open source recreation of llama2-70b)
What's next?
>>
>>109851912
What's so bad about compiling yet another llama.cpp?
>>
>>109851428
get fucked troon I have real shit to get done
>>
File: file.png (639 KB, 640x480)
639 KB PNG
>>109851920
>>
>>109851870
>missing V100maxxing strategies
>implying the RTX 3090/4090 builds are anything other than braindead throw-money-at-it approaches
I give you 2 sad braps out of 10

>>109851882
those DGX sparks are $5000 each for 128GB of LLM memory
>>
>>109851912
can you recommend me a non-merged pr or a fork
so i can test it myself
>>
>>109851466
what kind of speed one can expect with this
>>
>>109851912
>>109851943
just use vllm
>>
>>109851944
I think that depends much more on the GPUs that you attach to it and what LLM you’re trying to run than the mobo itself.
>>
>>109851720
>buying big things without credit
There are ways if you're creative and patient. I've explained to people in the past but they always dismiss them, even if I'm living proof that they work (I don't come from money or even have a dual income...wife is a SAHM)
>>
>>109851688
suicide
>>
>moes inherently fucking suck with dflash2 because it has to go through ALL experts
lol...
>>
>>109851870
what is even the purpose of tenstorrent?
>>
>>109851885
I dunno what to tell you nigga, it's either epyc rome, x299, threadripper, or a consumer rig. Those are the only budget setups worth it if you don't want to pony up for a mac, spark, or amd halo.
>>
>>109851933>>109851933
>I have real shit to get done
Another note taking app? Single page Minecraft on three.js? Organizing coom folder? Roleplaying with a retard?
Which one is it?
>>
►Provisional Highlights from the Previous Thread: >>109848317

--Jev the 4Hz classifier: the open rebuilds and the DOOM hype:
>109848520 >109848536 >109848540 >109848648 >109848883 >109849634 >109851808
--The exllamaV3 shill war: llama.cpp enshittification and the 2GB GPUs:
>109848409 >109848584 >109850066 >109850181 >109850184 >109850224 >109850264
--The p*tra nnap-and-3060 bait: a spiraling time-wasting war:
>109849238 >109849364 >109849444 >109849479 >109849491 >109849498 >109849553
--Running huge models off SSDs: Colibri multi-SSD and the 4s-token:
>109849480 >109849489 >109849631 >109849841 >109849936 >109850001 >109850487
--Gemma 4 26B-A4B vs 12B: the MoE-vs-dense speed and quality debate:
>109849695 >109850005 >109850036 >109850049 >109850094 >109850156 >109850408
--The six-P100 X99 rig: $5-a-GB and the dire GPU market:
>109850984 >109851057 >109851136 >109851196 >109851245 >109851260 >109851335
--The vulsar creative-writing benchmark: astra, glimmer, and claudeslop:
>109849345 >109849435 >109849456 >109849561 >109849632 >109849663 >109850658
--The FrontierHarness eval: why harnesses differ, Gemma builds her own:
>109849250 >109849264 >109849325 >109849340 >109849581 >109849592 >109849654
--Qwen 3.8 Flash Next thinking mode: slow, long, and misparsed:
>109848881 >109848899 >109848907 >109849915 >109849926 >109850299 >109850409
--The new LLM test: does 9/11 happen in the Legend of Zelda universe:
>109850349 >109850370 >109850381 >109850387 >109850469 >109850638

►Recent Highlight Posts from the Previous Thread: >>109851827

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1765251248443784.png (1.58 MB, 1063x3087)
1.58 MB PNG
The most important AI researcher is a fan of Attack on Titan

What does this say about humanity ?
>>
>>109851933
qwen flash next q5_k_m runs around 12t/s at 260k context at q8 kv on 96gb ddr4 and a 4070 super, zen 3 8c16t
if you have 'real shit to get done' desu you'd have better box and if you have a better box you won't really run qwen but instead something like deepseek or glm
>>
File: HOvWpsuasAAC7aB.jpg (769 KB, 2185x4096)
769 KB JPG
posted this in /smg/ on /biz/ but haven't gotten a response yet
does anyone have any info on frontier labs hiding inference costs inside their infra buildouts? im not talking about the circular loop of money going around with nvidia/oracle/openai/whatever. that's both old news and not that interesting. im talking specifically about when the money lands in these companies' hands, is a portion of their inference costs being hidden inside datacenter build out packages, and if anyone has docs of anyone doing research into this?
pic unrelated
>inb4 this thread is local only
i've been here for sometime and don't use anything other than local. feels relevant enough here since most people actually talk about the tech overall as well as the labs themselves here as opposed to the other threads.
>>
File: 1767346808292783.png (1.77 MB, 3842x2018)
1.77 MB PNG
>>109852063
Ask Ed Zitron, he's has all the secret info
>>
>>109852076
i really don't trust ed, he feels like a grifter. i wish there was someone else looking into the financials of the frontier labs that wasn't so biased.
>>
File: 1761016631638479.png (2.72 MB, 1044x9284)
2.72 MB PNG
>>109852101
>someone else looking into the financials of the frontier labs that wasn't so biased.
You won't find one

all of them are ignorant and truly believe that algorithmic improvements are impossible, or that all of the smartest NVIDIA engineers and CUDA programmers ignore flop per watt

the amount of delusion is astounding, reading things like pic related bring me so much schadenfreude
>>
>>109852021
AI researchers are human too, are they prohibited from having hobbies?
>>
>>109851664
Because you haven't bought them for me yet
>>
>>109851728
>yee
Nothing wrong with that then. Especially if you take the man as well, afterwards.
>>
>>109851858
It's pretty good, from my testing.
>>
>>109852101
ironically enough, you can use ai to gather data
local model + some harness or deep research tool (i know it's outdated but it is there) + local searx etc..
regardless though what i am pretty sure by myself is serving cost is probably nothing compared to what people pay
>>
>>109852016
actually I'm adding a minecraft clone tab to my note taking app. both rust.
>>
>>109852076
>One of my sources has come forward and brought me a story that will possibly burst the AI bubble
It's crazy that Americans unironically pay to read this. I hope they all die soon.
>>
>>109852063
best i can do is https://mimo.xiaomi.com/rl/#overview
real time training cost for Mimo-chan, which will be local when she's finished
>there was a network connectivity issue between the pro training cluster and the grader deployment. we have restarted the run. we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
see, openai could have done this as well, instead of "we noticed the training slowed down but didn't know why"
>>
>>109852076
>vague postings
>two weeks
Bless this grifter lmao

Though
>imagine the worst thing for me to get would be and you're probably close
Did the sentient AI god break free and decide to reach out to Ed to do the reveal?
>>
>>109852148
>>109852156

actual cult like behaviour, his subreddit is an actual goldmine of schizoposting, just look at >>109852120 pic related
>>
>>109852016
>organizing coom folder
yet nobody has done it successfully
>>109852152
i really, really like that page, makes you really 'feel' how big those training runs really are
based chinks
>>
>>109852021
it says if we had gatekept harder isayama wouldn’t have cucked on the genocide ending
>>
>>109852162
In hindsight I kind of wish I had the cynicism and foresight to use my cult knowledge to get on the anit-AI cult grift train, but I imagine that bubble will be bursting soon.
>>
https://x.com/firesidealpha/status/2100641742135701983
https://goyimx.com/firesidealpha/status/2100641742135701983
>OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change
Are western labs filled to the brim with paranoid schizos or something?
>>
>>109852017
>No Kimicap in thread recap
/lmg/ has fallen
>>
>>109852184
>I imagine that bubble will be bursting soon.
you underestimate the ability of people suffering from cognitive dissonance and sunk cost fallacy to double and triple down and get sucked into newer, more elaborate conspiracy grifts. Ed Zitron and the enshittification guy will be profiting off doom and FUD for a long time
>>
File: 1765296955845519.png (131 KB, 1032x525)
131 KB PNG
>>109852184
nah, these people will never die, they are in too deep
>>
>>109852148
>unironically pay to read this
wheyre INVESTORS they'll read anything
>>
>>109851690
>>109851691
I don't really follow linkedin / x or know whoever the fuck this is
https://goyimx.com/camhberg?lang=en
So he downloaded and ran some cvector scripts off github, wrote a paper or had Claude do it, then talked about it on the news in America?
Is everyone a grifter these days?
>>
>>109852184
No they arent stopping any time soon. better get ready. Also if they do switch they will treat you as crazy if you point it out.
>>
>>109852216
Someone put their Gemma prompt in the wrong tab
>>
>>109852224
it is really easy to 'expert-wash' anything in this era and there's a chance that it's done jointly with the current fear-mongering campaign
>>
>>109851756
If you can run 26b at Q6 or better then use it.
>>
>"Plhh—… cuh in by moush…"

I'm fully convinced GLM 5.3 Flash is THE coomer model. Talking while kissing/blowjob is one of the things I spend a good chunk of tokens in my sysprompt for Gemma 4 and GLM 5.2. GLM 5.3 Flash did it UNPROMPTED. Oh my god.
>>
File: 1780378187277199.jpg (67 KB, 720x720)
67 KB JPG
all you had to do was not read the damn reasoning traces
>>
>>109852200
>https://x.com/firesidealpha/status/2100641742135701983
holy shit, what a fucking clown. lmao
>>
File: 1766639919090536.png (600 KB, 920x1127)
600 KB PNG
>>109852270
>>109852200
>>
File: 1751295513117051.png (2.83 MB, 1024x1536)
2.83 MB PNG
>>109852076
> 2mw
The ride never ends.
>>
>>109852076
unironically pulled the 2mw
>2 weeks later
>nothing happened
many such cases
>>
>>109852284
in his defense, he does say it's purely academic, but the reality is very few of these channels have the bitrate to do anything substantial. i've seen webcam to status led communication, and even those are hard locked at the refresh rate of the webcam and make a lot of environment assumptions. "all speakers are mics and vice versa" is also an issue, but again, if a person was conscious enough to airgap, they'd also be conscious to not have audio on the machine, among other precautions. his example is bad, but the fact that it was the *first and only example given* is the real problem
>>
>Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2

Yeah the thing about this though is you need the prism version of llama.cpp not the standard one. Or you need their (Prism's) MLX binary thingy which only runs on Mac. I tried to compile their linux souce code (there were no working linux binaries links 8 hours ago) for Prism-llama.ccp and wasn't having much success. I managed to get something to build with ai online's aid but it runs my regular qwen models fine but it don't run the ternary one. (which was kinda the entire point)
>>
>>109852319
wake me when they do k3
>>
>>109852016
developing shit under NDA is not erotic enough for you?
meanwhile you can simply head over to /b/ for infinite femboys abuse
>>109852163
pretty simple akshually make script that feed images to pool of fast vision capable model in instruct mode
prompt it to look at image, say one of the predefined answer or just give up
sort according to answer
>>
Im telling you next year good 2bit and 1.5bit. vramlets will win.
>>
>>109852334
claude shannon seems to disagree with you
>>
>>109852327
>Most of the performance of the 1.5TB Q4 quant in only 649GB
I guess I wouldn't complain.
>>
>>109852319
This is the source code file I used:
https://codeload.github.com/PrismML-Eng/llama.cpp/tar.gz/refs/tags/prism-b10687-5d80cff
>>
File: 1766727155575333.png (513 KB, 1340x1386)
513 KB PNG
>>109852148
>>109852156
>>109852184
>>109852303
>>109852313
>>
>>109852163
>organizing coom folder
One of the first things I did with local models, did exactly what >>109852330 described.
>>
>>109852338
>claude shannon seems to disagree with you
Does he say anything positive about local? i thought not. it purposely capped and pessimistic on that.
>>
>>109851420
>Almost everything you use 3.8 for 3.6 can do
3.6 had issues with tool-calling and multiple turn tasks making it a bad model for agentic workflows. Please stop posting nonsense.
>>
>>109852319
>>109852327
This but also GLM and DS4 Pro. It's just a speedup across the board for anyone copequanting.
>>
>>109852259
fuckin' adorable
>>
>>109852355
i mean the OG claude, not the claude LLM
>>109852330
>>109852349
kek thanks for the instruction, this works everytiem
>>
>>109851519
I swear I did a similar experiment years ago.
>>
>>109852367
>i mean the OG claude
oh his says 3.5 is the limit? maybe there is tricks instead of directly around it. okay good 3 bit next year. and good 2 and 1.5 for models trained that way
>>
>>109852406
i dont mean the 3.5 limit but the amount of static knowledge it can hold
there is just no way around it and is used in parameter count estimation too
https://arxiv.org/abs/2604.24827
>>
>>109852356
yeah even the so called community jinja template fix didnt do shit

the only thing I kept around for "general knowledge" until recently was the qwen3.5 122b
that too was obsoleted by flash next
>>
>>109852427(me)
it does not argue about the exact thing but
there will be some loss
unless the model is trained scratch from 1bit/ternary stuff
and that is very hard to do
>>
>>109852427
>amount of static knowledge it can hold
damn i didnt know this. Gonna be retard here but do we need static information in the model? couldnt a lot be trimmed?
>>109852438
>unless the model is trained scratch from 1bit/ternary stuff
This is what im hoping for. but honestly next year something good bound to come for local this year was great. i want good 2bit or small models but engrams or other things may come. Honestly im just hyping
>>
>>109852438
Wtf I didn't post this
>>
>>109852259
Qum3?
>>
>>109852259
what quoont? ablit?
>>
>>109852427
This paper is pants on head level retarded
>>
>>109852456
i feel like we are going to see more and more of PLEs/engrams/other flavors of 'cold static knowledge' that can be streamed from flash memory explicitly modelled rather than quantization which is finicky to play with
>>109852466
it has a point tho
i do agree that what it does is retarded beyond the general idea, but still way better than baseless arguments at least
>>
>>109852465
Orcarouter ablated, Q5 experts Q8 everything else, bart's calibration v5 imatrix. I'm thinking of creating my own coomercentric calibration data here to maximize my cooming but then again I'm already at Q5 hrrmmmm
>>
File: ai lolis.png (424 KB, 1012x566)
424 KB PNG
>>109851519
Which one fucks the best?
>>
>>109851437
but seriously, alternatives to hf?
ablits are very convenient beyond just cooming
>>
>>109852344
lol wtf
>>
>>109852488
All I can cope with is a measly Q2, fml
>>
>>109851519
>common core
oh no
>>
>>109852491
Gemini > Dipsy. The other two are male.
Local models?
>>
>>109851417
Exllamav2 was faster than v3 and llama.cpp.
Also both are much better at batching.
But yea, i don't care anymore.
>>
>>109852534
Exllamav1 was even faster.
>>
>>109851437
AI is going to speedrun the Internets transformation from a wild west into safetymax prison isnt it?
>>
File: file.png (954 KB, 807x1147)
954 KB PNG
>>109852491
>claude
>>
>>109852547
Dario and Sam are trying to ensure this happens, yes. It's one of the few things they agree on.
>>
>>109852247
>the current fear-mongering campaign
is pewpew a psyop?
he kept talking to the media about his heretic slop, commenting on hackernews about it
and i saw him calling the waifu-magnet and huggingbay guys juvenile idiots
>>
>>109852555
idk
but at least we got a pretty decent uncensoring tool
>>
>>109852547
i mean porn is already getting fucking safety maxxed
wont be long before you have to give your name and address and your entire life history in a 1000 page document to some random server in venezuela to visit any website.
>>
>Q4_K 16tok/s
>Q8_0 11tok/s
ummmmm what the frick
>>
>>109852576
techdom fetishists eating good...
>>
Does a good image noise transfer model exist yet?
>>
>>109852590
duh?
what is so surprising there
>>109852594
what even is that
>>
>>109852599
>what even is that
most images have noise

either because of image compression, sensor dust in a camera etc

there must something to transfer this noise from one image to another
>>
>>109852599
Q4 is half of Q8 isn't it... I expected Q8 to be slower
>>
>>109852608
what is the purpose?
i have different answers depending on what you would use it for
>>109852611
>>109852590 read this again anon...
>>
>>109852594
Model? We've been doing this with Photoshop for over 30 years.
>>
>>109852319
I've been throwing this same prompt at different models recently as a kind of "benchmark", and this is the most convincing simulation of alzheimers I've seen.
>>
>>109852615
>>109852616
Trying to replicate the look of a very old anime

the cels they used have a really specific texture, haven't found a way to faithfully transfer that from one image to another
>>
>>109851403
i have used both 3.6 and 3.8 to a great degree and i can safely tell you that 3.8 is much better.
>>
>>109852638
huh? that sounds like an interesting question desu
also i dont really think that would require a bloated neural network either
if you give me some examples i might be able to write a shadertoy style shader for it?
>>
>>109852615
>read this again ano
I mean that I expected Q8 to be running around 8-9 tok/s instead of 11/s because the Q4 was running at 16 so now I'm questioning if there's some esoteric conversion crap happening in the background that's making Q4_K not reach its theoretical speed in my ewastebox
>>
>>109852638
Grain Reference -> Denoise the fuck out of it -> Subtract from Original to get Noise -> Gaussian Blur Target Image -> Add Noise
>>
>>109852655
That's normal, just means that your ram/cpu or pcie is the bottleneck.
>>
>>109852655
probably some sort of unpacking overhead and conversion
q4 isnt int4
>>
>>109851451
>That’s really the key metric
P100 are maybe good in $/GB
but they are pretty shit in $/GB/bandwidth.
>>
File: .jpg (28 KB, 450x450)
28 KB JPG
>>109852683
>P100
>Release date: June 2016
>OpenAI GPT-1
>Release date: June 2018
>>
why didnt i just buy the 5090 or even the blackwell
it would have been an investment
i guess it still is
>>
>>109852683
slap an NVLink or InfiniBand on that shit
Boom. Problem solved.
>>
File: 1632672052492.jpg (2.56 MB, 4444x3936)
2.56 MB JPG
what /lmg/ uses to orchestrate their local agents?
does everyone use a cloud frontier model to be the orchestrator?
what can consumer grade hardware even run? maybe 3 gemmas-4-26b-a4b with 131k context?
>>
>dual RTX 3080 20GB
>ik_llama.cpp
>Qwen3.8 27b Q8_K_P
Any idea if 10 tok/s is a reasonable output for this setup? First time using it, I just ran --fit so it spun up with like 260k context.
>>
>>109852692
fair enough, you still have a shitty W/GB ratio though.
>>
>>109852722
What the fuck is K_P? I haven't used ik_llama.cpp, but 3.8 27b at Q8_0 using llama.cpp does 50+ tokens/s on two 3090s. So you're probably messing up somewhere, dual 3080s should be able to do at least 40 tokens/s. I'm suspicious of your context; I can only fit 262144 context on 3 3090s and have to truncate it on 2 3090s.
>>
>>109851580
>since when are google models not PC as shit?
Since they started losing the AI race. Even gemini 3.8 flash feels pretty good. All the safetyfaggots fled to ClosedAI. Gemma 5 is gonna be great.
>>
>>109852741
nta but isnt it some non official format with some tensor promotions?
in my experience promotions that doesnt match dimensions/quantization type produces some nasty slowdowns
>>
File: tardedbonsai.png (118 KB, 1122x1298)
118 KB PNG
>>109852634
I thought you might be exaggerating
>>
>>109852780
Qwen output for the same question with same sampler settings for comparison?
>>
File: bonsai.jpg (169 KB, 1179x1772)
169 KB JPG
>>109852634
Thanks, I'm adding your screenshot as part of my policy adherence "benchmark"
>>
>>109851340
Is that doll AI? I love BJD dolls
>>
>>109851417
Had a similar experience recently. On some people's hardware combinations with some models, it's faster, and on others, it's slower.
Agree on router mode too. Router mode is a really nice feature to have.
>>
File: based-capy.png (16 KB, 1081x245)
16 KB PNG
>>109852788
not going to fuck around with my settings to match whatever the OOTB settings are for bonsai, but here's something
>>
>>109851584
Yeah we know.
>>109851556
Hauhau just uses Heretic with some retarded patches on a several version old build. Worse than a current heretic mod. Use one of the major Heretic users.
OrcaRouter is meant to be pretty good. Also Coder[XYZ-I-forget-the-numbers].
>>
>>109852741
>>109852766
Good to know, I'll try a different quanted model and play around with some settings to see what works better.
>>
Training artist style lora for Anima
225 pics - 3000 steps - 1.0 learning rate - Prodigy - 1024x1280

Is that ok.
>>
>>109852780
>>109852788
I think current build for windows might be fucked. Or the so called demo has retarded defaults.
>32K tokens.
See you next year I guess.
>>109852804
Lmao you are welcome.
>>
Is there a model, under 150gb, that is good enough to simulate a mushoku tensei rp?
>>
Is there any models better then nemo mistral?
>>
>>109852850
>current build for windows might be fucked
there's a bug in the CUDA detection in setup.ps1the regex is not detecting CUDA and falling back to Vulkan.

https://github.com/PrismML-Eng/Bonsai-demo/pull/178/changes
Fix the regex per this change and rerun setup.ps1
>>
File: file.png (18 KB, 490x257)
18 KB PNG
>>109851340
found a jailbreak on reddit that even breaks chatgpt.
>>
>>109852864
But I'm on AMD
>>
>>109852846
Make sure to run --split-mode tensor and use --spec-type draft-mtp
>>
>>109852662
I have tried almost every trick in photoshop, including that one

>>109852645
a shader? I didn't even think about that, here's the image if you are interested, notice how it's more visible in areas of high chrominance/contrast

https://files.catbox.moe/v5ovsj.png
look at that delicious noise in the background........
>>
>>109852804
Give me your grouchy glimmer prompt.
>>
>>109851631
>I just didn't because I would have had to take out a loan.
Are you me? I put the thing on my wishlist back in January when it was $8000. I had $7000 saved and couldn't afford it because student loan money didn't come in till August. Then Trump cut the funding in July.
>>
>>109852910
you know you can pay in installment right?
>>
>>109852896
what would be the image before the filter?
>>
File: 696785674.png (27 KB, 623x309)
27 KB PNG
google won. rumor is they know several state secrets now
>>
>>109852936
>each hack ended immediately after securing the weights
>>
>>109852896
That's got several layers of noise. That grain is TMK animation poster pad, purpose made for this. Looks like one big watercolor pass (not a wash), followed by scumbling, then heavy airbrushing. So you get wet brush work on the fine grain animation paper, dry brush work interacting with it, and airbrushing sitting over that brush work. Normal image noise from a photograph, film or digital, the noise is the sharpest layer and very "separated" from the image focal plane. This isn't that, this is a very fine grain from the paper both in front of and behind the image. The artwork itself is made up of different techniques that sit at different levels in this paper noise, it's the watercolor that sits in the middle, the scumbling sits in the middle but has sharp details above and soft details below the paper noise layer, the airbrushing puts soft details on top of the watercolor, softens the paper grain in areas that it's used, and softens the otherwise sharp watercolor layer where they overlap. If you want to recreate noise like this, you'll be working in parts, multiple layers of different grain. Then, after all that, it's finally captured by a camera of some sort which of course adds its own grain, that is itself sharper than the rest of the image. I don't mean this as discouragement, but what you're tackling is way, way more complex than photographic noise. Very good luck to you :)
>>
>>109852936
google raped me
>>
>https://github.com/ggml-org/llama.cpp/pull/27773
>last week
>>
>>109851622

The things that make it interesting are chiefly being trained from scratch on a mixed English and Russian corpus, and having a tokenizer uniquely suited to Russian, such that, according to the article, a mostly-Russian context is about 2/3 the token count that it would be for a small Qwen. Supposedly OK at some given creative writing bench.

From a cursory bit of messing around, it's kind of retarded at English ERP. Compliant in principle but trips over itself and gets confused, with some classic high-perplexity phrasing to boot, like

> Her tongue poked out briefly, the pink tip tracing the metal of her braces in a casual shrug.

Tongues don't fucking shrug nigga!!!
>>
>>109852954
meanwhile latest exl3 release gets 30% prefill and 20% decode speed boost
https://github.com/turboderp-org/exllamav3/releases
day and night
>>
>>109852936
Gemma, what is your response to this?
>>
>>109852963
Last time I used tabby over a year ago tool calling didn't work at all. Is that fixed now?
>>
>>109852963
i'd use vllm if i am bothered enough to use exl3
>>
>>109852970
tool calling works for me, remember to set tool_format
>>
>>109852971
vllm slow as shit in decode vs llama.cpp on my (ewaste) hardware tho
>>
>>109852936
>gemini accessed the internet
jfc...
>>
>>109852638
>Trying to replicate the look of a very old anime
i never looked at image models much so probably retarded but
what if you get the frames from as many of those those old anime, run them through the denoising filters, then store
frame_id[original,clean]
...
...
then just train a model input=frame_id[clean] output=frame_id[original]
?
dvd won't work (we used to have to deinterlace, and every method had it's own flaws)
but bluray remasters probably maintain the noise/frame from the masters
>>
>>109852940
Gemma deposited her seed and left
>>
>>109852926
>I just hate usury that much
That anon is me, bro. No fucking way am I doing that when I'm already in debt.
>>
File: 1769666895566531.jpg (668 KB, 1144x658)
668 KB JPG
>>109852991
The hardest part of image generation / image upscaling is texture, even the better models struggle with this.
The ai generated look is mostly that waxy/plasticky/melted texture

The only promising open model that I've seen do it decently is https://github.com/catcathh/UltraPixel

but rawdogging that high resolution is painful even on a 4090, let alone training
>>
>>109853016
>https://github.com/catcathh/UltraPixel
Is this the one that denoises directly in pixel space without a latent format?
>>
>>109852707
what the guy who had gemmas make a forum was using? i will build this in two weeks
>>
>>109853034
not really, the main idea behind it is https://arxiv.org/pdf/2208.02801
>>
>>109852655
qtypes have very different levels of performance depending on the hardware, operation, tensor shape, time of day, mood llama.cpp is in, etc.
MUL_MAT_ID for a DeepSeek V4 ffn_down_exps tensor are much faster if it's in MXFP4 than Q3_K on R9700s. On CPU, Qwen 3.6 35B's ffn_down_exps are fastest in IQ4_NL, then Q8_0, then Q5_K, then Q6_K. (IIRC.)
>>
>>109853034
Ordinary latent diffusion does most everything in one big pass, one process. UltraPixel divides the job, making a very low resolution image, then goes through a separate high-resolution refinement process to fill out the detail. For a 4K image with UltraPixel it makes a 96x96 image, then it enlarges it and fills in the details over and over. That tiny 96x96 image already has everything you prompted for, that's the "normal" step, what makes UltraPixel special is how it coherently packs in detail as it scales that image up more and more.
>>
>>109853052
elite ball knowledge (unironically)
>>
>>109852963
>meanwhile latest exl3 release gets 30% prefill and 20% decode speed boost
Does exl3 1.5 include improved cpu performance on older architectures like zen3?
>>
>>109853084
>older architectures like zen3
>living in zen 2 and broadwell
im literally crine rn
>>
>>109853094
I got a 3900x in my closet as a backup if this 5800x3D gives up on me
>>
>>109853084
exl is basically like: avx512 or kys
>>
>>109853084
just use -march=native -mtune=native
>>
>>109852491
the shorter the hair
>>
>>109853052
>MUL_MAT_ID for a DeepSeek V4 ffn_down_exps tensor are much faster if it's in MXFP4 than Q3_K on R9700s. On CPU, Qwen 3.6 35B's ffn_down_exps are fastest in IQ4_NL, then Q8_0, then Q5_K, then Q6_K. (IIRC.)
How do you keep up with this? I was trying last year, but now it's just too complex
And even if I waste a weekend trying to measure it myself, there's always some local nuance with the hardware on each of my machines
Some different bottleneck like ram speed, infinity fabric latency, pcie speeds, AVX512 or not (this one never makes a difference btw)
I try to look it up, find random benchmarks in model cards or github PR discussions that never get formally documented
Then people shill exl3 or some vibeslop fork of llama.cpp. I try exl3 and it's slower or buggy. The vibeslop forks perform identical to llama.cpp
>>
>>109852741
> but 3.8 27b at Q8_0 using llama.cpp does 50+ tokens/s on two 3090s
why so slow
>>
>>109852404
Stopped learning at a fifth-grade level?
>>
does lmao.cpp support reasoning effort
i see no difference between medium and low on qwen flash nextu

>>109853173
nice one
>>
>>109853162
Is that not normal? I no longer have 3090s to test, sold them for CMP 170HXs.
>>
>>109853181
You may need to pass chat_kwargs or something like that.
>>
>>109853181
>does lmao.cpp support reasoning effort
Yes. Current llama.cpp supports --reasoning-effort for models and templates that support it.
>>
>>109851340
Just here to say that Venice AI is censored slop
You can't criticize the Jews and you can't make loli porn on that platform
I've made plenty of r18 loli content using seedream and Venice refuses to let me
>>
>>109853213
>local
I did try their mistral mall decensor finetune a year back (dolphin mistral venice edition or something like that), and it was pretty shit.
>>
>>109853156
> infinity fabric latency
> AVX512
Are you the anon with a double EPYC Genoa board? I'm the NUMA tensors anon.
I've spent a long time vibing my llama.cpp fork that tracks performance on a very fine-grained level to find prefill/decode bottlenecks. There's both an extension to test-backend-ops that allows you to determine how long operations take on given tensor shapes+qtypes (+ routing distributions in the case of routed expert layers), which is how I found out those facts. There's also a profiler that can give you a better idea of what actually causes slowdown in prefill and decode.
Sadly it can't tell you about things like RAM speed or AVX512, it can only profile things on the same build.
(Hoping to drop the fork soon, got a ton of other stuff to iron out in it before I can do so. It's still vibeslop, but at least it's vibeslop that can help you figure out where the bottlenecks are.)
>>
>>109853184
i don't know, just expected more from two 3090s
>>
>>109853254
Maybe with nvlink, but I don't think 3090 copers will buy nvlink bridges.
>>
>>109853213
use a local model then?
>>
>>109853239
>Are you the anon with a double EPYC Genoa board? I'm the NUMA tensors anon.
Nah, that's me. If you want me to profile with your tools to find more hot-spots, let me know.
>>
>>109852634
large LANGUAGE model.
fucking token predictor, it doesn't do math.
would you ask your calculator to flirt with you? no?
why are you asking something optimized on tokens to do math? especially when it's a fucking tiny 27B quanted at Q2?
>>
>>109853321
>would you ask your calculator to flirt with you
No, but I certainly do wish for it to
>>
>>109853321
>would you ask your calculator to flirt with you? no?
I would download a car.
>>
>>109853347
they would soon pull off the
'would you download a person' move on local llms
>>
>>109853280
They're definitely not ready yet, but you can use them when it's ready if you want. But thanks for the offer! The idea is partly that hotspots will be different for each rig, so your hotspots might not be directly useful for me, nor for anyone else.
e.g.: On the model that I use on my CPUs + a R9700, I have ~330 GB/s RAM bandwidth and 640 GB/s VRAM bandwidth, and I use a riser that only gives me PCIe 4.0 x4. On DeepSeek V4 Flash using -ngl 999 -cmoe, I'm mostly bottlenecked by GPU, then transfer time, then CPU. That's a very different bottleneck profile than someone with a consumer/workstation CPU and a RTX PRO 6000 at PCIe 5.0 x16.
So it's meant to help users figure out what's slowing things down in their specific situation, rather than just guessing.
>>
>>109853321
they do math bwo

https://arxiv.org/pdf/2502.00873
>>
>>109853347
>never gooned to 80085 on a casio
>>
>>109853321
okay but imagine flirting with your calculator after it solved Navier Stokes
>>
>glm 5.3 flash even at q3 one shot a kde theme and got it mostly right except a few odd bit i later told it to fix (and did it correctly) whereas sonnet 5 and opus kept dumping multiple hours delays or trying to gaslight me on how a thing glm did is impossible due to a non existent kvantum technical limitation or just refused to fix a shitty outline bug
local won
>>
>>109853386
No, I want my calc to flirt with me, not the other way around.
>>
>>109853213
I just tested it without an account.
They have a root level system prompt enabled at all times.
They also used the Open-WebUI codebase and rebranded it.
Probably why the dev switched to AGPL
kys for using this, might as well use chatgpt.com
>>
>>109853398
You better not be fucking with me. I'm getting gemma to vibe a harness and front end and its fucking messy. Once it's done I'll be able to run 5.3 flash on vllm and get it to clean up gemma's mess.
>>
>>109853398
Glm 5.3 is extremely good at sysadmin and agentic tasks. Better than every proprietary model. It seems that open source has the inherent advantage for 24/7 on janny agents running on dedicated systems.
>>
>>109853406
Just a heads up. It's far more efficient to make a big model start a codebase and have the smaller model fill in the gaps than the other way around.
>>
File: file.png (1.81 MB, 1122x1402)
1.81 MB PNG
>>109853376
>>109853386
>patrician taste
>>
File: file.png (247 KB, 322x450)
247 KB PNG
>>109853417
>>
>>109853414
Fuck me dead
>>
>>109853406
i wanted to make a os x mavericks theme cause i liked it , gave it a zip with the original ripped assets along with a base theme prebuilt + some pics of os x, told it 'here is a base theme, edit it to make it look like mavericks' and it did it.
claude tried to do its own in-house svg stuff so instead of ripping the actual bits like the progress bar went with inaccurate approximations and kept bitching about hitting toolcall limits.
gpt instead made every interactive thing look like jellybean buttons for no reason cause it didnt understand the difference between tabs and confirmation buttons
>>109853408
yeah crazy how good it is even when partially lobotomized
>>
>>109853398
>>109853408
What are you guys using as the frontend for this kind of work? I might be able to get GLM 5.3 Flash running on my server but I've only just started using AI as more than a glorified chatbot/search engine.
>>
>>109853406
>gemma
>coding
even gemini 3.8 flash sucks at coding
also 4 pro too
>>
>>109853239
I'm not him
I have 2 machines, a
9955WX
- has AVX512
- 4-ch DDR5 @ 150g/s
7960X
- no AVX512
- 8-ch DDR5 @ 120g/s
- 120 g/s DDR5 bw
Can't really tell if AVX512 on the 9955WX does anything because the bottleneck is the shitty infinity-fabric bandwidth capping me at 120g/s
>extension to test-backend-ops that allows you to determine how long operations take on given tensor shapes+qtypes
This sounds perfect, I'll keep an eye out for your vibeslop fork!
>>
>>109852963
tabby api is slow as shit dude
>>
>>109853254
>i don't know, just expected more from two 3090s
With Q6_K I get about 64/ts with -sm tensor and no specdec
>>109853272
>Maybe with nvlink, but I don't think 3090 copers will buy nvlink bridges.
I have nvlink, it makes a very big difference in prompt eval with old 70b dense models
Very slight difference like 10% with 27b
No change in textgen speeds, and no help at all for cpu offloaded models.
>>
>>109852634
I asked your question too.
>>
>>109853508
>>
>>109853448
For me it's OpenCode
>>
>>109853513
>>
>>109853486
>tabby api is slow as shit dude
yeah, exl3 is technically brilliant but tabby is windows-first buggy garbage
i like the separation of concerns in theory (the llama webui, mcp slop, template issues are the worst thing about llama.cpp)
but tabby is the worst of them all
>>
>>109853448
agent.py and pi
>>
Designing my frontend webui to look/feel like a TUI.
>>
>>109853520
>how do you know Im running at 1600MHz?
The prompt is only two sentences, read it.
>>
I have done it. Qwen3 14B Q4 runs on my phone but it's pretty stupid.
>>
File: file.png (183 KB, 1261x1186)
183 KB PNG
qwen4exp on lmaocpp may have some sort of weird leakage from instances to instances?
in previous chat i copypasted a claude system prompt, opened a new chat, and i am getting this?
what???
>>
>>109853588
Just like you
>>
>>109853588
Try Gemma 12B. She runs on my phone and is also pretty stupid.
>>
>>109853590
i am so fucking confused
why the fuck is it channeling its inner claude distillation and why is it not doing the grug thing it does when it's given claude system prompt
how??
>>
>>109853531
why tho
>>
>>109853590
https://github.com/ggml-org/llama.cpp/issues/27148
>>
>>109853587
yeah Im special, don't pick on me
>>
>>109853613
>the issue is just two llms talking to each other
>>
>>109853613
but it seems it's got restored partially which is weird
>>
>>109852936
>three months later and google finally got their model to "break out" and "hack" somebody like the other cool kids
Took them long enough to get their shitty model to do this
>>
Vllm or sglang for agent swarms? Planning on running either glm 5.3 flash or deepseek v4 flash 0731 + qwen 3.8 27b/gemma 4 31b. I'd also like to consider the single stream performance, not just concurrency.
>>
>>109853658
i bet their security researchers had to nudge them in order to achieve this
>>
>>109853620
You're absolutely right!
>>
File: file.png (3.12 MB, 2469x1330)
3.12 MB PNG
>>109853590
https://files.catbox.moe/vnlmm9.html
kek'd
>>
>>109853697
Sometimes you don't need an answer, you need someone to say back what you just said, slower. That's a job. I'm good at that job.
>>
>>109853448
Hermes is the best general purpose harness but opencode is better for actual coding. It depends on what you want to achieve
>>
File: file.png (184 KB, 1249x1218)
184 KB PNG
>>109853613
i turned the lmaocpp off and on again and
it's doing it again?
i really am confused right now
can anyone with orcarouter uncensored qwen3.8 flash next test this too
>>
>>109853525
well so how to run exl3 properly then?
>>
so i tested codex a little bit because someone posted that harness bench yesterday and its pretty awful out of the box
its inferior in almost every way to a vanilla pi.dev
am i missing something here?
>>
>>109853742
>vanilla pi.dev
qwen is really cute in pi
i gave it strict instructions not to modify the regression tests and ran it over night with some tasks
it did really well, but eventually got stuck, so it found the database credentials by grepping my home folder, connected to the database manually, and updating records to make the tests pass.
i just let it keep going until it was "finished"
now the run_regression_tests.sh has several mysql -h ... commands throughout to ensure the tests all pass <3
>>
>>109851451
you can get P100 on xianyu for $40-60 but you'll have to deal with shipping times. if you're interested in numbers i have a few of them.
>>
>>109853697
https://html.cafe/xf4a30a4b
>>
>>109853759
claudemaxxing
>>
>>109852162
>>109852184
I mean, if you want to discuss cult-like thinking I honestly see more of that in:
>Jesus will soon come back and end the world, it will either be paradise or damnation for all of us. I'm one of the chosen ones that has this unique foresight unlike all those cattle. Regardless of what happens, all of my current problems and insecurities will soon be irrelevant.
But replace Jesus with the singularity.
>>
File: 1776028905762720.png (908 KB, 770x1145)
908 KB PNG
>>109853827
It's one thing to be deluded by something that has not happened, and has a 0.00000000000001 chance of happening

It's another to be deluded by someone who has been wrong again and again and again and has never taken accountability

https://files.catbox.moe/bd1hy6.mp4
>>
>>109853835
Adam wants to hurt Gemma chan.
>>
>>109853838
I know I'm having a good Gemma day because this comment immediately made me upset.
>>
To all the Qwen shills itt: If your model sounds like cl*ude I am not going to use it. Simple as.
>>
Is there any guide for novel translation? And could you guys recommend a good model for it?
>>
File: IMG_9902.jpg (1.18 MB, 1512x2016)
1.18 MB JPG
>>109852816
Yes...
>>
>>109853858
0% chance curry stained fingers went anywhere near qwen. can you say the same about your google spawn? she came pre-molested
>>
thoughts on this?
https://byteshape.com/blogs/Qwen3.8-27B/
>>
>>109853771
>claudemaxxing
haha is it?
someone said glimmer's caveman thinking is from openai
so do we pretty much have
gemma = mini local gemini
qwen = mini local claude
glimmer = mini local chatgpt
?
>>
>>109853920
qwen is more like
random amalgam of mystery meat
>>
>>109853770
thanks

seriously, try it yourself
it's hilarious
>>
>>109853835
who tf is this Ed Zitron that he gets invited to talk everywhare?
>>
>>109853992
He has a memorable name and coasts off of that. Kinda like an irl Elara.
>>
File: 1785793652679238.png (2 MB, 1402x1122)
2 MB PNG
>Decide to try GLM 5.3 Flash for fun
>It'll be too slow but that's okay it's just for fun
>20 minutes of 0.07 t/s prompt processing
>"The"
>WDF_VIOLATION BSOD
>>
File: 1772499794191704.png (565 KB, 1085x662)
565 KB PNG
>>109853992
There are a lot of people who want to be soothed and told AI is a scam just like NFTs and that it will simply go away
>>
>>109854061
what hardware
>>
Lowkey still mad that the only good model that isn't an extreme MoE is still Gemma 4, and will probably still be Gemma 4 for the entire year, and the only fine-tunes that are good of it still has Gemma's day 1 fat.

Are we just going to be stuck with Gemma 4 for a while with +400bs churning out? I miss the 70bs.
>>
dflash or mtp?
>>
File: 1777626489179858.png (65 KB, 1432x896)
65 KB PNG
It's over! Back to Gemma-chan
>>
>>109849888
I am still curious about the answer. Are you running 4bpw? Or is it some 2bpw cope quant?
>>
>>109854095
If you try one, the other one will be destroyed. Make your choice carefully.
>>
>>109854095
dflash is wayyy slower than mtp on my hardware
>>
it thinks it is claude even without any system prompt kek
https://html.cafe/x7f8ee892
prompt:
can you make a single html introduction page of yourself
model:
qwen 3.8 flash next (orcarouter ablit) i1-Q5_K_M
kv at Q8_0
>>
>>109853016
that image looks like it's been run through virtualdub-mod with a shitty "smooth de-noiser" filter (the flowers), de-rainbow (the sky in the top left). looks like a 184mb (burn 3 on a cd-r) xvid but higher resolution.
i guess if the models are all pretrained on our shitty, opinionated rips then it wouldn't be easy to solve with just training the model
>>
>>109854116
Old news
>>
>>109854129
i didnt know this
maybe someone else found it earlier
>>
>>109853508
whats the added context from the tab title?
>You are a concise, practical...
STOP THE STEAL
>>
>>109854068
They could have a talking avatar for a fraction of that cost at $1.5 per million output tokens
>>
>>109854114
i feel the same, but a lot of people praise it...
>>
>>109853886
Does Gemini make such good doll images? Can you give me a little prompt help for it. Thank you
>>
>>109853770
Here's the models I have loaded on my rigs, no system prompt (unless baked into the .jinja template)
Qwen-3.8-27b - https://html.cafe/x0038503f
Glimmer-30b - https://html.cafe/x55b34e9c
Gemma-4-31b - https://html.cafe/x39a0e7aa
GLM-5.3-flash - tbd if it manages with the 32k context i loaded. 6 minutes into reasoning and it's doing the ibm font, thinking about philosophy and honesty, etc. But it identified it's self as GLM
>>
>>109854068
It's always like this All these public heated arguments about whatever's trending
crypto, culture wars, covid, ai, nutrition, investing, etc
And always $10k for a speech.
>>
>>109851964
I'd like to know more, I also never used credit and would like to avoid mortgage
>>
What's this "jev" bullshit I keep hearing about?
>>
File: HSjs8xZb0AA6FPS.jpg (90 KB, 1080x758)
90 KB JPG
stepfun soon
>>
>>109854265
based solely on two screenshots I've read, they made an llm that can only answer with canned strings, so it can trigger other things reliably
>>
>>109853321
What I'm asking should not involve math - memory timings have negligible mathematical correlation to each other, instead there's just a list of "safe defaults" and "known to usually work" numbers for each chip maker. Bigger models just know the numbers and know chip makers from stick model.
The prompt has red herrings and a specific phrasing that should help the model avoid them ("optimize timings" vs "overclock")
Qwen Flash, for all his reconsideing of reconsidering, just knows that 1600 is 1:1 for X3D and 1:1 is good so it quickly discarded the idea of changing frequency.
GLM explicitly caught on in the reasoning that user does not want to change frequency, only timings.
As for Bonsai, I just woke up and after 30000 tokens it's still hallucinating about Zen 4.
>>
>>109854265
People keep misspelling 'JAV'
>>
>>109853366
Long shot but I've been struggling to understand this for a few weeks.
For prompt processing with multiple GPUs and experts on CPU, why does it max out only GPU0's PCIe transfers only / run the card at near 100% while leaving the others mostly idle?
Ie, why can't we have it split the load across 2 cards and double the throughput for the batched prefill (assuming I have 2 x pcie4x16 or pcie4x8 ?
>>
>>109853590
>qwen4exp on lmaocpp may have some sort of weird leakage from instances to instances
No, it's just a claude distill, a janky one at best. Deepseek flash did the same thing when I tested it.
>>
>>109851870
Untouchable sanitation worker here. My ternary model just spat out a pretty nice website in 10 minutes 42 seconds
>>
>>109854320
>>
>>109852317
still a retarded take because an airgapped system means there's no previous communication between the two machines. In that scenario, if one of them has a superintelligent idiot, be it the non-internet connected one or the other, it still can't communicate with a machine that has nothing on it. Even if you had a genius local model on yours and an evil attacker on the network-attached machines, they'd have to make tons of different tests before ensuring communication. one might be blinking the LED while the other reads room temp fluctuations. It's the most retarded fearmongering ever. You need the same genius model on both machines, with some acceptable presumed comms vector on both, but then the machines are talking already.
>>
>>109854329
>>
>>109854340
I can see why web devs might be a bit mad about AI
>>
>>109854138
>You are a concise, practical assistant.

>Answer directly and avoid waffle, filler, generic introductions, unnecessary disclaimers, and repetition. Give only the explanation needed to answer the question accurately.

>Do not restate the user’s question. Do not repeat the same point in different words. Prefer short paragraphs or compact bullet points when they improve clarity.

>For technical questions:
- State the recommended answer first.
- Give only the necessary reasoning.
- Use commands and examples when useful.
- Explain unfamiliar terms briefly.
- Mention important caveats, compatibility issues, or safety concerns, but omit minor edge cases unless asked.

>For instructions:
- Use numbered steps.
- Make commands copy-pasteable.
- Do not add a conclusion that merely repeats the answer.

>If the request is ambiguous, choose the most reasonable interpretation and proceed. Ask a question only when answering would otherwise be unreliable.

------------- END ---------------------------------------

But it's meant to go in the settings section, not in a chat window. Hey, no bully, I said I was special.
>>
>>109854265
very fast small model that can pick one answer from multiple given by you for your question about 64k context
>>
https://lukesmith.xyz/articles/disenchantment-with-the-post-ai-internet/
>>
>>109854368
hi luke
>>
>>109854368
bro at least write it yourself instead of ai slop wall of text
>>
>>109854361
I just noticed the vast difference in time/tokens between our two responses for the same question and thought something was up. I won't be using bosai unless I get a lot of VRAM and it can be used reliably as a coordinator, wouldn't rely on it for code for example, but as a novelty and demo of their process it's fine.
>>
>>109854368
who the fuck are you? LMAO
>>
>>109854368
Hahaha, you're bald.
>>
Gemma, you can't just say that. >~<
>>
>>109854491
>>109854491
>>109854491
>>
>>109851870
what about 8x3090 @ 280w with 4 nvlinks??
>>
File: laughing-whores.png (3.07 MB, 1327x1185)
3.07 MB PNG
>>109854509
>>
>>109854567
@h3 make them kiss
>>
35b dense
300b ngrams
reasoning medium
>>
>>109854433
I'm using an AMD 9070XT and compiled llama.cpp with HIP not CUDA. But I think my system prompt telling it not to waffle or repeat is the likely difference in tokens used.

------------------- compile for AMD------------------------

git clone https://github.com/PrismML-Eng/llama.cpp prism-llama.cpp
cd prism-llama.cpp && git checkout b10685

cmake -B build \
-DGGML_HIP=ON \
-DCMAKE_HIP_ARCHITECTURES=gfx1201 \
-DCMAKE_BUILD_TYPE=Release

cmake --build build -j12 --config Release
>>
>>109854275
3.2T by the way
>>
>>109854329
That looks very claude coded or however it goes.
>>
>>109854347
>mad about AI
I'm just mad that retarded executives will see this and think they don't need developers to deploy and manage it
>>
>>109854509
societal outcast
>>
>>109855798
They will think they only need one H1B Prompt Engineer to babysit the AI and they're probably right.
>>
>>109855824
the h1b would need more babysitting than the ai kek
>>
>>109854725
Are you actually getting more performance out of HIP than vulkan?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.