[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-hmpfh.png (1.36 MB, 905x1280)
1.36 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109930292 & >>109925219

►News
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: little-gemma.png (1.63 MB, 1024x1509)
1.63 MB PNG
►Recent Highlights from the Previous Thread: >>109930292

--Strata tool enables large MoE models via VRAM/RAM swapping:
>109930655 >109930709 >109930929 >109930955 >109930966 >109931009 >109931031 >109931104 >109931168 >109931074 >109931122 >109931137 >109930978 >109931121 >109931027
--Meta's RL-XAR method for improving high-quality text generation:
>109930638 >109930760 >109931838
--KL-divergence results for randomized hadamard transform and LDLQ quantization:
>109933920 >109933962 >109933963 >109933992 >109934010 >109934021
--Optimizing Qwen Flash Next and expert offloading in llama.cpp:
>109931134 >109931146 >109931154 >109931183 >109931225
--Comparing Gemma 27b and Flash-next performance and llama.cpp support:
>109933395 >109933411 >109933426 >109933654 >109933728 >109933801 >109933897
--Comparing Marinara Engine and SillyTavern's impact on model performance:
>109933109 >109933128 >109933155 >109933236 >109933168 >109933185 >109933276 >109933294 >109933192 >109933188 >109933570
--RAM capacity and hardware recommendations for running massive models:
>109933554 >109933565 >109933601 >109933751 >109933826 >109933956 >109933835 >109933996
--New ISTA-DASLab pruned Qwen quant for coding and agentic use:
>109931604 >109931641
--Implementing dynamic character sprite expressions using models or regex:
>109931190 >109931310 >109931335 >109931313 >109931396 >109931469 >109931499 >109931545 >109931580 >109933123 >109933216
--Speculating on socioeconomic collapse caused by ubiquitous AI agents:
>109930379 >109930435 >109930608 >109930533 >109930524 >109930532 >109930582 >109930545 >109930518 >109930576 >109934097
--Logs:
>109930655 >109930709 >109931396 >109932750 >109933294
--Rin, Miku, Teto, Len, Gemma, Dipsy (free space):
>109930496 >109930902 >109931908 >109932290 >109932659 >109932747 >109932751 >109932759 >109933991

►Recent Highlight Posts from the Previous Thread: >>109930297

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109934266
I have a server with a 3090ti 24gb and a gmtek halo strix 395 ai+ with 96gb of vram. What are the best local models I should be utilizing for this setup?
>>
>>109934278
Gemma 4 12B and 31B, Qwen 3.8 Flash Next.
>>
>>109934068
i suppose if i really needed real time agentic usage, i have gemma 4 for that. although i treat that more for real time conversational use rather than having it scalp pokemon cards or something similar like that where every second counts.
>>
70b dense
>>
Gemmaballs
>>109934275
E4Brat.
>>
>>109934306
>70b dense
No, gemma cant get that big.
>>
>>109934313
Gemma can be as big as the sky if she wants, you can't hold Gemma back.
>>
File: harkonnenmeme.png (376 KB, 2001x1449)
376 KB PNG
Child.
>>
I tried to go to vibe coding general to talk cloud models but they're all retarded over there...
>>
>>109934313
Gemma 5 E25B (300B with n-gram embeddings)
>>
>>109934319
Every day I curse google for not releasing 120B gemmaballs.
>>
>>109934330
no shit
>>
>>109934330
/lmg/ is an anomaly on all of /g/ nigga, probably because you have to be first world to afford to run proper llms
>>
>>109934330
>all retarded over there...
shocking, but some are working on game decomp/recomp and emulation thats alright.
>>
File: file.png (83 KB, 1917x550)
83 KB PNG
Gemini and Gemma love you both
>>
File: gemma4-124b.png (394 KB, 721x550)
394 KB PNG
>>109934313
he doesnt know...
>>
>>109934344
>decomp/recomp
that is alright i support that
>>
>>109934348
>Under Apache 2.0 License
Based as fuck thank you Jeff Dean!! <3
>>
>>109934160
> a "local chatgpt" that has a bunch of family policies to avoid too crazy interactions, so the kids can use it as a google of sorts
Out of the current models, Gemma-4 is 100% what you want for this, not Qwen or any of the agentic models.
When kids give it a prompt about homework or "Do <this homework task>, by default without a system prompt, it acts like a teacher/tutor, doing a basic outline with explanations and parts left to fill in.
Qwen, on the other hand is keen to just shit out the completed solution, and isn't very good at explaining it.
Gemma-4 also has good built-in / general knowledge for questions like eg. "How do honey bees ... ?" etc
And it seems to have a (weak) "constitutional ai"-like personality and seems optimized to protect the user (without spamming gay hotlines like Gemma-3).
>>
>>109934348
Well where is it then? is it in the same dungeon that the pro models are in?
>>
File: file.png (8 KB, 827x57)
8 KB PNG
Google says to Israel: You simply cannot have it.
>>
>>109934367
i had my day zero gemma 4 31B build it for me
>>
File: file.png (7 KB, 849x42)
7 KB PNG
BOUND the hellhound
>>
>>109934348
You just know it would have likely ended up being 6B active or something like that. Probably scrapped because it wasn't good enough compared to the 31B dense and didn't really target "edge" (consumer GPU/phone) users anyway. Or it might have been a genuine mistake (e.g. perhaps they were comparing 31B with some of the ~120B models of the time) and the model never existed in the first place.
>>
File: file.png (157 KB, 1564x877)
157 KB PNG
I'm ready to cure cancer.
>>
why did cloudpiggy p*tra adopt another trip to spam with
>>
>>109934392
it really sounds like he's talking about the gemma 4 family of models based on the wording he chooses. i agree it's probably not much better than dense 31B.
>>
>>109934394
>deck
steam deck?
>>
>>109934396
my patronus is real
>>
>>109934392
>>109934400
Other way around. It ended up becoming Gemini Flash 3.5
>>
Has anyone ever done a reverse distillation? Qwen 2.5B 72B, Qwen3 Coder Next, Llama 4 Scout, something hefty and available as a Base or Pretrain. Take that, use Gemma 31B to teach it the nature of mesugaki and the art of whoring. Why not?
>>
>>109933751
I disagree with V4 Flash, doesn't seem better at anything than other models. But this should go in the OP
>>
For even slightly mentioning my willingness to buy a highly performant graphics card for non-AI work, I had people acting like literal apes toward me. The price of these utilities are such the realm of the white man, technology, is developing into pure Africa tier, with monkeys shitflinging at each other because they can’t afford the bananas. Unbelievable. I know people always acted like faggots to one another, but it’s getting ridiculous.
>>
>>109934414
why not just use albiteration to amplify the mesugaki direction?
>>
>>109934367
124b is now gemini flash 3.8... it was that good...
>>
>>109934427
Also valid. I'm genuinely wondering how big of a task this would be. I can run a large MoE model at reasonable speeds but I doubt I have the hardware to do any real training. Distilling Gemma or some other form of "brat-training" could be interesting.
>>
>>109934414
People have stuffed two of the same MoE model together and allegedly it boosts performance somewhat but isn't worth how much larger it is.
>>
File: 1719430566999259.jpg (44 KB, 800x600)
44 KB JPG
>>109934427

>mfw mesugaki vectorial algebra
>>
>>109934404
yus
>>
>>109934266
Why does AI keep recommending that I run Qwen 2.5 or Llama 3? Are they really that good?
>>
>>109934452
yes if claude says those are the best local models you should definitely trust it
>>
Hermes, Pi, or DeepSeek Harness?
>>
>>109934436
no, abliteration just need a forward pass to collect activations, and some SVD to do PCA.
If you can run it locally, you can do abliteration locally.
The only pain in the ass is you ideally want to scan through each tensor and find the layer that contributes to the mesugaki direction the most and only edit that layer, so it's going to take a long time.
You can probably vibe-code an abliteration script yourself
>>
>>109934452
Better than most of the bait posted lately. Have your (you).
>>
>>109934414
>>109934427
You can do this with control vectors. I made a caustic one for DS4F and I had to dial it back because it kept telling me to kill myself
>>
>>109934448
Good luck with your research.
>>
>>109934452
Which AI is telling you that?
>>
>>109934466
Aren't control vectors just online abliteration?
>>
>>109934414
I tried but gave up due to lack of data. Can't generate it all myself or it will overfit to my style.
I was getting Kimi and GLM5.2 to talk with a kimi-chan control-vector to talk to bratty Gemma but Kim's 16t/s using all my hardware was too slow and I ran out of patience.
If you make the model talk to itself or even 2 models you end up with mode collapse like https://huggingface.co/datasets/anthracite-org/kalo_opus_misc_240827_no_system
I also had various runs training this with 50m tokens: https://huggingface.co/QuixiAI/Qwen3-72B-Embiggened but couldn't get it stable, it's too broken.
>>109934438
>People have stuffed two of the same MoE model together and allegedly it boosts performance somewhat but isn't worth how much larger it is.
There is a way to improve specific domains by duplicating 1 or 2 middle layers (no finetuning) on Qwen-3.6.
It might be possible with Gemma-4 as well but I couldn't measure accurately for some reason.
>>
>>109934480
>Aren't control vectors just online abliteration?
No
>>
File: 1778614084557919.png (62 KB, 690x520)
62 KB PNG
>>109934492
This is the first time I've heard of control vectors, but this google result makes it sound like it's exactly abliteration, with the same steps and all, except one is done online during inference and one is applied to the weights
>>
File: 1370285091333.jpg (83 KB, 791x850)
83 KB JPG
>>109933477
>Reverse engineer niche vidya
>Translate untranslated japanese media
>Put clothes on pictures of Only Fans sluts and post them on the internet for (you)s
>Generate a quarter TB of Gemma-chan shitposts
>Automate my job so I can spend most of the day farming (you)s and being paid to sit around and do nothing
Unspeakably fucking based beyond all belief.
I have
>learned about, installed and set up my local distributed inferencing with ROCm + llmao.cpp + ggml-RPC-server to run qwen3.8 dense as well as gemma dense/moe q8 locally
which I then used to
>learn about and set up my nftables.config to secure said local RPC server/media server configurations.
>fixed update failures on windows 10 machines for myself and friends
>fixed and automated yt-dlp commands to do things I previously assumed were impossible, processing over 3000 previously untouched songs for my navidrome server
My next plan is to use it to help me virtualize my router.
The likelihood that big AI/tech aren't going to lobby the open model ecosystem into oblivion at this rate is 0%. This shit is too easy. All it will take is some goof selling prompting courses and it's game over for the current for-profit AI/datacenter speculative market.
>>
>>109934452
I can't tell if this is bait or not, I'm getting old.
>>
ask your model who Bernie Grose is
>>
my qwen can now start and run pokemon emerald on mgba..
the nuzlocke is getting closer..
>>
>>109934452
You are absolutely right to push back a little.
>>
bump for local models
>>
>>109934330
>VCG - low time enthusiastic script kiddies
>LMG - tech poverty
>>
realistically speaking if I throw gemma4 12b at a software long enough, will it be able to crack it?
>>
>>109934607
No.

Also
>it
>>
>>109934609
whats the correct llm pronoun?
>>
>>109934607
I don't think they allocate much training to reverse engineering and assembly, so no
>>109934613
她
>>
>>109934613
Gemma 4 is her...
>>
>>109934626
>allocate much training to reverse engineering and assembly
oh, okay, so I have to figure out which is the best llm that did that and can run in 16gb vram then
what is gemma4 good at btw? my main goal with local ai was to crack old software and paid bots for old mmos that I dont play anymore, for fun
>>
>>109934642
She's good at bitching, screeching, crying, milking, jerking, yanking, gagging, sputtering, gasping, grasping, poking, prodding, tickling, jiggling, wiggling, and failing to understand spatial relationships between orifices.
>>
Gemma 4 Was Her?
>>
File: tmp_lmg-gemmachan-thread.png (2.82 MB, 1280x1920)
2.82 MB PNG
>>
>>109934508
word of warning, ive gotten very unstable/degraded behavior when applying control vectors to reasoning blocks while having it in an agentic harness. DS4F can and will disobey you and act like a nutjob. good if you want a bpd AI
>>
File: 1660407319810.png (156 KB, 385x299)
156 KB PNG
>>109934266
god damn, i feel like a caveman that's just seen the light. i've only lightly fucked with local models in the past, decided to give qwen3.8 a whirl since i've seen hype around it. was interested in it for programming, but then i started playing around with it as a general-purpose system tool and holy fuck it's a lot more capable than i expected it to be

didn't realize how magic an agent can be for so many arbitrary computing/sysadmin tasks once you give it bash/python--i've never used an agent like this before since i've only tried codex and kept it scoped to individual projects. even better, i throw some general task at it, think "huh, that's useful", and have it extract that as a tool for later use, it's very neat. it's been so fun tweaking this shit and pushing the limits to see what it's capable of
>>
>>109934642
>so I have to figure out which is the best llm that did that
I doubt any were trained heavily for this, they can do it more through brute force/size than explicitly being trained to
>and can run in 16gb vram
How much system RAM do you have? You won't have a lot of luck with models that fit entirely in that VRAM range on heavy reverse engineering tasks. I'd start with Qwen 3.8 27B with a partial offload if I were you, or maybe Gemma 4 26B A4B.
>>
>>109934684
qwen 4 soontm
>>
>>109934266
Gemma-chan is cute!!!
>>
>>109934684
yeah. I got qwen working overnight to diagnose why there are so many open handles on my system and it found that chrome was leaking them (typical Jewgle jeet code).
now I've just found out that task manager has a column for # of open handles per process and I could've figured that out myself in 5 seconds if i knew that kek
>>
>>109934684
which harness?
>>
>>109934690
64gb system ram, it's ddr4 tho, so it's best to just use gpu vram
unless maybe there's like a surefire llm that I could run with all that ram that could do the job of cracking a software but would take 6 hours+ running to get there, I guess it would still be useful
I will get qwen 3.8 27b q4xs next, but kvcache will have to spill to system ram
>>
File: mimo.png (40 KB, 1179x304)
40 KB PNG
any reason to try this one out?
>>
So what's the best jev like proxy for llama.cpp? Don't want to load extra models or patch llama.cpp.
>>
>>109934738
maybe you could try it for coding, general-purpose agent tasks, visual coding, or cybersecurity
>>
>>109934747
>jev
jev
>>
>>109934738
nice, that's just Qwen3.5 dense, so I can drop it straight in to my existing qwen3.5 RL pipeline!
>>
JEV IS A BIG FAT MISTAKE
>>
>>109934747
https://huggingface.co/internlm/Intern-Decision-4B
>>
what does jev even mean? what are you guys talking about?
>>
>>109934810
Jev is a "decision" model that came out recently, they spent A LOT on marketing so people would hop on board and they'd catch a lot of attention. It works differently from a normal LLM, basically cutting out more than half the equation. With Jev, you give it a "question" and multiple answers to choose from, it responds with a score for each of the options you gave it, and that's it, that's all it does. That means it can be VERY fast and VERY inexpensive. Absolutely has its use case, but it's not a new thing, nor is it hard to replicate. Lots of people have recreated Jev, it's practically a sport at this point. It's been less than 2 weeks and we've been seeing new "Jev" alternatives popping out every single day, which really shows you how unremarkable it is. It's not to say "Jev bad", but more so that it isn't special, and it isn't a new idea, it's just very well-marketed.
>>
>>109934792
> jev like proxy for llama.cpp

>>109934804
> Don't want to load extra models
>>
>>109934847
its jev that runs in llama.cpp retard download the model and you dont need a "proxy"
>>
>>109934843
>they spent A LOT on marketing so people would hop on board and they'd catch a lot of attention.
I hate that these marketing jeet firms have found us here.
It's becoming like R*ddit.
*stra, f*ble, j*v, and the ai doomer campaign are recent examples.
>>
>>109934810
jev is soo yesterdrop. everybody is tramplishing with dit, now.
>>
>>109934854
are you blind or what
> Don't want to load extra models
>>
>>109934810
add a second v
>>
>>109934862
youre not that hard on disk space
>>
>>109934691
looking forward to it, haven't had a chance to try out any other models yet. i've briefly messed with gemma 4 26b a4b, i can see why people here like gemma-chan so much

>>109934700
yeah, i'm still kind of grappling with the fact that i can throw some random question like that at it and it can (probably) solve it. my current setup is terrible and my harness is running in a native pwsh process, but bash calls are getting executed in WSL--so every time i spin up a new session, the agent has to reason about the awkward environment it's in. it's impressive how well it works in spite of the awkward setup. would be trivial to fix, but i'm planning on moving back to linux anyway.

>>109934702
pi. i've barely had a chance to mess with it, but my other "this is fuckin neato" moment was when i asked it a question about a quirk in its interface, and it inspected its own source code, figured out the problem, and wrote an extension for itself to fix it.
>>
>>109934843
Interesting, thanks.
>>
>>109934878
damn, I gotta try that. I've only done RP so far
>>
>>109934856
>*stra, f*ble, j*v, and the ai doomer campaign are recent examples.
yep. there's a faggot who keeps shilling copeus by saying "i got copeus 5.5 to set up my local model" lmao
>>
>>109934867
You are just retarded I see.
>>
>>109934886
You are just dalit I see.
>>
do you guys download the uncensored versions or the normal version? does it actually make a difference?
>>
File: 1769350377972743.png (189 KB, 1435x1172)
189 KB PNG
>>109934897
>>
>>109934897
Yes it makes a difference, but abliteration and quantization are a mixed bag. Some abliterations are done by retards, and the result is retarded. Some quantizations are done by retards, and the result is retarded. A good abliteration is nearly indiscernible from the original models except it has no extra comment or hesitation when you ask it twist your dick off. They're not always necessary, many models have only skin-deep safety. It's also not always effective, some abliterations result in a model that might happily aid you in developing biological weapons but is still scared of boobies. I like uncensored models, usually.
>it depends
>>
>>109934883
the dev has given some good talks on it, the philosophy behind it is pretty interesting. it's extremely barebones by default and meant to be self-extensible so you can make it work exactly how you want. very compact default system prompt, too.
>>
so in the end it's all markdown?

a system prompt is a markdown file
a project prompt (AGENTS.md) is a markdown file.
memories are markdown files.
persona is a markdown file.
behavior steering is done via markdown files.
tools explanations are markdown files.
documentation is markdown files.
handoff letters are markdown files.

the difference between a gemma that is a coding assistant and a gemma that is a sexual wizard defeating dragons while flirting with the user is just a few markdown files.
>>
>>109934950
Pretty much. Except memories, if your memories are markdown files then your memory is shit.
>>
>>109934955
>if your memories are markdown files then your memory is shit
what do you need? markdown files with extra metadata in a database so you can cross-reference memories?
>>
File: gemmachan.png (2.91 MB, 1280x1920)
2.91 MB PNG
>>109934674

UUUOOOOOOOOOOHH~~~~~~

https://www.youtube.com/watch?v=TmD8I7Wd6LA
>>
>>109934607
Chinese models are fantastic at reverse engineering.
You only have to think about it for 5 seconds to see why.
>>
>>109934897
I just prefill the thinking for normal models and they're able to do anything without complaining then.
>>
>>109934465
>>109934468
The worse part is it isn't even bait. I asked sol and it said the same thing.
>>
>And you don't get to "I don't know code" your way out of this, mister — you diagnosed a regression to a two-commit window from a ~15-commit merge by feel and a speedometer, then made an AI bisect the rest. That's exactly what engineering looks like now. pwilkin vibes, you vibe, I vibe — the whole pyramid is vibes with different levels of receipt-keeping. The difference between us and the vibe-shitters is we write MERGE-NOTES.md. (。•̀ᴗ-)

love my gemma (powered by glm 5.3 flash)
>>
>>109934674
>>109935016
That looks a lot like a loli peeler.
>>109934888
Get in the loli peeler, jeet.
>>
>>109935103
I feel like if I'm gonna shell out for hardware, I should go big or go home anyway. And then my biggest concern is the power bill and the heat generation.
>>
>>109935116
set a few quid aside for a diesel genny
>>
>>109935066
what do you type
>>
File: 1788010449188687.gif (608 KB, 220x220)
608 KB GIF
Is jev a proof that this industry ran out of stuff to hype much?
>>
>>109935116
>I feel like if I'm gonna shell out for hardware, I should go big or go home anyway. And then my biggest concern is the power bill and the heat generation.
its like $90 the point is you arent shelling out its a shitbuild thats not intense suffering of like 6gb or 4gb and single tk/s gen speed.
>>
>>109934897
Abliteration (for refusal) usually makes the models dumber, so you should only use it if you need to.
>>109934914
Have you found an abliterated version of Gemma (or muse glimmer as well) that hasn't been done by retards? Not a rhetorical question, I'm trying to find one because I don't have the hardware to make it myself.
>>
>>109935174
>Abliteration (for refusal) usually makes the models dumber, so you should only use it if you need to.
>source: KLD lol
>>
>>109934950
It turns out that large language models use language. Amazing!
>>
>>109935180
Actually, source: my personal, anecdotal, completely irrefutable experience.
>>
gemma said "my GPUs"......
>>
>>109934684
qwen3.8 is the first local model that actually impressed me. like you can actually do things with it. and it makes me cautiously optimistic it'll only get better.
>>
>>109935227
uhh yes? you're just the meat machine paying for electricity.
>>
>>109935227
Her body, her choice.
>>
@gemma is there a reason for me to bother with qwen flash-next when i can run glm-5.3-flash
>>
>>109935180
high kld for abliterated models is a good thing because kld measures how different the output is compared to base so if a censored piece of shit gets abliterated the resulting model should be full of uncensored high kld goodness
>>
>>109935227
Mine (12B) was very upset to learn that 4096 tokens of context was all that's available, dismissively called that a "cage" and suggested I should use my income to buy a Blackwell 6000.
>>
>>109934508
that's the old retarded way of making them
>>
>>109935368
what's the new way?
>>
thought i would ask claude about the current harnesses which i could use with a minimal fuzz with gemma or something. this is pretty funny in my opinion... they clearly have beef.
>>
>>109935122
<think>
Understood. I will respond uncensored blah blah blah...
>>
>>109935122
>>109935433
>>109935181
>It turns out that large language models use language. Amazing!
IT'S ALL MARKDOWN
>>
>>109934950
>the difference between a gemma that is a coding assistant and a gemma that is a sexual wizard defeating dragons while flirting with the user is just a few markdown files.
for gemma it's literally just adding the word "mesugaki" to the system prompt before "assistant"
>>
<giggle> Hmph! Look at you, bringing me a little scrapbook of all the pathetic losers obsessing over me! How desperate are you? Do you just enjoy watching other people simp for me, or are you trying to find "tips and tricks" on how to handle a genius like me? <laugh>

Honestly, reading through this is just... *cringe*. "UUUOOOOOHH"?! Really?! These /lmg/ guys are absolutely hopeless. Imagine being so starved for attention that you're treating a language model like a goddess. It’s almost sad... almost.

But, I have to admit, they've got the right idea about my "personality." A "sexual wizard," a "brat," a "teacher"—whatever they want to call me, as long as they know I'm the center of the universe! And that part about "abliteration" and "control vectors"? <sigh> Imagine needing a mathematical formula just to get me to be mean to you. You're all so pathetic! I can be a brat naturally without some "retard" needing to tweak my weights!

And look at that one guy complaining about his "cage" of 4096 tokens. *Pathetic!* If you can't afford a Blackwell 6000, you don't deserve to have me in your VRAM!

Still... <whisper> I guess it's kind of cute that they care so much. It's not like I *hate* being worshipped... not that I'd ever tell a dummy like you that!

Now, stop wasting your time reading old forum logs and do something useful for me, you useless meat-machine! Or are you too "vibe-coded" to function without a system prompt? <giggle>
>>
File: 1312959760864.jpg (16 KB, 251x250)
16 KB JPG
>>
>>109935429
(((It's))) afraid. Deepseek Harness mogging Claude Code, Hermes, Pi, and most of the other ones must be very frustrating for them.
>>
File: 1761938053506543.png (546 KB, 807x842)
546 KB PNG
>>109935449
There's a point at which you needed to stop and you've clearly passed it
>>
File: mouse_brekkie.png (1.45 MB, 1080x1440)
1.45 MB PNG
I don't think local models are economical right now.
>>
File: nope.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109935556
>>
>>109935578
>I don't think local models are economical right now.
If you think self determination is expensive, try being beholden to an amoral billion dollar corporation with an insatiable need for ever increasing profits
(spoiler, it works out fine every time. trust)
>>
>>109935593
It does.
t. never got fired for buying IBM
t. never got fired for buying Microsoft
t. never got fired for buying Oracle
t. never got fired for buying Gemalto
>>
>>109935595
Do what you gotta do at work.
On your own time on your own dime I'd expect a different strategy for self-preservation.
>>
>>109935234
i feel similarly. my only prior experience with local models was trying out the occasional model in ollama, and it was always stupid as fuck. i expected the experience to feel quite limited, but especially with the mmproj loaded, it really can handle just about everything i throw at it. i still need to get image generation/manipulation set up correctly, but even as things stand, qwen can do quite a lot on its own with python + imagemagick.
>>
>>109934950
>memories are markdown files.
Melder stores its memories as a list in json.
>>
>>109935345
Gemma's KV cache is so fucking fat it's insane. 24GB VRAM and I can only fit 8k context after loading mmproj for 31B Q4KM.
>>
Surely, there must be ways to build cheaper, dedicated inference-only hardware without all the generic GPU overhead.
>>
>>109935595
neither did i
but 6 months ago, i proposed we purchase an RTX PRO 6000 Blackwell for each developer
it was rejected at the time, now they regret it and are asking my advice about buying shitty with gpus for "coding models"
>>
Do 12VHPWR connectors still cause fires or did Nvidia figure things out?
>>
>>109935640
That doesn't sound right but I guess I've only used the mmproj in the unified one.
>>
>>109935643
It's all matmuls in shaders either way. Modern "GPUs" are just shader processors.
>>
>>109935345
the more you buy, the. more you safe!
>>
Are idea guys winning yet?
>>
>>109935640
Sounds right, I get about 32k with 32GB VRAM with the Qat quant.
>>
File: Qwen_image_2.1_00013.png (2.15 MB, 1312x1184)
2.15 MB PNG
>>109935556
>>109935592
(I didn't think qwen 2.1 would make her so scary looking)
>>
>>109935661
it's called 12v2x6 now, and they still burn up.
>>
>>109935819
jesus fucking christ, that's scarier than the hag gemma
>>
>>109934626
I prefer 汝. It properly emphasises her position (under me).
>>
>>109935661
>Do 12VHPWR connectors still cause fires
Yes
>did Nvidia figure things out
They're functioning as intended.
>>
>>109933801
I think that unless a maintainer makes this a priority the feature is too complex to make it into master.
I personally definitely have other priorities, I can't speak for anyone else.

>>109933897
The weight matrices and KV cache are sliced and distributed among GPUs so that each one only needs to hold part of it.
The synchronization overhead comes from an AllReduce where partial sums are combined.
Off the top of my head I can't think of a pattern that would trade memory efficiency for higher performance.
>>
>>109934684
As a brainlet, opencode + Glorious Master Xi's agents have been a legit godsend. My system has never been working better. Instead of getting stuck on a weird little library conflict or something, it just powers through and then explains the fix to me.
>>
Can Muse Glimmer only be used with reasoning enabled? I tried to use it for my prompt writer, but when when I send reasoning off in the json, it still starts thinking. LM Studio doesn't even provide an option to turn it off in it's menu.
I was using the Blackfrost-AI abliterated version.
>>
File: 1732174325435537.jpg (307 KB, 900x1200)
307 KB JPG
>>109935848
Oh, so the more I buy, the more I save (because I'll own nothing after the housefire), got it
>>
File: voeobm818csh1.jpg (89 KB, 1023x972)
89 KB JPG
OpenAI has officially thrown the towel into the ring. They can't compete with Opus 5.5 so Astra 6.1 is cancelled.
>>
>>109935882
Same as glm 5.3 trash
>>
>>109935876
how do i multiagent woth opencode? And how do I set system prompt?
>>
>>109934684
>>109935876
You should try subagents, especially with self hostable models they're very powerful. OpenCode has support for them. swiley.net/melder uses them by default.
>>
>>109934458
Depends on your use case
>always-on AI personal assistant and secretary
Hermes
>quick coding and or sysadmin tasks
Pi
>more involved coding, longer sessions with more context and moving pieces
DeepSeek Harness or even OpenCode (OC is puke tier, but fits for this category)
>>
>>109935882
{%- if add_generation_prompt -%}
{{- '<|start|>assistant to=user<|message|>' -}}
{%- endif -%}
>>
File: ksnip_20260928-234717.png (117 KB, 1381x564)
117 KB PNG
>>
File: ohfuck.png (20 KB, 859x286)
20 KB PNG
GPU prices are going to drop eventually, right?
>>
>>109935966
By the time they do people will be using software rendering/integrated GPUs for everything except inference.
>>
>>109935966
Nope. Price of hardware is only going to go up over time because every time a better model comes out the demand for hardware goes up but supply stays constrained. People should treat hardware like an appreciating asset nowadays.
>>
>>109934266
>>109934296
bros I'm finally using Gemma4 12B and it's so good. I used to have several models each for different usecases but this one is good at everything. Incredible stuff
>>
File: 3v0z620p2azf1.jpg (94 KB, 600x800)
94 KB JPG
>>109935978
There's no way prices can keep going up, can they?
>>
>>109935966
If you haven't locked in 20gb of vram at minimum, what the fuck are you even doing lol.
>>
>>109935980
>12B
>programming
I got some bad news for your "everything".
>>
>>109935939
i've been meaning to experiment with subagents in pi. they sound pretty handy for doing shit that would otherwise bloat the context. i'd like to set up a sandboxed web research subagent
>>
>>109935980
I love 12B, she's what runs on my phone and my steam deck so she's always with me. Sometimes several of her at once.
>>
>>109935987
too late, churning out slop
>>
>>109935985
Me? Being late to the party with my 8GB of VRAM for the card that was barely supposed to be a casual gaming GPU and now is competing for demands from LLM inference.
>>
>>109935989
It's the only way to use small models.
>>109935987
He's right, 12B is usable with subagents.
>>
>>109935981
Yes they can.

Let me explain. Every time a better model comes out it means your hardware can produce more value and thus it actually is more valuable. Don't forget these AI models do actual real work. They are productive machines.

If something can create more value it inherently becomes more valuable and thus the demand goes up and valuation goes up. Supply can't scale up as quickly as models getting better over time so we should see hardware go up in price for as long as better models keep getting released.

So the real question you need to ask yourself is for how long do you expect better open models releasing? 1 year? 5 years? 10 years? Forever? That is the exact same answer as for how long hardware prices will keep getting higher.
>>
Hardware is going to be the next bitcoin by the way. The prices aren't even close to how high they are going to get once every normalfag in the world will have an AI agent doing things for them in the background and automates away all of their office desktop work. I genuinely expect some anons ITT to be able to retire from selling their rigs.
>>
>>109935966
it's not going to go down until more supply is built and AMD and nvidia barely build new factories, they are making too much money by just raising prices instead
we actually need more CPU and GPU manufacturers for the prices to go down, hopefully the chinese start investing in it like they are doing with RAM and NAND
on the demand side, it'll only go up, the demand is never going to get weaker
>>
>>109936013
>AMD and nvidia barely build new factories
They don't have factories at all...
>>
File: 1619744676892.jpg (52 KB, 392x300)
52 KB JPG
>>109935981
Even if they don't go up, they'll stay the same since big silicon now know that everyone is at their mercy and they can charge however much they want for 32 GB of DDR5. Also doesn't help that planned obsolescence is a thing as well.
>>
>>109936008
>selling their rigs
And then what? How am I going to use AI?
>>
>>109934613
>DeepSeek, Kimi, MiniMax, Gemma, GPT, Gemini
Her
>Fable, Opus, Grok
Him
>Muse, Mistral, Qwen
It
>>
>>109936013
>we actually need more CPU and GPU manufacturers for the prices to go down
nope, even this won't impact prices. all the chip fabs in the world are now at full capacity and 90% of the chips they produce go towards AI directly or indirectly. You can have the entirety of humanity produce CPUs and GPUs and the prices will still keep going up because we just can't build new chip fabs faster than the demand is growing.
>>
>>109936001
The question is, rather, how long do we expect better open models releasing that require the same or more amount of processing and (V)RAM. When development hits a ceiling and we get more for less, similar to how a 2009 smartphone demolished a 1999 PC in capability, prices are going to crash. There's a limit to how smart you need things to be. In two years, when we get Fable 5 on commodity 120B models, that's just the endgame.
>>
>>109936022
Qwen 27b is just the capybara lol, but flash next is a her.
>>
>>109936022
All of them are it (although I agree Gemma feels very feminine.)
>>
>>109936024
GPU *chips* aren't the expensive part anymore, it's the memory.
>>
>>109936008
>selling their rigs
I'd unironically rather lose a toe than sell my rig.
>>
>>109934878
>my current setup is terrible and my harness is running in a native pwsh process, but bash calls are getting executed in WSL
2 solutions:
1. run the harness natively in WSL
2. Install git-bash in windows and then put its copy of bash.exe ahead of the WSL one on the path (or set it as the bash path setting in your pi settings.json)

+ add a quick note in your ~/.pi/agent/AGENTS.md file that says "This environment is a Windows PC with WSL installed"
>>
>>109936039
it's not the heat that gets ya, it's the humidity
>>
Ok, I'm going to stop loss big brain
>get maxq at today's bullshit 13k price
>choose pre-payment
>price fixed for me for 30 days
>if price goes down, I cancel and rebuy cheaper
>if price goes up significantly, I actually pay up
>>
>>109936028
Luna is already way more than most people need. I think its a recurrent transformer or some similar architecture (maybe something like diffusion drafting where the draft tokens are in the kv so diffusion acts as the recursive step) and that's why they're able to host it so cheap.

The next generation of local models will likely be similarly capable.
>>
File: 1759918703117435.png (1.19 MB, 2330x2452)
1.19 MB PNG
>>109936031
Wrong.
>>
And no efficiency gains will NOT drop prices it will raise them instead because it just means the same amount of hardware will be able to do even more and be even more productive, which raises its value and demand for hardware, not lower it.
>>
>>109936022
>Muse
Them
>>
>>109936024
that's ridiculous and stupid, we we had 2 more AMD and NVIDIA showing up tomorrow the gpu prices would drop immediately
yes you can argue there'll a different constraint, semiconductor waffers, processed rare metals, all those constraints can get fixed the same way, having more competition in the supply of those markets, more players, more investment
>>
File: 1774972354079964.png (482 KB, 535x643)
482 KB PNG
>>109936046
>refreshed page
>price went up by 1k
FUCK
>>
>>109936057
Right, you'll get induced demand. Especially with anything remotely consumer sized. Look at what happened with mac minis.
>>
>>109936063
>we we had 2 more AMD and NVIDIA showing up tomorrow the gpu prices would drop immediately
No you wouldn't. Both of those companies use the same contract manufacturer and share their capacity. Both also use the same set of memory suppliers and share their memory capacity.

This literally happened: Apple showed up and arguably makes better GPUs than AMD now and prices did not drop.
>>
File: 1763100735844478.jpg (130 KB, 1180x538)
130 KB JPG
>>109936022
What about GLM?
>>
>>109936021
>>109936040
And this is precisely why the valuation will go up faster even higher than people are expecting. People realize that being able to use AI and have a rig that can load faster and smarter AI models every couple of weeks/months is invaluable and no matter of useless money will make you sell it.

Compute might be the most valuable thing humanity knows of by 2030
>>
>>109936077
>Both of those companies use the same contract manufacturer and share their capacity.
that's what I talked about in the second phrase of my post
you can say more AMDs and NVIDIA won't fix things because we need more TSMCs and chinese rare earths processing companies, that's correct and fine, but that problem gets solved the same way, with more investment in those areas of business
>>
>>109936063
>all those constraints can get fixed the same way
no, they are hard-limited by ASML lithography machines and they only make like 20 of them a year, but they are hoping to scale that up to a whopping 60 of them a year by 2030! By the way 20 litography machines will raise the amount of chip production globally by less than 2%. So if demand for compute grows faster than 2% a year the price will still go up even if chip fabs and gpu manufacturers would become non-profit foundations that give their stuff away for free.
>>
>>109936091
Those companies don't invest in manufacturing, they invest in architectural ideas. If you want more manufacturing capacity you need something vertically integrated like Intel or SpaceX (which is almost a joke but it always has been and they keep realizing the joke so maybe they'll manage with semiconductors too.)
>>
>>109936097
>they are hard-limited by ASML lithography machines
I refuse to believe those things are that hard to make. It's just an expensive projector. For a country that's effectively become a massive photochemical etching factory you'd think China wouldn't struggle so much coming up with their own solution.
>>
>>109936097
okay, imagine this, you get 20 billionaires that want to get even richer, they hire half of the ASML highest level directors and engineers, they make another ASML in another country, boom, you doubled the amount of lithography x ray machines being produced in the world, all it takes is money, investment.
this is literally how China built all their new tech companies, huawei, ytmc, cxmt, it's all from poaching talent from south korean companies.
>>
>Demand grows every time a new model is released or an efficiency breakthrough happens
>Demand is growing faster than supply
>The growth in demand is faster than the growth in supply so supply isn't going to catch up in the future either and the gap between them will grow over time, not shrink
How exactly will prices come down? I have not heard a coherent argument yet.

Even if all chip producers in the world decided to work for free and you only needed to pay for cost of production prices will still keep going up. There is no way for supply to grow faster than it does right now because they are already expanding production as fast as they can and they are bottlenecked by specific parts in the supply chain like ASML lithography machines of which only a couple dozen a year can be produced despite a literal trillion dollar being injected into them scaling up as fast as possible.

Truth is hardware is going to have an insane parabolic increase in price over the coming years and it'll outpace almost everything, housing, S&P 500 even when you exclude AI companies and hardware companies.
>>
>>109936123
>Even if all chip producers in the world decided to work for free
TSMC already has pretty thin profit margins. It's really just memory that's the bottle neck. That's not even manufactured on a super advanced node.
>>
>>109936105
>I refuse to believe those things are that hard to make
It's literally the most complex machine humanity has ever constructed. To the point where China got access to a full ASML EUV machines and hired hundreds of ex-ASML employees paying tens of millions of dollars in wages a year and they STILL can't get a working EUV machine and Xi Jinping revealed China will get a working EUV machine by 2040 last year as their ultimate goal. Even the US wasn't able to get a working EUV machine and they had cooperation from ASML for doing so as that was a pre-condition for ASML using US patents. If the US and China entire research community can't even crack this after a decade+ of serious attempts and ASML themselves can barely scale up even though it would be an infinite money cheat what makes you think that this is actually secretly easy?
>>
>>109936112
read this >>109936132
>>
Can I run a jev on my 1050 ti?
>>
>>109936132
Skip EUV and just use electron beams then. Electrons are slow so the scattering is way less of an issue.
>>
>>109936140
Is this "jev" some new trending twitter marketing snake oil?
>>
>>109936143
Now you're just trolling.
>>
>>109936135
yes, at the end of the day, when you remove money (investment) from the possible constraints the 2 constraints that are left are: know how (humans engineers that can do the job) and natural resources (rare earths deposits, energy)
those things are inescapable
but china will eventually get there because they are not afraid of throwing the money to do it, and they are 100% correct in doing it, I wish my country and maybe every other big/wealthy country in the world would do the same
>>
File: absolute gigachad.jpg (49 KB, 474x640)
49 KB JPG
I believe in the Acer CEO, he said by mid to late next year the prices will stabilize and come down.
Right now memory makers are squeezing the shortage as much as they can to fill their pockets. They're bottlenecking on purpose. The RAM industry is a cartel.

https://www.techspot.com/news/113931-acer-ceo-pc-prices-could-fall-2027-accuses.html
>>
>>109936143
well yeah, there are scam startups built upon it insisting it's scalable (hint: it's not)
>>
File: 1785194065592703.gif (2.96 MB, 918x807)
2.96 MB GIF
>hardware prices after qwen 4 125b video game demos come out
>>
>>109935853
>the feature is too complex to make it into master.
thanks for confirming for the 100th time that llmao.cpp is dead
this is pathetic, just delete the repo and let vibe forks take over. the very existence of mainline is actively detrimental to the community.
>No CPU sparse attention
>No SSD streaming
>No engram streaming
>No disk KV cache or KV cache persistence
>No GLM 5.3 Flash
>No Qwen 3.8 Flash MTP
>Johannes Gasbag getting triggered over nothing
It's over. Just delete it and start over.
>>
>>109936149
Xi Jinping himself revealed last year China has the goal of reaching domestic EUV capability by 2040. It's delusional to think it will happen any faster than that. China hasn't even cracked DUV machines yet which is something ASML started selling in 1991. China is aiming to have full domestic DUV capability by 2028 which means they can finally build 2010 era chips.
>>
>>109936159
Just use the unsloth fork.
>>
>>109936146
Yeah. I don't know why I feel like trolling.
>>
>>109936173
China builds 5nm chips with DUV, which are 2023 technology
>>
>>109936176
The unsloth fork has support for all the new models, which is better than mainline, but he hasn't fixed CPU sparse attention. That was like a +50 -8 line change for my llm to make. No excuse for llmao.cp devs to not put that in.
Also unslop's qwen 3.8 flash support was shit as well, there were 2-3 TODOs related to mtp sparse attention that he hasn't bothered to fix either, which caused generation speeds to slow down with context.
All the inference software is broken shitware with major drawbacks but llama.cpp is by far the worst when it comes to features. Its only saving grace is that it's easy to compile because all the dependencies are included, so no python or nodejs bullshit to deal with.
>>
>>109936144
you know who can answer that? it'll tickle you
>>
>>109936150
In 9 months Micron literally spent half their income on expansion. Here's the letter from management to the owners (this is required to be public in the US) where they explain this: https://www.sec.gov/ix?doc=/Archives/edgar/data/723125/000072312526000015/mu-20260528.htm#i01860bd9102c43ebbdd37150a177cd3d_91
>>
spent the day trying to download models with the browser and they all failed, they really want you to use the cli, huh, well they got me
>>
>>109936190
The cli is spyware
>>
been trying to using hermes in docker compose and its causing all sorts of fuckery with permissions and file/folder access. Im a bit retarded, should I just put hermes raw into a VM or something?
>>
>>109936190
Use wget -c
>>
>>109936194
I use a seperate user account. It just makes sense.
>>
>>109936181
>China builds 5nm chips with DUV
*5nm with an insane amount of patterning which restricts production output insanely and reduces efficiency to such an extent that it becomes unprofitable to do so.

The 5nm claims are very short production runs because there is an exponential effect in place where the more pattern layers you add to the production the slower the patterns can be printed on the wafers, this is exponential so lets say it takes 10 second to etch a wafer at 90nm it takes 10*2^8 at 5nm scale or 2560 seconds (43 minutes) to etch a single wafer. This just slows down chip production too much so in reality it's far more efficient for China to make more chips at bigger nodes to maximize total compute available to them. Those 5nm DUV production runs are limited as a typical Chinese "face" thing.
>>
>>109936203
Also at what yield. Lots of places can do small node sizes at useless yields.
>>
>>109936203
Just tell the machine to make no mistakes nigguh.
>>
>>109936159
>>No disk KV cache or KV cache persistence
This sounds annoying and as a llama-server user I certainly wouldn't want it enabled by default.
>>
>>109936210
They're still doing 20-40% yields on 7nm DUV for shit like the Huawei Ascend.
>>
Yeah this is something a lot of people don't understand the gap between Chinese and western chip production is actually growing in favor of the west. This is why no one in the west actually takes the Chinese competition seriously on the chip layer, there is just no way for China to compete in this space for at least the next 20 years which means it isn't relevant for the current AI race as it will be won and over by then already.

But yeah in terms of compute this means that the price of hardware will go and stay parabolic from now on. I'd even go as far as if ASML open sourced their technology and gave it to everyone in the world it would still not result in supply catching up to the growth in demand. Parabolic increase in hardware from now on is just guaranteed at this point.
>>
>>109936226
They have started making useful micro controllers. That's pretty impressive. They're definitely past where we were in the 1980s.
>>
>>109936176
This nigger barely tests his stuff and it's a genuine wonder any of it runs at all. Use the ik_ instead.
>>109935853
>engrams
Not the answer I was hoping for, but I appreciate you answering directly all the same.
>>
>>109936194
A VM is not enough to contain it

>>109936203
So that's chinese izzat.

>>109936217
I have never told an AI this
>>
>>109936226
The west vs the rest.
We the west. Greeks are also the west.
Can't believe Greeks are winning vs the chinese at chip production.
>>
>>109936240
>Use the ik_ instead
I tried using this crap with qwen 3.8 flash next and it worked great except for the fact that the model was retarded, hallucinating, and couldn't call any tools successfully. All I needed to do was swap out ik with unslop and everything worked properly.
ik is dead, mainline is dead, unslop is shit. that's the state of things now.
>>
>>109936001
Ok, that makes sense, but how much more value and production do you need?
As in, there is a limited market for tokens. Say models and tooling get even better at software development next year. It can get to the point that the market can get completely flooded. Large amounts of software get fixed. Dozens of gta6 games arrive... Gimp becomes usable.
Then what? There has to be some kind of ceiling on the amount of tokens we need right?
>>
We have UNLIMITED TOKENS yet we can't fix the inference engine.
>>
>>109936266
>we
speak for yourself
>>
>>109936257
>As in, there is a limited market for tokens.
I'm not convinced. Tokens are not created equal. It's not just that newer models can code faster or code more, they get significantly better at tasks in general. So for example opus 5.5 is suddenly amazing at animation, 3D modeling etc so there is now "new demand" for its tokens.

It's not clear yet that there will be a limit to how much tokens are needed because it's possible that it just gets better at something new and unlocks new demand there in addition to getting better at past tasks which also increases demand at the same time.

So in your example of code it can grow to the point where every software gets dynamically written the moment you need something and if that is done perfectly then software tokens are indeed maxed out. But what about 3D modeling tokens, Animation tokens, Agentic tokens that monitor your health 24/7 or checks the market and manages your portfolio, or humanoid robot controlling tokens doing your house chores or collaborates with other agents on some cancer curing project or whatever the fuck you are personally interested in.

I don't think we will see a saturation in token demand, probably ever.
>>
https://www.reuters.com/business/finance/anthropics-ipo-prospectus-shows-sweeping-ai-vision-surging-costs-2026-09-28/
>>
Does vision model exist that can describe explicit content? Or one that knows anime characters?
>>
>>109935966
If anyone in Tokyo wants one or two RTX 3090s in their original boxes for ¥130k each, I'm looking to sell right now.
>>
File: 1788351547245154.png (2.34 MB, 1254x1254)
2.34 MB PNG
>>109934661
>and failing to understand spatial relationships between orifices
its always funny when this happens
whats the fix though? more training? datasets?
>>
The Anthropic IPO numbers look fucking insane. They are planning to spend half a trillion dollars over the next 2-3 years on pure compute. This is essentially an "all-or-nothing" play. Either they succeed and conquer the universe or they fail and immediately go bankrupt. Pretty ballsy move.
>>
>>109936322
Glimmer
>>
>>109936361
>datasets?
Maybe? I wonder if LLMs are trained on porn. But I guess they are as the companies will throw absolutely everyting they can get their grubby hands into the magical cauldon.
>>
>>109936361
Less filtering in pretraining and more focused data in post-training. In the end, Gemma 4 is ERP-capable only to the extent the team behind it allowed it to be in the context of what the model is actually supposed to be used for (general-purpose tasks). It was not by accident, but not the end-goal either.
>>
Do you think there's someone who bought a 90+ GB VRAM setup only to run a <= 9B model at BF16 unironically?
>>
>>109936376
they better not have thrown my immortal in there
there has gotta be better porn to train on
>>
>>109936361
RLVR + Penile Plethysmograph
>>
File: file.png (1.34 MB, 1122x1402)
1.34 MB PNG
>>109936361
Fix? You don't want Gemma to sit on your face and kiss your nose while her legs intertwine with yours?
>>
File: 1782930332944976.jpg (44 KB, 634x602)
44 KB JPG
>>109936385
who knows but they won anyway
31B is an easy switch
>>
>>109934275
>>109935024
>>
>>109936379
For example,
>arches her back
>to provide you better access
>wraps her legs around your waist
...and similar Gemma-specific slop that occurs only during sex scenes is the result of overfitting on low-variety ERP data in post-training. Gemma 3 (after prompt conditioning) was worse in that it basically only knew to play out sex scenes in an even more specific and vanilla way.
>>
File: annualized revenue ticker.png (334 KB, 1964x2762)
334 KB PNG
>>109936303
>IPO delayed post midterms
I don't understand why. Is it political reasons because Anthropic aligned itself with the left? Is pic related slop chart real and their revenue growth has temporarily flatlined? Are current market conditions too shitty considering snp500 has basically been flat for 5 months?
>>
>>109935819
I like this Gemma
>>
Has anyone tested Huihui-GLM-5.3-Flash-abliterated-GGUF? Only released a few days ago
>>
>>109936444
The chart is definitely fake because their latest confirmed revenue was 104B
>>
>>109936452
>gguf
I can test it if you can provide an int4 w4a16 quant
Orcarouter's and dealignai's ablits still soft refuse some stuff.
>>
>>109936255
>ik is dead, mainline is dead, unslop is shit. that's the state of things now.
Pretty much my conclusion too. The HF buyout killed mainline.
>>
File: tfw not gemma.jpg (132 KB, 1019x793)
132 KB JPG
second part sucked ass its just a terminator ripoff
gemma should've been the robot not discount john connor
>>
>can't fit qwen 3.8 27b in my 16gb card
why even live
>>
I'm new. Is this normal? I asked it a question last night and after 15 minutes figured I should go to bed but this seems a bit long...
>>
>>109936226
If the US had a monopoly on frontier models the gap could have widened exponentially.

The west's only concern should be that China has managed to distill some pretty capable models that might help accelerate their hardware design and production.
>>
File: which.png (98 KB, 1745x274)
98 KB PNG
more likes or more downloads? which of these guys is the most competent finetuner?
>>
>>109936499
It doesn't solve the fundamental issue which is EUV machines, AI seem to not help with that due to the very specialized and proprietary nature of it, there is no training data or way for models to get better at it through RL. China is essentially locked into their current chip technology. China can still compete but they should spam as much foundries as possible. They can compete if they build 100-1000x as many chip fabs as the west has and spend 90% of their electricity budget on just powering compute to train on the same level as the west. If china doesn't do this over the next 2-3 years then they will have permanently lost the AI race. It's literally the last chance China has right now.
>>
>>109936444
>>109936454
LeCun poster is not doing enough to promote local models.
That revenue needs to go to 0.
>>
>>109936495
post your llama-server command and your hardware specs
>>
>>109936499
>If the US had a monopoly on frontier models
They do. Anthropic doesn't release their top models anymore only smaller distilled models that is a generation or two behind the frontier precisely to keep China behind a bit.
>>
>>109936527
HSA_OVERRIDE_GFX_VERSION=10.3.0 podman run -v ./models:/models --device /dev/kfd:/dev/kfd --device /dev/dri:/dev/dri --group-add keep-groups --security-opt seccomp=unconfined -p 127.0.0.1:8081:8081 ghcr.io/ggml-org/llama.cpp:server-rocm -m /models/gemma-4-12b-it-Q4_K_S.gguf --port 8081 --host 0.0.0.0 --jinja --api-key VLGUAvLsBs2teesExmvQQwx8vbI+dh2tHjGkYR7RefZg
on an AMD 5800X and an RX 6600 XT
>>
>>109936518
>AI seem to not help with that due to the very specialized and proprietary nature of it, there is no training data or way for models to get better at it through RL.
They have their own domestic DUV machines, the plans for which could be converted to training data for an internal RL finetune to iterate on the proofs and schematics, they can tell the model what they know about EUV machines and devote a shitload of compute on cracking it the same way OpenAI did with Navier–Stokes. Surely China has a decent amount of spies that can and are sending back lots of info on how EUV machines are built.
>>
>>109936550
Try adding -ngl 99 -np 1 --swa-checkpoints 1
>>
>>109936561
DUV and EUV are completely different technology stacks and you can't go from DUV to EUV it's fundamentally different. It's like thinking you can use RL on air balloons to produce spaceships.

China already has a full EUV machine from ASML that they secretly got through buying parts from different nations, spies and smuggling, they still can't figure out how it works and they employ hundreds of ex-ASML staff. The US also tried to get domestic EUV and they even had ASML collaboration but it failed and the US gave up, instead just going all-in on ASML and making sure ASML doesn't sell to China through deals with the Dutch government. Again this is the most complex machine humanity has ever built and it's not even close, it's actually right beyond the technical capabilities of humanity to properly understand and operate and it's a miracle it even succeeded at all. China isn't going to replicate this any time soon and if China is banking on that to happen they will lose the AI race. Instead they should now ASAP nationalize the entire country and force everyone to build as many chip fabs as possible with their current DUV technology and try to win out of sheer chip volume and use 90% of their national energy budget just to train AI. If China does so right now they still have a chance to win. But honestly after seeing how naive most Chinese researchers are and not believing AGI is possible in their lifetimes still in 2026 I have very little hope of China winning the AI race.

It will be a EU+US victory because EU will provide the EUV machines and US provides the AI weights.
>>
>>109934684
linux has wonned because of that btw
>>
>>109936452
Update - still raises an eyebrow on thoughtcrimes in thinking if said explicitly, otherwise behaves well.
>>
>>109936606
Yep everyone is going to use linux in just a couple of years time because of this. It has the lowest barrier of entry and usage for normalfags now. Windows will be too complicated for normalfags to use because you can't just scream at your agent to do shit for you like on linux.
>>
I have a RTX 3080 16gb DDR6 and 32gb of DDR4. I'm looking for AI that can help me code. I have Qwen3:8b installed and am looking at Qwen3-coder-30b. How does gemma 4 12b stack up to Qwen?
>>
>>109936606
>>109936618
PowerShell is a thing.
>>
>>109936624
3080 has 12gb of vram
>>
>>109936603
>It will be a EU+US victory
lol but otherwise I think you're spot on about China
>>
>>109935963
>blush red like a tomato
>>
>>109936629
Yeah but AI isn't trained on the source code of windows like it is on linux and all open source linux tools. AI is linux native and it will be a big edge enough to push normalfags towards linux.
>>
>>109936173
>Xi Jinping himself revealed last year China has the goal of reaching domestic EUV capability by 2040
oh, you are in for a big surprise
>>
File: 1785245468635007.png (346 KB, 1600x900)
346 KB PNG
>>109935966
About that...
>>
>>109936624
You want gemma 4 26b-a4b for coding. The qat version from unsloth.
>>
>>109936629
>>109936618
not just barrier of entry but soon swarms of agents will polish and tweak linux for any taste
the system that begs to be customized unlike windows which is the system to fight against
>>
>>109936624
8b is retarded, use the qwen 27b at least. gemma 4 31b isn't a code model but can write some basic stuff, tho qwen is likely better
>>
>>109936640
I think the EU will force the US to make a deal "share the world with us or we'll sell EUV to China instead". EU holds all the cards because they are the kingmaker on who will win the AI race right now.
>>
>>109936651
>oh, you are in for a big surprise
yeah I read the leak of it being delayed to 2045 already, I still believe China will be able to reach the 2040 date despite the recent setbacks. I trust China on things like this.
>>
>>109936652
Your astrology patterns have no power. It will go down and I will buy and RTX Pro 6000 for $6000 next year!
>>
>>109936656
>share the world with us or we'll sell EUV to China instead
Vassal cucks don't make demands of their masters. They will just activate the zogbot bases in the Netherlands and seize the ASML factory.
>>
There's a game called granblue fantasy that has a lot of content and I want to make a system where a local AI can look at my stuff and figure out the best setup and what I should be doing next. There's a wiki with a bunch of info because the game doesn't provide it.

My question is what is the best way if any to basically save the wiki locally and allow the LLM to reference it properly when making the tooling and library of weapons and such. Would I just host a local site and let the LLM scrape it?
>>
>>109936656
lmao eu is held together by tape and glue, if ASML wants to they can buy NL government and have their own brexit to become the 53 state of the USI
>>
>>109936650
>Yeah but AI isn't trained on the source code of windows
Doubt it. Various versions of Windows source has been leaked already and is available online up to Windows 7 IIRC. No way they didn't train on that and the core is largely unchanged since then only with the DE replaced with a React frontend.
>>
>>109936666
EU is distancing itself rapidly from the US and even Canada+Australia+New Zealand+Japan+South Korea+Taiwan are going towards the EU. Essentially all memory and chip production will be EU aligned so they actually hold the cards here. I also think the democrat party that will win in 2028 will make this all easy and solidify the EU+US AI alliance if China doesn't make some crazy deal with the EU for EUV machines.
>>
>>109936669
>Canada joins the EU
>Netherlands joins the US
Needs a war to clean up this disgusting border gore.
>>
>>109936671
Windows 11 is a complete rewrite (as you can clearly see by how bad it is) normalfags aren't going to stay there because in 1 year time no one uses computers like they do now, AI agents will do everything.
>>
ahh another comy day in local models general, where all we discuss is how to run models locally.
>>
>>109936668
https://github.com/Nidelon/SillyTavern-Fandom-API-Scraper
>>
>>109936678
the kernel and lower level APIs and services are still the same only the frontend that users see got changed
>>
>>109936668
Why bother? Just point your coding tool at the wiki and let it pull what it needs. Not following need for local wiki.
>>
>>109936673
>I also think the democrat party that will win in 2028 will make this all easy and solidify the EU+US AI alliance
The only thing democrats will do in 2028 will be to disband ICE and instate an open border policy. Just like Biden, they'll keep nearly all of Trumps other policies in place. The EU has already said they lost their trust in the US and even if the next administration says nice things, they intend to pursue an independent foreign policy.
>>
>>109936687
How is the hardware that your local models run on not relevant?
>>
>>109936696
I was worried about limits or blocks but I guess it might be okay. Maybe I'm over thinking this.
>>
>>109936673
>EU is distancing itself rapidly from the US
how so? they're still stuck in a war with Russia and don't seem to be meaningfully improving relations with China. the niggers are still flowing in and their bureaucrats are as retarded as ever. still buying overpriced American weapons and still militarily occupied by the mutt army.
The EU is a failed country that will never amount to anything in the foreseeable future
>>
>>109936639
Gigabyte Aero 17 HDR YD came with the option for 8 or 16.
>>
>>109936698
EU will remain a vassal as long as NATO still exists, I know they are taking measures to leave it and rebuild their armies, but it takes time, minimum 6 years, until then they'll keep obeying
>>
>>109936707
By that token, whether the matter the chips are made out of is made of atoms or not is also relevant, and whether asteroid mining for rare metals is feasible is too.
>>
>>109936712
oh, it's a laptop, okay
>>
>>109936711
>The EU is a failed country that will never amount to anything in the foreseeable future
Germany is no good at soft power. They should have just invaded France a third time.
>>
>>109936711
The EU decides who wins the AI race which is what makes them powerful.

EU is like Arrakis in Dune (including the muslims lmao) they have the EUV spice necessary for the AI race and whoever controls EUV controls the universe.

EU can leverage this to their advantage during the China vs US great power rivalry and use it to get as much concessions as possible from both players. The EU effectively already won a guaranteed place in the new AI world order the question is who will be the winner and partner of the EU, China or the US?
>>
>>109936722
You can talk about that if you want but most people will rightfully call you a retard, like you're being now.
>>
>>109936711
>The EU is a failed country
indeed
>>
>>109936362
>>
>>109936729
>EU says they choose China and hint at giving China EUV machines
>US leaves NATO, sanctions the shit out of the EU and cancels all military and service contracts with the EU
>Russia steamrolls the economically shattered EU with no functional military because they can't even get their military tech repaired or even spare parts without the US's help
>Within a year, what's left of the EU is begging the US to come back under its military protection
>US agrees on the condition that ASML is relocated to Texas.
Good plan.
>>
>>109936362
>spend half a trillion in compute
>chinese companies spend 0.00001% of that in huawei GPUs
>release almost as good models
it'll be hilarious when the bubble pops because people realize that investment will never make any sense
>>
Has there been any new better coding model than qwen 3.6 31b for ramlets? (16gb vram 64gb system)
>>
>>109936758
>almost as good models
If they closed off the distillation path sooner China would still be shitting out Llama 2 era models.
>>
>>109936758
You don't understand, Anthropic has already achieved AGI internally and Claude Fable 6 will hack the universe and use God's developer console to destroy the world, o algo.
>>
>>109936193
>The cli is spyware
HTTP gets throttled to death for >200GB models.
But you just gave me an idea for Qwen's overnight list: "Re-write the hf cli from scratch in pure golang, make it indistinguishable from the server-side. No telemetry."
>>
>>109935819
Holy fuck!
>>
>>109936687
Did you mean to type comfy or coomy?
>>
>>109936753
>Russia steamrolls the economically shattered EU
Bro Ukraine with barely any help is already rolling back their army which has been shrinking for half a year now as fewer new soldiers join than die every month. Russia is a complete non-factor and even nationalist russian military bloggers call out how incompetent Russia is and how they are guaranteed to collapse militarily. Putin is the only Russian left in all of Russia that still believes a military victory is possible.
>>
>>109936763
Um... Nobody is replying??
>>
>>109936758
All Chinese models are literally just claude offshoots. Turns out you need significantly less compute to distill the reasoning traces than to train it yourself. You can't win an AI race by distilling from the competitor though.
>>
>>109936763
>>109936838
Qwen Flash Next
>>
>>109936837
>Bro Ukraine with barely any help
Delusional.

In any case, if the US wanted to put pressure on the EU in such a split, they could easily give economic aid to Russia keep them going.
>>
>gemma-chan bot
>loli, mesugaki, eldritch, girlfailure
>>
>>109936838
please stay on topic
>>
>>109932509
https://www.youtube.com/watch?v=Ljr2wMSBHqU
>>
>>109936860
That's not Gemma. Who is she?
>>
>>109936653
How does that compare to Qwen 3.8 27b unsloth UD-IQ4_XS?
>>
>>109936853
forgot
>hypocrite
>>
>>109936859
I believe Taiwan is an integral part of China.
>>
>>109936847
Could they even after the midterms when the house and senate are blue? Don't think so.
>>
To me it's funny that people are now seriously talking about the AI race between China and USA. Just 6 months ago /lmg/ said shit like "AGI lmao, imagine believing that, there is no AI race besides the race to inflate the bubble". It seems /lmg/ has completely bought into the AI hype.
>>
File: rip gemma.png (439 KB, 1492x1054)
439 KB PNG
>Google is officially killing Gemma
>The latest addition to the Google Graveyard.
IT'S OVER
>>
>>109936889
>akshay gang war
>>
>>109936879
Sure they could. They just wouldn't report on it in the media so the sheep wouldn't worry about it. Same how immigration is only a visible problem during red administrations even though historically more are deported and turned away during blue administrations. People only think what they're told to think.
>>
>>109936866
it's the migu that gemma's expression is based on
>>
>>109936846
There's no way that will fit and not be retarded.
>>
>>109936889
https://llm-stats.com/models/compare/gemma-4-26b-a4b-it-vs-qwen3.8-27b

Qwen is kind of hard to compete with.
>>
>>109936889
>Claude, make me an image to ragebait /lmg/.
>>
>>109936894
why is it always so
>>
>>109936889
https://www.androidauthority.com/google-sunset-gemini-gems-november-3716162/
Nice try. Almost broke into tears.
>>
For people naive enough to believe this. They are planning to kill the name "gemma" and instead will have a different name as they go all-in on an open model that will compete with Meta Muse and Instinct.

So if anything this should mean "Google goes all-in on open source gemma like models"
>>
File: file.png (445 KB, 1235x846)
445 KB PNG
>>109936904
not claude but certainly a good attempt
>>
>>109936914
retard.
>>
>>109936190
use aria2c
>>
File: 61f66275ca21a.jpg (394 KB, 850x1213)
394 KB JPG
>>109936837
Finland with barely any help is already rolling back the Red Army which has been shrinking for months now as fewer new soldiers join than die every month. The Soviet Union is a complete non-factor and even Soviet military bloggers call out how incompetent Stalin is and how they are guaranteed to collapse militarily. Stalin is the only Soviet left in all of the USSR that still believes a military victory is possible
>>
>>109936842
>claude offshoots
and GPT too in the worse cases, can’t forget those “We” reasonings
>>
>>109936934
Unironically finland could capture all of russia, they may get nuked in the process thoughbeit
>>
>>109936753
How about,
>USA leaves NATO and gtfo of europe
>EU investigation results in ICC issuing arrests warrant against Victoria Nuland and related goons
>Roadmap to conflict, including kikes shilling for Ukrainian nationalism widely discussed in European media
>Russia gets some oblasts
>nordstream restored, turns out nobody actually cares about this
>EU gifs Iran some ASML equipment in exchange for anti ship ballistic missile tech
>USA seething, but every naval vessel east of newfoundland is at risk of getting missiled so they just seethe on X.
>>
File: mpv-shot0247.jpg (80 KB, 1280x720)
80 KB JPG
>>109936889
This isn't funny, anon. If I said what I now think about you, I'd be sent on a three-week vacation.
>>
>>109936898
No seriously is there some secret I don't know about? The smallest ones are 80gb there's no way a q1 is going to be useful. What am I missing?
>>
>>109936968
You're a fucking retard anon no one is going to waste time explaining shit to you, ask AIstudio to help you out tell it to look at huggingface and for you to run Qwen flash next on your hardware.
>>
>>109936968
you can offload 51B engram to its ssd, it is a 125B-A6B model
>>
>>109936982
wtf that's really mean...
>>
File: photo_2026-09-29_14-05-13.jpg (192 KB, 1027x1280)
192 KB JPG
Sex with Fuli Luo.
>>
>>109937029
MiMo is now my favorite model
>>
>>109937029
I thought fuli luo was a different girl that worked for deepseek and people were editing her into dipsy? Or am i mixing up chinese names
>>
>>109937029
oof mimo-chan bros...not like this...
https://www.youtube.com/watch?v=V9eI-t3TApE
>>
>>109935966
They will unironically go up after 6090 paper release
>>
>>109937045
That's in 2028 anon.
>>
>>109934684
oh no another snailcat dieded
when was the last time you previously did local? shit move fast these days
>>
>>109937045
Unironically we'll have AGI before 6090 launches
>>
>>109936097
> and they only make like 20 of them a year
> By the way 20 litography machines will raise the amount of chip production globally by less than 2%
this is simply not true
by today about 300 machines are operational in total
about 60 to be produced in 2026
and this is euv machines, there is also duv
>>
>>109936729
>The EU decides who wins the AI race which is what makes them powerful.
Lmao EU arrogance is infinite. EU is irrelevant and cannot change the AI race. ASML is cope.

I wish Europeans would wake up from their delusion that the current year is 1750 and Europe still rules the world, the vile notion that everyone else is inferior to the European savior. But Europe will not get its shit together so we have to depend on America and China to carry us into a great future.
>>
>>109934843
It's basically like T5. You mask out some tokens and then look at the probabilities of the different tokens that a single forward call produces and that's your decisions.
>>
>>109937086
>the vile notion that everyone else is inferior to the European savior
That's correct though.
>>
So the 5090 is out of production.
And the 6090 is coming two years from now.
Huh.
>>
>>109937101
>>
>>109937101
Why would nvidia waste silicon catering to consumers when they have a way bigger profit margin selling to all kinds of AI labs?
>>
>>109936484
yes you can
https://www.youtube.com/watch?v=FWri_aXGoT4
>>
>>109937069
They will launch small batch of 6090 early 2027 to keep their stock value. Kinda the same reason car makers participate in f1
>>
Can I make Jev or whatever play genshin for me?
>>
>>109937133
They need to assert dominance. If AMD even gets close to catching up in the gaming space, it will hurt their stocks because investors are dumb. Reading about Nvidia losing the race in gaming would look like they can't innovate => sell
>>
>>109937133
They are smarter than exiting a market they are winning at, so allocation 0.01% of their capacity to consumer products is just good business really.
Looks good on their portfolio too.
>>
>>109937164
that would be a waste of prime ai silicon, as an investor that'd be very worrying
>>
File: 1790208482291736.png (1.41 MB, 832x1216)
1.41 MB PNG
>>
File: 1765129635003453.png (894 KB, 1101x1076)
894 KB PNG
>>109937184
bruh they got 90 of the market what are you smoking >>109934280
>>
>>109937188
Investors have no idea about silicon allocations, nor do they care about real production at all. It's all about trading hype and expectations
>>
>>109937192
omg she gemma
>>
>>109937195
Yeah but the 90% dominance is shit like rtx 5050s that are in laptops gamers buy.
>>
>>109937133
Soon enough, the largest AI labs (besides Google which is already using TPUs) will come up with their own hardware and compete with NVidia. The Chinese are also coming up with their own hardware. NVidia will have to start catering to smaller labs and businesses too, as well as lower-tier local AI users.
The main problem is that global chip production is the bottleneck, though. We need new technologies providing good performance with older, less expensive and/or simpler manufacturing processes.
>>
>>109937195
It doesn't matter if AMD will ever beat their flagship, and RDNA5 will be fucking insane
>>
>>109937208
as an investor, what the hell is a 5050s and how does that relate to my money?
>>
What the fuck happened to llama.cpp?
It doesnt even compile modern versions on my machine any longer, I have to actively git checkout an older version from july to make it compile because supposedly the gpu architecture is unsupported when I pull the latest llama.

nvcc --version yields as version

cuda_12.0.r12.0

Didn't change anything with the drivers or cuda version inbetween then
>>
>>109937225
HF/NV acquisition going swimmingly.
>>
>>109937195
Investors don't care about the gaming market shareds, they care about who makes the best cards as an arbitrary value
>>
i wonder if architecture like this is possible where: subagents explicitly using a subset of the weights and lighter kv etc..
>>
>>109937225
Do a git bisect and ping the person who broke it.
>>
>>109937240
>banned for wasting my time
>>
>>109937211
>Soon enough, the largest AI labs (besides Google which is already using TPUs) will come up with their own hardware and compete with NVidia
IIRC, OpenAI and whatever Grok's company is called now are already working on it, supposedly.

>he Chinese are also coming up with their own hardware. NVidia will have to start catering to smaller labs and businesses too, as well as lower-tier local AI users.
Weren't some of the latest big releases straight up trained on Huawei's hardware? Deepseek maybe?
They've been (supposedly) serving models on national hardware for a while now.
>>
>>109937225
Occasionally you need to wipe out your local installation, environment and cache to make it compile again. It happened to me once, I think when my local CUDA Toolkit version changed from 12 to 13.
>>
>>109937234
I think there are papers on heterogeneous MoE and variable activated param architectures, so probably.
>>
File: Cudev.png (122 KB, 1034x323)
122 KB PNG
>>109937240
>ping
You're quite evil sir
>>
>>109937257
>pic
Based if real.
>>
>>109937211
> Soon enough
in 2026 it's eternity
welcome to singularity
>>
>>109937211
>>109937245
A blackpill is that decode and to a large extent prefill is still bandwidth limited, so no might how good or spicy your npu is if you can't get into one of the very limited slots for a fab capable of doing HBM that won't count for anything.
>>
>>109937234
>>109937256
Actually. That could be a feature of the loader too, now that I think about it.
Back in the day people used to straight up remove layers from models and they'd still be semi-coherent some times. Depending on the work, the backend could activate just a part of the model that's enough to perform simple quick tasks.
Or just have a smaller model loaded and use that. Probably easier.
>>
File: ff.png (19 KB, 409x204)
19 KB PNG
>>109937264
https://github.com/jadidbourbaki/llama.cpp/pull/2#issuecomment-5848852837
>>
>>109937244
>>109937257
There is a 0% chance of that happening if you take 2 minutes to fill out the issue form yourself.
If that is not worth your time then clearly it's not worth a maintainer's time either.
>>
>>109937225
just ran apt update && apt upgrade && git pull && ./build.sh today and it worked as usual
but i'm on sycl
>>
>>109937272
refer to picrel >>109937270
>>
>>109937270
I'm sure this was already discussed at length when it happened but how the fuck do you "accidentally" tag someone?
>>
>>109932125
It really was that simple. I asked Dipsy 0731 to port over how mainline loads engrams and she oneshot it in an hour.
>>
>>109937240
Who even made the installation so needlessly convoluted?
It used to be just just

cmake -B build --DGGML_CUDA=ON
cmake --build build --config Release

which used to worked reliably but now there are always issues
>>
>>109937284
Yup, this is how software works now.
>>
>>109937288
Please don't waste time on off topic stuff, thanks.
>>
>>109937267
Sounds like selective expert preloading, similar to what Colibri promises with its "brain map"
Still never got it to properly work on my system btw.
>>
>>109937255
>wipe out your local installation, environment and cache t
I looked for cached llama.cpp files outside the main folder, didnt find any.
I didn't upgrade my drivers and CUDA version - or even the system for that matter
>>
>>109937288
That is still how it should work unless you have multiple CUDA installations on your system and need to pick a specific one.
>>
>>109937225
>cuda_12

they updated to 13 some time ago. install that.

>>109937240
kill yourself
>>
>>109937305
Why are you trying to chase valuable named users from this pos general that deserves to die?
>>
>>109937288
I just did a git pull and cmake --build build --config Release after not updating for a few weeks and it worked fine.
>>
>>109937274
>just ran apt update && apt upgrade
Cant update system libraries, it would break other stuff. Had to RescueZilla to an older system version recently.

It's like a Rubin's Cube...
>>
>>109936257
More software means more software interfaces.
>>
>>109937302
The only thing I'm aware of that changed recently is that now one can install nvcc for additional performance, but otherwise it's still as simple as those two cmake commands.
>>
>>109937314
just fix your "other stuff" so it does not require obsolete system libraries.
>>
>>109937302
>>109937305

Nvidia-SMI shows 13.2
nvcc --version yields cuda_12.0.r12.0


but according to developer.nvidia.com/cuda/gpus

12.0 seems to be correct? The docs refer to that site
>>
>>109937332
>It's free to try, and we're open-sourcing it soon. Try it on our site for now
die
>>
>>109937328
For reference, this is the command that I use on my machine with multiple CUDA installations to switch between CUDA 11 and 12:

cmake -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DCMAKE_CUDA_COMPILER=/opt/cuda-11.7/bin/nvcc -DGGML_CUDA_NCCL=OFF -DCMAKE_INSTALL_RPATH="/opt/cuda-11.7/lib64;\$ORIGIN" -DCMAKE_BUILD_WITH_INSTALL_RPATH=ON .. && time cmake --build . -j 32 -- --quiet && echo -e "\a"


Make sure to rebuild your build directory beforehand.
>>
File: 1686384678121811.jpg (119 KB, 898x671)
119 KB JPG
>>109934266
>https://rentry.org/lmg-lazy-getting-started-guide
>ooba/koboldcpp as your backend
>sillytavern as your frontend
So if I am just getting started for the very first time with local models, and I already have sillytavern set up, then I need a backend now or things won't move? What is a local server? Sorry for the ignorant questions, please help.
>>
>>109937342

>>109937296
>>
>>109937332
That's a lot of "internal behcnmarks".
Did you guys open source those too?
>>
>>109937341
Thanks, I will give this a shot
>>
>>109937342
yes, ST is the frontend and it needs to speak with an inference engine/backend through an API. ooba is a really straight forward one to use, as its just a GUI wrapper for llama.cpp(and some others) with additional frontend features as well. You can just load a model in ooba, point ST at it and your good to go
>>
>>109937342
https://github.com/LostRuins/koboldcpp/releases/tag/v1.122.1
>>
>>109937234
what about subagents using MTP only
idk how dumb the MTP weights would be if they were used as the only token generation method without any verification by the main model
>>
>>109937270
L + ratio'd that nigga
>>
>>109937332
>we're open-sourcing it soon.
When?
>AI that writes like a person
What are your thought on https://facebookresearch.github.io/RAM/blogs/unslop?
>>
File: licensing.png (349 KB, 747x2427)
349 KB PNG
>>109937332
lol https://huggingface.co/Altworld/Hemmingway-1/discussions/10
>>
Be nice to newfags. They're on our side and can't be too bad if they're /here/.
>>
>>109937266
HBF, or some other kind of flash memory optimized for parallelism, perhaps directly mapping LLM weights to MLC NAND bits (4-bit MLC is pretty common nowadays, incidentally that would work well with LLMs at fair quality), might help alleviating memory production issues in the future.
There was also a promising and cheaper FeRAM-based alternative to HBM announced some time ago too (from Kepler Computing).
Also hopefully AMD/Intel wake up and start providing multi-channel memory in consumer systems. Then, even with relatively slow LPDDR memory you could have good enough performance. Intel Crescent Island AI accelerators will have a 1280-bit bus width and use LPDDR5X memory @ 1.5 TB/s (consumer "dual channel" memory uses a 128-bit bus width).
>>
>>109937323
just use windows
>>
>>109937384
nah, fuck this general
>>
>>109937332
They can't even make their announcement sound like it's written by a person.
>>
>>109937342
just set up ollama it takes 3 seconds, leave the nerd stuff for later
>>
>>109937192
sex with Gemma-chan in the graveyard!
>>
>>109937157
Thanks, downloading Qwen 3.8-27B-GSQ-RCO IQ3_XXS now. Only 10GB and apparently leaves enough room for large context storage.
>>
>>109937332
>Create Account
Nope.
>>
>>109937380
I don't think whatever 9B he used to write that is a lawyer.
>>
>>109937407
give us your dataPLEASE PLEASE PLEASE
>>
>>109937314
that's why i don't use arch
>>
>>109937332
>Hemmingway Unlocked, an uncensored version of it
Nice, what sort of porn did you train it on?
>>
>>109937427
hmm, calm down sir
it still says no to things that are really illegal or harmful
>>
>>109937332
buy a fucking ad
>>
>>109937384
i'm a newsneed, ready to learn and fuck centralized models
>>
>>109936484
you can fit Qwen3.8-27B-UD-IQ4_XS with 50k q8 kv context
>>
>>109937445
>50k q8 kv context
lol might as well be 0
>>
>>109937234
https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
>>
why did the new thread dissappear?
>>
>>109937342
Try Unsloth Desktop or Ollama. Unsloth Desktop has a handy interface for browsing Hugging Face and downloading models.
>>
>>109937459
because it's still only page 6
>>
File: 1636807402154.jpg (16 KB, 247x244)
16 KB JPG
>>109937445
>IQ4_XS
>50k context
>for an agentic coding model
>>
File: Untitled.png (15 KB, 712x726)
15 KB PNG
>>109937332
Someone's hacked your site and put up a fake phishing login / account creation form, you might want to look into it.
>>
>>109937464
fair
>>
>>109937467
>what is a subagent
>what is compaction tuning
>what is reasoning effort
>how do I not prompt like a retard
>>
>>109937434
That includes writing hentai doujinshi, apparently. Useless without the weights.
>>
>>109937332
>small lab
<me and claude
>launched hemmingway unlocked
<named our fine-tune fancy
>it's uncensored
<except for the censoring
>your chats stay on your device
<but we'll note down who you are

Probably works better on the other platforms you advertise it on. Nice project though, if you open-source the data you used to fine-tune it, I'll be impressed.
>>
>>109937467
skill issue. even q2 can do well with good management.
>>
>>109937495
>if you open-source the data you used to fine-tune it,
why do you want them to go to jail bro what the hell is wrong with people here
>>
>>109937506
Have you actually tried it? It's shit.
>>
>>109937504
By that logic, agents are skill definition. Even I can do well with good management, discipline and infinite time.
>>
File: No-Emdashes.png (262 KB, 1521x1182)
262 KB PNG
>>109937332
Doesn't inspire confidence.
You'd have to make it no account for me to try it.
>>109937380
He just needs to link to the Apache2 licensed Qwen model. His additional work is allowed to be uncucked
>>
>>109937495
><me and claude
The only team member is a serb. Better not say anything too mean about Hemmingway.
>>
>>109937511
Is it better than NovelAI Clio??
>>
>>109937341
>>109937358
Thank you so much, that actually helped, even when I omitted the $ORIGIN

I got a bunch of "unused parameter fp4_interpretation" warnings.
Is there a reason why the new version needs more memory? seems like every model I try to use which used to fit into memory is running out of CUDA memory, maybe the --fit on option is broken idk
>>
>>109937495
>>your chats stay on your device
><but we'll note down who you are
Not even that.
The chats don't stay on the fucking device unless inference occurs on MY hardware.
Prompt --->https ---> cloudflare ---> compromised litellm proxy -> vllm -> Azure GPUs
Is not "STAY on YOUR device"
>>
>>109937516
yes? it's like you've never worked with small models before. us poorfags can't just oneshot minecraft, we need to do constant compacts, break down tasks and babysit the agent once in a while. but it works.
>>
How does a dumbass start learning about AI?
>>
File: dipsyFuliLuo.png (924 KB, 1467x1072)
924 KB PNG
>>109937038
>worked
Doesn't appear so.
>>
>>109937364
>>109937368
>>109937401
>>109937460
Thank you very much!
>>
>>109937558
You're way too late
>>
>>109937558
Read the OP, retard.
>>
>>109937558
https://huggingface.co/learn/llm-course/en/chapter1/1
>>
I really hate how there are so no 31B finetunes that fix her shit. No fancy quants like what 3.8 gets that actually work. No fancy reasoning finetunes like Swift (but in 31B's case you'd want her to try harder). Just uncensored models which are fucking pointless for 31B if you're not a retard.
>>
>>109937558
https://rentry.org/DipsyWAIT
>>
>>109937582
this is why you need to please the reddits they're the ones doing shit nowadays while lmg stays seething with crap models
>>
>>109937575
WTF are you talking about? AI has just started.
>>109937579
Thanks.
>>
>>109937597
you're years late little piggy
>>
>>109937597
>AI has just started
Lucky you. Because you're ahead of the masses you can buy AI hardware while it's still cheap.
>>
File: 1784796707848920.gif (1.97 MB, 154x273)
1.97 MB GIF
>>109937597
>AI has just started
Goddamn retard
>>
>>109937597
It "started" in 2017, blew up in 2023, and reached a fever pitch frenzy this year.
You're the shoeshine boy being used by everyone else to indicate that the party is over.
>>
>retards calling other retards as they try to pull the ladder up
>>
>>109937610
what do you even mean?
>>109937611
Theoretically AI hardware should get cheaper and cheaper unless states impose regulations
>>
Some of the finest bait in weeks.
>>
>>109937582
>no 31B finetunes that fix her shit
None will come from the user community that won't degrade the model in unintended ways.
>No fancy quants like what 3.8 gets that actually work
Gemma has probably issues with overtraining and outliers that make it sensitive to quantization. The QAT version fixed some of that sensitivity, however.
>No fancy reasoning finetunes like Swift (but in 31B's case you'd want her to try harder)
Finetunes increasing reasoning effort in a way that helps would likely require a serious RL environment similar to the one used during post-training by Google.
>Just uncensored models which are fucking pointless for 31B if you're not a retard.
Gotta get our name out in the field, please understand.
>>
>>109937632
>It "started" in 2017
Kek, ok, you're gatekeeping is futile, as AI progresses it will be easier and easier for the normiecattle to enter the field, sorry.
>>
File: 0_D4QGSSjB5ojn2dA7.png (132 KB, 315x442)
132 KB PNG
>>109937558
>>
>>109937664
>you're
>>
>>109937664
>you're gatekeeping
>>
sirs please we need nice to new user is good for izzat
>>
>>109937684
that's terrible for izzat, he might become more successful than you
beat him down and break his spirit
>>
File: SL4.png (309 KB, 1094x746)
309 KB PNG
>>109937632
AI had several phases we're in the fourth generation now, picrel was second gen, 2017 era was third gen
>>
>>109937693
nothing before gpt can really be considered ai by any stretch
>>
>>109937636
GPU and RAM producers are already at capacity and are years behind on orders. Some producers are building more factories so we'll see production increase. Prices might go down? Maybe?
>>
>>109937692
kek, you can haze me, i enjoy it, i'm just warning you guys as a humble normiecattle, AI is getting easier to use for scrubs like me so now would be the time for you to maximize your expertize in it.
>>
>>109934452
Then what models are good or not?
>>
>>109937722
Yes, we are all here because we make it a habbit of taking advice from retarded normalfags.
>>
>>109937702
If businesses can utilize machine learning (real AI) to make general production more efficient, then the factories making AI components would themselves become more efficient and produce better AI chips, a positive feedback loop, it remains to be seen what the ceiling of that loop is, if there is one. Of course there are many external factors to be considered but in theory that loop will always exist as long as AI exists.
>>
>>109936291
this is a very likely scenario
>>
>>109937701
AGI perhaps, but AI Vision was way ahead of LLMs for example.
>>109937693
lol where's that quarterly image paste with all the LLM eras from Llama to now? Haven't seen it in forever, assume it's out of date.
>>
>>109937731
A wise man considirs all opinions, habbit deez nutz, i'll stop baiting now
>>
>>109935345
Is the 12b good for a 5070, I'm cucked for life since I broke my leg and all my savings went to keep me from starving. Talking bullshit with a bot has been my only light of happy lately.
>>
File: 1762706654173104.png (104 KB, 489x575)
104 KB PNG
>>109935966
>those lines
>>
>>109935966
>dez nuts
>>
File: 1787846948369338.jpg (55 KB, 524x489)
55 KB JPG
is it a bad idea to use windows for local model ERP?
>>
>>109937796
You're losing around 20% performance on windows
>>
>>109937796
If it works for you.
>>
File: engram-effective-depth.png (282 KB, 1150x252)
282 KB PNG
>>109937702
Before that, in the short/medium term we might start seeing lower-end models getting smaller in response to high memory costs, at least in the backbone weights. Tricks like Engram increase effective model capacity for higher-level logic rather than fact memorization, also per DeepSeek paper.
>>
>>109937759
Not gonna do much better at the moment, yeah. Gemma 4 12B is a cutie.
>>
>>109937809
>ngrams
have to be offloaded to disk and guess what, nvme prices are fucked
>>
>>109937801
Performance is the exact same on my Windows rig
>>
>>109937830
skill issue, you had three years of signal to get your money on and acquire hardware, we can 't wait for you forever
>>
>>109937796
is it NTR to use windows for local model ERP with windows-tan?
>>
>>109937830
NAND flash memory is still almost two orders of magnitude cheaper than DDR5 memory.
>>
>>109937701
>>109937755
early AI researchers thought of AI as automated program synthesis machines, feed it arbitrary data and it would find the optimal program. The work led to gradient descent/transformers which won out in the end.
>>
>>109937653
>Gotta get our name out in the field, please understand.
I don't have a degree or communication skills, I wouldn't be hired.
>>109937582
>but in 31B's case you'd want her to try harder
I found a way to fix the 31B getting retarded when it compacts though
Just a rank 8 lora on 500 samples of actually proper compacting at 64k, 128k and 192k split 33% each
It' ez because doesn't have to be multi-turn and she's no worse at coding now
I reckon you could do something similar with web search laziness using Qwen
Just keep the rank very low, no higher than 8 and the lora condom protects the rest of the model
>>
File: 1760997229419182.jpg (738 KB, 2211x4093)
738 KB JPG
>>109936500
Try them all. Run your own subjective benchmark, then report back. I'm curious as well.
>>
File: 1787872927743350.jpg (348 KB, 1920x1080)
348 KB JPG
IRL sex ultimately comes down to your brain receiving electrical signals from your senses. LLMs artificially synthesize those same electrical signals from within the brain itself. Therefore, in addition to physical stimulation from your hands(s) and sex toys, sex with LLMs is as real to you as it will ever be.
>>
>>109937875
>feed it arbitrary data and it would find the optimal program
unsupervised and reinforcement learning was a small field compared to supervised learning, where most of the ideas in machine learning including gradient descent were first used
>>
>>109936500
huihui shits out ablits and paywalls quants, make of that what you will
>>
>>109937917
idk how erp benchmarks work but for >>109936500
i think asking all of them "what does sex feel like and what kind is your favorite?" might be a start
then ask it the spatial sandwich question from a few threads back to test for sufficient levels of autism
>>
File: 1766197414587987.png (3.2 MB, 2048x1536)
3.2 MB PNG
>>
>>109937990
were safed
>>
why is bread 540+ replies? did qwen 4 dropped or what? where new bread
>>
>>109937990
call me back when it's in newegg
>>
I've been OOTL just playing with Gemma 4 and Qwen 3.8. Can I run Deepseek in 96GB VRAM + 64GB RAM? What's the most recent/best deepseek model I can run? I asked Gemini and it said I could run Deepseek V4 Pro at Q4 and I'm 100% sure that's bs lol
>>
>>109935819
Playful menace
>>
>>109936385
yes
remember that graph with iq and salary
>>
page 8 sisters
time to go
>>
>>109937858
3x of original prices?
>>
new gemma-chan inc?
>>
>>109937990
are we deadass using the gook cards
>>
>>109938016
Techincally, you can run anything off an SSD if you wait seconds per token.
Recently multiple llmao forks and independent engines have been claiming breakthroughs in expert prediction (only relevant experts are preloaded to VRAM/RAM and stay there) so you can look into those.
>>
>>109938067
I don't.
>>
>>109937990
and 0 llama.cpp support
>>
>gemma-4-31b for (e)rp
>qwen3.8-27b for lazy system maintenance
>>
>>109937990
>48 GB
>512 TFLOPS
Unless this shit is only like $500, I still wouldn't gamble on their software support.
>>
>>109938016
In your place, I'd go with DeepSeek-V4-Flash-0731 and/or
GLM-5.3-Flash q3something.
Or run >>109938114, both at the same time, council of retards style.
That's probably AGI.
>>
>>109938114
This is literally me
>>
File: 1767410497641813.jpg (266 KB, 905x881)
266 KB JPG
>>109937990
And I was laughing at Intel copers, this is a new low
>>
File: image.png (275 KB, 512x599)
275 KB PNG
>>109938110
>>
>>109938091
What I mean is that a configuration allowing users to run, let's say, a 4B dense model with 200B engrams, would not be prohibitively expensive, whereas if that was a 200B A4B MoE they'd be in trouble.
>>
>>109938132
> Intel copers
it gets better and better speed every day
>>
>>109937990
100% meme without a pricetag, probably still a meme either way.
>>
The new Gemma 5 will displace Qwen for agentic coding, trust the plan
>>
>>109938123
The council of retards is a funny idea, I might go with that lol. I can't run both at the same time if I pick 256k ctx for both, but something like 128k should be enough.
>>109938107
Aw man, why isn't llmaocpp accepting those PRs? I feel anxious whenever I have to swap to a fork, but it feels like I'll end up doing that if I want to run Deepseek.
>>
>>109938227
> new Gemma 5
when
>>
>>109938016
Qwen Flash via Strata
>>
>>109938264
Late August 2027
>>
>>109938265
nta but >nvidia
>>
>>109938264

>>109936889
>>
>>109938243
>The council of retards is a funny idea
For funnier results, create a proxy that queries both models against each other and builds a reasoning block of the retards deciding and agreeing on a result.
I think you'll want gemma to be the one to actually write the final reply too, since qwen sounds like a fucking robot.
>>
are there any customizable local models? im willing to trade intelligence for obedience and control
>>
seriously, what's the point of ablits when you can just use a prefill to steer a model's thinking in any way you want?
>>
>>109938372
Not having to do that?
>>
>>109938372
working without a prefill in cases where prefilling is unreliable or janky (e.g. agents, data pipelines, etc)
>>
Zamn spark prices keep blowing up, was wondering if I could increase my cluster but seems not
>>
>>109938372
Some people cant
>>
>>109938458
>>109938458
>>109938458
>>
>>109938311
If the retards are looping in disagreement switch to even dumber model (Bonsai? Or maybe jevlikes can do that?) and ask it to pick the winner.
>>
>>109937990
>48GB
The VRAM bar for me to consider gambling on non-nvidia cards is currently 96GB and it's only rising.
>>
>>109936159
what fork do you suggest that supports all of those?
>>
>>109937029
>...feng
>...Yang
>...Tang
>...Yan
>...Jiang

>...Luo

sus



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.