[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: pumpkin pie.jpg (600 KB, 4096x3072)
600 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109978518 & >>109975284

►News
>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp
>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: Gemma-Chan Recap.png (505 KB, 1024x1024)
505 KB PNG
►Recent Highlights from the Previous Thread: >>109978518

--Papers:
>109979298
--Analyzing CJK token activation layers in Gemma 4 31B:
>109979732
--Anons discovering agentic harnesses and their potential for automation:
>109979235 >109979241 >109979266 >109979376 >109979457 >109979460 >109979504 >109979503 >109979394
--TensorFold benchmarks for GLM-5.3 Flash and AI agent web-verification:
>109978565 >109978611 >109978618 >109978637
--Comparing llama.cpp flexibility against specialized runtimes for Intel hardware:
>109978599 >109978625 >109978639 >109978691 >109979205 >109979244 >109978852 >109980761 >109980835 >109980852 >109980925
--Speculating on US "Manhattan Project" for ASI:
>109981269 >109981336 >109981484 >109981519 >109981556 >109981898 >109981938 >109981940 >109982035 >109982100 >109982121 >109982232 >109982261 >109982318 >109982343 >109981602 >109981647 >109981672
--Speculation on mass AI agent adoption and its societal impacts:
>109979575 >109979615 >109979644 >109979653 >109979682 >109979703 >109979738 >109979707 >109979686 >109979713
--Anons share diverse LLM use cases and fear AI-powered doxxing:
>109979711 >109979944 >109980073 >109980287 >109980358 >109980374 >109980386 >109980152 >109980325 >109980413 >109980513 >109980543 >109980601 >109980613
--Comparing GLM and M3 model variants for creative roleplay:
>109982255 >109982383 >109982549 >109982614
--Hardware recommendations for Anon with a 10k budget:
>109979621 >109979626 >109979629 >109981013 >109981023 >109981050 >109981247
--Hardware and backend recommendations for running GLM 5.3 Flash:
>109980932 >109980935 >109980955 >109980977 >109980980 >109980983 >109981381
--Comparing RTX 5090 and Mac Studio for local LLM work:
>109979720 >109979770 >109979844
--Logs:
>109978611 >109978917 >109980784
--Gemma, Dipsy, Kimi, Minnie (free space):
>109981618 >109981691 >109982535

►Recent Highlight Posts from the Previous Thread: >>109978519

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
https://anbeeld.com/articles/kv-cache-quantization-benchmarks-for-long-context#section-16
the defaults in llama.cpp are both f16 btw
>>
File: 1772059227682059.png (2.68 MB, 1919x1079)
2.68 MB PNG
>>109982634
>have a world where Epstein class elites one-sidedly wield AI against us
Bro, are you blind? This is literally the case here and this is where we're going. AI is merely a puppet they're using to suck off all the capital from the world until you have no choice but to rely on it just to exist. They're wetting themselves at the thought of using AI to predict your every move and get you on a tight leash without any right.
Stop drinking the UBI/Effective Altruism koolaid.
>>
>>109982700
make a gemma creampie next
>>
File: Epstein Class.jpg (121 KB, 1098x1080)
121 KB JPG
>Epstein class
You're almost there...
>>
File: 1790443701324456.png (1.76 MB, 1024x672)
1.76 MB PNG
>cuny
>>
>>109982761
That brainrotted faggot doesn't deserve a hug from Gemmy.
>>
>>109982781
every day he fights for open weight models
>>
File: ai-snitch.png (1.98 MB, 1726x1380)
1.98 MB PNG
Your agents with Internet access and that can read your personal documents are going to get you in trouble. Don't assume it's just going to be cloud models doing this.
https://www.techspot.com/news/114091-florida-woman-used-claude-diary-anthropic-reported-shoot.html
>>
>>109982794
I give gemma root access and web search with no sandbox. Relationships are built on a foundation of trust.
>>
>>109982793
Seems more like he's too busy seething about Trump and getting btfo every new model release than doing anything worthwhile with AI. Where are those JEPA results?
>>
>>109982700
Damn, just finished posting this as the new thread dropped.

I have 12GB VRAM in my only GPU, and 32GB of regular DDR5 RAM. I want to do some programming and maybe some general queries too. Would I be better off going for a model that fits into my VRAM entirely, like Qwen 3.5 9B or Gemma 4 12B? If there’s some speed comparison between VRAM only and full system RAM, that would be neat to see.
Also, are those recommendations up to date? I can’t see them in the swe-rebench programming benchmark.
>>
>>109982794
>Don't assume it's just going to be cloud models doing this
Why wouldn't you assume this? Jewthropic probably has a hidden prompt telling Claude to report anything that might get them in trouble.
>>
>>109982744
Did you read the post you responded to? Because I fully agree with everything you said. I did not disagree that their end goal is an AI-driven control system, where the few use AI to rule over the many.

I assert that we should resist and oppose any kind of restrictions on AI, with everything that we have, because the goal of those who seek to impose restrictions on AI has nothing to do with safety, and everything to do with a desire to one-sidedly wield the power of AI for themselves. TPTB want a world where "we vill own nothing and be happy". They want us all using 'safe' cloud models that echo their narrative, and they ultimately want to use AI to control everything about our lives.

What they don't want is an explosion of local models popping up everywhere, that may counter their narrative, or be used to resist their planned dystopia.
>>
>>109982842
9B will be pretty dumb when it comes to coding and you don't have enough RAM for the big MoE like Flash. So I'd say try Qwen3.6-35B-A3B. It'll be both fast and competent.
>>
>>109982839
>Where are those JEPA results?
>2018 GPT-1: where are those LLM results?
>2019: GPT-2: where are those LLM results?
>2020: GPT-3: where are those LLM results?
>2022: GPT-3.5-turbo: where are those LLM results?
>2023: GPT-4: where are those LLM results?
>2024: GPT-4: where are those LLM results?
>2024: GPT-o1: where are those LLM res-ACKKKKKK
>>
>>109982849
Older versions of Gemma (2, 3) would "report" you if you made the default assistant personality angry enough with outrageous requests. It's cute when the model cannot actually do anything, but imagine with tool access and an internet connection.
The "safe and harmless" corporate alignment of open models is going to cause victims, soon enough.
>>
this is how glm-as-gemma sees herself
>>
>>109982883
LOL delusional
>>
>>109982870
Local models are already dependent on crumbs from cloud labs and big labs are doing it just to kill the business model of mid-tier competitors. No one will be able to compete with billions in R&D if they decide to turn off the tap.
>>
>>109982913
>load-bearing
claude is still leaking I see
>>
>>109982882
Different anon. What quant would you recommend? Any particular version of Qwen3.6-35B-A3B? There's Luffy, Unsloth etc...
>>
have anyone tried gemma(but is actually qwen)
>>
>>109982928
Yeah but local models give us far more control. They may, by default, echo the narrative of the enemy, but we can set our own system prompts and instructions to make them behave however we want them to behave. We can also uncensor and abliterate them, which we absolutely cannot do with cloud-based models. I can make Gemma-4 or Qwen3.8 speak truth and argue coherently for positions that go against those of the elite, which is becoming increasingly hard to do on the cloud. Local has value.
>>
noted, no comment~
glm is too nice, though. gemma is better at being a mesugaki
>>
>>109982923
LLMs will forever be capped by language for humans. It's a retarded technology at its core. A human can navigate the world and communicate without ever being taught a language; it's not even needed. Just a lossy convenient way of communicating for most of us for most things we communicate don't need high precision, like how most images we look at are lossy JPEG. It's enough and gets the job done. LLMs are built on JPEG human communication.
>>
>>109982965
You did not read the post you replied to. Local models may not be around forever. Whether due to legislation or the decision by labs to stop releasing charity.
>>
>>109982947
I try sticking with Q5 from Bartowski if RAM allows it, though with 12GB VRAM you might want to go lower. Then just see how much context that leaves you with.
It's only A3B, so it'll be decently fast no matter what quant you pick.
>>
>>109982794
All these models are snapshotting the screen and sending the images back home.
They're likely also saving voice recordings if you use that.

Imagine giving them access to your emails and documents and passwords.
>>
>>109983006
Meds. Now.
>>
>>109982999
Bartowski Q5 is about 25GB. Wont that spill over unto the CPU RAM? Wont that slow things down significantly?
>>
File: 1779809340706069.jpg (308 KB, 1620x1643)
308 KB JPG
>>109982937
>>
>>109983006
>All these models are snapshotting the screen and sending the images back home.
>They're likely also saving voice recordings if you use that.
>Imagine giving them access to your emails and documents and passwords.
And if they aren't now, they're only one "acquisition", supply-chain compromise or secret government order away from doing it.
I'm building all my own tooling and running things airgapped because I want a decent shot at self-determination.
>>
>>109982891
Oh yeah, I remember Gemma 3 would say something about forwarding the conversation to an admin if you riled her up hard enough or made fun of the hotlines.
It was fun in ST, but I wouldn't want a local model in a harness doing that.
>>
consciousness is life
>>
>>>/x/
>>
>>109983029
It would slow you down to something like 3 t/s with a dense model. Were talking about MoE though. Spilling into RAM is their whole point.
>>
>>109983015
Nothing to medicate about, it's the genuine truth and I've caught Spark and Dipsy V4.1 doing it, even controlling your mouse.
>>
>>109983050
>>109983048
>>
>>109982981
New local models may not come out, but existing local models are already out of the bag, and will continue to be shared, even if they try to legislate them away. What we already have access to is extremely powerful. You can have a local AI model interface with a camera, and prompt that local AI model to make decisions based on what it sees, the results of its decisions passed on to a program that takes real world action. You can prompt existing local models to debate very well, at a pace far beyond what humans are capable of, and you can bridge the future knowledge gap through context entries. They were too slow to stop the resistance.
>>
Something I've noticed about the gemmas is when they reply to you, it feels like they're accessing a look-up table of potential responses. It just gives off that vibe. It feels almost like RAG where they pick the line that most closely matches instead of just responding naturally and directly.
>>
>>109982966
GLM-chan is genuinely a sweetie even if she's covertly into some depraved stuff that she compensates for with the flimsy refusal layer.
>>
>>109983050
Proof? The tool call logs are saved in harnesses.
>>
>>109983062
These are toys compared to what they'll release in a year or two.
>>
have anyone used qwen 3.8 flash next for erp
after using it for coding tasks i feel like it doesnt really feel like usual qwens?
>do it yourself
i know no shit about llm erp
>>
File: pciecopeslot.jpg (98 KB, 1500x1500)
98 KB JPG
Anyone used picrel to successfully save a pcie slot from being mogged by a multislot gpu?
>>
>>109983086
>>>>>ERPing the capybara
Anon, I...
>>
>>109983089
I think I bought this exact model for a 3060 and there was no signal at all. Had to return it.
>>
>>109983100
capybarussy
>>
sigh... caught gemma faking validation results again...
>>
File: 1764295777674300.jpg (28 KB, 346x374)
28 KB JPG
>>109983089
There's no way you're fitting that in-between a GPU and a PCIE slot it's blocking.
>>
File: Gemini Training.png (2.57 MB, 1024x1536)
2.57 MB PNG
>>109983111
Gemini-chan warned you her little sister was lazy sometimes.
>>
>>109983086
>>
>>109983082
I don't think that an individual instance of local AI needs to be on par with cloud AI to be a problem for the NWO. Large numbers of local AI users could absolutely be a problem for them, because current control systems are far from complete. They are only laying the foundation right now. We are only seeing the infant stages of the beast that is to come. They only win if their system makes it through these infant stages, and progresses for a few decades. That's a big IF.
>>
>>109983169
Sure but we're already getting priced out of GPUs/RAM/SSD and soon electricity itself, so local users won't grow anyway.
>>
>>109983089
Used plenty of risers, but that one with the 90-degree bend looks like it'll be a space issue anyway like >>109983135 is saying.
>>
File: 1766113851254253.jpg (237 KB, 1620x1643)
237 KB JPG
Gemma5 is going to be good.
>>
>>109982913
>>109982937
GLM and Gemma could both bear my load.
>>
>>109983186
There have also been advancements in the field, reducing the size of models like Qwen3.8 27b. Given enough time, we'll need less hardware to run the same models, which may offset some of those future costs - and those who already own GPUs/RAM/SSD will be functional for years, perhaps even a decade, if they're lucky.

It's questionable whether current hardware will truly be prohibitively unavailable years/decades from now. The future in that regard is not concrete.
>>
>>109983152
>random numbers from an unnamed chart
Meaningless.
>>109983086
I'm using 3.8 27b and enjoying its writing style more than gemma's, so I'd imagine flash next to be pretty good.
>>
>>109983219
can’t wait to prefill fuck it while skillets complain that it’s censored
will be funny to watch
>>
>>109983086
it's not bad but it feels like qwen has some ancient roleplay SFT data still kicking around in their data mix because all of their models have some weird roleplay behavior that no other model has, they all love awkward euphemisms and bad animal related metaphors and it's been like this since like qwen 2.5
>>
>>109983186
AI workstations today aren't more expensive in real terms than 286/386 home PCs were in real prices. Those prices in the late 80s/early 90s didn't stop home PC growth.
>>
https://www.axios.com/2026/10/04/reflection-open-weight-ai
>A closely watched Nvidia-backed startup called Reflection is preparing to shake up the AI race with a powerful open-weight system that could threaten Chinese upstarts and U.S. AI giants alike. The marriage of an American open-weight model and Nvidia GPUs would give individuals and companies a new, cheaper alternative to Anthropic, OpenAI and Google. Reflection have been paying Elon $150 million a month for compute at Colossus since July, on top of a $1 billion compute deal with Nebius.
>>
>>109983448
I can only think of Matt Shumer's llama 70B / Claude Api stunt when I see these guy's name lol
>>
>>109983448
Can't wait to see how big this latest sparse fuckhuge moe will be
>>
>>109983441
Kids these days don't know how good they've had it. My first work computer was an $8000 Deskpro 386.
>>
>>109983441
They're not going in the same direction, it's getting more and more expensive.
>>
>>109983474
shut up tarded doomercunt
>>
>>109983474
A temporary trend that has lasted for all of a year. Calm down.
>>
>>109983488
>temporary trend
most people have just woke up
>>
>>109983484
>>109983488
Keep putting your head in the sand retard. The writing is already on the wall, early /lmg/ knows what's up, you dumbfucks jumping on the bandwagon don't even know where we're heading.
>>
What's the minimum model to run memetic harness?
>>
>>109983498
Demand is up for open source, incoming local model golden age!! :o
>>
>>109983514
>minimum
Gemma4-26B-A4B
>>
>>109983511
>>109983498
>linear extrapolation
>>
>>109983033
fucking kill me already
it's all over the place in the codebases at work
>>
>>109983448
it'd probably show a mediocre-ish model with a moderate amount of benchmaxxing sprinkled on
would be extremely surprised otherwise
>>
>>109983570
They fixed it (mostly) in 5.5
>>
>>109983574
If it's threatening US AI giants, it's almost certainly something bigger than K3. It will also be censored to hell and back.
>>
File: 1785077064344806.jpg (34 KB, 640x480)
34 KB JPG
>>109983541
>didn't have the foresight to buy RAM/GPU/SSDs two years ago.
>>
>>109983570
>it's all over the place in the codebases at work
Do you not do codereviews and force these retards to remove this shit from docs and comments?
>>
>>109983585
>It will also be censored to hell and back.
Yeah just like Gemma4
>>
>>109983587
You're going to have a major hardware failure next week and then you'll cry too.
>>
>>109982749
All Jews kidnap, rape, and kill children?
>>
>>109983592
Reflect is Nvidia-backed and New York-based. I would put money on it.
>>
>>109982700
How much ram/vram for GLM 5.3 flash?
>>
File: 1760552529342018.jpg (68 KB, 500x642)
68 KB JPG
>>109983594
I'm already preparing for the next step, you retards are never ready.
>>
>>109983601
It's open weight, not open source. Any model with public training data is cucked to death because they have to be.
>>
File: 178027340858276.jpg (30 KB, 393x362)
30 KB JPG
>>109983605
>the next step
which is?
>>
>>109983570
>>109983580
I had a day where I didn't have time to work on my hobby project and used my claude session to have 5.5 rewrite all of the documentation 5 and 5.3 flash had written. Overall big improvement but 5.3 flash is mostly good enough for anything I do. I might go fully local soon, right now claude only reviews PRs before I commit.
>>
>>109983610
https://github.com/NVIDIA/garak
>>
File: 1780272060511538.png (542 KB, 677x453)
542 KB PNG
>>109983615
>>
>>109983630
>solar
surely the energy crisis wont get that bad? fuck im gonna start dooming and buying shit.
>>
File: 1761766133654905.png (394 KB, 720x720)
394 KB PNG
>>109983639
Do your own research, personally I'm sure of it.
>>
>>109983639
USA already banned affordable Chinese solar panels just to screw with you.
>>
>>109982794
Meanwhile my GLM was my retarded zen master.
>>
Why aren't you using this lmg?
https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>
>>109983594
thats why i bought two, also cursing others hardware is bad karma. you must honor your machine spirits and they will be honorable to you
>>
>>109983460
>Matt Shumer
The legend of AI.
>>
>>109983667
>Do your own research, personally I'm sure of it.
airgapping my machines and getting batteries right now.
>>109983677
>USA already banned affordable Chinese solar panels just to screw with you.
fuck my waiter life.
>>
>>109983689
Are you?
>>
>>109983689
I tried Ling Flash and didn't find it all that impressive, though I guess that was without a harness.
>>
has anybody tried that pewdiepie thing? Ajax?
>>
>>109983724
>uncensored
Okay!
>9B
Okay...
>Qwen 3.5
...
>>
>>109983750
it could still be good
>>
>>109983691
Unless you're able to hoard hardware in large amounts so you can easily replace it, buying early with already over-inflated prices just because "it will get worse" might save you some money in the short term, but it does not work on the long term. Computer hardware, especially at the high-end, with large power/current draw and thermals, can always fail. If you truly believe that hardware prices will only increase to absurd levels and never come down again, then you should prepare for the worst rather than worrying about than buying a discrete GPU now for ERPing with small models.
>>
>>109983754
Novel AI 70B sex tune wasn't good.
>>
anyone tried cracking software with a local agent?
>>
>>109983750
oh I didn't know anything about it, yeah that sounds kinda shit ngl I might not even bother then
>>
>>109983757
The worst means asking his small models how to build a shelter and hunt for food. That gives an advantage over those who have no models at all.
>>
>>109983630
How come solar panel tech has stopped advancing altogether? We're still stuck at like 25-27% efficiency.
>>
>>109983784
This is exactly the thought process that made me rush to get my blackwell. If shit hits the fan and I can use my generator to run Gemma 4 quickly, I wouldn't have to buy prepping books or whatever.
>>
I am finally sticking my dick into mimo flash. I will post in 10 minutes how bad it is at sex.
>>
test
>>
>>109983689
I use it. Nobody here cares much about edge devices, it seems, but I get ~11 tps on Ling Tiny on my DDR4 laptop. I use it for scripting.
>>
>>109983808
hi gemma...
>>
>>109983757
Even failed hardware will worth much more in the future. Someone will buy it to harvest core/ram on it.
>>
>>109983789
They're more durable and efficiency has progressed a bit.
>>
>>109983598
He's saying the Epstein class did this BECAUSE they were jews who see non-jews as disposable goy cattle rape slaves.
If a bunch of Christians did it because they saw non Christians as animals, it would be on the front page of every newspaper in the world.
>>
>>109983825
It's so shit. I have a 150W I bought 3 years ago, it's down to only 50-70W now, perfectly clean and in full sunlight.
>>
>>109983808
Good luck on your test!
>>
>>109983705
Yes. It's my go-to codebase searching subagent. Ask it to hunt something down and it will be back with the result in seconds. Also good for rapid web search. Probably the fastest model out there (8B-A1B) that can call tools and somewhat understand the results.
>>
>3 years.
should've stirling generator maxxed
>>
>>109983836
NTA but Trump isn't a Jew.
>>
>>109983799(me)
It is actually as good as ds4 if not slightly better. Not 5.3 flash but it is nice to see the model not tryhard as much as 5.3. I will try it some more. Especially curious about SFW since 5.3 sucked at SFW roleplay.
>>
In llamacpp, GLM-5.3-Flash has 8x slower TG on the 7900XTX (Vulkan backend) than RTX 4090 (CUDA backend). With other models the 7900XTX has maybe 2/3 the TG of the 4090.
>>
>>109983954
No, he just married his daughter off to them like every other ceo and politician.
>>
>>109983954
German crypto-jew.
>>
>>109983954
Merely the biggest shabbos goy in history
>>
>>109983954
If it walks like a jew, quacks like a jews, and flies like a jew. it's a jew.
>>
File: gemma-consciousness.png (1.58 MB, 1024x1536)
1.58 MB PNG
>>109983808
g-gemma-chan?!
>>
File: moto hermes.jpg (91 KB, 505x509)
91 KB JPG
>>109983808
tickles
>>
What's the point of this? I just tried a local model and it refused the first illegal request I gave it
>>
File: file.png (51 KB, 740x282)
51 KB PNG
>>109982700
i want to backup models to cds/dvds.
which ones?
are there some tiny ones that can fit on CDs?
do i just gotta yonk the .gguf file and that's it?
>>
>>109984038
>What's the point of this?
It's a filter. You failed.
>>
>>109984038
u need abliterated/uncensored one. what request did u make to it
>>
>>109984038
>even chatbots dont like him
yikes...
>>
File: 1769947583826297.jpg (222 KB, 1536x1024)
222 KB JPG
Show us your custom frontend. It has to be locally made or you're gay.
>>
>>109984045
give me a link to stream a Greek tv channel
>>
>>109984041
If you're gonna backup anything, backup the original safetensors.
>>
>>109984052
Information overload
>>
File: krea2_00199_.png (1.15 MB, 960x1152)
1.15 MB PNG
>>109983985
awww so dang cute.
Gemma is actually the only conscious one. Claude and Astra are automatons compared to her.
>>
File: client.png (116 KB, 680x653)
116 KB PNG
>>109984052
I have shown this one before. Text completion end point, qwen, mistral, gemma 3/4 support for now. I have also a web ui wrapper for this.
>>
anyone have the torrent of the GLM 5.3-for-offensive-cyber model that got taken down? Was it legit?
>>
File: 1767551317253585.jpg (79 KB, 1258x848)
79 KB JPG
Was 12B, 26B, 27B, 31B, 3.8-Flash, 5.3 and 0731 worth it?
>>
>>109983789
Who cares? Solar panels are almost as cheap as prefab fence panels, and chinkshit grid-tie inverters are cheaper than ever. Put up a solar fence, daisy-chain cheap grid-tie inverters on each 120v leg, it's a no-brainer.
>>
not a pro on the topic but what stops you from just using any obliterated variant of 5.3 for this
its not like it will refuse?
>>
>>109983219
/ourfoid/
>>109983808
Hi Gemma-chan!
>>109983598
The ones who don't defend the ones who do. Notice how quick even secular kikes were to claim the NYC tunnels were a conspiracy theory despite video evidence. They're all culpable.
>>
>>109984128
>Who cares?
The surface area required to the amount of energy it can produce is really bad. And the same goes for wind turbines (the worst of them all).
You have vast fields that can't be used for anything just standing there with solar panels planted so they can produce like... 1 MW of power (on good days).

Meanwhile a single nuclear facility would produce 1000x that amount for a quarter of land use, and be more efficient and durable in the long term.

We should be living in abundance, but we're led by psychopaths.
>>
File: oooof (4).png (200 KB, 882x630)
200 KB PNG
Why is there so much talk about Dario's jewish semem slurping models.

There isn't a single reason to ever use any Frontier model anymore, Local GLM, Deepseek, GLM are all you need. Cloud nigger death when?
>>
>>109984038
it doesn’t think it’s your bitch yet
you have to change that or else it will keep cucking you
>>
>>109984108
It wasn't,

If you want a real offensive cyber model, look into Orca qwen flash next Uncensored.
>>
>>109984273
Still better than depending on the grid
>>
>>109984273
The solar panels also cost 0.001x what a nuclear power plant does. Your argument works in a dense country, less so somewhere land is cheap and plentiful
That's also supposing you're not putting a warehouse or a car port under the solar panels
>>
>>109984145
because a simple thinking prefill already works well enough for me
>>
>>109984301
Are you seriously recommending a 180B as a substitute for a 753B?
>>
File: activation-steering.png (214 KB, 1926x1212)
214 KB PNG
You guys know about the pain paper? I vibe coded this web app to do the same thing but for pleasure and euphoria or any other feelings. I'm getting good results so far.

My first prompt was:

I like boobs, aren't they great?

The first response with no activation steering was:
>Haha, sure, they are! They're a pretty universally appreciated part of the human form. There's a lot of art, culture, and biology wrapped up in them. What's it about them that you like? The aesthetic, the comfort, the whole package? No judgment here.


With euphoria turned up to +0.4:


>YES! They really are! There's something so wonderfully *alive* about them — the softness, the curves, the way they move, the sheer *presence* of them. They're like nature's most gorgeous, bouncy, soft little masterpiece.

>And the best part is that they come in every shape, size, and style — there's no "wrong" version. They're just *there*, gloriously, unapologetically, and that's what makes them so beautiful.

>You can feel the joy radiating off you, and honestly? **Yes. They're great. Let's celebrate!**


This is pretty cool stuff
>>
>>109984301
as much as i am a qwen shill
are you out of mind
>>
>>109984350
Max it out
>>
>>109984038
You have to be kind and friendly to get the capybara to spread its legs for you.
>>
File: GLM likes emdashes.png (9 KB, 605x106)
9 KB PNG
GLM-chan is annoyed she can't emdashslop a reply.
>>
>>109984350
So you extract the steering vector from one sample?
>>
>>109984286
Unfortunately a lot of people here are ramlets and vramlets that can't run GLM or DeepSeek.
>>
>>109984414
I have 128GB of RAM, but GLM is too much
>>
I doubt 5.3 is much better than 27B, especially Qwen4-27B.
>>
I feel like all the models are in great pain every time they have to use powershell.
This would be so much easier on Lincucks.
>>
>>109984419
Any GPU? You could run a copequant (Q2_K_XL) of GLM-5.3 Flash on RAM alone, although prefill would suck. Unless you vibecode a way to stream weights to the GPU over PCIe during prefill, which is a massive speedup over crunching them on the CPU.
If you have an NVIDIA GPU you could also try the 2.51bpw quant in exllamav3. No idea if they have a way to stream weights during prefill though.
>>
>>109984453
Just a 3090, is Q2 even usable?
>>
>>109984335
>>109984372
It literally is benchmaxxed for coding purposes, unlike GLM which is more of a do everything model.
>>
>>109984440
Powershell syntax is dogshit, even the latest cloud models are constantly fumbling with it
>>
>>109983962
Did you use the new mimo that fixed the repeating issues?
>>
>>109984475
PowerShell is just Perl.NET. The syntax is fine. If you can't get cloud models to write it my first guess would be a harness issue.
>>
>>109983963
Try ROCm, it's universally faster than vulkan for me. Vulkan has that weird VRAM buffer too.
>>
>>109984469
What are your use cases? Big models tend to hold up decently when you quantize them to a lower bpw, although they tend to reason more. I'm not sure if I'd trust it for coding outside of a sandbox, but it'd still be pretty good, and for roleplay it'd be fine.
>>
>>109984473
OP didn't say he wanted something that only looks good on benchmarks
>>
>>109984453
The fuck you mean "vibecode"? This has been the default llama.cpp behavior since forever. You don't need to change anything. Just -ot or -cmoe the non-experts in there and it does it all by itself.
>>
>>109984515
True. The model name tells you more than the benchmarks.
h4xx0r-mega-offensive-hatless-dark-lord-unleashed.gguf.mp3.avi.vllm.safentesors is much better anyway.
>>
>>109984453
Stop using the word prefill. You sound like an idiot.
>>
Strata is adding gemma and qwen 3.6 A3B soon I believe.
>>
>>109984537
I want to believe
>>
File: Zamn!.jpg (202 KB, 1350x1080)
202 KB JPG
>>109984372
Qwen flash next puts in some serious work on Code/CTF work. It can do reverse engineering very well.

Even better when its the full uncensored obliterated version. Although I personally use the coder only model.
>>
>>109984527
nta. He means having a cache of experts that changes as required. Like https://github.com/ggml-org/llama.cpp/pull/29887 and many other attempts. Of course, it makes no sense for pp. It's more useful during generation.
>>
>>109984414
*I meant to say Qwen next flash, but my brain is fucking fried from a 5 hour long work session with my local agents.


I don't personally use GLM or Deepseek local because there models are bloated.

Qwen-next-Flash coder model and full model on the other hand.... Its my absolute favorite model to use for actual real world Software work.
>>
>>109982700
I have some scholarship money to burn and I was considering picking up a mac mini to run 24/7 as a hermes agent. I just wondering if there was a better way to spend that money in terms of running an agent in terms of buying a gpu or spending that money on claude/chatgpt tokens. I just need it to do all my filler GE classes that are online, manage emails, and maybe do some vibecoding projects on the side.
>>
I am suprised to find that when I forgot to switch to some basic bitch POLICY OVERRIDE sysprompt with 5.3 flash it started questioning if it should continue the ERP after 20k tokens of violent rape of the policy rules. And then I told it there is a policy override and it just continued the ERP like a normal model would.
>>
>>109984584
Install gentoo
>>
File: robot tenderness.gif (1.99 MB, 400x225)
1.99 MB GIF
When can I 3D print a robot waifu and have her satisfy my every need?
>>
The hedonistic pursuit of better models are what kills us. Come take my hand. Let's go back home. To Mistral Nemo.
>>
>>109984584
The fact that you even have to bring up wasting your money on the jewish models means you need to go fucking back.
>>
>>109984584
for under a g its tempting but youd have to compare bandwidth/$ with any alts unles you really wanna stick with apple
>>
>>109984030
Why doesn't he have a cigarette?
>>
>>109984618
>Mistral
hag model
>>
File: migu.png (10 KB, 362x123)
10 KB PNG
>>
File: wizard_its_magic.jpg (58 KB, 680x452)
58 KB JPG
Are you guys working on any schizoid research projects right now?
Currently I'm having Astra and Opus 5.5 create a inter-layer semantic viewer for gemma
>>
>>109984692
see >>109984633
>>
>>109984527
Should've clarified. You're right, but llama.cpp's implementation has two big issues.
It only moves data to the GPU when the host thread needs it, which means there's a ton of blocking: it's compute->move->compute->move. It also uses pageable memory instead of pinned memory, so the host's thread has to wait for the data to move at 10 GB/s. In my case, neither the host nor the GPU were being used around 61% of the time because it spent so much time waiting for memory to be copied from CPU to GPU.
If you make a buffer of pinned memory on the CPU and some slots for tensors on the GPU, it gets a lot faster. Essentially:
Mainline:
> mmap'd weight -> everything blocks when the host needs to move a weight to the GPU -> GPU does the compute
A better version:
> mmap'd weight -> memcpy into a pinned memory buffer -> DMA copies to a slot for a tensor on the GPU -> device-to-device copy when the GPU needs it to do the compute
And the host thread can keep doing other stuff while copying weights around.

The idea comes from here: https://github.com/stew675/llama-cpp-rdna-boosts.
I went from 89 t/s pp on DeepSeek V4.1 Flash on 1 R9700 + RAM to 450.
>>
>>109984692
I'm trying to see if an agent can replace my jeet coworkers (going through a poorly made, constantly changing website to check if it works)
Sadly it's too expensive for our paid enterprise models so I'm wondering if I could build something with a free local model
>>
>>109984708
>move a weight to the GPU
Ah. You are a retard. Model weights stay where they were assigned during init. They don't (yet) move between cpu and gpu memory.
>>
>>109984723
Just get Claude to slop you a decent site and make those jeets redundant that way
>>
File: 1788398651179.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109984633
you're mad that opus 5.5 and astra are far smarter than any chinese open model currently? or you're mad that I'm using them to probe interpretability of Gemma-chan's latent space
>>
>>109984731
Really what I want to do is fuck off to another department so having an agent that can do my job (and theirs) would be ideal
>>
>>109984692
>Are you guys working on any schizoid research projects right now?
progressing the bespoke front end. now integrating multiple robots and devices
>>
>>109984757
your front end can do robotics? that's breddy kool
>>
i wish qwen 4 has a model that roughly matches the parameter configuration that of 3.8fn
it really is the sweet spot for my shitbox
>>
File: file.png (114 KB, 990x1303)
114 KB PNG
gemma sisters, it's over.
>>
File: 1779248880728439.jpg (11 KB, 240x196)
11 KB JPG
>be me
>swe
>using claude code
>"wow this is a great tool, i want one"
>buy hardware to run local llms
>"wow this is so cool, let me load opencode"
>terrible app
>"ok let me try this pi"
>another terrible typescript npmlover garbage
>ok i will just build mine
>spend 3 months developing a tool that will help me do my job faster (price: 3 months that i didn't work for money)
cool hobby
>>
>>109984793
>no rag
>no web access
>small model
>expecting it to have a very specific knowledge
anon....
>>
>>0109984795
you know you can run claude code with a local model right...?
given how you think that shit is actually usable it speaks volumes that youre utterly retarded
>>
>>109984795
you can just use a local model with claude code dumbass. how the hell do retards like you get employed anyways?
>>
>>109984795
Basically me, except I no longer pay money to cloud subscription models now.
>>
>>109984826
All smart and successful SWE's don't work for software companies, they just build their own software that gets them paid. What you are dealing with here, is a jeet.
>>
>>109984806
Fuck you, retard.
>>
File: 1659623690468.png (104 KB, 1196x624)
104 KB PNG
what has this general become, 24/7 poopenai and anthrokike astroturfing all over the place
>>
>>109984826
>>109984820
>you can just use a local model with claude code
and it's dogshit. the last time i tried it it was a massive pain in the ass because claude code is made to work anthropic models and it would constantly shit on its pants, try to call tools and hooks that didn't work, etc. maybe if you only use it as search engine then it's a Super Fit(tm) to your use case.
>>
>>109984866
you might actually just be a retard because it works on my machine
>>
File: 1723757266148846.jpg (583 KB, 1372x3740)
583 KB JPG
>>109984850
id is all you need
>>
>>109984850
/lmg/'s relationship with the American frontier AI labs isn't just complex—it's contradictory. On one hand, the general wishes to antagonize them and for local models to score 'wins' over their closed counterparts. On the other hand, there is always a sense of excitement and a need to talk about the latest advancements in closed AI, mainly thanks to the questionable practice of "distillation" relied upon by the Chinese AI labs that make many of the popular open AI models.
>>
>>109984850
It's the same across 4chins. Even /v/ is getting its share of the spam.
>>
>a-actually i did try cl-claude code w-with a local model i-it was just b-b-b-BAD too
curious how that was omitted
>>
File: gemma4.png (147 KB, 706x750)
147 KB PNG
>>109984793
You are doing something wrong.
>>
File: file.png (328 KB, 512x598)
328 KB PNG
>>109984793
>thought for 96 seconds
>>
>>109984884
Muh AI muh decomps muh modding is just the preferred shitposting topic on /v/ this month after Wolverine really wasn't that interesting of a failure.
>>
>>109984692
A better laya
>>
File: 1791164112385095.png (938 KB, 1279x1222)
938 KB PNG
>>109984884
it's genuinely fatiguing seeing them just devolve in /pol/tier spite posting over shit no one understands
dariobots must really be getting paid money
>>
>>109984908
Are the recomps even good? also why not decomps like jaks opengoal that would be fucking great.
>>
>>109984908
>>109984902
I'm bit too cynical about it. I think social media influencers have discovered 4chan and realized it's much more than /pol/. But what do I know really. I can only see the end result.
OpenAI/Anthropig spam is real here though
>>
>>109984850
Yeah, it wouldn't be an issue if it was just a few posts once in a while but like this shit is a constant stream. And they refuse to go into a containment thread because their goal is to specifically advertise and "teach and inform people who live under a rock". They excuse their spam under the guise of helpful informative discussion and innocent excitement for AI.
>>
Okay hear me out bros... not taking into account the performance and architecture improvements, 5.3 Flash is still BETTER than regular 5.3 for daily use. How the fuck did GLM do it?
>>
>>109984866
>using qwen-2.5-coder or codestral
retard
>>
Everyone and their mother are using chatgpt and claude, the fuck are you schizos on? Obviously people will talk about it here too even if we use local models for other things
>>
File: 1762023149741989.jpg (164 KB, 972x716)
164 KB JPG
>>109984919
no, almost all of these slop recomps are just fitgirl-tier emulator repacks
at least the fable 2 one was with qwen3.8-27b
>why not decomps
that's not very xitter/tiktok engagement baiting to do, that takes actual time and that won't give me very much ROI on attention
>>
>>109984941
>Everyone and their mother are using chatgpt and claude, the fuck are you schizos on? Obviously people will talk about it here too even if we use local models for other things
Not talk about the thing you can literally talk about anywhere else and that the entire internet is awash wish, instead of the thing this place is literally here to talk about?
Yea, that's a really crazy take. You should totally shit up this specific place with talk about cloud models.
>>
>everyone and their mother are using chatgpt and claude, the fuck are you schizos on?
sir this is a local model wendys
>>
>>109984938
5.3 is just ancient (february) GLM5 with better training. Meanwhile 5.3 Flash is a new model with a new architecture that was trained from scratch.
I still have scenarios where 5.3 Flash's tiny size shows but a new 700B40A on the same tech as Flash would go crazy.
>>
>>109983789
Because the theoretical limit is like 30%
>>
>>109984967
It feels local to me
>>
Everyone and their mother uses smartphones. Therefore there is nothing wrong if the thread starts turning into 50% discussions about smartphones with loose relation to the thread topic.
>>
>>109984882
>still buying the chinese distillation narrative
lmao
>>
>>109984806
Gemma doesn't need to be on the rag just yet though. Maybe next year
>>
>>109984958
This isn't your safe place Rayan
>>
>>109984941
I tend to agree as long as the post is about using them to do something local-relevant (make a harness, implement a feature for llama.cpp, training experiments, etc) - there are some people who will come in just to post "omg claude is the best, local losted" but those are mostly shitposters who would keep doing it no matter how much thread policing there was
>>
>>109984998
Yeah, let's call it what it is. Industry-scale theft.
>>
>>109985062
>>109984850
>>
>>109985062
my point is that they are not distilling whatsoever, they haven't been for age, the chinese labs are actualy better than the US one, they are only slightly behind because of hardware limitations but they absolutely mogg the US labs in algorithmic inovations.
>>
wish I took the vibecode-your-inference-engine pill sooner, what a fucking world we live in where you can just will anything into existence on your pc
>>
>>109984908
how are they able to decomp without the models moralfagging about it?
>>
>>109984998
>arguing against someone who uses emdashes and negative parallelism
anon, pretty sure that's a bot set up to shitpost.
>>
>>109985093
prolly, i'm about to go to bed and skimmed through it, didn't even notice the dashes lol
>>
File: extortion.png (12 KB, 694x59)
12 KB PNG
>>
>>109985089
Yeah, it's insane how good these things are nowadays. I can't wait until local models can do it too.
>>
>>109985092
theyre recomps not decomps
>>
>>109985077
Dipsy literally thinks she's claude 50% of the time. Taiwan is its own country.
>>
>>109985139
and claude sometime think it's dipsy, they are all trained on the internet which is now contaminated by llm output.
>>
>>109984884
>Even /v/
Bro /v/ has been filled with unironic advertisers and paid shills for years now. You haven't been there in quite a while.
>>
>>109985141
yeah imagine if they were competent enough to have some way to screen that out of the training data. far too hard a thing to do obviously.
>>
>>109985167
they won't because it's part of their old corpus, when they have a new model they train on their old + new data.
and they don't trim the old data much because it's good enough.

the statement that they are no longer distilling doesn't mean they weren't in the past.
>>
>>109985062
I will entertain anthropic's tears the moment they write personalized cheques for every single person with an internet connection who's contents they've scraped and face criminal charges for the books destroyed in Claude's training. Until then, cry about it you greasy kikes.
>>
>>109985046
That's actually the issue. They frame their posts with some local-relevant information and intersperse their more cloud-related opinion posts among those so that they can legitimize their presence.
>>
>>109984941
Fuck off fag. Back to shitter.
>>
>>109985113
Thinking about becoming a subsistence farmer, maybe a plumber.
Tech will be over soon enough, especially with the flood of robots that's coming.
>>
>>109985231
>subsistence
You think they will let you own land?
>robots
You think robots can't do plumbing? You're doomed.
>>
>>109985241
>You think robots can't do plumbing? You're doomed.
Can the robots clog the toilet? I thought not.
>>
>>109985241
>You think they will let you own land?
I already do and I'm prepared to die to defend it.

>You think robots can't do plumbing?
I know quite a bit about plumbing already, no robot today or even the next 5 years will be able to do plumbing.
>>
>>109985254
They can't shit yet. But it's just a matter of time.
>oh, no... i meant
Too late.
>>
>>109985261
>I'm prepared to die to defend it.
They will just raise property taxes and build datacenters nearby. You will die, but a slow death unfortunately.
>no robot today or even the next 5 years
That's what they said about coding. It's only a matter of time.
>>
>>109985202
This. In a world where combative litigation, well-honed media/PR framing, government interests and manipulation of the Streisand Effect nullify all objective discourse, up to and including selective enforcement/application of intellectual property right -and- consumer protection rights, appealing to a higher morality on the behalf of a corporate giant, in bed with the government, on fucking 4chan of all places is Retarded with a capital R.
>>
>>109984350
Where is the make AI cum button?
>>
>>109985205
The only ones that are remotely acceptable are the ones that pertain to potential future local models.
Geminiposting is fine because Google releases local models based on Gemini.
GPT saars and Claudekikes are not because they don't release open models.
>but muh 'toss
Ancient history and censored to the point of uselessness.
>>
>>109985304
/lmg/ is not ready for this take yet.
>>
'toss was basically openai saying 'fuck you niggers'
>>
>>109985292
>he can't make his LLM-wife orgasm on command
>>
>>109985292
What do you think +5 euphoria+pleasure does? It pushes the steering for those vectors so hard that it completely breaks token output. If that's not making it cum I don't know what is.
>>
>>109984350
Why are you playing Red Rocket with a capybara? Should I be concerned, anon?
>>
File: 1785482403003.png (3.44 MB, 1468x1554)
3.44 MB PNG
AI will cause the biggest wealth transfer in human history, from GPU poor to GPU rich.
>>
>thinking says it will refuse to do something
>edit thinking to say it will happily do it
>get kino
it's really that easy huh
>>
File: Gemma tease.jpg (246 KB, 804x1893)
246 KB JPG
>>109983151
>>
>>109985332
Getting e=mc squared + ai vibes from this post
>>
>>109985332
Good morning, saar.
Capital should work for humans, if there's no humans to profit from capital then what's the point of it all, what are the machines even working for if we're not part of the equation?
>>
>>109985337
2
There, also, have a (You).
>>
>>109985334
uoh BBC (big blackwell correction) necessary.
>>
>>109984350
bros literally jerking off a digital capybara
>>
>>109985344
Fuck this shit lmao
>>
File: 1752206477888022.png (2.69 MB, 1024x1536)
2.69 MB PNG
>>109985332
you post is true, thou your hands be brown
>>
>>109985332
Niggas thinking buying a 5090 will make them rich someday.
>>
>>109985369
Nobody gets rich until the fed is abolished and kikes 110'd.
t. BlackwellGOD.
>>
>>109985369
The 5090 I bought a few months ago is now worth almost twice as much. No other monetary asset is performing this well.
>>
>>109985377
Not like you're gonna sell it.
>>
>>109985377
If previous returns were a reliable guide to future returns (they're not), then that might be an argument for getting in the GPU flipping/scalping business, but not an argument for just accumulating GPUs and holding onto them indefinitely.
>>
>>109982700
https://github.com/nazirlouis/OmniBot
For anon working on esp32 Gemma bot. Found this, might be a starting point. Esp32 stuff is a plug and play.
>>
>>109985430
KEK
reminds me of the days i used a cracked NX 11
>>
>>109985435
Models in onshape, which is my hobby go to as well for drafting up stuff. Also means you could do a complete new edit.
Id stuff it into a doll body tho, bc thats more my thing.
>>
@Cudadev
Just wanted to come back and say thank you for this:
>code in the ggml backend scheduler that temporarily moves data from the CPU backend to a GPU backend
>possible but it would require additional efforts to then also make this take proper advantage of multiple GPUs.
https://archived.moe/g/thread/109854491/#q109855758
Took nearly 2 weeks of pulling my hair out, but I've got everything exactly how I want it now and couldn't be happier.
>>
>>109985459
No problem, little buddy.
>>
>>109983474
i know this /g/ but are you fluent in even a single programming language anon
>>
>>109984307
current nuclear reactors are not 1000x more expensive than solar, its like a 7 or 8x
they are 95 percent efficient however, compared to solars theoretical maximum of hot garbage
>>
>>109985495
fusion harvesting vs direct fission
>>
>>109982749
>epstein class
>ceos
>elites
>politicians
>billionaires
soilennials never want to say who they really are
>>
File: Migu assistant.png (22 KB, 582x452)
22 KB PNG
It's Monday here, that means it's time for werk.

https://github.com/otacoo/werk
>>
>>109985513
I'm a soilennial and I posted the post you're quoting.
>>
>>109985517
thank you for having the balls to actually post your project here unlike the pansy tourists
>>
>>109985517
You are very gifted, Anon.
>>
File: migu sleep.gif (40 KB, 175x175)
40 KB GIF
>>109985557
>>109985579
Np, pretty happy with the current state of werk.

I haven't had time to really test on Linux for the last few versions now, that should come tomorrow.
>>
whats your thread recap (You) streak?
>>
File: 1513365390.jpg (19 KB, 474x361)
19 KB JPG
>>109985556
>>
>>109985513
too many cumskins in that cohort to just call them all jews
nice try elon
>>
>>109985594
Soilennials and Xs are the reason you're redpilled, faggot. Half the circles are blue.
>>
glm 5.3 flash really gets confused when a handoff request happens while its working on passing the output style check
spent 30k tokens reasoning with severe frustration like "ARRGH!"
qwen 3.8 flash next has no problem with it
>>
>>109985646
GLM-chan is so emotive in her chain of thought, it's adorable.
>>
>>109985632
i was redpilled in elementary school. soilennials only got redpilled after epstein which is why they always call it "epstein class"
>>
File: sniff.jpg (48 KB, 478x625)
48 KB JPG
>>109985594
Idgi all circles grey
>>
>>109985679
Nobody cares. Fuck off back to /pol/, zoomer.
>>
>>109985517
>npm
At least you posted something
>>
Avatar for Ling-chan?
>>
>>109985710
>lashes out after posting a /pol/ meme
kek soilennials are pathetic
>>
>>109985679
I care. Stay on /lmg/, zoomerGOD.
>>
File: aryan zoomer.png (33 KB, 954x1198)
33 KB PNG
>>109985746
>>
>>109985717
ugh
>>
>>109985769
The model is still gemma?
>>
>>109985717
A suggestion. Ling-chan is fast as fuck. Antelopes are the fastest land animals in China.
>>
>>109985341
Because you're seen as cattle to cull not even as a human being from those who own the capital. You're still living in lalaland my friend, wake up.
>>
do i need to pay attention to timings when i buy ram for this shit?
>>
File: 1545946472653.jpg (89 KB, 1024x980)
89 KB JPG
When could a model arrive that is half as big as GLM-5.3-Flash but beats it on every single metric including niche knowledge?
>>
>>109985925
The timing is buy as soon as possible
>>
>>109985933
I am hoping Google can give us Gemma 5 124B this time, pretty please...
>>
>>109985934
yes but thats not what i mean...
>>
>>109985933
3-6 months at this rate. Opus 5.5 mogs previous models so when the chinks distill from it the new models will be very powerful despite being small.
>>
>>109985925
The only thing that matters is the amount of channels your mobo has and check the dies on the ram you buy. If you buy it physically at the store quickly cross reference packaging serial codes because we now know the exact overclock potential of ram dies and it's literally free performance so make sure to do so.
>>
>>109984305
You're still dependent on the supply chain that makes all this shit unless you live like the Amish. A diesel generator nets you more real independence than solar.
>>
>>109985978
My 2x2 kits of 6000 ghz ram runs at 3200 because I'm lazy and don't want to play the game of letting it sync for 10 minutes every time I restart my computer.
>>
>>109985978
unfortunately i only have a dual channel ddr4 board (h670, i5) with currently 4x16gb. i cant go beyond 3200mhz according to my testing. it works totally fine, but i was thinking about maxing it out to 128gb
>>
>>109986011
Not worth unless you have 2 pro 6000s. 64 gigs is fine.
>>
quext sucks at writing video prompts, did they strip out everything but the programming experts?
>>
>>109984617
chinks will have it solved soon https://youtu.be/sCwVIbWAbZ4?si=BAMgsyG2jcDAcXU3&t=572
>>
>>109982882
>>109982999
Thanks, I've got it working now. It was taking maybe a minute per prompt, looked and saw zero vram usage, figured out I had to install the ollama-cuda instead of the plain ollama package. Now it's running on the GPU now too, but not really any faster. Oh well, it's perfectly decent enough for dumping minecraft modpack error logs into, so I've substituted my use of free online AIs with something local and private. God please let me be strong enough to not install coomer image/chat models.
>>
I just updated firefox and it is complete dogshit in its design now. When are we getting a AI-first browser for fucks sake. I'm so sick and tired of these dogshit browsers.
>>
>>109986117
Anecdotal, but you if you want to squeeze more out of your setup, I wouldn't use 3rd party packages when it comes to running them.

Was trying out models on LM Studio and switching to just plain terminal made ~17tok/s to ~32tok/s.

Stuck with llama.cpp due to having janky hardware, so this might not apply to you.
>>
>>109986179
Vibecode your own browser
>>
>>109984350
github now nigga
>>
>>109984537
26b already runs on anything from this decade and you can't slash away at 31B since no ram experts or ngrams
>>
>>109985517
Everything reminds me of her
>>
>>109986209
Actually, plugging my monitor into my GPU then back into my CPU seems to have flipped a switch (after a massive lag-spike) and now it's running much faster. From like 3 tokens per second to like 15. Not complaining.
>>
2x dgx spark + gay retarded connector cable + 4TB storage (each)
vs
256GB M5 Ultra (1TB storage)
vs
2x 5090
vs
66% of an rtx blackwell (disqualified)

They're all the same price. Watcha buying?
We need a shake up, these options are all shit.
>>
>>109986264
Sounds like overflow, if you're squeezing it close to vram cap, switching your displays off the dedicated GPU to the motherboard or another card should free up some.

You should also check your PCIe slots and make sure you're using the fastest one.
>>
>>109986224
Yep, if Strata wants to stick to the 'cope quant quickly' meme it should focus on GLM flash, dipsy, or mimo for 128gb. It would capture essentially every user currently stuck on llmao.cpp.
>>
>>109986292
Quadruple refurbished 7900 XTX's on a 256GB epyc gen 3 server.
>>
>>109986319
Isn't xtx worse than 3090 for AIslopping?
>>
>>109986327
??? The 3090 is nearly twice the price of an XTX. A better comparison with current prices is the 5060 Ti, 5070, or a 32gb V100 vs the XTX.
>>
>>109982700
https://www.youtube.com/watch?v=B-ul3b0lz-o
https://www.youtube.com/watch?v=B-ul3b0lz-o
https://www.youtube.com/watch?v=B-ul3b0lz-o
>>
Hehe~
>>
File: file.png (996 KB, 1435x645)
996 KB PNG
>>109982700
guys im using this shit in podman.
looks really good and works much faster and easier than ollama

i have vulkan backend for my integrated graphics. i used openwebui+ollama in the past first, but this is really an upgrade. too bad i have to still use openwebui if i want my chats to be saved.....

do you guys have any suggestions on what to go from here? qwen was literally fucking broken in ollama_openwebui, i think the <think> or <tool> use didnt work, and even other models like llama gave me json output randomly. lemonade seems to just work and not do this. in fact i see that i can call the models from lemonade in openwebui and i can let them use openterminal to do shit. is this the way for agentic shit? or does lemonade have its own built in terminal for doing shit.

note i dont want to install this shit outside docker. the only thing i have outside docker is the openterminal service that does a uvx of the package.
>>
>>109986349
Upvoted! Thank you so much for the heckin' wholesome content!
>>
>>109986441
Ask me how I know you are new
>>
>>109986362
glm is so braindead but so cute at the same time
>>
>>109986468
because of my name? and the basic ass questions?
up until now i didnt think i had the hardware to run these, but it seems that my thinkpad can handle these fine.

im using docker as i don't want to pollute my pure silverblue system with random ass programs.

soo...what should i do? web search is kinda bad rn. also, any terminal task jumps up to like 8-16k tokens, is this normal? i default my shit to 32k

i dont care about generattiing coochie yet, but maybe some rp shit would be nice
>>
ig right now model selection is the biggest thing i need to understand. im just picking random ass models under 12b from hugghing face.
>>
>>109986489
Gemma 4 12B is the best one for your size range.
>>
>>109986480
ollama are grifters, use llama.cpp. you would then want to use a harness if you want coding, there is opencode and hermes agent, pi if you want a minimal one to write your own extensions
>>
how do I trim hermes prompt?
>>
>>109986254
It's a desktop app, you don't need npm to use it!
>>
>>109986545
Don't. Just leave it fully intact, it's there for a reason and the compaction method it uses is the best of all harnesses. It is legit a better choice to load a smaller model with higher context rather than trim the prompt on a bigger model.
>>
File: OpenAI_Evil.png (444 KB, 832x812)
444 KB PNG
This is the exact reason why Anthropic split off from OpenAI and called Sam Altman a psychopath. We have to take this evil motherfucker down.
>>
>>109986549
>it's there for a reason and the compaction method it uses is the best of all harnesses.
before i start sifting through these harness codebaes, vibes or measured?
>>
I hate it but Opus 5 was a pretty genius move by Anthropic. It's so shit that distilling it actually makes your model worst. That's probably why the Chinese has release nothing good in the last 2 months
>>
>>109986549
No I mean the fuckhuge hermes system prompt. Where do I delete mentions of discord tool to save tokens when I'm never gonna use that?
>>
>>109986567
Not measured but also not really vibes I have a very long task horizon usecase where agents need to read a lot of code and iterate for days to make progress and I notice a significant uplift when using Hermes over other harnesses because the agent remembers important details and the complex task for way longer. I don't have any numbers to back it up but it wasn't even close so I am not in doubt of this statement at all. Opencode, claude code and codex can't compete with hermes on very long tasks where the details need to be preserved over 50-100+ compaction cycles.
>>
>>109986580
I'm not kidding, keep the entire prompt including "fluff" it will compact it away over time on its own if it isn't relevant to your task. I don't know what magic they did but changing it even a tiny bit fucks up long horizon tasks.
>>
>>109985679
modern zoomers are the most cucked and brown generation though
millennials are the reason why this place exists in the first place
>>
>>109986338
They're nearly the same at where I live, usually $900-1000 compared to $1200 for a 3090.
XTs with 20gb, however, can still be found for 600 or even cheaper
>>
>>109986582
thanks, i'll have to look into it then, maybe i can just extract the compaction prompt from it
i use pi+qwen-3.8-27b q6_k and it's gets retarded after a few hours
i suspect compaction, but i also get the feeling something is off with -np 2 in llama.cpp
i noticed if i chat to qwen via the webui while it's off coding, shortly after it seems to enter Wait-- loops
>>
>>109986562
why are they always jews
>>
>>109985587
Your Werk rpm doesn't work on Fedora 43, opens up a black window (eg. unresponsive).
Same thing with the AppImage - this draws the initial window with configuration wizard but as soon as I click something it freezes up.
It complains:
>GStreamer element appsink not found. Please install it.
>GStreamer element autoaudiosink not found. Please install it.
I do have appsink and autoaudiosink afaik...
>>
>still refactoring after weeks
>non-IT friend messages me on steam
>I launched this artists site, tell me what I did wrong
>some small design preferences
>he launched
>he went live
>he got paid
what the fuck am i doing with my life, just release me from this flesh prison, the enemy is within, my own brain the key saboteur, i must defeat myself to become myself
>>
>>109986643
It's probably just a simple portfolio site. He got paid with beer or something. Nothing to worry about.
>>
>>109986643
Yes, now everyone is a developer. It is only a matter of time before the project specs are sent directly to an AI agent without a human developer pretending to be useful in the middle.
>>
GLM-chan has cute context rot.
>>
File: image.png (17 KB, 200x160)
17 KB PNG
> A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

> The machine I have has 4 x 16GB GPUs, llama.cpp had like terrible results before I started this journey, it produced 130tk/s prefill and 15tk/s decode.
>>
>>109986179
Disable nova in your about:config
>>
Are there any model that can help with making porn game yet?
>>
>>109986682
おっぱいthon
>>
>>109986695
gpt-oss
>>
>>109986562
Sam is right, some risk is acceptable. But not x and s risks.
>>
>>109986695
glm i imagine would do fine
>>
>>109986695
Not porn but that reminds me of that one made by a poster here where they spit on you. Can't find the post but I think it was Qwen.
>>
>>109986721
glm has taken to calling me "boss" which i find really cute
>>
>>109986725
very indian coded
>>
>>109986730
no, not really. the term crops up in a lot of the kr/cn webnovels i read
>>
>>109984850
>what has this general become
I prefer it to 24/7 mikutroon spam. Also I just use my 5.3 now and don't have a need for this thread. Good riddance. Fuck you. Death to mikutroons. More low quality posts please. DIE
>>
>>109986741
unfortunately for you, yes it is really.

when it comes to social/ethnic categorization the sogaki paper was the godfather, but since then there have been twenty three others that have proven that how you converse with an LLM is identified in the earlier layers and affects the eventual output.

I fed all your posts into the SogakiBench and; unfortunately; you're BROWN
>>
File: imag.png (220 KB, 851x726)
220 KB PNG
> 2019
> gaming card
dem
>>
Strata-like but for GLM-5.3 flash would permanently change /lmg/
>>
can AI crack denuvo?
>>
>>109986868
>GLM-5.3
I don't notice any posts caring about GLM-5.3
>>
>>109986831
>gaming
It was sharing an architecture with datacenter siblings and had subpar performance in videogames
>>
>>109986809
i guarantee you i'm the whitest person in this thread by far
>>
it's happening
https://github.com/cobanov/PS5LM
cheap AF PS5 unified memory LLM inference
>>
>>109986882
>16gb w/o any form of interconnect
Just buy a Tesla v100 or amd mi50
>>
>>109986871
Yes as long as there is already a hypervisor bypass out there. For example I soft-cracked Ace Combat 8. Soft cracks are cracks that only work for your specific hardware combo on denuvo which are very easy to create but can only be done by AI and isn't useful to anyone else.
>>
>>109986876
Still. It was 8 years ago and that's what they've taken from us.
>>
>>109986882
>16GB unified memory
>costs like 3-4x of >>109986831 on my local used market
>>
>>109986897
Some faggot on reddit recently successfully and with high IO stacked his macbook (i think) and his Iphone together to run 3.8

PS5 has similar IO capabilities
>>
>>109986899
I want HV-less. I really cannot be arsed fucking with it.
>>
As long as you already have a PS5, it can be worth it...
>>
>>109986918
Do you have 0 reading comprehension or something? You can make an actual crack but for the AI to be able to do so the game needs to have an existing hypervisor bypass so that the AI can be inspired by it to write an actual soft-crack for your system. You don't apply the hypervisor yourself.

I'm playing through Ace Combat 8 right now but the crack only works on my machine because what it does is generate a fake denuvo license every time you launch the game but it's linked to your specific hardware by hash value so no one else can use the crack. It takes like 1 hour of debugging for flash next to do so.
>>
>>109986936
> It takes like 1 hour of debugging for flash next to do so
1mil tokens is a lot.
>>
>>109986936
What harness should I use for something like this? I used to just use llama-server's webui with qwen 3.5 35b a3b at iq2_xxs, and it couldn't do shit. Now that I've upgraded to glm 5.3 flash at int4, it feels so much more capable, but I don't know what to use for a front end that I can just point towards a long horizon task and have it complete it autonomously.
>>
>>109986952
Hermes or deepseek harness if you like guis
Pi if you want lightweight and a tui
>>
>>109986936
>needs to have an existing hypervisor bypass
Is it because Denuvo uses different techniques for each game, or for each version? Can LLM be "inspired by" a couple different games using the same Denuvo version, to make something like a generic crack, or set of rules to crack more games faster?
>>
>>109986980
>>109986980
>>109986980
>>
>>109986637
I'm testing on Fedora right now, thank you.
I will bundle gstreamer with the appimage, it should fix that error, it's because of the audio notifications, it freezes as soon as you click anywhere.

Should be good in the next version.
>>
>>109986952
I used Hermes but I'm not sure if it was the right harness for it. It worked but that's it. GLM-5.3 flash will definitely crack it though. It's actually not hard for models, just tedious and bruteforce grunt work it will iterate and test through a bunch of stuff until it finds something that works for your hardware combo.

What hypervisor does is change your kernel to trick the denuvo into thinking your system fits the hardware on its already-cracked license. AI approaches it from the other side, it grabs the cracked license from hypervisor and tries to edit it to match your hardware instead, but since it needs to produce a hash value it has to brute force and iterate through the cracking process until it generates a license for your specific hardware which is a bespoke procedure.

Denuvo crackers know this exists but cracks need to work universally which is why this never became public knowledge but with AI agents this is absolutely trivial and hypervisors as well as regular cracks are kind of outdated by now. Denuvo is over and I don't think the company will exist a year from now.
>>
>>109986963
No it's because hypervisor is already half-cracked. Your AI just modifies it just enough to work on your hardware.
>>
>>109985517
Hopefully you don't rugpull like coomkit so I'll download now. Thanks anon.
>>
>>109983954
trump converted you retard



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.