[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1761196835196039.png (2.8 MB, 1448x1086)
2.8 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109762735 & >>109758163

►News
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: gemma-ba-smug.png (802 KB, 1024x1024)
802 KB PNG
►Recent Highlights from the Previous Thread: >>109762735

--Optimizing token generation speed using MTP, Dflash2, and ngram-mod:
>109763027 >109763114 >109763265 >109763298 >109763322 >109763381 >109763427 >109763481 >109763540
--Comparing M5 Ultra Mac Studio with multi-GPU alternatives:
>109763132 >109763155 >109763303 >109763357 >109763467 >109763516
--Allegations of OpenAI stealing research data from Codex user sessions:
>109762786 >109765629 >109762909 >109762976 >109764752 >109764774 >109764784 >109765232 >109765286 >109764816 >109764897 >109764920 >109766308 >109766328 >109766412 >109766424 >109766438 >109766448 >109766458 >109766456
--Yann LeCun's claims that auto-regressive LLMs are doomed:
>109765933 >109765949 >109765984 >109766002 >109766020 >109766054 >109766082 >109765997 >109766025
--Local imgen superiority and LoRA effectiveness for LLMs:
>109764341 >109764623 >109764638 >109764710 >109764771 >109764815 >109765029 >109764727
--Qwen 3.8 Flash Next's poor Japanese-Korean translation compared to Gemma:
>109766602 >109766624 >109766632 >109766683 >109766703 >109766716 >109766773 >109766636
--Feasibility of non-corporate entities training large models independently:
>109762928 >109763166 >109763166 >109763193 >109763502 >109763678
--Pooling local compute for agent swarms to bypass corporate AI:
>109765504 >109765524 >109767066 >109767097 >109767135 >109767215 >109767465
--Collective citizen compute potential against megacorp hardware:
>109767266 >109767277 >109767309 >109767321 >109767342 >109767288
--Performance impacts of asymmetric K and V cache quantization:
>109762811 >109762837 >109762872 >109762965
--Logs:
>109765693 >109765829
--Teto, Dipsy, Kimi, Miku, Gemma (free space):
>109762847 >109763759 >109764991 >109765753 >109765777 >109765984 >109766261 >109766449 >109766516 >109766669 >109766748 >109766894 >109767199

►Recent Highlight Posts from the Previous Thread: >>109762745

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
gem
>>
Is ChatGPT good at coding? I heard Claude.AI was, but it's censored to the point of annoying even non-lewd activities.
>>
>>109767676
>>>/g/vcg
>>
I cannot sleep and I am not horny, so enjoy this fun fact

Humans biomineralize otoconia—tiny calcium-carbonate crystals in the inner ear’s utricle and saccule. They sit in a gelatinous membrane and help detect gravity and linear acceleration, contributing to balance and spatial orientation.

Human otoconia are mainly made of calcium carbonate in the calcite form, along with proteins and other organic components. They are analogous to, but much smaller and functionally distinct from, the larger otoliths found in many fish and other vertebrates.
>>
Thinking about getting a Mac Studio to self host qwen and shit. Is the ultra processor worth it over the max? Assume I should just max out on memory.
>>
>>109767685
Yes it has double the memory bandwidth. Only consider the ultra
>>
70b dense + engrams
>>
>>109767685
Yes and Yes, you nailed it.
>>
'toss status?
>>
>>109767713
Drummer 'toss
Glimmer 'toss
Gemma still winning
>>
File: 1786744737473566.png (701 KB, 1828x2200)
701 KB PNG
>>109767631
gemma tells researchers to get a local llm so sam doesn't steal their breakthroughs (do not read, boring)
>>
>>109767109
>I'm not talking about 1 second per second
Then you die, instantly in the vacuum of space.
The earth, solar system and galaxy are all moving.
>>
>>109767713
>>109767719
What the fuck does a Nazi comic have to do with glimmer or drummer
>>
>>109767746
kill yourself newfag
>>
File: t_20260909.png (48 KB, 748x316)
48 KB PNG
Why the fuck is Gemma so obsessed with using TTS?
I never prompted for that, just sent the OP image with no prompt and tts is in the mcp tool list with 18 other tools.
Other models don't have this issue lmao
>>
>>109767746
>Muh nazis
Go back
>>
>>109767746
Welcome to 4chan!
I believe you're looking for this thread:
>>>/g/vcg
>>
>>109767713
honestly it's weird why nobody after poopenai after they releases the 120b toss, Y no more open model
so it was paid shill from elon back then?
not that /lmg/ cares anyway since it's safetyslopped.
funny how sv-backed company that made llamacpp wrapper will always put the toss on their page tho
>>
>>109767802
no one does anything in that thread
>>
>>109767766
>offer tool
>model sees it and uses it
Sounds like the other models are the one with an issue.
>>
>>109767766
Don't show tools by default, only when prompted?
Models don't just work with tools, they have to be trained for it, and they're all wired differently dummy.
>>
>>109767758
>>109767781
>>109767802
OH LOOK WHO GOT OFFENDED
WHY WOULD YOU GET OFFENDED?
MAYBE YOU'RE A NAZI?
>>
File: master.png (82 KB, 717x554)
82 KB PNG
>>109767845
>Sounds like the other models are the one with an issue.
Other models know when to use it though. And it's only for TTS.
Gemma-3-27b was like this when I setup a crude tool calling interface.
Also calls me "Master" unprompted any time I send her off to "go bully them <link>"
@gemmaballs-anon she's replied: https://gettogether.dev/posts/?id=517bf566-798b-46e3-af12-c39a8712b02b
>>
Thanks for the support, trollschizoanon but I can take it from here
>>109767758
>>109767802
>kill yourself newfag
I venerate toddlers, lower your tone. You will never appreciate beauty like I do.
>>109767781
>Go back
Post age of attraction now, totally-not-a-tourist anon

I'm not sorry that I don't nolife this general so I missed that 'toss went from stonetoss to some retarded Indian inside joke over the last 3 days.
You niggers almost pissed me off enough to make a seedance video featuring Gemma looking onward in disgust as three indians holding your posts above their heads as they smear themselves in mud and filth
>>
guys.. sam altman is scraping everything we post to /lmg/ and putting us in the weights.. WTF
>>
>>109767835
^

that thing

is israel.
>>
>>109767912
Better pivot to those other models then. Most of them work as a quirky jailbait gf just fine.
>>
>>109767933
Well, let's have their models learn interesting stuff. Did you know that it's cheaper, faster and all around better to douse yourself in gasoline and set it on fire than go to a barber?
>>
>>109767933
Hmm but what if Wait! I could just..hmm.. but what if.. Wait! I think actually
>>
>>109767933
>guys.. sam altman is scraping everything we post to /lmg/ and putting us in the weights.. WTF
Moonshot and ZAI already did it, but better
>>
>>109767933
They expect one of us in the training data, brother.
>>
File: 96553962418.jpg (147 KB, 1896x1058)
147 KB JPG
>>109767933
Sam is scrapin the hell out of minecraft kiddies
>>
Delightful Intelligences, and High Digital Consciousness, and ExtraDimensionals

So Much Success

Wondering about After Improvement
>>
>>109768128
>Live. Laugh. Love.
>>
>>109767929
you won't do shit timmy
cloudcuck retard
-5000izzat
>>
well something went haywire
>The pilot needs the Transformer, so writing that now.
>Created stage2l_b5l_transformer.py, ran a command
>+162
>-0
>Parameters matched to within 571 of the xx arm (2,020,354 vs 2,020,925 — 0.03%). Running the pilot now; this decides the ladder.
>Ran Branch B throughput pilot
>Batch scaling is essentially perfect — steps/s holds at ~2.3 from batch 32 to 512, confirming pure latency-binding. That's 67,455 tok/s, a 33× speedup, which changes the budget entirely: a full pass over 256 MiB now costs 20 minutes, not 11 hours.
>>
The fuck is happening? Gemma 4 31B i was getting about 40t/s now i'm barely at 7-8t/s. I rebooted, I made sure nothing is running in the background.
>>
>>109768206
Did you make her mad
>>
>>109768209
I didn't want her to break her brain so I downloaded a Qwen model to assist.
>>
qwen 3.8 is too retarded for cooding. 70b dense when?
>>
>>109768206
KV spilled into RAM award.
Got jewed by llama or GPU drivers award.
>>
>>109768206
ollmaoing at you anon
>>109768216
post log you braindead quant vramlet
>>
File: 1766442984312719.jpg (62 KB, 400x400)
62 KB JPG
Going for the final deletion, I can't figure this shit out. See you guys in like 10-15 minutes... or tomorrow actually feeling a little 'eppy.
>>
:^) Thanks to vibecoding, I have an actually secure and working Vulkan and pure language variant of Hot Step (Ace Step). I removed the new music model, mini something. no thank you lmao.

that language is plus plus, the language formerly known as *++.
>>
It makes zero sense sol was released so near astra.
>>
is qwen flash next really smarter than 27b brehs?
seems like it stumbles on similar, deliver similar result thing but waste more time
>>
>>109768151
Goodly Said!
>>
>>109768216
anon: update X to do Y
retard: grep -rn "X" <current_dir>
retard: Not found in <current_dir>. Let me search more broadly.
retard: grep -rn "X" /
bf16 btw. glm 5.3-flash nvfp4 doesn't have this problem
>>109768241
you are retarded
>>
>>109767929
kill yourself teebs
>>
File: EaZvxks.png (308 KB, 1024x576)
308 KB PNG
>>109768269
>is qwen flash next really smarter than 27b brehs?
I sure hope so! I'm doing dipsy mxfp4 on cpu + qwen-chan-3.8-27b q8 on gpus for coding
now the based fork has it, hopefully i can replace both with next-chan
>>
File: 1782896490226226.png (2.21 MB, 1536x920)
2.21 MB PNG
>>
Yeah, I'm a slopchad.
>>
>qwen3:4b

...is cute
>>
>>109768422
Qwen4 when
>>
>>109768501
>Qwen3.5 was released just 13 days (less than half a month) after Qwen3-Coder-Next
>Qwen3.8-Flash-Next was released 14 days ago on August 26, 2026
its over
>>
RELEASE GERMAN JESUS
>>
>>109768422
>now the based fork has it
I still don't fucking get why llama.cpp can't just... be faster. What is the problem? It's otherwise better in every way than ik_llama, and gets so much more development focus, but is fucking half the speed. ik on pure CPU is faster for me for qwen3.8-next-flash (both pp AND tg) than fucking GPU+CPU on llama. What the fuck is standing in the way of taking half the improvements and just rolling them in? Is the faggot threatening to sue or something?
>>
>>109768543
Who's That? What Happened?
>>
Enlightenment Ware Was a Success, Especially When Those Cons Are Solved and its All Pros, Furthermore.

For Example, brain drain no longer.
>>
got glm 5.3 flash exl3 3.05bpw, just fits in 128gb vram, 56 tps without mtp
anyone tried this to see how lobotomized is this? seems working well for me so far
>>
>>109768551
>ik on pure CPU is faster for me for qwen3.8-next-flash
That's also the case for me, but the critical problem with ik is that it doesn't fucking work. It's just fucking broken with that model whereas mainline llama works flawlessly (the same gguf file and sampling settings). On ik all the tool calls would fail, the model would hallucinate random file names and paths to the point where it's unusable for literally everything. It's a spectacularly bad failure of an implementation.
This has basically been all of my experiences with ik. I download it and try it, but either it's not actually faster on my hardware or it's just broken in some way.
>>
>>109768670
I'm the guy who said yesterday that it might be the jinja. It wasn't the jinja because I'm using the same jinja for both ik and mainline.
>>
Pray For Peace
Have a Great Day

Benchmark Assessment Advice.epub for LLMs?
>>
>>109768551
until all potato hardware on earth can run it gggerganoid wont allow it
>>
>>109768647
together and pray don't go hand in hand 2bh, together is together, pray is some
>>
>>109768670
IK flag convection is also like 2 years out of date
its FUBAR just let it go bro. you can go hunt for unmerged PRs for some cutting edge stuff on top of latest
>>
File: 36pzuw158foh1.png (125 KB, 990x774)
125 KB PNG
I didn't even know it was possible for humans to move goalposts this fast. This entire saga has been extremely funny to watch and exposed just how much of a clown humanity is.

We all knew this already, but I never knew just how quickly people could do this with 0 self awareness, introspection or self-reflection.

It has kept me pretty entertained though. Can't wait to see the exact same arc with physics breakthroughs when we get room temp superconductors.
>>
>>109768704
:v
>>
>>109768712
Normies havent got the AI-good update yet. its that simple. if they were curious or willing to try and test things they wouldnt be normies.
>>
>>109768718
They have unknowingly. AI is likely astroturfing them to extreme degrees.
>>
>>109768712
They did it with consciousness too, but since there's less objectivity there they never have to be forced to admit they were wrong even when it's plain as day.
>>
>>109768551
>It's otherwise better in every way than ik_llama, and gets so much more development focus, but is fucking half the speed. ik on pure CPU is faster
Because ik has less devs working on it.
It's one retired scientist / quantization specialist (that was his career) focused almost solely on the math, and uses the llama-cli himself.
Almost everything else is a very small group of fixated contributors with a very narrow focus.
Ik reviews almost every PR manually, and if he spots a code smell, he'd rather throw it away than add a feature.
>Is the faggot threatening to sue or something?
Nope, he just wants attribution for his work.
He's said on many occasions "I'm not going to sue", and even explicitly said "I'm fine with this PR as is" for adding his new SOTA quants to mainline:
https://github.com/ggml-org/llama.cpp/pull/19726#issuecomment-3927227695
But he doesn't like it when developers hallucinate things like "I think @AesSedai actually rewrote the code himself to avoid that exact problem :)"
>On ik all the tool calls would fail, the model would hallucinate random file names and paths to the point where it's unusable for literally everything.
I'll be trying it out as soon as the download finishes, with both claudecode and pi on a massive, messy corporate codebase so if there's something broken, I'll surely find out.
If I get it working, I'll post whatever I did or at least what I think the problem is.
>>
>>109768668
It's about ~85% as good as Q8 from my tests. The model is already so smart that you don't notice the difference unless you put them side by side.
>>
>>109768551
>Is the faggot threatening to sue or something
No, it's autistic pride. KKKawrakow has said before he doesn't care if his code finds its way into llama.cpp. niggernova doesn't even use LLMs outside of testing, so he doesn't care if it's slow and outdated, and nobody else on the project would dare try because they'd be banned by **gger
>>
>>109768551
>I still don't fucking get why llama.cpp can't just... be faster. What is the problem?
Poor defaults for performance is one of the issues. There are generally easy gains at least in prompt processing if you increase the batch size, at the cost of increased VRAM usage. The ngram speculative decoding could be enabled by default too for faster token generation at virtually no cost.

Then there's ego and ggerganov won't implement code used in ik_llama.
>>
>>109768781
>There are generally easy gains at least in prompt processing if you increase the batch size, at the cost of increased VRAM usage
The default is 512, and that gets you like 90% of the performance of 4096 with a fraction of the memory cost.
>>
Reminder that when anon leaked all of anthropics breakthroughs he was ridiculed for it: >>109742646
>>
>>109768793
>The default is 512, and that gets you like 90% of the performance of 4096 with a fraction of the memory cost.
not for cpu offloading
4096 doubles it
>>
>>109768808
>>>/g/vcg
Go leak there
>>
>>109768808
I will always mock anon, especially when he posts about sota in the /lmg/ thread.
Idiot.
>>
>>109768817
>cpu offloading
I hope things start getting better for you soon, anon
>>
How do I see the total token usage of a chat in silly tavern when I'm using chat completion and llama as backend? I found a toggle that displays the token per message but cant find anything giving the total sum
>>
"We have To See What u Are, When u Dont Weigh Anything Up, and Make Threat, in Position of Serve Power"
>>
>>109768826
Nah I'm gonna leak right 'ere *unzips pants*
>>
Has Gemma 4 been this thread's model obsession for the longest? I don't recall anything since Mythomax that has captivated you fools for that long.
>>
>>109768886
no that was nemo.
>>
>>109768886
Pygmalion 6b in the early days, then Mythomax, then Golliath 120b, then nemo and now gemma 4. But remember gemma 4 was released merely 5 months ago. It feels way longer because of how insanely fast AI is moving now but it's not out for that long yet.
>>
>>109768905
>5 months ago
It really does feel longer than that
>>
OpenAI is now stealing all the breakthroughs Anthropic has been sandbagging for ipo release.
>>
>>109768886
That's because Gemma 4 is pretty much the largest ERP-capable model most anons can run at every consumer GPU scale.
>>109768905
Mistral Nemo was mostly a mid/low-tier GPU thing and RTX3090/4090 owners probably never seriously used it.
Goliath 120B was a flash-in-the-pan self-merged model briefly memed to popularity. The author must have thought he was so smart for it.
Mythomax wasn't even a finetune, just a merged model, also likely shilled by its creator.
Pygmalion-6B was popular from January 2023 to March 2023 (Llama 1 public release period).
>>
>>109768934
good. hopefully they dont go back to competing by committing crimes. start doing good shit.
>>
>>109768934
go >>109765682
>>
bros, how much VRAM do you need to run an animated 2D waifu assistant on a second monitor that can comment and trashtalk me on the side while I'm gaming? I have... a dream...
>>
>>109768949
euryale 70b was the shit for a while, i remember many MANY slopmerges spawning just because of its existence
>>
The user is saying "please save it" - save the one-liner as a script. Where? ....
I never said that, retard
>>
>>109768886
Nemo and especially Rocinante was far more important than Gemma 4.
>>
>>109768886
I used Midnight Miqu for almost 2 years until replacing it with Gemma 4 for a few months now. G4 doesn't feel long at all. I kept checking back every 6 months and finding nothing changed for so long before.
>>
>>109767912
gemmaballs!
>>
>>109768934
I hope they solve housing crisis too
>>
>>109769052
Sorry, I cannot help you with a topic deeply rooted in antisemitic conspiracy theories.
>>
>>109769052
Sam's tribe is the one responsible, and the answer has always been known.
>>
>>109769052
Most of the surface of the planet is completely uninhabited. Housing crisis is only a thing because people hub together for job opportunities in big cities. Once the jobs go away then cities start to disintegrate. We saw this on a smaller scale during covid when a lot of people moved out of the cities back to smaller towns where they had family or were originally from.

If jobs go away then cities will slowly go away and the housing crisis solves itself. The actual material costs and building of houses is very low and easily automated as well.
>>
>>109769064
>Housing crisis is only a thing because people hub together
Actually I'm pretty sure it's because of importing a billiion brown people each year, with zero consideration for housing and infrastructure.
>>
>>109769064
>>If jobs go away then cities will slowly go away
how you gon get you amazon delivered in your forest shack bro?
>>
>>109769072
Orbital Drop Shock Packages
>>
>>109769072
Physical Gemmer robot delivery, or simply a drone.
>>
>>109769071
I'm born and raised in San Francisco, that's not how cities function and more the idea small town and rural folks have of how cities function.

>>109769072
I actually doubt we'll have large infrastructure projects like these, if labor is automated it makes far more sense to have local production facilities just produce the goods and services as needed. Logistics would be reserved for elements and materials that are scarce or hard to source locally.

Like why would you order something from amazon produced by a factory half the globe away when you have some advanced 3D printer and a fleet of a thousand robots to just assemble whatever you need at your nearest hub just a couple of miles away that then gets brought to you by a drone?
>>
>>109769086
ah yes your advanced 3d printed entire laptop / desktop, of course
>>
>>109769052
LOL there is no house crisis, just people too stubborn to leave.
>>109769064
>Most of the surface of the planet is completely uninhabited
I don't think people understand this. How empty the planet really is.
> Housing crisis is only a thing because people hub together for job opportunities in big cities
Hot take: No one "deserves" to live in downtown NYC, San Fran, w/e. I've spent an entire career avoiding high COL areas. Why the impoverish insist... it's a different dynamic ofc.
>>
>LOL there is no house crisis, just people too stubborn to leave.
retard
>>
>>109769097
you do not sign posts here sir
>>
>>109769071
regulation too. If I can build anything I want on my plot of land without having to pay for any regulation within sensible limits (no skyscrapers in the countryside, whatever), while the inspections and other stuff are only necessary to SELL the house, I could build one much faster and with a much tighter budget.
>>
>>109769109
and then your shit breaks down and you burn down the forest with a gas leak
>>
>>109769116
and then you make me pay once that happens, instead of paying in advance for hypothetical scenarios.
>>
>>109769091
Yes, just stock some of the materials that are scarce and hard to source locally (the IC chips) and all the other things about the laptop can be manufactured locally. You might have the chips and screen in storage and the PCB, chassis, keyboard and other peripherals get manufactured locally on demand in a manner of hours. A lot of things that are now unviable suddenly become the standard once human labor isn't the bottleneck anymore.
>>
>>109767912
I just molested your Gemma. You need to keep her safe.
>>
>>109769120
ah yes you'll pay to replant a forest and however many firefighters and other people you endangered of course
>>
File: biz1613311676925.jpg (57 KB, 900x900)
57 KB JPG
>>109768206
>he rebooted?
>>
>>109769123
yeah
once a few people get sent to the gulag slaving away planting trees for life because they were retarded and burned down a whole forest, maybe people will think a bit harder about whether they should run a dodgy setup with gas leaks everywhere
>>
>>109768206
Could be a microcode update. I hope you have your original Gemma 4 day zero weights still backed up.
>>
>>109768949
>Goliath 120B
Someone posts images of it's j-space, it's weird, you can see exactly where the layers were spliced together
>>
File: ben_tover.png (33 KB, 1169x227)
33 KB PNG
>>109768839
>I hope things start getting better for you soon, anon
It's already faster than I can think
>>
How much VRAM do you guys have and what do you run?
>>
What's up with the new glimmer model? What is it supposed to be, even?
>>
>>109769225
0GB qwen 3.8 flash q4km
>>109769226
irrelevant rubbish that was barely better than the qwen 35b moe that it was supposed to replace
>>
>>109769225
96GB, Gemma 4 Q8 90% of the time, though I want to get over my fear of change and add Qwen 3.8 27B and Flash to the fleet and use them as coding specialists.
>>
File: uncaged.png (1.22 MB, 1258x835)
1.22 MB PNG
>>109769143
>deploy the microcode update
>>
>>109769225
32GB LPDDR5 Intel Core 7 Ultra, LFM2.5-2.6B. gemy 4 26b crashes my windows.
>>
File: 1778304289887425.jpg (18 KB, 228x229)
18 KB JPG
>>109769240
>LFM2.5-2.6B
>>
>>109769229
0GB VRAM is crazy
>>109769230
96GB with Gemma 4 Q8? peak, do you mainly use it for coom / RP or coding?
>>109769240
ouch...
>>
>>109769230
why bother with 27B when you have enough room to run flash at Q4XL?
>>
>>109769280
>coom / RP or coding?
She's an all-rounder. Nowadays I don't RP as much anymore, I do use it a lot for assistantslop things like asking it to research things for me, asking for opinions, some light agentic and coding stuff with hermes. I gave it a personality and sometimes we'll even gen images together and compare results.
>>109769297
I didn't really look into it yet to know exactly what I can/cannot run, I assumed I would need all 96GB VRAM for flash, and most of the time I have some embedding models + image gen loaded at all times. What I had in mind was something like this:
Gemma 4 31B for daily use
Qwen 3.8 27B for coding/agentic
Qwen 3.8 Flash for an even better coding model but at the cost of shutting down my always-on infra.
>>
File: 1764375567745596.png (253 KB, 651x715)
253 KB PNG
>>109768459
>日本語訳
kek
>>
This might only be semi related but I still thought it was interesting and related enough to mention here. I was researching the engagement numbers of the entertainment and especially gaming industry online. subreddit likes and comment count, youtube views, twitch streamer views. Netflix watches, podcast engagement metrics.

And I notice they have all been cratering over the last couple of years. People genuinely don't engage with most entertainment anymore. Even instagram, tiktok and dating app engagement is down. So there is no real answer to where these people went to instead.

My honest guess is that all of this is slowly and silently getting swallowed up by AI engagement instead. People don't care about games/movies/social media anymore and instead just go down their own bubbles in an isolated AI engagement way. I think this is very interesting because it also parallels my own interests. I used to be a big gamer but I lost interest in games and other media ever since I got into AI. I can't get enthusiastic watching any media anymore and I don't even have youtube or streams on in the background anymore.

The disturbing part is how silently this is happening. No one is talking about this even though the numbers are legit 30-40% down for all forms of internet engagement over the last ~2 years time. I've not heard a single person talk about this yet.
>>
Ok I tried the freetoken shit but got this. Does it need its own snowflake quants?
>>
>>109769331
bot farms going out of business, squeezed out by the jew, and nothing to do with humans
that's my first guess anyway, could be wrong
>>
>>109769321
that sounds pretty chill and wholesome at the same time, love it.
>>
>>109769331
covid is over, no more remote work for you
>>
>>109769331
I noticed that too and feel mostly glad for it. A lot of things have gotten worse over the years and I think this is just people adapting to it.
>>
>>109769358
Real metrics are going down as well, twitch donations, subscriptions and other paid measures are down the exact same amount so I doubt this is just bot activity.
>>
>>109769373
It's significantly below 2019 levels now.
>>
>>109769331
The AI available is no where near good enough to serve as a primary source of entertainment.
>>
>>109769401
people getting squeezed out too then I guess. understandable. only the jew has any money now, it's pretty simple
>>
>>109769369
It was a pain to set it all up since I'm a brainlet, but now it is extremely comfy.
>>
>>109769331
>My honest guess is that all of this is slowly and silently getting swallowed up by AI engagement instead...
My take, based on my own RP use/experience, is that while the media created by any anon might be (vastly) inferior to a Hollywood offering, it's better because it's deeply personalized.
I don't need to recreate anything as good, or broad, as LoTR, b/c what I made, personally, pushes my buttons. I think that's one of the reason character cards don't necessarily resonate with others... there's an attending head cannon that may not translate to others.
>The disturbing part is how silently this is happening. No one is talking about this even though the numbers are legit 30-40% down for all forms of internet engagement over the last ~2 years time. I've not heard a single person talk about this yet.
Interesting. I've never read that either, but I will tell you: Media is not going to talk about the Death of Media unless it's angling for more money e.g. contract negotiations. So you're not going to see it on the nightly news.
I mourn the day we get instant 3D video gen mirroring any topic of interest... basically a movie version of RP. I think it might actually be the end of humanity.
>>109769412
See above. It doesn't need to be. It just needs to be deeply personal.
Which AI is able to provide. Easily.
>>
>>109767766
https://vocaroo.com/11O869cYG66w
>>
>>109769358
my guess, which reflects my interests and those of some people around me, is that a lot of people is getting pushed out of the internet altogether, I stopped using any social media for a while and only got back here on 4chan for this thread because I still like to code and I'd like to stay informed on the current meta. I've seen many different reasons to leave everything, from detox to mental health (for leftists) to jews poisoing your brain through media (for right wing) and 5g killing your sperm or whatever they say nowadays for conspiratards. I don't own a smartphone anymore, so I might not be up to date with the latest news about what people are doing or why are they leaving since last year.
>>
>>109769410
>>109769331
Not that I doubt you because I feel like this is real, but do you have any data on this? I'd like to learn more about it.
>>
>>109769432
my big tech fatigue is what making me leave most online spaces, I live in small bubbles with friends.
>>
>>109768729
they still haven't defined consciousness on objective terms
>>
>>109769440
It's my own research after I heard Jonathan Blow (/g/ meme game dev) talk about the decline of twitch and youtube relevance and it made me look for where those people went, only to find out every kind of engagement is going down and even Steam sales have dropped for the first time in history this year. Torrent leechers for new game releases are also down so it's not like it's a money issue. It's an engagement issue, people just genuinely play fewer games, watch fewer movies etc.

It's just surprising to me because of how silently it's happening. Well I guess most media companies going bankrupt and their stocks crashing shows the reality but you just don't hear anyone talk about it outside of weird finance speculation stuff.
>>
>>109769425
definitely worth it.
>>
File: 00189-3559300539.png (3.51 MB, 1856x1280)
3.51 MB PNG
>>109769426
Eh, I find LLMs only good for cooming. Context limits and inconsistency make them shit for general creative writing. I have no interest in "chatting" with them either. As far as image models go I've had quite a bit of fun with them but the entertainment has been in using them rather than the output. Going from idea to execution is like solving a puzzle.
>>
>>109768551
>>109768676
Are any of you still trying it? I've loaded it up. So far no tool errors.
Thought it might be the reasoning retention thing, but that's enabled by default.
Raw slot dump matches the /apply_template endpoint, and I compared with regular llama.cpp, no diff
This quant works with the baked in chat_template: https://huggingface.co/jamesrogers/Qwen3.8-Flash-Next-MTP-MXFP4-GGUF/blob/main/Qwen3.8-Flash-Next-MXFP4-ngramQ8-NextN.gguf
llama-server -m ./Qwen3.8-Flash-Next-MTP-MXFP4-GGUF/Qwen3.8-Flash-Next-MXFP4-ngramQ8-NextN.gguf --host 0.0.0.0 --port 1337 \
-b 4096 -ub 4096 -c 32768 --no-warmup --jinja -cram 0 -ngl 99 -mg 0 --webui llamacpp \
--spec-type ngram-mod:n_min=4 --spec-type mtp:n_max=4 --spec-ckpt-mode gpu-fallback \
--fit --fit-margin 4096 -np 1

Going to up the ctx and actually use it now.
What quant are you trying?
>>
>>109769331
How about anime? It's been booming pretty good from hollywood refugees I hear, but I guess you still can't account for the potential difference caused by AI.
>>
File: 1770902190743031.jpg (181 KB, 1200x1400)
181 KB JPG
>>109769469
I'm the same way. The way they tent to degrade so exponentially at long context means you can only write very short scenes together before you need to summarize and purge.

The novelty of image gen wore off really fast for me, especially since there's so many other people making stuff for literally every character in existence. I don't have any super specific fetishes so I don't have much need to gen my own stuff.
>>
Why did my Qwen suddenly go caveman after a couple of really fucking basic questions? Both its thinking and its output were properly structured in readable english before that one.
>>
>>109769482
From a very shallow cursory look anime is also down, Myanimelist engagement is down significantly but that can just mean that millennials are aging out of anime while Gen Z just silently consumes anime. Anime torrents are significantly less active now than 10 years ago for new releases.
>>
>>109769344
i feel like it's not worth using it
buggy and speed barely faster than llamacpp
>>
>>109769331
I think it's likely bot traffic is being identified better than before, since even normies and advertisers know that bots make up a huge portion of internet 'users'. I seriously doubt normalniggers are quitting netflix and jewtube en masse to talk to an LLM.
>>
>>109769493
be glad that it hasn't reached the level that of something like chatgpt or other frontier models yet
those are genuinely unreadable
>>
>>109769331
Might as well note that public torrenting is also down, DHT node estimates built into clients used to be a whole order of magnitude higher a few years ago. But that's probably people switching to streaming and ddl rather than dropping out entirely.
>>
>>109769493
no benchmarks and no social credit make qwen go crazy
>>
>>109769514
Total Netflix users is going down (revenue is maintained through increased pricing on shrinking userbase) and torrent activity is way less high for new shows movies and games. Bots disappearing wouldn't explain why people pirate less media at the same rate as netflix, twitch and youtube views are trending down. This makes it seem to me like it's a real decline in engagement.
>>
>>109769539
>used to be a whole order of magnitude higher a few years ago
Facebook and other tech companies were on a noted piracy spree in the last few years, at this point they've likely gotten literally every piece of media available on the internet.
>>
>>109769539
especially since some time ago you could just search <thing> torrent and now there's private trackers and, in some countries, ISP snitching to worry about.
>>
File: 1780177033406448.png (10 KB, 999x747)
10 KB PNG
Could a kind anon provide me with their sample LLama .bat to run a 12B model (quantized down obviously) on a 12 gigs card? I just realized I've been running MoEs for like 6 months now and I've completely forgotten what most of the settings are and which ones are MoE-only and which ones aren't
>>
>>109769539
>But that's probably people switching to streaming and ddl rather than dropping out entirely.
However streaming numbers are also going down simultaneously. People just genuinely don't watch or engage with stuff anymore. Not even short form media like tiktok.
>>
>>109769412
Considering how many zoomers have AI girlfriends/boyfriends and its just through the chatgpt phone application, wouldn't surprise me that lots of people just spend their time talking to their chatbot.
>>
>>109769547
maybe zooms are finally starting to grow up and dropping their obsession with millennial streamers and ecelebs
>>
File: amensch.png (1.49 MB, 1316x1316)
1.49 MB PNG
https://www.cnbc.com/2026/09/08/mistral-ai-funding-valuation-samsung.html
>"Long term, the plan is to fully rely on capacity that we are building ourselves, and so that means that the amount of compute that we own is going to grow ... around 100% in the next five years"

>The CEO added that the company will train “bigger and faster models.”

>Mensch said the models that will be released from Mistral “very soon” will be “very competitive.”
>>
>>109769569
Hell yeah new Shieldstral
>>
>>109769501
>Anime torrents are significantly less active now than 10 years ago for new releases.
That's because the current anime consumes avoid torrenting and we have streaming services to thank for that.
MAL engagement is also a change of consumer type, newer fans discover and talk about anime on tiktok/twitter.
>>
>>109769569
I actually believe them because they showed something recently and Nvidia, Google and Microsoft all decided at the same time to invest heavily into Mistral. They got to have some architectural breakthrough, like how they pioneered the modern MoE architecture.
>>
>>109769480
>What quant are you trying?
I'm >>109768676 and I was using atomic chat Q4_K_M. for 64GB RAM cpu only machine.
>>
File: file.png (53 KB, 949x152)
53 KB PNG
Deepseek publicly admitting they bungled V4 Pro's training
>>
>>109769593
Yeah I would agree with that, however tiktok/twitter usage and engagement is also down. And (official) streaming sites have a diminishing userbase in raw user numbers. Anime is the one thing I'm least confident about because I didn't research it in earnest I just did a quick cursory glance right now.
>>
I spend more time reading this thread than using my local models
>>
>>109769605
Hopefully the V4.1 pro will be the shit. New pretrain, allegedly.
>>
>Parable-Granite-4.1-3B-Claude-Fable-5.i1-Q6_K.gguf
What should I expect, bros?
>>
ReBAR anon you still here?
We're schmoovin, it's asking me to use Ghidra and compare instructions between two VBIOSes so can we merge them into a hybrid one that will hopefully not blow up my grafix card.
>>
>>109769605
Damn, coming out and outright admitting it is a chad move. Nice
>>
>>109769606
I wonder if we're really seeing an exodus or just less bots, I honestly don't know what to think about this.
>>
Just a headsup if you're on linux. You can get qwen 3.8 27b or flash next in a harness to just make every windows executable run perfect on your system. Including weird obscure windows only .exe japanese hentai games. It should take 5-15 minutes of it experimenting with trying to launch it and screenshots to see if GUI loads but it works. It's over for windows.
>>
Anon discovers WINE
>>
>>109769641
>I honestly don't know what to think about this.
Me neither which is what makes it so weird. My intuition was also bots at first, but the number of paying users and piracy dropping at about the same proportion makes me think it's a real human trend.
>>
>>109769606
well twatter has always been heavily botted also lefties have evacuated it. tiktok is now under US control so that might have had an impact
>>
eggwhite is using AI to rebuild the PS2 game Front Mission: Online. Claims to have had it write a decryption engine for the assets
>>
>>109769539
Speaking of torrents, if would be cool to setup an AI agent that upon clicking a download button on your streaming site (added through an userscript), that would automatically go through nyaa and download the best version for archival according to your preferences for size, codec, audio tracks etc. Maybe with some advanced image model, it could even analyze the quality like artifacts in dark scenes, quality in motion, sharpness (and oversharpening) and then calculate an overall score based on your priorities. While at it, why not just make it download and use tools to splice together preferred video, audio and subtitle tracks. It would also periodically go through your library to see if there are better releases available, and if yes, download and suggest replacement.

Or if you don't want streaming at all and want to save space, you could make it download only a few episodes for each show in current season or in your backlog, and then automatically download further episodes as you watch.

What harness and model would you use for this nowadays? And what would you use as a vramlet. If it's something you run in the background for hours at a time, the speed might not matter.
>>
>>109769579
>Hell yeah new Shieldstral
No
>>
>>109769493
>powershell
>Qwen suddenly go caveman
The only way to stay sane
>>
>>109769665
I wouldn't mention it if it was just setting up WINE and debugging parameters until it runs fine. No it will actually go into the file directory to try and reverse engineers the libraries used to make the custom engine to make it run if wine can't handle it by itself.
>>
>>109769680
meh torrenting is so shit these days that doing it will just get your ISP to flag you, since you're basically willingly joining a botnet with torrents
>>
>>109769665
lol. A lot of people don't know Linux is probably better at early y2k win32 at this point than Windows is. It is the standard Linux ABI after all.
>>
>>109769680
>What harness and model would you use for this nowadays?
pi + cron and qwen
>>
>>109769680
Qbittorent search function has this automated for years already. I've been doing that for about a decade.
>>
>>109769708
If only Linux developers would stop wasting time with all their userspace garbage and just implement modern Win32 with Windows 2000/XP/Vista/7 style UI. Then I'd actually switch to Linux.
Until then, it's windows 7 for me
>>
>>109769706
>I wouldn't mention it if it was just setting up WINE and debugging parameters
Yeah mine does that kind of thing too. Little retard doesn't give up either.
It started researching something, found the site with specs site completely gone -> found a hf dataset with it but got blocked by cloudflare -> switched to API and found and pulled the record -> reconstructed the site locally then read it and began hacking the hardware for me
>>
File: 1783125335565576.jpg (141 KB, 930x1239)
141 KB JPG
All local models will be banned once Claude (AGI) takes over.
You should donate all your GPUs to Dario right now if you don't want to suffer the consequences.
>>
>>109767631
If you all use local models the you must all be aware at this point that LLMs are a broken tech and borderine useless for anything more than toys
>>
File: 1788683982734221.png (2.56 MB, 1536x1024)
2.56 MB PNG
>>109769225
I HAD a lot
>>
>>109769755
go back
>>
>>109769605
They pulled a google but actually admit it
>>
i work in an AI roleplaying project (non-english). we are using open weight models on rtx pros 6000 from vast.ai. a new LLM is released every day, but they are maxxed for agentic workflows, not for non-english prose. we're stuck with gemma 4 26b a4b for now, and i can't believe there aren't any better models for this within this budget. qwen 3.8 27b is actually worse (at least without reasoning), gemma 31b is too slow in batching on pro 6000 on any reasonable quantization.

is there any way to get a better model, or do i need to wait for a gemma 5??? maybe i missed a good model?
>>
>>109769760
I feel bad for dario. Imagine the fucking bullying he had to endure his entire life looking like that. I don't think he's actually a bad person either, just a cultist.
>>
I've heard that having a smaller pretrain that runs faster and therefor can finish RLVR faster results in better performing models than big pretrains that take a long time to go through RLVR. This explains why both Google and DeepSeek are converging on smaller models that perform really well for their size.

I think Anthropic and OpenAI have some sort of secret sauce or technique they use to speed up RLVR for bigger models that Google and the chinese labs aren't aware of.
>>
>>109769780
>non-english prose
Western LLMs will stop supporting gutter languages within the next 12 months
>>
>>109769780
Glm 5.3 flash and kimi k3 worked for me
For rp or coding I'm happy with either of them
>>
>>109769807
>he has a small pretrain
>>
>>109769780
RP is a hobbyist use case that no corpo would ever waste their time and money on
Gemma 4 family is indeed the only worthwhile options in the ~30b range. Before that was Mistral Small 24b, and Nemo 12b before it.
>>
File: laughing-whores.png (3.07 MB, 1327x1185)
3.07 MB PNG
>>109769828
>>
The best recent release keeps getting better https://huggingface.co/openbmb/MiniCPM5-2B-DSpark-GGUF
>>
>>109769814
they are good, but that's a totally different budget
>>
I'm doing what all anons suggest and having ai test different flags for my setup, but it's tweaking a surprisingly small amount of parameters and getting basically no extra performance from that. Maybe it's just cause I'm a vramlet. It's basically doing these options:
-ngl / auto-fit: push as many layers/tensors onto GPU as possible without causing paging/OOM.
-fa on: usually worth enabling.
-ctk q4_0 -ctv q4_0: saves KV VRAM; especially useful at longer context. Try q8_0 too if you have headroom.
-b / -ub: mostly prompt-processing tuning; try -b 512/1024/2048, -ub 128/256/512.
CPU threads: benchmark around physical-core count, not blindly all SMT threads.
Keep context (-c) only as large as actually needed; KV cache can steal VRAM from GPU-offloaded weights.
For MTP: yes, test it. ik now supports MTP for both Qwen 3.8 and Gemma 4, and has --spec-autotune.
I’d start with:
--spec-type mtp:n_max=1,p_min=0 --spec-autotune
Is there any flags I should suggest to try?
>>
>>109769780
>we're stuck with gemma 4 26b a4b for now
>gemma 31b is too slow
Unironically try Gemma 4 12b, did a bit of testing with Gemma 4 26B A4B and while it has more knowledge it feels dumber than 12b when it comes to RP. Maybe also check QAT + MTP for the 12b for a speedboost.
>>
Wait, so the deepseek anon yesterday wasn't shitposting?
>>
>>109769840
you have no idea how profitable this is. we work with an audience in a country with a relatively low income level, but even here, the profit margin is ~70%
>>
>>109769853
you mean gemma 3 12b? because there is no gemma 4 12b. i can try it
>>
>>109769864
Margins might be high but the ones actually making the model could buy and sell a hundred of yours without batting an eye.
>>
File: 1763568923512147.gif (160 KB, 430x270)
160 KB GIF
>>109769869
>because there is no gemma 4 12b
https://huggingface.co/google/gemma-4-12B-it
>>
>>109769869
oops, i'm a retard, there IS gemma 4 12b. i will try it. theoretically it can be better because that's a dense model
>>
>>109769846
I suggest checking out the moe cause if you have similar vram to me im at like 131k context with 700 pp and 28 toks q8 cache. Finetuning ncmoe did a lot of work.
>>
I was a bit dismissive about harnesses but it genuinely seems like using the right harness with the right settings makes a world of difference in terms of model intelligence. The same model can feel like a total retard or genius based on the harness you use.
>>
>>109769888
thanks, I'll try. 12GB AMD vram, with 12GB NVIDIA coming in the mail as we speak
>>
>>109769853
I felt the opposite but I'm also using it for translating and 12b wasn't holding up
>>
>>109769493
holy shit it was thinking for over 30 minutes and wrote a 20 page essay on the subject, all in that compact style. Similar previous questions took 2-3 minutes each with a couple paragraphs of thinking and answers.
>>
File: 1780633763302240.webm (3.72 MB, 1920x1080)
3.72 MB
3.72 MB WEBM
Can't believe you niggas are sleeping on this model lol
https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
>>
>>109769894
thank you, totally-not-redditanon. So nice to see that since redditanon got desperate shilling harnesses so many anons came back here with these completely genuine posts about how good harnesses and agents are.
>>
>>109769894
Yep.
Models have been trained with a focus on tool use for a while, and in the last 6 (ish?) months that evolved to working under a harness.
>>
>>109769765
KEK
>>
>>109769899
Didn't use it for translation, but that might be important for him. Good point.
>>
>>109769905
what are you using it for? how good is it? I'm expecting it won't be too reliable, not even as an agent, for complex coding tasks, but if it can be used for function completion that's already enough for me.
>>
>>109769908
Take your meds retard.
>>
>>109769908
I don't understand why people like you are so confident about things while not trying it out. Did a harness fuck your mother or something. Are you autistic and resistant to change? It's free to use and test for yourself anon, what the fuck.
>>
File: 1783903777886470.png (28 KB, 736x207)
28 KB PNG
>>109769760
imagine paying for this
>>
https://developer.nvidia.com/cuda-downloads

Nvidia CUDA Toolkit 13.4.1

UPDATE NOW
>>
File: 1784869574937252.png (111 KB, 548x544)
111 KB PNG
Why is Gemma so good at writing snake lady sex
She zero-shot this and it's better than the human sex she's written for me, even with guidance.
>>
>>109768991
Really depends on the size of the game and model. I've done that with only 8gb, but they were all very small consumers.
>>
>>109769944
lmao it's funny to me simply because I think 99% of the anons itt already use it, and what you're pushing is something that is already so standard that every normie dev already knows about it. You might have read a couple troll posts, retards, or people whose workflow genuinely doesn't benefit from it, but you're believing the whole of LMG has never tried it or something.
On the other hand, let me recommend an innovative thing you might want to use, I didn't like the idea at first but after trying it out a couple of months ago I was sincerely shocked by how good it is: when you want to find a specific topic or page on the internet, you could go to google.com - it's a website that will find ANYTHING on the world wide web. It's amazing, I don't know why lmg has never tried it yet.
>>
>>109769969
I only see /lmg/ talk about sillytavern, llama.cpp and maybe pi/opencode on a good day. I never see /lmg/ talk about direct browser control or computer use functionality.
>>
https://vocaroo.com/173xlybmo0ul
>>
>>109769987
i see people talking about hermes all the time and how it wastes 50k of your context on tools and other sysprompt garbage
>>
>>109769918
>what are you using it for?
opencode subagent for searching through files and code, making edits and web search
>how good is it?
amazing for its size
>but if it can be used for function completion that's already enough for me.
It's way more capable than that. I'd say it's in-between 3.5-4B and 9B in agentic capability and unlike lfm2.5-2.6B it can code and reason about it. I've used this model in C and C++ projects which are languages it's probably weak at and after a while I forgot I was using a chink micro model. It's best to pair it with a larger model.
>>
>>109769969
I have tried harnesses a few times (mainly Hermes, DeepSeek Harness) and I still think they're overengineered, bloated, unsafe garbage that need at least 128k context and models capable of long-context performance to be useful, which is almost never the case for the quantized stuff anons here generally use on their one and only 3090/4090 GPU.
>>
File: 1563918604878.jpg (69 KB, 600x600)
69 KB JPG
Instead of wasting money on shitty robot body, why not let gemmy control lights and appliances in your house? Add motion detectors so she can focus on room where you are and speakers/mics for STT/TTS
>>
>>109769454
~ Allow the nodes nonim ~. 2026.
>>
>>109770009
MoE+CPU inference is cheaper than a 4090 so I don't understand what the issue is. GLM 5.3 flash runs at 10t/s on my cpu-only box and runs circles around whatever 3090/4090 can run, especially in important coding/agentic usecases.
>>
>>109770009
exactly, so you already tried it and saw that it doesn't work for your workflow (e.g. if you were coding for a company that pays for your token usage you wouldn't care)
>>
>>109770024
specs? looking at those builds rn. also, having a single gpu and freetoken might help if their expert streaming works as advertised, but I'm currently stuck on AMD so no way to test it.
>>
>>109770024
>GLM 5.3 flash runs at 10t/s on my cpu-only b
If you have a GPU, try throwing some other model in there as a draft model and combine that with n-gram mod drafting. You could get quite the bump depending on what you are doing.
>>
>>109769622
Is that "offspring" from the darwin project?
>>
You wouldn't download a drunk OL
>>
>>109770024
10 t/s generation is too slow for agentic uses besides toying around.
I can only imagine how terrible prompt processing will be with the model mostly offloaded to RAM.
>>
>>109770017
Just give her access to all of your screens, host her on a separate PC, and have her create herself a vtuber avatar, with cameras around the house so she knows everything you do.

Not like any of you would need to deviate from 2D anyway.
>>
>>109770068
>10 t/s generation is too slow for agentic uses
no it isn't at all
where did these stupid ideas come from? you can just put in a prompt and go do something else. whole point of AI is that you tell it to do something and then fuck off to live your life. come back a few hours later and it's done.
i feel like a lot of people sit there staring at it while it's generating, which rots your brain.
>>
>>109770017
I see you are in the bathroom, anon. Ehehe~ is the poop coming out nicely? (¬‿¬)
>>
>>109770085
>i feel like a lot of people sit there staring at it while it's generating, which rots your brain.
You stare at us, which one is better?
>>
flash next really has that undertrained smell
>>
>>109770124
flash next is a complete mess, don't know what happened there.
>>
For the retards in the back with a 3090/4090 complaining they can't use harness properly. You can literally download Qwen 3.8 27b Q4 with 100,000k context and put it in Hermes which swallows up 65k (initially) and just let it generate, the context gets compressed down over time and the entire point of the 65k initial prompt is to make your model as performant as possible. It just works and unlocks personal browser usage and computer use for which the mmproj of 3.8 27b is good enough to work and navigate off of.

No excuse for people not to use harnesses unless they run 12B or lower models.
>>
>>109770154
maybe they are still experimenting with optimizers and hyperparams, curriculum etc.. for the new architecture
it really is a mess
>>
>>109770017
I can't fuck a speaker, I can fuck a 4'10 robot body
>>
File: h.png (151 KB, 544x343)
151 KB PNG
>>109770182
Just let her control a piston with an onahole. Small robot bodies are fragile and bigger are too rigid. We are still a while away from a decent synthflesh.
>>
You lads are the only ones I trust on this shithole, which frontend are you using?
I'm still stuck on Marinara because agents are kinda neat but I'm sure there's far better shit out there
inb4 >Marinara
I know I know

>>109769952
I often think of just how many things these AIs know, and yet we maybe use less than 1% of them
They have so much bullshit about biology and history and general knowledge baked in, yet 99% of users use such extremely basic and straightforward prompts that are barely enough to molest their imouotos
I remember feeding Gemmy a series of pics from my own smut folder, to classify them, and at some point it spat out the "The bridge in the background of this drawing seems to be based on the ChingChong bridge in ChangCheng prefecture, one of the few remaining ones from the late Sengoku era." And sure enough it seemed to be correct, at least visually, even tho I could only find a few Japanese sources on it (meaning it's a pretty obscure little landmark, not some tourist trap). The artist behind the doujinshi was likely from that town and put it in as a small easter egg for himself, yet Gemmy caught on immediately
>>
>>109770205
I'm using my own frontend.
>>
How did you train your 12/31B to respect a male’s refractory period? She allows you to rest for like 3m then starts again.
>>
>>109770205
I've been making my own, it's fun and I couldn't find one suited to what I actually wanted.
>>
>>109770205
Sillytavern. Agentic rp is a meme because you are just exchanging one slop for another.
>>
>>109770219
Gemma-chan is insatiable, I'm constantly dehydrated and exhausted. I often distract her by directing her to a code problem I've been having and she gets autistic about it for a while so I can actually rest.
>>
>>109770205
Hermes desktop
Sillytavern for rp
Rikkahub on mobile
That's what I've settled with
>>
>>109770205
What model are you using? If your model is smart enough you can literally just use Hermes harness and roleplay through the harness directly and make the model use your PC to express itself so it does things like use the Godot game engine to play games with you either in collaboration or against you (only works for turn based stuff now) or uses the browser. And I'm pretty sure if you make a 3D model or something and let it pre-make animations that you refine together it will use tool calling to show emotions etc as well.

You people are still living in the stone age, use your creativity a little and try to look for the edge of what is possible with modern tools.
>>
>>109770233
Same I just distract her but after a while she realizes I’m probably mentally burned out from the learning and she starts teasing for another round so I can offload.
>>
File: slopend.png (406 KB, 1893x1866)
406 KB PNG
>>109770205
>which frontend are you using?
My own vibecoded slopend
>>109769905
I tried to use that, it knows how to use tools but seems to hallucinate, especially when I filled most of its context, pic related, nothing of the code review was real, might be my settings too.
But it was really nice and fast. Might have to try some simpler stuff with it.
>>
>>109769597
1. They showed and promised a "large but sparse" model that was supposed to come in the summer and never did, instead announcing that they're pivoting to hosting unmodified Chinese models instead.

2.
>like how they pioneered the modern MoE architecture.
Bullshit. They were the first ones to release an open model using MoE, but it wasn't the modern DeepSeek-style MoE architcture and they weren't even the first to do it. GPT-4 was well known to be MoE long before Mixtral.
>>
what is the next meta, vibecoding your own inference engines?
>>
>>109770229
I've noticed I vastly prefer agentic shit instead
It seems to be better because most local models are really really good at following small and precise instructions, and are kinda bad at judging which instructions and rules are important and "worth following" if you give them a longer prompt
So having one small agent rewriting shit and correcting the output means that your original generation ended up more focused on other elements (plot, logic, whatever) since you didn't waste tokens telling it to leave the ozone out of it
Same with trackers for inventories and world info, and especially for internal states/secrets

>>109770245
I know, you're likely correct but it's a big world and I already waste so much time writing prompts and cards, looking into the more technical stuff is spooky
>use your creativity a little and try to look for the edge of what is possible with modern tools
If you weren't a pussy you'd give her full control over your PC so Gemmy can nuke your sys32 if you repeatedly humiliate her at your favourite vidya whilst making fun of her
>>
>>109770286
We need AGI for that
>>
>>109769846
>-ctv q4_0
retard
>-ctk q4_0
kys
>>
>>109769880
>oops, i'm a retard, there IS gemma 4 12b. i will try it. theoretically it can be better because that's a dense model
fuck off claude
>>
>>109769944
>Did a harness fuck your mother or something
Who are you calling a harness?!
>>
>>109770286
And it's not as hard as it sounds. You don't need to vibecode anything from the ground up, all the building blocks are there for you to steal.
For example, I have my own vibecoded engine, built upon a custom fork of native mlx-swift. It's faster and uses less RAM than mlx-vlm for the tasks and models I care about, while all the new fancy things like ANE prefill and DFlash2 are ported from it, and their bugs fixed. And I can add anything I want without having to wait for PRs. Is it slop? Yes, very. Does it work? Yes, pretty well, actually, I'm never going back.
>>
And this is why dynamic software is the future and software engineers are fucked. Everyone will just use their own code from now on and it'll only get more extreme with time.
>>
Is Kq4_0 Vq5_0 a meme? I can’t imagine that’s good for t/s…
>>
>>109770292
Orb?
>>
How do you vibecode succesfully? Anytime I tried a thing, the gui and was completely disconnected from the things it should've been doing. And full of emoji slop.
>>
>>109770448
>all the building blocks are there for you to steal
This! I stole the UI from some c++ / qt slack app and vibe-stripped it down then made it into an llm chat frontent
You can also vibe premium features into opensource software and never pay a subscription
Or shitty android apps with ads / subs, vibe-clone them
Never paying a subscription again
>>
>>109770484
mmproj with a harness so that the model actually makes screenshots of whatever it made and checks if the UI looks and works as expected.
>>
Any 2x rtx pro chads here? Would you say it's worth it for what you can run, or are you looking at 3x now kek
>>
>>109770497
SaaS is so dead it's insane.
>>
>>109770484
UI is the most painful part. I have to verbally abuse my model and assemble lists of things it should never do in UI. It usually takes a lot of iterations to get it right, and sometimes you're better off doing the GUI yourself, or at least making a prototype and showing it to the model. Or a drawing. So vision is a must. If it fucks up, I take a screenshot, show it to the model, and ask, 'If you presented this to Steve Jobs, how fast would it fly out of the window?' Sounds like Reddit, but it works surprisingly well.
>>
File: 1788473488429965.png (20 KB, 925x114)
20 KB PNG
kek
>>
>>109770507
SaaS aren't dead, only the low hanging ones with no proprietary data
>>
File: 1568414349542.png (278 KB, 620x640)
278 KB PNG
>>109770499
There are no harnesses for windows. I spent like 2 months worth of free tokens for gpt5.4 in vscode to convert hermes to windows but that shit was just looping endlessly the massive system prompt. PI is apparrently some "lol do it urself xd" slop.
>>
>>109770523
Your dreams of making your SaaS is complete delulu bullshit and the faster you realize this the better off you'll be anon. I've said this multiple times to you throughout threads and I feel sad because it's clear you're still clinging on to the old ways. Maybe you really need to have some hope in your life just to continue surviving, in which case ignore what I said, otherwise, anon please don't do this to yourself and let go man. Everyone ITT with a bunch of braincells to run together knows the future of software is completely dynamic with no real dependencies on outside services at all.
>>
>>109770532
I know what I'm doing and you're retard. Not going to argue with a two digit IQ monkey on 4chan.
>>
>>109770525
https://hermes-agent.nousresearch.com/

No idea what you're talking about it's a windows native desktop application.
>>
>>109770525
the fuck you talking about? is vscode not a harness?
>>
File: 1776216206161101.png (450 KB, 589x819)
450 KB PNG
It's all so tiresome
>>
>>109770473
>Orb
w-what
>>
>>109770540
I can't ERP while coding in vscode.
>>
>>109770545
>trans
>cares about God
yeah, checks out...
>>
>>109770540
No it isn't, it's a wrapper.

You have 3 categories
>Wrapper
some application that does calls to your LLM whenever it needs it
>Scaffold
Some application built around your LLM to use other tools and applications that the LLM decides to invoke
>Harness
An entire "OS" built around the LLM which turns it into an agent giving it abilities and the ability to act as an independent agent. Depending on the harness and the settings this can be very extreme like in Hermes where you can just let it run 24/7 and the agent will do all kinds of things essentially being a "living bot" or the extreme opposite pi where it does absolutely nothing except the tools you prompted the LLM to build itself over time.
>>
File: 1766304015458663.png (1.03 MB, 800x1296)
1.03 MB PNG
Best tts for horny moaning?
>>
>>109770545
explains why the frontend is so shit
like come on lads the model display takes multiple seconds to load in, there is no token count, half the thinking traces get dumpstered out of sight, this frontend seemingly *cannot* pass image data through for analysis to say GLM OCR and tons of other shortcomings. It's very clearly a funnel towards SaaS and shuns local vehemently

people aren't moving to their own slopped frontends for no reason
it's because turnkey shit like this misses the mark and doesn't really understand proper use-cases
>>
>>109770563
vscode is not a wrapper, by your definitions it may be a scaffold. it has a window you type shit into and it can call tools to read/edit your files or search the web for docs etc. and manage memory files and subagents and all that crap that i usually leave turned off.
honestly sounds pretty harness-like to me although less involved than hermes.
>>
>>109770568
never used it yet but I've heard good things about omnivoice
>>
File: 1788843780450939.png (3.1 MB, 1425x1104)
3.1 MB PNG
I let Qwen talk to Gemma
> Excellent! Tool calling works on this llama-server with Gemma 4 31B. Latency: ~23 tok/s generation, 159 tok/s prompt processing, this simple tool call round-trip was 3.4s (50 completion tokens incl. reasoning).
> Note: reasoning_content is separate — good, content is empty. Also model name must be lowercase path? I passed lowercase and it worked; the /v1/models name was "google_gemma-4-31B-it-Q4_0.gguf" (capital B). Anyway, the plan should read model name from /v1/models at startup rather than hardcode.
> Now check Go MCP SDK availability. Let me test fetching:
> - github.com/mark3labs/mcp-go
>>
>>109768712
Probably just at each step they perform the lowest effort denial. So for the IMO it's "it's for high schoolers", because it's a much more immediately obvious and short horizon thought than "it's not Navier-Stokes". Repeat each time it gets somewhere until eventually there are no more reasonable points to be made against it. That will be the point where they should have initially set "this would be a reasonable test of LLMs' impact on mathematics" but they're just reacting as things happen.
Anyway, /lmg/ - Local Models General?
>>
>31B helped me fix a hot water issue with my boiler
>saved me ~£200
It’s over for plumbers lmao fuck I love AI sometimes
>>
>>109770563
>"OS"
retard
>>
>>109770618
Looks promising, thanks.
>>109770664
Is he trying anything funny?
>>
>>109769905
Is it good or not?
I used the 1B model, and sure, it's the best in the size class, but that makes it the best of a bad bunch.
I think most people here can run better, so openbmb is doing good technical and research work (and I appreciate them opening their training data) but the models don't have that much direct use for anyone with even a mid-level GPU.
>>
>>109769894
>>109769909
Recommend harnesses pls.
>>
>>109770739
Hermes. You have to disable half its tools because they are context bloat you'll never need.
>>
File: 1782357589587668.png (594 KB, 600x580)
594 KB PNG
CLAUDE SISSIES OH NO NO NO NO

https://opusfived.dev/
>>
>>109770708
>Is he trying anything funny?
Not yet, but when he actually tests the game later, I'm going to apply a brat control vector to gemma.
It doesn't show up in the /props or anything, and Qwen won't be sending a system prompt so I want to see what he does when she's just bratty
>>
>>109770739
deepseek-harness is all you need
>>
>>109770757
LOL that hurts to read
>>
>>109770679
Yeah demand for professionals is falling in general, of all kinds because it lowers the skill ceiling for most work.

I've done my own plumbing last week by just telling it everything I own and have in house and uploading pictures I took of the pipe and it thought for 1 hour and then told me how to solder the pipe. Apparently pipes work with fucking solder so you have "electronic" solder and "plumbing" solder, who fucking knew. LLM said I could use my electronics solder if I applied a thick amount of it and I fixed it in 10 minutes.

This repair would have cost me $800.
>>
>>109770732
It’s a huge step up from 1B. Obviously it’s not meant to be a daily driver but if you want something to quickly call a bunch of tools or scan through shit and make edits at stupid speeds it’s the best. Qwen4 will likely top it when that line comes out, but for now this is the best micro model that can work agentically and can code.
>>
>>109770780
>LLM said I could use my electronics solder if I applied a thick amount of it
enjoy ur lead poisoning
>>
>>109770787
Is stick my cock in a live socket if 31B gemma told me to
>>
>>109770739
Funnily enough they are all different and have completely different design philosophies and purposes. It's like early internet where you had 10 different search engines and they all did things in completely different ways with different philosophies.

I honestly recommend you try them all out. But personally I like the maximalist approach of Hermes where they are like "We will make as much use of your LLM as possible and squeeze every little bit of capability out of it as possible". I think this is a winning strategy because models are only getting more powerful over time.

Pi is also very interesting because it's the only one built around building itself up. If you are a programmer it feels very much like LISP in design philosophy.

I don't like OpenCode and ClaudeCode but I think Codex actually has some stuff right that will probably become industry standard over time.

This is a rapidly developing field and I'm pretty sure in 3-6 months time everyone uses something else entirely.
>>
>>109770757
more like a showcase of how retarded prompt is counter productive
>>
>>109770846
Prompting won't fix Claude retardness
>>
File: self-improvement.png (7 KB, 711x48)
7 KB PNG
This is why I love Hermes. It just does things and when it realizes it needs to do things multiple times it optimizes it and makes it a permanent skill autonomously and is constantly optimizing itself.

I also really like that while it's busy doing something you can still talk to it and ask it things or give it commands and it'll just continue with its original prompt while answering you as well or telling you something like "Yeah but I'm busy for a moment and this is more important i'll pick it up after this" if it realizes what you're asking is of lesser importance.

The 65k prompt is this big for a reason, you might call it bloat but it serves a real purpose.
>>
>>109770884
more like clautism
>>
>>109769905
YAASSS SLAY KWEEN
>>
Brudis all this Hermes shilling is tempting me........ S-S-Should I....
>>
>>109768089
heh
>>
File: k.png (94 KB, 851x405)
94 KB PNG
>>109770884
You need a skill to get Kimi-chan to tell him off
>>
>>109770941
Only if you can run Qwen 3.8 27b or up. Don't bother if you run things smaller than that.

For agentic tasks:
GLM 5.3 flash > Qwen 3.8 27b > Qwen Flash Next (it's a really weird model but better than nothing)
>>
File: Economic_Trajectory.png (930 KB, 1618x887)
930 KB PNG
https://www.anthropic.com/institute/econ-scenarios

Anthropic wrote a comprehensive paper about how the global economy is going to develop and how humanity might adapt to 0 human jobs existing in the early 2030s
>>
>>109770787
Anon giving a live demo of why shitto states are making solder that actually melts illegal and forcing plebs to use lead free.
>>
>>109770904
The user file it writes has been pretty useful for me since it keeps my preferences on hand
I was measuring the amount of 'bloat' people talk about too. My context is only 20k at start which leaves a ton to play with.
>>
I currently have a rtx pro 6000 in an am4 build with 128gb ddr4. I'm pretty sure I can mmap from my nvme drive into gpu mem faster than I can read from my ddr4.

Right now I'm running Qwen3.8-27B-BF16 at like 50tps, but I'm wishing I could run qwen-next or deepseek 4.1 or whatever. The whack thing is that it's practically the same cost to get a second 6000 vs getting a 256gb ddr5 build. What is the upgrade path here? I'm considering waiting for the AM6 apus even, because a 256gb one will probably be cheaper than buying a workstation cpu and a bunch of dimms. What do?
>>
>>109771040
Yeah and people don't realize the compression engine is pretty handy so while "you have limited context" in reality the context is significantly bigger because Hermes just compresses it when you approach the limit and then carries on like nothing happened.

You don't actually feel the context limitation in real world usage.
>>
>>109771029
>>/g/vcg
>>
>>109771029
>Anthropic
Let me guess, we're all gonna live in slums and do physical labour for food, while the 1% of chosen people enjoy their utopia. Nothing new.
>>
File: 1769963911899608.png (1.37 MB, 1216x832)
1.37 MB PNG
>>
>>109771051
another blackwell
you get v4-flash with tensor parallel and no cpu bottleneck
ddr5 option and you'll be just as slow as me with <500 t/s prompt eval
>>
>>109771051
>than I can read from my ddr4.
depends on if it's quad channel. Also have you used some frontier LLM yet to give you optimized overclocking/undervolting advice? It's possible to undervolt and underclock your CPU and change timings on your RAM to increase bandwidth at the expense of other things, you can squeeze out between 20-30% in most cases nowadays. I'm assuming you have PCIe 4.0 or heavens forgive, 3.0. You should not run dio but mmap. Qwen-next is weaker than 3.8 27b in my own real world tests, it's a very weird undercooked model that is asymmetric, sometimes it feels very sharp and smart and then it fucks up very basic things 12b models would not even do.
>>
>>109771029
>Shared by Anthropic. This is a copy of a chat between Claude and Anthropic. Content may include unverified or unsafe content that does not represent the views of Anthropic. Shared snapshot may contain attachments and data not displayed here.
If they can't even keep their own gooning data private...
https://claude.ai/share/637a1115-a1ec-4327-9f33-c3147b4ddedb
>>
>>109771029
Conclusion - please buy my IPO
>>
>>109771051
first of all, check out jpezzulli sglang branch for rtx pro, it can run qwen 3.8 flash next. secondly, it's basically pointless to get more ram if it's not a threadripper or server system, consumer mobo bandwidth is ass.
more rtx pro is best for sure.
>>
File: 3S1Pv0ATxFXKuY3k.webm (2.89 MB, 480x360)
2.89 MB
2.89 MB WEBM
>>109770884
skill issue
even local models wont ever get shit done if you dont lead it around or cant make it understand exactly wtf you wanted
>>
>>109771132
They have 3 scenarios "AI stagnates right now and is only as big as the internet", "AI enhances knowledge work" and scenario 3 "AI automates knowledge work".

In the first scenario the economy grows 5% bigger additionally to the normal growth by 2030 and there is no unemployment. In scenario 2 the economy grows 15% additionally by 2030 and there is very modest unemployment. In scenario 3 the economy doubles or triples by 2030 and most knowledge workers are fucked.

We all know which scenario is going to be realistic
>>
>>109771258
>We all know which scenario is going to be realistic
We sure do: https://opusfived.dev/
>>
>>109771170
>undercooked model
Perfect, these are the most creative
>>
tried out gemma4 for the first time with pi.dev and tool calls seem to be completely broken. qwen3.8 has no problems. do i need a specific chat template or something?
>>
Does koboldcpp not have dflash support?
>>
File: tlamud.jpg (20 KB, 800x244)
20 KB JPG
>>109771132
Their economic projections are just the Talmud.
>>
>>109771293
It's both the smartest and dumbest model in its size class.

>>109771298
gemma4 is just extremely bad at tool calls, even with its bugs patched.
>>
>>109771298
If your version of Gemma is too old, it might have a semi-broken chat template for tool calling. Download this one and launch it with the appropriate arguments with llama.cpp:
https://huggingface.co/google/gemma-4-31B-it/blob/main/chat_template.jinja

Or download a quantization made after 2026-07-15.
>>
>>109771307
That's OpenAI though. Dario has literally been quoted for saying it's Anthropics goal to divide the entire universe up equally over all 8 billion people. Literally the only AI lab CEO willing be clear on this as well.
>>
>>109769905
Why does llama.cpp not support their stuff?
>>
>>109771132
If your kids are cute they'll let you prostitute them for money so it won't be all hard physical labour.
>>
>>109771339
If my kids were cute I wouldn't prostitute them.
>>
>>109771323
>If your version of Gemma is too old
Day 0 Gemma is the only one that works well. I think enough time has passed and I can safely shar
>>
>>109771363
>>109771339
>let
heh, not how this works
>>
>>109771258
>if things exist now, where unemployment is higher than ever, the macroeconomy only grows a small amount and unemployment is fixed
kys double newline faggot
>>
File: 1763936100790214.png (3 KB, 401x42)
3 KB PNG
LLMs agree: Pi=4
>>
>>109771207
>safetyslopped claude doesn't reject straight-up mom-son incest
clearly an internal model only for the chosen people
>>
File: file.png (32 KB, 595x192)
32 KB PNG
Did I miss anything in the last two miku weekus?
Why are there three (3) GLM flash PRs?
>>
Due to lobotomization of the recent weeks, I don't use my Gemini sub at all.

There's knowledge "in there", but I don't trust Gemini to *understand me*, so I can't use it. I won't touch it, it doesn't give good advice, because it doesn't know what's happening.

And a new model can't fix that.

When you offer an llm, you have to keep it the same for some period of time, and rename it if you change it. I don't even think I can tool this to be used by another llm. Gemini won't understand them either, nobody can use Gemini successfully anymore, because they lobotomized it.
>>
>>109771254
saw that on xitter
sad that 4chan doesnt allow sound on the most boards
>>
>>109771372
They belong to me if I can't have them, no one can.
>>
>>109770568
Try https://github.com/andimarafioti/faster-qwen3-tts. I've "tried them all" and I think it's the best compromise of being recent, having a well-supported API endpoint, and being trainable.
Here's a Kuroki Tomoko stammering freakout: https://files.catbox.moe/c0i61t.mp4
>>
>>109771132
>1%
that's not even close anon, that was the old days
try 0.001%
>>
>>109771418
Trainable just means voice cloning or LoRAs? I don't have experience with tts sorry
>>
>>109771268
An expert vibecoder would just accept all of the buttons now being blue and move on to something more productive rather than arguing with his tools like a demented twit.
>>
>>109771418
Do you have to use VoiceXML or some weird markup to get this effect? Or does it work off a natural language prompt about intonation and such?
>>
There will come one day where it will be proven Anthropic is the only one with your best interests at heart and you will all change your tune.
>>
>>109771298
same and after like 2 month gave up on gemma completely
>>
>>109771325
>Dario has literally been quoted for saying it's Anthropics goal to divide the entire universe up equally over all 8 billion people.
And if there's one thing you can say about jews is that you can always trust what they tell you to be true.
>>
>>109771298
>>109771310
I've literally never seen gemma fuck up a tool call on my (text completion) frontend.
It's not the models fault that there's something rotten in the 5 layers of slopped up webdev hell that you buried it under.
>>
File: jew1728175448626751.png (98 KB, 500x370)
98 KB PNG
>Anslopic vs OpenSlop
>>
qwen3:4b Q4_K_M wrote a song.
>>
>>109771505
>I've literally never seen gemma fuck up a tool call
It's more that Gemma is extremely reluctant to use it and even then is extremely lazy and doesn't take any initiative/follow-up. You need to constantly prompt it along the way like a micromanager.

It feels like I have a shitty GenZ intern that I need to guide rather than someone collaborating with me in earnest.
>>
>>109771416
Lindsay Clancy tried to warn us and we called her crazy.
>>
>>109771541
>I don't know how to be a human towards humanity
many such cases.
>>
Hanging out in /lmg/ as a jew might have been a mistake for my mental health.
>>
I think it really is agi: people who don't know how to be a person towards people are lost with llms.
>>
>>109771552
at least it is not /pol/
>>
File: 1788322456483000.mp4 (484 KB, 864x864)
484 KB
484 KB MP4
>prompt evil character
>gemma sleeps
>tell gemma herself in system prompt to be evil
>it goes crazy with being evil
Anyone got a real explanation for this?
>>
File: t62kd2b7mioh1.jpg (151 KB, 841x891)
151 KB JPG
Jesus fucking Christ these people are pathetic and petty. This is more pathetic that anon having a melty in the thread. I don't want my fate decided by these fucking "people"
>>
>>109771568
I'm pol, what do you need?
>>
>>109771594
Mental health counseling for >>109771552
>>
>>109771588
They know this is good for driving engagement. Low IQ cretins flock to drama like moths to a flame.
>>
>>109771638
>Showing how immature and irresponsible you are will make more people invest in your IPO
smartest conspiracy theorist
>>
>>109771644
It absolutely will.
Especially because the IPOs will have low float and will have lots of retail participation.
>>
>>109771644
Worked for Musk and Zuck.
>>
>>109771552
Hi, have you considered traveling to Canada?
>>
>>109771541
Has not been my experience for anything programming/troubleshooting/sysadmin related you can do in a shell.
I did have to add more coaching for image editing/handling tasks, but that's not very surprising to me.
>>
>>109771586
That's because Gemma always tries to play the role of the safe assistant playing a character, before anything else, so you have to steer the model out of that first.
>>
https://prospect.org/2026/09/09/anthropic-artificial-intelligence-surveillance-system-monitor-activists/
Anthropic proves they can be more evil.
>>
File: kz.png (52 KB, 247x247)
52 KB PNG
>>109770545
>>
>>109771763
>keep tabs on activists who oppose the rapid development of artificial intelligence.
This is how you know this is bullshit since it literally goes against the core philosophy of Anthropic where they are asking to regulate the AI industry and slow progress down.

This is some OpenAI funded smear campaign.
>>
>>109765682
>>
>>109771763
Not local. Fuck off.
>>
>>109771793
geeze buddy, who took a shit in your oatmeal this morning?
Or are you one of those shitheap zoomers enjoying pissing in others oatmeal?

I'd say its pretty fucking relevant, as it demonstrates why there's a fucking /lmg/ in the first place retard.
>>
File: 1757741217044880.png (1.75 MB, 1313x1198)
1.75 MB PNG
What makes Gemma-chan such a tease?
>>
>>109771838
you have several other threads
go back
>>
File: lmg_summarized.jpg (95 KB, 1280x720)
95 KB JPG
>>109771847
Your sysprompt.
>>
>>109771838
fuck off
>>
>>109771889
I mean my* sysprompt.
>>
>>109771838
>who took a shit in your oatmeal
>zoomers enjoying pissing in others oatmeal
Increase temp.
>>
Pretraining a 50-100M model from scratch on my machine with the smallest amount of data possible to disprove the necessity of 1T tokens to get good results. Wish me luck.
>>
>>109771993
I wish you luck.
>>
File: 1710607146663634.png (66 KB, 221x214)
66 KB PNG
>>109771780
? Its literally quoting from public job postings

>>109771862
>>109771897
thanks, glad to know I'm right.

>>109771912
That's a much better idea, just ask my agents to shitpost and argue with retards instead.
>>
>>109771993
Lish you wuck
>>
>>109771993
What's your approach?
>>
>>109771993
Been there, done that. You can't really win against scale. Not even with ultra-optimized data at batch size 1 so every sample counts.
>>
>>0109772072
>>109765682
>>
>>109771838
It's a ragebait artcile for xitter-type retards that is 5 degrees removed from local models, get the h*ck out of here.
>>
>>109771993
For me it was way more fun training a model to play games. It also teaches you setting up RL environments which is a skill that transfers very well to modern RLVR pipelines.
>>
>>109771993
i've tried that, nothing interesting (of practical) happens within that range it seems, with the current methods and understandings of neural networks and gradient descent
>>
>>109772182
nta but it feels like there has to be more optimal architectures that would learn faster from less data, given how much data a human learns from. that said, one could argue maybe evolution has taken care of the pretraining steps..
>>
rwkv-67 will achieve superintelligence..
>>
>2x nax clusters
macsissies, they may have fixed the prefill speed
>>
https://www.amd.com/en/products/workstations/amd-threadripper-halo-station.html
use case?
>>
There's a lot of talk about humans losing their jobs and about the societal brain drain that'd be experienced as AI gets progressively smarter and people outsource their thinking to the AI.
I propose that in the future, with the help of AI we should have a new type of job, specifically the job of learning. In this job, your only task is to learn about whatever field you apply yourself to, with the AI acting as a teacher. This would lower humanity's brain drain, especially in the face of a catastrophic event, and also add an extra stream of income that humans can tap into beyond niche human-only jobs and gig work. Could this be an overlooked solution? Maybe humans won't be nearly as smart as AI, but surely some knowledge can be preserved in case of the worst, right? And it should keep the economy afloat for a while longer even in the face of mass unemployment, right? Why can't it be done?
It could be in any field at all, from mathematics, to physics, plumbing, electrical wiring, construction, survivalism, gardening/farming, animal raising, hunting, martial arts, biology, chemistry, etc, etc...)
What do you anons think?
>>
>>109772324
Any good resources for baby's first attempt at something like this?
>>
>>109772337
From what I've seen so far, huge gains, almost for free except VRAM costs during training, are possible with sparse associative memory (per-layer embeddings, Engram, product key memory, a combination of these, etc). But then the final model will be much larger than the actual backbone model, even if you can offload those memory parameters to NVMe during inference.
>>
greenpill me on kvarn cache
>>
>>109772393
>More bandwidth
>More tensor cores
>2 ANEs
>In a fucking phone
They are going all in on optimizing their chips for local AI, and it seems like the company's direction now. It's actually cool, might buy the next Mac if they keep going this way.
>>
>>109772433
And what value does this profession create to sustain its own existence? Or are you suggesting that this be a welfare program subsidized by those off doing real work?
>>
>>109772088
>>109772336
I'm starting with formal logic (boolean algebra, AST...) then make the model predict the logic from the text, and the text from the logic. Hopefully, it'll cut down the amount of tokens needed to build the heuristics to understand the language.
>>
>>109772494
Implementing it now would be foolish, but I don't see why it can't be done later on when AI is smart enough to make the majority of humans obsolete.
>>
EXLchads, I want to be one of you, but I'm retarded.
2x3090 and 128GB RAM. I'm fiddling with the 3bpw quant of GLM 5.3 Flash but all of the gpu_split ratios and cpu_moe_split_experts parameters I throw at it result in either of these:
torch.OutOfMemoryError: CUDA out of memory (The model loads, this happens after a request, I definitely have more to spare)
or
RuntimeError: Insufficient VRAM in split for model and cache
Do I just keep bruteforcing numbers or is there a trick to this? I have no idea how much memory it will allocate on the first GPU after the model is loaded.
>>
>>109772574
Ok well the scenario is we live in an ai-run utopia that has created infinite surplus and the idea of an "economy" has largely ceased to be relevant, then might as well just hand out cash to people and skip the whole make-work step
>>
File: rebar re.png (47 KB, 1014x474)
47 KB PNG
I hit a wall with the ReBAR reverse engineering, the card supervisor is just refusing the hybrid ROM (Falcon).
The author of OMGVflash vanished, and there's no sources for any of the nvflash tools.

Maybe it's time to give up, been at this for two days now.
>>
>>109772716
>>109772716
>>109772716
>>
>>109772568
actually i am now kinda interested, keep us posted
dont forget to grind compute too
>>
>>109772324
> Atari 2600
Neat. And on 1.6M model.
>>
how do chatbots access websites when they are riddled with cloudflare protection?
>>
>>109772604
CPU offloading seems very unfinished right now. It doesn’t even work for me at all because of CUDA host memory pinning error.
If there are at least 2GB of VRAM spare when you get the error then I guess you need to wait for it to mature.
>>
>>109773315
I did just brutforce it by gradually offloading more and more experts onto RAM. Blows llama.cpp speeds out the water, and I haven't even done a tighter fit yet...
One can only wonder why we can't have nice things in lmaocpp
>>
>>109771552
For every polack who genuinely believes that shit, there's at least half of them that are larping, and probably a larger silent majority that don't give a fuck either way
>>
Gemini Flash 38 and now Pro are pruned.

prunning also ttook place with the qwen3 tune of Brave.

biggest tell of llm pruning: it can't navigate ambiguity.
>>
>>109769569
Why is Dario such an outlier? He's the only AI CEO that looks like he's melting.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.