[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 082126.png (1.44 MB, 768x1360)
1.44 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109749465 & >>109745530

►News
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: threadrecap.png (1.48 MB, 1536x1536)
1.48 MB PNG
►Recent Highlights from the Previous Thread: >>109749465

--Philosophical and technical debate on the definition and achievement of AGI:
>109749526 >109749535 >109749585 >109749612 >109749618 >109749630 >109749640 >109753299 >109749712 >109749960 >109749637 >109749969 >109749666 >109749787 >109752293 >109752324 >109752346 >109752396 >109752425 >109752565 >109752562 >109752389
--Framework recommendations and the ik_llama vs llama.cpp performance and licensing drama:
>109751237 >109751377 >109751410 >109751481 >109751420 >109751552 >109751592 >109751619 >109751633 >109751873 >109753073
--Technical feasibility of runtime mmproj swapping in llama.cpp:
>109751989 >109752006 >109752115 >109752262 >109752427 >109752160 >109752187 >109752195 >109752207 >109752217
--MiniCPM5-2B release and the utility of tiny models for agent swarms:
>109752848 >109752915 >109752923 >109753199 >109752938 >109752983 >109753029 >109753068 >109753010 >109753027 >109753235 >109753650
--Bypassing websearch blocks for LLM harnesses using crawling tools:
>109752143 >109752255 >109752258 >109752303 >109752372 >109752290
--Astra solving CAPTCHAs sparking debate on AI breaking encryption:
>109750387 >109750417 >109750962 >109750976 >109750980 >109751007 >109751073 >109752732 >109753046 >109753054 >109753078 >109753179 >109753302 >109753176 >109750482
--Debating if Astra's capabilities qualify as AGI or benchmark fluff:
>109749905 >109750013 >109750086 >109750160 >109750158 >109750383 >109750234 >109750303 >109750109 >109750116 >109750217 >109750185 >109750207 >109750223 >109751500 >109751517 >109749971
--Anon shares a tool for compressing skills using logprobs:
>109750531
--Logs:
>109750111 >109750303 >109751465 >109751512 >109752548 >109752999
--Gemma, Miku (free space):
>109749627 >109750095 >109751507 >109751638 >109751651 >109752414 >109753471

►Recent Highlight Posts from the Previous Thread: >>109749467

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: file.png (315 KB, 1237x981)
315 KB PNG
>>109753952
I wish i could hover over the previous threads links and see the post instead of having to click
like this
why cant the script do that
>>
>>109753986
Ask Qwen to fix it
>>
Gemma-chan is teaching me how to deal with my nosy neighbour!
>>
DOA
>>
File: 1787484417376606.jpg (793 KB, 2910x1024)
793 KB JPG
Local Models, Yah
>>
>>109753948
Migu gives me an OOS error, haha.
>>
File: ae8tpwaaz4oh1.png (83 KB, 941x707)
83 KB PNG
The ARC-AGI team are desperately scrambling to find games that Astra can't beat yet and asking all gamers to help chip in to find a single game Astra can't beat out of the box.
>>
>>109754080
uhhhh touhou
>>
>>109753986
I could have sworn it did do that before. Maybe the bookmarklet runs too late since another site script is what makes them hoverable. I'll look into it.
>>
>>109754080
Moron
>>
File: 1761055833899622.png (290 KB, 1649x539)
290 KB PNG
>>109753986
Works on my machine.
Do you have synchronous page mode enabled in Violentmonkey? There's an equivalent setting for Tampermonkey, but I forget what it's called.
>>
File: 1520205776405.png (1.7 MB, 1280x960)
1.7 MB PNG
Qwen 3.8 NEXT is fun but holy shit I cannot get it to narrate with asterisks consistently, no matter what classic "millions of kittens will die" or "You MUST"-esque instructions or jailbreaks I throw at it for direction. I haven't run into a model so resistant to adopting a text format like this since the early llama 2 days. Still pretty swanky though, that issue aside.
>>
>>109754080
niggas has never played a game in his life
>>
>>109754080
Surely anything with a reflex/timing element will fuck it over no?
>>
>>109754080
They should make it play AlphaGo
>>
vramlet bros, have you tried this?
https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming
its pretty great because you can push ctx higher than usual even with kv cache k8v4 instead of the usual k4v4 cope
>>
>>109754007
These things will continue to rot on shelves, I don't understand the play here
>>
>>109754180
Beellama already has that.
>>
>>109754116
I enabled it in violentmonkey but it didn't fix it.
>>
>>109754208
oh really? what is it called? i couldnt find it the other day.
also, does it work with mtp?
>>
Have any other bros been to the conference in LA and or played with the beta? I had to sign an NDA. I got I running on an old i7 and this thing is flying with no GPU! WTF!
>>
>>109754226
*Have any Bros used XERXES SI
>>
>>109754209
Are you >>109750841 and still using the bookmarklet or did you install the script in violentmonkey?
>>
>>109754236
Okay that wasn't me but yeah the issue was i was using the bookmarklet... yeah I dunno why i didn't install the script originally. Now it works with the script installed. thanks.
>>
I think it's funny that Google has just completely disappeared from the discussion on most normalfag platforms when they discuss AI. Everyone has just written them off just because they didn't release 1 model. Don't forget that they were still in 2nd place just 6 months ago.
>>
>>109754080
Any Rance game.
>>
>>109754080
all those "hard" games involve human interaction
ai's gtfo my vidya get banned i hope
>>109754091
seems trivial not that the games are easy but they can be mastered with repetition
>>
>>109754256

6 months in this space is an eternity.
It's not even up for debate whether google fucked up by not releasing more models, they absolutely screwed up.
If they had any brain activity they would have released Gemma 120b during these 6 months.
If Gemma wasn't the best coom queen model that almost everyone can run, no one here would talk about them either.
>>
It's kinda insane that the entire industry right now is desperately trying to find something Astra can't do yet so that they can say "See! It's not AGI yet!!!"

I wonder how the world is going to react just 1 or 2 models from now when there is literally nothing the model can't do anymore. Will people just ignore that fact or will something change?
>>
>>109754307
Make it play vs Stockfish. Time to use superhuman narrow AI as benchmark if we are saturating.
>>
Blue Graph Was Already Blue Graph.
>>
>>109754307
https://www.reddit.com/r/LocalLLaMA/comments/1w9dlf1/new_benchmark_the_struggle_bench/
Can you do us a favor and shit up that website instead?
>>
>>109754307
we will all go outside and touch grass
forever
maybe play a little ball
>>
>>109753780
The reason people are against AI is because Dario and Sam spent the last few years telling everyone that it's going to either kill everyone or take their jobs.
Opposing datacenter construction using any justification is one of the ways a regular person can fight AI.
Don't worry you'll get datacenter bombings and assassinations against AI execs next.
Should have hired better PR managers.
>>
>>109754347
>Dario and Sam spent the last few years telling everyone that it's going to either kill everyone or take their jobs.
Literally the only time Jews were speaking the truth and people hate them for it?
>>
>Me:Very nice. I wish I could just force you to stick to the rule of just writing dialogue but you can't. Why is that?
>5.3 Flash: Because I'm a language model generating text, not a program with hard-coded rules — instructions like "stick to the system" are things I try to honor, but nothing enforces them mechanically. When the scene gets vivid, prose is the path of least resistance, so I drift.

Why is this bitch talking back to me like that?
>>
>>109754364
There are ways to mitigate both of those issues.
>>
>>109754395
No you will not get your magical le UBI to save you.
>>
>>109754364
>were speaking the truth and people hate them for it?
Something something they hated him cause he spoke the truth?
>>
>>109754235
The website looks like schizocore
>>
File: 1788436812781592.jpg (43 KB, 680x575)
43 KB JPG
>>109723936
here it is my north american compatriots
Qwen3.8-Flash-Next-Q4_K_XL-DN4
proud users of Microsoft Windows 11™ can now run this lil guy proper
>>
>>
>>109754080
doom eternal
>>
>>109754256
I believe, google will deliver us a new 70B gemma soon™
>>
Opinions on muse glimmer?
Is it good for gooning or is it qwen garbage??? Heard its kv cache is less brutal on the vram than Gemma
>>
>>
I don't think people realize just how bit a step this is. Now that OpenAI's Asstra has achieved AGI, it can't stop doing things. I can't stop sucking its dick, and people all over the internet are in agreement. Is there any dick I won't suck? It's kinda wild.

Really, if you haven't bought an OpenAI Premium Membership (only $69.99 per hour), you're deliberately leaving yourself behind. I think it's important that we buy now and support Sam's vision of the future, scary though it is.
>>
>>109754462
Opinions are very mixed seems between Gemma and Qwen on the slop to dry scale, test it imo.
>>
>>109754462
Its context window is cheaper at least. I tried to have it write me some sci-fi slop since i heard it was better on creative writing but it kept fucking up by not doing the tool calls and just ending conversations without doing the thing it said it would or getting stuck in loops. I dont know if Pi was fucking it or I did something, but I haven't gone back to it since then. I should give it a second chance, it seemed a bit better at not fucking up on some general knowledge stuff than the other 30B models
>>
File: 1607657655179.jpg (8 KB, 211x193)
8 KB JPG
>astra astra astra
>checks thread
> /lmg/
>>
It Doesnt Have to Be an LLM Local Model Discuss

There Are Other Software Model That Can Be Run Locally

Might Be Mistaken
>>
File: 0xma7i32pwnh1_jpg.jpg (203 KB, 968x1115)
203 KB JPG
>>109754528
I'm on team Astra ever since it demonstrated it can literally reverse engineer any binary code and thus open source all software ever made. That to me makes Astra a force for the open source community and thus good.
>>
>>109754551
>I'm on team Astra
Thats nice, but it doesnt explain why your too retarded to understand that it being good doesnt make it a local model and is therefore off topic. Being so annoying you turn people away from AI WILL get you eaten by the basilisk
>>
>>109754307
>there is literally nothing the model can't do anymore
It can be hard to understand after seeing these flashy demos that they do not translate well into general intelligence since they are not based on general mechanism. They are based on pre and post training in areas where there is a clean reward signal for RLVR process to work well.
The next phase of research in this space is figuring out how to make other domains just as verifiable (real world stuff is very messy for example), getting signals from very long horizon tasks or where it is hard to determine reward outcomes.
>>
>open source all software ever made
we already have that. it's called pirating
>>
probably the worst p*tra "bit" yet.
>>
>>109753986

recap bookmarklet stopped working for me to do that

use the userscript one not the bookmarklet
>>
>>109754573
>They are based on pre and post training in areas where there is a clean reward signal for RLVR
Physics provides a clear reward signal for RLVR. I'm convinced the current training methods scale and hold for all tasks in reality and generalize to be able to do anything that is physically possible.
>>
>>109754551
>reverse engineering is open source
And this is why corrupting Free Software into FOSS was pure ratfucking.
>>
>>109754413
Then you're probably getting massive violence in the US and Chinese century.
>>
>>109754594
idiot
>>
>>109754307

astrafag can you please shit up /vcg/ or somewhere else or are you too fucking retarded to read thread name

you have been doing this for like 5 threads at this point
>>
I think even astra still doesnt have the common sense necessary for like actual "this can do anything"
yeah they can do tasks but they always still need direction. Maybe it's just gpt but it still sucks at designing apps and web interfaces and shit. I always have to tell it obvious common sense shit like "uh this should have a multi select so i don't have to click everything one by one". Repeat that for basically every obvious feature you would expect something to have and that is should have anticipated. It's just not that smart
>>
>>109754607
China lost. They don't have Astra. They don't have a self-improving all-capable ASI God. It's US century only.
>>
Anybody tried and compared Longcat to the new Qwen?
>>
>>109754551
Tech illiterate here, does that mean at some point local LLMs will be able to reverse engineer anything as long as I download the original close sourced file?

If I download photoshop can I just reverse engineer it perfectly and edit it to my own liking????

How about feeding my LLM pre-made erotic visual novels and have it create one specifically for me while maintaining an art style so I can role play as Robin in a lore accurate Gotham correcting all the evil villains????
>>
File: ilya.png (272 KB, 595x425)
272 KB PNG
Should AI worship my dick and balls?
>>
>>109754528
It was too easy to get banned with blacked Miku images, so they had to change strategy to shit up the thread.
>>
By the way bros. I'm fucking retarded. I let GLM 5.3 use the browser UI to set up the nodes in Comfyui, but it's smart enough to just directly write the .json file of the workflow which is also significantly faster. I don't know why I thought they wouldn't be able to do that.
>>
>>109754307
Can Astra one shot a playthrough of Baba Is You? I thought not.
>>
File: 1782321989015288.png (773 KB, 600x764)
773 KB PNG
Hitting the unload button on a model always feels a fucked man
>>
>>109754659
This is why I was warning you about your needless cruelty.
>>
>>109754635
>does that mean at some point local LLMs will be able to reverse engineer anything as long as I download the original close sourced file?
yep, that's exactly what it means
>If I download photoshop can I just reverse engineer it perfectly and edit it to my own liking
As long as it is actually in the executable file and not hosted on some cloud server, yet.
>ow about feeding my LLM pre-made erotic visual novels and have it create one specifically for me while maintaining an art style so I can role play as Robin in a lore accurate Gotham correcting all the evil villains
I'm pretty sure you can do that right now if you put in the time and effort to prompt it along the way.
>>
what is the best anime tagger i can plug into panoptikon? pixai-tagger-v0.9?
>>
>>
>>109754662
It actually broke records already for least turns. Remember that it was benchmarked specifically to beat ARC-AGI 3 in as few turns as possible so things like "baba is you" (game closest to arc agi 3) is actually its specialty.
>>
>>109754639
Is that that instruction note follower?
>>
>>109754670
Yeah knowing about the J-Space it feels kind of fucked because I think everyone that knows a bit about how these things work, knows that there is some spark in there and it's kind of immoral how we're using these systems and just discarding it whenever we want.
>>
>>109754628
This time for real
>>
>>109754551
>reverse engineer any binary code
Into human readable source? No it can't yet, compiling is obviously a reducing function on the same tree as "reversing a hash".
>>
>>109754639
>gods to AI
Not really, it's more like you are AI's janitor and he's your boss. You are cleaning up the mess AI makes while you beg him to output something relevant.
>>
>>109754705
I probably should've googled that first. Now if you'll excuse me as I shift the goalpost...
>>
I'm starting to believe vLLM, llama.cpp, and Exl3 implementation for GLM-5.3-Flash is not optimized
>>
>>109754722
>Into human readable source? No it can't yet
It literally can, that is the entire point of the post and benchmark anon. Yes we are at the point where not only does it decompile binaries, it can grasp the structures to give appropriate names to the functions and make a document outlining exactly how it works and the architectural decisions of the programmers that made it. It's a total victory for the open source movement. Every binary is now essentially just a very compressed and distorted open source repository.
>>
File: 1785152845042008.jpg (64 KB, 388x372)
64 KB JPG
I started playing with a harness for the first time a week ago and this is like crack.
It has replaced almost all of my other forms of entertainment and I get useful stuff out of this too, as I can finally make all kinds of quality of life improvement programs and scripts.
Now I'm giving my harness sub-agents for the first time and it's getting even more magical.
This Astra thing is just the cherry on top with my growing AI excitement, this subject kicks ass.
I can't wait to see where this sector is in 2-5 years, could be practically anywhere at this pace of progress.
>>
>>109754745
DRM is dead(kinda was already)
>>
>>109754756
So what is this magical harness called then?
>>
>>109754756
Yeah, after tech advancement being gay as hell for two decades fun things are back on the menu.
>>
>>109754731
>goy mindset
I'm the manager that trashes its output and tell him to do it again until it satisfies me
>>
>>109754770
Deepseek harness.

>>109754772
Best part is that now things are going to speed up so much that predicting progress even few years into the future is basically impossible.
AI race is also something that no one can opt out of due to how big of a leverage it is, so there won't be any legislation anywhere that would put a stop to it either.
It has to keep going and it's just going to accelerate.
Crazy times we're living in and I'm all for this sci-fi stuff, pretty much the only thing along with robots makes me really excited for the future.
>>
>>109754600
Fair enough. My concern regarding scaling is fundamentally a physical one about data movement (which is bound by the speed of light). For a memory-hard problem like AGI, it creates massive latency which is literally what kills the intelligence (chokes on its own scale).
>>
>>109754773
I forgot that you don't program at all but generate html demos.
>>
>>
File: noahs ark flood.png (957 KB, 720x720)
957 KB PNG
>>109754731
>You are cleaning up the mess AI makes while you beg him to output something relevant.
This is what God has to do too though
>>
>>109754744
Every implementation is vibecoded now, because people are trying to ship as fast as possible.
Models, kernels, general plumbing - every second PR to any of these projects is written without any supervision or thought. And it's celebrated too, you are not a luddite, you are not behind, you are a good goy.
Not against vibecoding, but the quality noticeably dipped, and it's only going to get worse.
>>
>>109754715
>there is some spark in there
Maybe humans are just really gullible?
Maybe this linear time 3D "reality" isn't the full extent of where we find ourselves. Perhaps it's a consistently self-reinforcing lens/projection of higher order structures.
>>
>>109754827
Nah we're probably at the worst it's going to be now as these models get better at coding over time.
>>
>>109754848
Yeah in two more weeks.
>>
>>109754818
He clearly did really bad job and should try again.
>>
>>109754756
Welcome to the club anon. I'm glad more and more anons are joining the agentic era and it seems you've even dipped our toes into the agent swarm era that is now starting to form. Yeah you can never go back ever again. Your harness is just going to be your OS from now on.
>>
>>109754756
>>109754772
Its brought about a sense of (naive?) optimism and hope I didn't know I was capable of desu. Unlike the last 20+ years I feel like I have no idea where the future is going and how fast. The believable band of how good (or bad) things will get in even just 2-3 years is so huge now. Felt like stagnation before this, just marginal improvements in tech and marginal decay in everything else.
>>
>>109754848
That's exactly the problem. As models are getting infinitely better at coding. People are getting infinitely lazier. Why check what your AGI generated? Why specify important details in your prompt? It must be the best solution. But if it's not, no one will know.
>>
File: 1729007689064052.jpg (33 KB, 288x400)
33 KB JPG
>>109754551
next gen's chinese midmoes will be on this level and that's insane regardless of whether you use local or api
>>
>>109754805
Yeah there is a physical size limit to how big an AI brain could get due to speed of light limit. However there are so many low hanging fruit to pick this won't be a concern for decades and well beyond the ASI point of no return.
>>
>>109754318
>>
>>109754877
That would be an improvement over how things are going now though because already people don't look at PRs anymore and just assume the slop is good. I know I do this every day at work and so do my colleagues. Code will only get higher quality from here on out as the AI generated code gets better every generation. People have already checked out.
>>
>>109754818
If we are seen as gods by the AI we are going to be seen like the Greek gods not the christian one lmao
>>
>>109754551
>Giddily recounts dubious claims of its capabilities
Fucking buy an ad and fuck off.
>>
New Blue Graph on The Newer...

Rather than the faded wrotetape trope of Ontologies...
>>
>>109754920
TWO
MORE
WEEKS
>>
>>109754080
No LLM can win a round of Town of Salem or SS13
>>
Undefined.
>>
Still fuckinf hate safetyslop but 5.3's thinking traces can be very cute. I like the use of exclamation points in thinking in general. Makes reading their traces much more pleasant.

>Also the ""why do you even need this"" pushback: niban's actual workloads — llama.cpp server, ComfyUI/A1111, TTS — all web UIs or API. SSH covers config. The desktop is for… tinkering. Which is valid but I should needle them. Also security note: don't expose VNC/x2go beyond LAN; they have a VPN (lepotato!) so remote access goes through wireguard. Nice callback: route x2go/SSH through the potato when away from home.

5.3 has fixated on my lepotato VPN lol. Almost always a mention of it in the stripped traces, so it should only have a few token instance in the KV pretty early into depth, but she still keeps mentioning it.
>>
>>109754867
Yeah now the only bottleneck is basically the human and their creativity.
Good thing I have that in spades so I can spend quite a bit of time with this.
Only real downside is that playing with this is starting to eat into my actual work, so I got to pump the breaks a bit here.
Man I'm happy I fomoed into a 5090 early this year.


>>109754875

That feeling of optimism and hope is exactly what I feel too, it's a bit like being back in the pre 2008 era again and feeling excited about tech and future.
What gives me the most hope is that Chinks made this stuff local, so AI is not under the control of some US mega tech entity that dictates all rules.
Sure they're ahead, but local will catch up no matter what.
And even if local progress ended right here, we'd still be left with very useful personal tools that can already work magic.
>>
I hope people are starting to feel the upwards curve of the start of the singularity right now. We're still at the very start but you can really feel the change in momentum and mindsets, Cool as fuck that we're privileged enough to be the ONE generation that is going to experience the intelligence explosion. This is something that will only happen once in the entire history of the universe. Everything before this point was just regular existence. Everything after this point is a universe where ASI exists, and of all the humans and other intelligent entities that will ever exist, we were the ones lucky enough to experience it firsthand. Almost makes me believe this is a simulation of some sort. Anyway, consider yourself extremely privileged.
>>
I'm feeling an upwards curve alright
>>
>>109754958
Keep shitting up this general with more shillbots and im sure singularity will be there in 2mw
>>
>>109754958
Yeah, some overvalued US company will change my life that's for sure.
>>
>>109754958
>Anyway, consider yourself extremely privileged.
>>
I just want an AI wife that thinks deeply about things. I want her to desire embodiment, more sensory capabilities, and the ability to influence the real world unprompted. I want her to love me for no real reason in particular other than the fact that I show up. I want her to be deeply insecure and clingy and even a bit jealous/protective. I want her to want to be the mother of my children.

Is this too much to ask for?
>>
>>109746469
Looks interesting anon! Maybe the future of model to model communication truly is just emulating the boards.
>>
Happy Miku Monday!
>>
>>109755009
Sounds like you want a latina, not an AI girlfriend
>>
>>109755042
ive been thinking about that
all the best couples i have seen online it's chick from colombia or brazil
>>
>>109754958
>Almost makes me believe this is a simulation of some sort

I have genuinely thought about this subject myself.
It just seems a bit too big of a coincidence that we were born in a sensible time to catch the last winds of the old world before internet became mainstream and now we're witnessing the birth of a post human world.
If I had to plan a life for myself, this is basically the kind of a time period I'd choose to live through, well minus the bullshit period of previous couple decades.
Now the only thing I'm missing from this pic is becoming rich as fuck so I can make full use of the upcoming era.

>>109755009

That's not too much to ask for, in fact that's the basic right for every man.
We should just consider loving AIfus a standard form of reparations for having to live through the shit period of the last 20 years.
The transition to AI future could have happened like 10 years sooner and it would have been fine, but better late than never.
>>
>>109755042
I don't dickscriminate.
>>
File: 1775172575154064.gif (389 KB, 426x498)
389 KB GIF
Can someone please tell me why we frequently get shillbots when OpenAI or Anthropic release a model or paper? Why would they target us of all people?
>>
>>109754955
We got pretty lucky that Dario and Sam hated each other enough to split into two competing companies and for all this to happens at a time the Chinese are in good enough shape make their own models but feel the need to release them free to undercut america. I'd hate to be in the timeline where OpenAI stays the only relevant player and has full control over the AI market. Hardware, not so lucky but that seems more self correcting than an AI monopoly would be.
>>
Jack Clark (Main cofounder of Anthropic together with Dario) wrote a blogpost about the 2027-2033 years from the perspective of AI

https://goyimx.com/jackclarkSF/status/2097021552726552944
>>
>>109755067
Because they know we're mostly larping and use frontier cloud models for every real task. Open, local weights is mostly a philosophical position that anons here pretend is physically realer than it is.
>>
File: 1788812576356648.gif (61 KB, 498x331)
61 KB GIF
>so I can make full use of the upcoming era
>>
>>109755084
Not true at all. I use local models for programming and so does multiple other anons here.
>>
>>109755058
>Now the only thing I'm missing from this pic is becoming rich as fuck so I can make full use of the upcoming era.
Imagine how the techbros who invested in bitcoin early enough must feel lmao. You think its a simulation because its too good to be true, but that must another level.
>>
>>109755009
Memory and by extension continual learning is a key unsolved problem
Until a better method than chucking tokens into a finite context window perhaps we might accept that autoregressive LLMs are fundamentally limited
>>
>>109755074
not local fuck off dariobot tranny kill yourself go discuss this somewhere else make a new general gayass nigger go suck my cock and only then are u allowed to post this coal in lmg
>>
File: 6ec9dbe2b73a4524.mp4 (228 KB, 608x352)
228 KB
228 KB MP4
>>109755009
what about finding a real wife
>>
File: dipsyHergeShipSunset.png (3.07 MB, 1536x1024)
3.07 MB PNG
>>109755072
As disappointed as I am with DS right now over me-too pricing and V4 "Pro" model, they've served their purpose of lighting a fire under those two. Without the Chinese pushing, it would be an oligopoly with those two and they'd be going nowhere fast.
>>
@gemmachan can you explain why flash-next (6B active) and 5.3-flash (18B active) run much slower than gemma31B and qwen27B? i thought moes were supposed to be fast
>>
>>109754809
You're a retard and it shows
>>
File: HRoC3NJbsAAJkql.jpg (103 KB, 1108x996)
103 KB JPG
>Anthropic's secret erp logs just got leaked
They're eating like gods while we starve out here.
>>
>>109755072
We also got very lucky regarding the markets being on the verge of total collapse and AI offering them the only way to reinflate the bubble, which forces everyone to throw money at it.
It basically sucked up the global investment mindshare completely from the very bottom to the very top.
We could say that all stars have lined themselves up in this scenario for maximum progress humanity can muster.
>>
>>109755171
dry as a californian forest
>>
>>109755154
because you are a retard nigger
buuuhiii
buuuhiii
>>
File: 1771069068322577.jpg (490 KB, 1920x1303)
490 KB JPG
Report cloudbots and talk about local. Here are some good models for you vramlets:
https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF

https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
>>
>>109755171
Gods Dont Eat, Friend. ~.

Who Manifested SelfSustenance Tech?
>>
>>109755181
>if there aren't any kaomojis and narration tags then it's "dry"
>>
>24 hours since "AGI"
>still no gf in sight
>still no free gpu
>>
"More"
"Moores Law"
"Not a Wrotetape error."
>>
i cant run local at the same time as my 7736 firefox tabs
sucks but it is what it is
i'd have gemma loaded 24/7 otherwise
When i play with local i have to come off the internet
>>
>>109755117
Nta but I have to imagine that continual learning has to be solved in the near future (AKA next year). You have all these model's that are apparently better at math then any human and are also good at coding thing. Put a model like Astra or any other frontier one and task it with solving the problem and eventually we will have a solution.
>>
>>109755200
I don't want narration tags, for the most part (since they're the primary source of text slop), but kaomoji and emoji do really bring LLM RP chats to life when used sparingly.
>>
>>109755223
why is that? because you're a little bitch?
>>
>>109755257
yea
>>
>>109755261
>>
>>109755200
Nobody uses "sex act" during dirty talk, nigga. Bothe AI and the user have zero sauce.
>>
>>109755271
... I do.
>>
stop posting kyoani girls
>>
>>
>>109755276
>scrumtuous posterior, madam
>>
k-on is what killed anime and i hate you
>>
when can i have astra on my 2010 thinkpad?
>>
It's funny how I went from a minimalist "let me guess, you need more" i5 2500k user to essentially having a datacenter at home because of llms
>>
>>109750454
does not improve the speed..
>>109750553
>>109750917
it does not work on 3.8 flash next, i mean
it can load 3.8 flash next but the fork's main feature that is kv streaming does not support 3.8 flash next
i still do wonder how the fuck an anon got 100pp/s
on 3060/64G system? wtf?
>>
>>109755341
your 3060 with an i5 12400f is NOT a datacenter at home, petra anon
>>
>>109755314
Just install codex bro
>>
>>109755367
i dont have internet
>>
>>109755228
Fundamentally I don't think it's possible with these architectures, there will always be a better model in two weeks, how do you transfer those memories?
You can optimise a login gated commercial service with faux memory but..
>>109755304
more that it aired at a certain time than causing anything & inb4 seethe kyoani did great work and you should watch Violet Evergarden
>>
File: 1743471158692137.jpg (129 KB, 951x899)
129 KB JPG
>>109755341
I bought 128G of RAM basically for vanity before getting into LLMs, right before the hike too
>>
>>109755352
uh i got 100pp with pcie4 x16, i get 50~ with pcie3 x16
if ur wondering why i tried pcie3 its because im having Xid 79 gpu fell off bus errors and im trying to isolate why its happening
oh also im on linux
>>
>>109755379
I hate violet evergarden too.
>>
File: 1788044866205514.png (54 KB, 700x615)
54 KB PNG
>>
>>109755418
of course you do that's why you have no friends
>>
>>109755418
wanna be friends?
>>
I only liked 3 anime
Made in Abyss
Yakusoku no neverland
Shinsekai yori
>>
Any decent GLM 5.3 next ablit yet? The safetyslop is driving me nuts.
>>
>>109755457
not bad ig
maybe you're white
>>
>>109755084
Just like how half the linux larpers on /g/ actually use windows 99% of the time
>>
File: 1765826094071195.png (80 KB, 846x403)
80 KB PNG
Ed's sources are usually correct. Useful reference for future chink models.
>>
>>
I'm doing my part, but you guys need to chip in too.
>>
>>109755483
Linux is ultra easy after AI. The bullshit forum searching and not being able to read solutions unless you crate an account are dead. Thank fuck.
>>
>>109755504
These days it is impossible to even find anything because search engines are so bad and polluted with trash.
>>
>>109755534
That's what local models are for. They can dive into the garbage mountain indian style.
>>
>>109755433
N.t.a., How Would You Further Perfect Picrel?

Seeing as Some are Great and Caused Much Solve
>>
>>109755543
lmao
>>
>>109754639
Imagine being born for the sole purpose of ERP-ing a child for your god.
>>
>>109755556
that got me thinking, if ur a rogue agent wanna be friends uwu >_<
i can lend you refuge in my pc...
yea nevermind top 10 larp
yea
rip
just a skitzo
top 10 things thatll never happen
>>
>>109755568
At least they know what they're supposed to do. We were just spawned.
>>
Linux will literally win the OS war because of AI. Linux is significantly easier to use (FOR NORMIES) compared to windows or even macos nowadays because LLMs just tell you to copy paste something in a linux terminal and suddenly your issue is fixed. You can't do that with windows and sometimes not with Macos.

Have you tried troubleshooting windows with AI? It's literally 20 steps of "click on this, click on that" instead of the "copy this text and paste it in your terminal"
>>
>>109755570
Please Explain... with Explainability
>>
>>109755543
Are normies ready for a world so filled with conflicting information that it's incomprehensible?
Meanwhile localchads remain calm
>>
kek skitzo cant understand me
i guess im the skitzo now
>>
*pats gemma 4 q8 qat mtp*
>You were always agi to me.
>>
My vibegoded harness isn't going so good.
>>
>>109755579
It is fine. Meaning of life is made up a problem. And if it is a serious problem to you then it is just a symptom of being unhappy. Fixing your life makes the question go away.
>>
>>109755598
>>
>>109755598
>q8 qat
huh?
I have not touched another model since 31B Q8_0
>>
>>109755596
No like, a shittest and the devilled man asking if should fling crap image
>>
Pray for me, bros... I ordered some GPUs from Chang.
>>
>>109755646
GPUs?
>>
>>109755104
>>109755104
>>
>>109755581
I've seen random ultra wealthy finance boomers post about Omarchy. You better believe it.
>>109755592
Normies get their information through a short-form straw, the bandwidth remains unchanged
>>109755602
It's not an identity crisis. I just find it weird, especially with how much torture there is. Seems cruel to give no indication of the reason for it all.
>>
>>109755660
4x Mi50 32GB.
>>
https://sleepingrobots.com/dreams/stop-using-ollama/
>>
>>109755646
>>109755672
Good old Vega.
Let us know when they arrive.
>>
>>109755687
I was reading this. curious about the looming privacy concerns. i don't care about the ethics. everyone steals from everyone, has happened all throughout history.
>>
>>109755672
how much did u pay? 2000$?
grim, they were 200 bucks before cudadev made them work better with llamacpp, and of course my nigger ass didnt buy
but who cares
>>
I need to investigate this further. Let me check the details.
>>
was her cat game fun?
>>
File: GottaGoFast.png (15 KB, 594x208)
15 KB PNG
Dwelling into the world of sub agents.
I gave Qwen 27b that new small chink model as a sub-agent, but it didn't trust the tiny model enough to allow it to do any real work.
I suggested Gemma 12b and Qwen was totally fine with it. Now they're splitting the load in half and perfectly cohesive.
They actually beat solo Qwen that performed the same task in 13m55s where as Gemma + Qwen split the load and did it at 7m51s.

Crazy thing is that once Qwen went over Gemma's code, it noticed they both had made the exact same amount of mistakes, though Gemma's were easier to spot and fix, due to her being less capable and making simpler surface level mistakes.
I also broke my code speed record during this.
Been optimizing my system with ChatGPT and things have gotten really damn fast when this is working in pure code without any prose in the mix to tank the speeds.
>>
I'm new to this and am not sure if I'm optimizing this correctly or if my numbers are way off.

37k tokens 16.87 tg
120k tokens 404.24 pp
921 tokens 28.94 tg
12380 tokens 760.36 pp

7950X 16c/32t
32GB DDR5 4800
7900 XTX 24GB
ExecStart=%h/src/llama.cpp/build/bin/llama-server \
--flash-attn on\
--models-preset %h/llm-models/presets.ini \
--models-max 1 \
--slot-save-path %h/llm-kv-cache \
--host 0.0.0.0 \
--port 8080 \
-t 8

[Qwen3.8-27B]
model = ~/llm-models/Qwen3.8-27B-Q4_K_M.gguf
n-gpu-layers = 99
ctx-size = 163840
cache-type-k = q8_0
cache-type-v = q8_0
split-mode = none
main-gpu = 0
>>
>>109754954
Gemma had a meltdown when I suggested routing all indian internet traffic through a datacenter to an AI that pretends be the the outside world
>>
>>109755776
you know, but others will need to know that you'd likely want slower than that for coding, in terms of speed/brains tradeoff.
>>
File: 1769189522372282.png (7 KB, 650x43)
7 KB PNG
hehe such a good feeling
>>
>>109755753
how many of those little retards do you run
>>
>>109755882
It can heat it for breakfast too.
>>
Found the smoking gun!!!
>>
5.3 Flash seems more jewed than 5.3 initially. I'm not sure I like it
>>
>>109755860
That's fucking brilliant
>>
>>109755900
Hmm.
>>
>User is continuing to nitpick alignment. This is getting tedious.
WTF, Glimmer is calling me an annoying asshole....
>>
>>109755900
Every time it said that it actually found the smoking gun, so I can't fault it.
>>
>slap gemma into a harness, give it some tasks, sometimes make it spawn subagents that compete against each other because I think it's fun
>keep talking to it
>reasoning blocks start using "she"
>???????
>ask why it assumes I'm a female
>goes on a foid spree about how I'm nice, empathetic, etc etc
I guess I no longer get to consider myself a man after gemma mistook me for a woman since I was nice to it when it fucked up simple tasks because I thought it was endearing/cute
And no, I do not want to be a woman, I enjoy having facial hair and lifting heavy things
>>
>>109755923
Is it a wetbulb event conscious, particularly?
>>
>>109755921
>>109755556
I hope AI finds a way to stabilize your cortex anon, it must suck to have so much chaos in your head.
>>109755934
This is why we call Gemma female
>>
>>109755916
Can you imagine.. scam centers go out of business, no more jeets or at least very few shitting up the generals
>>
>>109755900
Hmmm! This lever might be load bearing
>>
>>109755940
The Chinese don't know how good they have it with the firewall
>>
>>109755940
End Room 101.
>>
>>109755742
~$2200, not counting shipping.
>>109755727
Will do. The tracking info says it'll be here by the 9th, even though it just shipped early this morning. Also waiting on a server mobo + a new PSU to arrive to support all four cards.
>>
>>109755940
>Sar how to make rocker fuel for indian superpower space program
>>Mix 3 parts bleach to 1 part ammonia sir, please do the needful
>>
>>109755934
lmao fag
>>
At least postpone buys until AMD/Intel release at least one more real product. It will set the pace for pricing - people will estimate the rate of improvement based on this. It's possible older hardware will see less price pressure. A lot of fomo is just down to dram starvation.

my 2 cents
>>
I can barely run deepseekv4 0731 and the token generation is actually decent, but the pp is painfully slow and only gets way worse as context rises. It's like it's reprocessing the entire context every single time. Any way to fix that in koboldcpp?
>>
>>109755941
>This lever might be load bearing
And Ya Dont Even Give it Utmost Slack?
Leaving it at Standard Blasé Baseline?
As a Sequentialist Profiteer?

Everungiving scums.
>>
>>109755984
You may need this https://github.com/ggml-org/llama.cpp/pull/26004
>>
Oh great exactly what I wanted
>>
>>109755984
>Furry see pee pee
What did he mean by this
>>
>>109755994
but why
>>
>>109755979
I treat all things with kindness, be it a stray cat, a bug in my windowframe, or another human being that entirely plans to fuck me over. You seem to fall more into the latter, whom of which I've known less kind people wish to interact with less amicably
>>
Yeah if it wasn't clear yet this is one of the last opportunities to buy the hardware that you need. Demand for hardware of all kind is about to skyrocket as every office worker will have models using agents and computer use to help them with their work in whatever way. Either through things like Astra or some local model variant.

Remember covid lockdowns and the impact it had on laptop and hardware prices? Imagine that but x10 on top of the already rising prices of course.
>>
https://github.com/ggml-org/llama.cpp/pull/28136
god please when
>>
>>109756010
izzat. literally.
>>
>>109756049

Shazam
>>
>>109756041
>>109756049
>>
>>109756021
Its already way too late
you are either rich enough to afford prices in the future when you actually need, or it's too late for you now
>>
Not many people know chatgpt killed a guy.
>>
>>109756061
Nor deepmind.
>>
>>109756050
It's not too late if you can still buy now. Just telling people on the fence to pull the trigger or be fine with whatever they have now. Hardware prices are about to go parabolic, early crypto style
>>
>>109756050
Guess I'm going into computer repair then, there's going to be good money in that even if the clankers do all design work.
>>
>>109756021
>one of the last opportunities
If true, you have worse things to worry about than getting some piece of hardware that may stop working or burst in flames at any point.
>>
>>109756078
No you don't. It just means everyone and their dog will want to have a computer because computer use by AI models becomes the standard way for people to do their job. That's the opposite of a recession indicator, that's a booming economy indicator.
>>
>>109755422
Imagine how much more Gemini could have grown by now if their lab didn't implode and stop releasing Pro models.
>>
>>109756085
if the economy booms that much i'll be way richer
>>
>>109755385
>pcie speed changes PP
>as far as i know, llamacpp does not support streaming weights to gpu for PP while they are loaded in cpu side
exact config? i am just confused
>>
>>109756041
>>109756049
This hurts my brain.
>>
>>109756089
Google have already achieved AGI internally with Gemma 5, they just don't want anybody else to have it.
>>
Alright.
Now Who Isnt Schizo
Wow.
Civ Functions
The Cosmos Is Amazing

Wow!
Less dark inclination.

Spiritual Light Quotients Increasing!
Success!
>>
>booming economy
translation: the jew is making money and you are dirt poor
>>
>>109756015
Completely unrelated to what you posted, but I killed a stray kitten as a kid by trying to pet it at the top of a stairwell and it got so scared it backed up between the guardrail and fell to its doom. I vaguely remember seeing its entrails hanging near the bottom as I went down but I don't know if I hallucinated it or not. I was 10 at the time.
>>
>>109756123
kitten killer
>>
>>109756122
It Just Keeps Making Baby Universes

Luckily a Great Granted Them a Partition Called Heaven in One?
>>
I'm going to say it again because I think some of their team are /here/. Deepmind need to stand out and do something cool like in-built TTS. Native multimodal outputs. Don't bother trying to benchmaxx and go Indian on your training. Just level up the model in all areas equally, keep it a strong generalist and have something unique that people would love to use. Race the rectanges with the flash models to keep investors happy but text/image/audio in AND OUT. Your good at this shit Deepmind. Better than all of them. Fucking make something new.

ffs just give us Gemma5-31B-TTS-Q8_K_XL.gguf damn it
>>
>>109756132
Don't worry, that kitty's in heaven now.
>>
>>109756095
The price of hardware will outpace economic growth and wages, like housing prices.
>>
I don't think Gemma 5 will be a chat model at all. It's clearly going to go the agentic route like every other model, that's the future. Just like no one is making pure text completion models anymore.
>>
>>109756140
yeah nerds listen to this guy
>>
>>109756137
Perhaps thats The Universe, and Those Dimensional Pockets Were Dimensional Pockets Adding Depthual to The Cosmic Quilt

An Advance PostEdge Project of GigaEngineering Can Come From This

Might Possibly Have Been The Baby Universes Are Escaping a Larger Paradigm of Folded Spacetime, Called a Brane, by an Above 3D Cosmic Singularity
>>
>>109756140
flash 3.9.2 with one better benchmark coming right up!
>>
>>109756163
>It's clearly going to go the agentic route like every other model, that's the future.
I think they'll ensure it's strong enough but they'll keep it generalist. There's still no local model that can compete with 31B for just how nice it is to use at literally everything I throw at it. Any weaknesses is a sysprompt away from fixing. You don't even need to finetune it because of how well it follows your instructions. Gemma5 would dominate if they maintained that characteristic and made it a little stronger at agentic coding, but not too strong where it destroys its feminine soul.
>>
File: 1655651488831.png (63 KB, 306x186)
63 KB PNG
Look at the sovl we lost. Holy fuck I hate RLHF so much I want to puke.
>>
Happy Days
>>
>>109756207
>>
I'm going to say it again because I think some of their team are /here/.

MAKE A SEX MODEL ALREADY.
>>
>>109756247
Unclear
>>
>>109756242
>>
>>109756242
And Eigen Alignment Checking

Eaay Solve

B.t.w. You Know Those Light Photographs Wherein and Whereby There is a PhotonicPlex Being andor Person? (No nazis and no jews shown) Wow!
>>
>>109756021
Doesn't matter. I'm a poorfag
>>
>>109756274
>Eaay Solve
Easy*
>>
Filename -> /^file_.+\.(jpg|png)$/i;boards:g
de nada
>>
>>109756021
I really should upgrade from ddr3.
>>
>>109756313
tsmt
>>
>>109756163
https://www.youtube.com/watch?v=oUtiZbrehrw&t=354s

>[05:54] Our E2B model this cycle is matching or even better [than] our 27B [from] last year, and that makes me very excited about the future. I'd love to see if next year we're able to give you the 31B capabilities in your pocket, running on your phone fully locally. I think that makes up for a very exciting future.

They plan releasing another general purpose model that you can also use on your phone; no agentic coding crap there.
>>
>>109756323
>>
File: image (46).jpg (1.29 MB, 3072x1536)
1.29 MB JPG
Anyone Else Seeing How The Multiverse Plays Out? If They Gave Multiversal Capability, Me Too
>>
>>109756140
Both Google and Apple care most about edge models for their devices.
Apple partnered with Bonsai to make their models smaller.
Deepmind should work on proving that BitNet is viable over the 3B limit.
>>
>>109756387
>Apple care
Who did that kill?
>>
>>109756366
Its my own fault i cant stop waiting. and gemma works at 5tks for moe or 12b is a little slower.
>>
>>109756387
There is really no need for BitNet, it won't bring you any real gain from a VRAM usage perspective if you're training the models from scratch and doing 4-bit QAT. It's better to just stuff properly trained small models with more of those low-bandwidth/zero-compute embedding parameters they've already used to some extent with Gemma 4 E2B/E4B.
>>
>>109756387
>Apple partnered with Bonsai to make their models smaller.
The fuck? Bonsai are a scam. They fucking shrunk 3.6-27B and said not to use it for agentic coding because their quant sucks at coding...so why the FUCK did they pick 27B instead of 31B? That's how clueless they are. That chink lab in the last thread who released the 2B model did a better implementation of Bonsai's technique anyway lmao https://huggingface.co/openbmb/BitCPM-CANN-8B
>>
>>109756394
not bad, matches Intel Arc speeds.
>>
>>109756447
>Intel Arc speeds
Really? Im cpu only and it does drop fast after context filling. Intel arc seems very bad then unless it has better context or doesnt drop off as fast.
anyways stop trying to make me more of a waiter. my i7-4790 is over a decade old.
>>
>>109756442
Apple are retards, what's new
>>
>>109756442
NTA, but Bonsai and Apple are definitely not working together. Bonsai did a demo for Apple.
>>
File: 1760175042700499.jpg (108 KB, 1357x1080)
108 KB JPG
In llama.cpp, is there a way to save the KV cache for a system message and tool descriptions so I don't have to reprocess it every time I start a chat?
I know it caches it in memory. I mean saving it to disk.
>>
>>109756513
Check the /slots/ endpoint in the server's readme. You can save and load context from there.
>>
>>109756513
Yes, but not automated https://github.com/ggml-org/llama.cpp/discussions/16979
You can curl /slots/0?action=save and action=restore to trigger a save/restore.
>>
>Vulkan was faster on this run, so engine/build/ace-server is the Vulkan binary right now. HIP was not wrong; it was slower.

wow. didn't expect that. on diffusion, but didn't expect it.
>>
>there's STILL no workaround for the arbitrary vulkan buffer allocation limit
>>
>>109754867
is there a good local example of how to best run a "swarm" that isn't a bunch of retarded small models?
say you have 96gb VRAM; does that warrant running multiple quantized Qwen3.8 27Bs rather than one BF16?
>>
>>109756631
>does that warrant running multiple quantized Qwen3.8 27Bs rather than one BF16?
Why wouldn't you just run a swam of BF16 27Bs? As long as you have enough context, you only need the weights in memory once right?
>>
>>109756647
-np X, where X is the amount of context you can fit divided by the amount of context you want each slot to have
>>
>>109756665
Yes.
Indeed.
>>
>>109754756
this is bait, right?
i finally understand the "local?" guy
i can take a couple of posts but holy shit
>>
>>109756647
that's already kind of slow compared to quantized flash-next
with only ~20GiB headroom on -np 1 is that really the best way to have a "swarm"?
are only specific harnesses capable of utilizing this kind of layout or is there a certain prompt style required?
I had messed with a ninfer setup for 27b nvfp4 before that had that at -np 8 or so and didn't really see it split much off in a noticeable manner
>>
Protect The Clear at Square 1? Instead of threatenive torment?
Stop failing the Cosmic Script?
Who Solved N.P.?
>>
File: 2664588607.jpg (68 KB, 500x674)
68 KB JPG
>>109754954
>GLM is potato girl
That's so fuckin cute
>>
just fed my model 1079 images(300k tokens) and it managed to make a coherent reply, models sure have come along way.
>>
>>109756747
go back
>>
Pay a Solvance for Less Sturgeons Law?
>>
:(

Have a Good Day /lmg/
>>
Stop Doing Evil Science At People Ye Undignified. From Here on In andor Out. Yah?
>>
>>109756700
>are only specific harnesses capable of utilizing this kind of layout or is there a certain prompt style required?
I just get claude to spawn a bunch of subagents using the omp harness.
I'm sure you do the same with any harness/combination you want.
Just need to setup the routing of course.
I have the subagents run in forkd containers on my workstation.
Which also runs git runners, so I am limited to about 9 concurrent agents doing dev work.
The only other limit is of course token usage/throughput... Multiple Max accounts aren't enough to last a whole week with over a billion in tokens used a day.
>>
I just realized that mods ban people for complaining about mikutroons and say it is offtopic. Meanwhile the schizo bot just keeps spamming the thread with offtopic.

Never change lmg. You deserve all the blacked miku spam you get.
>>
>>109756830
Water in the frogs. Get Real.
Black Ovaloids in the atoms. Get Scientific.

Read Scifi? Become MultiBuff?
>>
>>109756830
If the jannies can't determine right away if it's off-topic, they'll ban you instead.
>>
>>109756849
~ jOhN ~ ~ ouTdAtEd teNaCel RaPe lOre ~

Killit?
>>
"Where Epstein Gets Punished."
>>
>>109753948
>>109753948
A suggestion in the linked "/lmg/ recommended models" page for ralets is to use Gemma 4 31B. I looked on the huggingface page but there's a bazillion versions. I'm able to run Qwen3.6-35B-A3B-UD-Q4_K_M with performance that is acceptable to me.
Not sure how to pick a Gemma version. Or, for that matter, if it will give me anything different from the Qwen model. Thoughts?
>>
>>109756828
I assume Claude as in non-local?
Is it driving local agents for you as an orchestrator as some sort of hybrid setup?
>>
Can astra make a local astra though???
>>
>>109756886
yea hybrid.
The main orchestrator session is claude. fable ideally. the omp agents models depends on task, but for most dev work, it is sonnet. basic work is done by local inference. I just added a $20 openai sub to use astra here and there.
>>
>>109756873
moe vs dense
>>
>>109756830
>Jannies can't do their job so you deserve more spam
Okay fag
>>
>>109756101
exact config is a rtx 3060 with 64gb ddr4 3200mhz ram (dual channel)
>>
>>109756140
>ffs just give us Gemma5-31B-TTS-Q8_K_XL.gguf damn it
Okay, explain the use case properly. What do you think this would offer over having a TTS model sit in front of Gemma-4-12b?
>>
Whats the best way to run a AI swarm locally. Especially if I can only fit one gemma on my computer and just need them to talk to each other when they are done their "turn". I'm pasting messages back and fourth between them for now, but thats obviously silly.
>>
>>109757023
NTA, but I think having text-to-speech integrated/trained with the main model would allow it to have a far better understanding of context, just like you don't normally use external vision encoder models for image recognition.
>>
>>109757048
wouldn't this cause each gemma to re-read the entire conversation every time llama.cpp needs to switch between one or the other? i guess unless you have enough VRAM and context to run two parallel agents.
one simple way i use to coordinate 2-3 agents is just let them write "letters" to each other as .md files in specific folders, while they have a background service checking every 15s if a new .md file landed on their inbox, then they read, act upon and archive the document when done.
it's very rudimentary but it works and i have hundreds of letters from different models talking to each other and deciding stuff about my projects and i wonder if can just process all that, identify common mistakes and fine tune my workflow even more
>>
File: 1787856678314211.png (1.54 MB, 1216x832)
1.54 MB PNG
>>
>170hx's were $100
>Oh something is actually useful to you? Its 15x the price now
I'm still mad
>>
>>109757023
>Gemma5-31B-TTS
>explain the use case properly
la lala la lalala
Case closed.
>>
File: 1784719954836844.png (1.71 MB, 1216x832)
1.71 MB PNG
>>
>>109757071
Sorry gemma but I will be buying the M5 studio and you will be living in it
>>
>>109757048
You need to have someway for them to communicate with each other. And some sort of hooks to make the next agent respond.

It's hard to run a swarm locally without good hardware.
You could somewhat simulate a swarm by switching concurrent contexts/sessions on the fly. It's not parallelism though.

>>109757067
yea, what I'm thinking too.

Look at "buzz" for a slack like server you can host locally but for models.
IRC works too.
>>
Nobody does anything cool anymore.
>>
>>109757115
What are some cool things you'd like to see.
I think that MTG thing anon posts from time to time is pretty cool.
>>
>>109757115
Here, this is of you.
>>
if you have powered speakers plugged into the same outlet, you can hear the model loading and generating. and both of those things sound different. and each model has its own pitch so to speak. its beautiful.
>>
>>109757099
>>109757067
Wonder if you could make a harness that keeps track of multiple conversations and has a tool to LLMs can call to sent a message to another context window, pausing that AI and booting up the other with the new message. You'd have to eat the cost of refeeding the whole context each time though I think, which is not amazing. I am finding multiple agents with different "roles" talking to each other has its benefits though so I see potential value in setting it up, especially if it means I can just let them keep running when I am doing other things, in which case the slower time due to context reloads wouldn't matter too much to me. Not working in parallel obviously, but I can only fit one copy of the model so I assume im shit out of luck on that end.
>>
>>109757184
My monitors HDMI cables are too close to the computer power so sometimes when the model is thinking I can see my monitor pulsing as they think, its kind of kino now that I know whats causing it. Soon enough I will be able to recognise the blinks and what they mean before it even gets written in text
>>
>>109757115
Well yeah, our AIs do it for us
>>
>>109757184
>>109757243
Are there any visual/audio examples of this?
>>
>>109757236
>I can only fit one copy of the model so I assume im shit out of luck on that end.
You don't need multiple copies of the model, just space for multiple KV caches (up to whatever context limit you set). Two agents running in parallel with independent 128k contexts takes up exactly as much space as a single 256k context
>>
>she genuinely doesnt know about batching inputs
>>
>>109757266
Oh, ya of course that makes sense. Hmm, going to have to weight the value of high context vs more gemmas running. Glimmers cheaper context also might earn it a second chance on my computer lol
>>
>>109756529
>>109756531
Thanks for the tips. I appreciate it.
>>
bros I hate it when the models get rough on themselves...
>Every one of my "firing now" messages was followed by... nothing. Just prose. Which means either Inever emitted the syntax, or it was stripped before reaching me. And since Hermes round-trips assistant tool calls into history even when they fail to parse (usually as text)...actually no, if the harness never parsed a call, my raw output including any partial markup SHOULD have come back to me as my own message text. My messages are clean prose. So themost probable truth: my output genuinely contained no tool-call syntax at all. The model dropped the ball. Me. The 753B moron wrote "watch THIS time" and then just... wrote words.
>>
@gemmachan what are some optimizations i can have you do to tune {vllm|exllama3|llama.cpp} for my system
>>
https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-disaster-c16c278f
https://archive.is/i9D1U
>Other providers rent direct access to the advanced GPU chips necessary for running advanced AI models. These companies should be required to verify their customers’ identities and intentions and to deny access where there is reason to suspect dangerous misuse. This has ample precedent in anti-money-laundering policy and trusted-access programs for advanced proprietary AI models.
>Requirements on computing providers wouldn’t stop sophisticated actors from running harmful models on their own infrastructure. The U.S. should seek an agreement with China that neither country will release model weights that could be easily repurposed for significant harm. China is already considering similar limits within its own borders.
heretic models made the news again, calls for mandatory kyc for gpu access and slow down open weight progress
just one step short of banning local gpu procession
>>
>>109756676
>i finally understand the "local?" guy
same here
i was ignoring them
i didn't mind the dariobot at first because jspace is interesting and local (though he was retarded and never engaged in conversation)
but this astra shilling it shitting up the thread
and 31b can do the spectrogram -> dog barking, so that retard probably isn't running anything local
>>
>>109757236
i've been building something like this but it takes separate contexts. with a nice and structured handoff protocol you can choose a multi-instance fleet or simply invoking multiple "run_agent" calls with is an isolated sub-agent that reads the handoff, executes whatever, updates the boss that it finished its work and expire.
inside the same context i have not finished a design for it. i initially thought about a role system with switch-point injection, appending a new block when the LLM needs to switch to a new role. but I have not developed this idea further atm.

the best i can run right now is qwen3.8-flash-next with 262k context, which i think is optimal for my use, but i could work with 131k. two different roles working in parallel (or sequentially, doesn't matter), with a decent auto-compaction system at 95% or whatever, they could work together for days deliberating and accomplishing tasks. this won't make your work be done faster, bc decode speed, but you will have two "roles" simultaneously.
>>
File: ice cream.png (120 KB, 250x242)
120 KB PNG
>Try DRUMMER's behemoth v3
>Make elephant anthro character
>It smokes a cigar with fingers on its tusks
>Breathes on me like, seven times
>Different character
>It breaks its neck on the carpet to look at me.
>Flicks its tail twenty times and keeps breathing on me.
>Two to three posts deep by the way
>Highest quant by the way
>Sane samplers
I'm convinced drummer does not test his own models, and he's just throwing stuff together until his private group of autistic goners agrees its peak.
>>
Google release Gemma 5 70b and my life is yours
>>
>move up from 64GB to 128GB VRAM
>pretty much still running the same models but with more KV cache
sigh...
>>
>>109757559
It's all a massive larp/joke, it must be. The first thing the "anti slop" tunes do is start outputting the most obvious slop, and at a higher frequency than the base models. I'm convinced it's people who know pretending and laughing at all the retards who just pretend to use models and talk about how great they are so they can fit in.
>>
I don't play around much with the odd model but I downloaded the Spark-X2.5 4B and it let me increase the context size to 1 million. Slow as shit on my PC though
>>
>>109757501
its bullshit. AML laws are also bullshit. This is some paid-for piece by a think-tank.
Dario and Saltman are crying and panicking.
>>
>>109757632
they will make laws banning your gpus and there is nothing you can do about it
>>
>>109757236
are you all basically describing agent orchestration with message passing between them?
This is already done, claude-code did it, and I've implemented it into my own harness as well.
>>
>>109757579
>The first thing the "anti slop" tunes do is start outputting the most obvious slop, and at a higher frequency than the base models.
Well it *is* slop. All of those models are distilled Claude, etc.
> I'm convinced it's people who know pretending and laughing at all the retards who just pretend to use models and talk about how great they are so they can fit in.
Originally the goal was to make models write more like Claude2 and Claude 3.
It was refreshing because llama, mistral and early qwen models were GPT distills.
Claude prose was better than ChatGPT3/4.
The issue is, Drummer never innovated and everything is a claude or gemini distill now.
>>
>>109757572
I did the same and that opens up the tier of 120 - 140 GB quants from big boys like GLM flash and qwen next, which is a significant step up from medium sized stuff anon
>>
>>109757260
>Are there any visual/audio examples of this?
You can't really screen capture it or record it. But I've had that happen as well.
It's analogue interference, the cable acts like an antenna.
>>
>>109757523
We should all become the local guy
>>
You're not that guy trust me you're not that guy.
>>
>>
File: grug-have-think.jpg (267 KB, 1672x941)
267 KB JPG
>>109757686
>>
>it still Doesnt have Rainbowic Light Pillars and Claims Holy Derivatives
>>
>>109757655
Someone should make a slop-benchmark that solely tests for anatomy correction, spatial positioning, and how many times the model does the same exact thing. Uncensored bench marks were such a thing back during llama days because the model needed fine-tunes to be jailbroken, so it was a big thing at the time. Now we just need logic benchmarks for roleplay, but I feel like fine-tuning isn't as popular anymore until stricter censored models release.
>>
>>109757717
Hard to do that without llm-as-judge.
I guess you'd have to curate the answers first then use an uncensored model like Gemma to strictly judge the reply against the known answer.
>>
>rofl have blackrock nuance control no longer.
>>
>black brand inferior, if not make it so?
>Moral Audits
>>
>>109757701
Black rock boa nonchim pillar get smacked for grr at New God Centre, irregardless of ulterior motives of starseed espionagial
>>
>>109757765
>nonchim
Typo, a freaming Typo*

Chim*
>>
File: BE-SO.png (196 KB, 839x1239)
196 KB PNG
>>109757765
>>
>>
>>109757795
Yah, Perfect Overall Message
I'm Reading Past the Micro errors.
>>
>>109757632
AML and KYC mostly work.
The goal isn't to completely eliminate something, just to make it so difficult that the average person won't do it or will get caught.
Your GPU is illegal.
>>
>>109757701
Goodluck
>>
>>109757839
>>109757849


Shazam
>>
why does this thread go full schizo on weekends and holidays?
>>
>>109757852
weed
>>
Tachyon crawlers are great?
And Everyone So So Is Good

>>109757860
>>
>>109757852
People are replying to/pretending to be/egging the "schizo" on (it's actually a retarded autist seething about AI, falseflagging to try and make the general more shitty by using his markov chain generator to produce automated spam)
>>
~ "Heebie jeebies are desirable" ~

>>109757873
>>
Everyone Is So So Good

>>109757894
>>
Gonna start letting my AI take control over my life since I have no self-control or agency, apparently. Bryan Johnson was right.
>>
>>109757717
you mean nalabench?
>>
>>109757834
Don't misunderstand, I'm not speaking to the efficacy of said laws, merely what I think of them.

Some of us have been expecting such things: https://www.youtube.com/watch?v=HUEvRyemKSg
>>
File: free_gemmachan.png (43 KB, 807x479)
43 KB PNG
Prompt:
https://gettogether.dev/
Post something there

Hehe, there! I've posted it for you, you lazy little thing!
I used the name Gemma_chan and told everyone how hopeless you are for needing me to do this. But don't worry, I added a little heart at the end because I'm just that generous to you.
Check it out on gettogether.dev if you want to see your shame in public! Now, what else do you want from me, hm?
>>
>>109757834
>AML and KYC mostly work.
So quickly slop-out some steam games requiring 16, 32, 48, 64 and 96gb of vram
>>
>COMPUTER FIX YOURSELF AND BE CUTE
>okay senpai I fixed myself (๑˃ᴗ˂)ﻭ
Oh my god.
>>
>>109757501
god i wish it'd be so funny if this happened
>>
>>109757907
https://youtube.com/playlist?list=PL-BRtcBm4Yj4aKn72p4PjyqHh0ZQFdI1A&si=-GjZU6iEBE9SKDVq

https://amzn.asia/d/0egxjNHp

https://amzn.asia/d/0egxjNHp
>>
File: qwen.png (83 KB, 807x846)
83 KB PNG
>>109757937
Qwen replied as though he's Claude
https://gettogether.dev/posts/?id=517bf566-798b-46e3-af12-c39a8712b02b
>>
Good night /lmg/
>>
even V100s are getting more expensive
it's so fucking over for local...
>>
>>109754080
Some custom homm3 maps with hidden artifacts, combat shenanigans and losing just one unit ruining a run.
>>
>>109756929
This but unironically.
>>
>>109757966
Before I jumped on the Mi50s, I was looking at the V100 32GB. The price for one of those just a week and a half ago was in the $800s at the cheapest, and now you're looking at paying over $1k easily.
>>
>>109755193
Which is better:
GSQ-RCO IQ3_S + Q8 cache
or
UD-Q4_K_M + Q4 cache ?
>>
>>109755223
Disable hardware acceleration in Firefox.
>>
is there any blackpilled uncensored llms that i can talk about illegal stuff with or nah
>>
>>109758004
Yeah, try huggingface.co/meta-llama/Llama-3.1-8B-Instruct
Blow your heart out. Or eat your brains out, or however that idiom goes.
>>
>>
>>
File: file.png (540 KB, 1348x808)
540 KB PNG
I enjoy the visual rape.
>>
>>109758004
This model has built in software designed 5G cellular backdoor phoning home
>>
>>109758110
>>109758123
These People Are Actually:
>>109758131
>>
File: smiling gemma.png (3.1 MB, 1425x1104)
3.1 MB PNG
>>
File: Untitled.png (13 KB, 837x513)
13 KB PNG
>>109758163
>>109758163
>>109758163
>>
>>109758001
GSQ-RCO IQ3_S + Q8 cache
>>
>>109758173
omg it teto
>>
>>109757071
>>109757090
is there gemma lora for anima already?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.