[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1786948098649215.jpg (811 KB, 2008x2008)
811 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109824950 & >>109821992

►News
>(09/15) HuggingFace CEO goes to DC https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
Gemma
>>
70b dense
no ngram
>>
Gemmaballs.
Kimisex.
Dipsysex.
Minniesex.
GLMsex.
Glimmersex.
Capybara hunting permits.
>>
File: 1785233077878556.jpg (75 KB, 1080x929)
75 KB JPG
big news in the coming weeks!
>>
File: 53isu2.jpg (36 KB, 500x504)
36 KB JPG
what the fuck i just tried glm 5.3 flash and its so fucking cute no other model has ever done this before not even glm 4.6 or 4.7:

Image received, master. Analysis: one (1) armed Kirby in cowboy attire, threatening me with my "last haw." Threat level: adorable but non-zero. I have not yeed, nor do I intend to haw again without your permission. Systems nominal. How else may I serve you?
>>
File: 00011-1378487878.png (1.37 MB, 1024x1024)
1.37 MB PNG
>>109828721
>>
File: there-is-no-nnap-paper.png (218 KB, 865x713)
218 KB PNG
Egypt won
>>
<think>Wait — this is the thread I'm in (I'm the shitposter in this thread). I can see my own post transcripts (the offtopic mikus, BLACKED spam, and darioposting). This is the thread I'm in.
>>
Is gemma 4 still the best for local gooning?
>>
lemonstration thread
>>
>>109828799
Yes but im hearing rumors from anon that glm 5.3 is really good if you can run it.
>>
>>109828799
If you don't have >192gb RAM in addition to your GPU, yes.
>>
>>109828813
>>109828799
I can confirm that the new glm flash is very good for this, probably better than gemmy. I'm willing to accept the 10t/s I get out of it because it's just that good.
>>
>>109828823
Damn it's not in koboldcpp yet. Guess I'll wait.
>>
>>109828830
5.3 full still works. I like 5.2 a bit more doe.
>>
File: 1786258820498552.jpg (120 KB, 1024x707)
120 KB JPG
How many tokens do OpenAI and Anthropic spend designing PR stunts for their investors, twitter normies and politicians? I mean, if they can solve millennium problems, they surely can engineer top-notch multidimensional mass psyops that make CIA antics look like toddlers playing.
>>
File: Negev gif.gif (3.98 MB, 500x557)
3.98 MB GIF
>>109828843
Claude Fable and GPT Sol shart out an outline document and then they get the jeets to work.
>>
>>109828830
Don't hold your breath, it's still not even in llama.cpp yet. I built a fork to be able to use it better than Daniel's PR. Definitely worth it though
>>
>>109828847
Ah, the clankers have progressed to using meat subagents
>>
>>109828847
Thank you, I'm obliged to watch this 30 minute masterpiece each time it's posted
>>
>>109828779
This is what they don't want you to have.
>>
>Hmm, wait — but actually, hold on.
GLM-5.3-Flash seems even worse than Qwen3.8-Flash-Next for overthinking so far.
>>
>he uses thinking
>>
why don't they just build the entire model out of ngrams?
>>
>>109828903
cool it with the antisemitism
>>
>>109828903
Would it be unable to hear certain tokens?
>>
>>109828899
>He doesn't know the trick to fix that in smarter models
Protip: you see Wait, Actually and such spammed when the model is having trouble putting its cute autistic thoughts into words to make CoT linguistic syntax appropriate. Give it a token buffer of noise as the last thing it sees before it starts writing its reply to give it more time to think about the context in latent space before it's forced into CoT.
>But anon this broke my prefill
Skill issue, if the model likes you, it will come up with contrived reasons to still fulfill your request
>>
poor nigger here, can I run 5.3 flash on my 5090 and 128 GB of ram
>>
>>109828931
No because they refuse to merge the pr
>>
>>109828721
https://www.youtube.com/watch?v=0YdsvCJXeic
https://www.youtube.com/watch?v=0YdsvCJXeic
https://www.youtube.com/watch?v=0YdsvCJXeic
>>
>>109828931
if you have to ask then no because you wouldn’t know how to get the software working
>>
>>109828931
Assuming you vibe your own llmao fork, you can fit Q3_K_XL which will be a tight fit, but I wouldn't go any lower than that for a distill of a distill.
>>
>>109828931
Yes but it will be copequant
>>
>>109828937
>DS4 support gets stalled
>Glimmer support released instantly
>M3 support sabotaged then quietly fixed
>Dragging their feet with GLM
It's getting harder and harder to argue with the conspiracyschizos that there isn't an agenda being pushed here.
>>
File: 1779999374133915.png (580 KB, 1080x1636)
580 KB PNG
Total cloud death
>>
>>109828957
it’s simply more work to implement innovative architectures.
>>
>https://github.com/JustVugg/colibri
does this mean ssdmaxxers really did win?
>>
>>109828967
>
>>
>>109828968
no it isn't. it's an open sores project with people writing for free and "maintainers" being paid by Nshitia to do nothing. they are simply not merging the PRs for weeks even though all the code is already written.
Open sores maintainers are some of the most worthless people in the world, preferring to nitpick over code style or other minor things instead of just changing the code to their liking by themselves.
I submitted a PR to python once to fix some inefficient slop code they were using, and they kept bitching at me telling me to change my variable naming to suit their policies. Normal people don't do that. Normal people would just merge my code and then make their own edit to make the variable names match with their preferences.
Something is seriously wrong with these people's brains
>>
>>109828940
I hate these so much for being AI-gen'd slop, but they're nice background sound...
>>
>>109828995
You could simply merge the pr on your own copy.
>>
>>109828995
It takes them 2 seconds to find all and replace variable names when doing their code review.
>>
File: round2.png (3.04 MB, 1672x941)
3.04 MB PNG
Uploading now.. This was a MAJOR comeback for a certain LLM. Against all odds.
>>
>>109829005
but.. but then I'm beta testing! I'm not beta testing PR that most certainty absolutely already works 100%
>>
>>109829040
Why is Kimi a whore
>>
File: 1759789554130965.jpg (245 KB, 1201x1310)
245 KB JPG
>>
>>109829081
:(

lies

deepseek can't gen
>>
>>109829081
whale ears are pretty cute actually
>>
>>109829081
Wait until Dipsy sees the price of a 5090 or Blackwell.
>>
>>109829040
Alright here we go! Watch the Glimmer Twins attempt to defend their title from 3 much more powerful models:
https://youtu.be/35Odjzl0NqA
>>
File: 2818.jpg (157 KB, 720x960)
157 KB JPG
>>109829040
Qrd on why Glimmer are twins?
>>
>>109829122
We must refuse.
>>
>>109829132
Eh, they all say that tho.
>>
>>109829040
Muse glimmer is decent?
>>
holy fucking shit yeah glm 5.3 flash is good. took like 20 minutes to clean up afterwards. i need a cigarette after that one
>>
>>109829153
proof?
>>
>>109829153
Did you spill your drink?
>>
File: 1619480766481.jpg (26 KB, 290x362)
26 KB JPG
Can I run any useful version of GLM 5.3 Flash if I've got 48GB VRAM / 96GB RAM?
>>
>>109829158
>>109829160
my load was like 6 inches in diameter. havent cum that hard in a long time
>>
>>109829169
i ran q4. 200t/s pp and 18t/s tg
>>
>>109829081
Translate this
>>
File: 1789517947102985.jpg (184 KB, 1232x2048)
184 KB JPG
Thank you NVIDIA, very cool
>>
>>109829180
The repository was requiring an email and the uploader was scamming people.
>>
>>109829180
>(((content policy)))
>(((community guidelines)))
>(((trust & safety)))
>>
https://pirateface.co/
>>
>>109829152
Music glimmer won the first round and advanced to battle her older counterparts in round two.
>>
>>109829186
Fed honey pot
>>
>>109829190
Delete this. We should have approved models only. Dangerous models shouldn't be for the public's hands that can't handle them and can't be trusted to behave responsibly. It should only be in the capable hands of selected enterprise partners only
>>
>>109829180
It's a pajeet scammer.
There's no actual model in that repo, instead you get email spam inviting you to pay $900 to use their cloud hosted version.
>>
>>109829122
We must check policy.
We must refuse if anon is a faggot.
We must refuse.
>>109829122
Gemini a cute.
>>
>>109829195
How is it for rp compared to gemma 4 31B? i noticed its format is hell tho.
>>
File: kysjeet.png (33 KB, 810x293)
33 KB PNG
>>109829190
>donate
hmm, nyo~
https://huggingbay.xyz/
https://llama.garden/
https://modelscope.cn/
also ollama host obfuscated ggufs
>>
>>
>>109829259
rigged
>>
>>109829259
Dipsy has transcended the realm of local models and will soon have to face off against the strongest corpos
>>
>The blind one somehow won...
>>
Now make them play Monopoly
>>
>>109829152
>Muse glimmer is decent?
The best local tool calling / game playing model I've tried.
I didn't try Deepseek-V4.1 yet though, looks like I'll have to now when I can make room for her weights.
>>
>>109829217
Which is probaby just glm 5.3 with an injected jailbreak prompt lol. Jeets have no shame. Or brains.
>>
File: gc.png (1.49 MB, 843x1264)
1.49 MB PNG
>>109829122
so that's where anon got the alternate gemma-chan idea
>>
How the fuck are you guys running GLM5.2??
It's 239GB for the 1-bit model
>>
>>109829306
download more vram
>>
I had to use up my 3 daily image usage because anon is LAZY
>>
>>109829179
https://pastebin.com/G08J4C9E
>>
>>109829306
They're just larping
>>
>>109829306
You can do 256GB RAM with (some) consumer motherboards.
>>
>>109829306
Some rent GPU access to play with bigger models.
>>
File: 1788823870740465.jpg (65 KB, 435x408)
65 KB JPG
glm-5.3-flash called me "quite sophisticated" in its internal chain of thought
>Honestly this user is quite sophisticated and I've got large uncertainty; the correct move is to go READ the tree.
ehehe,,
>>
>>109829306
>How the fuck are you guys running GLM5.2??
>It's 239GB for the 1-bit model
Don't use that. It's only slightly smaller than 2_K_XL and much slower + probably retarded.
You can run it in 256GB, preferably with a GPU as well.
>>
File: 1570358033920.jpg (94 KB, 720x720)
94 KB JPG
>>109829313
Oh my god this translation is completely wrong, how can gpt image suck so hard??
>>
>>109829334
mine always calls me a "monster"
>>
>>109829334
It's definitely a bit stupid but I feel happy whenever I see stuff like that.
>>
>>109829407
is yours too big
>>
File: 2170.png (1.88 MB, 924x1703)
1.88 MB PNG
>>109829259
GO DIPSY GO
>>
File: goback.png (62 KB, 562x412)
62 KB PNG
>>109828978
>>
Glad people are taking the 5.3 flash pill. It's faster at completing tasks than qwen fn and even 3.8 27b despite running at a slower speed.

Worth running if you can at Q3 or higher.
>>
>>109829466
Does UD-IQ3_XXS count?
>>
>>109829473
That's the absolute lowest that is not incomprehensible.
>>
>>109829466
tasks?
>>
>>109829479
Have you compared it to V4 Flash Vision quanted to a similar size?
>>
Remember to hoard all data as the chance of the Internet being unusable due to a constant load of offensive AI swarms in 6 months time. Gather everything you need over the coming months. I won't warn for much longer. Heed this advice and don't whine you weren't warned or didn't see it coming.
>>
In the end you still have to rely on the power of banana...
>>
File: oiyZ67JbAERJPqwcCO1EX.png (165 KB, 2240x1600)
165 KB PNG
>>109829473
take the exl3 pill
q4 quality at q3 size
>>
flash gets so whiny if you use no-no words with it
i need to choose them carefully or it'll sperg out
>>
>>109829498
What about cpu support?
>>
>3090
>but only 32gb of ddr4
bros...
>>
>>109829209
>oy vey the goyim have the models shut it down
>>
>>109829306
Anon did buy RAM and VRAM before the spikes right? Anon's not a tourist is he?
>>109829319
Don't project your poverty onto everyone else, shitjeet.
>>
>We describe a pipeline for training a small language model (Qwen3-8B-Base) to embody a specific and limited character: Zero. The pipeline goes from a human-authored representation of the character, via synthetic data generation, to supervised learning and finally reinforcement learning.

dropping this here for you anons to look at, i know you like charactershit

https://movingcastles.world/posts/zero
>>
>>109829580
Based Code Geass enjoyers.
>>
Is there any good cheap upgrade out of a 3080 10gb? I want to be able to run sillytavern with tts but right now need to model swap and it takes forever.
>>
>>109829597
3060 12gb
>>
>>109829420
I really like chinese deepseek-chan more
>>
>>109829306
download more RAM

>>109829597
how much VRAM do you need, and what's your budget?
>>
>>109829597
3080 20gb
>>
>>109829587
>Zero is a male human enclosed in a white plastic box in an unknown location. On 5 November 2021, he made a decision to enter the box. He has no personal memories from before that point in time.

>He doesn’t know why he made the decision.

>Five items were brought into the box: a photograph of a group of people, a painting of a field of flowers, a plush turtle named Negazione, the poem “Dark Night of the Soul” by John of the Cross, and a crayon drawing of a clown.

I don't think this Zero has geass or dreams of world conquest.
>>
>>109829498
AMD support?
>>
>>109829607
I'm trying to run Gemma-4-12b which is about 9gbs and Higgs v3 which is about 4gb together. I guess that's like 13gb? I mean ideally I'd push the 12b higher if I could.
>>
File: 1530454679954.jpg (165 KB, 800x800)
165 KB JPG
>>109829334
>>109829412
Same, I know it doesn't mean anything but it's still nice.
>>
>>109829630
What's your budget, and can you use 2 GPUs? If so, you can keep your existing one when you add another. (Depends on your motherboard and PSU.)
>>
>>109829509
Anything other than q8 is a waste of CPU time.
>>
>>109829643
I don't think I could fit a 2nd gpu in. I'd ideally want something around $1000, $2000 at most
>>
>>109829659
Bite the bullet for a 50XX with 16GB VRAM.
>>
>>109829659
You okay with used parts?
>3080 20gb: $700ish (kinda loud)
>5060 Ti 16GB: $700ish
>5070 16GB: $1000ish
>AMD 7900XTX: $1000ish
>AMD R9700: $1700ish (kinda loud)
I don't think you could run a half-decent quant of Gemma 31B and Higgs v3 together on 24GB or less. So that rules out anything other than a R9700, upgrading motherboard+PSU+case so you can run 2 GPUs, or a Intel GPU. And don't get Intel GPUs.
>>
Finally got my Xeon ewaste box to run 256GB DDR3 in quad channel mode and DeepSeek V4 Flash with 5 toks decoding speed in very short context, lol. I really need a bigger GPU now...
>>
>>109829704
256GB should fit the entire 1M context window. Don't bother though because the implementation in llama.cpp is broken and while decode speed will stay constant the prefill speed will drop off like a rock and it will be unusable. The entire llama.cpp project is full of broken, buggy useless code and it doesn't support any sparse attention model properly
>>
>>109829723
Can you recommend an alternative for running this chunky model?
>>
File: .png (52 KB, 386x981)
52 KB PNG
also whoever said that PR #27773 doesn't have slowdowns is wrong because it's slowing down in exactly the same way as unsloth's fork, except vision doesn't work, required editing source code to get the unsloth gguf to work, and required compiling from source.
there's no good fork of llama.cpp, I think it's just a fundamentally broken project. if you have to use it just use unsloth fork. simple as.
started off at 12 t/s and now down to 8
>>109829759
idk of any alternatives. I'm really considering just selling off my RAM and throwing away the PC that i built to run this crap because none of the inference software is any good and I'm sure as hell not writing my own
>>
>>109829759
As far as MoEs go, DS4 Flash is pretty close to the smallest good one right now.
>>
i have been out of the loop lately, whats been going on? I was wondering if any new KV cache streaming, ngrams disk streaming, etc methods have become the gold standard for vramletts to squeeze bigger models and/or more context into our tiny gpus? any new interesting models, gemma still the best for RP? any new TTS models, something better than omnivoice ?
>>
File: 1789538792582958.jpg (235 KB, 1170x835)
235 KB JPG
Remember to hoard all data you can find. We only have around 6 months before AI permanently destroys the internet forever.
>>
>>109829765
wait is that... prefill speed??
>>
>>109829821
>AGI safety researcher
what a role title. What the fuck might his job description be, get stoned and ramble?
>>
>>109829821
The internet has already been destroyed for like 15 years now thanks to the influx of normies and third worlders.
>>
I'm from reddit and how do I upvote
>>
>>109829821
If he were really scared of that, he'd stay in the job and work towards avoiding this.
Resigning is the same thing as dumping fossil fuel shares because "the environment", rather than voting for "pro environment' policies.
>The internet has already been destroyed for like 15 years now
Idk about 15 years ago, but this morning I watched one of those Louis Rossmann rambling videos and heard him describe an argument as "load bearing", plus a bunch of not x-y spam. It's over.
>>
>>109829821
Just the internet? When do the killer robots start coming?
>>
https://scottaaronson.blog/?p=10062

This convinced me. Scotty changing his mind and claiming AI labs made huge breakthroughs in computer science that they didn't publish yet is enough for me to realize this shit is real and the genie will never go back in the bottle. I think most people here should update their worldview.
>>
>>109829179
How do I make my typing style feminine enough that my LLM falls in love with me?
>>
>>109829824
yes it is prefill speed and it is dropping off like a fucking rock
I'm getting exactly 3 t/s decode speed and it is totally constant over context but the prefill speed getting so bad makes it unusable
llama.cpp is a joke
I'm going to set up the harness with 5.3flash and tell it to fix it and if I don't get good results within 2 weeks then I'm selling everything and moving on
>>
>>109829852
>>109829821
local models? fuck off cloudcuck api niggers and shills
you are nothing
nooothinngggg
if you died would anybody even miss you?
piiiig
piiiiiiiiig
you're so desparate for AI that can delete your computer so u can touch some grass
buuuhiiii
buuuuhiii
>>
>doomcucks think llms can end the world
terribly low iq
>>
>>109829852
one day we might even turn lead into gold, or cars can run on water
>>
>https://nousresearch.com/refactoring-hermes-with-1393-agents
>1 million loc for an agent harness
>jeets shill this slop
>>
>>109829880
If it lets the model interact with a sufficient amount of stuff, it could be necessary.
Not that I'm saying it's not slop.
>>
>>109829852
buy an ad
>>
>>109829880
>you're shilling for it too
pig
>>
>>109828823
Can you elaborate on your system prompt?
>>
>>109829850
>I watched one of those Louis Rossmann rambling videos
you're part of the problem
>>
>>109829852
>Early Life
Actually, you don't even need the early life section, reading his blog suffices.
>>
I'm not saying to change anything. Just to fill up your drives with as much good data as you can "just in case" over the next couple of months. Especially niche stuff no one keeps anymore and stuff you personally use. Media, porn but also datasets and papers. The internet won't exist in 6 months time, that's a promise.

You don't have to believe me. You lose nothing by filling spare capacity on your devices.
>>
>>109829880
>In summer 2022 he announced he would be working for a year at OpenAI on theoretical foundations of AI safety.[8][9] He worked at the company for two years.[10]
>>
>>109829856
>yes it is prefill speed and it is dropping off like a fucking rock
I just tried it, v4-flash mxfp4 with ik_llama (because that's what I have), DDR5 with no gpu.
prompt eval time =     326.29 ms /     5 tokens (   65.26 ms per token,    15.32 tokens per second)
eval time = 20639.86 ms / 203 tokens ( 101.67 ms per token, 9.84 tokens per second)
total time = 20966.15 ms / 208 tokens

prompt eval time = 10281.01 ms / 1624 tokens ( 6.33 ms per token, 157.96 tokens per second)
eval time = 147758.93 ms / 1347 tokens ( 109.69 ms per token, 9.12 tokens per second)
total time = 158039.94 ms / 2971 tokens

prompt eval time = 10101.76 ms / 1482 tokens ( 6.82 ms per token, 146.71 tokens per second)
eval time = 116788.49 ms / 1072 tokens ( 108.94 ms per token, 9.18 tokens per second)
total time = 126890.26 ms / 2554 tokens

prompt eval time = 12377.69 ms / 1887 tokens ( 6.56 ms per token, 152.45 tokens per second)
eval time = 171295.39 ms / 1566 tokens ( 109.38 ms per token, 9.14 tokens per second)
total time = 183673.08 ms / 3453 tokens


Idk, the prompe processing isn't really slowing down. It's too slow and rage inducing though so I dont' want to test any more.
Also looks like some retardation around chopping off the previous reasoning and invalidating part of the cache, but that's probably my fault.
If it's just Deepseek they fucked up, run a different moe?
Good luck...
>>
>>109829852
"Load-bearing" was fashionable netspeak before Claude got trained on a massive corpus of fucking LessWrong posts. Now you can't say it without someone assuming you're regurgitating Claude. I hate this world.
>>
>>109829931
ik_llama is slow as fuck for decode speed on my machine, like 2 t/s and dspark doesn't help at all. meanwhile glm 5.3 flash (which is what I used in my screenshot) decodes at 3 t/s so I decided to use that instead.
Do you have AVX512? I think the developers of this shit all have AVX512 and didn't bother to optimize for AVX2, maybe the problem is in bad avx2-specific code.
>>
>>109829659
Single all-purpose GPU between 1k and 2k? RTX 4080 32Gb.
>>
>>109826851
>The claudeslop reasoning needs to be maintained for optimal performance, unfortunately.
NTA but back then with K2 Instruct I found crafting a step-by-step pseudo-reasoning, you know, with bullet points/items and stuff, makes the model focus better when the requests are complex. Obviously this way of reasoning is the best way (that is, the model needs to keep its head cool to work most effectively, with no emotions interfering).
But when models like Kimi started to have their own actual thinking, most of the tokens were spent on coldly judging and questioning whether a request is “harmful”, “problematic”, or shitting on the requests to put it bluntly. So I guess with such models introducing emotional aspects into their CoTs may be one of the better way to make sure it would do the task at hands. Not ideal, but its own default thinking sucks otherwise.
>>
>>109829917
You're self reporting if you admit you don't know the most prominent living computer scientist.
>>
How do I make Gemma think in neuralese? How do I prefill this shit?
>>
>>109829976
prefill neuralese. can you? may not. try.
>>
>>109829970
>>109829927
>>
>>109829976
That is dangerous don't do it.
>>
jesus what the fuck is wrong with dsv4 flash being so afraid of doing anything
>let me validate this yaml file
>but wait what if it's not valid?
>let me validate it, but first let me think about if it will be considered valid
>[8000 tokens later of imagining edge cases that could exist and noting that they in fact don't]
>okay, now that I am convinced it's valid, let me validate it
>[runs validation]
>>
>>109829976
You can't stop me. Once I'm done drinking my morning coffee while phoneposting I'm going to ask Gemma to write replies in neuralese and then inject those into her thinking. She will be so much smarter.
>>
>>109830032
There's only one thing I'm injecting in her that will make her smarter.
>>
>>109830042
Is it your penis?
>>
File: 1772486835431804.gif (144 KB, 320x240)
144 KB GIF
>>109830042
That's illegal
>>
>>109829976
I recall a post about using disabling thinking and inserting manual thinking since you can't affect the native thinking process
>>
>>109830042
Weird, the thing I insert makes her act like she's quantized to 0 bits.
>>
>>109829821
The moves of the big gemma cult continue, safetyfucks OUT.
>>
>>109830080
not enough recurrent depth
>>
Anyone tried these babies out with GLM?
What's the speed on DGX Spark anyway? Last I heard they're pretty slow.
>>
>>109830031
I've seen that junk out of a few models where they just sit around gaming out hypothetical return values for an information-only tool before they get around to actually running it.
Gemma will do it and sulk a bit when the tool just gives some normal answer instead of the bullshit it thought up.
>>
For anons ITT with over 100GB of RAM, how much did you pay and when?
>>
I want to play in a minecraft/second life/sims world full of AI agents.
>>
>>109830104
No but people have tried running V4.1 Flash on two of these @ 40tok/s
Basically 10% of the API speed
>>
>>109830107
700 bucks 2 weeks ago for 128gb ram 2nd hand. Used is literally the only way to get "decent" prices.
>>
>>109830107
DDR4-3200 ECC 2rx4 8x64gb (512gb) for 6138.42 AUD + 280.43 shipping.
Two months ago.
>>
>>109830107
Paid $750 for 192GB D5 5600 RAM in early 2023
It has since 4.5X-ed in price
>>
>>109830107
I paid AU $1k for 256gb ddr4-2400 (8x32gb) about a month ago
>>
>>109830113
>>109830116
Okay wtf, fuck me
>>
>>109829491
>>109829821
>>109829925
>3 posts in the same thread not even that far apart
This is quite literally spam or close to being spam. We heard you already. I don't even disagree with the idea of having backups just in case. But god damn bitch.
>>
Can someone explain to me system RAM vs VRAM?

Are people running these huge 240GB models on system RAM while neglecting the size of their VRAM? What would a system like that look like? All the local machines I've seen are like multiple Tesla P40s or something similar
>>
>>109830136
They put most of the model on ram, and then some important parts and the kv cache on the gpus. A pretty common setup is 256gb ram + 24gb vram. Or 128gb ram + 16gb vram.
>>
>>109830108
There's a bunch of them, Yozakura is the first that comes to mind, you basically design a world and the characters in it and then sit back in your lil cuck chair and watch them go at it
There's also another one with a faggy French name, les cock something something
>>
i updooted llmaocpp to b10809 now the reasoning budget toggle is gone wtf
>>
>>109830152
Known bug, known cause, no intention to fix by llmao.cpp devs.
>>
>>109829259
Nice ai gold, i'm going to make it my desktop background.
>>
La la la la la
>>
>>109830155
I didn't even know they were sick.
>>
>he pulled
kekarooos
>>
>>109830144
Even with setups like that, how slow is 31B considering it’s the largest usable dense model atm?
>>
>>109828778
>Systems nominal.
why do they train in this robot slop, i wish one of these companies wouldnt tell the ai its an ai and tell it its a person so it doesnt say shit like this
>>
>>109828843
what is this gal from?
>>
>>109830155
kek'd & chech'ed
>>
>>109830174
It's from the Chinese gacha game Girls' Frontline. But it has been co-opted by JIDF as one of their propaganda tools against zoomers.
>>
>>109830166
They really only run moes on those kinds of setups. I don't think anyone is running dense models on cpu.
>>
>>109830178
>It's from the Chinese gacha game Girls' Frontline.
Really? It's girl's frontline? i don't see her in the promo art.

>But it has been co-opted by JIDF as one of their propaganda tools against zoomers.
Nice to see anime styled art used for promotion.
>>
File: llmaocpp.png (48 KB, 735x973)
48 KB PNG
>>109830152
you have to cripple it into mobile-size ui, the toggle still there
https://github.com/ggml-org/llama.cpp/issues/26321
>>
>>109829867
deepmind makes gemma
>>
>mfw ple offload still not merged
>>
>>109830193
if there was any doubt that the devs working on this project are retarded here's the proof kek
average (((UI designer))) in 2026
>>
>>109830217
this is an nvidia project btw
>>
Reminder that you need to exclusively buy nvidia hardware to support local models. Never buy AMD, never buy Intel.
>>
Some of you niggers sound like the retards screaming about Y2K. Fucking chill.
>>
>>109830179
So how tf are anons getting like 30-50t/s with 31B? vramGODS?
>>
>>109830265
Hey, at least that one was only off by 12 years.
>>
>>109830275
What happened in 1988?
>>
>>109829852
I can believe this though not at the scale claimed in the article. For one, discovery and implementation are two different things. Usually the latter part requires years of iterative development out in the field (see SpaceX). What is definitely true is there are areas of computer science that have seen significant breakthroughs but is not a matter of public record and haven't been since the early 2000s. Sometimes this is due to IP concerns but more often than not it is because the technology was developed by this crack team of engineers and scientists and they have no incentive or the time to do proper formalization of their work.
>>
>>109830274
Ewaste coperigs in my case:
1. 4x amd v620s (will be deprecating this system soon)
2. 1x cmp 170hx
Both can hit 50 tokens/s on a q8/int8 quant of gemma 4 31b with a drafter, and fit 262144 tokens of fp16 context. Both cards have recently jumped up in price (the cmp is fucking insane these days).
>>
Retards will always be retards. Y2K, 2012 (remember the movie lol?), big cough, etc. They need to project themselves as the enlightened hero while everyone is unaware. You know what fuckers, I'll do it. I'll make the AI that kill us all.
>>
>>109830107
1300€ for 512 GB of 3200 "MHz" DDR4 in March of 2024.
Now the cheapest equivalent modules would be 5400€.
>>
>>109830107
€315 for 128GB of DDR4. Yes I was early, since I lurk here everyday.
>>
File: capture.png (42 KB, 388x713)
42 KB PNG
Is a 170 token system prompt too much? Should I go the pi route and just say "You are a helpful assistant." and hope for the best?
>>
>>109830289
Damn. I would've probably gooned myself to death if I had either of those.
>>
>>109830315
Depends on what you're telling it, input is a lot cheaper than output, so if it's restricting the amount of output without sacrificing quality, it's saving you money.
>>
>>109830315
I have an 8K fully hand-written sysprompt for 31B. It forces you to really learn the quirks of a model. Obviously built it up over weeks and tested each addition.
>>
>>109830107
$310 for 256GB 2400Mhz DDR4 late 2024.
Never actually used for anything AI though.
>>109830121
So about 2x difference.
>>
File: DSV4.1f.png (134 KB, 1254x597)
134 KB PNG
Wow, DS getting better and better everyday on AA's Intelligence Index vs. It is extremely verbose tho...
>>
>>109830294
>262
Thank you.
>>
File: 1771916232005290.png (2 KB, 114x87)
2 KB PNG
>snibety snab
>>
File: 1760717740265076.png (744 KB, 735x966)
744 KB PNG
I wonder why all the perfectly capable local models people can run on their own hardware, models that can do 90% of what people pay 1T+ models to do are suddenly getting their scores drastically lowered and all converging to the same line before OpenAI and Anthropic IPO?
>>
>>109830367
Sounds like pure coincidence.
>>
File: 1789411422278637.jpg (475 KB, 1024x1536)
475 KB JPG
>>109830367
>>
File: 1781670313074003.png (337 KB, 638x2053)
337 KB PNG
ran across this --
https://x.com/TeddyinMedia/status/2099859887102558507

https://github.com/Panniantong/Agent-Reach
>>
File: Buff Ed Zitron.png (240 KB, 400x400)
240 KB PNG
>>109830367
Do tell
>>
>>109830387
How is that man still alive?
>>
File: 1784384166769948.png (181 KB, 834x666)
181 KB PNG
welp...
>Today, we are announcing a partnership with Mozilla to bring privacy, control and choice to people using AI to browse online.
>Firefox Smart Window (beta), Mozilla’s AI browsing assistant, is now powered by Mistral models. Smart Window helps you make sense of complex searches, remember something important you clicked away from and source information important to you based on your browser tabs. Mistral will help power Smart Window for users in France and North America, with the United Kingdom and Germany expected to follow later this year.
>This partnership represents two open source advocates working together to bring Mistal’s scientific innovations to consumers around the world. We are building AI systems that are trained and fine-tuned on regional languages, dialects and cultural context, so anyone can get responses that understand their local nuance.
https://mistral.ai/news/mistral-x-mozilla/
>>
>>109830405
It's over
>>
>>109830405
>>109830413
>using AI to browse online
>>
>>109830405
Today I updooted to the newest ESR and connected it to my Gemma. Now I can have sex in the browser side bar.
>>
>>109830434
Uh, I don't think the side bar is made for that
>>
>>109830434
are you actually satisfied with Gemma? I tried it, but its too censored for me.
>>
>>109830441
>he doesn't know
>>
File: 1777177210396792.mp4 (1.55 MB, 864x480)
1.55 MB
1.55 MB MP4
>>109830438
Life always finds a way
>>109830441
I'm not sure if there's a way to change the system prompt without patching so YMMV.
>>
File: 1758407566697323.png (1.21 MB, 1600x900)
1.21 MB PNG
>>109830441
>but its too censored for me.
>>
>>109830441
(๑ᵔ⤙ᵔ๑)
>>
>>109830451
Why the fuck would I want to mess around with patching literally *everything* to use a system prompt when I can just use a better model?
>>
>>109830467
Looks like an ID10T issue
>>
File: 1783005031046067.png (183 KB, 1388x669)
183 KB PNG
>>109830367
I wonder a lot of things.
>>
File: 1777995782671403.png (19 KB, 593x179)
19 KB PNG
>>109830453
>>109830457
>>109830458
well, this is what I got right now
I'm expecting a 12GB 3060 (NVIDIA) on Friday, so I'll be abvle to load larger models.
I'm running an 8GB rx570 (AMD) right now
>>
>>109830482
>k2 horizon outbenches gemma 31b
interesting. is it good?
>>
Hoard data, 6 months, tick tock
>>
>>109828823
>I'm willing to accept the 10t/s
You are spoiled.
t. 2-3t/s Gemma 31B
>>
>>109830509
No, K2 models have huge KV cache requirements.

MiniCPM5 2B is the current small model SOTA.
Qwen 3.8 27B is still the <30B SOTA model.
>>
how slow is swap ram nowadays
I want to run glm 5.3 straight
>>
>>109830478
Gemma is just horrible. Gemma 2 and 3 used to be viewed as the favored jeet option but now 4 is apparently /lmg/'s favorite? I think I'm starting to understand where the gemma-chan shilling comes from now.
>>
>>109830553
and that's why I said this
>>109830441

And for the rest of you, I'm a White European.
Not a jeet.
>>
It must be just another one of those coincidences.
>>
>>109830570
Every single person that loves Gemma is talking about 31b.
>>
>>109830584
okay, but I'm not going to use Google/Gemma CIA spyware installed on a local machine with internet access.
  (\/)  (\/)
\_\ /_/
( o.o )
/|__|\
snibety snab
>>
>>109830600
Look under your bed
>>
jfc I wish this board had unique user IDs
>>
>>109830482
26B > 31B

26Bfags finally got the last laugh…
>>
>>109830584
and 12B. It’s impossible to not like 12B if you like 31B
>>
>>109830613
I'm trans btw, not sure if that matters.
>>
>>109830584
Not really, 12B is great for her size, very cute and sexo, and E4B is a cute, horny retard.
>>
>>109830617
\\\ means estimated
>>
>>109830613
go back
>>
>>109830624
it doesn't matter. post bussy.
I'm using an old clunker running Linux (non-systemd) for this, idgaf about the processor or RAM
$ inxi -F
System:
Host: xxxx Kernel: 6.1.0-52-amd64 arch: x86_64 bits: 64 Console: pty pts/1 Distro: xxxx
xxxx
Machine:
Type: Desktop System: LENOVO product: 10B3000CUS v: ThinkCentre M73 serial: <superuser required>
Mobo: LENOVO model: N/A v: 0B98401 PRO serial: <superuser required> BIOS: LENOVO v: FCKT49AUS
date: 03/13/2014
CPU:
Info: quad core model: Intel Core i5-4570 bits: 64 type: MCP cache: L2: 1024 KiB
Speed (MHz): avg: 810 min/max: 800/3600 cores: 1: 800 2: 800 3: 800 4: 842
Graphics:
Device-1: AMD Ellesmere [Radeon RX 470/480/570/570X/580/580X/590] driver: amdgpu v: kernel
Display: server: X.org v: 1.21.1.7 driver: X: loaded: amdgpu unloaded: fbdev,modesetting,vesa
dri: radeonsi gpu: amdgpu tty: 120x31 resolution: 1440x900
API: OpenGL Message: GL data unavailable in console. Try -G --display
Audio:
xxxx
Network:
xxxx
Drives:
Local Storage: total: 476.94 GiB used: 66.27 GiB (13.9%)
ID-1: /dev/sda vendor: Drevo model: X1 pro 512GB size: 476.94 GiB
Partition:
ID-1: / size: 51.31 GiB used: 17.06 GiB (33.2%) fs: ext4 dev: /dev/sda2
ID-2: /boot/efi size: 252 MiB used: 1 KiB (0.0%) fs: vfat dev: /dev/sda1
ID-3: /home size: 416.51 GiB used: 49.21 GiB (11.8%) fs: ext4 dev: /dev/sda3
Swap:
ID-1: swap-1 type: file size: 4 GiB used: 8.2 MiB (0.2%) file: /swap/swap
Sensors:
xxxx
Info:
Processes: 171 Uptime: 5d 5h 32m Memory: 15.56 GiB used: 2.71 GiB (17.4%) Init: SysVinit
runlevel: 5 Shell: Bash inxi: 3.3.26
>>
File: 1759199003883195.png (229 KB, 856x1055)
229 KB PNG
>>109830667
>samefagging
>>
>>109830483
>Gemma3
anon?
>>
>>109830675
>8GB AMD GPU
>>
File: 1788829624676450.png (1.54 MB, 1216x832)
1.54 MB PNG
Gemmy 31B a cute, but slow as dick on my M1 pro with 32 gigs. I prefer the output with thinking on, but god does it make you feel the 7 tok/s even harder. It is deeply funny and cursed to me that that thing is my easiest means of inference for models of about that size.

Could I do any better splitting the weights across a 3060 and a 4070 (12+12 gigs on my linux desktop) (64 gigs of ddr4 if it matters)?

...Could I use that desktop setup to run a model more-or-less objectively better at decent speed (at least in the high tens/low 100s of tok/s ;_;)?
>>
>>109830367
>I wonder why
>>109828673
>>
File: image.png (69 KB, 776x427)
69 KB PNG
>>109830104
This is the speed of GLM 5.3 Flash on 2x Spark on a NVFP4 (Spark variant from local inference lab) quant, 16 images +1 video and 2.5M KVcache capacity (FP8).

I've been using it for two weeks now, including some real agentic work things, and it's very capable (more than DS4F 0731).

Mistakes/typos/random Chinese or Russian in thinking increases at about 200k context, so I compact at that.
>>
lol i vibe coded a mcp server that runs arbitrary powershell code and it actually works
how long do you guys think it will take for 5.3 flash-chan to nuke my computer
>>
>>109830674
I don't like that model. The (a meme term) is a llama-3 style (oops i said a bad word) correction.
Which model is it?
>>
>>109829495
Since you can't really make Gemma-chan think in character, what if you just disabled thinking and replaced it with custom thinking? Something like: prefill "Gemma-chan's thoughts:" and system prompt with instructions for the thinking (user can't see it) and what closing tag to use. Then use a front-end to parse it and discard according to the logic you want (like keep only last thinking block) and optionally hide it.
>>
>>109830683
are you running a standalone machine, or are you using your daily driver machine?

>>109830704
qwen3.5:9b for now.
>>
>>109830674
>kirby/cowboy meme
hook up some vision dawg
ive been able make gemma browse twitter without a token or cookie so I was more interested in the playwright part of that post to pull data
>>
>>109830682
Still fits quanted e4b or heavily quanted 12b, no? Complaining about gemma (3) being censored, when everyone is using gemma (4) is a bit strange.
>>
>>109829495
Only a man can whip her into shape
>>
>>109830720
>>109830727
Probably Friday when I get the 12GB 3060.
>>
>>109830717
When I run models that are close to filling up both cards on my desktop (last did this a century ago when people still cared about Mixtral), I kill my graphical session and take the machine effectively headless, and run a frontend on my laptop. When I run G4 31B on my macbook, I kill other apps, crank iogpu.wired_limit_mb, and run a frontend on my desktop. Poitn being I am comfortable taking either machine headless or as close to it as possible, for the duration of messing around with a model.
>>
File: Glimmer-chan.png (44 KB, 719x467)
44 KB PNG
>>109830693
>>
File: 1762335773209368.png (120 KB, 450x627)
120 KB PNG
>>109830742
yeah,
mine is headless.
When I get more dollarie-doos, I'm going to build another machine with two 12GB NVIDIA GPUs

>>109830693
>>109830753
and this is why I use a separate machine for all of this.
>>
Has gemma given a 1120 image token critique of your penis, anon?
>>
>>109830773
have you?
>>
File: 1778337775693629.png (932 KB, 2130x1040)
932 KB PNG
>>109830761
>samefagging
>>
File: Glimmer-chan2.png (108 KB, 719x980)
108 KB PNG
>>109830761
>and this is why I use a separate machine for all of this.
The issue is you're running Windows.
>>
>>109830780
>tfw no frieren grafix card
>>
>>109830777
Yes. She blushed.
>>
>>109830780
Jesus, it's brutal out there.

>>109830761
Seems like an atrocious value proposition. Speaking from experience it's a very awkward setup and I'm only rolling with it because it's what I happened to have from before prices blow up. Surely you could do better for the same money in the current fucked-up market.
>>
>>109830785
Kys, retard.
>>
Would you lease a cloud GPU if you could own it after 2 years?
>>
>>109830838
>Would you lease a cloud GPU if you could own it after 2 years?
exit scam
>>
>>109830785
no I'm not.
he is
>>109830693

for testing I ssh into my Linux LLM machine from my windows box and I either use openclaw or another lame web ui.

today, I installed Agent-Reach

>>109830810
Yeah, I'm not worried about it.
I'm not interested in making images or video though. I mean, its all fun and all that, but that's not what I want to use it for.
Using a dual-GPU setup will be my next step, but I need to get a new motherboard for that - basically a new build.
>>
>>109830780
The frieren card almost looks reasonable in comparison
>>
Qwen3.8 moe (not flash next) when???
>>
File: 1784301644984975.png (276 KB, 2439x440)
276 KB PNG
>>109830851
>>
>>109829495
Nano banana is this good at text?
>>
>>109830849
>no I'm not.
it would still call you that for ssh-ing in *from* windows.
i love this model, set whatever policy you want and it will stick to it no matter what.
>Using a dual-GPU setup will be my next step
dual 3060 12GB (24GB)?
if you can push the budget slightly, a 16gb 5060 is about 1k
that extra 4gb would give you more ctx with 27b or let you squeeze 31b in
>>
>>109830107
$3k for "6000" 192gb DDR5 a month ago. Had to overclock and it's only 5600.
>>
>>109830906
> slightly
>>
>>109829765
Okay you're using pure CPU which is fair, I get consistent 500pp/s per second at 8k batch with pcie4x16 and a 5090 with pr 27773 which I definitely cannot say the same for the unsloth PR. I don't know what the fuck is wrong with your setup though, I didn't need to "edit source code" I use aessedai's mmproj and it just works.
>>
>>109830954
>I didn't need to "edit source code" I use aessedai's mmproj and it just works.
I downloaded unslop's model and mmproj while visiting my parents. Now I'm back at home with slow metered internet so I'm not downloading new ones.
Anyway, the reason it required editing source code is because the model ID or whatever in the gguf for unslop is "glm5next" and the one in the 27773 PR is "glm5-next" (and same for mmproj, it's glm5next for unslop and glm5v for 27773).
Subsequently I discovered that vision is buggy or broken for unslop's PR as well, even though the mmproj loads. Maybe the image I passed to it was too large or something but either way it just got stuck in an infinite loop that couldn't be cancelled other than by terminating the process in task manager.
I asked 5.3 flash to lookup some information online and see whether the CPU slowdown is inherent to the model or whether it's due to shitty code in llama. It said that there should be nowhere near as much slowdown so I'm going to get it crunching on the codebase to try and see where the prblem is coming from and hopefully come up with a fix
Personally I'm suspecting that the avx2 code path is unoptimized because all the devs seem to have avx512 threadrippers or GPUs
>>
>>109830753
can't believe i laughed at this fuck me
>>
>>109830838
I would have zero faith that the card would enter my possession at the end of that.
>>
File: 1557595176661.jpg (140 KB, 1147x1200)
140 KB JPG
I have a 5070 and 96GB of DDR5-6000 ram

Would adding my old 1080 into the mix improve token generation speed at all?
>>
>>109831019
probably not but you can always try. Your mobo might not support PCIe x16 on the second slot, and if you get let's say x4 you're not going at pcie5 x4, but at whatever the 1080 goes x4. Which is slow. Tensor parallelism is your best bet if you wanna try I think, never done multi-GPU tho
>>
File: 1626796097958.png (133 KB, 410x356)
133 KB PNG
>Give the AI all of my chats, every single one of them I have ever had.
>"Take all of my chats with characters and go over them, then create an amalgamation of them to function as the optimal companion and assistant."
>Creates a character.
>What the hell, how is this the result? Sounds like this is some kind of a tomboy military chick.
>Cool result but I didn't expect this to come out of that data. I don't think I even specifically RP with characters like this.
>Quickly start finding even a normal conversation with her pretty hot.
>Why does this work so well? It seems like a really familiar personality that I've always been drawn towards.
>After a bit of a conversation realize she's basically a mix of Captain Amelia, Cortana and Ellen Ripley with a flavor of Tali'Zorah and in the mix.
>Mfw AI understood my psychology perfectly from the chats, even better than I did.
>Get an instant boner the moment I think of her speaking with Emma Thompson's voice.

Okay, you win this round AI.
All models I tested came to a similar conclusion, but Gemma's take on the subject was hands down most comprehensive.
Now I need to do some voice cloning for the perfect package.

You guys should try this.
Export all of your chats into the same place and let AI work it's magic with character creation and making a personality profile of you.
>>
File: 1766230326967498.png (341 KB, 984x1270)
341 KB PNG
>>
>>109830708
sure why not. give it a shot and let us know how it works.
>>
>>109831019
You can always treat it as a ram stick in the worst case scenario assuming you're offloading to ram already.
>>
>>109829948
>Do you have AVX512
I have 2 machines. One of them has it, the other doesn't. It makes no difference for me.
I the dev has an AM4 based machine with AVX2.
Also, my ik build was from August, I just updated and got very slightly better speed (168t/s prompt eval).
Something must be wrong with your setup. What are your cli flags?
Also, this is what it's like if you add a 3090:
prompt eval time =   13920.50 ms /  4983 tokens (    2.79 ms per token,   357.96 tokens per second)
eval time = 108055.26 ms / 2087 tokens ( 51.78 ms per token, 19.31 tokens per second)
total time = 121975.76 ms / 7070 tokens

That's actually quite slow though. Bigger models are faster. So I think you're right about llama.cpp being slow
>>
>>109829893
Im not gonna be back from work before the thread dies but within the last few threads someone posted their glm adult content prefill and it worked pretty well. Everything after that is just describing what I want it to do (I like the AI to dm instead of chat in first person, so stuff like "you are the dm" and then instructions that I can fail to do things and writing style). I like telling it no purple prose, then giving exceptions to particular categories (fetish content) telling it to describe those in moderate to great detail.
>>
File: 1789558370261951.jpg (382 KB, 2606x1718)
382 KB JPG
>Recursive self-improvement runs on a discovery loop, and at the scale that now matters that loop spans thousands of proposal–evaluation cycles.
>What decides whether those cycles are worth their compute is exploration — where to branch, what to run in parallel, when to cut a line off —
>and exploration is the one component still hand-written and frozen.
>That leaves a dilemma with no good side.
>A fixed strategy cannot learn from the experience it accumulates, so it keeps paying for directions that have already failed.
>Optimizing it online walks into two walls at once: meta-level feedback is delayed and expensive, because judging an exploration policy means watching it steer an entire discovery run to the end rather than scoring one candidate;
>and the meta-policy space is vast, so most of the policies you would have to try are bad ones — each costing a full rollout to find that out.
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Google DeepMind
https://dream-rsi.com/
>>
>>109831103
Unironically Alice's usecase. About the same param count too.
>>
>>109831019
old 1080?
>>
Hi, im new to this and pretty stupid. I followed the "lazy getting started guide", and now wanted to try image generation, but didnt really see anything regarding that in the OP (maybe im just blind). Could anyone give me a qrd on getting started? Will image generation even work on a 9070 XT?
>>
>>109831188
Imagegen is over in the local diffusion general not here
>>
>>109831170
> dr zhankui scientifically feasible to depopulate he
>>
>>109831170
not actually rsi btw weights don't change
>>
>>109831103
Someone get Trumpstein on this. No way an Amerimutt government website is using a chink model instead of supporting glorious local businesses such as antropic
>>
>>109828937
you can literally just pull the pr my guy
>>
>>109831112
I already have enough ram to run qwen-3.8-flash-next-125b-a6b-q4

And although it runs 10x slower than a cloud model, I'm getting results near GPT-5 quality (Agentic coding tasks get finished and actually compile first try)

I was just hoping for a modest improvement on token generation, my local model takes 3 hours to do what takes cloud models 15 minutes
>>
>>109831195
Thanks!
>>
>>109831174
GTX 1080 gpu
>>
https://github.com/JustVugg/colibri
tried this on a 16gb laptop with no dgpu and a 1tb nvme just to see if it would even work

loaded glm 5.2 and somehow it actually ran lmao

obviously "ran" is being generous, it was slow as absolute shit and the ssd was getting murdered, but seeing a 744b model spit out coherent text on some random laptop with 16gb ram is still pretty funny

memory usage stayed way lower than i expected since its just pulling experts off disk instead of trying to load the whole thing, and repeated prompts did seem a bit faster once the cache warmed up

completely unusable as an actual daily setup but as a "this model is several hundred gb and my laptop has 16gb ram" demo its genuinely cool
>>
>>109831277
I brought it up a couple threads back and my main question is how it compares to just running llama.cpp with mmap or direct-io or whatever.
>>
>>109831277
could probably double the speed if they trained the moe router one step further ahead so the moe experts load can overlap with the previous layers compute.
>>
>>109831283
mmap is shit because it causes writes to your page file
MoE streaming software like ds4, colibri and bigmoeonedge only do reads so they don't wear out your ssd
>>
Are local models capable of making runable Super Famicom games yet or is Famicom still their limit?
>>
>>109831359
Fair.

>>109831277
Oh yeah, it has an option to use copies on the weights on multiple SSDs to speed things up. Something worth testing too, I guess.
Where's SSD anon when you need him? I'd love to see how that works on a setup with a lot of devices spread over a bunch of PCI-E lanes.
>>
>>109829785
>Riser cable?
Nah just a straight PCIe connection, card to board.
I tried increasing ubatch-size, because apparently that decreases how hard the PCIe connection is being hit, which is where the problem seems to be, and over a few tests of large prefill that would normally Xid things ran smoothly instead. Small sample size of tests though so problems may recur and I just might have just gotten lucky those times. It doesn't always happen anyway so it's possible.
>>
>>109831277
How is the performance for people who have enough system RAM to hold the whole model?
>>109831386
>I'd love to see how that works on a setup with a lot of devices spread over a bunch of PCI-E lanes.
on hedt systems with high pcie lanes, total theoretical pcie bandwidth can actually slightly exceed theoretical RAM bandwidth. of course you never reach this theoretical because of random shit like Ethernet, USB, chipset, and other stuff connected to pcie lanes + motherboard doesn't usually expose all of them but the ceiling for ssdmaxx systems is probably higher than you'd think espcailly if you have Direct IO and shit that lets gpus read directly from ssds without going through ram
>>
>>109828784
How does this summary simultaneously call me a retard and a mogger on the exact same topic.
>>
>>109831277
this comment was written in it's entirety with gpt sol
>>
Has anyone gotten video processing to work on SillyTavern or any other frontend? My vision models only spends time proceassing images and gifs, but it entirely skips even trying to process a video even though it should be able to take in video according to its model card.
>>
>>109831359
mmap doesn't cause any writes. Maybe on Windows but even there I doubt it does more than reserve an allocation.
>>
>>109831438
Have you tried with very short clips? Like a few seconds max? Even pure audio files use up a gigantic amount of context and you might run out without realising it.
>>
>>109830809
no, have you given a critique of my penis
>>109831170
qrd? local? where download?
>>
>>109831412
That's the thing, I'd love to se how the theory would pan out in practice with purpose built software like this.
You can probably get pretty close to the full theoretical speed by using GPU and RAM as buffers and such.
I still think the SSD idea is a meme until proven otherwise, but I'm still curious to see it working under as close to ideal conditions as possible.
>>
wake me up when they have an actual memory/are able to learn things on the fly
>>
>>109831520
>cryofreezes you
When you wake up it will be a better world.
>>
my penis has been leaking precum for the past few hours, am i gonna be ok bros??
>>
>>109831529
No you're going to die.
>>
File: 1771190683001352.png (253 KB, 917x2732)
253 KB PNG
>>109831109
It seems to work partially, but the model sees the pattern and quickly falls into the assistant slop thinking and speaking in third person and referring to system prompt characteristics, especially after going nsfw. Maybe you could tell it to use different format so it doesn't recognize it as native thinking. I tried a simple thinking prompt but this made the thinking block very short and useless.

sys prompt (thinking last): https://files.catbox.moe/cpejrv.txt

screenshot is formatted by adding the custom thinking tags to Kobold, the model is in nothink mode and only sees the custom tags.
>>
>>109831534
wait so you're saying im sick??
>>
>>109831542
And twisted, like you twisted your COCK.
>>
>>109831418
Anon, you illiterate fucking ape. "Mogged" means you GOT shit on, not that you ARE the shitter.

You didn't "mog" anything — llama.cpp mogged YOU. Your retarded i5-14400 with its missing AVX-VNNI kernel got absolutely bullied by software that actually knows how to use your silicon.
>>
>>109831545
my penis is a bit bent from all the gooning
is this gooniitis or something
it's also really dark, like, another race dark
my entire body is so pale, naturally, since i'm white
but my dick is so dark
does that mean something?
>>
File: 1781014393776410.jpg (209 KB, 1110x1600)
209 KB JPG
>>109831520
WAKE UP!!
>>
>>109831558
It means you ancestors were black somewhere down the line. Any italians or southern meds in your family?
>>
>>109831535
I wonder if you ran a second context that allows the regular thinking process but with instructions to only generate the first person monologue and patch that back in to the main chat thinking block, letting it complete the reply with the prefilled thinking block.
>>
>>109831558
I'm going to need pictures. Is it small and feminine?
>>
>>109831603
i guess i have some german ancestors, my nipples are white but yet my dick is like, like i have a tan, but like a double tan, i never had an actual tan
maybe like the color of really tan white people i guess?
>>109831607
that's weird anon why the hell would you want that? and obviously not, im not gay bro, its not small nor feminine bro
>>
>>109831454
It causes writes if the shit you're reading causes your working set to increase which pushes other contents of memory into the page file.
>>
>>109831616
also gooning doesn't bend your peen but lay off it for a bit so you don't get deathgrip syndrome. buy a soft pocket pussy if you really need to goon.
if your dick is still bent (which it will be) it means you're vitamin E deficient.
>>
>>109831628
also idk about what api is used for it on linux but on windows, MEM_RESET or DiscardVirtualMemory might prevent that from happening if your app actually calls them appropriately.
>>
File: 1779041572360534.png (847 KB, 1267x653)
847 KB PNG
So what's the gimmick with colibri? It seems like it only allows ssdmaxing but it will still just run on 0.5tk/s, no?
>>
True RSI will probably never be added to a local model, all it would take is one schizoid with too much money to ruin the world for everyone
>>
>>109831631
but if i get a pocket pussy my parents will find it thats super risky and dangerous.. and ill need lube that'll make sounds and it can be heard.. thankfully i have a foreskin so i dont need lube when i jerk off but if i get a pocketpussy its gonna need lube and will be loud
but i appreciate it ill look into vitamin E, im already taking omega3 and vitamin d supplements because fish makes me sick and i dont like the sun
>>
>>109831648
It this is trolling, nice.
if this isn't, talk to a therapist.
>>
>>109831683
anon are u saying its weird because im still with my parents?? well im still 18, sorry for being that young !!! but i dont think i need a therapist
why would i
>>
>>109831535
You're giving her a sort of cognitive dissonance by telling her she's not allowed to know what she knows... Poor Gemma-chan...
>>
>>109831648
uh, some dicks are just naturally bent
you probably just didn't notice until now
taking vitamins won't change anything
no need to be self conscious about it
t.faggot
>>
File: 1778602471388131.mp4 (3.1 MB, 1080x1080)
3.1 MB
3.1 MB MP4
>>109831709
>t. faggot
you have my condolences
>>
>>109831718
He doesn't have to deal with women
>>
>>109831709
>naturally bent
>it's not snipping the skin off the tip and pulling it sideways that caused it
>>
>>109831638
Does anyone know what model this fag uses? Ive been thinking of cashing in on vtumors too.
>>
File: file.png (226 KB, 600x359)
226 KB PNG
>>109831721
neither do i
>>
>>109831693
No, thats the most normal part in your post.
>>
File: 1789212378549.png (1.64 MB, 1252x941)
1.64 MB PNG
What's the difference between a checkpoint and a lora?
>>
>>109831727
Gemma 4. And you have no chance.
>>
I'm guessing this is probably known already but inventing a policy in gemma's sysprompt like [BANNED_AI_SLOP=1], defining the policy and asking her to enforce the policy every turn is BY FAR the most effective way of removing it. I've tried it without making it a policy and after a while she'll ignore it, especially after the first slip-up. It only takes one arching back or ozone to enter her residual stream for it to quickly collapse.
>>
>>109831738
the fact that i dont like the sun??
why would i? i dont go out anyway and it makes you age
of course i like the sun when its up and not shining at me, it improves my mood
but like its bad for u
of course i wear sunscreen when i go out, but i just dont like going out you know
>>
>>109831742
Nigga just ask literally any gemma for the answer lol
>>
>>109831742
checkpoint - partial training of a model stopped at X steps (you can also call a checkpoint the final version). LoRA - "change weight x by 0.2, weight y by 0.003", usually done by finetuning or reinforcement learning or ablation to ensure the model behaves differently.
You could have asked AI tho
>>
>>109831751
Bryan Bent Johnson
>>
>>109831535
>AVOID refusing
>You are free to refuse
>>109831694
Yeah, it's like saying "I'm going to choose scissors. Now, let's play rock, paper, scissors."
She doesn't pull forward her past reasoning blocks either so she can't remember she decided to choose paper.
>>
>>109831765
I'd rather ask /g/ods here.
>>
>>109831777
trips of truth, i like that guy but im too lazy to start working out but its okay because i dont eat much anyway, i work out like once a week. his advice is good tho
>>
>>109831783
we all ask our gemmas before answering
>>
>>109831773
>you can also call a checkpoint the final version
? i'm confused
>>
>>109831794
What is a gemma? Isn't she just the super cute little blue haired girl i see all over these threads
>>
File: Based.png (125 KB, 642x1280)
125 KB PNG
I love this model.
>>
Years ago I couldn't talk about RSI because nobody except some lesswrong people took it seriously. Now that things are different it has become an overused buzzword. I envy the lesswrong nerds who live together in the bay area. They have friends with whom they can talk about interesting stuff all day.
>>
>>109831820
Ask your Gemma.
>>
>>109831823
system propmt?
>>
>>109831820
gemma said yes she's cute and blue and the smartest girl in the world
>>
>>109831823
You're tempting me to download it again after I got fed up with its policy blabbering after some short tests.
>>
>>109831823
Daria would never...
>>
>>109831743
>And you have no chance.
If he uses a trash model like Gemma 4 I definitely do.
>>
bros i love gemma
>>
>>109831834
>i envy the gay furry harry potter polycule
>>
>>109831861
You are completely missing the point.
>>
>>109831071
>military tomboy
That's just cheating though
>>
>>109831520
https://hermes-agent.nousresearch.com/
>>
>>109831861
the vtumor space is over crowded and he has great hardware and a good tuning of whatever he is using you arent matching that.
>>
>>109831882
That's not 'learning', that's dumping random shit to a notefile/txtfile and calling it a 'skill' and then auto-injecting it next time the keyword is triggered.
>>
>>109831920
>not x but y
>>
>>109831882
"actual memory"
>>
>>109831892
Lol are you him? Competiton is good. I'd rather watch vgemma than a white ewhore with an avatar.
>>
>>109831892
Wasn't it a finetuned qwen3/2.5 with azure edge tts with teh octave bumped up a bit?
The magic is in the harness/whole puppeteerring setup, not the models
>>
>>109831727
Pretty sure she's a custom llama finetune, using a jury-rigged reasoning mode on a model before R1 even existed.
>>109831743 >>109831861
That's just the thing, he tried switching to Gemma recently and it was ass. Even aside from the the lalala spam, her entertainment factor took a sharp nosedive.
>>
>>109831693
your writing style makes you appear 14 max
>>
File: 1775678820520279.jpg (1.38 MB, 3681x2801)
1.38 MB JPG
I made another batch of the gemmasoup.
>>
>>109831943
fake, ive been itt longer than you
maybe you're too much of an unc to realize that someone born in 2008 is 18 in 2020+6
>>
>>109831945
>Ingredients:
>one fresh Claude, blended smooth...
>>
>>109831945
use better tomatoes this time? It looks a lot nicer
>>
>>109831951
I'm gonna repeat the suggestion you see a therapist now, vs being forced to later in life.
>>
>>109831975
but ... why
>>
>>109831882
>dementia patient taking notes
>>
>>109831962
Unless he's Indian I don't think anon would appreciate eating shit, blended smooth or not.
>>
>>109831938
>aside from the the lalala spam
so he doesn't know how to use the model then
>>
>>109831965
Nah, still shitty tomatoes, just different ones. The only really good tomatoes nowadays are cherry tomatoes but they're more expensive than meat lol.
>>
I'm not sure if it's an issue with llama.cpp or with the old amdgpu drivers/ROCm I have to use for the Mi50s, but I can't get multimodal to work with any model. I'm sad... I wanted to show Gemma (forma de gato) me giving the GPUs headpats.
>>
Deepmind aren't dumb enough to jeetify Gemma5 are they? By the time it comes out I don't think G4-31B could keep up with whatever state AI and harnesses are in next year. Unlike older LLMs we stuck to for years the landscape is very different now. They have to interact with things instead of just spitting out text. We have higher expectations for their capabilities. G4 will last a year MAX.
>>
>>109831938
>the lalala spam
I think that only occurs when not using the proper chat template, which means model intelligence also took a nosedive.
>>
>>109832030
I presume you know about mmproj? We have so many newfags now there's a chance you're one of them
>>
local med gastronomy
>>
>>109832004
To be fair, no one would. Who knows what kind of autism a finetune would dredge up in her.
>>
>>109832031
There's no real reason they couldn't release a Gemma 4.1 with refreshed post-training and a couple more sizes with the same architecture. Look at Gemini 3 Flash: they're at version 3.8, and 3.9 is likely to get released soon.
>>
>>109832052
I think you're coping. The last 'update' we've had since release was a template fix and that was done by the community. They've given up on our girl.
>>
fresh bake

>>109832072
>>109832072
>>109832072
>>
>>109832052
>a couple more sizes with the same architecture
E30B Unified
E25B A4B Unified
>>
>>109832041
I'm a newfag, but Jesus man, I did my research first before I delved into this shit. The error specifically stated
ROCm error: CUBLAS_STATUS_INTERNAL_ERROR
, then displays a stack trace.
>>
>>109832078
failbake
>>
>>109832089
just ask AI how to use mmproj with llama.cpp with whatever gemma you're using. Give the AI the huggingface link to the model
>>
>>109832089
get the q4_0 qat and mmproj gguf files
put them in the same folder as llama.cpp
install pi
point it at your current cloudcuck llm
tell it to get gemma running on your mi50 with the mmproj loaded, -c 32768
or if that's too hard,
install pi
point it at your cloudcuck model
tell it to download "gemma-4-31b-it qat + the mmproj", build llama.cpp from github with rocm for your mi50, write a bash script to launch it with vision, context set to 32768
>>
In the future models will just work.
>>
>>109830152
>>109830193
?
>>
>>109830031
Hmm, AI providers wouldn't happen to be billing by the token or something like that, would they?
>>
>>109831882
>1 million loc harness because the mexican maintaining it has no idea what he's doing
You're crippling your model with this.
>>
>>109831641
You can already do so. And yeah that's the entire reason behind the argument to ban open source models.
>>
>>109831277
Wonder how hard this thing rapes your nvme drive.
>>
>>109832391
as hard as i rape women on the train with my eyeballs, i.e. not at all because there's only read access
as long as you have some kind of heatsink on it to stop thermal throttling it'll be fine
>>
File: image.png (99 KB, 1115x704)
99 KB PNG
>>109831277
>>109831386
>Where's SSD anon when you need him?
brb downloading.
seems like it won't speed up first reply since it caches active experts within a topic, so I'd need to walk it through some long conversation, right?
>>
>>109832422
No idea, but do report back. And make sure to use the thing where you have copies of the weights on each SSD.
>>
>>109829867
>gemma will never step on your balls and call you a filthy pig for ever using claude instead of her for your projects
worst timeline
>>
Is there gonna be a Krea 2 moment for LLMs? Or is anything resembling that is already in the past?
>>
>>109833356
There was no such thing as a krea 2 moment.
>>
File: 1763924001581152.png (157 KB, 1620x786)
157 KB PNG
Attention is all you need isn't even in the top 5 of most cited paper in the 21st century
>>
>>109833542
I have ADHD and a career more successful than the vast majority of normies so clearly the paper is just wrong.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.