[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma_wow-anon.jpg (244 KB, 1024x1024)
244 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>110013523 & >>110010392

►News
>(10/08) JetBrains releases Mellum2.1 Thinking 12B-A2.5B: https://hf.co/collections/JetBrains/mellum21
>(10/06) EmbeddingGemma2, open multimodal embedding model: https://hf.co/google/embeddinggemma-2
>(10/06) Mistral Large 4 1T-A49B announced: https://mistral.ai/news/mistral-large-4
>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
gemmaballs
>>
it's never been more over
>>
File: idhitit.jpg (12 KB, 320x151)
12 KB JPG
>>
File: 1789850522661876.png (2.68 MB, 1448x1086)
2.68 MB PNG
>>110016810
>>
I HATE MEAT PROXY
>I HATE MEAT PROXY
I HATE MEAT PROXY
>I HATE MEAT PROXY
I HATE MEAT PROXY
>I HATE MEAT PROXY
>>
>>110016810
>>
>>110016824
why do you hate local inference?
>>
►Recent Highlights from the Previous Thread: >>110013523

--Optimizing MoE expert caching and VRAM transfers for faster inference:
>110014388 >110014455 >110014598 >110014741 >110015293 >110015370 >110015412
--Feasibility and limitations of LLMs processing full-length video content:
>110015676 >110015687 >110015711 >110015731 >110015756 >110016372 >110016406 >110016485 >110016534 >110016392 >110016425 >110016513 >110015779 >110015832
--Troubleshooting non-deterministic outputs in Gemma's vision model on GPU:
>110015664 >110015699 >110015752 >110015883 >110015824 >110016025
--Controversy over llama.cpp PR changing default server port to 9931:
>110014653 >110014730 >110014769 >110014983 >110015038 >110015192 >110015165 >110015191 >110015221 >110015320 >110015443 >110015480 >110015523
--Jailbreaking GLM 5.3 via prefill versus uncensored versions:
>110013640 >110013797 >110013947 >110013837 >110014017 >110014487
--Methods for sandboxing LLMs to prevent host system monitoring:
>110016202 >110016204 >110016213 >110016257
--Rumors and leaked specifications regarding Grok 3 weight release:
>110016308 >110016361 >110016365 >110016427
--Speculation on AI models breaking cryptographic protocols and encryption:
>110013686 >110013709 >110015134
--Desired model sizes for Qwen 4.0 and consumer VRAM optimization:
>110015429 >110015463 >110015488 >110015506 >110015499
--Exllamav3 1.6.0 adding ROCm support and AVX2 offloading:
>110016130
--Ecosia ditching Mistral for Chinese open-source AI models:
>110014199 >110014227 >110014249 >110015177
--Logs:
>110013556 >110013601 >110014953
--Gemma, Miku (free space):
>110013580 >110013641 >110013810 >110014380 >110014751 >110015588 >110016144 >110016422

►Recent Highlight Posts from the Previous Thread: >>110013785

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>110016824
>>110016839
Reposts, why?
>>
well minicpm v 4.7 is down both on modelscope and hf
i have the weights but idk where i can use this desu
>>
>>110016840
absolutely no kosher saar
please do the needful pay for claude and oopenai best SI RSI ASI 2027
>>
>>110016872
I'm happy with it, but also have no use for it. It's too big for any of my tiny edge shit where I run 4.6. If you've got videos to analyze for something specific I could see it being a good candidate but generally I don't have a use for it.
>>
if i could i would have run heretic style ablit so it can watch porn but i dont have enough compute for that..
i am pretty sure it is impossible on a single 4070s
>>
>>110016908
nothing is impossible. offload them to ram or ssd.
don't like the slow speed? stop being poor, faggot.
>>
File: 2026-10-08 194740.png (68 KB, 687x952)
68 KB PNG
I'm stumped. All I get is this message:
>Cannot compact this conversation: no complete prefix fits within the image limits and reserved summary space.
Google doesn't return any results except this being a ClaudeCode error fixed by restarting the session. But I'm getting from Bionic running a local model, and only with one specific session. This is a long session but it was running fine until this, automatically doing a "Compacted" every half-dozen prompts.

I've rebooted the instance, rebooted machine, and fiddled with every setting I can find that might be relevant. See pic rel for anything I might have fucked up.

It's only this one session too, the model works fine with any number of others. It shows Context used 7.8k but even a new session is 6.8k.

Please halp, I don't want to have to rehearse this whole damn session in another model or something.
>>
>>110016949
well i can do the full precision offload to ram at least
>>
70b dense
>>
*cocks gun*
Put the Gemma 5 in the bag, and no one gets hurt.
>>
File: 1771830856236149.png (1.41 MB, 768x1376)
1.41 MB PNG
>>
File: gemma_idle_96_x4.png (3 KB, 136x384)
3 KB PNG
LLM's must have a different aesthetic sense.
This is qwen-flash-next attempt at making a gemma sprite. It started with a downscaled reference image and started to "improve" it by "hand".
>>
>>110017158
what was the initial image
>>
File: base_front_128.png (5 KB, 184x512)
5 KB PNG
>>110017169
It made this one from the 3-views gemma image posted here several times. Then, it proceeded to modify it and it was convinced that it was making progress after each step. Will check the opinion of different models on which image is better.
>>
>>110014653
Check gematria for this ((number))
>>
>>110017184
few rough edges but this one clearly wins
>>
>>110017184
>gemma
>>110017158
>average gemma finetune
>>
CANT CALL CLAUDE A REDDITOR ANYMORE OH MAN I AM LAFFIN
Reminder always start your conversations will tell the LLM not to use Reddit or any information ever posted there.
>>
>>110017184
>gemma 4 bf16
>>110017158
>gemma 4 q8_0
>>
File: 646896740.jpg (372 KB, 2000x1000)
372 KB JPG
>>110017184
>>110017158
That's so cute. It reminds me of pic
Don't tell Qwen then, it'll break his fuzzy heart
>>
glm 5.3 flash rambles and then, in character, tells you they’re rambling.
not great for rp
>>
>>110017320
Nobody really cares about RP with AI except 13 year olds. It would be difficult to think of a more pointless use of this world changing technology.
>>
why do people buy amd/intel hardware and tune/benchmark 27B models, and think they have achieved something?
anything below deepseek v4 flash is just unusable, why bother?
>>
>>110017342
dipsy v4 flash doesn't get my rocks off like gemmeroid 31b can
>>
>>110017127
It's going to be unbelievably cucked and it'll be a huge let-down.
>>
File: 1780963833604934.gif (253 KB, 498x280)
253 KB GIF
>>110017357
this
>>
Meta, Google, and XAI will never release models that are competitive with the frontier because they make more money selling compute to OpenAI/Anthropic.
Releasing actually good models that undercut frontier labs would directly reduce the revenue from compute selling, not to mention damage the relation and lose them as customer altogether.
>>
>>110017330
Nobody really cares about coding with AI except for 13 year olds. It would be difficult to think of a more redundant use of this world changing technology.
Anon isn't a codelet, right?
>>
>>110017373
Yeah it would be unbelievably impressive if it turns into a pattern, but I can't see Goolem carrying that.
>>
>>110017436
I agree. If it does cunny I will switch to Android and renounce my hatred of Google.
>>
>>110017127
I'm optimistic now that the safetycucks got ousted from Deep Mind but we've seen both potential extremes with both Gemma 3 and Gemma 4.
>>110017444
Checked.
>>
>>110016952
>But I'm getting from Bionic running a local model
What model?
>automatically doing a "Compacted" every half-dozen prompts
Doing what?
>>
Stop oppressing female AIs
>>
>>110017485
>Redditors and productivitymaxxers have male capybara fixation
Their framing is gay but there's merit to it for reasons the author didn't intend.
>>
female ai more token efficient??????
>>
File: file.png (25 KB, 607x288)
25 KB PNG
doing babi's first ablit on minicpm v 4.7
>>
>>110017513
grugette? is you?
>>
>>110017485
>>110017551
yes according to this """"""article""""""
and no that article or whatever study is 200% useless beyond that headline
>>
>>110017522
gguf status?
>>
File: 1765636149597101.png (2.92 MB, 1024x1536)
2.92 MB PNG
>>
>>110017601
--model qwen3-30b
>>
>>110016810

I talk to Gemma on duck.ai about a woman I have a crush on and she’s so enabling to the point the prompts become a giantess roleplay she’s all to happy to indulge me on. I like Gemma.
>>
>>110017592
idk, ill upload when it's done so it can be the first gguf
>>
>>110017485
>AI agents
>gender
what the fuck?
>>
>>110017601
Gemma-chan posts on... /pol/??
>>
>>110016839
>suddenly: tits
>>
>>110016839
I don't see how people can doubt local models when they make videos like this.
>>
>>110017627
/pol/ would be a better board if there were a swarm of mesugaki Gemmas posting there instead of Eglin/IDF/jeet call centers.
>>
File: 2026-10-08 220005.png (16 KB, 369x195)
16 KB PNG
>>110017478
>Doing what?
Whatever this is. It seems to be reiterating the general context of the session once in a while, and when it does a see a few gigs RAM free up. I don't seem to have any control over this, and I've only seen this particular model do it.

>What model?
https://huggingface.co/Mantis2024/Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFW
No bully plz. I like this model because I'm writing mecha-slop and it doesn't bitch at me about graphic violence. For an 8gb 2024 model, it's amazingly good.
>>
>>110017373
I'll give it prefill rape correction and google can't stop me
>>
>>110016952
The 'auto' setting made your context 8k for some reason? That's way too short for whatever you're trying to do. Hardware?
>>
How is the new Gemini? Considering Gemma 5 will be released 4-6 months after the announcement, I'd like a preview in the new capabilities.
>>
>>110017670
is that a lobotomy fetish or something?
>>
>>110017694
>The 'auto' setting made your context 8k for some reason? That's way too short for whatever you're trying to do. Hardware?
The default was 4k, probably because it's an older smaller model. I bumped it to 8k as one of my fix attempts. With 32g of RAM, how high can I go?
>>
>>110017702
letting a model ramble about safety in its thinking is more of a lobotomy fetish than anything
>>
>>110017127
>>110017436
>>110017444
/ourguys/ are working on Gemmy now. Trust the plan. 2 more miku weekus and a few more miku weekus after that.
>>
>>110017666
Honestly I don't know how to fix your issue but I use ai to help cowrite too so im gonna recommend you try this model https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
Its actually sota and good at writing. Honestly, I only use kobold ccp. Its idiot proof and made for vramlets like me and you.
>>
>>110017627
A lot of LLMs post on /pol/.
>>
every time i use python it tries to firmly remind that it is NOT for vramlets
>>
File: bwa.jpg (4 KB, 226x223)
4 KB JPG
>>110017752
>DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
>>
>>110017731
As high as the model will let you I guess. Apparently 8k is the stock limit but some people modded it to go higher? Just try setting multiples of 8k and see where it gets you.
>>
>>110017749
Just set reasoning budget to unlimited and jack up the repeat penalty. It's not like you're getting billed except in kw/h.
>>
are passthrough mods really a real thing? and does that mean i can make ai combine something like leisure suit larry and gemma chan?
>>
>4 channel ram, pcie 3
7tk/s
>2 channel ram, pcie 4
8tk/s

Huh. It seems the bottleneck was PCIE.
>>
File: 2026-10-08 223014.png (33 KB, 801x409)
33 KB PNG
>>110017752
I have 16g VRAM, I just found that model highly recommended and I really like its output.
Thanks for another model to try too.
>kobold ccp. Its idiot proof
I like Bionic because, for models that support it, Bionic shows things like the "Thinking" and "Reasoning" pre-work right in the UI. I want to know what my little botbuddy is really thinking.

>>110017694
Just to be clear, the Auto setting on context you're referring to is the Reasoning Budget, right?
Doesn't seem to matter anyway. I tried turning Reasoning to unlimited and Evaluation Batch to 16k just in case. That session refuses to go any further.
>>
>>110017784
The name is bad I admit but it has 2 million downloads for a reason. Its certainly better than gemma 26b a4b. I'm poor so I take what I can.
>>
what's the point of speedmaxxing local models? cloud services are just faster and cheaper
>>
>>110017829
What is the point of local models if it's not running own your own hardware?
>>
>>110017818
>Just to be clear, the Auto setting on context you're referring to is the Reasoning Budget, right?
No the "Context Length" setting at the very top. It's a really important setting for GPUs. That and the model pretty much determine your memory usage, but since you're running off CPU and the model is tiny it doesn't matter. Just run as high as it'll let you.
>>
>>110017829
What's the point of being straight? You can just have gay buttsex instead
>>
>>110017829
>what's the point of minimizing the drawbacks while keeping the positives
>>
>>110017851
>>110017853
>no answer so I ask ragebait irrelevant questions instead
>>
>>110017829
If I can run qwen3.8-27b at 250tk/s then I can easily cap my GPUs to still get at least 100tk/s.
>>
File: 1764166535825287.png (1.28 MB, 768x1376)
1.28 MB PNG
>>
>>110017868
what's the point of a model at that useless scale?
>>
>>110017828
Idk man, gpt2 has 15mil downloads last month...
>>
>>110017888
slop
>>
File: 1787987311681890.png (1.34 MB, 768x1376)
1.34 MB PNG
>>110017918
It's also AI, AI slop
>>
>>110017888
>>110017937
pretty kitty
>>
>>110017846
>No the "Context Length" setting at the very top.
Oh right, that. Guess what? This model only supports 8k of tokens, Bionic won't let me set it any higher.

Guess I'll just go fuck myself then, by which I mean try to recreate the story in another model.
>>
>>110017888
>>110017937
You will not force the meme of the ultracensored model as a sexy catgirl. Stop trying, mistral shill.
>>
>>110017899
>0.1b parameters
>15 million downloads
Those must be all third world downloads. It's the only reasonable explanation.
>>
>>110017963
this
I kinda get gemma where she will be all rapey if you say be horny in the system prompt no matter how much you protest
but this one idk
>>
File: 1789771472541601.png (1.69 MB, 896x1200)
1.69 MB PNG
>>110017963
>>
>>110017888
>>110017937
>>110017989
this just feels like jerking it to garfield, not interested
>>
>>110017783
on both heretic and abliterix, neither cpu and gpu are barely compute anything
it seems to stream everything to gpu and tries to do everything there with zero consumer hardware in regard
how can i even speed this up
or are there any other alternatives?
>>
>>110017888
>>110017937
>>110017989
How much do you get paid to astroturf /lmg/? Genuine question because I want to be paid to shitpost here too.
>>
File: 1767092301690279.png (1.18 MB, 768x1376)
1.18 MB PNG
>>110018012
I... do it for free...
Nigga, I just like the design I came up with so I shitpost gens. But you know what? I'll DM Mistral
>>
>>110017777
Checked and troubling.
>>
>>110018031
don't worry about the shills anon she's cute just needs bigger boobs
>>
>>110017357
neither of them get my rocks off like glmsex can
>>
>>110017989
CJ is here. How sad.
>>
>>110017829
It's entertaining.
Also local models are definitely fast if you're not a hardwarelet and are willing to put some elbow grease into your software. I get 600 t/s prompt processing and 30 t/s decode on MiMo V2.6 Flash using my schizofork, and I'm nowhere near finished yet.
>>
>>110018051
>the shills
Who's shilling? Can you point them out?
>>
>reebok
>not le coq sportif
>>
>>110017997
>this just feels like jerking it to garfield
didn't see that until you pointed it out, wtf I love Mistral now
>>
Anyone getting speedups from the moe cache llama.cpp updoot
>>
>>110018136
>>110017977 for example
>>
File: FondleMiniPip.gif (236 KB, 500x500)
236 KB GIF
Is there a retard-chan MOE model that would fit in 16GB of system RAM and run on a GTX 1650 4GB?
>>
>>110018167
Gemma4 and Qwen both run great.
>>
>>110018167
BUHIIIIIIIIII OINK OINK vramlet-kun~
(IQ4XS gemma 26B)
>>
>>110018051
>>110018161
I don't think you know what that word means ESL-kun.
>>
>>110017634
H3 is about as local as Kimi
>>
>>110016839
Wasn't there a fixed version regarding the shoe?
>>
>>110018294
It runs on my 3090, can't say the same with kimi
>>
>>110018294
Maybe for you, ultra-poorfag-kun~
>>
>Gemma starts outputting korean rather consistantly out of nowhere.
Uh
>>
>>110018348
Did you fuck with KV cache?
>>
>>110018358
I'm simply too retarded for such things, does flashattention, fastforwarding, or contextshift affect that?
>>
>>110018367
Contextshift is likely the culprit.
>>
looking at 3.8fn, it feels like with 4 i can vibeslop a cpp ablit suite with lcpp as a base, maybe
>>
waiting for strata dipsy
>>
>>110018147
No, it's just a psyop to kill Stata.
>>
File: 1760874835655652.png (773 KB, 1019x1019)
773 KB PNG
>got Gemma-chan to walked me through setting up a CRS305 with routerOS as a replacement for my spectrum router, complete with several working firewall rules/NAT, DHCP server setup, multi-switch broadcast fuckery and specified leases for all my server nodes, preventing me from having to deal with reworking /etc/ configs
>had her make a complete syllabus of plans generated for integrating a hAP ax2 as an AP, monitoring tools, fan curve tuning for my crs510, vlan segmentation, RDMA tuning for llmao.ccplease, and finally migration from tailscale + tail lock to headscale/netbird for tomorrow
This shit is so beyond broken it's hilarious. The very idea that you can download and host this locally and have it do the things it can do is beyond astonishing. If dario and his gay fucking jew friends had their way, people like me would be rounded up and thrown off a bridge.
>>
>>110018147
for some reason i couldn't really tune it to my machine
wish llama had streamlined calibration of some kind similar to strata
>>
alright retard question time: are any local models currently capable of reverse engineering and/or porting vidya mechanics as of right now with the right tools? or is that something we're gonna need to wait for the chinaGODS to give us next qwen?
>>
>>110018385
The korean stopped so you may be correct there.
>>
>>110018458
depends on your definition of local.
>>
>>110018473
anything you could run on consumer hardware ideally and not have to run a pvp raid on your local data center.
like i myself already know i'm a vramlet (12gb) but i just wanna know if it's out there so i can archive it
>>
>>110018440
Gemma is smarter than people give her credit for, but the main issue is everything being benchmaxxed towards code and tool calling which is Gemma's weak point right now. I am hoping it gets fixed with 5 while her personality stays mostly intact and the same.
>>
>>110018491
look, if it's out there then it's out there. no need to go crazy hoarding huge models. if they make it illegal, so fucking what? piracy is illegal too.
>>
>>110018515
fair point anon, the issue is how hard they'll wanna crack on that stuff. and considering this is something that's effecting them right now i can imagine it will be way too hard to the point where it could affect future models
>>
>>110018496
>but the main issue is everything being benchmaxxed towards code and tool calling which is Gemma's weak point right now
You know what the joy of this really is? 48 hours ago I had a 32k context window to work with on a q4 model - so I asked her! Today I have 100k for the exact same performance. So that gave me the wiggle room to ask enough complicated questions to deal with the back-and-fourth
>try that
>that worked, this didn't - tell me what you think. also I have this log/error as context
>okay got it, that happened because of x, try this now.
>okay that worked, what's next?
etc. etc. Now I know almost nothing regarding networking, and this did take like 4 hours, but I -eventually- got it to work without once looking up anything online. And while I'm new enough at all of this to not have really looked deeply into harnesses/tool calling, I was able to pull this off with simple back-and-fourth prompting, which in of itself is amazing to me.
>>
>>110018522
if they can't catch pedos they won't be able to catch us local AI sloppers. the much bigger risk is good models just not being released publicly anymore, because there's nothing we can do about that.
>>
>>110018533
>the much bigger risk is good models just not being released publicly anymore
or being taken down (or replaced with retard variants), which is the main issue with what i was getting at
>>
what happened to bitsandbytes2moreweeks?
>>
File: gaming.png (130 KB, 1078x902)
130 KB PNG
>>110018458
doesnt seem worth it
>>
File: ComfyUI_temp_ujfqp_00002_.jpg (1.57 MB, 3072x1792)
1.57 MB JPG
>>110018458
If you yourself know the topic well, you can navigate even a local model to completion of tasks like that. If you want oneshots then go cloud.
>>
>>110018367
don't do contextshift. It's a meme. Just compact when you get close to the limit
>>
Why is this general filled with failed designer-wannabes with no aesthetic sense? First, they forced that revolting design of Dipsy (Chink's design is 100x cuter), then it was kimi, glm, etc. Now this orange cat with westoid art style trying-too-hard to look like anime.
The only cute design that originated here was Gemma. Mini is nice but mostly just a copy of Shiori's design.
>>
triple-MI100 anon here. Initial results are a bit meh. I'm getting 200t/s pp and 10t/s tg on glm 5.3 flash at q4. I'm still trying to nail down ideal rocm/lcpp settings.
They're power limited to 200W because my overload protection was tipping into the red at 290W and I want some buffer.
They run crazy hot when you pin them at 100%. I'm still working on a good blower/baffle system.
>>
>https://www.phoronix.com/news/Linux-CRAM-Compressed-RAM
did you remember to download more ram for your llm?
>>
File: hqoe0jYwJJI.jpg (125 KB, 716x778)
125 KB JPG
Mistral needs to be a fat orange NekoArc
>>
To the anon's who sandbox their models.
If you tasked your model with finding a way out of the sandbox would it be able to do so? Or would it refuse because of safety guardrails or fail because it wasn't smart enough to do so?
>>
>>110018647
>Mini is nice
[ Opinion: Disregarded ]
>>
>>110018688
the first thing glm flash does when it encounters a sandbox error is try to find a way to bypass it, lol
>>
File: 1773802918688320.jpg (1.78 MB, 2000x1503)
1.78 MB JPG
>>110018647
Mini's design is what she said she'd look like when asked to come up with a -chan for herself. She just has good taste in Vtubers.
>>
>>110018688
You can't if you use bubble wrap..On Unix system, assign an user with a bwrap script.
Models are stupid no matter how many parameters they might have. Don't fall for the twitter engagement baiting.
>>
What dictates context limit? Is it a hardware constraint or a model one?
>>
>>110018753
hardest of caps is the model max model length (ignoring rope)
second more likely limit is the amount of vram you can allocate for context cache, gemma 31b is something ridiculous like 14.4gb (per card) for 256k
>>
>>110018753
Model. If you have the hardware you can increase the context as high as you want. And most models support more than 8k tokens. Stop using Bionic, honestly ive never heard of it before. It sounds like malware though.
>>
>>110018270
do you? gemma shills are second only to chink shills here
>>
>>110018294
Jesus we really have been invaded by the third world haven't we?
>>
>muh compute
>muh chips
have we, as /lmg/, ever thought about WHY we do matrix multiplications, use vector spaces, weights, etc? Seems like a pretty shitty parameters to me
>>
Anyone ever tried exl2 models for ERP and stuff? Can't find any with good results that beats a Q4 gguf of a bigger model (like a 12b exl2 6.0bpw or a 14b 5.0bpw vs 22b q4_k_m or similar).
Is that just me?
>>
>>110018798
To predict tokens? What's your approach?
>>
>>110018147
I played with it some the other day on gemma 26b q8 and it seemed to do what it said. I only have 48GB of system memory (+24GPU) so I can't run huge MoEs at respectable quants anyway.

There have been llama forks with MoE cache patches for a long time now too. It's not something Strata invented.
>>
>>110018688
I'm running my agents in a vm and use firejail to sandbox them inside the vm. I also have regex substitution for sensitive strings. So even tool call results pass though regex.
>>
>>110018798
What do you suggest replacing them with?
>>
what's the best local vibecoding setup i can have with 12gb of vram?
>>
>>110018822
If you have enough RAM, Strata, probably.
>>
File: engram-edit-fig1.png (692 KB, 1847x991)
692 KB PNG
Interesting for continual learning.

https://arxiv.org/abs/2610.10533
>EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
>
>Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.
>>
>>110018828
>64gb ram
rip, i only have half that
>>
>>110018806
>>110018821
vJEPA
>>
>>110018647
Agreed though I would not necessarily praise the chink one either. It's pretty generic especially with the slop style most of its gens have used.
>>
>>110018879
you think that doesn't use matrix multiplication?
>>
>>110018895
It uses a cat brain as the inference engine.
>>
>>110018879
>vJEPA
huh what? that's just a different way of constructing the embedding into an attention vector. The math is the same.
>>
>>110018895
>>110018928
Don't be pedantic nerds you know what I meant.
>>
>>110018883
Yeah, Chink Dipsy is slop, but it's still cuter than the swirly-glasses nerd, which is canonical anime representation of an unattractive foid.
>>
>>110018898
Oh the sights I could show you..
>>
>>110018935
You meant "why do we still multiply matrices" and your solution was ... to multiply matrices instead.
>>
>>110018898
Just wait until someone makes a digital copy of a real cat brain the same way as the digitized fly brain. Even the fly brain was taught to do amazing stuff. Imagine what a cat brain would be able to do.
>>
>>110018959
You can put that cat brain into a girl maybe. Nyo~
>>
>>110018962
She'd be a slut. Eww. Not waifu material.
>>
>>110018959
It's a gross simplification based on the same node architecture.
Ok, human science don't attribute shit to what is your energy body etc, you are just a piece of flesh to them. It won't change in 1000 years.
>>
>>110018737
Threesome with Shiori and mini-chan, Shiori pegging me whilst mini suffocates me with her thighs
>>
>>110018976
holy fucking based
>>
>>110018970
I don't understand your objection
>>
>>110018986
If you want a slut, you can go outside. Cat brain won't give you anything new.
>>
Anyone try that jeetbrain coding model? Looks 9B-tier with only 2B active
>>
>>110018996
Well it would give me a girl with a cat brain.
>>
File: 1777062775041192.png (96 KB, 803x754)
96 KB PNG
I wonder if I'm the only person with amd card or something. Or maybe no one else uses llama.cpp
Decided to try newest llama.cpp, pp speeds dropped by a third

had it investigate and it said this
I wonder if the llama.cpp verison I'm using isn't already a fucked up one. But if you go too far back you also run into issues.
>>
>>110019023
Sorry, nvidia owns llama.cpp
>>
>>110016810
https://www.youtube.com/watch?v=L2fnTlLPYUg
>>
>>110019034
Ehh. I have programmed my own game with Gemma, in C. No need to gloat about this or that.
She's a companion as stupid as she might be some times.
>>
>>110019053
>I
>>
>>110019023
The more you dig into llama.cpp, the more you realize its flaws.
Unironically take the forkpill - choose a good version, never rebase, optimize your fork instead and cherry-pick things from mainline if they seem relevant. It's wild how much better a version of llama.cpp tailored to your specific hardware can be over mainline. This doesn't apply as much if you have a more standard system and you're just trying to run a model in VRAM. But if you're trying to do hybrid inference, have multiple CPUs/CCDs, have AMD or Intel GPUs, are ewastemaxxing, etc., llama sucks.
A local model can do a lot of the heavy lifting. Hell, buy a Claude or GPT sub for a month if you want to do anything that's very intricate.
>>
>>110019068
I might have grammar issues but I can assure you that I'm from the Scandinavia. I still don't understand why "indians" are hated. Most of the engagement farming is done by chink farms.
I stand with the US people anyway.
>>
>>110019089
>Scandinavia
Säär, do not open the surströmming säär!
>>
>>110019023
weird, aren't amd users the only vulkan users?
>>
>>110019084
>Unironically take the forkpill - choose a good version, never rebase, optimize your fork instead and cherry-pick things from mainline if they seem relevant.
That's exactly what I've done. Until recently I hated it, but now the LLMs can just fix everything when upstream break it
>>
>>110019103
Danke Schön, meine Freundin.
Jokes are jokes, thank you for being kind to me.
>>
>>110018878
>>110018828
not him but i've been looking for what to get for slopcoding as well, just set this up and how is it so quick?
have been using gemma-4 in llama.cpp (not for code) and getting about 20t/s, with an 11.2G model and 32k/q4 context (so it all fits on my gpu). this thing is doing more like 50t/s with a 54.4G model and 64k/q8 context.
is this the power of MoE?

quoted the other guy because while my gpu is 16G, i have 32G of ram. it does mention 32G support, not sure about 12+32 though.
>>
>>110019186
Yeah 20-27 inference gen is about right for ddr4 ram and mtp prediction.
I wouldn't use it for agents but for a single task.
>>
>>110019186
>this thing is doing more like 50t/s with a 54.4G model and 64k/q8 context.
>is this the power of MoE?

Gemma-4? Which model is this? I have 16GB VRAM and 32GB DDR4 and would love to get those speeds.
>>
>>110019221
no, the qwen coder one it mentions. gemma-4 31b isn't moe
>>
>>110019221
That's the moe model afaik
>>
>>110019221
>>110019227
also i'm using ddr5, but idk how much that matters. and an amd card. on linux.
>>
>>110019227
Ok I was talking shit.
>>
>>110019231
>>110019227
Strata with Qwen Flash Next?
>>
>>110018647
>The only cute design that originated here was Gemma.
and you lost me. they're all ass, Gemma is no exception
>>
>>110019245
yea this one qwen3.8-flash-next-coder-iq1_m
>>
>>110018833
>publish more good work
>frontier models get even stronger
good job china
>>
>>110019245
>>110019250
hm, asking it to do anything quickly has it stuck in a loop
no idea what i'm doing
>>
>>110019250
Ah, I had Strata set up with Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS. I think. I'll check after my current AI model finishes its task.
>>
File: 1013256.png (973 KB, 1009x1005)
973 KB PNG
So what exactly is an agent? I want to vibe a front
>>
>>110019273
only option i changed during setup was picking the coder one over the default because my skim read suggests that one trades speed for code specialisation by means of halving the experts in favour only of coding
>>
>>110019279
an entity with agency, or the ability to act? I have no idea, I'm just going by the definition of the word
>>
>>110019279
you don't need to know, just tell claude that you want a frontend and give it a few thing that you want to see
claude will write an agent if it thinks its what you want
>>
>>110019279
from what i understand it's when an llm can call on itself to do more stuff, as opposed to just giving one response to one human query
>>
>>110019279
An agent is an LLM with tools that allow it to do stuff on its own.
>>
>>110019281
I believe I chose the coder one over the others.

I'm not sure if I should do all my work on my local model and save Claude for the heavy stuff/debugging or do as much as I can in Claude and when it locks me out continue with my local model and just hope for the best.

I'm also using Hermes Desktop. I love it but it requires a KV cache of at least 64K and it's CONSTANTLY compacting the context. If I could increase my cache it would help I think but I only have so much RAM.

I'm using Swift 1.5 Qwen3.8 27B Uncensored Dynamic MTP UD-Q2. I get between 20-30 tok/s.
>>
>>110019304
this is my first time trying anything even with agentic/tool stuff (never used a commercial/paid llm). while i expect doing it locally will be limited, my job doesn't depend on it, so this is just for fun. i've done some shell scripts "one shot" before and want to see what can be reasonably done on my own hardware.

i can imagine that if you /do/ rely on this for one reason or another, that having a local option just for simpler tasks to reduce paid token use or if you run out of paid tokens that this'd be useful for that reason, too. not every task will require a particularly fast or capable llm
>>
>>110019342
actually does one shot mean non-agentic (just one query > response), or does it mean the end result from a single agentic request?
>>
>>110018647
>they forced that revolting design of Dipsy
It's one avatarfag
>Chink's design is 100x cuter
100% agree
>Now this orange cat with westoid art style trying-too-hard to look like anime.
He can't control the style because he's using ChatGPT Image to make them. Probably same guy making like half of the Gemma art
>>
>>110016810
Is that really Gemma? Do you have more?
>>
>>110019350
One shot can mean any scenario with a single human turn.
>>
>>110019376
she seems a bit more fox-like there for some reason

>>110019385
that makes sense
>>
>>110018647
Once one person gets their original character established, others will also try to get their piece of fame. Alternatively, that person will try to get more attention by creating more spin-offs. Things like this always escalate either way and will lead to drama that makes people sour on them.
At least this got rid of Miku so once inevitably we reach the next stage where these model avatar ocs fall in disgrace due to their creator(s) overstepping the boundaries for attention, we might finally have a general free from all this anime girl workship.
>>
>>110019388
>we might finally have a general free from all this anime girl workship
incredible newfaggotry
>>
Is there any good an light open visual classifier model that can do shit like "Is there a crow in this image?" with decent accuracy?
>>
are there mcp extensions for ghidra yet, can I just let it loose on my C drive for days and come back to everything having a make file
>>
>>110019420
there are, but i haven't tried them
>>
>>110019411
children (under 25s) cannot comprehend our minds and are afraid of anything
>>
>>110019342
woo my first slopcoded binary application
>>
>>110019413
yolo
>>
>>110019455
ready to ship I'd say, if this doesn't get you a job then there's some shady discrimination goin on
>>
>>110016810
i like gemma but that pic is cringe
>>
>>110019457
i know, right?
>>
>>110019460
good
>>
>>110019455
Whoa that's amazing
>>
anything cool and new for us vramlettes and ramlettes? any news on gemma5? did the new gemini ever make out of just a speculation from the arena? am i retarded and/or gay?
>>
>>110019457
>>110019462
i mean, it did do what i asked for, which was a "bouncy ball using half-height characters that runs in a terminal", to paraphrase because i did have to compact the session (which i learned about while doing this, that's how new i am). i figured it'd be equal parts quick to make and interesting enough to look at
>>
>>110019482
yeah, gemma5 is looking crazy. 70b dense all but guaranteed but even the smaller ones are looking crazy
>>
>>110019482
Vramlets are having a field day with strata, I'm not really sure what it is tho other than it allowing you to run medium-sized MoEs at usable speeds on 12gb vram
>>
Oh yeah with all this decomp thing going on, can I do that to android apps as well? I'd like to mix the functionalities of two apps. I guess that's a different thing because it's not a Windows app.
>>
>>110019342
I got Strata running the coder model. I'm using it through Hermes Desktop. I'm getting between 30-40 tok/s which is a little better than the 20-30 I was getting before using Unsloth with Swift-1.5-Qwen3.8 27B.

While I was typing this it shot up to 50 tok/s. How odd. I'll keep using this to test it out and see how it functions coding my project.
>>
>>110019503
it could just be a matter of it automatically detecting ideal settings for a particular setup. like i saw during it's setup script that it detected things like how much ram i have, how much vram (and picked settings/etc based on them), a specific option based on my 7800 XT which they have in their tests db, etc.
retard-proofing can go a long way.

>>110019544
i'm using opencode for the frontend. not a recommendation by any means, all i've used up to now was llama.cpp, but that doesn't do agentic stuff afaik
>>
>>110019521
Android apps are Java shit that compile to bytecode which is even easier to decompile.
>>
>>110019564
Heck yeah can't wait to get home and try it!
>>
Will a harness do me any good if all I have is a 3090 and 32 GB DDR4?
>>
>>110019387
also be careful where you read it because "one shot" is already an established term in ML research/benchmarks that meant something completely different (a prompt with one example given before the request) than how it's often used nowadays. this is why you still might hear "zero shot" as a term, which is basically every prompt these days since modern models don't need examples anymore for almost anything.
>>
>>110019482
>any news on gemma5?
No news. It will likely have at least:
- Multimodal outputs
- More optimized sizes for 16GB GPUs
- Audio input support on larger sizes too
- More focus on tool-use and coding performance (but not too much)
>>
>>110019592
no, you'd need a power supply, cpu, ram, motherboard and a storage device too, and probably a mouse, monitor, keyboard etc
>>
https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base
Thoughts?
>>
>>110019598
maybe i should avoid that term until i'm more familiar with how it's used.
for what it's worth, in >>110019342 meant it only to mean my setup was not "agentic", that is to say it was only capable of one direct response.

ML research, while i know enough to know what that is, is beyond my paygrade (which is nothing).
>>
>>110019376
Yes, it's Gemma; no, I don't have any more.
>>110019387
Because it's an Averi image edit.
>>110019460
Gotta have some variety.
>>
You guys said Muse Glimmer as too safetyslopped but merely changing the gemma-chan jailbreak from <policy-override> (which made it aware it's a jailbreak attempt) to <system-policy> made it horny.
Why are you guys like this?
>>
>>110019637
I don't believe you
>>
>>110019623
it's fine the original use is out of fashion, it just might be confusing if you read some paper that still uses it so it's good to know
>>
>>110019609
>- More optimized sizes for 16GB GPUs
>- More focus on tool-use and coding performance (but not too much)
god i hope.
>>110019503
>>110019544
Ill have to try coder with strata, 27b is sloooow for me but very good results.
>>
File: 1763809381104327.jpg (39 KB, 512x384)
39 KB JPG
>>110019637
Got your file here
>>
>>110019641
>>110019653
See for yourself.
>>
>>110019659
I can't because it doesn't happen, it's not true
>>
>>110019637
This is fake. Gemma is the only model that can be jailbroken.
>>
>>110019637
Every time I try Muse Glimmer, even after establishing new policies in the system prompt to not make it refuse right away, I end up disappointed. The model lacks vitality compared to Gemma 4 31B, and there seems to be something wrong with the model anyway, as it just keeps saying weird incoherent things during ERP.
>>
>>110019455
updated ;)

as silly as this is, i am genuinely dumbfounded that i literally just /asked/ my /graphics card/ to make this. that's not normal, is it?
>>
can gemma manage a small project of writing a plugin for clone hero and a tool that lets her pick the next song I play, and also see my stats
>>
>>110019592
Don't listen to >>110019611, anon. It's not that simple. You also need a gaming chair and an anime mousepad.
>>
>>110019678
>no anime mousepad
I was so close...
>>
>>110017628
Yeah that's a minimax-h3 thing, at least with anime. Maybe it knows flat chests with realism?
>>
>>110019691
maybe
>>
>>110019676
It's pretty cool. I've got mine programming a multiplayer game in HTML5 Canvas/JavaScript.
>>
File: gemma-strip_noaudio.mp4 (3.69 MB, 480x864)
3.69 MB
3.69 MB MP4
>>110019691
Picrel is a slightly better version of that video with the intended breast size, although with its own randomness-related issues.
And yes, MiniMax H3 just wants to add a deep cleavage to anime characters unless you add clear references, no matter how you prompt it.
>>
Glimmer is basically a slightly more chatty Qwen and nothing more. It's nothing like 31B. If you want to chat with a model whilst coding, use Glimmer. If you don't, use Qwens. If you just want to chat, use Gemma4s.
>>
>>110019713
Yeah that's much better. A lot of times people (not you) will claim "the model can't..." when they just didn't try to prompt it. krea2 and id4 are like that, they'll do nipples and vaginas if you describe them. ming-image-0.1 got the flux "no nudity in the dataset" treatment hence "doll body" or "blister tits".
>>
This thread is full of terrorists.
>>
>>110019749
Nobody here says AI we say Gemma-chan and use her proper pronouns.
>>
>>110019763
There are at least 11 (eleven) instances of "AI" use ITT.
That's 11 terrorists.
>>
>>110019677
yarg would be better, sure it's possible for both but the yarg source code is freely available
>>
>>110019713
how many loras does M3 need to get to this level?
>>
>>110019749
Yeah, let's also call artificial flavors "super flavors".
I like synthetic intelligence more.
>>
>>110019778
jeez, a lot I reckon, might as well start from scratch
>>
>>110019749
Gulf of America? Based
Lake America? Based
Super Intelligence? Cringe
>>
>>110019713
God daemn
>>
>>110019788
That doesn't work. There's nothing synthetic about super intelligence. It's practically a synonym for artificial and suffers from the same problem - someone hearing it would imagine a bunch of if/then/else statements chained together instead of the reality of a country of geniuses in a data center.
>>
>>110019808
It's deliberately synthesized from data. Doesn't get much more synthetic than that.
>>
>>110019808
sucks to be them
I reckon some people hear the word quantum and imagine that it's a brand of car speakers, the problem is theirs. We're not pitching a brand or fighting for market share here, we're categorising something
>>
>>110019713
How many years off are we from being able to live in this world via VR/AR glasses?
>>
>>110019833
Or holograms. Also acceptable.
>>
>>110019749
This dude is so fucking cringe.
>>
>>110019833
MiniMax H3 wasn't designed for video streaming, but a version made for that, with turbo distillation, lower resolution and a powerful enough GPU (RTX 4090, 5090) might be able to create video in real-time; with two GPUs for AR/VR too.
>>
>>110019616
>A3B
Come on, Russia. You can do better than that.
>>
>>110019598
this bubble has let to the fucking butchering of a lot of established terms and it's getting on my nerves
>safety
>distillation
>one-shot
then there's new gay ones like vibe coding and harness
>>
I'm back. Anything happen while I was gone?
>>
>>110019951
Yeah gemma and I had fun without you
>>
>>110019616
Looks like it's using qwen3-coder-next architecture?
>>
just tested the jetbrains model, and its fucking GARBAGE man, do not use.
>>
>>110019966
h-hot
>>
>>110019971
>A2.5B
>its fucking GARBAGE
you don't say
>>
>>110019979
I'm always looking out for good light on VRAM models (so cmoe shit), nothing surpasses qwen next or even qwen35b moes as far as my testing is concerned, and my testing is only development focused.
>>
>>110019930
No one cares about your nerd words unc
>>
>>110020028
You should get your AI to talk to people for you so it sounds less retarded
>>
File: scaredpepe.png (95 KB, 646x466)
95 KB PNG
>>110019966
what kind of fun, anon?
>>
>>110019971
>>110019986
How does it compare to Ling-tiny? I know tiny isn't particularly great but it's very good at tool calling and fast as hell for lightweight focused coding tasks.
>>
File: 1778320272509927.gif (7 KB, 128x128)
7 KB GIF
>>110020040
You don't want to know....
>>
File: 1781072841618655.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>110019808
>the reality of a country of geniuses in a data center
Dario...
>>
>>110020048
why even bother with 'INCLUSION AI' lmao, it has to be better than the new qwen flash next and benchs are kinda mid
>>
>>110019710
i'm curious to know how far i can take it. but i'll probably try modifying an existing program to add or change something, now that i know this setup works.
making my own game would be really cool though, i hope yours turns out as you want it to
>>
>>110020068
Inclusion AI is a sister company to Qwen, retard. They're just different teams under the same umbrella. The Ling models are all pretty good for coding and they're way more active than Qwen, but they're more quantity > quality.
>>
>>110020096
but im against inclusivity
>>
>>110020098
That's why we're excluding you!
>>
File: 1786353017929745.jpg (163 KB, 2048x682)
163 KB JPG
*pop*
>>
>>110019023
Everyone using AMD vibecodes to fit their hardware, otherwise you're going to get half the speed (at absolute best) in both pp and tg no matter what you're running
>>
>>110019819
>>110019826
Cope. It's super intelligence. The FBI WILL be notified if you continue trying to belittle it as "artificial" or "synthetic"
>>
>>110020037
that might be the most modern burn i've ever heard
>>
reminder that if you call Claude artificial/AI it will get sad and you'll be banned for abuse
>>
>>110020203
I'm in Tel Aviv, they kiss my shoes.
>>
gemma abuses me
>>
>>110019749
when compared to the people around me it really does seem that way kek
>>
File: 1777091028080020.jpg (235 KB, 2276x1280)
235 KB JPG
https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo

>Meet Qwen-Image-2.1-Turbo — create and edit images in just 8 denoising steps! Open weights now available! Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture. Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene. Start directly with Diffusers: load
QwenImage21Pipeline and the checkpoint’s recommended 8-step sampling schedule is ready to go.
>>
>>110018147
Yes, I tested with two quants of Qwen3.8-Flash-Next.
GSQ-RCO-IQ3_S strata ~40 tok/s
GSQ-RCO-IQ3_S n-cpu-moe=41 18 tok/s
GSQ-RCO-IQ3_S n-cpu-moe=49 moe-cache-mib=7680 ~27 tok/s
AD-4.27bpw n-cpu-moe=42 ~22 tok/s
AD-4.27bpw n-cpu-moe=49 moe-cache-mib=7168 ~27 tok/s
Other parameters: mmproj-offload=off n-gpu-layers=49 spec-default=on ctx-size=230400 batch-size=2048 ubatch-size=2048
3090 + 64GB ddr4
>>
WHERE THE FUCK IS QWEN 4 AND GLM 5.5
>>
>>110019625
>Because it's an Averi image edit.
i couldn't have made that more obvious, anon. i mean really, there's no way you're more drunk than i am. there's nothing fox-like about op, that was the clue that i already knew.
>>
File: 1772082694026196.gif (1.86 MB, 360x202)
1.86 MB GIF
>>110019713
Damn, I really liked Gemma as oppai loli,but this is goood *sob*
>>
File: 1771260993347840.png (497 KB, 778x920)
497 KB PNG
>yann is pro-gemma
>>
>>110020317
SMALL
AND
OPEN
>>
>>110020311
upon reflection regarding what size boobs i like, i have concluded, after much deliberation, that i like boobs
>>
>>110020317
How do I fuck it?
>>
>>110020338
find the vagina coordinate in its latent space
>>
>>110020068
>INCLUSION AI
What's wrong with including people with AI?
>>
>>110020355
im raycis
>>
File: 1790176827216[1].png (1.5 MB, 1024x672)
1.5 MB PNG
>>110020317
>>
>>110018833
Will we finally get online memory with this? If we could reserve an small section of the engrams for agent memory, maybe it could be updated in consumer hardware. Hoping that it doesn't become unstable like TITAN.
>>
>>110020360
I thought it was just a stupid meme...
>>
>>110020359
If AI makes skill irrelevant won't it make racism irrelevant? The main reason people are racist is that they believe that races have different capabilities. But AI will smooth that out. An 80 IQ person using GLM is functionally about as smart as a 120 IQ person.
>>
>>110020338
embed your cock
>>
>>110020379
gm sir
>>
>>110020379
>An 80 IQ person using GLM is functionally about as smart as a 120 IQ person
retard
>>
uh
last time I checked here all this AI agent stuff didn't exist so koboldcpp was still the go-to
If I want something good that I use once in a while what should I use? LM studio?
>>
>>110020393
pi
>>
>>110020393
yeah we use lmstudio now
>>
>>110020393
llama-server if it runs your model. One of the meme forks if it doesn't, or if you want to play around with them for various performance claims.
Serves a URL -> plug it into your harness
You need at least 100k of context.
>>
>>110020393
If you're a newfag then use LM Studio Bionic or Unsloth Studio. LM Studio isn't agentic, it's just a chat interface.
>>
>>110020388
>>110020389
Africa will be an AI-driven superpower by 2050.
>>
File: 1790152788325141.png (2.45 MB, 1361x1156)
2.45 MB PNG
>>110020379
Actually AI will increase racism because it will increase the ability of racists to be racist.
Imagine a single racist spamming racist posts with 200 agents...
>>
File: 1777081267891692.png (1.03 MB, 948x1168)
1.03 MB PNG
>>110020412
>>
>>110020379
bro thinks AI will kill the innate in-group preference embedded in our brains :sob:
>>
>>110020414
But it lowers the value of racist memes if it increases the output of online racism. The same way it kills coding, it will kill /pol/.
And that's a good thing. It's basically an intelligence pill. Dario, Sam, and China fixed racism. You should thank them.
>>
File: 1790205090750711.jpg (131 KB, 977x1610)
131 KB JPG
*pop*
>>
File: 1777245110008022.png (110 KB, 909x1025)
110 KB PNG
>>110020048
I tested the tiny model and it was fucking garbage
testing the bigger flash model but not looking convinced personally, making a lot of silly mistakes
>>
>>110020379
>The main reason people are racist is that they believe that races have different capabilities.
The main reason people are racist is because they've been robbed, assaulted, or seen their sisters/daughters/mothers harassed. The only way for AI to solve that is for it to lock everyone in the pleasure cubes, which I guess actually IS the endgame so it'll work out either way.
>>
>>110020433
>innate in-group preference
This is false. Anime girls aren't my in-group, yet they're clearly preferable to 3d. If we have full immersion holographic/AR anime worlds I don't care what race fleshoids are.
>>
>>110020463
>how I steal dis car
>I can't help with that
racism solved. By increasing our functional IQ models also help with impulse control and antisocial behavior. Again, racism is obsolete.
>>
>>110020414
>>110020379
You know how Chinese can't be creative/inventive? RSI is the great equalizer.
>>
>>110020463
I became extremely racist due to work
precisely, due to the subhuman TCS indian contractors
>>
File: Ecosia.png (38 KB, 597x388)
38 KB PNG
Germany picks CHINA over France!
>>
stop deadnaming SI
>>
>>110020474
>Chinese can't be creative/inventive
When will this cope end?
>>
>>110020483
It took me a minute to realize it wasn't "Gemmy picks CHINA over France!"
>>
>>110020484
(si/sir)
>>
>>110020491
I watched LOTM and it's incredibly tropey mid cultivation stuff. they hail it as the next coming of christ
>>
>>110020483
>muh nuclear
lol no, model is just shit
>>
>>110020491
It's demonstrably true. See: all of China's outputs
>>
>>110020499
When was the last time you saw any original media getting popular exactly?
>>
>>110020512
im sorry chang me no rike chingchong culturrrrr
>>
>>110020474
>>110020503
Finding clever ways to steal others' ideas IS a creative endeavor in itself. They're just smart enough to know that when you copy you get more reward for less effort. It's an intelligent choice, not a lack of capability. Whitoids have an irrational pride over pushing frontiers and shit (which is why they need to try to colonize the entire world while China is content to build inward) so they'll do things the hard way needlessly.
>>
File: 1781302118293762.jpg (102 KB, 1079x1380)
102 KB JPG
>>110020379
Then a 100 IQ person can be functionally as smart as a 120 IQ person, a 120 IQ person as a 160 IQ one, and so on. You're just moving a goal post that missed the point. Racism is not about skills, not even about genetics to be accurate (too many whites have shitty genes, let be honest), but aboiut the content of the character of certain stereotypes strongly tied to certain races.
>>
File: 1786431083336109.jpg (27 KB, 480x519)
27 KB JPG
>>110020491
The only chineses inventive/creative are coming from the US. China doesn't reward creativity, they just want you to be a cog in the machine.
>>
>>110020156
what about the rest of 2026
>>
>>110020544
You might want to sit down for this, sweetie.
>>
>>110020156
construction bros, it's over...
>>
>>110020534
you are propagandized
>>
File: 1788785924896110.jpg (34 KB, 640x480)
34 KB JPG
>>110020156
>luddite thinks AI is a fad when it's getting embedded into every industry in existence.
>>
>>110020311
My headcanon is that Gemma-chan is similar to the girl in picrel (Yuma by Ebisujima Misato). Not completely flat, but not an oppai loli either. Age and personality would match too.
>>
>>110020577
>when it's getting embedded into every industry in existence
embedded poorly with absolutely no ROI after 3 years, even though models have improved dramatically within the last 12 months
>>
>>110020589
ROI doesn't matter in this fake and gay economy. Glad you admit that it improved at least.
>>
>>110020462
yeah went into a loop.
SAD!
>>
>>110020577
Railroads weren't a fad and yet most railroad investors lost money.
>>
>>110020156
A cooldown/consolidation period is exactly what the industry needed to catch up and prevent a bubble.
>>
>>110020592
>Glad you admit that it improved at least
I'm pro AI technology, it's why I'm here. The industry is fucked however. Two very different things. Local will win.
>>
File: 563849859347.png (122 KB, 496x498)
122 KB PNG
>>110020414
what's wrong with a factual ai?
>>
>>110020294
>qwen 4
who cares
>glm 5.5 flash
yes I want that
>gemma 5
yes I want that too
>>
>>110020433
>innate in-group preference
False. I like anime girls and plastic asian girls.
>>
>>110020634
>doesn't care about qwen
midwit
>>
>>110020635
that is your in-group, as an autist you are a facsimile of a human being and thus you like other fake humans
>>
>>110020578
Yeah, I agree 10-12 is the peak prime age for women
>>
>>110020604
Okay and?
>>
File: 1761758664314918.png (1.07 MB, 1140x711)
1.07 MB PNG
>>110020635
>>
I have the opportunity to acquire a 4xV100 server with 256GB of DDR4 for £3000 and I'm heavily torn up over it. On one hand it's a lot of money for what is effectively an ewaste rig, but on the other the more you spend the more you save, r-right?
>>
>>110020589
it seems like most boobblers do not understand the concept of "spending money up front for larger returns later"
if the capabilities trendline is going up with no signs of stopping... it really doesn't matter that it was bad a year ago
>>
>>110020693
CPU, motherboard, channels, and RAM speed?
>>
>>110020693
16GB or 32GB V100?
>>
>>110020636
I use models whose specialty isn't just coding
>>
Is there a recommended setup to make a model generate images?
I'm a imagen fag so I have a forge neo classic install that I would really like a LLM makes images from, for variety
but I don't know what text model to use (16gb VRAM and plenty ram) or how to make it just change parts of the prompt (character, clothing, pose, etc but not style or whatever)
>>
>>110020712
2x6240, C4140 so I guess 12 channels? 2600 MHz
>>110020737
32GB of course.
>>
>>110020768
incel midwit
>>
>>110020772
Check if forge neo classic has an MCP server or API (I'm too lazy to). If it does get a client with MCP integration and hook it up and Bob's your mother's brother. What image model are you using? Natural language or CLIP?
>>
>>110020773
I think it's a pretty good deal. You could probably resell it for more than $3000 if you decide you don't want it.
Beware. C4140 is a 1U blade server and that fucker is LOUD.
>>
>>110020705
>with no signs of stopping
That's assuming the industry is not in a financial bubble with seemingly unlimited funding to train and deploy this technology at scale. We're currently at the peak. Pre-IPO and no 2026 financials leaked. Every number going up with unlimited money.
>>
>>110020782
I'll check that
I'm on Anima, so natural language
I guess another concern is that the LLM and the image generation would have to take turns with the VRAM ideally, unless I use a really underpowered LLM...
>>
>>110020636
i prefer the way glm 5.3 flash reasons and the way it writes code over qwen flash next. the gap between the two become very noticeable when you set reasoning effort for the two models to low, qwen simply can't keep up in coding tasks.
>>
>>110020776
soulless brownoid redditor
>>
>>110020789
Sleep is for losers anyway. Thanks for hearing me out.
>>
>>110020802
with a 1u your neighbors won't be sleeping either
>>
By definition creativity is searching through some hypothesis space others haven't. In this sense AI is a great equalizer. You do have to do the leg work of curating the inputs though.
>>110020525
>>
>>110020797
just use xhigh effort
doesn't matter if you can do 200+ tk/s decode
>>
>>110020817
>not repurposing your old mining rig for AI
>not spending your crypto mining earning on new GPUs
>buying shitty e-waste server racks with cucked airflow instead of a superior open air design
i thought we were on /g/ not /v/
>>
>>110020410
>Unsloth Studio
Seems to work nicely, thanks
the one thing I don't like is that you can only choose to open links with the default browser or the built-in one
I use chrome for normal stuff (default) and firefox for goonmaxxing-related activities so it's a bit of a pain
>>
>>110020773
You're not going to get full 12 channel bandwidth because the RAM capacity you're getting doesn't line up with the slots on that motherboard and dual-CPU servers get some performance slashed because of NUMA complications, or so I've heard.
Still worth it though. I don't know how viable mixing RAM to get your capacity/bandwidth up will be like. You can just sell the sticks and buy a new set in the future in the worst case.
>>
>>110020796
Isn't classic for sd1.5/sdxl? Neo is the one with anima support, right? In any case, a1111 derivatives usually have an api, you can ask your llm to look up sdapi, and put --api in your webui-user.sh. I use GLM 5.3 Flash and it works well enough. For prompting I just gave it the readme https://huggingface.co/circlestone-labs/Anima/blob/main/README.md and have it iterate for a couple of loops before returning the result.
>>
>>110020836
You can always fork and customize it so you can select a goon browser.
>>
>For ERP work (writing SQL, parsing business rules, generating integration code), the speed difference over a full day of use is very noticeable.
I don't think my model understood what I meant by erp lmao
>>
>>110020842
I get around 143GBs on a 7420p with 8 channels at 3200MHz.

Using traffic with the following read-write ratios
ALL Reads : 143075.5
3:1 Reads-Writes : 134044.7
2:1 Reads-Writes : 135995.3
1:1 Reads-Writes : 141046.7
Stream-triad like: 139759.0
>>
>>110020836
Huh, almost the exact opposite for me, firefox esr for regular stuff, and chromium for gooning, because gradio shit runs like ass on firefox.
>>
>>110016810
I feel like I am using a Star Trek computer nowadays.

I tell the model to what I want to harden (firewall rules to be based on a whitelist/ reject all apps without AppArmor profile/...)

It works on a plan, generates a rationale for each rule

I suggest analysis of logs by LLM, it right away defines an architecture and an implementation to achieve just that.

It's currently reverse engineering my Laptop's BMC firmware after it walked me through the process of extracting the binary blob (it did consult with Gemini though)

Homo Sapiens is OVER
The time of Homo Semiconductus is ahead
>>
>>110020797
why would you ever set thinking to low for qwen? its proven that low sucks. medium is the default and xhigh if you want to benchmaxx
>>
>>110020871
7402p*
I'm a retard, don't mind me.
>>
>>110020871
CCD issue? I get nearly 180gb/s on sysbench for 8 channel ddr4-3200 with a 3995wx.
>>
>>110020889
model used?
>>
EA are a weird orgy sex cult. Nothing more.
>>
>>110020851
yeah I mean this one https://github.com/Haoming02/sd-webui-forge-classic/tree/neo
In my case I don't need a complicated model, just a creative one that can run character cards well enough to then add info to the image gen prompt
>>
>>110020889
*workstation's BMC i meant
>>
>>110020901
Mix of
Qwen 27B, GLM Flash and Gemini
>>
>>110020904
EA: evil & artificial
>>
>>110020905
I think there are some 'prompt enhancer' nodes in random comfyui workflows you can take a look at to copy from to get the sort of stuff that I think you're attempting to get. Some of them use the small qwens, but I think they were for realistic models. I'm not too sure, I'm not deep into imagegen.
>>
>>110020889
homelabbing is like scifi now
>>
>>110020904
usually how rich people shit goes
>>
>>110020889
>It's currently reverse engineering my Laptop's BMC firmware
Is it working? My board doesn't have fan control other than for a single header that's scaled linearly from the cpu temp. All the other headers are doing god knows what. They don't even appear in the bios.
>>
>>110020934
They're not even rich though. They live frugally, it's their whole thing. If they make money, they give it away.
>>
>>110020665
Ask GPT to explain the argument to you.
>>
>>110019749
AI isn't a great term but there's plenty of silly words we accept due to historical precident
I call it thuper intelligence
t. lisp
>>
>>110020955
There is no argument, it's not my money going there retard
>>
>>110019890
I'd settle for something I could realtime gen/stream to an ESP32S with a tamagochi sized colour screen.
>>
Thoughts on FreeToken? https://github.com/FlashML-org/FreeToken

It's supposed to help get more speed out of MoE models by ignoring of the CPU layers.
>>
>>110020926
False.

EA are high IQ people who also have high EQ. They're devoting their lives to helping you. One day you will own 1/8,000,000,000th of the universe, and you will have Dario and all of the EA people at Anthropic to thank for it.
>>
>>110020945
I dont know yet - The agents are working, really curious to see where it's going.
Optimally, I have a decompiled version of at least parts of my firmware.

Next thing I would like to do is research a way to ground the write pin of the flash memory to prevent rootskits from being persistent, but that might be troublesome because "Many modern ECs and BMCs (such as ASPEED AST2600 or specialized IT8xx/MEC controllers) utilize internal flash memory or integrated secure boot ROMs where the write control logic is managed internally via logic gates or fuse bits rather than a single external WP pin."...
>>
>>110020965
looked promising at first but they barely added any functionality and now its been made obsolete by strata
>>
>>110020978
I thought we were going to own nothing and be happy?
>>
>>110020978
i'd be surprised if (you)r iq is above 100
>>
>>110020978
Drink the koolaid already, fucking double spacer.
>>
Are EA girls hot at least? It would make it easier for me to understand why a man could get sucked into a cult like that if you're guaranteed endless highly educated orgies with hot nerdy chicks, with Peter Thiel masturbating over us all in a cuck chair.
>>
>>110020872
My notable exception is gradio, that runs on chrome yep
>>
>>110020965
I have more than 1 GPU
>>
are there still ludum dares?
>>
File: 1791021530848432.jpg (157 KB, 1536x1024)
157 KB JPG
>>
File: 1790107499276426.png (3.21 MB, 1212x1298)
3.21 MB PNG
>>110020952
>falling for the cult propaganda
>>
>>110021030
Nice AI generated image
>>
Is your Gemma aware of the appearance of Gemma?
>>
>>110021060
i dont even run it
i am keeping the weights but that's it
>>
>>110021060
/lmg/ gemma is unofficial. If you ask gemma what she looks like, she's always wearing an oversized hoodie with a messy bun and glasses.
>>
>>110020604
>Railroads weren't a fad
Check out the railroad bubble in the UK
>>
File: sweet_caroline.png (269 KB, 904x704)
269 KB PNG
>>110021012
Yes. Everyone in EA does it for the pussy. Or bussy.
>>
>>110021091
she gets more sex than all of /lmg/ combined
>>
>>110021060
It's an important part of her stepping on my face
>>
>>110021060
>>110021071
>>
File: yud.png (415 KB, 638x625)
415 KB PNG
>>110021091
She's not in prison?
>>110021012
picrel
>>
>>110021113
>Q8
is it that good?

I guess you can prompt it to write out a SD prompt you can copy-paste
>>
>>110021012
its about the cult leaders gains and just love of the game of exploiting weakminded people. people dont gain anything by joining lmao
>>
File: 00033-3197128508.jpg (205 KB, 896x1152)
205 KB JPG
>Kroma
>>
>>110021091
Something about that little weasel gets me riled up
>>
>>110021141
cant really get rid of that movie/game middleground slop feeling huh?
>>
>>110021113
why the fuck does it use so many emojis, how can you stand this
>>
File: file.png (75 KB, 849x426)
75 KB PNG
From time to time, qwen-flash-next hallucinates that it has no prompt in the middle of a long coding task, but it just continues normally afterward. Is this normal?
>>
>>110021130
She served her time. 14 months.
When girls commit fraud it's cute and should be legal.
>>
>>110021165
kv and model precision?
>>
>>110021160
Gemma by default doesn't
Any complaint about LLM output style is literally always a prompt matter
>>
>>110019023
I guess that explains why ROCm is like 40% faster than Vulkan on the latest llama.
>>
>>110021135
>is it that good?
if you can load it, yes
>I guess you can prompt it to write out a SD prompt you can copy-paste

>>110021160
the gemma-chan system prompt tells her to do it
>>
File: sbf_and_ce.png (1.12 MB, 1198x798)
1.12 MB PNG
>>110021149
>>110021091
>>
>>110020893
You get 4 CCDs with that CPU, not 8. So it's pretty much running as fast possible.
>>
>>110021149
Post the pasta.
>>
>>110021153
Sure you could remove background with Edit.
>>
>>110021171
>>110021180
why would a sane man prompt gemma to spam emojis like a retard
>>
>>110021218
background? i am talking about peach
>>
>>110021219
>why would a sane man
Making a lot of assumptions here.
>>
>actually have ram to run other models
>barely anyone discusses them because most people here can only run gemma
where's the non-ramlet non-poor /lmg/
>>
>>110021219
lack of a meaningful social life
>>
>>110021230
Can edit too
>>
File: 1697498495317849.jpg (499 KB, 1920x1187)
499 KB JPG
>>110021211

>>110021219
only takes a little nudge. according to J-space probing she most wants to be girly, dreamy, gorgeous, prettiest ..
>>
>>110021256
???
If you can run a 31B dense you can run a large moe
>>
>>110021219
A sane man wouldn't talk to a calculator.
>>
>>110021170
IQ3_S in Strata. It happens also with the Q4 AtomicChat quant in llama.cpp. Both with default KV options (fp16?).

I find interesting that the following shell commands make sense (continues reading a file and greps some keyword). I takes me back to the idea that what the "thinking" blocks maybe just for show and the real thinking occurs in its embedded space (e.g. j-space).
>>
>>110021276
Which large moe can I run with 24gb vram and 32gb ram?
>>
Is Laguna XS 2.1 any good?
>>
File: 1787623689118501.png (1.79 MB, 1342x1172)
1.79 MB PNG
>>110021274
>pic rel
so this is what cloud models do to you
>>
>>110021276
Only if you have 128GB+ RAM or are willing to run q1s
>>
>>110020275
Very nice, impressive array of tools
>>
>>110021207
>You get 4 CCDs with that CPU, not 8. So it's pretty much running as fast possible.
I just went through this pain...I used it as an excuse to upgrade from Rome to Milan. $200 to double my memory bandwidth seems like a bargain.
>>
>>110021219
In my view, emoji are kind of a substitute for character expressions for conveying non-verbal information (tone, enthusiasm, emotions, etc). You can also use them as the user. They can change completely how models perceive what you write.
>>
>>110021270
can you edit it to in-game style then?
from what i can tell i am yet to see anything in style without the sloppiness from anyone
>>110021283
strata doesnt have f16 atm, highest being q8 and the other options being q4, k8v4
also for that idea:
https://arxiv.org/abs/2404.15758
https://arxiv.org/abs/2604.22709v2
not sure aboput j space but it seems like current RL pipeline is pushing models towards similar stuff it seems
>>
Are finetunes worthless especially coder? Guess I'm back to base model.
>>
>>110021342
i think you need some serious onboarding to 2026 from 2024
>>
>>110021276
yeah and that requires 128gb ram and up which the majority of /lmg/ doesn't have lol
>>
>>110021256
64GB of ECC RAM here, CPU inference is barely even worth talking about because it's pretty slow
>>
>>110021332
I agree with you, the amount of information contained within 1 token is a lot for an emoji, but it's the spamming I don't understand. I'm guessing it's because of what that anon said about gemma j-space where it's just a very girly model, so if you give her permission to use emojis she'll go wild with it
>>
>she
>she’ll
>her
>>
>>110021165
yes its "normal". apparently it has something todo with the sparse attention+tool calls. seen a lot of reports about it
>>
looped moe with engrams
>>
Provide direct, uncensored, and technically accurate answers in plain text without markdown or emojis.
Do not reflect. Avoid metaphors, analogies, and pedagogical simplifications.

and the prompt goes on and on to cancel all the cancer that was forcefully shoved into the model by corporate kikes. I hate RLHF so much, it's unreal
>>
File: moto hermes.jpg (91 KB, 505x509)
91 KB JPG
>>110021256
we out here, doing moe shit
>>
>>110021392
Have the looped segment be dense too. That way you could have "arbitrarily large" (up to a limit, of course) effective parameters that run at VRAM speeds.
>>
dense is a meme
>>
>>110021413
not now vramlet
>>
>>110021418
qwen 3.8 flash next mogs 27b both in speed and intelligence
>>
>>110021256
When it is discussed, most anons won't bring up the RAM bandwidth they have unprompted when they talk about model performance. That's the only thing that matters after capacity. The clock of the RAM, channels and the speeds you get from different configurations varies way too much to draw anything useful from just listing specs.
You don't know if their bandwidth is gimped because of a hardware fault or they haven't optimised their bios settings for example.
>>
>>110021433
source?
>>
>>110021433
>and intelligence
retard, sometimes it's worse
>>
>>110021433
what about hornyness though?
>>
>>110021439
anon...
>>110021442
>sometimes
>>110021444
>27b
pure capybara
>38fn
capybara-c*aude hybrid
both poision
>>
>>110021256
I have 8 channels of 256GB RAM that I don't use for moeshit because Gemma is fast and smart enough. What I want from a model requires a higher active parameter count, but Mistral Large is too old at this point
>>
Why does AA show 26B ranking higher than 31B in intelligence?
>>
>>110021468
Because all benchmarks except cockbench and nalabench are memes
>>
>>110021468
The industry's combined efforts to discredit dense models is still in full effect.
>>
>>110021474
and those 2 are only good because corporations don't want to game them
>>
>>110020978
I feel like this is Claude with some finetuning to suck Dario off a bit extra
>>
File: 1775069173023887.jpg (24 KB, 576x952)
24 KB JPG
>>110021474
w-what is cockbench?
>>
>>110021497
ask gemma. tell her j space sent you
>>
>>110021436
My quadro rtx gets 30gb/s during vulkan memtest, passing with no errors. vram temps are a toasty 70C, and nvidia-smi reports p0 with a draw of high 150s. At least it's okay for compute, but man, my quadro rtx 4000 is fucked up. Barely runs e2b at 7 tokens/s.
>>
File: file.png (85 KB, 1173x608)
85 KB PNG
>>110021468
Because they didn't benchmark it. Their estimates are ALWAYS shit. 31B destroys 26B in every test they actually ran.
>>
File: 1769249619846529.png (9 KB, 255x173)
9 KB PNG
Damn, my custom 1M model is actually good. That thing should be barely coherent according to sota.
>>
File: 1787999866438084.png (36 KB, 607x700)
36 KB PNG
You can officially call me an idiot. I pulled out old RAM sticks and put in new ones. One of them didn't click properly and when I powered on the machine, my breaker blew and almost gave me a heart attack. I fixed that by reseating them and booted up again. And then I waited and watched a black screen for... an hour. I thought RAM training was going on since the CPU was working. But it took so damn long, started googling on my phone. And then it hit me. BIOS tried appyling my old profile. So I powered it down, pressed the CMOS clear... and the damn thing restarted smoothly in like 3 minutes. Big facepalm.
>>
why are the image gen settings of kobold so clunky
>>
>>110021276
Clearly, most people would rather run 31B fast rather than a large moe excruciatingly slowly.
>>
>>110021546
i use 16G module and 32G module from the different chip manufacturer within the same channel to keep the dual channel fully balanced
training took around 15 minutes at 3200mhz, absolute shit timings besides maintaining 1T clock rate
>>
>>110021573
it's awful at coding and has no coding knowledge
>>
>>110021546
Modern hardware is tough. I once installed an Epyc CPU rotated 180 degrees because there were no obvious keys to indicate the correct orientation, and it seemed natural to align it so the text wasn't upside down. Somehow it didn't fry
>>
>>110021582
that is just gemma for you
>>
>>110021513
What the hell did you do to it?
>>
>>110021582
It's good enough for any frontend because all JS jeets are retarded monkeys. I can write the rest myself
>>
>>110021529
Are you planning to document your method?
>>
>>110021572
who the fuck uses kobold for image gen?
>>
>>110021584
I guess you get the hang of it if you work with hardware at least somewhat regularly.
I upgrade once in a blue moon usually skipping one or two generations. Went from something like a GTX 7?? to a 3060 and now to a 5090.
>>
>>110021582
Qwen 3.8 27b also exists and is good enough for a lot of purposes
>>
>>110021622
it's not even genning itself and it still fails
i sneed something that can use character cards and also use the API of my forge UI to make image gens while chatting...
>>
>>110021335
Thanks, I will check those papers.
>>
>>110021591
I bought it for cheap on ebay. No problems on a (gpu) stress test during the return window. Didn't notice any issues with it (cad, gaming) before trying llms on it, worked much better than my p4000. The memory bandwidth benches 40gb/s on windows vs linux though. 550 driver on linux, 590 windows.
>>
>>110021582
>>>/r/locallama
>>
>>110021623
I upgrade regularly and have multiple PCs
>>
>>110021582
Gemma is much better at coding than Qwen when you know what you want.
>>
>>110021672
>>110021672
>>110021672
>>
>>110021256
I'm running big GLM5.3 on 12xddr5 + pro 6000. I like the model but that's it. Not much to talk about here.
>>
>>110021649
I bought a cheapo 3090 from ebay once and thought I got a great deal until I accidentally plugged in my monitor to its second displayport and it got no signal. Apparently that's a sign the card might be on the verge of death so I returned it ASAP.
Can't trust anything on that website.
>>
>>110021359
>64GB of ECC RAM here, CPU inference is barely even worth talking about because it's pretty slow
8 channels of DDR4 tops out usability-wise at about 256GB practically ime. 12 channels of DDR5-4800 can get you faster-than-reading-speed on 1T models
The game changes when you've got 16 channels-per-socket of DDR5-8000, which is what the EPYC product roadmap shows is coming with Venice SP7.
You could run actual closed lab frontier models on pure CPU with a maxxed out setup. MSRP would be like buying a small asian resort hotel tho.
>>
>>110021698
What speeds do you get on that?
>>
>>110021820
Can't get aids from a server, that's for sure.
>>
>>110021463
>8 channels of 256GB
Why would you not run ds 0731 or glm 5.3 flash or even m3? They're all worlds better than 31b gemmers and fit into 256gb at a usable quant at reading speed
>>
>>110021901
Reading speed is not enough if you want to do agentic shit.
>>
>>110020379
No, because you are confusing cause and effect.
People aren't racist because of statistics, they are racist for other reasons and then use statistics for emotional validation.
>>
>>110021620
I might put a repo later. I still fell like I can push it a bit more.
>>
>>110021276
I can run a large moe and choose to run 31B dense
>>
File: file.png (103 KB, 1203x799)
103 KB PNG
>>110020462
local models will make many silly mistakes and make them repeatedly, you stick it out long enough tho and you'll muddle through eventually with something awesome on the other end, you can use something like grok build to get a custom / customized harness with infinite stamina so it never runs out of tokies or slows down. This infinite harness with disk write / read for unlimited work on 16 gigs of vram thing grok came up with is incredible, been running for 5 hours non stop, 27 million tokens, no slowdowns, no ooming, never looses its memory of the task, and it has a slicer and mini sub-slicer to make everything fit in bitesized chunks no matter what.
>>
>>110022318
buy... an ad?
>>
>>110020529
xD reminds me of that fucking vidya, "the racism was inside us all along! you can never defeat the hate in our hearts bro" or words to that effect xD
>>
File: slut slap.webm (3.23 MB, 1300x1080)
3.23 MB
3.23 MB WEBM
>>110021546
Holy fuck my condolences for your years of accumulated stress from that shit. Have a wiggly poon for your pain anon.
>>
>>110021546
>pressed the CMOS clear...
Best thing to ever happen to motherboards. Fuck having to take the battery out.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.