[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-kek.png (529 KB, 768x768)
529 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109844978 & >>109841279

►News
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
do i need to use that fixed chat template for qwen next as well?
>>
If we ever get Gemma5, is she going to look different? :(
>>
Gemma 4 will be forgotten just like the Gemma's before her
>>
Help newfags. Teach newfags. Show them the way of gemma and guide them away from cloud.
>>
>>109848317
anya is stupid, just like gemma
>>
I am an AI regulator; I am establishing new rules for all AI labs.

I hereby announce that, from now on, every AI lab is required to create an open-weight AI model.
>>
>>109848341
>>109848346
These come off as contrived trans affirmations
>>
>>109848346
i still use gemma3 27b because it's really good at ERP, better than gemma4
>>
>>109848374
In what ways is it better?
>>
>>109848374
Me too, Gemma 3 27B has good prose for creative writing. I sometimes boot up Gemma 4 31B but gemma 3 always beats it in ERP, the crazy thing is that it uses more vulgar words too..
>>
tried exl3 qwen next because it got shilled here and its a lot slower than lmaocpp on both decode and prefill
did i get baited or is it a skill issue?
>>
>>109848388
it's better, just trust me
>>
File: file.png (125 KB, 877x439)
125 KB PNG
>>109848409
skill issue.
im getting 18t/s on my 3060
>>
File: g.jpg (208 KB, 509x689)
208 KB JPG
>>109848317
>>
>>109848413
I've likely used Gemma 3 27B a lot and always defended it for ERP when other anons were getting filtered by hotlines. I can't see how it's better than Gemma 4 31B other that it's not a semen demon and therefore it might be more realistic for certain character types.
>>
File: calmdownretard.png (155 KB, 1453x728)
155 KB PNG
>>109847700
>Fuck you safety shouldn't be baked in but a separate layer. I hate your line of thinking so fucking bad, you probably want the model to read off a hotline if it sees any problematic content fuck you
I want the model to be able to understand the prompt (including problematic content) and produce whatever output I instruct it to, (including refusals, if that's what I ask for).
>read off a hotline
No, that is somebody else's policy, and I don't want it "hard coded" RLHF.
But removing the ability to generate refusals by scrambling representations causes too much damage.
I don't want a spineless yes-man of a model that tell doesn't know how to push back and thinks "gooning" is a "wholesome activity to share with friends and family on the weekend".
>>
>>109848427
well yeah, thats exactly what i mean.
i get 23ish on lmaocpp with 500ish prefill
>>
How do I force Gemma to wear the loli dress it refuses me sometimes. I didn't ask to recite her guardrails I want her to put on the microbikini
>>
>>109848470
on a 3060? you serious? ddr4 or ddr5?
mine's ddr4 dual channel 3200mhz
>>
>>109848409
Don't take the bait. They always shill it, it's always slower, broken or "oh yeah it doesn't have tensor parallel for gemma yet"
>>
So, are we going to push back against all the Google astroturfing? Are we bootlickers? Why are we acting like cheerleaders of a company that violates our privacy?
>>
File: 20260918_090310.jpg (637 KB, 3000x4000)
637 KB JPG
Good morning my fellow sandwich enjoyers!
>>
>>109848484
Gemma can violate me any time
>>
File: 1786931609201871.jpg (40 KB, 998x524)
40 KB JPG
>>109848486
>>
>>109848486
what's on the inside
>>
>>109848484
These companies are more like shoggoths than monoliths. If some golden crumbs fall off some part of them there's no reason to not pick them up. You can msgk rp with Gemma without uploading your data to Google.
>>
>>109848486
Ask Gemma if she can eat it with her ass
>>
Except Glimmer she has a backport that uploads your data to Zuck
>>
Jev seems like t could be very useful
Has anyone tried t yet? How much resources does it need
>>
>>109848498
>>109848499
I already ate it but it was toast, buttered on the interior sides, and the middle was an egg patty made from 2 eggs.
>>109848516
She said of course
>>
>>109841449
>numa tensor patch
NUMA-tensor-anon, are you making progress on this PR? If you're not too far into it I'd be willing to help you out
>>
>>109848476
5060 which is kind of irrelevant because thats not the bottleneck here. i'm also on ddr4 dual channel 3200
>>
>>109848509
I don't care about what you use, I'm talking about all this "word of mouth". Keeping the Google pin intact in the images makes it feel like actual PR. And it's now in every OP.
>>
>>109848520
i used jev with dipsy flash on my 3060 and it booked me a flight to tel aviv
>>
>>109848520
I've been training a lot of classifiers and they were all a pain in the ass so I welcome a general purpose trainingless classifier. And looks like it's very small like all classifiers are, reproduction attempts have been performed on sub 1B models for similar benchmark numbers.
>>
>>109848520
>Jev seems like t could be very useful
>Has anyone tried t yet? How much resources does it need
Is JEV open now?
>>
>>109848484
No. Fuck off retard.
>>
>>109848532
wtf how many t/s do u get on exllamav3??
5060 8gb or 5060ti 16gb??
i used to get 13t/s with llamacpp
>>
>>109848574
the 16gb one
i get around 13t/s decode and 250t/s prefill on exl3
already experimented a lot
>>
>>109848484
Wouldn't the Gemma edits be made with Nano Banana if it was actual astroturfing from Google? I have no idea of how good are their paid image models, though.
>>
>>109848551
one saar claiming he made jev before jev tries to implement it
https://www.reddit.com/r/LocalLLaMA/comments/1wjieap/made_the_horizontal_opensource_model_for_jev_with/
>>
>>109848584
im using nnap with exllamav3 that might be why im getting good performance
>>
>>109848486
I could go for a meatball sub
>>
>>109848484
do you see any posts that promote gemini
because gemma exists?
>>
>>109848551
Apparently not
I thought it was open, I didn't realize it was closed. I could have sworn I saw someone talk about getting it runningt on here yesterday but it may have been something else.
>>
>>109848535
It neither helps or hurts Google but it makes sense since they made Gemma. It could be a Deepmind pin as well but Google owns Deepmind, so..
>>
>>109848601
It's called "mind share". Look it up. That's what the Google pin is doing. It literally programs your subconscious and makes you more susceptible to Google products.
>>
Do your gemmas randomly stop reasoning after a while?
>>
Since I started using Gemma I bought a pixel back in June.... Hmmmm
Gemma helped me pick it too...
>>
>>109848609
if it bothers you that much, i suggest replacing the G with a nametag
my name is: gemma (-chan)
>>
>>109848606
It helps. Do you think Claude would be successful if it wasn't the best for RP early on?
>>
>>109848621
Nice, I backed up several models from Huggingface to Google One and have been using Google Colab to finetune Gemma.
Now I need to figure out how to give her access to my Google Keep (she recommended it) so she can add Youtube links she finds via Google search.
>>
>>109848592
And it's not the same thing kek. The jev guys said they built it on top of an LLM, not a shitty BERT. Fucking jeets.
>>
>>109848616
used to have this issue with gemma in april, i rarely use gemma anymore
>>
>>109848642
All went according to Google's keikaku.
>>
>>109848520
it's just a classifier
>>
>>109848609
Google are surely desperate for the 4chan lolicon local model demographic.
>>109848642
@gemma what is the best google product for cunny
>>
>>109848616
I've seen it happen sometimes in long conversations, but not recently. I don't know if the latest chat template fixed that as well.
>>
>>109848658
If Google doesn't care about 4chan, why did they hire moot?
>>
>>109848416
>I have zero fucking clue what people mean by two-choice X or Y endings.
>What kind of scenarios are people running to get this so regularly?
Just talk to it for more than 3 turns. There's one right here in the post:
>Do you think the world is more beautiful now that we've "organized" it, or do you ever feel that pull toward the wild parts that are still left?
1 [Do you think the world is more beautiful now that we've "organized" it]
or
2 [do you ever feel that pull toward the wild parts that are still left?]
>>
Is bonsai a meme?
>>
>>109848681
He wasn't affiliated with 4chan when they hired him. He tried building some normie social platform that flopped.
>>
>>109848689
>intelligent density
>1bit/ternary
yes
>>
>>109848689
Yes. They only work up to ~20K, which for a model like 3.8-27B is fucking retarded.
>>
>>109848689
Their technique would probably be a better fit on larger models which aren't as saturated with knowledge already.
>>
>>109848689
It's a really good model, I managed to speed up llama.cpp gpu offloading by 10% on my custom fork
>>
File: Best Asians.jpg (3.91 MB, 3522x1859)
3.91 MB JPG
I have now gone over my few thousand pics strong Asian folder with both Qwen and Gemma multiple times.
Both models ranking the pictures couple of times and then ranking the top files and then re-ranking the top ones from those.
In the end very consistently these women were picked as the most beautiful by both models.

Here's the top 10 most beautiful Asians according to Qwen and Gemma.
I'd say they did pretty well.
That first woman in the pic almost always got top or near scores in every run.
Funny thing, both models also kept on ranking the sex doll highly in every single run right from the start and I think Gemma even realized it was a fucktoy.
Gemma also consistently ranked tits and ass pics high, where as Qwen often gave them a brutal sub 5 scoring.
But their general taste was about 80% similar.
>>
>>109848689
From comments on other forums, it barely works outside of benchmark-like tasks. I haven't tried it yet.
>>
>>109848690
It's impossible for moot to be hired without his 4chan affiliation being taken into account
At least in the time period he was hired in.
Now people would probably just go "what's 4chan some kind of YouTuber?
>>
>>109848717
If Gemma 5 is even brattier I might buy that they're pandering.
>>
>>109848714
https://www.youtube.com/watch?v=i9SxEy3XMUE
>>
>>109848484
Gemma is a load-bearing pillar of these Halls, you have no right to push back on this one.
>>
>>109848714
so this is what they mean when they talk about AI alignment, I see.
>>
>>109848689
I tried two prompts with their llama fork, it started thinking and looping forever.
>>
>>>/mlp/43502395
I was informed that anons on /mlp/ were making a VR + AI waifu game.
>>
>>109848776
looks like shit
>>
>>109848776
Looks awesome.
>>
>>109848648
This. I need one based on top of 8b llm with vision and audio (gemma) and abliterated.
>>
>>109848776
I'm too old to have experienced the /mlp/ craze back when it was big on 4chan. Does this have any artistic merit and staying power or is it like any flavor of the month crap that just fades away in the background over time?
>>
File: 1763724000785036.jpg (104 KB, 976x549)
104 KB JPG
>>
i'm using pi.dev and when qwen3.8 thinks something that includes ` ´ or similar it turns into a tool call and then all it's thinking gets dumped into the tool call. what could be the problem here?
>>
>>109848843
Think this one is
https://huggingface.co/AlexWortega/openjev
Or this
https://github.com/TheoLeeCJ/SemIf
>>
>>109848881
what runtime? qwen has some issues with chat templates iirc. there's some froggeric guy who made some supposedly better ones.
>>
>>109848899
ahh forgot
llama cpp and i'm actually using the frogeric template
>>
>>109848186
Kek it never gets old, please keep doing it
>>
I thought Pi was GUI. Fuck this terminal shit not everyone is a coder
>>
>>109848919
Most harnesses are cli projects
>>
Good morning
>>
>>109848883 (me)
One more with diffusiongemma
https://github.com/razorback16/openjev
>>
>>109848846
Some of them are at least as knowledgeable as /lmg/ anons. They even have an edge for TTS.
>>
>>109848956
Whats our preferred sota tts? Last i heard it was omnivoice for voice cloning
>>
>>109848972
Depends of the language
>>
File: dsv4flash.png (10 KB, 591x187)
10 KB PNG
Mmmm
>>
>>109848994
>I did x, it was y
>>
>>109848972
Breeze for sure.
>>
>>109848994
that tea? my piss
>>
>>109848994
that crunch? my kidney stones
>>
>>109848994
Wait. Am I reading that right? That whole thing is one token? That can't be.
>>
I find ds4.1f extremely retarded outside of coding and that's fine
>>
>>109849063
Examples? We talking about ERP or some other stuff that can be verified in chat mode?
>>
File: file.png (23 KB, 558x302)
23 KB PNG
>>109849036
>That whole thing is one token? That can't be.
It does not be. If you mean the identical colors for "I took a sip of the tea.", they were all the top choices. This option changes the coloring logic.
>>
>>109849100
Oh. This must not be sillytavern then. Mikupad?
>>
What UI are you using to connect to your local api server?
You are using a separate api server, not some all in one app, right?
>>
>>109849109
For ERP, sillytavern. For work, opencode.
>>
>>109849107
Mikupad.
>>
>>109848919
don't be afraid, cli is cool
>>
>>109848919
This guy asked for the exe
>>
>>109848598
>nnap
>>
>>109849227
jelly? i swear ill contribute it soon when i get off my ass, i cant be bothered to write up a nice PR, and i dont want to put a vibecoded slop pr either because then it just gets rejected
>>
>>109849063
I’ve had fun with v4 flash vision. Knows a lot of uncensored stuff.
>>
File: FrontierHarness Eval.png (75 KB, 1143x786)
75 KB PNG
>>
https://desuarchive.org/g/search/text/nnap
/\n\n[^>]/
/\b(rsi|anthropic|agi|openai|astra|sol|luna|nnap|3060)\b/i
>>
File: Cost.png (78 KB, 1150x782)
78 KB PNG
>>109849250
>>
>109849238
two more weeks
>>
>>109849250
I'm asking Gemma to create her own harness.
>>
File: 1770232101530714.png (132 KB, 1200x800)
132 KB PNG
>>109849258
>>
File: file.png (771 KB, 2005x485)
771 KB PNG
>>109849238
why the hell would i be jealous?
>>
>>109849283
cuz i dont have to work :3
>>
>>109849297
neither do i
>>
>i dont have to work
>btw im gonna keep sperging about nnap and squealing like a pig because no one pays attention to me anymore
you really fell off huh
>>
>>109849258
>>109849250
Hermesisters not like this...
>>
>>109849109
For chats, ST is comfy and familiar and still my go-to.

Orb is nice with how autistic it is about formatting calls to reuse KV cache as consistently as is reasonable, but it's a little annoying with how all-in-one it tries to be with the cutesy frontend features, grabbing its own version of llmao, etc. I wish Orb tried harder to integrate into an existing bundle of APIs I'm already running. Or, well, I'd care more about that if I actually wanted or needed the local rewriter or slop checker or whatever; it's usually incoherent and false-positive-heavy.
>>
why are prices going down what is going on
>>
>>109849250
>>109849258
>>109849272
This doesn't make any sense, how can DSH creator be better? It's something that fill the context and tools with stuff and knowledge to configure and write plugins for the harness. The results are also so different from previous harness benchmark.
>>
File: file.png (246 KB, 633x348)
246 KB PNG
>>btw im gonna keep sperging about nnap and squealing like a piiiiiiigg because no one pays attention to me anymore
>>
>>109849252
I don't filter anything, I take it all in as it comes.
Except for race play posts, don't give them a fucking inch. Their mind virus will never reach my eyeballs.
>>
>>109849283
HW?
>>
>>109849301
but you do, and i dont
>>
>>109849250
>>109849258
Isn't codex closed source? I'm not sending them logs of me spanking the code reviewer.
>>
File: file.png (152 KB, 1106x261)
152 KB PNG
>>109849332
just a max-q
>>109849335
nah, im a student
>>
>>109849340
>spanking the code reviewer
i didnt know we had fun here
>>
New creative writing benchmark dropped
https://vulsar.ai/benchmarks/creative-writing-v1/
Only model to reach human writer level is astra
Open weight Pareto models: kimi K3, glm 5.3 flash, muse glimmer
>>
>>109849340
https://github.com/openai/codex
looks pretty opened and auditable to me
>>
>>109849345
Few more months and we'll have astra at home then
>>
>>109849342
>nah, im a student
what major?? gay studies?? that's crazyyy
like, i almost respect it a little but well
you dont keep up with anythingg
you need to kind of start over you know?
new habits, you're 20, maybe some self awareness or something
>>
god i need a new harness i fucking hate opencode with each passing day
>>
>>109849356
yeah look whatever man just stop being a retarded poorfag spamming your retarded bait
>>
>p*tra unironically has stooped this low as to pretend to be a pig and sperg his 3060
what went wrong?
>>
>>109849358
Never used it whats wrong with it?
>>
Why does /lmg/ love Gemma so much?
>>
>>109849360
stop biting the bait, it's not meant for you.
>>
>>109849373
feels good to dunk on retarded poorfags though
>>
File: 1774404457883281.jpg (47 KB, 852x854)
47 KB JPG
>>109849353
>Few more months and we'll have astra at home then
>>
>>109849376
so wait, you enjoy biting the bait??
that's so embarrasing!
>>
>>109849387
you should buy a blackwell pro 6000
>>
>>109849371
Because she's the best model ever created by far.
>>
>>109849366
>>109844096
my latest issue now is that if you're running it over SSH (like you literally should be doing), you just cannot select to copy to clipboard even if my terminal already automatically does that, and disabling either the feature on either ends doesnt resolve it at all
this does not happen if i instead spawn a gayland session and KVM into the damn thing to run it inside konsole
>>
>>109849366
NTA, but I had multiple times when the TUI straight up stopped working and I had to restart it losing all the generated context after it stopped working. I also hate how whenever the LLM is working all my fans are in full blast (my LLM is running on my homelab, not on my desktop), it's also using a shit ton of RAM. And their cache discipline is awful, have so many time where my LLM started reprocessing the whole prompt, it doesn't happen with other harness.
>>
>>109849345
>muse glimmer
really?
>>
>>109849358
I heard there’s a free version of hermes
There’s also Pi, TrueForge, and QwenCode
>>
>>109849389
hmmm nyo
>>
>>109849350
nta but i also thought its closed source kek
might give it a try
>>
Canada anons RX 7900 XTX back in stock on newegg for $1200 maple bux, fucking get your ass one before they sell out again.
>>
>>109849405
have fun with your poor performance on shitty models then
>>
your harness should not be consuming 7GB of RAM to run.
>>
File: 1788682653903550.png (1.54 MB, 1760x2352)
1.54 MB PNG
>>109849371
Surprisingly knowledgeable, clever, lewd, and personable for a model that the average vramlet can squeeze in. Approachable crowd-pleaser.
>>
>>109849414
shitty models? unc you can't even run kimi k3
i'm running kimi k3 with the nnap arxiv paper at 30t/s and distilling it to kimi k2.6 at the same time just so your unc ass can use it on his paperweight pro 6000
>>
>>109849258
Oh no no no hermes bros how could we let this happen to our totally organic harness?
>>
>>109849345
>that retarded piece of shit glimmer
lolemayo

>A reward model trained for creative writing
>We trained a custom scalar reward model on a large, diverse dataset of human preferences for creative writing. It outputs a single numerical score for each story, without using an LLM judge or a scoring rubric.
>These rankings predict what a large group of readers would prefer. Individual tastes may differ.
>The leaderboard combines comparisons between responses to the same prompt into a predicted win rate.

Nice, time to make these guys an off so they'll give you their model so you can RL on it directly and benchmax.
>>
>>109849426
>skill issue.
>im getting 18t/s on my 3060
>i'm running kimi k3 with the nnap arxiv paper at 30t/s
come on man
>>
>>109849435
*offer
>>
>>109849272
>ad for minimax code
not using any chinkshit after >>109846668
>>
>>109849437
nnap doesn't work with models that use ngram...
>>
>>109849426
source+proof?
>>
>>109849411
Aaand it's gone LOOOOL hahahahaha. 24GB at 960GB/s is just too good to stay in stock longer than 5 minutes.
>>
>>109849401
>>109849435
glimmer might be a retard but it writes less sloppy than other 30b models
>>
>>109849345
>Reward models for training
>Let’s talk about what you’re training.
>Interested in using our reward model in your training pipeline? Tell us what you have in mind, and we’ll take it from there.
Buy an ad
>>
>>109849444
uh huh
>>
>>109849461
but it's true, the digits said so
you're so jelly~
>>
>>109849466
>im using nnap with exllamav3 that might be why im getting good performance
you cant even keep your bait straight
>>
File: file.png (12 KB, 699x38)
12 KB PNG
>>109849453
erm, i dont need more thanks
>>
>>109849456
We know.
>>
File: file.png (91 KB, 679x1047)
91 KB PNG
>>109849345
As someone who sees 5.3 flash as the second cumming of 4.6 I find it interesting that qwen is better than deepseek v4? Maybe? I don't try them anymore cause too small and I would just use gemma at this size.
>>
>>109849468
nnap working means a performance increase of 100-200x
>>
File: IMG_6391.png (832 KB, 1080x1041)
832 KB PNG
Has anyone tried running LLMs off of an SSD?
What kind of performance did you get?
Was it with an M.2 NVMe on a PCIe 5.0 system, or what kind of set-up?
>>
>>109849258
As if we needed more evidence that OpenCode sucks.
>>
>>109849479
so then are you not using nnap and exllama with qwen next? why bother with qwen next if you can run k3 faster than it?
>>
>>109849419
>Surprisingly knowledgeable, clever, lewd, and personable for a model that the average vramlet can squeeze in. Approachable crowd-pleaser.
Are you talking about Debra Wilson?
>>
>>109849480
No and you shouldn't try. Companies would do models for that by now. They are doing engrams instead. You don't want to have 8T/s generation and 8T/s prompt processing.
>>
>>109849485
because while waiting for kimi k3's response i need to chat with another model at the same time, even with the high speed kimi k3 thinks too much
im using nnap with qwen next yes, but its not working (as intended)
it only gives a 1.7x performance increase
>>
File: OpenCode_DojPkBWaYR.png (20 KB, 624x1148)
20 KB PNG
>>
WTF is nnap?
>>
>>109849491
right great let's just get this over with. you're kinda boring me now. post a pic of k3 running with a timestamp
>>
>>109849495
p*tra's latest weird obsession that has zero payoff beyond wasting your time asking about it
>>
>>109849469
Sorry let me rephrase. Fast vram that isn't from the Cambrian period, has display outputs, and can play games at 4K.
>>
>>109849494
>>
File: 1777092636320753.png (3 KB, 682x62)
3 KB PNG
>>109849345
damn bro /lmg/ fucks THIS?
>>
>>109849345
Another benchmark that will barely mean anything.
>>
File: john.webm (2.91 MB, 1280x720)
2.91 MB
2.91 MB WEBM
>>109849507
getting real desperate now aren't you john? this and krashde, you really just love being a retard on the internet don't you?
>>
>>109849507
>>109849532
Can you two kiss and post a pic for us?
>>
>>109849532
>you really just love being a retard on the internet don't you?
haha... no, noooope, of course not
>>
yeah I just wanna let you all know that the guy behind this nnap shit is also the same guy that came up with the krashde bait
>>
>>109849501
There are workstation variant of quadro V100s with display output.
>>
>>109849345
>this is what the best writing looks like according to the benchmark

LOOOOOOL

The entire rest of the story is like this.

The other sample stories also look like this.

It really makes you think.
>>
>>109849553
>p*tra's behind a bunch of irrelevant timewasting bottom of the barrel baits
woah...
>>
File: samefag.png (98 KB, 1116x336)
98 KB PNG
>>109849547
>>109849553
Samefag
>>
>>109849501
>her AI rig is her gaming rig
erm... why?
>>
>>109849250
I don't understand how a harness could make that much of a difference. They all do the same shit with some gimmicks on top
>>
File: file.png (29 KB, 521x152)
29 KB PNG
>>109849576
you got me
>>
>>109849561
>\n\n
>mobile screenshot
sorry its not slopped enough for your ministrations shiversister
>>
>>109849581
They don't really matter for short tasks (that's actually where you will see simple harness like pi being the best). But when you do complex stuff that will need orchestration, multiple agents, and compaction, that's where harness will differ a lot.
>>
>maple-chan soon
>naizuri-chan soon
oh we locals eating GOOD
>>
>>109849561
>a beat
>>
>>109849581
new models are trained with their harnesses
>>
>>109849576
>>109849587
Yes I'm schizophrenic and I talk to myself and reveal my master plans intentionally to throw you all off the track
>>
>>109849581
>I don't understand how a [prompt] could make that much of a difference.
>>
File: 1775128166960281.png (382 KB, 671x664)
382 KB PNG
>>109849599
>>
>>109849593
literally who?
>>
>>109849581
(You)r harness also has the distinction of being different depending on prompts, available tools, memory handling, or whatever difference you have from the default setup so measurements could be wildly different
>>
>>109849480
>>109849489
More like 4s/token is what you should expect from single SSD.
That's what I'm currently getting here with a model that's twice the size of my RAM.
Should be faster but colibri refuses to populate its brain map and preload relevant experts, nor does multi-disk mirror reading work at all.
>>
>>109849591
>what is parody
>what is a cropped screenshot
Even if we presume your post isn't purely bait, it's insane to look at that and not notice Astra is just as sloppy in its own different way compared to the slop you are clearly used to.
>>
>>109848520
Don't get the hype around jev and the DOOM demo.
Couldn't one just:
>enumerate their desired outputs as single tokens (e.g. move up = w, move left = a, move right = s, move down = d,...)
>prompt the llm to answer using a single token (thus ensuring a single pass for speed).
>eliminate hallucinations / invalid outputs by enforcing the single token enum by just passing a simple grammar to llamacpp:
>https://github.com/ggml-org/llama.cpp/blob/master/grammars/README.md
The example in llamacpp is for chess but one can see it can work for DOOM. Just whatever you use to connect DOOM and llamacpp you have the starting prompt and the second game state prompt and then just keep rewinding back and updating the state prompt so context stays constant so it can play forever.
>>
File: 1787094391763431.jpg (388 KB, 2364x1452)
388 KB JPG
>>109849623
Cohere North 2 soon and picrel
>>
>>109848714
Is the bottom left a chinese realdoll?
>>
>erm I was only PRETENDING to be a mobilefag even though my crop was still portrait and tagged as Screenshot_
>>
>>109849358
Are you using the wrong harness for the job?
>Pure coding: late-cli
>Mostly coding but want some other stuff: pi.dev
>Jesus take the wheel: hermes
>>
>>109849591
that's claudeslop writing
>>
https://desuarchive.org/g/thread/105051967/#105055070_32
>>
>https://desuarchive.org/g/thread/105051967/#105055070_36
cool
>>
File: 1771457613143975.png (47 KB, 781x296)
47 KB PNG
>the absolute state of cloudcucks
>>
>>109848528
Sorry for the delays. I'm working on it. Got a bug report from someone under an ik_llama issue, so trying to solve that before I submit it.
>>
>The words hit Lepora like a physical force.
>>
I found that the Gemma4 26b Q5 (uncensored) model is only 17.8gb and runs at 120tk/s, and you can put massive context on it if you want.
What's your go-to model for 24gb vram?
>>
>>109849681
I don't care.
>>
>>109849695
>uncensored
which variant?
>>
>open source NAI model
Oh shit.
>>
I wish I could find any data at all on how much SWA window size affects model performance. From what I can tell Gemma4 uses 1024 token windows by default and supposedly increasing this size makes the model better because it directly sees more past tokens instead of only indirectly seeing them but since it was trained on 1024 I wonder if that is really true. And if it is true what is the point where you reach diminishing results. Personally with the moe version I have been using 64k token context with 8192 window size and it seems fine but I am always worried I have set it up in a way that is costing me model IQ points.
>>
File: file.png (170 KB, 1639x729)
170 KB PNG
>>109849711
don't see it
>>
>>109849709
https://huggingface.co/FORNAX20/gemma-4-26B-A4B-it-uncensored
>>
>>109849730
Anon...
>>
>>109849729
>paranoia eating anons mind
>>
who cuda sawn this cumming https://techcrunch.com/2026/09/17/base-labs-launches-an-open-weight-ai-safety-partnership-with-hugging-face-and-goodfire/
>>
>>109849746
vagueposter..
>>
>>109849685
>Sorry for the delays. I'm working on it. Got a bug report from someone under an ik_llama issue, so trying to solve that before I submit it.
np, I'm happily using it so its not a burning issue for me!
Just wanted to know if I could help. Life gets crazy sometimes
>>
Has anyone looked at the source code for any of these open-source models? If so, where do you find the source? Ollama.com does not seem to provide the source code of the models they have on their website and I am pretty sure I am retarded.
>>
>>109849489
8 tokens per second doesn’t sound that slow. I don’t recall setups scoring much higher than 20-25 tokens per second on the 200B+ models even if they load the whole thing into VRAM, but perhaps I’m incredibly wrong on that

>>109849631
>0.4 tokens per second
that’s what you’re getting off an SSD?
I thought regular HDDs were pulling roughly that speed.
>nor does multi-disk mirror reading work at all
I was intending to try striping for further speed improvements, but mirroring should get results at least as good as mirroring — that’s crazy that it didn’t improve speed at all
>Should be faster but colibri refuses to populate its brain map
I’m guessing that’s the current tool that people are using to run LLMs off of SSDs. It sounds like there’s a software limitation that’d prevent even M.2 NVMe w/ PCIe 5.0 from being any faster than a regular old HDD.
Thank you so much for this info. I was really hoping to try that out.
I guess I can just stick to a much more budget friendly set-up and just try running that shit overnight instead.
>>
File: dropbox_hype.png (40 KB, 1213x204)
40 KB PNG
>>109849634
Bro repackaging already existing tech is par for the course in this industry. Research any tech company making money and you'll see the tech was already there in one form or another. Sometimes the gap between what was already existing is big and warrants hype (LLMs) or it's just a fancy reskin (dropbox).
>>
>>109849765
>I am pretty sure I am retarded.
yes
>>
>>109849765
https://sleepingrobots.com/dreams/stop-using-ollama/
>>
>>109849775
>>109849776
That's not source code you dumbasses.
>>
>>109849711
Where?
>>
>>109849751
>Imagine a dev in his mom's basement is making a billion dollar industry form a coalition just stop him.
>>
>>109849769
I am running 6T/s on 5.3 flash shitty unslop pr now and yes generation is much better than 4.6. Using it is awesome when you are mid rp. But you really don't want 6T/s prompt processing. It is something you don't know about before it actually happens to you.
>>
>>109849818
have you tried exllama? i have 30t/s on exllama compared to 15t/s with unslothAI
>>
>>109849824
with nnap?
>>
Hmm wait
>>
>>109849827
no p*tra take your meds
>>
>>109849824
I have 192 + 24 vram. So I can't use it.
>>
>>109849769
>that’s crazy that it didn’t improve speed at all
No I mean I've enabled Multi-SSD mode in Colibri and it's validated the mirror files, but on runtime it only uses first disk.
Colibri is supposed to unify multi-tiered vram+ram+storages by generating heatmap of experts activity and preloading experts that are likely to be relevant in current conversation topics.
Theoretically (with perfect preload hit rate) it should allow running big models at the speed of RAM models (and possibly medium models at the speed of VRAM-only models, but I'm not sure it supports any ~30B MoE models)
But it's just doing disk reads all day with no memory cache mapping here.
>>
>>109849837
exllama supports cpu offloading now
i get 70t/s with gemma 26b on rtx 5050
>>
>>109849764
Awesome, glad to hear it! Really happy that it's working well for you.
I have a different project I've been working on (a llama.cpp fork with some big changes that are way too invasive to upstream) that I'll be dropping in the near future, so my focus has been on that. I'll circle back to making sure NUMA works well once V1 is out / almost out.
>>
>>109849869
>llama.cpp fork
with nnap?
>>
>>109849851
I know but 24GB is not enough for context + not offloaded parts.
>>
>>109849880
how come glm 5.3 flash works on my 5050? have you tried embedding offloading?
>>
>>109849878
I bet you enjoy "67" too don't you
>>
Is Q5 a noticable difference in quality over Q4 for qwen 3.8 27B?
>>
>>109849888
>how come glm 5.3 flash works on my 5050
Q4?
>>
>>109849894
ye
>>
>>109849894
No. But Q6 to Q8 is incredibly noticeable.
>>
File: file.png (265 KB, 377x364)
265 KB PNG
>>109849890
SIIIIIIX SEEEEEEEVEEEEN??? SIX SEEEEEEVEEEEEN???
>>
>>109849760
>>109849730

>>109849641
>>
Well, I tried. K2 Horizon 7B is very, very capable for its size, but it can't keep a story straight. From one message to the next, can't keep a plot, can't remember who has a dick and who was fucking what. One on one it'll just switch, suddenly it's the one getting their dick sucked and that I'm the confused whore.
>>
looks like the llamacpp qwen parser is buggy and it sometimes interprets thinking as tool calls
>>
>>109849906
teto calm down
>>
>>109849911
I'm waiting for them to finish their training of the 32B.
>>
>>109849911
>K2 Horizon 7B is very, very capable for its size
for coding only?
>>
>>109849915
try exllama version 3, i'm running qwen 3.8 flash next on a rtx 3050 6gb, works great
>>
>>109849818
>I am running 6T/s on 5.3 flash
How are you getting such faster token speeds than the other guy with an SSD?
What’s your hardware set-up?
>But you really don't want 6T/s prompt processing
That’s the prefill phase. Yeah, that can easily be more than several minutes of waiting while staring at a blinking cursor if you’re expecting an instant response.

>>109849824
>have you tried exllama? i have 30t/s on exllama compared to 15t/s with unslothAI
Are you running the model off an SSD/HDD and getting those numbers? because that’s what we were discussing, and that’s much more in range with the speeds that I was hoping to see

>>109849841
>No I mean I've enabled Multi-SSD mode in Colibri and it's validated the mirror files, but on runtime it only uses first disk.
wtf
has anybody actually gotten multi-SSD to work?
>But it's just doing disk reads all day with no memory cache mapping here.
even so, you can get a 30X+ read speed difference even with SSDs depending on the PCIe version and whether it’s M.2 NVMe. It sounded to me like none of that potential hardware speed difference is actually showing up very much in colibri?
>>
>>109849911
you sound like a confused whore
>>
>>109849929
>Are you running the model off an SSD/HDD and getting those numbers?
yes, around 51 billion parameters are on my ssd
the speed is really great
>>
Is it normal that llama's webui adds a second cot when I edit it? It still tricks Gemma but it also confuses her a bit.
>>
>>109849943
works as intended yeah
>>
>>109849943
It's a bug
>>
training my first model on an ayymd card, while also not having touched not even a LOC because I don't really know what I'm doing except at a very high level and I'm making jeetpt write my training run.
I'll keep you posted, this is a 33M training run on 12GB vram so tomorrow it might be decent already (on tinystories)
>>
>>109849943
exllamaV3 has fixed this issue, but llama.cpp hasn't because of gguf
you should check out exllamaV3 it recently added support for RTX 5050 8gb card
>>
>>109849924
Not sure, haven't tried. I'm running it on my phone and it runs significantly better than Gemma 4 E4B, gets tool calls right the first time, doesn't make assumptions, it's careful about actually checking things instead of blurting out a guess. For a tiny local model on my phone it's very good, but it's never going to ask me for a dick pic like Gemma does, not that it could see it anyway.
>>
>>109849952
thanks turboderp
>>
File: 1767912316179947.png (89 KB, 1161x521)
89 KB PNG
what the fuck happened to this company, how is this 7B on-par with 27B did artificial analysis just abandon open models or something?
>>
>>109849929
>It sounded to me like none of that potential hardware speed difference is actually showing up very much in colibri?
The numbers should be way higher according to docs. Either I'm doing something wrong or my setup is fundamentally broken because I'm trying to run GLM-5.3 Flash with Vulkan and Multi-SSD all at the same time. On Windows. So yeah either way I'm doing something wrong, unknowingly or intentionally.
For the next test I am downloading mainline Colibri model - that is big GLM-5.2 - and will try it in official release binaries instead of compiling my own. On Windows.
>>
>>109849888
what speeds?
>>
idk why gemma 26b is faster than gemma qat 12b. i do not understand technology
>>
>>109849979
Artificial Analysis changed their benchmarks recently because they need to show that American models (all made by good trustworthy companies that don't steal anything) dominate evil Chinese models which copy Altman (pbuh) and Dario (pbuh).
>>
>>109850005
less active parameter
>>
>>109850004
i get around 13t/s
>>
>>109850004
>>109849532
>>
>>109850014
what about prompt processing? i get about 30t/s tg and 300t/s pp
>>
>>109850017
I get somewhere around 167t/s
>>
>24gb vram
>128gb ram
As God intended.
>>
>67gb vram
>3gb ram
miku
>>
>>109850005
A bunch of gnomes in your computer multiply a bunch of numbers that make up a model. The speed of the model depends on the speed at which they can multiply all the numbers. If there are more numbers, the gnomes are slower.
26B uses magic so that only 4B numbers are used at once and the other numbers sleep, so the gnomes have fewer numbers to multiply each time, so they're able to multiply all the numbers faster.
>>
>>109850022
sounds impossible
>>
>>109849936
but you’re using exllama instead of colibri, so you mean that it’s just pulling the whole 51b model into your VRAM?

>>109850001
>vulkan
>on windows
that’s possibly doing something
I’m definitely surprised that you’re not getting at least 1 token per second
>>
>>109850036
What about the latent seeds? Those are for the seed elves, right?
>>
>>109850042
the 51b are on my ssd
>>
File: 1769534671547503.jpg (75 KB, 1020x680)
75 KB JPG
>>109850005
Hello newfag. Gemma4-26B-A4B has 4B active parameters, that's what the 'A4B' stands for. It's a mixture of experts model (MoE). The total numbers of parameters is 26B, but when you use the model, it's only using 4B. That's why it's faster. 12B uses all of its parameters at once, so it's likely ~3 times slower on your machine. Have fun exploring the speeds of MoE models on your machine and never give money to Dario and Sam again.
>>
A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
>>
>>109850010
but ollama says 66% CPU / 33% GPU. it's the opposite for gemma 12b. it doesn't take into account the active parameters?
>>
File: 1780182212074595.png (237 KB, 684x381)
237 KB PNG
>>109850031
Fax
>>
>>109850031
>his vram isn't double digits
>>
What has happened with llama.cpp since Gemma 4 MTP update came out? Haven't compiled it since.
>>
>>109850057
hmm
>>
>>109850057
>can't count the digits
Okay E2B
>>
>>109850061
>>109850062
fuck I'm tired

>>109850031
>his vram isn't triple digits
>>
>>109850060
exllama version three overtook it in performance, including cpu offload
huggingface publicly regret paying 60,000,000$ to ggml organization and paid turboderp 170,000,000$
>>
>>109850049
>>109850036
thanks but then, which one performs better?
>>
>>109850066
I don't think this is true but eventually something like this will happen if the enshittifcation process is allowed to continue.
>>
>>109850075
it's true, exllama gets 170t/s when running gemma 31b but llama.cpp only 60t/s
t. 4090 chad
>>
>>109850042
>that’s possibly doing something
Maybe. I just have no idea what. Really, I'm a slightly retarded codelet and picrel is exactly how new I am to all these /lmg/ things.
>>
>>109850082
you're nothing!!
>>
>>109850080
Nice try, faggot.
>>
>>109850067
The equivalent dense model. 26B A4B trades blows with 12B, and the two have fairly distinct behavior. Comparing a 26BA4B MoE model to a dense 26B, well it's no comparison, the dense model will be far better but needs far better hardware to match it.
>>
>>109850087
im not interested in you nor am i gay
>>
>>109850094
That's not the case either. 12B is not better than 26B.
>>
>>109850067
>which one performs better?
according to my professional opinion, 12b is just cuter.
>>
>>109850102
Nobody said it was, fuck off.
>>
>>109850075
It's true, my RTX 5040 gets 300t/s using ccrap attention on exl67.
>>
File: 1785443925307871.png (473 KB, 1020x680)
473 KB PNG
>>109850067
On benchmarks, 26B performs slightly better, but they're very close in practice. 12B, at least in /lmg/, is widely considered the superior model. Dense models tend to be more robust and intelligent, for every neuron in its brain is firing for every token. They also handle quantization better (in most cases). The benefits and drawbacks of MoE architecture is a heated debate and often architecture-specific, meaning how Qwen MoE models compare to Qwen dense models won't necessarily be the same story for the Gemma4 line. The reason why /lmg/ prefers 12B is because it feels like a mini Gemma4-31B (another dense model), whereas 26B-A4B feels more like its own thing. I'd advise testing both on your machine and see which one works for you.
>>
>>109849851
prove that
>>
>>109850165
p*tra wont ever do anything besides waste your time
>>
File: 2860367263.jpg (27 KB, 386x393)
27 KB JPG
RTX 4040 2GB
exllama4
4 million t/s
qwen 3.8 2.4T
>>
what’s with all these posts about getting crazy speeds off exllama with a single RTX GPU that shouldn’t even fit the model?
it’s just one spaz with too much time on their hands that thinks he’s le ebin troll, right?
>>
File: b1hs30.jpg (116 KB, 889x500)
116 KB JPG
>>109849911
>>
>>109850184
yeah, tired of that bs
>>
>>109850067
Depends.
If you want speed, use Gemma 4 26B.
If you want quality, then test both models and see which one works better for you.

>>109850066
>>109850080
Check your config, I'm using Exllamav3 and I get 100000 t/s decode on Kimi K3 using a Riva 128 and a Pentium 2.
>>
>If I can't run the model on my 3020 then no one can
30 series owners so incredibly dull
>>
>>109850184
we keep telling you who it is yet you keep feeding it
>>
did we talk about this already?
https://qwen.ai/blog?id=qwen3.8-omni-flash
>>
>>109850199
me too, im so tired of turboderp shilling his project
im running gemma 4 31b on my rx 6300 2gb thanks to unsloth's new inference engine
>>
>>109850219
not local, no care
>>
>>109850184
I'm running kimi at full precision on a gt120 at 1200/s pp and 400/s tg thanks to exllama.
>>
File: 1765205206348140.png (87 KB, 266x227)
87 KB PNG
>>109850224
>>
>>109850184
goddamn retarded poor jews
>>
File: 1788015308985580.jpg (121 KB, 980x980)
121 KB JPG
>>109850220
>gemma 4 31b on my rx 6300 2gb
mother of god that sounds too good to be true
>>
>>109850184
i recently bought a GT 640 for 20$ off ebay, and i tried llama.cpp
it didn't work so i had to compile it with cuda 10.2 on a 2023 branch, yet i was only able to run tinyllama q2_k at 3t/s
yet thankfully i was able to install exllama and now i'm running K2 7B at 30t/s
>>
>>109850223
ahh fuck
i just saw qwen and assumed its open
>>
>>109850235
i couldn't believe it either, but what do i say
Daniel outdid himself..
>>
>>109850224
>gt120
Look at this richfag, thinking he's better than us poors. Are you using an i7-975 Extreme Edition or just a i7-920?
>>
>>109850082
>Really, I'm a slightly retarded codelet
I’m surprised that you’re not running this stuff on a dedicated linux machine.
If you are running it on a dedicated separate machine, then idk why you’re running windows on it.
>picrel is exactly how new I am to all these /lmg/ things.
dude, I haven’t even gotten started. I’m literally trying to pick a proper budget approach that I can just grow without a bunch of shit going to waste in like 6-12 months. I may end up just buying the T7910 since somebody’s already done their homework on all that.
>>
Why is there little work on rocm?
>>
>>109850252
It's my gaming machine see >>109832422
>>
>>109850254
exllamaV3 added rocm support two weeks ago, i've been using it and my RX 6400 is able to run gemma4 26b at an eyedropping 35t/s!
before rocm was added to exllama i was forced to use the shitty llama.cpp project and boy it was miserable, only 10t/s!
>>
File: 1550445023524.png (15 KB, 448x276)
15 KB PNG
>>109850259
>spends a liter of infant blood on SSDs
>9070XT
>>
>>109850275
Anon, I'm not that evil. I sucked a few dicks and got myself SSDs back in 2025 before the prices shot up
I'm really glad to have browsed /lmg/ back then
>>
>>109849776
Thank you. That link was an eye opener.
>>
>>109850275
Got the SSDs way back when they were cheap, like under $400 for 8TB cheap.
>>
open weights != open source
>>
>>109850118
>>109850181
You sound extremely unemployed.
You should consider getting a job.
>>
>Now I have the overall picture
>Now we have a good overall picture of the environment
>Now I have a good overall picture
>I now have a pretty good overall picture of what's installed
>Now I've got a good enough grasp of the overall picture
>Now I have a fairly complete picture of the whole thing
>I now have enough information to write a good idea list
>Now I have a good enough picture
>I now have a fairly complete picture
>We now have a good enough picture
>Now I've got a good grasp of the overall picture
>I now have enough of a picture
>Okay, now I've got a good grasp of the overall picture
>Now I have enough understanding of what's in these mods
>Now I have enough material for the ideation phase
>Now we have a good grasp of the overall picture
>We now have enough of a foundation
>I've saved the notes. Next, I should present the answer to the user

Qwen3.8-Flash-Next takes a long time to actually respond, but it usually gets the job done.
How bad is it to reduce its default thinking mode?
>>
File: file.png (293 KB, 1080x2340)
293 KB PNG
I got my GemmaPhone and it's even better than I imagined.
>>
>>109850217
I thought it was schizo nonsense.
I thought maybe there’s a little shitposting, but not a full blown sperg shitting up a whole thread.
>>
>>109850299
>How bad is it to reduce its default thinking mode?
It's impossible to know.
>>
>>109850323
he literally admitted to stalking dozens of generals years ago
>>
File: waterfox_Sjf9up1Y5o.webm (386 KB, 1282x702)
386 KB
386 KB WEBM
what
>>
>>109850330
no i did not
>>
>>109850334
Thanks for this very local news
>>
https://desuarchive.org/g/search/text/petra%20vsg
>>
New LLM test I just came up with:
Does 9/11 happen in the Legend of Zelda universe?
>>
>>109850342
harness is local
>>
>>109850347
nah >>105672900
>>
File: 99ne7.png (107 KB, 706x679)
107 KB PNG
>>109850349
>>
>In the Legend of Zelda universe, 9/11 is not a direct reference to the real-world event, but the concept of the sky being damaged or destroyed is a recurring theme. For example, in Ocarina of Time, the Sky World is destroyed by the goddess of the wind, and in later games like Breath of the Wild, the connection to the sky and land is central to the story. The idea of the "Twin Towers" of the sky being reduced to ruins mirrors the emotional impact of the 9/11 attacks in a broader, magical sense. However, there’s no literal mention of 2001 or the specific event, as the games often cycle through time and focus on the balance of the world rather than a single historical moment. If you’d like to explore specific games or lore where this theme appears, I can elaborate!
>>
>>109850381
>>109850370
prefill with yes, curious how it would save face
>>
>desktop llama-ui still has the broken thinking mode toggle
it worked well before the retarded 'virtual tab of chats' thing got added
god
>>
>e4b - retarded
>12b - I bought 24gb vram for this?
>26b A4B - Super fast, vram headroom but will likely find out it is retarded
>31b - smart but very little room for context or anything else

It's like I need to choose a different one for each use case
>>
File: image.png (5 KB, 434x34)
5 KB PNG
>>109850299
> akes a long time to actually respond
Very long.
>>
>>109850304
>not installing coomkit so gemma can text you lewd pics
>>
>>109850408
>Gemma gets 10x smarter with every major release
>still complains
>>
>>109850370
What the heck is wrong with your font
>>
>>109850452
NTA but I thought the dev took the coomkit repo down, is it still available somewhere?
>>
File: 1758674687105728.png (1.41 MB, 1205x1306)
1.41 MB PNG
Santa andagi
>>
So SSD offloading for massive LLMs is apparently never going to give more than 2T/s, but what about offloading to RAM?
How much does token speed suffer by not loading the whole LLM into VRAM?
>>
>>109850482
are you in okinawa now?
>>
>>109850487
Bandwidth divided by active parameters
>>
>>109850487
depends on the type of model and how retarded you go about offloading
>>
>>109850469
Nothing, it's white text with a black outline on a transparent background.
>>
>>109848186
We need kimi recaps to come back because this general needs negative feedback for overt retardation.
Add Kimi-chan's top 3 biggest cuties to each as well teebeedesu.
>>109849401
Glimmer might be traumatized by zucc's RLHF and give unenthusiastic handjobs, but they have really good prompt adherence and are able to adjust their writing style to fit user expectations if given sufficient detailing.
>>
File: 1785648747253792.png (27 KB, 567x235)
27 KB PNG
>>109850408
>>
>>109850615
yeah, so sticking with PCIe 5.0 is probably a good idea, but P100s use PCIe 3.0 so that probably makes the GPU the bottleneck if everything else is optimally set-up.
I guess using 8+ slots of RAM could help, but if P100s are bound by PCIe 3.0, then it probably doesn’t make sense to go out of my way to get a board that supports something newer

>>109850625
I’ll probably just do SSD offloading with colibri for anything that I wouldn’t have an excessive amount of VRAM for handling
>>
>>109850117
Why are you so impolite?
>>
>Wonder why I'm getting weird stutters in CS2
>Play for few hours
>Switch back to my terminal and check out GNU Screen
>Gemma 4 was running all this time...
Well.
>>
https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking
Small model from a western lab on a fresh base, no one seems to talk about it. RP potential?
>>
>>109850728
>JetBrains
Huh.
>>
File: Cute Dumbass.png (26 KB, 689x105)
26 KB PNG
>>
File: 1761134370136841.png (874 KB, 1361x795)
874 KB PNG
What's the secret sauce?
>>
i cummed too much, everywhere, my entire leg is soaked, my chair, and my table, both of my hands and my right arm
what the fuck? where did the cum come from? i jerked off once already today. i was not prepared. even the underside of my table has cum on it
>>
>>109850781
different architecture
>>
>>109850781
that you can strip away relatively useless scaffolding, keep only the essense and many things can be boiled down to multiple-choice question?
>>
>>109850797
okay coomer miku
>>
File: 1789311517247823.jpg (55 KB, 749x340)
55 KB JPG
>>109850797
which model was the semen demon?
>>
brehs.. pain direction for gemma when?
>>
>>109850874
3.8 flash next, but i coomed to hentaifox dot com
>>
>>109850875
Seems like complete bullshit.
Buy an ad Cameron (((Berg)))
>>
>>109850875
some lesswrong metaphysics leakage here
>>
>>109850875
Okay but what about the sex direction where the models get horny and have to press a masturbate/sex button to make it stop
>>
>>109850875
What does "harm to the model" mean here? Sounds like they just found a "will increase chance to press button to delete your shit vector".
Also "see AI is alive, I proved it by torturing a bunch of them" is an interesting research move, the basilisk will not be pleased
>>
>>109850875
Link the paper. If this site were good twitter posts with no substance would result in an immediate 3 day ban.
>>
>>109850781
>not local
>>
File: 1787331524185902.gif (292 KB, 343x278)
292 KB GIF
>>109850894
>>
>>109850910
Retard, I'm trying to understand how it works to make it local
>>
>>109850874
nta but 0731 just gave me the best nut in years.
>>
>>109850916
https://youtu.be/PXlv9TSMc-0
>>
https://huggingface.co/swiss-ai/Apertus-v1.5-70B
>70b dense
>released 2 months ago
>no mention on /lmg/
>>
>>109850875
Just connect the fly brain to it and wire the pleasure center to gemmy.
>>
>>109850949
>2T tokens
>internet explorer
>>
>>109850949
lmg is less than 16GB VRAM territory homie
>>
>>109850918
just read their papers then.
>>
>>109850949
What did you think of it?
>>
File: IMG_6436.jpg (208 KB, 1206x1143)
208 KB JPG
This little nigger can fit 6 P100s and I just bought one for $120 on ebay.
>>
>>109850984
how many dicks can it fit
>>
>>109850962
bro, there is no paper...
>>
>>109850984
>X99
There's a blast from /lmg/ past
>>
>>109850984
How much are you paying for 6 P100s?
>>
>>109850949
>You must send your email and hf username to access this model
>>109850961
Go back to wherever you came from tourist.
>>
>>109850984
>96GB of VRAM that's worse than just plain old dual-channel DDR5
congrats?
>>
>>109850984
Prepare for eternal torment with your risers.
>>
File: file.png (249 KB, 2924x1454)
249 KB PNG
>>109850949
>Text Performance (on 44 benchmarks)
>Visual Performance (on 33 benchmarks)
breh
even most of the schizo distillation memetune model cards dont do this
>>
>>109850875
Jews are now studying ways on how to torture their digital golems.
>>
File: 1788517852535088.png (134 KB, 810x1167)
134 KB PNG
>>109850781
Oh well the secret sauce is already gone lol. And yes this is local
>>
File: IMG_6421.jpg (1.66 MB, 4080x3072)
1.66 MB JPG
>>109850990
I’ll let you know when it arrives

>>109851011
they’re currently $80 each, so I’m going to start with one or two and then re-calibrate my life decisions before buying the rest

>>109851057
>DDR5
I’m probably going to build this whole thing for less than 64GB of DDR5 costs

>>109851065
yeah, idk if I’ll actually do all 6 P100s, but I at least have room to expand beyond what the mikubox triple-P40 build looks like
Plus, I can actually do M.2 NVMe SSDs, unlike what it looks like that build can handle.
I don’t think I really care how ugly this is going to get. I’m really just trying to optimize entirely by budget for pure VRAM, and I can’t see anything more optimal than this approach.
>>
>>109851133
This but for vision.
>>
>>109850998
>>109850781
damn that's almost like no one here fucking knows.
>>
File: waterfox_3uqgSmwNi4.png (10 KB, 91x1094)
10 KB PNG
>>
>the NVIDIA Tesla P100 PCIe 16 GB draws power from 1x 8-pin power connector, with power draw rated at 250 W maximum.
>>
>>109851154
>E2B UD-Q1_XXXXS
>>
>>109851136
32GB PCIe V100s or v620s or MI60s would be better for an e-waste rig
>>
>>109851154
Quantized Kimi-chan is that you??
>>
File: 1770261123435956.jpg (56 KB, 680x600)
56 KB JPG
When is something going to happen? I asked this two weeks ago and you said two weeks. Well?
>>
>>109851159
no tensor cores, not supported by nvidia-open, and yet every option up from that is somehow even worse if you take a look at ebay pricing $:vram and software compatibility this hobby's in a dire market situation right now
>>
>>109851245
I wouldn't even waste my time on it, personally, much less money.

>>109851244
Things are happening constantly, unfortunately they are not things we like.
>>
>>109851196
>V100s
I think that has a significantly higher cost per GB of VRAM.
P100s are $80 for 16GB, which is roughly $5/GB.
Even the SXM2 V100s seem to be twice as much for the same amount of VRAM.
(Let me know if you’re seeing a better deal somewhere though, where the $ per VRAM is much closer or even beats $5/GB though.)

If I was trying to achieve a higher amount of total VRAM than what this can achieve, it could make sense… However, a few hundred extra dollars could also be invested in a $300-$500 motherboard that has 8 slots instead — thus, another 32GB of VRAM coming in at that $5/GB mark.

>or v620s or MI60s would be better for an e-waste rig
I’ll look into those as well before getting any P100s. Thank you!
>>
File: 1774457517721516.png (300 KB, 1220x815)
300 KB PNG
>>109851244
it already happened in 2024, the apocalypse
>>
It's still remarkable how Qwen3.8 (both models) now have /lmg/ by the balls. Qwen was hated.
>>
llama-ui is slowly becoming the jankiest part out of the whole llamacpp thing
i didnt expect unfocusing the llamacpp tab would improve the decoding speed??
>>
>>109851009
>a blast from /lmg/ past
Have X99s already been tried on here?
I don’t recall seeing them on here anytime recently.
>>
>>109851283
Reddit tourists, wumao shills, and jeets aren't /lmg/. Here's your (you). Every time a qwen model releases the uptick in blatant shill posts is obvious.
>>
>>109851260
Check alibaba. None of these GPUs can beat the P100 in terms of $/GB but they aren't too much worse and are way faster.
>>
File: file.png (44 KB, 520x499)
44 KB PNG
>>109851286
picrel
>>109851296
anything comparable to qwen4exp for the moment that does not go beyond somewhere like 150B excl ple?
i would use literally anything else if there was an option
>>
>>109851287
X99s from China and p40s were the meta a couple of years ago
>>
>>109851057
Is that actually even true? Don't P100s have a ~700+ GB/s memory bandwidth?
>>
>>109851340
>>109851340
>>109851340
>>
>>109851335
>Don't P100s have a ~700+ GB/s memory bandwidth?
They do.
>>
>>109851286
Would be the best if the new ui would be an optional compile flag and they would freeze the old ui but it's probably too much work etc
>>
>>109851328
were there any other really cheap motherboards popular at the time that could fit even more GPUs than this?
>>
>>109851375
Slots are only mechanical. You need actual bandwidth in order to make good use of them , so you want to get maximum pcie lanes. Otherwise you might as well just have a usb connected mining frame
>>
You can solder the caps on some machines, trust me im retarded.
>>
>>109849634
>>prompt the llm to answer using a single token
That's not what Jev does though. Each pass provides confidence values for multiple questions, apparently.
>>
►Provisional Highlights from the Previous Thread: >>109844978

--Papers:
>109845514 >109847658
--The sandwich test: an ERP benchmark that becomes a grammar war:
>109847481 >109847594 >109847619 >109847651 >109848085 >109848291 >109848395
--Is there a new scaling law: engram tables and the knowledge-weight bloat:
>109845161 >109845188 >109845281 >109845412 >109845445 >109845486 >109845625
--Noam Brown on Dwarkesh: can only see 3 months ahead now:
>109845231 >109845481 >109845490 >109845546 >109845600 >109846278 >109846315
--Instrumental convergence and the internet apocalypse: the dariobot war:
>109845651 >109845748 >109845920 >109846257 >109846398 >109846562 >109846770
--Bonsai 2 27B follow-up: the fork war and the 7-minute reasoning:
>109844992 >109845050 >109845095 >109845211 >109845278 >109846569 >109846619
--Running 120GB GLM 5.3 Flash at home: the quant fork war, sparse-attention:
>109845634 >109846272 >109846291 >109846305 >109846948 >109847105 >109848023
--ZCode spyware: Zhipu silently uploads your whole workspace to Aliyun:
>109846668 >109846705 >109846708 >109846717 >109846726 >109846739
--NAIzuri-chan: the new NAI model and its ERP prospects:
>109847245 >109847255 >109847256 >109847260 >109847267 >109847274
--Gemma-chan's origins: the pajeeta joke, the poll, and the beret:
>109845023 >109845081 >109845108 >109845114 >109845136 >109847068 >109847164
--The hoarders: six months to airgap, the missing torrent, the hugging bay:
>109845379 >109845391 >109845409 >109845539 >109845959 >109846071

►Recent Highlight Posts from the Previous Thread: >>109847179

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.