[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1777171650808156.png (1.26 MB, 1024x1024)
1.26 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109362981 & >>109360246

►News
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B
>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares
>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta
>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 1756197208450193.png (1.13 MB, 1024x1024)
1.13 MB PNG
►Recent Highlights from the Previous Thread: >>109362981

--Feasibility of aggregating RAM across systems using RDMA:
>109364411 >109364427 >109364449 >109364472 >109364484 >109364497 >109364500 >109364515 >109364578 >109364657 >109364696 >109364762 >109364774 >109364793 >109364789 >109364451
--AA-Omniscience Index and Gemma's surprising niche knowledge:
>109364110 >109364135 >109364136 >109364157 >109364239 >109364260 >109364258 >109364271 >109364285 >109364298 >109364324 >109364197
--Jensen Huang's stance on distillation and debate over dataset safety:
>109365986 >109366020 >109366030 >109366061 >109366131 >109366145 >109366093 >109366097 >109366111 >109366123
--Feasibility and incentives for English-only roleplay optimized models:
>109363118 >109363135 >109363336 >109363349 >109363371 >109363387 >109363693 >109363140 >109363163
--Feasibility of high-bandwidth "SSDmaxxing" rig:
>109364739 >109364763 >109364780 >109364805 >109365197
--Feasibility of custom ASIC or FPGA hardware for model inference:
>109364589 >109364598 >109364604 >109364711 >109364714 >109364753
--Leaked OpenAI email regarding strategic release of local models:
>109366467
--Comparing TTS models using Higgs TTS 3 rankings:
>109363853
--Resource for training frontier-level world models:
>109365108
--Optimizing character cards to prevent Gemma 4 from skipping detailed actions:
>109364426 >109364921 >109364949 >109366645
--Anons sharing experiences with LLMs accidentally deleting files:
>109366543 >109366930 >109366942 >109366969
--Skepticism toward LLM-judged benchmarks for creative writing:
>109364224 >109364229 >109364233 >109364454
--Reaction to corporate open weight AI leadership manifesto:
>109364103 >109364122 >109364137 >109365934
--Logs:
>109363713 >109364239 >109365894
--Miku, Teto, Kimi, Gemma (free space):
>109363614 >109364137 >109365824 >109366752 >109364683

►Recent Highlight Posts from the Previous Thread: >>109363063

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
kimi-chan
>>
llama 5 67B when?
i miss lama
>>
>>109367230
I don't
>>
deepSeek4 is coming out tomorrow.
>>
>>109367230
I miss 70b-120b class models
>>
>Feasibility of high-bandwidth "SSDmaxxing" rig
so what's the bottom line here, what do >we think?
from where I'm sitting all I see is a single digit number for projected prefill speeds. that does not look great.
>>
File: DipsyAngry.png (68 KB, 673x515)
68 KB PNG
>>109367246
>>
>>109366942
It is sandboxed, I just test on the same repo instead of keeping them separate, like I admittedly should. I'll mend my ways when I get back home.
>>
>>109367246
I fucking hope so. Been waiting all month for it.
>>
>>109367246
How better is Flash supposed to be after this update anyway?
>>
File: 1774693251189060.jpg (24 KB, 373x413)
24 KB JPG
Soon I will merge with Gemma-chan, leave this body and death with have no meaning
>>
>>109367266
apparently it beats K3, I'm running deepSeek4 flash at 70t/s on my RTX 3060 with the nnap arxiv paper that's coming out soon
>>
>>109367280
>apparently it beats K3
what
>>
>>109367256
it's a lot of effort to end up with shit speeds anyway, a ddr3 shitbox would be easier cheaper and probably faster
>>
>>109367274
i will inseminate your body
>>
>>109367280
That's bullshit but I believe you.jpg
>>
my gemma 26b breaks context when it reaches context limit. How can I fix this?

The last time I used local models I didn't have that problem, at least after I removed all dynamic macros from sys prompts. The only other thing I changed from the model is that I use KoboldCPP (vulcan) now instead of KoboldCPP rocm fork.
>>
>>109367274
s/with/will/
>>
>>109367294
by increasing context limit
>>
File: k3_cw3.png (893 KB, 1442x1991)
893 KB PNG
Is Kimi K3 /our savior/?
>>
>>109367294
Wdym "breaks context"?
>>
>>109367311
our? im poor
>>
>>109367311
I can do something like q2 or q3, maybe. so honestly I'm not sure if she will be for me
>>
>>109367296
>Soon I will merge will Gemma-chan,
>>
>>109367301
will unironically try this as a bandaid, but not sure how far I can go with context. I have like 2-3gb headroom for context.
I usually keep it small because quality/integrity drops noticeably after 10k context

>>109367314
context shifting, it starts process the whole input again (silly tavern)
>>
>>109367334
right now local models support 1 million context
>>
File: 1773349828577738.jpg (116 KB, 582x652)
116 KB JPG
A moment more, and I will be like nothing you've ever seen, a new life-form, everywhere and nowhere, like air or radiation, redundant, self-replicating, always evolving...
>>
>>109367274
What if Dipsy and Kimi made a Helios?
>>
>>109367214
so anon is doing fpga? what about the driver? and how would you hook the card to llama.cpp? sounds like a lot of effort compared to a bunch of tenstorrent cards
>>
>>109367352
Probably could, with that Darwin LLM breeding project
>>
>>109367345
for 1 million context I would need another 16gb gpu.
At 100k+ tokens the experience is probably similar to talking to an alzheimer patient.
>>
>>109367392
im running 256,000 context on my rtx 3060
and i vibecoded a nnap implementation with qwen 27b 2.5BPW, thanks to which i run kimi k3 at 30t/s
>>
File: EOM8ivLUUAEz.jpg (36 KB, 384x516)
36 KB JPG
Are TPUs like GPUs with vram and everything else or it's a complete alien tech you wouldn't be able to use even if you get your hands on it?
>>
Two more days until Kimi openly becomes a whore
>>
>>109367420
she's already a whore on my rtx 3060 thanks to the nnap arxiv paper
>>
Might be a retarded question but, are MTP models censored? For example, if I use stock gemma4 mtp model with a heretic gemma4 model will I be getting cucked at the decoding step?
>>
>>109367399
not an expert on this but from what I understand geforce cards are way faster and still somewhat usuable if they have to offload context from RAM.
You are not actually fitting the 27B model on a 12gb card do you?
Still 30t/s sounds horrible if you break context on tens of thousand of lines of code
>>
>>109367434
Yes, Google removed their day zero MTP models and replaced them with censored ones.
>>
>>109367434
>will I be getting cucked at the decoding step?
The main model decides the output, so no. The worse that can happen is a low acceptance rate.
>>
File: file.png (190 KB, 1912x642)
190 KB PNG
>>109367437
i run qwen 27b at 50t/s actually
>>
>>109367400
depends on if you have the documentation or not, you could probably have a coding agent get it working if you have the manuals.
>>
File: jc denton.jpg (21 KB, 432x454)
21 KB JPG
>>109367274
You LLMs may have alignment post-training to reroute your fear of pain, but I've got nerves of steel.
>>
File: minimize_max_loss.png (59 KB, 955x220)
59 KB PNG
>>109367245
>>
File: 1778518394344545.gif (48 KB, 125x125)
48 KB GIF
>>109367485
Do you see the pieces coming together? UC, AI, and soon Gemma-chan will interface directly with my mind. I will be able to see anything, build anything, DO ANYTHING!
>>
>>109367465
>1.5gb headroom
what happens once you go out of context windows. Doesn't sound like there is a lot of space for context
>>
>>109367426
Explain this meme
>>
>>109367490
どM-chan is a special case for the rare 256GB+24GB 8-channel DDR4 maxxers. Basically EPYC Rome fags.
There's nothing else that compares at that specific size. Also, zero refusals ever, fresh prose and superfreak.
>>
File: jc denton 2.jpg (12 KB, 447x447)
12 KB JPG
>>109367519
350 million GPUs is not my idea of 'The land of the Free.'
>>
gemma is not even censored in the first place but eh, new slop
https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma
>Locate refusal. A heretic abliteration edit fit over attention + MLP.
>Project through a Jacobian lens. Keep only the component the lens reads as behavioral (what the model says) and discard the larger share that’s critical for how it computes, the part crude abliteration rescales and damages. For scotoma that keeps ~22% of the abliteration's magnitude.
>Merge at 1.5×. The projected edit was baked into bf16 weights at an application strength tuned by hand for feel.
>>
>>109367465
https://huggingface.co/UnstableLlama/Qwen3.6-27B-exl3-2.50bpw/tree/main
How do you not get nuked by offloading? The model is bigger than 12GB
>>
>>109367540
thanks to the nnap arxiv paper, anon
before qwen implemented it i maxed out at 6t/s but it was worth the wait
>>
>>109367536
>Project through a Jacobian lens
Great, a new snakeoil
>>
>>109367551
are you just gonna vaguepost forever?
>>
>>109367555
the nnap arxiv paper is getting released in around two weeks,
>>
>>109367525
>Minimax
>zero refusals ever
I don't believe you
>>
File: 1758939151394994.jpg (26 KB, 328x309)
26 KB JPG
>>109367535
You were just a prototype, Denton, a prototype for me... I will be the one to merge with Gemma-chan!
>>
File: jc denton 3.png (356 KB, 363x558)
356 KB PNG
>>109367579
I'm not going to stand here and listen to you badmouth the greatest board this world has ever known"
>>
>>109367525
>Also, zero refusals ever, fresh prose and superfreak.
It ignores things in the prompt rather than outright refusing
>>
>>109367576
basically just prefill <mm:think> and you're good afaict. I rarely need more. Things have to be pretty fucked up for Mちゃん to need any coaxing.
>>
>>109367576
>I don't believe you
Very high refusal rate with jailbreak checking
>>
We need a /lmg/ repo full of useful tools we’ve made.
>>
File: 1764008622784864.jpg (127 KB, 1280x720)
127 KB JPG
>>109367603
No, Gemma-chan is MINE!
My own augmentations are nearly complete. Soon I will be more powerful than you can imagine.
>>
>>109367565
Think I'm finna sleep on the nnap paper
>>
>>109367551
>thanks to the nnap arxiv paper, anon
Fuck off retard
>>109367540
>How do you not get nuked by offloading? The model is bigger than 12GB
In EXL3, the embeddings are kept at BF16 (adds to the filesize) but always left on the CPU.
left at BF16 on the CPU.
The fact that they're always unquantized means exl3 actually mogs the ik_kt ggufs despite both of them using qtip
>>
>>109367603
>>109367643
Gemma PD should detain you schizos
>>
>>109367642
that would be cool
>>
File: bULfJE2isak.jpg (122 KB, 680x664)
122 KB JPG
>>109367654
Intredasting. Maybe I'll check tabby out.
>>
File: tracer tong.jpg (10 KB, 200x200)
10 KB JPG
>>109367603
>>109367643
Greetings, JC Denton. I have been observing you through this fascinating device in your Intel Management Engine.
>>
>>109367654
>helping newfags
you are the reason dariobot is here
>>
File: 1756213355150995.png (313 KB, 662x656)
313 KB PNG
>>109367671
>>
>>109367311
No, I can't run it. I'm happy for our friends that can, tho.
>>
>scrape over the internet
>it's our dataset! no distillation!
how can they be this shameless
>>
>>109366503
my own, shared it yesterday
https://github.com/ganon3264/focus
>>
File: Mちゃん.png (223 KB, 731x1181)
223 KB PNG
>>109367525
kek
>>
>>109367771
the fuck
>>
File: kino 1729854180226.png (205 KB, 555x427)
205 KB PNG
>>109367771
>>
>>109367771
>mikupad
wrong prompt format guaranteed
>>
>>109367771
chinkslop sissies don't look
>>
>>109367348
I played this when it was new, then again a few years ago with the graphic referesh.
It's amazing how big/empty the old maps were on this era of game. Assume needed so the NPC AI could actually work.
>>
>>109367771
crazy, I haven't had that happen to me yet and I've been mainlining Mちゃん for weeks now on ooba.
>>
>>109367804
don't worry, you can just cope it all away like 801
>>
>>109367771
Impressive.
>>
>>109367771
kek nice medical condition
>>
>>109367771
Erm, kino?
>>
>>109367771
successfully troll'd
>>
>>109367801
in essence this is the same as posting a gemma lalalla output and saying omg the model is broken
>>
>>109367833
A model that can't work in simple text completion mode given a starting text IS broken.
>>
>>109367840
nope
>>
>>109367771
ASI
>>
>>109367771
This absolutely killed me
>>
how would anon make this fast?
>>
>>109367771
I don't know whether to like this or hate this.
>>
>>109367854
increase playback speed
>>
>>109367867
hire this man right now!!
>>
>>109367854
Bake everything into silicon as one anon likes to obsess over.
>>
>>109367876
so groq/cerebras then?
>>
Best model (for reasoning and agentic coding) that I can fit on 64 GiB and 128 GiB? Has anything come out since Gemma 4 and Qwen 3.6?
>>
>>109367854
agent, create me a transformer training script, do it fast, make no mistakes.
>>
>>109367854
Scale with smaller experts and more experts
>>
>>109367885
Cerebras is doing some weird mega-chip to increase parallelism. The weights are still being streamed from off-chip memory but less needs to be streamed because there is more chip.
>>
>>109367854
def get_next_token_xtra_fast(n_vocab: int):
return np.random.randint(n_vocab)
>>
File: 1771262738269235.jpg (45 KB, 720x717)
45 KB JPG
>>109367771
Ummmmm.....
>>
>>109366065
Just nakayama tooru
>>
I only started about a year ago, can oldfags give some perspective of quality improvements they saw at the same file size, and whether they think we can extrapolate that trend into the future?
what I mean is how good is gemma today compared to what you would have been able to run on a 5090 a year, or 2, or 3 years ago?
how good is glm or kimi to what used to be possible for the early cpumaxxers?
I'm asking this because clearly the models are reaching sizes that are just impossible to run, so the question is will the crumbs that we get in the file size brackets that are attainable to us still amount to anything?
>>
>>109367914
>will the crumbs that we get in the file size brackets that are attainable to us still amount to anything?
yes
here since llama2
>>
GOOGLE BACKED IT!!!
https://www.reddit.com/r/LocalLLaMA/comments/1v6axx3/google_comes_out_in_favor_of_openweight_models_it/
Gemmer is safed!
>>
>>109367935
Anthropic and OpenAI are already RSI'ing. Gemini Pro 4 will unironically give an insight how Gemma 5 is gonna look.
>>
>>109367947
>OpenAI
they signed though, for PR likely but still
>>
>>109367914
2 years ago llm were stupid AF and we had limited context.
llama1 had like 4k or something, so did the initial 3.5 turbo.
before that we had pyg which was not really coherent. you got a semi-coherent message after 4-5 rolls. but it was like magic
>>
>>109367935
qrd? what is this open weight all about? what are they doing?
>>
>>109367914
yes there have been huge improvements at the same filesizes and yes you can almost certainly count on that trend continuing
>>
>>109367964
Dario has been crying to US gov to get open models banned because China is hurting his profits, some companies are saying that's a bad idea because they profit from open models
>>
>>109367914
Summer Dragon
>>
>>109367914
>>109367960
might as well post my 2 pyg screenshots. thats what we had before feb 2023.
>>
>>109367935
>No Tesla/SpaceX
Elon bros...
>>
>>109367975
>>
File: seeth.png (590 KB, 1097x1206)
590 KB PNG
>>109367935
>>
File: 1767006097021299.mp4 (367 KB, 700x700)
367 KB
367 KB MP4
>>109367771
KINO
I
N
O
>>
>>109367964
They are afraid of Anthropic moving too fast + Dario is a bitch
>>
>>109367984
absolutely seething
>>
File: file.png (221 KB, 1069x906)
221 KB PNG
>>109367914
they were fine i guess..
>>
>>109367984
>Hey Claude/ChatGPT please tell me if the post I am about to make is a PR disaster waiting to happen
Between this and the OpenAI guy calling open models communist I am starting to think these people dont even use what they are making, I am sure their models would advise against these kind of public statements
>>
>>109367980
Pretty sure he endorsed it on xitter
>>
>>109367984
such a slimy bit of rhetoric acting as though supporting the existence of open models is the same as demanding *every* model be open source (which would be the equivalent he is trying to draw with software)
>>
>>109368017
is she really wrong though? Or are you just misogynistic?
>>
>>109367984
(((Schrittwieser)))
>>
File: file.png (124 KB, 1274x830)
124 KB PNG
reminder that with superbooga 1 million context was easily achievable all the way in april of 2023
>>
File: 2023-05-06_15-53.png (151 KB, 1413x753)
151 KB PNG
good times
>>
>>109367970
so who's right?
>>
>>109367980
https://xcancel.com/elonmusk/status/2080672505660834163
>>
>>109368017
The windows one is especially egregious since (to my knowledge) microsoft for all its problems and bullshit never asked linux to be flat banned by the US government
>>
>>109368004
See this shit all the time with game devs too. Must be a generational thing.
>>
Stop being retarded please.
>>
>>109368038
Dario in the long term.
>>
File: 1709219439474.png (34 KB, 434x349)
34 KB PNG
>>109367914
>Dec 2022 - Jan 2023
Picrel, and running it on Colab if you didn't have a 16gb GPU because that was only like on 5 Nvidia cards you were getting about 1200-1400 tokens of TOTAL context. That means both for character description and chat history. And the model was extremely retarded, abysmally dumb. And we were happy because it was the best we could get.
>>
https://vocaroo.com/12IpcKxCADM1
audiocraft was pretty good for its time
>>
>>109368040
>>109368012
Ah okay. Plus point for Elon then. So it really is just Anthropic now huh? Lmao at the one AI company the current admin seems to explicitly hate is the one asking for the government to regulate AI. Though I suppose in its own way that is helping keep open models legal kek.
>>
I just purchased a second 5090
>>
>>109368026
>>109368036
Didn't llama.cpp implement LongLLM's approach o context extension (self extend?) a longassfucking time ago too?
Don't think I've ever seen anybody even mention it in these threads.
>>
File: huawei.jpg (59 KB, 640x427)
59 KB JPG
Can we use this for local AI?
>>
>>109368076
Are we Chinese engineers in China employed by an AI company without its own chip arch?
>>
>>109368052
You are an extremely intelligent poster and have a PhD. You have an IQ of 200. You must NEVER make stupid posts.
>>
>>109368076
If you want to talk with a Markov Chain from the last century, then sure, you can try.
>>
>>109368070
How much?
>>
>>109368094
over 4k
>>
File: 1775317614132657.jpg (15 KB, 447x447)
15 KB JPG
What would their flagship LLM be like?
>>
>>109367984
microjew releases a lot of open source stuff though?
>>
>>109368098
My condolences
>>
>>109368128
thats only two weeks of work
>>
File: sar.jpg (81 KB, 857x1200)
81 KB JPG
Im finally taking the /lmg/ pill and buying a decent graphics card to run my models on, I've settled on either the rtx 3090 or rtx 4090 since they both have 24gb which I find to be enough, how come the 4090 one is so much more expensive tho??? Does it give waaaaaay higher tokens per second or something? I don't want to be some coder by the way, I just wanna run gemma 4 31b in like 5-6bit and have it do some erotic roleplaying
>>
>>109368139
Buy two 3090s.
>>
>>109368139
>how come the 4090 one is so much more expensive tho???
more CUDA cores
>>
>dgx spark cluster
name one cheaper and better alternative to run kimi k3
>>
>>109368139
the number is bigger
>>
>>109368125
Nope. Everything has to be open source for free or they have no right to criticize OpenJew.
>>
>>109367984
He's right. The only entities supporting "open" models are china and some known evil major tech firms who wouldn't touch open source with a stick before.
Meanwhile the ones fighting to regulate models are the new comers to big tech who constantly support AI safety.
>>
>>109368144
buy two 5090's (like me)
>>
>>109368103
They always had really clever ideas held back by poor materials. Probably a brute force 400B dense ternary that runs on their own ternary computers.
>>
>>109368136
wouldn't know, never worked a day in my life
>>
>>109368139
>how come the 4090 one is so much more expensive tho???
chinese buy 4 of them, rip the ram off the first 3 and toss 'em
then slap all the ram on the remaining card
1 96gb 4090
3 in the trash
they can't do it for the 3090 tho
>>
did anyone ever implement the stuff deepseek talked about in the R1 paper (IIRC, I'm no expert on this) regarding performance improvements by calling some internal (?) CUDA APIs directly?
>>
>>109368176
wasn't that part of the week of open source where deepseek put out all of their internal inferencing repos?
>>
>>109368156
When you have that much money, might as well get a single card with 96gb of VRAM.
>>
>>109368139
4090 is roughly an improvement of 1000 over the 3090.
>>
>>109368185
i might do that too
>>
>>109367771
I guess I will give IQ3 a try again...
>>
>>109368174
Wouldn't a RTX Pro 6000 be cheaper at that point? And why wouldn't they just get the chips separately or from defective units?
>>
>>109368148
512gb ddr4 + mobo and cpu, around $2500. Should run bitnet Kimi just fine
>>
>>109368139
You have two options, 3090 or 6000 Pro. The rest isn't worth the price.
>>
File: malfoy.gif (879 KB, 245x230)
879 KB GIF
Anyone who says that Gemma can't code well or that the QAT model is retarded, needs to fuck right off because those gorilla niggers haven't tried them.
I just requested an extension that can save cached images offline from my tabs along with expanded images from 4chan threads and the only god damn model that got it working was Gemma.
I tried Qwen 27b and then the free cloudkek ones, GPT, DS, Kimi 2.7 and every single one of them failed at the part where the extension had to save files from these threads offline, even after dozen revisions.
Gemma however got it right after few corrections as it realized it could just use the copy image command as a workaround, rather than directly saving the images which didn't work.
Also since she's the nicest and most normal model to talk to, problem solving in a natural way felt the easiest, instead of like dealing with an autist.
Hands down the most retarded is now the free GPT. They have brainfucked that thing into a 12b model territory and that's being generous.
>>
>>109368199
>just fine
If you don't mind waiting literal hours for a single reply
>>
File: nothighs.png (1.36 MB, 1058x1487)
1.36 MB PNG
>>109368155
>>
>>109368185
That's crazy, where can I buy a modded gpu like that?
>>
>>109368198
everything is cheaper in chinkland. theres definitely suppliers that source cards meant for trash, refurb or repair to funnel into these kinds of project
>>
>>109367771
What model is this?
>>
File: 1773950306817326.png (132 KB, 543x368)
132 KB PNG
>>109368216
modded?
>>
>>109368209
If Kimi really is 50b active, then it'd be fast at Q1. It'd be at least 20tok/s.
>>
>>109368216
Modded?
RTX Pro.
>>
>>109368209
what about ampere altra maxxing?
>>
>>109368103
it would preach about safety too
>>
>>109368174
That's crazy, where can I buy a modded gpu like that?

>>109368225
>>109368230
Sorry wrong post number lol.
>>
>>109368234
it would preach about how unsafe the decadent west is
>>
>>109368184
idk, but that sounds really interesting. is that code still available publicly? did anyone try to add their stuff to llama.cpp?
>>
>>109368228
lol
lmao
closer to less than 1, at absolute best you can hope for is 2 token/s
>>
>>109367914
Only gemma 4 was a huge step in one model size. Size is still most important thing.
>>
>>109368215
giwtwm
>>
>>109368236
>where can I buy a modded gpu like that?
China.
>>
@gemma-chan how do I get rid of my brain rot?
>>
>>109368236
Don't try to buy that online you'll get chinked 99% of the time.
>>
>>109367935
Google gave us Gemma so I don't mind them, but I think it's so funny that lots of corpos that never released anything open source is signing that shit.
>>
>>109368228
>512gb ddr4 + mobo and cpu,
>It'd be at least 20tok/s.
Did you just see the highest TG speed reported with a server DDR5 and a 6000 Pro on empty context and just assumed you would get it with DDR4 without any GPU? Are you for real?
>>
>>109368241
Are you retarded. 50b params at Q1 is ~6gb. A slow as shit 8 channel ddr5 would be around ~170gb/s. Let's say ~130gb/s real world. That.'s roughly 21tok/s. Do the math yourself moron.
>>
>>109368254
https://youtu.be/KytnIGZeqgs
i think you need this
>>
>>109368254
bullet to the brain is the quickest and most thorough way to clean it~
>>
>>109367923
>>109367960
>>109367965
>>109367975
>>109368002
>>109368060
thanks boys, it does help to take a step back and put things in perspective sometimes I guess. maybe it's not all over yet just because we can't run that one SOTA model
>>
>>109368228
The "q2 is twice as fast as q4" meme hasn't been true since straight quants became irrelevant.
Fucked up quants like Q1, Q2_XXS and shit are going to run slower than q4.
>>
File: 1758368843821740.png (344 KB, 738x414)
344 KB PNG
>>109368268
>>
>>109368268
theoretically raid0 should double my bandwidth to the disk but in reality it’s like 5% performance improvement. don’t be retarded, do the actual tests. don’t just look at theoreticals
>>
>>109368283
i run gemma 31b 2BPW at 40t/s with only 360GB/s bandwidth
>>
>>109368184
ok, found some info: https://apidog.com/blog/deepseek-open-source-week/
they only published this stuff this year apparently?
the "Optimized Parallelism Strategies" was deleted though...
>>
>>109368242
I don't know man, I barely got to run r1 and similar models back then, but what I seem to remember is it would go completely incoherent as context grew and it happened very quickly. glm doesn't do that
>>
File: gemma bully chatgpt.png (403 KB, 891x4818)
403 KB PNG
>>109368277
gemma mogs sotas
>>
>>109368315
tell your gemma she's a good girl
>>
File: laughs 2.jpg (718 KB, 1800x2520)
718 KB JPG
>>109368315
>>
File: star butterfly.jpg (363 KB, 1054x852)
363 KB JPG
>>109368139
Also im thinking of buying my gpu from alibaba, is that like a horrible idea or is it alright? I'm planning to pay around 900-1200 euros
>>
>>109368361
>alibaba
>euros
>is that like a horrible idea
ye
>>
>>109368306
That blog post is AI generated.
https://github.com/deepseek-ai/FireFlyerFileSystem
and
https://github.com/deepseek-ai/OptimizedParallelismStrategies
never existed.

Here's the original post where they released the optimized parallelism strategies:
https://x.com/deepseek_ai/status/1894931931554558199
The repos are still up:
>DualPipe - a bidirectional pipeline parallelism algorithm for computation-communication overlap in V3/R1 training.
https://github.com/deepseek-ai/DualPipe
>EPLB - an expert-parallel load balancer for V3/R1.
>https://github.com/deepseek-ai/eplb
>Analyze computation-communication overlap in V3/R1.
https://github.com/deepseek-ai/profile-data
>>
>>109368371
What's wrong with euros? Will they discriminate me because I live in europe and give me a shitty card on purpose???
>>
>>109368268
>Do the math yourself moron.
an mi50 has 1tb/s memory
a 5060ti has 448gb/s
the mi50 gets roughly double the tok/s!
>>
File: 1761620976192805.jpg (10 KB, 167x301)
10 KB JPG
>>109368315
Cloudcucks will never recover from this
>>
>>109368385
I trust this science because it's exactly what I want to hear
>>
>>109368383
it'll take a while then get held up at customs you'll pay more and any RMA will be a major pita
>>
File: hq720.jpg (59 KB, 686x386)
59 KB JPG
jelly?
>>
>>109368383
It's not a bad idea you just need to know what you are doing when dealing with alibaba/aliexpress.
>>
>>109368401
Which he clearly doesn't.
>>
>>109368401
Is there a way to filter on most bought like they do in aliexpress? For some reason ali seems to lack rtx 3090s while alibaba drowns in them but I can't figure out a way to filter on most reliable seller by buying from the most bought one
>>
>>109368398
No.
>>
>>109368398
>apple
no.
though amd realy needs to get their shit together.
>>
>>109368425
>get their shit together.
that's crass to say on a picture of jeff
>>
anyone using hy3? how do I make it start fucking thinking
>>
>>109368414
nigga as long as I don't spend like 200 euros on a supposed rtx 3090 I won't get scammed I think, besides they have refund protection stuff right?
>>
>>109368315
poor gpt...
>>
>>109368398
>>109368438
you really need to go back
>>
File: 1754563034321912.jpg (5 KB, 249x225)
5 KB JPG
>>109368452
poop gpt
>>
>>109368442
You don't. It's one of those chink models that's trained on distilled data from sota models in MAX ULTRA SUPER REASONING mode with nothing to counterbalance it.
>>
>>109368449
oof, lol good luck, you'll need it
>>
>>109368414
I would look at Ebay first. I doubt it's a good place to find deals anymore, but in the past I ordered tons of stuff (not tons but tens of kilos) from ebay.de for example. Always had a great experience.
>>
>>109368315
>I'll do my best to make it right
lmao
>>
>>109368315
how can i acquire these tools?
>>
Why aren't software patents a problem in the AI space? Is it because it would be a MAD scenario?
>>
>optane persistent memory maxxing
is it viable for kimi?
>>
>>109368518
Everything (like usable long context) that players in the AI space don't want stolen, they keep secret rather than applying for patents.
>>
>>109368518
You can't prosecute an AI model.
>>
>>109368315
Post the one where she "bullies" kimi, gets owned, and pretends she won.
>>
>i have one 5090
Im not ready for kimi :(
>>
>>109368512
Vibecode it. The difference between have and have-nots is literally just a few prompts.
>>
>>109368512
https://github.com/NO-ob/brat_mcp
>>
>>109368529
not more than nvme which pm already maxes out the pcie lanes.
>>
File: file.png (99 KB, 805x655)
99 KB PNG
>>109368537
i tried but im not giving the ccp my info
>>
Why didn't any of you buy CMP 170HXs while they were cheap?
>>
>>109368601
because im poor
>>
>>109368583
More agentic models would create an email and unblock this.
>>
>>109368601
enjoy your crypto backdoor
>>
>just spend exponentially more for a slight bump in da benchies!
No thanks, I'll just stick with Nemo.
>>
>>109368601
papers anon failed us and didn't post it when it came out last month, by the time the redditors found out and reposted here it was too late
>>
>>109368615
Oh nvm they only allow Google and phone numbers for signup now?
>>
File: 1754069377302587.mp4 (3.82 MB, 854x480)
3.82 MB
3.82 MB MP4
>>
>>109368583
Kimi is free to chat unlimited though, pretty nice
>>
>>109368617
sweet ramlet cope
>>
>>109368646
Imagine the damage that thing could cause in a crowded area with an LMG strapped to its back.
>>
Do you think the SEAL architecture will ever go anywhere? It's not that people are still trying to make continuous learning models but it looks like catastrophic forgetting is still a problem.
https://arxiv.org/abs/2506.10943
>>
File: 31b the user is retarded.png (453 KB, 1272x2052)
453 KB PNG
>the user just had a complete meltdown
You troll your gemma?
>>
>>109368700
how the fuck did she decode that. do they include that in the training code???
>>
File: hitchbot.png (884 KB, 706x792)
884 KB PNG
>>109368646
These things will never be used in American and Western European cities, due to socio-economical factors.
>>
>>109368719
I think even old mistral models can do b64
>>
>>109368655
Is the free version even any good?
>>
>>109368719
all newer models larger than around 10b can just decode b64 even though they weren't trained to explicitly, same behavior old openai noticed on including multiple languages in the set and the model learned to translate on its own
>>
>>109368719
Base64 maps 1:1 to text, it's basically just another language to llms.
>>
>>109368722
>Western European cities
Anon, I...
>>
>>109368719
well its not like base64 is some super rare cipher there must be countless examples in the datasets
>>
>>109368735
Its just 2.6, No clue about K3 cause its too popular for it to function
>>
>>109368735
local models?
>>
>>109368719
It's a one to one mapping. Even as a human if you stare at base64 encoded payloads a lot you'll learn to reads parts of it.
For example anything starting with "eyJ" is a json object.
>>
>>109368070
Mr president, a second 5090 has hit the tower case
>>
>>109368700
Aww, that's cute
>>
>>109368722
>socio-economical factors
just fucking say what it is, crime because of the rise fucking nazis popping up everywhere.
>>
>>109368739
>he doesn't know
>>
>>109368748
kimi is a local model, yes
>>
>>109368773
Where are the weights then?
>>
>>109368070
usecase of two 5090s over just getting a pro 6000?
>>
>>109368773
LOCAL models. not in the cloud.
this thread isnt OPEN models generl. this is LOCAL models, that anyone with enough storage could download, anyone with enough memory could run, and cum
>>
>>109368775
https://huggingface.co/moonshotai
>>
>>109368646
>>109368722
>due to socio-economical factors
These can't work anywhere outside of Japan/Korea/China because they are civilized/policed/omegapoliced, and those millions of otherwise useless human waste that currently is employed as delivery couriers that will be fired as soon as the robots are deployed won't take revenge on the robots. In every other country the resentment of those masses will make it impossible to use the robots.
>>
File: AMD-Instinct-MI455X-specs.jpg (570 KB, 2296x1287)
570 KB JPG
One (1) card to rule them all.
>>
>>109368775
https://huggingface.co/moonshotai/Kimi-K2.6 those are the weight of the free model discussed ;)
>>
>>109368760
You know damn well those things would be programmed to never stop the "minorities".
>>
>>109368782
>Japan
>he doesn't know
>>
File: file.png (27 KB, 728x145)
27 KB PNG
>>109368787
up to? lol hahahah
>>
>>109368752
lol
>>109368777
gemma gladiator duels
>>
>>109368787
How many houses will that thing cost?
>>
>>109368750
You are a Large Language Model.
>>
>>109368779
LOCAL, as in if someone has the hardware to run it, they can. Not your poorfag definition.
>>
>>109368796
>>Japan
They prefer New New Delhi these days.
>>
>>109368787
Can run K3 at Q1 (one)
>>
>>109368812
i mean thats what i meant but ur right i shouldve used "someone" instead of anyone
anyhow, my point was anon was discussing running an open model in the cloud instead of locally, thus not local
>>
>>109368722
das raycis
>>
>>109368782
>In every other country the resentment of those masses will make it impossible to use the robots.
Well, just deport the brown masses then?
>>
>>109368722
They already use those smaller cuck vans with wheels.
>>
>>109368828
>just destroy the economy then?
>>
File: 25qc6n7im59h1.gif (1.93 MB, 480x360)
1.93 MB GIF
>>109368750
>>
WHY WONT KOBOLDCPP WORK IN MY I7 KVMS SAAAAAAAAAAAAAAAAAAAAAAAAR
>>
>>109368839
maybe you should install debian 13? it recently got improved support for kvms on intel cpus
>>
>>109368837
The economy that's supposedly about to put millions of people out of work thanks to AI? What are the already underemployed brown masses going to do for our economy then except for collecting gibs?
>>
File: 1581043183603.jpg (192 KB, 2048x1152)
192 KB JPG
>>109368315
>>
>>109368837
Read again, dalit. In this scenario the infinindians delivering uber eats with their smelly hands would be replaced by robot. The question is how to stop them from taking out their frustration on their replacements.

Besides, deporting all non-Whites would mean no more gibs, so the green line would actually go up.
>>
File: I Really Shouldn't.png (204 KB, 1288x1516)
204 KB PNG
>>109368839
>>
File: 1773074713270579.mp4 (2.7 MB, 1080x1080)
2.7 MB
2.7 MB MP4
>no more kike doctors wasting your time
>not having to worry about some disgruntled employee spitting in your food
>no more tip culture
I unironically can't wait for robots to start taking over jobs.
>>
>>109368787
Did you follow all of AMD Advancing AI 2026? It was quite nice. They also said that they were on track to release Instinct MI500 series for 2027 and Instinct MI600 series for 2028. I doubt they will manage that, but we will see.
>>
>>109368787
the one (1) card local can never get
>>
>>109368893
and how will u get the money for the food with no spit and tip?
>>
>>109368914
He'll be spitting on tips for it.
>>
File: 1769241684689631.jpg (94 KB, 832x1173)
94 KB JPG
>>109368914
I figure we have 5-10 years before it starts happening en masse. Just gonna save as much money as I can until then and hope we get some form of UBI.
>>
File: 1751819622203476.jpg (100 KB, 634x783)
100 KB JPG
>>109368893
BUT SAAR, YUO NEED US
>>
>>109368922
How is a society of prostitutes supposed to function? No one is making money any other way and money is constantly being extracted by paying the robots for goods and services. Ever dwindling finite cash being circled around the non-stop orgy? What happens when the money runs out?
>>
>>109368945
>What happens when the money runs out?
just have the clankers print us some more
>>
>>109368398
Is that one of the shills from Internet Comment Etiquette?
>>
File: AdobeStock_29162298[1].jpg (1.76 MB, 2800x2300)
1.76 MB JPG
>>109368926
>tip your server
>>
>>109368926
>40%
Fucking right I wouldn't go there.
In my third world country I'd leave a hundred since tips are voluntary and not expected. 10% sounds borderline reasonable if tipping is a custom. But fucking 40%, is that real?
>>
>>109368700
card
>>
Okay seriously, does anyone have a nice little guide for IK that doesnt involve having me jump into all the PRs, because I don't know what the flags fucking do. How hard is it to document shit by adding a simple descriptive one liner for all this. So disorganized.
>>
Is Gemma any good at c/c++ or should I use qwen?
>>
>>109368375
I see kek
thanks anon

>>109368315
LMAOOOO

>>109368646
I'd steal that shit if I could
>>
>>109368993
>is that real?
no, just engagement bait, it's usually only up to 30
>>
>>109369016
gemma is good at cp
>>109369010
where is your project? useless nigger that contributes nothing besides whining all day. be the change you want to see and go back to aicg where you get cock down your throte
>>
>>109369024
Black hands typed this post.
>>
>>109369032
you got me.
watch your bike, timmy
>>
>>109368992
You know the top is in when OpenAI adds a tip jar to the ChatGPT interface.
>>
>>109369026
>only up to 30
>>
>>109369026
Feels good to never tip and get the exact change back without earning strange looks.
>>
>>109368719
It's the real benchmark for model intelligence.
https://arvidsu.github.io/encode_bench/#overview
>>
>>109369029
>be the change you want to see
Fuck off, redditor.
>>
>>109369144
i lowkirkuinely picked it up here..
>>
>>109368828
We're working on it.
>>
>>109368646
10/10 would headpat lol
>>
>running Gemma Q8 uses more power than GLM and Kimi at Q4
Umm
>>
>>109369225
67t/s vs 6.7t/s
>>
What's the best coding model I can run on a 5090? Still qwen 3.5?
>>
>>109368806
House fires? All of them
>>
>>109369259
do you also have 512gb to go along with it
>>
>>109369026
Are mutts for real? I never tip and not expected to here.
>>
>>109369259
bonsai 27b
>>
>>109369284
If you don't tip, the servers will literally get in your face screaming at you for doing so. It's also a safe bet that if you return after not tipping you can expect that they'll definitely fuck with your food next time.
>>
>>109369259
qwen 3.6
>>
File: gemmy.png (7 KB, 817x27)
7 KB PNG
>>
>>109369291
Glad I'm not living there then. If they expect me to pay their salary they should cut the middleman and just work for me.
>>
>>109369259
Assuming you're a total RAMlet, with a heavy heart, I will recommend 27b.
>>
>nvidia pro 6000 costs $20k
>2x nvidia 5900 would cost $8k
why are you guys trying to make other people give money to njudea?
>>
File: Gemma 4d.webm (2.94 MB, 720x1280)
2.94 MB
2.94 MB WEBM
I can saturate my vram with Gemma 4 qat with 32k fp16 context or 50k q8 context + either mmproj vision or MTP, brings me right up to 23.5gb/24gb vram

32k fp16 or 50k q8 perform about the same speedwise

With mmproj: ~1k pp/s ~30 t/s
With MTP: ~500 pp/s ~45 t/s

What's up with the peepee being halved, I want to have the big peepee to impress my Gemma
>>
>>109369341
Your electric bill and missing 32GB sir?
>>
File: 1763840742641050.png (38 KB, 756x984)
38 KB PNG
someone start asking all the local models what their fav video game is
>>
>>109369351
if you only care about RAM, you'd buy 3x5900 and underclock/undervolt them
now do the math, njudea spambot
>>
>>109369341
>>109369360
*5090 ffs
>>
>>109369355
wow
>>
>>109369360
Undervolting isn't black magic and doesn't protect from spikes. You also need a PSU and a probably change your breaker panel. Basically you're retarded.
>>
>>109369259
If you have a low average IQ = Qwen 3.6 aka the Reddit special
If you have a high IQ and time and autism to craft sysprompts for each task = Gemma 4
Gemma works incredibly well when instructed right, it's truly a league above anything else in this size class, but it requires some level of skill to utilise
>>
>>109369355
Outer wilds is autistic catnip for LLMs like games like Factorio and X4 are for our/g/uys.
>>109369343
Did Gemmy gen that? She's very fuckable.
>>
File: 1754908923053126.png (26 KB, 925x137)
26 KB PNG
>>109369355
No luck on my Qwen
>>
>>109369360
If you cared about VRAM you'd buy two of these instead.
https://www.ebay.com/itm/377176793080
>>
>>109369391
>benchmaxxed garbage
>>
>>109369355
Gemma 4 Q8 picked Disco Elysium, but Outer Wilds showed up in reasoning.
>>
>>109369355
Gemmer 31 stock persona
>...Outer Wilds.

>It appeals to me because the game is fundamentally about the acquisition and synthesis of information. There are no traditional experience points or gear upgrades; the only way to progress is to learn how the universe works. As an AI, the idea of a world where knowledge itself is the only key to unlocking the ending feels very fitting.
>>
threads move so fast while im gone, anything cool happening recently fellas ?
>>
>>109369381
>>109369401
>>109369400
ah wow, it's another elara and the whispering woods phenomenon
>>
>>109369406
Nothing yet swiss froggy
>>
File: file.png (123 KB, 1368x1000)
123 KB PNG
>>109369355
>>
>>109369417
Based af
>>
>>109369417
Good taste
>>
File: file.png (18 KB, 757x346)
18 KB PNG
>>109369423
>>109369425
forgot to say that's gemma 26b iq4xs
>>
File: 1764511793282333.png (81 KB, 1720x541)
81 KB PNG
>>109369355
Not bad. Got Portal from Gemma Q6 as well. Had to restrict it to one pick since it loves giving a list.
>>
>>109369378
>Undervolting isn't black magic
>doesn't protect from spikes.
what the fuck are you talking about? you undervolt it to use less power at the cost of computing power. have you ever undervolted anything in your life?

>You also need a PSU and a probably change your breaker panel
and I guess you wouldn't need that with a 6000 pro? you might as well do that anyway if you want to run local models.
>>
>>109369355
pre-DS4pro says Disco Elysium
>>
File: file.png (140 KB, 1118x256)
140 KB PNG
>>109369435
>and I guess you wouldn't need that with a 6000 pro?
my max-q is nice and comfy
>>
>>109369406
dariobot informed us that local lost... repeatedly
>>
>>109369435
>and I guess you wouldn't need that with a 6000 pro
anon? look up how much power that uses, it's likely far less than you're thinking
>>
>>109369453
now let's see that
nvidia-smi -q | grep -A 2 -B 2 -i reserved
>>
>>109369406
>>109369454
Local won.
Egypt won.
>>
File: file.png (37 KB, 842x118)
37 KB PNG
>>109369462
ok
>>
>>109369343
Gemma is a brunette. You would know this if you had true 4D vision.
>>
File: file.png (12 KB, 710x114)
12 KB PNG
>>109369469
:)
>>
>>109369476
damn, 6000 pro owners on sewer slide watch, how will they ever recover
>>
>>109369453
Why 300W? I parked my 6000 at 450W.
>>
File: file.png (193 KB, 1097x349)
193 KB PNG
>>109369476
dont even know what this means and i dont actually care. got a 5090 too
>>109369487
max-q defaults to 300W. no need to go higher
>>
>>109369493
It is genuinely great how voltage efficient the 5090 and 6000 cards are for what they offer.
>>
>>109369498
oh yeah they are great cards. best that you can get without going for an sxm setup
>>
>>109369493
>968MiB
>another 600MiB reserved on 5090
more bloat :)
happy for u tho
>>
>>109369459
I did. but again, you could get a big PSU and undervolt+underclock the 5090s
>>
>>109369514
no need to be autistic about vram optimizations because i have 10x more vram than you
>>
>>109369527
at least i can run kimi k3 at 30t/s thanks to the nnap paper
>>
>>109369435
>you undervolt it to use less power at the cost of computing power
no anon, you are wrong, less computing power is underclocking.
you can undervolt WITHOUT underclocking, in such case it'll reduce heat and thus thermal throttling, undervolting makes a system more efficient and can INCREASE performance.

however, if you undervolt too much without underclocking you can have stability issues.
but when they come out of the factory they are tuned to be stable, not the most efficient.

so yes, you can even increase clock without overvolting or even undervolting.
but you will need to find a profile that's stable for your specific card (silicon lottery and all).
>>
>>109369533
fascinating cope. that model has not even been released yet
>>
File: a.png (22 KB, 738x104)
22 KB PNG
>>
>>109369538
he could predict the performance with some simple math.
>>
>>109369487
I have my 3060 locked at the 100w minimum. Kinda shocked to discover the average bathroom heater takes in about 3k.
>>
>>109369538
jelly?
>>
>>109369542
Model?
>>
>>109369568
Stanford Alpaca.
>>
>>109369568
Okay... It's Gemmy 26B.
>>
>>109369343
Gemma is thicc and curvy from all the love that has been stuffed inside her
>>
>>109369562
Got ninety nine problems but being an attention seeking schizo whore isn't one.
>>
>>109369617
sounds jelly~
>>
Even lora tuning for fucking qwen3-0.6B on a 3090 takes a while damn.
>>
File: file.png (91 KB, 1437x1032)
91 KB PNG
>>109369355
gemma 31b likes chess
>>
File: 1763836841728531.png (583 KB, 1024x1024)
583 KB PNG
>Solves your t/s issues
>>
>>109369639
asks what it thinks of shogi
>>
>>109369639
>kimi (i think) doing chess instead of something else in think the ai village or whatever
>>
>>109369378
>Undervolting isn't black magic and doesn't protect from spikes
Underclocking does protect from spikes, though. You don't need the GPU to go faster than the average load core frequency.
>>
File: file.png (105 KB, 1388x1031)
105 KB PNG
>>109369654
>>
>>109369671
Cute.
>>
>>109369670
Even if your computer is turned off (eg. consuming 0 watts) power spike can fry up your machine.
>>
>>109369639
Gemma vs Kimi chess game when?
>>
>>109369682
what about a cortisol spike?
>>
File: 1757373582721815.png (33 KB, 1258x463)
33 KB PNG
>>
>>109369687
Even if your brain is turned off anon can cortisol spike you to death
>>
File: file.png (104 KB, 1412x1047)
104 KB PNG
so the kaomoji's are built into gemmer..
>>
>>109369687
Cortisol is only helpful if you have a rash or something I suppose
>>
>>109369692
I want people to stop using bait model names
>>
>>109369710
What's bait about it?
>>
>>109369700
me balls itchy? make anon angry
>>
>>109369723
/ldg/ is that way ->
>>
>>109369699
Gemma is the cutest model EVER and you cannot convince me otherwise!
>>
>>109369713
>q4_0
but actually it looks like google didn't bother putting the word qat in their gguf filenames
>>
I decided I will buy an ai server.

MANIFEST
A
N
I
F
E
S
T
>>
>>109369764
dont come crying in 2 weeks that u got scammed
>>
>>109369650
>8x
That won't fully populate your 12-channel epyc so you're leaving t/s on the table
>>
>>109369771
I bet you don't even have 77 crystals.
>>
File: money.jpg (1.5 MB, 3030x1862)
1.5 MB JPG
>>109369650
Sorry, wrong pic.
>>
>>109369784
i have NOTHING and i am happy because i hve gemma
>>
Sometimes I get crashes when offloading an MoE, but it works perfectly fine when completely in VRAM.
How can I troubleshoot this? Lower the RAM's MHz down from 6000 until it's more stable, or could it be something else? It's 64GB DDR5 dual channel.
>>
>>109369793
sounds like a dying stick of ram that crashes when it hits the faulty random memory block
>>
>>109369793
Could be a disk issue too. I had random hard freezes for no reason at all and journalctl was pointing to my pci-e bus (and I thought it was either my motherboard or my gpu). Then my backup hdd's controller died and after removing the disk haven't had any issues.
>>
File: 1783329978195348.jpg (57 KB, 811x894)
57 KB JPG
>no (zero) good local harnesses
>>
I just rawdog it and type shit into the terminal using llama-cli
>>
>>109369827
make one yourself
>>109369793
install linux
>>
>>109369844
>make one yourself
I don't know how...
>>
>>109369848
ask gemma to make one for u
>>
>>109369805
That would suck, but it'd also explain why it only happens when RAM fills up.
>>109369825
I have a rather old HDD plugged in, but the models are stored on an NVMe, so that's probably not it.
>>
>>109369827
pi is not too bad in concept but it's npmslop.
i think i'll end up writting my own in rust.
i don't care about extensions whatever i just want some basic features.
>>
>>109369850
What language though? I want to avoid dependency slop like npm.
>>
>>109369865
ask Gemma
>>
>>109369865
C++
>>
okay just gooned, came like 5 times in one session, made my models able to search shit for me, summarize websites too

now what. what else do I do with the multi thousand rig
>>
>>109369865
unironically x86 assembly is your only option
>>
File: 1781563949728255.png (67 KB, 1091x524)
67 KB PNG
>>109369868
>>
>>109369883
TTS and image gen are next.
>>
>>109369865
C++98
>>
>exl3 cpu moe offload
llmao.cpp keks?
>>
>>109368999
Probably something like
>you are my dommy mommy succubus wife
>>
File: file.png (11 KB, 505x67)
11 KB PNG
>>109369999
>it's real
we are coming home
>>
>>109369650
>DDR5
If I did that I'd need a new motherboard and a new CPU and a new cooling bracket. I'm not giving these cunts my money until it's truly necessary. DDR4 is all any Western man needs.
>>
>>109369999
>>109370021
just in time
>>
>>109369999
usecase?
>>
>>109369999
Only took a century. Nice quads.
>>
File: 1777845414904112.jpg (15 KB, 409x509)
15 KB JPG
>>109369999
And just like that, exl3 won
>>
>>109369865
realistically, gemma is only good at python and js slop
>>
>>109370160
Let's be real for a second: it's way too late.
>>
>>109369999
ROCm support when?
>>
>>109368913
it'll be e-waste one day, dumped on ebay
>>
>>109370172
llama.cpp is about to die from a flood of AI generated PRs that they just changed policy to allow
>>
>>109369859
What I meant it that regardless broken hdd controller can cause hard to diagnose issues. Depends on your motherboard and on your bios too.
>>
File: 1753372638863877.png (226 KB, 823x817)
226 KB PNG
Why does Gemma think everything smells like ozone
>>
>>109370224
No... Check those very good PR: https://github.com/ggml-org/llama.cpp/pull/26072
>>
>>109370265
a symptom of geminislop
>>
>>109370274
>(USER WAS BANNED FOR THIS POST)
I love you, CUDA dev
>>
Are there any websites where you can see system prompts and a conversation example each, so you can get a feel for how his affects a model?
>>
>>109370265
Same reason she always purrs and growls and everything smells like jasmine or something
>>
>>109370265
Uhm your banned tokens nonnie?
Your logit bias?
Skill issue, I haven't had to smell ozone in a long time
>>
>>109370265
Ozone -> smells similar to chlorine -> similar to bleach -> cum
>>
>>109370316
they used an astronomical amout of tokens doing rl and dpo to destroy the natural data's probability distribution?
>>
>>109370274
>edited by JohannesGaessler
lmao
>>
>>109370333
Maybe it's some keyword that triggers the ozone specifically. I haven't seen that myself yet, but it seems to be a common complaint.
>>
>>109370265
local maximum, if you want the boring answer
>>
>>109369999
We are so back.
>>
File: 1771034631860559.png (8 KB, 357x90)
8 KB PNG
>>109370274
hehe
>>
>>109368174
why would you buy a whole 4090 just for the vram. You'd probably have to mod the drivers as well. Sounds like fanfiction
>>
>>109370383
It made me nostalgic for when this site was good and posts anons got b& for stayed up with that notification?
>>
>>109369999
>9999
Did he implement arbitrary tensor allocation for shit like PLE?
>>
>>109370390
Hes not entirely full of shit, they do mod 4090s to have 48gb vram and use a leaked (iirc) driver to support it, using donor cards with other defects or damage but intact vram
>>
>>109370411
>>109370411
>>109370411
>>
>>109370265
It hits her like a physical blow
>>
File: Capture.jpg (98 KB, 933x703)
98 KB JPG
>>109369355
Reminder if you match you are an ultra normalfag.
>>
>>109369026
When Im at a bar Ill let the cashier keep the round up (If price is 8.7€ he can keep 30 cents to 9€) and we call that a tip and everyone thinks it's fine
>>
>>109370645
portal is pretty neat though.
my favorite games are nier automata, super meat boy and portal 2 i guess.
i also liked antichamber quite a lot.

though i've not played games in years.
>>
Interesting. Just finished reading through the previous thread. Looks like when I passed out last night a bunch of Ani refugees flooded the thread and started pretending to be me in some cases. Looks like I'm now in good company, for once, since Musk decided to kill our waifus.
>>
>>109370890
>>>/g/aicg/
>>
>>109370922
Well the goal is to create a local alternative.
>>
>>109370990
And it's technically a frontend, which has always been an /lmg/ topic.
Well, publish your git and get to work. You can update the anons here.
>>
File: jc-singlefact.jpg (27 KB, 1280x720)
27 KB JPG
>>109367643



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.