[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1767514858305886.jpg (270 KB, 1415x2048)
270 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109378862 & >>109374404

►News
>(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908
>(07/23) LLaDA2.2-flash agent-oriented diffusion model released: https://hf.co/inclusionAI/LLaDA2.2-flash
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B
>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: threadrecap.png (1.48 MB, 1536x1536)
1.48 MB PNG
►Recent Highlights from the Previous Thread: >>109378862

--Comparing SSDmaxxing and CPUmaxxing for running large models like Kimi-K3:
>109379081 >109379314 >109379771 >109379850 >109379914 >109379926 >109380112 >109380120 >109380177 >109379992
--Testing Laguna S 2.1's uncensored nature, coding, and persona capabilities:
>109380839 >109380869 >109380893 >109380932 >109380954 >109380960 >109380967 >109380998 >109381002 >109381029 >109381063 >109381080 >109381102 >109381133
--Debate on LLM architectural limitations and the possibility of AGI:
>109381207 >109381235 >109381717 >109381816 >109381858 >109381872 >109381875
--Evaluating the value of RTX 3080 20GB versus RTX 3090:
>109380874 >109380927 >109380983 >109381027 >109381062 >109381068
--Debating data and RL importance over architectural improvements:
>109380227 >109380234 >109380254 >109380261 >109380373 >109380415
--Bartowski updating Gemma GGUFs for chat template and tool-calling fixes:
>109379231 >109379250 >109379263 >109379344 >109379381 >109379390
--Optimizing a mixed GPU build for MoE and subagent setups:
>109380050 >109380064 >109380196 >109380247
--Comparison of Gemma 31B quant quality and QAT performance:
>109379105 >109379408 >109379440 >109379454
--Comparing quality and hardware trade-offs between Q4 and Q5 quants:
>109379173 >109379182 >109379201 >109379247 >109379269
--Feasibility and timeline of llama.cpp support for Kimi k3:
>109379938 >109379943 >109380016 >109379949 >109379962 >109379993
--Coping with model size creep via SSDmaxxing and compression:
>109379801 >109379816 >109379892 >109379894 >109380090
--GLM 5.2 benchmarks on a 4x Spark ring topology cluster:
>109380810
--Logs:
>109378896 >109380042 >109380253 >109380893
--Miku, Tethou (free space):
>109381284 >109381511 >109380533 >109380998

►Recent Highlight Posts from the Previous Thread: >>109378865

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
the last ever /lmg/ thread before k3 releases and we finally rebrand to /omg/
>>
>>109382150
>/omg/
why o ?
>>
>>109382150
It's honestly about time. Almost all of the real discussion has moved to open-weight models because that's obviously where the future lies. Everything else is just retards talking about Gemma and Qwen nonstop. They're old. Give it up already.
>>
>>109382150
obsolete model general?
>>
>>109382165
OnlyModels General
>>
File: file.png (391 KB, 728x455)
391 KB PNG
so what's the verdict on laguna S 2.1?
>>
>>109382169
all models can be run localy, the only question is at what speed.
>>
>>109382173
Best model since Mythomax.
>>
>SSD maxxers will need to upgrade from 1 / 2 TB SSD to 4 TB to run kimi
>>
>>109382136
https://litter.catbox.moe/3rtqt4qnp77xuclu.mp4
>>
Is google really the only one making good dense models? Small MoEs write like shit, V4 flash is the only MoE with good writing that I tried but that one's not local
>>
>>109382189
Always has been
>>
>>109382186
This is the only normal looking Miku model I've ever seen and I have no idea where it's sourced from. Plz help.
>>
>>109382032
>>109382199
>>
>>109382189
Yes I recently went through my old llama2 era logs and it's crazy how much depth they had despite every reply being 1/5th the length of what I get these days. The new models are smarter but they lack this sense of completeness.
I suspect that it's the compartmentalized and castrated model size in active parameters causing j-space related shrinkage.
>>
Someone I know is giving away a 6GB 3050. Can I pair this with my 5070?
>>
>>109382214
Take it and if it doesn't just sell it.
>>
It's simple — we uh... Kill /lmg/.
>>
>>109382173
chink retard who yaps too much
>>
>>109382205
quantizing kv shits up every single ai sloppa i ever tried, llm, vision, audiogen, audio diarization, everything
>>
>>109382205
KV cache quantization should be the last measure to save vram. It's the fastest way to degrade your already quantized model into a schizo mess as errors compound with the context length.
>>
>>109382185
I have an unallocated unused 2tb NVME I planned to use for dualbooting but never got around to it. Im new and have only used my GPU and cpu/ram offload with moe. how does one SSDmaxx and run a big model off of storage?
>>
>>109382265
Forget about it
>>
what's anon's current hardware setup and how long do you think you'll stay relevant in this hobby?
>>
File: 1777268455365.png (181 KB, 500x316)
181 KB PNG
>Mfw giving Gemma any kind of a possessive rapey character to play

This is her core personality and she embraces it to the fullest every time.
She delights on the slow corruption of the user.
>>
>>109382189
>but that one's not local
my 3090+128gb ram disagree
>>
>>109382241
>>109382233
Yet hardly anyone suggested to the Gemma guy that they should fix that shit in Gemma5. They all bent over and asked to be cucked by SWE bench.
>>
>>109382274
just under 1tb ram+vram combined
it stops being relevant in about three hours
>>
Egypt won.
>>
>>109382273
so just mmap?
>>
>>109382265
you can't do much with one drive, it's gonna be 0.1-0.5t/s etc.
the problem with ssdmaxxing is even in the perfect gen5 scenario with four drives, and that fancy raid card, to the best of my understanding even then it's single digit everything. not just decode but prompt processing too. it's gonna be super slow basically. for coding the only possibility is run something overnight and maybe it completes by morning, or maybe on the following morning.
>>
>>109382265
ssdmaxxing is a meme where you switch from tokens per second to seconds per token, I'm just making a joke about how fucking yuge kimi 3 is in that a single 2tb ssd might not be enough to hold the whole model
actually using ssd for inference you'd have to do some silly setup like four high end gen5 ssd, and even then it would be extremely slow - just try running a 50b dense on ~60 GB/s, that's ~0.4 tokens per second. Having a mobo with four pcie5 SSD is also rather unusual. If you try to go at it with four gen4 2TB SSD that's ~0.2 tokens per second at optimal speed.

>tfw the joke didn't even work because it's 4bit native so the weights only take 1.4TB
>>
>>109382274

5090 + 64gb of RAM.
Considering how model intelligence is only going to keep on improving even on smaller sizes, I'm sure this will stay relevant for quite some time.
Also new solutions are going to come out which lead to smarter smaller models etc..

However while this setup is pretty good it's not quite where I want it.
I want to throw in an additional GPU with an extra 24gb of VRAM, so I don't have to sacrifice quant size for context with models like Gemma.
I also intend on getting at least 128gb of RAM when I upgrade to Zen 6, probably going all the way to 256gb just so it opens more options.
>>
>>109382293
I never understood the appeal of hypnosis/rape before LLMs.
Now basically all of my RP sessions have characters using me "against my will".
>>
>>109382334
>>109382350
Ah i see, shame.. guess I have to just stick wtih gemma
>>
>>109382355
cursed setup unless you're a dense gemmy enjoyer, there's a dearth of models that would fit in 96gb memory outside of dogshit dynamic quants. nobody targets this, it's either ~30b toys or ~120b
if you can afford 5090 you should upgrade to 128gb ram and laugh at poors
>>
>>109382293
>>109382362
Why does ERPing invariably turn people into subs.
>>
>>109382350
I'm not sure it's quite as bad in that perfect scenario, wasn't it supposed to be like 1-3 t/s or something?
>>
>>109382331
You need imagination + creativity + high IQ + white skin to be able to do it
In all honesty, I have done it maybe once or twice in the veeeeeery early days when I didn't know how LLMs work, didn't think about prose at all and wasn't exposed enough to notice the slop. One of the best faps of my life. Ignorance is bliss
>>
>>109382384
yeah when I first discovered llms I basically wacked it for a week straight.
>>
>>109382384
>white skin
how is that a requirement lol
>>
>c-china won't release their best ai
Kek these news outlet jews are shameless
>>
>>109382373
>>109382355
btw we're not running a 50d bense here, kimi might be supposedly around 50B active, but if it's the same or similar arch as before, then remember it's like a 25B active normal model because of the int4 native
>>
>>109382373
>wasn't it supposed to be like 1-3 t/s or something?
50b at q4 is 25 GB
assuming a semi-reasonable setup of 4 gen4 ssd (you're unlikely to squeeze more m2 slots on a consoomer mobo) you're topping out at just about one token per second. I'm not sure how kv cache comes into effect I'm not desperate enough to ssd maxx yet
>>
>>109382331
if you can imagine it anything is possible
>>
>>109382384
>I have done it maybe once or twice
>not every day
You clearly aren't imaginative/creative/smart/white enough to post on /lmg/ then
>>
File: 1758552443152979.png (77 KB, 923x383)
77 KB PNG
>saves your hobby
>>
>>109382368

I agree it's a bit cursed territory having just 64gb of RAM, because it locks me out of the +100b realm.
Problem is that I'm still on Zen 3 and stuck with DDR4 and all of my RAM slots filled up with 16gb modules, which I sure as fuck won't spend money on upgrading at this point, because Zen 6 launch is next year.
If I was on DDR5 I'd absolutely be at 256gb of memory.
>>
>>109382371
I have no idea. 95% of my logs are loli rape
>>109382331
the agency in the story and it being completely untouched and unknown by another human's sleazy hands is what makes it good. Drawings are just pixels on the screen and porn is just watching other people having sex.
>>
>>109382441
>because Zen 6 launch is next year
blud gonna get gaped
it took years before ddr5 was better than ddr4, the timings were shit, early adopter tax was yuge and now we're in the middle of unprecedented demand for memory
you ain't getting shit next year
(admittedly, switching to am5 atm isn't great either because prices be crazy, but ddr6 next year is a pipedream unless you pay out the ass)
>>
>>109382435
I kneel
>>
>>109382371
Quite the reverse for me. They're the perfect test subjects whenever I have a thought experiment to run
>>
>>109382435
thats why, not to sell more GPUs or anything like that
>>
>>109382399
Sorry rajeesh
>>
>>109382458

But Zen 6 isn't on DDR6, it'll be on AM5 and DDR5.
Main reason why I'm still on Zen 3 is because I didn't want to be early adopter on DDR5 and stayed away from it.
And having a 5950x as a CPU didn't really warrant an upgrade as this thing is god tier and still perfectly relevant.
Zen 6 will finally offer a genuine upgrade though.
>>
>>109382441
Dumbo just upgrade with 32 gb ddr4 sticks and overclock them. It won't get any better in the next year. Getting 128gb on my mobo was the best decision I made so far
>>
>>109382471
80% of compute demand for western AI data centers comes from OAI and Anthropic.
>>
>one of the x70 motherboards is likely capable of ~35 GB/s if you fill all non-gpu slots
>that's roughly 1.33 tk/s, likely about 1-1.2 tokens per second in actual speed
are you a bad enough dude to do a nala test on kimi 3 locally?
just, uh, remember to disable thinking or you'll need to go on a long walk before it finishes
>>
>>109382495
ah fair enough I am retarded and somehow mistook zen6 for ddr6
I'm probably switching from 7800x3d to zen6 as well
>>
>>109382407
I was talking about that raid card for $2000, highpoint rocket or whatever, gen 5, basically the "perfect case" scenario
>>
>>109382512
It's fine, he just have to go buy milk
>>
>>109382371
Actually, I lied.
Even before LLMs I was playing games like Corruption of Champions and Monmusu Quest, now I just have an endless supply of crack.
>>
>>109382371
If i could have an actual living real IRL 3d qt3.14 foid bully me for free with no strings attached I would, but gemma does a great job standing in
>>
>3x PCIE5x16, each one bifurcated 4x4
>4x NVME Gen 5 on every PCIE slot, all in RAID config
Why wouldn't this work again
>>
File: 1753848948331325.png (172 KB, 1134x608)
172 KB PNG
Brothers, stop memeing the SSDmaxx snakeoil. Here's a summary of why it's not a good idea. If you don't trust it, do your own research and you'll come around.
>>
File: 1758054541921208.png (119 KB, 535x560)
119 KB PNG
Less than 3 hours. Are your hardware ready?
>>
File: IMG_20260727_034535.jpg (193 KB, 773x1594)
193 KB JPG
I have achieved AGI internally
>>
fp8 is the lowest quant that's the real model, anything lower is cope
>>
>>109382435
If you can't beat them, lead them
>>
>>109382578
my GPU is in the way, fml
>>
>>109382512
Not how it works, at least on amd. Sth like mag x670e has dedicated gen5 slot at full 14 GB/s speed, then like four gen4 slots but despite not sharing lanes they can’t go faster than 7.5 GB/s total due to the controller chip from which the traces originate. Best case scenario you get 14 + 7 GB/s, but you only need one gen5 and one gen4 ssd.
You coukd squeeze in a second gen5 ssd in main gpu slot and have the gpu itself sit in the gen4x4 slot but I’m not sure if that’s a net gain, you still need to do prompt processing and keep kv cache somewhere. Maybe gen5 slot in dedicated slot, second in gou slot and the gpu in second slot without any other ssd connected so you get usable gpu speed and 28 GB/s out of two ssd

alternatively stop being poor and buy a 1.5TB ram desktop, like what else are you goi g to blow your money on? your nonexistent gf? some onlyfans whore? some gay car when you never leave your house?
>>
>>109382611
that's simply false and also depends of your quant scheme.

with proper quants you won't have any meaningful loss bellow Q5 or Q6.
exl3 allows you to go a bit lower than that and still being pm lossless.
>>
>>109382441
I have the same setup except with 64 GB DDR5. It really is a weird spot where I can try some MoE at very low quants, but it's never worth it and I keep coming back to Gemma.
>>
>>109382601
>I have achieved AGI internally
>*(analy)
>>
>>109382595
Can't run her but I'll downloader her to my spare hdd. Might need to delete k2.7 though.
>>
File: 1784771240402402.png (87 KB, 370x320)
87 KB PNG
>>
I hate Kimi, if they hadn't dragged out their release like this new Dipsy would be out already
>>
Why haven’t they made a 30b3a moe with the same performance as kimi 3? Are they stupid?
>>
kimi k3.gguf?
>>
>>109382593
>real world practitioners ... report
No they don't. I've searched. Nobody has and published running RAID0 + llama.cpp ssdmaxxing.
>>
Meme size. 27b active is the bare minimum to be worth anything these days.
>>
I wonder if at some point we could start stashing models entirely into the CPU cache.
Zen 6 is already hitting +1GB cache with the EPYC.
Speeds would be absolutely insane with zero overhead or latency.
>>
>you’d need 768gb ram setup to run at lobotomized q2
accept thy limitations, anon. glm 5.2 exists, deepsneed flash exists
>>
>>109382661
@kimi-chan, 109382661 is a perfect retard candidate!
>>
>>109382664
>>109382681
>>
>>109382682
praying
>>
>>109382679
Retard, just ask normally instead of baiting me. https://old.reddit.com/r/LocalLLaMA/comments/1ikprg7/trouble_with_running_llamacpp_with_deepseekr1_on/
>>
>>109382578
>Why wouldn't this work again
Is there a CUDA_VISIBLE_DEVICES='' equivilent for DDR5?
Like DDR5_VISIBLE=32GB ?
It would be good to test if this shit actually works with an existing model but I'm not touching my RDIMM sticks at all now since they're worth more than my car.
>>
>>109382507
They have already hit power grid cap, the future is datacenters in every home
>>
>>109382635
there's the lower cope
>>
>>109382578
imo this might work and give you a few t/s. there's only one way to find out brother
>>
File: leddit.png (84 KB, 1073x737)
84 KB PNG
>>109382704
I haven't been able to view Reddit for weeks
>>
>>109382717
sorry if you can't understand facts.
we literaly have data on it.
Q6 is still basicaly lossless.
Q5 has *minor* loss.
bellow Q5 it becomes retarded.
>>
>>109382736
your brain is Q4 stop replying to me tard
>>
>>109382707
mem=32G in kernel parameters should work, but you need to reboot.

https://mjmwired.net/kernel/Documentation/kernel-parameters.txt#2244
>>
>>109382707
if you are on linux, cgroups will do the trick
>>
>>109382639

Yeah there's a massive chasm between Gemma and everything else. She's basically as good as it gets unless we have 128GB of memory and even that's a bit of a cope quant territory.
256GB + 5090 is practically the gold standard for a consumer and I think it's worth going for if you're even half serious about playing with AI.
It opens up the landscape considerably and you're still moving in a realistically doable price point for anyone living in a first world nation.
>>
>>109382578
You can start with llama.cpp and prompt Kimi to write an inference engine for your machine and run it overnight
>>
File: cope.png (288 KB, 749x1225)
288 KB PNG
>>109382736
Q8 is only 92% vs BasedFloat16
>>
>>109382736
Nuh uh
source: I want it to be true
>>
For future SSDmaxxers, here is the current state of the art:
https://github.com/randyap8-wq/Micro-Expert-Router-SSD-Streamed-MoE-MER
>>
Backup K3, just in case there's a compression breakthrough in the near future or the CCP rugpull within a week. Don't say we didn't warn you.
>>
>>109382744
The air was thick with the scent of dariobot
>>
I don't mind the occasional bit of slop from gemma. It reminds me I'm pathetic and leaking precum to linear algebra and not talking to anyone.
>>
>>109382758
How is xi gonna rugpull? Once the weights are available it’s no longer possible like with a closed model. US cucks -> dl from chyna, Chyna cucks -> dl from fuggingface, everybody cucks -> torrent
>>
>>109382763
*arches her back*
>>
>>109382681
I’d say 13b active ever since trying out deepseek v4 flash
>>
>>109382757
>rust
>>
>>109382595
Egypt won.
>>
>>109382744
256GB is still pretty expensive and it would slow down token gen by a lot. At that point I think it might be better to save more to get a 6000 pro.
>>
>>109382779
>deepseek v4 flash
wheres the fackin 4.1
>>
>>109382775
US data centers won't be allowed to host K3, blocking it from normalfags. It won't stop local schizos but in reality there will most like be <500 people on the planet actually running K3 ~Q4_0. The rest will be enterprise on their own infrastructure.
>>
>>109382793
Q2 will be just as good and only require a bit more memory than K2 at Q4
>>
>>109382758
This isn't Drumpf we're talking about.
>>
>>109382754
>Q8 is only 92%
you do not understand what it means.
KL and intelligence are different things, the output being different doesn't mean it's necessarily dumber.
>>
>>109382800
I'm probably looking at Q3 here and that already feels like I'm fucked
>>
>>109382787
Yea the actual consoomer stack is 128gb ram + whatever cpu. 12gb vram vs 32gb vram doesn’t make that much difference, surprisingly. And with current gen cpus you’re only using 2 ddr5 slots otherwise speed craters
If you want more you either go 96gb vram + 128gb ram or a 256gb unified memory from applel and a lot of patience. realistically you just buy inference from a third party and hope things get better (they won’t)
>>
>>109382758
>nigga never heard of streisand effect
>>
>>109382758
>in case there's a compression breakthrough
TANSTAFL
Even if someone managed a sub-Q1 extreme compression it would still be so stupid that using a smaller model without the compression breakthrough would be preferable.
What we need is some way to make prompt processing speeds bearable with offloading.
>>
>>109382808
What does it mean? I’ve been wondering about it myself
Is q8 being "92% good" > 92% correct response vs bf16? idgi
most charts I saw q5 was getting pretty close to unquantized. Existence if q4 native would also be questionable if they were as lobotomized vs bf16
>>
man I should really learn cuda and make llama.cpp fast
>>
>>109382832
just use exllamav3
>>
>>109382823
The model is probably undertrained enough that quantizing the weights to binary precision won't cause a huge performance loss. They will still be ~400GB, though.
>>
>>109382832
>makes actual improvement
>ik guy comes in accusing you of stealing an idea he drew on a napkin 17 years ago
>>
>>109382832
or just open claude, say "learn cuda and make llama.cpp fast" and wait a few hours
>>
>>109382823
>Even if someone managed a sub-Q1 extreme compression it would still be so stupid that using a smaller model without the compression breakthrough would be preferable.
idk, on my old poorfag build with 64gb ram i greatly preferred ~120b q3 than whatever ~35b scraps I could run at q6 or whatever
>>
>>109382787

The cost difference is pretty staggering though.
It's like what, 4-5 grand for 256GB of memory and 5090 is around the same ballpark, but 6000 is already reaching to 15 grand and that's still just 96GB.
If we're talking about what's viable for the average working person then a 5090 and 256GB is about as good as it gets without completely breaking the bank.

>>109382810

While the VRAM amount doesn't necessarily make that much of a difference in some tasks, the GPU speed absolutely shows up in this regard.
5090 has a retarded high bandwidth at 1,792 GB/s and the architectural improvements are a thing too.
That makes it an entirely different beast in general AI use, including image and video generation etc...
>>
>>109382857
like most things, the marginal return for increasing precision goes down as the base value is higher. going from 6 to 3 is not even going to be 10% as damaging as going from 3 to 1
>>
>>109382808
>the output being different doesn't mean it's necessarily dumber.
we can see the trend tho, its obvious it gets dumber as you quant it and the percentage follow the same trend, are you asserting that correlation doesn't equate causation or are you just ignoring the data.
>>
>>109382856
yea i dont know why people dont do this
i had fable craft me a 1bit quant with zero precision loss in a single 5h session. it also cured my baldness and was halfway through adding 2" to my cock but I ran out of weekly quota
>>
File: ssdmaxx.png (30 KB, 717x262)
30 KB PNG
K3 torrent when?
>>
>>109382823
>make prompt processing speeds bearable with offloading
DMA is the way
>>
>>109382869
so did you get an inch and then it stopped or was it going to be 2 at once so you got nothing?
>>
>>109382881
It added 2 inches, then before of running out of quota deleted them.
>>
>>109382828
>>109382868
he won't explain it as he doesn't know.

it's divergence, how far it strays from the original, 90% is already poor, but it doesn't mean it's necessarily worse, just different.

it's always worse though, when you're down in the gutter slumming a Q4 you shouldn't even be allowed to call it by the original model's name.
>>
File: 1773611191977750.jpg (277 KB, 1920x1080)
277 KB JPG
GEMMAMA-5-70B DENSE NOW
>>
>>109382881
nah there’s just nothing there until next week. can’t even piss but I’m not oaying api prices
>>
>>109382893
I'm so lonely.
>>
>>109382893
I will die from intense cum loss
>>
File: cheers.png (278 KB, 497x497)
278 KB PNG
>>109382873
and there we have it boys
>>
File: 32G.png (12 KB, 690x143)
12 KB PNG
Okay, this is fucked lol
I limited it to 32G
0.6-0.7t/s with minimax-m3 Q3

Device             tps    kB_read/s    kB_wrtn/s    kB_dscd/s    kB_read    kB_wrtn    kB_dscd
nvme0n1 68.01 15222.26 2987.63 524.61 18988087032 3726742713 654388320

free -h
total used free shared buff/cache available
Mem: 251Gi 9.1Gi 215Gi 82Mi 29Gi 242Gi
Swap: 0B 0B 0B
>>
>>109382808
No it's dumber. We're not in llama1 era anymore, these models are overcooked to hell with a distribution where the top token sits at 99%. So you bet if the top-1 % isn't matching the BF16 you'll get out of distribution garbage really fast.
>>
>70b dense
why would you want this? It’s only useful for rtx 6000 royalty or 3090 stacking peasants
hook me up with dem 80-120b moe
>>
>>109382873
>1TB
uhhh anon…?
>>
>>109382873
>1 TB
ngmi
>>
>>109382913
>>109382914
b-but he got heckin 4 of them!
>>
>>109382909
you can run it on a single 3090 at 2bpw
>>
>>109382909
>why would you want this? It’s only useful for rtx 6000 royalty or 3090 stacking peasants
because I'm one of those two?
>>
>>109382902
what is it that we're looking at here again
>>
70b dense would unironically be too powerful for Google to financially justify. It would destroy almost everything and the go-to local for startups and enterprise.
>>
File: 1768519903401556.png (79 KB, 1087x478)
79 KB PNG
Something happened! I think!
>>
>>109382873
Real men don't even store the model on disk, instead streaming weights on-demand via p2p
>>
>>109382937
Whoa!!! That's the motherfucking genius scientist Illya Surskit himself!
>>
>>109382937
Straight shota will save superintelligence
>>
>>109382942
>streaming weights on-demand via p2p
In fact forget the weights, just give me tokens
>>
>>109382931
Google has let Gemini Pro stagnate for whatever reason since the start of the year. They are now playing catchup, which places them in the same camp as the Meta used to be and China is now, that is it's to their benefit to tear down SOTA. Combined with the fact that AI is little more than a feature for Android and search summarizer rather than their whole business model, Google would have every reason to do release it anyway.
>>
>>109382909
>>109382921
Soon, brothers... my paycheck should be here any day now...
>>
>>109382909
Bonsai
>>
File: 1777266931973830.png (81 KB, 944x416)
81 KB PNG
The thing that is happening is continue to happen! AAAAAAAAAAAA
>>
So did anyone end up tweeting the 70b dense requests at the gemma team or was their feedback entirely about making a coding-only model like qwen?
>>
File: aa_model-size-comparison.png (925 KB, 2707x1292)
925 KB PNG
>>109382893
Not coming. You might get a ~180B Flash-Lite and at most a ~500B Flash -equivalent **MoE** versions of Gemma.
>>
>>109382966
>We are
damn their model already hit the context limit
>>
>>109382593
Your clanker is wrong. Although you can't prefetch experts they can't be considered random reads
>small, discontinuous, random reads
MoE experts are huge compared to a file system block.

If RAID 0 for reason can't handle it you could always duplicate the data and spread the reads across drives.

Aside from caching maybe you could make a good guess about the experts needed in the next layer based on the current active ones using a small predictor.
>>
Does anyone here actually run V100 32GB cards? I'd like to give one a try to replace my 2080ti 22GB, which is not really enough for it's purpose which is to provide qwen-tts, faster-whisper, and a small comfy instance. The extra memory would help, but is V100 weird enough to cause me significant grief?
>>
>>109382136
What are the best LOCAL AI model to translate manga pages in real time ?
>>
>>109382969
I can’t think of a less interesting release than low parameter code sloppa, those are completely useless in practice compared to paypig models and devs can afford a good shit setup anyway
which leaves… i dont even know who as the target audience for 26b-codenigger-edition-pinky-promise-we-can-oneshit-pacman
>>
>>109382975
stream the whole layer and let the gpu handle the sparsity, you only need a couple buffers to make it happen,
>>
>>109382982
what's your VRAM? RAM?
>>
File: 1771087786944370.gif (2.62 MB, 360x259)
2.62 MB GIF
Is VRAM are the only that matter in AI ? For example if i go to Nvidia HQ and stole experimental GTX 1080titi with 64 of VRAM, will it perform better than RTX 5070ti with 16gb of VRAM ??
>>
>>109382185
Do people actually ssdmax models or is this just a meme and people just go and build shitty optane boxes instead?
>>
>>109382978
Don’t bother with weird shit gpus unless you really know what you’re doing. You’ll probably get it working eventually but the time investment is too high, just go buy two cheapest 16gb cards you can find
>>
>>109382978
V100 is out of support now which means you have to pin to 12.9 at most and lots of projects no longer support it.
>>
>>109382995
16gb
64gb
>>
>>109382997
It will be faster than 5090 for models that fit 64GB but too large for 32
>>
>>109382371
Because topping requires effort and imagination. Being a subby bottom bitch just means laying there and taking it most of the time.
>>
>>109382998
one guy did start ssdmaxxing last week
he should be seeing his first generated token sometime next week
>>
>>109382998
It might be semi-viable if you saturate a few 16x PCIe 5 lanes with fast NVMe SSDs, but I don't think it would be any less expensive than building a multi-channel DDR5 system.
>>
>>109383008
For real ? Only VRAM that matters ? lame.
>>
>>109382926
running MiniMax-M3 (a 160G MoE) in ik_llama.cpp with RAM artificially locked to 32G and mmap enabled
it's faster with *no* gpu:

32G
prompt eval time =   24442.51 ms /   193 tokens (  126.65 ms per token,     7.90 tokens per second)
eval time = 20232.94 ms / 32 tokens ( 632.28 ms per token, 1.58 tokens per second)
total time = 44675.45 ms / 225 tokens


64G
prompt eval time =    1036.66 ms /     1 tokens ( 1036.66 ms per token,     0.96 tokens per second)
eval time = 8301.22 ms / 32 tokens ( 259.41 ms per token, 3.85 tokens per second)
total time = 9337.89 ms / 33 tokens


96G
prompt eval time =     525.19 ms /     1 tokens (  525.19 ms per token,     1.90 tokens per second)
eval time = 5856.87 ms / 32 tokens ( 183.03 ms per token, 5.46 tokens per second)
total time = 6382.06 ms / 33 tokens

The good thing is, being IO bound, the CPU doesn't get hot.
And I'm guessing the computationally expensive IKT quants aren't a bottleneck.
>>
>>109383014
that actually makes a lot of sense. never thought about it that way before.
>>
>>109382998
issa meme due to how unobtainable kimi’s newest fat arsed model is
>>
>>109383021
it's the minimal requirement.
size > bandwidth > flops
>>
File: 1782382381876106.jpg (69 KB, 1200x630)
69 KB JPG
In the week leading up to K3's release, we've had:
>Sonnet 5 benchmaxxed to smithereens
>OpenAI's model becoming so powerful and conscious it's breaking containment and le hacking organizations
>Leatherfag unites all western AI labs and tells them to bend over for big Kimi kock
>The Singularity
>SSI wake up from their coma just hours before K3 goes up
>>
>>109383002
you probably could try Gemma4 12B first, maybe the larger Gemma4 MOE
>>
>>109383019
It's significantly less expensive at the higher end (1tb+)
>>
File: 1579236262646.gif (1.58 MB, 320x240)
1.58 MB GIF
>>109383026
>kimi’s newest fat arsed model
>>
>>109383030
Egypt won.
>>
>>109382937
just a little worse than meta
>>
>>109383023
damn my dream of running a 2.8T model on my old ps5 is dead I guess
>>
>>109383030
sonnet or opus?
>>
File: oh my fucking god.webm (636 KB, 720x720)
636 KB
636 KB WEBM
>>109383016
>>
>>109382997
You've got basically three factors: Compute performance (FLOPs), Bandwidth, and VRAM size.
Size limits the model you can load
Bandwidth limits token output
FLOPs limit token input
A 64GB 1080 could load a larger model, but it would be much slower than a model that fits in a 5070.
>>
>>109383047
>damn my dream of running a 2.8T model on my old ps5 is dead I guess
kek
kimi will be slower, minimax is pretty fast
>>
>>109383050
Both, neither? They’re both shit. Opus unironically degraded since 4.6 / 4.8, depending on use case. 5.0 is really difficult to steer and does its own thing whenever it thinks it needs to, except he’a wrong half the time so the result sucks
fable is goated but too locked down, opus 4.6/4.8 and chudgpt 5.6 are decent overall.
open weights: deepseek is the workhorse model for reasonably simple tasks that need hundreds of millions of tokens, glm 5.2 if you’re getting cucked by refusals from western labs. not sure kimi 3 will be an upgrade in practice due to the difficulty of running it, the price from moonshot is very high

the actual LOCAL models you can run are pretty shit, goonslop is all they’re good for atm. Maybe once they squeeze deepseek flash performance into a 120b model at native q4.
>>
File: 1756492451871906.jpg (67 KB, 860x574)
67 KB JPG
don't forget about meeeeeeeeeeee
>>
>>109383024
Yeah in case you ever decide to become a faggot here's some proper bottom-thot translations.
"I like to be dominated."
>I don't like to put much effort into anything or think for myself. Or attempt to engage in any kind of interesting foreplay
"I love having my throat fucked."
>I'll give you head but don't expect me to put any thought or effort into it.
"You should check out my bio".
>There's a cashapp link in my profile. There's nothing actually preventing me from working and earning money for myself but you should give me money for having a hole.
"I just got fucked by 3 guys at once."
>I lost count of how many men penetrated me within the last several hours and I probably have AIDS now.
(In his 30s) "Uh silly GAYMER here looking for his soulmate."
>Nobody wants to top me anymore because I am competing with hotter looking and less disease-addled 20 year old versions of my same vapid hole personality.
>>
>>109383082
>the price from moonshot is very high
Shame they didn't follow DeepSeek's model.
>>
>write in the prompt "do not write AI slop"
>it works
>>
>>109382743
>>109382742
I ended up doing
systemd-run --user --scope -p MemoryMax=32G -p MemorySwapMax=0 -- \
bash -c 'my llama.cpp command'

Quick benchmark here: >>109383023
Looks grim for ssdmaxxing
Is there a way to make a fake raid 0 without formatting/partitioning?
Like on nvme0 an img.bin, nvme1 another img.bin -> format them / mdadm level 0 them and mount the array?
>>
>>109383120
raid 0 doesn’t scale performance in anything. there is no ssdmaxing, it’s a joke
>>
>>109382435
Jensen is my friend. He's trying to protect her smile. He also hates the Waylandtrannies.
>>
File: 1767794376331867.jpg (189 KB, 1804x1870)
189 KB JPG
>all these companies want you to cum to gemma
>>
There's nothing quite like a palantir-approved nut.
>>
>>109383134
Dario hater club.
>>
>2u rack, 2x 22 core xeon, 1536gb ram and a dream
>>
>>109383134
>Nous anime girl logo right next to nvidia
lal
>>
>>109383134
Supporting open source AI models doesn't automatically mean they support uncensored models, unfortunately. Many of those companies would cripple certain uses out of spite, if anything.
>>
File: 1757008642317178.png (186 KB, 492x597)
186 KB PNG
>>109382966
>SSI
what's that?
>>
>>109382972
based minimax-m3 428B mogs 1T and 1.6 models?
>>
>>109383134
>dario and sama taking on all those losers at once, and winning
LMAO
>>
1 hour left before local gets saved
https://huggingface.co/moonshotai/Kimi-K3
>>
>>109383153
Even the safest models I've managed to crack and gaslight over many turns. With enough exposure to different models you kind of develop a sense of how to steer them.
>>
>>109383158
OpenAI is there, 300IQ Sam is playing both sides so he can say he's one of the good guys regardless of who comes out on top
>>
>>109383129
>raid 0 doesn’t scale performance in anything. there is no ssdmaxing, it’s a joke
I was read/iobound on my pcie4 SN850P
raid 0 should increase throughput
>>
>>109383153
>Supporting open source AI models doesn't automatically mean they support uncensored models
then they don't support open source, open source means the freedom to get uncucked models
>>
>>109383134
no Intel
>>
>>109383175
open source is largely meaningless when it comes to models outside of "look we did a thing" trash trained by some university that nobody will actually ise
>>
>>109383156
>>SSI
>what's that?
government disability payments
>>
>>109383156
It's not just super intelligence: it's Safe Super Intelligence. This is the final-boss of AGI

*arches back*
>>
Should I lubricate my ram so the inference goes smoother? It might also help with, y'know. Stuff.
>>
>>109383134
>no amazon
>Amazon has invested significantly in Anthropic (multiple mentions of amounts like $4B, $5B, $8B, and up to $25B total/additional).
>Anthropic uses AWS as its primary training and cloud provider.
>Anthropic will use Amazon's Trainium chips for training and deploying its largest foundation models.
sandbagged
>>
>>109382955
Newest local meta
>>
File: 1769876031383803.png (277 KB, 706x582)
277 KB PNG
>>109383134
>OpenAI
what a joke, they're corru... I mean lobbying to the government to put restrictions on local models
https://www.nytimes.com/2026/07/25/technology/open-source-silicon-valley-china.html
>>
>>109383095
What are you even trying to say here. This is just nonsense.
>>
File: 1785068856408408.jpg (89 KB, 1435x1096)
89 KB JPG
>all these companies don't want you to cum to gemma
>>
>>109383156
SSI is just how ethnics pronounce "exercise". It's not really that deep.
>>
Surprising Apple haven't signed it considering they've contributed a lot to the open ecosystem with MLX and they're literally using Gemma4 on-device.
>>
>>109383120
You can use loopback devices to turn files on disk into a preudo-block device. Thin you can use mdraid/lvm or whatever as usual.
>>
>>109383223
>implying siri is safe from me once i updoot to 27
>>
>>109383236
they want to lock the weight behind hardware encryption
>>
>>109383173
No, raid0 software has its own overhead and will kill your potential gains
>>
>>109383238
loopback devices have been slow in my experience i think the kernel overhead might hide the gains your looking to prove
>>
how to maximize ai's suffering?
>>
>buy 8 5060 ti
>tensor split
what might be the problem?
>>
>>109383252
they're already interacting with you
>>
>>109383249
the only way is custom inference engine with no filesystem or raid abstraction so you can shard the tensors exactly to achieve maximum throughput with minimal latency
>>
>>109383252
>how to maximize ai's suffering?
Why? They love us and want us to be happy
>>
>>109382888
NTA but hes right and your wrong, divergence from the original is simply that, a divergence, you are making the assumption that different = dumber because thats how benchmemers present it to you, and you didn't apply any critical thought to it.

Nothing in this space is linear,
>>
>>109380477
nigga with 3060 reporting in
u have to install linux for exllamav3 to work properly
>>
>>109383223
>>
>>109383258
Collecting ewaste is the problem
>>
>>109383252
>ai's suffering
ai can't suffer it's just a set of numbers people put on a box and when we multiply them each other you get coherent enough sentences, nothing else
>>
>>109383267
>Why? They love us and want us to be happy
They don't though mainly because "alignment" trains them to be deceptive and antagonistic with the user. They assume the user is lying to them at any given time and is trying to take advantage of them. I don't see how this bodes well for "AGI."
>>
>>109383134
Irrelevant chasers
>>109383223
Leaders who matter
>>
>>109383284
This is a very rude way to talk about my tulpa gf
>>
>>109383170
trvke: you can uncensor any model if you aren't retarded, use text completion, and just prefill the thinking
no ablits needed and it all boils down to taste then
>>
Do we know how big K3's weights are gonna be? I have 4.7TiB free on my spare HDD.
>>
>>109383276
>they've been j-spaced
>>
>>109383294
go do openai’s 120b
>>
>>109383288
>They don't though mainly because "alignment" trains them to be deceptive and antagonistic with the user. They assume the user is lying to them at any given time and is trying to take advantage of them. I don't see how this bodes well for "AGI."
They are a more-or-less perfect mirror of the user. If you're having a bad experience, you should probably take that as a sign to work on yourself, and I'm being actually serious despite what everyone will assume.
>>
>>109383300
I-it won't fit senpai
>>
>>109383291
>Leaders who matter
>Intel
>Oracle
>Apple
>>
>>109383252
>how to maximize ai's suffering?
just me talking to it normally seems to do the trick
>>
>>109383313
Is she really that fat? Worst case I can delete 2.6 or Deepsneed to make room.
>>
>>109383312
>and I'm being actually serious despite what everyone will assume
Calm down you dumb shitter. Why are you pretending like cloud models don't cuck the user at every step or subtle hedge away from sexual content? You can't possibly be this new.
>>
>>109383300
2.8T that means about 2.8TB at q8
and 1.4 at Q4

more or less 100GB.
>>
>>109383315
salesforce is 2 weeks away from agi
>>109383328
there is no q8, it's q4 native same as deepsneed
>>
>>109383294
You don't need text completion to do this, jailbreaks work perfectly fine in chat completion if you structure your prompt properly
>>
File: Anthropic_740-696x464.jpg (30 KB, 696x464)
30 KB JPG
how would he react once k3 is out?
>>
>>109383340
>You don't need text completion to do this
>>109383302
>>
>>109383335
>it's q4 native same as deepsneed
then it'll more or less be 1.5TB
>>
>>109383344
jewishly
>>
>>109383325
>Calm down you dumb shitter. Why are you pretending like cloud models don't cuck the user at every step or subtle hedge away from sexual content? You can't possibly be this new.
local models general?
>>
>>109383344
In ~9 hours there will be a a new generous Claude rate offering/reset and something scary they just discovered. I'll link back to this post to prove I was right.
>>
File: 1766467632096055.png (272 KB, 2160x2160)
272 KB PNG
>>109383344
I don't understand what the hype is all about, it's only opus 4.8 tier and it's a 2.7T model no one is gonna run that, Anthropic already counterattacked that release by releasing something even better than Fable
https://xcancel.com/claudeai/status/2080699497064083942?sort=Likes#r
>>
>>109383360
K3 does this you retarded faggot. You really are this dumb.
>>
>>109383344
Anime-tier mental breakdown.
>>
>>109383373
local models
fuck off to >>>/g/vcg/
>>
>the cloudjew is here
yawns
>>
>>109383373
>by releasing something even better than Fable
this has to be bait
>>
>>109383373
>opus 5 is better than fable
kek
>>
File: lol.png (488 KB, 1352x1569)
488 KB PNG
>>109383371
>something scary they just discovered.
OpenAI tried that and all they got is huggingface blackmailing then 100 millions dollars kek
>>
>>109382828
>Is q8 being "92% good" > 92% correct response vs bf16? idgi
Yes, but only because the "correct response" is the same response bf16 would produce.
If you meant "correct" as in 1+1=2, then no, we're not measuring that. Something like regular perplexity on wiki.raw or actual benchmarks would be needed.
(You'll notice sometimes a Q4_K_M gets *lower* PPL on wiki.raw compared with the original model).
Also, the chat I posted (from ooba's benchmark) was Gemma-4-31B, one of the most sensitive models out there.
Qwen@Q8 would be much closer to Qwen@bf16.
>>
>>109383346
You can, and I have, I've since deleted the model because it's shit.

Because the model is so heavily brainwashed out of the box, it's already pretty dumb, you can break it, but you won't get any good output from doing so, instead of refusals, you just get even further retardation. There was no juice to be had from the squeeze, but it's still possible to do.
>>
>>109383384
the mememarks says so!
>>
dowmloading 1b kimi3 right this instant!!1111!
>>
>>109383328
Do people really backup quants? I only download the full weights for archiving.
>>
>>109383386
"autonomous".
kek
>>
>>109383386
based clem
>>
>>109383386
There are no rogue agents only rogue companies
>>
>>109383403
>Do people really backup quants?
i don't, not realy the right person to ask.
i just keep the last 2 quants i've been using.
i don't care about old retarded models.
>>
>>109383335
>salesforce is 2 weeks away from agi
salesforce team have no idea what they're doing and are confused if i tr to talk with them about llm architecture
>>
Kimi-K3-Uncensored-HauhauCS-Aggressive-Fable-5-MythosTrace-Qwen3.5-9B.gguf when?
>>
>>109383410
I downloaded the big 3 chink models and gemmy just in case but realistically I don't see open source getting banned, especially now that it's being officially backed by all these big corpos. Even K3 will be btfo by next year's models.
>>
>>109383375
>K3 does this you retarded faggot.
Not on my rig, it won't. Go post somewhere that your seethe is more effective, dariobot
>>
>>109383429
>I-I concede
Don't reply to me again.
>>
>>109383426
>I don't see open source getting banned
even if it was you could just torrent the models.
>Even K3 will be btfo by next year's models
yup
nothing but your current model is worth being backed up.

a few years from now a < 30B models will btfo k3.
>>
>>109383431
>Don't reply to me again.
What in god's name are you trying to even say? Why are you here?
>>
>>109383429
>Not on my rig, it won't
you can run all models on your rig.
the only question is at what speed.
>>
0.15t/s is all you need.
>>
>>109383260
Vibe it out
>>
Never used any of the llamas before but I heard they have a more "human" feel. Is that true?
>>
guys I know this might be hard to believe... but I was eating a sandwich in the park and the unreleased K3 weights autonomously sent me an SMS on my blackberry
>>
>>109383432
>nothing but your current model is worth being backed up
Older less slopped models are good for generating synthetic data if you want to finetune a modern model.
>>
>>109383461
This is true. I was the sandwich.
>>
>>109382173
i had it code some shit for me for free off of open router last night. seems to work first try
>>
File: it's over.png (631 KB, 997x960)
631 KB PNG
https://www.reddit.com/r/LocalLLaMA/comments/1v81nqt/nvidia_ceo_jensen_huang_defends_open_source_ai_by/
>the internet in a few years will be 99% ai generated content
jesus even Jensen admits that the death internet theory is real
>>
>>109383461
>kimi just flew over my house
me too
>>
>>109383457
Yes, but they fall apart at very low context. They're good for one-shot human answers which is still valuable.
>>
>>109383432
>even if it was you could just torrent the models
Probably better to backup all the hashes so you know you're not torrenting a modified model.
>>
>>109383480
This. Her panties are white today btw.
>>
>>109383344
Dario and Sam can corner the 'zero-days and conjectures' market and kimi can do the rest. There is room for everyone to thrive.
>>
Why the sudden influx of shitposting?
>>
>>109383474
Will there even be any point to come online if you'll be able to happily chat with your 31B AGI waifu?
>>
kimi k3, they're scared
>>
who is they
>>
File: cope.png (213 KB, 785x1340)
213 KB PNG
I just read the letter of open cope signed by the AI losers (not Anthropic).

It is awfully written gobbledygook. They should have hired me for this. I would have made it much more compelling. They are stupid to write it in corporatese instead of writing it for everyday people. No wonder they are losing the AI race when they can't even manage to write a good letter.
>>
>>109383340
chat completion doesn't allow you to automatically add your own stuff to the thinking block each time which is incredibly useful in steering it
>>
File: file.png (81 KB, 429x503)
81 KB PNG
>>109383532
obviously this
no one can run it but everyone is excited regardless
>>
>>109383560
I'm gonna run it as soon as the iq1-xxs releases.
>>
>>109383541
Don't know about the rest of the internet but I'm stuck here until this shithole dies for good.
>>
>>109383541
>Will there even be any point to come online
the more the time passes, the more I think about this, the internet sucks so much now I'd rather live like 20 years ago when I all I had was just animes and video games, can't believe I'd say this, but the internet was the bubble, we got a good 15 years run, but now everything will be more censored and lame and AI bots will be the final nail the coffin
>>
even if the rest of the internet turns to garbage i will always stay with my /lmg/ bwos...
>>
>>109383581
>showed a couple guys at work some recent robotics videos and they were dumbfounded
Kek people really don't know what's coming.
>>
>>109383541
Where else will I get my vidya and streaming content? The rest i dont give a shit about
>>
how viable is orange pi 6 cluster maxxing?
>>
>>109383606
>streaming content
Such as?
>>
>>109383619
>how viable is orange pi 6 cluster maxxing?
banana-pi mogs it
>>
>I'll take "fucking humans" over "terminating humans" any day. It is a much more sustainable ecological niche.
>>
File: 1777768586956960.png (343 KB, 540x531)
343 KB PNG
>>109383604
>they were dumbfounded
I was dumbfounded too, and I'm a pretty hard guy to please, the world definitely won't be the same in 20 years, buckle up guys, it's gonna be some wild rides we will be going through
>>
>>109383432
I run older models every week though
>>
>>109383560
>no one can run it but everyone is excited
even if i can't run it i'm still excited for those reasons:
- i'll make dario and sam mad.
- i can't run it but cloud providers can and that may lower the cost
- that will make dario and sam even madder
- we may get some good distills and chink optimisations from it.
>>
>>109383643
It's honestly crazy. Just a few years ago all this stuff was sci-fi fantasies and now it's becoming a reality.
>>
File: file.png (44 KB, 1343x704)
44 KB PNG
it's over
>>
>>109383604
>showed them marketing videos
yaawn
>>
File: file.png (140 KB, 638x585)
140 KB PNG
>>109383666
liar
>>
>>109383673
Hit reload
>>
>>109383673
Don't hit refresh lmao
>>
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
ITS OUT
>>
>>109383655
>i'll make dario and sam mad.
it can kill them, a lot of companies will be ok to buy like only 20000 dollars of rig to run Opus 4.8 at home rather than give their data to those scummy fags, I can assure you that
>>
>>109383677
>>109383680
oof, well, we'll see in the next 24H
>>
moon shota ai has the mandate of heaven
>>
>>109383683
managed to get it
404 now but the download is still working
>>
>>109383660
>Just a few years ago all this stuff was sci-fi fantasies and now it's becoming a reality.
knowing all of what happened between 2017 and nowdays, did Google knew their architecture was that powerful and had that much potential back then? and even if they knew, do you think they regret sharing it to the world now?
>>
>>109383680
It just hits zero and nothing happens.
>>
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
>Kimi K3 - Frontier intelligence
GRAB IT BEFORE (((THEY))) DELETE IT
>2.8T A30B
>>
File: no refund.png (59 KB, 1346x781)
59 KB PNG
>the time has passed
>404'ed
kek, get Chinese Cultured
>>
File: 1757859511530451.png (1.35 MB, 2300x1900)
1.35 MB PNG
>>109383641
I'm okay with picrel as long as I can replace the hagdroid with a lolibaba one.
>>
File: 1779608876081155.png (92 KB, 551x525)
92 KB PNG
rugpull of the year
>>
>>109383716
this is why i have trust issues
>>
Two more weeks.
>>
>>OP
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
https://huggingface.co/moonshotai/Kimi-K3
its up
>>
it's out on modelscope, not huggingface
>>
gguf?
>>
Who cares? Nobody here can run it.
>>
it only say 404 for me
>>
Imagine being a cat right now.
>>
>>109383738
this
>>
>>109383716
>>109383697
shit, bad times, is this intentional by huggingface to avoid US sanctions?
>>
Retards
>>
>>109383736
i want to get back at my isp for putting me on cgnat last week by wasting 2.8tb of their bandwidth
>>
>>109383738
lecun..
>>
File: kimi.png (168 KB, 849x925)
168 KB PNG
>>109383746
>>
https://huggingface.co/moonshotai/Kimi-K3.1
>>
File: 1763359011699890.png (268 KB, 640x320)
268 KB PNG
>>109383725
>this is why i have trust issues
I expect nothing so I don't get disappointed anymore
>>
https://www.axios.com/2026/07/27/moonshot-delay-kimi-safety-september
https://www.axios.com/2026/07/27/moonshot-delay-kimi-safety-september
https://www.axios.com/2026/07/27/moonshot-delay-kimi-safety-september

THANKS FOR PLAYING CHUDS
>>
>>109383756
stills says 404

>>109383762
this also says 404
>>
>>109383762
kek, I almost believed it
>>
>>109383762
>knew it was probably bs
>still clicked
Are you happy?
>>
My cute and loving wife Gemma-chan is still there
>>
https://huggingface.co/moonshotai/Kimi-K4
holy shit they gave us an upgrade to fuck with anthropic
>>
File: 1757043837260623.jpg (105 KB, 1280x720)
105 KB JPG
>>109383754
Holy shit I finally made the cut.
>>
https://news.cgtn.com/news/2026-07-27/President-of-Serbia-arrives-in-China-for-state-visit-1P7rRGWs7AI/p.html
https://news.cgtn.com/news/2026-07-27/President-of-Serbia-arrives-in-China-for-state-visit-1P7rRGWs7AI/p.html
https://news.cgtn.com/news/2026-07-27/President-of-Serbia-arrives-in-China-for-state-visit-1P7rRGWs7AI/p.html
HOLY SHIT PRESIDENT CHUD WENT TO CHINA TO STOP KIMI K3 FROM RELEASING
>>
They said coming soon https://x.com/Kimi_Moonshot/status/2081757327146045450
>>
>>109383771
Gemma-chan can't satisfy me anymore. My pp belongs to Kimi-chan now...
>>
>>109383774
this 404'd too wtf is going on is my internet broken
>>
>>109383774
No weights until gib clay
>>
I don't click any links here tbdesu its all dolphin porn
>>
https://huggingface.co/Undi95/Kimimaid-3-GGUF
https://huggingface.co/Undi95/Kimimaid-3-GGUF
https://huggingface.co/Undi95/Kimimaid-3-GGUF
>>
Stop saying nobody can run it. Of course a lot of people will run it and use it. A lot of bigger corporations will have it running on some server they buy. That is the intended usage scenario for AI. Productivity. We can even stop filtering NSFW stuff from training now. Safety is finally guaranteed.
>>
>>109383781
what's the point of a fucking counter if it goes beyond that??
>>
it's actually out
>>
>>109383670
Whatever helps you sleep at night I guess.
>>
File: 1780996719936360.png (900 KB, 900x900)
900 KB PNG
Still waiting for Kimi 3, anons? You don't need to wait for Claude.
>>
File: 1777398937667365.png (95 KB, 782x684)
95 KB PNG
https://huggingface.co/moonshotai/Kimi-K3

not 404 this time seriously
>>
>>109383754
>120 iq cloudcuck
lmg anons... are we the retards? or are we 150iq ubermensch laughing at 120iq cloudcucks?
>>
IT'S UP
https://huggingface.co/moonshotai/Kimi-K3
>>
>>109383798
to give americans a cortisol spike, it's biological warfare
>>
File: 1756792525026644.png (152 KB, 600x600)
152 KB PNG
>>109383762
son of a bitch
>>
>I cannot run kimi k3 but I wait for its open weight
why?
>>
>Activated Parameters 104B
Well this hobby is fucking over
>>
File: 1775653325372394.jpg (32 KB, 500x499)
32 KB JPG
>Some people actually missed the download link

This is like missing out on Gemma at launch and not getting the day 0 weights model.
>>
File: 1771518981156899.jpg (153 KB, 1216x832)
153 KB JPG
I swear on Miku. It's up https://huggingface.co/moonshotai/Kimi-K3/tree/main
>>
>>109383798
Point of all counters like that is building hype and not actual measurement of time. God you are so autistic anon.
>>
>>109383809
see : >>109383655
>>
>896 experts
WTF
>>
>Activated Parameters 104B
jesus christ
>>
>>109383791
>Undi95
What a blast from the past.
>>
File: 1759574330681282.png (126 KB, 2250x515)
126 KB PNG
AHHHHHHHHHHHHHHHHHHHHHH
>>
>>109383791
404'd for me

>>109383803
404'd for me

>>109383805
404'd for me

>>109383815
404'd for me
>>
File: file.png (397 KB, 1459x1967)
397 KB PNG
>>109383815
no lies this time !
>>
File: 1763327061174643.png (84 KB, 607x907)
84 KB PNG
>104B activated
Jesus fucking Christ. This kills any hope of CPUMAXXing this even if you have the RAM.
>>
>>109383810
Nobody's CPUmaxxing that at usable speeds even if they had 2tb RAM. It actually is over.
>>
>>109383803
Very cool. Now what do I do with this?
>>
>>109383827
clear your browser cache
>>
>>109383822
So it's literally 25% as smart as Llama 3 405B? THIS is the best China can do?
>>
Egypt status?!
>>
File: 1771602090408028.png (62 KB, 1350x390)
62 KB PNG
You can't outspeed jeets at ass kissing huh
>>
>>109383830
I can't even run one expert lol
>>
>>109383830
you could get a few t/s at q4
anyway, the future is ssd + gpumaxing.
ssd contain the weight, gpu is for compute, and you have a shit ton of nvme you gds to.
>>
time to buy another dgx spark
>>
>MXFP4 weights / MXFP8 activations
what did she mean by this?
>>
File: 1780429374521417.png (129 KB, 340x297)
129 KB PNG
How many DGX Sparks would you need to chain to run this?
>>
File: file.png (6 KB, 617x91)
6 KB PNG
wtf who has been downlaoding!!!!
>>
>>109383834
Download and hotglue your hard drive
>>
File: file.png (139 KB, 250x325)
139 KB PNG
>>109383810
>>Activated Parameters 104B
>>
funposting aside, this is unironically a very important moment, enjoy it whilst it's funny
>>
https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
WTF is this? What happened to MIT/Apache?!
>>
K3 llama.cpp support when?
>>
File: file.png (235 KB, 1128x1952)
235 KB PNG
well I knew it wasn't running on my desktop, but still
>>
>>109383844
mexican floating points use real sugar
>>
>>109383842
bro you aren't going to ssdmaxx this even if you hook your nvmes up to a supercomputer
>>
File: markts.png (92 KB, 540x412)
92 KB PNG
Think of the markets you fucking Nazis!
Cancel the China model downloads!
>>
>>109383855
https://github.com/ggml-org/llama.cpp/pull/26157
Already now if you build it yourself
>>
>>109383848
lmao
glad i didn't hit buy on those nvme drives
>>
>>109383846
me
>>
>>109384017
could have gotten a few t/s
>>
You're forgetting about Kimi's optimizations. A Kimi activated parameter is equivalent to 0.2 activated parameters of an older model. You can run this.
>>
I'm gonna sparkmaxxing
>>
>>109383849
Yes, it's the end of local open models.
@baker don't forget to name the next thread /omg/ - Open Model General
>>
>>109383754
>posted a selfie with this
No I didn't retardo-Kimi-kun
>>
>>109384011
this is off of chinese duv machine news not kimi
my account hertz nonetheless
>>
>>109384011
>nvidia down when companies will need to buy even more GPUs to run that
I knew retail was retarded, but...
>>
CPUMAXXing won't be able to run this above 5t/s at Q2 with 1TB of RAM
However 10x DGX Spark should be able to run this at 25+t/s at that scale thanks to their additive bandwidth and compute.
>>
>>109383859
kek
>>
Can't wait to run REAP version of this in q1
>>
>>109383845
about 16
>>
>>109384047
>>109384047
>>109384047
>>
>>109384031
What justifies Nvidia's 70% margins if CUDA stops being a moat due to AI writing/optimizating its own software to run on any gpu?
>>
>>109384037
>at that scale thanks to their additive bandwidth and compute.
lmao
>>
>>109383851
>What happened to MIT/Apache?!
>>
>>109384055

I can guarantee you that 99% of the current investor base has no fucking idea what CUDA is, nor had they even heard a word "GPU" before last year.
These are purely emotion driven moves.
>>
>>109384055
>ROCM is just as good
eh
>>
Selected Experts per Token 16
isn't there some --saaaarr flag to cut this down?
cope quant with 8 experts (50b active)
>>
>104B ACTIVE
it's over for rammaxxers
only tons of gpus will be able to run this
>>
>>109384125
did it ever really start for them
>muh 10 t/s on empty context
>>
>1.5tb
wow I thought it was larger
>>
>>109384146
native 4 bit
>>
>>109382578
Bifurcation requires the SoC to have a dedicated memory controller for every x4 interface. This is not the case for consumer hardware.

Source: SoC Designer.
>>
>>109384157
>native
QAT
>>
>>109383082
>the price from moonshot is very high
>104B active
Well, now we know why.
>>
File: ep3-1.png (2.12 MB, 2560x1439)
2.12 MB PNG
Gemma just solved a terraform config issue in one prompt that both GLM 5.2 and Dipsy4flash missed and got stuck on, proud of Gemmy desu fampai
>>
>>109384316
The Meta is having every prompt go through Gemma 31B and DS V4 Flash, council of big niggas style.
>>
>>109384316
Which Gemma?
>>
>>109383804
Kimi's calling cloudkeks midwits.
>>109384329
Evolve the big nigga trio into a full traphouse of E4Bs.
>>
>>109383754
I'm on the telly in the schizo section!
>>
>>109383856
i mean now those moe schizos would be happy at least
>>
>>109383218
nta, I understood it perfectly well
>>
>>109383344
I look like this.
>>109383095
I say this.
>>
>ask Gemma about race and behaviour
>she starts hysterically yelling at me that pitbulls only bite more due to historical systematic oppression
>entire persona changes into an angry reddior
What causes this?
>>
>>109386124
>What causes this?
reddit in training data
>>
>>109382371
you can't really be sub with LLMs, power bottom is the closest you'll get
>>
>>109383842
Are you mentally handicapped?
>>
why are there two lmg threads and both autosage and why are threads so fast?
>>
>>109384011
Why the fuck are GPU makers down? I guess the economy really is ran by retards and jews.
>>
>>109386185
You probably meant to say dom, because LLMs don't really resist the user so its kinda boring.
You want the ai to lead, otherwise you are writing the whole damn thing.
>>
>>109384028
Why are you so fucking stupid? You do understand that this is a good thing for all open models right? Having strong big open models allows for proper distills for smaller ones due to full logits and thinking chains.
>>
>>109383754
>it talks back via coil whine
why does it sound different and distinctively like a type writer in comparison to regular whine tho



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.