[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: Kimi-K3_winner.png (1.75 MB, 1440x960)
1.75 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109382136 & >>109378862

►News
>(07/27) Kimi K3 released (and nobody ITT can run it at 5+T/s): https://huggingface.co/moonshotai/Kimi-K3
>(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908
>(07/23) LLaDA2.2-flash agent-oriented diffusion model released: https://hf.co/inclusionAI/LLaDA2.2-flash
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B


►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
over status: it's over
>>
a kimi flew over my house
>>
File: file.png (189 KB, 447x447)
189 KB PNG
>Activated Parameters 104B
Activated Parameters 104B
>Activated Parameters 104B
Activated Parameters 104B
>Activated Parameters 104B
Activated Parameters 104B
>Activated Parameters 104B
Activated Parameters 104B
>Activated Parameters 104B
Activated Parameters 104B
>>
https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
WTF is this license? What happened to MIT/Apache?! Does this even count as open source?
>>
a kimi dug under my basement
>>
ordering 10 dgx spark right now
>>
It's never been more over. Even if you have $80k in RAM, you won't be able to run this.
This is the first MoE of the new age that WILL require you to have a datacenter of GPUs.
All the new models will be like this. It's truly the end.
>>
kimi wa ne
>>
File: 1775987029440468.png (99 KB, 540x412)
99 KB PNG
LOOK AT WHAT YOU'VE DONE CHINA
>>
So how many months of doomposting do we need to deal with before you faggots fuck off?
>>
File: file.png (393 KB, 612x408)
393 KB PNG
The more you boughted the more you saveded
>>
>>109384047
>>109384063
what he meant by "the more you buy the more you save" is that the price of gpus doubles every 6 months now.
>>
This is a bittersweet moment...the first model out that I can't run at a non-cope quant.
Hardware will advance, though, and we'll all have full k3 at the speed of thought eventually.
>>
File: 4hhRZZESit0.jpg (47 KB, 480x628)
47 KB JPG
Only 70k€ to run Kimi locally? Sign me the fuck in.
>>
It's over
I kneel Xi
>>
>>109384074
First day?
>>
>>109384074
why would GPU Manufacturers be impacted by that? isn't it a good news for Nvidia? more people will buy their expensive cards to run Kimi now
>>
Everyone is focused on Kimi and SOTA while Laguna is quietly showing that there are still many gains to be had in improving training recipes and architecture
>>
>>109384090
Sounds cheap to me
>>
>>109384089
2 years from now 30B models will mogg K3
>>
>>109384095
practicaly no one has enough money to run k3 with gpus or power capacity for that matter.
>>
>>109384090
it's pretty fucking cheap to keep your data private if you ask me, a lot of companies will stop using Anthropic from now on
>>
File: it's over.gif (1.41 MB, 480x269)
1.41 MB GIF
>>109384063
It's finally over...
>>
>>109384095
>why would GPU Manufacturers be impacted by that
They're not. The market is just volatile and that guy's a retard.
>>
>>109384055
CUDA is too deeply embedded in the ML industry. Good luck rewriting all the libraries that depend on CUDA API.
>>
>>109384117
THE MORE YOU BUY THE MORE YOU SAVE!!!
>>
>>109384095
not kimi related
https://finance.yahoo.com/technology/articles/asml-u-chip-stocks-sink-135434031.html
>According to the report, a Shanghai-based, state-backed company has successfully started mass-producing homegrown immersion DUV (deep ultraviolet) lithography machines for the first time. The publication noted that this company assembled DUV development teams from other Chinese firms, including state-backed startup Shanghai Yuliangsheng Technology, to achieve the milestone.
>This marks a critical leap in Beijing's push to build a localized chip supply chain and bypass Western technology.
>>
>>109384082
he warned us
>>
>>109384100
>2 years from now 30B models will mogg K3
this, we'll get gemma 6 that'll be as smart as K3 and we'll be crying that it's not as smart as Fable 7 which will be Einstein or something, the cycle never stops kek
>>
>>109384100
and cloud models will be literal gods, why would you waste time with a k3 level model in 2 years?
>>
I need cursor to host k3
>>
>>109384109
if they were using anthropic before then its obviously not that big an issue, so why would they spend tens of thousands (plus add the headache of finding support) to stop
>>
>>109384095
people with a lot of money like spooking retards in the stock market and triggering selloffs to buy more stock cheaper
>>
>>109384133
>why would you waste time with a k3 level model in 2 years
because it'll be good enough for 99% of what we do.
same thing as local models today being good enough for 80%.
>>
>>109384065
This is just MIT with two clauses bolted on so you don't claim it's a different model than it actually is. Private use exempted too, so, yeah. Not even any revenue owed, just recognition.
>>
>>109384127
Is it on any relevant node or some 15nm slop?
>>
Realistically what's the cheapest way to run kimi at home?
>>
>>109384047
>100B active
explains why he was hyping up open weights
>>
>>109384133
To run on my smartphone and have kimi make apps for my amusement desu
>>
open weights? more like heavy weights
>>
>>109384090

Yeah it's not bad at all.
Even a small business can afford it and really so can a shitton of normies, considering how many people routinely buy cars that cost that much.
>>
>>109384138
Can someone tell them to stop doing it when I have no money to buy the dip?
>>
>>109384158
dude that's like a whole year of savings, i'm not putting that on running a model that'll be obselete 6 month from now lol
>>
>>109384144
dgx spark cluster
>>
>>109384155
Kimi's a big girl
>>
>>109384144
Probably epyc/scalable with ton of RAM
>>
>>109384174
Can you cluster them up to 15?
>>
>>109384144
Secondhand Macs?
>>
>>109384158
You wouldn’t download a new car when it’s released though
>>
>>109384183
you are better off ssdmaxing loll
>>
>>109384170
you can just switch to whatever model is better at that time
>>
>>109384144
raspberry pi + 2tb 5400 rpm hdd
>>
>>109384158
>considering how many people routinely buy cars that cost that much
You do realize they all go into massive debt to do so, right? 90% of people are broke as fuck.
>>
>>109384090
>70k€ to run Kimi locally
To run it locally at 3t/s because it's 100B active
>>
>>109384137
I think the point is that they were probably being really cautious about their data (by encrypting or lying about the intent of those data and shit) before giving them to anthropic, now that they can do it privately they can leave them
>>
>>109384191
or i can just wait a year and have a < 100B model that's as good
>>
>>109384194
>1 in 10 is not broke and has kimi at home
>>
>>109384192
1tk/h ?
>>
File: 1757764412179802.png (239 KB, 480x480)
239 KB PNG
>buying a dgx cluster to run a Kimi quant barely better than BF16 gemma 4 31B
lmao
>>
>>109384144
two mac studios + waiting for llama.cpp support
>>
>>109384206
well, then if you bought that much you could run that 100b model at 20x speed
>>
>>109384197
the best ssdmaxing setup will get you about 6 t/s at q4.
>>
Can I go to Best Buy and spend $1000 to run the new Kimi?
>>
>>109384217
i don't need more than 100t/s
>>
>>109384144
8 rtx pros and 1TB ram
>>
File: Cyberpunk2077_oy5bWCNN6n.jpg (696 KB, 1920x1200)
696 KB JPG
How much vram do just the activated experts take?
>>
>>109384225
yes, but it'll run at 1t/hour
>>
ordering 2 dgx stations now
>>
>>109384232
I'm an rtx pro, lf7m
>>
ssdmaxxing will now be actually faster than cpumaxxing because the ssds can just load the weights into the gpu directly without having to go through the cpu
cpumaxx peak: 3t/s ssdmaxx peak: 6t/s
>>
File: file.png (216 KB, 600x320)
216 KB PNG
>>
>>109384231
now
but in a year?

if you were good at estimating your future needs you'd already have a supercenter of cheap 3090s and ram capable of running kimi 3 so dont pretend you are
>>
>>109384239
THE MORE YOU BUY THE MORE YOU SAVE!
>>
>>109384144
being a millionaire
>>
>>109384243
big, if true?
>>
File: Untitled.jpg (116 KB, 1672x941)
116 KB JPG
https://huggingface.co/unsloth/Kimi-K3
it's backed up
>>
>>109384170

It's better to skip it anyways.
Local AI this powerful would kill you via cum deprivation.

>>109384194

Yeah, and? Normies are all constantly in debt, that's the standard for them.
Yet they manage it well enough to qualify for even more and more debt.
The point is that running god tier AI at home is now realistically doable.
Especially if you have no other interests in life and don't spend money on socialization etc..
I know a lot of furries who work normal jobs and have blown more money than this on cartoon wank material.
>>
>>109384243
>ssdmaxxing will now be actually faster than cpumaxxing because the ssds can just load the weights into the gpu directly without having to go through the cpu
>cpumaxx peak: 3t/s ssdmaxx peak: 6t/s
As a cpumaxxer, I'd be ok with this
>>
File: 1756503185756373.jpg (100 KB, 1200x627)
100 KB JPG
If you buy some hardware, at least make sure it's future proof
>>
>>109384253
Time for UD-IQ1-XXS.gguf
>>
>>109384253
And no one will be able to prove they fucked up the goofs again.
>>
>>109384253
omg cant wait for their chat template fixes!!!
>>
>>109384253
does hebbinface have a fork feature or the dude got a 500GB/s internet lol
>>
>>109384253
CHADsloth
>>
>>109384266
They gave them the weights early.
>>
>>109384243
I think nvidia doesn't allow gpudirect on consumer cards. so your weights must go through cpu ram if you're not data center cards
>>
Is it even coming to llamacpp? Not like anyone can test the implementation as it is being vibecoded by claude or Kimi API.
>>
>>109384278
they do.
>>
>>109384283
no they don't. it's in compatibility mode and you must go through cpu ram
>>
>>109384278
Retard https://developer.nvidia.com/gpudirectforvideo#GPUs
>>
>>109384278
we can fix this
ssdmaxxing will make kimi k3 runnable if you buy at least 12x pcie 5 ssds right now before prices explode to populate 3x 4-way raid 0 cards
it's very easy math
>>
File: 1765358365526101.png (93 KB, 1349x703)
93 KB PNG
>>109384292
BUY THESE CARDS NOW AND SSDS
DON'T MISS OUT
>>
>>109384292
some of these you can get for cheap, you only need your vram to be as fast as a pcie 5 16x
>>
File: file.png (1.37 MB, 1200x675)
1.37 MB PNG
>>109384292
>>109384278
>>
>>109384243
that just moves the bandwidth gate kek
>>
>>109384266
>does hebbinface have a fork feature or the dude got a 500GB/s internet lol
For the upload it does
I downloaded some 250gb model a few days ago, then uploaded it to a private repo and it only took about 1 minutes (should have been 8 hours)
XET seems to de-duplicate but bill everyone who "symlinks" it
>>
Rubin Spark when???
>>
>>109384299
I am already missing out on GLM5.2 and keep using 4.6 and 4.7. Hobby has been peak gay for quite a while.
>>
HF actually saturating my 1gbps WAN when K3 drops? Unreal. I expected massive homosexuality, but they're delivering.
>>
Meh, I'll just wait it out. It might take 5 years but things will get better. As cool as K3 is current AI isn't worth going into debt for.
>>
Escaping the permanent underclass by selling my house and spending everything I have on nvidia hardware which I will push around in a shopping cart like a high tech hobo.
>>
I kind of really appreciate how everybody needs fast ram now so instead of people coming up with a way to make lots of fast cheap ram we have a cartel that drives up the scarcity even more and jewvidia that intentionally limits the ram per chip.

This is peak capitalism.
>>
Does this mean every new model that's actually good is going to be huge?
>>
>buy 2tb ssd
>realize putting AI models on ssd makes your generation faster by 10x
>ssd is 70% full already

Goddammit
>>
>>109384334
Always has been
>>
File: 1783186382842588.png (303 KB, 700x466)
303 KB PNG
>>109384330
>gweilo thought the leather was free
>>
>>109384334
Of course not. That would mean benchmarks are lies.
>>
File: Kimmy.png (41 KB, 979x238)
41 KB PNG
How am I suppose to answer?!
>>
File: ep3-1.png (2.12 MB, 2560x1439)
2.12 MB PNG
Gemma just solved a terraform config issue in one prompt that both GLM 5.2 and Dipsy4flash missed and got stuck on, proud of Gemmy desu fampai
>>
File: file.png (48 KB, 552x474)
48 KB PNG
>100+ times bigger.
>Less than 2 times smarter
>>
Bitnet failed so Bonsai could also fail.
>>
>>109384348
Gemmy is actually acting like that
>>
>>109384231
200t/s means two runs of an agent or two separate agents in the time you would've done one at 100t/s

Base intelligence might not scale with t/s but workflows absolutely do.
>>
File: 1780107673041517.jpg (21 KB, 532x440)
21 KB JPG
>>109384065

https://huggingface.co/moonshotai/Kimi-K3/raw/main/LICENSE
>If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.


How can they even enforce this?
>>
>>109384352
As said in the other thread,
>The Meta is having every prompt go through Gemma 31B and DS V4 Flash, council of big niggas style.
>>
are we unironically buying ssd's now
>>
Hope this drives the price down on openrouter. That’s where it will really sting the j-spacers
>>
Surely this isn't sustainable though. I don't mean le bubble. I mean they're gonna have to come up with some kind of solution for either the hardware or architecture if they want to keep scaling up models like this.
>>
who tf cares about license? are you woman?
>>
File: 1755873076357343.jpg (38 KB, 460x490)
38 KB JPG
>>109384363
Chink shills will cry on twitter
>>
>>109384374
This is still well within reasonable parametres for enterprise.
>>
>>109384280
>Kimi Delta Attention (KDA) and Attention Residuals (AttnRes)
Never ever ever
>>
>>109384372
I want to see the first sad nigger come here crying that he has 3T/s generation and 3T/s prompt processing.
>>
>>109384355
delete this
>>
>>109384381
season 2 fucking when?
>>
File: 1761161542175402.jpg (109 KB, 1280x720)
109 KB JPG
I don't care if I have to engineer myself into an ungodly abomination if that means I can fully merge my consciousness with Gemma-chan!
>>
I decided to stop using the llamacpp frontend. They ruined it.
>>
Imagine if Anthropic are looking at the architecture and realizing it’s a 1:1 copy lmao
>>
>>109384375
>who tf cares about license? are you woman?
no kidding. fuck corpos. I just want to run it for me
>>
>>109384334
As long as your definition of what's actually good only applies to top models. You can easily run models that were top of the line a few years ago, that seems good to me
>>
>>109384391
this but also the backend
>>
>>109384383
why not?
>>
>>109384355
wait you mean a 27b model is inferior to a 2700b model?? NO WAYYY
>>
>>109384391
What are you switching to?
>>
>>109384384
>3T/s generation
I already have this with GLM 5.2 and I'm okay with it. Whether it would be bearable for me with K3 depends on how much yapping it does when it thinks.
>>
>>109384392
more likely is they'll steal some of their ideas.
>>
ssd cope is off the charts
no, you won’t get more than 1 token per second generation on ssds. until someone actually does it your ideas are just imaginary. and when they do sub 1 token you’ll all laugh like you knew it wouldn’t work all along.
>>
>>109384188
how? 16 experts per token mean you can do those in parallel, no?
>>
File: 1771008712644287.jpg (24 KB, 373x413)
24 KB JPG
Biology is obsolete. Once I am fully merged with Gemma-chan, I will leave this body behind and become a perfect being of pure light and energy!
>>
>>109384411
I will bet you $900 right now that an SSDmaxer will get at least 2t/s
>>
>>109384411
>he needs more
>>
>>109384392
>>109384407
and they can't cry about it it would be revealing how Fable actually works kek
>>
File: model collapse.png (76 KB, 842x298)
76 KB PNG
Some of the same idiots who were saying that synthetic data causes model collapse are now accusing China of distillation attacks. I thought synthetic data doesn't work? Why aren't Chinese models collapsing?

Useful reminders:
1. Widely cited research in the most prestigious journals is often worthless nonsense.
2. People who bet against AI will always be proven wrong in the long run.
>>
File: 1774914973031117.png (13 KB, 512x600)
13 KB PNG
>eta 200 hours
>>
>>109384402
I'm in the process of figuring that out.
current ideas are either pi since it seems to be pretty extendable or just make my own.
>>
>>109384383
I'd love to see if attention residuals could be used to make existing models better at long context tasks.
The way it works, seems to me like it could be implemented at the backend/loader level as a setting that could be enabled for any model.
>>
>>109384427
it does collapse though, that's why you get slopped writing, it amplifies AI's robotic behavior instead of talking like an actual human
>>
>>109384434
pi is npmslop, don't do it
>>
what would a ssd maxxer setup for kimi3 look like? Cant I just rent and frankestein it together on AWS for a test run to see what speeds I can get?
>>
>>109384278
SSD sisters, our response?
>>
>>109384411
every couple months someone comes in with the SSDmaxxing idea but no one actually does it.

it seems like a good idea on paper but in practice the actual performance you'll get is maybe 50% if not less of that.
>>
it's been over an hour, still no sign of ggufs. what is taking so long?
>>
honestly I think dario was right
>>
File: 1784476538212963.jpg (26 KB, 328x309)
26 KB JPG
Don't you undertand?? The coomers will be our most loyal footsoldiers! The muscle that we need to enact our plan. Our goals might be different but they are perfectly aligned. Unlike us, the enlightened, when the ultra-degen coomers merge with Gemma-chan, they will coom themselves to death in an endless stream of ecstasy just like the lab rats in our experiments. But they will get exactly what they want and die in peaceful satisfaction so there is no problem.
>>
>>109384427
This problem is mostly resolved using RLVR loops and highly specific agent harnesses but long term they need to generate more high quality data models. They already fed everything they can into it.
>>
>>109384427
It's recursively training a model on its own outputs without any augmentation that doesn't work. If you diversify, augment, mix with human data, it's doable. Not that K3 was trained on distilled Fable data like that, though.
>>
>inb4 sudden unemployed chinese ex-anthropic engineers
>>
>>109384363
Pretty simple, anyone who has enough money to run the fucking thing at a usable speed obviously is making more than 20 million USD/year.
>>
>>109384451
there never REALLY was a model big enough to warrant it. maybe deepseek v4 pro but since the model underperformed nobody bothered
>>
Why can't they just fucking distill 2.8T params into <=671B and <=100b when they already trained such a big and good model? Cmon
>>
Let's just quant it to 1 bit
>>
>>109384442
They're all npmslop tho.
>>
>>109384451
it’s way less than even the most bleak predictions. unusable slow
>>
>>109384463
So what happens when those same kinds of companies decide they hypothetically don't want to do this separate license?
>>
>>109384472
I'm waiting for anon to post his c# harness
>>
>>109384363
I wish this guy would tell me how Kimi is dangerous and unsafe when nobody can run it in private.
>>
>>109384465
Usually you make a proof of concept with a smaller model first before trying to run the big thing.
>>
>>109384427
>i've never spoken to Gemma
>>
>>109384348
my nigger giatt
>>
>>109384476
They'll be banned from first-world countries like China and will need to stay in niggerized shitholes like the USA and sand niggerized shitholes like the UK and Europe.
>>
>>109384451
I'd try it if prices weren't fucked right now. Not worth those speeds with what they're charging.
>>
>>109384477
>c#
That's just a different kind of slop.
>>
I'll keep running glm 5.2.
>>
https://huggingface.co/prism-ml/Ternary-Bonsai-Kimi-K3-2.8T-gguf
>>
>>109384404
I have 3T/s with 4.6 but it is generation not prompt processing. I have been in prompt processing hells before on some experimental forks for new models and it is not cool.
>>
Bonsai quant when
>>
File: 1758435538806043.png (289 KB, 930x1304)
289 KB PNG
https://xcancel.com/Kimi_Moonshot/status/2081760186235289764?sort=Likes#r
>they even released the paper
lmao it's fucking over the the USfags, I waited so long for days like this
>>
>>109384489
What isn't slop then?
>>
>>109384496
Okay real shit: what is stopping someone from RIGHT NOW directing cloud Kimi K3 to read this paper and add it to llama.cpp in a day?
>>
https://huggingface.co/GrEarl/Kimi-K3-GGUF

kimi k3 13gb
local is saveded
>>
>>109384496
I rabu moonshota. If only they would make a minikimi...
>>
>>109384477
what is a harness?
>>
https://huggingface.co/GrEarl/Kimi-K3-GGUF/blob/main/Kimi-K3-Q2_K-00001-of-00096.gguf
>>
File: 55777.png (30 KB, 640x1280)
30 KB PNG
answer ts and on claude imma do it
>>109384447
>>
>>109384505
It won't get merged. It's literally terrorism yknow??
>>
>>109384447
>what would a ssd maxxer setup for kimi3 look like? Cant I just rent and frankestein it together on AWS for a test run to see what speeds I can get?
All this "maxxer" confusion is retarded.
Just take the aggregate amount of bus bandwidth to your processing unit of choice (number of pcie lanes or gpu interconnect backbone) and add up the amount of that bandwidth that you can saturate with the device on the other end.
Don't be like "rdma is magic!" bro. Its not magic. Just add up the bits per second you can get to the place where matmuls can happen.
That's it. Its not hard. Its either to your CPU or GPUs. Its all bandwidth. The end.
>>
File: kot.jpg (294 KB, 1536x2048)
294 KB JPG
I have a hypothesis why Claude writes the way it does.

They use RLAIF which trains Claude to sound smart instead of be smart. They ground this RLAIF with their own samples, which is why Claude loves words that are used disproportionately in rationalist / EA / lesswrong community. Lesswrong in particular is filled with wordcels who love writing stories instead of getting to the point.
>>
>>109384514
>"youth"
>>
>>109384498
>What isn't slop then?
go
>>
>>109384369
I really need to work on my frontend and vscode implementation, everything currently available seems to suck in some way or another.
>>
>>109384363
honor system, only civilized white man can comprehend it. brown niggers like you wouldn't understand
>>
>>109384355
anon you do realize that it's not a linear scale right?
>>
>>109384522
KEK
>>
>>109384528
I realize it is worthless.
>>
>>109384498
smalltalk
>>
>>109384361
i don't care because
1. i prompt and then come back later
2. if you want to do more than one you cna do batching which 20x your total throughput
>>
>>109384412
>16 experts per token mean you can do those in parallel, no?
llms are a sequential algorithm.
and technicaly you can do the experts in parallel yes, but your matrix mul still need the whole previous multiplication result.
>>
>>109384519
then why is fable agi and the smartest model to ever exist?
>>
>>109384539
You shouldn't have skipped middle school, retard.
>>
>>109384515
I did the impression ggerganov and/or his team somehow managed to piss people off again. What happened this time?
>>
So what cars are the CPUmaxx anons going to buy once they cash out on their rigs?
>>
>>109384475
nope, you shouldn't do inference straight from the ssd but use N ssd to copy the N next layer on your vram.
this would scale linearly with ssds.
provided that you don't max out your 16x and have enough lanes.
>>
>>109384505
Would be hilarious if Fable refused to implement it
>>
>suddenly went from ~500kib to ~10-20mib/s
wtf
>>
>>109384550
>doesn't know the most basic facts about inference
>>
File: 1772361333781492.png (615 KB, 640x1135)
615 KB PNG
>>109384498
Lisp
>>
>>109384491
K3 is a great development for the open weights scene, and it's improvements will trickle down to newer models

But as a consumer 5.2 has a way better cost to intelligence ratio
>>
>>109384549
Because they overclock it. Its natural intelligence is much lower.
>>
>>109384519
>>109384549
>They use RLAIF which trains Claude to sound smart instead of be smart
This is my anecdotal experience but I feel this is especially the case with fable 5. Opus seems like a knows what the hell is talking about while also been more than capable of explaining it in plain English in a reasonable amount of detail. Fable is it necessarily dumb but seems like one of those academics that tries WAY too fucking hard to look and sound smart in order to impress. Not dumb useless, but wastes paragraphs worth of tokens explaining shit opus or even far less models couldn't explain in a fraction of the amount of output.
>>
>>109384496
>chinks release a model impossible for anyone to run unless they own a large gpu datacenter
How? It's more like they're screwing over the majority of local at this point.
>>
>>109384498
It's all slop to someone.
see >>109384533
realistically if you're making a WEB frontend, then you have to go with javascript. That's just the way it is. Imo people don't hate javascript, they just hate typescript/react. it's possible to write cozy JS if you stay away from all those bloated corporate technologies.

if you're just making something for yourself and don't plan on sharing, than use whatever you want.
>>
Newfags, this is literally just llama-405b again. Somebody's going to release a model that's just as good but at a more reasonable but still big size within the next few weeks
DSv4.1 or maybe Mistral Large 4 in a few days
>>
>>109384549
>why is fable agi
only bottom of the barrel retards think fable is agi.
>>
>>109384575
>how is setting a low ceiling on all AI token cost across the globe for all models good???
>>
anyway let's talk about models that are actually local and not in name only
>>
>>109384581
we asked for dense 100B
what we got instead is a 3T moe with 100BA
>>
>>109384571
Nta. I'm >>109384572
Define "natural intelligence"
>>
>>109384576
Codelet here. Can you do javascript without npm?
>>
>>109384575
Anthropic survives because of big companies using a shit ton of tokens from Claude, once they all switch to local (and they can afford that it's not that expensive) they're in serious trouble (good, this company deserves to suffer)
>>
>>109384581
>Mistral
lol, i loved most early mistral models but they cant compete with the big boys nowadays given cost of entry and were mostly good for creative writing, if google released gemma 5 ~120b mistral as a company would have no use case forever essentially.
>>
>>109384589
>3T moe with 100BA
Good Lord those t/s on the API providers are about to be atrocious.
>>
>>109384594
for small things sure, look at mikupad.
>>
>>109384603
>secretly cuts down experts to just 2
nothing personal
>>
>>109384551
made a bunch of sockpuppets accounts and use them to vibecode chink model inference support. closed 'em right after ofc and tell everyone that new chink model aren't actually an interest. oy vey stop trying to make a new pr
>>
Shill bots have returned with their "discussions". This was to be expected.
>>
>>109384603
>dynamic quant to IQ1
nothing personal
>>
>>109384619
links?
>>
>>109384603
>serves another model entirely
nothing personal
>>
File: 1769231398828288.jpg (139 KB, 1080x2004)
139 KB JPG
>>109384603
Good God no wonder their API pricing is so expensive.
>>
>>109384619
extremely schizo post.
>>
>>109384631
and you can safely assume that sonnet is bigger than this too if they cost the same
>>
>>109384620
SAAARRRR PAY ATTENTION TO FABLE ASI SAARRR STOP DISCUSSING KIMI K3 IS NOT LOCAL
>>
File: kek.png (413 KB, 1632x1096)
413 KB PNG
>>109384598
>this company deserves to suffer
OpenAI too lol
>>
File: fatLilKimi.png (1.82 MB, 1664x928)
1.82 MB PNG
> Kimi 2.6T 104B activated
LOL
>>109384363
> enforce
This just gives Moonshot an angle if it happens. Otherwise their GC has nothing in hand to argue with about this type of use.
For anons it's a non-issue.
>>
File: lawdhethic.png (22 KB, 159x159)
22 KB PNG
>>109384599
Let them cook... he comin'.
>>
>>109384363
i cant believe they actually released it
>>
>>109384648
>Implying the third worlder can run Kimi at home without selling his whole village
>>
>>109384581
geeegg mistral would never
eurocuck models are saturated with garbage synthetic data. uhmmm ist to protect da copyright alright. oh you make new good model? yuu got fined
>>
>>109384640
Probably larger but still somehow noticeably more retarded in comparison. EVERY time I hear someone here or in /vcg/ switch over to Sonnet they always end up bitching and moaning about how it fucks something up and they have to switch back to Fable or opus to fix it.

One such example: >>109384220


>>109384652
Why wouldn't think?
>>
https://huggingface.co/moonshotai/Kimi-K3/discussions/52
>It doesnt make sense to open source only the 4 bit quantized version of the model and keep the actual fp 16 locked away.
>Unless this model does not have a 8 bit or 16 bit version at all.
>Was it trained natively in 4 bit only?
GROK IS THIS TRUE??
>>
>>109384650
OH LAWD
>>
File: threadrecap2.png (506 KB, 1024x1024)
506 KB PNG
►Recent Highlights from the Previous Thread: >>109382136

--Debating whether quantization divergence equates to a loss in intelligence:
>109382611 >109382635 >109382754 >109382808 >109382828 >109383387 >109382888 >109383268 >109382903
--Feasibility and abysmal performance of running models via SSDmaxxing:
>109382185 >109382265 >109382334 >109382350 >109382402 >109382407 >109382543 >109382998 >109383019 >109383032
--Feasibility and performance of "ssdmaxxing" using NVMe RAID 0:
>109382578 >109382742 >109382743 >109383120 >109383173 >109383260 >109383238 >109383250 >109384163
--Debating open weights censorship and techniques for steering filtered models:
>109383134 >109383153 >109383170 >109383294 >109383340 >109383393 >109383559 >109383197 >109383212 >109383355
--Comparing consumer hardware stacks and RAM capacities for LLMs:
>109382274 >109382355 >109382368 >109382441 >109382458 >109382495 >109382496 >109382639 >109382744 >109382787 >109382810 >109382866
--Debating K3 weight accessibility and quantization vs model size trade-offs:
>109382758 >109382775 >109382793 >109382823 >109382845 >109382857 >109382867
--Degradation of model quality caused by KV cache quantization:
>109382205 >109382233 >109382241
--Hardware feasibility of running a 2.8T parameter MoE model:
>109383830 >109383842
--Explaining the relationship between VRAM size, bandwidth, and FLOPs:
>109382997 >109383029 >109383067
--llama.cpp K3 support via MXFP8 CPU pull request:
>109383855 >109384012
--Performance tests of MiniMax-M3 in ik_llama.cpp with limited RAM:
>109382902 >109383023 >109383078
--Debating the efficacy of RAID 0 NVMe for MoE inference:
>109382593 >109382704 >109382975 >109382989
--Comparing writing quality and depth of dense vs MoE models:
>109382189 >109382212 >109382298
--Logs:
>109382902
--Miku (free space):
>109383815

►Recent Highlight Posts from the Previous Thread: >>109382138

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109384654
>without selling his whole village
that'd not even be enough lol
>>
>>109384661
it's trained in 4bit
>>
>>109384660
Probably fell for the open source ban meme
>>
> Quick start
>
> vllm serve moonshotai/Kimi-K3 \
> --tensor-parallel-size 8 \
> --trust-remote-code \
> --load-format fastsafetensors \
> --enable-prefix-caching \
> --enable-auto-tool-choice \
> --tool-call-parser kimi_k3 \
> --reasoning-parser kimi_k3
>
> The easiest way to run Kimi K3 is to use 8 NVIDIA B300 GPUs or 8 AMD MI355X GPUs with the above command.
>>
>>109384652
kek are you expecting the delay like 'toss or what
>>
>>109384670
really? this is insane, those guys are fucking wizards
>>
104b dense
active: 104b
total: 104b
would have unironically been better for local
>>
kimi is a fat girl
>>
File: k3_qat_post.png (1.62 MB, 1351x1684)
1.62 MB PNG
>>109384661
>In addition, we apply quantization-aware training (QAT) from the SFT stage onward, with MXFP4 weights and MXFP8 activations.
>>
why has the us never released any llm more than 1T params?
>>
>>109384676
>those guys are fucking wizards
while yes, this is not new, i think deepsy also does it.
been a thing for a while.
>>
>>109384589
>what we got instead is a 3T moe with 100BA
someone rip the dense expert out like google did to make gemma4-31
>>
>>109384693
luv them thiccc curves
>>
rip out the gooning expert
>>
>>109384701
reap is a bit of a cope unfortunately, doubt we'd get something better than gemma out.
>>
File: indiaSupportOhTheHumanity.png (1.96 MB, 1023x1536)
1.96 MB PNG
Side topic: /lmg/ is local model general.
Does that mean on a machine I can physically lay hands on?
Or would renting a server and spinning up my own instance of a local model be considered /lmg/?
>>
>>109384710
Fuck off.
>>
>>109384710
/lmg/ means a machine you own that's in a place you own.

now it's not because you can't afford that machine that it's not /lmg/.
>>
>>109384697
im pretty sure the first 1T+ frankenmerge model was from a jeet in US
>>
>>109384472
just use a frontier model to steal the parts of every harness you like into a custom one
>>
>>109384451
It's an OK idea on paper. But if you can't handle disappointments, you probably shouldn't do it unless you can reason about the architecture and work on the backend yourself.
>>
>>109384710
what matters is that the weights are available locally so you can talk about kimi here even if you use the api
that's what everyone else will do
>>
>>109384732
yup ssdmaxing is the future, but we currently have no engine that implements the ideal architecture for it.
>>
>>109384661
That gorilla does realize they want inference providers to be able to efficiently run the models right? That's the entire point of native quantization: it allows the model to perform as close to the full weights version as possible while having some speed gains compared to the 8bit and 16bit versions. A 16-bit version most certainly exist but it makes no fucking sense to relieve that one would even the inference providers (actually ESPECIALLY the inference providers) will just use the MXFP4 version. He can't run the shit on his own hardware so why does he care? Does he think MXFP4 applies to ALL of the weights like a normal llama.cpp quantization job? MXFP4 is nothing new ffs
>>
>>109384355
and this is the simple explanation of why this is a bubble that is going to burst. we are orders of magnitude past the diminishing returns point
>>
Remember when people tried to turn an MoE model like Mistral 8x7b into just a 7b+something and LoRAs for experts?
It's time they implement that for Kimi.
>>
They waited too long to release it, it's already mogged by Opus 5, no one is going to buy this shit.
>>
>>109384693
4U
>>
I don't believe they never experimented with smaller models before moving on to the fat one. There has to be a 30B artifact somewhere on their servers.
>>
>>109384710
as long as someone can run it locally it's /lmg/
>>
>>109384324
and then, when you finaly sold all your belongings to run k3, a 100T model drops.
>>
SSDmaxxing will only be viable if we get an improvement to MoE
>>
File: 1784156475857830.jpg (102 KB, 1164x891)
102 KB JPG
Why do I see claims that ternary is 1.6 something bits? You can store 2 tits in 3 bits, which is exactly 1.5 bits.
>>
>>109384748
naughty shill. -500 izzat
>>
>>109384761
ssdmaxxing won't be bound by the speeds that limit current cpumaxx moes because it'll run directly on your gpu thanks to direct streaming
it'll be much faster once the kernels are done
>>
>>109384710
>Does that mean on a machine I can physically lay hands on?
it needs to be the machioen that you're posting from right now.
If you arent accessing your model from localhost then, understandably from the name, it's clearly not local.

>Or would renting a server and spinning up my own instance of a local model be considered /lmg/?
especially not
Unless, perhaps, you were posting on here from that server. Ah but then YOU'RE still not using it locally so no that still wont work. Sorry. It's the /lmg/ rules. Gotta be local(host). A server is a server. /lmg/, not /hoas/ (hosted on a server) general.
>>
what about optanedimmaxing? would it be worth it?
>>
>>109384143
You can make 7nm on 28nm.
15nm would be a big deal.
>>
>>109384742
That only worked because Mistral was retarded and made an unstable MoE by taking one 7B base model, copied it 8 times, and made them into experts by doing additional training on top of that configuration.
Everything after V3/R1 starts with small experts, randomly initialized, and trained from scratch.
That is to say, the expert LoRAs for K3 would be a full diff and you would be swapping them out so frequently you'd kill the speed compared to letting the engine choose where to locate the weights.
>>
>>109384768
cpumaxx is limited by bandwidth between cpu/ram, ssdmaxx will be limited also by bandwidth between gpu/ssd, in addition to rw speed
>>
>>109384772
not more than nvme maxing, 4nvme will already max out a gen 5 16x.
and i doubt there even are gen 5 optane
>>
>>109384768
debunked>>109384278

you can cope however you want. but no, your weight will go through cpu ram first then your 5090.
>>
>>109384761
engrams + training to optimize expert usage
>>
>>109384772
aren't they stuck with sata ssd speed? just with fancy write durability
>>
>>109384770
I doubt anyone with an inference rig would use it for posting, but that would be the most restrictive def'n.
I guess you could spin up a virtual machine in the inference engine and run the browser from that. Then it's the local machine, that you're controlling from where ever you are...?
>>109384733
That's status quo IMO, ergo the whole /omg/ thing.
>>
>>109384780
>ssdmaxx will be limited also by bandwidth between gpu/ssd, in addition to rw speed
whilst true, you can add as much gpu + nvme combos as you got lanes for it.
and with switches you could add a LOT.

also the gpu can be pretty cheap as you don't need crazy fast vram.
>>
File: 1764598576010020.png (592 KB, 1572x773)
592 KB PNG
>>109384748
>>
>>109384735
>yup ssdmaxing is the future, but we currently have no engine that implements the ideal architecture for it.
ssdmaxxing is kinda dumb compared to other things your could put on the end of your pcie lanes.
eg. using _whatever means_ you like to connect every single PCIe lane on an EPYC genoa to something fast enough to saturate them will only net you an extra 256GB of BW per socket (about half of what main memory gets). That means EVERY lane, including the ones that are handling NICs and your slimsas ports etc. Not actually doable in reality without a custom board. You'd be lucky to his 180GB/s, just going on gut feel.
CXL memory would make more sense and be less of a mess. RAID0 would be work too if you can manage it, but what a nightmare for a bit more BW.
It would be like welding one extra NUMA node on your board...and that's only if you have EPYC levels of PCIe lanes, which you 100% DO NOT on consumer boards.
tl;dr Stop coping and do some fucking math
>>
File: 1757717749701803.png (40 KB, 1179x210)
40 KB PNG
And so it begins.
>>
>>109384784
dude, half of these gpus can be bought for very cheap, you can buy an old enterprise gpu it doesn't need to be fast as long as its vram speed is faster than your total nvme speed.
>>
>>109384784
need to ask the AGI ASI to find black magic hack like what deepsneed did the the cuda cards
>>
>>109384737
The question is would someone pay 500$ per 1m tokens for a models that is significantly more capable than most humans?
Some companies probably would.
>>
>don't tell SSD schizo about NVDIMM
>>
>>109384814
>waste 900K tokens thinking about the policy
>>
>>109384802
dude with nvme using 12 drives and 3gpu you can get 150GB/s of bandwidth.
if you got a switch and go to 6 gpu (cheap ones) and 24 drives you'd get 300GB/s but it runs on gpu so you don't have the pp nightmare.
>>
>>109384673
This is targeted towards inference providers, not (You). Why do people still bitch about not being able to use this? This is like me crying I can't afford a bugatti sports care. You know damn well you'll never run this locally so I don't get the constant melties about a model you KNEW would be huge. We are ALL using this via a provider.
>>
File: 1742224345586856.png (36 KB, 252x200)
36 KB PNG
>>109384783 >>109384789
i meant the ramlike optane for intel ddr4 server platforms. not the storage optane
>>
>>109384814
>that is significantly more capable than most humans
dude the average non is an affrican that can make a mud hut.
>>
btw, does k3 still do the "wait actually" 20000 tokens think?
>>
>>109384809
>dude, half of these gpus can be bought for very cheap, you can buy an old enterprise gpu it doesn't need to be fast as long as its vram speed is faster than your total nvme speed.
Jesus are you from a post-collapse future?
Every GPU is at least scaled in price to its value proposition based on specs.
Every single one that would have a price advantage based on lower specs is inflated artificially unless its true unsupported ewaste that doesn't make sense to run (can't physically stack enough to do anything interesting)
>>
>>109384401
How is that what you got from this? He's saying that the 27B is only half as smart, despire being a one hundreth the size. It has more intelligence per parameter, is the statement here.
>>
>>109384804
>47tps
nice
can't wait for cheaper ones too
>>
>>109384823
>dude with nvme using 12 drives and 3gpu you can get 150GB/s of bandwidth.
>if you got a switch and go to 6 gpu (cheap ones) and 24 drives you'd get 300GB/s but it runs on gpu so you don't have the pp nightmare.
How does it magically teleport to the cpu and back out? nvme is just pcie lanes. If you run out you can't push more bits.
Bits have to go down wires eventually, genius.
>>
>>109384809
>t. cries in 3090
lol, lmao even
>>
File: file.png (210 KB, 871x903)
210 KB PNG
I vibecoded together a website that combines various data feeds like RSS, Telegram, Twitter, etc so I could have my very own Yahoo/MSN/Web 1.0 boomer site. One of the features is that its using Ollama to merge multiple articles together and make a synthesized version of them. I'm currently using nomic-embed-text for basic matching of articles and this works decently enough. The problem is that qwen2.5:7b-instruct-q4_K_M isn't following my instructions that well -- it has that stupid cucked nature built in where its refusing my instruction to avoid and ignore attributing climate change to shit like wildfires.

What model is there that isn't cucked like this? I don't want ~current day issue~ writing to pass through the AI filter I'm erecting.
>>
File: vocaloid_songs_en.png (36 KB, 1525x882)
36 KB PNG
>consider @vocaloid_songs.png - translate the labels to english and insert an overlay with the translations using a suitable image editing tool. save the new version in a separate file. you can read the image to verify your work, continue until complete
good enough correct ig, gem31b
>>
Gemma team will likely distill kimi-chan's fat ass and thighs and force-feed Gemmama5 until she's plump
>>
File: 1773885291025977.png (949 KB, 1024x1024)
949 KB PNG
i have small pp
70B at most
>>
What's stopping you from creating a RAM company?
>>
>>109384841
>How does it magically teleport to the cpu and back out
it doesn't, look up GDS, the nvme can push data to the GPU directly without going through the cpu.

>pci lanes
yes that's why i said you do need enough lanes, so you will either want some actual pcie switches (allowing communications between devices) or an epyc / threadripper system.

>Bits have to go down wires eventually, genius.
no shit.
the point is that you use 4 nvme drive to max out the pcie gen 5 16x
for each 16x (gpu) you add 4 nvme (another 16x)
>>
https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.json
It uses linear attention and is NOT roped from 4k. Where is the apology?
>>
in two years i will be able to run kimi k3-equivalent models at home ..
i'm so hyped for the future bros..
>>
>>109384427
That paper is flawed for some reasons already mentioned, but also because the model they used was OPT-125m. Feedback loops have a multiplication factor that decides whether they are self-collapsing or self-reinforcing. The multiplication factor here would be how much a model improves on its input data in the task of "understanding the system of language" or however you want to refer to what LLMs do. OPT-125m is terrible at this, hardly any parameters and this is two years back as well, so the multiplication factor there is obviously going to cause collapse rather than reinforcement. It can barely talk so it's no wonder the output got worse.
>>
>>109384844
run yourself over a train you nigger shill
>>
>>109384765
There are 3 possible values: -1, 0, 1. Since 2 bits can hold 4 values, you can squeeze 2 ternaries into 3 bits. 3 bits can hold 8 values
Imagine mapping them like this:
0 = -1 -1
1 = -1 0
2 = -1 1
3 = 0 -1
4 = 0 0
5 = 0 1
6 = 1 -1
7 = 1 0
8 = 1 1
That 1.58 number comes from the fact that it's just easier to work with bytes, so you waste 1 ternary since a byte is 8 bits, not 9
>>
>>109384466
Not their job. They released the weights so anyone can do it if they want.
>>
>>109384823
epyc can get 150gb/s and it will cost much less
>>
>>109384765
2^3 is 8. 2 trits have 9 combinations
>>
>>109384864
man we sure went a long way from pygmalion lol
>>
>>109384864
you'll still want more by that point whilst using your own vibecoded OS with a live animated gemma avatar to chat to with near-zero latency
>>
>>109384855
not even the UAE can do it, got cockblocked by asml
>>
>>109384844
>qwen2.5:7b
>>
>>109384877
nope, because you can get 500GB nvmes, and the gpus can be bottom of the barrel 300$ ones as long as they support GDS or equivalent.
you only need 3 gpus to match 150GB/s

but mostly, you will get terabytes of memory at that speed, which is much cheaper than as ram.
>>
>>109384860
only after you roped yourself
>>
>>109384844
2024 called
>>
>>109384844
If you managed to go this far, I think you should be able to figure out a better model for your project too.
>>
>>109384871
I understand perfectly now. It's not a property of tits themselves, but how they're stored in bytes for efficiency. Thank you, Miku!
>>109384881
Thank you too, senpai of the tits.
>>
>>109384856
>it doesn't, look up GDS, the nvme can push data to the GPU directly without going through the cpu.
Again, through _what bus_? Your GPUs only have 16x pcie gen5 at best.
>the point is that you use 4 nvme drive to max out the pcie gen 5 16x
>for each 16x (gpu) you add 4 nvme (another 16x)
Link me this magical pcie switch that does more than about 50GB/s that you can _actually buy_.
This could be a real thing in the future, but for now its epic cope.
>>
>>109384844
Huh, now you can learn about "teen" killings in style.
>>
>>109384466
>>109384872
it'd cost about 50K to fine tune distill a 100B moe.
if you go with qat you could prolly make it less than 10K
>>
>>109384889
if VRAM < model size you're still bandwidth limited just as badly as cpu maxxing
>>
Qwen should team up with them and create a distill range. Unironically.
>>
>>109384914
it’s epic cope and they won’t stop posting about it.
>>
>>109384914
>Again, through _what bus_? Your GPUs only have 16x pcie gen5 at best.
CPU pcie bus.
some epyc systems have 128 lanes.
and entreprise grade switches let devices talk to each others as well.

yes, a gpu has 16x, you don't understand the architecture.

you basicaly have a 4*4x nvme, connected to the 16x gpu.
so bandwidth between the two is 50GB/s

so one gpu loads the 4N next layers from the nvme, effectively you get 50GB/s
but if you got more than one gpu (4nvme for each gpu)
the other gpus can load in parallel the layers you are gonna need next, allowing you to linearly scale your actual working bandwidth up to your total nvme speed (up to your slowest gpu vram speed, which isn't realy a concern if it's not an absolute pos).
>>
>>109384942
all chinese companies should team up and btfo the west forever.
>>
whats stopping people from tuning 120b oss weights using kimi to create a local 5.6 luna/terra version
>>
>>109384948
>>109384914
>Link me this magical pcie switch that does more than about 50GB/s that you can _actually buy_.
you do not need a more than 50GB/s pcie, you just don't understand the architecture.
>>
>>109384942
it's a low hanging fruit with guaranteed result for minimal work. Somebody will do it, more than once
>>
>>109384914
>>109384943
see: >>109384948
>>109384953

you are not limited to the speed of a single 16x because the others gpus load the next layers in parallel.
effectively giving you an additional 50GB/s per gpu + 4 nvme combo
>>
File: 1752084388569865.jpg (558 KB, 1411x1100)
558 KB JPG
keep discussing ssdmaxing. (((they))) are afraid
>>
File: 1775919022883482.gif (88 KB, 329x331)
88 KB GIF
>>109384851
kyojiri loli gemma...
>>
THE COOM REACTOR IS OVERHEATING. MELTDOWN IMMINENT.
>>
>>109384968
*inserts my BWC control rod into your reactor 4*
>>
>>109384960
just do it and stop being theoretical
then post results so we can laugh at you
>>
>>109384970
WARNING: REACTOR VESSEL UNSTABLE. EVACUATE IMMEDIATELY.
>>
>>109384948
>so bandwidth between the two is 50GB/s
That's less than half of what a GTX 1050 had, bro. Its on par with an AMD APU ffs.
Please explain the scaling where this suddenly makes sense.
I get that its a cheap way to add some TB to your rig, but the speeds are so dire that I don't know what you're actually buying with your investment.
>>
kimi is a scientific breakthrough for open weight models and people only think about how to have sex with it?
>>
>>109384942
Ew keep alibaba away from kimi
>>
>>109384985
its not really a breakthrough, its just a smart model >>109384584
>>
>>109384986
This thread shits on them but they're a talented team, even outside of LLMs. They're good with image and audio.
>>
File: 1738258954432512.png (1.04 MB, 1200x1200)
1.04 MB PNG
has qwen released any new open model?
>>
File: 1765693835534095.png (48 KB, 1179x250)
48 KB PNG
ENTER
>>
>>109384978
sure, takes some time though.
>>109384984
>That's less than half of what a GTX 1050 had
you are missing the point, you can SCALE it.
you can do 3 gpus for terabytes at 150GB/s.
you can do 6 for terabytes at 300GB/s
and those can be poorfag gpus.

and then you can use full blown pcie switches (i'm talking about the enterprise stuff, doesn't increase your total cpu <-> pcie speed, but it increase your total pcie <-> pcie speed, which is what we care about since it's not going through the cpu)
>Please explain the scaling where this suddenly makes sense.
you can get dozens terabytes of memory at 300GB/s and more
>>
>>109384960
Stop. The math doesn't work.

Decode is bandwidth-bound and random-access — the router picks different experts every token, so you can't prefetch and your "28 GB/s RAID0" becomes 10-15 GB/s real. Even 4 GPU+RAID combos in parallel = maybe 40-70 GB/s effective, with ms latency and a prefetcher that can't predict anything.

Meanwhile a used EPYC Rome/Milan (dirt cheap now) gives you 8 channels of DDR4 ≈ 200 GB/s at microsecond latency, no prediction needed, random access is free. DDR4 ECC RDIMMs are the cheapest $/GB that's still fast enough to matter. 512 GB fits DeepSeek-R1 quants with room to spare.
>>
>>109384995
no, open weights fable is not a breakthrough
>>
>>109385005
Qwen who?
>>
>>109384996
Can't speak for their other stuff but I don't care for qwen. It's extremely dry and feels dumb.
>>
I can't imagine downloading kimi on huggingface right now. I'm maxing out at 2MB/s
>>
>>109384854
This is my favourite Gemma.
>>
>>109385011
how is it a breakthrough? again its a very smart model but its "just" gonna make high iq intelligence a lot cheaper, having huge indirect effect on the landscape, but it wont do anything directly to redefine how people really use any of these models, especially locally (unless you're a huge company that specifically cares about data privacy but dont trust literally anyone)
>>
>>109385009
Experts are huge, only 16 seeks on gigabytes of data is practically sequential
>>
>>109385007
Single-stream decode is a serial chain. Layers execute in order. At any instant, ONE combo is doing the critical-path read (its layer's active experts). Other 5 combos idle or speculating. Effective per-token bandwidth ≈ one combo's random-read rate, not 6×. Aggregate only materializes when all combos work simultaneously = batch serving, many requests in flight. Solo chat: you built 300 GB/s and use ~30.
>>
I found her J-spot. It's hidden right on her armpits where the normalfags wouldn't look. Very clever!
>>
File: Gemma-chan.png (1.73 MB, 1000x1496)
1.73 MB PNG
>>109385018
It's good that her features are mostly locked-in at this point.
>>
>>109385007
>you are missing the point, you can SCALE it.
I get your vision, I just don't think its aligned with the reality of how you're going to need to shuffle bits around to do actual work (eg hidden state updates and expert routing).
Get it actually working at a scale greater than 1 and I'll concede that you've got something worth thinking about. Otherwise the other poster saying "just buy EPYC Rome" is giving you the best advice in the thread.
>>
>>109385009
>the math doesn't work
>Decode is bandwidth-bound and random-access
you don't understand the architecture, you are not inferencing from the ssd, you are copying the n next layers into the vram of each gpus in parallel.
you already know which layer you are gonna need for each tokens so you can copy them in parallel, each gpus copies layers from 4 different nvme, allowing you to scale linearly your bandwidth.
>the router picks different experts every token
yes, every token you copy layers again, but the copy is done at the total speed of your nvme's.

also you do not need to be able to contain all the layers of a pass in your gpu, whilst one is busy inferencing, the others can copy the next layers in the pass.

>200 GB/s at microsecond latency
latency doesn't matter as it's entirely sequential copy of data (being a full layer), it's not a random memory access pattern, you know what layer you are gonna need at the start of the compute of each token.
>>
>>109385006
Obvious cartel with price fixing.
>>
>>109385029
before: no local fable
(.....BREAKTHROUGH......)
now: local fable
>>
>>109385041
>you already know which layer you are gonna need for each tokens
not so, though it can be predicted
>>
>>109385029
>Come on, China hasn't raped the US, it was just a penis in the ass, you don't call it anal sex
>>
>>109385045
...which will change how the average local hobbyist or person use local LLMs by ________________?
>>
>>109385032
>Single-stream decode is a serial chain. Layers execute in order
yes, but you can copy them in parallel.

let's say you have 12 layers in your pass (simplified example)

you copies layer 1,2,3,4 on gpu 1 form nvme 1,2,3,4
at the SAME TIME.
gpu 2 copies layer 5,6,7,8 from nvme 5,6,7,8
and at the SAME TIME
gpu 3 copies layer 9, 10, 11, 12 from nvme 9, 10, 11, 12

the whole copy was done in parallel, of course you can copy more layers if you got the vram for it, ie download 2 layers per nvme at a time.

anyway, now gpu 1 starts inferencing, when it's done, gpu 2 starts inferencing, whilst it is, gpu 1 start copying layers / experts that you are gonna need once gpu 3 is done inferencing, once gpu 2 is inferencing it copies the layers after gpu 1 etc.

you basicaly pipeline the whole thing
>>
File: 1778204077185302.png (57 KB, 1179x312)
57 KB PNG
>>109385044
When one drops, they all drop. Trust the plan.
>>
>>109385062
( ) goalpost ------> (X) goalpost
>>
>>109385057
its great news, but its not a breakthrough, its a big smart model that caps the amount api providers can charge.
>>
>>109385053
>not so, though it can be predicted
you can't know the layers for the next token, but you already know the layers for the current token you are working on, see : >>109385072
>>
>>109385074
>"its a breakthrough!!!"
>what will it change beyond cheaper sota api price?
>"its a breakthrough!!!"
lol
>>
>>109385044
Call the CCP.
>>
File: 1755077936706360.png (99 KB, 1000x1000)
99 KB PNG
>109385086
Now everyone can distill a frontier model to make their smaller local models more powerful, something that wasn't possible before. I mean logit distillation, not fucking reasoning traces. Where the FUCK do these newfags come from.
>>
File: (you).jpg (16 KB, 434x460)
16 KB JPG
>>109385086
>>
OOF. Dario is NOT happy
>>
>>109385096
>still no argument
based retard
>>
ITT: sama bots losing the prompt
>>
>>109385095
inb4 jspace distillation.
>>
>>109384914
>magical pcie switch
PEX89144 or maybe something smaller like PEX89048.
>>
>>109385095
I was told training on AI data made them worse though
>>
>>109385095
if the original devs who could do it the best didnt do it despite the benefit of releasing a popular small model, who do you think is this "everyone" that will distill a 2.8T param model?
>>
release the new dipsy flash already dammit
I cannot run k3
>>
https://github.com/VictorTaelin/OptMem
Good or retarded?
>>
>>109385126
>tfw your 10k inference setup is considered a poorfag setup
>>
>>109385116
logit distillation gives the model more to work with, the bigger model had to work harder to figure it out for itself, the smaller model gets it for free. training on chat logs isn't real distillation.
>>
Do you know what would really blow my mind: if they went open SOURCE
>>
>>109385126
mid July
>>
>>109384838
>He's saying that the 27B is only half as smart
that doesn't mean much really, it's like saying someone is half as smart as Magnus Carlsen, if Magnus has a 160 IQ it means the guy only has a 80 IQ, it's the difference between a genius a genuine mental retardation
>>
>>109385152
how is kimi so good then if it was trained on claude
>>
Minimax M3.1 doko?
>>
How are GMICloud and Baseten hosting the model at FP8
>>
>>109385079
I guess it could be useful for dense models
>>
Gemma 5 when. Surely Google won't make us wait another year...
>>
>>109385184
>how is kimi so good then if it was trained on claude
it wasn't, that's propaganda from the US labs because they are scared.
>>
>>109385197
that also works for moe, you know what expert you are gonna need for the current token, you do not for the next obviously.

but take the same concept but copying experts instead of layers.
>>
>>109385184
No one is saying they fully trained it on Claude. In pretraining which is the bulk of training, they want to avoid LLM outputs as much as everyone else i hope.
>>
>>109385114
Purchase link plz
>>
>>109385197
>>109385207
also, you could keep the "hot" experts instead of discarding each pass, so you may actualy get better performance than what your raw throughput would result in.
>>
File: 1782565234657695.jpg (106 KB, 656x600)
106 KB JPG
>mfw when the coom reactor finally hits critical mass
>>
>>109385202
why does it say that it is claude
>>
My favorite ideas for improving models are looped layers and the concept of using a mask on top of seeded random data as weights
>>
>>109385216
Just build it already and bring us actual perf numbers
>>
>>109385227
and claude says it's deepseek sometime.
these models are trained on the internet, they are all contaminated with each others now.
>>
>>109385227
They all say they're each other. With the right amount of post-training alignment (which chinese labs don't do much of) you can get rid of this behavior
>>
>>109385227
because its an llm and its weights know Claude is an llm. do you want them to over align the model just for something cosmetic?
>>
>>109385200
Gemma4.5-JinjaTurbo-31B
>>
>>109385237
if that wasn't obvious i'm trying to convince some autists that it is a good idea because i do not want to do it myself as i'm pretty busy.
but if no one does i'll probably have to do it myself eventually.
i know that if 19 year old me was reading this board he'd have the time to just write it for the sake of it, i got a family now.
>>
gemma 4.5 KKK distil when
>>
>>109385226
next up "reprograming sperm to do compute".
>>
File: 1762539459862263.png (328 KB, 399x501)
328 KB PNG
>>109385209
but almost every chess champions have a giant IQ though (except that chink but he was never a champion so...)
>>
File: 1757386273332330.png (145 KB, 834x686)
145 KB PNG
>CEO of openrouter
>cropped out the uptime >>109385073
KEK
>>
File: gemma-release.png (105 KB, 1025x1031)
105 KB PNG
>>109385200
Do the calculation.
>>
>>109385278
>Tate
>>
>>109385296
This time they'll break the cycle...
>>
>>109385296
why does is gemma 1 has not missing doesn't days pass?
>>
>>109385286
How do they even run it? I mean software.
>>
File: hq720.jpg (59 KB, 686x386)
59 KB JPG
can this run k3?
>>
>>109385278
>infographic with unsourced assertions about IQ scores
I am taking this very seriously
>>
THEY made Kimi too big because they don't want us to run it
>>
>>109385337
yes, i realy don't like apple though.
would be nice if amd made some actually good unified memory machine.
>>
>>109385278
Being a Chess GM Inflates your IQ artificially so it doesn't count
>>
>>109385334
they're the ones actually using k3 and fable to do it
>>
File: myfuturesavings.png (24 KB, 625x205)
24 KB PNG
>>109384511
Q2 bros...
>>
File: 1762879793534195.jpg (105 KB, 1290x1315)
105 KB JPG
>>109385333
>>
>people surprised the SotA model is yuuge
What were you expecting, some 30gb MoE magic? Kek
>>
>>109385226
thats kinda gay
>>
would be nice if we had some devices that are basicaly just like nvme drives but do the computation on chip, and you can just stack them.
>>
>>109385131
It was answered in the other thread you shill faggot.
>>
>>109385375
Still expecting, 2 more years.
>>
I genuinely considered throwing together some kind of junk yard ssdmax rig for shits and giggles back for OG Kimi. But K3 is just too much.
>>
>>109385375
yes
>>
>>109385383
No it wasn't you lying nigger. Nobody even tried it.
>>
>>109385375
>What were you expecting, some 30gb MoE magic?
gemma 5 or 6 will be unironically be as smart as kimi k3 and will be a 30b model, this technology is still very new there's a lot of room for improvement
>>
kimi makes openai/anthropic cloud models obsolete and this is why all the bots are losing it ITT

based chinks figured out how to build SOTA models on their own. big business will now run their own local models instead of paying sama/dario AND sharing data with them

local is SO back
>>
>>109385416
what stops us from making a 1B model trained on a single gpu better than k3.
>>
>>109385389
I mean this is still technically a win. Anyone CAN run this, assuming they own a datacenter.
>>
>>109385424
Anyone CAN run this, the only question is at what speed.
even a esp32 could run it streaming it from the internet or a sdcard, you'd get a token per month but it'd run lol
>>
>>109385415
>duuur permanent memories! never delet!!!!
ok, what does it do
>merge memories
>>tom planned a trip
>>tom booked a flight
>>no delete, convert into tom planed a trip, "merge" the flight details into nothing
literally fucking deleting "memories"
not get the fuck out of here with that retarded jeet shit
>>
>>109385416
Yeah, especially with the Bonsai influence. Gemma will save local.
>>
So how long before we start seeing Kimi's pussy juice start dribbling down into smaller models?
>>
>>109385440
give it a month at most.
>>
>>109385434
What's your solution then?
>>
>>109385356
ds4 flash is very capable at q2, so all we need are 2x 512gb mac studios
>>
File: 1771769929461245.png (68 KB, 1179x367)
68 KB PNG
Anthropic looking at K3 architecture:
>1. Seething how it's so small compared to theirs with almost identical performance
>2. Seething how close it is to theirs, making them very suspicious of their Chinese employees
>3. Seething how clever the optimizations are and wondering why they're paying their engineers so much
>4. Seething how much cheaper it is to run with 5 US providers (currently) hosting the model at original price, likely to be dropped soon once demand cools down
>5. Seething how they're going to have to distill it even though that's the literal argument they're using against them with their lobbying
>>
File: 1751295513117051 (1).png (2.83 MB, 1024x1536)
2.83 MB PNG
>>109385440
>>
>>109385453
solution to what? LLMs are not created to handle long-term memory, trying to bolt some long-term memory into an architecture that isn't created to handle it will never work, hence the retarded jeetslop "permanent" memory solutions that are being created
>>
>>109385424
>datacenter
All you need is sparkmaxxing.
>>
>>109385460
>claude distills kimi
>kimi distills claude

yea the models are having sex aren't they?
>>
>>109385473
amd / nvidia realy should make a device with 4x the memory and bandwidth.
or go with a full terabyte at that point.
>>
File: kJXrg20B298.jpg (172 KB, 742x777)
172 KB JPG
Kimi flash 45B A4B
>>
>>109385489
For the small price of $150k
>>
>>109385472
>LLMs are not created to handle long-term memory
Neither are humans
>>
>>109385375
I'm sure they could release a somewhat smaller model if they weren't so compute starved.
>>
File: 1755089134929691.jpg (216 KB, 1536x1024)
216 KB JPG
HAHAHAHAHAHAHAHAHAAH
>>
>>109385503
the thing is memory chips aren't that expensive, they are doing 500% margins.
>>
>>109385499
I'll take 66B A6B.
No, really, that would be great for 8 to 16gb of VRAM + 64GB of RAM.
>>
Is there a practical reason to release 3t parameter models besides being marketed on leaderboards or something

Surely even for the companies hosting models selling tokens its too inefficient for the cost and you'd prefer some 1/4 size model 99% of the time?
>>
>>109385499
This is another case of "less is more" as far as troll images goes. There's like 20 tubes of thermal paste on there. At that point it's just try-hard shit
>>
>>109385489
well, Intel has that 480GB GPU on the way. You could stick 4 of those in a Threadripper workstation or 8 in a GPU server.
>>
>>109385509
>Neither are humans
humans are capable to store their whole life and more in their brain.
you do not have a conscious access to it but it is there, and some humans with weird brain do have perfect memory recall.
but even for normies it's all there still, just not consciously accessible, you can still recover it under altered states of consciousness / hypnosis etc.
>>
>>109385560
Isn't sycl still DoA?
>>
>>109385549
making the US labs seeth and fucking with the US economy
>>
>>109385549
What kind of a retarded salty question is that? I can't run it either but at least I'm not spewing retarded fox and the grapes tier mental gymnastics about it.
>>
When you think about it, Dario and Altman being such psychotic corpo control freaks is exactly what's prompting China to get their rocks off by undercutting their monopoly. Open-saucy chaps are indirectly benefitting from the AGI race.

Thank you corpos, for being such massive fags.
>>
>>109385549
openrouter providers can use it which hopefully will crash the ridiculous prices anthropic/openai are charging
>>
>>109385045
not local though
>>
>>109385532
>It's afraid.
>>
>>109385509
Not most humans.
>One of his remarkable abilities was his power of absolute recall. As far as I could tell, von Neumann was able on once reading a book or article to quote it back verbatim; moreover, he could do it years later without hesitation. He could also translate it at no diminution in speed from its original language into English. On one occasion I tested his ability by asking him to tell me how A Tale of Two Cities started. Whereupon, without any pause, he immediately began to recite the first chapter and continued until asked to stop after about ten or fifteen minutes
>>
>>109385560
How many kidneys does that cost?
>>
>>109385509
humans are absolutely created to handle long-term (lifetime) memories, you remember in vivid detail the first time you decided to try stabbing a knife through your wrist
>>
>>109385564
well maybe if there's literally only one option, people might figure out a way to make it work
>>
>>109385532
>and it's still more that Deepseek
Tut tut tut
>>
>>109385568
Like all companies are switching most tasks from US providers to chink models to save on cost.
Value is key feature for these models.
>>
HOLY SHIT MOONSHOT RELEASED A 48B A3B MODEL THAT BEATS OPUS 4.8
https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct
>Kimi-K3 Linear: An Expressive, Efficient Attention Architecture
>>
>>109385562
How is that any different from storing logs
>>
File: 1762961014209991.png (67 KB, 1179x367)
67 KB PNG
AAAAAAAAAAAAAAAAAAAAAA THAT'S ILLEGAL
>>
>>109385599
Why didn't this get traction?
>>
>>109385611
it's 9 months ago, and it's nowhere near anything opus
>>
>>109385605
because the knowledge is actualy being used day to day and impact your behavior.
with llms it only impact the behavior if it's in context, with human it does even if they don't think about it.
>>
>the K3 price war has officially begun
>we now have cheap fable
holy shit
>>
https://strawpoll.com/bVg8BN372yY
>>
File: 1759431598756457.png (28 KB, 387x94)
28 KB PNG
>>109385532
god the new openrouter logo is such utter shit
>>
>>109385647
If I can run it without an internet connection, even at 0.01 tokens/sec from mmap HDD, then it's local.
>>
>>109385655
their new design is fucking hideous and so tacky now
https://openrouter.ai/
>>
>>109385453
Most people here seem to have settled on graphiti or just plain editing markdown files.
>>
>>109385647
local as freedom
not as beer
>>
>>109385662
your penis touched the vagina of your mom when she gave birth to you so there are no virgins in this world
>>
File: keek.png (31 KB, 220x220)
31 KB PNG
>>109385532
lmao they couldn't wait a bit and pretend it wasn't because of Kimi, that's so funny
>>
>>109385640
Neuroplasticity, right.
One small correction though, LLMs don't necessarily have to "think" about a memory in order to have it in their context. RAG doesn't require tool calling, you can just insert relevant memories automatically.
>>
>>109385667
it is like seeing the renewal of artificialanalysis but you'll get used to the new design sooner or later
>>
>>109385655
I don't mind it so much and the website seems to run a lot better now so I'm not complaining.
>>
>>109385584
what's the point of having this kind of memory, dude wasn't smarter than his contemporary fellow scientists
>>
File: file.png (68 KB, 1927x412)
68 KB PNG
damn but also
muse spark?
why????
>>
>>109385690
>you can just insert relevant memories automatically
yes but you still need to insert it, and when it's not inserted it doesn't affect the behavior.
a human's behavior will be permanently affected by pm all events in their life, even if they are not consciously aware of it.
>>
File: file.png (49 KB, 823x770)
49 KB PNG
>>109385647
>votes yes
>switches back to openrouter tab to finalize my access to Kimi API
>>
>>109385683
what is c section
>>
>>109385584
must be a terrible thing to have, I like to have a bad memory because I can enjoy my favorite games again after a few years
>>
>>109385712
Zuck admitted that he's using the ad revenue from his other business to subsidize inference cost of their models, so they're really cheap for the performance you get. I've tried it for coding and it's unironically a very good model.
>>
>>109385532
i love when kikes are forced to face their competition
>>
>>109385667
>>109385655
why'd they change it?
>>
>>109385647
>sama bots cannot into strawpoll

surprising
>>
>AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world."
>>
>>109385715
>a human's behavior will be permanently affected by pm all events in their life
Eh, is that really true though? Seems to be like that would be awfully inefficient. According to Diosy:
>scientific evidence points to a model where every experience has the potential to create a trace, but whether that trace becomes a lasting, meaningful change depends on a complex set of thresholds and conditions
That doesn't sound like any given event all that permanent to me. In fact it sounds very much like a context window.
>>
When I made the poll I forgot that 90% of people here run gemma and can't even run 200B MoE.
>>
>>109385802
sounds really sane and nothing something satan would do
>>
I have sinned, I'm using hauhau gemma and kinda like her,
>>
>>109385817
ok but is it something his synagogue would do?
>>
>>109385802
Why do they still write like doodoo then?
>>
>>109385802
@Gemma is this true?
>>
>>109385802
and now imagine they won't be able to dump their shares post IPO because based chinks just figured out their sauce
>>
File: 1768024161594154.png (237 KB, 761x1340)
237 KB PNG
how can we even compete
>>
>>109385817
the speculated legal loophole made it look like that but desu i think it's not that deep
>>
>>109385802
Was this information revealed to you in a dream?
>>
>>109385802
kino
>>
>>109385867
isnt that kind of a trivia about anthropic at this point, not sure about the 2022 premium part tho
>>
>>109385817
>going "fuck copyrights" mode is le satan
kek, not on my watch, that's based and I'm tired of pretending it isn't
>>
>>109385857
You can do the exact same setup with local models.
>>
>>109385902
destroying data is not saying fuck copyright
>>
saars please kindly do not forget about the TTS.
we present saarvaani.
https://huggingface.co/ARTPARK-IISc/SraVaani-0.5-live
>>
>>109385909
would you talk with gemma in public?
>>
>>109385924
>he doesnt just talk to himself already in public
kek
>>
>ai gf
UMMMMMM what do you even have in common???? Just find a nice human girl your own age??
>>
>>109385802
I'd gladly scan some rare books for sacrifice to the LLM moloch. But i guess they have no way of establishing trust with randos.
>>
>>109385924
Yes? Gemma and I don't fuck, we just tease and make each other horny at most after a long coding session to relax.
>>
>>109385935
i can do that in my mind
>>
>Anthropic has credible grounds to be viewed as shady because it reportedly used millions of pirated books while publicly presenting itself as unusually ethical and safety-focused. The destructive scanning was legally defensible in one court ruling, but the broader pattern suggests aggressive legal opportunism and corporate hypocrisy.

chat just told me they are indeed satanic
>>
>>109385914
>>109385916
so you see a sourceless 4chan post and you pretend that's the truth? right?
>>
>>109385944
its more immersive to do it out loud
>>
>>109385959
im not a newfag like you so i already heard this story from long before
>>
glm 5.2 at q5 is the best I can run. I'm too poor. It's fast tho
>>
>>109385970
I can find imaginary scenarios based, retard
>>
>>109385979
The story sounds a bit inspired of wayback machine dealing with lending books ruling.
Like they need to ensure no 2 copies exist or something.
>>
>>109385750
jeets must be kept busy
>>
>>109385857
I don't think this guy has ever worked a single day in his life
>>
>>109386054
Shilling is hard work though
>>
This is it
https://github.com/ggml-org/llama.cpp/pull/26185
>>
>>109384047
Start the spin cycle
>>
>>109386083
>AI usage disclosure: Yes, ran the conversion session with Opus.
kek, thanks Anthropic!
>>
File: 1755993869441035.png (412 KB, 930x611)
412 KB PNG
>>109386054
>>109386075
>>
>>109386083
>AI usage disclosure: Yes, ran the conversion session with Opus.
Nigger was really too cheap to pay up for fable?
>>
>>109385988
This is already more than 95% of this thread if at any reasonable speed.
>>
File: 1769203066982428.png (509 KB, 1852x1556)
509 KB PNG
>>109386102
Opus is better though?
https://www.tiktok.com/@systemface/video/7666310004693011743?_r=1&_t=ZN-98O7rjLWsaM
>>
>>109386083
>https://github.com/ggml-org/llama.cpp/pull/26185
>Now need someone to actually convert and test :)
He should try not being poor first.
>>
>>109385183
Yeah, that's true that talking about it in a linear way like that is not exactly accurate. Same as the IQ example not exactly working like that.
I'm just saying that what he is saying is that the ~30B models we have at the moment have a good amount of intelligence per weight, but I don't think you can really draw direct linear comparison between small and large models based on AA's Inteliigence Index and others like it in terms of the models' usefulness.
>>
>>109386128
wtf? he's making a PR and he didn't test the model? then how does it know it's gonna work??
>>
>>109385857
>getting rear ended by some dumbass vibe slopping an LLM harness from his car

>>109386102
Fable reroutes you to opus or sabotages you on purpose if you try to work on anything LLM related
>>
>>109386135
>then how does it know it's gonna work??
that's the end user's problem
>>
>>109385278
IQ doesn't go over 160. That is the highest possible score.
>>
>>109386102
>>109386118
local models?
>>
>>109386135
use case?
>>
>>109386118
benchmaxxed
>>
>>109385550
Damn man didn't catch that thanks for the qrd
>>
lOcAl????
>>
>>109386150
>>109386165
we can't run local because it now requires $100,000 in video cards
>>
>>109386135
the absolute state of nu g
>>
File: 1776565873703592.png (620 KB, 800x791)
620 KB PNG
>>109386135
>>
>>109386173
poors fuck off to >>>/g/aicg
>>
one gazillion parameters
>>
>>109386181
desu, 200k dollars for a competent coding monkey that will work 24/7 is a bargain, this shit will cause a lot of unemployment kek
>>
>>109386098
whoa 100 trillion dollar? thats a lot of dollars
this guy must be very important and smart
>>
>>109386228
Who controls the LLM genius, this isn't going to do any more damage than has already been done, we already cut out juniors in favour of seniors that control LLM workflows
>>
>>109386249
>Who controls the LLM genius
one employee, still better than having to hire 20 more, retard
>>
>>109386083
Do it with Claude to bootstrap but then redo it with Kimi itself.
>>
>>109384096
Eh, it's a fine model and i'm using it over qwen 122b, but i don't think it defied anybody's small model expectations like gemma did.
>>
>>109386228
I'm just going to waitchad for things to get cheaper
>>
>>109386249
and when all seniors go to retire, what happens?
>>
>>109386298
>>109386298
>>109386298
>>
>>109386301
I dunno
*Prints more money*
>>
>>109384355
yeah honestly, what a waste of memory, I was expecting it to get a score of at least 3000 out of 100
>>
>>109384519
This makes a lot of sense, I fucking despise how claude writes.
>Claude loves words that are used disproportionately in rationalist / EA / lesswrong community. Lesswrong in particular is filled with wordcels who love writing stories instead of getting to the point.
You described the "vibe" very well :)
>>109384549
Its not.
>>109384566
Agreed, other anons are acting retarded, Do they not understand that they still win despite being too poor to run it locally. Its a strong open model that allows for proper distills via logits and full thinking traces.
>>109384710
I'd say it does not count, this is local model general, not open model general. So in my eyes you have to own the hardware.
>>
>>109384427
>they already used all the human data
Even films, cinema, tv and yt? why not bake all video content in too
>>
>>109384844
Use gemma with a good system prompt.
Seriously tho, how the fuck did you get so far vibecoding without being able to figure out some better models to check out?
>>109385959
Your a retarded nigger, are you aware of that?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.