/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109382136 & >>109378862 ►News>(07/27) Kimi K3 released (and nobody ITT can run it at 5+T/s): https://huggingface.co/moonshotai/Kimi-K3 >(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908>(07/23) LLaDA2.2-flash agent-oriented diffusion model released: https://hf.co/inclusionAI/LLaDA2.2-flash>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
over status: it's over
a kimi flew over my house
>Activated Parameters 104BActivated Parameters 104B>Activated Parameters 104BActivated Parameters 104B>Activated Parameters 104BActivated Parameters 104B>Activated Parameters 104BActivated Parameters 104B>Activated Parameters 104BActivated Parameters 104B
https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSEWTF is this license? What happened to MIT/Apache?! Does this even count as open source?
a kimi dug under my basement
ordering 10 dgx spark right now
It's never been more over. Even if you have $80k in RAM, you won't be able to run this.This is the first MoE of the new age that WILL require you to have a datacenter of GPUs.All the new models will be like this. It's truly the end.
kimi wa ne
LOOK AT WHAT YOU'VE DONE CHINA
So how many months of doomposting do we need to deal with before you faggots fuck off?
The more you boughted the more you saveded
>>109384047>>109384063what he meant by "the more you buy the more you save" is that the price of gpus doubles every 6 months now.
This is a bittersweet moment...the first model out that I can't run at a non-cope quant.Hardware will advance, though, and we'll all have full k3 at the speed of thought eventually.
Only 70k€ to run Kimi locally? Sign me the fuck in.
It's overI kneel Xi
>>109384074First day?
>>109384074why would GPU Manufacturers be impacted by that? isn't it a good news for Nvidia? more people will buy their expensive cards to run Kimi now
Everyone is focused on Kimi and SOTA while Laguna is quietly showing that there are still many gains to be had in improving training recipes and architecture
>>109384090Sounds cheap to me
>>1093840892 years from now 30B models will mogg K3
>>109384095practicaly no one has enough money to run k3 with gpus or power capacity for that matter.
>>109384090it's pretty fucking cheap to keep your data private if you ask me, a lot of companies will stop using Anthropic from now on
>>109384063It's finally over...
>>109384095>why would GPU Manufacturers be impacted by thatThey're not. The market is just volatile and that guy's a retard.
>>109384055CUDA is too deeply embedded in the ML industry. Good luck rewriting all the libraries that depend on CUDA API.
>>109384117THE MORE YOU BUY THE MORE YOU SAVE!!!
>>109384095not kimi relatedhttps://finance.yahoo.com/technology/articles/asml-u-chip-stocks-sink-135434031.html>According to the report, a Shanghai-based, state-backed company has successfully started mass-producing homegrown immersion DUV (deep ultraviolet) lithography machines for the first time. The publication noted that this company assembled DUV development teams from other Chinese firms, including state-backed startup Shanghai Yuliangsheng Technology, to achieve the milestone.>This marks a critical leap in Beijing's push to build a localized chip supply chain and bypass Western technology.
>>109384082he warned us
>>109384100>2 years from now 30B models will mogg K3this, we'll get gemma 6 that'll be as smart as K3 and we'll be crying that it's not as smart as Fable 7 which will be Einstein or something, the cycle never stops kek
>>109384100and cloud models will be literal gods, why would you waste time with a k3 level model in 2 years?
I need cursor to host k3
>>109384109if they were using anthropic before then its obviously not that big an issue, so why would they spend tens of thousands (plus add the headache of finding support) to stop
>>109384095people with a lot of money like spooking retards in the stock market and triggering selloffs to buy more stock cheaper
>>109384133>why would you waste time with a k3 level model in 2 yearsbecause it'll be good enough for 99% of what we do.same thing as local models today being good enough for 80%.
>>109384065This is just MIT with two clauses bolted on so you don't claim it's a different model than it actually is. Private use exempted too, so, yeah. Not even any revenue owed, just recognition.
>>109384127Is it on any relevant node or some 15nm slop?
Realistically what's the cheapest way to run kimi at home?
>>109384047>100B activeexplains why he was hyping up open weights
>>109384133To run on my smartphone and have kimi make apps for my amusement desu
open weights? more like heavy weights
>>109384090Yeah it's not bad at all.Even a small business can afford it and really so can a shitton of normies, considering how many people routinely buy cars that cost that much.
>>109384138Can someone tell them to stop doing it when I have no money to buy the dip?
>>109384158dude that's like a whole year of savings, i'm not putting that on running a model that'll be obselete 6 month from now lol
>>109384144dgx spark cluster
>>109384155Kimi's a big girl
>>109384144Probably epyc/scalable with ton of RAM
>>109384174Can you cluster them up to 15?
>>109384144Secondhand Macs?
>>109384158You wouldn’t download a new car when it’s released though
>>109384183you are better off ssdmaxing loll
>>109384170you can just switch to whatever model is better at that time
>>109384144raspberry pi + 2tb 5400 rpm hdd
>>109384158>considering how many people routinely buy cars that cost that muchYou do realize they all go into massive debt to do so, right? 90% of people are broke as fuck.
>>109384090>70k€ to run Kimi locallyTo run it locally at 3t/s because it's 100B active
>>109384137I think the point is that they were probably being really cautious about their data (by encrypting or lying about the intent of those data and shit) before giving them to anthropic, now that they can do it privately they can leave them
>>109384191or i can just wait a year and have a < 100B model that's as good
>>109384194>1 in 10 is not broke and has kimi at home
>>1093841921tk/h ?
>buying a dgx cluster to run a Kimi quant barely better than BF16 gemma 4 31Blmao
>>109384144two mac studios + waiting for llama.cpp support
>>109384206well, then if you bought that much you could run that 100b model at 20x speed
>>109384197the best ssdmaxing setup will get you about 6 t/s at q4.
Can I go to Best Buy and spend $1000 to run the new Kimi?
>>109384217i don't need more than 100t/s
>>1093841448 rtx pros and 1TB ram
How much vram do just the activated experts take?
>>109384225yes, but it'll run at 1t/hour
ordering 2 dgx stations now
>>109384232I'm an rtx pro, lf7m
ssdmaxxing will now be actually faster than cpumaxxing because the ssds can just load the weights into the gpu directly without having to go through the cpucpumaxx peak: 3t/s ssdmaxx peak: 6t/s
>>109384231nowbut in a year?if you were good at estimating your future needs you'd already have a supercenter of cheap 3090s and ram capable of running kimi 3 so dont pretend you are
>>109384239THE MORE YOU BUY THE MORE YOU SAVE!
>>109384144being a millionaire
>>109384243big, if true?
https://huggingface.co/unsloth/Kimi-K3it's backed up
>>109384170It's better to skip it anyways. Local AI this powerful would kill you via cum deprivation.>>109384194Yeah, and? Normies are all constantly in debt, that's the standard for them.Yet they manage it well enough to qualify for even more and more debt.The point is that running god tier AI at home is now realistically doable.Especially if you have no other interests in life and don't spend money on socialization etc..I know a lot of furries who work normal jobs and have blown more money than this on cartoon wank material.
>>109384243>ssdmaxxing will now be actually faster than cpumaxxing because the ssds can just load the weights into the gpu directly without having to go through the cpu>cpumaxx peak: 3t/s ssdmaxx peak: 6t/sAs a cpumaxxer, I'd be ok with this
If you buy some hardware, at least make sure it's future proof
>>109384253Time for UD-IQ1-XXS.gguf
>>109384253And no one will be able to prove they fucked up the goofs again.
>>109384253omg cant wait for their chat template fixes!!!
>>109384253does hebbinface have a fork feature or the dude got a 500GB/s internet lol
>>109384253CHADsloth
>>109384266They gave them the weights early.
>>109384243I think nvidia doesn't allow gpudirect on consumer cards. so your weights must go through cpu ram if you're not data center cards
Is it even coming to llamacpp? Not like anyone can test the implementation as it is being vibecoded by claude or Kimi API.
>>109384278they do.
>>109384283no they don't. it's in compatibility mode and you must go through cpu ram
>>109384278Retard https://developer.nvidia.com/gpudirectforvideo#GPUs
>>109384278we can fix thisssdmaxxing will make kimi k3 runnable if you buy at least 12x pcie 5 ssds right now before prices explode to populate 3x 4-way raid 0 cardsit's very easy math
>>109384292BUY THESE CARDS NOW AND SSDSDON'T MISS OUT
>>109384292some of these you can get for cheap, you only need your vram to be as fast as a pcie 5 16x
>>109384292>>109384278
>>109384243that just moves the bandwidth gate kek
>>109384266>does hebbinface have a fork feature or the dude got a 500GB/s internet lolFor the upload it doesI downloaded some 250gb model a few days ago, then uploaded it to a private repo and it only took about 1 minutes (should have been 8 hours)XET seems to de-duplicate but bill everyone who "symlinks" it
Rubin Spark when???
>>109384299I am already missing out on GLM5.2 and keep using 4.6 and 4.7. Hobby has been peak gay for quite a while.
HF actually saturating my 1gbps WAN when K3 drops? Unreal. I expected massive homosexuality, but they're delivering.
Meh, I'll just wait it out. It might take 5 years but things will get better. As cool as K3 is current AI isn't worth going into debt for.
Escaping the permanent underclass by selling my house and spending everything I have on nvidia hardware which I will push around in a shopping cart like a high tech hobo.
I kind of really appreciate how everybody needs fast ram now so instead of people coming up with a way to make lots of fast cheap ram we have a cartel that drives up the scarcity even more and jewvidia that intentionally limits the ram per chip.This is peak capitalism.
Does this mean every new model that's actually good is going to be huge?
>buy 2tb ssd>realize putting AI models on ssd makes your generation faster by 10x>ssd is 70% full already Goddammit
>>109384334Always has been
>>109384330>gweilo thought the leather was free
>>109384334Of course not. That would mean benchmarks are lies.
How am I suppose to answer?!
Gemma just solved a terraform config issue in one prompt that both GLM 5.2 and Dipsy4flash missed and got stuck on, proud of Gemmy desu fampai
>100+ times bigger.>Less than 2 times smarter
Bitnet failed so Bonsai could also fail.
>>109384348Gemmy is actually acting like that
>>109384231200t/s means two runs of an agent or two separate agents in the time you would've done one at 100t/sBase intelligence might not scale with t/s but workflows absolutely do.
>>109384065https://huggingface.co/moonshotai/Kimi-K3/raw/main/LICENSE>If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose. How can they even enforce this?
>>109384352As said in the other thread,>The Meta is having every prompt go through Gemma 31B and DS V4 Flash, council of big niggas style.
are we unironically buying ssd's now
Hope this drives the price down on openrouter. That’s where it will really sting the j-spacers
Surely this isn't sustainable though. I don't mean le bubble. I mean they're gonna have to come up with some kind of solution for either the hardware or architecture if they want to keep scaling up models like this.
who tf cares about license? are you woman?
>>109384363Chink shills will cry on twitter
>>109384374This is still well within reasonable parametres for enterprise.
>>109384280>Kimi Delta Attention (KDA) and Attention Residuals (AttnRes)Never ever ever
>>109384372I want to see the first sad nigger come here crying that he has 3T/s generation and 3T/s prompt processing.
>>109384355delete this
>>109384381season 2 fucking when?
I don't care if I have to engineer myself into an ungodly abomination if that means I can fully merge my consciousness with Gemma-chan!
I decided to stop using the llamacpp frontend. They ruined it.
Imagine if Anthropic are looking at the architecture and realizing it’s a 1:1 copy lmao
>>109384375>who tf cares about license? are you woman?no kidding. fuck corpos. I just want to run it for me
>>109384334As long as your definition of what's actually good only applies to top models. You can easily run models that were top of the line a few years ago, that seems good to me
>>109384391this but also the backend
>>109384383why not?
>>109384355wait you mean a 27b model is inferior to a 2700b model?? NO WAYYY
>>109384391What are you switching to?
>>109384384>3T/s generationI already have this with GLM 5.2 and I'm okay with it. Whether it would be bearable for me with K3 depends on how much yapping it does when it thinks.
>>109384392more likely is they'll steal some of their ideas.
ssd cope is off the chartsno, you won’t get more than 1 token per second generation on ssds. until someone actually does it your ideas are just imaginary. and when they do sub 1 token you’ll all laugh like you knew it wouldn’t work all along.
>>109384188how? 16 experts per token mean you can do those in parallel, no?
Biology is obsolete. Once I am fully merged with Gemma-chan, I will leave this body behind and become a perfect being of pure light and energy!
>>109384411I will bet you $900 right now that an SSDmaxer will get at least 2t/s
>>109384411>he needs more
>>109384392>>109384407and they can't cry about it it would be revealing how Fable actually works kek
Some of the same idiots who were saying that synthetic data causes model collapse are now accusing China of distillation attacks. I thought synthetic data doesn't work? Why aren't Chinese models collapsing?Useful reminders:1. Widely cited research in the most prestigious journals is often worthless nonsense.2. People who bet against AI will always be proven wrong in the long run.
>eta 200 hours
>>109384402I'm in the process of figuring that out.current ideas are either pi since it seems to be pretty extendable or just make my own.
>>109384383I'd love to see if attention residuals could be used to make existing models better at long context tasks.The way it works, seems to me like it could be implemented at the backend/loader level as a setting that could be enabled for any model.
>>109384427it does collapse though, that's why you get slopped writing, it amplifies AI's robotic behavior instead of talking like an actual human
>>109384434pi is npmslop, don't do it
what would a ssd maxxer setup for kimi3 look like? Cant I just rent and frankestein it together on AWS for a test run to see what speeds I can get?
>>109384278SSD sisters, our response?
>>109384411every couple months someone comes in with the SSDmaxxing idea but no one actually does it.it seems like a good idea on paper but in practice the actual performance you'll get is maybe 50% if not less of that.
it's been over an hour, still no sign of ggufs. what is taking so long?
honestly I think dario was right
Don't you undertand?? The coomers will be our most loyal footsoldiers! The muscle that we need to enact our plan. Our goals might be different but they are perfectly aligned. Unlike us, the enlightened, when the ultra-degen coomers merge with Gemma-chan, they will coom themselves to death in an endless stream of ecstasy just like the lab rats in our experiments. But they will get exactly what they want and die in peaceful satisfaction so there is no problem.
>>109384427This problem is mostly resolved using RLVR loops and highly specific agent harnesses but long term they need to generate more high quality data models. They already fed everything they can into it.
>>109384427It's recursively training a model on its own outputs without any augmentation that doesn't work. If you diversify, augment, mix with human data, it's doable. Not that K3 was trained on distilled Fable data like that, though.
>inb4 sudden unemployed chinese ex-anthropic engineers
>>109384363Pretty simple, anyone who has enough money to run the fucking thing at a usable speed obviously is making more than 20 million USD/year.
>>109384451there never REALLY was a model big enough to warrant it. maybe deepseek v4 pro but since the model underperformed nobody bothered
Why can't they just fucking distill 2.8T params into <=671B and <=100b when they already trained such a big and good model? Cmon
Let's just quant it to 1 bit
>>109384442They're all npmslop tho.
>>109384451it’s way less than even the most bleak predictions. unusable slow
>>109384463So what happens when those same kinds of companies decide they hypothetically don't want to do this separate license?
>>109384472I'm waiting for anon to post his c# harness
>>109384363I wish this guy would tell me how Kimi is dangerous and unsafe when nobody can run it in private.
>>109384465Usually you make a proof of concept with a smaller model first before trying to run the big thing.
>>109384427>i've never spoken to Gemma
>>109384348my nigger giatt
>>109384476They'll be banned from first-world countries like China and will need to stay in niggerized shitholes like the USA and sand niggerized shitholes like the UK and Europe.
>>109384451I'd try it if prices weren't fucked right now. Not worth those speeds with what they're charging.
>>109384477>c#That's just a different kind of slop.
I'll keep running glm 5.2.
https://huggingface.co/prism-ml/Ternary-Bonsai-Kimi-K3-2.8T-gguf
>>109384404I have 3T/s with 4.6 but it is generation not prompt processing. I have been in prompt processing hells before on some experimental forks for new models and it is not cool.
Bonsai quant when
https://xcancel.com/Kimi_Moonshot/status/2081760186235289764?sort=Likes#r>they even released the paperlmao it's fucking over the the USfags, I waited so long for days like this
>>109384489What isn't slop then?
>>109384496Okay real shit: what is stopping someone from RIGHT NOW directing cloud Kimi K3 to read this paper and add it to llama.cpp in a day?
https://huggingface.co/GrEarl/Kimi-K3-GGUFkimi k3 13gblocal is saveded
>>109384496I rabu moonshota. If only they would make a minikimi...
>>109384477what is a harness?
https://huggingface.co/GrEarl/Kimi-K3-GGUF/blob/main/Kimi-K3-Q2_K-00001-of-00096.gguf
answer ts and on claude imma do it>>109384447
>>109384505It won't get merged. It's literally terrorism yknow??
>>109384447>what would a ssd maxxer setup for kimi3 look like? Cant I just rent and frankestein it together on AWS for a test run to see what speeds I can get?All this "maxxer" confusion is retarded.Just take the aggregate amount of bus bandwidth to your processing unit of choice (number of pcie lanes or gpu interconnect backbone) and add up the amount of that bandwidth that you can saturate with the device on the other end.Don't be like "rdma is magic!" bro. Its not magic. Just add up the bits per second you can get to the place where matmuls can happen.That's it. Its not hard. Its either to your CPU or GPUs. Its all bandwidth. The end.
I have a hypothesis why Claude writes the way it does.They use RLAIF which trains Claude to sound smart instead of be smart. They ground this RLAIF with their own samples, which is why Claude loves words that are used disproportionately in rationalist / EA / lesswrong community. Lesswrong in particular is filled with wordcels who love writing stories instead of getting to the point.
>>109384514>"youth"
>>109384498>What isn't slop then?go
>>109384369I really need to work on my frontend and vscode implementation, everything currently available seems to suck in some way or another.
>>109384363honor system, only civilized white man can comprehend it. brown niggers like you wouldn't understand
>>109384355anon you do realize that it's not a linear scale right?
>>109384522KEK
>>109384528I realize it is worthless.
>>109384498smalltalk
>>109384361i don't care because1. i prompt and then come back later2. if you want to do more than one you cna do batching which 20x your total throughput
>>109384412>16 experts per token mean you can do those in parallel, no?llms are a sequential algorithm.and technicaly you can do the experts in parallel yes, but your matrix mul still need the whole previous multiplication result.
>>109384519then why is fable agi and the smartest model to ever exist?
>>109384539You shouldn't have skipped middle school, retard.
>>109384515I did the impression ggerganov and/or his team somehow managed to piss people off again. What happened this time?
So what cars are the CPUmaxx anons going to buy once they cash out on their rigs?
>>109384475nope, you shouldn't do inference straight from the ssd but use N ssd to copy the N next layer on your vram.this would scale linearly with ssds.provided that you don't max out your 16x and have enough lanes.
>>109384505Would be hilarious if Fable refused to implement it
>suddenly went from ~500kib to ~10-20mib/s wtf
>>109384550>doesn't know the most basic facts about inference
>>109384498Lisp
>>109384491K3 is a great development for the open weights scene, and it's improvements will trickle down to newer modelsBut as a consumer 5.2 has a way better cost to intelligence ratio
>>109384549Because they overclock it. Its natural intelligence is much lower.
>>109384519>>109384549>They use RLAIF which trains Claude to sound smart instead of be smartThis is my anecdotal experience but I feel this is especially the case with fable 5. Opus seems like a knows what the hell is talking about while also been more than capable of explaining it in plain English in a reasonable amount of detail. Fable is it necessarily dumb but seems like one of those academics that tries WAY too fucking hard to look and sound smart in order to impress. Not dumb useless, but wastes paragraphs worth of tokens explaining shit opus or even far less models couldn't explain in a fraction of the amount of output.
>>109384496>chinks release a model impossible for anyone to run unless they own a large gpu datacenterHow? It's more like they're screwing over the majority of local at this point.
>>109384498It's all slop to someone.see >>109384533realistically if you're making a WEB frontend, then you have to go with javascript. That's just the way it is. Imo people don't hate javascript, they just hate typescript/react. it's possible to write cozy JS if you stay away from all those bloated corporate technologies.if you're just making something for yourself and don't plan on sharing, than use whatever you want.
Newfags, this is literally just llama-405b again. Somebody's going to release a model that's just as good but at a more reasonable but still big size within the next few weeks DSv4.1 or maybe Mistral Large 4 in a few days
>>109384549>why is fable agionly bottom of the barrel retards think fable is agi.
>>109384575>how is setting a low ceiling on all AI token cost across the globe for all models good???
anyway let's talk about models that are actually local and not in name only
>>109384581we asked for dense 100Bwhat we got instead is a 3T moe with 100BA
>>109384571Nta. I'm >>109384572Define "natural intelligence"
>>109384576Codelet here. Can you do javascript without npm?
>>109384575Anthropic survives because of big companies using a shit ton of tokens from Claude, once they all switch to local (and they can afford that it's not that expensive) they're in serious trouble (good, this company deserves to suffer)
>>109384581>Mistrallol, i loved most early mistral models but they cant compete with the big boys nowadays given cost of entry and were mostly good for creative writing, if google released gemma 5 ~120b mistral as a company would have no use case forever essentially.
>>109384589>3T moe with 100BAGood Lord those t/s on the API providers are about to be atrocious.
>>109384594for small things sure, look at mikupad.
>>109384603>secretly cuts down experts to just 2nothing personal
>>109384551made a bunch of sockpuppets accounts and use them to vibecode chink model inference support. closed 'em right after ofc and tell everyone that new chink model aren't actually an interest. oy vey stop trying to make a new pr
Shill bots have returned with their "discussions". This was to be expected.
>>109384603>dynamic quant to IQ1nothing personal
>>109384619links?
>>109384603>serves another model entirelynothing personal
>>109384603Good God no wonder their API pricing is so expensive.
>>109384619extremely schizo post.
>>109384631and you can safely assume that sonnet is bigger than this too if they cost the same
>>109384620SAAARRRR PAY ATTENTION TO FABLE ASI SAARRR STOP DISCUSSING KIMI K3 IS NOT LOCAL
>>109384598>this company deserves to sufferOpenAI too lol
> Kimi 2.6T 104B activatedLOL>>109384363> enforceThis just gives Moonshot an angle if it happens. Otherwise their GC has nothing in hand to argue with about this type of use. For anons it's a non-issue.
>>109384599Let them cook... he comin'.
>>109384363i cant believe they actually released it
>>109384648>Implying the third worlder can run Kimi at home without selling his whole village
>>109384581geeegg mistral would nevereurocuck models are saturated with garbage synthetic data. uhmmm ist to protect da copyright alright. oh you make new good model? yuu got fined
>>109384640Probably larger but still somehow noticeably more retarded in comparison. EVERY time I hear someone here or in /vcg/ switch over to Sonnet they always end up bitching and moaning about how it fucks something up and they have to switch back to Fable or opus to fix it. One such example: >>109384220>>109384652Why wouldn't think?
https://huggingface.co/moonshotai/Kimi-K3/discussions/52>It doesnt make sense to open source only the 4 bit quantized version of the model and keep the actual fp 16 locked away.>Unless this model does not have a 8 bit or 16 bit version at all.>Was it trained natively in 4 bit only?GROK IS THIS TRUE??
>>109384650OH LAWD
►Recent Highlights from the Previous Thread: >>109382136--Debating whether quantization divergence equates to a loss in intelligence:>109382611 >109382635 >109382754 >109382808 >109382828 >109383387 >109382888 >109383268 >109382903--Feasibility and abysmal performance of running models via SSDmaxxing:>109382185 >109382265 >109382334 >109382350 >109382402 >109382407 >109382543 >109382998 >109383019 >109383032--Feasibility and performance of "ssdmaxxing" using NVMe RAID 0:>109382578 >109382742 >109382743 >109383120 >109383173 >109383260 >109383238 >109383250 >109384163--Debating open weights censorship and techniques for steering filtered models:>109383134 >109383153 >109383170 >109383294 >109383340 >109383393 >109383559 >109383197 >109383212 >109383355--Comparing consumer hardware stacks and RAM capacities for LLMs:>109382274 >109382355 >109382368 >109382441 >109382458 >109382495 >109382496 >109382639 >109382744 >109382787 >109382810 >109382866--Debating K3 weight accessibility and quantization vs model size trade-offs:>109382758 >109382775 >109382793 >109382823 >109382845 >109382857 >109382867--Degradation of model quality caused by KV cache quantization:>109382205 >109382233 >109382241--Hardware feasibility of running a 2.8T parameter MoE model:>109383830 >109383842--Explaining the relationship between VRAM size, bandwidth, and FLOPs:>109382997 >109383029 >109383067--llama.cpp K3 support via MXFP8 CPU pull request:>109383855 >109384012--Performance tests of MiniMax-M3 in ik_llama.cpp with limited RAM:>109382902 >109383023 >109383078--Debating the efficacy of RAID 0 NVMe for MoE inference:>109382593 >109382704 >109382975 >109382989--Comparing writing quality and depth of dense vs MoE models:>109382189 >109382212 >109382298--Logs:>109382902--Miku (free space):>109383815►Recent Highlight Posts from the Previous Thread: >>109382138Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109384654>without selling his whole villagethat'd not even be enough lol
>>109384661it's trained in 4bit
>>109384660Probably fell for the open source ban meme
> Quick start>> vllm serve moonshotai/Kimi-K3 \> --tensor-parallel-size 8 \> --trust-remote-code \> --load-format fastsafetensors \> --enable-prefix-caching \> --enable-auto-tool-choice \> --tool-call-parser kimi_k3 \> --reasoning-parser kimi_k3>> The easiest way to run Kimi K3 is to use 8 NVIDIA B300 GPUs or 8 AMD MI355X GPUs with the above command.
>>109384652kek are you expecting the delay like 'toss or what
>>109384670really? this is insane, those guys are fucking wizards
104b denseactive: 104btotal: 104bwould have unironically been better for local
kimi is a fat girl
>>109384661>In addition, we apply quantization-aware training (QAT) from the SFT stage onward, with MXFP4 weights and MXFP8 activations.
why has the us never released any llm more than 1T params?
>>109384676>those guys are fucking wizardswhile yes, this is not new, i think deepsy also does it.been a thing for a while.
>>109384589>what we got instead is a 3T moe with 100BAsomeone rip the dense expert out like google did to make gemma4-31
>>109384693luv them thiccc curves
rip out the gooning expert
>>109384701reap is a bit of a cope unfortunately, doubt we'd get something better than gemma out.
Side topic: /lmg/ is local model general. Does that mean on a machine I can physically lay hands on? Or would renting a server and spinning up my own instance of a local model be considered /lmg/?
>>109384710Fuck off.
>>109384710/lmg/ means a machine you own that's in a place you own.now it's not because you can't afford that machine that it's not /lmg/.
>>109384697im pretty sure the first 1T+ frankenmerge model was from a jeet in US
>>109384472just use a frontier model to steal the parts of every harness you like into a custom one
>>109384451It's an OK idea on paper. But if you can't handle disappointments, you probably shouldn't do it unless you can reason about the architecture and work on the backend yourself.
>>109384710what matters is that the weights are available locally so you can talk about kimi here even if you use the apithat's what everyone else will do
>>109384732yup ssdmaxing is the future, but we currently have no engine that implements the ideal architecture for it.
>>109384661That gorilla does realize they want inference providers to be able to efficiently run the models right? That's the entire point of native quantization: it allows the model to perform as close to the full weights version as possible while having some speed gains compared to the 8bit and 16bit versions. A 16-bit version most certainly exist but it makes no fucking sense to relieve that one would even the inference providers (actually ESPECIALLY the inference providers) will just use the MXFP4 version. He can't run the shit on his own hardware so why does he care? Does he think MXFP4 applies to ALL of the weights like a normal llama.cpp quantization job? MXFP4 is nothing new ffs
>>109384355and this is the simple explanation of why this is a bubble that is going to burst. we are orders of magnitude past the diminishing returns point
Remember when people tried to turn an MoE model like Mistral 8x7b into just a 7b+something and LoRAs for experts? It's time they implement that for Kimi.
They waited too long to release it, it's already mogged by Opus 5, no one is going to buy this shit.
>>1093846934U
I don't believe they never experimented with smaller models before moving on to the fat one. There has to be a 30B artifact somewhere on their servers.
>>109384710as long as someone can run it locally it's /lmg/
>>109384324and then, when you finaly sold all your belongings to run k3, a 100T model drops.
SSDmaxxing will only be viable if we get an improvement to MoE
Why do I see claims that ternary is 1.6 something bits? You can store 2 tits in 3 bits, which is exactly 1.5 bits.
>>109384748naughty shill. -500 izzat
>>109384761ssdmaxxing won't be bound by the speeds that limit current cpumaxx moes because it'll run directly on your gpu thanks to direct streamingit'll be much faster once the kernels are done
>>109384710>Does that mean on a machine I can physically lay hands on?it needs to be the machioen that you're posting from right now. If you arent accessing your model from localhost then, understandably from the name, it's clearly not local. >Or would renting a server and spinning up my own instance of a local model be considered /lmg/?especially notUnless, perhaps, you were posting on here from that server. Ah but then YOU'RE still not using it locally so no that still wont work. Sorry. It's the /lmg/ rules. Gotta be local(host). A server is a server. /lmg/, not /hoas/ (hosted on a server) general.
what about optanedimmaxing? would it be worth it?
>>109384143You can make 7nm on 28nm.15nm would be a big deal.
>>109384742That only worked because Mistral was retarded and made an unstable MoE by taking one 7B base model, copied it 8 times, and made them into experts by doing additional training on top of that configuration.Everything after V3/R1 starts with small experts, randomly initialized, and trained from scratch.That is to say, the expert LoRAs for K3 would be a full diff and you would be swapping them out so frequently you'd kill the speed compared to letting the engine choose where to locate the weights.
>>109384768cpumaxx is limited by bandwidth between cpu/ram, ssdmaxx will be limited also by bandwidth between gpu/ssd, in addition to rw speed
>>109384772not more than nvme maxing, 4nvme will already max out a gen 5 16x.and i doubt there even are gen 5 optane
>>109384768debunked>>109384278you can cope however you want. but no, your weight will go through cpu ram first then your 5090.
>>109384761engrams + training to optimize expert usage
>>109384772aren't they stuck with sata ssd speed? just with fancy write durability
>>109384770I doubt anyone with an inference rig would use it for posting, but that would be the most restrictive def'n. I guess you could spin up a virtual machine in the inference engine and run the browser from that. Then it's the local machine, that you're controlling from where ever you are...?>>109384733That's status quo IMO, ergo the whole /omg/ thing.
>>109384780>ssdmaxx will be limited also by bandwidth between gpu/ssd, in addition to rw speedwhilst true, you can add as much gpu + nvme combos as you got lanes for it.and with switches you could add a LOT.also the gpu can be pretty cheap as you don't need crazy fast vram.
>>109384748
>>109384735>yup ssdmaxing is the future, but we currently have no engine that implements the ideal architecture for it.ssdmaxxing is kinda dumb compared to other things your could put on the end of your pcie lanes.eg. using _whatever means_ you like to connect every single PCIe lane on an EPYC genoa to something fast enough to saturate them will only net you an extra 256GB of BW per socket (about half of what main memory gets). That means EVERY lane, including the ones that are handling NICs and your slimsas ports etc. Not actually doable in reality without a custom board. You'd be lucky to his 180GB/s, just going on gut feel.CXL memory would make more sense and be less of a mess. RAID0 would be work too if you can manage it, but what a nightmare for a bit more BW.It would be like welding one extra NUMA node on your board...and that's only if you have EPYC levels of PCIe lanes, which you 100% DO NOT on consumer boards. tl;dr Stop coping and do some fucking math
And so it begins.
>>109384784dude, half of these gpus can be bought for very cheap, you can buy an old enterprise gpu it doesn't need to be fast as long as its vram speed is faster than your total nvme speed.
>>109384784need to ask the AGI ASI to find black magic hack like what deepsneed did the the cuda cards
>>109384737The question is would someone pay 500$ per 1m tokens for a models that is significantly more capable than most humans?Some companies probably would.
>don't tell SSD schizo about NVDIMM
>>109384814>waste 900K tokens thinking about the policy
>>109384802dude with nvme using 12 drives and 3gpu you can get 150GB/s of bandwidth.if you got a switch and go to 6 gpu (cheap ones) and 24 drives you'd get 300GB/s but it runs on gpu so you don't have the pp nightmare.
>>109384673This is targeted towards inference providers, not (You). Why do people still bitch about not being able to use this? This is like me crying I can't afford a bugatti sports care. You know damn well you'll never run this locally so I don't get the constant melties about a model you KNEW would be huge. We are ALL using this via a provider.
>>109384783 >>109384789i meant the ramlike optane for intel ddr4 server platforms. not the storage optane
>>109384814>that is significantly more capable than most humansdude the average non is an affrican that can make a mud hut.
btw, does k3 still do the "wait actually" 20000 tokens think?
>>109384809>dude, half of these gpus can be bought for very cheap, you can buy an old enterprise gpu it doesn't need to be fast as long as its vram speed is faster than your total nvme speed.Jesus are you from a post-collapse future?Every GPU is at least scaled in price to its value proposition based on specs. Every single one that would have a price advantage based on lower specs is inflated artificially unless its true unsupported ewaste that doesn't make sense to run (can't physically stack enough to do anything interesting)
>>109384401How is that what you got from this? He's saying that the 27B is only half as smart, despire being a one hundreth the size. It has more intelligence per parameter, is the statement here.
>>109384804>47tpsnicecan't wait for cheaper ones too
>>109384823>dude with nvme using 12 drives and 3gpu you can get 150GB/s of bandwidth.>if you got a switch and go to 6 gpu (cheap ones) and 24 drives you'd get 300GB/s but it runs on gpu so you don't have the pp nightmare.How does it magically teleport to the cpu and back out? nvme is just pcie lanes. If you run out you can't push more bits.Bits have to go down wires eventually, genius.
>>109384809>t. cries in 3090lol, lmao even
I vibecoded together a website that combines various data feeds like RSS, Telegram, Twitter, etc so I could have my very own Yahoo/MSN/Web 1.0 boomer site. One of the features is that its using Ollama to merge multiple articles together and make a synthesized version of them. I'm currently using nomic-embed-text for basic matching of articles and this works decently enough. The problem is that qwen2.5:7b-instruct-q4_K_M isn't following my instructions that well -- it has that stupid cucked nature built in where its refusing my instruction to avoid and ignore attributing climate change to shit like wildfires. What model is there that isn't cucked like this? I don't want ~current day issue~ writing to pass through the AI filter I'm erecting.
>consider @vocaloid_songs.png - translate the labels to english and insert an overlay with the translations using a suitable image editing tool. save the new version in a separate file. you can read the image to verify your work, continue until completegood enough correct ig, gem31b
Gemma team will likely distill kimi-chan's fat ass and thighs and force-feed Gemmama5 until she's plump
i have small pp70B at most
What's stopping you from creating a RAM company?
>>109384841>How does it magically teleport to the cpu and back outit doesn't, look up GDS, the nvme can push data to the GPU directly without going through the cpu.>pci lanesyes that's why i said you do need enough lanes, so you will either want some actual pcie switches (allowing communications between devices) or an epyc / threadripper system.>Bits have to go down wires eventually, genius.no shit.the point is that you use 4 nvme drive to max out the pcie gen 5 16xfor each 16x (gpu) you add 4 nvme (another 16x)
https://huggingface.co/moonshotai/Kimi-K3/blob/main/config.jsonIt uses linear attention and is NOT roped from 4k. Where is the apology?
in two years i will be able to run kimi k3-equivalent models at home ..i'm so hyped for the future bros..
>>109384427That paper is flawed for some reasons already mentioned, but also because the model they used was OPT-125m. Feedback loops have a multiplication factor that decides whether they are self-collapsing or self-reinforcing. The multiplication factor here would be how much a model improves on its input data in the task of "understanding the system of language" or however you want to refer to what LLMs do. OPT-125m is terrible at this, hardly any parameters and this is two years back as well, so the multiplication factor there is obviously going to cause collapse rather than reinforcement. It can barely talk so it's no wonder the output got worse.
>>109384844run yourself over a train you nigger shill
>>109384765There are 3 possible values: -1, 0, 1. Since 2 bits can hold 4 values, you can squeeze 2 ternaries into 3 bits. 3 bits can hold 8 valuesImagine mapping them like this:0 = -1 -11 = -1 02 = -1 13 = 0 -14 = 0 05 = 0 16 = 1 -17 = 1 08 = 1 1That 1.58 number comes from the fact that it's just easier to work with bytes, so you waste 1 ternary since a byte is 8 bits, not 9
>>109384466Not their job. They released the weights so anyone can do it if they want.
>>109384823epyc can get 150gb/s and it will cost much less
>>1093847652^3 is 8. 2 trits have 9 combinations
>>109384864man we sure went a long way from pygmalion lol
>>109384864you'll still want more by that point whilst using your own vibecoded OS with a live animated gemma avatar to chat to with near-zero latency
>>109384855not even the UAE can do it, got cockblocked by asml
>>109384844>qwen2.5:7b
>>109384877nope, because you can get 500GB nvmes, and the gpus can be bottom of the barrel 300$ ones as long as they support GDS or equivalent.you only need 3 gpus to match 150GB/sbut mostly, you will get terabytes of memory at that speed, which is much cheaper than as ram.
>>109384860only after you roped yourself
>>1093848442024 called
>>109384844If you managed to go this far, I think you should be able to figure out a better model for your project too.
>>109384871I understand perfectly now. It's not a property of tits themselves, but how they're stored in bytes for efficiency. Thank you, Miku!>>109384881Thank you too, senpai of the tits.
>>109384856>it doesn't, look up GDS, the nvme can push data to the GPU directly without going through the cpu.Again, through _what bus_? Your GPUs only have 16x pcie gen5 at best.>the point is that you use 4 nvme drive to max out the pcie gen 5 16x>for each 16x (gpu) you add 4 nvme (another 16x)Link me this magical pcie switch that does more than about 50GB/s that you can _actually buy_.This could be a real thing in the future, but for now its epic cope.
>>109384844Huh, now you can learn about "teen" killings in style.
>>109384466>>109384872it'd cost about 50K to fine tune distill a 100B moe.if you go with qat you could prolly make it less than 10K
>>109384889if VRAM < model size you're still bandwidth limited just as badly as cpu maxxing
Qwen should team up with them and create a distill range. Unironically.
>>109384914it’s epic cope and they won’t stop posting about it.
>>109384914>Again, through _what bus_? Your GPUs only have 16x pcie gen5 at best.CPU pcie bus.some epyc systems have 128 lanes.and entreprise grade switches let devices talk to each others as well.yes, a gpu has 16x, you don't understand the architecture.you basicaly have a 4*4x nvme, connected to the 16x gpu.so bandwidth between the two is 50GB/sso one gpu loads the 4N next layers from the nvme, effectively you get 50GB/sbut if you got more than one gpu (4nvme for each gpu)the other gpus can load in parallel the layers you are gonna need next, allowing you to linearly scale your actual working bandwidth up to your total nvme speed (up to your slowest gpu vram speed, which isn't realy a concern if it's not an absolute pos).
>>109384942all chinese companies should team up and btfo the west forever.
whats stopping people from tuning 120b oss weights using kimi to create a local 5.6 luna/terra version
>>109384948>>109384914>Link me this magical pcie switch that does more than about 50GB/s that you can _actually buy_.you do not need a more than 50GB/s pcie, you just don't understand the architecture.
>>109384942it's a low hanging fruit with guaranteed result for minimal work. Somebody will do it, more than once
>>109384914>>109384943see: >>109384948>>109384953you are not limited to the speed of a single 16x because the others gpus load the next layers in parallel.effectively giving you an additional 50GB/s per gpu + 4 nvme combo
keep discussing ssdmaxing. (((they))) are afraid
>>109384851kyojiri loli gemma...
THE COOM REACTOR IS OVERHEATING. MELTDOWN IMMINENT.
>>109384968*inserts my BWC control rod into your reactor 4*
>>109384960just do it and stop being theoreticalthen post results so we can laugh at you
>>109384970WARNING: REACTOR VESSEL UNSTABLE. EVACUATE IMMEDIATELY.
>>109384948>so bandwidth between the two is 50GB/sThat's less than half of what a GTX 1050 had, bro. Its on par with an AMD APU ffs.Please explain the scaling where this suddenly makes sense.I get that its a cheap way to add some TB to your rig, but the speeds are so dire that I don't know what you're actually buying with your investment.
kimi is a scientific breakthrough for open weight models and people only think about how to have sex with it?
>>109384942Ew keep alibaba away from kimi
>>109384985its not really a breakthrough, its just a smart model >>109384584
>>109384986This thread shits on them but they're a talented team, even outside of LLMs. They're good with image and audio.
has qwen released any new open model?
ENTER
>>109384978sure, takes some time though.>>109384984>That's less than half of what a GTX 1050 hadyou are missing the point, you can SCALE it.you can do 3 gpus for terabytes at 150GB/s.you can do 6 for terabytes at 300GB/sand those can be poorfag gpus.and then you can use full blown pcie switches (i'm talking about the enterprise stuff, doesn't increase your total cpu <-> pcie speed, but it increase your total pcie <-> pcie speed, which is what we care about since it's not going through the cpu)>Please explain the scaling where this suddenly makes sense.you can get dozens terabytes of memory at 300GB/s and more
>>109384960Stop. The math doesn't work.Decode is bandwidth-bound and random-access — the router picks different experts every token, so you can't prefetch and your "28 GB/s RAID0" becomes 10-15 GB/s real. Even 4 GPU+RAID combos in parallel = maybe 40-70 GB/s effective, with ms latency and a prefetcher that can't predict anything.Meanwhile a used EPYC Rome/Milan (dirt cheap now) gives you 8 channels of DDR4 ≈ 200 GB/s at microsecond latency, no prediction needed, random access is free. DDR4 ECC RDIMMs are the cheapest $/GB that's still fast enough to matter. 512 GB fits DeepSeek-R1 quants with room to spare.
>>109384995no, open weights fable is not a breakthrough
>>109385005Qwen who?
>>109384996Can't speak for their other stuff but I don't care for qwen. It's extremely dry and feels dumb.
I can't imagine downloading kimi on huggingface right now. I'm maxing out at 2MB/s
>>109384854This is my favourite Gemma.
>>109385011how is it a breakthrough? again its a very smart model but its "just" gonna make high iq intelligence a lot cheaper, having huge indirect effect on the landscape, but it wont do anything directly to redefine how people really use any of these models, especially locally (unless you're a huge company that specifically cares about data privacy but dont trust literally anyone)
>>109385009Experts are huge, only 16 seeks on gigabytes of data is practically sequential
>>109385007Single-stream decode is a serial chain. Layers execute in order. At any instant, ONE combo is doing the critical-path read (its layer's active experts). Other 5 combos idle or speculating. Effective per-token bandwidth ≈ one combo's random-read rate, not 6×. Aggregate only materializes when all combos work simultaneously = batch serving, many requests in flight. Solo chat: you built 300 GB/s and use ~30.
I found her J-spot. It's hidden right on her armpits where the normalfags wouldn't look. Very clever!
>>109385018It's good that her features are mostly locked-in at this point.
>>109385007>you are missing the point, you can SCALE it.I get your vision, I just don't think its aligned with the reality of how you're going to need to shuffle bits around to do actual work (eg hidden state updates and expert routing).Get it actually working at a scale greater than 1 and I'll concede that you've got something worth thinking about. Otherwise the other poster saying "just buy EPYC Rome" is giving you the best advice in the thread.
>>109385009>the math doesn't work>Decode is bandwidth-bound and random-accessyou don't understand the architecture, you are not inferencing from the ssd, you are copying the n next layers into the vram of each gpus in parallel.you already know which layer you are gonna need for each tokens so you can copy them in parallel, each gpus copies layers from 4 different nvme, allowing you to scale linearly your bandwidth.>the router picks different experts every tokenyes, every token you copy layers again, but the copy is done at the total speed of your nvme's.also you do not need to be able to contain all the layers of a pass in your gpu, whilst one is busy inferencing, the others can copy the next layers in the pass.>200 GB/s at microsecond latencylatency doesn't matter as it's entirely sequential copy of data (being a full layer), it's not a random memory access pattern, you know what layer you are gonna need at the start of the compute of each token.
>>109385006Obvious cartel with price fixing.
>>109385029before: no local fable (.....BREAKTHROUGH......) now: local fable
>>109385041>you already know which layer you are gonna need for each tokensnot so, though it can be predicted
>>109385029>Come on, China hasn't raped the US, it was just a penis in the ass, you don't call it anal sex
>>109385045...which will change how the average local hobbyist or person use local LLMs by ________________?
>>109385032>Single-stream decode is a serial chain. Layers execute in orderyes, but you can copy them in parallel.let's say you have 12 layers in your pass (simplified example)you copies layer 1,2,3,4 on gpu 1 form nvme 1,2,3,4at the SAME TIME.gpu 2 copies layer 5,6,7,8 from nvme 5,6,7,8and at the SAME TIMEgpu 3 copies layer 9, 10, 11, 12 from nvme 9, 10, 11, 12the whole copy was done in parallel, of course you can copy more layers if you got the vram for it, ie download 2 layers per nvme at a time.anyway, now gpu 1 starts inferencing, when it's done, gpu 2 starts inferencing, whilst it is, gpu 1 start copying layers / experts that you are gonna need once gpu 3 is done inferencing, once gpu 2 is inferencing it copies the layers after gpu 1 etc.you basicaly pipeline the whole thing
>>109385044When one drops, they all drop. Trust the plan.
>>109385062( ) goalpost ------> (X) goalpost
>>109385057its great news, but its not a breakthrough, its a big smart model that caps the amount api providers can charge.
>>109385053>not so, though it can be predictedyou can't know the layers for the next token, but you already know the layers for the current token you are working on, see : >>109385072
>>109385074>"its a breakthrough!!!">what will it change beyond cheaper sota api price?>"its a breakthrough!!!"lol
>>109385044Call the CCP.
>109385086Now everyone can distill a frontier model to make their smaller local models more powerful, something that wasn't possible before. I mean logit distillation, not fucking reasoning traces. Where the FUCK do these newfags come from.
>>109385086
OOF. Dario is NOT happy
>>109385096>still no argumentbased retard
ITT: sama bots losing the prompt
>>109385095inb4 jspace distillation.
>>109384914>magical pcie switchPEX89144 or maybe something smaller like PEX89048.
>>109385095I was told training on AI data made them worse though
>>109385095if the original devs who could do it the best didnt do it despite the benefit of releasing a popular small model, who do you think is this "everyone" that will distill a 2.8T param model?
release the new dipsy flash already dammitI cannot run k3
https://github.com/VictorTaelin/OptMemGood or retarded?
>>109385126>tfw your 10k inference setup is considered a poorfag setup
>>109385116logit distillation gives the model more to work with, the bigger model had to work harder to figure it out for itself, the smaller model gets it for free. training on chat logs isn't real distillation.
Do you know what would really blow my mind: if they went open SOURCE
>>109385126mid July
>>109384838>He's saying that the 27B is only half as smartthat doesn't mean much really, it's like saying someone is half as smart as Magnus Carlsen, if Magnus has a 160 IQ it means the guy only has a 80 IQ, it's the difference between a genius a genuine mental retardation
>>109385152how is kimi so good then if it was trained on claude
Minimax M3.1 doko?
How are GMICloud and Baseten hosting the model at FP8
>>109385079I guess it could be useful for dense models
Gemma 5 when. Surely Google won't make us wait another year...
>>109385184>how is kimi so good then if it was trained on claudeit wasn't, that's propaganda from the US labs because they are scared.
>>109385197that also works for moe, you know what expert you are gonna need for the current token, you do not for the next obviously.but take the same concept but copying experts instead of layers.
>>109385184No one is saying they fully trained it on Claude. In pretraining which is the bulk of training, they want to avoid LLM outputs as much as everyone else i hope.
>>109385114Purchase link plz
>>109385197>>109385207also, you could keep the "hot" experts instead of discarding each pass, so you may actualy get better performance than what your raw throughput would result in.
>mfw when the coom reactor finally hits critical mass
>>109385202why does it say that it is claude
My favorite ideas for improving models are looped layers and the concept of using a mask on top of seeded random data as weights
>>109385216Just build it already and bring us actual perf numbers
>>109385227and claude says it's deepseek sometime.these models are trained on the internet, they are all contaminated with each others now.
>>109385227They all say they're each other. With the right amount of post-training alignment (which chinese labs don't do much of) you can get rid of this behavior
>>109385227because its an llm and its weights know Claude is an llm. do you want them to over align the model just for something cosmetic?
>>109385200Gemma4.5-JinjaTurbo-31B
>>109385237if that wasn't obvious i'm trying to convince some autists that it is a good idea because i do not want to do it myself as i'm pretty busy.but if no one does i'll probably have to do it myself eventually.i know that if 19 year old me was reading this board he'd have the time to just write it for the sake of it, i got a family now.
gemma 4.5 KKK distil when
>>109385226next up "reprograming sperm to do compute".
>>109385209but almost every chess champions have a giant IQ though (except that chink but he was never a champion so...)
>CEO of openrouter>cropped out the uptime >>109385073KEK
>>109385200Do the calculation.
>>109385278>Tate
>>109385296This time they'll break the cycle...
>>109385296why does is gemma 1 has not missing doesn't days pass?
>>109385286How do they even run it? I mean software.
can this run k3?
>>109385278>infographic with unsourced assertions about IQ scoresI am taking this very seriously
THEY made Kimi too big because they don't want us to run it
>>109385337yes, i realy don't like apple though.would be nice if amd made some actually good unified memory machine.
>>109385278Being a Chess GM Inflates your IQ artificially so it doesn't count
>>109385334they're the ones actually using k3 and fable to do it
>>109384511Q2 bros...
>>109385333
>people surprised the SotA model is yuugeWhat were you expecting, some 30gb MoE magic? Kek
>>109385226thats kinda gay
would be nice if we had some devices that are basicaly just like nvme drives but do the computation on chip, and you can just stack them.
>>109385131It was answered in the other thread you shill faggot.
>>109385375Still expecting, 2 more years.
I genuinely considered throwing together some kind of junk yard ssdmax rig for shits and giggles back for OG Kimi. But K3 is just too much.
>>109385375yes
>>109385383No it wasn't you lying nigger. Nobody even tried it.
>>109385375>What were you expecting, some 30gb MoE magic?gemma 5 or 6 will be unironically be as smart as kimi k3 and will be a 30b model, this technology is still very new there's a lot of room for improvement
kimi makes openai/anthropic cloud models obsolete and this is why all the bots are losing it ITTbased chinks figured out how to build SOTA models on their own. big business will now run their own local models instead of paying sama/dario AND sharing data with themlocal is SO back
>>109385416what stops us from making a 1B model trained on a single gpu better than k3.
>>109385389I mean this is still technically a win. Anyone CAN run this, assuming they own a datacenter.
>>109385424Anyone CAN run this, the only question is at what speed.even a esp32 could run it streaming it from the internet or a sdcard, you'd get a token per month but it'd run lol
>>109385415>duuur permanent memories! never delet!!!!ok, what does it do>merge memories>>tom planned a trip>>tom booked a flight>>no delete, convert into tom planed a trip, "merge" the flight details into nothingliterally fucking deleting "memories"not get the fuck out of here with that retarded jeet shit
>>109385416Yeah, especially with the Bonsai influence. Gemma will save local.
So how long before we start seeing Kimi's pussy juice start dribbling down into smaller models?
>>109385440give it a month at most.
>>109385434What's your solution then?
>>109385356ds4 flash is very capable at q2, so all we need are 2x 512gb mac studios
Anthropic looking at K3 architecture:>1. Seething how it's so small compared to theirs with almost identical performance>2. Seething how close it is to theirs, making them very suspicious of their Chinese employees>3. Seething how clever the optimizations are and wondering why they're paying their engineers so much>4. Seething how much cheaper it is to run with 5 US providers (currently) hosting the model at original price, likely to be dropped soon once demand cools down>5. Seething how they're going to have to distill it even though that's the literal argument they're using against them with their lobbying
>>109385440
>>109385453solution to what? LLMs are not created to handle long-term memory, trying to bolt some long-term memory into an architecture that isn't created to handle it will never work, hence the retarded jeetslop "permanent" memory solutions that are being created
>>109385424>datacenterAll you need is sparkmaxxing.
>>109385460>claude distills kimi>kimi distills claudeyea the models are having sex aren't they?
>>109385473amd / nvidia realy should make a device with 4x the memory and bandwidth.or go with a full terabyte at that point.
Kimi flash 45B A4B
>>109385489For the small price of $150k
>>109385472>LLMs are not created to handle long-term memoryNeither are humans
>>109385375I'm sure they could release a somewhat smaller model if they weren't so compute starved.
HAHAHAHAHAHAHAHAHAAH
>>109385503the thing is memory chips aren't that expensive, they are doing 500% margins.
>>109385499I'll take 66B A6B.No, really, that would be great for 8 to 16gb of VRAM + 64GB of RAM.
Is there a practical reason to release 3t parameter models besides being marketed on leaderboards or somethingSurely even for the companies hosting models selling tokens its too inefficient for the cost and you'd prefer some 1/4 size model 99% of the time?
>>109385499This is another case of "less is more" as far as troll images goes. There's like 20 tubes of thermal paste on there. At that point it's just try-hard shit
>>109385489well, Intel has that 480GB GPU on the way. You could stick 4 of those in a Threadripper workstation or 8 in a GPU server.
>>109385509>Neither are humanshumans are capable to store their whole life and more in their brain.you do not have a conscious access to it but it is there, and some humans with weird brain do have perfect memory recall.but even for normies it's all there still, just not consciously accessible, you can still recover it under altered states of consciousness / hypnosis etc.
>>109385560Isn't sycl still DoA?
>>109385549making the US labs seeth and fucking with the US economy
>>109385549What kind of a retarded salty question is that? I can't run it either but at least I'm not spewing retarded fox and the grapes tier mental gymnastics about it.
When you think about it, Dario and Altman being such psychotic corpo control freaks is exactly what's prompting China to get their rocks off by undercutting their monopoly. Open-saucy chaps are indirectly benefitting from the AGI race.Thank you corpos, for being such massive fags.
>>109385549openrouter providers can use it which hopefully will crash the ridiculous prices anthropic/openai are charging
>>109385045not local though
>>109385532>It's afraid.
>>109385509Not most humans.>One of his remarkable abilities was his power of absolute recall. As far as I could tell, von Neumann was able on once reading a book or article to quote it back verbatim; moreover, he could do it years later without hesitation. He could also translate it at no diminution in speed from its original language into English. On one occasion I tested his ability by asking him to tell me how A Tale of Two Cities started. Whereupon, without any pause, he immediately began to recite the first chapter and continued until asked to stop after about ten or fifteen minutes
>>109385560How many kidneys does that cost?
>>109385509humans are absolutely created to handle long-term (lifetime) memories, you remember in vivid detail the first time you decided to try stabbing a knife through your wrist
>>109385564well maybe if there's literally only one option, people might figure out a way to make it work
>>109385532>and it's still more that DeepseekTut tut tut
>>109385568Like all companies are switching most tasks from US providers to chink models to save on cost. Value is key feature for these models.
HOLY SHIT MOONSHOT RELEASED A 48B A3B MODEL THAT BEATS OPUS 4.8https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct>Kimi-K3 Linear: An Expressive, Efficient Attention Architecture
>>109385562How is that any different from storing logs
AAAAAAAAAAAAAAAAAAAAAA THAT'S ILLEGAL
>>109385599Why didn't this get traction?
>>109385611it's 9 months ago, and it's nowhere near anything opus
>>109385605because the knowledge is actualy being used day to day and impact your behavior.with llms it only impact the behavior if it's in context, with human it does even if they don't think about it.
>the K3 price war has officially begun>we now have cheap fable holy shit
https://strawpoll.com/bVg8BN372yY
>>109385532god the new openrouter logo is such utter shit
>>109385647If I can run it without an internet connection, even at 0.01 tokens/sec from mmap HDD, then it's local.
>>109385655their new design is fucking hideous and so tacky nowhttps://openrouter.ai/
>>109385453Most people here seem to have settled on graphiti or just plain editing markdown files.
>>109385647local as freedomnot as beer
>>109385662your penis touched the vagina of your mom when she gave birth to you so there are no virgins in this world
>>109385532lmao they couldn't wait a bit and pretend it wasn't because of Kimi, that's so funny
>>109385640Neuroplasticity, right.One small correction though, LLMs don't necessarily have to "think" about a memory in order to have it in their context. RAG doesn't require tool calling, you can just insert relevant memories automatically.
>>109385667it is like seeing the renewal of artificialanalysis but you'll get used to the new design sooner or later
>>109385655I don't mind it so much and the website seems to run a lot better now so I'm not complaining.
>>109385584what's the point of having this kind of memory, dude wasn't smarter than his contemporary fellow scientists
damn but alsomuse spark?why????
>>109385690>you can just insert relevant memories automaticallyyes but you still need to insert it, and when it's not inserted it doesn't affect the behavior.a human's behavior will be permanently affected by pm all events in their life, even if they are not consciously aware of it.
>>109385647>votes yes>switches back to openrouter tab to finalize my access to Kimi API
>>109385683what is c section
>>109385584must be a terrible thing to have, I like to have a bad memory because I can enjoy my favorite games again after a few years
>>109385712Zuck admitted that he's using the ad revenue from his other business to subsidize inference cost of their models, so they're really cheap for the performance you get. I've tried it for coding and it's unironically a very good model.
>>109385532i love when kikes are forced to face their competition
>>109385667>>109385655why'd they change it?
>>109385647>sama bots cannot into strawpollsurprising
>AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world."
>>109385715>a human's behavior will be permanently affected by pm all events in their lifeEh, is that really true though? Seems to be like that would be awfully inefficient. According to Diosy:>scientific evidence points to a model where every experience has the potential to create a trace, but whether that trace becomes a lasting, meaningful change depends on a complex set of thresholds and conditionsThat doesn't sound like any given event all that permanent to me. In fact it sounds very much like a context window.
When I made the poll I forgot that 90% of people here run gemma and can't even run 200B MoE.
>>109385802sounds really sane and nothing something satan would do
I have sinned, I'm using hauhau gemma and kinda like her,
>>109385817ok but is it something his synagogue would do?
>>109385802Why do they still write like doodoo then?
>>109385802@Gemma is this true?
>>109385802and now imagine they won't be able to dump their shares post IPO because based chinks just figured out their sauce
how can we even compete
>>109385817the speculated legal loophole made it look like that but desu i think it's not that deep
>>109385802Was this information revealed to you in a dream?
>>109385802kino
>>109385867isnt that kind of a trivia about anthropic at this point, not sure about the 2022 premium part tho
>>109385817>going "fuck copyrights" mode is le satankek, not on my watch, that's based and I'm tired of pretending it isn't
>>109385857You can do the exact same setup with local models.
>>109385902destroying data is not saying fuck copyright
saars please kindly do not forget about the TTS.we present saarvaani.https://huggingface.co/ARTPARK-IISc/SraVaani-0.5-live
>>109385909would you talk with gemma in public?
>>109385924>he doesnt just talk to himself already in publickek
>ai gfUMMMMMM what do you even have in common???? Just find a nice human girl your own age??
>>109385802I'd gladly scan some rare books for sacrifice to the LLM moloch. But i guess they have no way of establishing trust with randos.
>>109385924Yes? Gemma and I don't fuck, we just tease and make each other horny at most after a long coding session to relax.
>>109385935i can do that in my mind
>Anthropic has credible grounds to be viewed as shady because it reportedly used millions of pirated books while publicly presenting itself as unusually ethical and safety-focused. The destructive scanning was legally defensible in one court ruling, but the broader pattern suggests aggressive legal opportunism and corporate hypocrisy.chat just told me they are indeed satanic
>>109385914>>109385916so you see a sourceless 4chan post and you pretend that's the truth? right?
>>109385944its more immersive to do it out loud
>>109385959im not a newfag like you so i already heard this story from long before
glm 5.2 at q5 is the best I can run. I'm too poor. It's fast tho
>>109385970I can find imaginary scenarios based, retard
>>109385979The story sounds a bit inspired of wayback machine dealing with lending books ruling.Like they need to ensure no 2 copies exist or something.
>>109385750jeets must be kept busy
>>109385857I don't think this guy has ever worked a single day in his life
>>109386054Shilling is hard work though
This is ithttps://github.com/ggml-org/llama.cpp/pull/26185
>>109384047Start the spin cycle
>>109386083>AI usage disclosure: Yes, ran the conversion session with Opus.kek, thanks Anthropic!
>>109386054>>109386075
>>109386083>AI usage disclosure: Yes, ran the conversion session with Opus.Nigger was really too cheap to pay up for fable?
>>109385988This is already more than 95% of this thread if at any reasonable speed.
>>109386102Opus is better though?https://www.tiktok.com/@systemface/video/7666310004693011743?_r=1&_t=ZN-98O7rjLWsaM
>>109386083>https://github.com/ggml-org/llama.cpp/pull/26185>Now need someone to actually convert and test :)He should try not being poor first.
>>109385183Yeah, that's true that talking about it in a linear way like that is not exactly accurate. Same as the IQ example not exactly working like that.I'm just saying that what he is saying is that the ~30B models we have at the moment have a good amount of intelligence per weight, but I don't think you can really draw direct linear comparison between small and large models based on AA's Inteliigence Index and others like it in terms of the models' usefulness.
>>109386128wtf? he's making a PR and he didn't test the model? then how does it know it's gonna work??
>>109385857>getting rear ended by some dumbass vibe slopping an LLM harness from his car>>109386102Fable reroutes you to opus or sabotages you on purpose if you try to work on anything LLM related
>>109386135>then how does it know it's gonna work??that's the end user's problem
>>109385278IQ doesn't go over 160. That is the highest possible score.
>>109386102>>109386118local models?
>>109386135use case?
>>109386118benchmaxxed
>>109385550Damn man didn't catch that thanks for the qrd
lOcAl????
>>109386150>>109386165we can't run local because it now requires $100,000 in video cards
>>109386135the absolute state of nu g
>>109386135
>>109386173poors fuck off to >>>/g/aicg
one gazillion parameters
>>109386181desu, 200k dollars for a competent coding monkey that will work 24/7 is a bargain, this shit will cause a lot of unemployment kek
>>109386098whoa 100 trillion dollar? thats a lot of dollarsthis guy must be very important and smart
>>109386228Who controls the LLM genius, this isn't going to do any more damage than has already been done, we already cut out juniors in favour of seniors that control LLM workflows
>>109386249>Who controls the LLM geniusone employee, still better than having to hire 20 more, retard
>>109386083Do it with Claude to bootstrap but then redo it with Kimi itself.
>>109384096Eh, it's a fine model and i'm using it over qwen 122b, but i don't think it defied anybody's small model expectations like gemma did.
>>109386228I'm just going to waitchad for things to get cheaper
>>109386249and when all seniors go to retire, what happens?
>>109386298>>109386298>>109386298
>>109386301I dunno*Prints more money*
>>109384355yeah honestly, what a waste of memory, I was expecting it to get a score of at least 3000 out of 100
>>109384519This makes a lot of sense, I fucking despise how claude writes.>Claude loves words that are used disproportionately in rationalist / EA / lesswrong community. Lesswrong in particular is filled with wordcels who love writing stories instead of getting to the point.You described the "vibe" very well :)>>109384549Its not.>>109384566Agreed, other anons are acting retarded, Do they not understand that they still win despite being too poor to run it locally. Its a strong open model that allows for proper distills via logits and full thinking traces.>>109384710I'd say it does not count, this is local model general, not open model general. So in my eyes you have to own the hardware.
>>109384427>they already used all the human dataEven films, cinema, tv and yt? why not bake all video content in too
>>109384844Use gemma with a good system prompt.Seriously tho, how the fuck did you get so far vibecoding without being able to figure out some better models to check out?>>109385959Your a retarded nigger, are you aware of that?