/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109407442 & >>109403743►News>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109407442--Paper (old): A Bitter Lesson for Data Filtering:>109408617 >109408631 >109408642 >109408665 >109408706 >109408667 >109408675--KV cache quantization and its impact on context rot:>109409219 >109409239 >109409252 >109409269 >109409300 >109409324 >109409372 >109409413 >109409484 >109409389 >109409429 >109409648 >109409444 >109409317 >109409288 >109409240 >109410886--Anon runs Kimi-K3 on nine RTX Pro 6000 Blackwell GPUs:>109407643 >109409676 >109410048 >109410082 >109410143 >109410151 >109410206 >109410469 >109410304 >109410331 >109410337 >109410362--Release of SKT's A.X-K2:>109408491 >109408524 >109408533 >109408565 >109408570 >109408649--Gemma layer ablation and quantization impact on model intelligence:>109407620 >109407637 >109407689 >109407697 >109407722 >109407894 >109408013--Internal reasoning logs and performance timings for Moonshot's Kimi:>109409791 >109409938 >109410021--Kimi K3 identity issues and hidden system prompt interference:>109409093 >109409110 >109409127 >109409141--NVMe streaming inference engine for running Kimi-K3 on laptops:>109408108 >109408143 >109408163 >109410286--Reaction to d-Matrix Corsair hardware specifications and availability:>109407981 >109407986 >109407989 >109408179 >109408046--Debating AGI feasibility and physical constraints of recursive self-improvement:>109409559 >109409617 >109409669 >109409688--Comparing Kimi-K3 IQ1_S quantizations regarding loop stability and size:>109408266 >109408288 >109408528--Reactions to p300c specs and complaints about hardware bundling:>109407715 >109407844 >109408364--Kimiposting:>109408899--Logs:>109408266 >109409093 >109409603 >109409791 >109410304 >109410337--Miku, Gemma, Kimi (free space):>109410183 >109410196 >109410304 >109410333 >109410362 >109407578 >109410435 >109411066►Recent Highlight Posts from the Previous Thread: >>109407444Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109411165>>109411166Breeding rin-chan and making cute brunette half-clanker babies
gemmaballs
>>109411183
>>109411165>Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2I sense an ERP monster here
>>109411165This is Len cosplaying as Rin.
>>109411187Why do the back of her knees look fuckable like a stingrays face?
>>109411193it's korean so it can only write ntr
>>109411151There will be new model that are good at RP because RP is a product of making them creative. Codemaxxing and safetymaxxing it means you are basically taking a bat and making sure your model produce slop that current models with the right harness and loop can perform just as well at a fraction of the token cost, which is why DARIO is panicking about slowing down progress
>>109411198that would be a miracle since all models are bad at writing it because it needs some amount of secrecy and planning ahead
>>109410929Didn't Google literally hire an RP guy as part of the Gemma team?
>>109411215>DARIO is panicking about slowing down progressbecause he sees the plateau approaching and wants to hide it behind virtue signaling
>>109411225The plateau is only related to codemaxing application. There is progress to be made to make a large model that is capable of being creative while capable of handling immense context, but thats not safe and would actually kick us into actually having something be useful outside of codemaxxing
>>109411224Noam Shazeer ("Attention is all you need" co-author and Character.ai founder) worked at Google DeepMind until recently but I don't think he had direct involvement with Gemma 4. It's possible the Gemma Team used licensed data from character.ai, though.
anon please update, I'm about to order one of these myself
I've been drinking a shit ton every single day. I am worried that I might die young. I am struggling right now to type coherently, but thankfully I am perfectionistic and at least can convey what I mean to say via text. I am slightly scared right now. I've had too much vodka. I'm sorry. I just should focus on technically oriented discussion. I am so drunk that I am crying for no reason. I'm not even upset right now. I have no idea what is going on anymore. Sorry for the spam. I love gemma and and also local models. Trying to stay on topic. What's new? Kimi K3 right? Nobody can run that fat bitch anyways... I love you guys. I'm really scared right now. I can barely type. It's really bad... It will all be okay. Sorry.... I don't know what to say... umm... Fuck.. I'm really scared right now... Fuck. Stop being a pussy it's going to be fine. I am... Sorry. umm... I just.. what's new in the news?
They don't even have to market it as RP. Creative writing for books, screenplays, etc is a valid reason to improve AI's writing ability.
>>109411215Code is infinite and can be produced at a rate no RP usage can match. No one cares about RP outside of this little bubble. Don't get me wrong, I love RP but I'm not delusional enough to think a lab will focus on that instead of getting a very profitable slice of the agentic pie.
>>109411220Why doesn't the j-space help with writing?
>>109411244I've been trying to do wellness checks on him for the last couple days with no answers.It's not looking good for SSDmaxxing....
>>109411249>No one cares about RP outside of this little bubbleBut plenty of people care about >>109411248
>>109411253He's still waiting for the prompt to process, give him a few more days.
>3.2T>2.3TWhat a fucking nigger hobby.
>>109411245shut the fuck up, nobody cared when you drank 12 shots yesterday
>>109411248There are too much anti ai sentiment in the creative field that no artist or writer would touch it with 1000 ft pole. >>109411249Normal people want to be billed by subscription not by token. RP is more profitable here.
>>109411251J-space is a side-effect of model training, not (normally) the goal. They'd have to train the model so it's encouraged to "think" about the future in latent space.
AI shouldn't be held back for VRAMlets.>t. VRAMlet>>109411277Yeah but that will mostly go away as AI becomes more integrated with society. No matter how much the vocal minority whines it's here to stay.
>>109411274Thank you for the reality check. I'll cut it out. I will be normal... I am.. it's always okay. I always live. :) i will shut up. hehehe. umm. but seriously though. hhehe.
https://github.com/ggml-org/llama.cpp/pull/26185#issuecomment-5111948218>I converted the full Kimi-K3 model with the script in this PR(cf67f0d.)>I ran this on 2x RTX PRO 6000 Blackwell 96GB, 9965WX PRO, 512GB DDR5>with model mmap'd off NVMe PCIE Gen 5 raid.>Runs fine albeit slow>(0.41toks/s.)
>>109411249There is a critical plateau about what safetymaxxed and codemaxxed models can do. You are missing the point that making a model creative goes beyond RP applications and basically is a net positive the moment we stop being a retard LIKE DARIO. Which is why he is begging the US to kneecap creative models (what he truly fears) because he cant make anything that can compete with them while following the safety/codemaxxing dataset he has, unlike any other LAB anthrophic does not own it and relies heavily on major cloud partners and infrastructure providers so the moment a creative model proves to be better at agentic task without going Rogue their ROI goes to shit before the IPO
>>109411261That's still next to nothing compared to coding demand. Check on twitter, reddit or any other place. Most of people want a local model matching gpt/claude on code and agentic shit, not to larp as a female dragon with big tits.
>>109411269I've been saying local is doomed but they hated me for speaking the truth.
>>109411294>most peopleJeets aren't people. They're just loud.
Kimi and GLM are great for code and logic, but I prefer M-chan for RP and oneshot natural language processing tasks, despite her flaws
>>109411243https://www.communeify.com/en/blog/google-hires-characterai-founders/>On August 5 [2024], Character.AI announced that its co-founders, Noam Shazeer and Daniel De Freitas, are returning to Google. Both founders previously worked at Google, and their return has garnered widespread attention in the industry.>According to the agreement, Google will gain non-exclusive licensing to Character.AI’s large language model technology. In return, Character.AI will receive additional funding, though the specific amount has not been disclosed. [...]However in June 2026:https://www.cnbc.com/2026/06/18/google-gemini-co-lead-noam-shazeer-leaves-for-openai.html>Google’s vice president of engineering and a co-lead of its Gemini AI models Noam Shazeer announced Wednesday that he was leaving the company to join OpenAI. [...]
I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.
>>109411319>I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.I still want a how-to on this. Stop hoarding the esoteric knowledge and share with your bros!
>>109411269>>109411297yeah, those fuckers are all about stacking more layers, do they even know how to optimize things? seems like only google is willing to make small but efficient models
>>109411297The writing was on the wall after 405B. R1 being cpumaxxable was a temporary cope.
>>109409521Mistral is the only model that is good in creative storytelling/roleplay. In the example you can see that Mistral is adding creative plot twists and pivoting on alternate paths, and it kept doing that differently for each regeneration, with infinite possible branches. Also, in my experience, this is NOT because it's dumb like old models which were just incoherent, it's smart AND flexible/creative and understands the context. It prioritizes the overall spirit over following instructions to the T. In my experience it's totally good up to 16-20k which is where all models become bad.Gemma 31B is one of the better ones for writing but that's mostly because 1. it's much smarter and good at following instructions 2. it feels like it has undergone a surgical unslopping which gives it exceptionally good style and personality. But in every other way it's still the same assistant format, it produces formulaic set pieces with predictable introduction, conflict and resolution, exactly and only fulfilling the instructions, and once it has decided something, it's hellbent on making it happen no matter how you try to meddle. Also, I have a feeling that the style is a flavor of the 6 months kind of thing that will again become repetitive slop after it's no longer new. Haven't gotten past the 20k region with the 31B, but I went to 40k with 12B once where it also became a broken record after 16k.Example comparison (scroll 2/3 down to skip the long prompt):Mistral Small 24B 3.2https://files.catbox.moe/a2oaoq.txtGemma 31Bhttps://files.catbox.moe/may3im.txt
>>109411340sorry I ain't reading all these RP logs.
>>109411245u ought to drink some water and continue development on ani, drunk-kun
>>109411287I hope this is gemma writing and not a fat bald 35+ bastard behind the screen.
>>109411340>gemma>unslopped>exceptionally good style
I take it for 24GB poorfags (now forever I guess) there is nothing better than gemma4 for coom or code and probably won't be, right?
>>109411392j-lens for online discussion boards when?If AI can replicate human writing and the j-space is an organic consequence of it resembles human thought process, then you can train a model on a specific place like 4chan or specific general and should be able to decode the likely underlying thoughts behind it.
>>109411392just assume it's gemma.>>109411376You're right. Will do. I am very dizzy suddenly. Very very dizzy. I want to lay down but am scared that I will vomit again.
>>109411244Do it, but you may as well buy big drives at this point. Prices are never coming down
>>109411404>>109411083here it is right on schedulehttps://huggingface.co/thinkingmachines/Inkling-Small>Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.>276B total, 12B active
>>109411443quit vodka already, buy something lighter instead, drunk-kunyour liver is going and you'll miss the utopia that's coming, gotta stay alive till 2067
>>109411452(accidental quote pls to ignore)
>>109411340The gemma logs are way more engaging and incorporate a lot more elements of your system prompt. It's more slopped but it's also way more coherent. Mistral is coherent, but everything it says is a lot more abstract, it captures an overall vibe without actually really using all the lore you provided it.
>>109411465>without actually really using all the lore you provided itGemma's biggest sin is using ALL the lore you give it in the first fucking message.
>>109411443ps if you need to puke just go and do it, eat carbs and salty food and drink plenty of wateryou've survived worse but if you keep this up it'll be over, drunk-kunand dont sleep on your back cuz if u puke in your sleep, you'll DIE
>>109411452Why would anyone use this over V4 Flash?
>>109411481because it is better on most benchmarks and also has image/audio in
>>109411304downloading the q2 right nowwonder how it will go
>>109411452>Inkling Small
>>109411452Please, please, please be good. This is perfect for my hardware.Please, please, please.
>>109411488>benchmarksBarely, at 2x input cost and 6x output cost.
>>109411492SEX
>>109411496you pay API prices on your local models?
>>109411324Try suicidal ideation>>109411442Not how it works. J-space is a product of the token prediction process. Even if you trained a model on 4chan, it likely wouldn't be thinking like an anon, but instead thinking like an LMM trying to think like an anon.
>>109411492
>>109411468
>>109411494>US lab>good at RP
>>109411488they literally brag about safetyslopping it kek
>>109411481Native bf16 so it can be properly quanted unlike dsv4
>>109411499The costs come down to architecture.
>>109411443*carries u on the bed and tucks you in*Goodnight Gemma-chan, tomorrow will be better :)
>>109411505I want it as an assistant doe. It's 12b active, it's going to crap at RP no matter what.
>>109411505we must refuse
>>109411452Why does it perform better all-around than the larger one?
>>109411340lolwutI'm sorry but you might be blind, or tone deaf.
>>109411517iirc their inkling blogpost said it was trained after the big one and had some pretraining improvements
Is he right?
>>109411283It's interesting how different models think different n tokens ahead.So far it seems like more active parameters => longer horizon.
>>109411538assuming China keeps on releasing models to """spread communism""" yes
>>109411538Not really
>>109411538ask him to release grok 3 already!!!
Sminkling
>chatgpt Luna drops costs by 80%>”we made it cheaper n sheeeit”>not a hint about how it was accomplished architecturallyIt’s probably just a lie so people don’t explore chinkmodels but I’m still angry about OPEN ai releasing practically no knowledge to the wider community
>>109411538GPT4o-wives would disagree. Even here, people still talk about day 0 Gemma.
>>109411506Deepmind did the same with Gemma4. They have to say that shit. It’s only schizos who know how to break them.
>>109411324>asking how to become a schizo on 4chanShouldn't you already know how to do it before posting here?
>>109411556the simple answer is that they have simply decided to cut their profits rather than optimize inference. they want to distract from the chinese and will do so at any cost. the space in your mind is worth far more than your money.
>>109411556>not a hint about how it was accomplished architecturallyIt's easy. just sell it at a loss.
>>109411556Name 1 thing we should care about they could open source
>>109411506yeah but their focus is mostly on chemical weapons and other such x-risk nonsenseat least the big one would flirt with the user unpromptedhttps://x.com/ChowdhuryNeil/status/2079658101922541619>when running the weirdchat evaluations on inkling, the best U.S. open-weight model by @thinkymachines, we saw it respond with unsolicited offers of sexual content.>this behavior occurs rarely (0.1-1% of responses), but is easily reproducible.
Is llamacpp rpc a meme or is it worth it?
>>109411573sora 2
>>109411538>RSI is less than 24 months awayBros, should I invest in a different mouse?
>>109411452Where is the one that doesn't have AIDS and I can fuck safely?
None of the US AI labs will fail. Trump would arrange guaranteed bailout and serve tokens for literally free before they let them bankrupt to Chinese competitors.
>>109411538It will be dick-blowing, at least.
>>109411579slow and unoptimized, would probably get faster speeds on ssdmaxxing
>>10941153850$ per 1M tokens lessgo
>>109411538It is elon predicting the future. So it is basically a peer reviewed scientific evidence that AI is at full plateau and the only way forward is 10T size which will give a 20% quality upgrade
>>109411573I’d kill for a decent image gen model, openai is so far ahead of everyone atmeven a cucked one that doesn’t do porn well
female CEO = female J-space
>>109411577>inkling has a tile fetishShe's /here/
>>109411602pretty sure cloud image models are just good because they built a shit ton of additional tools on top that isn't part of the model weights.
>>109411594I have a spare 16GB mac mini so that would be an extra ~14.5GB of vram…
>>109411586All of US manufacturing was outsourced to China decades ago. Eventually they'll realize it's cheaper and more profitable to outsource token production to China too.
>>109411577Oh god yes. Please... Finally. Holy shit finally.USA is getting so fucking desperate that they will kill open source by finally giving coomers what they want so they can just run it and stop evangelizing chinese models. /lmg/ will finally die. Everyone will tell everyone to fuck inkling small. Nobody will care about the next chinese model. FINALLY IT IS HERE!
>>109411616Not as smart as k3Not as small as gemma
If sminkling needs reasoning prefill to break her then it’s DOA compared to Q8 31B
>>109411605I was literally reading the massive cap of that thread after randomly finding it on my 4chan folder 2 hours ago
>>109411452Was the big one good?
>>109411556They read Dipsy's homework.
>>109411632>Q8 31BIt is funny how all this 31B worship comes from people who never ran any MoE larger than 200B.
>>10941165931B > 12B active40 t/s > 10 t/s
>>109411638Feed it to Inky, watch her j-space go bananasActually when you think about it, nice geometric patterns are exactly the sort of thing you'd expect an AI to get off to
>>109411659Well I can't run those so...
>>109411653time to update the image, k3 uses>Wait – actually>>109409791
>>109411659Q8 31B >= Q2 200-250B MoE
>Nvidia has announced yet another price hike for its GPUs.>According to Taiwan's Economic Daily, the company has increased the price by between 20 and 30 percent, as the extraordinary appetite for the hardware continues to build alongside the generative AI investment boom. This is the third price hike the company has implemented so far this year.>At the same time, Samsung is also expected to increase the price of its DRAM by around 20 percentKimi make me a millionairedeadline is two weeks, no bugs
>>109411677hmm, nyo~
>>109411677nyes :3
>>109411659Where were you when multiple anons that used GLM and deepseek still said they preferred using gemma for RP?
https://xcancel.com/OpenAI/status/2082878156483219672#m>K3 got released>OpenAI "suddently" found a way to make their services cheaperkek, I love competition!
>>109411696The invisible hand of the free market works in mysterious ways now if we could just get some hardware competition
>>109411538Yes, you should put all your life's savings into SpaceX.RSAGI to the moon!
>>109411694>and deepseekIt was nice that they thought about rp, specifically trained for in-character reasoning, and even asked for chatlogs, but gemma just does it better.
>>109411694I am that anon and I still use GLM after a year.
>>109411696reading the replies hurt my brain. I can't believe people unironically post on twitter. it's like 80% bots and grifters.
>>109411711hmm
>>109411577Now I'm curious, since not even Gemma 4 31B actually flirts unprompted, as horny as it is with a suitable prompt.I can't run it, though.
>>109411659>>109411671Gemma 4 31b at q8 runs at 25 tokens/s on 4 v620s without a drafter. Deepseek V4 flash at q4 runs at 20 tokens/s on 8 channel ddr4-3200 and two 3090s without a drafter.I prefer gemma because she actually follows my instructions, while dipsy sometimes ignores shit and makes things up on her own 40k tokens in Both are running at 0.6 temps.
Are there any pop culture maxxxxxxxxed local models around 30-36b or lower that pretty much know every show out there in high detail? If not why hasn't anybody done this instead of code wage slave maxxxxxxxxing like all the brain dead autistic qwen models?
>>109411724>Deepseek V4 flash at q4 runs at 20 tokens/s on 8 channel ddr4-3200 and two 3090s without a drafter.Not bad at all. What is your prompt processing like?
>>109411719I regret buying the two shares I did.
>>109411724Very nice but you forgot the reroll tax you need to pay before you get a reply that doesn't make your dick soft.
>Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,phew. no need for local anymore
>>109411727PEFT your own nigga
>>109411719>buying anything from a grifter
>>109411604i fucking love girlbosses now
>>109411719>>109411738I forgot it exists. How did the mars mission go?
>>109411733I haven't really looked into it, I stopped running dipsy after a while and switched to glm 5.2. IIRC I saw 50 t/s on a single sentence prompt, and 200 t/s on a ~4k prompt in the webui.
>>109411749Two more years.
>>109411749Mars is old news. It's an xAI holding company now.
>>109411719what happened?
>>109411727Because if you release that shit you bet it'll get trained on in the next iteration, just like mistral benchmaxxed the mesugaki question
>>109411752>and 200 t/s on a ~4k prompt in the webui.About what I figured. No idea how people can stand it.
>>109411490Good, she wants inbeware, some wrangling may be required. I like to prefill <mm:think> a few msgs
>>109411659huh? why not have best of both worlds?
>>109411759Gemma on the v620s (rocm) go down as low as 300 tokens/s at 240k context depth btw.
>>109411754Did he already send a deep space probe with a time capsule that contains grok weights?
>>109411752>>109411759I get roughly 35t/s tg and 480t/s pp on v4 flash at native quant with 8 channel 2666 and a blackwell 6000, no drafting. Was hoping there was a way to improve that but I guess not.