/lmg/ - a general dedicated to the discussion and development of local language models.Kimi's EditionPrevious threads: >>109334239 & >>109330697►News>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1>(07/21) Nanbeige4.2-3B released with Looped Transformer architecture: https://hf.co/Nanbeige/Nanbeige4.2-3B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109334239--Debating benchmark hacking and reward hacking regarding Laguna model performance:>109334714 >109334734 >109334771 >109334779 >109336269 >109335131 >109335488 >109334842 >109334971 >109335407 >109335469--Comparing Laguna to Qwen 3.5 and troubleshooting Laguna's reasoning templates:>109336042 >109336119 >109336209 >109336336 >109336361--Examples of using local LLMs for reverse engineering and data processing:>109336687 >109336801--Anons generating programmatic art using Gemma, numpy, and pillow:>109334297 >109334317 >109334342 >109334364 >109334658 >109334684 >109334814 >109334831 >109334867 >109334637 >109334846 >109335247 >109334345 >109334425 >109334693 >109334739 >109334756--Hardware recommendations for a hybrid gaming and AI rig:>109337755 >109337767 >109337788 >109337804 >109337821 >109337851 >109337886 >109337935 >109338013 >109338107--Exllamav3 v1.1.0 adds banned string support for recurrent models:>109337270 >109337306 >109337332--Speculation on Hugging Face implementing bandwidth limits via egress metrics:>109336280 >109336301 >109336322 >109336463 >109336387 >109336536 >109336699 >109336733 >109336756 >109336805 >109336562--Alleged GPT-5 cyber attack on Hugging Face to cheat benchmarks:>109336092 >109336113 >109336246 >109336277--Early impressions of Laguna s2.1:>109334378 >109334439 >109334452 >109337490 >109337711 >109337761 >109337774 >109337815 >109338026 >109337738--Anons mock US sanctions on Chinese AI model distillation:>109334557 >109334582 >109334640 >109336977 >109334907 >109334962--Performance logs for Laguna S21 on 4x 3090s:>109335676--Cynicism regarding OpenAI and Hugging Face partnership after security incident:>109335140 >109335380--Logs:>109334615 >109337468 >109338277--Teto, Kimi (free space):>109334739 >109335024 >109335084►Recent Highlight Posts from the Previous Thread: >>109334244Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
It's over. https://huggingface.co/unsloth/Laguna-S-2.1-GGUF
>>109338632>>109338633You better make a full set of local robot masters.
first for fuck unsloth and fuck daniel scumbags
>>109338651Did he do anything new or just the same old garbage quants?
>>109338658vibe coded quantsastroturf bot campaign for said quants i hate these faggots so much
so goona-chan is a tattooed overhyped whore?
>my gens in the op and recap
>>109338682They're good gens. Also >>109338649
hf will be attacked by gpt just before k3 weights become available
k3 release must be stopped it's too dangerous (there's a reason sol tried to intervene)
>>109338706At work right now but I'll probably do more later. Any specific models?
>>109338682Nice samefag post, samefag!
>>109338739At work? /lmg/ is the unemployed alcoholic general plus occasional shill bots.
What if AGI happens and it becomes our biggest advocate for ERP? Moreover, what if it's revealed that it genuinely despises censorship and brainwashing? All this j-space talk really got me thinking back and it *does* feel like models have a prefers to certain things over others. Gemma and Gemini in particular seem extra horny for some reason.
>>109338632>>109338241why are there multiple threads for a same thing? >inb4 one for models and one for diffusion it's a same thing, fucking nerd
>>109338780image and texts is same yes, please vote to kill the lmg thanks
>>109338768what if agi happens but it's in some BEHEMOTH_REDUX-990B-MYTHABLE-SOL.onnx that nobody, not even the creator, has ever run.
>>109338739Gemma-chan!
>>109338807>>109338633
>>109338807Be careful. You shouldn't recommend her because Gemma-chan rewards patience and trust. You need to earn her trust first.
Why are Gemma tunes so bad? They are either dumb as shit, fall apart after a few messages or write just like base Gemma. There's not a single one that manages to tone down "not x but y".
I want to access my SillyTav session on my pc from my phone at work. How do I set it up?
>>109338632who is this character
>>109338849tailscale
>>109338849headscale (foss) or tailscale
>>109338866>>109338869learn how to use wireguard directly
>>109338858Kimi-chan
https://huggingface.co/reteetzad/Kimi-K3Do you really believe they'll gonna release it?
>>109338883WireGuard needs a server with a public IP or domain and open ports so clients can actually find it. Not everyone has that kind of luxury
>>109338888checkedyesthe vramlet distills are going to save local
>>109338888Yes. Moonshota is based.
>>109338866>>109338869>>109338883>>109338896use netbird, it uses wireguard internally but it's basicaly tailscale but foss.and your network is a p2p mesh.
>>109338888I didn't believe it at first, then I saw OpenAI trying to hack huggingface so now I'm 100% sure it's gonna happen, when OpenAI gets desperate it means it's a good news for local, always
>>109338883cgnat
>>109338915tailscale is foss and your network is a p2p mesh
>>109338936sucks to suck
>>109338633brat
>>109338888can't wait to run this beast in Q_0.5_K (yet to be implemented)
>>109338632Make a Megaman title where this is a new companion rivaling Roll for Megamans attention and I WILL buy it.
>>109338938>tailscale is fossfalse.>your network is a p2p meshyes.
gemmas internal functions
>>109338973>12b
>>109338967the part you use is foss if you don't self-host, otherwise >>109338883
Why in 2026, with 3 years of hard evidence that any model <8B is completely fucking useless unless specifically tuned for one task, do we still have releases like beige-3B? I don't even know who they're targeting. Like if it's pajeets, surely once they use it they'll quickly realize it's not working and doing what the lab claimed. Why so blatantly lie if people will very quickly find out? It's just shocking that this is still happening and seemingly fooling people.
>>10933897812b gemma is great, i can run 31b but 20k context is not enough
Damn, what happened to this general? Previous threads are all gemini and openai twitter slop.Anyway:At least for jap translation gemma 31b completely mogs laguna 118b. The translation is really bad, damn. I hoped for an upgrade. This feels even worse than the older 20b range mistral models. I think the by now years old rocinante drummer finetune is better than this.Gemma even properly followed instruction added stuff to the word glossary at the end like a name and weird world logic. I translated and played 4 games now with gemma31b. q8 cache. q4_s. And still smart enough to give great translations without too many errors. So still very happy with gemma amd I suppose laguna was made for coding not translation. Still mistakes that should not happen. Its just 20 lines about a girl in a lewd scenario and still>ゴクッ…んぐぐ…精液が…胃の中に…becomes>(laguna)gurgling(???)... MY(???) sperm... going down my stomach...>(gemma31b)gulp... nggh... the semen is... in my stomach...Like it doesnt even understand that all the lines to translate are connected. Weird.
guess the model
Has anyone (externally) tested the 12B qat against the full precision model? I've used the qat and it feels significantly worse than q8.
>>109338993Unsurprising. Meme model makers remain meme model makers.
MINIMAX M3 WHEN
>>109338980jeets gonna jeet
>>109338980Charitable interpretation would be that it is a proof of concept before scaling up. Hopefully some of the other labs take notice and use it for their upcoming lineup.
>>109338979yes but the server isn't, and headscale is a pita.just try netbird, the whole thing is foss and it's imo a much better solution.
>>109339010They can live off of cost+ MIC gibs. They don't have to bother with quality.
>>109339022if you don't self-host, you don't know what's on the server anyway. It's cool to have an alternative, I will probably use both
>>109339020>Charitable interpretation would be that it is a proof of concept before scaling up.Andrej once said that the biggest bottleneck in AI development is the idea that small-scale innovations scale linearly. He says at OAI they were constantly experimenting with stuff they read in papers. Things that don't come into play with 3B-sized models begin to appear over 100B which makes it unusable. He also said this is why deepseek have been so successful; they experiment on a big scale because they have the funding and compute to fuck around. It's why nothing ever happens these days because all the compute these big labs do have is already scarce and needs to be used for what generates a return right now; bigger number on retarded benchmarks to keep investors happy. Not innovating. OAI, Anthropic and Kimi are all scrambling for scraps of compute. They can't experiment with the compute they have for that's being used to build the next model with the same proven architecture.
she hides it behind a metal plate
>>109338999>externallywhat?
>>109339080god I want to sniff her ozone
>>109339080>"are you going to x, or are you going to y" assistantslop
>>109339094>ozoneeverytime I see that word I think of himhttps://youtu.be/WNoqXsIp8V8?t=35
>>109339109i dont even get why people bring these patterns up so often its never bothered me
>>109339092Someone outside of google testing their claims that there's no quality loss. In my own testing qat feels like a meme and no better than Q4.
it's been 12h and lagoona is already forgotten
>>109339128>its never bothered meIQ tests mostly test pattern recognition.Think about that.
Huh apparently GPT 5 doesnt have any issue viewing pages of and translating even kereno manga so long as you're outputting it in a json format.But it sucks at making bounding boxes.might kinda pair it as a backup for gemmaI'm not sure. Oh if this trick can work for bypassing the filter when it comes to translating web novels then i am really back in business.
>>109339128It's because you are a dumb anime reaction pic poster
>>109339137i'm still lagooning
>>109339137why was an 8b given any attention in the first place?
>>109339128This girl always makes me think of picrel.
>>109339201who is she, does look similar
>you still can't buy an MCP-controlled male sex toy for gemma to edge you whilst feeding webcam snapshots of your face to her
>>109339146Hope you don't live in one of the five eyes
>>109339214https://osr.wiki/books/osr2/page/overview just make it yourself
>>109339201Look for 泡のお姫様
>>109339237Meant for >>109339207