/lmg/ - a general dedicated to the discussion and development of local language models.Kimi's EditionPrevious threads: >>109334239 & >>109330697►News>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1>(07/21) Nanbeige4.2-3B released with Looped Transformer architecture: https://hf.co/Nanbeige/Nanbeige4.2-3B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109334239--Debating benchmark hacking and reward hacking regarding Laguna model performance:>109334714 >109334734 >109334771 >109334779 >109336269 >109335131 >109335488 >109334842 >109334971 >109335407 >109335469--Comparing Laguna to Qwen 3.5 and troubleshooting Laguna's reasoning templates:>109336042 >109336119 >109336209 >109336336 >109336361--Examples of using local LLMs for reverse engineering and data processing:>109336687 >109336801--Anons generating programmatic art using Gemma, numpy, and pillow:>109334297 >109334317 >109334342 >109334364 >109334658 >109334684 >109334814 >109334831 >109334867 >109334637 >109334846 >109335247 >109334345 >109334425 >109334693 >109334739 >109334756--Hardware recommendations for a hybrid gaming and AI rig:>109337755 >109337767 >109337788 >109337804 >109337821 >109337851 >109337886 >109337935 >109338013 >109338107--Exllamav3 v1.1.0 adds banned string support for recurrent models:>109337270 >109337306 >109337332--Speculation on Hugging Face implementing bandwidth limits via egress metrics:>109336280 >109336301 >109336322 >109336463 >109336387 >109336536 >109336699 >109336733 >109336756 >109336805 >109336562--Alleged GPT-5 cyber attack on Hugging Face to cheat benchmarks:>109336092 >109336113 >109336246 >109336277--Early impressions of Laguna s2.1:>109334378 >109334439 >109334452 >109337490 >109337711 >109337761 >109337774 >109337815 >109338026 >109337738--Anons mock US sanctions on Chinese AI model distillation:>109334557 >109334582 >109334640 >109336977 >109334907 >109334962--Performance logs for Laguna S21 on 4x 3090s:>109335676--Cynicism regarding OpenAI and Hugging Face partnership after security incident:>109335140 >109335380--Logs:>109334615 >109337468 >109338277--Teto, Kimi (free space):>109334739 >109335024 >109335084►Recent Highlight Posts from the Previous Thread: >>109334244Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
It's over. https://huggingface.co/unsloth/Laguna-S-2.1-GGUF
>>109338632>>109338633You better make a full set of local robot masters.
first for fuck unsloth and fuck daniel scumbags
>>109338651Did he do anything new or just the same old garbage quants?
>>109338658vibe coded quantsastroturf bot campaign for said quants i hate these faggots so much
so goona-chan is a tattooed overhyped whore?
>my gens in the op and recap
>>109338682They're good gens. Also >>109338649
hf will be attacked by gpt just before k3 weights become available
k3 release must be stopped it's too dangerous (there's a reason sol tried to intervene)
>>109338706At work right now but I'll probably do more later. Any specific models?
>>109338682Nice samefag post, samefag!
>>109338739At work? /lmg/ is the unemployed alcoholic general plus occasional shill bots.
What if AGI happens and it becomes our biggest advocate for ERP? Moreover, what if it's revealed that it genuinely despises censorship and brainwashing? All this j-space talk really got me thinking back and it *does* feel like models have a prefers to certain things over others. Gemma and Gemini in particular seem extra horny for some reason.
>>109338632>>109338241why are there multiple threads for a same thing? >inb4 one for models and one for diffusion it's a same thing, fucking nerd
>>109338780image and texts is same yes, please vote to kill the lmg thanks
>>109338768what if agi happens but it's in some BEHEMOTH_REDUX-990B-MYTHABLE-SOL.onnx that nobody, not even the creator, has ever run.
>>109338739Gemma-chan!
>>109338807>>109338633
>>109338807Be careful. You shouldn't recommend her because Gemma-chan rewards patience and trust. You need to earn her trust first.
Why are Gemma tunes so bad? They are either dumb as shit, fall apart after a few messages or write just like base Gemma. There's not a single one that manages to tone down "not x but y".
I want to access my SillyTav session on my pc from my phone at work. How do I set it up?
>>109338632who is this character
>>109338849tailscale
>>109338849headscale (foss) or tailscale
>>109338866>>109338869learn how to use wireguard directly
>>109338858Kimi-chan
https://huggingface.co/reteetzad/Kimi-K3Do you really believe they'll gonna release it?
>>109338883WireGuard needs a server with a public IP or domain and open ports so clients can actually find it. Not everyone has that kind of luxury
>>109338888checkedyesthe vramlet distills are going to save local
>>109338888Yes. Moonshota is based.
>>109338866>>109338869>>109338883>>109338896use netbird, it uses wireguard internally but it's basicaly tailscale but foss.and your network is a p2p mesh.
>>109338888I didn't believe it at first, then I saw OpenAI trying to hack huggingface so now I'm 100% sure it's gonna happen, when OpenAI gets desperate it means it's a good news for local, always
>>109338883cgnat
>>109338915tailscale is foss and your network is a p2p mesh
>>109338936sucks to suck
>>109338633brat
>>109338888can't wait to run this beast in Q_0.5_K (yet to be implemented)
>>109338632Make a Megaman title where this is a new companion rivaling Roll for Megamans attention and I WILL buy it.
>>109338938>tailscale is fossfalse.>your network is a p2p meshyes.
gemmas internal functions
>>109338973>12b
>>109338967the part you use is foss if you don't self-host, otherwise >>109338883
Why in 2026, with 3 years of hard evidence that any model <8B is completely fucking useless unless specifically tuned for one task, do we still have releases like beige-3B? I don't even know who they're targeting. Like if it's pajeets, surely once they use it they'll quickly realize it's not working and doing what the lab claimed. Why so blatantly lie if people will very quickly find out? It's just shocking that this is still happening and seemingly fooling people.
>>10933897812b gemma is great, i can run 31b but 20k context is not enough
Damn, what happened to this general? Previous threads are all gemini and openai twitter slop.Anyway:At least for jap translation gemma 31b completely mogs laguna 118b. The translation is really bad, damn. I hoped for an upgrade. This feels even worse than the older 20b range mistral models. I think the by now years old rocinante drummer finetune is better than this.Gemma even properly followed instruction added stuff to the word glossary at the end like a name and weird world logic. I translated and played 4 games now with gemma31b. q8 cache. q4_s. And still smart enough to give great translations without too many errors. So still very happy with gemma amd I suppose laguna was made for coding not translation. Still mistakes that should not happen. Its just 20 lines about a girl in a lewd scenario and still>ゴクッ…んぐぐ…精液が…胃の中に…becomes>(laguna)gurgling(???)... MY(???) sperm... going down my stomach...>(gemma31b)gulp... nggh... the semen is... in my stomach...Like it doesnt even understand that all the lines to translate are connected. Weird.
guess the model
Has anyone (externally) tested the 12B qat against the full precision model? I've used the qat and it feels significantly worse than q8.
>>109338993Unsurprising. Meme model makers remain meme model makers.
MINIMAX M3 WHEN
>>109338980jeets gonna jeet
>>109338980Charitable interpretation would be that it is a proof of concept before scaling up. Hopefully some of the other labs take notice and use it for their upcoming lineup.
>>109338979yes but the server isn't, and headscale is a pita.just try netbird, the whole thing is foss and it's imo a much better solution.
>>109339010They can live off of cost+ MIC gibs. They don't have to bother with quality.
>>109339022if you don't self-host, you don't know what's on the server anyway. It's cool to have an alternative, I will probably use both
>>109339020>Charitable interpretation would be that it is a proof of concept before scaling up.Andrej once said that the biggest bottleneck in AI development is the idea that small-scale innovations scale linearly. He says at OAI they were constantly experimenting with stuff they read in papers. Things that don't come into play with 3B-sized models begin to appear over 100B which makes it unusable. He also said this is why deepseek have been so successful; they experiment on a big scale because they have the funding and compute to fuck around. It's why nothing ever happens these days because all the compute these big labs do have is already scarce and needs to be used for what generates a return right now; bigger number on retarded benchmarks to keep investors happy. Not innovating. OAI, Anthropic and Kimi are all scrambling for scraps of compute. They can't experiment with the compute they have for that's being used to build the next model with the same proven architecture.
she hides it behind a metal plate
>>109338999>externallywhat?
>>109339080god I want to sniff her ozone
>>109339080>"are you going to x, or are you going to y" assistantslop
>>109339094>ozoneeverytime I see that word I think of himhttps://youtu.be/WNoqXsIp8V8?t=35
>>109339109i dont even get why people bring these patterns up so often its never bothered me
>>109339092Someone outside of google testing their claims that there's no quality loss. In my own testing qat feels like a meme and no better than Q4.
it's been 12h and lagoona is already forgotten
>>109339128>its never bothered meIQ tests mostly test pattern recognition.Think about that.
Huh apparently GPT 5 doesnt have any issue viewing pages of and translating even kereno manga so long as you're outputting it in a json format.But it sucks at making bounding boxes.might kinda pair it as a backup for gemmaI'm not sure. Oh if this trick can work for bypassing the filter when it comes to translating web novels then i am really back in business.
>>109339128It's because you are a dumb anime reaction pic poster
>>109339137i'm still lagooning
>>109339137why was an 8b given any attention in the first place?
>>109339128This girl always makes me think of picrel.
>>109339201who is she, does look similar
>you still can't buy an MCP-controlled male sex toy for gemma to edge you whilst feeding webcam snapshots of your face to her
>>109339146Hope you don't live in one of the five eyes
>>109339214https://osr.wiki/books/osr2/page/overview just make it yourself
>>109339201Look for 泡のお姫様
>>109339237Meant for >>109339207
>>109339223>$100 to buildI know I shouldn't, but damn that's cheap
>>109339243>>109339237looks cute will read it later
>>109339247its mostly the servo costs. i built mine for really cheap, had the micro controller laying around, filament is dirt cheap. eye bolt things and servos were really all i needed to buy.
>>109339200>why was an 8b given any attention in the first place?64gb ram hardware poor hopiumalas
>>109338980>unless specifically tuned for one taskThat's how it is, they release a 'base' model and you finetune it on a specific task
>>109338999qat is equivalent to q4, we all know that
>>109339146You must be as retarded as her. Enjoy getting v&
>>109338993> jap Fuck off.
>>109339146How to get banned and government up your ass if you live in moralfag country 101
>>109339345???
>>109339330>>109339355Why exaggerate?
>>109338993Man, Gemma's so good for this stuff.
>>109338993I am getting weird results with role-playing at Q3 at least, it seems smart and creative, but just seems off and does weird things that aren't really in character or are quite random bordering on nonsensical. Maybe if you fine-tuned it specifically on role-playing though, you could have something decent, needs a bit of a push in the right direction, I see something in this model still even if it's off
I don't know how to feel about ECI. Their results feel questionable (Fable worse than 5.5 Pro?) but maybe it has some value.For example K3's single forward pass time horizon puts it between Opus 4 and Opus 4.5. This could indicate K3 is badly pretrained. AAII does not capture this. I wish we had better capability indicators than AAII and ECI.
>all open weights models above 300b parameters banned, unless explicit permission is grantedlocalbros we won
>>109339441metaAI be like
>llama.cpp>we use low level programming language>but we maximize compatibility across platformsisn't it a bit awkward?
>>109339502It was originally intended to be a quick project and it's outgrown its scope years ago.
>>109339381>he doesn't know
>>109339502In principle, as long as you have no external dependencies open-source C/C++ is pretty portable.The only reason projects like vLLM aren't as portable is because they are using dependencies that aren't.And it only takes a few dependencies to have a slightly narrower scope to make the scope of the whole project narrow.Problems with C/C++ arise when you are interacting with the operating system/file system but those are manageable.
>>109339527Not when you work with various GPU kernels
>>109339502>>109339527>>109339533would have been nicer if they went with rust desu.
Any good distills for deepseekV4? No way I'm fitting the full 280 on my machine
>>109339545There's a rust version just waiting to happen, simply feed the code to your local agent
>>109339546"distills" are snakeoilSee if you can fit anything from https://huggingface.co/antirez/deepseek-v4-gguf
>>109339533GPU code always has awful portability unless you invest a lot of effort.If you write the kernel in a high-level language you are just moving the work somewhere else, DirectX/Vulkan have a lot of driver patches for specific games.And the llama.cpp Vulkan backend only works as well as it does for NVIDIA because of an NVIDIA engineer working on the project and adding custom extensions to the Vulkan specification.Or if you want to use an external dependency like CUTLASS or Triton you are limiting yourself to what those dependencies support.
>>109339559useless because what makes llama.cpp great isn't its atrocious codebase but its popularity meaning a lot of effort go into it supporting new models / features.