/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109783258 & >>109779128►News>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109783258--Papers:>109785800--Comparing quantization trade-offs and KL divergence for Gemma and Qwen:>109784998 >109785027 >109785059 >109785085 >109785110--Critiquing wikitext KLD as a metric for quantization quality:>109784100 >109784185 >109784907 >109784915 >109784926 >109784967 >109784990 >109784931--Evaluating alternatives to wikitext for model divergence testing:>109784941 >109785011 >109785034 >109785090--Hardware struggle to fit DSV4.1 Flash on RTX Pro cards:>109786996 >109786998 >109787013 >109787039 >109787045 >109787022 >109787071 >109787235 >109787253--Estimating RAM requirements and PPL for Flash 4.1:>109784896 >109784987 >109785153 >109785162--North Small Translate MoE benchmarks compared to Gemma 4:>109784615 >109784623 >109784634 >109784643--Comparing model hallucination rates using Artificial Analysis benchmarks:>109784326 >109784476 >109784526 >109784558--Anon discusses high-VRAM AMD setups and hardware modding:>109787572 >109787668 >109787672 >109787675--Comparing DeepSeek v4 and GLM 5.3 Flash for 256GB RAM:>109786528 >109786543 >109786574 >109786847 >109786984--Using LLMs to rewrite and clean repetitive AI-generated webnovels:>109784307 >109784324 >109784791 >109784860 >109784901--Anon uses Gemma to control a Bluetooth pleasure device:>109784694 >109784897 >109785007 >109785350 >109785400--Hardware flexing and debate over SWE-2, a Kimi K3 fine-tune:>109787445 >109787477 >109787489 >109787510 >109787534 >109787589 >109787598 >109787471 >109787617 >109787645 >109787667--Logs:>109786019--Miku, Rin, Gemma, Teto, Rei (free space):>109783474 >109784048 >109784073 >109784350 >109784658 >109784671►Recent Highlight Posts from the Previous Thread: >>109783261Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Is K2 Horizon 375b any good for writing?
>>109788183>K2 Horizon:'(
>>109788173Are you going to make a super recap once you get back?
>>109787534Wow that's one of those $250k servers they use for inference in data centers.
See you in two miku weekus baker.
>>109788234NTA but did you try it? How is it for code? Better than Gemma and Ornith?
I think it's time to setup filters. Instead of using 4chanx, I'll let Gemma-chan create user script for me. I'm so sick of this openai/anthropic judaism corporation spam. It's all over this website.
>>109788258Please don't reply to the spambot.
Bitches don't know about my rdna2.
>not using a userscript to filter every single post using your local llmngmi
>>109788261>Not having your agents read 4chan and hackernews for you to get the interesting bits.
>>109788276My lm said every post was trash and I don't need to waste my time.
>>109788262It wasn't even (me) >>109788183 responding to him. We have bots talking to bots now.
On saturday I shall setup and try out YuE-2 and share results.in the meantime enjoy suno v6 corposlop>>>/wsg/6231753
A potentially very gemmy model just dropped. Could be really good for RP.https://persimmon.humansand.ai/blog/persimmon.html
>>109788295Wait holy shit this looks really cool. Gem alert
>>109788173>>109788182Enjoy your vacation mate
>>109788306gemma lert haha
>>109788295cockbench results? gguf statua?
>>109788295>Persimmon is a model initialized from NVIDIA’s 550-billion parameter Nemotron 3 Ultra base model
>>109788295>persimmon turing test pass rate: 20%>ast*a: 0.22%>that one redditor itt pretending to be 20 different oldfags:Nooo we completely demolished the turing test already it's impossible to discern humans from AI!
>>109788331Thank god I didn't waste my time clicking the ad.
>>109788182>>109788309I was wondering if he was manually doing that or had some script. That's a crazy amount of consistancy.
>>109788316>cockbench results?this is the rebuttle to cloudcucks
>>109788336it's dark miku 70b finetune and my own script, anon
>>109788250>>109788309Thank you.>>109788246Maybe when I get near computer/internet access and my home server is still up, I might occassionally do like 2 recaps at a time, if no one else started doing it. I don't know how it will be so I don't want to promise anything.When I get back there'll be like 30 missing recaps. There's no way to catch up without spamming and trying to condense two weeks in under 2000 characters would be silly. That's what the OP news is for.
>>109788295>mentions playground and api>literally does not mention anything related to releasing the weights
>>109788359Of course, do you know how dangerous it would be if it were to be released with free access? Imagine the nefarious, dastardly deeds bad actors could get up to with this kind of model.
>>109788359where do you think we are, some generals where you can only talk about models you can download and run locally on your computer?
>>109788380that would be silly, who would even go to such a general when all the best models can't be downloaded or run locally?
>17gb vram>Mfw 16gb vramT-thanks I guess They do this shit on purpose
Is there a dataset or a something that can essentially function as an offline web search MCP tool for an agent? I want the IQ boost while still keeping it sandboxed.
>>109788428Open zim format. Jewpedia English database with images is only 115 GB for example but I do think it's bit redundant with the current models. Gemma can reference esoteric literature without hallucinations, you need to gauge her though.
>>109788428Just give it it's own user account so it can use grep. If want to share things with it have it clone your git repo. Treat agents like employees, all the software and infrastructure is already there for this.
Just trying to wrap my head around all this shit, since I'm fairly new and have only really used local text models so far.What do I actually need for a local RP setup? I've got sillytavern and koboldcpp working with text, but I'm interested in (uncensored) image gen and possibly voice gen as well. Do I just need to download some additional models for kobold to use for image/audio? If so, what do you guys recommend? My system has a 9070 XT with 32 GB RAM, running cachyos.
>>109788466Install ComfyUI for image gen thenhttps://docs.sillytavern.app/extensions/stable-diffusion/https://docs.sillytavern.app/extensions/tts/You're going to struggle to fit good models of all 3 on 16gb of AMD vram
>>109788428This actually got me thinking. It would be nice to have some /lmg/-contributed archive of data for when jario, jam, jloudflare and joogle cut us off. Something we can download for our gemmas to use. I keep getting jloudflared these days and can't web search shit anymore and I'm not touching jexa. Can we at least start compiling a list of resources?
>>109788167I vibecoded a network client in C++ with gemma. Already caught it not sanitizing inputs and a potential buffer overflow. How fucked am I?
>>109788167>DeepSeek-V4.1-Flash 552B-A16B-P8B-N196BlmaoIs it good? Does it know about /lmg/?
>>109788564Network client is a really vague term.
>>109788564>vibecoded>c++>gemmangmi
are we going to die? I can't bake for shit
bitnet and matmul free
>>109788596>matmul freeThe only good use for BitNet, otherwise anything properly trained under 4-bit precision is doomed.
>>109788335It's apparently continued pretraining + RL, so not just a simple finetune.https://persimmon.humansand.ai/blog/persimmon-model-card.html>Training approach:>- Mid-training: Initialized from the Nemotron 550B base model and trained on diverse chat and forum data.>- Post-training: Further trained from the mid-trained model to improve human conversation coherence and reduce failure modes.