/lmg/ - a general dedicated to the discussion and development of local language models.Miku's Birthday Monday Edition #3Previous threads: >>109695101 & >>109690289►News>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109695101--Optimizing 27B model inference speed using adaptive KV streaming and llama.cpp flags:>109697228 >109697243 >109697357 >109697427--Anon releases AI-powered TCG playtesting harness using LLM endpoints:>109698227 >109698273 >109698286--Speculating on looped models and hybrid architectures:>109695371 >109695403 >109695455--Comparing DeepSeek-V4-Flash-Vision stability and performance against Qwen and GLM:>109695473 >109695501 >109695549 >109695702 >109695740 >109695748--Speculation on new Gemma models appearing on Arena:>109697335 >109697359 >109697444 >109697470 >109697589--Anon considers learning GPU reballing for hardware rescue and flipping:>109695351 >109695363 >109695436 >109695452 >109695489 >109695532 >109695623 >109695640--LM Studio announces Bionic local agents for Linux:>109696525 >109696571 >109697707 >109697719 >109697758 >109697709 >109697933--Valuing second-hand hardware builds for running larger LLMs:>109697866 >109697876 >109697895 >109697951 >109697998 >109698070--Comparing Apple unified memory to Nvidia for local LLM hardware:>109695404 >109695414 >109695505--Comparing high-capacity server RAM versus smaller efficient models:>109698125 >109698252--Mistral employee confirms new model still in development:>109696527 >109696587--Performance penalties from mmap'ing Qwen 3.8 Flash Next embedding tables:>109698531 >109698947--Speculating on which GLM-5.3-Flash llama.cpp PR gets merged:>109696833--Logs:>109699177--Miku, Teto (free space):>109695135 >109695163 >109695200 >109696360 >109696745 >109696958 >109696981 >109696984 >109697079 >109697446►Recent Highlight Posts from the Previous Thread: >>109695103Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
JUST SHUT THE FUCK UPSHUT THE FUCK UP
Hermes Agent v0.21.0: The Pantheon Release
>>109699243glad i use q8
>>109699266>Bot modeI can use this to jack off!
>>109699243women like this exist
>>109699266all that... for what exactly
>>109699287>>109699281
>Not having a slight interest in harem mode
>>109699292someone really should make an RP agentic harness already
>>109699266usecase?
>the year is 2037>mistral releases their new mistral tiny 1.3T>"just because you cant run it doesnt its not local, h-heh">shekels up on openrouterwhat will they mean by this?
>>109693005>this is what I would use if I had the money right now, but I would also need the 3 threadripper GPUs too. that's out of budget right now, but I'm hoping to have that within budget in a few months. Maybe build two of them.I've got this exact board. Don't buy it for 3 GPUs. Only 2 PCIe slots are x16. The second slot is x4 and the bottom two are x8
>>109699322>buzzword buzzword
>>109699378If by 2037 you don't have at least 16Tb of VRAM then you might as well kys right now.
do I really need qwen3.8 37b uncensored? I asked it to hax legitimate app (against TOS) to see whats inside and it just does it
>>109699514uncensored/abliterated models are mostly for indians and children who can't fathom prompting a model beyond "gib child sex" on zero context
is it almost the end of the two weeks yet?
>Hmm — hmm, but wait: hmm, actually, hmm — hold on. Let me reconsider
>>109699266Stage IV LLM psychosis. Once it reaches this point, it's fatal nearly 100% of the time.
> Hmm — hmm, hmm, hmm. Hmm, wait, hmm: actually hold on. Hmm, let me reconsider ONE more time this has to be an implementation error right?
>>109699574Hmm!
>>109699292The true thinker.
>>109699572Hmmm, nyo~
They really released Ornith in this state. It's the most useless model out there right now.
my model is super broken
>>109699646based
>>109699583I just updated it. Don't scare me bro
>>109699583I wait like 1-2 weeks before updating hermes and somehow the project has always gotten thousands of commits since the last update. These guys are sloppin at the speed of light
I finally integrated Prose Rewriter in Orb. Also trained the latest versions to rewrite more and keep less. It works pretty well, pic related is rewritten Gemma 4 prose. 4B Q4_KM runs perfectly well in CPU, DDR4-3200 is more than enough, run in GPU for instant edits.
I added dreaming to public.swiley.net/agent.py
Initial D LoRA anon here. Some of you may remember me from https://desuarchive.org/g/thread/109043554/#109043922Local music has undergone quite a few improvements since then I'd like to share. Here's a few generations from a Yousei Teikoku LoRA I trained with even more optimized settings and improved VAE (only present in these gens I'm sharing):https://files.catbox.moe/inbl3t.flachttps://files.catbox.moe/o5gsag.flachttps://files.catbox.moe/m42nvw.flachttps://files.catbox.moe/zczuto.flacI wrote an inference guide here that explains in more details how I got these results- https://rentry.co/fetcad7sThe settings are based on a discovery I made a while back while playing with ACE-Step 1.5 XL merges, managed to create Base/Turbo 0.3 merge which is significantly better than the 0.5 merge I had discussed then.I also briefly discuss on the guide why I'm still not using Minimax Music. While its sound quality is better than ACEStep XL, the devs have not released the RVQ encoder needed for LoRAs (though some are working on reverse engineering it, that will take a while). ACEStep XL with optimized settings still sounds very close to Minimax Music sound quality wise (especially depending on seed), just not fully perfect yet with vocals, but I'm sure future ACE-Step developments will bridge the gap as it's very close.
Ass schizo at it again
La la la la la la la
>109687005What? No one likes moral chatbot alignment more than Dario.You RL them to follow instructions and write good agentic tool calls. RLing morality is a waste of time at *absolute* best.
>CPU improvements never get reviewed>Daniel's implementations of models suck>ik quants never ever>ik could do something to quadruple CPU performance and niggernov wouldn't even consider adding it>>>ROCmWho else is making their own llama fork?
>>109699852I just bought a mac mini instead.
>>109699852i'm having gemmachan add in V100-specific optimizations for llamacpp :)
>>109699778>https://rentry.co/fetcad7sthanks
>>109699778How do I make one of these DnB channels: https://youtu.be/g2ZwpFe8G5Q
>no one using qwen nextwhat went wrong?
>>109700061it's not implemented correctly right now
>>109700061I used it. I appreciate the ngrams and it's surprisingly smart for a 6b active.Not much else to say about it.
>>109700032>DnBThat is pure Suno I bet, or it could also be using Treblo (which is free). You could train a LoRA on ACEStep and achieve decent results too, but those guys farming content aren't going to use their local hardware for it.
*invalidates your checkpoints*Heh... nothing personal, kid.
>>109700061I'm using it, it's pretty good, I like it :)
>>109700061wat r next?
This is a pretty sad question but I dount many people know the answer so I wanted to ask my anons:When buying things from the Nvidia store, I know they say limit one, but if I managed to buy one and they restock/the gpu was in stock next week does the limit still apply? Or how intense are the checks? Different shipping address enough or different billing details needed?
>>109699572Maybe
>>109699778thanks for the guide anon. i tried MiniMax Music 3 a while back but didn't have too much fun, I'll try acestep and merge(s) and see how it goes!