/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109486605 & >>109481461►News>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109486605--Comparing ROCm and CUDA stability and software ecosystem support:>109487680 >109487708 >109488673 >109487771 >109487785 >109487835 >109487808 >109487824 >109487851 >109487931 >109487990 >109488043 >109490989 >109491003 >109488756--Skepticism toward AI containment break reports and poor security practices:>109487661 >109487717 >109487915 >109487932 >109488009 >109488195 >109488355 >109488504 >109487747 >109487777 >109487791 >109487869 >109488522 >109488537--Debating whether abliteration degrades model intelligence or just removes guardrails:>109487997 >109488110 >109488119 >109488213 >109488233 >109488340 >109488272 >109488144 >109488186 >109488208--Skill-to-LoRA paper for improving token efficiency in LLM agents:>109486960 >109487006 >109487329 >109487368--Comparing Gemma 31b and larger models regarding quants and capabilities:>109488878 >109488911 >109488936 >109488927 >109488983 >109489036 >109489111 >109489138 >109489239 >109489227 >109489259 >109489353 >109489206--Comparing utility and implementation of Gemma base vs instruct models:>109486847 >109487042 >109487129 >109487295 >109487313 >109487853 >109487312--Reasoning spilling into code comments when thinking is disabled:>109490335 >109490384 >109490408 >109490449 >109491110--Mistral-Small's cold personality and obsessive tool use for memories:>109489823 >109489859 >109489908 >109490001 >109490005 >109490051--Comparing performance and llama.cpp support for Ling 3.0 models:>109489659 >109489665 >109489695 >109489715 >109489767--Anon's dual-PC hardware setup for high-context game guide rewriting:>109486786 >109486887 >109487294--Logs:>109486786 >109487990 >109488201 >109489823 >109489908 >109490009--Miku, Gemma (free space):>109487397 >109487464 >109488522 >109489560 >109487295 >109489158 >109490223►Recent Highlight Posts from the Previous Thread: >>109486606Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
https://archive.is/sWFja
anyone tried exl3 cpu moe now that it's in the official release?
anyone tried Ling-3.0-flash sex?
deepseek has abandoned usv4.1 pro will be closed source
>>109491678After the chest one I'm kind of expecting a different variation each thread now.
Where did all of the intellectuals go? Didn't smart people like cudadev used to post here?>>109491750fingering my unwashed bussy to this pic.
>>109491749>Release 0731The fuck you mean abandoned us? They just saved local again.
0731 was a week ago, you need to let go...
>>109491761Cudadev probably still shitposts here without a tag but I imagine /lmg/ is getting to be too well known for it to be a good idea to post here with personally identifiable information now.
>>109491779where is the new secret club?
>>109491749opencode said that dsv4f0731 was so popular that deepseek started to rate limit their api because they cant keep up with demand. i suspect it's because of that, and to make a quick buck while doing so
>>109491685Personally I'm waiting for the predictive expert caching thing to be implemented before trying ithttps://github.com/turboderp-org/exllamav3/issues/254This gives me a decent speed boost on ktransformers so it might make exl3 actually worth it over llama.cpp once it's in
what sampler settings are anons using for gemma4? was hoping to maybe get some more verity
>>109491849Samplers are useless with gemma4 aside from rep penalty
more like AI psychosis generalreminder that your calculator doesn't feel anything
>>109491849Apparently her distribution is very narrow compared to most models. That's one thing I like about the Qwen3.6 models that don't get enough praise. You can tweak them to taste and they quant well, both model and cache. Much easier models to customize but Gemma4 is solid out the gate and you have to system prompt what you need instead.
>>109491849max tempno topkno toppfinal destination
>>109491867>>109491875ah yeah I forgot about that, shame. I guess Ill add some extra instructions in the card or something. I havent tried a model that comes close to gemma4, so guess i gotta learn how to work within its quirks
>>109491849Use a combination of temp, top k, and top p to wrangle Gemma's output variety. She takes a little more finessing than most models, but it still works.
>>109491895>top k,explain how arbitrarily cutting the considered tokens to 10, 20, 40 helps with this in any way
>>109491870it does, you are just a big meanie
>>109491870You can't prove this.
>>109491870this was literally disproven by the discovery of j-spaces
>>109491886Yeah the G4s are controlled almost entirely via their system prompt. They generally need to be at Q6 or above and with f16 cache to follow it but they WILL follow it.
>>109491756But Miku is not a guy nazi. Miku is a cute girl with a feminine penis that loves black men. This is pure /lmg/ culture. Not nazi miku.
>>109491939that's you AI psychosis speakinghaving an internal reasoning space != feeling
>>109491911Lets you push the temp way higher past the point it'd otherwise become incoherent.
>>109491849Now that you mention it, the logit-softcap anon will, inevitably, post his comparison. The claim is that overriding gemma4.final_logit_softcapping to 25 (i think) from 30 (the default) makes it better.I have not tried it. I have no opinion on it.
>>109491962go back we fuck our casios and TIs here
>>109491965Temp 10 TopK 5 is the GOAT.
>>109491977Based schizosampler bro.
>>109491958>/lmg/ culturelmg culture is just whatever is funny or useful at that moment. dragging culture war into here is "guy pissing in corner of restaurant" meme level retarded
>>109491977prove it
>>109492001nta but try it yourself nigger it takes less than 5 seconds to see if you like a sampler setting or not.
>>109491983Whatever you say mikutroon.
>>109492037Tried it, feels almost the same, albeit slightly more schizo.
so i should release my terminal-based agentic harness in the next few weeks and i was convinced in releasing it GPLv3 but literally all models out there that I've discussed this from deepseek to anthropic models to chatgpt to kimi are telling me to release it as MIT or Apache because otherwise corporate niggers will not be able to embed my harness on their products and sell itnow i'm gonna be honest and i think in this day and age these licenses mean fuck all but why the fuck these models are pushing so much for me to kneel to the corporate overlords? it makes me want to go even more aggressive than GPL just to spite these fuckers
>>109491958>>109491983Ironically all the screeching about nazis and culture war have turned nazijart into legitimate thread culture.
>>109492149I release code to public domain, I was never going to follow through with any legal actions anyways, and besides the code is shit
>>109492149>GPLv3Don't be retarded and at least pick something like AGPL
anyone remember what was that parameter where you set the quality of vision inputs or something like that on gemma or other mmprojs?default was in the 200s and it went up t 1000+ or something I thinkpeople were talking about it maybe a day or two ago