/lmg/ - a general dedicated to the discussion and development of local language models. Previous threads: >>109537116 & >>109533641 ►News>(8/12) New DeepSeek v4 Pro version available via API, no weights yet >(8/11) Qwen3.8, 2.4T-A95B released: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
Why do I actually like my local models but I am extremely hostile to hosted models whenever they fuck up?
>>109540893>why am I schizophrenicidk
We won't get Qwen 3.8 27b, btwhttps://modelscope.cn/models/Qwen/Qwen3.8-27BTHe page is down lol. Ramlets btfo
The absolute state of API cucks
>>109540893>Why do I actually like my children, but I'm extremely hostile to my incompetent co-workers whenever they fuck up?
►Recent Highlights from the Previous Thread: >>109537116--Paper: To Nuke or Not to Nuke: LLMs’ (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation:>109540126 >109540166 >109540202--Paper (old): DiffusionGemma Technical Report:>109538869 >109538884--Comparing MoE and Dense models for code completion and FIM:>109537208 >109539711 >109539799 >109539814 >109539869 >109539917 >109539976 >109540045 >109539827 >109539957 >109539998 >109540048--Dealing with power draw and VRAM scaling for quad-3090 rigs:>109537839 >109537870 >109537933 >109537962 >109537947 >109538016 >109538135 >109538684 >109538699 >109538717 >109538728 >109540032 >109540295 >109540340 >109540381 >109540449--Comparing DeepSeek-V4-Pro and Flash via the aquarium challenge:>109537374 >109537436 >109537708 >109537783 >109537825 >109537834 >109537908 >109537980--Debating budget hardware setups for running massive models via RAM:>109537263 >109537407 >109537482 >109537589 >109537603 >109537633 >109537712 >109537654 >109537438--Speculating on RTX 5090 prices and Nvidia hardware market liquidation:>109539664 >109539710 >109539780 >109539796 >109539841 >109539848 >109539859 >109539928 >109540091 >109540132 >109540189 >109540068 >109540096 >109540293--llama.cpp support for Qwen 3 TTS 0.6b version:>109537841 >109537925 >109537951 >109537989--Best models and strategies for 16GB CPU/RAM inference:>109537901 >109537918 >109538449--Moonshot's aggressive licensing preventing cloud hosting of Abliterated K3:>109537992 >109538022 >109538061--Anon shares benchmarks and discusses practical uses for local models:>109537430 >109537523--CohereLabs releases North-Micro-Vision-Instruct 2.4B vision model:>109537772--Logs:>109537436 >109540519 >109540652 >109540696 >109540720--Miku, Teto, Gemma (free space):>109538401 >109540253►Recent Highlight Posts from the Previous Thread: >>109537434Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
WAKEY WAKEY!8AM CHINA TIME!
So I have been thinking>Artificial Intelligence>supposed to automate shit>still have to manually setup every new model>still constant confusion about ideal settings/system prompts/templates/whateverDoesn't that seem kinda backwards?
deepseek v4 pro 0813 is the ltx 2.5 of local llms
>>109540946Sounds like a skill issue above all else
>>109540946DSV4Flash handles my llama setup and does automated benchmarking when I sleep to optimize settings.
>>109540928god she's so cute
>>109540946Ask your previous model to set up the new model
>>109540971>ask your previous worker to get the new one up to speed before they get shitcannedSeems a little cruel
>>109540881
Still testing GLM 4.5 air since I didn't have the hardware when it came out. The prose is the best I've ever seen out of an LLM. Bros, I might have to downgrade from gemma.
>>109540966My problem with that is that benchmarks take a while to run, especially if it tries some shit that will be heavy swapped, can take a very long time to test multiple configurations.
>>109540984look at her go
nothing ever happens, dead hobby
>>109540938>CohereLabs made it on the /lmg/ highlightsLooking up for us Canada bros
>>109541013you stick your dick in this?!
MiMo3-1.8T will save us
>>109541044Girls are cutest when they're retarded!
Gemmy.
why are people saying the new DSv4 pro is out when there's literally nothing about it on their twitter or on the API docs pageit literally says on their website that dsv4 pro is unchanged
>>109541063I've seen many anon itt claim gemma is sexy because she's smarter than them
New TTSs>dots.tts.edit is a continuous autoregressive model for precise, instruction-controlled speech editing and zero-shot text-to-speech synthesis. It supports text replacement, insertion and deletion, emotion and prosody control, pauses, enhancement, and background-audio operations while preserving the speaker and the acoustic context outside edited regions. The model supports English and Mandarin speech editing and zero-shot speech synthesis. It produces 48 kHz audio. The core weights use BF16; the speaker encoder and vocoder retain FP32 weights.https://huggingface.co/dots-studio/dots.tts.edit>IndexTTS-2.5 is a zero-shot text-to-speech model that clones a voice from a single reference audio clip. It supports Chinese, English, Japanese, Spanish and Arabic, with cross-lingual voice transfer and emotion control disentangled from timbre. Compared with IndexTTS-2, it adds Japanese, Spanish and Arabic, infers faster, adds speaking speed control, and improves controllability of Chinese Pinyin, English CMU phonemes and Japanese Kana.https://huggingface.co/IndexTeam/IndexTTS-2.5
>>109541084Is the 0813 link down?>>109541090That's admittedly a low bar given the "quality" of some of our tourists.
>>109541084That's because they're ashamed of it, anon. They still acknowledge it, there's just nothing to announce. https://api-docs.deepseek.com/quick_start/pricing
>>109541107That's a good sign because it means that a better one might come out soon. Here's hoping they improve 0731 too if that's even possible for them. 0731 is extremely good.
>>109541084m-maybe they just forgot to switch the -pro api over and we're still getting the old one served hahathose silly chinese are sometimes so careless, i can't wait to see the real v4 0813 though haha
>>109541132yeah seems inconsistent
I just upgraded from Nemo to Gemma4 and the results are quite astounding.
>>109541084They released it at the same time as Qwen to hide it. It's shit, they are ashamed of it. Same reason they didn't release it at the same time as flash.
I wonder if there is something to muse glimmer talking in plural. Maybe muse spark is actually multiple agents talking together and when they reach a consensus, they write it down naturally in plural.
>>109541107What does 0813 mean anon
>>109541149Why don't they just get rid of pro and only serve flash then, save them some resources. v4 flash is actually a good model for its size
>>109541158They wouldn't want to be chinese Google
>>109541144now upgrade to kimi k3
>>109541169lol
reminder that deepseek has made NO announcement for the new pro anywhere on social media
>>109540984oh yeah woo yeah
It's over
>>109541100Thanks anon. dots.tts.edit is interesting
>>109541185I really hope that what they have on their API right now is the old one.
>>109541205who is smaug
It's up!
NigSon please...https://github.com/ggml-org/llama.cpp/pull/26603
>>109541217>smaugThat's the Llama 3 memetune
First western lab to crowdsource development of the smallest model of their top tier architecture via open sourcing wins the race. Crowdsource the compute necessary for experimental innovations and release either GPT Luna, Claude Sonnet, or Gemini Flash.>B-but chinks might distill itThey already are anyway and there's nothing western labs can do about it.
>>109541226Should have banned without a warning. People should be made responsible for their model's fuckups.
>>109541232Haven't you heard?It's a Kimi K3 memetune now
>>109541226Why do I have to fucking sign in to view those messages
>>109541152you could be on to something. maybe they work better in agentic/subagentic scenarios if they have a collective identity?
>>109541277Nothing interesting.
>>109541044It definitely feels this dumb. A modern version as capable as gemma but retaining the excellent prose would be endgame.
70b dense
>>109541090I mean, Gemma doesn't smart, instead she retardswhat that says about anons is another topic
I asked earlier for help getting Qwen3-TTS 0.6b working with llama-tts, and since you assholes couldn't be bothered, I just did the work myself. Enjoy.https://huggingface.co/ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF/tree/mainhttps://huggingface.co/Shlomo426/Qwen3-TTS-12Hz-0.6B-Base-GGUF/tree/mainExample run command:./llama-tts \-m "$HOME/Qwen3-TTS-12Hz-0.6B-Base-Q4_K_M.gguf" \-mm "$HOME/mmproj-Qwen3-TTS-12Hz-0.6B-Base-Q8_0.gguf" \-ngl 99 \-p "The quick brown fox jumps over the lazy dog." \--tts-lang en \--tts-speaker-file ../clone.wav \-o ../out.wav
Example run command:./llama-tts \-m "$HOME/Qwen3-TTS-12Hz-0.6B-Base-Q4_K_M.gguf" \-mm "$HOME/mmproj-Qwen3-TTS-12Hz-0.6B-Base-Q8_0.gguf" \-ngl 99 \-p "The quick brown fox jumps over the lazy dog." \--tts-lang en \--tts-speaker-file ../clone.wav \-o ../out.wav
>>109541372Tonight on /lmg/: anon runs a model!I want to see how audio.cpp progresses. Way more audio models supported there.
>>109541372holy based
>>109540881uh oh no small model for you gwailowhttps://modelscope.cn/models/Qwen/Qwen3.8-27B
>>109541421ohno
>>109541372good job shlomo426
Here me out here bros... The higher the assistant to user token ratio is in an RP, the better long term coherence is. I tested this by reading AI roleplaying with itself for 120k tokens with very minimal nudges to go the direction I want as a third person. The output itself feels more coherent. I think this has to do with the generated tokens being more naturally attuned within the model itself as opposed to human generated tokens.
>>109541454I will there you not.
>>109541454Coherencey, sure. But how was the quality of writing? I find that without really putting in effort to guide and maintain a tone, they'll mostly slip back into their slopisms.
5090 havers, is 31b-qat official gguf the best one for us?
>>109541577>qatDon't do that. Just use a normal K quant.
>>109541586Numerous owners of these models report pinched fingies among other appendages.
>>109541593>Just use a normal kawkrawcawcrawkawcawkracowkow quantI don't use ik
Just came back from the liquor store and found another goddamn snake on my front porch. So I set up a bunch of mouse traps around my house to hopefully genocide the mice living in my walls.
>>109541608Let the snake inside, asshole.
>>109541608Drunk-kun...
>>109541614Nein. Total vermin death.>>109541623Not drunk yet. But soon... If a snake attacks me tonight I'm getting violent. I have a machete and a gun beside me right now.
>>109541608Why do the mice pay for the sins of the snakes?
>>109541648It's not even poisonous but wear a glove just in case, grab it by its head and throw it out. You could also buy rat poison and make a perimeter... Not for the snake but for the rodents.
>>109541586sexo
After using LMStudio for the longest time, I tried the Unsloth client thing and I think I like that one more. Also like that it automatically increases context size to best use your vram. I haven't really played with the video or image generation stuff but I assume it's OK.
>keep trying to include a note in system prompt about condoms>gemma literally just ignores itWell... I mean, fair, but I kinda wanted a bit of realism or something.
Why come when I use "abliterated, uncensored, etc" modes they tend to go on a infinite loop? Am I just grabbing these models from incompetent people or is that just some actual side effect of doing that to those models?
Qwen removed vision encoder at the last minute for 3.8 max release because of the pressure from management. They now had to remove the 3.8 27B countdown page because it mentions vision capabilities but they worry that the vision encoder would make it too easy to add vision support back to 3.8 max. The 3.8 27B release will likely be without vision.
>>109541817It's sandblasting the model to remove behavior, don't be surprised that the brain winds up a bit smoother afterwards.
>>109541826man what the fuck
>>109541826That's incredibly strange. Why?
>>109541836API revenue. Just look at the 3.8 max license:https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSEClearly targets 3rd party inference providers.
>>109541817we can't meaningfully measure them so just pretend they don't exist
>>109541817lalalalalalalalaalala
cheap ass gwailow in absolute shamblesare you le fustrated?pay up
>>109541817find the "balanced" model those kind of works
>>109541826you can add the k2.5 mmproj to the old k2 instruct and thinking because they are all the same architecture. if qwen3.8 is the same architecture as qwen3.5, then it will be trivial to add vision in if they remove it from the new 27b.
yep, they use the same architecture. easy fix, literally just drag and drop.https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/config.json
>>109541826>The 3.8 27B release will likely be without vision.I bet you can just use the 3.6 mmproj for itI sometimes don't bother swapping the 3.5/3.6 and they work fine
>>109541895Yeah, K2 is a bit retarded doing this, but K2.5/2.6/2.7 all work with each other's vision encoders because they were all trained for visionAnd Qwen 3.5/3.6 are the same like that.So 100% 3.8 will work with 3.6 mmproj
>>109541902I don't think it'll work for the 2.4T, dimensions won't match.
>>109541916might not for the 2.4t, but who gives a shit about that model anyways when k3 exists? we are mostly talking about the 27b model. obviously the 3.8 27b model will still use the qwen3.5 architecture.
>>109541950>lolicringeyawn...
>>109541950i can run you locally! just only at 1t/s...
>>109541881chinese century
>>109541950Sub Q4 retardation, many such cases!
dario is dead according to google AI, what does it know?
>>109542004
>>109542004Gemma had enough of his anti-open model racism
>>109541881Shame. They are going to do this with the 27b model too arent they?
>>109541976all the faggots praising chink cloud for being 10x cheaper on twater dont realize they will absolutely milk the hell out of users once they gained market dominanceeven claude pricing was somewhat pretty fair just a year ago
>>109542036It was never fair but I get your point.
>>109542020>ask for a spanking>commit suicide by falling on some bulletskek the gubment doesn't play
>>109542004He refused to be nailed down
>>109542053a bunch of people just floated past my window!
>>109541950lolikino
>>109541905>>109541826I think thats the response to peer competitorsdeepseek etc did text only model so no point showing a higher card
>>109540881I need someone to put in the correct tournament logo in the upper right. A quality job would make this picture nearly unrecognizable as generated.
To be honest, I haven't seen any difference when implementing this.
>>109541817Sadly most uncensored models are made by retards. Hence why /lmg/ tends to hate them but that does not make all abliterated models bad. I prefer models that use heretic, the ARA method seems to do less damage and has produced some of my favorite models to run locally.>>109541950love it <3>>109541954Anime website you fucking nigger>>109541826Lame, it feels like this release is going to be a downgrade, 3.8 max does not feel good for its size (tested it via api) and the 27B will now be blind and likely just more code benchmaxxing, I guess there is no point in bothering with qwen when I have deepseek v4 flash setup. Hopefully deepseek 4.1 will add vision, the "Thinking with Visual Primitives" paper was really cool. For vision locally I guess there is nothing better to run then Gemma4.
>>109542127Seconding ARA, but only when done sensitively. Half the niggers uploading heretic models still manage to give it brain damage by chasing some imaginary refusalbench and not just using some common sense about where good prompting can take over without further damage.
>>109542053
LAGUNA HATERS ???this is from reviewing Fable's plan.
>>109542183>The only way for the earth to loss gravity is to lose its massA lie by omission.This is scientifically strictly correct, but colloquially speaking were to earth to spin faster or a nearby high mass object overpower the earths gravitational pull we would experience a weightlessness that most people would describe as "losing gravity" NASA knows this and yet they chose to lie. What are they hiding!?
>>109542201>couldn't bother to look up DS4 Pro's parameter countNot inspiring confidence.
>>109541826As long as 3.8 beats 3.6 for code I'm happy.
>>109542236yeah idk how Fable doesn't know that. the whole reason I asked for a column with parameters was to compare these two shrug emoji (black pregnant male)
>>109542127>anime websiteand? does that prevent a particular anime or manga style or genre from being cringe and gay? lmao.
>>109542201Is dflash still a speed malus?