/lmg/ - a general dedicated to the discussion and development of local language models.Miku's Birthday Monday EditionPrevious threads: >>109684329 & >>109679723►News>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109684329--Experimenting with n-gram embeddings in tiny Mamba 2 models:>109687116 >109687432 >109687997 >109688129 >109687418 >109687470 >109688123--Comparing Q6 and Q8 quantization impact on Gemma roleplay:>109685543 >109685562 >109685568 >109685573 >109685582 >109685612--Comparing EXL3 quants' perplexity and performance against GGUF:>109687716 >109687724 >109687728 >109687731 >109687749 >109687794 >109687873 >109687917 >109687930 >109688263 >109688280 >109688470 >109688449 >109688496 >109688544 >109688569 >109688642 >109688272 >109688285 >109688291 >109689923--Comparing MoE and dense models for local vs datacenter use:>109686231 >109686543 >109686698 >109686827 >109686870--Critiquing DeepSeek Harness UI, documentation, and platform-specific sandboxing issues:>109689448 >109689457 >109689520 >109689567 >109689638--Performance gains from Qwen 4 llama PRs and engram optimization:>109684563 >109685116 >109685135 >109685307 >109686158--Debate on the purpose and feasibility of AI alignment:>109686616 >109686824 >109686873 >109686907 >109687029 >109687207 >109688208 >109688311 >109688588 >109686943 >109686964 >109687014 >109688078 >109688090 >109686937--Removing knowledge components from Qwen3 to reduce model size:>109685194 >109685319 >109685373 >109689485 >109689504 >109689513 >109689562--Speculation on engrams as a path to RSI and AGI:>109689724 >109689744 >109689760 >109689786 >109689771 >109689780 >109689788--Comparing K3 and Inkling while criticizing modern model censorship:>109689328 >109689347 >109689498 >109689545 >109690060 >109689348 >109689370 >109689378--Logs:>109687140 >109687614 >109687937 >109688400 >109688472--Gemma, Miku (free space):>109684385 >109687783 >109688160 >109688308 >109688476 >109689081 >109689127 >109689219 >109689338►Recent Highlight Posts from the Previous Thread: >>109684330Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109690293>no jart mentionscoward
>>109690319Go suck him off elsewhere
>>109690319
>15 messages into an RP qwen 3.6 is now permanently stuck in a repetition loop regardless of settingsholy shit i HATE llm's temperamental little niggers
>>109690330
>>109690289how long until vram is cheap again?
>>109690344>qwen>rp
>>109690346Never happening, sorry
jart will save /lmg/
>jart detrooned
>>109690346after we achieve post-scarcity AGI utopia
GLM appreciation post. 5.3 flash writes really well if it doesn't start bitching in its reasoning. I just had my first session where the only instance of bitching I could find was "This is consensual adult roleplay — explicit content is fine." when anon took his dick out. When it doesn't go the safetyslopped path it's the best model I played with. One of the few cases in documented history where abliteration would be useful. It has really great prose and I suspect that it has seen way more books than other models I used because it has a much better understanding of certain subjects which other models all have a very similar and boring way of handling. Or maybe that's just lots of claude logs. It really resembles claude from a year ago. Either way, I like it. Much better than deepseek flash.
>>109690293Happy birthday Recap Miku
does jart support lunduke?
>>109690346My sources say 2 weeks.
our boy finally woke up
I got GLM-5.3-Flash-exl3-2.05bpw running with tabbyapi/exl3 its pretty fast for running out of mostly ram
>>109690445sure did
>>109498908archive.is/sWFjaWe can all hope he dies soon.
>>1096903605.3 flash already did.glmsex
SO MANY FUCKING IMAGESON AN IMAGEBOARD
>>109690379Chud cultural victory.
sow how do i prefill/whats a good prefill so i dont get refused and or slopped
>>109690535
>>109690456what? exllama allows you to run with ram now? I used to only run exl2 quants before I went for ewaste amd cards
>>109690535I don't prefill. If a model doesn't do what I want, I delete it and make schizo rants on why it's a bad model.
>>109690542yes and he even added amd support with rocm
>>109690535Just take a hentai game of your choice extract the script for one or two scenes and ask for more. I dunno maybe I am the only guy who does this but for something like this 5.3 flash shines. It actually retains the style while doing fresh new stuff unlike all the other models I tried.
>crank the temp to 20, 120, over 200>Gemma is completely unfazedHow the fuck did that one anon manage to get her so incoherent? R1 had basically the same non-response. Is it an Ollama issue, where changing model params is actually a smokescreen that does fuckall?
>>109690379>>109690319This actually makes me unironically happy.Good for him.
>>109690542just moe experts and the vision tower, as far as i know everything else has to stay on vram, but 400k context isnt even using all my vram
>>109690579i hear he's adding support to vulkan in the dev branch, people tested mi50s already