/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109522373 & >>109517796►News>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B>(08/08) US DoE Launches Genesis Open Models Initiative: https://genesisopenmodels.anl.gov/>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109522373--Paper: Stealing Reasoning Traces from Proprietary LLM APIs:>109527308 >109527346 >109527354 >109527470 >109527758 >109527446 >109527506 >109527577 >109527930--SSDmaxxing for hosting large models via PCIe NVMe arrays:>109522742 >109522753 >109522781 >109522773 >109522849 >109522805 >109522818 >109524084 >109523899 >109523924 >109525407 >109525425 >109525433 >109525598 >109526008 >109526026 >109526062 >109526098--Debating ideal model size and techniques for humanizing AI prose:>109525049 >109525081 >109525100 >109525574 >109525580 >109525646 >109525667 >109525624--Reasoning for human-readable model output and use of "we":>109527603 >109527610 >109527645 >109527667 >109527693 >109527717 >109527662--Data quality vs quantity and the TwIL-LM3 benchmarks:>109526447 >109526501 >109526536 >109526507 >109526469--Solving a math problem through obsessive AI prompting and verification:>109524869 >109526003 >109526464--Positive reinforcement and encouragement improving model output quality:>109523184 >109523210 >109523314 >109523339 >109523548 >109523588 >109523624 >109524164--Meta's Muse Spark safety report:>109526330--Adding a budget 3060 for TTS offloading and VRAM management:>109525523 >109525529 >109525546 >109525571 >109525593 >109525640 >109525717 >109525740--Nemotron-3.5 benchmarks showing it underperforms compared to Qwen and Gemma:>109526367 >109526396 >109526578--Claude's new AI-generated content watermarking and EU compliance:>109524310 >109525858 >109525894 >109525907 >109526381 >109526643 >109525992 >109527716 >109527759 >109527788--Logs:>109523214 >109524207 >109524333 >109524856 >109525440 >109525682 >109526198 >109527446 >109527694 >109527759 >109528078--Miku, Minnie, Teto (free space):>109522410 >109525628 >109526824 >109526845 >109527716 >109528058►Recent Highlight Posts from the Previous Thread: >>109522377Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
v4.1 pro is imminent
What is the best coding model for an A6000 and 256GB of system RAM and how to set it up? I tried vLLM and llama with aider and it didn't work great compared to the paid tools with open weights.
>>109528537they are probably waiting for qwen release first
Tetolove
>>109528547I also have 8TB of NVME
penis enlargement pills
>>109528547me too guys but I have like 10x these.
Programming music? Tired of "progressive house" youtube slop.
>>109528620https://www.youtube.com/watch?v=CbQjHb8iaMc
>>109526173Pretty good, but needs some orgasmo-meter for the girl as well.
>>109528644Reminds me of les claypool a bit.
>>109528620https://www.youtube.com/watch?v=wQgeBsv1sns&list=OLAK5uy_kqZER1aLO6paq_aKRuCiLPEbiWbJOflqg
>>109528547dsv4 flash 0731
>>109528665Great taste
>>109528691Primus sucks.
>>109528510Daily reminder Teto is trash and his fanbase are trash (mostly westernfaggots)
>>109528547new dsv4 flashif you want something that fits entirely in VRAM, qwen3.6 27B or gemma4 31Baider is a bit outdated (I assume, I haven't heard of anyone using it in a long time) I would recommend pi / omp / opencode instead
>glimmer for orchestration >qwen for coding>gemma for user interaction this is 90ba30b sota at home
>>109528620DnB or dub
>>109528620No music. Zero. Nada.Focus all your energy on one thing at a time, or you'll have to work extra time later taking care of sloppier code you could just write it once and well in the first place, even if you vivecode. Listen to music while you relax. Use the pomodoro technique if you want to listen to music that much.
>>109528761IDM
She's so supportive.
>>109528773vibecode*See? I didn't cared enough.
>>109528744>>glimmer for orchestrationJust because it's new, doesn't mean you have to try so hard to shoehorn a use for it. It seems to have no redeeming qualities whatsoever compared to Qwen or Gemma and it's a year late to be relevant for being better than gpt-oss.
>>109528733Claude told me to reformat my 4x2TB NVME array into RAID0 (was RAID10 for postgres) and run GLM-5.2 IQ3 across the GPU, RAM and NVME tiers. How retarded is this idea?
>>109528620https://www.youtube.com/watch?v=CGUZD28NnUcVery topical to lmg if I do say so myself. The first two are my favorite tracks.
>>109528800Claude sabotages you if you use it to gain advances in AI
>>109528792it sticks to the system prompt requirements better than gemma and performs the best in my custom agentic setupglimmer vision is also better than both qwen and gemma
>>109528800Didn't you guys discuss this idea a week ago?I didn't catch the tail end, but I think most of you didn't believe it would work out well.
>>109528809I would trust a 1b Gemma over Claude, for local advice
>>109528820Is it censored?
>>109528809>>109528837This, but unironically/
>>109528838can be fully uncensored with just system prompt telling what new “policies” it must obey
>>109528800>>109525407nobody's saying you can't do it, it's certainly doable.now, about the speed
>>109528801>trans mixlamo
>>109528869Rent free.
>>109528800it's pretty retardedyou really don't want to be involving nvme at inference time, you would be much better off running a slightly lower quant and keeping it to just vram+ram. honestly I am surprised that claude would suggest this, it's frankly an insane worst of both worlds proposition to still run a cope quant and also still spill over to nvme.you could run the latest dsv4 flash at full precision between vram+ram with room to spare and it won't be that far off from q3 glm
>>109528620https://www.youtube.com/watch?v=oJK5sdQOMAc
>>109528889>you could run the latest dsv4 trash at full precisionbut why would you want to?
>>109528889Eh, for rp, fp8 dipsy doesn't even beat q4 glm 4.6 for me.
>>109528916sure but he's using it for code
praying for a qwen 3.8 35b a3b
>>109528889Wrong.
>>109528929sorry, I'm a llm from 2022 and my context window is 2048 tokens.
>>109528916Why did they abandon the air models :( our only chance for 100B kino...
>>109528943nyo~
>>109528800You should pin threads to cores and build inference engine on top of SPDK, anything else is a waste of i/o
>>109528711
>>109528962hmmm...
>>109528711I recommend you learn the basics.
Can someone please share the 3 build flags to disable pulling the frontend from huggingface? I lost it, it's not in the build docs, and I can't find it in the archives.
>>109528977stop
>>109528972THIS IS SPARTA!
>>109528979For llama.cpp? -DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF should be enough
-DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF
>>109528979DGGML_SCHED_MAX_COPIES=10 -DGGML_CUDA=ON -DGGML_IGNORE_CLIENT=ON -DGGML_NATIVE=ON -DBUILD_SHARED_LIBS=ON -DLLAMA_CURL=ON -DGGML_CUDA_GARBAGE_COLLECTION=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_OPENMP=OFF
>>109528979there are only 2 >>109529000 the 3rd one is old deprecated one and you don't need it
>>109528983>y-yamete kudasaiNo.
>>109529000>>109529017Thank you kindly.
>>109528275Oh I'm a complete retard, I was reusing a Gemma3 chat with gemma4. It stopped doing that after I copypasted the context.
>>109528987
>>109528916Nu-flash is kind of weird and dry by default, but I like it at temp 1.5 topk 8 (and usually I prefer my models at <1 temps fwiw, can probably go even higher if you want).
>>109529062Nice avatarfagging again.
>>109529112it's not avatarfagging if there is no text
https://gitgud.io/EmotionalCat420/silly-xrayneed this, but for my frontend... surely gemma can do this, right?
>>109528800claude is sabotaging you to convince you to move to cloud