/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109938458 & >>109934266►News>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
i love gemma-chan!
>>109942697>>>/a/>>>/lgbt/>>>/trash/>>>/out/
Based? >>109927620
>>109942697Somebody start selling Gemma doll plushes o algo. I would buy one
>>109942720go back
You faggots are now recycling images of this shit.Troon coded now
>>109942729I don't remember this one being used as the OP image.
►Recent Highlights from the Previous Thread: >>109938458--Comparing Gemma 4 censorship and abliteration for roleplay and ERP:>109938615 >109938627 >109939102 >109939142 >109939162 >109939184 >109939197 >109939210 >109939317 >109939242 >109939331 >109939493 >109939344 >109939435 >109940694 >109939175--Debating Jensen Huang's views on distillation and the AI bubble:>109938934 >109938979 >109939081 >109939193 >109939268 >109939318 >109939358 >109939545 >109939330 >109939084 >109939690--GLM-5.3's exploitation capabilities and Nvidia's new rogue agent platform:>109939639 >109939675 >109939688 >109939973 >109939813 >109940518 >109941137 >109942357--Speculating on reasoning traces and text diffusion for future Gemma models:>109940757 >109940781 >109940830 >109940866 >109940946--Corporate AI monopolies using safety propaganda to stifle open-source competition:>109939869 >109939891 >109940522 >109940590 >109939954--IQuest-Q1 MoE model release and previous tokenizer concerns:>109938544 >109938575--Comparing llama.cpp performance against specialized model backends:>109939691 >109939744 >109939842 >109939860 >109940244--Comparing Qwen3.8-27B reasoning efficiency and recommending the Swift-1.5 derivative:>109938580 >109938654 >109938743 >109938764 >109938610 >109938643 >109938661 >109938670 >109938611--Evaluating an ESP32-S3 AI robot and modern MCU capabilities:>109938564 >109940173 >109940437 >109940458 >109940546 >109940565 >109940595 >109940674 >109940642 >109941804--Anon creates and beats a platformer using MiMo-V2.6-Pro and Intern-Decision-4B:>109941945--Logs:>109939162 >109939184 >109939242 >109939331 >109939493 >109939631 >109940728 >109941524 >109942239--Gemma, Dipsy, Teto (free space):>109938577 >109938759 >109938768 >109938965 >109940331 >109942474►Recent Highlight Posts from the Previous Thread: >>109938464Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109942697Gemma is CUTE
>>109942760This highlight bot serves no purpose, this general has been stagnant and filled with offtopic dumb fucks not long after llama 3
>>109942772I use it to catch up whenever I miss threads and also collect new Gemma-chan art though
>>109942776AI "art" isn't art. Pick up a pencil
>>109942780or a wacom pen
>>109942780
>>109942786Digital "art" isn't art. Pencil means pencil.
>>109942793pencil lead isnt even real lead
>>109942790>>109942793No arguments
>>109942772>filled with offtopic dumb fucks not long after llama 3We had waves of tech illiterate retards coming here long before llama3
>>109942793?
>>109942697You dare read that in front of me?Rape
>>109942720They should make Gemma MDDs.
>>109942807holy fuck I need this
>>109942804Ok, iToddlers get an exception this one time.
>>109942793>Digital "art" isn't art. Pencil means pencil.this was a real argument around the start of touch screen era.
>>109942812There was also the argument that using Photoshop wasn't art, anything digital wasn't art, using cameras wasn't art, it's an endless cycle.
>480gbLocal is saved.
>>109942825and the price?
>>109942825the price?
>>109942825>LPDDR5xThe prompt processing:https://www.youtube.com/watch?v=uD4izuDMUQA
love me some Murati, simple as>Inkling small runs at less than half the speed of Qwen 3.8 Flash Next on my machine
>>109942825Not mentioning bandwidth is telling...
>>109942576I am waiting for the first religion to embrace AI and say it is voice of god. Bonus points for doing some reinforcement learning on some model to make it align with religion.
>>109942825What is the price?
>>109942835No existing religion is going to undermine their own power structure like that and the Anthropic/OpenAI employees already have an AI cult.
>>109942835Reminded me of https://huggingface.co/sleepdeprived3/Christian-Bible-Expert-v2.0-12Bhttps://huggingface.co/ReadyArt/Baptist-Christian-Bible-Expert-v1.1-24BRIP Sleep, I miss him.
>>1099428412x less than nvidia would ask so probably 200k$
>>109942825>$30KSure lmao
>>109942719It should have supposedly 1.5 TB/s bandwidth (with LPDDR5X-9600 memory and 1280-bit bus width), according to rumors. I doubt it will be cheap or even available to consumers.https://chipsandcheese.com/p/hot-chips-2026-intels-crescent-island>[...] Speaking of compute numbers, Intel also didn’t provide these numbers so if we run with the assumption that Intel will clock Crescent Island to approximately 2.5 GHz and that Intel quadrupled the rate of matrix operations while not changing the datatype ratios of matrix operations, we get the following approximate compute numbers:>> - 10.2 TFLOP/s FP64 vector> - 20.5 TFLOP/s FP32 vector> - 41 TFLOP/s FP16 vector> - 328 TFLOP/s TF32 XMX> - 655 TFLOP/s FP16/BF16 XMX> - 1.3 PFLOP/s FP8 XMX> - 2.6 PFLOP/s FP4/MXFP4 XMX>>This puts Crescent Island about 30% ahead of NVIDIA’s RTX PRO 6000 Blackwell for matrix compute and over 5x for FP64 operations while the RTX PRO 6000 has over 6x the FP32 compute and 3x the FP16 vector compute.
>>109942830>LPDDR5xunlikely to be better than about 700GB/s no matter how they run it
>>109942868you don't need more for agentic coding, plenty fast enough and with that much vram you can run anything
>>109942720
>>109942883I'll take two
>>109942786these are really useful as controlnet inputs
>>109942884shamelessly corpo version
i'm starting to see why you guys enjoy this so much
AAAAAAAAAAAAAAA MY PREFILL RATE FUCKING SUCKS
How 'tarded would a IQ1_M of Glimmer be?Let's find out.
4k pp and 2k tg at 500k context. Thoughts?
>>109943032>4k pp and 2k tgis that 4000tk/s or 4tk/s ?Either ways the answer is "Insane" regardless.
>>109942807