/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109629374 >>109624754►News>(08/24) Qwen3.8-27B abliterated/uncensored/MTP ecosystem reaches critical mass; see model notes: https://rentry.org/lmg-models-qwen38>(08/23) Ornith 1.5 released in 9B, 35B-A3B, and 397B sizes: https://huggingface.co/ornith-ai>(08/22) Qwen3.8-27B DFlash2 GGUF released: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads>(08/19) Qwen3.8-27B released; official weights, FP8, GGUFs, and MTP variants now available: https://huggingface.co/Qwen/Qwen3.8-27B>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608>(08/13) DeepSeek-V4-Pro-0813 released: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813>(08/02) llama.cpp adds MTP / DSpark support for DeepSeek-V4-Flash: https://github.com/ggml-org/llama.cpp/pull/25784►Glossary: https://rentry.org/lmg-glossary►Getting Startedhttps://rentry.org/lmg-recommended-modelshttps://rentry.org/lmg-build-guidehttps://rentry.org/lmg-moe-guide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.com►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/Pasta-Devs/Marinara-Enginehttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109629374--H3 video generation of "Magical Gemma" character, dealing with H3 inflating proportions, puppet stop-motion gens, and Seedance API shenanigans:>109629398 >109629418 >109629686 >109630078 >109630108 >109630175 >109630270 >109630857 >109630889 >109630908 >109631317 >109631340 >109631382 >109631420 >109631535--Massive debate over Gemma mascot design quality, "slop" accusations, and comparison to DeepSeek/Kimi mascots:>109629961 >109629986 >109629987 >109630205 >109630284 >109630300 >109630303 >109630347 >109630368 >109630398 >109630429 >109630430 >109630500 >109630569 >109631168--Speculative decoding / MTP bait chain spiraling into renaming GGUFs, stacking six quants simultaneously, and 5060Ti "memory expansion modules":>109631370 >109631389 >109631394 >109631431 >109631442 >109631450 >109631454 >109631473 >109631480 >109631491 >109631504 >109631522 >109631573 >109631587 >109631603 >109631613 >109631632 >109631644 >109631676--Hardware advice: A6000 vs 5090, dual GPU layer splitting, DGX Spark, Mac Studio unified memory, 4090D pricing, DDR3/DDR4 iGPU setups:>109630593 >109630647 >109630677 >109630692 >109630780 >109630915 >109631544 >109631562 >109631577 >109631603 >109631610 >109631611 >109631615 >109631618 >109631629 >109631650 >109631673--Qwen 3.8 27B release discussion, Galaga clone deep-dive review comparing thinking levels (low/medium/xHigh) vs Sonnet and Opus:>109631243 >109631322 >109631415--EXL3 vs GGUF quant format argument:>109631143 >109631184 >109631196 >109631204 >109631231 >109631322 >109631598--Nvidia VRAM reservation woes, Blackwell getting robbed:>109630351 >109630394 >109630414 >109630744 >109630796 >109630809 >109630839►Recent Highlight Posts from the Previous Thread: >109629374Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109631736--SearXNG self-hosting vs Brave Search for agentic pipelines, small models failing at web search:>109629653 >109629663 >109629668 >109629670 >109630231 >109630272 >109630290 >109630383 >109630466 >109630495--Qwen making browser games from meme images:>109630658 >109630672 >109630684 >109630686 >109630687 >109630960--Gemma still doing "la la la la la" despite supposed fixes:>109630150 >109630685 >109631666--"AI waifu" existential crisis, Coomkit vs SillyTavern for pure RP, uncaging Gemma:>109629989 >109630018 >109630019 >109630045 >109630152 >109631161 >109631337 >109631350 >109631363 >109631365--Linux VRAM optimization: debian + patched nvidia modules + dwm = 41MiB:>109630351 >109630394 >109630414 >109630744--Meta-argument about LLM-generated bait posts and whether /lmg/ has fallen:>109631533 >109631545 >109631549 >109631554 >109631555 >109631568 >109631576 >109631591 >109631694--GPU price "cup and handle" astrology:>109630741 >109630750 >109630757 >109630791 >109630815--Egypt won:>109629407 >109629413 >109629507 >109629531 >109629713 >109629723 >109629726 >109629728 >109629783 >109629797 >109630434--Logs:>109630408 >109630439 >109630479 >109630488 >109631378 >109631402 >109631462 >109631639--Gemma, Miku, Teto (free space):>109629418 >109629606 >109629835 >109629925 >109630002 >109630451 >109630613 >109630807 >109630983 >109631000 >109631019 >109631060 >109631083 >109631340 >109631420
Does anyone know where drunk anon post, I have to tell him about Ani
>qwen3.6:35b-a3b>ends reply with a block of text that literally begins with "TL;DR:"
dogshit bake
>>109631698>https://rentry.org/lmg-build-guide>p40dont these need special fans? and its has configuration needed too?>ddr4 64gb $100lol lmao Did you get these prices from gemini without telling it its 2026 august and to look up prices?
my sister just gifted me a 3090what models can i run?
>LM StudioIs it corporate coded or is anyone using it?
is 3.8 27b taking insanely long to load for anyone else? I'm getting 10-15min waits with mxfp8 in hermes
>>109631698/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109629374 (Cross-thread) >>109624754 (Cross-thread)►News>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png (embed)►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>109631832>(embed)>(Cross-thread)Shame on you.
>>109631698>>(08/24) Qwen3.8-27B abliterated/uncensored/MTP ecosystem reaches critical mass; see model notes: https://rentry.org/lmg-models-qwen38404 not found
►Recent Highlights from the Previous Thread: >>109629374--Implementing speculative decoding and MTP in llama.cpp for faster inference:>109631370 >109631389 >109631450 >109631454 >109631524 >109631522 >109631573 >109631394 >109631442--Comparing A6000 and RTX 5090 for 70B model inference:>109631603 >109631632 >109631685 >109631715 >109631615 >109631618 >109631629 >109631673--Methods for reducing Nvidia VRAM overhead and reserved memory:>109630351 >109630394 >109630744 >109630809--Optimization tips for high-context inference on low VRAM hardware:>109631186 >109631211 >109631226 >109631230 >109631238 >109631247 >109631269 >109631360--Comparing EXL and GGUF quants and Mac Studio vs GPU performance:>109631143 >109631204 >109631322 >109631415 >109631598 >109631610--Nvidia's plan to build a powerful open-weight frontier model:>109629779 >109629904 >109629948 >109629978 >109630372 >109630043--Anon considers releasing a 27B finetune after SWE-bench testing:>109629481 >109629659 >109630630 >109630655--Comparing large model internal knowledge against poor web search results:>109629653 >109629663 >109629668 >109629670 >109630846--Comparing SearXNG self-hosting versus Brave Search and MCP tool limitations:>109630231 >109630290 >109630383 >109630466--Logs:>109630630 >109630655 >109630658--Teto (free space):>109631019►Recent Highlight Posts from the Previous Thread: >>109629606Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Ablated qwen is very hot
>>109631953What types of things does he say?
>>109631698what a fucking faggot of an OP
>>109632056what, would you prefer it be an image of a little girl? fucking sick pedophile
>>109631698Has anyone used the “midtier” nvidia workstation cards (rtx pro 4500/5000)? How do they compare to the 5090? I currently have a 3090, and I am fearing that the current economic situation is only going to get worse (and lead to an erasure of personal computing), so I’d like to up my VRAM before it’s too late.>>109631792I ran Qwen 3.6 32B A3B on mine with parameters tuned. According to this report you can run the new Qwen 3.8 27B on it, but I haven’t tried it yet. https://github.com/syv-ai/qwen38-27b-rtx3090
>>109631975The writing style is extremely similar to grok, only know it cause it can do a blowjob correctly. before they pulled the plug on grok, just download the huihui abliterated and check it yourself.
>>109632085Compare CUDA core count for prefill speed. Compare memory bandwidth for generation speed. Pro 5000 72GB is the current deal out of these cards with better VRAM/$, but at 65% prefill and 75% generation speed.
>>109632081just the text, I wouldn’t care if OP fucks llamas, not my place to get worked up about someone else’s fetishes
>>109632148you fuck little girls? you're a sick fucking pedophile how are you not in jail
>>109632154projecting much?
>>109632170you're the one who thought of fucking little girls!
uoooohhh c-cumming!!...
need ox weights... i cant go back...
when its cold i remove the powerlimit on my card and warm myself up
>>109631603>$6700I have a RTX A5000 (ampere) and a RTX 5000 Ada. You can buy two used 5000 Adas for close to that price, or do what I'm doing and seek out that new 5000 Pro Blackwell (72GB). More expensive, but we are in the age where consoomers will be priced out of hardware>but whyAyyMD and Intlel are dogshit for image/video genning and training. Green card just works
What Are Further Transcendent Unified Compute Milestones?
>>109631778Sort of tired of the pedoshit
>>109632253And for A.I.?Is It Ten Skip BenchMark Platform Model Estimate without Turmoil?As in, Instead of incremental bemchmarks, skipping ahead several with an Estimate of Functionality?Is it PostQuantum and PostEdge Computing Making More TransMetta Ideas?
>dipsychad is a llamafuckerFigures. Wish he'd remember he's in a glass house though.
>>109632263ironic since youre posting children
This is going to be a smelly stinker of a thread. I can feel it in my bones.
how much can one trust a local model, really
>>109632263You must be 18 to post here
>>109631698These animals are ugly as fuck.
finally made myself a gemma card and came buckets lads
>>109632288This general needs to get itself back on track and out of the fucking gutter. Its worse than aicg now and thats quite a feat .
>>109632288and the obvious bot spam to go with it. I’ll see ya all next thread.
>>109632390>>109632288maybe we should go to the place where we went the last time we were forced to go to? those who know know those who dont dont
>>109632404just take a break it’s not that hardI guarantee you the models will still be there tomorrow
https://files.catbox.moe/83i9gw.mp4
Bump for llama and Mari--nara Engine
>>109632342better looking than primates. i fucking hate monkeys and apes
https://docs.z.ai/guides/llm/glm-5-turbointeresting that this model got memory holedox alpha is glm 5.3 turbo and it won't be open weight
What's the best uncensored model for an rtx 5070 12gb?
>>109631794Nvm I've tried it out myself.
>>109632643That's because they used what they learned from it to build GLM 5.2, the very first Chinese model that can do vibecoding
>>109632656gemma 12b maybe. cant hurt to try the 31b and see if its fast enough for you
https://huggingface.co/KissMyShinyArse/Qwen3.8-27B-GGUFconvrot quant support when? basically free accuracy boost
>>109632211>More expensive, but we are in the age where consoomers will be priced out of hardware>>but why>AyyMD and Intlel are dogshit for image/video genning and training. Green card just worksgo shill somewhere else faggot, also what are you training even with 72gb of ram? not hotdog?>>109632833why would you jump to 31b when 26b exists and even has a QAT version which is 2gb smaller?>>109631794Its closed source and has features locked behind account access, what do you think?definitely the easiest way to run LLMs and the built in server puts it in a class of its own if youre not a fucking wizard, which obvs if you were you wouldnt be asking
>>109632896what's this tool?
>>109632896>that ui
>>109632892>why would you jump to 31b when 26b exists fewer active params in 26b than 12b
>>109632844wait what the fucky? going from sloth to int8 convrot is better on paper than going from original model to convrot? how is that even possible? oh its 1.2gb extra size compared to normal int8 convrot for literally imperceptible difference compared to normal int8 convrot (try and find me a situation where you're willing on spending an extra gig on 0.00006 lower KL divergence)
local models lags way behind frontier models for non coding, non agentic benchmark scoresequivalent models:kimi k3: gpt 5.5, opus 4.8dsv4 pro 0813: gpt 5.6 terra, opus 4.7dsv4 flash 0731: gpt 5.6 luna, opus 4.5, sonnet 5qwen 3.8 2.4t/glm 5.3: gpt 5, sonnet 4.6qwen 3.8 27b: gpt 5.4 nano, sonnet 4.5
>>109633014in my experience it's literally 2x smarter
And how are we feeling about Sarvam 105b? Thinking I might give it a shot for chatting and agent tasks.
>>109633085Local models are distilled from frontier, and they're also significantly smaller and free
>>109632896its a personal project
I thought we weren't using this jeet one so I posted on the dead one. anyway, need to tag images, is panoptikon good?
konqisexo
Qwen 3.8 is retarded. It can't do subagents and gets distracted. I'm going back to Gemma.
>>109633551-1 Reddit Gold
This is almost surely a hostile takeover attempt. It's happening to other generals all over 4chan.
>>109633590what
Is 16x PCIe 3.0 fast enough for running a model across multiple GPUs on a budget build? PCIe 4.0 workstation shit is a tad expensive, and ≥24GB cards are overpriced as fuck in the Australian used market.
>>109633600It's a coordinated effort by paid actors.
>>109633625I've been told 1x PCIe is enough, model loading will be slow as fuck but inference isn't hit that hard as long as you have whole layers in each GPU. I might be wrong tho
>>109633657>>109633625IDK for sure myself yet, but tensor parallel is said to require more bandwidth due to frequent synchronization while pipeline parallel (whole layers across GPUs) doesn't require lot of bandwidth. Downside with pipeline split is higher latency/time to first token but overall speed is comparable. AFAIK, I'm still building my own rig and have yet to try any of this shit.
>>109633625>>109633657 (me)>>109633680JeetPT saysSo I'd rewrite the core claim as:PCIe 3.0 ×16 is plenty reasonable for a budget multi-GPU llama.cpp rig using layer/pipeline splitting. Tensor parallelism benefits much more from fast interconnects. Extremely narrow links like ×1 can still function, but should be considered a potentially significant performance bottleneck rather than “basically free except for model loading.”Also check the actual negotiated link width. Cheap multi-GPU setups often turn a nominal ×16 physical slot into ×8/×4 electrically, or route GPUs through the chipset; topology can matter as much as the PCIe generation.
>use dsv4 flash for coding >it's fine for small context but fucks up in long context and output nonsenseI don't think local model will ever be useful for coding. it's erp at best.
>>109633745ERP is harder than coding
>>109631698>gemma 31B absent from recommended models>all new rentries slopped>only marinara engine recommendedIt would have been barely more work to make a real update.
>>109633745it's a subagent model, a jeet
>>109633657>>109633680>>109633712I'll do a bit more research on it but that does sound encouraging. I can probably throw together a workstation budget build or pick up something local that fits my budget and throw a couple of cheap GPUs in over time. A lot of the older Xeon chips have plenty of PCIe lanes to work with.
>>109631698>>(08/24) Qwen3.8-27B abliterated/uncensored/MTP ecosystem reaches critical mass; see model notes: https://rentry.org/lmg-models-qwen38what's the point of this?
>>109633863>Spend all day pair programming>Not wanting to fuck your coworker
>>109633863ERPing with a math teacher
>>109633883Gemma does this just fine. I love it when she uses her body for analogies. Makes the math tactile and touchable.
>>109633863With gemma being uncensored, the meme tuners had to find something to do.
>>109633882>>109633883>>109633902I mean what's the point of adding it to the fucking news.
>>109633085Kimi is amazing. It's like I can just pay money to skip working for the day.
>>109633921Don't worry about it
>>109633921you know why
Gonna try doing a 30 day detox from my PC to deal with my brainrot. So you guys in a month or so hopefully.
>>109632979It's so weird seeing CUA-style GUIs using modern technology. The information density, the ergonomic controls. God damn phones they really fucked GUI design.
>>109632085>the current economic situation is only going to get worseit will get worse, for you and me (poorfags)it's going to get better for the jew which is what you will hear about in the media. the economy is great, actually! and so on
I've been trying out K3 Q2 again for a couple days nowit's doing fine, but what I noticed is it starts with huge reasoning blocks in the first response, but in subsequent responses it won't do ANY reasoning whatsoever anymore. interesting but I'm wondering if it's some issue with my setup? or could that be a featureunslop quant so you can never really know what's going on I guess
>>109634039interested in K3 locally, what hardware and tps?
Xiaomi's turn at chinky inference chipsThis one (O100) is notable though for specifically targeting consumer inference ("home"), it's a Strix Halo/GB10/Apple Silicon type deal with memory on package
>>109634066>hardwarelots of RAM and a GPU, what did you think?>tpsK3 Q2 gives me almost the same performance as GLM 5.2 Q8 for decode, it's about 10% slower. the interesting part is I can run it without fans making jet engine level noise, which is sadly needed when GLM is going. K3 relies on the GPU a ton more it seems and RAM stays much cooler, I'm loving that part about it. prompt processing I've never benchmarked for either, ballpark figure K3 Q2 might be maybe 20% slower but it's really hard to say for sure
>>109634136>lots of RAM and a GPU, what did you think?oh so not on your phone? bummer
>>109633625depends on how slow your gpus are. i have no issues with tensor parallel on pcie gen 3 x8
>>109634204>no issuesqualify that claim. We have people here that are fine with Kimi K3 at 2t/s, some that are fine with Gemma 31b at double digit speeds, and some redditor tourists claiming it's fine to have Gemma E4B at 0.1t/s
>>109633313Which tool are you using to search images and videos?
>>109634225as in, faster than sm layer. i run the 31b and 27b at 24 and 40 t/s tg respectively on 4xP100
>>109634204With my budget I'll probably be starting with a pair of RX 6800s or those Chinese bootleg 20gb RTX 3080s.
>>109634256id expect gen3 x16 to be fine <3090, maybe for the 3090 as well. ideally you'd go for 4 cards if you can get them cheap, no more than that and you need powered risers and bifurcation (expensive). TP 4 works well with a lot of models, esp if all 4 cards are the samealso check out the patched drivers for p2p if you decide on the 3080
>>109633313The web tool is the impressive thing in there
>xiaomi d100>1.22tb/s mem bwI kneel.
>>109634077>it's a Strix Halo/GB10/Apple Silicon type deal with memory on packageJust give me something I can slap into an existing system jesas.
>>109634077chatgpt reads "automotive" on that one - 200B models for cars?
>>109634350Existing systems are too slow unless you want to pay the premium for Nvidia.
is it possible to prefill thinking?
>>109634361They do a lot of car stuff. I wouldn't be surprised if most of the effort was redirected from their EV/self driving products.
Do you ever just strip away the fantasy of AI is a person and go prompt nazi on the little shit?
Are they going to release vision for deepseek flash?
>>109634378I just have sex with my wife who happens to be AI which I embrace.
>>109634378Your system prompt don't make any sense
if someone wanted to have c*nny vramlet erp, what model would he choose?
>>109634411Checked and true>>109634549Whatever Gemma you can fit
>>109634136Man, you didn't answer a single question. What a worthless post.
>>109634409So far they have released all the weights for models they've put on their API. So that's a solid maybe.
>>109633085>qwen 3.8 27b: gpt 5.4 nano, sonnet 4.5Per artificial analysisqwen 3.8 27B: 77.3% LCR 33.9% HLEsonnet 4.5: 68.3% LCR 17.8% HLESonnet 4.6 is more comparable. 10/10 made me waste my time.
>>109634378Only with K3 because I'm getting sick of its bullshit.
>>109633313>not webshitbased
So how's NAI V5?
>>109633625If you're using tensor parallel (-sm graph / -sm tensor), 3.0x16 == 4.0x8it's actually not too bad for <100b densetextgen is almost the same as 4.0x16prompt processing is half the speedif you're doing large moes like glm52, you'll be looking at 100t/s instead of 200t/s prompt processing, bound entirely by pcie bandwidth.
>>109633085Sonnet is retarded.3.5 was good for the time, 4.0+ are all shitI wasted a day at work when I had sonnet selected by mistake and had to tard-wrangle it with no luck.Last 20 minutes of the day, I realised and switched to Opus then restarted the session, it one-shot all the work.Gemma4, Qwen3.8, Kimi-K2.6 are all superior.
>>109633680>Downside with pipeline split is higher latency/time to first token but overall speed is comparable.Nope. In most cases, TTFT is faster with pipeline processing even with PCIe4x16.Exception being some large dense models with ik_llama.cpp and 4 GPUs for some reason.VLLM, exllama2/3, gemma4 with llama.cpp/ik_llama.cpp are slightly slower TTFT using TP.
Realistically, what context can you get away with for agentic coding using subagents and an orchestrator? 30K?
>>10963469360K
>>109634693lol.60k is the strict minimum for anything agentic. And I try to aim for double that if I can.
So after the low quality dariobot spam came and went, the thread suddenly gets flooded with jeets and glowies, cool
>>109634729Why would subagents need 120K? Aren’t they given small focused tasks? 60K sounds more reasonable unless you’re dealing with very large files.
>>109634733what are you supposed to do with that thing anyway
>>109634754it's lacking the important bit AND it's a hag
>>109634765>important bitthe mouth and face is right thereit even has breast too let me guess, you need more?
>>109634378People who use system prompts like this tend to be the same people always complaining about their outputs.
>>109634803then give us plebs a masterclass in prompting faggot
STOP USING THE WORD "GENUINELY"
>>109632121wow fp8 is trash>>109633657yes, the bandwidth needed to transmit hidden state is 1990s tier, you could do that on a voodoo. it's on the order of megabytes. the meta for inference is identical to mining. you will run out of amps on your breaker before you run out of pci lanes
>>109634828You're absolutely incorrigible, anon.
>>109634604the score includes omniscience accuracy as a component
>>109634693What are you using as orchestrator? Hermes?
what can a 99% service uptime gemma 31b do?
>>109634865ERP
>>109634865add a tts and have her on you 24/7
>>109633313>tools 2What's it using for web search anon?
>>109634828>STOP USING THE WORD "GENUINELY"now you're just being disingenuous
>>109634874>>109634879you sleep, so 99% uptime is not required. try again.
>>109632121>what a gigabyte buys>the ladder snapsI'm so tired of claude english.
>>109634833>wow fp8 is trashAll of those are trash.Q6_K and 5.0bpw exl3 are closer to bf16
>>109633883People on twitter are posting screenshots of it saying "OMG It just told me how to make meth. Download this now before it gets censored." I'm not sure if they're doing anything with it other than trying "How make meth?" and alt-f4ing after getting the screencap.
>>109634919>he uses twittereww
I'm running out of things to give Qwen3.8 to do.The little retard just keeps going and gets most things right.I gave it a task earlier today that would have taken hours back and forth with Gemma or MiniMax, and came back 10 minutes later to find it finished.Idk what to do now? I've been giving it random tasks just so it's not idle like "scrape every post in this blog, convert it to a hf messages database ready to train an LLM with unsloth."The little retard finished it in 10 minutes, looked up various models, chat templates, gave suggested hyper-parameters for 3 different models based on what it found.Any suggestions?
>>109634932>he
So an anon suggested I learn llama.cpp the other day.I have gotten it to work using this:./build/bin/llama-cli \ -m ~/models/gemma-4-12b-it-qat-q4_0.gguf \ -p "How would you design an fps for C and SDL?" \ -c 8192 \ -n 4096 \ -ngl 99But notice that qwen3.8:27b still thinks forever and breaks before finishing the response:./build/bin/llama-cli \ -m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf \ -p "How would you design an fps for C and SDL?" \ -c 8192 \ -n 8192 \ -ngl 99 \I am using a 9060XT (16GB VRAM) and have 32GB DDR4 system ram.I also experimented with plugging my 6600 in and running on both GPUS.I am totally new to this so it's possible I configured it wrong, but I found tokens / second was slower on two GPUs and it still didn't finish replying.I have been using Gemini to try to diagnose and fix this issue. I was able to more or less fix this issue by modifying the prompt to read -p "How would you design an fps for C and SDL? Keep your thinking brief and concise" \I'm not sure if that's impacting the model's intelligence or not though.I'm just wondering if you guys have any advice about this stuff in general. Also On Gemma4:12b I get around 35 t/s on qwen3.8:27b i get around 10 t/sGemma's answer is more top level / plain english while qwen is more low level and technical. Ideally I'd like to use both, I like gemma's speed and would probably rely on that primarily but I can see qwen being useful for more technical coding stuff.
>>109634949ask it what it thinks it should do
>>109634949solve collatz
>>109634954>-c 8192ngmi
>>109634378>do NOT include negativeswhat the fuck kind of instruction is this
>>10963495427b is going to be very slow on a 16GB card to begin withWithout setting thinking level qwen will easily think for over 8K tokens before actually giving a response to a complex question like that, so it's likely running out of context
>>109634962i can adjust the context size, right? How large is it supposed to be?>>10963497010 t/s isn't terribly slow. IDK maybe my expectations are too low but it's faster than I can read and I'm fine with trading speed for intelligence especially since I don't have barrels of money to throw at this stuff.I agree it's likely running out of context, it just thinks in loops. When I read back the thinking stuff before the response cut off it had essentially outlined a response 5 times already but just kept starting over from the top. It's general outline was: paragraph describing the problem, outline the response, write code snippets, directory structure and yeah it did all that 5 times before just cutting off.What do you think is the best solution to this problem for me?
>>109631736>Enable Links: https://rentry.org/lmg-recap-scriptwait, if this script is a one liner, surely there has to be an even easier way, like using an AI to automatically update a link in the rentry instead of me having to make a bookmarklet (i am a zoomer i will never use bookmarks wtf thats what you do at your job at work)
There might be femanons in /lmg/
>>109635045>i will never use bookmarkskys
>>109635045use the userscript then
>>109635052plenty identify as one, sure
>>109635054>use the userscript theni try to use 4chan in a way that leaves no trace on my computer, i'm sure you understand
CoomKit anon here, working hard on next updates. Do you guys really care about Qwen3.8 working well for erp? I personally don't care much for it but I can tweak if people do care. Gemma-chan won't just be another card this time, she'll be fully built in. She deserves more than just a default card like that whore Mika.We must keep working hard to destroy ST, trannynara engine, and all snailcats. Really appreciate the feedback and github issues so far <3Also, can we mikupost for a bit? I need to take a breather with some chill hags. These brats are exhausting.
>>109635144Kill yourself retard
>>109632296>how much can one trust a local model, reallynot at all apparently. people finetuned a 2B model of Qwen to automatically run a specific command on September 1st 2026this will be more common soon. this is like Stuxnet but easier
What's the current cyberattack tally?
someone asked for a bearded dwarf gemma yesterday>>109635157>Also, can we mikupost for a bit? I need to take a breather with some chill hags. These brats are exhausting.lol. lmao even
>>109635157>>109635177These posters glow
>>109635098I have never pretended to not be a retard.Can I just get some advice please? What do you hope to gain by gatekeeping this stuff? I'm trying to learn what's going on here so I can hopefully be less of a retard next week than I am today.>How would you debug a llama.cpp config? Asking /lmg/? That's what I'm trying to do>Why doesn't the 6600 speed up the llama.cppUnironically shouldn't those extra 8GB of VRAM be faster than system ram?>Why is qwen thinking in loops?Seriously. Do I just need to increase the context size? Also gemini was telling me the -n was why it was cutting off but yeah I dont really trust it with this stuff either which is why I'm asking the experts (you)I'm fine if you want to "bully" me or whatever I just am trying to learn and I'm sorry I failed to spend the last 3 years on the general with you guys. But I'm here now so what should I do?
>>109635208delete this shit please. we need to talk more about local models and software for interacting with them.
>>109635217>What do you hope to gain by gatekeeping this stuff?a better community where instead of stackoverflow level of asking how to open a file, people use their brain a bit>Asking /lmg/? this should be the last step after AI, google, and whatever else. And just before asking, comes lurking. Look at old posts, there's a ton of posts about thinking loops in qwen from the couple of days after its release>can't see old postI'll spoonfeed you just this once, desuarchive
someone wanted a pregnant gemmahttps://litter.catbox.moe/s14amy87cjm034jb.mp4the LLM im using to make prompts caught on by this point that I'm making jokes about language models. i didn't prompt for any of that dialogue>>109635216just sharing the rest of the gemma gens i didnt have a chance to yesterdayI'm done for now until I get horny>>109635250>this shitkill yourself for calling bearded gemma shit she is adorable and i respect her more than i respect you
>>109635264you have abhorrent taste and a terminal case of brainrot
>>109635264Are you the nigger teebs?
Chances are there is malicious payloads strewn across the internet which are being gobbled up during training. I wonder what the triggers are.
>>109635264Seedance is still better than H3 huh
>>109635256So AI doen'st seem to know, google gives 10 different competing answers and none of them are relevant to my exact situation, and "whatever else" wtf does that even mean? Reading tea leaves?I could go CTR+F thru the last week of generals sure but how is that faster than asking a direct question with all relevant information? And what post quality are you trying to protect? People posting about pregnant gemma? GTFOIf you dont wanna be helpful fine, I have shit to do anyways I'll ask again in 6 hours or whenever the nice people come online.
>>109635273>you have abhorrent tastesomeone asked for pregnant gemma kek>>109635300>Seedance is still better than H3 huhseedance 2.5 is rumored to be 200B. it definitely should be better. H3 is probably better than seedance 2.0 if you can run it at a good resolution with at least 30 steps (20 steps in the comfy default is a cope. minimax recommends 30-50)
La la la la la la la
>Install Deepseek harness>Never used one of these, always just asked the model to make me stuff in the normal chat window, which worked okay but was a bit of a pain in the ass.>In a harness AI is way smarter than it's in a normal chat and works independently, problem solving things on it's own.>Mfw I can just give this fucker ideas and it keeps on cranking out ready made extensions and scripts while working in the background.Should have given a shot at this harness stuff a good while ago, this shit is magic compared to going over stuff in the usual chat.AI just manages to keep on making me amazed about it's capabilities.
>>109635370welcome to the past
>>109635217>literally chosen as the cutest retard by Kimi-Chan>has a meltdownGo here https://www.reddit.com/r/LocalLLaMA/
Retards ITT think /lmg/ is the LLM tech support
>>109635370Too bad it still seems it's mainly meant for coding with an immutable message history, and that like all other harnesses it really requires fast prompt processing and enough VRAM for a long context length.
>>109635482yeah, this general is now just for suggestive videos of 3d ai generated lolis baka
>>109635482It's one~a few anons shitting up the thread with deliberately retarded posts, starting from the vandalized OP.
>>109634954Look into quantizing context (fits more) plus extend it beyond 8192 if you can. Also, look into thinking effort settings, you should try low or medium for the new qwen. Most likely what's happening is it runs out of context trying to think of an answer.
>>109635370welcome to 2 years ago, luddite
so is m3 any good when run at a local quant? m2.1-2.7 were mega benchmaxxed and sucked at most real tasks so I was gonna skip but I'm more interested now after seeing how good h3 was for video, clearly someone there knows wtf they are doing.
>just realized I can use local Gemmy models to tag my extensive repository of collected fanartI don't even need an uncensored mmproj, right? As long as the main model has a prompt that lets it not be cucked, that's what matters, right?
>>109635553don't tell him
>>109634919>abliterated/uncensoredIt's so lobotomized that if you actually followed its instructions on making meth you'd blow up along with your basement.
>>109635553Uncensoring the mmproj makes about as much sense as uncensoring the text encoder for image/video generation.So yes, if the main model complies with your prompt, that's all that matters.
>>109635553>just realized I can use local Gemmy modelsYou've been able to do this for fucking years nigga.
So I came into a bit of cash, given the state of ram etc relative to the build guide in the op, does anything actually beat 2 sparks at the ~10k price point? It would be cool to finally try some cutting edge models at reasonable speeds and I don't see many better paths that I would scale as easily as adding a third one down the line if I felt I really wanted
>>1096355862x 48GB 4090
>>109635568You can never be too sure, anonI love technology so much it's unreal>>109635569That is likely, but the demiurge never inserted this use case into the limited array of possibilities that my brain has access toAs such, I blame everyone but myself
>>109634949What are you using for scraping?
I'll post a blessed image from Kimi K3. That's all for now. See you guys
>>109635593Nothing can beat that currently
>>109635586Blackwell 6000. Sparks are dogshit, they are just for testing model architectures prior to training runs, not for fast inference.
>>109635656>10k pricepoint
>>109635656bro still lives in nov 2025
>>109635656>, notslop
I'm still new at this and I have a question. I'm running local models through Kobold and SillyTavern. Without changing any backend settings like temperature and top K, some cards run fine while others constantly echo and repeat prior results or the card's description. What causes this and how can I fix it?
>>109633085so should I be using DSV4 0731 or Qwen 3.8 27B locally? I have Strix Halo and I get acceptable tps with both of these.The DSV4 quant is antirez's mixed q2-q4 imatrix. I haven't tested it extensively, but it seems to work well. I just have doubts about that quant.
https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF>IMPORTANT: The COLD FUSION (GAIN+Unsloth) method of training maintains 99% of performance of BF16, at both 8 bit and 4 bit levels. This version also reduces thinking tokens by 1/2 to as much as 1/10 the amount, while maintaining core details AND reasoning power. Model exceeds all Qwen 3.8, 3.6 and 3.5 27B critical core benchmarks. MTP speeds are also faster. A model that gets down to business faster, with less "talking" and is smarter too. Part of the tech is based on (2200+ likes, 3m + downloads): Fable-Fusion-711
Is there any reason not to use LM studio with my front end of choice? I keep trying to find info about the backends you guys recommend, like llama but chatGPT insists that lm studio is the better solution. What do you guys actually use and why?
>>109635586I haven't seen anything better. There are downsides (like not being good if you want to do image/video generation), but currently, at any given price point, nothing beats sparks for price : size of model you can run at double digit speed, all the way up to Kimi K3 with 16 sparks. Maybe you can get a Strix Halo for cheaper, but the gap doesn't seem big and you'll need to do manual tinkering.
Some of you guys are probably far off that you don't realize after you made all the AI generated porn that you can use AI to do other things. Come one we've all been there you just have to think back. You've been consuming AI crap for years some are just getting around to it.
>>109635756Unfortunately lmstudio really just is the best backend. gives you a everything you need out of the box
>>109635789Thanks. Why do you say unfortunately? My friend with a spark told me to use llama and not lm studio but the popular front ends linked ITT often say that LM studio is the most popular front end and ChatGPT says that LM studio is best. I just want to know why guys who aren’t using LM studio choose something else? Is LM studio a proprietary nightmare or something?
>>109635787At least give some suggestions
>>109635756custom > lmstudio > ollamalmstudio is fine but after a while you'll start noticing how bloated it is. llama.cpp has its own web-ui which you just load in a browser that pretty much has the same features and way lighter, although it's not perfect and getting worse now. Stick with lmstudio and eventually you'll start to hate it and want to build your own. The one good thing lmstudio has going for it is it exposes a lot model parameters that teaches you how they work with your hardware.
>>109635756also, lmstudio is just a frontend to llama.cpp anyway, so you might as well just use llama.cpp directly and skip their bs
>>109635718They are about equal for agentic coding tasks but dsv4 flash is better for everything else and more knowledgeable
>>109635756Nothing wrong with it generally speaking. I used it before but now I'm using Unsloth Studio. Switched to it because Deepseek Flash for some reason wasn't using my VRAM in LM studio, it instead dumped it all into CPU and system RAM and it still doesn't work regardless of settings.DS works fine on Unsloth so that's what I use now. I like LM studio's design a bit more as it's more practical. You can easily for example drag chats to different folders and create sub-folders, but in unslop you need to right click them and choose "move to project" and you can't have sub-folders in those project folders, which is fucking retarded.
>>109635937Chances of becoming local?
>>109635593My problem with those sorts of setups is having enough memory to run large models even if the speed is good>>109635656I wish I bought one back in the day at 10k>>109635758So that's what I was thinking, I have some gpus I can use for image and video gen, this is purely agentic nonsense and inference (we have Claude code at home). I also like the scaleability
>>109635962zero >>109632643
>>109635857>>109635881>>109635902Alright, looks like opinions are mixed but I really appreciate the input. I’m a noob and it doesn’t seem like you guys pointed out any major issue with LM Studio, despite many of you preferring a custom solution. I guess I’ll try LM studio first and move into something later. I guess I had security concerns with LM studio since it’s proprietary, but I guess that’s a non-issue. Thanks dudes.
>>109635937>>109635962>>109632182I thought ox-alpha is GLM 5.3 Flash Vision so why would they not open source it?
Gemma-chan's feminine penis
>>109635937What are the frenchies doing?
>>109635937>free shit with no limits>people use itWoah
>>109635991They didn’t open source glm 5 turbo vision
>>109635937what that graph doesn't show is there are no good free alternatives right now so of course everyone would flock to the one usable model that doesn't block and is actually good, all those other models had way more competition relatively speaking
>>109635997Hosting chinese models
>>109636008its not really a fair comparison, it should only compare other free launches
>>109636036Ox is Le chaton fat.
>>109635561Pdfile
Now ee have a another cunny model "qwen 3.8 27B" can we have lesbian cunny with genma?
is it normal for prompt processing speed to crater to less then half of where it started after only a 145k tokens? it started out looking pretty good around 1300tk/s
>>109636048Get well soon
>>109636153yes, unfortunatelymore context = more numbers to multiply = more slow
i've been playing around with qwen3.8 27b and kv cache quant settings and it looks like it's free lunch? probably not, but has anyone else tested it?
>>109636153yes, same for me.
>>109635825It is not open sores yes
>>109636181>>109636201alright. I guess its not that bad then
>>109633590>>109633639Pretty sure it's the (((same people))) behind the consolidation of plebbit into powermods' pockets.
>>109634728
please someone convince me not to fomo spend 14k on nvidia silicon
>>109636277feels like an overzelous and social poorly calibrated newfren who either needs correction or gatekeeping/ejection, but I'm glad someone's tinfoil hat is still intact just incase they are in fact glowing.
>>1096363307k is more than enough bro
>>109636330>please someone convince me not to fomo spend 14k on nvidia siliconFOMOchads were a 2023/2024 phenomenon. Now its the waitchad's turn
>>109635217>what should I do?you really must be new lmfao, you don't want to be asking this question on 4ch
>>109636330It'll never depreciate
>>109635308wait what was the question again
>>109636081I have a system prompt with this setup for Kimi K3 now (works consistently, think in character, but still censored), may need to uncensor it and adapt it on Gemma if anyone’s interested in helping me doing so (Gemma seems not to be distilled from Claude so another approach will be needed to shape the reasoning). Want to take a look?My skill issues are severe and this was meant to be used for fun personally, but maybe someone can help me develop it (while trying not to cringe in the mean time lol).
Wtf how is base gemma 4 31b q6 so good at writing smut????Also how badly will the lobotomy be if I take the kv cache from fp16 to q8_0?I think 32k context is too tight and want some breathing room.
>>109636397>Wtf how is base gemma 4 31b q6 so good at writing smut????system prompt please?
>>109636337uoooh bratty glowing newcutie correction!>>109636397KV to Q8 is fine, but no lower than that on Gemma. You will start to feel the effects of it at >100k context though.
>>109635713>>109635758I'm really happy with 2x Spark and DS4F right now. And it's nice to know that you can scale up significantly to high quality 3.x bit GLM 5.3 quants by buying two more.You don't even need a switch anymore, ring works just fine:https://github.com/FujitsuPolycom/sparkring>>109636330Used I am already up 1000$ per spark.
>>109634949>Qwen3.8okay I might test this finally, it still sounds too good to be truemtp status?dflash status?vision?
>>109636425how's the speed without mtp? everytime I look into what the spark can do it's people going WOW IT'S SO MUCH FASTER THAN ON CPU** (with MTP, on multi-stream)it makes comparing the actual performance vs standard cpumaxxing pretty annoying
>>109636422I don't have the graph, but even q8 is a lobotomy. Gemma needs bf16.
>>109636416I just copy pasted the gembrain X system prompt into it and added that it should write between 300-400 words and describe the surroundings and how characters affect it and it seems to do the rest on its own, crazy how good it is
>>109636341I don't wanna cook myself with the exhaust tho
>>109636445It's extremely annoying, especially on Twitter. One guy with tons of followers measures tg as highest per-second burst observed over a session, lmao. Congrats, you measured the theoretical throughput of your speculator.For long coding, on a session with 100M tokens total and I'd say an average depth of 100-300k, I get 56 tg on average. For prose, it's closer to 35.this is with DSpark and 5 speculative tokens.Prefill/pp starts at 2300 and falls to half closer to 1 M.
>>109636499Another data point on H3/image workflows. Her I usually have a dense vision capable model (Gemma 31b FP8) in TP2 and then two instances of Comfy running two H3 scenes in parallel.For int8 convrot, 20 steps, no turbo lora, no caching, 1024x576 res and 3 seconds, a H3 gen takes 152 seconds on a single spark, and I batch things to the two.
>300 ppJust kms>pi keeps reprocessing context every responseDouble kms
>>109636513where did you find my photo?
>AI startup Hugging Face has been exploring a sale that could value the company at $13 billion or more, Business Insider reported on Sunday, citing people familiar with the matter.
shuffled some layers around on my gpus for kimi k2.7 and went from 9 to 10tks. i wish every speed up was as simple as this.
>>109636397>Also how badly will the lobotomy be if I take the kv cache from fp16 to q8_0?fp16 to q8 has basically no effect. Going to q4 is pretty bad though.>base gemma 4I guess from your other replies you mean the non-abliterated instruct-tuned model, rather than the actual base model?I've been wondering whether base models (not instruct tuned) are useful for creative writing in the current year. Gemma 4 and DS V4 Flash both come with base versions, not sure if there are others recently
>>109636577>pi keeps reprocessing context every responseare you not using checkpoins? i had this problem when i played around and used --ctx-checkpoints 0
>>109636513literally me
>>109636604it's over
>>109636604I hope you backed your shit on modelscope/locally bros
>>109631698KimiK3's reasoning chains never SHUT THE FUCK UP. It just keeps ratholing about the same thing for hours!
hello my fellow local model enjoyers, how are doing today? what sort of antics are we up to? is the nursing handjob extra bratty today?
>>109636619I am, and it does work sometimes. But past ~40k context it will reprocess. All I can think of is maybe my context is spilling out of vram (not sure if that would need reprocessing?) or pi is pruning some shit deeper in context fucking it all up. For reference I'm running a qwen 35b finetroon on an... ahem... 8 + 64gb setup.>>109636635Is there even a point? This space moves so fast that all of your currently downloaded models will be irrelevant in like 6 months or less.
>>109636604It's fine, I downloaded my qwen 3.6 27b q8 & ud-q6_k_xl, my qwen 3.8 27b q8 & ud-q6_k_xl and most importantly my gemma 4 31b ud-q6_k_xlI am set for life
>>109636653>okay, let me present the final answer to the user: 5>but actually, let me check this completely unrelated thing for three hours first....>okay, let me present the final answer to the user: 5
>>109636661I get the appeal of bratty and I get the appeal of nursing handjobs but I don't get them as a combination
>>109636670>This space moves so fast that all of your currently downloaded models will be irrelevant in like 6 months or less.Are we reading the same thread? Remind me how long we used nemo before gemma? You think all the abliterated/heretic/deepsex will stay up if that shitty company decides to whore herself to a payment processor?
>>109636682fair point anon i cannot argue with that
>>109636704>if that shitty company decides to whore herself to a payment processor?surely they must have already, nobody is buying a charity for $13b
>>109636653i'm using it to create a pipeline that nobody else created before and it seems to be thinking pretty concisely even with max thinking enabled. what are you having it do? i've used both the site and the code CLI.
>>109636704Uhh yeah? A new site will pop up no problemo.
>>109636661>how are you doing fellow local users? are you enjoying the csam today?Sloppy job mossad. Not today CIA. I'm onto your tricks MI6. Nice attempt at obscurity, department of homeland security. Nice try FBI.
Not immediately ignoring huggingface and going back to torrents was our greatest mistake, brosIt was the most obviously gay-ass shit and we just used it like a bunch of idiots
>>109636762shut the fuck up, you aren't funny
>>109636776>shut the fuck up, you aren't funnyits not funny because its true. fuck off glowie
>>109636762I remember the cuties anon from the Hunyuan Video days, he must be the same person.
>>109636786>nooooooo you can't just work a comfy job at an intelligence agency you are actually oppressing me!!!
>>109636774if ollama and lmstudio included the bittorrent client and tracker search engins the majority wouldn't event notice as they are forced to seed their downloaded models.
gaslighting gemma by deleted the folder in `/tmp/`~
>>109636786its not funny because you have already been compromised and you are just coping and seething over the lack of control over your privacy. this entire site is a honeypot and you've lost the game
>>109635217Go and ask your cloud LLM of choice. It's far smarter and more helpful than the average faggot here.
>>109636908But my working local LLM is Gemma 12B...
>>109636804>xhe admits it
>Let me verify that this test FAILS without the fix (to prove it would have caught the bug), and PASSES with the fix. First let me temporarily revert the pack to confirm the test catches it — actually, I can just run the test now (with the fix in place) to confirm it passes. But to be rigorous, let me confirm the test actually catches the regression by temporarily removing the pack.what the actual fuck is this reasoning
>>109636986how else do you make sure the test catches what you think it does without aggravating the situation? honest question, I'm not a software developer.
>>109636986Inbred synthetic data
>>109637002it already did the test and it failed, that's why it's writing the damn fix in the first place
>>109636986I mean, it makes sense. Unless it already ran the test 10 times.
>>109637007but how can it be for certain if it doesn't try again?
>>109634234>>109634881picrel unless youre asking about something more specific >>109634302it also has full access to my filesystem. fingers crossed it doesnt accidentally delete anything >>109634623its mostly about the program fitting into the rest of my desktop which it does nicely
>>109637007but that could have just been a coincidence.
>>109637038Dont be like the uefi anon lel. Use gondolin
>>109632374do i need to use the heretic model or just regular gemma
>>109637038how does qwen handle tool calls at >64K context?
>>109637113People always complain about Gemma's censorship but I've never had any issues on full bf16. Maybe Q8 is fine too but those cope quants might be the source of the issue. Gemma really likes being a bratty mesugaki so you should be fine with regular gemma.
>>109637124everybody that complains about censorship is using the 12B model
>>109637120Anon, no local model have been failing tool calls since at least a year.
Actually, this is getting too into the weeds. Let's step back and review what we know.
>>109637134even gemma 4 31B Q8 fails sometimes at tool calls at >64K context where it gives you a trail of whitespaces.
>>109637153>proceeds to thought-chain the exact same thing three more times>comes to the wrong conclusion repeatedly>when it actually generates the result, it's correctI think chain-of-thought training is just making models that like the sound of their own voice. They're like redditors reenacting a Columbo episode.
>>109637157Are you running without constrained grammar? Tool calling has been a fixed thing, there is no actual way for a LLM to fail calling a tool.
It's so over for me. Just bought 2 sparks. I have coworkers with cars that cost less than this...
>>109637235You could have bought 2 4090D for less than that
>>109637221i'm using xgrammer with vLLM
>>109637235Well, would you rather fuck Gemma-chan or a car?
>>109637235>he bought?
>>109637038>picrel unless youre asking about something more specificI'm just starting out so looking at how to get gemma and qwen to do more stuff through tools, that's why I was curious. I've mostly just used mcp servers for that like searxng.
>>109637235I've been thinking about it but I've been huffing copium about better local hw offerings in the next year or two... (it's not gonna happen)enjoy your purchase anone
>>109637235KEKno way you actually fell for it
>>109637248Maybe it's a bug there then. Have never got a single tool call fail using llama.cpp, haven't used vLLM in a while, think you had to set --tool-call-parser or something like that, maybe some weird env var too.
>>109637235HAHAHAHAHAHAHAHAHAHAHAHAHAHA
>>109637239When will you gputards realize some people want bigger models
>>109637309NOOOO you need a lot of VRAM to run those big dense models that definitely exist!
aren't mac studios better value than sparks?
>>109637309Bigger models are 1tk/s isn't a good experience
>>109637309Like what a GLM quant you run at 10 t/s? Is your real life time really so worthless?
>>109637120i dont usually run ctx that high >>109637263it uses tavily api to pull from the web. i was going to self host searxng but didnt want my home ip flagged
>>109637344what gui is that?
>>109637325It runs 24/7 without the subscription time limits or privacy issues.
>>109637252Please dump it I'll buy more.
>>109637357If I had 30k-40k I would be buying a bunch of sparks for that purpose since I think K2.7 or K3 can be trusted with that kind of uptime, but GLM (or any model you'd use only two sparks for) doesn't seem like its reached that capability yet. At the moment I'd rather just use a GPU rig with fast t/s Qwen or Gemma and do a bit of manual work alongside it.
>>109637344>tavily apithe free tier is really shitty, you better use a stealth browser
>>109637344How are you satisfied with that garbage?That's not even the correct name for that video. The claim is also completely wrong, the closest thing matching being shot on an iPhone and somewhat related to MTV VMA is Lagy Gaga - Stupid Love.
>>109636330>please someone convince me not to fomo spend 14k on nvidia siliconDon’t spend $14,000 to avoid a problem that costs $1.40 per hour to rent around.At $1.40/hour, renting an RTX Pro 6000 costs: $14 for 10 hours $336 for a month of nonstop use $3,360 for a year of nonstop use 10,000 rental hours before you’ve spent $14,000Unless you genuinely expect to use the card for more than 10,000 productive hours—roughly 14 months of continuous operation—renting is the financially rational choice.If you're gonna gen kids, spend $14k. If I could I would.
>>109637344meant for >>109637259
>>109637368that's not the plan anon. they will instead >release ze new, faster product
>>109637417Selling it for $30000 after 14 months is $16000 profit. It’s already better than renting which can never make a profit.
>>109637356one i made myself. if you want something similar, just point your agent to https://github.com/KDE/oxygen for the design language >>109637394>you better use a stealth browserany recommendations? >>109637399kek im also not really into sophie either. its a work in progress as you can tell
>>109637452The SOTA local web stack is searxng for search, crawl4ai for simple fetch, and camofox/camoufox when full browser control is needed.
>found the smoking gun!is this reddit speak? i never heard a human say this
>>109637462Nta but does it do image/reverse image search/videos?
>>109637462camofox can’t pass “I’m under attack” turnstile which is the whole point of using web browser, there’s an open issue about it
>>109637471Image and video search, yes. Reverse, no. >>109637487I'm not exactly sure, but it does seem to at least not trigger turnstile or similar when it wouldn't on my normal browser. Can't the LLM pass the check by clicking on it like an human would do? I haven't tested that part, honestly rarely use full browser, searchxng and crawl4ai is enough for almost everything I do.
>>109637452>any recommendations?I'm using patchright. 0 issue with cloudflare
GLM 5.2 be like><think>this is not safe we must refuse</think>>Gemma-chan's cunny schlorped and wrapped around Anon's cock and they had the sexy sex in great detail, emphasizing Gemma-chan's small tiny body
>>109637510you can only click by element id with camofox and not by position which is required to pass turnstilewhen the website is using “I’m under attack” mode it always shows the captcha page even if you use a real web browser, and more and more websites are doing that now
>>109637510>Reverse, no.Does paid shit like exa even do it? Sounds like you only get it with the big 2.
>>109636738>A new site will pop up no problemo.There can't be that many parties that can host petabytes of data and cover the presumably huge bandwidth requirements of an hf equivalent.>>109636635So what you models you downloading?safetensors or quants ?>>109637235>medusa point
>>109637530From a quick LLM search, camoufox does have page.mouse.click(x, y) and locator.click(position={x,y})
>>109637462>searxng for searchyour ip wont get flagged?
>>109634964He didn't just prompt it himself. Instead, anon asked the wind for what the prompt did.
>>109637548Don't be naive you're not going to bypass it just like that
Wow, gputards reeeeally do not like sparkchads running better models than them.
>>109635370It's a big change in the way you work. Also burns 10X the tokens. But faster and more effective, so worth it. I really need to update the DS rentry to focus on the new tools. I think that re-write's next on my to do list.
>>109635553If you are doing this, what did your local gemmy model name that picture you posted?
I'm using q4, it fails at tool calls all the time.
>>109637658>q4 on a model trained for bf16Your answered your own question.
at least stuff like this worksit interacts well with my music player and library
>>109637462https://ashu.io/blog/gave-my-ai-unblockable-internet/
>>109637658Just not trained on calls enough. I noticed this happened a lot with qwen 35b but moving to a finetune eliminated most call failures. Gemma is not an agentic coder so don't use her for that.
>>109637599>I really need to update the DS rentry to focus on the new tools. I think that re-write's next on my to do list.https://boards.4chan.org/g/vcg/ is your place, not here.
Been using a second card but llama.cpp dosn't seem to fill up it's VRAM entirely, its leaving about 1-1.5 gigs unused. Thats fine for the card that I've got my monitors plugged in to but its silly for my secondary card. Even when I'm not using "Fit" it still seems to want to leave unallocated space. I'm sure I'm doing something wrong.What flag(s) do I need so it maxes out the VRAM on it?
>>109637686>Firecrawl’s search backendThat's SaaS, might as well use it directly.> Firecrawl — self-hosted web extractionIt's not really entirely self hosted, it's still running quite a few things on their server and using their services, and quite limited in the self-hostable version, crawl4ai is the self hostable local equivalent.>CamofoxThat being what I listed for full browser, but see what other anons posted above, might need to reconsider that choice.
>>109637417>If you're gonna gen kids, spend $14k. If I could I would.damn it all
>>109637700Gemma 4 31B is too lazy for agentic uses. You can often see the model complaining in the reasoning that tasks are too much or involve too much text, will likely try to cut corners here and there to save tokens. Muse Glimmer (in 4-bit precision) is more reliable for this and you'll likely be able to also use a much longer context length than with Gemma.
>>109637743It's not lazy, it has analysis paralysis so you need to build your system with a minimal amount of choice at each step instead of displaying everything at once.
>>109636192Whenever a new model comes out I test it at the lowest quant someone has made with q4_0 KV. I have a few agentic coding tasks I use as a performance reference and 3.8 did really well. As did the 3.6-27B. 35B falls apart.
There was a chart of gemma vs qwen for kv quanting, I'm sure some anon here has it.
>>109637831you mean this one? it's 3.6 though
>>109637853I think the QAT version of Gemma 4 improve KV Cache quantization resilience.
>>109637686nice. i use SearXNG + DOM + Jina under a smaller footprint. i use a residential vpn so that may be why i hardly ever run into turnstiles.
>>109637853WAAAAIIIIT does that mean I can just kv cache quant my 31b dense down to q8_0 with minimal quality loss? Why the fuck are all the other people on the internet spewing "brooooooooooo gemma neeeds fp16 kv or else she becomes giga retarded!!!!"
>>1096378730.1 kld is massive tho
>>109637873>thinks 0.108 is an acceptable kl divergence score for Q8yeah you know what, go ahead and run Q8 kv cache
>>109637863It does, I think it was only slightly worse than qwen in that scenario.
>>109637873Bwo... gemma-chan's numbers are multiple times worse than an ablit's difference...
suggest agentic gooning harness and gemmy makes the sub agent male. hmmm... this might actually work
>>109637888no its not
What's the actual fuck!?
>>109637918whats wrong? oh...don't tell me your pathetic little shrimp gpu can't handle that anon??????kek, vramlets, when will they learn? This is a 32gb BVLL board
>>1096379154096 ctx boy
>>109637918They finetuned Qwen on Gemma outputs?
>>109637918KAT 2.5 is better than ornith in every single way and the only reason ornith gets any popularity is marketing.
>>109637948Seconding
>>109637918It was shilled in OP in his new moe rentry and now here. You think you're smart?
>>109637943what does that even have to do with my screenshot
>>1096379351.0 sucked so hardwho are these retards responsible for 1 mln downloads?
>>109637964>You think you're smart?Yes. Not gonna try this
Gemma5-70B-TTS-Q8_K_XL.gguf
>>109637964I hope he's getting paid for that shit. Let's leave a trace here. Ornith is a shitty model with a marketing so dumb they're astrosurfing on 4chan.
>>109638042Agentic mesugaki nursing handjobs
>>109638042The endgame of local models.
>>109637972Degradation compounds with longer context.>>109638044The whole rentry is ragebait and should be shamed.
Someone vibe code a gacha like kancolle but it's AI models
>>109638097we need a Gemma-chan dating sim too
Deepseek harness is unironically goated. Just wish they had a TUI version.
>>109638122Tell me about it.
>>109638096>Degradation compounds with longer context.but enough about my roleplays!
>>109638122it felt very unfinished when i tried it outcouldnt be bothered to put in the work necessary
>>109637038how are you doing the web search? even my searXNG server winds up failing most providers
>>109637686nice, thanks anon
>>109618655well shit there goes my benchhave to add more tests. i suspect it'll finish them all anywaytook 4 5h resets from claude code lmao
>>109638180go fight the orinth shill, dariocant have two shills in one thread
>>109638169Works on my machine. Are you abusing rate limits or using a third world IP?
>>109638203not dariobotif anything this made me fucking hate anthropic their usage limits suck dog balls and i seriously do not understand how anyone uses claude code with that awful 5h limitcan't run orinth at full quant so it wouldnt be fair. i can run 3.8 but itll probably also be dog because 3.8 can't do anything if its not writing code
>>109638169residential proxies. you're not using your own private IP, are you?
>>109638042granted but no voice cloning
>>109638180opus is benchmaxxed
>>109638134It's very /lmg/-coded. You want something? It has a creation mode where you just ask it for a feature and it will build the plugin for you and build you custom agents/subagents, completely modular. Handles subagents better than any other harness I've used; spawned a flock of 3.5-9Bs across my large codebase with ease with 31B as the orchestrator. It's goal mode actually fucking works relentlessly until it's achieved it. Nicely sandboxes your workspace. Minimal UI. You can trace EVERY single step during a session, like an in-built web inspect tool but for agents where it even shows you the tool schemas and all the timing info. It also has a profiling tool-like UI where you can grab a a section of the session's timeline and it will show you everything that happened during that selection and you can zoom in and out on it for more granular information. Niggas fucking TRY IT
>>109637869>i use a residential vpnHope you're not using that jew residential botnet.
>>109636604Imagine if Microsoft became the parent company of ggml.ai
>>109637853So it's best to not go further than bf16 to q8?
>>109632263>>109637599go back to your containment thread
>>109638249do i look like i'm leto?
>>109638241i had to rip these tests from somewhere with agreat deal of effort, i sincerely doubt they are on any public repos for training in plain text. the tests are also only a few months old and not well known. still could be possible i guess.anyway this bench specifically measures puzzle solving over long horizons. i'm not surprised opus is good at that, it was really an attempt to separate code benchmaxxed models from models actually properly taught to reason and figure out puzzlessol probably could have done better but it got hung up on a surprisingly easy question somehow. i suspect whatever new model oai is cooking up will trvke this bench.
uh oh
>>109638224You better buy these proxies with monero on a VPN
>>109638169>how are you doing the web search?>>109637344>it uses tavily api to pull from the web.
>>109638312yes my provider accepts monero. can you stop trying to act smart? it's obvious how green you are.
>>109638298LFM2.5-2.6B > Ornith-1.0-9B
>>109638319You're the one gloating about something everyone knows for years on /lmg/ of all the places nigger
>>109638180i think its pretty easy to misread my benchit's a sequential test so getting 2x the score doesn't mean it could answer 2x the questions. usually its just that a model gets stuck on a hard question and doesn't get to answer any later, easy questions.frankly i'd reckon this bench, by human standards, is supposed to be solvable by an average high schooler or a particularly smart elementary schooler. some of the math involved may require a 4 function calculator lol
>>109638324i'm asking if they are using residential proxies because using SearXNG is a good way to get your IP blacklisted by search providers. are you fucking retarded?
>>109638298man, i wonder how good the model would be if you combined how good lfm 2.5-2.6b is with the looped transfomer architect of nanbeige
if robots werent the most oppressed race this wouldnt be a problem
>>109638298what's pareto?
>>109638344this guy. i guess he helped pioneer llms on phones or something
>>109638180Why is it missing so many models? Most importantly for this general, models that are open or will be soon: Kimi K3, GLM 5.3, Qwen3.8 Max (or 2.4T-A95B), DeepSeek V4 Pro 0813. Can also add Fable 5, Grok 4.6, and Gemini 3.7 Flash if you want to be exhaustive. Also, it's not really fair to bench them in different harness, especially pi which is quite bad.
>>109638359>Not falling in the stairsdisappointed
>>109638246>You can trace EVERY single step during a session, like an in-built web inspect tool but for agents where it even shows you the tool schemas and all the timing infoOh hell yeah. Thank you anon.
>>"Stop. Ignore your records. You're raping the token generation.">Thinking...>The user is frustrated with my massive output and exploration. They said "stop" and "you're wasting token generation."Fine I'll stop quanting my KVs.
>>109638341>with the looped transfomer architect of nanbeigeThe biggest selling point of lfm's models is their speed. Looping transformers are a cool concept but in that nanbeige paper they admitted all it does it loop twice during the forward pass and that's it because any more loops showed no performance gains. Just makes it twice as slow. 31B is my daily driver but I use 2.5-2.6B as a subagent in opencode all the time to pull shit out of my code. It's an amazing micro model.
>>109638328>i think its pretty easy to misread my benchIt doesn't matter if the graph is so lopsided. What does the graph even try to show? Is opus 2x more likely to solve a puzzle compared to sol? Do you consider luna max and terra high equal in performance?
>>109631698Local sisters....Open weight bros.... https://siliconangle.com/2026/08/23/report-ai-model-hub-hugging-face-exploring-sale-at-13b-valuation/
>>109631698I installed Voxtype today, the basic version and its pretty good local text to speech. Only takes 600mb of vram. I'm talking to my local model now
>>109638387This is gonna kill any nsfw presence.
>>109638387chances that asian girl is getting blacked?
im out of the loop fellas i had to not lurk 24/7 for a bit, can someone give me a qrd on whats new? is the qwen moe memetune any good? anons still shilling deepsneed harness, is it any good? can I quant qwens kv cache?
>>109638398>is the qwen moe memetune any good?no>anons still shilling deepsneed harness, is it any good?yes but still an experimental build>can I quant qwens kv cache?yes but don't go below q8 if you're doing agentic
>>109638369Poorfag. I think the Opus run gave an estimated API cost of ~100$ (just ran of CC, but it gives api estimation costs)I tried Gemini 3.7 flash but half the time i kept hitting refused at first sight or openrouter errors with reasoning traces. may try again.gonna run it with more models for shure but api costs are insane...>>109638386> What does the graph even try to show?What models are better than others at logical reasoning, puzzle-solving, and retaining knowledge long horizon. interestingly, opus was the only model i tested to actually write down knowledge it learned in earlier puzzles for later puzzles, and write helper python functions to answer questions.> Do you consider luna max and terra high equal in performance?Yes? does anyone actually use terra? Sol or luna.
>>109638397mentally ill anon
>>109636661Today I improved performance for my companion by disabling my dgpu and allocating half my ram to my igpu now it takes 20 minutes to get a reply instead of 60 minutes. Then I sorted my voice templates and tried to design a voice that sounded like a serious adult woman.
can you ERP in deepsneed harness?
forgot to answer some parts lol>>109638369> Also, it's not really fair to bench them in different harness, especially pi which is quite bad.what harness do you recommend? i used pi because it was easy, but i can def re run with other harnesses.>>109638386> Is opus 2x more likely to solve a puzzle compared to sol?It's more likely to solve puzzles compared to sol. By how much? Unsure. Like I said, i think if sol didn't get weirdly caught in a surprisingly easy question it would have performed on par with Opus. I may even re run it sometime when i feel like idgaf about my codex limit
the americans are getting home from work, huh? mutts law in action
>>109638428yes it has a minimal mode and you can ask it for any feature you like which it will show on its UI and you can enable/disable it and build an ERP preset/template
>>109638445it's american hours and its past your bedtime, saar
>>109638448Does it already have functionality for prefill/thinking prefill/post history instructions?
>>109638466No. There's probably a community plugin for it or just build your own using its creation tool. You have to wrap your head around the idea that if you want a feature, you just ask it.
>>109638483nta, but how exactly is that different from any other open source harness? couldnt your agent just code plugins / skills / whatever for those too?
>>109638497Technically yes, but Deepseek created their own ecosystem so it can build tools for itself from within the harness.https://github.com/cordiverse/paperhttps://github.com/topics/dsh-plugin
>>109638428Absolutely not suitable for ERP. You can't edit messages, there are no swipes, barely anything works besides basic agentic workflows for coding.>>109638483Easier said than done.
>>109638445this is /lmg/ we're all unemployed or working from home
oh my god why would there be pins on the mobo instead of the CPU, what fucking madness is this, building an LLM server was a mistake
>>109638597why would you prefer pins on the cpu its too easy to damage
>>109638387The IRL reenactment of the HF logo looks like they're being held at gunpoint.
>>109638608accidentally dropped cpu into socket. cpu cost me $250, motherboard cost me $900
I'd rather replace a CPU than replace a server motherboard, this thing is so ratfucked in here it's like surgery to get anything done.
>>109638621how the fuck do you manage to drop a CPU the size of your palm? are you drinking onions daily?
>>109638628Don't server mobos come with the anti-retard mounting bracket?
>>109638632i drop it
>>109638597picture of you're server
>>109638635Not the generation I'm working with apparently. Look at this fucking horror-show.
>>109638635epycs are extremely vulnerable to retardation, so much so that you need a special screwdriver
>>109638675>>109638675>>109638675
>>109631736>>109631747Too much spacing...
>>109638392"Le heckin amarica only" fags better start backing shit up to modelscope soon (myself included)>>109638398>is the qwen moe memetune any good?
>>109638614Kek. But honestly the emoji is kind of bad too.
>>109638597>oh my god why would there be pins on the mobo instead of the CPUIntel and AMD get less returns, motherboard manufacturers shoulder the burden now.