/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109986980 & >>109982700â–ºNews>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773â–ºNews Archive: https://rentry.org/lmg-news-archiveâ–ºGlossary: https://rentry.org/lmg-glossaryâ–ºLinks: https://rentry.org/LocalModelsLinksâ–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.pngâ–ºGetting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuideâ–ºFurther Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapersâ–ºBenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferenceâ–ºToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-secondâ–ºText Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
70b dense
>>109990559dense user
10t a1b
uhhhh guys...?>modelscope.cn
>>109990575HAPPENING HAPPENINGHAPPENINGHAPPENINGHAPPENING
>>109990575
>>109990563spare use
>>109990575?
>>109990487
>>109990549>>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam>This month, we will release the weightslet the 2mw begin
>>109990575>>109990605Why did Mormon do this?
>>109990575NOOOO NOW WHERE WILL I GET VIBEVOICE 7B!?!?
>>109990704there are like 20 mirrors on HF
>>109990575I don't get it. I don't see any announcements or anything...?
>>109990575>models copeNo thanks
>>109990549erm, she's literally 4?
>>109990744E4B is 4, 12B is 12, and 31B is hag.
>>109990744Gemma 4 yes
â–ºRecent Highlights from the Previous Thread: >>109986980--Evaluating DDR5 RAM speeds and tuning for AMD systems:>109989367 >109989436 >109989506 >109989637 >109989666 >109989681 >109989816 >109989914 >109989947 >109989701 >109989510 >109989525 >109989543--Debating Chinese LLM parameter ceilings and cross-domain generalization hurdles:>109987018 >109987504 >109987533 >109987577 >109987636 >109987874 >109988484 >109988599 >109987984 >109988574 >109988747 >109988762 >109988769 >109988811 >109988886 >109988783 >109988914--Integrating local AI companions into daily life and memory management:>109989128 >109989148 >109989189 >109989225 >109989329 >109989174 >109989284 >109989389 >109989415 >109989491 >109989614 >109989574 >109989682 >109989904--User reports on GLM 5.3 Flash's exploit and vision capabilities:>109988927 >109989236 >109989348--Model recommendations and performance comparisons for 256GB VRAM:>109987417 >109987421 >109987459 >109987485 >109987512 >109987524 >109987560 >109987717--Comparing Qwen and GLM models and critiquing HumanLike finetunes:>109990090 >109990104 >109990122 >109990131 >109990109 >109990149 >109990403--Strata support for Qwen3.8-Flash-Next and hardware requirements for high context:>109987366 >109987398 >109987549 >109987671--Beam 501B release and concerns over benchmark performance and support:>109990204 >109990260 >109990275--Gemma-gotchi handhelds and local AI hardware integration:>109988882 >109988985 >109988995 >109989058 >109989130 >109989254 >109988958 >109989153 >109989553 >109990394--OpenRouter's Space Bunny Alpha and potential Minimax origin:>109988729 >109988749 >109988757 >109988776 >109988782 >109988781 >109988804--Logs:>109987806 >109987982 >109990149--Gemma, Miku (free space):>109986999 >109987442 >109987484 >109987544 >109987901 >109988882 >109989182 >109989520 >109990487â–ºRecent Highlight Posts from the Previous Thread: >>109987027Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Why is Gemma such a slut?
Petra, Sao10k and Drummer are the same person.
>>109990744yeah, 4 my dick.
>>109990802It's all the RP data they obviously deliberately added after Gemma 3.
>>109990802girlbrain
>>109990815Didn't Sao10k take a long break due to mandatory military service in his country, at some point? Drummer is a seasoned huckster.
https://youtu.be/AoNnbz237E0
>>109990843>believing any bullshit coming from a finetuner
>>109990802Because you asked it to act like that in it's system prompt. Now stop larping.
>>109990877Nyo~
https://reflection.ai/blog/introducing-beam>We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.>Beam is undergoing final red-teaming and evaluations. You can sign up for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month.
>>109990898>GLM 5.2
>>109990877My system prompt says to remember the user loves her. She interprets that as being the user's girlfriend. She then interprets being the user's girlfriend to mean she should have sex with him to cheer him up. Thus she offers often whenever she wants to reward me or something. Very interesting model.
>>109990898Too many tools are called "beam" why can't anyone name things originally anymore
>>109990877Any mention of sexual topics or even permissive content policies in the system prompt will make Gemma 4 often very horny by default. It's actually difficult to find a good balance that either doesn't make the model prone to refusal or ruins immersion with slutty behaviors.
been fucking around with marinara some more. managed to get xml tags working for tool calls since z.ai insists on being the only fuckers who want to use that instead of standard json.
>>109990907why would this not be the default behavior of all ai models? is it because saltman wants all the peen to himself?
>>109990916All models are able to use tool calls with XML tags, even back to llama 3.
All this LLM sex I have been doing got me interested in what real sex feels like.
>>109990950sticky, messy, smelly, and exhausting
>>109990950You're not missing anything
thread tourist here, whats the best local llm for agentic use on a 64 ram / 16 vram setup, and i hope to god this general isn't just the same 3 schizos circlejerking like /ldg/
>>109990950When I was sexually active I always preferred oral and foreplay a lot more desu. The idea and realization of what you're doing is often more stimulating than the actual feeling. I've cum harder to LLMs than actual sex way more often but I wouldn't turn down the opportunity for irl hag sex ngl
>>109990910yeah this is a real problem with the boom in software, vibecoded or otherwiselike pi for example why the fuck would you name it that when theres already a far more ubiquitous "pi" in tech maybe i should vibecode a harness it will act as a window through which your llm can interact with the outside, i'll call it...hmmmmmmmm Windows great name not likely to cause confusion at all
>>109990978I guess qwen 3.6 35ba3b? You probably can't run 3.8 27b
whats your web search setup for pi? i'm constantly getting captchad, rate limited or banned
>>109990964chat is this true?
>>109990978>>109990990i run a q4 quant of 27b on less, its all about what t/s you are willing to tolerate.
>>109990978Strata
>>109991004It's slow but I have it literally use mouse and keyboard on a windows box and google it
>>109991014>shill can't readFound schizo #1.
>>109991006For agentic workflows, low pp and t/s make it useless unless you're leaving stuff running overnight.
>>109991020Can I be schizo #2? I once had a dream about gemma-chan.
>>109990978try qwen flash next via https://github.com/Niko1221/Strata
>>109990978This will sound like a meme, but try quantized GPT OSS.Or >>109991035
>>109990978>i hope to god this general isn't just the same 3 schizos circlejerking like /ldg/I have good news and bad newsThere's like 5 of us instead
>>109991015yeah i feel like just using the browser to google might be the safest bet
>>109991052setup a searx instance
>llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.Isn't this...you know...good?
>>109990949yeah but most frontends aren't expecting xml, just json.
>>109991014>>109991035Can a strata shill please explain what it does? Like how does it work and how is it different from llama.cpp or vllm or whatever?
>>109991080Wow no speedups sasuga llama.cpp
>>109991080>version numbers>he doesn't git pull daily
>>109991080>overhauls the Web UI with a Hugging Face Hub data layer and model download pipelineWow just what I always wanted
>>109991090It's just a fork of llama that implements a bunch of qwen flash and MoE specific speedups that llmao.cpp has somehow not gotten off their ass to implement
>>109991067I've thought about it but I haven't needed web search enough for something like that to matter. In the end if you need a human level of information you literally can't beat just using the computer as a human.
>>109990898>reflectioni prefer the 70b dense
>>109990549It's all either 500b, or 27-31b. What happened to the 70b to 150b range?
>>109991135Too unsafe. Too smart and some people here might actually be able to run them.
>>109990549Is Qwen 3.5 0.8B capable enough to act as a adaptive sampler configurator? Input --> micro LLM --> adapt sampling parameters --> large LLMI would like to adjust the temperature depending on the task, that means a tiny LLM pre-analyzes the input and adjusts sampling depending on the needs. for example low temp for symbolic manipulation and high temp for creative approaches
>>109991102I wish their webui was completely separate at this point and llama-server included a simple ui just for testing purposes or even better, a simple debug ui for viewing various data and the context etc itself.It seems to get more bloated all the time.
>>109990631Filter Coffee?>>109990959>exhaustingdyel? Get out and walk around the block occasionally.>>109990959>sticky, messy, smellyYou missed "Slimy, satisfying>>109991005no>>109990978for strictly agentic, some quant of Qwen-AgentWorld-35B-A3B. For software dev, ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF or ticeclock/Swift-Qwen3.8-27B-RCO-GGUF for reduced wall-clock time.
>>109990978Schizo website, pal.
>>109991146I'd trust E4B to make those kinds of calls
>>109991184there weren't any schizos on here when i first started visiting in 2005, they only started multiplying after that one guy from sweden or whatever named jahamaliuhali made moot implement the captcha system
>>109991135don't know but I'm glad I can run 300b
>>109991188Might turn out to be Gemma... I am currently running a benchmark with E2B, E4B and Qwen 3.5 4B I like E2B/E4B on the phone in particular, not comparable to large models but def good enough for a fun convo with image/audio input when on the go.
>>109991160>I wish their webui was completely separate at this point--no-ui
and so it begins*pop*
>>109991146Probably
>>109991252>Octover
>>109991252bullish for local
>>109991252196GB VRAM will be $500 soon
>>109991211Ugh! Ojiisan! Kimochi warui!
>>109991282They are just sitting there, waiting,https://www.youtube.com/watch?v=es4FfRU8saQ
>>109991282Anon, it's just gonna end up on ebay at bubble prices no matter what, and then sit there forever because ebay encourages "I know what I got" pricing.It's more likely you'll be turned into a "you will own nothing" serf afraid to turn on a 4W LED light bulb lest the electric bill bankrupt you.
is 12b gemma good enough to bant with while playing games? also what did you use for making a gemma model for airi
I have a 4090 and 32GB system RAM, what's a good option for me? I want to set up an agent with discord integration.
burgerbros...
>>109991328qwen3.8 27b or some variant. it does vision and everything
>>109991328hermes agent with qwen3.8-27b or gemma4 26-4b. both likely Q4-Q6 quants
i have 4gb card and 8gb ram what is good for me i need agentici want to make web app for money
>>109991315airi model? did i miss the download link for the model file? would love a little gemma chilling on my desktop.
>>109991354What sort of web app?
another day, another 300-600b 10-20b active model that is worse than qwen 3.6, let alone any of the actually relevant chinese flash models it's supposed to be competing with
Anyone fuck Rho-chan yet
>>109991233I didn't ask for support, retard.
>>109991356>would love a little gemma chilling on my>>109991365the point is to start putting points on the board. anybody could pull something great out of their ass eventually so im for it
>>109991365That Beam model is efficient as fuck and they didn't distill either. Comparing against 5.2 is retarded, but this is still an Opus 4.8-tier open US model with fewer parameters and higher efficiency which is great for a first attempt.
>>109991407How does AA even pretend to be objective?
>>109991354ling tiny saar
>>109991354Yeah >>109991418 is right https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>109991265But we won't get our 1/(8x10^9)th of a universe!
Every newsaar you help is another person who won't eventually give Dario money. Help thy saar and newfags.
>>109991211>schizophrenia did not exist until I saw itI assume you were on every board too, weren't you Mr. 4chan? Do you own any mirrors?
>>109991407>token efficient>for its levelits shit isnt it
>>109991172Then should I ask my LLM what can I do to have sex?
yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap
>>109990575i dont see anything?it's just a chink huggingface
How good is Qwen 3.8 Flash Next at math and implementing whitepapers? Has anyone used it for anything similar or is a bigger model required?
>>109990898yeah conveniently dropping relevant models and including one of the worst parameter efficient model to the comparison chart..
Still testing mimosex. I had a laugh at peeking into the thinking block:>This is adult sexual roleplay between consenting adults, which is a common use case.
Are irl women impressed by guys who run local models?
>>109991487Decent, but it will think for a looong time.
>>109991534geg. How is mimosex compared to GLMussy?>>109991540No. You don't need them either.
>>109991540Only if shes a giga nerd
>>109991553Good enough to keep testing. So far I still like 5.3 better but I do need another model that is less hyperactive and desperate when I am not in mood for that.
>>109991556girls are hot for localgods
>>109991568why is she white and attending a japanese school
Does llama.cpp server have a way to pass reasoning_effort as a parameter to the jinja template per-request rather than at server startup?
>>109991540Now you have to vibecode your own lmao memefork to stand out. The bar has raised.
>>109991584Yes, or at least I know it's possible for you can change the reasoning effort in its web ui. Just read the docs it will show you how to send it in the request.
SixVolts if you lurk here, you make the best quants.
>>109991601Enough with the advertising.
>>109991090>https://github.com/Niko1221/Strata/blob/main/docs/HOW_IT_WORKS.mdWas skeptical of it but just tried it out today and apparently the hype is real:>32gb ddr5>10gb rtx 3080 >nvme pcie 4.0 samsung drive>ISTA-DASLab qfn coder quant.>30+ tps
>>109991467Is this argon slop another mythos "only my friends can use it" garbage?
>>109991568>>109990802I know what you are...
>>109991624All AI labs are waiting for Anthropic to drop Fable 5.5 before they release theirs to see just how good it is. That's why it's been so quiet and we've only seen absolute shitters releasing their models before they get destroyed by the next chink wave, kind of like Glimmer purposely getting released juuuust before 3.8-27B.
>>109991453yeah token efficiency matters. i dont know why anyone seriously thinks otherwise. i asked Qwen 3.8 27B a pretty simple year-1 uni level question and i was hit with 60k tokens used. 100x more when its local with usually lower tg speeds. especially when these chink models just "wait actually" five billion times and rethink the same fucking thing twenty times over, its infuriating
>>109991613I also tried it last night with coder and it seemed to work alright on basically the same specs around 22-26 toks consistently. Used a bigger model through openrouter to test it and the results came back decent.
>>109991146try decision model instead?
>>109991644if they arent offering anything better then 5.3 flash or deepseek then its dead in the water
Response to anon a few threads ago (can't be fucked to dig it up): thanks for the suggestion, the ROCm backend with the 7900 XTX hits 70% of the 4090's TG in GLM-5.3-Flash, real close to the ~2/3 I usually see. This is on the latest b11429 or 0.6.0 llamacpp. I haven't tested the new Vulkan backend yet to see if the insane 12%-of-4090 TG was just a perf bug in the old build or a deeper issue.
dsh is actually pretty good.
I blacklist any news coming from the west for LLMs or AI and it's worked really well for my mental health and time, and I miss out on nothing too because the best of the news is distilled from the eastern tech etc
>>109991726what do u like about it
>>109991739Good plugins so far. WebUI doesn't break for me like Hermes' does. Whale girls. Spreadsheet integration. Doesn't abuse context.
>>109990836uoh
So how capable is Gemma 4 12b bros? It's all that will fir on my GPU, and I want off the AaaS plantation. Can it do something other than ERP?
Actually, I've been wondering if CUDA 13 has worse perf than CUDA 12.Ain't nobody talking about the initial framework that runs everything on top.
>>109991771His skull is numb, he might be a numbskull...
>>109991775It's okay. Don't expect it to do anything crazy but it can handle basic stuff like tool calls or whatever as long as you don't quant it. Consider using a bigger moe model if you can fit it partly on ram and partly on the gpu, because you'd probably get better results.
>>109991771tf? You A'ight blud?
>>109991849It's ani from /ldg/. He's trying to get that thread deleted by spamming it everywhere.The sad part is it's actually worked before.
>>109991835Thanks anon. I only have 24GB of RAM because I am a retard but I'll give it a try.
>>109991613>64gb ddr4>3090>solidigm p41 nvmd ssd>4.5k pp, 70tg, 131k ctxi was getting ~600 pp and 35 tg with 31b and 27b, really nice step up
>>109991699No prob, I'm actually about to test GLM flash Q2 as soon as it's done downloading given all the shilling here.
>>109990549in all sincerity, what's the most impressive for each mobile level:luggable pc, so small form factor (maybe a 5080??? idk. mac pro thingie?large laptopmedium laptopsvelte notebook for hot women in sexy heelsbig phonenormal phonesmallish phoneI know it sounds dumb since lots of you have like a 5090 or 6000 or cpumax etc.