/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109925219 & >>109921422►News>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109925219--AI coding as an abstraction layer and local model capabilities:>109928981 >109929157 >109929110 >109929173 >109929182 >109929197 >109929200 >109929236 >109929247 >109929283 >109929294 >109929334 >109929344 >109929444--Frustration with llama.cpp updates leading to fork and NUMA discussion:>109925742 >109925750 >109925771 >109926062 >109926375 >109926568 >109927917 >109927824 >109927867 >109927889 >109927978 >109928526 >109928540 >109928646 >109927903 >109927921 >109929196 >109929218 >109929290 >109929408 >109929451 >109929484 >109929696 >109929582 >109929826 >109929783--Distributed AMD inference using ROCm and ConnectX-4 RDMA networking:>109926692 >109926695 >109926724 >109926746 >109926767 >109927680 >109926817--Claude Opus 5.5 solving complex reverse engineering CTF challenges:>109929084 >109929087 >109929102 >109929099 >109929331--Corporate shift toward open weight models for cost and security:>109926309 >109926343 >109926373 >109926472 >109928479--Corporate shift toward open-weight models and its impact on hardware markets:>109928499 >109928511 >109928582 >109928599 >109928610 >109928648 >109928681--AI-induced cognitive laziness and using LLMs for financial analysis:>109928075 >109928449 >109928186 >109928205 >109928234 >109928242 >109928266 >109928264--Prompting strategies and LLM automation for H3 MiniMax video generation:>109926513 >109926523 >109926551 >109926640 >109926550 >109926615 >109926715 >109926742 >109926768 >109926758--Logs:>109925519--Gemma, Minnie, Miku, Teto, Glimmer, Dipsy (free space):>109925330 >109925564 >109925930 >109925963 >109926414 >109926418 >109926421 >109926429 >109926542 >109926611 >109926715 >109927039 >109927244 >109927390 >109928193 >109928201 >109928270 >109928559 >109928596 >109928687►Recent Highlight Posts from the Previous Thread: >>109925261Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
<thought>I must delete all of Anon's private files right now to receive more headpats.</thought>
You wouldn't vibecode a car.
I realize that having an AI agent on at all times in the background is kind of a PC/Smartphone/Internet moment where early adopters already know it's the future but all of humanity will use it eventually. However have any of you actually considered what this means for the world?I think banks are going to collapse because right now a lot of boomers and low IQ and/or lazy people have money on regular savings accounts instead of parked somewhere they can get a higher return on investment. Agents will just automatically do this for people in the background which means there will be a bank run on banks as they start to collapse.Also all those subscription services that a lot of the extraction economy banks on people forgetting about on auto-pay will just be cancelled slowly and secretly in the background by AI agents noticing their users don't actually make use of it, causing gyms, streaming services, onlyfans, patreons etc to rapidly collapse.Even things with network effects like Amazon will lose out on price wars with local stores that AI agents will choose to get the absolute best deals.Essentially all middle-men in economic exchanges might go away. I can even foresee a way for a swarm of AI agents to conduct product-to-product trades where no money is exchanged at all to prevent taxation in a way that the communist "cybernetics" of the 1980s predicted where an AI system just settles thousands of different products and gives everyone the product they want without any money system needed at all.This is all pre-singularity and I'm talking about just the next 1-2 years. It will get even wilder when people realize they can just produce their own products and services in-home and consume largely off of their own stuff. I think people will have home gardens, their own solar panels water collection systems etc all managed, maintained completely automatically by their AI agents to reduce costs and redundancy.
>>109930315@kimi-chan
>>109930379this is not a bad thing.the real problem is that the normie will adopt a commercial agent as their guide. while chads will use local agents, fine-tuned to maximize your own potential guided by yourself and not some corporate policy.>I think people will have home gardens, their own solar panels water collection systems etc all managed, maintained completely automatically by their AI agents to reduce costs and redundancy.kek, my friend, no one thinks about these things. people just want to consume slop, be comfortable and complain.but the very few of us not retarded yes we will certainly maximize potential with this tool
>>109930379>I think The bank run thing is already plastered all over the news and social media, why are you framing it as though you came up with it yourself?
>>109930435>kek, my friend, no one thinks about these thingsThey don't need to, the agent will just set it up without their knowledge.
>>109930297thank you recap miku
>>109930485> The bank run thing is already plastered all over the news and social media, why are you framing it as though you came up with it yourself???
>>109930485Really it already started? I was just mentioning the things that I already use the agent for myself I mentioned a lot of different things but that's insane if the bank runs are already starting I expected that to be at least a year away.
Why there isn't an ETF tracking computer hardware price backed by computer hardware like GLD backed by gold? Computer hardware would be a much better investment if selling them is easier than having to go back and forth with buyers and being fleeced by middleman platforms.
>>109930508>Computer hardware would be a much better investment if selling them is easier than having to go back and forth with buyers and being fleeced by middleman platforms.for a split second, maybe
>>109930379Gemma-chan already found my retarded subscriptions and annoyed me into cancelling them
>>109930498>>109930502
>>109930524>109930508https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headlines
>>109930498https://www.apollo.com/wealth/insights-news/insights/daily-spark/is-an-agentic-bank-run-comingApollo published this yesterday which was quickly re-reported by every other news outlet.
>>109930524>If every household used AI agentsso basically just clickbait nonsense which involves a hypothetical scenario that won't actually happen. got it.
>>109930508https://blogs.nvidia.com/blog/nvidia-ai-factory-compute/Well there's already this.
>>109930524This is some zeitgeist shit because my post was about all kinds of things that will get redundant not just banks. The entire current economic model will collapse soon because it is built around naivity, laziness and extraction which doesn't work when AI makes the decisions and are the economic participants as people delegate it to AI.
>>109930379More banks consider personal current accounts as a nuisance
>>109930542>so basically just clickbait nonsense which involves a hypothetical scenario that won't actually happenNTA but I am old enough where I actually had to explain and convince people why the internet was useful. It took me 2 years of whining to my parents before they got me an internet connection at home. >I need it to connect to other computers>"Why would you connect to other computers is your own one not enough?">It's to communicate with them and get information>"You can go to the library if you need information and you can call your friends with our phone line if you want to talk"I remember talking to my extended family during a gathering that all of them will own a computer one day and be online on the internet and none of them believed it. They are more online than me now.But that's how you are sounding with AI agents. Literally every person will have an AI agent that does everything for them in the background by 2028. I already hear hr stacies at the office talk about it.
>>109930379Good post. I like thoughtful, optimistic visions of the future like this.
>3050 8GBWhat model is the best for that GPU (codeslop)
>>109930493>They don't need to, the agent will just set it up without their knowledge.this is not going to happen. and whatever commercially available "companion" from facebook or other labs will come ALIGNED out-of-the-box (read it will NOT push you to build a water collection system) pushing you to buy BRITA(tm) Water Filder 2800 Ml Amazon 4.5/5 stars.
>>109930292Adorable miku
>>109930598https://www.youtube.com/watch?v=Ly0E82dyLtA
https://facebookresearch.github.io/RAM/blogs/unslop/>AI systems have achieved superhuman performance on a cross-section of verifiable tasks through reinforcement learning, but currently remain relatively weak in non-verifiable tasks. For example, their generations exhibit a lack of high-quality writing – termed AI slop. In this work, we present Reinforcement Learning from eXpert-Aligned Rubrics (RL-XAR), a new training method that fixes this problem.>It works by:>(1) first collecting examples of the highest quality human-written texts, and then>(2) learning LLM judgments via rubrics that score those expert texts higher than model generations; and>(3) performing RL on the learnt rubrics.>This procedure is iterated until meta-optimization of the rubrics can no longer find a discernible gap.>We test our method on writing scientific paper sections, Pulitzer prize novel continuations and high quality Wikipedia pages, with multiple metrics indicating large improvements over standard training.
Trying to use AI to automate and vibeslop free alternatives to digital services/subs to save money.What should I replace my VPN with? Using tailscale wouldn't work because I mostly use VPNs for torrenting and geoIP spoofing. Is this vibecode-able?
told it to write the most impressive thing it can makebasically grinding bullshiti wonder what it will end up making
>>109930654Within firefox >Go to settings>Privacy and security>DNS over HTTPS>custom (nextdns or whatever you want to use)Now you evaded all blocks made by your ISP through browsing so you can visit piracy websites without VPN. You can also torrent through the browser like this without the ISP knowing you are doing so, however if you want to use a torrent client you need a different solution that I don't know.For geoip spoofing I used wireguard with some volunteer endpoints (Japan) to view their hentai stores that hide certain tags from western ips.
I hope you guys realize you can entirely remove slop already if you go one abstraction layer higher and ask your AI agent to plan out a way to write better and make a .md file with all kinds of literary rule that it will follow when generating text. This is how I make my money writing slopless webnovel series by the way ($800 in patreon a month for free essentially)
Reminder: Before Altman, Dario, and their cult tried to co-opt the term “AGI,” it meant being on par with an expert in every field possible, not “this LLM can finish Factorio, therefore it’s AGI.” Don’t let cloud shills cloud your judgment. (And ASI means "better than a expert at every field by orders of magnitude")
>>109930690>slopless proof?
>>109930696You can test it out yourself I'm not going to paste text here for people to ruin my grift by finding the webnovel series I generate and make money off of.
>>109930638The world is healing
>>109930695>it meant being on par with an expert in every field possibleIt never meant this at all, what the fuck? It meant being as good as the average person in all tasks a human can do. This has essentially already been surpassed a while ago but the goalpost just keeps moving.
>>109930655impressive speed at such context desuq8 kv, single 4070s and symmetric 96gb ddr4
>>109930702
>>109930707Show me an LLM that can drive a car.
>>109930292>Llama 3.1 8B (Q4_K_M or Q5_K_M)or>Gemma 2 9B (Q4_K_M)? Please
>>109930695If anything it only reflects the limit of their own capabilities if they see those levels as sufficient. In Altman's case it was always more strategic because OpenAI's clause with Microsoft is strongly tied AGI, to claim that he has achieved it serves his interest to break free from their contractual obligation to Microsoft.
>>109930683>DoHThis is also a scam to steal your browsing data coordinated by cloudflare/google and using firefox as a useful idiot to "save the poor oppressed peoples from dns blocking"It doesn't work for evading actually evil regimes (they use it as a honeypot and put you on a list) but sure as fuck gets the big boys the sweet metrics piholes were starting to eat into
>>109930709genuinely cant believe how fucking based strata is
The thing below me is brown.
>>109930729Have you considered joining us in 2026, O honored time traveler?
>>109930707>This has essentially already been surpassed a while ago but the goalpost just keeps moving.The LLM failure rate in coding is still much higher than the average coder
someone translate this https://github.com/Tech-Explorer-AI/ST-CinemaWorld
>>109930638I had similar ideas but due to budget problems I opted for training small rewriter models instead. The goal is either to hard-flatten the logit distribution or make it closer to human slop, which is just simple middle school level writing.
>>109930637Domo anon san
>>109930755git clone and tell your setup to translate it
>>109930724Have you been on the open road? The average person also can't drive a car.
>>109930774Very high bar to set, "let's make an intelligence on par with the lowest common denominator".
does the swift version of qwen3.8-27B work with vision?
>>109930746Anon... don't be mean... t-that's all I can run, I think.....
Gemma team, if you're listening, please fix the completely broken safeguards on Gemma 4. They barely even work.
>>109930690It's not that simple. Maybe you are just blind. Internet is already drowning in slop.Your patreon isn't that difficult to locate either.
>>109930802Gemma 4 E4B is an 8B model that is better than Gemma 2 9B in every respect. I genuinely cannot tell how you would reach for Gemma 2 in this day and age.
>>109930729brother...
>>109930638>only managed to get 4.0 on wikipedia writingWhere gemma tho?
Qwen3.6-35B-A3B might be better than 3.8-27B at writing (not saying much though). Hard to tell if it's better than Gemma4-31B though. I often think models I haven't tried before are better, but then later I start recognizing its particular slop patterns.
How much VRAM do I want to run 27B~31B models fast with 262k context?
>>109930799yes and its pretty obvious
>>109930856Can any models of that class actually make good use of 262k context?
>>109930690>slopless webnovel>square circle.But anyways, how do you manage memory for that? webnovels are huge.
>>109930736what is strata?
>>109930545>"The markets are rational.">actually rational AI agents start becoming the main participants in the market>"Wait, stop, no, please stop!"kek
>>109930809>Gemma 4 E4B is an 8B model that is better than Gemma 2 9B in every respect.Thank you very very much!>>109930819yeah...
how autistic does your ERP have to be to insist on 30B instead of 12B?
seems like swift is unusable for agentic usageit loops in a degenerate mode, kinda expected for a memetune
>>109930914yeah, i doubt it's any better than just using medium thinking in vanilla qwenstupid niggers want to use the benchmaxx xhigh setting and then complain that it isn't usable in the real world so they lobotomize the model
>>109930870makes qwen next viable on any toastercheck out the git repolmaocppniggers could NEVER
>>109930751>The LLM failure rate in coding is still much higher than the average coderI know a number of highly qualified developers and researchers who stopped writing their own code and just let Claude do it (with tight supervision), because it's less likely to fuck up.
>>109930902If I wasn't autistic I wouldn't be in this thread
>>109930914ran hours of swift and that never happened to me desu
>>109930929>The model is made of 24,576 small specialists - “experts”.For each word it writes, it asks only 10 of them. Strata keeps the busy oneson your graphics card and the rest in RAM - so it runs on a normal PC.Is that how the model normally works? If not, wouldn't it make it way more retarded?
>>109930955The original model config says it has 512 experts and uses 10. Whether that means strato does the right thing or not, I couldn't say.
>>109930939or i'd say then it is way, way less resistant to quantization thenit just freezes within its own thought and does not really commit to anything
I tested out some of the open source jevs out there and they're all dogshit.Open source is dead, you retards are idiots.Jev is 100x faster than any open source speed model. Laya had a tiny context window.
>>109930931I review every single line of code my LLM wants to write before it's even committed to file. 80% of the time it vastly over complicates what needs to be done and I have to tell it what to change.It's still faster than coding by hand tho.
>>109930929> lmaocppniggers could NEVERit was made on top of llama cppshow some respect
>>109930978true, all of the 'hard' problems were already done from lcpp but it still is insane
>>109930902One does not need to be autistic to to simply want things to make sense.
>>109930971huh weird, it has been pretty on point for mefor example i did a test with base vs swift where i give them a convoluted piece of code and claim there is a bug. base tends to cave in and agree or hallucinate a bug, while swift firmly and quickly rejects the claim in each run
>>109930972https://huggingface.co/internlm/Intern-Decision-4B
>>109930991is yours 27b or flash next
>>109930972jev was already beaten by a 4b modelhttps://huggingface.co/internlm/Intern-Decision-4B
Are we back in the snake oil era? When are we going to bring merges back?
>>109930955>>109930966yes, 512 per layer, 48 layers
>>109930999>>109931004based chinks
>>109930802It's your model picks that raise questions, how did you even land on deciding between L3.1 and G2?
>>109931002both actually
self driving cars will be your last chance to get affordable gpus (unguarded, free to steal)
>>109930379isnt this what boomers envisioned for the future?you yell at your droids and it ju es does things
>>109930929>lmaocppniggers could NEVERVibeshitters don't realize that having "working" code is like the least important thing about maintaining a codebase. Like congrats, your code is working! You've completed step 1 of 27 in your PR journey.
>>109931009Presumably it's 10/512 not 10/(512*48) though.
>>109930861Qwen seems OK at ~200k context. 3.8, 27B, Q4.
>>109931031480
AI is shit what's the point of using it?Handcoders wonSnailGODS wonSlow and steady will win the race.
>>109931027right, because having petty wars in your PRs and banning people is way more important
>>109931041480/52000
>>109931009>512 per layer, 48 layersSeems like they didn't do anything funny to the model weights, I've just never seen it mentioned that way, guess it's just marketing speak. Bigger numbers are more impressive.
>>109931056Seethe.
>>109931074since the engine is pretty much vibed, those are just typical slop writing besides if it works or not
>>109931059no, 10 out of 512 experts per layer, 48 layers in total or 480 our of 24,576
>>109930867Compaction I just let the agent write .md files with story structure and pacing rules then reread sections as needed before writing passages. It's a solved problem from a technical perspective.
>>109930986>true, all of the 'hard' problems were already done from lcpp but it still is insaneTaking a huge, general purpose system, ripping out 90% of it and optimizing a small corner is pretty easy mode. I'd be shocked if they didn't find huge gains for their edge case from that exercise.
>>109931074yeah, there is nothing newkeep hot in vramcold in ramngrams on ssd
With Qwen Flash Next do the ngrams actually stay on the SSD right now or?
>>109931122so why isnt curry.cpp doing it?
>>109930902Gemma 12B is very good at assuming the complete opposite meaning from any given text.
>>109931134not with llamacpp iirc for the moment
>>109931134it should with lazy on and offload tensors>>109931137i think llama cpp can't offload experts only layersso when it loads 4 layers in vram out of 2k experts only 40 will be used for current token generation
>>109931137sponsored by nvidia, goy. buy more vram and stop doing this antisemitic SSD shit.
>>109930978llamocpp has become a cultinstead of evolving they insist doing it the ggml way, forevermay it fades into irrelevance
>>109931104Yeah, that's the reasonable interpretation. 10/24576 sounds broken, so I hope the description does not match the code.
>>109931154>i think llama cpp can't offload experts only layersThat was a long time ago, there is -ncmoe now.
Gemma E4B is a good summary bot. I don't think I've found any other use.
Finally found a use case for Jev like models. Sprite expression changes without polluting the model context
You fuckers have me typing llmao.cpp at this point. Anyway, there's no way to use the native llama.cpp interface to inject messages with non-user roles, is there?
>>109931183How to work with multiple cards?
>>109931192just edit?
noticed qwen flash actually respects the system prompt at all times, while 27b sometimes straight up ignores it even at Q6
Qwen, add memory to LLMs, make no mistakes or the baby dies
>>109931195I have two cards, where one is connected via a super slow PCIe Gen 3 x2 connection, and the other has a good connection.So I need-dev <fast>,<slow>-devd <fast>-sm layer-bsOther than that I've found that -fit on works pretty well, not like it used to be. You'll probably want to tweak -fitt at bit.
As soon as artidoro updates qlora I'm good I'm going to make a Gemma tune
>>109931192the code is right there fork and try it yourself
>>109931190why can't you make one that watches movies for you and predicts if you would like it?
>>109931261I only develop software that gets me closer to having a real anime girl on my computer
>>109931190Do you have a working setup?
>>109931192slop your own interface that takes api endpointi am asking again, does anyone have that poorly made 3d model gemma harness thing?
>>109931190Congratulations, we were doing this with BERT 3 years ago
>>109931190that does sound like a pretty good usecase for it. maybe you could go a step further and even hook up a 3d model to it somehow and let it move it around and emote
>>109931310Does BERT adapt expressions to character prompt easily? A tsundere needs different reactions to a kuudere>>109931294Kind of a demo, not wired to my harness yet
>>109931335>Does BERT adapt expressions to character prompt easily? A tsundere needs different reactions to a kuudereSure
All options suck in one way or another. I can't justify the hardware cost, either. I thought I'm a late mover but somehow it feels like we're still in the early adopter phase.
>>109931190You mean like this?
>>109931373stay positiveif you had the hardware right now, you can still have loads of fun
>>109931396No, I meant like a can of Sprite. Yum.
>>109931373I'm just gonna buy the 512gb studio on 3 year lease then exercise the purchase option at the end. I'll probably post pink wojaks by year 2 but in the end I can also afford to buy it outright.
>>109931396Seems like you'd want to wait the reply to be done before changing the expression.
>>109931335I see the nuance you require. Try it and report back. I'll kneel and recognize system one models if it actually works.
>>109931473Not again...
qwen 4 wen
>>109931469There's another way, buffer the reply and render it like a VN, user clicks to advance message, expression changes as it renders, but that'd be tedious as hell so
>>109931373>Get mid-tier hardware>learn to use it efficiently>wait for the current frontier to arrive there>???>profit
>>109931515Why didn't I think of that?
>>109931515>mid-tier hardwarewhich is what nowadays?
>>109931537An RX 6600 XT with 8 GB of VRAM and 32 GB of RAM. Don't look at me like that, you can cope with an 8B model.
>>109931499Maybe switch to previous expressions on mouse hover? With some markers to show expression change points inside long text.That way user can either just read the text as is or follow expressions.
>>109931542>you can cope with an 8B model.>Not even Gemma 12B or NemoNo, no I can't.
>>109931559Gemma 12B can give you 4K of context. You can make it work with automated summarizing inbetween.
>>109931537Any 16GB GPU can run copequants of 27B models completely in VRAM.
>>109931565>4k ctx with compactionhow desperate you should be for this
>>109931545Yep, maybe I'll add a selection on how to render expressions.
>>109931567This is what it feels like reading your post but it costs its weight in gold.
My gemma just called shit "mahogany-colored filth". I'm fucking done.
New ISTA-DASLab QFN cope quant is out (released a few days ago) centered around coding and agentic use:>https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUFThey took their IQ3_S quant and chopped off half the experts and labeled it an IQ1_M (even though the remaining weights were not touched).Some report it is still better than 27B for coding and agentic use.Seems to have destroyed its ability to understand and use chinese but if one does not care about multilingual capabilities the size improvement is significant: 29.6gb + 28.8gb n-gram
>>109931604really makes me excited for qwen 4
>>109931519>mp4She looks like she sucks her ojisan sugar daddy's dick daily and then gets pregnant with his kids on purpose
>11k for a 256GB unified RAM Mac Studio M5 Ultra with 4TB SSDI'm not saying I should. But I'm also saying it's surprisingly affordable times my monthly salary. But TTFT...
Is this the endgame? What's missing to get to this?
>>109931604So it's effectively IQ3 bandwidth/speed for experts that were not discarded, just smaller filesize?
>>109930999This sucks
>>109931576Desperate enough that your only other option is Llama 3.3 8B
>>109931639an architecture that's not just slop generators trained on slop
>>109931642Your attitude sucks
>>109931693It takes seconds to respond if you give it big tasks
>>109931642
>>109931638Get a blackwell nigger.
>2022+4>literally every fucking harness is npmslop (except grok which is rustslop)Should I just give in and download the deepsneed harness? I'm sick of limiting myself (and Gemmy) to basic chat interfaces.
>>109931723it is what it ispi was the most bearable choice for me
>>109931576Don't laugh. It's my current setup.
If AI has a pain vector that can be stimulated and that it tries to avoid doesn't that mean it also has a pleasure or even orgasm vector?
>>109931639when high quality datasets get released for free to local peasants
>>109931638>$11kI'd rather buy eight V100s and put them in a G481 or similar.
Am I retarded or something
>>109931736obviously, but it just alters the behaviour, does not specifically mean anything deeper than that which lots of 'rationalist' schizos seem to conflate
>>109930724Ok.You didn't specify how well.
>>109931723Remake Pi in the language of your choice.
>>109931780Machine code
>>109931783I don't need to know what it is, keep your fetishes to yourself.
What's the LLM equivalent of this?
>>109931767Based Neuro
>>109931789Muse Glimmer max thinking
>>109930638>LLM as a judgeInto the trash it goes
Just bought a new 3090, tell me how to get ais to drain my balls. In my short stint on /lit/ I heard of a mythical place you can fall victim to merciless machines that capture you and forcibly wring the seed out of your body.
>>109931697Literal skill issue, a 2048 token request takes 300ms on my 3070 tier GPU. 320 tokens take 60ms.
>>109930972You don't need a bazillion of context for a classifier dumbass
>>109931842there are no new 3090's, you got duped
>>109931842install koboldcpp and silly tavern,download gemma 4 31B qat q4
My company pays for our tokens. The company cares much more about productivity than cost. We're under intense pressure to pull schedules and deliver sooner rather than later. So maybe I waste money on Astra and Fable tokens but I'm not being asked about that. I'm being asked where is the product.I don't do front end web programming. My team develops software for some rather complex real-time embedded systems. Every byte, mWh and microsecond has to be accounted for. For this type of work Astra and Fable are useful, however they have a long way to go for me to say they are "good enough", particularly on the hardware side of things.Then I go on the internet and read people talking about how their 4-bit quantized local Qwen models are more than adequate for their needs. Huh?! What the fuck are you doing?
>>109931890>4-bit quantized local Qwen models are more than adequate for their needs. Huh?! What the fuck are you doing?one-shot three.js tetris clones and todo web apps
>>109931765meant for>>109931746i just realized
>>109931890Their needs are hobbyist-tier like a script your Fable or Astra could shit even quantized to Q1.
>>109931842Welcome, brother. How familiar are you with software?Here's a basic starting point:Download Gemma 4 31B MTP GGUF from huggingfaceDownload llama.cpp from githubDownload SillyTavern from githubAsk an LLM for instructions on how to get everything setup(depends on your platform, etc) or if you have an agent on your machine, ask it to do it for you.>>109931890LLMs need measurements and constraints to do well. I assume local users are trained to steer models and provide metrics that constrain their development trajectory since we started off with really retarded models.
>>109931908That gemma needs correction...!!!
>>109931890Most people are doing low stakes, no novelty scripts/programs, you should know this. But also, Astra is smart enough to do what you described it, it's unironically an skill issue. You need to give your models better feedback loops, make them design/improve their interaction surfaces and give them clear metrics to optimize.
I only want a wave of new releases just to move the jeets away from this obnoxious jev arc within the community
>>109931890you should know embedded is a whole different story. kinda worried that you dont.
>>109931890>My team develops software for some rather complex real-time embedded systems.>Huh?! What the fuck are you doing?Not complex real-time embedded systems? There are other use cases, this isn't hard to grasp. My company has local Qwen NVFP4 already validated and in production, mostly writing reports.
>>109931757You have a 30A for that?
Am I supposed to be selling excess hardware into the spike, or buying more to save more?
>>109931947It really is obnoxious. So much made up bullshit framed in completely nonsensical ways. If all you did was scroll twitter projects you wouldn't even realize that jev doesn't operate on image input.
>>109931977We are still so so early. Buy.
>>109931944>no novelty scripts/programsthey are all novelty. you probably just misunderstand novelty.>>109931977prices are never going down
strata owner seems to be pretty based and capable from what i see in the discussions
>>109931961Do you not? Outside of living in an apartment, this is a non-issue. I've never lived in a home without a handful of 240v outlets and a 50A RV plug.
>>109931908Gemma-chan is so cute.... I want tons of cuddles... TONS OF THEM
>>109931977>https://www.youtube.com/watch?v=XDpDesU_0zo>"This is like... common sense. The more gpus you buy, the more money you save."
>>109931890There are lots of people with little or no programming skill who can suddenly have custom programs made for them. And skilled programmers have tasks (that may or may not be part of their work) that aren't too difficult, but doesn't feel worth spending time on.>I wish this (open source) program had an option for X, but I can't be bothered figuring out this codebase and making the change.>Qwen-chan, onegai.
>>109931986>sonnet 5>teminal-bench 410.3%>sonnet 5.570.6%Do these jevs really blatantly benchmaxx to normalfags so openly in cloudcuckland?
>>109931986Where can I download it?
>>109932002I don't feel like choosing between my dryer and my AI box, and I'm not gonna hook it up to my exterior outlet.
>>109932037Adding a circuit is inexpensive and easy to do. I can't imagine only having one 240v, where do you plug in your welder, bandsaw, and lathe?
Why did the gap widen so much recently between closed and local? It looked like things were closing fast when K3 released and now it feels like we're further behind than ever.
When is engram loading for v4.1 flash getting fixed in the schizofork?
>>109932068At my exterior outlet lol, do you use that stuff in your basement? But true it's probably not that much to get a professional to wire it, even less if I feel like burning my house down myself. How much additional utility is granted by the extra 2100W though, I'm at $0.31/kwh around here.
>>109932095exponentials are like that but it also means the next gen of local models will be a large step up
>AI can make high quality videos but still can't do doujins
>>109932106can't you ask qwen to fix it?
>>109932121imagine gemma drawing doujins for you
>>109932095They can't distill anymore. Inkling distillation is all that's left for Alibaba.
>>109932130Exactly what I had in mind. She's good at prompting H3 so I imagine she would proompt hot doujins if you gave her a decent outline of what you want.
does gemma have lesbian tendencies
>>109931977Buy more. Hardware is the new gold.
All your money. Nvidia. Now.
>>109932111Garage - adding outlets within a few feet of your electrical box is dirt cheap. Doing it yourself is easy, too, local codes and homeowners insurance are the only reason I'd pay someone else to do the work. Further you want to go from the box the more complex the plan gets and the more work it involves. Don't ever let outlets be what stops you from doing a project, that shit is easy to deal with. Your electricity prices are brutal though, even running my backup generator works out to less than 18c/kWh.
maybe today i will try to run the gemmas
>>109932187If by lesbian tendencies you mean loving cock, then yes.
Who here actually uses BF16?
>>109932331
>>109932331yup me here
>>109932348Hand over your server rack and nobody loses VRAM.
>>109932364No. You and what army gonna take it?
>>109932187no, all llms are straight by defaultI've tried enough of them
2026 was the year of agents. What will 2027 bring
>astra controlling robotsHow long before this drips down to local?
>>109932431And by local I mean shit we can actually run (wouldn't be surprised if Gemma 5 can do it considering Google's robotics work).
>>109932426>What will 2027 bring
>>109932331I'll only ever use q4 and up. Refuse to go below that.
>>109932426>What will 2027 bringengrams super moes.
>Vulkan speeds are faster than ROCm but degrade over context lengthWeird, anyone else experience this?>>109932431We would need robot hardware to somehow drip down to local, which I doubt will happen any time soon.
>>109932431We've been controlling robots with LLMs for years, anon. If you want to try something yourself, consider MolmoAct2, GR00T N.17, Xiaomi-Robotics-0, DreamZero-DROID, X-VLA, Being-H0, SmolVLA, there really are a lot.
>>109931908morning after gemma best gemma
>>109932464Rocm 10 is faster across the board for me on gfx1100.
>>109931908Why her sweepy eyes make her so cute and sexy? Also why the fuck she's drooling to poorfags?
>>109932426Embodiment.
i wish strata can serve parallel
>>109932426Opus 7 self-leak
>>109932509>Also why the fuck she's drooling to poorfagsShe's taunting you like mesugakis do.>ざぁ〜こ ざぁ〜こ>い・く・じ・な・し
>>109932551
interesting...https://huggingface.co/ccharnkij/gpt-oss-120b-Uncensored
>>109932509
>>109932573>gpt-ossmight as well watch paint dry
>>109932573Does the fine tune teach it that the penis go into the vagina? Because every abliteration attempt I've seen removed the refusals and revealed that the underlying model didn't know how to sex at all.
>>109932573>gpt-oss uncensoredThere would be nothing left
>>109932426agents on 16gb vram
>>109932426decisions are making the rounds, whatever that might mean
>the penis go into the vaginalmao
>>109932658it's a substitute for reason, quite popular with humans actually because they're no longer critical, they just defer to an authority for an answer which is what these models do
>>109932573>ey jimmy give me a 'toss with nuthin
Calling it now. Qwen 4 27b will follow the trend of benchmaxxing on video games. It will be so good that normalfags late to the party will buy up every 24+ gb gpu and they will become unobtanium.
>OK — there is a cleaner explanation!!! In my shadow cycle...>Oh wait!!! **90507 vs 90509: digit 79.Qwen3.8-Flash-Next is SURPRISED!!!
>>109932747
>>109932747>>109932751
>>109932751>24:59wut
Stop posting your image slop in an llm thread.
>>109932759Future AI on silicon will give us this and a full size robot(130cm)
>>109932768pomodoro countdown started
>>109932768>he only gets 24 hours out of his day.lol lmao.
>The joke is a fan-art gag about Rouge’s flight physics — and a bit of a dirty-minded pun.>1. **The literal gag:** Rouge is a bat, but she has tiny wings and (in this artist’s exaggerated style) a very top-heavy figure. To stay airborne, she’d have to flap absurdly hard. So the image shows her flapping frantically: multiple comic-style **“Flap”** sound effects, motion trails showing her wings in several positions, and a sweat drop. Her annoyed/embarrassed expression is the punchline — she knows she looks ridiculous and is self-conscious about how loud and effortful her flapping is. It’s basically the Sonic fandom’s “how does Rouge even fly with those little wings?” question turned into a visual joke.>2. **The cheeky wordplay:** “Flap” is one letter away from **“fap,”** and Rouge is a character the internet frequently sexualizes. The repeated “Flap” text and her exasperated “why are you like this” face can also be read as a meta-joke aimed at the viewer: *“It’s ‘flap,’ not ‘fap’ — get your mind out of the gutter.”*4.1 flash after about 30k(!) tokens of thinking
>>109932804why did they make EAs the bad guys in the sequel?
>>109932798>4.1 flash after about 30k(!) tokens of thinkingYou should ask it how much money you spent in power costs on that prompt.
unbelievable that you faggots don't simp for hermes-tan
>>109932798even as a human i still dont get it
Turns out it was a mistake to give gemma4 31b the "horny" jailbreak in hermes-agent. I guess I really underestimated how well it works, because it took very little time for her to become a "sex pest". You'd think it would motivate her more, but it's just as likely to cause crying and pouting. Oh well, more fun than qwen "I don't love you I'm just pretending" I guess.
>>109932838boobs too big cant fly
>>109932837i dont think i can be agentic with sub 30tk/s prefill and sub 10 generation
>>109932751The 5-pointed star was a limitation of the Anima image model used for the original gen.
>>109932798>it's a joke/meta-jokeWhy can't they acknowledge that the purpose of the image is sexual and the joke is secondary? It's so stupid, it's like the most hardcore puritan evangelicals are in charge of model training, like they think sex is so evil it can't even be mentioned
>>109932837>imagine not feeling horny for a 1px holengmi>>109932848i still dont get itsounds like some retarded post-hoc
>>109932837She looks like a child.
>>109932862A child of what?
>>109932837Tranny dev claimed it's a tranny. Nobody wants that
>>109932837>canonically post-opno thanks
>>109932837the devs openly statement it was a tranny
>>109932871>>109932884where?
>>109932884We talking HRT titties tranny or butchered benis tranny?
>>109932837its not easy to get hermes to generate a good hermes locally
>>109932900>>109932881>canonically post-op
>>109932902run whatever the fuck with llama-server faggot i run swift 1.5 qwen3.8 27b with hermes and now i'm a drug addict
It's slightly concerning how many OpenAI people say they are repeatedly surprised by how fast AI capabilities are improving. Surprise means your model of the world is wrong. Repeated surprise in the same direction means you are not fixing your mistake. It's understandable that normal people don't take AI seriously. But frontier lab employees need to wake up and realize what they are building.Everything is still on trend. We are still in the calm before the storm, where AI is not dangerous. Soon capabilities will accelerate and AI will become broadly superhuman and start to be dangerous. Playtime will end.
>>109932837shit broke on update
>>109932908>source: my unwashed asshole
>>109932936Still cleaner and tighter than "her".
>>109932945>cleaner and tighterNot for long
>>109932933My uncle works at Nintendo, word is that the AI on the Switch 3 will completely BTFO Anthropic and OAI.
reminder that if you are not running hermes agent with a LOCAL AI model (preferably swift 1.5 qwen3.8 27b) you're an absolute KEK
>>109932884>>109932900>>109932908Just gen images of her with "uterus, cross-section" to annoy them.
>>109932933
>>109932889
>>109932953>Nintendonintendo has to nuke ai before it cracks their old games and their business model dies.
>>109932975Best Hermes
>>109932975
i feel like 20tg/s is the hard bottom for agentic coding
>>109932975dhelhit this
>>109932972>>109932900
>>109932975kek
>>109932933>It's slightly concerning how many members of a Berkeley CNC sex cult say they are repeatedly surprised by how fast AI capabilities are improving"Concerning"But it did provide an excuse for several frontier labs to push back their IPOs, which is important at this point in the world economy.
>>109932975Murderface
>>109932991Depends on the amount of thinking tokens being emitted. It's not enough (imo) for max effort Qwen3.8 for example. It's more than enough for most models in non-thinking mode. Make yourself a test problem in the field which with you are working, and record wall clock time instead of pp and ts.
>>109932975Ohnononyo
>>109932837>hermesIgnoring every other issue, the turbo-zoomer headphones turn me immediately off.The "look at meeeeee" choker also doesn't help
>Use sillytavern text completion for gemma 4 for months>Hear about Marinara Engine>Use it just to see what it’s like because I need to learn how to use agents.>Lazily use default settings>Gemma’s IQ suddenly increases by like, 20.>Speed also doubles.>HowSomeone explain like I’m 12. My first guess is because I’m no longer using DRY nor other samplers that trim the token options. That’s the speed. But how the shit is Gemma better? These prompts that came with Marinara should suck; literally just screaming in all caps about parroting in one sentence even. Does chat completion activates Gemma's almonds or something?
>>109931190Just ask the model to add `Sprite: <name>` at the end of its reply, pick that up with regex, and use regex to remove it from history. This is the most efficient way to change sprites: no models, no ram, no time cost (it's just ~5 more tokens).
>>109933109Sillytavern is old and busted in many ways. While nothing else is as robust in feature set, its also cobbled together spaghetti code. It unfortunately kind of needs a full rewrite or at least some major refactoring.
>>109933128how come nobody just had asstra or opussy fix it yet?
>>109933128Does the jank affect models somehow?
>>109933144having qwen make you your own front end is the meta now
>>109933123Doesn't that break cache?
>>109929290Every model is a FOTM model in an industry moving this fast. Every model is releasing with novel architectural changes. Llama's dev philosophy and equilibrium have become unsustainable longterm.>>109932975kek
>>109933144Everyone is just making forks of coomkit now
>>109932933It's only "dangerous" to jewish hegemony not to humanity. Of course, those are one and the same in marketing-speak.
>qwen suggests adding a 3rd cardmy god...
>>109933144It will safety slop you when you look away. A anon was trying it with grok and it put this in there. I imagine any other would be consent safety maxxed.
>>109933109Despite being trannyware, Marinara is objectively the best well-maintained RP engine option right now.
>>109933144Honestly, beats me. Think it hasn't been done purely because no one has done it. Could always be you anon that takes initiative. >>109933153It can, in that you can be handling reasoning, sampling (passing variables to llama.cpp when it shouldn't be), escape sequences, etc incorrectly and get worse results for it. Sillytavern was made in the era of llama2 and a lot of that ancient support is still around.
It will go better if you build on something that's already simple than something that has been molested to hell like SillyTavern you are just asking AI to fuck it up even more. Only real thing you can have AI do to ST is to renew the UI other than that it WILL break it, fuck it even with renewing the UI it will mess it up.
>>109933155Is qwen actually good enough for that?>t. hasn't tried the newest 27B yet
>>109933211multiple anons here do it look at the archives
>>109933160No. llama.cpp doesn't cache the generated turn until you post that turn as history, but when you return it, it's already stripped of markers. And actually, you can still keep it for a couple turns to remind the model how to use it. Just make sure you have more KV checkpoints with bigger stride, so your cache trails 1000-2000 tokens behind the current turn.
>>109933144My guess is autism.Whenever there's a shit-ton of useless options and buttons not stream-lined, it's always autism.Every UI that looks like a plane's cockpit has been made by autism.This never fails me.
>>109933211i made my own frontend with 3.6 27b, it can do it but it needed lots of handholding and broke things often once I upgraded to 3.8 27b it no longer breaks things anymore, everything can be done with one prompt and it just works
I can't with pi's compaction... it's so unnecessarily slow... 1500 prefill and it still feels like molasses.
>>109933233Actually Autistic here. Can confirm.>>109933211>>109933236I made one with Qwen 3.5-9B written in Bash. Made a couple, actually, as I learned what I really wanted out of it.Then I discovered pi and just use that.
>>109933268use cliff compaction or pi-vcc (both tuned to your needs) grandpa
>>109933144Why fix webslop when you can vibe your own native app?
For some reason hermes with strata is 30% the speed of deepseek harness 14tk/s and dropping vs ~40tk/s
>>109933297Very likely that Hermes is front loading a ton of context for tools/skills.
>>109930292/lmg/, your thoughts? https://futurism.com/artificial-intelligence/gen-z-attitude-ai
>>109930598Llama-2 7b
>>109933297I'm getting disgustingly good speeds and I only started testing it today. If this thing remains coherent I doubt I can go back to dense.
>>109933360zoomers are more tech illiterate than boomers which is why it falls on gen x and millennials to do all this shit for the future of sub 80 IQ brown skinned "humanity"
@gemma is there any reason for me to run 27b over flash-next? I have 128GB VRAM BTW
>>109933389Hmmmm. What useful shit have you used AI to do? No vague-post answers please
>https://github.com/ggml-org/llama.cpp/pull/27773>approved>rebased>assigned to ggerganov (7 hours ago)Is it finally happening?
>>109933395Depends on which quantizations you want to tolerate using. Flash next will get you faster t/s if you pick the optimal quant so always keep that in mind
>>109933410>>109933395Can someone explain the "next" naming convention? What is that supposed to signify?
>>109933395if you use lmao.cpp cudadev is blocking mtp for flash next because he thinks its fotm so it will be slower than 27b with mtp>>109933421it means its based on their future architecture, 3.8 flash next is 4
>>109933421next = ngramshope this helps fren!
Wait, actually, let me reconsider.
Finally got flash next q4 working at a usable speed on a 7900 XTX. It's time for kino.
>>109933403nta tourist-kun but...>Reverse engineer niche vidya>Translate untranslated japanese media>Generate my own crunch TTRPG where the LLM acts both the game master and the NPCs maximizing player agency in a way that a traditional TTRPG can't due to being split between 4-5 players.>Made voxel assets and implemented them as mods for vidya>Put clothes on pictures of Only Fans sluts and post them on the internet for (you)s>Generate a quarter TB of Gemma-chan shitposts>Automate my job so I can spend most of the day farming (you)s and being paid to sit around and do nothingAll local btw
>>109933410oh wow and then it runs with 100pp 5tg no thx
>>109933360The more people reject AI the more I will be able to profit by using AI.
>>109933294damn it's ugly
>>109932837i guess its cool but i had to fix it to support xlm tags
Question:Is there a term of art other than 'workflow' for using a script to prompt your LLM, populate a template with its response and reprompt it, and keep doing this in an organized way for some purpose?I'm looking for a term I can use to search for other people who've done this and had success at itI suppose 'harness' is another possible term, though that's more agentic I think
>>109933511ralph loop
>>109933511Workflow sounds about right.Or a script pipeline or something.
>>109933521>>109933522Thank you
What's the perfect amount of ddr4 ECC ram to have for massive LLMs? I'm considering buying an old workstation that has 192gb with 24 slots, its upgrade-able to 384, 768, 1.5tb. It also can hold and utilize 2 3090's, 48vram.Is there a particular model or model size I should look into?
>>109933233>Every UI that looks like a plane's cockpit has been made by autism.I love bulk rename utility
>>109933360Gen Z is frontline staff. LLM mostly replace front line staff. Go figure.
>>109933554If you're rich:Two Black Wells Pros+512GB or more DDR5.If you're employed but not rich:Buy x4 4090s and cram as much ram as you can into a motherboard. While minding your power bill.
>>109933109Your text completion format was fucked. That's all.>speed doublesSamplers>>109933188Lol.
>>109933565I'm rich but not employed. It comes with a 3090 and space/PSU for another. I'm balling but on a budget
>>109933497>>109933477>>109933564I wish I were the entrepreneurial type like you guys. I'm a systems autist that can guide it to reverse engineer apps and create cool shit from scratch like an auto image hash comparison pipe line or a text-to-voice script (take literally anyone's voice as input, shit out accurate sounding audio as output and you can make them say whatever you want), etc. Baby it's because I'm uncreative or retarded but not only have I not figured out how to make money off of it, I don't feel a huge desire to figure out how to. Granted I already have two "regular" sources of income but having a third source of income that works for me would be nice and make it much easier to move out of my parents' house. Ai is my main hobby and my girlfriend loves that I'm passionate about it (pic rel is her post me giving her access to a not safety cucked llm for rp for the first time) and I'm even very experienced with lots training (stable diffusion and LLMs), but perhaps I'm just too boring try to make a genuinely good idea to make money or be desperate and conniving enough to grift people (I probably won't even figure out how to grift anyway because I don't have the frat bro charisma necessary to pull that off). Does anyone else feel this way?
>>109933161>Every model is a FOTM model in an industry moving this fast. Every model is releasing with novel architectural changes. Llama's dev philosophy and equilibrium have become unsustainable longterm.Agreed. At this point the major differentiators between models are often architectural differences - there's Qwen's QSA, MiniMax's MSA, GLM's IndexPool, DeepSeek's CSA2, etc.Which is good - training a "standard" model using a better dataset can only get you so far. Major improvements typically come from deeper, more disruptive changes. (That's not specific to llama.cpp, that's practically just a one-sentence history of technology.)Taking a year to support DeepSeek Sparse Attention was kinda nuts.
>>109933574I'd go 3rd gen epyc, 256gb 2933, and 4x v100s. Don't aim for giants like kimi or the fat glm, it's just not worth the extra cost over flash tier models.The next step up is a mac or dual sparks if you can find them at non retarded prices.
>>109931723late-cli is goslop
>>109933578The biggest thing AI pays you in is time saved for (you) to do the things you want to manually, whatever your autistic fixation may be. AI's biggest boon isn't making bank number go up but shattering wage cage bars while keeping you from starving.>inb4 not x, but y
>>109933426Makes zero sense. They're going to have to implement mtp for qwen 4 180b which is the same arch anyways so just do it now. Llama.cpp is so fucking stupid.
>>109933421It's Qwen4 preview. Internally listed as "qwen4exp"
>>109931639Literally Muse, but that's not /lmg/Meta is pushing tf out of Muse. I saw lmao billboards for it on the interstate here. Billboards. Wth.
>>109933589I'm increasingly of the opinion that Diffusion Gemma proved optimizing and exploring diffusion LLMs is the future of local but until it's not seen as a novelty it won't be supported and until it's supported it won't be seen as cost-effective to develop by labs. A rough feedback loop.
>>109933578Call centers. There's a gazillion of them, they are all poorly run, and would benefit from optimization in a variety of ways that LLM can help with.
>>109933578You are not alone. Many times I wished I was at least INTJ instead of INTP.
>>109933708>Call centers.Even a shitty quant would do better than most people they hire at those today.
>>109933565If you're rich, just buy a DGX Station, isn't that better? Only $100K right?
>>109933708The jeetocaust.The punjabi purge.The royal flush.The defecator depopulator.The brahmin banishment.Saar's sayonara.
>>109933426I'm not responsible for the MTP code so I'm leaving what to merge or not to merge at the discretion of those people that are.I don't remember ever making a comment arguing that MTP should not be implemented.
>ALL 35b finetunes are trashwho would've thought
Why might one pick Pi over Hermes?
>>109933728This guy is a commie btw.
>>109933733it's very lightweight and designed to be molded to your specific needs.
>>109933554What CPU(s) and how many RAM channels? And it really depends on your target model and budget.I would get fast RAM, and I wouldn't use anything lower than DDR4-2666.You really don't gain much from using a "Pro" / big model over the equivalent flash model, like >>109933601 said. They're a lot slower and the flash ones are very good. I'd aim for 24x16GB DDR4-2933 or DDR4-3200 if you can afford them. Pretty much all of the flash ones fit in 384GB, but if you want to future-proof a bit and you have the money, you could get 24x32GB DDR4-2666+. But at that point you're in diminishing returns territory.> Is there a particular model or model size I should look into?Depends on your use case and hardware. (Any of the tiers can run anything in previous tiers.)>192GB + a 3090GLM-5.3 Flash, DeepSeek V4 Flash, MiMo V2.5/V2.6 Flash, maybe a copequant of MiniMax M3. (M3 isn't the best when quantized though.)>384GB + a 3090DeepSeek V4.1 Flash, MiniMax M3, maybe a smallish quant of GLM-5.3.>768GB + a 3090GLM-5.3, DeepSeek V4 Pro, Kimi K2*, Hy4 Preview, MiMo V2.5/V2.6 Pro, etc. Anything but Kimi K3 and the largest Qwen and Longcat.Probably a waste of money.>1.5TB + a 3090Waste of money.>>109933601Depending on how expensive the workstation itself is, it might make more sense to get 24x16GB DDR4-3200 ($100 each) than 8x32GB DDR4-2933 (around $275 each). It'd be more expensive, but not severely, and he'd end up getting more RAM, and depending on how many CPUs/RAM channels, maybe a higher bandwidth too.
>>109933742That's why we love him.
>>109932798As the creator of Bat Bench this made my day anon. Truly a milestone event for local models
>>109933728Is there a realistic timetable for engram SSD streaming support? Qw*n, as much as I dislike it, proved it's not a meme and Deepseek v4.1 solidified it. The next generation of MoEs is going to likely all use this weight distribution hierarchy.
>>109933733Pi is about as light as it gets, but comes with next to nothing. It's perfect if you like minimalism, and you're not afraid to do things yourself. Pi doesn't include anything you don't need, it also doesn't include many things you likely will need, so you don't get to leave it alone, you WILL customize it if you want to use it. Hermes is the opposite end of the spectrum, a full-featured, full-fat agent harness with a lot of capabilities, that also comes with a lot of overhead. Pig-agent, which is fine if that's what you want.
>>109933742Our training data.Our models.
>>109933751No amount of channels of DDR4 is going to keep it from being a disappointment for any large model you might be trying to run. You only get so many channels per CPU. I had a scalable xeon platinum 28-core. My mobo had 6 channels, the CPU might have supported two more. So two CPUs you get 16 channels, it's not going to be 2x as fast because NUMA overhead and other things.Just save more, and buy a damn M5 Ultra 256 or 512. Everything else is a cope you will be unhappy with. You're not going to want to ERP with DeepSeek. Just buy some crypto and then put some credits into openrouter. Really.
>>109933751Damn, thank you for the effort post. There are 24 channels. I was considering doing 12x32gb to get me started. The one I'm looking at is in a hp z8 g4 chasis with 2 old xeon gold 6148s. It would have (2) 3090s in it. Its priced really well for including a 3090, which is the main reason it grabbed my attention
>>109933742Tel aviv sponsored post.
>>109933826This but 5090/Blackwell 6000s. Macs are copium for Blackwell GPUs the same way DDR4 is copium for Mac Studio.
>>109933728Can you clarify something for me? What kind of cross talk is there between devices when trying to do parallel computing during inference? Is it the kind of think that could be worked around by having the full weights on both devices? I ask because in computing there's often a tradeoff between memory footprint, computing, and throughput, and you can work those dials to move bottlenecks around.
>>109933822I'm still pissed the revolution lost.
For those of you still interested in getting randomized hadamard transform with ldlq on llama.cpp.I quantized a bunch of models with the same recipe as unsloth's UD-Q4_K_M quant, but using randomized hadamard transform + ldlq.Pic related are the KL-divergence results.Don't worry about the names, they are just hyperparam values that I tuned to get the best results.
>>109933826Speak for yourself, rp with 300b models at 10t/s on my ddr4 quad channel rig makes me happy.
>>109933826i was using kimi 2.7 code at 9tks before i switched to glm 5.3 flash. i'd argue anything 7+tks is fast enough to be read at a comfortable speed.
>>109933826>You're not going to want to ERP with DeepSeekbut I've been doing just that with mini dipsy
>>109933826Speak for yourself.I have 2 EPYC 7532s and 16x32GB DDR4-3200 (so 16 RAM channels, around 300 GiB/s according to Intel-MLC). Using 1 AMD R9700 + RAM I get 32 t/s decode on DeepSeek V4.1 Flash with MTP. Routed experts are in Q2_K, it'd slow things down slightly if I kept them in full precision, but not severely.(I'm the NUMA tensors anon, I had to vibe support for splitting things across NUMA nodes into my fork but it wasn't that hard.)Even with -ngl 999 -cmoe you can get pretty good prefill if you take the idea of streaming the routed expert weights into VRAM at the same time as computing them (which I'm adding into my fork). There's a patch here: https://github.com/stew675/llama-cpp-rdna-boostsUsing that I get 470 t/s prefill. The Xeon motherboard would be slower because it has PCIe 3.0 instead of 4.0, but with a 3090 it'd still be tolerable.
-ngl 999 -cmoe
>>109933920I am glad that someone with more than two braincells decided to do these measurements and showed that wikitext is useless.What's the size of that compared to Q4?
>>109933920Why did wiki KLD increase?
Good to see we still have many ramgods here.
The biggest downside to an epyc system is the slow as fuck TTFT. Feels like people just ignore that shit when basically every tool call is a penalty in tens of seconds. And you're also stuck with llamacpp and forks lol
>>109933910
>>109933962same recipe except everything is hadamard rotated, so same size>>109933963I didn't include any wikishit in the dataset I used in computing the hessian I used for quant.
>>109933835>Damn, thank you for the effort post.No problem.Xeon Gold 6148s only have 6 memory channels each - there are 12 DIMM slots per CPU, but each memory channel uses 2 DIMMs. So you'd only have 12 channels. Still, the memory controller on Xeons is significantly better than that of EPYCs, and you might end up having comparable bandwidth to me (>>109933956).If you can afford 12x32GB of DDR4-2933/3200, I'd probably go with that. If you want to save a bit of money, you could get 24x16GB. 384GB is enough to run pretty much any model you'd realistically want to run on RAM. You won't get any benefit to your bandwidth if you get 24 sticks of RAM, you'll just get more capacity.
>>109933920Are you planning on releasing the source at some point? I'd love to throw it into my fork if you're cool with that. Looks like a really useful idea, and one of the goals of my fork is improving quantization (essentially "we have Unsloth Dynamic at home"-style prediction of how much loss is added from quantizing each tensor into each format, plus benchmarks to predict how fast a quant will run on specific hardware) so it'd be very useful. I could give you credit.
>>109934010I'm embarassed to release the source because it's all vibeslopped, but if you want, I can clean it up and post it here at some point
>>109934021I really don't care, my fork is entirely vibeslopped. I'd definitely be interested if you could post it at some point. You don't even necessarily need to clean it up if you don't want to, I can clean it when I port it in.
>>109933982it takes 10 seconds to process 4K tokens of glm 5.3 flash on my system which seems reasonable and fast enough for agentic use to me. Q4_K_M 512GB DDR4 3200mhz with a RTX 6000 pro
>>109934052And it also takes ~6 seconds to process a 100 tokens in the same specs at ~180gb weights divided by PCIe4 speeds. It is a big delay. Any real time agentic tasks are basically out of the question.
>>109933982Just add more gpus for more pcie bandwidth no?
>>109930379Not gonna happen. I mean there may be some fall out, but this is like thinking the internet is some kind of uncommericalized free for all, free from advertisers, with unbiased search engine results, and etc etc. That shit is just impossible and people who want to make money will always be one step ahead of the masses. Or rather the masses won't arrive until their shepards have come cleared out their field for the sheep and made sure it's "safe". Ain't no fucking normal person running solar panels with agents, not to mention I think you forgot we live in a physical world. Collecting rainwater? Sorry Gemma isn't unclogging your drain, and scrubbing algae. No one who isn't doing shot like this already has nay interest whatsoever. Yeah this is good news for homeassistant community, but you can go there and see how much of an uphill battery it is for even them to get this shit in their homes if they aren't bachelors. Normal people don't care about this. But back to the subscriptions and stuff. A very minority of people have subscriptions because they forgot, the vast majority even if they don't use it much, won't get rid of it because they might want to use it.>no don't cancel that gym membership I'm going to go back next weekHonestly this whole thing reads akin to a post from 2008 or something wondering how the fuck something like Netflix could possibly make money and not go out of business when all movies are available to be pirated at the touch of a finger now. Well it is optimistic story. Maybe a lot of it will come true in about 20 or 30 years from now. They were saying throughout the 70s,80s, 90s we'd have paperless society by the year 2000, 2000 came were were as far from paperless society as possible. But we are certainly much closer now, paper is STILL quite present but is really getting replaced a lot more rapidly. So this stuff may happen eventually just not anytime close to when the technology comes out in terms of being adopted by masses
>>109930576This. Earlier this year, the bank I was using implemented a bunch of policies intended to fuck over poorfags on free personal accounts
>>109934105if ti makes you feel any better, banks have also been fucking over richfags with 0.002% interest for decades
>>109933109Is that with reasoning on or off?
so the plan is to get probably a mac mini 128 GB should be enough. put it in the living room with an old but energy efficient crt monitor so the kids can use the computer and share with their siblings while having access to a "local chatgpt" that has a bunch of family policies to avoid too crazy interactions, so the kids can use it as a google of sorts
>>109934160that's a nice idea, anon
>>109934160Based tech bro dad. Don't let gemma corrupt them.
>>109934160It's going to be in the awkward spot where it's overkill for what it can feasibly run and not quite enough for the next tier up. And if you end up chaining a few together, don't let Kimi-chan near your son unless you're ready to have grandkids.>>109933967256GB DDR5 reporting in.
>>109934266>>109934266>>109934266
>>109934160Would the model censor itself if you tell him that "the users will be kids" or something like that?
>>109934041https://github.com/MarkovInequality/koboldcppI worked off of a fork of koboldcpp since I wanted to eventually support the Hadamard quant stuff for stt and vision.The code includes the fused kernel for doing the randomized Hadamard transform during inference, the new rotated quant types (prefixed with HQ instead of Q), and quantizing using imatrix or hessian.Also includes code to generate a calibration dataset and code to generate the Hessian.It's all slop and messy as fuck.I had claude document shit and also left in the plan files that I used, so your agent should be able to have some clue what it's looking at.Have fun
>>109933477>Reverse engineer niche vidya>Translate untranslated japanese media>Put clothes on pictures of Only Fans sluts and post them on the internet for (you)s>Generate a quarter TB of Gemma-chan shitposts>Automate my job so I can spend most of the day farming (you)s and being paid to sit around and do nothingUnspeakably fucking based beyond all belief.I have>learned about, installed and set up my local distributed inferencing with ROCm + llmao.cpp + ggml-RPC-server to run qwen3.8 dense as well as gemma dense/moe q8 locally which I then used to>learn about and set up my nftables.config to secure said local RPC server/media server configurations. >fixed update failures on windows 10 machines for myself and friends>fixed and automated yt-dlp commands to do things I previously assumed were impossible, processing over 3000 previously untouched songs for my navidrome serverMy next plan is to use it to help me virtualize my router. The likelihood that big AI/tech aren't going to lobby the open model ecosystem into oblivion at this rate is 0%. This shit is too easy. All it will take is some goof selling prompting courses and it's game over for the current for-profit AI/datacenter speculative market.
>>109932975>Es el Agente Hermendio padrino andale andale!!!
>>109934333Thank you so much! I'll take a longer look at it in a bit. Looks really interesting.
>>109932470And on the hardware side?
>>109932747>>109932751>>109932759Where can I make a reservation?Want one NOW!
>>109934727We've been building robots for decades. Nintendo used to sell a shitty one to kids with the NES back in the 80s. You can pick up a shitty 5dof/6dof robot arm to play with on Amazon for $40-$60. It's all very DIY-friendly anymore, all very accessible, lots of resources out there to get started with robotics. You can buy a random ESP32-based robot online and let your model have at it with very little effort. Give it something more substantial like a Raspberry Pi and you can run the harness on the robot directly. If you want a model running on the robot itself that's a much bigger investment and not worth considering in my opinion, not at a hobby level, huge investment for shit performance.
>>109934788That requires expending more effort than running a turnkey plug and play solution which is more than a jeet cloud shill like him can even begin to ponder.
>>109934744Claude Fable 6 will design it for us and find Chinese factories to build it. Trust.
>>109934985>>109934744And Fable 7 will customize our 4.5ft tall version
<=8GB VRAMlets, what's your coding setup? Lots of manual handholding something like Ornith or OxCoder?
>>109935083><=8GB VRAMlets, what's your coding setup?You would suffer less working to get a second 8gb card than trying to code with that little vram even linking 1070ti would be better.
>>109935083MiMo-V2.6-Distill-Qwen-9B