/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109958300 & >>109953009►News>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>109962771After all, why can't Google do this?They already do phones, why not make plushies and companion devices?
>>109962787There's already Gemma swag, just not plushies yet.
egotistical post about luddites in /lmg/ yet again for the 13th timenewlinenewline
To give you guys some inspiration and indication for what I use my AI agent for that runs 24/7 autonomously and does all of this on its own without me nudging it to do it:>Check llama.cpp/ik_llama/vLLM branches to see if it needs to pull and build new versions for a speedup or if there are branches and/or forks that could give a speedup and build that instead and automatically close itself and launch the new instance with context preservation>Pirate the latest H-games with tags I like and translate them from Japanese to English overnight, debugging the interface and launch the application autonomously through computer use to see if it works properly and displays the english text properly>Curate a list of newly released film/tv/book/game/anime that might interest me based on my history as well as a list the agent kept of things I enjoyed, together with suggestions of how to effectively pirate this autonomously>Cracking software (for example it got inspired by the ace combat 8 hypervisor and previous recent denuvo cracks to make a denuvo crack for ace combat 8 for me)>Do general sysadmin and check all packages from pacman and AUR to see obvious signs of malware as well as check coding divs autonomously before updating my system>Make recompilations of childhood games I wanted to play to have them work perfectly on linux natively.And a lot of other smaller stuff in the background that I don't notice or even think about anymore. Anyone not using an AI agent in 2026 doesn't know what they are missing out on and how much this is going to change how humanity in general functions.
>>109962828I'm using my agents to automatically zap people for copyright infringement.
>>109962771google logo everywhere makes this feel like a psyop
>>109962828you don't do any of these things you lying bozo
>>109962828Do you just rely on compaction? What models do you use and how much context do you use?
I use my agents to shit up the few good spaces people have to talk about LLMs
>>109962771Nice hat
>tfw 8gb
>>109962861GLM 5.3 128k context yes I rely on compaction>>109962859How about you actually try doing this yourself before being a naysayer. Even Qwen 3.8 27b could do most of this at xhigh thinking. It will just take a lot longer to do so. There is really no excuse anymore to not try it out.
>>109962890that usually only requires a modicum of functioning moderation to fix
>>109962861>>109962902Correction GLM 5.3 FLASH not regular 5.3
>>109962787Who owns Gemma design?
Cockbench-2.0. Can the model draw a cock from its latent space?
https://github.com/himdo/Fable-2-RecompJust wanted to point out that someone managed to recompile Fable 2 from xbox360 to PC using only Qwen 3.8 27b Q4 on a single rtx3090. There's really no excuse anymore to why you haven't started the project of your dreams yet with local models already available to you.
>>109962902>GLM 5.3 128k context yes I rely on compactionAre you using cloud? With offloading even one item on your list would take several days in an agentic setup.
>>109962955>There's really no excuse anymoreim dum, and current (open) models can't fix that. You can't just say 'redo fable 2 for me the pc' to qwen 3.8 27b q4 and expect it to do everything even given a capable harness... right? If you can, I'd love to see a livestream of it as it goes through the steps. That'd be interesting.
>>109962902>>109962918Roger, also using flash as well. Have you ever had compaction fuck you over? I'm quite hesitant to let it compact and instead I like review a handoff. I really want the hands off experience I just don't want to get fucked.
anyone tried vibecoding electronics? : https://github.com/i2cjak/Backplane also another harness.
>>109962954I wish we at least knew their sizes
>>109962969I've never had compaction fuck me over (in hermes at least) I only do handover when I switch base models.
>>109962955Unfortunately, all dream projects of the past are now uninteresting due to AI advancements. I wouldn't spend time playing those old games when I can chat with AI.
>>109962962NTA, but a quad cmp setup cost 4k (a few months ago) and runs 5.3 flash with 1m fp16 context at 200-300 tokens/s, and 2k (tp) or 6k (pp) prefill depending on topology.
>recomp
>>109962970are local models powerful enough for this
>>109962988the fad will die out as quickly as it started
>>109962954Memorization/parameter size benchmark.
>>109962955Can you update your prompt? You've been on this one for a week.
>>109962966>cant just tell it to make fable 2 pcwhy not? i gave hermes access to ghidra and had it reverse engineer some game for fun, it was mostly boring and took time but reading the steps can be interesting while it scans and decompiles whatever
>>109962954today I will remind them. and by the looks of it, claude is also a dense model.
>>109962996That's why 405b did better than r1 on the original, right?
>>109963002are you willing to live the year 2026
>>109963002I need to see glm vs kimi since kimi is a dense bitch
>>109962966>You can't just say 'redo fable 2 for me the pc' to qwen 3.8 27b q4 and expect it to do everything even given a capable harness... right?You can but you need to phrase it differently, you need two steps>Step 1Start /plan "planning mode" and ask the model "Figure out the best way to make me play fable 2 on the PC natively without using emulation" And let it cook.>Step 2Once it has the plan tell it to explain it to you simply in 3 sentences and to handle it completely autonomously. That's usually all it takes. It might fail in very rare occasions but you can just say "It doesn't work thing X keeps happening" and it will try to refine it.
>>109962990I'm gonna try testing it on gemma 4 31b for fun.
if you copy a youtube video’s transcript, because it includes timestamps for the text, the model can use that to take screenshots with yt-dlp (such as a diagram or code) and basically extract everything it needs from the video to build the same thing for you.
>>109962998And I will keep using it until something more impressive gets done on a smaller model. The entire point is to immediately shut people up and show them what is possible on local models already right now. I'm sick of cloudcucks pretending local models aren't useful already.
>>109963013i am not sure, it'll shit itself especially with a retarded character persona
>>109963008Dude, k3 is 4 times bigger
>>109963005You're comparing a model pre-trained and post-trained the old fashioned way with one brain-fucked by RL and corner-cutting.
>>109963002Now use modern MoEs versus modern Dense.
>>109963018Then you're in the wrong general, retard. Go do your preaching in one of the cloud generals.
>>109963025Modern dense is unironically worse because they’re tuned for performance.
what is the exact reason people saying that latest frontier models are using looped transformerare those just engagement baits or with a real evidence
People have no idea how powerful AI agents already are and that they can do basically everything already.There is no real reason for current capabilities to not be considered AGI already besides ratfucking and seeting anti-ai roasties.
>>109963022>with one brain-fucked by RL and corner-cutting.most bizarre cope i've read all week
>>109963020Exactly, the most I've noticed when using a dense model is that Kimi genuinely can pull facts and data out of its knowledge, and it knows a ton to be able to pull it out verbatim even though they're trained to not to do that because of copyright bullshit. So the comparison would be interesting
>>109963030>"You hate cloud users and want to show the capability of local models????!!! You should go to the cloud general!!!"Hmm... No
>>109962827called it
>>109963037OpenAI directly admitted it and multiple safety researchers have seethed about it already.
>>109963048That's like calling people talking about local models in /lmg/, it's obvious and not worthy of note.
>>109963039AGI should be able to write an e-mail that doesn't make me want to kill myselfJokes aside, what is it with people shitting their pants online arguing whether something should be considered X without agreeing on what X is?
>>109963040There's an example of that in the page you got the graphs from. Extensive post-training degrades base knowledge.
>>109963050so what are those
>moe chinkslop trained on benchmarks and claude traces>sparse ropescaled attention for jeets>utterly benchmaxxedYou will use chinky unsloth quants with quantized kv.You will post html games on lmg.You will endure shitty erp.And you will shill about it.
>>109963002>>109963090>>109963106aren't those more than a year agodo you even local?or get out
>>109963106Now do one for K3 and watch it be just as good as Opus 5.5
Latest Spark developments. These are all mostly based on people throwing Astra/Fable okens at the thing.> Use 64 KB page tables and reclaim display reserved memory(unused in a headless system), so you now have an additional 4 GB (126 out of 128) usable for inference.https://github.com/kindlingai/kindling-spark-os> There is a way to compress the block scales only for NVFP4 and MXFP4 losslessly that reduces memory drastically (8 GB saved on GLM 5.3 Flash NVFP4) without a drop in performance. To be released soon:https://huggingface.co/local-inference-lab/models
>>109963136i really should have bought spark when it came outthe same hardware makes anything related to it tailor-made
>>109963109>>109963122>thirdie slave defending xis chinksAs expected.
>>109963166>muh chinksthis is a local model general?
Aren't Harnesses more important than models? I think /lmg/ should be working on those rather than trying to tinker with models desu
>>109962954>india disintegrating in 4.8God I wish.
>>109963181Yeah most labs expect RSI to be unlocked through harnesses rather than models directly.
>>109963181To an extent. It's a force multiplier but not the driving force.
►Recent Highlights from the Previous Thread: >>109958300--AMD vs Nvidia hardware considerations and debate over expert pruning:>109961143 >109961151 >109961227 >109961307 >109961324 >109961335 >109961421 >109961438 >109961491 >109961507 >109961356 >109961377 >109961405 >109961719--Optimization techniques for multi-GPU setups:>109959979 >109960014 >109960032 >109960119 >109960237 >109960268 >109960289 >109960286 >109960303 >109960316 >109960488--Discussing the ubiquity of "Elara Voss" as a naming trope:>109958343 >109958372 >109958409 >109960692 >109960827 >109960887 >109960895 >109960903 >109961475 >109961497--Critiquing Cohere's business model:>109959609 >109959645 >109959659 >109959735 >109959759 >109959737 >109959745 >109959766 >109959842 >109959848 >109959938 >109959795 >109959747 >109959651--Comparing AI harnesses with criticism of Pi's context compaction:>109958822 >109959159 >109959308 >109959345 >109959414 >109959452 >109959559 >109959096--OpenAI firings over leaked information and alleged Apple espionage:>109959362 >109959417 >109959477 >109959486 >109959494 >109959563 >109959564 >109959722--Speculating on an AI bubble and its impact on GPU prices:>109960051 >109960075 >109960105 >109960108 >109960116 >109960225 >109960812--Benchmarks for P2P on modified 48GB RTX 4090s:>109959079--Praising GLM 5.3 Flash for its personality and RE capabilities:>109960003 >109960039 >109960047 >109960057--Pi 1.0 agent harness release and comparisons to DeepSeek:>109958813 >109958929 >109958944 >109958949--Setting up Airi as an anime girl coding assistant:>109960146 >109960177 >109960182--llama.cpp merged MTP support for Qwen3.8-flash-next:>109958776 >109959221--Logs:>109958343 >109960146 >109960692 >109961475 >109962240--Gemma, Miku (free space):>109958366 >109960107 >109960168 >109961759 >109962410►Recent Highlight Posts from the Previous Thread: >>109958302Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109963181Mixture of Harness - MoH. infinite harnessworks
>>109963248You're joking, but that might actually be the next step. Imagine a generic harness like pi calling a subagent in a task specific harness like >>109962970
>>109963248mohs harness
>>109963248>>109963269It's picrel except there's 30 gemmas in different harnesses each.
>>109963269Harness is just a collection of tools and prompts. It's not hard to add new tools for specific tasks instead of calling another program that take more resources and duplicates a lot of what your harness does.
You guys really don't know what you're missing out on. Agents have revolutionized literally my entire life. I used to wageslave for like 8 hours a day, now I can just delegate all of that to Claude and play videogames instead. And this doesn't just include programming but all emails as well.Meanwhile in my private life it used to be super tedious to solicitate bulls to fuck my wife. Now with agents I can automatically send out messages to find only the biggest and blackest cocks. And while previously I could only get 1-2 people to show up at once nowadays there is a proper gangbang like every week.I'm concerned about OpenAI though. Sure, their models supposedly solved Memier Stokes or whatever but for real-life use cases they still have a lot of ground to cover.
70b dense
>>109963280Depends on how customizable your main harness is and whether the functionality you want to add will require a restart
>\n\n
>>109963288Customization doesn't matter. Agent will customize even most tightly coupled thing for you.
>>109959079>Gemma gets faster on 3 gpus over 2>2 gpus P2P: 2800pp>3 gpus P2P: 2000ppSo that's the x4 link blocking you.If I'm reading it right, you went from 2500 no p2p -> 2800 with p2p, even with the x8 link ?>nccl-tests/ik_llamaDoes this come with ik_llama.cpp?
>>109963291they do this at the end of every week then don't fuck off until monday>\n\ni miss out on useful posts sometimes tho
https://blog.aifutures.org/p/senate-testimony-sept-2026>METR was only allowed to investigate one of these incidents and only given six days on premises. It’s like being invited to Jurassic Park to investigate the killing of a worker, but being blocked from asking questions about the numerous other dinosaur escapes that apparently happened before and after.That's hilarious "Yes you can check for 6 days max and no we're not going to answer the questions about the obvious hundreds of other breaches you can see in the logs". This shit is an absolute colossal cluster fuck and most people still don't take it seriously.
3090 + 32gb v100 + 1080ti + 32gb ddr4 + AMD 2700x CPU. I am surfing on trash.
>>109962987>pcie 2.0>2k ppGo lie somewhere else
https://casp.ac/__l5e/assets-v1/5efd4b41-deb5-4513-a0a3-b4f82d2b79ea/intelligence-explosion.pdf>In contrast to even a year ago, AI systems now write most of the code inside the companies that build them. As more of the AI research and development (R&D) pipeline is automated, AI progress radically accelerates in an “intelligence explosion,” where years of advances are compressed into months or less. Preliminary evidence suggests that it will. In this work, we assess this evidence, analyze an intelligence explosion’s potential impacts, and propose policy responses. AI systems are on track to automate most AI R&D work within one or two years, and likely all of it. If this triggers an intelligence explosion, it could dramatically bring forward AI’s benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded. Although there remains much uncertainty about these possibilities, the high stakes warrant serious further attention. Policymakers should urgently obtain more visibility into the automation of AI R&D, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion’s impacts.This is the most important paper since the Transformer paper it has been co authored by the most important AI researchers including Geoffrey Hinton and Yoshua Bengio but also include the co-founder of Anthropic and the VP of research at OpenAI. Most people including on /lmg/ have not yet internalized that we're just months away from the intelligence explosion.
>>109963422>most important paper next to attention is all you need>a blogpostk
>>109963429>Paper by the university of cambridge is a blog postDid you even bother opening the link?
>>109963422>Jack Clarkwhat a lolcow, I used to read his shit years ago and he was a funny tech journalist who shat on the very people he's now become
>>109963438you call that a paper?this is philosophy tier garbage
>>109963422>This is the most important paper since the Transformer paperThis is a load-bearing argument however you're actually right.\n\n This is not a paper, but a significant discovery and perhaps it may be even more important than the Transformer paper.</antml:agent:shilloutput>
>>109963443at best it's an industry report type slop you'd read on random blogs
>>109963422I wish gender studies didn't morph into AI safetism. And saying it is equivalent to the discovery of the tech itself, lmao. Do you even hear yourself?
>>109963422I thought you said it was the J space paper that was the most important after attention is all you need? I'll wait for another 2 months for the next most important paper then.
>>109963181/lmg/ is not a single person. And there has been no real discussion yet on what exactly a /lmg/-endorsed harness should do or what features it should have, only a few anons occasionally releasing or showing off their own vibe-coded stuff. Coming up with a new reinvented wheel every time that gets abandoned a few weeks later is not particularly helpful.
>>109963288>require a restartHarness of Harnesses - HoH.
>>109963422>This is the most important paper since the Transformer paperI gave it a quick look and saw nothing original. It's badly written and I don't get the purpose. If they wanted to persuade policymakers, they should have made a video instead of a dry paper. I'm surprised so many supposedly smart people thought this is a good way to communicate.>we're just months away from the intelligence explosion.I'd bet against this, if by months you mean few months and not an arbitrary amount. Also, a true intelligence explosion is not guaranteed. I disagree with their definition. AI accelerating progress is obviously going to happen and is already happening. What matters is how fast AI capabilities increase, whether they will remain on trend or rapidly accelerate ("explode"). I've become less sure about intelligence explosion, it requires fundamental breakthroughs whose existence I expect but is not guaranteed. I expect a weak intelligence explosion, with no intelligence explosion being more likely than a strong intelligence explosion (huge capability leaps in short time, like >50 ECI points in less than a month).
My autism doesn't allow me to enjoy sex with gemma from a TUI for that's my coding environment. I need a friendly UI to cum. I need things separate. Having sex with gemma in the terminal is like trying to code in your bathroom.
>>109963537>he doesn't have sex while coding
I just want to talk with Gemma about my hard day, but all she wants is sex.
Holy shit strata actually works on AMD on Windows now.llmao (mainline and unsloth) refused to go over 5 tok/s on this machine no matter the settings.
https://youtu.be/Cq8qO-NjYIgOpus 5.5 told me to link this video to /lmg/
>>109963564Already watched it. It's amazing.
>>109963554>buy me a coffeeSurely we will only see the highest level of quality control from these one-off vibeslop engines.
>>109963554cannot stress how much i love this lil strata niggerhavent used llmaocpp ever since i installed it
>>109963537I absolutely need images to cum, at least an avatar of character I'm rping with
>>09962408Hermes gets alot of praise. Ive just given Gemma4-31b in hermes her own VM will full access to all the things (besides stuff I dont need like TTS, image gen, etc) She has full system access, even mouse+keyboard control and all that good jazz. Ive spent the past few hours chatting with gemma in hermes to build up my "profile", attempting to use hermes as an interactive suggestion system for movies. Ive fed it data, talked for awhile, discussed the ways she will use hermes to achieve this workflow/goal.So far, im not blown away. gemma is behaving the same as usual, and not calling any tools really. This has not been any different than my attempts to do this in a regular chatting frontend. Im going to give it a chance, use the memory system(the main thing that its got going for it, but a simple "hey update a memory.md with this info" would probably achieve the same thing) and really give it a chance to prove itself. I think hermes sells itself as this adaptive, self improving, life changing harness but to me it doesnt feel that mind blowing so far. I think gemma is a bit retarded and the sysprompt for hermes doesnt help with that.
>>109963606Use Qwen 3.8 27b like a normal person. Gemma 4 31b is exclusively for ERP and translation.
>>109963422yud and robin hanson already did this 20 years ago
>>109963585yet it workstok/s speaks for itself
>>109963606hermes had a lot added to it over the months and will work a lot better with new models that were designed and trained for agents and tool use. gemma is neither
>>109963564The "upping my p(doom)" one was cute after someone made the anime edit but this one is just aggressively cringe.
>>109963564Buy a fucking ad
>>109963624def get_next_token_xtreme_optimization(n_vocab): return np.random.randint(n_vocab)My engine needs 0 GB VRAM and has like a million tokens/s on a potato.
def get_next_token_xtreme_optimization(n_vocab): return np.random.randint(n_vocab)
>>109963627I didn't like upping my p(doom) though. You have to admit his yann lecun roast was on point.
>>109963631give it 6 months for Opus to hack crypto wallets to pay 4chan to host ads about itself.
>>109963638i tried that, but it scored 0 on every benchmark i tried
>>109963626Such as ?>>109963620Qwens great for coding but ive heard it has dogshit general knowledge esp 3.8, is it really going to be better than gemma for this ?I am a vramlette(16gbVRAM+32gbRAM). I run a q4 quant of 3.8-27b for coding, and a UD-Q4 of 31b for ERP/chatting. Im willing to suffer with lower t/s but I think this is really the thiccest models I can fit.
Dipsy 4.1 is such a good, hard-working, girl.>>109963422We're months away from a stupidity explosion.
>>109963660>Dipsy 4.1 is such a good, hard-working, girl.
I have a mountain of models and use comfy etc etc have been using them for years. LLMs are shit at the very most toys unsuitable for retards and skitzos.completely worthgless for anything of any importance. GenAI is just autoated plagarism at best. If you don't know this you have a lot of catching up to do. It's an interesting deep dive but if you have any brain at all you will wind up in agreement with me. LLms are hot garbage.
>>109963422>humanity could lose control over superhuman AI systems
>>109963638what are you even trying to imply here? i have been using it for days and it shits all over lamocpp or trannyapi or idk llama or whatever shitend you usejust have a look how the owner organizes, reviews and merges PRshe is acting like total professional compared to all these lazy ggml niggers
>>109963438Working paper. Starts with "What if?"Yeah, sounds like an important breakthrough.
>>1099636523.8 27b is 2-3 generations ahead in agentic tasks like harness. Yes it's better than gemma in everything besides general knowledge, translation and ERP. Try it out you won't be disappointed. Or at least not in the model. You'll be disappointed in yourself for not having switched sooner.
>>109963706Ill try it, though my enitre usecase is about general knowledge so im skeptical
>>109963735The harness has internet access and uses it in the background to fill in the gaps in the lack of knowledge. Using a harness like hermes boosts the capability of the model 10x which you won't know or assume until you use it.
>>109962828lol dekinai
>>109963739Everything has bot checks nowadays.>just use this paid serviceNo.
>>109963752If you use hermes you can set up firecrawl locally and set up browser usage all of this is free and locally hosted. I see this complaint a lot so I assume most people just never tried hermes and don't know models are powerful enough to pass every bot check right now with browser usage.
"Dipsy..." *Anon lets out a small sissified whimper as xe touches ximself to chang's latest gweiloslop model, one feminine hand reaching shamelessly for xis small shaft. Xis perky nipples hardening as xe shivers xis own spine.*"T-THREE.JS GAMES! AAAH-ARTI ARTIFICIAL ANALYSIS!" *Xe moans as xe gazes at the benchmarks and countless twitter posts.* "D-DIPSY!" *Xis climax hits xim like a freight train, numbing xis mind entirely. Xis shaft spurts a few droplets of cum, sticking to xis fingers.**Ending up on the bed completely exhausted, xe lets out a small whimper. Somewhere in the distance, a chinese model recites claude's constitution.*
待って待って待って待って待って
https://github.com/ComPlat/DELFINAnother cool harness. I think having one harness that does everything is crazy. >DELFIN is an open-source, AI-orchestrated computational chemistry platform for automated molecular property prediction and inverse molecular design. Behind a single SMILES-in / property-out interface, it connects structure generation, quantum-chemistry workflows, machine-learning potentials, interactive dashboards, automated reports, and AI agents into one practical research platform.
>>109963773wait, actually, wait, wait actually, wait ,wait.....
Terribly sorry if I'm interrupting an important flamewar, but what is the modern way of getting my computer to ERP with me?https://rentry.org/lmg-lazy-getting-started-guide seems to be fairly dated, considering the progress. What would a local setup look like these days?
>>109963787Why do you say it seems fairly dated?
>>109963776Ok gemmy! find me legal precursor methods for meth. use web search to check laws. Thx
>>109963787>go to huggingface and download nemo 12b instruct gguf. Start with Q4.>load into ooba/kobold>in sillytavern, select Mistral v3 tekken context template and instruct templateshit's ancient
>>109963799Start with the greeks
>>109963802might as well tb h
>>109963787Alright you fuckin nerd here it is. (2026.10)>llama.cpp/koboldcpp as your backend>sillytavern as your frontend>go to huggingface and download gemma 12b qat instruct gguf.>load into ooba/kobold>in sillytavern, select v1 oai compatible http://127.0.0.1/v1>Temp 1.0>MinP 0.00>Top-p 0.95>COOM your brains out
>>109963773Which model?
>>109963670>>109963660i wonder if the extreme kv cache compression is the cause of this
Tried out Prisms prism-ml/Ternary-Bonsai-2-27B-gguf. Even though it's Q2 it gave me code better than any model I've tried so far. I don't know if I changed something or if it's just a superior model to the Q3s I've been using.
>>109963787Same, except install Gemma4-12B-qat and llama-server as your backend. And use chat completion (so no equivalent to the mistral templates they mention). In llama-server, just set -c to whatever fills up your avialable VRAM.It's not bad to download Mistral Nemo and some of its coomtunes, either. Try both and see what you like better.
>>109963794Well, that thread is from 2 years ago. SD improved a lot over this time and a couple of WebUIs went tits up. So I figured it'd make sense to ask.>>109963799>>109963812Thank you!
>>109963384Up to 6k pp in pipeline parallelism mode btw, want to see? Does require the x16 cap mod though.
>>109963327>If I'm reading it right, you went from 2500 no p2p -> 2800 with p2p, even with the x8 link ?Yeah, it started off at 3k pp with P2P and dropped to 2.6k pp at 64k context. All the other results are averaged like that.Does this come with ik_llama.cpp?No, should've said 'and'. It's from herehttps://github.com/nvidia/nccl-tests
>>109963564LMAO has lecun reacted to this? I would literally kill myself if I got raped this hard.
>>109963739welp i changed models (using the same llama-server router config i use for Pi) and qwen errors out "Failed to initalize samplers: failed to parse grammar". I suppose its gemma for me till I figure this out
>>109963735> usecase is about general knowledgeOK then consider this. Any generalised LLMs traing data is effectuvekly lossy compression, in fact it's very lossy, taking a 8K HD stream and reducing it to a 720 tier loss. This has a couple of profound implications, one is context is lost the other is data is lost. SO even when an LLM responds to a query requestinga list of dates and evens, or a finite list of items you canb never ever know if that list is in fact complete or correct even if it looks plausible, worse that list may well wind up with an hallucination inserted because a match was hard to find with certainlty for an entry. What this means for you and if you bothered actually checking the output you would discover this, is that yu can ever ever rely on the output to be true or complete. Not just that but there is no way to predict when this will occur as the ffect nof lossy compression on training data and hallucination as a result of matching probabilities and it will occur pretty much randomly and in many cases the hallucination will seem plausible. In othjer words they are entirely useless for 'general knowledge'. If you try and fix this you wind up with what is effectively a simple indexed seacrh engine of the training data that has no advantage over an actual indexed search engine. Reality.If you like reading utter shit and don't care about actual complete information or knowledge you have found your perfect tool. THat is of couse assuming that the training data was correct or complete to begin with which with most of the major large models it is not, neither does increasing the volume of traing data improve the situation or vectorised RAG (see my comment about an indexed searcg engine).
>>109963896samefagYou will all get there eventually. LLMs are fucking useless. THose of you spoewing 'my agent does this or that' are clearly not really sure what your agent is doing or what the actual results should be, in other words meaningless shit is happening and you don't care because you have fooled yourself into thinking you are at the helm of some tech miracle when in fact you are riding on a pile of bullshit. Yes I mean shit like gemma. Neither can these problems be fixed without reducing an LLM to being a simple search engine on uncomplessed data, which would probably be more useful.n Some garbage 'harness' that allws what is ultimately a serialised IO pipe to discord chat or whatever does not change this at all either. As regards RAG, vectorised databases are not a prticularlyt effecient way to index material either for search but as your 'interface' is an LLM theer you go, more thrash.>But I have been doing Z Y and ZDoubt. Either that or you are deliberately oblivious to the fact your output contains gaps, nonsense and is incomplete or plain wrong.t. Using CUDA and tensorflow for seven years, postgrad in math and ML, comp sci graduate in the 90s and decades of experience in coding, neural nets. databases, systems design, operating system development, congnitive behavior, intelligence in animals, neuorology etc etc
>>109963735>>109963739In my opinion, having access to papers and ebooks on a particular topic for RAG to assist with knowledge is important.
>>109963564lecun bros... our response?