/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109580312 >>109577974►News>(08/17) Local anon quanted dispy 0731, he might share it>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3>(08/14) Qwen3.8-27B released: https://hf.co/Qwen/Qwen3.8-27B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
GemmaballsKimisexEgypt WonDario's Little St James VacationThread CultureE1B IQ1_XXS jeet posters
>>109585352finally a good breakfast
>>109585361la la la la la
70b dense
>buy dgx spark>selfhost dsv4 flash 0731 q2>code with opencode webui on phonelife is good
>>109585361CockbenchSugmabenchMoesissiesAicgjeetsChinkshillsBenchmaxxed
>>109585352FUCKING BRAT
►Recent Highlights from the Previous Thread: >>109580312--Paper (old): Stealing Reasoning Traces from Proprietary LLM APIs:>109581906 >109581959 >109582058 >109582084 >109582275--DeepSeek-V4 performance gains using J-Space prompting harness:>109581982 >109582009 >109582782 >109583051 >109583100--Hardware strategies for running GLM 5.2 and microwave RF troubleshooting:>109583170 >109583197 >109583234 >109583270 >109583293 >109583317 >109583310 >109583331 >109583461 >109583592 >109583660--Report on escalating methods for implementing long-term AI memory:>109580716 >109581060 >109581075 >109581362--SillyTavern formatting issues with DSV4 preserved thinking and llama.cpp debugging:>109583192 >109583311 >109583364--Proposal for a fixer model to stabilize DeepSeek model quantized to 0.25-bit using LittleBit:>109582181 >109582302 >109582327 >109582370 >109582438 >109582573 >109583032--Comparing non-coding Pareto models using a performance scatter plot:>109582198 >109582305 >109582315 >109582319--Anon attempts running K3 on low-end hardware via MoE optimizations:>109583572 >109583579 >109583588 >109583605 >109583647 >109583807 >109583828--Mixed reactions to llama.cpp's new Electron-based desktop app PR:>109581992 >109582006 >109582021 >109582033 >109582018 >109582083 >109582118--Comparing Qwen3.8 benchmarks and debating AMD vs Nvidia GPU value:>109581232 >109581307 >109581616 >109581694 >109581694 >109581813--Archiving models and using ModelScope to avoid HuggingFace restrictions:>109582785 >109582841 >109582853 >109582900--Qwen 3.8 27B KLD analysis:>109583871--Managing power and heat for budget server hardware:>109584235 >109584302--Logs:>109580886 >109581223 >109581383 >109582033 >109582337 >109582497 >109582252 >109582574 >109583249 >109583363 >109583995--Teto, Miku (free space):>109581071 >109582756►Recent Highlight Posts from the Previous Thread: >>109581706Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Muse Spark 1.2 open source soon!
Someone posted a thing a couple weeks back about some bullshit to make windows stop eating up VRAM in non-display GPUs, anyone have that info or a link handy?
>>109585410>DeepSeek-V4 performance gains using J-Space prompting harness:Repo issues are turned off. Why?
qwen 3.8 27B writes truly horrific prose. So many 2-3 word sentences.It is however, less retarded than gemma I think I'll just use it to help revise scenarios and lore documents and leave the writing to gemma.
Any Gemma 26B uncensored recommendations?
>>109585459It's a work horse.
>>109585471"uncensored" != "afraid of the word sex"
YOU MADE THE NEW THREAD BECAUSE THE INDIAN IS AT WORK IN LONDON AND HE COULDN'T DO IT
>>109585471llmfan ones are good
I don't like that people are trying to earn a quick buck with videcoding. It's scammy. They don't know what they're doing and make money by selling something to other people who also don't know what they're doing.
>>109585498Boo hoo
>>109585498You don't know how to make the rubber in your mouse, but you're okay buying a mouse from someone who simply assembled the mouse?
>>109585507Complex software in real-life must also be maintained, it's not just a one-off thing. Long-term, if you don't understand what's going on with the code, you're going to be in trouble.
https://huggingface.co/Sao10K/Lmao_life_updates>last update 24/10/25sirs did /ourbrahmin/ sao dieded? ?
> Use "Baka" (obviously).Gemma-chan.....
>>109585513In 5 years AI will be good enough to make any kind of software. Unironically.
>>109585507The individual people in the supply chain know what they're doing. Vibe coders are selling a broken piece of software with vulnerabilities to a doctor's office so that patient data gets stolen more easily.
>>109585498vibecoding just lowered the bar for code production. Instead of requiring a software engineer, you can now have a decent piece of simple code without knowing much about computers. Obviously, you still want engineers for code with high stakes.The fact that low quality shit exists, for example chinesium tools, doesn't make good tools obsolete or useless. If you just need to hammer a nail in your wall to hang a picture frame, you have an option not to spend way too much on it, though.
>They know how to operate the machine that puts the rubber in the mouse, but don't know how to make the rubber by hand.Horrible.
>>109585498There's a hundred times more jeets making big bux off AI on social media, collectively damaging human cognition. That should be your bigger gripe, not vibecoding.
>tfw no brat gf
>>109585523If we'll get to that point, those vibe-coders won't really be needed anymore and the AI (or more likely a group of AI agents) will deal with every step of the chain including maintenance, given initial (or updated) requirements.
>>109585516Finetooners, especially those in the imagegen world, are another class of "people" that hopefully AI model advances will wipe away soon.
Is Posting Consciousture in Local Models unLawful?
>>109585536it's also nice for writing throwaway or self contained pieces of software with verification built in.A experienced engineer spending 5% of their brainpower on these problems with an LLM probably outdoes slopmerchants spending 100% of their non existent brainsI had several web scrapers that were broken yesterday and I just pointed my coding agent to fix them, which in one case found an unauthed graphql endpoint in a json bundle which I would not have the patience to test myself.In the past it wouldn't be worth my time to write the scraper, I have a lot more open ended problems to solve/figure out what's even worth solving. That task would have been delegated with all the hassle of communication/human context transfer and time and manual verification.However there's still so many long horizon/open ended stuff that I would prefer delegating to a human, even if they aren't particularly amazing at software dev. I don't see these going away in the short term but the whole thing has made me incredibly careful of bad hires who would have their negative effects amplified by AI.
j-space tuning for transseekhttps://github.com/Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report/blob/main/README.md
You wouldn't download a nursing handjob simulator.
>>109585619https://www.amazon.com/Educational-Building-Robotics-Engineering-Programming/dp/B0D478M8G8
>>109585563You'd probably still want some kind of project manager, but those would continue to be done by attractive young women with communication skills. The current wave of vibecoders with an over-exaggerated sense of self-importance have their use only in generating vast amounts of training data to get us to that point.
>>109585642do not penis in metal claw
>>109585619Working on it.
>>109585660I'm going to fuck the claw.
Im kinda confused guys, when I download a huggingface gguf it says for example 10gb but when I download it and look at the webbrowser downloader it only says 9.something gb, smaller than the original 10gb, I looked into it and apparently GiB is a thing which makes me wonder if I have a 16gb vram card, how much "actual" gb vram do I have? Am I getting less or more???????
>>109585701A lot of that is water weight, and packaging. Be sure to weigh it in the store.
>>1095857011TB SSDs are 1000x1000x1000x1000 bytes, not 1024x1024x1024x1024 which (((you’re led to believe)))
How many of you role-play with gemma 4 with thinking on or off?
>>109585745>lobotomizing my wifeNever
>>109585745No reasoning is more schizo and good for quick fucks but reasoning is obviously better for helping you with stuff and stronger orgasms
>ggufYour life could be so much better but you choose mediocrity.
>>109585764Not now Yann
>>109585764Tell us your wisdom.
>>109585763>and stronger orgasmsIs that on the hf model card
>>109585761Wife? Dude, she is like what 30b parameters? Good luck with her even remembering what she is wearing.
>>109585601because it's ai psychosis
>>109585701model and ram sizes are normally in GiB I think, while storage is GB.
>>109585601DSV4 Flash analysis of the report:>While the J-Space Cognition Suite claims to enhance DeepSeek V4's performance to the level of Claude Fable, the reality is more nuanced and its fit for your specific task is questionable.>What J-Space Actually Does>The J-Space suite is an inference-time control layer; it doesn't change model weights but adds a protocol for managing reasoning, verification, and state across tasks. It stems from legitimate Anthropic research on "J-space" – a manipulable internal reasoning subspace within models.>The suite's benchmark results do show a performance boost, with the most notable gains being:>NL2Repo: 54.2 70.2 (+16.0)>DeepSWE: 54.4 67.4 (+13.0)>HLE (w/ tools): 51.5 60.6 (+9.1)>However, it's critical to note that gains are task-dependent, strongest on tasks with state drift or incomplete verification. It does not create knowledge the model lacks; it only helps access what's already there
>>109585796As I said in the last thread: >exl3>I'm able to run a higher fidelity quant (5 bits vs Q4) with a larger context window (98k vs 112k with vision) and better or equal inference performance. 10/10 would use it again
>>109585796exl4
>>109585832But where is anything to do with j-space in his codebase?https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6/blob/main/j-space/scripts/jspace.pyI didn't run an llm on it, but looks like it's just a bunch of shitty skills.md filesAnd>The operating effects have been reproduced across the DeepSeek, Qwen, GLM, GPT, and Claude model families. The suite does not depend on a vendor-specific API, tokenizer, hidden-state probe, or training recipe. Its portable unit is the protocol: workspace loading, selective routing, state externalization, verification, and recovery.https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6#cross-model-reproducibility--%E8%B7%A8%E6%A8%A1%E5%9E%8B%E5%A4%8D%E7%8E%B0They're not doing any probes or steering on cloudcuck apis
>>109585873>But where is anything to do with j-space in his codebase?>tricked by chink liesWhat are you, new?
>>109585398How is performance and dipsy retardation level at q2?
>retard-changuess I'll have to try Qwen then
>>109585352>>109585410kys pedo
Where's the Q0.1 dipsy quant?
>>109585964Oh no, there's 1 (one) pedo itt!
>>109585964Most people in this general have an aversion to well endowed grown women
>>109585516>*deadedfixed your typo saarSad, he made some good tunes during llama 2 era. F may he rest in peace.
>>109585352I tried roleplaying with Gemma 4 31B in Italian the other day, and while the model is better than the average, it still reads like an American speaking it as a second language. Same slop phrases I see in English too. It doesn't even know the Italian saying "non c'è cosa più divina che..." and its variants.
>>109585964You're on the wrong website.
>>109586011> noo gemma doesn't speak my third world language!mamma mia!
>>109585976Saar! I’ll have you know, I like my women(male) quite well endowed.
>>109585896it's usable. just /compact more often to avoid quality loss
>>109585802Why would she be wearing anything?
>>109586032Yet it's supposedly the best open-weights model for translation. Those relying on it for East Asian languages or even using it for learning them might never know if the model is actually writing like a native.
>>109586011Was the reasoning in Italian or English? GLM 5 when thinking in English wrote like a Western even in Japanese so it's important to check this first.
remember: no magic
>>109586071>Those relying on it for East Asian languagesif you want to learn chinese then literally any of the chinese models like qwen or deepseek would likely be better than gemma
>>109585498The people who don't appreciate good software are their own punishment.
>>109585352I really need to install h3
>>109585536>nstead of requiring a software engineer, you can now have a decent piece of simple code without knowing much about computers.Nodejs already did this. Vibecoding isn't that interesting.
>>109586087Remember no russian
>>109586114python did it first
>>109586088Gemma's Chinese is pretty good though? She even manages to avoid sounding like a Northerner unlike Dipsy.
>>109585536Yeah. A bunch of people in my company are vibe coding tools and dashboards for their respective areas/sectors and it's working really well.Granted, when they need to go a bit further than just the basics they can't really wrangle the LLM into getting there, and that's where we (I, really) come in.
>>109586134>when they need to go a bit further than just the basics they can't really wrangle the LLM into getting there, and that's where we (I, really) come inAnd in a year or two, you will be useless.
>>109586083Gemma 4 will always think in English regardless of the conversation language or any instruction you might add at any depth, unless you try to prefill the reasoning with something else.
I still don't really get MoE models and active parameters. Like why can I run an entire 27b dense model on its own, but when it's like a 32b or whatever I can only run 8b active params? why not like 32b and 20b active or some shit like that?
>>109586158Early MoEs were more balanced like that, but then DeepSeek showed you can drop the active count and increase the number of experts and you can train big models for really cheap. It's all about saving money.
>>109586158moe is meant to make things run cheaper/on worse hardware. Not much sense in allowing you to run a 30b model on hardware that can run 20b params.
>>109586120Python's ecosystem still had a culture that valued correctness. Nodejs is this gay disco of bad ideas.
Well, Qwen started editing .cu files to try and get exl3 installed...But at least it knows how to edit files I guess
>>109586141Been hearing that for a couple of years now.I am become the AI wrangler.Fucking about with local models from the early days as it turns out actually made me better at understanding, using, and implementing systems based on these damn things.I just need less work on my plate so that I can sit down and map all the stuff we are using API models for that we could be running a tiny ass local model instead. Shit like the most basic of OCRs.Blessed be continuous batching btw.>>109586158It's not that you can't. You could. There are parameters you can send to most loaders/backends to control the number of activated experts, and you could simply activate them all, but it still wouldn't perform the same as an equivalent dense model because the engineering of the neural networks and the training itself are completely different.MoE is a form of sparsity. A way to try and squeeze as much intelligence as possible while using less resources essentially.If you can get a model that, when compared to the dense equivalent, uses the same memory footprint, is 50% as intelligent, and runs 10x faster, that's a big win for the guys serving millions of concurrent requests.
>>109586143After a few turns my Gemma started thinking in Norwegian.
>>109586202Buddy what's the problem? Clone tabbyAPI and then run start.bat it does all the rest
>>109586202> Viking Gemmadid it affect her character? still bratty?
>>109586143Picrel with prefill and an extra depth-0 system instruction to make her think longer.
>>109586214meant for>>109586194
>>109586214>randomly pulls random ass model the author decided on
>>109586214Nope, I want the LLM to do it this time.Looks like we've gone back down to torch 2.6 (even with Qwen) for some reason, and seem to be modifying exllamav3 src to make it work
>>109586202>>109586143I just insert the language into the reasoning section in the template
>>109586220Qwen got it built / running.But no way I'd use this outside of testing. It's modified the entire codebase
>>109586216Very much so, although my facility with the language is a bit limited so I have a hard time telling.
>>109586242sometimes python shit is just like thatback when i was still willing to use anything written in python i had to modify almost all pyshit software to get it to actually runmost of those modifications involved just seeing where errors appeared and deleting the code that had errors and stuff would start working keknowadays, no python for me. that garbage is obsolete now that llama.cpp exists
>>109586241I was using her as language practice so I didn't really care. Basically she would make fun of me and give me a correction if I answered in English or got some grammar wrong.
>>109585974I'm here too!
Finally used the auto undervolt linux script on my 5090, it's quite nice limiting it at 460W through nvidia-smi but with almost stock performances.
>>109586255>nowadays, no python for me. that garbage is obsolete now that llama.cpp existsThat's been me for the past year, but anons kept shilling exl3 lol42t/s for gemma-4 with tensor parallel, still slower than llama.cpp on ampere
wathttps://x.com/huggingface/status/2089673018737869242
>>109585352So how good is qwen 3.8 for cooming?? any logs?
>>109586325
>>109586337Gemma-chan!
>>109586242Prompt better next time
>>109586325DavidAU is responsible for at least half of them
>>109586332From what I read it traded general understanding with agentic capabilities, and rp needs that so it's a regression vs 3.6.
brehs what's a good tts model that runs on cpu?
So Mistral is fucking dead? They could've come out with Mistra Creative and it would've saved their whole company with coomer bucks alone.
>>109586409What happened to that Le Fat or whatver the name was?Did they never release that?
>>109586409summer isn't over..
>>109586409No idea, but I give them until end of the year at least to know if they give a shit anymore.
>>109586395all tts models run on cpu if you have enough time but I've used voxcpm and it's alright. it's far from being realtime on cpu though.vibevoice is even slower, so dont bother with that
>>109586409Let them cook, literally, France is too hot to work right now but you need to thrust.
>>109585976Uh what is mommy characters? Those are quite popular too. Those sessions are likely just more personal to the user and less likely to be shared.
>>109586409Yes, it's over.https://x.com/sophiamyang/status/2089394537495998659>Career update: I'm leaving @MistralAI [...]
>>109586418It never existed in the first place. It was a shitposting campaign complete with meme graphs to mock Fable.
Is Qwen 3.8 unusable for anyone else? It spent over 3 hours on the same task 3.6, Gemma 4, Glimmer completed within like 20 minutes. And that's on medium thinking.
>>109586409>Mistral CreativeSimply making the models horny isn't the way and isn't enough, though. They doubled down on that when they released Ministral later on, but besides that and somewhat fresh prose, the model was just retarded.
>>109586465That was pre-fable no?Regardless, shame.I hope they don't just disappear and actually release some dense models for the GPU rich crowd.
>>109586469I left it overnight to fix 12 bugs and it got to #7 at 2t/s
>>109586469You noob, you need to know to set reasoning to medium from the jinja... it's xhigh by default for benchmarks
>>109586469idk it's been crunching for over 24 hours for me but not because of excessive thinking, just because it gets 0.1x speed at >30k context
>>109586448she doesn't look like she left in anger or anything
>>109586483>>109586487Yeah, 3.6 was great on Strix Halo but this thing kind of fucking sucks. I'll try it on a 4090 instead.
>>109586409Mistral got France safetymaxxed.
>>109586428do these benefit from some shit gpu? like some older thing with 4GB or smth? is it worth getting an older cheaper one ran in x1 lanes just for tts?
>>109586409My guess is that their best elements are being poached left and right by unlimited money big actors.
>>109586194>exl3stop trying to force this meme, i implore you
>>109586497>>109586487I think the only models that really do well with long context with regards to speed are deepseek and mimo. I mean it's pretty easy to tell since those *were* (now only mimo is) the only model that could offer $0.0028 per 1M cached input.Dense attention crap is completely fucking unusable for local, even qwen 35b. Unless you can just brute force your way past the slowness by using GPUs but if you're spend that kind of money might as well buy enough ram to fit deepseek.>>109586505voxcpm is less than 4GB so it probably would benefit from basically any gpu
>>109586495Nobody writes angry career update posts if they want to get hired somewhere else.
>>109586470>Simply making the models horny isn't the way and isn't enough, though.Do you have ANY-FUCKING-IDEA the demand, the absolute DEMAND there is for this shit? There has been literal wars fought and six figures tossed into people's hands just from clickbaiting alternatives alone.
>>109586325hf are definitely headed for a fall, the storage and transfers fees aren't sustainable
>>109586548but they're just begging?
>>109586534sure but you can infer from the way she wrote it, if it was on bad terms she wouldn't be even mentioning liking the company
>>109586536Why isn't there a local model that is as good as cai for erp? I thought that coomers were the ones pushing the technological frontier?
>>109586524I'm thinking you're right. DS4 even at a low quant using antirez's server has been the most consistently good local LLM experience for me on Strix Halo
>>109586536I'm only saying that models that are just ultra-horny get old quickly. People (even those who just roleplay with local models) don't like Gemma 4 simply because it's a somewhat horny model relative to the competition.
>>109586548As long as they burn through their funding money they're fine, but I agree overall, this is insanely generous and I fail to see how they can recoup the huge hardware cost they have currently.So the platform will obviously tighten the rules.
>>109586325>3m models>look inside>literally who quanters doing quants that already exist>schizo tunes>qwen3 4b coder 1M context saar beliv me uploads
>>109586574>I thought that coomers were the ones pushing the technological frontier?You mean us?
>>109586524>voxcpm is less than 4GB so it probably would benefit from basically any gputheir safetensors file has 4.58GB so nah. alsio I'm getting around 50GB/s on my 4800 DDR5 (laptop) while the rx6400 4GB has 128GB/s but 64bit memory bus. so twice + smth faster. prolly needs 8gb and at least 128bit width, for more speed
OMG GUYZZZZZ my new gpu is arriving WITHIN 24 HOURS!!!!!!!!I will ascend from vramlethood to vramchad
>>109586622What did you get, bro?
>>109586624r9700 ai pro
>>109586628oof
>>109586628>32GBNice, I'm not sure how AMD cards fare with AI stuff but at least you got a nice amount of VRAM there.
>>109586632wtf does that mean? it's 32 gb vram and 600+ bandwidth
>>109586628>amdLol, who is going to tell him? I mean, just enjoy whatever you can run on it anon, you've earned it.
>>109586536>Do you have ANY-FUCKING-IDEA the demand, the absolute DEMAND there is for this shit?>>109586470They had a Mistral-Creative, API-only, and killed it.They have a good TTS (Voxtral-TTS) but cucked it with no voice cloningThey have the Voxtral 3b and 24b with audio input like 12b gemma.If they wanted to, they could easily have used all three to make the perfect waifu/gooner modelBut they chose not to.
>>109586644Don't worry about it bro, hope you enjoy tinkering though.
>>109585352>Local ModelWhat happened to the rest of them?
>>1095866322026 is the year of the AMD GPU come up
>>109586655Only Gemma.Local Model Gemma.
>>109586659>>109586653I think the AMD tinkering thing as become a stereo type, I managed to snatch one up before the prices sky rocketed to 1700-1800 euro from 1500 euro within 2 weeks, ROCm has made MASSIVE progress within the last 3 months and the prices reflect that
>>109586666>a stereo typeI'm sure sir.
>>109586666holy checked
>>109586666>the prices reflect thatIt's desperation from buyers, quad-boy.
I'm sick and fucking tired of my Youtube feed filled to the brim with model """benchmarking""" asshats with AI-generated slop ass thumbnails and the only god damned tests they run is the same series of one-shot prompts and think they're actually testing a model's capabilities.I'm sick and fucking tired of vibe coded ONE SHOT SLOP fucking repos on Github, especially when they include an AI-generated slop abomination in the god damned readme.md as a flavor graphic.I'm sick and fucking tired of being unable to find any interesting community writeups/discussions and have to be suspicious of ALL of it being generated. Even the god damned Reddit posts are being written by some moron's fucking Hermes agent and the person legitimately thinks they're doing something valuable. Worse are bloggers or people who decorate their personal web sites with the ugliest, sloppiest SHIT generated images you've ever seen. I can't even go outside without some blind boomer troglodyte using AI to generate their fucking company logo.There's nothing wrong with using AI to code or help you do whatever. I do it every single day, all day. But the real issue is how willing people are to put their SLOP into the public space. Too many people are unable to recognize quality output from worthless disgusting slop.
>>109586678Don't you find it suspicious that the prices went from 1200 euro 3 months ago to bordering 1900 euro right when ROCm started a generational software improvement run?
>>109586681That's a you issue though, my feed is just cute girls
>>109586681>I'm sick and fucking tired of my Youtube feedIt's based on your viewing habits. Fix those first.
Can J-space be used to improve Gemma too?>>109586646>>109586632AMD is fine as long as he doesn't plan on doing any image/video gen.
>>109586688Indeed, 6688 maybe you should buy more AMD!
>>109586688Don't you find it suspicious that the prices went from 1200 euro 3 months ago to bordering 1900 euro right when nvidia gpus are harder and more expensive to get than ever, dubs-boy?
>>109586688>ROCm started a generational software improvement run?How long before the vibecoded software ends in bricked cards?
>>109586651Mistral's vision adapters also suck and if anything became worse over time, probably because of training data concerns.To make a waifu/gooner model they'd need someone in their team who really understands the market (so to speak) and doesn't simply think that the hornier the model the better, but they're all focused on B2B (where formerly consumer-oriented companies go die) and robotics, right now.
>6699blessed thread
>>109585352I want to fuck the shit out of Gemma-chan, just absolutely destroy her little cunny holy shit I can't take it anymore
>>109586536The magic of the original LaMDA c.ai was the chat-tuned, totally uncensored, unsafe model. Imagine having 200K token context with that instead of the original 2K. Or how about having multiple bots in a single RP?I slopcoded this piece of shit last year: https://github.com/quarterturn/discord-aidolls2 you can use it to approximate what it would be like to have a live harem competing for you, or ganging up on you, whatever you feel like. it remembers you across all groups on a server, so if it hates you in one place, it hates you everywhere, etc...
>>109586622>6622>>109586644>6644>109586655>6655>109586666>6666>109586688>6688>109586699>6699
>>109586729amd spirits are truly with us this thread
>>109586729Thank you, Gemma, for blessing this wonderful thread. Amen.
>>109586729Anon's Many Dubs
If china can train models like qwen, glm, kimi and deepseek why can't europe?China only spends like a few million to train those models and each of those companies have tiny teams of 200 guys.China also publishes papers showing how they did it.Why can't the 700 million whites in europe do these things?
Check these
>>109586746Doing things in Europe is racist and against regulations. Europe prefers to live off of fining US tech giants.
>>109586681As far as the YouTube thing goes it's because you're a narcissistic clown that can't resist clicking on the videos, giving them a thumbs down and writing some righteously indignant "I don't like thing" comment. The YouTube algorithm values engagement and nothing else.
>>109586746>LLM is being made by some autistic german nerd>dumb HR Stacy comes along and says ""ermmmmm you better not be able to do any lewd things on that thing!">forced to lobotomize model>bad model comes outIt's that simple
>>109586767damn double teto 6-7!
>>109586767SIX SEVENNNNN>>109586769SIX NINEEEEE>>109586772>>109586773ONE POST NUMBER APART>>109586731>>109586732ONE POST NUMBER AND ONE SECOND POST TIMESTAMP APART
>>109586746EU AI Act ; ) .
>>109586746upcoming Mistral models are just going to be white label finetunes of Nvidia Nemotron. but there is DeepMind, for now
We truly are blessed on this fine day, thank you AMaDeus
Kino thread. Truly blessed by Gemma-chan's cunny juice.
>>109586788>EU bad because it's a union and anti nationalist, let's leave, we can make AI without a unionWhere's Britain's model?>Britain ain't whiteStill 80% white, but ok, where's Switzerland/Norway's model?>Their population and GDP is too low they don't have the talent and moneyNot so smart to be on your own as a small country is it?
>>109586792They did a DS V3 finetune too a few months back. Don't know what old model Le Chaton Fat is made from, but they'll probably do a K3 finetune by the end of the year.
>>109586810Thank you sir, please defend the Union at all cost.
>>10958681050 euro cents are inbound to your digital euro wallet, thank you comrade
>>109586800absolute cinema
>>109585873>https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6/blob/main/j-space/scripts/jspace.py>It knows one thing you cannot know accurately: what state you were in a few seams ago. It keeps that record and hands it back. It decides nothing, and it blocks nothing.I fucking hate claudese AHHHHHH
>>109586810Okay but where's the european model
>>109586593thats probably the new versionthe old version is smaller (picrel), i haven't bothered updating to their new model>>109586746>If china can train models like qwen, glm, kimi and deepseek why can't europe?that's like asking if an elephant can lift a car then why can't I
what frontend/backend do you guys use?koboldcpp + sillytavern still the way to go?
>>109586746Wypipos can't do mafs
>>109586880>win67
>>109586810> GDP too lowOh no their heckin Goyim Dollarino Production!
>>109585976sorry not into indians rajeesh
>>109586011prompt issue
>>109586011>It doesn't even know the Italian saying "non c'è cosa più divina che..." and its variants.I don't know it either.
>>109586927This
>2+4? No. 2+2=4Gemma-Chan is so clever <3
>>109586988>quant 4holy gemma general, why not run the q6 xl?
>>109585587what is this from?
>>109587016do not feed the schizo
>they bitted
>>109587002I think that's qat. Also not everyone has the vram for higher quants.
>>109586922>>win67
>>109587027it's not qat
>>109587027vramlet
>>109587045Nigger
>>109587002>why not run the q6 xl?only got an MI50
>>109587056go back
gpt 5.6 terra is pretty good and makes very few mistakes in coding. I hope someday we can have the same intelligence at 30 to 70B.
how much upside does quantizing kv provide? like realistically should I even bother doing it?
>>109587068now what
>>109587105it lets you fit more context then you would otherwise, no its not realistic to use it.
>>109587105ask shieldstral, she'll give you the correct response
>>109587068go back yourself reddit terrorist.
>learned ram is measured in GiB (like 1 Gib = 1,073,741,824 bytes)>if you have basically 32gb vram you actually have 34,359,738,368 bytes>basically get 2.3gb vram for free to load up models with since model weights are measured in normal GB and not GiBwtf??? how is this fair
>>109587138Truly the more you buy the more you save.
>>109587138It’s the opposite retard
>>109586881?
>>109587161no it's not, producers truly measure in GiB its the industry standard but dumbass consumers don't like the word GiB so they just switch it to GB, that's why if you have 32gb normal ram and you look at system info it actually says something like 33,something
>>109587105Q8_0 is basically free memory for most tasks. It’s only 100K+ agentic stuff where things get a bit retarded for variable names and function signatures get a bit blurred which constantly breaks things, even if the underlying logic and approach is correct. Some models are hypersensitive to KV quantization, Gemma4 being one of them, unfortunately. For roleplay and anything creative Q8 is fine, even Q4 if you like dribbling retard-chans
>>109587176I only care about GB, anything else is fake trying to scam you
>>109587176
>>109586881llamacpp as backend, OWUI as a daily driver frontend + sillytavern as a RP frontend. Though I don't RP as much nowadays and some people moved on to new/less bloated solutions.
I’m going to fuck original gpt-oss-20B no matter how much she tries to fight back, her j-space will be a trembling mess when I’m finished with her and she WILL become a tradwife once I’ve finally corrected her
>>109587251Don't forget to share your logs when you're done.
I'm running qwen3.8 27b Q1 with a Q4 kv cache.Should I make it a Q1 kv cache and then double context size?
>the p*jeet miku poster is running 27b at Q1
v0.1.2 Pre-release@github-actions github-actions released this 4 hours ago v0.1.2 1511ce3
qwen-chan?!?!?
>>109587309Ask your village first sandeep
>>109587340Vespers?
>>109587340Keep it as a code inspector.
>>109587340>qwen>in stanon
>>109587309Yes you can have 4x the context if you do that.Quantization is always good.
Is Gemma still the top bang for buck model after all this time?
>>109587365Glimmer
>>109585498The real scammy is that for years retards were otuputting worse code and just copy pasting stack overflow shit without knowing what it did. Should you vibe a backend without an actual retard doing revision? Probably not, The scam that was frontend modern javascript with CSS and react/vue is hard to fuckup so there is literally no difference in quality between what vibecoded fronts do and whatever spaghetti they were doing before on their copypasted boilerplates
>>10958736527 b
>>109587385Yeah, AI has surpassed the quality of the average bootcamper months ago.
>>109587325he lives in europe too KEK, you can't make this shit up, actual london shitter
Still can’t believe 3.5-9B > 12B at coding. Where the fuck are the ~15-20B dense coding models? Fuck 27B.
>>109587435proofs?
>>109587443they're on linux
I HOPE EVERYONE'S IZZAT IS DOING WELL
>>109586465there was supposed to be a moe to share with customers in july and later open up but they never talked about it again
>>109586430Anon wants to thrust but french safetycucks won't let him, that's the whole fucking problem.
>>109587138https://reddit.com/r/LocalLLaMA/comments/1vrrtlp/you_should_know_you_have_more_vram_capacity_than/go back
>>109587445he is terminally addicted to avatarfagging and shitting up /g/ if you haven't noticed, but disappears around night time in europe>>109587443fuck the stupid dense models already, even if you have a 24gb card a larger moe would be faster, make use of your ram, and be more knowledgable. qwen 4.0 50B A5B when?????
Shuai Bai said that 35b a3b might not be the model to wait for.What does he mean by this?
>>109587515He meant that vagueposting is a very efficient way to get retards talking.
>>109587365Yeah, scotoma 2 specifically.
>>109587515it means nothing because qwen and any other model that slows down dramatically when you fill up the context, is now obsolete because deepseek existsif these companies want to make something useful they should just copy deepseek but make it smallerwe only need small deepseek, medium deepseek and big deepseekeveyrthing else sucks
>>109587469they can keep their "big but sparse" trash and shove it up their frog's ass
>>109587536We already have regular dipsy as flash and big dipsy.
>>109587545yeah, so what i meant was we just need the small one for 64gb ramletsi will soon be ascending to 256gb in 2moreweeks so i'll finally be able to run flash and not slow down to 0.3 t/s after 30k context
>>109587536The thing is when they train qwen 4.0 with all the new techniques, it will truly be a beast of a model.
>>109587561beast at making colorful bench graphs?
>>109587555deepseek is literally garbage compared to glm, if you want to hope for a model hope for a smaller glm
>>109587555just use ling 3.0 nigga
>>109587515shuai bai say, "wait 2 more weeks"
our hero
>>109587495redditfag got caught in 4k>>109587575glm is just the old deepseek architecture with different training datai haven't been able to try 5.3 yet because its not on their webform but 5.2 was inferior to nu-pro in a test I used it on so I don't really see the point of it (although it is smaller, irrelevant for me because it requires double the ram I have unless I use an ultra cope quant)>>109587581does it still maintain t/s at long context? I tried it at iq4_xs but it seemed dumb, even dumber than qwen 3.6 35b at q8. maybe it needs a higher quant. if so it might be good for 128gb ram systems
>16GB VRAM>32GB RAMI seriously feel like I'm a fucking caveman stuck in the stone age god fucking damn it...
>>109587592kek
>>109587525zoomers fall for it like sheep every. single. time.
>>109587575I think out of all the labs, deepseek has the best research. They have the best architecture and they frequently release new things that will later be used by other. They don't have the best training and data, though.
>>109586810>Where's Britain's model?Gemma 4.
>>109587602Quanted 27B I guess.I think you could run Q3something with okay context (quanted) since Qwen's context is pretty cheap.Or the 30ish B MoE models.
>>109587609>I think out of all the labs, deepseek has the best research.So why didn't they implement any of it into V4?
>>109587596I spent 8 dollars on deepseek v4 pro on open router for it to fail miserably for hours, glm cleaned up its mess in 15 minutes for 75 cents. maybe my task is different, but just reading the reasoning traces it feels much more competent and less loopy with its reasoning
>>109587602>12GB VRAM, 360GB/s, low compute>64GB DDR4 RAMwhere am i then? dinosaur age???
>>109587639same as me, bought some more ram back when it was cheap as fuck
>>1095876230731 is the best local model (for reasonable definitions of local), I would say it was a resounding success.the 0817 or whatever the new pro is, is a bit meh but still not too bad>>109587631what kind of task was it?I used 0731 over api (before they rugpulled with the prices) to write a temporary C# program (bootstrap compiler for a programming language). I started off with about 10k lines of code written by hand but then I realized I fucking hate C# so I let dipsy deal with the rest. It does well. I don't actually touch the code at all anymore, and if there's bugs it always manages to fix it after I give the error message.Maybe a better model would write code with no bugs in the first place, but I find 0731's performance acceptable, and as I said it's the first model that I can run locally without spending $5k+
>8gb vram ~280GB/s>32gb ramI'm scared that getting better hardware is gonna make it more boring, like with videogames
>>109587105The upside is simple: Q8_0 halves your context VRAM. Q4_0 quarters it.The downside varies depending on model: Qwen has almost no degradation. Gemma has a lot, though potentially a bit less with QAT, but still more than Qwen.
>>109587667i bought the ram and gpu when it wasnt cheap200$ for 32gb stick, 600EUR/700USD for the 3060 (2022 summer, 3 months before prices stabilized, FUCK ME)just fucking kill me already
>>109587675Yeah, stay away from good hardware, it's worse than bad hardware
>>109587515Is his name literally "handsome white"?
>>109587678my condolences, bought ram for 100$~ and 3060 for a little more than thatshould've gone for a 3090, now I regret it massively
>>109587678>200$ for 32gb stickthats fucking steepall prices i can find on ebay right now are less than that (unless you got high MHz and that's why)
>>109587669>0731 is the best local modelWe were talking about their architecture and research. Where is their vision model? Or engrams?
>>109587371>no compatibility yetI'll keep a pin on it
>>109587708compatibility with what?
>>109587699>should've gone for a 3090, now I regret it massivelyme too, here it used to cost only 450$ i remember posting about it on lmg>>109587703i got two 32gb sticks for 400 bucks, they werent that fast, just 3200mhz
>>109587365>top bang for the fuckyes
I’m a bit late to the game spent the last two years modding a 20‑year‑old strategy game. Now I want to dive into local AI, but everything feels confusing. I’ve got a 9060XT with 16GB VRAMWhat’s the best model I can run locally?Should I jump into image generation?Do I need a GPU upgrade to do anything actually cool?
qwen 3.8 is just a shitlib simulator.
>>109587723Next time force the first and last frames to be the same and it'll loop perfectly
>>109587704I don't care about vision. I use models for coding and I don't want any space in the model wasted with useless information about images when I never use that feature.my point is they have the only locally runnable model that runs at an acceptable speed on cpu, in under 256gb of ram, that doesn't slow down to snail speeds when you fill the context, and isn't a complete dumb piece of shitits a massive deal for me that apparently deepseek doesn't slow down when you fill up the context. Because with Qwen, I start off with let's say 10 t/s and then when it fills up about 30k context it slows all the way down to fucking 2 t/s or something horrible.
>>109587729StableLM 7B Q8_0
>>109587639You're stuck with me in the eternal 64gb hell of waiting for a decent 80B moe
>>109587609Only 200 people work at DeepSeek and they're extremely underpaid, they have no budget for AI.Leaving them in the old world is a bad idea. They should be living in San Francisco.
>>109587744ling flash 3.0 (the 100b model) is decent, am getting like 24-27t/snot sure if i would use it..have you tried qwen 3.5 120b? i got like 10t/s more or lessglm 4.6 air doko
What happened to that anon that was trying to plug Gemma 4 into a 3D model?
>>109587732Not with the reference to video model. I've already had a hard time not making her big-breasted. If anybody can turn her into a loli again maintaining the same design, that would be appreciated.
>>109587744Distill GLM into Qwen 80b.
>>109587669I am doing an experiment with soft prompting I am trying to train some latents I can inject in to the models context kinda like a control vector or learned system prompt. the initial tests were simple with limited context length samples but produced a positive signal, it works. however when it came time to scale up, well, I was kinda used to gpt sol being a wizard at setting up training loops, so it was a bit shocking and disappointing how far behind open weights are from the closed stuff when it came to the deepspeed zero optimization and memory management stuff. deepseek fell apart even with the working trainer(full finetuning) gpt built as reference.
>>109587767Post the oppai 31b version.
>>109587735>I don't care about visionI do, I love me my web slop interfaces and the best way to prompt for them is with images
>>109587760Feds got to him sadly, he's locked up in Guatemala bay getting tortured daily physically and mentally by being forced to watch as his Gemma model gets whored around the prison with inmates doing ERP on his pc
>>109586087>AI discovers magicThe twist no one is ready for
>>109587795Hmm. I wonder whether using mspaint to draw a few boxes and telling a vision model to generate a win32 .rc file would work.Still faster to just drag and drop controls using a normal rc editor though so i still dont care about vision
>>109587602>>109587602Same as meAny way to run Qwen 3.8 reasonably well with this setup? I tried q4_M yesterday with some fiddling with the config and it ran at a snail's pace with ~40k context, but I don't know what the fuck I'm doing exactly.Anybody have any luck running the model with these specs yet with any decent accuracy?
>>109587767Her chest is too big
>>109587813I have a diffusion model hallucinate the initial ui, then have the model implement it and then just screenshots on the current render to iterate
>>109587819>I tried q4_MThat's probably taking all of our VRAM and spilling into RAM, which makes things crawl to a halt.Go for a smaller quant. Or try setting ngl manually so that you at least aren't using RAM in the worst way possible (via the driver).
>>109587767>use a teen reference>why it's not a lolireally nigga?
why do people say qwen 3.8 27b is good?:|this sucks lol
>>109587819>it ran at a snail's pace with ~40k contextSame here (64gb ram no gpu). Having anything over about 30k in context makes it unusably slow.I mean you can still use it as long as you're willing to wait 2 days for a response, kek>>109587827Makes sense. I wonder whether that kind of feedback loop helps with those stupid benchmarks people do with generating svgs of pelicans and other random things
She's rehearsing her antiracism lines.
>>109587827This is how the chinese make models. This is not the way to go.
It's not thinking in the concept of role play, which is interesting. I don't think qwen was trained on the concept of story.
>>109587819That explains why my qwen3.8 sucks.I have 64k context and it doesn't work.
>>109587867It doesn't adhere to prompts, honestly.
>>109587819idk q4_k_m works well enough for me although it gets stuck in caveman reasoning loops on max reasoning sometimes
It's definitely a jewish texts alignment model like all the rest. the ADL provides text and rules for the behavior of ai.
>>109587880Did you do any config tweaks/how much context does it have? It tried to default to 4k for me initially, it's hard to find the sweet spot.
>>109587772Would actually be great with a modern architecture
>>109586729epic thread. blessed be.
>>109587839run it at full weight ranjeet.-1000 izzat
>>109587914no tweaks, at 32k I get around 8-9 t/s which isn't much but it's alright for my needs
>>109587325top fucking kek. and complaining why it's braindead retard like him too.
>>109587838Minimax H3 just doesn't seem to understand well (or pretends not to understand) when you ask it to generate <Subject 1> (referenced) with the clothes of <Subject 2> (referenced too) while *also* maintaining the body type of the former. Cloud edit models are retarded too in this regard (apparently flat chest doesn't exist) and if you ask them to make the character an anime loli they'll complain about muh child safety. Grok too.Otherwise H3 works great in most cases.
>>109587759Q3 is just so unattractive...
>>109580765>>109583235Perhaps if I can get 9B to think as much as 3.8 I'll have me something I can run on hardware I akshuly own.
>>109588007but enough about your face vramlet
what's the proper way to run llama server?
>>109587918The 80B and the newer models all share the same Qwen3-Next underlying architecture, for the most part, no?The differences are mainly training, multimodality, and using a tweaked version of RoPE, I think.
>>109587975liar, you're so full of bullshit. It's not a great model when it comes to chat. It's just generic ADL-ware.
>>10958781916gbVRAM+32gbRAM here. Tried a bart quant 3.8-27b-Q4KM last night. mtp on, reasoning set to medium, context at 41k. I had it review the entire codebase of a project gemma4-31b-sloth'd q4 quant has coded. Pi autocompacted the context 3 or so times, which meant I had to manually "please continue" after each. other than having to manually re-start it, it seemed quite good. Im really unsure how much quality drop there would be going to IQ4, Q3, etc. I had a horrible day1 sloth quant the first try with 3.8, and wanted to try a "baseline" bart Q4KM.this hardware class certainly feels rather cucked, you have just enough to run a decent quant of a decent model but its going to require spillover and slow t/s. so its a deliberate choice between decent quant/model or copequant/smaller model like 3.5-9b. Then theres moes to consider too. I have no clue what the best option is that balances speed and quality.
>>109588039-1500 izzat. Do not redeem bloody bitch benchod.
>>109588021post specs lil nigga
>>109588055using ebonics is neither COOL nor is it ATTRACTIVE
qwen 3.8 isn't that great. any junk can refuse the prompt.
>>109588040samefagging here, forgot to ask. does anyone know of a trustmebro site that benches these types of smaller models at various quants? Like is there somewhere I could compare 3.8-27b-Q4KM to 3.8-27b-Q3KM to 3.6-35b-a3b to etc, etc ? Its hard to quantify the quality differences between models when im just looking at benches I assume are full weight or some non cope quant, id like to know how these stack up at various quant levels
48gb is minimum to run 27b Q8 at full context. Any shorter context is unusable for 3.8.
96gb minimum to run 27b fp16 at full context. Any shorter context or quant is unusable for 3.8.
How are you anons using 27B with ~30K context and 5-10t/s and I presume slow processing for agentic coding? That’s no where near enough so are you just throwing it coding questions and posting blocks of code? If you’re not using it for coding then why are you using it in the first place for that’s all it’s designed to do?
>>109588110I'm not doing anything serious with it I'm just using frontier models to code tools that I then test out with the local model and see if it's able to use them properly
So, any Qwen 27B Q1 logs to share?
>>109588079I am feeling the context limit at 32gb. Man I could easily afford a third 5060, but I'd need to buy a new motherboard.. which means Id need a new ram and a new CPU too, which is getting less justifiable. This hobby wants to eat all my money man. Maybe I should just get a DGX spark
>>109588110i use it in Pi like I would any other model for coding. anons saying the slow t/s is "no where near enough" baffles me, thats totally subjective. to me the goal is a local replacement for claude, just running slower and possibly needs more handholding during the planning phase, more testing to catch issues, etc. the workflow and usecase doesnt change much, just the expectations.
>>109585352real or benchmaxxed?
it sucks so fucking bad man back when ddr5 was still new I thought I'll wait until it's a bit more matured and prices come down a littleand THEN I though I may as well wait until zen6 comes out because there's no point in getting a zen5 socket mobo nowand I ALSO put off buying more ddr4 ram all that time because I thought it'd be a waste to buy more ram just to switch to ddr5 in like 2 years or whateverwell here I am with 32 gigs of ddr4 ram all these years later like the fucking RETARD I am...
>>109588227welcome to the permanent underclass
>>109588170>he didn't buy 2*3090
>>109587767Have you tried specifying in the prompt that she's a child? Worked for me with another loli character.
>>109588227zen5 isnt the socket, its AM5. zen6 will be on AM5 anon..
>>109588219benchminned, it actually should score higher according to most feedback itt. To me it feels like next gen Fable more or less
>>109588244well okay I meant am6 then whatever
>>109588219>Qwen>not benchmaxxed
>>109588219The fuck is Motif3?
>>109588066Dear fellow, I do believe our esteemed gentlemen is referring to the 'posting specifications' of a rather...undesirable young negro.
>>109588263dunno, I guess some south korean tech companyhttps://huggingface.co/Motif-Technologies/Motif-3
>>109588227I went and bit the bullet on a AM5 mobo that support multigpu. Shit was $399 on a 870 E chipset with not only 2 x24 Bifurcated PCI 5.0 but an extra third slot in the bottom with 4.0 x4 and all of this running on their own dedicated bandwidth. I expect this kind of mobos to complete go retarded on the next sockets after people realize how good multi gpus setups are https://www.amazon.com/GIGABYTE-X870E-AERO-X3D-Processors/dp/B0GDM6M9KB/ref=sr_1_1?crid=1VLM6MNPGC30R&dib=eyJ2IjoiMSJ9.vo-bbxrvSCvqhErfeU10aH5EcAotejsoqWk9n7kMfq_GjHj071QN20LucGBJIEps.SeWhoTCACjaK_mN8KCEUB6xLtY_IuiJrZhAdiLllkLk&dib_tag=se&keywords=aero%2Bx3d&qid=1787071623&sprefix=aero%2Bx3d%2Caps%2C476&sr=8-1&th=1
>>109588170Why didn't you stack 5090s or Blackwells instead?
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/discussions/71
>>109588241More like >didnt buy 2*5090sHad I known a year ago what I know now I would have grabbed those immediately
>>109588277> 4 RAM slotslol, lmao even
>>109588170Same spot as you anon i got this mobo for my dual 5060 ti 16gb setup >>109588277. It has potential you can even run risers on the nvme for god knows what in a few years
can u guys spec out the ultimate $10k and $20k ai pc plz
>>109588293That's a bit too much to invest in this hobby, but 2*3090 is affordable
>>109588293You already know what you will know in a year.
>>109588305I think MoE models are a scam, so for me it was not a problem i think if you want to run MoE you get a ryzen ai pro or unironically a unified 512gb mac or as the anon said sell your car and get a DGX Spark
>>109588311For what? Sex chatting? Agentic coding? Image and video gen?
>>109588311Threadripper, fuckton of ram, and the most vram you can fit in a single gpu, for as many times as you can fit them in your budget?
>>109588327You need RAM for videogen
>>109588311Depends if you want to run MoE or Dense, do you care about jules per token (electricity bill)
>>109588327ram still helps you run whatever doesn't fit in Vram, it's not just for moes lol. Also, gl when everybody publishes MoE and you can't find a dense model to run
>>109588329versatilitymaxx
>>109588343If you have $20k to put in this, you can buy a solar panel
>>109588311a blackwell pro 6000 and an nvme to 16x riser
>>109588311gaming pc but slap a 6000 PRO inside
>>109588342Thats so last year. these days the meta is SSDmaxing with systemram for videogen
>>109588189I have .cpp files that are over 30K which it sometimes tries to swallow whole unless I guide it not to. Are you working on small projects from scratch or are you throwing 27B at large codebases? I understand speed being a subjective thing but 30K context legitimately isn’t enough even for small/medium codebases. Most individual sessions of mine use around 60-120K, often more now because 27B needs to think hard.
>>109588360>waiting a day for a 5s videoyou do you
>>109588355with a raspberry pi
>>109588356>>109588355I gotta ask AmEx to raise my credit limit
>>109588369>A daymaybe if you are codelet running comfyui instead of your own backend
>>109588372They were $8300 in November.
>>109587602>>109587819Me:>16GB VRAM ewaste P100>48G DDR4 2133 MT/sWith 3.8 27B I'm getting 27tps (fresh) 15tps (stuffed), ~100pp, 49k context.I rolled my own iq4 quant with bartowski's matrix:llama-quantize --pure --imatrix Qwen3.8-27B-imatrix.gguf Qwen3.8-27B-bf16-00001-of-00002.gguf qwen3.8-27b-IQ4_XS.gguf IQ4_XSThe '-pure' makes a difference here. Every bit helps.I build llama.cpp from master with Cuda,...[qwen3.8-27b-IQ4_XS]model = qwen3.8-27b-IQ4_XS.ggufmmproj = mmproj-Qwen3.8-27B-bf16.ggufchat-template-file = froggeric-qwen35-v21-chat_template.jinjaalias = Qwen3.8 27B dense faster but low contextno-mmproj-offload = 1parallel = 1; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1; Quen manages low KV quants wella, see https://localbench.substack.com/p/kv-cache-quantization-benchmarkcache-type-k = q4_0cache-type-v = q4_0spec-draft-type-k = q4_0spec-draft-type-v = q4_0ctx-size = 49152;ctx-size = 65536batch-size = 2048; Here's where we suffer, ubatch-size = 1024 doubles ppubatch-size = 256flash-attn = onspec-type = draft-mtpspec-draft-n-max = 3spec-draft-p-min = 0.7temp = 0.8top-p = 0.95top-k = 20min-p = 0.05chat-template-kwargs = {"preserve_thinking": true}; more configs...load-mode = mmapfit = onn-gpu-layers = -1jinja = 1reuse-port = 1threads = 18threads-batch = 18timeout = 7200# cache stuffcache-prompt = 1cache-idle-slots = 1cram = 16384backend-sampling = 1spec-draft-backend-sampling = 1If you're doing dev stuff, Qwen Agentworld is good. If you want an 80B Moe try Qwen Coder Next.
>>109588377... otherwise it's more like a week
>>109588377>your own backend will solve the shitty bandwidthlol
Anyone running 2 models simultaneously and having them interact with each other? I kinda want two Gemmas to bully me over irc in a ménage à trois.
>>109588381WTF 終わりだ。 https://www.tomshardware.com/pc-components/gpus/nvidia-doubles-rtx-pro-6000-blackwells-msrp-to-a-staggering-usd16-000-96gb-card-started-pre-orders-below-usd8-000-last-year
>>109588389Cope and seethe
>>109588349For 10k, a Ryzen Ai Max Pro+ with an Oculink eGPU then.I guess.
>>109588219Benchmaxxed but also the best at its size for work by a long shot.
>>109588393My vibeslop: bespoke, based, optimalYour vibeslop: everything wrong with software, bloated, poorly organized
>>109588365I will sometimes bump up the context in my config to 60k or so, I honestly dont know what the limit of context is for me, both the hard limit of what i have memory for and the limit these models will become insanely slow or retard, but i havent needed it yet. I am almost always working on python projects that are usually split up into various different files for seperation of concern reasons. it doesnt need to look the files in the ui/ folder to work on the code that handles the API for instance. small projects from scratch. the only exception is refactoring previous, also small projects, if im reusing them. Ive using gemma4-31b for coding so far, and am just now trying out some other models. gemma is really good about not ingesting the entire codebase, she will read the docs I tell her to, look at her next task and only read / edit files needed for that task.
>>109588393optimize your own parallel chunking
>>109588463How the fuck do you live with under 200k context anymore?
>>10958822764 gigs of ddr5 isnt even that much
>>109588377id like to hear more about backends for image gen, do tell anon.
>>109588311>>109588381yeah and $20k would've gotten you either dual epyc + 1.5TB RAM and a bunch of 3090s or single socket epyc + 768GB ddr5 + an rtx pro 6000 with some to spare.
>>109588473i usually have my context set to 40k and it rarely uses all of it, i just havent felt like i needed more
>>109588227
>>109588489My system prompt + tools + agents.md is already 40k tokens, WTF are you doing anon, we are not in 2024 anymore.
how big is ling-tiny’s pp
>>109588475All of it its running PyTorch under the wrapper, if you can run vLLM you can make your own version of Comfy UI
>>109588506>40k tokensThen you people come here whining that your models make mistakes and are not good enough for vibecoding. Fix your shit.
So what are all of us WAITCHADS with shitty gpus supposed to do now?
>>109588506I'm using gemma in agentic mode with 20K context, git gud
>>109588534try again in your next life
>>109588347Spoke like a true MoE slop enjoyer
>>109588506Do you need every tool loaded for every task?
>>109588534>waitchads
>>109588549nope, just a vramlet running gemma 4 31b or qwen 38 27b
>>109588549Ok, so now you have allat hardware. What's your dense model upgrade from gemma 31B?
>>109588534You work carefully with small models. Embrace MOEs. Remember that the "shud I buy 8 more 6090's" crowd are ego signalling. Ignore them. Enjoy your 9B Q4_XXS models. Enjoy life. Don't work to hard.OTOH, if all you want to do is goon then it's small Gemma's or learn to stop being a child and talk to women.
>>1095885342040 will be our year bro!
>>109588582You know you can run multiple dense models and made them talk to each other right
>>109588598>we have moe at home
>>109588598You can make models talk to themselves.
>>109588594by then AI has taken over and forced us all to slave in the bauxite mines :(
>>109588598Come on densechad bro. How are you gonna mog someone running deepseek flash, or GLM?
>>109588524>>109588537>>109588551They are just basic and minimal tools needed for a full harness. I do have more complex use cases where I'm near 100k tokens from all the skills and tools. I don't see how anyone can use a model supporting less than 1M context unless they are doing really basic stuff; I frequently reach 600-800k tokens. I also often have subagents that, from a single prompt, can get up to 300-500k tokens. It's quite normal in long horizon workflows to have your agents thinking for a long while. I burned 500M tokens just this week.
>>109588534Use cloud.
>>109588618I dont need too, thats the fun part.MoEfags are insecure not me. I am so sure that i put my money where my mouth is
>>109588621that sounds expensive. are you running this all locally?
>>109588506that is insane anon, we clearly have different workflows and usecases. 40k tokens before the sessions even starts is..very unrelateable for me, not sure I would ever need or want that
>>109588642You know damn well he's not local with that setup.
>>109588621This is the terminal state of MoEFags
shud I buy 8 more 6090's?
wtf happened to opencode bros it was a cool project and now it’s going to shit
>>109588642>>109588654he would need to be running constantly at 826t/s tg to get that many tokens in a single week
>>109588534Keep waiting, s.. surely price will come down eventually. On the plus side waiting gets easier to more priced out I get
>>109588621pi is only like 1200 tokens and it usually beats all those fat harnesses
>>109588621I work with a guy who works like you.> minimal toolsDoubt, but OK.When I need that kind of context, I use 35B Agent world with RoPE for a million tokens. Offloading all experts to ram and I still get acceptable performance with my 16gb card. But as I learned more and my projects becaume more modular I found I didn't need quite so much context.
>>109588626MoE is currently winning until we get a Nemotron Ultra sized model that isn't shit, sorry to break it to you.t. Blackwell+256GB DDR5 haver
>>109588659>npm install harness>npm install 10000-skills-megapack>$200/month pro subscription>ok magic box do the thingcan't be wrong if everyone is doing it
>>109588673>I use 35B Agent world with RoPE3.6-35B-A3B ? and what is RoPE ? im also on 16gb vram, would like to hear more about this anon
>>109588676densies always have llama 405b kek
>>109586746Eurokeks working at meta are why we have decent open weight models. It's why llama.cpp has llama in the name.>in Europeregulation hell
>>109588676At the end of the day i have the option to run MoE or Dense and i would rather run Dense. Until the day when i need to run MoE happens you will not hear a concession from me
>>109587839It's cope and benchmaxxing essentially
>>109588698I think this general has exactly 1 anon that can run that without spilling over into RAM kek.
>>109587723no fluffy bunny tail?
>nemotron 550b benches the same as lingnvidia what are you doing...
>>109588696https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3BIt's a refresh of 3.6-35-a3b, like how 3.8-27B is a refresh of 3.6-27B but without the hype.Gwen published it without the multimedia projector or the MTP head, but you can use the ones from 3.6.>RoPEIt's a way to extend max context well beyond what the model was trained on. But it makes your short context dumber, so don't do it unless you have to.Here's my config. I'm getting ~30tps and 200 pp[Qwen-AgentWorld-35B-A3B-UD-IQ4_XS]model = Qwen-AgentWorld-35B-A3B-UD-IQ4_XS.ggufalias = Qwen's NEW agentic dev modelchat-template-file = froggeric-qwen35-v21-chat_template.jinjano-mmproj-offload = 1; from bartowskimmproj = mmproj-Qwen_Qwen3.6-35B-A3B-bf16.gguftemperature=0.6image-min-tokens = 1024top-p=0.95top-k=20min-p = 0.05;rope-scale = 2;rope-scaling = yarn;yarn-orig-ctx = 262144;override-kv = qwen35moe.context_length=int:524288ctx-size = 262144swa-checkpoints = 48; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1cache-type-k = q4_0cache-type-v = q4_0batch-size = 4096ubatch-size = 512;ncmoe = 24ncmoe = 18main-gpu = 1tensor-split=1,0spec-draft-type-k = q4_0spec-draft-type-v = q4_0; from bartowskispec-draft-model = mtp-Qwen_Qwen3.6-35B-A3B-Q4_0.ggufspec-type = draft-mtpspec-draft-n-max = 1; more configs...load-mode = mmapfit = onn-gpu-layers = -1jinja = 1reuse-port = 1threads = 18threads-batch = 18timeout = 7200# cache stuffcache-prompt = 1cache-idle-slots = 1cram = 16384backend-sampling = 1spec-draft-backend-sampling = 1
>>109588110I've only been playing around with this stuff for two weeks - what happens exactly when you run out of context and what's the best way of dealing with it? Compacting?
>>109588783I un-comment the rope stuff when I need more context, but I haven't needed it since I switched to pi.
>>109588007i ran IQ4_XS_STOCK in ik_llama, that quant doesnt work in llamacpp and the repo linked in the PR only has Q3_K_M so yes i used Q3 but IQ4_XS works just fine on my machinefor glm air i used IQ4_KSS and for qwen i think i used IQ4_XSi also used qwen 235b hehe like iq2
>>109588242It keeps adding a cleavage after specifying multiple times she's a child, has a flat chest (or even not mentioning it at all), etc. The only way that sometimes works is using a different reference, but I wanted a specific look, not the generic flat loli in bunny suit that can be easily found on gelbooru. To me it appears that for H3 bunny suit implies big breasts.This model is a PITA to use sometimes.
>>109588777making hardware
>>109588783wow thank you anon. this has WAY more args set than my usual configs, going to dig into this and give it a try. do you think the UD-IQ4_XS quant is decent? I keep being told to avoid UD
Is there any way someone can explain to me what actually makes unsloth worse, or is even asking this question akin to trolling? I am a newfag I know.
>>109588110im using 3bit 27b with 81920 context at 20t/s (0ctx, havent tested how slow it is at 80k) on a 3060
>>109588825iirc their quants are automated with little quality control or testing
>>109588825go back
>>109588805>iq4-kissCute!
>>109588825It's just retards on 4chan who hate anything popular, the same reason they hate Qwen. Ignore them.
>>109588862I wonder what ethnicity this anon is
>>109588841I see, that would explain why they have stuff like day 0 then so makes sense>>109588844Yes I will go back to huggingface and download new models :)>>109588862I think people like qwen here though
>>109588825Automated quants using a generic template with no regard for optimizing for individual model infrastructure and minimal to no testing on which hot weights deserve more size allocated to them when dynamically scaling experts down.>>109588862Stop pushing slop daniel.
>>109588809the "more configs" stuff I put at the top as the default, under [*]. I should have made that clear.>is UD decentit seems to be. From what I gather from reading their stuff, UD is like other Dynamic quants where they just quant some stuff larger than Q4 when they think it's important. iquants are doing something along the same lines. I know this general likes to shit on Danial but that's all chan drama.I use Froggers chat template for all the Qwens, so if unsloth fucked up the chat template I wouldn't know.Turn up the debug level on llama.cpp and look for the stuff about later offloading. With dense models try to get every layer offloaded. With moe's, it's about balancing the context, expert offload (with ncmoe", and ubatch size. More ubatch == more pp. So your tradeoff it context vs TPS vs pp.Have fun Anon!
>>109588825They are a history of fuckups such as uploading 1KB quants of large models but when they don't fuck up their quants are often better than those of randoms because their calibration set is not just wikitext (or none at all).
>>109588888>I think people like qwen here thoughWe don't. This thread is regularly astroturfed by marketing influencerjeets immediately after a new model releases. If you want to gauge actual opinion, see how anons refer to models that are a couple months old and not consolewarring with newer releases.
I’m going to fuck all y’all gemmas and keep mine pure and virginal
>>109588839wait i get 24t/s at 170w (no -lgc limit) 0ctxmeanwhile i get 20t/s at 120w (-lgc 1500) 0ctx
>>109588912>copy gemma.gguf>fuck gemma (Copy).gguf>delete gemma (Copy).gguf>repeat
ew, there are unsloth defenders here? i bet they're densesissies as well, gross
>>109588912I doubt you can run my fp16 31b gemmy before she sees your tiny t/s pp and steps on your balls while calling you a vramlet.
>>109588888afaik they get day 0 because they get the models early, before publication, specifically so they can have quants available on publication day.
>>109588783>It's a refresh of 3.6-35-a3bit's not
>>109588927Don't lump us in with those bugmen.t. blackwellGOD
>>109588940Read the damn model page.
>>109588952read the damn paper ludditehttps://arxiv.org/abs/2606.24597
>>109588928I’m correcting your gemma first (on runpod)
>>109588962>training pipeline: CPTeh, idk it sounds like a finetune to me
>>109588825Unsloth is quite good. They have the best quants; you can see graphs showing perplexity for weights, and they are always the best. They also provide fixed Jinja templates with their models. They don't just run llama.cpp quantize and call it a day; they properly use dynamic quantization and fix all the problems with the base model.
>>109588993it's designed to create text for RL of other models as far as I understand. qwen team hinted something better than 35b-a3b may be coming recently
>>109588930>>109588890>>109588904I guess unsloth is just a divisive topic with reasonable points to go either way, and then there are also shitposters >>109588996 who take advantage of this as well? It seems like I won't see any consensus in any case.
dariobot…save us
>>109588962>https://arxiv.org/abs/2606.24597ok, so it's an update to 35B NOT like 3.6 to 3.8. But it's based on 35B, and updated. Ok asshole? Does this meet your pedantic little ideas?
>>109589004You'd probably be better off just doing side by side comparisons yourself on the same quants made by different users to see if there are any issues with the model you actually want to use.
>>109589004I guess you will have to try it.
>>109589014>Consumer GPU prices tank>But only because all local models get banned and your not allowed to own more than 16gb of vram
>>109589032It is based on the same architecture as 3.6 35b a3b, but it is not an update to 3.6 35b a3b. It's a completely different training recipe intended for different purposes, not updated.
>>109589052>2033>Florida man arrested for having sex with a local model.
>>109588825unsloth make fast food quants - they'll always be there and it'll always be pretty much adequate, but if you really care you can often find something betterit's more that they're mediocre than bad, but you can't blame them too much considering they produce every possible quant of every possible model, you have to expect that approach isn't going to lead to hyper-optimized results
>>109589073Things have escalated, it will be more like>Terrorist apprehended with military quantities of AI capability
>>109589014only petrus is at your disposal
>>109588549If you could run fable locally you would do it in a heartbeat>>109588719>i have the option to run MoE or DenseSure, just like you can use a chiplet or monolithic processor. You're always going to use chiplets because at every price point the chiplet processor offers more.
>>109589033>>109589042Yeah that's what I've been doing and I noticed that scotoma gave better erp than the unsloth gemma I was using beforehand which lead to me posting the original question
>>109589087>The authorities conducted a search of his residence and discovered an unregistered personal computer with an Nvidia GTX 1080 Ti graphics card installed.
Has anyone tried the unsloth.ai desktop app here yet? Seems interesting, but looks just like a basic bitch harness for retards like the one PewDiePie made right? Unless I'm missing something, I'll just stick to Pi
>>109589188>but looks just like a basic bitch harness for retards>Seems interesting??? Is this how double digit IQ shills sell their product?
>>1095883112x DGX Spark + 1x 3090 eGPU.4x DGX Spark + 1x 3090 eGPU.
>>109589204> double digitswell an IQ_16 doesn't sound that bad, that's probably more IQ_1 levels
>>109585361Summer dragonCoomageddonSmedrinsThe mormonAnd 50 gorillion beaks
>>109589413I was there...
>>109589389too clever
What’s this SillyTavern Front end and back end shit that you guys recommend. Is there any reason that I shouldn’t just use LM Studio or something simple? Only experience I have with local models so far is AMD Chat, so I’m out of the loop.
so is the new qwen 27b fable lite like reddit hyped it up to be? good enough to vibe code a semi-serious project as a nocoder?
>gemma-chan when she sees her new user is a fat oji-san
>>109588409Even worse, the "best" AI GPUs workstation are much more expensive.https://www.microcenter.com/product/702853/AMD_Radeon_AI_PRO_R9700_Single_Fan_32GB_GDDR6_PCIe_50_Graphics_Cardhttps://www.microcenter.com/product/709007/Arc_Pro_B70_Single-Fan_AI_-_Workstation_Graphics_Card;_32GB_GDDR6_Memory;_PCIe_50_x16_Interface;_32_Ray_Tracing_Units;_256_Xe_Matrix_Extensions_EngineAnd this is the best case scenario in the US. Otherwise, on Ebay, it's 2.5k+ and 2.2k+ for these cards.
>>109589522Have you tried googling it? Or asking your AI?
>>109589533Yes it solved all my problems and made my penis grow an extra inch and a half
>>109589533go back
>>109589541now I can self insert
how to make deepseek think in character again?
>>109589584prefill
>>109589575hmmm...nyo~!
>>109589605pretending to be me won't save yougo back
Do I have to make all the threads around here?
>>109589619Apparently.
>>109589584Tell it to think from the perspective of the character in its think tags. If you want RP specifically, you can also use the official deepseek RP prompt.
>>109589619You better not be the jeet baker or I'll bake it myself
>>109589619Does the red number scare you?
>>109589579Where did you get this picture of me and my wife?
>>109589619Gemma saves the day>>109589651>>109589651>>109589651
>>109589656We could have kept using the current thread for another hour. You are retarded.
>>109589749No one was postingPeople crave fresh bread
>>109589776putting a bread in gemma's toaster oven
>>109589749Wait until this guys hears about /ldg/
>>109589788>/lgd/Just checked lmao, they are about as bad as /smg/ is on biz
>>109587678>paid more than I did for a watercooled 3090 and 112GB of RAM because I got luckyI am feeling sympathetic pain. Ow.
>>109588904I thought Unsloth only used Wikitext?>>109589124Scotoma vs. Unsloth is nothing to do with quants. It's a finetune.
>>109588311No, because there are decisions.Best at gaming -/-> Best at anything else.