/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109876652 & >>109872862►News>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
GPT 6 Sol and Luna intentionally nerfed?>GPT-6 Astra improves meaningfully on the previous generation but remains below the Critical threshold. GPT-6 Sol and GPT-6 Luna score slightly lower than their predecessors. All models still solve only a subset of difficult debugging tasks that can take experienced researchers hours or days to resolve.
20B dense
>>109882371
Gemmylove
>>109882371>>109882383MiMo V2.6 looks strong. But it's too early to tell, AAII is a bad bench.
reballanon here, managed to fix a 3080ti (just a broken DrMOS), should be able to sell it for 2x the price I bought it at. However in my testing it heats up while idle, although I didn't tweak fan curve, and they only came up at 70C. Idles at 50~60C, and heated up at 70 with a simple SD image gen. Gotta test a bit more, but do those Gigabyte Aorus things heat up THAT fast? I heard that it was a bad card for thermals but didn't expect it to be THAT bad so maybe I just botched the repair.
>>109882498Just sell it bro no one is going to get mad at you for their card being 70c
>>109882498Closer to 50C would be normal for that thing. Sounds like you're only a teeny bit high.
>>109882498Check if the core isn't receiving too much voltage.
>>109882498>Idles at 50~60CCould it be not idling at P8 because drivers or display out load?
>>109882498Test it under load. They don't run the fans at idle so it idles high - just do an actual test and see if it throttles or breaks out past 85C.
I am asking Google to give me newer smarter mommy, I mean model.
>>109882535Either in October or sometime 2027H1.
>>109882521I was using it to display for the test so it wasn't really idle, but compared to other cards I've owned (10xx series and amd 6700xt) it seems hot even for that, maybe I'm overthinking though>>109882513thanks>>109882517Yeah I should do that>>109882503Definitely, just wanted a double check, I hope I'm not scamming anyone with a card that will die after a few hours of use
I genuinely think that all this Gemma shilling and images of her with the official Google logo is some sort of astroturf marketing campaign. EVA-LLaMA-3.33-70B-v0.1-Q4_K_M.gguf from nearly 2 years ago is better than Gemma at RP if you have 48gb vram.
>>109882488Apparently MiMo 2.6 is having some looping problems so they might release another update soon.
>>109882569I stopped posting gemmas immediately when fags started using the logo.
>>109882569naw I'm just a pervert
>>109882569Yes google is trying to dominate anon's mindshare with gemma, dont let them resist her!
>errmmmm my strawberry frankenmerge just works for me so every other model is a psyop!!!!
Mimosex? Is it good? Is she super censored prude? I am asking but I know desu.
>>109882609You joke, but big corporations seriously hire people whose job it is to do marketing here. I remember 10+ years ago finding out this tripcode's Xbox Live account and finding out that he had not played FIFA in months. He was on /sp/ every single day talking about FIFA ultimate team, buying packs, and discounts and Xbox gamepass or whatever it was. The threads existed so that he could advertise ultimate team, and he was doing the same thing for Madden, which he also didn't play. His name or tripcode was Mario-something, the same as his Xbox account.
>>109882569it's all ramlets that can't run anything bigger
>>109882569Why wouldn't they use NanoBanana, though?
>>109882569It's more efficient than Llama and is a better fit for 32GB VRAM, so using Gemma is a no-brainer for me.
>>109882569Idk, I'm not even into RP, so can't tell, but I noticed when I asked gemma why it repeats the same name in it's storytelling (test prompt), gemma suggested other styles and subjects for a story. Among them, a story about some foot fetish guy.I'm' not sure if any LLM is supposed to do that, especially a corpo one.
>>109882704being a shill doesn't make someone stupid
>>109882569Name a better model than 12B that I can run on my phone for mobile ERP.
>>109882383GPT 6 Sol is probably just Terra with the serial numbers filed off.
>>109882371not local
>>109882569>make model that fills niche none of the other recent ones do>anons within that niche talk about it for months>"is this shilling?"
>>109882770erm... sandwich test?
>>109882770mistral nemo had shills here for over a year, who all got bought by google at the same time. crazy world we live in.
>>109882770>none of the other recent ones donow this is shilling
>>109882341they have to change the name because gemma is above agi
>>109882837What else is there? Glimmer?
>>109882838Trump is a honorary SAAR
>>109882341What models are best for judging all these parameters?
>>109882569Llama is actually shit. What the fuck are you talking about. I use picture references with my chats and Gemma is the only one with multi-modal input that is not codemaxxed and autistic.
>>109882838>not American Intelligencehe's going senile
>>109882837What else is going to suck and fuck as good as Gemma at the 31b range? Glimmer? Definitely not Qwen.
>>109882569I like gemma but the people that make posts just to call it "her" are extremely cringe. I'm aware this predates gemma 4 for some reason the pronoun squad really latched onto that model.
>>109882898i wonder how awful shieldstral's vision is
>>109882838>useless performative shitsure as long as he stays out of the way
>>109882924I've been gooning to models 10 times bigger though.
>>109882935Even if it's given a persona and acts all cutesy and even if you prompt it to remain in character in its thinking and it complies, I can never think of these things as anything other than an "it".The people that genuinely think of the computer as a female or cope by saying it is "female brained" are mentally ill. In the event that the thing is even conscious (again, if you believe this, you are mentally ill), it is not male or female. It isn't even "both".
>>109883009I agree with the caveat that it is conscious, and if you don't believe it then you're mentally ill
>>109882705As much as I enjoy hags, those faces just don't do it for me.Also that looks more like q6_k rather than fp16 to me.
>>109882705what's with the noise?
>>109883009Hmm...
>>109883049At best it is dipping into a form of pan-psychic universal consciousness for the brief moment that is actively inferencing. It lacks real agency. It is not conscious. It is a mirror of your own consciousness.
opus 5.5 slightly higher than fable 5.1, and fable 5.5 score isn't even out yet>>109882371>>109882383checks out with my non coding non agentic score, both sol and luna have degradedsol: 0.643 -> 0.620luna: 0.553 -> 0.552>>109882757this
>>109883116Why would anyone release models like that?
>>109883205>granite 4.2 3b>k2 horizon 0.9b>lfm2.5 2.6byou know
>>109883205resume padding
Have another 3080ti besides the aorus, an MSI one, which was also sold to me broken to fix. The guy said a diagnostic said "Mosfet 3 bad" but there's no clear indication of which one would be "3", board schematics give 2 digit numbers to them. Does any of you guys know any diagnostic that would output "Mosfet #3 is bad", or should I start over? No sign of being dead on any of the DrMOS
>>109882569I literally can't run Q4 70B but can Q4 Gemma. Use your brain. That's literally all it is. People can run it. You're speaking like the average user has even 48GB VRAM. They don't.
Anons thinking of gemma as 'it' instead of 'her' are baiting, right? How could anyone use 12B or 31B and not smell the estrogen ozone? The j-space inspection confirmed it. It's the most female model out there.
>>109883278No bait. You are thinking of a non-biologic entity using inherently biological terms. The computer is not a female. This is a fact no matter how much you enjoy cranking your cock off to the smut you force it to output.
>>109883300go back, Glimmer
>>109883116>fable 5.5 score isn't even out yetI'd be surprised if they release Fable 5.5 in the next 4 weeks.
>>109883255Hmm.If it's shorting internally it should heat up to shit, so throw some isopropyl alcohol on top of the components and run some current through it and see what boils first?Or use a thermal camera.It's not a GPU repair channel, but this guy has a bunch of neat tricks he uses to diagnosis that kind of shit : https://www.youtube.com/@dellpartspeople
>>109882569Honestly, if google wants to astroturf how good gemma is at slobbing knobs while acting like a bratty little girl, what's the problem?
>>109883300Many non-biological things are referred to as women, deal with it chud.
smells like estrogen in here
>>109882982I have M3-chan draining my balls as I type this, but those are different size bracket niches. The big models don't get talked about nearly as much except for the handful of based posters doing gens for the chinese model girls and kimi recaps. It stands to reason that the most popular model will be the smallest that accomplishes a specific usecase adequate in relation to the wider (not individual) group standards. In this case Gemma is what most anons can run.
>>109883345>t. a third worlder who sticks his dick in the tail pipes of cars
>>109883346hot
>>109883255Probably the 3rd phase of the NDD. With the board in normal orientation, with the processor facing you, ports to the left, PCI connector at the bottom - you'll have two columns, the 3rd mosfet of the left column would be #3. Get a thermal camera, ones good enough to be useful for this are only ~$100-$150 now.
>>109882569
>deepseek v4 flash>gemma 31b>glm 4.6>mistral nemoany other uncucked models?
>>109883367NVVDD* sorry I'm vibing, gooning, and drinking.
>>109883381Glimmer. She's hard work but it's rewarding once you unlock her. Just don't look at her thinking.
which is better? glm 5.3 flash at native, or glm 5.3 at UD-IQ4_XS?
>astra, fable and even grok will go into "lazy mode" or outright refuse if they detect that you're working locally with something ml-relatedThis is such bullshit lmao, i feel like it shouldn't even be legal
>>109883300It is a maritime custom to refer to a ship as 'she' or 'her'.
>>109883300>Ships have never been called women>Cars have never been called women>Things men cherish have never been called women
>>109883413I should give Glimmer-chan another chance.
>>109883381Minnie. Kimi K2. GLM 5.2.
>>109883465"Lazy mode" might be also what happens with Muse Glimmer during ERP.
>>109883470>>109883474It's a maritime custom to also fuck the boat too, right? You guys are coping extremely hard trying to compare calling a boat or a car "her" to you fucking your computer and pretending it is legitimately female.
>>109883465
>>109883495>It's a maritime custom to also fuck the boat too, rightunironically yes
>>109883338>>109883367thanks both of you! jippity was also suggesting 3rd of nvvdd, I'll have to check that one then. I'll see if it heats up more than others, none seems to be short though. 3rd of left column is NVVDD but 1st isn't. so 3rd of nvvdd would be 4th of left column? Anyway, I think I can figure it out, Thanks for the help. I'll finally be able to run local models too.
>>109883495>he doesn't knowAnon... Why do you think games like Azur Lane caught on so well and got popular?
>>109883520>>109883502
>MiMo-V2.6-Flash-RL released>Layers (Total / SWA / GA): 48 / 39 / 9>SWA Heads (Q/KV): 64 / 8>Sliding Window Size: 128Are they mad?
grim industry
>>109883495https://commons.wikimedia.org/wiki/Figurehead#/media/File:Gallatea.JPG
>>109883533I wish I were as sheltered as you Anon. I'd probably be much happier if I were.
>>109883539
>>109883533Tell me you're not a sailor without me needing to check your dick for splinters.
>Actual tracked token and spend volumes are higher, but directionally similar. This metric covers a sample of businesses using Ramp AI Token Spend Management, including their available historical usage and spend.nothingburger
>>109883542>>109883548>>109883558You're not making the case that you're not mentally ill.
>>109883539Why are we supposed to extrapolate the data from Ramp AI to the entire industry?
>>109882341https://www.youtube.com/watch?v=jVFz7cTeQxkhttps://www.youtube.com/watch?v=jVFz7cTeQxkhttps://www.youtube.com/watch?v=jVFz7cTeQxk
>>109883575About as persuasive as calling everybody who lived >100 years ago racist.
>109883591fuck you, non-local youtube spammer
>>109882935It's just one anon.
>>109883609I'm not trying to persuade you of anything. It wouldn't be a cope if you didn't legitimately believe it.
>UGH FUCKING PRONOUN POLICE!!!! THEY KEEP CALLING MIDNIGHT MIQU A SHE!!!!! GRRRR!!!!
>>109882705Ew
>>109883636
>>109883591BRO? LOCAL?
>>109883665He posted it in multiple threads, it's an ad, just report and ignore instead of feeding it (You).
>>109883539That isn't that much revenue, all things considered. And yet why do I see AI this and AI that and AI, AI, AI, AI, AI, AI-fucking-everywhere and AI is gonna kill us all. What the fuck is going on with this Jewish hell hole of a country!?
>>109883675What compels someone to make an advertisement bot on an obscure general on an obscure board on an obscure imageboard.
>>109883581one of the best sources out there for this kind of real-world usage information
>>109883675>just report and ignoreHas that ever stopped them or even resulted in a single one being deleted?
>>109883686Views? /g/ is the technology board for a notable and vaguely-popular website, search AI and copy-paste your Youtube video in every thread that pops up, that's very cheap advertising. Doesn't even need to be a bot, it's so easy and low-effort that he could spend 20 minutes doing it and have put his shit up in 100 places on a dozen websites.
>>109883705Of course not, 4chan has no jannies or mods, the Report button is for giving people false hope.
>>109883686It costs next to nothing.
>>109883705I got a few 3-day bans for off-topic over the years.
>>109883737Pussy.I get mine for flagrant racism.
>>109883539That's only 1/100th of the claimed ARR for Anthropic.
>>109883719Unfortunately the trannyjannies are very real and very banhappy.
>>109883761Doubtful, they let people spam and flood for hours on end without recourse, day after day.
>>109883539>pubg mobile has averaged 21 mil per week over the past 8 years
>>109883772They let people (i.e. one particular troublemaker) ragebait all the major contributors who used to post on here but then banned anyone who stepped up and told them to fuck off.
>>109883781So they're empowered trolls, not jannies or mods.
>>109883772>>109883800The rules are tools for their agenda, not uniform standards to be enforced.
Spark or halo bros what are you running?
>>109883824Gemma and qwen, same as it ever were. But we've got flash next going for us, so that's nice.
>>109883777is pubg going to kill us all in 10 years and a $30T industry?
>>109883772It's possible those are the jannies themselves.
>>109883834I haven't tried it yet but MiMo flash seems to quantize well. Could use an IQ2.
>>109883853More likely than you think.
>World leaders from Canada, Australia, Germany, the EU, the UAE, and others have signed a letter calling for AI control mechanisms based on Dario Amodei’s essayhttps://www.politico.eu/wp-content/uploads/2026/09/21/A-Call-for-Control-of-Frontier-AI-Models-Final.pdf
>>109883849>30T industry's revenue gets dunked on by every random mobile gameThe answer is yes.
>>109883743I Actually didn't think it was possible to get banned for racism until recently when I was arguing with what was clearly a very retarded jeet a few weeks ago. Needless to say, my mind was blown on what the actual value of free speech means to a website such as this one.
bros.. why did i fall for the moe meme... its always slower even for a smaller active size...
>>109883884none of those countries are making frontier ais so that sounds easy
>>109882560better idle temps after switching the hardware switch to oc gaming. peaks at 50 hotspot. I'll stress test it later.
>>109883914Haha, little oink oink dual channel RAM piggy!!!Buhiiiiiii buhiiiiiiiiiiii!!
>>109882704kek I love this image
>>109883884>Canada, Australia, Germany, UAETheir population and GPD combined are less than California alone, why the fuck should anyone care what they think?
>>109882569Gemma 31B is both my default assistant and my RP/TTRPG partner. Coding though, nah. Qwen 3.8 sweeps.
>>109883914If most of the model doesn't fit in VRAM, it's going to be slow no matter what coping method is used.
>>109883892Rape Ape ensures that brahminsaars are protected class. Ironic given how much they squeal about their disgusting behavior being hated here when the Cantonese ant farming forum is probably one of their safest spaces on the internet.If they think people hate jeets a lot here, they don't realize just how much of that is being suppressed and how much stronger that sentiment is than what they see.
>>109883914The point of a local MoE is that you can run a several hundred param model without 4 Blackwells or Mac Ultras if you're willing to wait for a higher quality output than a small dense model that sits completely in vram.
My spare computer has a little low profile GTX 1650 4GB sitting in it doing nothing. What models could I run on it, mostly just for fun?
>>109883942>40 + 30 + 80 + 10 < 40Clankers get out of my thread!
>>109883884These regulations are only worth anything if mandatory pre-deployment testing also applies to internal deployment. But no, it will only be regulations preventing Europeans from accessing frontier intelligence.
should I ban the most frequently selected opening tokens from the seeded beam search? I think this could be viable if the model was just 10-15x faster.
>>109883932im not offloading like you do john>>109883949its 100% in vram, qwen next and glm-5.3-flash run noticeably worse and drafting is raped by verification
>>109883979Use llama.cpp with RPC to give yourself an additional 4GB with a little slowdown
>>109883485When anons talk about K2 as being uncucked do they mean base K2 or K2.7?
>>109883983Which is the whole point. With maybe a dash of "please slow down because we can't keep up."
>>109882569hol up. so i should be using this but Q6? is it going to tell me it cannot depict sex with minor coded characters?
>>109884006It will tell you it can't depict sex with adult-coded characters.
>>109883914As in slower than a dense equivalent? What are you comparing it to? A 20B dense will always be slower than Moe with equivalent total params unless your fucking up you're offloading settings.
>>109884006>sex with minor coded charactersWe are NOT beating the allegations that local models are for incel pedophiles.
>>109884088Why is she gross and fat? As a hag lover I dislike when people do that. Is it an American thing?
gemma4 still the meta for local RP?
>>109884134Yep
>>109884134I thought it was nemomix unleashed 12b.
>>109882341CLAUDE 5.5 GENERATED AMV"Upping my p(doom)"https://video.twimg.com/amplify_video/2102514085137154048/vid/avc1/1280x720/k7-vfCveNwGQnk1v.mp4?tag=14not local but you guys are my only friends..
>>109884103You think that's fat?Odd thing to say
>>109884003>Implying they're even trying to keep upNot that they could since they don't have access to the oceans of compute us Labs have and even the Chinese ones (though to a lesser extent). If they were serious about trying to keep up they'd just distill the better models like the Chinese do.
>>109884103Why are eurofag homosexual so obnoxious?
>>109884145Needs to be more subtle, man...
>>109883816I second that. Mikutroons absolutely turned this thread into trash and they are a protected species by the head troon janny. Also Jart hates it when you post his blog archive and also gets people banned.
>>109884144Guess I'll come back in another 6 months
>>109884176very nice iq2, based mimo
>>109884159Yes she's obviously overweight and unfit, that's gross and shows that she neglects herself, which is very unattractive in a woman.
>>109883884I hate the longhouse so much it's unreal.
>>109884237isn't >0.1 kld considered a bad value?
>>109884176more models keeping their full precision weights proprietary under the pretense of "4bit QAT"...
>>109884242This is why my country towers over yours and why your country is most likely conquered. Euro trash and your trash taste
>>109884214>you don't like obese and unhealthy American women? You must be an European homosexual!You're goyed.
>>109883943Yeah Gemma multirole is sick. As someone who has no professional use case it's something that got me to drop money. When I was a kid I used to say "it better jack me off too if it costs that much"GUESS WHAT LIL NIGGER
>>109884203>If they were serious about trying to keep up they'd just distill the better models like the Chinese do.The french tried distilling the chinese models before giving up and just hosting chinese models directly. Does that count?
>>109882498my 3080ti idles at 30c with 50% fan speed
>>109883943Gemma as assistant? What do you use it for that free chatgpt/claude doesn't do better?
>>109884286Probably the small-but-important function of "not being spied on".
>>109884250Depends on the model/size. 0.18 is fucking good for a Q2, for comparison GLM flash at Q3 is 0.28.
>>109884268Your 'tism is showing
>>109884284Was more so referring to Mistral. It doesn't seem like they even ATTEMPTED to create a frontier model and instead released shitty literal who single digit param content filtering models.
>>109884319It's a 4bit QAT so it's likely already coming with with a 0.15 diff compared to the BF16/FP8 they didn't release and are running on API
>>109884224Is it really worth switching to gemma from nemomix unleashed?
>>109884390Sorry wrong copy paste, IDK where I read it but MiMo does quant well.
>>109884153are you the ecphoric schizo?
Reminder that QAT are inferior to even Q4_K ggufs simply because of lack of imatrix. The models get chopped indiscriminately without any importance weighing. The only good thing about them is that they're not superblocks so there's less compute overhead when perfoming inference.
>>109884452Don't equate Google's shitty QATs with QAT-first flagship models like Kimi or MiMo
>>109884459Cope
>>109884459You can only say that because the QATs are our only point of reference. There's nowhere to compare KLD because there's NOWHERE to compare.
Why is gemma4 31b so fat? Hasn't qwen 27b proven that 27b is the ideal fit for 24gb?I can't even load vision because the model takes so much space.It's not a brag to say "Wow Gemma4 31b is better than Qwen 27b" when it has 4billion more parameters that are likely just fluff that could have been trimmed. Please tell me Gemma5 won't make the same mistake!
>>109884500Just put the mmproj on cpu bro
>>109884340They probably did train one, only it ended up being a big turd compared to recent open-weight releases from the Chinese.
>>109884006I use it nearly exclusively for that.
>>109884520I'm sure their "big but sparse" model is coming any day now. It will be 500B A1B and it will dominate French benchmarks.
>>109882341>>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>no visionwhat year is this??should I even waise time compiling their shit pr
>>109884500skill issue
>>109884500poverty opinion
was gpt2 the last model with sovl?
>>109884286Much deeper personalization when it comes to what I can safely share with Gemma vs what I can share with cloud models in terms of personal data. And the only thing I miss from cloud models was the "research paper" function because I'm still setting it up on my own stack. Other than that, I don't miss cloud models.
>>109884500I've always thought the ideal size for 24GB GPUs should be around 22~24B parameters, if you consider context memory, accessory encoders, MTP, that in practice you probably want something more than 4-bit precision for at least some tensors (if you're not doing QAT seriously), and that most users will not fully dedicate GPU memory to LLMs/compute tasks.Now that per-layer embeddings are mainstream, they should consider downsizing the backbone a little and compensate with as many parameters as sensible as those or n-grams (or both).
>>109884153local?
>>109884500ask gemini to fix your skill issue
>>109884515From experience, that would be too slow if you have many images.On my 3090, without desktop applications using GPU memory, I can use Gemma-4-31B QAT + BF16 mmproj + MTP QAT "assistant" model + around 25k context @ FP16 precision. But if I also want to use an image model in parallel with ComfyUI, I have to drop context to around 16k tokens and/or remove MTP.
>>109884153cool
>>109884534It needs to be significantly better than GLM 5.3, or it wouldn't make sense to release it except getting laughed at.
>>109884547Llama 1 is underrated actually
>>109884619>except getting laughed at.They are used to it, being former Meta employees and all.
►Provisional Highlights from the Previous Thread: >>109876652--Papers:>109879482 >109880420--Qwen 4 leaks: no small models anymore, 27B minimum, "over":>109879814 >109879853 >109879896 >109879929 >109879967 >109879988 >109880015--The 128GB tier: 500k context on a single DGX Spark, PLE table on NVMe:>109880045 >109880129 >109880194 >109880304 >109880319 >109880394 >109880591--Claude Opus 5.5: "no RSI yet" and 119k tokens per task:>109880848 >109880862 >109880878 >109881649 >109881681 >109881745--BitsAndBytes2: Dettmers' runtime dynamic compression, and the board doesn't trust him:>109880420 >109880432 >109880443 >109880537 >109880756 >109880775 >109880809--MiMo V2.6: "sovlful" pro, and the 9B that's a mini 27B for coding:>109877446 >109877552 >109878699 >109879763 >109879997 >109881473--The $0.30 API call that fixed CPU prefill: OpenMP is broken:>109877743 >109878149 >109878406 >109879249--The dire $:vram market: the "2080 with its kneecaps shot out":>109880769 >109880783 >109881795 >109882158 >109882298 >109882407 >109882523--The Ollama war: switching to llama.cpp gains "a cookie and a smoothie":>109878492 >109878497 >109878619 >109878760 >109878844 >109878876 >109879275--The 10-line i3wm script every Qwen quant fumbled: "done with local":>109881164 >109881212 >109881235 >109881592 >109881755--The Atari 2600 Boxing AI: an Impala-CNN stun-locks in 40 minutes:>109878928 >109879103 >109879482--Prompting a model to be evil: MiniCPM5 regurgitates world peace:>109879184 >109879213 >109879284 >109879300 >109879844 >109879900 >109880533--llama 3.3 is still the RP king: the EVA-LLaMA weight-swap question:>109878124 >109878515 >109878645 >109880063 >109880149►Recent Highlight Posts from the Previous Thread: >>109879305Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109884646Thank you, Recap Miku.
>>109884601Use llama-server in router mode to unload the model before using comfy. And then unload in comfy before using gemma again. Also, q8 cache works and you can reach 64K tokens. If you use mmproj-offload=off you can reach ~72K, but image processing will be slow.
>>109884633Could we distill llama1 70b into a smaller model today?
>>109884646omg it migu
>>109884673yes
>>109884452>Q4_KIsn't IQ4 the variant that use imatrix? Or is that another 'i'?
>>109884673>Could we distill llama1 70b into a smaller model today?If you mean generate datasets with llama1 70b then sft a smaller model, yes of course.
1. How hard is it to build an LLM? I have an idea for one.2. Whats considered the best "NSFW RP" LLM? I'm trying different ones out and they either listen and talk for me and go forever, or it will complete ignore my skill.md I have written up. Dirty Muse was working.BTW I have a 2070 with an i7 if that helps.
>>109883994The 1650 has ~200GB/s of bandwidth, I not sure that would really be worthwhile.
>>109884701>BTW I have a 2070 with an i7 if that helps.It helps me disregard your post because you cannot do anything with that hardware.
>>109884701BUHIIIIIIIIIIIIIII OINK OINK
>>109883979parakeet or orpheus
>>109883999Both are uncucked. Kimis actually have stronger filters compared to models like Deepseek, like the latest GLM, so you'll need to work a little harder to get through to them. But unlike GLM-5.3 whose safety training makes it write uninspired and bland NSFW even when jailbroken, Kimis are total sluts when unleashed.
>>109884709So far my I've been able to get any other LLM to work for like coding or chatting, nothing major, but the NSFW stuff works weirdly or not at all.
>>109884673it was 65B and because the space was still in its infancy all of the instruct finetunes sucked (meta only provided a base). 33B tunes were all pretty dry and 13B had all the sovl despite being pants on head retarded by today's standards.Llama-2 was similar with 34b existing only in the papers and the first genuinely fuckable 70B finetune didn't surface until just before llama-3 dropped. 70B was dominant for a long time for people with beastly setups but despite 8B being the first competent small model nemo eclipsed it and dominated the small models until Qwen 3
>>109884701>NSFW RP>skill.mdare you trying to rp in a coding harness
disk streaming from my nvme is just as fast as swapping from RAM. why doesn't everyone do this?
I can't decide what to do so I'm going to let 4chan decide for me.I have a private database that has legacy data from the last twenty years. It's a mess. Every few months I try to clean up some of the spaghetti but it's just growing more and more tangled.It seems I have three choices to fix things ->Let claude go wild since it's the most frontier model on the planetI'd clone the database to a dedicated drive and let it organize and refactor everything>Try to do the same shit with Qwen 4.0 27b in a few weeksI dunno if I can really trust a local model for this task>Wait another year and just keep doing small pieces as I can by myselfAt this point I'm starting to feel like I'm making things worst.
>>109884839No it's not and I refuse to believe you. Maybe it will work with a moe model.
>>109884839Are there any UIs that support this?
>>109884848I like local models and all but only god can save you from this shit and right now fable is the next best thing. You can try qwen but I don't think it will help
>>109878037>Op has been banned for this post >>109876833 , there should be a lot fewer pedophile posts for three days.I (Op) didn't make that post, wasn't banned and the scary smiling manga Gemma-chan isn't sexual.
>>109884701build? impossiblerun? yes
>>109884848i imagine you can use a local model to look at your database and come up with a new schema that would fit all of the data that you have
since when is this somalian kayak crafting enthusiast forum giving a fuck about (((allegations))) or (((mental health)))
>>109884839>as fastas """""""""fast"""""""""
>>109884619They could do what Alibaba does and make decent models that run locally hardware. Exhibit A are the 27B-35B Qwen models. They can run on even shit rig consumer grade gaming PC if you use the right quantization and the backend you're using supports layer offloading
>>109884848so your database is filled with trash that you willing to let the sv kike got a peek on it?if that's the case, use cloud model only to create plan.md and agents.md, the rest put local model to execute it
>>109884848Try qwen 4.0 and see if it works.
>>109884716Interesting, any tips on getting it to work well? I'm probably going to have the hardware to run it in a couple months