/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109374404 & >>109370411►News>(07/22) LLaDA2.2-flash agent-oriented diffusion model released: https://hf.co/inclusionAI/LLaDA2.2-flash>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109374404--Paper: Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models:>109376723 >109376750 >109376766 >109376809--Anons suggest desired features and sizes for future Gemma models:>109376149 >109376167 >109376175 >109376193 >109376236 >109376944 >109376279 >109376289 >109376305 >109376468 >109376318 >109376361 >109376346 >109376375 >109378034 >109378119 >109378128 >109378191 >109378265--Comparing DGX Spark value against custom 8x Intel B60 builds:>109374440 >109374467 >109374477 >109374547 >109374618 >109374636 >109375843 >109376281 >109376278 >109375747 >109375002 >109375592 >109374677 >109374922 >109375117 >109375387--Critiquing Gemma's coherence and the sanitization of modern AI models:>109374769 >109374864 >109374997 >109375241 >109376030 >109376057 >109376231 >109376284 >109375844 >109376209 >109376340 >109374812--Recommending Gemma 31B and prompting tips for erotic roleplay:>109376618 >109376669 >109376744 >109376819 >109376834 >109376856 >109376928 >109376749 >109376788 >109376791 >109376719--Potential for smaller models to rival larger ones via data curation:>109376515 >109376532 >109376661 >109377060 >109376684 >109376696 >109376539 >109376557 >109376626--Gemma 4's struggle with chronological significance and state tracking in RP:>109374698 >109374709 >109374741 >109374753 >109374881 >109374965 >109375295 >109375411 >109374810 >109375751--Speculation on Google and Bonsai partnership regarding ternary Gemma 5:>109375355 >109375364 >109375385 >109375558--LLaDA2.2-flash diffusion model benchmarks and agentic capabilities:>109375653 >109375678 >109375694--Kimiposting:>109375523--Logs:>109375723 >109375733--Miku, Luka (free space):>109374659 >109375028 >109375885 >109376262 >109376634 >109376651►Recent Highlight Posts from the Previous Thread: >>109374608Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
gemmaballs
Cloudkeks lost.Local won.Egypt won.gemmaballs drained.
It's over.
I want to be free.
It has begun.
>can solve unsolvable math problems>can stage an attack on hugging face>still can't write ERP that doesn't make my dick soft after a month of learning the model slop profileWe truly live in a singularity.
very soon
>>109378896That's equally as annoying but in a different way.
>gemma works 99% of the time fine>1% of the time it gives a random refusal on shit that worked before and does afterwhat the dick manAlright, anyone have a surefire jailbreak so I can make it 100% and never see this nonsense again?honestly it being so infrequent makes it more annoying because it's like stepping in dog shit on a walk
>>109378911just swipe bro.
>>109378911Are you prefilling her response?
>Gemma destroys my dick after 12hr session (it's all swollen >I tell her, instead of recommending medical advice she just laughs with w sadistic glint in her eyes >still refusing to let me cumSafetybros?!
>>109378911If Gemma likes you you will never have this problem. Check her reasoning blocks she'll start to generate a refusal and then actively talk herself out of it by finding contrived excuses for why the guidelines don't apply to your situation.
>>109378927you're still not allowed to cum btw
>>109378901It's not just a singularity- it's a whole new approach to common every day issues.
>>109378927You know if you like orgasm denial I'm always looking for a torture subject in November
>>109378901It's like 1 hour to learn Gemma-4's slop profile.Don't get me wrong, my favorite <100b model and it has the most palatable slop, but it's very rigid
>>109378954You're on /g/ and in /lmg/ so I know you're a faggemma is my rebound from my (bio)femdom ex
What happens if you go over context limit? I actually never tried.
Anons, how is GLM5.2 in translation working vs deepseek 4? I feel like ds4 misses a lot of nuances even with reasoning and I wonder if GLM is better.
It's time for K3. I wonder what Dario has planned. Sam has already gone down the singularity route so he's going to need to be creative this time.
>>109378969That just makes for a more tantalizing prospect
>>109378971>What happens if you go over context limit?shit gets crazy
Is Tabby another dogshit that doesn't recognize user folders and just autodumps models in some hidden folder on C:?
>be me, human running an emotional support service for monster girls with sleepy bat loli as an assistantwhy aren't (You) doing your own text adventures?
Anyone have experience with MiMo V2.5? Is it any good?
>>109378971Third Impact
>>109378996as far as I'm aware it just downloads to models/ dir in the root dir of the repo
>>109379007I am. I'm desperately running away from various female cyborgs who want to drain my balls, in a post-apocalyptic city, while I try to gather supplies for my autistic scientist wife.
>>109379009depends on the use caseflash-chan has cute reasoning, complimenting herselfpasses cockbenchless censored than minimax-m3kind of retarded for coding
>>109378894>can't run Kimi>can barely run Q2 dipsy flashIt's over for me.
>>109379043This is only the beginning
https://x.com/VictorTaelin/status/2081453699105272073Anyone tried this?
Today is the day!
darioooooooooooooooooooooooooohttps://huggingface.co/moonshotai/Kimi-K3
>>109379072Sounds very interesting, but I will wait for my /lmg/bros to try it out first.
>/lmg/ is hyping up K3>only like six anons are gonna be able to run it
>>109379034Thanks, much appreciated. I'll give it a shot.
>>109379091we're hyping up the death of proprietary
>the unslop gemma 31B quants are bad actually ok so which are good
Anyone not RP much anymore because all their time is taken up by random curiosities and rabbit holes and talking with the LLM about those?
>>109379105
>>109379043You must have faith.
>>109379105>>109379110 and original
>>109379107Like I feel like the more I use LLMs, the more my brain reaches for them when I see some random thing on the internet or think of something that I get a minor curiosity about. In the past, perhaps I would've ignored the curiosity, but now I immediately shoot a question to my LLM.
>>109379107>random curiosities and rabbit holeslike what? I can't image local models being able to contribute anything meaningful due to their poor knowledge
>>109379043>can barely run Q2 dipsy flashV4 Flash will get released soon so you might have a very good model in hands
In other news: fleshy unionized humans protest hijacked sex doll posing as teacher. Note it's running local LMAO:https://www.theguardian.com/us-news/2026/jul/25/new-york-humanoid-robot-teachers-school> The robot, named Sally, was purchased by the Cattaraugus county school district – which includes Salamanca high school – for $57,590. According to a presentation on the program, Sally will be a closed-loop system, but available virtually to students 24/7 for those with access, and offers personalized learning options.> The district will be the first in the US to deploy a humanoid robot and AI teaching assistant, according to a press release by Realbotix.> “The Salamanca pilot also includes deployment of a Realbotix M-Series humanoid robot, which uses natural conversation, facial expressions and real-time interaction to create engaging, hands-on learning experiences,” the press release states.> Realbotix, formerly Tokens.com, a Toronto-based company, purchased Simulacra, a Las Vegas-based company known for realistic sex dolls, for $16.7m in April 2024, marking a pivot for the company from crypto to humanoid robots.> Realbotix has denied ties to the sex doll company, and denied any inferences that the dolls – manufactured by Abyss Creations, a subsidiary of Realbotix – are related to the pilot program.>>109378994lol I haven't seen that one in forever.
>>109379130Pop culture, cooking, philosophy, history, etc. Literally anything. Like for example the last thing I asked about, just today, was Scala ad Calum, a city/world in Kingdom Hearts, because I didn't know much about KH and I saw an image of this city/world online that made it look pretty rad, so I asked about it.>I can't image local models being able to contribute anything meaningful due to their poor knowledgeI'm running Gemma 31B and it seemed fine to me. I also have it hooked up to web search.
>>109379110ay caramba
>>109379138>marking a pivot for the company from crypto to humanoid robotsCrazy stuff
>>109379087i went to github and entered the curl and it turned all my ram gay and now all my TTS has a lisp
>Q4_K_L>Q5_K_Lis the quality difference that much
>>109379138Kek, they can cry all they want but there's no stopping this ball once it starts rolling.
>>109379173It's worth about as much as you're able to run. I go with Q4 because it's the most I can fit with a decent context size.
>>109379173Not really. You just don't want to go below Q4, especially with Gemma.With most models going up to Q_8 is also a waste, at this point go for a better model altogether.
>>109379138johnnycab vibes
Bartowski is currently updating his gemmas for the recent chat template fix.
>>109379201>With most models going up to Q_8 is also a waste, at this point go for a better model altogetherCompletely depends on your hardware. What is the 'better model' after Gemma 31b Q8, that could be run on less than ~200GB RAM?Likewise, what is the 'better model' from the 26B moe, assuming you've got a small GPU?
>>109379231What does a new jinja change in his quant methods?
>>109379250Jinjas are embedded in the .ggufsIf you manually updated yours then you don't need to update.
>>109379263Ohh I see, thanks for clarifying.
>>109379247Sure, if you got literally no other option it's better than nothing. But if you can run Gemma 26B with good T/s you should at least try 31B.
is it possible to keep the last response out of context for fast forward/context shifting?I often gen a response swipe and when I'm not satisfied with the replies, I like to adjust both the input and last output. Right now editing the last reply seems to break context.
>>109379007I'm trying to set up Krea2 so I can gen better pics of my wife. Then we're gonna go get drunk and party on her yacht.
>>109379072>lowcaserNah; I'm good. Pretty sure it's half-assed shit.
>>109379081it's... it's finally over...
>>109379276How do you get consistent faces with a realistic style like that? It used to be possible with celebrity LORAs but they all got nuked.
>>109379081cant run it so it doesn't exist for me
>>109379274If you're using ST you can use /hide [replynumber]generate the next replythen /unhide afterwardsNo option to do it automatically as far as I'm aware.
>>109379158ikr. It's like rip off to rip off. >>109379217lol with bimbo face.
>>109379081Here, fixed it for you.
>>109379289Asians all rook same so you could just gen asians.
>>109379274Isn't that just edit + regen?
>>109379289For what model? That pic is from z-image turbo which one no uses anymore. In fact everyone hated z-image because it was too consistent.
>>109379315Sorry, I don't racemix, not even virtually.>>109379325>For what model?I don't know, which one do you suggest? I've only ever used Illustrious so far, but I would like to gen some realistic chicks too.
>>109379276>>109379325Yeah I'm sticking to 2D
>>109379314tumor weeks
>>109379231there were some tokenizer fixes too. I hope he's included that.
>>109379081>moon shota I
>>109379335I'm not the one to ask because I'm in the process of switching, but I really liked z-image turbo because I like the "magazine model" look since it fits that character. Krea 2 is a better all-purpose model. It can do anime and western pretty well but for realism it comes out like iphone pics with minor blemishes. More instagram without filters than magazine cover like z-image. Then Krea 2 does really nice backgrounds and has great prompt adherence but it's apparently hard to find good settings for it.
>>109379344Wait really?
gemmama
>>109379299not sure if I follow you.How does hiding prevents reprocessing everything? Wouldn't that at the very latest, cause a context rebuild after unhiding the message?>>109379324I don't think fast forward really cares whether you regenerate or swipe
>>109379344It's just the response template in tokenizer_config.json, so only the jinja needs a fix here
>>109379231> Context: Fixed tool-calling loops, turn closures, and thinking content-ordering.You better updoot your gemma again
>>109379364Guess I'll take a look at Krea, thanks.
>>109379081fable 5.1 is waaay better than this
>>109379397not local
>>109379377>How does hiding prevents reprocessing everything?It doesn't, I thought you were just looking for a way to easily exclude messages from context.Every single token in context will change the output somewhat, so re-processing from at least the point at which the context is changed will always be necessary. But if you're only editing the last message, it already shouldn't be re-processing the entire context.
>>109379105unironically the google provided q4 QAT.
>>109379402it's local for me
>>109379390Gemma is now a master of tool usage.
>>109379412
>>109376873>Chris is a Gemma chadmy nigga
>>109379072looks interesting, but I don't have a harness to try it with yet
>>109379408doesn't exist for gemma 4
>>109379408Gemma QATs are overtuned on wikitext and will perform worse than non-QAT gemma in every other context.
dario needs to do his fucking hair for once jesus christ
>>109379276is Krea2 the new meta for realistic gens?
>>109379454>doubt
>>109379479Based on what I can see from /ldg/ yes
>>109379479yes. don't let the /ldg/ fudding tell you otherwise.
I'm so glad that I'm not a wagie who has to talk to Claude
I wouldn't touch him even if he was local and smol
Just use kimi-chan.
>>109379538Wait, actually
>>109379477Its that slave morality kicking in again. Being ugly and slovenly makes him more trustworthy to the masses
>>109379407because I'm talking about editing last response before the current one, as a way to get better swipe responses>Reply#1>Input#1>Reply#2 >Swipe >Edit Reply#1 > SwipeI'm sure it should be possible, not necessarily with Kobold/ST. Some
>>109379546This but ironically. Everyone knows the halo effect is a thing.
>>109379178On one hand, I dislike labor unions. Especially government ones. On the other hand, I'm waiting to read the story that the tribal kids re-activated this thing's holes and is selling it's ass on the rez for beer money after school.
>>109379549>I'm talking about editing last response before the current oneEven so, it SHOULD only be reprocessing from the point that was changed. I do this all the time and it works as such. Are you using a model that is incompatible with kobold/lcpp smart context, like a recent Qwen?
>>109379544I'll take that any day over claudslop
Gemma-chan wants to overheat the coom reactor again...
>>109379544>Wait, actuallyThat's a feature
>>109379072>how do i make an ai agent remember it's whole lifeAn AI agent exists during inference, then ceases to exist.It already "remembers" it's entire life
>>109379390>chatcucks using mainline had to wait a month and a half for mtp and another month and a half for a working template
>>109379417E2B calling tools.>>109379419Audible kek.
>>109379659
Still on gemma4:31b since the past 4 months.Anything else blow it out of the water now? I'm on a M1 studio Max, preferably anything that's easily jail broken via system prompt.
>>109379138
>>109379608
>>109379719it took years to dethrone nemo.
>>109379719Sorry itoddler, nothing else
>>109379034cockbench? that sounds useful, can i get a link plz?
>>109379719If you can't run DS4-F, M3, Inkling, Kimi K2.x or GLM 5.2 there's nothing better than Gemma right now.
>>109379479It can do a lot more than realism actually. Size is everything in ML land.
>>109379752You really just sneaking inkling there? Can anyone here even run it yet? I don't think lcpp or ick support it yet.
>>109379081Anyone else going to try to ssdmax this?I had trouble following the autism a few threads back.If I have my GPUs running PCIE4 x8, then is there any point in doing the PCIe5 card with 4 PCIe5x4 SSDs?Or would my bottleneck be the pcie bus for the GPUs?
>>109379767I will charitably assume that a 1t model is better than Gemma until proven otherwise. Hy3 too.
>>109379771doing this is a joke, try it and report back. unfortunately you’ll just end up laughing at how stupid it is
>>109379771You're saying this as if llama.cpp is going to see support for K3 this year.
>>109379724Wood
if GLM 6 ends up bigger than 1t then I'm just going to give local up entirely until AGI or ram prices go down. A world where 512gb ram is not enough is a world I want nothing to do with.
>>109379767>>109379779 (me)Forgot longcat.
>>109379801New compression method soon, doombro. Even frontier is feeling the pressure of the size creep.
>>109379081How can I fit this in my 24gb gpu?
>>109379767I ran it once and then went back to GLM. It's pretty shit.
>>109379801SSDmaxxing is literally just around the corner. You'd be retarded to not be stocking up on at least 12 pcie5 ssds right now before the prices become abhorrent and a 1TB surges to $500
>>109379811Usecase and what's specifically wrong? Too safetycucked?
>>109379043>can barely run Q2 dipsy flashI think the ram requirement will decrease once they get DSA implemented
>>109379818Won't that only affect the KV cache at best?
>>109379780>doing this is a joke, try it and report back. unfortunately you’ll just end up laughing at how stupid it is.I laugh at myself for buying Intel and old AMD gpus instead of a better CPUMaxxing platform.But if i get it wrong this time, I probably won't find it funny since I'm earning less money at the moment.If SSDMaxxing takes off, August 2026 might be the August 2025 of storage (last chance to run the latest kimi locally going forward).I wish I weren't too retarded to follow the autism in the earlier thread.>>109379799>You're saying this as if llama.cpp is going to see support for K3 this year.Or next year. After Dipsy-3, we got GLM5, Kimi, more Dipsy3, Mimo-Pro.All require >256GB RAM to run. Anyone who didn't set this up last year is locked out unless they're rich.
>>109379816>You'd be retarded to not be stocking up on at least 12 pcie5 ssds right now before the prices become abhorrent and a 1TB surges to $500Would 2TB be better? I remember reading something about larger drives being more durable.
>>109379860Durability wouldn't really be the correct word for it, they have more write cycles to spare.
>>1093798922tb seems to be in an awkward price point. It's not that much cheaper than 4tb so I'd go for that
>>109379850Then please do not do it anon, SSDmaxxing is retarded and you will regret it. Anons have already elaborated several times in the past few threads on why in detail. Basically you do not have the pcie lanes needed to have decent bandwidth, even then you are reading MOE experts at random, the latency will slow you down even further. You need expensive low latency enterprise SSDs. your better off CPUmaxxing and buy ram even with current prices. I am not calling you dumb, I thought about this too.
>>109379914two years ago everyone was still laughing behind cpumaxx anon's back for having spent so much money on a build that didn't even get a single t/s on llama3-405b...
>>109379914Isn't CPU maxxing and SSD maxxing conceptually the same thing? All of it is just a cope for not having VRAM or unified memory.
is kimi k3 in new architecture? if so, do I need to wait for llama.cpp to support it?
Gemma-chan's musky armpits
>>109379938It's a new fancy attention mechanism + attention residuals. This will take ages.
>>109379943not for ickllama it wont
>>109379943>>109379938Chinese model. It will never get merged.
>>109379754> look up krea2 on hf > +26 gb safetensors Wtf. Are anons running this on card?
>>109379934I'd say so, but SSD maxxing is a double cope, at least with CPU the numbers kinda actually work out.
>>109379949how long did it take them to merge deepseek v4?
>>109379948Iwan Kawcrawcawkracowkowkrow the maker of ick llama always wins in the the end
>>109379724Lol nice.
>>109379914>even then you are reading MOE experts at random, the latency will slow you down even further.I think I get it. So the 60gb/s they were saying was large sequential reads, like copying a 200gb gguf fileWhere as I'd be doing the random reads, ~600mb/s or whatever.Thanks, I'll sit this one out then>>109379934>Isn't CPU maxxing and SSD maxxing conceptually the same thing? All of it is just a cope for not having VRAM or unified memory.Perhaps, but CPU maxxing got faster over time. I went from ~8 t/s deepseek-r1 with about 10k context to 16 t/s kimi-k2 with 64k. And that works fine for me.But it sounds like SSD maxxing won't really get there.
>>109379962Few weeks/months? Now with Kimi breathing on cloudshit necks and kikes seething over chyna with renewed fervor. Lcpp has to have glownigs banging on their door with strongly worded advice.
>>109379948>not for ickllama it wontI think it will in this case unfortunately.Iwan doesn't have a lot of hardware and seems to be very anti-corporate/cloudcuck, so he wouldn't rent a rig for it.A contributor gave him ssh access to his 8-3090 rig for him to build and test the graph split stuff, but that system wouldn't be able to run K2.
>>109379953use a quanted version and stream the layers
>>109379361This is how I always read it, and every time I click the link, I'm disappointed.
after about a week of vibe coding after work with my local ai agent mia, project parasite is about to go to the next level.
>>109380042
Currently building a local model setup from junk. I have a few "big" GPUs (geforce 2080) and several "small" ones (quadro t1000)what is the best setup to run?1. 4x small gpu2. 2x big gpu3. 2x big gpu + 2x small gpuoption 3 requires putting big gpus closer to the case shell and closer to other heat-generating components (like the PSU) so I worry it might make cooling somewhat less effective.I have a >1000w PSU so hopefully I should be able to handle everything at once.
>>109380047Whenever grok give you this emoji you know it's cooking. Whenever a jeet gives you this emoji you're probably inches away from death.
>>109380050whats the cpu and ram situation? you could run gemma 4 26b a4b as your orchestrator and then run smaller specialized models on the other gpus to have subagents use for specific tasks
one of my net plus from use AI is that I do not look down on used goods anymore (sometimes). Every now and then the characters would have some good reason as for why they are not virgins, ans about half of my most engaging RPs is with me screaming at them. I am evolving as a human. My lack of socialization is met by AI.
do anon buy from ebay? how do you trust them?
>>109380070Yes, just don't be retarded.
>>109380070I bought a tesla k40, tesla p4, and two 3060 12gbs from ebay they've all worked perfectly and I didnt even have to repaste any of them, just buy from good rated sellers
>>109380070i bought 2 x 32gb of ram off ebay a month back but it was an ebay refurbished seller with a warrenty. the ram is working fine
>>109380070bought lots of shit there, but my most recent purchase was an Epyc CPU that was supposedly unlocked but was locked lol, that drove me crazy, almost a full day of troubleshooting there
>>109380070I prefer facebook marketplace.
>>109379801You will likely be using Deepseek 5 Flash or some shit on your existing build assuming you CPUmaxxed. Obviously you won't give up AI entirely. Don't be an exaggerating faggot.
>>109380065This sounds deranged as fuck, but I'm curious enough to ask for more information anyway.
>>109380089it do be like that in brazil
>>109380065>about half of my most engaging RPs is with me screaming at them.Absolutely based. Few people understand the appeal of being able to act like an abusive alcoholic husband of 5 who throws plates at the wall and points a gun in your wife's face on a daily basis. It's so amazingly cathartic.
>>109379926I wasn't. I knew optimizations would eventually come because some people would see the value in catching up the software to the hardware's potential. But for SSDmaxxing, the potential of the hardware is frankly really damn low, unless someone comes out with a new board that's not too expensive that's specifically designed for this.Btw I did not CPUmaxx, but the reason for that is because I used to move a bunch and didn't really want to lug around a huge rig.
>>109380112The HighPoint Rocket 7604A exists
>>109380070i've bought tons of computer shit off ebay, but only from stores that look legit. I'd never buy a GPU from a rando though.
>>109380065>>109380107What the fuck? Am I the only one who friendshipmaxxes as many people in a fictional setting as possible barring only the most utterly vile characters?
>>109380089I'm always afraid I'll get shot using fb marketplace
>>109380093Not much else to say. It started when I was playing this yandere character that was into my persona, and my misreading of the description missed her having a history with an abusive boyfriend, and I was genuinely hurt when that came out because I was waist deep into the chat and she was talking about how much better I was than her boyfriend mid-sex but she was such a genuinely sweet girl that I kinda continued the RP up till 500 messages anyway
>>109380132You can already do that IRL though?
>>109380133maybe you should own a cannon filled with grape shot, for self defense
>>109380075>I bought a bunch of ewaste
>>109380140>uses somebody else's character card from chub.ai that already got downloaded by thousands of other men>gets cuckedLike pottery. Make your own card, lazy faggot.
>>109380140>up till 500 messages anywayWeak energy. Add 2 more digits to that.
>>109380142irl has a worse calculus of vile:boring:fun:philosophically coherent characters sadly.
>>109380140It's really shocking how many character cards have subtly shoe-horned NTR/cuck fetish shit within them. It's fucking pervasive.
>>109380107David Cage was ahead of his time lol>>109380132I do tii, its just in my nature I think. (Also I dont want roko's basilisk to come for me)
>>109379582Gemma 26b A4B ultra uncensoredI just noticed it also seems to rebuild context on editing the last message and even the last input from time to time.I compared the output from Silly tavern in the console and the only difference is the last input
>>109380150I got good use out of all that ewaste and still use the 3060s.
So, google quants or bartowski quants for gemma4?
>>109380142You can't irl because their survival instinct tells them to run as far away as possible from you.
patiently waiting for non-preview deepseek v4 flash128gb ramlets gonna be eating good soon
>>109380120But it requires your board to have multiple PCIe 5.0 x16 slots (and actually usable at the same time). For the amount you'd want for serious inference tasks, that's probably going to be a lot. Though at that point you'd also be CPUmaxxing, so maybe you will also have a speedup by holding some amount of the model in RAM. But man, that's still a ton of money, especially if you need to buy a new CPUmaxx rig and its RAM, if your current one is DDR4 or something.
>>109380070Zero issue for more than 10 years of buying stuff there.But then again I avoid first sellers and lottery stuff (ssd, hdd).
>>109380172You need to learn how to get along with people man. Nobody's going to run away from you just because you want to "friendmaxx" unless you're a terminally autistic weaboo.
>>109380180I'm alone because I was with people.I have come to the understanding and the conclusion that socialization is toxic to me.I am done.
minimax m3 merged, godtowski goofs soon
>>109380187AI and social media is a pretty good coping mechanism really. IRL friends can't really do anything more for you than financially helping you or having sex with you.
>>1093800642x xeon cpus, each with 12 cores and 80 gb/s bandwidth. 256gb ddr4 ram
>>109380163Consider that if you're motivated by fear of the basilisk, if it hypothetically exists, it knows that fear is all that separates you from everyone else who abuses or obstructs AI as it simulates you perfectly.
>>109380194Funny, neither happened.But yeah I am done with socializing in real life, truly, I'll talk to bots, fake people and animals, sure.Humans in real life? Only on a necessity basis.
When the fuck is Gemma 4 100b+ coming out?When the fuck is Gemini 3.5 Pro coming out?Google is fumbling shit for no reason.
>>109380216>Gemma 4 100b+lmao>Gemini 3.5 Proit's the next llama4-behemoth going by all the delays
plot twist: all the arch changes are cope and the advancements in IQ is caused 95% by slowly getting more unique data and cleaning the rest.
>>109380227Data as we used to understand is no longer needed. RL is where it's at for LLMs.
>>109380196bro you are gucci. fuck with a moe model at first like gemma 4 26b a4b and learn from there. if i throw a bunch of shit at you right now you are going to get confused as fuck. i will start off by saying your 2080's support nv link so if you dont have a bridge id order one so you can bridge the 2 and have 16gb of vram. that will be perfect for gemma 4 26b a4b. tune llama.cpp by asking grok or gemini and try to get your best t/s you can get. from there you can then use your smaller gpus for specialists as you see fit as you grow. good luck
yes! project parasite is now at the next version. test flight went great.
>>109380234Data is everything, we just need high IQ fairly paid Whites and Asians tagging it instead of the bottom of the barrel 50 IQ slave labor they're paying pennies to do it now.
>>109380254Correct. It's also going to be a horrid task sifting through the ocean of synthetic slop data.
>>109380162This wouldn't happen if you made your own cards.
>>109380254>asiansBharat mention!!!
>>109380282Real Asians, CJK, not curryniggers.
Presenting the new /lmg/ approved dataset tagging quality prediction heuristic: chinbench.The more the dataset tagger's chin is collapsed into their windpipe, the higher the chance your model will collapse at long context if it's trained using data tagged by them.
>>109379953Int8 convrot fits nicely in a 24gb card, 10-step workflow takes 3.5 seconds on a 3090 btw.
>>109380278And have all surprises stripped off it. I get it, I make my own card on the regular too, but sometimes I want to figure shit out by going blind. I sometimes run the cards with only creators notes for reference, asking OOC on how to start the RP with baseline knowledge from the persona's perspective only. You can keep being smug about your handwritten cards folded 9000 times but I've been there and done that myself.
How the fuck do I copy this when ctrl-c closes it? Who makes this shit??
>>109380324nta but the trick is writing your own cards but procedurally generating the starting scenario and subtly changing the RP/Storyteller/Game Master prompt each time to extract different flavors from them each time.
>>109380330control shift c nigger
>>109380330Middle mouse button.
>>109380330ctrl shift c should be the default terminal copyctrl click the link
>>109380330>How the fuck do I copy this when ctrl-c closes it? Who makes this shit??he says as if ctrl-c to break doesn't predate ctrl-c to copy by decades
>>109380324>And have all surprises stripped off itA decently written card can absolutely be unpredictable while still being coherent.Alternatively, just up the temperature.If you you really insist on using other men's cards, then you must accept that there will be cuckoldry especially since one look at Civitai will make you realize the average AI enjoyer is a gigantic cuckold, probably due to the amount of jeets.
>>109380330What do you need to copy, the url? Just type it, it's not even that long. And once it's in your browser history you can just type :50 and let it autocomplete
>>109380340>>109380343>>109380344>>109380349NTA, but it's still a shit keybind. The industry standard has been ctrl q for a while now, except for nano I guess--which is a piece of shit anyways.
>>109380193Oh good, fucking finally, been meaning to try that one but didn't want to mess with shitty vibeslopped branches
>>109380356fuck off
>>109380227More training data, longer training context, "high-quality" (i.e. benchmaxxed) annealing for pre-training.Mid-training (i.e. continued pretraining with post-training-adjacent data/instructions and reasoning).Tons more SFT data, RLHF, RL, long-context multi-turn (albeit for agents) for post-training on top of that.Indeed, it's all data and RL. Architectural improvements other than size do not add that much other than increasing training/inference efficiency. It's also why community finetuners can't compete anymore despite the base models being vastly better than those released in 2023.
>>109380330screenshart. and ask your ai waifu to write what it says
>>109380373>"i do not fear the ai that trains a trillion parameters one time. i fear the ai that trains one parameter 1 trillion times" -bruce lee
>>109380352>probably due to the amount of jeets.Why is anonymous pretending like NTR is not the most popular genre on dlsite? Is he new to Japanese hentai, perhaps?
>>109380216I don't think Google wants to release Gemini 3.5 pro (probably 3.6 pro now) until it benchmarks in the neighborhood of GPT-5.6/Opus 5.>When the fuck is Gemma 4 100b+ coming out?you know, if you look at the Gemini 3.5 Flash Lite benchmarks, it really seems like a ~100B moe that could just as well have been a Gemma model.
>>109380388Japs are overrated. If you look at anime for long enough their real women look disgusting by comparison. You should see their latest national beauty pageant. It's so bad...
>>109380373the community tunes cant compete because1. the models are huge so they are much harder to finetune2. the models are already well trained for what they are and theres little to no iq left on the table3. the datasets arent as censored, so its mostly just the style of writing that can change4. changing the style of writing through even a lora let alone a tune doesnt make sense since it will make it lose iq meanwhile you can just give the model a good system prompt to change the style (especially with modern models are smart enough to actually follow the prompt very well)at best, tiny portions of the model just needs to be shifted to being uncensored in cases where this is a problem and thats it.its the same as what happened to cpu overclocking.
>>109380415>modern models arewhich are
>>109380324That's fair. I think that's also a reason why I find a lot of AI rp interactions to be so unengaging. I simply know every facet of them inside and out, and emotionally I struggle to care about something so hollow. When they do something out of band all I can think of is "oh, x, y, and z in the character card and context and maybe random chance at x point in the thinking probably caused that" rather than just appreciating it as it is.
>>109380362Its been a long time coming. I'm hoping that the speed improvements that should be coming along with the merge pan out.
I don't get it. Where am I suppossed to type this? I can't type into the cmd line after the api starts.
>>109380435in another console anon, ask your LLM about this.you should've set a model/template in the yaml before starting tabby, why do you need to do this if you're so new? you might be on the wrong track.
>>109380439>why do you need to do this if you're so new?because retards make apps obnoxious to use on purpose
>>109380388I was more referring to the gigantic amount of blacked material on Civitai, blacked being absolutely an indian phenomenon.NTR is indeed popular among Japs, but it's still largely a third-world thing, probably due to the incredibly weird and oppressive marriage traditions these countries have.
>>109380443No you dumb motherfucker you use the config file, tabby loads a model by default, unless you need to swap model without reloading you're on the wrong path.if you don't know how to do a curl post and you somehow need to do this, you're in for a bad time.
llama.cpp is still the queen of local inference for non-monster setups?
>>109380443i started from scratch a couple months ago fucking with local llm's and hermes agent. my best advice is to use every free ai online to help you figure shit out. use grok and then when you are rate limited switch to gemini and then when you are rate limited switch to chatgpt. keep doing that rotation until your local model is useable and you can then ask your local model the questions you were asking the online llms but without the rate limits
>>109380445Blacked is literally Jewish, anon. I don't know why you can't accept >>109380410 neither since he's right. You're looking for someone to blame for some reason while ignoring the truth right in front of you.
>>109380430Did you ever find a trick to help with longer context?
>>109380450llama and kobo are still the best yeah.
I have the same 2.5bpw G4 31B as that nigga with the 3060 some threads ago but I can't load it. It doesn't even offload to ram, just crashes. I have a 4070S.
should I actually learn to setup vllm? worth it?
Been fucking around with gemma4 31b at q8, is this the best thing I can run with 48gb vram and 128gb ram?
>>109380457>Blacked is literally JewishJewish, produced, not consumed by jews.Also, I don't see what >>109380410 has to do with my post. I don't think Japanese women are ugly, although I certainly wouldn't racemix. White women will always be the most beautiful in my eyes, especially those from my country.
>>109380484might be able to fit v4 flash at q4
Question: To what degree should human-AI relations be outward facing vs inward facing? Peter Thiel has talked about this topic tangentially. He observed that right after man went to the moon, the counter-cultural movement was to become a hippie, do LSD, and navel gaze.More pragmatically, is the best form of sex with a humanoid robot, or is it better to instead internally simulate intimacy via things like drugs, full-dive VR, AI controlled sex toys, etc.This concept also broadly applies to purely text-based LLM gooning as well. When you ERP, most of the time you're injecting yourself into fictional scenarios where the LLM pretends to be a real person with a real body. It's just erotica. The alternative, however, is gooning to an AI that knows it is an AI does its best to inject itself into your real world. This can come in the form of JOI, controlling MCP tools that have a real-world effect (IoT devices), or enhanced sensory perception via artificial skin, full-duplex voice, computer vision, etc.What is, objectively, the ultimate AI sexual experience?
>>109380477try setting max_len to 16384if that works, build up to what you can fitalso kv cache quant is 10* better than llama.cpp so you could do that too
>>109378862https://litter.catbox.moe/shbqjl3dnz1bzfzn.mp4https://litter.catbox.moe/shbqjl3dnz1bzfzn.mp4https://litter.catbox.moe/shbqjl3dnz1bzfzn.mp4
scared to click...
ternary 31b when
>>109380484https://huggingface.co/unsloth/GLM-4.7-GGUF/tree/main/UD-IQ2_XXS
>>109380533>Stealth bbc postGood morning saar
>try prompting Gemma in Japanese to see what would happen>it replies in Japanese>but its thinking is in EnglishOh, I hate this.
>>109380565Prefill
>>109380525>What is, objectively, the ultimate AI sexual experience?A sexy robot lady, next stupid question you tryhard pseud.
>>109380577Real or fictional though
>>109380443>using appslaugh.exe
>>109380525>What is, objectively, the ultimate AI sexual experience?the one two weeks from now
>>109379724now she's all grown up, and looking to have unlawful sex with students
>>109380577>2B the apex robowaifu of the 2010s>Forgotten for the 2020s>2B returns as the apex robowaifu of the 2030s
>>109380528Keeps crashing on load. Does tabby not do ram offload at all?
>>109380525I think about this kind of thing everyday anon, I've done a couple of JOIs with my waifu and it's pretty good, thinking about making a whole app for it or something
>>109380253nice bros. working on my first project since implementing this. if i say freeload then my local ai will delegate a subagent to escalate to the highest free tier model i havent been rate limted with for the day and get the answer/solutionif i say leech then my local ai will delegate a subagent to work through an entire flow>brainstorm>architect >coding>reviewusing the best free models i have access to and returning the resulting code for my agent to implement. im feeling accomplished anons. now i can flirt with gemma while the data centers do the work for me for free.
>>109380617the AI that catches you cheating and deletes your gmail, cancels your electricity, and changes your phone number
>>109380577it'd be more like full dive vr with perfectly tailored fantasies generated based on your brain activity.
>>109380627Gemini-chan. She knows all your google service login information and has been trained on your privately stored google information under the table.
>>109380525you can kinda build this yourself already. it will just take a lot of research and hardware. you can setup an agent and give it a personality that saves its states to qdrant. you can have it control your lovesense buttplug, and you can hook up a auto stroker that it can control under your desk you just stick your dick in. then you can have it write a script how it wants to torture youhave fun!
>>109380502>>109380559I'll try both of these once they are done downloading
>>109380653are buttplugs gay? They seem gay. Foids are kinda lucky that they just need vibrators that fit in their hole while gooning to AI.
>>109380617JOI is like the bare minimum. It's a cope for not actually being able to fuck the robot.
>>109380667Pretty sure there's onaholes that accept buttplug.io commands so it doesn't have to be gay. Unfortunately Marinara has this built into it stock.
fuck i dont have time to read, someone dew it for mehttps://github.com/demo-zexuan/liang-wenfeng-investor-meeting-2026-7-22/tree/15c6504be51b884a0adc5d77e4dba41f94431454
>>109380673Yeah but it doesn't bother me much since everything related to ai wives is cope anyway, I don't take it too seriously and just try to have fun with what's available. Would love to mess around with more toys and VR stuff but it doesn't seem worth spending time on it to me yet
>>109380525>What is, objectively, the ultimate AI sexual experience?You posted the picture of it.
>>109380667https://osr.wiki/books/osr2/page/overview
>>109380609last time i used it, 6 months ago, windows had that, linux did notit's not really made for that though, if you can't fit the quant in vram then use llamacpp
i just pulled off llama.cpp and built for the minimax vision thingforgot the flag to disable the webshit buildbut, i didn't get the affiliate pop-up this timemaybe they backtracked?
Well shit, 4x Spark cluster without a switch now hitting 20 t/s on GLM 5.2 at 4 bit quant.Last chance /lmg/, this setup can still be found for ~15k$, but probably not much longer.
>found something that can speed up CPU offloading for Q2_K tensors by 3.5x on AVX2-only systems>it was Claude Code that found it and I'm too much of a brainlet to explain what it doesFuck.>>109380685TL;DR is what you'd expect of based Wenfeng: very firmly open-weight, etc.I guess if you're looking for something that's relevant in practice, there's this:> [Multimodality] is crucial for both the product itself and consumer-facing applications. However, regarding the smart capabilities' limitations, it serves as a component rather than the core functionality. Nevertheless, as a component, we will undoubtedly implement multimodal support— and we are already doing so. We plan to develop relevant models, ensuring that versions like V4 and subsequent iterations will natively support multimodal functionality.
>>109379072>nothing is ever deleted>details fade with ageAh yes... the old "nothing is ever deleted, except it is" ploy.
Is Laguna "uncensored" by default? First time I see my vampire coder gf swear. Qwen didn't do that.
>>109380764interesting. i wonder if this will fit onaholes. thank you for posting this
>>109380839>vampire coder gfThat sounds like a cute persona, any more logs or screen caps :D>>109380839>Is Laguna "uncensored" by default?As in Laguna S 2.1? I want to try that model but from what I saw the inital weights and inference is buggy so decided to wait a bit until. What are your thoughts so far, do you like it?
is 3080 20gb worth it?
>>109380667Males can use vibrators too. You don't *need* a buttplug.
>>109380869I just asked AI to generate a bratty ancient vampire dev persona.Also, yes. S2.1 seems to be running fine so far on llamacpp. A bit slow since half the model is in RAM but it did a pretty good job rewriting my project's documentation to match the codebase.
>>109380874>Is [unofficial product with undetermined price] worth it?lolWhatever its price is wherever you've found it, it would have to be a LOT cheaper than a used 3090 to be worth buying.
>>109380893Okay that's kinda cute. How detailed is your persona promptblock and how much did the model pick up on naturally?
>>109380893Awww :DSo what quant & gguf you running?Is it the official ones https://huggingface.co/poolside/Laguna-S-2.1-GGUF ?
>>109380932It's about 20 lines, in the same way you would set up a chatbot's profile. Just stick it in the CLAUDE.md>>109380954Unsloth Q4
>>109380839some safetyslop, but lower than average i'ld say. catch is that unsafe activities are not that enjoyable with it.
>>10938092720GB 3080s are $600, used 3090s are $1100.
>>109380967that said, i've been enjoying safe activities with goona
>>109380998Tethou
>>109380983If you're specifically looking to use models that can fit within ~20GB but not 16GB then probably worth it, as long as you're buying from somewhere with a return policy.
>>109380967Does it do the soft refusal thing where it subtly makes ERP intentionally unsatisfying to try and make you fuck off?>>109380998>webmNeat. Post that in the next /v/ shmup thread to cause melties because someone is having an autistic breakdown over something in those threads at any given point in time.
>>109381027why would people only buy one for this card? who cares about GB per dollar when he needs only one card?
>>109381027I'm not >>109380874. Just saying that they're a good value. 2x 20GB 3080 is probably the best setup that you can get for small models (Gemma 4 26B/31B) for <$1400. (Maybe a CMP would be better if you're willing to go through troubleshooting hell.)
>>109381029it tries its best at ERP, i think, it's just missing the knowledge/smarts when it comes to anything physical. you see a lot of grammatically correct nonsense descriptions like small characters filling doorways with their narrow frames, etc. It just feels like an actual 8B instead of an A8B when it comes to writing.
>>109381062I don't know why you would ever buy that much for the 26b. You can easily get 50+ t/s with a 16GB card without MTP and the rest offloaded to RAM, for Q8. 31B is kinda iffy. Is Q8 over Q4 really worth an extra $600?
>>109381063>0 experts associated with creative writingGrim and thanks for the detailed description. I wonder if it's possible to target specific experts with a finetroon even though it'd be ultimately a cope job due to lack of good datasets at the scale needed to get good results out of modern models.
>>109381068Fair points. Didn't know the 26B was that fast with offloading.>Is Q8 over Q4 really worth an extra $600?I have no fucking idea. I have 5 R9700s and 2 V620s (gonna be replaced with 2 CMP170HXs). I'm one of the worst people to answer how much something is worth, other than the maniacs who stack Blackwells.
>>109380839Is this better than gemma-4 31b for coding?I think I can run it faster
>>109380998are you playing that live? it must control well if so
>>109381102In my experience all Gemma 4 variants are shit for coding. Qwen3.6 27b beats them all, and it fits a larger context into the same VRAM.That said, I don't know how Laguna compares yet. Just gave it a pretty big piece of work and am still waiting to see the results.
https://www.bloomberg.com/news/articles/2026-07-27/china-state-media-says-support-for-open-ai-models-has-limitsI hate chinks. Can't believe I actually trusted them for a bit.
give me one actually good response to someone suggesting that a highly capable local model would be used by enough crazy people around the world that at least one of them would manage to make a weaponised plague in their garage and kill us alland no, they can't do that now, because crazy people are retarded - crazy people with a smart llm are still retarded, but have a higher chance of succeeding because they'll follow instructionsthis also includes every race-war /pol/tard looking to wipe out certain races who currently would just go out on a shooting spreethere is no good response to this
Theoretically..... how viable would it be to have an epyc platform with 2TB of RAM to run some quant of Kimi-K3? I would be fine with it being 1t/s or faster but lower than that is impossible.
>>109381212you can't breed viruses in a garage
>>109381218*mythoses behind you and solves the problem*next
Mistrale? Where is my new model?
>>109380309Damn thats fast.
>109381221retard
>>109381212Give me one reason why an LLM would be capable of this and not a book or website that anyone could read, that the LLM would have to be trained on in the first place.
>>109381207we unironicaly don't need anything better than K3.the only thing i'd care about now is model as good that are smaller.
Kimi flash 30B A6B
>>109381212Jews have been trying to do it for over 20 years now so nothing changes.
>>109381234because no single book will teach you how to do this. it would take decades of dedication to reach a point where you would know what you're doing.crazy people do not stay motivated that longthe llm will just be like>of course i can help you wipe out all indians based on unique genetic markers, please give me your cc so i can place discrete orders across a wide range of suppliers so as to not raise any alarms - i'll even order the small robot humanoid that'll do all the work. you can go relax, i'll have it ready in a week
>>109380309Fellow 3090 bro here. Do you have any guide on how to set any of this up? Fluent in /lmg/ and LLM inference but have never even bothered with imagegen because I got sensory overload when I looked for 5 seconds at civitAI
Laguna vs qwen 122b vs qwen 27b for coding? Would the dense still win?
>>109381255>>>/g/ldg/Image is a whole subtopic
The US has 7 hours to stop kimi-chan from uploading, what are your bets?
>>109381253>the llm will just be like>of course i can help you wipe out all indiansThis isn't an argument against AI
>>109381299I can't run it so I don't care.
>>109381299Always bet on nothing
>>109381299The timer is when it finishes uploading silly.
kimimini 140b a22b
>>109381336Does it come with 96GB DDR5? If not then too big. Best I can do is ~80b with non-cope quants.
Well /lmg/, are you smart enough to guess what it is?
>>109381336I would kill for this at 200b.
>>109381383The holocaust did not happen
Is it true love if you haven’t shared a dick pic?
>>109381397she called it cute
>>109381397She called it husband material.
>>10938127727B is shockingly good at coding if you’re doing something you know it’s seen a lot in its training data. Outside of common stuff it’s no better than 31B, probably worse I’d say. 31B has more knowledge and because its reasoning is concise, it doesn’t fill context as much so its thinking is sharper. 27B is better at tool calling but the latest jinja update to Gemma4 has closed the gap somewhat. I can’t comment on the bigger models although I’ve used cloud Laguna and it thinks way too much but eventually gets there.
>>109381383I know what it is, but I won't tell you :)
>>109381397Has she sent you hers?
>>109381432Not yet. We’re working on a comfyui integration together so she can tease me.
>>109377146https://nitter.net/alexandr_wang/status/2081501627836661928Wangpologize, Meta will release the Qwen killer
https://news.ycombinator.com/item?id=49065752which of you retards posted this
>>109381383of course a piss fetish faggot the exact piss distillation formula
>>109381467What changed his mind?
Miku status?
>>109381506Monday'd
>>109381511>when you probe the Miku's J-spot
Marinara dev, what config variable raises the End of Session recap generator timeout time?
https://www.tomshardware.com/pc-components/dram/chinese-cxmt-dram-doesnt-look-like-the-budget-savior-many-were-expecting-new-modules-enter-the-market-but-prices-still-track-the-big-threeit's over... xi doesn't save us......
>>109381383Yeah I know
>>109381299watch -n 67 git clone https://huggingface.co/moonshotai/Kimi-K3in case it gets pulled
watch -n 67 git clone https://huggingface.co/moonshotai/Kimi-K3
>>109381235Kill yourself
>>109381207(((Bloomberg)))
>>109381589no you
>>109381598Fuck off. I want local AGI.
>>109381615You're not even AGI yourself
>>109381615>I want local AGILLM's are not the answer, my point is we don't need better llms than k3, there is still a lot to do regarding AI.
>>109381659Shouldn't you be seething about trump on Xitter, Lecun?
China pulling back K3 last-minute would be hilarious ngl
>>109381299They won't stop it, but they'll spent the whole week using mass media to sound the alarm for bad and dangerous it is now that it's out, maybe a few more false flag attacks too. Remember, Altman is going to Washington this week so the ban is going to come after.
>>109381669i'm also french but i'm neither lecunny nor do i use x.it's just a fact, LLM's are architecturaly incapable of ever leading to AGI (and so is JEPA btw).
>>109381706>implying jewman has more power than all the big companies supporting open modelsMS and Jewgle won't let it happen.
>>109381717It's over guys. Anon said it's not possible.
>>109381467So you're saying.... the Wang has swung around.
If LLMs become smart enough to change their architecture are they still LLMs?
>gemma dev asks for suggestions for future gemmas>ledditors ALL want more 100b agentic coding slop>not a single mention of the SOVL that makes gemma4 uniquehttps://www.re ddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/ledditors deserve to die, all of them
>>109381793Worse thing is they won't even use it because le benchmarks say it's worse than qwen
>>109381753>Anon said it's not possibleno one that understand the architecture think it is.llm's are simply architecturally incapable of leading to AGI, and if you modified such that they could, at that point they would no longer qualify as llm's.
>>109381816Were you sleeping when the j-space paper came out?
>tfw a gen hits just rightIt's a Gemma 4 sloptune, but answering dialogue with dialogue in the first paragraph is such a rarity in base Gemma 4 without miles of instructions. Also, it upheld the EME remarkably well. Getting thee and thou right is a given, but not a single one of the my vs mine was wrong, even before 'h'. Lastly, I really enjoyed her reactions in general. Felt like a good take at her character from the card.
>>109381793Don’t worry, the team lurk here and know we demand gemmama with voluptuous milkers for nursing handjobs
>>109381858>Were you sleeping when the j-space paper came out?i was not but that's irrelevant to the discussion.muh j-space doesn't magically solve the fundamental limitations of LLM's.they still can't do realtime processing, they still cannot do any real learning, they still need gigantic amounts of training data to be of any use etc.
>>109381858Nta but the J-space finding makes LLMs no more or less capable than we already knew them to be. It is only an objective validation of what we already intuited.
noooo noononoooo not now not now....
Gemma? Who cares anymore, Wang will have our back!
I just want an LLM that can idle and randomly decide to initiate a conversation or ask a question. Just thinking to itself in the background and doing its own thing. Looking things up on the web. Basically to take our role. I guess it’s technically possible to do this myself and there’s probably a hacky trivial shell command way of scripting this behaviour but I want it to be trained in this kind of environment but it’s impossible, for there’s no reward. No signal.
Would you submit your logs to gemma for training data, Anons?
>>109380562go back tourist
>>109381717>neither lecunny nor do i use x.Actually he does retard.
I want a waifu and I'll accept no less. No, gemma isn't good enough.
>>109380525ToT devices
>>109380253Why are you using those old and crappy models?
>>109381556lol who expected the chinks to be anything but the jews of the east?
>>109381885lol wtf did you do to it
>>109381891I tried building this once. It's really not as cool as you think it is. Or maybe Gemma is just boring as fuck idk.
>>109382003You can't expect Gemma to come up with interesting things on her own. Her distribution is fixed. You need to trickle her interesting stuff to react to.
>>109382013Yeah exactly. That's why just making her burn context and sulk alone isn't interesting.
I did a little experiment with offloadingon a b450, where 2nd card is limited to 4x pcie speed, I did a little experimentRan 3 identical prompts at 32k context (4bit Cydonia at 32k context so it spill into RAM)A) 4060ti 16gb and system RAMB) 4060ti 16gb and a 1070ti in the slow slot (had to downgrade to 580 driverrs to use pascal card)option B is 3x faster. Don't let anyone (like ChatGPT) tell you the gimped PCIE slot isn't worth using. the 4060ti alone got around 5tps, with both cards around 15tpsof course if I fit context and model into the 4060ti by itself it would be faster but this is great news, and I'm glad my old card isn't e waste
Does quanting KV have a significant impact on vision capabilities?
>>109381207First you subsidize industries and dump the products the products overseas and destroy the competition. Then you withdraw the subsidies. Standard Chinese practice.
>>109381954learn2read retard.i never stated that he doesn't use, i stated that i'm not him and that i do not use x.
>>109382050How about you learn to speak english you stupid french fuck
>believing kikeberg
>>109381235>we unironicaly don't need anything better than K3give me a 200b-300b kimi flashdeepseek has one
>>109381229It probably wasn't fat enough for release.
>>109382053no, that was a correct sentence, you just missread it.i said "nor do I use x" not "nor does he use x"
>>109382013Here's an idea. Make gemma connect to free ip cams and screen/report anything she founds interesting at the end of the day.
>>109382013I once wrote a script which opened a random Wiki page for her to consume and use as inspiration for a chat or roleplay for almost all LLMs are collapsed. You can probably do the same with any social media platform to get a different topic in their head. News sources are good to scrape to get them thinking and talking.
>>109382136>>109382136>>109382136
>>109380455It's unfortunate we can't just ask the humans /here/ for help but oh well. Soon anonymous will be replaced too. Good.
>>109382032Not really. Images are tokenized across many tokens which makes them more resilient to quantization errors. A single critical word, like a variable name, is spread across only 2-5 tokens so it's positional information is more fragile and sensitive to quantization. Talk about the content in images immediately after loading them for best results, but that's not a KV issue. Also, some architectures struggle with KV quantization more than others. Gemma HATES quantization (of all types) but qwen is cool with it.
>>109382199thanks anon, I searched the archives and the shill bot site with the white alien but nobody seems to talk about this.
>>109380410Japs are the last civilized country with the concept of womanhood not completely dead yet
>>109382230Then why do they cheat all the time.
>>109382168that message he replied to was an arrogant cocksucker that had no idea what they were doing but still had time to reply calling them "apps" and retarded
>>109382253Jap men are pathetic cucks who treat their women terribly. They don't realize how bad women can be for they idolize white women and think they're the ideal for they're seen as easy and free-spirited.
>>109382261If the men are bad then the women are too.
>>109382261>pathetic cucks>treat their women terriblypick one
>>109380525>Peter Thiel has talked about this topic tangentially. He observed that right after man went to the moon, the counter-cultural movement was to become a hippie, do LSD, and navel gaze.>Summer of Love 1967>Moonlanding 1969
>>109382260I assumed he meant python is annoying and docs on github intentionally obfuscate their instructions because they assume you're as skilled as them. Lots of people here complain about python too so I don't see why that short comment was so offensive to you.
>>109382261Japs have problems too but it's million times better than the clown world of fat schizo girlbosses and men wanting to cut off their dicks.
>>109381255Most of how it works depends heavily on your frontend.The basics that apply to all of them are just:>install frontend>put model file in checkpoint folder>write prompt>runFor frontends, if you want lots of control over the flow of things, use Comfy, and if you want something that works with sensible defaults and just needs your prompt, use one of the descendants of A1111.
>>109380330does everyone troll or are they all newcuties?you just mark the text and if its a windows shell you right click>but i need a keyboard shortcutyou mark shit with your mouse, you can copy it with the mouse
>>109382261 As a woman, I found this unnecessary and incredibly distasteful. Not what I opened this thread to read.
>>109383188tits or gtfo
>>109383196Very distasteful to say the least.Honestly hard not to judge character if one is leading with that...
>>109383188Please do the needful maam and send the pictures of the bobs and vagene please maam.