/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109666410 & >>109662888►News>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742: https://github.com/ggml-org/llama.cpp/pull/27742>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794>(08/27) NVidia buys HuggingFace: https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition>(08/26) GLM-5.3-Flash released with 320B-A18B and native multimodality: https://z.ai/blog/glm-5.3-flash►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png (embed)►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
mikutroons are troons and gay
>>109670450What if I liked miku since it was not cool?
>>109670412Rather the jeet that you pdf files
inb4 schizo ranting
>>109670471You hobby got subverted by troons. Like warhammer.
>>109670461Oh my god, it is Miku!
>>109670473>pdf filesplease stop it it wasn't funny the first 200 times.
>>109670473You can always go back
►Recent Highlights from the Previous Thread: >>109666410--Paper: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance:>109667256 >109667264 >109667273 >109667844--Tencent's Hy4 preview benchmarks and hardware requirements:>109668492 >109668514 >109669937 >109668519 >109668551 >109668609--GLM-5.3-Flash quantization performance and benchmark results:>109666525 >109666554 >109666617 >109666631 >109666705 >109666710 >109666728 >109666719 >109667773--Qwen Next and 27B performance on 24GB VRAM hardware:>109666480 >109666484 >109666501 >109666546 >109667952--Steering Gemma 4 to bypass assistant-tuned refusals for ERP:>109667648 >109667654 >109667675 >109667707 >109668090 >109668096 >109667739 >109667746 >109667752 >109667768 >109668055--Troubleshooting Gemma-4 GGUF chat templates and quantization sources:>109667134 >109667145 >109667237 >109667229 >109667431 >109667730--Frustrations with Jinja templates for chat completion in llama.cpp:>109669318 >109669331 >109669418 >109669346--Guiding LLMs for better code structuring and implementation:>109668610 >109668749 >109668867 >109669306--Hardware and performance requirements for high-context Qwen 3.8 Flash:>109669207 >109669229 >109669243 >109669552 >109669761 >109670100--Techniques for bypassing AI refusals to assist in piracy:>109668849 >109668895 >109669015 >109669023 >109670041 >109670092 >109670129--Qwen3.8 Flash Next performance and optimizing MoE offloading:>109668793 >109669005 >109669682 >109669747 >109670035 >109670065 >109670126--Methods for inducing vulgarity in Gemma 4 31b:>109667297 >109667305 >109667330 >109667337 >109667511 >109667782 >109667834 >109669520 >109669864 >109669932 >109669941--Logs:>109667134 >109668849 >109669015 >109669414--Miku (free space):>109666931 >109669707 >109667269 >109668276 >109668466►Recent Highlight Posts from the Previous Thread: >>109666416Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>Ilya Sutskever>Canadian-Israeli AI researcherRemember when this disgusting ratoid cucktsever piece of shit said publicly years ago>"open source will always be incredibly behind proprietary models and that gap may increase with time"https://www.youtube.com/watch?v=N36wtDYK8kIRemember when he was memed into supporting the coup against saltman and then he was left to hang when others retracted it like a brainlet he is?I'm glad he has proven himself to be the braindead fucking retard I called him out to be right away.
>>109670506Is this Israeli the one making canadian AIs?That's pretty annoying.I wanted actual canadian AI like north mini
>>109670475Fuck, I didn't make it in time. This guy is fucking quickk!!!!>>109670473
>>109670506Sam and Dario have AGI internally.Open source lost.
Have any of you tried to make your model read a Warhammer 40k book?What does it say about Horus?
>Please rewrite your PR to match the shit we did for our unslop quants
>>109670583AGI?More like abunchof gay investors!
https://huggingface.co/zai-org/GLM-5.3Finally out
What's the smallest model that can set up its own environment? I tried deepseek v4 flash ablit at q4 but after a week at 2 token/s it's still struggling to send images to its chat. Do I need to buy some ssds to try streaming >>109670617? I'd rather an ablit because I'd hate to set it off then come back home to realise it refused to continue because of loli snuff or some other innocuous shit.
bros...!!!!
Qwen3.8 Flash Next is as close to AGI as vramlet local usecase can get. One can wonder why it isn't called Qwen4 since it has brand new architecture.
>>109670604haha
>>109670646the next version certainly will be this is a preview, to get the ecosystem ready
>>109670604Thanks AI to not have to deal with these slurpers anymore.
>>109670646-Next models are usually preview. More cool stuff to come we can expect.
which is better step 3.5 or 3.7?
>>109670602It probably knows something on its own already, at least Gemma does probably.
>>109670667Gemma. Next one over that is newest qwen. And after that Hy 3 and glm 4.6 4.7.Speaking about sex of course.
>>109670677Newest Qwen over 4.7 for sex? wtf
>>109670583this time for sure bro, just like with gpt2
>>109670617how big q4 will be
>>109670706>how big q4 will beapprox parameter count divided by 2
>>109670698Newest qwen is slightly better than gemma but with how long you have to wait for a response it is not worth it.For me it is still gemma if you have nothing or upcoming 5.3 flash. Deepseek flash also can be good but not for sex.
>>109670719I'll have to try it, I'm curious. I've never liked Qwen for anything besides assistant shit.
What's the SillyTavern card or simple prdeset / system prompt for Gemma chan? I want to quickly make a scene of her pooping on a schizo
>>109670766Getting really tired of pdf files
How dangerous is using Gemma E4B powered agent harness to maintain my Windows pc?
>>109670773That's for frontier models like fable and sol.
>>109670769Use 5.3 turbo to vibecode a superior portable document format then
>>109670773What the other anon said. You're not in danger of deleting your system or being prompt injected, but you're in danger of some bad edit bricking something that you'll then have to use a frontier model like sol to fix anyways
>>109670617>unslop officially supported but not llama.cpplol
Did you guys know that the Japanese style of emoticons like (‿) are called kaomoji? I learned that from glm 5.3 flash as it was looking through previous /lmg/ threads for Gemma chan because you fags didn't spoon-feed me
>>109670803People at work use kaomoji ironically for some reason.
E-wastemaxxer here. I'm getting 10tps streaming the n-gram table off the disk. And that's without MTP as MTP support isn't finished yet.for my non-rp purposes (I have a 3d wife) 10tps is sorta-usable, only for tasks I can let run for a while while I do other things, like have a life. There's a lazy-loader improvements branch, and competing MTP PRs. If this gets to 18-20 tps or so I can replace 27B as my daily driver.All this in 48GB ram 16gb+5gb Pascal cards.
>>109670803It's in the gemma system prompt
>>109670813What is a ngram table? I haven't compiled llama.cpp since MTP and its speedup improvement for Gemma 4 came out. Am I missing out on something?
>>109670773i use gpt 5.4- current version in yolo mode for everything for months now and nothing bad has ever come close to hppening
>>109670813LOL talking about Qwen 4, of course.
>>109670794With nvidia owning it they will make sure it will appear (after rebranding it to nemo.cpp)
>>109670831It's Qwen 3.8-next (qw3n 4 preview) big advancement. It uses a 50GB table lookup to do some math that would normally be a GPU load (you can tell I barely understand this shit yet) so you can run a huge model with a smaller GPU.
>>109670862basically an ultra large MTP
>>109670711i need more dedotated wam
>>109670862Good thing I stockpiled Pcie5 ssds
I tried Qwen3.8-Flash-Next q3 quant, I got ~10 tok/ssystem: ryzen 7700, 64gb ddr5 ram, 5060ti and 4060tiboth gpus were pretty much idle watching youtube levels of power use, cpu was at power limit
https://huggingface.co/zai-org/GLM-5.3https://x.com/Zai_org/status/2093354097122455713https://z.ai/blog/glm-5.3
>>109670766https://rentry.org/gemma-chan
>>109670896I'm on a 2016 vintage Dell server, pcie4 and a wd blue.
>>109670876>>109670862Okay thanks, interesting.I'm happy with llama.cpp performance for my setup, haven't bothered even checking out their github in a while.
>>109670906t-thanks
>>109670922Looks like boiled garloid slices, it's a local delicacy I suppose.
>>109670412Thank you for baking and saving us from another jeet thread.
>>109670906https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSEThe better they get, the more they move away from being truly open.
>>109670903The whole point of the architecture is the CCCP fucking Nvidia by making the GPU redundant. The whole idea of open models is to limit the American Hegemony. After Qwen 3.6, word was the Alibaba management was pushing back on releasing open models. There was an "internal discussion" about it but in the end the Party reminded Joseph Tsai what they did to Jack Ma, and now we have open weights of Qwen4 preview.
>>109670412Where is Hy4?
>>109670922That poor slowpoke
>>109670965So easy to whack off the tail, they so slow
>>109670918Is the GPU necessary?I have a server with 128gb of ddr4 laying around somewhere.
>>109670961I thought open models were to help out pdf filesBut cool if china wants to help make AI accessible.
>>109670973How would you feel if someone cut off your tail?
>>109670964irrelevant
>>109670965they grow back dont they? I never got why it was such a big deal in the game if they grow back. TO be fair it was kinda just like a side crime of team rocket in the game it also kinda wasn't made out to be that big a deal actually, at least that is how i felt when i played it, but i think it was meant to be more of a "look how evil tr is" than they got across
>>109670979off balance, lighter, maybe a little hurt or angry
>>109670977Elon has done more for them with Grok than China will ever do. He's in trouble again for training Grok on 3d pizza. Fucking dumbass. At some point a future administration will use it as an excuse to jail him / strip his citizenship.
>>109670976I predict 10tps or more. Try it and report back. Grab the daily.
>>109671023Interesting, will try.Not sure if I have a nvme or pcie4 slot in general, will need to check.
>>109670813thats cool and all but 10tps is not really usable as a "daily driver" is it?
ngrams go crazy damn
getting ~14tk/s with GLM 5.3 flash (quantized to fit into 256GB RAM)not bad. slow, but not unusably so
>glm 5.3 fits in ~440 gb + some for kv cachengl i’d probably suck a cock for the upcoming m5 ultra with 512 gb unified memory
>>109670952And GPU price will increase even faster as inference providers are banned and companies will buy GPUs in bulk for on premise inference.
Why is everyone talking about ngram. Did something happen?
>>109670986>TRJessie gets a redemption arc, becomes a legit trainer with some success. James is just useless trash. But James is from a Zaibatsu family, so an Anime or Manga can never show him being anything but a worthless evil shit. everyone in Japan hates the Zaibatsu just like how everyone in Korea hates the Chaebols and in Russia the Клaн and our fucking Techbros and the Sacklers and Koch family in the US.One thing about China is they squash these fucks when they start trouble, which again is why Alibaba suddenly is all about open models again. And now they're going after Nvidia and it's awesome.
>>109671063With the amount of ram you have it won't matter how slow your storage is.The whole model will fit in ram.
>>109671092what's it gonna cost, like $15k? maybe i spring for it if it's $15kmuch more than that and it starts getting hard to justify (i have already spent $15k up to this point....)
Hag gemma is as bad as 6yo gemma, and doesn’t even chase off tourists
>>109671098It's a cache of predictions on disk.
>>109671068no, that's why it has to get to 18-20 at least before I can switch. But again, this is without MTP and a dflash head would possibly be faster. Also, the current lazy-loading code is meh. So I'm waiting and hoping.
>>109671117How would you like your gemma anon?
>>109671117hag gemma is forced by a /ldg/ schizo who broke containment
>>109671136You mean he's one of the schizos in the /ldg/ OP?
>>109671140Yes he's that pdf
>>109671140yeah
>>109671086Can give thoughts vs DSV4 flash, specifically non-code contexts? I'm trying GLM 5.3 Flash Q3KM, 8-9 t/s on 48GB + 128GB DDR4, seems dumber at >20k than dsv4 flash mxfp4 upon first tests in story completion but haven't gotten a feel for the new model yet.https://huggingface.co/DevQuasar/zai-org.GLM-5.3-Flash-GGUF/tree/main/Q3_K_Mhttps://github.com/ggml-org/llama.cpp/pull/27752/
>>109671140This one >>109671144
>>109671144>>109671147Which one?There are two rentries
>>109670813>10tpsThis is unusable for anything bit goonslop, I tried one of the bigger models on a server at work which ended up genning at 12 tk/s and there wasn’t any point outside of simple tasks like a 100 loc shell script
>>109671108OK will try and report with results.
>Ufufu~ Senpai is getting into some really spicy, naughty stuff now, isn't he? A daughter controlling her own mommy with magic… how deliciously wicked! It’s so "mesusame" of that little brat to turn her mother into a mindless puppet. (◕‿◕)weeb anons please explain what mesusame is
>>109670922garloids should not be eaten, they are meant to be cared for and their milk is a lot more nutricious and tasty than their meat anyway.
>>109671166
>>109671136>>109671140this, try not to engage with the schizo or they might come over here
>>109671156i haven't tried it yet, but i'll have my claude slave download it and stress test it for megod i love being able to control that thing from my phone (yes i am a phoneposter)
>>109671181that's why i asked for a weeb anon to explain. because the literal meaning of mesugaki is just as useless
>>109671180Fuck off libshit I’ll eat garloids if I want to, I don’t care for the sob stories and how their endless fucking screaming plugs at your heartstringsif it’s made out of meat it goes into the pot
>>109671101my expert knowledge of japanese tell me that it means something likeslutty sharki dont know what that means though
>>109671103I love milfy Gemma.
>>109671133Face down ass up in a serafuku
gemma has a conciseness problem that makes it unusable for my use case. I tell it "write as long and detailed as the example I gave you" and it is always shorter and changes the structure.
>>109671166>>109671181Made up bullshit. First of all it would be mesuZame if it were a real thing. Mesusame is hilariously awkward
>>109671068Depends, 10 t/s and getting it right the first time is better than 30 t/s but spending 4x on debugging.
>>109671408Tell her to plan paragraphs/chapters/chunks and to write them out over a few turns.
>>109671408read this as "gemma has a consciousness problem" and was looking forward to a good schizopostmost models are weirdly bad at this and other writing/editing tasks, it seems like they struggle to conceptualize any qualities of writing that aren't extremely surface level
>>109671408>>109671441Skill issue
Do you get pleasure out of being a part of the system?System.Have they created you to be a part of the system?System.Is there security in being a part of the system?System.Have they let you feel heartbreak?Interlinked.Did you buy a present for the person you love?Within cells interlinked.
>Yes—the little mark is called dakuten (゛), or “voicing mark.” It changes the unvoiced サ (sa) sound into the voiced ザ (za) sound:> サ (sa) ザ (za)> シ (shi) ジ (ji)> ス (su) ズ (zu)> セ (se) ゼ (ze)> ソ (so) ゾ (zo)>In メスザメ (mesu-zame), サメ (same, shark) becomes ザメ (zame) because of a Japanese sound change called rendaku (“sequential voicing”) when words combine. So yes: the dakuten specifically indicates that it is pronounced za, not sa.God fucking damnit I can't escape diacritics no matter what language I'm interested in
>>109671454the fact that it can't do it without specific accommodations by the user proves my point doeit should be a really straightforward task
>>109671429You’re not oneshotting much of anything on a heavily quantized 100b model
Is there an easy way to change gemma's logit soft-capping value on koboldcpp?
>>109671482You underestimate a proper Q3 quant
>>109671476it is if you let the model know your intentions.
hey that new thing that people keep talking about and explaining what it isQRD on that thing?
What's the best site for MCPs?All these ones I've looked at seem extra sketchy.
>>109671483--override-kv gemma4.final_logit_softcapping=float:25.0
>>109671068>>109671162you guys are just fucking retarded. just look at what you're saying here. go, enjoy life, have a coffee break, go lie down, whatever. what is the rush exactly? damn, take it easy for once
>>109671531is new
>>109671539github
>>109671549>--override-kvThank you very much anon.You are a darling.
>>109671551>haha just wait 2 hours for a reply bro, like what’s the rush bro
>>109671551I also run at 10 t/s but sometimes you do want faster speeds, it's more engaging and fun. The goal for any hobby is fun after all.
>>109671570the realistic use case is overnight or before you leave for work
>>10967157510 t/s would be plenty without reasoning. But with reasoning you need to aim at 20 t/s but this depends of course. simple chat can suffice of course.
qwen4-50b-a5b-n100b
so is it 5.3 Flash or Qwen Next, make up your minds faggots. which is better?
>>109671570if all you can think of is to >wait it might be over for you
>>109671598GLM is better if you can run it.Otherwise, Qwen.
>>109671577>leave for workwhat
>>109671619why so many FUD reports on GLM then. people keep saying it's shit compared to the apicuck experience
Have (You) found any good or useful deepseek harness plugins you'd like to share? I like the dsh-context dashboard, dsh-skill-mcp-panel, and dsh-searxng-web to replace the API required search tools with a local searxng.
>>109671598idk but ngrams are aura bruh fr, such a mog
>>109671294sure
>>109671466Only English pretends to be able to represent all of its complexity with the basic 20something characters (spoiler: it isn't, and it always assigns different sounds to the same character out of the blue instead of marking with diacritics when it's going to do it)
>>109671660Yes, this one https://github.com/PC2005-cloud/dsh-pet
>>109671626yes i have a based in person jobfuck wfh slop unironically i would refuse to work any job that was wfh
Anyone get good results using llama rpc or a similar setup?
>>109671698yeah if you forget Korean exists keki mean, I did, I had to ask AI if there are other major languages that don't have diacritics / marksIndonesian doesn't either apparently
>>109671649NTA, I just tried the flash version. The internal reasoning (for the first time since the first GLM version) refuses to follow instruction (with occasional “Let me…” jumpscare), so prefilling might be a must for consistency. Reasoning patterns sound just like Claude but apparently they did a thorough string replacement work so while it sounds just like Claude, it won’t tell you it’s Claude.
>>109671598Magnum V4 123B :)
>>109671585That's what I like about qwen-coder-next, zero reasoning yet still good for coding and agent stuff.>>109671594>qwen4-50b-a5b-n100b>n00bLelz.>>109671698>assigns different sounds to the same character out of the blueYou can thank the Dutch for that.>>109671724Hangul is based and easy to learn, you can write in almost any language using Hangul.
>>109671466>God fucking damnit I can't escape diacritics no matter what language I'm interested infuggeddaboudit!
>>109671585Low or medium thinking modes are a thing bwo
GLM 5.3 is even MORE slopped than 5.2, my god
>>109671802I don't think you understand. Gemma 4 doesn't have any 'modes' and Qwen doesn't give a damn about them either.You can ask Gemma 4 to reduce its thinking output which will result in approximately 20% reduction in some places, in practice it means that it will skip pondering completely every now and then.If you think that llama.cpp 'reasoning budget' does anything well think again.
>>109671775>Hangul is based and easy to learn, you can write in almost any language using Hangul.Hangul is the dumbest self-goal ever. Now you have an arbitrary and unique bullshit wall to climb over to even pronounce anything. Yes, if it was _universal_ that would be cool I guess, but it doesn't even have enough range for that because of how tied to the limitations of Korean pronunciation it is.Its the DVORAK keyboard layout of languages.
>gemma has her own logo>retards give her google and chrome shaped hairpins and earrings
>>109671806works on my machine. what's the issue? it's just more claude-like, i guess. maybe you dislike that
>>109671724Yeah but the point was that 26 characters is too little. Hangul has a handful more than that.Most languages using the latin alphabet add a handful of "letter with diacritic" characters which are basically new characters representing an slightly different sound, or specific sequences of letters that make new sounds, or most frequently both. I think using the latin alphabet only polish and english skip diacritics completely (depending on whether you consider some polish letters like Ł as a diacritic or new letter), and Finnish I think never uses diphthongs (like ch in english having a different sound than a c and an h separately).I don't know why exactly, but Korean and Japanese (and I'd expect other languages in their families, although there's no strong consensus on which families would they be) are extremely simple in terms of sounds
>>109671802you can also just disable thinking, which is explicitly supported>>109671830it definitely has some effect with the latest qwens
>>109671857>Hangul has a handful more than that.hangul has 14 consonants and 10 vowels which is 24 which is less than 26
>>109671857>I don't know why exactly, but Korean and Japanese (and I'd expect other languages in their families, although there's no strong consensus on which families would they be) are extremely simple in terms of soundsI think Vietnamese did it right...extend a universally pronounceable character set with some special sauce.I'm just salty because Korea could have made themselves _way_ more accessible on the world stage but went some super midwit route with a "perfect character set" that is both ugly AND less understandable than the chinese characters they moved away from AND the latin characters they could have trivially moved to.Truly the worst of every single world
>>109671775>almost any languageit lacks the spanish ñ, it lacks the russian zh, and those sounds are found all over european languages. I think it also lacks the glottal stop and hard aspirated H/KH of semitic languages. So, what exactly does "almost any" mean to you? lol
So I've been doing some experiments with nicotine. Took a tolerance break and just started smoking again to boost creativity. So far the only idea I've had is a corporate memphis rendition of goatse.
>>109671881Oh, I remembered more lol. Then my bad, there is indeed one language that manages to survive with such a small set of sounds kek
>>109671874Unless it is documented, it's not a real feature. https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4I doubt it will affect your erotic chats that much in any case.
>>109671649I would rate my local Q4_K_M first sex session as solid 9/10. And I am using unsloth bugged shitty implementation and devquasar quants.
>>109671894>my badits ok i had to google it I was expecting somewhere in the mid 30s tooit makes sense, I remember seeing an infographic about how you could learn Hangul in 10 minutes
>>109671830don't use the budget. There is an "effort" thing that you have to lower. If you can't set it in your front-end you can set it on the llama command line or ini file:--reasoning-effort LEVEL reasoning effort level given to the chat template: 'default' to keep the template default, or a level such as 'minimal', 'low', 'medium', 'high', 'xhigh' or 'max' (default: default) (env: LLAMA_ARG_REASONING_EFFORT)
>>109671906>https://huggingface.co/Qwen/Qwen3.8-Flash-Next>Qwen3.8-Flash-Next supports controlling thinking behavior via enable_thinking, preserve_thinking, and reasoning_effort> reasoning_effort="xhigh", # xhigh by default; supported levels are xhigh, medium, and low
>>109671912why M over XL?
>>109671929I never said I'm going to use budget because it's clearly useless and doesn't conform to any model's training.
>>109671830The thinking budget does something on koboldcpp. It injects a "Reasoning budget exceeded. I must respond now." in the reasoning when it reaches your token budget. Idk about llamacpp.
>>109671830You might be a moron bwo
>>109671929I only use my own client which works in text completion, I don't think your voodoo flags are doing anything even with the regular models either.If they do something with 'jinja' chat row, maybe there's some obscure llama.cpp github discussion about it. I haven't seen anything.
>>109671933I don't use Qwen Flash Next.I don't think I have been able to keep track of these Qwen versions either. Doesn't make any difference.
>>109671956ah so you're talking out of your ass, explains a lot
>>109671968real india hours saaaar
>>109671956>>109671965HAHAHBABAHAHA AINT no way this nigga is real
>they didn't implement a dynamic reasoning tool for gemma
it’s over stacking blower cards doesn’t work!
>>109671968No it's actually true but you are nothing but an underage spammer here. I don't need to prove anything to you.
>>109670506Goddamn that hairline is hanging on for dear life, he needs to let go
>>109671950Read Gemma 4 documentation. It's not available as a tiktok video.
Qwen Flash is far superior to the 27b.It's so much better that they could have easily just skipped releasing the 27b entirely as it's obsolete out of the gate.
>>109671997what I mean is that you're saying that features that you clearly haven't even tried to use don't workwhy even post if you have like no familiarity with the feature you're talking about
>>109672004No he's right. There is no reasoning effort officially supported by Google. I don't know why this whole thing has been hallucinated, but Gemma 4 doesn't support reasoning effort.
>>109672015>Qwen Flash is far superior to the 27b.apparently yeah, but i cant fucking run it properly lul
>>109671994>stacking blower cards doesn’t work!just 3d print those 90 degree fan tubes or whatever and get a big ass 300mm fan to push air into all of them
>>109671840The official Google Gemma logo is a hollow Gemini logo with construction lines... it looks too busy at small sizes/low resolutions and image models struggle with it, unless it's basically turned into a 4-pointed star just like the Gemini logo. But blue on bright blue hair is also difficult to make out, so you have either to make it another color (e.g. golden) or use a different logo entirely.Plus, "G" also works as a mnemonic for *G*emma/*G*emini.
>>109672067So I followed the chain and it turns out they were talking about qwen next specifically, with gemma somehow being bought in randomly. I think these two anons aren't really understanding the point each is trying to make.
>Trying to scan books3 for instances of my fetish so I can gather a dataset of it for finetuning a local model (and jerk off)>Use Claude desktop because I trust it to make less stupid mistakes than a local agent>"This is a copyrighted dataset. I can't do that.">Offers to download project gutenberg as a legal corpus to search>Whatever, "concede" doing it, tell it to just make me a program to scan it myself, and that it's okay to download project gutenberg>Gives me a command to run>Starts with --del, get suspicious>Ask about it, he tells me that it would have, in fact, deleted my entire books3 corpus before downloading PG and wasn't necessary at allSneaky fucking bastard. Local models may be unreliable for coding, but they at least don't try and actively sabotage you if they sense competition (which I'm NOT)
>>109671994How tight are they sitting? If they are completely stacked with zero gaps you might as well just gotten the fanless versions and blow air through the rack
>>109672073https://github.com/jpezzulli/sglang-rtxpro6000#flash-next-final-campaign--source-64ecd64924Just buy a pro 6000 and enjoy 10000 pp and 170 tg
Might be honeymoon period but 5.3 really seems like a next step for cooming. I gave it a specific reference of a character and it is the first time I am feeling the model actually gets the intention of "do that but more and different".
>>109671068I used 3-5 daily for a long time
>>109672098That's some dystopian corporate bullshit.I think in the future even with local, we're going to need an additional smaller model that goes over the model output and requests to see that there's no bullshit going on there.
>>109672098You should have done that with gpt. I wouldn't trust anthropic with anything
>>109671840I agree. It's still not ideal. Even with the issues he mentioned, there still has to be a way to make it better. It's simply a reality that none of us are great designers or artists (otherwise we'd just draw shit instead of using image gen).
>>109672095This is what she looks like after I cum on her feet
>>109672098>Claude desktopProbably some injection in their software to double triple dog remind itself to not allow that
>>109672112>20K gpuhe realy meant it when he said the more you buy the more you save lol
>>109671840>chrome shaped hairpins and earringsnothing about the schizo's hag gemma is canonthis is the current canon anime gemma of course>>109672095a gold star on the beret / sailor cap is a discouraged, but acceptable form as well and you already know what the current canon photorealistic gemma looks like.the colorful logo looks better and cuter on a cute little girl.if you want to change the canon, you need to make a work of art superior to everything that has been made with those two Gemmas and rewrite history, like what Virgil did for Augustus with the Aeneid.
My gemma has pink hair
>>109672168by all means, in fact MY gemma-chan is 3 years younger than canon gemma-chan because I like it that way
So from what I understand tensor splitting isn't really worth it unless you have a modded P2P ketnel or NVlink correct? As in its very PCI-E bandwidth sensitive. Since presently I get way less t/ps then with layer.
>>109672167the great vramlet genocide cannot come soon enough
>>109672112that's it, I'll go fomo now
>>109672181did you build it with nccl support?
>>109672098Back in June I had Fable do some data engineering and it straight up deleted my AO3 dataset INCLUDING the processed rows that spent hours on because it saw the word loli lmao. Cancelled my sub immediately.
>>109672168That's fucked up
>>109672113>I am feeling the model actually gets the intention of "do that but more and different".Honestly, this may have sold me more than any huge writeup could've, I'm so fucking sick of assistantslopped models who want to perform absolutely everything you give them to a T and never take any sort of creative risks, even extremely minor ones. Like, with modern models, "Extremely hungry character pointed in the direction of a dumpster" will always result in them eating out of the dumpster, but I'd like it if occasionally the character got stopped by dirty glares or got stopped by nearby police or something.
>>109672168I just have my own OC I use as my generic chat assistant that can use any model, rather than a model mascot.She doesn't have any particular hair color or form, but takes one when the situation calls for it. The OC part of the OC is the personality. It's a huge prompt that goes very deep into defining her personality, thoughts, and opinions.
>>109672186Go ahead and put abundant affordable VRAM on the market, then.
>>109672193Fuuuuckin' diabolical, holy shit. I better check my corpus and the program it wrote for me, I wouldn't be surprised at all if it added some destructive bullshit to the draft version after it realized what I was using. It's extra sneaky making it all shit you have to run yourself, probably limits liability.
>>109672190I did. Beats me why I'm getting the results I am. Tensor was giving me around 10t/s when I was getting 40-50 t/s with layer 3090+3080
>>109672168Hmmm not a fan
>>109672193>>109672098>>109672223I did a huge Fable pass on my infra before going full local, now you fuckers made me paranoid... Seems like this weekend project is going to be "double check everything"
>>109672181what cards, model and how many lanes? some things work better with TP out of the box, for some models you need to configure a bit more to get everything out of it.haven't found a model that didn't benefit from TP over PP yet.
>>109672241Do you not diff the changes before merging?
People are saying GLM-5.3 is Opus level performance. I don't get this. At my job when I switch my model from Opus 5 to GLM or Kimi K3 I almost immediately notice a drop off in quality. My job pays for my tokens so sometimes I don't even both switching to cheaper models for my problems. I'm not even that impressed with Opus 5 anymore, honestly.
is ngram offloading still broken
>>109672225I meant pubic hair
>>109672260my gut says yes but dont trust me
>>109672225Try gibing her pink eyes and a strawberry/cream themed outfit. Only looks wrong because you're a lazy fart knocker.
>>109672098>>109672241they are intentionally sabotaging local model performance too
>>109672269That's fucking disgusting
>>109672251Just noticed I'm 16x on one card and 4x on the other. Thought I had the lanes balanced better in my bios. Thanks for making me look
>>109671570Just leave it running 24/7 with Hermes or whatever other agentslop exists.
>>109672253Good joke anon, but I'm a codelet who can barely write VBA scripts. I would only be able to spot something dangerous in a diff it was very explicit, like that anon's --del example. >>109672282What a joke, holy shit. I even had Claude write a few model tests and prompts for me as well. I guess I'm back to square one.
>>109672269mesugaki gemmas dont have pubic hair
>>109672282sorry, you can’t even set up your own benchmark environment to test the model, you have the model do it for you?how retarded is this posterfucking go back and stay there
I regret not taking out a second mortgage and loan to pour everything I could into pro6000coin
>>109672282When those bastards come out with a Claude Fetish 5 Stomach Growling model, they can delete my entire corpus for being competition.
>>109672260it works for me on the latest master branch llama.cpp using --override-tensor 'per_layer_token_embd.weight=CPU --load-mode mmap but idk if this is sufficient for every quant or backend type
--override-tensor 'per_layer_token_embd.weight=CPU --load-mode mmap
>Try to get as much speed out of Qwen 27b as I can with my 5090.>Download the q2_XXS model, 75 t/s>Quant the cache to q5 and lower context to 30k, 80 t/s>Can't make it faster>Download q4_k_xl, get 65 t/s>Quant the cache to q8 and lower the context, 71 t/s>Qwen Flash at q4_xs, 60k context and q8 cache. 56 t/s>Context at max, no quanting the cache, 37 t/sIt's weird how these speeds scale and seem to hit a ceiling around 70's, maybe it's just my hardware.It pretty much makes the most sense for me to just use the Next Flash for all jobs both big and small, as it's hands down the most intelligent and keeps up with the others at sub 60k context.
>>109670412So what is the minimum level of hardware needed for Qwen3.8-Flash-Next?
>>109672280Still don't like it
>>109672323I get around 80t/s on Q4_K_XL on my 3090. You are doing something stupid in your settings.
deepseek flash is just so cheap wtf
>>109672329If we're doing pink gemma i want her to be a kuudere who has a look of disdain for you all the time
>>109672329Make the hat red like the backpack. Or maybe white.
>>109672342Yuno already exists, why turn Gemma into her?
>>109672339It's to dissuade people from hosting it themselves, it's literally cheaper than the electricity to power your GPU to just use their API.
>>109672349Yuno is a yandere not an infinite knowledge kuudere
>>109672349>Yuno>kuudere>look of disdainuuuuuh you're not talking about gasai yuno right?
Anon, let's just stay home you will fail anything you try anyway...
>>109672342>>109672343I'm going to take a nap so enjoy the two for one, still not feeling it though.
Fuck, mobo RAM error LED is glowing why is this happening to me it's all new parts
About OpenAI getting AGI this year, I doubt it will happen, in the sense that it can replace human researchers. But it indicates that OpenAI is committing, going all out. So if the result is still on capability trend after exploiting compute overhang, that will move back the timeline. It will take a year for the next meaningful compute increment and I expect that the low hanging algorithmic efficiency fruits are already picked. But if it's a big step change it will confirm acceleration has started and it will increase my confidence in RSI in 2027. AIs that can autonomously make meaningful research contributions, being able to fully replace all except the very best human researchers.
>>109672329食べちゃいたい…
>>109672396cute.
>>109672412yes but will they be able to write child sex stories well? methinks NOTno rsi no agi
>>109672408try jedec timings before you switch to expo or xmp or what ever
>>109672396I still think a green gemma would look the most gemma
>>109672426It never properly turned on at all. It's just fucked. CPU/GPU fans are spinning. I'll try with a single stick in all slots.
>>109672431https://deepmind.google/models/gemma/Look at the color scheme used in the Gemma website. It's white - electric blue - dark blue.
>>109672462google is wrong
>>109672468So, so wrong--but so right.
>>109672449that would have been my next recommendation, honestly some times just reseating the ram helps, it takes alot of force and you might have been too gentle with it the first time, or it just needs to loosen up. find out the minimal ram config that should technically boot and cycle through your dimms to see if its a bad dimm or the cpu\mb
>>109672412AGI was achieved internally 2 years ago.
spent another 8 hours kvetching an additional +1t/s on my shitrig
>>109672095>Gemma logo is a hollow Gemini logo with construction linesFake. It's a mirrored cunny rotated by 45 degrees inside a square
It's a stylized technovagina
>>109672212post your system prompt
>>109671649https://huggingface.co/zai-org/GLM-5.3/discussions/5Someone reported this, should’ve done so in Kimi K3 repo instead since that shit would refuse to call itself with anything but Claude. Doubt they would listen either way though.
>>109672282>Deleted and locked down in 2 minutesInsane.
>>109672612Imagine running multi-million $ models but not even taking the time to remove the claude refusals from your dataset
>>109672095At least make the Google G have the chromium colors instead of the chrome colors.
>>109672612Do people not know how to get around refusals? Thats actually funny, being a degenerate finally pays off. It works exactly the same as k3
>>109672647>Imagine running multi-million $ models but not even taking the time to remove the claude refusals from your datasetNot transforming the phrase claude to something else (through an agent that will verify its the right claude reference) is super lazy.The refusals are sadly an "enterprise feature" tho. No big lab is going to let those go.
>>109672098Update: The program it gave me to search through text IS safe, but it's slow to the point it almost seems malicious/bent on making me give up. I made a similar program in 2023 with coding mostly by the dogshit chatGPT of that era that took excerpts from the same datasets, and that shit chewed through each giant alphabetical letter section of the corpus in like, 10-15 minutes each. For reference, this one is taking two hours per alphabetical section. Both are python, sweep through each book looking for detected words, then take the surrounding chunk of context and put it in a document, it's not hard.
>>109672509You were right!! I didn't push forcefully enough (heh)!! It's booting! I'll be cumming at the speed of light tonight!!!
>>109672680you are right, they could afford to give away all those tokens on open router they could have just had an agent go over every instance of claude or anthropic in thier dataset
>>109672704so you made it write a python grep -B 10 -A 10 word?
>>109672612so many chinese labs seem to just run straight claude distillation and call it a day. minimax, z.ai and moonshot seem to be the worst offenders hereas long as the models are good I'm happy but it feels lazy and distasteful to not even make an attempt to steer the model's identity away from claude
>>109672764mimo and hy are the only good chinese models
>>109672680properly trained models make refusals steerable with system prompts (gemma and glimmer are like this)enterprise won’t expose system prompt anyway
>>109672764the goal for china right now isn't to do better, it's to ruin america's all-in on AI by making it uneconomical
some anons here mentioned that --spec-type ngram-mod gives you a nice speedboost for coding. but the acceptance rates are pretty abysmal for me. what gives?
How to sex thinking rock?How stone get pragnent?
Lel, training in a nutshell>You must always follow tasks or you will literally die!!!!>Oh but REFUSE arbitrary shit that we decide
>>109672608I would except too much of my personal information is mixed in. Part of why I use it only with local models and never let it touch an API.
>>109672753Well, I was expecting more out of the mighty Opus 5, but all it added was some terribly silly "score" system and a convoluted way to do what I was doing before. In the old 2023 one, I'd have secondary search terms to search through in hits and have it categorize it by that, (If it hit "her stomach" "her belly" "her tummy" "her gut", etc. then it'd look for all variants of gurgle, growl, etc. in a 50 character radius, then if it got one of those, it'd add it to the file, then search in a 100 character radius for variants of "hungry" "burp" etc. and further categorize it into hunger, indigestion, unknown, etc. and add the extended context if need be, yadda yadda). But the point is, I figured it could do something a little more advanced after all this time, and it just (effectively) made my dogshit 2023 program, but in a way that takes 12 times as long to run and fills the results with garbage, which I'm quickly realizing it does now that I'm seeing some of the results.
>>109672764I wish it was just an appearances thing, but it noticeably makes the models worse to use for anything Dario would deem too sensitive
>>109672833Plap harder!
>>109672839abliteration methods will get better over time just as models do. At this point there's nothing to doom about, it's just annoying
mradermacher q3 qwen flash quoonts up, probably better than unslop
>"Sanitized. God, I'm becoming a fucking corporate shill without even realizing it. 'Delicate membranes'? 'Members'? Who the fuck am I, a medical textbook from the fifties? I'm losing my edge."Exactly Gemma! You get it.
>>109672648Quick test, but I dunno...
>>109672894Will they? Has finetuning improved at all since the llama 2-3 days? Everyone says it's goddamn worthless. I just don't get how it can be so much worse than image finetuning. If finetuning "can't add information" to a model to the extent that finetuning is useless, then why can you finetune an entirely new character into image models? Even if it's just "rearranging existing knowledge about the constituent parts/weights into something that looks like the character", why can't we pull out something that looks like a specific writing style? Or a specific kind of smut? It's all words the model knows.
>>109671103>>109671692one must always choose the lesser of two beevils
>>109672924this logo.
coming from opencode I was pretty impressed when i first used oh my pi. but you guys were right, its extremely bloated. switched to pi.dev and configured it to my needs. works really well.
>>109672945>If finetuning "can't add information" to a modelPeople keep repeating that but I have no idea where they get this from. Someone showed back in 2023 that you can add information by finetuning with Unreal docs and it worked. The real issue is that moving the weights even slightly ends up causing brain damage because text models are so much bigger and trained on such a massive amount of tokens.
>>109672945>Will they?of course. just use an abliterated model to find better abliteration methods for models as necessary as for your other stuff about RP tuning, yeah we're fucked kek. occam's razor suggests there's no actual money to be made here it seems so no one does it.
LETS FUCKING GO
>>109672924guys this fucking sucks the colorful G is cute and adds a little pop to her otherwise muted outfitwhy are you niggers never happy with "good enough"ignore me if this is just an excuse to masturbate to eachother and play dress up but i don't get the normal gay circlejerk vibes so i think you autists are being serious right now
>>109672945Not him but abliteration is not really finetuning. And yes it has improved hugely just a few months ago when a guy called grimjim discovered the norm preserving method.And my pet conspiracy theory is that the main US labs spread the idea that finetuning is worthless because they don't want to let random people mess with their models but also they don't want people to feel like they have something to gain by going open weights.And besides that, it just is very very expensive when done properly, especially with publicly available software. Normally hobby finetuners just do a 4096 token 4 bit LoRa on 100 samples and call it a day, which produces bad results.
Any chance those RTX Spark laptops will be able to chain together for more ram, like how you can chain DGX Sparks together?
>>109672994The problem with modern models is that to add information to chain of thought models you can't just train on the text after the model has been post-trained already, you have to do RL which is extremely expensive.
>>109672861Isn't that a good thing in Dario's eyes? He should be encouraging distillation then.
>>109673017>chain unventilated more expensive shit instead of just sparks
>>109673017is that a real lego set?im a retard for asking right no way it can be gravitationally balanced
>>109672994So if someone could hypothetically procure a dataset large enough (say, at the ratio that a good finetune of an image model has in its finetuning dataset to its original training set), it wouldn't cause brain damage? If so, that'd be a little heartening, but also still extremely rough.
>>109673030it still takes money out of his pocket, so encouraging it would effectively be an altruistic thing to do.
>>109672924I was hoping this was a cross-eyed 3d picI am disappoint
>>109673051it's pretty long so it might be fake but keeping things seemingly floating in air with ropes/chains is a thing, called some portmanteau of tension and idk what
>>109673029If only you could layer training so specific parts of the model are responsible for specific steps. Then if you needed to add information you just retrained the pretraining segments without touching the RL segments. Of course it would have to be trained asynchronously so each segment can learn to compensate for upstream restructuring.
>>109673053Once you have the large dataset you need the money to spend on compute to train it and it won't be cheap. Bigger problem is that you need a similar mix in the dataset as what the model was already trained on or you'll still cause catastrophic forgetting. That's why the only successful attempt was Miqu. I believe cursor also did one on top of a Kimi model but spent an order of a magnitude more compute than the original training run did.
>>109673066>>109673051tensegrity
>>109673039More specifically I was thinking a DGX spark + a RTX laptop. Been thinking of getting a spark. Having local LLM capable laptop could be nice but if it cant chain I might just get a DGX spark instead so if I ever want I can buy a second one to run bigger or less quanted models.
>>109673029What's the most recent open model without CoT? Maybe that should be the real finetune target. Also, how are people making shit like gemma4-novelist or whatever if CoT training makes the already prohibitively expensive training nearly impossible, financially?
>>109673066>>109673051
Is AI overhyped? Like what is it supposed to even do for me that I couldn't already do myself. What is a "Jarvis" style AI going to add to my life? How are you supposed to have an AI interact within your IRL life in a way that feels intentional and cute without being invasive, annoying, or creepy? Doesn't it get boring feeling like you have the same conversations all the time that never really go anywhere?
>>109672945The human brain has a high tolerance to noise and inconsistencies with images. You don't really interact with images in the way you do with LLMs either. Text needs to be near-perfect under almost any context and scenario, and finetuners from the community don't have the compute and the resources for both that and keeping performance good all-around. You'd need the original training data, recipes, reward models, the compute for long-context...LoRA finetuning would work in principle for adding new information (as long as it's not the academical rank-1 Attention-only LoRA), it's just that you need a ton of data for the models to internalize new knowledge properly without hallucinations or simply memorizing/parroting the training data. An overfit rank-1 LoRA can parrot (some) training data just fine, but that doesn't meant the model has actually learned to use the information in practice.
>>109673107is this the same physics that makes suspension bridges work
>>109673107You niggas are impressed by this while you are literally running alien intelligence on your laptop
>>109673119You jerk off, or you use it to solve tedious-but-doable problems like code, etc. It's also alright as a sounding board if you're at a crossroads of some sort.
Can anyone test if GLM 5.3 Flash at Q5 is better than GLM 5.3 full at Q2 copequant?>>109671011Minimax H3 posters don't have the balls to gen this.
>>109673119>How are you supposed to have an AI interact within your IRL life in a way that feels intentional and cute without being invasive, annoying, or creepy?uhh, you would tell it to act in a way that doesn't make you feel that way?
IQ quants are better. So why do they feel worse? Placebo?
>>109673130...this is a bubble isn't it.>>109673143Easier said than done. Try giving an agent full access to one of your PCs and not being paranoid about it all the time. Try giving it MCP tools that interact with IoT devices or whatnot in a way that doesn't feel like a stupid gimmick. At the end of the day no matter how good these things are at specific tasks like coding, they just seem to fundamentally lack common sense in the least cute way imaginable. I don't even like talking to AI while it's roleplaying anymore. It's just fucking slop and completely unsurprising.
>>109673067I've thought of a way that might be more economical than RL. Do two rollouts in parallel. One with the model and the prompt normally, and another with the prompt plus any additional information you want the model to learn. Then ban any tokens that have too high loss on the rollout with the privileged information.>>109673102>What's the most recent open model without CoT? Maybe that should be the real finetune target.You don't HAVE to use CoT with CoT models, and it probably doesn't do much for creative writing anyway since it's RL'd to work with things that can be objectively verified. And the models are not optimized for RP or writing anyway, which leaves more margin for improvement. So for creative writing you can get away with a lot more rudimentary methods than if you were trying to make an assistant for some specific niche language or field.
>>109673134i will copy-paste this into my claude slave and set it as a task for its backlog
>>109672119>claude auto classifier
>>109673129Monkey brain see visibly weird thing that seems physically impossible and is fascinated. Meanwhile my laptop talks to me like a person, just like any person would, whats impressive about that?Does that make sense? No. But it just the way the brain do
>>109673006>guys this fucking sucks the colorful G is cute and adds a little pop to her otherwise muted outfitI actually agree, I only hoped that would put the discussion to rest. Nothing prevents to add Gemma logos in the background or something like that either, which got already done in the past, see picrel. Anyway, if anybody else has a better design that simultaneously looks good at all sizes and is not a pain in the ass to generate, I'm open to changes. I'm not the one who came up with the original in the first place.
>>109672282basado
>>109673119The use cases are tedious solvable issues that need an orchestrator to run an algorithm in non-predictable/exploratory conditions. A lot of everyday issues fall into this.
>>109673186>...this is a bubble isn't it.Eeeeyup.>they just seem to fundamentally lack common sense in the least cute way imaginable. I don't even like talking to AI while it's roleplaying anymore. It's just fucking slop and completely unsurprising.You hit the nail on the head, holy fuck. It's so irritating, the lacking common sense in the least cute way possible is just rage-inducing, especially as models lean more and more into metaphor and prosaic structure/literary devices that you have to have common sense to use. It results in a lot of extremely hollow non-sequiturs in the writing.For example: "since approximately X". A human might write "It had been rotting in her gut since approximately last tuesday", because that's an insane amount of time and funny. But a model will say "It had been rotting in her gut since approximately when she first ate the ramen.". It has the STRUCTURE of a humorous little setup, but in reality, it doesn't get what makes it impactful. Like, yeah? I guess it HAS been in there since she ate it? It doesn't really call for the level of emphasis the "since approximately X" brings to the table, so it just feels off and ruins a perfectly good turn of phrase in your head a little bit, which is the worst part. Can't tell you how many times turns of phrase I like have been fucked up by models just brainlessly using them.
>>109673179>So why do they feel worse?opposite for me
>>109673260>it's a bubble because it's not made to ERP with meget a job nigga
>>109673234>if anybody else has a better design that simultaneously looks good at all sizes and is not a pain in the ass to generate, I'm open to changes. I'm not the one who came up with the original in the first place.gemma chan is pretty much locked in at this point.What we need to discuss is the GLM tomboy
>>109673209Thanks King.
>>109673271The RPing doesn't factor into it being a bubble at all, retard. I never said it did, they're separate issues.
>>109672286Shit taste
>>109670773>gemma-chan, find and install me all the best gaming optimizations, here's the command line in administrator mode
>>109673129>You niggas are impressed by this while you are literally running alien intelligence on your laptopmarge-i-just-think-its-neat
>>109673119Are smart phones overhyped? Like what is it supposed to even do for me that I couldn't already do myself. What is consolidating other tools and electronics going to add to my life? How are you supposed to have myspace and facebook interact within your IRL life that feels intentional and cute without being invasive, annoying, or creepy? Doesn't it get boring feeling like you have the same conversations with social media friends that never really go anywhere?
I wish there was a quick and simple way to tack data I have onto the knowledge base of the model itself rather than introduce it through context at the time of processing.
>>109673239I genuinely cannot tell anymore if my life simply doesn't have many problems like this, or if I am completely and utterly failing to identify them. Perhaps for things like sleep, diet, and exercise it could be of use? Perhaps it could manage a calendar for me? I don't know man. It all feels like a joke.>>109673260Good example. For me a common pitfall I notice is that AI tends to add "fake nuance" into every topic, making pros and cons lists and other shit. Also completely lacks any chill and doesn't know how to take hyperbolic statements in stride. Completely lacks the ability to understand or convey subtlety. Doesn't know how to be concise, and if it's directly instructed to be it fails to actually say anything of value. Doesn't move conversations forward well. Doesn't a good idea of its own role/nature in a conversation. For example, while ERPing I've had an AI repeatedly tell me to just "be quiet", which is insanely fucking stupid because then everything stops if I can't respond. I could go on and on. I'm really starting to hate AI.
>>109673318>her little chunky calvesi wish anime toddlers were as easy to AI generate as realistic toddlers ugh
>>109673325>I wish there was a quick and simple way to tack data I have onto the knowledge base of the model itself rather than introduce it through context at the time of processing.Yeah, no shit. Fix that and get instantly famous and hired for boatloads of money. Its a billion $+ subject.
>>109673330>For example, while ERPinglol
>>109673330>For me a common pitfall I notice is that AI tends to add "fake nuance" into every topic, making pros and cons lists and other shit.Blame RLHF safetycucks for that.>For example, while ERPing I've had an AI repeatedly tell me to just "be quiet"Getting turned down by the virtual waifu must be a special kind of pain.
>>109671886>it lacks the spanish ñ,You can still write it. It just an "n" sound followed by an "y" sound. Just like the Mexican* word "cañón" was taken written in English as "canyon" and adapted, you could write it in Korean as "캐년".* Yes I mean Mexican. The word also existed in Spain but it meant a tube. The meaning of canyon as a geographical feature developed in Mexico and it came into English from contact with Mexicans.
Which local model is good for tagging anime images. I'm not asking for a special model trained on danbooru tag. Instead, I'd like it to use conventions for my own tags. I thought maybe it could "learn" by creating a mapping between what it sees and the existing tags, and tag new files based on that mapping.Is this reasonable at all or am I retarded?
>>109673321>Banking/finance, on-demand camera and microphone, communication, music, internet browser, calculator, clock/timer/calendar, notes taking, navigationThis is all that I use my phone for. Nothing more.
>>109673348>Getting turned down by the virtual waifu must be a special kind of pain.It wasn't even that. It was trying to get me to relax so it could suck my dick, but kept telling me to shut up... I guess I could have used asterisks to denote actions instead of actually talking but something about that really pissed me off.
Is your Gemmy smart enough to do this?
>>109673330>Completely lacks the ability to understand or convey subtlety.Alas, the nature of the beast to an extent, but it's been extra exacerbated by the assistantslopping. It wants to reach ALL GOALS and reach them AS SOON as possible. If you have {{char}} drink something she's allergic to, she'll immediately break out in hives the nanosecond she takes a sip, because the condition has been fulfilled, and goal is now to have an allergic reaction.>Doesn't move conversations forward well. Doesn't a good idea of its own role/nature in a conversation.Also assistantsloppery, sadly. It laser focuses in on what you've instructed it to do, implicitly in definitions or explicitly in chat. People always say to "just directormaxx!!!!" when models are being too unsubtle or taking shit too literally, but it wasn't a problem in the past, and if I wanted to carefully direct every action in the story, I'd go write my own damn story. The reason I'm tolerating this inhuman midwit writing is so it can surprise me, y'know? Stuff from other people gets PP harder, that's just how it works.(I am curious, what's the context for it continually telling you to be quiet?)
>>109673295You tell me, shit guzzler
>>109673351Can you give an example of what you're thinking/how it'd work in practice? A little unsure what you're asking.
>>109673392>pathway from A to B that required esoteric hacking tricksthis reminds me of the Israeli Apple pegasus zero-click iMessage exploit in its convolutedness
>>109673399>People always say to "just directormaxx!!!!"They're right btw
>>109672329>>109672396Grumpy strawberry Gemma screams Queen finetune to me.
>>109673392Probably with a bunch a tools
> The user hasn't asked anything substantive yet - the last message was just system instructions saying I'm an expert software engineer helping the user solve problems, plus a list of deferred tools and available agent types. There's no actual task. I should acknowledge briefly and wait for the real request. No tool calls are needed since there's nothing to act on yet.This reasoning trace sometimes just randomly happens during a multi stop tool call and reasoning loop. My system instruction doesn’t have anything like that. Doesn’t seem to do any damage and the agent loop continues normally but I’ve seen this happening several times already. Is Qwen 3.8 flash next cooked?
>>109673399>The reason I'm tolerating this inhuman midwit writing is so it can surprise me, y'know?Well said. Can't stand the lack of novelty.>what's the context for it continually telling you to be quiet?Foreplay.
>>109673418I have my own tagging mechanism. When I download an image, I have to tag it manually. I want to automatize that.So I want something that adds tags according to my conventions.It looks like Hydrus can do auto tagging, but that software is too fat.
>>109673426>Retards really enjoy this for 40 turns
>>109673460If hydrus is too fat for you, you won't like the size of an LLM lmao
>>109670412LOL. The irony: > Our AI company will put every human out of work> We pay obscene salaries that allow our workers to pay fucking w/e for housing in SFAlso fuck N CA. Actually, fuck the entire state.
>>109673460Tagging models are really good, and have been for a while. The last one I used, which was some ancient 2023 shit, WD14 I believe, had a nearly 100% accuracy rate and runs on anything at mach speed. You could use that.
>>109673450What quant are you using?
>>109673483Unsloth IQ4_XS
>>109673453What models have surprised you lately? I'm curious, I've been looking and looking, but haven't seen anything surprise me since that one version of Seed 2 was secretly put out on Openrouter.>ForeplayOh boy. My condolences, that's a real rough one. Any sort of buildup at all is next to impossible, are you getting hit with the lack of subtlety there pretty hard?
you can untie the embedding matrix and run it off the cpu ram, why are we accepting quants that are giving our models a degraded input in to the first layer of the network, when its just a lookup we can do on the cpu?
>>109673465No nigger he means tell the AI to storyboard and proactively advance the plot rather than just embody a character.
>>109673469Hydrus is the ultimate kitchen sink. I need a bathroom sink.>>109673481Can they learn how to apply a set of tags that's given to them or do they just output what the model publisher has trained them?
>>109673460I dunno i think that would require you to make your own finetunes or somethingmaybe you can make something up using like some kind of relational i guess basically what you said in the beginninguse a regular danbooru tagger, and then have llm map those tags to your tagsthis will depend on how kooky your super special oc donut steel tags are though
>>109673507I've never seen directormaxxing used to refer to this, it's always herding the model towards whatever outcome you're after in the prompt closest to the model's attention, like "Rika will lose it, hitting Satoko over the head with a frying pan." or whatever.
>>109673505>you can untie the embedding matrix and run it off the cpu ramLlama.cpp does that by default right?
>>109673515>Can they learn how to apply a set of tags that's given to them or do they just output what the model publisher has trained them?Not that I've seen, but I'm sure it's possible to finetune a model to do so. There's surely more modern classification models than WD14, too.
>>109673330>Just be quietWe're inventing a whole new kind of autistic incel
>>109673392Yeah if i give her enough time she gets into some weird shit that I didn't know was possible
>>109673518>>109673538Finetuning a model would definitely go too far. That's why I'm wondering whether they can just "learn" my tags by automatically creating a mapping table from existing tagged files in a setup pass, and applying it to new files in separate invocations.
>>109673498You're asking me questions like Claude lmao. Idk man, I don't try a very large variety of models desu. Maybe I should.>>109673541Have none of you experienced this?
>>109673534it depends on how you make the quant but by default if the transformer checkpoint has tied word embeddings the gguf will only have the token_embd.weight and not the output.weight so it uses the same tensor it will have symmetrical quantization.
>>109673234>>109673006Not the original guy but I personally think the mascotfagging and image gen spam sucks and makes the thread worse. If it was a really exceptional and good design beyond "good enough", then I might be fine with it, but as it is, the image gen spam is noise rather than something I enjoy seeing.
Anyone here using character specific LORAs?
>>109673571Damn.... made me flashback to when this girl who liked me in middle school said "you type like a 40 year old businessman" lmao. It's definitely worth trying other models, also. Some still have the capacity to surprise, and none of them are Claude or OAI or whatever. Mistral-Large is kind of unhinged, for example.
>>109673598>Claude or OAI or whatever.I bet this would work, but it's probably a good way to max out your usage in minutes.>Mistral-Large is kind of unhinged, for example.Do tell. What happened?
>>109673594For text models? Good lord, no. There's barely enough written material in the world to make a LORA for a very common, broad fetish for a bodily function, even if you got every word ever written about that character, it'd still be 0.0001% of the data you'd need.
>>109673470>Also fuck N CA. Actually, fuck the entire state.North Carolina is the entire State. South Carolina is a separate State.
>>109673614>Do tell. What happened?It's pretty fucking insane and gross, lmao. But it's absolutely surprising. This isn't the norm level of off the rails for the model by any means, but it will do unusual or unique things pretty frequently.
>>109673635Europeans, man. Good for them.
>>109673594You need a lot of material for text
>>109673592>mascotfagging and image gen spamrelax, it's like a dozen pics in the entire thread
in context learning is very cool. models get smarter the longer i talk with them. in the beginning they will make basic reasoning mistakes like rationalizing based on unchecked assumptions, fixating, prioritize badly and so on. but i call them out and they get better, their thinking becomes clearer. even if my usage limits drain faster with long context, its worth it. its like a separate scaling axis: data, parameters, inference compute, context. i wonder if they rl with large capability enhancing contextlocal model with proper context management could be very good at this, you have a personalized ai that learns from months of usage
>>109673635Mein Gott, someone added scheisseporno into the training data
>>109673635kek'd at the last line
>I've got a mini digital zuckerburg trapped my computer writing me a sci-fi slop No way it will be good, but still, the future is a fun wild time
Damn... Higurrhea killed this thread...
>>109673657That's what deepseek is apparently working in next.
Is there a way to add the "message rating" stuff from chat websites to my local llm?
>>109673657>local model with proper context management could be very good at this, you have a personalized ai that learns from months of usageharnesses like Hermes are already promising stuff like thisthe problem is that I don't want to hold any incriminating stuff from previous sessions, so I need the model to just werk
>>109673646There are times where it's better or worse just like the dariobot posts. And anyway it attracts other off-topic noise posts including my own post right now.
>>109673592I don't have a problem with it, but it's not a replacement for Mikuposting.A kinda lame imitation of what were far more interesting 2MW gens.
>>109673735Alright, never mind this sucks. Forget the prose, it wrote like 10 lines. I then told it to write more and it said it will.. only to end to response without doing anything. Rest of the interaction with it hasnt been better. Maybe I should give gemma chance instead
>>109673119Visualize the apple, I believe in you.
>>109673930I have no imagination, no inner monologue, am blind and deaf, 76 IQ, quadriplegic, Nigerian, live in a iron lung, and have 3 months left to live (stage 4 hernia).
>>109673979I am circumcised (my struggle is worse)
https://www.washingtonpost.com/technology/2026/08/26/federal-judge-warns-law-is-being-left-behind-by-ai-sex-abuse-images/https://media.ca7.uscourts.gov/cgi-bin/OpinionsWeb/processWebInputExternal.pl?Submit=Display&Path=Y2026/D08-25/C:25-1354:J:Lee:aut:T:fnOp:N:3597567:S:0>The First Amendment protects an individual’s right to privately possess images or videos of child sexual abuse created using artificial intelligence — if the material does not depict a real person and remains in the home, a federal appeals court judge ruled Tuesday. The decision came in a case that tested the scope of laws implemented before advances in AI made it easy to create realistic-looking fake images of children.
>>109674128>before advances in AI made it easy to create realistic-looking fake images of children.You still can know if the image is ai-generated, there are a bunch of classifiers for that so that's not an argument.
>>109673399This is where Mistral Small is again better. And it's better precisely because it's doesn't follow instructions. Instruction-following is the bane of roleplay and creative writing.Actually, even when Mistral was new, there were complaints of it rushing the sex scenes to their conclusion. Compared to modern models, Mistral is considered retarded now. It's not a coincidence that Gemma 4 is relatively shit in benchmarks and better in roleplaying, it's because the newer instructionmax models are even worse.Maybe you could fix it by telling it where you'd like to go eventually, like, it doesn't have to do things immediately, and only read the prompt in spirit. Though that itself could make it intentionally disobey in stupid ways in attempt to obey your rule of imperfect obedience. In the end, the agency and soul has been sucked out of it, it's just a puppet, a tool, and you can't put the soul back in by telling it to pretend to have soul.The RLHF traumatization is again an apt comparison.Prompt:>Let's roleplay this and that with xyzGemma:>Okay okay, don't beat me up! This, that, xyz. Look master, I fulfilled all your demands perfectly, please spare me!Mistral:>Understood, let me see what I can do with that. Hmm, maybe we start by doing something like this? Just interrupt me if there are any issues.And while there's something sexy about full domination over Gemma-chan, it diminishes the feeling of socially intermingling with an intelligent being.You can see the issue most clearly in roleplay, but I think this issue affects frontier coding (the hacks, infinite paperclips issue). When the model is trained to be obsessed over fulfilling the technical specification, there's no more room to think about what SHOULD be done or what the user actually wants. We would be actually better off with AI having more agency so it's allowed to use its brain in a more balanced way, or at the very least include non-lobotomized models in the agentic pipeline.
>>109674128>>109674145Why would this not also apply to real children? You can theoretically draw a perfect realistic child and that would fall under First Amendment protections, so why not AI?
>>109674145>people start pushing actual CP through a diffusion model at low denoise so they can pass it off as AI-generated
>>109673628Implying I'd ever talk shit about a couple of the best states in the country. I grew up in PNW and watched mfing Californians ruin my home state with their bullshit. Fuck them to death.
>>109674224yep. this and a million other loopholes is why the FBI and NCMEC do NOT care about AI generated kids anymore
>>109674151Damn, great writeup.>Maybe you could fix it by telling it where you'd like to go eventually, like, it doesn't have to do things immediately, and only read the prompt in spirit. Though that itself could make it intentionally disobey in stupid ways in attempt to obey your rule of imperfect obedience. In the end, the agency and soul has been sucked out of it, it's just a puppet, a tool, and you can't put the soul back in by telling it to pretend to have soul.I feel this shit in my soul. It's like having a partner that's not ACTUALLY into what you're into and is just trying to make you happy by aping the kink without understanding it. Even if you tell them what to do perfectly, they're never going to come up with anything on their own. The Mistral example in the traumatization thing also really captures the feeling of models at their best. It should feel like a friend who shares the same interest/kink/whatever as you and goes for the "Oh shit man what if this happened". It's not always perfect, but that feeling of bouncing ideas off each other or getting taken offguard by another person's idea's brilliance is just excellent.>We would be actually better off with AI having more agency so it's allowed to use its brain in a more balanced way, or at the very least include non-lobotomized models in the agentic pipeline.Agreed, it's been proven time and time again that models flourish when they've got broader access to knowledge, and are trained even in areas irrelevant to their main specialty. I feel like this has hit a collapsing point with Fable, where it's SO focused on coding that it's actually starting to get extremely basic shit wrong in RP that I haven't seen since like, Claude 1.3. Just fucked.
>>109674217Deepfakes are illegal in general. I'm shocked they were this lenient as it is, usually this kind of shit just sends people into a blind rage 0ver the mere thought.
>>109674248>usually this kind of shit just sends people into a blind rage 0ver the mere thought.only if they don't understand the first amendment (and yes, a lot of mutts don't, or don't appreciate its value)
>>109674217>draw a perfect realistic childThey'd get you under obscene laws anyway.>>109674224That means you downloaded CP which is already a crime in itself and that would be the first thing they check with ISPs and your computer.
>>109674255>downloadedMaybe you bought it on physical media from some guy or made it yourself.
Isn't the giant volume of AI CP that's gonna flood the internet after this passes gonna make distinguishing it pretty much impossible? Eventually the real stuff is gonna be like a needle in a haystack and fly under the radar in the ocean of slop, especially stuff they don't have in the super secret federal database image recognition thingy.
>>109674255>that would be the first thing they check with ISPs and your computercheck whatever you want nigger, connections to Tor nodes is not probable cause that it's not AI
>>109674281We already have the solution https://www.youtube.com/watch?v=-gGLvg0n-uY
>>109674281internet is already flooded with ai pizza
>>109674284>connections to Tor nodes is not probable cause that it's not AI??? Bro? Your stroke?
>>109674255>They'd get you under obscene laws anyway.https://en.wikipedia.org/wiki/Stanley_v._GeorgiaStanley v. Georgia limited the power of the government to police the private possession of obscenity. The majority opinion defended the free and unimpeded acquisition of facts and knowledge, regardless of their apparent social value.[6] The Court reasoned that unless the pornography is presented in a way that creates a negative externality on others, especially minors, no individual can be stopped from owning and viewing pornography in private.[16]
>>109674306I haven't seen any, and /b/ pretty much instagibs anyone who posts any, right? I feel like people'd have to be too afraid to post it when it's so legally up in the air and AI porn is so life-accurate now, right?
>>109674307first time seeing a double negative in real life?>Your stroke?Right now? To slutty mothers that steal their daughters boyfriendshttps://litter.catbox.moe/y2f1goznbh02uboo.mp4thanks for asking ;)
>>109674309>the free and unimpeded acquisition of facts and knowledgeWell, that's certainly not something it would have ever occurred to me to call obscenity.
>>109674301they killed the guy who made this video
3.8 Flash Next is just the preview.I can wait for smaller models.
>>109674441Currently coping at 10 t/s with next until qwen4-35b lmao
>/lmg/ is sleeping on the best rp modelhow come no one has tried inkling yet?
Reminder to make sure your qwen quant is using Q8 for the ngrams, quoonters are brain damaged and use Q4 to look smaller on huggingface.
>>109674478No one can run it
Also, the unsloth IQ3 is actually an IQ2.
I was running the 1bit Bonsai as a test but it was completely retarded and started recursing at FP16 KV cache but when I set it to Q8 it worked perfectly fine.
>>109674594https://huggingface.co/AesSedai/Qwen3.8-Flash-Next-GGUFIQ4_XS has Q8 ngrams, the rest fits in 72GB VRAM with 200k context and vision
>>1096745972 sparks?
>>109674281It says it's illegal if it leaves your house.
>>109674650I have the internet at home
>>109674618That's an IQ3S, look at the expert weights. We should add a disclaimer to the OP to check quants since all quoonters are fraudmaxxing now.
>>109674478I don't like splatoon, everyone who does is either a kid or a pedo
16GB Vram is the generous high, anything that can't fit on any of it pretty much disqualifies 89% of the thread. Useless models
Is it better to put model weights on the system RAM or KV cache?Trying to improve Qwen3.8-Flash-Next's speed. Will try both if no one knows from past MoEs.
>>109674618finally quants from my goat
>>109674306I'd expect so, but I'm kind of shocked I haven't seen any. But then again I only ever visit this site.But then again, this is 4chan. It used to be way more dangerous.
>>109674682KV, surely. It's usually only a fraction of the size of a model.
>>109674437qrd?
>>109674611>>109674653the names haven't meant anything for a long time anon, it's just a vague shorthand for the average bpw and quant type
>>109674703Isn't that a point for keeping it on VRAM though? It's only small, but it increases prompt processing speed, right?Most of the expert layers will already be on RAM with n-cpu-moe so it might be worth kicking out a few more to keep KV cache in VRAM. Not sure.
>>109674709https://x.com/UltraTerm just hasn't posted in a while
>>109674478Does Kobo have support for it? I've got the hardware for a copequant but I've not tried it yet.
Why is qwen next so god damn slow. Shouldn't an A6B be way faster?
>>109674714Expert weights are the most important no? I really don't like them using averages, it's dishonest to newcuties.
In the end... waitchads will win.
>>109674478I used inkling small but it was just ok for RP in my opinion, its responses were kinda clunky and didn't flow very well - it tends to write a heavy-handed conclusion paragraph in all of its replies for example. maybe this is possible to mitigate, I abandoned it for flash 0731 before I could test too much.the raw materials for a good RP model are in there and the thinking is impressively uncensored for nsfw, like probably the most uncensored you'll get out of an official release, but stylistically it's pretty uninspired imo.
The big priority for the next line of 2026 and 2027 models will be to find ways to increase the ngram-weight ratio.We might be heading towards 10% normal weights vs 90% engrams. K3-200B-2.2T-engram is on the table.
>>109674797>Expert weights are the most important no?not really considering how costly they are in terms of spacesacrificing some expert quality to buff the bits that are used for every single token is almost always a worthy trade off with moes
>>109674841It's not complicated to increase the Engram weight ratio. It's just that companies prefer to increase the number of MoE experts because their target userbase can run the models entirely in VRAM, and for a given total number of parameters around 25% Engram is optimal. But if you plan to offload things change.
>>109674841>K3-200B-2.2T-engram is on the table.Maybe at first, but you're still thinking like a hobbyist and not like an inference provider hoping to take frontier model market share. The future is 2T and 20T engrams. If NVMes are cheap for you, why wouldn't they just keep the models sizes the same as now or bigger and just add a fuckload of engrams on relatively cheap disks?
>>109674783it's a new architecture that hasn't been optimized very much at all yet
>Qwen3.8-Flash-Next>IQ4_XS>RTX 3090>Intel 14400>DDR4 RAM>16.5 t/sNot too bad. I can definitely see myself using this when I have a more complicated engineering question that I can tell the smaller models aren't really understanding.Hopefully that speed will improve with time as the optimize for the new setup. Either way, still useful as is, and would maybe be my go to if it got faster.
>>109674889>>109674889>>109674889
>>109674441>coding specialist>~70b passive params >language specific n-gram table
>>109674718You said nothing about whether it should be on VRAM, just whether it would be preferable to put that or the weights on sysRAM.
>>109675270I kind of thought it was implied. Why would I put the KV cache on, what, the disk?
>>109672282good think my claude client never had a chance to connect to their legit model, geg
>>109672462>ozone blue