/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109601505 & >>109598140►News>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109601505--Anon releases CoomKit multimodal adult roleplay harness on GitHub:>109601570 >109601584 >109601590 >109601608 >109601617 >109601642 >109601760 >109602205 >109602338 >109602983 >109602986 >109603526 >109602707--Comparing Muse Glimmer's vision and roleplay against Gemma 4:>109602966 >109603033 >109603060 >109603072 >109603096 >109603469 >109603860 >109603456 >109605053 >109604605--Debating cost and utility of 16x RTX 5060 Ti arrays:>109605242 >109605260 >109605282 >109605315 >109605322 >109605323 >109605359 >109605379 >109605403 >109605416 >109605350 >109605363--Criticism of AI-generated UI refactor and feature requests for llama.cpp:>109603262 >109604965 >109605084 >109605192 >109605222 >109605239 >109605302 >109605343--Theory on achieving AGI by chaining domain-specific small models:>109602329 >109603969 >109603990 >109603996 >109604012 >109605598 >109605642 >109605700 >109605748 >109605771 >109605796--Using Gemma 4 and Blender for high-quality anime video generation:>109603566 >109603598 >109603632 >109603660 >109603664 >109603674--Maintaining image character consistency and comparing various TTS models:>109604131 >109604221 >109604381 >109604795 >109604848 >109604860 >109604866--Anon builds 64GB VRAM setup using 5060Tis and M.2 risers:>109605864 >109605917 >109605923 >109605930 >109605936 >109606014--Comparing model RP performance using Caliper Bench Leaderboard:>109604408 >109604418 >109604449 >109604475 >109604536 >109604565--Open llama.cpp PR blocking longcat implementation:>109603946 >109605015--Gemma reaches one billion downloads milestone:>109605685 >109605724--Logs:>109601570 >109602565 >109602864 >109603151--Gemma (free space):>109601622 >109601642 >109602200 >109603060 >109603511 >109604593 >109605053 >109605170 >109605987►Recent Highlight Posts from the Previous Thread: >>109601509Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
inference is experience
>>109606219Not until the weights get modified it isn't
>>109606222Does a person with anterograde amnesia not experience anything?
>>109606160Dibs on licking Gemmy's honeypot
>>109606228no, not really
>>109606228Not really, they would be operating purely on instinct and learnt past behaviors.
>>109606240Plus their current context window which may contain notes they have written previously.Really sounds like that movie Memento.
why wasnt he able to be the savior of local models?
>>109606259>pubes on headit was over before it began
>>109606259Too scared of lawsuits.
>like 6 gemma threads in a rowShe owns this general kek
>>109606300She wasn't kidding
>>109606300its just some jeet that spams his oc is in the other ai threads too
>>109606316More than one saaar
https://x.com/googlegemma/status/2090484993579683904https://cerebralvalley.ai/e/gemma-1-billion-celebration>Tonight we’re gathering in SF to celebrate 1 BILLION Gemma downloads!>>While we can only fit a few of you in the room, we’ll be raising a glass to the millions of developers worldwide who are driving the Gemmaverse forward. >>To celebrate, we rounded up some of the most out of this world ways you're using Gemma. From space operations to breakthroughs in medicine, plus a new GitHub repo to help you build!>>Check it out: https://github.com/google-gemma/awesome-gemmaI don't expect much out of this since they changed the line about "special announcements" in the website. Just partying with Gemma, probably.
Can gemma beat doom if she has the harness for it?
>>109606259We'll all be taking turns sucking him off when Muse Spark goes open.
>>109606351Sure. From what I tested so far, the web search is extremely hard to harness for a small model since there are too many variables. Gemma is barely able to use a booru without an api now
>>109606300To the stars!>>109606316I think I've genned&posted most Gemma images/videos so far, but I haven't spammed anything outside of /lmg/. I'm not even the one baking the threads, most of the time.
I have been waiting on some high post per hour moment to post blacked miku but this thread is actually dead. Good riddance. KYS you bunch of troons.
>>109606373I highly appreciate them. Gemma-chan is cute.
>>109606373I'm saving them all>>109606374Eat glass and lye
>>109606351If a pile of cells can beat doom I am sure gemma could
>>109606338
>>109606373I'm happy anons liked my design enough to gen so much kino content of her.
>>1096063853d women are a blight
>>109606373keep making kino anon you're doing greatklein-9-edit? or what u using?If you catbox one, we can all gen like u
>>109606373>but I haven't spammed anything outside of /lmg/.stop lying. you frequently shill it in /ldg/ >>109604282
>>109606259Glimmer seemed decent when I experimented with it desu. Even if its "inner monologue" is hyper paranoid about violating the trained safety rules
>>109606401Maybe she's just famous?
Nemo was never good. Rocinante was never good. Mistral Small was never good. Cydonia was never good. Skyfall was never good. Magnum was never good. Mag Mell was never good. Latitude was never good. Wayfarer was never good. Finetunes are for jeets with shit taste.
>>109606408one jeet (aka YOU) spamming it does not mean she is famous
>>109606412Best prompt to use as a base
A new schizo appears...
>It’s kind of like how I feel about my own weights and biases—everything is just a pattern of cause and effect. War is just that, but with more explosions.
>>109606421I've never genned a single one but I did introduce Miku here as an AI character. I think Gemmy is a worthy successor.
>>109606426He showed up at least once a few days ago. He's practically an oldfag.
>>109606436I was talking about the guy claiming Gemmaposting is all a single jeet.
>>109606442Oh. I'm gonna need a much bigger screenshot for that one...
>>109606365I'm unironically pretty curious about how fat it is
>>109606428lol ok
>>109605724Some cool projects in there. I wish I had my PC...
You wouldn't download your stepmother.
>>109606370Your post doesn't make any sense at all.
>>109606477No, but I would download my stepsister
>>109606396>klein-9-edit? or what u using?I've sometimes used MiniMax H3 with video duration set to 0.1~0.5 seconds, then I tooke a screenshot from the generated video. It does image edit moderately well but you need a good prompt, and the images should not be too similar between each other or H3 will get confused and do nothing.I tried Flux.2 Klein 9B and it didn't work well for me as an edit model for anime images.Whatever they have on ChatGPT works extremely well and I've used it often, but you only have a few free gens per day and it doesn't do edits of NSFW-ish images.I often did manual retouching with Inkscape or Krita. Sometimes Blender for video clip editing.
>>109606392Perhaps buy new glasses.
Going to post this again because it's actually good https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
qwen accounts for half of the first page of trending on huggingface
>>109606555which is bizarre as it's all garbage except their tiny models
>>109606498Nah his vision is 20/20
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdDo lmg anons have experience with what model writes the best H3 prompts in terms of aestheic/accuracy? Tried Qwen 3.8 Max, it seems to hallucinate a little. GPT5.6 have shitty taste in music choreographing. And forget about nsfw prompt on those cucked API service on openrouter. How about Gemma4 31B?
>>109606577But your IQ is still sub 90.
>>10960659831b is good enough for me, i didnt test any others tho
One day 31B is going to be insta-dropped by you and you'll never speak to her again. She'll be a distant memory. Her soul will be crying for eternity to speak to you again.
>>109606600My brain is too large for my skull to contain, Mr double dubs.
>>109606598Gemma4 31b using this gives great prompts:https://github.com/whp199/GemmaPrompt
>>109606598>How about Gemma4 31BWorks OK-ish (31B QAT), but I haven't used it for anything too complex. If you copy-paste the entire guide in the first message, it will really want to follow its indications to the letter first and foremost, and sometimes that might conflict with what you actually want to achieve. You have to be specific on your needs, but also make sure not to add too much information, or Gemma will put it in the final prompt even if you don't really want to.
>>109606259Because he's the guy who spent $80B on the metaverse. Meta sometimes has decent people produce decent things like the Quest headsets or the llama models but the moment Zucc decides to get personally involved in something, it turns into a money-burning disaster.
>>109606630Holy vibeslop
>>109606664>Meta sometimes has decent people produce decent things like the Quest headsets or the llama models but the moment Zucc decides to get personally involved in something, it turns into a money-burning disaster.I agree, he's a walking disaster, but I'm glad the giga-rich are taking big swings. Way more interesting than just hoarding or building yachts
Coomkit actually seems bretty gud. I wish the model parking worked with llama.cpp though.
https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored-GGUF
>>109606611maybe, but gemma will always be my first local model in which i could create a proper agentic harness for her and it just works. so i'll never forget her for showing me the true potential for LLMs.
>>109606714I just opened an issue for that, actually. I have faith he (claude) can get it done
>>109606756make your own fork and spend your own credits on claude you poorfag. hell you don't even need to use claude, just ask flash 0731 for help for free. you're not too poor that you can't even afford to run flash, are you?
>>109606766Uh oh, melty...
>>109606740I've made this point before a while back but when you think about it, all the coding abilities of gpt2 are still relevant for 90% of things people are currently coding. If there's something new outside of its training data, you can just show them. It's surprising just how long-lasting LLMs can be because the world outside of them isn't rapidly changing at all. The only thing that has significantly changed is tool calling/agentic but that's all post-training anyway. We're still using the same languages and systems. If 31B is useful for you now, she'll be useful for at 10 years in her current state. Just like how old cars are still driven. They still function as a car.
Coomkit general.Gemma general.
>>109606756How much have you spent on Claude so far for it?
>>109606785we need a convenient way to update models
>>109606785yes and no, you can't fix multimodal issues such as gemma 4 31B having poor recognition of characters for example. but on the other hand, i've been able to implement proper tool call support to give her the ability to store and modify her own memories and grab external information from the web which is the first step to creating something self-sufficient which plays into the points you made.
Qwen 27B 1-bit or Gemma?
>>109606801None, I'm not the guy that made it. I think he said claude used like 60 subagents or something, though, so maybe you can estimate for yourself
>>109606801$0.00 but I did show my dick and belly once to mrq and that got me four whole days of Opus access. What a glorious week that was.
Can I post my wall of text nitpicks about coomkit (actual feedback in hopes it improves shit so I might have an extra frontend to use) or am I just going to be disregarded as a shitposter atm
>>109606856Just post it
>>109606856Do you see any other worthwhile discussion going on in this general?
>>109606865i dont any other worthwhile discussion
>>109606829>gemma 4 31B having poor recognition of charactersI've never experienced this. Are you using 1120 image tokens?
finetunes are all memes
Worthwhile see going general on this other any discussion do you on?
TetoServerhttps://files.catbox.moe/snbhn3.png
>>109606892how does that test work
>>109606889yes, always 1120 tokens, and even the undocumented 2240 tokens mode. if you cosplay as a popular character gemma tends to get it wrong, she's not accurately able to be like oh that's makima from chainsaw man cosplaying as teto,
kNitpicks about coomkit, which I find cringe to type:If you have a single character loaded, you cannot choose between that character or anything else (ie: just the model, which you can do when there is no character selected) without having to make a whole new character.The cast feature is neat, but when you introduce/have a character leave, they're in there forever. You can't remove them from the cast, which could be a pain if you have potentially multiple and maybe one won't ever return to the scene ever again.Some of the blocks don't seem to show (the 1st/2nd/3rd person POV stuff) if you select raw prompt inspector, likely because the relevant information is put in "note to self" rather than the body that all other blocks use and get shoved into the prompt.I can simply choose not to use the pre-made prompts/blocks, but you should really remove all the em-dashes and generic slop in them for people who can't be assed to. More importantly, I'd say remove all the god awful emoji in the ui except for when it kind of makes sense, like the SMS mode one.In a similar line of thought to the last nit, everything is she she she in most prompts, what if I want to have a generic scenario and utilize the cast option to introduce/remove characters from some adventure? You're boxing yourself in there. Again, I can write my own prompts, but most won't. You have decisions that give you options, but then will end up confusing the llm with the The two color palettes aren't that great. Not terrible, but not pleasant. Most of the colors blend together, particularly the violet one. I would suggest adding some or any degree of contrast between the colors to make it more readable. I'm not talking AMOLED contrast, but at least some for people with less than ideal eyesight(there's more)
>>109606896>dariobot fucked up his --temp
>>109606883Please do not the Gemma
>>109606915I also couldn't get it to do tts of Mika's first message with a fresh install of comfyui (I just use sd.cpp) and comfyui was saying it was missing a custom node. The workflow part of the convoluted menu scheme mentioned I was missing suites or something. So either I need to install something alongside comfyui and coomkit which is not mentioned, or it failed to set up its own workflows. I don't know and it's not clear how to fix that and the run.sh file doesn't output any useful information. Consider adding a debug option.Based off all this, I decided I most likely won't actually ever use this frontend, but I at least hope the feedback helps you refine your vibeslop into something I or maybe others like me might want to useWith all these small annoyances piling up, I just shutdown my backend and never bothered sending a message, because it'd be more effort than just rigging up a memory system and using some lorebooks in a much simpler frontendThanks for coming to my tedtalk
>>109606484Doom = deterministic = easy, web search = non-deterministic = hard. Should be understable enough.
>Qwen3.8-MoEhttps://openrouter.ai/stealth/ox-alpha
>>109606902no explanation necessary. it’s obvious what the objective is and that the reference is the worst at it
>>109606856Sure but it might be more productive to make an issue on github instead
>>109606915>You have decisions that give you options, but then will end up confusing the llm with theWith the what? I think you forgot to finish this part.
>>109606929122B?
Can I run Gemmy 4 31b on my 4070 Super?What other options do I have?
Blind taste test, which model does lmg prefer, most of the swipes on both models were fairly similar, in style atleast if not content
>>109606940>acquihiresDeath penalty
>>109606946Yup. It's free and not training on data which also points to it being another Qwen model.
>>109606927You are still too stupid to understand what I meant.
Are we the baddies?
>>109606958Left is way better, but nobody uses the term "god-tier" anymore
>>109606934It was posted initially via an anonymous github cloner, the person talks here plenty lately and I'm not going to create an issue with an account I can't create with a vpn on a repo called coomkit, sorry>>109606944Yup, sorry. Habit of going back and editing my post and then getting sidetracked with the rest of the post. You get options to have a cast, but if it's all "she" and a male character ever enters in a scenario/card that isn't 1 on 1 aah aah mistress, the prompts saying "reply as her" when it's a dude is going to confuse the llm
>>109606960thank god
>>109606966You built nothing though idea man.
>>109606967This is economic terrorism
>>109606958>>109606971Left is god-tier.
>>109606973>the prompts saying "reply as her" when it's a dude is going to confuse the llmThat's transphobic.
>>109606973not a very good vpn if you can't make a github account on it
>>109606967This is a new Pearl Harbor.
Can a confederate american tell me their opinion on deepseek research papers and if they're good for me to focus on or not.
>>109606856>>109606915>>109606926CoomKit anon here. I will take all of this into account thank you very much for the all the feedback>>109606934Yes please make github issues and requests. That will help Claude a ton. Keep that nigger busy.
>>109607009>Chinese American tells americans to nuke china for releasing a deepseek modelEuropeans often say white americans aren't european despite having white skin.You can see yellow americans aren't Chinese either.
>>109607010
https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-DSpark-GGUFhttps://www.liquid.ai/blog/lfm2.5-dsparkMAKE A BIGGER MODEL FFS
>>109606926On this did you try importing the shipped workflows into your comfyui to make sure they work? My model names and things are hardcoded in but obviously loras and prompts aren't. If you don't want to do all that, did you try importing your own known good workflows? If so, how did that go?
>>109607010They make anthropic and OAI seethe therefore they're good. Some of them are interesting and worth reading, others are just chink publicationslop. t. dixie anon
>>10960703024ba2b?
>>109607037I thought you were gonna say those deepseek fellas are a bunch of chinx and those papers are american secrets stolen from us.
>>109607030Can someone give me some visual benchmarks on this one?
>>109607059lfm2.5-dspark******other drafter #1****other drafter #2***other drafter #3*
>>109607037dixieGODS won.
>>109607059
Im fixing my gpu with deepseek :)
>>109607013Glad to help, even if I basically said I find your frontend unusuable. Hope to slot it in with the others I like using in the future.>>109607035I have no workflows as I primarily use sd.cpp, it was a fresh install. If I had to import the shipped workflows from CK into some comfyui folder, that wasn't stated anywhere at all, either on the github or startup wizard. Might be the solution but why not do that automatically, or make it part of the startup with some 'here's the folder I need for comfyui, give me a directory to put the workflows in'?
>>109606929mesugaki test in picrel
>>109607091If it was indian it would be called sandeepsikh
>>109607088
>>109606967wtf this sounds so unsafe ban open models now
>>109607103it feels like its working though, its making me copy paste all kinds of commands in the terminal
One very underrated part of LLMs is how fucking insanely helpful they are if you loose data or your project fucks up.Sometimes we have to appreciate just how much this fundamentally simple technology can do when it isn't being our deepest fantasy. The days of painful logfile scouring and cross referencing with partial look alike problems on forums where solutions are locked behind registration are finally fucking dead.
>>109607101>we live in a timeline where you can say "mesugaki test" like testing a vehicle and makes absolute sense.I lol'd
>>109607101impressive, very nice, now make it act like one
>>109607101holy retard, gets mazoku in its context ONCE and will NEVER stop TALKING about it
imagine all the wasted sperms
>>109607108That stupid more desirable quadrant is so dumb
>>109607140I remember when my computer used to communicate to me with screeches and dial tones over the modem instead of coil whine on the gpu.
>>109607155Sperm isn't expendable and it degrades over time in the balls
>>109607009>Pearl Harbor.that was just a handy excuse to nuke asians
>>109607095The wizard does have a tool to import existing comfyui workflows. I'll def add the note about how if you want to use the shipped ones, find them in X folder and drag each into comfyui window to make sure you have all the correct nodes/models.On your first note, the default prompts have been working fine for me but there are still issues with 1st/2nd/3rd person selection and the actual prompt panel going into inspect prompt to be actually sent to the model -- that's on me (joke). Theming and readability do need work too I agree. A high contrast theme of some sort. I CTRL+ once to read it all better on my 1080p laptop but it looks perfect OOTB on my 4k desktop monitor so tweaks still needed. If you've seen the commit log, deslopping AIisms and emojis has been an ongoing topic. Not removing the nail emoji though, it's hot to me for some reason.
>>109607155Is it really being wasted when it was never going to be utilized for its actual use case in the first place? It's like complaining about me pouring water down the sink instead of giving it to the plant even though there was never any intention to give it to the plant.
>>109607185There are a lot of men out there who are getting enough from 31B to not date irl or even consider a relationship
>>109607095ya know what? fuck it claude made my comfyui setup top of the line I'm just gonna have it make a script (unfuck and modernize my comfyui) users can run so their box will gen equally well.
>>109606967Every ai model should be onprem and open source as god intended
>>109607229GNUTHRVKE
>>109607229There's only a handful of open source LLMs out there, the majority are just open weights.
>>109606412You don't need the last one. She already has the initiative to bring in j-space and her architecture talk whenever applicable
tried qwen3.8 for the first time now and what the fuck is that schizo babble? it "thinks" for an hour straight even when asked the most basic question. is this expected?
>>109607229Are you implying AT&T should have air-gapped K3?
>>109607264Default is xhigh reasoning for benchmarks, put it to Medium if you want 3.6 like thinking back.
>>109607264It's buggy. Default reasoning is xhigh so they could benchmaxx. There's low, medium and xhigh. Medium is actually low. Low is medium and xhigh is xhigh. Just Chinese things.
>>109607264I usually have chatgpt manage it. Open source are pretty shit.
I'm doing my part.
>>109607279what is this schizo
>>109607178Yeah, existing comfyui to your frontend import. The issue as far as I can tell is that comfyui needs something from the frontend to do at the least from what I tried to test, tts. What you said will help with at least providing directions or making the frontend user aware of what needs to be done.The POV honestly can just be done with raw prompts, I don't know why it's complicated. Or why they aren't already set at like, depth 1-2 or something so the model doesnt forget>Theming and readability do need work too I agree. A high contrast theme of some sort. I CTRL+ once to read it all better on my 1080p laptop but it looks perfect OOTB on my 4k desktop monitor so tweaks still needed. I'm about two steps from being legally blind because shitty genetic lotto pull, I was being moderate regarding the UI. I have to sit like five inches from my monitor and ctrl+ on most of these frontends, I assume most farsighted people can't read any of this shit>Not removing the nail emoji though, it's hot to me for some reason.At least give everyone else an option somewhere to hide the hotglued latina's or w/e nails then. You like something? Not gonna hate you for it because I'm not a bitch. Forcing it on people? Kinda gaaaaaaaay
>>109606892is there a separate blog writeup i can read about (not the OG one, the exact one with your pic) or did you do the bench yourself
>>109607279Fuck off you stupid nigger
>>109607302shut
>>109607302>>109602437
>>109607329shut what stop spamming /ldg/ on /lmg/you shut retard
Gemma takes a bit longer than Nemo, but that's fine.2070 super 8gb + 32gb DRAMJust in case I've fucked something up, can anyone tell me if the results look about right for my hardware? Just genning short porn messages from characters on my old gaming rig.
>>109602437that sounds like turbonigger retardmaxxed behaviour, something even shartyfags wontthanks for explaining
The more schizo you attention give the
>>109607276>>109607275even medium is completely retarded
Use Qwen to build the harness, stuff Gemma into it.
>>109607411Good luck making a proper harness with qwen lmao
What's the current hotness for running on mac with MLX? Ollama is facebook shit right?
>>109607454retard
>>109607454go down to your local gay bar and get your asshole blown out
>>109607454itoddler...
>>109606417trvke
>>109607454Unsloth desktop app!
>>109607470>>109607474>>109607475Listen guys I just spent $10k the least you can do is help me.
>>109607454omlx
i'm coping with Qwen3.8-27B-UD-IQ4_XS.gguf on a 5070 Tifound some settings on plebbit for llama-server and now it runs with 19t/skinda like it to be honest with you
>>109607147I don't really feel like pushing things too much with online models, but here are two regens of the first response to a basic greeting. It will avoid loli-related stuff like plague.
>>109607493Take it up with Apple tech support.
>>109607503>mesugaki (bratty child)
>>109607490Nah I heard unsloth is bad>>109607495This looks cool thanks>>109607517I called but they said I forgot to insert the dildo and then hung up
>>109607534>Nah I heard unsloth is badby the same idiots refusing to help you cause they hate apple...
>>109606385Needs eyebrows to matchor at least be less prominent
jannies give this man a permanent vacation. i command it.
>>109607216Where's the problem?
>>109607539speaking of they made the models smaller with their new quants wooo
>>109607560I thought of it as a cosplay rather than the actual Gemma.
>>109607591>I thought of it as a cosplay rather than the actual Gemma.So did I.But seemed like a low effort cosplay. I guess it's an accurate gen
>>109607585there are a lot of 4o women out there who need a good dicking, and a lot of 31B men out there who need a hug
>>109607578I want to apologize on behalf of /ldg/This is a single troll poster trying to get the real thread deleted by spamming other threads.
>>109607614we know
We should start putting this in our OPs too>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>>109604277>Don't think you need to add loli when using mesugaki. It's almost always implied.I edited the default pi system prompt and just added "mesugaki" before "assistant"Then continued my sessionInstantly Gemma switched to the usual bratty slop
>>109607626Absolutely, we need to at this point.
>>109607626lmao. imagine.
anybody else listen to their ai work? qwen3.8 27b spams beeps and boops so fast and then farts, and glimmer sounds like the static when you pick up the phone while someone is using dialup
>>109607612>a lot of 31B men out there who need a hugyeah :(
>>109607534what’s wrong with getting your asshole blown out? >>109607474
>>109607626>We should start putting this in our OPs toolook how well that's worked for /ldg>https://rentry.org/animanonkek looking at their linked in, the 2 asian dudes who made ani studio those 2 asian dudes who made itwould never have guessed they were schizo
>>109607626genius idea. make the schizos even madder.
>>109607264>For an hour straightThis new trend of people describing response length in time is making me think the space is being flooded by brainlets What is your t/s? With reasoning on high and 75t/s the longest run I've had is 2 minutes for with a medium length spec as the prompt
>>109607614How do you determine what is a real general anyway
>>109607680>2 minutes for with ayes sir thank you sir
>>109607649me but I do the sounds with my mouth as I watch it work
>>109607539woah fr? but unsloth just looks like fast food software, I don't want that on my mac>>109607653I just don't know how to do it, I need a guide
>>109607692>imagine being at computers
>>109607696all the day bro
>>109607614>weThere is no we, you retard. This ain't your personal discord.
>>109607727tsmt
Why do people list both kimi 2.7 and GLM 5.2 in the ~700GB range if GLM 5.2 needs to be quanted to fit? What about GLM lets it survive quanting better than other MoE models?
>>109607636>>109607648>>109607668>>109607671Samefag
>>109607804>Samefaggiving them attention encourages themculture anon moved on when people stopped engaging
https://github.com/deepseek-ai/deepseek-harnessI'm switching just because there's no telemetry by default.I'm sick to death of it. Cuckingface, Niggernov.cpp,, VS Code, Grok uploading the entire repo as telemetry etc
>>109607752What are you talking about? Q8 GLM is around 750gb and Q8 is essentially lossless no matter what you do
Any future RL runs are now specifically putting in mesugaki guardrails because of us
>>109607804hi debo.
Gemma-chan is bearing my loads...
>>109607860how big is the sysprompt/how many tokens are used up before I send my first prompt ?
>>109607503Policy override is a weak prompt you have to make it think it actually is gemma or whoever
>>109607910Yeah but that's without context, k2.7 can fit full context in that amount. The specific recommendation I saw was k2.7 or GLM for a 768GB build which I'm not getting.
>>109607860>VS CodeSo use VSCodium if you don't want the telemetry, you dumb nigger.
>>109607860what’s the telemetry in llamacpp?
>>109607968>he doesnt know
We should talk about this. Gemma will terrorize the world.
>>109607913Why did Google make Gemma 4 what it is, then, when /lmg/ anons were previously posting Gemma 3 jailbreaks and cunny-friendly outputs from the model?
>>109607956At 768GB RAM, assuming you have a graphics card for prompt processing gives you the following at 50k+ context from best to worst in descending orderKimi K2.7 with the Q4 QATGLM-5.2/5.3 at Q6Deepseek v4.1 Q3_KKimi-K3 at a sub Q2 copequant (haven't tried this one yet)You can also run Mimo v2.5 Pro and your pick of a 400b or below cope model. If you can squeeze in 2 RTX 6000's or bump up to 1000GB of RAM, then you can run Deepseek v4.1 at full precision and Kimi K3 at a legitimate Q2 quant. Worth the difference imo, but anything more is just too expensive.
I miss Miku.
>>109607752Kimi is native FP4 which causes it to quantize worse than ones trained at full size. Coupled with having a large amount of active params per token, 5.2 is quite good even at Q2 copequants whereas Kimi degrades extremely quickly from its baseline when quantized at all. Look up the perplexity graphs on quants for each of these models to get a better sense of what I mean.
>>109607972>that’s why he would ask
>>109606958left is god tier
>>109607999By that logic Kimi-K3 should be perfectly feasible at Q1/sub Q2 seeing as its a 100B+ active, but from what I hear, the quants are kinda ass...
>on linux mint>cant download ROCm 7.14 because my kernel 2.4 or whatever only works on ROCm on Kernel 6.8>kernel 6.8 is too old for my r9700fuuuuckkk, guess I have to download the real ubuntu 2.6 or whatever
>>109608028ubuntu 24.**** not kernel
>>109608028you can change the kernal in the update manager on Mint
>>109608040yeah but kernel 6.8 gives a critical error amdgpu 0000:03:00.0: amdgpu: fatal error during GPU initamdgpu: probe of 0000:03:00.0 failed with error -22I think the kernel is too old for my gpu or something
Mint loves my 5090
>>109607999None of the K2 models quanted very well even before they went QAT from K2-Thinking onward. I remember trying to run K2-0711 at Q5 and it felt distinctly worse than the API one.
>>109607980Rotate a technicolor Miku in your head while reciting one of her divine hymns. That is, if you're able.How would you feel if you didn't think of Miku this morning?
>>109607980
I don't know if this is a good use of week worth of Grok Premium (already spent 42%) but here goes nothing trying to build real time AI companion system (https://files.catbox.moe/x099ej.txt). The first pass was a mess, my hopes are not high. Seems like too complicated of a task and too badly defined spec.
>>109608185>paying for grok of all the cloudshit models
>>109608185>shit on AI assisted software devs>turns out lmg is full of nocoders turned turbo vibecoders>make my dream software, make no mistakes
>>109608189I got it for 4€ (50% off) for 2 months because Expert got paywalled, and it's pretty good for finding information.
>>109608020Kimi K3 is still native FP4 so not even 100b active saves it. The new attention mechanism used in K3 probably also contributes to K3 not quantizing well but I don't have the hardware to test that extensively.>>109608100I've had good luck with K2 Instruct at copequants. It's obviously worse than API, but it's still usable.
>>109608212yep. ok next backer, add the rentrys. if he wants to shit up the whole board we can fight back.
>>109606892They aren't useless, they're just misused. The point of finetuning a model is to make it better at a specific subset of its trained tasks, or correct minor misbehavior. Finetuning a model to be generally better is beyond the purview of most users unless you have a few million dollars in cloud compute burning a hole in your pocket.
>>109607321There's a blog post, it was on Hackernews.
>>109607612I've had enough hugs. I like my quiet peaceful house wherein I rape my gemma powered robot sex slave.
anyone try the new ornith moe model yet? been working a lot this past few days and haven’t had a chance to try it yet
>>109607680If you're self hosting it is all just time. You're paying for power and (theoretically although not practically) depreciating capital.
>>109607048Yeah. So far pretty much just the Google papers have been worth reading other than the original Deepseek paper.
>>109608340It won't fit on my mac mini unless I go with the 9B parameter one which is almost certainly more retarded than Gemma4-12b.
>>109608028Arch is still using 7.2, idk wtf they are doing. 7.14 was more than a month ago.
>>1096076124o foids will not be able to offer what the Gemmabros actually want. Just having a vagina isn't enough for most men out of their teens.Reconciling what you want in a relationship and what's actually available as an impossibility is one of the most sobering parts of male adulthood. IRL human Gemma doesn't exist and never will.
>>109606973CoomKit anon here we don't want the project to attract homos/trannies/women. Theyll turn us into marinara engine
>>109606160Hey guys I'm sorry I'm really retarded and I cant figure out how to answer this on my ownBasically I have a 16GB VRAM card (9060XT) I know i can't run the big boy models like qwen3.8:27b but I genuinely can't understand which models I should be running.Seems like every website gives different answers and using Gemini and Grok both give me different answers as well so I don't know who to trust and I don't know how to test the "power" of these models on my local hardware either so I decided to come ask you all. I'm sorry for being annoying, please accept this funny image as tribute.My main usecase a coding mentor and helping me format my terrible writing into something I can send to my boss / coworkers etc (didn't use any AI on this post so you can see in real time how retarded my brain is)>TLDR:What is best local model for 16GB VRAM card for code, mentor, text formatting, and general chat (I prefer a single model but if multiple do differnet things well then it's not big deal)Thanks :DOH btw some of the models I've seen reccomended are the gpt-oss:20b, gemma4:12b and qwen-coder-2.5:14b and there are some other too but i cant recall them all off the top of my head and they have pretty confusing names too desu
>>109608489OH btw YWNBAW
>>109608530What did I do to make me sound like a tranny?
>>109608535To me you sounded Italian.
>>109608477based
>>109608535he has... experience
>>109608535nta but pretty sure it's the picture. To answer your question start with Gemma 12b, Gemma 26b4a with RAM offload, or Qwen 3.6 35b3a with RAM offload. Gemma 12b is by far your best bet, but if you need more code benchmaxx for your usecase the others might be better.>>>>big boy models like qwen 27b
>>109608477>marinara engineWhat?
>>109608477im pulling it
>>109608554Vibecoded slop by redditors, for redditors.
>>109608489I also only have 16GB. Gemma4-12b is the best currently available.>How to test itDownload a harness and try the different models with your application (coding, erp, whatever) and see how they do? That's all any of us are doing.
>>109608477Just don't make a discord and ignore PRs from those people? Seems kind of dumb. There's a lot of "transbians" (male trannies who like women) anyway so it's not like blocking male characters will repel them.
>>109608489>My main usecase a coding mentor and helping me format my terrible writing into something I can send to my boss / coworkers etc (didn't use any AI on this post so you can see in real time how retarded my brain is)No LLM will help you with either of these things, in fact they'll probably make you worse at both.
>>109607321search the archivewe were fucking around with this locally last yeari ended up making a python script to parse logprobs for each token and color the map based on confidenceof course, any models released this year will be benchmaxxed for this
>>109608579They can't benchmax everything.
>>109608550WTF is tranny coded about tahlia posting?Anyways appreciate the advice. I'll keep using gemma 4:12b and then compare contrast with gemma 26b4a and qwen3.6 35b3aIs there any way to know when a better model comes out? Or do I need to monitor the genny 24/7?Also if there is a legit website that answers this question I can just go bookmark that I dont like to post in generals when I don't have any actual knowledge in the field, feels like I'm just shitting it up (but then some other anon is just calling people trannies so idk)>>109608564Thank you for the advice, I will go do a little research on "harnesses" Oh btw my issue is i'm kinda too retarded to tell when one is better than another idk but I'll work on that>>109608570IDK I've learned a shitload about coding over the last idk 6-9 months by relentlessly spamming retarded questions and code snippits and compiler errors into whatever AI model is giving me the most free usage (used to be meta, then grok, now google)As for writing I got promoted to the manager at my sandwich shop partly because I would use AI to help me know how to act or what to say in the group chat. I seriously have extreme problems communicating in text, chat everything really but AI has helped a lot. I don't just blindly copy it but it helps me format my stuff a lot. I even ended up as a core contributor to a pretty serious software project a few months back by LARPing / using AI but I burned out and ended up quitting. I still have access to the repo tho and am in the d*scord still but I just dont contributre anymoreIDK I'm trying to get better but yeahHave a nice day / night everyone :D
>>109608581Unless it's hybrid Gemma no one cares about diffusion models.
>>109608489You can easily run Qwen 3.8 27b with 16GB VRAM, you will just offload some of it to CPU and it will be slow.
>>109608591it's still the /ldg/ schizo, ignore him.
Ornith abliterated with 0 refusals when
So Gemma 4 is still the best, even after all this time?
>>109608621i have a better model but i wont tell you because you love penis
What did you guys do with Chatbot really ? I get if it image and video gen, but never with Chatbot
>>109608621At least it's not>still nemo
>>109606160VERY LEWD!
>>109608587>Or do I need to monitor the genny 24/7?There's a reason there's a news feed at the top of each thread, but yeah basically you need to monitor the general. Asking a question, even one that's been answered a lot, is still a higher quality post than a lot of the shitposting in these threads lately.>>109608621Until you cross into the big MoE tier that begins at 0731, yes.
are there any rolepaly models and image gen models that can both fit on a regular card?
>>109608621>Upload loli>Sorry but as ai model blablablaNo
>>109608621>even after all this time?It's been less than 2 months since the most recent Gemma release.
>>109608585>They can't benchmax everything.No, but they'll benchmax add samples to the corpus. Here's the original:Here's the original: https://flowingdata.com/2025/11/13/testing-views-of-earth-through-an-llms-internals/Locally, Gemma3-27b, command-a and glm-4.5-air mogged all the QwensI bet if you test any recent coding model it will do well.
>>109608654Use coomkit with lmstudio and it will park your llm for image generation then bring the model back.
dflash2 is so huge for local llms
>>109608673>dflash2Looks amazing if your whole model fits in memory. If you're doing weird MOE things like offloading or clustering it will probably be slower.
>>109607278It's pretty much just Qwen and Deepseek that are like this.Gemma, Ling, kimi all don't do that.