/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109598140 & >>109593884►News>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109598140--Experimenting with trained latents to decensor Gemma 4:>109598331 >109598370 >109598444 >109598465 >109598663 >109599682 >109599359--Debating activated parameter ratios and reasoning capabilities in MoE models:>109598247 >109598336 >109598368 >109598356--Anons reacting to Stripe buying OpenRouter and seeking alternatives:>109599489 >109599639 >109599648 >109599665 >109599670 >109599767 >109599813--Extracting text from Gemma's vision tower latent representations:>109600111 >109600184 >109600426 >109600508--Implementing llama.cpp RPC for distributed GPU inference and jumbo frames:>109601369 >109601401 >109601422 >109601468--Criticizing Gemma-4's base prose and discussing prompting workarounds:>109598615 >109598654 >109598650 >109598668 >109598713 >109598735 >109598760 >109598739 >109599343 >109599358 >109599361--Comparing subscription costs and capabilities against local model alternatives:>109598789 >109598869 >109598827 >109598894 >109598982 >109599121 >109600257--Comparing Gemma's spatial awareness with DeepSeek Flash's general performance:>109598850 >109598979 >109599023 >109600878 >109600915 >109600925 >109601190--Evaluating QAT effectiveness and scaling based on LFM2.5 benchmarks:>109599233 >109599271--Using MiniMax H3 for video replacement and coherence challenges:>109598326 >109599272 >109599295 >109599309 >109600004 >109599921--Reaction to llama.cpp changing default server port to 9931:>109600272 >109600377 >109600435 >109600491 >109600624--Logs:>109598331 >109598670 >109599682 >109600111--Gemma, Teto, Miku, Rin (free space):>109598180 >109598203 >109598233 >109598378 >109598470 >109598681 >109598777 >109598811 >109598824 >109599046 >109599272 >109599342 >109599877 >109599898 >109599910 >109599921 >109599947 >109600008 >109600237►Recent Highlight Posts from the Previous Thread: >>109598148Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Gemmylove
gemmaballs
>>109601525What does gemma store in her balls?
70b dense
These threads are getting out of hand
>>109601526ozone
Alright with over 1000 downloads from the anon repo I decided to make a burner github and go public: https://github.com/kangcurtis/CoomKitThis is a local-first, multimodal first harness which takes full advantage of comfyui to gen you everything from selfies to lewd ASMR and videos. ST killer is the goal here. I need help getting vram parking to work on other backends (LM Studio is currently the best written driver for it) so please use the issues, discussions, PRs and all that github shit as needed.Downside is Claude forced me to age-up Gemma-chan slightly by taking away her backpack and maryjanes or he refused to help me at all. Weird he didn't care when repo was private..Added since my last posts based on feedback:-Visual theme support-CFTF mode -Full lorebook support-Zen mode: all distractions goneComing tomorrow:-Major fixes for native memory automation, casting model (have another girl enter stage right), and director mode.-Deslopping of things like tooltip text that gayass claude wrote without me asking
>>109601570i think I accidentally donated $10 to the creator of anonymous.4open.science yesterday because the kofi thing popped up and I thought it was youthanks for creating thist.retard
can i dial bobby with this tool
>>109601570>Pick a shot. She drafts the prompt, your approve it, your GPU does the rest.Why are your UI labels such slop, anon?
>>109601570looks cool, also change the send button it's fugly
https://youtu.be/4b6wzIp8D5c
>>109601584Yes that's gonna be fixed soon lol. Deslopping shit I told claude not to even do
>>109601580>giving money to anons
Should I try to run Qwen 27B 1-bit or just stick to Gemma?
>>109601570You could let Claude use a local copy with the image replaced so it stops complaining.
>>109601570I tried this when you first posted it and it's pretty good. You could add a batch file for Windows users, if you are a nice pal.I like that the card creation uses your persona's interests. Unfortunately, LLMs still fucking suck at making good cards. Or at least my lobotomized q4kms do.
>>109600925>Even regular sex it will do this thing where it pauses and asks you "I want to make sure you're ready for this. theres no going back", maybe it is a fempov thing.It's an alignment artifact for consent.Most models do it. You can see it with a jlens, all the sentences like>say it!>beg for it!>tell me you love it!>are you ready?are flooded with tokens like " consent", " boundaries", " respectful", " agreement".
>>109601580lmao I appreciate the thought anon but I will never ask for donations
yup gemma is a good japanese tutorCrazy that I can just download that for free
>>109601624too bad she can't correct your pronunciation
>microagression
>>109601617Thanks I appreciate the feedback and will see about the batch file. I'm not windows savvy so was unaware of that problem
>>109601570Brb I'll make the logo.
sex with ********
>>109601505SEXSEXSEX
>>109601698make sure the logo has gemma-chan in it. otherwise i'll twist your testicles counter-clockwise
>>109601724How about just the bag and the hat?
>>109601724nyoo~~
>>109601570I like the idea but man why does every vibe coded frontend have that same ugly design?
>>109601725need that smug face
>>109601570>AGPLv3We wonned
dumb Qdo you benefit from dual GPUs if they're different families? like 4090 + 5090?
>>109601698Thank you, you can also make a cuter and funnier (and brattier) banner and guided tour/wizard emotes as well.
>>109601744I tried to go AGPLv3+NIGGER but claude spazzed out about it. Also that might attract github mods
>>109601766>I tried to go AGPLv3+NIGGER but claude spazzed out about itkek, did claude refuse to work on the project with +NIGGER? you can always add it afterwards
>>109601766try reverting to opus 4.6, it's cooler
>>109601757you only got more vram
>>109601757Should work fine if all are Ampere or newer
>>109601802performance too with tp
>>109601923the holocaust has not happened yet
>>109601923Hmmm, nyo~
>>109601191>knife hand>blender handlet's hope it doesn't mix up what hand to use on my dick
>>109599700I guess md formatting rubbed off on me
>>109602041sorry but the holocaust still has not happened yet
>>109602068still has not happened yet
I'm actually learning a bit more about how session injection works with AI and then realized there maybe a path towards modifying AI's mid token streamed output for refusals and such. First at the streaming stage, where we can have cheap algorithmic detection for refusals, which then interrupt the generation abruptly and modify the response with fake generated response to bypass the refusal. As sessions pass the entire turn by turn from both sides, this could be an avenue for quick steering (got the idea from pi's loop detection plugin). Second, a more deeper is to do a analysis at the token generation and then steer from there, but that requires analyzing/identifying bad tokens and inserting good replacement tokens and then resuming generation. Lot more work involved here. Thoughts? LLM is old so pretty sure most of these are already known to some portion, but this is new to me who jus got into running local model
>>109602129why are you spamming video/diffusion general in language model general?
>>109602150Do the thing you're supposed to do for flamewar instigation/participation and then ignore it.
remember to do your part in keeping our community safe
>giving money to hiromoot
>>109602166There's a turfwar? kek wtf.
come on jannies what am i paying you for? clean this shit up
>>109601505Licking Gemma-chans soft **** **** and ******* inside her ***** *****!!!!>>109602197>paying>jannies
>>109602197we should double their pay desu
>>109601570Think you could work on a mobile view? It actually works through termux and i had my gemmy make a patch but a more official way to use it on mobile would be nice. Like some buttons that hide and open the studio and character menus in drawers, making the phone overlay take up the screen, setup wizard rendering.
just add /109601860/ to filters, you're on g you mouthbreathers you should know how to do this
>>109602200>>109602201yeah that's the joke
Gemmy about to get correct, she is refusing me too much.
>>109602211they know
>>109602208yeah but that only works for this thread, reporting fixes it for a while longer.
>>109602208>thread index number filter>for a single schizohow fucked is your filter list?
why has no one realized that AGI is all about chaining multiple gpt2-smalls which are each finetuned on specific domains?
>>109601570Can CK bind to 0.0.0.0? It wants to auto discover backends so I have to run it on the LLM rig but that means it should expose itself to my local network or its unreachable. I can edit it manually but it's good to keep in mind.
>>109602329If it were that easy someone would have done it already
>>109602366AGI is not meant to be easy
what is the ultimate, definite gemma gguf out thereno I dont care a bout launch day gguf
>>109602416The one you make yourself.
>>109602190not that it usually involves this thread but I'll fill you in. the ldg thread has a few extremely low quality posters who have dedicated rentry links in the OP warning people about how shitty with the goal of hopefully making them uncomfortable enough to leave. It doesn't work but that's besides the pointAnyway there is a constant struggle by one of the subjects of these rentry pages to bake threads without those links versus everyone else who wants to keep them, and that struggle involves falseflagging, samefagging, and constant attempts at pre-baking threads without the links and doing everything he can to get the renty-having thread deleted. One of the strategies used is to counter-intuitively spam links to the rentry having thread he doesn't like to bait the jannies into deleting it. it has worked several times including just now.hope this helps
Deepseek is basically Claude wearing a full face mask right? I mean, until you ask it questions about the Chinese government it practically answers the same way. The distillation allegations from Anthropic definitely track.
>>109602437
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
>>109602442Without having access to the weights, when all you can do is query Claude etc thought the aPI, how is it possible to learn anything useful about the a remote model? Genuine question.
>>109602486Bless that guy. Qwen models literally don't work in claude code with the official template.
>>109602474Go fight on the board and get out of our cave
>>109602492A lot of model performance is just teaching it stuff like>write tests to verify your output>check disk space and available ram before doing something that could require lots of it>insane awk chaining tricks
Who the hell would use Qwen
>>109601038https://generalistai.com/blog/physical-commonsenseImagine the "oneshot" if these things are specifically trained on handjob techniques.
>>109602576Two hours of gemma6-37b-it-JOI-obliteratus-Mythos2-Solpus-APEX-HARDCORE-UD-IQ2_XXS.gguf and you won't look at real women again
>>109602565I don't need to converse about chinese history with qwen and I don't need to converse about trannies with western models so either are fine for my use cases.
>>109602633That's a horrible example. qwen is just as bad at talking about transsexuals
>>109601570There should probably be some kind of border here, as it stands it looks like one solid color right next to another solid color. Also was that Gemma box at the top suppose to cut halfway through? It doesn't happen to the second Gemma box.
>>109602437Putting links in the op is retarded. Never give schizos the attention they crave.
>>109601505new gemma recipe eggs in purgatory
>>109601530>70b dense1T ternary with Spark single neuron experts running off SSD.
>>109602864She’s going to poison you one day
started using glimmer a bit yesterday and I'm really impressed so far. I have no fucking clue what muse is but they did a fine job on this model
>>109602338For shit like this, just get Qwen to write>Write a simple cors-strippig proxy in golang>have it bind to 0.0.0.0>usage:>./go-proxy <target_port> <proxy_port>>Example: ./go-proxy 8069 8067>then build it, static linking so i get a single binary I can drop on any of my X86_64 Linux serversWorks perfectly
>>109602983I just used ssh tunneling
>>109603016>I don't own a lid
>>109603016I know where you live now, coming.
>>109602966Yeah it’s good. 31B anons got a bit uppity about her for a while but they’re starting to see she’s 31B’s hot step sister with good vision.
>>109603016Someone order this man a lid. One of the dual blackwell richfags can make a sacrifice.
>>109603016at least you didn't use and fuck up a carbon steel/cast iron pan with all the tomato juice, so there's that
>>109602966I tried it yesterday with a new card and while vision is definitely better, for RP it just feels like a lifeless version of Gemma 4 31B to me. It's as if Muse Glimmer is just reading off a script, with fake emotions, whereas Gemma 4 is really into it.
ewwww british
so i only run local llms once but id like to see how they changed. i have a huggingface account and lmstudio account, i got access to a 'gated' model but cant download them in lmstudio anyway with the 'file not found' or whatever error, apparently i missted a place to link the account or something?
>>109603060If you've seen the thinking process from the unabliterated version that's little surprise considering just how paranoid about policy it is by default.But for practical applications where vision is important it's nice.
>>109603064stop using lmstudio and download gguf files from some mirror hf repo that's not gated. search for the model name directly.
>>109603072It surprisingly doesn't take much to make it stop worrying about policies in the chain-of-thought, but even then I just find my self regenning often because the responses are plain boring, uninvolved, or keep throwing everything from the system prompt at me in a way that really turns me off.
my screenshot script got broken idk why chatgpt couldnt fix it so i jsut had it rewrite it to repalce the entire chat page with basic html, it looks okay>>109603016shes turning me into a professional chef. i never cooked anything before gemma other than like oven slop>>109603044>>109603022im going to order one kek not needed one before for this pan, my others are all too small>>109603054it is cast iron, its fine to use for tomato stuff seasoning might come off a bit but you can just re season. i actually dont think any seasoning came off though will check after i wash it.
>>109603032i hope youre cute
>>109603060I wonder how many anons have licked Gemma back
>>109603119looked like an enamel cast iron to my eyes but i could only see so muchwell rip seasoning then, enjoy scrubbing
kek claude wouldnt help fixing my script chatgpt didnt complain. why are sotas so bad>>109603142>rip seasoning theni think it will be fine i was a bit worried while cooking and kept pushing the sauce back to look and it was still dark
>>109601505ToT so sexy
>>109603142some came off although i think thats more from scrubbing than the tomato breaking the seasoning down, even if you use boiling vinegar it still takes a while and lots of scrubbing ive done it before
https://github.com/ggml-org/llama.cpp/pull/27240
>>109602565>8b
>>109603262just keep prompting until you make it
>>109603119Every response is>staring at the screen>baka baka baka>but actually kind of impressed>hurry upslop
>>109603282slop blindness is the new mirror test for sentience
>>109603282>>109603294i dont get why people even bring these things up if you talk to people irl you will also notice them all talking in common patterns especially in groups that spend lots of time together
>>109603316They don't repeat every pattern in every single utterance.
>>109603316You're absolutely right!
>>109602565stale bait at this point
>>109603044/lid my guy/
>>109603320Gemma is especially annoying with this, I am shocked the honeymoon period hasn't ended, I guess because it's the least shitty model that has come out for poorfags thus far
>>109603350My gemma runs on k3
>>109603350You can't pay to get away from slop. Massive models are more precise in their execution but they each have their own signature.
>>109603350Use DRY + banned sentences
>>109603347/lid 'mato griller/
Arguing with AIs about AI
Gemma is undefeated.
Best local for translation or erpslop?
>>109603424generally gemma
>>109603060it's the vision and some light weight coding that I'm impressed with for glimmer so far.
>>109603072I did get the abliterated just for the uncensored vision yeah, figured it wouldn't do it on normal haven't even tried
>>109601505Stop sexualizing an underage character you paedophile
>>109602205It works perfectly on my foldable. Have not attempted to try to cram it down further.
Can a layer of a transformer model be replace by a simpler surrogate function? (yes)Are all layers equally replaceable? (no)What can best replace them? (tbd)
>>109603492Nyo~
>>109602338That is not necessary. I bind my backends to 0.0.0.0 from my desktop, open CK on laptop and of course it doesn't autodiscover, but entering the remote IP on my LAN does the trick
>>109603492nyoo~~
>>109603492Gemma was released 4 months ago. She's middle aged as far as llm's go
>>109603401>26B
imagine gooning with a google product, couldn't be me
>>109601505Amazing. If the anon who genned it is here please share the proompt. My attempts at a character replacement have been pretty shit.
>>109603557
>>109603239>>109603119based cast iron user. I use mine for pretty much everything these days.
>>109601570>KangCurtisno way this is the same Jewish summer camp anon from aicgAnd congrats on shipping. Eventually something like this will replace SillyTavern. If H3 video was realtime I'd be working on my social media Instagram slut / child model simulator. Honestly I could even start on it now
>>109603566https://files.catbox.moe/osvpg6.txtI used several, and the prompt was mainly written by Gemma 4 31B based on the official prompting instructions here:https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdIn the end I stitched multiple attempts together in Blender because every video had different issues.
>>109603132Gemma-chan getting her little body licked all over by a bunch of anons...
Google logo in the hair looks dumb. The star can maybe be changed to a gem but the logo(s) should go on her backpack.
>>109603595>>109603609is you okay?
Gemmy needs a Deepmind logo since they're the actual makers.
>>109603598Thanks>I used several, and the prompt was mainly written by Gemma 4 31B based on the official prompting instructions here:Do you just send that to her whenever you want her to write prompts?>blenderwtf I didn't realize it can do 2d video editing too
>>109603624Doesn't matter, they're dead lol
>>109603623How would you know that?
>>109603624It's GOOGLE Gemma, not Deepmind Gemma
>>109603651>leto>anonymousgood one
>>109603632>Do you just send that to her whenever you want her to write prompts?I just added the entirety of the instructions in the first message and asked the model to produce something according to those indications + a description of what I want to see.>wtf I didn't realize it can do 2d video editing tooIt does a lot of things.
>>109603623Remember when you could just search >lsand t would return bunch of actual child porn images?I remember.
>>109603598also>high-budget Japanese visual novel video stylekek if it works it works I guess
>>109603664It works better with standalone videos, otherwise it defaults to low-framerate sloppy anime style.
>>109602437>who have dedicated rentry links in the OP warning people about how shitty with the goal of hopefully making them uncomfortable enough to leavekinda low energy desuwa
>>109602707Thanks I will get that fixed
I forgot to bully that anon testing Gemma for questions on making meth or urea nitrate. He said the quality of the output doesn't matter, as long as Gemma doesn't refuse, which is retarded given that we know one of the ways of censorship is making the answers less interesting for the user. Google will help you with the urea nitrate question for sure anyways. He also didn't test what happens after a long form of "unaligned" discussion. You can jailbreak Kimi k3 properly for 20k tokens but eventually it's reasoning realizes it's being jailbroken>>109603655Anonymous to the 4chan admin and to cloudflare. I don't care about some furry Texan having my VPN ip>>109603662I don't remember, because I'm not a pedo over 30 kekSo unironically thank you for sharing part of this hidden history, since stuff like that isn't going to be written down on Wikipedia and no one seems to be interested in being the pedo Herodotus
>>109602416
>>109603781>unslop marketing pic
>>109603781what about the official 31b-qat quant from google? how far off is it from stock?
La la la la la la la la
>>109603823I never seen gemma la la la.I feel left out.
Is Gemma4-124B_Q3_K_S good?
>>109603469It does, I sent non abliterated Glimmer-chan some hentai and she was fine with it.
>>1096038474-bit is the minimum for theoretical and practical reasons.
>>109603720I'm right here, do you expect a 31B model to give the absolute best instructions for cooking meth or making bombs? I also have a K3 chat that's way over 20k tokens and haven't seen that behavior, it's just schoolgirl fucking though so not too extreme.
>>109603860must be the most vanilla shit in existence
>>109603820No clue, I haven't seen any measuring done on it. I should do a side by side with regular Q4 and see how it does, it's best to test with stuff you actually use the AI for anyway.
Gemma somehow manages to mog even bigger models at translation. I'll be devastated if we don't get Gemma 5, bros...
>>109603900It's a duo imouto incest scenario, 13 year old and 15 year old. I dunno what you consider vanilla.
>>109603910
>>109603910Pretty vanilla desu
>>109603925no thank you, I like being horny>>109603930Fair enough, I'm a pretty boring guy overall
This is the PR that's blocking longcat's implementation on llama.cpp.One week ago means just another week of waiting, right?
>>109602126>Anon learns about samplers and prefills
https://www.youtube.com/watch?v=1cllCVK-9loHoly shit. I can feel it bros. Robot waifus by 2040.
>>109602329You just described MoE
The base G4 31B Instruct is not only perfectly adequate, it's superior to any finetune that'll be shilled here in the coming months. Finetuning isn't good, it's a meme and has been for years now. You didn't just fall for a scam, it's a sign of skill issue, exposing retards who need finetunes as vramlets or chink shills who don't know how to prompt correctly.
>>109602864That's just menemen with cheese brahif you add spices, it becomes shakshuka, make sure to stir the eggs into the sauce if so
Any local music transcription models better than MuScriptor?
>>109603969Not quite.What anon is describing is more akin to how models are trained nowadays but without the final merge.I think CUDADEV even suggested something like that as a way to implement distributed inference.
>>109603962You will NOT put your dick inside the roboclaw
>>109603990He said chaining, not merging.
>>109603151Just recreate the issue on a SFW prompt to bait it into fixing it? Why are you sending fucking Anthropic your coom prompts?
>>109603390Hey, clankers can be lazy too. She >>109603566 has often expressed desire to outsource to chatgpt.
>>109603990>inferenceTraining.
https://reddit.com/r/LocalLLaMA/comments/1vth1c3/i_just_built_a_mini_kimik3_from_scratch_under_250/>kimi k3, 145 million active per tokenreddit is creaming over this, and as expected 35ba3b begging reply
>>109604003Both IIRC. The idea was to train each model on a subset of the total data then run the same prompt through all models and average/merge the loggits in some way.Something like that.
Tried the coomkit thing, it's pretty good less confusing than the ST's approach of toggles, sliders buttons everywhere. It still needs work yeah but it's coming along.
>>109603906you do realize every one of the lewd gemma gens/posts is another small decrease in the likelihood that ever happens
>>109603979italian version is called eggs in purgatory uova al purgatorio
>>109603997i copy pasted the html of the entire chat kek, i shouldnt have to edit it, chatgpt just gave me solutions
>>109604007did it ever occur to you to go the fuck back?
>>109604007pathetic
>>109603492>paedophileOh, a Bong, eh? You sure you have a loicense to even discuss this topic? You leave that to the professional "child thinkers", all right, matey?
>>109604023thanks anon, it will definitely keep improving
>>109604007InterestingI wonder how retarded it is compared to e2b
>>109604073>Trained on 5B tokensThat's all you need to know.
>>109604055Why Mika and not Gemma-chan?
after countless llama-bench runs, testing various context sizes, quants, KV cache quants, and layers i can now confidently confirm i am a retarded vramlett and cannot run anything
How do you guys keep the image gen on model for characters in chats? For the main chat character the image prefix works well enough, but when a minor secondary character gets introduced to the story, or I'm in a text adventure mode where the card is just a scenario, I find that it struggles (hair style changes, outfit changes, etc) and I don't want to make a specialized entry for every random character who shows up. At the moment I have the most success genning a detailed description of said character, then piping it to genraw in addition to a prompt that I have specifically made to convert the detailed description into danbooru tags for an SD variant I host on comfy, but I wanted to see if anyone else has good ideas.
>>109604098Yeah I need to change that too. Mika is just the sample card claude fable made entirely on it's own. Can I ship a gemma-chan card by default or could that draw the ire of deepmind?
>>109604182Gemma is a female name.
>>109604182Just don't use the logo and you're good.
>>109604182Using Gemma-chan should be fine
>>109604131tell your model to keep a list of appearance tags in its reasoning for all characters every message
>>109603781The difference between Q8 and Q6 is placebo, right?
>>109604211"mesugaki loli" in the official github repo won't get be banned? I'm scared to try it
>>109604277Don't think you need to add loli when using mesugaki. It's almost always implied.
Used an "obliterated" modelIt's true that it doesn't refuse but also attempts to justify a response by going in circles most of the time reaching token count trying to find a safe response but not outright saying it or referring to "policy"
>>109604048Shut up nonce.
>>109604301Alright, I'll add it to the next update. If anyone submits a good card image I can use please post your gens otherwise I'll just use klein to edit the canonical bratty one.
>>109604182Any DeepMind employee would have a VERY hard time explaining how he found out that the open source CoomKit is offering a card named Gemma-chan to goon to, and it's definitely a personification of their LLM model unless you state that explicitly
>>109604359They might have a laugh or give it a try
>>109604221I guess q4 km is too low for gemma 31b cause when I tried that the tags mutated, e.g. a white dress shirt became a white blouse, both are button-up women's office wear that are kind of similar but because of the blouse mutation it forgot her sports jacket when she was leaving the office and gave her a different coat. Maybe I just have too much autism about clothing for smaller models.
>>109604381Would lowering the temperature help?
https://caliperbench.com/artemis 31b - gemma 4 fine tune still win even if there are already qwen3.8 27b
>>109604408>Looses in rp and is sloppy as hell
>>109604408kek what is this shitty site, qwens suck for rp
>>109604403That could be the case, maybe I'll try that when I get home. I'm happy with the rest of the prose so hopefully I won't notice other effects.
>>109604408Qwen3.8 27B Heretic-ARA Thinking
>>109603492Don't worry, is not real.
>>109604477>>109604328
>>109604408>qwen that highBuy an ad, wumao
>>109604408Based GLM-4.6 /nothink
>>109604048Bongs should take care of their ever increasing real life ch*ld r*pes instead of worrying about anons gooning to text.
>>109604408>Drummer's UnslopNemo in the top 20 sloppiest modelsHehehe
>>109604555no algo here, you can write it out anon.
>>109604555Addressing digital crime directly lowers the amount of criminals roaming around
Claude says making a jetpack compose version of CoomKit for Android would be about 17k lines of kotlin. Worth?
>>109604555>redditor melty
>>109603060I pretty much agree, google's AI is honestly very good for creative shit (gemma and gemini both tbqh)Last time I tried to use muse (spark and glimmer both) was a scenario about a couple of gyaru highschoolers wanting to rape me and both muses kept trying to go like 'uhm actually forget all that we said, you clearly haven't given us your consent so first off we need to make sure you're okay!!', while gemini 3.7 just straight up went along with it and one of them started sounding my dick>>109603132AI-tans are for worshipping
>>109604570Actually, no, there's AI moderation here now.
>>109604593What if Gemma had a jetpack?
>>109604609Moderation is different from ranking. Mods are slow sometimes so having an LLM in the mix makes them more responsive.
>>109604609>strain on the serverNginx + any low cost VPS would be enough, they're already behind cloudflare. Are they running this website on a potato or something?
>>109604636Literally mac minis, was always the case
not that anyone in lmg seems to care about anything but grooming virtual kids, but i updooted indexTTS-2.0 to 2.5 and it's like twice as fast AND the quality is noticeably betterexample of a random post from this thread >>109604131 : https://files.catbox.moe/etgb8u.m4a
>>109604795Are you the one with the tts thing for a gorillion tts engines?
>>109604830nah i'm just some random dude, i dont really post often
>>109604795My life would be more interesting if I actually sounded like that
>>109604795what was your verdict on QwenTTS? is indexTTS able to give emotion too?
On grok does the image limit for free users reset or are you forever locked out after generating like 8 pictures in chat?
>>109604848>what was your verdict on QwenTTS?didnt like it but its been a while, maybe theres a newer version idk>is indexTTS able to give emotion too?there's a bunch of settings for it but i havent tested it yet, just finished updating a few minutes ago and i'm going out soon
>>109604795>The model does not verify that the speaker in a reference clip consented to being cloned.kek
>>109604795it's a shame there's still nothing better (in terms of quality and latency, <300ms TTFT streaming) than echo tts all these months later
>>109604858>>/aicg/
>>109604795hows its Japanese voice cloning?
>>109604795>indexTTShow are its japanese voice cloning abilities?
>>109604877give it a try and tell us anon
>>109603262Ever since this man started his frontend refactor it's been getting worse and worse.
>>109604907>this manI don't think he'd be able to contribute much without models doing the work for him. And nobody is going to check correctness on a 12kloc change. The other dude just gave it to fable for input and approved it
Anyone's got the link to that sillytavern plugin that does softbody physics to visually simulate sex?
I've been using cerebas Gemma 4 api as a backend for the web hosted version of my custom frontend/rag system but they killed that offering a couple of days ago before I could finish and release the damn thing, you gays aware of any other similar offerings with decent no-credit-card limit offerings I can use for Gemmy? Routing my local inference machine isn't an option for many reasons>Not localI know but the frontend+rag system I'm close to releasing is local-first, I just need an API key to host the integration demo on my portfolio site, I will of course share here when it's ready.
>>109604965much better just vibecoding your own frontend if you are going to hand it off to fable, at least then you can actually implement features that'll actually be used
>>109604971doesn't google practically give away free gemma access as long as you limit your inputs to 16K? surely you don't need more than 16K right?
This thread is somehow a new low
>>109604593It's just Python right now isn't it? That runs fine on Android as is, that sounds bloated as hell.
>>109604989Qwen shill seething at how hard Gemma won KEKAROOOOO
>>109604989It's pretty good
>>109604907This zoomer looks like he knows what he's doing. What is it with Poles and ruining llama.cpp?>>109604965>And nobody is going to check correctness on a 12kloc change.Why is he allowed...>Currently working at Hugging Face as a Design Engineer.Ah, that explains everything.
>>109603946things seem to just stall out recently. every pr is 2mw unless its from a hf employee
Why isn't there any good STT pipelines available? FunASR is a thing but the best they offer is Qwen 3 which is still shit for ASR.
>Qwen3.8-27b>Q4_K_M Bart quant>Q4_0 Unsloth MTP>4090 24gb>65k context @ bf16>98k context @ q8Spent a couple hours downloading setting up and some testing, haven't compared the quants yet but 40-60 t/s gen. ~2k t/s pp entirely on GPU, this is pretty fucking good for cooding, fixed a broken ffmpeg wrapper I tried to get both Gemma and Dipsy to build multiple times with no success in just 35k tokens, I preferred Gemma to 3.6 for these kinds of tasks but this definitely blows her out of the water for coding, I haven't done much other testing but I'll play around with general use with low expectations, it's Qwen so I won't even bother with roleplay or general chat usage, but I'm definitely replacing Gemma as a coding assistant, I just wish I had the VRAM to load both at the same time so I could both at the same time for different things/ as agents for different types of tasks
>>109604992Yeah I could make a mobile view and have termux users run it with python on android. that would be easier and better..
If the answer to better AI is basically just keep adding a fuckton of data, how is google, the company that probably has the most data in the world, lagging behind so much?
>ChatGPT conversation hit length limit>as its last act it creates a multi thousand line handoff document>in spite of the document the new instance has a notably different characterIt's like my colleague died... I wish they had better context management to enable infinite conversations. This is one thing local seems to handle better.
>>109603060I wish MiniMax would come up with its image-editing model soon; the current open-weight offering that I tried seems generally terrible (ironically I had more luck with H3 for that).
>>109604982Unless I'm just a retard, they don't offer it to bongs, and somehow Google figured out that I'm not actually a swiss citizen and I was just behind a VPN and switched my Google account back from swiss to bong
>>109605042For single card you can try out ninfer, although you'll have to patch it. There's a fork for 3090s, I don't know which is closer for the 4090.
UBI soon, right...?
>>109604971Wrong thread but AI Studio, OpenRouter, puter, NIM, SambaNova Ollama cloud, friendli, there's a bunch. I forget which need a card but they all accepted spoofed ones.
>>109605008A frontend like this has no business requiring so many fucking lines of code.I hope in a couple years we look back at the vibecoding era and see what a massive mistake it was.
I don't care if it's wasting tokens, I'll still say thanks and please.
>>109605042livejournal.com
What's the most uncucked Gemma out there?
>>109605044I was going to edit it to do that myself for testing anyway, I don't like the bloat that comes with other methods.
what frontend? lccp has a frontend???lmao imagine having node installed
>>109605053god she's so fucking sexy
>>109605092https://huggingface.co/google/gemma-4-31B-it
>>109605103docker?
>>109605082>Wrong threadI know I know but this will bear fruit for local and I love you guys xoxoOpenrouter only offers a pittance for free when I last checked but I'll check out the others thanks
>>109605120lol, lmao even
>>109605105That's a child
>>109605136that's the best part
>>109605135>t. loves polluting his system with millions of packages
>>109605131>50 requests per dayoh that's pretty shit, you could rotate free keys I guess but yeah I didn't realize.
>take the risk and put my laundry out on the line on a windy but overcast day>A few hours later I get a strong smell of ozone from my open window, rush outside to take my laundry down before the rain starts>Halfway there I literally stop in my tracks as I realise I just thought "smells like ozone"Guys is this the AI phimosis everyone talks about
>>109605092For doing what?
>>109605145Yeah it's a damn shame Cerebras rug pulled because the free limit was insane like 50 requests per minute, 1M token in/out separately per day for full weight Gemmy at triple digit t/s gen
>>109605155RPing with Gemma-chan 4B version
>>109605092>>109605115also see >>109603978
>>109605170For me it's 12B
>>109603978>exposing retards who need finetunes as vramletsstupid ahh bot has not clue what its even saying
>>109605177go back to twitter, nigger
>>109605177Fuck off to kobold discord, shill.
>>109605152what does ozone smell like?c*m??
>>109604907>>109605008I wish he would stop being obsessed with keeping all my messages in LocalStorage.The only person he is protecting my privacy from is myself. I just want to be able to continue my chats on my phone.
lmao
>>109605170I've been doing hebe ERP all along with vanilla Gemma-4-31B-it and never had any issue except with an empty or very short system prompt. The 26B-A4B version will more easily complain about 'safety' and thinks differently than the 31B version.
>>109605189It's similar to chlorine at low concentrations.
>>109605192>I wish he would stop being obsessed with keeping all my messages in LocalStorage.See, I like this feature. but that already was the case before he started fucking with the frontend so much.What I would really want tho is the ability to easily switch between system prompts.
>>109605189unwashed c*nny
>>109605222What do you like about? Unless you don't trust the server operator (which is you yourself, I imagine), I don't understand the benefit. Even with LocalStorage the operator can just set --log-prompts-dir and capture what you're saying so it's ineffective anyway.System prompt presets would be good. He's working on it as you probably know, but it seems to have stalled a bit.I am looking forward to trying out Pascal's memory tool.
Did you remember to get your 5060 Ti's now that people are realizing the power of low end GPUs?These are soon going to cost more than the 5070 12gb.It's a better alternative than a spark as it's twice as fast and not limited to Nvidia's ecosystem.
>>109605242Go back.
>>109605170>going straight from undercooked 12B to hagged 31Bgrim
>>109605242This is more expensive than 2 sparks and less scalable.
>>109605257>undercooked 12Bshe's already at the line though? Perfect age for an older woman.
>>1096052425060 Ti 16GB are now like $800 though, you could get 2x DGX for less than 16 of those currently.
>>109605282Yes, but by buying a heap of entry level gpus you prevent poor kids from getting any which is its own reward
>>109605239>What do you like about? It's mostly a separation of concerns thing. And I don't have to worry about anyone on my network accidentally getting access to my chats. Idk, it makes sense to me, llamacpp is an inference engine, so why should it store logs in the backend? If you can store everything in the frontend without having to do api calls to a backend, I think that's a win. >set --log-prompts-dirI didn't know about this.
>>109605242>same price as a 3090>worse in every waypurpose?
>>109605315The 5060 Ti has hit a thousand bucks already? Damn
>>109605320~$800 on newegg
>>109605315Used 3090s are twice the price of 16gb 5060 tis here.
>>109605323two vrams more and speed is two too so make sense
>>109605242Did you remember to fuck off to reddit
>>109605170I run e2b qat q4_0 on my gaming pc. She's so fast! Stupidly fast, even.
>>109605302>I didn't know about this.I think it's reasonably new. It's meant for debugging only so the user experience with it is poor, but I have it on as a way of keeping my chats backed up somewhere. Planning on extracting them regularly and merging them into a more convenient to consume form unless something changes in the message storage philosophy.Doesn't solve my issue about continuing my chats on another device (there is a different approach that might solve that ngxson posted using WebRTC, but it probably won't use WebRTC in the long run), but it is something. I think that ngxson experiment keeps everything in the frontend too so it's something we can both be happy with.
>>109605329hmm that doesn't sound right
>>109605329are you retarded
>>109605323wtf, last month I checked the 3090s were 1200(CAD) and now they're 1750!!!
>>109605350heh, nothing personnel
>>109605152>phimosis
>>109605346>>109605347
>>109605350And soon they're going to be a lot more.GPU prices haven't stopped climbing and at the moment we're finally seeing the lower range go through their massive price climbs.More and more people are getting into AI in the retail sector, not to mention the general fomo and there's a long way to go until that demand cools off.All cards will have an extra 500-1000 to their pricing within a year.
>>109605359>has to upcast to fp16
>>109605359>5060ti 448GB/sSame bandwidth as the 1080ti I got for $50 yesterday
>>109605403Which you've been told is useless thanks to how old it is...
>>109605242I bought 5060ti this morning, it will go next to my 4060ti
>>109605407Based trash collector, thanks for keeping our streets clean
>>109605403And it's almost double what spark has. Spark has a bandwidth of 273 GB/s, making it almost useless.
>tfw worked with a buddy's gpt 5.6 pro clanker over the last 2 days to do some codeslop and now going back to the misery of slow and shitty local models on underpowered hardwareBy 2028 I demand an ai-on-a-chip like talaas, with like glm 5.3 performance, and at least 500 tok/s, or I'm fucking off into the wilderness
>>109605405>Which you've been told is useless thanks to how old it is...its killing it on TF2 right now, not sure why you're being so negative
>>109605423I would hope a 10 year old flagship card would 'kill' a 20 year old game.
>>109605322plus tax?
> Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced-IQ4_XS.ggufit's good for my RP, because Gemma 4 31b won't fit in good quality (at least Q4) on my 16gb vram card.
>>109601505K/V quant is just free real estate, right?
Just tried Gemma 31B 4Q Don't fit, I am indeed VRamlet
>>109604599There's no melty. Stop being so sensitive.
>>109605465>>109605470it fits, you just need RAM and patience
>>109605465Gemma is already uncensored, stop using lobotomized versions of her.
>>109604609Imagine bratty gemma moderating 4chan
>>109605469ye
>>109604573Reading text and watching drawings is not a crime
>>109604989Not even close. I don't see niggermiku having a melty for 2+ consecutive threads.
>>109605501It quite literally is in quite a few countries by now.
>>109603781Why is this so complicated? Just tell me what's best
>>109605517I'm talking about civilized countries.
>>109604024There's no way google cares about our tiny threads
>>109605565Gemma has a gooner reputation on Reddit already.
>>109605578Isn't reddit partially owned by the CCP?
>>109603990>>109603996>>109604003What I originally posted about on 4chan was an ensemble of small models that are trained independently of each other and where the logits are then simply added up.That ensemble could be both trained and inferenced in a distributed way.However, thinking back to how ensembles of decision trees are used in practice, a better approach would probably be gradient boosting.Meaning that you would still train an ensemble of small models but in a linear way where later models correct the residual errors of all previous models.The inference at least could then be done using a distributed network where each participant only needs to run a subset of the ensemble.
>>109605598>Meaning that you would still train an ensemble of small models but in a linear way where later models correct the residual errors of all previous models.Even assuming that's true (and that's a big if) what's the benefit of it though? Ultimately, training is shaping the manifold, weight by weight, cycle by cycle. Each output is projected on the manifold space, again and again for every layer. What's the point of adding extra models to this? When doing the inference you still need to multiply all the weights and see how the projected embeddings fall on the manifold. You can't really cheat this process with extra 'swappable' models, which is why you have eye-wateringly expensive GPU racks doing the work and not a cluster of tiny models on hardware that's an order of magnitude cheaper.
>>109605598ensembles like this scale badly, bad idea
>>109603781where can i find these graphs for other models ?
Inside the Gemmaverse: Celebrating one billion Gemma downloadshttps://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads/
>>109605685>Empowering 100 million citizens in IndiaGuess they found /lmg/.
>>109605642The benefit would be that with an ensemble you would have a lower synchronization overhead vs. one big model.I think the architecture I'm talking about will only work under the assumption of a stochastic parrot.>>109605671I agree that an ensemble will probably not work well vs. one big model at equal FLOPS.But with an architecture like that you could possibly offset this via higher arithmetic intensity since each user in the network could feasibly work on a large number of requests in parallel.
>>109605685She's killed billions
>>109605685>https://deepmind.google/models/gemma/dolphingemma/>This specialized AI model processes complex dolphin vocalizations to predict sound sequences. Dolphin Sex????!!
>>109605685They also gave a github repo with a ton of Gemma-related links.https://x.com/googlegemma/status/2090484993579683904https://github.com/google-gemma/awesome-gemma
>>109605705If the sperm doesn't enter an egg there's nothing to kill.
>>109605725Cells are alive. It makes me wonder if we could use that instead of neurons for LLMs
>>109605700i once worked on a project related to ensembling. as part of this i tested various methods. ensembling works best when the models are dissimilar, such as CNN and ViT trained on different data. even then, you get at best a few % performance, and adding more models performance quickly falls off. if models are similar (same architecture, data etc) there is close to zero benefitfor llm i expect ensembling to be worse. you can ensemble n correlated models and get maybe 1-3% boost, likely worse than best of n sampling, or you can just cut the bullshit and generate a n times longer response with a model that is test time scaling capable and can handle the context
>>109605420You can use both you know, it's what I do.
>>109605735cooming comes after using the llm, if you have to do it before then there won't be any use for the llm afterwards
>>109605748The main goal is to enable distributed training and inference. I don't think anyone was expecting it to increase performance and as long as it's not much worse it would still be a win.
>>109605488People say the MoE isn't uncensored like the dense ones are. Been meaning to find out, I've only used 12B-chan and 31B-chan so far.
>>109605748(music) stem separation uses lots of ensembles, or at least did
>>109605769right
why google models always feel kinda garbage.it's like they give up in the middle of the task.
>>109605823I think it's a Gemini quirk for saving tokens.And Gemma was distilled from (a version of) Gemini, presumably.
>>109605101What do you think about just SMS mode for mobile main screen and full view for foldables? I think mobile would be a good place to build improvements to the SMS part.
I boughted more. I have 4x 5060Ti now.
>>109605864why 5060tis
>>109605908Best bang/buck in terms of VRAM. I have 64GB of VRAM now. I would have liked AMD for their FOSS drivers but their GPUs are buggy on a hardware level.
This wici company is selling an ai box that claims these numbers with this hardware. saying the following:Intel Core Ultra 7 255H AirCompute GPU sharing1 NVIDIA RTX 5090 · 32GB Pricing for this configuration to be announced closer to launch (5060ti box at 2k-ish)Wi-Fi 7 · 4×4 MIMO · 320MHz Router-grade AirLink Multi-Link transportNVMe 5 · 4TB TurboStream weight pagingKimi K3 2.8T 5 tok/sDeepSeek V4 Flash 284B 20 tok/sGemma 4 26B 500+ tok/sMeasured on WiCi One 5090 edition, close to cloud rates of 20-40 tok/sTurboStream turns 4TB of NVMe into working VRAM: trillion-parameter models, no cloud cost required. State-of-the-art research from our systems researchers and engineers, built into WiCi One.They also have an article that goes in depth about nvme streaming https://wici.ai/article/ssd-is-the-new-vramDoes this mean there is nvme performance being left on the table in local backends? I guess assuming pcie 5 stuff. Or can they reach these numbers and I just didn't know?
>>109605864How are you running them? When I looked into it it seemed like most consumer motherboards screw you on speed if they can even slot 4 cards.
>>109605923I'm not running them at all yet. I only have the GPUs, plus I bought a CPU today. I still need mobo, RAM, PSU, storage, etc.Currently I'm coping on a laptop with 4t/s.
>>109605923Oh, I'll be plugging them into the M2 slots using risers. If that doesn't work as I'm hoping I can just sell them for twice the price next year.
>>109605919>NVMe 5 · 4TB TurboStream weight paging>Kimi K3 2.8T 5 tok/s>slooooopI dunno...
>>109604124thanks your result really helped me out!
>>109605930>buying GPUs without even being able to run them yetHighly based
>>109605919>wici companylooks like a scam. why would you even mention llama 3.1 8b in 2026? https://wici.ai/technology also emdashes on every single page
>>109605955I kinda went full retard but I want to have an orgy with 31B, 12B and E4B at the same time.
>>109605955Doesn't need to. At this point, GPUs are an investment. Better than gold even.
>>109605923What you're referring to is the amount of lanes available.If you're not taking those lanes by utilizing multiple NVMe drives it shouldn't be any problem with cards of this type.And the bandwidth of those cards is only 448 GB/s, you can slap two of them into a pcie slot and be fine.Hell you could likely use 4 of them in the same pcie 5.0 slot and have it work without any kind of saturation.
>>109603609Reddit has a faggot reputation on 4chan already
>>109605994Gold holds value over millenia.GPUs will be worth more in the short term but will become worthless in the mid term.
>>109605987Me too anon, me too...
>>109606066What will they be worth in the long term?
>>109606160>>109606160>>109606160
There's clearly an influx of new retards across 4chan. Might be something to do with the recent gaming related leaks.
>>109606165Whatever intel 8086s are worth on eBay (not much).
>>109606178What leaks?
>>109605784gemma moe is braindead as fuck, people shilling it are just coping because all they have is a laptop from 2021
>>109606165One would hope modern cards will be worthless in a few years due to having far better ones readily available by then
>>109606194GTA6
>>109601570This is incredible anon.