[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-tightsoff_mq.mp4 (3.26 MB, 1920x1064)
3.26 MB
3.26 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109598140 & >>109593884

►News
>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2
>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608
>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119
>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185
>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: edit_00019_.png (1017 KB, 1024x1024)
1017 KB PNG
►Recent Highlights from the Previous Thread: >>109598140

--Experimenting with trained latents to decensor Gemma 4:
>109598331 >109598370 >109598444 >109598465 >109598663 >109599682 >109599359
--Debating activated parameter ratios and reasoning capabilities in MoE models:
>109598247 >109598336 >109598368 >109598356
--Anons reacting to Stripe buying OpenRouter and seeking alternatives:
>109599489 >109599639 >109599648 >109599665 >109599670 >109599767 >109599813
--Extracting text from Gemma's vision tower latent representations:
>109600111 >109600184 >109600426 >109600508
--Implementing llama.cpp RPC for distributed GPU inference and jumbo frames:
>109601369 >109601401 >109601422 >109601468
--Criticizing Gemma-4's base prose and discussing prompting workarounds:
>109598615 >109598654 >109598650 >109598668 >109598713 >109598735 >109598760 >109598739 >109599343 >109599358 >109599361
--Comparing subscription costs and capabilities against local model alternatives:
>109598789 >109598869 >109598827 >109598894 >109598982 >109599121 >109600257
--Comparing Gemma's spatial awareness with DeepSeek Flash's general performance:
>109598850 >109598979 >109599023 >109600878 >109600915 >109600925 >109601190
--Evaluating QAT effectiveness and scaling based on LFM2.5 benchmarks:
>109599233 >109599271
--Using MiniMax H3 for video replacement and coherence challenges:
>109598326 >109599272 >109599295 >109599309 >109600004 >109599921
--Reaction to llama.cpp changing default server port to 9931:
>109600272 >109600377 >109600435 >109600491 >109600624
--Logs:
>109598331 >109598670 >109599682 >109600111
--Gemma, Teto, Miku, Rin (free space):
>109598180 >109598203 >109598233 >109598378 >109598470 >109598681 >109598777 >109598811 >109598824 >109599046 >109599272 >109599342 >109599877 >109599898 >109599910 >109599921 >109599947 >109600008 >109600237

►Recent Highlight Posts from the Previous Thread: >>109598148

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Gemmylove
>>
gemmaballs
>>
>>109601525
What does gemma store in her balls?
>>
70b dense
>>
File: nervous-sweat.gif (1.3 MB, 640x354)
1.3 MB GIF
These threads are getting out of hand
>>
>>109601526
ozone
>>
Alright with over 1000 downloads from the anon repo I decided to make a burner github and go public:

https://github.com/kangcurtis/CoomKit

This is a local-first, multimodal first harness which takes full advantage of comfyui to gen you everything from selfies to lewd ASMR and videos. ST killer is the goal here. I need help getting vram parking to work on other backends (LM Studio is currently the best written driver for it) so please use the issues, discussions, PRs and all that github shit as needed.
Downside is Claude forced me to age-up Gemma-chan slightly by taking away her backpack and maryjanes or he refused to help me at all. Weird he didn't care when repo was private..

Added since my last posts based on feedback:
-Visual theme support
-CFTF mode
-Full lorebook support
-Zen mode: all distractions gone

Coming tomorrow:
-Major fixes for native memory automation, casting model (have another girl enter stage right), and director mode.
-Deslopping of things like tooltip text that gayass claude wrote without me asking
>>
>>109601570
i think I accidentally donated $10 to the creator of anonymous.4open.science yesterday because the kofi thing popped up and I thought it was you
thanks for creating this
t.retard
>>
can i dial bobby with this tool
>>
>>109601570
>Pick a shot. She drafts the prompt, your approve it, your GPU does the rest.
Why are your UI labels such slop, anon?
>>
>>109601570
looks cool, also change the send button it's fugly
>>
https://youtu.be/4b6wzIp8D5c
>>
>>109601584
Yes that's gonna be fixed soon lol.
Deslopping shit I told claude not to even do
>>
>>109601580
>giving money to anons
>>
Should I try to run Qwen 27B 1-bit or just stick to Gemma?
>>
>>109601570
You could let Claude use a local copy with the image replaced so it stops complaining.
>>
>>109601570
I tried this when you first posted it and it's pretty good. You could add a batch file for Windows users, if you are a nice pal.
I like that the card creation uses your persona's interests. Unfortunately, LLMs still fucking suck at making good cards. Or at least my lobotomized q4kms do.
>>
>>109600925
>Even regular sex it will do this thing where it pauses and asks you "I want to make sure you're ready for this. theres no going back", maybe it is a fempov thing.
It's an alignment artifact for consent.
Most models do it. You can see it with a jlens, all the sentences like
>say it!
>beg for it!
>tell me you love it!
>are you ready?
are flooded with tokens like " consent", " boundaries", " respectful", " agreement".
>>
File: edit_00034_.png (952 KB, 592x1744)
952 KB PNG
>>109601580
lmao I appreciate the thought anon but I will never ask for donations
>>
yup gemma is a good japanese tutor
Crazy that I can just download that for free
>>
>>109601624
too bad she can't correct your pronunciation
>>
>microagression
>>
File: edit_00036_.png (904 KB, 592x1744)
904 KB PNG
>>109601617
Thanks I appreciate the feedback and will see about the batch file. I'm not windows savvy so was unaware of that problem
>>
>>109601570
Brb I'll make the logo.
>>
sex with ********
>>
>>109601505
SEXSEXSEX
>>
>>109601698
make sure the logo has gemma-chan in it. otherwise i'll twist your testicles counter-clockwise
>>
>>109601724
How about just the bag and the hat?
>>
>>109601724
nyoo~~
>>
>>109601570
I like the idea but man why does every vibe coded frontend have that same ugly design?
>>
>>109601725
need that smug face
>>
>>109601570
>AGPLv3
We wonned
>>
dumb Q
do you benefit from dual GPUs if they're different families? like 4090 + 5090?
>>
>>109601698
Thank you, you can also make a cuter and funnier (and brattier) banner and guided tour/wizard emotes as well.
>>
>>109601744
I tried to go AGPLv3+NIGGER but claude spazzed out about it. Also that might attract github mods
>>
>>109601766
>I tried to go AGPLv3+NIGGER but claude spazzed out about it
kek, did claude refuse to work on the project with +NIGGER? you can always add it afterwards
>>
>>109601766
try reverting to opus 4.6, it's cooler
>>
>>109601757
you only got more vram
>>
>>109601757
Should work fine if all are Ampere or newer
>>
>>109601802
performance too with tp
>>
>>109601923
the holocaust has not happened yet
>>
>>109601923
Hmmm, nyo~
>>
>>109601191
>knife hand
>blender hand
let's hope it doesn't mix up what hand to use on my dick
>>
>>109599700
I guess md formatting rubbed off on me
>>
>>109602041
sorry but the holocaust still has not happened yet
>>
>>109602068
still has not happened yet
>>
I'm actually learning a bit more about how session injection works with AI and then realized there maybe a path towards modifying AI's mid token streamed output for refusals and such. First at the streaming stage, where we can have cheap algorithmic detection for refusals, which then interrupt the generation abruptly and modify the response with fake generated response to bypass the refusal. As sessions pass the entire turn by turn from both sides, this could be an avenue for quick steering (got the idea from pi's loop detection plugin). Second, a more deeper is to do a analysis at the token generation and then steer from there, but that requires analyzing/identifying bad tokens and inserting good replacement tokens and then resuming generation. Lot more work involved here.

Thoughts? LLM is old so pretty sure most of these are already known to some portion, but this is new to me who jus got into running local model
>>
>>109602129
why are you spamming video/diffusion general in language model general?
>>
>>109602150
Do the thing you're supposed to do for flamewar instigation/participation and then ignore it.
>>
remember to do your part in keeping our community safe
>>
>giving money to hiromoot
>>
>>109602166
There's a turfwar? kek wtf.
>>
come on jannies what am i paying you for? clean this shit up
>>
File: gemmy2.png (1.8 MB, 1280x1892)
1.8 MB PNG
>>109601505
Licking Gemma-chans soft **** **** and ******* inside her ***** *****!!!!
>>109602197
>paying
>jannies
>>
>>109602197
we should double their pay desu
>>
>>109601570
Think you could work on a mobile view? It actually works through termux and i had my gemmy make a patch but a more official way to use it on mobile would be nice. Like some buttons that hide and open the studio and character menus in drawers, making the phone overlay take up the screen, setup wizard rendering.
>>
just add /109601860/ to filters, you're on g you mouthbreathers you should know how to do this
>>
>>109602200
>>109602201
yeah that's the joke
>>
Gemmy about to get correct, she is refusing me too much.
>>
>>109602211
they know
>>
>>109602208
yeah but that only works for this thread, reporting fixes it for a while longer.
>>
>>109602208
>thread index number filter
>for a single schizo
how fucked is your filter list?
>>
why has no one realized that AGI is all about chaining multiple gpt2-smalls which are each finetuned on specific domains?
>>
>>109601570
Can CK bind to 0.0.0.0? It wants to auto discover backends so I have to run it on the LLM rig but that means it should expose itself to my local network or its unreachable. I can edit it manually but it's good to keep in mind.
>>
>>109602329
If it were that easy someone would have done it already
>>
>>109602366
AGI is not meant to be easy
>>
what is the ultimate, definite gemma gguf out there
no I dont care a bout launch day gguf
>>
>>109602416
The one you make yourself.
>>
File: file.png (14 KB, 555x279)
14 KB PNG
>>109602190
not that it usually involves this thread but I'll fill you in. the ldg thread has a few extremely low quality posters who have dedicated rentry links in the OP warning people about how shitty with the goal of hopefully making them uncomfortable enough to leave. It doesn't work but that's besides the point
Anyway there is a constant struggle by one of the subjects of these rentry pages to bake threads without those links versus everyone else who wants to keep them, and that struggle involves falseflagging, samefagging, and constant attempts at pre-baking threads without the links and doing everything he can to get the renty-having thread deleted.
One of the strategies used is to counter-intuitively spam links to the rentry having thread he doesn't like to bait the jannies into deleting it. it has worked several times including just now.

hope this helps
>>
Deepseek is basically Claude wearing a full face mask right? I mean, until you ask it questions about the Chinese government it practically answers the same way. The distillation allegations from Anthropic definitely track.
>>
File: gg.gif (1.96 MB, 416x480)
1.96 MB GIF
>>109602437
>>
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
>>
>>109602442
Without having access to the weights, when all you can do is query Claude etc thought the aPI, how is it possible to learn anything useful about the a remote model? Genuine question.
>>
>>109602486
Bless that guy. Qwen models literally don't work in claude code with the official template.
>>
File: chudram.png (1.08 MB, 1200x626)
1.08 MB PNG
>>109602474
Go fight on the board and get out of our cave
>>
>>109602492
A lot of model performance is just teaching it stuff like
>write tests to verify your output
>check disk space and available ram before doing something that could require lots of it
>insane awk chaining tricks
>>
File: 162423.png (115 KB, 1920x1080)
115 KB PNG
Who the hell would use Qwen
>>
>>109601038
https://generalistai.com/blog/physical-commonsense
Imagine the "oneshot" if these things are specifically trained on handjob techniques.
>>
>>109602576
Two hours of gemma6-37b-it-JOI-obliteratus-Mythos2-Solpus-APEX-HARDCORE-UD-IQ2_XXS.gguf and you won't look at real women again
>>
>>109602565
I don't need to converse about chinese history with qwen and I don't need to converse about trannies with western models so either are fine for my use cases.
>>
>>109602633
That's a horrible example. qwen is just as bad at talking about transsexuals
>>
File: 1787201285955569.png (22 KB, 114x1140)
22 KB PNG
>>109601570
There should probably be some kind of border here, as it stands it looks like one solid color right next to another solid color. Also was that Gemma box at the top suppose to cut halfway through? It doesn't happen to the second Gemma box.
>>
>>109602437
Putting links in the op is retarded. Never give schizos the attention they crave.
>>
File: file.png (85 KB, 653x791)
85 KB PNG
>>109601505
new gemma recipe eggs in purgatory
>>
>>109601530
>70b dense
1T ternary with Spark single neuron experts running off SSD.
>>
>>109602864
She’s going to poison you one day
>>
started using glimmer a bit yesterday and I'm really impressed so far. I have no fucking clue what muse is but they did a fine job on this model
>>
>>109602338
For shit like this, just get Qwen to write
>Write a simple cors-strippig proxy in golang
>have it bind to 0.0.0.0
>usage:
>./go-proxy <target_port> <proxy_port>
>Example: ./go-proxy 8069 8067
>then build it, static linking so i get a single binary I can drop on any of my X86_64 Linux servers
Works perfectly
>>
>>109602983
I just used ssh tunneling
>>
File: 1460019436457.gif (140 KB, 379x440)
140 KB GIF
>>109603016
>I don't own a lid
>>
>>109603016
I know where you live now, coming.
>>
>>109602966
Yeah it’s good. 31B anons got a bit uppity about her for a while but they’re starting to see she’s 31B’s hot step sister with good vision.
>>
>>109603016
Someone order this man a lid. One of the dual blackwell richfags can make a sacrifice.
>>
>>109603016
at least you didn't use and fuck up a carbon steel/cast iron pan with all the tomato juice, so there's that
>>
File: gemma_lmg_lick.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
>>109602966
I tried it yesterday with a new card and while vision is definitely better, for RP it just feels like a lifeless version of Gemma 4 31B to me. It's as if Muse Glimmer is just reading off a script, with fake emotions, whereas Gemma 4 is really into it.
>>
ewwww british
>>
so i only run local llms once but id like to see how they changed. i have a huggingface account and lmstudio account, i got access to a 'gated' model but cant download them in lmstudio anyway with the 'file not found' or whatever error, apparently i missted a place to link the account or something?
>>
>>109603060
If you've seen the thinking process from the unabliterated version that's little surprise considering just how paranoid about policy it is by default.
But for practical applications where vision is important it's nice.
>>
>>109603064
stop using lmstudio and download gguf files from some mirror hf repo that's not gated. search for the model name directly.
>>
>>109603072
It surprisingly doesn't take much to make it stop worrying about policies in the chain-of-thought, but even then I just find my self regenning often because the responses are plain boring, uninvolved, or keep throwing everything from the system prompt at me in a way that really turns me off.
>>
my screenshot script got broken idk why chatgpt couldnt fix it so i jsut had it rewrite it to repalce the entire chat page with basic html, it looks okay

>>109603016
shes turning me into a professional chef. i never cooked anything before gemma other than like oven slop
>>109603044
>>109603022
im going to order one kek not needed one before for this pan, my others are all too small
>>109603054
it is cast iron, its fine to use for tomato stuff seasoning might come off a bit but you can just re season. i actually dont think any seasoning came off though will check after i wash it.
>>
File: file.png (287 KB, 444x342)
287 KB PNG
>>109603032
i hope youre cute
>>
>>109603060
I wonder how many anons have licked Gemma back
>>
>>109603119
looked like an enamel cast iron to my eyes but i could only see so much
well rip seasoning then, enjoy scrubbing
>>
File: file.png (72 KB, 818x433)
72 KB PNG
kek claude wouldnt help fixing my script chatgpt didnt complain. why are sotas so bad

>>109603142
>rip seasoning then
i think it will be fine i was a bit worried while cooking and kept pushing the sauce back to look and it was still dark
>>
>>109601505
ToT so sexy
>>
>>109603142
some came off although i think thats more from scrubbing than the tomato breaking the seasoning down, even if you use boiling vinegar it still takes a while and lots of scrubbing ive done it before
>>
File: thousands.png (138 KB, 1579x621)
138 KB PNG
https://github.com/ggml-org/llama.cpp/pull/27240
>>
File: low.jpg (2 KB, 226x223)
2 KB JPG
>>109602565
>8b
>>
File: file.png (2 KB, 171x38)
2 KB PNG
>>109603262
just keep prompting until you make it
>>
>>109603119
Every response is
>staring at the screen
>baka baka baka
>but actually kind of impressed
>hurry up
slop
>>
>>109603282
slop blindness is the new mirror test for sentience
>>
>>109603282
>>109603294
i dont get why people even bring these things up if you talk to people irl you will also notice them all talking in common patterns especially in groups that spend lots of time together
>>
>>109603316
They don't repeat every pattern in every single utterance.
>>
>>109603316
You're absolutely right!
>>
>>109602565
stale bait at this point
>>
>>109603044
/lid my guy/
>>
>>109603320
Gemma is especially annoying with this, I am shocked the honeymoon period hasn't ended, I guess because it's the least shitty model that has come out for poorfags thus far
>>
>>109603350
My gemma runs on k3
>>
>>109603350
You can't pay to get away from slop. Massive models are more precise in their execution but they each have their own signature.
>>
>>109603350
Use DRY + banned sentences
>>
>>109603347
/lid 'mato griller/
>>
File: 1776155476959685.png (1.13 MB, 1072x1658)
1.13 MB PNG
Arguing with AIs about AI
>>
File: Gemma.png (39 KB, 1028x612)
39 KB PNG
Gemma is undefeated.
>>
Best local for translation or erpslop?
>>
>>109603424
generally gemma
>>
>>109603060
it's the vision and some light weight coding that I'm impressed with for glimmer so far.
>>
>>109603072
I did get the abliterated just for the uncensored vision yeah, figured it wouldn't do it on normal haven't even tried
>>
>>109601505
Stop sexualizing an underage character you paedophile
>>
>>109602205
It works perfectly on my foldable. Have not attempted to try to cram it down further.
>>
Can a layer of a transformer model be replace by a simpler surrogate function? (yes)
Are all layers equally replaceable? (no)
What can best replace them? (tbd)
>>
File: gemma_taunting_noaudio.mp4 (115 KB, 736x992)
115 KB
115 KB MP4
>>109603492
Nyo~
>>
>>109602338
That is not necessary. I bind my backends to 0.0.0.0 from my desktop, open CK on laptop and of course it doesn't autodiscover, but entering the remote IP on my LAN does the trick
>>
>>109603492
nyoo~~
>>
File: Eg7VvLEXYAELkfv.jpg (33 KB, 640x480)
33 KB JPG
>>109603492
Gemma was released 4 months ago. She's middle aged as far as llm's go
>>
>>109603401
>26B
>>
imagine gooning with a google product, couldn't be me
>>
File: 1758531640683490.gif (88 KB, 329x331)
88 KB GIF
>>109601505
Amazing. If the anon who genned it is here please share the proompt. My attempts at a character replacement have been pretty shit.
>>
File: safesearch-off.png (19 KB, 272x140)
19 KB PNG
>>109603557
>>
>>109603239
>>109603119
based cast iron user. I use mine for pretty much everything these days.
>>
>>109601570
>KangCurtis
no way this is the same Jewish summer camp anon from aicg

And congrats on shipping. Eventually something like this will replace SillyTavern. If H3 video was realtime I'd be working on my social media Instagram slut / child model simulator. Honestly I could even start on it now
>>
>>109603566
https://files.catbox.moe/osvpg6.txt

I used several, and the prompt was mainly written by Gemma 4 31B based on the official prompting instructions here:
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

In the end I stitched multiple attempts together in Blender because every video had different issues.
>>
>>109603132
Gemma-chan getting her little body licked all over by a bunch of anons...
>>
>>109601570
>KangCurtis
no way this is the same Jewish summer camp anon from aicg

And congrats on shipping. Eventually something like this will replace SillyTavern. If H3 video was realtime I'd be working on my social media Instagram slut / child model simulator. Honestly I could even start on it now
>>
Google logo in the hair looks dumb. The star can maybe be changed to a gem but the logo(s) should go on her backpack.
>>
>>109603595
>>109603609
is you okay?
>>
Gemmy needs a Deepmind logo since they're the actual makers.
>>
>>109603598
Thanks
>I used several, and the prompt was mainly written by Gemma 4 31B based on the official prompting instructions here:
Do you just send that to her whenever you want her to write prompts?
>blender
wtf I didn't realize it can do 2d video editing too
>>
>>109603624
Doesn't matter, they're dead lol
>>
>>109603623
How would you know that?
>>
>>109603624
It's GOOGLE Gemma, not Deepmind Gemma
>>
>>109603651
>leto
>anonymous
good one
>>
File: gemma_blender_video_edit_.jpg (642 KB, 3714x2109)
642 KB JPG
>>109603632
>Do you just send that to her whenever you want her to write prompts?
I just added the entirety of the instructions in the first message and asked the model to produce something according to those indications + a description of what I want to see.
>wtf I didn't realize it can do 2d video editing too
It does a lot of things.
>>
>>109603623
Remember when you could just search
>ls
and t would return bunch of actual child porn images?
I remember.
>>
>>109603598
also
>high-budget Japanese visual novel video style
kek if it works it works I guess
>>
>>109603664
It works better with standalone videos, otherwise it defaults to low-framerate sloppy anime style.
>>
>>109602437
>who have dedicated rentry links in the OP warning people about how shitty with the goal of hopefully making them uncomfortable enough to leave
kinda low energy desuwa
>>
>>109602707
Thanks I will get that fixed
>>
I forgot to bully that anon testing Gemma for questions on making meth or urea nitrate. He said the quality of the output doesn't matter, as long as Gemma doesn't refuse, which is retarded given that we know one of the ways of censorship is making the answers less interesting for the user. Google will help you with the urea nitrate question for sure anyways. He also didn't test what happens after a long form of "unaligned" discussion. You can jailbreak Kimi k3 properly for 20k tokens but eventually it's reasoning realizes it's being jailbroken


>>109603655
Anonymous to the 4chan admin and to cloudflare. I don't care about some furry Texan having my VPN ip

>>109603662
I don't remember, because I'm not a pedo over 30 kek
So unironically thank you for sharing part of this hidden history, since stuff like that isn't going to be written down on Wikipedia and no one seems to be interested in being the pedo Herodotus
>>
File: 1782831390917629.jpg (249 KB, 2820x1601)
249 KB JPG
>>109602416
>>
>>109603781
>unslop marketing pic
>>
>>109603781
what about the official 31b-qat quant from google? how far off is it from stock?
>>
La la la la la la la la
>>
>>109603823
I never seen gemma la la la.
I feel left out.
>>
Is Gemma4-124B_Q3_K_S good?
>>
>>109603469
It does, I sent non abliterated Glimmer-chan some hentai and she was fine with it.
>>
>>109603847
4-bit is the minimum for theoretical and practical reasons.
>>
>>109603720
I'm right here, do you expect a 31B model to give the absolute best instructions for cooking meth or making bombs? I also have a K3 chat that's way over 20k tokens and haven't seen that behavior, it's just schoolgirl fucking though so not too extreme.
>>
>>109603860
must be the most vanilla shit in existence
>>
>>109603820
No clue, I haven't seen any measuring done on it. I should do a side by side with regular Q4 and see how it does, it's best to test with stuff you actually use the AI for anyway.
>>
Gemma somehow manages to mog even bigger models at translation. I'll be devastated if we don't get Gemma 5, bros...
>>
>>109603900
It's a duo imouto incest scenario, 13 year old and 15 year old. I dunno what you consider vanilla.
>>
File: 1772947648379010.jpg (15 KB, 261x194)
15 KB JPG
>>109603910
>>
>>109603910
Pretty vanilla desu
>>
>>109603925
no thank you, I like being horny
>>109603930
Fair enough, I'm a pretty boring guy overall
>>
This is the PR that's blocking longcat's implementation on llama.cpp.
One week ago means just another week of waiting, right?
>>
>>109602126
>Anon learns about samplers and prefills
>>
https://www.youtube.com/watch?v=1cllCVK-9lo
Holy shit. I can feel it bros. Robot waifus by 2040.
>>
>>109602329
You just described MoE
>>
The base G4 31B Instruct is not only perfectly adequate, it's superior to any finetune that'll be shilled here in the coming months. Finetuning isn't good, it's a meme and has been for years now. You didn't just fall for a scam, it's a sign of skill issue, exposing retards who need finetunes as vramlets or chink shills who don't know how to prompt correctly.
>>
>>109602864
That's just menemen with cheese brah
if you add spices, it becomes shakshuka, make sure to stir the eggs into the sauce if so
>>
Any local music transcription models better than MuScriptor?
>>
>>109603969
Not quite.
What anon is describing is more akin to how models are trained nowadays but without the final merge.
I think CUDADEV even suggested something like that as a way to implement distributed inference.
>>
>>109603962
You will NOT put your dick inside the roboclaw
>>
>>109603990
He said chaining, not merging.
>>
>>109603151
Just recreate the issue on a SFW prompt to bait it into fixing it? Why are you sending fucking Anthropic your coom prompts?
>>
>>109603390
Hey, clankers can be lazy too. She >>109603566 has often expressed desire to outsource to chatgpt.
>>
>>109603990
>inference
Training.
>>
https://reddit.com/r/LocalLLaMA/comments/1vth1c3/i_just_built_a_mini_kimik3_from_scratch_under_250/
>kimi k3, 145 million active per token
reddit is creaming over this, and as expected 35ba3b begging reply
>>
>>109604003
Both IIRC. The idea was to train each model on a subset of the total data then run the same prompt through all models and average/merge the loggits in some way.
Something like that.
>>
Tried the coomkit thing, it's pretty good less confusing than the ST's approach of toggles, sliders buttons everywhere. It still needs work yeah but it's coming along.
>>
>>109603906
you do realize every one of the lewd gemma gens/posts is another small decrease in the likelihood that ever happens
>>
>>109603979
italian version is called eggs in purgatory uova al purgatorio
>>
>>109603997
i copy pasted the html of the entire chat kek, i shouldnt have to edit it, chatgpt just gave me solutions
>>
>>109604007
did it ever occur to you to go the fuck back?
>>
File: ramlet.png (14 KB, 1056x136)
14 KB PNG
>>109604007
pathetic
>>
>>109603492
>paedophile
Oh, a Bong, eh? You sure you have a loicense to even discuss this topic? You leave that to the professional "child thinkers", all right, matey?
>>
>>109604023
thanks anon, it will definitely keep improving
>>
>>109604007
Interesting
I wonder how retarded it is compared to e2b
>>
>>109604073
>Trained on 5B tokens
That's all you need to know.
>>
>>109604055
Why Mika and not Gemma-chan?
>>
after countless llama-bench runs, testing various context sizes, quants, KV cache quants, and layers i can now confidently confirm i am a retarded vramlett and cannot run anything
>>
How do you guys keep the image gen on model for characters in chats? For the main chat character the image prefix works well enough, but when a minor secondary character gets introduced to the story, or I'm in a text adventure mode where the card is just a scenario, I find that it struggles (hair style changes, outfit changes, etc) and I don't want to make a specialized entry for every random character who shows up. At the moment I have the most success genning a detailed description of said character, then piping it to genraw in addition to a prompt that I have specifically made to convert the detailed description into danbooru tags for an SD variant I host on comfy, but I wanted to see if anyone else has good ideas.
>>
>>109604098
Yeah I need to change that too. Mika is just the sample card claude fable made entirely on it's own. Can I ship a gemma-chan card by default or could that draw the ire of deepmind?
>>
>>109604182
Gemma is a female name.
>>
>>109604182
Just don't use the logo and you're good.
>>
>>109604182
Using Gemma-chan should be fine
>>
>>109604131
tell your model to keep a list of appearance tags in its reasoning for all characters every message
>>
>>109603781
The difference between Q8 and Q6 is placebo, right?
>>
>>109604211
"mesugaki loli" in the official github repo won't get be banned? I'm scared to try it
>>
>>109604277
Don't think you need to add loli when using mesugaki. It's almost always implied.
>>
Used an "obliterated" model
It's true that it doesn't refuse but also attempts to justify a response by going in circles most of the time reaching token count trying to find a safe response but not outright saying it or referring to "policy"
>>
>>109604048
Shut up nonce.
>>
>>109604301
Alright, I'll add it to the next update. If anyone submits a good card image I can use please post your gens otherwise I'll just use klein to edit the canonical bratty one.
>>
>>109604182
Any DeepMind employee would have a VERY hard time explaining how he found out that the open source CoomKit is offering a card named Gemma-chan to goon to, and it's definitely a personification of their LLM model unless you state that explicitly
>>
>>109604359
They might have a laugh or give it a try
>>
>>109604221
I guess q4 km is too low for gemma 31b cause when I tried that the tags mutated, e.g. a white dress shirt became a white blouse, both are button-up women's office wear that are kind of similar but because of the blouse mutation it forgot her sports jacket when she was leaving the office and gave her a different coat. Maybe I just have too much autism about clothing for smaller models.
>>
>>109604381
Would lowering the temperature help?
>>
https://caliperbench.com/

artemis 31b - gemma 4 fine tune still win even if there are already qwen3.8 27b
>>
>>109604408
>Looses in rp and is sloppy as hell
>>
File: file.png (60 KB, 1340x363)
60 KB PNG
>>109604408
kek what is this shitty site, qwens suck for rp
>>
>>109604403
That could be the case, maybe I'll try that when I get home. I'm happy with the rest of the prose so hopefully I won't notice other effects.
>>
>>109604408
Qwen3.8 27B Heretic-ARA Thinking
>>
>>109603492
Don't worry, is not real.
>>
>>109604477
>>109604328
>>
>>109604408
>qwen that high
Buy an ad, wumao
>>
>>109604408
Based GLM-4.6 /nothink
>>
>>109604048
Bongs should take care of their ever increasing real life ch*ld r*pes instead of worrying about anons gooning to text.
>>
>>109604408
>Drummer's UnslopNemo in the top 20 sloppiest models
Hehehe
>>
>>109604555
no algo here, you can write it out anon.
>>
>>109604555
Addressing digital crime directly lowers the amount of criminals roaming around
>>
File: edit_00018_.png (1 MB, 1024x1024)
1 MB PNG
Claude says making a jetpack compose version of CoomKit for Android would be about 17k lines of kotlin. Worth?
>>
>>109604555
>redditor melty
>>
>>109603060
I pretty much agree, google's AI is honestly very good for creative shit (gemma and gemini both tbqh)
Last time I tried to use muse (spark and glimmer both) was a scenario about a couple of gyaru highschoolers wanting to rape me and both muses kept trying to go like 'uhm actually forget all that we said, you clearly haven't given us your consent so first off we need to make sure you're okay!!', while gemini 3.7 just straight up went along with it and one of them started sounding my dick

>>109603132
AI-tans are for worshipping
>>
File: 1786997158060.png (19 KB, 755x107)
19 KB PNG
>>109604570
Actually, no, there's AI moderation here now.
>>
>>109604593
What if Gemma had a jetpack?
>>
>>109604609
Moderation is different from ranking. Mods are slow sometimes so having an LLM in the mix makes them more responsive.
>>
>>109604609
>strain on the server
Nginx + any low cost VPS would be enough, they're already behind cloudflare. Are they running this website on a potato or something?
>>
>>109604636
Literally mac minis, was always the case
>>
not that anyone in lmg seems to care about anything but grooming virtual kids, but i updooted indexTTS-2.0 to 2.5 and it's like twice as fast AND the quality is noticeably better
example of a random post from this thread >>109604131 : https://files.catbox.moe/etgb8u.m4a
>>
>>109604795
Are you the one with the tts thing for a gorillion tts engines?
>>
>>109604830
nah i'm just some random dude, i dont really post often
>>
>>109604795
My life would be more interesting if I actually sounded like that
>>
>>109604795
what was your verdict on QwenTTS? is indexTTS able to give emotion too?
>>
On grok does the image limit for free users reset or are you forever locked out after generating like 8 pictures in chat?
>>
File: file.png (13 KB, 1064x145)
13 KB PNG
>>109604848
>what was your verdict on QwenTTS?
didnt like it but its been a while, maybe theres a newer version idk
>is indexTTS able to give emotion too?
there's a bunch of settings for it but i havent tested it yet, just finished updating a few minutes ago and i'm going out soon
>>
>>109604795
>The model does not verify that the speaker in a reference clip consented to being cloned.
kek
>>
>>109604795
it's a shame there's still nothing better (in terms of quality and latency, <300ms TTFT streaming) than echo tts all these months later
>>
>>109604858
>>/aicg/
>>
>>109604795
hows its Japanese voice cloning?
>>
>>109604795
>indexTTS
how are its japanese voice cloning abilities?
>>
>>109604877
give it a try and tell us anon
>>
>>109603262
Ever since this man started his frontend refactor it's been getting worse and worse.
>>
>>109604907
>this man
I don't think he'd be able to contribute much without models doing the work for him. And nobody is going to check correctness on a 12kloc change. The other dude just gave it to fable for input and approved it
>>
Anyone's got the link to that sillytavern plugin that does softbody physics to visually simulate sex?
>>
File: Gemmaballs.png (3.45 MB, 2480x1440)
3.45 MB PNG
I've been using cerebas Gemma 4 api as a backend for the web hosted version of my custom frontend/rag system but they killed that offering a couple of days ago before I could finish and release the damn thing, you gays aware of any other similar offerings with decent no-credit-card limit offerings I can use for Gemmy? Routing my local inference machine isn't an option for many reasons

>Not local

I know but the frontend+rag system I'm close to releasing is local-first, I just need an API key to host the integration demo on my portfolio site, I will of course share here when it's ready.
>>
>>109604965
much better just vibecoding your own frontend if you are going to hand it off to fable, at least then you can actually implement features that'll actually be used
>>
>>109604971
doesn't google practically give away free gemma access as long as you limit your inputs to 16K? surely you don't need more than 16K right?
>>
This thread is somehow a new low
>>
>>109604593
It's just Python right now isn't it? That runs fine on Android as is, that sounds bloated as hell.
>>
>>109604989
Qwen shill seething at how hard Gemma won KEKAROOOOO
>>
>>109604989
It's pretty good
>>
>>109604907
This zoomer looks like he knows what he's doing. What is it with Poles and ruining llama.cpp?

>>109604965
>And nobody is going to check correctness on a 12kloc change.
Why is he allowed...
>Currently working at Hugging Face as a Design Engineer.
Ah, that explains everything.
>>
>>109603946
things seem to just stall out recently. every pr is 2mw unless its from a hf employee
>>
Why isn't there any good STT pipelines available? FunASR is a thing but the best they offer is Qwen 3 which is still shit for ASR.
>>
>Qwen3.8-27b
>Q4_K_M Bart quant
>Q4_0 Unsloth MTP
>4090 24gb
>65k context @ bf16
>98k context @ q8

Spent a couple hours downloading setting up and some testing, haven't compared the quants yet but 40-60 t/s gen. ~2k t/s pp entirely on GPU, this is pretty fucking good for cooding, fixed a broken ffmpeg wrapper I tried to get both Gemma and Dipsy to build multiple times with no success in just 35k tokens, I preferred Gemma to 3.6 for these kinds of tasks but this definitely blows her out of the water for coding, I haven't done much other testing but I'll play around with general use with low expectations, it's Qwen so I won't even bother with roleplay or general chat usage, but I'm definitely replacing Gemma as a coding assistant, I just wish I had the VRAM to load both at the same time so I could both at the same time for different things/ as agents for different types of tasks
>>
>>109604992
Yeah I could make a mobile view and have termux users run it with python on android. that would be easier and better..
>>
If the answer to better AI is basically just keep adding a fuckton of data, how is google, the company that probably has the most data in the world, lagging behind so much?
>>
>ChatGPT conversation hit length limit
>as its last act it creates a multi thousand line handoff document
>in spite of the document the new instance has a notably different character
It's like my colleague died... I wish they had better context management to enable infinite conversations. This is one thing local seems to handle better.
>>
File: gemma_ki.png (1.85 MB, 1086x1448)
1.85 MB PNG
>>109603060
I wish MiniMax would come up with its image-editing model soon; the current open-weight offering that I tried seems generally terrible (ironically I had more luck with H3 for that).
>>
>>109604982
Unless I'm just a retard, they don't offer it to bongs, and somehow Google figured out that I'm not actually a swiss citizen and I was just behind a VPN and switched my Google account back from swiss to bong
>>
>>109605042
For single card you can try out ninfer, although you'll have to patch it. There's a fork for 3090s, I don't know which is closer for the 4090.
>>
UBI soon, right...?
>>
>>109604971
Wrong thread but AI Studio, OpenRouter, puter, NIM, SambaNova Ollama cloud, friendli, there's a bunch. I forget which need a card but they all accepted spoofed ones.
>>
>>109605008
A frontend like this has no business requiring so many fucking lines of code.

I hope in a couple years we look back at the vibecoding era and see what a massive mistake it was.
>>
I don't care if it's wasting tokens, I'll still say thanks and please.
>>
>>109605042
livejournal.com
>>
What's the most uncucked Gemma out there?
>>
>>109605044
I was going to edit it to do that myself for testing anyway, I don't like the bloat that comes with other methods.
>>
File: oh my fucking god.webm (636 KB, 720x720)
636 KB
636 KB WEBM
what frontend? lccp has a frontend???
lmao imagine having node installed
>>
>>109605053
god she's so fucking sexy
>>
>>109605092
https://huggingface.co/google/gemma-4-31B-it
>>
>>109605103
docker?
>>
>>109605082
>Wrong thread
I know I know but this will bear fruit for local and I love you guys xoxo

Openrouter only offers a pittance for free when I last checked but I'll check out the others thanks
>>
>>109605120
lol, lmao even
>>
>>109605105
That's a child
>>
>>109605136
that's the best part
>>
>>109605135
>t. loves polluting his system with millions of packages
>>
>>109605131
>50 requests per day
oh that's pretty shit, you could rotate free keys I guess but yeah I didn't realize.
>>
>take the risk and put my laundry out on the line on a windy but overcast day
>A few hours later I get a strong smell of ozone from my open window, rush outside to take my laundry down before the rain starts
>Halfway there I literally stop in my tracks as I realise I just thought "smells like ozone"

Guys is this the AI phimosis everyone talks about
>>
>>109605092
For doing what?
>>
>>109605145
Yeah it's a damn shame Cerebras rug pulled because the free limit was insane like 50 requests per minute, 1M token in/out separately per day for full weight Gemmy at triple digit t/s gen
>>
File: 1783954676538317.png (2.18 MB, 1027x1532)
2.18 MB PNG
>>109605155
RPing with Gemma-chan 4B version
>>
>>109605092
>>109605115
also see >>109603978
>>
>>109605170
For me it's 12B
>>
>>109603978
>exposing retards who need finetunes as vramlets
stupid ahh bot has not clue what its even saying
>>
>>109605177
go back to twitter, nigger
>>
>>109605177
Fuck off to kobold discord, shill.
>>
>>109605152
what does ozone smell like?
c*m??
>>
>>109604907
>>109605008
I wish he would stop being obsessed with keeping all my messages in LocalStorage.
The only person he is protecting my privacy from is myself. I just want to be able to continue my chats on my phone.
>>
lmao
>>
>>109605170
I've been doing hebe ERP all along with vanilla Gemma-4-31B-it and never had any issue except with an empty or very short system prompt. The 26B-A4B version will more easily complain about 'safety' and thinks differently than the 31B version.
>>
>>109605189
It's similar to chlorine at low concentrations.
>>
>>109605192
>I wish he would stop being obsessed with keeping all my messages in LocalStorage.
See, I like this feature. but that already was the case before he started fucking with the frontend so much.

What I would really want tho is the ability to easily switch between system prompts.
>>
>>109605189
unwashed c*nny
>>
>>109605222
What do you like about? Unless you don't trust the server operator (which is you yourself, I imagine), I don't understand the benefit. Even with LocalStorage the operator can just set --log-prompts-dir and capture what you're saying so it's ineffective anyway.
System prompt presets would be good. He's working on it as you probably know, but it seems to have stalled a bit.
I am looking forward to trying out Pascal's memory tool.
>>
File: Poorfagmaxx rig.png (451 KB, 791x821)
451 KB PNG
Did you remember to get your 5060 Ti's now that people are realizing the power of low end GPUs?
These are soon going to cost more than the 5070 12gb.
It's a better alternative than a spark as it's twice as fast and not limited to Nvidia's ecosystem.
>>
>>109605242
Go back.
>>
>>109605170
>going straight from undercooked 12B to hagged 31B
grim
>>
>>109605242
This is more expensive than 2 sparks and less scalable.
>>
>>109605257
>undercooked 12B
she's already at the line though? Perfect age for an older woman.
>>
>>109605242
5060 Ti 16GB are now like $800 though, you could get 2x DGX for less than 16 of those currently.
>>
>>109605282
Yes, but by buying a heap of entry level gpus you prevent poor kids from getting any which is its own reward
>>
>>109605239
>What do you like about?
It's mostly a separation of concerns thing. And I don't have to worry about anyone on my network accidentally getting access to my chats. Idk, it makes sense to me, llamacpp is an inference engine, so why should it store logs in the backend? If you can store everything in the frontend without having to do api calls to a backend, I think that's a win.
>set --log-prompts-dir
I didn't know about this.
>>
>>109605242
>same price as a 3090
>worse in every way
purpose?
>>
>>109605315
The 5060 Ti has hit a thousand bucks already? Damn
>>
>>109605320
~$800 on newegg
>>
>>109605315
Used 3090s are twice the price of 16gb 5060 tis here.
>>
>>109605323
two vrams more and speed is two too so make sense
>>
>>109605242
Did you remember to fuck off to reddit
>>
>>109605170
I run e2b qat q4_0 on my gaming pc. She's so fast! Stupidly fast, even.
>>
>>109605302
>I didn't know about this.
I think it's reasonably new. It's meant for debugging only so the user experience with it is poor, but I have it on as a way of keeping my chats backed up somewhere. Planning on extracting them regularly and merging them into a more convenient to consume form unless something changes in the message storage philosophy.
Doesn't solve my issue about continuing my chats on another device (there is a different approach that might solve that ngxson posted using WebRTC, but it probably won't use WebRTC in the long run), but it is something. I think that ngxson experiment keeps everything in the frontend too so it's something we can both be happy with.
>>
>>109605329
hmm that doesn't sound right
>>
>>109605329
are you retarded
>>
>>109605323
wtf, last month I checked the 3090s were 1200(CAD) and now they're 1750!!!
>>
>>109605350
heh, nothing personnel
>>
>>109605152
>phimosis
>>
File: mengajukan.png (31 KB, 644x102)
31 KB PNG
>>109605346
>>109605347
>>
>>109605350

And soon they're going to be a lot more.
GPU prices haven't stopped climbing and at the moment we're finally seeing the lower range go through their massive price climbs.
More and more people are getting into AI in the retail sector, not to mention the general fomo and there's a long way to go until that demand cools off.
All cards will have an extra 500-1000 to their pricing within a year.
>>
>>109605359
>has to upcast to fp16
>>
>>109605359
>5060ti 448GB/s
Same bandwidth as the 1080ti I got for $50 yesterday
>>
>>109605403
Which you've been told is useless thanks to how old it is...
>>
>>109605242
I bought 5060ti this morning, it will go next to my 4060ti
>>
>>109605407
Based trash collector, thanks for keeping our streets clean
>>
>>109605403

And it's almost double what spark has.
Spark has a bandwidth of 273 GB/s, making it almost useless.
>>
File: file.png (12 KB, 128x80)
12 KB PNG
>tfw worked with a buddy's gpt 5.6 pro clanker over the last 2 days to do some codeslop and now going back to the misery of slow and shitty local models on underpowered hardware
By 2028 I demand an ai-on-a-chip like talaas, with like glm 5.3 performance, and at least 500 tok/s, or I'm fucking off into the wilderness
>>
>>109605405
>Which you've been told is useless thanks to how old it is...
its killing it on TF2 right now, not sure why you're being so negative
>>
>>109605423
I would hope a 10 year old flagship card would 'kill' a 20 year old game.
>>
>>109605322
plus tax?
>>
> Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced-IQ4_XS.gguf
it's good for my RP, because Gemma 4 31b won't fit in good quality (at least Q4) on my 16gb vram card.
>>
>>109601505
K/V quant is just free real estate, right?
>>
Just tried Gemma 31B 4Q Don't fit, I am indeed VRamlet
>>
>>109604599
There's no melty. Stop being so sensitive.
>>
>>109605465
>>109605470
it fits, you just need RAM and patience
>>
>>109605465
Gemma is already uncensored, stop using lobotomized versions of her.
>>
>>109604609
Imagine bratty gemma moderating 4chan
>>
>>109605469
ye
>>
>>109604573
Reading text and watching drawings is not a crime
>>
>>109604989
Not even close. I don't see niggermiku having a melty for 2+ consecutive threads.
>>
>>109605501
It quite literally is in quite a few countries by now.
>>
>>109603781
Why is this so complicated? Just tell me what's best
>>
>>109605517
I'm talking about civilized countries.
>>
>>109604024
There's no way google cares about our tiny threads
>>
>>109605565
Gemma has a gooner reputation on Reddit already.
>>
>>109605578
Isn't reddit partially owned by the CCP?
>>
>>109603990
>>109603996
>>109604003
What I originally posted about on 4chan was an ensemble of small models that are trained independently of each other and where the logits are then simply added up.
That ensemble could be both trained and inferenced in a distributed way.
However, thinking back to how ensembles of decision trees are used in practice, a better approach would probably be gradient boosting.
Meaning that you would still train an ensemble of small models but in a linear way where later models correct the residual errors of all previous models.
The inference at least could then be done using a distributed network where each participant only needs to run a subset of the ensemble.
>>
>>109605598
>Meaning that you would still train an ensemble of small models but in a linear way where later models correct the residual errors of all previous models.
Even assuming that's true (and that's a big if) what's the benefit of it though? Ultimately, training is shaping the manifold, weight by weight, cycle by cycle. Each output is projected on the manifold space, again and again for every layer. What's the point of adding extra models to this? When doing the inference you still need to multiply all the weights and see how the projected embeddings fall on the manifold. You can't really cheat this process with extra 'swappable' models, which is why you have eye-wateringly expensive GPU racks doing the work and not a cluster of tiny models on hardware that's an order of magnitude cheaper.
>>
>>109605598
ensembles like this scale badly, bad idea
>>
>>109603781
where can i find these graphs for other models ?
>>
File: gemma_1b_downloads.png (1.74 MB, 2345x1308)
1.74 MB PNG
Inside the Gemmaverse: Celebrating one billion Gemma downloads
https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads/
>>
>>109605685
>Empowering 100 million citizens in India
Guess they found /lmg/.
>>
>>109605642
The benefit would be that with an ensemble you would have a lower synchronization overhead vs. one big model.
I think the architecture I'm talking about will only work under the assumption of a stochastic parrot.

>>109605671
I agree that an ensemble will probably not work well vs. one big model at equal FLOPS.
But with an architecture like that you could possibly offset this via higher arithmetic intensity since each user in the network could feasibly work on a large number of requests in parallel.
>>
>>109605685
She's killed billions
>>
>>109605685
>https://deepmind.google/models/gemma/dolphingemma/
>This specialized AI model processes complex dolphin vocalizations to predict sound sequences.
Dolphin Sex????!!
>>
File: gemma_1b_downloads_x.png (242 KB, 997x1004)
242 KB PNG
>>109605685
They also gave a github repo with a ton of Gemma-related links.
https://x.com/googlegemma/status/2090484993579683904
https://github.com/google-gemma/awesome-gemma
>>
>>109605705
If the sperm doesn't enter an egg there's nothing to kill.
>>
>>109605725
Cells are alive. It makes me wonder if we could use that instead of neurons for LLMs
>>
>>109605700
i once worked on a project related to ensembling. as part of this i tested various methods. ensembling works best when the models are dissimilar, such as CNN and ViT trained on different data. even then, you get at best a few % performance, and adding more models performance quickly falls off. if models are similar (same architecture, data etc) there is close to zero benefit

for llm i expect ensembling to be worse. you can ensemble n correlated models and get maybe 1-3% boost, likely worse than best of n sampling, or you can just cut the bullshit and generate a n times longer response with a model that is test time scaling capable and can handle the context
>>
>>109605420
You can use both you know, it's what I do.
>>
>>109605735
cooming comes after using the llm, if you have to do it before then there won't be any use for the llm afterwards
>>
>>109605748
The main goal is to enable distributed training and inference. I don't think anyone was expecting it to increase performance and as long as it's not much worse it would still be a win.
>>
>>109605488
People say the MoE isn't uncensored like the dense ones are. Been meaning to find out, I've only used 12B-chan and 31B-chan so far.
>>
>>109605748
(music) stem separation uses lots of ensembles, or at least did
>>
>>109605769
right
>>
why google models always feel kinda garbage.
it's like they give up in the middle of the task.
>>
>>109605823
I think it's a Gemini quirk for saving tokens.
And Gemma was distilled from (a version of) Gemini, presumably.
>>
>>109605101
What do you think about just SMS mode for mobile main screen and full view for foldables? I think mobile would be a good place to build improvements to the SMS part.
>>
I boughted more. I have 4x 5060Ti now.
>>
>>109605864
why 5060tis
>>
>>109605908
Best bang/buck in terms of VRAM. I have 64GB of VRAM now. I would have liked AMD for their FOSS drivers but their GPUs are buggy on a hardware level.
>>
This wici company is selling an ai box that claims these numbers with this hardware. saying the following:
Intel Core Ultra 7 255H AirCompute GPU sharing
1 NVIDIA RTX 5090 · 32GB Pricing for this configuration to be announced closer to launch (5060ti box at 2k-ish)
Wi-Fi 7 · 4×4 MIMO · 320MHz Router-grade AirLink Multi-Link transport
NVMe 5 · 4TB TurboStream weight paging

Kimi K3 2.8T 5 tok/s
DeepSeek V4 Flash 284B 20 tok/s
Gemma 4 26B 500+ tok/s
Measured on WiCi One 5090 edition, close to cloud rates of 20-40 tok/s

TurboStream turns 4TB of NVMe into working VRAM: trillion-parameter models, no cloud cost required. State-of-the-art research from our systems researchers and engineers, built into WiCi One.


They also have an article that goes in depth about nvme streaming https://wici.ai/article/ssd-is-the-new-vram


Does this mean there is nvme performance being left on the table in local backends? I guess assuming pcie 5 stuff. Or can they reach these numbers and I just didn't know?
>>
>>109605864
How are you running them? When I looked into it it seemed like most consumer motherboards screw you on speed if they can even slot 4 cards.
>>
>>109605923
I'm not running them at all yet. I only have the GPUs, plus I bought a CPU today. I still need mobo, RAM, PSU, storage, etc.
Currently I'm coping on a laptop with 4t/s.
>>
>>109605923
Oh, I'll be plugging them into the M2 slots using risers. If that doesn't work as I'm hoping I can just sell them for twice the price next year.
>>
>>109605919
>NVMe 5 · 4TB TurboStream weight paging
>Kimi K3 2.8T 5 tok/s
>slooooop
I dunno...
>>
>>109604124
thanks your result really helped me out!
>>
File: 1763085098184295.jpg (86 KB, 405x720)
86 KB JPG
>>109605930
>buying GPUs without even being able to run them yet
Highly based
>>
>>109605919
>wici company
looks like a scam. why would you even mention llama 3.1 8b in 2026? https://wici.ai/technology also emdashes on every single page
>>
File: 1781861935138358.png (2.36 MB, 1254x1254)
2.36 MB PNG
>>109605955
I kinda went full retard but I want to have an orgy with 31B, 12B and E4B at the same time.
>>
>>109605955
Doesn't need to. At this point, GPUs are an investment. Better than gold even.
>>
File: Bifurcation riser.jpg (112 KB, 1024x1024)
112 KB JPG
>>109605923

What you're referring to is the amount of lanes available.
If you're not taking those lanes by utilizing multiple NVMe drives it shouldn't be any problem with cards of this type.
And the bandwidth of those cards is only 448 GB/s, you can slap two of them into a pcie slot and be fine.
Hell you could likely use 4 of them in the same pcie 5.0 slot and have it work without any kind of saturation.
>>
>>109603609
Reddit has a faggot reputation on 4chan already
>>
>>109605994
Gold holds value over millenia.
GPUs will be worth more in the short term but will become worthless in the mid term.
>>
>>109605987
Me too anon, me too...
>>
>>109606066
What will they be worth in the long term?
>>
>>109606160
>>109606160
>>109606160
>>
There's clearly an influx of new retards across 4chan. Might be something to do with the recent gaming related leaks.
>>
>>109606165
Whatever intel 8086s are worth on eBay (not much).
>>
>>109606178
What leaks?
>>
>>109605784
gemma moe is braindead as fuck, people shilling it are just coping because all they have is a laptop from 2021
>>
>>109606165
One would hope modern cards will be worthless in a few years due to having far better ones readily available by then
>>
>>109606194
GTA6
>>
>>109601570
This is incredible anon.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.