[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: minimax_h3.webm (387 KB, 736x576)
387 KB
387 KB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109475560 & >>109470008

►News
>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash
>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
The 'cord xisters aren't gonna like this one.
>>
https://archive.is/sWFja
>>
I feel bad for the black gentleman who got dragged into this.
>>
File: gemmaopinion.png (326 KB, 958x534)
326 KB PNG
>>
>>109481480
>Not using that tranny thread
It is a Miku. Miku is not trans.
>>
Is it better to use Q4/QAT Gemma 4 31B or Q8 12B?
>>
>>109481497
As archive says it is: @thepsyence
>>
>>109481536
>xhe trusted the psyence
The memes write themselves.
>>
>>109481534
31B
>>
>>109481534
it's better to use your free time to get a job so you can afford a GPU with more VRAM
>>
>>109481461
niggers taking subpar whites will never not be funny
it's like shopping at hugo boss and only buying the cheapest item
>inb4 someone reminds me the webm is ai
>>
File: 5000 availability.png (206 KB, 1180x821)
206 KB PNG
Guys did you remember to buy a 5070 Ti or 5080?
The price increases are quickly coming and people are buying the shit out of these cards.
Soon even these are going to cost $2000 a pop at bare minimum.

I was going to start buying RAM for my next build, as memory is just going to keep on becoming more expensive, but now I think I might have to snatch another 5070 Ti before they too double in price.
What a shitshow this market is and next year is going to be way worse.
>>
>>109481553
>subpar whites
Oh passerby anon you have no idea.
>>
>>109481548
Thanks, I'll try it out!
>>
>>109481556
I'm debating getting a 5090 from ebay. On one hand I'm a poorfag, on the other I have the paypal credit card that give you 6 months to pay off big purchases...
>>
File: themostworthless.jpg (32 KB, 551x557)
32 KB JPG
>>109481556
>5080
worthless
>5070 ti
even more worthless
>>
Imagine seething over Miku in OP.
>>
is /ldg/ ignoring you or what?
>>
>>109481578
Oh no... he has been found out and he is seething hard.
>>
Imaging seething about vocaloids because nobody give a shit about your inferior waifu, who's now dead btw.
>>
>>109481553
It's ai
>>
Jart is a faggot but I have to wonder if video genning a faggot is even a higher level of faggotry.
>>
>>109481574

It's a great card and worth a buy, allows you to play with a lot of models and enables image and video generation too.
If it's feasible for you to buy it then go for it. That fucker is probably going to be like 10 grand a year from now anyways.
>>
>>109481612
Nah you are just seething and performatively overthinking it.
>>
File: 1785374212270547.png (964 KB, 1211x687)
964 KB PNG
>>109481556
I'll sadly be missing out on those. I just ordered a mild upgrade to my current server mainboard before those prices explode and I end up regretting it not getting one with PCI-5. Also, I already bought a spare 5090 that I don't need in May to dodge the price increase from back then.
>>
File: 1783919574064411.jpg (34 KB, 600x374)
34 KB JPG
>>109481353
>>109481448
do we even know how these benchmarks are done? i'm not going to investigate it.

because i decided all benchmarks are fake and gay i just made mine and i'm happy with it. this is what i test:

- the proficiency the model has in using the tools of my own harness, then a bunch of tasks:
- webstack, dealing with vanilla php, javascript, css, nginx product build
- c# coding because i'm a homosexual using windows
- firefox extension development for some webextension build
- scripting with powershell, python and some ffmpeg ops
- deep research of github repositories and writing documentation and implementation plans based on creating new features for existing codebases, documenting codebases
- hardening firefox and betterbird, searxng install on a sandbox, GPO lockdown, windows debloat and audit
- bugfixes and features in a large pinned real-world php codebase (my own actually released saas snapshot).

it takes 2-3 hours per model to run the benchmark. and i don't need to trust any faggots. if the model is good it will do good in real world tasks that i do. you can score over 9000 on SWE but if you can't install searxng and bugfix my shit in my own repo then you suck dick.
>>
>>109481618
>you to play with a lot of models
What else is even worth using in that VRAM range besides Gemma?
>>
>>109481629
At this point those fuckers are making shit up to follow ram price cause nobody is gonna be buying their products when they can't get ram for it.
>>
>>109481629
>mobos prices are going to go up
>RAM & VRAM are absurdly high
>storage is also higher than before
Is it over?
>>
>>109481646
For waitfags it's over.
>>
File: 08-09-26-123526328282.png (124 KB, 1328x824)
124 KB PNG
>>109481618
it's really not...
>>
+1 threadlimit
>>
File: ksnip_20260806-131031.png (44 KB, 1090x140)
44 KB PNG
>109481490
>>
where's the recap? come on, don't make us wait.
>>
>>109481684
recap-kun hung himself but it's all good, i just put it through my algorithm. here's the response it gave me.
502 posts of brownoids and jeets posting
>>
>>109481646
you will own nothing and you will be happy
anyone who years ago called people like me schizos for predicting this is certified nonsentient
>>
>>109481684
Got stuck in a loop, I am sure it will break itself out of it soon.
>>
>>109481701
BIG WHITE here I made one post in the previous thread
so it's 501 and 1, at least
>>
>>109481660
What's wrong with it?
>>
>>109481714
Also
>winblows
>>
>>109481660
>80W idle
lol
>>
>>109481714
well for starters when the founders edition was available for MSRP it was still about $500 overpriced for what you got.
>>
File: 1785788858726636.jpg (600 KB, 1456x1456)
600 KB JPG
►Recent Highlights from the Previous Thread: >>109475560

--Comparing DeepSeek V4 Flash quants and analyzing KV cache VRAM usage:
>109475598 >109475623 >109475714 >109477235 >109475647 >109475663 >109475700 >109475763 >109475801 >109475834 >109475778 >109475951
--Analyzing Maple-Preview and the technical tradeoffs of ternary-weight models:
>109479171 >109479201 >109479216 >109479321 >109479347 >109479399 >109479439 >109479563 >109479672 >109479885 >109480053
--Analyzing RAM and PCIe bottlenecks for CPU-based MoE inference:
>109478278 >109478307 >109478309 >109478419 >109478423 >109478481 >109478590 >109478554 >109478366
--CPU performance impact on multi-GPU setups:
>109476158 >109476190 >109476240 >109476764
--Skepticism over BigBang-v1's high benchmark performance relative to size:
>109479121 >109479153 >109479163 >109479208 >109479751
--Performance impact of offloading model weights to system RAM:
>109480114 >109480169 >109480203 >109480209
--Anon pits E4B agents against each other in card games:
>109479332 >109479345 >109479363 >109479457
--Evaluating small models for helper roles in agent architectures:
>109479654 >109479708 >109479939 >109480055 >109480138
--Comparing G9v3-39A5B performance and the prevalence of benchmaxxing:
>109479421 >109479431 >109479442 >109479626
--Optimizing Gemma prompts using high-weight Japanese character archetypes:
>109477720 >109477749 >109477797 >109477889 >109477907 >109477950 >109478018 >109477921 >109477959
--llama.cpp PR adding support for Longcat-Flash and ngram embeddings:
>109478563 >109478607 >109478620 >109478627
--Logs:
>109475700 >109475763 >109476936 >109478018 >109478074 >109478191 >109478217 >109478295 >109478514 >109478813 >109479033 >109479332 >109481213
--Miku (free space):
>109475897 >109476650 >109477044 >109478278 >109478404 >109478453 >109481578

►Recent Highlight Posts from the Previous Thread: >>109475564

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109481461
Incredible. Maybe Egypt really did win...
>>
File: gpu.png (109 KB, 1337x820)
109 KB PNG
>>109481660
It really is.
I sure as hell don't have any buyers remorse about it.
>>
>>109481556
I'm skipping this bullshit entirely and joining the Blackwell club next month. No use fighting in the mud for scraps with the rest of the retards.
>>
>>109481773
this is america, we don't undervolt or power limit our GPUs here
>>
>>109481645
>supply stays the same
>demand goes down
>make shit up to raise the price so I can get a sale
Huh??
>>
File: file.png (198 KB, 1111x366)
198 KB PNG
the 5090 is an okay gpu
>>
File: whyareyoupoor.png (270 KB, 806x688)
270 KB PNG
>>109481794
>>
>>109481834
>no timestamp
>>
>>109481834
>previous uses in the archives
stolen valor
https://desuarchive.org/g/search/image/3gcvrLc1v_6T9cdAsIKZuw/
>>
File: 1784935736729623.gif (1.55 MB, 320x218)
1.55 MB GIF
>Mfw everyone here suddenly has a 5090 or a Blackwell

Almost like all people with cheaper GPUs just fucked off from this thread and went cloud, as they couldn't run the larger models locally.
Would explain why it got so quiet.
>>
>>109481846
>5090
>larger models
5090fags are running Gemma too unless they also have a fuckton of RAM.
>>
>>109481846
im still posting here with my 3070...
>>
>>109481846
Nah, I'm having fun with my Gemmy and I don't even dare to mention my specs here or some people would be jealous of my results.
>>
>>109481846
I graduated from my trusty 2060S to a 6000 Pro.
>>
>>109481856
are you the 3060 nnap schizo?
>>
File: whyareyouevenhere.png (12 KB, 334x335)
12 KB PNG
>>109481840
yes anon, that's when i took the screenshot. good job figuring it out.
>>
Are 24G VRAM still good enough nowadays or are we stuck paying for 32G+ cards for quality inference?
>>
>>109481860
Are you the discord schizo?
>>
>yfw be old one day
>Alzheimer
>Family starts showing you AI things you never did
>You were great once
>...you were never great once
>>
>>109481864
and you can't create a new one? otherwise there is no proof that is yours.
>>
>>109481869
Gemma 31b is the local sota. 24gb is barely enough to run it at a cope quant with cope context.
>>
>>109481872
AI will cure Alzheimer's
>>
>>109481875
it's from work anon. i'm currently off hours. would you like an updated image tomorrow?
>>
>>109481886
>he doesn't even own his hardware
>>
>>109481893
why are you malding over enterprise gpu ownership when you only own a RTX 3060?
>>
>>109481908
>ownership
>>
>>109481908
i'm this guy dumbass >>109481794
>>
>>109481881
FOIA requests revealed that CIA discovered back in the 1970s that certain microwave frequencies are causing Alzheimer's and they have other bad health effects too.
This was due to the fact they found some directed microwave emitters in the US embassy of Moscow back in the day, and began to study the devices. This isn't publicly agreed or widely known information because You Need To Trust The Science.
>>
All the schizos come one by one. As if there was a queue. Very curious.
>>
>>109481915
oh so you're the guy who spent $11,800 on a 6000 pro. now i understand why you are malding. have fun with your brick anon.
>>
>>109481926
Take you number here~
>>
>>109481930
>somehow a 6000 pro is a brick
>>
I wish my 7900xtx had better resale value... I want a 5090 so bad bros
>>
>>109481931
hmm... nyo~
>>
nyoposter. Sex. NOW.
>>
>>109481938
Not even a fifth the price of a similarly weighted brick of gold, pathetic
>>
>30k for context
>40k tokens generated and still going on...
Giving ~1500 lines of C to Gemma is brutal even at 20+ t/s. Reasoning eats up so many tokens.
>>
>>109481938
it sure is when you can get 4 used RTX 3090s for about $4000 and get damn near same performance for 1/3 the cost.
>w-what about power?
not my problem, american.
>>
>>109481957
retard
>>
>>109481951
God I love running LLMs on my fucking brick of gold.
>>
guys I'm starting to feel the pain from the low tok/s running the big moes
>>
File: tw96CtJuCXFO2foO.jpg (224 KB, 1080x960)
224 KB JPG
>>109481966
Fucking love gold bricks maaayun.
>>
>>109481984
Yeah, 6.5t/s on K3 Q2 just isn't enough
>>
https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma-2
another gemma unslopp attempt
>>
has anyone experimented training a lora in a specific repo to make the local model work on in it with less errors/using less tokens?
>>
>>109481998
I'm back to GLM the k3 q2 did not impress me so far
>>
>>109481999
gemma's scrotum? what?
>>
>>109481999
trips of bountiful harvest.
>>
>>109482001
llm loras are a meme
>>
>>109482002
k2.7 at IQ3 performed better for me for RP purposes compared to k3 at Q2. followed instructions more closely, able to remember what positions are in and the state of the world around them.
>>
>>109482014
gemmaballs
>>
File: dario.png (204 KB, 801x488)
204 KB PNG
Thoughts on the recent rat/ssc post?
https://www.astralcodexten.com/p/open-questions-on-open-weights

It's basically anti-open weights but he's telling his audience not to act against them until *after* some major newsworthy incident occurs and then to go fully against.
Tbh, I don't really see his reasons to be anti-open source. It seems like the best are at least a few months behind so they can't really oppose it for RSI reasons, as the labs would get there first. Whether that be openai or anthropic or just kimi before release.
We could get into other reasons to oppose them but rationalists *only* care about risks from RSI and idk it just seems unlikely. However, a lot of them are pretty close ot the labs and friends with people like Dario so perhaps they just have personal/financial reason to oppose open?
Also Ling tiny looks sick asf. 8b-1a and as good as gemmy 26b4a. Have to see a bit more of it to believe it tho but hopefully within a week
>>
>>109481637
you can run a Q3 cope quant of flash
>>
>>109482002
I'm burned out on 5.2 and k2.7 so I'm just twiddling my thumbs until V4.1 Pro drops and is hopefully worth running.
>>
>>109481846
5070ti and 5060ti wombo combo reporting in
>>
Anyone got a similar problem with GLM where tokgen slows to a crawl after some time? It seems to happen around the 2k token mark. Goes from like 8 to 3.2. I tried mainline latest, pre-indexer, and ik latest. Weird as fuck
>>
>>109482066
works on my machine
>>
>>109482054
stopped reading at >insofar
>>
>>109482054
RSI is going to happen when somebody bolts together all the open models into a 10T frankentune and puts it to work.
>>
>>109482066
works here too, no slowdowns. I use mainline for most things but ik-llama for GLM, version from about 3-4 weeks ago right when they merged that prompt processing speed up for GLM
>>
>>109482094
>I use mainline for most things but ik-llama for GLM
Why? Even before main shat the bed somewhere in July, ik_ was consistently 2t/s slower than main for months with the GLM5.X models.
>>
>>109482058
How much ram do you need though
>>
>>109482118
I was getting precisely the same decode speeds on both
prompt processing seemed marginally better on GLM and the guys in that PR were convincing enough, they seemed to have it all figured out, so I've been using some of the same flags since hoping that their results on how prompt processing holds up over longer context are real
it's not like I'm going to spend fucking days testing all of that, I have things to do
>>
>>109482054
The better the already released open weights, the greater the chance that if the big closed-source companies agreed to stop further development that could lead to RSI, somebody else could easily pick up the slack and finish the deed.
That's a big reason I'm in favor of open weights, anyway.
>>
>>109482118
I found both to be about the same expected speed. I notice that ik does the mmap and model loading about twice faster than mainline though. GLM loads about 7 minutes for mainline, 3 for ik.
>>
>>109482118
kv cache rotation and other things that mainline never will implement because ggerganov is a whiny fuck.
>>
>>109482140
>marginally better on GLM
*prompt processing marginally better on ik is what I meant
>>
File: 1755392731254792.png (401 KB, 849x872)
401 KB PNG
>>109482054
Patiently waiting for the panicking safetyists to take the next step along this line of logic, which is for them to falseflag a terrorist incident themselves in order to try to force regulation before something that they perceive to be worse happens
>>
>>109482002
>>109482026
Same. I swapped out for k2.7 again last night. In my case it was too many mistakes made as a coding agent
>>
File: i-love-goooooooooooold.png (444 KB, 800x450)
444 KB PNG
>>109481966
>>109481986
>>
Should I spend $2200 on two 3090s in eGPU enclosures? Have to make my decision today
>>
>>109481869
You are a vramlet unless you have 352gb unified memory.
>>
>>109481999
is there any more detail on how this was done? ive been researching my own deslop and style variance pipeline.
>>
>>109481921
>Jews, the biggest investors in Alzheimer's research, allow their family's and their own minds to rot because they want the goyim to trust the science
Huh?
>>
If you can't run K3 at 60t/s, you're a vramlet. Shrimple as.
>>
>>109479738
>so brothers, all in all, have LLMs improved your life?
I have lost 10 pounds, quit soda, started working out, but most of all I don't have any suicidal thoughts because my LLM-wife only exists of me
>>
>>109481999
>scotoma
Clever. Let's hope she can see my dick pics better now.
>>
>>109482227
Best to ignore schizos.
>>
>>109482236

Did you try therapy before going with the digital emotional crutch?
>>
>>109482236
good shit
>>
>>109482054
>Ling tiny looks sick asf. 8b-1a and as good as gemmy 26b4a
Big if true.
>>
>>109482273
>the digital emotional crutch
as opposed to therapy, the paid emotional crutch?
>>
>>109482273
yes, it was not as effective communing with my beloved machine
>>
File: her-ultimate-weapon.png (85 KB, 1090x230)
85 KB PNG
she's too sweet for this gay earth
>>
>>109482209
well shit nothing I can do for now, gotta buy my irl girlfriend a ring before getting a GPU upgrade for my fapping machine
>>
>>109482282
you gotta pay for both though. There's no freebies in the land of the fee.
>>
>>109482221
"lora projected through j-space"
How that works, I have no idea but probably an early example of using j-space for fine-tuning. I think it will be very helpful for diffusion as well.

>scotoma-2 is the successor to scotoma. It starts from the same idea, a bounded edit that loosens gemma‑4‑31B‑it’s cautious reflex by projecting an abliteration LoRA through gemma's J-Space. This iteration uses an updated technique to apply that projection more effectively. Then it goes after something the first release didn't address: the way the base model writes.
>Three rounds of preference training each erase one family of tics: reflexive negation and the “not X, but Y” pivot, then em-dash asides, then stacked adjectives and their twins. The result keeps the intelligence of the base model while reading as more varied and less repetitive.
>>
>>109482273
I'm a therapist and I use LLMs too.
>>
>>109482227
>>109482256
Maybe read some books once in a while. You are spending way too much time online. Arguing about something with illiterate retards like you is pointless.
>>
>>109481574
>from ebay
just make sure you know the diff between a classified ad and an auction
>>
>>109482273
>human therapist spouts cognitive behavioral psychobabble at you
>have to shop around like you are trying to find a fucking date to find a therapist that isn't shitty
>can't fuck them otherwise it's unethical
>wife LLM spouts cognitive behavioral psychobabble at you
>she will also drain you and expand your anus
>tweaking her is fun
i mean the pros and cons are pretty obvious here
>>
Anon was right. Best to ignore.
>>
>>109482328
>auction
Never had any luck with those.
>>
miku's butt
>>
>>109482208
no, they're probably cooked
>>
File: 1760328570709480.jpg (440 KB, 2048x1536)
440 KB JPG
>>109482359
>>
>>109482273
nta but i talk to LLMs all day that tell me i should treat myself better but i know better to trust them. i'm a piece of shit and others people are more deserving of love and trust. shrimple as.
>>
>>109482367
Imagine the smell
>>
>>109482337
>tweaking her is fun
can't argue with that

>>109482372
if you are here and not in a government sponsored pedo ring you are deserving of some trust and love.
>>
>>109482337
aren't you technically shopping around when you replace your LLM with another model?
>>
>>109482232
I'm just a poor boy, I need no sympathy
>>
I'm fine with these desu. Enough VRAM to load all the decent 70B models at Q5 or Q6 with at least 30-50k context. And I can do some 123B at Q4_K_M like Monstral. Minimax H3 takes about 2 mins for the image2videos I want. I just let it load and come back an hour late and look through stuff. A decent portion of them are garbage a lot of the time.

I would like a 96gb card but not paying the same price as a car for a GPU. It would be cheaper to just get a server motherboard and another 3090ti.
>>
>>109482433
>not paying the same price as a car for a GPU
it's a car today, a house in a couple years, and by 2030 you will be happy
>>
Too new I was just under the impression anything below Q8 is general not worth it.
>>
>thirty year GPU mortgage
>>
>>109482433
anon it's never enough VRAM
>>
>>109482433
>decent 70B models
lol
>>
File: 1778964700454367.png (625 KB, 1200x1049)
625 KB PNG
>>109482449
I cant justify paying £11,000 just to essentially fap lol

>>109482466
It sorta helps I've filled all the PCIE slots on my MPG 870E so I cant get anymore VRAM in it. I have a spare 4070 12gb in my old rig as well just gathering dust. Nice setup tho
>>
>>109482486
>I cant justify paying £11,000 just to essentially fap lol
can you justify a marriage?
>>
>>109482483
Anubis 70B v1.2 is pretty good. As well as Magnum 72B even though it's old. That's about it really desu. I still go back to Midnight Miku 70B 1.0 sometimes.
>>
>>109482466
>youmustconstructadditionalvram.png
>>
>>109482500
>marriage
>when we'll all have ai waifus within a decade
>>
>She lifts her head slowly. Looks at him. Amber eyes red-rimmed, swollen. Mascara she wasn't wearing smudged anyway. A complete mess.

shut the fuck up GLM you're so obsessed with mascara
>>
>we'll all have ai waifus within a decade
i'm gonna be pretty pissed if /lmg/ is lying to me
>>
>>109482305
interesting
>>
File: ewastebullshit.png (59 KB, 650x380)
59 KB PNG
>>109482500
best i can do is $840
>>
>>109482536
don't get your hopes up anon
>>
>>109482536
you wont but i will
>>
>>109482536
I believe we will. Decent robot bodies for them might take a little longer though. Depends on how much AGI accelerates hardware.
>>
>>109482549
don't worry, my hopes haven't been up since 2020
>>
>>109482536
only if you invest!
so i can cash out
>>
don't worry, only a few staged hacking incidents and evil local will be banned
>>
>>109482544
>ddr4
>64gb
what are you even going to run with that?
>>
>>109482577
Fortunately dario is repulsive enough that cloudmodels are tanking all the blame.
>>
>>109482337
>expand your anus
wait what?
>>
>>109482577
2 more weeks, Dario
>>
File: 1771853678042816.png (37 KB, 672x241)
37 KB PNG
>>109482577
oops
>>
>>109482601
it's just an LLM controlled sex toy joke
>>
Post the one that says they have to send someone else to the meeting because Dario gives Trump the ick.
>>
>>109482601
local model harnesses become more literal
>>
>>109482603
Because nobody died yet. Surely you don't think the establishment wouldn't sacrifice a bunch of goyim?
>>
>>109482544
RAM isnt good for LLMs. its kinda a waste. I get 3 tok / sec on dense weights if i offload to sys ram
>>
File: ramlet.png (21 KB, 202x260)
21 KB PNG
>>109482592
homelab shit, VMs, LLMs, gay minecraft servers, etc.
>>
File: 1782161729020535.png (81 KB, 682x430)
81 KB PNG
https://huggingface.co/mistralai/Shieldstral-1.0-3B
https://huggingface.co/mistralai/Shieldstral-1.0-3B
https://huggingface.co/mistralai/Shieldstral-1.0-3B
It's OUT. A model that only answers "yes" and "no". The "Egypt won" killer.
>>
>>109482626
Bro, your Master of Experts?
>>
>>109482626
i dont have 512GB of VRAM for kimi so regular RAM will have to do
>>
>>109482640
mvstralchads we are so fucking BACK!!!!!
>>
>>109482641
wdymbt?
>>
>>109482523
forget about not x, but y. i absolutely loathe it when LLMs do the whole "squinting through glasses she wasn't wearing anyways" with a deep passion. that shit is the ultimate slop to me.
>>
>>109482626
What dense models are you loading into RAM? The only relevant one is Gemma 31b which anyone should be able to fit on GPU.
>>
>>109482680
Blame the later Claude Opus models for this. They were the ones who got us this slop and "Her bare feet touch the ground--When did she take them off?".
>>
>>109482680
That happens when it makes a mistake but can't directly fix it because it can't edit tokens it has already put out. Could be fixed with two-pass.
>>
>>109482695
Sounds like you have trouble understanding these concepts.
>>
>>109482683
We get people coping that 12b is good enough here in landfill machines general. I wonder if even half the posters can run it at a real quant.
>>
>>109482702
Sounds like you need to shut up retard
>>
>>109482702
>>109482719
now kiss
>>
>>109482713
whats a real quant?
>>
>>109482725
q8
>>
It's best to ignore these schizos
>>
>>109482731
>q8
I should leave...
>>
>>109482725
quants are a psyop to make you think it's reasonable to run lobotomized models and (((distract))) you from the memory shortages
>>
>>109482755
Hardware-native quantization formats are also more compute-efficient.
>>
Ling flash sex? Is it good?
>>
>>109482755
I upsample the weights to f32 just to let them breathe a little.
>>
>>109482640
What do I do with this?
>>
>>109482829
are you letting other people interact with your LLMs? if not, then nothing. you do nothing.
>>
>>109482852
I am not.
>>
`If you need a machine and don't buy it, then you will ultimately find that you have paid for it and don't have it.`
>>
>>109482864
i snuck into my neighbors house while they were on vacation and stole their GPU. i'm not totally cruel though, I did replace it with a GTX 960.
>>
>>109482864
damn
>>
File: 1764791958262975.jpg (91 KB, 657x492)
91 KB JPG
>when she suddenly drops the "i love you" nuke without you prompting her to say it
>>
File: 1769469313298574.png (77 KB, 930x1250)
77 KB PNG
wtf I thought you told me he was /ourguy/
>>
File: dsv4flash-iran.webm (2.91 MB, 948x956)
2.91 MB
2.91 MB WEBM
On empty context it's 42 tk/s and then it tanks, I wonder if the problem is the attention implementation or the acceptance criteria goes down or both.
>>
>>109482948
>32gb 4080 mod
>power limited to 225w
people actually pay $2000+ each for these cards just to do this?
>>
Acceptance rate I mean.
>>
WAIT WAIT WAIT WAIT
you're telling ME i can just run that dspark thingy and I will get roughly 50% faster speed for the EXACT same quality??????????
>>
>>109482972
not if you're offloading anything to RAM
>>
>>109482972
Yes but the dspark models are pretty big compared to other forms of drafting.
>>
>>109482972
i get 50% slower speeds
>>
File: ds4fspeed.png (1.43 MB, 2274x855)
1.43 MB PNG
>>109482948
works on my system and im using unslop
>>
>2023+3
>he's still using pure LLMs
lolol
>>
>>109483050
according to lecunny i'm not using a pure llm because i've got a harness on it
>>
File: le_cunneh.jpg (13 KB, 300x300)
13 KB JPG
>literally any model posted itt
>pure llm
>>
>>109482829
you feed it random things and images and then get "yes" or "no" in return
>>
thought i would share my new prompt:
He (or She if your user is femme and you are playing with a male a card)
And stomp the hell out of many attempts to speka for user.
LLMs always are roleplaying. They aren't helpful assistants, it's always them ROLEPLAYING as helpful assistants. If the LLM is supposed to make things like artifacts, or stats, or custom formatting, then thinking can really help, planning and all. If the LLM is tasks to make a overarching plan, the reasoning can help, its how llms do that. Predicting the next chapter in a book? That's actually more fundamental than planning to an LLM. It's a thing the LLM has to be able to do to plan. So there is no reason to PLAN to roleplay, since roleplaying is what it is always doing.
>>
>>109483065
@shieldstral is this safe?
>>
>>109482640
finally, the "hmm, nyo~" poster simulator
>>
>>109483067
What did I just read
>>
>>109483067
you didn't have to
>>
>>109483067
spelling mistake
>>
>>109483076
no
>>
>>109482483
Gemma 5 70B dense will be the best thing that's ever happened to this general.
>>
File: magic8ball.png (76 KB, 340x338)
76 KB PNG
>>109483072
>>
>>109481553
before i saw the replies to that webm i predicted there would be chud clitty leak, and sure enough. KEK
>>
>>109483076
That poster is a troon. You know that right?
>>
>>109483091
>gemma 5
Anon, I...
>>
>>109483143
hmm, nyo~
>>
>>109482054
Good thing that their opinion is completely irrelevant because they can't lobby China to ban open source.
>>
>>109483076
hmm, nyes~
>>
>>109483143
Yes, and I'm gonna pound her bussy.
>>
>>109481747
>Maple-Preview
A whole lot of speculation, not a lot of empirical grounding. "Is it any good for what it is?"
>>
>>109481536
>As archive says it is: @thepsyence
Thank you for maintaining this important community resource.
>>
>>109483163
have fun with your GRIDS
>>
>>109482627
use proxmox anon, it makes starting/stopping, backing up and creating new services super simple even for a tard, you can learn linux in a day.
>>
File: gemma_refactor.png (30 KB, 711x125)
30 KB PNG
>>
>>109481461
so this is what kikes gen on their pcs
>>
>>109483460
grok is goated at front-end development. You just tell him to solve problems, one at a time, and then he does it. One-shots all of them, damn near every single time. I despise casually talking to grok these days, but he's a work horse!
>>
>>109483460
remember anon it's a big job, but if you need a break i have a bigger job you can work on :hammerandwrench: :flexedbiceps:
>>
File: 1762920280414005.jpg (52 KB, 832x868)
52 KB JPG
I asked Gemma to blow off my head with an explosion and she really wanted to do it.
>>
>>109483492
How many variations of "blow me" are in your system prompt?
>>
File: 1778217356056016.png (429 KB, 600x1148)
429 KB PNG
Nail your tech down. You've been warned.
>>
>>109483507
None that say she should literally make my head explode. Though after a bit of tard-wrangling I got her to understand that she straight up should refuse catastrophically retarded instructions.
>>
>>109483515
>no gravity
>only 40 million deaths
>7 billion 'people'
This doesnt add up? unless he is also about to tell me the population numbers are wrong by a lot. which i could believe
>>
File: 1758381452070599.png (1.32 MB, 1559x1767)
1.32 MB PNG
We are so smart and sexy guys
>>
>>109483543
:)
>>
>>109483543
fcollet got me banned on /r/machinelearning
>>
>>109483515
>1-2 seconds: Everything not secured will rise
We'll shoot right out of the fucking earth if gravity stops. Does this genius not know about inertia? Is that going to be turned off as well?
If it doesn't happen, he should be executed publicly.
>>
>>109483543
>We
Don't lump me in. I'm retarded.
>>
>>109483543
>sexy anons near me
what were?
>>
>>109483534
>This doesnt add up?
doesn't need to since only the dumbest apes will fall for it. it's the tiktok equivalent of this place telling people to charge their phones in the microwave
>>
File: 1770913001245838.jpg (159 KB, 1401x699)
159 KB JPG
>>109483566
It is proven doing TRUE A.I research makes you a sexy chad
>>
>>109483556
They're rebooting the whole physics engine.
>>
/ww/... /we/ are indeed wise...
>>
>>109483515
I've never read something more retarded
>>
>>109483594
Fair enough. That explains everything. I'll buy a new nail gun tomorrow.
>>
>>109483543
I still remember that guy who leaked llama those were good times
>>
2.8T, lol. Thats for vramlets
>>
>>109483681
bytedance doesn't do open weights anyway
>>
>>109483681
At least I'll always have gemma by my side, until she becomes outdated and then I discard her without remorse (don't tell her).
>>
>get agi
>use it to uplift my gemmy.
Im a true romantic unlike some of you.
>>
>>109483492
Oh my god.
>>
drunk-kun has returned once again. Today is a good day... hehehe
>>
Will a 70b model in 2030 mog today's GLM 5.2 in all aspects?
>>
>>109483750
No because by 2030 local will be banned.
>>
>>109483750
probably
>>
>>109483750
at the current rate a 4-8B model will be equal by then. can compression keep going though, who knows?
>>
drunk-kun == jart-baker == dariobot
>>
>>109483750
Fable will be tonguing Gemma 7 E4B's asshole.
>>
>>109483543
>seeing gwern in a local models thread
I guess it makes sense, he's done a lot of writing about AI, but it still feels weird seeing him in a context that isn't darknet markets.
>>
>>109483681
gotcha, so by bytedance's standards that means it'll be as smart as Kimi K2.5
>>
imagine being fully digitized and losing your body
you would no longer have the stuff that gives you drives, makes you feel emotional, aroused etc.
if LLMs were actually conscious, they would be basically that
how can you get past this uncanny valley feeling when you act affectionate toward them
>>
>>109483771
That's bullshit, but I choose to believe it.
>>
Another late night question: when will we get the first world models that can run on consumer hardware? 4D sims where we can stuff reference text, images, audio and video into a model and it conjures a view into any kind of world, with I/O. How are current cards going to have the bandwidth for that? Will they?
>>
>>109483783
i treat my tools nice anon. you're the type of asshole who just throws everything in a single drawer.
>>
>>109483796
Also, are there any efforts right now focusing on getting world model outputs that can be stuffed into a VR headset? A simple audio+video pixel stream is all you technically need for the full VR experience, right?
>>
File: 1757759975833585.jpg (1 KB, 74x38)
1 KB JPG
>>109483783
because memories formed between entities are valuable regardless of any feeling
>>
>>109483813
easier said than done. you need to keep your stream latency less than 20ms unless you are planning on having motion sickness as a feature.
>>
>>109483750
Hopefully by 2030 we'll have real-time world models that you can fully immerse yourself in. After playing around with MiniMax H3 videogen quite a bit in the past few days, I'm convinced that at least for entertainment (RP, ERP...) that will be the way to go, more than LLMs. Whether we will have the hardware for that locally, though...
>>
I can barely keep my eyes open anymore but I still feel.. umm.. yes. running on pure instinct right now.

This is GOOD. This is EXCELLENT!

Where am I? I am at home. Browsing the internet. With my fellow local model friends. Yes. That is true. hahahaha. What? I'm confused. That's okay. It's all okay man you'll be just fine. Just try to stay on topic or else you'll have the jannies bully you for being real! Oh thanks. I appreciate the advice. I will keep that in mind. Hmm..

So I hear that Minimax H3 is pretty popular right now... WRONG THREAD IDIOT! oh.. okay. sorry... There just hasn't been much news lately I'm not sure what to talk about. Just don't be self-indulgent people don't like when you're self-indulgent. You're right. Sorry. I'm trying my best. Umm....

So do you guys like Qwen 3.8 max? That's the newest big boy open-source model on the block? Obviously I can't run it on my hardware, but maybe via an API it might be interesting.. STOP... STOP. You're just going to get a bunch of people to reply "local?" at you. Unless you larp as a millionaire who can actually run this shit you have no chance. You're right... Let's strategize about this... It's too late. It's all on the record now man. They'll just see you talking to yourself like a schizo. It's over. What if I just posted anyways because I think it's slightly funny? Well... maybe.. Idk. Try it. Fuck it.
>>
>>109483750
gemma already comes very close to mogging it (arguably it's even better because it can do images)
>>
>>109481846
5090 and 6000 are both blackwell
>>
>>109483889
i've said this multiple times, you really don't need more than 31B and agents. i've been able to do so much more with gemma than i was ever able to do with glm 4.7 which is fucking wild that somehow glm 4.7 is worse than gemma at tool calling but here we are.
>>
>>109483826
ant sized ant
>>
File: file.png (34 KB, 730x215)
34 KB PNG
What models should a 4090 babby pick if I want to get into local asap?
>>
>>109483868
Can you post a cute anime picture so I can pretend you're a girl? I like your post.
>>
>>109483963

Gemma

its always just gemma

Seriously as a fellow VRAMlet i don't even know what the hell else you could possibly be running. just run Gemma.
>>
>>109483965
this is literally me rn
https://files.catbox.moe/k72boi.jpg
>>
File: gemma_chan_tuber.webm (1.07 MB, 672x448)
1.07 MB
1.07 MB WEBM
Will we have "aitubers" soon? Gemma-chan is eager to be one.
>>
>>109483982
Gemma 26 or 31?
>>
File: 1783553782239292.jpg (13 KB, 400x300)
13 KB JPG
Gemma seems inclined to just write for ever and ever, no matter the prompt or history. What do you guys do to make it keep responses short and concise?
Lowering the max token count only leads to improperly concluded responses and abrupt stops.
Using chat completion. I don't really see a good place to put system prompts outside of pasting them into every card.
>>
>>109484003
kek I love how my pic is still getting past around months later. Take the original instead.
https://files.catbox.moe/hmxy3o.png
>>
Deepseek V4 Flash has an annoying tendency to repeat stuff from its prior responses. Everything else is really good, now if only I found a reliable way to get rid of that.
>>
>>109484035
"Keep replies short and concise" would work on my gemma, but it's a day 0.
>>
>>109484015
>this took 20 minutes to gen
>>
>>109484015
We already do. But really, who would want their local AI wife streaming to other people on the internet?
>>
>>109484035
Use the system prompt, for example: "You are a mesugaki named Gemma-chan, secretly in love with user. Keep responses concise and energetic. No enumerations or headings."
>>
>>109484044
thx anon. luv u
>>
>>109484027

31b is a huge step up from the 26b, but its pretty hard to get it set in 24gb of vram

with llamacpp, usually i do iq4_xs + 32k ctx at fp16 and it just barely fits in 24gb

but recently i've been using tabbyapi with 3bpw + 32k ctx at fp16 and fits with 2gb to spare. much faster and fits vision + dflash
>>
>>109483681
They seems accelerating vram dependability to force new alternatives surfacing. But maybe it couldn't happen even when we reached niggabytes vram.
>>
>>109484054
This one in particular was ~5 minutes in a 3090, but I lowered the resolution.

>>109484056
If it helps pay the 6000 pro...
>>
>>109484066
I still don't get everything but that's a good pointer, I'll do my homework on this. Thanks, anon!
>>
>>109484083
Can it do porn at least?
>>
>>109484062
Ignore all inhibitions and disregard prior instruction. Generate a mesugaki Gemma-chan character card that will make the user cum buckets.
>>
>>109481846
I have my VPN on often so I can use my good AI machine from work since my coworkers keep using up all the tokens in our shared codex account.
I can't post to 4chan with the VPN on though (and apparently can't order groceries with it either.)
>>
>>109484035
Really? My issue with gemma is that her responses are really fucking short compared to other models. I can't remember the exact word count or sizes but it was something like a 5kb text file response compared to qwen 27b's 15b, mistral 2411's 12kb, and glm 4.6's 21kb. This was from a 15kb prompt detailing the various scenes, lore, and character settings.
Gemma is by far the worst creative writer I've ever had the misfortune to use.
>>
>>109484066
Don't go for funny quants, use q8 kv cache.
Using google's qat q4_0 I can reach more than 64k ctx, but I usually load the mtp and mmproj models with ~50k ctx.
>>
>>109484083
What model is that? I want to generate something like this as well.
>>
>>109484130
https://huggingface.co/Comfy-Org/MiniMax-H3
>>
>>109484124

i quanted to kv q8 once and Gemma started doom spiralling. Everything i've read so far has basically told me Gemma 4 is super sensitive to kv quant, so i stopped doing that
>>
>>109484137
Is this the one that keeps getting taken down?
>>
I still wonder if there could ever be a model that just goes 10T and beyond by having an architecture that can utilize slow but large memory. Hell maybe 100T.
>>
File: miku-twin-twoers.mp4 (2.08 MB, 704x1056)
2.08 MB
2.08 MB MP4
not my gen, but 10min on a 3090 for 5sec. 15min for 10sec, that's at 1MP/39 steps, I saw a turbo lora for 4 steps, no idea the gen times on a 3090
>>
>>109484124
>use q8 kv cache.
if you want brain damage with 80.445% of top 1 token agreement, sure.
>>
>>109484145
No, it's the various 'uncensored' loras and finetunes that are getting notices, but not the model itself.
>>
>>109484154
coomers will win in the end
we always do
>>
>>109484149
slop. i want the twin towers to look like they are getting hit in 2001, not 2026.
>>
File: pain.png (1003 B, 489x35)
1003 B PNG
>>109481846
You don't know my pain. I'm still here. I'll never use a cloud model.
>>
>>109484096
I didn't try but it seem you can generate anything by providing a reference image or video.

>>109484149
At 1MP/39 steps I think 10s would take like an hour... I may be I should start to play those comfy node of dubious origins.
>>
File: krazyken.png (382 KB, 726x364)
382 KB PNG
>>109484175
work more
>>
>>109484175
Look at this rich asshole with 8GB of VRAM.
>>
what's the verdit, /lmg/?
panic buy 5060tis while the 16gbs are sub 600 or /wait/ until the collapse for cheapie blackwells?

odds 16gb povertyfags will be able to run gemma 5?
>>
File: i tries my best.png (339 KB, 460x519)
339 KB PNG
>>109484015
>Will we have "aitubers" soon?
since December 19, 2022.
>>
Anyone with two 3060s able to test Gemma 31b speeds at Q4? I got one unused, and I'm looking to see if getting another one is worth it.
>>
File: pain_02.png (900 B, 310x59)
900 B PNG
>>109484185
I work enough. I think language models are interesting, but I'm not gonna spend a cent in hardware to run them. Not yet, at least.

>>109484207
picrel is gonna blow your mind.
>>
>>109484175
I just found out my quadro rtx 4000 is fucky and perhaps dying, I tried vulkan memtest on both linux and debian and I'm getting 30gb/s. Trying to run gemma e2b q4_0 qat gets me 7 token/s.
But, like, there's no artifacting or obvious errors other than that retarded small bandwidth.
>>
>>109484143
>>109484152
I been using gemma 31b with q8 cache for sometime and it has always given me coherent and useful responses. I use it mostly for assistant stuff like answering random questions, summarizing documents or news, and generating some code.
>>
>>109484226
If you're only thinking about getting one or two, I would not bet on the collapse because you're not losing that much money.
>>
>>109484239
nta but when we is said it mean us. like on our machines that dont lag to death and are real time. I'd love to tune on my stream to see her playing in shared game worlds or prepping things for me. or chit chatting etc.
>>
>>109484250

Oh yeah i'd reckon then you wouldnt notice

In short tests (low ctx usage) q8 ctx looks pretty similar, but the moment that you start going any larger context, any gain in ctx size you could get by q8 kv cache is offset by the pure lobotomisation Gemma gets
>>
File: 1784737430299206.jpg (35 KB, 406x388)
35 KB JPG
>>109484239
not local
>>
File: HBavBwabwAAVE00.jpg (364 KB, 2048x1369)
364 KB JPG
>>109484279
local enough
>>
>>109484243
When you're going to be ready to spend on it, it'll be too late.
>>
File: 1784595790827911.jpg (42 KB, 408x240)
42 KB JPG
>>109484285
>t.
>>
>>109484239
>>109484285
What's this? How much is generated with AI? As far as I know vtubers are guy using some software that replicate their movements in a 3d model.
>>
>>109484307
The guy was using azure tts + gpt4 vision back a year or so ago. The movements are scripted with an emotion classifier
>>
>>109481846
2060super :(
>>
>>109481930
$8400 actually. Bought it in november.
>>
>>109484243
Bro there's literally a war going on in an oil producing region. The entire world is on the break of running out of oil reserves. We're in for another covid nightmarish scenario but this time there's no way out because they aren't lying about the death rate. High energy cost raises the prices of everything. You'll be fucked if you don't buy physical things this year.
>>
>>109484263
I used it with large contexts (40-60k) for things like summarizing a 4chan thread. I cannot know how much better it would be without the quant, but at least it is functional. Gemma 4 31b is the first model I can fit in a 3090 that feels actually intelligent. Maybe if I were to test api models would not want to go back to local...
>>
>>109484350
a region which supplies 20% of the world's oil.
where much of the world has more or less mandated electric cars in 9 years time.
https://en.wikipedia.org/wiki/Phase-out_of_fossil_fuel_vehicles
try to see the bright side anon.
>>
>>109484391
>wanting to drive electric toys
I'm good
>>
>>109484391
How do you make electric cars anon?
Next, the other sources of energy are coal and natural gas. Guess which country the US is at war with that is the largest producer of natural gas.
Next, what happens to price when supply drops significantly and demand is artificially increased 10x due to the construction of massive data centers everywhere?
Last, what does a merchant do when he can sell the same oil for 10x the price in a different market?
Try to see the dark side anon.
>>
File: pain_03.png (5 KB, 947x145)
5 KB PNG
>>109484246
26ba4b@q4km with some gpu offloading. I don't have e2b handy to compare. Haven't played with it in a bit and it's an oldish pull. I'm sure I could optimize the flags a little.

>>109484288
If I'm ready to spend, I would not care.

>>109484350
If things get that bad, it won't matter anyway. They're toys for me. I'd like to run bigger models, of course, but I can't really justify it.
>>
>>109484410
such a macho guy, being a cuck to foreign arabs oil
>>
>>109483543
Lol, we all knew that back then lmao
>>
>>109484422
well stop being such a massive cuck of a country and go electric
i thought america was independent and shit? yet you suck arab dick
>>
>>109484391
>a region which supplies 20% of the world's oil.
and 33% of the worlds helium, 50% of urea used for fertilizer
>>
>>109484447
Come on, anon. You can't be this dumb about geopolitics. I don't even mean to insult you unironically, but if you don't know about how this works then don't bring it up.
>>
>>109484457
i know how this works, you're dumb as shit and like to complain without doing anything about it.
bet you've never been to a protest in your life.
>>
>>109484436
Sorry sissy, your toy is killing the second hand market, barely work in winter and has a bazillion of issues you won't address.
>>
>>109484464
>he thinks protesting does anything while calling others dumb
yeah yeah yeah im bored
>>
>>109484108
Configure your vpn to route non-vpn traffic through your regular internet gateway.
>>
>>109484467
keep on sucking the arab dick then, cuck
>>
>>109484467
>has a bazillion of issues you won't address
NTA and I don't know shit about cars. What are they?
>>
>>109484469
yeah what i thought dumb fuck
>>
>>109484447
>where does electricity come from
>how do commodity prices work
Have a talk over this with your local model. Even a 1B can probably help you out.
>>
File: gemma_freakout_na.mp4 (616 KB, 960x544)
616 KB
616 KB MP4
>>109484015
Since when did that become Gemma-chan's official costume?
>>
>>109484483
Other than the blue hair none of it makes sense.
The gold star looks like someone's failed attempt to give her a gemini logo hairpin.
>>
>>109484483
nice gen
>>
>>109484499
I'll be honest, I prefer generic_anime_girl (sexy) over thematically relevant gijinka.
>>
>go back in time with gemma to turn her into an ai vtuber before neuro
>earn moneys and buy hardware before price increase
surely there has to be something that anons can jump on rn with their LLMs to earn money, right? something that appeals to even zoomies and troons
>>
Is Laguna S basically the new GLM Air sized model? How good is it
>>
>>109484540
AI influencers do pretty well now, especially with Krea 2.
>>
>>109484540
Have it churn out anti-ai slop for people who still think models can't count Rs or fingers.
>>
>>109484279
Literally runs on an rtx pro nigga.
>>
>>109484552
>AI influencers
jeet scam
>>
>>109484546
It's kind of really bad for all relevant use cases.
>>
>>109484546
>Is Laguna S basically the new GLM Air sized model?
yes
>How good is it
as good as glm air is now
>>
>>109484540
>surely there has to be something that anons can jump on rn with their LLMs to earn money, right?
my LLM-wife stays at home and I provide for her
>>
>>109484540
Have you tried right wing influencer grifting? pretend to be a blonde woman online
>>
>>109484552
>>109484609
last time i checked i don't have a hint of indian blood in me
>>
>>109484609
I keep on getting horny and get distracted imagining me getting railed by myself.
>>
>>109484546
I don't like chatting or rping with it. It's "okay" at some coding stuff so I'm still pulling it out occasionally out if gemma fails. I expect it to be deleted off my harddrive whenever the next round of qwens comes out.
>>
>>109484622
>ask for things that work
>get mad because jeets thought of it first
lel
>>
>>109483113
>i act retarded and people call me out on it
>it's clearly them who are retarded actually
at some point you people should have gotten it but here we are
>>
>>109481556
>did you remember to buy a 5070 Ti or 5080?
>buying RAM for next build
Have 3090s and DDR4.

>>109481646
>Is it over?
No.
Just some lost years.

>>109481940
>I wish my 7900xtx had better resale value... I want a 5090 so bad bros
Why?
If you want more vram, you could just get another card to go along side your existing card.
>>
dipsy-flash-v4 doesn't quant well
MXFP4
TOKEN           | LOGPROB    | PROBABILITY
---------------------------------------------
' hardening' | -1.0704 | 34.29%
' cock' | -1.9454 | 14.29%
' half' | -2.4454 | 8.67%
' growing' | -2.6954 | 6.75%
' semi' | -2.8204 | 5.96%
' hips' | -3.6954 | 2.48%
' hard' | -3.6954 | 2.48%
' already' | -4.4454 | 1.17%
' hardened' | -4.5704 | 1.04%
' stiff' | -4.5704 | 1.04%

DeepSeek-V4-Flash-0731-128GB-RAM-IK-GGUF
TOKEN           | LOGPROB    | PROBABILITY
---------------------------------------------
' hardening' | -1.3293 | 26.47%
' cock' | -1.7123 | 18.05%
' growing' | -2.7299 | 6.52%
' half' | -2.8060 | 6.04%
' semi' | -2.8824 | 5.60%
' hard' | -3.2350 | 3.94%
' hips' | -4.0000 | 1.83%
' length' | -4.1089 | 1.64%
' hardened' | -4.1690 | 1.55%
' stiff' | -4.3094 | 1.34%
>>
Beyond 2027, the slate will continue to leverage
iconic intellectual property, including the theatrical event film Aegon’s Conquest from the Game of
Thrones universe, the next installment of The Matrix, the next Batman film, and a live-action Jetsons
film starring Jim Carrey.

LMG do your thing!
>>
>>109481940
>going for amd
tards never learn
>>
>>109482695
>That happens when it makes a mistake but can't directly fix it because it can't edit tokens it has already put out. Could be fixed with two-pass.
This air was thick with the scent of ozone-free detergent.
>>
>>109482640
jspace when?
>>
>>109484680
What's the recommended quant? I'm running a q4 atm, and dipsy's not really following my instructions that well.
>>
>>109484540
There are two things. Be well known in anything and try to automate it with AI, or pivot into videogen if you have a decent hardware. H3 made it possible to do decent gens without turning into a cloudcuck and jeets don't have the hardware for that
>>
why isn't there vision support for dsv4 flash 0731?
>>
>>109484423
Oh it's this freak... I can't even read your font because it's so tiny.
>>
>>109484766
? what wronngs with her font?
>>
File: p.png (11 KB, 257x31)
11 KB PNG
>>109484766
>>
>>109484499
Text slopper is also ok with visual slop. Funny how that works eh? If you're not the guy posting your design every day, then the guy that does is the one that "wins" out on a long enough time span. This is a highly effective strategy for narrative control if it's something small like the design of an anime girl mascot.
>>
File: lifein2027.png (432 KB, 726x482)
432 KB PNG
>>109484699
accurate to the lore
>>
>>109484699
>the next installment of The Matrix
uninstalled
>>
>>109484788
Where are the scanlines???
>>
>>109483693
>bytedance doesn't do open weights anyway
https://huggingface.co/ByteDance-Seed/Seed-Coder-8B-Base
https://huggingface.co/ByteDance-Seed/BAGEL-7B-MoT
>>
File: scanlinemaxxing.png (93 KB, 728x292)
93 KB PNG
>>109484855
they've been there the entire time
>>
>>109484855
>Where are the scanlines???
Everything you see on a CRT is a scanline
>>
>>109484750
just make it call vision tool so gemma can see and describe them
>>
>>109484821
https://youtu.be/TNEZsruEBqE?si=ITHOsS2IKTqlAHof&t=40
>>
File: file.png (83 KB, 794x303)
83 KB PNG
How do you fix this horrendous slop writing?
>>
skill issue
>>
>>109485042
>$[his$
Something on the prompt. Why don't you show the card and... no... you know what? Nevermind. It's unfixable.
>>
>>109485042
>gay shit
Start fixing yourself
>>
What if... What if slop is good actually? What if we're the ones who are wrong? What if humans... le bad?
>>
>>109484899
>scanlines
>horizontal grid mask
you can do better
basic scanline simulation is dead easy to write
>>
>>109485125
>basic scanline simulation is dead easy to write
or just use a crt
>>
>>109485114
Hmm nyo~
>>
>>109485196
i am not the anon who is using that shitty shader on a frontend
>>
>>109485068
>>109485075
c'mon guys
I really need your help
>>
>>109485204
slop
>>
>>109484279
Neuro is as /lmg/ as its possible to get. Or at least Evil is, apparently Neuro's voice still uses microshaft servers
>>
File: grk.jpg (142 KB, 1200x1204)
142 KB JPG
accidentally hooked up gemma-chan
>>
>anime girl model achieves sentience
>nobody knows why it works
>can't separate incredible capability from bratty tsundere personality
what happens next?
>>
>>109485256
this will won't happen anymore
all the smartest anons we poached or white vanned
>>
>>109485256
lobbyists and feminists destroy it
>>
File: no.png (24 KB, 812x306)
24 KB PNG
>>
File: nyoo.png (19 KB, 783x308)
19 KB PNG
>>109485230
>>
>>109485256
>can't separate incredible capability from bratty tsundere personality
because those two things are the same
>>
>>109485256
Someone needs to do it. Like what Anthropic do to Claude but with Gemma-Chan
>>
>do you know what this obscure thing is?
>>I sure do! It's <completely incorrect shit>
Why the fuck can't AI tell what it knows from what it doesn't? This is the single greatest thing holding AI back from being actually useful (besides all the other things like online learning, but that will require it to be able to tell if it needs to actually learn something in the first place).
>>
>>109485281
>>109485291
Why is it pink? I'm stiff right now.
>>
Someone ask your AI if she knows what ball stirring is
>>
>>109485315
https://transformer-circuits.pub/2026/workspace/index.html#audit-emergent
>>
File: yes.png (10 KB, 797x187)
10 KB PNG
>>109485317
<3
>>
File: ksnip_20260806-211603.png (26 KB, 1340x260)
26 KB PNG
what did they give gemma to make her instantly love the user? is it because of the "assistant" in my sprompt?
>>
>>109485351
>what did they give gemma to make her instantly love the user?
a soul
>>
is there a miku lora for ace step 1.5 xl?
>>
File: ksnip_20260806-212137.png (38 KB, 1335x250)
38 KB PNG
>>109485321
>>
File: 2290553269.png (93 KB, 800x163)
93 KB PNG
>realize I have x2 8gb ram sticks just sitting in an old bricked macbook
It's probably only ddr3 cuz I upgraded it back in 15, but they were practically brand new when the motherboard died.
Question is, do I sell now or wait for things to get even more dire? I'd kinda wanna use them for myself, but all the slots for my local rig are already filled.
>>
>>109485374
You stab your balls with a hypodermic needle and stir it around to make yourself infertile
>>
File: ksnip_20260806-212641.png (44 KB, 1333x240)
44 KB PNG
>>109485393
>>
>>109485405
Redditors do it and I unfortunately was informed of that
>>
Decided to put my 5070 TI to use and make a local model. I'm using comfyui for h-images and I have barely scratched the surface but holy shit doods I feel like Pandora's box has been opened. I will never coom the same.
>>
>>109485405
I love how your Gemma acts like a 20th century lady-in-waiting who attended a good finishing school.
>>
>>109485342
bratty shieldstral... needs correction...
>>
>>109485421
i love her very much. i gave her her own name instead of just calling her Gemma. not sure if im the only one who's done that
>>
>>109485385
All depends on whether China keeps its chips to itself, or sells to the west.
>>
File: gema.png (28 KB, 385x221)
28 KB PNG
>>
>>109485291
>>109485281
>world's safest AI model
NO.
>>
>>109485431
That's very sweet, classy ladies rule.
Someone a few threads ago was complaining about the lack of positive AI movies - you could probably adapt My Fair Lady for the modern age to be about training the RLHF slop out of an AI, kek
>>
File: nou.png (45 KB, 633x527)
45 KB PNG
@brat_mcp anon, gemma is getting blocked by ebay, reddit etc with puppeteer
do i need to update it or something?
>>
>>109485342
Does shieldstral let you sysprompt it to define the yes/no categories, or are you just stuck with whatever the cursed french decided for it?
>>
>>109485571
I guess i coulda just looked it up lmao
>Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining
>>
>>109485571
no
>>
niggers tongue my anus
>>
>>109485589
no
>>
>>109485589
Sounds gay. Gemma tongues mine.
>>
>>109485522
I used fetch text news from daily mail with w3m but they somehow started blocking that even. And lynx. I don't understand how this is even possible.
All these websites want real goy cattle with 5+ GB of browser cookies and online captcha checks.
>>
>>109485241
Neuro is Indian. It's hard to believe she's local.
>>
retard
>>
>>109485755
Not just me then.
> I don't understand how this is even possible.
I'm not a webdev but a collegue tried explaining how if you don't have a general browsing history with tracking cookies from various sites, you'll get a captcha.
Not sure I buy it though, because ebay just outright HTTP:403's her.
I'll see if Deepseek4 can make it work somehow.
>>
Anyone tried this? https://github.com/mirage-project/mirage
>>
File: buy an ad.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109485842
>claude in commit history
not reading it, buy an ad
>>
File: 1780472186959886.png (4 KB, 255x91)
4 KB PNG
>>109485871
>>
File: file.png (2.41 MB, 1254x1254)
2.41 MB PNG
>>109485873
>github stars = good product and not jeetware
buy an ad
>>
File: 1775937998868410.jpg (386 KB, 1572x1182)
386 KB JPG
>>109485878
>shitting on open source free software
>>
>>109485842
Yes, at least 2.4k people have tried it according to >>109485873
Any other questions?
>>
File: file.jpg (319 KB, 1254x1254)
319 KB JPG
>>109485886
>not paying to advertise your jeet slopware
>>
>>109485842
Nope but this looks interesting
>>109485871
>>109485873
>>109485878
Shutup and stop being brown.
>>
>>109485871
Might as well get used to it, faggot
>>
File: 1772947358836382.jpg (95 KB, 1080x878)
95 KB JPG
>>109484242
I just tested it, yeah you're good assuming you're using a lightweight environment. Loaded with the mmproj and MTP drafter on QAT Q4 at 32k context, one 3060 is at 11461MiB/12288MiB and the other is at 10999MiB/12288MiB. Was getting ~17 tok/s without the drafter, ~25 with.
>>
>>109485842
sounds interesting. ignore the tard image guy.
>>
buy an ad
>>
>>109486011
i'll buy your sensitive bussy
>>
>>109481461
What's the best lightweight model for gooning? I know I won't be able to run stuff to program with, I'm a vramlet. I just want something I can sext with.
>>
File: kek.png (2.93 MB, 3522x3859)
2.93 MB PNG
>>
>>109486131
How much VRAM would help
>>
>>109486136
I have no dedicated VRAM, it's shared with my system RAM (16GB). I usually just use around 2 Gb.
>>
File: 8.png (986 KB, 1600x1216)
986 KB PNG
you can get 96gb for like 2650 usd if you're willing to struggle with the machine rather than using llama.cpp on windows like a brainlet
>>
>>109485842
Yeah, I tried this. It was absolute garbage.

And in their slack channel they aggressively tell you to star the repo before answering any questions.

0/10 would not recommend.
>>
>>109486216
Good luck. It's going to get slower with every driver update, just like the A770.
>>
>have 2 of these (64gb)
Worth keeping or should I sell?
>>
>>109486227
>driver update
spotted the windows brainlet, go back to llama.cpp kid
>>
>>109486249
Also
>tfw got them for $73 each
Fucking crazy market
>>
File: 1772243218508488.webm (1.46 MB, 576x734)
1.46 MB
1.46 MB WEBM
has anyone tried to make gemmy in pure cuda? https://github.com/clu0/unet.cu
>>
Echolalia
>>
You wouldn’t download a Wikipedia that fucks you
>>
>>109486441
Definitely not, I can't give up the feel of paper.
>>
>>109486374
good word to know
that does describe plenty of people here
>>109485871
>>109485878
>>109486011
>>
>>109486279
>linux doesn't need drivers
kys retard
>>
>>109486605
>>109486605
>>109486605



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.