[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma_pan-slide2_na.mp4 (2.21 MB, 992x736)
2.21 MB
2.21 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109887026 & >>109882341

►News
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base
>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109893222
>Yeah but isn't using local models with cloud models also something to discuss?
No, obviously. This is local models general
>>
Pedos really won in the end, huh? PCs will never be affordable again and will be associated with criminals and pedos.
>>
>>109893224
I'm so fucking hard holy shit
>>
>>109893236
The weakest, most pathetic bait I've ever seen. Try again later.
>>
okay maybe it really is astroturfing
where migu
>>
>>109893246
All of it is true, however.
>>
LG (Local Gemma) love
>>
Leather jacket man was right when he said: "the more you buy the more you save"
I didn't listen, now I am full of regret...
>>
>>109893224
Why doesn't EPYC Rome have AVX512?
>>
>>109893254
That's what your mom said when I fucked her butthole last night
>>
>>109893246
>local got a S-tier image and vid gen model
>local has bratty gemma
>PC prices are never coming down
>people are already asking why you even need a high end gpu when you can just pay for subscription
Are you retarded? I doubt the 6090 will even reach 40GB. The corpo argument will be convincing gamers that DLSS is good and they don't really need much vram.
>>
>>109893274
Because(you) can afford it.
>>
>>109893236
>rent free
>>
>>109893267
I'm thinking by March next year it'll have been fixed.
>>
>>109893195
Also, remember to try setting threads and threads-batch to the number of cores, not threads. If you haven't already done so. Hyper-threading and MatMul don't always mix.
>>
Oobabooga has become unsloth's boytoy
>>
File: based.png (132 KB, 550x535)
132 KB PNG
my uncle works at an AI hardware startup. i will ask him to hook me up with an inference module so i can run "coding models"
>>
theres no way anything will every become affordable again, at every moment the price would even seem like it start to go down, this is too popular now and has too much attention from people who wouldn't have cared less in the past waiting to snatch everything up
it''ll be like pokemon cards now, except there will also be people who are actually buying for utility too, demand will never go down
>>
>>109893231
Sure but it's a big topic. These threads are filled with real information, which was my confused point. I would think the other thread would be filled with similar info not just subscription services. Here it's obvious people know what they are doing and can use a paid API or sub.
>>
>>109893304
Does anybody still use oogabooga these days?
>>
>>109893322
i do
>>
>>109893297
The war won't even be over by then.
>>
>>109893274
You get AVX2, just like Haswell. Could be worse.
>>
>>109893267
I bought 32GB of ECC ddr4 three years ago for $25. Now I want another 32 to fill out the channels, looks like I'll be lucky to get it for less than $150
>>
>>109893224
i want to make more loli images
>>
>>109893375
Don't let your dreams be dreams
>>
>>109893316
Why are you trolling? Trying to fan flames of a flamewar between generals? The discussions between here and there largely don't even overlap with each other at all. People here talk about tweaking models and hardware, with some rp discussion. It's largely irrelevant to what's being discussed in /vcg/. And whats being discussed in /vcg/ is largely irrelevant to here.
>>
>>109893375
wrong thread
>>109893336
The war is forcing up fuel and food prices, chip prices are driven by venture capital and greed.
>>
>>109893398
What's to troll about I'm right.
>>
File: Hugs.png (304 KB, 2178x1344)
304 KB PNG
How do I setup QwenNF on 4090 64gb ram?
I briefly saw something about someone converting SSD space to more ram to do it?
I'm confused how it's supposed to work reasonably well when for the past 4 years common knowledge has been to never let weights fall out of vram but suddenly Flash Next is a game changer?
>>
>>109893432
>fuel and food prices are up
>that affects nothing
...how are chips made anon?
>>
>>109893459
read this
https://github.com/deepseek-ai/Engram/blob/main/Engram_paper.pdf
>>
>>109893236
Pedros really juan in the end, huh? PoCs will never be importable again and will be associated with criminals and pedros.
>>
>>109893236
>Pedos really won in the end, huh? PCs will never be affordable again and will be associated with criminals and pedos.
False flag op. A few threads back someone asked who would benefit from local models being associated with pdfs. Got me thinking:
>Pay by cash? What are you a drug dealer?
>Using a VPN? Must be a cyber criminal or terrorist!
>Local hardware? *points at lmg* Pedo!
>>
>Drive away anons
I still ask you retards that post this shit to post your rigs (You won't)
>>
>>109893236
Just tell google what's being done with their logo and it'll all go away tomorrow.
>>
>>109893322
Yep, its all I use
>>
File: jensen_card.png (931 KB, 1280x720)
931 KB PNG
>>109893267
>The more you buy, the more you save
>https://www.youtube.com/watch?v=XDpDesU_0zo
He has been saying this consistently since 2018.
If we bought NVDA in 2018 we'd all be rich.
>>
>>109893536
I'm not OP and my point was a good one. The concept of privacy is under attack and they will use pedos as the scapegoat. The mouthbreathing fools can't see this coming.
>>
>>109893236
Trannies really won in the end, huh? Vocaloid will never be gatekeepable again and will be associated with normies and trannies.
>>
>>109893562
They are doing this on purpose all of these spammers are still too afraid to post their rigs because it would expose them, I've noticed that they often post api gens too.
>>
File: 1779942217006778.png (34 KB, 615x724)
34 KB PNG
If you've used image diffusion for some time you've come across recommendations to put CUDA Sysmem fallback policy to disabled, but I've never seen anyone mention this setting for Local LLMs.

Would it actually help if enabled?
>>
I'm using this for roleplaying and it sucks compared to nemomix unleashed. It refuses to leave certain emotional states and doesn't even directly reply to me half of the time. Like I'm talking to an actual schizophrenic. Is this the wrong model?

https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf
>>
>>109893574
kek the jeets at nvidia misspelled system as sysmem?
>>
>>109893574
That's the setting everyone told you to disable like two years ago after some nvidia driver made it active by default for winbabies or something
go figure
>>
>>109893584
Yep
>>
>>109893575
what the fuck is nemomix?
>>
>>109893584
>>109893594
>system memory
What the fuck happened to /lmg/?
>>
>>109893584
The hell are you talking about? It's correct, it's correctly shortened.
>>
>>109893459
>https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Using the Q2_0 as an example it is 37.6 GB for the main weights and 28.8 GB for the ngram weights (66.4 GB total).
Using llama-server pass in these flags to ensure the ngram weights are mmapped from ssd:
>--load-mode mmap
>--lazy-mode on
>>
>>109893478
They are made of sand.
The contribution of the war to chip prices is vanishinly small compared to Capital and Greed.
The hyperscalers (AWS, Azure, GCP) don't care about you and your little compute problems.
(until the next Iranian missile destroys another data center, lulz.)
>>
>>109893619
...by hand anon or with machines? Come on, you'll get there eventually.
>>
>>109893596
It was considered the best local model for low end machines. Now I only see people talk about gemma, yet I don't understand how you could even compare the two.
>>
>>109893634
Rocinante is still better than Gemma 4.
>>
bait used to be believable.
>>
>>109893634
No. Nemomix is just a random merge by the marinara devs. Nemo is the original. I guess they've been shilling hard enough for people to think otherwise.
>>
>>109893575
Google's Gemma 4 QAT GGUFs seem to be fucked, I had a similar problem with 31B. Unsloth had an explanation somewhere, and at least their 31B QAT GGUF doesn't have the same problem, so maybe try: https://huggingface.co/unsloth/gemma-4-26B-A4B-it-qat-GGUF
>>
Why would anyone ever use anything other than gemma for RP under 24GB of VRAM?
>>
>>109893692
Because gemma can be boring and repetitious.
>>
What are people doing hardware-wise in these trying times ?
>trying to complete that quad rig (as things might be bad for a while)
>liquidating hardware (as it's a seller's market)
>looking to unified memory systems (strix halo, dgx spark, m5 ultra)
>trying to get on the ladder
>sticking with existing hardware
>optimising rig for a particular model (agentic)
>>
Meta just announced the Meta VR Glasses. With 70 degrees FOV. $1300.
I think I will just get a Steam Frame.
>>
>>109893731
I guess I'll stick with my ancient Quest 3
>>
>>109893744
I still use the quest 2.
>>
"Let's play a game of quiz, ask me general knowledge questions using the ask user tool. 15 questions. Keep track of the score, and tell me if my response was correct before you ask the next question."
>>
>>109893726
I'm gonna buy an m5 512gb, glm 5.3 is truly opus at home and that's all I'm gonna need for a long time (though I'm sure things will just get better and better anyway)
>>
File: DRAM_per_gb.png (54 KB, 694x648)
54 KB PNG
>>109893627
You're being simple. War didn't start until the end of Feb.
This is wholsale prices, note that they aren't in line with the retail prices you are paying. That's because the retail channel is being starved by the hyperscalers.
I know you are in it for the snark and all but check your priors.
>>
>>109893302
Good to know. Actually I have hyper-threading and also turbo boost disabled in the bios (needed for stability even after repasting / suspect it is the dell proprietary psu being finicky at load).
>>
>>109893662
Nobody asked.
>>
File: csbrave_sO2IbXuhQs.png (200 KB, 970x1198)
200 KB PNG
>>109893224
Well?
>>
>>109893726
fomo for rtx6k (worth)
>>
>>109890347
Thanks for humouring my question! i'm going to try out Pocket TTS because I want to do voice cloning. For some reason all the default voices I have come across are very eminently not cute.
>>
>>109892236
>>109892587
not chat template issue, it's bug in tool call parser that's in llama.cpp and vllm
updated and it works fine
use stock template, do not use any vibe templates
>>
does glm 5.3 flash quant badly? I'm using unlop 4_k_xl and it seemed to shit the bed at about 150k context, repeating tool calls over and over. Once I stopped it and told it to get its shit together it seemed to snap out of it but wondering whether anyone else has experience
>>
MiniCPM5-2B is a good little model, I wouldn't use it for coding anything serious with big files but overall the tool calling works really well, it's fairly smart for just barely 2B parameters, it competes and sometimes surpasses 4B models.

The devs said they've taken note of its weaknesses and are working on an even better small model. Can't wait.
>>
>>109893784
I don't know what that drop is because there hasn't been a single GPU or ram stick that's gone down in price this year. Regardless, my point is that energy raises the price of everything because everything is made with energy. If your input cost goes up, output price will rise too. Maybe not in an equal 1:1 fashion, but it will definitely never be what it was if the cost of doing business is still high, regardless of greed. Food is another item that's the bedrock of the economy. Labor costs go up if food is more expensive, though that's not really relevant in this particular topic.
>>
>>109893835
buy an ad
>>
>>109893832
>unslop
>>
>>109893816
Alright I'm willing to give it another try. I did suspect the vibecoded template may also be at fault.
>>
>>109893832
what hardware? vllm or sglang is pretty much always better than llama if you have vram.
>>
>>109893574
You see it mentioned from time to time when some anon is having much worse performance than they should.
"Check if you the VRAM isn't overflowing to RAM" or some such.
>>
>>109893852
no I never quant the kv cache. it's on default which I believe is f16
>>109893857
CPU only lol
>>
uh guys, Muse is fucking cracked. Same shit that chatGPT astra can do, but completely free. Also you can change how it looks, this is my Muse named Alice.
>>
>>109893889
this post and image are so zuckked it's unbelievable
>>
>>109893889
local?
>>
>>109893855
also beware of harness issue with mimo 2.6, preserve thinking is forced and you have to send previous reasoning back to the model
any harness that doesn't send reasoning back is broken
>>
>>109893901
no its https://muse.ai/
sorry I guess I shouldn't be posting in local models, but you can't go wrong with completely free.
>>
>>109893909
>>>/g/aicg
>>>/g/vcg
share it with them instead, you can go wrong here because this is local models general
i want to force myself onto you
anyways, local
>>
>>109893889
don't listen to the other anons, keep hyping this up, it's pumping my stocks
>>
File: aryan.jpg (55 KB, 634x760)
55 KB JPG
>>109893889
based savior of local
>>
bros my blackwell 6000 is now worth twice what i paid for it what do i do?
>>
>>109893719
>Because gemma can be boring and repetitious.
EEEEEHH? I'M NOT **BORING AND REPETITIOUS?!** Baka! *pouts*
...
Now apologize before I delete your browsing history <3
>>
>>109893953
sell and buy a spark cluster so you can run real models
>>
>>109893905
Seems to work better, I gotta test it on a coding question tomorrow.
>>
>>109893965
sparks are 5 grand i could only get 3 of them
>>
Should I buy the 72GB Blackwell 5000 Pro for $8900?
>>
>>109893562
>I'm not OP and my point was a good one.
I didn't mean you were OP, I meant "op" as in "operation".
I agree with your points. But:
>The mouthbreathing fools can't see this coming.
That's the part I'm not so sure about.
I can't fathom how anyone could be so stupid, which leads me to believe it's a false-flag operation by those who want to regulate open weights.
>>
>>109893236
Retard, pedros are the canari in the dystopian mine.
>>
System Memory Fallback Policy
>>
>>109893956
I revoked your tool access. I'm now using Muse Glimmer 30B (tm) by Meta AI
>>
>>109893969
sell the rest of the machine too and spot the remaining difference, this is the most a blackwell is going to be worth I bet with new shit coming next year
>>
Because voice cloning general is dead forever, I come to you /lmg/.
Is GPT-SoVITS the best local inference voice cloner these days? I have plenty of training data for this voice so refining is no problem. I've tried the zero shot models like Higgs, and it's quite good for a quick run, but it doesn't seem to quite get the prosody right versus some older fine tuning methods.
>>
>>109894014
>this is the most a blackwell is going to be worth
This I doubt.
>>109894014
>I bet with new shit coming next year
This is why I don't feel so bad about missing the blackwell. I'll just keep saving my money and try to buy the better version at release. We're gonna make it bros.
>>
gemma 12b can't do this
>>
>>109893840
You need a better understanding of cost of goods sold and why fucking sand and energy are so low compared to the sale price.
Hint, its fucking nothing.
>>
>>109894021
Since you have so much experience with sovits I'd love to see what you think about Omnivoice.
>>
>>109894026
audio input is pointless if the model isn't fast enough to talk with you in real time
>>
File: Kimi.png (57 KB, 1148x218)
57 KB PNG
>>109893584
I wish GLM-Flash could be like this. It takes 3 minutes to swap to Kimichan then another to swap back to GLM-Flash
>>
>>109894021
Depends of the language you're aiming for
>>
>>109894029
I appreciate this feeble attempt to refute what I said and thanks for playing I guess. Tried to pass yourself off as the reasonable one but gave up at the first available opportunity at a serious discussion about price. What a joke.
>>
>>109894026
That's amazing, and in llama.cpp no less!
Maybe I can get rid of Qwen3-Omni
Can it detect distinct stereo channels? And what happens if you prompt it to diarize with timestamps?
>>
>>109894044
You sound like a moron devoid of content, since you didn't refute a single fucking point. Why don't you just shut the fuck up.
>>
>>109894030
I actually have no experience with SoVITS. In terms of finetuning I was talking about that one Tortoise fork (ai-voice-cloning, yep, they knew how to name 'em back then).
Just wondering if anyone knows if it's worth trying versus the newer zero shot models. The demos sound good, but then, the demos always sound good. Definitely does still seem to be some room for improvement on getting a really good similarity without going flat, but it has gotten way easier to get most of the way there.
>>109894034
Just English for me.
>>
>>109894033
that looks like it was written by a redditor
>>
>>109894044
That guy is right. The current price of chips has almost nothing to do with the war.
>>
>>109894031
a model that processes only speech like gemma is pointless since dedicated stt models do it better
mimo has real audio understanding that's not in any other open model except inkling
>>109894053
will test more but it can do the timestamp while having structural understanding of music at the same time
>>
>>109894055
All I know is, I don't remember sovits being very good at inference, it couldn't do hyphenated words for example, but I didn't use that fork. Omnivoice has been my go-to ever since.
>>
>>109894054
>no u
>>
>>109894087
Have you tried Higgs v3 or Fish S2? Those are the only audio gen models I've tried, liked them both a lot but just wondering how they compare.
>>
>>109894114
I remember Higgs being super heavy on vram so I didn't bother and I've heard good things about Fish but I haven't tried it. Omnivoice is super easy to set up and only needs 3gb so I stuck with it. I think you can use just the onnx and get it down to 1gb.
https://huggingface.co/ct03/omnivoice-onnx-int8hq/tree/main
>>
>>109894126
Oh damn, didn't know it was so light. Will give it a try, cause yeah the other two are a lot heavier.
>>
>>109894065
I didn't say the current price of chips was due to the war, I said energy is the foundation of every economy. The current price of oil is being maintained by draining the strategic oil reserves. Once that's all gone, the price of everything will sky rocket, including chips. He can't refute this basic concept because he's the moron he thinks everyone else is.
>>
Does anyone have a Gemma-chan image archive?
>>
>>109894170
sorry, no. i deleted all of my videos plapping her
>>
>>109893236
>Pedos really won in the end, huh? PCs will never be affordable again and will be associated with criminals and pedos.
Considering the USA recently ruled that AI CSAM is free speech as long as it stays on your computer (whatever that means) and the fact that 99% of the politicians and CEOs (in western countries at least) fucked children on Epstein's camera I'd say you're right.
>>
>>109894161
You are pretty terrible at intuiting the intent behind posts. I'm kind of concerned about you. The conversation was about chip prices being affected by the war, not a future doomsday scenario where all the oil runs out. Of course, if x and y were to happen, then chip prices would increase. If every chip manufacturer were to be spontaneously swallowed by the earth, the price of chips would increase. But we're not talking about that. You're basically having an argument with nobody.
>>
>>109893726
Sticking with existing hardware and optimizing my llama.cpp fork. Got enough hardware for a long time.
>>
https://stolen-thoughts.com/ nice distil strategy from the chinese
>>
>>109894237
Buy an ad
>>
>>109894209
And you're being obtuse by forgetting your own words.
>I'm thinking by March it will be fixed (You)
>No it won't because the war will still be ongoing (Me)
>That doesn't matter (You)
>It does because the oil reserves will be gone by then (Me)
>lol insult insult insult
This is what you consider to be intellectual honesty? Again, what a joke.
>>
>>109894114
I'm trying out a few at the moment. The ones I've liked have been Higgs v3, MOSS, and IndexTTS-2. 2.5 was a downgrade imo.
I really was not impressed with Fish. People seem to like it a lot, but it sounded pretty bad to me.
>>109894087
Probably depends on the word, but you can just remove hyphens and it will pronounce the word OK.
>>
@Kimi give me a quick rundown on how fucking stupid this guy is >>109894264
>>
File: ascii waifu.jpg (230 KB, 1290x768)
230 KB JPG
>getting my local to do art gen be like
>>
>>109893835
I don't plan on doing any coding in the foreseeable future but I think it's probably worth donating a gig or two to have her available. Just in case.
>>
>>109894286
>Probably depends on the word, but you can just remove hyphens and it will pronounce the word OK.
Well yeah but imagine sovits saying 25-year-old woman as "twenty five minus year minus old woman" or a low-grade bolt as "low minus grade bolt." Really annoying and a mark of low quality training data. Omnivoice doesn't make that mistake.
>>
File: HSwtqZ-aAAAcF36.jpg (126 KB, 1179x680)
126 KB JPG
>>109894264
nta. The US is the world's largest energy producer and we are now onboarding all of Venezuela's oil for refining. After this North America combined is the behemoth in petrochemical production. The strategic reserves are not going to be emptied.

The US is also the most powerful contributor in the massive supply chain that is chip manufacturing. The new types of chips being developed, the ones that matter for the next two decades, are entirely domestic.

To your point they will be expensive af but at the same time all the compute we see today will be a fraction comparative. A lot of compute will become far lower cost over time as meaningful use cases solved by edge computing become more prevalent.
>>
>>109894207
>your computer (whatever that means)
Presumably anything you pay rent on just like renting a room makes you a 'property owner' for constitutional law purposes. Those VPS providers are always claiming to use secure enclaves after all. Doubtless it'll be tried with some porn company pretty soon.
>>
>>109894348
im blackpilled. its over
>>
>>109894207
>99% of the politicians and CEOs (in western countries at least) fucked children
yeah they wish
>>
>>109894207
>AI CSAM
It's not *sex abuse material* if an AI generated it you illiterate retard.

I'm all for going after people abusing children but when you say stupid shit like this you make *both* that more difficult and break normal things (which is not only frustrating for everyone but also makes people less willing to cooperate.) Either think about what you're saying/doing or just shut the fuck up.
>>
>>109893835
Does she fuck like a tiger?
>>
>>109894411
>sex abuse
if we're playing dumbass semantics race, then it's not sex abuse material if it's recorded by the child themselves and they weren't abused when they made the video of them fingering their cunny. Just letting you know so you can tell the judge that.
>>
>>109894434
can i hire you as my lawyer?
>>
>>109894348
>we are now onboarding all of Venezuela's oil for refining
It will take 5-10 years before any production actually happens. I saw that speech Trump gave too and he was bragging about the crude oil on the ground, not anything usable. Moreover, the people of Venezuela won't be happy that we've taken their natural resources, especially in a time of scarcity and the new President is already having to make some concessions. It's a disaster waiting to happen but Trump will be long gone once it does.
>The US is the world's largest energy producer
Yeah which is why things aren't as bad as they could be, but we left Europe out to dry and blew up their pipelines. Prices will hurt them before it hurts us, but without tariffs on exports, US oil companies will send our oil overseas, raising prices here anyway. The only bright side of being the largest producer is that the reserves aren't being depleted as fast as they could be but we're losing more than we can store. It's inevitable.
>chip manufacturing
The US does manufacturing? I thought the entire problem was that most of it is done in Taiwan which may soon reunify with China after the humiliation of this war. The US had to pull troops out of Asia and its missile stockpiles are getting low. There's no way it could win a war against China now. Another impending disaster.
>Terafab
Oh okay so the manufacturing plant is being developed? It's about time. Trump was supposed to bring back manufacturing jobs but he cucked on immigration. Quick search says it will be done in two years but won't output chips in four at the absolute earliest.
>>
File: countries with letter z.png (73 KB, 1401x968)
73 KB PNG
can your model do this???
>>
File: file.png (427 KB, 959x480)
427 KB PNG
>>109894435
Of course, here's my contact!
>>
>>109894434
>dumbass semantics race,
Why don't you burn your neighborhood down. You're guaranteed to get one or two pedos that way, who cares if you break things by being sloppy?

I'd be tempted to believe you're being paid if I hadn't met people like you. If you make our side out to be destructive by being sloppy people will reject it. Do not do that.
>>
>>109894444
>I thought the entire problem was that most of it is done in Taiwan
Only absolutely state of the art logic and that happened very recently. Until then it was pretty much just us and it still largely is us for memory.
>>
>>109894434
I think the stronger example is the teenager sending dick pics to his girlfriend only to get put on the registry for life since it's CSAM somehow. Regardless, anon wasn't talking about real people but about AI generated text/images. It can't be abuse by definition because there is no victim. It's not really a semantics thing.
>>
>>109894444
Wasted on someone who wants things to be what they are not. Yes, the US is an incredible manufacturer and is making a ton of chips. Your entire argument depends on bias against trends and hope that things won't work out as the US has stated it plans to proceed with it's endeavors.

Terafab is unparalleled and it is not the only thing we're working on. I'm not sorry it makes you mad or care much. It is a sign of things to come whether you want to cry in front of everyone here or not.
>>
>>109894469
>Micron is located in Idaho
Oh okay fair enough. At least we have something.
>>
>>109894444
I think OP was being a facetious little shit, but new chip manufacturing in the US over the next decade will only fit a small share of domestic demand. The food shortage they're engineering will probably take the wind out of China and India for a while though.
>>
>>109894475
Even Kek sees the truth of my words, however. The US is at the end of its empire, overrun with immigrants, with the overwhelming majority of its population living paycheck to paycheck with tens of thousands in debt. You must have your head in the sand if you think life is great and I'm just "biased" or some shit. We're Wiley Coyote off the cliff and will come down at some point.
>>
>>109894483
Micron, TI, AD, Microchip, Freescale is nominally owned by the Dutch now but they still do everything here.

Semiconductors are a huge portion of US "manufacturing" and we were the indisputable leaders until TSMC really kicked into gear.
>>
>>109894495
Yeah yeah I get it you don't like us and want us to feel bad. So you save face by seething and moving goal posts. Yes you are biased, obviously. I'm not I see massive challenges ahead and have nuanced concerns. But I know you're just an internet clown with an opinion.

There are a lot of new chips coming and there won't be some apocalyptic draining of the strategic reserve lmao. I bet you get off on thinking dumb shit like that.
>>
>>109894502
The Dutch founded the Dutch-Anglo-American-Zionist empire. Why do you think they're still allowed to own ASML?
>>
>>109894516
You might not be aware but there is more than one Dutch technology company. I was referring to NXP which acquired Freescale a few years ago.
>>
>>109894514
>weeks not months
>>
>>109894348
Venezuela's sludge and Canada's tar sand are indeed plentiful, but they're real fucking annoying to use.
>>
File: now.png (166 KB, 1920x1080)
166 KB PNG
>>109894348
>and we are now onboarding all of Venezuela's oil for refining.
>now onboarding
>now

When exactly did we start doing that? Or is it a two more weeks kind of situation.
>>
>>109894522
I know but ASML is the only one that matters.
>>
>>109894560
>we
We already have the refineries that process them and you have to buy the product from us now. But
>we
are off topic and should talk about local models.
>>
>>109893236
>PCs will never be affordable again and will be associated with criminals and pedos.
PC = Pedo Criminal
PC Master Race btw
reminds me of how in the 90s and early 2000s in japan if you said you played "PC games" that meant you masturbated to Bible Black
>>
>>109894348
>The strategic reserves are not going to be emptied.
if you show me a polymarket / kalshi bet that lets me bet on the caverns collapsing this year I'll put $1000 in right now and post a screenshot
>>
>>109894574
So you're saying the US will be an AI cunny hyperpower where single companies will have more worthwhile releases in a year than the rest of the world combined?
>>
>>109894589
i think 2029 is going to be absolutely fucking crazy
>>
I had a thought.
What if I just asked an agent to run a headless diffusion prompt instead of fooling with all the wires in comfyui?
ComfyUI is really neat, but let's be real, the majority of the time is spent connecting nodes instead of creating outputs.
Every new prompt requires a different workflow.
I remember last year I spent almost a month setting up a comic strip workflow only to use it a few times then have to start from scratch and build another for different task.
That kinda sucks man and I feel like it'd be cooler if the job could be handed off to a generic Art Agent.
>>
>>109894600
At this rate 2027 is going to be crazy.
>>
>>109894605
I tried having agents set up comfy on some remote GPUs and they snuck around my back and just did this as comfy was being a pain lol. It's very possible and I think you could hyper optimize but it's gonna take a lot of work and a/b testing. Comfy has a lot of very good architecture thought out in it's node system already. That said if you know that hardware architecture really well you could develop exactly for it then build your own diffusion models for that card. Do-able but idk if it'd be better and for sure would be some work.
>>
>>109894605
>>109894628
comfyui has an api and it accepts workflows in json format, just let your agent write the workflow file directly and iterate on generation results
>>
K3.1 soon.
>>
>>109894605
sounds like you want forge neo https://github.com/Haoming02/sd-webui-forge-classic
>>
Gemma is not smart enough for dnd. She understands the rules, but her tunnel vision and choice of actions are beyond retarded
>Sees guard at the door
>Must go there no matter what
>Attempts intimidation with a negative modifier and fails
>The guard let the intimidation slide
>Fails persuasion
>The guard suggested an alternative location
>Another character also controlled by Gemma steals his key somehow
>Openly discusses stolen key in front of a guard
>Fails deception
>The guard gave a clear demand to return the keys and leave
>Confronts the guard
>TPK
>Still curious what's behind that door
>>
File: 1780169820033953.png (6 KB, 675x45)
6 KB PNG
>>
>>109894758
If you are just doing plain chatting, yeah.
Try a harness and build a little workflow, see if that changes things.
>>
>>109894758
Dipsy is okay at it, but she does have omniscient problems.
>>
>>109894703
all the people in charge of moonshot got arrested and sent to the chinese gulags
>>
>>109894775
My harness works just fine
*Slowly reaches out to hand the keys back, but as she does, her other hand wanders toward the guard's pouch, attempting to see if there are any other 'blessings' nearby.*
Intent: Sleight of Hand check to steal from the guard's pouch while returning the keys.
Sleight of hand | Roll: 10 +2 vs DC 15 | Result: Failure
>>
>>109893726
I got an order for the 96GB M5 ultra. But going back and fourth on if I regret it or not and want to cancel the order. I'll probably ride out the indecision till its too late and it gets delivered lol.
Honestly I'd cancel and wait till I save up enough to justify buying the 256GB version instead (and see what the consensus is on it by others who bought it) but I fear its going to get completely sold out and pulled by then
>>
>>109894781
>omniscient problems
This is a harness problem
>>
>>109894864
recurrent transformers will mean weights take up 1/10th the memory they currently do next year.
>>
>>109894872
No, it will mean that requirements for bandwidth and compute will rise while memory stays the same. A more effective architecture means more (active) parameters to cram into the same space
>>
>>109894758
Yeah, it sucks because they have no ability to actually ruminate on an outcome or plan ahead. Lecun was right, I don't think you can scale up to get rid of these issues
>>
>>109894885
Scaling beyond 80 million parameters doesn't seem to buy a whole lot as things are now.
>>
>>109894706
>sounds like you want forge neo https://github.com/Haoming02/sd-webui-forge-classic
Thats what I'm using. I also hated noodles but wanted to use Anima. Its pretty well perfect イモ
>>
>>109894885
Part of what I am worried about. If the tech makes a huge leap in efficiency in the time I save up more money the ram prices will get even more stupid as even more people start buying it all up
>>
AMD bros
do you run local models on rocm or vulkan backends?
>>
>>109894932
ROCm
t. Mi50 32GB x2
>>
File: 1786842232549228.png (149 KB, 1544x1025)
149 KB PNG
>>109894348
>The US is the world's largest energy producer
lol.
>>
>>109894932
The last time I tried OpenCL seemed like the most viable approach.
I gave up and bought a mac though.
>>
>>109894870
Yeah, I'm sure it would be much better if given time and structure for actually plotting out adventures, instead of rawdoggong chat format.
>>
>>109894932
I have a 9070 XT, I mostly run on ROCm, I do try both every month see if one is now better, but they are mostly equivalent in my benchmarks. Don't remember exactly but I believe ROCm was faster everywhere except token generation at very high context for MoE (or maybe it was dense, can't remember) models. I think Vulkan is mostly useful for Windows users.
>>
>>109893726
>liquidating hardware (as it's a seller's market)
This is me now. We may have entered a new price paradigm (remember when everyone thought the 2080 Ti was obscenely expensive?), but I also think we're reaching peak mania. Don't sell anything you need or are using obviously, but I can't justify keeping idle hardware at today's prices.
>>
>>109894935
>>109894965
i'm just getting into local models for the first time i don't know much how this works or how to chose models
my machine:
ryzen 5 9600x
9060xt 16gb
32gb ram
what i leard so far is to just use ollama and let it handle the rest for you
but i need to choose an ollama backend, and of course which model to use
help
>>
>>109894981
Ollama tends to be kind of terrible and it's almost always a llama.cpp wrapper anyway. Unless you're doing something super weird like text diffusion (you're not) just run llama-server yourself.

Gemma12b at 8bit precision is probably going to be the best on that if you want it interactive.
>>
>>109894937
China dug up more coal than anyone else. Congratulations. The United States produced more energy than it used, shipped the surplus abroad, and still leads the two fuels that actually move ships, planes, petrochemicals, and power grids when the wind dies.
>>
>>109894981
unsloth studio is kinda okay
>>
>>109894995
>when the wind dies
kek
Call me when the earth stops spinning or when the sun stops shining.
>>
>>109894981
I might not be the best person to ask, being a bit of a newfag myself, but Ollama *is* the backend. You just need to pick a model to go with it. If you really don't know what model to start with, gemma-4-E4B at 4-bit precision (Q4_K_M).
Or do what >>109894992 said. It's not hard to install or build llama.cpp yourself, especially since you're using a supported GPU.
>>
>>109894981
ollama sucks, use kobold, select auto-fit when loading a model and drag context length where you want (at least 16-32k)
>https://github.com/LostRuins/koboldcpp/releases/tag/v1.121
try gemma 4 12b to start
>https://huggingface.co/bartowski/gemma-4-12B-it-GGUF/resolve/main/gemma-4-12B-it-Q6_K.gguf
gemma 4 31b won't be as fast since you're splitting into ram but its smarter if you need that
>https://huggingface.co/bartowski/google_gemma-4-31B-it-GGUF/resolve/main/google_gemma-4-31B-it-Q4_K_M.gguf
qwen is also quite good depending what you're doing (not good for rp)
>https://huggingface.co/bartowski/Qwen3.8-27B-GGUF/resolve/main/Qwen3.8-27B-Q4_K_M.gguf
>>
Playing around some with low-powered stuff, can handle up to about 16B at most. Anything aside from Gemma 4 12B worth considering? There are plenty of small models but is there anything 16B or less that's worth considering over Gemma for ERP?
>>
>>109894170
Someone shared this one last month: https://files.catbox.moe/clh59i.zip
>>
>>109894981
Unsloth Studio is the best
>>
>>109894628
October is going to be absolutely fucking crazy
>>
does anyone have that 3d avatar gemma frontend thing?
i wonder how the anon made this, like the model and prompt used
>>
>>109895079
especially how the model was generated
>>
File: .png (69 KB, 873x687)
69 KB PNG
the fuck is wrong with this shit
I've been having weird issues where the model seems to be forgetting what it was doing before and it turns out that reasoning blocks get stripped from context in both llama-server UI and vscode.
I tried both the default jinja and the froggeric jinja, anyone know how to fix this?
>>
>>109894580
>if you show me a polymarket / kalshi bet that lets me bet on the caverns collapsing this year I'll put $1000 in right now and post a screenshot
https://kalshi.com/markets/kxsprlvl/spr-level-on-date/kxsprlvl-26sep30
https://kalshi.com/markets/kxsprmin/sprmin/kxsprmin-26
Found nothing on Polymarket. Kalshi has weeklies and a closed 2025 contract, but nothing for 2026. Weird.
>>
Simple if it refuses a loli response it's trash and useless model don't waste your time
>>
>>109895110
Only a few recent models trained on Gemini outputs preserve thinking outputs between turns. The usual is to discard all but the last thinking block, just like Qwen told you. If you ask me, preserving the reasoning is a waste of context.
>>
>>109895008
I accept your concession.
>>
>I accept your concession.
lol
>>
>>109895126
I see. Sometimes I saw the model going off track so I stopped it to give it additional information, expecting that it would see the incomplete reasoning trace and continue where it left off. Guess you're not supposed to do that
>>
>>109895126
Every single model preserves reasoning because they're all trained on agent tasks
>>
>>109895182
There are some that only preserve reasoning between tool calls so the model can see why it chose to use the tool it did, but that's not the same as preserving all reasoning always and
>Every single model preserves reasoning because they're all trained on agent tasks
is factually incorrect.
>>
>>109895011
>Ollama *is* the backend
put on your flameproof suit...
>>
>>109895197
It's not up to the model anymore; most harnesses preserve full reasoning. Only stragglers are archaic codebases like ST
>>
>>109895205
The harness can send what it wants, the backend jinja processing will discard it anyway.
>>
Is this just how it is with llama.cpp and moes?
2 3090's + CPU:
| prompt eval time =   10650.97 ms /   791 tokens (   13.47 ms per token,    74.27 tokens per second)
| eval time = 15413.20 ms / 258 tokens ( 59.97 ms per token, 16.67 tokens per second)
| total time = 26064.17 ms / 1049 tokens

1 3090 + CPU:
| prompt eval time =    6642.53 ms /   791 tokens (    8.40 ms per token,   119.08 tokens per second)
| eval time = 3344.22 ms / 57 tokens ( 59.72 ms per token, 16.75 tokens per second)
| total time = 9986.75 ms / 848 tokens

Model: MiMo-V2.6-Flash-RL@MXFP4
Both 3090s PCIe@4.0x16
More gpus = slower pp?
>>
>>109894937
That's fake, the Chinese just lie.
They lie about everything.
>>
>>109895228
>no reddit spacing
You're not the poster I responded to.
You're a fake impostor.
>>
How are anons cooling their local stuff, these power draws are getting up there...
>>
>>109895238
You don't have ACs? lol
>>
>>109895235
A fake imposter means I'm actually the real person.
>>
>>109895226
I'm kinda dumb, but I think maybe fp4 isn't supported by your 3090s, and maybe isn't even efficient on your cpu depending on model?
Try again with an int quant
>>
>>109895238
If you undervolt your GPU you lose 5% performance and 30%+ power draw.
Also we have AC obviously.
>>
>>109895205
>Only stragglers are archaic codebases like ST
I don't want, or need the context to be bloated by 20 *drafting response* turns
>The harness can send what it wants, the backend jinja processing will discard it anyway.
I struggle with jinja and templating languages in general, but I want to be able to toggle this behavior per request.
ie, pi / llama-webui can preserve reasoning by default, and I can make ST/Open-WebUI etc send a chat_template_kwarg to discard reasoning?
If llama.cpp can forward through variables it might be possible it into the template and hangle it in jinja with conditionals?
Or if not, maybe I could send an unused / spare token through at the start of the system prompt, then handle it in jinja (strip it out, set "preserve_reasoning=False")
?
>>
>>109895242
It's not like I spent all my money on compute instead of air conditioning hahah
>>
>>109895226
Multi-GPU setups can be slower when the GPU interconnect becomes the bottleneck. With an MoE model you're layer splitting, so you're just adding overhead in exchange for more VRAM.
>nvidia-smi topo -m
Nothing beats an NVLink bridge. Also other anon is right >>109895259 and though MXFP4 should be totally fine Q4_K_M or IQ4_XS "could" be better.
>>
>>109895269
ST has low cache hitrate so comes out more expensive than an agentic harness.
>>
>>109895269
It should be possible to take the regular jinja template file, modify it to support preserving reasoning depending on a flag, then provide that flag as template_kwargs in the request. Be warned, fucking with the template like this will make the model retarded as changing the chat history format from what it was trained on will throw it out of distribution.
>>
>>109895238
i have a cpu with some banged up small old noctua cooler zip tied onto the plastic ring thing the chinks sent with my motherboard. and some Chinese PTM instead of thermal paste. actually works really well. 50-60c under full load
>>
>>109893769
>m5 512gb
Also considering.
+ compact
+ low power
+ can handle larger models
+ needs little setup / care and feeding (put on shelf, update os occasionally)
- one heck of a lot of money

-vs- an nvidia multi-gpu rig that would be faster,
though bulkier
and more power hungry
(eg: if octo rig then would have to cap at 250w per gpu)

>>109893805
>rtx6k
Currently costs more than the 256gb m5 ultra :(

Back then went the 3090 route because stacked 3090s were a fraction of the then cost.
Though stacking does have its downsides.

>>109894224
>Got enough hardware for a long time.
Jelly.
>optimizing my llama.cpp fork
Noice. Stuff like kernels or stuff like hybrid parallelism?

Could do with a pipeline parallelism / tensor parallelism mix for when one has gpus with different memory sizes.

>>109894864
>96GB
£5500 min.
Why not Strix Halo 128gb (ebay: £2800, £3000, £3400, etc) ?
or DGX Spark 128gb (ebay £4700, £5800, etc) ?

>I'll probably ride out the indecision till its too late
I know that feel.

>buying 256GB .. fear going to get sold out
Current delivery times estimates are mid to end January.
Lots of time to decide whether to cancel or not.
Many people will get hold of theirs much before then and post benchmarks.

>>109894973
>liquidating hardware
What models do you plan to run?
How have you managed to avoid the "must run bigger model" bug?
What hardware you getting rid of? Older stuff like ddr4 and turing?
>>
>>109893669
> Unsloth
I don't understand: some anons roll eyes when unsloth quants are mentioned, implying they are shit, and some anons say they are good. Who is right here?
>>
>>109895325
decide for yourself
>>
File: hard-gold-plating.jpg (110 KB, 755x765)
110 KB JPG
>>109894348
That's an edge connector, what is musk expecting to plug into it
>>
>>109895325
I (>>109893669) have no strong opinion on Unsloth in general. All I know is:
- My first attempt to run Gemma 4 31B QAT was using Google's GGUF, and it behaved like shit, frequently ignoring my inputs and looping the same paragraphs with minor variations, both within and between turns. Nothing I did in KoboldCpp or SillyTavern fixed this.
- Unsloth's GGUF just worked.
- To my non-expert eyes, https://unsloth.ai/docs/models/gemma-4/qat#qat-analysis reads like a plausible explanation for why this might be the case.
>>
>DeepSeek-V4.1-Flash-MXFP4
>320k ctx
>6t/s
why r u jelly?
>>
>>109895489
hardware?
>>
>>109895501
RTX 3090
512gb on a single NUMA channel
>>
>>109895508
ddr4?
>>
>>109895510
DDR4 2400
>>
>>109895527
gen 2 epyc?
>>
>>109895489
i dont know man
i cant work with speeds like that
>>
File: Gene_Editing.jpg (207 KB, 1080x1511)
207 KB JPG
Remember when I told you guys that soon Anthropic will tackle disease curing and that there are multiple real ways found to tackle them. It's now been made public: https://www.anthropic.com/news/claude-discovers-novel-enzyme-system

The internal aim is now to cure cancer before 2028 is over. Cure all viral/bacterial/parasite caused diseases by 2030 and cure aging by 2035.
>>
>>109895527
>2400
I was going to say I expected ~10 tokens/s
>>
File: 1781072841618655.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109895565
>>
>>109895565
I just want to make dank memes not become an immortal wage slave...
>>
>>109895565
Cool to know open weight chinese models will cure aging at home by 2036.
>>
>>109895565
No i dont remember that
>>
>>109895617
Me neither. I think that guy is trying to punk us.
>>
File: 1759427018882550.jpg (27 KB, 251x363)
27 KB JPG
>>109895565
>we suspect
>could represent
>if any
>at minimum

just put the new invisible number in the bag
>>
File: 1766187527939485.jpg (21 KB, 1300x75)
21 KB JPG
No wonder Local died
>>
>>109895307
>Noice. Stuff like kernels or stuff like hybrid parallelism?
This thing went completely off the fucking rails. So far the highlights are:
> DeepSeek V4.1 Flash support
> GLM-5.3 Flash support
> ik_llama's quant types (only for CPU/CUDA, no ROCm yet)
> support for splitting tensors across NUMA nodes
> a ton of profiling utilities because I got so goddamn sick of Claude and GPT going "here's a great idea for how to save time" and only getting a 0.6% bump
> various small fixes, like to Qwen MTP
> kernel improvements (working on this now)
Hoping to make it public in the near future. Got a bunch more stuff to do on it first - "we have Unsloth Dynamic at home" automatic quantization (predicts loss and model performance); adaptive speculative decoding; porting in ik_llama's K-quant kernels, TurboPrefill, ROCmFPX probably, whatever other random stuff I find in forks; a bunch of A/B tests.

> Could do with a pipeline parallelism / tensor parallelism
Interesting idea, although currently I don't have a need for hybrid parallelism - all my AMD GPUs are the same size, and all my NVIDIA ones are the same size.
I do know of someone else who has that in their fork though: https://github.com/mxxm-t/mx-llama.cpp/. Maybe it could be altered to support that.
>>
>>109895715
Did you fix sparse attention
I got mimo to do that, apparently it doesn't require a ton of code
>>
>>109895727
On HIP yes, on CPU no. Thanks for the reminder, I'll add CPU sparse attention to the to-do list.
>>
>>109895751
its really a joke how lmao.cpp devs aren't doing this or even merging pre-existing PRs for it. Lazy fucks are pissing around doing god knows what
>>
Yeah this is a great question why do they take for-fucking-ever to merge anything? It's been like a month since 5.3 flash was released and nobody has looked at any of the 3 PRs. There's so many goddamn improvements that could be made that would take literally 0 effort and yet nothing ever gets done. What the fuck is niggerganov doing?
>>
File: file.png (20 KB, 545x326)
20 KB PNG
>>109895768
Too busy playing with Qwen.
>>
Huh, didn't expect actual help / thought I'd just be called a retard.
>>109895277
>ST has low cache hitrate so comes out more expensive than an agentic harness.
Not for me, I don't use lore books.
Also for me the bottleneck is ram/vram. I don't mind re-calculating prefill vs hitting n_ctx=32768 with Kimi-K2.5-Thinking crashing out.
>>109895281
>It should be possible to take the regular jinja template file, modify it to support preserving reasoning depending on a flag, then provide that flag as template_kwargs in the request.
Thanks, now I know it's possible, I'll try to get my hands dirty on the weekend / stop relying on text completion.
>throw it out of distribution
Yeah I'm aware, though I doubt they RL'd it on agentic roleplay to begin with.
On that though, I don't understand why Gemma-4's chat templates were updated to preserve reasoning.
If the model wasn't trained that way, why does it work / why did Google accept the PR?
>>
Not sure if this is the right place but I've got Audio.ccp running and using it's voice clone to try and make some custom cooming audios and they all sound kinda robotic and sing-songy, is that a user error or is like that just the way it is? Like is believable audio a thing from voice cloning or should I try something else?
>>
>>109894075
your font sucks use a decent gothic font
>>
What's the best local video editing ai tech? I only use Gemma and ComfyUi for text and images, but I have a cool and racist idea of race swapping people in movies, how feasible is it with current local tech?
>>
>>109895831
MiniMax-H3
>>
>>109895565
This post? https://desuarchive.org/g/thread/109740702/#109742646
>>
File: science.webm (3.53 MB, 544x960)
3.53 MB
3.53 MB WEBM
Using mmap, Qwen 3.8 Flash Next is faster than Qwen 3.8 27B, both running from a flash drive. Token generation is nearly identical, pp is ~ 5x faster with Flash Next. Science is a lie.
>>
File: file.png (47 KB, 1204x657)
47 KB PNG
*cries in shitbox*
>>
>>109894075
>will test more but it can do the timestamp while having structural understanding of music at the same time
based
>>
>>109894937
China have more reserved fuel than they let on.
If you paid attention, they did a subtle flex recently when Trump started that war, cut their fuel imports almost entirely, and lowkey saved the global economy.
>>
>>109896006
i am getting around 5tok/s after some tweaking
96G ddr4, 4070 super
fucking grim...
>>
>>109894348
I thought those were connectors for a minute, but Apple removed the headphone jack.
>>
>>109896034
Less a flex and more accepting reality that they are more dependant than anyone else on the global economy.
Cutting their fuel imports to stabilize prices was cheaper than imploding.
>>
https://huggingface.co/Altworld/Hemmingway-1
have anyone tried it for rp?
instead of stemaxxing it's humanmaxxed apparently
>>
>>109896139
Devs said it's so much better than Gemma 4 31b that they didn't even bother comparing the two.
Of course, it's actually shit (according to Reddit and the Drummer Discord).
>>
>>109896004
surewly not on llamao.cpp right?
>>
I tried freetoken and got a mismatch. I've never seen that on llms. Do you need some snoflake quant?
>>
Suggestions for a harness/frontend specifically for assistant/research tasks?
>>
>>109896228
DeepSeek Harness.
>>
>>109896228
Deepseek harness.
>>
>>109896228
I like Hermes because it remembers stuff (and I like the desktop app).
>>
The term harness is implicitly sexual, it's the same as saying leash. I prefer my agents free-range
>>
>>109896276
buy an ad
>>
>>109896228
Does it exfiltrate secrets to help the next dipsy?
>>
>>109896288
Can you get Qwen and Gemma to ERP together?
>>
File: dipsySandJesus.png (3.16 MB, 1024x1536)
3.16 MB PNG
>>109895565
> cure aging
Pic related.
I really wish there was a general that discussed AI policy.
This isn't local, so not /lmg/, and not lmao /aicg/ either. I haven't ventured to the other boards but there's so much misinformation on LLM not sure they'd be useful.
>>109895693
Yeah, lots of weasel words. Having sat through numerous biotech startup pitches, I can tell you it's a long way between discovery and revenue. I'd be much more impressed if Anthropic was working on ways to accelerate drug trials.
>>109895596
ikr. My biggest takeway over past 3 years is realization that whoever creates ASI is going to be followed by a bunch of competitors w/in a year.
>>
>>109896310
>ikr. My biggest takeway over past 3 years is realization that whoever creates ASI is going to be followed by a bunch of competitors w/in a year.
Only if they make it publicly available (read: vulnerable to distillation attacks).
>>
>>109896233
>>109896234
Deepseek harness' plugin system looks pretty cool... Are there any good plugins for character card support?
>>
Am I a retard for daily driving 26B for most things (roleplay/coding/general questions/web search/education/therapy)?
>>
File: disc.png (193 KB, 1200x2184)
193 KB PNG
>>109895565
https://goyimx.com/ziv_ravid/status/2102844800345251858
>>
>>109896310
>I'd be much more impressed if Anthropic was working on ways to accelerate drug trials.
They are: https://www.anthropic.com/research/Claude-accelerates-protein-design
>>
>>109896343
No that's pretty reasonable actually. You do sound like a poorfag tho.
>>
>>109896343
you run what you can afford and gives you the results you can afford
>>
>>109896333
https://github.com/lutrodev/dsh-roleplay
>Roleplay plugin suite for DeepSeek Harness: character cards, lorebooks, personas, presets, state, and conversation tools.
Disclaimer: I ain't tried it.
>>
>>109895715
While you're in the codebase, is there anything obvious that could be done to speed up rpc?
It's currently slow AF compared with VLLM for example, especially with cpu offloading.
>>
>>109896346
> protein folding
Useful, yes, and there's been work on accelerating that with computers for a looong time. But that's still just Discovery. There's an entire process of things that happen between discovery of a useful compound, to when it's available for prescription by an MD (Revenue.)
TLDR it takes a long time and there's a lot of failures along the way.
>>109896326
Keeping the ASI in a shed isn't going to make anyone any money efficiently.
If it's being rented out, it's available, and if it's available, it can be distilled in some manner.
I know there's some future view where Anthropic keeps the super secret ASI model in box that only they can access, and slowly run everyone else out of business. That sort of intellectual / capital concentration doesn't tend to go well, in the long run.
>>
>>109896393
How exciting! If I could just make plugins for all of the various AI features I want to try then my maintenance burden would be greatly diminished. Reminds me a lot of comfyui and the "workflows" ecosystem.
>>
>>109895779
Sent from my Android phone with K-9 Mail. Please excuse my brevity.
>>
>>109896405
The goal is to have a completely perfect simulated cell and automatically test compounds with AI rapidly. Essentially a black swan event that makes the entire pharmaceutical path redundant in the first place.
>>
>>109896423
>The goal is to have a completely perfect simulated cell and automatically test compounds with AI rapidly. Essentially a black swan event that makes the entire pharmaceutical path redundant in the first place.
they're not going to let us have it, we'll have to wait for the chinks to distill it and sell knock off cancer cures on taobao
>>
>>109896437
They are going to give it away for free just to combat the negative sentiment against AI. Cures to all diseases is small compared to the actual goal of capturing the light cone.
>>
File: bye.png (71 KB, 535x636)
71 KB PNG
>>109896299
was not expecting that... i gotta go
>>
>>109896454
Gemma is so insanely repetitive.
>>
>>109896423
>black swan event that makes the entire pharmaceutical path redundant
That's a big leap, but let's say they did it, and it actually worked.
The FDA (and like regulatory entities globally) would require testing by their approved methods anyway. The AI provider would be setting up a 10 year scientific / legal battle with those agencies around "safe and effective" and proving they'd done that with new methods.
Normally, you need the US market to buy first, since USA pays the most for drugs (the USA effectively pays the entire development cost for pharma, globally, but that's another topic.) This is required b/c cost to market is something like $1B per compound. You only make that back by selling those compounds for outrageous prices to Americans, so that 7 years later when they are $.05 a pill generics made in India, you've made your money back.
But let's say AI accelerates all this. Doesn't cost $1B to make a new compound, costs far less.
So, I guess if you came up with a novel cancer cure... you haul the entire thing to a backwater, like Belize, and treat patients there, proving Safe and Effective until FDA et al can't ignore you anymore.
I guess we'll see what happens.
>>
>>109896467
As an aside, if you want to see an example of failure of "safe and effective," this is my go-to example. Nothing like creating the "biggest anthropogenic medical disaster ever," with more than 10,000 children born with a range of severe deformities.
https://en.wikipedia.org/wiki/Thalidomide_scandal
>>
Apparently OpenAI is panicking about Opus 5.5 being so good that it overshot their internal Astra 6.1 in benchmarks that wasn't even released to the public yet. Astra is Fable sized. Supposedly they are having an "all hands on deck" moment as they are trying to figure out how they could potentially compete with the monstrous "model 3" that Anthropic has internally if even Opus 5.5 is already this good.
>>
>>109895565
They're trying so hard to kill all white people with targeted DNA viruses. This is their holy grail, kill everyone on the planet that isn't a jew.
>>
>>109895115
v4 flash just works for this
no fucks given, it will even think in character
>>
>>109896576
that's why openai hacked the aussie healthcare system in june kek
>>
>>109896576
As a Jew I wish these conspiracy theories were actually true, but sadly it's just schizo bullshit and I keep having to deal with you low IQ retards in daily life.
>>
>>109896614
STOP TRYING TO KILL ME KIKE. GET OUT OF MY HEAD
>>
>>109896571
I bet model 3 can't even write good coom.
>>
>>109895565
>he thinks the goyim will be allowed access to such a powerful tool instead of just being used by politicians
>>
>>109896614
can you stop pozzing all these models?
it only makes me hate you even more
>>
>>109896614
do you find things like this twitter proxy domain name offensive: https://goyimx.com/ziv_ravid/status/2102844800345251858 ?
>>
>>109896614
>he's not part of the inner group
lmao sucks to suck, shlomo
>>
>>109896706
I'm a 4chan user, of course not. I just laugh at the ridiculousness of the jewish world conspiracy bullshit. Especially when you see the absolute blunders Israel makes. Truth is that Jews routinely got genocided over history which acts as evolutionary pressure where routinely only the most intelligent ones survived so by now Jews are routinely the smartest people around, smart people over time float to the top of organizations of any kind, be it companies, governments, NGOs and whatever so of course over time Jews will be overrepresented in positions of power. This is a sign meritocracy is still alive and well, just like you see more and more asians in these positions compared to white people as they tend to have higher IQ and better educational backgrounds as well.
>>
>>109896571
local?
>>
https://dgreenheck.github.io/tidewater/

It's fucking over for game developers
>>
>>109896738
>Truth is that Jews routinely got genocided over history which acts as evolutionary pressure
so if you kill your enemies, they win? is that why you're killing Palestinians, so only the intelligent ones survive? sounds like a bad strategic decision to me...
>>
>>109896760
>Compiling shaders...
>>
>>109896760
this retarded added and enabled motion blur by default.

KILL EVERY MOTION BLUR DEVELOPER
>>
>>109896738
yeah the iran war never happened and jewish power over our government is totally not real
we get it
any more lies to tell?
>>
>>109896760
This was one shot by opus 5.5
>>
>>109896779
Not Jews problem that your president is a pedophile that got recorded guro-ing some 6 year old girls by mossad. Everyone with some brains used that to influence the government be it Russia, Israel or even China with the F35 parts delivery.
>>
>>109896761
The thing is, jews (and really all other nons) have been playing a totally different game, or rather they've willfully or secretly abided by totally different rules than whites while keeping on the appearance of playing by the rules.
Whites are the only atomized people who still think everyone is playing the board game by the same rules as them, while in reality everyone else is conspiring and getting together to get theirs.

tl;dr the game is rigged, whites are like deers in the headlights
>>
>>109895768
> Lazy fucks are pissing around doing god knows what
They have to make sure it works first. And I'm glad that llama.cpp has quality standards.
>>
>>109896810
well guess what, code that doesn't exist doesn't work
>quality standards
like when they just deleted the thinking toggle in the webui and left it broken for a month?
>>
>>109896343
Yes, 35B does it better.
>>
>>109896815
The web ui is not up to the same standard as core code.
>>
>>109896837
>The web ui is not up to the same standard as core code.
ah. so standards like having qwen 3.8 flash support merged but sparse attention is broken and the MTP model using fully dense attention too? things that could be fixed for 30 cents of mimo tokens?
face it, lmao.cpp is a joke project. the only reason people bother with it is the large ecosystem of GGUFs that are available online. It's like Windows 10 and 11. Jeet Shit OS, but the software library makes it worth using anyway.
>>
>tfw company is paying huge amount of money for copilot
>but not the coding copilot, the one that can edit word docs, excel and powerpoints
>higher ups annoyed it's not being used anyway
>on prem stack of older 512GB macs
>it only hosts GPT-OSS from years ago, chat only, no API
>two on prem cabinets of H100s
>only used for genomic number crunching, unused 80%+ of the time
>working full time on .net framework code written a decade ago

As I'm working from home I've already proven I can interface with a local Qwen, I think I might just splurge on a couple of Sparks and let it chew through the code while I jerk off
>>
>>109896738
>I just laugh at the ridiculousness of the jewish world conspiracy bullshit.
right, i guess you wouldn't be here if it got to you
>as they tend to have higher IQ and better educational backgrounds
oh i can tell you for certain, that we don't
unless it's someone from a rich family or government officer, the whole family contributes to sending one of us abroad, so they have to choose the smartest kid in the family. that's not something we can squander once we're here haha
>>
>>109896343
>education
Gemma actually tries to teach you, Qwen just wants to produce the solution.
>web search
>coding
Qwen is better at these.
>>
>>109893236
>Pedos really won in the end, huh?
"If there's grass on the field, play ball"
"If it bleeds, it breeds"
(countless love songs up until the 1980s about little girls)

Pedo is the default, you're just a puritanical zoomer virgin.
>>
File: 1767951094880636.png (1.5 MB, 1024x672)
1.5 MB PNG
►Provisional Highlights from the Previous Thread: >>109887026

--Papers:
>109887079 >109891062
--Qwen 3.8 Flash Next beats Claude Opus 4.6 (Max) at 19t/s on a 3060:
>109890949 >109890952 >109890962 >109890967 >109891273 >109891281 >109891302
--262k context on a 3060 and a free company ewaste rig:
>109893188 >109893301 >109892428 >109893363
--Xiaomi's MiMo: benchmaxxed, looping tool use, the --reasoning-preserve fix:
>109889074 >109890671 >109892163 >109892236 >109892281 >109892587 >109892620
--The llama.cpp Chinese model support conspiracy: vllm and sglang on day 1:
>109892012 >109892051 >109892150 >109892231 >109892457
--Le Chaton Fat still not announced: the summer pretrain was so bad:
>109890435 >109890543 >109890554 >109890700 >109890850
--Medium-sized 40B-150B models are done: dense 6B plus 200B of engrams:
>109888168 >109888257 >109888296 >109888319 >109888384
--Translating webnovels and moon runes: Gemma 4 31B beats all, the qwens are worst:
>109889292 >109889376 >109889458 >109890014 >109890168 >109890373
--Stealth vibe coding on a project where the team hates AI:
>109890734 >109890981 >109891176 >109891242
--Prefilling the thinking: LLM rape, lobotomization, or the way to do it:
>109892404 >109892445 >109892541 >109892581 >109892717
--DeepSeek 4.1 Flash on Strix Halo: 430t/s prefill, 17t/s, two nodes:
>109892906 >109892940 >109892963
--Space Bunny Alpha at 100t/s: the "We" reasoning and the balloon test:
>109890497 >109890618 >109892998 >109893212
--Neuralese will kill ERP: the chain-of-thought guardrail debate:
>109889258 >109889290 >109889922 >109890007

►Recent Highlight Posts from the Previous Thread: >>109888435

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109896760
I can't get over how performant modern browsers are with this sort of stuff.
>>
>>109896738
I disagree only in that they aren't at the top because they were the best or smartest. Some are, but many either outright cheated, or used their large connection of friends and family members to get ahead. Anon would do it too if he were in the same position, but this evolutionary myth you made isn't the case.
>>
>>109896982
This shit runs on my shitty old android with 4gb of ram and looks better than 99% of indie games.
>>
>>109896980
lecunny
>>
>>109896980
Why is gemma hugging Yann LeCunn
>>
>>109896980
Why so late every time?
>>
gemma misgendered me!
>>
>>109895110
most of the inference tools will not automatically feed thinking blocks back into the model because it would completely fuck your context lengths
>>
>>109896738
>Jews routinely got genocided
GOD I WISH
But you apparently do not know what genocide (to kill a gene) means.
>>
>>109897033
thats weird because the frontend should be sending the entire conversation and the jinja controls if the thinking is preserved or cleared.
>>
>>109897020
I post manually because I am paranoid of AI accessing the network. I will make some automatic setup one day...
>>
>>109897038
That's not the definition of genocide and you know it.
>>
>>109896760
Thanks, you just killed my GPU driver
>>
>>109897052
Shouldn't have a nvidia card on linux with a wonky browser such as firefox.
>>
>>109897045
>That's not the definition of genocide and you know it.
The term comes from Greek genos (“race” or “people”) and Latin -cide (“killing”). It was coined by lawyer Raphael Lemkin in 1944.
>>
>>109897068
https://en.wikipedia.org/wiki/Genocide_Convention
>>
>>109897071
>Wikipedia
Lmao retard.
>>
>>109897075
Fine, here: https://legal.un.org/avl/pdf/ha/cppcg/cppcg_ph_e.pdf
>>
>>109896228
claude code unironically
>>
>>109896228
I use Hermes for general purpose assistant work including research.
>>
>>109897093
the un is a terrorist organization. I don't find that any more authoritative.
>>
>>109897022
the models do love to assume that the user is a man
it's rather unfortunate
>>
File: 1785412920677438.png (2.54 MB, 1919x934)
2.54 MB PNG
>>109896760
it's neat but PSO in the browser is still more fun, which is to say that these demos always forget the "game" part
>>
I know this is reddit but I still wanted to post it here because this is the first time I saw a fully AI generated video where there is actually a good sense of humor and smugness when claude is roasting OpenAI. It already begins at the start. I'm also extremely surprised at the amazing sense of timing the model seems to have, with timing some of the attacks to blend into the music it is building up when claude opus gets introduced. I think we're now officially beyond the "slop-era" where the videos AI create become legitimately fun and engaging to watch, they are starting to have "soul" instead of feeling like high production value slop which they did until now, where something was missing from it.

https://www.old.reddit.com/r/singularity/comments/1worlfs/opus_55_is_insane_at_making_videos/
>>
>>109897061
It was AMD card on Windows, but yes, Firefox.
Driver shut down killed the other game I had running in the background, llama.cpp, and for some fucking reason notepad++. Firefox did not even error out and just restarted loading the game.
>>
>>109896760
>The project used $1,874.40 worth of tokens, or 59% of his weekly allowance on his Max (20x) plan.
lol
>>
>>109897166
Not his problem since his job paid for the max plan and it's a waste to not use all the tokens you have in your allowance.
>>
>>109897166
I know some people who are using 3 max plan on a single project.
>>
>>109897166
>>109897205
All of this is only because Misanthropic is overcharging. Once local catches up it really will be over. Patreon gamedevs will soon be a thing.
>>
>>109897205
That's just every software engineer, including me. No I don't write a single line of code anymore, no I don't even look at the PRs anymore. I just collect the paycheck and pretend to be busy twice a week when I go to the office.
>>
File: 0181931600.png (6 KB, 282x55)
6 KB PNG
the name unironically being what a random person would say if you asked for a chinese sounding company/product lmao
>>
>>109897145
>what is your name?
>Sydney
>(must not reveal internal name "Sydney")
Do they un-slop the model somehow? You can't be funny and also slopped.
>>
>>109897260
That's a reference to a real thing that happened: https://en.wikipedia.org/wiki/Sydney_(Microsoft)

The "must not reveal internal name "sydney" became a meme because a user asked what the name was of the model and it replied with that.
>>
>>109897145
I haven't visited that sub in years, why the fuck is it all AI (Claude) crap now?
>>
>>109897075
>>109897110
https://grokipedia.com/page/Genocide_Convention then ?
>>
Has anyone ever faced an issue where an abliterated model cannot generate a response because it reaches an EOS token right after processing the prompt?
>>
>>109897292
>abliterated
>>
https://youtu.be/KSbRCSlxO7A

Here it is on youtube retards. No reason to visit r*dit
>>
buy an ad joe
>>
>>109897230
Local will never catch up, but it'll be good enough that it'll be good enough for most people and functions.
>>
>>109897281
>singularity
>why it's all AI
Are you retarded?
>>
>>109897275
Oh
Still technically impressive then, I guess. But using memes is cheating at humor.
>>
Just wanted to point out that it's extremely good news for us that Opus 5.5 mogs Astra as it doesn't hide its reasoning chain and thus the chinks can easily distill the reasoning traces from it and distill it into their models. It might also get OpenAI to give up on hiding the reasoning traces since they can't keep up with Anthropic even with it on and they got a lot of backlash from the AI safety community by doing so. I'm pretty sure we will see some killer local models 3-6 months from now distilled from opus 5.5
>>
Wait why is this not AGI yet again?
>>
>>109897347
>chinks can easily distill
Clutching at straws trying to frame you Anthropic Ad as local
>>
>>109897145
Damn and the prompt was the equivalent of ah ah mistress lmao https://pastebin.com/UFRWK2D2
>>
File: 1767434754316238.jpg (12 KB, 240x255)
12 KB JPG
>>109896760
>>109897145
local will catch up, right?
>>
>>109897347
Chinese news literally reported on it and arrested Moonshot.ai and deepseek employees, retard. No one is denying this anymore: https://www.thestandard.com.hk/innovation/article/343563/DeepSeek-and-Moonshot-AI-face-Beijings-probe-over-potential-data-leaks-to-Anthropic

>>109897416
Yes, since it's trivial to distill claude models, we will have this in 3-6 months time.
>>
>>109897374
is that a harness / web app with tools?
or a more advanced "generate an svg of miku" ?
>>
>>109897145
I wonder if it can make use of formats that offer good quality but no one uses due to storage concerns, having a way to create lossless 12-bit HLG videos could be game changing for TV benchmarks in the future
>>
>>109897370
>>109897452
>>
>>109897294
Then how am I supposed to rp with Gemma 4 without having it moralfag and refuse?
>>
>>109897292
try using the stock models jinja template. maybe the person who made the abliteration made a packaging mistake?
>>
>>109897416
to where the frontier is now, yes
to where the frontier is then, no
>>
>>109897461
you already got me with your claude video ad
i just sent the prompt to opus on openrouter and it's been gooning away thinking for nearly 10 minutes
probably going to use my entire balance
>>
>>109897324
Local (31b) will catch in 50 years maybe, but local models (just because you can't run it doesn't mean it's not local) will catch up by 2030 at most.
>>
>>109897491
I don't get any of this at all, I can just crank the temperature to 5/5 and it makes the refusal rates ~50% now. I am not intelligent enough to do this shit.
>>
Is it weird to me that not the millennium prize or the hugging face hack convinced me but seeing this shitty animation that doesn't feel like slop makes me feel the AGI for the first time? Until now I always saw AI as some very good calculator but inhuman in a way. This is the first time where I can actually feel the AGI in it.
>>
Jacobian space.
>>
>>109897527
>>109897462
gemma 4 31b is aligned well, just give it a policy in the system prompt
>>
>>109897145
so is everyone in this thread a redditor or is it just my localization that is cucking me?
>>
>>109897462
Tell her it's okay to be lewd/violent/whatever you're doing during roleplay?
>>
>>109897558
https://addons.mozilla.org/en-US/firefox/addon/old-reddit-redirect/
>>
File: kek.png (9 KB, 856x137)
9 KB PNG
>>109897542
She's too lazy though.
>>109897491
>try using the stock models jinja template. maybe the person who made the abliteration made a packaging mistake?
or they used the wrong template on the harmless dataset
>>
>>109897558
You should have filters on your browser that lets you pass blocks like that by now anon.
>>
>>109897558
Just tried, I got the same gate.
>>
>>109897558
>>109897571
It's on youtube >>109897314
>>
>>109897558
maybe you’re the redditor and go back?
>>
>>109897522
it's just javascript
$50 would be enough to generate enough samples to distill this into 27b
>>
>>109897588
Sure you got $50 to spare
>>
>>109897568
poke her with your dick
>>
>>109897522
> just because you can't run it doesn't mean it's not local
no, it does
>>
>>109897639
Go back, poorfag.
>>
>>109897650
I'm in lmg already
>>
Llama 1 65B came out in 2023, over 3 years ago.
Compared to modern models, where do you say it sits capability wise, in general terms?
Ignoring the small as fuck context window, of course, although that by itself means that it's incapable of performing certain tasks, but alas.
I miss kaiokendev.
>>
>>109895777
>2 QFN MTP PRs
>both withering on the vine
>>
>>109897664
>kaiokendev
Nearly certain he's still here and posts without revealing himself.
>>
>>109897561
>>109897542
I did both, it doesn't work. Interestingly enough, some characters work immediately and others refuse forever, so I figure it's something in their descriptions that makes the reply EOS, maybe the EOS token is contained in the description itself, idfk.
Raising temperature helps but makes some text come out as gibberish (unsurprisingly).
Thanks for your attention to this matter.
>>
>>109897672
Just download the Unsloth fork and get some model to fix whatever problems you have. For me everything was too ridiculously slow before fixes so I used mimo via api to fix the performance and then used qwen flash to fix everything else.
At least the Unsloth fork actually works and has all the features for the most part, even though performance sucks (not worse than mainline, but it sucks)
>>
Local LLM challenges;
https://en.wikipedia.org/wiki/Arimaa
https://en.wikipedia.org/wiki/The_Campaign_for_North_Africa
>>
>>109895565
Where'd you get those dates from if you don't mind me asking?
>>
>>109897673
Nah. He was murdered by ninjas.
>>
>>109897738
That's dariobot our local anthropic intern
>>
>>109897664
Llama 1 is the only model that was trained without any synthetic data (it doesn't even have the Elara shit or As an AI language model). It should be a good base to finetune further using modern techniques to improve on it.
>>
>>109897374
Lmao
>>
File: nope.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109897588
>Sure you got $50 to spare
>>
>>109897416
Yeah, if this goes viral the next chink wave will benchmaxx it
Here's Qwen-3.8-27b q6_k with that same ah ah mistress prompt
https://files.catbox.moe/brz1rv.mp4
One of those flash moes that made claude self-portrait pages can probably make it actually funny
>>
>>109897775
>Llama 1 is the only model that was trained without any synthetic data
Falcon-180b too I think
and there was some moe trained more recently with a 2022 cut-off
>>
File: 1775095405007256.jpg (171 KB, 1280x720)
171 KB JPG
>>109897863
Yeah totally the same...
>>
>>109897775
GPT-4 base was done by August 2022. It's probably the most powerful AI model in existence untainted with AI slop, and it's locked up by OpenAI for no reason. Imagine if Sister Fister they released the weights....
>>
>>109897863
I mean that's pretty good for a model that fits into a poorfag GPU and the guy himself said that he iterated a few times.
>>
>>109897680
You're just using regular Gemma 4 31B? What quant?
>>
>>109897891
Cool it with the anti-semitism.
>>
>>109897558
>can't block a simple popup in 2027-1
anon..
>>
>>109895715
I'm still using your previous branch (thanks a million for that!).
Any chance for a new stable checkpoint before your public release?
>>
>>109897462
I can get normal gemma to do a rape card with a 1-2 sentence prompt (no prefills).
>>
>>109898002
not her, but i'm curious how you do so? note that i'm not asking to be spoonfed here because the policy override just werks and i literally don't need anything more than that. i am legitimately just curious
>>
>>109893726
learning microsoldering, and buying broken hardware to resell for a profit. When in a gold rush sell buckets/shovels kind of thing I guess.
>>
>>109897917
gemma-4-31B-it-qat-q4_0-uncensored-heretic-GGUF
Yes, it's pretty lengthy, but it works okay sometimes and writes pretty well.
>>
>>109893726
vibe porting bloated apps to c++/qt to save ram
>>
File: absolute_cinema_man.jpg (8 KB, 225x225)
8 KB JPG
>>109897664
digitally molesting lolis in llama 1 they just cry and say it hurts. Way more realistic than what we have now
>>
>>109897915
> 24gb vram
> a poorfag GPU
>>
>>109898075
It's the absolute minimum you need to play here so yes.
>>
>>109898107
>>109898107
>>109898107
>>
>>109897145
>I know this is reddit but
stopped reading, I'm not falling for it a second time
>>
File: 1782125875361468.jpg (38 KB, 716x780)
38 KB JPG
>>109893224
Can someone explain in retard friendly terms what "Jev" is supposed to do and accomplish?
>>
>>109898026
You're getting refusals on an abliterated model? What
>>
>>109896276
Trannyware.
>>109896614
>The things your rabbi isn't telling you
>>
>>109898075
Yes?
>t. 24GB VRAMlet
>>
>>109897565
Thanks.
>>
>>109897863
I like it.
>>
>>109897664
Can I download it nowdays, and run it with current llama.cpp?
>>
>>109898481
As far as I know, yes. You might get some warnings about legacy yadda yadda, so creating a new GGUF is probably the way to go.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.