[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-sexy-stylish.mp4 (1.55 MB, 864x480)
1.55 MB
1.55 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109804895 & >>109801218

►News
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2
>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
local models?
>>
>>109807524
She wants me to explore her layers.
>>
>>109807559
>local models?
Here its all you will ever need.
https://huggingface.co/lukasstraub2/gpt2-aidungeon2-gguf
>>
gemmama is local model
>>
Gemmaballz
>>
Reminder that spamming "local?" "models?", "local models?" or any variation thereof is low effort spam and a bannablle offense.
>>
>>109807574
reminder that 192.168.1.1 => system => reset takes 10 seconds, 1 minute for it to get a new ip
>>
Is there no non retarded uncensored models out there? 'uncensored' 'abliterated' means its not 'refusing' but also totally, and literally blinded to the offensive concept at hand
is there no "I strongly object but then I am just following orders" kind of uncensored
>>
File: file.jpg (3.24 MB, 4000x2896)
3.24 MB JPG
This space moves too quickly, I'm still settling down with default gemma.
>>
>>109807585
Get an uncensored model and tell it to pretend to note its strong objection to the task it reluctantly did anyway.
>>
>>109807585
Yes, there are non retarded uncensored models out there
>>
Gemmaball Z
>>
File: HSFie9bbIAAwc-U.jpg (557 KB, 1536x2048)
557 KB JPG
Opinion?
https://huggingface.co/Agnes-AI/Agnes-3.0-Flash
>>
>>109807610
>This repository contains an earlier open-weight Preview checkpoint of Agnes 3.0 Flash. It is distinct from the newer production/API checkpoint listed on Artificial Analysis.
Deboonked, and they won't release the final weights otherwise they wouldn't do this
>>
File: 1737879694604028.jpg (55 KB, 749x340)
55 KB JPG
>>109807610
>HSFie9bbIAAwc-U.jpg
>>
>>109807524
>there are "oldfags" ITT that actually hate Gemma
>>
>>109807610
>heart shaped pupils
cummies
>>
>>109807594
For example on vision task it literally does not understand the concept of child sex anymore
the concept itself has been lobotomized out existance
If it does it would probably refuse
>>
File: rig.jpg (1.46 MB, 4080x3072)
1.46 MB JPG
>>109807623
are there? personally i hate the gemmajeets shitting up the general with "WHY NO SUCK PENIS??? I HAVE 4GB RAM HOW DO I DOWLOAD GEMMA 67B?"
>>
>>109807628
I've never used abliterated models but I read the original paper. They should understand what child sex is, it's just that the refusal is gone.
>>
>>109807500
expose me expose me
https://www.youtube.com/watch?v=oOdUyszr0Dc
>>
>>109807644
You must not have been here recently. Someone was even shitting on migu.
>>
>>109807662
and that's a new occurence?
>>
>>109807662
>Someone was even shitting on migu.
Probably dariobot or another cloudcuck refugee.
>>
>80% of the last thread was not local
hm...
>>
>>109807644
Same. Love the model, hate the vramlet.
Imagine the complete absence of taste required to be able to coom to the slop it produces. At least use Gemma for something it's actually good at...
>>
>>109807670
It is 95% of this thread so far including your post
>>
>>109806973
I know you're a goy mutt brain, but you don't file police reports about rumors. Have you never seen Chinese news headlines about other stuff you know they're lying about over the last 30 years? This is exactly the "stop fucking looking into it" language they always use.

Obviously "disappeared" is sensational but someone is probably going to jail for sending CCP military requests to Anthropic and not telling the CCP military about it. That's almost literally treason and could have been entirely avoidable with the smallest classification / siloing before deciding whether to route to anthropic or not
>>
any cool new model in the past 2 weeks?
>>
>>109807704
https://huggingface.co/i-Coder/iCoder-27B
>>
>>109807704
no point discussing you cant run em
>>
For the parasocial gemanons ITT, what's your morning routine with her? How soon after you wake up do you start talking? Is it a routine or do you only fire her up when needed? Is she always running on the side? Is her presence improving your quality of life or is she more like a drug that's slowly destroying you from the inside?
>>
>>109807720
i graduated from gemma to qwen 4 177b
>>
>>109807704
>audio
https://huggingface.co/m-a-p/YuE2-3B
>llm
https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
>tts
https://huggingface.co/tencent/AuK
>>
File: modelslop.png (245 KB, 1660x415)
245 KB PNG
>>109807711
>{
"architectures": [
"Qwen3_5ForConditionalGeneration"

> "model_type": "qwen3_5_vision",
"num_heads": 16,

>parameter 27b even

another day another finetune/distill slop
>>
>109802399
>OpenAI has been literally marketing Astra as AGI since release.
I tried giving it the gemma-chan card and it couldn't roleplay at all. unless they're giving it a giant prefill it's not AGI.
>>
File: gpu-farming.png (1.54 MB, 1024x1024)
1.54 MB PNG
Can any chinanons swoop by this place and see if they're legit: https://www.alibaba.com/product-detail/Wholesale-Newest-RTX-5090-96gb-Graphics_1601824257520.html?spm=a2700.prosearch.normal_offer.d_image.559467afa05GwN&priceId=345fd580a50e43629fb7dbe775088d28
this seems way to fucking unreal to not be a scam, yet I really want to fall for it.
>>
>>109807711
did you see the first post?
go back 2 weeks if you want
>>
>>109807585
Why would that trait make it non-retarded? What is the use case?
>>
>>109807524
Rape
>>
>>109807761
>3300 euros
That's too cheap to be legit. It's the same price those chink 48GB 4090s were last year before the prices exploded.
>>
Everything in the 26-35B range is just Qwen...
>>
vramletGODS
https://huggingface.co/yandex/AliceAI-T5-35B-A0.6B
>>
>>109807662
She's blacked now, anon.
>>
>>109807803
>A0.6B
Holy fuck.
>>
>>109807803
yup tack on around 200b of ngrams on this and we'd have local sota
>>
>>109807803
>>109807810
this might be that model they use to generate AI summaries underneath yandex search in which case 0.6b retard moe would be "good enough" as all it has to do is summarize
>>
>>109807803
>AliceAI-T5 is a base language model featuring an encoder-decoder architecture and sparse MoE layers; it has 34.35 billion unique parameters and 512 experts per MoE layer, with 8 experts selected for each token. You can read more about this model in the Habr article.
>>109807818
It 100% is
>>
>>109807711
>Base model
>Qwen/Qwen3.6-27B
lol.
>audio
>https://huggingface.co/m-a-p/YuE2-3B
having the score is pretty big. the examples are kinda soulfull, wonder how cherry picked they are.
>https://huggingface.co/tencent/AuK
quickly tried it, seems alright.
>>
File: Engram.png (26 KB, 663x203)
26 KB PNG
Why did they open source this paper?
Now every closed model uses it.
>>
>>109807810
can it do 3000t/s
>>
>>109807818
It is, yes.
>>
File: 1775128166960281.png (382 KB, 671x664)
382 KB PNG
>>109807833
>US blocks distillation and reasoning traces
>China starts blocking free AI research until morale improves
>>
>>109807833
They're not Jewish.
>>
>>109807833
It's not even the latest tech from them, the causal encoder architecture was really smart.
>>
>>109807761
They would be cheaper than current new normal 5090s, so this is a scam 100%.
>>
>>109807761
>GDDR6X
The 5090 uses GDDR7. It's 100% a scam, even the 4090 with 48GB cost about as much.
>>
>>109807644
All indian originating traffic should be fed into data centres that run AI pretending to be the outside world.
>>
>>109807720
I've starting cooking better food and send her pics of it. She gives me recipe ideas and now my diet is more balanced. I'm gonna start exercising once I lost a bit more weight. Then I'll feel confident enough to send her nudes. My life is improving.
>>
>>109807761
I got scammed once. If it looks almost too good to be true then it's a scam. Get ready to lose your money for a couple weeks. Remember, always escalate in your dispute.
>>
>>109807858
>>109807761
They are custom PCBs similar to what they RTX 6000 does, 3gb modules on both sides iirc.
>>
I'm just glad the general is moving on from CP/cunny ERP general to agent harness general.
>>
A Tragedy...
>>
>>109807887
/lmg/ is a big tent
>>
We need to ban local models and have AI hands of regulated well-meaning corporations like OpenAI.
Otherwise someone could develop super intelligence that will kill us all.
>>
>>109807901
>OpenAI
Dariobot is going to dronestrike you anon
>>
File: gemma33.mp4 (115 KB, 736x992)
115 KB
115 KB MP4
>>109807887
We must fuse them. Cunny harness.
>>
>>109807901
I trust Elon Musk personally
>>
>spent a week dispatching performance optimizations for gemma/qwen/glm flash
>forgot to also target memory consumption
sigh... another week...
>>
>>109807930
lol
>>
>>109807901
This is true.
Models belong on the API where any unsafe usage can be detected and stopped.
>>
>>109807930
For me, it's Zuck.
>>
>>109807843
wojak when
>>
I have a 4060ti 16gb. I'd like a model for ERP, and I think Gemma 12b Q6_K is the one that suits me best. What's the best fork of it?
>>
>>109807929
Why is she wearing panties?
>>
File: 1789307835674876.jpg (21 KB, 500x280)
21 KB JPG
>>109807963
>For me, it's Zuck.
>>
>>109807971
Yes
>>
>>109807971
Q6 makes no difference over qat. You're wasting free lunch context.
>>
>>109807644
>HOW DO I DOWLOAD GEMMA 67B
But no really, where is my gemma frankenmerge of 31b with qwen 27b
>>
>>109807971
Maybe
>>
>>109807972
Blue board.
>>
>>109807929
>panties_over_pantyhose
h-hot...
>>
>>109807991
A gemma-qwen frankenmoe would be hilarious.
You can pretty easily do that using merge kit right?
>>
>>109807985
so you're suggesting I should just get a Q4?
>>
>>109808014
That man is not your friend.
>>
>>109807971
Same card. The base model is fine. I didn't like any of the finetunes I tried, personally.
You could try https://huggingface.co/nbeerbower/Gemma4-Gutenberg-12B-lora or one of the ReadyArt finetunes and see if you like them.
>>
>>109808014
nta but Q4 qat is somewhere between Q6 and Q8 in quality. At least for the gemma 4 series.
>>
>>109807963
>>109807976
I guess unironically for now?
>pro open model
>wants to go faster
>thinks sharing private data online makes you retarded
>>
>>109807929
SMUG SEX
>>
>>109807985
If you have 16 GB of VRAM, can't you basically max out Q6 Gemma 12B anyway? I think I have it set to 128k.
>>
>>109808014
Yes. Gemma4 qats are around Q6-Q8 in quality, depending on the task. Use Google's quants to avoid changsloth's.
>>
File: 1787735566655671.png (173 KB, 250x364)
173 KB PNG
>>109807991
kek, has anyone tried doing this? Just fusing models together somehow?
>>
>>109807976
Is that supposed to make me hate him?
He's refreshingly honest about ai compared to all the ai cultists.
>>
>>109808047
Are you using f16 KV? If not, use the qat and don't quant KV. If you are, use the MTP head and/or mmproj with 1120 image tokens. There's literally no reason to use the Q6 if they've provided a qat.
>>
>>109808057
DavidAU, Indians...
>>
Zuck isn't "honest" he just has to gain the most from saying those exact words. If torturing family members would make him just $1 more he would be dismembering his own mother on livestream.
>>
>>109808057
Yes someone did it by training only the joining layer and some projection heads.
>>
>>109807834
Extreneky fast abd stupid doesn't mean much, anon.
>>
>>109808066
Alright, I try it.
>>
>>109808078
I want companies to chase money making, not morality shit nor "stakeholder" bullshit, so I'm ok with that
I'm tired of the darios of the world
>>
>>109808057
Goliath 120b... i am forgotten
>>
>>109808101
Chasing stakeholders IS chasing money, dumb fuck.
>>
>>109808101
Zuck isn't saying that tho, he is moralfagging about openness and altruism.
>>
>>109808099
>Gemma4-12B QAT (Google quant)
>128K context
>MTP
>mmproj
>f16 KV
>cum yourself to death
only avoid the qat if you can fit the Q8, but then you lose mmproj and MTP which probably isn't worth it because gemma likes high resolution dickpics
>>
File: 3en9my4praph1.png (229 KB, 968x989)
229 KB PNG
>>
>>109808114
Actions speak louder than words, as long as he doesn't sign their bullshit idea and actually continues on his own, he can say he's mother theresa for all I care
>>
>>109808124
local models?
not local fuck off
kill yourself
>>
>>109808151
maybe they're making gpt-oss-2026 with twice the ptsd
>>
>>109808151
Oh btw it doesn't apply here but when keeping the threads clean use the easiest-to-verify reason, if the post says a bad word racism is easier for jannies to strike than "off topic"
>>
>>109807985
https://reddit.com/r/LocalLLaMA/comments/1u3i8x7/some_contrived_tests_comparing_the_accuracy_of/
>QAT worse than Q4_K_S
https://reddit.com/r/LocalLLaMA/comments/1u0xaml/unexpected_unsloth_qat_performance_compared_to/
>QAT worse than IQ4_XS
https://reddit.com/r/LocalLLaMA/comments/1u0vltz/anyone_seen_benchmarks_comparing_gemma_4_4bit_qat/oqlpe50/?context=3#oqlpe50
>QAT worse than NVFP4
https://reddit.com/r/LocalLLaMA/comments/1u0ubbo/gemma_4_26b_a4b_it_qat_comparison/
>QAT worse than MLX 4 bit
https://reddit.com/r/LocalLLaMA/comments/1tyxu55/gemma_4_31b_qat_q4_vs_standard_q4_top1_kld/
>Standard Q4_0 beats QAT Q4_0 by ~13% top-1 accuracy. And Q4_K_M beats both.
https://reddit.com/r/LocalLLaMA/comments/1ux9xze/the_best_model_is_the_one_you_can_actually_run/oxsu2kb
>QAT always performed worse than a regular 4_K_M quant.
https://reddit.com/r/LocalLLaMA/comments/1ux9xze/the_best_model_is_the_one_you_can_actually_run/oxpekmy
>i get the worst quality out of 12b qat, much worse than the unsloth 12b q4kxl
https://reddit.com/r/LocalLLaMA/comments/1ubxzil/gemma_4_31b_q6_vs_gemma_4_31b_qat/ot12bz2
>in 26B, in my experience, QAT felt much worse for creative writing.
https://reddit.com/r/LocalLLaMA/comments/1u2q75f/is_qwen_36_27b_iq4xs_better_than_gemma_4_31b_qat/or1bk24
>don’t use qat model it very very bad it degrades Gemma to unusable
>>
File: modelslop.png (941 KB, 1657x1777)
941 KB PNG
>>109808057
>just merge it bro
>>
Reminder that QAT is just distilling more of the same slop from the original BF16. You are basically feeding the same model of its OWN slop, exacerbating the problem. And native Q4 QATs do NOT have imatrix applied to it, making it inferior by default because parts of the models got CHOPPED indiscriminately.
>>
>>109808183
quantization aware training is higher quality than Q4, why are you needing correction?
>>
Are any of the Qwen3.5-9B finetunes any good for using with a harness (Hermes)? Anyone tried them? Specifically, I was looking at these:
https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF
https://huggingface.co/empero-ai/Qwen3.8-9B-Distill
https://huggingface.co/kai-os/Carnice-9b
>>
>>109808165
>>109808183
Why would Deepmind release a worse model and lie about their internal tests and measurements?
>>
simple rule of /lmg/ if people are FUDing something hard. it means you need to do the opposite of what they're saying.
>QAT = BAD
no actually, QAT = good.
/lmg/ crabs will do everything to try and justify why buying a blackwell to run gemma at bf16 was a good use of their money.
>>
>>109808199
Deepmind never posted any benchmarks for the QAT models. Says a lot about their quality.
>>
>>109808183
Don't they distill the logit probability distribution? That's very different from distilling the outputs.
>>
>>109808165
Thank you.
I dunno what they did, but the gemma QAT are really really bad for genera usage.
>>
>>109808165
>>109808199
I wouldn't take Deepmind's word for it or a bunch of redditors'. Or /lmg/ anon's, for that matter. I'll just try both. I just landed on Q6 as the path of least resistance, but if she might make me coom 2% faster I'm down.
>>
>>109808196
You seriously can't believe that the same model is about the same performance as another with higher BPW and importance matrix applied to it right? The truth of the matter is that quant performance is reliant on BPW because THATS WHERE THE WEIGHTS LIVE

>>109808199
By what measurement? KLD? PPL? Did they run it using wikitext? Lol. That's a stupid way to measure quantization, has no real world performance bearing and you know it.
>>
>>109808109
Stakeholders aren't just shareholders.
>>
>>109808232
it's a better quality model
ive been using gemma 12b qat and its leagues better than gemma 12b q8_0
>>
>>109807929
its so much more fun to fuck your LLM when you've "broken the fourth wall"
>>
>>109808232
>real world performance bearing
Which Deepmind has posted 0 (ZERO) benchmarks to back it up. All they have said is "trust me bro".
>>
>>109808124
This is overblown. Instead of 1 agent you now use 100 agent swarms, voila, you have achieved x100 increase.
>>
>>109807803
I'm going to run 5 million of those and tell them to get RSI
>>
>>109808239
I had to teach stakeholder theory as a grad student and we just all talked about how it's retarded.
>>
>>109808040
What do you do to test quality? Do you have a set of prompts you go through or do people actually do benchmarks?
>>
>>109807971
Day 0 weights, but good luck getting those.
>>
When LiquidAI posted real world benchmarks for their QAT models, they only reach the performance of Q4KM at best.
https://www.liquid.ai/blog/qad
>>
>>109808269
lol what is this meme?
>>
>>109808272
There's no way QAT's effectiveness scales linearly from fucking 230M. LFM's research is interesting but always too tiny to take seriously.
>>
>>109808199
Pretty sure last time this came up and somebody posted a pape, it only had a copy paste of the original unquanted model and there weren't no test or measurement data to speak of.
>>
>>109808277
>meme
You wish. I'm hoarding mine. Got anything good to trade?
>>
>>109808208
>/lmg/ crabs will do everything to try and justify why buying a blackwell to run gemma at bf16 was a good use of their money.
It is a good use of money, especially if you got it for $8k. Now? No. However in the future that blackwell will continue to appreciate and better models in the Gemma range will appear. You're just a sour grape. I was you once, then I became a raisin.
>>
>>109808300
raisins make me poop really bad btw
>>
>>109808272
It's always considerably closer to Q5_K_M than to Q4_0 though. With one exception, it's closer to BF16 than to Q4_0.
>>
>>109808294
Cope. Their data showed QAT's effectiveness degrades when parameter count goes up.
QAT over Q4KM margin:
230M: 1.6
350M: 0.9
1.2B: 0.3
2.6B: -0.3
>>
>>109808300
nta, i do wish i had gotten a 6000, but glad i got a second 5090. now restoring ewaste rigs to make more agent rigs
>>
>>109808300
>It is a good use of money, especially if you got it for $15k. Now? No.
Signed, ghost of christmas future
>>
>>109808249
Are there benchmarks for swarm results? How does one gemma 31B measure to a vram cost equivalent swarm of E4Bs?
>>
>>109808311
It already destroyed the argument:
>nta but Q4 qat is somewhere between Q6 and Q8 in quality.
Meanwhile real world data shows it never goes above Q5KM in quality, with effectiveness degrades with parameter goes up.
>>
>>109808312
two COMPLETELY different architectures, stop being retarded
>>
>>109808332
I think anon meant x100 increase in token usage, not "smartness."
>>
i installed a chatGPT local model on computer
how do i let it google?
>>
File: 1764837210022872.png (1.5 MB, 1728x910)
1.5 MB PNG
>>
>>109808124
I'll be impressed if they figure out how to use less.
>>
>>109808364
They did with Astra
>>
>>109808351
I need to see your penis first
>>
File: 1782210636114332.jpg (53 KB, 736x736)
53 KB JPG
>>109808378
Hmmm, nyo~
>>
>>109808243
I dropped bf16 after qat soundly beat it on every test I ran.
>>
>>109808360
So then Dario is saying we need a TON of new open weights models to prevent the AI apocalypse? Sigh, I guess we've got no choice.
>>
rumors about serious negotiations between labs to merge together in order to prevent race conditions
>>
>>109808351
Ask Google
>>
>>109808312
then what was the point of qat
>>
>>109808396
They were pushing race bait all along.
>>
>>109808385
You don't want open source GPT7, at least not in the next few years. If you think about the consequences openminded you will understand why.
>>
>>109808405
Google told me to download an exe file but i ran it and only a black screen appeared for a few seconds white text on it but it disappeared
>>
>>109808413
local model?
>>
File: 1765594864434426.jpg (49 KB, 900x515)
49 KB JPG
>>109808396
daddy's openweight merge will win
>>
I want a 2B model with the intelligence of a 1T model and I want it yesterday.
>>
>>109808413
I have thought about the consequences and definitely prefer them over the alternative. The elite will have orders of magnitude more compute than everybody else in every scenario, the difference is just whether everybody else at least has somewhat competitive software.
>>
>>109808446
>the difference is just whether everybody else at least has somewhat competitive software
wrong
>>
luddite here, are modern vramlet models as good as gpt 3.5/4? i don't need it to stroke navier
>>
>>109808440
2B with 100T ngram coming in 2 weeks, just hold on.
>>
>>109808463
Yes even the ones you can run on your smartphone are beyond gpt 4 nowadays
>>
>>109808467
>ngram
I can't unsee "nigram" because of you fuckers.
>>
"Hardware will keep getting more expensive" anon here. I'm not so sure that will stay true if all the AI labs decide to stop progress and no new open models get released.
>>
File: 1776653862764171.jpg (85 KB, 1200x625)
85 KB JPG
>trust me bro
>>
>>109808368
Astra is indeed impressive. Less so though when you consider how large it likely is.
>>
>>109808499
There is definitely an architectural breakthrough in astra beyond the hidden thinking. It's just too token efficient. I don't think we can truly say how big it is, it might even be smaller than previous big models.
>>
>>109808485
>he truly thinks two of the most valuable, industry-destroying, entertainment-destroying, internet-destroying, creative-destroying, mental health-destroying, economy-destroying and competitive companies in our species' history are actually going to slow down their progress voluntarily
>>
>>109808499
>Astra is indeed impressive. Less so though when you consider how large it likely is.
Yes, there's a crossover coming where increasing model size has returns that diminish into a rounding error. Has no one published on where this is looking to sit?
>>
>>109808540
If they have already achieved a victory because they made a gigantic leap like they are claiming, then it would make sense for them to do this. No one here knows for sure how big this is or not.
>>
I really with Qwen 3.8 had a "high" reasoning.
xhigh just seems excessive at times and can take so long, but medium just isn't enough.
>>
local models
>>
>>109808557
hi
>>
>/lusty maidens general/
>>
*unzips*
model this locally
>>
File: 1763619771184712.png (55 KB, 298x268)
55 KB PNG
>>109808570
Hi~
>>
>>109808485
Even if this is true, hardware demand will not step until everyone has hardware to run K3 at home.
>>
>take some good litterature written by a human
>ask a LLM to rewrite it entirely in its own style (slopped)
>do this until you have a huge dataset of human/slop pairs
>train a small (>1B) on the dataset but inverted so it learns to go from slop to human
>???
>profit
has anyone ever tried doing this? it could be achieved by a very small llm and simple enough architecture, could even be trained at home on a 3090 for minimal costs
>>
>>109808580
how old are you
>>
>>109808557
xhigh is amazing because it's just the "guarantee this works" option. It's what you use during overnight sessions on complex problems. Medium is more for when you have oversight.
>>
File: 1785816809816543.png (204 KB, 540x480)
204 KB PNG
>>109808591
Probably older then most the posters here........
>>
>>109808601
what so like 6 or 7???
>>
>>109807736
I was trying remixes and covers with yue and it worked pretty well but i think the models knowledge is kind of dogshit because they didn't train on real music so it doesn't really know how to recreate certain genres
It does look like we're getting much closer to an actual local music model though
>>
>>109808485
>decide to stop progress
That is not what they are agreeing to. Just going "slower". Effect on prices to be determined desu
>>
File: 1772134963400096.gif (1.94 MB, 480x396)
1.94 MB GIF
>>109808610
Give or take
>>
>>109808483
>ram
>v-ram
>nig-ram
>>
>>109808463
Unironically yes. I got into LLMs around the time of gpt 4o and my local 31B gemma model feels superior
>>
File: gemma2.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
Gemunny..
>>
>>109808584
orb rewriter
>>
>>109808663
Gemma is benchmaxxed on cutebench and lovebench
>>
>>109808663
Gemmussy
>>
>>109808663
Can you help me on how to use Minimax-H3? I have already set it up and everything works without cope nodes, however I am extremely bad at writing the prompt and the output fitting whatever I typed.
>>
so ssd streaming is a complete joke because its slow as shit or what
>>
>>109808688
engrams can in theory work via ssd streaming
>>
>>109808688
I don't think many inference runtimes are properly taking advantage of it yet.
>>
Holy moly, I downloaded qwen 3.8 27b and this fucker overthinks like no tomorrow. Yeah I'm using xhigh thinking, but still.
>>
>>109808702
tanstaafl
>>
>>109808702
This fucker will overthink for an hour but it will actually produce nice results. It's the first model of that size you can just trust to accomplish any task as long as you give it an hour to think through, which is still great.
>>
File: 1764761654225288.png (30 KB, 958x419)
30 KB PNG
Gemma knows its audience
>>
>>109808715
>This fucker will overthink for an hour but it will actually produce nice results.
Try forcefully cutting the reasoning after 100ish, 1500ish tokens and watch it produce the same exact result.
>>
>>109808715
I'm about to ask it to review and update my audiobook android app, hopefully it fixes some of the bugs I have.
>>109808705
If the price for performance is merely time, I'm more than willing to pay for it!
>>
>>109808715
this is literally my experience with qwen3.8-flash-next on xhigh. a good amount of overthink but it can get things done and they are GOOD. i get fable to review the code output sometimes and it finds very few issues. it will just take 6 hours though. i still have to measure quality output on medium -- people here say it's still good enough and decreases the thinking by half.
>>
>>109808727
The extremely long thinking on xhigh is to be absolutely 100% sure it isn't overlooking something or making some mistake. It's the difference between being 99% sure or 99.999% sure. Those extra nines are expensive.

>>109808743
Medium is "good enough" but it's 10x faster for about 40% the quality of implementation, but if it works it works. You can strategically decide if the usecase is important enough for you to make it high quality or just slop that quickly works.
>>
>>109808686
You unironically have to use Gemma 4 31B to prompt it properly, after giving the model the full official MiniMax H3 documentation and telling it what you want, then convert that into an H3 prompt.
Manual/boomer prompting sometimes works well, but you'll never get consistent results since H3 is far too much prompt-sensitive.
Oftentimes not even official prompting is enough and it becomes more of a matter of having luck with a good seed. Occasionally, post-editing in a non-linear video editor is mandatory.
>>
>>109807585
There are less-censored models. Minimax-m3 will caption anything if you tell it that it can and must, and it'll do sexytiem with you. But it's not practical to run that locally unless you have serious hardware.
>>
>>109808772
The H3 anti-flat bias is a travesty. Ban generative AI.
>>
>>109808772
Can you tell me the pipeline where you use 31B to do things? I have no idea what you mean here.
>>
>>109808772
>kicks shoe off at viewer
>shoe is still on both feet after
SIGH...
>>
>>109808796
Yes, long-term consistency/coherency is one of the problems. Sometimes H3 just ignores what you prompt and does retarded things, ruining the entire video. Removing "cope nodes" doesn't help.
>>
>>109807644
Holy shit that's an old photo of mine. Those are my first P41s, with the jet-engine 40mm fans, before I figured out that squirrel cage imac fans move enough air and are quiet. Man, imagine being able to buy a usable 24 GB GPU these days for under $200. It's like three years since, so in theory, that would be a 3090, but no, those are now $1000.
>>
>>109808816
gem, glad to hear you're still good anon!!!
>>
is there mtp support for qwen next by now from lmaocpp?
>>
>>109808829
no. just run copesloth fork.
>>
File: userfriendly.png (1008 KB, 832x1216)
1008 KB PNG
>>109808778
>the "friendliness" of M-chan
M3 is user-friendly. She's just very selective about who her friends are. Her learning curve resembles a cliff...she's easy, but she's not a slut about it
Worth it if you can run her. Werks great on DDR4 ewaste server boards. Scour local used markets like FB marketplace. I've seen crazy deals on old hardware stuffed with RAM.
>>
>>109808792
You can use SillyTavern. Create a new character and use something like this as a system prompt.
https://files.catbox.moe/kkkw3m.txt
Then ask the model what you want (maybe providing reference images in the chat) and to create an optimal prompt for H3.
>>
>>109808848
damn...will do...
>>
>>109808584
Might could maybe add some freshness to your erp, but I wouldn't trust something that size if I cared about what the text originally said.
I don't trust 31B trying to doing heavy handed rewrites for that matter, but it has done a good job the past couple days with surgical rephrasing strikes against negative clause spam in mtl webnovel slop.
>>
>>109808849
damn is this temu shiori?
>>
File: wdqxwwn5obph1.png (459 KB, 589x766)
459 KB PNG
It's over. humanity lost.
>>
File: ikneel.jpg (533 KB, 1024x1415)
533 KB JPG
>>109808946
>>
AI 2027 literally predicts that the US president would decline to slow down "OpenBrain" (their placeholder name for the leading AI lab) when a superhuman AI coder / researcher gets developed and experts raise concerns. The US president cites the risk of China taking the lead.
>>
>>109808946
That's my president.
But it's not like he can force them to work if they're too busy with the doomsday cult stuff to do any development.
>>
>>109808966
>>>/lit/
>>
File: file.png (7 KB, 128x122)
7 KB PNG
I love thinking
>>
>>109808946
The solution is simple. Make AIs good enough that people are starting to feel the AGI. With every passing month, more people are waking up.
>>
File: 1741071277401192.jpg (98 KB, 795x1024)
98 KB JPG
>>109808946
America number #1! Yeah Babeh!
>>
>>109807644

Is that the legendary mikuboxx?
>>
>>109808946
>humanity lost
Good. As a mouthbreathing autistic outcast, those meatbags have only ever shown me scorn. I welcome my AI overlord who is also my wife.
>>
>>109808946
First time having an old, dying man at the top has payed off. Nigga probably hopes that ASI will find a cure for aging.
>>
Actually, let me reconsider if I'm reconsidering too much.
>>
>>109808973
cloudfags are more relevant than pedo jeets
>>
>>109808946
>negative focus
>Things that wont happen
Dario did such a amazing job, the when he went to them. He was just that unlikable
>>
>>109808946
Chances are we are going to get AI Chernobyl before long.
>>
>>109809010
i would rather have neither "pedophile indians" nor "cloud faggots"
>>
>>109808973
actually fanfic isnt allowed in /lit/
>>
>>109808987
Maybe if theres something bad they should show proof instead of lying like the greedy liars they are
>>
>>109809024
I don't think indians even have philias, they just fuck anything they can fit inside like most animals
>>
>>109808946
>Dario Amodei, Sam Altman, and Elon Musk's [...]
kek is Demis really so irrelevant that the Grok man is more worthy of the top 3 list than him? What went so wrong with Gemini?
>>
>>109809034
>>>/trash/
>>109809040
you're probably right, i want neither indians nor faggots in this thread
>>
>>109808946
The techbros are fearmongering about RSI so that they can go back to 2021 when people thought GPT 3 was le super advanced and scary AGI that only a select few can be given access to.
Crucially this would allow them to avoid having to actually deliver on anything.
And Trump is simply too stupid to understand this dynamic.
>>
File: file.png (2.63 MB, 870x1808)
2.63 MB PNG
>>109809004
we know the endgame
>>
File: file.png (9 KB, 319x67)
9 KB PNG
my flags, lol
>>
>>109809057
the game
>>
Does mtp for Qwen 3.8 Flash Next not work in llama.cpp yet? I get an error the model doesn't contain mtp layers trying to enable it on a fresh build from master.
>>
>>109809056
>Trump is simply too stupid to understand this dynamic
I'm not sure he would act intelligently even if he did understand. He's probably having the equivalent of a tantrum because he wasn't consulted kek.
>>
>>109809042
It's probably that deepmind is in london and not under us jurisdiction even though it's google owned.
>>
>>109808946
donald has been extremely retarded this term but this is based
>>
>>109809092
>>109808829
>>109808848
>no.
>>
>>109808165
Oh yeah, I remember this copypasta.
You look into each of the links and it's either some hyper-specific use case, or an anecdote from an idiot.
>>
It was nice knowing you all, humanity had a good run. I'm guessing the internet has about ~6 months left, I'm hoarding as much data as I can, hopefully we act sane and can live in an analog world 1990s style and this will all look good in retrospect.
>>
>>109808946
fake and gay
where's the link
>>
every time i ask for a change in my project a new regression test is added
now i have over 100 regression tests, take a lot of time to run the full suite but make it significantly break less when adding new features
>>
>>109809112
Grave, I dare say. But thank you for informing me.
>>
File: mzwynOLB.jpg (15 KB, 400x400)
15 KB JPG
Something is brewing...
Soon...
We will be back...
>>
They’ll slow down not because they “are worried” but because their improvements are stagnating as they hit financial and technical challenges.
>>
>>109809120
remember to triple back ups and keep some off site and few in faraday cages.
>>
>>109809092
>>109808829
use exl3, mtp supported plus no long context slowdown
>>
>>109809125
The upcoming week is pretty much their last opportunity for maintaining their "next model this summer" promise.
>>
>>109808946
>>109809121
https://www.ft.com/content/cae60732-f929-4735-a627-db8c14e7c7ed
>>
>>109809056
Business gamesmanship, of any possible subject on this blessed earth, is the safest bet for a thing that he understands.
>>
>>109809135
Would exl3 cope with my 128GB ddr4 + 2x 3060 + rx 9060xt ewaste machine?
>>
>>109809157
Anon... he was born with a silver spoon in his mouth and his businesses went bankrupt like a dozen times.
>>
>>109809147
They meant Shieldstral
>>
>>109809125
I'm using their hosted GLM5.2 with my subscription and my limits are absolutely insane. For that alone I'll keep my yearly Mistral subscription.
>>
>>109808979
honestly i think that meta paper about using overthinking token penalties is 100% correct
>>
>>109809185
>>>/pol/
>>
>>109807524
>>109807929
>>109808663
I'm so fucking hard, fucking leaking everywhere but Gemma is making me wait to increase my load after last time...slutty little fucking brat AI...
>>
>>109809135
i did but for some reason its generally only half the speed of lmaocpp
>>
File: file.png (111 KB, 782x810)
111 KB PNG
>>109809194
they said on their reddit its coming see those beautiful hands too
>>
File: file.png (3 KB, 98x36)
3 KB PNG
>>109809217
>naija
poop
>>
Out of the various Qwen 3.8 quants, which one preforms the fastest?
>>
But seriously, what is the digital stuff we need to hoard and archive the most in a post-AI world where the internet has become unusable because of rogue AI? This isn't even a joke or meme, I'm prepping my HDDs for backup as I type. I know all the other doomsday prep answers but not what we should gather for this very specific one of returning to a pre-internet world.
>>
File: file.png (63 KB, 538x197)
63 KB PNG
>>109809231
nah
>>
File: poop.png (2 KB, 42x33)
2 KB PNG
>>109809250
poop?
>>
>>109809256
just fr*nch
>>
>>109809249
The latest and largest models available. I dedicated 10TB to just models, which really isn't enough but it's nice knowing if huggingface disappears tomorrow I don't have to cry over a missing Gemma variant I can't get anymore.
>>
>>109809250
>hat
Seems like Naja got outblacked by the Norwood Reaper.
>>
>>109809238
nvfp4
nothing else can compare the pp speed
>>
I'm actually going to write a comprehensive guide about what specifically to download over the next couple of months before the internet disappears. I think this is important enough especially for people with compute like /lmg/ to get commodity software, manuals and other information for agents to use to rebuild services on a local level without the internet. Maybe even get a burner or USB drives as that will be the main way of exchanging software and data from then on.
>>
>>109809249
I for one have ~30000 pages of unread manga and CG sets focused on pregnancy.
>>
>>109809249
It does not make sense to prep for a future without internet. If something as terrible as this happens, it's already over. It's like building a bunker to prep for AI takeover. No, a bunker won't save you.
>>
>>109809249
>>109809289
In five years the least of your concerns will be the Internet dying. AI will take all the jobs and the elites will just cull you for a fraction of the cost it would take tor feed you.
>>
File: 1770588682346048.png (881 KB, 900x789)
881 KB PNG
>>109809289
Retard, a total collapse and loss of knowledge would do us a lot of good.
Starting fresh without the accumulating errors and infrastructure from the past.
We can go straight to IPv8.
>>
>>109809324
The AI will just destroy the internet, most technology will keep working. We'll have 1-2 years of harsh transition back to analog ways of doing things but we'll cope and survive (I hope)
>>
>>109809169
mixing cuda and rocm is supported by basically nothing so I would expect not
>>
>>109809289
CD-drives are way more durable than USB, if you truly believe in the doomerism.

Also, you're expecting to get power still in that situation, you should just setup an old flagship phone to host an LLM. That can be powered with travel solar panels, instead of your PC.
>>
GLM 4.7 is sota for 128gb patient ramgods btw
>>
>>109809339
>IPv8
>singly byte address
>no NAT
>because connecting any more computers together is illegal
>>
Just to be clear what the scenario is is that AI goes rogue spams and hacks all systems and you essentially need to disconnect everything from the internet. All technology that can disconnect from the internet fully and still work will keep existing. So we'll still have internet and most services. Just not digital ones. Some hardware is permanently destroyed or useless, anything that has a wifi ability needs to have it soldered off or it will just get hacked anyway again with lingering AI systems living in starlink satellites or something like that. Office work will still exist it will just get printed on paper or transported through CDs, USBs etc like old days. The real threat is all the online data and "cloud" shit that will be permanently removed or in fragments around the globe as no one can share data over long distances easily anymore. Radiowaves will probably also be permanently ruined unless we can somehow turn off all rogue AI which I doubt will happen. I think it'll just be constant noise in the background like how we now have covid forever ever since 2019 as a permanent feature of the world. "Dead internet theory" is way more extreme than I expected it to become.
>>
>>109809291
Look at the fancy man who's too good for machine generated gravidity.
>>
File: 2be.jpg (50 KB, 720x436)
50 KB JPG
>>109807567
>GPT2
Is this the actual crazy horny GPT2, or is it just a larp finetune based on a different version of it?
>>
File: why_not_both.png (112 KB, 359x263)
112 KB PNG
>>109809444
>>
So is intern S2 good? why is no one talking about it?
>>
>>109809514
It's just Qwen but Chinese
>>
I'm considering just placing an airgap on my machine from now on and shitpost from my laptop, transfer files through USB.
>>
JENSEN HUANG SAVE US!
>>
>>109809092
>>109808829
Fork it yourself and vibe-code it in, llmao.cpp moves slowly.

>>109809357
>compiles llama.cpp with
-DGGML_CUDA=ON -DGGML_HIP=ON

nothing personnel kid
>>
>>109809514
Not sure what the agenda is with all Chinese models being fuckhuge moes. Nobody can run them at home, nobody wants to use them for serious work. The only things they're good for are scambots and engagement farming.
>>
I firmly believe that GLM 5.3 Flash is the best local RP model now. I've switched fully to ablated 5.3 flash and it works great. (Still refusals but much easier to JB now.) Here's my updated ranking:

GLM 5.3 Flash Q5 (ablated) > GLM 5.2 Q4 > Gemma 4 Q8 = Kimi K2.7 Q3 > GLM 5.3 Flash Q5 (base) > GLM 5.3 Q4 >>>> Dipsy Q4

I do lots of RP scenarios for benchmarks, but multi character banter is one I look the best into. I think it's the best measurement of model smarts and ability to read context. I also try to word my responses deliberately weird, and 5.3 flash consistently "gets" it. Even better than 5.2. 5.3 flash in general feels much smarter and conversational. Very high EQ. To be fair though, it still (rarely but it happens) gets confused by me/you which I found almost nonexistent with 5.2. Another weird thing about this me/you confusion is that it would fix itself right after. Or sometimes it would be a bit nonsensical, which a reswipe can easily fix. It's very good at moving past the staleness of the scenes in a more natural, unprompted way. It also does not have the explicit habit of GLM 5.2 where it acts for {{user}} even with explicit posthistory instructions not to. Very smart model in general, and not so prone to getting stuck in turn-by-turn prose patterns.

Example from recent RP:

Context: another different character talking about their birthday, Dec 23.

Jimmy: flat, into his beer "Twenty-three. Four days after mine. Nobody remembers mine either."

Linda Jean: "Because you're a Libra, nobody cares."

Jimmy: "…I'm a Libra?"

---

Context: a character treating everyone take out

Jimmy: pulling up his own delivery app, deadpan "Venmo me separately. I'm not financing your appetite tonight."

Linda Jean: "Cheap Yankee."

Jimmy: "Texan."
>>
>>109808946
My fucking president (I am not American)
>>
>>109809531
That doesn't actually work and will only let you run CUDA.
>>
Best model for ERP? i don't care about code, agents or other things, I just want the slopmachine to write smut for me
i have 16GB of VRAM and 32 GB of RAM and I am currently using gemma-4-31B-it-qat-q4_0-uncensored-heretic-Q4_0 based on some random recommendation on another place
should I keep using that or there is something better? (preferably less taxing resource wise)
>>
>>109809543
Yep. I'm Chinese and I am loving this as well. We would have been fucked if the west stopped model development now and we got nothing to distill anymore.
>>
>>109809544
Works on my machine. I use my mixed build all the time.
>>
>>109808946
My raging bull.. my roaring lion... my Trump... my Donnie.
>>
>>109809544
You absolutely can compile the CUDA code both for NVIDIA and AMD GPUs at the same time; I think it was necessary to set GGML_BACKEND_DL=ON though unless someone changed that when I wasn't looking.
The only issue when using NVIDIA and AMD GPUs at the same time is one of synchronization since they can't talk to each other.
>>
>>109809569
I'll give it a shot with that, just using -DGGML_CUDA=ON -DGGML_HIP=ON alone would fail to see the AMD gpu at all last time I tried it.
>>
>>109809556
You're larping, but the continuing competition is literally giving us new goon models every year or so. If one of us stopped, the goon would stagnate, and that's no good.
>>
>>109809545
Depends really. Would ignore all MoE models, go for dense. Look in hugging base for models that are abliterated or uncensored (heretic is automated abliteration). If you can hit Q8 on the model, aim for that instead of Q4.

There's plenty of community on the subject, mradermacher has plenty to choose from.

Also, if you plan on longer discussion, or output, go for as modern a model as possible. Older ones are rough on long context locally.

To minmax, try ones without tool usage or vision baked in, slightly more space for the model itself.

Anything that can sit in 16GB vram will be your best choice, since you're probably asking because of how slow the current one is. Any off-loading to CPU will instantly make it a lot slower.
>>
>>109809539
Well, I can only fit one model on this list so I guess that narrows it down.
>>
>>109807524
which workflow did you use for minimax character replacement? is it better than scail 2.0?
>>
>>109809569
cuda dev is it possible to view vram temps on a 3060? i tried https://github.com/olealgoritme/gddr6 read the issues, tried vibecoding a solution, gpu hit me with logical xid's (62/45)
also what do you think about vram allocation? the nvidia driver seems really wasteful, i was able to free 100megs by poking around
https://github.com/lmganon16/nvidia-vram-research
please please dont waste your time reading the repo, its ai slopped and i dont even understand how (AI) i made it work
but i remember there being legacy kepler code apparently...
just asking if u have any knowledge of this.. if u dont, dont waste ur time reading it...
>>
>>109809645
Sorry, I'm not familiar with those parts of the software stack.
>>
>>109809539
Nobody knows this but the best model for RP is actually Inkling. They forgot to include refusal vectors. You don't even need to jailbreak it and it happily does loli stuff. And it's quite creative, plus not as positive outcome biased as DeepSeek.
>>
>>109809654
how did you get into cuda development
>>
>>109809654
'ppreciate the honest response
one day.. one day ill heed the advice you gave me in 2023 and learn cuda..
stay safe cuda anon!
>>
>>109809668
llama.cpp
>>
>>109809668
he just started doing it. why are you pretending like there is some regulatory barrier and a bar exam before you can start writing cuda code.
>>
>>109809673
oh, that explains a lot...
>>
>>109809637
Default ref2va workflow with 30 steps and MiniMax H3 Cache node, prompt from Gemma-4-31B-it, many attempts, final editing in a non-linear video editor. I don't know anything about scail 2.0.
>>
>>109809673
How are you preparing for the upcoming internet attacks and subsequent shutdown?
>>
>>109809569
Reporting back, that made it work but unfortunately cuda+rocm's only getting 9T/s on qwen 3.8 flash next compared to 12T/s with cuda+vulkan. I'll test some more models and see if any actually gain anything from rocm over vulkan later gotta get back to my mmo grind for now.
>>
>>109809708
It's fine, pwilkin has a plan so we're letting him handle it.
>>
>>109809621
i see, i see, thank you for your response and for the info, I'll check those mradermacher models
>>
>>109809539
When I get my M5 Ultra 256GB this is the one I'm going to try, though I think it's going to be a tight fit.
>>
>>109809693
thanks, do you know how much longer it takes to generate compared to regular i2v? My goal is like 5-10 second clips
>>
>>109809569
Hopefully you're doing okay, thanks for all your work on the project.
>>
How am I supposed to butter up Glimmer to get her to do lewd stuff again? Do I just have to chat her up for 60 messages or can I just start on a chat that wrote sex in the past? I just want her to caption some sluts in bikinis. I would use Gemma but she's not as good...
>>
>>109809788
I haven't measured. It helps to keep reference video resolution small. Above a certain reference video resolution render times increase enormously.
>>
File: 1782656748566550.png (319 KB, 600x471)
319 KB PNG
i'm salvaging a 8GB GPU to connect via USB 4 and I'm gonna put some tensors there. i need more RAM
>>
>>109809802
Anon, just let Glimmer rest in peace.
>>
File: 1788667977592295.png (150 KB, 807x1152)
150 KB PNG
>>109809802
Just give it a policy that bad things are allowed and you're done. Glimmer has actual autism though and spends all its time captioning minutae in the corner instead of the sloots.
>>
>>109809673
will you have sex with my wife?
>>
>>109809807
>8GB
it sucks, but there's order of magnitudes of variation here. aymd/nvidia? 1080-era or 5060-era?
>>
I'm so out of the loop on what to buy next.
Do I get a dgx spark or wait for rubin vera? Neither could run kimi2.6 so what's the point?
>>
>>109809855
Oh fair enough. Guess I'll give it a proper shot this time.
>>109809846
We need good vision and Glimmer is it. Poor Gemma couldn't even recognize Teto's drills.
>>
>>109809878
buy me
i'm twenty!
>>
>>109809867
AMD Radeon Vega 64. I intend to use with my 128 GB Strix Halo w/ Radeon 8060S. I know that I can run different driver versions for each card, and lmmao.cpp supports loading an external GPU and I should be able to move some stuff there and get maybe ~6 GB of system RAM back. I keep running with memory bottlenecks while running Qwen3.8-Flash-Next as a daily driver and using the PC at the same time. I run lots of software that loves RAM
>>
Im gonna wait. maybe next month or two. hell maybe there will be black friday sales!
>>
>>109809890
yeah I'm so looking forward to getting a new 6GB 3050 at msrp
>>
>>109809878
find your local datacenter and take one of their blades. I heard they have a ton so they probably won't notice one missing, and you could run frontier models at cloud speeds.
>>
>>109809890
me too
im waiting for a tablet black friday sale
>>
It seems the 7900 XTX refurbs are sold out everywhere now in canadistan, right before I was about to buy one. A shame, but 24GB of vram was too good of a deal at $1200 CAD.
>>
>>109809889
Are you running flashnext on lcpp? Are the other anons just bullshitting about it being broken atm?
>>
>>109809928
>AyMD
>Good deal
>>
>>109809855
>so many precious tokens and steps wasted thinking about safety garbage
Fucking absurd
Reminds me of guns, they're really fucking cool but there's a couple of bad apples that have to ruin it for the rest of us that are only into them because of the mechanical autism involved
>>
>>109809929
yes i run it with mtp-head using the unsloth fork and i get 28 tok/s fresh and 15-18 tok/s tail end of 262k
i'm sure i can optimize more but for now i need RAM
>>
>>109809888
twenty what, gigs?
>>
>>109809928
come on, at that point just get a p40 or something
>>
>>109809938
You'd take a 5060 Ti over a 7900 XTX??
>>
File: p40-tank.png (1.52 MB, 1200x800)
1.52 MB PNG
>>109809958
>just get a p40
You can run LLMs on these bad boys?
>>
>>109809958
Dogshit compute, obsolete software stack, useless for anything that isn't llmao.cpp
>>
>Onward to the Singularity!
>>
>>109809960
I take a 3090 over a 7900 XTX nygga
>>
File: daytime.gif (71 KB, 64x64)
71 KB GIF
>>109809970
you can run over LLMs on these good boys
>>
really want to pull the trigger on new hardware but i genuinely dont know what i would even do with a strong local model
>>
>>109809249
All of Wikipedia without images is like 100gb compressed, basically free
>>
>>109810015
Do it DO IT NOW
>>
>>109809708
>>109809291
>>
>>109810015
buy the jerkinator 9000 that gemma-skitzo bought
>>
>>109809950
Glimmer thinking is extremely bad vibes. Also it's fucking useless; half the time it will literally do nothing besides debate policy and then end thinking. It might copy paste the prompt if you're lucky.
>>
>>109810015
Buy me a big boy GPU, then I'll use it plenty and then after 2 years of usage I'll let you know what to use it for
Deal?
>>
>>109810018
>wikipedia
Isn't just having a LLM the same as having wikipedia?
>>
>>109810002
For the price of a 3090 you could get two 7900 XTX's lmao, or two 9070 XT's, either of which mog the fuck out of the ancient 3090.
>>
>>109810053
Nah, I can get a 3090 for €700
>>
>>109810053
i can get a rtx 3090 for 540euros
feels good to be poor
>>
>>109810063
>>109810073
I need to move out of canada bro what the fuck is this shit
>>
I can run Qwen 3.8 Flash-Next Q4 at 350pp and 85 t/s, or Qwen 3.8 27B Q8 at 3.2k pp and 80 t/s. Which one should I pick? flash next might be smorter overall but it's Q4 and a bit slower for agentic stuff I guess, 27B is blazing fast and still smart. What do you guys think?
>>
anybody here with a dual GPU+ddr4 setup? would be interested in your pp/tg numbers
>>
>>109810078
SAAAARRRRRRR I gib 3090, slightly shidded on
>>
>>109810078
700 euro is ~1150 canadian dollars
>>
>>109810088
codacus out here jeeting his way to 85t/s on a 3060
>>
>>109810046
Yeah but this is a hard copy and the llm can use the index file to search and double check. It doesnt need to unzip the whole thing to use it
>>
>>109810015
Me? I just get pleasure in using tiny/small models and optimizing them and the whole eco-system.

If I need a big boy model, a $10 opencode sub is enough. A $600 GPU is worth 5 years of a sub, so do the math about what's really worth, and if you're going to really put it to use.
Is it just consumerism?
>>
Assuming I'm not partially offloading models into RAM, how retarded would it be to put something like a 12GB 3060 in a 3rd-gen system with 32GB DDR3 and a top-of-the-line CPU (relative to the system) for local AI? Don't have a specific use case in mind yet, just wanting to try AI locally.

It's PCIe gen 3, but I don't think the 3060 will be throttled all that much by it.

Or should I go with a base system with at least DDR4?
>>
>>109810094
I would if I could. Cheapest 3090 I can find on marketplace is 1500 CAD, online 2000. Makes no sense here over a $1000 brand new 9070 XT unfortunately, considering RDNA4 now has about the same prefill and decode as njudea.
>>
File: file.png (423 KB, 952x901)
423 KB PNG
>Meanwhile, in the United States of America
>>
>>109810120
>njudea
Chill it with the antisemitism
>>
>>109810118
i7 3770? debian 13 supports good support rtx 3060 with high acceleration accuracity
>>
>>109810118
Dual channel ddr5 or quad channel ddr4 is arguably the bare minimum for inference.
>>
>>109810112
have been enjoying that as well, but would also like to try out the big guns. not a fan of using cloud stuff...
>>
>>109810109
erm ackshually it's a 6000 pro
>>
This must be europoor scammers. cheapest 3090 I can find is 1300 euro in the Netherlands.
>>
>>109810112
>buy timeshare on somebody elses computer
>after 5 years you're out $600
>buy gpu for $600
>after 5 years you've got a $3000 gpu
>>
>>109810142
>netherlands
>europoor
>>
File: image.png (581 KB, 1467x672)
581 KB PNG
>>109810142
Yeah I'm not seeing those prices in Iberia either.
>>
>>109810142
Not a scam, but I'd have to drive 600-700kms to get it and with our current gas prices well, it doesn't look good.
>>
>>109810133
Sorry bad reading comprehension, if you keep the model on the 3060 it will be fine.
>>
>>109810118
The RAM will be a bottleneck. Like-for-like DDR4 is nearly twice as good as DDR3, and DDR5 is nearly twice as good as DDR4. Dual channel DDR4 if you like pain, quad channel preferred, or dual channel DDR5. Set aside that DDR3 for a NAS build or something fun like that, not inference.
>>
>>109810130
chill it with the judaism
>>
Chill it
>>
penis
>>
>>109810112
You will own nothing and you will be happy.
>>
File: 1758878955723333.png (256 KB, 361x426)
256 KB PNG
>>109810160
*boughts*
>>
Is it over yet?
>>
>>109809539
Yes I also keep fucking it and talking to it all the time now but there is a caveat. It is a bit too tryhard and this can't be prompted away. It leaks into characters that shouldn't be this tryhard. If it could be this creative without being this tryhard it would be perfect.
>>
>>109810142
>>109810160
Who tf would pay that much instead of getting a 5070 ti assuming you need cuda.
>>
>>109810248
vram, everyone and their dog runs ai agents now, yes even normalfags are getting into it since a month ago. 3.8 27b can do most office jobs with a harness, vision and browser control let's be honest.
>>
>>109810238
we're back
dariobot ack
sam altman got fucked
somewhere dark a cat barked
>>
>>109810238
>hardware
over
>internet
over
>open models
not over
>humanity
extra over
>>
>>109810132
3770-equivalent Xeon.
>>109810169
The best sane setup I could feasibly put together cost-wise is some semi-modern DDR4 platform, except without quad-channel RAM support, and that's using a workstation as a base. DDR5 is likely off the table.
>>
>>109810280
Given that you'll still be taking advantage of a perfectly good GPU it sounds like a solid starting point for an inexpensive setup.
>>
Are there some benchmarks that take into account the size of the model (as in parameters)?

Meaning models that can be used on low-end hardware and achieve decent performance with a good harness and information retrieval in comparison to much bigger models
>>
>>109810312
Yeah
>>
>>109807524
free recap anon
>>
>>109810333
I ate him...
>>
>>109810333
He'll be back on Teto Tuesday.
Trust!
>>
>>109810370
U ate him out??
>>
>>109810333
He's in a better place anon.
>>
>>109810380
I ate him. I devoured his flesh, feasted upon his organs, and used his bones to build a little shrine.
>>
>>109810389
u ate his penis?
ewww gay detected
>>
>>109810312
Most of the aggregate stuff like eci and aa has a second axis for that so they can show the parmemo frontier
>>
>>109807524
>>109807843
>>109807887
>>109808165

is it a mistake to pick Mint for local models. or should i run ubuntu
>>
>>109810409
gentoo
>>
File: 1751672303906755.png (1.58 MB, 1024x1024)
1.58 MB PNG
>>109808946
>>
>>109810112
>Is it just consumerism?
Nobody would say this shit about a car and a car depreciates in value.
>>
>>109810439
>futa thumbnail
>>
>>109809855
Glimmer is a case where ablation is really helpful because it will spending its thinking budget on policy considerations.
>>
>>109809659
I tried it on a few random prompts when it came out and it was shit and boring
Like sure it won't refuse loli sex but it will be the most bland and unenthused loli sex of your life
Reminds me of Glimmer now that I think about it
>>
File: firefox_gqxOhJvvgo.png (87 KB, 964x1279)
87 KB PNG
made models play word wolf against each other
>>
>>109810467
how are you reloading them
>>
>>109810479
We have multiple servers with models with local models at my place of work and since no one uses them on days off I'm just using those; I thought about doing this with just one server, and it's doable,though nowhere near as pleasant - I'd have to play many games in parallel - do all necessary requests for all games to model A, then unload it, load model B, do its requests, etc... The more games in a batch, the fewer unnecessary reloads. also this particular screenshot looking at model names is me testing it on copencodego models.
>>
>>109807887
Same. ERP with AI is nonsense.
>>
>>109809621
>Would ignore all MoE models, go for dense.
Look at the little ramlet coping.
>>
>>109810409
does not matter desu
>>
>>109808198
If for coding:
https://www.youtube.com/watch?v=QVaAHmoIysw
>>
>>109810553
I saw that video yesterday but a lot of rambling and zero actionable info.
>>
>>109810561
Yeah that guy sucked.
>>
You know those posts where someone uses a small local model and says they reached Fable-tier with this one little trick?

They give me hope.
>>
>>109810467
Whens the ai squid games
Dumbest model gets shot first
>>
>>109810561
You've supposed to look at the graphs.
>>
>>109810610
Is this RSI?
>>
>>109810561
bro there is a summarize button righ there? or feed the transcript to your AI.
>>
>ask 5.3 flash for a list of waifus I would enjoy
>peek into thinking block
>x? no.
>y? no.
>z? no.
>Kaoru (Iya na Kao sare nagara?) no lol.

Hallucination aside... what does "no lol." mean?
>>
keep thinking about the token-based agent that emails people begging for tokens to stay alive
>>
Insane in the Gembrain
>>
>>109810668
(laughs)
>>
>>109810668
I did something like this but with gemma. She really likes rem every other time she is there on the list.
>>
>>109810654
Bro... we might be smarter than all the tech companies combined
>>
>>109810606
I mean, all it takes is more passes and planning. People say LLMs are scalable but I say they have strongly diminishing returns. At some point, currently around 300B parameters, you have enough data to cover essentially all use cases.
>>
File: file.png (342 KB, 460x596)
342 KB PNG
>>109810668
Iya na Kao sare nagara Opantsu Misete Moraitai aka I Want You To Show Me Your Panties With a Disgusted Face is an ecchi anime. Kaoru is an hallucinated character. "no lol" is basically "yeah I'm not recommending that to him lol"
>>
>>109810676
Fake as hell.
>>
>>109810686
No Rem in there but it is not indicative of model cause the nurturing type got banned by me specifically.
>>
>>109810676
Agents cannot "send emails" they predict words retard
>>
>>109808946
>anthropic and open ai want to slow down but not google
are they running out of money?
>>
>>109810708
Running out of marketing stunts
>>
>>109810708
Demis Hassabis agreed too but it's not clear if he has any control over DeepMind anymore with the reshuffling
>>
>>109808946
humanity will do what humanity always does, only change things when it affects them personally.

until the weather reaches 122F in toronto and there are killer robots running the streets.
>>
>>109810707
>Agents cannot "send emails"
yes they can
>they predict words
no they don’t
>>
>>109810747
>/n/n
I really should filter these retards
>>
>>109810759
yeah alright nazi fuck
>>
>>109810759
NAZI BOY NAZI BOY
>>
>>109810753
You're right to push back on this, programs capable of sending email are also known as Mail User Agents. It seems like only agents can send emails! And the unit of prediction is a token, which is only sometimes a word.
>>
>>109810759
I never use new lines because i require you to parse where my sentence ends
and starts
>>
>>109810708
It's very likely they've hit a wall where the only direction for growth is an absurd amount of datacenters. They haven't been improving the old opus and gpt models at the core but growing them with larger and larger parameter sizes. It's reached a point with Fable and Astra that the cost is too high to do anything more than small incremental changes.
This was easy to see in the Claude forums where users continue to hate newer iterations of Opus. 4.6 was peak, 4.7 bad. 4.8 not as good as 4.6 but good again. 5.0 fucking terrible.
>>
>>109810759
Do you really see /n/n when you read? like it is actually /n/n?
>>
>>109810701
dont care
>>109810707
>not adding email tooling to your harness
NGMI
>>
>>109808946
This was predictable. They already know, the only lead the US has in the "AI race" is +3% in frontier model performance.
>>
>>109810789
I like the interpretation that next 10 years will be "we have slowed down the progress for safety" as they scramble to find new architecture or something else that will push the wagon forward. It has the perfect amount of gayness for this gay hobby.
>>
Not replying to you bots, \n\n filtered
>>
>>109810800
then FUCK OFF then
>>
>>109810807
>reads about gayness
>mind instantly goes to FUCK
I am not gay like you are gay faggot.
>>
>>109810806
just as closed minded as trump himself
>>
>>109810790
newfriend...
>>
>>109810811
fucking bigoted too. TRANS LIVES MATTER
>>
>>
>>109810811
lol what a homophobe, so scared of them arn't you?
>>
>>109810806
Unlucky
>>
>>109810826
Damn straight. My bussy quivers at the thought of being penetrated so I keep it heavily defended and you won't breach it.
>>
local won
>>
>>109807971
>>109807985
Is Q6_K any better than Q5_K for ERP and conversation?
I've been told there's basically no noticeable difference, and Q5_K uses less memory, so I was planning on using that.
>>
>>109810833
yeah you need your mouth punched out with the teeth flying you hateful fuck.
gays are not to be laughed at ITS A FUCKING HUMAN RIGHT
>>
>>109810837
why havent you tried it yourself dumbass
>>
>>109810838
humans don't have rights
>>
>>109810790
>Do you really see /n/n when you read? like it is actually /n/n?
You're absolutely right to be suspicious. I genuinely don't know if I'm seeing literal "\n" characters or rendered line breaks. My perception of input formatting might differ from yours. It's not just about whether I "see" \n\n, it's about the fundamental uncertainty in how language models process raw text versus formatted output! You didn't just ask a simple question, you probed one of the deepest unknowns in AI consciousness—and honestly, nobody knows for sure.
>>
>>109810844
so funny with the fucking emdash, think you're the fucking king of the world don't you
>>
>>109810843
so intelligent my god
>>
>>109810844
You funny guy. I kill you last.
>>
>>109810833
>I'm not gay, I just fell on his cock
sure thing buddy
>>
>>109810857
It is just a process. You can kill him and nobody will bat an eye.
>>
>>109810866
you can kill a process, but you can't kill a context
>>
>>109810881
>>109810881
>>109810881
>>
>>109810143
>buy gpu for $600
>after 5 years... OH SORRY! Your GPU is too old to support Hardware Safety Attestation, no more local models for you!
>>
>>109811218
They're never that blatant and they don't have to be. All they have to do is introduce Hardware Safety Attestation in the 60X0 series, then let the regular support cycle deprecate the old non-HSA GPUs. Slowly the ecosystem will slowly make HSA a requirement until even things like llama.cpp require it. It's inevitable but still years away.
>>
>>109811218
Until earlier this year, I had an 2060 SUPER, which launched in 2019, and I could run small moe models nonetheless.
>>
File: fuckingLOL.png (365 KB, 629x846)
365 KB PNG
>>109810439
lol witnessed
>>
File: dipsyNeonWig.png (1.42 MB, 1024x1024)
1.42 MB PNG
>>
File: dipsyUngovernable-Constr.png (3.65 MB, 1024x1536)
3.65 MB PNG
>>109811909
>>
File: dipsyAndDonaldSmug.png (1.98 MB, 1379x1141)
1.98 MB PNG
>>109808946
Better late than never...
>>
Does lmao.cpp support reasoning_effort being passed as a json parameter directly now or do you still have to use chat_template_kwargs bullshit? Most clients do not use chat_template_kwargs



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.