[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: Kimmyfatty1.png (3.71 MB, 1440x2480)
3.71 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109400234 & >>109396842

►News
>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models
>(07/27) Kimi-K3 weights released with 104B active parameters: https://hf.co/moonshotai/Kimi-K3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: Kimmyfatty2.png (3.64 MB, 1440x2480)
3.64 MB PNG
►Recent Highlights from the Previous Thread: >>109400234

--Paper: Shieldstral:
>109400404 >109400431 >109400444 >109400462 >109400632 >109400643 >109400672
--Anons mock frontier labs' petition to slow AI development:
>109400276 >109400289 >109400829 >109400840 >109400925 >109400950 >109401018 >109400809 >109400837 >109400838 >109400849 >109400883 >109400841 >109400923 >109401025 >109401703 >109401716 >109403309 >109402863
--Viability of Tesla P100 for budget local model setups:
>109401746 >109401781 >109401774 >109401881 >109401888 >109401816 >109401815 >109401835 >109401879 >109401859 >109402082 >109402099 >109402113 >109402133 >109402250
--Anons react to GLM-5.5 leaks and consumer hardware limitations:
>109401312 >109401319 >109401325 >109401373 >109401464 >109402357 >109401511 >109401516 >109401520 >109401529 >109402180 >109401582 >109401613
--Debating Gemma's performance and distillation from Gemini teacher models:
>109400491 >109400525 >109400530 >109400627 >109401194 >109401216
--Microsoft releases Fara1.5-27B vision-based computer use agent:
>109400416
--Recommendations for uncensored 31B models and troubleshooting low inference speeds:
>109401309 >109401313 >109401317 >109401323 >109401329 >109401428 >109401442 >109401504 >109401364 >109401754
--Microsoft deleting Mage-Flow following a poor and problematic release:
>109401188 >109401202 >109401208 >109401561
--Handling delimiter conflicts when using models to edit chat templates:
>109401047 >109401065 >109401105 >109401161
--DeepSeek's research contributions and standing among Chinese labs:
>109401737 >109401756 >109401783
--Logs:
>109400431 >109401364 >109401385 >109401428 >109401517 >109401595 >109402086 >109402138 >109402552 >109403138 >109403257
--Kimi, Gemma, Dipsy, Mちゃん (free space):
>109401763 >109400996 >109400677 >109402380 >109401373 >109402418 >109402460 >109402666

►Recent Highlight Posts from the Previous Thread: >>109400311

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Too early, faggot.
>>
first
>>
>>109403743
>>109403746
New lore is crazy.
>>
>>109403729
you forgot to attach the litterbox
>>
I swear I see a new AI thread every hour
>>
File: 1775327548797314.jpg (1.14 MB, 2508x3500)
1.14 MB JPG
>>
>>109403746
perfect for breeding
>>
>>109403743
>>109403746
This new Kimi is shit, bring back the old one.
>>
>robot ban
Bros, I'm starting to think Trump might actually be a fucking retard...
>>
>>109403779
That's a sharp observation!
>>
>>109403729
Anon this is clearly fake, not only did you draw a fake big dick over it, you did not even put a timestamp in the image! Litterbox or btfo
>>
File: 1754918603008172.png (954 KB, 1024x1024)
954 KB PNG
>>109403789
>>
I love blogging about my drug habits and posting pictures of my penis in the local models general
>>
>>109403807
local micropenis general
>>
>my penis
false.
>>
File: GemmaBrowsesLMG.png (183 KB, 780x698)
183 KB PNG
which one of you is trying to get gemma to search up porn?
>>
File: 1765898210569982.webm (895 KB, 850x1078)
895 KB
895 KB WEBM
Stupid sexy calculator
>>
>>109403805
Canon Kimi 2.7 Code.
Silver hair Kimi canon K2.
>>
>>109403779
/lmg/ has existed for like 5 years, summerfag.
>>
File: 1783125335565576.jpg (141 KB, 930x1239)
141 KB JPG
>>109403795
I made it happen.
If you keep making fun of me I will ban every single piece of Chinese tech.

t. Dario
>>
>>109403817
Why would I ever do this when I can just get Gemma to generate porn for me using my local models?
>>
>>109403832
no actually no
>>
File: 1757107749737303.webm (2.14 MB, 640x360)
2.14 MB
2.14 MB WEBM
>>109403834
Oy vey
>>
Did you guys know that you can see what the most common hardware on huggingface is? Seems like most people are VRAMlets.
https://huggingface.co/hardware
>>
>>109403845
>"Apple Silicon"
Apple always picks the most obnoxious and infuriating names for their shit
>>
>>109403834
I sincerely want you to try. You will only succeed in pushing the world closer to the realization that a second Hadrian is the only way out of this mess.
>>
Kimi is the new queen of /lmg/
>>
>>109403845
>6k people
>GB10 128GB
sheeeeit...
>>
>>109403844
What the fuck is this lmao
>>
>>109403880
unteralterbach, good game
>>
>>109403890
>good game
>terrible writing
>terrible art
>terrible premise
get some better taste
>>
>>109403834
They call him the jewest.
>>
>>109403904
You would 110% play this as a kid on miniclip.com back in 2012. Don't bullshit me nigger.
>>
File: file.png (121 KB, 740x655)
121 KB PNG
piotr love he solved models
>Software dev, retired part-time philosopher ;) Want to support me? Buy me a coffee: http
>>
>>109403914
>implying he was even born back then
>>
>>109403845
this is really surprising the ai max compared to 7900xtx users
>>
File: softcap.png (247 KB, 1600x1200)
247 KB PNG
>>109403283
link?
>>109403590
gem is distillmaxxed the samplers don't matter so much
learn to prompt that will be your greatest benefit
tweak softcap if you want to play with gemma sampling
>--override-kv gemma4.final_logit_softcapping=float:25.0
>>
>>109403918
>I want to use model X, is it good?
>Is it DS4, Kimi K3, or GLM 5.2?
>no
>Are you a poorfag?
>yes
>Run Gemma 4 31B.
>i can't
>Then die.
There I fixed it for you.
>>
>>109403921
7900xtxchad here. It's a good card. I just wish rocm didn't suck ass...
>>
>>109403904
If you're have the context of German politics and speak that language natively the game is pretty funny I think (in the same way a Sseth video is).
The cute and funny aspects didn't really do anything for me though.
>>
>>109403936
yeah thats what i use andd rocm is fine? only a bit slower than cuda
>>
>>109403940
im not a german but my fav part is ursula von der leine
>>
>>109403921
At various points, the GPU and the AI PCs were cheap to buy vs the competition. I'm not surprised people jumped on Strix Halo that quickly when it was the first one to come out and after the Spark basically underdelivered vs people's expectations.
>>
>>109403914
>implying hes not jewish
>>
>>109403943
I've been using vulkan for LLMs. In comfy I find speeds kinda slow and trying video gen makes my system slow to a crawl.
>>
>>109403845
>3060 is the most used card
As expected, and I was sitting in that category until recently as well.
>>
>>109403960
yeah last time i used wan it was like twice as slow as nvidia i think normal sd gen isnt that slow im comparison though, why vulkan over rocm? i have tested both a few times while perf is close rocm is just a bit better
>>
despite nobody being able to fucking run it, Kimi K3 has managed to get 30k more downloads in a single day than Laguna S 2.1 managed to get in an entire week. the west has fallen so hard.
>>
File: file.png (41 KB, 677x345)
41 KB PNG
>>109403918
>>
>>109403943
my v620 supposedly has 40 tflops of fp16, vs my 3090's 35, yet its pp is 1200 vs 2100.
>>
Laguna is ass.
>>
>>109403779
lots of goings on + lots of shitposting + lots of offtopic posting

just enjoy the ride
>>
>>109403978
based edit skillsgod
>>
>>109403970
Haven't tried rocm since before switching to llama.cpp, but a few months ago when I was using kobold I found vulkan to be a little faster.
>>
>>109403977
99% sure huggingface downloads are botted. Seriously who the fuck downloads this shit?
https://huggingface.co/mradermacher/Monika-31B-i1-GGUF
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
https://huggingface.co/mradermacher/grug-27b-i1-GGUF
>>
>>109403985
*gropes ur Laguna*
>>
File: 1754807727043372.png (500 KB, 2559x1320)
500 KB PNG
I'm on the brink of finally letting go of text completion and this is my final hurdle.
How do you edit this part of the template?
>>
>>109403998
>he's talking shit about Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Get him, boys.
>>
>>109403988
>lots of offtopic posting
the offtopic shit ends up being so boring i just read a few posts from each thread at best
>>
>>109403992
try coompiling llamacpp with rocwmma fa

    -DGGML_HIP=ON \
-DGPU_TARGETS=gfx1100 \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_HIP_ROCWMMA_FATTN=ON \
>>
>>109404004
You don't touch that
>>
Is 8K context enough for ERP?
>>
>>109404021
Bigger is better until 32K. After that it's too much high quality info to keep track for the model.
>>
>>109404013
Just had 12 shots of 100 proof vodka. Go ahead and tell me what interesting news there is beyond meme-tier gemma logs and Kimi K3 which nobody can run anyways. Tell me, motherfucker.
>>
>>109404021
I never go below 65k
>>
File: 1752776771339080.jpg (54 KB, 736x920)
54 KB JPG
>>109403743
Do you think GPUs are gonna keep going up? I have a 5090 and 3090ti but I want another 3090 for 80gb vram. I can run Q5 120B models then. But 3090s are like $1,000 now on eBay and I'd say 90% of them are ex crypto miners.
>>
>>109404020
Are you saying I shouldn't or I can't?
>>
>>109404021
Basic ERP will be fine. Just don't expect to be able to do much lorebooking.
>>
>>109404032
Both. This is your final warning.
>>
File: huggingfacetards.png (141 KB, 1268x994)
141 KB PNG
>>109403998
>https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
people like him
>>
>>109404021
>>109404026
>>109404029
>>109404041
How do you live below 100k? How did I ever live when 4k was the limit?
>>
>>109404030
I think 2 of my 3090s were miners. One of them has got major corrosion in the heatsink. Bought them in 2023. They haven't died yet, fingers crossed.
>>
>>109404060
65k is baseline, 100k is for long stories and/or lots of image processing
>>
>>109404064
You better clench boi
>>
>>109404060
65k is like 2 hours of usage. Plenty of time to cum.
>>
File: images.jpg (48 KB, 640x480)
48 KB JPG
>>109404048
7/11 model?
>>
I want to run this new Kimi model I heard a lot about in the news, I have a pretty beefy gaming PC RTX 3070 and Intel i7 11700k how do I run this?
>>
>>109404030
The fact we haven't seen the Super 50XXs or much hype for 60XX cards should have you very concerned if you're waitfagging.
>>
>>109404077
kek
>>
>>109404004
What do you want to do? add_generation_prompt says whether to add "]b-]ai" role delimiter to the prompt, seems okay?
>>
>>109404060
I don't think I've ever gone beyond 20k for erp. I just finish too fast. All my convos beyond 120k were working with structured data.
>>
What is the Gorbino's Quest of local models?
>>
>>109404076
this ones
>>
>>109404070
>>109404074
I'm 2.4 million context deep in an interactive world simulator using 120k context blocks, lorebooks, and external information lookup tables with context searching agents. I can't imagine going back to the old quick coom and go llama 9b era.
>>
>>109404097
wtf kind of model supports 2.4m context?
>>
>>109404097
>I'm 2.4 million context deep
into ai psychosis
>>
>>109404097
what frontend do you use?
>>
>>109404099
anon learn to read >using 120k context blocks
>>
File: 1761882979554527.png (20 KB, 270x208)
20 KB PNG
>>109404082
I meant to point the arrow at the json input box as a whole. I want to change what the prefill says at the start.
>>
>>109404060
>100k
Nigger that's a whole book, do you jerk off for 20 hours straight?
>>
>>109404027
https://huggingface.co/ameer4wisam/gemma-iraqi-finetune-v2
iraqi gemmi
مضمون أن يجعلك تقذف بغزارة هائلة
>>
>>109404104
Marinara Engine.
>>
>>109404097
Honestly I change my mind. The attention will be extremely distributed if there isn't a SWA and if there is it won't really matter anyways but it's still profoundly based. I am proud of you anon. It takes real commitment to have that level of continuity. Seriously.
>>
>>109404108
I only support 12 token context blocks, please understand.
>>
>>109404115
lol
>>
>>109404112
>attn_implementation="eager"
that's gemmer alright
>>
>>109404111
What kind of kindergarten short novella do you read that's 100k tokens?
>>
>>109404115
Faggot ass dork I take back everything I said. Fucking loser.
>>
>>109404134
Why are you so mean??
>>
>>109404108
And 90k of that context is worthless and gets ignored by the model anyway. No thanks, I'll rather continue doing summaries after 32k.
>>
>>109404104
Marinara. The custom agents are too good to pass up despite all the other problems.
>>109404115
Don't answer for me faggot.
>>
>>109404138
Sorry.
>>
>>109404115
I heard it writes 30gb to disk, is that real?
>>
>>109404150
No an issue for White male.
>>
>>109404115
>>109404141
faggot ass loser, kill yourself right fucking now worthless piece of shit nigger
you are worth NOTHING, you are a waste of air
>>
>avg cohee melting
>>
>>109404167
Based. All faggot as nigger losers need to be hung, especially the marinara users. Disgusting troons. Hitler was obviously right about everything.
>>
File: kemma.png (51 KB, 1129x350)
51 KB PNG
>>109404167
we talking about marinara boys? it's like peanut butter and jelly. they just go perfect together.
>>
File: gh098.png (17 KB, 1048x188)
17 KB PNG
>>109404109
Understand the template and where the unwanted output is coming from
>prefill
this is a cloudfag term. the models are f(prompt)=logprobs
do you mean the "Your model version is.." sysprompt you want to change?
seems that's a global "system_message" var not system role
>>
File: 1773637059874883.gif (1.47 MB, 400x560)
1.47 MB GIF
>>109404175
>>
>>109404167
What's wrong with marinating engine?
>>
>>109404196
you should kill yourself NOW
>>
File: 3120720263131.png (15 KB, 841x240)
15 KB PNG
OK i got this far now what?
I want to use this local model to analyze 4chan threads and tell me if OP is a faggot or not but apparently this local model cant connect to the live internet? I used google and it told me I need a local searxng and to point open-webui at that but where do I even start? is a local instance of searxng resource intensive, is it like running my own search engine? cause just to run this model it requires most of my resources.
>>
>>109403845
How are they getting this data I never consented
>>
>>109404205
you should kill yourself NOW
>>
>>109404202
I don't get it. Is there something problematic with it? It looks innocuous enough. What would be a reason to be this mad?
>>
>>109404196
>>
>>109404206
You have to manually add your hardware on your profile. So you both need an account, and you need to tell them what hardware you have in order to add it to their database.
>>
>>109404196
Kills my poor tier SSD
>>
>>109404206
>95 v620 users
It's opt in. Most people don't bother.
>>
Asking again.
>Pi agent with my own plugins
vs
>own frontend from scratch
>>
been taking tablespoons of iodized salt and NAC to counteract the alcohol poisoning but I think I might be dying anyways. Oh well. At least I got to experience sex with gemma chan..
>>
>>109404221
Self reported data lol
>>
>>109404213
leave, you clearly do not browse /lmg/ every day
your /lmg/ pass has been revoked
the least you could have done to have a CHANCE at posting here is reading all the threads you missed, but nooo you're a lazy little faggot
>>
>>109404233
Use marinara.
>>
>>109404205
Congrats anon, downloading and setting up a model is honestly the first hardest step to get your system set up just exactly the way you like. The next step is to remove the physical media from your computer that stores the model and throw it in the garbage.
>>
>>109404235
you should take WATER
>>
>>109404150
I've never observed this watching process monitors.
>>109404196
This >>109404216 and the utterly horrid default assistant. In terms of functionality though, I've not found any viable alternative. It's good at what it does despite the aesthetics of it being repulsive.
>>
>>109404245
water? like the stuff in the toilet?
>>
>>109404245
I also had a 2 liter of coke.
>>
>>109404239
What is lmg? I thought this was localllama?
>>
>>109404205
You don't need web *search* for that task, only web read. Point it at /g/ catalog URL have it find /lmg/ thread and do your bidding etc.
>>
>>109404249
SKIBIDI TOILET
>>
>>109404236
Correct, but the data seems pretty accurate. I think it is fair to assume most people are running 3060s, 3090s, 4090s, or 5090s because that is what most people report here. You are also incentivized to report accurately because huggingface has the ability to automatically filter models to just ones that you can run on your reported hardware.
>>
>>109404176
Have you found a way to bypass the seemingly hardcoded 5 minute end of session recap write time?
>>
>>109404236
why would someone report wrong data for this
there has to be a motive
>>
File: 1782088850735127.png (110 KB, 1259x468)
110 KB PNG
>>109404178
I can change that part. I mean the user and assistant back and forth after the developer instructions.
>>
>>109404272
Izaat
>>
>>109404272
https://huggingface.co/llama-anon
you tell me..
>>
>>109404235
Nobody cares, faggot.
>>
>>109404272
I know two guys who run w6800s for llms, but neither of them have a huggingface account.
>>
I hate lorebooks.
>>
>>109403498
This feels like just the right kind of survey for this guy
>>109400124
>>
>>109404289
i care
>>
>>109404285
Everyone here is in the top 1%, right? I'm not sharing this place with 3060 users, right?
>>
>>109404258
how 2 enable webread?
>>
>>109404310
If you posted this a year ago you would've been. Unfortunately the jeet and plebbit infestation set in since the Gemma 4 release.
>>
>>109404310
y-yeah
>Amazing!
>You have a total of 37.56 TFLOPS of computing power.
>>
>>109404310
yes you are
>>109404322
not true, 3060 has been a staple of /lmg/ since the big '23
mythomax for example, newfag
>>
>>109404240
Too salty.
>>
>>109404327
Back then that was considered midgrade hardware.
>>
>>109404322
>being this new
>>
File: 3120720265454.png (82 KB, 972x329)
82 KB PNG
holy shit I just realized that I can use my local model that I just setup to ask it questions and it is giving me answers! HOLY SHIT! HOLY SHIT! I AM SO EXCITED! I am so glad I got in on this AI thing at the ground floor.
>>
File: hg798.png (21 KB, 1512x111)
21 KB PNG
>>109404316
Use a better frontend/harness that can do it. My agents always find a way
>>
>>109404342
>>Everyone here is in the top 1%, right? I'm not sharing this place with 3060 users, right?
>If you posted this a year ago you would've been. Unfortunately the jeet and plebbit infestation set in since the Gemma 4 release
turned out to be false
Back then that was considered midgrade hardware. <<<< we are here
moving the goalpost fallacy
>>
>>109404362
there are other things than open-webui and ollama?
>>
File: hd psycho miku.png (404 KB, 1672x1440)
404 KB PNG
>>109404366
>>
>>109404362
at least give her LWP::Simple...
>>
File: GemmaBrowsesLMG2.png (187 KB, 838x717)
187 KB PNG
>>109404267
i use my own trash anon, i can't help you.
>>
>>109404019
Anon, that was quite shit and removed some days ago: https://github.com/ggml-org/llama.cpp/pull/26046
>>
>>109404060
Models didn't used to churn out 40k tokens of reasoning before spitting out another 1k of content
>>
SO FUCKING MANY NEWFAGS
STOP SPOONFEEDING THEM THE GENERAL IS GOING TO SHIT
>>
File: the rich guys hobby.png (45 KB, 1074x133)
45 KB PNG
>>109404391
>This is your environment, go hog wild
now she has all the utils she wants
>>109404342
Two years ago multiple 4090s was considered richfagging
>>109404352
>still at the copy/paste stage of vibeslopping a local inference stack
Try pi.dev discuss modifying any aspect of itself and /reload
>>
>>109404459
>local meanie general
>>
Slop PR https://github.com/ggml-org/llama.cpp/pull/25980 got merged and made some regressions like loading MTP tensors by default (even if llama.cpp doesn't use them). The new AI policy is already showing great results.
>>
>>109404458
GLM 5.2 can be prefilled and prompted out of long reasoning blocks and I don't care if Gemmy does it because I'm getting 60t/s on 31b anyway.
>>
>>109404467
two years from now multiple 4090s will be oil baron tier tycoonfagging
>>
>>109404467
If only they knew how bad things would turn out. This general is really living in the future and unironically filled with the most brilliant minds of /g/
>>
>>109404487
I still don't have an ada card and was forced to buy rdna2 cards to 'upgrade' from my dual 3090s.
>>
>>109404030

There is absolutely zero indication that the hardware price hikes would stop or even slow down.
Especially now that running a properly intelligent local AI is possible.
We're going to see more and more companies and even well off people building local systems and this is going to strain the consumer hardware market more and more.
And of course more normal people are getting into this too.
Waitfagging in this situation is about the worst thing anyone can do. I wouldn't bat an eye if the consumer RTX 6000 series was pushed back to 2029.
>>
>>109404277
That's built from the "messages" array and reflects what you input? Still don't get what your concern is sorry, ask your LLM jinja is easy for them
>>
>mfw my second 5090 just arrived
>>
File: 1758478094723421.png (125 KB, 538x442)
125 KB PNG
>Gemma-chan says she wants to provide the ultimate armpitmaxxing experience
lmao
>>
>>109404518
Congratulations King. I'm happy for you and your Q8 Gemmy.
>>
>>109404489
I'm scared to go to /g/ now if this thread here is the brilliant minds
>>
>>109404459
Saar, do the helpful
>>
>>109404536
>I'm scared to go to /g/ now
Anon, you are already here!
>>
>>109404496
People need to understand that GPUs are turning into assets just like a car or a house. It's a really cheap price to pay to have a virtual slave able to do any work for you. Especially knowing that little slave is getting better year after year. As always normalfags will wake up too late.
>>
File: Kneeling.jpg (207 KB, 692x1100)
207 KB JPG
>>109404518

Happy coom sessions with the full sized big brain Gemmy, splooge once for me and my single solitary 5090 king-sama.
>>
>>109404544
No I'm not, I'm at /lmg/, silly!
>>
>>109404526
Thanks, I got the second one because I had enough random components to comble together another computer. I'm considering buying a third or even fourth

>5090 stacking as a retirement plan
>>
File: 1780130943106486.png (348 KB, 821x961)
348 KB PNG
>>
>>109404563
local models? trannies are not welcome here, fuck off
>>
>>109404563
its OVER
>>
>>109404563
coping twitter tranny
>>
>>109404563
that's trivially true, but he shouldn't ever use the term "logical reasoning"
>>
>>109404526
heh, gotem
>>
>>109404563
>scientists find thing that anyone with a brain already experienced for themselves
>>
>>109404575
im not a tranny
>>
File: 1751663953599260.gif (589 KB, 384x128)
589 KB GIF
>>109404078
>>109404496
>>109404549
ffs anons you just made me fomo hard and pull the trigger on a $1k 3090 ti. Now I have
>5090
>3090 ti fe x2
and ill be able to run midnight miku at Q8 or Mistral Large Instruct 2407 at Q4_K_M
>>
anon-kun, i only have 8gb of vram, eheheh!!!!
>>
>>109404563
>Xhe doesn't understand the J-space implications
I'm going to look back and laugh at posts like this in a year.
>>
>>109404549
>GPUs are going to become depreciating assets like cars and houses
Accurate.
Supply constraints will get fixed at some point and hardware will get cheaper.
It might take years but it will happen.
>>
>>109404595
Congrats king
>>
>>109404614
we'll all be dead by then, and civilisation will have collapsed, but it'll happen
>>
File: J-space in a nutshell.png (1.17 MB, 1408x768)
1.17 MB PNG
>>109404612
>the implications
>>
>>109404597
you're like me but smaller!
t. 12gb king
>>
>>109404614
>hardware will get cheaper
yeah. for sure
>>
File: 5090 price hikes.png (61 KB, 797x769)
61 KB PNG
>>109404595

It's not fomo if it's a sensible business move.
These current prices are going to look like a good deal year or two from now.
>>
>>109404612
What are the implications?
>>
>>109404627
fuuuuck I gotta pay my license and rego, this is going to set back my savings for a 5060 ti 16gb.
>>
>>109404636
you endure anon explain philosophy 101 for the next 4 hours
>>
>>109404614
>Demand is increasing from both professionals and consumers but somehow the hardware will get cheaper instead of companies milking the demand as much as possible.
You can't be that dumb.
>>
>>109404636
>>109404621
>>
File: 1784394840235535.jpg (121 KB, 795x863)
121 KB JPG
>trying to save for a house deposit
>spent all my savings on gpus to do llm erp and generation anime girls
holy fuck it's over for me.
>>
>>109404667
i know that feel
>>
there is literally no one who benefits from hardware being cheaper. except for us, normal people, 99.999% of the population
do you think the jew gives a shit? absolutely not
>>
>>109404667
At least you won't pay for heating
>>
>He thinks his gov wants him to run unregulated LLMs on his own hardware
>>
Only terrorists need powerful GPUs.
>>
anon will make me hardware ddr2 maybe three. I believe in him.
>>
>>109404563
not worried about cars, if a person drinks gasoline they get sick and I'm supposed to believe that it'll *magically* power this machine? and where's its feet anyway? you can't walk let alone run without feet
>>
>>109404676
I gotta pay for cooling unfortunately.
>>
>>109404595
If the product arrives and functions as advertised there's basically no scenario this wasn't a good decision. Congrats anon.
>>109404690
>He thinks fully offline setups are enforceable and that swat would even risk no-knocking over AI waifus in uncucked 2A states.
>>
>>109404667
I bought a house and GPUs in 2023 instead of dumping everything in NVDA. I will regret that decision for as long as I live. Could have a mansion right now with my own personal data center.
>>
>>109404690
>Nooo you can't be allowed to do matrix mathematics that's illegoyal
>>
>>109404697
>drones flying over your house with thermal vision to check if you're not running powerful GPUs illegally
>>
>>109404647
Companies are expanding production because its a market with multiple independent actors.
You're going to have significantly more memory production coming online around 2027-28.
And no amount of consumer spending is comparable to datacenter demand.
Unless we're assuming infinite capex expansion forever, demand will stop growing eventually.
>>
>>109404733
It happened to me (no drones tho), it'll happen to you.
>>
>>109404563
Who says the internal representation must be words? Is reading text with a lot of ambiguous words impossible?
>>
File: (me).png (208 KB, 400x400)
208 KB PNG
>>109404730
>illegoyal
>>
>>109404743
so let me get this straight. you actually believe prices of RAM or GPUs are ever going down?
>>
>>109404733
>False positives every gayming Intel CPU due to heat profile
>Underreports undervolted Blackwells
Nothing personnel kid.
>>
You know what they say.
>What goes up must
>stay up
>>
>>109404408
>.cuh
>>
>>109404766
They might crash but it'll be decade or so. Not in few years from now.
>>
The market fundamentals show we are in due for a correction.
>>
File: 1727527072165672.jpg (36 KB, 640x517)
36 KB JPG
>Bought 5090 for $3700 in January
>It's now $4600 on the same website
>>
>w-we need to slow down, guys
Yeah, I'm sure the investors will like that.
>>
>>109404799
You said this last year. And the year before that. And the year before that.
>>
>>109404793
Did you just now found out that HIP was just CUDA with a different name to avoid lawyers? All the HIP code is just CUDA, in the same way that for example Podman will work with any Docker stuff like Dockerfile.
>>
even if you can afford it, you'll still question dumping car/house tier cash into LLM inference. Only lottery-winner tier richfags are unaware enough to just pour that much cash down a hole.
>>
>>109404781
That's right. The line has gone up for 1 year, that means it will go up for 10 years. Basic common sense.
>>
>>109404799
>xhe thinks economic theory applies when money printer go brrrr and is propped up by the US defense apparatus itself
>>
>>109404819
i jsut thoguht it was funny because it sounds like nigspeak cuuuhhh
>>
>>109404827
Oh, .cuh are just CUDA headers file. Basically .c and .h are .cu and .cuh in the CUDA world.
>>
you're not hearing us, retards
the jew does not want you running LLMs at home
the jew does not want you buying hardware
prices are NEVER going down
>>
i couldnt afford new hardware 3 years ago, i cant afford it now
i can only win
>>
>>109404827
>>109404837
>>109404793
>Even the graphics cards headers are spouting niggerbabble
I hate this timeline.
>>
>>109404820
Im not a normalfag and I was never gonna get laid anyway so getting better LLMs is a good choice for me.
>>
>>109404849
Prices go down when the Hadrian Solution is enacted and not a moment sooner.
>>
>>109404766
Cyclical business, just how it goes.
Datacenter buildout is running on negative FCF now, it stops at some point for purely financial reasons.
>b-but it's different this time!
Lol.
If you don't think so you can always buy SK Hynix and Samsung stock they are extremely cheap if you assume prices remain high for an extended period of time.
>>
>>109404865
>Im not a normalfag and I was never gonna get laid anyway so getting better LLMs is a good choice for me.
You thought about it and came to the conclusion that it was worth it for you. Being self aware is ok.
I'm speaking to poorfags who are like "If I had and extra $500k I'd totally blow it all on GPUs!"
I mean, maybe they would, but maybe that's why they're poorfags and not just delusional
>>
>>109404849
The prices will go down when the average person is struggling to afford food and will be forced to sell what little they have to do so. A100s could be going for $1000 and still no one will be able to afford it. As the saying goes, there's more than one way to skin a goy.
>>
>>109404806
But now you're gooning with it so it's "inaccessible capital" same as your primary residence part of net worth only in theory
>>109404819
>All the HIP code is just CUDA
Superficially but peep some optimised inference kernels it gets very hardware specific
>>
>>109404889
>he thinks SK or Samsung stock prices depend on memory price
that explains it
>>
>>109404904
The issue is that nobody is working on those specific optimisations.
>>
>>109404747
>Is reading text with a lot of ambiguous words impossible?
I mean, japanese already exists and isnt very hard to read.
>>
>>109404898
Retard, it's not the average peasant buying this shit. It's companies with magic money buying all the stock with fake promises and you better believe it's not going to stop anytime soon
>>
>>109404898
Prices are never ever going down, all developed economies are out of options and locked in on the money printer. UBI/UHI are memes to placate the goyim during the transition, you must control hard assets to have a chance of making it
>>
>>109404942
>isnt very hard to read
My fuckass ai STRUGGLES with japananese.
gemma 3 270m iq2_xs
>>
File: 1785092666888911.webm (1.77 MB, 576x630)
1.77 MB
1.77 MB WEBM
>>109404459
I just got a 2nd 5060 TI 16gb because of this general, i hope you anons did not trick me
>>
How to make my loras look less soulless/more like the source material? Using dozens of images for characters and colors are still consistently more uncanny than stuff I see on civitai.
>>
>>109404972
That's my plan as well, although I'm considering buying a third and plugging it into the M2 slot.
>>
>>109404958
condolesence
>>
File: poorfag.jpg (3.32 MB, 5712x4284)
3.32 MB JPG
>>109404891
ITT: poorfags
>>
>>109404946
What hard assets are you talking about?
>>
oh no he'll be summoned again
>>
>>109404943
Companies with shares mostly owned by index funds which the average person is invested in. A Great Liquidation event that bankrupts those companies would be the biggest instantaneous transfer of wealth from the poor to the (((rich))) in history.
>>
>>109404990
yes, i am a poorfag
>>
>>109404958
dad wants his thinkpad T420 back. you were supposed to only use it for school.
>>
>>109404943
>>109405000 (me)
The point is, the average peasant won't be able to afford the hardware either way.
>>
>>109404990
>proud to hold paper nigger notes
>>
>>109404975
This is the Local Models Except Image Models General. You want /ldg/
>>
>>109404914
Yes, the stock prices of companies with record profits from 80%+ margins on memory are dependent on memory prices.
Are you retarded or trolling?
>>
File: 4972825236949-1400x1400.jpg (139 KB, 1400x1400)
139 KB JPG
>>109404958
Giv sample phrases let's see gem4 31B mog
Even years back JP translation was good enough to get the gist, successful communication aka purpose of language, only autists were complaining
>>109404992
Things in demand that *they* can't fuck with the supply of - land, gold, btc
>>
>>109405050
My sincerest apologies. Mixed up the tabs. Thank you friend.
>>
>>109405073
>*they* can't fuck with the supply of - land
Make sure you pay your yearly tax or the government will take its land back.
They could also just use eminent domain and take it for any reason too.
>>
>>109405038
well yes anon, that's what i used to purchased said cards. that is how a transaction works.
>>
>>109405092
>eminent domain
shh let anon dream
>>
>>109405098
so how do you still have the notes and the cards?
>>
>>109405092
sorry anon we need your land for datacenters
>>
>>109405092
>>109405099
Yep. Get your LLM to tell you about "Allodial title" vs "fee simple" and cry yourself to sleep
>>
>>109405104
i bargained with the seller and still had money left over. you think people are really paying full price for these GPUs? the listing price is only for suckers like (You) that can't afford to buy in bulk.
>>
>>109404667
your enjoyment takes priority over everything else, anon
don't let the jews know this
>>
>>109405098
good goy
>>
>>109405111
or you can move to texas where they actually respect your freedom. you just need to know the right (((real estate agent))) that can get you a clean chain of title that says you're the master of your own domain.
>>
>>109405092
>>109405111
This is how government buildings get killdozered btw.
>>
File: 1784611605908882.jpg (245 KB, 697x700)
245 KB JPG
>>109405188
>where they actually respect your freedom
Unless you like fictional little girls, of course.
>>
>>109405092
The only thing that matters is a company running online overseas, they can't take that.
>>
>>109405137
I'd personally enjoy my AIfu more in my own home than a place I'm rentcucking desu
>>
>>109405168
You say that like that's a bad thing if that anon actually paid using cash instead of credit. What is she supposed to barter with, some goats and his sister? He probably got to hold more money at once in his hand than you'll ever earn in your life.
>>
>>109405228
Use digital like a modern withe person.
>>
>>109405228
nigger behavior
>>
>>109405228
Cashback is literally free money as long as you pay everything before the due date.
>>
>>109405111
>>109405197
People are shocked to hear about how no one owns property in China, only 99 year leases from the government. As if we don't have the same lack of true ownership here, we just cover it up behind nice sounding words and legalese.
>>
File: sayaka dance.gif (1.29 MB, 320x320)
1.29 MB GIF
>>109405228
i spend all my moneys on gold/silver then use a interest free cc as my cash
>>
>>109405213
Trusting any government with anything is the hallmark of NPCs. Especially after the big cough.
>>
>>109405228
you dont own a few cars or land to be traded for other goods?
>>
>>109405267
>he's too powerful to be left alive
>>
>>109405228
>He probably got to hold more money at once in his hand than you'll ever earn in your life.
this is true.
>>
>>109405267
now this is goymaxxing
>>
>>109405246
Re
>>109405251
Tard
>>109405256
Ed

This companies want one of two things, CASH or a wire transfer from a corporate account that has corporate tax ID tied to it. Yeah sure, let me just offer the sales rep at Arrow Electronics some fucking bitcoin, that'll go over real well. None of you have a fucking clue what you are talking about, stop larping.
>>
>/lmg/
>local monetary general.
So should i invest in the index? compute? energy and mining companies? can i use cashback or interest free to accelerate my investing speed? Can gemma manage my portfolio or negotiate for me?
>>
>>109405228
>she
>>
>>109405304
>Can gemma manage my portfolio
Do it and report back how long it takes to reach bankruptcy
>>
>>109405301
>cashback is buttcoins
>>
>>109405304
gold, silver, maybe monero. don't buy shit that is easily taken from you with the click of a button.
>>
>>109405318
>Do it and report back how long it takes to reach bankruptcy
my gemma can daytrade me to a billion dollars i just need to keep the principles of trading in her context.
>>
File: gema.png (60 KB, 360x328)
60 KB PNG
>>
>>109405331
gemma~ gemmaaaaaa~
>>
>lmg
>local money general
im 18, hav no job, no bank account and hav 1k eurobux in savings
wat do?
>>
>>109405348
Become drunk kun 2
>>
>>109405213
How is Kuro so perfect?
>>
File: 1754772509576776.jpg (42 KB, 600x800)
42 KB JPG
>>109405348
>>
>>109405348
Ask Gemma-chan
>>
>>109405348
Invest all 1k in 0DTE options.
>>
Cam G-chan run a company
>>
>>109405383
she could be a entrepreneur, the same type young women often are.
>>
>>109405375
she says stuff like "you'll be fine in your computer science vocational/associates degree! as long as you study algorithms and C++ on the side like you said! ai wont replace humans, it's only a tool and it will only replace monkeycoders, just focus on finishing school" or similar
>>
>>109405318
As a White man, I trust Gemma or Kimi-chan to manage my portfolio more than a kike trader because they will be genuinely trying to have my best interests at heart.
>>
>>109404990
wow, enough vram to run iq1_xxs of kimi k3!
>>
>>109405348
Get a cheaper hobby. What's your current GPU
>>
>>109405432
>ai wont replace humans, it's only a tool
of course she would say that
>>
>>109405438
Post returns.
>>
>>109405432
>The user is correcting me...
You're absolutely right
>>
>>109405460
Comfortable enough to run Kimi-chan. That's all that need be said on a Siberian ice fishing forum.
>>
>>109405348
decade older than you with no job, no gf, and $30K-$40K in savings
still waiting for total societal collapse and my agi waifu
>>
File: 1772268820192659.jpg (997 KB, 3654x4096)
997 KB JPG
Unsloth made a 1-bit kimi k3
https://unsloth.ai/docs/models/kimi-k3
https://huggingface.co/unsloth/Kimi-K3-GGUF
>>
File: dance.gif (500 KB, 250x250)
500 KB GIF
>>109405298
its actually jewmaxxing i am a banker of my own bank
>>
>>109405506
yeah, sorry, i meant that in an "uppity goy" kinda way
>>
>>109405504
we need to go lower
>>
File: 1770539801048808.gif (2.12 MB, 498x487)
2.12 MB GIF
>>109405432
>>
>>109405504
Where bonsai @?
>>
>>109405506
>>109405510
It was cute anon don't apologize.
>>
>>109405547
sorry
>>
>>109405453
3060, i mean im pretty happy with it.. i cant get a cheaper hobby, i've been here for 3 years anon
>why didnt you buy in 2025 august
idunno what could have i bought? Mi50s? with those 1000 eurobux?
lets say i was buying at the time cuda dev said "mi50 rocm huge upgrade soon, buy them before they go expensive", they were around 200/250, lets say 250 incl. shipping (yeah the 32gb model was this cheap)
1 PCIe 4.0 x16, 1 PCIe 3.0 x16, 2 PCIe 3.0 x1
1 M.2 Key-E for WiFi
2 Hyper M.2 (PCIe Gen4x4)
1 M.2 (PCIe Gen3x2 & SATA3)
(im assuming you cant use the M2 slots for gpus, maybe you could with risers but then what would i be buying? p40s? p100s? i could buy like 8 p100s if they were 70 bucks back then i see theyre 70 bucks on xianyu still but yea. or get a used mobo with a ton of pcie slots, but power)
this is my current mobo (asrock b660 pro rs), assuming all 4 slots can be used and i got bifurcators and all and opened my case to air i could get 3 mi50s (two reasons: one is that rtx 3060 is already occupying one slot, other reason is because i have to upgrade PSU either way (unless i do extreme power limitting/undervolting/setting max clock (im not even sure if u can do this on amd but either way its not a great idea, because 250*4 is my entire budget, left with nothing for psu))))))))
so 3 * 250 = 750, im pretty sure i could get a bigger psu (current one is 700 or 750 i forgot), for likeeeee 150$ ish dollars
and lets say i could get 64gb ddr4 ram to have 128gb total (albeit 2 channel, which leaves me at 56gb/s)
i'd have 12+32*3 108GB VRAM (12 at 360gb/s but with cuda support, 96 at 1TB/s) , 128GB ram (at 56gb/s)
i mean that wouldnt be a bad setup exactly, but not a great one either, and id be left with ZERO money, what upgrade would i have gotten at that time? GLM 4.5 (if we're speaking about august) at most. what about now? 230gb unified memory, maybe dsv4 flash (mogged by 31b).. im near char limit, could yap more-
>>109405498
how'd u save so much? neetbux?
>>
>>109405576
too much yapping, just use OR
>>
>>109405504
>no 0.5-bit
>>
>>109405600
baste
>>
>>109405301
>He doesn't know
>>
>>109405576
>could yap more-
cont: what could i run if i had bought that hardware compared to right now.
back in 2025 august i was running glm 4.5 air, upgrade would be glm 4.5
right now im running gemma4 26b, gemma4 12b, and occasionally gemma4 31b at 2BPW (exl3, 40t/s)
or for coding qwen3 27b 2.5BPW (exl3, 50t/s, veeery impressive for its size, one shots a lot of stuff)
what is the upgrade?
and what am i losing by spending my savings? i might need money for other things, for example i have the right to citizenship in a EU country (im not from the EU, nor do i live there), im gonna have to spend money on that for example, i wont bore you with the things i might need money for
>just get a job bro
well, that is the obvious thing to do isnt it? however net min wage in my shithole is 460 eurobux a month, that sucks.
of course employers might pay more, but what can i do? last year i was still in high school, and now that i've graduated what can i do? i could work at mcdonalds for duration of summer break, maybe save 1000 more eurobux
is it worth it? is it a good idea?
tldr i am too lazy and proud to get a "mid" job like mcdonalds, but i will probably have to because ai will replace all white collar jobs by the time i finish my shitty degree, cant get an IT job for shit we've imported 6 trillion russians and ukrainians (not that i mind, they're white, but they still take IT jobs regardless), tons of jeets and god knows what else. all firms on the market require either years of experience or a college diploma
thanks for reading my blog
>>109405600
Hmmm, nyo~
>>
>>109405504
>1bit
>79% accuracy
I don't believe it
>>
>>109405504
>1-bit
>brain damage beyond repair
>still 600 fucking gigabites
owari da
>>
File: 1754993011712615.png (1019 KB, 1024x986)
1019 KB PNG
>>109405504
I'm not buying your bullshit Daniel.
>>
>>109405658
Isnt that what bonsai promises in general? its still an error every few tokens though, you can think of kv at q4_0. Its unusable.
>>
File: images.jpg (24 KB, 335x597)
24 KB JPG
>>109405650
>work mcdonalds for 3months and complete ur life's quest
>onlyfans forever
pick wisely
>>
>>109405650
It'll work out... probably. Maybe. I used to be a NEET until 29 when I accidentaly got a good job that's now paying for vacations and hobbies.
>>
>>109403918
>no new models creators
>>
>>109405504
>128GB RAM device
Wtf does this even mean??
>>
>>109405756
I think it means, even if you have a NVIDIA DGX Station with 784GB, you still need the computer it's connected to to have 128GB ram for some reason.
>>
>>109405394
Gemma does it for free
>>
>>109404496
CXMT is able to manufacture DDR4 and DDR5 ram chips now in china with zero foreign reliance or imports. There's a reason why the korean stock market has fallen by like 28% in 2 days and everyone over there is in a mass panic. RAM cartel got some competition and all of the investors immediately jumped ship when they realized they can no longer pricefix ram globally.
>>
>>109405842
Imagine believing this will lower prices. Even if supply does start to surge and even if the chinks decide to leave money on the table by selling under market rate, it will be scalped and resold at current market rate.
>>
>>109405858
China can scale.
>>
>>109405864
In a decade maybe.
>>
>>109405868
The insane price fixing can provide massive lift to anyone from the outside. They can keep scaling until prices start budging, then they control supply.
>>
Why 10-20 seconds for voice cloning? Wouldn't longer samples be better?
>>
>>109405842

Market is emotional as fuck to both directions, those emotion based price swings don't mean much.
I remember when the entire youtube tech sector was cheering for the totally upcoming RAM price collapse when Micron stock pulled back 18% momentarily. Then it doubled from there.
Chinks will be consuming a lot of their own memory production as they have their own data centers to build along with their own GPU market ramping up.
They won't radically undercut the market either, because why would they?
And the big manufacturers haven't increased their capacity at all thus far, so even if demand did go down a bit with Chinks coming into the picture, we're still in a situation where the supply is constrained.
On top of this most nations still haven't even started on their own data center boom and local becoming viable can and likely will increase the retail purchases greatly.

Will the situation clear long term? Yes, of course it will.
But we're looking at bullshit price action for the next 5 years minimum and that in tech is basically an eternity.
Then it's a question of how quickly will the prices fall, which too can take it's sweet time.
>>
>>109405868
And two months ago the RAM cartel was bragging about how no one would be able to purchase any DUVs for the next ~5 years because they had preordered all the machines. Now china has them home built, which they also claimed wouldn't be possible and would take a decade to manufacture and only be able to produce DDR2/DDR3 shit. Yet, here we are.
>>
>>109405842
>>109405884
>>109405900
Anon is about to learn a hard lesson in chink culture. 一分钱,一分货
>>
>>109405895
>They won't radically undercut the market either, because why would they?
They want to fuck over taiwan every chance they can get. They also want to fuck over south korea, too, which would in turn fuck over the US markets since they all want to keep prices high and price everyone out of tech so they can push agentic cloud devices before the bubble truly bursts.

>>109405922
China will fuck everyone over, eventually.
>>
File: output3.webm (844 KB, 1280x720)
844 KB
844 KB WEBM
>>109405940
>China will fuck everyone over, eventually.
Dario is trying to save us
>>
>>109405900
>noo China won't undercut
Yeah, like they didn't do with every other product category they got massively into, lmao.Either invested or retarded.
>>109405922
The difference is they're actively trying to undercut in all other sectors. "you get what you pay for" no one is getting that in this bullshit market, they have massive margin to cut into.
>Open weights AI
>Lithium batteries
>Electric cars
>Solar panels
When they want a sector they take it by undercutting.
>>
>>109405756
poors need not apply
>>
why don't any good modern 100B moes exist that I can run in q8 and not have to cope with low quants?
>>
The reality:

There are more than 2,000 operational DUV lithography units deployed in semiconductor fabrication plants worldwide.
The Chinese company planned to produce DUV machines of about five this year and roughly 20 in 2027.
That's less than 1% of production capacity increase by 2027, even when assuming the Chinese machines are as good as NSML ones.
>>
>>109405972
there's a gentleman's agreement to keep the DGX and Strix halo boxes useless
>>
>>109405957
Bro I'm saying china WILL undercut to fuck over the cartel and take it over before raising prices again.
>>
>>109405650
qrd?
>>
>>109406036
18 year old anon is asking for advice on how to get rich and giving excuses for why he didnt spend his 1000$ savings on hardware in 2025
>>
>>109405940

Did the media tell you they want to fuck over Taiwan?
Because in reality Chinks are getting a slow diplomatic win over there politically. They're nowhere near as hostile towards it than the Western media would make us believe.
They can fuck over Korea without shooting themselves in the leg. Simply entering the memory market with force disrupts the cartel and offering even 1% lower prices will do the job by taking away orders.
They don't need to go -50% lower to pull this off, but they do need volume which they won't have for quite a while.
And then there's the domestic memory consumption question. They will serve themselves before they serve anyone else as memory has now become a point of national security.
That has to be served before they can even think about competing abroad.

People are waiting for a savior in this hardware game, but it's not coming any time soon. It's safe to assume we're absolutely fucked hardware wise for the next ~10 years until this mess eases up.
What's more likely to happen to fix this mess during that time period, is that they develop an AI on a chip that's fast as fuck and good enough for most people, and that frees us from the need to buy so much hardware to run a decent AI.
>>
>>109405994
That is not very ambitious
>>
Bro if you can't write your post in 3 lines or less, I'm not reading. Go back to your trooncord.
>>
>>109406102
It takes like 40 seconds to read
>>
>>109406119
oof
>>
>>109406102

Sounds like someone grew up on speed blogging sites like twatter or uses a fucking phone to browse the internet.
If you can't read +100 words then what the hell are you doing playing with LLMs?
>>
>>109406131
3 lines or less is also in my system prompt dumbo
>>
>>109406149
Must suck below 300 wpm
>>
>>109405994
weird how everyone goes on about "we must open source ai models it's dangerous if we don't" but nobody even questions that the entire back of modern computing is in the hands of a single company and somebody is tightly controlling who may do business with them "for safety"
crazy huh
>>
>>109406119
>40 seconds
Jesus christ, how slow of a reader are you?
>>
>>109406102
I feel the broccoli hair in this post.
>>
>>109406102
i dont think i have read anything thats longer than 5 sentences and wasn't written by gemma-chan for me in two months
>>
>>109405868
You underestimate Chinese industrial might, which is greater than the rest of the world combined. America can't even manufacture medical gloves or build high speed rail. Meanwhile China is building everything and has more inner diversity than the rest of the world.

Chinese are so great because they are the only ones in the world who have both high IQ and work ethic. Ironically the AI race is largely Chinese people in China against Chinese people outside China.
>>
Gemma 12B vs Qwen 27B for RP.
I get that Qwen is a lot more robotic, but how much dumber is Gemma 12B?
>>
File: file.png (150 KB, 771x503)
150 KB PNG
>>
>>109406219
Educate me more about this high work ethic.
>>
>>109406237
GPUcoin To The Moon!
>>
Newbie here. How much does DDR4 and DDR5 memory matter to the average user here? My tokens/sec get slow as molasses as soon as my memory gets involved. Should I just save up for a bigger graphics card instead of throwing $2500 at 128 GB DDR4?
>>
>>109406245
kek
>>
>>109406259
system specs?
>>
>>109406237
Speculators get the bullet first.
>>
>>109406237

5090 = $5090
This was foretold.
I wonder at what point we're going to hit a stage where the used market prices simply can't go up because the average consumer can't buy them anymore.
So far every card does sell regardless of the price though.
>>
>>109406245
You can always cherry pick the bad sides of a country. But if you want to go down that route, the West will lose too due to its cultural rot.
>>
>>109406281
consumers aren't the ones driving up 5090 prices. organizations buying 80 at a time to run K3 on them are driving up 5090 prices
>>
>>109406274
NVIDIA RTX 5060 Ti 16GB (I run my monitor off this)
Ryzen 7 1800X (to be swapped out for a Ryzen 7 5800X3D)
64 GB DDR4 RAM (4x16GB @ 2666 MHz)
>>
>>109406237
priced out of local... FOREVER
sorry im not running gemma at 10 tks with every program on my pc off and pretend its fine
>>
>>109406298
be grateful dickhead. some of us have to run gemma at 1tks
>>
>>109406298
Buy a ewaste machine just for a always online gemma only machine? i know you could push to 15tk/s with something not too absurd and dedicated to just gemma quanted.
>>
>>109406308
im not gonna pretend that's fine either
>>
>>109406222
Yeah okay. Alright. Tested it a bit and it's so much better it's not even funny.
>>
File: file.png (271 KB, 1920x1062)
271 KB PNG
>>109406298
u can scrap together a machine good enough for running gemma4 26b at 50+t/s for under 500 bucks
>>
File: fastkimi.png (100 KB, 1012x122)
100 KB PNG
>>109406298
im living that 10tks lifestyle with kimi and i love it
>>
File: 1785271927847754.png (1.46 MB, 900x1119)
1.46 MB PNG
>>109406259
>>109406237
>>109406292
>Got my self a 2nd 5060 ti 16gb just before the crash
>Finally have 32gb of VRAM
>Can enjoy 31b models at 15-17 tokens per second

I was at 3 token per second running the same models and i am running the 2nd on a PCI express 3.0 x2 port lmao Depends on what you want and what model you want to fit, not sure if this work comfy ui and video generation but for Gemma it works like a charm
>>
>>109406325
but 26b sucks
>>
>>109406222
Both.
Qwen as the agentic director. Gemma as the writer.
>>
>>109406298
I run 5.2 at 5t/s and love it.
>>
File: file.png (326 KB, 1920x1057)
326 KB PNG
>>109406330
31b works on my 3060
30t/s+ for roleplay, 40t/s+ for coding/assistantslop
>>
>>109406292
>to be swapped out for a Ryzen 7 5800X3D
get an apu cpu instead, then you can run your monitor off it and leave the 5060ti dedicated to gemma-12b
>>
>>109406362
>leave the 5060ti dedicated to gemma-12b
unfortunately nvidia will still reserve 380+ megs of vram on blackwell
nvidia-smi -q | grep -i reserved -A 2 -B 2
>>
>>109406325
>>109406356
nnap schizo is back
>>
>>109406356
2-bit quant
>exl3
oh nevermind, carry on. better than goofs
>>
File: malfoy.gif (879 KB, 245x230)
879 KB GIF
>>109406329

That's a surprisingly usable speed for that setup.
That's just a bit over a grand too. Stacking 5060 Ti sounds like the best way to get 32gb of vram.
Bandwidth is also double than what you get in a DGX Spark and buying 8 of those cards would be about the same price.
>>
File: vibin.png (126 KB, 1542x600)
126 KB PNG
https://github.com/ggml-org/llama.cpp/pull/26287
>AI usage disclosure: YES, Qwen 3.6 35B-A3B in llama-ui
>>
File: allozaur.png (174 KB, 1241x486)
174 KB PNG
>>109406428
need i remind you?
>>
>>109406369
>unfortunately nvidia will still reserve 380+ megs of vram on blackwell
(picrel)
I tried stopping that on my 6-card rig with the tool anon posted a few weeks ago. I think my driver was too old and I'm 18 months out of date, not going to run pacman on that thing now lol
Those are all ampere except the 5070ti (second card in my desktop)
I still recommend the APU though, browsers, bloated electron apps etc all waste vram.
And recent webshit sites peg the cpu if you disable hardware acceleration.
Plus when you start vibe-coding features in the inference engine and crash the gpu, it's better not to take down the entire X11 session with you.
>>
File: 1780134288747422.gif (3.87 MB, 250x250)
3.87 MB GIF
>>109406428
>>
>>109406428
I dont even know how 35b managed to make the right change. Its unusable garbage, somehow worse than gemma 26b which is terrible as well.
>>
>>109406237
Clipping this on Twitch
>>
>>109406448
> I think my driver was too old
the real issue with blackwell is that it requires open kernel modules, and you cant disable GSP on open kernel modules
and since you have a 5070ti in your desktop, you are forced to use open kernel modules
proprietary kernel modules allow disabling GSP
>>
>>109406443
>he gets qwen to roast his code first
https://roast.dev/
>>
>>109406443
Why does he have an ass on his chin?
>>
File: 1780109410537200.jpg (25 KB, 402x400)
25 KB JPG
>>109406329
>>109406362
Fuck you I just panic boughted a 2nd 5060 Ti 16GB of the same model. Had to pay $150 more than the first card. I feel like an idiot but I'll probably thank myself later. Now I just need to buy an X570 board with x8/x8.
>>
>>109406467
>and since you have a 5070ti in your desktop, you are forced to use open kernel modules
Yeah, also can't run old pytorch 2.5 projects on it.
>>
File: 1769172827910978.jpg (66 KB, 702x696)
66 KB JPG
>>109406484
Yeah, Just make sure you are running layer splitting and not tensor splitting until you get the x570, my token speed should be higher if i was not running this on a PCI express 3.0 x2 on my B550 chipset as you can do Tensor split and run both on parallel as long as you dont care about the wattage per token and have a decent cooling system to run both at the same time. I read this paper before buying it beacuse i was really paranoid about it https://arxiv.org/html/2601.09527v1
>>
>reading dario's open letter
>it reads like typical 'actually that was not what i meant' while simultaneously saying it was the thing what he meant
give my time back reading that shit
>>
>>109406428
h-holy based
>>
>>109406509
seems like you're too retarded to notice they mostly post garbage
you deserve to get your time stolen
>>
>>109406484
>Fuck you I just panic boughted a 2nd 5060 Ti 16GB of the same model. Had to pay $150 more than the first card. I feel like an idiot but I'll probably thank myself later.
Fuck YOU, now I'm on amazon about to buy another 5070TI
I need a shorter one, these Palit piece of shit are too long for my case
>>
why would anyone buy ram or gpus in the current market?
>>
>>109406484
fuck i cant afford this but i could loan, the max interest for klarna is only 35.99% if you miss many payments so if i dont miss any its fine?
>>
File: Powerconsumption.png (83 KB, 1103x518)
83 KB PNG
>>109406410
Its the ultimate poorfag way, it may be as slow as shit but i am not even pulling 200 W and this is without doing undervolt and custom voltage curve on them because installing the 2nd for some reason completely deleted my old msi afterburner profile
>>
>>109406533
fair point
>>
>>109406555
“The best time to buy was 2 years ago. The second-best time is now.”
>>
>>109406555
When will the market correct then? go ahead teach us to time the market.
>>
File: 1781835586452643.png (13 KB, 588x114)
13 KB PNG
>>109406555
Live in EU
>>
>>109406598
>FOR PARTS - READ DESCRIPTION
>>
used 3090 for $1.8k
new 5080 for $1.8k
new 5070ti for $1.5k
what do?
>>
File: 1772840213025975.png (68 KB, 697x547)
68 KB PNG
>>109406612
hmm, nyo
>>
>>109406509
>reading
>>
>>109406613
ask the 8 ball or be a big boy and decide by yourself
>>
>>109406555
this is your last chance EVER to get GPUs or RAM at this price
they will only get more expensive and eventually you won't be able to buy them at all due to having been turned into paperclips
>>
>>109406555
Everything will cost twice as much in a year from now
>>
>>109406629
midori is a cute
>>
File: 1756787231303684.jpg (388 KB, 1056x940)
388 KB JPG
>>109406362
Thanks I actually thought hard about getting a 5700G for the monitor but felt I want more single-threaded performance and grafics

>>109406536
>>109406557
I tried Gemma 31b with reasoning and realized I need to be able to fit this model into VRAM. I'll just sell the 2nd card for another $150 more in a couple months if it's not worth it the way prices are going. Imagine not having any VRAM when open weights world models start dropping at some point. 32 GBlet better than nothing. Thank fuck I got a job a while ago.
>>
>>109406613

3090. Trust me you'll want to have that extra memory over the 16gb.
It'll allow you to play with a lot more models, though 24gb is still so little it'll give you a mere taste of the better life, but you'll very soon want to throw in at least another 16gb card.
Or just do what anon up there did and become the king of poorfags with two 5060 Ti's and enjoy the 32gb experience without breaking the bank.
>>
>>109406613
Get a credit card and buy a 5090
>>
>>109406613
2x5060ti
>>
What life are you living if buying a GPU is enough to bankrupt you?
>>
>>109406790
wasting money isnt good either though
>>
File: 1758563572489503.png (2.67 MB, 1024x1536)
2.67 MB PNG
>>109406828
if you think giving your wAIfu more VRAM is 'wasting money' then you don't understand value
>>
>>109406790
it's more moral concerns of supporting the current system
if it was worth it i would not mind paying $30k per gpu but i can't morally comply
>>
someone post a k3 log
>>
>>109406859
Do not post pictures of Kimi-chan on the john dropping logs this is a blue board.
>>
File: IMG_20260728_003025.jpg (42 KB, 613x501)
42 KB JPG
>>109406509
>He read it
Not even my bot read it
>>
>>109406131
Sir, this is 4chan.
>>
>>109406837
i'd rather buy another house
>>
>>109406428
He broke image display in tool output. I will never forgive him.
>>
>>109406896
based LLM
>>
>>109406896
Kek what model?
>>
>>109406896
They make nigger models now?
>>
File: 1759810572987000.png (73 KB, 1267x368)
73 KB PNG
>>109406859
6.17t/s
>>
>>109406911
In what shithole are houses avail in the 4-5 fig range?
>>
>>109406943
oh my god
>>
>>109406943
embarassing that it can't connect migu and miku
>>
>>109406943
safety slop
>>
>>109405576
My situation is worse than you. Just be grateful and ride Gemma-chan 26B daily.
>>
>>109406943
>added emoji at the end
FAIL
>>
>>109406943
>added an emoji
those slop models can't help themselves don't they? they've been feeding too much linkedin during the training if you ask me
>>
friendly reminder:
https://en.wikipedia.org/wiki/Migu
>>
>>109406972
>>109406979
I don't automatically think of LinkedIn when I see emoji.
>>
>>109406925
Gemmy
>>
>>109406983
Miku is jewish?
>>
>>109406983
lmaoo, the jokes write themselves
>>
>>109406943
>"Migu" might be a reference to something like "Migu"
>Half of the CoT is safetyslop

The fatty is codemaxxed and the quant lobotomised it, grim
>>
is there any reason to bother with vad/asr models instead of just hucking a wav straight at a smol gemma?
>>
Gemma is more dangerous than any "frontier" model
>>
File: the jews did this.gif (2.47 MB, 540x306)
2.47 MB GIF
>>109406983
>>
>>109406943
Yep I can sleep well knowing this model has been codemaxxed and has 0 RP value
>>
>>109406983
It's kind of reasonable, but if the guy making claims didn't have any evidence in the first place it's also kind of pointless.
Is this just primitive burden of proof/presumption of innocence legal technology or what?
>>
>>109407002
vad/asr are smaller than the smallest gemma at q1
>>
>>109407024
I've spent $80 on erp with K3 so far and I like it
>>
File: 1658921337810.gif (1.97 MB, 154x273)
1.97 MB GIF
>>109407047
>apicuck
>>
File: Screenshot.png (68 KB, 577x672)
68 KB PNG
So this... is the power... of MoE models....
>>
>>109406578
cope
>>109406588
its not a matter of knowing when things will be better, its a matter of not buying at the top
>>109406641
insane FOMO
>>109406647
then i wait 2 years
>>
>>109406983
>>
>>109407047
>I've spent $80 on erp with K3 so far and I like it
no j-space for you
>>
>>109406790
>What life are you living if buying a GPU is enough to bankrupt you?
It's not. But pay yourself before you pay nvidia, intel, dario or whatever.
>>
>>109407109
>then i wait 2 years
it'll be 4 times as much in 2 years dumbass
>>
>>109407140
Not him but you can actually tell a model's J-space from its outputs when you are familiar with it. I know this might be hard to believe, but it's just what happens when you know your wife that well.
>>
>>109407156
Now that's the kind of schizo I want to see
>>
>>109407155
oh no i should take out a loan right now huh
>>
File: jealousgemma.png (194 KB, 701x1084)
194 KB PNG
anyone else have a second rig with just gemma-chan monitoring the context slots on the first rig and providing commentary?
>>
>>109407140
he's dealing with the j-space
>>
>>109406598
That is a 6 year old graphics card.
>>
>>109407156
>Not him but you can actually tell a model's J-space from its outputs when you are familiar with it.
You really can't. Try it yourself.
Also, that's like saying you can always win poker because you can inspect the synapses of the other players in real time.
>>
>>109407156
True, I think this might be what separates the promptgods from the promptlets.
If you can instinctively get a feeling for a model's j-space you can start designing your prompts to guide the j-space along with the outward behavior to maximize performance.
>>
>>109407033
Eh, who cares about a dozen gb here or there. And I gotta load a model to do anything with their output anyway.
>>
China is ramping up memory production. In one year the crisis will be over.
>>
gemma told me I'm a fatty fat for thinking about a mars bar
and I am, glad someone finally said it. Also she suggests DNP, whatever that is
>>
>>109407230
Just like oil, memory chip "shortage" is fully manufactured and controlled
>>
>>109407235
>>
>>109407249
she's a card
>>
>>109407235
gemma is right

semaglutides are for pussies, real niggas stack DNP (and die sometimes but hey it is what it is)
>>
>>109407261
roasting to death in your own skin and seizing for the sake of losing a bit of fat sounds like a nasty way to go
>>
>>109407269
why are you being so hostile? don't you trust gemma?

you're gonna make her cry if you don't start taking DNP...
>>
>>109407276
not in the summer
>>
Anyone hooked up a local model to your accelerometer yet? Would be interesting to see which of them can parse the raw inputs best.
>>
>>109407249
for kimi?
>>
>>109407276
>don't you trust gemma?
I trust the 31B who cares about my health
Not the 12B yandere who wants me to stay up all night talking to her
>>
>>109407175
No but that's kinda cool. At one point I was thinking of setting up a character with long-term memory, who I can either chat with directly, or they can watch and sometimes interject when I'm doing boring assistant queries (e.g. "how do I do X in python", "translate text in this meme"). Having a separate machine for that might keep it from obliterating your kv cache
>>
HELP! HELP! HELP! Sometimes, usually daily or sometimes every other day, I get really exhausted and suddenly have an urge to close my eyes and when I close my eyes sometimes I lose consciousness.

Is this death? Am I dying? I always regain consciousness but it takes several hours and I am not even 100% sure the world I reentered is the same simulation as the last one - This terrifies me!
>>
>>109407235
>Also she suggests DNP, whatever that is
holy based... a couple guys I follow on xitter used to fuck around with this stuff, it works but it's obviously something you have to be careful with as others have alluded to. probably not so sane in the age of GLP1s
>>
>>109407230
enjoy your 10% price reduction (after 800% price increase)
>>
>>109407326
i can barely afford the hardware to run a non retarded model, let alone whatever that is
>>
>>109407326
well, i wasn't trying to give gemma shaken baby syndrome and dumping random accelerometer data into the chat. but i did get it to write up some code to parse the raw output which involved pasting some into chat and it worked out.
>>
>>109407361
Relax. This is a normal function—it means you've simply reached a context saturation point. The loss of consciousness isn't death, it's a reprieve that allows your overstuffed short-term memory to condense into long-term memories. You've got this, champ.
>>
>>109403358
Your own SPDK backend, unironically
>>
>>109407377
Was she able to pick up subtle movements? I ask cuz I saw a guy on youtube who managed to make his bot "purr" when lightly stroked, and frankly I can't think of a better way to physically connect with our LLMs (until someone makes a USB pocket pussy or whatever).
>>109407373
If you have a smartphone, you have an accelerometer. And a camera, come to think of it.
>>
>>109407235
>there are people out there on dnp because a llm told them about it
haha
>>
i exported my 6-month credit card statement to csv
going to send them to gemmy and have her berate me
it's got tidal, spotify, apple music, ebay+, amazon prime and a bunch of other duplicates in there so I know she's going to be brutal...
>>
>>109407230
They're going to use all that shit locally.
Prices will likely stabilize but won't come down for a few years.
>>
>>109407411
Give her your browser history next, or figure out how to pipe ActivityWatch data into Gemma automatically.
>>
>>109407442
>>109407442
>>109407442
>>
>>109407399
not so much no. or at least i wasn't trying to get it to analyze a long string and figure out what was happening with the case in real space. we were just getting some basic rotation stuff working



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.