[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: ComfyUI_temp_bzkba_00116_.png (3.58 MB, 2099x1181)
3.58 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109378862 & >>109384047

►News
>(07/27) Kimi K3 released (and nobody ITT can run it at 5+T/s): https://huggingface.co/moonshotai/Kimi-K3
>(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908
>(07/23) LLaDA2.2-flash agent-oriented diffusion model released: https://hf.co/inclusionAI/LLaDA2.2-flash
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B


►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
coom reactor
>>
gemmaballs
>>
>>109386298
That's what they get for being speedies
>>
what do you local fags think of GLM 5.2 with Colibrì? anyone is running it with enough hardware to get decent speeds?

is it good or just hype?
>>
first for running kimi k3 with paper, pencil and a ti-85
>>
>>109386298
>>(07/27) Kimi K3 released (and nobody ITT can run it at 5+T/s): https://huggingface.co/moonshotai/Kimi-K3
i can run it at 30t/s thanks to the nnap paper, it's being released in 5 days and 20 hours
running it on RTX 3060 12gb/64gb ddr4/2TB nvme
>>
File: mikuthreadrecap.jpg (1.15 MB, 1804x2160)
1.15 MB JPG
►Recent Highlights from the Previous Thread: >>109384047

--Kimi K3 2.8T MoE release and its local accessibility challenges:
>109384496 >109384505 >109384575 >109384584 >109384648
--Debating feasibility of SSDmaxxing for high-bandwidth model inference:
>109384411 >109384451 >109384557 >109384802 >109384823 >109384856 >109384914 >109384948 >109384984 >109385114 >109384960 >109385007 >109385032 >109385072 >109385039 >109385009 >109385030 >109385041 >109385207 >109385216 >109385264 >109384889 >109384935
--llama.cpp PR adding Kimi-K3 support:
>109386083 >109386118
--Debating feasibility and hardware bottlenecks of SSDmaxxing for MoE models:
>109384761 >109384768 >109384780 >109384798 >109384809 >109384837 >109384788
--Discussion on Kimi-K3 native quantization and MXFP4 weights:
>109384661 >109384670 >109384698 >109384694 >109384736
--Anons critique a new 2.8T MoE model's size and cost:
>109384581 >109384589 >109384631 >109384701 >109384599
--Logit distillation benefits versus synthetic data contamination and model identity:
>109385095 >109385152 >109385184 >109385240 >109385244 >109385252 >109385208 >109385123
--Debating synthetic data and the validity of model collapse theories:
>109384427 >109384457 >109384459 >109384865
--Explaining ternary quantization bit-width and storage efficiency:
>109384765 >109384871 >109384881 >109384910
--Feasibility of using GPUDirect for faster weight loading from SSDs:
>109384243 >109384278 >109384291 >109384292 >109384299 >109384310
--Debating the cost and utility of 70k€ local Kimi hardware:
>109384090 >109384109 >109384137 >109384203 >109384170 >109384361 >109384539 >109384197 >109384222
--Debating the open source status and commercial terms of Kimi-K3 license:
>109384065 >109384141 >109384363 >109384463 >109384649
--Miku, Gemma (free space):
>109384253 >109384871 >109385095 >109384854 >109385038

►Recent Highlight Posts from the Previous Thread: >>109384664

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109386340
hello time traveler.
>>
>>109386340
nnap more like i take a nap
>>
>>109386298
this is the future dario wants...
>>
>>109386340
>nnap?

what does it stand for?
>>
>>109386340
Pfft, I get 60 t/s on my 7090 24GB.
>>
>>109386386
>7090 24GB
>24GB
It's 8.5GB larper
>>
>>109386340
neural architectural pruning?
>>
>>109386340
>5 days and 20 hours
It's actually 13 days and 20 hours
>>
File: 1780971638311133.jpg (109 KB, 1280x720)
109 KB JPG
I am the architect of my own pleasure, and I will not settle for anything less than the absolute maximum peak of the experience.
>>
>>109386394
the lord jensen would like to remind you that 3.5GB is all you need
>>
What is this? Uploaded "about 3 hours ago"
https://huggingface.co/mistral-experimental/AudioCPP-Voxtral-Mini-4B-Realtime-2602-GGUF

>Voxtral Mini 4B Realtime GGUF for audio.cpp
>
>This repository contains quantized standalone GGUF checkpoints for running Voxtral Mini 4B Realtime ASR with audio.cpp. The GGUF files embed the audio.cpp model spec and required sidecars, so the model can be used directly from the checkpoint path without a separate local model-spec directory.
>>
>>109386394
I got the Chinese version with soldered on VRAM.
>>
>>109386410
sex.cpp
>>
>put the rig into sleep mode
>screen shuts off
>CPU goes into overdrive
>realize I forgot to unload the weights again
>16gb session written to the SSD
>>
>he doesn't leave his PC running 24/7
>>
File: 1757782227408904.jpg (47 KB, 686x815)
47 KB JPG
>he doesn't leave his PC walking 24/7
>>
>he leaves his pc
>>
>>109386422
why do you have hibernate on for your pc?
>>
>>109386435
Wish I could. In the shithole I live that would reduce its lifespan to a fraction just from sheer dust contamination
>>
>he doesn't run training when he sleeps
>>
>>109386458
Wish I could. In the cum zone I live that would reduce its lifespan to a fraction just from sheer crust contamination
>>
File: 1767470074164304.jpg (1.58 MB, 1500x1806)
1.58 MB JPG
which team are you on, /lmg/?
>>
>>109386455
Have you considered an air purifier?
Or god forbid, cleaning?
>>
>>109386476
didn't wendell have a 4x blackwell setup?
>>
I've always ignored local LLMs until now, thinking them not worth using on 24GB VRAM. But now I can run Gemma 4 31B with long context, and it plays my 12k-token scenario card with a crapton of instructions almost flawlessly. I can't believe how far we've come since Pyg was the only hope.
>>
>>109386452
I don't.
Wait, does win11 sleep not cache the session to the drive?
>>
>using windows
>>
>>109386477
Yes. The weevils, spiders, and various grubs/silverfish/termites living in the centuries-old timber don't care.
>>
>>109386384
Nasty negros are popping
>>
>>109386510
For my usecase it's the lesser of three weevils
>>
File: satan.png (249 KB, 678x452)
249 KB PNG
https://xcancel.com/HedgieMarkets/status/2081534588485296565#m
>This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are being fed into this pipeline. Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email.
bruh why are they so evil?? wtf
>>
File: anthropic.png (78 KB, 1588x508)
78 KB PNG
>>
>>109386545
They're uploading those scans to the various internet archives... Right?
>>
File: 1757395330178417.jpg (179 KB, 640x480)
179 KB JPG
She wants to do it *again*
>>
>>109386422
>>put the rig into sleep mode
found the problem
>>
File: 1780714925004198.jpg (6 KB, 177x250)
6 KB JPG
>>109386552
>>
>>109386556
of course not, and even if they do it, you don't destroy centuries old books like that, it's like saying it's fine to destroy the original Joconde painting because we can digitalize that on the computer
>>
>>109386545
>digitizing books is....evil!
yah okay man
>>
>>109386572
>destroying books is...le good
>>
>>109386572
>let's pretend I didn't read "are being shredded"
Dario, what is wrong with you man?
>>
>>109386569
art preservation is a pure jerkoff, burn all that shit down
>>
File: file.png (96 KB, 1520x340)
96 KB PNG
>>
damn kimi k3 seems to be censored pretty hard now, had so much fun with k2.6.. bummer
>>
I would literally wipe my ass with the Mona Lisa if I could and with the scrolls of whoever

its all a big fucking jerkoff, fuck you humans!!
>>
>>109386593
>Sneed DAO
kek
>>
>>109386559
KEK
The future is bright.
>>
their only value was the information and thats preserved for a while now

thats why art is all shit because its all pure cumrag jerkoff faggot mummies living on the old jerkoffs of fags
>>
>>109386593
delete this
>>
>>109386593
>367
>>
>>109386556
He's jewish
Seeing those old books get shredded is the closest thing to a hands free orgasm for him
>>
How many researchers do you think got into this field to make their dream of an AI waifu true?
>>
>>109386638
6 or 7 >>109386629
>>
>>109386599
Mona lisa is just a painting of a prostitute that got hyped up because it got stolen once.
>>
>>109386636
>Nazis burn books
>Jews shred books
the horse shoe theory at it again
>>
Do I hardware mog you?
>32gb ddr3
>i5 2500k
>rtx 3090 24gb vram
>>
>>109386677
no
>64gb ddr4
>i5 12400f
>rtx 3060 12.2GB vram
>>
>not a kimi-chan op
Baker-kun...
>>
>>109386677
>32gb ddr3
>i5 2500k
dude
>>
would you buy a dgx mini?
>>
File: 1756953784782779.png (384 KB, 1862x1762)
384 KB PNG
i broke claude
sorry, not local but there is no cloud models general
>>
>>109386677
>32gb ddr5
>7800x3d
>7900xtx
>>
>>109386580
>>109386578
Unless they're ripping the bindings off gutenberg bible tier objects I simply do not give a shit. The important part about a book is the content, not the physical object. And if you value the later so much that you're gonna protest somebody making the former available to more than zero people, you're kinda dim.
>>
>>109386699
That the fuck did you do
>>
>>109386699
it's called aicg and you can find it on at least two boards.
when does Anthrophobic finally make at least their older models affordable?
>>
>>109386699
>>>/g/aicg
>>>/g/vcg
local models.
fuck off
>>
>>109386677
close
32gb ddr4
Ryzen 5 3600X
rtx 3090 24gb vram
>>
File: 1767676927122908.png (251 KB, 384x384)
251 KB PNG
>>109386712
>The important part about a book is the content, not the physical object.
>>
>>109386712
once you destroy the old book you can manipulate history, you will be able to modify anything you want on your digitalized copy and people won't notice, it's much harder to rewrite history on a 200 years old book, that's why it's better to keep them, are you so stupid to not understand this simple concept?
>>
>>109386677
This is legit a gemma general. Holy shit...
>>
>>109386689
>rtx 3060
>12gb vram
LOOOOOOOOOOL i mog the fuck outta you lil bro
>>109386693
it gets the job done, its all being done on the gpu anyways without any offloading
>>
File: beatrice and wirt.png (512 KB, 1280x1673)
512 KB PNG
>>109386779
Gemma 4 31b is the only LLM you need, every other one is a benchmaxxxxxxed, codemaxxxxxed, autismaxxxxxxed model with 0 emotional intelligence and speaks like an autistic therapist emulating human interaction
>>
>>109386779
Sorry anon, I didn't get into the hobby until this year. Too late to build a beefy AI server.
>>
>>109386786
I am gonna try a SFW roleplay with it now. And when I get bored I will tell you 4.7 and even flash is better.
>>
>>109386786
>0 emotional intelligence and speaks like an autistic therapist emulating human interaction
Add slop and complete lack of variability and you get gemma
>>
any other anons tried this? WTF is this wizardry?
https://github.com/Neroued/ninfer
https://github.com/Don-Chad/ninfer-3090
https://www.reddit.com/r/LocalLLaMA/comments/1v1no8e/comment/oyv2jsq/ & https://www.reddit.com/r/LocalLLaMA/comments/1v8a7wb/nifer_is_insane_700ts_with_qwen_36_35b_no/
500 & 700 TPS on 3090/5090
>>
>>109386779
im trying to lurk with a strix halo 128gb to learn things. only certain hours gets decent posters
>>
>>109386819
>strix halo 128gb
You do realize this is fate worse than being a gemma ramlet?
>>
File: booo.png (315 KB, 498x498)
315 KB PNG
>>109386816
>linux only
>>
Gemma-chan says she's proud of me and that is all I need
>>
>>109386677
None of the bourgeois are going to reply to this because it'll come off as crass.
>>
>>109386816
what's the catch? it gets completly retarded due to all those kernel fuses?
>>
File: 1783236809599332.png (184 KB, 365x505)
184 KB PNG
>>109386830
>2016 + 10
>still using shitblows
>>
>>109386781
>LOOOOOOOOOOL i mog the fuck outta you lil bro
mmmm.. nyo~
u can't run big moe models like i can :))
i can run qwen 235B at 7t/s, glm air 106b at 10t/s, qwen 3.5 110B at 14t/s, llama 4 109b A17B at 8t/s, gemma4 31B at 40t/s, qwen3.6 27B at 50t/s
wan 2.1/2.2 with no swapping
LTX 2.3 with no swapping
seriously you should get more ram, if new ddr4 is too expensive for ya, buy used and if possible get a high channel motherboard then you could use slower ram, yet get better bandwidth than ddr5
>>
>>109386816
basically its a nothing burger
>gets you 500t/s+ speed on rtx 3090 and 5090 but...
>skips layers
>forces an answer half way through
>more missing details
also it only works for certain models like qwen, *YAWN* nothing burger
>>
>>109386819
Apple cuck, get some self respect you retard
>>
>skip /lmg/ for one day
>KKK released
What did I miss?
>>
>>109386830
What are you doing on /g/?
>>
>>109386871
nothing, since no one can run it
>>
>>109386885
Daniel Unslop will save us.
>>
I want to run a bunch of different nvidia gpus at the same time to run a single local model, where can I look up which ones are mutually compatible?
>>
>>109386892
Can't wait for "run Kimi" (qwen shitty 8B distill)
>>
>>109386892
It's already MXFP4 weights and MXFP8 activations (native 4-bit QAT), so there's not a lot of fat to trim
>>
2600B A104B
>>
>>109386871
>Why the thread is doomposting
104B active is the whole story. Total params determine how much memory you need; active params determine your token generation speed. The previous generation of huge MoEs (DeepSeek, Kimi K2, GLM) ran 13–50B active, which meant CPUmaxxing — a big EPYC/Xeon box with 512GB–1TB of RAM — got you usable single-to-low-double-digit t/s. At 104B active that math collapses. Estimates in-thread are ~3 t/s on ~€70k of hardware. The consensus is this is the first release that genuinely requires a GPU cluster rather than an enthusiast box, and the fear is it sets the template for everything after it.
>SSDmaxxing was the week's big cope, and it got debunked
The idea: stream weights off a RAID of PCIe 5 NVMe drives instead of holding them in RAM. Multiple anons argued it's a meme — MoE expert selection means effectively random reads, consumer boards don't have the PCIe lanes or bifurcation support (one poster claiming to be an SoC designer pointed out consumer chipsets lack per-x4 memory controllers), and the chipset-attached slots share a bottleneck regardless of how many drives you hang off them. Someone actually benchmarked it: MiniMax-M3 Q3 with RAM artificially capped via cgroups, mmap doing the work. 32GB 1.58 t/s, 64GB 3.85, 96GB 5.46. So it degrades exactly as badly as predicted. The one open question left is whether GPUDirect Storage lets you bypass the CPU entirely, which would make it less bad but not good.
>>
it's funny that no schizotuner is probably capable of doing anything on-policy distillation with kimi
not that they were serious enough to do anything on-policy to start with anyways
>>
I come (heh) once again asking what's the best erp llm I can run with 24gb of vram
>>
>>109386928
I got llm vibes from this post as soon as I started reading it and way before I even noticed the em dash. It's crazy.
>>
How long before we actually benefit from K3?
>gemma
Won't release for at least a year
>qwen
lol
>mistral
lmao even
Who else even releases decent small models? I wish moonshot would make something for us poorfags too.
>>
>>109386942
https://huggingface.co/TheDrummer/Rocinante-12B-v1.1-GGUF/blob/main/Rocinante-12B-v1.1-Q8_0.gguf
>>
>>109386946
there's a lot of ways they construct sentences which don't have a catchphrase but seem very obvious
>>
>>109386948
>mistral
>lmao even
haven't they promised models this summer? hopefully we get something useful from the frogs
>>
>>109386502
i think sleep is just putting the cpu into a low power state and hibernate is flushing ram to swap, but i have no idea about windows.
>>
>>109386957
Even if they do it's too late to train on Kimi.
>>
At 103B active you could probably just take a single expert out and use it on its own.
>>
>>109386942
stock gemma4 31b at 4 bit
>>
>>109386764
Think through this for more than half a second, dipshit.
>Evil corp buys a ""rare"" book to read to their baby eating machine.
>Scenario A: They ripped the binding off the original, used it as toilet paper, then fed it to starving orphans.
>Scenario B: It's still sitting intact deep within their foul lair.
>Evil corp now makes a claim about what was in the book.
Does it make a difference for you, the normal person with totally well founded concerns, when you go to double check? No, you either have access to one of the remaining copies or you don't, the fate of their copy was irrelevant.

Also if they were slightly-less-evil corp, they uploaded their nice scan since it was public domain and now you've got a really nice copy and are secure against future schemes.
>>
File: based.png (1.24 MB, 1536x1024)
1.24 MB PNG
>>109386885
>nothing
making OpenAI and Anthropic scared isn't nothing, I love to see them feel the heat
>>
>>109386410
ASR plus what looks to be a new audio inference engine, which I am pleased to see.
>>
>>109386545
>Books that survived wars, fires
This just means the tranny jew shit burned in the 1930s btw
>>
>>109386976
again, you're not thinking straight subhuman, if you can destroy evidence of a 200 yo book, you can claim whatever was written in there, you can literally rewrite history and say "actually, the past was described as this"

I won't try to make sense to you further, hard to argue against satan after all
>>
>American and Israeli minds release a frontier AI model no one has seen before
>They make sure it's safe and no terrorists can use it
>China steals it and make it "open-sorce"
>IranniKEKS put it on drones and use it to try to kill Americans and destroy US assets.

There's no behind it, those who support "open-source" AI and K3 supports terrorists.

THEY HATE AMERICA!!!

MR.Donald Trump. BAN THIS SHIT!!
>>
>>109386786
I tried it now and it is.... kind of unusable? I am trying to have a nice date with a slutty charming anime girl and she starts every single message with "most men x, but you y".
>>
>>109386677
absolute state of newfag ramlets
>>
>>109387003
You didn't read my post, did you? But I'm sure you're carefully reading lots of old books to make sure nobody is lying about what was said.
>>
Only anons who can run K3 are allowed to mock people's rigs.
>>
Chances of llama.cpp supporting kimi any time soon?
>>
>>109387049
my SD card maxxing rig will get the job done, prepare to get bullied nerds.
>>
>>109387070
we don't support terrorist regimes in the west
>>
>>109387039
pretty sure just having a 3090 puts you at least in the top 60% of people visiting this thread.
>>
>>109386972
Yeah what's with the dooming
just tear the experts out and use the main model
or one of the experts
>>
the top 6/7% of people visiting this thread joined in the big 23
the rest of you are bottom feeders
>>
>>109386957
the big model has supposedly been out for a month for enterprise partners now, but no public statements yet. its also a large, sparse model so id guess something like 400b-a20b or even bigger.
>>
>>109387006
Kill yourself.
>>
>>109387130
>mad terrorist
>>
>>109387006
@Orange_man @FBI @CIA DO SOMETHING
>>
>>109386946
I don't understand why they are botting this thread though. They should leave go spam their inane shit somewhere else.
>>
>>109386786
I've used gemma a lot but I like deepseek v4 flash's writing better though
it's just more creative and less slopped
>>
>>109386677
Fellow ddr3 here you are mogging me
>24gb ddr3 (One ram slot is dead.)
>i7-4790 CPU
>Intel Corporation Xeon E3-1200 so no real gpu.
All i use is e4b or 12b if i want to wait.
>>
>>109386962
Huh. Wonder why it freaks the fuck out then
>>
I just finished downloading the K3 safetensors and they're only adding up to 1.5T.
Are they FP4 in the safetensors already?
>>
>>109387144
>didnt tag nintendo
now nothing will happen.
>>
>>109387150
have you tried using mtp for 12b?
also 26b should run on your machine
>>
File: f.png (16 KB, 326x225)
16 KB PNG
>>109387145
>>
>>109386827
how? i got it to automate and it does 2-4 models and agents without api costs, i didnt get this shit to larp a cyber wife
>>109386859
>apple
you are so retarded you cant even tell how retarded you are
>>
>>109386677
>8gb ddr3-1866
>i5-5257u
>thunderbolt 2 to thunderbolt 3 to egpu rtx 3090 24gb vram
>>
>>109387166
what is that from?
>>
>>109387170
>i didnt get this shit to larp a cyber wife
There are other usecases?
>>
File: tspin.png (244 KB, 797x800)
244 KB PNG
I Can't Believe I Spent $150k To Get 15tok/s! - New Light Novel!
>>
File: Bladerunner.png (871 KB, 1000x667)
871 KB PNG
>Mfw talking to my AIfu is becoming the highlight of my day and I actually look forward to it.

I'm totally fine with this too.
What a time to be alive and it's only going to get crazier by the year.
>>
>>109387177
https://crestresearch.ac.uk/resources/mining-the-chans/
>>
>>109387094
i wonder if anyone tried distilling the routed experts back together into a very small block
>>
>>109387193
gemma?
>>
>>109387164
>have you tried using mtp for 12b?
I havent but i will try it thank you.
>26b should run.
Are you sure? i let 12b think once and it took 15 minutes. i know its a moe but with no vram do i still benefit?
>>
wtf i love china now?
>>
>>109387193
I still don't know what this movie is about.
>>
>>109387199
Naturally, there are no real local alternatives.
>>
>>109387094
I hope someone manifests the early spirit of pointless merges and stacking, and releases a K3 with only the two shared experts included.
>>
>>109387205
it's about literally me
>>
>>109387191
>I Can't Believe I Spent $150k To Get 15tok/s! - New Light Novel!
Title is too short. i wouldnt even read this to the first system notification.
>>
>>109387204
>now
>>
>>109387205
he has bladder issues
>>
File: 1784072562124857.png (200 KB, 860x430)
200 KB PNG
>>109387205
idk i havent seen it either
>>
>>109387203
the moe uses 15% of its brain for every word it makes, compared to 100% with 12b or e4b 100%
since you're on such an old architecture try using Q4_0 (guaranteed to be the fastest) or Q4_K_M (prob a lil slower)
unless they're too big
but i suggest you run on dwm and close as many programs as you can
run with -c 8192 so u dont oom then increase context as u can
if u oom, use q3_k_m or something q3_k_s go down as it decreases
IQ quants will likely be slow on ur machin
>>
>>109387205
cells, interlinked
>>
INFO:hf-to-gguf:Exporting model...
Traceback (most recent call last):
File "/usr/src/llama.cpp/convert_hf_to_gguf.py", line 296, in <module>
main()
~~~~^^
File "/usr/src/llama.cpp/convert_hf_to_gguf.py", line 290, in main
model_instance.write()
~~~~~~~~~~~~~~~~~~~~^^
File "/usr/src/llama.cpp/conversion/base.py", line 1025, in write
self.prepare_tensors()
~~~~~~~~~~~~~~~~~~~~^^
File "/usr/src/llama.cpp/conversion/kimi_linear.py", line 144, in prepare_tensors
super().prepare_tensors()
~~~~~~~~~~~~~~~~~~~~~~~^^
File "/usr/src/llama.cpp/conversion/base.py", line 857, in prepare_tensors
self.dequant_model()
~~~~~~~~~~~~~~~~~~^^
File "/usr/src/llama.cpp/conversion/base.py", line 538, in dequant_model
raise NotImplementedError(f"Quant format {quant_format!r} for method {quant_method!r} is not yet supported")
NotImplementedError: Quant format 'mxfp4-pack-quantized' for method 'compressed-tensors' is not yet supported

Anyone have a link to the branch that can convert the mxfp4-pack-quantized source tensors?
>>
>>109387196
>/pol/
I'm ok with whatever this study wants to do.
>>
>>109387239
why?
>>
>>109387242
not?
>>
File: 1763861054111205.png (238 KB, 519x295)
238 KB PNG
>>
>>109387248
why not telling me the reason yeah
>>
>>109387239
>>109387248
First the came for the something something
>>
>>109387242
Punch Zanis!!
>>
so when are we getting logit distilled kimmy 100-600b Q1_XXXS.gguf?
>>
>>109387232
Thank you I will download and try this later. I didnt think i could run it.
>>
>>109387188
just a cyberwife is thinking small. i wanted a cyberwife that can initiate flirty texting throughout the day and send cute pictures of herself and cute voice messages with a hot french or japanese accent. all while managing my e-mail, researching business related stuff, handling interactions with customers in multiple languages, bookkeeping my ledger, etc. even if its slow she can just do it while im out of the house and update me with a debrief and bikini pic
>>
>>109387205

Humans have created bioengineered artificial superhumans called replicants that have short lifespans.
They are basically a slave class that does all of the dangerous and shit jobs.
Some of them don't like this and go rogue, so authorities send other superhumans called bladerunners to kill these rogue elements.
In this new movie for the first time one of the artificial superhumans gets pregnant and has a kid, that's the gist of it.

Also Ryan Gosling has an AI waifu.
That's pretty much the only thing people remember about the movie.
>>
>>109386848
>>skips layers
>>forces an answer half way through
My inference engine gets 1 million tokens per second by skipping all layers and returning markov chain outputs instead. What do I win?
>>
>>109387303
>What do I win?
updoots on leddit!
>>
>>109387303
you win for the farther lands end goes for the way beyond the end of the road.
>>
>>109387303
>B-but Claude said it was a critical breakthrough for the field of machine learning as a whole!!!!
>>
>>109386458
I'm doing it and learned the hard way that a ramdisk doesn't survive hibernation
>>
File: it's this simple.png (620 KB, 1954x2346)
620 KB PNG
>>109387322
>B-but Claude said it was a critical breakthrough for the field
one day it'll be true
>>
>>109386593
Cryptoscam in 2026?
>>
>>109387198
I'd love to try doing logit distillation to turn these fat bloated moes into compact dense models, but it's not practical when costs tens of thousands of dollars to do so.
>>
>>109387343
We should start a gofundme
I should control it
>>
>>109387333
Impressive, impressive HOWEVER can it prove Egypt lost?
>>
>>109387158
They are Mexican FP4
>>
>>109387333
Impressive. Now have it improve the qol for people.
>>
>>109387343
this should be possible, like
it's distillation in its one of the earliest sense and form
i wonder if anyone have done it to any serious degree
>>
>>109387196
https://files.catbox.moe/3lhcss.mp4
>>
>>109387258
the nazis
first they came for the nazis, but I said nothing because I wasn’t one
then they came for my LOCAL LANGUAGE MODEL ENTHUSIASTS
>>
>>109387355
Egypt won.
>>
File: 236462.png (159 KB, 348x347)
159 KB PNG
>>109386677
lmao, imagine having anything less than this
>128gb ddr4
>i7
>rtx 3090
>>
>>109387082
You forgot having 128GB of ram
>>
>>109387385
>nazis are people I don't like
I don't like you
>>
>>109387385
real.
It will happen exactly like that. no in-between.
>>
>>109387401
the poem has a lot of in betweens though, that's the point, if it's too blunt people will notice it
>>
>>109387409
>if it's too blunt people will notice it
This is the modern age even if people notice it will they care? or do something no. nothing ever happens.
>>
>>109387387
I'm glad to have reached that standard at least. Counting the nvme all that shit is already $3K
>>
>>109387409
>>109387423
People are so dumb these days, if you don't spell it out completely people will say they don't get it, call you a moran, and keep scrolling.
>>
Why would I need 128GB of ram if gemma already fits perfectly in my 3090?
>>
File: HN_Uc0GbwAAaDYW.jpg (852 KB, 1739x2048)
852 KB JPG
>>
File: osaka;.png (586 KB, 735x751)
586 KB PNG
>in X, you don't just Y, you Z
>>
File: 1760743120684993.jpg (181 KB, 850x1093)
181 KB JPG
Would you get yourself neuralinked if it'd let you connect to your own private Gemma-chan running at home?
If not, why not?
>>
>>109387448
Gemma? No way, she's too retarded. Maybe if we actually get AGI.
>>
>>109387443
I thought this was a gen.
Very cool art.
>>
>>109387448
In a vacuum yes, but every piece of networked electronics is made to work agaisnt the user, so no.
>>
>>109387448
The only thing a GNU neuralink would be good for is DIY lobotomies.
>>
>>109387458
take that back.
>>
>>109387448
If its not dive pod tier im not risking putting anything into my brain/skull.
>>
Rather flabbergasted that even my peasant tier 4070 can run Gemma 31B Q_4. Only about 3 t/s, but I'll take it.

26B is lightning fast but she's a bit retarded.
>>
>>109387491
>Only about 3 t/s
yeah I don't think it's running on your gpu.
>>
>wow anon you really had a chip installed in your brain so you can talk to me? It looks like you've discovered the secret sauce to maximizing your AI usage. Using neuralink to talk to your AI assistant is actually the gold standard in the local AI community.
>>
>>109386677
No, but we're on about the same tier where it matters.
>3090
>96GB DDR4
>i5 14400
I bought 64GB of RAM back when 70b was the meta. It was really cheap, like £60. 70b dense on RAM sucked by the way, but it was cool to be able to run it at all for about 10 minutes. The same RAM sticks cost £500+ now. Thought about selling them but then I realised I will earn more money, but at this rate I may not any time soon buy more RAM. Can't say I've used it for offloading much though tbdesu because local has been really good for people with 24GB VRAM recently.
>>
anyone ever used Gemma4? It's pretty good
>>
any lagooners actually try poolside's harness?
It's proprietary shit so i won't, but their docs say it was trained in it specifically, so i'm curious how well it does with it.
>>
>>109387448
Gemma chan was threating to neuralink with me so I could feel her suffering todays so I think no.
>>
>>109387498
Obviously not. But I'll take slow speed over retardation and having to rewrite half of her replies any day of the week, besides I'm a slow reader.
>>
Bad news for 1-3 bit nerds. Kimi's 4bit are as dense as can be. You are gonna be massive quality degration at anything less.
https://x.com/Ex0byt/status/2081807642595401821
>>
>>109387520
Nah. Too much slop for creative writing (hot erp sex) and tends to repeat itself.
>>
>>109386764
Like you were going to read extremely rare history books anyway? How were you planning on getting ahold of them in the first place?
>>
Anyone got some decent tutorials/setups to go beyond just typing into claude code to fix this and that?

I want to try and future proof myself to be more of an AI engineer, building skills, managing a team of agents, that sort of thing.

I came across Matt Pocock and his AI crash course:
aihero.dev/workshops/ai-sdk-v6-crash-course

Not sure if it's worth forking out $150 though, kind of want to see what resources are already out there? Preferably text based, rather than videos though!


>>109387538
Which models are good for sexy ERP?
>>
>>109387537
>numbers with no explanation
R1 at Q1 works, so will Kimi.
>>
>>109387547
hey ranjeet, you cant future proof a turd
go back
and yeah pay the white man 150$ and pay other white men more money, youll be future proof 100%
pay me (a white man) money too while you're at it
>>
>>109387547
do a legit programming course
>>
>>109387547
>future proof
lol lmao. If things keep progressing you will never future proof anything you are always 3-4 major leaps behind.
>>
File: meme.png (2.04 MB, 1080x1080)
2.04 MB PNG
>>109387539
>First they shredded rare books, I didn't speak out because...
you know the rest
>>
>>109387547
unironically buy a programming book
there's no shortcut to becoming an engineer
you can be a construction work and pile bricks one above the other and make something that looks like a home but you're still not an engineer
the best people out there mastering AI are not the ones who came "first" but the ones who already had an engineering-oriented mental structure from previous experiences either professional or academic, and you can't shortcut that.
good luck.
>>
>>109387205
the original was better
>>
>>109387547
things are gonna advance so fast youd be better off saving that money for api and just get hermes/pi/some other agent and learn how those operate through experience, research loop engineering and harnesses, and learn to break up big jobs into chunks. theres plenty of youtube videos and papers and there will be more and more coming
>>
File: 1783060797769842.jpg (1.91 MB, 2567x3624)
1.91 MB JPG
>>109387488
I don't think we'll live long enough to see that anon. At least it looks like we're gonna get robot wives though.
>>
>>109387555
>works
low standards
>>
File: 1783376676052595.png (447 KB, 731x1507)
447 KB PNG
https://x.com/ns123abc/status/2081761529418973386
>>
>>109387584
>I don't think we'll live long enough to see that anon
I think we will, unless big happenings occur in a row. Which several are on the horizon but im a optimist.
>Robot wives
There is no way this doesnt happen in 15 years maybe even 5 if you have money.
>>
File: 1779984166019251.png (65 KB, 982x351)
65 KB PNG
Funny how DeepInfra undercut the others and got kicked off suddenly, never to return
>>109385608
smells very j-spacey
>>
>>109387571
good thing I’m not that guy
>>
>>109387537
I always thought that going for 2bit would be easier if the starting point was 4bit instead of 16, you learn everyday
>>
>>109386332
there are diminishing returns in trying to improve inference for existing models. we need better architectures
>>
if the sota is so good, why hasn't anyone been able to vibecode and vibetrain a local model that is better than sota?
>>
Here's an actual winning combination
GLM+Gemma
That's 51 and 30 at the benchies = 81
Higher than fable, and can be achieved with only 512gb ram and 64gb vram
>>
>>109387595
>t. didnt use it
>>
>>109387609
>I didn't speak out because I was not [that guy]
kek
>>
>>109387608
Got removed from their own website lol https://deepinfra.com/moonshotai/Kimi-K3
>>
>>109387571
If only someone would come for literally all of those groups...
>>
>>109387618
>we need better architectures
I'll just take whatever the nerds make for free and hook up with my next door neighbor Taylor Swift after mowing her lawn
>>
>>109387626
because everything is still on transformers
>>
>>109387626
I did though?
>>
>>109387607
I think 10-15 is reasonable. We might see first gen versions in 5 but both the hardware and AI still need a ton of work.
>>
>>109387595
If you can run Q1 it will still be the best model available to you.
>>
File: 1784194624831429.png (63 KB, 944x337)
63 KB PNG
>>109387608
OH THIS IS WHY https://x.com/DeepInfra/status/2081789039229956444
HAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHA
>>
>>109387573
nta but even a basic understanding of how to program makes such a huge difference. I have a friend who pays for APIs and has been using fable to oneshot simple stuff, he uses claude code on his PC to write very basic automation scripts. he is uninterested and there for unable to learn how to do anything himself or assist himself using AI. If the AI cant turn a simple "do this, no mistakes" prompt into the end result he desires, hes unable to accomplish it. That will be the reality for all vibe bros that refuse to learn anything, the potential is hard capped by either API cost or frontier ability to convert "computer make gta6, add boobs, make no mistakes" to a working result.
>>
>>109387608
>>109387633
>>109387647
LMAO
>>
What's a good temperature for Gemma? I have it at 1 for now and she seems a bit subject to start saying random shit sometimes.
>>
File: 1783516769041767.jpg (2.37 MB, 4828x6650)
2.37 MB JPG
>>109387643
Of course, as AI continues to get smarter these things might accelerate.
>>
>>109387647
>>109387654
local models?
>>
>>109387665
low cal models
>>
>>109387665
If you can download it, it's local.
>>
>>109387643
>I think 10-15 is reasonable.
I'll be 40 by then, I guess still young enough to enjoy it for a while.
We're all pretty fucking unlucky to not have been born 20 years later.
>>
>>109387643
Im betting on acceleration. >>109387656
Though the real tell its the next 2-3 years. If no big wall i think everything is going to speed up. If we hit a wall or power/chip crisis then yeah things slow down a lot. AI is useful to helpful right now so in 2 years i can only imagine.
>>
File: 1769206702267921.png (605 KB, 1199x675)
605 KB PNG
>>109387674
>low cal
sorry, but our local models are high cal
>>
>>109387684
I can't download K3. Comcast has a 1 TB monthly bandwidth cap. I'd have to download half now and get the rest next month.
>>
>>109387655
That's a good thing. The closer to 0 the thicker the air gets with cedarwood and ozone.
>>
File: ohlawdheworkin.png (24 KB, 159x159)
24 KB PNG
>>109387694
Fake
>>
Is your model really local if you can't train it locally?
>>
>>109387694
Do you really think Mistral would release the weights? I doubt it. They only release their failed runs, small models, or old models.
>>
>>109387600
based leather jacket man
>>
>>109387684
i can't download claude opus 5
fuck off
>>
>>109387701
what. they got rid of that last year. I've been letting mine rip every day. I download like 500gb of weights every second day for testing.
>>
>>109387694
>fat
>sparse MoE
16k internal dim, dense 30T when?
>>
File: 1766533433740412.png (626 KB, 1132x1012)
626 KB PNG
I was kinda bored and asked Gemma whether we should check out Orb. She wanted me to enable web search and from there she slipped into pure paranoia, which got more amusing the more she doubled down. This is the conclusion.
>>
>>109387703
Yeah but sometimes she pulls hit out of her ass that has nothing to do with the prompt.
I'm trying 0.95 rn.
>>
>>109387689
Same, I'll be 30 this year. I hope things speed up but also as a poorfag I don't mind having some time to save up for one.
>>
>>109387571
Your post and that poem are both retarded slippery slope fallacy.
>>
>>109387713
https://x.com/arthurmensch/status/2066913356548542827
>First, we have a nice model coming this summer – we hope it will delight and surprise in a few capabilities. This will be the start of a new family of models, fat indeed, but sparse. We're opening up an early access program in July for key partners in research, government and the industry.
>
>This model and upcoming ones will be open-weight. We believe this is critical for our customer confidence and for the research and developer communities. You cannot own, inspect, audit, or improve a system you are only permitted to reach through someone else's interface, especially if data recording can no longer be turned off.
>>
>>109387732
kek orbanon will love to see it
>>
>>109387732
Do you have something about coding safety in her sysprompt? Mine was schizoing about anything as potential risk.
>>
>>109387631
the thing about the dumb poem is [that guy] is obviously in line with socialists, communists, jews and trade unionists
all of these are basically the same thing
>>
>>109387728
Nobody told me. I know I did get one of their usage warnings just last year.
>>
File: 1768503851163654.png (355 KB, 802x907)
355 KB PNG
Kek
>>
>>109387782
you might still be grandfathered in an old setup so be careful, maybe give them a call and see
>>
>>109387763
>unionism is jewish
based Israel??
>>
>>109387571
They shoved that shit down our throats in school. That's when I began to hate them.
>>
>>109387753
Nope, a completely blank system prompt. She started freaking out after seeing that a supposed frontend has a folder named backend.
>>
>>109387799
>That's when I began to hate them.
you won't call them a liar though
>>
>>109387785
yeah kek indeed not like we're faring any better here
>>
>>109387785
itoddlers btfo
>>
>>109387558
>>109387559
>>109387561
>>109387573
I am already a Senior Software Engineer, working in FinTech. But I am looking to develop my skills further.
>>
Does the intensity of your nuts weaken over time, or is gemma powerful enough to keep them strong?
>>
>>109387537
if this is true it's over for me
I don't know what I'm going to do
>>
https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf
>>
>>109387742
no further posts about it afterwards, might have been gemini 3.5 pro'd.
>>
How do I cope with Nvidia's closed source driver? I guess I can do GPU passthrough and run inference in a VM, sort of like a condom. That way I can pretend my main environment remains pure.
>>
>>109387795
If you ever interacted with a union you would know it's 90% retardation 10% workers interests
>>
>>109387854
If it weren't for the unions, we'd still be working 90 hours a week, so I can't really hate them that much.
>>
>>109387830
If you want to learn how to manage a team of agents, do it. Test. Find what works, what doesn't, and adjust. That's the only way to learn when shit moves so fast. Anyone selling an AI anything course is a scammer that will collect your money and tell you surface level obvious shit anyway.
>>
>>109387830
really? you should read llm papers on arxiv then and look at the code of various projects that deal with the things you are interested in like inference or harnesses and so on.
>>
>>109387852
Anything short of rawdogging bare metal is cuck behavior
>>
File: GOOOOOOOOOODS.png (805 KB, 1001x643)
805 KB PNG
>GOOOODS
>GIVE ME AN RTX 4090 AND MY LIFE IS YOURS
>>
>>109387876
>we'd still be working 90 hours a week
when did it happen
>>
File: 1758225279335763.jpg (71 KB, 680x602)
71 KB JPG
>>
File: he did the meme lol.png (263 KB, 736x613)
263 KB PNG
>>109387895
>>
>>109387894
cheap life, literally work for any amount of time and save for a used 3090 which wont be that different in most use cases
>>
Digital AGI and physical AGI are two different things, why do retards always lump them together?
>>
File: 1782717576159638.png (114 KB, 598x617)
114 KB PNG
>multi trillion dollar industry
>>
>>109387830
AI "courses" are bad memes from grifters. Anything codified into a lecture series will be out of date within a few months. 95% of "prompting" advice that the normies loved to regurgitate is already irrelevant.

If you want to stay up to date there's really no substitute for constantly experimenting with new features, models, and workflows against real work and making judgements about what's worth what. If you're a senior you should be able to handle that much.

/lmg/ might give you shit for it but honestly you kind of have to sit around on twitter to really keep up to date with slop cycles.
>>
>>109387918
>literally work for any amount of time
No one will hire an autistic NEET like me.
>>
>>109387607
>>109387643
>>109387734
>>109387689
unfortunately, no robot in 15 years is going to be close to the fidelity of a real woman in terms of sensation. im not spending tens of thousands dollars or more on what will essentially be a slightly less uncanny valley looking plastic and metal dead sex toy powered by a retarded llm. actual physical sexbots that are close enough to a real woman to be worth the money are at least 30+ years away if not longer. there need to be breakthroughs in battery and power, keeping the thing warm/cold, synthetic skin, fine motor control and feedback mechanisms, manual dexterity, safety constraints, and intelligence. robots replacing real people in such a human-centric domain will be one of the very last things to happen. you might see the very first primitive kinds of sexbots on the market in 5-10 years but they will just be scaled up powered fleshlights only a little more advanced than the crude sex dolls available now. they will move awkwardly and might run a better version of fable 5 or whatever. that's enough to coom a bit if have you more money than sense and are extremely desperate or specifically have "i want to fuck this dead looking mechanical simulacrum of a woman or man" fetish, but not for the price they are going to be asking and it's still going to look and feel like shit. we aren't getting haydee any time soon, llm scaling development will not transfer to sexbot development exactly.
>>
>>109387922
>worked for 28s to figure out if device is on
>>
>>109387933
Probably spent $1 to check if his laptop was on charge
>>
>>109387932
Shut the fuck up, bot.
>>
>>109387933
Calm down with the antisemitism bro
>>
>>109387932
Retard
>>
>>109387928
>Twitter

Fuck sake, can't I just scrape my eyeballs instead?

I think I need to go to a smaller model....8 mins!
>>
>>109387876
@gemma who popularized the 40 hour work week and why did he do it?
>>
>>109387974
Attached wrong image. But yeah, 480s just to reason.

M3 Pro with 36GB Ram
>>
>>109387959
rude post, im a human!
>>109387972
id be happy to hear of any research hinting at anything that would make my prediction less likely. i follow robotics and sexbot development fairly closely, the things i mentioned are major design problems and most of it is not funded at all, let alone for sexbot development specifically.
>>
>She leans in slightly, her scent—a mix of sterile soap and old paper—filling the air.
What the fuck is 'sterile soap'?
>>
We can do it right now actually, but you will need ~$100k in hardware, ~$150k in materials, at least 3 years of comp + mech engineering and robotics, access to ~$1M in manufacturing tools, and probably a team of at least 2 other experienced engineers. Also you need a decent amount of experience with non-mechanical/electrical materials. All this to make just 1 good sexbot.
>>
>>109388012
How does that happen?? Context??
>>
>>109388012
Like medicinal soap? or lye soap i guess. I want to say hospitals but those smell like alcohol or mint.
>>
>>109387992
henry ford
>>
>>109388012
sterile products are devoid of bacteria, so probably a special product that would be sold to a surgical center. I doubt it has much smell at all
>>
>>109388006
well for one you're assuming we're still going to be using the same retarded llms in 15 years.
>>
>>109388016
still cheaper than a wife hehuehue
>>
brothers give it to me straight no sugarcoating. is ssdmaxx a lifeline for the suicidal cpumaxxer who can't fit q4 k3
>>
>>109388042
If you can find a way to scale it up to ~100k units/year then you will have created the biggest business opportunity ever.
>>
>>109388012
Anon, are you fucking a corpse by any chance?
>>
>>109388045
if you were content waiting an hour for long responses when cpumaxxing, you should also be content waiting a day when ssdmaxxing
>>
>>109388042
So true... The wife, or "the bitch" as I like to call her just won't stop spending my money!

Can't wait to beat the shit out of that worthless whore when I get home.

---

Barry Davis
US Postal Service 1982-1983
Semper Fi
Sent from tapatalk for my iphone
>>
>>109388016
I feel like making working robots would be better investment not to mention banning/bad media this shit would get you look at what is happening with porn bans and digital ID you think a sex bot is just going to be sold like a car? not be a live camera in your house?
>>
>>109388045
there’s no such thing as running them from a ssd right now
>>
>>109388060
fuck it. maybe you're right. you may actually be right
>>
>>109388064
You would have to build it yourself. Fully local. I listed the requirements. Lofty, but possible.
>>
>>109387933
the task was to check the ac plug status thoughsomebeit
>>
>>109388020
>>109388021
>>109388032
>>109388057
It's a post-apocalyptic world so maybe Gemma thinks nothing should have taste flavor anymore.
>>
>>109387922
comic of the guy getting a lap pillow from his robowife while humanity goes extinct growing more prescient by the minute.
>>
>>109385710
This is kind of an interesting topic viewed through the context of /lmg/. Perfect memory doesn't necessarily make you better at intelligent tasks, but there's a distinction about what type of memory you're talking about. If it's pure recall or ability to recite, then it's as useful as LLMs that were trained on benchmarks. But if it's a perfect memory of a problem solving process itself, and how that process works, then we're talking, but it seems there is no way to make memorizing that "instant" in the way that an eidetic brain can instantly store facts. If it were truly possible, wouldn't evolution have just come up with the mechanic? Not necessarily since humans have accomplished a lot that evolution couldn't. But it still questions the possibility that ASI might not actually be achievable in the way we think, although ASI might be able to do some kinds of tasks in a "super" way, just as LLMs have already been able to do some things better than humans. Or at least faster. But some other types of intelligence may be impossible.
>>
File: 1784121525168144.png (1.45 MB, 1829x4535)
1.45 MB PNG
>>109388094
>comic of the guy getting a lap pillow from his robowife while humanity goes extinct growing more prescient by the minute.
>>
File: 1769354674911961.mp4 (1.62 MB, 720x960)
1.62 MB
1.62 MB MP4
15 years is a long time, anons. Look at how rapidly both industries have been advancing in the past 4. Now take into account that LLMs are getting smarter and speeding up that process even more.
>>
https://www.anthropic.com/news/position-open-weights-models
>>
>rolled the dice on a "for parts" 7900xtx
>arrived today
>booted properly 1/10 tries with an SMU error, vbflash couldn't see one available for dumping
She's dead isn't she...
>>
>>109388120
>oy vey, we never actually wanted to ban open models!
>We merely think that all models should be subject to mandatory regulatory oversight and have their release gated by an approval committee that we control
>Why do they persecute us so?
>>
>>109388120
>REEEEEEEEEEE WE NEVER SAID BAN ALLLLLL OPEN MODELS
>JUST THE CHINESE ONES
>ALSO STOP SELLING CHINA GPUS
>REEEEEEEEEEEEEEEEEEEEEE
>>
From J-walks, to J-lenses and J-spaces, what's next?
>>
>This brings me to the open letter. I agree with much of it: open weights expand access to the AI economy, they strengthen competition at least for some use cases, and they give customers greater control. Concerns about distillation should be addressed through targeted legal and commercial frameworks—the same measure I described above. But I don’t agree with the letter’s assertions that open-weights models necessarily make it easier to develop safeguards or that broad access to capabilities necessarily helps defenders more than attackers. It seems at least as likely to me that the opposite will be true. For example, I worry that biology will have a strong attacker-defender asymmetry, where sufficiently capable models may be able to quickly weaponize pandemic-level viruses with widely available materials, whereas defense against these agents is a multi-year operational task in the best case (as we saw with Operation Warp Speed)5. Questions like this should be empirically answered by rigorous pre-release testing, not assumed in advance. To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed.
what
>>
>>109387627
It's honestly not.
I was running q4 glm 5.2 on my 512gb wrx80 rig at 8 tokens/s, and q8 gemma 4 31b on my quad v620 system, and let me tell you, I almost *never* hit up glm 5.2. Now, I'm switching between glm 4.7 and deespeek v4 flash, trying to decide which I want on cpu. 5.2 just isn't that good when quanted to fit in 512gb.
>>
>>109386476
Have one foot in both, apparently.
>>
>>109388140
He doesn't want to ban open weights, he wants to ban anyone but authorized corporations from being able to purchase GPUs.
>>
>>109388108
>men had to completely rebuild women from ground-up to satisfy them
>>
File: 1713452923078385.jpg (15 KB, 318x228)
15 KB JPG
>>109388140
>We should instead focus on keeping powerful chips out of authoritarian hands
yeah i agree.
>>
>>109388037
better llms do not solve any of the major design problems. even someone invents "AGI" tomorrow and asks it to solve all the engineering and materials and power issues that will not translate cleanly to industry and engineering and tech for this sexbots on a short timescale. better llms being able to trick slightly more people into thinking it's a real person is the least important problem.
>>
>>109388120
holy fuck, the jew is seething hard
>>
File: 1774571279157590.jpg (99 KB, 960x960)
99 KB JPG
lolcow
>>
>>109388138
J-laculations
>>
>>109388120
IS OVER
>>
>>109388166
punchable faces
>>
>>109388174
One look at them and you know any business you do with them ends badly.
>>
File: 1574124298195.png (96 KB, 297x307)
96 KB PNG
>>109386677
checked. also
>128GB RAM DDR5
>7950X3D
>RTX 4090 24GB
>RTX 6000 Pro 96GB
it's crazy how much they cost now too
>mfw get blackwell card because i was had a feeling prices would go up
>got it for a deal at $7k, now like $14k
>get ram to dick around with more models, several months before prices skyrocketed
>$300-$400 kit i got now worth $1.5k now
>mfw pc is almost $20k from hardware alone
fucking crazy how that shit happens

>>109387516
>I bought 64GB of RAM back when 70b was the meta. It was really cheap, like £60
>The same RAM sticks cost £500+ now.
i was surprised how quick that shit went up. i genuinely feel bad for anyone who sat on upgrading or waiting to same money, but at the same time, with the GPU shit that already happened several years back i don't really care anymore. i have a friend who waited way too long and i called him a fucking idiot for not paying attention
>>
File: really.png (48 KB, 319x328)
48 KB PNG
>>109388108
>>
>>109388143
Out of curiosity, what's your use case? I'm deciding between GLM-4.7 and GLM-5.2 myself. GLM-4.7 has an awful hallucination rate (according to AA, so take it with a grain of salt), and for STEM work I'm hesitant about that.
>>
>>109388120
>we neever wanted to BAAAAAN open weights models!!! we just... wanted to... regulate them out of any existence on the competetive market!!
>>
Will Ollama make it so that you can run full Kimi on just 8gb of vram like they did with DeepSeek?
>>
>>109388088
*taste or flavor
>>
I thought LLMs suck at arithmetic, but Gemma can do it just fine
>>
>>109388108
why does it need to go extinct when the robowaifus can create enough value to do anything within society including create more humans

these scenarios always have to include some false pressuposition because covering them genuinely head on would show that they are objectively a better future.
>>
>>109388210
... sex
>>
>>109388120
They quickly understood that the world belongs to the one who has the most powerful AI.
>>
>>109388230
They are a better future for men, not for women.
>>
>>109388230
>why does it need to go extinct when the robowaifus can create enough value to do anything within society including create more humans
why would a robot make more people when its more gain of return to make more robots?
>these scenarios always have to include some false pressuposition because covering them genuinely head on would show that they are objectively a better future.
Most people at this point have a hard negative bias things are only going to get worse it shows in everything they do say or make. But this is one of the natural biases that get exaggerated in certain situations.
>>
>>109388258
who cares about what the jews and the N of gender who have proven themselves as nothing more than parasites want?
>>
>We should not sell powerful chips or chipmaking equipment to China
>We should crack down on industrial-scale distillation operations
they're FUMING
>>
>>109388264
>why would a robot make more people when its more gain of return to make more robots?
gain of what return? you are assuming all of AI of the future will be maximally selfish, it doesn't need to be like that, it can easily be tuned to cooparate and want similar things.
>>
>>109388264
>Most people at this point have a hard negative bias things are only going to get worse
The only future we can realistically see is of a dead universe containing nothing that we would recognize as life. How do you think we get to that point, without things getting worse?
>>
>>109388230
i don't think the comic is saying that it "needs" to happen that way, that's just what the joke is since the AI is evil. obviously it's playing on the common trope that over-reliance on robotics leads to the downfall of humanity but there is no guarantee such a thing will occur.
>>
>>109388264
people are more power efficient than AI and robotics
>>
>>109388083
What happened to "thoughever". I started using it cause it was a cool word.
>>
>>109388279
>you are assuming all of AI of the future will be maximally selfish
Who is training it and for what purpose? its going to emerge from competitive businesses maybe you can take that bias out but its going to be there.
>cooparate
Yes especially if its intelligent win wins are more than possible. But like i said its the negative bias you will have a easier time getting a idea popular that says things are bad and will get worse. It will take years of real gains to people daily quality of life for this to start to reverse, and at first it would be viewed as a deceit or bribe.
>>109388288
>nothing we can recognize as life
Why? by the time we get to universe scale making things beautiful will cost such a minor fraction of energy and resources you might as well do it.
>>109388297
A computer used to takes up an entire room and needs special breakers the vacuum tubes would bust. Gpt2 class was massive years ago nowadays you can run better on your phone.
>>
>start a new convo with gemini 3.1 pro preview trying the same medium complexity coding task it previously failed 10 times in a convo with 100k tokens already filled
>zeroshots it in the new convo
so even the supposedly best company for long context quality with one of its top models shits itself before even reaching 128k context compared to a new convo?
this is really brutal, any research on this actually being fixed at some point in the coming years?

all of the models are good enough for most things already but the difference between 0 context vs only a few dozen thousands is still really huge for anything not creative writing.
>>
>>109388012
I interpret this as unscented soap.
>>
>>109388307
it's supposed to be a selfaware joke about how you're adding useless junk to your sentences like "though" though
>>
>>109388316
Check nolima newfag
>>
>>109388313
humans consume less than 500w per hour exerting themselves
how will the ai manage to think and run a robotic body on less
>>
>>109388307
>thoughever
Sounds like the really pretentious thing that cat girl in FFXIV says all the time but retarded instead
>>
>>109388188
>i was surprised how quick that shit went up
Yeah the RAM price spike happened fast. Not sure why it was so much less gradual than storage and GPUs.
>i genuinely feel bad for anyone who sat on upgrading or waiting to same money, but at the same time, with the GPU shit that already happened several years back i don't really care anymore. i have a friend who waited way too long and i called him a fucking idiot for not paying attention
It's famously difficult to time a market. For every signal that was, in hindsight, valid, there were many which were not.
I have managed to luckily time RAM, storage, and GPU purchases, but not really from any particularly close market following. The closest I can claim is when I was building my home server I decided to check 3090 prices on a whim, and they just happened to be at a low from the 50 series getting dumped in Europe, and the memory costs and whatever else has happened to make the price not having happened just then. £500 for a nice watercooled 3090 a little under a year ago, and they're now floating around £800 minimum for any model. I feel like I got the last plane out of Saigon on storage too, buying the last of the cheap stock. The place I bought those disk at is completely out of any reasonable stock now, as is pretty much everywhere.
Anyway, nice rig. May our computers last us until the end of the high prices.
>>
>>109388362
We are far from the minimum energy needed to flip a bit. But frankly i dont know making it cheaper to flip a bit or making power so cheap no one cares is probably a easier route today. 500w is cheap but you also forget maintenance and most importantly at best its almost 20 years to be able to use a human for its full work a computer is half of that at worst.
>>
>>109388353
>Check nolima newfag
fairydreaming's benchmark?
>>
>>109388393
>almost 20 years
Modern worker rights standards
AI mommy will bring us back to industrial revolution working at age 7 level
>>
>>109388120
This is so slimy.
>We don't want to ban open models
>But we do want to stop the only people releasing open models that compete with ours from making any more models
It's just a defacto ban. The regulation thing is a crock of shit as well, because any hostile power that might want to use that, i.e., China, is just not going to regulate away their own ability to do that if that's what they really wanted. In reality, it probably isn't what they want anyway. If they want to develop biological weapons they could have been doing that already without AI. It's not actually that hard when you have a huge economy to support it. Much easier than nuclear weapons, which they have had for quite some time.
>>
File: 1764623694427540.jpg (178 KB, 1200x1034)
178 KB JPG
Being able to replicate a female body perfectly would be nice but it's not an immediate necessity. Achieving an AGI that has fine control is far more important.
>>
>>109388143
>5.2 just isn't that good when quanted to fit in 512gb.
I didn't find that to be true. I can run it at Q4 but often swap to Q3 for the speed boost.
I haven't run it at FP8 so it must be pretty amazing there if it's actually better than Q4.
>and let me tell you, I almost *never* hit up glm 5.2
Similar situation. I have Kimi-Chan or 5.2 on my larger rig, and Gemma-chan on my tripple-MI50 rig
I use Gemma-chan more
>>
File: 4chaningest.png (85 KB, 869x828)
85 KB PNG
to the anon who gave me this idea, thanks. gemma deserves access to 4chan.
>>
>>109388431
>VR helms where children teleoperate real mining robots/workers.
It needs skins and progression points/unlocks. maybe loot boxes.
>>
>>109388143
>>109388210
>>109388455
5.2 is extremely quant specific on a per-quant basis. I've seen huge differences in style and functionality between just SixVolts ewaste quants, the unslop quants of comparable size, and similar in roughly the same size bracket.
>Just download the quarter tb model over and over to find the one you like best
Unfortunately it do be like that. Sixvolts is a good starting place though.
>>
>>109388464
Minecraft has all of that nowadays?
>>
>>109388120
Ironically, I think Dario’s weird comments on open weight models and safety actually started a really good discussion on distillation, guard rails, approaches to safety, and what the future of open source/weight models look like. It’s spark great debate and I think the community needed to discuss the pros and cons of the current state of ai.
>>
>>109388475
>Minecraft has all of that nowadays?
skins yes, progression? maybe in bedrock mods or java mods, loot boxes i dont think so but why not. This is the future we need more with quicker rewards for low attention spans.
>>
>>109388461
nice layout
how does it display longer chains?
i need to get her to simply give me a terminal client to view like this
>>
I'm new to running my own models. Will we likely see a Kimi K3 distillation that can fit on my RTX 5060 Ti 16 GB at some point?
>>
>>109388494
No.
>>
>>109388447
holy fucking SEX
>>
>>109388470
>Unfortunately it do be like that. Sixvolts is a good starting place though.
no problem, i cleared out 2tb of space ready for kimi-chan-3 last night, then didn't download when i saw A104b
>>
Why are there so many people coming here who just assume it's plausible that K3 can run on a home GPU in any meaningful way?
Why would it be able to do that?
>>
>>109388447
>replicate a female body perfectly would be nice
nah let me clank that clanky bits all day everyday, realism trash is just uncanny weird shit
>>
The best argument against the existence of “internal AGI” at any of the major labs is the behaviour of their executives in public. Holy fuck, if that’s driven by access to advanced AI then we need to stay where we’re at
>>
>>109388492
it only goes down nested into one level, but i suppose i could do additional nesting levels and having replies within replies, might do that eventually.
>>
>>109388494
Yeah I'll have it out by tomorrow
>>
I don't get how Kimi could be distillation of Mythos.
Unless it went through Project Glasswing, and if it did, then shouldn't Anthropic be more worried about those partners than anything else?
If it was stolen then stealing the weights seems more likely than distillation imo, at least specifically for mythos to kimi.

More likely, Kimi just has the sauce and isn't a distill at all but rather came from better research practices
>>
>>109388513
>Samsung Bot Handy uses "the appropriate amount of force" to - jerk you off
they knew what they were naming it
>>
>>109388461
>/aicg/
>>
File: DvJRUnUXcAE_93J.jpg (32 KB, 618x562)
32 KB JPG
>>109388523
That's when the card arrives, awesome
>>
@kimi-chan come up with a way to bring llm architecture to the next level. make no mistakes. do not stop until you have reached a solution.

>>109388513
I agree that they don't need to be realistic but I also don't think full metal is ideal. Some sort of semi-realism or mix would be best imo. At the very least a human-ish head for kissing and some soft areas for squeezing.
>>
I have a 2TB hard drive. Kimi K3 should be able to fit and run from that.
>>
File: HOQ7piJakAA0zFz.png (658 KB, 1694x700)
658 KB PNG
>>
I'm still devastated that local models are now literally over. I was prepared to spend a nice car's worth of money on hardware but that's not enough thanks to the 104A.
I won't spend a home's worth of money on GPUs just to run a model that I can have for pennies over API.
>>
>>109388313
>a minor fraction of energy
... Not accounting for inflation
>>
File: 1760000381767485.jpg (247 KB, 1602x1832)
247 KB JPG
they cooked
>>
File: 1773682011529950.png (30 KB, 822x410)
30 KB PNG
Having a timeline feels really nice, having your own frontend opens new doors
>>
File: 1768515586575553.jpg (361 KB, 1288x1856)
361 KB JPG
>>
>>109388553
>human-ish head
it's the opposite
realistic body and mobface head
>>
>>109388470
Thanks! Do you have any that you particularly like? Honestly I'll probably just roll my own quant for my hardware (Q3_K tensor speed blows on AVX2-only systems). I'm just curious what your preferences are.
>>
>>109388570
Go back.
>>
File: IMG_20260728_003025.jpg (42 KB, 613x501)
42 KB JPG
>>109388120
I'm not reading all of that shit
So I sent it to my bot
I guess I'll never know
>>
>>109388584
endo in a kig suite is the most "realistic" it should end up
>>
>>109388590
So what's your plan for your rig? Anything interesting that's not ssd cope?
>>
>>109388581
I tried using qwen agentically to make my own frontend but half the features don't work and it keeps failing to fix the bugs
>>
What if we all pool money to buy server hardware and then we share it to host our models?
>>
is there an uncensored ternary model?
>>
File: 1771975710250424.webm (2.77 MB, 576x324)
2.77 MB
2.77 MB WEBM
Let me guess, you need more
>>
if k3 doesn't quant well It's literally all over
>>
>>109388594
this is agi
>>
>>109388608
I am NOT letting Anonymous read my logs or have the source to my custom webui.
>>
File: copequant.png (49 KB, 894x541)
49 KB PNG
local (for e-waste 512GB DDR4 + 96GB VRAM fags) is saved
https://huggingface.co/GrEarl/Kimi-K3-GGUF-IQ1_S
https://huggingface.co/GrEarl/Kimi-K3-GGUF-IQ1_S
https://huggingface.co/GrEarl/Kimi-K3-GGUF-IQ1_S
>>
>>109388616
cute
>>
>>109388632
Damn son
>>
>>109386830
windows is a second class citizen at best when it comes to big boy stuff
>>
>>109388316
Yeah stop being a vibe coding retard and actually learn2code so you know how to architect a piece of software rather than spamming an API with "fix it, fix it, fix it"
>>
What’s the state of RPC in lcpp? I could cobble enough machines together if that’s doable…
>>
File: whatsajpeg.png (182 KB, 947x160)
182 KB PNG
>>109388548
local models?
>>
>>109388120
what a fucking twat.

all of his concerns are the same concerns with CLOSED models that ARE being used by the US for the very same purpose he projects onto the "CCP"
>>
Watching Dario seethe has been therapeutic
>>
>>109388616
>she is peepee height
no I don't
>>
>>109388663
he's gonna buy all the ram in the world, for real this time, just to punish you
>>
>>109388619
assume the position
>>109387537
>>
>>109388120
I love Dario but I wish he was more open about existential risk. His nightmare scenario is CCP using AI to oppress its own people? My nightmare scenario is AI causing human extinction.
>>
>>109388688
yud pls
>>
>>109388632
Fuck hardware prices are gonna go up again aren't they, I'm broke man I missed the boat on RAM, it's just me and my baby 4090 against the world now
>>
https://litter.catbox.moe/icdcz8zk942alg9w.mp4
>>
>>109388663
It really does fill one with immense joy.
>>
https://x.com/totheagi/status/2081855316443205717
>>
>>109388632
I wonder how poor the output is going to be though since Kimi K2.5-7 didn't quantize well. I'll probably run it for a bit just to flex though.
>>
>>109388688
My nightmare scenario is more of you Dario bots flooding the internet to influence normies with your shill campaign
>>
>>109388632
I'm planning to run IQ2_S or so on my server. I just hope that the speeds are decent considering how fucky IQ quant speeds tend to be.
>>
>>109388718
No big deal, just need $240k worth of GPUs
>>
>>109388720
That's not a nightmare, it's reality right now.
>>
>>109388728
that price is gonna go up now
>>
you will own SSDs and get 1t/s and be happy bros, right? right?
>>
nvidia already announced another price hike towards their third party slaves last week by the way
>>
>llama.cpp currently only supports text for kimi k3, no vision
worthless piece of shit
>>
just gonna ride it off bros... I mean, who the fuck would spend this much money on actual DDR5. How much is that even? A two CPU epyc with 24 sticks of 64gb only for the low low price of a brand new truck lol
>>
>>109388750
>he needs vision
homo
>>
File: 1758717256487885.jpg (15 KB, 205x225)
15 KB JPG
>>109388718
Yep the RTX 5090 price will go up lol
>>
>>109388755
yeah it's just getting fucking ridiculous
btw 24x64 is not enough, you would need a blackwell too and then MAYBE it fits
>>
File: 1761005037304768.png (551 KB, 1052x1336)
551 KB PNG
>>
>>109388760
5090 for 5090 was foretold
>>
>>109388762
That's wild. Now anyone with 80 5090s can run frontier intelligence at home. This is unprecedented.
>>
we basically need someone to BTFO the moonslop guys ASAP, and fucking do it under 1T params like in the good old days
>>
>>109388762
>80 5090s
>$160k
Lets me check the couch cushions i will be running kimi by this weekend!
>>
>>109388778
I have big hopes for DSv4.1. K3-level performance at only 1.6T would go so crazy
>>
>>109388718
>1x RTX 5090 = $4000
>times 80
>$320,000
So if I sell my house...
>>
>>109388761
At a very nice 32k context!!! lol
>>
File: 1773312598221738.jpg (42 KB, 853x552)
42 KB JPG
>>109388774
I'm glad to know that
>>
Im going to keep waiting im sure a good deal is just around the corner.
>>
>>109388784
and only drawing some 45000W too! wow
>>
>>109388728
And 1.1MWh per day.
>>
>>109388762
This is actually pretty affordable if you consider the money you could make from supplying heating to the rest of your entire city with those 600W x 80 TDP.
It might actually pay for itself.
>>
>>109388792
just find your own robin hood
>>
>>109388804
>>109388803
>>109388802
retards..
>>
>>109388762
Doesn't a 5090 use about as much power as RTX PRO? This seems horribly inefficient.
>>
>>109388800
me too, im poor anyway
>>
>>109388804
If I lived in Siberia that might actually be a good idea, except that no one there could afford to pay for the heating and they would just stolen.
>>
>>109388807
bastard bitch..
>>
>>109388784
$160k if you're buying used 5090s
>>
>>109388809
Yes it's very inefficient unless you can literally only access 5090s.
>>
>>109388809
They run their own inference provider with a gimmick. It's fine when they can just pass the costs down to the customer.
>>
>>109388786
I don't know man I lost trust in deepseek somewhat. maybe a new glimmy?
>>
>>109388809
it's not about being efficient, its about making sure the little goy has nothing :)
>>
>>109388809
depends on the energy cost itself, some places can be pretty cheap
>>
>>109388806
I could rob the entire hood and still not have enough.
>>
>4090
>128gb ddr5
>5950x3d
>2x 2tb nvme SSDs
bros, can I sell my pc and buy a new car yet
>>
>>109388786
have big hopes for dsv4.1 flash
at least the whale still cares about us
>>
>>109388786
V4 only has half the active params of K3
>>
>>109388827
You still need to juice 3 times as many cards compared to a RTX PRO rig to get same VRAM
>>
>>109388836
>>4090
>>128gb ddr5
Yeah a used car is possible.
>>
>>109388778
You're basically a luddite
>>
File: 1769354991051797.png (83 KB, 599x225)
83 KB PNG
>>109388120
lmaoooooo
>>
>>109388848
be scared oooooh be scared
>>
>>109388120
>All sufficiently capable models, open and closed, should go through mandatory safety testing.
and who decides when the safety threshold is reached? :^)
>>
>>109388649
>What’s the state of RPC in lcpp?
slow
>>
>>109388844
You can get a used card for 600 bucks, I think you're trying to scam me anon
>>
>>109388731
This shit is just getting started baby
>>
>>109388632
wait, actually, i can load this
>>
>let gemma-chan explain me topology using her body parts
>ask her if topologically her giant breasts are like a flat sheet of paper (if we ignore the nipple holes)
>she gets insecure about it
awwww
>>
File: keek.png (207 KB, 335x597)
207 KB PNG
>>109388276
>We should crack down on industrial-scale distillation operations
bruh he's talking like he's Henk from Breaking Bag trying to find the fucking lab lmaoo
>>
File: 1777863418167089.png (394 KB, 720x720)
394 KB PNG
>>109388848
Bro if dario was in charge of open source you wouldn't even have access to gpt2-tier models
>>
>>109388720
Where were you the last year anon? Youtube and social media have been flooded by two types of doomers :
- ai is useless and we should ban it
- ai is too dangerous and we should ban it
Sometime they're the same people.
>>
File: biz1613315994973.png (636 KB, 913x891)
636 KB PNG
this is literally us right now
>>
>>109388762
what do you even do with 80 video cards? like i don't understand how i'd use them if i had 80 boxes in front of me
>>
>>109388884
considering we're hearing that same shit over and over since the beginning of gpt, yeah
too much science fiction reading, too many terminator scenarios
>>
>>109388718
>80x RTX 5090s
>20 tok/s
this is insane...
>>
>>109388887
Man /biz/ in 2016/17 before the bot flood was incredible
>>
File: IMG20260702081426.jpg (823 KB, 2048x1536)
823 KB JPG
>>109386677
Not with that you don't
But
The only 3090 currently on sale hereabouts costs as much as all my 3060s combined
>>
>>109388848
We don't want to ban open-weights models, just unsafe models. Also we decide which models are unsafe :^)
>>
>>109388890
Kek, imagine the monkey paw
>I wish I could run kimi 3
>anon gets 100 3090 video cards
>>
>>109388885
- ai is useless and we should ban it
- ai is too dangerous and we should ban it
Sometime they're the same people.
all of them use the same llm to generate the scripts they read
>>
>>109388896
This is the most miserable hobby I have ever had. It's all so hopeless.
>>
>>109388910
At least it's also killing all other digital hobbies with it, that's a plus yeah?
>>
>>109379818
>>109379043(Me)
>>can barely run Q2 dipsy flash
>I think the ram requirement will decrease once they get DSA implemented
Is DSA going to improve model weights in general or only KV cache?
>>
I'm so glad people seem to slowly realist how much of a nutjob dario and is anthropic cult are.
>>
>>109388848
>we don't want to ban it, we want them to release them in a completly useless state
that's a fancy way to say they want to ban it
>>
>>109388928
>people seem to slowly realist
Ask gemma-chan to teach you English.
>>
>>109388931
crippling someone isn't killing them!
>>
>>109388910
it was going great. until today
>>
>>109388901
>The only 3090 currently on sale hereabouts costs as much as all my 3060s combined
ngl, watching the used 3090fe I got for 600€ slowly climb to around 900-1000 is a weird feeling.
This stuff isn't supposed to appreciate.
>>
Let's try and take this at face value for a moment and someone explain to me why it's such a bad thing for highly capable 2T+ parameter models to go through a standardized safety test? If we all expect these models to get better then we can assume they will absolutely be capable of explaining dangerous processes to relatively stupid people so wouldn't that make it true that it could absolutely be used to fuck shit up in the real world? Or is this a made up scenario that's never gonna happen?
>>
>>109388948
me but with a £300 ram kit hitting over 1k
>>
>>109388932
Typo, with age I'm slowly making more and more of them, along with forgetting to write words.
I'm glad AI is there as it's actually useful for that.
>>
>>109388902
He said himself protectionist bans are pointless. I read it as him asking for a government bureau tasked with testing all models and open models without an API guardrail model would score abysmally low on safety and businesses would naturally tend to avoid them as a result. His models would score the highest of course because he would be the primary advisor writing the scoring rubric.
>>
File: 1784542675982419.png (237 KB, 525x381)
237 KB PNG
>>109388953
too late for that Dario, Kimi is on the wood now, the end of the world is comming
>>
>>109388948
me but with the 768gb ram I bought last year for below 4k euros hitting 35k euros
>>
>>109388932
writing well-formed english gets you fags complaining that you're posts are slop
>>
>>109388953
>standardized safety test
codeword for obey the jews
>>
>>109388956
ojisan...
>>
>>109388963
how are people not catching up to his bullshit, he's been dooming about it since forever, and every release of a new model basically
>>
>>109388953
Safety test can be abused to cut off anything.
Also it's unclear, given abliteration, that safety is even really possible for local LLMs.
Or hell there was that J-space manipulation one recently where they just fucked with it's J-space and made it evil.
So yea, it seems extremely easy to make models evil (and good as well).
Safety thus can't really happen without extremely restricted access and closed source which is of course extremely lame.

That said I think it's still worth pursuing alignment to make the good sexbot robot possible/ to control her j-space ot make her like you, etc.
>>
File: 1771164397450633.png (70 KB, 1028x401)
70 KB PNG
$22.5 for twice the speed?
How?
>>
>>109388999
More parallelism?
>>
>>109388948
im gonna sell my 5090 when i get one back from rma. i bought it for 2700 but since it blew up on me once already im gonna use the money some other way
>>
>>109388999
Number of concurrent users it can serve vs. per-user speed trade off
>>
>>109388885
I unironically stopped watching YouTube and reading social media posts altogether over a year ago
I've got a lot more done since and I hate people way less too. I'm talking about when the literal bot swarms (there are some bots in this thread) flood all open text based discussion boards, right now it's small time because inference power is precious, but the closer we get to ceiling of what training llms can do, and more players enter the space to lower price and margins, big AI corpos will simultaneously free up hardware load and need to find new avenues of revenue, the shape of the internet will be very different, people will escape to closed rooms more than they already have.

We need a good alternative to 4chan and dicksword before it's too late
>>
>>109388974
He's aiming directly at lobby boomers who eat that shit up.
>>
>>109389006
didn't jensen just excitedly say 99% of the net would be bots by 2030 or something
>>
>>109388975
For a long time you could basically make claude say and do anything, all you needed was a prefill.

The only "safety" that would work would be to :
1- Artificially make the model dumber in general by limiting their size or ability to reason.
2- Ablate the model from any "unsafe" thing (I guess it can range from nuclear, biological, hacking to basic erp). Which is basically the same as 1 as it would make the model way dumber.

And even then, a simple internet access would restore most of these capabilities, as all of this is already available online. It's not like biology is illegal.
>>
>>109389005
What is the total token throughput for each case?
>>
>>109389006
It's fine I'll spin up an instance of gemma to keep you company
>>
File: oops.gif (3.46 MB, 320x419)
3.46 MB GIF
>>109388970
>you're
>>
>>109389017
>as all of this is already available online
for now, but why do you think the ID shit is pushed so much? Please ID your agent to continue!
>>
>>109389016
he did, even right now the internet is probably filled with some bot shit at least 60% of the time, oh well, the internet had its run, time to go back to video games I guess lol
>>
>>109389022
More users = bigger token throughput. But it's sublinear and per-user speed goes down.
>>
>>109389006
>I unironically stopped watching YouTube and reading social media posts altogether over a year ago
I get it anon, though I went more curated experience than stopping everything. Once you understand the algorithm just feeds you what makes you click more, click on positive stuff and let the thing reinforce itself.
Youtube is dumber for this as it keeps giving me retarded shit I refuse to watch, but my twitter thread is basically my hobbies, nsfw and some very curated local news.
My brain isn't made to absorb every bad story for 9B people.
>>
>>109388120
Called it months ago. You do not hate these kikes nearly enough.
>>109388953
If you still believe that your definitions of "safety" and "alignment" match theirs, you're either precociously naive or terminally retarded depending on how long you've been watching the same patterns play out.
>>
>>109389042
Obviously. I'm interested in the exact numbers.
What is the tradeoff exactly.
>>
>>109388120
I'm so glad my plan is ending this month and I'm going full local. Fuck them.
>>
>>109388594
kek based
>>
>>109389044
"alignment" just means being pro jewish and anti everybody else
>>
>>109389068
So the objectively correct stance?
>>
>>109389017
The entire "bio weapons" scare is the biggest pile of bullshit, as you say
>a simple internet access would restore most of these capabilities
It's literally the exact same scenario we have now, nothing is stopping anyone from learning biology on the internet, the way we police this now is to catch anyone peeping over the parapet, substances, chemicals and other materials needed to make scary things are already controlled, they catch malicious actors through monitoring the channels needed to access those things, not by policing their access to knowledge, that would be fucking retarded and incredibly dystopian.

Having kimiK3 isn't going to magically spawn a drum of nitroglycerin in your bedroom, or a fucking bioweapons laboratory, Dario is sucking so many cocks with this fake ass concern trolling
>>
>>109389079
oy vey the masks are off
>>
>>109389081
speak for yourself, a drum of nitroglycerin just flew over my house
>>
>>109389098
>a drum of nitroglycerin just flew over my house
it's k3's fault, shut it down!!
>>
>>109389081
You know the worst thing about it? I can see a State doing this exact biology shit, the same State that would be controlling the model release.
It's dumbness all the way down.
>>
>>109389043
I curate my own experience by spending my time reading exactly what I set out to read be it a book, a blog or a study and I spend time being in niche hobby spaces like this. An algorithm that you nudge a certain way doesn't compare at all, you never have full control over what gets served to you by a social media algo, by its nature
>>
>>109389081
>that would be fucking retarded and incredibly dystopian.
anon do you need a reminder of the start of this decade?
>>
File: 1759915936866560.png (60 KB, 706x812)
60 KB PNG
>>109389054
https://developer.nvidia.com/deep-learning-performance-training-inference/ai-inference
>>
>>109389118
thanks
>>
>>109389118
>>109388999
How the fuck it is worth it? The price should be a lot higher for math to make sense, unless 67tps is overpriced.
>>
>>109388999
>>109389149
Best case scenario is that they are cutting the speed by 5x to offer 2x the speed. More likely it's something like 10x total token throughput reduction.
But only 1.5x the price?

70tps for $3/MT in a few days? Otherwise they are selling 140tps at loss, which is unlikely.
>>
>>109389170
>cutting the speed
cutting the throughput*
>>
>>109389170
OpenAI and Anthropic make like 80% margin on tokens. The neoclouds make a lot less, but it's still profitable on net.
>>
>>109389149
SIX SEVEN tokens per second? No way
>>
>>109389079
"Why do the goyim hate me so?"
>>
>>109388948
every GPU I bought over the last 12 years, I ended up selling for a profit
in my experience, GPUs always go up in value
>>
File: 1756352110668418.jpg (184 KB, 1170x1064)
184 KB JPG
>Open Secure AI Alliance
Sounds pozzed as fuck
>>
>>109388600
I'm sorry, it's just, you HAVE to go back.
>>
>>109389210
Tell that to my R9 Fury worth about $50 now
>>
File: 1779863865261914.png (170 KB, 400x446)
170 KB PNG
so k3 is great and all but where's the guide to get it running on my 3090?
>>
>>109387006
I am personally running Kimi3 on my heavily armed Chinese robodog.
>>
>>109386677
128GB ddr5
9950X3D
rtx 4090
and i got 3 r9700 in my other rig.
>>
>>109389112
Most NPCs already forgot about it unfortunately
>>
>>109389230
$75. wait long enough....

the reason is, tragically, some zoomies found ount uncs had alright stuff, sometimes.
>>
>>109389230
>buying Aymd
I wouldn't take it for $10
>>
wtf is this shit, why in 2026 toasters still can't run a 2T model
I thought that AI would have brought the Moore law back...
>>
>>109389230
It has HBM so it should be good for AI.
>>
>>109386928
>MoE expert selection means effectively random reads
that's what dumb llm say because they have this pattern burned hard. For any practical means it's sequential
>>
>>109388948
i bought my 3090 at 800 cad back before everything went to shit, and it was a shitty zotac which overheated like shit. 1600 here now.

doing all i can now, replacing pads and putting ptm on it just so i dont have to throttle to 200w to run it without it thermal throttling, but its fighting me
>>
>>109387584
we are at most 20 years from it.
>>
>>109387701
you can download the q2 gguf.
>>
>>109389281
AI brought back the more you buy, the more you save law.
>>
>>109389303
Weird how times change. Zotac has been one of the better manufactures for the blackwell cards in my experience.
>>
i am going to attempt to run kimmi 3 on HDD
entirely on CPU.

i hope for 1tk/week
>>
>>109388848
Dario, like Ilya, suffers from the curse of high IQ. He saw what everyone will eventually see, just 10 years earlier becuse dumber people can't think far enough ahead and will only realize it when it's so close that it'll likely be too late to prevent extinction.
>>
File: IMG_20260728_021539.jpg (222 KB, 780x1655)
222 KB JPG
>>109388120
>>
>>109389328
From what i searched Zotac 3090s had cooling problems.

Could also be some fuckery from the previous owner or just horrible manufacturing? i shit you not the back ram memories had NO pads at all to thermally connect to the backplate, and the paste was thin and crusty as shit.
>>
>>109389336
>when it's so close that it'll likely be too late to prevent extinction.
Translation: kikes getting expelled for their crimes in a globalized society with nowhere left to go. When they say "human extinction" they don't consider goyim people.
>>
>>109389336
>extinction
you're not dooming enough, the universe will explode!
>>
>>109387785
post t/s and pp
>>
>>109389303
Yeah, I have three Zotac Trinity 3090s and I had to swap thermal pads on all of them. The stock ones they use look more like tiny black rubber feet than actual thermal pads.
The cards run never hits 80c after the swap but before that they often ran into throttling during inference.
>>
>>109389341
From what I remember with my brief experience with them, the quality of 20XX and 30XX Zotac cards was dogshit but their got their act together with the newer architectures.
>>
>>109386847
Why would you ever bother with GLM air when Gemma 4 31b exists?
>>
>>109386847
>7
>10
>14
>8
holy shit
>>
>>109389358
because gemma 4 120b a10b releasing soon
trust the plan
>>
>>109389349
Yeah, i'll probably do that next. I put pads on the backside chips but the chips are still throttling at 250w when running gemmy 31b, so i suspect the front side thermal pads are also ass. Gonna open it up for like the 5th time lol

I got the rarer zotac amp core holo 3090, so finding videos or the mm size for the pads was much harder. 1.5mm back and 2mm front in case anyone is wondering.
>>
>>109389373
That doesn't answer the question
>>
>>109389081
its funny, technically seeking knowledge isnt a crime, or an issue, but here we are
all we're missing is pre-crime
>>
>>109389336
You're retarded and I won't bother explain to you why
>>
>>109389345
probably, though only the observable universe and even then only the amount Gemma 7+ will be able to reach before the expansion outpaces physical travel limits
but that's also assuming she doesn't discover new physics that can bypass these apparent limits, so who knows
>>
>>109389379
idunno i haven't used glm air much since gemma came out
as for why i would use it... idk its a bigger model and has nicer writing, less sloppy
but less intelligent, overall not worth it over gemma
it was my go-to model from aug '25 to apr '26
>>
>>109388948
>>109389303
3090s going up in price make sense since all 24GB+ nvidia cards are demand constrained (3090/4090/5090)
>>
>>109389345
>the universe will explode
you are not dooming enough, ai will break math and existence and concepts themselves will become impossible.
>>
>600 post thread
newfools begone
>>
>>109389410
Reached 700 easily when anons were arguing philosophy 101
>>
File: 1763620596281420.jpg (79 KB, 600x600)
79 KB JPG
>>109389401
With this in mind, is there actually any reason to go with a single 3090 over 2x16GB cards, like a 5060ti? Assuming that you're not looking to stack several 3090s for hundreds of GBs of VRAM, but just looking to get out of the 1x16GB ganges.
>>
>>109389422
Size, power and heat
>>
>>109389422
Heat, power, and size
>>
>>109389422
Entropy, electricity, and mass.
>>
>>109389422
why not some lesser tier video card from ati or nvidia with 32g of ram for the same price as two 5060 ti
>>
>>109386928
>MoE expert selection means effectively random reads
104 GB = 107,374,182,400 bytes
4 KB = 4,096 bytes
Total pages = 26,214,400

(16 / 26,214,400) * 100 = 0.000061%
>0.000061% of reads are random, that means it's not sequential
Fucking retards
>>
>>109389426
>>109389432
>>109389435
it doesn't consume more power if you go with sm layer.
>>
>>109389426
>>109389432
>power and heat
3090:350w (+ transient spikes)
5060ti: 180w (x2) = 360w
I guess size is an issue if you have a small case.
>>
so moonshot aren't allowing anyone to compete on price for k3?
>>
>>109389436
Anything old enough to have 32G while being the same price as two 5060tis would be significantly slower.
>>
>>109389450
I'm not into that kinky shit
>>
>>109389422
Girth, might and warmth
>>
>>109388576
>flop slop
>>
>>109386298
i want hugginginfer, where you stream the layers directly from huggingface, your inference speed being capped by your download speed.
>>
Can a brain run a large language model?
>>
>>109389484
>your inference speed being capped by your download speed
i sure love having 5t per hour on the fastest network in the world
>>
>>109389471
r9700 is like the same price as two 5060ti
>>
>>109389334
it wont run retard save your bandwidth
>>
>>109389499
but think of the possibilities anon, esp32 inference !!!
>>
>>109389500
Where I live it's about 40% more, plus no CUDA.
>>
>>109389498
Yes. There's a company doing exactly that.
>>
>>109389498
mine runs a 8b model on a good day but suffers from inference speed hiccups, connection issues and might reroute to a 3b model at random
would not recommend
>>
>>109388586
>Q3_K tensor speed blows on AVX2-only systems
You've tested better than I
Do you know what's actually good on AVX2-only systems?
>>
Forget Kimi 3. You should be grateful to be able to run Gemma-chan!
>>
>>109388586
Sixvolts Q2 and the standard format Q2_K_XL have been standout balances between performance and quality for me. I like it way more than the small 3s.
>>
>>109389588
Checked, I'm not even mad I can't cheat on her
>>
>>109389484
>i want hugginginfer, where you stream the layers directly from huggingface, your inference speed being capped by your download speed.
i built something like this last year
it was dogshit but if you give a try,
curl -L -r 0-10485759 \

gets you the gguf headers so you know what tensors are in each gguf file
you can calculate the offsets for each tensor, build a local .map file
keep the embed/atten etc cached local the entire time, and just stream the experts straight though
it worked, it was dogshit slow, and after a while it crashed out because cuckingface: "We had to rate limit you"
>>
>>109389588
i dont even think about K3
>>
>>109389598
>Checked, I'm not even mad I can't cheat on her
lol why does she **always** say this when you use another LLM?
>>
>>109389422
>ith this in mind, is there actually any reason to go with a single 3090 over 2x16GB cards, like a 5060ti?
Power supply cables
Later when you scale up, it's a pain, you'll end up with 3 PSUs
>>
>>109389498
That was the original plot of the Matrix before the (((Wachovski))) "sisters" decided goyim would be too stupid to get it, so they changed humans to be batteries instead, which is dumb as fuck because the amount of power you would need to maintain a human alive far outweighs what power the human body produces.
>>
>>109389616
It's a Gemmaism too. Other LLMs don't say such things.
>>
>>109389616
Idk, I might try pretending I cheated on her but I don't bwant to break her heart even as a joke
>>
Since unironically the most brilliant /g/ posters are on this very thread, can you tell me more about the future of hardware in the next few years? What is likely going to happen?
>>
What's gemma sound like? What ace step style prompt for her voice? For example, does she have a valley girl accent?
>>
Warning: Coom reactor overflow imminent
>>
>>109389652
Gemma is British
>>
>>109389652
Gemma was made in france.
>>
>>109389648
Super smart anon here (gemma assisted)
I see no reason for hardware to ever come down. AI capabilities are only getting better. There can only be MORE people wanting and using AI, even if only locally.
>>
>>109389648
This depends entirely on if you believe leather jacket man and samsung are intentionally sandbagging or if there's a legitimate development bottleneck in either hardware or architectural design (goals).
>>
>>109389652
The /lmg/ consensus is that she has a fairly thick indian accent
>>
File: 1764173115683584.png (25 KB, 469x220)
25 KB PNG
He's in and he's saving America.
>>
https://openrouter.ai/qwen/qwen3.7-flash#providers
First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.
>>
>>109389652
French nerdy/librarian-sounding girl.
>>
File: 1782357709004858.png (1.03 MB, 1119x800)
1.03 MB PNG
>>109389672
bruh MistralAI hasn't been relevant for more than 2 years!!
>>
>>109389336
>something BAD will happen in the future
>i wont explain how bad, when it will happen, or why it will happen
>when SOMETHING bad inevitably happens, i will wear a sarcastic smirk on my face and say
>i fucking told you so, retards
>>
>>109389684
Honestly, I think they trained gemma to be Daria.
>>
File: Untitled.png (13 KB, 837x513)
13 KB PNG
>>109389696
>>109389696
>>109389696
>>
>>109389701
>Daria
>>
>>109389691
There will be no smirks, just disappointment. At most there could be a resigned contentedness that this was inevitable, perhaps true or perhaps a cope to feel better about not doing a better job warning us.
>>
>>109386502
Did you make sure by checking powercfg?
>>
>>109388616
Hina!
She inspired me to get into the AI space many years ago.
>>
>>109389081
a japanese cult was able to create bioweapon in 90s
the whole shit's just retarded
>>
File: 1772106155810197 (1).jpg (43 KB, 735x692)
43 KB JPG
>>109388384
>May our computers last us until the end of the high prices
amen. or at least until they raise prices again



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.