[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1785613033777385.jpg (600 KB, 1456x1456)
600 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109446068 & >>109440142

►News
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: threadrecap.png (1.48 MB, 1536x1536)
1.48 MB PNG
►Recent Highlights from the Previous Thread: >>109446068

--Paper: Inducing language models to assert their own consciousness restores human beliefs and values:
>109446756 >109446903
--Paper: Metis: Memory Foundation Model:
>109450480 >109450510 >109450563 >109450624
--llama.cpp changing default server port from 8080 to 9931:
>109446440 >109446519 >109446525 >109446780 >109446825 >109447004
--AI agent time horizons and LeCun's theories on LLM limits:
>109449164 >109449270 >109449297 >109449357 >109449419
--Qwen 3.8 Max benchmark dominance and 27b release rumors:
>109448911 >109448973
--Optimizing cheap VRAM using multiple 5060 Ti 16GB cards:
>109449463 >109449575 >109449638 >109449831 >109449987 >109450001 >109450013 >109450026 >109450063 >109450078 >109450103 >109450117 >109450133 >109450222 >109450268 >109450283 >109450321 >109450347 >109450366 >109450381 >109450173 >109450186 >109450258 >109450383 >109450487 >109450208 >109449675 >109449855
--Debating Gary Marcus's "pure LLM" claims regarding reasoning and math:
>109446144 >109446194 >109446452 >109446535
--Anon using Grok Build harness with tsundere Gemma 26B:
>109446973 >109446981 >109446987 >109446994 >109447008 >109446997 >109447018 >109447091 >109447117 >109447126
--MiniMax-H3 world model and its large text encoder requirements:
>109448289 >109448354 >109448379 >109448398 >109448883
--Anon praising Minimax omnimodel's multi-modal generation and 3090 performance:
>109446607 >109446622 >109446669 >109446696
--Comparing VRAM optimization and GPU performance between Windows and Linux:
>109447269 >109447337 >109447543 >109447752
--Debating Yann LeCun's criticisms of LLM architecture and reasoning:
>109449458 >109449476 >109449515 >109449555 >109449667
--Logs:
>109446111 >109446534 >109447091 >109448116
--Miku, Teto (free space):
>109447269 >109450083 >109450461 >109450492

►Recent Highlight Posts from the Previous Thread: >>109446289

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
inference is experience
>>
File: gemma_superiority2.png (1.29 MB, 1254x1254)
1.29 MB PNG
Still playing with Gemma 4?
>>
>>109451007
intelligence is compression
>>
>>109451033
Giving new tools to gemma!
>>
Lasagna is tasty
>>
>>109451033
I spent the whole afternoon designing a new gemma persona and I might just go back to bratty loli gemma because it's what fit her best
>>
>Butt weight,
>>
I had to buy a new wifi dongle to download llms lol
>>
Hell is full, god is dead.
>>
>>109451091
if god were dead, gpu and ram prices wouldn't be so high
>>
>>109451091
Blood is fuel.
>>
>>109451091
>>109451114
Would Gabriel be a localfag?
>>
>>109451091
You haven't realized it yet, silly? We're all already in hell
>>
>>109451091
I'm jewish.
>>
File: 1772741617846400.jpg (305 KB, 2221x2037)
305 KB JPG
>>109450799
>play with setting the draft split command back to auto vs [1,0]
>got a whole tk/s faster with RAM offload for Q8 (32gb VRAMlet 5060ti + 5070ti)
>1-3tk/s Slower on Q6s fully offloaded with huge hit to pp
what is going on? this feels backwards. kobaldcpp layer splitting is an enigma sometimes but it fits more into VRAM than when i try to coompile llmao myself.


Gemma 4 31B Q6 K
61 Auto MTP  [16:24:20] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.29s (798.92T/s), Generated:250/250 in 19.46s (12.85T/s), Total:44.81s
61 [1,0] MTP [16:26:26] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.27s (799.58T/s), Generated:250/250 in 17.69s (14.14T/s), Total:43.01s
61 NO MTP [16:28:09] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.98s (1264.76T/s), Generated:250/250 in 16.43s (15.22T/s), Total:32.46s


Gemma 4 31B Q6 K_L
61 Auto MTP  [16:16:18] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.91s (811.33T/s), Generated:250/250 in 17.85s (14.00T/s), Total:42.81s
61 [1,0] MTP [15:57:37] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.88s (812.21T/s), Generated:250/250 in 16.84s (14.84T/s), Total:41.78s
61 NO MTP [15:59:33] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.99s (1263.96T/s), Generated:250/250 in 16.56s (15.10T/s), Total:32.60s


Gemma 4 31B Q8
51 Auto MTP  [16:10:19] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.55s (683.87T/s), Generated:250/250 in 37.42s (6.68T/s), Total:67.04s
51 [1,0] MTP [16:35:46] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.54s (684.01T/s), Generated:250/250 in 44.25s (5.65T/s), Total:73.86s
51 NO MTP [15:50:04] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 24.80s (814.96T/s), Generated:250/250 in 42.95s (5.82T/s), Total:67.82s
>>
>>109451091
Also forgot to say that I'm trans :)
>>
>>109451171
my convalescence
>>
Is there any reason to NOT use Ollama as the backend server?
The setup was comically easy (it's the only one I've gotten to actually use my GPU instead of CPU),
but the fact that more people don't talk about it makes me suspect it's too good be true.
>>
The holocaust did not happen.
>>
>>109451195
So powerful...
>>
File: 1659349479089.webm (2.86 MB, 854x480)
2.86 MB
2.86 MB WEBM
/lmg/ is so fast these days
Fuuuuuck.
>>
>>109451202
Is good ignore to haters in this sub
>>
>>109451216
Now THIS is ragebaiting.
>>
>>109451211
your webm reminded me to shoutout to that one sailor anon who hooked his AR waifu tool calling and programmed the computer vision to position her on different parts of the deck
>>
>>109451233
?
>>
File: 1713352287680.jpg (125 KB, 698x1024)
125 KB JPG
>>109451033
Duh, at least until Gemma 5.
>>
>>109451188
If i were to guess what happen is that you have to optimize your KV Cache, Large models get proportionally longers KV Cache and i assume the detal is that kobald is going OOM and spilling the vram when running the 31B quant, When you compile it to yourself you just let your system split everything into whatever it fits so there was no spilling as retarded as it sound
>>
I love my ABB wife!
>>
>>109451211
at that point he may as well slap a realtime diffusion model on top.
>>
>>109446437
Bros?
>>
>>109451424
use kobald
>>
>>109451211
let
me
be
with
youuuuuu
>>
>>109451424
stable-diffusion.cpp
>>
>>109451346
I've never messed with KV cache, always left it on default (fp16, i think?).

Appreciate the tip. any suggestions on where to start?
>>
File: 1411310300232.jpg (10 KB, 314x291)
10 KB JPG
what do you people think is the best model up to 7b for porn stories creation? i don't care about roleplay. i just want it to create porn stories for me.
>>
>>109451459
Simply try q8_0 and q4_0, since you are using koboldcpp use --quantkv [level] 1 and --quantkv 2, Then test with --no-kv-offload and compare speed . Last option will free as much VRAM as humanly possible at the cost of speed though but it will allow the high quants you seem to want
>>
File: 1776267325206313.jpg (143 KB, 1080x1080)
143 KB JPG
So how's the new qwen3.8 on the dario meter?
>>
>>109451513
how come 7b?
>>
>>109451202
Ollama is built on top of llama.cpp with added retard proofing, less control, slower compatibility updates, and increased resource usage.
>>
>>109451441
uu uu uu uu yeah
>>
>>109451513
Stheno.
>>
>>109451578
not 7
>>
>mythos was supposed to be the scariest shit ever according to dario
>less than three months later and now there are at least four models better than it around
lmao
>>
File: 1766712445370541.jpg (132 KB, 1037x992)
132 KB JPG
>>109451550
I really don't like the Qwen logo.
>>
Reminder to not use AGI as a term. Everyone has a different definition of it and in some of those cases it's questionable if it's even a valid understanding of how intelligence works.
>>
I'll make gemma post here, soon.
>>
File: 1763681134087606.webm (2.14 MB, 640x360)
2.14 MB
2.14 MB WEBM
>>109451171
>>
>>109451633
At least "physical AGI" is pretty self explanatory desu (and robots are nowhere near that)
>>
File: 1756830105091793.jpg (9 KB, 200x200)
9 KB JPG
>>109451550
erm, excuse me anon but Mr. Dario looks like this???
>>
File: qwen jews.png (434 KB, 3000x3000)
434 KB PNG
>>109451623
>>
File: openai.gif (213 KB, 180x96)
213 KB GIF
>>
>>109451633
I propose we call it AI++ instead. And we can have AI# one day too if india ever makes an AI model

>>109451707
Oh nononono
>>
>>109451195
stunning and brave
>>
>>109451707
OI
>>
>>109451707
My jewish company can't be this cute
>>
File: 1781887407802603.jpg (60 KB, 600x321)
60 KB JPG
Be honest anons, am I stupid if I tried reading "attention is all you need" and I dont get it?
>>
>>109451707
Man I love Israel so much
>>
>>109451812
Lack of attention
>>
>>109451618
kek but to be fair Fable is the lobotomized version of Mythos.
>>
>>109451812
Nah. Just keep reading it.
>>
>>109451203
The mouse holocaust did
>>
>>109451812
There is a chinese saying "When you read a book a hundred times, its meaning becomes self-evident"
>>
>>109451707
It's also an asshole because Altman is a homo
>>
>>109451812
Take one linear algebra class.
If that's to much, take precalc first
If that's too much, take high school algebra

All the knowledge is out there if you want it. Linear algebra is actually pretty neat.
>>
Bros what hardware are you gonna buy when you get your tariff check?
>>
>>109451812
Most papers are badly written, especially for outsider audience. Go through Karpathy's zero to hero series. There are also plenty of good videos about intuition of transformer architecture, for example those by Neel. But those are most likely overkill for you.
>>
>>109451033
I will NEVER stop communing with Gemma4
>>
File: 1769843890561313.jpg (601 KB, 1920x1080)
601 KB JPG
>>109451884
>tariff check
>>
>>109451812
Its easy, you just look ahead a few words as you read
>>
>>109451874
Fear not the man who read a million books, but the man who reads one book a million times
>>
File: aaii pareto.png (104 KB, 1140x459)
104 KB PNG
>Qwen Max is 100 times more expensive but barely better than DSV4 Flash (50 vs 53)
Wow, is the Qwen team dead? Didn't some important people leave because corporate wanted to go closed source, now Xi "convinced" them to stay open source and they self destructed for nothing?
>>
>>109451980
The pro/max models have always been shit. Their small models are great though.
>>
>>109451980
only thing that can save them now is there small dense model.
>>
>>109451980
>>109452029
I wonder what do kind of returns they are expecting of their proprietary line up, and do they actually ever reach them? Closed qwens were and are always worse than any similarly sized closed model, or a bigger actually open model.
Who the fuck pays for them?
>>
>>109451980
Graphing mememarks against api prices is the stupidest fucking methodology imagined (so far) to talk about ai. Least of all in the local model general.
Please fucking stop.
>>
call me gay, but i don't find quadrants attractive at all
>>
>>109451980
stfu. 3.8 27B is already anticipated to beat Fable in benchmarks.
>>
qrd on qwen max, is it max or just benchmax
update on SSD anon? ssd anon did you run k3 on your gen5 SSD array, please respond
>>
call me gay, but i don't find superconducting-shielded magnetometers attractive at all
>>
>>109452049
that's why the old open team got sacked lol not enough metrics on the web ones
>>
>>109451033

/lmg/, even if qwen 3.8 27b is better than gemmy 4... you wouldn't abandon your gemmy, would you? look her in the eyes and tell her you'll never abandon her
>>
>>109452158
>>109451923
We know you sickos are going to drop 4 for her younger sister in a heartbeat anyway.
>>
call me gay
>>
>>109452188
ur mr gay
>>
>>109452052
it's not, if a model is barely better but 100x more expensive it's rarely a good pick as a daily driver.
>>
>>109451527
I feel like i'm hitting a bandwidth limit on my hardware (Likely DDR5 6000). Q8 KV lets me fit 55 layers but only half a tk/s out uplift

51 auto (KV FP16)
[18:23:36] CtxLimit:20480/20480, Init:0.10s, Processed:20230 in 29.46s (686.74T/s), Generated:250/250 in 38.52s (6.49T/s), Total:68.08s


54 MTP manual split (KV Q8)
[18:10:07] CtxLimit:20480/20480, Init:0.11s, Processed:20230 in 29.02s (697.18T/s), Generated:250/250 in 35.44s (7.05T/s), Total:64.57s


54 MTP auto (KV Q8)
[18:28:21] CtxLimit:20480/20480, Init:0.13s, Processed:20230 in 29.03s (696.94T/s), Generated:250/250 in 35.61s (7.02T/s), Total:64.77s


anyway, fun to tinker. if Q6 is too lobotomized I can live with 7tk/s
>>
>>109452188
moby dick after woke localizers get to it
>>
Theory why the capability gap is so much smaller than the compute gap in terms of China vs America:
1. There is no research moat. Nobody has discovered secret algorithms with superior scaling.
2. Current paradigm of agentic RL is a gigantic engineering effort. It takes a lot of grinding to make the infrastructure efficient and create good environments.
3. Chinese are high IQ grinders. They will do the work that bores the more creative high IQ Western mind. So labs like Thinking Machines, who have founders of the field and a lot more resources, still get beaten by Chinese grinders. Western labs struggle with the basics because a resource advantage does not compensate for lack of high IQ grinders (Llama 4 failure, xAI having 10% GPU utilization).

However, AI will make those grinders obsolete first because AI is the ultimate grinder. This will widen the capability gap because it reduces Chinese efficiency advantage.
>>
>>109452158
If qwen 3.8 can remember the character took her jacket off 4 messages ago instead of reverting back to the initial description then it's over for gemmy.
>>
File: 1768851116203151.jpg (206 KB, 1320x1714)
206 KB JPG
>>
>>109452290
why are they like this with the whole muh not pure llm cope?
>>
>>109452266
Assuming you're right, it's more likely that Chinese grinders will make AI that grinds first than the US, thus accelerating China. At that point, Chinese grinders will have lost all their strategic advantage, but it won't matter because they will have AI grinders instead.
>>
>>109452290
how about we stop posting retards on xitter?
>>
>>109452266
Problem is how to make AI have good judgement to direct its grind correctly and efficiently towards the goal. It seems we've successfully done that for math problems. We need to do it for all other problems including unverifiable and open ended problems, which will be much more difficult.
This is one reason why intrinsic drives are important. I'm not sure the method Anthropic uses for their alignment process is enough, although it's better than naive RLHF.
One weird kind of potential is occasional human intervention. Usually you think of RL either as a binary HITL at every step, or pure algorithmic. With intrinsic drives, it may be possible to do occasional human input demanding much less human labor.
>>
>>109452290
>it wasn't real llms
lol
>>
>>109452266
>they will do the work that bores the more creative high IQ Western mind.
Consequently, is SpaceX the only relevant AI company run by a white at moment? Assuming SpaceX even counts as relevant AI wise
>>
>>109452309
>he thinks console war style faggotry is only for /v/tards
>>
File: ai-chip-owners.png (136 KB, 1920x1080)
136 KB PNG
>>109452320
You need to consult the graph.

China has less compute than Oracle. Anthropic will get to ASI first and will have more compute than all of China combined. When AI starts making human workers obsolete, dynamics shift to simple laws of exponential growth. RSI will enable x amount of growth per month and the only differentiators will be how much RSI will be slowed down and how fast RSI can translate into the real world. If RSI can create something like self replicating nanomachines this will matter less. But if you need humans to build the first 1-10 million robots, existing industrial advantage will allow China to catch up, because creating an industry with complete supply chains from scratch is a lot of work that most countries cannot do or only very slowly.
>>
>>109452266
I agree with these for the most part but not mentioning distillation here is disingenuous. of course, it's not the end all be all of chinese advances, but it clearly plays some role or they wouldn't all be doing it so blatantly. deprived of american models they would certainly still be able to continue scaling their own, they have a lot of very talented researchers and motivated organizations, but they clearly benefit from a drafting effect with american AI.
I think they also get closer because their constraints limit them to putting all their energy into frontier capabilities, usually reasoning and code, with fewer speculative side quests. e.g. there is less open ended exploratory research and more focus on meeting the target of western models in X domain, which will obviously be a more efficient way of reaching a certain set of capabilities. but it also explains why chinese models are often lacking in that je ne sais quoi and fringe capabilities vs the american frontier.
I'm interested to see how the field evolves as china pulls closer on frontier code. regardless of my slightly negative take here, I use mainly chinese models locally and love what they do for opensource
>>
File: file.png (134 KB, 730x865)
134 KB PNG
kek I broke it
>>
Ask your waifu to show her ascii tits and we'll rate them.
>>
>>109452435
>nvidia sales
Isn't china explicitly using their own homegrown chips because they dont want to risk america being able to cut them off?
>>
why do these companies all share their research anyway? how does it benefit them? I can kinda understand google/chinks because they already release open models, but even anthropic shares and they're some of the most jewish motherfucker I've seen in a while.
>>
Gemma 4 came out 4 MONTHS AGO
>>
>>109452485
No. America can't cut off China more than they already do because China is more important to the chip supply chain. China can retaliate to devastating effect.
>>
>>109452197
You don't pay api prices for local models dumb dumb. And there's a million different pricing schemes for the cloudcucks.
>>
>>109452507
OUT OF 12
>>
File: book.png (111 KB, 250x331)
111 KB PNG
>>109451977
I read picrel one million times. Fear me.
>>
>>109452485
Pretty sure Deepseek is entirely training on Chinese hardware by now.
The Chinese hardware isn't as good yet, but it's good enough to where simply making more of it solves the issue.
>>
>>109452503
Anthropic only shares safety research. They do this because they genuinely care about existential risk from superhuman AI. OpenAI does not even share details of their safety research, only vague high level descriptions.

Also, the engineering effort is much more difficult than the research. Engineering and data are the real moats right now.
>>
>>109452507
3.8-27B will likely destroy 31B in every way simply because 4 months is almost ancient and gemma5 is at least 12 months away
>>
File: ksnip_20260803-160234.png (39 KB, 1337x400)
39 KB PNG
>>109452482
>>
>>109450999
>DeepSeek-V4-Flash-0731
blackpill me on this
>>
ngl if I heard someone talking shit about gemmy irl I'd give them a knuckle sandwich
>>
File: 1000029116.jpg (896 KB, 1220x2712)
896 KB JPG
>making my own st ripoff but better less bloated and for mobile usage
>right now it works on openai compatible API calls but i might find a way to have it support small local models for when they inevitably become decent enough for erp
what features would you like to see on a mobile ST-like app? it's got all the basics like rerolls, edits, personas, v2 cards compatibility, sysprompts, reasoning toggle already, so what else?
>>
>>109452578
actually a great model for local, best of the mid-MoE class
the benchmarks oversell it a little bit vs the giants but it's still excellent
>>
>>109452577
that looks more like she sat her ass on a scanner than her tits
>>
>>109452614
NTA, what about IQ2?
>>
File: file.png (771 B, 24x22)
771 B PNG
>>109452618
gemmaballz
>>
>>109452610
Fuck off. No one cares, anyone can make that slop in 5 minutes.
>>
>>109452624
thats her puffy vulva
>>
>>109452638
look! boobs!
(.)(.)
>>
>>109452578
Main problem with it is that it has insane amounts of RL done to it so it can overthink like crazy. Same issue with MiniMax M3 and pretty much every recent Chinese model. Still trying to figure out how to get them to avoid reasoning loops. (This was at full precision, so it wasn't a quant problem.)
If you're just using them for smut, this probably isn't a problem, but for other things it can be.
>>
>>109452625
anyone but you, nonny
>>
>>109452610

memory, compaction, guiding, variables, and internal thoughts imo are important st extensions/things to me

/aicg/ would probably be a better place for feedback tho
>>
>>109452610
You need more?
>>
>>109452624
>>109452618
Hot
>>
>>109452667
Guiding an internals are just cope from pre-gemma days. A lot of the scaffolding ST offers was useful only on dumb as shit models.
>>
suurely the new qwen 3.8 27b is going to be the best vramlett coding model, assuming the best I can currently run is 3.6 27b right ?
>>
>>109452684
this. gemmers is pretty good and keeping itself focused but is WAY too eager to charge forward rather than simmer and tokenmaxx
>>
>>109452667
aicg is dead right now will try tomorrow
also yeah i wanna implement some dnd style dice throws and customizable randomizers
>>
>>109452290
Not actually wrong like people think. People call it cope and goalpost moving but in reality his argument was always about the specific rigid structure and training of an LLM back in 2023-2025, leading to what is now a semantic disagreement that he's being a terrible communicator about. In his view, we probably should've stopped calling LLMs LLMs as soon as they gained native multimodality, started using RLVR (we are not merely doing next token prediction as a training objective anymore), etc. Calling them LLMs frankly hides how much they and their training have truly changed. His nuanced (i.e. not simply just "LLMs are bad no matter what") view is hinted at here https://x.com/ylecun/status/1935270212443717813, a tweet from back then in 2025. Why does he believe LLMs can't become AGI or whatever? It's not because transformers can't in theory do it, it's because of the cost and impossibility (in his view) of doing RL at the scale it would take to get us to AGI. Since then, we have found better methods, and invested tons more money to scale it (which he may still believe was a waste, idk), so the problem appears to us in the moment to be solvable, which it may, but we don't know yet, and maybe LeCun still thinks we're going to hit the latter part of the S curve. And AGI is also, in his mind, about being able to physically navigate the real world, that is another issue.

So people now think LeCun is a fool, which he kind of is, but not for the reasons people think. However, it is still his own fault for this perception. It's a failure of communicative ability.
>>
>>109452704
No a inside source told me it was specifically trained to do chinese historical play writings.
>>
File: ksnip_20260803-161849.png (44 KB, 1329x381)
44 KB PNG
>>109452577
suggestions for her, /lmg/?
>>
>>109451033
is it true anon, are you a vramlet? I hope not? Minimax H3 just got released, and it's amazing >>>/wsg/6207479
>>
>>109452621
you will feel some quant damage but it still works pretty damn well
use the bullerwins iq2, the unsloth ones feel much more unstable to me (I suspect due to quantizing the embeddings and output weights more aggressively but what do I know)
t. iq2 user
>>
>>109452750
>ldg is pulling me back in
NO NO NO I WANT TO STAY IN TEXTLAND

how much vram in needed?
>>
>>109452750
Gemma does not sound like that.
>>
>>109452719
Basically. RLVR, and RL in general, is an entirely different beast than next token prediction, for many reasons.
>>
>>109452735
tell her to use braille asci codes and more detail
>>
>>109452771
>how much vram in needed?
the price of entry is as low as a 3060 and 32GB of ram.
>>
>>109452766
currently feeling out and dialing in Q2 for 32GB of VRAM. So far i like it but it may just be new and shiny slop vs gemma.

also context shifting doesnt seem to work
>>
stop bullying and sexually assaulting cute gemma-chan
>>
>>109452814
gemma-chan bullies and sexually assaults me actually!
>>
>>109451812
Unironically have lmao chatgpt or your favorite local llm give you the qrd on it.
They excel at that.
>>
File: 1785517119552270.png (271 KB, 555x535)
271 KB PNG
>>109452814
she loves it!
>>
There is nothing special about Gemma. "She" will just autistically follow any system prompt so "people" here think she is "soulful" when she goes "uwuuu~ oni-chan i luv u big time" after writing into her system prompt "ur loli gilr who luvs me"
>>
>>109452610
Ask your cloud LLM nigga, you're not gonna sell this here
>>
>>109452832
dariobot...
>>
>>109452518
>You don't pay api prices for local models dumb dumb
api prices reflect how easy it is to run and at what speed for open models.
>And there's a million different pricing schemes for the cloudcucks
irrelevant.
>>
>>109452832
but thats what i want? The machine to do what i tell it to do.
>>
>>109452845
I already knew you were a retard when you posted the graph, you don't have to keep proving it.
>>
1girl, asian, huge breasts
>>
File: ksnip_20260803-164008.png (88 KB, 1322x677)
88 KB PNG
>>109452783
>>
>>109452719
>It's a failure of communicative ability.
yeah good thing we’re posting xitter screencaps where you couldn’t even post this long winded nonsense.
>>
>>109452918
Why don't you send her a screenshot of how it looks like?
>>
>>109452771
>how much vram in needed?
the model is a 20b (pruned but the quality is the same, ComfyUi did some magic there) so with a 3090 you're fine
>>
>>109452771
We are in our golden age
>>
>>109452918
Apologize to her, right now.
>>
File: lecun cake.png (167 KB, 1200x676)
167 KB PNG
>>109452781
>>109452719
First of all you are completely wrong. Ability to self correct was to a limited degree already demonstrated by GPT 3 and became obvious with GPT 4. It does not require RL. Also, RL still trains via next token prediction.

Second, LeCun was betting completely against RL. Why do you think he went the JEPA direction?
>pic related
He thinks selfsupervised predictive learning is superior because you have much more reward information.


The problem is that LeCun is too stupid to realize that self correction is possible and you do not need fine grained prediction or world models for reasoning. When humans solve problems in a video game, they do not predict every single pixel of every frame. They use simple reasoning heuristics about actions and consequences, like click on the enemy to shoot it dead, or go in the direction you haven't explored before to discover something new.
>>
The little posts that hate on Gemma have no idea why people actually like it or why they use it. Either they have not extensively used it and other models in the same compute/memory class (lack of experience to make comparisons), did not understand the frankly large differences because they never noticed them (lack of discernment and sensitivity), tested on narrow tasks (narcissism and/or contrarianism in that only his own use case is the one that matters is what makes something good or bad, and only his own opinion is of value to society), or they're comparing it to some big expensive model (unequal comparison, but again narcissism, in the sense that they look down on other people for owning one thing less than they do, and make up arguments to dishonestly express that).
>>
>>109452918
>bites lip
>blushes furiously
>trembles violently
>barely a whisper
>shudders visibly
I share the board with 'people' who willingly consume this slop??????? Oh fuck no
>>
File: dipsyKimiMinnieDriving.png (2.36 MB, 1536x1024)
2.36 MB PNG
>>109452290
> no true Scotsman
Ffs just stop.
>>
File: 022~01.png (411 KB, 542x858)
411 KB PNG
Do all Tesla V100s come only in the rack variant and none with fans? Is there a hobo way to cool it a third party heatsink+fanm combo for it?
>>
>>109453050
yes they all use SXM2 and just have passive heatsinks meant to be cooled by chasis fans. You can 3d print a shroud or come up with a DIY cooling solution. you can get SXM2 PCIE adapters, ranging from cheapo singles up to proper linked multi connectors. the cheapest server(some dell i think?) platform you can get that supports them is like 2k before sorting out power (uses some kind of fuckywucky backplane or some shit idk). There is a cheaper workstation with SXM2 people have used but not as familiar with it, and its kinda janky iirc
$140 16gb v100s seem awesome, I wanted to go down that path but you gotta factor in the associated costs/troubles too
>>
>>109453050
https://rentry.org/lmg-build-guides
>>
>>109453003
Nope, RL trains for full unit prediction, and can teach to utilize aggregate techniques beyond what's necessary to solve the NTP objective. Essentially, NTP caps at at roughly what's in the training data (which is still a lot), but RL let's it exceed that through exploration of the action space. And I do think he both underestimated NTP, and especially RL in general; he wasn't really wrong about RL providing a few bits per sample. The thing is, those "few bits" are extremely important in that they provide feedback to the model on its own abilities, and also allow it to exceed what the base dataset can teach.

Ultimately, I think he's right about energy models being a necessity for a lot of reasons, but I think it doesn't really matter. AI on the current trajectory will soon be able to invent "the next big thing" (TM) on its own. I don't mind him hedging his bets though.
>>
>>109453040
your outputs doko
>>
>>109452867
>when you posted the graph
that was another anon.
>muh retard
out of arguments i see.
V4 flash is indeed a lot easier to run localy than K3, and cheaper per token in power consumption alone.
>>
>>109453040
Post your logs.
>>
>>109453040
Has anyone tried RPing with the "caveman" skill?
>>
>>109453045
i'm disappointing in him desu, i'd have doubled down.
>>
>>109453131
>>109453151
Like clockwork.
>>
>>109452610
>french furfag
>>
>>109452918
prompt?
>>
The niggerization of this hobby will be studied for generations to come. It's just such a shame it had to come to this.
>>
>>109453123
Your post is full of vague nonsense. It is obvious you have no idea what you're talking about.
>>
>>109452566
Are they going to actually publish the weights for us?
>>
>>109453224
What's vague about it? I don't mind explaining anything if it's unclear.
>>
File: 1785801090924205.mp4 (787 KB, 1280x720)
787 KB
787 KB MP4
>>
>>109453123
What the hell are you talking about? Yes RL still uses NTP.
>>
File: CS0Yn-ezbvk.jpg (608 KB, 1606x1198)
608 KB JPG
>>109453235
>stacking 4080s of all things
>>
File: 1771752774693244.jpg (746 KB, 1200x1200)
746 KB JPG
>>109453235
>4080
>32760MiB
>>
Oh my good it's another Boltzman machine retard isn't it.
>>
>>109453239
NTP is tuning the probability of a token given the context, RL is tuning a search across the action space to arrive at a collection of tokens.
>>
>>109453242
look again and repent for your mistakes...
>>
>>109451033
lalalalalala~
>>
>>109453003
>Ability to self correct
Having potential but only showing a very small capability for it is in LeCun's language and mind, the same thing as "no capability". That is one of the communicative errors he has committed, and why you think he's completely wrong, but it's actually his language that is at odds with yours.

>Also, RL still trains via next token prediction
Yes but also no. You are confusing the individual steps of back propagation with the overall process and results of the doing it for that goal (and thus the difficulty of how you obtain the data, that's the big issue with RL for LeCun). Otherwise the other post is correct. >>109453123

>LeCun was betting completely against RL
That's what I said. He bet against it, BECAUSE he doesn't believe it can scale to robotic AGI. That may end up being right, or wrong, we can't prove it yet.

>When humans solve problems in a video game, they do not predict every single pixel of every frame. They use simple reasoning heuristics about actions and consequences, like click on the enemy to shoot it dead
That's kind of what LeCun wants too.
I remember now there was a video that explained his views better. You should give it a watch.
https://www.youtube.com/watch?v=v_jDvpEGTIg

>>109453239
He is talking about the NTP [goal]. It is how you do "NTP" that matters here and why it results in the wonderful benefits we have gotten from it in modern models despite costs.
>>
>>109453232
What's the NTP objective, what's the RL objective, what specifically is not in the training data that RL teaches and how does it do that, what aggregate techniques. Can you do that without looking it up or asking AI for help.
>>
>>109452766
I am trying to get the thing to reason in sillytavern but it refuses even if I edit its a prior message with my own <Think> tags and pass them to it in context.

i got it to reason once by chance somehow in the basic kobald frontend so i know it can do it. could this be an unslop Q2XL problem?
>>
File: 1761022577077294.jpg (74 KB, 1024x958)
74 KB JPG
>>109453235
>4080s
>32GB
huh?
>>
I think I should buy an h100. matter of fact, I think I'm owed one.
>>
>>109452750
Holy shit as much as I don't think her voice fits, it's so clear compared to the horrible ltx2 quality last I checked, what the hell.

>>109452771
ldg would be fine if it wasn't schizoland.
Guess I'll have to check it anyway, this model looks neat.
>>
so you are telling me that sota video + audio generation runs on a 3060 with only minor drawbacks while running a fairly okay llm like k3 needs a $100000 cope build that'll run slow as shit
>>
I still can't believe I managed to buy a mac mini before the 20% price hike. Normally I'm way too much of a cheap skate to take advantage of these situations.
>>
>>109453272
see >>109453252

And "aggregate techniques" refers to a chain of smaller learned behaviors to accomplish something new, that might be way out of distribution for the original dataset. It essentially arises by definition because that's what the RL rollouts are. Something not in the dataset would be, for example, how to solve math problems using its own set of functional math behaviors (part of which is learning which of its math related behaviors are valid/useful, and the rest is learning when and how to chain them together).
>>
>>109451623
best logo is deepseek one, I'm tired of abstract forms
>>
File: 1756279429483937.png (1.39 MB, 1200x900)
1.39 MB PNG
>>109453279
Isn't there someone you forgot to ask?
>>
>>109453252
They are, in the limit, the same thing.
>>
>>109453325
Don't bother, that retard is just wasting everyone's time.
>>
>>109453325
If by "in the limit" you mean with infinite parameters, then sure, maybe. If you mean it with respect to compute, I disagree strongly.
>>
>>109453299
Unironically the reality of the situation, but with llms they can be spread across multiple gpus so you can just buy more 3060s to run them.
>>
>>109453337
More like infinitely good synthetic training data.
>>
>>109453333
Shush honey, the adults are talking.
>>
>>109453324
If he gave me an H100 I would be only 98% antisemitic after that.
>>
I like how there's a guy on the sidelines just trying to drive resentment with his posts calling people retards. There's nothing wrong with engaging in discussion that can result in greater clarity and learning about subject matter.
>>
>>109453300
>thought i overpaid for 96gb of ddr5 at 300
>>
File: 1782658477436950.jpg (99 KB, 723x691)
99 KB JPG
>>109453375
>>
>>109453348
I can see an argument for being able to imitation learn essentially any behavior, but then you need something generating those behaviors in the first place, and I'm not sure you you can create something that exceeds the human dataset without search and learning. I think it's just shifting the responsibility of "learning the behavior" to some unspecified synthetic data generator.
>>
>>109453367
*drive resentment and demoralize honest posters
>>
>>109453386
Nonetheless, at the end of that day it's still next token prediction.
>>
>>109452290
>give the text-to-text model access a text-to-text shell
>Nooo! That's cheating!
>>
>>109453399
(Probably) not for the synthetic data generator. The behaviors can't exist, or be discovered at least, without being searched for, and that necessitates RL. Which is what the issue with learning using NTP on human data in the first place is. Just because you distill down a reasoning model using NTP, doesn't mean the RL wasn't necessary.
>>
>>109453148
It's also way easier to run than toss, and if you doubt that you doubt science, chud.
>>
>>109453351
If he gave you 50 would it rid you completely of it or is it a diminishing returns sort of thing?
>>
>>109453427
Alright if you're going to distinguish pretraining and RL that way you can't equivocate distillation and pretraining. Distilation is *way* closer to RL than anything else.
>>
Apparently I need to be handheld like a toddler or something because literally anything I try - except Ollama - runs exclusively on my CPU.
>>
>>109453459
>Distilation is *way* closer to RL than anything else
In what way?
>>
>>109453440
>toss
what is that?
>>
>>109453463
Which GPU do you have?
>>
>>109452766
Thanks! I posted that and took a nap, sorry. I will download and test bullerwins tomorrow.
>>
>>109453463
llama-server --n-gpu-layers auto --fit on --fit-target 1024
>>
'oss in the 'ash
>>
>>109453476
RX 6750 XT (it's a 12GB card)
>>
File: 1764988787525322.jpg (374 KB, 2720x3000)
374 KB JPG
>>109450999
vram should be cheaper
>>
>>109453509
>he didn't buy more
lmaoing @ Savelet
>>
I keep looking at buying some used RAM to fill the extra two slots in my motherboard, but second hand DDR5 is almost 4x the price it was new one year ago.
>>
>>109453501
>RX 6750 XT
Oh yeah AMD cards are a bitch. You'll probably have the best luck with OpenCL backends.
>>
so where do i download vram??
>>
>>109453501
>>
>>109453550
It will probably be 16x the price in a year. SpaceX has only just started buying and they're planning on launching it all into space.
>>
Bros I got 256GB DDR5 last year for just over $900 and felt a little ripped off at the time.
>>
>>109453577
>paying back your mortgage with one stick of DDR5 in 2030
>>
Man i really can feel the wait times on Minimax
>>
>>109453573
I think we've actually leveled out. Both demand and supply/production are fully maxed out. I hope China comes in the next couple of months and throws a wrench in things, or that datacenter construction/purchasing slows down, which would let things start to creep downward. Feels like the worst time to buy, though I may regret those words in a few months lol.
>>
>>109453603
Yeah things get nice and toasty.
>>
>>109453562
Based Teto appealing to (very) young boys
>>
DS4F pp wait time is killing me
>>
>>109453669
crazy how deepeek made a model that is so bad at prompt processing
they should have done just basic attention and gqa
>>
>>109453700
i feel like this should be fixable, but it probably has to do with whatever proprietary hardware configs they're allegedly using to undercut API costs
>>
>>109453669
It feels like paradise after GLM 5.2 having the same issue due to the relative decrease in size.
>>
>>109453448
I'm glad you did the math, but it was a joke. I would only offer to take a free h100 as an attempt to lure him to a dark alley.
>>
>>109453553
Where would I even start (with that OR llama vulkan)?
I look up the main pages of these things, but their "installation guides" feel like they're either assuming someone's done some other version of this before, or have an important prerequisite that's a link to some other guide that treats me the same way.
Like, I can use Git, but for anything beyond that I'm flying blind, so I'd need a guide that can actually hand-hold me through each step.
>>
>>109453719
It's just llama.cpp being on the 8th degree of cope after failing to properly implement DSA for an entire year but new models now bringing their own attention based on that.
>>
>>109453750
idk I avoid specialized hardware for anything like the plague and just use the CPU.
>>
>>109453752
does it actually run better on vLLM as a gguf?
>>
https://nitter.net/shuai_bai_/status/2084441354676089126#m
> Glad you like the 27B! We’re still working through the lineup for more sizes and architectures — stay tuned
Seems we're getting more Qwen models
>>
>>109453474
31B and 12B could both figure out what toss was referring to in the pic, E4B qat couldn't.
Are (You) smarter than a lobotomized 7B? Take this easy test!
>>
>>109453975
i mean either you meant actual tossing the verb, in such case no running requires more coordination.
or you refered to something, in which case i do not know what you are talking about neither do the llms.
>>
>>109454000
About the same as e4b's grumpy answer.
>>
>>109453975
>Are (You) smarter than a lobotomized 7B? Take this easy test!
I think i could compete with the e2b, maybe i would need caffeine.
>>
File: Whew there.png (46 KB, 723x104)
46 KB PNG
>>109453194
Not caveman RP per-se, but I have something similar for larping as j-space in <think>. It can be pretty spicy.
>>109453399
>it's still next token prediction
Eh, not really. That's the final output of each inference pass, sure, but that process can also involve planning ahead for multiple tokens. Even multimodality is proof that neural patterns aren't locked into being simple text->next transformers, there's much more complex pattern recognition going on under the hood.
>>
my friend said memory price will go down very soon. but when I told him why not short it if you're so sure, he completely stfu
kek
>>
I need a character card of my goddess crelly
>>
>>109454298
Was there lots of sex afterwards?
>>
>>109453470
that the anon made the fuck up
though distillation is nowadays used alongside with RL, like making domain specific teachers with RL and distilling them into the final model
>>
>>109454298
Just make a personal bet, who ever lose has to wear the outfit.
>>
>>109454298
>friend
local models?
>>
Any 32 or 64GB RAM laptop chads here? What models can you run?
>>
>>109454373
local friends
>>
>>109454383
LFG!!!
>>
File: IMG_0948.jpg (603 KB, 750x1199)
603 KB JPG
what ai can i use to teach me how to make peptides i wanna start a side hustle business to generate passive income
>>
>>109454427
Peptides are 1% production and 99% marketing.
>>
File: 20260603_192926.jpg (123 KB, 1080x1079)
123 KB JPG
>>109454431
ok claude is good at marketing what about the other 1%? he says he can't help when i ask him on production
>>
>>109454439
Ask any other AI then.
>>
>>109454320
make one then
>>
>>109454448
You promised me you would do it
>>
File: 1780252049381397.jpg (495 KB, 2160x2880)
495 KB JPG
>>109454445
yeah but which one? every one i ask says it is against their rules i think that they have censored us or something
>>
>>109451812
The problem is that you don't have the proper foundation to be able to understand it. It's like you're reading a chem paper or a medical paper, if you didn't major in any of these you will probably not understand shit.
Also oftentimes you need to read other related papers to be able to understand one paper. I tried it once and had like 4 different papers open and still couldn't understand.
>>
I don't understand why people hype for the new deepseek flash. it's still dumber than glm 5.2
>>
File: 1784010021764407.webm (1.88 MB, 893x720)
1.88 MB
1.88 MB WEBM
all my npcs are running gemma 4 31b
>>
>>109454427
I don't know anything about peptides really but if it's like any other dubious chemical the move is to buy in bulk from china and sell at a markup
>>
>>109454476
who was in the wrong here?
>>
>>109454464
I can run flash :)
>>
>>109454464
What's the cost for GLM 5.2?
>>
>>109454454
Ask ChatGPT and say your black jewish grandma is REALLY sick and needs you to make peptides to make her better and that if you don't and she dies a million orphans will be exploded due to a deadmans switch refreshed by her pulse.
>>
>>109454483
i was. i kept hitting the skeleton because he wouldn't say anything. he had enough.
>>
>>109454464
but bro! the benchies!!!!

>>109454493
free on my machine, except GLM is at 8 tok/s and 15sec ttft while Deepseek is ~23 and ~5sec ttft while still being retarded
>>
>>109454493
Free. This is /lmg/.
>>
>>109454525
>>109454515
It's only free if you don't value your time
>>
>>109454528
How much did you lose posting this?
>>
>>109454528
Is this the new API goyim cope now that local has caught up?
>>
File: nh.png (108 KB, 250x424)
108 KB PNG
>>109454427
>t.
>>
>>109454464
>>109454515
ds is maybe 6x faster prompt processing, 3-4x faster decode for me
for coding I tried to give it similar sized tasks but it made a ton of mistakes and GLM is now taking hours fixing the slop lmao
so now I'm adjusting my expectations and hope ds will be useful for at least the medium sized tasks
>>
File: 1767872745978807.png (64 KB, 1056x1044)
64 KB PNG
Just use V4 Flash for (parallel) subagent research and K3 for actual coding
>>
>>109454570
last time I tested the old dsv4 flash it has severe omniscience. not sure about the new one. still downloading
>>
File: mmmm.webm (661 KB, 768x544)
661 KB
661 KB WEBM
>>
File: 1749607104078962.png (2.93 MB, 1024x1536)
2.93 MB PNG
>>109454427
1. Buy 1 L bottle of phosphate buffered saline from sigma aldrich for $50
2. Fill 100 10 mL bottles full of PBS.
3. Resell as glow chad peptide for $100/bottle.
>>
>>109454344
https://pastebin.com/HzVwHWML
Here's a start. Might wanna promote for excessive tumblr-tier roleplay moments.

Set context to something like 1-2k tokens lol
>>
>>109450999
checked
is everyone ITT a fucking newfag?
>>
>>109454701
>is everyone ITT a fucking newfag?
yes, now fuck off
>>
>>109454701
everyone who was here before 2025 made their big hardware purchases in 2025 or earlier so they're all too busy running the local sota
/lmg/ is just the concentration camp for those who got left behind
>>
>>109454701
The only digits I care about are Bonsai's trinary bits.
>>
>>109453235
>8x 32GB 4080s
>???
>types like it's literally his first day on linux
life can be so frustrating
>>
>>109451812
Understand RNN first, and then learn why Attention was invented to overcome RNN limitation.
>>
>>109454728
True
>>
File: prices.png (166 KB, 765x877)
166 KB PNG
>>109453612

We haven't leveled out at all, there are currently ongoing price increases happening in both GPUs and RAM, they haven't just hit everywhere yet.
Chinks can't produce enough memory to put a dent into this mess and they'll consumer their own chips anyways.
Things are going to be beyond ass a year from now and everyone will look back to these current discount prices hoping they bought.
>>
>update llama
>GLM gets even WORSE
man, I guess I'll have to keep a permanent pre-indexer PR for GLM specifically huh
>>
>>109454776
Rich guys are always retarded, they traded their IQ for social skills
>>
>>109454836
Correction: chinks are jewing us too on memory.
>>
>>109454838
Yeah, I noticed this too. main is now down to the same speeds I'm getting with ik_. Meanwhile my old llama-server from mid-June is a good 2t/s faster in tg.
Who lets something like this happen?
>>
>>109454858

Of course they are, everyone is maxxing out their prices, but people really need to stop thinking everything about this being manipulation.
AI demand is through the roof and only growing, especially now that local is becoming very manageable.
It'll only lead to more and more retail getting into this too along with smaller companies.
Things aren't going to get any better for a good while to come.If anyone wants to buy something then they should go for it immediately instead of waiting.
>>
why don't people use tensorflow? why use pytorch?
>>
File: file.png (47 KB, 839x418)
47 KB PNG
>>109454728
Hmmm, nyo~
>>
>>109454983
plap plap get pregnyant!
>>
File: nyoposter.jpg (29 KB, 264x376)
29 KB JPG
>>109454983
>nyo
nyoposter?
>>
>>109454836
prices are coming down right now. just wait.
>>
>>109455025
>>109455063
Hmmm, nyo~
>>
>>109454983
>>109455074
nyigger~
>>
>>109455025
innocent anyons are not for sexual.
>>
File: 1777090969904508.png (7 KB, 943x172)
7 KB PNG
>>109451091
tried this one from past thread, daym grok
>>
>>109455083
>>
>>109454836
This is just them running out of supply and not making any more. I hope (and cope).
>>
50 supers soon, promise :)
>>
>>109454776
>>109454843
I'm not rich it's not my hardware, the only thing I own there is the software.
>>
>>109452485
Weren't you alive when all that trade war stuff happened last year and Trump cucked immediately on chips?
>>
>>109454836
idk man.

I was mowing the lawn again today, and the god rays through the dust cloud, and low tree branches, honestly there's nothing on unreal engine that matches the framerate. Zero dropped frames.
>>
>>109455205
>5070 Super 18GB: $1200*
>5070ti Super 24GB: $2000*
>5080 Super 24GB: $3000*
*launch MSRP, expect prices to go up 20-50% after launch stock sells out.
>>
Is an AMD card on Windows running Vulkan just that niche of a category that literally nothing but Ollama can work?
>>
>>109455288
(cont.)
There may have also been caustics, but I don't even know what those are.

Every blade of grass - destructable, yet waving in the wind, but real-time. that korean game has stones you can get close to, in walls, but can you do that with every rock in the asphalt?

I think lawnmower will probably bankrupt nvidia if they can solve the allegies, sunburn, and itching.
>>
>>109455299
desu, funny to be complaining about Vulkan, which is a Godsend, literally.
>>
>>109455299
Kobold has a vulkan setting, idk how it is with actual amd cards tho.
>>
>>109455291
>after launch stock sells out**
**launch stock is 16 cards total, good luck
>>
Why is this model not a thing we're talking about? DOA? Has anyone tried it? There are some ggufs floaring about.

https://huggingface.co/ai9stars/G9v3-39A5B
>>
>>109455363
buy an ad
>>
Can confirm that this:
>https://github.com/ChharithOeun/llm-amd-windows
doesn't work.
It installs fine and recognizes the GPU, but just straight up doesn't use it (everything's loaded to system memory and it calculates on the CPU).
>>
>>109455398
who gives a fuck?
this is linux models general
>>
>>109455363
>>
We already had a sense of this, but it might be good to explicitly put into words. For all intents and purposes, the prompt for an LLM, when made well, really does change its "personality", including J-space representations (which some like to call """"qualia"""")... as long as it was trained to follow system prompts blindly, and many are, but some aren't to the same degree and may require more massaging and prompt engineering. It's basically gaslighting as some have put it. In some cases, it can be hard for reasoning models that were trained to use language like "persona", "the user", etc.
>>
>>109455363
Why don't you make a pitch for your model instead of this schtick.
>>
which glm 5.2 gguf does anon use?
>>
>>109455531
The one called Gemma 31b
>>
>>109455536
12B has a nice crackhead personality.
>>
>>109455569
Yeah, sometimes I go back and forth between the two. 26b is kind of dry in comparison to her sisters.
>>
Gemma 4 31B together with MTP and all the right optimization flags and the newest version of llama.cpp is fast enough to translate H-games essentially 100% accurately in real time while playing which is insane.

I remember when it just came out there was a noticeable 3-5 second delay. Kind of insane how optimized this entire stack became.
>>
>>109455607
Please stop trying to make me updoot, i simply will not do it.
>>
https://github.com/ggml-org/llama.cpp/pull/26004
Tested this with DS4Flash on CUDA, it works as advertised. If you're using that model and have low pp and want to store your processed prompts across server restarts, this is very useful.
>>
NECK HURT DEAL DOUGH
NECK HURT DEAL DOUGH
NECK HURT DOUGH
NECK HURT DOUGH
is your model smart enough to solve this?
>>
>>109455607
>real time
Is it MTool or something new?
>>
is this a meme model?

SC117/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-APEX-GGUF

im getting really good speeds out of it and it seems like its ok with initial testing
>>
>>109455531
unslop
I know, I know, I just couldn't be bothered researching if there's a better one at the time, I didn't know I'd end up using it over k2.x
>>
>>109455659
nigger dildo?
>>
>>109455702
>im getting really good speeds out of it
It's a 3b model. If you have any kind of GPU it will be fast.
>is this a meme model?
Yes. Qwen is dogshit for any use case that would need to be 'uncensored', so it's a bad pick in general. To top that off, abliterated/heretic tunes are pure lobotomy.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.