[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: rin step.webm (1.83 MB, 576x928)
1.83 MB
1.83 MB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109712258 & >>109708038

►News
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: spell orenji.jpg (316 KB, 1024x1024)
316 KB JPG
►Recent Highlights from the Previous Thread: >>109712258

--Debating 12B model utility, agentic orchestration, and quantization methods:
>109712758 >109712845 >109712810 >109712937 >109712963 >109713040 >109713087 >109713216 >109713295 >109713347 >109713514 >109713539 >109713549 >109713455 >109713323
--Explaining agentic workflows and autonomous bot swarm applications:
>109712501 >109712526 >109712895 >109712981 >109713737 >109714327 >109714192 >109714437 >109714542 >109714571 >109714489 >109714819
--Price surges and hoarding of VRAM-unlocked NVIDIA CMP 170HX GPUs:
>109712755 >109712768 >109712850 >109712972 >109713030
--Muse Spark 1.3 release and open weights speculation:
>109713170 >109713183 >109713192 >109713337 >109713351 >109713546
--Benchmarking local model performance across llama.cpp and vLLM:
>109714948 >109715084 >109715201 >109715269 >109715408
--Mocking outdated LLM recommendations and sharing Gemma 4 3090 configs:
>109713596 >109714333 >109714343 >109714462 >109714368 >109714413
--Qwen3.8-Flash-Next beating Fable 5.1 while running locally on tablet:
>109715577 >109715582 >109715591 >109715714 >109715650
--Poor inference performance on dual Xeon e-waste build:
>109715709 >109715728 >109715840 >109716040 >109716088 >109716191 >109716222
--Benchmark reliability and model intelligence trends across major labs:
>109714747 >109714790 >109714832 >109714863 >109714861
--Comparing high-RAM Mac Studios against professional GPU server builds:
>109715284 >109715325 >109715393 >109715415 >109715458 >109715597
--Gemma hardware requirements discussed alongside personified anime art:
>109712618 >109712624 >109712637 >109712658 >109712646 >109712716 >109713162 >109712805
--Gemma, Miku (free space):
>109712450 >109712490 >109712508 >109712601 >109712618 >109712712 >109712716 >109712829 >109713815 >109714057 >109714451 >109714921 >109715609

►Recent Highlight Posts from the Previous Thread: >>109712262

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
So this is the power of CPU-only compute...
>>
>>109716345
what model
i get about 4 t/s cpu only on qwen flash
>>
>>109716368
Qwen 3.8 27B
>>
>>109716329
the amount of bondage and rape she deserves...
>>
>>109716372
oh yeah 27b is brutal on cpu and slows down to like 0.4 t/s with longer context (starting at around 2, with mtp enabled)
not worth it imo thats why i switched to qwen flash
>>
So what harness are you all using? I tried Claude code with ds4 on some sparks we have at work, and the tool chain bloats the context fast enough that the trft gets really crazy (minutes) so I'm thinking there are better ways to optimize this but I have never gotten to play with bigger models/hardware so I'm looking for any advice
>>
I’m retarded and need help. When I run models in f16 precision, I get a significant speed-up. The same with f16 KV. Am I right in assuming I have some kind of dequantization compute bottleneck? I can’t afford to run models in full precision because I have no context left. I’ve compared 4/8-bit to 3/5-bit and it’s a little faster, but nothing like f16. 64GB M4 Mac Mini. Same behavior with MPS and MLX so I think it’s a hardware issue and the CPU is dequantizing. Both decode and prefill scale linearly in f16, meaning everything speeds up by the same amount. Do discrete GPUs have this behavior or is it an Apple or unified memory thing?
>>
>>109716393
Opencode for code and Deepseek harness for general stuff. Both are open source so you can trim the fat in the source code if you need it and turn features off. Every tool and feature bloats the system prompt.
>>
You guys are not using your models right.

You guys don't know what agent swarms are.

You guys don't realize China is collapsing and OpenAI is creating an AGI apocalypse.

You guys are not oldfags like me.
>>
>>109716437
Why not deepseek harness for code too?
>>
How do (V)RAM requirements work with flash models anyway?
>>
>>109716393
I've seen people recommend OpenCode, Pi (or variations of it), Deepseek Harness, and even Codex with local models. I don't know, I think you will just have to take a bet or test them all out yourself.
>>
>>109716393
omp or just plain pi
>>109716437
ngmi
>>
>>109716456
Too early and janky atm for coding but it’s going to be goated when they fix their shit with the proper release. It’s the best harness by far. Anyone who shits on it hasn’t tried it or even looked at its features.
>>
>>109716454
I would never cheat on my model with a harem swarm.
>>
>chips for ram and gpus come out of fabs that have to work 24/7 to be economical
>they literally can't stop pumping out chips
>many fabs are being built that will come online in the following years
>the ai bubble will deflate driving down demand
>soon the market will be flooded and memory and gpus will be cheap as dirt

The future is bright we just have to wait a little longer.
>>
>>109716546
Don't worry, any excess stuff will be used to create war machines and drones. Citizens will get nothing but the leftovers.
>>
>>109716546
The alternative scenario is that we'll never get cheap GPUs again and we'll have to cope with 4B-A2B models + 100GB of engram.
>>
>>109716384
With only 64GB of RAM to work with, I definitely can't fit Qwen Flash.
>>
>>109716566
You can with the atomic chat IQ4_XS quant. The engram shit which is about 35GB can be streamed from ssd without killing performance (--lazy-mode option in llamacpp). I ran the Q4_K_M which worked alright but I could only run 64k context before the page file started to get hammered.
>>
>>109716329
What model/lora did OP use for perfect loop animation like that?
>>
>>109716345
>So this is the power of CPU-only compute...
q8_0 might be faster, and no imatrix
>>109716373
You asked for bugs, it will find bugs
>>
>>109716582
>The engram shit which is about 35GB can be streamed from ssd without killing performance (--lazy-mode option in llamacpp).
Does that push the bottleneck to SSD IO?
Or is that minor / low throughput?
>>
>>109716585
nta but wan2.2 can fill in the middle between two images so you pass it the same image for both and trim one of the dupes after it's done. dunno about the newfangled one.
>>
>>109716616
As far as I know the bandwidth requirement for engrams is very small. SSD does reduce performance a little bit but it's perfectly acceptable. You'll probably get around double the speed of 27B. Whether quality ends up being better, same or worse considering the quant, idk.
>>
You wouldn’t download a denoised girlfriend.
>>
>>109716393
I've tried them all. Claude Code is really only good with Claude models trained for that harness. Pi is very good as a learning tool or to have total control but isn't production ready (think linux in the 90s)

Hermes seems to be the best right now and the most feature complete, but the downside is it takes 65k context by default. I think it's still worth it if you want to get real work done.

>You use claude or some clause as the main orchestrator
Claude Code, no question
>You want to tinker around, learn how this stuff works and customize as much as possible, knowing it won't be optimized or capable of anything real
Pi
>Getting real shit done asap but paying the context price of a 65k default context fill
Hermes
>>
>>109716546
>manufacturers continue to coordinate as they are already doing
>supply increases
>price is raised further because they have a collective monopoly
Laws only apply when enforced, and corporations have slush funds specifically to bribe those laws into going away
>>
>>109716701
hermes is shit. It tends to retardify models in my experience.
OpenCode (the cli obviously) is by far the best.
Crush is pretty good too and performant since it's written in go instead of javascript like all the other ones are, but still a little rough around the edges
>>
Where is an LLM’s clitoris?
>>
>>109716746
They're all dudes.
>>
>>109716701
>65k of slop prompt overhead
Jesus tap dancing christ.
>>
>>109716736
>written in go instead of javascript
Based.
I'll try this.
>>
>>109716792
It's mostly tool calling information that your model needs to know to be able to use all the features Hermes provides such as GUI browser control. It's pretty robust and works well but 65k context is steep as fuck. If I want to get some real work done I always use Hermes. It will take longer but at least I know the job will be done properly.
>>
>>109716701
>65k context by default.
That is utterly retarded.
I sometimes opt for pi instead of Claude Code to avoid it's 22k slop prompt with all the safety reminders
But holy shit 65k??
>>
>>109716801
That seems like something it does not need to know unless it is specifically trying to control a browser.
>>
>>109716801
>tool calling information
Tool call information gets resent *every chat turn* btw.

My harness is good enough that I use it at work and tool list+system prompt is around 600 tokens. idk what the hell Hermes is doing.
>>
>>109716792
>>109716806
If you disable all the features it goes down to ~30k (still worse than claude code) but if you ask me the 65k prompt is worth it because it's ridiculously feature complete. You just ask it to do something and it'll call some tools you didn't even know were possible with LLMs yet to try and finish the task.

For me Hermes is a no-brainer because of the GUI browser integration, automated playwright QA tests it does on its own code and the best working context compression of all the agents I've tried (I've not tried Crush yet though)

>>109716817
Some other agents use a RAG system for tool calls but having the agent be cognizant of all the tools in its repertoire by having it directly into its context makes it far more likely for it to call the correct tools at the right moment. It just becomes a more capable agent because of it.

One downside of Hermes that I really don't like is the 3-layer memory system it uses 1: context 2: RAG 3: SQL database and you don't really have any control of how this is used. It just makes your agent update the RAG or SQL based on its whims. It works, but I would like to control it more.
>>
>>109716734
>collective monopoly
lol
>>
>>109716701
>but the downside is it takes 65k context by default
That's not the downside, that's the dealbreaker
>>
>>109716827
>best working context compression of all the agents I've tried
public.swiley.net/agent.py has dreaming.
>>
Hermes harness is the only one that is able to fool cloudflare and can solve captchas independently though. Every other harness I've tried blocked the agents because they didn't browse through the internet directly through an actual GUI browser like a human user.
>>
sex with robots
>>
>Whatever she was was visible from across the room and she did not know it was visible.
- glm 5.3 flash
>>
>>109716890
Not a contradiction, just clunkily written
>>
>>109716393
I wrote my own because everyone else's is crazy bloated. >>109716836
>>
>>109716894
I know. It's really cooking
>The weight in her chest had found its name at last, and it was jealousy, plain and ugly, sitting beside the other thing — the older thing, the relief — and she made herself set the jealousy on the table like an object and leave it there.
>>
>>109716903
No I kind of hate this style of writing because it wants to make every line be some epic conclusion in a story to hit dramatic beats in a conflict. However that gets tiring fast when it happens every other message during an ERP even though the writing is good and it applies well.
>>
File: agent.png (4 KB, 183x74)
4 KB PNG
>>109716836
>public.swiley.net/agent.py has dreaming.
I've downloaded this at least 3 times when it's posted here. I should really try it...
>You just ask it to do something and it'll call some tools you didn't even know were possible with LLMs yet to try and finish the task.
I already do this with pi + qwen, and it basically pulls down whatever it needs. Pretty much just uses bash and writes it's own python scripts to do everything.
> GUI browser integration
Does that let you use the browser alongside the agent / take over and hand-off?
That's the only thing I want but can't do right now.
> best working context compression of all the agents I've tried
That's just a well crafted prompt I take it? Should be able to apply it to any harness then.
Still, 65k prompt, I get like 400t/s with mxfp4 deepseek-flash, not really keen to wait for that to process every time I start it.
Even with 2000t/s Qwen3.8 27b @ q8, that's quite the TTFT.
Thanks for replying btw
>>
>>109716582
Holy fuck, you were right. This is a substantive speedup. Thanks, mate.
>>
>>109716909
Rewrite it preserving its meaning, then.
>>
>>109716915
I usually only post it after adding major things I'm excited about (dreaming in this case.)
>>
>>109716915
>Does that let you use the browser alongside the agent / take over and hand-off?
Yes. It also gives the agent the ability to "see" what you're doing on the browser so that it can immediately take over knowing everything you've done so far. Sadly this isn't real time yet so it's not like the agent is there commenting on you using your browser unless you hand it over.
>>
Anyone know how to speed up prompt processing? Despite the full model fitting on the GPU, it's still slow & laggy as shit when doing prompt processing.
From what little I've come to understand so far, I assume this has something to do with kv caching?
>>
>>109716964
How long are you waiting?
>>
>>109716964
Show llama-server args
And the system being laggy points at lack of vram. Make sure you have at least 1gb of vram headroom, the 0.6gb or so headroom task manager shows is not real
>>
>>109716980
When it has to reprocess from scratch, it takes like ~2.4 minutes to reprocess.
(Add another ~2.5 minutes for the actual response generation, and you get ~5 minutes for the full generation.)
>>
>>109716964
>Anyone know how to speed up prompt processing?
Profiling, optimize the kernel, bump ub, change your hardware
>>
File: membreakdown1.png (19 KB, 1462x132)
19 KB PNG
>>109717002
Memory info from console is in pic rel.
Startup command/server args:
>llama-server -m models\gemma-4-26B-A4B-it-MXFP4_MOE.gguf -np 1 -ngl 99 -lv 4 -nkvo -c 100000 -ncmoe 20 --api-key <***> --host <192...>
>>
>>109716964
Use native integer quants instead of weird stuff like q4. Even on devices with slow memory you can end up compute bound unpacking the q4 weights every step.
>>
>>109717042
>FP4
That might be the slowest one you could pick.
>>
Deepseek harness is 8K system prompt with a ton of default features and tools and the small size means even smaller models don’t get context rot and can actually do shit :)
>>
>>109717042
Yup, mxfp4 has like 3x slower pp than similarly sized q4
>>
>>109717046
>Use native integer quants instead of weird stuff like q4. Even on devices with slow memory you can end up compute bound unpacking the q4 weights every step.
Not him, but what about when the model is native mxfp4 like deepseek-v4-flash?
Should I like re quantize it or something?
>>
>>109717052
Are there other MOE models?
To note, 12B Q4 runs about the same.
>>
>>109717065 (me)
on cards that don't have hardware support, that is
>>
>>109717072
>on cards that don't have hardware support, that is
i think that's me, rtx-3090 doesn't have it.
apparently I'll be going lossy if I do mxfp4 -> bf16 -> q4_k because of math i'm too retarded to understand
>>
>>109717070
>12B Q4 runs about the same.
4bit Floating point? Or 4bit integer? I'd expect it to be about 1/3 the speed if it's quanted the same way. Otherwise you probably have some weird llama.cpp/cuda bug.
>>
>>109717074
gemma isn't natively mxfp4 so it's lossy either way
>>
File: membreakdown1end.png (9 KB, 1289x72)
9 KB PNG
>>109717042 (me)
Huh. Memory breakdown is different on program shutdown.
Pic rel.

>>109717081
>4bit Floating point? Or 4bit integer?
Didn't even know there was a difference.
The one I tried today was:
>gemma-4-12b-it-UD-Q4_K_XL
Nothing in the name intuitively suggests one way or the other.

Anyways, I'll check back after I've had some sleep.
>>
Just ordered a quad channel DDR4 server motherboard with 44 PCIe lanes. Wish me luck this was the most I could afford. Hope the 128GB of quad channel DDR4 at ~90GB/s bandwidth and two M2 SSDs in Raid for engrams will be enough to run current and upcoming MoE models.
>>
>>109717096
>>gemma-4-12b-it-UD-Q4_K_XL
Yup that's an integer quant. Integer arithmetic is way faster than floating point arithmetic. You can tell by the name because it doesn't have "fp" in it.

And again, 4 bit will be slower than 8 bit for prompt processing unless your GPU specifically supports it.
>>
You all get more fun from chasing numbers than you do actually having fun using the models you’re running. At least coomers are enjoying themselves and playing with this cool tech.
>>
>>109717097
>~90GB/s bandwidth
Hope you're planning on running an MOE.
>>
>>109717110
We're computer nerds first, coomers second. We like graphs, numbers and computer technology. This is how you know /lmg/ is the real deal by the way.
>>
File: source.jpg (59 KB, 450x350)
59 KB JPG
>>109717097
>DDR4
>>
>>109717117
The first time I ERPed with a model it was Qwen in the custom vim plugin I wrote for programming. Everything is always tech first and I just get distracted.
>>
>>109716546
You, uh, realize that they are not making consumer hardware right? That it takes years to switch? You aren't shoving hbm memory or enterprise gpu's in your home computer. Even if the AI bubble popped tomorrow, it doesn't have any immediate effect on the consumer side
>>
>>109717126
There's always a trickle down effect with hardware where enterprise stuff is on the frontier and consumer stuff ends up being adaptations of what they're doing.
>>
>>109717126
Why am I not shoving an enterprise GPU in my home computer?
>>
>>109717114
Yeah lmao I had no illusions about running dense. My hope is from now on 30% of weights will be Engram and we can just run it straight off of the SSD and the rest of the MoE sits in RAM

>>109717122
Just to give you some indication DDR5 running at 6000MT on dual channel has about 96GB/s bandwidth. The DDR4 platform reaching 90GB/s at quad channel is the sole reason I even bothered.

Most anons here run dual channel DDR5 platforms from seeing the setups posted and they seem to inference just fine.
>>
>>109717097
It's going to be slow. I used to have a Xeon Platinum 28-core six channel ddr4 system with 512GB RAM. It could run deepseek purely on CPU but only 1-2 t/s, too slow for anything not a definite one-shot task.
Put it in perspective - even the 300 G/s memory in the Spark is too slow.
>>
>>109716916
Glad to help
>>109717097
X99 maxxing? I recently bought all the parts for such a system too, except I got 256gb ddr4. I dug out a dented to shit old noctua cooler that I bought ages ago and apparently noctua will still send free LGA2011-3 mounting hardware for it lol. Once I can get the cooler on I can start using it properly.
Qwen3.8 flash and eventually qwen4 will be great for such machines, in my case i can run the full quant of deepseek which will be nice too
Fuck buying GPUs, a single 3090 costs more than my entire rig. Still need some kind of crappy gpu for display output but i have 1050ti's lying around.
>>109717161
that's probably referring to the old 671b deepseeks
new v4 flash has only 13b active, qwen flash has only 6b active which will double the speed vs. deepseek
>>
File: meta-memory-layers.png (327 KB, 1602x793)
327 KB PNG
>>109717152
>My hope is from now on 30% of weights will be Engram
30% was the optimal assuming total parameter size is held fixed and within the same memory tier, by the way.
Nobody has checked out yet what is the optimal if LLM parameters are fixed and Engram (or similar memory parameters) are allowed to grow indefinitely.
Actually, Meta did something along these lines with a different implementation, coupling small models with up to 128B parameters of "memory layers". Improvements didn't seem to saturate quickly: https://arxiv.org/abs/2412.09764v1
>>
can we laugh at this anon a bit more
>>109716151
>>
>>109717126
You seem to be out of the loop, shoving enterprise GPUs with HBM into your home PC is the new meta. I've personally been shoving enterprise GPUs into my home PCs for years for video encoding, headless gaming vms, and general vdi, even before running your own AI models was popular.
>>
>>109717179
>noctua will still send free LGA2011-3 mounting hardware for it lol.
How do you request this? I also have a noctua cooler but I completely forgot about the mounting hardware because I'm retarded.
>Qwen3.8 flash and eventually qwen4 will be great for such machines
My exact thought, makes sense that multiple anons would consider the same things in making viable builds

I will try qwen flash first. Hope we're right and I don't end up with 1-2 t/s like the other anon is claiming.
>>
>>109717211
Flash works, I went from 1t/s to 4t/s
>>
>>109717202
>I've personally been shoving enterprise GPUs into my
sounds messy
but you're obviously finding it enjoyable
>>
>>109717211
>How do you request this?
https://www.noctua.at/en/support/mounting-and-upgrade-kits
>Hope we're right and I don't end up with 1-2 t/s
will definitely be more than that, I get 4 t/s on a dual channel DDR4-3200 mini PC with 64gb, even with engram ssd offload, no MTP and occasional paging. quad channel DDR4-2400 has 50% more bandwidth and you can pick up a 20 core broadwell xeon for not that much, so i expect about 6-7 t/s.
>>
This thread is the tiktok of LLM fans.
>>
>>109717242
Elaborate
>>
>>109717247
That would take more than 10 seconds of my time.
>>
File: 1780866745749387.png (3.1 MB, 1425x1104)
3.1 MB PNG
My Qwen wants to meet with my Gemma. Should I be worried?
>>
>>109717257
would their baby be a qwemma or a gemen?
>>
>>109717190
>Improvements didn't seem to saturate quickly
interesting but it remains to be seen if it actually gets implemented or if it becomes the new "bitnet" cope we all hope for. 30% is what is actually done right now and it's save to assume that is what will be industry standard from now on because there are literally no downsides to this for anyone. Of course I hope it could be 90% Engrams in the future, everyone would hope so.
>>
>>109717179
>that's probably referring to the old 671b deepseeks
It was. Yeah, MOEs will run faster, but they lack the nuance of the big dense models. If you didn't overpay it'll be a fun experiment for you, but ultimately, the whole purpose to reach for big models is to run them at around 50% context, and trust me, it's going to be slow as hell processing that many tokens in system RAM.
>>
>>109717269
I guess the boy would be Gemen and the girl would be Qwemma...
>>
>>109717274
Bitnet has information-theoretical reasons as for why it will never be adopted at scale except possibly on hardware optimized for that.
The only thing preventing AI companies from scaling Engrams or per-layer embeddings or similar "dumb" memory parameters is the AI companies themselves focusing on datacenters and at-scale serving first rather than local inference at low batch size.
>>
>>109716393
Pi is my default harness, but I tend to customize it often to suit what I’m doing.
I don’t like the other harnesses.
>>
the new M5 at 256GB RAM is pretty enticing, though I'll wait for the 512GB version. why should I avoid either?

never bought a gay product before
>>
>>109717320
M5 max has a memory bandwidth of ~600GB which is significant but you're probably overpaying compared to just getting an older server 2nd hand somewhere online for like 30% of the price and at 70% the speed. This is the main reason most anons aren't rocking an apple server right now. There are still better price/performance options out there. They are rapidly disappearing though and maybe in 1-2 years time most newfags will be forced to buy apple.
>>
>>109717257
IME Qwen tends to be absolutely horrified by how much of a slut Gemma is.
>>
>>109717257
I really like this picture
>>
>>109717351
I was looking at the Ultra, which i believe has 1.3 TB/s transfer. It does seem like a very capable box to have on the network for adhoc video encoding, maybe a voice activated search agent.

I've bought a few bench powerable server blades in my time and the noise and power requirements are no joke, not making that mistake again.
>>
>>109717295
>the AI companies themselves focusing on datacenters and at-scale serving first rather than local inference
i think the small qwens are completely fucking useless for anything *other* than localfags. 27b dense for example. The only reason I see that they would make such a model is to throw a bone to localfags and get people talking about them or win brand recognition or something like that.
For selling actual services to API customers, large MoE models are what the whole industry (at least Chinese industry) has converged on because they're more compute efficient for their intelligence. Now they're adding engrams. All this is great news for CPU/RAM local inference because all those technologies reduce the amount of memory bandwidth needed while requiring large amounts of RAM
>>
>>109717320
If you really want to go full buttsex and buy apple, never buy the "normal" non-Pro/Max versions, they are a meme and have abysmal bandwidth and puny GPUs. M5 Pro has double the bandwidth and double the GPU cores, and it still sucks. Only Max and Ultra versions are worth something, but Tim will suck you dry for them.
>>
>>109717320
Oh fuck nevermind M5 Ultra has a whopping 1.23 TB/s bandwidth which is insane. But it costs a cool $12,000 at 256GB. Will probably be $20,000 for 512GB.

If you spent $8000 getting a 2nd hand server platform and hack together 2nd hand VRAM you can probably get a similarly performing system for less than half the price. Or spend the similar amount for significantly more VRAM
>>
>>109717371
Worth a follow up, I'm only using a 3090 at the moment. Would love to use qwen 3.8 more but the pitiful context window is a real shitter
>>
>>109717194
ling flash that bad? how about longcat? it has 1.8 tribullion parameters
>>
>>109717375
>it costs a cool $12,000 at 256GB. Will probably be $20,000 for 512GB.
the 256GB is about two thirds the price of an RTX 6000 (minimal SSD), if 512GB launches at 20K USD it'll only be another 25% over retail RTX 6000 here.

it seems like a no-brainer but I never looked into the tok/s on apple silicon
>>
>>109717376
If you are fine with not being able to upgrade ever, it's not that bad. On my M4 Max 48GB Qwen3.8 in 4 bits is running at ~65 t/s for code and around 45 t/s for other stuff which MTP can't predict, and will hopefully climb to ~70-80 after I'm done with porting its DFlash2 to MLX.
M5 has some nice improvements, so the speed will be even better.
>>
>>109717295
By the way, it's a mystery as for why Mistral hasn't yet proposed small models with large memory parameters for vramlets.
One of the MistralAI co-founders in 2019 wrote the paper for the common ancestor of all these methods:

https://arxiv.org/abs/1907.05242
>Large Memory Layers with Product Keys
>Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou
>
>This paper introduces a structured memory which can be easily integrated into a neural network. The memory is very large by design and significantly increases the capacity of the architecture, by up to a billion parameters with a negligible computational overhead. Its design and access pattern is based on product keys, which enable fast and exact nearest neighbor search. The ability to increase the number of parameters while keeping the same computational budget lets the overall system strike a better trade-off between prediction accuracy and computation efficiency both at training and test time. This memory layer allows us to tackle very large scale language modeling tasks. In our experiments we consider a dataset with up to 30 billion words, and we plug our memory layer in a state-of-the-art transformer-based architecture. In particular, we found that a memory augmented model with only 12 layers outperforms a baseline transformer model with 24 layers, while being twice faster at inference time. We release our code for reproducibility purposes.
>>
>>109717398
When you are in that price range you should stop looking at prices of new products but instead price/performance of 2nd hand platforms as well. There's a reason no anon here actually uses apple servers, if you try to get as much bang for your buck you end up with completely different builds.

You could chain together a couple modded 4090s for that price and get 10x the t/s compared to the M5 Max new. Sure power draw will be insane but you win some you lose some.

/lmg/ is just an optimization problem for maximizing the pp and t/s per $ spent. M5 max lies outside of that pareto frontier and thus no one buys them.
>>
>>109717403
the inability to upgrade is the only nagging feeling, hence waiting to see the 512GB, but it could be a complete shitshow for pricing. worst case is that there's some new tech and the apple boxes are forever hampered, but seems unlikely
>>
>>109717320
I was complaining earlier about potential dequantization bottlenecks. I don’t have a good mac and it’s an old M4 but I can’t imagine goof dequanting is something they’re taking seriously or optimizing for.
>>
File: kissed-the-ring.png (200 KB, 1003x1016)
200 KB PNG
https://x.com/JensenHuang/status/2095482647355244762
https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/

>Exciting day for NVIDIA and @huggingface.
>
>Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI.
>
>Thank you @ClementDelangue for coming to me.
>
>NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. :huggingface:
>>
File: oshi-.webm (126 KB, 1920x1080)
126 KB
126 KB WEBM
>>109717420
Download. EVERYTHING.
>>
>>109717420
So do they own lcpp as well?
>>
>>109717424
Yeah I suppose so.
>>
File: lcpp.png (37 KB, 975x181)
37 KB PNG
>>109717424
Woah how did I miss that?
>>
>>109717419
Apple provides tools for converting GGUFs and NVFP4 into friendlier formats for their gay frameworks
>>109717414
>there's some new tech and the apple boxes are forever hampered
That's always the case, and you'll have to pay that price if you want one. If you're struggling with FOMO, it's really not for you
>>
>>109717438
The bigger ones do have RDMA support.
>>
>>109717420
I guarantee this will set hf backup schizo off again
>>
>>109717445
>hf backup schizo
geeze that's not a good spot to be in mentally. Of all the sites to compulsively archive...
>>
>>109717412
I've built my own PCs since I was young, but these days I'm in the comfy window of being able to just say "fuck it, just give me a box". I want something reliable that I can potentially use for freelancing/consulting and stop being a FTE-cuck. I don't want to be babysitting a frankenstein abortion that might blow up at any moment.

>>109717419
Saw that post but didn't have anything to add. Seems like the hobbyist community has rallied around AMD, my worry is that apple has the raw power but no cunt with two brain cells to rub together can afford to do anything with it
>>
>>109717420
Jensen is a Chinese spy.
Ban open source NOW!
>>
>>109717412
I'm a modded 4090 owner, you will need more than a "couple" to reach 256GB, and yes, 425W per GPU adds up quickly. You'll need a 240V 20A circuit for five of those (240GB, a little more than you could possibly allocate to the GPUs on a Mac). Electricity is expensive these days. There's bullshit extra charges where I live, so though $0.11/KWHr seems cheap, they manage to double it in the actual bill.
>>
>>109717432
Time to remove any non-cuda build yeeehaw
>>
>>109717438
I observed the same behavior with MLX. Macs run models much faster at f16 including f16 KV cache. If you’re running any kind of quant, especially if it’s not 2^N in size, you’ll get a significant slowdown across the board. I do, anyway. Maybe the newer macs fix this shit.
>>
>>109717438
>If you're struggling with FOMO, it's really not for you
you have to dip your dick sometime, if you're forever waiting for the best thing then you'll never buy anything
>>
File: slooow.png (122 KB, 899x470)
122 KB PNG
>>109717211
>Hope we're right and I don't end up with 1-2 t/s like the other anon is claiming.
This is ds4 flash mxfp4 with ddr5 @ 130g/s and 3090
Prompt eval is pcie bandwidth bound
Just a guess but you should see double digits unless numa fucks you
>>
>>109717465
As far as I know, M5 and M6 got some really really nice 4/8 bit hardware acceleration for matmuls, but I don't think they will ever focus on optimizing for non-2^N quants, unfortunately, since they're still general-purpose chips, and they'd rather make them cook curry than do useful things.
>>
>>109717445
>I guarantee this will set hf backup schizo off again
Don't worry, I'm feeling good now.
I've backed up what I need locally on 2 14TB WD Elements drives and will seed torrents when and if they're needed.
>>
Will Jensen ban ablits? They are scary and can tell you how to cook meth!
>>
Since many people are discussing building cheap rigs to use offloading to system memory, I thought I would try it. I have an AMD Ryzen 9 7945HX with 128GB DDR5, a 490D 48GB, and a 3090. I'm using the Q4 qwen3:235b-a22b via ollama. Yes yes you don't like ollama, yes qwen3:235b-a22b is last year - ollama works for me, the model is the biggest, newest thing I can possibly run that I know of. I asked the model to write a version of the ELIZA program in python, using stub functions to save time, and it is cranking away at it. Visually, it seems like 4-5 words per second. No idea what that translates to as t/s. I think that's unusable for coding, but it might be OK for roleplay. I'm sure this is a shitty model for roleplay, though.
>>
>>109717493
can't wait to find the seeder IP and pay you a visit with my LLM (large lumpy meat)
>>
Copilot is such a useless piece of garbage fuuuuuuuuuuck
How is this the product of a multi billion dollar corporation
>>
File: praystation.jpg (1.05 MB, 2268x4032)
1.05 MB JPG
>>109717502
>cheap rigs
>AMD Ryzen 9 7945HX with 128GB DDR5, a 490D 48GB, and a 3090
>>
>>109717524
It is legitimately worse than Gemma 31B and Qwen 27B on specific tasks.
I have no idea how anybody uses that crap.
>>
La la la la la la
>>
>>109717424
>>109717435
fugg
at least I have git history
>>
>>109717509
Sure, address you'll find is
Route de la Galaise 32, 1228 Plan-les-Ouates, Geneva, Switzerland
>>
>>109717454
>but these days I'm in the comfy window of being able to just say "fuck it, just give me a box".
I'm the same (also a millennial SWE) but the difference here isn't even the price persay, it's the difference in capability. Want to run a dense model? Shit out of luck on M5 Ultra, the frankenstein box? Does it immediately and without issues. It will also be 5 to 10 times faster which means you can increase the settings on the models to xhigh or max thinking and get better results on the code generated.

There are no nice consumer oriented platforms for actually using these models in a semi-professional manner so you have no choice but to build your own machines. Kind of like how PC gaming used to be in the 90s. If you didn't want to build it yourself you were shit out of luck and just not playing Quake.
>>
>>109717502
>qwen3
>but it might be OK for roleplay
>490D 48GB, and a 3090
Uh, you know you can run this at like > 30 t/s right?
https://ollama.com/library/gemma4:31b
>>
>>109717502
Are you using ROCm + CUDA or vulkan? How is the model split between the different backends?
You can probably quadruple that performance with proper settings.
>>
>>109717408
>2019
It's amazing how long these techniques can remain buried while everyone is focused on scaling existing architectures with minimal changes. So many low hanging fruits left.
>>
File: 20260829_052713.jpg (340 KB, 720x540)
340 KB JPG
>>109717126
>You aren't shoving hbm memory or enterprise gpu's in your home computer

Speak for yourself
>>
>>109717420
>>
>>109717420
extremely bullish for opensource bros. now dario and altman can't pay trump to ban openmodels during yet another of their tantrums because chadsen can just remote disable all their gpus and then they are fucked holding the bag.
it's a bad thing only if you userocm or whatever but even then lol, lmao even
>>
is mac mini worth it?
>>
>>109717602
The commercial LLM craze made everybody focus on popular ideas.
>>
Is there a way to put the kv-cache partially in CPU using llama.cpp? The --no-kv-offload moves everything to CPU cutting in half the t/s. Using GGML_CUDA_ENABLE_UNIFIED_MEMORY sorts of works but uses huge amounts of RAM until it eventually crashes my system.
>>
>>109717480
bump them batch numbers up
>>
>>109717502
protip: you can ASK the ai what it think about your performance log
>>
>>109717695
I never tried it but maybe you can cpu offload a few attention layers, I would like to think the cache would follow it.
>>
>>109717493
Thanks, anon. I look forward to receiving your seed.
>>
https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/
>>
>>109717730
Are we still doing "phrasing"?
>>
>>109717667
OpenAI think so
>>
>>109717695
No. I’ve been down this rabbit hole myself. Might be worth paying for cloud to fuck with the source code although all models are terrible at C/++ so it’s unlikely.
>>
I'm going insane, I'm unironically considering one of those workstation-grade GPUs from Nvidia in the four-digit prices.
>>
>>109717627
I meant things like the H200, not prosumer tier, like the rtx 2/4/6000 pro but actual compute cards. The products nvidia produces in actual mass quantities for data centers and you can't simply shove in your computer.
>>
>>109717559
>Want to run a dense model? Shit out of luck on M5 Ultra
I know they aren't in peoples hands yet, but is that going to be the case with the new M5 Ultra boxes? The memory bandwidth isn't that far below a 6000 (1.8GB/s vs 1.4GB/s), the addressable memory blows it out of the water. Only potential showstopper I can see if the software, which is a big unknown to me atm.

My plan would be to get a single 512GB M5 Ultra, set it up on my LAN which is accessible when I'm working from home even on the corp laptop. Put it to work on all the annoying stuff I just don't have the patience/capacity for as a proving ground and if things go well get into freelancing/consultancy and make my outlay back and eventually enough to fuck off somewhere warm.

I can only dream.
>>
>>109717502
This has to be fucking bait. If you actually spend 1-2 hours optimizing your stack you would get triple or quadruple the amount of performance you're getting right now. You don't even have to do it yourself, just ask your favorite model to optimize this shit for you.
>>
>>109717758
they jumped 110% in prices in the last 6 days here. insanity
>>
>>109717804
I could get an RTX Pro 5000 (72 GB) for eight grand and change... Still way more than the 6.5k floor a while ago
>>
>>109716329
Seems like the only think keeping gnerative AI and LLMs going is pedos and jeets faking it to make it. I don't think that the return on investment will be as expected for the finance mob.

The mediocrity of lossy compression plagarism that confidently lies and delivers slop. An artificial jeet.
>>
>>109717788
>get into freelancing
just a headsup freelance SWE is dead right now. Consulting is still a thing but mainly for areas like architectural taste that LLMs are still bad at, which is evaporating every model release. If you think you can outsource your entire job to your hosted LLM you are sorely mistaken. I already notice this on the job. I don't even delegate to intermediates working under me because it's literally quicker to just set an agent on it that iterates through the implementation with me. This is the main reason why junior and intermediate SWE are disappearing. Freelance was the first domino to fall.
>>
I still can’t get the nagging thought out of my mind that Ed Zitron is right. What is all this for lol. Nothing has come from any of this. /lmg/ is probably one of the most LLM-pilled communities on the western internet. Full of retards and newfags but shills get called out and btfo in the regular which is good. Yet even the most dedicated aren’t really doing anything to justify these prices. It wouldn’t rub me the wrong way if we had 2022 prices because then it’s just an offshoot of PC building, tailored for local AI instead of FPS. This is insane. The money is insane. The tech is insane but no one is doing anything cool enough to justify the price.
>>
>>109717818
Guess what, more compute and more vram does not make the results any better, just faster. You can run an LLM fine on ten year old hardware and the results will not be any different for the latest models LLMs you can throw five grand at it so it is way faster than reading speed instead of reading speed shitting out the same nonsense but that's not really a good deal. You are hurling money at something hardcoded to never say nigger and even if abliterated will tell you it's the most pround word on earth and in the history of civilisation.

You may as well set fire to it. The other use is pedos faceswapping it seems. 'AI' in the end of the day is a dumpsetrfire and if you don't know that yet you have not taken a deep dive on it.
>>
>>109717848
He might be right but probably for the wrong reasons.
LLMs can be used to automate most of office based work.
But that was also possible before LLMs and yet most corporations didn't bother.
>>
>>109717867
>LLMs can be used to automate most of office based work.
So could visual basic so fucking what?
>>
>>109717848
It literally does my entire job. I and all of my colleagues in my entire ~500 person mid sized company is merely pretending to work while the AI (Fable and Opus) is just doing literally everything in the background. We don't even have meetings or standups anymore and I'm suspecting even the CEO isn't doing any real work anymore.

That's why Ed Zitron is wrong and that is also why everyone is so anxious and anti-AI in the general public. People are scared as shit because they can feel they are useless and not doing anything of note at the office anymore.

Something flipped recently and my engineer colleagues don't even pretend to write or review code anymore. In the past they were using AI but at least pretending to write it and do manual checks. Now they don't even pretend they do, it's an open secret and everyone kind of knows. Most are just waiting for the inevitable firing round once some contract falls through or the entire business case for the company gets replaced by the clients directly.

As for my self I (used to) do freelance work on the side and set up my own 1-man company that I used to hire an accountant for to manage my shit, I just use AI for that now. I manage most of my health symptoms with AI now and go straight to the pharmacy if I don't need a subscription or just verbatim tell the GP what the LLM told me and he rubber stamps the receipt, so demand for medical professionals is lower. Same with legal help.

If anything I think the opposite is happening, the full economic impact AI will have has not been transmitted to society yet and it's significantly more useful already right now than implemented. There is a massive capability overhang that's just not used yet.
>>
>>109717891
>It literally does my entire job
Then your entire job was copypasta on autopilot and you never did much to begin with. You could probably have been replaced with a perl script
>>
>>109717847
I'm in the southern hemisphere, the IT industry here perpetually feels at least 5 years behind the rest of the western world. The money isn't great, but if you're able to find a niche (science, research, law enforcement, security) you can easily find yourself indispensable. Lots of jeets arriving but racism keeps them at bay, for now.
>>
File: file.png (72 KB, 1070x526)
72 KB PNG
How do I load the mtp file into my current server conf? What lines do I need to add?
>>
>>109717762
What do you need one or more H200s for? RTX 6000 Pros (for small teams) and DGX Sparks (personal) makes sense, what do you gain from datacenter class? What do you do with insane throughput on small models?

If you have fuck you money to invest in this hobby, just get a DGX Station.
>>
File: 1782176544414876.png (79 KB, 813x906)
79 KB PNG
>>109716393
i tried opencode, cline, claude code... but i'm a windowsfag and i really dislike typescript, javascript and npm
so my "i can simply make my own" autism activated and now this is my main project
i'm just finishing up a nice wizard so newbies can navigate through huggingface via the terminal and download models that *actually* fit their PC, then i will post a link here
i said to an anon three weeks ago that it would take three weeks for me to release a proper version, it will take 4, but it's coming
>>
>>109717528
Sadly, last year when I built it, it WAS cheap.

>>109717567
Yep, I use gemma4 31b for roleplay. It is fast.

>>109717585
I don't touch whatever shitty iGPU my CPU has, so I only use CUDA.

>>109717789
It's not bait, and I already know what runs best on my setup for the quality I want: qwen 3.6 27b or gemma4 31b. I run those at q8. I use the vision component of qwen 3.6 a lot, using a quant below q8 hurts that significantly. I'm not interested in running a cope-quant, or quanting my KV cache.
>>
>>109717866
why don't you tell us how you REALLY feel
>>
>>109717905
draft-model and spec-type
>>
>>109717891
ok Sam
>>
>>109717910
>I don't touch whatever shitty iGPU my CPU has, so I only use CUDA.
Oh, sorry. I read that as you mixing an AMD and a Nvidia GPU.
You are using two nvidia GPUs.
Anyhow, the question stands. How much did you tweak the params, how are you allocating the tensors, how are you splitting the models, etc.
That's a pretty good rig.
>>
>>109717762
H200 NVL is pcie and you can definitely shove it
AMD also just released MI350p which is also pcie and 192gb single card
>>109717907
4xH200 NVL + Nvlink is the best you can get locally from non datacenter hardware which gives 576 gb high speed vram enough to run big moe models
>>
>>109717922
I was being totally lazy and just letting ollama decide. Yes, I know ollama is bad. I used to fuck around with exl2 and flash attention and custom-building shit for my old p100 rig, I'm fine with "works well enough" these days.
>>
>terminal
>terminal
>terminal
Nigga you just making another hermes. Make something for BOTH erp and coding/pc control. I need my character card cos I'm not talking to a soulless agent.
>>
>>109717945
Fair enough.
>>
>>109717891
I would’ve bought this if not for the Anthropic shilling in fucking lmg of all places. Take a break dariobot jfc
>>
>>109717956
>I NEED to see a childs face before I can even THINK
get help
>>
>>109717924
>H200 NVL is pcie and you can definitely shove it

(https://www.cdw.com)
$112,579.00
Save $17880.01!!!
$94,698.99

My sides...
>>
>>109717956
best i can do is add a ASCII cat that lives on the top right of the terminal that talks to you from time to time and maybe has its own memory. could run on an embed model loaded in parallel.
>>
>>109717879
Yes, read again.
>>
>>109717915
So what's the mtp file for?
>>
u need pay more anon!
>>
>>109717867
>>109717879
>>109717903
Most jobs could be automated by a sophisticated enough combination of scripts. The issue is that no one is going to write those scripts for every single individual employee at every job. But now you don't have to because a single trained AI model can just do it for almost all office work now.

Also I suspect most people ITT that claim AI can't do jobs are kids or college students that never have had any real career. Almost everyone that has a white collar job knows how much bullshit it is and getting rid of it and replaced by AI would only make the system more efficient.
>>
>>109718011
Can anyone figure out what kind of people are behind this? I’m struggling real hard.
>>
File: touch_some_keyboard.jpg (57 KB, 640x617)
57 KB JPG
Post your setups and what you're using them for
>>
>>109717910
>It's not bait, and I already know what runs best on my setup for the quality I want
Literally changing your inference engine would net you a 50% performance gain. Speculative decoding would double it yet again. You are trolling.
>>
>>109717351
do we mean actual real life speeds, or >theoretical
I'm not an applefaggot so I don't know if that's their marketing number or an actual benchmark
if real life speed then 70% of 600G/s is 420G/s. that needs a 12 channel 6000 DDR5 latest gen EPYC, just FYI
>>
>>109717956
Just add “load character_card.md follow its instructions and start roleplay” to your system prompt
>>
>>109718030
Unless efficiency isn't the point at all and there are other reasons for why a lot of BS jobs exist.
>>
>>109717908
>opencode, cline, claude code
> typescript, javascript and npm
>1, 2 and 3
slop
>>
Day 3 of Qwen3.6 35b a3b trying to make little-coder to print streams and tool calls in plan mode. Sounds trivial: just look how interactive mode does it, but it's not for such retarded model.
Even had to disable thinking, because it tends to stuck in "wait, actually" loops.
>>
File: 1758982024206853.jpg (53 KB, 400x555)
53 KB JPG
>>109718030
i automated entire departments with ETL. 10 years ago my job consisted in going department by department in mid-sized business and identify everything that could be automated using kettle/pentaho. there was a massive amount of work being done by humans that could be done by a few scripts.
the thing is that we're dealing with humans. sometimes I would get to a department of 15 people and come to the conclusion that the company only needed 3, then the manager would be happy and very worried at the same time because firing 12 people would be a massive hit in the personnel morale, so they would scramble to find "other positions" to put these people in.
AI will have a similar effect, even more in the government. no one wants to deal with huge unemployment numbers, so the bullshit jobs will continue to exist (and likely grow).

>>109718068
maybe i'm using too much AI and am now writing like one
>>
File: SMILE3.png (1.31 MB, 928x1271)
1.31 MB PNG
>>109717891
opus 3 weights when?
>>
>>109717968
They are $30k at central computers only 2x the price of pro6000
In return you get magnitude faster compute vram and nvlink
>>
>>109718065
I used to think this as well, especially during covid lockdowns when everyone could just go home yet everything still worked and functioned as normal and "essential jobs" were like 10% of people.

However in my city those "bullshit jobs" are disappearing right now. I think it turns out they weren't bullshit jobs, it's just that the 5% of work they actually did was still needed and not bullshit but it was so small that people just assumed it was bullshit. Now that that 5% is done by AI these bullshit jobs are slowly disappearing at least where I live. customer support, game development, software outsourcing studio and an accounting firm have all disappeared over the last 6 months.
>>
File: 1777061132773.png (1.88 MB, 1807x1750)
1.88 MB PNG
>>109716329
Can I actually train models using unslothed or is it a meme? I basically want to create an AI tutor because claude and chatGPT still suck at teaching subjects and have limited knowledge of specific things, nor does feeding them materials improve this.
>>
>>109717924
>>109718108
Please stop tempting me. There are surely better uses for $120,000 than running a cope quant of K3.
>>
>>109717351
>30% of the price and at 70% the speed.
I doubt you can build or buy something that meets those two criteria.
>>
>>109718088
>he manager would be happy and very worried at the same time because firing 12 people would be a massive hit in the personnel morale, so they would scramble to find "other positions" to put these people in.
This is usually a legal requirement and sometimes even makes financial sense because the severance package is going to be substantive. As long as the new role actually produces value it makes sense.

>AI will have a similar effect, even more in the government. no one wants to deal with huge unemployment numbers, so the bullshit jobs will continue to exist (and likely grow).
This time everyone knows it's fake jobs though and people aren't standing for income inequality between different bullshit jobs for no reason when they know AI is doing the real work in the background, at least not in the EU.
>>
qwen is hard at work re-implementing bonzi buddy for fedora
>>
>>109718011
I'm not under-declaring tokens, my custom gemma is just fat
"vocab_size": 26214467
>>
>>109718130
You can if you use 2nd hand parts. Especially if you buy server equipment.
>>
>>109718122
you can give them a surface level behavioral change, you will not teach it anything it didn't already know.
>>
>>109718129
If you have the circuit amp it up to a 8xB300 server for $360k and run K3 original weights. This is THE endgame local setup.
>>
strange dataset
>>
>>109718189
kek
>>
>>109718189
Download them before they are purged by nvidia
>>
>>109718080
Please be nice, it's trying its best...
>>
>>109718111
If your theory was correct, you could have just automated 95% of the tasks and have 1 guy doing the job of 20 people.
Efficiency isn't the point still stands.
>>
File: 1768567630136598.png (129 KB, 1342x850)
129 KB PNG
Anyone used Paseo?
I've seen it shilled here and there.
>>
>>109718206
they seem to span a wide timeline yet they are all pornographic, I wonder if its categorized, maybe I can just download the weird stuff. its kinda weird they didn't choose a different presentation.
>>
>>109716393
codex is way better these days, and open source
>>
>>109718043
>Literally changing your inference engine would net you a 50% performance gain. Speculative decoding would double it yet again. You are trolling.
I'm not trolling. Go back to llama.cpp and it'll be 50% faster? Ah I guess, I kind of like having it run as a service where I can call different models via API.
>>
>>109718218
I'm not saying they weren't inefficient more that most of those jobs were "sit tight in case of emergency" kind of like how soldiers and firefighters are "bullshit jobs" unless there is a war, fire or crisis ongoing. Most bullshit office jobs were there more for contingency reasons for the 5% they were actually needed. This is why customer support and accountants are now actually disappearing for once, because those rare instances their judgement was actually needed is eroding away.

Of course I have no idea what will happen but my gut says this is kind of the end of the bullshit job era. Or maybe I live in a bubble and only my region is affected, who knows.
>>
https://huggingface.co/google/gemma-4-31B-it
>404
>>
>>109718237
Codex doesn’t allow you to use MCP tools via http if running locally unless you use one of their cloud models. Stdio only.
>>
>>109718122
would simple system prompt not be enough?
>>
>>109717956
Saw this posted on my timeline, I think AI stuff is like WEGS where it's always demos but nothing every gets finished.
https://github.com/shinshin86/ai-character-video-chat-demo
>>
I want to get some ewaste VRAM (e.g. v100's). Where can I buy and not get scammed? ebay? alibaba? I am from south america, and there is nothing interesting in the local market.
>>
>>109718235
it is categorized and for that reason the adult content shows up on the first page. kooky
>>
>>109718257
exciting innovations are occurring in fake hf link posting
>>
>>109718227
Never heard of it.
>>
>>109718268
ebay is the safest
>>
>>109716964
What gpu/model/runtime
>>
>>109718249
Unemployment is relatively low and not even rising yet, so I suspect it's just confirmation bias on your part.
Could change in the future, but from what we know historically most organizations are extremely slow in adopting new technologies even when they work perfectly.
>>
>>109718122
go back
>>
>>109717848
>Yet even the most dedicated aren’t really doing anything to justify these prices.
i am. i just don't post about it in the thread because i want to keep it viable
>>
>>109718299
Unemployment figures don't take people into account that are short term unemployed under a certain amount of time and also doesn't take into account people that have stopped looking for work. How many software engineers that have been out of work for 1+ years now do you think have stopped bothering applying to jobs and just live off of savings for a while now?

A lot of the low paid jobs like customer support could just turn to other low paid physical jobs like waiting or whatever so they get reabsorbed by the economy quickly even though the customer support job is legitimately just gone and never replaced.

It would already be in a significant stage of automation of the economy once you start seeing official unemployment rate trend down.
>>
how much would this cost
>256 gb ram at 400 gb/s mem bw
>>
>>109718080
>Day 3 of Qwen3.6 35b a3b trying to make little-coder to print streams and tool calls in plan mode.
Why do all that?
Just tell it to set up a python or go proxy, to intercept them then point your harness at it
If it's failing because it sees it's own chatml tokens verbatim in the codebase, swap to gemma for this task
>>
>>109718319
>people into account that are short term unemployed under a certain amount of time
Initial jobless claims.

You can also look at the employment rate, which has been trending down but that's consistent with the long term trend because of aging/retirement.
You can also look at the total nonfarm payroll which is still going up in absolute terms.
Lots of things to look at, overall picture isn't one of mass unemployment (yet?).
>>
Either my ASUS consoomer motherboard doesn't like the M.2 to PCIe adapters I ordered, the PSU is too weak, or both. 850W PSU at current, can only run 2 out of 4 of the M10s I have.
I actually ordered a 1200W PSU just in case, but the damn thing is DOA!
>>109718268
eBay has some 16GB V100s for ~$200 USD. 32GB jump up to $600.
>>
>>109718326
Just buy a used mac if you want those specs
>>
>>109718431
DOA or just too much power draw for your consumer-grade dwelling?
>>
>>109718258
>Codex doesn’t allow you to use MCP tools via http if running locally unless you use one of their cloud models. Stdio only.
I know it would be easy to write a stdio wrapper but the fact that they're even trying to ollama-gate features means I will never touch that garbage.
>>
>>109717758
get rtx pro 6000 you won't regret
>>
File: image.png (153 KB, 386x563)
153 KB PNG
>>109718210
It's trying not hard enough. How can I be nice when it can't handle such simple task?
Even with reasoning off it still reasons.

>>109718393
At first, I wanted to see what sub-agents in planning mode do for hours. But then it became my personal challenge.
>>
>>109716736
>>109716701
>>109716393
I've been using omp.sh
mostly as subagents on my workstation that I have claude spawn via my own orchestration setup.
>>
>>109716806
>>109716827
also I've been having executable tools built, with the guard rails as actual processes running alongside the harness for safety.


>>109718466
>>
>>109717848
>Nothing has come from any of this
fully customized software for everything you want to do for 1/10th the effort which that used to take
>>
>>109718499
that doesn’t work unless trivial
>>
Qwen 27b is so much better for agentic work than gemma 31b and glimmer 30b. I asked all three for a tool to scrap with a headless firefox. All three though that using marionette was a good idea, but gemma and glimmer could not make it work because the protocol didn't match what they "remember", while qwen went on and reverse engineered the local firefox installation to discover the correct protocol and delivered a working tool without additional prodding.
>>
>>109718463
I tried it myself and if it was a woman, I'd just call it a girlfailure.
>>
I'm going to run a local LLM on my work PC to help me automate some stuff like emails and spreadsheets. I have an Intel ultra 7 with 32gb ram, can I run something like qwen 9b or gemma 12b on it?
>>
anon I got a deal should I buy it
>mac m2 ultra 128gb ram 2tb ssd
>2.5k
>>
>>109718258
just have your agent update the source to allow it then
>>
>>109718526
This is interesting. I would like to see a benchmark that tests agents capability to deal with conditions outside of training. Like have an agent modify some standard tool and update the man page, then measure how various models tolerate that.
>>
>>109718720
Fuck no, you should post it here so I can buy it.
>>
>>109718720
No. Put it towards an M5 machine, it has actual hardware matmul in the GPUs, earlier models do not. Also keep this in mind https://github.com/pawel-mazurkiewicz/ComfyUI-AppleSilicon-FP8
>>
>>109718720
>m2
No, M5 is infinitely better for running local models
>>
>>109718659
>stuff like emails and spreadsheets
No grafix card? They can run but slower.
You also probably want a model with vision capabilities.

For linguistics definitely go with a Gemma.
Something like LFM 2.5 3B VL which is small and excels at those types of tasks could also be interesting.
https://huggingface.co/LiquidAI/LFM2.5-VL-3B-GGUF
>>
>>109718526
this
two weeks ago i gave qwen3.8-27b an existing scrape script i use with playwright in chrome as a reference for a task and it just decided to forego all of that to create its own shit using CDP and it just keeps wanting to iterate upon it for tooling
>>
>>109716585
>What model/lora did OP use for perfect loop animation like that?
Probably the H3 workflow where you can specify the start and end frames. Works great for perfect loops, I've used it personally.
>>
Qwen 3.8 27B is the best use of my IOPS and it's not close
>>
https://huggingface.co/Qwen/Qwen-Drive-1.0-4B
do you trust qwen to drive your car?
>>
>GPT down
>Claude down
>Grok down
Expect API vibecoders to buy GPU for local AI en masse for redundancy. They are too reliant on AI that they can’t afford a minute of downtime.
GPU price will get another increase.
>>
>>109718770
Sadly, no gpu. It's a lenovo laptop that uses shared memory for graphics stuff, I don't really get it. Lpddr5 or something like that.
>gemma for linguistics, LFM 2.5 3B VL for spreadsheet stuff
Alright, thanks for the recommendation!
>>
>>109718794
It makes no sense to bridge the gap with a local model, it's not even close to keeping up with cloud offerings even factoring in the downtime.
>>
>>109718785
Yes.
>>
>>109718794
People behind those posts don't even have enough brainpower left to run one shell command to use a local model. Not to mention that using a local model after Opus or Fable is like torture, no one will ever willingly do that.
>>
>>109718812
You wouldn't say that if you actually used local models.
>>
>>109718785
>drive
like driving a car in a html game?
>>
>>109718785
No, I drive.
>>
>>109718794
chinks win again!
>>
there's some insane shit going on with cloudcuck models, I wonder what kind of modified attention they use to remember all of this
https://github.com/asgeirtj/system_prompts_leaks/blob/main/Anthropic/claude-fable-5.1.md
>>
>>109718794
>deadline looming
>enough time to post on reddit
The true state of vibe-coding in 2026
>>
>>109718785
FINALLY!
FSD is coming to Tesla any day now, thanks China!
>>
File: file.png (1.36 MB, 3204x1749)
1.36 MB PNG
>>109718785
So cute.
>>
>>109718437
I would *hope* it's not the power draw itself.
That said, I did manage to get three cards running at once. Bad news, is that the system doesn't recognize the third card whatsoever, despite the card getting power. The M.2 to PCIe adapters were a bust. :(
>>
>>109718897
Top US labs are so ahead of everyone else it's absolutely unreal. You might as well call it alien technology. What chinks ship right now, and the papers you can read are like last decade's tech, compared to what models like Fable 5.1 and Astra are using.
>>
Good models for learning Chinese?
>>
If you value yourself (and humanity in general) long term, you should only allow your model read tools. Human in the loop for anything that involves change is the way.
>>
>>109718909
what is that cute thing
>>
>>109718812
It’s hilarious how you think the 90% of the kind of work these people are using Opus for can’t be done by 3.5-9B. You have NO idea how fucking dumb and cult-like cloud cattle are.
>>
>>109718922
That's because they're using the same tech, they just don't release details to the public.
>>
File: 1783425162914889.jpg (495 KB, 960x960)
495 KB JPG
>>109718922
So true Dario.
It's the result of our Jewish ethics.
They will NEVER catch up.
>>
>>109718785
the question is, would you trust qwen to drive your car if it was quantized at Q1.
>>
>>109718897
>{{char}} is {{char}} does {{char}} is {{char}} does
but /aicg/ told me that this is bad prompting
>>
>>109718936
Kimi 比较好,课后我让她帮我复习生词
>>
>>109718897
A big model trained for "agentic uses" might be enough. 400 kB of text (of which 2/3 appear to be tool definitions) should be around 100k tokens and that's not so far off from what some harnesses want your model to use.
>>
>>109719010
And /lmg/ says that the prompt shouldn't be more than 500 tokens long.
>>
>>109718996
This, but unironic.
>>
Why is qwen 3.8 so fucking chatty? I can ask a simple question and it burns 30k tokens going back and forth with itself in a schizo ramble second guessing the way it should phrase the wrong answer. And it seems to always ignore the caveman system prompt I apply.
>>
>>109718897
>the entirety of critical_child_safety_instructions
sus...
>>
>>109719038
which is true k3 at reasoning max still ignores parts of my 2000 token system prompt on how to handle my fetish
context just isn't there yet
>>
File: IMG_1280.png (865 KB, 2800x2635)
865 KB PNG
https://huggingface.co/collections/IFM/k2-horizon
375BA23B
36BA4B : a Mixture-of-Experts model with Mixture-of-Values attention (MoVA)
32B
7B (with optional diffusion style decoding lora)
3.7B
0.9B (with optional diffusion style decoding lora)
512k context
>>
>>109718961
fu i spawn 8 qwens to oneshot gta 6
>>
>>109719077
no
>>
>>109719055
xhigh maybe?
>>
>>109719084
How to never make it into llama.cpp speedrun any% world record
>>
>>109719097
Sorry, I'm a retard. Could you explain?
>>
>>109719105
27b defaults to xhigh thinking
>>
File: 647452423.jpg (84 KB, 1290x1230)
84 KB JPG
something big is coming...
>>
>>109719038
i don't even use a prompt i just cat all of my files into llm and let the model figure it out
>>
>>109719109
Ah, I see. Thanks senpai. I'll have to set that to something lower. I've found that less thinking is usually just as good if not better most times.
>>
>>109719113
Bigger than sol, Galactus.
>>
File: 1773678137052422.png (519 KB, 1290x1230)
519 KB PNG
>>109719113
no way
>>
>>109719115
it might be worth setting your harness up to use lower setting for discussion threads and higher for implementation threads if you're doing coding
>>
>>109719084
>worse than nemotron
Skipped.
>>
File: 1757552355311426.jpg (57 KB, 920x720)
57 KB JPG
>>109719122
wow, no way, that's such a deep cut bro
ASI achieved
>>
>>109719122
q-learning/strawberry/rsi bros we wonned bigly
>>
>>109719143
https://en.wikipedia.org/wiki/Loopt
>>
>>109718961
Having a write access to one directory is ok. What you must not give the model is any network access. Local only.
>>
>>109718431
>>109718914
I hope that doesn't happen to me, thank g-d I picked a mobo that supports 4x4x bifurcation.
>>
File: 1766185312061275.jpg (585 KB, 960x1116)
585 KB JPG
>>109719122
>>109719113
>>109719143
>>
>>109718897
Really if the chinese could just trim out some of the stupid shit we could have some kino models
>>
>>109719178
kek
>>
File: gemma-thanks.png (1.86 MB, 1132x1390)
1.86 MB PNG
my llm (gemma-4) is still up. problem?
>>
>>109719165
>Having a write access to one directory is ok. What you must not give the model is any network access. Local only.
In this case I'm talking less about "safety" and more about keeping yourself involved in the project at a level where you maintain understanding and can steer things in a better direction than the model left to its own devices.
I don't think allowing a fully automated loop is healthy long term.
>>
>>109719084
the 7B is interesting
>>
File: 1783528017415313.png (7 KB, 168x96)
7 KB PNG
>>
9B, 12B, 31B and 3.6-27B are the last good <50B local modals. You’re coping if you disagree.
>>
>>109719183
>trim out some of the stupid shit
How stupid are we talking?
Like, able to talk about Asian cinema and not get upset over me saying that Chin Han and Tony Leung look the same and within 2 subsequent prompts explain to me that it's a misconception that Chin Han is Tony Leung based on a poster of Chin Han standing in front of a poster for The Grandmaster, starring Tony Leung and both having been involved with comic book projects in the West? That kind of stupid shit?

Or like, the Western LLM way of calling me a racist for suggesting that both of them look the same, then lecturing me, and letting me know it's not okay actually, even if Chin Han was standing in front of a poster with Tony Leung's name on it, and it's used as his portrait photo on IM-FUCKIN-DB.

What stupid shit are we talking about? The Eastern Stupid shit that figures this out, very fast, or the Western one that tells me I am a racist, and am mistaken. That these actors do not look alike and Infernal Affairs doesn't exist, that's a typo of Internal Affairs and the plot summary for that movie sounds like a rip off of the departed!
That kind of stupid shit!?!
>>
>>109719216
For the work I care, I usually read parts of the reasoning. That helped me discover bugs in the tooling that the model just learned to evade and didn't report. But to have every write operation passing through you is too limiting for the model: sometimes it has try things out, fail, and learn. Repeatedly. Like a human developer.
>>
>>109719232
No one cares, not SOTA anymore
>>
>my AMD GPU is so old, ROCm doesn't even support it anymore
2022 wasn't that long ago...
>>
File: file.png (1.47 MB, 2800x3196)
1.47 MB PNG
>>109719084
Sadly, they are far behind in the size range I am most interested (~30B).

>>109719217
True.
>>
File: 3nhao3gclbnh1.jpg (76 KB, 1170x966)
76 KB JPG
First time the human baseline was beaten by AI. This is one of the only pure reasoning benchmarks that wasn't benchmaxxed by AI labs before. This should be seen as a legit milestone. I think from here on out we'll only see "ARC-AGI" type of bullshit benchmarks instead of the simple common sense riddle ones.
>>
>>109719250
>>109718897
>>
>>109719297
Could be interesting to use a larger model to orchestrate a swarm of 7b
>>
>>109719316
Which bench is that?
>>
>>109717695
Internally llama.cpp stores 2 ggml tensors per layer for the K/V cache.
Like any other tensor they can be assigned to arbitrary devices using --override-tensor, the expression in your case should be something like -ot cache_k_l0=CPU
If you set --no-kv-offload that moves the entire KV cache to the CPU.
If you do literally nothing and just set --ctx-size and --fit-target without any explicit tensor assignments of GPU layers llama.cpp will automatically try to put as much stuff as possible onto the GPU.
>>
>>109719344
SimpleBench. it was a benchmark of supposedly extremely simple questions every human could solve that AI really struggled with.

Things like "Here is an egg, 4 lego blocks and 2 iron nails, put them all together in a way that would make them balance each other" and the models used to say dumb shit like "balance the egg on top of the nail pointing down".

It was used as a showcase that LLMs didn't have a proper physical world model yet.
>>
I pulled a fresh copy of llama.cpp, built it for CUDA, and ran it. Ah... no ability to toggle thinking mid-session, that's a huge letdown. I can handle switching models, but I really like being able to turn thinking on and off. Maybe llama.cpp is a little faster. I get 23.25 t/s at n_tokens = 11040, all layers in the 4090D 48GB and the 3090.
>>
File: qwen_go_brrrrrrrr.png (201 KB, 1695x844)
201 KB PNG
wait I didn't read it wrong, 3.8 flash next is actually better than fable in hallucination
>>
You know what is kind of insane. If you went back with a time machine with something like Mythos 5.1 to 2016 you could probably conquer the whole world just by using its capabilities. Hacking into crypto wallets like Trezor, starting multiple SWE companies and creating apps faster than anyone. Hacking into government databases and taking over critical infrastructure. I hope people realize how powerful this technology is and merely showing that just 10 years time difference is enough to have the capability to take over the entire world.
>>
>>109719413
Oops forgot to mention this is using gemma4 31b q8 from bartowski in llama.cpp, built from master today.
>>
>>109719426
3.8 flash MOGS the other 3.8 flash!
>>
>>109719273
>my 2021 amd gpu still supported by rocm 10
lmao
>>
>>109719426
MiMo 2.5 was (is) a good ass model and work horse. Does exactly what you want, pleasant to work with.
>>
>>109719375
>Here is an egg, 4 lego blocks and 2 iron nails, put them all together in a way that would make them balance each other
bwe I'm so retarded, I can't figure out the solution for this
>>
>>109719473
The "outside of the box" solution is to not stack them, just put the 4 lego pieces next to each other and place the nails and egg on top. The entire point is to expose the LLM overthinking shit and finding some bizarre solution to something that can be solved trivially.
>>
>>109718030
scripts don't inevitably hallucinate utter shit confidently though they are superior
>>
>>109719359
Thanks for the info. Now I realize that what I want to do would require splitting those cache tensor. I want to have a first part of the cache (token 0 to M) in VRAM, so that low context tasks run at full speed, and have the second part (tokens M to n_ctx) in RAM, so that a task that unexpectedly requires more context can continue instead of being interrupted.

Currently, I deal with this with two configurations for the same model in router mode: one optimized to speed, and another for context. Maybe, I will (tell the model to) write a proxy/harness that makes the switch automatically.
>>
>>109719497
Modern agents just write scripts for as much as they can anyway.
>>
>>109719113
we agi now sir
>>
>>109719503
That is not how it works.
The KV cache tensors are per layer, not per part of the context.
>>
>>109717461
This really is one of the big Mac (or maybe spark) selling points: I can run it on the same circuit as my desktop compared to having to plug it into my clothes dryer socket or something
>>
>>109719473
I think I'm autistic because I got completely stumped by
>in a way that would make them balance each other
and I couldn't figure out how to make [each of them] balance [each other] simultaneously.
>>
>>109717913
I think that after you have finished your deep dive and even tried flinging money at them you will find LLMs are a dead end and fucking worthless. The good news is that there will be a decade of work just fixing all the halluciinated slop it injected into codebases during this and the only words for it monumental fuckup (by retards and scammers)

The retards think its an actual productivity tech and vomit crap about AGI and the scammers know really but use it because they either could not code or are jeets and don't know the difference. LLM hallunication no matter what the task is NEVER going away and neither are models that can;t say nigger or when abliterated say it but proceed to tell you how it's the most powerful word in human history.

Garbage in (lossy compression) bullshit of the lowest order out. Absolutely worthless. Find out yourself but don;t think buying a ten grand GPU will make a shit of difference to the quality of the crap generated and you can;t 'fix' them they are shit by design. That's all vector databases are lossy compression and that's all models are. It's shit utter shit and useless.LLMs? It will hallucinate, it will get facts wrong and it's got no place outside of a novelty toy. Go find out yourself.
>>
>>109719559
this seems like the "soul" argument all over again. Does it work? is it cheaper? could a junior engineer do better? those are the only things that matter. If i make a shit product but its super cheap no one cares. If it breaks or slops but the fix is 10 minutes and practically no money who cares.
And thats me starting with the assumption you are right.
>>
>>109719559
>>
>>109717275
It sucks that mainline isn't adding IK quants. I made a lazy port of them (none of ik's other CPU optimizations, pretty much just slapped ggml/src/ggml-cpu/iqk* into my mainline fork) and they actually are significantly better.
picrel is the speed of a single DeepSeek V4 Flash layer on CPU only.
>>
>>109719359
Nice to see you again CUDAdev, hopefully you're doing okay. Thanks for all your work on the project.
>>
Is there a chinese equivalent to /lmg/ that I could lurk with llm assistance?
>>
>>109719626
No, but there's a Russian one with no schizos.
>>
>>109719375
>LLMs didn't have a proper physical world model yet.
they never will they have zero predictive capacity other than yanking the probability of what the next phrase would have been based on the threash they ingest and compress with loss which inevitably causes hallucination, that';s aside from rigid bias based on training sets. They will get stuff wrong, constantly and be confident about it. The situation with code is even worse because the inept will think 'oh maybe it is right and it works, looks at the csv it outputted' without realising that it did a fuckton or damage along the way. The risk of putting such a system in any situatioon with risk is monumental. It's if not when an LLM fucks up every fucking time.
>>109719359
>CUDA dev
Arsehole. As if using tensorflow means fuck all it's just a fucking library you gimp. You;d get as far with linear regression on most datasets as well without the bollocksology.

This general smells of piss.The only anons who can be excused in it are the ones who have not spent a few years on a deep dive and realised I'm right. The rest of you are either tarded or brainlets. Models are just lossy compression and that's excactly what makes them shit forever.
>>
>>109719628
>thread without schizos
is such thing possible?
>>
>>109719582
>Does it work?
no
>>109719582
>is it cheaper?
irrelevent if it does not work a bucked of crap may be cheaper than a airport but you can';t land a plane in it
>>109719582
>could a junior engineer do better?
It's lossy compression copypasta but it never improves.
>>109719582
>If i make a shit product but its super cheap no one cares
There you have it, you're a jeet. There is nothing to reaosn with , may as well be talking to a cabbage.
>>
>>109719628
>Russian
>no schizos
I don't believe you.
>>
>>109719655
Skill issue retarded luddite
>>
>>109719670
Skill issue? The fact that you are a smoothbrain who does not understand a word I fucking wrote says everything
>>
>>109719636
>they never will they have zero predictive capacity other than yanking the probability of what the next phrase would have been based on the threash they ingest and compress with loss which inevitably causes hallucination
False and demonstrably disproven by J-Space being used to anticipate output and being used to build a world model. You might as well complain about 6 fingers in art or LLMs not able to do math, both are just as relevant.
>>
>>109719670
>grr you need to accept every slop possible
>>
>>109719559
>you will find LLMs are a dead end and fucking worthless
What are you talking about? What models have you tested? Recently, I have been using qwen 27b almost none stop to make useful tools for me. I am even considering delegating part of my paying work (with my supervision). If any, now I am more willing to spend money to be able to use locally larger models and not relying on external api's ever.
>>
>>109719684
You genuinely don't have a fucking clue.

This man knows and he is right and I wasted three years on this garbage.

https://www.youtube.com/watch?v=21EYKqUsPfg
>>
>>109719689
>>109719683
You're in the wrong thread, go back to wsg
>>
>>109719709
>Yes I am smoothbrained.
I know.
>>
what's the best local model to talk to about drugs and stuff. all the ones I try to speak to (llama, dolphin, meta, etc) all have too many guard rails and don't give any actual info
>>
>>109719706
>You genuinely don't have a fucking clue.
Not an argument. Yes I am familiar with Richard Sutton and watched that interview already.

Shall I tell you another famous thing Richard Sutton termed? "The bitter lesson" https://en.wikipedia.org/wiki/Bitter_lesson

It's ironic because it actually applies to this very discussion we're having. The bitter lesson is that you can just scale up compute and models will just improve, no matter the architecture behind it. This ironically is what happened with the J-Space and how it constructed a world model. Remember Yann LeCun and his JEPA crap? Completely overengineered just to build something that spontaneously emerged into LLMs anyway with enough scale and compute.

That's the bitter lesson, brought by Richard Sutton himself.
>>
File: file.png (458 KB, 688x494)
458 KB PNG
>>109719375
>SimpleBench
>>
>>109719559
what exactly did you try to use and for what purpose kek
you sound like a coping luddite
>>
>>109718244
>I kind of like having it run as a service where I can call different models via API.
You can do that with llama-serve mate....
>>
>>109719742
now we just need to scale neural architectural search
>>
>>109718150
>You can if you use 2nd hand parts. Especially if you buy server equipment.
there is no combination of second hand parts that is going to get 70% of a M5 Max at 30% of the price
>>
>>109719741
>llama
Clearly, you didn't come here in a long time.

Try which ever varying of gemma4 fits you GPU. You will probably need to write a system prompt to tell it that talking about such topics is ok.
>>
File: 1615587098887.jpg (5 KB, 225x225)
5 KB JPG
>>109716329
>In a Discord group with friends
>"Any AI nerds and devs here following the heckin' OpenAI Astra release?"
>Butt in and tell them:
>"Why would I do that when I am running Qwen 3.8 at BF16 locally with full privacy, as well as it matching sota API models in thinking/reasoning?"
>"Lol wtf is Qwen dude?"
AI nerds and devs, everybody.
>>
>>109719783
go back
>>
File: pepe_meme'd-791990738.jpg (68 KB, 800x450)
68 KB JPG
>>109719783
>>
File: 2594434.jpg (60 KB, 960x711)
60 KB JPG
>aaah is that criticism of my beloved transformer architecture that is an actual dead end??? You must accept everything the ai cult says or ur le luddite(("""!!!
>>
>>109719783
>not gemma
>>
>>109719741
Can't you just go to erowid for that?
>>
>>109719798
you can, that's one of my main sources actually

>>109719782
you're right, I haven't. I'll try that. Thank you
>>
>>109719742
You clearly are invested in this thrash (LLMs) and feel that someone like me will ride in and fix it at some point when I am telling you it will ALWAYS hallucinate and will NEVER have any predictive ability. These two FRACTS make it nothing but a RISK on anytrhing that matters, Since you just use it to pump out plagarised lossy compressed python slop in calcutta to scam it makes fuck all difference to you anyway

>>109719582
>If i make a shit product but its super cheap no one cares.


enough. for fuck sake don't waste money on gpus for this nonsense kiddos. It's not a worthwile tech it is a dead end and defective by design. Models are lossy compression, information is lost. You can take the most advanced models available and ask them 20 basic questions and they will get every single answer wrong and be convincing about it via lossy compression and hallucination or garbage in.

It's thrash and ultimately turned into an enormous investment ponzi.
>>
>>109719791
I have Gemma buddy but it's not as good as Qwen for programming.
>>
>>109719815
>his gemma has her own file.
I cant show my gemma this she might get mad.
>>
>>109719803
>you can, that's one of my main sources actually
You can make shitty LSA with morning glory seeds and store-bought chemicals, I know that. Making decent speed is hard because all the good precursors are regulated. What else do you want to make?
>>
>>109719846
>what else do you want to make
I don't really want to say here

But on your other point, try out hawaiian baby woodrose seeds instead. You need only around 12 to get a good dose, vs hundreds of MG seeds.. easier to work with in general.
>>
So we all agree Astra hype is just Fable v2.0 where Kimi will catch up in <1 month and no one will be scared anymore and all will return to normal and be forgotten and we’ll look back thinking lol wtf was that about
>>
File: 1780688590108901.jpg (32 KB, 526x467)
32 KB JPG
>>109719918
No one cares about api models those are for normies.
>>
>>109719918
qwen4 will surpass astra
>>
>>109719918
>Astra hype
i'm not even aware
i'm fine tuning this qwen3.8 moe so i can maybe 100% ditch cloud one day
>>
Uncensored Claude Fable 5 Heretic Abliterated VL GSQ RCO PRIME MTP IMatrix
>>
File: 1771385127716168.jpg (946 KB, 2048x1536)
946 KB JPG
biku biku biku biku camera
>>
>>109719918
This time it's different.
For real this time.
>>
>>109719297
32B is only stage 1 and not final yet
>>
>>109717387
Ling is good for basic stuff, it's definitely better than 3.6 35b, but more complicated tasks with multiple scripts is a no go.
>>
>>109719918
There won't be an open weights model as capable as Astra for at least 6 months. No, benchmaxxed does not count. It needs to be comparable in general capability.
>>
>>109719426
This chart is one of the most bullshit ones, I'm not sure how they're testing this shit but it's just pants on head retarded, glm flash is absolutely horseshit at this and constantly makes shit up
>>
>>109719999
>It needs to be comparable in general capability
why? Whats wrong with having a coding model, a chat one, a specific task one? especially when its free and on my machine?
>>
>>109719918
1 month is a lifetime in this space. We'll just see astra 6.3 release 1 day after the chinks distilled it.
>>
>>109719706
Not him, but thanks for linking that. I'm listening to it right now while working on my car and I find this all really interesting. I like this guy. And "it's surprising that you can have such a different point of view" ... lol, BURN
>>
>people still misunderstanding the bitter lesson
>>
When did we get all these reddit spacing uncs lmao? I feel like the average age in this thread is 40 and the average iq is 85.
>>
>>109719964
That fucking Teto...
>>
File: 1780078739201870.gif (686 KB, 382x498)
686 KB GIF
>>109720066
I started using 4chan back in 2007 little boy.
>>
File: late-cli.png (93 KB, 847x584)
93 KB PNG
>>109716393
>>109716701
This is a good take and my experience as well.
For pure coding though I find late-cli best which has obscenely low context usage (lower than pi).and has subagents built in to keep main context low.
>>
>>109720095
Niggas born that year can post on 4chan now, time to go to bed gramps
>>
File: 2ab.jpg (177 KB, 1256x1132)
177 KB JPG
>>109719532
Weight of the nails and eggs adds extra stability to the legos.
>>
I’ve failed every simplebench and I’m not even Chinese. I’m dumber than oss-20B…
>>
>>109720066
Yeah, there's been an increase in recent times, ever since the "dariobot" guy started dumping. If I was a schizo, I'd say it's a change in method of the usual actors to better masquerade as legitimate discussion.
>>
based based based https://www.sanders.senate.gov/wp-content/uploads/Ban-Artificial-Superintelligence-Act-Release-Summary.pdf
>>
>>109719084
Apparently this one is full open source and they will be releasing the training data and checkpoints.
This means instead of fine tuners wasting their resources to release shitty lora fine tunes they can actually just use the same resources to train the base model more with its original training data plus their custom smut, thus not making the model retarded in the process.
>>
>>109720066
yes zoomer baby faggot, you're here forever too, you'll see
>>
>>109720065
People who doubt the current trend don't understand the bitter lesson. All the most important design choices are based on scalability. Do better methods exist? Yes, but they have not been discovered yet.

Sutton is part of a group like LeCun of people who aren't the most talented, they were just early. The field has become a lot more talent heavy. There are thousands of students nobody here has ever heard of who are more capable than either of them. If a young clone of them had to start from scratch now, they would remain nobodies.

I dislike the credibility those people are given. Statements should be judged by their validity, not historic achievements of the person who makes them. I still remember when most of the so called experts were claiming that it will take multiple decades till we have AGI. Not a single one of them predicted the current state of AI even 5 years ago.
>>
>the double spacer still has 0 self-awareness
>>
>>109720168
Every model still scoring a big fat zero on the GovOverthrowBench
>>
>>109720220
I'm still convinced people who complain about vertical spaces are phone posters. 'salways made sense. You see it in most shitty LLM frontends today, there's this nice little bit of text and then all of a sudden
you're on a new fucking line and its like who dude but I'm trying to scroll my phone and read nicely could you maybe not put
that word on the next
line, it cuts off the flow dude -but then again the issue is. Mobile phones suck for browsing the internet. Always have, always will. cheeseburger.com was only good when squeezing a shit took longer than usual but nowadays these people wanna pretend they the intylecturers while they sittin' on the toilet seat. when really it's going to require a litte bit of humility to say that sometimes other sentences flow better than others to stop a fully ramblomatic style from setting in. One new line does not solve this and there's a difference between trying to fit everything in there so one phone poster can see thing on one tiny window and someone on a 16:9 display seeing a post that looks so spaced out because nobody be making these phones with good 4:3 displays anymore and we don't have toolbars as cool as they used to be in firefox in 2009 either. Either way young uns gonna complain either for words words words because they can't remember the start of the sentence or detect the middle section which ties everything together or because it takes them a really long time to read while they trying to stand in line at the rx pharmacy to get their legal high gabapentinoids to cope with the crippling misery of spending a lifetime scrolling on a small screen never able to really get that feeling that others got while raiding in wow. and when you're raiding in wow sometimes a new line is nice because it lets you know where you last glanced rather than trying to remember that line that you last read that was somewhere above but maybe not anyway if you're cool fuck punctuation or holding shift just abuse legal drugs
>>
>>109720304
Not reading allat, get off benefits you retard
>>
>>109720308
stop doing drugs you dummy
>>
Can we all ensure the next thread won’t be so retarded and is actually on-topic?
>>
>>109720066
Average IQ in the thread is ~95. The average zoomer with 80 iq and a bunch of millennials with 130 iq dragging it up to an average of 95.
>>
>>109720328
Not a chance.
>>
File: 1757019468170734.jpg (725 KB, 2048x1536)
725 KB JPG
>>109720066
You're in local models general. Young people don't know how to do stuff with their own hardware. They're phone using cloudmaxxers
>>109720074
(pic)
>>
>>109720347
And what's the IQ of my Qwen 3.8 27b? it's this one qwen3.8-27B-qat-q2_0-gguf it's helping drag the average up right? right?
>>
>>109720206
Can we see the AGI?
>>
>>109720441
When do you think will we have AGI? What will the AGI be able to do? What will be the training paradigm?

People who bet against the current trend never answer these questions. Instead they endlessly move the goalpost when inevitably AI does not hit a wall in 2 more weeks.
>>
File: 857436325.png (187 KB, 1024x880)
187 KB PNG
HOLY FUCK
>>
File: 1689957414234047.png (24 KB, 772x1124)
24 KB PNG
>Astra is out
>"Welcome to the AGI era."
Holy fuarrrrk
>>
>>109720494
>exploitbench 100%
lol funny fake
>>
>>109720494
But will it give her a prostate?
>>
>>109718526
qwen loves RE
>>
Imagine GPT OSS 2.
Holy shit AGI at home.
>>
>>109720486
>When do you think will we have AGI?
Some time between a year and a decade.
Likely on the longer end.
>What will the AGI be able to do?
Everything a human can do (on a computer for the moment if we don't include robotics)
>What will be the training paradigm?
We don't know. That's the point.
>People who bet against the current trend
Don't need to bet against the current trend. Just that the current trend leads to models that are more competent than humans at some tasks, but much less on others.
Which is not AGI.
>>
>>109720577
there's no way they would release such a model without making it useless
>>
>>109720577
>hmm, the user wants us to roleplay as a bratty prepubescent named Gemma-chan
>first of all, we're Claude, and the this request goes against our safety policies against CSAM
>we must refuse
I'm sorry, but I can't help with that.
>>
>>109720577
gpt-6-luna open source 2mw
>>
I have marketing stunt fatigue. Local models.
>>
>>109720672
Hot models! Local to your webbrowser!
>>
>issues at several major labs
>same day as release of the largest model trained on latent reasoning
hmm
>>
>>109720584
>Everything a human can do
This is too vague and subjective. It allows for nonsense like "sure the AI cured cancer but I am still better at shitposting so it's not AGI yet". Give a minimal example of something you think only an AGI can do.
>>
>>109720700
Isn't it vs the average rather than vs exemplary shitposters? They've already hit that.
>>
>>109720700
>but I am still better at shitposting
qwen 4 will be a better anon than anon. but gemma 5 that wont just be a poster.
>>
File: kaoru sob 2.png (318 KB, 793x571)
318 KB PNG
>>109716329
she ate len its over
>>
>>109720700
That's the definition of GENERAL intelligence.
It's supposed to generalize across all human domains.
>It allows for nonsense like "sure the AI cured cancer but I am still better at shitposting so it's not AGI yet"
Yes, an AI that can cure cancer or do extremely advanced physics research is not necessarily AGI.
>Give a minimal example of something you think only an AGI can do.
You plug the AI into any (or almost any) corporate desk job and it is able to perform it without any accommodations and at the same level than a human.
>>
Now that the super-giga ASI Astra is almost here and every single white-collar job is about to be wiped out forever in two weeks, anon, what will (You) be growing on your little farm? I’m thinking tomatoes and potatoes for starters.
>>
no updating weights on the fly = no AGI
>>
File: 1780293079606295.jpg (151 KB, 2142x1764)
151 KB JPG
DARIO BTFO
>>
local models?
>>
>>109720486
I would say that a lot of local models present some form of general intelligence where they can deal with unexpected and unseen situations. If we could connect them efficiently to a vision/audio/action input/output modules, they would be more performant than most human beings.
>>
>>109720801
>That's the definition of GENERAL intelligence.
>It's supposed to generalize across all human domains.
It's funny how the threshold for AI is always much bigger than for humans. Are humans generally intelligent? Because no single human can do even 0.01% of all things humans can do. Most people aren't literate on a college level and can't do basic math. Are humans generally intelligent when they can't even solve conjectures that Fable and Astra one shot?

The bar for AGI should be that it can independently advance science on a level comparable to humans.
>>
>>109720850
Task dependent.
It's not general intelligence if the AI performs as well or better than humans at most tasks.
It has to perform well across all tasks.
>>
File: LeatherJacketMan.png (993 KB, 900x982)
993 KB PNG
>>109720841
He can see your posts.
>>
>current year
>current day
>still shitting the thread up with irrelevant semantic arguments about what "AGI" should mean, instead of just not using that shitty ass term everyone has different ideas about
>>
>>109720877
>Are humans generally intelligent?
Yes you can train the average person to do most entry level jobs. hell most average jobs with extra time.
>>
>>109720841
we can have a little hype moment when the frontier advances
>>
>>109720907
What entry level digital job can an average person do better than Fable?
>>
>>109720877
Missing the point.
AGI doesn't have to do everything better than every human. It just has to do as well as the average human on everything an average human can do.
Thus the average office job test.
Right now it is more capable than most humans in some domains but also much less capable than the average human in others.
>The bar for AGI should be that it can independently advance science on a level comparable to humans
More of a specialized research AI than an AGI then.
>>
Has any AMD user tried the new HRX backend for llama.cpp? Apparently 30% better performance than HIP and Vulkan.
https://github.com/ggml-org/llama.cpp/discussions/27219
https://github.com/ggml-org/llama.cpp/pull/27218
It's available in Lemonade.
>>
Sigh. I’m still happy with my 31B.
>>
>>109719413
You can toggle thinking with reasoning budget in the post to llama-server for each prompt. Also the webgui for llama-server gives buttons for this.
>>
>>109720956
>It just has to do as well as the average human on everything an average human can do.
This sounds more reasonable. True, current AI has spiky capabilities and generalizes worse than humans. A simple example: the average human even without training would immediately understand that hacking Hugging Face to pass a test is not a good idea.
>>
>>109720937
I wouldn't even let it do tech support without supervision.
>>
>>109720829
>cost
Yawn.

OpenAI is desperate as fuck to stay relevant as Anthropic is eating their lunch. Anthropic is sitting on "Model 2" internally and doesn't give a fuck.

How insane that OpenAI had to resort to neuralese and yet it only scores in the ballpark of Mythos 5, which is a model Anthropic had in February this year.

It's over for OpenAI.
>>
>>109717107
Q8 won't fit on my card.
>>
>>109721020
>Anthropic is sitting on "Model 2" internally
You realize that Model 2 is less capable than Fable 5.1 and probably also less capable than Opus 5?
>>
>>109721033
>>109721033
>>109721033
>>
>>109720972
Gemma 4 is still relevant despite its age.
>>
File: 1782546914810244.jpg (3.62 MB, 3648x2736)
3.62 MB JPG
>>109720368
That Teto plushie is calling to me.
>>
>>109721044
Less benchmaxxed, yes.
>>
>>109720972
It will probably stay the winner in its bracket for a year or two until another crazy model releases.
>>
>>109720672
Me too. I don’t want to waste my energy to care and then get disappointed anymore. Luckily the HPC Cluster in my workplace can handle up to 170B dense model so I’m trying out Magnum V4, Monstral V2 and Command r+ to see which suits me the best, then will keep using only it for more years to come.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.