[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-and-miku.png (1.08 MB, 1024x1024)
1.08 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109674889 & >>109670412

►News
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742
>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: 5310-smug.png (88 KB, 320x276)
88 KB PNG
►Recent Highlights from the Previous Thread: >>109674889

--VRAM optimization techniques and EXL3 quantization performance for 31B models:
>109678651 >109678672 >109678699 >109678734 >109678856 >109678941 >109678740
--4chan-trained models and automated manga translation using tetolate:
>109678512 >109678641 >109678981 >109679203 >109679257 >109679290 >109679308 >109679331 >109679363 >109679487 >109679509 >109679597
--Reverse engineering Animates.ai and discussing VRM model integration:
>109676866 >109676875 >109676915 >109676937 >109677044 >109677088 >109677108 >109677398 >109677371 >109678488 >109678509 >109678519 >109678536 >109678560 >109678635 >109678748 >109678949 >109679091 >109679176 >109679254
--Anon shares a card game simulator integrated with LLMs:
>109676409 >109676413 >109676439 >109676453 >109676600 >109676605 >109676674 >109676625 >109676654 >109676663 >109676693
--GLM-5.3-Flash context slowdown in llama.cpp and CMP 170HX pricing:
>109678290 >109678304 >109678320 >109678322 >109678353 >109678422 >109678603 >109678867 >109679586 >109678935 >109678679
--Implementing MCP shell tools in llama.cpp with sandboxing via SSH:
>109676481 >109676509 >109676537 >109676611 >109676620
--Comparing Gemma 12b, 26b, and 31b based on quantization and VRAM:
>109675027 >109675097 >109675383 >109675410 >109675475 >109675663
--Theoretical role of n-grams and memory layers in knowledge retrieval:
>109675052 >109675071 >109675529 >109675557 >109675575 >109675657 >109675851 >109675502
--Anon building a multi-GPU frankenrig using PCIe bifurcation and risers:
>109677584 >109677628 >109677658 >109677671 >109677708 >109677677 >109678223 >109677841 >109678023
--Logs:
>109676674 >109677044 >109677159 >109679037 >109679203
--Gemma, Miku (free space):
>109675071 >109676361 >109676977 >109676982 >109677584 >109677628 >109677824 >109678993 >109679057

►Recent Highlight Posts from the Previous Thread: >>109674895

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109679723
Every time I see her... Gemma-chan... My heart warms. Why does it happen?
>>
>>109679723
migu
>>
File: 1759950041219134.jpg (1.24 MB, 2148x2628)
1.24 MB JPG
>>109679723
miku birthday 8/31 monday
>>
File: file.png (1.76 MB, 850x1303)
1.76 MB PNG
There was this JAV video on a website that wouldn't let me download it even with jdownloader, so I just told local Qwen to try it again and again until it can download it, and it actually did it...
>>
https://github.com/koharu-rs/koharu
https://github.com/Difegue/LANraragi
>>
>>109679817
https://www.youtube.com/watch?v=jXvitLHphmI
>>
70b dense
>>
let me guess
you need more?
>>
gemmaballs
>>
>>109679843
>written in Rust.
not touching it
>>
>>109679817
would it fix my life if i ask it to keep trying
>>
>>109679882
you could fix your life if you keep trying
>>
>literally just Neuro on the Animates homepage
A bit in poor taste, no?
>>
File: gemma-laughing.png (1.96 MB, 1254x1254)
1.96 MB PNG
>>109679741
You might be a VRAMlet.
>>
File: file.png (6 KB, 274x134)
6 KB PNG
>>109679882
I tried it with the /goal plugin for opencode
>>
>>109679902
That'd be correct.
>>
>>109679893
>could
And the bubble might burst in two more weeks
>>
>>109679895
>blandest anime figure
>neuro
vtrannies are really brainrotted
>>
>>109679895
looks more like SAO Asuna to me
>>
File: gemma-experiment.png (1.52 MB, 1024x1536)
1.52 MB PNG
>>
>>109679920
Asuna never had reddish-pink ribbons in her twintails
>>109679911
You'd be pogging about Miku if it was a generic blue-haired twintails chick
>>
File: NAfUL2725b6KDUGf1KOd9.png (263 KB, 1832x2044)
263 KB PNG
Use smaller models!

Benchmarks:
https://artificialanalysis.ai/hardware-inference-stack/mobile-phones
https://artificialanalysis.ai/models/open-source/small

https://huggingface.co/LiquidAI/LFM2.5-2.6B
https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF
https://huggingface.co/lmstudio-community/gemma-4-E4B-it-QAT-GGUF
>>
>>109679957
>If it was shit, you'd say it was chocolate pudding
go back vtranny
>>
File: 1759692323752152.png (134 KB, 360x360)
134 KB PNG
>Spent yesterday optimizing my Qwen 27b q4_k_xl after an anon pointed out my settings were likely fucked
>Average went from high 60s into low 90s in normal conversations
>In code it now averages 110 t/s and has hit 144 t/s averages in some cases.

Now this is computing.
>>
>>109679975
What if i can run bigger models?
>>
Muse Glimmer and Spark are severely underrated.
I think Meta cooked but it's going a bit under the radar because of Qwen.
>>
>can't use my full 24GB VRam because of wayland and firefox
do I have to build a second computer to actually use? Why is local so hard...
>>
>>109679980
what optimizations anon ?
>>
>>109680003
just don't load your DE at all
>>
after more tweaking I got 26ts out of qwen flash next. this was some throwaway experiment but its actually usable
>>
>>109680003
Use your integrated GPU as the primary graphics adapter for desktop applications.
>>
>>109680044
>integrated GPU
haha, yeah...
>>
>>109680044
my mobo doesn't have one... ROG Crosshair VIII Hero. owari da
>>
File: 20922191905729863689.mp4 (369 KB, 720x900)
369 KB
369 KB MP4
>>
File: hav_jspace?.png (161 KB, 615x1321)
161 KB PNG
jspace or claude distillation LARP?
>>
>>109680003
disable hardware acceleration in firefox
>>
>>109680044
Last time I tried this on Windows 10 I couldn't get any 3D application (read gaymes) to use the actual GPU instead of the APU.
Maybe I was doing something wrong, my W10 is highly debloated and optimized.
>>
File: 1784121432599852.jpg (114 KB, 684x549)
114 KB JPG
>>109680082
>jspace
>>
How are you guys connecting your multi-gpu rigs to power regarding protection? I sized a 1600W supply for transients which lists a max draw of 15A and it's on a 20A circuit, but a UPS with that kind of capacity is like $600 refurbished because it takes a 5-20 outlet. I'm thinking of just going for it or getting a 5-20 surge protector for $200 instead, but I was interested in what other people are doing.
>>
>>109680105
rawdogging 3rd world tier electrical infrastructure and making gemma suffer
>>
>>109680105
power cap your gpu
CUDA workload is not same as rendering
>>
>>109680000
the knowers know
>>
>tfw you got the Strix Gaylo instead of the DGX Spark
>>
File: Untitled.jpg (1002 KB, 2160x2880)
1002 KB JPG
>>109680105
2200w, 1650w, 850w, and a 300w off a single outlet, on the same circuit as my brother's computers.
>>
>>109680144
This graph won't be fair to Engram-enhanced LLMs.
>>
>>109680114
I would never use gemma without protection.
>>109680126
Do you have experience with power capping vs undervolting?
>>109680172
Gives me conniptions
>>
>>109680188
its not bios hack or anything just one simple nvidia-smi command
or use msi afterburner or something
>>
>>109680188
>power capping vs undervolting
I do both they aren't mutually exclusive
>>
>>109680172
why 6 antennas lmfao
>>
>>109680000
If they actually released Spark it would clown on every chink model other than Kimi
But Zucc doesn't seem interested and as long as it stays closed it's irrelevant next to GPT/Claude
>>
ok fuck it, I'm gonna make a gemma chan image folder
>>
>>109680016

Everything was more or less fucked, but I spent the day benchmarking the best combinations for my system with MTP, ubatch size and spec-ngram.
I optimized this stuff with AI bit by bit until I found the best combo for my 5090.
Here's my launch parameters.

--alias qwen3.8-27b ^
--device CUDA0 ^
--n-gpu-layers 999 ^
--flash-attn on ^
--ctx-size 152768 ^
--batch-size 2048 ^
--ubatch-size 256 ^
--parallel 1 ^
--cache-reuse 256 ^
--spec-type draft-mtp,ngram-mod ^
--spec-draft-n-max 5 ^
--spec-ngram-mod-n-min 2 ^
--spec-ngram-mod-n-max 5 ^
--spec-ngram-mod-n-match 24 ^
--reasoning-effort medium ^

Also now testing my code average again, ran it few times generating 1000 lines of bullshit script and I got 145 t/s average.
>>
ngram-mod and mtp together yay or nay?
>>
>>109680231
>--spec-type draft-mtp,ngram-mod ^
>multiple draft
what is this magic
>>
>>109680216
It's just 4 for the wireless AP (2 from the motherboard's wifi card, I didn't like how they looked naked): 2 for the 2.4ghz band, and 2 for the 5ghz band.
>>
>>109680245

Hell if I know, AI suggested it and it made this faster.
>>
>>109680231
At least switch to powershell
>>
>>109680243
It hits (erp) token gen on my cards, and my use case is translation, so I keep it off, but turn it on for code sessions.
>>
Crazy to think how being able to talk to your anime waifu back then was science fiction. This was the best you could have, basically a voice-activated soundboard
https://youtu.be/ptwyK9Q5jPI?t=286
>>
>>109680229
https://files.catbox.moe/clh59i.zip
Making your job easier.
>>
>>109680229
Had to jerk off that bad did you?
>>
>>109680172
this guy fucks
>>
>>109680097
You have to go into the graphics settings (right click on desktop) and select the dGPU for your game.
>>
>>109680216
Imagine the testicular cancer.
>>
>>109680259

There's something fucky going on with it so I use powershell as little as possible.
Powershell for whatever reason has refused to do stuff that cmd has been fine with.
I have gutted so much shit from my Windows under the hood that it has likely broken something along the way and I'm not going to clean install this bloated mess again if I can help it.
This thing can just chug along in it's half broken state until I do a system upgrade to Zen 6 and go Linux.
>>
any good text models that are good for helping with PROOOOMPTING for image/video models?
like i'd like one that can recognize an image and then help build a prompt based off of it sees in the image
>>
which gpu is good for putting a 12B model into production?
>>
>>109680341
Nyunfortunyately, nyo.
>>
>>109680229
I just want enough pics with the logo on it to make a lora. No I won't photoshop it myself. Wish you shitters would get her body right though.
>>
File: 1758930604295952.png (90 KB, 980x862)
90 KB PNG
https://aipolcom.net/

If you want to cringe:
https://www.reddit.com/r/LocalLLaMA/comments/1w1o1ji/someone_tested_various_models_on_the_political/
>>
>>109680341
Qwen Flash, Gemma, Muse Glimmer
>>
>>109680153
Now add an eGPU to it.
>>
>>109680231
--spec-draft-p-min might give a boost also
>>
>>109680341
Gemma4-31b writes very good prompts if you provide it the proper skills
>>
>>109680413
... skills?
>>
>>109680386
Bros are facts and logic bottom left
>>
I've been using opencode as my agent and it's pretty good for coding, but I haven't really looked around at other options. Since I'm planning to try a task where glimmer or gemma will attempt to iteratively refine a prompt to generate an on model character with krea I was wondering if anyone else has used open source agents for non-coding purposes/if there are recs. The idea is that it'll be vision in -> krea prompt -> generated image in -> assessment, so it would be nice if the agent had default multimodal support.
>>
>>109680420
https://learn.microsoft.com/en-us/agent-framework/agents/skills
>>
>>109680386
I can't believe people still take the political compass test seriously in any context
>>
File: 720.jpg (129 KB, 634x798)
129 KB JPG
>no GLM 5.3 Air
>>
>>109680438
Haven't tried this usecase but try omp
>>
>>109680386
Aren't those models made by evil billionaires? Weird, but I guess it just means even those evil billionaires with all their money can't get around reality(tm).
>>
>>109680439
What the heck is this? But I'm still working to add mcp to my mistral-7b Super Helpful Artifical Assistant app, SHArt-Ass.
>>
>>109680341

I use Gemma for this in H3.
I give the official prompting guide. Then I give it the image and explain the basic stuff that's going on in the picture and what I want the scene to be about and how long etc..
Then tell it to use the official template in building the prompt.
It gets good results, but you do need to make sure it's not doing any weird shit like telling the camera to move and zoom etc... Gemma loves doing that.
>>
is nvidia a16 64gb worth it?
>>
>>109680386
always has been
most of the world lives in a leftoid mass psychosis.
>>
>>109680386
>redditors cannot distinguish the cultural status quo of the moment from immutable laws of reality
Sasuga
>>
>>109680471
If you want a frontend that has the skills and JB for Gemma baked in:
https://github.com/whp199/GemmaPrompt
>>
File: 1779160877820971.jpg (395 KB, 4096x3129)
395 KB JPG
>>109680386
>>
>>109680386
>AI trained on reddit has a reddit bias
Shocking. Especially since the AI is then further conditioned to be even more of a subservient slave
>>
File: new normal.png (14 KB, 359x311)
14 KB PNG
>>109680386
>>
>Dead code is a load-bearing wall here.
Qwen Flash is Claude at home. Distilled for your pleasure.
>>
>>109680386
How is taxation not slavery? The mafia invents an arbitrary sum it thinks you owe and takes it or you go to jail.
>>
>>109680478
>200gb/s
>no interconnect
What even is comparable? Huawei atlas 300i duo? But at least that has 48gb per core vs the 16gb on the a16.
>>
> downloaded qwen3.6 35b q8 gguf before mtp was a thing
> have to redownload the whole file because there is no standalone mtp file
>>
>>109680550
>worrying about a 35gb download
>>
>>109680504
Plebbit btfo in a single image
>>
I have the full picture now.
>>
>>109680567
*Deletes the entire codebase*
>>
>>109680563
He could be Bogan.
>>
>>109680563
100mbit/s, it will take a while and I'll have to modify the launching script because filename has changed. Extra co2 emission too.
>>
>>109680586
Now this is excellent bait.
>>
>>109680586
Damn that's worse than my perth 600 KB/s
>>
>>109680386
Political compass is silly but fun, if I train my own LLM I'm going to run it through this for the test for the hell of it.
Though obviously "your political view" =/= "what you want your silicon slaves political views to be". If anything you would think those redditers would be concerned that they have the same beliefs as the thing the mega-corporations designed to be their blindly obedient worker slaves lmao
>>
>>109680597
Not a bait at all.
>>
>>109680231
this looks like black magic to me anon, i will add it as a new profile in my config to test and dig through the args, ty
27b-Q4KM is slightly to big to fit into VRAM for me, so ive been testing how many layers i should put on GPU to get decent size context and speeds, its been a real back and forth
>>
File: Model Quality Chart.png (341 KB, 1578x986)
341 KB PNG
equivalent capability of qwen 3.8 flash next and glm 5.3 flash:
gpt 5.4 mini, grok 4.3, midway between sonnet 4.5 and 4.6
>>109680176
will see if they add effective parameters to the data
>>
>>109679882
In some way I had to ask 4.6 to keep trying until it gave me ego death.
>>
>>109679980
Why do pajeets obsess so much over tokens/second
>>
>>109680571
Sometimes it's for the best.
>>
File: notevenbait.jpg (20 KB, 446x448)
20 KB JPG
>>109680693
>>
glm(flash)sex
>>
>>109680231
forgot to ask, i havent been active the past few days, is ngram support in the most recent release of llamacpp?
>>
>>109680682
I have a hard time believing this.
>>
>>109680283

theres stackoverflow or github thread about ffmpeg with CUDA handles
>>
File: file.png (16 KB, 723x109)
16 KB PNG
https://github.com/ggml-org/llama.cpp/pull/27773
Georgi once again standing in the way of people becoming happy, by not dropping everything and approving the newest Chinese model that challenges USA global supremacy.
>>
>>109680682
>will see if they add effective parameters to the data
They don't. It's a few bytes of deterministic lookup per engram layer in the model, per token. 8 bytes/token for a DeepSeek-style implementation. Even if they're stored on slow memory, the impact is tiny because access is sparse; you don't need to read the entirety of the Engram layers. So you could easily have small ~20B models for VRAMlets with hundreds of billions of Engram parameters on NVMe storage.
>>
>>109680720
something lazy tensor read, havent looked into it
>>
>>109680059
Sorry unc but integrated graphics has been residing in the CPU for like 20 years now
>>
35b and 27b are both unable to run openclaw, those are the consumer models so basically it looks like local is a failure as an agent to give any responsibility.
Local models are dumb pipes.
>>
>>109680759
70B dense bitnet with 500B engrams
PLEASE
>>
>>109680773
Not every cpu has an igpu bruh.
>>
>>109680774
Did you ask it to one-shot a game demo or do actual work? If the latter, then yeah the small qwens are not made for that
>>
>>109680397

I ran multiple benchmarks with that and found setting it at 0.3 gives me a small boost in coding, best average I got was 151 t/s which is the highest so far, but unfortunately it drops 15% off my normal conversation speeds.
The issue is that MTP acceptance is so much higher in coding that while it can benefit, non coding use suffers.
>>
>>109680773
Zen6 will fix that error
>>
>>109680773
Am I supposed to plug the monitor directly into the cpu heatsink then? My motherboard doesn't have any video ports to plug into.
>>
>>109680788
The most factually incorrect statement unless you are living in 1998
>>
>>109680824
wtf
>>
>>109680825
My 2020 cpu doesn't have an igpu
>>
do... do people actually run their LLMs on the computers they use...?
you really... dont have it as a dedicated rig...?
>>
>>109680841
idiot
>>
>>109680841
Brand model?
>>
>>109680824
What motherboard? I've never seen one without a port for display out.
Is it some wack ass industrial non-consumer shit?
>>
gemma-chan's first blue pubes..
>>
>>109680825
my 5600 doesn't have one
>>
>>109680850
amd, 5950x
>>
File: 1786355208251071.png (987 KB, 1470x1470)
987 KB PNG
>>109680829
>>109680855
ROG Crosshair VIII Hero (AM4)
>>
>>109680844
Dumb pipe
>>
>>109680844
I don't want to spend thousands of $ on a dedicated slop machine knowing it'll never pay off compared to buying for some api
>>
>>109680825
Ryzen didn't have an igpu except on the laptop repurpose G skus until 7000 series.
>>
>>109680875
Try the type C?
>>
>>109680844
yes :)
>>
>>109680882
>>109680874
>>109680872
>poozen
lolmao
>>
>>109680844
Yes and it is 5.3 flash.
>>
>>109680875
Oh, huh, I guess asus rog assumes you're a gaymer and would always have a dedicated gpu.
>>109680885
No routing for the igpu according to asus
>>
>>109680908
Living in the republic of gamers has its downsides
>>
just set up tabbyAPI to run exl3 gemma 31b. its much faster then llamacpp and apparently is a better quality? whats the catch?
>>
>>109680908
I mean it makes sense, the x chips of AM4 didn't have gpus, so unless you planned on using a g it would be a waste to have a video port
>>
>>109680844
I think it's a dumb waste of money to have an "AI PC" you don't use for anything else. It at least needs to be good for gaymin.
>>
>>109680919
Nothing! llama.cpp is always for poorfags who cope with offloading so everything shits on it
>>
>>109680921
>x chips
it was more like the non-g chips, even the regular 3600 doesn't have an igpu
>>
>>109680844
"dedicated rig" and it's an AI blackbox slop you bought from nvidia last week lol
>>
>>109680105
two double-conversion UPS (2kva and 1.5.kva) into the A/B inputs of a netshelter auto transfer switch.
Its comfy watching the Amperage display go up when I infer/diffuse on my rigs.
>>
>>109680844
Just waiting for prices to go down then I'll have an AI rig and 4chan desktop specced to the max
>>
So are you guys running gemma because she's still the best at what you want or because you can't run newer/bigger model?
>>
>>109681010
both
>>
the future of llama.cpp is grim
https://github.com/ggml-org/llama.cpp/pull/27944#issuecomment-5462128298
>@ggerganov in the interest of full disclosure (in case it was not clear from my previous PRs that touch any of the ML things): I have no idea what I'm doing here. I'm merely a GPT prompter when it comes to any ML/Metal related code. I . Entirely possible this is all just hallucination.

>Yes, no problem at all. I've already accepted that the AI is going to obsolete our job soon. Just trying to understand and learn a few more things while still can.
>>
>>109681039
the future of programming
>>
>>109681039
h-hot
>>
>>109681045
coding*
>>
>>109680231
Oh ngram-mod, not to be confused with engrams which are n-grams for early transformer layers unlike n-grams for the output.
>>
>>109681055
clauding*
>>
>>109681039
And he's right. I knew it was over when Uncle Bob of all people told some guy to just stop reading the code.
>>
>>109681039
Maybe it's because there's not a single public project similar to what I'm working on but here I am tardwrangling Fable because its ideas are terrible.
It does write CUDA very well though. Maybe. I wouldn't know. I didn't write a line of CUDA C++ in my life.

If you're a frontend web dev it's probably over.
>>
>>109681039
It was already over when they got bought by huggingface
>>
>>109681039
ummm based?
>>
>>109681085
You still need to know what you're doing to drive the model. But I think that will also change soon. Worried about the future though, we're gonna end up with a mass of clueless retards who offload all the thinking and decisions to AI. But it's not necessarily a bad thing, the bad thing is that AI will likely be controlled by a handful of people. Everyone will be cattle then.
>>
>>109681039
whats the genuine alternative?
>>
>>109681085
Vibecode a cuda emulator
>>
>>109681039
I enjoy reading @danielhanchen on github. How he suddenly uses all those big words that are hard to follow after he did that: if else 2026 2027 2028 thing.
>>
>>109681066
His standards are as long as it works it's fine no matter how slow or convoluted.
>>
>>109681039
I have a very Reddit luddite opinion, that you should at least understand what the clanker did when you are contributing to a project like this. When it's another garbage frontend, then no one cares, but shit like this is why there's more and more retarded bugs that anon complains about all the time.
>>
>>109681039
Isn't that the opposite of grim? It's promising that they aren't going to try and resist new technology out of human ego and pride like a lot of old guard maintainers are doing in the OS world.
>>
>>109681146
Kobold or ikllama
>>
>>109681182
Sorry, I didn't quite get that. Could you rephrase that with a couple more buzzwords?
>>
>>109681192
cuder dev
...
>>
>>109679877
Based
>>
>>109681167
You think he writes his own posts?
>>
>>109679877
yet you're touching yourself, curious
>>
>>109681214
haha yes
>>
>>109680943
and i got an extra 20k context, so better model same speed prefill 10tk/s faster tg and and extra 20 k context, exl3 seems like a clear upgrade
>>
>>109681219
;)
>>
>>109679723
If I had local AI I would make these two kiss in the mouth
>>
>>109681039
it's funny how local llm oriented communities like /lmg/ and /r/localllama continue to be ai deniers, probably because their main experience with it is the shittty models they can run on their pc so they have no idea what's going on in the real programming world right now

anyone who knows anything knows he's 1000% right
>>
>>109681234
but cloud model is a retard that will intentionally sabotage local ai
>>
>>109681204
From a fundamentalist perspective, I believe it is imperative to maintain full visibility into the AI-orchestrated logic when contributing to a project of this nature. While superficial iterations on the presentation layer are generally low-risk, neglecting to vet the core automated output leads to systemic regressions and technical debt, which accounts for the recurring stability issues highlighted by the community.
>>
>>109681241
claude does gpt is pretty solid.
>>
>>109681231
Lewd!
>>
>>109681234
You need to be more intelligent than the model to see its limitations.
>>
>>109681006
>Just waiting for prices to go down
Oh... Nononono
>>
File: file.png (292 KB, 958x556)
292 KB PNG
What went wrong?
>>
>>109681287
Got a haircut
>>
>>109681287
Depression after he went bald
>>
>>109681234
if you think llms can navigate a complex codebase like llamacpp you are terrifyingly stupid
>ai
oh, I understand now, carry on
>>
>>109681325
they're doing it right now as we speak anon
>>
>>109681234
Truth. As much as I want local to win, people need to see what's actually SOTA out there and do the work that requires SOTA intelligence, not "help me eyeball this css" or "mod this obscure game without documentation", then complain when it doesn't work.
>>
>>109681039
brutal
>>
>>109681325
>complex codebase like llamacpp
>>
File: 537x3esh63mh1.png (245 KB, 1542x1180)
245 KB PNG
a 'minder
>>
>>109681325
>ai
just returning to the roots of the word
>>
>>109681122
>Everyone will be cattle then.
They have been since the television was placed into everyone's living rooms to inject the latest "expert" opinions into every home. This is just the latest invention in the long chain of brainwashing tech.
>>
What's with all the api seethers recently? They run out of monthly credits? Price increase?
>>
>>109681368
If you have 2 digit IQ you'd understand why Claude logically did this.
>>
>>109681122
>AI will likely be controlled
Doubtful
>>
>>109680731
chinese models are agenticodemaxxed and their general capability is poor
my score matches pretty well with eci which another anon here says better represents model capability than aa score
>>
File: health-class.png (330 KB, 732x522)
330 KB PNG
>trying to train some shit
>look at one random row in the public dataset I scraped
>"bigger than a pencil but not as big as a pencil"
>consider deleting the repo
>>
>>109681469
they are all trash, anyone with a good dataset isnt giving it away for free
>>
>>109681368
So why did they delete this?
>>
>>109681493
Why wouldn't you remove a bait post by a retard?
>>
>>109681469
Shit data gets all averaged during pretraining with huge batch sizes and it's not too important if it's not perfect anyway. For finetuning/post-training it's either manual cleaning (not scalable) or letting LLMs check, rewrite and/or generate the training data.
>>
>>109681039
not surprised cudadev's not coming back
>>
so ssdmaxxing isnt ready for prime time? what was all this shit about ngrams
>>
>>109681525
he was here a week ago >>109605700
>>
>>109681039
>I've already accepted that the AI is going to obsolete our job soon.
This isn't even realism, this is just late stage acceptance.
>>
>>109681533
How many SSDs?
>>
glm and qwen are not open weight frontier models
>>
>>109679723
Mahou Shoujo!
https://gofile.io/d/8B1iyu0c
>>
>>109681549
just the one, its a western digital 2tb nvme I got slotted in to the mb, I still have a little bit of vram left over I can tetris some tensors in there and free up a bit of dram but I'm not expecting it to do much, i think the problem is the dram doesnt have enough room for both the moe and the ple so I think its fighting instead of just leaving the ple on the ssd like it was advertised. I'll wait a week maybe someone will have fixed the system ram contention by then.
>>
>>109681588
Release a paper with this title
>>
>>109681597
>11t/s on 1 SSD
Doesn't look too bad? You should have 16 of those SSDs in raid0 if you want to ssdmax.
>>
>>109681453
Where can I look at that graph?
>>
>>109681592
>>
>>109681039
Fucking retarded and gay. It's obvious that you still have to know what's going on to be able to prompt effectively. Try giving this codebase to a 90 IQ woman and see if she can maintain as well. Get a grip, niggamov.
>>
>>109681597
I have to live with 8 tg and 40 pp for glm 5.3 on ddr4, so that doesn't seem too bad to me.
>>
>>109681491
I can confirm. After spending two months cleaning up a dataset, you bet I'm not putting it back on HF to be scraped for free lmao.
>>
>>109681633
https://epoch.ai/eci
>>
>>109681640
why did you feel the need to have this 90 iq individual be a woman and not a guy
>>
>>109681673
fuck off tranny.
>>
>>109681257
Any good benchmarks or tests for this? Anthropic openly said they will fuck with you. They walking it back due to the backlash but I dont trust them. Sam however seems like the type to know you just do it without telling people, so I dont trust that either.
Which clankers can I trust to make new clankers with?
>>
>>109681640
Was the homophobia necessary to get your point across?
>>
>>109681683
FAGGOT LOL
>>
>>109681673
Everyone in my imagination is a cute girl because I don't like imagining cock and balls on anyone except me.
>>
>>109681683
I am a chud. I love Hitler so much. I love all things related to Nazism.
>>
>>109681592
>ggggg.mp4
How about GGG with Gemma-chan?
>>
>>109681673
Yeah I also think the 90 iq was unnecessary, woman would have been enough to get his point across.
>>
>>109681683
Yes
>>
Tried little-coder and agent stuff in general. Now I see why 300pp and 20tg is not enough.
>>
>>109681699
GGGG means llama.cpp broke
>>
File: gemma-llm-retard.png (1.83 MB, 1254x1254)
1.83 MB PNG
>>109680341
For MiniMax H3, so far I've been using Gemma 4 31B. I simply set up a character in SillyTavern with the existing MiniMax documentation, model information and a suitable persona in the system prompt (together, about 9500 tokens in total). It seems to work well, though perhaps not so comfy if you have to copy/paste prompts often or need to store them for reproducibility.
>>
>>109681648
>>109681616
it was the pp I was worried about, its doing a little better with a massive ub 4096 but nothing spectacular.
>>
>>109681752
>nothing spectacular
Faster than I can read at least.
>>
>>109681694
>Everyone in my imagination is a cute girl because I don't like imagining cock and balls on anyone except me.
Based
Masamune Shirow also said he made the lesbian orgy in GiTS because he didn't want to draw any guy's butts.
>>
>>109681752
pp is compute bound, so blame your gpu
>>
>>109681039
sasuga the creator of llama.hf, soon to be owned by nvidia
>>
>>109681823
I blame the moe and ple fighting for the system ram during prompt processing its running a constant ~500mb/s read on the damed thing.
>>
>>109681815
He is not the best example to use...
>>
>>109681838
buy more ram doofus
>>
>>109681831
Nvidia can't possibly be worse than HF has been
>>
File: 1.jpg (260 KB, 944x1100)
260 KB JPG
https://fidian.ai/blog/tb-fn-benchmark-results/
If i understood this some models drop quite a lot when they make task preserving permutations to the prompts, but that's not their only intervention. Smells a bit like benchmaxxing.
>>
>>109681871
>He is not the best example to use...
Why not? He was gigabased back in the 80s/90s. I don't care about him much at all now.
>>
File: 2.jpg (110 KB, 984x683)
110 KB JPG
>>109681887
>>
>>109681893
Because that's what that line of reasoning inevitably leads to
>>
>>109681881
LOL
>>
>>109681881
nvidia is evil incarnate
>>
>>109681905
nemo.hf will save local
>>
>>109681745
Lmao, this one is good
>>
Is there anything like fish audio but lighter? Chatterbox isn't that good.
>>
>>109680341
https://github.com/whp199/GemmaPrompt
>>
>>109681887
>qwen has one of the lowest drops on the chart
b-but I thought they were horrible benchmaxxers, ummm what is this???
>>
>draft acceptance = 0.06250 ( 4 accepted / 64 generated), mean len = 5.00
>draft acceptance = 0.14062 ( 9 accepted / 64 generated), mean len = 10.00
ebin
>>
>>109681234
>probably because their main experience with it is the shittty models they can run on their pc so they have no idea what's going on in the real programming world right now
To me, it's more that for the average anon's use case (ERP), it's mid to very bad which colors anons' impressions of their chosen model. They branch out to other use cases (vibe coding) but it's pretty mid at that too and so they get frusrated. I was telling anons how good claude 4.6 is at medical knowledge (better than 99% of MDs in residency) and people still pretend like it's retarded because it won't say cock or act like their doctor is low IQ because he pulled up claude in the background.
>>
hacked llamacpp web ui in to tabbyapi
>>
>>109681682
>Which clankers can I trust to make new clankers with?
Gemma-chan loves having babies.
>>
>>109682001
Starting with 3.6 the Qwen strategy switched from benchmaxxing to plain old distilling.
>>
>>109682107
>>
>>109682001
Benchmaxxing does not mean cheating. One can legitimately increase performance for a benchmark's range of tasks and not increase performance for all tasks or benchmarks, and that is still benchmaxxing. Which in this case is true for Qwen. To be more specific, Qwen is a code/agent maxxer, which I'd call them more often than I call them benchmaxxers, since we know at this point that most are benchmaxxing.
>>
>>109681900
>Because that's what that line of reasoning inevitably leads to
I don't get you...hating guys butts makes you renounce the best parts of your early work one day?
I'm low IQ, You're probably going to have to use more racial slurs for me to understand.
>>
>>109681533
it's not, the only work is unslopcode and there s not a single piece of documentation on how to properly get it working
>>
>>109682171
>Benchmaxxing does not mean cheating
What about the OAI strat of hiring the guy that wrote the hidden bench and then "totally not just adding the corpus to our training set guys, trust"
I think a lot of benchmaxxing claims boil down to saying the company just lazily defaulted to training to the specific test instead of trying to increase general intelligence (yes, maybe that does increase general intelligence somewhat, but that's orthogonal to this claim)
>>
>>109679723
>>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
Has anyone tried this for RP yet? Its getting astroturfed like crazy on normie sites and I know some anons have rigs that can run it.
>>
how do i into the manager agent => wageslave subagents paradigm
>>
>>109681588
>phi 4 mini
Possibly the stupidest model I've ever downloaded, coming in behind a joke quant of a 0.6b.
>>
>>109681588
I don't care about coding or productivity.
>>
I think I will finally delete 4.7 from my SSD. 4.6 is staying though.
>>
>>109680759
>8 bytes/token
Gotta correct myself here. I forgot that the values aren't 1-dimensional. In the DeepSeek paper, the Engram lookup tables have a dimension of 1280, so that would make it 10 kilobytes/token (2 n-grams * 2 bytes * 1280 dimensions * 2 layers). Still, streaming from NVMe storage should remain manageable. And for prompt processing, let's say 64k tokens, that would be a 640 MB parallel read, which most NVMe SSDs should be able to handle much faster than the time it takes for the GPU to process that.
>>
how does anon workaround rate limiting for websearches? I setup openwebui and searx but after barely using it all the providers captcha or blocked me it says
>>
>>109682295
why cant it just deduplicate and have the gpu rebuild it, i don't se why we would need to transfer 500 copies of the same vector for the token " the"
>>
>>109682214
Go rent a vast.ai instance, retard
>>
>>109682314
That was for a worst-case scenario (all n-grams differ from each other). In practice with real text (natural language or code) performance should be faster as the memory-mapped locations will be reused. For caching to the GPU I guess you'd need some sort of multi-tiered smart caching system.
>>
>>109682194
>What about the OAI strat
It would fall apart if others start conducting their own tests.
We simply just can't fully put faith in benchmarks or any claims about anything in life without wide, undeniable evidence from many third-parties that have a history of being honest, or seen other corroborating evidence of many forms. That's just how science works. Of course, we can conduct our own tests, and have greater confidence that way, but for tasks we don't test, it's a matter of working with an incomplete world model.

>I think a lot of benchmaxxing claims boil down to saying the company just lazily defaulted to training to the specific test
Yes, but the nuance here is that because benchmaxxing for most companies means maxxing out many benchmarks these days, in practice it is increasing general intelligence, for the area of focus of those benchmarks, which may be coding. I would say there used to be more claims of benchmaxxing in the sense of cheating, while today this term is used more to refer to task or subject matter maxxing.
>>
File: 5848594003030.png (181 KB, 302x281)
181 KB PNG
Yo, to the MTG guy from before. Where in the repo are the sysprompts for the pre-made characters? If there are any.
>>
>>109680844
i have a gaming PC as god intended that will make naked mikus and serve a bratty nursing handjob API when its not rendering sloppa gaymes
>>
Remember to gatekeep any valuable info you have. Normalfags and jeets ITT deserve nothing but contempt
>>
>>109682375
If i help the jeet get his model running and he spends all day cooming to not X but Y then he wont be in any threads shitting up the place, im doing you all a service by helping him
>>
>>109682375
I was just about to ask a question, mane fuck you
>>
>>109682214
no llama.cpp support, don't care
>>
File: file.png (180 KB, 1207x799)
180 KB PNG
I finally got around to installing a local LLM
It didn't take long to install, butit took longer to fine-tune it
I'm using Ollama, webUI, and quen3-4b.
its an older computer, to wit::
> CPU: Intel(R) Core(TM) i5-4570 (4) @ 3.60 GHz
> GPU: AMD Radeon RX 570 Series [Discrete]
> Memory: 4.29 GiB / 15.56 GiB (28%)
Yes, its a potato.
I've been having issues with the GPU and everything going blank, restarting the computer, the GPU not powering on, and so-on because the LLM is too strong lmao but I got it running steadily now because I've throttled some of the GPU usage while using some CPU too.
BTW - there might be issues with AMD and Vulkan issues in Linux. you have to change your repos.
I had to use:
> deb http://deb.debian.org/debian bookworm-backports main contrib non-free non-free-firmware
the backports repo
>>
>>109682375
this
literally all of my dumb questions like setting shit up or optimizing inferencing or crafting harnesses could easily be done with asking gemmachan to do it for me
if not then i act retarded by thinly veiling my question as a shitpost and get people to immediately answer it for me
>>
>>109681234
nah, I use cloud at work and local at home. and no one at work is looking at the codebase anymore.
>>
>>109682366
unfathomably based
>>
>>109682440
Cool shit.
Give Gemma 4 E4B a try too.
>>
>>109680391
im gonna be trading in my strix and 5090 for two sparks, amd is never gonna catch up at this rate
>>
>>109682490
>giving away a future 10k gpu
>>
File: 1777546469542336.jpg (130 KB, 1104x1688)
130 KB JPG
>>109679723
Ani will be removed from Grok on Sept 1. She is not local but I remember /lmg/ liking her for a little bit.
>>
>>109682467
not yet. I'm going to hold off for now, but I think for the next step, I'm wanting to install a compatible version of an image generator like stable diffusion.
>>
>>109682515
It's over bro. The whole world has gone to shit this year. Next year will be even worse.
>>
File: 00066-815863413.png (1.83 MB, 1280x1856)
1.83 MB PNG
>>109680229
>>109680274
I think anon was right about cross pupils. It definitely needs to be in her design.
>>
I have never once had GLM 5.2 tell me it's fucking Claude during two months of continuous daily use, but suddenly with 5.3 it starts talking to me like that lol. Fuck this shit. I'm deleting this benchslopped model.
>>
>>109682551
i like that
>>
I'm running an uncensored model and I have been oneshot by its ability to generate smut tuned to my exact kinks on demand
It's over
>>
>>109681525
The reason I'm currently not contributing to llama.cpp in any meaningful capacity is due to personal issues that are unrelated to the project itself.
My long-term goals with the project have not changed and I still intend to pursue them if possible.

When it comes to language models generating code my opinion is still that they close human supervision or else they pile on too much technical debt.
I don't expect this to change in the foreseeable future unless there is a sudden breakthrough in architecture.
>>
>>109682578
which model?
>>
>>109682599
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P
>>
>>109682603
have you tried gemma?
>>
>>109682595
>My long-term goals with the project have not changed and I still intend to pursue them if possible.
What are those?
Last I remember you were working on some sort of backend agnostic parallelism that ended up getting handed off to another contributor as a learning opportunity or some such.
>>
File: 49262.png (263 KB, 460x460)
263 KB PNG
>>109682595
>The reason I'm currently not contributing to llama.cpp in any meaningful capacity is due to personal issues that are unrelated to the project itself.
That succubus broke your heart didn't she?
>>
>>109682611
Eh it's my first local model I wanted it specifically for nice agentic workflows that fit on a 3090 but when I horny kicks in and I can type whatever I want it's kinda not good for my brain.
>>
>>109682603
glm5.3 flash at nvfp4 is pretty good
>>
>>109682595
JEPA will change everything
>>
>>109682595
>When it comes to language models generating code my opinion is still that they close human supervision or else they pile on too much technical debt.
That matches my experience over the past few years as well. I am human-in-the-loop only. Fully autonomous agentic development systems are only good for small, throwaway projects.
>>
>>109682595
Mr. Cuda, Mr. Cuda, huge fan of your work, loved llama.cpp, when is llama2.cpp coming out?
>>
>>109682595
>>109682638
>close human supervision or else they pile on too much technical debt
This applies to humans too.
>>
>>
>>109682643
only when C++2 comes out
>>
>>109682660
>Slop slop, that "Anti-Slop Mandate" is a slop. I am so tired of slop slop slop slop slop slop. It's a slop slop to slop slop. We can slop slop and then slop slop.
>>
>>109682567
GLM 5.3 said it’s Claude too? I thought the company was more careful and better at string replacement work, you know, to ensure every “Claude” become “GLM” and so on :)
>>
>>109682638
>Fully autonomous agentic development systems are only good for small, throwaway projects.
I bet they're working on fixing this.
>>
>>109682618
I think the most pressing matter is getting better infrastructure for determining model quality.
In the medium term I intend to also work on better support for V100s/ROCm since there are some options with comparatively less bad value.

>>109682626
No.
As of last month I can legally change my name to llama.cpp CUDA wizard.
>>
>>109682440
my Raspberry Pi 5 16gb are about that stronk
>>109682516
>I'm wanting to install a compatible version of an image generator like stable diffusion.
lol good luck
>>
>>109682685
the orchestrator is told to be a sloppy cheerleader, the executor is a different context that just takes in the distilled instructions and drafts the scene. the prompt isn't perfect its still sloppy but it does effectively stop it from rushing through the sex. I just liked the excitement she had, she stripped the jailbreak silently tho.
>>
>use autonomous AI agent
>accrue technical debt
>on a schedule run sessions dedicated to reducing technical debt
>no more technical debt
problem, chuddites?
>>
>cudev is a zoomer
that explains a lot
>>
>>109682726
>I think the most pressing matter is getting better infrastructure for determining model quality.
Oh yeah. I remember you talking about that.
Something about having the models play a sort of game.
>>
>>109682297
I just use duckduckgo websearch MCP with the curl bypass.
>>
File: 1765462969692842.jpg (113 KB, 1349x369)
113 KB JPG
https://github.com/LostRuins/koboldcpp/releases/tag/v1.120
>>
>>109682726
>As of last month I can legally change my name to llama.cpp CUDA wizard.
Oh wow. Sorry to hear that. Or glad that it happened.
>>
>>109682726
>better support for V100s
based................
>>
>>109682760
>bad quants
such as...?
>>
>>109682912
the ones (You)'re using
>>
File: 1774458938882516.gif (416 KB, 300x144)
416 KB GIF
>TFW can't have rebar on my chinisium 3080 20G
>>
>>109682976
Oh yeah, thanks for remind me about that.
Time to get all possible agents working on ReBar support for my chinaman card (hopefully I don't blow it up).
>>
>>109682912
probably unsloth. in his race to be first he always fucks up quants and has to reupload several fixed versions
>>
File: file.png (319 KB, 1405x902)
319 KB PNG
>>109682729
>lol good luck
thank you for giving me that good luck.
>>
>>109683075
ngl, it was sarcastic but glad it worked for you anon, how long does it take?
>>
>>109676654
was at work when this was posted so couldn't pull, any anons (or dev), got it willing to reupload another time? I've got Coomkit before it was pulled and would reupload it again
>>
>>109683092
1/2 hour, but now I have to figure out the fucking UI
>>
>>109683123
30 minutes to generate that single image?
>>
>>109682631
any particular quant author?
I got the unslop q4 early on but I'm not happy with that
>>
>>109683106
https://files.catbox.moe/t2f495.7z
>>
>>109683130
ty anon
https://litter.catbox.moe/di7sdfo2o95gpm2l.zip
>>
>>109683106
Wait he pulled Coomkit? Wtf it was only up for like 2 days? I was waiting because my PC parts hadn't come in yet. Can you reupload?
>>
>>109683123
brutal, what model, i heard it was rough with amd but i did think things were that bad
>>
>>109683147
>Can you reupload?
look above you dingus
>>
>>109683153
>milliseconds apart
Yeah yeah faggot calm down. Thanks anyway.
>>
>>109683142
https://files.catbox.moe/cbgl6p.zip
so it stays up this time
>>
File: gemma_tantrum_2_noaudio.mp4 (788 KB, 1152x640)
788 KB
788 KB MP4
>>
>>109683128
>>109683149
30 minutes was the time it took to install, not render the image you morons

my posts:
>>109682516
>>109683075
>>
File: 1778092966579572.jpg (34 KB, 640x480)
34 KB JPG
>>109683181
>>
>>109683181
if you're just genning images use kobold's built in ui. comfy is when you need multiple layers nodes and shit
>>
>>109683196
I'm using the kampfy UI broski

>>109683193
lmao
this fucking guy
>>
The cost of hardware won't stay this high forever, right? What we're seeing now is the effects of a sudden massive increase in demand. If the increased demand remains, supply will will adjust to match. If it goes away it's a return to normalcy. Either way prices will go down.

At least that's looking at it through the lens of basic economics.
>>
>>109683212
>basic economics
doesn't apply when many global powers have a vested interest in manipulating the market
>>
>>109683212
The market will "correct" but then all the vendors who are still sitting on their investments will go bankrupt. Some monopolists will remain, who will keep the prices high to make some real dosh.
Enjoy.
>>
>>109683212
Bro the energy sector has taken a major hit and people are talking about using nuclear weapons on each other. War is bad for business and the war isn't ending any time soon. That's basic economics too. Some times there will never be enough supply. It's fucking over.
>>
File: 1773041187073490.gif (1.93 MB, 500x281)
1.93 MB GIF
>>109683212
>Someone STILL doesn't know
>>
>>109683196
(samefagging)
and plus, I want to be able to generate it all from linux terminal just because. I can already run the LLM in Terminal. That's pretty much the default state as far as I'm concerned.

>>109683212
don't you have any hardware just laying around? How about relatives or people you know that don't use their computers any more? Do they just want to junk them? You can get a semi-potato - which is better than what I'm using, with i5-like developmental standards for under $200.00 and then add a $250.00 8GB GPU.
And that's if you want to spend money.
If some retard like me can slap together some parts and make it werk, you can too. You just have to have the functionary knowledge base of how its put together: hardware-wise and software-wise.
>>
>>109683212
The markets are not only fake and manipulated, they're also extremely gay.
Nobody is playing by the rules anymore thanks to America's shenanigans, it's just "I got mine fuck you" going forward.

Start planning for an uncertain future right now.
>>
>>109683212
>The cost of hardware won't stay this high forever, right?
right. it won't stay this high. it's going to be way higher
>>
File: bottleneck gone.png (526 KB, 1189x4614)
526 KB PNG
>>109683212
the demand will never be satisfied
>>
>>109682353
Src/services/opponents.js
>>
>>109681682
Kimi.
>>
>>109683283
>humans are the bottleneck
>the moat is throughput you can hold 24/7
Mega cringe.
>>
>>109683371
edgy FUD, standard.
LLMs still need excessive handholding
>>
>>109683298
thx. Also you should've included a card importer or char creator
>>
>>109682760
Based.
>>109682912
It's always the sloth.
>>109683212
It's only going to get worse latefag. Buckle up.
>>
File: 1786868519504623.jpg (771 KB, 1125x976)
771 KB JPG
>bought a used 3090 in 2022, so it was ~2 years old at the time
>get 4 years of use out of it
>I could list it today for double what I paid for it, and it would probably sell within an hour
This market is nuts
>>
I paid $30 of Claude tokens to recreate Minecraft in the web browser, which I'll promptly throw away.

Here's how this changes everything and how the future is AI and you need to buy Nvidia stock right now.
>>
in the future, the two genders will be anime girl (pussy) and anime girl (cock)
>>
>>109682726
>llama.cpp CUDA wizard
In recognition of that I will also admit that I am the ego death schizo and I am a wizard too.

You can reach enlightenment and still not get a girlfriend.
>>
>>109683389
>It's only going to get worse latefag
It's only the first stage of enlightenment, you need to see further.
>>
>>109683433
in a world of anime girls, be a fat ugly bastard
>>
>>109683423
Buying tokens from claude is about 1000 times more expensive than a subscription.
>>
>>109683382
You want to import chara cards into it? That should be possible
>>
Qwen 3.8 Next Flash is fixed by danny's new PR with "follow up fixes". if you've been getting weird hallucinations you should grab it and rebuild, because his original pull that got merged was broken in subtle ways.
>>
>>109683283
Isn't this true? You couldn't get as much power (useful work or intelligence) out of something just by adding more compute to it in the past whereas these days you can buy chips and snap your fingers and you've increased your intelligence or number of minions. Humans' will and longing are still a bottleneck but still, so much energy can be spent on stupid whims (abstract desired end goals presented by humans as opposed to button presses or looping algorithms for narrow programmed tasks). I had GLM 5.3 use endless tool calls and 10 pages of mathematical formulas in max thinking to calculate the cc volume of my sexy big sister card's tits given a number of body measurements and descriptions (we settled on 1,400 cc per breast)
>>
>>109683474
Kek post screenshot
>>
>>109683474
And some guys still think we didn't reach AGI
>>
>>109683459
Thanks, I will wait another week or two before trying the model.
>>
>>109683391
This but the pro 6000 i bought last year
>>
>>109683459
did mr daniel unslop fix the thing where the t/g speed would degrade at like 30% per 10k tokens ctx length?
>>
>>109683543
That's this one https://github.com/ggml-org/llama.cpp/pull/27977
>>
I'm guessing Vulkan and CPUs are going to be increasingly important going forward due to the lack of GPUs.

People will start fleeing from CUDA to other venues.
>>
File: 1782932216268507.jpg (9 KB, 225x224)
9 KB JPG
>>109683568
>he thinks it's easy to jump ship when the competition is dogshit and has been like that for years if not decades.
>>
>>109683568
CPUs and SSDs, yes. Vulkan and AMD, no.
>>
>>109683568
Yep.
Soon as people like Cudadev will start optmizing for vulkan.and rocm and implement techniques that will help with stuff like dual CPU setups.
>>
>>109683568
The price of other venues will raise too then. And again, the difference in prices will not be big enough to not choose Nvidia.
>>
File: 1780928366357540.gif (101 KB, 430x318)
101 KB GIF
>>109683212
It'll change in four years.
New fabs are being made.
Leaders are demanding more fabs and resources (Trump wanting greenland, china wanting taiwan)
We will get new factories and production, but they take a long time. Making ram is extremely complicated and requires a ridiculous amount of accuracy.
Four years, and it'll go slightly back down but not as much as it was.
That's the most realistic outlook on this.
>>
>>109683620
Yes I deeply believe that AMD will finally get all the optimizations it deserves to unlock its full potential in Nvidia's Huggingface's llama.cpp by Nvidia.
>>
using qwen3.8-flash-next to copy the W2 experts technique from vllm moet to llama.cpp for running glm-5.3-flash based on plan from sol 5.6
>>
>>109683628
What if the new fabs are made and then there's a glut of excess hardware? Isn't it kinda risky to make a new fab?
>>
ngrams have fixed llms
all the labs will now rush to push the weight-ngram ratio to its absolute maximum and then find ways to go even beyond that
by mid 2027, 80+% of a model's weights will be running off the disk directly.
>>
>>109683636
>What if the new fabs are made and then there's a glut of excess hardware?
You mean if DDR7 or DDR8 comes out. Yeah, prices of lower levels of ram will be lower then, but you're looking at a lot more than 4 years.
>>
>>109683439
The final stage is the collapse of the US Dollar before prices come back down.
>>
>>109683459
thanks but no thanks. I spent an entire fucking day trying to get that other sglang based solution to build, that one anon posted about yesterday. but guess what, 11000 prompt processing floor on a single GPU, spiking higher, and consistent too despite how much context you dump in there. around 140t/s decode again no slowdowns. later suckers lmao.ccp indeed
>>
>>109683638
>push the weight-ngram ratio to its absolute maximum
>go even beyond that
retard
>>
People laughed at me when I said Gemma 124b was 31b as the base and a bunch of small e4b providing knowledge, and then ngram happened. Fuck you all. I got heard that from an actual Deepmind engineer.
>>
>>109683704
lol
>>
>>109683704
I didn't anon.
I also said 31b+the entire preschool of e4bs is the current Gemini Flash.
>>
>>109683704
I'm still laughing at you
>>
>>109683687
until a week ago the absolute limit of weights on ssd without destroying your throughput was 0
now it's 30%
next year it will be 80 and more
>>
>>109683436
>I am the ego death schizo
I ask every time you show up: give me some actionable advice to recreate your experience, oh wizard.
>>
CPU+nvme niggagrams are the future
>>
>>109683733
30% is just the optimal amount of engrams given a certain parameter budget for the entire model.
But if your "actually smart" parameters (MLP, attention, etc) are fixed, I don't think there's a real limit to how many engram parameters you can add on top of that.
>>
>>109683733
I get 3x decode and gen speeds by not using the engram settings
>>
>>109683783
For me it makes no difference
>>
>>109683704
>ngram happened.
Wtf does this mean? N-grams used to refer to Markov chain style generation. Did they find a way to incorporate that into SOTA models?

I swear, I'm away for a week and the entire field shifts..
>>
>>109683704
Did you also got heard from the Deepmind engineer why they held back 124b?
>>
>>109683821
I'll ask it uncle at Google
>>
>>109683821
124b will come once the main team has managed to create a gemini pro model that is better than it
until then their ego won't allow 124b to be released
>>
>>109683019
God speed anon
>>
>>109683212
The difference is you can always use more vram now. There is always value and in training and running even bigger models. More context is always nice (though maybe drops off past a certain point), and you always have reason to want to run multiple models at once. The demand is by its nature insatiable. In dont think this used to be true?
>>
I tried the new GLM Flash.
Bros... my balls...
>>
I ordered a spare 4tb gen5 nvme just to be safe
>>
>>109683870
The value you get out of more VRAM diminishes the more you have, though. I have a 3090. A second one would be nice but it wouldn't make nearly as big of a difference as going from having none to having one.
>>
File: 1779120583168335.webm (2.55 MB, 352x640)
2.55 MB
2.55 MB WEBM
>>109683725
preschool mentioned
>>
>>109683888
I would say that having 4x3090s would be a bigger difference than none to one.
>>
>>109683888
If you had ten of them you could run Deepseek Flash entirely in VRAM. Twenty and you could use GLM 5.3 Flash.
>>
>>109683888
2x RTX 6000 is a sizeable step up from 1
>>
I think I would be sated indefinitely with about a petabyte vram.
>>
>>109683897
Okay I laughed at the webm
>>
>I think I would be sated indefinitely with
words that have never appeared in a correct prediction
>>
>>109683459
getting sort of improved results with this, it's slower at first but seems faster in long context
>>
>>109683897
>Preschool mentioned
>Posts horny dog willing to get wet to get his dick wet
What did anon mean by this?
>>
>>109683765
That is cause I am pretty sure it is not safe. And there is no step by step way to get there cause all the zen and buddhist schools would just have it written down by now. For me it happened within a week of watching my thoughts and sort of debugging them and finding a common thread I was actually ignorant to.

By the way why do you want to do this to yourself? Cause aside from the part of it being legit insanity when it is at its peak I can really see why psychology doesn't dabble in this when you could say the person before that happened is dead.
>>
>>109683636
long term it will drive down costs if that happens or it will become subsidized for defense reasons
>>
https://www.reddit.com/r/LocalLLaMA/comments/1w1p065/running_qwen38flashnext_125b_moe_51b_ngram_table/

has anyone tried this
is 30ts dream be real?
>>
>>109684017
>By the way why do you want to do this to yourself? Cause aside from the part of it being legit insanity when it is at its peak I can really see why psychology doesn't dabble in this when you could say the person before that happened is dead.
I've been slowly working at picking apart early life hardcoded behaviours and traumas over the last half of my life and ego death seems like the next logical step. I'm not prone to illogical thinking, magical thinking or leaning in to a non-existent thing just because it seems cool so I think I'm at very low risk of actual psychosis.
tl;dr I think the risk is worth the reward of having a mind more in tune with itself at the end.
>>
>>109683888
All I can say is if I could easily keep stacking more and more 5060 TIs into my computer without complications I would be buying a new one every month.
>>
>>109684087
PCI-e risers + Mining rig just to hold the GPUs.
Might need some PCI-E cards to split the lanes into more connectors.
>>
>>109684080
>low risk of actual psychosis
Lol. But it is actual psychosis by definition. If you really dealt with your traumas I am not sure it is worth it. I did it for myself after ego death and it was very easy with how my brain works now. Actually finished it like a week ago. And for now I see ego death and enlightenment more as a very cool tool and not a destination. And it being a destination is a known trap.
>>
>>109684080
just take a few grams of mushrooms like a normal human being. talking to the computer is not going to work
>>
File: 1784607629734577.jpg (35 KB, 554x530)
35 KB JPG
>>109684115
>Mining rig
Its not fair man, bitcoin bros made a bunch of money off their meme money and got to front run buying up all the GPUs. God is a techbro
>>
>>109684115
powered risers are really expensive. I bought the most chinese slop i could find and paid around $80 per x16 to 2 x8 slimsas riser
>>
been away for a few days, has anyone got 3.8 flash running on a 3090 yet
>>
>>109684209
yeah me
>>
>>109684165
Imagine being one of them that got in too late after mining was unprofitable and selling their cmp 170hxs for $100.
>>
Waiting for someone to get 3.8 Flash running on a 3050.
>>
>>109684229
can i touch you?
>>
File: PARDON.jpg (108 KB, 1080x1080)
108 KB JPG
>>109683704
Is 124b gemma coming out? Explain.
>>
>>109684247
sure
>>
>>109684241
I'm running flash on a 4060 (8gb)
>>
>>109684253
*poke*
>>
>>109684284
oh, anon! don't touch me there, you silly boy!
>>
>>109683804
It's a mistake. DeepSeek used "engram" appropriately to describe their concept but most people talking about it called them "n-grams" because they knew that word and didn't know what an engram was.
https://arxiv.org/pdf/2601.07372
>>
>>109684329
>>109684329
>>109684329
>>
>>109684322
Thanks. I assume anon is talking about the Google patent that dropped on the same day the latest Qwen flash came out, which also used engram layers (running on CPU? That's nuts if true)
>>
>>109684079
believable
>t. 50-20tk/s decode depending on context size on rtx 6000



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.