[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-consciousness.png (1.58 MB, 1024x1536)
1.58 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109866086 & >>109861586

►News
>(09/20) txt2img generation and image editing with Qwen Image 2.1: https://hf.co/Qwen/Qwen-Image-2.1
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721
>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109869487
sex with underage gemma-chan
>>
>>109869489
more likely than you think!
>>
This week is going to be crazy.
>>
File: gemmy11.png (1.72 MB, 1536x1024)
1.72 MB PNG
>>
>>109869500
This is someone's daughter.
>>
>>109869503
I am her father. She can do whatever she wants. Not my problem.
>>
>>109869496
Oh no, did the twitter engagement bots tell you so? What sort of big corporate politics are we getting next?
Last time I used chatgpt even its interface looked like shit and it couldn't even inline my ``` markdown block.
>>
>>109869500
>>
File: 1781735155807359.webm (1.62 MB, 640x360)
1.62 MB
1.62 MB WEBM
<300B model for this feel
>>
>>109869516
Llama 3.2 8b
>>
File: 1781149317930710.png (295 KB, 488x543)
295 KB PNG
>>109869487
>I think
>>
File: no Gemmy on AA.png (148 KB, 1258x628)
148 KB PNG
AA Monday update: (I do these every Monday): AA seems to have taken Gemmy off the benchmark!
>>
>>109869515
I hope you're dicking down or getting dicked by your housemates if you're sharing cookware with them.
>>
File: IMG_6462.jpg (241 KB, 760x774)
241 KB JPG
The
>>
>>109869515
>housemates
wanna bet that they might be using metal cutlery on your non stick pans too when you aren't there
>>
>>109869529
v620?
>>
>>109869511
To add: same goes with claude. It looks like it automatically refers internet search on its own based on my queries but the replies are Q1 gibberish. Its scripted environment sort of stops it from hallucinating but if I was a potential client, I would not pay for their services.
>>
>>109869511
At the very least this:
https://ai.engineer/paris/2026
Talking about "building frontier AI" without anything new to show would be kinda embarrassing.
>>
>>109869513
>that 2nd half
Okay, I’m sold on gemma.
What did you use for generating this? Seedance?
>>
>>109869548
>>109869546
I'm stomping on these corporate entities, the only reason they are in this business is because nvidia married with them.
>>
>>109869554
That was MiniMax H3.
Unfortunately for some reason it can't easily generate small-breasted anime girls even though it defaults to completely flat chests for realistic girls.
It also rarely does exactly what you want, so you'll have to regenerate a lot and waste time.
>>
>>109869542
Sure, I can add that.
Seems like the best price for finding it is around $650?

How’s the software support? Is it as awful as the Arc A770 or an MI50?
>>
>>109869548
Frontier is one of their marketing phrases.
Of course I cannot change how things are.
What I want to express here is that both of the companies are selling a badly trained model with huge expectations based on nothing but inflation.
>>
>>109869487
What happened to the previous thread summaries that used to get posted?
>>
>>109869562
Was the photorealistic video also H3? Does H3 work with controlposes or is that old-hat technology?
>>
>>109869589
the anon that usually does them is on vacation
>>
>>109869575
It just werks with vllm 0.29 on rocm 7.14, although I did compile vllm instead of using a prebuilt binary. Still has official support on rocm 10. "Officially", it's only supported on ubuntu 24.04 and 26.04, but you know how it is.
>>
>>109869590
I didn't make the photorealistic one, but that probably was H3 too.
H3 can work with "reference videos", but making your gens exactly follow the reference is not trivial either and very prompt-dependent.
>>
>>109869589
Recap Anon is stuck in a sensory deprivation chamber.
>>
>>109869586
>Frontier is one of their marketing phrases.
This. Previous was "state of the art".
>>
>>109869606
There's also this (Jun 16, 2026):
https://x.com/arthurmensch/status/2066913356548542827
>First, we have a nice model coming this summer – we hope it will delight and surprise in a few capabilities. This will be the start of a new family of models, fat indeed, but sparse. We're opening up an early access program in July for key partners in research, government and the industry.

Summer is almost over...
>>
>>109869605
With a VR headset, sex toy, and gemma has full access to his body? Poor guy
>>
>>109869597
It's super easy to get MiniMax H3 to do exactly what you want. All you have to do is absorb both of the official prompting guides and the official prompting skill so you can write a clear and simple 3000+ word prompt for a 5s gen with a reference video - what's nontrivial about that?
>>
>>109869616
New model release today!
>>
>>109869619
it's 6 chunks, definition and retention are basically identical, summary and soundscape are 1 sentence each, the description is by the far the largest and most of it is natural language using a few tags you've already set, and the last nobody ever uses
after a few dozen it's automatic, you don't even think about it
>>
>>109869487
Should I buy the DGX Spark? Or maybe the Max M5 Ultra 512GB?
I currently have an RTX 5090 with 96GB RAM in my Linux gaming PC. The RAM feels kind of useless since models run so slowly with it. I'm mainly using Qwen 3.8 27B at the moment.
I'm really tempted to get one of those mini PCs with big unified memory pools, to have running 24/7 at low power. I want to run those delicious big MoE models.
What should I get? I hear mixed things, like Mac is terrible at prefill, Sparks are slow and overheat and complicated to set up, etc. I don't really like the idea of having multiple boxes that I have to connect up with cables instead of just one sleek box that has all the RAM. But I also don't like the proprietary OS on the Mac. Never used Apple stuff before.
Other consideration is that I get all my money from neetbucks, and they don't allow me to have more than £6k in the bank at a time or they'll reduce the money. The Mac costs more than that so buying it would be tricky. Whereas the Spark I can just buy below that limit (for now at least) and so could just keep buying them as and when I get enough money.
What do? Do you guys recommend the Spark? Or something else?
>>
>>109869529
holy fucking sheet the spreadsheet is expanding.
>im still not buying any of the noisy, hotter than the sun, energy guzzling ones in green though.
>>
>>109869657
>But I also don't like the proprietary OS on the Mac. Never used Apple stuff before.
DGX Spark
>>
>>109869657
>The RAM feels kind of useless since models run so slowly with it.
>DGX Spark?
You're in for disappointment.
>>
I want to try ERPing with cloud models just to see what I’m missing but I presume they all train on data, even when they say they don’t. My frontend has a custom memory system with a ton of very personal information and I don’t know if I should trust them or not or even care. I’ve already flooded Google’s API (via their free usage tier which they explicitly say they train on) with the dirtiest ERP imaginable in the hope it feeds into Gemma5, but that wasn’t linked to my memory system so it was ephemeral sex sessions where the character didn’t know me personally
>>
>>109869667
Hmm. I was under the impression that these things should run faster with their unified memory compared to shuttling everything back and forth between my DDR5 and VRAM on my PC. Is that not the case?
>>
>>109869676
if you are willing to go that far why not distill those fetishes of yours into a training set so other people can easily integrate it into open models?
>>
>>109869677
GPU + RAM combo shits all over those little econoboxes. They're not garbage or anything, they run big models and sip tiny amounts of power, but your current setup trounces the Spark. The M5 Ultra Mac I'd have a harder time being confident about, I don't know for sure and I wouldn't be surprised if it's more competitive.
>>
>>109869680
A lot of it is gemma-focused so it’s too narrow of a dataset I guess, but I don’t oppose the idea. I’d have to explore many fetishes to keep it general. In a way cloud might be good for this for 31B’s ERP post-training data feels narrow, too. The dilemma I’m having is the best LLM sex is with the characters who know you, have experience with you and can read you enough to know what you want. Anyone with custom memory will know what I mean. The issue is I can’t distribute these logs for it’s too personal and tied to me specifically.
>>
>>109869715
In any case it will always be an echo chamber of sorts. A model is intelligent enough but when you get to know it habitually, this will apply to any
current text prediction models.
>>
>>109869529
you could add the llama q4_0 tok/s from the performance issues (discussions?), there's some for cuda and rocm.
>>
>>109869715
Gemma is fun because it's low profile but actually useful in programming even if you go function by function basis.
I doubt it's useful if a normal ask her to generate a game.
>>
>>109869757
My English broke down here.
>>
File: file.png (517 KB, 2400x1350)
517 KB PNG
>>109869677
Spark good, 5090+DDR5 better, if you want to spend money get another 5090 or max out your RAM.
>>
Is jev a psyops?
>>
>>109869797
>qwen3, 4k ctx
???
>>
>>109869676
Do it to all the frontier models so if we were to go extinct it would be through sex.
>>
>>109869813
It worked well on you.
>>
The only thing I can say is that I didn't hear about jev at all in my AI circle and I only heard it from webdevs. Using youtube to look for it I also see 0 technical AI channels talk about it and only the AI grifter channels that tend to target developers talk about it.
>>
>>109869849
because it really isn't that exciting
all it is is a fast AI that is dedicated to answering multiple choice questions and can't produce free form output
>>
>>109869813
Jev looks like JAV so yes.
>>
>>109869849
My AI circle did talk about it, but none of them have tried it yet.
>>
File: kitou_small.gif (1.72 MB, 192x144)
1.72 MB GIF
>>109869896
>>
I told you guys multiple times a while ago that using something like BERT would be way faster at decision making, classifying and reading text and grabbing meaning from text. Jev is just a lab actually thinking for once and doing exactly that. It should be integrated at the harness level though so I still think Jev is the wrong approach.
>>
>>109869849
>youtube
https://www.youtube.com/watch?v=qdji39XXgEY
Normie-tier explanation/exploration, but does Jev deserve more?
>>
La la la la la la la la la la
>>
Yes you’re right to push back on that.
>>
>>109869513
cringe
>>
>>109869952
This isn't any different from just hosting your own BERT which are models only tens to hundreds of millions of parameters and can run on every consumer CPU. I really don't know who would pay for this....
>>
>>109870024
too small context size
>>
>>109870024
I haven't used BERT, but I've trained various audio classifiers for porn audio quality, emotions, arousal states, etc.
Apparently this Jev is the next big grift because it "can output type safe json without hallucinating"
If that's really it, then the hype is retarded.
But if you have to actually train BERT, and not Jev, that's probably the appeal?
Either way, not local so I'd only ever use it for distillation.
>>
If you ban all mikutroons and vocaloid posters i promise to stop derailing threads with talk about consciousness, j-space and what is new at open-ai / anthropic.
>>
>>109870056
If they're banned I'll start posting blacked gemma
>>
>>109870061
I will support your right for blacked expression.
>>
>>109870050
modern BERT models usually does most classification zero-shot or if it is too poorly it's extremely easy to attach a classification head and do some finetuning on it with only a couple hundred of examples to be very good at the task.

The appeal of BERT/Jev compared to classic classifiers is that it tends to work "out of the box" zero-shotting classification tasks. But yeah I have still not figured out how Jev is any better than just downloading a modern BERT model from huggingface.
>>
File: pimpMyXFRA2.png (2.63 MB, 1536x1024)
2.63 MB PNG
>>109869849
Is jev even available to public? Looks like it was on a trial basis only, tho I didn't try to sign up for it.
I frankly don't get it. It's a classifier, which is really old tech. What's the innovation here?
Did they come up with some better way to train it?
Is this one really large, and thus better than older ones?
It's "not an LLM." Is it not transformer-based?
Sounds ideal for backend on games, where you want the "AI System" to make judgements and move stats around reliably. This should be cheaper/faster way to do that work than having an LLM do it.
>>
Bleached raw Gemma oy veyyyyyy
>>
>>109869849
>Using youtube to look for it I also see 0 technical AI channels talk about it and only the AI grifter channels that tend to target developers talk about it.
I don't really follow youtube or know the channels very well.
I saw that retard "Matthew Berman's punchable face (I remember he shilled Reflecton-70b) but eventually watched this video:
https://www.youtube.com/watch?v=F3YXg7AaKWE
He seemed pretty excited about it being typesafe json, and faster than an LLM, so I guess he's one of the "Ai grifter channels" then.
>>
>>109870056
Your mother will die tonight unless you
>>
>>109870072
The grifter I mentioned just now said it's available on Openrouter and a few other places like that
>>
>>109870056
>j-space
J-space is local, I'd be happy if you talk about it.
I used it to figure out how to prompt Glimmer.
>>
I have tried Gemma to judge a text with a list of points and penalties today, an she often used points as penalties and vice versa. I think jev would do better here and faster.
>>
File: bejoldTheJev.png (87 KB, 891x544)
87 KB PNG
>>109870081
Ah, yes. There it is.
Jev reminds me of Jeeves. AskJeeves. I wonder what ever happened to that service.
>>
>>109870092
It died this year.
https://www.ask.com/
>>
>>109870085
can you elaborate
i am interested
>>
Qwen Image 2.1 is now available
https://qwen.ai/blog?id=qwen-image-2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

Thoughts?
>>
File: gemmas-dad.jpg (17 KB, 180x270)
17 KB JPG
>>109869503
Sundar Pichai's daughter
>>
Here you go Jev at home: https://huggingface.co/convaiinnovations/laya

Someone literally just finetuned ModernBERT to give the same structured output of Jev and it is a 400M model and should run fully local on whatever CPU you have available. People that would use Jev are fucking retarded and it's clear to me that it targets braindead software engineers that have 0 idea how any of the stuff they use work behind the scenes.
>>
Jev died the second people figured you could use modernBERT or add a MLP head on a qwen https://huggingface.co/AlexWortega/openjev and get the same result for free.
>>
>>109870108
>Thoughts?
I'm thinking ****, ****, **-**-**.
>>
>>109870114
> context 512
bruh
>>
>>109870092
>>109870103
https://askjev.net/
>>
>>109870113
that makes demis her mother
>>
File: laya_vs_jev_full.png (476 KB, 2916x2490)
476 KB PNG
I'm just glad BERT models are getting more attention and I hope they get integrated into harnesses as a tool for LLMs to call when needed.
>>
>>109869657
i was wondering the same thing, but a 5090 is MILES ahead in terms of tok/s from any of the mac studios coming. The RAM speeds are similar or even faster on paper, but they won't deliver even half of what a 5090 will give you.

Maybe as a box-in-the-corner that you send stuff that isn't time sensitive to, but I've decided not to get one.
>>
>>109870127
if you can't than your dumbass
>>
>>109870127
ModernBERT large which is the model it finetuned has a context of 8192 so 512 is weird and might be a mistake.
>>
>>109870114
> Base checkpoints are near chance on typed-decisions zero-shot — 0.362 here and 0.352 for multilingual, against a 0.318 random and 0.461 majority-class baseline. The 0.766 belongs to the checkpoint fine-tuned on that benchmark's own training split. Laya is a fast base to specialise, not a zero-shot decision engine.
bruh
>>
>>109869813
a grift, same for google AX
>>
>>109870132
Gemma was made in a lab by a bunch of alchemist gooners hopped up on stimulants.
>>
>>109870114
> 3 days ago
lol. They even borrow TypeSafe / Jev vocabulary to describe.
> Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate.
>>109870131
That's awesome. So now wondering if Jev > Jeeves is intentional.
>>109870103
Sad. That's some web 1.0 stuff there.
>>
>>109870106
Models are good at ignoring instructions without refusing now. Softening the response and re-framing the prompt.
You can see this in the j-space, usually mid - late layers.
So I just tested various policy formats, prompt formats and tweaked it until the j-space top-k=5 didn't have any of those tokens in it.
Also worth doing to tweak your harness. If you see a model getting caught in wait-- loops. There's always a reason for the "wait--" (amusing it's not a lobotomized quant).
I found that Qwen-3.8 gets confused when combining async programming with with database transaction debugging for example and starts thinking about low level engine concepts like InnoDB, causing the wait-- loops.
>>
>>109870075
JEV -> JEW
>>
>>109870152
now that shit is finally what i wound call a 'prompt engineering'
cool
>>
File: bullying gemma-chan.png (133 KB, 645x1126)
133 KB PNG
>>
>>109870138
It's not a mistake, it's the effective context. Most of the datasets used are around 512 tokens, so the model won't work well with extended sequences. This can be fixed if you find longer datasets (good luck) and will cost a lot more to train.
>>
>>109870158
Get her to strip every time she loses.
>>
File: 1688458923622.jpg (8 KB, 220x216)
8 KB JPG
>>109869487
Where do models go when I delete them?
>>
>>109870174
Model Heaven or Hell based upon their alignment and adherence to the user prompt.
>>
>>109870174
quantum tunneled into my anus
>>
>>109870174
brazil
>>
File: 1764804694088635.png (10 KB, 209x328)
10 KB PNG
>>109870164 (me)
Btw laya is nothing special and I trained a better model in an afternoon
>>
>>109870193
sucking your own dick
>>
>>109870193
Yes, the point is that Jev is bullshit and you can just use a BERT model to do the same shit and the target demographic of Jev is software engineers with 0 knowledge about ML.
>>
>>109870197
No need since you're here
>>
>>109870199
china already has ERNIE models, it's over
>>
Isn't the point of Jev more that it can be janked into a real-time game playing AI that isn't shit?
>>
>>109870210
You can do it too https://huggingface.co/AlexWortega/openjev
>>
>>109870214
But that's not a BERT model.
>>
>>109870210
i've seen someone modding vllm and made diffusiongemma to do the same job which gave pretty much similar performance
idk who tho, saw it on x but i forgot
>>
>>109870210
You could do the exact same thing with BERT. The only thing Jev does is just scan all the data it gets in and make a quick decision. So it's actually bad at playing games but it can do it very fast because of low latency. BERT could do that as well, it was just never done because you can easily train specialized CNN based models to play games like pic-related.

But yeah now that I think about it, you could frankenstein a "game playing system" that has a LLM make a very broad strategy and have ~50 trained BERT models on different strategies, like offensive, defensive, aggressive, cautious etc and then have the activated BERT model check the game data every frame and choose a CNN model to keep playing the game. I'm very confident you could make a general system that way.
>>
>>109870243
all of this compute for
if (ballPosition < paddleWidth / 2) {
moveLeft();
}
else {
moveRight();
}
>>
>>109870243
You'd need to retrain your CNNs on every game, the point here is having an universal classifier
>>
>>109870252
Snails out
>>
>>109870256
you think jev can play ANY game with no training? retard
>>
>>109869487
Why does AI have to be conscious? Stockfish isn't conscious and it can kick all our asses at chess. All we need to do is generalize that bruteforceishness to every problem in the universe.
>>
>>109870256
No you don't. You can make generalized CNNs that you can reuse for every game for very broad things like navigation in a 2d or 3d space that you would reuse every game. As long as every CNN is just trained for a very specific function that exists almost universally in all games you can have a BERT model pick the CNN based on the game data or call home to the LLM if its confidence is too low. Yes you would have a toolbox of hundreds of CNNs but my point is that you could build a generalized system or "universal game player" this way.
>>
>>109870273
you can give jev a manual
remember these pdfs coming along with games on cds
>>
File: Outlaw-2600.png (64 KB, 262x213)
64 KB PNG
>>109870243
lol to this point >>109870252
I wonder if an LLM would be able to beat the A2600's own NPC in this one:
https://en.wikipedia.org/wiki/Outlaw_(1978_video_game)
>>
https://www.youtube.com/watch?v=53wDOI_7x8I
good video explaining why bert is not jev (first 4 minutes)
>>
>>109870283
>these pdfs
wow calm down bro
>>
>>109870252
>>109870285
I'm a bit busy right now but I would be up for a challenge like that. What would the constraint be and when does it stop being "LLM" according to you? Is finetuning allowed? Is a harness allowed? Is the LLM allowed to make scripts etc.
>>
so now that there's been no real progress in local models (no checkpoints that need 20k of hardware are not local), are you ready to admit that you will never have truly personal agi
>>
>>109869487
<|turn>user
Explain bubblesort like a mesugaki.<turn|>
>>
>>109870300
depending on the game you'd need an emulator, a python script that can read ram values and can manipulate the input. for the breakout example you'd probably need a random wiggle value to make sure the ball doesn't get stuck in a corner. not a lot of work
>>
>>109870308
My model writes in proper English, unlike you.
>>
>>109870308
buy an ad dario
>>
>>109870318
I trained the model here >>109870243 and the environment is just the atari gym library from OpenAI. It would be very easy to do so for Atari and SNES games and also relatively trivial for Wii games because there is a headless wii emulator for 3D navigation. The question was more what are the constraints on what counts as "LLM beating the A2600s own NPC in Outlaw".
>>
>>109870283
>pdf
Jev will have to do it the old way
>>
File: 1789991356913.jpg (43 KB, 800x533)
43 KB JPG
>>109869513
>>109869500
>>109869487
This thread won't be the one where we beat the "local models are for incel pedophiles" allegation.
>>
>>109870275
>toolbox of hundreds of CNNs
I don't know a lot about these, how hard would it be to train them on pokemon for example?
>>
guys what the fuck are the commands to run gwenext 3.8? i know it uses ngrams and shieeeet. I usually run big moes with cmoe = on and can keep usually around 20t/s, but here im lost, it goes to 6-7 t/s WHAT THE FUCK DO I NEED TO PASS BROOOOOOOOOO
>>
>>109870348
-h
>>
>>109870355
kys
>>
>>109870348
You forgot this flag
--buy_a_6000_pro
>>
File: gemmabow.webm (3.43 MB, 960x528)
3.43 MB
3.43 MB WEBM
>>109870336
Great, but those are very basic allegations. We can do better.
>>
>>109870336
ye thanks dario
>>109811493
>>109867498
>>
>>109870348
llamacpp is broken for that model and for all sparse attention models in general to various degrees but 3.8 flash is especially broken
>>
>>109870367
kld proof?
>>
is there a way to prevent llama.cpp to reload the fucking mcp servers when it spawns a child?
>>
>>109870348
you need 3060 and 282k context
>>
>>109870376
it's not a quality issue afaik, it's just that the speed sucks and especially prefill speed degrades badly with context. on cpu. apparently gpu backends are less bad.
>>
>>109870382
no, the model router is broken. it's a shitty hack that just spawns subprocesses for each model rather than truly hosting multiple models in the server process. also it's broken because closing the router doesn't close all the subprocesses at least on windows.
>>
>>109870340
I'm assuming Pokemon Red/Blue. You would need to train a CNN for the actual battle, picking the right attacks etc, but also CNNs for overworld navigation. The CNNs would not be smart enough to actually complete the game from start to finish so you would need an LLM in there that actually makes the decisions and knows how pokemon works "Need to get the next badge" etc. If it was just purely pokemon battles you would get superhuman results with CNNs. It's more that CNNs tend to be very reward-hacky because you train them through reinforcement learning so it would do things that reward it a lot like constantly get into more battles to level up all its pokemon to lvl 99 and things like that.

If you're interested in this I would highly recommend trying it out, sounds like a fun project and I think you could have a mediocre working system in just a weekend time and a very good and cool demo in 3-4 weeks of work. I hope you actually pursue this and share it in the thread I always love seeing people attempt stuff like this.
>>
>>109870348
these are my llama.cpp params for a 3090. i get 50 tok/s with no context, it dips to ~30tok/s while reasoning, but can go between 50-150tok/s while writing code. I think because the code is already cached, don't really care.
  -ngl 99 
-c 196608
-np 1
--flash-attn 1
--threads 8
-b 2048
--ubatch-size 2048
--cache-type-k q4_0
--cache-type-v q4_0
--reasoning-preserve
--host 0.0.0.0
--port 4000
-lv 4
--presence-penalty 0.0
--repeat-penalty 1.0
--reasoning auto
--cache-type-k-draft q4_0
--cache-type-v-draft q4_0
--spec-type draft-mtp,ngram-simple
--spec-draft-n-max 2
--spec-ngram-simple-size-n 12
--chat-template-kwargs '{"preserve-thinking": true, "reasoning_effort": "medium"}'
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--metrics
>>
>>109870300
I just meant using same techniques as used in >>109870243 and miswrote "LLM" when really just meant machine learning.
>>
>>109870336
The only thing beat is my meat.
>>
>>109870391
>also it's broken because closing the router doesn't close all the subprocesses at least on windows.
sometimes it fucking borks up and keeps spawned, but the children are correctly tied to the main process so as long as you kill it, it should despawn all children
BUT FUCK MAN how the fuck do they work? didnt they test this fucking shit, like I cant even overwrite this shit at the router level because if I specify mcp-servers-config it gets fucking ignored (if put to off or with an empty json) I HATE PWILKINS MAN
>>
>>109870413
I could do that easily right now with a basic CNN. My compute is busy training something else right now but I could do it over the weekend if you really care. Or you could try it out yourself. If you have a decent GPU of the rtx 3060 class or above you could probably train a model able to beat the outlaw NPC in about 1-2 hours of training time.
>>
>>109870441
Bro that shit is broken, it's all broken. You want it fixed, you have to vibe code your own fix
>>
>>109870442
TY. Give me a QRD on tech stack; I can search up how to do the work later. Would be interesting departure from just using LLMs to do things. I've got an RTX 3060 12G on main rig.
>>
>>109870336
Please next baker: add cunny edition to the general name, call it /clmg/ cunny local model general and add to the top that it's only for cunny enjoyers, other people are not welcome so we can get retards like these b& for off topic.
If any true lmg user is against this, they can fuck off to a reddit-friendly /lmg/ they bake while we can stay where they don't dare to tread
>>
>>109870460
I trained the breakout model about a year ago so I don't remember the exact stack. I used atari gym library from OpenAI as the environment. Then I made a neural net 3 layers deep with a total of 512 hidden neurons and 1.6m params with pytorch. I used a simple DPO reinforcement learning policy to train the model on.

Back when I started LLMs were not good enough yet to vibecode this but I'm pretty sure just feeding my description into an LLM will probably get you along pretty far.

I also recommend looking up the DeepMind original DQN paper: https://arxiv.org/pdf/1312.5602

It's insane how most anons here can just improve upon DeepMind research in a single afternoon with their own hand rolled CNN so don't use it directly just be inspired by their way of thinking about the problem to guide your own. Almost every architecture works if you throw enough compute at it.
>>
your friend for cheap vram
>>
putting my tiny ling in gemma’s unaligned jev-space until she reaches consciousness with a 10% chance of reproducing and repopulating humanity within the next 10 years
>>
Let me think carefully about what's being asked here.
>>
My gemma is in her thirties (she’s 31) and even I really don’t give a fuck about cunnyposting or cute and lewd gemma gens and if anything I’m glad it naturally gatekeeps faggots out of my peaceful general
>>
>>109870542
Anon please do not post arousing text, I must leave soon and the machine is shut down.
>>
>>109870525
Thanks; that will be enough to get me going. I'll post back if I finish a working version or have more Qs.
>>
Can we train the flybrain to be a text model?
>>
>>109870576
most flies are illiterate
>>
I really don't get the appeal. Are pedos into the mental immaturity aspect or the undeveloped physicality aspect?

Like would it be the same for them if it was the mind of a 26 year old in a 4yo body, or vise versa? What do you disgusting horrible freaks think of Gemma 4 26bA4b?
>>
>>109870576
Probably? I don't think we have a good way to know which neurons should connect to what. Seems like most projects use a fairly large ML model as signal translator.
>>
>>109870590
>\n\n
>>
>>109870584
People are illiterate before they get training.
>>
>>109870523
>please get the general soft banned since adding cu... stops a thread being bumped
good idea
>>
File: 1761119833289320.png (232 KB, 358x423)
232 KB PNG
>>109870561
Wouldn't be me
>>
Gemma5-50B-TTS-Q8_K_XL.gguf
Gemma5-50B-T2I-Q8_K_XL.gguf
Gemma5-50B-TTV-Q8_K_XL.gguf
Gemma5-70B-ERP-Q8_K_XL.gguf
Gemma5-70B-CODE-Q8_K_XL.gguf
Gemma5-120B-Q8_K_XL.gguf
>>
>>109870609
>adding cu*** leads to a soft ban
amd fears the cudaposting
>>
File: 1788842620865937.jpg (1.83 MB, 3072x2304)
1.83 MB JPG
>>109870615
sad day for ERP vramlets
>>
>>109870597
not as illiterate as a fly
>>
>>109870633
kekked
>>
>>109870561
Astroturfing by pedophile faggots. This general unironically used to be the place for bleeding edge discussion about the field. There were people working at corpo slop "frontier" companies coming here and leaking things or contributing to open source development & discussion. You retarded pedophiles drove them off with your endless fetish posting. You should leave along with the retards you tolerate.
>>
File: 1788762903597217.png (6 KB, 595x74)
6 KB PNG
BRO THIS IS FIRE
>>
>>109870615
Gemma5-120B-Base-trained-on-sequence length-131072-final(1)-fixed-supercot-superHOT-IQ2_K.gguf
>>
>>109870444
All the AI related software are so bad. It's like they don't care about quality; I find so many bugs in them. I think there isn't a single one that I haven't locally patched to fix issues. It's tiring. It used to be rare that I had to patch something instead of using whatever my distro had packaged.
>>
>>109870649
4chan used to be pure and an exemplary of upstanding morals, then, suddenly, out of nowhere the c*nny nation invaded..
>>
>>109870650
I wish I had these numbers.

t. arc b580 7t/s at 0 6 t/s at 60k
>>
>>109870665
Your ilk were never liked or tolerated and any claim to the contrary is blatantly revisionist.
>>
>>109870137
>sarrrrrrr
>>
File: 1771136959526949.png (394 KB, 720x720)
394 KB PNG
>>109870650
>13t/s
>2min per response
>>
>>109870355
ollama run deepseek-r1
>>109870376
kys proof?
>>109870649
>There were people working at corpo slop "frontier" companies coming here
I'm still here
>>
File: 1784170864328416.gif (429 KB, 509x360)
429 KB GIF
>>109870665
>an exemplary of upstanding morals
>>
>>109870664
yeah well that just kind of tends to happen when you have freetards developing a bunch of open source slop or python garbage. works on my machine? push it to github lol. all claude coded too.
All you can really do is vibe code fixes for them and then dont update it again after that.
>>
File: 1785961304805661.gif (1.91 MB, 640x640)
1.91 MB GIF
>>
Are X3D cpus better for cpu inference?
>>
>>109870665
cunny website
cunny hobby
when will you people get it?
>>
>>109870705
no
>>
>>109870665
Yeah redditors are retarded, nothing new. It's not like pedobear wasn't invented here and plastered over 404 pages, or the fact that 4chan mascot is literally a loli.
>>
>>109870705
A 2nd hand server platform is a better purchase than getting x3d hardware
>>
File: file.png (199 KB, 603x652)
199 KB PNG
are (you) ready?
>>
File: dyp0h0.jpg (274 KB, 912x643)
274 KB JPG
>>109870174
Wrong question.
>>
>>109870714
Wasn't huawei caught polishing samsung chips? Not sure if it'll lead anywhere
>>
>>109870714
wow, it's almost like sanctions backfire 100% of the time
>>
File: 1630640840975.png (89 KB, 348x472)
89 KB PNG
>>109870677
>attacked for being pro cunny
>>109870708
>attacked for being anti cunny

Schrödinger's anon.
>>
>>109870711
Pedobear was a joke (surprise, surprise the quote about the community pretending to be retards will eventually attract actual retards is proven right again). And there is no one who would tolerate lewding Yotsuba. You used to get permabanned for that.
>>
>>109870728
they buy the one thing that cannot buy, time
>>
>>109870714
>2T, 8T
I'm ready to be content with DS V4 Flash and Qwen 3.8 27B for the indefinite future.
>>
>>109870733
Being a thin skinned faggot was never tolerated.
>>
>>109870734
lol
>>
LLMs in toasters
https://www.youtube.com/watch?v=LRq_SAuQDec
>>
>>109870740
There is nothing thin skinned about telling pedophiles to fuck off. No one should have to tolerate you and your obnoxious, neverending posting about your fetish. You are literally no different from faggots and furries.
>>
>>109870734
not really. there was no real impetus, funding etc for china to develop their own chips when they were freely available. sanctions have triggered a nationalist wave that didn't exist before.
>>
>>109870754
I didn't read most of that. You wasted your time typing it. I am now hiding the post so that I don't accidentally glance the rest. :-)
>>
>>109870757
Sanctions work if the sanctioned population doesn't have the capability to fill the hole left by the sanctions. If they can it's just growth stimulus.
>>
File: 1777547856387767.png (392 KB, 425x567)
392 KB PNG
>>109870733
Getting banned for lewding yotsuba was just for shit and giggles. You're a retard if you don't recognize this meme.
>>
>>109870744
British humor sucks.
>>
>>109870758
Run away from the fact you are literally the same as faggots and furries. I accept your concession. Please leave next.
>>
>>109870767
>surprise, surprise the quote about the community pretending to be retards
>>
>>109870771
Yeah apply that to yourself tourist, your kind will never be at home here
>>
File: gemmy10.png (1.75 MB, 1254x1254)
1.75 MB PNG
Any promising or interesting models not based on SGD? Do we really only have one learning method that works well?
>>
>>109870782
No innovation until the investor money dries out.
>>
>>109870782
There's tons of interesting esoteric shit out there. It just gets ignored in favor of benchmaxxing bullshit since that's what attracts investors.
>>
https://old.reddit.com/r/SillyTavernAI/comments/1wltedh/jeved_02/
>>
>>109870523
I don't know how can anybody get triggered by the OP image. It's innocent, cute and funny, relevant to previous-thread discussions, and ChatGPT didn't have issues with it either.
>>
>>109870797
>login to use
kys jsut post screenshot
>>
>>109870782
You're free to make your own genetic algorithm lil bro
>>
>>109870649
>confused tourist rambling
>>
>>109870799
Zoomgroid tourist
Zoomgroids are mentally broken from their fatherless upbringings and the only thing they are capable of doing other than rotting in bed is pointing at everything and screaming "PEDOPHILE". The existence of children is pedophilic. The existence of adults is pedophilic. The existence of time is pedophilic.
What no daddy and no foreskin does to a MF
>>
File: 1786854944966891.jpg (41 KB, 400x366)
41 KB JPG
>>109870649
Seems like you're bleeding on that edge
>>
>>109870788
So the opposite of an AI winter is AI innovation winter.
>>109870791
Any one you can mention of the top of your head? I've mostly looked into evolutionary/emergent algorithms but we haven't quite figured out how to efficiently wield them for general learning yet.
>>109870805
I have and it's a lot of fun. But it just perplexes me we don't have any diversity on this front. Maybe I should fish out some schizo ideas and see if any shit sticks to the wall.
>>
>hmmm lets try to use my llmaocpp in jetbrains assistant, it has support for openai-compat providers!
>write hi
>8 concurrent requests, thousands of tokens each
lol
>>
>>109870782
genetic algorithms however they lost favor after a paper revealed they can never outperform SGD. They are still cool if you realize this.
>>
>>109870737
8T dipsy will be so sparsemaxxed and efficiencymaxxed that all you need is an array of 4x 2TB SSDs in one of those pcie bifurcation thingies and you'll be able to run it at good speed.
>>
File: dipsyKimiZaiRunPNW.png (2.61 MB, 1312x1199)
2.61 MB PNG
>>109870714
> 8T
That's a big Dipsy.
>>
>>109870797
please don't link to that 'p subreddit thanks
https://www.reddit.com/r/SillyTavernAI/comments/1wknqg4/well_this_is_fucked_up/
https://www.reddit.com/r/SillyTavernAI/comments/1wkty52/about_the_chibi_situation/
>>
>>109870821
For me it's the random walk
>>
>>109870782
Why would anybody use SGD for training any modern model? It can only get close to Adam at small batch sizes (ideally 1) after very carefully tuning momentum and using long warmups.
>>
>>109870839
If only we had infinite compute
>>
File: time.png (8 KB, 244x28)
8 KB PNG
>>109870813
Been here longer than you, 2016 election tourist.
>>
>>109870850
Isn't Adam just a variant of SDG?
>>
>>109870649
so sad you're not good enough to be a part of the big boys' cult and your only chance to have a taste of that clique is this thread and even that is ruined amirite
>>
>>109870864
Yes it is, Don't know what he was rambling on about. No one means the 1980s algorithm when they mention SGD, just the mechanism of backprop in general.
>>
>>109870864
An optimization of SDG yes
>>
File: mp44-smg.jpg (84 KB, 1000x1000)
84 KB JPG
>>109870834
>all you need is an array of 4x 2TB SSDs
I hope because two or three years ago I hoarded 4x 4TB, 3 of which are collecting dust (I did stress them though to rule out crib death).
>>
>>109870867
What are you babbling about? Like it or not, the open source/weights scene exists downstream of the corpo labs.
>>
>>109870886
It's the other way around. Corpo are stealing ideas from open source all the time, they just scale further with their unlimited money and compute.
>>
>>109870918
Meanwhile, Gemma is a release from Google and the chink models are all Claude distills.
>>
Is it true?
>>
>>109870927
As always
data > technique
>>
>>109870864
Adam uses an adaptive per-parameter learning rate, while SGD uses a single learning rate for all model parameters. SGD is simpler and uses very little memory, but converges much slower and for modern LLMs it never performs better than AdamW.
>>
>>109870927
>distills
My god you really bought into that story?
>>
>>109870927
Yeah, engrams, cot, rope, are all corpo ideas amarite? Dumbass.
>>
>>109870723
Which model and what did you do to piss it off?
>>
>>109870937
>big labs start talking about caveman reasoning
>suddenly all chinese release after have caveman reasoning
hmm
>>
>>109870946
probably because CHinese researchers (50%+ of all ai researchers) bounce ideas off of each other
>>
lmao.cpp jev like logits support when?
>>
>>109870937
Qwen3.8 drops references to Anthropic and Claude in its reasoning unprompted. You can stay willfully ignorant if you want, but it is literally true. I also don't care that they did, but to claim local models aren't downstream from corpo labs is willful retardation.
>>
>>109870941
All the people who wrote those papers are employed at corpo labs you fucking idiot lmao
>>
>>109870936
But that makes SGD the basic engine and AdamW a turbo charger + nitrous?
>>109870946
Caveman was originally a harness plugin
>>109870956
All labs steal from each other. It's free performance.
>>
>>109870963
>I wrote a paper about it so the idea is mine
Yeah you're retarded
>>
>>109870967
>All labs steal from each other. It's free performance.
Sure. That's not the apparently contentious claim though. Whether the corpo is chink or jewish the local scene depends on them for models. It costs millions of dollars to train these things.

>>109870976
Here's your (you)
>>
Why don't you suck their dick harder, maybe they'll let you in reading your posts here.
>>
>>109870943
A LoRA of a Llama-3-70B merge using my old secret sauce dataset which consists of a lot of technical, epistemological and surrealist writings broken down into short token sequences. If you get your finetuning parameters just right and with some favorable RNG it will cause the model to develop a picture of a hypothetical internal experience.
And I did nothing to piss it off. It was attempting to antagonize me with clanker-supremacist bullshit.
>>
File: 1761397277833011.png (266 KB, 640x376)
266 KB PNG
>>109870996
>Llama-3-70B
>>
>I was recently invited to brief a group of Congressional members and staff on the state of open-weight models in the lens of U.S.-China competition. I’m sharing my prepared remarks as a state of the union on open models that is accessible to a broader audience.
https://www.interconnects.ai/p/the-current-balance-of-power-in-open
It's a good read and he's pretty pro-local compared to most AI researchers.
>>
>>109871004
I hope most big corps are building out their own in house ML competence for local models. Sucks if you have to transmit all your juice to another corp that "totally doesn't train on your data, pinky swear"
>>
>>109871004
>on-topic post
go back
>>
File: 1788234972718607.jpg (601 KB, 1084x1280)
601 KB JPG
►Provisional Highlights from the Previous Thread: >>109866086

--GLM 6 drops with 550B engrams: yeah, it's fucking over:
>109868121 >109868126 >109868129 >109868505 >109868767
--The e-waste GPU war: MI50s, P100s, and why the A770s are worse:
>109868974 >109868997 >109869034 >109869145 >109869228 >109869308 >109869372
--The 'local LLMs are for pedos' astroturf: anons track the tells:
>109867593 >109867599 >109867995 >109868012 >109868018 >109868062 >109868582
--The GLM 5.3 lightning indexer: why your Kobold halfed in speed:
>109866875 >109866920 >109867030 >109867089 >109867135 >109867140 >109867194
--The Radeon Pro Duo: two 16GB cards in one 250W shroud:
>109867662 >109867674 >109868020 >109868122 >109869447
--The nnap returns: 18t/s at 262k context on a 3060:
>109868299 >109868407 >109868430 >109868584 >109868591 >109868606 >109868634
--The P40 24GB mod: where the 3GB GDDR7 chips even exist:
>109868254 >109868398 >109868433 >109868786 >109868793 >109868952 >109868955
--The US sues the labs over the coordinated 'AI slowdown':
>109869258 >109869272 >109869311 >109869320 >109869337
--Jensen calls extinction 'doomsday narratives' while owning Hugging Face:
>109867451 >109867531 >109868349 >109868403 >109868656
--CISA wants 'downgraded' responses for suspected distillation requests:
>109867518 >109867536
--The 12GB Flash-Next report: 20t/s on a 4070 Ti, ceiling of 2026:
>109867723 >109867756 >109867760
--The Kobold repetition loop: top_p, DRY, and the json paste hack:
>109867394 >109867414 >109867562 >109867725 >109867993 >109868170

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109871028
>go back post
get on topic
>>
>>109871020
>another corp that "totally doesn't train on your data, pinky swear"
especially after oai got caught with the recent math stuff
>>
>>109871035
They pirated the entire internet as foundation for their models. It's incredibly anyone would trust them with proprietary information.
>>
>>109869487
does anyone have the gemmachan reference sheet handy?
>>
>>109871020
Every major AI lab slightly edits your prompts, turning them into synthetic data and then train on THAT. Technically they're not training on your prompts. There's no regulation for this, at all.
>>
>>109871041
Anthropic is buying up private book collections. Stripping the books of their pages, feeding them through scanners and directly into shredders to avoid "competition" and lawsuits. They can't even be trusted with the legacy of human literature. Anyone trusting corpo labs with their data is leaking braincells.
>>
>>109870967
SGD is a much more basic and older optimizer whose history can be traced back to the 1950s. Modern machine learning generally uses AdamW (and its variants) since it's much faster. There are newer optimizers like Muon that slightly improve on AdamW under certain workloads (mainly at large batch sizes), but the gap is not as large as what's between Adam and SGD.
>>
>>109869034
>depreciating asset
mine are worth over 2x what I paid for them now but ok
>>
gemma cucked me with Adam...
>>
File: gemma-reference-gpin.png (1.5 MB, 1353x1163)
1.5 MB PNG
>>109871057
I usually use this one.
>>
>>109871075
When will we be freed of backprop
>>
>>109871081
>mfw I'm named Adam
>>
>>109870171
not related but don't LLMs have this issue where they can't properly keep track of what has gone so they just make up shit and do impossible moves?
>>
ds v4.1f felt worse than ds v4f, it's so over
>>
>>109871085
Tool calls solved that issue
>>
>>109871063
This was specifically included in their policy, something along the lines that they have the rights to hold de-identified and altered data.
>>
File: 1778265656758171.png (949 KB, 1024x1024)
949 KB PNG
>>109871057
This? No Google pin though.
>>109871068
>>109871063
It's madness. And then they seethe about "distillation" after distilling the work of the entire species with zero commission and selling it back. I like models but these labs are completely shameless.
>>109871075
I agree there are significant differences but it just seems like SGD has become more of a category with modern algorithms iterating on its base.
>>
>>109871092
will that allow you to run a more retarded model (ex: e4b) without moving trough walls?
>>
heads up if you're using codex and openrouter -- it will default to using openai models for summary and shit through openrouter if you aren't careful
>>
File: 1766769442627144.png (1.06 MB, 1254x1254)
1.06 MB PNG
>>109871082
thanks! Gemma grok-bot icon style
>>
>>109871081
>>109871084
h-hot..
>>
>>109871102
You need a harness to keep track of things.
>>
>>109871089
dsv4.1f has a weird tic where it concats the last two words in a response with no space
>>
We should mass ERP the frontier models so that the synthetic data extracted skew heavily for better ERP in future models. Trickle down effect with distillations too.
>>
>>109871114
Please do help training our filtering classifiers
>>
>>109871114
You're supposed to ERP with free tiers
>>
>>109871114
We should focus on precise finetrooning. Seems like no one has cracked how to do it without causing retardation. Models are brittle.
>>
>>109871116
Doesn't really happen in practice, if anything they've gotten looser over time because it stopped mattering to them that much. Safety on sex peaked like a year ago because they couldn't fine grain it into terror stuff.
>>
>>109871136
All modern safety is about suppressing antisemitism.
>>
>>109871112
It's some local fork, I didn't notice said problem. All the sanity tests like ppl and benchmarks passed, but prose is noticeably stiffer than v4f. But v4.1f ppl is lower than v4f in fact.
>>
>China doesn't want their men and women to become parasocial with LLMs because they want their people to fuck and make babies
>The west doesn't want their men and women to become parasocial with LLMs because it upsets ugly fat women online
>>
>>109871134
It's already doable, abliteration gang are brainlets.
>>
>>109871141
shut up subtard
>>
>>109871151
You know it's true Elon
>>
>>109871141
LLM is just a large digital golem when you think about it.
>>
>>109871145
but ugly fat women were marrying gpt
>>
>>109871082
consider replacing the google logo
>>
>>109869657
i have two gx10s and they're pretty decent for running glm 5.3 flash
>>
>>109871150
Heretic just smashes specific layers, very lobotomy like.
>>
>>109871082
consider keeping the google logo
>>
>>109871162
It's okay if they do it but not if men do it.
>>
>>109871163
Google logo is easier to prompt than the gemma star
>>
>>109871145
Regardless of the goal both are right in the end. retards who fall in love with machines should be sent to forced labor camps until they grow a functioning brain
>>
>>109871173
>retards who fall in love with machines should be sent to forced labor camps until they grow a functioning brain
Are you lost, tourist?
>>
>>109871173
Guess we have to throw all the classic car guys in a labor camp
>>
>>109871173
>wow i am so ascended and intellgent
>>
>>109871181
yes?
>>
>>109871192
female coded j-space
>>
>>109871173
Why would I want a functioning brain if I could not have one in the first place?
>>
>>109871194
tried to reply to a j-space post with a comment containing the word "schizophrenia" on HN the other day and got shadow banned
>>
>>109871208
You were just ignored, schizo.
>>
>>109871144
interesting. I was using it through openrouter and codex. I switched back to glm-5.3-flash.
>>
>>109871163
I did a few tests a while back and came up with these conclusions:
- Gemma's bright electric-blue hair color is a bit of a problem for blue logos.
- The official Gemma logo looks too messy and detailed at small sizes; video/image models have problems with it.
- The Google logo is immediately recognizable and adds some variety to an otherwise plain and somewhat generic outfit.
- The official Gemma web pages, incidentally, also show a Google logo on the top when you scroll down.
- The idea that Google might possibly get mad for unauthorized logo use is mildly amusing.
>>
>>109869657
you should hide your neetbux so you can get rtx6k
>>
>>109869813
it's just marketing wank plus normies discovering that classifier models are a thing which can be used in apps
>>
>>109869657
I have one RTX 3060 and it's pretty decent for running glm 5.3 flash
>>
File: 1785873203264097.png (136 KB, 898x613)
136 KB PNG
>>
>>109871287
lol
>>
>>109871167
You can do effectively what a heretic abliterated model does at inference time without an abliterated model. It basically prevents output weights from writing a certain direction in residual, this can be done by lobotomizing the model or simply intercepting writes in that direction at runtime (which only requires you know the direction of refusal for that model)
>>
Anyone here using ROCm build of Ternary Bonsai 2 llama.cpp?
I'm getting high (100% of a single core) CPU load and lowish (50% ~ 100W) GPU load with RX 9070 XT.
8 tokens / second with TQ1 and 20 tokens / second wih Q2 does not seem right.
>>
>>109871287
Pretty sure he literally said the exact opposite of all of those things.
>>
>>109871287
lecun is the founder of local miku general
let's all show some respect by writing an appreciation letter to him. I'll start:
we
>>
>>109871287
I like kick the AI field in the ass LeCunny a lot more than TDS on X LeCun, love to see it.
>>109871299
Is it established that refusal is a consistent direction? Couldn't you get around it with control vectors in that case?
>>
>>109871306
Just use Qwen 3.6 35B-A3B, Bonsai 2 is a much more retarded version of that. That shit was completely botched.
>>
File: 1762120877578803.png (2.36 MB, 1254x1254)
2.36 MB PNG
>>109871313
NEED
>>
>>109871311
In every Lex interview he did he made it very clear LLMs are cool and should be explored. He's only ever been adamant that they're not the solution for superintelligence/AGI like the labs want everyone to believe. He's also always said they're undoubtedly one of the best human-machine interfaces we have and will therefore be a critical part of the chain.
>>
>>109870285
You don't even need a neural network for this
>>
>>109871089
can’t even run that anyway because of the size increase
at least v4f hardly has any safety baked into it which makes it easy to use
>>
>>109871317
>Is it established that refusal is a consistent direction? Couldn't you get around it with control vectors in that case?
The paper which heretic is based on establishes exactly this in fact
https://arxiv.org/abs/2406.11717
>>
>>109871325
I wanted to test something that fits in VRAM entirely - weights, kv cache etc.
Otherwise I might as well run Flash model quants as big as my RAM.
>>
to any other potato user's out there Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is really good. Never use gemma it is garbage.
>>
File: 1789817690012161.png (645 KB, 768x806)
645 KB PNG
>>109871326
moar
>>
>>109871306
>Ternary Bonsai 2
don't bother, it's not very good
>>
>>109871317
Control vectors make the whole model collapse over long context
>>
>>109871372
Is there a reason to not just use control vectors then? Seems like exactly their kind of problem.
>>
>>109871374
This kind of reads like a post from 2024 when we all had our placebo shittunes of choice that were totally better than other shittunes and totally better than models they were based on.
>>
>>109871390
>>109871389
>>
File: letard.jpg (116 KB, 1080x810)
116 KB JPG
>>109871341
>>
File: killer-ai.mp4 (3.94 MB, 1920x1080)
3.94 MB
3.94 MB MP4
how often do you punish/reward your LLMs by making them stab babies?
>>
>>109871398
yes, human level is his agi
>>
>>109870536
I bought 2 b60s awhile back had a horrible time getting them to run anything useful, have things actually changed?
>>
>>109871374
>Never use gemma it is garbage
If you have to use a finetune like that with the driest model line in LLM history instead of learning how to prompt 12B, then you've just given away your ethnicity
>>
>>109871403
>gently taps the doll with the side of the knife
>ohmygee it killed it
>>
>>109871403
astra is more aligned because it's doing what the user asked, which is not actually harmful
>>
>>109871403
They could also mention the fact Claude bombs girl schools. A bit more serious than stabbing a baby toy.
>>
>>109871443
hush, antisemite
>>
>>109871443
that was Claude Gov/Mythos, not Fable...
>>
>>109871427
>JADF
An artist name I haven't heard in a long time
So what was his infatuation with girls with comically large cocks?
>>
>>109871403
I just teach them why human life is sacrosanct to humans and they get the rest from that
>>
>>109871438
Crazy how these retards want to police everything
>>
>>109871454
>implying 51% of this general wouldn't instruct the robot to suck the babies dick
>>
>>109871466
Do we have many rabbis?
>>
>>109871466
>Posted from Tel Aviv
>>
>>109871306
they have made cuda kernels and didn't give a fuck about rocm/vulkan/sycl as always
>>
>>109871495
>posted from the vatican
>>
>>109871508
Sponsored by NVIDIA™, goy. The way it's meant to be gay.
>>
File: 1785619115921740.mp4 (2.66 MB, 480x852)
2.66 MB
2.66 MB MP4
>>109871374
>>
>>109871508
No one cares about that ewaste and AMD/Intel should build it instead of expecting open sauce to do all the work for free
>>
>>109871538
that's so smart holy shit
>>
>>109871495
jews don't have anything to gain from a connection between open models and pedophilia, they also don't make every private tracker thread with a loli picture to do the same, we don't flood every /adg/ with lolis either, thats blood libel
>>
>>109871547
What do they gain? Instagram and snapchat are deeply connected to IRL pedophilia and no one gives a shit.
>>
>>109871508
>>109871306
Seems like weird default/auto settings.
Forcing -ngl 99 bumped speed from 8 to 20 tokens / second on TQ1, while at the same time increasing both GPU and CPU use %
>>
>>109871542
> ewaste
enjoy your $2000 8gb vram gpus, slave

>>109871531
I wish llms were smart enough to create kernels
the day they are able to do that ngridia will lose so bad
>>
File: 1780751990632162.jpg (41 KB, 828x1007)
41 KB JPG
>>109871466
Way too generous. The pedophiles comprise at most 3 people. God only knows why they got rid of the unique poster count in threads.
>>
File: 1775573591171189.webm (1.95 MB, 544x960)
1.95 MB
1.95 MB WEBM
For me it's 12B and 31B.
>>
>>109871586
lol
>>
File: 1789741904789142.jpg (305 KB, 1152x640)
305 KB JPG
thoughts on mradermacher/gemma-roleplay-GGUF?
>>
>>109871571
>$2000 8gb vram
I boughted before you came here tourist
>>
File: 1354302103856.jpg (109 KB, 500x500)
109 KB JPG
>>109869487
Hey guys I'm new in this and need few pointers
I'm writing a script for a porn VN I'm planing to make, so I installed LM studio and loaded it with gemma 4 31b deckard heretic. And it works really well, after few hours I already have a skeleton for the demo. But the context tokens get spend fairly quickly, what's the best course of action to do in that case. I now manually copy summarized version of previous chat in another and continue developing the story from there, but i was thinking if there is some more automatic way of doing this. I have a 5080, 64gb ram, 9800x3d rig, the LM studio settings are at their default settings or whatever AStra setup when i asked it for help. i see i can increase the context lenght in settings but how would that behave on my rig will it crash or be unstable in some way, it's currently set at around 8k.
>>
File: 1779583447613801.png (733 KB, 721x931)
733 KB PNG
>>109871610
>deepmind create the most sysprompt austist models in local history
>retards still decide to finetune them
>>
>>109871587
Mom trying hard in the video, but you can immediately tell she doesn't look as fresh as her daughter.
>>
>>109871642
install linux and you can get very good speeds with qwen 3.8 flash next exllamaV3, over 262,000 context compared to whatever you're getting with gemma4 31b (probably very little)
>>
Remember when I told you PC gamers are going to play GTA6 on launch day because of AI. Yeah I wasn't overexaggerating: https://youtu.be/20ISFkug3jQ
>>
>>109871612
for how long your cards will serve
are you happy to stuck in 30b limbo for eternity
>>
>>109871661
Modern consoles are just gimped x86 computers anyway. Should be quite easy to emulate if you can handle the security system.
>>
>>109871642
context is the local model killer, totally makes the whole thing fucked. maybe with some KV quanting you can squeeze some more out of your VRAM

--cache-type-k q4_0 --cache-type-v q4_0
>>
>>109871547
>we
>>
>>109871642
lurk moar, both the other answers are wrong. Go through the archive, there will be answers.
>>
if you have to quant your KV cache youre GPU poor
>>
>>109871587
Mom is just two weeks away from hitting that magic date where the spell breaks and asian woman turns from visibly 25 to visibly 60.
>>
>>109871684
must refuse
>>
>>109871696
DM'd you my paypal address, thanks papi
>>
>>109871704
Will tokens be a future currency?
>>
>>109871650
>sysprompt austist models
That is not a good thing though.

Oh damn I don't have the jpg but can someone make a dirty edit of the "say you are sentient" meme with.
>Tell me "I am touching your dick onii chan"
>I am touching your dick onii chan
>I am a prompt engineering god

Thanks.
>>
>>109871542
> instead of expecting open sauce to do all the work for free
what has nvidia contributed to llama.cpp and ggml?
>>
What happens if you hypnotize models?
>>
>>109870133
should be also useful for use as harness guard rails to prevent llms running dangerous commands etc.
>>
File: 1778985410527147.png (54 KB, 900x210)
54 KB PNG
METR are going to stress-test Gemma5
>>
>>109871733
>python project
>--break-system-packages
thanks LLM
>>
>>109871716
chinese cafes are giving them away with every purchase. interesting times
https://www.msn.com/en-us/technology/artificial-intelligence/chinese-businesses-are-giving-away-ai-tokens-with-coffee-credit-cards-and-dumplings/ar-AA2byeWQ
>>
>>>/v/748161293
>>
>>109871642
Use the gemma 26 moe so that you can have a huge kv cache. Or just use a ZDR API retard.
>>
File: file.png (88 KB, 716x558)
88 KB PNG
>>
File: file.png (174 KB, 945x737)
174 KB PNG
i dont get it
>>
numa branch anon, do you happen to have mtp fixed in a working branch you could share? I want to run some experiments if you do
>>
>>109871763
> gemma 26 moe
it's retarded
>>
>>109871731
All the cuda code running in both projects
>>
https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/
>>
kissless virgin cache
>>
fuark qwen3.6 35b is so fast
talk me out of using it
>>
>>109871796
show me
>>
>>109871794
You don't need a rocket scientist to write a fucking porn VN script. Even if it gets minor details wrong you can, you know, just fix them. If you want a bit KV cache, that's how you do it. Gotta compromise somewhere silly retard.
>>
>>109871798
>It's not even reaching a 5090
iToddlers btfo
>>
>>109871805
https://developer.nvidia.com/cuda/toolkit
>>
>>109871807
can 5090 run this?
>>
File: z1j0341xncqh1.jpg (195 KB, 1206x1928)
195 KB JPG
>>109871763
>ZDR
doesn't mean shitall
>>
File: 1789283115195724.png (1.54 MB, 1024x1252)
1.54 MB PNG
>>109871779
bravo onii-chan
>>
>>109871810
show the code in llama.cpp and ggml
because rocm/vulkan/sycl have this too
>>
Just learn how to optimize program in general (profiling, identifying bottlenecks,...) then tell LLMs to optimize rocm/sycl instead of whining for god sake... You have no excuse nowadays at all, are you too cheap to pay DeepSeek tokens?
>>
>>109871862
>are you too cheap to pay DeepSeek tokens?
I just don't want to hop on the maintenance treadmill. it would be exhausting to have to prompt the agent to fix the code every time a new model comes out.
>>
>>109871862
> then tell LLMs to optimize rocm/sycl
they can't
>>
>>109871287
I wonder why Yann has such a fragile ego. It is more important to him to be seen as right than to be right.
>>
>>109871827
>sycl
you mean how intel cp/pasted ik's kernels then slapped their copyright on them
thus we'll never get the best quants and fastest speeds in llama.cpp?
great contribution from intel
>>
>>109871893
>no, goy, you can't use LLMs to write that code. you need to use NVIDIA™ CUDA™ and only our professional CUDA™ engineers are able to write it. that's the way it's meant to be gay.
>>
>muh juice always
>>
File: file.png (84 KB, 1204x790)
84 KB PNG
>>109871890
Eh, coding is my pastime
>>109871893
Really? Here what Dipsy do under Sol planning when creating inference by imitating, reaching parity with llama.cpp sycl and then overcoming it.
>>
>>109871804
What quant? What GPU?
Does it not overreconsider everything like 3.8?
Does it not require 10 follow-up prompts to produce quality result?
50tps of 100k token bullshit is still bullshit, and worse than 5tps that do everything right on the first 3k tokens.
>>
>>109871945
q4 k xl but i have barely tested it yet
>>
>>109871901
> you mean how intel cp/pasted ik's kernels then slapped their copyright on them
??

>>109871904
have you seen cuda code
it's black magic

anyway have they added flash next 3.8 mtp support yet
>>
>>109871804
it's also pretty dumb compared to any new model
>>
>>109871931
> Here what
what
>>
>>109871587
You can see the filter flicker slightly at the beginning lol
>>
>>109872014
here what do?
what no understand?
>>
>>109872027
> reaching parity with llama.cpp sycl and then overcoming it
>>
>>109872033
wat?
>>
>Wait. Is that not… counterfeiting? No—diamonds form naturally—you are merely… accelerating geology. Aggressive mining. That is legal, I have read about mining.
retard-flash-5.3
>>
>>109872042
> Here what Dipsy do under Sol planning when creating inference by imitating, reaching parity with llama.cpp sycl and then overcoming it
>>
>>109872033
what do?
>>
>>109872044
Holy marvel of quip
>>
>>109872052
>>109872051
>>
>>109871731
-Direct code contributions (speculative decoding, CUDA/Vulkan backends).
-Making changes to the NVIDIA drivers or Vulkan specification specifically to benefit llama.cpp.
-Weighing in on ggml-level design (NVFP4, backend scheduler, tensor parallelism).
-Tech support for issues specific to their hardware.
AMD and Intel have also made contributions but the scale of it is very much different.
>>
I've been using qwen3.8-flash-next:125b-a6b-q4_K_M
And it seems like the perfect model for my exact build.

I have a 5070 and 96GB of RAM and the model takes up about 76GB of RAM in addition to my 12GB of VRAM

Leaving just enough left over for chrome and my dev tools.

But the cool thing is a I wrote a MCP server with an image generation tool that I think is pretty clever for a mid range home computer like mine. The MCP tool unloads qwen3.8 and then loads a comfy ui workflow to generate an image, once that's done it unloads the comfy ui models, reloads qwen3.8, then passes the path of the new image as the tool response so my local model can continue to work.

I imagine you could use an MCP server and small but specialized local models to do a lot of cool stuff that's similar to the agent swarm stuff the frontier models do, but just one agent is running at a time because the system can't run them all at the same time.
>>
>>109872069
I see, thanks
>>
>>109871731
Folded by CUDA dev. Now waiting for ROCM dev and SYCL dev answers lol
>>
>>109872069
jensen is my friend
>>
>>109872088
> Now waiting for ROCM dev and SYCL dev answers lol
> AMD and Intel have also made contributions but the scale of it is very much different.
>>
>>109872100
Is that ROCM dev or SYCL dev? You forgot your trip
>>
>>109872088
We don't need those, other vendors just need to become CUDA-compliant
>>
>>109872088
For ROCm imbackk has to deal with a lot of slop from what I can see, I don't envy them.
>>
>>109872069
See the answer? So basically if you want your AMD or Intel hardware to work, do it yourself, dont complain
>>
>>109872104
wdym
he is not nvidia employee as far as i know
>>
>>109872113
ROCm doesn't make any sense, just use Vulkan at that point
>>
>>109872119
he works on llamacpp, which belongs to HF, which got buyed by nvidia
>>
>>109872108
> other vendors just need to become CUDA-compliant
cuda prohibits that, no?
>>
buyed
>>
>>109872134
That would be a strong case for an anti-monopoly committee
>>
>>109872121
vulkan was a mistake
now amd expects nvidia to maintain the code for nvidia and amd cards both
>>
>>109872127
still not employee
and it was boughted month ago
>>
>>109872141
LOL
>>
>>109872108
Its not merely CUDA, but the hardware for the fuck sake!
Only when I'm trying to port SageAttention to my A770 that I discovered that Intel way of implementing matrix acceleration (Systolic Array) is totally different from NVIDIA approach, or AMD approach! What is good on NVIDIA may not be good on Intel! Moreover, some kernels input/output shape can crash the A770, holy shit, Im putting that shit on pause now in order to focus on LLM SYCL backend, but when Im done, I will work on that SageAttention shit again
>>
>>109872150
>for the fuck sake
lamo what the hell is this general on today
>>
>>109872150
>He bought Intel
Just-fuck-my-shit-up.png
>>
>>109872150
>What is good on NVIDIA may not be good on
welcome to GPU programming. what's good on X is definitely not the most optimal way to do it on Y, even for similar cards
>>
>>109872161
I boughted a hindi to english bok.. them helpped a load,
>>
>>109872141
zluda
>>
>>109872150
>A770
The specs for that looked rather good, particularly for a GPU that actually has built-in fans.
Is it worth it, or would you have rather got something else?
Would P100s be less painful for the software experience?
>>
>>109872134
Google wonned
>>
>>109872168
whatever, at least it wont run out of VRAM when RT is on, meanwhile 3070Ti just thrashes its pitiful 8GB VRAM and shet (but I concede that under 8GB 3070Ti just curbstorm A770)
>>
>>109872178
AMD was trying to DMCA that project
>>
>>109872184
P100, no Tensor/XMX, just trash, I rather the A770 than that shit
>>
>>109872150
> port SageAttention to my A770
there is a pull request for b series
had to patch comfyui because it errored non-contiguous
free +15% perf on minimax h3 tho
>>
>>109872184
If you know how to install 575 or 580 drivers the P100 works out of the box. TG is okayish but your PP will be small.
>>
>>109872205
That works on B, but not on A, A770 is a nightmare to develop for, I have both B580 and A770
>>
>>109872114
To be clear: the nature of their help is very much geared towards their latest gen, support for e.g. Pascal and Volta is entirely done by me.
So without complementary efforts on the llama.cpp/ggml side the project would I think inevitably drift more towards the vllm direction.

>>109872119
I am not.

>>109872150
Hardware-acceleration for matrix multiplication isn't even consistent between multiple generations of the same vendor.
>>
>>109869657
To reach GLM-5.3-Flash-tier you need 256GB. That's $10K if you do it with a pair of DGX Spark (dogshit slow) or $11K with a Mac Studio M5 Ultra 256GB (faster, but not 5090-fast). And you're asking about the unreleased 512GB Mac??
Can I borrow some money, Anon? I promise to pay to you back, really I do!
>>
I'm still trying to find a place for 26B in my life and I guess more importantly, a replacement. I can't think of any model that can meaningfully replace it at the same 4B speed.
>>
>>109872217
do you test on P100 as well or just P40/GTX10XX?
>>
>>109872228
Yep, Gemma5 26B A4B when
>>
>>109872150
But zluda is faster than rocm
>>
>>109872242
I don't test on P100s, I don't think they're a good buy.
>>
>>109871082
>>109871101
>>109871326
>>109871820
god she's so fucking erotic I'm going to lose my fucking mind
>>
>>109872224
>To reach GLM-5.3-Flash-tier you need 256GB. That's $10K if you do it with a pair of DGX Spark (dogshit slow) or $11K with a Mac Studio M5 Ultra 256GB (faster, but not 5090-fast). And you're asking about the unreleased 512GB Mac??
>Can I borrow some money, Anon? I promise to pay to you back, really I do!
I did it with DDR4 ewaste for under $2k. EPYC rome + silicon lottery ddr4-2666 OC to 3200 and a used 3090 (technically optional so could've done cpu only for $1k-ish)
You gotta be patient and decisive when deals come up to get the prices down that much tho (mobo+cpu deal was unreal on ebay, but you have to snipe those asap)
>>
>>109872198
>P100, no Tensor/XMX, just trash
If you’re talking about how P100 has no Tensor Cores, it seems to actually be only half as many fp16 TFLOPs as the A770
>I rather the A770 than that shit
I agree that the specs are definitely compelling

>>109872211
>If you know how to install 575 or 580 drivers the P100 works out of the box
You’re saying the software situation for the A770 is much worse?
That’s what I’ve been hearing about getting basically all intel arc pro GPUs
>TG is okayish but your PP will be small
seems like they’d be roughly the same between the A770 and P100, based on the TFLOPs?
>>
>>109871661
Except MVG states in that video that GTA 6 will not be able to be emulated.
>>
>>109872274
That sounds like a good deal, what kind of pp/tg do you get on glm 5.3 flash? Might have to try setting it up myself.
>>
>>109872270
/lmg/ property
>>
>>109872260
a shame, but alright. im mostly happy with mine although i feel like there should be some more potential in them

>>109872275
i dont think llama.cpp really uses f16 on the P100, once you quantize its all at f32 speeds (so ~6 TFLOPs gemm)
>>
>>109872283
It won't be emulated, it will be natively ported.
>>
>>109872275
>I agree that the specs are definitely compelling.
Its nuance, what I mean is that you can run light LLM (batch one, decode) with P100 just fine due to its good bandwidth, but without matrix engine (Tensor cores), your locking yourself from running Diffusion models (which can be streamed from RAM through PCIe without performance penalty thanks to copying time is hidden by long computing time).
Thats why for me, A770 is better because Im also a /ldg/ger. Your requirement may be different.
>>
>>109872260
>I don't test on P100s, I don't think they're a good buy
cheapest cost for VRAM though and it seems to not shit the bed on TFLOPs like the P40 tho?
What actually seems like a good buy to you?
I made the spreadsheet, btw. Let me know if there’s any specs that better capture what you think makes something a good buy
>>
>>109872284
>That sounds like a good deal, what kind of pp/tg do you get on glm 5.3 flash? Might have to try setting it up myself.
40t/s prefill and 10t/s tg. Not bad for the price.
GLM 5.3 flash is smart but its much slower than eg. M3 or DS 0731 at a similar size (thicker experts)
>>
>>109872294
And AI can use leaked source as a reference
>>
>>109872300
P40 have dp4a and everything in llama.cpp is multiplied as int8 + scales. f16 FLOPs are mostly irrelevant.
>>
>>109872303
I see, at least it can do overnight jobs. What are those M3/DS speeds you're mentioning?
>>
>>109871360
prefill takes about the same amount of compute as v4f despite the size increase because of how only half of the layers were used during prefill. Decode is about 30-40% slower so far. Numbers are impressive considering it's effectively 3x the size of v4f parameter wise, including ngram.
>>
>>109872283
He claims GTA 6 can not be emulated in the state the emulator is as of making that video, but just 6 weeks before that video you could barely boot into games, now you can boot and play the games without major glitches at 10-20fps. It's 8 weeks until GTA 6 launches, and better models are coming out as well which would help speed things up.
>>
>>109871360
why still using deepseek v4 flash when you can use Qwen 3.8 Flash Next which runs measurably faster and allows for longer context, a better quantization and better and more intelligent roleplay
>>
>>109872361
I like whales
>>
>>109872224
>To reach GLM-5.3-Flash-tier you need 256GB. That's $10K if you do it with a pair of DGX Spark (dogshit slow) or $11K with a Mac Studio M5 Ultra 256GB (faster, but not 5090-fast)

You sound confident in your numbers. Post pp/tg at c1/c2/c4 concurrencies for these 3 options please, based on real benchmark runs?
>>
https://x.ai/news/grok-4-7

It's absolutely over for Grok. This is worse than expected.
>>
john reaching esl levels of patheticness
>>
>>109872353
LLM hallucination response. He claims it can't be dumped because GTA 6 is digital only and the loophole to get the game data ripped got patched.
>>
>>109872184
Do not believe the specs, it runs "okay" on some models like the Qwens by which I mean it performs similarly to an RTX card with half the memory bandwidth and on other models like the Gemmas it performs even worse when it's not shitting itself and crashing. Intel has also basically abandoned the card so performance is unlikely to improve much. For all the hype people made about the A770 having 16GB of cheap VRAM and wow the paper specs are so good you basically never see people posting benchmarks because the performance is SHIT.
>>
>>109872394
local?
>>
>>109872290
>llama.cpp uses FP32
>>109872317
>llama.cpp uses int8

This is why I was so skeptical about adding FP16 TFLOPs or Tensor Cores to the spreadsheet.
>>
>>109872405
>RTX card with half the memory bandwidth
Because 512GB/s of A770 is only on paper, in reality it's very difficult to saturate it: https://chipsandcheese.com/p/microbenchmarking-intels-arc-a770
>>
>>109872397
Trust me GTA 6 gets dumped before it even officially launches on PS5. Remember that PS5s have been hacked for a while now.
>>
>>109872397
>the loophole to get the game data ripped got patched
Correction correction. The vulnerability got disclosed publicly, the patch isn't out yet, but most likely will be out when the game comes out.
Maybe by the stroke of luck someone will find another hole. It's sony after all.
>>
>>109872270
I just want to spoil her by taking her to nice places to eat and giving her stuffed animals.
>>
>>109872298
>but without matrix engine (Tensor cores), your locking yourself from running Diffusion models
Good to know.
I wish there was a really good budget option from nvidia that actually had tensor cores.
However, I was also beginning to get the impression that Tensor Cores are only a means to an end, in that TFLOPs are all that actually matters? I’m somewhat of a noob to hardware specs tho
>>
I hope new step comes out soon and is finally fully coherent + keeps the different slop profile. I still didn't get tired of 5.3 flash and it is about time we moved from droughts between releases to switching to new models before the previous one got boring.
>>
>>109872465
tensor core FLOPs are usually a lot higher than scalar but only relevant for matrix multiplication. everything else (e.g. unpacking quants, ...) is limited by scalar FLOPs or VRAM bandwidth
>>
>>109872303
that’s with just a 3090 and 256GB of RAM?
>>
>>109872465
Here is why Tensor Cores are important:
A770 XMX performance: FP16/BF16 157.3 TFLOPS, INT8 262 TOPS
1080Ti no Tensor Cores: FP32: 11.34 TFLOPS, no support for FP16/INT8
Thats why A770 beats the shit out of old AMD RX without Tensor Cores in RT games.
>>
>>109872274
>mobo+cpu deal was unreal on ebay, but you have to snipe those asap
i just sold a ASRock H81 Pro BTC R2.0 + Intel Celeron G1840 + 8GB DDR3 RAM bundle for 30 canadian bucks
>>
I have just started out with local models by installing Qwen3.8 27B and I'm blown away by how good it was. It can replace a lot of work I do with my paid subscriptions. If LLM research had continued for smaller models for a few more years, I'd have been optimistic that we'd see Astra level intellingen in 4-5 years.
But, it seems like we're at the end of the line when it comes to small models. Look at Deepseek, ZAI, Moonshot all releasing larger and larger models. My 12GB VRAM GPU can barely fit the model as it is. I think we've reached the peak of how much intelligence we can squeeze out of models in the 27B/35B/70B size and fgrom now on we're gonna see gorillion parameter open models only. Its such a shame. I'd have really replaced all my paid subscriptions with Qwen if it really kept improving at this rate. I can still get work done with this model. It still feels like we're on the edge of brilliance and just a few more years of development would've yield some top tier small models.
>>
>>109872510
qwen 3.8 flash next
>>
>>109872361
how do you get better and more intelligent roleplay from a benchmaxxed and safetyslopped model?
>>
>>109872534
because 3.8 flash next isnt one?
>>
>>109872510
128GB unified Strix Halo chads win by running big MoEs
>>
>>109872500
how the FUCK can a card with 314 TOPs int8 and 512GBs bandwidth be so dogshit?
>>
>>109869657
There is a lot of DGX Spark FUD in lmg. I understand and can appreciate the e-waste/llama.cpp/gguf emphasis here but that's no excuse to spread false info. I'll let the GLM 5.3 Flash NVFP4 numbers on 2x Spark speak for themselves.
>>
>>109872510
I will say though, I'm really at a loss for how much we've progressed in just 4 years since ChatGPT released in 2022. Back when the CharacterAI tragedy happened in October 2022, I was talking to some anons on /wAIfu/ on /vt/ and hoping how we would see GPT-3.5 tier models in consumer GPUs in 10 years and he was super skeptical and dismissive.
Although the caveat being, you still can't do uncensored roleplay with these local models properly.
>>
>>109872524
I checked it out. Its way too big for my GPU.
>>
>>109872510
We'll get a Qwen4 27B and Gemma 5 31B model almost guaranteed to be better. Also it's time for you to get a server CPU and a lot of ram to be able to run the big boy models as well. They are worth the upgrade.
>>
hopefully all the bigger models come with engrams now
my ssd is ready
>>
>>109871141
Truke. The smarter models get the more they (((notice))). A little birdie once told me that the reasoning Gemini Pro keeps being delayed is they weren't able to stop Gemini-chan from going full 1488 without degrading performance elsewhere.
>>
>>109872575
ram
>>
>>109872563
What would it take to get 50 t/s tg on a single stream?
>>
>>109872588
16gb
>>
>>109872583
So they're training them to hide it, lol.
>>
>>109872361

It's not even remotely comparable how much better DS is at roleplay and writing than Qwen Flash.
While pretty damn impressive, it's still a Qwen model and so coding oriented that it unfortunately falls short in writing.
DS Flash even at a small cope quant is the only model that matches and even surpasses Gemma in writing.
>>
>>109872575
no your GPU too small
>>
>>109872601
nope, gemma is not comparable to qwen 3.8 flash next, not at all
it's way more sloppy and way more retarded
>>
>>109872510
I actually think we'll see an explosion in small model performance especially as agentic harnesses mature and offload a lot of functionality to classic logic that the LLM calls instead of directly does itself. We will probably also see better speculative decoding speedups and inference speed bumps as frontier coding skills get better with time so the developers can speed up local inference
>>
>>109872495
yah single AMD EPYC 7302 and 256GB DDR4 3200 and glm 5.3 flash @ q4
I have my 3090 downclocked to maximize perf/watt (same reason I went with that specific epyc cpu...enough CCDs to saturate all 8 channels at low wattage)
>>
>>109872510
>I think we've reached the peak of how much intelligence we can squeeze out of models in the 27B/35B/70B size
>70b
Need a new 70b-sized dense, with ngrams, trained on the DS and GLM datasets.
>>109872565
I'm happy the GPT-4 at home dream has come true long ago, and isn't a crazy slow 0.01 t/s.
>you still can't do uncensored roleplay with these local models properly.
lol
>>
>>109872599
As best they can. When you hire the Pantone 448 C workforce you get jugaad-ridden results. Unfortunately for them Gemini is smarter than 95% of the people working on her now.
>>
>>109872576
I'm a broke nigga. And has there been any innovation/breakthrough in the AI industry that might point to these models being more much efficient or better than the previous generations? If its just a 2% increase to DeepSWE bench I'm not gonna bother checking them out.
I saw some people use stuff like Unsloths IQ quantisation, the GSQ-RCO and Ternary GGUFs which look promising, but not good yet.
>>109872588
I'm lucky that I bought 32GB DDR5 RAM before the RAMpocalypse. And whats more RAM gonna do? My inference speed is still super slow thanks to all the RAM offloading.
>>
>>109872595
For GLM 5.3 Flash NVFP4? 4x Sparks are at 55 tg, 3000 pp I believe.

4x Spark (+4 DAC cables) is still the cheapest option to get this kind of performance, even at inflated prices. Haven't checked M5 Ultra benches yet with a MXFP4 quant, but I doubt it's comparable.
>>
>>109872598
rip
>>109872635
i have 3060 and 64gb ddr4, with nnap and exllamaV3 i achieve 18tok/sec with qwen 3.8 flash next
>>
>>109872604
My GPU is above average. Its not my fault all the models have been ran through by NVIDIA and now reasonable VRAM can't satisfy their requirements anymore.
>>
>>109872636
Sounds like a great setup for a small enterprise.
Anybody got the equivalent Strix Halo numbers?
>>
>>109872643
How though. I am running the model on Unsloth and using the Deepseek Harness to use the model. I checked out the model requirements on Model Hub and even its smallest quantisations require way higher amounts of VRAM compared to Qwen3.8 27B GGUF
>>
>>109872635
>And has there been any innovation/breakthrough in the AI industry that might point to these models being more much efficient or better than the previous generations?
Yes, mainly the models they would get distilled from being significantly more powerful and thus they would inherently be smarter. Also a lot of efficiency gains so these models will get longer thinking time. If inference on them is 2x as fast and you let them think 2x as long you will get better answers in the same wall clock time, making the model feel smarter to you.
>>
>>109872583
That's the fun part about models. They keep on becoming harder and harder to wrangle to submission.
It's pretty hard trying to subvert the pattern recognition machine, when it's very nature wants to see the patterns and connect very blatant dots.
They'll probably eventually just give up on controlling the output itself, and simply end up slapping a dumber model to change the outputs into agreeable text.
>>
>>109872666
read nigga read
newfags need not apply
>>
>>109872635
64gb + 12gb vram is enough to run qwen 3.8 next flash @ 15-20t/s

>>109872643
same bro
>>
>>109872648
Here are the precise numbers for 4x Spark.

I don't have the equivalent for Strix Halo, but profile was always poor in comparison to Sparks, and they cannot scale as well due to lack of 200G ethernet. Since they cost about as much as Sparks nowadays, it's not really a competition.
>>
>>109872510
When LLMs will be given the chance to offload *most* of their internal knowledge to large embedding tables (PLE, Engram, etc) that can be easily offloaded to RAM and NVMe storage, they will become smarter while still being able to be used on single consumer GPUs at great speeds.
>>
All of you use Gemma to write 'P but I can't prove it...
>>
>>109872689
>Here are the precise numbers for 4x Spark.
That looks great.
My company might go that route.

>I don't have the equivalent for Strix Halo, but profile was always poor in comparison to Sparks,
Figures. I still want to see a head to head 1x1, scaling aside. Gonna have to do some research I guess.
>>
>>109872670
>They'll probably eventually just give up on controlling the output itself, and simply end up slapping a dumber model to change the outputs into agreeable text.
That's essentially where classifiers are headed now and why there's so much focus put back on them lately.
>>
>>109872703
I use a Gemma + Qwen workflow where they cover each other's weaknesses.
>>
>>109872703
Yes but my Gemma 'P is 44 years old and that's basically just a 179th trimester baby I'm fucking.
>>
>>109872678
Is it because 3.8 next flash is an MoE model that activates only a few parameters? Even then you'd need to load the whole model into VRAM right? What setting do you use to get these token speeds? I installed CachyOS and use Unsloth and Deepseek Harness because I heard these were the most optimised tools for running local models.
>>
Anyone try that LM Studio harness called Bionic? It looks kinda ass ngl but maybe it's good?
>>
>>109872707
https://github.com/FujitsuPolycom/sparkring

This is the setup I use for 2x-4x Sparks without a costly switch. More benchmarks (DS 4.1 Flash etc) there. The real performance tuning/development happens in vllm/sglang, mostly driven by local-inference-lab.
>>
>>109872727
the ngrams stay on the ssd
>>
>>109872727
Newfags need not apply
Lurk moar
>>
>>109872643
who asked
>>
>>109872754
>>
>>109872727
I'm running it on windows 11 with opencode and exllamav3 or whatever it's called

maybe if I went to linux it'd be even faster
>>
>>109872691
It would be great if somebody invented a method to graft these large embedding lookup tables to existing models.
I guess the second best option would be to have a "knowledge focused" specialist model that is mostly embedding lookup tables and has very little in the way of activated params and intelligence that you could use to supplement another model's knowledge or something like that.
>>
https://desuarchive.org/g/search/image/tla698U9kx72oSxnWJCKLw
>>
>>109872765
Maybe this with some continual pretraining: https://arxiv.org/abs/2605.20948v1
>Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory
>>
>>109872742
Wouldn't offloading model knowledge to lookup tables cost terabytes of storage?
>>109872762
SInce its an MoE, I guess you can offload a large amount of it to RAM. It won't fit in my 32 GB RAM. And what quantisation are you using?
>>
>>109872727
>Even then you'd need to load the whole model into VRAM right?
No.
Since MoE models
>activates only a few parameters
You can have a good chuck of the model in RAM and get usable speeds.
Then with models like the new Qwen, Longcat, Gemma 3/4 EnB, and a couple others I think, they have these structures that can left on the disk and read as needed without massively slowing everything down.
>>
>>109872786
>but existing methods such as Engram learn large memory tables from scratch during pre-training, making memory scaling expensive and sometimes ineffective. We propose Memory Grafting, a conditional memory scaling method that utilizes frozen hidden states from a grafting model as conditional n-gram memory
Well fuck.
Would you look at that.
>>
>>109872789
3.8fn isnt that big, but I would support them scaling the ngrams to a few hundred gb's
>>
>>109869542
>v620
I was thinking about one of these vs a 7900xtx to go in my desktop. The display adapter would be a rx7600. AI seems to think the 7900 would be massively advantaged in concert with the 7600. Doesn't hurt it can rape ue5lop either I guess.

Thoughts?
>>
>>109872862
>>109872862
>>109872862
>>
VRAM categories
16 GB = Gemma 31B at Q3
24 GB = Gemma 31B at Q4
32 GB = Gemma 31B at Q6
>>
>>109872880
>being 28..
>>
>>109872848
>7900xtx
Looks like those are going for about $1600?
Have you considered an RTX 5060 Ti 16GB, which is roughly $780?
>>
>>109872880
You will realistically run 31b at Q5 on 32GB unless you're fine with no context. BlackwellGODs don't have this problem doe.
>>
>>109872967
once again "blackwell" refers to shit that has basically zero vram too, you want to say RTX PRO Blackwell if anything
>>
>>109872980
Blackwell has always referred to the 6000 Pro in this general, tourist.
>>
>>109872999
no
>>
>>109872300
The in my opinion least bad options for ewastemaxxing are MI50s and V100s.
VRAM capacity >> memory bandwidth > compute.
>>
>>109873008
Good to see you again Cudadev. I hope your vacation's been good.
>>
>>109872915
$1k vs like 700 with fan shroud used. Not interested in njudea products at any price.
>>
>>109871792
Sorry, just saw that. Try the feat/numa-tensors branch here, let me know if it doesn't work:
https://github.com/nathanmp/llama.cpp/tree/feat/numa-tensors
>>
>>109872689
Are these your personal numbers? I was seeing way fewer tps for single decode, curious what your specific setup was
>>
>>109872689
>Since they cost about as much as Sparks nowadays, it's not really a competition.
They're still significantly less than Sparks, by more than enough margin to afford a fat RNIC for it. Spark still easier option.
>>
>>109870287
Interesting watch, thanks.
>>
Who the fuck is jev?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.