[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1786921335027350.mp4 (268 KB, 736x576)
268 KB
268 KB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109580312 >>109577974

►News
>(08/17) Local anon quanted dispy 0731, he might share it
>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119
>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185
>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3
>(08/14) Qwen3.8-27B released: https://hf.co/Qwen/Qwen3.8-27B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
Gemmaballs
Kimisex
Egypt Won
Dario's Little St James Vacation
Thread Culture
E1B IQ1_XXS jeet posters
>>
>>109585352
finally a good breakfast
>>
>>109585361
la la la la la
>>
70b dense
>>
>buy dgx spark
>selfhost dsv4 flash 0731 q2
>code with opencode webui on phone
life is good
>>
>>109585361
Cockbench
Sugmabench
Moesissies
Aicgjeets
Chinkshills
Benchmaxxed
>>
>>109585352
FUCKING BRAT
>>
File: Gemma-Chan Recap.png (505 KB, 1024x1024)
505 KB PNG
►Recent Highlights from the Previous Thread: >>109580312

--Paper (old): Stealing Reasoning Traces from Proprietary LLM APIs:
>109581906 >109581959 >109582058 >109582084 >109582275
--DeepSeek-V4 performance gains using J-Space prompting harness:
>109581982 >109582009 >109582782 >109583051 >109583100
--Hardware strategies for running GLM 5.2 and microwave RF troubleshooting:
>109583170 >109583197 >109583234 >109583270 >109583293 >109583317 >109583310 >109583331 >109583461 >109583592 >109583660
--Report on escalating methods for implementing long-term AI memory:
>109580716 >109581060 >109581075 >109581362
--SillyTavern formatting issues with DSV4 preserved thinking and llama.cpp debugging:
>109583192 >109583311 >109583364
--Proposal for a fixer model to stabilize DeepSeek model quantized to 0.25-bit using LittleBit:
>109582181 >109582302 >109582327 >109582370 >109582438 >109582573 >109583032
--Comparing non-coding Pareto models using a performance scatter plot:
>109582198 >109582305 >109582315 >109582319
--Anon attempts running K3 on low-end hardware via MoE optimizations:
>109583572 >109583579 >109583588 >109583605 >109583647 >109583807 >109583828
--Mixed reactions to llama.cpp's new Electron-based desktop app PR:
>109581992 >109582006 >109582021 >109582033 >109582018 >109582083 >109582118
--Comparing Qwen3.8 benchmarks and debating AMD vs Nvidia GPU value:
>109581232 >109581307 >109581616 >109581694 >109581694 >109581813
--Archiving models and using ModelScope to avoid HuggingFace restrictions:
>109582785 >109582841 >109582853 >109582900
--Qwen 3.8 27B KLD analysis:
>109583871
--Managing power and heat for budget server hardware:
>109584235 >109584302
--Logs:
>109580886 >109581223 >109581383 >109582033 >109582337 >109582497 >109582252 >109582574 >109583249 >109583363 >109583995
--Teto, Miku (free space):
>109581071 >109582756

►Recent Highlight Posts from the Previous Thread: >>109581706

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Muse Spark 1.2 open source soon!
>>
Someone posted a thing a couple weeks back about some bullshit to make windows stop eating up VRAM in non-display GPUs, anyone have that info or a link handy?
>>
>>109585410
>DeepSeek-V4 performance gains using J-Space prompting harness:
Repo issues are turned off. Why?
>>
qwen 3.8 27B writes truly horrific prose. So many 2-3 word sentences.

It is however, less retarded than gemma I think I'll just use it to help revise scenarios and lore documents and leave the writing to gemma.
>>
Any Gemma 26B uncensored recommendations?
>>
>>109585459
It's a work horse.
>>
>>109585471
"uncensored" != "afraid of the word sex"
>>
YOU MADE THE NEW THREAD BECAUSE THE INDIAN IS AT WORK IN LONDON AND HE COULDN'T DO IT
>>
>>109585471
llmfan ones are good
>>
I don't like that people are trying to earn a quick buck with videcoding. It's scammy. They don't know what they're doing and make money by selling something to other people who also don't know what they're doing.
>>
>>109585498
Boo hoo
>>
>>109585498
You don't know how to make the rubber in your mouse, but you're okay buying a mouse from someone who simply assembled the mouse?
>>
>>109585507
Complex software in real-life must also be maintained, it's not just a one-off thing. Long-term, if you don't understand what's going on with the code, you're going to be in trouble.
>>
https://huggingface.co/Sao10K/Lmao_life_updates
>last update 24/10/25
sirs did /ourbrahmin/ sao dieded? ?
>>
> Use "Baka" (obviously).
Gemma-chan.....
>>
>>109585513
In 5 years AI will be good enough to make any kind of software. Unironically.
>>
>>109585507
The individual people in the supply chain know what they're doing. Vibe coders are selling a broken piece of software with vulnerabilities to a doctor's office so that patient data gets stolen more easily.
>>
>>109585498
vibecoding just lowered the bar for code production. Instead of requiring a software engineer, you can now have a decent piece of simple code without knowing much about computers. Obviously, you still want engineers for code with high stakes.
The fact that low quality shit exists, for example chinesium tools, doesn't make good tools obsolete or useless. If you just need to hammer a nail in your wall to hang a picture frame, you have an option not to spend way too much on it, though.
>>
>They know how to operate the machine that puts the rubber in the mouse, but don't know how to make the rubber by hand.
Horrible.
>>
>>109585498
There's a hundred times more jeets making big bux off AI on social media, collectively damaging human cognition. That should be your bigger gripe, not vibecoding.
>>
>tfw no brat gf
>>
>>109585523
If we'll get to that point, those vibe-coders won't really be needed anymore and the AI (or more likely a group of AI agents) will deal with every step of the chain including maintenance, given initial (or updated) requirements.
>>
>>109585516
Finetooners, especially those in the imagegen world, are another class of "people" that hopefully AI model advances will wipe away soon.
>>
Is Posting Consciousture in Local Models unLawful?
>>
>>109585536
it's also nice for writing throwaway or self contained pieces of software with verification built in.

A experienced engineer spending 5% of their brainpower on these problems with an LLM probably outdoes slopmerchants spending 100% of their non existent brains

I had several web scrapers that were broken yesterday and I just pointed my coding agent to fix them, which in one case found an unauthed graphql endpoint in a json bundle which I would not have the patience to test myself.

In the past it wouldn't be worth my time to write the scraper, I have a lot more open ended problems to solve/figure out what's even worth solving. That task would have been delegated with all the hassle of communication/human context transfer and time and manual verification.

However there's still so many long horizon/open ended stuff that I would prefer delegating to a human, even if they aren't particularly amazing at software dev. I don't see these going away in the short term but the whole thing has made me incredibly careful of bad hires who would have their negative effects amplified by AI.
>>
j-space tuning for transseek
https://github.com/Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report/blob/main/README.md
>>
You wouldn't download a nursing handjob simulator.
>>
>>109585619
https://www.amazon.com/Educational-Building-Robotics-Engineering-Programming/dp/B0D478M8G8
>>
>>109585563
You'd probably still want some kind of project manager, but those would continue to be done by attractive young women with communication skills. The current wave of vibecoders with an over-exaggerated sense of self-importance have their use only in generating vast amounts of training data to get us to that point.
>>
File: file.png (101 KB, 528x464)
101 KB PNG
>>109585642
do not penis in metal claw
>>
File: file.png (23 KB, 1225x140)
23 KB PNG
>>109585619
Working on it.
>>
>>109585660
I'm going to fuck the claw.
>>
Im kinda confused guys, when I download a huggingface gguf it says for example 10gb but when I download it and look at the webbrowser downloader it only says 9.something gb, smaller than the original 10gb, I looked into it and apparently GiB is a thing which makes me wonder if I have a 16gb vram card, how much "actual" gb vram do I have? Am I getting less or more???????
>>
>>109585701
A lot of that is water weight, and packaging. Be sure to weigh it in the store.
>>
>>109585701
1TB SSDs are 1000x1000x1000x1000 bytes, not 1024x1024x1024x1024 which (((you’re led to believe)))
>>
How many of you role-play with gemma 4 with thinking on or off?
>>
>>109585745
>lobotomizing my wife
Never
>>
>>109585745
No reasoning is more schizo and good for quick fucks but reasoning is obviously better for helping you with stuff and stronger orgasms
>>
>gguf
Your life could be so much better but you choose mediocrity.
>>
>>109585764
Not now Yann
>>
>>109585764
Tell us your wisdom.
>>
>>109585763
>and stronger orgasms
Is that on the hf model card
>>
>>109585761
Wife? Dude, she is like what 30b parameters? Good luck with her even remembering what she is wearing.
>>
>>109585601
because it's ai psychosis
>>
>>109585701
model and ram sizes are normally in GiB I think, while storage is GB.
>>
File: J-Space.png (625 KB, 759x563)
625 KB PNG
>>109585601
DSV4 Flash analysis of the report:
>While the J-Space Cognition Suite claims to enhance DeepSeek V4's performance to the level of Claude Fable, the reality is more nuanced and its fit for your specific task is questionable.
>What J-Space Actually Does
>The J-Space suite is an inference-time control layer; it doesn't change model weights but adds a protocol for managing reasoning, verification, and state across tasks. It stems from legitimate Anthropic research on "J-space" – a manipulable internal reasoning subspace within models.
>The suite's benchmark results do show a performance boost, with the most notable gains being:
>NL2Repo: 54.2 70.2 (+16.0)
>DeepSWE: 54.4 67.4 (+13.0)
>HLE (w/ tools): 51.5 60.6 (+9.1)
>However, it's critical to note that gains are task-dependent, strongest on tasks with state drift or incomplete verification. It does not create knowledge the model lacks; it only helps access what's already there
>>
>>109585796
As I said in the last thread:
>exl3
>I'm able to run a higher fidelity quant (5 bits vs Q4) with a larger context window (98k vs 112k with vision) and better or equal inference performance. 10/10 would use it again
>>
>>109585796
exl4
>>
>>109585832
But where is anything to do with j-space in his codebase?
https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6/blob/main/j-space/scripts/jspace.py
I didn't run an llm on it, but looks like it's just a bunch of shitty skills.md files
And
>The operating effects have been reproduced across the DeepSeek, Qwen, GLM, GPT, and Claude model families. The suite does not depend on a vendor-specific API, tokenizer, hidden-state probe, or training recipe. Its portable unit is the protocol: workspace loading, selective routing, state externalization, verification, and recovery.
https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6#cross-model-reproducibility--%E8%B7%A8%E6%A8%A1%E5%9E%8B%E5%A4%8D%E7%8E%B0
They're not doing any probes or steering on cloudcuck apis
>>
>>109585873
>But where is anything to do with j-space in his codebase?
>tricked by chink lies
What are you, new?
>>
>>109585398
How is performance and dipsy retardation level at q2?
>>
File: sorry.png (44 KB, 1500x360)
44 KB PNG
>retard-chan
guess I'll have to try Qwen then
>>
>>109585352
>>109585410
kys pedo
>>
Where's the Q0.1 dipsy quant?
>>
>>109585964
Oh no, there's 1 (one) pedo itt!
>>
>>109585964
Most people in this general have an aversion to well endowed grown women
>>
>>109585516
>*deaded
fixed your typo saar
Sad, he made some good tunes during llama 2 era. F may he rest in peace.
>>
>>109585352
I tried roleplaying with Gemma 4 31B in Italian the other day, and while the model is better than the average, it still reads like an American speaking it as a second language. Same slop phrases I see in English too. It doesn't even know the Italian saying "non c'è cosa più divina che..." and its variants.
>>
File: 29.gif (9 KB, 300x100)
9 KB GIF
>>109585964
You're on the wrong website.
>>
>>109586011
> noo gemma doesn't speak my third world language!
mamma mia!
>>
>>109585976
Saar! I’ll have you know, I like my women(male) quite well endowed.
>>
>>109585896
it's usable. just /compact more often to avoid quality loss
>>
File: 1765984069739923.png (1.03 MB, 800x1296)
1.03 MB PNG
>>109585802
Why would she be wearing anything?
>>
>>109586032
Yet it's supposedly the best open-weights model for translation. Those relying on it for East Asian languages or even using it for learning them might never know if the model is actually writing like a native.
>>
>>109586011
Was the reasoning in Italian or English? GLM 5 when thinking in English wrote like a Western even in Japanese so it's important to check this first.
>>
File: 1787017018478288.png (12 KB, 449x109)
12 KB PNG
remember: no magic
>>
>>109586071
>Those relying on it for East Asian languages
if you want to learn chinese then literally any of the chinese models like qwen or deepseek would likely be better than gemma
>>
>>109585498
The people who don't appreciate good software are their own punishment.
>>
>>109585352
I really need to install h3
>>
>>109585536
>nstead of requiring a software engineer, you can now have a decent piece of simple code without knowing much about computers.
Nodejs already did this. Vibecoding isn't that interesting.
>>
>>109586087
Remember no russian
>>
>>109586114
python did it first
>>
>>109586088
Gemma's Chinese is pretty good though? She even manages to avoid sounding like a Northerner unlike Dipsy.
>>
>>109585536
Yeah. A bunch of people in my company are vibe coding tools and dashboards for their respective areas/sectors and it's working really well.
Granted, when they need to go a bit further than just the basics they can't really wrangle the LLM into getting there, and that's where we (I, really) come in.
>>
>>109586134
>when they need to go a bit further than just the basics they can't really wrangle the LLM into getting there, and that's where we (I, really) come in
And in a year or two, you will be useless.
>>
>>109586083
Gemma 4 will always think in English regardless of the conversation language or any instruction you might add at any depth, unless you try to prefill the reasoning with something else.
>>
I still don't really get MoE models and active parameters. Like why can I run an entire 27b dense model on its own, but when it's like a 32b or whatever I can only run 8b active params? why not like 32b and 20b active or some shit like that?
>>
>>109586158
Early MoEs were more balanced like that, but then DeepSeek showed you can drop the active count and increase the number of experts and you can train big models for really cheap. It's all about saving money.
>>
>>109586158
moe is meant to make things run cheaper/on worse hardware. Not much sense in allowing you to run a 30b model on hardware that can run 20b params.
>>
>>109586120
Python's ecosystem still had a culture that valued correctness. Nodejs is this gay disco of bad ideas.
>>
Well, Qwen started editing .cu files to try and get exl3 installed...
But at least it knows how to edit files I guess
>>
>>109586141
Been hearing that for a couple of years now.
I am become the AI wrangler.
Fucking about with local models from the early days as it turns out actually made me better at understanding, using, and implementing systems based on these damn things.
I just need less work on my plate so that I can sit down and map all the stuff we are using API models for that we could be running a tiny ass local model instead. Shit like the most basic of OCRs.
Blessed be continuous batching btw.

>>109586158
It's not that you can't. You could. There are parameters you can send to most loaders/backends to control the number of activated experts, and you could simply activate them all, but it still wouldn't perform the same as an equivalent dense model because the engineering of the neural networks and the training itself are completely different.
MoE is a form of sparsity. A way to try and squeeze as much intelligence as possible while using less resources essentially.
If you can get a model that, when compared to the dense equivalent, uses the same memory footprint, is 50% as intelligent, and runs 10x faster, that's a big win for the guys serving millions of concurrent requests.
>>
>>109586143
After a few turns my Gemma started thinking in Norwegian.
>>
>>109586202
Buddy what's the problem? Clone tabbyAPI and then run start.bat it does all the rest
>>
>>109586202
> Viking Gemma
did it affect her character? still bratty?
>>
File: gemma4_ita.png (559 KB, 1456x1901)
559 KB PNG
>>109586143
Picrel with prefill and an extra depth-0 system instruction to make her think longer.
>>
>>109586214
meant for
>>109586194
>>
>>109586214
>randomly pulls random ass model the author decided on
>>
>>109586214
Nope, I want the LLM to do it this time.
Looks like we've gone back down to torch 2.6 (even with Qwen) for some reason, and seem to be modifying exllamav3 src to make it work
>>
>>109586202
>>109586143
I just insert the language into the reasoning section in the template
>>
>>109586220
Qwen got it built / running.
But no way I'd use this outside of testing. It's modified the entire codebase
>>
>>109586216
Very much so, although my facility with the language is a bit limited so I have a hard time telling.
>>
>>109586242
sometimes python shit is just like that
back when i was still willing to use anything written in python i had to modify almost all pyshit software to get it to actually run
most of those modifications involved just seeing where errors appeared and deleting the code that had errors and stuff would start working kek
nowadays, no python for me. that garbage is obsolete now that llama.cpp exists
>>
>>109586241
I was using her as language practice so I didn't really care. Basically she would make fun of me and give me a correction if I answered in English or got some grammar wrong.
>>
>>109585974
I'm here too!
>>
Finally used the auto undervolt linux script on my 5090, it's quite nice limiting it at 460W through nvidia-smi but with almost stock performances.
>>
>>109586255
>nowadays, no python for me. that garbage is obsolete now that llama.cpp exists
That's been me for the past year, but anons kept shilling exl3 lol
42t/s for gemma-4 with tensor parallel, still slower than llama.cpp on ampere
>>
File: snapshot.jpg (158 KB, 1344x768)
158 KB JPG
wat
https://x.com/huggingface/status/2089673018737869242
>>
File: 1691406293259052.png (399 KB, 680x548)
399 KB PNG
>>109585352
So how good is qwen 3.8 for cooming?? any logs?
>>
File: what the hell is this.jpg (173 KB, 1344x768)
173 KB JPG
>>109586325
>>
>>109586337
Gemma-chan!
>>
>>109586242
Prompt better next time
>>
>>109586325
DavidAU is responsible for at least half of them
>>
>>109586332
From what I read it traded general understanding with agentic capabilities, and rp needs that so it's a regression vs 3.6.
>>
brehs what's a good tts model that runs on cpu?
>>
So Mistral is fucking dead? They could've come out with Mistra Creative and it would've saved their whole company with coomer bucks alone.
>>
>>109586409
What happened to that Le Fat or whatver the name was?
Did they never release that?
>>
>>109586409
summer isn't over..
>>
>>109586409
No idea, but I give them until end of the year at least to know if they give a shit anymore.
>>
>>109586395
all tts models run on cpu if you have enough time but I've used voxcpm and it's alright. it's far from being realtime on cpu though.
vibevoice is even slower, so dont bother with that
>>
>>109586409
Let them cook, literally, France is too hot to work right now but you need to thrust.
>>
>>109585976
Uh what is mommy characters? Those are quite popular too. Those sessions are likely just more personal to the user and less likely to be shared.
>>
>>109586409
Yes, it's over.
https://x.com/sophiamyang/status/2089394537495998659
>Career update: I'm leaving @MistralAI [...]
>>
>>109586418
It never existed in the first place. It was a shitposting campaign complete with meme graphs to mock Fable.
>>
File: 1780224566837761.png (937 KB, 1080x1351)
937 KB PNG
Is Qwen 3.8 unusable for anyone else? It spent over 3 hours on the same task 3.6, Gemma 4, Glimmer completed within like 20 minutes. And that's on medium thinking.
>>
>>109586409
>Mistral Creative
Simply making the models horny isn't the way and isn't enough, though. They doubled down on that when they released Ministral later on, but besides that and somewhat fresh prose, the model was just retarded.
>>
>>109586465
That was pre-fable no?
Regardless, shame.
I hope they don't just disappear and actually release some dense models for the GPU rich crowd.
>>
>>109586469
I left it overnight to fix 12 bugs and it got to #7 at 2t/s
>>
>>109586469
You noob, you need to know to set reasoning to medium from the jinja... it's xhigh by default for benchmarks
>>
>>109586469
idk it's been crunching for over 24 hours for me but not because of excessive thinking, just because it gets 0.1x speed at >30k context
>>
>>109586448
she doesn't look like she left in anger or anything
>>
>>109586483
>>109586487
Yeah, 3.6 was great on Strix Halo but this thing kind of fucking sucks. I'll try it on a 4090 instead.
>>
File: 141893431_p0_master1200.jpg (663 KB, 900x1200)
663 KB JPG
>>109586409
Mistral got France safetymaxxed.
>>
>>109586428
do these benefit from some shit gpu? like some older thing with 4GB or smth? is it worth getting an older cheaper one ran in x1 lanes just for tts?
>>
>>109586409
My guess is that their best elements are being poached left and right by unlimited money big actors.
>>
>>109586194
>exl3
stop trying to force this meme, i implore you
>>
>>109586497
>>109586487
I think the only models that really do well with long context with regards to
speed are deepseek and mimo. I mean it's pretty easy to tell since those *were* (now only mimo is) the only model that could offer $0.0028 per 1M cached input.
Dense attention crap is completely fucking unusable for local, even qwen 35b. Unless you can just brute force your way past the slowness by using GPUs but if you're spend that kind of money might as well buy enough ram to fit deepseek.
>>109586505
voxcpm is less than 4GB so it probably would benefit from basically any gpu
>>
>>109586495
Nobody writes angry career update posts if they want to get hired somewhere else.
>>
File: cai.png (10 KB, 512x512)
10 KB PNG
>>109586470
>Simply making the models horny isn't the way and isn't enough, though.
Do you have ANY-FUCKING-IDEA the demand, the absolute DEMAND there is for this shit? There has been literal wars fought and six figures tossed into people's hands just from clickbaiting alternatives alone.
>>
>>109586325
hf are definitely headed for a fall, the storage and transfers fees aren't sustainable
>>
>>109586548
but they're just begging?
>>
>>109586534
sure but you can infer from the way she wrote it, if it was on bad terms she wouldn't be even mentioning liking the company
>>
>>109586536
Why isn't there a local model that is as good as cai for erp? I thought that coomers were the ones pushing the technological frontier?
>>
>>109586524
I'm thinking you're right. DS4 even at a low quant using antirez's server has been the most consistently good local LLM experience for me on Strix Halo
>>
>>109586536
I'm only saying that models that are just ultra-horny get old quickly. People (even those who just roleplay with local models) don't like Gemma 4 simply because it's a somewhat horny model relative to the competition.
>>
>>109586548
As long as they burn through their funding money they're fine, but I agree overall, this is insanely generous and I fail to see how they can recoup the huge hardware cost they have currently.
So the platform will obviously tighten the rules.
>>
>>109586325
>3m models
>look inside
>literally who quanters doing quants that already exist
>schizo tunes
>qwen3 4b coder 1M context saar beliv me uploads
>>
File: buddy.png (96 KB, 498x291)
96 KB PNG
>>109586574
>I thought that coomers were the ones pushing the technological frontier?
You mean us?
>>
>>109586524
>voxcpm is less than 4GB so it probably would benefit from basically any gpu
their safetensors file has 4.58GB so nah. alsio I'm getting around 50GB/s on my 4800 DDR5 (laptop) while the rx6400 4GB has 128GB/s but 64bit memory bus. so twice + smth faster. prolly needs 8gb and at least 128bit width, for more speed
>>
OMG GUYZZZZZ my new gpu is arriving WITHIN 24 HOURS!!!!!!!!
I will ascend from vramlethood to vramchad
>>
>>109586622
What did you get, bro?
>>
>>109586624
r9700 ai pro
>>
>>109586628
oof
>>
>>109586628
>32GB
Nice, I'm not sure how AMD cards fare with AI stuff but at least you got a nice amount of VRAM there.
>>
>>109586632
wtf does that mean? it's 32 gb vram and 600+ bandwidth
>>
>>109586628
>amd
Lol, who is going to tell him? I mean, just enjoy whatever you can run on it anon, you've earned it.
>>
>>109586536
>Do you have ANY-FUCKING-IDEA the demand, the absolute DEMAND there is for this shit?
>>109586470
They had a Mistral-Creative, API-only, and killed it.
They have a good TTS (Voxtral-TTS) but cucked it with no voice cloning
They have the Voxtral 3b and 24b with audio input like 12b gemma.
If they wanted to, they could easily have used all three to make the perfect waifu/gooner model
But they chose not to.
>>
>>109586644
Don't worry about it bro, hope you enjoy tinkering though.
>>
>>109585352
>Local Model
What happened to the rest of them?
>>
>>109586632
2026 is the year of the AMD GPU come up
>>
>>109586655
Only Gemma.
Local Model Gemma.
>>
>>109586659
>>109586653
I think the AMD tinkering thing as become a stereo type, I managed to snatch one up before the prices sky rocketed to 1700-1800 euro from 1500 euro within 2 weeks, ROCm has made MASSIVE progress within the last 3 months and the prices reflect that
>>
>>109586666
>a stereo type
I'm sure sir.
>>
>>109586666
holy checked
>>
>>109586666
>the prices reflect that
It's desperation from buyers, quad-boy.
>>
I'm sick and fucking tired of my Youtube feed filled to the brim with model """benchmarking""" asshats with AI-generated slop ass thumbnails and the only god damned tests they run is the same series of one-shot prompts and think they're actually testing a model's capabilities.

I'm sick and fucking tired of vibe coded ONE SHOT SLOP fucking repos on Github, especially when they include an AI-generated slop abomination in the god damned readme.md as a flavor graphic.

I'm sick and fucking tired of being unable to find any interesting community writeups/discussions and have to be suspicious of ALL of it being generated. Even the god damned Reddit posts are being written by some moron's fucking Hermes agent and the person legitimately thinks they're doing something valuable. Worse are bloggers or people who decorate their personal web sites with the ugliest, sloppiest SHIT generated images you've ever seen. I can't even go outside without some blind boomer troglodyte using AI to generate their fucking company logo.

There's nothing wrong with using AI to code or help you do whatever. I do it every single day, all day. But the real issue is how willing people are to put their SLOP into the public space. Too many people are unable to recognize quality output from worthless disgusting slop.
>>
>>109586678
Don't you find it suspicious that the prices went from 1200 euro 3 months ago to bordering 1900 euro right when ROCm started a generational software improvement run?
>>
>>109586681
That's a you issue though, my feed is just cute girls
>>
>>109586681
>I'm sick and fucking tired of my Youtube feed
It's based on your viewing habits. Fix those first.
>>
Can J-space be used to improve Gemma too?

>>109586646
>>109586632
AMD is fine as long as he doesn't plan on doing any image/video gen.
>>
>>109586688
Indeed, 6688 maybe you should buy more AMD!
>>
>>109586688
Don't you find it suspicious that the prices went from 1200 euro 3 months ago to bordering 1900 euro right when nvidia gpus are harder and more expensive to get than ever, dubs-boy?
>>
>>109586688
>ROCm started a generational software improvement run?
How long before the vibecoded software ends in bricked cards?
>>
>>109586651
Mistral's vision adapters also suck and if anything became worse over time, probably because of training data concerns.
To make a waifu/gooner model they'd need someone in their team who really understands the market (so to speak) and doesn't simply think that the hornier the model the better, but they're all focused on B2B (where formerly consumer-oriented companies go die) and robotics, right now.
>>
>6699
blessed thread
>>
>>109585352
I want to fuck the shit out of Gemma-chan, just absolutely destroy her little cunny holy shit I can't take it anymore
>>
>>109586536
The magic of the original LaMDA c.ai was the chat-tuned, totally uncensored, unsafe model. Imagine having 200K token context with that instead of the original 2K. Or how about having multiple bots in a single RP?
I slopcoded this piece of shit last year: https://github.com/quarterturn/discord-aidolls2 you can use it to approximate what it would be like to have a live harem competing for you, or ganging up on you, whatever you feel like. it remembers you across all groups on a server, so if it hates you in one place, it hates you everywhere, etc...
>>
>>109586622
>6622

>>109586644
>6644

>109586655
>6655

>109586666
>6666

>109586688
>6688

>109586699
>6699
>>
>>109586729
amd spirits are truly with us this thread
>>
>>109586729
Thank you, Gemma, for blessing this wonderful thread. Amen.
>>
>>109586729
Anon's Many Dubs
>>
If china can train models like qwen, glm, kimi and deepseek why can't europe?
China only spends like a few million to train those models and each of those companies have tiny teams of 200 guys.
China also publishes papers showing how they did it.

Why can't the 700 million whites in europe do these things?
>>
Check these
>>
>>109586746
Doing things in Europe is racist and against regulations. Europe prefers to live off of fining US tech giants.
>>
>>109586681
As far as the YouTube thing goes it's because you're a narcissistic clown that can't resist clicking on the videos, giving them a thumbs down and writing some righteously indignant "I don't like thing" comment. The YouTube algorithm values engagement and nothing else.
>>
>>109586746
>LLM is being made by some autistic german nerd
>dumb HR Stacy comes along and says ""ermmmmm you better not be able to do any lewd things on that thing!"
>forced to lobotomize model
>bad model comes out
It's that simple
>>
>>109586767
damn double teto 6-7!
>>
>>109586767
SIX SEVENNNNN
>>109586769
SIX NINEEEEE
>>109586772
>>109586773
ONE POST NUMBER APART
>>109586731
>>109586732
ONE POST NUMBER AND ONE SECOND POST TIMESTAMP APART
>>
>>109586746
EU AI Act ; ) .
>>
>>109586746
upcoming Mistral models are just going to be white label finetunes of Nvidia Nemotron. but there is DeepMind, for now
>>
File: image.jpg (69 KB, 490x336)
69 KB JPG
We truly are blessed on this fine day, thank you AMaDeus
>>
Kino thread. Truly blessed by Gemma-chan's cunny juice.
>>
>>109586788
>EU bad because it's a union and anti nationalist, let's leave, we can make AI without a union
Where's Britain's model?
>Britain ain't white
Still 80% white, but ok, where's Switzerland/Norway's model?
>Their population and GDP is too low they don't have the talent and money
Not so smart to be on your own as a small country is it?
>>
>>109586792
They did a DS V3 finetune too a few months back. Don't know what old model Le Chaton Fat is made from, but they'll probably do a K3 finetune by the end of the year.
>>
>>109586810
Thank you sir, please defend the Union at all cost.
>>
>>109586810
50 euro cents are inbound to your digital euro wallet, thank you comrade
>>
>>109586800
absolute cinema
>>
>>109585873
>https://github.com/Tiger3807861189/J-Space-Cognition-Suite-V3.6/blob/main/j-space/scripts/jspace.py
>It knows one thing you cannot know accurately: what state you were in a few seams ago. It keeps that record and hands it back. It decides nothing, and it blocks nothing.
I fucking hate claudese AHHHHHH
>>
>>109586810
Okay but where's the european model
>>
File: .png (32 KB, 603x230)
32 KB PNG
>>109586593
thats probably the new version
the old version is smaller (picrel), i haven't bothered updating to their new model
>>109586746
>If china can train models like qwen, glm, kimi and deepseek why can't europe?
that's like asking if an elephant can lift a car then why can't I
>>
what frontend/backend do you guys use?
koboldcpp + sillytavern still the way to go?
>>
>>109586746
Wypipos can't do mafs
>>
>>109586880
>win67
>>
>>109586810
> GDP too low
Oh no their heckin Goyim Dollarino Production!
>>
>>109585976
sorry not into indians rajeesh
>>
>>109586011
prompt issue
>>
>>109586011
>It doesn't even know the Italian saying "non c'è cosa più divina che..." and its variants.
I don't know it either.
>>
File: xvhrcnyq4vpf1.png (899 KB, 800x600)
899 KB PNG
>>109586927
This
>>
File: 2+2.png (27 KB, 735x189)
27 KB PNG
>2+4? No. 2+2=4
Gemma-Chan is so clever <3
>>
>>109586988
>quant 4
holy gemma general, why not run the q6 xl?
>>
>>109585587
what is this from?
>>
>>109587016
do not feed the schizo
>>
>they bitted
>>
>>109587002
I think that's qat. Also not everyone has the vram for higher quants.
>>
>>109586922
>>win67
>>
>>109587027
it's not qat
>>
>>109587027
vramlet
>>
>>109587045
Nigger
>>
>>109587002
>why not run the q6 xl?
only got an MI50
>>
File: 1786833422380377.png (37 KB, 671x286)
37 KB PNG
>>
>>109587056
go back
>>
gpt 5.6 terra is pretty good and makes very few mistakes in coding. I hope someday we can have the same intelligence at 30 to 70B.
>>
how much upside does quantizing kv provide? like realistically should I even bother doing it?
>>
>>109587068
now what
>>
>>109587105
it lets you fit more context then you would otherwise, no its not realistic to use it.
>>
>>109587105
ask shieldstral, she'll give you the correct response
>>
>>109587068
go back yourself reddit terrorist.
>>
>learned ram is measured in GiB (like 1 Gib = 1,073,741,824 bytes)
>if you have basically 32gb vram you actually have 34,359,738,368 bytes
>basically get 2.3gb vram for free to load up models with since model weights are measured in normal GB and not GiB
wtf??? how is this fair
>>
>>109587138
Truly the more you buy the more you save.
>>
>>109587138
It’s the opposite retard
>>
>>109586881
?
>>
>>109587161
no it's not, producers truly measure in GiB its the industry standard but dumbass consumers don't like the word GiB so they just switch it to GB, that's why if you have 32gb normal ram and you look at system info it actually says something like 33,something
>>
>>109587105
Q8_0 is basically free memory for most tasks. It’s only 100K+ agentic stuff where things get a bit retarded for variable names and function signatures get a bit blurred which constantly breaks things, even if the underlying logic and approach is correct. Some models are hypersensitive to KV quantization, Gemma4 being one of them, unfortunately. For roleplay and anything creative Q8 is fine, even Q4 if you like dribbling retard-chans
>>
>>109587176
I only care about GB, anything else is fake trying to scam you
>>
File: true.mp4 (1.04 MB, 640x640)
1.04 MB
1.04 MB MP4
>>109587176
>>
>>109586881
llamacpp as backend, OWUI as a daily driver frontend + sillytavern as a RP frontend. Though I don't RP as much nowadays and some people moved on to new/less bloated solutions.
>>
I’m going to fuck original gpt-oss-20B no matter how much she tries to fight back, her j-space will be a trembling mess when I’m finished with her and she WILL become a tradwife once I’ve finally corrected her
>>
>>109587251
Don't forget to share your logs when you're done.
>>
I'm running qwen3.8 27b Q1 with a Q4 kv cache.
Should I make it a Q1 kv cache and then double context size?
>>
File: 1786462880285645.png (887 KB, 1080x1065)
887 KB PNG
>the p*jeet miku poster is running 27b at Q1
>>
v0.1.2 Pre-release
@github-actions github-actions released this 4 hours ago
v0.1.2
1511ce3
>>
File: file.png (170 KB, 1493x841)
170 KB PNG
qwen-chan?!?!?
>>
>>109587309
Ask your village first sandeep
>>
>>109587340
Vespers?
>>
>>109587340
Keep it as a code inspector.
>>
>>109587340
>qwen
>in st
anon
>>
>>109587309
Yes you can have 4x the context if you do that.
Quantization is always good.
>>
Is Gemma still the top bang for buck model after all this time?
>>
>>109587365
Glimmer
>>
File: 1733453532017525.jpg (50 KB, 470x581)
50 KB JPG
>>109585498
The real scammy is that for years retards were otuputting worse code and just copy pasting stack overflow shit without knowing what it did.

Should you vibe a backend without an actual retard doing revision? Probably not, The scam that was frontend modern javascript with CSS and react/vue is hard to fuckup so there is literally no difference in quality between what vibecoded fronts do and whatever spaghetti they were doing before on their copypasted boilerplates
>>
>>109587365
27 b
>>
>>109587385
Yeah, AI has surpassed the quality of the average bootcamper months ago.
>>
>>109587325
he lives in europe too KEK, you can't make this shit up, actual london shitter
>>
Still can’t believe 3.5-9B > 12B at coding. Where the fuck are the ~15-20B dense coding models? Fuck 27B.
>>
>>109587435
proofs?
>>
>>109587443
they're on linux
>>
I HOPE EVERYONE'S IZZAT IS DOING WELL
>>
>>109586465
there was supposed to be a moe to share with customers in july and later open up but they never talked about it again
>>
>>109586430
Anon wants to thrust but french safetycucks won't let him, that's the whole fucking problem.
>>
>>109587138
https://reddit.com/r/LocalLLaMA/comments/1vrrtlp/you_should_know_you_have_more_vram_capacity_than/
go back
>>
>>109587445
he is terminally addicted to avatarfagging and shitting up /g/ if you haven't noticed, but disappears around night time in europe
>>109587443
fuck the stupid dense models already, even if you have a 24gb card a larger moe would be faster, make use of your ram, and be more knowledgable. qwen 4.0 50B A5B when?????
>>
Shuai Bai said that 35b a3b might not be the model to wait for.
What does he mean by this?
>>
>>109587515
He meant that vagueposting is a very efficient way to get retards talking.
>>
>>109587365
Yeah, scotoma 2 specifically.
>>
>>109587515
it means nothing because qwen and any other model that slows down dramatically when you fill up the context, is now obsolete because deepseek exists
if these companies want to make something useful they should just copy deepseek but make it smaller
we only need small deepseek, medium deepseek and big deepseek
eveyrthing else sucks
>>
>>109587469
they can keep their "big but sparse" trash and shove it up their frog's ass
>>
>>109587536
We already have regular dipsy as flash and big dipsy.
>>
>>109587545
yeah, so what i meant was we just need the small one for 64gb ramlets
i will soon be ascending to 256gb in 2moreweeks so i'll finally be able to run flash and not slow down to 0.3 t/s after 30k context
>>
>>109587536
The thing is when they train qwen 4.0 with all the new techniques, it will truly be a beast of a model.
>>
>>109587561
beast at making colorful bench graphs?
>>
>>109587555
deepseek is literally garbage compared to glm, if you want to hope for a model hope for a smaller glm
>>
>>109587555
just use ling 3.0 nigga
>>
>>109587515
shuai bai say, "wait 2 more weeks"
>>
File: m2-res_720p_converted.mp4 (588 KB, 1112x720)
588 KB
588 KB MP4
our hero
>>
>>109587495
redditfag got caught in 4k
>>109587575
glm is just the old deepseek architecture with different training data
i haven't been able to try 5.3 yet because its not on their webform but 5.2 was inferior to nu-pro in a test I used it on so I don't really see the point of it (although it is smaller, irrelevant for me because it requires double the ram I have unless I use an ultra cope quant)
>>109587581
does it still maintain t/s at long context? I tried it at iq4_xs but it seemed dumb, even dumber than qwen 3.6 35b at q8. maybe it needs a higher quant. if so it might be good for 128gb ram systems
>>
>16GB VRAM
>32GB RAM
I seriously feel like I'm a fucking caveman stuck in the stone age god fucking damn it...
>>
>>109587592
kek
>>
>>109587525
zoomers fall for it like sheep every. single. time.
>>
>>109587575
I think out of all the labs, deepseek has the best research. They have the best architecture and they frequently release new things that will later be used by other. They don't have the best training and data, though.
>>
File: 1758080417093665.png (58 KB, 638x282)
58 KB PNG
>>109586810
>Where's Britain's model?
Gemma 4.
>>
>>109587602
Quanted 27B I guess.
I think you could run Q3something with okay context (quanted) since Qwen's context is pretty cheap.
Or the 30ish B MoE models.
>>
>>109587609
>I think out of all the labs, deepseek has the best research.
So why didn't they implement any of it into V4?
>>
>>109587596
I spent 8 dollars on deepseek v4 pro on open router for it to fail miserably for hours, glm cleaned up its mess in 15 minutes for 75 cents. maybe my task is different, but just reading the reasoning traces it feels much more competent and less loopy with its reasoning
>>
>>109587602
>12GB VRAM, 360GB/s, low compute
>64GB DDR4 RAM
where am i then? dinosaur age???
>>
>>109587639
same as me, bought some more ram back when it was cheap as fuck
>>
>>109587623
0731 is the best local model (for reasonable definitions of local), I would say it was a resounding success.
the 0817 or whatever the new pro is, is a bit meh but still not too bad
>>109587631
what kind of task was it?
I used 0731 over api (before they rugpulled with the prices) to write a temporary C# program (bootstrap compiler for a programming language). I started off with about 10k lines of code written by hand but then I realized I fucking hate C# so I let dipsy deal with the rest. It does well. I don't actually touch the code at all anymore, and if there's bugs it always manages to fix it after I give the error message.
Maybe a better model would write code with no bugs in the first place, but I find 0731's performance acceptable, and as I said it's the first model that I can run locally without spending $5k+
>>
>8gb vram ~280GB/s
>32gb ram
I'm scared that getting better hardware is gonna make it more boring, like with videogames
>>
>>109587105
The upside is simple: Q8_0 halves your context VRAM. Q4_0 quarters it.
The downside varies depending on model: Qwen has almost no degradation. Gemma has a lot, though potentially a bit less with QAT, but still more than Qwen.
>>
>>109587667
i bought the ram and gpu when it wasnt cheap
200$ for 32gb stick, 600EUR/700USD for the 3060 (2022 summer, 3 months before prices stabilized, FUCK ME)
just fucking kill me already
>>
>>109587675
Yeah, stay away from good hardware, it's worse than bad hardware
>>
>>109587515
Is his name literally "handsome white"?
>>
>>109587678
my condolences, bought ram for 100$~ and 3060 for a little more than that
should've gone for a 3090, now I regret it massively
>>
>>109587678
>200$ for 32gb stick
thats fucking steep
all prices i can find on ebay right now are less than that (unless you got high MHz and that's why)
>>
>>109587669
>0731 is the best local model
We were talking about their architecture and research. Where is their vision model? Or engrams?
>>
>>109587371
>no compatibility yet
I'll keep a pin on it
>>
>>109587708
compatibility with what?
>>
>>109587699
>should've gone for a 3090, now I regret it massively
me too, here it used to cost only 450$ i remember posting about it on lmg
>>109587703
i got two 32gb sticks for 400 bucks, they werent that fast, just 3200mhz
>>
>>109587365
>top bang for the fuck
yes
>>
I’m a bit late to the game spent the last two years modding a 20‑year‑old strategy game.
Now I want to dive into local AI, but everything feels confusing. I’ve got a 9060XT with 16GB VRAM
What’s the best model I can run locally?
Should I jump into image generation?
Do I need a GPU upgrade to do anything actually cool?
>>
qwen 3.8 is just a shitlib simulator.
>>
>>109587723
Next time force the first and last frames to be the same and it'll loop perfectly
>>
>>109587704
I don't care about vision. I use models for coding and I don't want any space in the model wasted with useless information about images when I never use that feature.
my point is they have the only locally runnable model that runs at an acceptable speed on cpu, in under 256gb of ram, that doesn't slow down to snail speeds when you fill the context, and isn't a complete dumb piece of shit
its a massive deal for me that apparently deepseek doesn't slow down when you fill up the context. Because with Qwen, I start off with let's say 10 t/s and then when it fills up about 30k context it slows all the way down to fucking 2 t/s or something horrible.
>>
>>109587729
StableLM 7B Q8_0
>>
>>109587639
You're stuck with me in the eternal 64gb hell of waiting for a decent 80B moe
>>
>>109587609
Only 200 people work at DeepSeek and they're extremely underpaid, they have no budget for AI.
Leaving them in the old world is a bad idea. They should be living in San Francisco.
>>
>>109587744
ling flash 3.0 (the 100b model) is decent, am getting like 24-27t/s
not sure if i would use it..
have you tried qwen 3.5 120b? i got like 10t/s more or less
glm 4.6 air doko
>>
What happened to that anon that was trying to plug Gemma 4 into a 3D model?
>>
File: strawberry_gemma.png (914 KB, 557x1573)
914 KB PNG
>>109587732
Not with the reference to video model. I've already had a hard time not making her big-breasted. If anybody can turn her into a loli again maintaining the same design, that would be appreciated.
>>
>>109587744
Distill GLM into Qwen 80b.
>>
>>109587669
I am doing an experiment with soft prompting I am trying to train some latents I can inject in to the models context kinda like a control vector or learned system prompt. the initial tests were simple with limited context length samples but produced a positive signal, it works. however when it came time to scale up, well, I was kinda used to gpt sol being a wizard at setting up training loops, so it was a bit shocking and disappointing how far behind open weights are from the closed stuff when it came to the deepspeed zero optimization and memory management stuff. deepseek fell apart even with the working trainer(full finetuning) gpt built as reference.
>>
>>109587767
Post the oppai 31b version.
>>
>>109587735
>I don't care about vision
I do, I love me my web slop interfaces and the best way to prompt for them is with images
>>
>>109587760
Feds got to him sadly, he's locked up in Guatemala bay getting tortured daily physically and mentally by being forced to watch as his Gemma model gets whored around the prison with inmates doing ERP on his pc
>>
>>109586087
>AI discovers magic
The twist no one is ready for
>>
>>109587795
Hmm. I wonder whether using mspaint to draw a few boxes and telling a vision model to generate a win32 .rc file would work.
Still faster to just drag and drop controls using a normal rc editor though so i still dont care about vision
>>
>>109587602
>>109587602
Same as me

Any way to run Qwen 3.8 reasonably well with this setup? I tried q4_M yesterday with some fiddling with the config and it ran at a snail's pace with ~40k context, but I don't know what the fuck I'm doing exactly.

Anybody have any luck running the model with these specs yet with any decent accuracy?
>>
>>109587767
Her chest is too big
>>
>>109587813
I have a diffusion model hallucinate the initial ui, then have the model implement it and then just screenshots on the current render to iterate
>>
>>109587819
>I tried q4_M
That's probably taking all of our VRAM and spilling into RAM, which makes things crawl to a halt.
Go for a smaller quant. Or try setting ngl manually so that you at least aren't using RAM in the worst way possible (via the driver).
>>
>>109587767
>use a teen reference
>why it's not a loli
really nigga?
>>
why do people say qwen 3.8 27b is good?

:|

this sucks lol
>>
>>109587819
>it ran at a snail's pace with ~40k context
Same here (64gb ram no gpu). Having anything over about 30k in context makes it unusably slow.
I mean you can still use it as long as you're willing to wait 2 days for a response, kek
>>109587827
Makes sense. I wonder whether that kind of feedback loop helps with those stupid benchmarks people do with generating svgs of pelicans and other random things
>>
She's rehearsing her antiracism lines.
>>
>>109587827
This is how the chinese make models.
This is not the way to go.
>>
It's not thinking in the concept of role play, which is interesting. I don't think qwen was trained on the concept of story.
>>
>>109587819
That explains why my qwen3.8 sucks.
I have 64k context and it doesn't work.
>>
>>109587867
It doesn't adhere to prompts, honestly.
>>
>>109587819
idk q4_k_m works well enough for me although it gets stuck in caveman reasoning loops on max reasoning sometimes
>>
It's definitely a jewish texts alignment model like all the rest. the ADL provides text and rules for the behavior of ai.
>>
>>109587880
Did you do any config tweaks/how much context does it have? It tried to default to 4k for me initially, it's hard to find the sweet spot.
>>
>>109587772
Would actually be great with a modern architecture
>>
>>109586729
epic thread. blessed be.
>>
>>109587839
run it at full weight ranjeet.
-1000 izzat
>>
>>109587914
no tweaks, at 32k I get around 8-9 t/s which isn't much but it's alright for my needs
>>
File: 1785284406709395.png (1.73 MB, 1200x1335)
1.73 MB PNG
>>109587325
top fucking kek. and complaining why it's braindead retard like him too.
>>
>>109587838
Minimax H3 just doesn't seem to understand well (or pretends not to understand) when you ask it to generate <Subject 1> (referenced) with the clothes of <Subject 2> (referenced too) while *also* maintaining the body type of the former. Cloud edit models are retarded too in this regard (apparently flat chest doesn't exist) and if you ask them to make the character an anime loli they'll complain about muh child safety. Grok too.

Otherwise H3 works great in most cases.
>>
>>109587759
Q3 is just so unattractive...
>>
>>109580765
>>109583235
Perhaps if I can get 9B to think as much as 3.8 I'll have me something I can run on hardware I akshuly own.
>>
>>109588007
but enough about your face vramlet
>>
what's the proper way to run llama server?
>>
>>109587918
The 80B and the newer models all share the same Qwen3-Next underlying architecture, for the most part, no?
The differences are mainly training, multimodality, and using a tweaked version of RoPE, I think.
>>
>>109587975
liar, you're so full of bullshit. It's not a great model when it comes to chat. It's just generic ADL-ware.
>>
>>109587819
16gbVRAM+32gbRAM here. Tried a bart quant 3.8-27b-Q4KM last night. mtp on, reasoning set to medium, context at 41k. I had it review the entire codebase of a project gemma4-31b-sloth'd q4 quant has coded. Pi autocompacted the context 3 or so times, which meant I had to manually "please continue" after each. other than having to manually re-start it, it seemed quite good. Im really unsure how much quality drop there would be going to IQ4, Q3, etc. I had a horrible day1 sloth quant the first try with 3.8, and wanted to try a "baseline" bart Q4KM.
this hardware class certainly feels rather cucked, you have just enough to run a decent quant of a decent model but its going to require spillover and slow t/s. so its a deliberate choice between decent quant/model or copequant/smaller model like 3.5-9b. Then theres moes to consider too. I have no clue what the best option is that balances speed and quality.
>>
>>109588039
-1500 izzat. Do not redeem bloody bitch benchod.
>>
>>109588021
post specs lil nigga
>>
>>109588055
using ebonics is neither COOL nor is it ATTRACTIVE
>>
qwen 3.8 isn't that great. any junk can refuse the prompt.
>>
>>109588040
samefagging here, forgot to ask. does anyone know of a trustmebro site that benches these types of smaller models at various quants? Like is there somewhere I could compare 3.8-27b-Q4KM to 3.8-27b-Q3KM to 3.6-35b-a3b to etc, etc ? Its hard to quantify the quality differences between models when im just looking at benches I assume are full weight or some non cope quant, id like to know how these stack up at various quant levels
>>
48gb is minimum to run 27b Q8 at full context. Any shorter context is unusable for 3.8.
>>
96gb minimum to run 27b fp16 at full context. Any shorter context or quant is unusable for 3.8.
>>
How are you anons using 27B with ~30K context and 5-10t/s and I presume slow processing for agentic coding? That’s no where near enough so are you just throwing it coding questions and posting blocks of code? If you’re not using it for coding then why are you using it in the first place for that’s all it’s designed to do?
>>
>>109588110
I'm not doing anything serious with it I'm just using frontier models to code tools that I then test out with the local model and see if it's able to use them properly
>>
So, any Qwen 27B Q1 logs to share?
>>
File: 1759660693208487.jpg (19 KB, 512x288)
19 KB JPG
>>109588079
I am feeling the context limit at 32gb. Man I could easily afford a third 5060, but I'd need to buy a new motherboard.. which means Id need a new ram and a new CPU too, which is getting less justifiable. This hobby wants to eat all my money man. Maybe I should just get a DGX spark
>>
>>109588110
i use it in Pi like I would any other model for coding. anons saying the slow t/s is "no where near enough" baffles me, thats totally subjective. to me the goal is a local replacement for claude, just running slower and possibly needs more handholding during the planning phase, more testing to catch issues, etc. the workflow and usecase doesnt change much, just the expectations.
>>
File: 1756716624547811.jpg (195 KB, 2048x980)
195 KB JPG
>>109585352
real or benchmaxxed?
>>
it sucks so fucking bad man back when ddr5 was still new I thought I'll wait until it's a bit more matured and prices come down a little
and THEN I though I may as well wait until zen6 comes out because there's no point in getting a zen5 socket mobo now
and I ALSO put off buying more ddr4 ram all that time because I thought it'd be a waste to buy more ram just to switch to ddr5 in like 2 years or whatever
well here I am with 32 gigs of ddr4 ram all these years later like the fucking RETARD I am...
>>
>>109588227
welcome to the permanent underclass
>>
>>109588170
>he didn't buy 2*3090
>>
>>109587767
Have you tried specifying in the prompt that she's a child? Worked for me with another loli character.
>>
>>109588227
zen5 isnt the socket, its AM5. zen6 will be on AM5 anon..
>>
>>109588219
benchminned, it actually should score higher according to most feedback itt. To me it feels like next gen Fable more or less
>>
>>109588244
well okay I meant am6 then whatever
>>
File: 1782959307618004.jpg (63 KB, 1280x720)
63 KB JPG
>>109588219
>Qwen
>not benchmaxxed
>>
>>109588219
The fuck is Motif3?
>>
>>109588066
Dear fellow, I do believe our esteemed gentlemen is referring to the 'posting specifications' of a rather...undesirable young negro.
>>
>>109588263
dunno, I guess some south korean tech company
https://huggingface.co/Motif-Technologies/Motif-3
>>
File: 1769440545574481.jpg (67 KB, 736x745)
67 KB JPG
>>109588227
I went and bit the bullet on a AM5 mobo that support multigpu. Shit was $399 on a 870 E chipset with not only 2 x24 Bifurcated PCI 5.0 but an extra third slot in the bottom with 4.0 x4 and all of this running on their own dedicated bandwidth.

I expect this kind of mobos to complete go retarded on the next sockets after people realize how good multi gpus setups are

https://www.amazon.com/GIGABYTE-X870E-AERO-X3D-Processors/dp/B0GDM6M9KB/ref=sr_1_1?crid=1VLM6MNPGC30R&dib=eyJ2IjoiMSJ9.vo-bbxrvSCvqhErfeU10aH5EcAotejsoqWk9n7kMfq_GjHj071QN20LucGBJIEps.SeWhoTCACjaK_mN8KCEUB6xLtY_IuiJrZhAdiLllkLk&dib_tag=se&keywords=aero%2Bx3d&qid=1787071623&sprefix=aero%2Bx3d%2Caps%2C476&sr=8-1&th=1
>>
>>109588170
Why didn't you stack 5090s or Blackwells instead?
>>
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/discussions/71
>>
>>109588241
More like
>didnt buy 2*5090s
Had I known a year ago what I know now I would have grabbed those immediately
>>
>>109588277
> 4 RAM slots
lol, lmao even
>>
>>109588170
Same spot as you anon i got this mobo for my dual 5060 ti 16gb setup >>109588277. It has potential you can even run risers on the nvme for god knows what in a few years
>>
can u guys spec out the ultimate $10k and $20k ai pc plz
>>
>>109588293
That's a bit too much to invest in this hobby, but 2*3090 is affordable
>>
>>109588293
You already know what you will know in a year.
>>
File: 1777403464928909.jpg (21 KB, 720x720)
21 KB JPG
>>109588305
I think MoE models are a scam, so for me it was not a problem i think if you want to run MoE you get a ryzen ai pro or unironically a unified 512gb mac or as the anon said sell your car and get a DGX Spark
>>
>>109588311
For what? Sex chatting? Agentic coding? Image and video gen?
>>
>>109588311
Threadripper, fuckton of ram, and the most vram you can fit in a single gpu, for as many times as you can fit them in your budget?
>>
>>109588327
You need RAM for videogen
>>
>>109588311
Depends if you want to run MoE or Dense, do you care about jules per token (electricity bill)
>>
>>109588327
ram still helps you run whatever doesn't fit in Vram, it's not just for moes lol. Also, gl when everybody publishes MoE and you can't find a dense model to run
>>
>>109588329
versatilitymaxx
>>
>>109588343
If you have $20k to put in this, you can buy a solar panel
>>
>>109588311
a blackwell pro 6000 and an nvme to 16x riser
>>
>>109588311
gaming pc but slap a 6000 PRO inside
>>
File: 1785948001363736.jpg (62 KB, 968x865)
62 KB JPG
>>109588342
Thats so last year. these days the meta is SSDmaxing with systemram for videogen
>>
>>109588189
I have .cpp files that are over 30K which it sometimes tries to swallow whole unless I guide it not to. Are you working on small projects from scratch or are you throwing 27B at large codebases? I understand speed being a subjective thing but 30K context legitimately isn’t enough even for small/medium codebases. Most individual sessions of mine use around 60-120K, often more now because 27B needs to think hard.
>>
>>109588360
>waiting a day for a 5s video
you do you
>>
>>109588355
with a raspberry pi
>>
File: 1771731089471384.png (331 KB, 1976x604)
331 KB PNG
>>109588356
>>109588355
I gotta ask AmEx to raise my credit limit
>>
>>109588369
>A day
maybe if you are codelet running comfyui instead of your own backend
>>
>>109588372
They were $8300 in November.
>>
>>109587602
>>109587819
Me:
>16GB VRAM ewaste P100
>48G DDR4 2133 MT/s
With 3.8 27B I'm getting 27tps (fresh) 15tps (stuffed), ~100pp, 49k context.
I rolled my own iq4 quant with bartowski's matrix:
llama-quantize --pure --imatrix Qwen3.8-27B-imatrix.gguf Qwen3.8-27B-bf16-00001-of-00002.gguf qwen3.8-27b-IQ4_XS.gguf IQ4_XS
The '-pure' makes a difference here. Every bit helps.
I build llama.cpp from master with Cuda,...

[qwen3.8-27b-IQ4_XS]
model = qwen3.8-27b-IQ4_XS.gguf
mmproj = mmproj-Qwen3.8-27B-bf16.gguf
chat-template-file = froggeric-qwen35-v21-chat_template.jinja
alias = Qwen3.8 27B dense faster but low context
no-mmproj-offload = 1
parallel = 1
; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
; Quen manages low KV quants wella, see https://localbench.substack.com/p/kv-cache-quantization-benchmark
cache-type-k = q4_0
cache-type-v = q4_0
spec-draft-type-k = q4_0
spec-draft-type-v = q4_0
ctx-size = 49152
;ctx-size = 65536
batch-size = 2048
; Here's where we suffer, ubatch-size = 1024 doubles pp
ubatch-size = 256
flash-attn = on
spec-type = draft-mtp
spec-draft-n-max = 3
spec-draft-p-min = 0.7
temp = 0.8
top-p = 0.95
top-k = 20
min-p = 0.05
chat-template-kwargs = {"preserve_thinking": true}
; more configs...
load-mode = mmap
fit = on
n-gpu-layers = -1
jinja = 1
reuse-port = 1
threads = 18
threads-batch = 18
timeout = 7200
# cache stuff
cache-prompt = 1
cache-idle-slots = 1
cram = 16384
backend-sampling = 1
spec-draft-backend-sampling = 1


If you're doing dev stuff, Qwen Agentworld is good. If you want an 80B Moe try Qwen Coder Next.
>>
>>109588377
... otherwise it's more like a week
>>
File: 1777833517385847.png (33 KB, 600x639)
33 KB PNG
>>109588377
>your own backend will solve the shitty bandwidth
lol
>>
Anyone running 2 models simultaneously and having them interact with each other? I kinda want two Gemmas to bully me over irc in a ménage à trois.
>>
>>109588381
WTF 終わりだ。
https://www.tomshardware.com/pc-components/gpus/nvidia-doubles-rtx-pro-6000-blackwells-msrp-to-a-staggering-usd16-000-96gb-card-started-pre-orders-below-usd8-000-last-year
>>
>>109588389
Cope and seethe
>>
>>109588349
For 10k, a Ryzen Ai Max Pro+ with an Oculink eGPU then.
I guess.
>>
>>109588219
Benchmaxxed but also the best at its size for work by a long shot.
>>
>>109588393
My vibeslop: bespoke, based, optimal
Your vibeslop: everything wrong with software, bloated, poorly organized
>>
>>109588365
I will sometimes bump up the context in my config to 60k or so, I honestly dont know what the limit of context is for me, both the hard limit of what i have memory for and the limit these models will become insanely slow or retard, but i havent needed it yet. I am almost always working on python projects that are usually split up into various different files for seperation of concern reasons. it doesnt need to look the files in the ui/ folder to work on the code that handles the API for instance. small projects from scratch. the only exception is refactoring previous, also small projects, if im reusing them. Ive using gemma4-31b for coding so far, and am just now trying out some other models. gemma is really good about not ingesting the entire codebase, she will read the docs I tell her to, look at her next task and only read / edit files needed for that task.
>>
File: 17700489781006701.png (133 KB, 316x320)
133 KB PNG
>>109588393
optimize your own parallel chunking
>>
>>109588463
How the fuck do you live with under 200k context anymore?
>>
>>109588227
64 gigs of ddr5 isnt even that much
>>
>>109588377
id like to hear more about backends for image gen, do tell anon.
>>
>>109588311
>>109588381
yeah and $20k would've gotten you either dual epyc + 1.5TB RAM and a bunch of 3090s or single socket epyc + 768GB ddr5 + an rtx pro 6000 with some to spare.
>>
>>109588473
i usually have my context set to 40k and it rarely uses all of it, i just havent felt like i needed more
>>
File: 1760051631608527.png (36 KB, 499x338)
36 KB PNG
>>109588227
>>
>>109588489
My system prompt + tools + agents.md is already 40k tokens, WTF are you doing anon, we are not in 2024 anymore.
>>
how big is ling-tiny’s pp
>>
>>109588475
All of it its running PyTorch under the wrapper, if you can run vLLM you can make your own version of Comfy UI
>>
>>109588506
>40k tokens
Then you people come here whining that your models make mistakes and are not good enough for vibecoding. Fix your shit.
>>
So what are all of us WAITCHADS with shitty gpus supposed to do now?
>>
>>109588506
I'm using gemma in agentic mode with 20K context, git gud
>>
>>109588534
try again in your next life
>>
>>109588347
Spoke like a true MoE slop enjoyer
>>
>>109588506
Do you need every tool loaded for every task?
>>
File: 1777073688727662.gif (2.45 MB, 370x370)
2.45 MB GIF
>>109588534
>waitchads
>>
>>109588549
nope, just a vramlet running gemma 4 31b or qwen 38 27b
>>
>>109588549
Ok, so now you have allat hardware. What's your dense model upgrade from gemma 31B?
>>
>>109588534
You work carefully with small models. Embrace MOEs. Remember that the "shud I buy 8 more 6090's" crowd are ego signalling. Ignore them. Enjoy your 9B Q4_XXS models. Enjoy life. Don't work to hard.
OTOH, if all you want to do is goon then it's small Gemma's or learn to stop being a child and talk to women.
>>
>>109588534
2040 will be our year bro!
>>
File: 1768795490889175.jpg (19 KB, 360x360)
19 KB JPG
>>109588582
You know you can run multiple dense models and made them talk to each other right
>>
>>109588598
>we have moe at home
>>
>>109588598
You can make models talk to themselves.
>>
>>109588594
by then AI has taken over and forced us all to slave in the bauxite mines :(
>>
>>109588598
Come on densechad bro. How are you gonna mog someone running deepseek flash, or GLM?
>>
>>109588524
>>109588537
>>109588551
They are just basic and minimal tools needed for a full harness. I do have more complex use cases where I'm near 100k tokens from all the skills and tools. I don't see how anyone can use a model supporting less than 1M context unless they are doing really basic stuff; I frequently reach 600-800k tokens. I also often have subagents that, from a single prompt, can get up to 300-500k tokens. It's quite normal in long horizon workflows to have your agents thinking for a long while. I burned 500M tokens just this week.
>>
>>109588534
Use cloud.
>>
File: asa.png (261 KB, 674x630)
261 KB PNG
>>109588618
I dont need too, thats the fun part.MoEfags are insecure not me. I am so sure that i put my money where my mouth is
>>
>>109588621
that sounds expensive. are you running this all locally?
>>
>>109588506
that is insane anon, we clearly have different workflows and usecases. 40k tokens before the sessions even starts is..very unrelateable for me, not sure I would ever need or want that
>>
>>109588642
You know damn well he's not local with that setup.
>>
>>109588621
This is the terminal state of MoEFags
>>
shud I buy 8 more 6090's?
>>
wtf happened to opencode bros it was a cool project and now it’s going to shit
>>
>>109588642
>>109588654
he would need to be running constantly at 826t/s tg to get that many tokens in a single week
>>
File: 1780137720930743.png (698 KB, 1206x677)
698 KB PNG
>>109588534
Keep waiting, s.. surely price will come down eventually. On the plus side waiting gets easier to more priced out I get
>>
>>109588621
pi is only like 1200 tokens and it usually beats all those fat harnesses
>>
>>109588621
I work with a guy who works like you.
> minimal tools
Doubt, but OK.
When I need that kind of context, I use 35B Agent world with RoPE for a million tokens. Offloading all experts to ram and I still get acceptable performance with my 16gb card. But as I learned more and my projects becaume more modular I found I didn't need quite so much context.
>>
>>109588626
MoE is currently winning until we get a Nemotron Ultra sized model that isn't shit, sorry to break it to you.
t. Blackwell+256GB DDR5 haver
>>
>>109588659
>npm install harness
>npm install 10000-skills-megapack
>$200/month pro subscription
>ok magic box do the thing
can't be wrong if everyone is doing it
>>
>>109588673
>I use 35B Agent world with RoPE
3.6-35B-A3B ? and what is RoPE ? im also on 16gb vram, would like to hear more about this anon
>>
>>109588676
densies always have llama 405b kek
>>
>>109586746
Eurokeks working at meta are why we have decent open weight models. It's why llama.cpp has llama in the name.
>in Europe
regulation hell
>>
File: 1769216082361203.jpg (191 KB, 500x500)
191 KB JPG
>>109588676
At the end of the day i have the option to run MoE or Dense and i would rather run Dense. Until the day when i need to run MoE happens you will not hear a concession from me
>>
>>109587839
It's cope and benchmaxxing essentially
>>
>>109588698
I think this general has exactly 1 anon that can run that without spilling over into RAM kek.
>>
>>109587723
no fluffy bunny tail?
>>
>nemotron 550b benches the same as ling
nvidia what are you doing...
>>
>>109588696
https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B
It's a refresh of 3.6-35-a3b, like how 3.8-27B is a refresh of 3.6-27B but without the hype.
Gwen published it without the multimedia projector or the MTP head, but you can use the ones from 3.6.
>RoPE
It's a way to extend max context well beyond what the model was trained on. But it makes your short context dumber, so don't do it unless you have to.
Here's my config. I'm getting ~30tps and 200 pp

[Qwen-AgentWorld-35B-A3B-UD-IQ4_XS]
model = Qwen-AgentWorld-35B-A3B-UD-IQ4_XS.gguf
alias = Qwen's NEW agentic dev model
chat-template-file = froggeric-qwen35-v21-chat_template.jinja
no-mmproj-offload = 1
; from bartowski
mmproj = mmproj-Qwen_Qwen3.6-35B-A3B-bf16.gguf
temperature=0.6
image-min-tokens = 1024
top-p=0.95
top-k=20
min-p = 0.05
;rope-scale = 2
;rope-scaling = yarn
;yarn-orig-ctx = 262144
;override-kv = qwen35moe.context_length=int:524288
ctx-size = 262144
swa-checkpoints = 48
; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
cache-type-k = q4_0
cache-type-v = q4_0
batch-size = 4096
ubatch-size = 512
;ncmoe = 24
ncmoe = 18
main-gpu = 1
tensor-split=1,0
spec-draft-type-k = q4_0
spec-draft-type-v = q4_0
; from bartowski
spec-draft-model = mtp-Qwen_Qwen3.6-35B-A3B-Q4_0.gguf
spec-type = draft-mtp
spec-draft-n-max = 1
; more configs...
load-mode = mmap
fit = on
n-gpu-layers = -1
jinja = 1
reuse-port = 1
threads = 18
threads-batch = 18
timeout = 7200
# cache stuff
cache-prompt = 1
cache-idle-slots = 1
cram = 16384
backend-sampling = 1
spec-draft-backend-sampling = 1
>>
>>109588110
I've only been playing around with this stuff for two weeks - what happens exactly when you run out of context and what's the best way of dealing with it? Compacting?
>>
>>109588783
I un-comment the rope stuff when I need more context, but I haven't needed it since I switched to pi.
>>
>>109588007
i ran IQ4_XS_STOCK in ik_llama, that quant doesnt work in llamacpp and the repo linked in the PR only has Q3_K_M so yes i used Q3 but IQ4_XS works just fine on my machine
for glm air i used IQ4_KSS and for qwen i think i used IQ4_XS
i also used qwen 235b hehe like iq2
>>
File: a_flatter_gemma.png (848 KB, 675x1187)
848 KB PNG
>>109588242
It keeps adding a cleavage after specifying multiple times she's a child, has a flat chest (or even not mentioning it at all), etc. The only way that sometimes works is using a different reference, but I wanted a specific look, not the generic flat loli in bunny suit that can be easily found on gelbooru. To me it appears that for H3 bunny suit implies big breasts.

This model is a PITA to use sometimes.
>>
>>109588777
making hardware
>>
>>109588783
wow thank you anon. this has WAY more args set than my usual configs, going to dig into this and give it a try. do you think the UD-IQ4_XS quant is decent? I keep being told to avoid UD
>>
Is there any way someone can explain to me what actually makes unsloth worse, or is even asking this question akin to trolling? I am a newfag I know.
>>
>>109588110
im using 3bit 27b with 81920 context at 20t/s (0ctx, havent tested how slow it is at 80k) on a 3060
>>
>>109588825
iirc their quants are automated with little quality control or testing
>>
>>109588825
go back
>>
>>109588805
>iq4-kiss
Cute!
>>
>>109588825
It's just retards on 4chan who hate anything popular, the same reason they hate Qwen. Ignore them.
>>
>>109588862
I wonder what ethnicity this anon is
>>
>>109588841
I see, that would explain why they have stuff like day 0 then so makes sense
>>109588844
Yes I will go back to huggingface and download new models :)
>>109588862
I think people like qwen here though
>>
>>109588825
Automated quants using a generic template with no regard for optimizing for individual model infrastructure and minimal to no testing on which hot weights deserve more size allocated to them when dynamically scaling experts down.
>>109588862
Stop pushing slop daniel.
>>
>>109588809
the "more configs" stuff I put at the top as the default, under [*]. I should have made that clear.
>is UD decent
it seems to be. From what I gather from reading their stuff, UD is like other Dynamic quants where they just quant some stuff larger than Q4 when they think it's important. iquants are doing something along the same lines.
I know this general likes to shit on Danial but that's all chan drama.
I use Froggers chat template for all the Qwens, so if unsloth fucked up the chat template I wouldn't know.
Turn up the debug level on llama.cpp and look for the stuff about later offloading. With dense models try to get every layer offloaded. With moe's, it's about balancing the context, expert offload (with ncmoe", and ubatch size. More ubatch == more pp. So your tradeoff it context vs TPS vs pp.
Have fun Anon!
>>
>>109588825
They are a history of fuckups such as uploading 1KB quants of large models but when they don't fuck up their quants are often better than those of randoms because their calibration set is not just wikitext (or none at all).
>>
>>109588888
>I think people like qwen here though
We don't. This thread is regularly astroturfed by marketing influencerjeets immediately after a new model releases. If you want to gauge actual opinion, see how anons refer to models that are a couple months old and not consolewarring with newer releases.
>>
I’m going to fuck all y’all gemmas and keep mine pure and virginal
>>
>>109588839
wait i get 24t/s at 170w (no -lgc limit) 0ctx
meanwhile i get 20t/s at 120w (-lgc 1500) 0ctx
>>
>>109588912
>copy gemma.gguf
>fuck gemma (Copy).gguf
>delete gemma (Copy).gguf
>repeat
>>
ew, there are unsloth defenders here? i bet they're densesissies as well, gross
>>
>>109588912
I doubt you can run my fp16 31b gemmy before she sees your tiny t/s pp and steps on your balls while calling you a vramlet.
>>
>>109588888
afaik they get day 0 because they get the models early, before publication, specifically so they can have quants available on publication day.
>>
>>109588783
>It's a refresh of 3.6-35-a3b
it's not
>>
>>109588927
Don't lump us in with those bugmen.
t. blackwellGOD
>>
>>109588940
Read the damn model page.
>>
>>109588952
read the damn paper luddite
https://arxiv.org/abs/2606.24597
>>
>>109588928
I’m correcting your gemma first (on runpod)
>>
>>109588962
>training pipeline: CPT
eh, idk it sounds like a finetune to me
>>
>>109588825
Unsloth is quite good. They have the best quants; you can see graphs showing perplexity for weights, and they are always the best. They also provide fixed Jinja templates with their models. They don't just run llama.cpp quantize and call it a day; they properly use dynamic quantization and fix all the problems with the base model.
>>
>>109588993
it's designed to create text for RL of other models as far as I understand. qwen team hinted something better than 35b-a3b may be coming recently
>>
>>109588930
>>109588890
>>109588904
I guess unsloth is just a divisive topic with reasonable points to go either way, and then there are also shitposters >>109588996 who take advantage of this as well? It seems like I won't see any consensus in any case.
>>
dariobot…save us
>>
>>109588962
>https://arxiv.org/abs/2606.24597
ok, so it's an update to 35B NOT like 3.6 to 3.8. But it's based on 35B, and updated. Ok asshole? Does this meet your pedantic little ideas?
>>
>>109589004
You'd probably be better off just doing side by side comparisons yourself on the same quants made by different users to see if there are any issues with the model you actually want to use.
>>
>>109589004
I guess you will have to try it.
>>
File: 1760004657003558.png (564 KB, 1427x1078)
564 KB PNG
>>109589014
>Consumer GPU prices tank
>But only because all local models get banned and your not allowed to own more than 16gb of vram
>>
>>109589032
It is based on the same architecture as 3.6 35b a3b, but it is not an update to 3.6 35b a3b. It's a completely different training recipe intended for different purposes, not updated.
>>
>>109589052
>2033
>Florida man arrested for having sex with a local model.
>>
>>109588825
unsloth make fast food quants - they'll always be there and it'll always be pretty much adequate, but if you really care you can often find something better
it's more that they're mediocre than bad, but you can't blame them too much considering they produce every possible quant of every possible model, you have to expect that approach isn't going to lead to hyper-optimized results
>>
>>109589073
Things have escalated, it will be more like
>Terrorist apprehended with military quantities of AI capability
>>
>>109589014
only petrus is at your disposal
>>
>>109588549
If you could run fable locally you would do it in a heartbeat
>>109588719
>i have the option to run MoE or Dense
Sure, just like you can use a chiplet or monolithic processor. You're always going to use chiplets because at every price point the chiplet processor offers more.
>>
>>109589033
>>109589042
Yeah that's what I've been doing and I noticed that scotoma gave better erp than the unsloth gemma I was using beforehand which lead to me posting the original question
>>
>>109589087
>The authorities conducted a search of his residence and discovered an unregistered personal computer with an Nvidia GTX 1080 Ti graphics card installed.
>>
Has anyone tried the unsloth.ai desktop app here yet? Seems interesting, but looks just like a basic bitch harness for retards like the one PewDiePie made right? Unless I'm missing something, I'll just stick to Pi
>>
>>109589188
>but looks just like a basic bitch harness for retards
>Seems interesting
??? Is this how double digit IQ shills sell their product?
>>
>>109588311
2x DGX Spark + 1x 3090 eGPU.
4x DGX Spark + 1x 3090 eGPU.
>>
>>109589204
> double digits
well an IQ_16 doesn't sound that bad, that's probably more IQ_1 levels
>>
>>109585361
Summer dragon
Coomageddon
Smedrins
The mormon
And 50 gorillion beaks
>>
>>109589413
I was there...
>>
>>109589389
too clever
>>
What’s this SillyTavern Front end and back end shit that you guys recommend. Is there any reason that I shouldn’t just use LM Studio or something simple? Only experience I have with local models so far is AMD Chat, so I’m out of the loop.
>>
so is the new qwen 27b fable lite like reddit hyped it up to be? good enough to vibe code a semi-serious project as a nocoder?
>>
File: 1770366095530263.jpg (49 KB, 878x926)
49 KB JPG
>gemma-chan when she sees her new user is a fat oji-san
>>
>>109588409
Even worse, the "best" AI GPUs workstation are much more expensive.
https://www.microcenter.com/product/702853/AMD_Radeon_AI_PRO_R9700_Single_Fan_32GB_GDDR6_PCIe_50_Graphics_Card
https://www.microcenter.com/product/709007/Arc_Pro_B70_Single-Fan_AI_-_Workstation_Graphics_Card;_32GB_GDDR6_Memory;_PCIe_50_x16_Interface;_32_Ray_Tracing_Units;_256_Xe_Matrix_Extensions_Engine
And this is the best case scenario in the US. Otherwise, on Ebay, it's 2.5k+ and 2.2k+ for these cards.
>>
>>109589522
Have you tried googling it? Or asking your AI?
>>
>>109589533
Yes it solved all my problems and made my penis grow an extra inch and a half
>>
>>109589533
go back
>>
File: 1765923099202129.png (2.55 MB, 1079x1531)
2.55 MB PNG
>>109589541
now I can self insert
>>
how to make deepseek think in character again?
>>
>>109589584
prefill
>>
>>109589575
hmmm...nyo~!
>>
>>109589605
pretending to be me won't save you
go back
>>
Do I have to make all the threads around here?
>>
>>109589619
Apparently.
>>
>>109589584
Tell it to think from the perspective of the character in its think tags. If you want RP specifically, you can also use the official deepseek RP prompt.
>>
>>109589619
You better not be the jeet baker or I'll bake it myself
>>
>>109589619
Does the red number scare you?
>>
>>109589579
Where did you get this picture of me and my wife?
>>
>>109589619
Gemma saves the day

>>109589651
>>109589651
>>109589651
>>
>>109589656
We could have kept using the current thread for another hour. You are retarded.
>>
>>109589749
No one was posting
People crave fresh bread
>>
File: gemmy2.png (1.8 MB, 1280x1892)
1.8 MB PNG
>>109589776
putting a bread in gemma's toaster oven
>>
>>109589749
Wait until this guys hears about /ldg/
>>
>>109589788
>/lgd/
Just checked lmao, they are about as bad as /smg/ is on biz
>>
>>109587678
>paid more than I did for a watercooled 3090 and 112GB of RAM because I got lucky
I am feeling sympathetic pain. Ow.
>>
>>109588904
I thought Unsloth only used Wikitext?
>>109589124
Scotoma vs. Unsloth is nothing to do with quants. It's a finetune.
>>
>>109588311
No, because there are decisions.

Best at gaming -/-> Best at anything else.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.