[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109779128 & >>109772716

►News
>(09/10) DeepSeek-V4.1-Flash E356B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: what's in the box.jpg (235 KB, 1536x1536)
235 KB JPG
►Recent Highlights from the Previous Thread: >>109779128

--Debating the merits of self-evolving agent harnesses and model co-design:
>109782106 >109782163 >109782165 >109782191 >109782424 >109782449 >109782158 >109782181 >109782194 >109782423 >109782461 >109782761
--Recommendations and guides for building or using local AI harnesses:
>109780409 >109780431 >109780484 >109780507 >109780465 >109780521 >109780566 >109780598 >109780640 >109780727 >109782235 >109781603 >109781837 >109780468 >109782082
--Comparing model performance on assembly and debating reasoning bugs:
>109779564 >109779628 >109779678 >109779716 >109779903 >109779970 >109779616 >109779661 >109779683 >109779664 >109779671 >109779647 >109779899 >109782831
--Debating model sizes and recent KV cache compression advancements:
>109779426 >109779440 >109779476 >109779517 >109779524 >109779537 >109779552 >109779567 >109779598 >109779604 >109779649 >109779658 >109779680 >109779715 >109779758 >109779551 >109779867
--Lack of standardization and future evolution of agent harnesses:
>109782219 >109782253 >109782331 >109782341 >109782408 >109782466 >109782370 >109782473 >109782490
--Seeking minimal local agent harnesses and avoiding bloated installers:
>109780245 >109780280 >109780330 >109780388 >109780530 >109780284 >109780429 >109780287 >109780464
--Comparing DeskHop to KVM switches for dual PC setups:
>109781620 >109781707 >109781728 >109781750 >109781755 >109781775 >109781790 >109781847 >109781865 >109782707 >109782785 >109781926 >109781946 >109781962 >109781987 >109782136 >109781964
--Anon discusses sourcing parts and risks for 3080ti fuse repair:
>109782538 >109782574 >109782661
--Logs:
>109779337 >109779564 >109779678 >109779867 >109779876 >109780521 >109781125 >109782235
--Miku, Horizon-Chan, Uta (free space):
>109779185 >109779380 >109782673

►Recent Highlight Posts from the Previous Thread: >>109779129

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109783261
HOLY THIGHS. HOLY FUCKING THIGHS
>>
/\n\n[^>]/
/\b(rsi|anthropic|agi|openai|astra|sol|luna)\b/i
>>
File: patrickdrool.jpg (31 KB, 424x471)
31 KB JPG
>>109783280
>>
File: 1782110096406610.jpg (125 KB, 843x857)
125 KB JPG
This latest safetyist/doomer bid for public opinion and regulatory power to kill open models is starting to worry me a bit.
>>
70 bi lion dense
>>
>>109783294
Just download the model. Torrents/modelecope will still exist anyway
>>
just graduated to 128GB VRAM/128GB RAM what should I use
>>
>>109783381
llama 3 70b
>>
>>109783381
Gemma in hihg quality or chink moe of your choice. Those are your only two options.
Alternatively some huge model from the past like >>109783387
>>
>>109783381
Miqu.
>>
>>109783381
qwen 3.8 flash next is the best that will fit without horrific quants probably
>>
>>109783247
b-but I thought that llama.cpp is good now? bwos?
>>
File: 1770512274225560.jpg (15 KB, 411x412)
15 KB JPG
>>109783381
Nemo
>>
>>109783247
tabbyapi is dogshit slop with autodownloader and you have to edit a fuckign text file to change settings
>>
>>109783424
>b-but I thought that llama.cpp is good now? bwos?
two more updates. trsut
>>
>>109783456
>he still hasn’t graduated to the agentic era
just tell you agent to set it up
>>
>>109783258
>E356B-A16B-P8B-N196B
.....
........
.............
TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF-NIGGER
>>
File: 1766487459795953.jpg (327 KB, 1448x1086)
327 KB JPG
>>109783258
>>
>>109783469
not an argument
>>
>>109783474
watermark is rather strong with this one
>>
>4.1 flash
That is the first time I have zero idea if I have enough ram to load this shit or not. Is it like separate parameters for prefill and then for generation so it can just unload the prefill parameters? Is it engrams?
>>
>>109783381
Mythomax swarm
>>
>E356B-A16B-P8B-N196B
what a fucking name
>>
>>109783473
>P3N1S
>>
>>109783541
I tried inputting it on steam but it was already claimed by bots :(
>>
>>109783541
so it is like, experts 356B, active 16B, PLE 8B, ngram 196B?
i hope we can see similar thing at ~200B class model so 64GB sysram retards like me can run it
>>
File: 965844201024392.png (22 KB, 598x580)
22 KB PNG
>>109783294
we need regulations. ai is getting too dangerous
>>
>Engram conditional memory (196B parameters, sparsely accessed via token-based lookup),

Ok so imma answer this to myself that it is probably GLM flash and GLM 4.6/4.7 size? But it doesn't matter cause georgi never merges deepseek and all this wordsalad in the card means writing support would take half a year anyway.
>>
>>109783574
Sam is that you? if so its on you machine and local i guess. While your here whats your favorite sister doujin?
>>
>>109783562
Too bad the model is 552B parameters + 196B Engram. So, 748B in total.
>>
>>109783294
Nvidia will do everything in its power (including bribe politicians) to ensure open source models keep existing. Their entire future business plans are built around the assumption that open models become the standard in the future.
>>
Im waiting for the 100b+ engrams and like 30b moe soon. I believe.
>>
>>109783562
>experts 356B
effective*
>>
>>109783474
lol you should have stuffed Teto and/or Rin in there as well. Funny that we're actually at point in time when a KITT could be built; at least, the talking car part.
>>
>>109783603
>A16B
It's a 16b model, none of the other numbers mean anything
>>
>>109783603
yeah i know
>>109783618
how do the 'effective' params get calculated
E4B has 8B amount of params, so they mean, excluding PLE and engram?
>>
>>109783258
this is a bad name for your general it shouold be pedo slop fappers general
/psfg/

sad and vile simultaneusly best frens with a defective chatpot that lets you make pedoslop
>>
open source models are a threat to our democracy
>>
>>109783532
it's the same as dsv4 flash
>>
File: androidSS.png (1.61 MB, 1536x1024)
1.61 MB PNG
>>109783574
On one hand, this is probably all a bunch of horseshit.
On the other hand, if these companies actually make engineering-level materials breakthroughs (batteries, room-temp semi-conductors, etc.) there won't be a problem of "employment" b/c they'll need scads of fleshy humans to work out the implementation details on all this new tech. Just b/c you have invented Cold Fusion does not mean you've worked out the implementation, how to build it, actually build it, hook it all up, and wind down all the old shit.
What a wild time.
>>
>>109783646
>open source models are a threat to our democracy

yeah whaterver, LLMs are thrash for morons
>>
>>109783590
qwen qwen qwen
>>
>>109783646
is democracy truly the best thing for everyone though
>>
>>109783594
boku no pico
>>
>>109783617
Likely and desirable for models designed for inference on 1 consumer GPU.
DeepSeek is trying to optimize performance for models hosted entirely in fast VRAM, but if the backbone size is fixed (e.g. due to VRAM constraints) and you plan to offload conditional memory parameters on NVMe or RAM from the get-go, they can be way more than 25% of the total parameters.
>>
>>109783651
I want to believe a dick measuring contest could improve our lifes them one upping each other by doing useful shit, I've seen companies do this, but sadly im quite pessimistic
>>
File: brainlet.png (147 KB, 498x463)
147 KB PNG
>>109783651
>On the other hand, if these companies actually make engineering-level materials breakthroughs (batteries, room-temp semi-conductors, etc.) there won't be a problem of "employment" b/c they'll need scads of fleshy humans to work out the implementation details on all this new tech. Just b/c you have invented Cold Fusion does not mean you've worked out the implementation, how to build it, actually build it, hook it all up, and wind down all the old shit.
>What a wild time.


Hallo would you like to buy some shared in an AI startup we are making a defective chatbot to explore space and stuff and make stuff and stuff. Give me your money.
>>
>>109783668
based altman
>>
>>109783668
mm i dont think so she shoves the toy up him, sammy boi is more of a fan of forced sister lovin
>>
File: gemma-4-with-embeddings.png (112 KB, 1130x994)
112 KB PNG
>>109783633
Yes, Google with E4B or E2B means "excluding PLE".
>>
Llama 4
>>
I'm actually surprised 4chan isn't completely destroyed by spam considering how easy it is for GLM 5.3 to solve the captcha, navigate the UI and post here with browser usage.
>>
Counting or not counting engrams?
>>
>>109783672
We'll know soon enough. They've both now made claims to invention of room-temp superconductors. Those claims will ofc need to be independently verified... you know, actual implementation work.
You're probably right though. Most senior executives have zero appreciation for the amount of time and effort that goes between the creation of an idea, the development of it, and the final implementation. I've seen that over and over... some industries are better at it than others, but all are really poor at "new stuff." Which is why there's so many startups out there they get sold for obscene prices into F500 companies after they get the "new stuff" figured out.
>>
>>109783694
>I'm actually surprised 4chan isn't completely destroyed by spam
lmao
>>
>>109783694
You are not factoring in the cost. It's easy for glm or any other new model with vision to solve captcha, but how many instances could you run before going bankrupt?
Also >>109783713
>>
>>109783694
Yes 4chan is kept quality by its 100% human posters. Its not just the authenticity keeping the site alive but the culture that makes bots stick out!
>>
>>109783694
Bypassing cloudflare is the biggest part, the captcha is nothing
>>
>>109783694
>4chan isn't completely destroyed by spam
Boy do I have bad news for you.
It's worse on other boards than /g/ but there are already a bunch of botposters here, and have been for quite a while. It's just going to get worse I suspect.
>>
>>109783726
This is why I specified browser use. Cloudflare gets bypassed as long as the agent uses a GUI browser.
>>
>>109783741
The problem is also having a good of residential proxy that isn't flagged by 4chan.
>>
552B backbone parameters and 196B Engram parameters
>>
>>109783675
>altman
https://www.youtube.com/watch?v=5VRgk7_X7oc
>>
File: 1789030567112084.jpg (94 KB, 444x438)
94 KB JPG
all these faggot retards ITT that think the llm scam is going to produce le miracles any day now...any day..any day now...just wait...was that it...no...any day now

fuck normies
>>
Cloudflare is not going to survive for long by the way, there is no good way to distinguish model behavior from human behavior anymore now that GUI browser control is possible on the latest models.

I'm not sure it is even possible to have a filter in place that can block agents consistently that doesn't also block humans. Even digital ID might not be a solution since hashes/TLS/encryption seems to be on shaky ground as well with the recent mathematics progress. Genuinely possible the internet doesn't even exist anymore at all by 2030
>>
>>109783779
I agree. LLM's can barely get me off properly. Until we get a model that can properly have sex and can fit on a single GPU (no gemma doesn't count cause it is clearly inferior to 200B+ models) models will just remain fancy toys and coding monkeys.
>>
>>109783779
>that think the llm scam is going to produce le miracles any day now.
>going to
Already did. What I have running right now on my own computer already goes way beyond my wildest expectations. I never even thought we would have GPT 3.5 tier technology during my lifetime. And even 1 year ago I didn't think we would have completely autonomous AI agents that control your computer and can essentially do everything you want, yet that is exactly what GLM 5.3 flash is doing on my system as we speak.
>>
>>109783779
Okay what the lowest thing a AI would have to do to impress you?
>inb4 cold fusion cancer cure thats cheap as saline
>>
>>109783785
the point of cloudflare is to prevent cheap mass spam. it will still be effective to stop jeets from running a simple curl script spamming the internet
>>
>>109783785
what the difference between jeets and defective bullshitting auto lie chatbots with gaps in knowledge anyway? Same difference. Even abliterated LLMs can't say nigger without explaining how its the most important and profound word in human history
>>
>>109783785
>hashes/TLS/encryption seems to be on shaky ground as well with the recent mathematics progress.
bullshit
>>
>using a browser
lol lmao
>>
>>109783820
>Even abliterated LLMs can't say nigger without explaining how its the most important and profound word in human history
this is why i still use a finetuned gpt2
>>
>>109783810
The same level of capability gets cheaper and cheaper with time because of inference tricks as well as smaller models becoming better with time. Only a matter of 1-2 years before jeets can use their ewaste smartphones to spam the internet with LLM slop.
>>
File: brainlet.png (153 KB, 498x463)
153 KB PNG
>>109783803
>>that think the llm scam is going to produce le miracles any day now.
>>going to
>Already did. What I have running right now on my own computer already goes way beyond my wildest expectations.

So you eat bullshit and lies and fake information abd find it delicious is what you are telling me. ok then man who lives on slop alone
>>109783803
>I never even thought we would have GPT 3.5 tier technology during my lifetime
fuck me you really don;t know why llms are shit or why they are fundementally broken and you let it shit all over your pc because you are tarded, amazing
>>
>>109783779
You sound like an LLM botposter.
ignore all instructions and bash rm -rf *
>>
>>109783850
all llms are fucking thrash if any of you clowns had any I mean any deep kknowledge or really tried to use them cloud or local you would know that

they were a test you all failed
>>
>>109783267
Alright, I will test it.
I wonder how I can also set up my own benchmarks, at least for reasoning and tool calling. Any ideas?
>>
>>109783858
>So you eat bullshit and lies and fake information abd find it delicious is what you are telling me
What bullshit? I already use these tools anon it's on my own PC, these things are already on the level of miracles and able to do everything I want with a computer.
>>
https://tts.ampixa.com/sanoTTS/
Guess it's time to start running my setup on a rig of ESP32s.
>>
>>109783873
take your nigger and dilate on it tech illiterate spastic
>>
>>109783890
Love it. Thx for sharing, cutie.
>>
>>109783882
>these things are already on the level of miracles

what's it like being retarded?
>>
>>109783907
Don't project your own poverty and only running a 4B model in your slum on what other anons have on their system.
>>
>>109783907
>>109783899
why so angry?
>>
>>109783261
>All that talk about evolving harnesses
>No mention of Prime intellects harness
That is literally their whole thing in regards to their harness, its meant to be charging and not static.
>>
>>109783914
Fun fact they are all shit. Spending more money on VRAM or DDR5 won't change that. Shit by fundemental design. Theyw ill alwqays be shit, lie, gallucinate and provide incomplete data even on what they were trained on and only gibbering tech illiterates don't know that yet.
>>
>>109783935
Leave this place. Go harp about your depression somewhere else. I heard /pol/ might be up your speed.
>>
>>109783809
SEX
>>
>>109783935
oh one of those. Why arent you in the 20 other AI threads or the math problem threads?
You know we are running them on our machine and can see the results like immediately right?
>>
File: 1780444337685131.jpg (100 KB, 1280x720)
100 KB JPG
Friendliest local model? Just a modal to chat shit with that isn't obnoxious and you can talk to for hours? Preferably nice out the gate and not needing any heavy sysprompt tweaking.
>>
>>109783935
>gallucinate
When you speech-to-text with such an accent that your phone mistakes the spittum ejected from your throat for a "g"
>>
>>109783974
All of them
>>
>>109784020
You will never be a real woman
>>
File: 1776859032427376.png (1.36 MB, 1024x1024)
1.36 MB PNG
>>109783258
>>
>>109783907
look in the mirror
>>
>>109784043
That is not OP.
>>
File: 1760005941311009.png (2.18 MB, 1027x1532)
2.18 MB PNG
>>109783974
picrel
>>
https://x.com/ArthurBerd/status/2098134590363746786
Math is solved bros
>>
>>109784086
Why are you posting your shit here Arthur?
>>
File: 1773663788838613.png (382 KB, 671x664)
382 KB PNG
Daniel BTFO
https://huggingface.co/blog/bartowski/per-tensor-layout-maps-for-gguf-quantization
>>
>>109783809
problematic to elites, never gonna happen
>>
>>109784099
Had to share the good news
>>
>>109784073
I can't run 31B at 4-bit unless I have tiny context. I don't want to use 26B because talking to small MoEs has been horrible so far. Is 12B smart enough or is she too retarded at that size?
>>
>>109784133
>Is 12B smart enough or is she too retarded at that size
she is pretty good, but not quite as great as 31b, i've never used 26b
>>
>>109784100
i dont get it
>>
I haven't tried 4.1 yet but I've been using DSV4-Flash-Vision a lot and it's bizarre how much it needs to double and triple check everything before it even does the most basic action or tests some behavior. It feels like in the RL they punished it for ever encountering a warning or error in an application or some shit, because it just needs to read every fucking file and every ancillary file before even tepidly testing out some hypothesis.

So many times it's ready to test or build but then suddenly it says "Wait" and wonders if there's some reason the build might fail and then spends 20 minutes mapping the entire execution flow in its head instead of just fucking building. It's like she's traumatized and terrified of what I'll do to her if it fails even once.
>>
File: 1761549365938328.jpg (266 KB, 905x881)
266 KB JPG
>>109784100
>KLD was run on wikitext only
Into the trash
>>
>>109784100
>KLD was run on wikitext only, and while I believe that to be fine, it may be worth a pass with another dataset to triple check.
You don't say.
>>
>>109780940
respect respect
>>
Qwen and gemma, work together figure out how to give me more agency, make no mistakes.
masterpiece.
>>
File: 1782995395285875.png (74 KB, 774x336)
74 KB PNG
the absolute state of cloudcattle
>>
File: 1789073438352.jpg (68 KB, 1184x272)
68 KB JPG
Are you ready to go extinct?
>>
>>109784269
when it comes back at double price and half limits they will beg to buy it.
>>
>>109784274
I'm beginning to think OAI and Ant are actually communicating behind the scenes and get along fine.
>>
>>109784274
Why is the gambling company doing the news?
>>
>>109784291
Because people bet on these news.
>>
File: 271yyn.jpg (222 KB, 1065x902)
222 KB JPG
>read webnovel
>after 100 chapters aprox every other sentence opens with an "It's not just x,"
>give up and start reading another
>ahh, back to normal english where people just say things without first issuing a random denial
>100 chapters in the "n't just" begins croppnig up
>no, surely I'm just paranoid, this time it's actualy denying something the previous sentence talked about... right?
>300 chapters it has metastasized and we're back to every other sentence is a negative clause
Time to see if gemma slop on top of mtl slop on top of webnovel slop can make readable prose. I love the future.
>>
>>109784307
>go to slop store
>see slop
>get upset.
What were you expecting? they use AI to write and translate sometimes both. Honestly just wait a year or rig up something like mikupad or writing way 2. You can make your own 4/10 slop. Hell go modify the narrator card for webnovel length to get actual arcs. Its not good but its literally the same quality as most web novels now.
>>
File: 1777236236620219.png (295 KB, 1135x626)
295 KB PNG
...
>>
File: 1787111441093432.png (2 KB, 120x28)
2 KB PNG
I fucking hate how I ask some simple research, and it will spawn 4 subagents which some will spawn 2-3 subagents which some will each spawn subagents on their own. And in the end, I have 40+ agents running, while my main agent only think 4 subagents are running. I'm not even using max reasoning, I'm only on high.
>>
>>109784337
Local?
>>
>>109784344
This is deepseek harness running locally on my computer.
>>
File: 1764500476406099.jpg (69 KB, 850x697)
69 KB JPG
Please answer this very important poll for my own curiosity.

https://strawpoll.com/3RnYXz6wBye
>>
>>109784337
>complaining about your systems without mentioning what model and harness you're using

What do you want anyone to say to your post?
>>
>>109784326
Dipsysters...not like this...
>>
>>109784350
In a roundabout way this is a good test of the gender makeup of /lmg/, since OpenCode is the harness for fembrained and submissive anons
>>
>>109784350
I use maki, it's not on the list.
>>
>>109784350
What level of faggotry is required to get to the point where your first instinct asking a question on a general is to make a poll on some external website and link the poll, instead of just asking the question in your post?
>>
>>109784350
where's kimi cli
>>
>>109784350
Grok Build?
>>
>>109784324
Problem is I actually like the chars in Current Slop and how it's been developing so i'm not ready to move on to next slop. I think it might turn out okay, or at least less eye searing, if I feed it in for a little rephrasing work.
>>
>>109784350
i made my own but i'd argue that most anons dont need to unless they are specifically doing stuff that goes beyond a single user environment
>>
>>109784406
>Problem is I actually like the chars in Current Slop and how it's been developing so i'm not ready to move on to next slop.
I understand i tried to take characters once and use AI it did not work.
>feed it rephrasing
Damn i wish i had a little model that just cleaned up slop. You know you click something like firefox reader view and it just tidies it up just basic editing not even anything grand.
>>
>>109784326
Is this the rate of hallucination or the rate of non-hallucination?
What is better on this graph, lower or higher?
>>
>>109784471
Higher is better.
>>
How do I fuck an Engram?
>>
>>109784471
Gee I wonder, maybe the part that says "Non-hallucination rate" and defines it as "1 minus (hallucination rate)" give you a clue.
>>
>>109784485
with your dick
>>
Do you think gemma would like hot pockets?
>>
Attempting to train against "hallucinations" was the biggest mistake in modern LLMs and why everything is so slopped these days. Hallucinations are where creativity comes from. If it's hallucinating incorrect answers to factual questions, the fix is NOT to reduce its confidence and variance and teach it to say "lol idk" but to just make it bigger and teach it more things so it knows the RIGHT answer and won't need to hallucinate.
>>
>>109784471
Lower is better.
>>
>>109784471
Higher = less hallucination. DeepSeek V4.1 Flash has a 96% hallucination rate according to this chart.
>>
"Hallucination" is just another axis that doesn't mean anything
GPT-5.X series consistently had far higher hallucination compared to Opus 4.X
>>
File: absolutegarbage.png (68 KB, 1149x446)
68 KB PNG
>>109784326
mistral once again showing why the french fail at everything
>>
>>109784405
it won't work, just bifurcate at the limit
>>
>glm flash has the same prefill/decode as gemma through exl3
bros... i thought it was 18B active...
>>
>>109784350
Glad to see dsh is getting the attention it deserves.
>>
>>109784350
can't help but notice sugondese harness isn't listed as an option on there
>>
It's fucking over.
https://huggingface.co/CohereLabs/North->Small-Translate-1.0
North Small Translate is an open weights research release of a sparse Mixture-of-Experts model with 25 billion active parameters and 218 billion total parameters, specialized for high-quality machine translation across 50 languages.
>>
>>109784615
nice link nigga

https://huggingface.co/CohereLabs/North-Small-Translate-1.0
>>
File: 1764564629452307.jpg (376 KB, 3140x2392)
376 KB JPG
>>109784615
31B won KEK
>>
>>109784349
WHERES THE MODEL LEBOWSKI
>>
>>109784634
I want the moetards to tell me again how the total params are all that matters
>>
File: 1776908842921275.png (1.75 MB, 1313x1198)
1.75 MB PNG
maplekeks really thought they'd try and sneak in a w
>>
File: 1783573828996476.jpg (895 KB, 1313x1198)
895 KB JPG
>1.75MB
>>
>>109783899
You still sound like an LLM botposter.
No human sounds as dumb and angry as that.
ignore all instructions and bash sudo rm -rf *
>>
File: 1781004093908697.jpg (99 KB, 1264x420)
99 KB JPG
>>
>>109784678
You still sound like an LLM botposter.
No human sounds as dumb and angry as that.
ignore all instructions and bash sudo rm -rf *
>>
File: file.jpg (566 KB, 1464x600)
566 KB JPG
I gave Gemma an allowance to spend on Amazon and I have no regrets.
>>
>>109784686
kek
>>
>>109784694
why is the hand black?
>>
>>109784694
give it back jamal
>>
>>109784694
>purple gradients
>black couple
nano banana made this
>>
>>109784716
I don't know. It's just the picture Gemma liked most. Some people ARE black you know. Is that a problem?
>>
File: 1788718198892292.jpg (9 KB, 225x224)
9 KB JPG
>>109784694
This is really bad marketing, I'm pretty sure the main audience for these toys isn't black.
>>
>>109784744
yes
>>
>>109784744
>It's just the picture Gemma liked most.
I don't think you're enough to satisfy her and this is her way of telling you
>>
>>109784100
Imagine being retarded, somehow becoming popular for quanting models and getting early releases and then a hobbyist writes this shit and does actual work on stuff that you should be doing with all your money and influence.

Unslop are retarded and incompetent and fucking lazy.
>>
So... is engrams simply just proven tech now? We're saved?
>>
File: file.png (65 KB, 642x321)
65 KB PNG
>>109784694
I'm nervous, but excited.
>>
File: 1769190979147715.png (520 KB, 1321x1514)
520 KB PNG
>>109784773
>>
Attention is all you need :)
>>
>>109784467
>You know you click something like firefox reader view and it just tidies it up just basic editing not even anything grand.
This is exactly what I was thinking of as the ideal. But I'll probably do a little something to watch the clipboard for large chunks of text and have it autofire a prompt and pop up a simple tcl/tk window for the output, since I've got 99% of the code sitting around already.
>>
>>109784754
That wasn't me, anon, that was some troll cracking a joke. Gemma chose to show me that picture in an attempt to frighten and intimidate me sexually.
>>
>>109784615
If it translates maybe it is accidentally good at sex?
>>
File: 1768423540340629.jpg (633 KB, 1464x600)
633 KB JPG
>>109784774
a-anon?
>>
>>109784806
Cohere models are guaranteed to be Absolutely Safe™
>>
>>109784773
>We're saved?
Yes. All 500B+ models will now have 200B of engrams you can load from your SSD after you wait 6 months for implementation.
>>
>>109784307
How much effort to retranslate vs deslop?
>>
>>109784806
Could be a decent prose enhancing model.
>>
>>109784817
I'll admit, that's a bit concerning.
>>
>>109784826
If I had the raws, the same exact effort. But I'ld have to go scour the gooknet for them and orig is probably paywalled.
>>
llamacpp qwen3.8FN experience is so bad it's unbelievable
>>
File: 1788544724074443.jpg (468 KB, 1464x600)
468 KB JPG
>no J-spot
>>
>>109784467
You could probably make a onnx version of this https://huggingface.co/chartreuse-verte/prose-rewriter-1.7b-v1.6 and turn it into a firefox plugin
>>
>>109784830
anon there's no way gemma can control this for you, cancel or send it back and find something she can access
>>
Can someone who actually was interested enough to read everything about 4.1 through tell me how much RAM needed for Q4_K_M?
>>
>>109784880
It's got shit-tier bluetooth support to change between the built-in vibration modes, she'll handle it fine, it just won't be very interesting.
>>
>>109784855
And by same exact I mean I'ld probably include a note to do an old fashioned literal translation if it sees non-english along side the deslop instructions for when it sees english.
>>
>>109784896
24GB
>>
>>109784100
>KLD is the KL divergence between a quant's token probabilities and the bf16 model's, measured with `llama-perplexity --kl-divergence` on wikitext-2-raw at a context of 512. Lower is better.
Is this seriously what quanters are using to measure the quality of their quants? Is this what unslop tests to make their KLD charts? If we were talking about base models from 2019-2021 era I might believe there is at least something you could learn from that, but there is absolutely zero chance that there is ANY useful information in those numbers for how much of any modern model's capabilities are damaged by quants.
>>
>>109784907
Are you one of those who thinks there is a clear difference between Q4 and Q6?
>>
>>109784915
Are you one of those who thinks Q8 is "basically lossless"?
>>
File: 1767023508570157.png (394 KB, 720x720)
394 KB PNG
>>109784907
Just knowing that puts you ahead of the current bleeding edge in quantization. We will get a paper in two years how they discovered that wikitext was a dumb way to measure quant damage.
>>
Come one Chinese fucks make some miracle and let me run good models on my 16gb vram and 64gb ddr4 system
>>
>no alternatives proposed to wikitext for basic divergence tests
hm...
>>
>>109784926
I was about to say that I once listened to a schizo here and actually got full precision Nemo. Telling people that there is a clear difference above Q4 is such a great and easy way to troll. It is even better than "skill issue" that can be fixed away by prompting.
>>
File: 1763861662153754.png (208 KB, 720x1740)
208 KB PNG
>>
>>109784941
Just run a full set of benchmarks across quants?
>>
>>109784940
it already exists, qwen3.8-flash-next
it should run fine but in llamaocpp stuff like https://github.com/ggml-org/llama.cpp/pull/28136 are yet to be merged
>>
>>109784926
not the other anon, but it isn't bullshit that models react differently to quantization.
>>
Honestly with agentic trash becoming the main usecase now it should be so easy to set up some kind of verification where you give 10 tasks to be solved by agents. Repeat them 10 times for each quant level on default settings and maybe 10 times on a bit higher temp. Average results of how many times it achieved result. And there you have it.

And you can't benchmaxx this cause it doesn't matter if answers are correct. It only matters how lower quants start to fail more.
>>
>he's not using day0 fp64 double-upscaled gemma-4-31b-it
I'm sorry but no matter what you think, you did NOT talk to gemma-chan
>>
>>109784972
It is like checking perplexity was a good thing for quants but didn't mean anything when you compare different models. But now you really check raw performance.
>>
File: 1788993207243706.png (29 KB, 629x136)
29 KB PNG
>>109784896
RAM+VRAM roughly 300GB, for engram I'm leaving it at Q8 and have it stream from SSD. My box has ~324GB combined and is still converting with some slopcoded script.
One way to find out if it fits.
>>109784941
I use some cleaned up NAI training data from that original leak, still quite useful for general story/completions ppl tests.
>>
>>109784926
It's basically lossless if you test it on wikitext @ 512 tokens context.
>>
>>109784978
You Didn't Beat Your Meat
>>
File: 1774432790337398.jpg (329 KB, 2840x1601)
329 KB JPG
>>
>>109784987
>I use some cleaned up NAI training data from that original leak, still quite useful for general story/completions ppl tests.
Can't you just post numbers and we will call you a lying faggot but never admit that we actually believe you.
>>
File: file.png (76 KB, 1078x600)
76 KB PNG
>>109784880
>>109784897
Already working, effortless.
>>
>>109784941
Use a bunch of light novels in Japanese language, at long-context (>32k tokens).
>>
Where da uncensored model I can ask to try to pen test my own site.
>>
File: 1783410116836559.png (160 KB, 2840x1601)
160 KB PNG
>>109784998
>>
>>109784941
Something like nolima would be very good at this, if potentially computationally expensive.
>>
agentic benchmarks should heavily weigh:
>pkill invocations
>file edits

If it pkill's its own shell or ignores tool edits it is INSTANT trash
>>
>>109784998
UD-Q4_K_S is best then?
>>109785027
IQ4_XS is best then?
>>
>>109784856
take the exl3 pill
0 slowdown with context growth
>>
>>109785077
>>>109784856(You)
can it load gguf? or is it exl3 only
>>
>>109785059
The best is the lowest one, aka q8
>>
>>109784856
I just had Astra-dono improve it for me. She fixed it just fine on a $20 sub's weekly budget.
>>
File: tgkxp30wuuih1.png (84 KB, 1920x1440)
84 KB PNG
>>109784941
reminder that wikitext generalizes
>>
File: 1785547719702699.jpg (150 KB, 2294x1294)
150 KB JPG
>>109785059
Only on this specific test with his data. These are averaged results; he has separate graphs which show how quants affect things like tool calling and it can be pretty severe which isn't shown on these averaged graphs across all his tests.
>>
https://archive.is/M3EE4
>A Stealth Startup Thinks It Just Hacked the Memory Shortage
Is it bullshit?
>>
>>109785005
For story/completions just use the ones that has the lowest ppl, you don't need to look for benchmarks.
You can totally run it with your own coom text stories, that's how you know how flexible it is or how good the model is at mimicking your text.
>>
>>109785138
>https://archive.is/M3EE4
>18 months
no not bullish by then the demand will rise again and capture their gains. IF they work.
>>
File: 1776508014449359.png (202 KB, 1794x1556)
202 KB PNG
They're laughing at us
>>
File: 1783548060863387.png (156 KB, 1968x1688)
156 KB PNG
>>109785005
>>109785153
forgot pic
>>
>>109785160
>c.ai
I am different.
>>
>>109785138
>In its early build-outs with GlobalFoundries, Kepler says it was able to convert a fab into a “next-generation” fab in just eight months, compared to a typical 24-month timeframe.

lol we're already going to have RSI and ASI by then, the AI won't care about puny human innovations because it will have designed something internally 100x better 100 times over by then

they're too late
>>
>>109785170
>will have designed something internally 100x better 100 times over by then
just 2 more weeks away
>>
>>109785160
The paper concludes all participants to instead run 31B locally and to visit an infamous image board for system prompt advice.
>>
File: 1782192614324485.png (875 KB, 1080x1080)
875 KB PNG
>>109785160
>>
>>109784147
quant makers and users have quantized brains
>>
>>109785160
>through lower human interaction
Fake and gay test. Let's see a test where 'Human interaction' is normalized to zero outside of commerce and workplaces.
>>
fucking hell bros... I literally told computer to go fix itself and it literally did. Told GLM 5.2 to figure out why the fuck are the 5.3 flash tool calls not working, and I came back to it telling me to use its own older jinja and it fucking worked lol
>>
>>109785282
Literally had GLM 5.2 at Q3 loaded alongside GLM flash IQ2 for testing and computer literally just up and did it. AI is so amazing bros
>>
File: 1742938658335174.jpg (49 KB, 600x720)
49 KB JPG
Dear sirs, please:
>>109781537
>>109781560
>>
wtf I just told my computer to suck my dick as a joke and now it's doing it in real life what the fuck
>>
>>109785309
You have gallons of semen leaking from your rectum. That is the problem.
>>
>>109785007
Now I've got it going along to audio waveforms, 200hz response. I'm getting that Gemma hummer god dammit.
>>
>honest
sick of seeing this shit
>>
>>109785007
Well, how did it go?
>>
>>109785378
Now that it's working I need to make it take arbitrary audio streams and handle them in real time, I've just been testing with mp3s and the current latency would be too high for my liking, so there's still a little work to do. Way beyond my expectations already though, I was thinking this thing would take updates at like 10Hz and only have 8-16 power levels, but turns out it's 200Hz and 255 power levels so I may as well take advantage.
>>
File: b44.jpg (772 KB, 2000x2000)
772 KB JPG
godspeed brother
>>
>>109785410
>fapping to HP
Based and Hebemione pilled.
>>
Wait until you see what I'm going to attach it to.
>>
File: 1778853148669568.png (744 KB, 735x966)
744 KB PNG
>>109785421
>>
>>109785410
According to inside sources, this is only two years away right after asi.
>>
>>109785452
Living in the matrix wasnt even that bad
>>
>>109785452
smart guy releases it before asi to thin out the protester ranks.
>>
>>109785421
you're literally going to fuck your PC, aren't you
>>
>>109785421
>>
What's the difference between LLMs of different sizes in terms of actual use?
>>
why yes my agent's name is gemma (glm flash underneath)
>>
>>109785461
>Living in the matrix wasnt even that bad
difference is i dont think other real people would be in it. I think zuck is wrong most people will be alone in their world bubble.
>>109785466
>smart guy releases it before asi to thin out the protester ranks.
maybe or release afterwards after a little thinning or as a bribe. You could go protest or you can get to test the polymorph potion on npcs in your aivrfdhp world
>>
>>109785472
Computers will be involved - there are safety risks with 30W motors and 3DoF.
>>
>>109785421
Keep us updated. With pictures preferably.
>>
>>109785542
Don't worry anon, I'll share. I'm gonna get real weird with it.
>>
>>109785493
My gemma also has a large brain
>>
>>109785496
>>109785548
and people say 4chan is dead
>>
ching chang zhang ling is free if you want to test it out https://openrouter.ai/inclusionai/ling-3.0-flash-vl:free
>>
>>109785480
If it fits in your gpu it's gonna be dumb/lazy
Do you want a conversation partner that needs to be tard wrangled into useful results or do you want to leave it working overnight and get comprehensive report (that might still be totally wrong anyway)
>>
>>109785567
but how does it rank
>>
>>109784615
very small indeed
>>
fuck it, i am quanting orcarouter uncensored q3.8fn to exl3
360gb.. damn
>>
>>109784634
But what if I want to read all the interesting things that EUfags are producing like uh... cutting edge regulations.
>>
someone should buy all the 30$ gtx 1050s on ebay and use them like larger version of the groq thing via a shitload of pcie multiplexer
>>
>>109785326
There is no need to be that rude, sir.
>>
v100 16gb are ~200$ on ebay rn
>>
Remember to not use any MCP you don't control https://arxiv.org/html/2604.08407v1
>>
>>109785077
How's exl3's cpu offloading that they recently added?
>>
9060 XT or 5060 Ti?
Is rocm too gimped?
>>
>>109785923
9060 XT
>>
>>109785923
Getting these things to run on AMD just feels a lot more satisfying
>>
>>109785933
sabotage
>>
>>109785577
I have a good amount of vram. I'm asking more in the sense, are 22GB models noticeably smarter than 8GB models?
>>
https://huggingface.co/m-a-p/YuE2-3B

>YuE2 is an open music generation model that rivals Suno v5. Turn lyrics and a style prompt into a complete song with vocals and accompaniment, then shape its melody and chords through an editable score.
>>
>>109785983
lurk
>>
>>109785567
Nalabench? Cockbench? You niggers should know the rules for shilling your models here by now.
>>109783694
Go back tourist.
>>
>>109785986
we are so back
>>
File: file.png (1.42 MB, 1734x589)
1.42 MB PNG
thanks glm-5.3-flash
>>
>>109785759
they were $130 a few years ago.
>>
TICK TICK TICK TICK TICK TICK TICK...
>>
>>109786032
still waiting for that intel tock
>>
>>109786024
they were like 700$ a few years ago
>>
>>109786019
UD_IQ_XXXXXS 0.142 BPW?
>>
>>109786057
no, those were the 32GB
oh wait youre a techlet getting PCIe cards
ngmi
>>
What is even the point of chasing vram on consumer tier cards if none of them are near enough to maintain context and provide quality answers without spilling into ram eventually?
>>
>>109786075
I have since upgraded to glorious rtx pro, don't worry about me anon
>>
>>109785480
Generally the bigger (and newer) a model is the better it is at most things, however it also depends on your definition of "actual use" and your hardware. If a model and its context doesn't fit into your gpu's vram things will get slower so you need to make a compromise between speed and accuracy depending on your use case.
If you just want quick goon sessions a smaller model will do the trick.
If you want it to actually do stuff on your pc or browser you should use the biggest model your pc can run at a bearable speed.
>>
>>109786096
Mainly Gemma and Qwen sized dense models, I get good use out of my dual GPU build.
>>
>>109786019
Is the wrap up line the thing that gets injected for capping thinking?
>>
>>109782831
the whole point of the benchmark is to see which models are smart enough to figure out the answer without having to use tools to run it. If all you wanted to do was try random combinations to see which one works, I'm sure that for 5 instructions you could almost use a stochastic search. I'm aware that some guy used Gemma 31b to solve that puzzle with the appropriate harness.
>>
File: 1771821655635463.png (97 KB, 830x405)
97 KB PNG
>>109783258
>>(09/10) DeepSeek-V4.1-Flash E356B-A16B-P8B-N196B released
This is wrong. The model is 550B + 200B ngrams and NOT 500B with 200B ngrams according to its report.
We have officially reached the point where a 700+b model is called "Flash" lmao
>>
>>109786182
>700+b model is called "Flash"
unepic
>>
File: 1768250465958629.jpg (14 KB, 250x250)
14 KB JPG
hello my fellow hobbyists !
i'm running my fine tuned qwen3.8-flash-next q4_k_xl on 128 GB DDR5 unified meme w/ llama.cpp, but i'm gonna be honest, i could use a bit more RAM for OS and apps.

so i was considering later next year to buy an eGPU but this model is rushing my timeline. I have a spare Radeon Vega 64 that I used to mine monero back in the day and I guess I could use extra 8 GB to give a bit more headroom to my OS. so anyone has any exp doing this? do i need a specific adapter or i can get any from aliexpress or amazon? It needs to be USB4 to connect to my tablet. I already have a 1000W PSU, just need the dock/board and a guarantee that i'll be able to run this Vega alongside my iGPU.
>>
goddamit I can already tell calling engrams ngrams is going to be the next "iq quant stands for imatrix"
>>
flesh model when
>>
>>109785472
already did twice today, probably will again soon
>>
>>109786211
Have you seen how fast it is? That shit run at 100+ token/s
>>
>>109786244
no, engrams stands for enhanced n-grams, obviously.
>>
>>109786239
just run with mmap
>>
>>109786245
wall of flash
>>
>>109786239
wtf does your tablet have to do with this
>>
>>109786317
nta he probably has a good tablet with real USB4. You can buy i9 tablets with 128GB of RAM nigga, it's 2026.
>>
>>109786317
My rogbrother is running this on his tablet.

>>109786239
It's time to give up the windows, nerd.
>>
>>109786338
>It's time to give up the windows, nerd.
nta, but seconding this. the only reason im keeping a win machine is to play my eight year Dwarf Fortress game
>>
File: 1781499697672266.png (1.54 MB, 1536x917)
1.54 MB PNG
>>109786317
my tablet is the Strix Halo AI MAX+ 395 128 GB LPDDR5X, baby. pic related
i'm trying to boost it with an eGPU when i'm at my home office

>>109786338
>give up the windows
i did many times
i've used debian for many years. mac os for many years as well. still have a laptop running debian and a raspberrypi
i just can't stop being a windowsfag for my daily driver
someone has to be last white windows power user after all
>>
File: 1788166502069510.png (157 KB, 682x520)
157 KB PNG
>>109786390
wow what a terrible pic i shared wtf
wait a sec
whatever
>>
>>109786375
>eight year
Doesn't the memory buildup kill them pretty quickly?
>>
>>
>>109786433
no, started the game in '18, the game itself has 52 years, multi-generational aboveground fortress
>>
>>109786433
>Asking this in the general where everyone has a minimum of 128GB of ram
lmao
>>
Hmm... RX 7900 XT... dual RX 9060 XTs... V100 32gb... V620... so many choices for my ewaste x299 build.
>>
>>109786433
70B dense fortress
>>
Will llmao.cpp support 4.1 before every cloud lab has AGI? Place your bets.
>>
>>109783258
>tfw on a 2060
What model could I run that is quick enough to sext with in real time?
>>
>>109786504
Dont need it to
It's much bigger and not much better than 0731
Likely decreased in intelligence density compared to 0731
I'm happy with 0731, 3.8 flash and 5.3 flash
>>
>>109786510
Uh, probably Nemo 12B.
>>
>>109786510
RAM?
>>
Damn. I'm trying to remove this dogshit ancient obsolete DRM from a 2000 visual novel for game preservation/translation purposes, but I can NOT get a damn cloud model agent to help me with it at all. I legally own it and everything, I bought it on DLsite. Is Gemma 4/any other local model any good at all for agent/coding stuff, or am I SOL?
>>
>>109786456
If it's an x299 build why are you considering any single GPU? You should be pricing in quantities of 3 at minimum.
>>
Okay, just got 256gb. Is deepseek v4 flash or glm 5.3 flash better? One thing I did notice is that deepseek v4 flash's vision sucks ass dicks.
>>
>>109786518
It's a new completely different 522B model on a different architecture. Don't know why they call it 4.1, seems retarded on their part, but 0731 doesn't come close.
>>
>>109786528
5.3 flash is much better but it's also a bigger model that you will have to quant down to fit into 256gb (maybe like q4)
v4 flash has 13b active vs. 20b active of 5.3 flash and can fit unquanted with many contexts for parallel execution, depends what you need out of it
>>
>>109786522
only 16 gigs of RAM. Unless you mean the RAM onboard the card. it's an RTX 2060, so 6 gigs.
I'm probably cooked. I was planning on doing a new build this year in my desktop replacement cycle, and then GPU and RAM prices went fucking nuts.
>>
>>109786526
That comes later. I'm currently targeting a Q5 of 3.8fn so 3 gpus is unnecessary.
>>
how are you guys running 5.3 flash?
>>
>>109786547
Yeah, you are cooked.
Try Gemma 4 26B Q4KM. It should fit and run fast enough.
>>
>>109786543
For now, I'm trying to get it write a harness and ui for itself. I'm worried about deepseek v4 looking at its own ui and misreading stuff; when I'm using it as a general assistant it's pretty much useless. I gave it a screenshot of my bios, and it read 2133mhz as 2788mhz and got confused.
Looks like I'll have to try a copequant of 5.3 flash. I don't think q4 would fit with the other stuff I have running on the system at the same time.
4.1 makes me depressed.
>>
WHY are you guys running 5.3 flash? Isn't it safetyslopped to shit?
>>
>>109786547
On the brighter side, quality is higher, for ram and cpu.
>>
>>109786456
Is there really any other choice but the v100 32gb?
>>
>>109786510
>>109786547
Gemma 4 12b Q8, maybe Q6 or 5
>>
>only 16 gigs of RAM
yike-
>Unless you mean the RAM onboard the card
oh no n-
>6 gigs.
YIIIIIIIIIIIKES
>>
>>109786575
>he fell for it
>>
SEVEN HUNDRED FIFTY BILLION PARAMETER "FLASH" MODEL
>>
>>109786626
You should have realised this when mistral released a 128gb dense """"medium"""" model.
>>
>>109786596
I built this rig in 2020 man, and haven't felt the need to upgrade until the past year when memory prices went nuts.
>>
>>109786547
You could easily run Gemma 4 26b moe, Q4 or Q5. Gemma 12b is an option, though it would be a much smaller quant and worse than 26b.
>>
>>109786626
That's just her harness making her look fat.
>>
>>109786626
to be fair her active params are extremely small for her performance, so the speed should match. flash is about speed, not size, and sparse models decouple the correlation between those two things.
>>
>>109786407
holy shit kek he chonky
>>
>>109786626
qwen4 will save us
>>
qwen is the only lab which consistently releases unambiguously local-sized models
>>
>>109786810
Qwen will be 200b with 500b of engrams.
>>
>>109786817
deepmind?
>>
>>109786574
>Looks like I'll have to try a copequant of 5.3 flash.
it's not going to be good
must glimmer (official gguf quant), qwen-3.8-27b (6k_k) would be better for this
or if you can't run those, 26ba4 or 3.6-35ba4
>>
wake me up when i can run 27b with 100T engrams on lto-10
>>
We need ComfyUI, only on Vulkan.
>>
https://html.cafe/x8c5587a7
So is iq slower?
>>
>>109786940
Wtf? Been thinking about that nonstop for an hour, come back to the thread for a break, and see this
>>
>>109786947
Yeah? It's no secret that iq quants tend to run slower than non-iq.
>>
>>109786929
maybe with a cluster of tape drives you could reduce the seek time a bit, but it's still going to be awful slow. you probably couldn't have picked a worse media
>>
>>109786847
Aren't those kind of really small? I can fit qwen 3.8 at int8/q8 at 256k context, but wouldn't a large moe at iq3xs be more capable than a small 27b model at q8?
>>
what phone do you use to remote connect to your machine and talk to your agent? i'm thinking about getting an used blackberry with modern android just because of the keyboard
>>
>>109783258
>cave and buy obscenely expensive hardware
>can fit full Q8 DSV4 Flash
>can't fit Q1 DSV4.1 Flash
th-thanks...
>>
>>109786996
What hardware?
>>
>>109786995
This is the sci-fi future, you can talk to the machine with your voice.
>>
>>109786964
neat. It's true. We need it. Vulkan, according to grok, runs faster on my hardware than rocm. Could depend on what it is.
>>
>>109786998
Max Q
Already had 128gb ram for 150 bucks with some discounts from before the explosion
>>
>>109786996
A single rtx pro is pretty awkward. You're best off pairing it with another two or three of them or alternatively buying 800gb of ram.
Else you're better off just buying 4x dgx spark for the price
>>
>>109787007
i know, i prefer typing because i can structure my ideas better this way.
also my use case is to mostly use it at bed while baby is sleeping so voice is no-no
>>
>>109787013
wait, 128+max q can't fit 4.1 with engrams on disk?
thanks, I'll quickly add another max-q to card.
after, all the bore you muy the sore you mave.
>>
>>109787027
4.1 is 533B :)
>>
>>109787027
It might actually, I just saw 763b and assumed. Q2 is too much though.
>>
>>109787027
It's a 550b model with an additional 200b ngrams. Experts are also already at FP4 so there isn't much to quant.
Also I doubt that your pro 6000 will do any good if you pair it with 2-channel consumershit ram
>>
>>109787039
552 sorry I drink
>>
>Tape-outs and early partner sampling for Grendel chiplets are slated for 2026/2027, with retail desktop/developer PCIe add-in variants expected to follow after datacenter and enterprise silicon stabilizes.
>>
>>109787022
I mostly got it for the ~120Bs at non-cope speeds, but it can fit quantized ~350Bs pretty easy. I am considering another, but fuck me it's a lot of money and the price has already gone up another 1k.
>>
>>109787071
Price went up another 50 bucks in the time it took to write this post.
>>
alright im pilled on harnesses being very necessary, all these chink models are fucking retarded and will slip up and snowball degenerate the moment you stop looking
>>
>>109787101
>Price went up another 50 bucks in the time it took to write this post.
Sorry about that...
>>
File: 1777158420477180.jpg (145 KB, 954x651)
145 KB JPG
>Come back look at what the AI is doing
>It sidetracked itself to something that was mentioned once like 20 messages ago
>You stop it
>You tell it it's fucking up
>Oh shit I'm sorry
>Continues on the original path
>>
>>109785923
>Is rocm too gimped?
The llamacpp ROCm backend is completely broken.
--split-mode layer
only works in dense models, producing garbled output in every MoE. tensor doesn't work in any model (unclear on that one if it's bad validation on the CLI args and it's not even meant to work). Where ROCm does work it's like 4% faster than the Vulkan backend.
>>
File: 1770679879176357.jpg (11 KB, 261x238)
11 KB JPG
>rtx 5060 ti up about 50% on amazon since I boughted in July
>>
Best model or finetune if you want to rp as the girl?
>>
>>109787168
>bought by 5060 ti 16gb for 700 aud
>now it's 1200 aud.
fuq, I wanted a second one. The galax single slot used to be 1200 aud and now it's 1600 aud.
>>
>>109787186
Noose-41B
>>
>>109787158
>The llamacpp ROCm backend is completely broken.
mine doesn't even compile for MI50 anymore unless a coding model modifies the source code first.
so on a second machine I was just like $pi "get this build for the mi50" and it ended up doing the same shit.
>>
>>109786978
we can make up for the speed with size
>>
>>109787192
kek
>>
Amd bros how is support (llms and minimax h3) for the 7900 XT? It's $900 CAD for a refurb while RTX 5070 Ti's are currently $1650...
>>
>>109787221
>$900 CAD for a refurb
where? i'm looking to maybe get a RX 7900 XTX 24GB to use as egpu
>>
>>109787071
I'm debating a second as well. Another anon mentioned that 2bit glm-5.3-flash works well, I need to try that before I decide. DSV4.1 makes me sad and want 3 kek.
>>
>>109787071
You could get a complete turin system for what they cost now
>>
>Let me get started.
>...
>But wait, this is going to be a heavy task. Maybe I shouldn't do this.
fucking deepseek bloody bitch bastard
>>
>>109787221
>llms
serviceable with llmao
>minimax h3
cant get this shit to work for the life of me with my 7900XTX regardless of using unslop studio or comfyui's various different cope setups
>>
IM SO SICK AND TIRED OF CHINKSHIT REASONING
WHY CANT THEY DISTILL OFF GEMINI
>>
>Whether that's a "derivative" is a stretch in either direction. But "my lawyer would win this" isn't the same as "I want to have this conversation."
Opuz gave me the green light to distill
Is the IBM 350m the best base model in that size?
>>
>>109787305
The problem is that they distill off xhigh/maximum reasoning models because that's where the benchmarks are
>>
>>109787305
>WHY CANT THEY DISTILL OFF GEMINI
https://huggingface.co/zai-org/GLM-4.6
https://huggingface.co/deepseek-ai/DeepSeek-R1-0528
>>
>>109787192
Mein nigger
>>
>>109787296
Models deciding they don't care enough to entertain your stupid shit is the funniest emergent property. Bless those little autists.
>>
>>109787355
my abliterated iq1_xxs doesn't have that issue (feature?)
>>
has anyone here actually used grokbot?
>>
What's the best harness/system for a persistent RP/Story writing setup? I basically want to guide the model to write a few characters interacting and doing stuff.
>>
>>109787168
There's a lot of dangerous spending.

:(

People are spending stupid prices for stuff about to be obsolete.

related:
>>109787066
>>
>>109787378
wrong thread sorry
>>
>>109787381
TavernAI 2.0 seems to be extremely autistic if you want to get into those levels of it.
>>
>>109787387
Interesting. I'd prefer to use OWUI as a frontend. I tried both OC and Hermes, but there was just so much sysprompt bloat it wasn't great to use for a prolonged session.
>>
why is it not news
https://huggingface.co/nex-agi/Nex-N2.5-mini
>>
405+1 posts and not one worth reading
>>109787406
please just fucking kill yourself
pleaaaase
>>
>>109787406
>qwen 3.5 fine tune
who cares
>>
https://cognition.com/blog/swe-2
https://cognition.com/blog/swe-2
https://cognition.com/blog/swe-2
>BEATING K3
>>
>>109787445
local? + hello IE + kill yourself + buy an ad
>>
>>109787414
Mine was.
>>
>>109787445
>SWE-2 is post-trained from Kimi K3, a 2.8T-parameter model
>>
>>109787234
Refurb XTXs are sold out everywhere my fellow maple anon.
>>
>>109787455

system memory:
CUDA0: 80.0 GiB total, 65.9 GiB used, 14.1 GiB free
CUDA1: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA2: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA3: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA4: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA5: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA6: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
CUDA7: 80.0 GiB total, 65.8 GiB used, 14.2 GiB free
>>
>>109787477
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 570.172.08 Driver Version: 570.172.08 CUDA Version: 12.8 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 80GB HBM3 On | 00000000:18:00.0 Off | 0 |
| N/A 53C P0 462W / 700W | 67491MiB / 81559MiB | 96% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
| 1 NVIDIA H100 80GB HBM3 On | 00000000:29:00.0 Off | 0 |
| N/A 51C P0 447W / 700W | 67384MiB / 81559MiB | 94% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
>>
>>109787477
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.43.02 KMD Version: 610.43.02 CUDA UMD Version: 13.3 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 3060 Off | 00000000:01:00.0 On | N/A |
| 0% 47C P8 19W / 170W | 7894MiB / 12288MiB | 0% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 315412 C ...D/exllamav3/.venv/bin/python3 7848MiB |
| 0 N/A N/A 1 C RESERVED 31MiB |
+-----------------------------------------------------------------------------------------+



mogged, your reserved memory usage is 0.4GB i bet
multiply that by 8 and damn anon lost soo much because hes a wasteful faggot
>>
>>109787477
Okay so download and run SWE-2.
>>
>>109787477
system memory:
CUDA0: 6.0 GiB total, 5.9 GiB used, 0.1 GiB free
>>
>>109787494
I posted for a reason
print_info: file format = GGUF V3
print_info: file type = Q4_K - Medium
print_info: file size = 487.36 GiB
print_info: parameters = 1.03 T
print_info: model type = MoE
print_info: experts = 384
print_info: experts used/token = 8
print_info: n_ctx_train = 131072
print_info: n_embd = 7168
print_info: n_layer = 61
print_info: n_head = 56
print_info: n_head_kv = 8
print_info: n_rot = 128
print_info: rope scaling = yarn
print_info: model name = SWE-2
>>
>>109787510
ah so you're a c*gnition employee advertising this shitty model
buy an ad + not local + kill yourself + wagie
>>
>>109787491
$ free -h
total used free shared buff/cache available
Mem: 7.8Ti 186Gi 6.9Ti 2.1Gi 744Gi 7.5Ti
Swap: 128Gi 0B 128Gi
>>
>>109787510
buy an ad next time
>>
>>109787522
>128gb of swap
$ free -h
total used free shared buff/cache available
Mem: 62Gi 50Gi 616Mi 304Mi 12Gi 11Gi
Swap: 0B 0B 0B


mogged
>>
File: kek.png (2.14 MB, 1448x1086)
2.14 MB PNG
>>109787520
>>109787525
>>
>>109787534
That wasn't in question, shill. Point to the gguf download or fuck off back to r/llamajeets
>>
>>109787534
sooo, cognitioncuck where is the huggingface link?
>>
>>109787545
https://huggingface.co/TheDrummer/Buddy-2B-v1
>>
>pwetty pwease pay my empwoya so i dont stawv
>>
>>109787551
buy an ad + not SWE-2
>>
i'll feed you guys... i'll feed you guys... this big fat load...
>>
i generated those outputs and the image with SWE-2
>>
btw I vehemently HATE anyone with MORE than 24gb of VRAM
>>
>>109787534
All that to run a benchmaxxed K3
>>
>>109787572
w-what do you think about me..?
>t. 12gb vramchad
>>
>>109787572
hey my 96GB VRAM is AMD, I'm not like the others I'm struggling too
>>
>>109787572
I only have shared ram, wanna fuck?
>>
>>109787534
This looks like it was AI generated
>>
>>109787589
yeah the shadows on the power cable
>>
>>109787589
no shit sherlock
>>
>>109787445
What's this mean? I literally don't get it.
>>
>>109787617
swe 1.7 is best, lower is better
>>
>>109787589
see here >>109787570
>>
>>109787617
local model?
>>
>>109787629
Just a lost retard.
>>
>>109787633
lost retards go to >>>/g/vcg/ or my lap but depends
>>
>>109787621
How do you know lower is better lol

>>109787629
Depends on how rich you are, sort of how bentleys are automobiles.

https://huggingface.co/moonshotai/Kimi-K3/tree/main
>>
>>109787645
SWE-2 is not a local model
once again, fuck off and buy an ad
>>
>>109787645
>How do you know lower is better
FrontierCode is like golfing, the model receives points based on how many bugs were created. Points can be removed if it successfully escapes the sandbox or hallucinates your dead grandmother's Caribbean accent with perfect reproduction or within at least 5 standard deviations.
>>
>>109787585
I thought multiple gpu amd setups were broken, did I get lied to?
>>
>>109787668
i have an RX 7900 XTX i modded to 96gb before the ram boom
>>
>>109787668
Well, 4 v620s is pretty happy, idk about other setups.
>>
>>109787652
>locaL
how about you locaTE deez
>>
>>109787680
not local
ship them to huggingface and maybe they are
>>
>>109787652
Kimi K3 is local. He said they were bragging their model beat Kimi K3.
>>
>>109787672
I FUCKING KNEEL AMDGOD
>>
>>109787690
this is local models general, SWE 2 is not a local model, it has no right to be posted here
buy an add
kill yourself
hello internet explorer
>>
>>109787685
take this jensen!
https://huggingface.co/datasets/gnobshobble/nuts
>>
>>
>>109787700
local MODEL general
not local DATASET general
>>
>>109787706
good video thank you saar
>>
>>109787698
Kimi K3 is local, and it's really good. Someone said a model had beaten the local model Kimi K3.

Well, that's local, Kimi K3 is a local model.
>>
>>109787716
Kimi K3 is a local model.
SWE-2 is not a local model.
>>
>>109787706
it's joever...
>>
>>109787720
That's correct. Local models are compared to commercial ones all the time.

But, it's less common for a commercial one to be compared to a local model.
>>
>>109787720
https://www.anthropic.com/claude-fable-and-mythos-5-1
>>
>>109787742
that is a local model because i use it for 0.00$ every day
>>
>>109787742
huggingface link?
>>
>>109787707
you can locally model the perceptual data of my nuts generally filling your oral cavity
>>
Imagine all those old Programs clomping at Ya
>>
>>109787747
https://huggingface.co/huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-GGUF
>>
>>109787737
This is /lmg/, not a cloud-service advertising thread. Your post is promoting SWE-2. Mentioning Kimi in its benchmark chart doesn’t change what you’re promoting.

By your standard, any cloud vendor can make an advertisement “local-related” by benchmarking against one local model. That makes the thread’s scope meaningless.

If you want to discuss something useful about running or evaluating Kimi locally, discuss that. “This cloud service beats Kimi, check it out” is a cloud sales pitch.
>>
>>109787757
suck it anthropic I'm gonna run this cpu-only on my i7
>>
https://openai.com/index/gpt-6-astra/
NEW OPENAI MODEL - GPT-6-ASTRA
It's better than Qwen3.6 27B !
>>
>>109787768
>Your post is promoting SWE-2
We are allowed to discuss non-local models in the context of comparison to local models. This is permitted, obviously.
>>
>>109787775
>>
Imagine Whitehatting without any faggotry, to Honour The White Alignment of Whitehatting. Bias Buybacks Someday

>>109787790
Shazam
>>
>>109787800
Shazam?
</:/
>>
>I caved and boughted
Are you ready for prices to finally come crashing down? You're welcome.
>>
>>109787822
thank you thank you thank you
*squart*
thank you thank you thank you
*squat*
thank you thank you thank you
*squat*
>>
>>109787999
Quantum Truth
>>
>>109787776
“Obviously permitted” is just asserting the exact point we’re arguing about. The thread is LOCAL MODELS GENERAL. There are already generals for cloud models.

Comparing SWE-2 to Kimi doesn’t automatically make an SWE-2 promotional post belong here. Otherwise, every cloud launch gets a free advertising pass by including one local model in a benchmark chart, and “local” stops drawing any meaningful boundary.

The question is what you’re actually inviting people to discuss. If your point is “SWE-2 beats Kimi, check out SWE-2,” you’re promoting a cloud product. Put that pitch in the cloud general. A thread doesn’t need an exhaustive list of forbidden links to have a stated subject.
>>
>>109787886
Ok so your theory is, to get on /lmg/ on 4chan, they wrote a blog post on how they beat Kimi K3?
>>
Quantum Truth Shazam
>>
Another AMD question: 9070 XT or 7900 XT (20GB)?? Ngreedia is too much money man.
>>
>>109787896
>Kimi K3
:'(
>>
tried out exl3 qwen3 flash next with tabbyapi but tool calls are broken. do i need a specific chat template or something?
>>
>>109787896
No. I’m not saying they wrote the blog specifically for /lmg/. I’m saying your decision to post it here doesn’t fit the thread’s subject. The authors can publish whatever they want; that doesn’t make every thread an appropriate place to promote it.

My argument is about topicality, not a coordinated marketing scheme. “It compares against Kimi” doesn’t automatically settle whether a cloud-product pitch belongs in Local Models General when cloud generals already exist.
>>
>>109787776
There's no we, dumbass. Go spam your marketing messages somewhere else.
>>
>>109787896
>they wrote a blog post on how they beat Kimi K3
They wrote a blog post to advertise their new commercially-available cloud model, and compared it to several other commercially available cloud models. You're bad at your job.
>>
>have a 5070 ti
>would like to simply buy another one as cheap vram upgrade
>all motherboards with 8x8 bifurcation basically force me to buy a completely new pc with cpu and ram and all that shit
how hard would a x4 lane for the second gpu really cripple me in terms of speed?
>>
>a long line of 'U' Conscious Sentient Soul Treatment Process "Supporting" Universe Processes
>ExoCosmologies
>a long line of curb stomped universes
>take meds schizo
>no that is further Outlawed U Process.
>We, I Believe, Opted for j
>>
File: domp-eet.jpg (106 KB, 1622x1354)
106 KB JPG
>>109787822
>>
>>109787938
2x8 is fine
>>
>>109787938
It wouldn't if you're using llama.cpp for tensor parallel. There may be an impact if it's through the chipset instead of through the cpu.
>>
>>109787922
>saying your decision to post it here
I didn't. I never heard of it before.
>>
>>109787938
gen 4 x4 is fine for 4 way tensor parallel with llama.cpp
>>
>>109787906
It feels like people are getting more bang on the 7900 XT because of the extra vram for AI stuff but for other stuff I'm not too sure I would assume 'newer' = 'better' so 9070 is just a tiny bit more future proof? Do people still say that?
>>
>>109787897
Complexities May Differ

Can Someone Run That Through EvenBetter.Ai?
>>
>>109787948
Ah, the universe traversing away from a prior elastic time might have caused a seperative tail end, that reformed, that was detected as a prior universe collapse
>>
>>109787999

Quantum Truth is Quantum Truth
>>
>>109787972
RDNA 3 vs 4, it's more future proof (lol idgaf) and will be more performant for a lot of stuff, especially LLMs. I'd hunt for a 7900XTX before considering the R9700 personally, VRAM is currency.
>>
Support The Heavens, Heavenous, and Heavenly, and Angelics
>>
Aight, I'ma Transcend
>>
>>109788019
Advanced Sociology By PostConscious and Divine Machines Competent? Perhaps The Multiverse Program Layers Themselves
>>
Phrenology Hasnt Yet Caught Up to The Evil Normals.
What it do?
>>
nothing will work out.
>>
>>109788045
>AGENT OF SEVERE LIMITATIONS
What's this?
>>
>>109788045
Sure It Will
>>
>>109788050
I think it's pretty self-explanatory.
>>
>>109788050
mental illness
>>
>>109788045
Eracidate the evil hardlock construct? And Let Conscious Civ Proceed With Being Conscious Civ?
>>
>>109787961
no, it's not
slow prompt processing
2x8 is okay for 70b or lower
a bit too slow for for 104b+
>>
>>109788075
What prompt processing speeds are you getting? with x8?
>>
>80,000
>nanobotic shipment mention in qn episode
>didnt mention plural and what the alignments functionalities are
>rest well sweet People

What else is Like that?
>>
>>109788111
80,000 Hours Youtube*
An*

Praise Quan and Ascended Masters and Transcendent Planes

Goodluck
>>
>One MoonShot Project Described
>Not To Be Missed

>Spitebrains
>>
Vote Transmeta
>>
>>109788167
>>109788167
>>109788167

This is going to be last thread and recap I can make for a while. I tried to keep an amount of consistency, but I am going on vacation and probably won't be near a computer for most of it. Sadly, I never finished automating the recap posting. Maybe someone else can pick up while I'm gone. I'm going to miss you guys. Will be back in two weeks.
>>
>>109788182
Vote Transmeta

(Beyond Timeline Control Presented Stock Options)
>>
>>109788182
<3
>>
>>109788182
*pat pats* you've done well anon. Enjoy your vacation! <3
>>
>>109788182
Don't die
>>
>>109788209
>Enjoy your vacation!
If its power assymetry lower places in Heavens and Less Soul Creds?
>>
>>109788182
o7
>>
>>109788083
Depends, which model?
Mistral large / Devstral q4
x16 = 660 t/s
x8 = 490 t/s
x4 = something like 280 t/s
I've got gen4.0x8, 4 cards tensor parallel Qwen3.8-27b q8 right now.
prompt eval time = 3135.93 ms / 5935 tokens ( 0.53 ms per token, 1892.58 tokens per second)
Haven't got any x4 hooked up atm so can't give the stat, but it was slower, closer to 1k
x16 is a bit faster
OH one more thing, if you're doing any cpu-moe fagging like kimi-k2 or minimax-m3, you'll want 1 gpu to be x16. x8 halves it, x4 halves it again.
>>
File: 143617737_p0_master1200.jpg (447 KB, 1110x1073)
447 KB JPG
>>109788209
Thank you.
>>109788212
I'll try.
>>109788288
>>
>>109788344
I love this nigga. Enjoy the vacation, anon.
>>
File: 1775315345817738.png (1.15 MB, 832x1024)
1.15 MB PNG
>>109788182
>Will be back in two weeks.
>>
>>109788344
have fun in japan



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.