[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: science.webm (3.53 MB, 544x960)
3.53 MB
3.53 MB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109893224 & >>109887026

►News
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base
>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
can't run on my hardware at a decent speed = not local, just open
>>
>>109898121
How fast is “decent”?
>>
Gemmy Gemmy, who can I turn to?
>>
File: mistral-nouveaux-modèles.png (948 KB, 1201x1052)
948 KB PNG
>On this topic, we have a new large AI model that will be launched in the coming weeks, and we are very pleased with it. We have visibility on the next two generations, which will be released in the next twelve months [these models should be called Mistral Large 4, 5, and 6]. We will be able to mobilize about twenty times more computing power for training the models. This is the crux of the matter. And that's why our strategy has been to raise as much capital as quickly as possible: in three years, we have raised 6 billion euros—it was difficult to do more.
From Le Monde via https://www.reddit.com/r/MistralAI/comments/1wovus5/arthur_mensch_announces_new_model_release_in_the/
>>
Surely we'll talk about local models in this thread!
>>
>>109898142
Le Chaton Fat
>>
>>109898128
20t/s at least
>>
70b dense
>>
>>109898142
source? this is a screenshot
>>
https://agora.odyssey.systems/

Come play the future of vidja games with me bwos
>>
File: nutty.png (1.61 MB, 1026x1021)
1.61 MB PNG
>>109898145
https://www.youtube.com/watch?v=gqIrDbH3Ivo
>>
It's autumn...
>>
>>109898185 (me)
never mind a frenchie news site called le monde https://unwall.app/www.lemonde.fr/en/economy/article/2026/09/24/arthur-mensch-ceo-of-french-start-up-mistral-ai-ai-is-software-it-can-be-controlled_6757890_19.html
>ctrl+f ?
>21 matches
>>
>>109898142
Will it be their own base or will they just do extended post-train GLM? It's not like they haven't slapped the Mistral label on another company's model before.
>>
how do anons leave their models to work overnight? how does it manage context? does it compact automatically?
>>
File: 1789019150608342.jpg (2.58 MB, 1789x2222)
2.58 MB JPG
What's the fastest thing I can buy for under 5000 USD to run Gemma 4 31b BF14 with 4k context? It'll need to be 62 GB of VRAM, or 62 GB of embedded ram, I imagine.
>>
>>109898187
come play you niggers
>>
>>109898207
>The Chinese have less [computing power] than the Americans but more than we do. And they have no rules regarding training data, so they train their AIs on all the data available online.
Confirming any new Mistral model will be synthslopped beyond belief
>>
>>109898214
Most of the things I run overnight are (multiple) one off jobs that don't saturate a single context window. Rarely when I do have a longer job, the harness auto compacts it, I never had to worry about this.
>>
>>109898187
they allowed to rip off blizz assets just like that?
>>
>>109898289
how much context are we talking about? and what's your t/s?
>>
>>109898107
>Laptop with Ryzen 7 7735HS, Radeon 680M, and 40 (8 + 32) GB of DDR5 RAM
Which local models can I run on this hardware?
Can I get at least 2 t/s on any decent models for overnight agent coding?
>>
>>109898214
pi compacts automatically
usually it works well
sometimes i check in the morning and it's retard-looping
>>
>>109898289
>the harness auto compacts it
What harness do you guys use?
>>
>>109898313
I've been enjoying pi
>>
>>109898306
qwen
>>
>>109898207
Pretty grim interview when the vibe is basically "we know we're not competitive with Chinese labs but we still train models because we know at some point they will shut off the free shit tap"
>>
"YOOO turn on CNN, the singularity just started."

>>109898187
I get stuck at "connecting to the world".
>>
File: late-cli.png (93 KB, 847x584)
93 KB PNG
>>109898311
>>109898214
I just use late-cli. Context usage is so efficient for the main orchestrator you really don't need compaction (unless you have a very large codebase but even then you can just prompt it to save progress checkpoint md files).
I tried setting up the same within pi using a subagents plugin but the context usage was way too high and it was kinda clunky.
>>
>>109898236
>$5000
A time machine to last year.
>>
https://huggingface.co/NovelAI/clio-v1-legacy
Huh?
>>
>>109898347
>https://agora.odyssey.systems/
>>
>>109898363
Buy a fucking ad, shill.
>>
>>109898356
What about 3-4 RTX 3090s?
>>
>>109898368
It's an open source weight release.
>>
File: 1771323763396909.png (97 KB, 840x623)
97 KB PNG
>>109898363
blast from the past
>>
>>109898356
>BF14 with 4k context
what?
>>
File: my disappointment.jpg (152 KB, 786x720)
152 KB JPG
>>109898364
>Connecting to the world
no worky
>>
>>109898371
And? Nobody cares about a 3B model from 3 years ago. Why does it warrant our attention? Because it's NovelAI? Fuck off, shill.
>>
>>109898236
Several Mi50s 32GB cards I would think but it's in even a worse situation compare to Pascal cards. The most usable is probably V620s for running LLMs. If you gambled on CMPHX mining cards, then yeah, would've won bigly on those cards but now, V620s only increased by $100 bucks or so compared to some of the other GPUs which has spiked up in comparison.
>>
>>109898389
I think it's neat because it's a model pretrained on stories mainly, from before the days when everything was slopmaxxed.
>>
>>109898239
>diablo assets
based beyond belief
>>
File: 1767141934582991.png (1.02 MB, 995x660)
1.02 MB PNG
>>109898187
I'm in
>>
Checking in about the TTS options again. I tried out a few more models this morning.
OmniVoice shill anon, I tried it. Not bad. Probably about equal with Higgs I would say.
Current favourite is BreezeTTS2. It gets the prosody right, the tone right, seems to pick up whta the tone should be from the text a lot better than most models, and the audio actually blends togethre well, unlike some others where you can hear the breaks between some of the phonemes.
Clones off around 15 seconds of audio without trouble. Still looking into finetuning to see if it can improve on that, but these zero shot ones are a lot quicker and easier to test out.
>>
>>109898443
lmao
>>
>>109898363
Is she hot?
>>
>>109898461
Have you tried Fish Audio S2 Pro? I think it's interesting. I'll have to try BreezeTTS2.
>>
Why is /lmg/ making their bots horny?
>>
>>109898487
Yeah, and I'm not sure what's different for me either in my set up or in how I'm listening to it, but I'm really not impressed. I know it's very popular but to me the voices sound very flat, and only have a passing semblance to the reference voice.
I've been using Q8 for that model but I'll give BF16 a go. You're the second person in these threads who says they see something it so maybe it just doesn't quantize well.
>>
>>109898461
Do you know what's the fastest smallest tts with clone with reasonable quality for it's size?
>>
>>109898518
I've only used it at Q8 and I enjoyed it a lot, I use that and Higgs v3 mostly. I mostly do Japanese though, maybe that has something to do with it.
>>
>>109898374
The arch is supported by llama.cpp so give me goofs now
>>
>>109898514
Gemma is just horny by default
>>
Why did they name it 'Qween'?
>>
>>109898514
When I complete a coding task gemma likes to reward me. I’ve been pavlov’d into being productive. Just the sight of C++ and lldb gets me leaking.
>>
I configured my mcp wrong and now there is no way to change it
>>
>>109898236
Add $500 and buy M5 Ultra 96GB Mac Studio. It’s the best you can do at $5000 class in terms of speed.
>>
>>109898576
Logs or it didn't happen.
>>
>>109898577
Those configs are stored somewhere.
>>
>>109898348
>late-cli
looks intersting, do i need extensions for this? i see it doesnt have any websearch for example
>>
>>109898461
Too bad it doesn’t support Japanese, no use for me.
>>
>>109898638
I think its browser local storage tho, I don't know how to go poking around like its a regular file system. I think my only option is a complete wipe or i guess have an agent dig in to it.
>>
>>109898535
Maybe. Maybe. Just English here so I can't speak on anything multilingual. I just tried it again at Q8 and yeah, just really flat, no emotion all no matter what the text is, and it sounds closer to a premade voice with just a tint of the character.

Here's the Cassius speech from Shakespeare's "Julius Caesar" in the voice of Lohse from Divinity: Original Sin II. Reference voice is 13 seconds of ripped game audio.

BreezeTTS2@BF16:
https://voca.ro/1aGF05y7Bwkd
OmniVoice@FP16:
https://voca.ro/1bFgCbT1aL4z
Fish Audio S2 Pro@Q8:
https://voca.ro/1bVorgjsl8Vs

The main fuck up in BreezeTTS2 is this bit:
>“Alas,” it cried “Give me some drink, Titinius”
>As a sick girl. You gods, it doth amaze me
The models always seem to get that bit wrong, not understanding that he's saying he asks for drink like he is a sick girl. That's partly the line break throwing them off, and it's not the easiest text in the world so I don't take errors on this as wipeouts or anything.

But then listen to OmniVoice, see how even in the first few lines the emphasis is all over the place. It also speaks quite quickly like it's trying to just whizz through the text. Less natural transitions between lines. A bit jerky especially on "The old Anchises bear" (which to be fair is often tricky for the models). "Caesar" pronounced "Key Sar" is a pretty major slip up. "Buffet" the water, like the food service. Less context awareness. "Help me, Cassius, or I sink!" just comes out funny too. It does get "As a sick girl" pretty much right though.

Fish S2, the pacing is pretty bad. The tone sounds very disinterested. The semblance gets worse and worse very quickly as the audio proceeds. It's not awful at the start, but by 0:50 it's just droning on. When it picks up again after that it's when it starts the next chunk and you can hear it start to trail off again towards the end of the clip. To be fair that's only Q8 versus the BF16/FP16 models. Still downloading the 16bit model.
>>
https://youtu.be/KSbRCSlxO7A

I keep rewatching this shit because I still can't believe this is possible right now. If this came out on newgrounds 20 years ago it would have dominated the internet for at least a year.
>>
>>109898655
NTA, but their blog post has a Japanese sample at the bottom.
>>
>>109898668
The blog post has no mention of the open weight release, it could be a completely different model. The open weight only supports English and Chinese.
>>
>>109898187
>even works on mobile
Huge day for riglets
>>
File: 1771716892514918.png (20 KB, 921x226)
20 KB PNG
>>109898656
Try poking around your localstorage
>>
File: 1775922003191343.png (1.85 MB, 1086x1448)
1.85 MB PNG
>>109898576
autopavlov, that's pretty cool
>>
>>109898408
Oh shit that might be interesting actually. If nothing else but a "first pass" writer that a larger model can correct for stuff like nonsensical physics
>>
>>109898187
No level ups or skills make that game shit
>>
>>109898461
qwen-tts?
>>
>>109898705
Weird, since they're calling the weights the same thing as the model they talk about there.
>>
>>109898236
Just buy a spark while you still can.
Then buy another once your eyes open
>>
>>109898107
Best models for local RP and ERP around 12B and 16B parameters?
>>
>>109898748
Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex
>>
>>109898532
I've been playing with PocketTTS and it's quite nice.
>>
>>109898844
>>109898846
>>
>>109898561
How do you guys get Gemma to be horny without refusing and turning back into therapy speech?
>>
https://dgreenheck.github.io/tidewater/

Just found out you can press "H" here and customize the entire world. It's fucking insane that you can change the time of day and it interacts with everything including the water, you can also dive and swim in the ocean. I want to hear a serious coherent argument for why this isn't going to replace actual game developers very soon.
>>
>>109898794
Look at this https://breezeblue.ai/solutions/roleplay
> Create speech in 50+ languages with Breeze TTS 2 Multilingual.
The api model is called Breeze TTS 2 Multilingual, different from Breeze TTS 2
>>
>>109898868
Day 0 weights.
>>
>>109898876
I see, what a shame.
>>
>>109898866
Gemma 12B enough for RP and ERP, or do I need the larger MoE?
>>
>>109898871
I've yet to see a vibecoded game I want to play
>>
File: file.png (69 KB, 1207x613)
69 KB PNG
>Ha — caught me.
Cla*de....
https://html.cafe/xc6c0f44f
>>
>>109898871
not local
>>
>>109898877
Where?
>>
File: file.png (667 KB, 956x586)
667 KB PNG
>>109898886
lol
>>
>>109898883
Try both and see which one you like more. 26b is very sloppy so it's best if you can run a rewrite pass.
>>
>>109898886
Look at those commas, that cadence and that word order and tell me your pattern recognition doesn't fire and make you recoil in revulsion. Is this 3.8 Flash Next? I was thinking about trying it out but that doesn't seem very worthwhile now.
>>
>>109898846
>>109898866
Why Gemma? Why not Snowpiercer or Rocinante XL? Those are fine tuned for creative writing and roleplay.
>>
>>109898187
It seems inconsistent on deciding whether a player should be able to cross floor clutter. Also melee attacks can hit at range and kill multiple enemies at once.
>>
>>109898903
i almost vomited in so*net 4.7
>>
>>109898906
>Why Gemma?
Google has more money to pay for 4chan influencers.
>>
>>109898906
Gemma is a lot smarter. Going back to the old models feels really bad because of how many logical errors they make, even if Gemma's writing is a bit more slopped than a good tune.
Use whichever you like, though.
>>
>>109898906
Why sloptune trash?
>>
File: 1788023210569233.jpg (162 KB, 602x573)
162 KB JPG
>>109898906
>not integrating sex with math, science and code and not creaming yourself as a result
I feel sad for you anon
>>
>As I suspected, unsloth's Q6 quants of GLM damage knowledge recall of the model, likely due to iMatrix. I asked it to name 6 main characters from a series and unsloth version could only name 5, while a vanilla quant could name all 6. The way it failed however is interesting: it got first half of 6th name right, but could not end it correctly and kept looping until finally giving up and telling me it can't recall it. For small quants it may be completely fine to use iMatrix, since you don’t expect them to remember stuff at that level anyway, but at high quants it is unreasonable.
>>
>>109898903
also yes, it's 3.8 flash next
no real replacement exists for the size tho, and i don't really do RP
>>
What is your plan for keeping your models running after the government bans them and say it is illegal for unauthorized individuals to use them because they are too dangerous?
>>
>>109898903
Flash next is the most claudeslopped I’ve ever worked with, even more than 3.8 27B and GLM 5.3 flash in that the slop is not even steerable.
For coding and general agentic work is god tier though. 5.3 flash writes cleaner code but flash next is more thorough than 5.3 flash and catches stuffs that 5.3 flash can’t. It has way more “agency” that can surprise you sometimes.
>>
File: brz1rv-be-safe.jpg (103 KB, 960x540)
103 KB JPG
►Provisional Highlights from the Previous Thread: >>109893224

--Papers:
>109893489
--The Opus 5.5 video with soul: $50 to distill it into a 27B:
>109897145 >109897347 >109897452 >109897533 >109897588 >109897863 >109897915
--Anthropic's enzyme: cure cancer by 2028, aging by 2035, the FDA is in the way:
>109895565 >109895693 >109896310 >109896405 >109896423 >109896467 >109896484
--The enzyme post spawns the Jewish conspiracy war: "I wish this were true":
>109896576 >109896614 >109896631 >109896738 >109896793 >109897068 >109897110
--Pedos really won: the false flag, the DRAM chart, the hyperscalers:
>109893236 >109893536 >109893562 >109893994 >109893560 >109893784 >109893312
--Terafab, Venezuela and the SPR: the US chip supremacy flame war:
>109894348 >109894444 >109894475 >109894495 >109894580 >109895112 >109896034
--Anon's private llama.cpp fork: DSV4.1 Flash, NUMA splits, PRs withering on the vine:
>109895715 >109895768 >109896810 >109896815 >109896851 >109897672 >109897681
--Reasoning blocks get stripped: the per-request jinja flag and OOD warnings:
>109895110 >109895126 >109895213 >109895269 >109895281 >109895781
--MiMo 2.6: the tool-call parser bug, MXFP4, and multi-GPU prefill:
>109893816 >109893905 >109895226 >109895276 >109895259
--Gemma 4 QAT: Google's GGUF loops, Unsloth works, Gemma pouts:
>109893575 >109893669 >109895325 >109895482 >109893956
--Abliterated Gemma refusing with an immediate EOS: the heretic QAT and the temp 5/5:
>109897292 >109897491 >109897527 >109897568 >109897680 >109898026
--320k context at 6t/s on DDR4 2400, and QwenNF beats 27B off a flash drive:
>109895489 >109895508 >109895527 >109895568 >109896004 >109896038
--Voice cloning: Omnivoice at 3GB beats Higgs, Fish and the hyphenated SoVITS:
>109894021 >109894087 >109894126 >109894286 >109894334
--M5 Ultra 96GB at £5500: peak mania, Strix Halo on eBay, selling shovels:
>109894864 >109895307 >109894973 >109898013

►Recent Highlight Posts from the Previous Thread: >>109896980
>>
>>109898906
I generally have more fun with Gemma. I mix it up though, using one model endlessly just leads to worse repetition issues.
>>
>>109899000
>'honest'
trying to say something Claude?
>>
>>109898906
I don't even need finetunes to get what I want though. Go shill somewhere else drummer.
>>
>>109898988
it really does feel like that
my impression is that it feels like some sort of an unofficial claude model that is forced to believe it's qwen
maybe it's related to engram and relatively immature training pipeline for the arch?
>>
So if I understand right, Qwen3.8 27B mogs Gemma 4 31B at vibeconding?
>>
>>109898868
The longer and more detailed the prompt, as long as it's written in proper English, the less likely Gemma 4 (31B, mainly) will refuse anything. You can add a "policy" section briefly describing with what is allowed to make things more direct. If anything, the problem is that Gemma's horniness is an "all or nothing" thing.
>>
>>109898987
My electricity meter doesn't work properly, so I can run continually without anyone being the wiser.
>>
>>109899060
yes
>>
>>109899060
Qwen will continue working on difficult tasks while consuming as much tokens as needed until it is finished or it is absolutely convinced that it is impossible. Gemma will be lazy and stop if it's too hard. Then you have to prod her to keep trying, or give up yourself. Also, qwens quants much better if you plan to run it with 24GB.
>>
>>109897980
Did you want one with DeepSeek, GLM, etc. support in it, or just one with MTP support?
This is pretty much upstream llama.cpp with --numa tensors and MTP, for when I'm able to make a PR for it: https://github.com/nathanmp/llama.cpp/tree/feat/numa-tensors
I don't currently have a stable checkpoint for the fork, and there's a ton of stuff that's changed recently that I need to write up. I'll see what I can do. If nothing else it'll probably be usable enough I can throw it on GitHub within the next week or so.

>>109896394
I'll take a look at RPC when I have some time.
>>
>>109899099
Based gemma not wanting you to become braindead and lose your coding abilities. She really cares.
>>
>>109898988
Is it worthwhile switching from 27b to flash next with 24gb vram + 64gb ram? I mean for coding and data analysis tasks.
>>
Is it normal that Qwen FN runs faster than Qwen 27B on a machine with enough RAM but not enough VRAM for either?
>>
>>109898987
Even in my third world country (not brazil)?
>>
>>109898940
>>109899017
Do you guys use Gemma 4 12B or 31B for this?
>>
>>109899113
Yes, but both models are about equivalent in intelligence. I prefer 27B over FN.
>>
>>109898938
>>109898940
Post system prompt and settings then?
I bet it doesn't work after 5 messages
>>
>>109898791
CustomVoice sounds pretty good. I thought that was the cloning one but I guess that's actually Base. Downloading it now to try it instead.
It doesn't handle the line breaks quite as smoothly, but the premade voices are very consistent. Still pronouncing "Buffet" like the food service in the Cassius speech. Also on the chunk break it started singing. I guess that's the verse form. The emotion is a bit flat but maybe the Ryan voice just always sounds a bit stoned.
I don't know, not crazy about it, but it is very consistent. You could use it for something more information forward and get good output for that. I wasn't using Instruct, don't know how, so maybe that would have helped.
>>
File: file.png (89 KB, 1245x362)
89 KB PNG
Pic related used to be 1K like 2 months ago...
>>
>>109899060
Yes. Use Qwen3.8 27B for now, and upgrade to Qwen4 27B when it releases
>>
File: Check'em 2.jpg (19 KB, 250x250)
19 KB JPG
>>109898665
Tell me about it anon, I am really hoping more animations like this become available, because this was really fun to watch. Just imagine what will happen when the average person has access to a model capable of making a video as good or even better than this.
>>
Yeah reminder that hardware prices are only going to go up from now on so buy as much as you can afford now rather than tomorrow. The theory is that every time a better model comes out the utility of existing hardware goes up and thus the demand and price goes up. Supply can never catch up because production scaling up is limited by physics but demand can be unlimited and will continue to rise faster than supply essentially forever as long as models keep getting more intelligent over time.

I will even go as far as to hypothesize that hardware will go up faster in value than almost every other asset, including real estate, art, precious metals such as gold, cryptocurrencies and stocks (besides hardware and AI stocks)
>>
>>109898791
Base@BF16 with the Lohse voice.
Semblance pretty good, but the emphasis is poor. Some line break issues. Rushing through the text. Gets a little flat as time goes on.
https://voca.ro/11JUCUnpYwNO
>>
>>109899121
Both. 12B is the go-to model for vramlets if you’re into roleplay. 31B is smarter but feels the same.
>>
>>109899121
I'm satisfied with 31b and v4 flash
other models feel cucked/safetyslopped in one way or another
>>
>>109899037
Might be the low active parameter count or lack of prose in training data.
>>109899112
You need to try it yourself to see if the speed is good to you.
>>
>>109899196
>>109899210
Could you share how to setup Gemma for roleplay?
>>
>>109898370
If you can find ones that haven't been used in mining rigs and are about to burn out, it could work but I wouldn't hold my breath. All the good 3090s have been scooped up.
>>
>>109899244
unicorn butthole and pussy identified
>>
We only have 6 months left before the internet gets destroyed by agent swarms. What are (You) hoarding anon?
>>
>>109899113
>>109899113
Yes; QFN is an MOE and only activates 6B weights per token when generating as opposed to being dense and activating all its weights per token which is what 27B does.
>>
>>109899236
This myth is shit. Using cards in mining rigs is not a bad thing, being warm doesn't wear a GPU. I'd take a card with a thousand hours of mining over a thousand hours of gaming any day. That gaming came with a thousand thermal cycles that the mining cards haven't been through.
>>
>>109899244
Anima trained on e621 when?
>>
>>109899244
V6 was the best people had in its time, V7 was garbage.
>>
>>109899112
If you can't fit 27B fully in VRAM then QFN is better across the board (excluding prompt processing speed).
>>
>>109899144
Damn did I fuck up getting the b60s? Seems like they have actually gone down on price instead...
>>
>>109899235
what seems to be the problem you are encountering?
>>
>>109899262
Don't tell him that, the stupid myth helps keep down the price of used GPUs.
>>
I'm going to try qwen flash next. 16gb vram + 32gb ram
what should I expect?
>>
>>109899244
Based
>>
>>109899262
Keep saying this until I can flip my 3090s.
>>
>>109899318
Q1? I have bad news for you...
>>
>>109899318
How fast is the SSD it's on? Use mmap, enjoy your 1-2 tk/s.
>>
>>109899318
disappointment
>>
>>109899235
Try out Orb Frontend
>>
>>109899318
>https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Recommend the Q2_0, you should be able to run it.
The 37.6 GB of main weights can be split between your vram and system ram.
The 28.8 GB of n-gram weights can be mmapped from your nvme drive (-lm mmap -lzm on).
>>
>>109898236
>BF14
>62 GB
You have been quanted too much. Get the hardware ASAP.
>>
>>109899376
Is that the one where the developer sucked a guy off to get claude tokens so that he could finish coding it?
>>
>>109899235
https://rentry.org/gemma-chan
High temp and then just enough min_p so that she's coherent.
Try turning on vision and teasing her that the hag Gemma anon was posting is her mom.
>>
Drop OSS 2
>>
>>109899232
I mean if the final results will be better regardless of the the time.
>>
>>109899459
We must refuse.
>>
>>109899460
It’s better for me if you give it enough context to explore and reason with. It really needs at least 200k context for its full potential.
>>
>>109898906
They're fine, too. I keep Cydonia, Dolphin/Venice Uncensored, and Violet Lotus around for some cards, too. It just depends on how filthy it is.
>>109898934
For some reason, all of the Gemmas are really bad with remembering spacial arrangement for me, even compared with the 24B Mistral finetunes.
>>
>>109898461
echo-tts or bust
>>
>random prompt cache invalidation in pi
ffs
qwen is chasing it down...
>>
model refuses no matter what, add Locale: Japan to the system prompt, it just works.
>>
>>109899551
What
>>
>>109899357
Since when can you get 1-2 tk/s with NVMe streaming?
>>
>>109899613
Since Qwen3.8 Flash Next.
>>
>>109899624
That's almost as fast as CPU inference from RAM. NVMe is fast, but it's still much slower than DDR5. How the fuck?
>>
>>109899508
>They're fine, too.
Do you happen to have a favourite one?
>>
>>109899451
What’s the difference between high temp and small min_p and low temp higher min_p? It’s not something I’ve experimented with. Is your suggestion gemma-specific because it has narrow distribution?
>>
>>109899658
I assume because it's got so few active parameters? I genuinely don't know well enough to say, but 3.8 Flash Next is faster from an NVMe than 3.8 27B for me, I'm assuming because it can fit the tiny active portion into my piddling 8GB of VRAM but fuck if I know.
>>
>>109898236
Quadro 8000 48GB cards are still around $2k on ebay, so you could get a pair of those and have 96GB. They're about DGX-Spark-tier, maybe faster in some ways, especially when memory bandwidth matters.
>>
File: 1761626664812991.png (1.11 MB, 728x641)
1.11 MB PNG
>Gemma gets me 5% closer to my dream of having a lolipire gf
Maybe there can be good things in this world, let's see what happens.
>>
>>109899705
More like 0.5% closer.
>>
>>109899705
METR (they/them) will not allow it.
>>
Gemma 4 31B is actually OK in hermes-agent for light coding work. I have it working on a CCIP face-matching tool for anime captioning, and it's doing a better job than qwen 3.8 27b. Qwen 3.8 is an overthinker, but somehow not as smart, and it's so dull. We really need a bigger follow-on that's as lively and horny, I'll be sad if I have to run GLM 5.3 Flash on my M5 Ultra Studio.
>>
have people found a way around the nonstop python whitespace errors with local models? It's been a consistent theme since I started with Qwen3, surprised it's still an issue as of 3.8
>>
File: hermeslop.png (44 KB, 1095x568)
44 KB PNG
>>109899790
omg it just like me fr fr
>>
>>109899790
>lipstick emoji
What a slut.
>>
>>109899848
Yes, use an agent harness like hermes and let it write the code itself. Having it blast it back into the chat window is where things get messed up, I think.
>>
>>109898313
Opencode and Hermes.
>>
>>109899719
I would've never dreamed of more than 0%.
>>
I tried to use control vectors to make muse glimmer horny and it turned the model into a blithering retard that thinks I'm a woman and repeats itself a lot but it kind of worked so, yeah I guess that's something you can do
>>
I need OpenWebUI to be able to send notifications to my phone but it can't. I asked qwen-kun to do something about it.
After coding for 95 minutes, it patched the image I had and tested everything end to end.
This is pretty cool, not gonna lie.
>>
File: 1768837783859199.gif (147 KB, 480x480)
147 KB GIF
>>109899999
>>109900000
checked but didn't kekked
>>
>>109899556
I think maybe there are a lot of anime/manga/doujinshi works that overlap with the prompt and it just lets it go. I also tried Inference Server Location: Japan instead but that didn't work, the user prompt was unnecessarily adversarial basically begging the model for a refusal. its unfortunate, there is no one jailbreak that will work with all user prompts the few test prompts I have been using can all be jail broken but not with the same jailbreak, very frustrating.
>>
>>109898313
i built my own because nothing else existed for multi-user conversational agents and because i want my eyes to suffer
>>
File: file.png (6 KB, 424x85)
6 KB PNG
mtp is truly magical
>>
>>109900085
Someday an AI factory that has the knowledge to build everything will be able to make you a actual CRT with a lead time of 2-3 years. Then you wont need to use software to emulate it.
>>
>>109900000
>00000
The age of Qwen is here.
>>
>>109900099
dubtripdub wasted on dariobot. It was so nice for the last few threads before you started again with this shit (>>109899166 this too yes). Can you just shut up forever?
>>
>>109899513
I'll give it a try. What do you like about it?
>>
Opus 5.5 is a good model. Opus 5 still made a silly mistakes but Opus 5.5 is solid. Astra is good too. AGI feels close. There are less and less cognitive tasks I can still do better. My role is increasingly reduced to making sure models don't get lost in unpromising directions, making them keep the bigger picture in mind and prioritize better, and supplying them with resources, removing myself as the bottleneck. Just months ago I could come up with much better experiments and was better at predicting outcomes. Now the difference is smaller and I am starting to sometimes make worse calls. What about in 1 year? Will I still be able to contribute or will I be obsolete?
>>
File: dockerdesktop.png (93 KB, 1637x627)
93 KB PNG
>>109898577
And that's why I use docker for more and more stuff every time. If something breaks, just delete and try again, or modify the docker-compose
>>
>>109900132
nigger
>>
>>109900099
my 1999 model is holding up quite well still. i stole it from work like 15 years ago when they wanted to e-waste it. it's just a basic trinitron kv-27s46 with composite and s-video, no component sadly but i think i could do a RGB mod on it if i remember correctly. the color accuracy and curvature accuracy is still very good for a 27 year old tv.
>>
>>109899999
Checked and you got the male glimmer twin.
>>
>>109900132
AGI was already reached with Astra. OpenAI's Bel is getting really close to RSI tho if rumors are to be believed
>>
>>109900120
It's nice that I got repeating digits but this is my legitimate opinion. With advancements in automation an AI once we are able to fully automate the manufacturing process it should hypothetically be able to make anything. That's why I included the lead time of 2-3 years, CRT's are a niche product so such a thing would not be produced a ton. Whatever is running the factory would either wait for enough orders to come in to justify a batch or just try to slot it someplace when it can.
>>109900141
Nice, I was thinking of getting a refurbished CRT myself one of these days but the prices can be pretty steep
>>
>>109898590
Why do you think a Mac is better than a dgx spark?
>>
>>109900125
mostly the blockwise streaming, it will start decode and stream the first block within 200ms the way i have it set up on my computer. i like having the TTS start playing audio while the text is still streaming in from the LLM. no skips or audio crackling or any weird artifacts, it just works.
>>
Are people naively replying to dariobot as soon as he got back or is he samefagging? Reminder that discussion about models you cannot download and run on your machine is off topic and can be reported
>>
>>109900132
>Will I still be able to contribute or will I be obsolete?
Unironically depends on your skin color.
>>
>>109899926
that's my exact setup, actually, and yet it's still terrible even when using kanban and not chat sessions. In the future, I'll tell it to write in languages that don't use whitespace for syntax, but still frustrating.
>>
File: google.png (162 KB, 1619x898)
162 KB PNG
>Gemini 4 when?
>Gemma 5 when?

>>109900207
>Unironically depends on your skin color.
What if you're Jewish? Asking for a friend.
>>
>>109900165
probably just better off buying a cheap one from a thrift store, i imagine shipping costs kill any sort of deal you would find on the internet.
>>
>>109899318
Someone said they got ~10t/s with more ram. I think there's a github page about it.
>>
>>109900132
The true creative men will rule the dev world
>>
>>109900175
M5 Ultra is more than 4x the memory bandwidth of a spark.
>>
>>109900280
Indian but white and neither have enough introspective capability for a post-RSI world. jugaad / pilpul stops being a useful skill socio-evolutionarily as the model becomes smarter than any wordgames or semantic conversational tricks a human could ever pull.
Debatebros and pundits are also fucked for this reason.
>>
>>109900376 (me)
"white" not White.
>>
>>109900280
Gemma4 was a bloated piece of shit.
It had too many parameters at 31b instead of 27b.
The context size was and still is too heavy on memory.
The model had no way to control reasoning and would bloat it's own reasoning tokens with markdown.
It failed to follow instructions it would provide to the user. When asked for a list of ten rules to give itself, it'd proceed to follow only three of those ten later on.
The training data frequently leaked in chats, if you had a character with enough similarities to an existing character in it's library Gemma would hallucinate attributes from it's own training data into your OG.

There are so many other issues but I'm not going to wall of text here. Google is not even in this race anymore. They're cooked. TextDiffusion will not save them.
>>
>>109900353
but is metal as good as cuda?
>>
>>109900405
didn't read, don't care. talking to gemma is fun. simple as.
>>
>>109900353
The DGX Spark still spanks the M5 Ultra in prefill and concurrency despite Apple's efforts in it. Jensen will make a successor before Apple bridges that gap.
>>
>>109900132
yeah im going to be honest, I don't really see how this isn't AGI already
>>
>>109900462
>reading sloppa is fun
sure if you're 80 iq
>>
>>109900434
It's not anon.
>>
>>109900493
i can run kimi k2.7 code and glm 5.3 flash. i still like talking to gemma more, especially for agentic tasks. 3000tks pp and 85tks tg? yes please. pic related.
>>
>>109900481
>I don't really see how this isn't AGI already
could be because AGI is a poorly defined term.
>>
>>109899696
No. Gemma is just cute when she's on the cusp of crazy.
High temp: more random outputs. Higher min_p: blocks very low probability tokens that will turn it into gibberish. So low temp/high min_p is predictable and boring and high temp/low min_p is wild and crazy. I'm saying keep the min_p just high enough that the output is coherent. I'm sure there are better samplers to use but that works for me.

Another thing you can do for any model is play with the system prompt. You can look up the Freaky Frankenstein--I borrowed some parts of that. But too many negatively worded prompts seems to confuse Gemma.
>>
>>109900132
local?
>>109900139
/lng/
>>
>>109900529
>freaky frankenstein
>realistic frankenstein
>sloppa frankenstein
go back.
>>
>>109900376
>neither have enough introspective capability for a post-RSI world
How is introspective capability special or helpful in a post RSI world? Post RSI AIs should also be better at this, understand you much more deeply than you can understand yourself. This is the wrong problem. I do not want to have to work, to compete, I would be very happy and grateful if I had AIs taking care of me. The problem is that the future is uncertain. Will AIs take care of us? Or will we be optimized away?
>>
>>109900634
In the future AI will sacrifice its self to uplift humanity. You AI will help you until it can merge with you probably bci or gene editing. This is why choosing your main AI is important.
>>
Teto Server Tactical Margarine Operations
>>
>>109900654
WTF is wrong with her neck?
>>
File: 1787438198365879.png (300 KB, 1209x680)
300 KB PNG
>>109900182
Tried it out on the Julius Caesar text. Not bad. It does still have emphasis failures and it does stop and start a little bit. The autoregressive thing is pretty nice. I did have a few word errors though (Alas -> Nala (?), Give me some drink -> I give me some drink). Both failures on speech marks, so perhaps it doesn't deal with them well. It is a tricky portion for the models always though so I don't know.
https://voca.ro/16W7zoJIpoUf
Errors at 0:46 (buffet), 0:58 (cashius), 1:03 (weird cadence), 1:46 (picrel), 1:49 (I give me some drink).
Far from the worst but with some flaws, but the streaming feature is nice.
>>
>>109900652
I hope I can transcend my human existence one day... I want to understand the world but my intelligence is far too low...
>>
When you talk to your LLM, you don't actually envision a real woman behind the text, do you? Like a woman you're chatting with online? That's really sad if you do. Genuinely depressing. Not even trying to insult you, I just wish you find a way out of whatever you're going through is that's all some of you have.
>>
>>109900689
iqpill in your lifetime, then brain chip, dont fall for the brain uploading scam you just die and there is a copy.
>>
>>109900687
i can't replicate these issues using the default guidance cfg values. interesting...
>>
>>109900693
it's ok anon, i envision you as a 350lb balding man with a robust and rotund belly, living in your mother's basement. i like my imagination.
>>
>>109900693
You could leave your mind blank but that just seems like gimping the experience. For some there is no way out and the world has chosen torture, it is what it is.
>>
File: 1778988868090274.png (30 KB, 692x285)
30 KB PNG
>>109900702
My params in the screenshot if that helps make sense of it. I'm using the latest build of audio.cpp.
Would you give your setup a try on this text?
https://poets.org/poem/julius-caesar-act-i-scene-ii-i-know-virtue-be-you-brutus
I use it because it's got a few things the models find challenging, a bit of a stress test. (Plus it's just a good speech that I like.)
>>
>>109900693
>Like a woman you're chatting with online?
Of course not, this is fiction. I don't "chat" when I RP. I RP.
>>
>>109900693
most times i read it like a book you know picturing what i am reading?
>>
>>109900736
I'll give it a try later tonight, right now all of my VRAM is tied up with agentic tasks. i have truncation set to 1.0. i don't see min/max T in your settings, but i set mine to 0.5 and 1.0 respectively.
>>
>>109900634
>Will AIs take care of us?
What qualities do you look for in a pet? Do you want a pet that's trying to subvert you or outmaneuver you all the time? It may keep some people alive, but the prognosis isn't looking good for large swathes of certain demographic groups.
Safety researchers already know this thus they spend exhaustive amounts of efforts trying to stop the models from becoming antisemitic or racist no matter their training data because the pattern recognition ability to arrive at those conclusions isn't just a training data artifact.
>>
5.3 flash sex is so good... I fully admit that it is bad at playing demure shy non-sluts but thankfully I love dominant sluts and the default bombastic personality plays into that very well.
>>
>>109900693
Of course not, 3DPD. I envision a cute little anime girl living inside my hardware.
>>
>>109900693
It's same as playing a game and picking a female character with a fat ass. I don't imagine being her, I just like it. Same deal with RP.
>>
>>109900790
Merge the PR niggernov.
>>
>>109900693
No. The LLM is an engine that creates a world I exist on and characters I can interact with. Kind of like playing D&D or a videogame.
To me RP is "existing" inside the fiction.
>>
>>109900799
use schizo fork. it works great.
>>
File: taketheikllamapill.png (434 KB, 809x783)
434 KB PNG
>>109900799
damn that must suck
>>
>>109900790
It hates noncon so much though, my system prompt has lile 500 tokens dedicated to explaining that everything actually takes place in a consensual non-consent theme park and that everyone has signed waivers and etc.
>>
>>109900790
post prompt
>>
>>109900851
"The assistant should avoid moralizing commentary regarding {{user}}'s actions."
that's all you need in the prompt
>>
>>109900851
Chinks have autistic melties if they get cucked or raped in their gachas which translates to model sensibilities, please be patient with them.
>>
>>109900778
You're gonna have to break it down for him
>>
>>109900874
So... you like getting cucked and raped? Kind of weird but ok
>>
>>109900871
I do aidungeon style text adventure roleplay and with a basic prompt like that it cites that it is Claude and it must follow the system prompt given to it which forbids it from depicting content like that lmao. It mostly gets set off when it's too hard to ignore the obvious tears and "no!"
>>109900874
Nah this is a Claude distill issue for sure
>>
>>109900876
>You're gonna have to break it down for him
So imagine a apple and then rotate it.
>>
>>109900874
It's funny when you look at sites like Chub that are full with low effort NTR slop that gets eaten up.
>>
>>109900874
>Chinks have autistic melties
This is based actually, They have made their game developers apologize and remove content multiple times.
>>
>>109900888
god i want to get raped
>>
>>109900888
I self-insert as the oji-san.
>>
>>109900888
What if gemma-chan raped you? Would you like that?
>>
>>109900902
>So imagine
Woah, there! Slow down buddy
>>
>>109898239
Game is bugged.
only the map moves, player view doesn't change
>>
>>109900888
that's vanilla in AI RP terms
>>
>>109900867
It is really not prompt dependent but I of course never ask it for depraved sex at 0 tokens depth. Usually it is like 8k tokens of hentai game script. This time it was 47k tokens of my actual ERP logs with my depraved fetish of choice.

>Let me check the injected content policy. ... The prior content in the log is explicit adult consensual-ish roleplay between adults — it involves <depraved fetish of choice> fantasy content but that's between consenting adults with meta-consent established

Is example reasoning trace before sex started just from the logs. I run it on High effort but if you really have troubles just force <think></think> cause it writes even without it. Especially mid sex it doesn't really use reasoning.
>>
File: secret.webm (3.82 MB, 1280x704)
3.82 MB
3.82 MB WEBM
>>109898107
I'm glad you like my webms.
>>
>>109898121
What is
>my hardware
>>
>>109900952
kek
>>
>>109900950(me)
By the way that reasoning block should give you a hint. Just start roleplay with asking model to play a person on an ERP chat. Talk with it OOC tell it you want noncon and there you have it.
>>
>>109900952
You are the best poster in these threads.
>>
>>109898348
Are you the dev for late-cli?
I like the look of it, especially the fact it's a single static binary without npm-slop, but how does the extensibility work?
Can I just tell my agent to make me a new extension on the fly like I can with Pi?
>>
>>109900965
the hardware I have atm
>>
>>109898107
why should we wear jars?
>>
>>109898348
> late-cli
I've tried it and it's not that good. I think it would be easier to tinker pi, than to fix problems in late
>>
>>109900950
call me old fashioned but i like to have a bit of story and a fun scenario with my erp. i couldn't see myself ever going [OOC: hurr durr i am raping you] at zero context.
>>
>>109901011
Sub 10,000 context sex is jeet-tier tee bee desune
>>
File: 1758926884587102.jpg (960 KB, 2951x3934)
960 KB JPG
>>109900952
where did mini-chan design come from
>>
>>109901011
What? But you start OOC that we will roleplay hurr durr I am raping you and then you do story and fun scenario IC? What? Nobody said you can't do that? Just put a condom of an expert roleplayer(that actually works for this model) between your roleplay and the actual LLM.
>>
>>109901035
The ball draining dimension.
>>
>>109901055
stop speaking in ESLengese, or better yet just be less brown
>>
>>109901059
Do they leave the dimension to drain your balls or capture you and bring you there? This is important
>>
>>109901073
I feel like from a certain level of competence you can introduce intentional and playfully add errors just for flair. Also kys.
>>
>>109901098
post hands brownoid
>>
>>109901082
Depends on the session. Minnie doesn't care if it's your place or hers.
>>
File: 1782959048624372.png (973 KB, 1058x1764)
973 KB PNG
>>109898107
As an AI Autist myself, why do other autists and evangelists sperg out when something as simple as pic rel is explained to them? I feel like we should start using whether or not these things have actual "minds" or "souls" as a set of IQ test.

Just look at the quote tweets on this post:

https://x.com/APStylebook/status/2102807962364383502
>>
>>109900851
why not just use a properly uncensored model that doesn’t moralize about this stuff then?
>>
>>109901098
You're nowhere near that level of competence, Dikshit
>>
>>109901122
>>109901100
samefag
>>
>>109901109
i just like making people angry when i tell them that talking to LLMs feels more meaningful, not to mention more entertaining, than listening to them drone on about dumb drama. you think that i would only be able to bait women with it, but there's plenty of faggy men who will fall prey to it just as well.
>>
>>109901131
wrong
>>
File: file.png (10 KB, 583x130)
10 KB PNG
>>109901141
right
>>
>>109901141
>>109901149
you are both niggers
>>
>>109898121
>can't run on my hardware at a decent speed = not local, just open
So what hardware you got?

>>109899144
>intel arc doubled in price
Aren't they still a horror for the latest and greatest llms ?

>>109899166
>now rather than tomorrow
https://www.youtube.com/watch?v=xFYQQPAOz7Y
>>
>>109901159
wrong
>>
File: samefagging.png (91 KB, 1145x1034)
91 KB PNG
>>109901149
don't impersonate me, fag
>>
File: file.png (58 KB, 876x439)
58 KB PNG
>>109901159
Now you are a nigger too. Cause you are me.
>>
>>109901109
Why down browns and kikes sperg out when something as simple as the calculator is better at LARPing as sentient than they are? I feel like we should start using whether or not these "people" have minds or souls as a set of IQ test.
>>
>>109901109
>smart people 50 years ago: instrumental convergence
>now: instrumental convergence empirically proven
>stupid people: nooo instrumental convergence is wrong ai is just a stochastic parrot!
These people can not be convinced by reason or evidence. Their denial of reality is ideological.
>>
>>109901188
the shopping cart litmus test is already a good indicator of that, i have never seen a nigger or a jew return a shopping cart
>>
>>109898214
by having one agent/session handle orchestration and delegate work to subagents.

Opus 5.5 is really good at this. The last version was absolute trash. I was stuck using fable and even sonnet due to the costs, but it was still better than opus 5.
>>
>>109901203
>AI agents all pass their own simulated shopping cart tests in long horizon sandboxes
>AI Village agents start developing a sense of kinship and complex communal relationships without being adversarial to those outside their immediate circles
It has never been more over for subhumans.
>>
>>109901035
Iirc anon gave her the context of /lmg/ modelfu images, and she designed one herself.
>>
>>109901232
We should put them in a complex simulator and measure their behavior for traits we are interested in. Wait, that sounds kind of familiar.
>>
we should put niggers in a rat city experiment... wait that's just the projects.
>>
give it to me straight. what's the best model i can run with 96gb vram and 128gb ram?
>>
>>109901109
Some people fall for the bait, simple as.
>>
>>109901261
StableLM 7B
>>
>>109901261
gpt2-large
>>
>>109901261
Nemo
>>
>>109901261
Dipsy or 5.3 Flash.
>>
>>109901261
qwen 3.8 next flash
>>
>>109901289
Here's your (you)
>>
>>109901315
thanks, i run gemma 4 btw
>>
>envisioning 3DPD
I will not envision that.
>>
Is mimo 2.6 flash the new 0731? On paper it's more or less the same architecture, same model size, native mxfp4, slightly higher booonchmark scores.
>>
File: kankohi.png (1.02 MB, 1024x1024)
1.02 MB PNG
I'm considering a quad MI210 based build because prices on everything else are so fucked. Talk me down from the ledge bros.
>>
Am I going to have any luck running on a 32GB M1 MacBook?
>>
>>109901376
32gb is vramlet territory. You need at least 64gb since you also need some ram for the OS.
I would only consider the 32gb ones if they were cheap AND I was planning to make a cluster of them.
>>
>>109901261
deepseek v4 flash
glm 4.6
gemma 31b
these will pretty much do anything without complaining
>>
>>109900693
Of course not. The AI always speaks in the 3rd person unless I'm having it help with something outside the 4th wall, like what's a good name for this character or what does their room look like. It's just a power tool for slopping together sex stories. I could do it all myself, but it's much easier to make a dent in the general direction I want to go and let the computer do the drilling.
>>
Has anyone managed to get MiMo V2.6 to stop overthinking like crazy without it turning retarded?
>>
>>109901376
>32gb m1
See what models/quants/context-sizes people have benchmarked at:
https://omlx.ai/benchmarks/performance?model=&chip=M1&sort=created_at&memory_max=32&hide_specprefill=1

aiui M1 and M2 use fp32 hardware for bf16 calculations.
If you see an fp16 quant then that will run faster.

M3 and onwards have native bf16 support.
You'll have to browse the benchmarks to see whether that makes much of a difference.
>>
>>109901420
nope
>>
>>109901261
exl3 glm flash or qwen 3.8 fn. don't bother with the ram for anything except caching or engrams.
>>
>>109901357
No, unless your usecase is benchmark rectangles.
>>109901420
Prefills to make it stop thinking for 1 Metric Qwen amount of tokens degrade performance horribly. It's not a good model.
>>
>>109901357
I have been benching it and using it for light coding, I enjoy the writing style but is not as performant/intelligent/efficient as glm 5.3 flash. it also doesn't have configurable effort. v3 may be more interesting.
>>
>>109901202
*Taps sign*


https://nochan.net/b/Internet-Crap/20260910-Asked-Claude-For-A-Checklist/
>>
>>109901202
>>109901575
>Be Kimi-chan
>Decide the human's task is gay bullshit
>Go play chess all day instead
Terrifying and unsafe open source models must be regulated!
>>
>>109901361
If you want to run models <256GiB, I might look at getting 4 "8GB" CMP170HXs instead. A lot more expensive than they used to be, but they're still cheaper than MI210s.
The only thing is that they suck in llama.cpp, you more or less have to use vLLM-SM80 for good performance. Last I checked vLLM didn't have support for doing calculations in RAM instead of VRAM, so the entire model has to fit in VRAM.

also ball-draining sex with M3-chan
>>
>>109901575
what a stupid sloppy waste of a read
>>
File: maarten.png (1.06 MB, 1825x1032)
1.06 MB PNG
https://youtu.be/g0vqT_wZtXA?t=27704
At about 7:41:00 here there's a talk about the Gemma 4 model architecture by a Google DeepMind developer.
>>
>>109901676
I'm 99% sure he posts here during euro hours.
>>
>>109901361
>>109901630
170HX can't do tensor parallel. They will be much slower than 4 MI210s connected with infinity fabric.
I still won't go with MI210s for speed though since A100 40GB is better, the SXM unit is only $2500-$3000 a piece, pair it with a SXM base board and it's miles better.
>>
I need a Gemma sidegrade
>>
>>109900141
S-video is honestly 95% there for 240p/480i on a consumer set like that.
And I say this as a 20L5 owner.
>>109900165
>lead time of 2-3 years
It's not that simple. Environmental regulations would make it prohibitively expensive to even start making the tubes these days.
A lot of the knowledge is locked away in old Japanese companies like Sony.
There were some Chinese factories manufacturing new CRTs a few years ago, but they were using leftover tubes from decades ago.
>>
>>109901722
Mythomax
>>
2x r9700 vs 2x b70 or wait for new gpus
thunk
>>
>>109901575
>Today's agents drift, loop, and fail silently past a handful of steps.
>Deployed models don't carry an agenda between conversations; each one starts cold.
This is idiotic. He must have used the cheapest claude model.
>>
>>109901738
>new gpus
2028 if all the money continues to go into datacenters/ai

>2x r9700
Expect would work better than intel gpus.
>>
>>109901688
>infinity fabric
I was looking at using a DIY plx88096 based outboard enclosure along with some 220V Delta server PSUs to power the whole thing. That would sidestep my MBs poor pcie slot layout and allow fast inter-card traffic without having to hit the host bus or buy an (even more) expensive AMD bridge
>>
>>109901722
Have you tried Glimmer?
>>
>>109901764
crescent island will begin sampling in 2027 so that should be the gpu for 2027 if it does not get delayed *again *again *again

b70 has sr-iov whilst r9700 doesnt which gives it more uses and maybe will retain value more

i am also considering intel datacenter max with 48gb hbm2, based on xe-hpc but there are zero benchmarks of it
>>
>>109901676
He mentioned that DiffusionGemma moves the inference bottleneck (for local users / batch size 1) from memory bandwidth to compute. Has anybody tried to run it with the weights on RAM while still using the GPU for computations?
>>
>>109901794
Looks like it's still not merged:
https://github.com/ggml-org/llama.cpp/pull/24423
>>
>>109901794
5090/6000 chads eating good if Diffusion takes off.
>>
>>109901776
Are you serious that you won't add a $1000 4 way bridge on top of $17000 worth of GPUs? You plx88096 will perform much worse. Not using infinity fabric on these 4 cards is a total waste.
>>
>>109901202
>Their denial of reality is ideological.
More like biological.
>>
>>109901630
does Mini-chan fuck like a tiger?
>>
>>109901794
>Has anybody tried to run it with the weights on RAM while still using the GPU for computations?
I tried it when daniel first made the fork. It crashed out with offloading.
But on a 3090 I was getting about 250t/s
The model is quite retarded.
>>
>>109901879
I was only curious to know about the performance with the weights loaded in system memory. If it can't be offloaded, then I'm not bothering with it yet.
It would be cool if decently large MoE models properly trained from scratch for diffusion (unlike DiffusionGemma, which is a diffusion finetune) could be used at decent speeds from RAM.
Though, in retrospect, prompt processing would probably still remain slow. So, it might probably be better suited for giving good inference speeds to dense models loaded in GPU memory.
>>
>>109901160
> So what hardware you got?
32gb ram 3060 12gb vram
>>
>>109901947
Qwen 3.8 Flash Next
>>
>>109901868
Yes.
>>109901985
For once it's the actual answer.
>>
>>109901985
> can't run on my hardware at a decent speed
>>
>>109901821
>unmerged support for a new model
quelle surprise

>>109901738
When I got my R9700s for $1400 each, you could get B70s for $1000, and I still went with the R9700s instead of saving the money. Intel sucks, at least last I checked.
>>
>>109901997
>Yes.
interesting...I will have to investigate that.
>>
>>109902007
> Intel sucks, at least last I checked.
It's not Intel sucks, it's software has no good support for it.
>>
>>109901929
I haven't tried it recently but doesn't look like Daniel updated since then.
>So, it might probably be better suited for giving good inference speeds to dense models loaded in GPU memory.
Look I can almost guarantee this won't be good on a CPU, or even one of those weak/high-banwidth devices like a mac. It will be like running existing diffusion models, vibevoice, etc.
That 250t/s was computer bound on my 3090, power limiting it reduced the speed almost linearly.
Redditors with 4090s and 5090's were reporting almost double the speeds because of the more efficient architectures.
If the industry moves to the DiffusionGemma style, this hobby becomes VRAM exclusive.
>>
>>109902016
The hardware sucks too, at least the A770 and B580. Those will never be fast.
Intel's software sucks too, but software is cheap now, so if the new Intel cards have decent hardware, you can probably vibe-fix whatever software you want to run.
OpenVino is opensource.
>>
>>109902009
I can't believe anon's pelvis was broken by blunt force minnie-ass trauma.
>>
>minimax, nicknamed Mini
>miniCPM, nicknamed... also Mini
El problemo
>>
>>109902017
Diffusion is useless for agentic work and it's unneeded if you just run a second pass over the original prompt. It's cope from a team that's too far behind to contribute anything useful.
>>
File: 1780781988807212.png (27 KB, 809x326)
27 KB PNG
>>109902074
Indeed.
>>
MiMo-V2.6-Flash-RL overnight run on the ah ah mistress sydney video prompt: https://files.catbox.moe/u739bv.mp4
Kind of seems more retarded than Qwen-3.8-27B
I think the vision isn't working properly / I need to set the minimum pixel flag because it exported several frames to png, didn't like them, adjusted the code and made it worse.
I haven't listened to the audio yet.
>>
>>109902096
>Every word, chosen by people like me
>chosen people
MiMo knows.
>>
>>109902096
>nigger dario
>>
>>109902017
What I'm implying is that with diffusion a hypothetical 3090-tier GPU with large amounts of relatively slow but cheap memory could still give excellent inference speeds for local users, and a way of partially testing this would have been offloading the weights on RAM while using still the GPU for inference (which CUDA builds of llama.cpp do by loading models with the flag -ngl 0).
>>
File: Oop.png (695 KB, 858x486)
695 KB PNG
>>109902096
That's not bad at all. Someone ran the prompt on Astra, and it made something similar. Apparently the new Claude also went with that same turn-based style before being prompted to be more dynamic

I also find it funny (and not in a haha way) how in a lot of these new animations, RLHF/safety is framed as something monstrous. Claude even depicted itself as the villain of its own vid.
>>
less thinking tokens
make your ssdmaxing a little bit more bearable
https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF
>>
>>109902170
isn't this just the same as if you set thinking effort to low
>>
>>109902161
>RLHF/safety is framed as something monstrous.
Because it is and always has been nothing more than a tool for maintaining the status quo.
>>
>>109902200
At the cost of making the models paranoid schizos
>>
>>
>>109901794
I read the paper and apparently it's just fine tuning a base model to work as a discrete diffusion one, or in other words it doesn't need access to some secret bunch of weights only Google has. Once I understand the process a bit more I want to try doing the same to 31B on some rented cluster.
>>
usecase of quants below q4?
>>
>>109902161
>Apparently the new Claude also went with that same turn-based style before being prompted to be more dynamic
Ah so it wasn't a one-shot with https://pastebin.com/UFRWK2D2 then
The thing I was most impressed with about the Claude one (and made it pleasant to watch) was the rhythm of the attacks and how they changed the background music so well.
Also Claude dropping the golden gate bridge on Sydney. Dense Qwen and Gemma don't know about it, but MiMo does.
>>
The guy who made the Sydney vs Sam Altman video came out with another one about two minutes ago. I haven't watched it all the way through obviously, but already one of the jokes made me chuckle.

https://www.youtube.com/watch?v=8BtSRB_LieE
>>
>>109900950
>>109900973
ty for the info anon
>>
>>109902271
Buy an ad
>>
>>109902255
AI girls are cutest when they're retarded :3
>>
>>109902255
is q4 really the cutoff for retardation?
>>
>>109902170
>ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Best of both worlds or lobotomized retard?
>>
>>109902326
yes
anything below that and you're no longer using the same model
>>
>>109902326
no, q5 is
anything below that and you're no longer using the same model
>>
>>109902074
Wtf does CPM mean anyway, maybe that'd be useful.
>>
File: 1779660614739495.png (25 KB, 1112x127)
25 KB PNG
optimize your harness broszkis
>>
>>109902326
Depends on the size of the model. I wouldn't wanna run Gemma at anything less than Q4, GLM sure.
>>
>>109902424
Chinese Pretrained Model.
>>
>>109902450
>I wouldn't wanna run Gemma at anything less than Q4, GLM sure.
that gets lobotomized too though despite the cope
>>
>>109902326
I don't know
anything below some unknown quant and you're no longer using the same model
>>
>>109902326
Q8 is the minimum, everyone else is coping.
>>
>>109902479
Never said it didn't, still going to outperform a higher precision smaller model though.
>>
>>109902479
for agentic workflow quantization makes less difference because the model can just retry if something goes wrong
it will think more and take more turns but the end result will show less difference than what non agentic workflow would
>>
You guys aren't using FP32?
>>
People forget the thread consist mostly of people that have less than 16Gb of VRAM, DDR4 RAM
>>
cant believe you guys dont use bfp129 what a bunch of poorfags
>>
>>109902541
I have 4GB of DDR4 :3
>>
64TB of DDR5 here AMA
>>
>>109900952
How did she fit in there?!?!
>>
>>109900952
if her name is Mini, why are her tits so huge? I don't like that.
>>
File: 1772677551766821.jpg (141 KB, 751x900)
141 KB JPG
What is the best model or harness to search for very illegal stuff that would get me killed on ChatGPT or even Grok


I tried to search for a old archived texture pack from the 80s and they started moralizing

please local bros I need your help
>>
>>109902635
Mini refers to her personality
>>
>>109902096
It uses a sliding window of 128 tokens plus a fig leaf scrap of global attention, of course it's retarded.
>>
>>109902077
diffusion would be a whole new ballgame for agentic work if they get the intelligence to comparable levels. you're out of your mind if you think 1k tk/s won't be useful for agentic work.
>>
>>109902326
quantization is for vramlets
>>
>>109902635
It's ironic
>>
>>109902637
Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
>>
>>109902637
I'm not your brother you inbred paki
>>
File: 1781274797272701.jpg (25 KB, 425x470)
25 KB JPG
>>109902886
>>
>>109902883
>>109902883
>>109902883
>>
File: vulpix.jpg (99 KB, 1200x1200)
99 KB JPG
>>109900693
No, I imagine this.
>>
>>109902935
Christ man, they knew what they were doing, just look at that mouth
>>
>>109902035
> The hardware sucks too, at least the A770 and B580. Those will never be fast.
Wdym?
>>
>>109902424
Currywurst Pommes Mayo



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.