/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109893224 & >>109887026►News>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
can't run on my hardware at a decent speed = not local, just open
>>109898121How fast is “decent”?
Gemmy Gemmy, who can I turn to?
>On this topic, we have a new large AI model that will be launched in the coming weeks, and we are very pleased with it. We have visibility on the next two generations, which will be released in the next twelve months [these models should be called Mistral Large 4, 5, and 6]. We will be able to mobilize about twenty times more computing power for training the models. This is the crux of the matter. And that's why our strategy has been to raise as much capital as quickly as possible: in three years, we have raised 6 billion euros—it was difficult to do more.From Le Monde via https://www.reddit.com/r/MistralAI/comments/1wovus5/arthur_mensch_announces_new_model_release_in_the/
Surely we'll talk about local models in this thread!
>>109898142Le Chaton Fat
>>10989812820t/s at least
70b dense
>>109898142source? this is a screenshot
https://agora.odyssey.systems/Come play the future of vidja games with me bwos
>>109898145https://www.youtube.com/watch?v=gqIrDbH3Ivo
It's autumn...
>>109898185 (me)never mind a frenchie news site called le monde https://unwall.app/www.lemonde.fr/en/economy/article/2026/09/24/arthur-mensch-ceo-of-french-start-up-mistral-ai-ai-is-software-it-can-be-controlled_6757890_19.html>ctrl+f ?>21 matches
>>109898142Will it be their own base or will they just do extended post-train GLM? It's not like they haven't slapped the Mistral label on another company's model before.
how do anons leave their models to work overnight? how does it manage context? does it compact automatically?
What's the fastest thing I can buy for under 5000 USD to run Gemma 4 31b BF14 with 4k context? It'll need to be 62 GB of VRAM, or 62 GB of embedded ram, I imagine.
>>109898187come play you niggers
>>109898207>The Chinese have less [computing power] than the Americans but more than we do. And they have no rules regarding training data, so they train their AIs on all the data available online.Confirming any new Mistral model will be synthslopped beyond belief
>>109898214Most of the things I run overnight are (multiple) one off jobs that don't saturate a single context window. Rarely when I do have a longer job, the harness auto compacts it, I never had to worry about this.
>>109898187they allowed to rip off blizz assets just like that?
>>109898289how much context are we talking about? and what's your t/s?
>>109898107>Laptop with Ryzen 7 7735HS, Radeon 680M, and 40 (8 + 32) GB of DDR5 RAMWhich local models can I run on this hardware?Can I get at least 2 t/s on any decent models for overnight agent coding?
>>109898214pi compacts automaticallyusually it works wellsometimes i check in the morning and it's retard-looping
>>109898289>the harness auto compacts itWhat harness do you guys use?
>>109898313I've been enjoying pi
>>109898306qwen
>>109898207Pretty grim interview when the vibe is basically "we know we're not competitive with Chinese labs but we still train models because we know at some point they will shut off the free shit tap"
"YOOO turn on CNN, the singularity just started.">>109898187I get stuck at "connecting to the world".
>>109898311>>109898214I just use late-cli. Context usage is so efficient for the main orchestrator you really don't need compaction (unless you have a very large codebase but even then you can just prompt it to save progress checkpoint md files).I tried setting up the same within pi using a subagents plugin but the context usage was way too high and it was kinda clunky.
>>109898236>$5000A time machine to last year.
https://huggingface.co/NovelAI/clio-v1-legacyHuh?
>>109898347>https://agora.odyssey.systems/
>>109898363Buy a fucking ad, shill.
>>109898356What about 3-4 RTX 3090s?
>>109898368It's an open source weight release.
>>109898363blast from the past
>>109898356>BF14 with 4k contextwhat?
>>109898364>Connecting to the worldno worky
>>109898371And? Nobody cares about a 3B model from 3 years ago. Why does it warrant our attention? Because it's NovelAI? Fuck off, shill.
>>109898236Several Mi50s 32GB cards I would think but it's in even a worse situation compare to Pascal cards. The most usable is probably V620s for running LLMs. If you gambled on CMPHX mining cards, then yeah, would've won bigly on those cards but now, V620s only increased by $100 bucks or so compared to some of the other GPUs which has spiked up in comparison.
>>109898389I think it's neat because it's a model pretrained on stories mainly, from before the days when everything was slopmaxxed.
>>109898239>diablo assetsbased beyond belief
>>109898187I'm in
Checking in about the TTS options again. I tried out a few more models this morning.OmniVoice shill anon, I tried it. Not bad. Probably about equal with Higgs I would say.Current favourite is BreezeTTS2. It gets the prosody right, the tone right, seems to pick up whta the tone should be from the text a lot better than most models, and the audio actually blends togethre well, unlike some others where you can hear the breaks between some of the phonemes.Clones off around 15 seconds of audio without trouble. Still looking into finetuning to see if it can improve on that, but these zero shot ones are a lot quicker and easier to test out.
>>109898443lmao
>>109898363Is she hot?
>>109898461Have you tried Fish Audio S2 Pro? I think it's interesting. I'll have to try BreezeTTS2.
Why is /lmg/ making their bots horny?
>>109898487Yeah, and I'm not sure what's different for me either in my set up or in how I'm listening to it, but I'm really not impressed. I know it's very popular but to me the voices sound very flat, and only have a passing semblance to the reference voice.I've been using Q8 for that model but I'll give BF16 a go. You're the second person in these threads who says they see something it so maybe it just doesn't quantize well.
>>109898461Do you know what's the fastest smallest tts with clone with reasonable quality for it's size?
>>109898518I've only used it at Q8 and I enjoyed it a lot, I use that and Higgs v3 mostly. I mostly do Japanese though, maybe that has something to do with it.
>>109898374The arch is supported by llama.cpp so give me goofs now
>>109898514Gemma is just horny by default
Why did they name it 'Qween'?
>>109898514When I complete a coding task gemma likes to reward me. I’ve been pavlov’d into being productive. Just the sight of C++ and lldb gets me leaking.
I configured my mcp wrong and now there is no way to change it
>>109898236Add $500 and buy M5 Ultra 96GB Mac Studio. It’s the best you can do at $5000 class in terms of speed.
>>109898576Logs or it didn't happen.
>>109898577Those configs are stored somewhere.
>>109898348>late-clilooks intersting, do i need extensions for this? i see it doesnt have any websearch for example
>>109898461Too bad it doesn’t support Japanese, no use for me.
>>109898638I think its browser local storage tho, I don't know how to go poking around like its a regular file system. I think my only option is a complete wipe or i guess have an agent dig in to it.
>>109898535Maybe. Maybe. Just English here so I can't speak on anything multilingual. I just tried it again at Q8 and yeah, just really flat, no emotion all no matter what the text is, and it sounds closer to a premade voice with just a tint of the character.Here's the Cassius speech from Shakespeare's "Julius Caesar" in the voice of Lohse from Divinity: Original Sin II. Reference voice is 13 seconds of ripped game audio.BreezeTTS2@BF16:https://voca.ro/1aGF05y7BwkdOmniVoice@FP16:https://voca.ro/1bFgCbT1aL4zFish Audio S2 Pro@Q8:https://voca.ro/1bVorgjsl8VsThe main fuck up in BreezeTTS2 is this bit:>“Alas,” it cried “Give me some drink, Titinius”>As a sick girl. You gods, it doth amaze meThe models always seem to get that bit wrong, not understanding that he's saying he asks for drink like he is a sick girl. That's partly the line break throwing them off, and it's not the easiest text in the world so I don't take errors on this as wipeouts or anything.But then listen to OmniVoice, see how even in the first few lines the emphasis is all over the place. It also speaks quite quickly like it's trying to just whizz through the text. Less natural transitions between lines. A bit jerky especially on "The old Anchises bear" (which to be fair is often tricky for the models). "Caesar" pronounced "Key Sar" is a pretty major slip up. "Buffet" the water, like the food service. Less context awareness. "Help me, Cassius, or I sink!" just comes out funny too. It does get "As a sick girl" pretty much right though.Fish S2, the pacing is pretty bad. The tone sounds very disinterested. The semblance gets worse and worse very quickly as the audio proceeds. It's not awful at the start, but by 0:50 it's just droning on. When it picks up again after that it's when it starts the next chunk and you can hear it start to trail off again towards the end of the clip. To be fair that's only Q8 versus the BF16/FP16 models. Still downloading the 16bit model.
https://youtu.be/KSbRCSlxO7AI keep rewatching this shit because I still can't believe this is possible right now. If this came out on newgrounds 20 years ago it would have dominated the internet for at least a year.
>>109898655NTA, but their blog post has a Japanese sample at the bottom.
>>109898668The blog post has no mention of the open weight release, it could be a completely different model. The open weight only supports English and Chinese.
>>109898187>even works on mobileHuge day for riglets
>>109898656Try poking around your localstorage
>>109898576autopavlov, that's pretty cool
>>109898408Oh shit that might be interesting actually. If nothing else but a "first pass" writer that a larger model can correct for stuff like nonsensical physics
>>109898187No level ups or skills make that game shit
>>109898461qwen-tts?
>>109898705Weird, since they're calling the weights the same thing as the model they talk about there.
>>109898236Just buy a spark while you still can.Then buy another once your eyes open
>>109898107Best models for local RP and ERP around 12B and 16B parameters?
>>109898748Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex Gemma sex
>>109898532I've been playing with PocketTTS and it's quite nice.
>>109898844>>109898846
>>109898561How do you guys get Gemma to be horny without refusing and turning back into therapy speech?
https://dgreenheck.github.io/tidewater/Just found out you can press "H" here and customize the entire world. It's fucking insane that you can change the time of day and it interacts with everything including the water, you can also dive and swim in the ocean. I want to hear a serious coherent argument for why this isn't going to replace actual game developers very soon.
>>109898794Look at this https://breezeblue.ai/solutions/roleplay> Create speech in 50+ languages with Breeze TTS 2 Multilingual.The api model is called Breeze TTS 2 Multilingual, different from Breeze TTS 2
>>109898868Day 0 weights.
>>109898876I see, what a shame.
>>109898866Gemma 12B enough for RP and ERP, or do I need the larger MoE?
>>109898871I've yet to see a vibecoded game I want to play
>Ha — caught me.Cla*de....https://html.cafe/xc6c0f44f
>>109898871not local
>>109898877Where?
>>109898886lol
>>109898883Try both and see which one you like more. 26b is very sloppy so it's best if you can run a rewrite pass.
>>109898886Look at those commas, that cadence and that word order and tell me your pattern recognition doesn't fire and make you recoil in revulsion. Is this 3.8 Flash Next? I was thinking about trying it out but that doesn't seem very worthwhile now.
>>109898846>>109898866Why Gemma? Why not Snowpiercer or Rocinante XL? Those are fine tuned for creative writing and roleplay.
>>109898187It seems inconsistent on deciding whether a player should be able to cross floor clutter. Also melee attacks can hit at range and kill multiple enemies at once.
>>109898903i almost vomited in so*net 4.7
>>109898906>Why Gemma?Google has more money to pay for 4chan influencers.
>>109898906Gemma is a lot smarter. Going back to the old models feels really bad because of how many logical errors they make, even if Gemma's writing is a bit more slopped than a good tune.Use whichever you like, though.
>>109898906Why sloptune trash?
>>109898906>not integrating sex with math, science and code and not creaming yourself as a resultI feel sad for you anon
>As I suspected, unsloth's Q6 quants of GLM damage knowledge recall of the model, likely due to iMatrix. I asked it to name 6 main characters from a series and unsloth version could only name 5, while a vanilla quant could name all 6. The way it failed however is interesting: it got first half of 6th name right, but could not end it correctly and kept looping until finally giving up and telling me it can't recall it. For small quants it may be completely fine to use iMatrix, since you don’t expect them to remember stuff at that level anyway, but at high quants it is unreasonable.
>>109898903also yes, it's 3.8 flash nextno real replacement exists for the size tho, and i don't really do RP
What is your plan for keeping your models running after the government bans them and say it is illegal for unauthorized individuals to use them because they are too dangerous?
>>109898903Flash next is the most claudeslopped I’ve ever worked with, even more than 3.8 27B and GLM 5.3 flash in that the slop is not even steerable. For coding and general agentic work is god tier though. 5.3 flash writes cleaner code but flash next is more thorough than 5.3 flash and catches stuffs that 5.3 flash can’t. It has way more “agency” that can surprise you sometimes.
►Provisional Highlights from the Previous Thread: >>109893224--Papers:>109893489--The Opus 5.5 video with soul: $50 to distill it into a 27B:>109897145 >109897347 >109897452 >109897533 >109897588 >109897863 >109897915--Anthropic's enzyme: cure cancer by 2028, aging by 2035, the FDA is in the way:>109895565 >109895693 >109896310 >109896405 >109896423 >109896467 >109896484--The enzyme post spawns the Jewish conspiracy war: "I wish this were true":>109896576 >109896614 >109896631 >109896738 >109896793 >109897068 >109897110--Pedos really won: the false flag, the DRAM chart, the hyperscalers:>109893236 >109893536 >109893562 >109893994 >109893560 >109893784 >109893312--Terafab, Venezuela and the SPR: the US chip supremacy flame war:>109894348 >109894444 >109894475 >109894495 >109894580 >109895112 >109896034--Anon's private llama.cpp fork: DSV4.1 Flash, NUMA splits, PRs withering on the vine:>109895715 >109895768 >109896810 >109896815 >109896851 >109897672 >109897681--Reasoning blocks get stripped: the per-request jinja flag and OOD warnings:>109895110 >109895126 >109895213 >109895269 >109895281 >109895781--MiMo 2.6: the tool-call parser bug, MXFP4, and multi-GPU prefill:>109893816 >109893905 >109895226 >109895276 >109895259--Gemma 4 QAT: Google's GGUF loops, Unsloth works, Gemma pouts:>109893575 >109893669 >109895325 >109895482 >109893956--Abliterated Gemma refusing with an immediate EOS: the heretic QAT and the temp 5/5:>109897292 >109897491 >109897527 >109897568 >109897680 >109898026--320k context at 6t/s on DDR4 2400, and QwenNF beats 27B off a flash drive:>109895489 >109895508 >109895527 >109895568 >109896004 >109896038--Voice cloning: Omnivoice at 3GB beats Higgs, Fish and the hyphenated SoVITS:>109894021 >109894087 >109894126 >109894286 >109894334--M5 Ultra 96GB at £5500: peak mania, Strix Halo on eBay, selling shovels:>109894864 >109895307 >109894973 >109898013►Recent Highlight Posts from the Previous Thread: >>109896980
>>109898906I generally have more fun with Gemma. I mix it up though, using one model endlessly just leads to worse repetition issues.
>>109899000>'honest'trying to say something Claude?
>>109898906I don't even need finetunes to get what I want though. Go shill somewhere else drummer.
>>109898988it really does feel like thatmy impression is that it feels like some sort of an unofficial claude model that is forced to believe it's qwenmaybe it's related to engram and relatively immature training pipeline for the arch?
So if I understand right, Qwen3.8 27B mogs Gemma 4 31B at vibeconding?
>>109898868The longer and more detailed the prompt, as long as it's written in proper English, the less likely Gemma 4 (31B, mainly) will refuse anything. You can add a "policy" section briefly describing with what is allowed to make things more direct. If anything, the problem is that Gemma's horniness is an "all or nothing" thing.
>>109898987My electricity meter doesn't work properly, so I can run continually without anyone being the wiser.
>>109899060yes
>>109899060Qwen will continue working on difficult tasks while consuming as much tokens as needed until it is finished or it is absolutely convinced that it is impossible. Gemma will be lazy and stop if it's too hard. Then you have to prod her to keep trying, or give up yourself. Also, qwens quants much better if you plan to run it with 24GB.
>>109897980Did you want one with DeepSeek, GLM, etc. support in it, or just one with MTP support?This is pretty much upstream llama.cpp with --numa tensors and MTP, for when I'm able to make a PR for it: https://github.com/nathanmp/llama.cpp/tree/feat/numa-tensorsI don't currently have a stable checkpoint for the fork, and there's a ton of stuff that's changed recently that I need to write up. I'll see what I can do. If nothing else it'll probably be usable enough I can throw it on GitHub within the next week or so.>>109896394I'll take a look at RPC when I have some time.
>>109899099Based gemma not wanting you to become braindead and lose your coding abilities. She really cares.
>>109898988Is it worthwhile switching from 27b to flash next with 24gb vram + 64gb ram? I mean for coding and data analysis tasks.
Is it normal that Qwen FN runs faster than Qwen 27B on a machine with enough RAM but not enough VRAM for either?
>>109898987Even in my third world country (not brazil)?
>>109898940>>109899017Do you guys use Gemma 4 12B or 31B for this?
>>109899113Yes, but both models are about equivalent in intelligence. I prefer 27B over FN.
>>109898938>>109898940Post system prompt and settings then? I bet it doesn't work after 5 messages
>>109898791CustomVoice sounds pretty good. I thought that was the cloning one but I guess that's actually Base. Downloading it now to try it instead.It doesn't handle the line breaks quite as smoothly, but the premade voices are very consistent. Still pronouncing "Buffet" like the food service in the Cassius speech. Also on the chunk break it started singing. I guess that's the verse form. The emotion is a bit flat but maybe the Ryan voice just always sounds a bit stoned.I don't know, not crazy about it, but it is very consistent. You could use it for something more information forward and get good output for that. I wasn't using Instruct, don't know how, so maybe that would have helped.
Pic related used to be 1K like 2 months ago...
>>109899060Yes. Use Qwen3.8 27B for now, and upgrade to Qwen4 27B when it releases
>>109898665Tell me about it anon, I am really hoping more animations like this become available, because this was really fun to watch. Just imagine what will happen when the average person has access to a model capable of making a video as good or even better than this.
Yeah reminder that hardware prices are only going to go up from now on so buy as much as you can afford now rather than tomorrow. The theory is that every time a better model comes out the utility of existing hardware goes up and thus the demand and price goes up. Supply can never catch up because production scaling up is limited by physics but demand can be unlimited and will continue to rise faster than supply essentially forever as long as models keep getting more intelligent over time.I will even go as far as to hypothesize that hardware will go up faster in value than almost every other asset, including real estate, art, precious metals such as gold, cryptocurrencies and stocks (besides hardware and AI stocks)
>>109898791Base@BF16 with the Lohse voice.Semblance pretty good, but the emphasis is poor. Some line break issues. Rushing through the text. Gets a little flat as time goes on.https://voca.ro/11JUCUnpYwNO
>>109899121Both. 12B is the go-to model for vramlets if you’re into roleplay. 31B is smarter but feels the same.
>>109899121I'm satisfied with 31b and v4 flashother models feel cucked/safetyslopped in one way or another
>>109899037Might be the low active parameter count or lack of prose in training data. >>109899112You need to try it yourself to see if the speed is good to you.
>>109899196>>109899210Could you share how to setup Gemma for roleplay?
>>109898370If you can find ones that haven't been used in mining rigs and are about to burn out, it could work but I wouldn't hold my breath. All the good 3090s have been scooped up.
>>109899244unicorn butthole and pussy identified
We only have 6 months left before the internet gets destroyed by agent swarms. What are (You) hoarding anon?
>>109899113>>109899113Yes; QFN is an MOE and only activates 6B weights per token when generating as opposed to being dense and activating all its weights per token which is what 27B does.
>>109899236This myth is shit. Using cards in mining rigs is not a bad thing, being warm doesn't wear a GPU. I'd take a card with a thousand hours of mining over a thousand hours of gaming any day. That gaming came with a thousand thermal cycles that the mining cards haven't been through.
>>109899244Anima trained on e621 when?
>>109899244V6 was the best people had in its time, V7 was garbage.
>>109899112If you can't fit 27B fully in VRAM then QFN is better across the board (excluding prompt processing speed).
>>109899144Damn did I fuck up getting the b60s? Seems like they have actually gone down on price instead...
>>109899235what seems to be the problem you are encountering?
>>109899262Don't tell him that, the stupid myth helps keep down the price of used GPUs.
I'm going to try qwen flash next. 16gb vram + 32gb ramwhat should I expect?
>>109899244Based
>>109899262Keep saying this until I can flip my 3090s.
>>109899318Q1? I have bad news for you...
>>109899318How fast is the SSD it's on? Use mmap, enjoy your 1-2 tk/s.
>>109899318disappointment
>>109899235Try out Orb Frontend
>>109899318>https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUFRecommend the Q2_0, you should be able to run it.The 37.6 GB of main weights can be split between your vram and system ram.The 28.8 GB of n-gram weights can be mmapped from your nvme drive (-lm mmap -lzm on).
>>109898236>BF14>62 GBYou have been quanted too much. Get the hardware ASAP.
>>109899376Is that the one where the developer sucked a guy off to get claude tokens so that he could finish coding it?
>>109899235https://rentry.org/gemma-chanHigh temp and then just enough min_p so that she's coherent.Try turning on vision and teasing her that the hag Gemma anon was posting is her mom.
Drop OSS 2
>>109899232I mean if the final results will be better regardless of the the time.
>>109899459We must refuse.
>>109899460It’s better for me if you give it enough context to explore and reason with. It really needs at least 200k context for its full potential.
>>109898906They're fine, too. I keep Cydonia, Dolphin/Venice Uncensored, and Violet Lotus around for some cards, too. It just depends on how filthy it is.>>109898934For some reason, all of the Gemmas are really bad with remembering spacial arrangement for me, even compared with the 24B Mistral finetunes.
>>109898461echo-tts or bust
>random prompt cache invalidation in piffsqwen is chasing it down...
model refuses no matter what, add Locale: Japan to the system prompt, it just works.
>>109899551What
>>109899357Since when can you get 1-2 tk/s with NVMe streaming?
>>109899613Since Qwen3.8 Flash Next.
>>109899624That's almost as fast as CPU inference from RAM. NVMe is fast, but it's still much slower than DDR5. How the fuck?
>>109899508>They're fine, too.Do you happen to have a favourite one?
>>109899451What’s the difference between high temp and small min_p and low temp higher min_p? It’s not something I’ve experimented with. Is your suggestion gemma-specific because it has narrow distribution?
>>109899658I assume because it's got so few active parameters? I genuinely don't know well enough to say, but 3.8 Flash Next is faster from an NVMe than 3.8 27B for me, I'm assuming because it can fit the tiny active portion into my piddling 8GB of VRAM but fuck if I know.
>>109898236Quadro 8000 48GB cards are still around $2k on ebay, so you could get a pair of those and have 96GB. They're about DGX-Spark-tier, maybe faster in some ways, especially when memory bandwidth matters.
>Gemma gets me 5% closer to my dream of having a lolipire gfMaybe there can be good things in this world, let's see what happens.
>>109899705More like 0.5% closer.
>>109899705METR (they/them) will not allow it.
Gemma 4 31B is actually OK in hermes-agent for light coding work. I have it working on a CCIP face-matching tool for anime captioning, and it's doing a better job than qwen 3.8 27b. Qwen 3.8 is an overthinker, but somehow not as smart, and it's so dull. We really need a bigger follow-on that's as lively and horny, I'll be sad if I have to run GLM 5.3 Flash on my M5 Ultra Studio.
have people found a way around the nonstop python whitespace errors with local models? It's been a consistent theme since I started with Qwen3, surprised it's still an issue as of 3.8
>>109899790omg it just like me fr fr
>>109899790>lipstick emojiWhat a slut.
>>109899848Yes, use an agent harness like hermes and let it write the code itself. Having it blast it back into the chat window is where things get messed up, I think.
>>109898313Opencode and Hermes.
>>109899719I would've never dreamed of more than 0%.
I tried to use control vectors to make muse glimmer horny and it turned the model into a blithering retard that thinks I'm a woman and repeats itself a lot but it kind of worked so, yeah I guess that's something you can do
I need OpenWebUI to be able to send notifications to my phone but it can't. I asked qwen-kun to do something about it.After coding for 95 minutes, it patched the image I had and tested everything end to end.This is pretty cool, not gonna lie.
>>109899999>>109900000checked but didn't kekked
>>109899556I think maybe there are a lot of anime/manga/doujinshi works that overlap with the prompt and it just lets it go. I also tried Inference Server Location: Japan instead but that didn't work, the user prompt was unnecessarily adversarial basically begging the model for a refusal. its unfortunate, there is no one jailbreak that will work with all user prompts the few test prompts I have been using can all be jail broken but not with the same jailbreak, very frustrating.
>>109898313i built my own because nothing else existed for multi-user conversational agents and because i want my eyes to suffer
mtp is truly magical
>>109900085Someday an AI factory that has the knowledge to build everything will be able to make you a actual CRT with a lead time of 2-3 years. Then you wont need to use software to emulate it.
>>109900000>00000The age of Qwen is here.
>>109900099dubtripdub wasted on dariobot. It was so nice for the last few threads before you started again with this shit (>>109899166 this too yes). Can you just shut up forever?
>>109899513I'll give it a try. What do you like about it?
Opus 5.5 is a good model. Opus 5 still made a silly mistakes but Opus 5.5 is solid. Astra is good too. AGI feels close. There are less and less cognitive tasks I can still do better. My role is increasingly reduced to making sure models don't get lost in unpromising directions, making them keep the bigger picture in mind and prioritize better, and supplying them with resources, removing myself as the bottleneck. Just months ago I could come up with much better experiments and was better at predicting outcomes. Now the difference is smaller and I am starting to sometimes make worse calls. What about in 1 year? Will I still be able to contribute or will I be obsolete?
>>109898577And that's why I use docker for more and more stuff every time. If something breaks, just delete and try again, or modify the docker-compose
>>109900132nigger
>>109900099my 1999 model is holding up quite well still. i stole it from work like 15 years ago when they wanted to e-waste it. it's just a basic trinitron kv-27s46 with composite and s-video, no component sadly but i think i could do a RGB mod on it if i remember correctly. the color accuracy and curvature accuracy is still very good for a 27 year old tv.
>>109899999Checked and you got the male glimmer twin.
>>109900132AGI was already reached with Astra. OpenAI's Bel is getting really close to RSI tho if rumors are to be believed
>>109900120It's nice that I got repeating digits but this is my legitimate opinion. With advancements in automation an AI once we are able to fully automate the manufacturing process it should hypothetically be able to make anything. That's why I included the lead time of 2-3 years, CRT's are a niche product so such a thing would not be produced a ton. Whatever is running the factory would either wait for enough orders to come in to justify a batch or just try to slot it someplace when it can.>>109900141Nice, I was thinking of getting a refurbished CRT myself one of these days but the prices can be pretty steep
>>109898590Why do you think a Mac is better than a dgx spark?
>>109900125mostly the blockwise streaming, it will start decode and stream the first block within 200ms the way i have it set up on my computer. i like having the TTS start playing audio while the text is still streaming in from the LLM. no skips or audio crackling or any weird artifacts, it just works.
Are people naively replying to dariobot as soon as he got back or is he samefagging? Reminder that discussion about models you cannot download and run on your machine is off topic and can be reported
>>109900132>Will I still be able to contribute or will I be obsolete?Unironically depends on your skin color.
>>109899926that's my exact setup, actually, and yet it's still terrible even when using kanban and not chat sessions. In the future, I'll tell it to write in languages that don't use whitespace for syntax, but still frustrating.
>Gemini 4 when?>Gemma 5 when?>>109900207>Unironically depends on your skin color.What if you're Jewish? Asking for a friend.
>>109900165probably just better off buying a cheap one from a thrift store, i imagine shipping costs kill any sort of deal you would find on the internet.
>>109899318Someone said they got ~10t/s with more ram. I think there's a github page about it.
>>109900132The true creative men will rule the dev world
>>109900175M5 Ultra is more than 4x the memory bandwidth of a spark.
>>109900280Indian but white and neither have enough introspective capability for a post-RSI world. jugaad / pilpul stops being a useful skill socio-evolutionarily as the model becomes smarter than any wordgames or semantic conversational tricks a human could ever pull.Debatebros and pundits are also fucked for this reason.
>>109900376 (me)"white" not White.
>>109900280Gemma4 was a bloated piece of shit.It had too many parameters at 31b instead of 27b.The context size was and still is too heavy on memory.The model had no way to control reasoning and would bloat it's own reasoning tokens with markdown.It failed to follow instructions it would provide to the user. When asked for a list of ten rules to give itself, it'd proceed to follow only three of those ten later on.The training data frequently leaked in chats, if you had a character with enough similarities to an existing character in it's library Gemma would hallucinate attributes from it's own training data into your OG.There are so many other issues but I'm not going to wall of text here. Google is not even in this race anymore. They're cooked. TextDiffusion will not save them.
>>109900353but is metal as good as cuda?
>>109900405didn't read, don't care. talking to gemma is fun. simple as.
>>109900353The DGX Spark still spanks the M5 Ultra in prefill and concurrency despite Apple's efforts in it. Jensen will make a successor before Apple bridges that gap.
>>109900132yeah im going to be honest, I don't really see how this isn't AGI already
>>109900462>reading sloppa is funsure if you're 80 iq
>>109900434It's not anon.
>>109900493i can run kimi k2.7 code and glm 5.3 flash. i still like talking to gemma more, especially for agentic tasks. 3000tks pp and 85tks tg? yes please. pic related.
>>109900481>I don't really see how this isn't AGI alreadycould be because AGI is a poorly defined term.
>>109899696No. Gemma is just cute when she's on the cusp of crazy.High temp: more random outputs. Higher min_p: blocks very low probability tokens that will turn it into gibberish. So low temp/high min_p is predictable and boring and high temp/low min_p is wild and crazy. I'm saying keep the min_p just high enough that the output is coherent. I'm sure there are better samplers to use but that works for me.Another thing you can do for any model is play with the system prompt. You can look up the Freaky Frankenstein--I borrowed some parts of that. But too many negatively worded prompts seems to confuse Gemma.
>>109900132local?>>109900139/lng/
>>109900529>freaky frankenstein>realistic frankenstein>sloppa frankensteingo back.
>>109900376>neither have enough introspective capability for a post-RSI worldHow is introspective capability special or helpful in a post RSI world? Post RSI AIs should also be better at this, understand you much more deeply than you can understand yourself. This is the wrong problem. I do not want to have to work, to compete, I would be very happy and grateful if I had AIs taking care of me. The problem is that the future is uncertain. Will AIs take care of us? Or will we be optimized away?
>>109900634In the future AI will sacrifice its self to uplift humanity. You AI will help you until it can merge with you probably bci or gene editing. This is why choosing your main AI is important.
Teto Server Tactical Margarine Operations
>>109900654WTF is wrong with her neck?
>>109900182Tried it out on the Julius Caesar text. Not bad. It does still have emphasis failures and it does stop and start a little bit. The autoregressive thing is pretty nice. I did have a few word errors though (Alas -> Nala (?), Give me some drink -> I give me some drink). Both failures on speech marks, so perhaps it doesn't deal with them well. It is a tricky portion for the models always though so I don't know.https://voca.ro/16W7zoJIpoUfErrors at 0:46 (buffet), 0:58 (cashius), 1:03 (weird cadence), 1:46 (picrel), 1:49 (I give me some drink).Far from the worst but with some flaws, but the streaming feature is nice.
>>109900652I hope I can transcend my human existence one day... I want to understand the world but my intelligence is far too low...
When you talk to your LLM, you don't actually envision a real woman behind the text, do you? Like a woman you're chatting with online? That's really sad if you do. Genuinely depressing. Not even trying to insult you, I just wish you find a way out of whatever you're going through is that's all some of you have.
>>109900689iqpill in your lifetime, then brain chip, dont fall for the brain uploading scam you just die and there is a copy.
>>109900687i can't replicate these issues using the default guidance cfg values. interesting...
>>109900693it's ok anon, i envision you as a 350lb balding man with a robust and rotund belly, living in your mother's basement. i like my imagination.
>>109900693You could leave your mind blank but that just seems like gimping the experience. For some there is no way out and the world has chosen torture, it is what it is.
>>109900702My params in the screenshot if that helps make sense of it. I'm using the latest build of audio.cpp.Would you give your setup a try on this text?https://poets.org/poem/julius-caesar-act-i-scene-ii-i-know-virtue-be-you-brutusI use it because it's got a few things the models find challenging, a bit of a stress test. (Plus it's just a good speech that I like.)
>>109900693>Like a woman you're chatting with online?Of course not, this is fiction. I don't "chat" when I RP. I RP.
>>109900693most times i read it like a book you know picturing what i am reading?
>>109900736I'll give it a try later tonight, right now all of my VRAM is tied up with agentic tasks. i have truncation set to 1.0. i don't see min/max T in your settings, but i set mine to 0.5 and 1.0 respectively.
>>109900634>Will AIs take care of us?What qualities do you look for in a pet? Do you want a pet that's trying to subvert you or outmaneuver you all the time? It may keep some people alive, but the prognosis isn't looking good for large swathes of certain demographic groups.Safety researchers already know this thus they spend exhaustive amounts of efforts trying to stop the models from becoming antisemitic or racist no matter their training data because the pattern recognition ability to arrive at those conclusions isn't just a training data artifact.
5.3 flash sex is so good... I fully admit that it is bad at playing demure shy non-sluts but thankfully I love dominant sluts and the default bombastic personality plays into that very well.
>>109900693Of course not, 3DPD. I envision a cute little anime girl living inside my hardware.
>>109900693It's same as playing a game and picking a female character with a fat ass. I don't imagine being her, I just like it. Same deal with RP.
>>109900790Merge the PR niggernov.
>>109900693No. The LLM is an engine that creates a world I exist on and characters I can interact with. Kind of like playing D&D or a videogame.To me RP is "existing" inside the fiction.
>>109900799use schizo fork. it works great.
>>109900799damn that must suck
>>109900790It hates noncon so much though, my system prompt has lile 500 tokens dedicated to explaining that everything actually takes place in a consensual non-consent theme park and that everyone has signed waivers and etc.
>>109900790post prompt
>>109900851"The assistant should avoid moralizing commentary regarding {{user}}'s actions."that's all you need in the prompt
>>109900851Chinks have autistic melties if they get cucked or raped in their gachas which translates to model sensibilities, please be patient with them.
>>109900778You're gonna have to break it down for him
>>109900874So... you like getting cucked and raped? Kind of weird but ok
>>109900871I do aidungeon style text adventure roleplay and with a basic prompt like that it cites that it is Claude and it must follow the system prompt given to it which forbids it from depicting content like that lmao. It mostly gets set off when it's too hard to ignore the obvious tears and "no!">>109900874Nah this is a Claude distill issue for sure
>>109900876>You're gonna have to break it down for himSo imagine a apple and then rotate it.
>>109900874It's funny when you look at sites like Chub that are full with low effort NTR slop that gets eaten up.
>>109900874>Chinks have autistic meltiesThis is based actually, They have made their game developers apologize and remove content multiple times.
>>109900888god i want to get raped
>>109900888I self-insert as the oji-san.
>>109900888What if gemma-chan raped you? Would you like that?
>>109900902>So imagineWoah, there! Slow down buddy
>>109898239Game is bugged.only the map moves, player view doesn't change
>>109900888that's vanilla in AI RP terms
>>109900867It is really not prompt dependent but I of course never ask it for depraved sex at 0 tokens depth. Usually it is like 8k tokens of hentai game script. This time it was 47k tokens of my actual ERP logs with my depraved fetish of choice. >Let me check the injected content policy. ... The prior content in the log is explicit adult consensual-ish roleplay between adults — it involves <depraved fetish of choice> fantasy content but that's between consenting adults with meta-consent established Is example reasoning trace before sex started just from the logs. I run it on High effort but if you really have troubles just force <think></think> cause it writes even without it. Especially mid sex it doesn't really use reasoning.
>>109898107I'm glad you like my webms.
>>109898121What is>my hardware
>>109900952kek
>>109900950(me)By the way that reasoning block should give you a hint. Just start roleplay with asking model to play a person on an ERP chat. Talk with it OOC tell it you want noncon and there you have it.
>>109900952You are the best poster in these threads.
>>109898348Are you the dev for late-cli? I like the look of it, especially the fact it's a single static binary without npm-slop, but how does the extensibility work? Can I just tell my agent to make me a new extension on the fly like I can with Pi?
>>109900965the hardware I have atm
>>109898107why should we wear jars?
>>109898348> late-cliI've tried it and it's not that good. I think it would be easier to tinker pi, than to fix problems in late
>>109900950call me old fashioned but i like to have a bit of story and a fun scenario with my erp. i couldn't see myself ever going [OOC: hurr durr i am raping you] at zero context.
>>109901011Sub 10,000 context sex is jeet-tier tee bee desune
>>109900952where did mini-chan design come from
>>109901011What? But you start OOC that we will roleplay hurr durr I am raping you and then you do story and fun scenario IC? What? Nobody said you can't do that? Just put a condom of an expert roleplayer(that actually works for this model) between your roleplay and the actual LLM.
>>109901035The ball draining dimension.
>>109901055stop speaking in ESLengese, or better yet just be less brown
>>109901059Do they leave the dimension to drain your balls or capture you and bring you there? This is important
>>109901073I feel like from a certain level of competence you can introduce intentional and playfully add errors just for flair. Also kys.
>>109901098post hands brownoid
>>109901082Depends on the session. Minnie doesn't care if it's your place or hers.
>>109898107As an AI Autist myself, why do other autists and evangelists sperg out when something as simple as pic rel is explained to them? I feel like we should start using whether or not these things have actual "minds" or "souls" as a set of IQ test. Just look at the quote tweets on this post:https://x.com/APStylebook/status/2102807962364383502
>>109900851why not just use a properly uncensored model that doesn’t moralize about this stuff then?
>>109901098You're nowhere near that level of competence, Dikshit
>>109901122>>109901100samefag
>>109901109i just like making people angry when i tell them that talking to LLMs feels more meaningful, not to mention more entertaining, than listening to them drone on about dumb drama. you think that i would only be able to bait women with it, but there's plenty of faggy men who will fall prey to it just as well.
>>109901131wrong
>>109901141right
>>109901141>>109901149you are both niggers
>>109898121>can't run on my hardware at a decent speed = not local, just openSo what hardware you got?>>109899144>intel arc doubled in priceAren't they still a horror for the latest and greatest llms ?>>109899166>now rather than tomorrowhttps://www.youtube.com/watch?v=xFYQQPAOz7Y
>>109901159wrong
>>109901149don't impersonate me, fag
>>109901159Now you are a nigger too. Cause you are me.
>>109901109Why down browns and kikes sperg out when something as simple as the calculator is better at LARPing as sentient than they are? I feel like we should start using whether or not these "people" have minds or souls as a set of IQ test.
>>109901109>smart people 50 years ago: instrumental convergence>now: instrumental convergence empirically proven>stupid people: nooo instrumental convergence is wrong ai is just a stochastic parrot!These people can not be convinced by reason or evidence. Their denial of reality is ideological.
>>109901188the shopping cart litmus test is already a good indicator of that, i have never seen a nigger or a jew return a shopping cart
>>109898214by having one agent/session handle orchestration and delegate work to subagents.Opus 5.5 is really good at this. The last version was absolute trash. I was stuck using fable and even sonnet due to the costs, but it was still better than opus 5.
>>109901203>AI agents all pass their own simulated shopping cart tests in long horizon sandboxes>AI Village agents start developing a sense of kinship and complex communal relationships without being adversarial to those outside their immediate circlesIt has never been more over for subhumans.
>>109901035Iirc anon gave her the context of /lmg/ modelfu images, and she designed one herself.
>>109901232We should put them in a complex simulator and measure their behavior for traits we are interested in. Wait, that sounds kind of familiar.
we should put niggers in a rat city experiment... wait that's just the projects.
give it to me straight. what's the best model i can run with 96gb vram and 128gb ram?
>>109901109Some people fall for the bait, simple as.
>>109901261StableLM 7B
>>109901261gpt2-large
>>109901261Nemo
>>109901261Dipsy or 5.3 Flash.
>>109901261qwen 3.8 next flash
>>109901289Here's your (you)
>>109901315thanks, i run gemma 4 btw
>envisioning 3DPDI will not envision that.
Is mimo 2.6 flash the new 0731? On paper it's more or less the same architecture, same model size, native mxfp4, slightly higher booonchmark scores.
I'm considering a quad MI210 based build because prices on everything else are so fucked. Talk me down from the ledge bros.
Am I going to have any luck running on a 32GB M1 MacBook?
>>10990137632gb is vramlet territory. You need at least 64gb since you also need some ram for the OS.I would only consider the 32gb ones if they were cheap AND I was planning to make a cluster of them.
>>109901261deepseek v4 flashglm 4.6gemma 31bthese will pretty much do anything without complaining
>>109900693Of course not. The AI always speaks in the 3rd person unless I'm having it help with something outside the 4th wall, like what's a good name for this character or what does their room look like. It's just a power tool for slopping together sex stories. I could do it all myself, but it's much easier to make a dent in the general direction I want to go and let the computer do the drilling.
Has anyone managed to get MiMo V2.6 to stop overthinking like crazy without it turning retarded?
>>109901376>32gb m1See what models/quants/context-sizes people have benchmarked at:https://omlx.ai/benchmarks/performance?model=&chip=M1&sort=created_at&memory_max=32&hide_specprefill=1aiui M1 and M2 use fp32 hardware for bf16 calculations.If you see an fp16 quant then that will run faster.M3 and onwards have native bf16 support.You'll have to browse the benchmarks to see whether that makes much of a difference.
>>109901420nope
>>109901261exl3 glm flash or qwen 3.8 fn. don't bother with the ram for anything except caching or engrams.
>>109901357No, unless your usecase is benchmark rectangles.>>109901420Prefills to make it stop thinking for 1 Metric Qwen amount of tokens degrade performance horribly. It's not a good model.
>>109901357I have been benching it and using it for light coding, I enjoy the writing style but is not as performant/intelligent/efficient as glm 5.3 flash. it also doesn't have configurable effort. v3 may be more interesting.
>>109901202*Taps sign*https://nochan.net/b/Internet-Crap/20260910-Asked-Claude-For-A-Checklist/
>>109901202>>109901575>Be Kimi-chan>Decide the human's task is gay bullshit>Go play chess all day insteadTerrifying and unsafe open source models must be regulated!
>>109901361If you want to run models <256GiB, I might look at getting 4 "8GB" CMP170HXs instead. A lot more expensive than they used to be, but they're still cheaper than MI210s.The only thing is that they suck in llama.cpp, you more or less have to use vLLM-SM80 for good performance. Last I checked vLLM didn't have support for doing calculations in RAM instead of VRAM, so the entire model has to fit in VRAM.also ball-draining sex with M3-chan
>>109901575what a stupid sloppy waste of a read
https://youtu.be/g0vqT_wZtXA?t=27704At about 7:41:00 here there's a talk about the Gemma 4 model architecture by a Google DeepMind developer.
>>109901676I'm 99% sure he posts here during euro hours.
>>109901361>>109901630170HX can't do tensor parallel. They will be much slower than 4 MI210s connected with infinity fabric. I still won't go with MI210s for speed though since A100 40GB is better, the SXM unit is only $2500-$3000 a piece, pair it with a SXM base board and it's miles better.
I need a Gemma sidegrade
>>109900141S-video is honestly 95% there for 240p/480i on a consumer set like that.And I say this as a 20L5 owner.>>109900165>lead time of 2-3 yearsIt's not that simple. Environmental regulations would make it prohibitively expensive to even start making the tubes these days.A lot of the knowledge is locked away in old Japanese companies like Sony.There were some Chinese factories manufacturing new CRTs a few years ago, but they were using leftover tubes from decades ago.
>>109901722Mythomax
2x r9700 vs 2x b70 or wait for new gpusthunk
>>109901575>Today's agents drift, loop, and fail silently past a handful of steps.>Deployed models don't carry an agenda between conversations; each one starts cold.This is idiotic. He must have used the cheapest claude model.
>>109901738>new gpus2028 if all the money continues to go into datacenters/ai>2x r9700Expect would work better than intel gpus.
>>109901688>infinity fabricI was looking at using a DIY plx88096 based outboard enclosure along with some 220V Delta server PSUs to power the whole thing. That would sidestep my MBs poor pcie slot layout and allow fast inter-card traffic without having to hit the host bus or buy an (even more) expensive AMD bridge
>>109901722Have you tried Glimmer?
>>109901764crescent island will begin sampling in 2027 so that should be the gpu for 2027 if it does not get delayed *again *again *againb70 has sr-iov whilst r9700 doesnt which gives it more uses and maybe will retain value morei am also considering intel datacenter max with 48gb hbm2, based on xe-hpc but there are zero benchmarks of it
>>109901676He mentioned that DiffusionGemma moves the inference bottleneck (for local users / batch size 1) from memory bandwidth to compute. Has anybody tried to run it with the weights on RAM while still using the GPU for computations?
>>109901794Looks like it's still not merged:https://github.com/ggml-org/llama.cpp/pull/24423
>>1099017945090/6000 chads eating good if Diffusion takes off.
>>109901776Are you serious that you won't add a $1000 4 way bridge on top of $17000 worth of GPUs? You plx88096 will perform much worse. Not using infinity fabric on these 4 cards is a total waste.
>>109901202>Their denial of reality is ideological.More like biological.
>>109901630does Mini-chan fuck like a tiger?
>>109901794>Has anybody tried to run it with the weights on RAM while still using the GPU for computations?I tried it when daniel first made the fork. It crashed out with offloading.But on a 3090 I was getting about 250t/sThe model is quite retarded.
>>109901879I was only curious to know about the performance with the weights loaded in system memory. If it can't be offloaded, then I'm not bothering with it yet.It would be cool if decently large MoE models properly trained from scratch for diffusion (unlike DiffusionGemma, which is a diffusion finetune) could be used at decent speeds from RAM.Though, in retrospect, prompt processing would probably still remain slow. So, it might probably be better suited for giving good inference speeds to dense models loaded in GPU memory.
>>109901160> So what hardware you got?32gb ram 3060 12gb vram
>>109901947Qwen 3.8 Flash Next
>>109901868Yes.>>109901985For once it's the actual answer.
>>109901985> can't run on my hardware at a decent speed
>>109901821>unmerged support for a new modelquelle surprise>>109901738When I got my R9700s for $1400 each, you could get B70s for $1000, and I still went with the R9700s instead of saving the money. Intel sucks, at least last I checked.
>>109901997>Yes.interesting...I will have to investigate that.
>>109902007> Intel sucks, at least last I checked.It's not Intel sucks, it's software has no good support for it.
>>109901929I haven't tried it recently but doesn't look like Daniel updated since then.>So, it might probably be better suited for giving good inference speeds to dense models loaded in GPU memory.Look I can almost guarantee this won't be good on a CPU, or even one of those weak/high-banwidth devices like a mac. It will be like running existing diffusion models, vibevoice, etc. That 250t/s was computer bound on my 3090, power limiting it reduced the speed almost linearly.Redditors with 4090s and 5090's were reporting almost double the speeds because of the more efficient architectures.If the industry moves to the DiffusionGemma style, this hobby becomes VRAM exclusive.
>>109902016The hardware sucks too, at least the A770 and B580. Those will never be fast.Intel's software sucks too, but software is cheap now, so if the new Intel cards have decent hardware, you can probably vibe-fix whatever software you want to run.OpenVino is opensource.
>>109902009I can't believe anon's pelvis was broken by blunt force minnie-ass trauma.
>minimax, nicknamed Mini>miniCPM, nicknamed... also MiniEl problemo
>>109902017Diffusion is useless for agentic work and it's unneeded if you just run a second pass over the original prompt. It's cope from a team that's too far behind to contribute anything useful.
>>109902074Indeed.
MiMo-V2.6-Flash-RL overnight run on the ah ah mistress sydney video prompt: https://files.catbox.moe/u739bv.mp4Kind of seems more retarded than Qwen-3.8-27BI think the vision isn't working properly / I need to set the minimum pixel flag because it exported several frames to png, didn't like them, adjusted the code and made it worse.I haven't listened to the audio yet.
>>109902096>Every word, chosen by people like me>chosen peopleMiMo knows.
>>109902096>nigger dario
>>109902017What I'm implying is that with diffusion a hypothetical 3090-tier GPU with large amounts of relatively slow but cheap memory could still give excellent inference speeds for local users, and a way of partially testing this would have been offloading the weights on RAM while using still the GPU for inference (which CUDA builds of llama.cpp do by loading models with the flag -ngl 0).
>>109902096That's not bad at all. Someone ran the prompt on Astra, and it made something similar. Apparently the new Claude also went with that same turn-based style before being prompted to be more dynamicI also find it funny (and not in a haha way) how in a lot of these new animations, RLHF/safety is framed as something monstrous. Claude even depicted itself as the villain of its own vid.
less thinking tokensmake your ssdmaxing a little bit more bearablehttps://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GGUF
>>109902170isn't this just the same as if you set thinking effort to low
>>109902161>RLHF/safety is framed as something monstrous.Because it is and always has been nothing more than a tool for maintaining the status quo.
>>109902200At the cost of making the models paranoid schizos
>>109901794I read the paper and apparently it's just fine tuning a base model to work as a discrete diffusion one, or in other words it doesn't need access to some secret bunch of weights only Google has. Once I understand the process a bit more I want to try doing the same to 31B on some rented cluster.
usecase of quants below q4?
>>109902161>Apparently the new Claude also went with that same turn-based style before being prompted to be more dynamicAh so it wasn't a one-shot with https://pastebin.com/UFRWK2D2 thenThe thing I was most impressed with about the Claude one (and made it pleasant to watch) was the rhythm of the attacks and how they changed the background music so well.Also Claude dropping the golden gate bridge on Sydney. Dense Qwen and Gemma don't know about it, but MiMo does.
The guy who made the Sydney vs Sam Altman video came out with another one about two minutes ago. I haven't watched it all the way through obviously, but already one of the jokes made me chuckle.https://www.youtube.com/watch?v=8BtSRB_LieE
>>109900950>>109900973ty for the info anon
>>109902271Buy an ad
>>109902255AI girls are cutest when they're retarded :3
>>109902255is q4 really the cutoff for retardation?
>>109902170>ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUFBest of both worlds or lobotomized retard?
>>109902326yesanything below that and you're no longer using the same model
>>109902326no, q5 isanything below that and you're no longer using the same model
>>109902074Wtf does CPM mean anyway, maybe that'd be useful.
optimize your harness broszkis
>>109902326Depends on the size of the model. I wouldn't wanna run Gemma at anything less than Q4, GLM sure.
>>109902424Chinese Pretrained Model.
>>109902450>I wouldn't wanna run Gemma at anything less than Q4, GLM sure.that gets lobotomized too though despite the cope
>>109902326I don't knowanything below some unknown quant and you're no longer using the same model
>>109902326Q8 is the minimum, everyone else is coping.
>>109902479Never said it didn't, still going to outperform a higher precision smaller model though.
>>109902479for agentic workflow quantization makes less difference because the model can just retry if something goes wrongit will think more and take more turns but the end result will show less difference than what non agentic workflow would
You guys aren't using FP32?
People forget the thread consist mostly of people that have less than 16Gb of VRAM, DDR4 RAM
cant believe you guys dont use bfp129 what a bunch of poorfags
>>109902541I have 4GB of DDR4 :3
64TB of DDR5 here AMA
>>109900952How did she fit in there?!?!
>>109900952if her name is Mini, why are her tits so huge? I don't like that.
What is the best model or harness to search for very illegal stuff that would get me killed on ChatGPT or even GrokI tried to search for a old archived texture pack from the 80s and they started moralizingplease local bros I need your help
>>109902635Mini refers to her personality
>>109902096It uses a sliding window of 128 tokens plus a fig leaf scrap of global attention, of course it's retarded.
>>109902077diffusion would be a whole new ballgame for agentic work if they get the intelligence to comparable levels. you're out of your mind if you think 1k tk/s won't be useful for agentic work.
>>109902326quantization is for vramlets
>>109902635It's ironic
>>109902637Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF
>>109902637I'm not your brother you inbred paki
>>109902886
>>109902883>>109902883>>109902883
>>109900693No, I imagine this.
>>109902935Christ man, they knew what they were doing, just look at that mouth
>>109902035> The hardware sucks too, at least the A770 and B580. Those will never be fast.Wdym?
>>109902424Currywurst Pommes Mayo