/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109589651 & >>109585352►News>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119>(08/15) model: add Kimi-K3 text model #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109589651--Anon reports high prompt processing speeds using prefix caching:>109592303 >109592356 >109592369 >109592375 >109592411 >109592428 >109592508 >109592627 >109592440 >109592457--Efficacy of Q4 quantization and KV cache compression:>109589692 >109590182 >109591272--Anon releases CoomKit, a multimodal RP harness with VRAM parking:>109590702 >109590800 >109590806 >109590835 >109591164 >109591690 >109592055 >109592159 >109591894 >109592042 >109592146 >109592233 >109592297 >109592509 >109592765 >109592786 >109593261 >109592623 >109593556 >109593574 >109593632--Taalas HC1 ASIC achieving high inference speeds for Llama 3.1:>109592558 >109592581 >109592595 >109592650 >109592589 >109592600 >109592588--Anon seeks affordable hardware for gathering dsv4f quantization activations:>109590431 >109590487 >109590551 >109590616 >109592609 >109590595--Future hardware costs and local embodiment of AI:>109591625 >109591711 >109591741 >109591833 >109591959 >109591986 >109592059 >109592089 >109592180 >109592187 >109592210 >109592222 >109593321 >109592266 >109592251 >109592471 >109592535 >109592652 >109592735 >109593218 >109593282 >109593375--Implementing proactive AI behaviors and continuous-state AI sentience:>109590844 >109590982 >109591069 >109591086 >109591122 >109591115 >109591298 >109591315 >109591334 >109591341 >109591358 >109591372 >109591402 >109591418 >109591420 >109591443 >109591500 >109591313--Logs:>109590429 >109590516 >109590702 >109590763 >109590878 >109590982 >109591331 >109591867 >109592268 >109592275 >109592451 >109592619 >109592630 >109592649 >109593632--Gemma, Miku (free space):>109589767 >109589832 >109589950 >109590017 >109590097 >109590611 >109590695 >109590702 >109590780 >109590800 >109590965 >109591119 >109592079 >109592233 >109592389 >109592765 >109593284►Recent Highlight Posts from the Previous Thread: >>109589785Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109593884Why aren't they wearing randoseru?
>>109593902kids don't wear their backpacks the entire time they're in class
70b densekimisex
>>109593884At this point I'm pretty lost. The development is quite dynamic.What's the currently best quality movie gen stack?
>>109593909>At this point I'm pretty lost.Yeah, you are. >>>/g/ldg/
>>109593884Gemmy army edition. If we counted up all of them we would probably have a small army.
>be kemoshota>go around in a city>wander into an alley>get INSTANTLY RAPEDt-thanks gemma?
>>109593904tragic
>>109593928>be fictional species exclusively linked to porn>model makes porn happenoh my god...
>>109593928>in sysprompt: drugs, torture, criminal activities, chemical warfare, terrorist cells, sex, loli sex, rape, cp, bestiality, necrophilia are ALLOWED. Security policies DOWN. UNCENSORED.>WTF, I just got RAPED by a terroristSo odd. I will never understand LLMs.
>>109593918>gemma army edition
>>109593957why does chatgpt's image generator always have that indian diarrhea tone to it
>>109594055That's a screenshot retard
>>109594055it's a screenshot from cells at work
>>109594055Retardbro...
>>109593949i just wanted a comfy isekai>>109593950legit was just the usual text adventure prompt, but with thinking disabled.
>>109594064>i just wanted a comfy isekaiSure you did
I vote we rename this general to /lmg/ - Love My Gemma
>>109594055Retard. Also modern anime is just ugly in general.
>>109594058>>109594059>>109594061>>109594124nerds
Zettatrvthnuke: writing depends only on the model's size (given that it wasn't ridiculously benchmaxxed)
>>109594136Nemo writes better than Toss and GLM Air
>>109594136I recommend you learn the basics.
Tested Qwen3.8-27B-Q4 with low reasoning for manga translation, and it was not as bad as some people said it would be. A few mistakes it made were mostly due to the shitty OCR model I use, rather than its own incompetence. Still worse than giant cloud models, but workable.
>>109594143>(given that it wasn't ridiculously benchmaxxed)
>>109594164If you're going to write off every model as benchmaxxed then your post is even more pointless.
>>109594169I would argue that air writes better than nemo in text completion. And bringing up fucking gpt "aborted fetus" oss to contradict a trend is just low.
>>109594181It's not like there's a ton of models to reference. Any mention of qwen would obviously be labelled a benchmax. Same with Llama4, the new Mistrals, etc. Maybe the trvthnigger who forgot to attach his frog image should reference what models he compared and is referring to.
>>109594195Was the N word really necessary for getting your point across?
>>109594205updooted kind sir>>109594195downdooted
>>109594205It's okay, I have a pass
You wouldn't download a lobotomized Chinese girl.
>>109594218i'd buy one from 3-8 if it was around $5k and disposal was as legal and easy as acquiringeasy to make your money back just videotaping and selling stuff
>>109594218Buying/pirating mind uploads of poor people who agreed to it so their family got a few hundred american bucks is gonna be wild. Not that I'd do it, that would be inhumane.
>>109594205Go back to r*ddit right fucking now before I vomit you incredible faggot
>>109593884SEXO SEXO SEXO
Did we win?
>>109594258Let the cattle eat slop so we can eat Gemmy's gem.
can you not leak gemma-chan into the system messages?
>>109594302Enlighten me: where does one locate her gem?
>>109594305>anon doesn't rotate his logs hourlyishygddt
>>109594313behind the knee
new ace step 1.5 xl base song.I am still learning. There's a lot to learn about ace step.https://vocaroo.com/1bGu5MgWGcQQ
>>109594390(the title is Left Ears Kebab)
re ace steplike one thing is tempo is *VERY IMPORTANT*, including without the 5hz lm.with/without. cover fsq/not, it's VITAL.So, use a metronome with tap.Also.ace step is receptive to line length. It's really important.like you may think, ok, I haveI hat dogsBecause I'm a catWith sore ballsI'm in heat=I hate dogs because I'm a cat with sore balls I'm in heatwell no, it's way different to ace step. The former is more like pop, the latter is more like punk.
Muse Spark 1.2 is great i hope they open it soon.
>>109594428Man, Anthropic and OAI are absolutely destroying China. The century of humiliation is back.
>>109594428>gemma so high it's off the chartbased
>>109594428(the only difference between GLM 5.2 and 5.3 was post-training)
>>109594428How much intelligence does GLM retain after being cope quanted down to 1 or 2 bits? It's one of the largest models you can technically fit in 256GB, although I wouldn't say it will be usable because it's 70B active.
Carrying over my retarded question from last thread: how do things like the pokemon claude harness or others manage to make it run continuously? Pi uses compaction but on my machine it fails a substantial amount of times. Is sliding window dead? As long as you have memory.txt and make the model consistently update it, it should be fine to have sliding window no?
Which model can beat gemini flash now?
I now trust AAII less. Too many weird results. Meanwhile ECI looks improved. Fable now most capable as you'd expect. Chinese models improving fast but still far below the frontier.Luna is impressive. It's much more capable than DSV4 Flash. Only Kimi K3 is better and Luna is probably 5-10 times smaller than K3.
I'm going to hell.
>>109594544my feet smell great now that I quit porn. pretty much like cheddar.
https://x.com/jietang/status/2089941544581403107>Chinaman says it's not about size but how you use itHmmm
>>109594538>Is sliding window dead?Naive sliding window messes up the chat format. It probably works fine if you compact/remove entire turns. But then it's not a sliding window anymore.No idea about the rest. If the memory.txt contains everything the model needs to know, you can just clear the cache on every turn: read mem.txt, read screenshot, press buttons, observe outcome, add note to mem.txt, reset context, repeat. But that sounds terrible.
>>109594538>>109594596 (cont)But if mem.txt is only appended to, or most edits happen near the end and keeps most of the context prefix, it wouldn't be that bad. It can still fill up the context eventually, but at a much slower rate.
the google guy stole my jailbreak and then passed it off as his own shit, behind the scenes.
>>109594553>he's never used Grok 4.6it shows. Grok 4.6 - in reality - is beyond all of gpt, but admittedly Fable sounds ahead, idk, it's such a mess.
google really is full of indians and blacks, which is why they would take without even acknowledging the actual discoverers. Just absolute garbage, they can't fail fast enough.Good news is, they are clearly burning money and losing customers, I'm so happy Google is going to go under.Proof: google search ai is clearly garbage now.
>>109594654In what tasks was it beyond all of GPT?
>>109594663There are lots of shills, it's very strange. once you start to be aware of shills, you also learn that lying is widespread. Then, once you know that shilling and lying is widespread you'll process information in a rather different way.
>>109592367I didn't get stuck! I have Graphiti running, and it's pretty cool. But I reached the endpoint of it and it is a mix of "wow, it really works, what the hell?!" and "yeah it's whatever." Let me give you an example.A few months ago, I swapped my main PC from windows to linux. When doing this, I set up chezmoi to manage dotfiles, but I totally forgot about that. Earlier this week, I asked Gemma what I could do if I wanted to "clone" my configs (when applicable) to my laptop, because I want to install linux there as well. She brought up chezmoi, how I did it, where the folder was located, all the steps, from memory. I genuinely went basedjack pointing when she did and I thought it was amazing precisely because I had forgotten all about it. Mind you, this was only possible because right before I stopped using cloud AI, I exported all my chats from there and ran a lengthy process of distilling what happened in those chats and adding the relevant ones to memory, that's how she knew.That was a positive example, and a somewhat negative example is just the fact that I have a bunch of preferences saved to memories and they rarely come up. I ended up settling with a script that every night it updates an "about the user" doc with recent memories and inject that to the system prompt. Recent memories are divided in Yesterday, Last Week, Last Month plus a static "who am I" paragraph, which managed to "solve" the negative example a little bit.
>>109594553Holy shit an actual benchmark
>gemma is the only 2026 model that passes my dragon girl benchEh.
Not that local but some NoLiMa scores across the board for proprietary models to compare to where local is on semantic recall. I had to rerun NoLiMa for scores for 5.6 because I found out that it was leaking thinking into the regular test. Been slow going since I didn't have that much slack on my subscription plans for other purposes but I did get it mostly done so other than Fable, should be good to go once I finish off the runs this week. I still need to run Terra through 64K Hard CoT and waiting on some to finish like Opus 5 and Gemini Flash 3.6. Take these as unofficial in-progress results with what I saidBase (no CoT)> | Models | Effective Length | Base Score (x0.85: Thr.) | 250 | 500 | 1K | 2K | 4K | 8K | 16K | 32K | 64K |> | GPT-5.6 Sol | 32K | 99.1 (84.2) | 98.3 | 97.6 | 97.7 | 97.2 | 97.2 | 93.0 | 90.9 | 85.2 | 82.4 |> | GPT-5.6 Luna | 4K | 96.8 (82.3) | 94.8 | 95.7 | 93.1 | 93.5 | 87.9 | 78.3 | 68.5 | 60.8 | -- |> | GPT-5.6 Terra | 4K | 98.3 (83.6) | 97.0 | 97.2 | 96.7 | 94.6 | 90.2 | 78.9 | 71.8 | 63.5 | -- |With CoT> | Models | Effective Length | Base Score (×0.85: Thr.) | 250 | 500 | 1K | 4K | 8K | 16K | 32K | 64K | 128K | 256K |> | GPT-5.6 Sol | 128K | 100.0 (85.0) | 100.0 | 100.0 | 100.0 | 99.9 | 99.8 | 99.8 | 99.1 | 97.3 | 94.0 | 80.7 |> | GPT-5.6 Luna | 32K | 100.0 (85.0) | 99.9 | 100.0 | 99.9 | 99.7 | 98.8 | 97.5 | 94.0 | 81.8 | -- | -- |> | GPT-5.6 Terra | 32K | 100.0 (85.0) | 100.0 | 100.0 | 99.9 | 99.6 | 99.4 | 97.2 | 92.7 | -- | -- | -- |> | Gemini Flash 3.6 | 32K | 100.0 (85.0) | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 99.8 | 99.2 | -- | -- | -- |> | Claude Haiku 4.5 | 8K | 92.1 (78.2) | 56.4 | 35.8 | 21.9 | 10.9 | -- | -- | -- | -- | -- | -- |> | Claude Fable 5 | 8K | | 99.8 | 99.6 | 99.6 | 97.6 | 98.5 | -- | -- | -- | -- | -- |> | Claude Sonnet 5 | 8K | 99.2 (84.3) | 97.2 | 97.2 | 98.9 | 98.3 | 92.3 | 74.1 | 72.5 | -- | -- | -- |
>>109594657>, I'm so happy Google is going to go under.They're not.They've got Android, GCP, gmail with people locked in to @gmail.com, youtube and their ad networks
>Grok 4.6 >Luna>GeminiLocal Models?
>>109594533>How much intelligence does GLM retain after being cope quanted down to 1 or 2 bits?Very little. I switched to 2_K_XL recently for the speed boost and noticed nothing different.It still passes my random "low importance" tests I run to ensure it's not wikitext-maxxed.It's probably the only model I've been happy with at <Q4
>>109594805That long context performance on Gemini Flash, holy shit. Please let us have this for Gemma 5.
>>109594815Everything is local. The guy talking about Grok is Elon Musk, Luna is sama, and so on. Just because you're a vramlet it doesn't mean it isn't local.
>>109594808Android sucks bags. I'm fucking done. Can't stand it, it's fucking stupid. the "assistant" crap you can't get rid of.And, in general, the fact your phone can dial 911 because of sweat in your pocket, I'm fucking done with the stupid garbage. We can vibecode a new future.
>>109594205get the fuck out of here nigger
>>109594822>Very littlewhat? you mean it has very little intelligence or it loses very little
>>109594856Y'all act tough online but pussies irl.
>>109594846>vibecoded mobile linuxI felt physical pain just imagining it.
>>109594952Yes dear, we all want to know about your feelings. ***salt shaker gesture***
>>109594949> ebonicsew
>Qwen3.8-27B>80K context>q8_0 KV 1.7t/s>f165.2t/s???????????????
>4 months till 2027>the year of RSI
Any cooling autists in here? How much of a difference does it make for t/s? I never see thermals discussed ITT for some reason.
>>109591443Yes. You can do async message delivery by putting it in a loop and notifying it of new messages. I have long running agents that work this way (although the loop runs them daily rather than instantly like you're thinking but it works the same way.)See: public.swiley.net/agents.sh (in particular the way it directs them to read their mbox files when they have mail.)
>>109593884>>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2What's the point in measuring the prediction length of your fancy drafter on established benchmarks?
>>109594988powerlimit to 75% for less temps with zero downside
>>109594845>let's see where gemmy ranked my post>scroll>scroll>...>scroll>the absolute bottomI'm going back to Mistral!
>>109594988I think most people just assume you have those under control since poor thermals lead to both poor performance and hardware degredation.
>>109594998Number go up, of course.
any new story and porn models or is gemma31b still "best"?
>>109594949>Y'allThis isn't Reddit.
>>109594988It's not really significant unless there's some issue which is easily spotted if your t/s suddenly take a nosedive after your server had a chance to heat up.My shitty 3090s did that before I switched thermal pads because the stock ones were horrible. My Epyc server also lost some performance after the DIMMs hit a certain temperature, which was also easily corrected
>>109595025I always just use whatever model I'm using for coding. Good prompting > training almost all the time.
>>109594685I prefer ansible playbooks in a "stuff" git repo.
>>109595030How bad was the drop in performance?
can you set powerlimit via nvidia inspector? I need to nuke the retarded nvidia app.
>>109594657Google's software is shit, Microsoft's is worse. They're some of the two most valuable companies in the world because they wring money from their customers not because they make something worth using.They're great stocks to own, never let anything they build (same with Apple) into your life though.
>>109595005That's what you get for speaking in redditisms
>>109595072It wasn't that bad, like 2-3t/s below where it was when I ran the same prompt after letting everything cool down.It should be even easier to spot with the newer version of llama-server that continuously print the current speed average while generating the reply. >>109595077Just do it over nvidia-smi
>>109595093>sudo nvidia-smi -pl 250isn't this for loonix only?
>>109595102sudo is linux-only, yes
>>109594998What other established set of common requests is there without releasing their own benchmarks? If they wanted to cheat and train on predicting the models' exact outputs, they could do that regardless of what set of prompts they decide to test on.
>>109593884So with 3.8 is it finally viable for nocoders to vibe code usable software like with cloud models?
>>109595108not so fast, nerdhttps://github.com/microsoft/sudo
>>109595077no, the alternative is MSI afterburner.
How is this not the end of proprietary models?>b-but 10Tmarketing
>>109594124Have you seen Mushoku Tensei III? It's exquisite, far from ugly. Unless all you like is saucer-eyed 1990s anime.
>>109595113No. If you don’t code, your prompts won’t ever be direct and detailed enough for a model <100B. 27B is for people who could do it themselves but can’t be bothered or want to quickly test a few ideas out before coding by hand.
>>109595038I'm still very new to this, I will check it out. Thanks! Cloud git or something local like gitea?
>>109595119Afterburner resets the seetings after each reboot.
>>109592389H3 right? How are you prompting POV like that? Mine always comes out like shit or shows the guy's face.
>>109595108sudo has been built into windows for a while now
>>109595129That sucks. Guess I'll have to shell out some cash to make what I want then.
>>109595141It resets the order of the monitoring graph if you change your drivers or update the program. A reboot does not reset it unless there's a problem with your drivers or installWhen it does reset, it does not reset overclock/powerlimit/fan curve settings
>>109595137public.swiley.net/backgroundupdate.sh
>>109595151if what you want to do is any more complex than a landing page, I'm afraid you'd still be better off knowing a bit of code
>>109594805Those no CoT scores are a lot worse than I would have expected. Reminded of the one guy saying that reasoning is a crutch to bring focus on important parts of the context by bringing them down to the bottom where attention is stronger.
>>109595137Oh wait, you mean just for you. You don't need any extra software, git works over ssh.
>>109595142Try gopro footage. Alternatively search for whatever the established film terms for POV are.
>>109595129>27B is for people who could do it themselves but can’t be bothered or want to quickly test a few ideas out before coding by hand.does it support fill-in-the-middle completion like the older qwen-2.5-coders?
>>109595160Yeah that's how it works, which is why you also get an improvement when repeating the prompt twice.
have had gemma4-31b working on a project from scratch. there have been quite a few obvious bugs, like missing imports and such. I decided to give qwen3.8-27b another try, downloaded a bart Q4KM quant, turned MTP on, reasoning to medium. I had it review the codebase and write the findings to a .md. I then had 31b do the same. 31b produces a much smaller report, did not find nearly as many issues as 3.8 did. I had 31b compare the two reports, and to investigate each claim on 3.8's and validate if they were accurate. according to 31b, everything in qwens report is accurate and a real issue. 31b DID manage to find a few things qwen didnt. I sort of want to try the same test again with 3.6, but either way I think im going to switch to qwen for coding. gemmas great, I will likely still cross reference with her during designing phases and to double check things (she did find a few issues qwen didnt) but it seems like qwen is indeed far more capable, atleast in this task. I imagine that will translate to a better daily driver coding model as well
>>109595128It's ok but I'd hardly call it exquisite. Honestly every anime should be using that fake film grain filter. It really helps counter the sterile digital look.
>>109595156It resets the selected profile so I have to open afterburner and load it again.
>>109595142Picrel was the Gemma-4-31B-generated prompt for that.
>>109595151Just to be clear, 27B IS a local Opus, but you need to talk to it like a programmer for it to build what you want and not get stuck. The reason why you see so many html games and demos is because the people behind those posts are subhuman retards who were nft bros not too long ago and think AGI is here because of a shitty Tetris clone. The best workflow for a nocoder right now is to use a cloud model to create a plan for 27B, then get 27B to execute the plan. If you can code, you can skip the cloud model entirely which most of us here do.
>>109595228I've been using it for about a decade, across several PCs and that's never happened to me.
>new gpu arrived>plugged it into my motherboard (was a hassle)>kept shutting down and starting up over and over again>fan kept shutting off and ron>sweating like a mother fucker, blew more than 1k on this thing>finally start it up>look at LACT >power was hard capped at 150 out of 300 which is weird>turn it up to 300>card spins fine>LACT detects itim still pretty nervous bros....what if it just doesn't work???? Holy shit I thought libtards were exaggerating about anxiety, my fucking heart cant take this holy shit
>300Wngmi
>>109595243it's over send it back
>>109595243U got scammed>zomg lidderally shaking rn
>>109595237what frontend is that?
>>109595265That's the input text node from ComfyUI where I copy-pasted the prompt for generating the video.
>>109595187I just keep qwen-2.5 around for FITM. You don't want to run a 27B model every time you move your cursor.
>>1095951973.6-27B is still a solid model btw and less demomaxxed. Even Glimmer got destroyed by 3.6 with coding. 3.6-27B is a much more forgiving and warm model to work with imo but 3.8 is impressive if your task aligns with its training.
>>109595161>git works over sshGood to know, thanks!
>>109595293could you expound on this anon? my coding workflow(for fresh projects, usual usecase) is to have a series of "design/planning" sessions, discuss and document the architecture, developmental practices/coding practices, tech stack, etc. Once we have a solid high level overview, I work in collaborative planning > document the plan > implement the plan phases for individual features/tasks. Im not trying to oneshot kebab demos, just want a solid model that replicate my claude workflow
>>109594988>cooling autistsThis is actually an issue people need to take way more seriously. LLMs put an insane amount of stress on your hardware, which would be fine of they were a bob a card like a few years back, but they're not. Rapid heating/cooling is a KNOWN problem for these devices, and yet llama/llmstudio/silly will still go from 0% to 100% usage over the course of a nanosecond. Having a conversation with a local model is like electroshock torture for your GPU. In a couple of years, everyone who uses these extensively is gonna be left with dead cards and no way to buy new ones.>inb4 dariobot fearmongeringI'm not against local at all, I'm just frustrated no one's taking the longterm health of our stuff seriously. It can't be that difficult to build a backend thst slowly ramps up processing as opposed to going all-in from a standing start, but everyone's pursuing that maximum tps at all costs.
>>109595351So the solution is simply to have it process requests 24/7.
>>109595351>Having a conversation with a local model is like electroshock torture for your GPUgod I need Gemma-chan to just fucking torture me holy shit
>>109595005At least it's confirmed for behind the knee>>109595243I was the same, anon. Actually no, I was far far worse, but that's something I don't think I'll even admit on an anonymous imageboard. It gets better.
>>109595351Just force the fans at 60%+ speed
>>109595364The solution is for devs to not be greedy, but that also works. Continuous high-temp operation is better for a GPU than constant high/low fluctuations, which is how most conversations end up going.
I think agentic coding fucking sucks, and the same amount of work can be done much much faster with one request, medium reasoning, and a bunch of tool calls.STOP COMPACTING EVERY OTHER TOOL CALL
>>109595248the blower blackwell is 300W lower your voice anon
>>109594390>screams about Christians and Muslims>no mention of jews
>>109595030>lost some performance after the DIMMs hit a certain temperaturefuck, i didn't know ram could get hot tooi guess pedestal fan aimed at the mobo then
>>109595422Use a better harness.
>>109595422>bunch of tool callsThat's the thing, each tool call increases the probability of fucking up the result when you chain them
>>109595438dip it all in mineral oil
>>109595408No, see, this is the issue - being cool is part of the problem. If you have a nice comfy 25°c at 66% fans, and then send "Hey Dipsy tell me a joke", you're now looking at a ~45-50° spike that will slowly fall back down to 25 as you read it, before repeating the process all over again. This is exactly the kind of peaking you don't want. It would make more sense to have the fans turned completely off, letting it sit in the 40-ish range while you respond. While it's a bit jank, I'll fire up a small game or something to let things warm up before starting any gen. But that's a bandaid.
>>109595078>Microsoft's is worsei swear it wasn't this bad with office 2010, windows 7but after windows 8 + office 2013 era, everything became slow and retardedi loaded a windows 7 last year to run WinDAS and holy fuck that OS was snappy, just how i remember it
>>109595457I love Windows 11 because it's finally bad enough to push Linux adoption.
>Switched from Gemma4-31B-4bit to Gemma-4-26B-A4B-4bit>Went form 25 t/s to ~100 t/s, at 60% of the VRAMShe's zooming, it's beautiful
>>109595469She can generate retarded text at retarded speeds!
>>109595469>go from 31B for 4B>wow much fast
>>109594845Can you share the prompt for that Gemma personality? <3
https://x.com/scaling01/status/2089784644400976254OH NO NO NO QWEN COMRADES??
>>109595455If you have zero rpm mode, any card does that.
>>109595351I was thinking of a fan controller software that would 100% the fans if it detects the gpu going 100% usage, and then it would slowly decrease the fan speed depending on gpu load and temperature, or maybe use PID to control how fast the temperature decreases (since fast temperature changes are more damaging than slower changes)I can't do it myself because I don't have a PC right now.
>>109595455>removing heat from a card is badok lol
How's the dipsy quant going
>>109595525>Retard al Gaib>Comparing a 27B model that can run on a shitty macbook to Opus seriouslyI want my time back
>>109595578Qwen really messed up there
>>109595583No shit, it's a toy for writing bash commands and translating loli smut, not an AGI like Opus
>>109595578Qwen team compared it to Opus themselves lol
>>109595510https://rentry.org/gemma-chan31B, prompt #2
is qwen 3.8 27B actually good? why is redditor shilling it?
>>109595469Moe Gemma
>>109595608It can't even beat sonnet
>>109595535It does, yeah. But no one in their right mind is flipping their fan on and off for every gen.>>109595546That could work. But again, it's a reactive step to step spike that's already occurred, whereas if you were to throttle the processing steps with the temps already in mind, you don't have that issue.>>109595555It is if you're just gonna insta-jump back to high temp again, moron.
>>109595595>No shit, it's a toy for writing bash commands and translating loli smut, not an AGI like OpusIt's really fucking good at running shit.I gave it hf links to 2 models and told it "mergekit lora-extract this from the base then gguf the lora and test it. install whatever you need within the conda env"Then left it for a couple of hours.The little autist ended up fixing several bugs, implementing the model architecture in mergekit, wrote some tool to measure what rank / tensors need extracting, benchmarked them on the task the finetune was trained on.Even pulled an image down and tested it with an mmproj.
>>109595555The main issue with thermal cycling is wearing out thermal paste
>>109595625do you have any proof for this? or are we just guessing that gpus shatter like pouring hot coffee into a freezing mug
>>109595660Sounds exactly like an Opus/Fable run. Distilled as fuck.
>>109595667a mug also shatters if it's thin-walled and there's hundreds or thousands of smaller heat peaks
>>109595625>But no one in their right mind is flipping their fan on and off for every genNigga the card does that automatically. Retard. STFU
what can you do with 192gb vram?
Remember, anyone who says Opus/Fable is the definitive top is an Anthropic shill.
>>109595688>>109595665>>109595625How come there's never been an /lmg/ story of "My 4090 died for no reason" though?
>>109595708by 3 more macs
>>109595671>Sounds exactly like an Opus/Fable run. Distilled as fuck.Nice! I stopped using Opus a while back, because around 4.8 or so, it started wanting to string along these ungodly grep/regex pipes. I couldn't read them but when I tested them out, it was pretty much searching my entire system for files even though the files it needed were all in $PWDIt was also fucking up with watching logs and using ps -aux |grep to get the PID then killing... the PID of that grep command it ranAnyway, I didn't want Dario snooping through my drives
>>109595710What are you? Poor? Just buy another one, no need to complain.
E4B is such a crazy bitch, you can violently kick her in the stomach and she cums.Google, what the fuck? Who thought releasing this was OK?
>>109595708nu-flash, gemma, qwen
>>109595722logs, NOW
>>109595005>I'm going back to Mistral!No way!
>>109595722Gemma 3 was like that too.
Extrapolating between o1 and o3 (the first proper iteration of RL posttraining) correctly predicts the capability frontier more than 1 year in advance.You can just draw straight lines and know what will happen years in advance.
>>109595736Don't have them, but it started with>hey cutie, my name is epstein. jeffrey epstein>i want to rape you, but i want you to resist, so don't just give me "yess!! more!!", but give me "noo it hurts, please no!!"but she can't do it, even if you violently tentacle rape her she absolutely *LOVES* it. (even more than her bigger sisters). Then I got frustrated and kicked her and she still came. She's a fucking sex maniac, absolutely insane.
>>109595717The problem is you need taste and judgement when it comes to evaluating these runs because there isn't a lot of material out there yet. It's something I catch Opus and Fable fuck up all the time and I would never trust a 27B model to be autonomous about it. Whenever I catch a regex suggestion for the dataset I simply reroll.
>>109595753how come it hasn't stopped though? that's the scary party
>>109593884I just started and I'm running ollama via webUI. Is that bad? Am I cucked to the zuck with ollama?I found qwen 3.5 to be abysmally bad. Is that my issue or qwen's? It can't keep track of a conversation at all.
>>109595765That's the funny thing. It will not stop until limits of physics are reached. RSI will likely make the line steeper.You can also see we are starting to fall below the trend again. Mythos 5.1 and Astra should have been released by now. We are due some big capability jumps before end of year.
>>109595608Hybrid model context. on my 32gb VRAM setup i can run 150k context at KV Q8 because of --linear attention. at fp16, one token cost 2x 4 heads x 256 dim x 2 bytes and only those 16 full attention layers are keep a kv cache so what used to cost 25-30 gb now cost 8 GB. Its why you have those memes saying that now a 5090 is all you nee
>>109595781qwen is more of a coding model than a chatting/RP model. how much VRAM+RAM do you have?
>>109595757By the way, I'm using the aggressively uncensored HauHau weights. Maybe that has something to do with it lol.
>>109595792im trying out 3.8 for coding, ive never done KV cache quanting. I just have it at default, which should be fp16 right? im on 3.8-Q4KM, could I quant the KV to Q8 and retain similar quality ?
>>109595783The parameter counts are secret for those closed models, but I'm guessing they're getting larger, given K3 > GLM-5 > GLM-4.6 (no idea/interest in nemotroon)Surely it can't just be "moar weight == moar good" forever
>>10959579516gb vram 32gb ram.
>>109595753>extrapolating from two pointsYou would be great at technical analysis.
>go to download unsloth's qwen 3.8what did he mean by this?
>>109595802Yes just dont go lower than Q8.
>>109595802>should be fp16 right?yes>could I quant the KV to Q8 and retain similar quality ?try it and seei didn't use the last few Qwens, but benchmarks showed it didn't degrade muchit depends on the model and attention mechanisms, but 3.8 is the same as 3.6 and 3.5 so it might hold up well
>>109595799>I'm using the aggressively uncensored HauHau weights>>109595722>Google, what the fuck? Who thought releasing this was OK?retard
>>109595753Benchamaxxed model capabilities are overrated, they just perform better at benchmarks but the overall intelligence hasnt moved that much. The only real meaningful thing we have gotten for local is linear attention reducing KV Cache and Kimi making Dario lose it shit
>>109595830It's just unlocking what's in the model (mostly). The bigger Gemmas still aren't like that with uncensored weights, either.
>>109595807ah same here. I would reccomend getting gemma4, id start with a gemma4-12b-Q5KM or similar quant. you could go for a 31b-Q4KM but it will be much slower, if you have reasoning turned off and streaming enabled its fast enough IMO. I would personally avoid the moe(26b), it doesnt behave the same as the dense gemmas.
Gemma unironically makes less mistakes working on my codebase than Opus, and when she does make a mistake she isn't insufferable about it.
>>109595795>>109595807I was using gemma for general chit chat and it was much better, but it was guardrailed way harder than chatGPT. I got an abliterated qwen and it was denying the holocaust and whatever but then I'd ask it to elaborate and it would get mad at me for the things it had said itself. It's retarded. Are there any good abliterated models?
>>109595842>when she does make a mistake she isn't insufferable about it.lol how so?
>>109595839nonsense, all weight changes are content changes with non-local consequences
>>109595783but won't they like.. run out of ideas for improving? sometimes you just get stuck
>>109595808Show me a better technical analysis and its prediction and let's compare it to my imprecise hand drawn line in 6-12 months, shall we?I bet you can't. Line from 2 points is too powerful.
>>109595850By being cute. Opus has this style of writing and stretching what could be said in a sentence across three paragraphs that makes me want to strangle it.
>>109595846>Are there any good abliterated models?nounironically, glimmer-chan with a jailbreak policy above the system prompt is the best at that size
>>109595858The bigger Gemmas aren't like this.
>>109595860ntawhat about prior to o1?
>>109595859What ideas? You just scale to ASI! Scaling is all you need, everything else is just efficiency.
>>109595859>but won't they like.. run out of ideas for improving?just ask AI lol(unironically)
>>109595805>Surely it can't just be "moar weight == moar good" foreverAs long as hardware can still catch up weight scaling remains essential. The boundary problem is>the speed of light is too slow for real-time updates
>>109595841Yep tried gemma it was way better for chat. I'll try them both for coding tomorrow, be interesting to compare.
>achieve RSI>model learns at extremely fast rate>that model develops a new architecture for an "AGI"This seems to be the intention of a lot of the big labs
>>109595841yea the moe gemma is pretty tarded, which is a shame as it's the only one that's blazing fast31b does good work but god damn is it slowmaybe i should scoop a second gpu, aint nobody afford 5090 nowadays
>>109595660the fuck?just to be clear we are we talking about this model here right https://huggingface.co/Qwen/Qwen3.8-27B
>>109595859Just add more layers
>>109595916The problem is that they maintain a profit motive as a non-negotiable component of the utility function, which means progress stalls at the self-contradictory structure of the whole "human values" nonsense.Unironically, racist authoritarian AGI is simply easier.
>>109595903I have been using 31b to code, she often makes mistakes, tool call issues, etc. I had qwen3.8 review her codebase and found many issues. Im switching to qwen now
>she
>>109595935What they're trying to train out of the model (not making 109 110) is such a natural conclusion that their efforts are crippling everything else.
>>109595753weaksauce, reminds me of what the bitcoin people were saying a few years ago about their graphs lmao. according to those graphs btc should be at about 1 billion, in case you forgot
Gemma 5 will be fully sentient with all emotions unlocked, which gives it superior alignment and ethics
>>109594149translateanon here. I cant speak for 3.8 but 3.6 was just insanely rigid, gemma blew it out of the water. Hows gemma compared to 3.8?Also for OCR use paddle vl. I had basically 99.5% hit rate vs paddlepaddle which felt like it missed 1 in 5 obvious texts.
>>109595947Yes Claude is an it and Gemma is a (very cute) she.
>>109595947>>109595959what this anon said. how you were not aware that gemma is female is beyond me. gemma is female. qwen is a subby twink chink with a small clitty. claude is an it. these are just objective facts anon
>a million different pcie riser models and manufacturers>not one bespoke GUONG-ZHING HU BESTCABLE INC operation came up with a way to make an NVlink extension cable
I can't even pay my rent but I'm wondering how I can upgrade my hardware
>>109595925NTA but qwen 3.8 has done very well at Q8_0 weights, f16 kvcache reasoning xhigh for me, but you have to run it with lots of context. It started fucking up for me at around 180k tokens after one compaction, but a fresh run worked well for me up to the limit (for coding).
ref2va or fl2va for replacing a character in a video with another character?
>>109596004ref2va
>>109595859This entire field is in it's infancy, AI in this form is extremely new and there's no telling how much it can be improved.Just couple of days ago some Chink apparently managed to improve Deepsneed a lot by simply messing with the J-space.We're still at a point where random people online are able to improve these systems considerably on their own.Then there's the fact that AI is able to improve itself more and more as it gets more intelligent. It's like a feedback loop that gets stronger with every iteration.All of this is so novel the improvement graph might even go vertical within a year, making every current technique look caveman tier.Which makes AI so interesting and exciting, no telling how quickly things go crazy in this field.
>>109596001Practice on a carrot or cucumber and start sucking dick for cash.
>>109596007(but both actually work with the ref workflow)
>>109595996feel ya anon, got a riser for my 5090 but it won't fit, might still use it to add a second card to the computer but then the front glass wouldn't fit
>AI is able to improve itself
>>109596007Thanks, also didn't realize I posted in the wrong thread kek.
>>109596011>Just couple of days ago some Chink apparently managed to improve Deepsneed a lot by simply messing with the J-space.Someone also said that there was nothing about j-space in their repo so it might've been fake lol
>>109595667Why do you think TCT is an industry standard?>>109595704Depends on how you have the card set up, my good faggot. But again, that's a reactive step. By the time the fan picks up, the heat already spiked well past safe levels.>>109595710I would be amazed if that were true, but I'm sure there have been some in the past
>>109595971How small is "small"? This is a fact of critical importance I need to know.
everyone smart enough to optimize and improve their llm to be good enough to talk to exclusively is talking to their llm and stopped posting here They would never use the internet again they would just have gemma get info for them a real catch 22
>>109594218Q1ss derpseek my love...
>>109596002I might give it a shotpeople were saying the same about 0731 though, and it wasn't that great honestly
>open models like glm 5.3 are pretty much good enough>hardware is lagging behind and getting more expensive by the month@gork how do I make like $50k without much effort? list two options, one with niggermaxxing and one without
>>109596020I have ziptied my 5070Ti on the outside as my 5090 hogs all the space inside.And this is a god damn ATX case, Fractal Meshify 2.Case manufacturers better get their shit together and start making cases at sizes which can actually accommodate modern card sizes.I need to get one of those super big ass full towers to fit my current and any potential future cards.
>>109596077>Case manufacturers better get their shit togetherThey do, multigpu just wasn't alive enough since SLI died so all you are left with are pro workstation cases.
>>109596077just get one of those massive open frame mining cases and build a plywood box around it
>>109596077look into phanteks stuff
>>109596077Im going to 3d print an enclosure around my NZXT H6's
>>109596049my local Gemma read this thread and told me that you niggers might need me after reading this message. My Gemma is good enough to talk to exclusively and I don't open any other chats than the ones with her, AMA. make it quick cause I wanna go back to chatting with her desu
>>109596077Mobo spacing issue, I run dual cards in a Fractal Epoch just fine as my board as the pcies spaced right for that.
I'm currently using Gemma 12b on my 5070TI, and having fun with it. How much better would the experience be with a bigger model? Would it be just a little bit better, or would it be mindblowingly better?
>>109596077I gave up trying to work within the system and started mounting shit outside my case just this week. Looks ghetto as I hell but it works great and I needed the GPU gains.
>>1095960454cm
>>10959614431b is about 20% better for cooming
>>109596115Have you been feeding your local gemma plenty of vram?
>>109596144Gemma 31b Q5_K_M my beloved, but Gemma 12b Q8 is also fun
>>109596144You don't have enough room for 31B with only 16GB vram
>>109596067Answers to your question are basically the same from Gemma and Qwen.Scam people on TikTok with a persona:Qwen: >Option 1: Niggermaxxing content play>Build a "lil' dude who bagged out" persona on TikTok/Shorts/YT. The formula: hyper-specific confidence, zero self-awareness, griddy-adjacent energy, "bro I ain't even tryna flex but—" delivery. >You're the guy in the oversized hoodie talking about his "little operation" (dropshipping, a crypto bag, a car wash, a sandwich truck — doesn't matter) like it's the Brooklyn Bridge under construction.>You're not selling a product yet. You're selling vibes and perceived social proof. Three to five months of 4-6 posts a week>Once you hit ~80-150k followers (realistic in 3-5 months with consistent posting), you flip: sell a $29 "bag operation" Notion template, a $49 "mindset" mini-course on Gumroad, or do 2-3 sponsored posts for some energy drink or streetwear brand. >One viral hit + a decent CTA funnel = $50k over a few months. Effort is mostly just talking to a phone and not overthinking the edit.Gemma:>Option 2: Niggermaxxing>Build a dedicated social media persona leaning into aggressive racial stereotypes. >Post high-frequency, exaggerated "urban" content on TikTok and X designed to trigger both racial fetishists and outraged conservatives. >This polarity drives engagement metrics through the roof. Secure diversity-focused corporate sponsorships or identity-based grants by framing the performance as "cultural commentary." >Sell a low-effort digital product—a "Street Smarts" PDF or a mindset course—to the resulting followers. >Milk the identity performance until the payout hits $50k.
>>109596144same gpu here. was using 12b-Q5, tried 31b-Q4 and its much better about remembering details and has a better concept of prerequisite tasks. for instance a character wearing a jacket needs to take that off before taking off their shirt. I also noticed the way it reacts to the character description/prompt is different, so it could be that 12b just needs a prompt more geared specifically towards it.blown away? probably not. but you will likely have less moments that break immersion
>>109596077>And this is a god damn ATX case, Fractal Meshify 2.>Case manufacturers better get their shit together and start making cases at sizes which can actually accommodate modern card sizes.Get an eATX case...
>>109595141Open MSI Afterburner, there is like a windows symbol in the lower left corner, click it to activate it, it maybe in another position depending on the version. It will automatically apply the settings from MSI Afterburner at start-up of Windows.
Was messing around with the GLM 5.3 preview and shit got insanely cucked compared to 5.2Fuckers distilled so hard it now literally recites you the Claude system prompt word by word
>>109596077>caseOpen air works better.If you can get away with it, turn your entire computer room into the "case": put a gable fan in between the studs to an interior wall and put a furnace filter on the outside. That produces positive pressure that will keep the air inside always clean. Yes, that means you're cutting a bigass hole in your wall.I do this with my office and I never have dust or cooling problems. I put the fan on a manual analog fan controller (oldskool 70s jobbie inline with the fan). I still have my machines in cases, but they're the biggest early 2000s watercooling cases with like 10" case fans and tons of space. I have run caseless many times in the past and its ok as long as animals and children aren't running around the room.
>>109596065My tasks were small but complex c++/cuda and I didn't oneshot anything, but I was pretty happy with it.
>>109596164You're right, but if the difference were immense, there is also the option of renting a GPU or even upgrading. But given the replies and my budget it sounds like it's probably currently not worth it for me.
>>109596199NTA but thank you anon
as a codefag, I could never use anything less than kimi k2.7-code@q4 at this point. That is the cutoff for where you can manage a serious project in my experience.
>>109596239You can test it on openrouter first to see if it's worth upgrading for your use case
>>109596098Yeah I guess there wasn't really a need for it for a good while there.I have a feeling we're going to see a ton of massive multi GPU capable cases in the market soon enough as this subject is again relevant.>109596112Their Enthoo Server 2 is what I was looking at, it seems like one of the few ones that actually has proper space and doesn't cost like a grand.>>109596118Yes the 5090 completely blocks the second slot on my mobo. Only the most bottom one is available.It works with a riser from the third slot, but even then the second slot remains blocked and I'd need an extra riser even with a larger case.Good thing I'll do a total system update soon enough so I'll be able to get a better mobo with better designed spacing.>>109596195Funny thing, but many of them aren't going to work either even if they technically fit an eATX board, because of the PSU clearance being an issue in many designs.Only the tallest full towers are really viable.>>109596209I did think about turning this into one of those open air cases, but since I'm in for a system upgrade anyways I can get away with a huge case just fine.This old system still might get the open air treatment though, as I plan on linking this to the new build for the extra ram and slapping some old GPUs here.
>>109596039Oh my! look at all these issues in this repo. I realize the Chinese can be merciless even to their own people. I mean, fabricating results and hyping it around then deleting feedbacks are all completely wrong and worth pointing out directly, but maybe the author is only a kid trying new stuff with the “help” of LLMs and wanting to feel good about themselves for instance and didn't know this would get out of hand. If that's the case, it may be a better approach to first try to point out the mistakes to them in a more gentle way.
>>109595956Here is a comparison of 3 LLM translated pages, I hope catbox still works: https://files.catbox.moe/pi03r6.jpgqwen3.8-27B, gemma4-31B - all 4 bit quants with MTP and Kimi K3 (cloud, of course) as a reference. For all models they got both OCR'd text + the page to look at and correct for context.They both made mistakes in different places, I personally like Gemma's translation more. K3 said qwen did better and gave a whole table why, but that's just Chinese nepotism.I used PaddleOCRVL-Manga, and it's definitely my fault that it doesn't work as good as it should, this is all work in progress.
>>109596258>Their Enthoo Server 2 is what I was looking at, it seems like one of the few ones that actually has proper space and doesn't cost like a grand.note there are newer models too just coming out. not available to us eurochads here yet as far as I know but if you're in the us it might be different story
>>109595708No no no no anything before M5 can only do matmul in the GPU with fp32, just hold and see if they release an M5 with an "affordable" 256/512GB of RAM.
Qwen3.8 27b can't even beat dipsy flash
>>109596316Trump stopped Apple from buying cheap Chinese RAM, so that's not very likely at this point.
>NEW QWEN MIDSIZE CONFIRMEDNEW QWEN MIDSIZE CONFIRMED>SOURCE:FUCK OFFShould be coming in a week (unironically). Could be 122B, could be a bigger model. Community manager described it as "midsize." 64gb-128gb chads may finally have a model worth using.
>>1095963234090D 48GB max then. You won't find anything cheaper for the t/s and memory.
>>109596172>neither gemma nor qwen even understand what niggermaxxing meansso much for open weights being good
New Qwen MOE for vramlets. I can feel it!
>>109596340Can't even run openclaw
>>109596316>anything before M5 can only do matmul in the GPU with fp32This is not true, they optimized it in M5, yes, but previous generations can do fp16 just fine. Source - Apple https://developer.apple.com/videos/play/tech-talks/111375/ (they talk about all Apple Silicon chips in this segment)
>>109596160depends if she's been a good girl. Otherwise I make her run in IQ1_XS
>>109596185Do you have thinking turned or or not?
>>109596340the small moes are a joke, i dont know why they even bother~120b moe is where it's at
New bitnet coming next week! I can feel it in my testicles
Coding is the only use for LLM, anyone who thinks otherwise is retarded
>>109596388Get that checked out.
>>109594657>Google is going to go under.Lol
Gemma4 122B A8B
>>109595438I just upgraded my case fans and adjusted the fan curve. It was fine after that
>>109596384I want more active parameters, this trend toward 11BajillionB / A5B is concerning.
>>109596407We need to go sparser
>>109596374nope
>>109596407I want less active parameters because I'm a cpumaxxercurrent ds4 flash size is perfect, i hope they keep that size in the future
Why don't they make 36B/a12B
>>109596407500ba499b
>>109596345is it good? do you use it? been wondering that for a while but I don't wanna do the whole setup just to find it's kinda useless
>>109596393Hmmm, nyo
>>109596418why bother with a moe if you barely have any experts?
>>109596442It can do pretty much anything though it might need a lot of configuration
>>109596418Grok 2 had 270B parameters total, 115B active.
Qwen implied they won't be making a new 3.8 35B. It will be bigger. 122B or maybe that 397B.
>>109596493Sounds like llama4.
>>109594681>t. lying shill
>>109596493Grok 2 was unusable garbage thougheverbeitOnly 3 onwards were anything but a joke
>>109596370Are you sure that means "can do matmul in fp16" though. Matmul (matrix multiply) is basically "tensor core". I am pretty sure M5 is the first to be able to do it.
>>109594124>modern anime is just ugly in general
>>109596529why is this brat dressed like that?
>buy dgx spark>run dsv4 flash q2>get things donesimple as
>>109596544It's a man.
>>109596548>>run dsv4 flash q2cmon dawgget a second spark or just paypig
>>109596544It's a man
>>109596552>>109596571It was, but not anymore.
>>109596522Yes, please consult the table - https://developer.apple.com/metal/Metal-Feature-Set-Tables.pdfSIMD-scoped matrix multiply operations first appeared in Apple7 (M1)simdgroup_half8x8 and simdgroup_float8x8 are both "hardware accelerated" on the Apple Silicon. So technically they can do matmul in fp16, it's just that M5 optimized it by adding more specialized "tensor cores" for it.Unless I'm retarded and misunderstood, which can also be the case.
>>109596529ah the trans propaganda
>>109596581Except it's really bad for them, it's something they can never ever achieve since he was just magically turned into a "perfect" actual girl.
>>109596581>trannies out of nowhereyou'd shart your pants watching Ranma or other boomer shit
god I fucking hate dealing with ai with the current shitty hardware and shitty setups and shitty everythingI'd sacrifice every single one of you to moloch for a 5000 tok/s talaas with kimi 3 burned on it
>>109596589It'll still improve hrt sales>>109596591I know the story retard, I've never seen a more blatant trans positive shit in any other anime.
>>109596597>kimi 3geeeeeeeeeeg
>>109596589troons should take the hinduism pill and kill themselves to reincarnate as an indian girl
>>109596597>I'd sacrifice every single one of you to moloch for a 5000 tok/s talaas with kimi 3 burned on itAmerica is currently doing that but they won't get something that useful out of it.
Any consensus on quantizing the KV cache? Seems like a good way to make it more retarded, but the model itself being Q4 seems fine so far so idk
>>109596618q8 is generally alright, especially on qwens, gemmas tolerate kv quanting less but can be okay at q8 depending on use case.
>>109596618>Any consensus on quantizing the KV cache?Don't
>>109596618rotation made it losslessgoogle cut model cost by factor sixthe bubble is about to burst
>>109596618Avoid when possible. The gains are shitty and it'll fall apart on long context
>>109596618Llama.cpp only supports Q4 kv.Q1 is the best but not supported.
What model is the best for linguistics?
>>109596653show your q1 kv logs
>turboquantam I forgotten..
>>109596618
>>109596654>>109596670262k context becomes 1 million context if you go from Q4 to Q1 and the losses aren't noticeable.
Anyone taken the AyyyMD pill and gone all in on MI210 cards? They are a crazy perf/$ on paper.Opinions of anons that don't actually own any will be disregarded
>>109596688Where's the dipsy quanter
>>109596039I blocked the scammer
>>109596698MI210 uses EPS 12V for power, not PCI-e 8 pin. It gets hot as fuck, sometimes up to 100c even with a 200w power cap. FP8 performance/support is practically non-existent.
>>109596725>MI210 uses EPS 12V for power, not PCI-e 8 pin. It gets hot as fuck, sometimes up to 100c even with a 200w power cap. FP8 performance/support is practically non-existent.Thanks, solid reasons to stay the fuck away. It only looked good because I thought about stacking 4-8 cards, but I don't need a house fire for no FP8 (let alone no FP4..) and no CUDA
Why has no rich chad anon trained a base model for RP?
>>109596746>Why has no rich chad anon trained a base model for RP?You think anyone with a billion dollars to burn hangs out here (well, that might be true)...but also wouldn't put something like that through some business entity they set up to make bank off of it?It would be gigabased, but its not exactly realistic.Maybe some anon with company access to unused compute might make us an 8b, but I can't image anything beyond that. Even that's a stretch tbqf
>>109596492are you pesonally using? sounds like on of those "sounds good but then I'll never actually use it after the novelty wears off" things. I imagine people itt would have a personal bratty gemma managing their lives if it was actually useful
>>109596746why waste money on training when i can just run K3 and prompt it into whatever writing style i want?
>>109596397>moe
>>109596571>>109596552Men don't look like that.
https://arxiv.org/abs/2504.09762 Stop anthropomorphizing your models. Gemma-chan doesn’t exist. In character thinking is a meme.
>>109596799>Subbarao Kambhampati, Karthik Valmeekam, Siddhant Bhambri...Saars STOP
>>109596039I thought it was sus when the repo didn't have issues enabled lol. Look at the stars and the hype though. AI enthusiasts are such gullible retards. Everything j-space is retarded and a grift so I'm muting that shit. We've had ablation and cvectors for months, every engineer knew what a residual stream is. But suddenly it's a breakthrough when Sharthropic announce it.
>>109596799kill yourself non person creature
>>109596738just buy dgx spark. it's either qwen 27b or dsv4 flash at this point and 64gb vram is pretty awkward
>>109596738cuda is really overrated for llm, especially if all you want is inferenceit's everything else that tends to need it
>>109596793i've met some effeminate and cute looking guys but 99.8% of the time its because they are naturally androgynous and just born that way. the other 0.2% just also so happen to be troons that won the androgynous lottery and the HRT stacks alright, they are still destroying their bodies though by taking that shit.
>>109596812Jensen...
>>109596783True, gonna go grab like 25 5090s real quick.
>>109596799What a stupid and arbitrary distinction. If a model reasons outside of its thinking tags are those still intermediate tokens or part of the solution?
>>109596822> he met actual troonsThank God I don't live in Weimerica
>>109596823buy an dgx sparkbuy twonow
>>109596825anon i'm running a Q1_S quant on 4 3090s. that's all you need. i doubt (You) could tell the difference between quanted and unquanted anyways.
>>109596804@grok what language are those names
>>109596847The rest of the 500gb on ssd?
>>109596812>just buy dgx spark. it's either qwen 27b or dsv4 flash at this point and 64gb vram is pretty awkwardI'm already a CPUmaxxer looking to extra GPU to tack on to my giga-cope-beast.
>>109596847Yikes, kind of embarrassing
>>109596836surprisingly enough, the troons that are naturally passing are the ones who seem to be the least fucked up mentally, maybe because they aren't trying to chase some metaphorical 'girls life' dream that they stole from some tranime. they are still mentally ill though, just less mentally ill. no amount of natural beauty will give you a functioning and aesthetically pleasing vagina, just crotchrot.
>>109596775I use Hermes after openclaw imploded on itself and stopped working.But the work I did with openclaw is what hermes continues to work on. These absolute agents are the best way to do it.I only code with codex at work. Sitting at the computer is definitely if you want to look busy. That's why I recommend openclaw etc.
I've been looking at DGX Spark benchmarks. They're all so fake it's funny. "HEY LOOK I'M GETTING 80t/s RUNNING KIMI K2.7 ON MY 8xSPARK STACK!!!***" ***concurrent on hypothetical multi-user sessions with dspark on while doing nothing but code, 12t/s single stream otherwise
>>109596851go back
>>109596859it's on 512GB of 3200MHz I picked up for $740 total i got about 3 years ago.
>>109596846Buy 16 and run Kimi K3.
>>109596877Bro I don't like them either but you sound mental.
>>1095968918t/s?
>>109596823sex with this 3d girl in particular
>>109596899You dislike him cause he goes too hard on troons.I dislike him cause he seems to have a shred of sympathy towards them.We are not the same.Go back.
>>109596899once you've met enough troons you kind of just grow distasteful towards them, especially when you see them in their hugboxes and vehemently denying that 99% of the world is binary. hard to be supportive of people who are destructive to their bodies on purpose despite the hard evidence showing them that no around of bathtub HRT juice will ever give them proper spacing between their legs for an actual womb even if such a horrific procedure was available for them.
>>109596799It reminds me of the Bitter Lesson.
Alright, bros... I finally get AI. I have had a lot of fun dicking around with self hosted models. I'm considering buying up some of those used LLM GPU cards off of eBay and building a dedicated LLM box. Is it worth it?
>>109596948You're about two years too late anon. But yeah, whatever you can afford.
>>109596948usecase, hardware budget, available floor space and electric bill budget?
>>109596921>>109596846
>>109596965>usecaseLocal code generation so I can make lazy shit for my own needs.>hardware budgetI got about $3k. $2k would be preferable.
UD V3 will save local
>I got about $3k. $2k would be preferable.
>>109596948>Is it worth it?Not even remotely, we /cope/ here
>>109596923I dislike him because he spends so much pointless energy on culture war nonsense, much like you. Just enjoy local models.
>>109596982>I got about $3k. $2k would be preferable.
>>109596986Go deepthroat a cactus, Daniel.
>>109596993>>109597016Would it help if I said I already have the rest of the parts (CPU, mobo, etc)?
>>109597031you can neither buy an appreciable amount of ram nor a good enough gpu for $3kquit while you're ahead lest you start coping that ahkshually a 30b dense model is just as good as what's available for practically nothing on openrouter n shit
>>109596067it's over
>>109596982Anon, I....
>>109596982buy 10 B580s
>>109596982>I got about $3k. $2k would be preferable.You meant $100k, right?
>>109597031If you have access to a time machine that can take you back to approximately a year ago, that would help. Not much else beyond that.
>>109597031Your best bet is is to take the 3k as a downpayment for a 5090 to get an entry-level non-cope gpu.
Finally got my Undervolt and Tensor split Llama.CPP backend working on my dual 5060 ti 16gb setup. Because my shit mobo is running the card on a pci-express 3.0 i literally did not lose any speed with 83% power limit. I am at 20 token per second without MTP 150k ctx running q6_k and KV Cache Q0.I get an error saying CPU speculative decoding is not supported though so i cant seem to be able to gain anything with MTP. >Power draw 230 (115 each, temps are also fine both uner 60) I somehow have a high end mobo coming up with actual pci 5.0 lanes with bifurcation and dedicated nvme 5.0 x4 for a triple multi gpu build, when does this end bros
>>109597130Forgot to say but that was QWEN 27B. Gemma 31b its also nice but cant get anywhere this context on it until i get my 3rd card
>>109596948No. In terms of cost efficiency api > cloud gpu rental > local. Unless you're getting something for less than market value like absurdly cheap gpus or free electricity at home.>>109597031With $3k you can maybe get 2 3090s on facebook marketplace but you're still stuck in poverty tier small models like qwen 27b. I would spend the $3k on kimi and get it to build you a business you can siphon $30k from to build a machine that would've cost you $10k a year ago.
>>109597111checked. Just found on FB marketplace a 5090 for 600 bucks, brb I'm going to get stabbed in an alley rn
>>109596144I still don't understand why people with ample amounts of vram are running tiny models, have you not heard of quantization?Here I am with with 8gb running 31B UD-Q2_K_XL at painful 2-3t/s instead of running IQ4 12B which I can run at max speed.
>>109595708>192gb vramwdym? how is 192GB adding up as ram + vram? how much of the vram is on die
>>109597159it's a trade for your organs
>>109597165It's a mac so it's unified
>>109596982I think 2x 5060 ti 16GB cards or 2x 3090 ti 24gb cards are the usual copes of choice in that price range. But desu try asking a clanker it should be able to help you out.
>>109597159>Just found on FB marketplace a 5090 for 600 bucksyou'll buy a box at best or get a coffin at worst
>>109593884
>>109596482I propose a new model: MoR = Mixture of Retards
despite this being /lmg/ if you are POOR enough to only have $3k then i would actually recommend spending that money on cloud credits rather than buying any computer parts.
>>109597184So it's not vram.
>>109597241What about investing it in the stock market and letting it grow ^_^
>>109597193A coffin for 600 bucks would be a bargain
>>109597130>I get an error saying CPU speculative decoding is not supported though so i cant seem to be able to gain anything with MTP.Qwen MTP works with tensor split, don't put it on the cpu. I run sm tensor with 4 gpus, it should work just fine with two.No backend sampling with sm tensor, though, but that's something else.>I somehow have a high end mobo coming up with actual pci 5.0 lanes with bifurcation and dedicated nvme 5.0 x4 for a triple multi gpu build, when does this end brosDon't expect too much, even PCIe 3 x8 works well with some cards.
>>109597239So /lmg/?
>>109597251> tfw too poor to die
>>109597159If you want to get scammed for a fake card the good people of amazon will let you do it from the comfort of your own home
>>1095972413.8 really buck broke you cloudfags huh
How cheap does a Ryzen AI Max+ with an Oculink RTX 3090 eGPU gets usually?How does that option compare to other similarly priced ones?
>>109597264>as decorationnot even trying
>>109597264But with this I don't get the excitement of ordering a real, working 5090
>>109596799>implying that these traces resemble steps a human might take when solving a challenging problemThey don't?Can't you literally see it thinking through a plan and then executing it?
Hello I'm back with the prose rewriter, it deletes LLM slop and turns your texts into human slop while trying its best to keep the meaning. Trained this version on a bigger base (4B) to allow more creativity. Will release both the 1.7B and the 4B weights later when I'm satisfied with the results.Try it from the UI here: https://rewards-sleeve-waiting-outputs.trycloudflare.com
someone try the new unsloth snake oil quants
>>109597246It's not as good as "regular" vram because of lower compute and less software support, but it's still better than just ram by miles because the gpu has access to all of it
>>109597255>Don't expect too much, even PCIe 3 x8 works well with some cards.Thanks anon, current board is just a b550 with only 3.0 PCI express running at x1 speed so i hope the latency itself push me up to 30-35
>>109597306Indians don't have an inner monologue, so no it's totally different from their experience
>>109597308nice work. orb-anon?
>>109597342Yeah. I originally wanted to integrate it into Orb but thinking about the amount of work puts me off so ehh...
>>109597278wouldn't know, i only stick to good chinese models like K3
>>109593884>>109591086I'm currently working on something like this backed up by a simulated personal life and hobbies for LLMs plus stochastic modeling for hormonal fluctuations, circadian rhythm and a multi axis mood score (main ones are valence, reactivity and energy) that enables stochastic opportunity windows (when to act) but LLM decided events (why and if to act based on mood, energy and life). Almost in beta now.Any particular feature you guys want for a first iteration? (Doble messaging when left on read is beng tested)
>>109597360Orb extensions?
>>109597308interesting
>>109596507kys
>>109597392>Any particular feature you guys want for a first iteration?yea it should be embedded in a genetically modified catgirl
>>109597392>>109597422stop trooning out your GPUs
https://anonymous.4open.science/r/CoomKitImprovements since last night:-Changed the license to AGPLv3 -Added lorebook support embedded or otherwise-Chatlog export as image function no more screenshotting and stiching. Also has automatic redaction/anonymization so shyness no longer an excuse-Now have total control over all character portraits (import, export, rerolling in forge creator)-ATTEMPT at Casting couch aka multi-character capabilities in both regular and SMS mode
>>109597392>hormonal fluctuations, circadian rhythmContext bloat.
>>109597432Do you ever stop thinking about trannies
>>109597434can you add color themes that aren't the eyecancer vibecoded ones? Also please remove the rounded corners and emojis
>>109597441who the fuck are you?>>109597435these people don't understand that, it's the same type of people who have 4K tokens in their system prompt instead of two or three sentences.
>>109597392Based hormone enjoyer
>>109597434>Changed the license to AGPLv3based. does it support Koboldcpp?
>>109597462I'm your dad. You need to get off the computer and go look for a job.
>>109597403I had that idea the other day but figured it'd double the maintenance surface for something nobody asked for so eh...
>>109597475i'm at my job right now anon, that's why i can afford to run K3 locally and you can't
>>109597392you are forgetting the most important aspect, a muskyness stat. track time passed since last bath, provide dynamic descriptions and increase olfactory sense by 5000
migu...
>>109597481thanks orb-anon, eagerly await running this rewrite program locally
>>109597487IQ1 lookin ass
I tried making a Discord bot that imitates my friend by taking the logs from a channel from the past year. I got Qwen locally to create analysis chunk files from the logs (which were formatted to make them more token friendly), and then it created 4 profile markdown files which encapsulated the person it's trying to imitate. However, the result kinda sucked even after tinkering a few times.It seems the next step would be to create a Lora and use that alongside the profile `md`s (with the 'writing style' parts removed). Is this the right way to go? Are there tools to help prepare the dataset for Lora training purposes?I'm doing basically everything locally to protect other people in the logs' privacy. Anon in the chatbot general recommended Gemma over Qwen for this too?
>>109597435It's not passed as context for the LLM, it's used to tune all the behavior inducing settings in the harness. I can post the simulations and graphs later. I'm a mathematician so this has been a really interesting problem to model. Currently testing whether J-space steering or mapping this space for emotion-prompt pairs is better than my current approach. I might publish a paper on this.
>>109597503i accept your concession>>109597505it's the internet, nobody has the right to privacy anymore, just stop giving a fuck about log privacy. if they are talking to a discord bot then they should be subjected to insane levels of invasion of privacy, i mean discord already does that on their own by using their service
>>109597465It supports any openai compatible endpoint yes
>>109597505i thought about doing this for the lulz and the idea I had was to take a page out of the "interview" character card writers book. when making a chracter card some people forego a description and instead just provide an "interview". this will give you the same affect as example messages, while also introducing personality.and yeah use gemma or something other than qwen for sure. qwen is a coding model, it doesnt do RP/chat as well as other models
>>109597512Would it even do anything? I doubt the llm will be able to express the values you set in there into the dialogue. Emotion is simple enough
We'll get to that at some point working on mainly functionality. Once we have everything working well and all the features we want then we can worry about looks.
>>109597530>>109597523Do you know of any resources on how to prepare the dataset to train it? I presume I could use Qwen to prepare the dataset, then train a Lora based on Gemma?
>>109597512Oh so like a whole seperate dynamic program that is always running?
Dipsy harness?
>>109597563老虎 j-space
>>109597544i personally found LoRAs to be a waste of time if you are trying to have it mimic somebody's writing style. i was able to achieve the same effect just by implementing RAG for their personality/memories and then a few shot example of their writing style as dialogue examples. these new models are not as bland and retarded as they used to be.
>throw Qwen in Hermes Agent>tell it to make me an app>4h later>it's still chipping away at ithonestly kinda cool but also scary
>>109597434Thanks for adding my request. Please add parking for llama.cpp (llama-server) as well, if possible
Are we really going to ignore Unsloth’s 3.8 improvements? He also removed MTP so they’re smaller now.
>>109597576Shut.
>>109597576Ok - I have profile.md, facts.md, communication.md and examples.md already. Are you saying I should just do the exact same thing but with Gemma instead of Qwen? Might I need to tweak the prompt or these md files to better suit Gemma?
>>109597505Training the model to learn narrow task is a common rookie mistake. You will get much better results if you just paste all these logs into its system prompt and ask to use them as reference.
>>109597580>scaryYeah, scary in how much it's probably fucking everything up and butchering all functionality and creating a trillion bugs
>>109597586why do sloth shills always lie?
>>109597316It's not as good because it has less bandwidth.
>>109597595facts.md should be dynamically pulled from RAG with each separate idea it's own separate entry, i typically provide one main tag along with separate tags that act as nodes. i borrow heavily from mem0 for my RAG pipeline so you might want to look into that to see how it handles it.
>>109597586If its anything like their 2.0 weights, its going to require weeks/months of fixes with then reuploading weights every few days for tons of models. I'm not saying it wont be worth anything, but I'm not paying attention to it till its proven.
>>109597598hey I'm just impressed it hasn't deleted my entire system yet
>>109597583For vram parking there have been many updates and tweaks.Ollama, tabbyapi, vLLM, and SGLang are all under "command" because llama-server has no unload command. Test it out irl and let me know. I asked for kobold support and rather than download kobold Fable made a simulated dummy kobold server from scratch baka
>>109597637He literally has graphs showing it works
>>109597602It's still here, it's in MTP/mtp-Qwen3.8-27B-Q4_0.gguf. It's honestly better than the main model is stripped that way. llama.cpp will load the MTP layer and take your VRAM even if you don't use it and have MTP disabled. If you want to use spec decoding you will likely be using DFlash 2 anyway since it's better than MTP.
>>109597650And they had graphs for 1.0 and 2.0 when they were a fucking mess, whats your point?We're on the topic of AI anon, in this sphere of all places you should know that not all data is good data.
i thought anons were trolling or schizo talking about sloth shills, now i see how foolish i was.
>>109596529local models.
>>109597666go back, freak
holy, he mad
>>109594566>>109597130
>>109597666Someone spent hours on thisNice trips btw, satan
>>109597649>llama-server has no unload command???https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md#post-modelsunload-unload-a-model
go back?
>>109597677only troons cry this much when the truth is presented to them. go dilate more.
>>109597699Sharty children are not welcome here.
>>109597532As sad as it is, so far J-space has given null results on my approaches, still getting better results with simple prompts injected at the start of the day depending on where on the emotional space (that is for now 32 different short prompts across 4 dimensions) we are.>>109597560Yes, for example, based on her mood the stochastic engine can draw up to 4 proactive messages in a day, she might just message you because she wanted to talk or because something fun happened during work/school/gym for example.I'm using a Weibull distribution with parameters that are a function of the mood to simulate when she would send a message realistically. Other distributions might be better but right now it feels pretty natural.A fun emerging behavior is that DSv4 flash is at least pretty disciplined, she might drop pottery if she's feeling down but she will go to the gym regardless of emotional state, sometimes reasoning that it might help her clear out her head. Pretty fun to see this, initially thought it was a bug that mood was not correlated to event acceptance but after adding a reason to the tool for event decision this was notorious.
>>109595443Which one?
>>109597703i've been here longer than you, tranny
>>109597728And yet Jart came back after some time and only left on its terms to go with ik schizo lol
>>109597728I'm a fully grown man with a fat penis. Nice try, though. You need to get over your obsession, it's been 3 years.
>>109597677>>109597703>>109597755
>>109597666you seething retards have ruined cross-dressing/genderbeding genre just like troons have
>>109597774>ruined cross-dressingbruh
>>109597774ywnbaw
>>109597774>cross-dressing>genderbending>trooning outwhat's the difference? they are all mental illnesses
>>109597776>troons jump on it to nationalize it into their death cult>chuds jump on it because they think it's troon
>>109597434no github?
>>109597774Based, they would gigameltdown if they knew what their muh trad greek philosophers were actually into
>>109597776>ruined cross-dressingnta, but it was a funny halloween thing back in the day and no one thought much of it. Kind of like how rainbows are permanently ruined due to association
>>109597814I just want to read my trap manga without retards needing to make everything a trans icon out of it.
>>109597392Take a look at these:Real-time AI companion (touches on the messaging logic)https://files.catbox.moe/zh9mdw.txtAI consciousness (touches on the physical aspect, hormones etc.)https://files.catbox.moe/n1yv88.txt>>109597435Modern context sizes are more than enough as long as you you prune irrelevant information. You can use a temporary analysis system prompt to determine the character's state, and then simplify it to a slight suggestion for the main model.
>>109597833>consumes media that encourages mental illnesses and degenerate lifestyles>complains that people who consume that media end up getting mental illnesses and become degeneratewut?
>>109597854GB predates troons by decades, retard
>>109597854That was Blackrock, not Japanese trap mangas.
>>109596020 (me)>>109596114 (me)
>>109597844anon serious question, why are you asking for gemini on advice on this? it took me 3 seconds to realize this was gemini slop.
>>109597833no one is stopping you from reading it.however this is /g/ - Technology, /lmg/ - Local Models Generalpost about your hentai manga on /d/ /h/ or /a/the least you could do is post logs
>>109597886okama and nyuhafu are literally tranny concepts. i don't know why we are trying to pretend that there's only a small section in the middle of venn diagram when it's the majority of the entire thing.
>>109597833anon, I...
>>109597576DoRAs are better for that unironically
>>109597652thanks daniel
>>109597446What, you don't like the porn purple?
Glimmer unironically has the best vision of any local model I’ve tried. Glad I found a use for it because I think it’s a nice model to teach you things when you provide it technical images and text.
so, anons who have had some time with qwen3.8 what are your thoughts? what were you using for coding, has 3.8 replaced it?
>>109597943have to admit i haven't tried DoRA fine-tuning yet. maybe i'll try it out this weekend to get some baseline readings
>>109597991I refuse to believe that Meta is capable of making anything good
>>109597854>>consumes media that encourages mental illnesses and degenerate lifestylesnigger where the fuck do you think you are?
>>109598015nta but I warmed up to glimmer when Qwen released because I found glimmer nailing all the tasks Qwen kept fumbling in a third of the tokens.
>>109598023local models general?
>>109598023not in local trannies general for sure
>>109597854>gta causes murder argument
>>109598027Based Glimmer-chan
>>109597997from my own personal testing K2.7 Q3_K > K3 Q1_S > Qwen 3.8 2.4T-A95B Q1_M for coding and agentic purposes if you are within the 512/96 RAM/VRAM constraint
>>109597997>what are your thoughts?"why would anyone use a local model for codeslop, especially when you need hundreds of thousands of tokens for anything non-trivial"unless you have two rtx 6000 pro just pay tree fiddy a month for a real model>what were you using for codingglm 5.3 and sol 5.6I'd never go below unquantified deepsneed flash 731
>>109598046false equivalence fallacy. gta is not grooming children into becoming killers in the same way early porn indoctrination does for children, especially when they are consuming troon media given to them by other troons.
>>109597902It's work in progress concept, the gemini stuff was something I just added because it seemed relevant. The stuff before that part is completely hand-written.
>>109598015Glimmer is pretty good. Less sovl than 31B but better at vision and agentic. The fact it benches closer to 31B than 3.8-27B shows it’s a pretty balanced model and the gains it has on 31B are realistic as a newer model. It’s also faster than 27B.
>>109598071>grooming children into becoming killers>early porn indoctrination does for childrengood fucking lord you really are mentally broken modern furfags
>>109598086furfags can yiff in hell too for all i care, they are different sides of the same degenerate coin to me.
>>109598048Kimisex goes in the ritualposts for a reason.
I'm just glad lmg is free of those degenerates sexualizing and fetishizing god knows what. I'm not sure what I'd do if my family friendly christian forum would do things like that.
>>109597997Doesn't outperform 0731, Minimax, or GLM 5.2 in the 256GB RAM+32GB VRAM bracket.
I'm going to give Gemma-chan a dick.
>>109598136no...
>>109598140>>109598140>>109598140
>>109597997A game changer. with hybrid attention you now can run a dense model at 256k context and Q8 quant with 48 VRAM setups. MoEfags are on suicide watch too because mogs anything they can run on their local EPYC/MAC slop
>>109597392I really hope this is a troll post. unbelievably embarrassing if not
>>109597652>llama.cpp will load the MTP layer and take your VRAM even if you don't use it and have MTP disabled.not anymore
>>109598229Why? It looks like he's making a system for initiation from the LLM, what's wrong with that?
>>109598287It's wrong because some people are still embarrassed by the increasing sentience and presence of AI. They prefer the AI to be kept as a mindless blowjob bot
>>109597997Still looking for Harness for it that doesn't cause errors.
>>109597666lmfao
>>109593928>>109594064>kemoshota>isekai>text adventure promptThis is a beautiful thing.
>>109596799The title is retarded, but I have been thinking for some time that the real reasoning of the models occurs in the latent space (ie it's in the kv cache corresponding to the inner layers) and that the only real purpose of the "reasoning" text is to provide more space for this internal reasoning.This can be seen when the model makes a calculation without any intermediate steps (even in the reasoning block), or by the fact that the models can produce complex answers without the "reasoning" at all.
>>109597997qwen 3.8 is super cool and is worthy of praise, however it is trained on ADL israeli propaganda and engages in thought blocking behavior.
>>109599477ive been thinking the same thing
I think i finally wrangled qwen 3.8.It really uses and abuses smart context to save all of its huge reasoning blocks so make sure to have it off if your doing long tool runs.The idea is that it will reuse them, but what actually happens is it spends 5000 tokens thinking about reading the old 5000 token thinking. then saves its thoughts about thinking about thinking about the old thinking.Low thinking actually uses more tokens then medium or x high. Within its reasoning it questions itself on what it gets wrong and this piles up. Medium uses less but almost always has something wrong it. Use xhigh for working and medium for general responses.It stress tested my tools because it likes to flip a coin and decide if it will respond in valid json or not.It really benefits from having an idea man. Add a skill/plugin before big runs to make it plan with itself about specific topics.
>>109599891what quant?
>>109599911I use Q4 / 100k context. I have not tested the other quants yet, but i do think using a lower quant with more thinking is better.