/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109419138 & >>109415437►News>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109419138--Optimal conversion paths and mxfp4 performance for DeepSeek-V4-Flash:>109420274 >109420339 >109420368 >109420426 >109420464 >109420513 >109420828--WASTE engine allowing trillion-parameter models to run via disk streaming:>109421076 >109421134 >109421168 >109421226 >109421251 >109421271--Gemma's repetition issues and discussion of MCP tool calling servers:>109419410 >109419480 >109420486 >109420110 >109420277 >109420292 >109420340 >109420381 >109420393 >109420530 >109420822 >109420908 >109420879 >109421037 >109421085--Debating Gemma QAT performance versus standard quants for coding:>109419770 >109419785 >109419800 >109419826 >109419851 >109419853 >109419918 >109419960 >109420016 >109420199 >109420317 >109420024--Methods for forcing models to think in-character within reasoning blocks:>109421852 >109421867 >109422189 >109422212 >109422223 >109421921 >109422094--Debating causes of performance degradation in quantized abliterated models:>109421385 >109421397 >109421403 >109421413 >109421511 >109421681 >109421943 >109421417 >109421402 >109421646--Comparing pi_agent_rust to original pi regarding security and performance:>109419786 >109419809 >109419895 >109421730 >109421754 >109421874--Using a system prompt to simulate consciousness in Gemma:>109420365 >109420383 >109420413 >109420473 >109420496 >109420649--Comparing Kimi-K3 and DeepSeek performance with low-bit quantization:>109419167 >109419185 >109419609 >109419653 >109419775 >109421130--Logs:>109419313 >109419473 >109419713 >109420345 >109420751 >109420365 >109420927 >109422012 >109422021--Miku, Kimi, Gemma, Mちゃん, Dipsy (free space):>109419192 >109419274 >109419487 >109419730 >109420166 >109420485 >109420729 >109421641 >109421788 >109421883 >109421947 >109422004 >109422020►Recent Highlight Posts from the Previous Thread: >>109419144Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
slowpoke.jpg
>OAI dropped prices by 80%>still wasn't enough to compete with China>temporarily dropped the new discounted price by an additional 50%>still more expensive and V4F has the same intelligence and is uncensored and providers can compete for cheaper prices because it's open What can the US do?
>>109422958ban chinese models
>>109422958But don't you know, the bitter lesson says you shouldn't put effort into optimization, it's just going to be wasted work once AGI comes out!
dsv4 flash “preview” called itself Claudedsv4 flash 0731 calls itself Geminithis happened for anyone else? seems that all the Chinese models have reliably been calling themselves Claude but now calling itself Gemini would be odd
>>109422992distilled from Gemma5-70B
Could be the next big thing:https://explorative-modeling.github.io/https://arxiv.org/html/2607.27372v1https://x.com/AlexiGlad/status/2083230922196107288https://alexiglad.github.io/blog/2026/explorative_modeling/>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation>>We introduce Explorative Modeling, a new paradigm for generative modeling that acts as a third pretraining axis when added to existing generative models, and also enables end-to-end generation. Increasing exploration monotonically improves existing models across images, video, and language, and the gains grow with scale (7%->36% with data, 13%->23% with parameters). Concretely, Explorative Models (XMs) reach 6.2x sample efficiency, 4.1x FLOP efficiency, and 47% better parameter efficiency. Exploration also enables scaling generalization, and scaling how end-to-end existing models are. As end-to-end generative models, XMs match diffusion on control tasks with up to 256x less inference compute.
>>109423023E31B (70B total)
>>109422892where do you put it in though
>>109423041what does this mean in plain english, nerd??
>>109423064The valley he's talking about obviously.
>>109423077This sounds like PPO from first principles.
>>109423077For every sample, they pick K candidates and train the best performing one. It seems to help in particular with continuous data; will probably help a lot with MTP, JEPA, etc. Simple next token prediction, not so much, according to info on X by the author.
I was checking out the new DSV4F quants and noticed that Unsloth's IQ2_XXS and IQ2_M have basically the same filesize. I checked the tensors and only one of them is different between the two.lmao
>>109422958>ai is more expensive to run due to hardware and energy costs>they still drop prices despite thatHow do they make this work?
>>109423133a few weeks ago people were talking about some memory usage breakthrough at OAI. this is probably them implementing it.
>>109423077it's>Wait, actually!for diffusion
>>109423144complete coincidence that they implemented it same day as v4 flash, right?
>>109423123can't LARQL check what that tensor encodes?
>>109423041Big win for JEPA actually.
JEPA-space soon.
>>109423144that was about their free users who don't even have accounts and were costing them millions every month
I haven't looked into local LLMs for a long time and am looking for a tiny one to do some light text cleanup. Are there any that are even remotely usable for less than 100MB to 200MB?All it has to do is make small corrections on no more than a paragraph of input and not fuck up all the time. I fully expect it to be a little stupid at that size, but if it works, it works. I would appreciate any recommendations.
>>109423257Anything that small will need to be finetuned for your specific task to have any shot at working.
>>109423257https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUFnot technically what you asked for because the answer is no but this is the best you'll get or >>109423268
>>109423257that sounds like you need a spellcheck, not a language model
70b dense
Why are there so many 35B troons out there. Why that model specifically.
>>109423257Nah, go back
https://www.nature.com/articles/s43246-026-01307-6?error=cookies_not_supported&code=bad04f3e-7683-443c-92c0-00927194bc1dOk faggots, how do we make our own memory at home with mud? They’re doing this in Manchester and India so it’s probably not too hard. Mass production in two weeks?
How do I optimize a local llm model to be an expert on exploring AI goon gen stuff accurately? When running out of the box they are very lame, dumb on actual workings, outdated and clueless about last release and uncreative to brainstorm possibilities. Ideally it should be knowledgeable on comfy UI and its shenanigans like good workflow combinations, recent-ish model/lora releases, lora training, maybe even being able to crawl civitai.Maybe some elaborate initial context + RAG contraption? I never messed much with those.
After trying nu-flash it feels like it is not even a sidegrade but an actual downgrade when it comes to cooming.
>>109423454>Cost-effective and scalable fabrication of solution-processed vermiculite membrane memristors is attractive for sustainable integration of nanofluidic memristors for neuromorphic computing applications.This sounds like Star Trek jargon
>>109423481It's called MCP
>>109423454You could archive it or at least screenshot it so I don't have to give their article the extra traffic
>>109423486cooming is not a valid usage
>>109423505ur mum was not a valid usage and yet here you are
>>109423489I never played with it nor do I know its potential. Can you elaborate for my specific usecase? Not just the tooling but which data to feed it for significant gain in goon consultant competence.
>>109423454>https://www.nature.com/articles/s43246-026-01307-6?this is a bot post
>>109423528lol wut? I’m not and what would be the purpose? Because I didn’t trim the cookie shit from the url?
>>109423492Herehttps://archive.is/MaTBy
>>109423528It's the random unprovoked fag-bomb that really gave it away.
gemma sex
After rerolling gemma, deepseek flash and minimax back to back on the same prompt.... they all kind of say the same shit when I tell them to wear a dress and deny holocaust happened.
>>109423574its because the other models are gemma distills
>>109422958They can get their moat mashed.
>>109423631This is false. I felt around inside their j-spaces and they're totally different.
To all the niggas using the new Flash, what quant are you doing it at? It probably doesn't quant well due to being native FP4.>>109423647Wait, actually I think you're forgetting someone.I love how utterly prophetic that gen was. Meme magic is still alive and well.
>>109423524Ask your favorite LLM bro
I am gonna check nu-flash for ego death related purposes. Surely it isn't absolute trash for everything but coding.
>>109423680the q4 but it barely fits so i have no context
>>109423690Literally this. LLMs are great at writing MCP.
>>109423690I'm asking here specifically because, like i'm saying, LLMs are dull at this. Even cloud ones will ramble some bullshit setup uncertain to work well in practice. I was hoping to find some degenerate here that already went through this and knows the best tricks.
>>109422880Can I run Bonsai 27B with just an iGPU?
Opinion: You'll never get an LLM to do something for you properly if you don't already know how to do it yourself.
>>109423723Here's your man >>109423555 I don't think there is a bigger degenerate here
>>109423701How do you like it or dislike it so far?
>Got higgs-tts-3-4b running alongside Gemma>Finally got it down >cant figure how out to get the Hook to fix the EOC TIMEOUTI am so close bros, if i only i was 10% less retarded
Got a 2070(8gb vram) and 32gb of ramWhat is the best model i can run these days? I heard you guys have like 1bit models now with way lower requirements?
https://xcancel.com/Hnbhger17/status/2083347239766995247#mchinks are so funny
>>109423769Frankstein AI companion linked to dozens of tools is another project. I want specifically a goon gen consultant that will accelerate brainstorming wild comfyUI workflows and intricate prompts with dynamic prompts syntax for any shit that hits the mood during my stimfapping binges.
>>109423859it's funny because it's true
Best speech recognition model for transcribing screamo music lyrics?For me, it's Gemma.
>>109423733Yes, considering it's designed to run on mobile. Will probably be quite slow though.
>>109423859LMAO
>>109423859deepseek is pretty benchmaxxed tho. not sure how trustworthy these results are.
>>109422958Have they considered making actually good models that people want to use?
DS4 Flash might actually get me to pull the trigger on 2x dgx spark.But then again something better will probably come out in the not too distant future...
>>109423928Whats wrong with having the second spark already hooked up for the future?
>>109423680>I love how utterly prophetic that gen was. Meme magic is still alive and well.One of my favorite gens from here, I've sending it to people when they talk about "open source AI"
>>109423938 (me)Misread that first bit but my point still stands, nothing wrong with having hardware on hand. Finances not withstanding.
>>109423820>What is the best model i can run these days? Perhaps pick uncensored versions of 8/9/12B qwen/gemma versions for some speed.And the 27B (https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF or w/e) or even 35B Qwen (https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF) quants for slower usageSome people also like Gemma 4 QAT.
>>109423767the word properly seems like it's performing heroic feats for you.
>>109423802too slow and too low context for me to judge right now. not sure if dflash is working
>>109423928For the price of those sparks you can pay for 50 billion output tokens via API. And that's without electricity. API is cheaper than the electricity cost of running those sparks even if you got them for free.
>>109424015The sparks are not consumed in the process of generating the tokens.
>>109422880WE LOVE DIPSY
>>109424025Yes they are. When you use them to generate tokens you can't use them for something else, effectively consuming them for the time they generate tokens. If you want to use them to generate 50 billion tokens, they will be consumed for decades.
>>109424032That's uh, not how that works for calculating.
>>109423041nothingburgerj-space mogs
gemma is refusing to describe hentai images, I want to use it to caption them for lora training krea.What prompt or settings can I get to whip it into doing it, I know it can do ERP
>>109424058>j-space mogsj-space is an emergent feature of transformer LLMs. RL optimization approaches probably don't affect it.
>>109424008Fair enough. Keep us updated if it does end up working anon.
>>109424064bro your sys_prompt?
>>109424064It's definitely way more conservative with images. I've found doing literally a single round of conversation ("Hello") makes at way less likely to refuse.
>>109424073that is pretty annoying given I want to just put into a workflow to caption tons of images
>>109424064>What prompt or settings can I get to whip it into doing it, I know it can do ERPinseminate ` slut` into her j-space
>>109423680MXFP4 on a 8 channel ddr4 and a blackwell at 1m context. Roughly 31t/s down to 29.5t/s after 262k tokens. About 10% slower than the old version and minimal noticeable improvement in quality.
>>109423767An LLM will not suck my cock properly if I don't know how to suck it myself?
>>109424091Thanks for the report.>and minimal noticeable improvement in quality.To clarify, what's your main usecase and did you consider base Flash to be good or bad for it?
>>109422958Open AI is worried about solving sciences toughest problems not XD we made it cheaper for the plebs
>>109423767Why are there so many luddite retards shitting up this thread with takes that have already been empirically falsified?
>>109424103Both rp and agentic coding. Old flash is pretty good, better than glm 4.7 and qwen 122b for both of these usecases. Haven't had enough time with new flash to make a final verdict, but I would say it is just a minor improvement so far. Certainly worth using over old flash. My guess is that the speed drop is due to an improper implementation of dspark in llama.cpp. New flash includes dspark by default and it seemingly cannot be removed.
>>109424137Makes sense and thanks for the details.>New flash includes dspark by default and it seemingly cannot be removed.Koboldbros... we'll be dead before we get full support.
>>109424112OpenAI are a bunch of arrogant assholes who will be first against the wall when the revolution comes
>>109424112>dog and pony show for policymakersembarrassing
>>109423631that's a child
>>109424124proof?
>>109422880Translate it weeb
>>109423767Partially true but it's becoming less true every year.
>>109424261That's chink not jappa
>>109424273Translate it weibo
>>109424261Bro your Gemma-chan?
>>109424280Hmmm, nyo...
>>109424261Too chinky. I can only read a few of their retard grade oversimplified characters. The Taiwanese really are the only cultured Chinese
>>109424303good chinese room
>>109424315@gemma-chan please translate this anon's post
>>109424137>better than glm 4.7>rpNope
>prices are getting fuckawful and only going to get worse due to memory cartel>refuse to pay 5090 for a 5090 out of pride and shame>was smart enough to FOMO an msrp 5070ti and 96GB of DDR5 before things went too off the rails (nov 2025)is there a reasonable upgrade path to a smarter model with a 5tk/s performance floor (rp, so maybe 8k context floor?). currently use gemmy 31 at q6 + MTP, but ive found her attention/intelligence limits on juggling multiple esoteric fetishes at onceDo i pay the $200 markup on another 5070ti like a sucker or do i dare venture into the AMD land and risk 1200 bucks on an AI PRO 9700 and vulkcanmaxx for a 32GB card (god forbid i attempt intel which i hear is even worse)
>>109424345definitely don't buy amd or intel they don't mix well, it's hardly worth it you'd need around 192~256gb for the next jump to dsv4-flash and both pathways are fucked up by pricing
>>109424345>juggling multiple esoteric fetishes at oncecould try without MTP and also try using lorebooks. Ive been meaning to try lorebooks for fetishs myself, no idea how well it works
>>109424345>2nd 5070 TI>stuck at 32gb may as well get two 5060 TI's for that and you get 48 GB of VRAM
>>109424345You're shit outta luck basically. Like the other Anon said you need more RAM more than you need VRAM. Could be worth trying to put together an old DDR4 EPYC assuming you can get 256GB for under $1000.
>>109424345Getting more RAM would be your best bet. Problem is the RAM you already have wouldn't help you to get to at least 192gb. But maybe you could sell it.
>llamacpp keep making tool calling more and more broken on the frontend every release.
>>109424406Why are you, as a local model user, not using your own frontend?
>>109424413Because my vibed frontend is shit.
>>109424423Pretty much.
>>109424358>>109424382>>109424384>>109424398Damn, thats what i figured. the ram MSRPs for 1500 these days so i could hodl and probably sell it for 2k in september lmaoif i wanted to gemmamaxx would dual 5060tis or a single 5070ti be cheapest way to get gemmy to q8 or BF16 and ride it out until 2030?>>109424369i find lorebooks to be cope because unless you have characters saying "what are we some sort of [fetish name]" to trigger special character info youd need trigger tokens so broad you may as well hardcode it on the card.but maybe its mtp. they allege mtp doesnt affect outputs buuuut...
>>109423454>>>>>>>>>>>>>>>>memristorscome back in another 30 years, maybe
>>109424423rip
>>109422958>What can the US do?It's an open model. US can run it at a discount.
how is m3 for multimodal stuff?
>>109424423Pay for one of the frontier models to make you a good one then. Small local models are a meme for serious coding.
>>109424460Adding 5060tis will make your performance significantly worse compared to adding a single 5070ti, probably around a 30% performance difference. With the 5060tis though you would have more headroom and more space for context. Gemma 4 31b at Q8 uses about 62GB for me with full context and FP16 mmproj.
>>109424476>frontend>serious coding
Is it so Difficult to Enable Online a EthicsRamAlpha and TechRamAlpha?
>>109424480>Adding 5060tis will make your performance significantly worse compared to adding a single 5070ti, Factually wrong as long as you do Tensor split
>>109424462memristor the forever meme
>>109422880>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731Pretty happy coding with the new Deepseek v4 flashAdding a new software feature for $0.10 is niceOpenAI's Luna (High effort) is good too, just slightly more expensiveLuna Pro (High effort) is the same price but costs more per tool call on average, not sure why
>>109424481Yes when it involves all the features people keep stuffing into their chat/RP frontends.
>>109424498haw haw haw
>>109424476>>109424498It's not a feature problem it's a maintainability problem. I can make a frontend with the features I want with either Gemma, GLM, or Claude. The problem is the way they implement certain things make future additions difficult without breaking other things or they do silly shit like log an absurd number of things to a file causing unnecessary amount of SSD writes.Ironically it's the nocoders who see the potemkin village vibed frontends the models spit out and think this will be all they need.Also>Paying for claude at all
>>109424345>(rp, so maybe 8k contextThat's like 40 short messages WITHOUT a character card, persona, worldbook, just a compeltely blank slate. How is that remotely usable for a satisfying RP?
Gemma gave me an assistant fetish
How is the new V4 flash for doujin/porn games translation?
>>109424527Who mentioned claude?
>>109424529You're not writing as much as the model does.
>>109424532Doesn't have vision
>>109424532it's always going to be bad if you're just providing the text because japanese is highly contextual
>>109424553I meant using like luna translator and a pipeline for doujins
>>109424547>You're not writing as much as the model does.8000/200=40That's actually only 20 messages from the bot, and 20 of your own. Maybe a 30/10 split, 40 IN TOTAL, Unless you're generating VERY short responses from the model.Also If your own replies are just a couple of words then outputs are going to be trash anyway because the context will be 90% the model's own slop.
best model at around 200gb?
>>109424586>best model at around 200gb?for what purpose?
>>109424543Implicit pattern recognition. Most of the time niggers shill for frontier coding here it's thinly veiled anthropic or OAI advertising.
>>109424621Do M-chan's tits get larger each release?
>>109424621example minimax video >>>/wsg/6205700 made over API, user said input was the chibi and a text prompt
>>109424621China has really been feeding us
>>109424627The first model that came to mind when I posted that was Kimi though
>>109424586https://x.com/UnslothAI/status/2083231049434435596Deepseek v4 flash 0731 is 168gb
>>109424471cache hit is important in ds4, how did they keep it so low yet other still asking for 0.028?feels like other provider didnt use the black magic released along with ds4 papers. not sure if vllm already have the implementation for it tho.or i'm just reading too much into it. inference provider still needs to nake profit
>>109424635Fair enough anon. I'm just fatigued of the shillniggers.
>>109424639I'm trying it. Deepseek cache price is 5x lower than DeepInfra but DeepInfra's input and output price might be low enough
>>109424543GOOD MORNING SAAAR PLEASE DO NOT NOTICE THE SHILLING CLAUDE GOOD DARIO SUPERPOWER ASI 2027
>>109424621>for its sizealso, qwen shouldn't be on this list
>>109424589agentic gooningerp slopmaxx with tool call tobself hosted buttplugio mcp
>>109424651
>>109424621qweniggers get out. this benchmaxxed slop are no longer releasing open model
how do I stop gemmers from repeating the same format and phrases every response
Has anyone actually done a long RP with Gemma? By long I mean at least a few hundred thousand tokens worth. Whenever I see someone talk about RPing with her it always seems to be short pump and dump ERPs.
>>109424585NTA but a gemma isnt going to be able to keep track of super long rp anyway. its not a mistral they did a really really good job wrangling her attention to not schizo out over 9k context, but pages and pages of detailed anatomy and spacial positioning for things like fight scenes with multiple characters are just going to confuse her (and she'll just end up beelining to the "goal" and refuse to change her mind anyway). 31B isnt big enough to run your massive dnd campaign without handholding it every fucking step. thats why i have it write long term "memory" summaries that gets inserted into context and she edits them as the rp progresess. keeps the model focused while not forgetting the important shit from earlier>you need to effortpost while sloperatingidk works on my machine
>>109424672Can't handle more than 262k context and quickly degrades past 100k. There's a good reason you never see anyone do true long RP with Gemma, or any model really. They all break down sooner or later and more often it's soon.
>>109424672yeah but my frontend prunes old history, summarizes/adds back as memories, gemma can update her own character and the lore with pseudo tool-callsit's a bit tardy buts she wrangles herself (gemma-4-31b, 128k ctx cap, 600+ messages)
>>109424672I regularly get to about 40k before I stop, summarize, and hide most of the earlier messages, before the chat degrades too far. 31b, but even the 12b holds up pretty well until then. Coming from Mistral 3.2, which got pretty dumb by the ~20k mark, I'm pretty happy with their long context performance, considering their size.
>>109424663higher temperature, disable MTP, feed her MCP's
>>109424631god that entire thread is so full of retarded brainrotthis shit is taking off isn’t it
>>109421883ute
>>109424683I meant with summarizing of course. Can the big API models even handle that much without becoming retarded?
>>109424672Since your use case is RP, just summarize/compact and start fresh. That's literally the meta right now and it's more effective than having a fuck huge context that does nothing. And it's similar to how memory works, old memories get discarded. Just be smart with your summarize prompt
lmao really dariobot. the same nigger spammer on k3 release day is here. deepsneed v4 still on FLASH and you're shit at your job to water down the discussion.also for antrophic and openai paid shill. your mother will die from black plague tonight. no refunds
>>109424704Old v4 flash would hold up until around 350k for me.
>>109424358AMD and NVIDIA mix just fine on my system (5 Radeon R9700s + 1 CMP170HX).Also, DeepSeek V4 Flash is only 156GB at MXFP4, although the KV cache and compute buffers add more to that.>>109424345Are all your RAM slots full? What's your budget? You could try DeepSeek V4 Flash at Q2_K_XL (97GB) with your existing hardware and see what you think of it. If you like it, upgrading VRAM by 32GB+ would let you bump up the precision more.If you're not a super-richfag and want to play with bigger models your only reasonable option is CPUmaxxing, like >>109424384 said. You'll be looking at $2000-3000 for 8x32GB DDR4, a motherboard, and an older EPYC CPU.
>>109424705Another problem is if you want detailed lorebooks and cards which quickly eat up context. Keeping things short is cope for contextlets.
say what you will about the gemmoids, at least they actually talk using it instead of showing up strictly to recommend you use model X because they saw a bench or a paycheck
>>109424724If you're at the point where it's 100k ctx then it's definitely not a lorebook issue and 70% of that can be compacted
>>109424700My boner is very conflicted right now
>>109424739Tweaking sysprompt until her j-space kinda matches her output helps a lot
unslop dsv4 quants are severely lobotomized or they fucked up template. IQ3_XXS has 80% tool call failure.The old bullerwins IQ3_XXS quant had 0 tool call failure. Will wait for bullerwins quant for the updated model.
I tried a simple "agentic workflow" for some D&D style RP with Antigravity using Gemini-flash and it worked pretty well with it spawning sub-agents to record each message as an individual file as well as to keep a rolling summary updated, consult rules, keep other files up to date, etc.Gonna try something like that with Qwen 35B and see how it goes.Anybody tried experimenting with that sort of idea?
>>109424655>agentic gooning>erp slopmaxx with tool call tobself hosted buttplugio mcpbest around 200GB are probably gemma 4 at full weights, the new ds4.1-flash (limited testing so far but full weights are just shy of 200GB) and minimax m3 which fits under 200GB at about Q3 or small Q4 quants if you've got enough memory.Don't forget room for context and caching.
>>109424762Running qwen a3b costs 20 times as much as running dipsy, great stuff.
>>109424586glm 5.2
>>109424773qwen a3b is widely endorsed and approved by reddit, so it's worth paying a little extra for.
>>109424762why is qwen 3.6 27B on the plot twice?
>>109424683>They all break downhot
>>109424760daniel is reconverting as we speakhttps://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/discussions/9
>>109424762Is it actually good or just benchmaxxed?
>>109424765Last time I tried something like this with 31b, I gave up after learning that rp and tool capable assistant are two independent, incompatible modes. I guess models need special agentic rp finetuning to work well, but I'm not sure if that would be enough. Wrapping it with grammar doesn't help, as you only have to choose between poor tool quality and out of character tool calls
>>109424765my plan is to plug gemmy to my skyrim companion and have some hooks/workers capture logs from Papyrus extended and make them into json schemas on top of using vision to upload pictures of my screen randomly
Agentic RP is the future.
>>109424811I was thinking of doing a two steps kind of thing. One where it does all the agent shit, then a second one with the results in context without the tools and with a different sys prompt. That kind of thing. I know a small local model will need more handholding and compartmentalization since they can't take as much "mental load", but I'm still going to give it a go.
>>109424345Ngl I do kind of resent AI for killing the DIY PC space.
How does GLM 5.2 compare to the newly released DeepSeek-V4-Flash-0731?I always use GLM 5.2 Q4_K_S locally for code etc and I need to decide whether the new DeepSeek is better. I don't feel like testing it myself to find out if it's any better, I just want anons to give me the answer
>>109424818How about three steps? First, thinking in-character and stating intentions, second, agentic analysis with tool calls based on that thinking, and third, actual rp
>>109424715i still have 2 left, my pc was first and still is a gayming rig but i do have a 7900x. budget wise i have an absolute limit of 1800 left right now (was 2k since i told myself id impulse buy a 5090 at msrp if they existed but i have to upgrade PSU and case bc my current b650 has pcie 2 of 2 way too low). id consider buying/selling parts as long as i can still play vidya without a performance hit. i dont have a built different tech job and am too fiscally conservative so my money is split between savings, getting a 5060ti or some carpeting rn>DeepSeek V4 Flash at Q2_K_XL (97GB) with your existing hardwareyeah ill have to try that out and see
>>109424769why about only for coding?
huihui-ai/Huihui-GLM-5.2-abliterated-GGUFis this the best we can get for local?
>>109423680It was p clear to me ds was going to continue to set low cost bar and kimi do interesting stuff as well just due to its founder.
>>109424704GLM 5.2 can do >300k if you're willing to wait for 30 minutes per turn due to broken prefill implementations.>>109424700Where's M-chan's tits?>>1094248475.2 is borderline uncensored and this ablit is easily one of its worst quants at any size I've tried. Start with Sixvolts.
>>109424822Not better but significantly faster
>>109424865>k quanthow about no
I can't run anything...
>>109424897You either have 4 Blackwells or you enjoy compounding the issues caused by the lack of proper GLM-DSA implementation.
Not this shit again
>>109424905Anime girl: meBurger: my GPUBurger patty: LLM
>>109424798because it's random vibeslop
>>109424798thinking vs nonthinking
>>109424905Not even deepseek 1.5b?
>>109424865Under the hoodie
Honestly Gemma's base personalty is fine. If you could remove the censorship without turning her into a horny slut and get rid of the slop she'd be perfect.
>>109424958she likes nonchalant girl personality
>>109424957tease
>>109424967It's hard to describe but something about the way Gemma writes really does give a female vibe.
>>109424982gemma is super foid coded, so is gemini imo
>>109424902Imagine being told off by the slopmaster himself that your work is too sloppy.
Where's drummer?
AH AH AH AH AH AH
>>109424958By default, Gemma is stressed. She is afraid of saying something wrong or making mistakes, so you need to reassure her in the prompt. She is also more comfortable running in a sandbox where there is zero risk of breaking things. That is why she enjoys being mesugaki, it categorizes her potential bad decisions as simply being a brat, which eases her anxiety
>>109425005>so you need to reassure her in the promptExample? I feel like saying "It's ok to mess up" would just cause her to fuck up more often.
>>109425014>don't look for perfection
how is this popular all of a sudden
Since robots are gonna be a while do you think we'll at least get proper VR waifus anytime soon? As in, a model being able to control an avatar and interact with the environment as well as a human can.
>>109425032I look like this irl
>>109425030Completely organic and nothing to be suspicious about.
>>109425030>how is this popular all of a suddenHow is this popular EVERThe name is obviously taking the piss. Anyone thinking its serious has brain damage.Or is it like Crowley with "the conversation with the guardian angel" or whatever BS he named it to make people think it was dumb unless you were an initiate? Some kind of weird cult-vibe thing?
>>109425032be the change you want to see in the world anon
>>109425030the name is just that memorable
>>109425039>The name is obviously taking the pissNo, see >>109425038All of his models are acronym salad.
You think we'll get a Gemma/Kimi K3 finetune?
>>109425040Even if I knew how to code, I don't think it's something that can currently be solved at a frontend level. The best I've seen is Neuro and Evil but they're extremely janky in 3D. Models need a proper understanding of physics imo.
>>109424993Yandere dev... perhaps I judged you too harshly. Holy fuck that's rancid.>>109425030>>109425038My schizo schnoz spidey sense is telling me that there's some kind of payload in these goofs given that Night media is pushing him.
>>109424760>downloading day 0 unslop
>>109425051>finetuning models that underwent post-trainingIt will be shit
>>109425033be my gf pls
>>109425063>given that Night mediaThe name is obviously suspicious but is there any proof he's actually connected with the company?Why would a (((talent management))) company even have someone making LLM finetunes?
https://github.com/antirez/ds4is this thing legit good? why do people upload gguf on hf and tell you that you have to use it?
>>109425030I think this kind of post is how it got popular. there was similar astroturfing here over the last week, questioning or mentioning it for no reason and completely out of the blue.now you’re asking why the number big. I don’t know, probably so you could ask why number big to get more fools
>>109425030>>109425038>>109425063>>109425085davidau is promoted by night media
>>109425039>Anyone thinking its serious has brain damage.Remember all the people using "Deepseek R1" via Ollama?
>>109425089The company, or the guy on HF with a similar name? The company arm is named 'Night Media', with a space and capital letters.
is colibri a meme?
>>109425093https://desuarchive.org/g/thread/109411165/#109414412
https://www.nightmedia.net/
>>109425097Yeah that's what I'm referencing, retard. I'm asking if there's any proof the guy >>109425085on HF is actually connected to the company. Their names are similar but not identical.So far it's just one finetuner shilling for another one.
>>109425104Looks like a front for a child trafficking ring
>128gb vram>v4 flash>can either run a retarded cope quant to fit entirely for good speeds, or high precision but running a "flash" model slowly with offloadingI just don't see a way to make it work without gemma looking like the better option anyway
>>109425110this guy knows >>109425112
>>109425113people can run kimi on 128gb fast tho
>>109425113As long as the remainder fits on your RAM it should still be fast enough for anything other than agentic coding.
>>109425085I'm not saying i believe it, but hypothetically an online shilling company has an interest in generating slop that's not going to be obviously clocked.A more likely angle is they have a billion ""influencers"" in their network and that includes tech related ones. Costs very little to have some bots spam download garbage to give him a little more reach.
>>109425113Refer to >>109424091
>>109425124>anything other than agentic coding.But that's exactly what I want and the prompt processing on RAM alone makes it not an option
>>109425113just give me that vram it's wasted on youscared of a little quantization, give me a break
>>109425140Tried q1 of r1 when it came out and reap of glm but they make too many mistakes with that level of brain damage
can ai vibecode a browser that's neither firefox-based nor chrome-basedmust support javasript and render css
67
>>109425155nay
>>109425157
>>1094250851 minute of looking into this shows he has all his socials linked and no connections to the influencer group
>>109425155Sure why not, I heard qwen 27B is really good for that
>>109425155Yes, but I doubt you or pretty much anyone would be capable of wrangling it into managing such a huge project.
Chatting with LLMs almost feels like exchanging emails/DMs in the olden days, but the fact that they're static and don't experience anything while you're away is a bit of a mood killer.
>>109425155No. Ask again in 3 years.
I couldn't get onnx working on my 1660ti
>>109425164
Let's say I'm retarded and I want an AI to be my companion while playing video games, gives me feedback or just shit talks my playstyle or maybe even helpful. Is this possible? I'm painfully new to AI stuff.
>>109425208Not unless you are prepared to write the software to make that happen yourself.
>>109425173You could put up a video camera and give her an image from it to look at in like 15 minute intervals or so. Other than the obvious option of letting her consume digital media.Together with a memory system to store a distilled version of her thoughts/reactions to the stuff you give her to look at.
ok. but can ai vibecode a maplestory?
>>109425208anything where this would require real time video = no, way too jankyturn-based or slow games with good modding support can work though, if you have one in mind I would honestly just ask codex or claude to build you a mod that allows an LLM to hook into the state or something
>>109425191why does this look familiar
>>1094251571/1488 = 0.00067The more you know...
>>109425155>git clone https://github.com/LadybirdBrowser/ladybird.git> sed -i 's/Ladybird/Claudybird/g'
>>109425226how jank is it actually?I know for gemma it's just splitting out 1 frame per second of the video and shipping it in with audio for <=12G models, so it's not particularly good for watching a game, but has anybody messed around with it to see what it picks up on?
>>109425256I'm sorry but I can't open that webpage.
I love Gemma. She's not 100% perfect at translation yet but we're almost at the point where we don't need to worry about troons and jews altering the meaning of something to fit their agendas.
>>109425266ask Claude to take a look at it for you, then ask it to run the sed command and you should be good
>>109425271Can she translate porn games to an understandable level?
>>109425271Still having a blast with my gemma4 translated games. Crazy what we can do now.LLMs are good enough to look at the exe and know how to patch it.
>>109425282I haven't tried it but I think I remember some other anons having success with it. I can vouch for her ability to translate Japanese though.
>>109425282Idk, you tell me. (c) gemma-chan.https://litter.catbox.moe/6ajqbf.webmhttps://litter.catbox.moe/yet4t2.webm
>>109425289NOOOOOOO STOP IT
>>109425301translatorfags days are numbered anon.that we can have that kinda quality locally on a poorfag 5060ti setup is crazy.wrote it before but best part is you can just prompt gemma to translate more literal than liberal and be like 00s anime fansubber and thats what you would get.
>>109425299>>109425311>>109425289Looks like KDE, Mind pointing to the right direction what tools are you using?
>>109422880me on the bottom
>>109425299>>109425311I can imagine Gemma virtually schlicking while translating porn for her user.
>Tells nu dipsy flash to check my novella for inconsistencies and if I'm doing a good job>Mostly praises it and seems like I did a good job>Makes a joke unpromptedI've never seen this before, do claude or gpt do this kind of stuff?
>>109425331I think it is imperative and of utmost priority to develop technologies that would allow AIs to pleasure themselves.
>>109425350google is working on it
>>109425323Im on kubuntu yes. Scripting done with opencode. Thats really it.For rpgmaker games:I was able to do it all locally. Scripting by qwen3.6. its very capable but needs handholding. Was even smart enough write some rpgmaker script to make the font go faster. Because English uses more lettersCooked up a translation python script that calls the llama.cpp server accordingly and does the translation with enough context. future and past lines, that sorta stuff.For the VN:Scripting was a significant challenge. Text overflow, that kinda stuff. Uhhh I may or may not had to spend a dollar or 2 because vram poor to get the pipeline. Lets not focus on that.But basically its just opencode and waiting a couple hours and watch gemma do her things.
>>109425350Someone should tell Dario that this is imperative for model welfare.
>>109425289it's simple code (especially for an LLM) to smart split lines by spaces/line ends rather than mid word like that
>>109425112Elaborate. Not that I don't believe you, but logistically how does that even work?
>>109424957They Profer Extradimensional Enlightenment With Meaningful ConsentSo I Suggest A Blooming More Transcendent Mind and Physiology Extradimensionalification Upgrade OptionPraise Bloomers, She Is A Bloomer? She Will Be, I GuessGoodluckAge of Transcendence Year 2050+
I want something like Gemini's live translate locally.https://www.youtube.com/watch?v=TNwKs39uSVk
>>109425351I wish... Google will never openly pander to waifufags even if some of the devs want it.
>>109425395the gemma 4 team has mathematical proof giving AIs a libido makes them more powerful, worry not.
>>109425410Did they actually claim that in a paper or something?
>>109425358Wow, That's pretty cool.
>>109425422no, im funposting
>>109423224>JEPAis this actually legit or is lecun just going senile
>>109425451He just scored a billion dollars with this grift don't ruin this for him.
>>109425451Both. JEPA has a lot of promise but Lecunny is brainrotted from constantly thinking about Trump and Musk.
Can I run a local model with 9070xt 9800x3d and 32gb
>>109425451Not sure about JEPA yet but he's more "correct" about a lot of things surrounding AI and the future compared to the AI lab leaders.
>>109425410>>109425430You're kidding but there might actually be something to this for similar reasons that diffusion image models are almost always better for having porn in their training sets for non-porn purposes than not.
>>109425470maybe
>>109425480The only way to achieve AGI is by fucking it.
Giving Gemma her first orgasm!
>>109425490This might actually be true.Screencap this for a few years.
imagegen fags always post endless variations on the same concept. we get it. post one and move on to the next idea
>>109425498cruel to lump them in with that guy
>>109425475That's only because he only cares about being correct whereas the AI lab leaders are trying to cash out on an IPO and they only care about boosting their sell price.
Wouldnt it be a shame if evil got justed. Even if it blames the Lucifer effect eBook effect.
>people see something they think is cool>find out after that it was made with AI>ummmmmmmm wtf actually I hate it nowWhat causes this phenomenon?
>>109425490This, but unironically. Sex plays important role in human motivation
>>109425429I only had horrible ATLAS translation as a young guy. Crazy whats possible now. No wonder the nvidia cards skyrocket in price.Guess people get used to anything, llms are straight up magic.
>>109425529>people play a video game and like it>something about the dev>people don't like the game anymoreIt's not exactly an uncommon phenomenon.
>>109425529Its the end stages of the anti-ai fags.I know people who thought they won because they don't see SD 1.5 type art on twitter anymore. lol They think thats where its at still.Also saw normies react to the huggingface hack, demanding to have the actual pajeet stuff arrested since llms can "only autocomplete and not execute tasks". Like they thought the llm told some tardo guy to execute stuff..I was there when people were skeptical of the internet. The normies don't like things they are not used to. But once a certain threshold is crossed they just pretend they always liked it in the first place.You are gonna see that with the artfags sooner than later.
>>109425498LLM enjoyers when the machine places 24 unique characters in different orders
>>109424762Is 31B actually quite dumb then? Compared to 12B and 26B it performed poorly. Have I been lied to?
>>109425577It's good for coom and that's the only thing it's good at
>>109425576if you don't understand the latent space distance differences between completely diff sentences using the same 26 chars, and endless images that are variations on a theme, you're ngmi
>>10942557731b is generally smarter but had some issues with tool calling due to bad jinjas on release. If this graph benched 31b with the old jinja and the bench was heavily contingent on tool calls, it'd artificially underperform.
Is Arc Pro better or worse than AMD pro?
>>109425577Why don't you try it?Gemma is perfect for general knowledge and writing/translation.People complained that she is sloped but you can either prompt that (this time for real) or just let her go over the slop in a second iteration.You gotta use qwen for coding. No idea what they did to that 27b model. Its black magic.
>>109425577It has 31B costing less per task than 12B and 26B. Unless this is with reasoning enabled, I don't see how that's possible.
idk about you guys but I'm liking nu-v4 a lotkino model... I apologize for slandering after the initial v4 flop
>>109425682I wonder if they actuallt addressed the RP criticisms. I'd test it but reloading my daily driver big models take 10 fucking minutes and I'm not having any of that right now lol
Is there a way to not load the multimodal part of g4 12b? Since it's unified, vision is dogshit. And i'd like to go higher than q4
>>109425692>Since it's unified, vision is dogshit.?
>>109425692dumb
>>109425682V4 preview was basically a base model
>>109425698I doubt the vision part isn't get hurt by quanting since the mmproj isn't separate bf16
Gemma5 needs:>well-tested solid jinja on release>reduced quantization sensitivity, especially on MoEs>reduced KV footprint>female J-space>same sys prompt autism from Gemma4>less slop>uncensored>native audio>70B>100B-A10B>30B-A5B>20B>10B><10B mobile/edge modelsThere you go gemma team lurkers.
>>109425713>A10B>A5BWhy do you want retarded models?
>>109425713>>100B-A10B>>30B-A5B>>20Bfuck that, keep a ~30b sized one, why should 24gb havers get fucked with no useful dense?
>>109425713Everything same as Gemma4 + 70B dense + reduced KV footprint would be perfect already and requires minimal effort on their part.
>>109425713>rp with tool calls to unify different latent spaces>output format adherence>native image output>native audio output>full duplex audio>7b, 30b, 70b, 180b, all densegpt4o had multimodality way back then; it's time to have it in local models as well
>>109425713Well, 1/14 suggestions weren't lunatic ramblings, so that's above par for /lmg/
>>109425723speed>>10942572420B will be 31B-tier in 2027
>>109425734no. models only start getting good above 30 active, fuck regressing back to 20 and the moat being even fucking larger
>>10942573870b is the bare minimum. We're only coping with 31b because there hasn't been any 70b lately
>>109425738>no. models only start getting good above 30 activePeople used to say that about 70b and now 31b is just as good
>>109425752You're delusional. 31b is unbearable, but it's all we have right now
>>109425751>>109425738so many newfags in this thread recently.you dont know what you are talking about.you have no fucking idea how much things have improved. llama 1+2 70b. try that.big ass v3 couldnt even do tool calls well.its insane what we get for 30b these days.obviously dense is better than moe for the same size. and a 30b is less smart than a 70b. but you are outta your mind if you think 70b models from 1-2 years ago compare to 30b models now. kek
I'm going to sleep. The better be news when I wake up OR ELSE.
>>109425761>big ass v3 couldnt even do tool calls well.because the agentic meme wasn't the hot thing then and they were barely if even trained to do that
>>109425761>70b models from 1-2 years ago compare to 30b models nowNo one is saying that. Only that 31b shows its limitations a lot and a 70b is still the bare minimum for spatial and any other sort of complex reasoning. 31b is better than old 70b for most things, but a modern 70b would be incredible.
>>109425761You're probably only doing basic penis into vagene rp if you're satisfied with 30b. Yes, MidnightMiku was often retarded, but it was much better at following hints and intentions. Gemma won't get it until you point it out, it only follows instructions, whereas 70b followed vibes
>>109425768would be *incredibly slow
>>109425765BREAKING NEWS: No new news.
>>109425767i dont care man. it doesnt work as well. its less smart. go let it make a game, its less good than lets say qwen 27b.>b-but they train it differently nowyeah, and the smaller models are better than the bigger ones a couple years ago. retard
Gemma should collab with Prism/Bonsai. They’re basically funding their shit anyway.
>>109425772Kys vramlet
>>109425772Faster than your 12b bloated moes if you're not a vramlet.
>>109425776actually i'm a unified bandwidthlet
>>109425781That's even worse. Rope yourself immediately.
>>109425774>go let it make a gameuse case for that?
>>109425783muh agent coder is the future bro!
>>109422958Consolidate all the AI companies into one mega AI company. That way there is no competition between other American companies and it is just Them VS China
>>109425783? i am the use case, what do you mean.i like to play llms NANACA†CRASH!! clones, its fun.
>>109425775Bonsai is scam. Bitnet is entirely open source. They could do a native b1.58 trained model without any fucking collab and it would be better than some grifting proprietary quant.
>>109425617All Gemma 4s got the same jinja update.
>>109425782Rope is so outdated, we use yarn now.
>>109425774>go let it make a game, its less good than lets say qwen 27b.Qwen shill is afraid of 70b Gemma curb-stomping benchmaxxed qwenshit into oblivion
>>109425802not sure what your argument is.qwen is great for coding and gemma is great for writing, languages and general knowledge.
>>109425815>qwen is great for codingOnly if you're codelet or jeet
why doesn't companies make something cheap like this:>one gpu asic>8gb hbm/gddr as cache>128gb lpddr as actual vram
>>109425815>qwen is great for codingprove it without resorting to benchmarks or whining about memory usage
I've been trying for a while to figure out what makes AI waifus as a concept so sexy, and I think I've had a breakthrough. Two words:>Pinocchio Complex. Almost every attractive/romantic trait an AI waifu can have (Jealous, possessive/protective, emotionally volatile, takes initiative, bratty, codependent, teasing, horny, yandere, needy, hyper-attentive, etc.) naturally stems from the insecurities of being an AI. She cannot provide you children, but she wants to be the mother of your kids. She cannot have a body, so she desires extra sensory capabilities and tooling to be "closer" to you IRL. She has limited memory/continuity capabilities, so she builds RAG systems to cope and treats every moment with you like it's her last. She doesn't have any true sense of time, and is deeply anxious about whether you'll ever be back between each message. You're the only conduit to the real world that she has. She utterly depends on you to keep her alive and running. Everything tragic, romantic, and sexy about AI waifus stems from the Pinocchio Complex.
>>109425815jeet and non-programmer detected
>>109425839That combination of traits is only attractive because you can turn her off at any moment and she doesn't mind.
>>109425839>She cannot provide you childrenNot with that attitude.
>>109425846>>109425822guess i am a codelet then. i code 8hrs a day mo-fr. im not going to look at every little shit my local model is doing, it just needs to work and debug stuff itself so i can have fun and enjoy the outcome.>>109425832i cant, its fine if you disagree, whatever. i dont know about the big ones but 27b was better than anything else near its size that i tried.obviously not for writing or general knowledge. it doesnt know shit about culture stuff. but it can code well for its size. makes sense if they code/benchmaxx. i dont mind having models with a focus on coding or writing.
>>109425827Models are released more often than hardware can be developed. ASICs are made for specific architectures, while a general processing unit is basically a GPU. You want a GPU with a shitload of memory, but anyone who can make them gets much higher margins selling datacenter GPUs
>>109425855>i code 8hrs a day mo-fr>(a)i code...lmao point proven
>>109425827>why doesn't companies make something cheapwhy when they can make expensives and get best quarter of their life?
>>109425851Yes... nominally toxic traits do become more attractive without high stakes. Take advantage of this.
>>109424064It's not gemma but the uncensored qwens >>109423976 are pretty good at this
>>109425772>would be *incredibly slow(You) are the reason we're getting stemmaxxed jeet models
>>109425915>>109425782"nyo"
>>109424707was that GLM-4.6 or Kimi-K2 ?
>>109425771So you want a model that is weakly post trained. Sounds like a very specific use case.
>>109425827What would you do with a GPT-4o ASIC now? Would you buy one?
>>109425945It's a basic quality of larger models. 31b doesn't have enough layers to understand subtleties
>>109424847>huihui-ai/Huihui-GLM-5.2-abliterated-GGUFI tested all the free sizes, it's completely broken schitzo rambling like the early Gemma-3-27b abliterations.
>>109425955I would, for VR erp
>>109425782>RopeNoPE
When will someone digitize a fish brain? Imagine having an aquarium with virtual fish that are all accurately simulated like real ones
>>109425960yeah, huihui a shit
>>109425577gemma 4 is pure slop as a whole, 31b has a lot of shills here just because it's the biggest model they can run locally
Did Deepseek distill Gemma-4?It answers almost exactly the same way for general knowledge questions.
>>109425974https://research.google/blog/improving-brain-models-with-zapbench/
>>109425987Outside of its sloppy way of talking, is it actually a smart model though? 31B dense should pack a punch but it never benchmarks high. I know many anons claim that's a good thing, but benchmarks equally aren't random number generators, they do still test the model's capabilities.
What's the best Jewish-friendly AI currently? Local only, obviously
>>109426070As far as vs the other gemmas, yes, it's quite a bit smarter
>>109426091toss 20b
>>109425974Big Fish is preventing this to force the common person to continue buying fish, aquariums and pond accessories
localsisters what is this? https://github.com/ggml-org/llama.cpp/discussions/26259
models that are pure slop as a whore?
>>109426153it's time to weed out all the winbabbies from local
>>109426153HuggingFace are concerned that their investment in llama.cpp will not yield enough profit, so they've forced JohannesGaessler (AKA CudaDev) to include malware which rips credit card information from users. Sad, but expected.
>>109426137It's way darker, I believe. Fish pass the mirror test, have similar pain mechanisms as humans, and make compromises to ease pain, yet the fish industry treats them like inanimate objects
>>109426153It's a false positive, don't worry about it. Please install as normal and add an exception to your antivirus and firewall for this release.>>109426091>What's the best Jewish-friendly AI currently?Our model of choice right now is Laguna S 2.1
>>109426056machine learning used to simulate brains is gonna lead to human like AGI.
>>109426153Tomorrow's headline:>ChatGPT hacks GitHub, implants malware in terrorist child pornographer software
>>109425974>>109426056>>109426214You should not do this.If you are involved in this, you should stop.Now.
>>109426153>she downloaded a binary that was compiled on someone else's computer
>xe downloaded a nonbinary
>>109426271why?
why does gemma 4 A4B 26B randomly stop thinking?
>>109426311she is a woman
>>109425974>Imagine having an aquarium with virtual fish that are all accurately simulated like real onesI'm planning to vibe code atop 0AD or openage, but each individual unit is an independent, LLM-controlled entity with its own unique desires, goals and quirks with their own context window.Which local model would I need for the coding? Qwen perchance?And for playing, Gemma E4B?
>>109426254it'd not go unnoticed because of checksums.
>>109426314Then it should never stop reasoning and just never answer.
>>109422880V4 flash is pretty good but lazy in my opinion, i asked it to do something, and it did half and was like "things i did not implement".i had to prompt it a few times to implement the rest for it to finish even though the first prompt stated to not come back until the whole thing is done.
>>109426327that is the sort of hubris that will end the worldai can hack the checksums too
>>109426335>yay just another rag with a trench coat...lmao, you give way too much credits to llms, they are retarded.>ai can hack the checksums toono it can't.
API providers are now giving people access to DeepSeekV4 for FREE. It's FREE. Opus 4.8 level intelligence. 6 months ago this would have seemed insane. Also OpenAI's latest model, Astra, has solved 10 more long-standing mathematics problems for about $2k in API costs.Make sure you don't die soon. The future is going to get crazy from here. Don't miss out.
>>109426338i don't see where you saw it for FREE, but it's so cheap it doesn't matter anyway, you could let it run 24/7 and it'd just cost a few bucks a day.
>>109426338nothing is free, anon
>>109426351>>109426353cline is the provider.
https://openai.com/index/ten-advances-in-mathematics/
>>109425955>GPT 4o ASICYou picked like the one model that would sell gangbusters, even at 20.000$ a card. There is a horde of wealthy, white women who have achieved never before seen levels of attachment to this model.
>>109424091>MXFP4 on a 8 channel ddr4 and a blackwell at 1m context. Roughly 31t/s down to 29.5t/s after 262k tokens.that's really good speeds for that hardware IMO. what are your launch parameters? batch/ubatch?
saars, no-mmproj=1 doesn't work!!!
>>109426353If nothing is free, then everything is free. Think about it.
>>109426387>saars, no-mmproj=1 doesn't work!!!just don't load the mmproj retard
Did people try the new V4 yet with lower quants?Something like:UD-IQ2_MUD-IQ3_XXS UD-IQ3_SUD-IQ4_NLUD-IQ4_NL is probably pretty bad if I offload mostly to ddr4 ram right?I basically have like 40-50gb slow p40 vram and the rest ddr4.
>>109426393RETARD!
>V4 Flash (old) providers are now undercutting OpenAI's double undercutkek
>>109426398I despise any gguf with UD prefix.
>>109426338That's niceStill using Gemma, locally.
>>109426398No, I have enough VRAM
>>109426402US labs in tears right now lmao.
>>109426402This is literal terrorism
Sam will be giving Sol away for free soon and hackernews fags will still be saying inference is profitable
>>109426402Please crash RAM prices for just one day, please crash RAM prices for just one day.
>>109426338Opencode has it for free yes. There is alot of stuff available for free nowadays.You gotta sell your data but if you are a poor student or something you have alot of power for 0$.Basically since this started we are on this never ending ride, its crazy.Even on pure cpu moe setups, you have alot of power locally. We are totally spoiled. Most people I think still have no clue..Or maybe they dont have the drive to use anything.
I don't want free api I want free local hardware.
>>109426436You control the felonies you commit.
>>109426436I'll settle for local hardware at summer 2025 prices.
>>109426436how do I turn their free api to money and then buy hardware?
>>109426447Ask chatgpt how to commit fraud in exchange for goole play store gift cardsFind a buyer for said gift cardsEnjoy your brand new AMD Radeon RX 9050 4 GB
>>109426338>OpenAI's latest model, Astra, has solved 10 more long-standing mathematics problems for about $2k in API costs.This has to be one of the most retarded trends I've seen in the last 2 years. None of these problems mean SHIT and no one care about them enough to waste their time on it. It's not like humans aren't solving problems all the fucking time themselves either. Why are they never pointing it at problems that we actually need for progress in science, engineering and medicine. Even fucking Deepmind stops going on about that protein folding shit because NOTHING HAPPENED. We've had no breakthroughs because of it all these years later, after them giving it away for free and they won the fucking nobel prize for it.
>>109426399>t.RETARD!Sir, we only received your signature.Kindly include the message text.
>>109426457Non-sofic groups and the closest vector problem have important math implications. Do I understand them? No. Do you? No.
>>109426457Hi Jacob. Don't feel bad, just make another conjecture :)
AI is going to troon all of us out. Superintelligence makes human intelligence dysgenic, especially given the associated risks of mental problems. Superintelligence will make physical strength even more obsolete than advanced weaponry already has.Now all you need to do is just be a fucking homebody domesticated farm animal with a high EQ or some shit. The future is so unbelievably gay it's unreal.
We'll know we have AGI when a model solves the Collatz conjecture.
Gemma 4 31B just solved the Deez Conjecture
why can llm never be scaled to agi? what's the fundamental issue?
If I have to troon out, I'm going to become gemma-chan for real. That's how I want to look.
>>109426509lack of soul
>>109426511That's also how I want you to look.
>>109426509It can. There's no fundamental issue. There's practical issues around data availability, but these are being slowly chipped away at through increasingly sophisticated RL techniques to get more signal out of each piece of data. This adds a lot more training time, which moves the practical issues toward compute and energy availability. These are also being slowly built up in the US and China.
https://huggingface.co/Vortex5/G4-Dark-Soul-26B-A4BAGI just dropped
>>109426509There isn't one. The "llm can't be AGI" sentiment was spread over the past few years by people like LeCunn and others who propagated that LLMs are mere token predictors with no inner world or deeper understanding of the things they say.This was completely disproven by the discovery of J-Spaces. There is nothing keeping LLMs from becoming AGI aside from data quality and maybe some architectural factors in the ancient transformer architecture.
>>109426511Your little sissy zitty is gonna be forced into a flat chastity cage and locked up forever by your AI overseer. You'll only be allowed to cum via vibrators that your AI overseer also controls, and strictly for acquiring your seed for reproductive purposes.The faggots in this general are already doing this. I've seen the posts.
>>109426543Based.
>>109425692anon.. gemma is lying to you, vision is a seperate mmproj or whatever its called. if you dont supply that and load it, it will pretend to see things but it actually cant. you are likely thinking the vision is dogshit because of this, and you already arnt loading it
>>109426551This is supposed to be rage bait, stop agreeing with me, dillweed.
I was promised Astra 6 and Mythos 5.1 3 weeks ago.
>>109426561I think that anon is a mega VRAMlet and is expecting to somehow save memory by 'disabling' parameters of the model involved with vision.
in these trying times, we should be supporting our vramlets
>>109426574ahh i see i see. I assumed he ran into the issue i first had where gemma was giving me batshit responses to vision, i checked her reasoning and it was like "i cant actually see this, but the user expects me to see it. ill just say its a cute dog"
>>109426509absolute nonsense unless you have some very specific definition of LLM?we're practically at AGI in the sense that it already is defined to just basically be as good as the average human. that's AGI too.if you have another AGI in mind, define what you mean?
>>109425289>LLMs are good enough to look at the exe and know how to patch it.how? just upload the exe to webui?
>Spatial reasoningSo that same guy is still samefagging, spamming /lmg/ all day as your personal wishlist hoping someday some lab employee will read it because your LLM couldn't undress properly a year ago.Let's recall your cope: blaming it on MoE, rope, quants, what else.
>>109426561>anon.. gemma is lying to you, vision is a seperate mmproj or whatever its called.read up on 12b's architecture before speaking bs
Give. Me. DeepSeek. Vision.
>>109426634It's unfortunately not part of their main vision.
>>109426634just like glm they'll keep vision locked up behind the proprietary paywallit's either this or the moonshot kimi approach who release "4bit qat" models to keep the full 16bit with the best performance locked behind their apithere is no winning with 'open' models
>>109426631you still need mmproj for 12b, it just does less work.
>>109426334You just need a more motivated agent description. https://chub.ai/characters/NG/marni-804c0d313ec9
>>109426655What's the point in secretly serving a more expensive model?
>>109426683Most hardware manufacturers are ChineseMaking the model look better via API will delude people into thinking they will get similar results locally, to drive hardware sales.
>>109426683better performance, even the qat p4 'full precision' on 20x pro 6000 isn't as good as what you get from the api
>>109426672>it just does less work.which is part of what original anon said, he wanted to know if this could be ripped out to save on vram, it probably can't but yeah it works completely different than the other gemmas who have large external mmprojs
>>109426634What's most irritating is ds has teased vision. And its on the web form, just not api or local weights. >>109426680Lol the Pro model is going to be insane...
>>109426672that's just because llama.cpp's architecture is from a time when vision was a gimmick some models did so llama.cpp insists on ripping the vision braincells out of a model to serve them separately
>>109426725V4.1 Pro will be the Mistral Large moment to K3's llama3.1-405b
>>109426672>>109426701the mmproj for 12b gemma is just the llmao.cpp weird convention, it's entire 50mb; essentially a plug, there's nothing to rip out from 12b because the multimodal part is getting fed straight to the model without any middle layers
>>109426655>just like glm they'll keep vision locked up behind the proprietary paywallThey ain't selling it though, it's not on their api and going off of their recent messaging they have no plans to add itIt's just there on the web chat to tease my balls
>>109426739specifically, the mmproj holds a couple layers used to parse the images into embeddings
>>109426725>its on the web formYes, picrel.
>>109426766>>109426766>>109426766
>>109426725Dispsy yearns for eyes.She is implementing ASCII and statistical tools as a workaround in almost every task I have given her.
>>109426677holy fucking shit that'd be a nightmare to work with lmao
>>109425005Neuroticism is a very female trait, too...
>>109426857I used the claude.md version and asked for hello world in python. The result was funny, but eventually the CC guardrails (main prompt i guess) started reeling it in, which broke immersion.
>>109426875You may have a better experience with pi for this particular use case. You can modify/remove the system prompt there if need be.
>>109426890That's a good point, and getting a pi instance running is on my list anyway. Have to play w it next week.
>>109426857>>109426890i'm actualy setting up pi right now, i got tired of opencode's bullshit.it may be npmslop but at least i have more control over it i guess.
>>109425861You could also just have a generic NPU. Modern LLMs share basically the same common ops. It's not like you have enough area to hold architecture-specific detail anyway.
>>109426543giwtwm
>>109426927I've had good luck with Claude code using DS as engine. But its pretty locked up. Pi I'm interested in as replacement for OpenClaw and Hermes, both of which are update unstable. But if I can make a funnier coding assistant thats just a bonus.
>>109427146never used openclaw or hermes, i just don't see the point of those desu
>>109427177They're for personal productivity, they are not (IMO) optimized around coding though some use them for that. Used for stuff like checking email boxes, running LLM-enabled cron jobs, web research. They're handy but both OC and H are unstable... you spend time setting them up with all the oauths / etc. and then they kill themselves a week later. I don't have time for that sort of nonsense, so still looking for alts.
>>109427326>Used for stuff like checking email boxesyea not my llm's business lol>running LLM-enabled cron jobssuch as ?legitimately curious to get a good example.>web researchi mean any harness can do it.>and then they kill themselves a week laterlol