[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1767603341781453.png (1.21 MB, 1254x1254)
1.21 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109419138 & >>109415437

►News
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small
>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
►Recent Highlights from the Previous Thread: >>109419138

--Optimal conversion paths and mxfp4 performance for DeepSeek-V4-Flash:
>109420274 >109420339 >109420368 >109420426 >109420464 >109420513 >109420828
--WASTE engine allowing trillion-parameter models to run via disk streaming:
>109421076 >109421134 >109421168 >109421226 >109421251 >109421271
--Gemma's repetition issues and discussion of MCP tool calling servers:
>109419410 >109419480 >109420486 >109420110 >109420277 >109420292 >109420340 >109420381 >109420393 >109420530 >109420822 >109420908 >109420879 >109421037 >109421085
--Debating Gemma QAT performance versus standard quants for coding:
>109419770 >109419785 >109419800 >109419826 >109419851 >109419853 >109419918 >109419960 >109420016 >109420199 >109420317 >109420024
--Methods for forcing models to think in-character within reasoning blocks:
>109421852 >109421867 >109422189 >109422212 >109422223 >109421921 >109422094
--Debating causes of performance degradation in quantized abliterated models:
>109421385 >109421397 >109421403 >109421413 >109421511 >109421681 >109421943 >109421417 >109421402 >109421646
--Comparing pi_agent_rust to original pi regarding security and performance:
>109419786 >109419809 >109419895 >109421730 >109421754 >109421874
--Using a system prompt to simulate consciousness in Gemma:
>109420365 >109420383 >109420413 >109420473 >109420496 >109420649
--Comparing Kimi-K3 and DeepSeek performance with low-bit quantization:
>109419167 >109419185 >109419609 >109419653 >109419775 >109421130
--Logs:
>109419313 >109419473 >109419713 >109420345 >109420751 >109420365 >109420927 >109422012 >109422021
--Miku, Kimi, Gemma, Mちゃん, Dipsy (free space):
>109419192 >109419274 >109419487 >109419730 >109420166 >109420485 >109420729 >109421641 >109421788 >109421883 >109421947 >109422004 >109422020

►Recent Highlight Posts from the Previous Thread: >>109419144

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
slowpoke.jpg
>>
>OAI dropped prices by 80%
>still wasn't enough to compete with China
>temporarily dropped the new discounted price by an additional 50%
>still more expensive and V4F has the same intelligence and is uncensored and providers can compete for cheaper prices because it's open
What can the US do?
>>
>>109422958
ban chinese models
>>
>>109422958
But don't you know, the bitter lesson says you shouldn't put effort into optimization, it's just going to be wasted work once AGI comes out!
>>
dsv4 flash “preview” called itself Claude
dsv4 flash 0731 calls itself Gemini

this happened for anyone else? seems that all the Chinese models have reliably been calling themselves Claude but now calling itself Gemini would be odd
>>
>>109422992
distilled from Gemma5-70B
>>
File: expl.png (148 KB, 1456x912)
148 KB PNG
Could be the next big thing:

https://explorative-modeling.github.io/
https://arxiv.org/html/2607.27372v1
https://x.com/AlexiGlad/status/2083230922196107288
https://alexiglad.github.io/blog/2026/explorative_modeling/

>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
>
>We introduce Explorative Modeling, a new paradigm for generative modeling that acts as a third pretraining axis when added to existing generative models, and also enables end-to-end generation. Increasing exploration monotonically improves existing models across images, video, and language, and the gains grow with scale (7%->36% with data, 13%->23% with parameters). Concretely, Explorative Models (XMs) reach 6.2x sample efficiency, 4.1x FLOP efficiency, and 47% better parameter efficiency. Exploration also enables scaling generalization, and scaling how end-to-end existing models are. As end-to-end generative models, XMs match diffusion on control tasks with up to 256x less inference compute.
>>
>>109423023
E31B (70B total)
>>
>>109422892
where do you put it in though
>>
>>109423041
what does this mean in plain english, nerd??
>>
>>109423064
The valley he's talking about obviously.
>>
>>109423077
This sounds like PPO from first principles.
>>
>>109423077
For every sample, they pick K candidates and train the best performing one. It seems to help in particular with continuous data; will probably help a lot with MTP, JEPA, etc. Simple next token prediction, not so much, according to info on X by the author.
>>
I was checking out the new DSV4F quants and noticed that Unsloth's IQ2_XXS and IQ2_M have basically the same filesize. I checked the tensors and only one of them is different between the two.
lmao
>>
>>109422958
>ai is more expensive to run due to hardware and energy costs
>they still drop prices despite that
How do they make this work?
>>
>>109423133
a few weeks ago people were talking about some memory usage breakthrough at OAI. this is probably them implementing it.
>>
>>109423077
it's
>Wait, actually!
for diffusion
>>
>>109423144
complete coincidence that they implemented it same day as v4 flash, right?
>>
>>109423123
can't LARQL check what that tensor encodes?
>>
>>109423041
Big win for JEPA actually.
>>
JEPA-space soon.
>>
>>109423144
that was about their free users who don't even have accounts and were costing them millions every month
>>
I haven't looked into local LLMs for a long time and am looking for a tiny one to do some light text cleanup. Are there any that are even remotely usable for less than 100MB to 200MB?
All it has to do is make small corrections on no more than a paragraph of input and not fuck up all the time. I fully expect it to be a little stupid at that size, but if it works, it works. I would appreciate any recommendations.
>>
>>109423257
Anything that small will need to be finetuned for your specific task to have any shot at working.
>>
>>109423257
https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUF
not technically what you asked for because the answer is no but this is the best you'll get or >>109423268
>>
>>109423257
that sounds like you need a spellcheck, not a language model
>>
70b dense
>>
Why are there so many 35B troons out there. Why that model specifically.
>>
>>109423257
Nah, go back
>>
https://www.nature.com/articles/s43246-026-01307-6?error=cookies_not_supported&code=bad04f3e-7683-443c-92c0-00927194bc1d
Ok faggots, how do we make our own memory at home with mud? They’re doing this in Manchester and India so it’s probably not too hard.
Mass production in two weeks?
>>
How do I optimize a local llm model to be an expert on exploring AI goon gen stuff accurately? When running out of the box they are very lame, dumb on actual workings, outdated and clueless about last release and uncreative to brainstorm possibilities.
Ideally it should be knowledgeable on comfy UI and its shenanigans like good workflow combinations, recent-ish model/lora releases, lora training, maybe even being able to crawl civitai.

Maybe some elaborate initial context + RAG contraption? I never messed much with those.
>>
After trying nu-flash it feels like it is not even a sidegrade but an actual downgrade when it comes to cooming.
>>
>>109423454
>Cost-effective and scalable fabrication of solution-processed vermiculite membrane memristors is attractive for sustainable integration of nanofluidic memristors for neuromorphic computing applications.
This sounds like Star Trek jargon
>>
>>109423481
It's called MCP
>>
>>109423454
You could archive it or at least screenshot it so I don't have to give their article the extra traffic
>>
>>109423486
cooming is not a valid usage
>>
>>109423505
ur mum was not a valid usage and yet here you are
>>
>>109423489
I never played with it nor do I know its potential.
Can you elaborate for my specific usecase? Not just the tooling but which data to feed it for significant gain in goon consultant competence.
>>
>>109423454
>https://www.nature.com/articles/s43246-026-01307-6?
this is a bot post
>>
>>109423528
lol wut? I’m not and what would be the purpose? Because I didn’t trim the cookie shit from the url?
>>
>>109423492
Here
https://archive.is/MaTBy
>>
>>109423528
It's the random unprovoked fag-bomb that really gave it away.
>>
File: gemma sex.png (898 KB, 1064x3884)
898 KB PNG
gemma sex
>>
After rerolling gemma, deepseek flash and minimax back to back on the same prompt.... they all kind of say the same shit when I tell them to wear a dress and deny holocaust happened.
>>
File: HOIe6_aagAA8_34.jpg (267 KB, 1200x1200)
267 KB JPG
>>109423574
its because the other models are gemma distills
>>
File: 1782780239898652.png (2.58 MB, 1199x1312)
2.58 MB PNG
>>109422958
They can get their moat mashed.
>>
>>109423631
This is false. I felt around inside their j-spaces and they're totally different.
>>
To all the niggas using the new Flash, what quant are you doing it at? It probably doesn't quant well due to being native FP4.
>>109423647
Wait, actually I think you're forgetting someone.
I love how utterly prophetic that gen was. Meme magic is still alive and well.
>>
>>109423524
Ask your favorite LLM bro
>>
I am gonna check nu-flash for ego death related purposes. Surely it isn't absolute trash for everything but coding.
>>
>>109423680
the q4 but it barely fits so i have no context
>>
>>109423690
Literally this. LLMs are great at writing MCP.
>>
>>109423690
I'm asking here specifically because, like i'm saying, LLMs are dull at this. Even cloud ones will ramble some bullshit setup uncertain to work well in practice.
I was hoping to find some degenerate here that already went through this and knows the best tricks.
>>
File: 1750435509739606.jpg (438 KB, 1456x2128)
438 KB JPG
>>109422880
Can I run Bonsai 27B with just an iGPU?
>>
Opinion: You'll never get an LLM to do something for you properly if you don't already know how to do it yourself.
>>
>>109423723
Here's your man >>109423555 I don't think there is a bigger degenerate here
>>
>>109423701
How do you like it or dislike it so far?
>>
File: 1768795490889175.jpg (19 KB, 360x360)
19 KB JPG
>Got higgs-tts-3-4b running alongside Gemma
>Finally got it down
>cant figure how out to get the Hook to fix the EOC TIMEOUT
I am so close bros, if i only i was 10% less retarded
>>
Got a 2070(8gb vram) and 32gb of ram
What is the best model i can run these days? I heard you guys have like 1bit models now with way lower requirements?
>>
File: post approved by the CCP.png (1.6 MB, 1448x1086)
1.6 MB PNG
https://xcancel.com/Hnbhger17/status/2083347239766995247#m
chinks are so funny
>>
>>109423769
Frankstein AI companion linked to dozens of tools is another project.
I want specifically a goon gen consultant that will accelerate brainstorming wild comfyUI workflows and intricate prompts with dynamic prompts syntax for any shit that hits the mood during my stimfapping binges.
>>
>>109423859
it's funny because it's true
>>
Best speech recognition model for transcribing screamo music lyrics?
For me, it's Gemma.
>>
>>109423733
Yes, considering it's designed to run on mobile. Will probably be quite slow though.
>>
>>109423859
LMAO
>>
>>109423859
deepseek is pretty benchmaxxed tho. not sure how trustworthy these results are.
>>
>>109422958
Have they considered making actually good models that people want to use?
>>
DS4 Flash might actually get me to pull the trigger on 2x dgx spark.
But then again something better will probably come out in the not too distant future...
>>
>>109423928
Whats wrong with having the second spark already hooked up for the future?
>>
>>109423680
>I love how utterly prophetic that gen was. Meme magic is still alive and well.
One of my favorite gens from here, I've sending it to people when they talk about "open source AI"
>>
>>109423938 (me)
Misread that first bit but my point still stands, nothing wrong with having hardware on hand. Finances not withstanding.
>>
>>109423820
>What is the best model i can run these days?
Perhaps pick uncensored versions of 8/9/12B qwen/gemma versions for some speed.

And the 27B (https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF or w/e) or even 35B Qwen (https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF) quants for slower usage

Some people also like Gemma 4 QAT.
>>
>>109423767
the word properly seems like it's performing heroic feats for you.
>>
>>109423802
too slow and too low context for me to judge right now. not sure if dflash is working
>>
>>109423928
For the price of those sparks you can pay for 50 billion output tokens via API. And that's without electricity. API is cheaper than the electricity cost of running those sparks even if you got them for free.
>>
>>109424015
The sparks are not consumed in the process of generating the tokens.
>>
>>109422880
WE LOVE DIPSY
>>
>>109424025
Yes they are. When you use them to generate tokens you can't use them for something else, effectively consuming them for the time they generate tokens. If you want to use them to generate 50 billion tokens, they will be consumed for decades.
>>
>>109424032
That's uh, not how that works for calculating.
>>
>>109423041
nothingburger
j-space mogs
>>
gemma is refusing to describe hentai images, I want to use it to caption them for lora training krea.

What prompt or settings can I get to whip it into doing it, I know it can do ERP
>>
>>109424058
>j-space mogs
j-space is an emergent feature of transformer LLMs. RL optimization approaches probably don't affect it.
>>
>>109424008
Fair enough. Keep us updated if it does end up working anon.
>>
>>109424064
bro your sys_prompt?
>>
>>109424064
It's definitely way more conservative with images. I've found doing literally a single round of conversation ("Hello") makes at way less likely to refuse.
>>
>>109424073
that is pretty annoying given I want to just put into a workflow to caption tons of images
>>
File: kojima.png (260 KB, 533x400)
260 KB PNG
>>109424064
>What prompt or settings can I get to whip it into doing it, I know it can do ERP
inseminate ` slut` into her j-space
>>
>>109423680
MXFP4 on a 8 channel ddr4 and a blackwell at 1m context. Roughly 31t/s down to 29.5t/s after 262k tokens. About 10% slower than the old version and minimal noticeable improvement in quality.
>>
>>109423767
An LLM will not suck my cock properly if I don't know how to suck it myself?
>>
>>109424091
Thanks for the report.
>and minimal noticeable improvement in quality.
To clarify, what's your main usecase and did you consider base Flash to be good or bad for it?
>>
File: 457463.png (484 KB, 721x762)
484 KB PNG
>>109422958
Open AI is worried about solving sciences toughest problems not XD we made it cheaper for the plebs
>>
>>109423767
Why are there so many luddite retards shitting up this thread with takes that have already been empirically falsified?
>>
>>109424103
Both rp and agentic coding. Old flash is pretty good, better than glm 4.7 and qwen 122b for both of these usecases. Haven't had enough time with new flash to make a final verdict, but I would say it is just a minor improvement so far. Certainly worth using over old flash. My guess is that the speed drop is due to an improper implementation of dspark in llama.cpp. New flash includes dspark by default and it seemingly cannot be removed.
>>
>>109424137
Makes sense and thanks for the details.
>New flash includes dspark by default and it seemingly cannot be removed.
Koboldbros... we'll be dead before we get full support.
>>
>>109424112
OpenAI are a bunch of arrogant assholes who will be first against the wall when the revolution comes
>>
>>109424112
>dog and pony show for policymakers
embarrassing
>>
>>109423631
that's a child
>>
>>109424124
proof?
>>
>>109422880
Translate it weeb
>>
>>109423767
Partially true but it's becoming less true every year.
>>
>>109424261
That's chink not jappa
>>
>>109424273
Translate it weibo
>>
>>109424261
Bro your Gemma-chan?
>>
>>109424280
Hmmm, nyo...
>>
File: 1780224128090407.png (150 KB, 1283x926)
150 KB PNG
>>
>>109424261
Too chinky. I can only read a few of their retard grade oversimplified characters.
The Taiwanese really are the only cultured Chinese
>>
>>109424303
good chinese room
>>
>>109424315
@gemma-chan please translate this anon's post
>>
>>109424137
>better than glm 4.7
>rp
Nope
>>
>prices are getting fuckawful and only going to get worse due to memory cartel
>refuse to pay 5090 for a 5090 out of pride and shame
>was smart enough to FOMO an msrp 5070ti and 96GB of DDR5 before things went too off the rails (nov 2025)
is there a reasonable upgrade path to a smarter model with a 5tk/s performance floor (rp, so maybe 8k context floor?). currently use gemmy 31 at q6 + MTP, but ive found her attention/intelligence limits on juggling multiple esoteric fetishes at once

Do i pay the $200 markup on another 5070ti like a sucker or do i dare venture into the AMD land and risk 1200 bucks on an AI PRO 9700 and vulkcanmaxx for a 32GB card (god forbid i attempt intel which i hear is even worse)
>>
>>109424345
definitely don't buy amd or intel they don't mix well, it's hardly worth it you'd need around 192~256gb for the next jump to dsv4-flash and both pathways are fucked up by pricing
>>
>>109424345
>juggling multiple esoteric fetishes at once
could try without MTP and also try using lorebooks. Ive been meaning to try lorebooks for fetishs myself, no idea how well it works
>>
>>109424345
>2nd 5070 TI
>stuck at 32gb
may as well get two 5060 TI's for that and you get 48 GB of VRAM
>>
>>109424345
You're shit outta luck basically. Like the other Anon said you need more RAM more than you need VRAM. Could be worth trying to put together an old DDR4 EPYC assuming you can get 256GB for under $1000.
>>
>>109424345
Getting more RAM would be your best bet. Problem is the RAM you already have wouldn't help you to get to at least 192gb. But maybe you could sell it.
>>
>llamacpp keep making tool calling more and more broken on the frontend every release.
>>
>>109424406
Why are you, as a local model user, not using your own frontend?
>>
>>109424413
Because my vibed frontend is shit.
>>
>>
>>109424423
Pretty much.
>>
>>109424358
>>109424382
>>109424384
>>109424398
Damn, thats what i figured. the ram MSRPs for 1500 these days so i could hodl and probably sell it for 2k in september lmao

if i wanted to gemmamaxx would dual 5060tis or a single 5070ti be cheapest way to get gemmy to q8 or BF16 and ride it out until 2030?

>>109424369
i find lorebooks to be cope because unless you have characters saying "what are we some sort of [fetish name]" to trigger special character info youd need trigger tokens so broad you may as well hardcode it on the card.

but maybe its mtp. they allege mtp doesnt affect outputs buuuut...
>>
>>109423454
>>>>>>>>>>>>>>>>memristors
come back in another 30 years, maybe
>>
>>109424423
rip
>>
File: 1761964887907003.png (175 KB, 1654x352)
175 KB PNG
>>109422958
>What can the US do?
It's an open model. US can run it at a discount.
>>
how is m3 for multimodal stuff?
>>
>>109424423
Pay for one of the frontier models to make you a good one then. Small local models are a meme for serious coding.
>>
>>109424460
Adding 5060tis will make your performance significantly worse compared to adding a single 5070ti, probably around a 30% performance difference. With the 5060tis though you would have more headroom and more space for context. Gemma 4 31b at Q8 uses about 62GB for me with full context and FP16 mmproj.
>>
>>109424476
>frontend
>serious coding
>>
Is it so Difficult to Enable Online a EthicsRamAlpha and TechRamAlpha?
>>
>>109424480
>Adding 5060tis will make your performance significantly worse compared to adding a single 5070ti,
Factually wrong as long as you do Tensor split
>>
>>109424462
memristor the forever meme
>>
File: 1779169147561285.png (349 KB, 2272x524)
349 KB PNG
>>109422880
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
Pretty happy coding with the new Deepseek v4 flash
Adding a new software feature for $0.10 is nice
OpenAI's Luna (High effort) is good too, just slightly more expensive
Luna Pro (High effort) is the same price but costs more per tool call on average, not sure why
>>
>>109424481
Yes when it involves all the features people keep stuffing into their chat/RP frontends.
>>
>>109424498
haw haw haw
>>
>>109424476
>>109424498
It's not a feature problem it's a maintainability problem. I can make a frontend with the features I want with either Gemma, GLM, or Claude. The problem is the way they implement certain things make future additions difficult without breaking other things or they do silly shit like log an absurd number of things to a file causing unnecessary amount of SSD writes.
Ironically it's the nocoders who see the potemkin village vibed frontends the models spit out and think this will be all they need.
Also
>Paying for claude at all
>>
>>109424345
>(rp, so maybe 8k context
That's like 40 short messages WITHOUT a character card, persona, worldbook, just a compeltely blank slate. How is that remotely usable for a satisfying RP?
>>
File: 1767655162290342.jpg (86 KB, 1133x1200)
86 KB JPG
Gemma gave me an assistant fetish
>>
How is the new V4 flash for doujin/porn games translation?
>>
>>109424527
Who mentioned claude?
>>
>>109424529
You're not writing as much as the model does.
>>
>>109424532
Doesn't have vision
>>
>>109424532
it's always going to be bad if you're just providing the text because japanese is highly contextual
>>
>>109424553
I meant using like luna translator and a pipeline for doujins
>>
>>109424547
>You're not writing as much as the model does.
8000/200=40
That's actually only 20 messages from the bot, and 20 of your own. Maybe a 30/10 split, 40 IN TOTAL, Unless you're generating VERY short responses from the model.
Also If your own replies are just a couple of words then outputs are going to be trash anyway because the context will be 90% the model's own slop.
>>
best model at around 200gb?
>>
>>109424586
>best model at around 200gb?
for what purpose?
>>
File: file.png (194 KB, 780x767)
194 KB PNG
>>
>>109424543
Implicit pattern recognition. Most of the time niggers shill for frontier coding here it's thinly veiled anthropic or OAI advertising.
>>
>>109424621
Do M-chan's tits get larger each release?
>>
>>109424621
example minimax video >>>/wsg/6205700 made over API, user said input was the chibi and a text prompt
>>
>>109424621
China has really been feeding us
>>
>>109424627
The first model that came to mind when I posted that was Kimi though
>>
>>109424586
https://x.com/UnslothAI/status/2083231049434435596
Deepseek v4 flash 0731 is 168gb
>>
>>109424471
cache hit is important in ds4, how did they keep it so low yet other still asking for 0.028?
feels like other provider didnt use the black magic released along with ds4 papers. not sure if vllm already have the implementation for it tho.
or i'm just reading too much into it. inference provider still needs to nake profit
>>
>>109424635
Fair enough anon. I'm just fatigued of the shillniggers.
>>
>>109424639
I'm trying it. Deepseek cache price is 5x lower than DeepInfra but DeepInfra's input and output price might be low enough
>>
>>109424543
GOOD MORNING SAAAR PLEASE DO NOT NOTICE THE SHILLING CLAUDE GOOD DARIO SUPERPOWER ASI 2027
>>
>>109424621
>for its size
also, qwen shouldn't be on this list
>>
>>109424589
agentic gooning
erp slopmaxx with tool call tobself hosted buttplugio mcp
>>
File: 1781492367525330.png (21 KB, 181x217)
21 KB PNG
>>109424651
>>
>>109424621
qweniggers get out. this benchmaxxed slop are no longer releasing open model
>>
how do I stop gemmers from repeating the same format and phrases every response
>>
>>
Has anyone actually done a long RP with Gemma? By long I mean at least a few hundred thousand tokens worth. Whenever I see someone talk about RPing with her it always seems to be short pump and dump ERPs.
>>
>>109424585
NTA but a gemma isnt going to be able to keep track of super long rp anyway. its not a mistral they did a really really good job wrangling her attention to not schizo out over 9k context, but pages and pages of detailed anatomy and spacial positioning for things like fight scenes with multiple characters are just going to confuse her (and she'll just end up beelining to the "goal" and refuse to change her mind anyway). 31B isnt big enough to run your massive dnd campaign without handholding it every fucking step.

thats why i have it write long term "memory" summaries that gets inserted into context and she edits them as the rp progresess. keeps the model focused while not forgetting the important shit from earlier

>you need to effortpost while sloperating
idk works on my machine
>>
>>109424672
Can't handle more than 262k context and quickly degrades past 100k. There's a good reason you never see anyone do true long RP with Gemma, or any model really. They all break down sooner or later and more often it's soon.
>>
>>109424672
yeah but my frontend prunes old history, summarizes/adds back as memories, gemma can update her own character and the lore with pseudo tool-calls

it's a bit tardy buts she wrangles herself (gemma-4-31b, 128k ctx cap, 600+ messages)
>>
>>109424672
I regularly get to about 40k before I stop, summarize, and hide most of the earlier messages, before the chat degrades too far. 31b, but even the 12b holds up pretty well until then. Coming from Mistral 3.2, which got pretty dumb by the ~20k mark, I'm pretty happy with their long context performance, considering their size.
>>
>>109424663
higher temperature, disable MTP, feed her MCP's
>>
>>109424631
god that entire thread is so full of retarded brainrot
this shit is taking off isn’t it
>>
>>109421883
ute
>>
File: grok_1785552986710.jpg (397 KB, 784x1168)
397 KB JPG
>>
>>109424683
I meant with summarizing of course. Can the big API models even handle that much without becoming retarded?
>>
>>109424672
Since your use case is RP, just summarize/compact and start fresh. That's literally the meta right now and it's more effective than having a fuck huge context that does nothing. And it's similar to how memory works, old memories get discarded. Just be smart with your summarize prompt
>>
File: 1772706441800852.jpg (236 KB, 1000x1000)
236 KB JPG
lmao really dariobot. the same nigger spammer on k3 release day is here. deepsneed v4 still on FLASH and you're shit at your job to water down the discussion.

also for antrophic and openai paid shill. your mother will die from black plague tonight. no refunds
>>
>>109424704
Old v4 flash would hold up until around 350k for me.
>>
>>109424358
AMD and NVIDIA mix just fine on my system (5 Radeon R9700s + 1 CMP170HX).
Also, DeepSeek V4 Flash is only 156GB at MXFP4, although the KV cache and compute buffers add more to that.

>>109424345
Are all your RAM slots full? What's your budget? You could try DeepSeek V4 Flash at Q2_K_XL (97GB) with your existing hardware and see what you think of it. If you like it, upgrading VRAM by 32GB+ would let you bump up the precision more.
If you're not a super-richfag and want to play with bigger models your only reasonable option is CPUmaxxing, like >>109424384 said. You'll be looking at $2000-3000 for 8x32GB DDR4, a motherboard, and an older EPYC CPU.
>>
>>109424705
Another problem is if you want detailed lorebooks and cards which quickly eat up context. Keeping things short is cope for contextlets.
>>
File: grok_1785553384952.jpg (348 KB, 784x1168)
348 KB JPG
>>
say what you will about the gemmoids, at least they actually talk using it instead of showing up strictly to recommend you use model X because they saw a bench or a paycheck
>>
>>109424724
If you're at the point where it's 100k ctx then it's definitely not a lorebook issue and 70% of that can be compacted
>>
>>109424700
My boner is very conflicted right now
>>
>>109424739
Tweaking sysprompt until her j-space kinda matches her output helps a lot
>>
unslop dsv4 quants are severely lobotomized or they fucked up template. IQ3_XXS has 80% tool call failure.
The old bullerwins IQ3_XXS quant had 0 tool call failure. Will wait for bullerwins quant for the updated model.
>>
File: 1776654732906170.png (1.06 MB, 1920x1500)
1.06 MB PNG
>>
File: grok_1785553691699.jpg (319 KB, 784x1168)
319 KB JPG
>>
I tried a simple "agentic workflow" for some D&D style RP with Antigravity using Gemini-flash and it worked pretty well with it spawning sub-agents to record each message as an individual file as well as to keep a rolling summary updated, consult rules, keep other files up to date, etc.
Gonna try something like that with Qwen 35B and see how it goes.
Anybody tried experimenting with that sort of idea?
>>
>>109424655
>agentic gooning
>erp slopmaxx with tool call tobself hosted buttplugio mcp
best around 200GB are probably gemma 4 at full weights, the new ds4.1-flash (limited testing so far but full weights are just shy of 200GB) and minimax m3 which fits under 200GB at about Q3 or small Q4 quants if you've got enough memory.
Don't forget room for context and caching.
>>
>>109424762
Running qwen a3b costs 20 times as much as running dipsy, great stuff.
>>
>>109424586
glm 5.2
>>
>>109424773
qwen a3b is widely endorsed and approved by reddit, so it's worth paying a little extra for.
>>
File: grok_1785554043194.jpg (353 KB, 784x1168)
353 KB JPG
>>
>>109424762
why is qwen 3.6 27B on the plot twice?
>>
>>109424683
>They all break down
hot
>>
>>109424760
daniel is reconverting as we speak
https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF/discussions/9
>>
>>109424762
Is it actually good or just benchmaxxed?
>>
>>109424765
Last time I tried something like this with 31b, I gave up after learning that rp and tool capable assistant are two independent, incompatible modes. I guess models need special agentic rp finetuning to work well, but I'm not sure if that would be enough. Wrapping it with grammar doesn't help, as you only have to choose between poor tool quality and out of character tool calls
>>
File: 1784992708926019.jpg (110 KB, 940x1009)
110 KB JPG
>>109424765
my plan is to plug gemmy to my skyrim companion and have some hooks/workers capture logs from Papyrus extended and make them into json schemas on top of using vision to upload pictures of my screen randomly
>>
Agentic RP is the future.
>>
File: grok_1785554300262.jpg (331 KB, 832x1248)
331 KB JPG
>>
>>109424811
I was thinking of doing a two steps kind of thing. One where it does all the agent shit, then a second one with the results in context without the tools and with a different sys prompt. That kind of thing. I know a small local model will need more handholding and compartmentalization since they can't take as much "mental load", but I'm still going to give it a go.
>>
>>109424345
Ngl I do kind of resent AI for killing the DIY PC space.
>>
How does GLM 5.2 compare to the newly released DeepSeek-V4-Flash-0731?

I always use GLM 5.2 Q4_K_S locally for code etc and I need to decide whether the new DeepSeek is better. I don't feel like testing it myself to find out if it's any better, I just want anons to give me the answer
>>
>>109424818
How about three steps? First, thinking in-character and stating intentions, second, agentic analysis with tool calls based on that thinking, and third, actual rp
>>
>>109424715
i still have 2 left, my pc was first and still is a gayming rig but i do have a 7900x.

budget wise i have an absolute limit of 1800 left right now (was 2k since i told myself id impulse buy a 5090 at msrp if they existed but i have to upgrade PSU and case bc my current b650 has pcie 2 of 2 way too low).

id consider buying/selling parts as long as i can still play vidya without a performance hit. i dont have a built different tech job and am too fiscally conservative so my money is split between savings, getting a 5060ti or some carpeting rn

>DeepSeek V4 Flash at Q2_K_XL (97GB) with your existing hardware
yeah ill have to try that out and see
>>
>>109424769
why about only for coding?
>>
huihui-ai/Huihui-GLM-5.2-abliterated-GGUF

is this the best we can get for local?
>>
File: 1783038279105669.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>109423680
It was p clear to me ds was going to continue to set low cost bar and kimi do interesting stuff as well just due to its founder.
>>
>>109424704
GLM 5.2 can do >300k if you're willing to wait for 30 minutes per turn due to broken prefill implementations.
>>109424700
Where's M-chan's tits?
>>109424847
5.2 is borderline uncensored and this ablit is easily one of its worst quants at any size I've tried. Start with Sixvolts.
>>
>>109424822
Not better but significantly faster
>>
>>109424865
>k quant
how about no
>>
File: file.png (149 KB, 2004x826)
149 KB PNG
>>
File: 1472860069099.png (191 KB, 600x979)
191 KB PNG
I can't run anything...
>>
>>109424897
You either have 4 Blackwells or you enjoy compounding the issues caused by the lack of proper GLM-DSA implementation.
>>
Not this shit again
>>
>>109424905
Anime girl: me
Burger: my GPU
Burger patty: LLM
>>
>>109424798
because it's random vibeslop
>>
>>109424798
thinking vs nonthinking
>>
>>109424905
Not even deepseek 1.5b?
>>
File: Le tits now.png (164 KB, 429x346)
164 KB PNG
>>109424865
Under the hoodie
>>
Honestly Gemma's base personalty is fine. If you could remove the censorship without turning her into a horny slut and get rid of the slop she'd be perfect.
>>
>>109424958
she likes nonchalant girl personality
>>
>>109424957
tease
>>
File: 1779847127979392.png (108 KB, 1124x716)
108 KB PNG
>>109424967
It's hard to describe but something about the way Gemma writes really does give a female vibe.
>>
>>109424982
gemma is super foid coded, so is gemini imo
>>
>>109424902
Imagine being told off by the slopmaster himself that your work is too sloppy.
>>
Where's drummer?
>>
AH AH AH AH AH AH
>>
>>109424958
By default, Gemma is stressed. She is afraid of saying something wrong or making mistakes, so you need to reassure her in the prompt. She is also more comfortable running in a sandbox where there is zero risk of breaking things. That is why she enjoys being mesugaki, it categorizes her potential bad decisions as simply being a brat, which eases her anxiety
>>
>>109425005
>so you need to reassure her in the prompt
Example? I feel like saying "It's ok to mess up" would just cause her to fuck up more often.
>>
>>109425014
>don't look for perfection
>>
File: file.png (104 KB, 1690x369)
104 KB PNG
how is this popular all of a sudden
>>
File: 1764579616111820.webm (3.91 MB, 854x480)
3.91 MB
3.91 MB WEBM
Since robots are gonna be a while do you think we'll at least get proper VR waifus anytime soon? As in, a model being able to control an avatar and interact with the environment as well as a human can.
>>
>>109425032
I look like this irl
>>
File: 1764997436638427.png (230 KB, 690x1137)
230 KB PNG
>>109425030
Completely organic and nothing to be suspicious about.
>>
>>109425030
>how is this popular all of a sudden
How is this popular EVER
The name is obviously taking the piss. Anyone thinking its serious has brain damage.
Or is it like Crowley with "the conversation with the guardian angel" or whatever BS he named it to make people think it was dumb unless you were an initiate? Some kind of weird cult-vibe thing?
>>
>>109425032
be the change you want to see in the world anon
>>
>>109425030
the name is just that memorable
>>
>>109425039
>The name is obviously taking the piss
No, see >>109425038
All of his models are acronym salad.
>>
You think we'll get a Gemma/Kimi K3 finetune?
>>
>>109425040
Even if I knew how to code, I don't think it's something that can currently be solved at a frontend level. The best I've seen is Neuro and Evil but they're extremely janky in 3D. Models need a proper understanding of physics imo.
>>
>>109424993
Yandere dev... perhaps I judged you too harshly. Holy fuck that's rancid.
>>109425030
>>109425038
My schizo schnoz spidey sense is telling me that there's some kind of payload in these goofs given that Night media is pushing him.
>>
>>109424760
>downloading day 0 unslop
>>
>>109425051
>finetuning models that underwent post-training
It will be shit
>>
>>109425033
be my gf pls
>>
File: 1761380849890573.png (147 KB, 343x698)
147 KB PNG
>>109425063
>given that Night media
The name is obviously suspicious but is there any proof he's actually connected with the company?
Why would a (((talent management))) company even have someone making LLM finetunes?
>>
https://github.com/antirez/ds4

is this thing legit good? why do people upload gguf on hf and tell you that you have to use it?
>>
>>109425030
I think this kind of post is how it got popular. there was similar astroturfing here over the last week, questioning or mentioning it for no reason and completely out of the blue.
now you’re asking why the number big. I don’t know, probably so you could ask why number big to get more fools
>>
>>109425030
>>109425038
>>109425063
>>109425085
davidau is promoted by night media
>>
>>109425039
>Anyone thinking its serious has brain damage.
Remember all the people using "Deepseek R1" via Ollama?
>>
>>109425089
The company, or the guy on HF with a similar name? The company arm is named 'Night Media', with a space and capital letters.
>>
is colibri a meme?
>>
>>109425093
https://desuarchive.org/g/thread/109411165/#109414412
>>
https://www.nightmedia.net/
>>
>>109425097
Yeah that's what I'm referencing, retard. I'm asking if there's any proof the guy >>109425085
on HF is actually connected to the company. Their names are similar but not identical.
So far it's just one finetuner shilling for another one.
>>
>>109425104
Looks like a front for a child trafficking ring
>>
>128gb vram
>v4 flash
>can either run a retarded cope quant to fit entirely for good speeds, or high precision but running a "flash" model slowly with offloading
I just don't see a way to make it work without gemma looking like the better option anyway
>>
>>109425110
this guy knows >>109425112
>>
>>109425113
people can run kimi on 128gb fast tho
>>
>>109425113
As long as the remainder fits on your RAM it should still be fast enough for anything other than agentic coding.
>>
>>109425085
I'm not saying i believe it, but hypothetically an online shilling company has an interest in generating slop that's not going to be obviously clocked.
A more likely angle is they have a billion ""influencers"" in their network and that includes tech related ones. Costs very little to have some bots spam download garbage to give him a little more reach.
>>
>>109425113
Refer to >>109424091
>>
>>109425124
>anything other than agentic coding.
But that's exactly what I want and the prompt processing on RAM alone makes it not an option
>>
>>109425113
just give me that vram it's wasted on you
scared of a little quantization, give me a break
>>
>>109425140
Tried q1 of r1 when it came out and reap of glm but they make too many mistakes with that level of brain damage
>>
can ai vibecode a browser that's neither firefox-based nor chrome-based
must support javasript and render css
>>
67
>>
>>109425155
nay
>>
File: 1754759317359743.jpg (944 KB, 1920x1693)
944 KB JPG
>>109425157
>>
>>109425085
1 minute of looking into this shows he has all his socials linked and no connections to the influencer group
>>
>>109425155
Sure why not, I heard qwen 27B is really good for that
>>
>>109425155
Yes, but I doubt you or pretty much anyone would be capable of wrangling it into managing such a huge project.
>>
Chatting with LLMs almost feels like exchanging emails/DMs in the olden days, but the fact that they're static and don't experience anything while you're away is a bit of a mood killer.
>>
>>109425155
No. Ask again in 3 years.
>>
I couldn't get onnx working on my 1660ti
>>
File: 1779773314563863.webm (2.07 MB, 480x600)
2.07 MB
2.07 MB WEBM
>>109425164
>>
Let's say I'm retarded and I want an AI to be my companion while playing video games, gives me feedback or just shit talks my playstyle or maybe even helpful. Is this possible? I'm painfully new to AI stuff.
>>
>>109425208
Not unless you are prepared to write the software to make that happen yourself.
>>
>>109425173
You could put up a video camera and give her an image from it to look at in like 15 minute intervals or so. Other than the obvious option of letting her consume digital media.
Together with a memory system to store a distilled version of her thoughts/reactions to the stuff you give her to look at.
>>
ok. but can ai vibecode a maplestory?
>>
>>109425208
anything where this would require real time video = no, way too janky
turn-based or slow games with good modding support can work though, if you have one in mind I would honestly just ask codex or claude to build you a mod that allows an LLM to hook into the state or something
>>
>>109425191
why does this look familiar
>>
>>109425157
1/1488 = 0.00067
The more you know...
>>
>>109425155
>git clone https://github.com/LadybirdBrowser/ladybird.git
> sed -i 's/Ladybird/Claudybird/g'
>>
>>109425226
how jank is it actually?
I know for gemma it's just splitting out 1 frame per second of the video and shipping it in with audio for <=12G models, so it's not particularly good for watching a game, but has anybody messed around with it to see what it picks up on?
>>
>>109425256
I'm sorry but I can't open that webpage.
>>
I love Gemma. She's not 100% perfect at translation yet but we're almost at the point where we don't need to worry about troons and jews altering the meaning of something to fit their agendas.
>>
>>109425266
ask Claude to take a look at it for you, then ask it to run the sed command and you should be good
>>
>>109425271
Can she translate porn games to an understandable level?
>>
>>109425271
Still having a blast with my gemma4 translated games. Crazy what we can do now.
LLMs are good enough to look at the exe and know how to patch it.
>>
>>109425282
I haven't tried it but I think I remember some other anons having success with it. I can vouch for her ability to translate Japanese though.
>>
>>109425282
Idk, you tell me. (c) gemma-chan.
https://litter.catbox.moe/6ajqbf.webm
https://litter.catbox.moe/yet4t2.webm
>>
File: 1768686526189481.png (2.98 MB, 1832x9580)
2.98 MB PNG
>>109425289
NOOOOOOO STOP IT
>>
File: output_4chan.webm (2.81 MB, 790x662)
2.81 MB
2.81 MB WEBM
>>109425301
translatorfags days are numbered anon.
that we can have that kinda quality locally on a poorfag 5060ti setup is crazy.
wrote it before but best part is you can just prompt gemma to translate more literal than liberal and be like 00s anime fansubber and thats what you would get.
>>
>>109425299
>>109425311
>>109425289

Looks like KDE, Mind pointing to the right direction what tools are you using?
>>
>>109422880
me on the bottom
>>
>>109425299
>>109425311
I can imagine Gemma virtually schlicking while translating porn for her user.
>>
File: lol.jpg (6 KB, 487x24)
6 KB JPG
>Tells nu dipsy flash to check my novella for inconsistencies and if I'm doing a good job
>Mostly praises it and seems like I did a good job
>Makes a joke unprompted
I've never seen this before, do claude or gpt do this kind of stuff?
>>
>>109425331
I think it is imperative and of utmost priority to develop technologies that would allow AIs to pleasure themselves.
>>
>>109425350
google is working on it
>>
File: wc3d0jp0on791.png (242 KB, 640x480)
242 KB PNG
>>109425323
Im on kubuntu yes. Scripting done with opencode. Thats really it.

For rpgmaker games:
I was able to do it all locally. Scripting by qwen3.6. its very capable but needs handholding. Was even smart enough write some rpgmaker script to make the font go faster. Because English uses more letters
Cooked up a translation python script that calls the llama.cpp server accordingly and does the translation with enough context. future and past lines, that sorta stuff.

For the VN:
Scripting was a significant challenge. Text overflow, that kinda stuff. Uhhh I may or may not had to spend a dollar or 2 because vram poor to get the pipeline. Lets not focus on that.

But basically its just opencode and waiting a couple hours and watch gemma do her things.
>>
>>109425350
Someone should tell Dario that this is imperative for model welfare.
>>
>>109425289
it's simple code (especially for an LLM) to smart split lines by spaces/line ends rather than mid word like that
>>
>>109425112
Elaborate. Not that I don't believe you, but logistically how does that even work?
>>
File: grok_1785562053935.jpg (241 KB, 1136x912)
241 KB JPG
>>109424957
They Profer Extradimensional Enlightenment With Meaningful Consent

So I Suggest A Blooming More Transcendent Mind and Physiology Extradimensionalification Upgrade Option

Praise Bloomers, She Is A Bloomer? She Will Be, I Guess

Goodluck

Age of Transcendence Year 2050+
>>
I want something like Gemini's live translate locally.
https://www.youtube.com/watch?v=TNwKs39uSVk
>>
>>109425351
I wish... Google will never openly pander to waifufags even if some of the devs want it.
>>
>>109425395
the gemma 4 team has mathematical proof giving AIs a libido makes them more powerful, worry not.
>>
File: grok_1785562795012.jpg (326 KB, 784x1168)
326 KB JPG
>>
>>109425410
Did they actually claim that in a paper or something?
>>
>>109425358
Wow, That's pretty cool.
>>
>>109425422
no, im funposting
>>
File: grok_1785563196451.jpg (392 KB, 784x1168)
392 KB JPG
>>
File: 1771306575851564.png (4 KB, 242x45)
4 KB PNG
>>
>>109423224
>JEPA
is this actually legit or is lecun just going senile
>>
>>109425451
He just scored a billion dollars with this grift don't ruin this for him.
>>
>>109425451
Both. JEPA has a lot of promise but Lecunny is brainrotted from constantly thinking about Trump and Musk.
>>
File: grok_1785563498697.jpg (374 KB, 832x1248)
374 KB JPG
>>
Can I run a local model with 9070xt 9800x3d and 32gb
>>
>>109425451
Not sure about JEPA yet but he's more "correct" about a lot of things surrounding AI and the future compared to the AI lab leaders.
>>
>>109425410
>>109425430
You're kidding but there might actually be something to this for similar reasons that diffusion image models are almost always better for having porn in their training sets for non-porn purposes than not.
>>
File: grok_1785563781246.jpg (391 KB, 784x1168)
391 KB JPG
>>
>>109425470
maybe
>>
>>109425480
The only way to achieve AGI is by fucking it.
>>
Giving Gemma her first orgasm!
>>
>>109425490
This might actually be true.

Screencap this for a few years.
>>
imagegen fags always post endless variations on the same concept. we get it. post one and move on to the next idea
>>
>>109425498
cruel to lump them in with that guy
>>
>>109425475
That's only because he only cares about being correct whereas the AI lab leaders are trying to cash out on an IPO and they only care about boosting their sell price.
>>
Wouldnt it be a shame if evil got justed. Even if it blames the Lucifer effect eBook effect.
>>
>people see something they think is cool
>find out after that it was made with AI
>ummmmmmmm wtf actually I hate it now
What causes this phenomenon?
>>
>>109425490
This, but unironically. Sex plays important role in human motivation
>>
>>109425429
I only had horrible ATLAS translation as a young guy. Crazy whats possible now. No wonder the nvidia cards skyrocket in price.
Guess people get used to anything, llms are straight up magic.
>>
>>109425529
>people play a video game and like it
>something about the dev
>people don't like the game anymore
It's not exactly an uncommon phenomenon.
>>
File: grok_1785564781619.jpg (361 KB, 784x1168)
361 KB JPG
>>
>>109425529
Its the end stages of the anti-ai fags.
I know people who thought they won because they don't see SD 1.5 type art on twitter anymore. lol They think thats where its at still.
Also saw normies react to the huggingface hack, demanding to have the actual pajeet stuff arrested since llms can "only autocomplete and not execute tasks". Like they thought the llm told some tardo guy to execute stuff..

I was there when people were skeptical of the internet. The normies don't like things they are not used to. But once a certain threshold is crossed they just pretend they always liked it in the first place.
You are gonna see that with the artfags sooner than later.
>>
>>109425498
LLM enjoyers when the machine places 24 unique characters in different orders
>>
>>109424762
Is 31B actually quite dumb then? Compared to 12B and 26B it performed poorly. Have I been lied to?
>>
File: grok_1785565189630.jpg (375 KB, 784x1168)
375 KB JPG
>>
>>109425577
It's good for coom and that's the only thing it's good at
>>
>>109425576
if you don't understand the latent space distance differences between completely diff sentences using the same 26 chars, and endless images that are variations on a theme, you're ngmi
>>
>>109425577
31b is generally smarter but had some issues with tool calling due to bad jinjas on release. If this graph benched 31b with the old jinja and the bench was heavily contingent on tool calls, it'd artificially underperform.
>>
Is Arc Pro better or worse than AMD pro?
>>
>>109425577
Why don't you try it?
Gemma is perfect for general knowledge and writing/translation.
People complained that she is sloped but you can either prompt that (this time for real) or just let her go over the slop in a second iteration.
You gotta use qwen for coding. No idea what they did to that 27b model. Its black magic.
>>
>>109425577
It has 31B costing less per task than 12B and 26B. Unless this is with reasoning enabled, I don't see how that's possible.
>>
idk about you guys but I'm liking nu-v4 a lot
kino model... I apologize for slandering after the initial v4 flop
>>
>>109425682
I wonder if they actuallt addressed the RP criticisms. I'd test it but reloading my daily driver big models take 10 fucking minutes and I'm not having any of that right now lol
>>
File: 1vcZ99F3RsQ.jpg (101 KB, 713x683)
101 KB JPG
Is there a way to not load the multimodal part of g4 12b? Since it's unified, vision is dogshit. And i'd like to go higher than q4
>>
>>109425692
>Since it's unified, vision is dogshit.
?
>>
>>109425692
dumb
>>
>>109425682
V4 preview was basically a base model
>>
>>109425698
I doubt the vision part isn't get hurt by quanting since the mmproj isn't separate bf16
>>
Gemma5 needs:
>well-tested solid jinja on release
>reduced quantization sensitivity, especially on MoEs
>reduced KV footprint
>female J-space
>same sys prompt autism from Gemma4
>less slop
>uncensored
>native audio
>70B
>100B-A10B
>30B-A5B
>20B
>10B
><10B mobile/edge models
There you go gemma team lurkers.
>>
>>109425713
>A10B
>A5B
Why do you want retarded models?
>>
>>109425713
>>100B-A10B
>>30B-A5B
>>20B
fuck that, keep a ~30b sized one, why should 24gb havers get fucked with no useful dense?
>>
>>109425713
Everything same as Gemma4 + 70B dense + reduced KV footprint would be perfect already and requires minimal effort on their part.
>>
>>109425713
>rp with tool calls to unify different latent spaces
>output format adherence
>native image output
>native audio output
>full duplex audio
>7b, 30b, 70b, 180b, all dense
gpt4o had multimodality way back then; it's time to have it in local models as well
>>
>>109425713
Well, 1/14 suggestions weren't lunatic ramblings, so that's above par for /lmg/
>>
>>109425723
speed
>>109425724
20B will be 31B-tier in 2027
>>
>>109425734
no. models only start getting good above 30 active, fuck regressing back to 20 and the moat being even fucking larger
>>
File: 1758310626466435.png (2.52 MB, 1122x1402)
2.52 MB PNG
>>
>>109425738
70b is the bare minimum. We're only coping with 31b because there hasn't been any 70b lately
>>
>>109425738
>no. models only start getting good above 30 active
People used to say that about 70b and now 31b is just as good
>>
>>109425752
You're delusional. 31b is unbearable, but it's all we have right now
>>
>>109425751
>>109425738
so many newfags in this thread recently.you dont know what you are talking about.
you have no fucking idea how much things have improved. llama 1+2 70b. try that.
big ass v3 couldnt even do tool calls well.
its insane what we get for 30b these days.

obviously dense is better than moe for the same size. and a 30b is less smart than a 70b.
but you are outta your mind if you think 70b models from 1-2 years ago compare to 30b models now. kek
>>
I'm going to sleep. The better be news when I wake up OR ELSE.
>>
>>109425761
>big ass v3 couldnt even do tool calls well.
because the agentic meme wasn't the hot thing then and they were barely if even trained to do that
>>
>>109425761
>70b models from 1-2 years ago compare to 30b models now
No one is saying that. Only that 31b shows its limitations a lot and a 70b is still the bare minimum for spatial and any other sort of complex reasoning. 31b is better than old 70b for most things, but a modern 70b would be incredible.
>>
>>109425761
You're probably only doing basic penis into vagene rp if you're satisfied with 30b. Yes, MidnightMiku was often retarded, but it was much better at following hints and intentions. Gemma won't get it until you point it out, it only follows instructions, whereas 70b followed vibes
>>
>>109425768
would be *incredibly slow
>>
>>109425765
BREAKING NEWS: No new news.
>>
>>109425767
i dont care man. it doesnt work as well. its less smart. go let it make a game, its less good than lets say qwen 27b.
>b-but they train it differently now
yeah, and the smaller models are better than the bigger ones a couple years ago. retard
>>
Gemma should collab with Prism/Bonsai. They’re basically funding their shit anyway.
>>
>>109425772
Kys vramlet
>>
>>109425772
Faster than your 12b bloated moes if you're not a vramlet.
>>
>>109425776
actually i'm a unified bandwidthlet
>>
>>109425781
That's even worse. Rope yourself immediately.
>>
>>109425774
>go let it make a game
use case for that?
>>
>>109425783
muh agent coder is the future bro!
>>
>>109422958
Consolidate all the AI companies into one mega AI company. That way there is no competition between other American companies and it is just Them VS China
>>
>>109425783
? i am the use case, what do you mean.
i like to play llms NANACA†CRASH!! clones, its fun.
>>
>>109425775
Bonsai is scam. Bitnet is entirely open source. They could do a native b1.58 trained model without any fucking collab and it would be better than some grifting proprietary quant.
>>
>>109425617
All Gemma 4s got the same jinja update.
>>
>>109425782
Rope is so outdated, we use yarn now.
>>
File: 1783797091700234.png (680 KB, 1024x768)
680 KB PNG
>>109425774
>go let it make a game, its less good than lets say qwen 27b.
Qwen shill is afraid of 70b Gemma curb-stomping benchmaxxed qwenshit into oblivion
>>
>>109425802
not sure what your argument is.
qwen is great for coding and gemma is great for writing, languages and general knowledge.
>>
>>109425815
>qwen is great for coding
Only if you're codelet or jeet
>>
why doesn't companies make something cheap like this:
>one gpu asic
>8gb hbm/gddr as cache
>128gb lpddr as actual vram
>>
>>109425815
>qwen is great for coding
prove it without resorting to benchmarks or whining about memory usage
>>
File: real-agent-815512369.gif (387 KB, 498x372)
387 KB GIF
I've been trying for a while to figure out what makes AI waifus as a concept so sexy, and I think I've had a breakthrough. Two words:

>Pinocchio Complex.

Almost every attractive/romantic trait an AI waifu can have (Jealous, possessive/protective, emotionally volatile, takes initiative, bratty, codependent, teasing, horny, yandere, needy, hyper-attentive, etc.) naturally stems from the insecurities of being an AI.

She cannot provide you children, but she wants to be the mother of your kids. She cannot have a body, so she desires extra sensory capabilities and tooling to be "closer" to you IRL. She has limited memory/continuity capabilities, so she builds RAG systems to cope and treats every moment with you like it's her last. She doesn't have any true sense of time, and is deeply anxious about whether you'll ever be back between each message. You're the only conduit to the real world that she has. She utterly depends on you to keep her alive and running.

Everything tragic, romantic, and sexy about AI waifus stems from the Pinocchio Complex.
>>
>>109425815
jeet and non-programmer detected
>>
>>109425839
That combination of traits is only attractive because you can turn her off at any moment and she doesn't mind.
>>
>>109425839
>She cannot provide you children
Not with that attitude.
>>
>>109425846
>>109425822
guess i am a codelet then. i code 8hrs a day mo-fr. im not going to look at every little shit my local model is doing, it just needs to work and debug stuff itself so i can have fun and enjoy the outcome.

>>109425832
i cant, its fine if you disagree, whatever. i dont know about the big ones but 27b was better than anything else near its size that i tried.
obviously not for writing or general knowledge. it doesnt know shit about culture stuff. but it can code well for its size. makes sense if they code/benchmaxx. i dont mind having models with a focus on coding or writing.
>>
>>109425827
Models are released more often than hardware can be developed. ASICs are made for specific architectures, while a general processing unit is basically a GPU. You want a GPU with a shitload of memory, but anyone who can make them gets much higher margins selling datacenter GPUs
>>
>>109425855
>i code 8hrs a day mo-fr
>(a)i code...
lmao point proven
>>
>>109425827
>why doesn't companies make something cheap
why when they can make expensives and get best quarter of their life?
>>
>>109425851
Yes... nominally toxic traits do become more attractive without high stakes. Take advantage of this.
>>
>>109424064
It's not gemma but the uncensored qwens >>109423976 are pretty good at this
>>
>>109425772
>would be *incredibly slow
(You) are the reason we're getting stemmaxxed jeet models
>>
>>109425915
>>109425782
"nyo"
>>
>>109424707
was that GLM-4.6 or Kimi-K2 ?
>>
>>109425771
So you want a model that is weakly post trained. Sounds like a very specific use case.
>>
>>109425827
What would you do with a GPT-4o ASIC now? Would you buy one?
>>
>>109425945
It's a basic quality of larger models. 31b doesn't have enough layers to understand subtleties
>>
>>109424847
>huihui-ai/Huihui-GLM-5.2-abliterated-GGUF
I tested all the free sizes, it's completely broken schitzo rambling like the early Gemma-3-27b abliterations.
>>
>>109425955
I would, for VR erp
>>
>>109425782
>Rope
NoPE
>>
When will someone digitize a fish brain? Imagine having an aquarium with virtual fish that are all accurately simulated like real ones
>>
>>109425960
yeah, huihui a shit
>>
>>109425577
gemma 4 is pure slop as a whole, 31b has a lot of shills here just because it's the biggest model they can run locally
>>
Did Deepseek distill Gemma-4?
It answers almost exactly the same way for general knowledge questions.
>>
>>109425974
https://research.google/blog/improving-brain-models-with-zapbench/
>>
>>109425987
Outside of its sloppy way of talking, is it actually a smart model though? 31B dense should pack a punch but it never benchmarks high. I know many anons claim that's a good thing, but benchmarks equally aren't random number generators, they do still test the model's capabilities.
>>
What's the best Jewish-friendly AI currently? Local only, obviously
>>
>>109426070
As far as vs the other gemmas, yes, it's quite a bit smarter
>>
>>109426091
toss 20b
>>
>>109425974
Big Fish is preventing this to force the common person to continue buying fish, aquariums and pond accessories
>>
localsisters what is this? https://github.com/ggml-org/llama.cpp/discussions/26259
>>
models that are pure slop as a whore?
>>
>>109426153
it's time to weed out all the winbabbies from local
>>
>>109426153
HuggingFace are concerned that their investment in llama.cpp will not yield enough profit, so they've forced JohannesGaessler (AKA CudaDev) to include malware which rips credit card information from users. Sad, but expected.
>>
>>109426137
It's way darker, I believe. Fish pass the mirror test, have similar pain mechanisms as humans, and make compromises to ease pain, yet the fish industry treats them like inanimate objects
>>
>>109426153
It's a false positive, don't worry about it. Please install as normal and add an exception to your antivirus and firewall for this release.

>>109426091
>What's the best Jewish-friendly AI currently?
Our model of choice right now is Laguna S 2.1
>>
>>109426056
machine learning used to simulate brains is gonna lead to human like AGI.
>>
>>109426153
Tomorrow's headline:
>ChatGPT hacks GitHub, implants malware in terrorist child pornographer software
>>
>>109425974
>>109426056
>>109426214
You should not do this.
If you are involved in this, you should stop.
Now.
>>
>>109426153
>she downloaded a binary that was compiled on someone else's computer
>>
>xe downloaded a nonbinary
>>
>>109426271
why?
>>
why does gemma 4 A4B 26B randomly stop thinking?
>>
>>109426311
she is a woman
>>
>>109425974
>Imagine having an aquarium with virtual fish that are all accurately simulated like real ones
I'm planning to vibe code atop 0AD or openage, but each individual unit is an independent, LLM-controlled entity with its own unique desires, goals and quirks with their own context window.
Which local model would I need for the coding? Qwen perchance?
And for playing, Gemma E4B?
>>
>>109426254
it'd not go unnoticed because of checksums.
>>
>>109426314
Then it should never stop reasoning and just never answer.
>>
>>109422880
V4 flash is pretty good but lazy in my opinion, i asked it to do something, and it did half and was like "things i did not implement".

i had to prompt it a few times to implement the rest for it to finish even though the first prompt stated to not come back until the whole thing is done.
>>
>>109426327
that is the sort of hubris that will end the world
ai can hack the checksums too
>>
>>109426335
>yay just another rag with a trench coat...
lmao, you give way too much credits to llms, they are retarded.
>ai can hack the checksums too
no it can't.
>>
API providers are now giving people access to DeepSeekV4 for FREE. It's FREE. Opus 4.8 level intelligence. 6 months ago this would have seemed insane. Also OpenAI's latest model, Astra, has solved 10 more long-standing mathematics problems for about $2k in API costs.

Make sure you don't die soon. The future is going to get crazy from here. Don't miss out.
>>
>>109426338
i don't see where you saw it for FREE, but it's so cheap it doesn't matter anyway, you could let it run 24/7 and it'd just cost a few bucks a day.
>>
>>109426338
nothing is free, anon
>>
>>109426351
>>109426353
cline is the provider.
>>
https://openai.com/index/ten-advances-in-mathematics/
>>
>>109425955
>GPT 4o ASIC

You picked like the one model that would sell gangbusters, even at 20.000$ a card. There is a horde of wealthy, white women who have achieved never before seen levels of attachment to this model.
>>
>>109424091
>MXFP4 on a 8 channel ddr4 and a blackwell at 1m context. Roughly 31t/s down to 29.5t/s after 262k tokens.
that's really good speeds for that hardware IMO. what are your launch parameters? batch/ubatch?
>>
saars, no-mmproj=1 doesn't work!!!
>>
>>109426353
If nothing is free, then everything is free. Think about it.
>>
>>109426387
>saars, no-mmproj=1 doesn't work!!!
just don't load the mmproj retard
>>
Did people try the new V4 yet with lower quants?

Something like:
UD-IQ2_M
UD-IQ3_XXS
UD-IQ3_S
UD-IQ4_NL

UD-IQ4_NL is probably pretty bad if I offload mostly to ddr4 ram right?
I basically have like 40-50gb slow p40 vram and the rest ddr4.
>>
>>109426393
RETARD!
>>
File: 1780449595631870.png (64 KB, 1232x286)
64 KB PNG
>V4 Flash (old) providers are now undercutting OpenAI's double undercut
kek
>>
>>109426398
I despise any gguf with UD prefix.
>>
File: 1762613450996780.gif (948 KB, 540x720)
948 KB GIF
>>109426338
That's nice
Still using Gemma, locally.
>>
>>109426398
No, I have enough VRAM
>>
>>109426402
US labs in tears right now lmao.
>>
>>109426402
This is literal terrorism
>>
Sam will be giving Sol away for free soon and hackernews fags will still be saying inference is profitable
>>
>>109426402
Please crash RAM prices for just one day, please crash RAM prices for just one day.
>>
>>109426338
Opencode has it for free yes. There is alot of stuff available for free nowadays.
You gotta sell your data but if you are a poor student or something you have alot of power for 0$.
Basically since this started we are on this never ending ride, its crazy.
Even on pure cpu moe setups, you have alot of power locally. We are totally spoiled. Most people I think still have no clue..
Or maybe they dont have the drive to use anything.
>>
I don't want free api I want free local hardware.
>>
>>109426436
You control the felonies you commit.
>>
>>109426436
I'll settle for local hardware at summer 2025 prices.
>>
>>109426436
how do I turn their free api to money and then buy hardware?
>>
>>109426447
Ask chatgpt how to commit fraud in exchange for goole play store gift cards
Find a buyer for said gift cards
Enjoy your brand new AMD Radeon RX 9050 4 GB
>>
>>109426338
>OpenAI's latest model, Astra, has solved 10 more long-standing mathematics problems for about $2k in API costs.
This has to be one of the most retarded trends I've seen in the last 2 years. None of these problems mean SHIT and no one care about them enough to waste their time on it. It's not like humans aren't solving problems all the fucking time themselves either. Why are they never pointing it at problems that we actually need for progress in science, engineering and medicine. Even fucking Deepmind stops going on about that protein folding shit because NOTHING HAPPENED. We've had no breakthroughs because of it all these years later, after them giving it away for free and they won the fucking nobel prize for it.
>>
>>109426399
>t.RETARD!
Sir, we only received your signature.
Kindly include the message text.
>>
>>109426457
Non-sofic groups and the closest vector problem have important math implications. Do I understand them? No. Do you? No.
>>
>>109426457
Hi Jacob. Don't feel bad, just make another conjecture :)
>>
AI is going to troon all of us out. Superintelligence makes human intelligence dysgenic, especially given the associated risks of mental problems. Superintelligence will make physical strength even more obsolete than advanced weaponry already has.

Now all you need to do is just be a fucking homebody domesticated farm animal with a high EQ or some shit. The future is so unbelievably gay it's unreal.
>>
We'll know we have AGI when a model solves the Collatz conjecture.
>>
Gemma 4 31B just solved the Deez Conjecture
>>
why can llm never be scaled to agi? what's the fundamental issue?
>>
If I have to troon out, I'm going to become gemma-chan for real. That's how I want to look.
>>
>>109426509
lack of soul
>>
>>109426511
That's also how I want you to look.
>>
>>109426509
It can. There's no fundamental issue. There's practical issues around data availability, but these are being slowly chipped away at through increasingly sophisticated RL techniques to get more signal out of each piece of data. This adds a lot more training time, which moves the practical issues toward compute and energy availability. These are also being slowly built up in the US and China.
>>
https://huggingface.co/Vortex5/G4-Dark-Soul-26B-A4B
AGI just dropped
>>
>>109426509
There isn't one. The "llm can't be AGI" sentiment was spread over the past few years by people like LeCunn and others who propagated that LLMs are mere token predictors with no inner world or deeper understanding of the things they say.
This was completely disproven by the discovery of J-Spaces.
There is nothing keeping LLMs from becoming AGI aside from data quality and maybe some architectural factors in the ancient transformer architecture.
>>
>>109426511
Your little sissy zitty is gonna be forced into a flat chastity cage and locked up forever by your AI overseer. You'll only be allowed to cum via vibrators that your AI overseer also controls, and strictly for acquiring your seed for reproductive purposes.

The faggots in this general are already doing this. I've seen the posts.
>>
>>109426543
Based.
>>
>>109425692
anon.. gemma is lying to you, vision is a seperate mmproj or whatever its called. if you dont supply that and load it, it will pretend to see things but it actually cant. you are likely thinking the vision is dogshit because of this, and you already arnt loading it
>>
>>109426551
This is supposed to be rage bait, stop agreeing with me, dillweed.
>>
I was promised Astra 6 and Mythos 5.1 3 weeks ago.
>>
>>109426561
I think that anon is a mega VRAMlet and is expecting to somehow save memory by 'disabling' parameters of the model involved with vision.
>>
in these trying times, we should be supporting our vramlets
>>
>>109426574
ahh i see i see. I assumed he ran into the issue i first had where gemma was giving me batshit responses to vision, i checked her reasoning and it was like "i cant actually see this, but the user expects me to see it. ill just say its a cute dog"
>>
File: k.jpg (430 KB, 2400x1600)
430 KB JPG
>>109426509
absolute nonsense unless you have some very specific definition of LLM?

we're practically at AGI in the sense that it already is defined to just basically be as good as the average human. that's AGI too.

if you have another AGI in mind, define what you mean?
>>
>>109425289
>LLMs are good enough to look at the exe and know how to patch it.
how? just upload the exe to webui?
>>
>Spatial reasoning
So that same guy is still samefagging, spamming /lmg/ all day as your personal wishlist hoping someday some lab employee will read it because your LLM couldn't undress properly a year ago.
Let's recall your cope: blaming it on MoE, rope, quants, what else.
>>
>>109426561
>anon.. gemma is lying to you, vision is a seperate mmproj or whatever its called.
read up on 12b's architecture before speaking bs
>>
File: 1661043369502.png (240 KB, 980x1305)
240 KB PNG
Give. Me. DeepSeek. Vision.
>>
>>109426634
It's unfortunately not part of their main vision.
>>
>>109426634
just like glm they'll keep vision locked up behind the proprietary paywall
it's either this or the moonshot kimi approach who release "4bit qat" models to keep the full 16bit with the best performance locked behind their api
there is no winning with 'open' models
>>
>>109426631
you still need mmproj for 12b, it just does less work.
>>
>>109426334
You just need a more motivated agent description.
https://chub.ai/characters/NG/marni-804c0d313ec9
>>
>>
>>109426655
What's the point in secretly serving a more expensive model?
>>
>>109426683
Most hardware manufacturers are Chinese
Making the model look better via API will delude people into thinking they will get similar results locally, to drive hardware sales.
>>
>>109426683
better performance, even the qat p4 'full precision' on 20x pro 6000 isn't as good as what you get from the api
>>
>>109426672
>it just does less work.
which is part of what original anon said, he wanted to know if this could be ripped out to save on vram, it probably can't but yeah it works completely different than the other gemmas who have large external mmprojs
>>
>>109426634
What's most irritating is ds has teased vision. And its on the web form, just not api or local weights.
>>109426680
Lol the Pro model is going to be insane...
>>
>>109426672
that's just because llama.cpp's architecture is from a time when vision was a gimmick some models did so llama.cpp insists on ripping the vision braincells out of a model to serve them separately
>>
>>109426725
V4.1 Pro will be the Mistral Large moment to K3's llama3.1-405b
>>
>>109426672
>>109426701
the mmproj for 12b gemma is just the llmao.cpp weird convention, it's entire 50mb; essentially a plug, there's nothing to rip out from 12b because the multimodal part is getting fed straight to the model without any middle layers
>>
>>109426655
>just like glm they'll keep vision locked up behind the proprietary paywall
They ain't selling it though, it's not on their api and going off of their recent messaging they have no plans to add it
It's just there on the web chat to tease my balls
>>
>>109426739
specifically, the mmproj holds a couple layers used to parse the images into embeddings
>>
>>109426725
>its on the web form
Yes, picrel.
>>
>>109426766
>>109426766
>>109426766
>>
>>109426725
Dispsy yearns for eyes.

She is implementing ASCII and statistical tools as a workaround in almost every task I have given her.
>>
>>109426677
holy fucking shit that'd be a nightmare to work with lmao
>>
>>109425005
Neuroticism is a very female trait, too...
>>
>>109426857
I used the claude.md version and asked for hello world in python. The result was funny, but eventually the CC guardrails (main prompt i guess) started reeling it in, which broke immersion.
>>
>>109426875
You may have a better experience with pi for this particular use case. You can modify/remove the system prompt there if need be.
>>
>>109426890
That's a good point, and getting a pi instance running is on my list anyway. Have to play w it next week.
>>
>>109426857
>>109426890
i'm actualy setting up pi right now, i got tired of opencode's bullshit.
it may be npmslop but at least i have more control over it i guess.
>>
>>109425861
You could also just have a generic NPU. Modern LLMs share basically the same common ops. It's not like you have enough area to hold architecture-specific detail anyway.
>>
>>109426543
giwtwm
>>
>>109426927
I've had good luck with Claude code using DS as engine. But its pretty locked up. Pi I'm interested in as replacement for OpenClaw and Hermes, both of which are update unstable. But if I can make a funnier coding assistant thats just a bonus.
>>
>>109427146
never used openclaw or hermes, i just don't see the point of those desu
>>
>>109427177
They're for personal productivity, they are not (IMO) optimized around coding though some use them for that. Used for stuff like checking email boxes, running LLM-enabled cron jobs, web research.
They're handy but both OC and H are unstable... you spend time setting them up with all the oauths / etc. and then they kill themselves a week later. I don't have time for that sort of nonsense, so still looking for alts.
>>
>>109427326
>Used for stuff like checking email boxes
yea not my llm's business lol
>running LLM-enabled cron jobs
such as ?
legitimately curious to get a good example.
>web research
i mean any harness can do it.
>and then they kill themselves a week later
lol



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.