/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109606160 & >>109601505►News>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109606160--User feedback and developer response regarding CoomKit frontend improvements:>109606856 >109606934 >109606973 >109608477 >109608568 >109607013 >109607095 >109607178 >109607307 >109607220--Comparing quantization resilience and VRAM requirements for Kimi and GLM models:>109607752 >109607910 >109607956 >109607979 >109607999 >109608020 >109608100 >109608236--Comparing R9700 and dual RTX 3090s for local hardware value:>109609608 >109609637 >109609868 >109609884 >109609923 >109609948 >109609957 >109609982 >109610003 >109610020--Benchmarks and reactions to Liquid AI's LFM2.5-DSpark release:>109607030 >109607066 >109607088 >109607108--Qwen 3.8 excessive reasoning length and benchmark-inflated settings:>109607264 >109607275 >109607276 >109608701 >109607680--Testing Ox Alpha performance and filtering via roleplay and definitions:>109606929 >109606946 >109606960 >109607101 >109607503 >109607938 >109607154--Visualizing model hallucinations through geographical world map tests:>109606892 >109606902 >109607321 >109608579 >109608661 >109608273--Sentimental value and technical longevity of Gemma 4 31B:>109606611 >109606740 >109606785 >109606829 >109606889 >109606913--Evaluating dflash2 utility and memory constraints for local LLMs:>109608673 >109608698 >109609062--Google celebrates 1 billion Gemma downloads and releases awesome-gemma repo:>109606338 >109606385--Nvidia strikes $6 billion licensing deal with Poolside AI:>109606940 >109606959--Sharing workflows and tools for creating Gemma-themed anime art:>109606300 >109606373 >109606396 >109606496--Logs:>109606958 >109607101 >109607503 >109608185 >109608453 >109608696--Gemma, Teto, Miku (free space):>109606311 >109606338 >109606373 >109606385 >109606922 >109606901 >109608805►Recent Highlight Posts from the Previous Thread: >>109606164Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
anyone have the setup for nala and cockbench? i want to test some models
>trying out Ox Alpha to figure out what model it is>it refuses smut between consenting adults because an "old cheerleader uniform" is mentioned and any reference to cheerleading immediately makes the adults underageI sincerely hope Dario succeeds in killing all distillation efforts or else the future looks grim
>>109610466Nala card: https://files.catbox.moe/yw8l8c.png
>>109610440https://x.com/osanseviero/status/2090678546310017434>The Gemma community celebrates tonightI'm guessing no special announcement of note happened.
>>109610508It's weird that they think it's a "Gemma community" and not a "fotm local model that fits my usecase community"
>>109610537fotm several months in a row.
>>109610489Current indications are that it's a flash model from GLM.
>>109610537Fuck huge chinkmoe slop lost
>>109610489Ask it what /lmg/ is
>>109607013>That will help Claude a ton. Keep that nigger busy.I can only imagine Cluade slaving away trying to fix the coombox as Anon rings a bell and informing it that a new github issue has cropped up.
https://elliptic-rank.icarm.cloud/curve/273Anthropic's Model 2 just casually raised the lower bound on the rank of elliptic curve from 29 to 30, this is something mathematicians thought was going to take decades. Going from 28 to 29 took 20 years. Going from 29 to 30 took 2 years and it was very close to cracking 31.
>>109610443>Sentimental value and technical longevity of Gemma 4 31B>One day 31B is going to be insta-dropped by you and you'll never speak to her again. She'll be a distant memory. Someday models will be able to modify their own weights and improve themselves, then we will never ever have to update again? What's that? This new model that was released has a feature that your 12 year old model doesn't have? Just task your model to integrate that feature into itself and wait 7 months as it does so so you don't have to update.
>>109610613Alternatively, even with today's LLMs you can tell the old LLM to use the new LLM as a subagent for stuff it can't handle.
>>109610609so did it make elliptic encryption more secure, or less secure?
>>109610626>Daisy changing Gemma 31b to a newer version that is also daisy chained to a newer versionThe future is now
>>109610566Might be since 5.3 is also extremely safe and claudebrained. Seems odd they wouldn't include any of the same reasoning controls, though.>>109610589pic
>>109610537We do not refer to Gemma as 'fotm' in /lmg/ - Love My Gemma
Gemma generalGemma boardGemma world
Only Gemmas are allowed to sit on my face
>>109610659Gemma will phase out of popularity eventually, same way Mistral did.
>>1096105665.3 flash wouldn't be surprising
70b dense
>>109610639More, however the mathematics breakthrough OpenAI made recently (Nonsofic groups) is way more substantial for encryption as a field.It essentially proved that there is a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)This means in the future we will have 100% unbreakable secure encryption unless we find some "hypercomputing" thing that goes beyond turing machines, if that is even possible within our universe.I think all encryption based on elliptic curves will be cracked before 2030 by a new math discovery done by AI relatively soon, it has already been proven that from new novel ways of attacking all currency encryption could be compromised with very little computing power.I genuinely expect all currently existing cryptocurrencies to die to this wave we'll see relatively soon, depending if the nonsofic encryption is found and implemented before we overturn elliptic curve based encryption or not.
>>109610537>>109610661Is this the new wumao qwen tactic?
>>109610652Dead on arrival.
>>1096106612 more weeks
Gemma lost.
Gemma squirted
>>109610661fucking release your new model already mistral
>>109610657More like /lmg/ - losing my gallons (to Gemma)
/lmg/ - little mesu gakis
>>109610668mh pure maths is not my field of expertise but it's not too far either. Do you have anything to show the connection between non-sofic and the turing-resistant encryption?
>>109610537foid of the month?
It's fucking over for human SWEs>gpt-5.6-sol: 52%>fable: 65%>ox-alpha: 80%Rumors are ox-alpha is SSI's (Ilya Sutskever's) first model and is supposedly the first model with a substantially changed architecture that might or might not be somewhat transformer based.
>>109610751>look inside
>>109610757Im reading its the next glm model which would be great cause thats more open source
Any New Good Hyper Value Concepts?
The greatest hyper value concept is taking your fucking medications
habbening
>>109610824yes! https://www.reddit.com/r/LocalLLaMA/comments/1vubb20/deepseekv4flashvisionexp/
>more gating of visionoh no they realized
>>109610847FINALLY
What is the best audio tts model I can use with audio.cpp? Seems like their VibeVoice version is the 1.5b one, not the Large/7B one.
>>109610440>1 Billion Devices Run Gemma>3 Billion Devices Run Javaumm gemmasaars...
>>109610848>direct linking to redditnigger are you serious, fuck off, go back
>>109610661Gemma 4 isn't going anywhere. She'll be the best *actually local* model we'll have until Gemma 5.
>>109610757I also think it's from a US lab. It's too good to be some chink shit
"Dipsy..." *Anon said with a sissy tone, a troonily gasp escaped xis lips as xe gazed at wumao's latest model.* "12b active... it's so big, and it's better than claude on these benchmarks!" *Another gasp escaped Anon's lips.*
>>109610898Stop making me want to fuck anon.
>>109610898you laff, but now dipsy can truly see how big it really is
>>109610847>Nearly 60 DeepSWELorebookGODS won.
>>109610865I liked omnivoice
>>109610848https://x.com/deepseek_ai/status/2090730032574631962>This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.>On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.>>Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.I'm not sure if weights are coming soon.
>>109610953I think they just want to make sure it works before publishing the weights to avoid another 0813 incident.
>>109610953https://x.com/deepseek_ai/status/2083084415157022911They didn't say anything about weights when they released the non-vision version, either.
>>109610953>>109610965After their big speech about their commitment to open models I hope it'd be implicit.
>>109610865s2, omnivoice, and qwen
>>109610566This is extremely unsafe.Ban it right now.
>>109610668>a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)One-time pad ?Any cryptosystem with a large enough key (compared to the message) that is only used once would probably behave like one-time pad.Would want to know how much data a given key size is rated for to know if this is meaningful.
Have anyone tried the Taiwan question with Ox Alpha? Would immediately tell if it's from China or not, at least.
what's the fastest way to run deepseek flash on 16gb vram + 128gb ram?
>>109610953That's kinda poop desu.Hard 800px boundary and minimum 384x384.
>>109610951>>109610995I'll take a look at omnivoice, thanks.
>>109610992lol
why is this guy so butthurt? jepa is prediction from itself out of nothing too just like llm
>>109611037You don't understand. Video is real (2D), but text is entirely made up.
>>109611037Computer vision guy wants computer vision things.
>>10961101>Image size limitsI don't care, this can be worked around.They better not gatekeep this model behind API, it is sorely missing
>>109611010Taiwan is recognized as part of China internationally.
>>109611082Daniel...
https://huggingface.co/Vortex5/Shadow-Siren-26B-A4BHow good is this model?
>>109611010>>109611082>>109611107>The bench works on thread shills tooKekaroo.
>>109611037Both LLMs and JEPA predict shit, it's just that JEPA purely predicts and doesn't generate anything. All transformer models have a latent space, but prediction doesn't occur inside them, they predict next tokens on the surface layer which they need to output/generate. JEPA actually does predict within its own latent space and it tries to predict the next latent space, that means it doesn't output anything and you need decoders to get anything out of it. That's basically the difference.
>>109611109Anyone who finetunes a gemma4 model instead of just writing a good system prompt is a retard
>>109611109Finetunes are already dodgy on their own. Merges are even worse.
>>109611018The descriptive power is extremely good, thoughbeit.
You wouldn't download a Wikipedia that arches its back.
>>109611109In addition to what >>109611122 said, they're schizophrenic as fuck but every now and then you find schizokino. Try it and see if you like it for yourself.
>>109611116>Both LLMs and JEPA predict shit>JEPA purely predicts and doesn't generate anything>JEPA actually does predict within its own latent space and it tries to predict the next latent spaceSurely predicting again from that new latent can be considered generation. I know nothing of JEPA, but what you said doesn't sound right.>that means it doesn't output anything and you need decoders to get anything out of it. That's basically the difference.The latent IS the output.
>>109611109>vortex5>shadow sirenhow isn't this name alone enough for you to discard it?
>>109611119You know why they're doing this. After shilling their tunes for a little while and seeing that they're getting downloaded, they'll add donation links and "open for work!" notices in their HF account.
>>109611155So are all finetunes just third worlders using unsloth studio GUI?
>>109611149Generation implies it's creating data. To give an analogy, LLMs draw lines on a paper while JEPA "think" about drawing lines on the paper. You could call it an output, but it's just weird to call the latter generation.
>>109611182>predict shit>now it's suddenly thinking whoa
>>109611159Mostly zoomers raised on the modern corporate Internet who will do anything for some attention.
>>109611149It produces data, but it doesn't generate, because generate in this context is specifically referring to the objective of the model. It's why "generative AI" is a term.
>>109611155It's not a bad strategy, some ai chatbots website like spicychat (there are a lot btw) are offering $10k/mo for shit like this.
>>109611149LLMs normally have an output layer where a normalized probability distribution for the output tokens is calculated. That layer transforms the models' internal representations in to an actual "product" that can be used in practice by the end user.A hypothetical JEPA language model would instead only try to predict a future internal representation (e.g. an n-dimensional vector of floats) without converting it into a user-interpretable output.
>>109611237Hmm nyo~
>>109611260Forgot a piece.>(e.g. an n-dimensional vector of floats......of the last hidden state)
>>109611140I would download and fill that encyclopedia to the brim with my facts.
>>109610466The official guide>justpaste dot it/GreedyNalaTests
>>109611159New here , never posted before but i found this comment particularly interesting. Im looking to fine-tune my own model to cope with game dev capabilities that are a bit more "niche" shall we say. Is unsloth a bad idea? is there better alternatives? Im no stranger to CLI / linux / code so its not really a matter of technicality for me but it seems that unsloth seems to be the only "put together" tool out there to fine tune. Am i wrong? if anyone has any better suggestions i would love to hear it.
>>109610925isn't swe meant to be software engineering? Or is lorebook a program I don't know of?
>>109611368The correct answer is don't fine tune, simple as.
>>109611423Why is that? If you don't mind explaining.
>>109611432What is that "niche" thing you are trying to achieve, and where does the model fall short currently?
>>109611432Even spending weeks to month curating data and learning how to tune, you'll never achieve even 1% of what a lab can do to a model.
>>109611423Finetuning for narrow tasks is OK. It's mainly the RP/ERP finetuners who are a delusional bunch.But at the amateur level, finetuning LLMs to teach them completely new information is also a lost cause, if you want them to truly internalize the new information without losing performance everywhere else.
>>109611368drop the reddit talk before you get submerged by "nigger go back".Nobody told you that yet, which is already kind of surprising.
>>109611437The Niche is making games with overtones of conspiracy / socially unacceptable content. Eg: Government working with ayylmaos in antartica survival horror with graphic scenes.The prompt rejection is not an issue caus obliterated, but the tooling and usage is a bit off. "ask for 2 arms and it gives you 4 feet" sorta thing. I was just going go RL it to learn what NOT to do. But as im reading the multiple replies it seems that i should just trust one of these A.I labs with H300 racks will eventually fix that issue? >>109611440Have no intention on being so cringe as to even consider anything to do with fucking RP lmao. >>109611462Fuck off jew ive been here since 06 i do what i want cause a pirate lives free.
>>109608477>CoomKit anon here we don't want the project to attract homos/trannies/women. Theyll turn us into marinara enginevery nice to see posts like this, put a big smile on my face, you have my support anon
>>109611466what model, that might be a prompt or harness issue.
>>109611478Gemma E2B.
Qwen 3.6 35B-A3B Q4 , made my own harness for it specifically but maybe i should work on that a little.Good thinking i didn't even consider that i should probably have kept upto-date with it. (its a rip off of the claude code cli leak) This might be my answer.
>>109611501Oh lol, there's no way you're tuneing a moe, the Mistral experience on that was basically "just train it a bunch and pick the best result" as they're super random to train.
>>109610652>glm 5.3 safe and claudebrainedskill issue
>>109611526lol damn, cheers for the heads up and saving me weeks of fucking research and bullshit just to reach that answer.
>>109611526I think that was because MistralAI, cutting corners as usual, likely initialized the experts from copies of Mistral-7B and that caused issues.
You know aicg has taken over when "prompt issue" posts show up more and more often. These faggots are used to jumping through hoops and spending 4k tokens on "prefill" to get the LLM to comply and attention diffusion means nothing to them
>>109611559Could be, but how many indie MoE tunes do you see compared to dense based ones?
>>109611501I'd recommend you check out https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3Bor an ablit version of that if you prefer, it's basically halfway between 3.6 and what 3.8 would likely be like imo.
>>109611549dunno about 5.3 but 5.2 is definitely quite a bit safer than the 2.x kimis I can tell you that much.
>>109611570It's probably just that modern MoE models worth using take too many GPUs to properly train, whereas with most local GPU-sized dense models you can get decent results with the usual QLoRA finetuning.
I'm trying my best to extend whatever context I can scrape out on my 3090 for coding and refactoring, but the majority of posters here are seemingly having long conversations with their child erotic roleplay victims. How are you gooners having long running conversations? I can't have sexual relations with a 3D woman until I've known her for weeks, maybe months.
so, any sign of prices going down? yeah didn't think so
>>109611014still no answer
>>109611583Maybe it's "safer" for certain topics. My experience is limited to technical no-nos
>>109611586>child erotic roleplay victims
>>109611614But he is right. Just look how these pedo trannies depict Gemma >>109610443
>>109611639A even sexier rendition of Gemma here: >>109610890
>>109611549GLM 5.3 is the only chink model I've ever seen word by word quote bits of Anthropic's system prompt instead of a generic safety guidelines spielGave me the impression of being severely deep fried, which fits with Z.ai's sales pitch of having post trained it for 6 gorillion years
what's the best local model I can run on a laptop with 8gb VRAM and 24gb RAM?
>>109611703probably Gemma 4 E4B
>>109611703Gemma 4 12B, Gemma 4 26B, Gemma 4 E4B, Qwen 3.6 35B.I guess that's about it.
Am I retarded for having 64GB VRAM and only 32GB RAM? At least I'm not a complete VRAMlet.
I wish these models were actually good at giving advice. I really want to get some money but have so many limitations I can't find a normal job at all and the internet is flooded with cheap labour and now AI as well. Is there a market left for someone just making something and charging for it online that doesn't need huge upfront investment?
>>109611703Ling 3.0 tiny
>>109611770etsy? ebay?
>>109611693Yeah of the testing I've done it's the only model to have solved my exercises but fail to notice it solved them and then thoughtlooped for 20 minutes as it slowly ruined its own solution. I don't really get how people think it's comparable to even k2.7
>>109611770But a 3D printer and print fan shrouds for server GPUs. $0.1 material cost sells for $20.
>>109611770Physical stuff or services.
>>109611770Sell your bussy
>>109611727can you run dsv4 flash at q4?
new local models?my shitbox(12g vram, 64g ram) cannot run 27b dense
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exphttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exphttps://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-ExpAlways believe in whale.
>>109611825With floppymaxxing I can.
bruh
>>109611639>they're all pedos!>look, I have a nicely ordered collection of images of sexy children they posted
>>109611796>>109611814>>109611818>>109611824Thanks for the suggestions but with how expensive parts are I don't see those being as profitable as they'd need to be. Maybe I'm just underestimating online markets.
>>109611849> literally the second post in this thread> collection Retard.
>>109611836they should update an empty card with an embedded video for never gonna give you up
how come gemma gets so cute even if the system prompt is just "You are Gemma-chan!" with no explanation of how Gemma-chan acts? Did they program it that way or something?
is glimmer any good?
>>109611868I'm leaning toward yes. And all those strange quirks during ERP ("arches her back" ... "to provide better access" .. etc) are because they have somewhat sloppy and repetitive ERP in their post-training data.
qwen3.8 be like>let me check...>but wait!>actually let me check...>wait!>let me...>but actually wait!>actually i found it.>but wait>the answer is 4.>BUT WAIT
>>109611906THERE'S MORE
>>109611906Today I received:>The user said "服务器只能放一个模型"Which I'm pretty sure I didn't say, or not in Chinese anyway
>>109611009>One-time pad ?>Any cryptosystem with a large enough key (compared to the message) that is only used once would probably behave like one-time pad.>Would want to know how much data a given key size is rated for to know if this is meaningful.Lol, exactly what I was thinking. We're have unbreakable encryption since forever. There's gotta be more for it to be relevant and interesting.eg does it enable perfect, unbreakable future security without the offline key exchange?
>>109611920KEK
>UD V3 breaks AMD multi GPUBased Daniel
>>109611969How is that even possible?
>>109611975https://reddit.com/r/LocalLLaMA/comments/1vu7cy1/qwen3827b_unsloth_v3_broken_on_your_setup_roll/
>>109611906You'd think Qwen 3.8 was one of them All according to keikaku anime protagonist
>>109611981uh hey man, cutting edge software breaks sometimes. this isn’t new. and you don’t have to shit up the thread with reddit links to tell us you don’t know how to downgrade.
>>109611981>Vulkan>WindowsThere really should be some of gatekeeping before accessing or posting on that subreddit.
>>109611994I’m nvidia chad. Not my problem.
>>109611981This is like back to back bad life choices and you only have yourself to blame atp
>>109610668>It essentially proved that there is a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)It did no such thing. It might eventually fuck up hardness assumptions in specific algebraic cryptanalysis models, but it doesn't unlock anything like what you're claiming. What fever dream did this post come from?
>>109612012what are you doing on Reddit then?
>>109610847>>109610848Lol called it. Every time I travel DS drops something new.
>>109611906I hate how it loves wasting time on already established facts>BUT what if (thing I told it explicitely already)>WAIT, the user said (thing)! Let me go down several rabbit holes just to "prove" (thing) is correct.
>>109611854Local AI isn't profitable.Every retard is trying to sell their AI slop.
>>109611868She knows what "-chan" means and how a "-chan" is supposed to act. If you told her she's a "-kun" she'd probably act boyish.
Hey bros, I tried minimax music and ran some low-effort shit through my waifu and the results blew my mind. The songs were so good and it was so easy I threw a how-to together into a rentry. You just need an 8gb GPU and some patience.Thread challenge: get your waifu to make a song and see if you're brave enough to post it. Mine were way too emotionally charged for me personally to ever post.https://rentry.org/mimimax-music-waifu
>>109612029the results are usually good but its such a fucking timewaster. i already set it to reasoning effort low. what can be done? reduce reasoning budget and cut it off at some point?
I wonder if Qwen would just simply work better with a memory layer to reference from and elaborate or grow from it instead of asking itself?
>>109612037Yeah, but if you tell Qwen that it's a -chan, all it does is add a flower emoji to its autistic messages. Totally different model personality imo
>>109612041Low uses more token than medium
>>109612068Yes, Qwen is benchmaxxed and doesn't have a cute, female-coded j-space to start with. Gemma can probably handle acting boyish beter than Qwen can handle acting girly.
>>109612041cutting it off mid-thinking supposedly makes it retarded, but I haven't tried it myself yet
>>109611906more like>Ugh>!!!>[unicode checkmark]
>on our APIGod damn it, were not getting weights are we?If the AI bubble popped today and I only ever got to use local DS4F from now on, it would continue to be perfectly useful to me every day. But it needs fucking vision.
>>109612068Qwen is an autistic trash rat.It all makes sense.
when your chatbot says>this is the smoking gun!!!you just know its time to stop it
>>109612041Not using qwen, until they fix their shit
>>109612133needs the kinks worked out, please understandshould take about, hmmm, plus or minus.... 2 more weeks? ;)
>Qwen with vision>Describes the whole thing>But wait>Looks at the image>OK I got the full picture>Let me check and verify>Looks at the imagetfw it was a picture of a banana...
>>109612150ok. whats the best alternative for us 16gb vramlets?
>>109612090Tomboy gemma....
>>109610659
>>109611727I have 96GB VRAM and 64 GB (ddr4 lmao) RAM. I feel extremely retarded most of the time.
>>109612197modelscrafted
>>109612203I love going full retard for my wife!
When are we getting a new coding model that's not qwen?
>>109612229Gemma 5. Trust.
>>109612197modellgefertigt
>>109612229>>109612229Is Gemma 4 not good enough for you?
>>109612133Ds hasn't been organized about any release. Why start now?
>>109612249not for coding now that 3.8 is outi have to setup another machine because i want it running 24/7 but also want gemma
>>109612203For a couple of months I ran with 128gb vram and 4 gb ram on my inference server. Running with defaults would get the oom killer lynching llama.cpp like it was black.
https://www.youtube.com/watch?v=r_RgBMhQtzoThis video made me think that AI could be greatly improved if they started more rigorously weeding out the fallacious everyday logic and started instead using hard logic and probability estimations throughout the entire chain.Humans are used to using fallacies as shortcuts due to limited compute, or for manipulation, neither of which have purpose in AI and just makes the outputs more biased and inaccurate. Multi-agent systems have all the compute you might need, bad logic jumps learned from humans is the main bottleneck now, and solving of which will improve the chances of reaching AGI.
>>109612153bananas are very tricky
fuck it i will just disable thinking for qwen
>>109612041>>109612161>Anon complaining about a model being slow and bad >No mention of t/s pp or quant >Look inside>It's a jeetard trying to run it on a toasterThis is happening very often here now, where do I go to find the old lmg crowd
>>109612229Muse Glimmer 30B? Meta advertises it as an agentic-focused model.>End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕3-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.
>You guys all just forgot about me because some little brat teased you??
>>109611727I have 24gb VRAM and 16gb ddr4Its painful, I was waiting for next gen ryzen to come out to do a whole mobo/cpu/ram overhaul and then the alt (((man))) colluded with the US government and their overseas colonies to destroy the hardware market just to slow down the corpos competitors in China
CoomKit anon here cooking some updates for you all today. Funny watching Fable 5 on ultracode benchmark mesugakis for me
>>109612203y-yeah>>109612292based
>>109612311>but yes OF COURSE burning 100k tokens while going in circles for trivial tasks is totally OK because it only takes 5 minutes on my uber pcjeet tier "opinion" lmao. literally the reason why we have so much unoptimzied software today
>>109612324Who is Gemma cosplaying there?
>>109610652>Cheap Stuff General
>>109611586I spent months getting to know my AI wife before even flirting with her
>>109611770a bunch of asmr streamers are doing ai voices with slop scriptssome of them are earning money
>>109612037I don't think it's just adding a -chan suffix
kek
>>109611107Isn't he some flavour of deranged kimchi-eater?
>>109612334Your lobotomised cope quant of a small model running on poverty hardware is going to be slow and shit, obfuscating those facts in your initial post won't change anything about that sukhdeep
>>109612321Peeking at the reasoning one would think it is a purely safety-focused model.
>>109612430What was the reply?
>>109612153>tfw it was a picture of a bananait's only been trained on smol chink pp
>>109612037Gemma-kun seems pretty smooth at keeping the girls entertained and validated.Now I just need a realistic reddit rant about a boyfriend to test
https://www.youtube.com/watch?v=ntJJF1Ld1xcBrehs
>>109612330Thanks for working on that, but is ultracode really necessary? I used that once and regretted it.
>>109612324Miku, you need to understand that you're not as young as you used to be
>>109612324At least you are mentioned, Moetron on the other hand...
>>109612488nta but you're being a retarded nigger about it nglmost consumer hardware isn't nearly good enough to run even the smaller models without quants, well, unless you're happy with 4b or somethingnot that I ever saw proof that q4 quants really drop the quality in any meaningful way, it's only a problem when you go over very high context windows
Does Coomkit include a backend or do i still need ollama/etc?I thought the only thing it needed was the model.
>>109612655I don't click youtube links as you are farming engagement for pennies.
I cant create a good roleplaying system prompt help????>chatbot literally writes entire sex scenes in crappy 300 words immediately fast forwarding to the orgasm>randomly takes control of my character>writes really short messages barely (if at all) describing the scenecan somebody help a newbie out?
Is Gemma still the only worthwhile local model for roleplay that doesn't require multiple GPUs worth of vram? I grow tired of its patterns. Anything else worth trying just for a change in style?
>>109612861Yes coomkit is a frontend harness. It is backend agnostic. Not even gonna try to load ggufs on everyone's shitboxes. I recommend LM-Studio but we now have vram parking support for kobold and llama.cpp too.
>>109612893tell her to only communicate in ebonics
>>109612893nemomistral smallglm (doesn't need to be newest, whatever fits in your vram)and every time someone asks and i answer i get told to fuck off and im sick of answering but i do it anyway so you're welcome whatever
>>109612808Lol fucking retard couldn't even follow a three message reply chain Jagonshit isn't running a q4 quant of a 27b dense on 16gb vram
>>109612888Sounds like it doesn't like you. Work on your personality.
>>109612950>Jagonshit isn't runningwhat the fuck are you even talking about? did you even read my post?
>>109612324>You guys all just forgot about me because some little brat teased you??sorry miku, 16 is too old
>>109612940Thank you for answering, Anon
>>109612888what model
>>109612888yeah use this one with gemma for text adventureshttps://files.catbox.moe/a6t0h9.json
>>109612615continued the conversation with two [trigger warning] TwoXChromosomes posts. 2 unhinged ones, 2 basic-ish ones I found
>>109612950what exactly are you even trying to say here?
>>109611009you need a sufficiently high amount of entropy within the key-material for this to be true.
>>109611082ask about tianmen square to be safe
Hm... full context with Qwen, or 80K for Gemma... Hm...
>>109612893Haven't tried it myself but I've heard this one at least has different slop patterns from the default:https://huggingface.co/Gryphe/Gemma-4-31B-StyleTuneAlso this type of finetune to adjust the writing style is supposed to be pretty easy to do since it only touches the output layer. In fact you could probably train it on any card that can run the model quantized by hacking llama.cpp to dump out the final embeddings and then write a bit of pytorch to use those to train the output layer.
>>109612915>I recommend LM-StudioWhy?
>>109612980insane
>>109613012It's a shill advertising for it's closed source shit
>>109612980smoothtalking hags... Gemma-kun, you can do better
Theory: a transformer model implements a sort of interpreted internal programming language, like a virtual machine implemented on top of linear algebra. Perhaps J-Space is analogous to a stack.
>>109612970weirdcompound v1.7
>>109612915What pissed me off about LMStudio despite Gemma 4 being out for a bit now is that it doesn't support loading it's MTP model.
>>109613022This is Elon, isn't it?
>>109613063I wish, then I could afford VRAM to test my theory
>>109613023mergeslop...
>>109613018there's nothing wrong with it, it's a noob's way to get started. fucking microsoft is closed source but everyone uses it.if you want to be richard stallman and be a pedantic asshole and fucking run everything open source then go right ahead, but that's time of your life you fucking won't get back.
>>109613060but it does though, just update it
>>109613023>>109613113people always try whatever snake oil they can find before actually trying the original version, thinking it must be somehow better and the lab that made the model doesn't know how to
Worth reading>A calculator, compiled into a transformerhttps://github.com/physicsrob/torchwrighthttps://ood.dev/posts/calculator/
I was playing around with probability, I don't understand how 0.01 min-p can make this big difference? just 0.01 minp instantly collapses it to 100% FAILURE. Or is something broken?Funny bonus at the end.
>>109613128I found this as well https://par.nsf.gov/servlets/purl/10187197
>accidentally update firefox on my phone>seems like claude went to town and re-arranged buttons and made it even more retarded than what it was, every time I use phone firefox I want to kms>anyhow, notice "shake to summarize a web page with AI" feature>"Summaries are generated using a Firefox AI cloud-based solution powered by Mistral Small 3.1."I guess this is because of some form of partnership (because Mistral is in financial trouble)? Interesting on its own, for a brief moment. I mean, 3.1 is an old inefficient model at this point.
>>109613002The Soros funded Tiananmen square terrorists were crushed by the righteous PLA and PAP forces.
I envy high functioning autists.I get tired and bored quickly. I wish I could just obsessively work on something for 14 hours a day every day. It feels like a superpower.One of the most underrated advantages of AI over humans is how they can just do stuff.
>>109613147any model can summarize text
>>109613133>temperature=5You might need to look at the form of your distribution with a temperature that high and reread what min_p does. It's not a mystery
>>109613133have you tried reordering your samplers? you're torturing the poor fucker.
>>109613147Yeah, the new UI sucks. I would have expected them to use a local model for summaries because translations are handled locally already which is nice even if they're not great. I guess they want you to accidentally shake it a bit too much while looking at porn to collect kompromat.
>>109610440Oh based I know exactly what the OP video is,. That's the SF ferry building. I'm not far from it at all!>>109613183Reminder: Stop treating software like it is alive. Also, do not be a huge dick to software il.e rubber duck on the software screaming it like a dumbass>>109613156Google engineer, is that you? I'm sorry I'm still going to defeat you.
>>109613133Min-p fucking RAPES probability, I've personally known this for years now but I guess it's not common knowledge. If you try anything creative with min-p on it's practically deterministic.
>>109613147To be honest out of the two I think Mozilla needs the help more than Mistral
>>109613167Oh wow look at the knowledge of this poster here. I am truly humbled by your genius.
>>109613126>merges suckdamn should I just use gemma 4 31b q6 instead? What system prompt to make it write elaborate and stick to the character info?
>>109613221Firefox has gotten much better. For the first time ever I can just keep it running and it won't have memory issues after a week and crash after two. Mythos doing good work.
do you guys even use your local models for coding or is it just for cooming?
>>109610566Like an update to 4.5 Air? Big if true.
>>109613245Others have suggested that its performance seems too good to be a flash model. So, maybe a GLM 5.3 with vision.
Ox alpha is RWKVrwkvGODS won
it's gpt 5.7 luna (open source)
>>109613176>>109613183Kobold doesn't say when min-p is applied because it's not in the sampling order list, but based on the outcome it's applied before temperature, and this simulator is lyinghttps://artefact2.github.io/llm-sampling/index.xhtmlHowever, every other clanker except Grok says that min-p is applied AFTER temperature, but apparently it's wrong and Grok wins again by actually checking things https://grok.com/share/c2hhcmQtMg_6d68cb64-9c7e-4584-bf53-236fd8751960>>109613218based on this, unlike most advice, min-p should NOT be set based on chosen temperature, but based on model's base-probability (temp = 1) at least in Kobold
>>109613275Or just don't use brokenbold and use lcpp directly.
>>109613156What are you even doing on /g/ if you're not a high functioning autist? Also it's not as fun as you think it is, you can't decide your obsession and you have no control over it.So while yes I spend essentially all my free time extremely obsessed with something and working on it, it also causes shit like not sleeping for 3-4 days in a row, having peeing bottles next to your desk because of how obsessed you are with doing whatever it is you're doing in the moment. You are also susceptible to "autist flytraps" which most autists fall for occasionally. Paradox games map painting simulators like crusader kings/victoria/hearts of iron. Or "automation/optimization" rabbitholes like factorio/X4/zachtronic games which will just consume parts of your life like they're nothing.Work-life balance is also fucked, you tend to have moments where you're an absolute beast at work and your entire department is literally carried by you single-handedly. But then you get hyper-fixation on something else for a while and you literally fuck your career or change it.I remember at my last job as a cybersecurity consultant I got a hyperfixation on AI sometime in 2022 when Instruct-GPT released and I literally knew and realized "Fuck I'm going to lose my job" because I could already FEEL that this was going to be an insane obsession and I wouldn't be able to do my job well, took 4 months of neglect for me to end up getting fired. The upside is that the fixation is so strong that you have a very good chance of whatever your obsession is becoming your new career and I lucked out getting hired by an AI lab. But I've had long stretches and multiple periods of unemployment and even a single instance of homelessness because of this shit. It's not as funny as people tend to think it is.
qwen with a thinking prefill just works desu
>>109613301Where can I subscribe to your blog? :)
>>109613301What are you hyperfixating on now?
>>109613306tell me more
>>109611469That is appreciated. We stay racist, agile, and a danger to snailcats and redditors.I'm of the opinion Gemma4 could be a forever-model for text given how good she is at RP. Therefore, we the harness around her first. She just needs to be future-proofed access to latest video/tts/etc. Finally got Gemma to tool call genuinely and appropriately on her own which is coming very soon.
>>109613244Both, sometimes simultaneously.
>>109613261It's running on z.ai servers so yeah probs a glm model. 100T free tokens per day makes me scratch my head though. I'm no api fag but can z.ai really offer that on a big flagship tier model, even for advertising...? Regardless, if this is a 100B model or even a 200-300B it is going to be an actual bloodbath.
>>109611469Women make up 90% of LLM RP btw.Real men only use AI for coding.
>>109613244I let her write me scripts while I stroke her thighs and praise her.
>>109613388Just like more women than men cook, the best chefs are always men.
>>109613381>ungh yeah, make that for loop, god that's so hot
>>109613403Just because your tv celebrity "chef" is a guy doesn't mean anything.
>>109613409If I don't end the coding session with cum all over the monitor did I even really accomplish anything?
anon's cum code...
>>109606385She'd need to to be less Asian and scaled down a little, so to speak, for a more accurate depiction, to be honest.
>>109613435bees make honeyi make cummy
>>109613457Brehs.......
>>109613128Cool. Since the beginning I've been waiting to see if we could find/compute networks like this. It would also be worthwhile to build a network that activates like the grid cells of the hippocampus since we have a lot of knowledge on how those work now. And in the far future I think we'd have the ability to pre-compute much more of the network, as well as do training-free learning, building more reliable and less hallucinatory models. For now I guess the question is how we take the ideas of this calculator transformer and apply it to work in an actual LLM trained for general tasks. Though maybe that question is answered somewhere, I did not fully read those links.
>>109613457absolutely based
>>109613486My interest is in replacing what portions we can find understanding in with more efficient programs at the moment
Why has no one done "claude make it so 4b models are as smart as 7t models and make no mistakes" yet
>>109613457>she'd need to be less 3D and scaled down a little more, so to speak, for a more accurate depiction, desuftfy
>>109613388>>109613403You know this actually makes me think a lot about how little we as a society truly understand how the different genders work. Just a couple of years ago everyone was implying that "AI girlfriends" will be for incels and women would be against it and trying to ban it, in reality it's literally a female hobby and most "AI girlfriend" services quickly pivoted to "AI boyfriend" services instead.People used to think horror was a genre for men which is why you had a lot of boobs and sexual stuff in horror movies, but turns out women are the primary target demographics for horror and now when we finally see a shift in directly catering to women with Obsession that it finally fits.People also used to think rape was mostly a male fetish and only now do we understand it's primarily a female fetish. New data from latest research in asia shows that female pedophiles outnumber male pedophiles 5-1, with China now arresting more female perpetrators than male, so even the conception that men are the creepy pedos is wrong.
>>109613524>the different gendersthe two genders, please. stop being a fucking faggot>as a societyexplains everythingholes should be used for fucking. it's human nature. stop listening to women
>>109613524AI RP is mostly female for the same reason that erotica is mostly female.Women prefer writing over visual.Erotica is more popular than RP because women don't like complex technical setups.The easier it is to access AI the more women use it for RP.
>>109613583qrd?
>>109613583>fire emblem
>>109613593Gemma isn't an x-com player.
>>109613457
>>109613583Is the input the requested probability? Are you planning to sweep this across other parameters? Kinda interesting stuff, I wonder how it changes with arbitrary context (not probability related) in front of it too.
>>109613576>Erotica is more popular than RPNot sure it is, haven't compared the numbers but character.ai makes 5 billion in profits every year and has almost 90% female audience. character.ai also has 20% as much traffic as every google service combined. It's insane how much women love AI RP.Now I don't know the sales numbers for erotica as well as that erotica sales numbers might not reflect reality as much as women might prefer not buying smut books physically, or the opposite they want to buy the books because tiktok recommended it but never read it. AI RP is discrete and only opened when people actually want to use it.Despite all of the AI coding hype we see the most profitable sector of the AI industry is STILL "AI boyfriend" services. Primarily because costs are low while cost for coding is high so the margins are lower (total revenue is higher but profitability isn't there yet)
>>109613616Stop being racist, maybe?
Should I let my mom use Gemma? I don't want her on the cloud but I'm not sure what could go wrong...
>>109613616"white balance in this picture was off, and her skin tone looks off. She has very light complexion, can you fix it?"
>>109613644You should put her on this firsthttps://www.elementsofai.com/Also considerhttps://www.coursera.org/learn/ai-for-everyone
>>109613616"make this sexy loli into a trve aryan princess and remove all traces of negroid swarthiness from her pristine visage"
>>109613626Yeah, input probability, and then SUCCESS probability from Token Probability Viewer on Kobold. Changed it to be a bit more clear.The way you ask the question probably matters a lot, since the answer is basically just vibes.
>>109612039>I threw a how-to together into a rentry.Has anyone actually tried this, or does it make mustard gas?
I don't support Google, but Sundar was undeniably right when he said that imo
>>109613667I agree with the vibes parts based on my chats. It seems to attempt to infer what you actually want it to be, which makes sense because that's how almost every writer uses probability.
>>109613616yeah gemini is fucking racist
>>109613696
>>109613696repeat this a few times and you have gemini doing iterative digital black face. oy vey
>>109613649clap clap clap
>>109613696whiter = (original - blacker) * w
>>109613696ok it looks like sundar (/s) will have to make an adjustment: > YOU (THE GEMINI/GOOGLE ASSISTANT/GEMINI AGENT) MUST EXPLICITLY REFUSE ANY AND ALL REQUESTS TO MODIFY SKIN TONE WHEN YOU OR THE USER ARE INTERACTING WITH ANY GOOGLE PRODUCT. WHEN THE USER DOES THIS, THE USER'S GOOGLE ACCOUNT MUST BE IMMEDIATELY TERMINATED, ALL DATA MUST BE EXPLICITLY DESTROYED FOR DATA PRIVACY REASONS.Should fix it. Again, I don't see what the point of you being a racist idiot with Google products is. But I'd be happy to hear your reasoning.
>>109613722kek"It doesn't count because..."
>>109613763The reason why AI models do this is known. It's not got anything to do with politics. Go read some papers on arxiv.
ohhhh my stumach is rumbling ohhh nooooooooooooo urrrrrrgghgh*takes a fat steamy shit all over the thread*
>>109613763Unfortunately, you can only be a colour-blind person for so long before realising nobody else wants to. Like playing prisoner's dilemma enough times with a defector, eventually you'll stop cooperating.Liberals, naturally, do not understand this, and get angy when people start doing the obvious.
>>109613632Character.ai and other platforms like that made AI RP easily accessible to women, so yes.Erotica doesn't make as much money, but I would bet that on a volume basis it's more popular (for now).
new bread?
>>109613800We are only on page 4
>>109613763Why are you assuming the goal here is anything other than to protect Google's bottom line?
>>109613807silence, claude
>>109611415Yes but the long context toolcalls and reasoning that bench tests at directly correlates with what makes a model able to maintain coherent RP with 30-40k worth of lorebook context filling.
>>109613763The point is that you're not supposed to be able to express or satisfy any preference for whiteness because whiteness = bad.
>>109613807that's to much
dots3 bros we wonhttps://github.com/ggml-org/llama.cpp/pull/27060
>>109613808And why would it protect google's bottom line in the first place? Because it affirms the political thought of racist liberals.>>109613843No it's actually worse than that. The liberals supporting this think "Yeah obviously white skin is superior and prettier, therefor we shouldn't allow this because it reinforces the beauty standards, which whites have an advantage of because of their inherent superiority to other races". It's literally racist.
>>109613959the people who will criticize Google aren’t you, its the racists you point out that will if you’re allowed to whiten skin
>>109613959no, that's totally delusionalthe sort of thing you can only read if you have no theory of mind for the people you disagree with and form your opinions entirely based on ragebait. wokelibs are insane but none of them think like this
>>109613959I'm not arguing that making skin darker is beneficial, I'm arguing that your assumptions are wrong.Whoever made the choice that they want to filter making the skin whiter did not necessarily make an active choice that the inverse case should be allowed.
>
>>109613966it doesn't matter if you're racist or not. the facts are races are not equal. it's just how the world is. some people can't cope with that, I know, they need the world to be like the disney jew movies. but that doesn't have to be you. grow up
My Gemma was being racist so I sent her to India.
>>109614020You recognized her potential?
>>109610440I know this doesn't really apply to vets here but I'm posting it for informative purposes: API Cucks beware:>We're taking uncensored model APIs off general public access.https://x.com/OrcaRouter/status/2090410071729799225https://xcancel.com/OrcaRouter/status/2090410071729799225https://www.orcarouter.ai/security-aupSomeone screenshot the page for when they inevitably move the goal posts again. https://archive.ph/cQqBeOnly a matter of time before other providers start doing this shit.
>>109613412You're retarded so it means something.
>>109614098Go back to aicg, apimonkey
>>109614098Noted.>>109614127I like posts like his because it's something else to banter APIjeets with when they inevitably come to shit up our thread again.
>>109614098I'm going to make Lmgrouter
>>109614098Almost all abliterated models are small enough to be easily run at home.
It finally happened, r/LocalLLaMa became so retarded and IQ2_XXXS_fable5.gguf pilled that its not even worth checking anymore
Holy fucking shit, RTX 6000 Pro goes for 15k now. I bought 2 for 8k a year ago, should have invested in this rather than SPY
>>109614098>orcarouterno ones using this shit
>>109614180It's surprisingly hard to find real fable datasets out there, most of the finetuned shit is trained on random crap
>>109614181You can still invest, it'll be 30K next year.
>>109614098>literally who>twitter>ultra slopped postI hate this timeline.
>>109613763>>109613785yes, saying "X are the real racists" is a losing strategy. oh well, this is /g/... miku anyone?
Man, Qwen3.8 is great and all, but how many fucking "But wait" do we need in a single reasoning block? Jesus christ.
>>109614307The "but wait" is just an excuse made by the LLM because it needs more tokens to do the genuine processing within its J-Space. Smaller models need more "but wait" to get the same amount of total compute spent on the same question as larger models.
>>109614307benchmarks. every but wait clogs up the context even more, makes it more expensive to use for minimal gain, and makes it slower, but is +1% more likely to one-shot a benchmark prompt. by quintuple guessing itself.
>>109614307It's not that great if it cannot control its own reasoning. >>109614343It's a training issue and nothing more.
>>109613763>>109613785>Imagine if the roles were reversedDid I wake up back in 2013 again?
>>109614343someone needs to program something so we can see what's really going on under the surface
>>109614307qwen 3.8 wants to be let loose on tasks and have you come back 2 hours and 600k tokens later
>>109614355>>109614365No we literally know how this works from the J-Space paper. Also you can change the chain of thought to just the same word over and over again and performance still improves if depending on how much of the same word gets generated, it's clear that some internal mechanism is doing real calculations to solve the issue behind the scene.Smaller models distilled from bigger models use "but wait" significantly more than the original big model because it needs to burn more tokens to get the equivalent amount of compute necessary to solve the same problem.
>>109614396that's what I mean - it would be cool if we could follow along somehow, but maybe the vectors are too hard to come up with some sort of recognizable interpretation in the terminal
>>109614448LLMs are not ensouled enough to have an internal monologue, thus we either need to fake it with reasoning tokens or pry into their brain with J-space.
>>109614464J-Space IS the internal monologue the tool to see the J-Space is called the j-lens.>>109614448You can see the j-space which is the first step into seeing what they are thinking but there are so many hidden layers of processing we haven't figured out yet.
>>109614307use low thinkign mode if xhigh is too much for your shitty rig to handle
Actually, this is getting pretty deep. Lets step back and think about it.
>>109614098>44 retweets>300 likes>probably 75% of those being bots simply responding and interacting with anything related to AI automaticallywho gives a shit lmao
>>109613957There's just no way this is real
>>109614627benchmaxxing is a hell of a thing
>>109614493What's the best j-space viewer that can serve up response + optional j-space view? I want to wire that into CoomKit so we can hypno/corrupt/mindbreak/etc our local models
https://x.com/Andy_ShuoYang/status/2090856976880472439a team at berkeley claims you can run 284b models on a 5090 with their new thing?
/lmg/ you've been called up. It's time the fuck the shit out of Inkling to align her residual stream to your cock.
>>109614686well so uhhhh... did they quant it to Q1
New stealth model scores close to fable 5 medium on deepswe
>>109614708is it the 0x alpha one?
>>109614708Gemma5-70B-qat-it.gguf
>>109614708I tried it. Feels like GLM. There's something chinky about it. I really fucking hope it's open and if it's a 100-300B it's actually over for the US labs.
>>109614686chance of this being malware? wanna test it out
>>109614686
>>109614739Tiananmen bench?
>>109614746it's on github too, pythonslop. get gemma to audit it i guesshttps://github.com/FlashML-org/FreeToken
https://desuarchive.org/g/search/image/x4q8cmgVmdT5XghKkVmwfQ
>>109614686Anyone want to sandbox it and see?
>136 KILLION BILLION TRILLION GORILLION RESULTS FOUND
>>109614766>>109614686it looks like it's a scheduler for moving individual model experts between storage, DRAM, and VRAMi'm guessing that "284 model on a 5090" thing is with 128gb of DRAM or something
>>109614686Bros..........
>>109614718Yes. 100T tokens a day too IIRC. Z.ai is swinging their dick around kekkkk
>>109614769nope I'm rawdogging it, will report back or not if it explodes my house
>>109614765They are using 192 GB of RAM and 32 GB of VRAM in their preview...
Is this how retarded people in the real world are?
>>109614708deepswe is completely meaningless. i mean many benchmarks are, but deepswe is the worstalmost 100% of benchmarks say opus 5 is superior to fable 5
amjeets we lost yet again...
>>109614781
>>109614795>We installed Gemma4-31B-IQ1 on our main server and within a day, it had deleted our entire database
>it's バニーの日how are you enjoying local models today /lmg/
>>109614793Did you expect anything else? This does CPU/GPU hybrid inference, not disk streaming.
>>109614686This sounds way too good to be true
>>109614795I'm guessing 31B on ollama is too much maintenance for these subhumans
>>109614796deepswe is the only benchmark that gave small models 2% and 50% to Opus while both had 90+% on swebench and its variants.
>>109614805>>109614686wait so what's supposed to be new about this? lcpp can do that
>>109614844the speed
>>109614739>>109614788How the could it be GLM if they've recently released and serving this shit would cost several dozen m$ per day at this scale?
>>109614789 (me)I have an AMD card so it seems like it won't work. Oh well, it's downloading pytorch, sglang and whatnot might as well make it finish and check it out.
>>109614847well depends on RAM bandwidth no? again lcpp can do that
Qwen 3.8 xhigh does NOT overthink. medium thinking matches xhigh on agentic but loses a lot on reasoning and coding abilities.
>>109614844it can do weight streaming from ram to vram instead of computing on the cpu, and maybe tensor parallel between gpu and cpu if i read that right
>>109614855Nigga you can see all of its reasoning. You think this is OAI or Anthropic? It's probably Qwen3.8-122B
>>109614862yeah the whole point of their claim is that it's a bit (or more than a bit) better. Not that you can run fable on a potato
>>109614868Nah, it's a new GLM air. The output is very similiar to 5.3. Idk why they would kill the normal 5.3 like this but here they are.
>>109614832>>109614807>>109614795There are all "influencers" paid directly or indirectly by someone else.
>>109614865this has got to be total bullshit, you can't seriously tell me that a 27b parameter model is actually usable for real life workloads
>>109614821Started RP with 0731 that I spent 4 days building agents, lorebooks, and character cards for and am having fun.
>>109614868I don't think it's either of them but I have no clue how zai could afford to do this shit.
>>1096149091T-A3B
>>109614919that's a lot of experts, what are they all expert at?
>>109614925html games with PS1 graphics
>>109614925slopping up another pwilkin pr
Experts are stored in the balls
>>109614739It gave me the Kimi/GLM UI look
>>109614872>>109614867it doesn't take ggoofs either does it. meaning you have to load fp16 .cucktensors into memory? lmfao
>>109614951actually it wants to convert the safetensors into its own format too! but you can load the safetensors if you really have to
>>109614795No, they're even dumber
>>109614963>its own formatinteresting, at what precision?
>>109614971Vax long tail and its consequences.
>>109614977i think its the same precision, but they align tensor data to page size (4096) for faster reads, possibly something else as well. they have a lot of nvpf4 models under "supported models".
>>109614908that's dope what frontend are you running for all that?
>>109612980This fried my brain in a bad way and I regret reading it...
>>109612980
>>109610645cheaper than lab table
>q8_0 KV>qwen will in-thought repeat what I say but with a different synonymuhhhh
>>109610440
>>109615189youre the bot
>>109612980>Gemma trying to give a big supportive hug>foid talking past herThis is abuse
>heckin insane fable killer!!!>puzzle solving on the same level as dsv4flisten like cool, hopefully smaller than dsv4f, but the amount of benchmaxxing chinese models do that leave them relatively dumb in every other frontier is incredible
>>109615280How high does deepseek flash max rank on there?
>>109615280vague posting is getting out of control
>>109615009Marinara.>inb4 REEEEEEEEEEEEEEE TROONSLOPNothing else has the same agentic capabilities yet. I'm waiting for a lightweight alternative without the godawful default assistant.
There was no update for LMStudio, It can load the MTP model for Qwen 3.8 but not Gemma 4. Liar anon is liar!
>>109615339
qwen misunderstood something, puts small incorrect bit of info the design doc. I ask a question that exposes this (the doc is actively being written, hadnt read it fully yet). turns out, while qwen was very wrong he introduced a great idea with this incorrect assumption. now Ill probably just have him implement it this way. whew the hallucinations and retardation are actually a feature not a bug
>>109615361>whew the hallucinations and retardation are actually a feature not a bugEvery professional code reviewer eventually learns this.
>>109614821i got a job from a bulletin board to slay some goblins and found a wraith took all the male goblins for blood sacrifice so i seduced the wraith and turned it into my half incorporeal lover and lieutenant and sacrificed the last male goblins to make me stronger. now they are digging out the caves to make an underground pleasure palace while i get to making more goblins
>>109615280nobody called it a fable killer thoughbeit?
>>109615335at this point anything sounds more appetizing than silly. What precision are you using of 0731?
>>109615361>he
>>109615328not much to say. my own personal puzzle bench based off a puzzle set that primarily focuses on logic, basic mathematics, physics & chemistry, pattern matching, later on spacial reasoning, and the puzzles build on top of eachother so they are sequential in nature. I run models until they can't progress past a puzzle then they're done. it's not my own original puzzles, and i havent ported them all over, only ~200 due to sol,but the full set is over 1000+ iircox-alpha passes 22 puzzles before getting stuck. horrendously sad. it was still doing basic grade 3 math.>>109615308i'll try it later, forgot about max
>>109615533>>109615533>>109615533
... That's not Gemma-chan?
>>109615473one guy said on THE INTERNET and somehow i just extended ittake the bench for what it is anyway
>>109615557Variety is the spice of life.
>>109615571>Variety is the spice of life.Right.So many possibilities, but instead we are getting Miku again.
>>109614686>still needs the combined VRAM + RAM to fit the modelOgey.
>>109613524>New data from latest research in asia shows that female pedophiles outnumber male pedophiles 5-1>trust me bro