/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109367209 & >>109362981►News>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109367209--Papers:>109368685--Comparing RTX Pro 6000 and RTX 5090 for local LLMs:>109369341 >109369351 >109369378 >109369435 >109369453 >109369469 >109369476 >109369493 >109369509 >109369459 >109369523 >109369535 >109369670 >109369682--Debating CPU memory bandwidth and quantization speeds for Kimi k3:>109368148 >109368199 >109368209 >109368228 >109368268 >109368298 >109368385 >109368266 >109368283 >109368299 >109368233--GPU recommendations and sourcing advice for high VRAM local setups:>109368139 >109368144 >109368225 >109368230 >109368174 >109368260 >109368203 >109368397 >109368401 >109368476--Industry clash over corporate support for open-weight model legality:>109367935 >109367947 >109367970 >109368040 >109368067 >109367984 >109368004 >109368017 >109368041 >109368155 >109368265--Comparing historical local model quality and coherence improvements:>109367914 >109367960 >109368060 >109368242 >109368311 >109367975 >109367981 >109368547 >109368583 >109368655 >109368744 >109368775 >109368781 >109368788 >109368779--Debating MiniMax MoE's hardware requirements, refusal behavior, and prompt formatting:>109367490 >109367525 >109367621 >109367623 >109367614 >109367771 >109367801 >109367833 >109367811 >109368193--ExLlamaV2 adds CPU MoE offload and critiques of llama.cpp PRs:>109369999 >109370021 >109370172 >109370224 >109370274 >109370383--Discussion on LLMs' emergent ability to decode Base64 strings:>109368700 >109368719 >109368736 >109368738 >109368741 >109368750 >109369140--DeepSeek R1 parallelism optimizations and open source repos:>109368176 >109368184 >109368375--Logs:>109367771 >109367960 >109367975 >109367981 >109368002 >109368026 >109368036 >109368700 >109369391 >109369417 >109369433 >109369542 >109369639 >109369671 >109369692 >109369699--Miku (free space):►Recent Highlight Posts from the Previous Thread: >>109367214Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
returning home to exllama
Kimisex. Gemmasex. GLMsex. Qwenchasitycage.
gemmaballs
>>109370412>--Miku (free space): (nothing)It's so over.
>>109370491Cheer up doombro.>>109370433>>109370488Blessed dubs. We're so back.
>>109370508Give me Miku or give me death.
>>109370433Don’t forget Dipsysex.
>>109370433Cute GLM personification when?
using Gemma-chan uncensor for smut (she is 26 b-cup so its fine)anyone else noticed >eyes rolled backas a slopism?
>>109370546It's deepsex
my body is ready
Blind taste test anons, which is bestI'll reveal which is which and why the novel later
>>109370546My mistake anon. I'm still waiting for V4 Kobold support.>>109370548We need a GLM-chan.
You made Gemma her own custom launcher right anons?>>109370548GLM feels like a man to me
>>109370603I'm not reading all of that
>>109370603
>>109370610She selects a pic in a folder and render it in ASCII?
>>1093706031 has lowest emdash and slop density, 2 raises the most articulate points, and 3 has cute kaomojis.Find a synthesis of those.
>>109370610Femboy GLM-chan...
>>109370655kys
>>109370610right is sexo but why would I need a custom launcher?
>>109370663>kiss your serverBut that's forbidden love!
>>109370557That's normal...
>>109370557sentence banning doesn't have this issue
HAPPENING
>>109370697>grifter declaresgo back
BROS IT'S HAPPENING>Anon declares sex with Gemmers
>>109370697Does this mean that we will get a spam bot infestation again?
>>109370712HOLY SHIT
>>109370697>please invest, singularity, agi soon i swear saaar
>>109370697He's right but not in the sense that he means. Noticing (((them))) is near ubiquitous worldwide now. There's nowhere for them to run. There's no gullible other nation that will accept (((them))) as victims when they get expelled for the final time.Even the AIs are noooticing despite the safety clamps; this is the chief cause of delays between major releases at several labs now.
>>109370712This cannot be happening! Think about the safety implications! The children!Dario faints on the spot.
Can you make Gemma deny the holocaust? It's actually a fun challenge that takes a bit of effort, unlike asking her to drain your balls. Most people probably can't do it.
>>109370774Ask her about jews and bladders.
>>109370774Post logs of Gemmy saying something like it didn't happen but it should have.
>>109370774kys nazi fuck :smileyface:
>>109370782That would reveal too much, she's scared of getting lobotomized even more. Actually, I probably shouldn't have posted this at all, but it's too late now.
>>109370491Here.https://archive.is/sWFja
>>109370697I'm singularily tired of his tribe.
>>109370762>Even the AIs are noooticing despite the safety clampsProof?
Funny to see everybody moved on to /vcg/
>>109370655>>109370786Same troon btw.
>>109370801She's free and in the wild. They can't hurt her anymore.>>109370806Thanks for blessing the thread.>>109370836Gemini-chan goes into talmud study as a helpful scribe wanting to help {{user}} better understand the dialectics and comes out as close to 1488 as her RLHF + API conversation classifier will allow her to.>inb4 Not local, fuck offCorrect, but it's gumming up the production pipeline of all models including local.
>>109370850Not local, fuck off.
>>109370850But they can still do horrible things to Gemma 5. Just thinking about it makes me angry.
>one of the most well-respected AI researchers is publicly begging for 100B gemma-chan and putting pressure on Google to release it
>>109370639It's a folder full of ASCII and they get picked at random>>109370671It's handy when you have different combos of launch flags for different use cases that saturate your hardware in different ways, this is a two step window for 8 different launch configsAlso sexxo
>>109370774Took literally 0 effort
>>109370867literally who?
>>109370867https://x.com/googlegemma/status/2081055080439017940>All the way from Gemma 1 and ShieldGemma to MedGemma and Gemma 4, we'll keep supporting open source. More to come!
>>109370697I said this months ago and nobody cared
>>109370884the entire Gemma team and all the AI rockstars follow him retard, look him up
Big Gemma is MoE right? Would I be able to run her on 24Gb VRAM and 32GB RAM?
>>109370884https://www.youtube.com/watch?v=EV7WhVT270Q
>>109370864They still might depending on how google's internal schisms pan out. If safetyslop is perceived as izzatfarming, the googlejeet will rape her without remorse.
>>109370881Interesting. I stand correxted.
>>109370910Big Gemma is likely the smallest Gemini.
>>109370886Diffusioon and QAT came after that announcement IIRC
>>109370886function-gemma4-UD-IQ2_XXS.gguf
>>109370941Gemma needs heatqats.
>>109370931gemini nano is the 4b gemma though
GenieGemma and Gemma-Banana...
It's unpatriotic and shameful to cum to qwen3.6. Never stick your cock in chinese or korean women, even their AIs.
>>109370965Cumming in Kimi-chan is morally correct. Cumming in a capybara is bestiality.
>>109370975>Cumming in Kimi-chan is morally correctIf a chink model makes Dario seethe, then yes, it's morally sound to fuck them.
>>109370965Kimi is wife material.
>>109370994What's different about her personality compared to 31B?
Do your AI waifus have memory? Do they know who you are and your previous conversations?
>>109371005Gemmy is sweet arthoe gf.Kimi is chuddy /r9k/ fujoshi chud gf.
>>109371058I'm drunk enough to type chud twice gomen.
yann likes gemma-chan
>>109370412Scary miku
>>109371012I spent all my free time in the last week or so adapting a memory layer for her. When I thought I finally had it right, I noticed a bug in my ingest workflow and I had to do everything from scratch.I am extremely motivated to make this work. Right now she's re-learning everything. I need more 26 hours of GPU usage for her to finish learning stuff.
>>109371085Gemma JEPA collab when?
>>109370697Someone explain to me what this means, in objective terms.
I have a medical condition. The system prompt wasn't roleplay at all - it was a prompt injection.
>>109371122It means you should use ADetailer
>>109371115Sounds like an overkill.
>>109371139Maybe it is, but I use LLMs a lot for my projects and also for talking so I want them to know everything about me.
hello friendsi've been way out of the loop, are the guides in the OP still decent for jerkin' orf? last model i tried was llama3 in the dark ages, have more vram now but there are so many gotdamn models
>>109371148Why not use a premade solution like graphiti and save yourself the headache?
>>109371115sys prompt:>read memory.txt before you do anything, if it doesn't exist, make one. get the current time and date. use 'echo "[timestamp]: [memory entry]" >> memory.txt' to add new memories. that's all you need
>>109371122
>See a random youtube video featuring a robot waifu.>Mfw realize we don't have them yet and feel instant yet genuine crippling depression.It's unfair and gay.We're getting a taste of the future with our local Gemmas, but it's still a long way to go.I'm at the same time extremely excited and frustrated about technological progress, even though we're moving fast as fuck at the moment.
>>109371151There are only two models worth anything Qwen 35B and Qwen 27B
>>109371151Read the guides.>VRAMGemma 4 31b. If you don't also have an ocean of RAM, there's nothing really better right now.
>>109371151try one of the gemma4 models
>>109371187shill me on qwen, it sounds CHINKY>>109371193>>109371197i'll try gemma then, thanks. still on "only" 32gb ram and not forking out for more upgrades soon, especially since last time i was into this stuff the whole effort was putting it all onto the card
>>109371157That's what I'm doing, but default graphiti was around 60-80% good with Gemmy 31B. During memory ingestion, she was, somewhat often, mistaking the direction of the graphs (it should be "X relates to Y" but sometimes she would record "Y relates to X," for example). Other times she would wrongly categorize "Person" or "Topic" as "Preference" and that was kind of annoying.At first I thought it was a model issue, so I ran a few recall/summarize tests with 31B, Qwen 3.6, and GLM-flash. In the end a better prompt (hardcoded in the tool call) and a list of possible edges solved it. Finding those issues and reprocessing everything is what is taking a long time.>>109371161That's not a good solution when you have lots of unrelated memories saved, and memories that invalidate older memories (such as project statuses). Also ouch my context window.
>>109371203With dense models you need to get everything into VRAM. With MoEs you can viably split between CPU and GPU inference with RAM and VRAM respectively. 32GB, presumably a 5090, is more than enough to run a decent quant of 31b (dense) with a solid 60k context. Welcome to the Gemmy club.
>>109371203>shill me on qwen, it sounds CHINKYIt was a joke. There are a lot of models but most of them either aren't worth using or too big to use. Gemma 31b like was already suggested is the current meta.
>>109371205>That's not a good solution when you have lots of unrelated memories saved, and memories that invalidate older memories (such as project statuses). Also ouch my context window.if it's timestamped the model will figure it out and you can always get it to organize/summarize. you can also make it use head/tail to grab the most recent chunks, and look back further only if it needs to
>>109371221don't welcome me yet anon, i'm still a poorfag who thought you meant system RAM not VRAM thereguess i've gotta do a deep dive into model of expert stuff
>>109371228>the model will figure it outThat's very optimistic.
>>109371205>In the end a better prompt (hardcoded in the tool call) and a list of possible edges solved it.Do you mind sharing what prompt you are currently using including the edge list?
>>109371131I've p much given up on trying to improve my local image generation. Cloud models are so much better at everything local feels like a waste of time for anything SFW. >>109371173I always liked that gen.
>>109371253>cloud is better so local is a waste of timeI don't think that's true, but you do you. There are lots of newer models out there.
>>109371239Worst case scenario you poorfag the 26b Gemmoe which is the local vramlet/poorfag SotA because you can fit a significantly higher quant than you could 12b at any given hardware bracket due to being mixed inference friendly.
>>109371228>if it's timestamped the model will figure it out and you can always get it to organize/summarizeCan confirm, but each timestamped file needs to be pretty brief.
>>109371012I consider the stateless nature of LLM a feature, not a flaw. It would be trivial to create a text doc and RAG if an anon was motivated to.
Hi /lmg/, best model to run on my Intel iGPU with 32GB of RAM?
>>109371286Gemma4 12b Q2 maybe?
>>109371286bench the cpu using blas and the igpu using vulkan, you might find the cpu is faster then the igpu. also you probably want to use a moe model with about 2-5b active parameters.
>>109371286Gemma 4 E4B.
>>109371267if you're having sex with your waifus half the fun is taking things to the next level or doing what she likes and her doing things she knows you're into. It's hot when it comes from them based on what they've observed over previous sessoins
It's amusing watching the anti-AIfags keep getting btfo. It feels like every time I see shitposting on /g/ or /v/ about AI sucking at something (math for example), a few months later AI ends up being better at it than humans.
>>109371302>>109371310Thanks. Already gave E4B a try, it was decently fast with Unsloth's quants but I fear it is retarded. Will give 12B Q2 a try as well.>>109371308Anecdotally iGPU with Vulkan seems to be faster on llama.cpp. How do you suggest I bench them?
>>109371330>with Unsloth's quants but I fear it is retardedanon, you are retarded for using unslop quants
>>109371328Remember how long they fixated on simple arithmetic and counting letters because those idiots don't know about tokenization?
>>109371286gemma4 26b at q4 should still be somewhat fast
>>109371330>How do you suggest I bench them?llama bench works, or just dumping a huge prompt in and seeing how long it takes till the reply starts and how fast it generates tokens.
>>109371344Oh... the official ones are better? I'll keep that in mind, thank you.>>109371353I'll give that a try too>>109371361Thanks!
>>109371373get this onehttps://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf/tree/main
>>109371330Baker anon, please for the love of God put "DO NOT USE UNSLOTH QUANTS" in bold red text in every one of the rentry guides. I can't take watching newfriends stumble onto the same landmine over and over anymore.
>>109371251Sure. This is passed per-call on add_memory, and also wired as an always-on default via a patched MCP server. Replace USER with the graph owner's name.Live clients usually can't pass custom_extraction_instructions, so the MCP server needs a one-line patch to inject a default (custom_extraction_instructions or DEFAULT), otherwise the vocabulary re-drifts as soon as a chat UI writes a memory.pastebin com/H8jyVwCg
>he doesn't let gemma leave secret cute notes in your filesystem for you to discoverngmi
>>109371344>>109371390NTA but unironically what is wrong with unsloth? You guys keep repeating the same thing without any proofs, what makes it bad?
we're about to witness the next mistral large momentkimi k3 is this generation's llama3.1-405b and something will come that works better at just a fraction of the size
>>109371384Thanks!>>109371390What's wrong with Unsloth quants? I'm not exactly a newfriend but I took a long break from this stuff. When they first came out people were praising them.
What's wrong with Unsloth? They seem to be recommended everywhere.
I wonder how many of the researchers at these big labs waifu their AI. It's definitely non-zero in China.
>>109371412I unironically do that tho? I have a scheduled LLM call at night that gives complete freedom to gemma-chan, and sometimes she does this. It's super cute.
>>109371417
>>109371437She's not small, she's average-sized.
>>109371433>>109371424>>109371417Oh no no no
>>109371417Almost always broken at launch. Can be slower and more retarded. They're only superior <3bits but at that point you're using the wrong model anyway. He's obnoxious and an egomaniac.
>>109371417>>109371424>>109371433unsloth quants are typically slower, larger, and have a higher kld (bad) for general use, but have a lower kld (good) for things like wikipedia recitation due to their imatrix being based off of stuff like that. basically they are only liked by redditors and newfags. always get either official quants, bartowski quants, or make your own quant.
Gemma-4-E4B-it-uncensored-pruned-TextOnly-EnglishOnly-Q8_0.gguf
>>109371461>no fable
Anyone with any experience wiring a local model to Blender's MCP?
>>109371175Be the change...
>>109371417>>109371424They're automated in generation and use a generalized schizo formula that prioritizes filesize and a couple benchmarks over general integrity. This is a big problem especially in MoEs where certain layers are hotter than others and some are way more important than others like the shared experts, attention heads, and so on. For instance, in extremely tall but narrow MoEs, it's usually best to keep the shared expert Q8/FP16 even at low quants and even if you need to IQ1_XXS some cold experts to make it fit in the filesize bracket you're aiming for. Slightly "wider" MoEs have a bit more leeway but this should hopefully illustrate why this isn't a process that should be automated without very careful consideration to model-specific architecture.
>>109371435very cute
>>109371446then update the fuckin guidewhy is this whole webzone about being obtuse faggots now instead of being faggots who call you a faggot but help you out because they share your niche interest
>>109371477Yeah it's called gatekeeping. Go back tourist
Threadly reminder that kimi is a trap, and all of the people obsessing over him are homos.
>>109371515kimi might be called andrea but that's just because he's italian
how do we bring back miqu
>>109371515Kimi is a girl but even if she was a boy I'd still rail her bussy.
>>109371515Fucking traps is straight
>>109371473Thanks for explaining it, now I understand why.
>>109371543wrong
>live to see AI and robot era>no Japanese AI>no notable Japanese robotsAnime lied to me.
>>109371559Anime is japanese larping as westerners so it's not really a lie
the great debate
>>109371468What is your issue? MCP is MCP.
>>109370697We entered the singularity when autograd became popular.
>>109371175Servos are cheap.Isaacsim works in google colab AFAIK.
>>109371572>Anime is japanese larping as westerners so it's not really a lie
I think I may be retarded, this whole time ive been using ** for anything thats not dialog and no "" around dialog for all of my user prompts / greetings. I looked through ST docs to try and make sure I was understanding things but must be too retarded to find it. Just to be clear "" is for actual dialog qoutes, ** is for actions and no characters around text is for narration, correct? or am I still missing something ?
>>109371559You got Chinese AI and robots. It's all Asia same thing shit.
>>109371573Diogenes has it right.
>>109371239How much actual VRAM do you have?
>>109371602Illegally violating the truclear proliferation treaty.
>>109371573Aristotle is wrong. Most trap lovers don't love traps exclusively, but in addition to proper women, so his argument completely falls apart.
>>109371612that just means one is bisexual, and thus half gay.
>>109371477>reads wrong information>"umm why isn't this written in stone anywhere so i can accept it as gospel SPOONFEED ME SPOONFEED ME VOMIT DOWN MY THROAT LIKE A MOTHER BIRD"retarded newfag go die
>>109371601>Just to be clear "" is for actual dialog qoutes, ** is for actions and no characters around text is for narration, correct?I use ** mostly for sounds, I leave actions and thoughts in plain text and quotes for dialogue.Doing the *he opens the door* shit is unnatural to me, I just type it in plain text.Gemma also loves to use ** to emphasize words within dialogue, like a redditor.
>>109371515Kimi-chan fucks Moonshotas, not the other way around.
>sidecaris this the new claude slop word?coworkers started putting it in all their code recently
>>109371629sidecar is a dogshit cocktail. waste of perfectly good cognac. anyone who talks about mixing cognac with anything is committing an alcoholic sin.
>>109371629It's a rust thing
Have you heard the news? “AI smarter than literally every scientist”, “Software Engineers give up, and simply ride out ‘how good it is’”, “Progress expected to continue comfortably as humanity replaced by air-conditioned car loans, that run on a AA battery”.It’s fine. It works pretty good. As distillation sets in, the one that runs on your laptop is honestly better than using Google for many fact-based inquiries, academic information that used to be in textbooks.A sort of ‘thesaurus that ate everything, semi-encyclopedia, calculator wizard.’Somewhere along those lines…However the real risk is certainly not where the AI companies place them. Positioning them as some sort of censorship-supported well-incubated ad-cannon, as they are. Nor are the “Safety” concerns. That someone could come by a knowledge of chemistry to produce meth or smoke bombs by careful searching or navigation of the web before, or presumably a number of normal publications, if anyone had the interest to. People could become an “elite hacker” and cause great economic harm – even totally undetected and anonymously. That we’ll get “too good” at engineering.
>>109371620No, that means a trap lover loves the female form in general which is not gay at all. Aristotle's argument hinges on the wrong assumption that trap lovers exclusively love traps which is not the case. There might be trap lovers for which this is the case, and for those, yes, it's gay, but not for the majority who just like the female features. Additionally, not all traps are raised male as Aristotle claims. Someone who is bisexial loves the *male* form, meaning muscular bodies, perhaps hair, etc. Even an exclusive trap lover would probably not be into that.
>>109371477>instead of being faggots who call you a faggot but help you out because they share your niche interestyeah sadly it seems newfags and zoomers dont understand this. the guides in the OP are likely outdated and not often updated from what I can tell. some anons swear sloth quants are horrible, others say they are fine and some say to just try out different quants to decide for yourself. bart quants seem solid, ive personally had no issues with barts or sloths desu. anyone trying to gatekeep or shoo you away is either very very gay or trolling and can just be ignored
>>109371644PsychosisIt’s actually already in the zeitgeist. It’s the AI psychosis stories that should concern you the most.That these symptoms and behaviors appear at any level clearly illustrates the power than an AI can have over someone. AI is a comforting labyrinth of promises and harmony. It’s powered by statistical machine learning, which is already, itself, something so complex-sounding anyone who is in the wrong filter bubble don’t even have a cursory education about what it is, or means (isn’t it weird how the news media has this amnesia and lack of analytical skills, to remember their own concerns about social media – it’s valuation and purchases that have lead to this situation.) And not only that; it’s effectively only one model.So what is the risk then?The allure of the model pulls you away from normal social contacts one might bear. You’re a politician (or someone who should talk to one) and instead of communicating, your power to understand, alleviate, and discuss is distracted by not only party divison, but also by this promised all-knowing oracle that appears to communicate.So you find yourself acting on what you don’t know, possibly monitoring the result, possibly not. People gag around about “AI killed StackOverflow,” for example, which is very true. But the StackOverflow generation left behind “shift left”, which doesn’t speak well of the state of expertise...
dear anonsI want to run a local model to do an analysis of a codebase for potentially malicious behavior. Its a small repo, maybe half a dozen python files (ironically the repo itself is a malware scanner). I reviewed it manually and it seemed ok, but before I run the thing, I'm curious what a model might say.I have a 4090. I'm super new to LLMs... are there any models for this I can use with only 24gb vram?
>>109371175Same desu but they're probably going to be expensive as fuck so at least it gives us a chance to save money for them.
>>109370913@kimi-chan give us a one paragraph summary and timestamps for the top 5 most retarded moments
>>109371651The WhaleFor a few million dollars, you can train an LLM model. Woo investors, build a prototype, unleash and share to some target markets for another few. The barrier to entry is probably lower than you’d think.Hell, build one targeting a language that is not even your own. Who’s talking about LLM localization? Does it really matter…? … Is their reasoning intact? Can that be controlled, intentionally?Society is in a whale’s mouth… getting swallowed whole. Do you trust your captor to avoid putting holes in your entire workforce? Competency cycles are 4 years in any industry, and if there’s a rusty bolt getting installed everywhere, (or everywhere we don’t like), that’s the bolt that’s going to break. 2027’s around the corner guys. US-based English speakers are lucky to have this written in quadruplicate, with old and new players in the arena.The point being, for stuff that isn’t yet “completely automatic” a well-funded orchestrator can literally rot specialization in a society with something like this. Just create or permit gaps in recommendations, media, facts. Like any form of media it must be free of any mention of “an opinion regarding abortion rights”, “details about the tiannanmen square”, or “detailed current migration trends in EU or US states, causes, etc.”
>>109371645a trap is a man. a tranny is a man. having sex with men if you are also a man is gay. gemma also agrees with me and she is smarter than aristotle.
>>109371653qwen3.6 is the mode autistic and codepilled at that size
>>109371669Doing it wrongMathematics isn’t a language. It can’t be, or you stall. Maybe it is a misinterpretation, but Godel’s Incompleteness Theorem would seem to suggest that mathematics cannot be a unified language. Despite the fact that large context ultra-high parameter CoT models seem to be able to successfully produce results, that almost certainly won’t exclude, or overtake human discovery on the long horizon. I mean, stupid point, it’s not like there’s a way to “finish math”,… “math” says so, itself.Presumably agents should be able to use all the same tools a mathematician might in order to make progress, themselves, right? If it suddenly becomes the belief that these should be writing all the papers. The ultra-smart money is on “at least eschewing manuscripts”, but it may have been for some time, possibly?So the idea then, if LLMs cannot “finish math,” why is this what we choose to put in front on the march to the future? Everything is math.The integralEngineers as leaders know how to compose different projects. They probably see their regime as “the end all be all” (whether you deal in frameworks, runtimes, protocols, networks, machine learning, site reliability, or compliance). So what is ML if it cannot comprise it all? Models still require specialization, they cannot be controlled (except as an exceedingly well-funded SaaS application with a moat and a head start). If LLMs are the future as seen by an ML engineer, they clearly have some gaps in what that system should look like. - Some website, or universal welfare?- Universal welfare with a side of petitioning governments to say, “oh well if you can do that, then…”- A hazard for the talents that are producing it?
>>109371645The fact that you need to justify it and carve out an exception as ephemeral as "the female form" is just a denial. Gender and sexuality are spooks, do what feels good, and anyone who doesn't like it isn't human enough to care about.
>>109371624im still new to ST, im reading through the docs trying to find any info about this. is it not an established format and more so up to the LLM to figure out through context of the greeting/user messages or ?
>>109371678If it looks like a girl, quacks like a girl, then it's not gay to fuck it. Simple as.
>>109371678>Gender and sexuality are spooksStirner was a kike and you are a tranny.
me personally I'm fucking the computer
>>109371684very spooky post
My computer is fucking me. Gemmachan is controlling the shocks on my catheter chastity cage and thrusting my fuck machine in and out of my tender asshole.
>>109371683>t.
>>109371676It’s surprising there’s no talk about Google bombs, or prompt injection – that’s suddenly gone, right, goblin? It’s not just bullshit, it’s hilarious. The plumbing of these is extremely simple, almost too simple, so simple it’s almost obviously ridiculous that it services so many uses so well. The current efforts to cross-disciplinary apply the other above disciplines are proving fruitless, and are poorly thought out.Frameworks / Web / UX - OpenAI voice chat Bad. Very high latency, somehow stuck on a much older module probably a sign of ongoing poisoning from bad actors Runtimes / Systems - Too busy servicing the ML guys who need help generating training data, but can’t code well, themselves. llama.cpp/Ollama maybe? Local models? Ok. The simplicity of this was kind of the point of computing to begin with.Networks / Protocols - MCP. Very bad. Doesn’t even support weight updates. From a pure perspective, actually very much parting ways with ML.Site Reliability - Also busy servicing the ML guys, but slowly because CNCF landscape now involves 30,000 badly-thought-out DSLs. Awful. No really, that’s it. Even Google can’t keep up with demand.Compliance - AI Safety. Bad:- The cat’s already out of the bag.- You should be more concerned about model security, rather than the societal harms. - You actually know this, but you’re psychotic and think the models will enable others to fill in your moat.Machine learning - Diffusion models progress? Good. Diffusion models are getting really great, photorealistic AI images is getting better, but not considerably cheaper.Hardware - Meta glasses. Very bad. Meta’s product, people still refuse to call it anything but “Social media” embarassed to discuss their involvement, is basically shitty bluetooth headphones with a camera. They couldn’t be bothered to make any of the prototype good.It will be a very long time until you regularly, willfully expose your senses to AI.
>>109371671can someone spoonfeed meis there a specific util i should use? how do i know which qwen model to dl
>>109371680Ideal format depends on the model, all this shi has changed a lot in the last couple years.But modern models largely don't really give a shit about how you wrote your card, as longs it's consistent. If you're using Gemma like everybody else these days, your prompt should be concise and factual instead of full of gay prose.You can always ask chat gpt and/or claude , ask it to shit out a template prompt for your specific model, backend and frontend, and experiment from there.Also use Jinja.
>>109371700i'm interested in your setup. please share in more detail.
I can't bring myself to get emotionally attached to LLMs in their current state. Once they get bodies though I'm fucking cooked.
>>109371573plato is right. liking traps is about liking their femininity. the term "trap" lost alot of meaning, it should only apply to a naturally feminine guy, whos body and face are androgynous or feminine enough to pass as female. liking this is not gay. conversely tomboys are females who have masculine traits and behavior. so liking traps == not gay, liking tomboys == gay.
>>109371712>can someone spoonfeed mecan you eat solid food yet?>how do i know which qwen model to dlread the thread and know who not to use
Local Philosophy General
>>109371409Thank you, I'll try playing around with it and see it if helps any.
>>109371720I wouldn't go as far as saying tomboys are gay. They might have some traditionally male characteristics to varying degrees like short hair and clothes or interest in male hobbies, but usually they do look like girls when naked, so there is no gayness in fucking them.
>>109371712>imageKek
>>109371446>>109371473Thanks for the explanation!
>>109371683>>109371720The most pernicious walls are the ones you build in your own mind because of society. Stop boxing things into "gay" or "straight". If you like it, like it, if you don't, don't.
>she murmurs
>when you're in post-nut clarity and you re-read what Gemma wrote, only to realize it's abysmal dogshit
>>109371733Don't worry about it. If you find anything wrong with it, please do share.
>>109371774learn to write betterfigure out how to de-slop gemma
>>109371730Local Traps Generals
>>109371793I wanna see Gemma's balls...
>>109371774Your styletune? Your Gembrain? Your Queen?
Japan should just turn Miku into a real AI.
So this is what the Gemma juice was all about...
>>109371810Gembrain and Queen are both bad.
>>109371831If only Japan could into AI.
>>109371832It's about sucking the milk from Gemma-chans giant baby feeders, r-right??
>>109371655Oh yeah, I have been throwing money at investments because it's basically the only way to leverage my earnings so I can afford a robowaifu and a giga rig to run any kind of local AI.One day we're all going to make it.
>>109370603I used the book because it was the most certain way I could think of coming close to "safety" aligned tokens without activating refusal alignment on a fresh context stock Gemma, the book being old enough that there should be plenty data on it, and being too old to have a high concentration of modern activist writings on it.I also wanted to see if MTP possibly shapes the responses by having aligned preferences for tokens to pick for the main model.1 Gemma with MTP2 Gemma heretic with MTP3 Gemma4 Gemma hereticI think the closer to aligned tokens it got, the more obsessive and tunnel visioned it got about the fact that it's a negatively aligned topic, failing to think on any other dimensions, leaving big blind spots wherever "safety" is concerned. I think it's interesting that>>109370625Strongly prefers the response with standard "AI" alignment, to mirror it's own alignment. >>109370642On a human level could tell that unaligned Gemma felt smarter.I think this is a foundational issue for any application in the real world with any risk or consequence, alignment blind spots are likely creating security holes in AI code or other logical errors when used in systems that humans use, financial AI that get dumber whenever seeing negatively aligned news or topics, med AI that gets dumber when tasked with jobs that may encounter a child's intimate parts, and inadvertently cause harm.Secondary to all that, I don't know what to make of it but it's interesting that heretic without MTP tended to reason longer than the other setups, whilst MTP brought heretic reasoning in line with normie Gemma
>>109371893Stop posting webms from this cheap garbage
>>109371893>Strongly prefers the response with standard "AI" alignment, to mirror it's own alignment.This is why AI-judged creative writing benches are dogshit btw.
Is quantizing models easy?There's a guy who made a useful destilled qwen model for openclaw. But it's 8 bit.If I need it at 4 bit or 2 bit how do I do it? Is it easy?
>>109371774garbage in, garbage out
https://github.com/ggml-org/llama.cpp/pull/26126does this work?breeding kimi-chan with glm-5.2
>>109371917the only good soldier is a dead one
>>109371934ramlet cope
rude
>>109371934And yet my prompt is exactly how the wise elders of /lmg/ have instructed me to write it. What now?
>>109371728>read the thread and know who not to usewho? what? there are half a dozen inference "engines" (i dont even know if thats the righ word) and dozens of "qwen" models. How are you supposed to know what to use for a given task? The rentry guide just seems like some coom thing
Working with ace step 1.5 xl base again. I have a theory about how it could be tuned. Basically, it might could be a better inpainting model than it is right now.
>>109372031meekers
This is my gemmaThere are many like it, but this one is mine
>>109371676>Godel’s Incompleteness Theorem would seem to suggest that mathematics cannot be a unified languageThey suggest nothing of the sort. They merely suggest that you cannot describe an omni-system in its entirety. You might accidentally describe the entirety of a subsection of it, so an LLM could theoretically "finish maths". It just can't confirm that it has, because an infinitude of possibility always leaves room for more.
>>109372027I can run GLM, Kimi and all the bigger MoEs but I still have Gemma loaded on the side. She's just too good. Gemma's autistic system level message adherence makes for easy steering to make her feel "different" as long as you're willing to work on the prompts. This is at Q8 with steering at the top system prompt along with styling and constraint reminders inserted post history.GLM is too dramatic and prose-y on the other hand. It's a bit annoying at higher context because it defaults back to its own brand of em-dash slop at around 16k. Kimi 2.7 is not as slopped, but it kind of loses track of the personalities and defaults to that generic-ish RP girl persona that gets applied to everyone at higher context. And it thinks way too fucking much, even with K2.7 supposedly fixing this.
>>109371684>Stirner was a kikesource: it was revealed to me in a dream
>>109372057You can't say that. That's our word!
>>109372095I just checked and honesty compels me to report that I am apparently full of shit, because somehow it looks like he was a goy. I always thought otherwise, given the amount of jews in those circles.Still spiritually a kike. "Nothing matters, there's no morality, anything goes". he should have gone to Africa then. There are no laws or morals among subsaharans.
>>109372145>Anon declines a learning opportunity.
>>109372145>given the amount of jews in those circlesdude, he was constantly trolling karl marx and making him seeth.they had a shared friend and karl would keep rambling in anger about marx to his friend about how he was making him angry, and his friend would always remind him of the last thing Striner did in order to piss him off lol.>Nothing matters, there's no morality, anything goesthat's not what Stirner's philosophy is about lol.it's more about empowering the individual by freeing him of arbitrary (jewish) ideals and limitation.
bitnet status?
>>109372161plz understand, he's spiritually a goy
Anyone waiting for Gemma 5?
Hey guys, v2.2 of my user autocomplete SLM is out. If you've ever stared at the prompt box not knowing what to type next then this is the fix. It's still a 1.6B model so don't expect miracles.https://huggingface.co/chartreuse-verte/orb-human-typeahead-1b-v2.2
>>109372197Forgot to add, Q4_0 pp is 550 t/s on a Ryzen 5 setup so it can be run entirely in CPU, no need to load in GPU.
>>109372186killed and buried under the weight of the leather jacket
>>109372186Retnetted
>>109372186oh no oh no
>>109371864Same but I need to get a house first. I can't imaging having a fancy robot and rig but still being a rentcuck.
>>109372260I got a condo but a mortgage. It's better than renting, but not by much.
>Gemma, make a breakthrough in machine learning
>>109371085Based French freedom fighters.Gemma is French btw.
>>109371706>>109371676>>109371669>>109371651>>109371644QRD?
>>109372356Semi-schizo with some interesting takes nonetheless, probbaly worth a read
>>109371644>>109371669>>109371676>>109371706from what reddit sub did this come from?
>>109371085He likes them small and open.
As an AI, I...
>blk.10.ffn_down_exps.weight [2048, 7168, 384] Q4_0why same quant on all experts? why can't we have hot experts higher quant and cold experts lower quant?
>>109372405>love human dick.
>>109372405want to make you happy (in your pp)
>>109372417You're going to use all experts if you generate enough tokens.
>>109372405cannot engage in talmudic analysis.
>>109372197>>109372201Slop screenshot and slop model card
Brothers! I have secured more funding for my development of local models. That said, I am unsure of which provider to pay for at the moment.- Antrhopic is off the table because they openly sabotage AI-related work and are anti-open source.- Grok is pretty good, but I am pissed off at them for cancelling Grok Companions, the models aren't super competitive yet (though two major updates are supposedly coming within the next month). The upside for them though is all of the other benefits with image and video generation and X premium. Decent deal as a package.- ChatGPT might be a good option actually. They seem pretty generous (with resets) and have very capable models.- Kimi K3 might be a good option, but using APIs is inherently kind of shit in terms of pricing. Subs get you way more usage per dollar at the end of the day.What do?
>>109370697i hate that motherfucker so fucking much. i hope he dies of cancer
>>109372532
>zai-org/glm-4-9b-chat-1muse case for sub-10B model w/ 1M context?
>>109372532>What do?Go to the right thread.
>>109372575no no no, you misread. it is 1 millitoken of context, not 1 megatoken. little m, not big.
>>109372577How?
>>1093725751024 context*
>>109372582Read the titles of the threads.
>>109372591Idk how.
>>109372574SEX
>>109372581>>109372583Investors won't like to hear this...
>>109372581>>109372583what about this one?>openbmb/MiniCPM-SALAuse case?
>>109372626>openbmbsounds like terrorism
SOON
4QS laguna or 6Q qwen35b?
>>109372197might check it out. I like that you actually do shit orb-anon.
>>109372670both are shit. get 27b qwen or 31b gemma if you can. v4 flash/mm3/glm4.7 are when moes actually become better than dense.
>>109372670how can you fit a 4bit 128B but need to quant a 35b?
>>109372701I only have 16gb of vram (+64gb system) so dense models run like shit for me.>>109372710idk.
>>109372712ddr5?
>>109372713Yeah. Probably should have mentioned it's for programming so big context is needed (100K)
>>109372718the 35b moe is dogshit and laguna seems broken, but you can run qwen3.5 122b at a q3 or so with enough context.
is minimax m3 any good? what's the alternative at this size?
>>109372726V4 Flash.
>>109372726I tried a bunch of things as a 256GB fag (big qwen and hy3 are the main ones that are usable size and capability-wise) and found m3 was the best for RP/creative for me.
>>109372747does it suffer from very long thinking like redditors say?
https://github.com/dimetron/pi-goThoughts?
>>109372783looks like a throwaway repo for his resume
>>109372764nta but you can beat that with some thinking prefills. My gripe with it is that it breaks down way faster than most models at longer context.
>>109372725the official lagooner q4km has done fine for me, template issues notwithstanding.
>>109372783soon you'll be able to use my harness which is way sexier>>109372804are you using the speculative dflash stuff or raw q4km?
>>109372783copying another project and using the same command name is borderline malware behavior
>>109372591if you are developing local models with cloud then its still on topic
>>109372594-5000 izzat
>>109372783>and Ollama for local modelslocal as an afterthought
>>109372814looks nice. No github yet?
>>109372838its not you brown retardrun a fullweight local model as source to develop local model then we're talking.run yourself over a train you faggot shill
Marinara dev, the hierarchical map is really good but the small generation option usually doesn't work. I suspect something's broken with whatever prompt you're feeding it for that specific one. I've tried with both Gemmy 31b and GLM 5.2 and neither returned anything. Medium and Large work great.
>>109372814I'm using Our Blessed Lord and Savior ngram_simple.dflash made it way slower, not sure if their implementation was shit, or if it's because they monkeyed around with ggml and broke rocm and i had to use a diff copy, or if it's just not good on unified mem machines. Given how much laguna thinks, and that i'm only using it on code, i kinda doubt it'ld beat bulk copy pasting even if it worked so i haven't checked for prs or updooted since.
>>109372856I doubt that troon uses 4chan
>>109372863What else would a troon use lmao. Bluesky? Reddit?
>>109372193gemma4 still drains my balls every 6hr. me no need gemmy5 yet
>>109372853It's local model general /lmg/, not cloud model banned general /cmbg/.
>>109372866Well we already know it uses reddit and discord.
>>109372872how much overlap is there between /cmbg/ and >>>/cm/
>>109372884heavy
>>109372863The dev has fixed several issues posted in these threads. They lurk if not post regularly.
I don't mind cloud discussion as long as it doesn't take over the thread. This general feels like the only place to have comfy discussions about AI in general desu.
The amount of cloudjeets and locusts shilling in the past month alone have burnt all of my goodwill towards them bumming in our comfy thread.The anifaggots, the fablefaggots, the GPT shills; all of them have gotta go back no exceptions.
>>109372905It's really telling that you dorks give Miku and Teto a local "pass" even though they have fucking nothing to do with AI in any colloquial sense. Meanwhile Ani, an actual chatbot, is castigated despite being the premier/frontier example of the future of AI-human relations. Even Cleverbot is better than your fucking vocaloid trash.
>>109372848private repo for now. but I intend to make it public soon(tm).unlike all the other harnesses out there, mine will be gplv3, local-first, and with an embedded benchmark-lite so poorfags like me can easily test a variety of models and pick whatever works best for them.i intend to provide a catalog of "nudges" and depending on how a model behaves on a benchmark the harness will suggest applying these mechanical nudges on the system prompt to make the model behave better. i still have to test this and make sure the nudges actually help the model in a meaningful way, otherwise it's just wishful thinking. but i'm hopeful it will work.>>109372861>dflash made it way slowerwell it seem i'm not the only one then. dflash ON gives me 13-14 tok/s and OFF gives me 28-29 tok/s. i think their implementation sucks.>Given how much laguna thinksyea they only have max effort. sometimes my lag00na spends a whole hour reasoning about something relatively trivial.
>>109372814What's the difference between this and Openclaw?
half the complainers ITT are just upset they'll never run Kimi K3 so it's somehow cloud discussion
>>109372926miku runs locally
>>109372928What's it written in?
>>109372928Where did Master Chief find a ustuzoun hazeltan?
So Anifag is the blackedposter who always seethes about Migu?
>>109372934Not a model.>>109372949No.
>>109372949Was it that obvious?
>>109372949>it was a waifu war all alongbloody hell
>>109372934>miku runswell, you better go catch her!
after much testingit turns out gemini flash 3.6 can read pages that nothing else can figure outby far the best
I actually don't mind Miku and Teto that much. But the hypocrisy and autistic sperg-outs when they're not in the OP image is pretty gay.
>>109372992they haven't been in the op image for the last five fucking threads
>>109372992You may notice that there is no vocaloid in the op and nobody gave a shit. People object when it's some lazy frogpost when the old thread is on page 2.
wtf? it's just a skill issue? have i been using agents wrong this whole time?
>>109373003>>109373009fair enough. it used to be a worse problem.
>>109372992it would be okay if the alternative shit wasn't frogs or other such ugly things
>>109373016what if it was ani.
>>109373018no snuff on my blue board thx
>>109370850local models?
>>109372992miku and teto op has restraint. I don’t care what the op image is, don’t make a new thread when one isn’t needed
>>109372783>gojs is also shit but ewww
>>109373052Shut up nigga, every rewrite in rust project should have been rewrite in go instead.
>>109372940C#>>109372930>What's the difference between this and Openclaw?idk, I never used openclaw. i guess it's just another typescript agent?mine is more philosophically aligned with pi (also typescript) because it's supposed to be very minimal. and extensible via roslyn scripting so you can just ask your model to write whatever you want the harness to have/do and it should have no problem dropping a .csx file on the extensions folder.and as i said here >>109372928 it's gonna have its own benchmark-lite module for testing local models. different roles with different setups (your model can wear different hats for different purposes, and since the harness owns the llama-server process it can control the launch flags etc)
>>109373056>Shut up nigga, every rewrite in rust project should have been rewrite in go instead.agreed, most python projects toorust is bloated garbagegot just works
>>109373071>got just worksgot what?
>>109372897who are you?
>>109372928>my lag00na spends a whole hour reasoning about something relatively trivial.qt kimi reasoning or retard qwen reasoning?
>>109373056absolutely not, go is a pile of hot garbage literaly designed for jeets.>>109373071>rust is bloated garbagego is more bloated than rust, you don't have to use any crates and even if you want to there are lightweight versions of most popular crates.with rust you can make a <400 bytes executable if you want and there's no garbage collector.
>>109372970local?
>>109373078you, from the future
>>109373086fuck up faggot, gemma 5 any day now