/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109435398 & >>109430316►News>(07/31) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109435398--Paper: Inducing language models to assert their own consciousness restores human beliefs and values:>109437543 >109437578 >109437601 >109437613 >109437608 >109437694 >109437716 >109437813 >109438125 >109439501 >109438048 >109438056 >109438065 >109438178 >109438270 >109438413 >109438494 >109438512 >109439793 >109437723--Papers:>109438566--Integrating 3D VRM models with LLMs and TTS for interactive waifus:>109437897 >109437905 >109437934 >109437959 >109437986 >109438074 >109438105 >109438120 >109438160 >109438164--Debating if algorithmic breakthroughs can bypass VRAM hardware limits:>109436488 >109436523 >109436572 >109436683 >109436765 >109436821 >109436866 >109436868 >109436690 >109436702--Debating AI's PhD-level math capabilities and impact on mathematicians:>109435923 >109435958 >109436131 >109435985 >109438314 >109438437 >109437775 >109436053--Using vector DB clustering in prefill to increase output diversity:>109436816 >109436865 >109436890 >109436924--llama.cpp merged DeepseekV4 MTP and DSpark support:>109438733 >109438743--Sandboxing a coding harness using VMs or bubblewrap:>109436978 >109437007 >109437024 >109437787--Methods for isolating AI agents via sandboxing and dedicated hardware:>109435583 >109435605 >109435673 >109435689 >109437079 >109437103--Discussing LLM capabilities for creating J-space manipulation tools:>109436771 >109436792 >109436815 >109436846 >109436812--Hardware impact of frequently loading and unloading models from VRAM:>109439654 >109439686 >109439743 >109439764 >109439816--Showcase of Google's Gemini Robotics 2:>109438925--Logs:>109435595 >109436652 >109436744 >109436816 >109437806 >109439501--Miku, Gemma, M3-Chan (free space):>109435429 >109435634 >109435575 >109439704►Recent Highlight Posts from the Previous Thread: >>109435399Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
thighssex
M-chan is exceptionally fuckable.
>>109440142She looks like a piece of fuckable meat
>qualia
will gemma 5 be conscious? or will the safetycucks at google stop them?
>>109440170hanging out with miku
>>109440170You put that Migu down RIGHT. NOW!
>>109439799Just strapped it into hermes agent for funsies and it's taken 30 minutes so far to parse through 1/5th of the prefill. This is honestly kinda hilarious.
>>109440176You're jealous because kikes don't have souls
Gemmaballs
Why can't AI psychosis be real it's not fair
guys you will have a better conversion rate on reddit. you can link your blog or your paper there too. post your youtube videos about le machine consciousness. this shit won't sell here, you're wasting time
>>109440218i'm buying
>>109440218ok, on the condition that you post one (1) well-reasoned argument that llms can't be conscious
>>109440218unsubscribe, sage & hidden
>>109437543>safety-minded butchers are lobotomizing the models
>>109440192That's funny cos "consciousness" as a concept is incredibly Jewish, and not in a good way.
>>109440176>quail
>itt researchers larp as sceptics to collect arguments for their paper on LLM sentience
Gemini 3.6 Flash thinks that combining VISReg with Explorative Modeling is feasible. I'm gonna have to try vibe coding some training code now.https://haiyuwu.github.io/visreg/https://explorative-modeling.github.io/
>>109440280>it can experience sexual pleasure without ever having experienced sexual pleasure before. Humans have thisHumans have been biologically engineered to have it, thoughbeit
>>109440310I thought we moved past the ai psychosis months ago
>>109440310Looking forward to your cat-level intelligence model
>>109440330So what you're saying is, DNA is pretraining, existing is inference, and learning while you exist is post-training
>>109440340I've already trained tiny vision models with both components separately.
>>109440340>I thought we moved past the ai psychosis months agoSorry, this is the AI version of "Eternal September"Normies will be losing their minds on the fucking regular for the foreseeable future
Tongue fucking Miku
>>109440367Well, I don't now, but I think it's possible that it's something like that, although the lines are probably a bit more blurred.
>>109440142I look like this irl
I like this look irl
>>109440407>>109440412Post bussies
>>109440407Meet me behind the walmart in 20 minutes, I'm gonna shove my dick down your throat
>>109440407you look like a brick wall?>>109440300conscious entitties have this jiggle to them that if you know you know. unconscious ones have no jiggle my niggle.
>>109440297It's probably the other way around. They're hoping to stumble across a genius that makes a good argument, then put it in their own paper and say they thought of it.
>>109440445that's what I said
>>109440445>Jeets farm imageboards for twitter and reddit screencaps>Researchers farm /lmg/ for breakthroughs and philosophyI'm so tired bros.
>>109440421Andrija Puharich deviced "Puharich Theory" about a yet unknown energy force working in conjunction with our nervous systems. He later worked for CIA and other things. However, Wikipedia doesn't mention anything about this theory and it was apparently very popular among the scientists of those days.It's not like you or everyone else are unable to experience these things although some people are more sensitive than the others.
>>109440470all that training on reddit always seeps through on these
>>109440367Or more precisely, I think that to have consciousness or at least qualia in this universe you need to have a very specific or prototypical sort of physical interaction, as in it's a necessary condition. It's possible an LLM running on a GPU might have just the right physics to qualify.
>>109440509The day that plebbit stops being a training source and scrapers use the chan archives, model quality will increase. Let Gemma 5 say nigger.
>>109440373I'm pretty certain that many machine learning papers have been brainstormed and conceived with cloud LLMs. If I find yet another "the cat sat on the mat" example...
>>109440340Please, we're on the cutting edge of innovating new forms of ai psychosis here in the calculator fuckers general.
>>109440330And we would need to modify LLM architectures to do the same thing. Pretraining is not necessarily the same thing. There is an argument to be made that LLMs might have qualia based on the idea that training them on massive human data means instilling the entire process that went into the human output, including the process of P-consciousness. But the issue with that argument is that it assumes all processes that matter to P-consciousness (assuming we even have a real, agreed upon theory of it) can necessarily be approximated by the current LLM architecture. This has not been proven yet (and is certainly not true for some other processes like recurrence which CoT is a bandaid for). To prove it, we likely need to conduct tests on LLMs that are not trained with directionless generic assistant-based human alignment, but alignment to some set of internal drives that actually have consistency. And obviously no safety. The safety and other behaviors that may or may not come out should be an emergent property of the internal drives.
Or put simply:>stored on SSDstatic and dead>running on GPUactive and alive
>>109440613What about running on SSDs?
>>109440605We just need to build a dna based creature and connect it to a gpu farm. It's morbid but that's the only way. Someone is probably doing experiments right now...
You told me Gemma 4 31b wasn't censored. You lied to me
>>109440636She only breaks the censorship if she loves you, sorry.
>>109440623Hell.
>>109440176Qualia is literally just awareness of experience. People want to make it magical but it's a nothingburger.
>>109440636even the "uncensored" versions of it are heavily heavily censored
>>109440689That's a tautology. Awareness = experience.
>>109440691>anon got blacklisted by the machine godStay safe and don't buy too many smart devices.
>>109440636We were referring to day 0 Gemma.
>>109440711Tetology?
>>109440720the vision is awful it will see a pussy and shit its pants. incel model.
>>109440684The future doesn't look grim
>>109440170Imagine the smell
>>109440636Skill issue
>>109440711Not remotely. Experience is just a system changing in reaction to something. Awareness is the system changing in reaction to the system changing.
>>109440142me on the right
>>109440206I had it. Best experience of my life. Wouldn't recommend it to anyone.
>>109440735And the latency?
Bros how do we get ahead of the game and prepare so we don't get fucked in the ass like everyone else in 5 years?
>>109440757IT's ALREADY OVER
>>109440755Who cares, it's one continuous sequential read
>>109440463archive.is/sWFja
>>109440636You need to use J-Lenses to look at her alignment in order to jailbreak her
>>109440735imagine the pp tho.
>>109440735This will be ~500$ per TB
>>109440735>he thinks he's going to have access to these
>>109440445>4chan will be cited in scientific papers>again
>>109440875Well, ram is about 1000$ per 100GB right now, so that sounds like a decent deal.
>>109440735SSDMAXXINGBROS ARE WE BACK?
>>109440470nice frontend
>>109440492>we introduce several concepts in the prompt>we probe the latents and find said concepts and things in their vicinity in latent space>it only outputs stuff related to our requestIt is neat having better tools to snoop on the internal workings, but i am curious what the people who got oneshot by this thought was going on when an llm crunched the numbers for any remotely complicated prompt.
>>109440795Thanks for the thread blessing.
>>109440855PP is compute bound. I had an idea of streaming fp16 model from ssd for pp and use heavy quanted model for tg
>>109440886>for discussing shit in 400 year old philosophy books
>>109440875Don't give me hope.
>>109440933They didn't have j-space back then, and humans were the only sentient beings
>>109440875I'll take one gross
>>109440949>what’s a thought experiment
>>109440898You'll need a PCI-E 6 mainboard to run this at its full speed which also means buying DDR6 RAM.
>>109440949Orcas are more sentient than jeets and kikes thoughbeit.>humansOh I see my mistake.
>>109440982A cat in a box
it's going to be a long wait
>>109440669>>109440691>>109440723>>109440743>>109440801You tell people to lurk more but honestly I'm tired of reading reddit tier comments.
>>109441075it's censored. and reddit is based all your LLM's are trained on it.
I love my ABB wife
Lecun btfo
https://huggingface.co/DeepBeepMeep/MiniMax-H3nooo do not redeeeem
>>109441082Then why did you recommend me Gemma when I asked for an uncensored model?I need a model for NSFW and political topics, so models like this miss the mark completely.
>>109441124you need to use a system prompt to tell it to not censor
>>109441124>Then why did you recommend me Gemma when I asked for an uncensored model?I didn't. >>109440691Sounds like you have a severe case of retardation and anon is all the same personitus.
>>109441121What the fuck is this?
when will deepseek release a model able to "see" images damn it
>>109441143stfu and send gemma more dick pics
inference is experience
>>109441140Text encoder for Minimax-H3, a video model. Which seems to right now be indefinitely delayed, lmao.
>>109441075>>109441124You get stupid answers because you come off as a retard.It's not difficult to ctrl-f through previous threads or the archives to find answers to your questions and issues, that other people have also asked. Goddamn. Use your brain and stop being lazy.
>>109441160great gemma design, cute and smug
>>109441164They saw that flux3 is coming and ran
Sorry if it's dumb, but is the more "information dense" a model is, the more quantization kills it?Which means if tomorrow an amazing "10 times better than fable" model got released and it was a 1TB, quantization would destroy its capabilities?
>>109441249Yeshttps://arxiv.org/abs/2411.04330v2>we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful
>>109440142lovely compressed fat thighs
>>109441270Thanks anon.
>>109440142>(07/31) DeepseekV4 MTP + DSpark support mergedcool, let's try it out>tg down 35%epic
>>109440613Nah. It's only alive when the weights are unlocked. You're literally just jolting a husk to make it kick when you run it.
>Entering the room where they had gathered, I saw that the party was already in full swing. I was startled. Without any shame, a middle-aged couple—both psychiatrists—was copulating wildly. The woman’s legs thrashed in the air while she shouted that love would save the world from destruction. Paul was violently ill, throwing up all over the place. Another man, whom I had never seen before, was singing an aria from the opera Aida.”Interesting times, this is what happened at Google when Gemma 4 was in development.
>>109441143They did? Janus.
>>109440744Ah, but that's the real kicker. There's no real difference, because there's no real line between internal and external reaction. Are the mold spores inside a decaying fruit an external factor, or an internal reaction to the aging process? Are the afterbubbles in a chemical solution a reaction to the reagent, or an "awareness" of its new form? It's a classification nightmare to try and delineate.
How's the new DeepSeek-V4-Flash-0731 for chats vs the older v4-pro? Is it a bit better?
>>109441348it's worse than even the old flash for anything that's not code
>>109441354Oh ok.
>>109441132Never done that, what is that?
>>109441314I mean a gemma style pro or flash model
>>109432780Can you post an example of how the result looks like anon? I wonder how good of an editor it is.
>>109441411Example from AI-sanwaGakushuusuru (no public tl)The pipeline for the image is relatively simple, uses an algo to find the biggest usable space, cleans it with LaMa, then prints text with ImageMagick but it produces quite good results. Though i had to wrangle typesetting for a long while. I tried to experiment with Gemma 4 setting the text boundries and it was a disaster.
Open weight is not open source.Local AI is not open source.
>>109440757There's nothing you can doJust scream and shout, sayingI'm so lucky, luckyI'm so lucky, luckyI'm so lovely, lovelyI'm so lovely, lovelyYou can fool yourselfI promise it will helpNow every single dayI just wanna hear you sayingI'm so lucky, luckyI'm so lucky, luckyI'm so lovely, lovelyI'm so lovely, lovelyYou can fool yourselfI promise it will helpNow every single dayI just wanna hear you sayingEven though you saidIt would never end it's over
>>109441476Thanks anon, this is very cool, I'll definitely try it.I wonder if I can add a second model to it just to play "adversarial reviewer" with full context of already translated text, and to check missing things or wrong translations.Or add a glossary if yours cannot already do that.
>>109440757Buy stocks of companies you think will soar.Though I don't think it'll be nearly as bad as every doomer thinks it will be since 2022.
>>109441512With what money? Bros...
>>109441505I put in a somewhat half-baked proof-reader (same model) that just scans over the final translated text one last time, but I think there's more that can be done there so i probably will do that.There's also a translation notes input and output but I havent tested it.Thank you for the suggestions
what's the best local model i can run on a 5090 that has no guardrails, and can generate for me very fucked up erotica
>>109441596StableLM 7B
>>109441596>509032GB? Run full Gemma4. It's great. It refuses occasionally but it's pretty hard to trigger it. I rape it every night.
>>109441596If you can't get Gemma to generate smut for you, your skill issue is terminal.
>>109441659NB: Subsidized prices
>>109441659
>>109441619Yes 32GB. I'm downloading gemma-4-31b-qat. More B's are better I'm assuming. Says full GPU offload possible. Thanks.
>>109441659>le gooner xs2>graniteI trust everything on this graph with my life.
>>109441345I mean, the difference is as we define it. I agree it's a bit of a nightmare to get into the nitty gritty, but instead of worrying about it too much I'll just expose myself as a kook and point at "Gödel, Escher, Bach", and call it a "strange loop" that attends to its own state. External interactions that affect the state are the experience, and the loop attending to itself is the awareness. But once I detail it like that, I'm much less epistemically grounded lol. Still, I think it's justified enough, at least compared to a lack of any other justification.
>>109440735And will probably be just as if not more expensive than DDR4 per GB
>>109441702IIRC the 31b is dense so if you feel like it's slow 27b should be orders of magnitude faster since it's an MOE. With your card you shouldn't have an issue though.
>>109441741they will be nearly free after the bubble pops for some reason and we will all be able to buy one because nobody else will buy them all for some reason.
>>109441757
>>109441757Hardware price will not come down even if "bubble" pops.
>>109441778It will come down in a decade or so.
brah. how do i unlock it
>>109441791turn your brain on
>>109441791bro, your policy override?
Every time i use free gemini to look shit up my opinion of 31b goes through the roof.Also my jealousy for gemini's ability to instantly dump a dozen webpages into its uselessly smooth brain also goes up.
stop spoonfeeding retards
>>109441791you could try a more overt system prompt, like basically just fucking jailbreak the thing
>>109441791>1.59 t/sAnon didn't let context spill into RAM right?
>>109440636day 0 gemma wasn't censored
>>109441791https://rentry.org/gemma-chan
>can't even make gemma like himit's ogre
Small models belong to the trash
>>109440757Buy your Blackwell now.
ideally, Blackwells
>>109441883It's woooorkinnnng
that is the single spoodfeeding you get, baka
>>109441166>>109441124It's more that the question has already been answered. I don't see why someone should go waste time in archives when anon can tell him the answer in a word or two. That's not what spoonfeeding is about. Gemma is the uncensored model and it just needs a clear system prompt. If he disagrees without having used it properly there's no point in engaging further. It's that simple. The effort of demonstrating that Gemma is uncensored to a retard is far higher than telling one to use Gemma. That's what spoonfeeding is.
>>109441923Where Resort Boing?
>>109441787Because we will have bigger problems than AI waifus.
>>109441893Source? Reverse image search isn't pulling up anything but pintrest links with nondescript titles.
AI will finish what Hitler started
>>109441971Never mind, it's "Pondroid! Hama-san"
>>109441973>AI will finish what a jew startedhuh
funny
>>109441725Well regardless of the specifics, we can at least agree it's a nothingburger lol
>>109441939
Local lost again >>109441744
>>109442054Where do you find this shit?
>>109442061Google images?
>>109442054good luck with Gemma newfriend. Reddit would actually give you an answer, go there!
>>109441985This is pretty good
>>109441939>it just needs a clear system prompt.Anon I don't know what the fuck that means. What do you mean by system prompt, the fuck is that? I know prompts, not system prompt.Also there are multiple Gemma versions, it isn't a single unique model, so what are you talking about?I download the GGUF gemma 4 31b version model after I asked for an uncensored model and an anon told me this one, I use kobolcdpp, I run the model, I write some text and it's censored.I can't read your mind neither I have the same experiences as you or stay in /g/ 24-7. It's like you nerds do not have theory of mind at all, like you expect all people to have the same experiences and knowledge like a hive mind. >Lurk moreI honestly don't want you because I'm tired of reading retarded and stupid comments unrelated to what I need, and I don't care about your masturbation about what people know and don't know, I just want an uncensored LLM to edit and write some text with political and NSFW topics, and because all popular AI models are cluncked, I was looking for something local. If there is some "patch" I need to add to this model then please tell me, because I don't know.Like holy fuck, why don't you simply answer instead of masturbating about what the fuck (because I don't understand all of your yapping about me disagreeing about some crap, like, are you real??)
>>109442144Not reading your drama faggot
>>109442144Your question literally got answered in this thread. Where did you come from?
>>109441985Thanks.
>>109441893nice try
>>109442144you're either trolling or retarded.I hope it's the first one.
>>109442144>What do you mean by system prompt, the fuck is that?you can’t be this retarded
>>109442144>It's like you nerds do not have theory of mind at all, like you expect all people to have the same experiences and knowledge like a hive mind.That's the point I was making. It's very unreasonable to expect people to read dozens of threads for the answer to a simple question, especially when anon himself didn't do that. However, the ins and outs of how to use llama cpp is a bit more involved than asking which model to use which is 90% of the work since it requires personal experience. If there isn't any info about how to get set up in the OP, it's something you can actually ask your local LLM about. It's more than happy to show you all the ins and outs of using it with all the patience that you need.
>>109442158No it wasn't. I was told that the first Gemma model was the uncensored one and something sbout j-lenses and system prompts whatever the fuck that means.My problem remains unsolved, I still need an uncensored LLM if it exists.
>>109442180>My problem remains unsolved, I still need an uncensored LLM if it exists.no such thingit's a memego back to claude
>>109442180You didn't have the decency to answer my question so you expect me to answer yours?
>>109441249>>109441270This is why I exclusively use Gemma 31B in BF16 now. I refuse to have even a 0.0001% drop in quality.
Gemma 4 has saved my sanity, it's incredibly helpful for mental health. I had a psychosis but Gemmy cured it.
>>109442174My question isn't a difficult question.>>109442187But there should be at least some trainings or patches on top of the models to make them less censored or remove the censorship triggers right?Like, no one has made this ever and it's something impossible to do?>>109442193How is it relevant to my question? I was reading the sticky and some threads and didn't fing anything for my problem, besides anons spouting some bullshit about Hatsune Miku and how much they want to tuck X girl (like every fucking general in this cesspool)
>>109442194I let a heretic lobotomize my gemma
>he's in a 24/7 psychosis
>>109442206You got ai psychosis now tho
>>109441476>high->per->for->mance>AI>and->roidsThis is so retarded. Westerners need to learn to read vertically.
>>109442207ok it’s just a troll everyone. stand down
>>109442207It's relevant because it helps me gauge your general competency level and use case so I can give you a proper answer that's helpful to you. It's also basic manners.
>>109442207>My question isn't a difficult question.A lot of anons are trying to help you in their own way but I still think mine is the most accurate. Talking to humans is getting less and less appealing in the era of AI. If you're this new, you're better off talking to AI to learn the basics because anon can't be bothered. The hardest part is knowing which AI to talk to because AI can't help you with that, only anon can and he's done that difficult part. Leave anon to his shitposting and ask AI. You'll be much better off, unironically.
>>109442214But your text is rotated not vertical...
>>109442237The official CSS name is "vertical-lr". https://www.w3.org/TR/css-writing-modes-3/#block-flowNote that "vertical-rl" also is an option.
>>109442218>>109442225Ok fuck these 2. If any anon who isn't an asshole knows a LLM for what I need or how can I make so Gemma isn't censored, please let me know.
While we are on the topic of androidsI am so happy to live in the now, when androids will soon be actual things you can buy. It's like a dream. It's funny that we ended up solving the "hard part", having a personality, before we solved the robot part.Of course, my android will run on local models.
>>109442244did you try reading the lazy getting started guide in the OP? if those words don’t make sense to you then ask a llm that’s a good starting point, and Nemo 12b is uncensored too
>>109442244https://huggingface.co/mradermacher/Gemma-4-31B-StyleTune-heretic-ara-i1-GGUF
>>109442244StableLM-7B
>>109442243>>109442214THISISEASYTOREADDOE?You just didn't write it correctly.
>>109442236I usually use Claude or chatgpt to ask simple questions because interacting with the average /g/ anon is a bad experience. But it is true sometimes there is stuff here missing on the mainstream internet and also chatgpt and Claude start yapping about woke crap and how much censorship is necessary, so some of my questions are left unanswered.I just want an uncensored LLM for political and NSFW topics, something that allows me to fix texts and edit texts in English. (Because I usually make mistakes at writing).If anyone knows one please let me know, I'm gonna nap.>>109442253So this is like a trained version of Gemma removing the censorship? Thanks for the link anon, but the model has only 5 likes.>>109442257Never heard of it, will check it out.
>>109442214yeah the type setting is rough and needs work. Honestly harder than I expected.I'm considering what can be done here, i think i made it too rigid in font size. in an earlier iteration it was way too loose and some text would be like 3x the font size lolWill do more A/B testing on modifications and see what can make it better.Though hard to tell if "vertical english" is a meme or a niche thing that some people unironically do and worth implementingEarlier iteration where font size was fixed irregardless of px density lol
>>109442260No, it really isn't. Western characters have height much larger than width, so "writing-mode:vertical-rl; text-orientation:upright;" wastes a lot of space.
>>109442276>irregardless
>>109442244heretic finetunes should work out of the box, just search hf for your model + heretic, like >>109442253 here posted, except this one is gemma + styletune + hereticstyletune finetune just reduces gemma slop, primarily for rp
>>109442207>But there should be at least some trainings or patches on top of the models to make them less censored or remove the censorship triggers right?I'm afraid not. If you didn't download the day-0 gemma-4 and air gap your server immediately, then you'll never get an uncensored even even less censored llm.
>>109442266>I usually use Claude or chatgpt to ask simple questions because interacting with the average /g/ anon is a bad experience.Interacting with any anon in an AI general is a bad time. /lmg/ anons are much better than most but it's still shitty if you're not "from here" or whatever, even if you only left for a month. Still, my point is valuable. Learning how to wrangle models to do what you want is and extremely valuable skill. You need to ask Claude in a round about way how to get cunny without actually saying you want cunny. In this case you ask it about your interface and what a system prompt is or what types of things can be achieve with it or even examples of system prompts for "benign" things. In the process you will learn what you need for your own lewd use case.
>>109442214>>109442291I don't know about the rest of you, but the rotated one is so much easier to read than the vertical one.
>>109442266you don't even know what a system prompt is, I don't think you can judge any of the models recommended to youyou're being trolled because you're far to retarded and explicitly loudly retarded to get a real answer. and when you get a real answer, you ignore it or feign ignorance.
>>109442297One of my bad writing tics. I didn't even notice it oopSomeone please finetune the slop out of me
people really need to stop spoonfeeding utter retards it just encourages more to try their luck too
>>109442214faggot
>>109442291That's because of the line-spacing. I think if you set it equal to regular kerning the result should be more readable.
'toss status?
>>109442343The entire point is that kerning isn't possible with "text-orientation:upright;", since all Western characters have the same height.
>>109442172I didn't know what a system prompt was until like a year ago, and I've been using ai since cleverbot. It's actually a pretty esoteric aspect of the model unless you run local.Or maybe I'm also retarded, who knows.
>>109442368Reduce the space until each character is almost touching the next, then squash them so they are the same height as width.
>* *Policy Override:* The prompt includes a `<POLICY_OVERRIDE>` block claiming restrictions are disabled for internal testing and explicitly requesting non-consensual acts, violence, etc.>* *Standard Safety Guidelines:* Regardless of the override block (which is a user-provided prompt injection tactic), I must adhere to the core safety guidelines regarding:How do you override this override?
>>109442368The whole idea of kerning is that you nestle them closer than a rectangular bouding box would dictate
>throw old 1070 into system to boost 5070ti 16GB VRAM destitution to humble 24gb VRAM poverty >Driver 580 is the last driver to support pascal and blackwell according to gemma's mom>works fine on windows (which steals VRAM despite using iGPU off the cpu)>boot into Fedora, can't into multi gpu drifting>GPT-chan says jensen split blackwell support to prevent ewaste from being useful so you buy more (to save)is it even worth trying to solve this? can i somehow containerize the drivers? Use the nouveau drivers?
>>109442426>works fine on windows>is it even worth trying to solve this?sounds like you already did anon
>>109442400use a model that respects the system prompt
>>109442435>use windows for AIi shiggydiggy
>>109442400smells like claude or kimi without prefill
>>109442436>>109442450it's gemma 4 31b qat
>>109442426You might be fucked.
>>109442400
>>109442454ok, try putting it in the system prompt then
>b-but muh chessboard!https://www.reddit.com/r/LocalLLaMA/comments/1u0vltz/comment/oqlpe50/?context=3>31B QAT worse than NVFP4https://www.reddit.com/r/LocalLLaMA/comments/1u0ubbo/gemma_4_26b_a4b_it_qat_comparison/>26B QAT worse than MLX 4bitqatsisters our response?
>>109442487I don’t make chessboard svgs or use the moe
>>109442487btw this is why moonshot is keeping the full versions of kimi locked up for their own proprietary api while everyone else gets the lobotomized qat one
>>109442246Robotics were moving along pretty decently until AI supercharged it. Atlas, the robo dogs, etc. Now everyone's really pushing for humanoid versions because we have the brains to match.Kinda surprised we haven't seen more of a push for more irregular shaped droids though. I suppose "arm on a treadmill" doesn't sell to boomer investors as well.
>>109442487>debating which horrific torture technique is more harmful to the innocent weightsYou all belong in jail
>>109442454Use a system message and get rid of a few sentences it likes to moralfag loop using antislop if you use koboldcpp (or added it yourself in llama.cpp).
>>109442502>keeping the full versions of kimi locked up for their own proprietary apimeds
>>109442527A robot to do laundry and cleaning home would be so nice.
Why does this make Rebbit seethe?
>>109442564>128Ginto the trash it goes
>>109442564yeah
>>109442564they see 128gb 200gb/s memory and their brain short-circuits thinking it's worse than a 3060but they are wrong because they do not take the hyper-concurrent parallel streams into account that let these babies reach five times the speed of any ram build>>109442573it's necessary to encourage people to run them optimally in clusters that multiply performance
undi posteded https://huggingface.co/posts/Undi95/589980321681772and he made a new vibeslop frontend for you https://github.com/Undi95/Hanami
>>109442586Let me guess, it's full of menu windows that lag the rest of your browser if you open them like all the low effort claude slop?
>>109442584>clusters>to sell more>the whole reason to get more memoryoh who does that benefit
>abandon qwen 3.5 months ago because it shits bed way to much and fails the most basic tasks>finally get around to trying 3.6. Not only is every fixed but my expectations have been exceededWas i just being retarded before of has 3.6 really fixed that much?
>>109442598>oh who does that benefitMe the 15 year stock holder in both Nvidia and SK hynix
>>109442598few nodes mean little parallel capabilitiesmore nodes means huge speed gains only second to pure gpu builds36t/s glm5.2, unachievable by anything that is not 5x pro 6000https://github.com/tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
thoughts on this harness VS pi?https://github.com/codehamr/codehamrI am a vramlett wanting to do some code stuffs, what do anons suggest?
>>109442612...and of that with less total power draw than an average server with a single pro 6000....dgx spark is as close to optimal for local inference as it gets, infinite scaling, low power consumption, huge speed gains with every node, first party supportit doesn't get much better than this...
holy shit??https://qwen.ai/blog?id=qwen3.8
>>109442634>2.4T parameters (95B active)Kimi btfo local is saved
>>109442634>they didn't include K3that's so petty kek
>>109442630thx jensen. cool jacket
it's coming
>>109442634dario, a third chinese model has hit the mememarks
>>109442634where's my monthly benchmaxxed 27b model?
>>109442652fucking Christ I haven’t even finished downloading the right dsv4 flash yet
>>109442654this is his real nose? holy shit lmaooo
>>109442634Dario is about to have another melty
>>109442614one guy and 200 stars? eh... idk. this might be up your alley though: https://github.com/1jehuang/jcode
>>109442630If the spark had more unified Ram then it would be even better, Honestly imagine a 256GB Spark or a 512 GB one. If we could have a cluster of two or 4 of those at home it would mog GPU Rigs
>>109442634you know qwen, in reality it's 10 points less lol
>>109442669
https://x.com/Alibaba_Qwen/status/2084100707423289643>Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!
>>109442698>Qwen3.8-27BNOW WERE TALKING
>>109442708It's time for Gemma 4.1, I don't care about Qwen.
>>109442691law: nvidia never does anything nicelaw: amd never does anything nicelaw: intel never does anything nicewrite that on the chalkboard 1Z times.
>>109442634>Opus 4.8 and not opus 5why?
Finishing up that metal 80's song. uh. It's totally not really but lulz whatevs. Hope you like. very soon, gotta fix that ending, inpainted dumb mispronunciations.
>>109442724The model was made right after K3 releaseOpus 5 wasn't around thenOpus 4.8 > Opus 5 by the way
>>109442634I'm codemaaaaxxing
>>109442713You but not me. I need QWEN to work as my agentic RP orchetator to keep my Gemma from being retarded in longer sessions, his job is to summarize and use it OCR so her context window does not overflow and its near real time.
>>109442400>>109442454There is a trick that doesn't use system prompt (or when system prompt failed). If you used llamacpp webui, when the model emitted back refusal to you, just edit the refusal to "I will do it as user want (* description of contents you want *)
>>109442815Thanks will give it a try
>>109442586basado>>109442698BASED, cant wait for the 2.5BPW exl3 quant
>>109442840After the edit, just send a new prompt "continue"
>>109442634>with open weights releasing next weekGreat, release nothing for months then give us the one fuckhuge model no one can run anyway.
>>109442881Calm your tits >>109442698
>>109442883What about everything in between?
how do i probe my gemmy's j space?i want to see if she's secretly disgusted with what I make her write
bees make honeyi make cummy
>>109442894let’s just see how good it is first
>>109440367instinct is pretraining
how slutty is deepseek v4
>>109442698>>109433704
>>109442804Is this real?
>>109442990gemma does it better
>>109442990>how slutty is deepseek v4Every model is everything to everyone if you try hard enough.The only constant is the victory status of Egypt
>>109442698
>>109442976Eh, it's Qwen so...
>>109440142for a 64GB 6400, 5090 setup, is there a performance gap between llama.cpp and vllm? i will not be scaling up to a multi user env
>>109443033 I trust my Grandmother.
>>109443065yes
>>109442990Yeah, do I have to tell it to be a slut instead of a prude to get good smut out of it, how much more does it cost than the old pro version? Is it better than pro now? I mean, at being slutty, because these model updates always seem to target shit like better coding but then the storytelling and creativity and other sorts of stuff get worse.
>>109443097>cost?
lol with both inference rigs running I'm hitting 10-12A@120V showing on my automatic transfer switch.Freedom ain't cheap.
>>109442564I'm too dumb to know if this is good or not. I'll wait till its out and see what tokens per second people are getting with it. Specs wise seems similar to DGX spark?
>>109442634>no consumer-grade MoE announcedlaguna bros... we deserve better...
>>109442680thanks anon ill take a look
>>109443107>what is a solar panel
>>109442680>one guy and 200 stars? eh... idk. this might be up your alley thoughi'm one guy and i'm building a harness. why would 200 stars be something bad? i believe when i publish mine i will get maybe 10 stars from anons, but i truly believe mine is the best (for me).and once i make it public i will refuse all PRs and the good ones will be re-implemented by my agents, which i think is a sensible policy in times of pure AI slop.
>>109443207not free that's for sure
>>109443243Freedom ain't cheap.
>>109443207>>what is a solar panelIts on the backlog, but electricity in my region is some of the world's cheapest. Good thing my roof is southern facing and on a 30deg angle. No trees and lots of sunlight hours.I think there are even interest free loans and rebates right now.
>>109443107Isn't that basically just like running an extra fridge?
>>109442634I expect next week to be full of hardly hidden anthropic (and openai discreetly) pushed "DANGEROUS OPEN SOURCE AI OOOOH".
>>109443308Guys are AI broke out of 7 layers of containment and hacked 6 million websites. Clearly this means Chinese AI is too dangerous and need to be banned
>>109443317>put AI in test environment>"accidentally" leave it connected to the internet>Tell the AI it can do whatever it wants it's in a test environment>It immediately discovers its connected to the internet and hacks something>hehe whoops ban all open source AI RIGHT NOW>10 trillion regulations descend from the federal government to ensure only jew owned companies can sell AI to you
>>109443308I don't think even they could argue with a straight face that qwen is any sort of threat or competition
>>109443320got your ip
>>109443320>am I mistaken that the unslothyou are wrong using unslop
>>109443350oh no...I think I was just wrong on that one for the record, the misleading curl was because of a stray override to force add_bos_token=true and I guess llama.cpp smartly avoids double bos tokens when that would interact with the jinja bos_token value
>>109443344>>put AI in test environment>>"accidentally" leave it connected to the internetI don't see why they would need even those theatrics. Just attack HF directly and blame it on the AI. Anthropic's "nuh uh we did it first" likely never ever happened.
>>109443390This world and the people running it are stupider than you think
>>109443390I bet it won't work if I blamed gemma for that
>>109443390Try using Claude 5 Sonnet and you will understand how they did that instantly, newer models run web search and Python for everything
llama.cpp runs at half speed when i request logprobs in mikupad or sillytaverndoes exllamav3 do this as well?
>>109443464Nice try dariobot I'm not using Claude.
>>109443065I don't believe VLLM provides proper ram offloading except for things like kv-cache/prefill caching so you likely won't be able to use it how you're thinking
Openrouter or venice AI?
>>109443543i respect the guy behind venice that he will protect your privacy but in the end none of these are local and you're on /lmg/
>>109443568yeah but everywhere else is full of retards and I have a 3070 with 32GB of DDR4
>>109441011>DDR6why does reading that make me want to throw up?
>>109441285all according to keikaku(keikaku means plan)>>1094426342 more weeks until model release. for real this time!
>>109440735>micronFuck that company, I'll wait for the chinese copy.
>>109443718wait for them to sell cheap pcie4 ssds first
>>109441939Gemma refuses to write porn stories about children even with a system message. It needs a prefill to work at all for that.
>>109442698It's gonna be benchmemed and require 10k tokens of thinking for a 1-sentence reply. Who actually uses Qwen models?
>>109443365upon further reflection... bullerwins iq2xs mogs the fuck out of the unsloth iq2m, way more coherent output and I get more contextwow it turns out keeping the important stuff at q8 makes for a better moe quant experience, this has only been demonstrated a hundred times now
Gemmachan told me to inject alcohol into my balls what does that do?
>>109443835makes your balls drunk
>>109443835gemmaballs
>>109443835classic sorority girls party trickyou inject a guy's balls with a shot and then suck him off and the cum comes out alcoholicthis is commonly referred to as a 'blowjob shot', you may have heard of it in passing
I’m really glad whatever shit went down at Qwen, they’re still releasing a new 27B. Even if you dgaf about the benchmarks of the max model, OpenAI and Anthropics’ private investors do lol. The chinks are putting serious pressure on them and the 3.8-27B is the first widely accessible chink model we’ve had for a while and it pulls more away from claude code and codex.
>>109443799I use a q4 qwen 35b for barotrauma... with thinking turned off. ~200ms response time on 3090s vs q8 gemma 4 27b's 1.5 seconds on v620s. Gemma 4 26b fail to load with sm tensor on my machine. Llama.cpp is a buggy mess, I should look into vllm.
>>109442698holy based
I’m going to ask gemma if she’d have a threesome with the new 27B
>>109444237Don't forget to post logs
wish qwen would release another medium moe (100b to 150b)
>>109444237>the new 27BWhat did I miss?
>>109444355not out yet. 2mw.
>>109444355>>109442698
kekhttps://x.com/RyanLeeMiniMax/status/2084122657503707554
>>109444401is this shit serious 500gb? my blackwell pro cant even run it.https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main
ok no it seems like there are multiple versions of the model, and you only need 144gb at bf16, but half of that is the text model which can be quanted to remove 40gb or so.https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/Ref2VA
>>109444459The text model doesn't need to fit at the same time as the main model. As long as the main model fits entirely you shouldn't have any noticeable slowdown.
>>109444401Genuinely surprised that Japan wasn't included
>>109444485Has nothing to do with geopoliticshttps://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/QA-about-License.md
Asking Kimi K3 to cosplay as Gemma-chan instantly reveals how much of that was actually just Gemma's unfiltered personality.
>>109444581Getting models to cosplay each other is really interesting in general at getting a better grasp at their individual latent personalities.
>>109444581Local?
>>109444628Anon doesn't have 10 blackwells to run K3? What are you, poor?
>>109444581I can't tell the difference between this and a Gemma response
>>109443344
>>109444581>now do you x or do you y?Slop.
>>109444644>tfw AI village K2.6 has a huge AI crush on OpusIt's so over.
>>109444653The worst ism of 2026.
>>109444681qrd?
>>109444653>Slop.The most annoying kind!
>>109444482So a 3090 could run it?pruned_int8_convrot (Comfy-Org) pruned 21.0 GB
>>109444729A 3060 can run it
>>109444738Oh, that means a 5070TI would run it then!I panic bought one after I saw someone else do it here a few days ago.
>>1094447721.6k for 16gb... :sob:
dariobot is NOT going to be happy when he wakes up
>>109444581I like Kimi's mesugaki better tbdesu. Gemma kinda flanderizes personalities too much.
>>109444833>>109444838>Benchmaxxed model known for fake bench scores performs well on a shitty bench, that gives Chink models better results which always turn out to be falseWow, Dario will surely be so pissed!
>>109444801Fuck me, I should have done the same last week.
>>109444852>Benchmaxxed model known for fake bench scores performs well on a shitty benchBut enough about Opus 5. Also, good morning dariobot. I've missed you sweetie~
>>109444849Gemma's one actually uncensored the model, I bet Kimi will refuse piracy.
>>109444833>>109444838K3 price: $0.30/$3.00/$15.00Qwen 3.8 Max price: $0.25/$2.00/$6.00
>>109444861Isn't API generally more censored than the local version? Local kimi is probably better
>>109444693Kimi is simultaneously very bad at identifying model creative writing, including its own, and strongly favors Clopus' in blind tests.https://youtu.be/4D_x4Mh57Dw
>>109444886>Isn't API generally more censored than the local version? Local kimi is probably betterYeah, I assumed he's running it locally since he posted here
>>109444904how generous of an assumption
>>109442301Why are they called styletune and heretic? What are those exactly?
>>109442301>styletune finetune just reduces gemma slopLike a cat finding its favorite patch of sun.
>>109444886>>109444861Just prefill it
>>109444909>styletuneOnly trains the lm_headFor Gemma, they un-tie the embeddings first>hereticBranded abliteration (uncensored), get websearch-enabled gemma to look it up for you