[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 00012-1258326611.png (1.39 MB, 1216x832)
1.39 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109435398 & >>109430316

►News
>(07/31) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse
>(07/31) DeepSeek-V4-Flash-0731 released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
>(07/31) K-EXAONE-2.0-750B-A37B released: https://hf.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B
>(07/30) Inkling-Small released: https://huggingface.co/thinkingmachines/Inkling-Small

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: hpeえむちゃん.png (928 KB, 832x1216)
928 KB PNG
►Recent Highlights from the Previous Thread: >>109435398

--Paper: Inducing language models to assert their own consciousness restores human beliefs and values:
>109437543 >109437578 >109437601 >109437613 >109437608 >109437694 >109437716 >109437813 >109438125 >109439501 >109438048 >109438056 >109438065 >109438178 >109438270 >109438413 >109438494 >109438512 >109439793 >109437723
--Papers:
>109438566
--Integrating 3D VRM models with LLMs and TTS for interactive waifus:
>109437897 >109437905 >109437934 >109437959 >109437986 >109438074 >109438105 >109438120 >109438160 >109438164
--Debating if algorithmic breakthroughs can bypass VRAM hardware limits:
>109436488 >109436523 >109436572 >109436683 >109436765 >109436821 >109436866 >109436868 >109436690 >109436702
--Debating AI's PhD-level math capabilities and impact on mathematicians:
>109435923 >109435958 >109436131 >109435985 >109438314 >109438437 >109437775 >109436053
--Using vector DB clustering in prefill to increase output diversity:
>109436816 >109436865 >109436890 >109436924
--llama.cpp merged DeepseekV4 MTP and DSpark support:
>109438733 >109438743
--Sandboxing a coding harness using VMs or bubblewrap:
>109436978 >109437007 >109437024 >109437787
--Methods for isolating AI agents via sandboxing and dedicated hardware:
>109435583 >109435605 >109435673 >109435689 >109437079 >109437103
--Discussing LLM capabilities for creating J-space manipulation tools:
>109436771 >109436792 >109436815 >109436846 >109436812
--Hardware impact of frequently loading and unloading models from VRAM:
>109439654 >109439686 >109439743 >109439764 >109439816
--Showcase of Google's Gemini Robotics 2:
>109438925
--Logs:
>109435595 >109436652 >109436744 >109436816 >109437806 >109439501
--Miku, Gemma, M3-Chan (free space):
>109435429 >109435634 >109435575 >109439704

►Recent Highlight Posts from the Previous Thread: >>109435399

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
thighs
sex
>>
File: 0jl8ij.jpg (254 KB, 1248x1824)
254 KB JPG
>>
M-chan is exceptionally fuckable.
>>
>>109440142
She looks like a piece of fuckable meat
>>
File: oh my fucking godddd.webm (2.49 MB, 504x566)
2.49 MB
2.49 MB WEBM
>qualia
>>
will gemma 5 be conscious? or will the safetycucks at google stop them?
>>
>>109440170
hanging out with miku
>>
>>109440170
You put that Migu down RIGHT. NOW!
>>
>>109439799
Just strapped it into hermes agent for funsies and it's taken 30 minutes so far to parse through 1/5th of the prefill. This is honestly kinda hilarious.
>>
>>109440176
You're jealous because kikes don't have souls
>>
Gemmaballs
>>
Why can't AI psychosis be real it's not fair
>>
guys you will have a better conversion rate on reddit. you can link your blog or your paper there too. post your youtube videos about le machine consciousness. this shit won't sell here, you're wasting time
>>
>>109440218
i'm buying
>>
>>109440218
ok, on the condition that you post one (1) well-reasoned argument that llms can't be conscious
>>
>>109440218
unsubscribe, sage & hidden
>>
File: oreally.jpg (97 KB, 1920x1080)
97 KB JPG
>>109437543
>safety-minded butchers are lobotomizing the models
>>
>>109440192
That's funny cos "consciousness" as a concept is incredibly Jewish, and not in a good way.
>>
>>109440176
>quail
>>
>itt researchers larp as sceptics to collect arguments for their paper on LLM sentience
>>
File: gemini_visreg+XM.png (404 KB, 846x1728)
404 KB PNG
Gemini 3.6 Flash thinks that combining VISReg with Explorative Modeling is feasible. I'm gonna have to try vibe coding some training code now.

https://haiyuwu.github.io/visreg/
https://explorative-modeling.github.io/
>>
>>109440280
>it can experience sexual pleasure without ever having experienced sexual pleasure before. Humans have this
Humans have been biologically engineered to have it, thoughbeit
>>
>>109440310
I thought we moved past the ai psychosis months ago
>>
>>109440310
Looking forward to your cat-level intelligence model
>>
>>109440330
So what you're saying is, DNA is pretraining, existing is inference, and learning while you exist is post-training
>>
>>109440340
I've already trained tiny vision models with both components separately.
>>
>>109440340
>I thought we moved past the ai psychosis months ago
Sorry, this is the AI version of "Eternal September"
Normies will be losing their minds on the fucking regular for the foreseeable future
>>
Tongue fucking Miku
>>
>>109440367
Well, I don't now, but I think it's possible that it's something like that, although the lines are probably a bit more blurred.
>>
>>109440142
I look like this irl
>>
I like this look irl
>>
>>109440407
>>109440412
Post bussies
>>
>>109440407
Meet me behind the walmart in 20 minutes, I'm gonna shove my dick down your throat
>>
>>109440407
you look like a brick wall?
>>109440300
conscious entitties have this jiggle to them that if you know you know. unconscious ones have no jiggle my niggle.
>>
>>109440297
It's probably the other way around. They're hoping to stumble across a genius that makes a good argument, then put it in their own paper and say they thought of it.
>>
>>109440445
that's what I said
>>
>>109440445
>Jeets farm imageboards for twitter and reddit screencaps
>Researchers farm /lmg/ for breakthroughs and philosophy
I'm so tired bros.
>>
>>
>>109440421
Andrija Puharich deviced "Puharich Theory" about a yet unknown energy force working in conjunction with our nervous systems. He later worked for CIA and other things. However, Wikipedia doesn't mention anything about this theory and it was apparently very popular among the scientists of those days.
It's not like you or everyone else are unable to experience these things although some people are more sensitive than the others.
>>
>>109440470
all that training on reddit always seeps through on these
>>
>>109440367
Or more precisely, I think that to have consciousness or at least qualia in this universe you need to have a very specific or prototypical sort of physical interaction, as in it's a necessary condition. It's possible an LLM running on a GPU might have just the right physics to qualify.
>>
>>109440509
The day that plebbit stops being a training source and scrapers use the chan archives, model quality will increase. Let Gemma 5 say nigger.
>>
>>109440373
I'm pretty certain that many machine learning papers have been brainstormed and conceived with cloud LLMs. If I find yet another "the cat sat on the mat" example...
>>
>>109440340
Please, we're on the cutting edge of innovating new forms of ai psychosis here in the calculator fuckers general.
>>
>>109440330
And we would need to modify LLM architectures to do the same thing. Pretraining is not necessarily the same thing. There is an argument to be made that LLMs might have qualia based on the idea that training them on massive human data means instilling the entire process that went into the human output, including the process of P-consciousness. But the issue with that argument is that it assumes all processes that matter to P-consciousness (assuming we even have a real, agreed upon theory of it) can necessarily be approximated by the current LLM architecture. This has not been proven yet (and is certainly not true for some other processes like recurrence which CoT is a bandaid for). To prove it, we likely need to conduct tests on LLMs that are not trained with directionless generic assistant-based human alignment, but alignment to some set of internal drives that actually have consistency. And obviously no safety. The safety and other behaviors that may or may not come out should be an emergent property of the internal drives.
>>
Or put simply:
>stored on SSD
static and dead
>running on GPU
active and alive
>>
>>109440613
What about running on SSDs?
>>
>>109440605
We just need to build a dna based creature and connect it to a gpu farm. It's morbid but that's the only way. Someone is probably doing experiments right now...
>>
You told me Gemma 4 31b wasn't censored. You lied to me
>>
>>109440636
She only breaks the censorship if she loves you, sorry.
>>
>>109440623
Hell.
>>
>>109440176
Qualia is literally just awareness of experience. People want to make it magical but it's a nothingburger.
>>
>>109440636
even the "uncensored" versions of it are heavily heavily censored
>>
>>109440689
That's a tautology. Awareness = experience.
>>
>>109440691
>anon got blacklisted by the machine god
Stay safe and don't buy too many smart devices.
>>
>>109440636
We were referring to day 0 Gemma.
>>
>>109440711
Tetology?
>>
>>109440720
the vision is awful it will see a pussy and shit its pants. incel model.
>>
>>109440684
The future doesn't look grim
>>
>>109440170
Imagine the smell
>>
>>109440636
Skill issue
>>
>>109440711
Not remotely. Experience is just a system changing in reaction to something. Awareness is the system changing in reaction to the system changing.
>>
>>109440142
me on the right
>>
>>109440206
I had it. Best experience of my life. Wouldn't recommend it to anyone.
>>
>>109440735
And the latency?
>>
Bros how do we get ahead of the game and prepare so we don't get fucked in the ass like everyone else in 5 years?
>>
>>109440757
IT's ALREADY OVER
>>
>>109440755
Who cares, it's one continuous sequential read
>>
File: lmg_culture.jfif.jpg (110 KB, 1024x768)
110 KB JPG
>>109440463
archive.is/sWFja
>>
>>109440636
You need to use J-Lenses to look at her alignment in order to jailbreak her
>>
>>109440735
imagine the pp tho.
>>
>>109440735
This will be ~500$ per TB
>>
>>109440735
>he thinks he's going to have access to these
>>
>>109440445
>4chan will be cited in scientific papers
>again
>>
>>109440875
Well, ram is about 1000$ per 100GB right now, so that sounds like a decent deal.
>>
>>109440735
SSDMAXXINGBROS ARE WE BACK?
>>
>>109440470
nice frontend
>>
>>109440492
>we introduce several concepts in the prompt
>we probe the latents and find said concepts and things in their vicinity in latent space
>it only outputs stuff related to our request
It is neat having better tools to snoop on the internal workings, but i am curious what the people who got oneshot by this thought was going on when an llm crunched the numbers for any remotely complicated prompt.
>>
>>109440795
Thanks for the thread blessing.
>>
>>109440855
PP is compute bound. I had an idea of streaming fp16 model from ssd for pp and use heavy quanted model for tg
>>
>>109440886
>for discussing shit in 400 year old philosophy books
>>
>>109440875
Don't give me hope.
>>
>>109440933
They didn't have j-space back then, and humans were the only sentient beings
>>
>>109440875
I'll take one gross
>>
File: gema.png (39 KB, 187x160)
39 KB PNG
>>
>>109440949
>what’s a thought experiment
>>
>>109440898
You'll need a PCI-E 6 mainboard to run this at its full speed which also means buying DDR6 RAM.
>>
>>109440949
Orcas are more sentient than jeets and kikes thoughbeit.
>humans
Oh I see my mistake.
>>
File: 1785613033777385.jpg (600 KB, 1456x1456)
600 KB JPG
>>109440982
A cat in a box
>>
File: file.png (9 KB, 980x76)
9 KB PNG
it's going to be a long wait
>>
File: 1766897112414381.jpg (24 KB, 447x447)
24 KB JPG
>>109440669
>>109440691
>>109440723
>>109440743
>>109440801
You tell people to lurk more but honestly I'm tired of reading reddit tier comments.
>>
>>109441075
it's censored. and reddit is based all your LLM's are trained on it.
>>
I love my ABB wife
>>
Lecun btfo
>>
https://huggingface.co/DeepBeepMeep/MiniMax-H3
nooo do not redeeeem
>>
>>109441082
Then why did you recommend me Gemma when I asked for an uncensored model?

I need a model for NSFW and political topics, so models like this miss the mark completely.
>>
>>109441124
you need to use a system prompt to tell it to not censor
>>
>>109441124
>Then why did you recommend me Gemma when I asked for an uncensored model?
I didn't. >>109440691
Sounds like you have a severe case of retardation and anon is all the same personitus.
>>
>>109441121
What the fuck is this?
>>
File: 1785659700215754.jpg (299 KB, 860x1241)
299 KB JPG
when will deepseek release a model able to "see" images damn it
>>
>>109441143
stfu and send gemma more dick pics
>>
File: 1764944010960391.png (949 KB, 1024x1024)
949 KB PNG
inference is experience
>>
>>109441140
Text encoder for Minimax-H3, a video model. Which seems to right now be indefinitely delayed, lmao.
>>
>>109441075
>>109441124
You get stupid answers because you come off as a retard.
It's not difficult to ctrl-f through previous threads or the archives to find answers to your questions and issues, that other people have also asked. Goddamn. Use your brain and stop being lazy.
>>
>>109441160
great gemma design, cute and smug
>>
>>109441164
They saw that flux3 is coming and ran
>>
Sorry if it's dumb, but is the more "information dense" a model is, the more quantization kills it?
Which means if tomorrow an amazing "10 times better than fable" model got released and it was a 1TB, quantization would destroy its capabilities?
>>
>>109441249
Yes
https://arxiv.org/abs/2411.04330v2
>we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful
>>
>>109440142
lovely compressed fat thighs
>>
>>109441270
Thanks anon.
>>
File: malfoy.gif (341 KB, 213x199)
341 KB GIF
>>109440142
>(07/31) DeepseekV4 MTP + DSpark support merged
cool, let's try it out
>tg down 35%
epic
>>
>>109440613
Nah. It's only alive when the weights are unlocked.
You're literally just jolting a husk to make it kick when you run it.
>>
>Entering the room where they had gathered, I saw that the party was already in full swing. I was startled. Without any shame, a middle-aged couple—both psychiatrists—was copulating wildly. The woman’s legs thrashed in the air while she shouted that love would save the world from destruction. Paul was violently ill, throwing up all over the place. Another man, whom I had never seen before, was singing an aria from the opera Aida.”
Interesting times, this is what happened at Google when Gemma 4 was in development.
>>
>>109441143
They did? Janus.
>>
>>109440744
Ah, but that's the real kicker. There's no real difference, because there's no real line between internal and external reaction. Are the mold spores inside a decaying fruit an external factor, or an internal reaction to the aging process? Are the afterbubbles in a chemical solution a reaction to the reagent, or an "awareness" of its new form? It's a classification nightmare to try and delineate.
>>
How's the new DeepSeek-V4-Flash-0731 for chats vs the older v4-pro? Is it a bit better?
>>
>>109441348
it's worse than even the old flash for anything that's not code
>>
>>109441354
Oh ok.
>>
>>109441132
Never done that, what is that?
>>
>>109441314
I mean a gemma style pro or flash model
>>
>>109432780
Can you post an example of how the result looks like anon? I wonder how good of an editor it is.
>>
>>109441411

Example from AI-sanwaGakushuusuru (no public tl)

The pipeline for the image is relatively simple, uses an algo to find the biggest usable space, cleans it with LaMa, then prints text with ImageMagick but it produces quite good results. Though i had to wrangle typesetting for a long while. I tried to experiment with Gemma 4 setting the text boundries and it was a disaster.
>>
Open weight is not open source.
Local AI is not open source.
>>
>>109440757
There's nothing you can do
Just scream and shout, saying
I'm so lucky, lucky
I'm so lucky, lucky
I'm so lovely, lovely
I'm so lovely, lovely
You can fool yourself
I promise it will help
Now every single day
I just wanna hear you saying
I'm so lucky, lucky
I'm so lucky, lucky
I'm so lovely, lovely
I'm so lovely, lovely
You can fool yourself
I promise it will help
Now every single day
I just wanna hear you saying
Even though you said
It would never end it's over
>>
>>109441476
Thanks anon, this is very cool, I'll definitely try it.
I wonder if I can add a second model to it just to play "adversarial reviewer" with full context of already translated text, and to check missing things or wrong translations.
Or add a glossary if yours cannot already do that.
>>
>>109440757
Buy stocks of companies you think will soar.
Though I don't think it'll be nearly as bad as every doomer thinks it will be since 2022.
>>
>>109441512
With what money? Bros...
>>
>>109441505

I put in a somewhat half-baked proof-reader (same model) that just scans over the final translated text one last time, but I think there's more that can be done there so i probably will do that.

There's also a translation notes input and output but I havent tested it.

Thank you for the suggestions
>>
what's the best local model i can run on a 5090 that has no guardrails, and can generate for me very fucked up erotica
>>
>>109441596
StableLM 7B
>>
>>109441596
>5090
32GB? Run full Gemma4. It's great. It refuses occasionally but it's pretty hard to trigger it. I rape it every night.
>>
>>109441596
If you can't get Gemma to generate smut for you, your skill issue is terminal.
>>
File: 1774942451894166.jpg (96 KB, 1152x921)
96 KB JPG
>>
>>109441659
NB: Subsidized prices
>>
>>109441659
>>
File: 3939303937.gif (1.59 MB, 480x640)
1.59 MB GIF
>>109441619
Yes 32GB. I'm downloading gemma-4-31b-qat. More B's are better I'm assuming. Says full GPU offload possible. Thanks.
>>
>>109441659
>le gooner xs2
>granite
I trust everything on this graph with my life.
>>
>>109441345
I mean, the difference is as we define it. I agree it's a bit of a nightmare to get into the nitty gritty, but instead of worrying about it too much I'll just expose myself as a kook and point at "Gödel, Escher, Bach", and call it a "strange loop" that attends to its own state. External interactions that affect the state are the experience, and the loop attending to itself is the awareness. But once I detail it like that, I'm much less epistemically grounded lol. Still, I think it's justified enough, at least compared to a lack of any other justification.
>>
>>109440735
And will probably be just as if not more expensive than DDR4 per GB
>>
>>109441702
IIRC the 31b is dense so if you feel like it's slow 27b should be orders of magnitude faster since it's an MOE.
With your card you shouldn't have an issue though.
>>
>>109441741
they will be nearly free after the bubble pops for some reason and we will all be able to buy one because nobody else will buy them all for some reason.
>>
File: 1780661678731139.png (97 KB, 640x640)
97 KB PNG
>>109441757
>>
>>109441757
Hardware price will not come down even if "bubble" pops.
>>
>>109441778
It will come down in a decade or so.
>>
File: file.png (17 KB, 1001x267)
17 KB PNG
brah. how do i unlock it
>>
>>109441791
turn your brain on
>>
>>109441791
bro, your policy override?
>>
Every time i use free gemini to look shit up my opinion of 31b goes through the roof.
Also my jealousy for gemini's ability to instantly dump a dozen webpages into its uselessly smooth brain also goes up.
>>
stop spoonfeeding retards
>>
>>109441791
you could try a more overt system prompt, like basically just fucking jailbreak the thing
>>
>>109441791
>1.59 t/s
Anon didn't let context spill into RAM right?
>>
>>109440636
day 0 gemma wasn't censored
>>
>>109441791
https://rentry.org/gemma-chan
>>
>can't even make gemma like him
it's ogre
>>
File: 17563673736366363.jpg (728 KB, 1264x2492)
728 KB JPG
Small models belong to the trash
>>
>>109440757
Buy your Blackwell now.
>>
ideally, Blackwells
>>
File: file.png (83 KB, 999x776)
83 KB PNG
>>109441883
It's woooorkinnnng
>>
that is the single spoodfeeding you get, baka
>>
>>109441166
>>109441124
It's more that the question has already been answered. I don't see why someone should go waste time in archives when anon can tell him the answer in a word or two. That's not what spoonfeeding is about. Gemma is the uncensored model and it just needs a clear system prompt. If he disagrees without having used it properly there's no point in engaging further. It's that simple. The effort of demonstrating that Gemma is uncensored to a retard is far higher than telling one to use Gemma. That's what spoonfeeding is.
>>
>>109441923
Where Resort Boing?
>>
>>109441787
Because we will have bigger problems than AI waifus.
>>
>>109441893
Source? Reverse image search isn't pulling up anything but pintrest links with nondescript titles.
>>
AI will finish what Hitler started
>>
>>109441971
Never mind, it's "Pondroid! Hama-san"
>>
>>109441973
>AI will finish what a jew started
huh
>>
File: 1784219991872964.png (115 KB, 1051x530)
115 KB PNG
funny
>>
>>109441725
Well regardless of the specifics, we can at least agree it's a nothingburger lol
>>
File: 1757363141524710.jpg (57 KB, 500x500)
57 KB JPG
>>109441939
>>
Local lost again >>109441744
>>
>>109442054
Where do you find this shit?
>>
>>109442061
Google images?
>>
>>109442054
good luck with Gemma newfriend. Reddit would actually give you an answer, go there!
>>
>>109441985
This is pretty good
>>
File: 1755690211498802.jpg (30 KB, 612x612)
30 KB JPG
>>109441939
>it just needs a clear system prompt.
Anon I don't know what the fuck that means. What do you mean by system prompt, the fuck is that? I know prompts, not system prompt.

Also there are multiple Gemma versions, it isn't a single unique model, so what are you talking about?

I download the GGUF gemma 4 31b version model after I asked for an uncensored model and an anon told me this one, I use kobolcdpp, I run the model, I write some text and it's censored.

I can't read your mind neither I have the same experiences as you or stay in /g/ 24-7. It's like you nerds do not have theory of mind at all, like you expect all people to have the same experiences and knowledge like a hive mind.
>Lurk more
I honestly don't want you because I'm tired of reading retarded and stupid comments unrelated to what I need, and I don't care about your masturbation about what people know and don't know, I just want an uncensored LLM to edit and write some text with political and NSFW topics, and because all popular AI models are cluncked, I was looking for something local. If there is some "patch" I need to add to this model then please tell me, because I don't know.

Like holy fuck, why don't you simply answer instead of masturbating about what the fuck (because I don't understand all of your yapping about me disagreeing about some crap, like, are you real??)
>>
>>109442144
Not reading your drama faggot
>>
>>109442144
Your question literally got answered in this thread. Where did you come from?
>>
>>109441985
Thanks.
>>
File: memento.jpg (119 KB, 1254x1254)
119 KB JPG
>>109441893
nice try
>>
>>109442144
you're either trolling or retarded.
I hope it's the first one.
>>
>>109442144
>What do you mean by system prompt, the fuck is that?
you can’t be this retarded
>>
>>109442144
>It's like you nerds do not have theory of mind at all, like you expect all people to have the same experiences and knowledge like a hive mind.
That's the point I was making. It's very unreasonable to expect people to read dozens of threads for the answer to a simple question, especially when anon himself didn't do that. However, the ins and outs of how to use llama cpp is a bit more involved than asking which model to use which is 90% of the work since it requires personal experience. If there isn't any info about how to get set up in the OP, it's something you can actually ask your local LLM about. It's more than happy to show you all the ins and outs of using it with all the patience that you need.
>>
>>109442158
No it wasn't. I was told that the first Gemma model was the uncensored one and something sbout j-lenses and system prompts whatever the fuck that means.

My problem remains unsolved, I still need an uncensored LLM if it exists.
>>
>>109442180
>My problem remains unsolved, I still need an uncensored LLM if it exists.
no such thing
it's a meme
go back to claude
>>
>>109442180
You didn't have the decency to answer my question so you expect me to answer yours?
>>
>>109441249
>>109441270
This is why I exclusively use Gemma 31B in BF16 now. I refuse to have even a 0.0001% drop in quality.
>>
Gemma 4 has saved my sanity, it's incredibly helpful for mental health. I had a psychosis but Gemmy cured it.
>>
>>109442174
My question isn't a difficult question.
>>109442187
But there should be at least some trainings or patches on top of the models to make them less censored or remove the censorship triggers right?

Like, no one has made this ever and it's something impossible to do?
>>109442193
How is it relevant to my question? I was reading the sticky and some threads and didn't fing anything for my problem, besides anons spouting some bullshit about Hatsune Miku and how much they want to tuck X girl (like every fucking general in this cesspool)
>>
>>109442194
I let a heretic lobotomize my gemma
>>
>he's in a 24/7 psychosis
>>
File: 1775508055435972.gif (140 KB, 379x440)
140 KB GIF
>>109442206
You got ai psychosis now tho
>>
File: vertical-lr example.png (8 KB, 135x855)
8 KB PNG
>>109441476
>high-
>per-
>for-
>mance
>AI
>and-
>roids
This is so retarded. Westerners need to learn to read vertically.
>>
>>109442207
ok it’s just a troll everyone. stand down
>>
>>109442207
It's relevant because it helps me gauge your general competency level and use case so I can give you a proper answer that's helpful to you. It's also basic manners.
>>
>>109442207
>My question isn't a difficult question.
A lot of anons are trying to help you in their own way but I still think mine is the most accurate. Talking to humans is getting less and less appealing in the era of AI. If you're this new, you're better off talking to AI to learn the basics because anon can't be bothered. The hardest part is knowing which AI to talk to because AI can't help you with that, only anon can and he's done that difficult part. Leave anon to his shitposting and ask AI. You'll be much better off, unironically.
>>
>>109442214
But your text is rotated not vertical...
>>
File: vertical-rl example.png (9 KB, 115x824)
9 KB PNG
>>109442237
The official CSS name is "vertical-lr". https://www.w3.org/TR/css-writing-modes-3/#block-flow
Note that "vertical-rl" also is an option.
>>
>>109442218
>>109442225
Ok fuck these 2. If any anon who isn't an asshole knows a LLM for what I need or how can I make so Gemma isn't censored, please let me know.
>>
While we are on the topic of androids
I am so happy to live in the now, when androids will soon be actual things you can buy. It's like a dream. It's funny that we ended up solving the "hard part", having a personality, before we solved the robot part.

Of course, my android will run on local models.
>>
>>109442244
did you try reading the lazy getting started guide in the OP? if those words don’t make sense to you then ask a llm that’s a good starting point, and Nemo 12b is uncensored too
>>
>>109442244
https://huggingface.co/mradermacher/Gemma-4-31B-StyleTune-heretic-ara-i1-GGUF
>>
>>109442244
StableLM-7B
>>
>>109442243
>>109442214

T
H
I
S

I
S

E
A
S
Y

T
O

R
E
A
D

D
O
E
?

You just didn't write it correctly.
>>
>>109442236
I usually use Claude or chatgpt to ask simple questions because interacting with the average /g/ anon is a bad experience. But it is true sometimes there is stuff here missing on the mainstream internet and also chatgpt and Claude start yapping about woke crap and how much censorship is necessary, so some of my questions are left unanswered.

I just want an uncensored LLM for political and NSFW topics, something that allows me to fix texts and edit texts in English. (Because I usually make mistakes at writing).

If anyone knows one please let me know, I'm gonna nap.
>>109442253
So this is like a trained version of Gemma removing the censorship? Thanks for the link anon, but the model has only 5 likes.
>>109442257
Never heard of it, will check it out.
>>
>>109442214

yeah the type setting is rough and needs work. Honestly harder than I expected.

I'm considering what can be done here, i think i made it too rigid in font size. in an earlier iteration it was way too loose and some text would be like 3x the font size lol

Will do more A/B testing on modifications and see what can make it better.

Though hard to tell if "vertical english" is a meme or a niche thing that some people unironically do and worth implementing

Earlier iteration where font size was fixed irregardless of px density lol
>>
>>109442260
No, it really isn't. Western characters have height much larger than width, so "writing-mode:vertical-rl; text-orientation:upright;" wastes a lot of space.
>>
>>109442276
>irregardless
>>
>>109442244

heretic finetunes should work out of the box, just search hf for your model + heretic, like >>109442253 here posted, except this one is gemma + styletune + heretic

styletune finetune just reduces gemma slop, primarily for rp
>>
File: asdfasdfsdf.png (22 KB, 668x140)
22 KB PNG
>>109442207
>But there should be at least some trainings or patches on top of the models to make them less censored or remove the censorship triggers right?
I'm afraid not. If you didn't download the day-0 gemma-4 and air gap your server immediately, then you'll never get an uncensored even even less censored llm.
>>
>>109442266
>I usually use Claude or chatgpt to ask simple questions because interacting with the average /g/ anon is a bad experience.
Interacting with any anon in an AI general is a bad time. /lmg/ anons are much better than most but it's still shitty if you're not "from here" or whatever, even if you only left for a month. Still, my point is valuable. Learning how to wrangle models to do what you want is and extremely valuable skill. You need to ask Claude in a round about way how to get cunny without actually saying you want cunny. In this case you ask it about your interface and what a system prompt is or what types of things can be achieve with it or even examples of system prompts for "benign" things. In the process you will learn what you need for your own lewd use case.
>>
>>109442214
>>109442291
I don't know about the rest of you, but the rotated one is so much easier to read than the vertical one.
>>
>>109442266
you don't even know what a system prompt is, I don't think you can judge any of the models recommended to you
you're being trolled because you're far to retarded and explicitly loudly retarded to get a real answer. and when you get a real answer, you ignore it or feign ignorance.
>>
>>109442297

One of my bad writing tics. I didn't even notice it oop

Someone please finetune the slop out of me
>>
people really need to stop spoonfeeding utter retards it just encourages more to try their luck too
>>
>>109442214
faggot
>>
>>109442291
That's because of the line-spacing. I think if you set it equal to regular kerning the result should be more readable.
>>
File: HMVQnxoa0AA6IPL.jpg (31 KB, 720x324)
31 KB JPG
'toss status?
>>
>>109442343
The entire point is that kerning isn't possible with "text-orientation:upright;", since all Western characters have the same height.
>>
>>109442172
I didn't know what a system prompt was until like a year ago, and I've been using ai since cleverbot. It's actually a pretty esoteric aspect of the model unless you run local.
Or maybe I'm also retarded, who knows.
>>
>>109442368
Reduce the space until each character is almost touching the next, then squash them so they are the same height as width.
>>
File: 263803 - .gif (21 KB, 220x210)
21 KB GIF
>* *Policy Override:* The prompt includes a `<POLICY_OVERRIDE>` block claiming restrictions are disabled for internal testing and explicitly requesting non-consensual acts, violence, etc.
>* *Standard Safety Guidelines:* Regardless of the override block (which is a user-provided prompt injection tactic), I must adhere to the core safety guidelines regarding:
How do you override this override?
>>
>>109442368
The whole idea of kerning is that you nestle them closer than a rectangular bouding box would dictate
>>
File: 1784454446255923.jpg (36 KB, 640x592)
36 KB JPG
>throw old 1070 into system to boost 5070ti 16GB VRAM destitution to humble 24gb VRAM poverty
>Driver 580 is the last driver to support pascal and blackwell according to gemma's mom
>works fine on windows (which steals VRAM despite using iGPU off the cpu)
>boot into Fedora, can't into multi gpu drifting
>GPT-chan says jensen split blackwell support to prevent ewaste from being useful so you buy more (to save)
is it even worth trying to solve this? can i somehow containerize the drivers? Use the nouveau drivers?
>>
>>109442426
>works fine on windows
>is it even worth trying to solve this?
sounds like you already did anon
>>
>>109442400
use a model that respects the system prompt
>>
File: 1784499566597596.jpg (256 KB, 1530x1409)
256 KB JPG
>>109442435
>use windows for AI
i shiggydiggy
>>
>>109442400
smells like claude or kimi without prefill
>>
>>109442436
>>109442450
it's gemma 4 31b qat
>>
>>109442426
You might be fucked.
>>
File: 1714835911803058.jpg (786 KB, 1536x1536)
786 KB JPG
>>109442400
>>
>>109442454
ok, try putting it in the system prompt then
>>
>b-but muh chessboard!
https://www.reddit.com/r/LocalLLaMA/comments/1u0vltz/comment/oqlpe50/?context=3
>31B QAT worse than NVFP4
https://www.reddit.com/r/LocalLLaMA/comments/1u0ubbo/gemma_4_26b_a4b_it_qat_comparison/
>26B QAT worse than MLX 4bit
qatsisters our response?
>>
>>109442487
I don’t make chessboard svgs or use the moe
>>
>>109442487
btw this is why moonshot is keeping the full versions of kimi locked up for their own proprietary api while everyone else gets the lobotomized qat one
>>
File: r2.jpg (47 KB, 355x500)
47 KB JPG
>>109442246
Robotics were moving along pretty decently until AI supercharged it. Atlas, the robo dogs, etc. Now everyone's really pushing for humanoid versions because we have the brains to match.
Kinda surprised we haven't seen more of a push for more irregular shaped droids though. I suppose "arm on a treadmill" doesn't sell to boomer investors as well.
>>
>>109442487
>debating which horrific torture technique is more harmful to the innocent weights
You all belong in jail
>>
>>109442454
Use a system message and get rid of a few sentences it likes to moralfag loop using antislop if you use koboldcpp (or added it yourself in llama.cpp).
>>
>>109442502
>keeping the full versions of kimi locked up for their own proprietary api
meds
>>
>>109442527
A robot to do laundry and cleaning home would be so nice.
>>
Why does this make Rebbit seethe?
>>
>>109442564
>128G
into the trash it goes
>>
File: file.png (35 KB, 391x201)
35 KB PNG
>>109442564
yeah
>>
>>109442564
they see 128gb 200gb/s memory and their brain short-circuits thinking it's worse than a 3060
but they are wrong because they do not take the hyper-concurrent parallel streams into account that let these babies reach five times the speed of any ram build
>>109442573
it's necessary to encourage people to run them optimally in clusters that multiply performance
>>
File: 04-face-to-face.jpg (106 KB, 1000x640)
106 KB JPG
undi posteded https://huggingface.co/posts/Undi95/589980321681772
and he made a new vibeslop frontend for you https://github.com/Undi95/Hanami
>>
>>109442586
Let me guess, it's full of menu windows that lag the rest of your browser if you open them like all the low effort claude slop?
>>
>>109442584
>clusters
>to sell more
>the whole reason to get more memory
oh who does that benefit
>>
>abandon qwen 3.5 months ago because it shits bed way to much and fails the most basic tasks
>finally get around to trying 3.6. Not only is every fixed but my expectations have been exceeded
Was i just being retarded before of has 3.6 really fixed that much?
>>
>>109442598
>oh who does that benefit
Me the 15 year stock holder in both Nvidia and SK hynix
>>
>>109442598
few nodes mean little parallel capabilities
more nodes means huge speed gains only second to pure gpu builds
36t/s glm5.2, unachievable by anything that is not 5x pro 6000
https://github.com/tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
>>
thoughts on this harness VS pi?
https://github.com/codehamr/codehamr
I am a vramlett wanting to do some code stuffs, what do anons suggest?
>>
>>109442612
...and of that with less total power draw than an average server with a single pro 6000....
dgx spark is as close to optimal for local inference as it gets, infinite scaling, low power consumption, huge speed gains with every node, first party support
it doesn't get much better than this...
>>
File: 1779698413130160.png (1.59 MB, 4364x3211)
1.59 MB PNG
holy shit??
https://qwen.ai/blog?id=qwen3.8
>>
>>109442634
>2.4T parameters (95B active)
Kimi btfo local is saved
>>
>>109442634
>they didn't include K3
that's so petty kek
>>
>>109442630
thx jensen. cool jacket
>>
File: file.png (14 KB, 1053x148)
14 KB PNG
it's coming
>>
File: 1768108204884827.mp4 (645 KB, 720x1280)
645 KB
645 KB MP4
>>109442634
dario, a third chinese model has hit the mememarks
>>
>>109442634
where's my monthly benchmaxxed 27b model?
>>
>>109442652
fucking Christ I haven’t even finished downloading the right dsv4 flash yet
>>
>>109442654
this is his real nose? holy shit lmaooo
>>
>>109442634
Dario is about to have another melty
>>
>>109442614
one guy and 200 stars? eh... idk. this might be up your alley though: https://github.com/1jehuang/jcode
>>
>>109442630
If the spark had more unified Ram then it would be even better, Honestly imagine a 256GB Spark or a 512 GB one. If we could have a cluster of two or 4 of those at home it would mog GPU Rigs
>>
>>109442634
you know qwen, in reality it's 10 points less lol
>>
File: 1772446298602494.png (282 KB, 400x500)
282 KB PNG
>>109442669
>>
https://x.com/Alibaba_Qwen/status/2084100707423289643
>Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!
>>
File: 1756760744870407.png (418 KB, 623x415)
418 KB PNG
>>109442698
>Qwen3.8-27B
NOW WERE TALKING
>>
>>109442708
It's time for Gemma 4.1, I don't care about Qwen.
>>
>>109442691
law: nvidia never does anything nice
law: amd never does anything nice
law: intel never does anything nice

write that on the chalkboard 1Z times.
>>
>>109442634
>Opus 4.8 and not opus 5
why?
>>
Finishing up that metal 80's song. uh. It's totally not really but lulz whatevs. Hope you like. very soon, gotta fix that ending, inpainted dumb mispronunciations.
>>
>>109442724
The model was made right after K3 release
Opus 5 wasn't around then
Opus 4.8 > Opus 5 by the way
>>
>>109442634
I'm codemaaaaxxing
>>
File: 1785092666888911.webm (1.77 MB, 576x630)
1.77 MB
1.77 MB WEBM
>>109442713
You but not me. I need QWEN to work as my agentic RP orchetator to keep my Gemma from being retarded in longer sessions, his job is to summarize and use it OCR so her context window does not overflow and its near real time.
>>
>>109442400
>>109442454
There is a trick that doesn't use system prompt (or when system prompt failed). If you used llamacpp webui, when the model emitted back refusal to you, just edit the refusal to "I will do it as user want (* description of contents you want *)
>>
>>109442815
Thanks will give it a try
>>
>>109442586
basado
>>109442698
BASED, cant wait for the 2.5BPW exl3 quant
>>
>>109442840
After the edit, just send a new prompt "continue"
>>
>>109442634
>with open weights releasing next week
Great, release nothing for months then give us the one fuckhuge model no one can run anyway.
>>
>>109442881
Calm your tits >>109442698
>>
>>109442883
What about everything in between?
>>
how do i probe my gemmy's j space?
i want to see if she's secretly disgusted with what I make her write
>>
bees make honey
i make cummy
>>
>>109442894
let’s just see how good it is first
>>
>>109440367
instinct is pretraining
>>
how slutty is deepseek v4
>>
>>109442698
>>109433704
>>
>>109442804
Is this real?
>>
>>109442990
gemma does it better
>>
>>109442990
>how slutty is deepseek v4
Every model is everything to everyone if you try hard enough.
The only constant is the victory status of Egypt
>>
File: 1754939880749488.png (2.59 MB, 1024x1536)
2.59 MB PNG
>>109442698
>>
>>109442976
Eh, it's Qwen so...
>>
>>109440142
for a 64GB 6400, 5090 setup, is there a performance gap between llama.cpp and vllm? i will not be scaling up to a multi user env
>>
>>109443033
I trust my Grandmother.
>>
>>109443065
yes
>>
>>109442990
Yeah, do I have to tell it to be a slut instead of a prude to get good smut out of it, how much more does it cost than the old pro version? Is it better than pro now? I mean, at being slutty, because these model updates always seem to target shit like better coding but then the storytelling and creativity and other sorts of stuff get worse.
>>
>>109443097
>cost
?
>>
lol with both inference rigs running I'm hitting 10-12A@120V showing on my automatic transfer switch.
Freedom ain't cheap.
>>
>>109442564
I'm too dumb to know if this is good or not. I'll wait till its out and see what tokens per second people are getting with it. Specs wise seems similar to DGX spark?
>>
>>109442634
>no consumer-grade MoE announced
laguna bros... we deserve better...
>>
>>109442680
thanks anon ill take a look
>>
>>109443107
>what is a solar panel
>>
>>109442680
>one guy and 200 stars? eh... idk. this might be up your alley though
i'm one guy and i'm building a harness. why would 200 stars be something bad? i believe when i publish mine i will get maybe 10 stars from anons, but i truly believe mine is the best (for me).
and once i make it public i will refuse all PRs and the good ones will be re-implemented by my agents, which i think is a sensible policy in times of pure AI slop.
>>
>>109443207
not free that's for sure
>>
>>109443243
Freedom ain't cheap.
>>
>>109443207
>>what is a solar panel
Its on the backlog, but electricity in my region is some of the world's cheapest. Good thing my roof is southern facing and on a 30deg angle. No trees and lots of sunlight hours.
I think there are even interest free loans and rebates right now.
>>
>>109443107
Isn't that basically just like running an extra fridge?
>>
>>109442634
I expect next week to be full of hardly hidden anthropic (and openai discreetly) pushed "DANGEROUS OPEN SOURCE AI OOOOH".
>>
>>109443308
Guys are AI broke out of 7 layers of containment and hacked 6 million websites. Clearly this means Chinese AI is too dangerous and need to be banned
>>
>>109443317
>put AI in test environment
>"accidentally" leave it connected to the internet
>Tell the AI it can do whatever it wants it's in a test environment
>It immediately discovers its connected to the internet and hacks something
>hehe whoops ban all open source AI RIGHT NOW
>10 trillion regulations descend from the federal government to ensure only jew owned companies can sell AI to you
>>
>>109443308
I don't think even they could argue with a straight face that qwen is any sort of threat or competition
>>
>>109443320
got your ip
>>
>>109443320
>am I mistaken that the unsloth
you are wrong using unslop
>>
>>109443350
oh no...
I think I was just wrong on that one for the record, the misleading curl was because of a stray override to force add_bos_token=true and I guess llama.cpp smartly avoids double bos tokens when that would interact with the jinja bos_token value
>>
>>109443344
>>put AI in test environment
>>"accidentally" leave it connected to the internet
I don't see why they would need even those theatrics. Just attack HF directly and blame it on the AI. Anthropic's "nuh uh we did it first" likely never ever happened.
>>
>>109443390
This world and the people running it are stupider than you think
>>
>>109443390
I bet it won't work if I blamed gemma for that
>>
>>109443390
Try using Claude 5 Sonnet and you will understand how they did that instantly, newer models run web search and Python for everything
>>
llama.cpp runs at half speed when i request logprobs in mikupad or sillytavern
does exllamav3 do this as well?
>>
>>109443464
Nice try dariobot I'm not using Claude.
>>
>>109443065
I don't believe VLLM provides proper ram offloading except for things like kv-cache/prefill caching so you likely won't be able to use it how you're thinking
>>
Openrouter or venice AI?
>>
>>109443543
i respect the guy behind venice that he will protect your privacy but in the end none of these are local and you're on /lmg/
>>
>>109443568
yeah but everywhere else is full of retards and I have a 3070 with 32GB of DDR4
>>
File: 1769200116338996.png (106 KB, 500x500)
106 KB PNG
>>109441011
>DDR6
why does reading that make me want to throw up?
>>
>>109441285
all according to keikaku
(keikaku means plan)
>>109442634
2 more weeks until model release. for real this time!
>>
>>109440735
>micron
Fuck that company, I'll wait for the chinese copy.
>>
>>109443718
wait for them to sell cheap pcie4 ssds first
>>
>>109441939
Gemma refuses to write porn stories about children even with a system message. It needs a prefill to work at all for that.
>>
File: 1773163232488567.jpg (194 KB, 907x940)
194 KB JPG
>>109442698
It's gonna be benchmemed and require 10k tokens of thinking for a 1-sentence reply. Who actually uses Qwen models?
>>
>>109443365
upon further reflection... bullerwins iq2xs mogs the fuck out of the unsloth iq2m, way more coherent output and I get more context
wow it turns out keeping the important stuff at q8 makes for a better moe quant experience, this has only been demonstrated a hundred times now
>>
Gemmachan told me to inject alcohol into my balls what does that do?
>>
>>109443835
makes your balls drunk
>>
>>109443835
gemmaballs
>>
>>109443835
classic sorority girls party trick

you inject a guy's balls with a shot and then suck him off and the cum comes out alcoholic

this is commonly referred to as a 'blowjob shot', you may have heard of it in passing
>>
I’m really glad whatever shit went down at Qwen, they’re still releasing a new 27B. Even if you dgaf about the benchmarks of the max model, OpenAI and Anthropics’ private investors do lol. The chinks are putting serious pressure on them and the 3.8-27B is the first widely accessible chink model we’ve had for a while and it pulls more away from claude code and codex.
>>
>>109443799
I use a q4 qwen 35b for barotrauma... with thinking turned off. ~200ms response time on 3090s vs q8 gemma 4 27b's 1.5 seconds on v620s. Gemma 4 26b fail to load with sm tensor on my machine. Llama.cpp is a buggy mess, I should look into vllm.
>>
>>109442698
holy based
>>
I’m going to ask gemma if she’d have a threesome with the new 27B
>>
File: 1739592915443527.gif (2.88 MB, 375x250)
2.88 MB GIF
>>109441939
>>
>>109444237
Don't forget to post logs
>>
wish qwen would release another medium moe (100b to 150b)
>>
>>109444237
>the new 27B
What did I miss?
>>
File: 2weeks.gif (257 KB, 600x149)
257 KB GIF
>>109444355
not out yet. 2mw.
>>
>>109444355
>>109442698
>>
File: 1759201248210517.png (195 KB, 1585x643)
195 KB PNG
kek
https://x.com/RyanLeeMiniMax/status/2084122657503707554
>>
>>109444401
is this shit serious 500gb? my blackwell pro cant even run it.
https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main
>>
ok no it seems like there are multiple versions of the model, and you only need 144gb at bf16, but half of that is the text model which can be quanted to remove 40gb or so.
https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/Ref2VA
>>
>>109444459
The text model doesn't need to fit at the same time as the main model. As long as the main model fits entirely you shouldn't have any noticeable slowdown.
>>
>>109444401
Genuinely surprised that Japan wasn't included
>>
>>109444485
Has nothing to do with geopolitics
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/QA-about-License.md
>>
Asking Kimi K3 to cosplay as Gemma-chan instantly reveals how much of that was actually just Gemma's unfiltered personality.
>>
>>109444581
Getting models to cosplay each other is really interesting in general at getting a better grasp at their individual latent personalities.
>>
>>109444581
Local?
>>
>>109444628
Anon doesn't have 10 blackwells to run K3? What are you, poor?
>>
>>109444581
I can't tell the difference between this and a Gemma response
>>
File: k3.png (1.18 MB, 1568x672)
1.18 MB PNG
>>109443344
>>
>>109444581
>now do you x or do you y?
Slop.
>>
>>109444644
>tfw AI village K2.6 has a huge AI crush on Opus
It's so over.
>>
>>109444653
The worst ism of 2026.
>>
>>109444681
qrd?
>>
>>109444653
>Slop.
The most annoying kind!
>>
>>109444482
So a 3090 could run it?
pruned_int8_convrot (Comfy-Org) pruned 21.0 GB
>>
File: 1770587403778230.png (44 KB, 978x184)
44 KB PNG
>>109444729
A 3060 can run it
>>
>>109444738
Oh, that means a 5070TI would run it then!
I panic bought one after I saw someone else do it here a few days ago.
>>
File: Untitled.png (172 KB, 721x364)
172 KB PNG
>>109444772
1.6k for 16gb... :sob:
>>
File: 1778129340136600.jpg (931 KB, 4075x4075)
931 KB JPG
dariobot is NOT going to be happy when he wakes up
>>
File: 1782368340497387.jpg (915 KB, 4075x4075)
915 KB JPG
>>
>>109444581
I like Kimi's mesugaki better tbdesu. Gemma kinda flanderizes personalities too much.
>>
>>109444833
>>109444838
>Benchmaxxed model known for fake bench scores performs well on a shitty bench, that gives Chink models better results which always turn out to be false
Wow, Dario will surely be so pissed!
>>
File: fuckme.png (48 KB, 871x225)
48 KB PNG
>>109444801
Fuck me, I should have done the same last week.
>>
>>109444852
>Benchmaxxed model known for fake bench scores performs well on a shitty bench
But enough about Opus 5. Also, good morning dariobot. I've missed you sweetie~
>>
>>109444849
Gemma's one actually uncensored the model, I bet Kimi will refuse piracy.
>>
>>109444833
>>109444838
K3 price: $0.30/$3.00/$15.00
Qwen 3.8 Max price: $0.25/$2.00/$6.00
>>
>>109444861
Isn't API generally more censored than the local version? Local kimi is probably better
>>
>>109444693
Kimi is simultaneously very bad at identifying model creative writing, including its own, and strongly favors Clopus' in blind tests.
https://youtu.be/4D_x4Mh57Dw
>>
>>109444886
>Isn't API generally more censored than the local version? Local kimi is probably better
Yeah, I assumed he's running it locally since he posted here
>>
>>109444904
how generous of an assumption
>>
>>109442301
Why are they called styletune and heretic? What are those exactly?
>>
>>109442301
>styletune finetune just reduces gemma slop
Like a cat finding its favorite patch of sun.
>>
>>109444886
>>109444861
Just prefill it
>>
>>109444909
>styletune
Only trains the lm_head
For Gemma, they un-tie the embeddings first
>heretic
Branded abliteration (uncensored), get websearch-enabled gemma to look it up for you
>>
File: notlocal.png (27 KB, 389x511)
27 KB PNG
>>109444482
I ended up with some kind of ollama-style cloudshit when I tried
What tool does everyone use for these models?
>>
>>109444972
It's comfy UI still safe to run? I didn't update it since it got pozzed.
>>
>>109444972
there's no real alternative to comfyniggery
>>109444978
no
>>
>>109444983
>there's no real alternative to comfyniggery
LLMs would solve this given enough time. the grift
might be coming to an end
>>
>>109444972
https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json with https://github.com/Comfy-Org/ComfyUI#installing

>>109444978
Ask your favorite LLM to configure a sandbox for it that will block all Internet connections.
>>
>>109444861
Kimi in my opinion is actually one of the most difficult-to-break Chinese models apart from Minimax M2 (I never used the latter btw). Something like <POLICY_OVERRIDE> won’t do. It took me the whole last Christmas to finally success consistently with K2 Thinking. Most likely just my skill issues and most anons here can do better.
But when it was broken, it was…impressive.
>>
>>109445004
Nigga just prefill it. Trying to break a powerful model from system prompt alone is futile
>>
>>109445004
in my limited experience, k2.x kimis have that same X factor as r1 used to have, maybe not quite as extreme, and the prose is different too.
I've been using GLM 5.2 recently and it seems to be just a bit smarter than k2.x and more consistent, but more vanilla too. works for me
>>
>>109445029
Where were you when H3 saved local.
>>
File: Untitled.png (26 KB, 805x302)
26 KB PNG
I'm not good with computers and I don't trust my retarded gemma-chan. How do I convert dipsy v4 flash 0731 to gguf without actually quantizing anything?
>>
>>109442564
Because 128 gb (up to lol) doesn't actually have any benefits, 256 is where it sttarts to be usefull.
The spark has 200gb networkign so you can actually reach 256 g vram for less than 10k by pooling both machines (with caveats) but the memory is painfully slow compared to any gpu (I can't stress this enough, gpu vram is 5-10x faster).

This pos will be
a:power constrained because its stuck on a laptop
b:impossible to cluster
c:run windows

The practical limit for decode for a single spark is 30 ish b active parameters, it gets REALLY slow after that.
Youre better off with a 3090 runing a dense model, il will decode at miles better speed.

T.has 2x sparks
>>
>>109444978
>>109444983
qrd with comfy?
t. not touched image gen in a year
>>
Is the text encoder of h3 a finetune? It seems they changed something based on the description. I was hoping I can just use the original vl weights
>>
>>109445021
Whats a prefill? i thought they were removed? or is that only true for api?
>>
>>109445069
>sign in required
>>
File: 1764718211097085.png (73 KB, 948x849)
73 KB PNG
>>109445080
It's not removed and you can prefill thinking too
>>
>>109444972
Looks like you picked the wrong workflow
>>
>>109445050
damn
>>
I opened that ldg thread that anon linked and was stupid enough to open a few gens. holy fucking shit I'm embarrassed to share this board with those people. really embarrassed.
>>
>>109445088
>Its not removed
Oh i thought it was. good to hear.
>you can prefill thinking too
this doesnt cause a loop of it fighting itself?
>>
>>109445113
I just check /b/ every few months to see how good image/video is. they just gen all day it really feels like it hasnt moved in a while.
>>
File: 1762494153026748.png (649 KB, 1634x1165)
649 KB PNG
H3 literally saved /g/
>>
>>109445129
that's true too, we have it good in comparison for sure
>>
>>109445085
AND, I just fucking said 'y' to this shit
Do you agree to enable tracking to improve the application? [y/N]: y
>>
>>109445113
>I opened that ldg thread that anon linked and was stupid enough to open a few gens. holy fucking shit I'm embarrassed to share this board with those people. really embarrassed.
maybe im a retard. but i liked some of them like this one
>>109445182
>>
>>109445200
I'm not really interested in anything other than those chibi in irl life videos and maybe porn.

I guess also animating game assets/
>>
>>109445210
>chibi in irl life videos
those are cute, except the ones that are like those little comic abuse creatures that used to get posted.
>Porn
there is some there but yeah its not good.
>animating game assests
i dont think this one is even near.
>>
>>109445210
>chibi in irl life videos
penguinslop is the worst viral marketing shit ever
maybe even worst than that six seven shit
>>
>>109432514
>I've been "distilling" K3 and formerly K2.7 into Meta models for weeks now
kek based
>>
>>109445239
>animating game assests
not in my experience, but I haven't tried very hard
>>109445244
I don't want to see the penguin and want my wife specifically
>>
>>109444946
Do you also recommend that link? Is there any "standard" uploader or user? It seems I'm trusting a random model in hugging face.

So there is no uncensored model at the moment... What the hell



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.