Discussion and Development of Local Image, Video, and Music Models and SoftwarePrevious: >>109510784https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbo>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Zhttps://huggingface.co/Tongyi-MAI/Z-Image>Qwenhttps://huggingface.co/collections/Qwen/qwen-image>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>Chromahttps://huggingface.co/lodestones/Chroma1-Basehttps://rentry.org/mvu52t46>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
Real bread
>mfw Resource news08/09/2026>Kroma v0.2 — Krea 2 fine-tune (full model) https://huggingface.co/lodestones/Kroma>krea2-turbo-bbox https://huggingface.co/jimmycarter/krea2-turbo-bbox>Kroma v0.2 Quanthttps://huggingface.co/silveroxides/Kroma-Quant/tree/main>Spectrum for Ideogram 4https://github.com/Nif00/ComfyUI-Spectrum-Ideogram4>ClipProj — MiniMax H3 conditioning from a Qwen3-VL-4Bhttps://huggingface.co/NicoLab28/ClipProj-MiniMax-H3>ComfyUI-SigmaSync-LoRA: Sigma-aware model-only LoRA strength schedulinghttps://github.com/capitan01R/ComfyUI-SigmaSync-LoRA>NexusBTA v0.2.44 adds MiniMax H3 supporthttps://github.com/JpAndreBTA/Nexus-BTA/releases/tag/v0.2.44>Experimental MiniMax H3 single-image VAEhttps://huggingface.co/Mamad8/MiniMax-H3-Image-VAE>MiniMax H3 REF2VA w4a8https://huggingface.co/realrebelai/Rebels_w4a8s08/08/2026>Kijai: MiniMax H3 Ref Lora Rank 256 bf16https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)https://github.com/Amduraznak/minimax-h3-fp16-fix>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopointshttps://github.com/fengwang/Cosmos3-Nano-WebUI>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patcheshttps://github.com/charlie12345/R9700AIProComfyUIPatch>MiniMax-H3-Pruned-GGUFhttps://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF08/07/2026>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely localhttps://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha>LIGHTX2V 4-step Turbo Minimax H3 lorahttps://huggingface.co/lightx2v/Minimax-h3-Turbo>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRAhttps://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA>Sage Ready: Local-only installer and readiness checker for SageAttentionhttps://github.com/CosmicFungi/Sage-Ready>Wan 2.2 Animate 2 14Bhttps://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
I'm a bit confused anon, so which model is the best between t2v/i2v and ref2v for h3?
>>109512077You should've used >>109511826
>>109512091>ClipProj — MiniMax H3 conditioning from a Qwen3-VL-4BThis is ultra retarded.
>mfw Research news08/09/2026>MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradationshttps://arxiv.org/abs/2608.00736>Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimateshttps://arxiv.org/abs/2608.03284>Test-Time Curriculum for Open-Set AIGC Detectionhttps://arxiv.org/abs/2608.00559>Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Meshhttps://arxiv.org/abs/2608.00094>Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognitionhttps://arxiv.org/abs/2608.04623>Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detectionhttps://arxiv.org/abs/2608.04394>Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Modelshttps://arxiv.org/abs/2608.05834>EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuninghttps://arxiv.org/abs/2608.01142>DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computationhttps://arxiv.org/abs/2608.01343>Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioninghttps://arxiv.org/abs/2608.00994>MiniWorld: Democratizing the Training of Video World Models from Scratchhttps://arxiv.org/abs/2608.01127>GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compressionhttps://arxiv.org/abs/2608.03517>Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Groundinghttps://arxiv.org/abs/2608.03471>Messages, Not Tokens: Grounded Coresets for Faithful VLM Compressionhttps://arxiv.org/abs/2608.02134>OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Modelshttps://arxiv.org/abs/2608.03812
im gonna find inner peace by buying a 32gb vram card, downloading qwen image 20b and using my hyper specific art style Loras from my favorite artists and goon my brains out
>>109512093they both have different usecases, you can't use references on t2v/i2v
>>109512098Why are you reading the schizo links?
>>109512098Yeah>We finally have an unretarded model with a powerful text encoder that lets it understand what the user is trying to do.>Quick, make the text encoder retarded! It could save up to 2% gen time for the text encoding step that only runs once! This will surely help.
>>109512119Are they not automated?
If I used the quantized heretic text encoder, will Minimax know what sex is?
>>109512132there's only one way to find out
>>109512129No he does it manually by hand, all of his thread rituals are manual endeavors. There's a reason why he's in OP and he does it in multiple threads.
kino, you can do so much with references.The entire video must forcefully match the exact art direction, PlayStation 1 graphics, fixed-camera style, and pre-rendered 2D backdrop appearance of <Picture 1>Hatsune Miku is holding a silver pistol, and running around a room in a mansion, and walks up the staircase.img source is re1 original.https://files.catbox.moe/w5vslt.mp4
>>109512077Why did this cuck of an OP generate a man groping his waifu?
breast thred
/a/ autists are filled with so much hatred. feels good.one of them even picked up on one of ani's boogeymen labels 'wanschizo', probably without even knowing what wan is.
>>109512149that's exactly what I want for their upcoming image model, to use a reference and perfectly transfer it to another image's style
>>109512112well you can use an audio and of course images but of course they don't 'work the same way, the audio tho works like LTX in that you can provide a small sample of a voice and it uses that sample as a refence for the talking in the prompt. And unlike LTX if that is used for someone not known (some voices are bad even if supposed to be the person) then that can be used with people in the prompt it does know
>>109512151We're not using your thread, debo.
>>109512163look man I love my aislop but you gotta stop being obsessed like this
>>109512171>look man I love my aislop but you gotta stop being obsessed like this
i wish i could hook my computer into my brain to give it more memory to generate stuff
>thread is ending>2 or 3 bakes>enter "non-troll bake">walls of garbage links>someone starts complaining about /a/>100 posts in>finally some discussion beginstiresome
holy shit, almost like the game itself minus the gun sounds (could prompt those)The entire video must forcefully match the exact art direction, PlayStation 1 graphics, fixed-camera style, and pre-rendered 2D backdrop appearance of <Picture 1>Hatsune Miku is holding a silver pistol. A nearby door is broken open and a group of human zombies walks through. Miku shoots the zombies with her pistol and they fall down.https://files.catbox.moe/mgz89e.mp4
>>109512163
>>109512188am I in that book?
>>109512185>100 posts in*29
>>109512098what makes it ultra retarded?
https://d.uguu.se/uBArvTUD.mp4
>>109512195who?
>>109512091>>109512101go back to your containment general
>>109512201Can you even run H3?
>>109510922Oh, sorry I missed this. I took a "break" from what I was making to generate something safer to share for you.This is safe, right?https://files.catbox.moe/mu5hz2.webm
Has anyone tried to take the missle knows where it is transcript and gen something with it yet? id be curious to know what h3 does
What are games that really light enough to play while genning
>>109512186yes. the sound often is the weakest. it already seems to have some capabilities to match externally provided sounds to video tho.so perhaps even before finetunes that make it better you could already supply audio made/generated elsewhere.
>>109512244Slay the spire on you igpu
https://files.catbox.moe/dtra7d.mp4 anons... Something went wrong...
>warninghmmm
>>109512255What do you mean? How do you typically float around? Head first? That's dangerous.
>another bake with no collage and bakers shitty gen.Grim times for /ldg/
>>109512276Guess you shoulda baked then, huh?
How many of you guys are using the heretic text encoder for H3?
>>109512250I could even feed game sounds as a reference, there is so much shit this model can do
>>109512280You're only slightly less worse than the rentry schizos.
>>109512276worse, the OP video is like 4 days old https://desuarchive.org/g/thread/109477686#109478380
>>109512293
>>109512282snake oil.
>>109512280>heh.... YOU should have baked then...>NOOOOOO WHAT THE FRICK WHY DID YOU BAKE THAT?>SPAM SPAM SPAM SPAM!!!!!!!!tiresome
>>109512276I gave you 3 minutes and you didn't bake, dumbshit autist.
Anyone ever find out why cumfart only uses like 60% of vram??
>>109512313I wasn't there, and you're pretty delusional if you think it's just 1 person that cares about the collage.Like the other anon said, Why do you insist on not baking with one?
>>109512326I don't care how many people care about it. You're all the same brand of boring autist.I don't give a fuck about your shitty collages and only care about having a non-debo thread active. If nobody else bakes, I do as the failsafe.Don't like it? Cry about it. Learn to realize there are bigger problems in life than /ldg/ not having a fkin collage OP lol.
>>109512326go and join debo in the other thread anon.
>>109512345I can smell the cheeto dust coming from this post
>>109512345Fedora energy
>>109512355RED 40 BABY
>>109512355>>109512358So head on over to the debo thread. What's stopping you?
>>109512355>>109512358>I can smell the cheeto dust coming from this post>Fedora energy
>>109512106good luck, I got two 5090s last year, and now with the same prices I paid, I can get 80% of one
>>109512244>https://slither.io/I'm going through my switch backlog.
>>109512369really wish i could use AI to filter out all basedjak/jak posting. Even frogposting was never this fucking annoying.
>>109512345uh oh melty
voicecloners ww@?https://n.uguu.se/NXNuXaNp.webmstill need to figure out how to get rid of that subtle oscillating undertone. maybe it's just the compressed sample, but not sure.I wonder if you can adjust the audio quality via the prompt.
to the anon earlier maybe someone else already implemented the multi image loader to your satisfaction:https://github.com/Deno2026/comfyui-deno-custom-nodes#deno-minimax-h3-multi-reference-image-loader
>>109512345based
>>109512345>I only care about having a non-debo thread active>Learn to realize there are bigger problems in life than /ldg/ are you a woman or a tranny?
>>109512385>englishdropped
wake up samuraihttps://files.catbox.moe/372yob.mp4
>>109512385if you don't succeed to your satisfaction (i couldn't really, maybe it's early support or maybe the audio model just isn't so close to SOTA) perhaps use another voice cloning TTS https://github.com/diodiogod/TTS-Audio-Suite and supply the audio.
>>109512399eh, someone on /vp/ requested the english voice so that's what i used. normally i'd go japanese myself, and did with this one:https://n.uguu.se/tdFbLerH.webmCan recreate May's original voice very authentically, even with a seductive tone. Looking forward to doing more voice cloning in H3.
>>109512414not gonna lie this video is way betterno idea why the subtle tease of her just opening the top was enough for me but holy by god i want more(also i do prefer may's english purely because of nostalgia)
>>109512399agreed, jp va are just that good>>109512414better
>>109512407I've actually got a decent amount of experience with tts models, and actually downloaded a bunch right before minimax launched and blew my dick up.They're really much of a muchness. For video synchronization there's certainly no reason to use a dedicated tts model over minimax, which is trained to sync with lip movement.I'm sure some nerd could figure out how to optimize it in post running a pass thru Adobe Audition or Tenacity or something.
>>109512379must mean its working if you're getting this angry
https://github.com/jpietek/PenguinBurnerSomething I recommend anyone to do : powerlimit the card to like 80-90% then run penguinburner to undervolt it, performance will be almost the same and the card will be more stable in general.
>>109512437You can undervolt in MSI Afterburner. Why would you install this just to do that?
>>109512441It's for Linux
Does anyone make 3D stereo images?I prompted krea 2 to make "3D side-by-side binocular stereogram stereographic image"The result was...it looked like it was going to work. There was clearly depth to the image but it would vary between being correct for cross-eyed viewing and parallel viewing at different parts of the image. My guess is the only reason it doesn't work is because the data wasn't properly tagged between these two for training.There are programs to convert images to 3d but they are based on lidar which can't do stereo correct transparency, refraction, reflection, etc. Plus finding a one with good inpainting for occluded areas has been a pain. It would be nice if I could just make stuff 3D in one shot.
>>109512163Stay here in your containment thread wanschizo. You can slop and slop and slop to your heart's content. No need to spread like a cancer trying to metastasize
>>109512437No thanks I did it in lact
>>109512441I'm genning on a headless linux.
>>109512452Anon, you do realize you're both stupid AND autistic, right?
Witcher 4 showcase.>>>/wsg/6211190
>>109512427>I'm sure some nerd could figure out how to optimize it in post running a pass thru Adobe Audition or Tenacity or something.but do you need to with all the tools that are available from local AI and conventional sound libs even in that custom node pack? https://github.com/diodiogod/TTS-Audio-Suite#features ... of course use any others but this is really quite nice rn since almost all the best stuff is in there with more or less the configuration you'd want.
>>109512460Someone should make a checklist of all your insults
>>109512431whats working is zoomer fags constantly forcing shit that is unfunny reddit dog shit.
>>109512467Anon, you don't know anything about the dynamics of /ldg/. You're just a low-IQ autist from /a/ who started copying ani's boogeyman labels.There's a rentry link about ani in the OP of this thread. Perhaps you should enlighten yourself before exposing yourself as an even bigger moron ;)
>>109512474Actually I take it back, it should be a bingo card. Yeah, that would work much better.
>>109512480Relax anon, we can all see you're quite stupid.
>>109512462How did you make 44 secs vid ?
this is how we do action in Ugandahttps://files.catbox.moe/vqcm7y.mp4
>>109512487It's stitched. Supposedly you can just gen much longer videos than 15 but i don't have the VRAM for that.
>>109512485Who is "we"
>>109512495This is what was promised to us by mark zuckerberg
>>109512495holy shit this dude is living his best life
>>109512495kys
Any implied sex h3 gens?
>>109512196told you
>>109512495oh my
>>109512499lets see your gens, if you're so great
>>109512495>you can see the woman walk all the way up from the water to the viewer impressive
>>109512507
>>109512493I have a 5060ti 16gb and 64gb ram and do 20 sec gens which are just that bit longer to avoid sped up talking to try to fir it in. But really it is all about the prompt fitting to the length of the gen length anyway, so timestamps would make it easier to do it, the prompt enhancer (or using a another LLM even online) can on the most part do it all for you via a basic prompt
>>109512512>>109512499
>>109512495What Epstein could have been if he wasn't the way he was
>>109512516mad
>>109512499>>109512522
>>109512345>>>/wsg/6211203
>>109512530keeeeeeeeeeeeek
minimax does a good vj emmie clone and I dont even have the best cropped cliphttps://files.catbox.moe/yv4udn.mp4
>>109512530my sides
>>109512532impressive, I need to get in on that
>>109512531>keeeeeeeeeeeeek
Adding a 9 second video reference triples the generation time...
>>109512495this is fucking amazing
i forgot to mention he should be screaming the words, not just screaming.