Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109474113https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>>109475019animal abuse :(
>>109475019Any way to make VAE decode faster ?? It took betweek 30 seconds and 1 minute :(
>>109475037those blacks are working hard. it's not abuse
>>109475043>Any way to make VAE decode faster ??good news for you anonhttps://github.com/Comfy-Org/ComfyUI/pull/15334
good morning sars, turbo is out? is it good, we back?
>>109475037>mfw Resource news08/05/2026>Inline Studio v1.2.62 - Minimax H3 Lora training still onlyhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3https://huggingface.co/Kijai/MiniMax-H3-TAE>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inferencehttps://github.com/6somehow/DAC-SPADE>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generationhttps://github.com/yizzz927/CAPE-T2V>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusionhttps://github.com/jd-opensource/JoyAI-Video-Edit>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttps://github.com/YangYangGirl/ParVL>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diethttps://huggingface.co/JamesZar/OliveGemma-3B08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES timehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editinghttps://github.com/IntMeGroup/MIEScore>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videoshttps://rathgrith.github.io/PeCA>Kandinsky WM 1.0: A family of models for Physical AIhttps://github.com/kandinskylab/kandinsky-wm08/03/2026>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>109475028typical retard judging a model from only a single output >Still trying to make a case for wan over h3.youll learn its pointless to try to beat that kind of thing into anonif you know its (being any model) is good then you know itll eventually proliferate
>mfw Research news08/05/2026>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrievalhttps://arxiv.org/abs/2608.03120>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Planehttps://arxiv.org/abs/2608.03422>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformershttps://arxiv.org/abs/2608.03082>Can T2I Models Draw from the Right Frame of Reference?https://arxiv.org/abs/2608.03357>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editinghttps://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0>Self-Supervised Representation-Guided Generative Dataset Distillationhttps://arxiv.org/abs/2608.03218>Latent Reward Registers for Diffusion Preference Alignmenthttps://arxiv.org/abs/2608.03929>UniWorld-Design: From Pixel Generation to Layer-Native Designhttps://arxiv.org/abs/2608.03971>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Bindinghttps://arxiv.org/abs/2608.03708>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editinghttps://arxiv.org/abs/2608.03059>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blendinghttps://arxiv.org/abs/2608.03269>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfoldshttps://arxiv.org/abs/2608.03135>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Modelshttps://arxiv.org/abs/2608.03057>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMshttps://arxiv.org/abs/2608.03112>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuninghttps://arxiv.org/abs/2608.03875>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understandinghttps://qwen-3d.github.io>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardwarehttps://arxiv.org/abs/2608.03649
>>109475055saar no cumfy saarport yet
>>109475052I hope comfy is paying kjGOD well
>>109475055>>109475066samefag
how do i merge the lora into the model so i dont have to keep reloading it?
>>109475071I actually made every post in the past 3 threads, saar.
>>109475019we are so close to working realtime fmv games.
8/6 and the default 12/3 sigma shift makes the subjects act like crackheads, like completely off the rails adhd mode. together w/ the turbo lora (the HF one) at 8 steps multires and a boomer prompt.fuck that shit, bypassed.
>>109475066what's this then? https://huggingface.co/QrusherZA/H3_Turbo_ComfyUIjust tried it, it works, dunno how good tho
>>109475052>>109475092>I put swapped the current VAE for this one, but my vids just returned as blackis it because of sageattention?
>>109475090first I see of this saar, thank you may blue goddess give you many handjob
How make a the sex with the big boobed anime woamn with "H 3" ?
>>109475101wait until brahmin sir wakes up and spoonfeeds us
Is camera shake broken in Minimax? Whenever i've prompted for an unstable camera/shaky camera, it's more like a vibrating camera with very fast jitter. I just want a camera that's like a person holding a handheld camera. The documentation is useless at providing information on this.
>>109475123using sigma shift? it seems to speed up certain things
>>109475123Prompt for Michael J Fox holding the camera.
fucking hate comfyui so much, everytime i have to update for new models my old reliable workflows break and i have to spend an hour fixing it
>>109475123werks for me
Tom cruise as a vampire goes hard nglhttps://files.catbox.moe/hx5ka2.mp4>>>/wsg/6208955
>>109475132>using sigma shift?Yeah.I guess I'll try it without sigma.
>>109475135beg ani to make a new release
Minimax is really Seed dependent. Dont waste your seeds if you got the good one
>>109475153wrong
>>109475153probably not correct
Still using sigma shift with the turbo lora?
>>109475052>vae decoding speedupwow thanks for saving me the 2s after my 15 minute gen
>>109475147Buffy is 5' 4"Tom Cruise is a midget.
>>109475175Movie magic
>>109475153even if that's true, that's a good thing. nobody wants a boring model
>>109475174the cherry on top is more artifacts! :D
>>109475174Yeah the VAE is super important I don't know why that's the thing people want to butcher for nominal speed increases.
>>109475153If youre not chase the gap that exists you no longer a racing genner
>>109475193people are retarded, you have to assume most things posted went through at least 4 different quality raping settings
Turbo lora fucking RAPES the audio
>>109475206it does, it's an unfinished lora, we have to let the poor lad finish the job
>>109475206true. but it's WIP, right?
Welp, i tried running H3 Int8 on a 3060 12gb and 24gb of sys memory.Was paging hard and thrashing my ssd at 64% system memory usage, so not sure if I want to keep using it
bored.flux3 waiting room
>>109475239>24gb of sys memorydon't you need 32?
Can i finally make on the spot cnc videos of abi shapiro or is that still a pipedream
>>109475239I have less system ram than you and have no issues. you must be on winblows.
>>1094752442x8, 2x4
>>109475244I was told 640k was enough for everybody.
>>109475244no, I have a 3060 and 16gb of ram. works fine for me including the non-pruned version.
>>109475256>16gb of ram. works fine for meneat.
so no video extension ala ltx yetthat's fine, I can wait
https://files.catbox.moe/f9ncl2.mp4haha cool
H3 is """"usable"""" on a wide range of hardware. It just depends how low your standards are
>>109475250Yup, that sounds about right.Been struggling with comfyui not using system memory and using the pagefile instead for a while now. Thought i fixed the issue with krea2 but now it's happening again...
>>109475239...why does it use the ssd? my ram is at 80%. is comfy fucking retarded?
>>109475283I mean its less about standards and more about paitence. its not too bad for me. 10 minutes for 10s? not ideal but hey it works well
pokemonGOD anon, what are your settings?
>>109475256Oh yup. I bet if you check task manager, the drive with your pagefile (assuming windblows) will be getting thrashed during inference.If not, I would like to know what black magic you are using to run the int8 model with 16gb.
I feel like people blame comfyui when its likely just windows being a giant steaming piece of shit>>109475307>If not, I would like to know what black magic you are using to run the int8 model with 16gb.no black magic I just use linux. it just works I think mostly due to dynamic vram which everyone shits on here or some reason. I use zram swap
I hope that how you guys pick up girls.https://files.catbox.moe/kadplw.mp4
>>109475328>masterpiece
so would you dudes who tested minimax h3 say that this is good enough to produce real kino?like do you think this is good enough that people can produce their own short movies, tv shows and such?also any other AI tools you'd recommend that could help with such?krea2 model also seems decent enough that it can produce something consistent enough.
>8gb vram>64gb ddr5>loonixCan I run it?
>>109475328score_9
>>109475322ok? why can't comfy make it work properly on piece of shit windows then
>>109475301>...why does it use the ssd?I'm assuming it needs 32gb to swap the whole model in memory, so 24gb means it has to constantly move layers in and out between ram and disk
>>109475345>like do you think this is good enough that people can produce their own short movies, tv shows and such?I can see that yes >>>/wsg/6208850
>>109475354you answered your own question so why are you asking
>>109475351>8gb vramwhy?
>>109475322>just windows being a giant steaming piece of shit100% Windows fault. it's what made me completely drop it. It just can't handle heavy memory workloads. Like it's broken at it's core.
>>109475358i mean i assume it doesn't happen on wan2gp
>>109475356maybe a good technical showcase but painful to watch
Idk if it's something I'm doing wrong but this node rapes FLF generations.
>>109475365this is what happens when the entire OS is scaffolded off of code written in the 90s
>>109475359laptop, please understand saar
>>109475365windows is made for normies. it doesn't need to be anything else other than that.>>109475368blame whatever you want. the model, comfy, or anything else other than the root cause of the issue. no skin of my back.
>>109475372it does, I prefer to wait a bit more and use Spectrum but the quality is there >>109474992
>>109475379off*anyways gn saars
>>109475368wangp is more comfy than comfy tbqhdesu
vid gen has made me realize i am not very creative
>>109475390fact
some say the text encoder being censored or uncensored doesn't matter-- are there any comparisons? Idk what to believe
>>109475351Maybe you can but it's not gonna be worth it. IMO genning is only fun if it's fast. Otherwise it's a slog and a waste of time.
>>109475403yes has been done to death at this point man...
>>109475396im just making my ocs moan n shit
>>109475403"Uncensored" text encoders are a meme and do literally nothing, has always been the case, stop giving attention to retards.
>>109475422proof?
It doesn't know that a penis is supposed to be rigid and instead makes it move like it's some kind of sea cucumber.
>>109475430for my eyes only, chuddie
>>109475429I'm something of an uncensored text encoder my myself, reggin.
>>109475396we got so used on using shit models that can only do 1girl that we don't know what to do anymore when we finally get a model that can do anything