Discussion and Development of Local Image, Video, and Audio ModelsPrevious: >>109881361https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GPNeural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Qwen Image 2.1https://huggingface.co/Qwen/Qwen-Image-2.1>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/neo_collage>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
gm saars
>mfw Resource news09/22/2026>CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generationhttps://zshyang.github.io/CoaG>AniPrO: Interpretable Anime Image Provenance Detection via Multi-Dimensional Semantic Reasoninghttps://github.com/YAN-LIU05/AniPrO>AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transporthttps://github.com/51xOne/Alignmorph>Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scenehttps://sunyangtian.github.io/Mira-Scene-web>Rethinking Vision Architectures with Gated Linear Attention and KANhttps://github.com/mehizelali/linear-kan-transformer>SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Modelshttps://github.com/aliathar1401/SK-Stars-shroom-visions-2026>Accurate Motion Estimation with Bézier Control Point for Efficient Frame Interpolationhttps://github.com/SHH-Han/ABC-Inter>Qwen Image 2.1 for intel Mac's with AMD GPUhttps://github.com/haseebeqx/qwen-image-intel-mac>China's Alibaba unveils new powerful chip and ambitious AI model planshttps://tech.yahoo.com/ai/articles/chinas-alibaba-unveils-powerful-chip-071601752.html>Photoshoot: Build a person once, then shoot a whole serieshttps://github.com/ralksta/ComfyUI-Photoshoot09/21/2026>Qwen-Image-2.1: Compact, Efficient, and Unified Image Creationhttps://qwen.ai/blog?id=qwen-image-2.1>Qwen Image 2.1 Prompt Enhancer — ComfyUI Custom Nodehttps://github.com/benjiyaya/ComfyUI-Qwen-Image-2.1-Prompt-Enhancer>Qwen-Image 2.1: ComfyUI Repackhttps://huggingface.co/Comfy-Org/Qwen-Image-2.1>Qwen-Image-2.1 GGUF quantized files https://huggingface.co/leejet/Qwen-Image-2.1-GGUF>Spectrum for Qwen2.1 (>2x Speedup on a 3060 12gb)https://github.com/awdqwdasdg/Comfyui-Spectrum-Qwen2.1>Supra2-IMG: Text-To-Image • 100M Parametershttps://huggingface.co/SupraLabs/Supra2-IMG
>mfw Research news09/22/2026>Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominancehttps://arxiv.org/abs/2609.24215>IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Promptshttps://arxiv.org/abs/2609.24228>An Efficient and Effective Watermarking Scheme for the Protection of the Intellectual Property Rights of Video Generative Modelshttps://arxiv.org/abs/2609.23586>VISTA: Video-Injected Stylized Text-to-Animationhttps://arxiv.org/abs/2609.23817>SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265 $ Single-GPU Acceleration of Visual Generationhttps://arxiv.org/abs/2609.23153>PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraininghttps://arxiv.org/abs/2609.22789>RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modelinghttps://arxiv.org/abs/2609.22947>Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanismshttps://arxiv.org/abs/2609.23658>PETR: Prompt Ensembling with Training-free Routing for Vision-Language Modelshttps://arxiv.org/abs/2609.23600>Rethinking Diffusion Segmentation: When Does It Rely on Its Noisy State, and Does Diffusion Matter?https://arxiv.org/abs/2609.23967>Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generationhttps://arxiv.org/abs/2609.22916>Hierarchical Prompt Learning for Hyperbolic Vision-Language Modelshttps://arxiv.org/abs/2609.24276>Classifier-Free Guidance in Flow Matching: Non-Autonomous Potentials, Overshoot, and Posterior-Mean Controlhttps://arxiv.org/abs/2609.24287>LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mambahttps://arxiv.org/abs/2609.24337
>sdg: schizo image gens>ldg: schizos arguing in text
>>109886072
>>109885340>>109885587still hoping for answers on this
>>109883871This is great, probably the last "free lunch" we're gonna see from H3. Personally this will cut 40 seconds from each of my gens. It makes me excited enough to try out some 1366x768 stuff again tomorrow>>109886017>japs are fucked in the headThere are 300,000 self hating Japanese in Japan that hate themselves because they can't understand and come to terms with why their wartime atrocities were so much more fucked up than anyone else's. Like of course they deny shoving bamboo sticks up Chinese 3 year olds to dilate them so the Japanese army could rape them, who wouldn't? And it's ended up as a unique self-hating ideology, similar but distinct to Brazil's mongrel complexBut to Japan's credit they seem to only rape others mostly (yeah that gang raped for 34 days and murdered girl but that happened like once ever) and their crime statistics are enviable (yeah I know it's weird with most crimes being solved by confession etc but still) so maybe we're the fucked in the head ones where we obviously can't handle some forms of artistic expression but they can This is related to local diffusion models because Comfy Inc is in Japan
What would anon wish for me to generate?
what is the current status on video models? t. hasn't been genning in almost 6 months.
>>109886086Thoughts on qwen image 2.1 deeby? (Sorry if you already voiced them)
>>109886112kinos
>>109886113Minimax H3 is better than Sora 2. Hopefully you have at least 16gb of vram and 64gb of ram>>109886112>What would anon wish for me to generateNon-sexual but interesting POV stuff, like working on fixing a satellite on a space station while in orbit, or being one of the Muslims walking around the black cube
>>109886114I haven't touched it personally. looks like a great edit model based on the stuff people have been posting though
>>109886120>Minimax H3 is better than Sora 2.fucking seriously? how? I have 24 GB of VRAM and 128 GB of ram.
>>109886133unironically yes
>>109886087Set seed to fixed. Change one thing in the prompt at a time so you can see how it changes. Change scheduler and sampler individually to see how those affect the output. Each model has its quirks so you need to isolate the variables by using a fixed seed until you know approximately what prompts work and which ones don't.
>>109886120sorry best i can do is 1girl
Someone else bake before the schizo baker gets banned again
>>109885630I don't think anon realizes how good that specific gen is. >>109886174The eyes are really pleasant.
>>109886141right but as far as the literal prompting how verbose does h3 need to really be? does it need to be as literal and autistically crafted with an llm like krea and qwen or does broad stuff work? to me its felt like a coinflip
>>109886212I don't do t2i h3. too weird and variable. Always first frame ref image. Much easier to control.
>>109886133>fucking seriously? how?They generated videos filling in the knowledge gaps with Seedance 2.5 and cloud video models and it mostly worked out. >>109886133>I have 24 GB of VRAM and 128 GB of ram.You're in for a treat then. Text to video is actually good now. Use the FL2VA model without any image references for text to video. Don't expect faces to look good at resolutions smaller than 720p though>>109886148Wtf bruh why even ask then. At least do pov featuring a 1girl you should still get turned on by that >>109886212>right but as far as the literal prompting how verbose does h3 need to really beYou're done with manual prompting. You are supposed to feed the official prompt guide (not the one with "ref" in the name, that's for the reference model, the other one) to an AI and ask it to make T2VA prompts featuring xyz that follow the structure
>>109886057>>109886064thanks!
>>109886236>>109886242im not doing t2v for h3, im doing ref2v and its still a fucking mess