Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109478939https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg
>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
First for based Will Smith gen
>mfw Resource news08/06/2026>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Modelshttps://zhouhyocean.github.io/uniworld-view>OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Filmshttps://xin1u.github.io/OminiVR_PAGE>DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Modelshttps://github.com/Zhong-Chenchen/DIVE.git>Multi-View Face and Gesture Animation with Dynamic Gaussianshttps://dfki-av.github.io/MVFGA>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbothttps://empaava.top>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generationhttps://github.com/Aoko955/Flash-VAED08/05/2026>Inline Studio v1.2.62 - Minimax H3 Lora training still onlyhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3https://huggingface.co/Kijai/MiniMax-H3-TAE>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inferencehttps://github.com/6somehow/DAC-SPADE>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generationhttps://github.com/yizzz927/CAPE-T2V>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusionhttps://github.com/jd-opensource/JoyAI-Video-Edit>ParVL: Parallel Scaling and Expandable Compute Allocation for MLLMshttps://github.com/YangYangGirl/ParVL>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diethttps://huggingface.co/JamesZar/OliveGemma-3B08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md
>>109480220kino
based
>mfw Research news08/06/2026>When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusionshttps://arxiv.org/abs/2608.04820>HelloWorld: Enabling Socially Interactive Characters in Video World Modelshttps://arxiv.org/abs/2608.05070>OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editinghttps://arxiv.org/abs/2608.05049>ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routinghttps://guoxu1233.github.io/ContextMaster>STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Modelshttps://arxiv.org/abs/2608.04887>Simile Understanding in Text-to-Image Models: An Evaluation Frameworkhttps://arxiv.org/abs/2608.04750>ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generationhttps://arxiv.org/abs/2608.04436>CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Modelshttps://arxiv.org/abs/2608.04302>Rethinking Pixel Mean Flows via Interval Denoiserhttps://arxiv.org/abs/2608.04818>Persistent Object Narratives for Token-Efficient Video Language Modelshttps://arxiv.org/abs/2608.04866>Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Modelshttps://arxiv.org/abs/2608.04349>Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roleshttps://arxiv.org/abs/2608.04483>Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Modelshttps://arxiv.org/abs/2608.04454>When does training on downscaled images yield the same gradients?https://arxiv.org/abs/2608.04448>Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detectionhttps://arxiv.org/abs/2608.04935
>>109480214https://old.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/
>>109480228finally
>>109480247kek
>>109480236>>109480242thanks!>>109480230kys
>literal tranitor bake looooole
theres way too much coomer stuff in the data set. random penis in the video., what th heck?
i hope you all behave yourselves
https://streamable.com/idwah1
I don't get it, Gen works great at 0.4mp, bump to 0.7, prompt adherence jumps out the window.
>>109480220Real thread>>109480276>>109480276>>109480276
>>109480259Please don't dig into him too hard, if you do he'll start crying and pissing himself
>>109480236>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
>>109480280> trolling outside of /b/
nah JSID bruh its over
>>109480287>ID 190, how are you today? >i am very good sir, thank you for asking. and how are you today?>i'm great, so i was wonde->you can open bobs?>excuse me? >show milk! open bobs!
>>109480220>>109480236>>109480242Fuck off debo
>>109480276>>109480276>>109480276
Ignore debo thread.Everyone here:>>109480276>>109480276>>109480276
non-debo bread>>109480276>>109480276>>109480276
the turbo lora makes better audio for war kinos surprisingly. still can't make alarm sounds sadly. something is wrong with the h3 dataset
>>109480378>still can't make alarm sounds sadly.such a strange limitation. There is so much free stock audio of alarms
who is ready for flux3?
>>109480392i might have to use the video2audio feature with ltx since it makes really nice realistic alarms and other combat stuff
>>109480412isn't there foley models that can make cool sound effects or do you just want to skip over editing?
Was there some prompt guide for H3?
>>109480378The parachute was pretty bad lol, I think they only trained it on gliders.
>>109480436i don't want to put that much work into it. if you want it to be realistic, then you have to do a lot of editing on the sounds themselves so they match the acoustics of the video
>>109480455what do you mean? the forward momentum would be from ejecting out of a fast moving jet. i do agree that it isn't really realistic if you see how a real ejection seat works. i tried to prompt something similar but it kept glitching out
>>109480244Someone go to reddit and tell me if they say anything interesting.
>>109480550we got:>thank you cocksucking>cumfart cocksucking>industry plugsnothing of note
the other thread feels like a bot that loops through the same convo
>>109480327
>>109480622What one?
>>109480634based
>>109480647I couldn't quite get the eyes right and gave up lol. close enough
>>109480657have you still never got into vidgen?
>>109480729I might play around with h3 when a decent turbo lora drops. I have a 4070 w/ 32gb ram so I'm just barely on the edge of workable hardware
>>109480763that should be enough with the 8int diffusion and text encoder if you enable disk memory
>>109480763just leave genning and go goon or something.Once I see h3 got the idea semi-right, queue like 10 clips and fuck off.when you're back, it's done.