Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109752182https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>mfw Resource news09/07/2026>Sol-H3 Speed-of-Light MiniMax-H3 on an 8× NVIDIA B300 Blackwell Systemhttps://nvlabs.github.io/Sana/Sol-Engine/Sol-H3>MageTrail: Danbooru/E621 Full-Finetune of MageFlow 4B https://huggingface.co/RicemanT/MageTrail>Learning 3D Editing without Paired Supervision via Generative Prior Distillationhttps://github.com/thiamine128/PriorEdit3D>UniMate: One Unified Model to Animate Diverse Skeletonshttps://linzhanmou.com/unimate>VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognitionhttps://github.com/FlamieZhu/Vicinal-Consistency-Alignment>WorldSculpt: Generating Compositional Worlds from Grounded Videoshttps://alaya-lab.github.io/WorldSculpt>AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memorieshttps://zunwang1.github.io/AnchorWeave>ComfyUI-InpaintCanvas: Krita-style inpainting without leaving ComfyUIhttps://github.com/DenRakEiw/ComfyUI-InpaintCanvas>ComfyUI-Ref2VA-VSA: Ultra-Fast Character Video Generationhttps://github.com/Kablex/ComfyUI-Ref2VA-VSA09/06/2026>ComfyUI-Viggle-Animate-H3https://github.com/Saganaki22/ComfyUI-Viggle-Animate-H3>uncomfymcp: MCP server for generating images on ComfyUI from a chat client.https://github.com/aschet/uncomfymcp>MiniMax H3 Semantic Bridgehttps://huggingface.co/speach1sdef178/MiniMax-H3-Semantic-Bridge09/05/2026>Intern Lumina U2: Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generationhttps://internlm.github.io/InternLumina-U2>Musk’s xAI loses court bid to block Minnesota's AI ‘nudification’ banhttps://www.reuters.com/legal/litigation/musks-xai-loses-court-bid-block-minnesotas-ai-nudification-ban-2026-09-04>Add Sparse Attention nodehttps://github.com/Comfy-Org/ComfyUI/pull/16072>MiniMax-H3 FL2V 8-Step Motion Enhancerhttps://huggingface.co/rzgar/minimax-h3_fl2v_8Step_motion_enhancer>Qwen3.8-Flash-Next-NVFP4 https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4
>mfw Research news09/07/2026>Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matchinghttps://arxiv.org/abs/2609.04283>ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Featureshttps://arxiv.org/abs/2609.04649>RefDiT: Local Attribute Guidance in Reference-Based Image Generationhttps://arxiv.org/abs/2609.04976>Importance-Aware Low-Rank Distillation of Diffusion Transformershttps://vislearn.github.io/SVDtrunc>Measured Sliders: Learning Continuous Controls from Differentiable Image Measurementshttps://arxiv.org/abs/2609.05234>PAPT++: Risk-Aware Adversarial Tuning and Generation for Single Domain Generalizationhttps://arxiv.org/abs/2609.04837>AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editinghttps://arxiv.org/abs/2609.04709>Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generationhttps://arxiv.org/abs/2609.04282>LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Renderinghttps://arxiv.org/abs/2609.04939>Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrievalhttps://arxiv.org/abs/2609.04800>WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editinghttps://arxiv.org/abs/2609.05171>FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencodershttps://arxiv.org/abs/2609.04276>Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inferencehttps://arxiv.org/abs/2609.05275
>>109757421>>109757414
>>109757430https://files.catbox.moe/sdeh8d.mp4
https://files.catbox.moe/6khu5e.mp4
>>109757430You've never noticed it generated the wrong type of keyboard?
Any YTP Minimax LoRAs?
>>109757571This is from the qwen prompt but you have to know the esoteric lore that debo was a failed DJ before going full NEET so it stabs extra deep.
>>109757589>failed DJ>gens black charactersdebo is 100% black. no way this guy is white.
I think I fucked up. the SaveVideo node in comfyui isn't saving any output on the server from clients genning on the local lan (with --listen). filename_prefix was the same for all of them, from the workflow.there wasn't a warning in the logs that it couldn't write or anything, and genning from a browser on the host works fine. is this because an output directory wasn't set at launch?
>>109757608the SaveVideo node ignores your output directory and write to the comfyui root for some reason
>>109757597No he's a sheltered Californian NEET
>>109757613What the fuck
>>109757617whats with the wheelchair
>>109757631Don't click random catbox links anon
>>109757639A anon asked if he was disabled due to the amount of time he spends in the general and I made a wheelchair gen to make fun of him and he flipped the fuck out and threatened to kill me while posting base XL gens at default resolution.I then posted it after the migration and he also flipped the fuck out so now I just post it because it hurts him.It does make sense seeing how often he's here and how he acts.
>>109757615oh jesus. it looks like it created a single file in the base output directory and then overwrote it instead of incrementing the filenames. I guess the clients couldn't read the path on the host?how do I fix this
It's kind of sad both the rentry schizos have threatened physical harm to me over bantz. Not once have I ever called for violence to either of them but they are so fragile gens and simply questioning their behavior pushed them that far.Too sensitive for the internet desu
>>109757641>>109757631For him to delete it, it must have been bad. What was it?
>>109757716Just something that doesn't belong on a Blue Board. The other one is still up, lazy jannies.
I am using ref 2 video , i have done one video with background music now making part 2 , what is the best way to make the audio so it is seamless? or is this better after i made the videos i need
>>109757716wrong board.
3 fucking days of failed gens.I should go back to LTX at this rate
>>109757744Based. Model?
>>109757673use a different node ig?
https://files.catbox.moe/amorcq.mp4
>>109757869I was waiting for him to slay the monster but this is fine, too.
>>109757761fixed it, I'm just a retard. the prefix was set to %date:yyyy-MM-dd%-user/ instead of %date:yyyy-MM-dd%/user.I changed it on the host early on and didn't ask anyone to refresh.
https://files.catbox.moe/h60qz5.mp4 https://files.catbox.moe/g1p627.mp4
>>109758018beautiful..
life must be tough and hard being a vramlet and ramlet in q3 of 2026. i hope your hardware is ready for flux3.
>>109758071
>>109758078>All the VRAM>20B PrunedGenuinely, why? Speed? Fitting the whole thing in one GPU? I'm curious, not criticizing.
https://files.catbox.moe/844nyk.mp4
>>109758083as opposed to what?
>>109758112As opposed to the full un-pruned complete 33B model.
>>109758117why would you want to use that?
>>109758125It produces consistently higher quality results.
>>109758133no. the people that pruned it found that the weights removed had very little to do with the generation side of things. you would only want to use the unpruned model if you are trying to do further training
latest miniconstruct is ready but i'm literally being filtered by git semantics
>>109758154I've been lied to.
>>109758180note that if you are saving seeds, the pruned one will make a different generation from the same seed that you used on the 33b one
>>109757250https://rentry.org/neo_collageI updated anon's collage script. It should now work better on lower end systems and smoother on higher ends, 5 minute max video length, less jank, a fresh interface, and some other changes. The original script recorded canvas playback in real time while this one decodes source frames, composites them, and encodes explicit timestamps. It should no longer drop frames and slow FPS.
>>109758292For editing and auditing purposes, please please @require mediabunny from a CDN and make the userscript code unminified.
>>109757250OK, so this is probably an obvious idea but has any company actually done the following? Take an already trained model and a diverse data set of various images (photos, paintings, illustrations, ...) and have the AI prompt itself to recreate the image from scratch (i.e. not I2I) using itself. Given that its prompted output image can easily be compared mathematically against the original to determine how close it came to matching it, this should allow it to quickly get better like neural nets are wont to at these kind of optimization tasks. By the end of this automated training you should have a model that can take any input image and give you prompts that, whether they make "sense" to a human or not, result in output that matches the original.So what you get out of that is an AI prompt engineer or sorts that can look at picture A and help you figure out the kind of prompts that would get picture B to look like it (style-wise or composition-wise or whatever).
>>109758292The op is proud of his low effort slopscript and will probably ignore this
>>109758489*of sorts
I like Krea 2
>>109758425Done.>>109758492OP and the author of the script in OP are clearly two different people.
Brainlet here. I'm getting decent 5 second clips with coomer loras on H3 so now I'm taking a shot at extending them. Do I need to write individual prompts for every single five second video extension?I'm using wangp, I assume comfy could let me automate this 100% for a 45 second clip without me writing these in between prompts
>>109758526>Do I need to write individual prompts for every single five second video extension?You can do either. Separate prompts are recommended when the action changes over time; one reused prompt works well for an ongoing activity. You have a setting, "How to Process each Line of the Text Prompt". With "All the Lines are Part of the Same Prompt", the same prompt conditions each window, with overlap carrying context forward. Good for someone walking through a forest, dancing, a long sustained landscape shot, that sort of thing. With "Each Paragraph Separated by an Empty line will be used for a new Sliding Window of the same Video Generation", you still paste everything into one prompt box and run one job, but you package it differently. One complete H3 prompt per window, no blank lines inside a window’s prompt, one blank line between window prompts, include ALL required H3 fields/sections in each window. Sliding window generation just automates the process, you can reproduce the same results without using sliding window generation at all, but it saves a lot of effort.
>>109758519Your script is broken at least on Tampermonkey due to the way it bundles requires. You need to add a leading semicolon on L24.;(() => {minor nits:- I'd suggest making the "collage -" button green to make it visually more obvious the image was selected than the + turning into a -- I'd suggest making selected gens have a green background in the review UI to make selected more obvious than a ticked checkboxOtherwise it looks good to me, I deprecated https://rentry.org/ldgcollage_v2
;(() => {
>>109758593>Your script is broken at least on Tampermonkey due to the way it bundles requires. You need to add a leading semicolon on L24.Copy the new version I just changed. I ran into the same issue it should work now. >minor nits:Good ideas, the first one I was thinking of as well. I'll add those soon. Maybe more stuff.
I have to fix the aspect ratio stuff as well. Hold on.
>>109758601>>109758607Cool, no rush.
>>109758526Comfy or Wan can automate that 100%, and both are fucking terrible at it. They work by rewriting your prompt, that's how it can be automated, you're letting the clanker take your prompt and decide how to break it up into several smaller ones, which they are very bad at in my experience. A frontier model with a real H3 prompting skill can do it, but I don't expect most people to use Astra to write smut. Though I do recommend it, because she's a dirty whore.
Good news everyone: MiniConstruct 0.4.0 is out. It is a MONSTER of a release.https://github.com/InsertSpice/MiniConstructWhat's new:>New 'creative interpretation' toggle - you can adjust how much creativity the LM is allowed to have when generating the prompt.>Extensive benchmarking and strengthening of the system prompt, ensuring the model can understand exactly how to incorporate references into the output prompt>this now includes reference sheets for subject identities, saving you from using multiple images to support all angles of a subject>Projects have officially transitioned from IndexDB to the filesystem.>New story planner: rather than generating single prompts, plan and generate a larger story across a sequence of generations, each with their own incoming state, outgoing state, camera shot/intent.>Don't want to fill in all that boilerplate? No problem, just click a button to have the LM generate it>Iterative revision features for story control. Don't like how a certain generation is composed? Just tick it's box and revise.>ComfyUI integration is now all systems go. You can now use MiniConstruct as a front-end for generating videos in ComfyUI.
idc
>>109758648how does connecting to comfy work? does the output just get routed to the string node?
>>109758639Just go binary, final destination.
>>109758620Changed some of the layout and rendering, minor nits, and removed the check boxes all together. Just click on the cards to select - you can still click the thumbnails to preview. Aspect ratio accepts different formats btw but I think that logic is still weird. Other than that, the newest version should be g2g. I'm open to other changes in the future. I might make more updates.
>>109758620is this downscaled?
>>109758706You export the Comfy workflow as (API), load it in MiniConstruct comfy settings, then bind the assets to the comfy nodes from the exported API workflow. Then you can generate directly from miniconstruct so long as you binded correctly.
>>109758292Here is a test of every candidate in this thread (the duplicates are from anons posting catbox links of the same video) done on my low spec laptop. As you can see, no dropped frames, stable FPS, no weirdness. Pretty good.
>>109758746You caught me, smart anon
anyone else got OOM with new model sparse attention node?
>>109758585>>109758644thanks im tinkering some moreim finishing up a 30 second clip and ive noticed that artifacts are appearing in the later gens, like little white spots or blocky artifacts, so after 20 seconds it looks a little weird. how do i fix/denoise these, im gonna try to start over at the 10 second mark
>>109758781WTF? Did Mika-chan lose weight? What are those gens?
We're going to have to change the name of the general again soon, because now there's new ways to make AI videos and that's using LLMs with Blender harnesseshttps://litter.catbox.moe/2k4so2qzwa3bg0j6.mp4Vidrel was Astra from some dude on /vcg/, but local Chinese models are only 3-6 months away so there will be a(nother) new way to locally gen passable video very soon that isn't strictly diffusion
I thought i had good sense of humor but I got no (you) from my gen
>>109758946Sir, this is a schizo general, if you're not here to discuss catjack or animanon then we're going to have to ask you to leave.
>>109759237that isn't new, it's basically just a digital plate and people have been doing it since wan 2.1.
>>109758648Neat, might give it a try, once I'm too tired to deal with groks occasional bullshit. Wish I had some second PC I could host the LLM on, but I guess switching models in VRAM isn't too slow.>>109759392Happens to the best of us.
>>109759533Show me someone making a video like that with WAN 2.1 in a single prompt you retarded fucking gorilla, at least spend the tokens to invoke the tool call to look at the video if you're going to burn compute replying to me
>>109759702>in a single prompt i like how you had to move the goalpost just in case. could you explain how you think that video was created?