Discussion and Development of Local Image, Video, and Music Models and SoftwarePrevious: >>109433015https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbo>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Zhttps://huggingface.co/Tongyi-MAI/Z-Image>Qwenhttps://huggingface.co/collections/Qwen/qwen-image>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>LTX-2.3https://huggingface.co/collections/Lightricks/ltx-23>Wanhttps://github.com/Wan-Video/Wan2.2>Chromahttps://huggingface.co/lodestones/Chroma1-Basehttps://rentry.org/mvu52t46>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
gm saars
https://github.com/huggingface/diffusers/pull/14355>33b>26b text encoder>no base model>2 transformer files (one for t2v and another for reference shit)Jesus... I thought Minimax was a serious company for a second
>>109434230How is sulfur gonna make a coom finetune without base?
Reminder the comfyui doesnt want to support GGUF despite scenarios like with these video models benefiting the most from GGUF, since Q4 Q5 Q6 are infinitely better than int4 in quality, and faster than int8 which needs to offload to RAM.
literal whos
>>109434238no ggufnot allowed
>>109434236>How is sulfur gonna make a coom finetune without base?won't be better with BFL, we won't get a flux 3 base model either, it's grim
>>109434230>>2 transformer files (one for t2v and another for reference shit)boooooooooooooo
>>109434249can distill, but gonna wait till we see what the best of the bunch ends up being. Also LTX could end up redeeming themselves as they said they would always release the base already
>>109434230if they only gave us the guidance distilled version, does that mean that the API is using the full version?
>>109434267with jews, you win
Chinese Culture
>>109434279literally lol
>>109434230>>109434249If flux 3 dev ends up the same size or like within 20B id rather undistill that (if that's even possible...)
>>109434238>Q4 [] infinitely better than int4 in qualityCitation needed, especially for convrot int4>and faster than int8Also citation needed. In my experience accelerated quant with offloading still beats non-accelerated quant that fits into VRAM, and by big margins too.Provide evidence for the claims you are shitting up the thread with (You won't because you are a raped schizo).
I hope it knows Michael Jackson
>>109434293at this point it's more likely to get LTX 3.0 than being able to undistil a video model kek
>>109434212>mfw Resource news08/01/2026>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generationhttps://explorative-modeling.github.io>EU to get power to enforce rules on AI starting todayhttps://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786>FameGrid Auto Color for ComfyUIhttps://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui07/31/2026>One-take Creation, Flexible Referencing: Introducing Seedance 2.5https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5>ReToken: One Token to Improve Vision-Language Models for Visual Retrievalhttps://github.com/avaxiao/ReToken>RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generationhttps://github.com/liuxiaobo66/RefineSVG>ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generationhttps://github.com/H-EmbodVis/ROAD>PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interactionhttps://physomni.github.io>DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localizationhttps://github.com/anonyme610/dinolizer>Inline Studio v1.2.6 - Flux 2 & Minimax H3 APIhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.6>SAM 3.1 Multiplexhttps://huggingface.co/Sparknight/sam3.1-int8-int4-convrot>Suno Loses AI Copyright Lawsuit to German Music Rights Society GEMAhttps://variety.com/2026/digital/news/suno-loses-ai-lawsuit-gema-1236825010>AI labels to be compulsory on authentic-looking content under EU rules https://www.theguardian.com/technology/2026/jul/31/ai-labels-to-be-compulsory-on-authentic-looking-content-under-eu-rules07/30/2026>AnimeGen: AI Models for Anime Video Generationhttps://huggingface.co/collections/aidealab/animegen
Dedistilling is a massive fucking cope anyway.It completely destroys model quality. You need massive, millions of images/videos massive finetuning to recover from that. There are like only 3 people currently training at that scale: nutbutter, td_russell, a certain furry lolcow
>mfw Research news08/01/2026>VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generationhttps://arxiv.org/abs/2607.23472>Layering Virtual Try-Onhttps://arxiv.org/abs/2607.22924>What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Featureshttps://stevencylu.github.io/PeakPatch>S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Imagehttps://arxiv.org/abs/2607.28164>Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillationhttps://arxiv.org/abs/2607.24731>FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attackhttps://arxiv.org/abs/2607.27800>UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routinghttps://umi3d-project.github.io>UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Modelshttps://arxiv.org/abs/2607.23373>Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steeringhttps://arxiv.org/abs/2607.26411>LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detectionhttps://arxiv.org/abs/2607.25962>MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformershttps://arxiv.org/abs/2607.28589>What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restorationhttps://arxiv.org/abs/2607.28526>Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradationhttps://arxiv.org/abs/2607.24440>Same Predictions, Different Reasons: The Effect of Quantization on Model Explanationshttps://arxiv.org/abs/2607.22872>Unifying Adversarially Robust Model Experts in Vision-Language Modelshttps://arxiv.org/abs/2607.27897
>>109434312most people don't use undistilled models, and most people don't train models, and ostris will have a training adapter for it tomorrow afternoon anyway.seems like a completely pointless fud.
>>109434324None of them ever trained a video model, much less a massive one btw.It's joever if no base.
>>109434248>>109434328Whatever this artstyle is, I hate it
>>109434329STOP SPAM
>>109434334He can't yap about creating the most ebinest training adapter ever as much as he wants but the loras won't be in good quality.
>>109434334>ostris will have a training adapter for it tomorrow afternoon anyway.the same ostris who made an adapter for flux.1, and it worked like as>most people don't train modelsthe guy that made sulfur was waiting for Minimax so that he could train it as well, he won't be able to do it because those chinese fucks don't respect us enough to give us the base model, say what you want about the LTX jews, but they gave us a base model at least
>>109434345*can yap
>>109434322>>FameGrid Auto Color for ComfyUIDo you even check what you share on your spam? This is just a vibecoded slop custom node made to advertise some grifter AI site
POST GENS
>>109434360the fuck is this schizo doing here, get the fuck back to your containment zone
>>109434357so don't use it? not everything is for you
>>109434345>>109434347>you can't train z-image turbo! literally no one has every trained z-image turbo, and z-image turbo is bad! again, seems like a pointless fud.
>>109434357It's a vibecoded bot that probably reads top posts on r/StableDiffusion and searches a few key words on arxiv.The malware spammer faggot in fact, doesn't check what he spams.
>>109434364>you can't train z-image turboyou literally can't yeah, name one succesful serious finetune of zimage turbo
>>109434300You'd think in a thread about AI people like you would know that you can simply use AI to verify whether a statement about AI is true or not.
>>109434338it's ludo doe
>>109434371he could just put it in a rentry instead of three huge walls of text if sharing news was the goal but its obvious its not
>>109434364Lots of ZIT loras commonly give body horror. They are also nightmare to combine. (Inb4 esoteric block swap combo cope) It's needlessly difficult to train, I sincerely doubt you trained any, if you call concerns about training on distilled models a "fud".
can h3 do this?
you already scrolled past anon. the news can't hurt you anymore
>>109434376What are you rambling about faggot, post your proof or get the fuck out.If you mean you asked bunch of loaded questions to Gemini Flash to confirm your priors, you are Indian.
>>109434363just like when you shared literal malware for weeks?
>>109434360based gguf supporter
>>109434392Don't see you providing proof to the contrary. Not-so-subtle way of admitting defeat, bud.
>>109434379i call general meltdowns anytime a new model pops up a fud. for 4 weeks straight "someone" shat up this general all day everyday about how bad krea is and how goated zit was, literally the best model, super easy to train, best finetunes, super based! now zit is slop trash that is untrainable because a new video model is distilled.
>>109434397>getting mad at text you can simply ignore on 4chinz
>>109434406>how goated zit wasfor inference yes, nothing was said about training, maybe you're just too stupid to understand the nuance
total namefag holocaust
>>109434406its always the same guy. i dont know why you retards keep replying to him since his typing style is obvious
>>109434406ZIT mogs Krea was one schizo spamming the general while posting next to no gensAnd ZIT was in fact difficult to train.How are these supposed to be contradictory?
>>109434423>And ZIT was in fact difficult to train.what the fuck are you on about? its the model that instantly had a fuck ton of loras basically day 2
So apparently Q8 > FP8 according to this guy using Wan2.2 and LTX for comparison. And use Q4_K_M for efficiency. Any lower is just coping. https://www.youtube.com/watch?v=octq1gJ0dkM
>>109434434>lorasare you the anima cope anon as well?
>>109434404>Make nonsensical claim that goes against established facts>Asked to provide any evidence whatsoever to back up said the claim>Nuh-uh! It's your job to provide evidence on the contrary to disprove my claims that I provided zero reasons for you to believe that they are true chud!I can take my time to provide the numbers but why would I bother to do so with a schizo or troll? I know int8 runs faster than q quants on my system. You are just some random fag on the internet and nowhere as important as you think you are.
Does the cover mode in ace step 1.5 actually do anything?
>>109434446are u implying the word "training" u used meant not what everyone does (train loras) but the 0.00000000001% of people are doing, full finetunes?and even for those why would people do full finetunes of a model that is turbo, after base was announced for it?
whats the difference between Q8 and INT8?
>>109434465Only chuds use GGUF.
>>109434446now you are just being silly
>>109434434>Lots of ZIT loras commonly give body horror.>They are also nightmare to combine.The difficulty was training loras that didn't suffer from these issues due to training on a distilled base you pedantic midwit.
>>109434465Q8 is an illegal format DO NOT use
>>109434458At 0.5 Cover strength the outpost songs and singing is similar to input song and singing. So it does a cover of a song, duh?
>>109434469>chuds*chads, sorry for the typo
>>109434465Q8 is unaccelerated and slow.Int8 is accelerated (on Turing and newer) and fast.Naive int8 has lower quality than q8.Convrot int8 has higher quality than q8 (And also faster).
>>109434465Q8 trims layers, int8 converts to integers. The CPU has a better time with ints so it's better when offloading but if you can fit the whole thing in vram them who cares
>>109434475>Lots of ZIT loras commonly give body horror.someone claiming one of the easiest models to train, zit, was actually 'hard to train' is enough to spot a retard trolling but now with another low iq troll lie like that ill just take your concession.
>>109434485>Convrot int8 has higher quality than q8citation needed.
>>109434485>Convrot int8 has higher quality than q8can you show a grid or something? It always looks worse for me
>>109434485>Convrot int8 has higher quality than q8*loud incorrect buzzer*
>>109434487so i should use int if i am using mixing gpu and cpu memory?
>>109434499can we see the images?
>>109434504yes unless you are at capacity even with that then you would use Q4
>>109434505nyo
>>109434455>i can but i won'ttee~hee same here :3
>>109434499Link to repo?
I can do a direct comparison with Krea2 between int8 convrot and q4_k_m gguf
>>109434512thanks
>>109434528No one would dispute int8 would have higher quality in that case.Tell us how much RAM and VRAM you have and then post inference times for both though. (Total and per step time.)
>>109434499Liar faggot.https://github.com/BobJohnson24/ComfyUI-INT8-Fast/blob/main/Metrics.mdOnly Chroma and ZIT had better quality on GGUF. HiDream, Qwen, Klein, Anima all had better quality with Convrot.That's why you didn't put the link.RAPE AND GAPE ALL GGUF FAGS
>>109434445Yes and GGUF Q8 being better than FP8 is the case on almost any model quantized from a more precise one - video, image, llm.
>>109434561>Only Chroma and ZIT had better quality on GGUF.and you don't care about that subhuman?
>>109434535int8_convrot (original turbo model)50s inference time on a 3060 12gb with 24gb of sys memory and a i7 67008 steps, euler ancestral, simple scheduler, cfg 1.03:4 ratio at 2mp (1216x1664)6.2 seconds per iteration
>>109434567Nyt you prompt separately for those stickers?
>>109434580how much memory did it use?
>>109434571Idc about Chroma, and I consider ZIT obsolete at this point but nevertheless bigger issue is that it doesn't debunk my claim (FACT) that convrot int8 has better quality q8 on average. And your cherry picked cope post was pathetic.
>>109434583They don't have their own bboxes. They're part of the bike bounding box prompt and high level description
>>109434580And Q4?
>>109434567nice
>>109434580Q4_k_m gguf of original turbo model (from realrebelai)same settings1m9s inference time at 8.7 seconds per iteration>>109434587int8?all 11.5gb of vram and about 50% of sys memory with 16gb paged (not max)
>>109434419Cope. Your model sucks.
I said 5 minutes per second, I was off by a factor of 2.5no, the other direction
>>109434617and how much did q4 use?
>>109434617Yes, even q quant half the size is still slower due to need to dequant.Even if you were spilling into system memory with int8 it would still be faster (up to a point)
>>109434607Being against ZIT is like being antisemitic, transphobic, racist, and homophobic, all ZIT posters are trans women that you should respect, sweaty.
>>109434617q4_k_m gguf uses about 9.5gb of vram and 35% sys memory during inference and 65% after vae decode.
not that ever seems to care any more (or ever did) but what's the word on flux 3, no ETA?
>>109434655DOA
God I love blonde fennecs, I will make one my wife.Fucking hate silver cats though, will behead with the swiftest of ease
>>109434609No loras?
>>109434655I may know things that others do not but I cannot say exactly
>>109434655Within the next few weeks and months because they're "rigorously safety testing" it.
>>109434655Expectations are low, it'll be huge and it'll be pozzed to hell and back
>Minimum requirements for 480p video on this model is a 3060 with 12GB vram, 32GB of system ram and a good nvme SSD. We tested generating a 5 second (124 frames) 480p (864x480) video on this system and it took a bit less than 9 minutes end to end (20 steps). I can pretty much guarantee it will also work on 8GB vram too but we did not test that.480p 5 seconds 20 steps is half a minute per step. This is almost certainly 8 bit quant too, I dread it if this was int4, so we will need a good step distill lora for sane inference speed/resolution on mid range cards.
>>109434680low res will be garbage regardless, it always is
>>109434637>Even if you were spilling into system memory with int8 it would still be faster (up to a point)Absolutely fucking not. The moment you spill even 10% of the model into RAM your speed craters unless you are running 9000 MT/s quad-channel RAM on a Threadripper setup. Spending a little compute on dequanting is a cheap and quick. Moving assloads of data over PCIE is not.
>>109434655Black Forest Labs seeing how much they can mindrape, censor, and lobotomize a model while it still outputs something ok is hard work, anon, it takes time.
>>109434680he really needs to stop posting ads in the main sd reddit
>>109434212>>109434212 [Posting this here in case if the thread got archived.]This is my battlestation setup right now about Stable Diffusion. [sd-webui-forge-neo]Are there new updates I need to download for SD? Or new .safetensors files for new and better default generations? Or is the IllustriousXL still the best at this time?
>>1094346803090 is 3x the speed of 3060 at least so im good
>>109434687That's not how any of this shit works retard. The slowdowns are fairly linear to the percentage offloaded. It's a cool bonus of how fused ops work in diffusion. What you said only applies to auto-regressive models like LLMs.BF16 ZIT with circa 2gb offloading was still faster than q8 ZIT fitting fully on my system.
>>109434680just text to video?
>>109434702>slowdowns are fairly linear to the percentage offloaded>linearlmao retard
>>109434668My own lora trained off 1000+ amateur photos from 2009-2011. An attempt to try take the overly cinematic edge look off Ideogram. Picrel is with no lora
Minimax is gonna be shit.LTX Is still the GOAT. The ZIT of video genning.
>>109434707Your schizophrenic sense of humor does not make the statement incorrect.
>>109434687the int8 model is over 12gb, so I am definitely "spilling" over to system memorybut it's still faster and better quality than the q4 gguf
>>109434707Hi ani
>>109434655since when BFL ever deliver? it's gonna be a bloated piece of shit like flux 2 dev, and because they never give base models we won't even be able to make serious finetunes of it (if it turns out to be good and small enough by some miracle...)
>>109434718>better qualityBy what metric? Your two images are practically identical, and what differences there are might as well be seed variations (because empty latents will be different with different models even with the same seed value).
>>109434718that is what we were saying the whole time. Q8 has higher quality, int8 better speed, q4 absolute lowest
>>109434698but you'll probably want to bump the resolution so that's at least a 1.5x slowdown
>>109434680>Don't be scared to try our default templateWill it be 70% useless bloat like all the recent templates
>>109434760yes, but it's all hidden in a single group node