Smell Your GPU EditionDiscussion and Development of Local Image, Video, and Music Models and SoftwarePrevious: >>109430870https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbo>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Zhttps://huggingface.co/Tongyi-MAI/Z-Image>Qwenhttps://huggingface.co/collections/Qwen/qwen-image>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>LTX-2.3https://huggingface.co/collections/Lightricks/ltx-23>Wanhttps://github.com/Wan-Video/Wan2.2>Chromahttps://huggingface.co/lodestones/Chroma1-Basehttps://rentry.org/mvu52t46>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
What the fuck is this collage?
why no kino in collage?
Blessed thread of frenship
so in 14 hours it comes out?
>>109433028this goes hardmuch love from iraq
>>109433015what the fuck is that faggollage
>>109433028>so in 14 hours it comes out?yep, glad that I will sleep soon, I know I'll wake up with a kino presenthttps://modelscope.cn/models/MiniMax/MiniMax-H3
>>109433015>>109432905>>109432910>>109432959Once again I'm asking you for a wallpaper that I can use in normie space without getting shot or cancelled.Pic rel is current pape.
>>109433045you're gonna get hilariously trolled and owned and made fun of by normies if they see that shit on your computerit's so blatantly AISLOPPA.
another trollbake, yawn
sometimes you need a strict baker to clean anons palette
>>109433066open your mouth so i can clear your palette
>it's even better than acestep at making musickekhttps://litter.catbox.moe/00k8tl4h39e0p7nd.mp4
>5 1girls, asian, singget new material, white incels.
>>109433100now the question is whether you can get consistent music if you string together a bunch of clips
>>109433045
>>109433124>gaming chair
>>109431232
>she doesnt own a Gaming(tm) brand chair
>>109433145>turned an oven dodger into a rice cookeri dig it
Which krea mix are the cool kids using these days
how is the chroma krea2 lora?
>>109433242I recommend it
>>109433205flashback to when we had a schizo spamming george floyd
>>109433028then in 48 hours the comfy nodes will be ready
>>109433258Fake AI
>>109433260krea2 knows about George Floyd too
>>109433260I'm sure he'll come back to try H3 out
>>109433258He has the smile of a king and gentle person
It is a shame they are releasing the model on Sunday, the day set aside to worship the Lord. (I almost said the Lord's day but that would have been a mistake as every day is the Lord's).
>>109433318chang cares little of God and has historically released on the sabbath
what's the first thing you will generate when h3 drops?
>>109433332ur mom lol
>>109433332Did they say how big the model is?My poor 3070 still has life left in it
>>109433332pornography of ur mom
Skateboarding is always a good test
>>109433332some yugioh dueling kino
>>109433332ur mom lol but probably princess peach undressing or something basic to start with
>>109433340no, i don't know why they are keeping it a secret
>>109433364>we know it can run on an 8gb 3060>but still no parameter/file size very strange indeed
>>109433369we cant know the filesize until the moment of release due to the quantum degradation and bit entropy
>>109433344can you make one where she dismembers a luddite?
>>109433376you scare me a bit not gonna lie
>>109433384her farts probably STINK like SHIT also elongated torso
3090 status: readyfeed me h3
>>109433369Wait I have a chance?
>>109433369hmm, 38gb for fp8god knows on the TE, maybe 20gb
>>109433340nobut we've had quite viable VRAM->RAM swap offloading for most models so far, most likely that works again in some way or another.if the model was completely too large for almost all consumer cards, I think they'd have warned.
>>109433015Haven't been keeping up, will MiniMax H3 be able to do porn or is it useless?
>mfw Resource news08/01/2026>FaceFusionYu: ComfyUI integration of the FaceFusion 3.5.4 processing enginehttps://github.com/thenotrealuser/facefusionuy>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generationhttps://explorative-modeling.github.io>EU to get power to enforce rules on AI starting todayhttps://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786>FameGrid Auto Color for ComfyUIhttps://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui07/31/2026>One-take Creation, Flexible Referencing: Introducing Seedance 2.5https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5>ReToken: One Token to Improve Vision-Language Models for Visual Retrievalhttps://github.com/avaxiao/ReToken>RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generationhttps://github.com/liuxiaobo66/RefineSVG>ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generationhttps://github.com/H-EmbodVis/ROAD>PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interactionhttps://physomni.github.io>DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localizationhttps://github.com/anonyme610/dinolizer>Inline Studio v1.2.6 - Flux 2 & Minimax H3 APIhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.6>SAM 3.1 Multiplexhttps://huggingface.co/Sparknight/sam3.1-int8-int4-convrot>Suno Loses AI Copyright Lawsuit to German Music Rights Society GEMAhttps://variety.com/2026/digital/news/suno-loses-ai-lawsuit-gema-1236825010>AI labels to be compulsory on authentic-looking content under EU rules https://www.theguardian.com/technology/2026/jul/31/ai-labels-to-be-compulsory-on-authentic-looking-content-under-eu-rules07/30/2026>AnimeGen: AI Models for Anime Video Generationhttps://huggingface.co/collections/aidealab/animegen
if ur gpu has never genned 24/7 for at least 2 weeks, lower ur tone when speaking here
only oldfags remember flux3
>mfw Research news08/01/2026>VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generationhttps://arxiv.org/abs/2607.23472>Layering Virtual Try-Onhttps://arxiv.org/abs/2607.22924>What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Featureshttps://stevencylu.github.io/PeakPatch>S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Imagehttps://arxiv.org/abs/2607.28164>Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillationhttps://arxiv.org/abs/2607.24731>FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attackhttps://arxiv.org/abs/2607.27800>UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routinghttps://umi3d-project.github.io>UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Modelshttps://arxiv.org/abs/2607.23373>Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steeringhttps://arxiv.org/abs/2607.26411>LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detectionhttps://arxiv.org/abs/2607.25962>MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformershttps://arxiv.org/abs/2607.28589>What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restorationhttps://arxiv.org/abs/2607.28526>Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradationhttps://arxiv.org/abs/2607.24440>Same Predictions, Different Reasons: The Effect of Quantization on Model Explanationshttps://arxiv.org/abs/2607.22872>Unifying Adversarially Robust Model Experts in Vision-Language Modelshttps://arxiv.org/abs/2607.27897
>>109433440>>109433446instead of spamming this trash every time with duplicate info that most dont care about, make a rentry and link it at the top and stfu
>>109433440>>109433446fuck off debo
>>109433440>>FaceFusionYu: ComfyUI integration of the FaceFusion 3.5.4 processing engine>https://github.com/thenotrealuser/facefusionuy>404
>>109433439probably not extensively on release. chinese company expecting public attention.
>>109433482weird. they must have pulled itI'll remove the link, thanks
https://github.com/Comfy-Org/ComfyUI/pull/15210lmao, the PR for minimax is here, but he closed it, he probably made a mistake, let's see>It uses Qwen3-VL-32B as the encoder (50 layers of it)>There is a message in this commit: "model must be a split MiniMax H3 transformer">Minmax H3 is about 28B-33B dense model >50 layers.>Main DiT blocks:>4 Attn, 5376 x (56 heads, 128 dim)>2 MLPs, 5376 x 14336.>AdaLN 6×(2688×5376×3)bruh...
predict how big the shitstorm tomorrow will be. i already expect anons to upload shitty ltx gens claiming theyre h3
>>109433504I can't run that wtf
>>109433504yeah I confirmed the sizes with Claude
>>109433509ever heard of krea?or qwen imagethat, again
>>109433504yay local
>>109433504My 3060 (8gb) is going to crush this shit
>>109433504nah that's too big I'm out, back to flux 3 waiting room
just got hit by the hardest oom yetshit shut down my computer
krea2 boys eating good :)https://civitai.red/models/1055425/ilulu-miss-kobayashis-dragon-maid-illustrious-krea-2?modelVersionId=3189920
>>109433531Is that dynamic enough for you, pal?
>>109433531ltx in comfy makes my pc bluescreen, no idea why
>>109433536kek
>>109433532look at that fucking trigger
>>109433536lmao
>>109433504Get multimodal'ed, Flux 3 dev will be as big too btw!
>>109433332definitely porn.
>>109433504that's too big, I doubt sulfur will make a finetune of that, would be too expensive
>>109433553sar please understand we must having use trigger
>>109433504I can probably run q4 with 24gb vram and 128gb ramwonder how that will be
>>109433504
>>109433586>q4the quality will be so ass though, I kinda wanted the model to be small enough so that I could run it on int8 and get that nice x2 speed increase
i have hope
>>109433504"the Z-image turbo of videos" they said, it's gonna be small and optimized they said...
>>109433553it's almost like people read no tutorials on how to train a lora. Might as well just forego the lora and just use the trigger as prompt.
is comfy still giving out 300 credits per month? how much is a h3 video?
>>109433600>"the Z-image turbo of videos" they saidWho?
>>109433504>28B-33BLooks like LTX and Wan will never die kek.
>>109433332A girl getting punched in the face.
>>109433602nope and they upped the cloud price
>>109433504LOL BODIED YOU RETARDS. only the brownest of localkeks fail to see it
>>109433609Wan 2.2 is 14Bx2
>>109433332>what's the first thing you will generate when h3 drops?nothing, I can't run that big boi after all, oh well >>109433504
>>109433124Workflow pls
>>109433615yeah, that means you can put only 14b to your memory, unload it, then put another 14b, you can't do that when it's a unique 33b model
>>109433504If I were them I wouldn't release anything, what's the point? who can run that??
>>109433622why haven't they figured out how to load groups of layers at a time?
>>109433614>>109428169>he meant you can use ur 3060 to render the nodes 2.0 ui at 30 fps with browser accel while the video generates in comfy cloud
>>109433504Why is everyone dooming? The only thing that matters is the VAE compression ratio. If it's the same as LTX, then the model is almost as fast as LTX. You stream the weights off RAM / SSD using dynamic VRAM (which LTX also requires on most setups).
>>109433596even at that quant it would still be leagues ahead of ltx and wan considering how big of an improvement it is
>>109433614Comfy is so funni, next time a 100b model will appear he will say it can be run on a gtx1600 (thanks to their MAGIC optimisation which is... putting everything in the fucking cpu) lmao
>>109433614the browns are the only ones who will have a problem. they will run to youtube or reddit and download pardeeps retarded as fuck workflow with 40 custom nodes, and a special seedvr upscale polish pass at the end to really slop the tits off of it. then they will complain that they are ooming "their" h100s.
>>109433634yep that's what i'm thinking so far
Reminder the comfyui doesnt want to support GGUF despite scenarios like with these video models benefiting the most from GGUF, since Q4 Q5 Q6 are infinitely better than int4, or int8 with offload.
>>109433632>You stream the weights off RAM / SSD using dynamic VRAM (which LTX also requires on most setups).dynamic vram doesn't work, I always OOM with that shit, I think I'm gonna take a break, ComfyUi is pissing me off, and the model we have so far have all been disappointing
>>109433504Are 5090chads safe or ...?
my pagefile is ready
>>109433636Going to have to start measuring /its in hours
>>109433647I don't think so, if it's a 33b model then on int8 it's probably 34gb just to load it, now you account the memory to make the video and you overflow your VRAM space
>>109433657back to slopping on wan
>>109433504looking at the code, it's probably a non-distilled model, how many steps would we need? 30? jesus...
>>109433402is this a k2 lora?
Let's not start doomposting. LTX 2.3 requirements look hefty but it's still pretty speedy on modest hardware
Can tomorrow get here faster please?https://files.catbox.moe/sxv4wt.mp4>>109433674It will be 1/2 the speed of LTX but about the same size.
>>109433674this. the fudders don't want you to have hope for the new age of kinos
once again, local hardware stagnation continues to hold local back. these requirements are quite small for a 2026 model. 96GB cards should be the norm by now, but there is a coordinated effort to kill off consumer hardware to force reliance on centralized API models.
>>109433504>you need a RTX6000 to fully load this model on 8bitkek, I'll pass, that model is good, but I'm not spending 13000 dollars for this
>>109433622yes you can retard. Its called offloading. Comfy does it automaticlly.
Its about LTX size model wise, but it has half the temporal and spatial compression so it will be about half as fast. It will be worth it though.
>>109433697>Comfy does it automaticlly.it doesn't work subhuman, his automatic offloading shit always OOM my card, Comfy has no idea how to make a good memory managment system, I used to use multigpu to manually do that and control that shit but I can't anymore because it doesn't work with dynamic vram anymore, FUCK THIS PIECE OF SHIT OF A SOFTWARE, AND FUCK YOU
>>109433712uhh, chuddie, stop using UNSUPPORTED GGUF NODES!
>>109433712stop being poor. You need 64GB+ ram for local video. Vram 12GB+
holy meli
>>109433504I really didn't expect the model to be so big, I thought they were smart enough to understand that 20b is the absolute limit, do they really believe there will be a lot of people that will be able to make loras from a 33b model? lmao
>>109433712wan2gp chads can't stop winning
>>109433720I have 64gb of ram, the ram isn't the issue, it always completly fills my vram space and I fucking OOM crash, "automatic offloading" my ass
Kroma RL, seems like he is actually doing a RL himself this time. Seems to massively increase qualityhttps://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.1-rl-mild.safetensors
>>109433726then you are doing something dumb cause it works fine for me. I can do full fp16 wan2.2 easily without enough ram to fit both / te at once
>>109433722Why should they give a shit? Doesn't matter for API inference powered by B200 or whatever the fuck. They release model's as ads basically. If people can't run it despite being open weights and have to paypig for API, that's even better. They both farm cred and still get money this way.
>>109433720>>109433726u need 96gb+ for video if you dont want to touch the disk
>>109433684>96GB cards should be the norm by nowyeah but we don't live in a normal world, we are at the mercy of Nvidia's monopoly so...
>>109433733>If people can't run it despite being open weights and have to paypig for API, that's even better.what?
>>109433727His schizo training might actually be salvaged this time if he is going to do proper RL, but I wouldn't get my hopes up.
>>109433684there's very little any of us can do about that
he is testing a RL atm at different settings
>>109433504BFL please don't be as retarded as them, please save us...
>>109433657There is no reason not to run Q4 these days. It's like 90% of the quality of Q8. Especially in cases where "quality loss" is abstract or subjective like in image diffusion you won't even notice the difference.
>>109433727Does anyone in his furcord here know the difference between rl and rl mild btw?I am assuming more and less aggressive rl, but a proper clarification on what these precisely entail would be nice.
>>109433749regular left, full RL mid, RL at lower strength right
>>109433749>the skin is even more plasticgood job kekestone!
>>109433727apparently there's another one that is just "rl" one hour earlier.>>109433743he basically always managed to train a checkpoint in the end
>>109433710># frames 17k+5 <-> latents 5k+2, 16x spatialFrom the PR. It's a weird ass VAE that maps 17 frames to 5 latent frames, with 16x spatial compression. So somewhere in between the compression of Wan (4x8x8) and LTX (8x16x16), but closer to LTX. It will be several times slower than LTX but probably still faster than Wan per step.
>>109433759are you blind. The RL has far far better skin
>>109433752>Especially in cases where "quality loss" is abstract or subjective like in image diffusion you won't even notice the difference.Q4 literally stops listening to your prompts but go off king
>>109433671Which part?
>>109433710>Its about LTX size model wisenot at all, LTX is 23b, that model is 33b
>>109433750>filenameDesu it would be fine by me if they also give us a kino size distill like Klein again.I am not sure if they would though.
>>109433773no the TE is. People are guessing on the actual model
>>109433769that image is 2 years old
>>109433752>>109433769I bet int4 convrot mogs it while running at least double the speed.
>>109433778>no the TE is.the te is 26b, it's a truncatured version of qwen 3 32b (50/67 layers) >>109433522
left regular, mid full RL, right RL at lower strengthIts a massive improvement
>>109433779And nothing changed about Q quants.They are ass and dead end for diffusion.
>tfw nochekaiser881 going full hard on krea2.tick-tock anima bros...you guy is moving on hehehe. https://civitai.red/user/nochekaiser881/models?periodMode=published&period=AllTime&baseModels=Krea+2
>>109433774how long did it take for them to release klein after flux 2 was released?
>>109433786These comparisons would be far more meaningful if you also include images without any lora
>>1094337912 months
>>109433769Out of all of them, Q4 and Q5 are the only ones who adhered to the prompt. Everything else put her in Tokyo and gave her dark skin instead of light black (grey).
>>109433752holy vramlet cope
>>109433786Wow all three are slop !
>>109433791A month I think?
>>109433793the left one
>>109433769anything to do with AI is way too jeeted at this point to take anything at face value. >>109433786what a fucking mess
>>109433797Sorry I only have 32G and I refuse to offload for marginal gains.
>>109433774>>109433791the problem is that when they made that flux 2 announcment, they already said on their blog that they were gonna make a klein version, the flux 3 blog only has "dev", we're fucked lol
>>109433786The middle one is legit amazing. This seems to have fixed krea's skin texture issues
kek, aged like wine
>>109433683Why do people keep trying to realismmaxx Krea when Ideogram is right there
>>109433808its over
>>109433816ideogram knows a fraction as much / is boring nsfw wise. Krea can do booru stuff with realism style
who does this fudder work for? he does his localkek spam literally every time a new model releases which deprecates the api ones
>>109433789>the model that already knows all the characters doesn't get as many character loras>guy known for training character loras is training them on a model that actually needs themas you can clearly see this means that anima is shit and a dead end
>>109433808>>109433818dev can be of resonable size though, dev 1 was 12b, dev 2 was 32b, I doubt they'll gonna go for ~30b again, they got clowned so hard they probably learned their lesson
>>109433826i don't see any localkek spam?>every time a new model releases which deprecates the api onesoh that must be why, because that never happened
>>109433801Add kroma lora without the lr into the mix too then
H3 is gonna be amazinghttps://files.catbox.moe/pclaeg.mp4
>>109433834klein deprecated nana banana and now these new models will btfo seedance
>>109433504I always refuse to go under 8bit for diffusion models, so I think I'm gonna pass on this one
>>109433786>SD1.5 vs realityDreamshaper1.5 vs perfectbodyMysticAnime1.5
>>109433848>klein deprecated nana banananow that's a quality bait
>>109433861cope. we will be getting videos while you get fired from your call center because no one cares about your fud anymore
the chinese really don't care lol. Its not gonna take much to train this into a hentai modelhttps://files.catbox.moe/q5pdbt.mp4https://files.catbox.moe/sxv4wt.mp4https://files.catbox.moe/wz7x8g.mp4https://files.catbox.moe/7yr24t.mp4
>>109433848true, klein is the top ranked image edit model in the world..on the plastic safetymaxxed leaderboard!
>>109433868I'm sure the finetuners will be delighted to spend tens of thousands of dollars to train a model that can be correctly run only by 3 people.
Never mind he is doing some incomprehensible bizarre nonsense again.
>>109433877those are just different ranks it looks like
>>109433860It brings me great paint to see trainers with such poor aesthetic sensibilities. It's all so ugly.
>>109433848what even happened with nbp? for awhile there was the "saar you can use to make ai influcers!"i feel like now the only time i hear someone mention nbp is in this thread saying it's actualllllly still super sota!
>>10943387630B is easily trainable. I was ready for a 100B. Its called offloading.
>>109433877peak pseudbabble slop. actual competent people would just called it "test version 2"
>>109433876nah. This should train much faster than LTX did and on 16GB vram with offloading. LTX had to be taught from the ground up since it knew so little
>>109433891>Its called offloading.by the time you finish your training we'll be on Flux 7 already kek
>>1094339041 block on device vs all blocks on device is a 5% slow down with DDR5.https://files.catbox.moe/mp08fi.mp4
>>109433889>what even happened with nbp?still the best model, I don't know why people glaze GPT Image 2 so much, it's slopped and the noise is so uncanny, it's so noisy how can people not notice that??
>>109433904at least we'll have krea chroma partial epoch 2 around then (512x512 hires training run)
>>109433258Happy for him he totally deserves :)
>>109433910this is for lora training btw. Its like a 20% slowdown for full finetuning. Still not much of a deal
>>109433847ayy lmao
>>109433877I lurk the chroma discord. He is truly doing bizarre nonsense again. His "RL" is training 2 models off of the base, one on good data and one on bad, and then taking the difference and applying that delta to base. He made some schizoposts about how that's contrastive learning and that's all RL really is (it's not).
>>109433504I guess I'm stuck with LTX and Wan for the rest of my life lol, looks like all the new video models will be giant one, oh well, I think it's a sign that tells me that this hobby isn't for me anymore.
>>109433912after looking up some recent nbp gens, i'm less than impressed. it looks like shit compared to ideo and krea.
>>109433932>nbp looks like shit compared to krea(You)
>>109433927"just train on the entirety of civitai and put it in the negative prompt" logic
>>109433927that works well. You can do RL training in either direction. Doing both is even better.
>>109433927Lmao, thanks for the info anon.Another month, another doomed kekstone failbake.
>>109433939it's cool that you can use your phone to gen stuff, but that doesn't make it good.
>>109433504I miss the days when I was optimistic about the future of diffusion models, Z-image turbo showed us that you could make a small but powerful model, looks like they didn't get the memo, they catched the LLM virus, stacking moare and more layers is all you need!
>>109433927>>109433945stop responding to your own posts
why do people say the russians fixed anima
>>109433951they managed to finetune it to produce good art instead of vectorized slop
>>109433950
>>109433956where did you see the outputs from their model
>>109433504they should do like Kimi K3, train their model on 4bit, so that they can go for big model but it won't be memory expensive for inference, still doing some bf16 in the year of our lord 2026 is sooo lazy
Wait, is H3 a all in one model? qwenVL as a video + audio + TE? So its 32B all together?
>>109433504Do we know if it's distilled or not?
>>109433968no, H3 is video + audio, and then there's a separate TE that is qwen 3 VL 32b
>>109433968
>>109433968No LMAO.It's using 32b version of qwen3vl as text encoder.So in total probably like 60 billion weights.So like 120gb when all weights are in bf16.
>>109433970all will be revealed soonâ„¢. looking forward to seeing what ostris says about how easy it is to train.
>>109433644shit like this makes me want to root for an alternative, but there's nothing unfortunately...
comfy said 32GB of ram + 12GB of vram would be enough
>>109433982wonder if that is with int4 convrot or something though
>>109433982>comfy said
>>10943398210 minutes for 480p + 5 seconds kek
Reminder for chuds itt.
>>109433982comfy also said that GGUFs are a meme so I don't tend to take this retard's opinions too seriously
>>109433964it was shared in a previous thread, it's in a telegram link
>>109433988worth it if it one shots 5 seconds of kino
>>109433994GGUFs ARE a meme. Just offload and int8 convrot is both faster and higher quality
>>109433982to scroll around a 9-node workflow at 10FPS maybe
>>109433998>Just offloaddoesn't work, it OOM, and this retard killed multigpu (the only node that lets you manually offload) with his dynamic vram meme>int8GGUF has 6bits, 5bits, 4bits, and they're good quality, int8 doesn't solve anything, and you know that, stop being disingenuous for a second will ya?>higher qualityQ8 is superior
kino transparency>>109433995I'll check it out, thanks.
>>109433982who gives a fuck about what Comfy says, he's literally paid to shill and implement those models, of course he's gonna find excuses
>>109433982sounds great, did he elaborate for what? surely it's not so good that it's essentially enough for 25s at 2k?
>>109433998ggufs are small, thats a big boon when you are minmaxing a new workflow and trying to do other shit on the side.
>>109434004it does not OOM for me meaning its something on your end / a custom node of yours fucking it
how many of you faggots are in lodestone's furry discord? everything that happens there just happens to appear in this thread minutes later like >>109433749>>109433786
>>109434013>it does not OOM for megood for you, but your experience isn't anyone's experience
>>109434016we keep our finger on the pulse of the local AI community
>>109434018its anyone who does not have a bad custom node it seems
>>109434004>int8 doesn't solve anythingi don't agree with the other anon that gguf are useless but int8 convrot is actually quite obviously quite effective for quantization too.
>>109433504more info herehttps://github.com/huggingface/diffusers/pull/14355
H3 is 26B
>>109434030it's not an unified model, you have one transformer for the t2v and i2v process, and another transformer for the reference process, that's really disappointing, I expected a unique model that does it all
>>109433982Comfy hyped the fuck of SD3 and look where it is right now
it's so funny seeing so much FUD about a model that isn't even out yet.If the model is good people will use it, if it's bad they wont. there's nothing else to it.
>>109434034no>>109434030>One packed token sequence carries text, conditioning, audio and video rows through a shared 33B transformer
>>109434042Levels of kino seen since
>>109434046>If the model is good people will use it, if it's bad they wont.Flux 2 dev (32b) and HunyuanImage 3.0 Instruct (80b) were good models, and they are dead, try to guess why
>>109434048*not seen since
>>109434042iirc there were also <added safety> reasons why it sucked MORE at release than a week beforeeither way this doesn't mean hardware specs are completely wrong?
>>109434030>>109434041looks like it's guidance distilled, it's like flux dev lol
>>109434053"Good" isn't just about the output.
>>109434053They were not good both were sloppa as hell even anima do a better job
>>109434062doesn't really matter considering most models run at CFG 1 anyways now
>>109434062Then... How do users prompt to remove unwanted bullshit from a gen?>inb4 you don't.
>>109434062https://github.com/huggingface/diffusers/blob/e1b518dfd5e390e7ba09a79a1d39fe1c6cb52dc1/docs/source/en/api/pipelines/minimax_h3.mdyep, confirmed, it's a guidance distilled model>>109434079it's a bad news for the trainers though, there's no base models, good luck finetuning that, we all know how well it went when we tried finetuning flux 1 dev kek
>>109434084NAG
>>109434084>Then... How do users prompt to remove unwanted bullshit from a gen?>>109433975
Alright, I made my own Kroma comp, since anon wouldn't do as I wanted.Just so that you would evaluate properly: Snow White is supposed to make a weird expression and the woman is supposed to lick red lollipop seductively, so it undoes krea's suppression of expressions, kinda.
nag for vid wen
>>109434086>it's a bad news for the trainers thoughostris will be training it tomorrow. may as well wait for the final verdict.
>>109434062video models have to be distilled otherwise they would be slow as fuck
>>109434099already exists
>>109434096>Just so that you would evaluate properly: Snow White is supposed to make a weird expression and the woman is supposed to lick red lollipop seductively, so it undoes krea's suppression of expressions, kinda.why would you post that instead of a pastebin of the prompts? are you retarded? also your outputs are burned to a crisp.
>>109434102>video models have to be distilledI don't mind turbo models, but you need a base model for the training, that's why LTX has a base and a distilled version
>>109434102distills should always be loras. Doing otherwise is intentionally sabotaging finetuning
>>109434084the English language? if you want no talking say silent, no hair bald, no women OP etc
>>109434113You can go Klein route and just release both
>>109434123or LTX
>>109434112So it's a big model and you won't be able to make a serious finetune from it? Sounds like a great deal!
Lets see what Flux / LTX does then.
why are you worried about finetuning h3? it doesn't need to be finetuned if it has good prompt adherence
>>109434156>it doesn't need to be finetuned if it has good prompt adherenceyou know it won't know coom
>>109434156>it doesn't know horse cocks and gaped assholes!
>>109434156these vramlets can barely run the model and they are crying about finetuning, what finetuning? some guy who downloaded some porn videos and kinda trained a shitty model?psa, videos are good as your gens
>>109434156>you can run that 33b goyim!! just trust the dynamic vram magic!!>you don't need those pesky negative prompts goyim, guidance distilled is fine!>you don't need to train the model goyim, what's the usecase for that??smells like a lot of cope ngl
>>109434147That's not quite a thing between the people that rent more srs hardware for tuning, and the people that just have more system RAM to do more offloading ramtorch style.
>>109434167>>you don't need those pesky negative prompts goyim, guidance distilled is fine!what do you mean? NAG works
>>109434168>That's not quite a thingit is, guidance distilled models are impossible to finetune, we tried with flux.1 dev, no success
Ostris will make an adapter to train it. Worry not.
>>109434173isn't z image turbo guidance distilled?
>>109434178he already tried on flux.1, it was called un-distilled flux dev, he didn't go far with that>>109434180z-image turbo is guidance distilled + steps distilled, so yeah, that one is also completly impossible to finetune, that's why people wanted Z-image base so bad
Krea2 is nice.int8 convrot works well on a 3060 12gbgenerated at 1080p and upscaled 2x using wan 2.1 upscale vae in just over 1min
>>109434184i think you are a little out of the loop lol
10 mins on 3060 was with 30 steps btw. Step distill should make it much faster later
>>109434195>10 mins on 3060 was with 30 steps btw.480p, 5 seconds kek
>>109434195That's completely irrelevant, what's suspicious is that he didn't say anything about speed on actual real gpus like 4090 or 5090 which tell me all gonna run like shit because ram swap regardless.
>>109434212>>109434212
>>109434186yeah I really like it. closest we are going to get to style freedom like the sd1.5 days.
>>109434211"like shit" is a bit of an overstatement for how the ram swapping usually worked on video models. slower yes, but typically maybe twice as slow, not 50x sloweri guess we'll see.
>>109434225Just the normal turbo version is proving itself to be fairly flexible and surprisingly coherentprompt "Anime style, abstract."
>>109434289Here is a style explorer btw for ideashttps://lumenastrum.github.io/clio-style-preview/gallery/