Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109503671https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg
>>109505489>no will smith gifnot a real bake, I'm leaving!
>>109505497see you soon
I was genning with Wan 2.2 just 2 weeks ago, didn't even know about LTX (didn't visit /g/)Had a fun and frustrating week with LTX discovering its strength and audio for the first timeAnd now just a week after, I'm playing with H3 and finally body horrors of LTX are gone, no more missing fingers or fucked up tits growing from the back or face on the back of the head.But checking back on my 2 week old gens, honestly, impressive as H3 is, I2V+text doesn't quite get you what Wan did where I could prompt something simple along>real life videoand get picrelated. (18 seconds, Wan SVI workflow, the speed of of gens is comparable)Minimax really wants to change body shape, create face that is ugly, ot keep it videogamey anyway. Tits physics is also not there. Wan: https://files.catbox.moe/qb567e.webmThough still, this is a huge for us and big leap with for me in 2 weeks discovering I can do audio locally
>>109505495>has a literal crime recordLmao.
>>109505497>will come back begging for bbc pregnant fart fetish futa porn
new kroma version is outhttps://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.2-turbo.safetensors
https://files.catbox.moe/30vzn4.mp4
>>109505509my theory is that they started the h3 training on well-captioned 3d rendered stuff so they could get the model to learn movement quicker, and then they finished with mostly real stuff. that is why the less common real stuff gets automatically turned into video game graphics
>>109505523>Doesn't track the eye movementsGarbage, improve prompt
>>109505525nope. They said they did hardly any post training, just a ton of pretraining on varied stuff. Its the opposite. The look your thinking of is gained by tons of post training / RL
>>109505521>26.3 GB>turbowhat the fuck?
>>109505509>>109505525> well-captioned 3d rendered stuff makes sense, I've seen some gens where the movement of realistic subjects look very stiff
larryvrh/MiniMax-H3-Turbo-LoraDoes this work on ref model?
>>109505539fp32, wait for fp16 I guess
>>109505521Who cares? Minimax image will soon be out and it'll destroy anything.https://xcancel.com/MiniMax_AI/status/2086253065657790895
>>109505554not sure if they even could beat krea.
>>109505552that cant be a lora. is it a full distilled model?
>>109505552I'd like it more as a lora
>>109505559it'll definitely destroy klein, edit models really need to be improved
>>109505509maybe there is, something good about dual model for high noise/low noise split technique after all
>>109505533oh. so they didn't really write about their training regiment? i'm guessing they wouldn't want to write down about how they got their copyrighted data lol
it's been 40 minutes, where is kroma-v0.2-turbo-int8-convrot?
>>109505592convert it yourself
>>109505592it would take all day to download it thobeit
>>109505497i'll stay here because fuck the collage
>>109505554this, i already deleted most everything and will delete more once it arrives. Almost everything gone, controlnets, facedetails sdxl illustrious wan ltx every into the trash.
It's been 40 minutes, where is my 2 second long gen?
https://www.reddit.com/r/StableDiffusion/comments/1vj5scs/deroping_minimax_h3_fast_motion_to_reduce/https://github.com/matlowai/ComfyUI-MAINodes>H3 smears bursty motion: backflips, fast sword arcs, whip-fast reversals. The cause is structural. One latent token spans four pixel frames, and at high motion speed those four frames need four distinct poses that a single token can't hold. Re-denoising the affected region doesn't help, because the missing poses were never generated in the first place.>This pipeline works around that at inference time. It re-generates the clip as a slowed-down version of itself, seeded from the original. Frames where motion is too fast get held (repeated) so the model has more temporal room, the result is generated video-to-video from that retimed init at partial denoise, and the original frame rate is recovered afterward by dropping the held frames. The oracle that decides where to slow down reads the clip's own latent. No extra model, no training.
>>109505708it makes inference slower though right?
>not being born rich to enjoy this shit on 2x 6000 prosKek, what a meme life.
>>109505708Why have you not yet purchased an advertisement?
>>109505644don't make you jump scare you anon, I can be scary sometimes but not all the time anon.
>not being born actually rich to enjoy this shit on a rack of DGX B300s
>>109505710it's ok anon, in a few years they'll manage to make something as good as H3 but on a 6b model
Under 60 seconds for 10 seconds square video with the turbo lora and only losing some sharpness and audio quality. All the accelerators suck. They give you turbo lora quality and shave off 10-20 seconds from base.
>>109505710you dont need to be rich. get a credit card, max that shit out and make small $200 payments monthly. treat it like a car payment. YOLO
i have learned almost every horror this model has. it knows a lot and a lot of audio also
>>109505554i only believe when they release it
>>109505718If I had access to something like that, I would also train a bunch of complex Loras and share for freeThe wrong people are rich
>>109505733lollmao even
>>109505733>The wrong people are richyou always have to sell your soul to be rich so yeah...
>>109505733I would also try to save local musicgen
>>109505710I don't think anyone but researchers who get GPU access for free are genning with that
>>109505738where do I sign?
>>109505728or just dont pay it until its passed between so many debt collectors they will let you pay pennies for it. just hit em with the ol "ahh jeez duuude.. i got these pills man i can do like...$5 a month?"
>>109505708Why do people has to include 70MB of shit their repos?
i daydream about winning the lottery, setting up a dedicated gpu farm and living the rest of my days just gen'ing.
>>109505785many such cases
>>109505554Isn't H3 a distilled model? It will just be pure slop just like Flux 1 dev. H3 is currently slopped for pure T2V too, just slightly harder to notice on a model this big.
>>109505816it is only CFG distilled, the flux dev models are CFG and step distilled
>>109505816>Isn't H3 a distilled model?no. if it was, we wouldn't need turbo loras
>>109505495fuck off and kys
>>109505554> soon > china
if using int8cr its probably best to disable fp16 accumulation if you have it as a cli argument since it will only lower the calculation precision of things other than the main model, that actually have fp16 operations, like the vae, which you dont want lowered.
>>109505840Yeah, hopefully we are not into another "chinese culture" cycle again like we had with the z-image releases (where the Edit model ended up never being released btw)
>>109505834it is distilled, it's guidance distilled, the turbo lora adds a steps distillation process on top of it
why are models so lenient on nipples, but so consistently fucks up how they look with that like weird double ring looki understand the gooner shit being filtered out or censored or whatever but why is the one thats more acceptable so consistently fucked?bottom text
>>109505862>the Edit model ended up never being released btwthat's because migu is still alive... >>>/wsg/6208995
>>109505889this
>>109505816nobody asked them why they decided to not release the base model on that reddit AMA?
>>109505509>wan>https://files.catbox.moe/qb567e.webmGuys how long will it take to get us soft tits lora like this for H3?I'm getting tired of air balloons in h3 that look and behave exactly the same across gens
>>109505521my outputs are all weird
>>109505509>LTX discovering its strength?
>>109505910>my outputs are all weirdwhen will you guys fucking learn that kekestone is a fraud?? holy shit he's been pumping shit models for more than a year at this point and you still believe he can pull it off? kek
>>109505920his butt buddy S1LV3RC01N seems pretty knowledgeable. he should just take over
>>109505869we've seen <safety> like you wouldn't believe and countless faces have been palmedin the end the "gooner shit" is what actually works for everyone from <questionable> up
If you want your fetish trained into sulphur now is the time btw.
>>109505916After 2 years of Wan prompting, with LTX I could now request something much more specific to happen with actions ordered in prompt (often including abominable body and limbs movement, but still), while Wan prompting was vague and it would simply ignore the specifics. Also wan gens looked way too similar and seed was barely relevant.It's all still fresh in my head and I can compare as I used all 3 models in the past 2 weeks
>>109505940Consider purchasing an advertisment.
>>109505940lemme sneak some war kinos into there
>>109505940brb uploading some loli
>>109505943Oh, and 20s music + image to video was amazing first time experiencing in LTX, with fast movements and looked great physics loras, while Wan was really reluctant on doing any fast movements at all even with lightspeed loras
Did anyone try training any video model on goon compilations? Any LoRAs of that? I doubt either Wan or LTX could learn those properly so I'm thinking of training H3 for it.
>>109505907They don't always behave the same. If you're doing i2v I think it depends on the medium of the image, for example 3D renders look stiff because it's emulating relatively stiff 3D animation.
Gening is nothing more than degenerate gambling.>Bro I swear bro next sees is the one, just one more step bro, just one more lora bro, just one more seed bro
>>109505907if h3 think the subject is real life, the breasts jiggles fine. maybe try to prompt "subject's breast is soft" or some shit.
>>109505986wan learns really rather well, but I suspect so does h3 (haven't tried).obviously it depends on what you mean by "goon compliation". perhaps you want it to at least sort-of have some more limited focus unless you have too much time/compute to caption and train.
Why is ComfyUI's default H3 workflow using nearest-exact instead of lancoz?
>>109506021IDK, I switched it to lanczos without issues.
>>109506021processing timeit's too intensive to use anything other than bicubic scaling>>109506006>he hasn't founded the preview settings in the settings
>>109506021>comfy raping your input image more than any optimization to save 2ms of cpu time
https://www.youtube.com/watch?v=S9O3FPumX4Q
Damn I fucked up ref2 video. Tried two reference images for characters and a 15 second animated video and it just showed the reference image 1 as the first frame, played like 4 seconds of the original video, and then morphed the characters into weird proportions over top which kinda matched but the background was all fucked up. I was using the default workflow + spectrum + sage. 9 minute gen time...
>>109506006yeah and?
>>109506038noob
Hitler, Flux employee, finds out Minimax is better at generating Mikuhttps://files.catbox.moe/0hht5s.mp4
>>109506037Open-Weight Strategy: MiniMax chose to release H3 as an open-weight model to foster innovation, allow local deployment, and give businesses the flexibility to adapt the model to their specific security and data needs (3:28 - 4:28).Multimodal Generation: H3 is a general-purpose model capable of text-to-video, image-to-video, first- and last-frame generation, and in-place video editing (2:51 - 3:08).Native Audio Sync: One of the standout features is its ability to generate synchronized stereo audio (including dialogue and sound effects) simultaneously with the video (5:51 - 6:02, 23:42 - 24:25).Performance: The model is considered a significant step forward in the open-weights space, performing on par with or exceeding some state-of-the-art closed-source models in specific arenas like video editing (7:30 - 7:56)Optimization for Consumers: H3 has approximately 60 billion parameters, which would typically require 120 GB of memory (BF16). ComfyUI and the community enabled it to run on consumer hardware through techniques like quantization and fine-grained offloading, which keep only the necessary computations on the GPU while offloading the rest to system RAM (25:29 - 27:18).Resources for Developers: MiniMax provides a Context-IR API to optimize prompts for users working locally, helping the model better interpret complex cross-modality references (9:01 - 9:21).Future Developments: A 2K-resolution regeneration API is currently available through the MiniMax hosted platform, with ongoing development for broader integration (19:55 - 20:20)Rapid Iteration: The team highlighted the community's impressive speed, noting that quantizations and hardware-specific support (like MLX) were delivered within 48 hours of the model release (10:17 - 10:46).Prompting Advice: The ComfyUI team suggests using the official prompting guides and, for dialogue, explicitly including the spoken text in quotes within the prompt to improve lip-sync consistency (28:12 - 28:50)
>>109506038> <Picture 1> and <Picture 2> reference and represent <Subject 1>. <Picture 1> is directly integrated as the first frame>...video starts at 00:00 with <Picture 1> ...
>>109506048add english subtitles and this would be peak kino
>prompt camera pan to show her ass>she start shaking her ass unprompted
>>109506065It knew you are black(also brazilian).
>>109506065What did you expect, you used a picture of a negress
there we go, got the proper first shot, used a 0 to 3s: timestamp.https://files.catbox.moe/z8kg34.mp4
Saar, realism, photographic saar.
with all the security concerns using these indian/chink nodes with volatile code, docker setup would be perfect if you don’t want to dual boot. it was also genning somehow 10% faster than on bare metal windows when I tried it.Unfortunately vfs volume mounts are painfully slow and model load times atr terrible, getting like 50MB/s disk read speed off my SSD, while normally they load at 2.3GB/s in comfy portable.Is there a way to fix this without storing models inside the docker container /g/?
>>109506128
>>109506128you have to store the models inside WSL and enable it in docker desktop.
fucking bodied that freakhttps://files.catbox.moe/6nw0nn.mp4>>>/wsg/6210814
>>109506146>don't want to dual boot>>109506155the reason I specified>without storing models inside the docker containeris my hoarded models are scattered across different SSDs and HDD and wired in comfy's extra_model_paths.yamlthere is no way to fix slow docker mounts? Why are they slow in the first place, seems like a bug.
>>109506038it's not on you anon, the model is scuffed. it doesn't get it unless it's a simple put Picture 1 in Picture 2 situation.
>>109506181>>don't want to dual bootdid i say duel boot?
why were people shilling res multistep? euler looks way better
oh yeah i get infinite hangs at step 0 if i try a big gen on my 5070ti+32gb like 10s@1.5MP but somehow the same runs fine when it's an instance claude span up from WSL2, which is extra weird since it just spins up powershell from there. i'm definitely finding comfyui is a bit painful to prompt H3 with since there's so much boilerplate text you can fuck up which varies between t2v, i2v (f2v, fl2v) and that's just on the normal model, if one of the a1111-likes has a good H3 prompting experience i'm quite keen for that. there's enough variety to it that it's not even particularly trivial to vibecode a good prompt builder ui, though a simple one can streamline the process a bit. and having a local text model write/check your prompts is also a pain since it's yet more loading and unloading per iteration of a prompt
>>109506181storing it in wsl is not inside the docker container. >there is no way to fix slow docker mounts?no.
https://files.catbox.moe/4jd3nl.mp4
>>109506197proof? did anyone do a comparison?
>>109506197Multistep for non-turbo, euler for turbo. That simple.
>>109506220>Multistep for non-turbo, euler for turbo.multistep works fine on turbo though?
FL2VA vs Ref2VA, which is better?
>>109506216>did anyone do a comparison?yes. i did
>>109506226depends on your needs and prompting skills
>>109506228forgot the>proof?part
>>109506235it is not possible for me to prove it because i can lie. try it yourself. it's not hard. you CAN use H3 on your computer, right?
>>109506246showing a side by side comparison would be good enough, im already genning many queued videos
>>109506226ref lets you do pretty much anything you want, though I sometimes struggle to prevent it from carrying over certain qualities of the reference
>>109506197Shut up Leonardhttps://files.catbox.moe/4ul931.mp4>>>/wsg/6210819
>>109506252ok, test it out when you finish
>>109506233>>109506253I wanna make placeholder idle animations for torso+face portraits. Using ref image
>>109506263ref >>109506056
>>109506252>25 stepsinteresting, have you found reliable improvements from this? i've often tended to find if i blind A/B test, running models at the top end of their recommended step range or a 10-20% above it is a good way to get more reliable gens, but i haven't tested many things with H3 yet since there are so many variables and it takes multiple minutes per gen
>>109506263I grant you the authorization. Next.
>>109506271>running models at the top end of their recommended step range or a 10-20% above it is a good way to get more reliable gensyup, i didnt do direct comparisons but essentially 5 extra steps arent gonna take too long and will probably iron out some more things, on 3090 128gb ram for now i settled for i2v 0.7MP 8s 25 steps which take 10min per gen, sage 2.2 and default spectrum.I can probably decrease some of these things but i want to gen at the higher end of quality so that i get used to good gens before dropping things down so i can then know how much will the later gens deviate from the known good ones.
do you guys use any uncensored loras for minimax or you reckon the base model is uncensored enough?
Github is down or what? Can't access it today
>minimax_h3_ref2va_pruned_int8_convrotis this the way to go?
>>109506373yes
do I need both nodes for the sageattention 2.2.0 workflow? Minimax H3 Mem Eff Sage attention AND patch sage attention KJ node?or just one of these?
>>109505910the "turbo" is the base model it seems accidentally. Use turbo lora on it and it looks good. Way better than before with details
>>109506396model -> patch kj -> mem effalthough in my case even without those it seems like the speed is the same if you already have use sage as cli arg
troll bake
>>109506171link to mod???
>>109506373pruned is worse quality btw. watered down
>>109506429i believe sageattention is a global setting so toggling it on with a launch flag does the same thing as the node that toggles it on, and it stays on for ALL your workflows in that running comfysession if the node turned it on since the node really does just flip the switch. other sage node stuff is separate.
>>109506449literally the same outputs
>>109506449
nah that unc is fr but you dont notice it until you go back to your raw as fuck gen and compare it to cope and prune
>talking in third person
>talkingunc turn off your narrator this is a text board
>>109506453no, it's not, otherwise there'd be no reason for a "prune", anon. original int8-convrot is 34gbpruned is 21gbthat's a whole 14gb chunk of sexy data missing, anon.you think that's "literally the same"? why do you think half the thread has garbage gens.
>>109506409how bizarre
>>109506480nice animation, still not using troonix though
>>109506487the pruned part literally did nothing, that's why they pruned it
>>109506487>that's a whole 14gb chunk of sexy data missingthat did fuck all and was substituted with a lookup table
You can actually remove the diffusion model altogether and substitute it with your imagination, only retaining the llm.
Why is Julien such a desperate, worthless bottom feeder?
>>109506487prove it, use both and show us the difference on a same settings gen
so, how do you dudes pass the time while waiting for gens to finish baking in the GPU oven?
>>109506528
>>109506541I accept your concession.
why haven't the comfy team used this technique in the past to dramatically reduce model weights?
>>109506547bodied that freak
naw bro busted out the capital letter and period cuh think he tuff naaaahhhh :skull:
>>109506554shut the fuck up
>>109506550it only works when the model maker did a big oopsie
>>109506550because its not a usual prunning but instead removal of a particular part of the model that the creators put in that was known to be shit architecturally, its not random 13b worth of random model weight data that got pruned
cuh crashing out :sob:
>>109506538get started on the next one
A few days ago, I already complained that the forum has turned into a showcase for TikTok slop.Some people have brushed it off, saying it’s always like this for a few days - I strongly disagree. H3 is so easy to use that any kid can churn out this slop in bulk, and as long as it gets likes and isn’t removed, more slop will keep coming.Until the mods do their job, this spam will go on forever. Valuable or interesting posts are buried under all this trash.People are now even calling their trash “another AI slop spam” right in the title, and yet the garbage stays online.It’s annoying, and the TikTok crowd will, of course, disagree with me.
>>109506590oh hi fellow ledditorhttps://www.reddit.com/r/StableDiffusion/comments/1vjmp5z/ai_slop_spam_continues_mods_do_your_job/
cuh a reddit tranny ahh nga i knew it fr
>>109506602Shalom (blessed day) saars
zoomspeak is basicallygoo goo gaa gaa>look im pretending to be retarded
>>109506644>goo goo gaa gaathey literally made a meme with a character they called goo goo gaa gaa, I'm not jokinghttps://www.youtube.com/watch?v=ziIk72mUz5c
>>109506644you're a redditor
which model is the current sota for 2.5d anime gen?
>>109506667Wait for the H3-derived image model
>>109506661>they
>>109506490
>>>/wsg/6210839
>>109506678yes, "they", the zoomers
>>109506677what about now?>inb4 no answer
>>109506690the youngest gen z are 14, they're not the ones watching an anime penguin girl go gugugaga
>>109506667Anima
>>109506709they're retarded enough to watch that shit, regardless of their age, a truly wasted generation
>>109506661you know it's a bunch of 40 y/o single mums that watch this crap the most
>>109506709>the youngest gen z are 14And the oldest are 30, unc
>>109506709true that, gugugaga is peak millennial boomer uncslop. on god, fr fr, no cap.
>>109506720don't you have clouds to go yell at?
unc is crashing out bruh
>>109506745>peak millennial boomer uncslopI've grown with final fantasy X, gta san andreas, tekken 3, you've grown up with goo goo gaga and body type A/B in games, we're not the same
>>109506755at least it isn't bother CJ crash out over ani
>>109506758>final fantasy Xnot 6 and 7. >gta san andreasnot the gta demo from pc gamer with a trainer to remove the time limit.>tekken 3not virtual fighter in the local arcade.always a bigger unc in the retirement home.
>>109506767what is blud even typing
https://files.catbox.moe/do9bqu.mp4 Is there any way to control the speed? I sometimes get gens look too much slowmo
>>109506771ff7 sucks, it got popular only because it was the first one that got released in europe
spectrum node rapes my system, everything just keeps freezing
>>109506776video sigma shift down a bit, like from 12 to 10ish
>>109506776finally some good shieet
>>109506785whats video sigma?
>>109506791sigma balls hah gottem ha ha
>>109506788>>109506776>pov fingers change into done nails
>>109506791
>>109506803true heterosexual men only watch lesbian porn, why would you want to see another naked man while jerking off? are you a faggot?
>>109506808That's a big node...Come on... say the line...
>>109506776nice
>>109506538Hop on the treadmill, i actually get fitter and become a more powerful gooner the slower the gen time.Ironic isn't it.
>>109506813reddit ahh nga
>>109506825https://www.youtube.com/shorts/lrrMJqjgsIM
>>109506820you sound like you were molested
>>109506833reddit nga crashing out nah son im crine :sob:
i really need a new prompt bot for h3. grok is uncensored and has given me some great prompts, but it constantly baits for payment. what are you guys using?
a new email address lmao
>>109506843Gemma4; on another syatem
>>109506856this
>>109506814Sir, this is /ldg/ we dont have lines here, just spaghetti.
>>109506843LM studio with gemma-4-26b-a4b-it-ultra-uncensored-heretic, it needs a bit of refining and extra instruction apart from throwing the two prompt structure guides at it but it's solid AND it can be put into ram if you want to run it concurrently with H3 and have enough ram to do so.
>>109506776u need to shorten the video length or add more actions in a long video. sometimes the gen become slow-mo because there are not much action going on.
>>109506864>>109506856what vram are we looking at? i've used my old server pc with a 1060 for prompting previously, but that was just normal descriptive text. i doubt it would be able to run a model smart enough to format this stuff properly.
>>109506820>>109506839>ahh>ngapot calling the kettle
do you put the optimization stack into both the guider and scheduler, or just the guider? i've seen workflows with both.
>>109506915both
>>109506884First make a tool to help with formatting the prompt with Claude or Chatgpt. You can even make it a standalone html page if that uses the browser local storage. Then you just need a local llm to handle the filling in the various sections of a prompt and not generating one wholesale. I'm not sure how well Gemma 4 will do generating a complete prompt, especially at low quants necessary for a potato pc.
>>109506884for minimax?12gb + 24gb for vae's and main model6gb + 32gb for clip/text encoderI have to restart comfy to free up memory for gemma 4 12b qat, but it seems to work alright at a temp of 0.7/0.8. Sit's in 12gb vram with 85k token memory and the right settings in lmstudio
https://files.catbox.moe/b6e7t0.mp4
4060ti 16gb 64gb sage+spectrumr2v0.8MP 14 seconds gen time: 40 minutesvae decode time: 10 minutestotal time: 50 mint2v0.8MP 9 seconds gen time: 14 minutesvae decode time: 3 minutestotal: 17 minthese vae decode times are making me angry.
what even is non_diegetic_music
>>109507069it's music that's not diegetic.
>>109507069search on reddit cuh
Nvidia drivers suck ass on Windows and Linux today
>>109507038>vae decode time: 10 minuteshmm, that seems crazyare u using --vram-headroom 1? try with thati also assume ur using sage 2.2, cu130 pythorch 2.10+, int8cr
>>109507069>diegetic>(narratology) Of or relating to diegesis; existing within a fictional universe (rather than as background), and able to be perceived by the characters.
>>109507069diegetic = part of the world, e.g. when in a movie or game you hear jazz playing but then there's an actual jazz band there on stage and the characters in the world can hear the music, it's diegetic. whereas non-diegetic means the music is just playing over the video. you remember all that buzz a few years ago when dead space came out with its "diegetic ui elements"? you can't have forgotten already, that was only 2008
>>109506843Get a cheap 16gb card like 5060/70 ti and let Gemma 12b run on it for captioning purposes.Makes life a lot easier having a dual GPU system for this AI stuff.
>>109507038There was a vae released the other day that cut times, idk if it was comfy or KJ that released it, it didn't work for me, maybe it's been updated.
>>109507085>are u using --vram-headroom 1no only --disable-pinned-memory>i also assume ur using sage 2.2, cu130 pythorch 2.10+, int8cryes
>>109507083Speaking of Nvidia driver, do i use the Studio or Game Ready drivers for local diffusion?
What's the best uncensored/nsfw realistic image generator like now?
>>109507099>a few years ago>only 2008Brootal uncpill
>>109507121try just removing --disable-pinned-memorythen try keeping it removed and adding vram headroom 1then try keeping --disable-pinned-memory and adding vram headroom 1
>>109507104I also tried it, didn't work for me either.
>>109507149There's a pinned memory fix in yesterdays stable comfy release for Linux.
>>109507149why vram headroom and not reserve vram?
>>109507132Nta but 2008 feels very recent to me as well and I was only a teenager back then.
>>109507167Did you love being a teenager in 2006?https://youtu.be/cydBbZaPBU0?si=wpdOon5WXZI2Pj15I was 9 in 2006 kek
>>109507163from what i heard its better, from what i know it should be similar/the same, but i dont know for certainalso it seems like the problem will be fixed soon https://github.com/Comfy-Org/ComfyUI/pull/15316
>>109507182>I was 9 in 2006 kekunc
>>109507185and https://github.com/Comfy-Org/ComfyUI/pull/15446
>just add the flags bro>no I don't know what they do lol
>>109506843>what are you guys using?https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUFlooks like it's also becoming one of the most popular uncensored models in general now. but it does work pretty well for both captioning images and inventing prompts (or combinations)
>>109506538
>>109506104fegelein!!! fegelein!!!
Does the reference model also want timestamped instructions?
>>10950720036
i've been trying to make a custom node for reference sets to easily switch between characters and concepts. however, the output seems to be the real bottle neck.from my experience so far:>image batch: has to be unpacked - stretches every image to the same aspect ratio before outputting, ruining them>image list: has to be unpacked - crops every images to the same aspect ratio, ruining them>having a ton of outputs in the node: will require a bunch of manual connecting and disconnecting to the main h3 r2v node depending on how many images are usedi could just connect all 9 reference images permanently and max the node out, but i'm worried that if i only have one reference image, it's going to get fed 8 additional black squares or whatever it decides to do with the additional lazy inputs that arent being used instead of only using the one selected image. does anyone have any suggestions?
>>109507223what's that supposed to mean?
>>109506883Maybe? I can roll the same prompt 4 times and half of them are slowmo, and rest are fine.https://files.catbox.moe/dfzqw0.mp4
>>109507236ok Jensen, you won, time to empty my wallet... and my balls!
>>109507038https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main>minimax_h3_video_vae_int8_convrot.safetensors
>>109507228My problem with reference is they always create entire new scene when i mention it in the prompt.
>>109507217https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md ctrl+f "000">subject_definitions: here you basically declare each input as <Audio 1> or <Picture 1> or explain that <Subject 1> is the fat guy in <Picture 1>>summary: pick from the list of source usages in the prompt guide and explain in brief how they relate to each other>retention_analysis: similar but defining how strictly they're to be preserved e.g. along the lines of "copy and paste this in" vs "use the outfit" or "play the exact audio" vs "copy the voice timbre">detailed_description: basically the same shots + timings as you're used to in i2v>overall_soundscape: same>non_diegetic_music: sameit's a huge pain in the ass, much moreso than i2v, so you're best off feeding the instruction md to a smart model (point cc or codex at it, for example) to get your base prompt set up each time you vary the input connections to the node or want to change the way you're using them, then tinker by hand from there. If it's NSFW to the degree that you think it'd be annoying to wrangle claude into helping, just build a sfw thing of the same shape and have the template for that set up. after a while you'll have a library of images with the metadata saved ready to help you where you personally remember what format each was doing. the amount of possible combinations of usages is huge so it's not solvable via others' templates, maybe someone will figure out a nicer ux that builds text chunks for you but until then this is at least not too bad
>>109507258Tried it. Only got like 5 secs speed up. Not worth it
>>109507228if the image list already only provides activated images - what if you fork the next node to take image list?
>>109507261This is why Minimax prompting is worse than LTX. they cant be simple. I think this is what filtered most people.
>>109507275All of the simple prompts work great for me on H3, maybe stop quanting the TE to 4 bits and using turbocope loras?
>>109507272yeah i was just looking at the ref2va node and i think that might be an easy option
>>109507275>the thing that's objectively smarter and more detailed is worse because it's smarter and more detailedvramlets aren't the ones that ruin it for everyone, it's literal retarded assfaggots that can't read instructions (or just steal from other people's workflows like a normal person)>>109507281the fucking nvfp4 works perfectly fine, it's why everyone uses it.
>>109507275Could be worse. At least we aren't animating bounding boxes.
Considering my output folder is 403 GB of PNGs, I'm finally going to accept the devil into my life and run a recursive python script that converts everything to WEBP and deletes the originals. Quality 95.Any last words?
>>109507286Im fucking horny. I dont want to prompt a literal words from lord of the rings novel. I just want to prompt "1girl, anal, threesome, blowjob" and be done with it
>rapes the text encoder with the toyest of toy quants to save 2% of gen time>complains the text encoder of H3 is worse than fucking LTXbrown moment
can minimax not use a singular reference sheet instead of individual images?
don't forget to preserve the workflowsyou can also use a script that uploads everything to a catbox account or similar
>>109507300no one said that
>>109507291you can save ~40gb without losing any data or workflows if you runoxipng.exe -o max -Z --zi 15 --alpha --preserve -r .if you want to compress it much more while losing information then convert it to jxl, although you'll have to vibecode the program that does that while also copying over the embedded workflow
https://civitai.red/images/139098095
>>109507301what do you mean? you can do character sheets
>>109507312so instead of using multiple ref images, I can just take a single image that contains all the references and it'll use the entire sheet without issues?
>>109506871
>>109507331well it gets scaled according to the resolution you are generating at so you can't cram the entire story line into one giant image probably, but a single image that has multiple perspectives of the character works fine
Minimax dont understand what 3d animation is.If youre reference image is anime it will force you to low anime frame like crazy.
>>109507345>well it gets scaled according to the resolution you are generating atthat's dumb as fuck what the hell why would references need to be the same res
>>109507352Based
does the text 2 video model know more celebrities than the reference 2 video model?
>>109507352Based complaint I meant
>>109507353it scales automatically. you don't have to do it
>>109507291>>109226527
>>109507376I know, that's not what I'm sperging out for
>>109507383>90%
>>109507376yes but the problem is that this means a low res character sheet, so small details will be missed. seems better to use individual ref images in that case since each individual image is higher res
Playing around with the T2V H3 model now for a bit and really struggling to get truly dark scenes. Even if I prompt stuff like "dark room, darkness, dimly lit, video shows a very dark room, chiaroscuro" etc. it always wants to add studio lights on the people in the shot. Any anon with tips here?
>>109507338heh>[INFO] Prompt executed in 375.94 seconds
this is a really good first frame for a video ideabut im out of ideas, brain focused on coom. fuck.
>>109507387i don't know. it probably feeds that as one of the frames to act as context for the rest of the video, but cuts it back out at the end
>>109507207kek
>>109507408Tarantino slobbering on dead nigger toes
>>109507408you could have a "ghost hunter" style clip where famous black entertainers like sammy davies junior, fats domino, michael jacksons bodies are kept for him, why? we dont know, but they spook him in the styles that made them famous?>>109507474maybe he bought their feet at an after death auction, like a bone from jesus finger or similar.
i heard europoor are burning in the summer meanwhile my gpu doesnt go past 50c during 100% usage. interesting
>run at half power>wonder why its only 50c
>>109507534we know the 1050 Ti is not power hungry, anon
>>109507549kek, bodied that freak
>>109507549europoor coping
>>109507338kek
>>109507573kek, bodied that freak
>>109507549>>109507543AI only need VRAM. Technically if 1050TI has 32gb of VRAM, it will performs as fast as 5090 (This is why RAM price skyrocketed)
>>109507573>>109507591>flexing a 1685 mhz gpudefinitely an indian
>>109507573>doesn't actually show which GPU he haslmao>>109507595troll
>>109507606you arent flexing anything cause your gpu died in the heat, europoor
>>109507613Bro 2080 ram gets soldered to 26gb by chinks to become an AI card
>>109507617irrelevant now quit trolling
>>109507573Let's see the actual numbers under H3, lil tyrone
>>109507631i gave them
>>109507573this u?
I just installed comfyui, how do I get started? What model should I use?
>continuous shot, fixed camera>H3 still makes cutsbruh, what's the point of a 32b TE if it doesn't listen to this basic shit??
>browns wake up>start wanking over hardware with ESLmany such cases, please fuck off or post gens and fuck off.
>>1095076243090 24gb peforms MUCH faster than 5070ti 16gb. Cmon now. ONLY VRAM matters for AI
>>109507653>brown is so low iq he thinks him saying something proves itThanks for exposing yourself.
why has no one done this yet? @china
>>109507677>>109507677>>109507677>>109507677
>>109507662>t. sub 24+64 v/ramlet
>>109507670because China isn't composed of brainlets who listen to AI hallucinations
>>109507670@china @xi @police @elonmusk
>>109507679troll bake, ignore
>>109507695Make proper
>>109507664the 5070ti is MUCH faster than a NVIDIA P40
>>109507677>>109507677>>109507677
>>109507679>>109507728This is Debo thread. Avoid this and go to >>109507737>>109507737>>109507737>>109507737
>>109507745>1 image>5 minutes latetoo slow