Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109533474https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
stinky
>there's no official comfy workflow for krea2 raw>find some on civit>the results are worse than turbo and don't match the intended style, which turbo matchesemmm... Does anyone have a working one for the raw model? I tried playing around with cfg, negative, samplers, and schedulers, but I can't make it work
>>109536207i think its just my mega autistic fetish for hime cuts/bangs to be honest
>>109536214Given the little time between the latest gen and the collage, you're the one who made the simpson gen
>>109536243try dual sampler, do first few steps in raw and pass latent to another sample with turbo and finish that way there's a raw to turbo lora so you don't have to load both models, only connect it to 2nd sampler
>>109536214me want the miku one
>>109536262nope, I am simpsons anon testing ref2v, havent baked in ages
>>109536207do you know where i can find high quality namine cutscenes and screenshots of her in game. There really isn't much of her in the kh2 and dream drop distance fmv. her look in kh3 is way too different than the previous games. i just need to have about 30-35 images of her to be in the safe spot.
>>109536275sure, but isn't this sidestepping the problem? I expected raw model to simply be better than turbo in every way, especially with loras that are trained on it
kek, just stitched the 2 5s clips with the vhs nodes/merge images/concat audio. 2 voice clones can be tricky in one scene but you can just do 2 cuts.https://files.catbox.moe/3o15uc.mp4
amdbros how are we coping with the h3 speeds?
>>109536324Won't work if we need a scene where two characters interact in a more flowing way than exchanging one line of dialogue each but it's a decent hack. That said, did you already try asking the model to just do Skinner and Chalmers voices?
>>109536306Not sure. You might have to look for non-official content if KH3 isn't usable.
>>109536329ebay sell amd gpuebay buy nvidia gpusimple as
>>109536313from my testing it isn't better in every way, it's only better in output diversity
>>109536341the problem is the model uses the first input for voice cloning, it can work with 2, but you need to use certain tags to avoid the model getting confused, even if you say character b uses voice 2.there is a way, but sometimes it will flip the voices. in any case, this is pure gold. it could clone gilbert gottfried's voice ffs.
>>109536347>my gpu is going for barely 1k on ebay>not even enough to get a used 3090 anymore which would be a downgrade for gayming>not nearly enough to get a 5090it's so ogre
Blessed thread of frenship
>>109536306maybe in KH re chain of memories? the card game one?
>>109536214Wow, comfy is a bad friend. this is bad pr
>>109536306haven't watched it but maybe there's enough in this?https://www.youtube.com/watch?v=khuXhyIbvDk
Blessed thread of schizos
sneedvr2?.. nah..more like OOMVR2 haha lol fuck hurry up with my upscaler chang.
>>109536324and the great thing about china training on all content is it gets all the details right, I just used a skinner reference image and it still gets the shading right and so on, in motion.western companies are too cucked to do this, or censor stuff too much like anatomy.
>>109536382If you haven't realized it yet, the OP is still being made by a sharty troll
>>109536382AniStudio stands on top as the best app for accessing and using the latest AI models
>>109536366get a 5070ti. I can gen whatever on a 4080, same vram. if you have enough system ram you can handle whatever.
For that anon who asked about doing ref2v + first framehttps://files.catbox.moe/1xpczk.mp4
>>109536413>just did a rapid fire nute at that last frame
>>109536350Yeah, that's why I want to use raw, but I feel like my problem is beyond that, as seen in the example of ratatatat74 lora I grabbed from civit. You can see how bad raw looks
>>109536411Bane?also nice 360
>>109536402Catjack putting his shitty ani obsessed gens in the OP means it's the sharty?
>>109536411she's a pretty cute girl
>>109536434pretty sure he is just a sharty retard they don't even like
>>109536411that was me anon. but your video has no metadata :(
>>109536434>>109536443This has merit but I think CJ is too un-ironic that just comes from being mentally unstable
>>109536454shhh be silent lil lolcow
>>109536446sorry rajeesh
Brown clownery about uncensored text encoders got so bad the main guy himself had to step in.
>>109536354nice
>>109536446Apologies, didn't want you to see the cringe DVa porn I was using before. I literally just asked my prompt creator system prompt to make a video where a girl does a vlog about her action figure and specified I was uploading a reference.https://pastebin.com/dDy1v5PM
>>109536373never actually played the game, will view the cutscenes but need video that ultra enhanced and not to blurry.>>109536387fuck... would have been a perfect video if was at 1080-2160p and didn't have that annoying disney and square enix watermark at the bottom. Might have the find a way to upscale and clean up the screenshot if this is it. I guess i got lucky with kairi and her fmvs. she had multiple body views and poses which is important for a well functional lora. namine is too much of a neet lol.
>>109536496most people here have already been saying this for years
>>109536496>censored models need to understand what problem content looks like in order to identify ityou'd think this would be obvious
>>109536496it's almost like you should never use any workflows made by a jeet
>>109536472ok lolcowjack
>>109536589most people don't think for themselves they just use whatever workflow has the most upvotes on reddit
>>109536472>I'm the big daddy lolcow around hereis he finally self aware?
>>109536621he's the only real lolcow here for sure
>>109536496One of the most dumbest things I get seen spread, up there with the retarded use of jumbled up numbers and letters for activation triggers (which still persists...)
microsoft PR teamhttps://files.catbox.moe/nyzrzg.mp4
>>109536639>One of the most dumbest things I get seen spreadFor me, it's:>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>20 mins total to gen 1MP 10sec H3 last time i tried>now it's 20 mins total to gen 0.8mp 10sec??? why is cumfartui like this
>>109536653Use voice reference.
>>109536668>pulling
>>109536677I did, may need a slightly longer clip of skinner
>>109536668I had to revert to comfy v0.31.1 because the newer updates completely slowed down my gen times
I shoulda bought an RTX 6000 PRO when I saw them dip to sub $8000 last year...
>>109536519oh shit it works, as long as you don't describe the first shot to be something different than what's in the picture, thanks anon i didn't realize
>>109536561You can always just look for higher quality videos of the specific scenes after finding the ones you need.
>>109536739live and learnhindsight is 20/20keep on truckin'an eye for an eyewe will bury you
>>109536739Seriously. Considering that local video can only go up from here I'm already thinking I have to turn my discretionary income into a savings account for a dedicated inference rig next time we see a major launch
>>109536739elon musk said money will become irrelevant in a few years thanks to AI, so don't worry about it
>>109536742So all you have to do is say use <picture 1> as the first frame?
>>109536668i didnt even update comfy or change anything, and i'm having this happen>26min total time to gen 0.5mp 20s video>run again>50 minutes passed so far, not even 80% done???I dont get it, sometimes it's like the sampler step just doesn't give a shit and hangs
>>109536785that is a memory leak if you do subsequent runs
>>109536785I've seen the same issue where gen size+length+cope combos I know work in a given timeframe just fuck up and hang at like 65%. I've tried either restarting comfy or my whole PC when this happens but it definitely seems like an issue.
>>109536776I think the relevant parts (at least in my toaster's output) are:- Define the picture in subject_definitions as the first frame- Add 'keyframe completion' to the summary- Note how the picture is used in retention_analysis- Restate the picture's role in [Shot 1]It's a lot to remember which really makes me think the guide was always intended to be fed to toasters
>>109536410>5070tiI'd be tempted to do it if I had enough money to build a second system to use as a headless AI server. I do have some spare DDR4 RAM but it's slow (2666mhz) so I doubt it would be useful.
>>109536739why, 16gb is enough. get a 4080 or 5070ti.
https://files.catbox.moe/a83qoj.mp4
>>109536826hey that cameraman is a pervert!
any way to stop characters from speaking gibberish?
>>109536826too much, have her wearing something transparent at least
xbox after hiring a jeet CEO:https://files.catbox.moe/f9fqh2.mp4
>>109536848Give them actual dialogue or specify in the prompt that there is no dialogue?
>>109536826Is this Jennifer Connolly
>>109536848noise. any kind of noise, wind, television, radio, static, coughing, birds chirping, tyre roll, footsteps, fans, vaccuums, you gettin the picture?fill overall_soundscape and it'll prioritise that right under dialogue
overall_soundscape
>>109536923kino
>>109536826>https://files.catbox.moe/f9fqh2.mp4holy moly
>>109536815>It's a lot to remember which really makes me think the guide was always intended to be fed to toastersI mean the guide reads more like a prompt than instructions for a human. Or maybe they've just been vibecoding so much they no longer see any difference between man and machine when giving instructions.
microsoft PR part 2https://files.catbox.moe/36ih00.mp4
>>109536826Fashion designer getting really lazy.
>>10953666826 mins to 1 hour
>>10953666820 minutes?!?!? sheesh. You using loras or anything?
>>109536214Why https://huggingface.co/Comfy-Org/MiniMax-H3 and not https://huggingface.co/MiniMaxAI/MiniMax-M3 ?
>>109537021you shouldn't run unauthorized models on your GPU. run comfyui's with a hidden watermark including your UUID.
>>109536422I've had the same problem. wanted to use raw for a long while but with the turbo lora set to low strength anatomy and adherence to the prompt collapses
>>109536739ram is just as equally important than vram anon. I'm able to generation 1mp resolution with a 5090/128gb ram in 4.5 minutes with the FL2VA model on wan2gp. 1080 native res for h3 is impossible for me.
>>109536797i've never had this happen with 24gb vram and 128gb rammight be a low memory issue
Thank you for your attention.https://files.catbox.moe/vv1kir.mp4
gonna have to download pussy lora, fuck...
>>109537075>128gb ramwell yeah I figured, I only have 32
>>109537042I better trust the Chinese on this one, m'lad.
>>109537097boing boing
>>109537097you need a continuation where a clown shows up and ties those boobs together or to some other balloons
>>109537097this is the true boob physics test btw
>>109537101your dad.
BREAKING NEWS:Ostris has given up on creating a Turbo LoRA for H3. His GPU is too busy working on other thingshttps://xcancel.com/ostrisai/status/2087550974319886492
>>109537200Ostris is a hack imo.
>>109537097MY EYESGET THE BLEACH PLEASEEEEEE
>>109536608>>109536621>>109536636>>109536660you must surely have realized by now that the only consequence to your incessant pissbaby tantrums is our amusement, right?
>>109537228you are weakand so not prepared
>>109537257Is that Akali?
>>109537214what does ostris even do for the community? serious lora trainers avoid his stuff like the plague and total noobs can just ask an LLM to do everything for them
>>109537291yeah, it's a bit of an underbaked lora cos kree is so large to train on but was the idea
>>109537298well looks good to me anon quality work
>>109537292He made like 1 bad lora for krea and comfy sent him a blackwell so everyone thinks he's hot shit apparently?
>>109537097hot
>>109537352now make freeman get him with the gravity gun
anyone else running into an issue where voices come out super echo-y using an audio reference?Audio ref: https://files.catbox.moe/p8bvys.wavoutput: https://files.catbox.moe/6ujw1q.mp4Pic rel is the image ref
>>109537372>Pic rel
>>109537372can't account for the quality but she's not moving her mouth, sounds like ADRdid you say that she's the one talking?
>>109537388Yeah that's a separate issue I'm also having, probably related to the prompt. I tried to be very clear that she's talking but for whatever reason half the time she doesn't
for anyone who was using heretic:https://www.reddit.com/r/StableDiffusion/comments/1vmdxzk/psa_im_the_creator_of_heretic_and_i_advise_you_to/
Anyone knows whats causing sampling to freeze my pc? it’s not like sampling itself is hanging, 40s/it but my pc is unusable and can’t preview due to thisThis is something new, either comfy update or cope nodes
>>109536668figured out what my problem wasit's the int8 fast node, i have no idea what could actually cause the slowdown from before but it raped comfy itself when i pushed to 1.0MP so i removed it, and now i'm doing 1mp/10sec for only another 30ish second speed penalty. fine by me, 1MP feels like the bare minimum if you want coherent faces for more than 1 character reference.Apologies, comfydevs. You're off the hook for this one.
>>109537404I use it for the prompt enhancer, though. He says I'm based and heretic pilled.
>>109537401if your overall_soundscape is empty then put something in there, like her slappin her arsedon't ask me why but it generally helps
>>109537333thanks
>download pussy lora>all shots now have extreme fast dolly zoomsI hate lora bakers....
>>109537528but saar you want to see vagene, yes? this why I give you this
>>109537547Metriod Prime but real?
>>109537564worse part is the vagene is already passed as reference, it just still fucks up.
Does LTX 2.5 support coom?
>>109537626I don't think I've seen anyone post gens from LTX 2.5 yet. I don't want to bother downloading it I'm already happy with H3.
>>109536668I had performance regressions that persisted though comfy restarts and only reverted by a reboot. It could be Windows but i want to blame cumfart.
>>109537585are you at least genning at a minimum of 0.7mp? i know some of you niggers are judging gens lower than 0.5mp for some reason. I passed it a reference to a specific pussy for a character and it got it almost 1:1, i think if i pushed that to the full 1mp it'd get it fine.>>109537639it's like russian roulette but not russian roulette at all because the gun is pointed at either cumfart or windows.
Can someone explain why the resolution affects the model's intelligence?
>>109537626it's no different from LTX 2.3 they just panic released this model because H3 became the new top dog
>>109537528also this basically adds pussies to any crotch region. now my guys have pussies too.
>>109537528I think it's gonna take a few more weeks before we start getting some good loras
>>109537664square peg, round hole
>>109537664>resolution affects the model's intelligence?Anon pretend like genning at higher res is better but I always get the best prompt adherence at 0.3mp. Anything higher and coherence drops out the window.
>>109537693Opposite for me. At low res it's constantly fucking up details.
>>109537693You didn't try enough. Higher MP has better spatial reasoning and movement.
>>109537693press x to doubt.model native and recommended resolution is 0.98 mp
>>109537678that's pretty funny though
>>109537693increasing the megapixels exposes the undefined behavior you left into to the prompt
What are some good video upscalers in ComfyUI that work with 16gb of vram? Tried SeedVR2, but it needs too much vram.
>>109537693you get better prompt adherence with a better prompt, dumboit's cope to say the resolution has anything to do with that
>>109537732I prompt in Rust thobeit
trying to enlarge the thumbnail and watch the full size video from the asset bar gives me this error on firefox. anyone know what's causing it and if there's a fix?
Less resolution means more attention heads per pixel. so better gens.
>>109537737>it's cope to say the resolution has anything to do with thatyes you're absolutely coping if you're telling us 240p provides better prompt adherence than 480p or 720p. but the big thing to remember is the reference model is broken and looks shittier than first frame last frame, so actually pushing it higher than 0.7mp is probably a waste of compute. Have to wait for chang to rugpull us - i mean provide an updated model and the upscaler.
>>109537755download video and watch in mpv with your profile=high-quality, duh
>>109537773I've just shoved the fl2va model into all my gens, ref mode or not, I think as long as you don't do audio or video it works just fine.
is Kroma worth it yet?
>>109537737>>109537773Can't you just test this by using the same prompt and seed and only changing the resolution?
>>109537817>testing
I asked this in another thread but since these autists can't agree on one thread:Is Krea or flux 2 dev better for creating realistic images? I just got into this shit with minimax and now I'm messing with photos.Just upon some testing with flux2 dev, the images look less convincing than videos produced in with minimax, and it's harder to get them to adhere to the prompt compared to minimax.
>>109535025Thoughts on the latest pokegen?https://d.uguu.se/LuBnRDpE.webm
>>109537847best one so far.
>>109537847>boobs>hipsShe is supposed to be 10 years old, retard. Try again. 0/10
>>109537847>>109537855is your /a/ thread that empty, freak?
>>109537866i haven't checked yet.
https://files.catbox.moe/cisuuk.mp4
>>109537885The Cryptkeeper looks weird without the janky, puppet-like movements, but I love it
>>109537664It's about what the model expects, you'll almost never see a TV show or film in 3:4 portrait, so it's harder to get that film quality at that resolution.
>>109537885That's great
>>109537973cool
>>109537911Yeah, maybe thats something I can prompt for. This was my first attempt at using the reference workflow since I was spending all my time trying to optimize the text/image to video for my 3060. Pretty impressed by it. 2 images and an audio reference took about 20 minutes, which is maybe 5 or so minutes longer than the other model, but worth it, it seems. Still gotta do more tinkering I think.
>>109537973is that a white Natalie Portman?
>>109535025What did I say?
>>109537865Anon, you're retarded and don't know what you're talking about
Whats the benefit of using the full bf16 model for minimax?
>>109538097it is more betterer
>>109537097Suddenly, I'm scared of clowns with balloons again.
>>109538085>cherry-picking a frame>can't even deny that proportions are completely wrongOfffuckingmodel
>>109538127Relax autist, you were just wrong.
>>109538097training
>>109538121a clussy with those would hit hard
>>109537847movements still look stiff and artificialI don't think we're there yet for anime unlike 3D stuff
>>109538185
>>109537847nice
>genning a 10 second video is only a few seconds slower than genning a 5 second videowhy is this the case with H3?
>>109538224damn that banner really enhanced this image.
>>109537810Briefly tested it, got meh results from Turbo, it's slightly less slopped but I think the 0.1 LoRA was better coherence wise.
Why does Spectrum tend to fuck up dynamic camera movement?
>>109537372With no audio reference it's waaay better quality. I genned this one at .8 MP but even at .5 like the previous one it's night and dayhttps://files.catbox.moe/hs2ghb.mp4
Thoughts?
>>109538103
>>109538218Y-yeah, only a few seconds...
>>109538256because it's shitjust use turbo with 8 steps + sage attention
Are there any local models for generating high quality audio? H3's audio is kinda meh.
>>109538303acestep xl. it's not as good as H3, imo.
>>109538276I hope you're not spending hours on that shit
>>109538284Unironically I don't need more
>>109538321It's just a low-IQ autismo thinking he's "getting back" at the pokemon poster.
>>109538339risky handjob but worth it
>>109538284i lived that way for years.. except my desk was just a card table
>>109538339Nice! That hot air effect is very neat
>changing the sageattention model from the fp16 down to auto (which probably brought it down to fp8) made all gens significantly worsewow, it really does make a huge impact. vidrel. absolutely unwatchable trash quality i've never gotten until now. huge shout out to that singular (1) anon from probably 20 threads ago that pointed this out, after enough testing off and on i can definitely say it matters.https://h.uguu.se/UUpLgssH.mp4
Ok, why the fuck does Comfy keep doing this? Just randomly it makes the icons in the mask editor black on a dark-grey background.
>>109538309H3 is not a specialized music model though, so while its generations are good, the actual usability as a music model is nowhere near ACEStep. Minimax does have a superior music model in their API but they haven't mentioned much in terms of open sourcing it. Alibaba is also sitting on their Qwen Music model which is franjly unlikely to ever come out, so ACEStep remains SOTA for local music (and it's among the best with a LoRA(.
>>109537973Prototype 3 looking good
>>109538366Sorry, but acestep produces unlistenable garbage 90% of the time. H3 is by far better at producing good results despite literally not being made for it.
>>109538353You know everyone can tell when you post wanschizo?
>>109538384You know you and ani are the only ones who use this label? and you copied it from ani, a long-time lolcow for this thread, and there's a rentry link in the OP about him.It's actually remarkable how stupid you are, anon. You have completely transitioned into an ACK-tier schizo and you don't even realize it. You have been successfully mind-broken and humiliated beyond repair. This is a good thing.
can you leave /adt/ alone Mr. Catjack? thank you
>>109538383Well. I would expect a 33B multimodal model trained on a similar dataset as Udio/Suno (namely, the entirety of Youtube) to perform much better than a 6B model trained on a smaller subset plus synthetic data. And apparently the audio portion of Minimax was also trained at higher quality so it captures more details as opposed to ACEStep did in its training (though ACEStep devs have a similarly high quality model in their API). We still rely on the Chinese to open source these models. H3 as is, is completely useless against Suno and Udio, while ACEStep is useful.
Been trying something different for an hour now instead of just gooning stuff.Here's the biggest fail - it's a laugh:https://files.catbox.moe/s9wigo.mp4I'm going to stick with gooning. kek
>>109536214>Minimax H3Oh great another model that requires 24+GB of VRAM to use.
>/ldg/ - schizo argument general
>>109538429I use it on my 3060 12gb vram and 16gb total ram. 90/it s at 0.5mp
>>109538339awesome>>109538380Slipped some stuff from Resident Evil remake into the dataset
>>109538441>90/it s at 0.5mpouch.
>>109538441>0.5mpWhy even bother at this ridiculous low res? Might as well use WAN or LTX
>>109538360except it looks more like shockwaves than rising hot air
>>109538447>>109538450That's not bad, though. 90% of the gens being posted here are 0.4-0.7 and I see people getting 100-200/it a second.>Might as well use WAN or LTXThose are bad, though. Unusable.Are you the jewish ltx niggas trying to fud? It's not gonna work, schlomos.
>>109538450stop baiting
>>109538479>90% of the gens being posted here are 0.4-0.7Yeah don't worry, looks like the new FUD tactic is to claim gens aren't usable at anything below 1mp.still, 90/it is pretty slow compared to what you get on a 3090.
>>109538499>still, 90/it is pretty slow compared to what you get on a 3090.I should mention I'm using references on the rev2v model. It's not slow enough for me to cry about it.
>>109538499can someone give me a qrd on what you guys mean when you said "fud" it's obviously bad but there are so many bad faith posters in this thread i never know whether something you guys say means anything
>>109538508fondling ur dik
>>109538508>there are so many bad faith posters in this thread i never know whether something you guys say means anythingThen bing/yahoo/jeeves it. It's a pretty old.
>>109538427Just generate at higher MP count, that's the only way to get the SOTA unfortunately
>>109538427not bad for a fail. i've seen worse from anime studios
>>109538524Also with no speed copes like sage attention etc...
>>109538427paste the prompt, im bored and wanna see what i get rolling seeds
I gen at 4MP, I like to get the most out of these models
>gen keeps getting to 80% with no issues and then fucking dying
>>109538538I could get 3.15 seconds of 4mp footage on my hardware, but my monitor is only 3.6mp
>>109538508>there are so many bad faith posters in this threadThat's pretty much it. trying to make anons feels like they're missing out or are wasting their time.
I gen at 0.4MP, I like to use these models
>>109538561ACK
>>109538561>bro is genning on his iGPU
Using right mouse click to edit image and create a mask using mask editor in comfyui is bugged or maybe I'm a retard. I paint on the mask layer, and after saving and re-opening in mask editor the source image is just altered and all fucked up with missing parts, and the mask layer is untouched.Also, what is the best model/workflow for image to image goon stuff? Using Juggernaut9 now and some example workflow from the ComfyUI_IPAdapter_plus repo, but my results are pretty shit.
>>109538561Sad, 0.6 MP is bare minimum for acceptable quality gens (and even that is not enough sometimes)
>>109538555see >>109538596
>>109538602did you tell it to use SD3 as a reference?
So what local model is the best prompt helper for lewd H3 gens? I would have guessed gemma4 but I've heard ppl here saying it can never remember the prompt structure?Will Qwen3.6 fare better?
>>1095385990.6 MP is the true and tested res if you're 24GB 3090 Chad. Below that drop in quality is very noticeable, but 0.4 MP is alright on a simple prompts, single subject gens, not far way etc... but it's limited. I'd even say bump it up to 0.8 or 1MP if only going for 5 secs.
i used to gen 1mp but after updating it either OOM's or tries to allocate to an illegal memory address or some shit and crashes my entire pc, so i just do 0.7 now
>>1095385960.4 is great for remixing old/lowres clips.
>>109538725I gen at 0.2 because I'm impatient
>>1095387420 is even faster
>>1095386430.6mp is fine on 16gb vram, t. 5070 ti
clanker said 0.75 mp is the sweetspot.
What is the redpill on changing samplers and schedulers in Minimax H3? What can I expect from er_sde, euler, euler a, dpmpp etc?
Honestly lower MP shouldn't even be as bad as it is. Seems like they didn't do SFT on the lower resolutions, which is fair, but they should release a higher quality lower resolution version. Finetunes could also help.
>>109538776how the fuck would it know anything?
>>109538790I dunno seems pretty accurate.
Ok anons I've had enough of this woman and her plants and my gens keep crashing at .9mp so here ya go. Same seed. ll run with spectrum cope.Prompt: https://pastebin.com/jRUBMiZ7Obviously her accent and facial aesthetics are all over the place, which is funny, but generally the details all seem to be there even in the lowest res shot. Even in the .2mp, all the camera motions show up, all visual cues are there, and the layout of the apartment and plants is consistent from starting to ending shot. The only detail that seems to sometimes vanish is the water dripping out of the pot, it's gone in the .3 but retained in the .2I don't see any evidence here that the model's ability to arrange shots or follow prompts is much worse at lower res, but judge for yourself..2 mp: https://files.catbox.moe/84dva9.mp4.3 mp: https://files.catbox.moe/f0wy5w.mp4.4 mp: https://files.catbox.moe/vnjgmq.mp4.5 mp: https://files.catbox.moe/mkecwr.mp4.6 mp: https://files.catbox.moe/u695bh.mp4.7 mp: https://files.catbox.moe/67v9q2.mp4.8 mp: https://files.catbox.moe/4rwz7e.mp4
>>109538788What we currently have is the H3-Base, which is meant to output at 768p. The second-stage 2K module is coming later.
>>109538797>a few drops run over the rim of the pot..8mp doesn't do it while .2 does.
Any good loras/finetunes/models/workflows/whatever for making CRPG custom portraits and avatars? Image gen is crazy good for making custom characters but I'm just not sure how to best nail the "painterly" or "digital art" style that most of these games tend to use. Most stuff seems to be optimised for either anime or photorealism, and any attempt at a "painterly" style just ends up being overly detailed.Hell, even just using a photorealistic model to get a base image and re-diffusing that into something stylised would work, but whatever's best really.Just not sure where to start, as all my genning has been basic photorealistic stuff, never messed much with loras.Any tips or pointers?
>>109538797doing gods work anon.
How to prompt for wet plapping sex noises?
>>109538847>>109538797Yea, physics knowledge seems to have improved on 0.8 since all the ones below do it, or maybe it's seed. Anyways, 0.5 MP is the point where the quality becomes acceptable, at 0.6 it's good, and above that it's excellent.
>>109538850https://huggingface.co/lodestones/Kroma
>>109538861just prompt for a mukbang. same thing
Do you think the Minimax developers will release an update for the H3 models? R2V is good but it has issues so it'd be cool to see them release a fixed version
>take 40% pay bump for new job to buy better equipment>actually have responsibilities and can't fuck around genning all day
>>109538921>Not being a high functioning NEET with a source of income from somewhereNgmi
>>109538921good for you.
>>109538921Just make AI do your job like everyone else.
>>109538876>he doesn't know about turbo upscale pass...
I would never work full time, only part time. I refuse to give my soul and 60% of my valuable time to any employer.
so now that dust can be said to be settled. Is there actually a workflow for krea 2 raw without the turbo lora that works?
can I get good gens on a 12gb gpu or do I have to load up on cope nodes just to get runny shit
>>109538960The trick is to find a full time job that doesn't actually require much of your time
>>109538964How many times do we have to tell anons that the devs themselves are telling you RAW is not mean't to be used for genning???
>>109538968"Good" is relative, you'll struggle to get high quality at 10s or more with that much VRAM, but you can probably get reasonable lengths of .5mp if you're patient. Some cope is pretty imperceptible in the output despite anon's insistence
>>109538960im seriously blown away by how such long endless hours are normalized.The only cope reason to ever do that is if you are slaving away for a family, but normgroids will go and actually sacrifice getting a family in order to work their shit job. Baffling.Everyone thinks they are 'getting this bread' and they all end up broke by old age anyway, or at very best with a house they will have to refinance for old age care, and that's about it.
Is it a bad idea to lower the fps from 24fps in H3? I don't really want smooth motion for all my gens but I remember back on wan lowering fps generally wasn't a good idea, because the generated videos were still trained for 16fps, and lowering it just rendered a choppier version.
>>109538978Depends on the type of cope. If you're genning anime you can get away with Turbo LoRAs, for realism you can only get away with sage and few other copes.
>>109539003You can't do it. the model always assumes its 24fps. if you lower the FPS you'll get audio issues.
https://files.catbox.moe/etg68i.mp4Finally nailed it. You can't do more than 15 seconds with r2r. You can get away with it using t2v or i2v, not r2r though. 41 minutes with sageattention 0.4mp.
>>109538974but turbo lacks variety. the devs said the recommend turbo for normies because you can get good results in lightning speed. for anautist willing to spend 30x the time raw should be perfectly usable, but not a single workflow has ever been shown.
>>109539033ok? do you know how to gen? just plug RAW in the sampler and gen at 50 steps euler/simple??? why do you need a workflow for that?
>>109539010My bored ass is gonna try some fixed seed cope comparisons to try and explore these questions.
>switch to FL model>same references>Gen instantly 5x better while still using the referencesWhat did they mean by this?
>>109539060"Stop using the ref2v unless you absolutely need use a video ref" should honestly be in the pastebin
>>109539060Have you tried using the minimax h3 FL ref hybrid model?https://github.com/scottmudge/ComfyUI_MinimaxH3HybridLoader
>>109539060kwablearn2ref
>>109539060>>Gen instantly 5x better while still using the referencesHow are you using references? Character/subject references or keyframes?
>>109539068stop shilling this thing, the only examples you provided made the gens clearly worse.
>>109538850>picrel with krea 2 without any lorasit's just a character portrait with a frame, you could prompt for that without much issue. as for the artstyle and (lack of) detail, there's style loras and detail sliders (can increase but also decrease details), so it should be doable. currently there's a shitton of loras for the krea 2 model, you may want to check civitai and see if something catches your eye.btw here's the prompt i used (note the beginning and end of the prompt); https://pastebin.com/L0CS4ETm
>>109539050are you dumb or just retarded?
>>109539077both.
>>109539069The problem is if you feed audio as reference to the ref model it just kills all the other audio you might be prompting for.
every workflow ive tried (apart from the H3 default ones) feel like a bit shit. Like they dont fulfill their fancy promises. Are there any that are actually worth using?
>>109539173Default wf + sage attention and disable dynamic memory and smart memory
>>109539173Mine.
h3 is awesome, but it just takes too long to actually have fun with it. It can do cool things, but is it worth waiting an hour for mediocre quality that’s 95% likely to fail anyway?We’re just not there yet.vidrel 0.1mp kek
Do any of the pussy loras produce natural looking pussies without fucking up the rest of the image?
>>109539197I have waited four minutes for gens that made me go "oh damn," but maybe that's just me
>>109539197Umm I only have to wait like 3 minutes.
>>109539197>waiting an hourspecs?
>>109539080NTA but I've used it, it's noticeably better than ref.
Test
I swear comfyui sometimes failes to regiester newly switched reference nodes or caches data or something (I'm not smart enough to know). I'll try to do something like a basic face or clothing swap for an hour and it won't work but then it works like a charm upon restarting comfy.
>>109539280>but then it works like a charm upon restarting comfywith the exact same settings?
ideogram 4 can only nail the sims4 look every other seed, and the prompt can nudge the style with off with just one adverb, disappointing
>>109539244Yeah, same. 30 version is the best. It still follows tyhe refs while being high quality as fl2va model.
>>109539290Yes. I seems to be related to how comfyui processes the references after I've been switching them around with the same workflow for a while.
>>109538587Besides these issues, the output is also just fully black 90% of the time, and then sometimes magically it works again.
>>109537413bump. anoyone?
What are the best ways to link clips/prompts together while keeping things consistent?https://github.com/jlucasmcrell/ComfyUI-H3-Multishothttps://github.com/NikoDemon80/ComfyUI-H3-Motion-ContextI found these are they the only ones I should be looking into?
>>109539370
>>109539384Thats mine, I’ve been at it the whole day. I’ll post about this in the next thread
>>109539402so you would not recommend it anymore?
I hate nodes so fucking much. Visual vomit.
>>109539295>>109539244You made me download it.I can't tell the difference from just using pure FL.It doesn't even reference the audio properly unlike the ref model.Literal hallucinated slop snake oil.
>>109539410it’s been much more tedium than just this to stitch audio and video, but i made it work
Not happy with how this came out but I don't feel like waiting 12 minutes per gen to fix it right now.
>>109539443it's comfy
>>109539468Honestly for stitching I just use kdenlive.
>>109539466Use ref guide from Minimax page
>>109539466>It doesn't even reference the audio properly unlike the ref model.Are you using any cope nodes or lora? I don't have this problem just using sage.
i hurt myself today
>>109539486>subject_definitions:><Subject 1> is Yuuka, whose appearance comes from <Picture 1>, featuring long purple hair, a white circular halo, and being completely naked except for a blue necktie. with the voice of <Audio 1>>retention_analysis:><Audio 1>: reference - <Subject 1> follows <Audio 1>'s voice timbre.
I can't believe there are anons who don't have LLMs build UIs for them. It's so easy now, you can make it build anything you want.
>>109539491That's as many as four 4000s. And that's terrible
>>1095392433090, 32gig ram, 18 sec, 0.7mp110it/sno sage2 or whatever, im lazy.I want to experiment with prototypes and techniques, and try out more dynamic camera movements.Plus, REF needs to work better, etc., etc.H3 just isn't ready yet to seriously attempt something like anime. Otherwise, I'd love to put in the effort and deliver
you think H3 has good likeness until you bump it up to 3mp. that shit makes you salute the monitor and wish you had a pro 6000
>110it/sidiot
>old hardware >new model >"this model sucks its not actually good" lol
>>109539503I don't use cloud models and imo local llms aren't good enough for that yet if you're a nocoder.
>>109539491knew this would fucking happen kek
>>109539521Seedance 2.5 pro is in another universe anyway. What's the point?
>>109539549H4 will mog it
>>109539513>H3 just isn't ready yet to seriously attempt something like anime.I think it is ready after seeing that anon and his BLAME! kino desu
>>109536306Lora training ?Download her KH3 3d models and take the screenshot from every angle
>>109539491That's pretty bad and it's going to get worse.Mark my words 5090 is going to hit 10 grand next year when things get really bad with the availability as more and more people start building local AI.I'm so fucking happy I got myself a 5090 just as the prices exploded. This thing is the best purchase I have made hardware wise.I should probably legally have to marry this card considering how much cum it has drained from my balls through AI generations of all types.
>>109537426Enable ref_image_size to Max and have a hi res of 360 view of your character face referencesIts slower, but it much faster than setting your megapixels to 1.0
>>109539513>H3 just isn't ready yet to seriously attempt something like animethat's like saying, "we're not ready to seriously attempt something like anime because we only have pen and paper." when it comes to talent, the tools don't matter nearly as much. anons have already made kino anime scenes with H3. if someone genuinely wanted to make a full episode, they theoretically could
>>109539491saw this going for as low as ~7k, shit's crazy
>>109539491>even if you could afford it, we're just making it more expensive lmaoi hate these kikes and slant eyed kikes so much fuck
>>1095376930.3mp is shit for realism
>>109539491when will china gives us an alternative to these nvidiakikes?i'm just tired man...
>>109539491china will save us, no biggiehttps://www.reuters.com/world/china/china-begins-making-homegrown-duv-chipmaking-tools-information-reports-2026-07-27/
>>109539491>GPU costs more than my car lmao
>>109539629you couldn't get your hands on one any way if they kept it at the same price
>>109539653they were sitting on goybay for like $7k-8k >>109539617sitting there, staring menacingly.
>>109539646trump will hit these with 500% import tariffs. you'll never even see one, let alone be able to afford one
>>109539669and yet anon didn't buy
Just use 8step turbo lora and set it at 0.7mp + RTX upscale 1.25/1.5x.Enable preview VAE to see if your gen fails or not.Done. You just created a porn machine. I even tried it at 20 seconds on my 5070ti. 10 to 15 minutes generation. No cope nodes except Sage attention.
>>109539671that fat orange jew couldn't stop brown people from pouring in from the southern border, he's not gonna stop giga autists from building waifuboxes either.the next prezzo might thoughhowever.>>109539676sh-shut up, i kick myself for this every day in regret.
>>109539681nice video example
>>109539688If i posted it i got an IP ban so no
>>109539701box it
>>109539701no one gets IP bans for catbox links
I'd feel bad about how this stuff is gonna kill the blender porn animation scene if it wasn't for the fact that half the cunnyfags disappeared and the ones that are left are slow to release, only do renders, or suck at animating.
>>109539709Does local AI really render Blender redundant? What about 3D modeling and rigging for game assets?
So it's settled then.
>>109539709Why do you care? Those people need to get a job.
>nodesYou're not supposed to use the sage/flash attention flag?
>>109539720>What about 3D modeling and rigging for game assets?This will be pretty irrelevant in 5 years. The more the tools get better making everything easier thus lowering the bar also lowers the wages and increases the pool of people fighting for these spots. Art niggas getting BTFO in every way.
>>109539720Modeling and rigging? No, not yet as far as I know. Porn animations? With H3 pretty much. >>109539722>why do you care?I like hand-crafted stuff made by people who enjoy the same things as me.
>>109539720Not sure about rigging but I know there's an open source model for motion capture, conceivably you could use H3 to generate a movement then use the motion capture model to map it to a character.
>>109539709Art has always been gatekept in some way, whether by the extreme skill required or by wealth. Right now, the main barrier to entry is money. If you had, say, £100k to throw away on a GPU workstation, you could probably dominate the 3D porn industry. I'm surprised some rich fag hasn't done it already
Drawing isn't irrelevant yet, and I suspect it won't be for along time.Comics and manga still can't be created effectively with AI. Sure nanobanana can create pages, but not a sequence of 20 pages (a chapter) that maintain consistency and tell a story. Doing this with AI would unironically take longer than just drawing the manga, and that's saying a lot because it takes a long ass time to draw lol.Also drawing is still better for creating original characters. You could of course combine the skill with AI to create loras of your characters, but it still starts with drawing.
Am I going to have to use audio references for sex noises?
>>109539739>lowers the wagesNo, there will just be less jobs available and only a select few (mainly those already in the industry with connections) will get to direct the AI. Everybody else looking to enter will be shit out of luck.
>>109539721the only use for these is if you use them as a NSFW prompt refiner in the same workflow.
>>109539721>So it's settled then.??? was there every any confusion about that
>>109539781Probably, h3 does really good with vocalizations but does pretty bad with any sort of body sounds.
>>109539785Yes. Absolutely. 100%
Is there a prompt for light skin in h3? I tried light, fair, pale but it always gives me at least a tan.
>>109539782>No, there will just be less jobs available and only a select fewLess jobs per company/studio yeah but the wages are for sure going to lower due to how easy to replace the people will be. The easier you are to replace the less you'll get paid. This is pretty basic and no the current people in the industry will not survive they will be priced out by indians.
>>109539720It's all going to go away. I can already import characters and animations and make basic scenes in blender using Claude prompts. In a few years you'll just gen a video and you'll be able to turn that video into a blender scene with fully rigged models, all local. That is if anyone even cares about the scene being in a 3D engine anymore.
>>109539794Porcelain.
>>109539791I'm not a datascientist but the idea of taking something that another model was trained on and using it as an input, then changing a significant portion of it and it's outputs seemed like a really dumb idea.
>>109539806you underestimate the stupidity of the average genner.
>>109539801Blender has an MCP server, why not just use that?
>>109539810Cool but its face looks too goofy
Are we going to clear Impossible mode?
>>109537214Ostris quietly fixes bugs without acknowledging them on his github repo. At least close the issues when they're fixed instead of letting them expire, so people will know if a certain bug has been fixed. He's acting like he's the person finding and fixing all of the bugs himself.
>>109539810I thought he was making himself smile. that would be much more creepy than that funny little kraken xeno baby face
Is Minimax the model for doing blender porn? Haven't seen anything impressive posted in these threads in that style
>>109539851Yeah, lemme whip up a gen real quick.
Bros I'm so done with the ref model being so inexpressive. But FL can't clone voices for shit.
>>109539759You know what really gatekeeps art? Imagination and intelligence. We're already in a state where any idea of gatekeeping in art should have disappeared but it didn't go anywhere.That's because most people are dumb as shit and incredibly lazy. Even with all of the tools available for them in image generation that can be run locally, very few people have managed to make anything even remotely interesting.Turns out that you need not just imagination, but also fundamental understanding of things like cinematography to get anything interesting done, along with having a good visual library that allows for you to come up with stuff.Artist with access to 3D will always beat a non artist using the same tools and this won't change.
Im a 3d artist and i can see this will be the end of me.I still hate at Fanbox for not allowing AI
someone make a new thread
Reminder that gatekeeping is a good thing. Just look at the absolute state of this website as an example.
i tried to gen 4mp, 30 seconds with h3 to see if i could just let it run for the rest of the day and overnight for some morning kino, but it just gave me a blue screen lol. my 3090 did not like that one bit
>>109539904>this will be the end of me.Only if you let it. Use those skills to take advantage of these new tools and make some cool shit.
>>109539911that's a little overkill. In my testing I found that going slightly over your hardware limits with dynamic vram enabled produces garbage on every step, going a lot over just crashes comfy, and going way over apparently kills windows
>>109539916Im ready to jew this shit out anon. It just Fanbox is not allowing AI anymore
>>109539842
>>109539948kek, you're so close to making it creepy then you do silly stuff lolI like the gen, though
>>109539958how would you make it creepy and not silly?
>>109539948It's a shame Stuart Gordon is dead. Imagine what he would do with this tech.
>>109539933What sites do? I think patreon does but their content rules are strict.
>>109539948hehe
>>109539966Keep the lifeless eyes.
>>109539972Unifans. Koikatsu animators goes there after the Fanbox cuckery, but no one uses those. Well you will see me on the street in few years. Spare me a few bucks anon.....
i've generated everythingwhat now?
Let's say your generation takes 500 seconds, then you use the cope nodes to bring to down to 300 seconds, then you increase steps to bring it back up to 500 seconds.Will your quality now be higher or lower?
>>109540007Generate everything again, but add "huge tits" to the prompt
>>109540007generate everything again, but include yourself (you) in the gens
>>109540018it added a flock of big bouncy birds to the video
>>109539597>i had my ref image size set to matchwhy the fuck, thank you king.
>>109540031that can't be healthy for my brain, but fine, ok
>>109540018>>109540031there are two wolves inside you
which text encoder do I want with h3? theres 15GB, 27GB, and 51GB. Is this loaded in vram or ram? I have 24GB vram / 64GB ram
>>109540044Its very useful for realism. But not so for animes.When i set it on my gen times goes from 300 seconds to 400 seconds (12 secs/0.7mp) but its worth it
>>109540063Does it help with image quality or just with adherence to the reference?
>>109540057which one do you think jackass?
>>109540070no idea. i only try it on face and it works
>>109540017shut the fuck up.
>>109539851https://files.catbox.moe/qj7xcn.webm
Iron man tetsuo>>109539948this is great
>>109540075not the 51GB
>>109540092Just closed the blender tutorial tab I had open for a month
Ok H3 is perfect at low res but as you climb up to 768p I'm tried of my outputs AI slopping and ignoring my prompt. It's also seed dependent but I don't have time to farm.Is there a workflow where I can gen a low res latent/video then feed that as a source to gen a high res one? I don't know if that involves reference to video. I've had workflows like that before with wan back when the idea was to start with low res with no speed ups as a first pass, didn't know if that worked with H3.
>use chroma to gen some 1girls for your gens>start getting attached to them
Can others confirm that the res model completely kills any prompted audio?
>>109540130You're saying prompt adherence gets WORSE at higher res?
>>109540057qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors Already good. I tried the 27gb and theres no difference
>>109540130>>109540144Sounds like bullshit but I don't believe it.
>>109540144I had to watch a lady water plants for like an hour to conclude with weak confidence that at least it doesn't get much better at higher res.
Baker?
>>109540144Yeah. At low res H3 is a champ and will do almost anything and it looks real, just blurshit. But at high res some prompts will work great and I can gen 1080p all day and all outputs will be acceptable, but most prompts have dice roll seed dependent outputs at high res. This means I can't do the prompt I want and I have to go back to the prompt I got working and just settle for it.
>>109539851https://files.catbox.moe/1d1g6i.mp4It will be one we get a proper sex lora so it doesn't butcher genitals.
>>109540183I have found that at low res any face that isn't pretty close in frame gets significantly distorted. Other stuff looks ok and faces that are at least medium shots or closer look fine, but there are definitely things that don't look "real" at low res
>>109540195>>109540195>>109540195>>109540195
>>109540144>>109540183Also let me be specific, the prompt adherence isn't bad. It just ignores a few things. The real problem is sloppification where people look waxy and the composition / light realism sucks. The subject and setting and what they are doing will still generally be what I ask for but things like camera and lighting go wrong.
>>109540208Is it just above 1MP, or some cutoff that you see it happening or just steadily worse up the scale?
>>109540193Yeah very low res has fucked faces like you said. I think the tipping point for sloppification risk is somewhere around 640p. So if you want to quickly see if a prompt is going to work at high res I do 672^2
>>109540227I think its a cutoff. If there's also a gradual decay it's not as noticeable as the cutoff. I did some genning and every 640p gen and below was fine but going to 672 or up was slopped. This isn't consistent at all. It depends on the prompt. Some prompts I will never see a real looking output at high res and some prompts are solid.
>>109539513You're generating relatively complex scenes at subpar settings anon. Of course it'll suck at anything less than 1 MP, everything lower is just a cope.
>>109540259Also w/ a 3090, you're doing something wrong if best you can get is 0.1 MP in one hour, you can do 5 secs of 0.6 MP faster in 4 mins on a 3090.
>>109540146thanks
So is local vid gen officially better than photo gen?
>>109540426Either can be used to make videosPhoto gen is better for any content in any styleVideo gen is more limited and harder to make loras for, but one place it seems to have an advantage over photo gen is correctness. I don't see anatomy gaffs with video gen, or incorrect reflections and refraction. Image gen isn't bad at this anymore and will probably be perfect in the near future but it's weird that video gen got ahead of it when it started with nightmare will smith eating spaghetti.
>>109540490>I don't see anatomy gaffs with video gen, or incorrect reflections and refractionYeah this is basically what I'm talking about.
>>109539683>>109539671>>109539646oh this is my favorite copie. China is more greedy than anyone else in the world, they'll just you even more
>>109539491GOT MINEFuck I feel like gen x cancer (I am not a cancerous boomer jr) ugh what an awful thought
>>109541072i was gonna buy one about 2 years ago when they were like $6k but couldn't talk myself into it even though i could easily afford it.. now im kickin myself
>>109538384He really didn't like this post
>>109539785there wasretards thought it could generate prawns for themhe covers it in his post targeting such audiencehowever heretics and alike do have a valid use instead of vanilla especially corpo qwen>what usenone if you have not tested these things thoroughly
>Video Model Showdown: Open-Source vs. Paid AI Video Modelshttps://xcancel.com/ComfyUI/status/2087585018130710975