[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1786248182339509.jpg (182 KB, 576x512)
182 KB JPG
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109503671

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
>>109505489
>no will smith gif
not a real bake, I'm leaving!
>>
>>109505497
see you soon
>>
I was genning with Wan 2.2 just 2 weeks ago, didn't even know about LTX (didn't visit /g/)
Had a fun and frustrating week with LTX discovering its strength and audio for the first time
And now just a week after, I'm playing with H3 and finally body horrors of LTX are gone, no more missing fingers or fucked up tits growing from the back or face on the back of the head.
But checking back on my 2 week old gens, honestly, impressive as H3 is, I2V+text doesn't quite get you what Wan did where I could prompt something simple along
>real life video
and get picrelated. (18 seconds, Wan SVI workflow, the speed of of gens is comparable)
Minimax really wants to change body shape, create face that is ugly, ot keep it videogamey anyway. Tits physics is also not there. Wan: https://files.catbox.moe/qb567e.webm

Though still, this is a huge for us and big leap with for me in 2 weeks discovering I can do audio locally
>>
>>109505495
>has a literal crime record
Lmao.
>>
>>109505497
>will come back begging for bbc pregnant fart fetish futa porn
>>
new kroma version is out
https://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.2-turbo.safetensors
>>
File: HOPJ27Ga0AAtyTg.jpg (119 KB, 640x1216)
119 KB JPG
https://files.catbox.moe/30vzn4.mp4
>>
>>109505509
my theory is that they started the h3 training on well-captioned 3d rendered stuff so they could get the model to learn movement quicker, and then they finished with mostly real stuff. that is why the less common real stuff gets automatically turned into video game graphics
>>
>>109505523
>Doesn't track the eye movements
Garbage, improve prompt
>>
>>109505525
nope. They said they did hardly any post training, just a ton of pretraining on varied stuff. Its the opposite. The look your thinking of is gained by tons of post training / RL
>>
>>109505521
>26.3 GB
>turbo
what the fuck?
>>
>>109505509
>>109505525
> well-captioned 3d rendered stuff
makes sense, I've seen some gens where the movement of realistic subjects look very stiff
>>
larryvrh/MiniMax-H3-Turbo-Lora
Does this work on ref model?
>>
>>109505539
fp32, wait for fp16 I guess
>>
File: I love China so much.png (122 KB, 1641x439)
122 KB PNG
>>109505521
Who cares? Minimax image will soon be out and it'll destroy anything.
https://xcancel.com/MiniMax_AI/status/2086253065657790895
>>
>>109505554
not sure if they even could beat krea.
>>
>>109505552
that cant be a lora. is it a full distilled model?
>>
>>109505552
I'd like it more as a lora
>>
>>109505559
it'll definitely destroy klein, edit models really need to be improved
>>
>>109505509
maybe there is, something good about dual model for high noise/low noise split technique after all
>>
>>109505533
oh. so they didn't really write about their training regiment? i'm guessing they wouldn't want to write down about how they got their copyrighted data lol
>>
it's been 40 minutes, where is kroma-v0.2-turbo-int8-convrot?
>>
>>109505592
convert it yourself
>>
>>109505592
it would take all day to download it thobeit
>>
>>109505497
i'll stay here because fuck the collage
>>
>>109505554
this, i already deleted most everything and will delete more once it arrives. Almost everything gone, controlnets, facedetails sdxl illustrious wan ltx every into the trash.
>>
It's been 40 minutes, where is my 2 second long gen?
>>
https://www.reddit.com/r/StableDiffusion/comments/1vj5scs/deroping_minimax_h3_fast_motion_to_reduce/

https://github.com/matlowai/ComfyUI-MAINodes

>H3 smears bursty motion: backflips, fast sword arcs, whip-fast reversals. The cause is structural. One latent token spans four pixel frames, and at high motion speed those four frames need four distinct poses that a single token can't hold. Re-denoising the affected region doesn't help, because the missing poses were never generated in the first place.

>This pipeline works around that at inference time. It re-generates the clip as a slowed-down version of itself, seeded from the original. Frames where motion is too fast get held (repeated) so the model has more temporal room, the result is generated video-to-video from that retimed init at partial denoise, and the original frame rate is recovered afterward by dropping the held frames. The oracle that decides where to slow down reads the clip's own latent. No extra model, no training.
>>
>>109505708
it makes inference slower though right?
>>
>not being born rich to enjoy this shit on 2x 6000 pros
Kek, what a meme life.
>>
>>109505708
Why have you not yet purchased an advertisement?
>>
File: images-4249855935.jpg (23 KB, 600x600)
23 KB JPG
>>109505644
don't make you jump scare you anon, I can be scary sometimes but not all the time anon.
>>
>not being born actually rich to enjoy this shit on a rack of DGX B300s
>>
>>109505710
it's ok anon, in a few years they'll manage to make something as good as H3 but on a 6b model
>>
Under 60 seconds for 10 seconds square video with the turbo lora and only losing some sharpness and audio quality. All the accelerators suck. They give you turbo lora quality and shave off 10-20 seconds from base.
>>
>>109505710
you dont need to be rich. get a credit card, max that shit out and make small $200 payments monthly. treat it like a car payment.

YOLO
>>
i have learned almost every horror this model has. it knows a lot and a lot of audio also
>>
>>109505554
i only believe when they release it
>>
>>109505718
If I had access to something like that, I would also train a bunch of complex Loras and share for free
The wrong people are rich
>>
>>109505733
lol
lmao even
>>
>>109505733
>The wrong people are rich
you always have to sell your soul to be rich so yeah...
>>
>>109505733
I would also try to save local musicgen
>>
>>109505710
I don't think anyone but researchers who get GPU access for free are genning with that
>>
>>109505738
where do I sign?
>>
>>109505728
or just dont pay it until its passed between so many debt collectors they will let you pay pennies for it. just hit em with the ol "ahh jeez duuude.. i got these pills man i can do like...$5 a month?"
>>
>>109505708
Why do people has to include 70MB of shit their repos?
>>
i daydream about winning the lottery, setting up a dedicated gpu farm and living the rest of my days just gen'ing.
>>
>>109505785
many such cases
>>
>>109505554
Isn't H3 a distilled model? It will just be pure slop just like Flux 1 dev. H3 is currently slopped for pure T2V too, just slightly harder to notice on a model this big.
>>
>>109505816
it is only CFG distilled, the flux dev models are CFG and step distilled
>>
>>109505816
>Isn't H3 a distilled model?
no. if it was, we wouldn't need turbo loras
>>
>>109505495
fuck off and kys
>>
>>109505554
> soon
> china
>>
if using int8cr its probably best to disable fp16 accumulation if you have it as a cli argument since it will only lower the calculation precision of things other than the main model, that actually have fp16 operations, like the vae, which you dont want lowered.
>>
>>109505840
Yeah, hopefully we are not into another "chinese culture" cycle again like we had with the z-image releases (where the Edit model ended up never being released btw)
>>
>>109505834
it is distilled, it's guidance distilled, the turbo lora adds a steps distillation process on top of it
>>
why are models so lenient on nipples, but so consistently fucks up how they look with that like weird double ring look
i understand the gooner shit being filtered out or censored or whatever but why is the one thats more acceptable so consistently fucked?
bottom text
>>
>>109505862
>the Edit model ended up never being released btw
that's because migu is still alive... >>>/wsg/6208995
>>
>>109505889
this
>>
>>109505816
nobody asked them why they decided to not release the base model on that reddit AMA?
>>
>>109505509
>wan
>https://files.catbox.moe/qb567e.webm
Guys how long will it take to get us soft tits lora like this for H3?
I'm getting tired of air balloons in h3 that look and behave exactly the same across gens
>>
>>109505521
my outputs are all weird
>>
>>109505509
>LTX discovering its strength
?
>>
>>109505910
>my outputs are all weird
when will you guys fucking learn that kekestone is a fraud?? holy shit he's been pumping shit models for more than a year at this point and you still believe he can pull it off? kek
>>
>>109505920
his butt buddy S1LV3RC01N seems pretty knowledgeable. he should just take over
>>
>>109505869
we've seen <safety> like you wouldn't believe and countless faces have been palmed

in the end the "gooner shit" is what actually works for everyone from <questionable> up
>>
File: 1754768905239866.png (124 KB, 1950x505)
124 KB PNG
If you want your fetish trained into sulphur now is the time btw.
>>
>>109505916
After 2 years of Wan prompting, with LTX I could now request something much more specific to happen with actions ordered in prompt (often including abominable body and limbs movement, but still), while Wan prompting was vague and it would simply ignore the specifics. Also wan gens looked way too similar and seed was barely relevant.
It's all still fresh in my head and I can compare as I used all 3 models in the past 2 weeks
>>
>>109505940
Consider purchasing an advertisment.
>>
>>109505940
lemme sneak some war kinos into there
>>
>>109505940
brb uploading some loli
>>
>>109505943
Oh, and 20s music + image to video was amazing first time experiencing in LTX, with fast movements and looked great physics loras, while Wan was really reluctant on doing any fast movements at all even with lightspeed loras
>>
Did anyone try training any video model on goon compilations? Any LoRAs of that? I doubt either Wan or LTX could learn those properly so I'm thinking of training H3 for it.
>>
File: 61004582548272.mp4 (3.96 MB, 640x832)
3.96 MB
3.96 MB MP4
>>109505907
They don't always behave the same. If you're doing i2v I think it depends on the medium of the image, for example 3D renders look stiff because it's emulating relatively stiff 3D animation.
>>
Gening is nothing more than degenerate gambling.
>Bro I swear bro next sees is the one, just one more step bro, just one more lora bro, just one more seed bro
>>
File: MiniMax_H3_00363_R.mp4 (3.74 MB, 832x1280)
3.74 MB
3.74 MB MP4
>>109505907
if h3 think the subject is real life, the breasts jiggles fine. maybe try to prompt "subject's breast is soft" or some shit.
>>
>>109505986
wan learns really rather well, but I suspect so does h3 (haven't tried).

obviously it depends on what you mean by "goon compliation". perhaps you want it to at least sort-of have some more limited focus unless you have too much time/compute to caption and train.
>>
Why is ComfyUI's default H3 workflow using nearest-exact instead of lancoz?
>>
>>109506021
IDK, I switched it to lanczos without issues.
>>
>>109506021
processing time
it's too intensive to use anything other than bicubic scaling

>>109506006
>he hasn't founded the preview settings in the settings
>>
>>109506021
>comfy raping your input image more than any optimization to save 2ms of cpu time
>>
https://www.youtube.com/watch?v=S9O3FPumX4Q
>>
Damn I fucked up ref2 video. Tried two reference images for characters and a 15 second animated video and it just showed the reference image 1 as the first frame, played like 4 seconds of the original video, and then morphed the characters into weird proportions over top which kinda matched but the background was all fucked up. I was using the default workflow + spectrum + sage. 9 minute gen time...
>>
>>109506006
yeah and?
>>
>>109506038
noob
>>
Hitler, Flux employee, finds out Minimax is better at generating Miku

https://files.catbox.moe/0hht5s.mp4
>>
>>109506037

Open-Weight Strategy: MiniMax chose to release H3 as an open-weight model to foster innovation, allow local deployment, and give businesses the flexibility to adapt the model to their specific security and data needs (3:28 - 4:28).
Multimodal Generation: H3 is a general-purpose model capable of text-to-video, image-to-video, first- and last-frame generation, and in-place video editing (2:51 - 3:08).
Native Audio Sync: One of the standout features is its ability to generate synchronized stereo audio (including dialogue and sound effects) simultaneously with the video (5:51 - 6:02, 23:42 - 24:25).
Performance: The model is considered a significant step forward in the open-weights space, performing on par with or exceeding some state-of-the-art closed-source models in specific arenas like video editing (7:30 - 7:56)

Optimization for Consumers: H3 has approximately 60 billion parameters, which would typically require 120 GB of memory (BF16). ComfyUI and the community enabled it to run on consumer hardware through techniques like quantization and fine-grained offloading, which keep only the necessary computations on the GPU while offloading the rest to system RAM (25:29 - 27:18).
Resources for Developers: MiniMax provides a Context-IR API to optimize prompts for users working locally, helping the model better interpret complex cross-modality references (9:01 - 9:21).
Future Developments: A 2K-resolution regeneration API is currently available through the MiniMax hosted platform, with ongoing development for broader integration (19:55 - 20:20)

Rapid Iteration: The team highlighted the community's impressive speed, noting that quantizations and hardware-specific support (like MLX) were delivered within 48 hours of the model release (10:17 - 10:46).
Prompting Advice: The ComfyUI team suggests using the official prompting guides and, for dialogue, explicitly including the spoken text in quotes within the prompt to improve lip-sync consistency (28:12 - 28:50)
>>
>>109506038
> <Picture 1> and <Picture 2> reference and represent <Subject 1>. <Picture 1> is directly integrated as the first frame
>...video starts at 00:00 with <Picture 1> ...
>>
>>109506048
add english subtitles and this would be peak kino
>>
>prompt camera pan to show her ass
>she start shaking her ass unprompted
>>
>>109506065
It knew you are black(also brazilian).
>>
>>109506065
What did you expect, you used a picture of a negress
>>
there we go, got the proper first shot, used a 0 to 3s: timestamp.

https://files.catbox.moe/z8kg34.mp4
>>
Saar, realism, photographic saar.
>>
with all the security concerns using these indian/chink nodes with volatile code, docker setup would be perfect if you don’t want to dual boot. it was also genning somehow 10% faster than on bare metal windows when I tried it.
Unfortunately vfs volume mounts are painfully slow and model load times atr terrible, getting like 50MB/s disk read speed off my SSD, while normally they load at 2.3GB/s in comfy portable.

Is there a way to fix this without storing models inside the docker container /g/?
>>
>>109506128
>>
File: screenshot.1786269726.jpg (39 KB, 580x218)
39 KB JPG
>>109506128
you have to store the models inside WSL and enable it in docker desktop.
>>
fucking bodied that freak
https://files.catbox.moe/6nw0nn.mp4
>>>/wsg/6210814
>>
>>109506146
>don't want to dual boot
>>109506155
the reason I specified
>without storing models inside the docker container
is my hoarded models are scattered across different SSDs and HDD and wired in comfy's extra_model_paths.yaml
there is no way to fix slow docker mounts? Why are they slow in the first place, seems like a bug.
>>
>>109506038
it's not on you anon, the model is scuffed. it doesn't get it unless it's a simple put Picture 1 in Picture 2 situation.
>>
>>109506181
>>don't want to dual boot
did i say duel boot?
>>
why were people shilling res multistep? euler looks way better
>>
oh yeah i get infinite hangs at step 0 if i try a big gen on my 5070ti+32gb like 10s@1.5MP but somehow the same runs fine when it's an instance claude span up from WSL2, which is extra weird since it just spins up powershell from there. i'm definitely finding comfyui is a bit painful to prompt H3 with since there's so much boilerplate text you can fuck up which varies between t2v, i2v (f2v, fl2v) and that's just on the normal model, if one of the a1111-likes has a good H3 prompting experience i'm quite keen for that. there's enough variety to it that it's not even particularly trivial to vibecode a good prompt builder ui, though a simple one can streamline the process a bit. and having a local text model write/check your prompts is also a pain since it's yet more loading and unloading per iteration of a prompt
>>
>>109506181
storing it in wsl is not inside the docker container.

>there is no way to fix slow docker mounts?
no.
>>
https://files.catbox.moe/4jd3nl.mp4
>>
>>109506197
proof? did anyone do a comparison?
>>
>>109506197
Multistep for non-turbo, euler for turbo. That simple.
>>
>>109506220
>Multistep for non-turbo, euler for turbo.
multistep works fine on turbo though?
>>
FL2VA vs Ref2VA, which is better?
>>
>>109506216
>did anyone do a comparison?
yes. i did
>>
>>109506226
depends on your needs and prompting skills
>>
>>109506228
forgot the
>proof?
part
>>
>>109506235
it is not possible for me to prove it because i can lie. try it yourself. it's not hard. you CAN use H3 on your computer, right?
>>
File: 1768662091638257.png (14 KB, 295x438)
14 KB PNG
>>109506246
showing a side by side comparison would be good enough, im already genning many queued videos
>>
>>109506226
ref lets you do pretty much anything you want, though I sometimes struggle to prevent it from carrying over certain qualities of the reference
>>
>>109506197
Shut up Leonard
https://files.catbox.moe/4ul931.mp4
>>>/wsg/6210819
>>
>>109506252
ok, test it out when you finish
>>
>>109506233
>>109506253
I wanna make placeholder idle animations for torso+face portraits. Using ref image
>>
>>109506263
ref >>109506056
>>
>>109506252
>25 steps
interesting, have you found reliable improvements from this? i've often tended to find if i blind A/B test, running models at the top end of their recommended step range or a 10-20% above it is a good way to get more reliable gens, but i haven't tested many things with H3 yet since there are so many variables and it takes multiple minutes per gen
>>
>>109506263
I grant you the authorization. Next.
>>
>>109506271
>running models at the top end of their recommended step range or a 10-20% above it is a good way to get more reliable gens
yup, i didnt do direct comparisons but essentially 5 extra steps arent gonna take too long and will probably iron out some more things, on 3090 128gb ram for now i settled for i2v 0.7MP 8s 25 steps which take 10min per gen, sage 2.2 and default spectrum.

I can probably decrease some of these things but i want to gen at the higher end of quality so that i get used to good gens before dropping things down so i can then know how much will the later gens deviate from the known good ones.
>>
do you guys use any uncensored loras for minimax or you reckon the base model is uncensored enough?
>>
Github is down or what? Can't access it today
>>
>minimax_h3_ref2va_pruned_int8_convrot
is this the way to go?
>>
>>109506373
yes
>>
do I need both nodes for the sageattention 2.2.0 workflow? Minimax H3 Mem Eff Sage attention AND patch sage attention KJ node?
or just one of these?
>>
File: ComfyUI_00075__2.png (2.68 MB, 1792x1152)
2.68 MB PNG
>>109505910
the "turbo" is the base model it seems accidentally. Use turbo lora on it and it looks good. Way better than before with details
>>
>>109506396
model -> patch kj -> mem eff
although in my case even without those it seems like the speed is the same if you already have use sage as cli arg
>>
troll bake
>>
>>109506171
link to mod???
>>
>>109506373
pruned is worse quality btw. watered down
>>
>>109506429
i believe sageattention is a global setting so toggling it on with a launch flag does the same thing as the node that toggles it on, and it stays on for ALL your workflows in that running comfysession if the node turned it on since the node really does just flip the switch. other sage node stuff is separate.
>>
>>109506449
literally the same outputs
>>
File: 129.jpg (55 KB, 680x772)
55 KB JPG
>>109506449
>>
nah that unc is fr but you dont notice it until you go back to your raw as fuck gen and compare it to cope and prune
>>
>talking in third person
>>
>talking
unc turn off your narrator this is a text board
>>
File: 1762867827997553.mp4 (467 KB, 800x672)
467 KB
467 KB MP4
>>
>>109506453
no, it's not, otherwise there'd be no reason for a "prune", anon.

original int8-convrot is 34gb
pruned is 21gb
that's a whole 14gb chunk of sexy data missing, anon.

you think that's "literally the same"? why do you think half the thread has garbage gens.
>>
>>109506409
how bizarre
>>
>>109506480
nice animation, still not using troonix though
>>
>>109506487
the pruned part literally did nothing, that's why they pruned it
>>
>>109506487
>that's a whole 14gb chunk of sexy data missing
that did fuck all and was substituted with a lookup table
>>
You can actually remove the diffusion model altogether and substitute it with your imagination, only retaining the llm.
>>
Why is Julien such a desperate, worthless bottom feeder?
>>
>>109506487
prove it, use both and show us the difference on a same settings gen
>>
File: i_00063_.png (1.3 MB, 768x1376)
1.3 MB PNG
so, how do you dudes pass the time while waiting for gens to finish baking in the GPU oven?
>>
File: bc936j.png (218 KB, 1000x600)
218 KB PNG
>>109506528
>>
>>109506541
I accept your concession.
>>
why haven't the comfy team used this technique in the past to dramatically reduce model weights?
>>
>>109506547
bodied that freak
>>
naw bro busted out the capital letter and period cuh think he tuff naaaahhhh :skull:
>>
>>109506554
shut the fuck up
>>
>>109506550
it only works when the model maker did a big oopsie
>>
>>109506550
because its not a usual prunning but instead removal of a particular part of the model that the creators put in that was known to be shit architecturally, its not random 13b worth of random model weight data that got pruned
>>
cuh crashing out :sob:
>>
>>109506538
get started on the next one
>>
A few days ago, I already complained that the forum has turned into a showcase for TikTok slop.

Some people have brushed it off, saying it’s always like this for a few days - I strongly disagree. H3 is so easy to use that any kid can churn out this slop in bulk, and as long as it gets likes and isn’t removed, more slop will keep coming.

Until the mods do their job, this spam will go on forever. Valuable or interesting posts are buried under all this trash.

People are now even calling their trash “another AI slop spam” right in the title, and yet the garbage stays online.

It’s annoying, and the TikTok crowd will, of course, disagree with me.
>>
>>109506590
oh hi fellow ledditor
https://www.reddit.com/r/StableDiffusion/comments/1vjmp5z/ai_slop_spam_continues_mods_do_your_job/
>>
cuh a reddit tranny ahh nga i knew it fr
>>
>>109506602
Shalom (blessed day) saars
>>
zoomspeak is basically
goo goo gaa gaa
>look im pretending to be retarded
>>
>>109506644
>goo goo gaa gaa
they literally made a meme with a character they called goo goo gaa gaa, I'm not joking
https://www.youtube.com/watch?v=ziIk72mUz5c
>>
>>109506644
you're a redditor
>>
which model is the current sota for 2.5d anime gen?
>>
>>109506667
Wait for the H3-derived image model
>>
>>109506661
>they
>>
File: 1777885300138334.mp4 (241 KB, 608x512)
241 KB
241 KB MP4
>>109506490
>>
File: tsunbot.webm (1016 KB, 544x736)
1016 KB
1016 KB WEBM
>>>/wsg/6210839
>>
>>109506678
yes, "they", the zoomers
>>
>>109506677
what about now?
>inb4 no answer
>>
>>109506690
the youngest gen z are 14, they're not the ones watching an anime penguin girl go gugugaga
>>
>>109506667
Anima
>>
>>109506709
they're retarded enough to watch that shit, regardless of their age, a truly wasted generation
>>
>>109506661
you know it's a bunch of 40 y/o single mums that watch this crap the most
>>
>>109506709
>the youngest gen z are 14
And the oldest are 30, unc
>>
>>109506709
true that, gugugaga is peak millennial boomer uncslop. on god, fr fr, no cap.
>>
>>109506720
don't you have clouds to go yell at?
>>
unc is crashing out bruh
>>
>>109506745
>peak millennial boomer uncslop
I've grown with final fantasy X, gta san andreas, tekken 3, you've grown up with goo goo gaga and body type A/B in games, we're not the same
>>
>>109506755
at least it isn't bother CJ crash out over ani
>>
>>109506758
>final fantasy X
not 6 and 7.
>gta san andreas
not the gta demo from pc gamer with a trainer to remove the time limit.
>tekken 3
not virtual fighter in the local arcade.
always a bigger unc in the retirement home.
>>
>>109506767
what is blud even typing
>>
https://files.catbox.moe/do9bqu.mp4 Is there any way to control the speed? I sometimes get gens look too much slowmo
>>
>>109506771
ff7 sucks, it got popular only because it was the first one that got released in europe
>>
spectrum node rapes my system, everything just keeps freezing
>>
>>109506776
video sigma shift down a bit, like from 12 to 10ish
>>
File: based.png (1.32 MB, 1125x1096)
1.32 MB PNG
>>109506776
finally some good shieet
>>
>>109506785
whats video sigma?
>>
>>109506791
sigma balls hah gottem ha ha
>>
File: 1671465111031801.png (909 KB, 1076x942)
909 KB PNG
>>109506788
>>109506776
>pov fingers change into done nails
>>
File: 1776501800349814.jpg (87 KB, 1676x564)
87 KB JPG
>>109506791
>>
File: yup, that's me.png (464 KB, 1080x607)
464 KB PNG
>>109506803
true heterosexual men only watch lesbian porn, why would you want to see another naked man while jerking off? are you a faggot?
>>
>>109506808
That's a big node...
Come on... say the line...
>>
>>109506776
nice
>>
>>109506538
Hop on the treadmill, i actually get fitter and become a more powerful gooner the slower the gen time.
Ironic isn't it.
>>
>>109506813
reddit ahh nga
>>
File: 1784188939598520.jpg (25 KB, 962x144)
25 KB JPG
>>
>>109506825
https://www.youtube.com/shorts/lrrMJqjgsIM
>>
>>109506820
you sound like you were molested
>>
>>109506833
reddit nga crashing out nah son im crine :sob:
>>
File: 1784408495869996.png (22 KB, 574x258)
22 KB PNG
i really need a new prompt bot for h3. grok is uncensored and has given me some great prompts, but it constantly baits for payment. what are you guys using?
>>
a new email address lmao
>>
>>109506843
Gemma4; on another syatem
>>
File: MiniMax_H3_00077.mp4 (3.78 MB, 816x1440)
3.78 MB
3.78 MB MP4
>>109506856
this
>>
>>109506814
Sir, this is /ldg/ we dont have lines here, just spaghetti.
>>
>>109506843
LM studio with gemma-4-26b-a4b-it-ultra-uncensored-heretic, it needs a bit of refining and extra instruction apart from throwing the two prompt structure guides at it but it's solid AND it can be put into ram if you want to run it concurrently with H3 and have enough ram to do so.
>>
>>109506776
u need to shorten the video length or add more actions in a long video. sometimes the gen become slow-mo because there are not much action going on.
>>
>>109506864
>>109506856
what vram are we looking at? i've used my old server pc with a 1060 for prompting previously, but that was just normal descriptive text. i doubt it would be able to run a model smart enough to format this stuff properly.
>>
>>109506820
>>109506839
>ahh
>nga
pot calling the kettle
>>
File: 1774539624307049.png (83 KB, 1191x583)
83 KB PNG
do you put the optimization stack into both the guider and scheduler, or just the guider? i've seen workflows with both.
>>
>>109506915
both
>>
>>109506884
First make a tool to help with formatting the prompt with Claude or Chatgpt. You can even make it a standalone html page if that uses the browser local storage. Then you just need a local llm to handle the filling in the various sections of a prompt and not generating one wholesale. I'm not sure how well Gemma 4 will do generating a complete prompt, especially at low quants necessary for a potato pc.
>>
>>109506884
for minimax?
12gb + 24gb for vae's and main model
6gb + 32gb for clip/text encoder

I have to restart comfy to free up memory for gemma 4 12b qat, but it seems to work alright at a temp of 0.7/0.8. Sit's in 12gb vram with 85k token memory and the right settings in lmstudio
>>
File: 146787697723408.png (831 KB, 544x960)
831 KB PNG
https://files.catbox.moe/b6e7t0.mp4
>>
4060ti 16gb 64gb sage+spectrum
r2v
0.8MP 14 seconds gen time: 40 minutes
vae decode time: 10 minutes
total time: 50 min

t2v
0.8MP 9 seconds gen time: 14 minutes
vae decode time: 3 minutes
total: 17 min
these vae decode times are making me angry.
>>
what even is non_diegetic_music
>>
>>109507069
it's music that's not diegetic.
>>
>>109507069
search on reddit cuh
>>
Nvidia drivers suck ass on Windows and Linux today
>>
>>109507038
>vae decode time: 10 minutes
hmm, that seems crazy

are u using --vram-headroom 1? try with that

i also assume ur using sage 2.2, cu130 pythorch 2.10+, int8cr
>>
>>109507069
>diegetic
>(narratology) Of or relating to diegesis; existing within a fictional universe (rather than as background), and able to be perceived by the characters.
>>
>>109507069
diegetic = part of the world, e.g. when in a movie or game you hear jazz playing but then there's an actual jazz band there on stage and the characters in the world can hear the music, it's diegetic. whereas non-diegetic means the music is just playing over the video. you remember all that buzz a few years ago when dead space came out with its "diegetic ui elements"? you can't have forgotten already, that was only 2008
>>
>>109506843

Get a cheap 16gb card like 5060/70 ti and let Gemma 12b run on it for captioning purposes.
Makes life a lot easier having a dual GPU system for this AI stuff.
>>
>>109507038
There was a vae released the other day that cut times, idk if it was comfy or KJ that released it, it didn't work for me, maybe it's been updated.
>>
>>109507085
>are u using --vram-headroom 1
no only --disable-pinned-memory
>i also assume ur using sage 2.2, cu130 pythorch 2.10+, int8cr
yes
>>
>>109507083
Speaking of Nvidia driver, do i use the Studio or Game Ready drivers for local diffusion?
>>
What's the best uncensored/nsfw realistic image generator like now?
>>
>>109507099
>a few years ago
>only 2008
Brootal uncpill
>>
>>109507121
try just removing --disable-pinned-memory
then try keeping it removed and adding vram headroom 1
then try keeping --disable-pinned-memory and adding vram headroom 1
>>
>>109507104
I also tried it, didn't work for me either.
>>
>>109507149
There's a pinned memory fix in yesterdays stable comfy release for Linux.
>>
>>109507149
why vram headroom and not reserve vram?
>>
>>109507132
Nta but 2008 feels very recent to me as well and I was only a teenager back then.
>>
>>109507167
Did you love being a teenager in 2006?

https://youtu.be/cydBbZaPBU0?si=wpdOon5WXZI2Pj15

I was 9 in 2006 kek
>>
>>109507163
from what i heard its better, from what i know it should be similar/the same, but i dont know for certain

also it seems like the problem will be fixed soon https://github.com/Comfy-Org/ComfyUI/pull/15316
>>
>>109507182
>I was 9 in 2006 kek
unc
>>
>>109507185
and https://github.com/Comfy-Org/ComfyUI/pull/15446
>>
>just add the flags bro
>no I don't know what they do lol
>>
>>109506843
>what are you guys using?
https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF

looks like it's also becoming one of the most popular uncensored models in general now. but it does work pretty well for both captioning images and inventing prompts (or combinations)
>>
File: MiniMax_H3_00154_.mp4 (1.52 MB, 672x1216)
1.52 MB
1.52 MB MP4
>>109506538
>>
>>109506104
fegelein!!! fegelein!!!
>>
Does the reference model also want timestamped instructions?
>>
>>109507200
36
>>
File: 1764328922411682.png (188 KB, 945x1052)
188 KB PNG
i've been trying to make a custom node for reference sets to easily switch between characters and concepts. however, the output seems to be the real bottle neck.

from my experience so far:
>image batch: has to be unpacked - stretches every image to the same aspect ratio before outputting, ruining them
>image list: has to be unpacked - crops every images to the same aspect ratio, ruining them
>having a ton of outputs in the node: will require a bunch of manual connecting and disconnecting to the main h3 r2v node depending on how many images are used

i could just connect all 9 reference images permanently and max the node out, but i'm worried that if i only have one reference image, it's going to get fed 8 additional black squares or whatever it decides to do with the additional lazy inputs that arent being used instead of only using the one selected image.

does anyone have any suggestions?
>>
>>109507223
what's that supposed to mean?
>>
>>109506883
Maybe? I can roll the same prompt 4 times and half of them are slowmo, and rest are fine.
https://files.catbox.moe/dfzqw0.mp4
>>
File: 1777252143082712.png (156 KB, 447x447)
156 KB PNG
>>109507236
ok Jensen, you won, time to empty my wallet... and my balls!
>>
>>109507038
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main
>minimax_h3_video_vae_int8_convrot.safetensors
>>
>>109507228
My problem with reference is they always create entire new scene when i mention it in the prompt.
>>
>>109507217
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md ctrl+f "000"
>subject_definitions: here you basically declare each input as <Audio 1> or <Picture 1> or explain that <Subject 1> is the fat guy in <Picture 1>
>summary: pick from the list of source usages in the prompt guide and explain in brief how they relate to each other
>retention_analysis: similar but defining how strictly they're to be preserved e.g. along the lines of "copy and paste this in" vs "use the outfit" or "play the exact audio" vs "copy the voice timbre"
>detailed_description: basically the same shots + timings as you're used to in i2v
>overall_soundscape: same
>non_diegetic_music: same
it's a huge pain in the ass, much moreso than i2v, so you're best off feeding the instruction md to a smart model (point cc or codex at it, for example) to get your base prompt set up each time you vary the input connections to the node or want to change the way you're using them, then tinker by hand from there. If it's NSFW to the degree that you think it'd be annoying to wrangle claude into helping, just build a sfw thing of the same shape and have the template for that set up. after a while you'll have a library of images with the metadata saved ready to help you where you personally remember what format each was doing. the amount of possible combinations of usages is huge so it's not solvable via others' templates, maybe someone will figure out a nicer ux that builds text chunks for you but until then this is at least not too bad
>>
>>109507258
Tried it. Only got like 5 secs speed up. Not worth it
>>
>>109507228
if the image list already only provides activated images - what if you fork the next node to take image list?
>>
>>109507261
This is why Minimax prompting is worse than LTX. they cant be simple. I think this is what filtered most people.
>>
>>109507275
All of the simple prompts work great for me on H3, maybe stop quanting the TE to 4 bits and using turbocope loras?
>>
File: 1785692335925607.png (168 KB, 1244x1008)
168 KB PNG
>>109507272
yeah i was just looking at the ref2va node and i think that might be an easy option
>>
>>109507275
>the thing that's objectively smarter and more detailed is worse because it's smarter and more detailed
vramlets aren't the ones that ruin it for everyone, it's literal retarded assfaggots that can't read instructions (or just steal from other people's workflows like a normal person)

>>109507281
the fucking nvfp4 works perfectly fine, it's why everyone uses it.
>>
>>109507275
Could be worse. At least we aren't animating bounding boxes.
>>
Considering my output folder is 403 GB of PNGs, I'm finally going to accept the devil into my life and run a recursive python script that converts everything to WEBP and deletes the originals. Quality 95.

Any last words?
>>
>>109507286
Im fucking horny. I dont want to prompt a literal words from lord of the rings novel. I just want to prompt "1girl, anal, threesome, blowjob" and be done with it
>>
>rapes the text encoder with the toyest of toy quants to save 2% of gen time
>complains the text encoder of H3 is worse than fucking LTX
brown moment
>>
can minimax not use a singular reference sheet instead of individual images?
>>
don't forget to preserve the workflows
you can also use a script that uploads everything to a catbox account or similar
>>
>>109507300
no one said that
>>
>>109507291
you can save ~40gb without losing any data or workflows if you run
oxipng.exe -o max -Z --zi 15 --alpha --preserve -r .

if you want to compress it much more while losing information then convert it to jxl, although you'll have to vibecode the program that does that while also copying over the embedded workflow
>>
File: 1780707386692669.jpg (146 KB, 407x1214)
146 KB JPG
https://civitai.red/images/139098095
>>
>>109507301
what do you mean? you can do character sheets
>>
>>109507312
so instead of using multiple ref images, I can just take a single image that contains all the references and it'll use the entire sheet without issues?
>>
File: MiniMax_H3_00181.webm (2.45 MB, 1152x640)
2.45 MB
2.45 MB WEBM
>>109506871
>>
>>109507331
well it gets scaled according to the resolution you are generating at so you can't cram the entire story line into one giant image probably, but a single image that has multiple perspectives of the character works fine
>>
Minimax dont understand what 3d animation is.
If youre reference image is anime it will force you to low anime frame like crazy.
>>
>>109507345
>well it gets scaled according to the resolution you are generating at
that's dumb as fuck what the hell why would references need to be the same res
>>
>>109507352
Based
>>
does the text 2 video model know more celebrities than the reference 2 video model?
>>
>>109507352
Based complaint I meant
>>
>>109507353
it scales automatically. you don't have to do it
>>
>>109507291
>>109226527
>>
>>109507376
I know, that's not what I'm sperging out for
>>
>>109507383
>90%
>>
>>109507376
yes but the problem is that this means a low res character sheet, so small details will be missed. seems better to use individual ref images in that case since each individual image is higher res
>>
Playing around with the T2V H3 model now for a bit and really struggling to get truly dark scenes. Even if I prompt stuff like "dark room, darkness, dimly lit, video shows a very dark room, chiaroscuro" etc. it always wants to add studio lights on the people in the shot. Any anon with tips here?
>>
File: MiniMax_H3_00084.webm (2.68 MB, 1440x816)
2.68 MB
2.68 MB WEBM
>>109507338
heh

>[INFO] Prompt executed in 375.94 seconds
>>
File: 1782348923179931.png (2.12 MB, 1024x1024)
2.12 MB PNG
this is a really good first frame for a video idea

but im out of ideas, brain focused on coom. fuck.
>>
>>109507387
i don't know. it probably feeds that as one of the frames to act as context for the rest of the video, but cuts it back out at the end
>>
File: i_00087_.png (1.29 MB, 768x1376)
1.29 MB PNG
>>109507207
kek
>>
>>109507408
Tarantino slobbering on dead nigger toes
>>
>>109507408
you could have a "ghost hunter" style clip where famous black entertainers like sammy davies junior, fats domino, michael jacksons bodies are kept for him, why? we dont know, but they spook him in the styles that made them famous?
>>109507474
maybe he bought their feet at an after death auction, like a bone from jesus finger or similar.
>>
i heard europoor are burning in the summer meanwhile my gpu doesnt go past 50c during 100% usage. interesting
>>
>run at half power
>wonder why its only 50c
>>
>>109507534
we know the 1050 Ti is not power hungry, anon
>>
>>109507549
kek, bodied that freak
>>
File: s.png (67 KB, 692x624)
67 KB PNG
>>109507549
europoor coping
>>
>>109507338
kek
>>
>>109507573
kek, bodied that freak
>>
>>109507549
>>109507543
AI only need VRAM. Technically if 1050TI has 32gb of VRAM, it will performs as fast as 5090 (This is why RAM price skyrocketed)
>>
>>109507573
>>109507591
>flexing a 1685 mhz gpu
definitely an indian
>>
>>109507573
>doesn't actually show which GPU he has
lmao
>>109507595
troll
>>
>>109507606
you arent flexing anything cause your gpu died in the heat, europoor
>>
>>109507613
Bro 2080 ram gets soldered to 26gb by chinks to become an AI card
>>
>>109507617
irrelevant now quit trolling
>>
File: 1769623089542495.png (52 KB, 452x552)
52 KB PNG
>>109507573
Let's see the actual numbers under H3, lil tyrone
>>
>>109507631
i gave them
>>
File: lel.jpg (99 KB, 1048x400)
99 KB JPG
>>109507573
this u?
>>
I just installed comfyui, how do I get started?
What model should I use?
>>
>continuous shot, fixed camera
>H3 still makes cuts
bruh, what's the point of a 32b TE if it doesn't listen to this basic shit??
>>
>browns wake up
>start wanking over hardware with ESL
many such cases, please fuck off or post gens and fuck off.
>>
>>109507624
3090 24gb peforms MUCH faster than 5070ti 16gb. Cmon now. ONLY VRAM matters for AI
>>
>>109507653
>brown is so low iq he thinks him saying something proves it
Thanks for exposing yourself.
>>
File: image.png (104 KB, 1107x992)
104 KB PNG
why has no one done this yet? @china
>>
>>109507677
>>109507677
>>109507677
>>109507677
>>
>>109507662
>t. sub 24+64 v/ramlet
>>
>>109507670
because China isn't composed of brainlets who listen to AI hallucinations
>>
>>109507670
@china @xi @police @elonmusk
>>
>>109507679
troll bake, ignore
>>
>>109507695
Make proper
>>
>>109507664
the 5070ti is MUCH faster than a NVIDIA P40
>>
>>109507677
>>109507677
>>109507677
>>
>>109507679
>>109507728
This is Debo thread. Avoid this and go to

>>109507737
>>109507737
>>109507737
>>109507737
>>
>>109507745
>1 image
>5 minutes late
too slow



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.