[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Thread Challenge Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109531268

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
blessed thread of frenship
>>
tfw i accidentally bumped the denoise down to .9 and only noticed it just now
>>
>>109533474
dogshit collage?
yep... that's the real /ldg/ thread for sure!
>>
I can't get my clip extension seamless....
>>
File: 000300.mp4 (428 KB, 736x736)
428 KB
428 KB MP4
>>
>>109533529
holy shit! lmao!
>>
>>109533529
wtf is wrong with you anon
>>
File: z-image_00641_.png (2.31 MB, 1432x1432)
2.31 MB PNG
>>
>some anon said to use a particular sage attention patch
>i've been using it while tweaking everything else to try and get more speed
>turning it back to auto just cut 100 seconds off a gen
ok
>>
>>109533519
ok I think it's because of my 2pass workflow. hopefully not feeding the starting frame to the second pass will fix the problem.
>>
test
>>
>>109533534
Have you ever watched Simpsons?
>>
File: 00042-3291262571.jpg (410 KB, 2880x1920)
410 KB JPG
>>
File: 1774422088093088.jpg (35 KB, 400x387)
35 KB JPG
>>109533474

>Face expression change
>Face start drifting

How to stop this bros ?
I dont have a neutral face on my character
>>
>>109533566
model? use a strong reference of just the face. or have expressions on your reference sheet.
>>
thanks for bakering
>>
>going through my old gens for i2v
>rediscover my anima kino
>rediscover my z image kino
>>
File: 1755460828348415.jpg (117 KB, 1000x915)
117 KB JPG
>>109533570
do i really need to have extremely large resolution of face reference with every face expressions on every angle ?
>>
Is it a known issue that audio references make the audio worse?
>>
I think anon lied. The difference between 20 steps and 40 steps is marginal. 40 steps takes me no time at all so I'm not coping.
>>
My dick really needs a break from this stuff
>>
>>109533590
same, it's been bad anon.
>>
>>109533575
typically no?
>>
File: 1764077990157250.mp4 (277 KB, 768x544)
277 KB
277 KB MP4
>>
File: 00045-2600317715.jpg (481 KB, 2880x1920)
481 KB JPG
such i beautiful model. three years later and this is what I'm able to generate locally on my own hardware. still can't believe people still using illustrious and pony in august 2026. I just hope for the best for the future of this model and the krea team releasing more open source models in the future. Very curious if flux 3 and the rumored minimax image gen model will be as good as krea2.
>>
>>109533590
I'm spending too much time on high effort gens to make coomer shit.

Any recommendations of low effort gens that yield maximum gooning energy?
>>
>>109533599
even this is better than cumfart
>>
>>109533577
When I was using an audio reference, any background noise audible in the reference kept making it into the gen even when I explicitly said to use the reference only for vocal timbre. It also felt harder to get any other audio cues to play. So, yes.
>>
>>109533580
steps... look.

You have really sigmas, that's pretty much it. the model makes assumptions about it, so the reality is they have to come in order, and possibly at a minimum rate. or otherwise, eg, the image becomes washed out.

anyway, think about sigmas, not steps, because this can help you understand why you might be wasting your time - if you are only marginally increasing the sigmas your gens need, you'll have only small improvements.

roughly, the sigma corresponds to the detail level. Even on more advanced models, it's true, mostly.
>>
>>109533618
That depends entirely on your kinks anon....but something I like to gen is women in simple positions (on her knees, fours, or back) and R2V that shit with her saying something sexy, gets me off everytime.
>>
File: 1776878284078862.mp4 (693 KB, 512x800)
693 KB
693 KB MP4
>>
>>109533519
Did you turn off spectrum?
>>
Is there any way to run Minimax H3 on a 12GB card? OR is there any way to split it between two 12GB cards? Sorry if this has been asked I've been out of the loop for about a month.
>>
>>109533653
I don't use any cope nodes.
>>
>>109533667
Adding to this I have 1 AMD and 1 Nvidia card
>>
>>109533667
run GGUF models?
>>
>>109533667
Yes, someone got it working on 6gb that I saw
>>
File: more_yarr.mp4 (2.41 MB, 1056x608)
2.41 MB
2.41 MB MP4
>>109533590 >>109533593
computer says no

>>109533667
sure. obviously not with every possible comfyui workflow and settings / uses the lower memory profiles on wan2gp, but it does work.
>>
File: image.png (180 KB, 848x477)
180 KB PNG
>>109533667
>>
>>109533617
>Is it just me or is the ref2va model noticably more prone to artifacts/issues for the same settings (sampler/steps)?
1.2MP at 8 seconds looks way better on the i2va model compared to the ref2va.
>>109533642
>it is https://old.reddit.com/r/StableDiffusion/comments/1vl3ed0/having_bad_ref2va_quality_compared_to_fl2va_try/
also needs own turbo lora

Even for IMG2Vid models ?
>>
I am your favorite genner's favorite genner.
>>
File: H3_noaudio_00002_.mp4 (1.19 MB, 736x1280)
1.19 MB
1.19 MB MP4
>>
>>109533702
you war kino anon?
>>
>>109533667
You're going to spend more time making your PC work at outputting shitty videos 30 minutes for 8 seconds clips than actually working and get a better GPU. All good things take money... cars, girls, guns, house.
>>
>>109533708
What would the best gpu be for pure genning? I don't plan on gaming
>>
>>109533694
anon, I have pictures of my crush, I have her voice, I can put her in any outfit I want. This is unsafe for my large dick.
>>
>>109533711
The zenith right now is RTX Pro 6000. RTX5090 second place. Purely because of VRAM size.
>>
>>109533726
I should have asked for budget card.
Couldn't justify getting them sadly
>>
>>109533711
probably something like a set of B300? maybe it eventually gets difficult to load something sensible after you have a few TB of VRAM and LLM + image model + video model at high resolution are actually covered.

i think no one here has the best still useful hardware.

>>109533712
imagine how much worse it is if you like many anime girls
>>
>>109533712
would it be safe for my small dick?
>>
>>109533711
1st anything that has more than 64gb of VRAM
2nd anything that has 32gb of VRAM
3rd anything that has 24gb of VRAM
4th anything that has 16gb of VRAM

Technically if GTX 750ti has 128gb of VRAM it gonna mogs everything including RTX PRO 6000

VRAM is king, anything else dont matter
>>
>>109533734
>B300
More expensive than the 6000 Pro where I look
>>109533748
I hope there will be more hardware scenes to get more ram on your current CPUs Got a 16GB Card right now only
>>
Images are still welcome here as long as they are good btw
>>
File: 1775937057453740.mp4 (1.2 MB, 512x800)
1.2 MB
1.2 MB MP4
>>
File: 17796374457.jpg (5 KB, 132x137)
5 KB JPG
>>109533694
>>
>>109533702
Bruce Genner
>>
File: Begone vramlets.png (97 KB, 1139x484)
97 KB PNG
>>109533711
the card I have. Good news, you can game on it too!
>>
>>109533763
>More expensive than the 6000 Pro where I look
it's ball park ~4-5x better in key specs with only ~2-2.5x the power consumption and in the end hardware is now nearly entirely priced for AI

of course since you mentioned "budget" later that's obviously not it.
>>
>>109533743
Asian dicks matter too anon :3
>>
>>109533787
>Good news, you can game on it too!
how many fps do you get in tetris?
>>
>>109533704
prompt, catbox, social security number, meaning of life?
>>
File: H3_noaudio_00011_.mp4 (1.92 MB, 1152x896)
1.92 MB
1.92 MB MP4
>>
Can you give a video reference of someone doing something but have h3 incorporate from a different angle?
>>
File: H3_noaudio_00012_.mp4 (1.85 MB, 1152x896)
1.85 MB
1.85 MB MP4
>>
https://h.uguu.se/LoCqFrzG.mp4
>>
>>109533609
>>109533564
Do namine
>>
>>109533857
https://huggingface.co/matlod/minimax-h3-turnaround
>>
>>>/gif/31030304
Armor break!
>>
File: continue.png (909 KB, 2677x1641)
909 KB PNG
The target video is direct continuation of <Video 1>. <Video 1> containing initial motion plays out and seamlessly continues.
>>
File: Misty-cut-NS-00015_.webm (1.84 MB, 672x960)
1.84 MB
1.84 MB WEBM
thoughts on the latest pokegen?

https://h.uguu.se/AqaEaAnP.webm
>>
>>109533922
>>>/gif/vdg
the greatest thread on this site
>>
>>109533955
So good bro, please use a tripcode so I can highlight all of your future posts man
>>
>>109533962
this so much this
>>
>>109533962
>>109533973
no, debo
>>
>>109533938
now this is my kind of autism!
The issue I run into tho is the first frame will be perfect, but then the later generated frames will have a slight morph "back into place". but maybe feeding it more frames is the way to stop that from happening?
>>
>>109533962
kek you almost got him
>>
>>109533955
We can see the transition, this is what I'm trying to avoid. but It's never as bad as this tho.
>>
>>109533975
Can we make a deal?
I'll go by debo if you go by pebo?
>>
>>109533986
Anon, you are very clearly retarded.
>>
>>109534014
>Spam same gens over and over
>Gets mad when called out
>Someone actually provides a fair view on a gen and he spergs out more
ok
>>
>>109534018
No, I'm calling you retarded because you clearly are retarded.
>>
>>109534023
You called some one else that not me.
>>
This is why I put black guys in his gens.
>>
>>109534025
Anon, you're retarded.
>>
File: ed5.png (131 KB, 680x1112)
131 KB PNG
>This is why I put black guys in his gens.
>>
File: 1775918123017393.png (15 KB, 128x128)
15 KB PNG
ref_image_size "Max" increase the face accuracy using image face reference..... At the cost of generation getting slower by 25%

I need 4 steps turbo lora to be useful so bad. 8 is too much
>>
>>109533977
As you can see I’m feeding it entire last second, it works almost perfectly. Sometimes a slight morphing transition happens but it could be because of the audio I have. gotta test more
>>
You realize that being a known 4chan schizo is actually a really bad thing right?
>>
>>109534018
wanschizo will always drop whatever he's doing to have a 4chan catfight, he just loves 'em
>>
>>109534040
It's the most prestigious role on me CV mate.
>>109534042
Flags or IDs would be a Godsend. Even better with just an AI Board at this point that contains either of those
>>
>>109534039
I'll try feeding more frames then. When I initially tried with 5 the morphing was worse than with 1.

I also didn't use your prompt. But for sure injecting frames in the latent directly works a lot better than just feeding a reference or first_frame
>>
i dont like pokemon spammer cause i asked him to do my favorite pokegirl but he ignored me so yeah he is a fart face
>>
File: _00167_.mp4 (890 KB, 736x544)
890 KB
890 KB MP4
>>
>>109534040
Clearly you know from your own experience.
>>
>>109534026
Based
>>
File: ldg.png (19 KB, 1893x79)
19 KB PNG
>>
File: 70338.png (32 KB, 855x196)
32 KB PNG
>>109534042
Anon, how does it feel to be completely mindbroken?
>>
>>109534050
> I also didn't use your prompt.
You basically need to describe the Video you’re injecting and starting from to gaslight the model into thinking it came up with the latent on it’s own
>>
>>109533891
Looks cool but not sure it's what I want. I want to apply the action to a different character from another perspective
>>
>>109534087
>"I made this"
>>
>>109534098
hey i have a question: why did you use two quotation marks next to each other? regards, anon
>>
File: MiniMax_H3_00661__1.webm (1.95 MB, 896x704)
1.95 MB
1.95 MB WEBM
>>
It's save to update comfy to 0.32? Didn't some anons had problems yesterday?
>>
File: image7.png (634 KB, 1939x1296)
634 KB PNG
>>109534120
why are you doing this?
>>
File: SwSh_Lass.webm (1.66 MB, 896x576)
1.66 MB
1.66 MB WEBM
>>
ltx2.5 still jeeted, you get indian woman singing music sometimes
bad physics
ear rape
censored on violence and sex
does not know famous people, could not do obama for instance, knows trump however
only thing going for it is it is pretty fast
>>
Are you guys looking forward to wan 3?
>>
There is a soft version of Minimax that only uses 6GB of GPU

https://civitai.com/models/2835250/minimax-h3-fastest-workflow-or-6gb-vram-16gb-ram

thoughts?
>>
>>109534161
seems indian desu
>>
>>109534135
so its got nothing going for it
>>
How do I disable the api nodes in comfy? I added --disable-api-nodes to my launch flags, but it still shows me api workflows.
>>
or south asian. still, the dregs of humanity. who cares what they think?
>>
>>109534161
I used it thanks to its R2V nodes on I2V models and disable most of cope nodes
>>
>>109533955
[spoiler]Are you ok with Gou as a femboy?[/spoiler]
>>
>>109534232
HAHAHAHAHA LOOK AT THE NEWFAG AND LAUGH
>>
>>109533955
on the right track anon
>>
>>109533474
>it’s called comfy but you a degree in computer science and subscription to an Ilm to use

Seriously why is it so complicated
>>
>>109534266
Just treat it like a puzzle game. Once you have your save file, aka workflow, it can be very convenient.
>>
>>109534266
>comfy
don't believe his lies
>>
File: x.mp4 (2.11 MB, 768x1376)
2.11 MB
2.11 MB MP4
>>109534266
the things it abstracts over are worse. you'll manage without the degree in computer science.
>>
>degree in computer science
You just need to not give up. Some pick it up quicker, others slower.
>>
I think this is the most snobbiest general on this site , no one helps one another just judges , I was called an Indian the other day because I made a spelling mistake !
>>
>>109534317
that would never happen on any other board, take my word
>>
>>109534266
it's not. just watch some beginner jewtube tutorials and starts with something more simple like image models.
>>
>>109534297
comfyUI is dog shit and confusing as hell i've tried multiple times to understand how but everytime output is dog shit or itll requirea bout 69999 fucking plugins or whatever the fuck.
>>
>>109534337
You're welcome to sit on the cuck chair while the chads have their turn with latest model releases on comfyui.
>>
>>109534349
I just use wan2gp+swarmUI. plus why would i really need to get up to date right away models anyway. It doesn't really improve anything at least as of lately
>>
So is a ref2va workflow the best way to get good results? The default comfy workflow for i2va is kinda inconsistent imo, especially for porn.
>>
>>109534337
the problem is you, BUT no one says you're not allowed to use for example wan2gp if you prefer that. it's not like it won't make images/videos.
>>
i think the glowniggers are messing with me. ltx keeps generating naked girls when i am trying to make the music video prompt
>>
>>109534337
I agree but there's really nothing better right now and it's hard for new projects to gain traction because the community decided on comfy. Kinda like sillytavern.
>>
>>109534357
Yeah. Once you figured you character replacement, the whole internet nsfw catalog is open to you.
>>
>>109534357
both can give good results but obviously ref2va is the more natural one for working with references

the upstream separation also represented in comfyui makes sense
>>
>>109534368
the problem isn't me comfyUI is fucking confusing as hell. I've also tried to use other peoples workflows and I honestly still don't get what the fuck is going on because of how much shit you need to set up to even get a proper output. Outdated plugins incompatible plugins plugin browsers plugin monitoring just loads of fucking bullshit.
>example wan2gp if you prefer that
thats what I've been using. Just don't fucking tell me comfyUI isn't extremely annoying to use. Its a load of bullshit!
>>
>>109534370
Dont ever do T2V with LTX
>>
>>109533667
you need better card
can try multigpu or raylight nodes but it will be problematic

new card
or very low quant (bad quality)
>>
>>109534389
Comfy's compromises favor extensibility over stability, and the AI space moves so fast that this is the model that won over monolithic UIs with poor extensibility. Work ethic and associates selection also played a big role. It's not ideal, but it's better than the alternatives.
>>
anon that genned >>109533087
asking genuinely, what are some custom nodes that do image to mesh, and then ideally mesh to rigged mesh? i figure i should just fuckin commit to it and swap it out when i have something made by a real person
>>
I was busy for a few days, is there a consensus on turbo loras yet? Are they usable?
>>
>>109534438
For ref model no. For the i2v model yes.
>>
I love competition
How is the new LTX?
>>
>>109534449
same as the old LTX
>>
>>109533667
You can try Select Clip Device, Select Model Device, and MultiGPU CFG Split.
>>
LTX Chads.
We won.
>>
>>109534407
t2v is the only way to get kino
>>
>>109534450
unfortunate
>>
Will flux really mog h3?
>>
>>109534479
it already does
>>
>>109534479
in some ways probably, but other ways no chance
>>
>>
Spectrum is causing artifacting
>>
>>109534449
kino as heck
>>
So can we admit that LTX has just killed Minimax? Flux 3 will piss on its ashes.
KINO
>>
>>109534479
same opinion as the other anon: seems unlikely but maybe it does <some stuff> better

idk, pole dancing gender swapped jesus christ.
>>
>>109534545
proof?
>>
File: MiniMax_H3_00112_e.webm (690 KB, 960x544)
690 KB
690 KB WEBM
>>
File: 85547.webm (3.99 MB, 420x224)
3.99 MB
3.99 MB WEBM
THANK YOU ZEEV
>>
File: 1758041506906389.gif (227 KB, 352x240)
227 KB GIF
Lower Shift video values makes motion faster on turbo 4step loras.... umu umu.....
>>
>>109534161
>take default workflow
>add cope nodes
>set generation license fee
>spam /ldg/ to advertise
>wait for retards to fall for it
>>
I still cant solve the door creaking sounds on porn vids
>>
>comfy's content security policy doesn't include media-src so disabling api nodes also bricks uploading videos
Thanks. To fix edit server.py.
>>
File: MiniMax_H3_00107_.mp4 (1.95 MB, 864x480)
1.95 MB
1.95 MB MP4
>>
File: MiniMax_H3_00119_e.webm (2.48 MB, 544x960)
2.48 MB
2.48 MB WEBM
>>109534706
Stop and try to gen cool stuff instead then.
>>
how do I keep the camera motionless in r2v?
>>
>>109534827
>still shot
>>
>>109534733
very nice shroom

>>109534827
fixed camera or still shot both tend to work (in detailed description or summary idk which place is correct tho)
>>
>>109534706
turn off you§re sound
>>
So when do you think we'll start getting some good loras for h3?
>>
Does H3 not understand what from above and top view means? Is there some term for it I'm not aware of?
>>
File: 74346.webm (2.65 MB, 420x224)
2.65 MB
2.65 MB WEBM
>>
Guys, LTX is our ally. I know H3 is much better, but please throw a few gens LTX 2.5's way here and there, just out of the goodness of your hearts
https://www.reddit.com/r/StableDiffusion/comments/1vma6pc/ltx_is_our_ally_its_two_cakes_dammit/
>>
Does nvidia driver update help with gen time if the installed driver is from march? I have cuda 13.
>>
>>109534986
>the Israeli mode is our greatest ally
lmao. In a China vs Israel competition I side with China, simple as
>>
>>109535013
this. china never called me goyim
>>
>>109533955
>File(s) expire after 3 hours.
ffs
>>
anyone know where ican find the ModelAttentionBackend node to replace sage with comfykitchen?
>>
>>109535025
nothing to worry about, the schizo will reupload it for you in the next thread
>>
>>109535023
based gweilo
>>
>>109534986
I would but for whatever reason trying to use that model or any derivatives on comfy causes an instant bluescreen, and I'm not ready to commit to checking if it's a drop-in replacement on w2gp
>>
>>109534977
China rulez
They're flooding the U.S. with open weights models and cheap fentanyl.
What does the U.S. do for China?
>>
>>109535049
buys their tat in bulk
>>
it's pretty fun. the wan2gp dev hasn't added it yet but check back once in a while to see if he updated it
>>
>>109535068
>>109535045
>>
File: 1783046490443157.jpg (14 KB, 470x164)
14 KB JPG
What happens If i enable comfy kitchen while --SageAttention is on ? Does the Comfy Kitchen still working ??????????????????????
>>
>>109535074
don tdo it creates mustarg gas
>>
ref model can clone gilbert gottfried, if it can do that it can do anything.

https://files.catbox.moe/32pvfv.mp4
>>
>>109535074
NO ANON DON'T
>>
>>109533874
starting the image hunting process for both her and aqua. larxene might be a very harder than i thought and i don't like her kh3 look.
>>
File: 00012-1605986821.png (596 KB, 640x512)
596 KB PNG
>>
>>109535074
one anon tried it earlier, we never heard from him again
>>
Is there an alternative to ComfyUI Lora Manager that doesn't try to call home?
>>
so what's the current cope node meta? i used comfy kitchen, but it made my gens 2x slower
>>
>>109535034
after u updated to latest comfy kitchen
>>
File: fw5x961spsih1.png (371 KB, 1440x3120)
371 KB PNG
Our response, H3bros?
>>
>find a random hentai sketch
>chuck it into h3 with a random model
still gets me that this works
>>
>>109535175
juriisdiction
>>
File: MiniMax_H3_00126_e.webm (2.76 MB, 544x960)
2.76 MB
2.76 MB WEBM
>>109534779
restructured my prompt a tad
>>
>>109535184
>random hentai sketch
as in just a hentai image?
>a random model
as in a woman's picture?
that's pretty cool, what positions are you having good success with?
>>
I need to talk to someone about my thyroid
>>
>>109535175
i said it like eight or nine times already but lmao this is so pathetic
>>
>>109535175
>>109535193
It's antisemitic to have three "i's" in a word because it's too close to trinity symbolism, so they add an extra "i" to sanitize it
>>
>>109535197
I wonder why people post stuff like that.
I get gooning - it satisfies an evolutionary sexual instinct.
But you only share something like that in the hope that someone will pay you attention.
It's so sad on a human level.
>>
>>109534986
>I checked reddit and have an opinion now
lmao
>>
>>109535175
>ermh, ackshually, minimax isn't available in the majority of the world, only ours
>NO DON'T GO DOWNLOAD IT FREELY FROM HUGGINGFACE! STOP!
>>
>>109535175
https://files.catbox.moe/03pjmw.mp4
>>
>>109535200
Kinda wild stuff. One with tentacles, one with a sex machine, it requires a little finesse in the prompt but it really does just werk, I can get it to pick up the pose as well as the non-human elements in the image
>>
>>109534983
that looks shockingly good quality for 400*200
>>
>>109535269
Do you always get generic squeaky voices?
>>
can I use the fl2v turbo lora for the ref model as well?
>>
File: 00015-2167880119.png (544 KB, 640x512)
544 KB PNG
>>
>>109535278
I get some okay variance. I use an LLM prompt enhancer based on the guides that pretty consistently writes out something that creates emotion, so if you're doing it manually just add some descriptors of the emotion behind the voice. You can also use a voice reference but I found that to be slightly more flat.
>>
>>109535289
What I do is just plug the fl2v model into the ref to video conditioning node and load the lora normally. It does in fact work.
>>
gonna download ltx 2.5 just to try a few complex gens that h3 managed to get, laugh at it, and delete it
>>
>>109535272
that's one nice thing about ltx. the blurry vae decoder actually makes low resolutions look better compared to h3 being sharp but dotty
>>
>>109535175
>LTX: Israeli
>MiniMax H3: Chinese
>>
>>109535306
what? there is no way that just works
>>
>>109535292
SD 1.5? lmao
>>
>>109533564
>>109533609
thank you kairichad. hope you end up doing other kingdom hearts girls, aqua is one i wanted to try out when krea 2 first came out because no model can handle the complexity of her outfit's little heart emblem thing.
>>
>>109533553
do you mean the comfy kitchen one?
>>
>>109533609
we want to see you do some h3 gens sasori
>>
>>109535175
I didn't know I have 4 GPU's and 115GB of VRAM
>>
okay boys it's been half a day. bottom line on the new jew model?
>>
>>109535314
Here's a complex one
https://www.reddit.com/r/StableDiffusion/s/4UxRLeBmUA
>>
I saw one LTX video gen, it was like Will Smith spaghetti in early days.
>>
>>109535353
People should stop using copyrighted characters when comparing these two models, it serves no real porpouse but making LTX 2.5 struggle even more with something it's not allowed to use.

Just do something with generic non-copyrighted characters, H3 completely obliterates 2.5 anyways so why not just do motion those comparisions? It's not like the result is gonna change, 2.5 is still ass at motion.

I know the LTX devs didn't play fair for their comparision either (such as including queue times as generation speed time for closed models which is so damn moronic and dishonest) but we should be better than that and be more fair for comparisions.

EDIT: in case it wasn't clear i'm not defending LTX i've been shitting on them since 2.5 dropped still terrible and those bogus charts they made, my point was show a fair and clean non-copyrighted comparision focused on just the motion capabilities to show how much H3 destroys 2.5 and how bad 2.5 is at that, so people have no excuses like "oh but it doesn't know naruto obviously it's going to look bad"
>>
>>109535333
Try it. Prompt and connect images exactly as if you were doing normal ref2v. It might not behave correctly with video and audio, so it'll depend on your use case, but with two images it can absolutely get there.
>>
>>109535351
you can tell it's good since the fud spam is happening here
>>
Is H3 just bad at 2D animation or is it a skill issue on my part? If it's the former can that be fixed with a lora?

>>109532049
>someone used my gen
>>
Kairianon please upload your gens to a catbox gallery or something
>>
ummm bros it's been like 2 weeks where are the good nsfw loras for h3??
>>
>>109535380
>It's not fair to use a feature that makes one model superior!!!!
>>
>>109535480
I'm starting to think we won't get any...
>>
>>109533575
Make sure the character references are within the output resolution, at extremely high resolution it's just going to get crushed on down scale.
>>
>>109535480
They're not good but there's a few on civit. It's only gonna get better also it's been like 10 days
>>
File: 1772737848754644.png (36 KB, 1151x170)
36 KB PNG
Not familiar with the training process. Does this mean it's almost done and they'll possibly release it by next month?
>>
>>109535480
If you need loras for H3 you're braindead retarded and should consider applying for gibs.
>>
>>109535547
No, H3 is great but it's pretty bad at porn and 2D.
>>
File: 00000-614496023.png (2.12 MB, 1536x1920)
2.12 MB PNG
>>109535442
i'm not the kairi anon but do you need the tags he shared last thread?
>>
>>109535547
yeah bro it's so much easier to find a reference porn video and cut 10s out of it or just put up with mangled looking genitals crashing together when you just text prompt lol
>>
>>109535563
Thank you for confirming you're beyond saving levels of retarded.

>>109535575
Another down syndrome retard to the list.
>>
>>109535571
No I just want to save all of his gens and not miss out of future ones, but sure the tags would be nice too.
>>
>>109535579
post one, just ONE single nsfw h3 gen you made
>>
so what were those DOA models comfy was talking about?
zedit maybe?
>>
>>109535579
Thanks for confirming you're a nigger.
>>
File: DARKNESS for breakfast.png (3.48 MB, 1536x1920)
3.48 MB PNG
>>109535582
sure, bro tagged the fuck out of it.

*DEEP BREATH*

kairi, Kairi (Kingdom Hearts), Kingdom Hearts 1, 1girl, solo, female, girl, youthful, young teenage girl, 14 years old, fair skin, light skin, blue eyes, short hair, red hair, hair bangs, layered hair, hair between eyes, red eyebrows, black choker, necklace, teardrop pendant, collarbone, sleeveless, white cropped tank top, white cropped tank top purple trim and purple straps, black undershirt, yellow sweatband on left forearm, purple armband on left upper arm, yellow and black bracelets on right arm wrist, navel, midriff, dark purple belt, silver belt buckle, light purple short skirt, light purple mimi-skirt, scalloped hem on skirt, flower petal cutout pattern on the scalloped hem of skirt, side slit on skirt, side slit on both sides on skirt, purple bike shorts, purple bike shorts worn underneath her purple skirt, oversized shoes, chunky shoes, white shoes, yellow laces on shoes, blue and purple soles on shoes, slender build, slender legs,
3d, 3d render, cinematic render, high fidelity, cgi, cutscene, videogame cinematic cutscene, cinematic scene, cinematic cgi graphics,
>>
>>109535640
Thanks anon. Also SEXO. I should really find some good clips of Kairi's voice (Hayden, of course)...
>>
>>109535666
you're welcome secret satan.
also had no idea Hayden Panettiere played a character in Disney's Fillmore, damn i should've guessed the girl she played sounded identical to Kh1 kairi.. what a cute voice.
>>
File: 743462.webm (3.82 MB, 420x224)
3.82 MB
3.82 MB WEBM
>>
>>109535197
Looks cool.
>>
File: 1762737793890814.png (606 KB, 1491x1055)
606 KB PNG
https://reddit.com/r/StableDiffusion/comments/1vecz21/minimax_h3_can_use_a_storyboard_as_a_visual/
>can use a single image as a storyboard and H3 understands everything
How is this model so fucking good?
>>
File: 00021-1850056409.png (3.14 MB, 1536x1920)
3.14 MB PNG
>>109535710
these are really cool, they're like cheesy 2000s music videos.
>>
>>109533575
Are you genning at low resolution anon?
>>
File: 1761554226863207.mp4 (3.82 MB, 2048x1130)
3.82 MB
3.82 MB MP4
>>>/wsg/6212817
>>
>>109535773
So apifags are poorfag brown pajeets and cucks? I don't get it.
>>
>update comfy
>update nodes
>same workflow is now 2x slower
what do?
>>
>>109535786
>update comfy
lol
>>
File: 4862.png (13 KB, 680x907)
13 KB PNG
>>109535786
>>
Does anyone have a simple workflow for long generations like minutes worth, I tried nicodemon80's one but its a mess to me, I tried getting a clanker to make one but it's beyond it. Just want clip1->video1 clip2->video output appended to video1 clip3->video output appended to video 2 and so on.
Little help?

>>109535500
A very anti-racist stance applied to technology, +1 upvotes to that particular redditor lol
>>
File: 644678.webm (3.43 MB, 420x224)
3.43 MB
3.43 MB WEBM
>>109535737
yeah i am trying to find the right settings that make it go crazy with effects
>>
>tfw you realize you forgot to change the aspect ratio 70% into a gen
>>
ref2va is so fun. just combine a few clips and you have a video.

Microsoft explaining their code issues:

https://files.catbox.moe/1o4w1h.mp4
>>
u need another program for local llm to generate prompt for h3? u can't use comfyui to do it?
>>
>1k buzz EA for this actual slop
what the fuck is this guy smoking lately
>>
File: rice dong.webm (1.53 MB, 1280x896)
1.53 MB
1.53 MB WEBM
>>109535904
It is fun as fuck. You can just drop a 4coma and make it animate panels from reference.
>>
>>109535914
all I did was add a voice reference and an image, and then do this:

Use <Picture 1> for the physical identity of Superintendent Chalmers, with the voice of <Audio 1>.

[Scene Description]
The Simpsons style, 2D animation. Cel-shaded animation, clean lines, vibrant flat colors.

and then poof, character and voice work great.
>>
How do you get the ref model to take the artstyle from one of the ref images and apply it to the whole video? When I try it ends up taking the artstyle of one of the other input images instead.
>>
>>109535913
that's because anima users don't know how to train loras. any lora is a luxury to them
>>
>>109535737
>>
File: 1775792243897767.mp4 (395 KB, 640x832)
395 KB
395 KB MP4
>>
>>109535928
I haven't found a way to do this super consistently, reference transfer with multiple image seems to have some trial and error. If you're improving your prompts on a toaster you could yell at it to make the role of the ref image as a style guide more clearly defined, or you could do it yourself by restating the role of <Picture X> in as many places as possible. No guarantees it'll help but it's what I've been experimenting with when I've had issues like the one you described.
>>
>>109535946
>>109535961
holy cum
>>
Can someone please slutify her?
>>
If LTX were smart, they'd focus purely on specializing in video upscaling. Let China handle the video models, because they're clearly better at it than LTX and it's not even close. Regardless of who makes the best video models, upscaling will always be useful
>>
File: locust.jpg (13 KB, 269x188)
13 KB JPG
>>109536012
>>
>>109536014
If they didn't have jewish brains focused on profit yes i think that's where they'd be useful. But it'll probably never happen.
>>
>>109535928
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
>>109536014
Chinese have their own upscaler.
>>
>>109536012
Be brown elsewhere
>>
>>109536027
Their models would probably not be open if they focused on a single use case. They would be another Topaz (who focuses solely on image/video upscaling and they are SOTA in that regard)
>>
Is anyone even using pure text to video?
>>
File: 00057-2182618659.jpg (633 KB, 2880x1920)
633 KB JPG
>>109535442
finally uploaded it but had to omit certain "young" tags to not get the lora shadow banned like the lunafreya one.
https://civitai.red/models/2853096/kairi-kingdom-hearts-1-2002-krea-2-lora?modelVersionId=3222105
>>109535640
glad you like the lora.
>>
>>109536022
Is there a different board i should ask?
>>
>>109536054
well yeah, real children would be unethical
>>
>>
>>109536057
Oh it's YOU the fella posting the queen's blade girl from before, and you have that blue dragon lora too. cool.
Do you mostly use krea 2 center? i keep encountering issues using it. May just have to stop using forge and go back to cumfartui..
>>
>>109536054
I generate cunnies on text2video. Quite good with innie pussy lora.
>>
>>109536072
his name is sasori, show some respect. also he's black irl
>>
>>109536054
yes
t2v is fucking powerful with this model
the possibilities are endless
>>
>>109534132
holy ass
>>
>>109536088
proof?
>>
>>109536070
if you're the same person who made the original mario video (because who else would care) , you should tell a short story over the course of 20 years by slowly adding a new shot every time a new SOTA video model comes out
>>
>>109536053
>>109536014
>>109536043
Speaking of upscalers, is this still the local sota?

https://github.com/ByteDance-Seed/SeedVR
>>
>there's no official comfy workflow for krea2 raw
>find some on civit
>the results are worse than turbo
emmm... Does anyone have a working one for raw model?
>>
>>109536101
no, RTX upscaler from Debo is better
>>
>>109536099
im actually nta but remembered the clip last night and wanted to try recreating it in h3
>>
>>109536101
Not worth the performance hit in my opinion but I'm a vramlet.
>>
>>109536029
Yes, I'm aware of the guide. I used that formatting.
>>
>>109536057
>>109535961
>>109535946
>>109535737
>>109535640
really liking the Kairi genns
please keep it up
>>
>>109536101
Cracked topaz is the sota. Good luck to find it though
>>
h3 r2v like to give people tanlines
>>
>>109536142
cracked topaz is giga slow. they don't focus much on optimizing their models because they expect you to use their api
>>
File: 00085-639946258.jpg (675 KB, 1920x2880)
675 KB JPG
>>109536072
yes I'm that anon that made the leina and bouquet lora. I indeed mainly use the krea 2 center semiraw checkpoint. Many of the other shitmix checkpoints are way too unstable and break the visual look and consistency for non-photorealistic characters. forge neo works perfect fine on my own end.
>>
>>109536142
How come nobody has reverse engineered it yet? Does it comes with weights, or is it an algorithm (or both)?
>>
>>109536088
>>109536097
Almost everything I'm doing is functionally t2v since I'm mostly just using a character reference or two. Designing the scene and handling the shots are all on the model, I haven't done a single literal i2v task since the first time I ran the model and that's all I was used to doing.
>>
>>109536129
>2.2 <Picture N>
Use a standalone <Picture N> when the reference image itself serves as a shot's first frame, keyframe, last frame, edited keyframe, or composition anchor:

><Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.
If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding <Subject N> definition.

>When an image acts as a storyboard or shot-planning reference, state which shots it maps to and what planning information it provides:

><Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.

It is this key part tho: 'If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding <Subject N> definition.'
>>
>>109535736
This is actually pretty cool, but you still need to create the storyboard for it.
>>
File: Krea2_turbo_00786_.png (1.29 MB, 1024x1024)
1.29 MB PNG
>>109536157
i like the luna one
>>
>>109536171
Write the storyboard in an abliterated LLM, generate it in Claude, hand it to minimax, profit. Actually I guess Claude would probably rap your wrists for doing no-nos
>>
>>109536157
i probably just need to update or somethin' then.
>>109536175
this bih here is literally just namine. but cuter i think.
>>
https://x.com/banodoco/status/2085089486032027790
Useful video for proompting.

>>109536171
I'm curious how detailed the storyboard needs to be. I can draw but I'm /int/ at best right now. If I can do rough sketches for a storyboard and pair it with high quality refs for the characters and scene I'll coom.
>>
simpsons on windows 11

https://files.catbox.moe/5p9tia.mp4
>>
>>109536192
>you're international at best
suka
>>
is ref2v + first frame impossible?
>>
File: 1763448632425942.png (1.88 MB, 2006x1714)
1.88 MB PNG
>>109536186
>namine. but cuter
Blasphemous
>>
>>109536201
just ask ref2v to do that. with respect.
>>
>>109536214
>>109536214
>>
>>109536200
kek
>>
what's a good test resolution to see if a gen is worth using at higher mp?
>>
>>109536224
0.3 is a good test or fast meme size
>>
>>109536224
anon has said that the model is "stupider" at lower res so if it's a complex prompt maybe not. I might try testing some .3 vs .5 today with the same seed and see if that's at all true
>>
File: migu-dj.mp4 (2.42 MB, 736x576)
2.42 MB
2.42 MB MP4
https://files.catbox.moe/u7l58a.mp4
I used the sageattention node. It cut the gen time in half to about 1:22. I made sure the input fps matched the output 24 fps. It doesn't look like you can push r2r past 15 seconds, it starts to fuck up right at that point, but otherwise it came out very well.
>>
>>109536175
make her clothes 半透明
>>
>>109536070
>Mario noir
Already been done.
>>
>>109536012
You can't slutify a slut anon
>>
>>109536054
Pretty sure that was a general specifically for that on /b/, but I think jannies banned it because we are "le wholesome reddit website xD" now.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.