[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 9f5fa7.gif (768 KB, 240x240)
768 KB GIF
Video Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109444052

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>LTX-2.3
https://huggingface.co/collections/Lightricks/ltx-23

>Wan
https://github.com/Wan-Video/Wan2.2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
It's dangerous when models like this get released. I always have to put a small cut on my foreskin to keep my self from gooning for 18 hours a day
>>
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
speaking of will spaghetti...

https://files.catbox.moe/7czozm.mp4
>>
>>109445731
>trolling outside of /b/

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 8kl8sy.gif (935 KB, 201x202)
935 KB GIF
>>
File: 87363616.mp4 (3.73 MB, 834x480)
3.73 MB
3.73 MB MP4
>>109445731
>>
>>109445748
why is it crunchy?
>>
>>109445760
amateur mistake, not cooking the pasta enough
>>
File: 1364457243237357346.png (19 KB, 331x307)
19 KB PNG
which one for 3090?
>>
>>109445789
where is sage attention 2? that's the highest the 30 series can go
>>
>>>/wsg/6207145
Stealing seedance prompts from xitter because I'm uncreative.
20 mins on 16gb/64gb. Lightning loras can't come soon enough
>>
>>109445793
I have sageattention-1.0.6, is this a version issue?
>>
>>109445547
>windmill poles alternated between having a white triangle and not having one
>zoom in on the rocket launcher
>two white triangles in a row come out
Literally unusable.
>>
>>109445760
It's called el dente you racist
>>
>>109445748
but then...drama ensues!

https://files.catbox.moe/ezxagv.mp4
>>
https://litter.catbox.moe/6wcv5kv68fqhe9nh.mp4
https://litter.catbox.moe/pltgjlr5b45jet3z.mp4
h3 maintains the body structure from the reference image very well, impressive.
>>
File: stretchy skull.jpg (101 KB, 1014x1024)
101 KB JPG
H3 is the first video model that does my niche obscure uhhh interest out of the box, it's so fucking over for me
>>
>>109445826
Which?
>>
>>109445822
Nice
>>
File: 726752.mp4 (3.38 MB, 512x800)
3.38 MB
3.38 MB MP4
It does a pretty solid loop with first and last frame.
>>
audio on /g/ now
AUDIO ON /g/ NOW
>>
160 seconds (10s/0.2mp), i2v, the spaghetti man gets revenge

https://files.catbox.moe/8610a2.mp4
>>
>>109445856
yaoi now
>>
Day 1 with zero optimizations and I'm already very happy. Localbros feastin good
>>
The catbox servers are gonna be working extra hard the coming days
>>
Hatsune Miku is watching will smith eating spaghetti on a TV screen. the camera zooms out to show an anime style Hatsune Miku looking at the TV saying "pasta, dayo?"

65s, 0.2mp/5s

https://files.catbox.moe/x1sk8o.mp4
>>
Thank you for your service LTX.
Its Minimax Time.

5080. 266s

files.catbox.moe/k04335.mp4
>>
do you think downscale->upscale workflows like in LTX will be possible? and turbo when?
>>
>>109445907
someone mentioned the minimax team are working on releasing an upscaler for this
>>
>>
i'm impressed that very small resolutions are not completely fucked like in other models, actually useful for prototyping
>>
is a thread on wsg justified yet
>>
File: x.mp4 (299 KB, 608x352)
299 KB
299 KB MP4
>>109445936
you do you
>>
>>109445936
there's already an ai thread on /wsg/
>>
>>109445936
there's already an AI video thread for wsg, you can upload your minimax kinos here
>>>/wsg/6191738
>>
>>>/wsg/6207151
>>
FUCKING WAGIE LIFE I HAVE TO GO TO WORK I WANT TO TRY IT OUT TOO BWEEEEEEEEH
>>
im trying to make soft smut but DAMN minimax gets confused and is translating the english text to japanese midway through
like it still sounds hot
it does moans and everything too lmao
>>
>>109445826
what's your niche obscure interest
>>
>>109445903
kino
>>
File: miketyson.png (658 KB, 717x1131)
658 KB PNG
bye bye ltx, we wont miss your melted plastic ass
>>
What do BFL do now?
Release their safetymaxxed model to wide mockery? Uncuck it? Or just quietly pretend they never announced an open weight version
>>
anyone tried alien pushups yet
is it even worth it
>>
t2v test at 0.4 mp

Will Smith is standing on stage at an award show, holding a bowl of spaghetti with lots of red sauce. he starts eating it with his hands.

https://files.catbox.moe/e5018y.mp4
>>
So fudnon was again full of shit? Shocking.
>>
>>109445997
if its good at actual kinos then i think it will have a sizable market
>>
>>109446028
he's hardware fudding instead now
>>
>generate a bukkake with i2v
>10s at 1mp

5min for initializing, shoving everything off into ram and then about 10min to gen the entire video, vram during genning was around 65% for some reason.

96gb ram, 5090.
>>
https://files.catbox.moe/u23jll.mp4
https://files.catbox.moe/df1hyw.mp4
looks like using CFG zerostar can improve the gens
>>
>minimax can't be used in the US, UK or EU
>/ldg/ full of gens
...
Hello, saars!
>>
kek the reference model is wild, heres a proof of concept test at 0.2

prompt: <Picture 1> is saying hello to <Picture 2> who walks on camera.

https://files.catbox.moe/78lz2k.mp4
>>
>>109446050
gm saar
>>
>>109446056
india runs this bitch now
this is our territory benchods
>>
>>109446050
we all signed the waiver
>>
When I use the ref model, how can I make it generate something in the art style of one of the images.
I already tried "<Picture 1> is used for the art style" and similar prompts, but it keeps using generic 3D animation styles instead.
>>
>>109446055
0.4 mp, you can see the quality getting notably better:

https://files.catbox.moe/0m9nmv.mp4
>>
>>109445750
spaghetti eating will smith?
>>
the setting is a western saloon. <Picture 1> is dressed as a police officer and is pushes <Picture 2> into a jail cell, and <Picture 1> closes the jail door to lock <Picture 2> inside.

lmao this is a new level of i2v shenanigans, reference genning. I guess this makes loras unnecessary cause you can provide the source material.

https://files.catbox.moe/mkb2m5.mp4
>>
Anyone try > 10 seconds? It gives me cuda memory access errors but its nowhere close to maxing out my VRAM.
>>
was wan 2.2 also this slow on release but then the turbo chads made it faster?
>>
omg, reference model is 100% worth trying.

the setting is downtown Shanghai. <Picture 1> is dressed as a police officer and pushes <Picture 2> onto the ground, and <Picture 1> kneels on the head of <Picture 2>. <Picture 2> says "I cant breathe!".

that was 0.2 to test the idea. now I will try a higher quality gen.

https://files.catbox.moe/zahbcx.mp4
>>
>>109446120
yeah it was. most people pretty much ignored it before the turbo loras came out because we didn't want to wait 50 minutes for a 5 second vid without audio
>>
>>109446120
we had teacaching before that
>>
>>109446126
gen at 0.2/0.3mp for quick memes or to test ideas, 1.0mp is gonna be long even on a 5090.
>>
>>109446120
H3 is next gen fr
>>
5090 with 64GB of RAM, initial testing with I2V gens:
>5s 480p video in ~60 seconds
>10s 480p video in ~160 seconds
Pretty good for trying out prompts before committing to higher resolutions. The best thing about this model is the prompt adherence, it's fucking nuts.
>>
How good is H3 i2v for gooning? do we still have to wait for lora?
>>
>>109446128
I don't think it's worth testing, it's handled everything I've thrown at it so far
>>
>>109446124
do "i caint breef"
>>
does more steps give better quality? i noticed comfy changed the default steps to only 20, anyone tried comparing like 50 steps vs 20 or so?
>>
>>109446135
https://files.catbox.moe/8yokpe.mp4
I did put this through a cheap upscaler as I gen on such low res. No lora or end image this was just from a start image
>>
>>109446124
0.5mp test. kino! this reference type of workflow is wild cause it's like t2v but you can add stuff the model doesn't natively know.

https://files.catbox.moe/8hovhu.mp4
>>
>>109446151
>it's like t2v but you can add stuff the model doesn't natively know.
you can do that with ltx too. i wonder if they used the same trick for this model
>>
Mr President, KJ God has hit a second optimization
https://files.catbox.moe/i7cplj.mp4
>>
>>109446166
speak of the devil >>109446127
>>
>>109446135
It's good for lots of things but not penetrative sex. All kinds of fetishistic opportunities here.
>>
File: 1771375901280769.png (24 KB, 433x288)
24 KB PNG
>>109446166
those were his parameters
>>
>>109446181
Instead of sage?
>>
No Spaghetti + Quality Control
>>109446173
>>109446173
>>109446173
>>
>>109446166
That's neat, but gotta test with stuff that has more movement to really tell the difference.
>>
>>109446190
no, you can use both of them
>>
the setting is downtown Shanghai. <Picture 1> from Jojo's Bizarre Adventure glows with a gold aura and summons a stand that punches <Picture 2> several times, causing him to fall onto the ground, and <Picture 2> says "I cant breathe!".

no stand, but holy kek

https://files.catbox.moe/7pgaay.mp4
>>
>>109446192
who cares nigga
>>
>>109446198
Does the order matter?
>>
>>109446192
dude, you're a bit late but ok, i'll move, even though rebaking this late is deletion bait and can split the conversation
>>
>>109446202
I care, now fuck off
>>
>>109446190
it skips steps where it determines that nothing of substantial value happens
might not skip anything for fast motion stuff
>>
>>109446208
ok you go cry in your splinter general while we have fun in here
>>
>>109446180
>but not penetrative sex
damn that's too bad. Looks like I have to keep using LTX 10eros
>>
>>109446207
Video script is bugged on my end (hence green thumbnail), I expected someone who got it working properly to bake instead, but everyone just gave in to the obvious troll?
>>
They 100% trained the model on hentai
>>
Can H3 do feet and or footjobs
>>
>>109446259
short answer: ya
>>
the setting is a cyberpunk city at night. the man in the white jacket from <Picture 1> is standing on the road looking up at a gigantic hologram of the red hair anime girl in <Picture 2>.

the audio can be fixed with prompting but still, pretty neat

https://files.catbox.moe/ebcik1.mp4
>>
>>109446276
and this time 0.4mp

https://files.catbox.moe/bxa5ld.mp4
>>
>>109445826
Wow, this guy stood up while the scan happened...
>>
>>109446259
Feet, yes. https://files.catbox.moe/584ncx.mp4
>>
>>109446338
My feet will get cramps doing shit like this
>>
>>109446342
Hydrate more. Hydrate also means electrolytes. Research it.

You may also need stretching. Could also have spine issues with ischias nerve messing with your muscles.
>>
File: 1774757581615588.mp4 (1.09 MB, 640x640)
1.09 MB
1.09 MB MP4
https://files.catbox.moe/fwvi3y.mp4
bros we UNIRONICALLY WON
>>
>>109446345
Thanks
>>
gore test https://litter.catbox.moe/uczoltkh8y5u5g45.mp4
>>
>>109446359
won... but at what cost?
>>
>>109446359
imagine the potential
>>
>>109446338
Which SCP is this?
>>
Can H3 do bestiality?
>>
you mean black people?
>>
0.3mp/10s + easycache (test)

https://files.catbox.moe/mptd9n.mp4
>>
FUCK JANNIES
USELESS RETARD ANIMAL
>>
>can generate full blown hentai sex scenes with penetration and moans, zero need in loras
>just werks even with a 3060 12gb
jesus fucking christ
>>
>>109446359
>AI yuri
now we're talking!
>>
>>109446192
i told you it was deletion bait, oh well, we'll make a proper bake later
>>
Janny = tranny! Confirmed.
>>
File: download.jpg (46 KB, 600x603)
46 KB JPG
i told you to behave yourselves
>>
>>109446479
>User impersonating as 4chan staff
>>
>>109446469
bro im crying, I do softcore because im not homosexual, but we've never wonned so fucking hard
>>
File: MiniMax_H3_00023_.webm (3.79 MB, 1024x1024)
3.79 MB
3.79 MB WEBM
This is me shooting jannies for being trolls.

I haven't laughed this hard at genning videos in a long time, h3 is amazing.
>>
>>109446485
>we've never wonned so fucking hard
true that, that moment feels better than Flux.1 and ZiT moments
>>
do i delete all my wan and ltx shit to make room for minimax?
>>
>>109446120
wan 2.1 took 50min for 5s on a 3090 q8 max res for a few days on release
>>
can someone address the elephant in the room? how is LTX's audio so shit while H3's is so good, even though H3's audio uses fewer tokens?
>>
>>109446512
before teacache or lightx loras it was pretty slow but the quality was better too.
>>
>>109446515
better quality data
>>
>>109446515
it's stereo and it's processed inline, not just layered on top afterwards
>>
>>109446515
sounds the same to me
>>
where is turbo lora
>>
>>109446515
because ltx trained on veo 3 which had dogshit audio and h3 trained on seedance which has better audio
>>
File: 1755897982122228.png (1.4 MB, 1024x1024)
1.4 MB PNG
>>
File: 1781039138009176.png (271 KB, 1731x787)
271 KB PNG
holy shit, it works (and is a big speed up):
>>
>>109446528
Yes where turbo lora
>>
File: 1769690475804977.png (1.01 MB, 1024x1024)
1.01 MB PNG
>>
File: 1775653929727687.png (1.21 MB, 1024x1024)
1.21 MB PNG
>>
>>109446544
first test at 0.3mp/10s, 143s (4080 16gb/64gb ram). so heres the trick, use the patch sage kj node NOT the startup args or it will use too much vram (in my case). then do easycache in line.

the pink hair anime girl plays her rock guitar and plays a rock song, as she walks around the stage.

https://files.catbox.moe/v9u0ov.mp4
>>
File: 1774514951031550.png (1.15 MB, 1024x1024)
1.15 MB PNG
>>
The VAE decode for me is either near instant, or takes like 15 minutes. My guess is it's because I have 32 GB RAM. Anything I can do to fix this? First gen was long, then short, then long again.
>>
>>109446561
update comfy and pytorch
>>
File: awij.gif (2.1 MB, 319x320)
2.1 MB GIF
H3 WON. SEEDANCE LOST.
>>
>>109446566
>pytorch
good point, thanks anon
comfy is on nightly already
>>
>>109445801 Took 19:41
Picrel with easy cache took 13:28
>>
Can I throw manga into the ref h3 and tell it to animate it in order?
>>
is sampler preview not working for anyone else? its showing the first frame the entire time
>>
File: 1770360074967968.mp4 (63 KB, 498x340)
63 KB
63 KB MP4
>>109446359
kinda mid

this is seedance
>>
File: gta2.mp4 (1.35 MB, 864x480)
1.35 MB
1.35 MB MP4
>>
eyeroll is another thing it can do that wan and ltx can't
>>
>>109446579
35 stars status?
>>
>>109446532
>h3 trained on seedance
is this how api cucks cope?
>>
0.4mp, easycache skipped 8/20 steps, sageattn doing its thing, 209s. but yeah, kijai sage node + easycache with the posted settings is a big speed increase. even with no turbo lora (yet).

https://files.catbox.moe/but02o.mp4
>>
>>109446515
LTX did it really stupidly by doing 2 separate streams side by side instead of 1 stream together.
>>
>>109446580
Did it just know that car from GTA or did you include it as a reference?
>>
File: honey.mp4 (576 KB, 608x352)
576 KB
576 KB MP4
>>109446574
i think so but i'm not yet exactly sure how to best prompt it

maybe you tell it to "cut to Image 2" and so on?
>>
>>109446444
From limited testing - probably yes, at least in i2v.
>>
>>109446593
It knew the car
>>
that first countdown ended about 20 hours ago, I should probably sleep but damn I can't
>>
>>109446511
LTX yes, WAN maybe keep if you have very specific loras. It can do SOME nsfw so you kinda have to test them or wait to see what other anons test.
>>
added teacache in the subgraph
result was blurry

where do I add sageattention?
last time I tried using sage was 1 year ago on wiindows, I failed to make it work.

Now I'm on Linux
>>
File: 2026-08-03_0.mp4 (1.31 MB, 1212x676)
1.31 MB
1.31 MB MP4
Fun meme model.
>>
>>109446614
>Now I'm on Linux
bruh
>>
>>109446614
It's a command line argument to comfyUI and you have to install the python packages.
>>
>>109446561
that means you are overfilling from vram
>>
>>109446614
You just install it and use the --use-sage-attention flag. You'll see picrel when starting comfy if it's working
>>
>>109446614
I have sage right after easycache, on auto preset

using patch sage kj node (after installing sage/triton via instructions)
>>
>>109446581
stop lying
>>
>>109446593
>>109446601
It's a bit different, the stallion in SA doesn't have the 'cadillac' wings on the back, but I can see how they ended up there, still very much looks like an SA car tho, very very good results
>>
Has anyone gotten a vid longer than 12 seconds to work?!?!
>>
>>109446629
So what can I do? Gen lower res?
>>
>>109446638
is it really so hard to believe
>>
>>109446614
>>109446618
>>109446619
do I add it right before teacache? what do I choose?
>>
>>109446642
Yeah I've done a few 15 second gens at 480p
>>
>>109446642
even 30 sec works.

>>109446643
for vae decode? Use lower tile size. Or just have other stuff using vram closed
>>
>>109446644
show an eyeroll example
>>
https://rentry.org/ldgcollage_v2
(Fable) added support for catbox/litter, added failsafes to prevent video failures if tabbing away during processing

as always:
>Maintain Thread Quality
>>
>>109446658
(Unfortunately I don't know why anon got a green thumbnail, maybe he can vibedebug that on his side)
>>
will smith photobomb

https://files.catbox.moe/fgt9ft.mp4
>>
>>109446654
I'm not gonna post the one I have, bear with me just 30 seconds of this current gen
>>
>>109446614
how do you use this with res_multistep?
>>
>>109446670
stop doing will smith videos. you are making a certain group of trannies have a melt down
>>
>>109446681
>a certain group of trannies have a melt down
you mean trani?
>>
>>109446685
whoever it is that tried to splinter the thread earlier
>>
https://files.catbox.moe/a8g288.mp4
hmmmmm
pronunciation on some words is ass, but I suppose audio isn't the main draw. you can always edit audio after the fact.
>>
ok, easy cache is worth it with 0.2, 0.2, 0.9 as the values. The default hurt quality. Default res + frames is only 1 min on a 3090 now
>>
Any way to preview the video during generation? Model preview override doesn't show anything.
>>
should i not use <5s durstion? i get yellow squiggly line artifacts, i also use q2 of TE, but that shouldn't matter
>>
k finally got it to work
https://files.catbox.moe/23zd3k.mp4

someone said you should use res_multistep for better quality, where is that?
>>
File: nigga...png (225 KB, 498x447)
225 KB PNG
>>109446712
>i also use q2 of TE, but that shouldn't matter
>>
I tired doing an image edit using H3 and I can't get it to work at all. It's completely broken for me but some anons in here have shown it's possible. How are you guys doing it?
>>
>>109446709
Mine worked out of the box.
>>
>>109445997
Too late to abort at this point. They were actually about to release a video model before, and made announcements about it on their homepage, then Wan was released and they just memory-holed it.

This time though they can't really bail, so they will have to release something, and yes, it will be safetymaxxed, likely not even as good as Minimax, and all NSFW training (and thus overall community support) will happen on Minimax so BFL will have another DOA model release.
>>
I'd be really interested in this node for h3
https://github.com/kijai/ComfyUI-KJNodes/blob/main/nodes/ltxv_nodes.py#L178
I don't think it should take much to get it to work
>>
todd...

https://files.catbox.moe/43hdnd.mp4
>>
>>109446717
ok, fixed, it doesn't. just needed to remove easycache. and fuck you nigga, using non MoE 32b te is retarded as fuck
>>
Easy Cache fucking ruins the output. Shaved maybe 30% off the gen time but it looks way worse.
>>
>>109446715
default setting in comfyui wf
>>
>>109446736
no, use 0.15, 0.2, 0.9, even highly detailed fast movement is hardly effected there
>>
>>109446630
I am getting a deja vu feeling.
If I go ahead I will probably end up going in circles and messing up my comfyui
>>
>>109446654
https://files.catbox.moe/0ryo6e.mp4
>>
>>109446748
its 100% worth it, its a 2x speed up and less vram used. Dont use the start up flag though, use the kj sage attention patcher node
>>
>>109446748
https://github.com/DazzleML/comfyui-triton-and-sageattention-installer

this worked well for me
>>
so is turbo lora ever coming out or are we doomed?
>>
>>109446751
ltx can do that
>>
>>109446755
Bro calm down. Turbo loras always take a while, it'll probably be a week.
>>
>>109446761
get fucked, you want proof but don't provide it yourself
you've exhuasted my good faith, good day.
>>
it knows what a ahegao / orgasm face is...

>>109446755
devs said they were gonna release a faster sparse attention method that the model was already trained with. And if they or lighttricks does not release a speed up lora then someone else I know will in a few weeks
>>
>>109446755
it hasn't even been 24 hours
>>
the anime girl fires a massive blue beam from her wooden magic staff, that hits the world trade center towers in new york city, causing a huge explosion, and the towers collapse.

anime girls did 9/11

https://files.catbox.moe/1xjb8r.mp4
>>
>>109446740
It's not really the motion it's just the clarity of the video, it looks like it's way lower resolution or encoded with really crappy quality settings. I'll try those settings.
>>
>>109446779
funny, she doesn't look jewish
>>
in the official workflow anyway to have a set seed? I hate sub graphs like you wouldn't believe
>>
>>109446770
calm down. i wanted to see your definition of an eye roll since i have certainly had those kinds of eye rolls in videos before
>>
Do I have to update comfy for this?
Is pozzed comfy safe to use?
>>
File: 2026-08-03_1.webm (466 KB, 752x416)
466 KB
466 KB WEBM
1013s -> 850s with easycache
>>
round 2, gonna try higher quality next. (did 0.3)

https://files.catbox.moe/524n8v.mp4
>>
>>109446789
I said good day, sir.
>>
reminder that you will want to use a prompt enhancer of some type to get api level quality. Short prompts work but wont be as good
>>
>>109446748
if that isn't a cli consider switching to a cli, it can install/check/debug whatever.
>>
>>109446793
If that's the same prompt it's way worse, just looks like shit and the transformation at 3 seconds.
>>
just keep in mind the model is much faster at lower resolutions, 1.0 on a 5090 is probably a long time even. you can get quick memes even at 0.2 with i2v or t2v. but yeah, the quality goes up if you bump it a bit.
>>
easycache literally makes my gen have mustard gas artifacts
>>
eh hehehEarg
>>
>>109446812
No, it’s not the same prompt. I like that one much more it’s less cheesy.
>>
on the 5080, im getting 10 seconds, 0.4mp, 20 steps at just around 220 seconds
>>
don't wanna belabour the point but I really love this model, and I've barely scratched the surface
thanks miniclip you commie toads
>>
>>109446709
>>109446725
my bad, I'm using a fresh comfy install and I forgot to set the preview method in the settings
>>
https://github.com/sumeetprashant/ComfyUI-SolAttn
>>
>>109446841
set it to what?
>>
>>109446847
sumeetprashant
and
claude
SIRS
>>
>>109446848
auto worked for me
>>
>>109446359
>>109446473
>>109446579
>>>/lgbt/
>>
>>109446872
yuri is the straightest genre tho, you faggot
>>
>>109446865
damn, that's what mine's set to and I only get the first frame
>>
>>109446755
Distilling a model takes time.
Relax doomtard
>>
>>109446883
yeah it's weird for me too it's all fucked up in the main view but it works fine in the subgraph
>>
File: MiniMax_H3_00022_.webm (1.44 MB, 1248x1664)
1.44 MB
1.44 MB WEBM
>>
>>109446876
Yet it is the genre with biggest tranny audience
Curious.
>>
>>109446876
Technically true, as much as its main demographic likes to pretend otherwise
>>
Can it generate penas and vagene sirs?
>>
>>109446901
proof?
>>
how to replace a character in an existing video with a reference?
>>
>>109446896
hey it's matches her real smile
>>
sage attention takes a minute less to generate but it makes my GPU fans ramp up like a jet engine... I'm scared bros...
>>
anyone got a good sys prompt for prompt enhancing minimax slop?
>>
>>109446908
it can't generate on it's own but it doesn't make eldritch horrors from i2v
>>
can you use minimax as txt2img model by just doing a single frame of video? I think you could do it with WAN in the past
>>
>>109446940
no chud wait for flux 3
>>
>>109446695
most noticeable it makes audio worse
>>
0.3mp, this time just trying with sageattn kj node linked to basic guider, 82 seconds on a 4080. with easycache it was 60s. this way has more quality though in motion (skips less steps). so depending on the type of gen add easycache for an extra speed improvement. this is with 0.3/20 steps though (just sageattn)

the white hair anime elf with turns to face the camera and takes off her white coat, to reveal a white bikini. she smiles and waves to the camera, while remaining silent.

https://files.catbox.moe/17skwt.mp4
>>
>>109446947
What happened to the belt?
>>
>>109446946
I noticed that. Sounds almost as bad as LTX.
>>
https://files.catbox.moe/7c917v.mp4
hopefully we can find fixes to improve length and resolution for ramlets
>>
>>109446949
>What happened to the belt?
its cosmetic
>>
>>109446949
magic happened of course! it was a magic belt.
>>
>>109446908
no, scuffed
>>
for now, easycache + sage is a good tradeoff for speed, at least till we have a proper turbo lora or whatever. it's not even the end of day 1. we'll get it within a week or two.
>>
File: MiniMax_H3_00037.mp4 (1.68 MB, 640x1152)
1.68 MB
1.68 MB MP4
>>
>>109446940
wan didn't batch frames, you could output a single one in a latent chunk but ltx and h3 don't work like that, which makes them faster so don't knock it
>>
pp in vagene status?
>>
>>109446989
chin is on point
>>
>>109446992
it can do the acts but but the genitals look bad.
>>
the first half was not quite right, but she focused at the end. with easycache + sage kj node, 0.3 mp/10s

https://files.catbox.moe/0loe4y.mp4
>>
File: 1784293399359722.jpg (204 KB, 1553x1600)
204 KB JPG
Can someone point me towards a 16gb vram workflow for i2v?
>>
File: MiniMax_H3_00028_.webm (1.36 MB, 1184x1760)
1.36 MB
1.36 MB WEBM
>>
File: ComfyUI_2026-08-03_00017_.png (1.91 MB, 1024x1536)
1.91 MB PNG
>>
>>109447021
i'm using the ones comfyui comes with.
>>
>>109447021
we are all using the same wf
>>
>>109447010
if loras are trainable and work well, we finna eat good
i'm already impressed how much this model does just out of the box
>>
>>109447036
>>109447035
so just lower res/frames until it works?
>>
damn, easycache is kinda crazy it's almost a 50% speedup in my case
wish it didn't scuff the audio as much though...
>>
>>109447036
>we
>>
This took about 10 mins on a 5090, not bad at all.

https://files.catbox.moe/s6dx8o.mp4
>>
>>109445801
My 5080 could make this in just 20 mins? Unreal
>>
>>109447043
Or high res if you have the time.
>>
>>109447042
loras should be trainable and easier than wan 2.2 since it doesnt have low/high models
>>
>>109447060
well it's unusable as it is without the helper lora, so fucking get on it, chop chop
>>
>>109447060
WIP
https://github.com/AkaneTendo25/musubi-tuner/issues/106#issuecomment-5162508544
will be a bit before int8 training / using the pruned model
>>
surprised it stayed sfw desu

https://files.catbox.moe/rlioxt.mp4
>>
https://litter.catbox.moe/f7o7y6i5bwjk7aoc.mp4
https://litter.catbox.moe/gf9mov23q6mu4don.mp4
>>
>>109447029
Is this an H3 t2i workflow?
>>
It's great with references, one image with two characters and another in a separate image, perfect. Making someone appear midscene is a challenge wit i2v, but with references it's easy.

https://files.catbox.moe/qzl7i4.mp4
>>
Or just use LTX.
>>
>>109447060
if this model were split in two then it would be much more accessible to train, but it would take twice as long
it's a trade off but cheaper hardware is the clear winner
>>
File: MiniMax_H3_00030_.webm (1.34 MB, 1280x1632)
1.34 MB
1.34 MB WEBM
>>
>>109447084
nah, the 2 model setup is utter garbage and wastes so many params. It would be like 80B total if they went that route
>>
>>109447077
so is the reference just like "character from <picture 2> comes in" kinda stuff?
>>
>>109447074
krea2
>>
>>109446912
aw fuck it worked:
<Video 1> is the source video for the editing task.
<Audio 1> is the synchronized audio track of <Video 1>.
use the character from <Picture 1>, entirely replacing the original leftmost character in <Video 1>.

borrowed part of the prompts here:
https://www.modelscope.ai/models/MiniMax/MiniMax-H3
>>
>>109447105
https://files.catbox.moe/jj4va3.mp4
>>
https://litter.catbox.moe/2q2f6arjhz5u9dcf.mp4
>>
File: MiniMax_H3_00042.mp4 (2.37 MB, 768x1376)
2.37 MB
2.37 MB MP4
>>
>>109447092
Pretty much. At the beginning I made a declaration and then in the sequence I gave the rest of the details in order. "Use <Picture 2> as the deformed man that arrives for CUT 3" then "the deformed man runs towards the girl lying on the floor, he falls to his knees and screams "No!!" in horror while raising his fists to the sky."
>>
>>109447120
cute
>>
int8 vs int4 comparisons
https://litter.catbox.moe/mqx0qc7meny95zs0.mp4
>>
easy cache comparison
https://litter.catbox.moe/6tpt9397mmqe4h8z.mp4
>>
>>109447137
Which one is which? What is the third one? Why is size/speed difference so tiny? Are the GB values size of model or VRAM use? Hello? Any explanations whatsoever???
Holy fucking shit stop being an autist for two seconds and think about how your post will be interpreted.
>>
the reference model is fun too.

the setting is a cyberpunk city at night. the man in the white jacket from <Picture 1> is standing on the road looking up at a gigantic hologram of the anime girl in <Picture 2>

picture 1 is literally me, 2 is miku. cinema kino.

https://files.catbox.moe/jspwpr.mp4
>>
>>109447159
Look at the bottom retard
>>
File: MiniMax_H3_00043.mp4 (2.28 MB, 768x1376)
2.28 MB
2.28 MB MP4
first try, got what i wanted. h3 is really a game changer
>>
>>109447168
Yeah you're ill we got it the first time
>>
File: MiniMax_H3_00032_.webm (1.34 MB, 1248x1664)
1.34 MB
1.34 MB WEBM
>>
>the chink spammer is also the nigga tongue spammer
oh
>>
File: ComfyUI_2026-08-03_00021_.png (1.61 MB, 1024x1536)
1.61 MB PNG
>>
>>109447165
Was chopped my browser's player.
Regardless you were testing this apparently: https://huggingface.co/tsolful/Minimax_H3_INT4Mixed/tree/main
Both versions of int4 here are heavily mixed. A bit important detail to omit.
I should still thank you though, I was curious about mixed int4/int8 inference speeds. Didn't expect them to be so close to just running int8 altogether. Disappointing.
>>
>>109447179
what makes you say that?
>>
File: 7454675.webm (3.85 MB, 420x291)
3.85 MB
3.85 MB WEBM
yaaaaaas queen
>>
where gguf
>>
>>109447201
Make your own
>>
>>109447196
boudicca bros we won!!
>>
>>109447201
cumfart banished them and nobody liked that
>>
File: 634564142.mp4 (3.34 MB, 608x864)
3.34 MB
3.34 MB MP4
>>
File: 1771551842073882.png (35 KB, 1385x218)
35 KB PNG
>>109447201
you dont need those chud
>>
>>109447214
>>109447214
>>
>>109447201
literally just free up some SSD space. Its GGUF is both worse and slower than just streaming the weights. Int8 it works on at low as 8GB vram
>>
>>109445740
DANGEROUS! ahahahahaha

i fucking love china
>>
YOU CAN USE A BJ VID AS A REF + a character ref
>>
>>109447272
proof?
>>
newfaggot here, I haven't been in the habit of browsing ldg

Has wan 2.2 been surpassed yet, for open source localhost models? I know there are alternatives but it's ultimately more of the same right?

Anything localhost that feels like grok imagine, wan 2.5 and the other modern closed source models? According to my online research wan 2.2 is still all we've got really

Some custom checkpoints seem to be almost as good as wan 2.5 at creating motion tho
>>
>>109447286
Read the thread redittard
>>
>>109447313
I don't browse reddit (unlike you), and I doubt i will find the answer to my question by reading the thread. I'll have AI take a look at it as I'm sure the answer is not there and I will not waste my time reading it myself
>>
>>109447272
>>109446607
>>
>>109447448
You're wasting everyone time, shoo shoo
>>
>>109447286
Just use MiniMax H2 and wait for loras
>>
>>109447459
>us local difusers and our secret club
>>
>>109445789
just put it on auto if you mean the node.
>>
>>109445793
if on windows you can download the latest version if not on windows you have to build it, its a pain in ass trust me and requires some changes to some code file to even build it...

I was looking at it last night actually because you do need the lastest version to use kj's magic nodes that reduce peak memory.
>>
>>109447522
yeah thanks, I'm checking this one now, I ask the AI from time to time to check for news but this wasnt on its radar yet
>>
>>109447558
cumguzzling retard
>>
>>109447557
might not be necessary on some linux but was on my arch because version of cuda or gcc.
>>
>>109447137
Please do ref2va too. I honestly don't understand the point if fl2va because you can just instruct ref2va to use one pic as first and second as last, it understands it.
>>
File: 1773505305638338.jpg (75 KB, 1125x1125)
75 KB JPG
>>109447286
>
>>
>>109447159
Jesus Christ, you are retarded nothing more to be said.
>>
https://streamable.com/j5pisb
>>
>>109447806
shhizos
>>
>>109447806
I like how 1girl became a countable noun.
>>
i love new model releases, everyone flocks to ldg again and starts posting meme gens and everything is fun again for a while
>>
File: MiniMax_H3_00010_.mp4 (2.22 MB, 864x480)
2.22 MB
2.22 MB MP4
>>
File: MiniMax_H3_00011_.mp4 (1.63 MB, 448x672)
1.63 MB
1.63 MB MP4
who keeps making the retarded threads



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.