[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109465146

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109466297
thanks for the bake anon, godspeed
>>
<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> plays a rock song on her guitar, then hits <Subject 2> in the head with the guitar several times, breaking the guitar. <Subject 1> steps into a cardboard box while looking shy.

https://files.catbox.moe/3knoxq.mp4
>>
>>109466320
valid crashout desu, he tried to swallow her with his body or something!
>>
>>109466297
>/g/ still doesn't have sound in the year 2026
come on this is unnaceptable :(
>>
>>109466332
Literally 0 reason not to have it sitewide
>>
>>109466334
they could prevent 'muh jumpscares' or whatever other retarded motivation they have by having the videos start muted by default
but it's /g/ we're talking about :)
>>
>>109466338
>by having the videos start muted by default
Which is already how it works on the boards that have it. Actually less than zero reason because it's extra shitcode that makes the site worse for no fucking reason whatsoever.
>>
I'm goin for 14 seconds 1MP, might be a while
>>
>>109466334
the new site janny broke the code for timestamps and tripcodes. i don't think he knows how to enable permissions for sound
>>
File: MiniMax_H3_00060_na_u.mp4 (3.27 MB, 960x544)
3.27 MB
3.27 MB MP4
Another thing to work on: environment consistency over longer times / multiple shots.
>>
im surprised how clear the audio is in this model. and this was at 0.3 and still pretty clear!

<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> plays a rock song on her guitar while standing near <Subject 2>, then hits <Subject 2> in the head with the guitar several times. <Subject 2> falls on the floor and says "bocchi, I cant breathe!"

https://files.catbox.moe/2ugwwz.mp4
>>
https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cache

suggests it's better than easycache. no tests/doesn't elaborate.
>>
>>109466363
>no tests/doesn't elaborate.
I hate when that happens, or else it's some vague posting king, or else it's some giant LLM wall of text that has nothing of substance in it
>>
>>109466357
Kino
>>
>>109466362
also, 185 seconds with 0.3mp/10s and the patch sage kj + spectrum nodes, this is my setup in the reference workflow:
>>
>>109466352
>i don't think he knows how to enable permissions for sound
there's no excuse we live in the LLM era he can ask claude Idk
>>
>>109466375
replace the patch sage note with MiniMax H3 Mem Eff Sage Attention Patch node.
>>
with this reference model you dont even need loras although obviously loras are best for 1:1 characters. but look at bocchi go!

https://files.catbox.moe/2bznao.mp4
>>
this model is surprising me. i can actually prompt for the tank commander to peek down into a detailed tank turret with crew members doing their job
>>
>>109466391
ah, I have that in my i2v model workflow but forgot to add it in the reference one, ty anon
>>
>>109466393
>with this reference model you dont even need loras
Ikr, it's amazing at keeping the character consistency, Klein dreams of being this good, people on twitter have to bully the Minimax CEO into making an image/edit model, he would nail that shit too
>>
>>109466348
I hope it's just the preview fucking up but it's showing nothing but noise
>>
>>109466401
adding that to the reference workflow speeded things up as well, I guess it optimizes the vram or memory or whatever. pretty cool how rapidly people are optimizing this, then again not surprising since it's a genuinely good model and not slop.
>>
I havent gen'd since an old wai SDXL a few years ago and was wanting to pick up some current models. Is krea2 for realistic and anima for 2D the current meta or ?
and do I really have to make an account to download krea2? I cant find it on civitarchive, and its cucking me on HF
>>
https://files.catbox.moe/zonwbf.mp4

punching aside, I love how this model can actually make realistic sounds, the guitar sounds like a guitar etc.
>>
LTX and wan might as well just close up shop, and BFL, other than klein edit 9b and krea 2, wtf do you need other than this

https://files.catbox.moe/b5kvq7.mp4
>>
>>109466393
Can you give multiple references for the same character? Like a front view and a back view
>>
https://huggingface.co/Kijai/MiniMax-H3-TAE
This shows noise for me. kjnodes is up to date. anyone else?
>>
>>109466430
yes, front back and side keeps it from guessing
>>
hahaha, this is a whole new level of shitposting potential.

<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> plays a rock song on her guitar while <Subject 2> plays the drums, <Subject 1> is singing "he can't breathe, he just wants some fent, who knows where it went, George Floyd.".

https://files.catbox.moe/xr8sv0.mp4

>>109466430
yep, you can give camera directions and even specify shot by shot, ie 0 to 3s: view from behind subject 1, etc
>>
>>109466437
The structure is even more sophisticated in the example at the bottom of the guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
bocchi concert part 2, I didnt even prompt the camera directions.

https://files.catbox.moe/qtdmje.mp4
>>
Okay, minimax is pretty fucking great
https://files.catbox.moe/ysfmy6.mp4
>>
Just deleted wan and ltx lol
>>
>>109466408
Okay it came back, kind of it's still noisy but there's something in there. I don't think the preview was designed to deal with latents this big
>>
>>109466486
based
>>
>>109466313
I dunno how it performs but you can combine sol and sage.
They modify different things and are in theory compatible.
>>
File: 2309498385.jpg (43 KB, 1000x591)
43 KB JPG
getting some good music from this
>>
File: MiniMax_H3_00064_na_u.mp4 (3.56 MB, 960x544)
3.56 MB
3.56 MB MP4
Continuity is difficult. I think I will need to try feeding it a whole storyboard.
>>
Any resources to getting started with minimax h3? I never opened comfyui in my life.
>>
>>109466560
Click the template button. Load the minimax h3 t2v template. Download the models. Lurk these threads for tips
>>
<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

this is the key to getting interactions in the reference workflow. then just say <Subject 1> or 2 does (whatever).

https://files.catbox.moe/z0543u.mp4
>>
>>109466558
we don't have the SCAIL2 type equivalent to this yet but even just passing references to past scenes and characters and making use of 20s+ segments should do a lot to get "continuity"?
>>
>>109466572
Does it still work if you do something like "<Subject 1> is the anime girl in <Picture 1> and <Picture 2> and <Picture 3>"?
>>
>>109466560
as the other anon said

if you prefer a simpler UI for now wan2gp also supports it, but you won't be able to use various of the nodes people here use.
>>
KINO

<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is the city Tattooine in Star Wars. <Subject 1> walks towards <Subject 2> and says "you tried to sell my lightsaber for fent, and now you must die.". <Subject 2> holds up a lightsaber with a red beam and says "I WILL get high, Skywalker." in a black man's accent.

https://files.catbox.moe/9wmqx4.mp4
>>
>>109466588
once you establish picture 1 (image 1) as subject 1, always refer to them as subject 1. like a variable or whatever. the model is assigning <subject 1> to that image. if you use multiples for one name it might get confused
>>
great more floyd gens
really funny lol!!!!!!! so much creativity!!!!!!!!!!!!!!!
>>
>>109466596
I hope you're making other stuff too, you don't have to try to impress us
>>
>>109466600
I meant for using multiple reference images for the same character. I've seen the version where you specify that the face is from <Picture 1> and the body from <Picture 2>, but I was wondering if there was a way to get it to use multiple images for the entire character.
>>
>>109466596
https://files.catbox.moe/v2qwhg.mp4
>>
>>109466608
In this house George Floyd is a hero end of story.
>>
>>109466608
we tell stories of our heros todays so that they may be remembered as gods tomorrow
>>
>>109466613
I am, Miku and Floyd are my "hello world" test case just to test concepts.
>>
>>109466618
Better than he ever was
>>
>>109466391
nta but can't both not be used? hmm I'm currently using both do they conflict or something? I'm always looking to squeeze a bit more out of my shit hardware.
>>
>>109466415
ok well back to wai SDXL i guess. anyone got a link to the local tag DB with autocomplete?
>>
File: videoframe_72.png (480 KB, 736x416)
480 KB PNG
>the sand people will be back and in greater numbers
>>
>>109466638
XD
>>
>>109466638
go back dumbass
>>
i still dont know what wai stands for
t. newCHAD
>>
>>109466642
>>109466644
:( not very nice!
>>
>>109466655
Waifu
>>
shieeeeeeeeeet

https://files.catbox.moe/4c7nwn.mp4
>>
File: H3_Combine.jpg (195 KB, 1470x888)
195 KB JPG
>>109466558
>>109466574

Chaining shots together is an option. I still need to git gud at shot cuts and transitions.

https://streamable.com/l6j9uy
>>
>>109466655
Lora shitmix of obsolete model popular with Bharatis and Hispanics.
Nothing you really need to know newGOD
>>
>>109466636
You just need one and that one is optimized for minimax.
>>
>>109466678
kk
>>
Got minimax h3 and it seems time to get back into this again.
Still hate writing prompts.
>>
>>109466678
Well whats the goto for animoo then? anima?
>>
>>109466663
We'd have killed back in high school for this tech
>>
>>109466703
I am rawdogging short lazy prompts and it is working good compared to how low effort I am with it.
Just Video:1-2 sentence description followed by Audio: 1 sentence description, no scene/second breakdowns.
>>
>>109466711
Yeah.
>>
i've been enjoying H3 so much that i downloaded it a second time
>>
>>109466728
don't be greedy. leave some for the needy
>>
What do you think about integrating RTX Super Resolution in your minimax workflows? The video looks good to me, but I don't know why it failed to reintegrate the audio.
>>
>>109466727
does it need a lora for the sexo stuffs?
>>
>>109466664
if you are serious about slopping(juding by your shit i assume you are) you might want to look into divinci resolve and more traditional ways of putting together longer videos. probably speed up your work quite a bit and give you more creative space to work in.
>>
>>109466744
It knows most sexual concepts, it just struggles with uncensored female genitals.
A lot of body horror if you try to generate vaginas
>>
File: MiniMax_H3_00065_na_u.mp4 (3.83 MB, 960x544)
3.83 MB
3.83 MB MP4
>>109466664
>>109466574
I'm generating in single 30 second chunks at a time, but it already forgets the structure of a complicated environment when it switches scenes within the same generation. I think it might be better to do only single cuts per generation and splice, but that would be a lot more work.
>>
File: MiniMax_H3_00166.mp4 (3.47 MB, 1376x768)
3.47 MB
3.47 MB MP4
>>
>>109466760
It wasn't trained with 30s cuts, I don't know what you expected
>>
>>109466744
I generated a lot of hardcore shit with it without loras, so no.
>>
Black Forest Labs insider here. We're fucked.
>>
Anyone get this to work on AMD?
>>
there is a latent upscaler, it has some problems though
https://github.com/Tr1dae/ComfyUI-MiniMaxH3_LatentUpscaler

i seems with this model earlier steps do more with the video and latter ones with the audio.
>>
>>109466772
on god.
isn't that shit releasing in a few days?
>>
>>109466767
Pretty but needs a "TRVTHNVKE" somewhere, probably on the wings.
>>
>>109466769
from what we know, the total video length is not a factor, but the length of an individual cut within a sequence may be limited.
>>
>>109466574
can't you just like do 30 seconds, then feed it the last 10 seconds of the first video and extend it 20 seconds and repeat? I thought the model was capable of that, i'm guessing you would also need to feed in character reference sheets and props. I haven't even begun to test those features out yet.
>>
File: 1773910715944006.jpg (1.32 MB, 1248x1824)
1.32 MB JPG
Jesus fucking christ, i cannot get enough of this model
https://litter.catbox.moe/p05yi0i570j87o02.mp4
>>
>>109466767
Shows how it struggles with the sequence of events, with the ground already burning in advance.
>>
File: MiniMax_H3_00167.mp4 (1.91 MB, 1376x768)
1.91 MB
1.91 MB MP4
>summary: mayli from facialabuse.com is sitting in a cafe alone. watching her own video on a laptop.
>the camera start with medium shot on her sitting on the table. after 2 seconds it shows her watching video on laptop, in camera third person watching her.

h3 t2v doesn't know who is mayli. trash model
>>
>>109466792
Yes, but the problem is that even within the first 30 seconds it already loses track of where it was. Doing one scene (cut) at a time and editing together seems like the safe way to go.
>>
File: 3157140495.gif (1.13 MB, 498x498)
1.13 MB GIF
>>109466767
i haven't tried any nuclear kinos yet with it
>>
>>109466760
try timestamps like 0 to 3s: (action) for every shot up to 30s, it may work but idk, works for 5-10s
>>
>>109466560
Comfy is really not that hard to operate, just use template workflows and remember the general control flow, then play with nodes and see what happens when you make a mess out of them.
>>
>>109466811
anon, there is an entirely separate reference model...just plug them in and say <subject 1> is the woman in <picture 1> and you are good to go
>>
this is wild, the possibilities are endless.

multi reference drifting!

image 1 is anakin, image 2 is floyd, image 3 is a random background from jojo part 7.

The video takes place entirely within the setting, and environment shown in <Picture 3>.

<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

<Subject 1> is facing <Subject 2> who is 10 feet away and says "you tried to sell my lightsaber for fent, Floyd.". <Subject 1> takes out a green lightsaber and <Subject 2> takes out a red lightsaber, and <Subject 1> and <Subject 2> have a high intensity lightsaber fight.

https://files.catbox.moe/xtmxgc.mp4
>>
File: MiniMax_H3_00066_na_u.mp4 (2.03 MB, 960x544)
2.03 MB
2.03 MB MP4
>>
File: jojosbr14.png (941 KB, 1406x1038)
941 KB PNG
>>109466834
reference background image:
>>
>>109466844
cute!
>>
File: MiniMax_H3_00035_.mp4 (1.36 MB, 864x480)
1.36 MB
1.36 MB MP4
>>
>>109466827
>>109466569
>>109466593
Thanks for the tips, I had no idea comfy had templates. I'll download it right now! I'm excited to test it
>>
>>109466486
WAN yes
But for static animations LTX still good
>>
File: MiniMax_H3_00169.mp4 (3.19 MB, 1376x768)
3.19 MB
3.19 MB MP4
>>
fent ball run. I added a music track from the show.

The video takes place entirely within the setting, and environment shown in <Picture 3>.

<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

<Subject 1> and <Subject 2> are riding horses side by side. <Subject 1> and <Subject 2> ride their horses very fast as a crowd watching in the stands cheers, and <Subject 1> crosses the finish line first, with <Subject 2> close behind him.

https://files.catbox.moe/9i347q.mp4
>>
>>109466767
Some anon's mom really shut it down geg
>>
2026 truly blessed, we got krea and minimax and it's only august
>>
>>109467036
we also got anima in january iirc
and we got the reject to mostly stop trying to shill his trash
>>
Have any of you cucks mastered ref prompting in Minimax yet?
>>
It feels like I was meant for more in life
>>
>>109466943
How are you getting such a clean video at that res anon?
>>
>>109466420
>other than klein edit 9b and krea 2, wtf do you need other than this
Anima
>>
>>109466297
Does anybody know the workflow or steps needed to get Minimax H3 to gen as crisp as this vid? Ive been trying things all day and no dice
https://files.catbox.moe/pj2jjl.mp4
>>
>>109467087
close up being crisp is expected
for non-close up the only thing i can think of is higher sampling resolution
fancy postprocessing video upscalers might help as well (not sure)
>>
>>109467077
It's the default workflow with sage attention.
>>
File: 1784898685333460.jpg (246 KB, 1400x1143)
246 KB JPG
>>109464587
Thanks i ended up doing this
>tfw faces visible for the first time
:o
>>
Is Qwen3 obsolete? Should I get rid of the models?
>>
>>109467098
Yeah Im trying upscalers atm, but ive been hitting dead ends for 7 hours
>>
>>109467119
Yeah, especially qwen3vl_32b
>>
>>109467060
Yeah and your mom is incredible; Rimming niggers, Asian gang-bang, pool sized bukaké she can't get enough.
>>
Any fag wants to test res_2s beta57 version of any of his already posted clips for comparisons?
>>
Btw - you can somehow fit int8_convrot qwen text encoder in 32 + 64, but with UnloadModel node
>>
reference model is so fun.

The scene is set inside the wooden horse-cart from <Picture 1>. The original blonde Nordic character from <Picture 1> must remain seated on the bench exactly as shown. Introduce the new character from <Picture 2> as <Subject 1>. The blonde Nordic character from <Picture 1> is <Subject 2>. Position <Subject 1> sitting directly beside <Subject 2> on the same wooden carriage bench. <Subject 1> turns to <Subject 2> and says "where is the fent at, nordic brother?". <Subject 2> says "we only have skooma" in a swedish accent. The carriage bumps and sways down the dirt path, causing both characters to physically react to the bumpy ride together while looking at each other.

https://files.catbox.moe/5wfgc1.mp4
>>
how clean do voice samples need to be to be usable?
>>
File: ComfyUI_7785.jpg (1.14 MB, 2328x2328)
1.14 MB JPG
>>
Can confirm MiniMax H3 Sigma Shift is a no-op if you apply it to guidance only, it uses the default values in that case.
An anon asked a few threads ago, just confirmed my guess now
>>
>>109467153
Not that clean. Im making a zero shot cloning workflow atm, its really good so far Ive tested
>>
File: ComfyUI_7787.jpg (1.27 MB, 2328x2328)
1.27 MB JPG
>>
>>109466375
Try it with a video ref and watch it drag
>>
>>
>>109466608
That's SAINT G. Floyd to you
>>
>>109467128
I use "4x_NMKD-Siax_200k.pth"
You should try it and show us the result.
>>
>>109467185
more reason to hate gingers
>>
anyone know, what existing TV media does minimax have really good knowledge about apart from The Office and Seinfeld?
>>
File: MiniMax_H3_00068_na_u.mp4 (1.92 MB, 960x544)
1.92 MB
1.92 MB MP4
>>
>>109467153
not at all, I've used clips with background music and it turned out pretty good. Granted it wasn't a voice I've heard a lot and could recognise as perfect but it was well in the ballpark
>>
>>109467206
It does better if the voice is linked in the prompt to the character that you're generating also
>>
Why is video generation so complicated? I just installed comfyui to generate porn on my new GPU but wtf how many tools, plugins, extensions, models, hours of learning do you need to get something decent?
>>
So when will minimax be able to do nsfw? Will the creators successfully shut it down?
>>
>>109467185
>didnt bite the connector
smart cat
>>
>>109467212
it works without doing that?
news to me
>>
>>109467218
If youre using an audio reference, why wouldnt it?
>>
>>109467225
because you haven't told it to?
I would suspect it works better if you do
>>
>>109467213
Beg coding with ChatGPT doesnt help that much, and people are coy with the best workflows sadly. Fair enough, but very annoying when people send you down bad paths for shits
>>
Do people use "director" nodes now or are they a meme?
>>
File: MiniMax_H3_00075.mp4 (3.32 MB, 960x640)
3.32 MB
3.32 MB MP4
>>
>>109467235
some kind of boiler plate for the syntax?
>>
>>109467238
kek
>>
>>109467213
Lot of things go into making even just porn. You also need to learn about cameras, lenses, lighting, color, perspectives and so on. To make something good you still need a film school crash course basically.
>>
just watch porn and learn the angles
>>
How do I prompt H3 for POV video???
>>
>>109467252
have you tried POV
>>
<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of icecrown citadel in the game World of Warcraft. The player character is a warrior with a large sword. The warrior charges after a raid boss with the appearance of <Subject 1> and hits him with the sword several times.

im dying. and that's actually part of the raid. gonna try higher res next.

https://files.catbox.moe/cjctnw.mp4
>>
We can finally get the collabs we were robbed of
https://files.catbox.moe/ckfdr2.mp4
>>
virgin porn watchers vs gigachad sex investigators
>>
Whats the best workflow atm?
>>
>>109466793
What resolution is that video?
What is your hardware?
How long did it take to gen?
I2V or R2V?
>>
>>109467255
yes
>>
>>109467265
currently this node setup seems to work better than easycache for me at least:
>>
>>109467282
Cheers
>>
>>109467256
KINO, now this is classic wow.

https://files.catbox.moe/jjk9of.mp4
>>
>>109466391
I see just ltx and wan mem Eff Sage Attention Patch node.
updated comfy yesterday
>>
>>109467256
I'm noticing those machine elf hexagons all over my gens as well
>>
>>109467294
update kj nodes
>>
File: ComfyUI_7797.jpg (1.19 MB, 2328x2328)
1.19 MB JPG
>>
>>109466415
I use Krea 2 for anime too. Go find the krea 2 fp8 scaled version on HF or get the int8 convrot but I couldn't get that running.
>>
https://i.4cdn.org/wsg/1785931146429739.mp4

>>109467242
yes, they look like audio/video editing timeline UI.
>>
>>109467238
Reminds me of a web game I played as a kid called Fly Like a Bird 2 where you shit on people as a bird in a city. Pretty insane that it ran in my browser on my 2007 Pentium

>>109467282
Nta but thanks. Will try spectrum over easycache shortly. Minimax H3 is so good I actually want to make sfw videos with it to see what it can do. It mogs seedance 2.0
>>
>>109466338
Could literally just have a sound icon on the thumbnail for the "jumpscare" warning
>>
File: 1767421254242619.gif (447 KB, 1392x908)
447 KB GIF
360 seconds for 13 seconds video at 0.6 mega pixels

1 hour = 3600 seconds

3600 / 360 = 10

10 gens every hour

10 x 18 = 180

180 gens / day

What can i do for 180 gens ??

We need more optimizations. I wonder in time Minimax will be optimized so you can gen under 250 seconds ??
>>
lmao, not exactly what I wanted but it DOES know a lot of games.

<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of the game super mario 64. Mario is running through a level with grass and coins. Mario runs and leaps on <Subject 1> and a coin appears after Mario jumps on his head.

https://files.catbox.moe/rrkqla.mp4
>>
>>109467267
0.5 mp
3060 12gb
around 13 minutes
i2v
>>
Can you run this on 8gb vram 32gb ram on Windows?
>>
everyday posting
https://files.catbox.moe/538de0.mp4
>>
>>109467314
>Fly Like a Bird 2
i also thought that exact thing, but i didn't say anything since i didn't expect anyone here to know about it
>>
so long gay bowser...

https://files.catbox.moe/hokybc.mp4
>>
>>109467326
thanks, anon. That's about what I'm getting, as well
>>
>>109467315
>Could literally just have a sound icon on the thumbnail for the "jumpscare" warning
Any of the top 5 AIs right now would be able to fix all of the sites problems and add all the missing features in like 4 hours
At this point 4chan iws defined by the fact that you can only post 1 image per reply
>>
>>109467192
Kill yourself
>>
>>109467235
It's cool what you can achieve with noodles, but If you are going to have to edit something from separated clips you should use a proper editing software, there are free ones like Davinci Resolve, I personally use Shotcut. There might be niche situations where a director workflow can achieve what simpler workflows but outside of that simple is better.
>>
>>109467185
Always an orange fucker
>>
>>109467352
gm saar
>>
loli
>>
>>109467392
I think the director nodes are good for filling in the gaps. Like using the end frame of clip1 and the beginning frame of clip3 to generate clip2 with FL2V.
>>
Been traveling the last few days. So is minimax gods gift to coomers or another nothingburger?
>>
File: lush.jpg (21 KB, 500x528)
21 KB JPG
rate my song
late 80s metal
>>>/wsg/6207803
>>
>>109467433
I haven't genned a single coom clip. It's fun model though.
>>
File: 1759967567955188.jpg (237 KB, 2067x875)
237 KB JPG
>>109466297
Help me connect Spectrum nodes because i have no idea what i am doing here
>>
it knows EVERYTHING.

<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of the nintendo 64 game the legend of zelda ocarina of time. Link opens a door in a dungeon and on the other side is <Subject 1> who is 10 times as tall as Link. Link slashes him several times with his sword and then <Subject 1> disappears, and a heart container falls on the ground where he was standing.

https://files.catbox.moe/9wjqug.mp4
>>
>>109467258
why won't they die?
>>
>>109467435
i rate it k for kino
>>
>>109467433
>>109467436
https://files.catbox.moe/gcupql.mov
Mate made this. Dont know what the work flow is, but it seems like a gem
>>
File: 1768207589549749.jpg (394 KB, 2067x875)
394 KB JPG
>>109467444
>>
>>109467474
Thanks bro
>>
>>109467408
good morning saars
https://files.catbox.moe/lcvwdh.mp4
>>
>>109467472
Oh it doesn't autoplay .mov files
>>
>>109467472
Based and goonerpilled
>>
>>109467502
https://files.catbox.moe/fsytg4.mp4
Here we go, Misato titties
>>
>>109467472
I've seen plenty of good goon gens made with it.
>>
File: videoframe_95.png (387 KB, 736x416)
387 KB PNG
the reference model can do style transfer too. this is KCD for example.
>>
>>109467507
What I havent seen is the workflow that gets it to this level
>>
>>109467379
Can I fuck your mother first? And your sister? I already did you father.
>>
>>109467483
is spectrum actually worth it? Please report back, anon
>>
>>109467517
You are brown
>>
like this guise?
H3 mem eff sage attention before spectrum?
>>
>>109466953
>>109466827
>>109466569
>>109466593
https://files.catbox.moe/bac7ag.mp4
Ha, it worked.
>>
>>109467515
is that musa?
>>
>>109467525
Yup, that's why you father has aids now. He was kind of a slut you know?
>>
>>109467534
Now unpack that subgraph and start adding optimizations to generate things faster
>>
The default shift/scheduler/step count combo feels ass to me.
What are (You) using?
>>
>>109467546
>unpack that subgraph
Holy shit, that's a thing?! What the fuck. The rabbit hole goes deeper than I thought.
>>
>>109467534
congratulations newGOD
post some kino in the thread
>>
Henry and Floyd's adventure:

https://files.catbox.moe/b1au0e.mp4
>>
how do I load loras for minimax? which node?
>>
File: 1134703996.jpg (34 KB, 474x461)
34 KB JPG
i been tweaking the same prompt for 5 hours now and i think it regressed
>>
>>109467572
>xe overfit on local prompt minima
ngmi
>>
>>109467572
2 hours but still having fun.
>>
>>109467572
Im at 7. As soon as I find the right one, im spamming it in every thread
>>
>>109467572
you need to use an llm for optimal h3 prompts
>>
File: image.png (53 KB, 665x425)
53 KB PNG
>>109467570
>>
>write elaborite prompt using the guide
>model completely ignores the prompt and does random bullshit
>?????
>Wtf? is this normal?
>>
I used pizza tower as a style reference and this happened:

https://files.catbox.moe/h5mkda.mp4
>>
>>109467582
the llm isn't good enough for kinosovl
>>
>>109467600
maybe it needs better director.
>>
>>109467590
Don't waste time writing it yourself. Get a clanker to do it.
>>
File: trump ai.png (26 KB, 789x232)
26 KB PNG
Are we entering a local golden age?? Seems like APIs are getting raped in the ass, with OpenAI having to drop prices massively in response to chinese models, and xi has openly endorsed the open weight approach. qwen is also releasing their Max LLM locally, something they haven't done before.
But where is Qwen Image and Wan???
>>
>>109466999
trips of double ice coffee
>>
>>109467620
>Seems like APIs are getting raped in the ass
When deepseek v4 pro drops we gonna see anal gapes so big that no goatse could come close.
>>
>>109467620
looks interesting
>>
>>109467585
won't let me hook it
also the h3 mem eff sage patch is not working
>>
>>109467521
Just tried it. Nope. Probably my workflow fault. But i heard Spectrum lowers quality
>>
>>109466811
try kelly balthazar
>>
File: 00109-3598072604.png (540 KB, 640x512)
540 KB PNG
>>
Fuck off back to your containment thread retard
>>
File: H3_00009.mp4 (1 MB, 736x416)
1 MB
1 MB MP4
>24GB vram
>64GB ram
>R2V
>736x416
>177 seconds to gen

audio: https://files.catbox.moe/9b5g8u.mp4
>>
>>109466703
Just feed the H3 prompt guide to your clanker and talk to it how you want your scene to play out.
>>
>>109467527
order doesnt really matter lmao
>>
Reference model is good as a whole but wtf is up with using video references? I haven't got it to work once so far even with your basic jeet character replacement scenes.
>>
>>109467527
is the spectrum apply node snake oil or legit? i don't trust anything that isnt from KJ
>>
>>109467668
If you're having prompt adherence issues you need to adopt the official format.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

I recommend plugging the guide's source into your favorite AI assistant (e.g. Fable) and asking it how to properly prompt for your desired result.
>>
>>109467672
absolutew snake oil. even easycache is better. you just should wait for KJ to wake up.
>>
>>109467435
impressive
but too noisy and more like early 90s
>>
>>109467679

I have already tried doing so, and I'm either getting weird outputs with my <picture 1> subject or getting <video1> thrown right back out at me. I've tried using it for both NSFW and SFW clips and can't get anything out of it.
>>
fuck im a total promptlett barely able to gen images and trying h3, the audio its putting out is so fucking cursed kek. this is alot of fun though, 4:3 at 0.2mp genning 10s in under 2min just to practice prompts
>>
https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cache
this does not effect audio
>>
>>109467691
I do feel it's quite finicky so it may not be just you.
>>
style transfer test, game is resident evil 1 original: pretty cool that it works.

<Subject 2> is the man in <Picture 2>

REFERENCES:Use <Picture 1> exclusively as a visual art style, texture, and color palette reference. Use <Picture 2> exclusively for the facial identity and clothing structure of <Subject 2>.

INITIAL PLACEMENT & STYLE:The video begins with <Subject 2> positioned standing on the left in the retro video game scene shown in <Picture 1>. <Subject 2> is fully integrated into the scene from the very first frame, completely transformed into the exact pixelated, retro art style of <Picture 1>.

STYLE RENDERING:The entire video must be rendered completely in the exact retro video game art style shown in <Picture 1>. Apply the pixelated textures, low-resolution aesthetic, specific color grading, and retro rendering artifacts from <Picture 1> globally to every asset in the video.

SHOT & MOTION: <Subject 2> is walking around a large mansion, towards the door at the left of the room. <Subject 2> is fully transformed into the retro video game art style, looking like a playable game character.

using a template from google ai search, seems to work.

https://files.catbox.moe/rxpjl9.mp4
>>
https://www.reddit.com/r/StableDiffusion/comments/1veb4bn/i_created_a_sysprompt_for_minimax_h3_to_emulate/
https://gist.github.com/Naxdy/43b7422a1e4a79fb8b0489c6c39eaace
For the local bros this sysprompt.md works perfectly with Gemma 31b q4 to create H3 prompts from just single line prompts
>>
File: 00008.webm (1.15 MB, 560x858)
1.15 MB
1.15 MB WEBM
What are the good prompt settings for animating figurines in minmax? Should it be 3D CG or Live Action? Any other parameters?
I tried live-action and it looked a little weird and the output sometimes had errors.
With wan2.2 it just werked pretty well out of the box (webm related).
>>
>>109467642

>also the h3 mem eff sage patch is not working

it needs minimum i think 2.0 Sage Attn. That's my problem with it and I cannot get 2.2 working on Linux, I may have to compile it but im not in the mood for fixing shit if it goes wrong today. I don't even have GCC installed (I do but apparently it doesn't exist ...)
I got the 2.2 version from comfys repository, that didn't work for me.
>>
>>109467709
Oh so I need sage 2.0
tried installing it but failed
>>
>>109467721
just give them an action and specifically say to keep them in the style of an anime figurine or sculpted model or something.
>>
>>109467721
Throw the prompt guide into a clanker and tell it to rewrite your prompt in that style, then just tell it what you want.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
same re1 style transfer prompt but 10s test:

https://files.catbox.moe/do0ecs.mp4
>>
>>109467634
>bruh is so crazy you see video of sinde sweeny kung fu fight batman!
and then everyone forgets the model exists in a week because normies aren't wasting 20-40 dollars to gen a few seconds of video.
>>
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
https://huggingface.co/moonshotai/Kimi-K3
>>
>>109466297
Can Minimax H3 run on a 4060 8GB? Whats the hardware requirments?
>>
>>109467750
same prompt again but with link to the past:

https://files.catbox.moe/fgmy5f.mp4
>>
>>109467747
Man I don't know about using an LM to write my prompts, I've always written them by hand because I desire control.
Like, I know what I want, but the guide says you should specify whether the video is 3d cg, live action, 2d animation, etc. Moving figurines is a bit of a tricky one to quantify.
>>
I'm trying both qwen3tts and fish s2. I think claude has cucked me and I can't get what I'm after.
I want to take a voice, preset and/or voice cloned and customize it's emotions for sentences. Do I have to git gud or is there something better?
>>
genuinely confusing how we got something like H3 released open weight. this model has zero limits
>>
I used the word fentanyl in the prompt and the video became worse quality, sort of wobbly. I stopped using it and it became better.

>>>/wsg/6208528
>>>/wsg/6208526
>>
>>109467784
alright that one made me laugh
>>
>>109467828
yeah, it's even better in cunnygen than Wan 2.2. i'm kinda blown out
>>
we have AI salty bet now. just pick 2 images and have them fight. make the prompt ambiguous so the AI decides the winner. "one of them falls down", etc.

<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "I can't breathe!". <Subject 1> uses a peace sign gesture with one hand and smiles.

https://files.catbox.moe/c90sc5.mp4
>>
I NEED BETTER MINIMAX WORKFLOW
>I NEED BETTER MINIMAX WORKFLOW
I NEED BETTER MINIMAX WORKFLOW
>I NEED BETTER MINIMAX WORKFLOW
I NEED BETTER MINIMAX WORKFLOW
>I NEED BETTER MINIMAX WORKFLOW
>>
>>109467862
use dasiwa's workflow from civitai
>>
its SOOO fucking fast now
https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cache
https://litter.catbox.moe/p6c7n2nu49gwhi9g.mp4
https://litter.catbox.moe/bsrgokundp30mjc0.mp4
https://litter.catbox.moe/zjqkhr6vvsp27fmn.json
>>
>>109467866
100%|| 20/20 [02:03<00:00, 6.18s/it]
[MiniMaxH3-Cache] Skipped 10/20 steps (2.00x speedup).
That was 2MP btw
>>
>>109467862
just wait until we have nsfw and turbo loras
>>
File: 11111.mp4 (2.02 MB, 1056x608)
2.02 MB
2.02 MB MP4
https://files.catbox.moe/svh1zf.mp4
>>
PSA

https://files.catbox.moe/9pa6ao.mp4
>>
asuna from blue archive fights crime:

https://files.catbox.moe/rax8j0.mp4
>>
>>109467866
>>109467871
Tutorials ??
>>
added this lora loader, hope i added it right
>>
>>109467898
literally has a WF and link to the new node
>>
Convrot is pretty much a necessity right? I tried the non-convrot models and motion definitely felt stiffer, but that was back in only my first handful of prompts in Minimax.
Do all of you guys run convrot models?
>>
>>109467866
>2 days ago
>zero activity
snake oil, kills the quality
>>
>>109467911
bf16
>>
>>109467911
no reason not to use cr
>>
>>109467911
yeah even with klein edit and krea it's faster and bf16'ish quality, fast and quality so why not
>>
>>109467916
it truly does not
https://discord.com/channels/1076117621407223829/1534342635597070466
>>
>>109467916
it's a chinese virus, don't download it
>>
File: 1760983495971671.png (21 KB, 507x295)
21 KB PNG
>>109467928
>>
>>109467925
>no reason not to use cr
Well, there is. The non-convrot nvfp4 model is like 8gb smaller.
>>
File: 36646.webm (756 KB, 448x256)
756 KB
756 KB WEBM
>>
>>109467936
its literally the bandoco server
https://discord.gg/4gVXrgsxW
its the "speed up H3" under H3 resources channel
>>
>>109467794
The datasets behind the models are all tagged by LLMs, so its not a bad idea to use one to fluff up your prompt
>>
File: MiniMax_H3_00017_.mp4 (1.69 MB, 480x864)
1.69 MB
1.69 MB MP4
>>
>>109467960
KEK
>>
File: 1646333888461.png (340 KB, 787x720)
340 KB PNG
I just deleted all my wan models and loras.
They served me very well, but it was finally time to let them go.
>>
>>109467966
how do i unsubscribe from this blog
>>
>>109467954
Ok, and what do you recommend for an abliterated/uncensored prompt writer? I generate lewd and I am not tolerating rejections.
>>
>>109467978
gemini 3.5 flash
>>
>>109467943
and the q1 is 3x+ smaller than that, almost like quality below int8 goes to absolute unusuable shit that no speed increase matters
>>
>>109467960
average arch user
>>
>>109467982
>cloud
no
>>
File: 1766577676228938.png (216 KB, 327x316)
216 KB PNG
>>109467892
>>
>>109467983
then why comfy defaults workflow to nvfp4 qwen?
>>
lmao you can literally plug in any image and it will work, kek

<Subject 1> is the dog in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "Okay I wont shock dogs any more!".

https://files.catbox.moe/e927ds.mp4
>>
>>109467983
The question was not related to quants, it was related to convrot.

minimax_h3_fl2va_pruned_nvfp4.safetensors 12.5 GB
minimax_h3_fl2va_pruned_nvfp4_convrot_int8.safetensors 20.1 GB

Same compression, one has convrot and is 8gb bigger.
I'm just saying there technically is a reason to use the non-convrot, but I'm curious what kind of effect it has.
>>
>>109467978
Gemma 4 heretic or uncensored models I guess
>>
File: 6924671.mp4 (3.76 MB, 1184x800)
3.76 MB
3.76 MB MP4
>>
so is stable audio better than acestep ???
>>
can we have separate vid gen and image gen generals?
>>
>>109468027
MoE?
>>
>>109468066
>>109468066
>>
>>109467978
DS4 Flash. Just write an appropriate system prompt, together with the mass of text from the prompt guide that will suffice to make it comply even without a prefill. If it somehow still refuses whatever abominable prompt you give it, just add a prefill. Without closing the think tag so it can still continue thinking.
>>
>>109468002
for h3 the text encoder is huge and 4 bit quants are ok while being small which most people need since they are v/ramlets, in basically every other case ever those are dogshit quants that should never be used, and if you even have some ram, you shouldnt use it even for h3 since int8cr will be quite a lot better
>>
>>109468013
its not the same compression if it says int8, although i dont know why it also says nvfp4, maybe its mixed quants, maybe they forgot to rename properly
>>
File: file.png (1.71 MB, 692x1100)
1.71 MB PNG
>dynamic vram and hip backend works properly with the new comfy-kitchen merge
thank you to everyone who kept pestering comfyanon to take a look at it, i can finally gen my ex having dinner with me
>>
>>109467728
>>109467730
>>109467642
https://huggingface.co/Kijai/PrecompiledWheels/blob/main/sageattention-2.2.0-cp312-cp312-linux_x86_64.whl
>>
>>109467702
How to connect it bro ??
>>
>>109468218
>spoonfeeding
>>
>>109466357
biggest issue is that you can't use it to make an existing video longer without the audio getting cut off afaik.
>>
>format spoken dialogue as instructed in minimax's prompting guide
>english comes out like gibberish
audio isnt very great on this model huh
>>
>>109468360
it is kind of tricky though at least right now to build it on linux, i posted a solution to build from source in the next thread.
>>
>>109468443
I tend to have problems with audio too.
I also think lot of the "free optimizations" people are running might also be playing a role.
>>
>>109468360
And the issue is the devs removed the wheel for linux and now it can't be built from source either due to newer version of gcc or cuda toolkit what ever but there are prebuilt on other sites. I keep it noted because I knew once the new big video models dropped anons would get stuck.
>>
>>109468443
I think audio is one of its best features, but as always, it's not an exact science and prompts, even when done exactly as in the guide, are not reliable. Hell I've had more luck with short, dumb prompts and it nailed the audio right away.
>>
>>109467595
>I used pizza tower as a style reference and this happened
Minimax H3 is starting to depress me, because I can no longer blame the model for not obtaining what I want. Anything that I don't get to see with AI now is a result of my own skill issues
>>
>>109467620
People who are surprised by trumps take on this forget the JD Vance is the only VP who had ever said the words "open source" in an official communication. This is one of the good things about the government being controlled by tech oligarchs
>>
>>109467921
>bf16
There is exactly one consumer card that can run H3 at bf16 and you should still run it in int8 convrot because you save on time and the degradation is unnoticeable, including double pendulum effects, before 10 seconds into the video or so
>>
>>109467620
This only applies to US based models, not everything open source. Likely that there's going to be a ban/restriction on Chinese models.
>>
bros, how many steps are you running H3 for?
>>
>>109467960
I need the prompt for this pls so I can make him follow a girl wearing a skimpy outfit through a city street thanks
>>
>>109467794
Just get it to spit out the prompt and then change what you want. Are you retarded? Why waste time handwriting boilerplate shit?
>>109467978
Probably an abliterated model like https://huggingface.co/huihui-ai/Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated-MTP-GGUF if you really need to do it on your own hardware. Otherwise, you can just fill in the gaps that the LLM shits out.
>>
Can't find KJ's "MiniMax H3 Mem eff Sage attention" repo on github.
pls halp.
>>
>>109469204
It's the kj nodes repo, just copy and replace it in the comy custom nodes folder
>>
What's the strat to using H3 as an image editing model?
>>
>>109469273
make videos 0.1 second long?
but you're right. I remember there was a paper where they compared models and it showed that video models had the best understanding of 3d. They basically became 3d models under the hood to be able to make sense of how objects move in space.
>>
or you could probably make your video 1s long and set the framerate to 1 frame per second and then instead of save video have a save image. I'm just guessing here though.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.