[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 17858867852228672.mp4 (1.1 MB, 480x266)
1.1 MB
1.1 MB MP4
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109461903

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Wan
https://github.com/Wan-Video/Wan2.2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
>>109463498
I admit, it's really funny
>>
>>109463498
>trolling outside of /b/
like the usual you know what to do

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109463516
>I admit, it's really funny
>>
    raise ValueError(f"MiniMax H3 requires {FPS} fps")
ValueError: MiniMax H3 requires 24 fps

really nigga
>>
I admit, it's really unfunny.
>>
I don't mind new spaghetti, posting the same thing over and over was boring
>>
>>>/wsg/6208177
>>
>>109463545
>I don't mind the lack of collage
you're not a true /ldg/ user
>>
>>109463498
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>BREAKING NEWS!
BREAKING NEWS!
>BREAKING NEWS!
BREAKING NEWS!
the 10Eros creator has abandoned LTX and is now working on the H3 lora
https://huggingface.co/TenStrip/LTX2.3-10Eros/discussions/66
>>
>>109463573
duh
>>
>>109463573
he should use the compute to make a distill version
>>
>>109463573
isn't that a violation of the licence or something? that ryan guy will knock at his door :(
>>
>>109463573
uhhh the license doesn't allow this
>>
>schizo wastes his time troll baking
heh
>>
>>109463625
no one else wants to make a bake, so he has no resistance
>>
File: MiniMax_H3_00129.mp4 (855 KB, 1056x608)
855 KB
855 KB MP4
t2v
>>
File: awij.gif (2.1 MB, 319x320)
2.1 MB GIF
>>
havent gend since using wai v14 awhile back, is anima the new replacement for that? whats the go-to for realistic/3d now, krea2? chroma?
>>
>>109463617
easy, just pretend that it has nothing to do with h3
>>
h3 is also secretly an image edit model
>>
>>109463645
>interpolated frames
zoomer or boomer? because millennials dont interpolate and disable smooth motion on any HDTV in their vicinity.
>>
>>109463669
is it good at it though?
>>
>>109463669
how do you set it up to do that
>>
>>109463666
satan sure loves these slop threads
>>
>>109463696
It took a character and made a completely unrelated scene while looking pretty faithful
>>
File: H3_mute__00005.webm (3.76 MB, 1120x880)
3.76 MB
3.76 MB WEBM
Spectrum node blurs the hands a little more or less. But the speed savings is about 1 minute less for 1MP 8sec. I'll take it. Effectively bumping my 0.75MP to 1MP gens for free.
>>
File: 1757183542616185.png (184 KB, 498x377)
184 KB PNG
>>109463707
>>
tired of these negro themed thread. anime thread when?
>>
>>109463625
It's his own personal hell
>>
File: Krea2_turbo_hgdf_00046_.png (1.47 MB, 1664x1216)
1.47 MB PNG
Gonna try running H3 on a 3060 12gb and 24gb of sys memory.
Wish me luck.
>>
>>109463732
>>>109439582
>>
>>109463744
would be good if he wouldn't involve us into his mentally ill introspection though
>>
>>109463718
It's not safe for work when I didn't ask it to be. I might need to use a llm to do the prompt for the i2v model
>>
>>109461396
seems like it was trained on brown dust 2 or grils frontline animations
>>
now I want to root for Flux 3 dev because the minimax faggots won't allow NSFW loras on civitai
>>
>>109463762
>It's not safe for work
>catbox
>>
>>109463573
See the last comment. Hes not abandoning it
>>
Oh shit I've been doing my H3 gens at 20 seconds but I see the native limit is 15, and I just kicked off a 24 second gen. Oh well, going that far past the context limit might be OK or it might go crazy. I guess I am spoiled by the 30-second limit of LTX-2.3. FWIW 20 seconds with H3 has been perfectly fine.
>>
>>109463815
i'm doing 20 second generations and it looks normal to me
>>
for people posting their problems but have the hardware, update your shit, clanker output below

Your old runtime had three major H3 memory optimizations disabled:
Torch 2.5 was too old for DynamicVRAM.
CUDA 12.4 disabled Comfy Kitchen’s native H3 quantized kernels.
H3 consequently loaded approximately 20 GB onto a 24 GB GPU, leaving almost no sampling workspace before the server crashed.
I upgraded the complete matched stack:
Torch 2.10.0+cu130
Triton-Windows 3.6.0.post26
SageAttention 2.2.0+cu130torch2.10
bitsandbytes 0.49.2
Comfy Kitchen 0.2.26
The fresh startup now confirms:
DynamicVRAM support detected and enabled
Comfy Kitchen CUDA backend: disabled: False
Native convrot_w4a4_linear and int8_linear available
LoRA Optimizer: Triton=True
SageAttention and bitsandbytes GPU kernel tests passed
ComfyUI is running at port 8189 with 22.72 GB initially free
This is the configuration Comfy recommends for current CUDA quantization and DynamicVRAM. ComfyUI recommends CUDA 13 PyTorch, and its DynamicVRAM announcement specifically recommends Torch 2.10/CUDA 13. Triton 3.6 is the documented match for Torch 2.10 in Triton-Windows.
>>
>post gen on twitter
>zero engagement
>someone steals it and then posts it
>gets millions of impressions
>they get all the credit while i get nothing
what does this mean?
>>
File: 754742.webm (3.81 MB, 576x320)
3.81 MB
3.81 MB WEBM
>>109463874
>>they get all the credit while i get nothing
now you know how nikola tesla felt
https://files.catbox.moe/vl3eij.txt
>>
>>109463874
You lack vitamin C. "Connections"
>>
File: 338593745.mp4 (3.7 MB, 896x480)
3.7 MB
3.7 MB MP4
>>109463492
Basically, yeah.
>>109463493
I tried easy cache, but found the quality degradation a bit much.
>>
>>109463929
hnngggg
>>
File: sussy baka.png (36 KB, 220x244)
36 KB PNG
>>109463874
>steals
anon, AI models work by stealing the whole internet, don't try to virtue signal you can't win that game
>>
File: MiniMax_H3_00024_.mp4 (1.37 MB, 544x800)
1.37 MB
1.37 MB MP4
>>109463492
Why am I getting deja vu? Did someone else prompt for this exact thing or have our timelines shifted and I'm the one prompting this?

https://files.catbox.moe/m2vn1n.mp4
https://files.catbox.moe/x4wkat.mp4
>>
File: 7547422.webm (2.9 MB, 320x576)
2.9 MB
2.9 MB WEBM
ok time to figure out a new prompt. i'm thinking tank battle
>>
File: videoframe_8782.png (1.47 MB, 1920x544)
1.47 MB PNG
Nothing / Sol attention
656sec / 477sec
https://files.catbox.moe/zvyjjb.mp4
Don't think Sol Attention is worth it.
>>
>>109463874
he has a platform and you don't
>>
>>109463532
just patch out that check and see what happens.
someone also said the fps was hardcoded in a math node in the default workflow, should change it there too.
>>
>>109463982
So many vramlets are so dense its unbelievable, they want fast gen times and 4k quality, you can't have it both ways
>>
>>109463929
Soundtrack
https://youtu.be/go6-MliYCAs?si=44HS-nC4eT249Lau&t=22
>>
>have a "platform"
>instant popularity on every post no matter quality

>don't have a "platform"
>every place either your post will have zero people see it, or some fat reddit mod will remove it
>>
>>109463874
you gotta watermark your content anon, you can even add text in your videos with h3, its really easy now
>>
>>109463994
>you can't have it both ways
yes you can, one day we'll get 6b models as good as seedance, there's still a lot to improve, this technology is still very new
https://explorative-modeling.github.io/
>>
File: 1750448866791591.jpg (115 KB, 800x868)
115 KB JPG
>>109463874
>be me
>scrolling through twitter
>see a nice gen posted by some loser
>retweet it
>gets billions of engagements
>profit$$$
>>
>>109463994
>Noooo! You're not allowed to run optimizations!
>You must run everything at bf16!!!
>>
It's interesting that you can use multiple images for a single subject.
><Subject 1> is the woman whose face comes from <Picture 1> Her body shape comes from <Picture 2>. She is wearing the outfit from <Picture 3>.

It actually worked really well. Originally I'd been trying to use the outfit as a subject and reference that in the prompt but this is more straightforward and reliable. Cropping close on the references helps a lot, especially with the face.
>>
>>109463874
add a second save video node, one for social media uploads with a watermark png + tiktok styled outro slapped on :^)
>>
>>109463874
copyright claim him and steal all his revenue.
>>
>>109464010
>>109464021
I'm not saying that you can't run optimizations, it just baffles me how many anons are crying about quality while they generate videos with 4 attention patches, and 20 steps, I mean how dumb are you thinking you can get HD quality with 20 steps
>>
>>
>>109463827
I noticed it kind of did nothing motion-wise during the first few seconds, so it's definitely running out of motion vectors at 24 seconds. 20 seems to be fine, I agree.
>>
>>109464041
ngl I won't complain too much, we already have sageattention that's a 30% increase, and using int8 convrot is a x2 speed while having the quality of Q8, I hope there's more optimizations like that that exists but for the moment, the remaining optimizations are memes who tank the quality of the video
>>
How much worse are the pruned weights to the not-pruned ones?
>>
needs some refinement on the prompt but we are getting there...

https://files.catbox.moe/ghap4c.mp4
>>
>>109464041
>it just baffles me how many anons are crying about quality
Was he crying? he literally said it wasn't worth it.
>20 steps
It's a 20 step model retard?
>>
>>109464063
the quality is at the same level, it's like having another seed
>>
>>109464066
>we are getting there...
nah, this model is DOA >>109463782
>>
>>109464056
did you add more things to your prompt when you extended the time? you might need to pad it out with filler actions to make the model have more stuff to distribute across the timeline. also, i make sure to prepend "at 00:00:00" before the motion section of the prompt so the scene composition stuff doesn't get used as motion data
>>
>>109464071
>Was he crying? he literally said it wasn't worth it.
learn to read idiot
>It's a 20 step model retard?
20 steps is the bare minimum to get decent quality, there is a even a big difference between 25 and 20 steps
>>
are there any defaults like changing from simple scheduler and res multistep, 20 steps that people have noticed are upgrades
>>
File: MiniMax_H3_00025_.mp4 (2.43 MB, 800x800)
2.43 MB
2.43 MB MP4
>>109463976
Damn, that's insane I2V
https://files.catbox.moe/20parh.mp4
>>
>>109464094
so far this beats res multi for me.
>>
>>109464066
I am shocked we got a model that can do this out of the box, how did this even pass safety filters
>>
>>109463976
>>109464105
don't hesitate to share your gens here too >>>/wsg/6208172
>>
>>109464115
>safety filters
lol. lmao, even.
>>
>fresh comfy install
fuck, I love not having 20 fucking node conflicts every time I start this piece of shit.
>>
File: SODIMMs.jpg (1 MB, 1894x1633)
1 MB JPG
$500 CADollarydoos for both
memtest showed it as okay so I think I did well considering current retail for them is about 2x that price not including tax.
>>
I won't be sleeping for the next few days.
>https://files.catbox.moe/1dhruw.mp4
>>
File: h3_sampler_speed.jpg (165 KB, 996x502)
165 KB JPG
>>109464111
Not for me. dpmmp_2m_sde is longer to gen and also broke the sound on Ref model.
>>
>>109464140
Why do these DIMMs look like they're from 1996?
>>
File: 9.png (2.69 MB, 1086x1448)
2.69 MB PNG
>>
File: thank you chinks.png (325 KB, 979x1107)
325 KB PNG
>>109464163
same
>>
>>109464166
same here, tried a few different combos and none worked as good as the default one.
>>
File: MiniMax_H3_00004.webm (3.14 MB, 864x480)
3.14 MB
3.14 MB WEBM
Got H3 running on a 3060 12gb and 24gb of ram.
This only took 6mins, but I think the source image hurt the quality and I have only just started figuring out the prompts.
>>
File: 1765560096177339.png (85 KB, 1075x920)
85 KB PNG
>>
File: SODIMMs.jpg (1.17 MB, 2003x1692)
1.17 MB JPG
>>109464169
Bought them off someone who was upgrading 2 laptops (I'm guessing with 64GB sticks instead).
>>
what does upping the steps do? It just seems to add more time
>>
>>109464163
this shit is so stupid
>>
>>109464196
begone nufag!
>>
>>109464115
This aint your pussy anthropic or openai model
>>
>>109464196
h3 brought so many retards is unbelievable
>>
>>109464166
ah, haven't tried the ref model yet. The sound is very good for me.
>>
>>109464174
I keep forgetting people also post images in here.
>>
i dont get it. some videos the model refuses to do character swaps with.doesnt matter how i prompt it. the character is clearly visible so i dont know why its not working.
>>
File: 2219745421.mp4 (3.62 MB, 640x640)
3.62 MB
3.62 MB MP4
>>
huh? spectrum only works with these?
I hate vibeslop.
>>
File: H3_webm__00005_.webm (1.28 MB, 1424x496)
1.28 MB
1.28 MB WEBM
>>109464242
You need to use official prompt for Ref. It can get complicated. Also, it works on char swap for hentai and porn.

Scroll down to bottom of page for official example.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
File: 1391607963618.jpg (24 KB, 387x461)
24 KB JPG
>>109464283
>allowlisted
>>
>>109464297
>works on char swap for hentai and porn
proof??
>>
>>109464316
using my allowlisted sampler to make the most racist img2vid edits you have ever seen
>>
>>109464215
>is unbelievable
esl detected
>>
Is it just me or does H3 obliterate faces at lower resolutions?
>>
>>109464331
kino
>>
File: 18.gif (3 KB, 67x72)
3 KB GIF
Does anyone know of a good tool or workflow for pixel art?
>>
>>109464341
obviously, models aren't good if there's not enough pixels to work with
>>
>>109464341
I use 0.5MP and it works fine. Did you change any other params?
>>
>>109464200
4 u
>>
>>109464242
yep thats my experience as well. random, pure luck. depend on reference video
>>
>>109464324
I may be a coomunist regarding AI prompting, but I don't help other men cum, faggot.
>>
>>109464361
You need to read
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
>>109464341
>for best results: >0.4MP + 1:1 or 16:9
>>
>it's because all his gens are cunny
>>
>>109464368
>>109464297
Yes, I've done that, even copy-pasted your exact prompt from one of the videos you posted.

I'm telling you; It does NOT work on some reference videos not matter how well you prompt it. It must have some safety feature baked in that's blocking it. I can't think of any other reason. There's no way it's failing to recognize the subject in some of my reference videos.
>>
>>109464355
I'm taking a high res image, scaling and cropping it down to the same size as the video. but the first frame of the video looks so much worse than the scaled and cropped image.
>>
>>109464361
I wonder if higher res ref videos make a difference. i'll try that approach.
>>
>>109464394
that's normal. it gets encoded and then you see the decoded version with all its compression artifacts
>>
>>109464390
Alright? How about you post the ref video for anons to have a crack at it if we feel like it.
>>
>>109464368
i tried "<Video 1> is the source video for the editing task.
Use the character from <Picture 1>, entirely replacing the original character in <Video 1>"
it doesn't work most of the time, i've tried like 5 porn clips, it just generate the same porn clip with slight change
>>
>coomers too dumb to prompt
>>
File: 1629080241563.jpg (49 KB, 954x694)
49 KB JPG
>>109464409
no

>>109464414
this x1000. feels completely random which videos it works for.

subject_definitions:
<Subject 1> is the character from <Picture 1>.
<Video 1> is the source video for the target video edit.
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.

summary:
[video editing + reference generation + audio reuse] The target video is an edited version of <Video 1>. The female in the original video is completely replaced by <Subject 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): attribute_transfer - All physical features from <Picture 1> is transferred to the female in the target video.
<Video 1> (cut and pacing structure): fully_preserved - the temporal structure and framing are kept, but the female is completely substituted.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track, including all vocalizations and sound effects.

detailed_description:
[Shot 1] The shot establishes <Subject 1> lying on her stomach.
>>
>>109464454
AI was supposed to make programming obsolete. but now we have to start programming again?
>>
>>109464478
No one can magically read your mind, you have to give instruction in some kind of form.
>>
Does anything about MoE reduce the required amount of VRAM needed to run models?
>>
>>109464478
Autists must learn to communicate in some form or another
>>
File: director-node.png (83 KB, 858x786)
83 KB PNG
>>109464478
That's not programming. If you're too lazy to setup a scene, just using H3 Director.

https://github.com/seesee75-commits/ComfyUI-MiniMaxH3-Director
>>
>>109464486
It just runs way faster ram lil especially for sys RAM bound bitches
>>
>>109464497
That's what I thought. I don't understand where all the hype/fear over Kimi comes from when it doesn't really reduce RAM demand.
>>
>>109464505
google and alibaba are ahead of the curve focusing on consumer hardware
>>
>>109464414
after the 8th attempt it finally did the swap, but when the sampler finished, it fucking error'd out because the ref video didnt have audio so the audio decode node didnt save the file.

i was absolutely pissed. i asked claude if it was possible the unsaved latent still existed somewhere, and it fucking found it in my ram and saved the output.

god damn i love AI
>>
File: anime.webm (2.7 MB, 1365x2048)
2.7 MB
2.7 MB WEBM
>>
>>109464528
>it fucking found it in my ram
wth
>>
File: mask__00050_.png (766 KB, 1728x1344)
766 KB PNG
>>109464454
Try bg removal node to completely isolate the subject to designate it for replacement.

https://github.com/1038lab/ComfyUI-RMBG
>>
>>109464529
4steps ?
>>
she looks like my microwaved sausage that exploded near one end
>>
>>109464529
try this.
>>
>>109464529
I look like this irl
>>
https://files.catbox.moe/5mp9ml.mp4

H3 is almost too eager to do h-scenes. I prompted for her to take off her blouse at the end, but instead everything came off. At least it does kissing, licking, and sucking properly. Can't say that about LTX-2.3.
Bye, LTX-2.3. You have been replaced.
>>
File: screenshot.1785898242.jpg (115 KB, 808x366)
115 KB JPG
>>109464533
yes, im not kidding.
>>
>i look like a 4 minute microwaved sausage
the state of these anons
>>
File: p56.jpg (86 KB, 405x720)
86 KB JPG
>>109464545
i thought claude was a meme
>>
>>109464545
How can you stand they way it writes?
>>
https://files.catbox.moe/k13wdi.mp4
> <Video 1> is the source video for the editing task.
Use the character from <Picture 1>, entirely replacing the original character in <Video 1>
somehow it combine picture 1 and 2, but still the character still does random thing
>>
>>109464554
are you joking? claude has been the best at coding for more than a year
>>
File: dmc.webm (2.7 MB, 1365x2048)
2.7 MB
2.7 MB WEBM
>>109464529
>>
Absolute kino >>>/wsg/6208249
>>
>>109464561
Buy an ad, Dario.
>>
File: peach.webm (2.68 MB, 2048x2048)
2.68 MB
2.68 MB WEBM
>>109464529
>>109464563
>>
>>109464561
i got bored of making projects a long time ago so i never bothered to look into this stuff. they also cost money so i don't want to give them anything
>>
>>109464554
Absolutely not. Buying a pro Claude subscription was probably the best purchase I've made this year. It's like a $300k salary senior dev as your co-worker. There's no way these models are going to remain this cheap once the masses get on it.

>>109464559
? doesn't bother me. you can modify the way it speaks if you want
>>
>>109464560
>mfw
>>
File: i have autism.jpg (60 KB, 543x642)
60 KB JPG
how do i double keksampler refine
>>
>>109464571
>$300k salary senior dev as your co-worker
oh no no no.....
>>
File: 1772858338104586.png (318 KB, 2472x1269)
318 KB PNG
>>109464478
>have to start programming again?
No, just ask Claude to make you a local prompt enhancer
https://files.catbox.moe/lw0x8n.mp4
>>
>>109464575
Just connect latent to another Ksampler. Picture is NOT for H3.
>>
>>109464544
>hand came off
AAAIIIEEEEEEEEEEE
>>
>>109464580
what is actually enhancing the prompt, h3's own text encoder?
>>
Uh-oh! Disney lawsuit incoming! https://litter.catbox.moe/ykvkfy4cr3xavqf0.mp4
>>
>>109464592
screenshot shows qwen3-vl 8b
>>
Anyone brave enough to gen Minimax at 8 steps or even 4 steps ? Hows the results ?
>>
>>109464580
nice. gonna give this screenshot to claude and tell him to copy it ;)
>>
File: 1769981142564871.png (365 KB, 760x1195)
365 KB PNG
>>109464599
>Disney lawsuit incoming
it's already a thing kek
>>
>>109464603
garbage
what did you expect
>>
>>109464603
anon, this is not a distilled model. anything below 20 steps is going to look like shit.
>>
>>109464609
Use 8/4 steps lora duh
>>
>>109464615
you skipped the prerequisite of making one first
>>
>>109464608
wtf are they gonna do tho? it's a Chinese company.
>>
>>109464620
Ill give it 3 days
>>
>>109464622
The lightx2 lora took like a month before it was decent on wan 2.2.
>>
>>109464622
I'm saying nothing until it happens.
buuut. cutting down the gen times to 1/5th would shift to testing the limits of ones own hardware instead of their patience, and I'd very much like for that to happen
>>
>>109463885
unironically cool, the flying is really fun to watch

if you'd like my advice though, make that he lands in oblivion and now he has to kill elves and stuff
>>
File: MiniMax_H3_00030_.mp4 (2.1 MB, 832x1248)
2.1 MB
2.1 MB MP4
>>109464105
Genning at 1MP seems to help/reduce artifacts that occur with lower resolutions on some more complicated gens with more movement
https://files.catbox.moe/mfcqfe.mp4
>>
>>109464626
it was never decent, it made outputs stretched and slow
>>
File: 11111.mp4 (2.02 MB, 1056x608)
2.02 MB
2.02 MB MP4
https://files.catbox.moe/svh1zf.mp4
>>
File: 644354.webm (421 KB, 448x256)
421 KB
421 KB WEBM
>>109464634
>if you'd like my advice though, make that he lands in oblivion and now he has to kill elves and stuff
i can go back to that prompt later and do some cool stuff with the stranded pilot. working on some new kinos. ltx struggled with this kind of thing
>>
>>109464644
I like these.
>>
File: MiniMax_H3_00031_.mp4 (1.61 MB, 928x672)
1.61 MB
1.61 MB MP4
>>109464636
Pure kino at 4:3
https://files.catbox.moe/7s3dli.mp4
>>
File: d04fih.jpg (121 KB, 736x736)
121 KB JPG
https://files.catbox.moe/d04fih.mp4
>>
>>109464653
4:3 was the analog standard so it's good to use that aspect ratio if you want vintage kinos
>>
>>109464224
me wondering you say this to clown the person for posting a image and not a video or because the thread is 90% schizobabble
>>
is wan2gp faster than comfy for h3? Anyone done any comparisons yet?
>>
Finally got everything set up to spit out the quality I was looking for. 22 minutes for this 10 second gen at 0.8 mp upscaled with RTX-SR on a 5090.

What can I say except, you're welcome.
https://files.catbox.moe/2o7xlw.mp4
>>
File: HKOvk5MXgAAiCqV.jpg (293 KB, 1773x2000)
293 KB JPG
>>109464694
>>
>>109464694
I doubt it can be faster, however certainly for anyone wanting to use it without having to mess around updating and all the different nodes then you can certainly end up sitting for a long time with comfy thinking what the hell is wrong. Wan2GP pretty much will work straightaway. But I sure people have tweaked the settings and got the correct nodes to get it running faster on comfy
>>
>>109464694
probably not. wan2gp doesn't have support for the new optimizations yet. it takes about 5 minutes to finish a 20 second generation at 320p if you want to compare that
>>
What's the current optimization meta? Sol attention, sigma shift and spectrum?
>>
>>109464733
only spectrum gives you an optimization that won't kill the model's quality
>>
>>109464742
As does genning below 2mp. Everything is a trade off obviously, people are willing to give up quality for speed
>>
>>109464725
fuck them elves
https://files.catbox.moe/jt14o3.mp4
>>
>>109464751
>people are willing to give up quality for speed
of course, but if the quality gets down so much you're back to LTX level there's no point in doing that
>>
File: thank you china.png (766 KB, 941x627)
766 KB PNG
>>109464759
holy shit this model is so uncensored
>>
>>109464765
the greater knowledge makes it worth doing if you are making t2v stuff
>>
>>109464730
>you can certainly end up sitting for a long time with comfy thinking what the hell is wrong
That was what lead me to use Wan2GP in the first place for LTX2.3, but clankers are better now and they can hold my hand through it
>>109464731
>it takes about 5 minutes to finish a 20 second generation at 320p
on what GPU? I'm on a 3090 Ti and I don't think I could even generate something that long with comfy
>>
>>109464796
NVIDIA GeForce RTX 3080 Ti
>>
>>109464796
>but clankers are better now and they can hold my hand through it
wow, you can waste your time and money on tokens as well as comfy headaches instead of software that just works!
>>
>>109464765
people can run prompts at lower quality to try them out. The counter argument why run at higher quality and wait far longer and the gen be shit anyway? At least you can run again at higher settings pretty safe the prompt wasn't a dud. Also for posting on 4chins then why go for something high quality when it probably wont' even be archived? I have to say for all the talk of people running gens fast on their 5090's with 120gb of ram there isn't a lot of posting. Almost as if minimax lasted a day, but I am sure there must be a reason for the lack of posting, maybe they don't want their masterpieces stolen?
>>
>>109464820
>there must be a reason for the lack of posting
no one posts on a will smith bake, that's the golden rule
>>
>>109464810
If you're doing 20 second clips on 12GB vram, that's incredible.
>>109464819
Well, that's why I'm asking.
>>
>>109464831
>If you're doing 20 second clips on 12GB vram, that's incredible.
yeah but i still desire a distilled version of this model. ltx was letting me do 20 second 720p videos on this card. 320p is the limit right now for this model
>>
>>109464831
>If you're doing 20 second clips on 12GB vram, that's incredible.
You can, it just takes a lot of time and by a lot of time I mean 45 minutes for 0.5MP@24fps
>>
File: 638575472.mp4 (3.65 MB, 1504x640)
3.65 MB
3.65 MB MP4
>>
so whats the current go-to prompt to feed into a slop model for h3
yes im that lazy, no i dont have to be
>>
>>109464826
Well I was referring to /wsg/ or /gif/. I so called seismic shift in genning and pretty much just a few extra gens on /aicg/? I would have thought it would be flooded with more than just the Seinfeld ones. And just to think when HailuoAI released their video genner maybe 2 years ago there were gens all over the place and like the Dall-E threads on /aco/. There actually is a significant lack of interest in AI video and music genning, the audio element is a big factor as only /wsg/ and /gif/ and a few others allow audio
>>
>>109464759
i'm honestly shocked by how much it knows and how well it follows your prompt
you can literally throw anything at it
>>
File: xx.mp4 (3.05 MB, 1010x1128)
3.05 MB
3.05 MB MP4
https://huggingface.co/Kijai/MiniMax-H3-TAE
>>
>>109464885
>you can literally throw anything at it
that's what happens when you give a model a good dataset and not some corporate slopped shit, based chinks
>>
>>109464866
The novelty wore off years ago. Nobody is actually that interested in proompting videos that are still uncanny and subpar compared to actual art.
>>
im having fun.

https://files.catbox.moe/1l0y0k.mp4
>>
>>109464892
>Nobody is actually that interested
>>109464895
>im having fun.
the duality of an anon
>>
>>109464895
>suddenly, GTA5 highway
>>
>proompting
>>
>>109464841
I don't see the point of gens longer than 10 secs desu. More than enough info in 10 secs and you can do that comfortably on 3090. Don't forget we're still pushing 30 series cards for several years now, they've long been outdated so if gen times are too long or you really want 20 secs then just upgrade>>109464844
>>
if you get distorted audio using something other than res_multi it's because you need to play with the sigma shift.

I had mine at 8/6 and I didn't have any audio issues.
>>
>>109464895
Me 2 anon, me 2.

>>>/wsg/6208291
>>
>>109464853
do a flip
>>
H3 video reference takes so fucking long to gen
>>
How many times have you gooned to h3 so far since release? for me it's 5
>>
>>109464885
It's insane. It just follows your prompt, no fighting back, no struggling, no refusals. It just does what you tell it, which is so nice for a change
>>109464888
>trips confirm
>Xi Jiping is based
>>
>>109464904
that was my first thought too. what makes it so gta5?
>>
>>109464919
Zero.
>>
>tfw went from 60s/it down to 12s/it

size of the input ref video actually matters. if you're genning at 0.4mp just resize the reference to <= 0.4mp and it flies
>>
>>109464919
0. It's the only model ever released where I'm having fun not gooning
>>
>>109464892
In a zoom zoom world you would think the 5-20 sec clips would be perfect as they scroll from one to another. Then again they can watch some fag on Twitch play some slop for hours so hey ADHD has ups and downs. But you might be correct in that if anything has "AI" attached to it then it get's a massive thumbs down on even bothering to watch because they know it was reasonably easy to gen. As for "art" most of the crap on 4chins isn't art, it's real life slop.
>>
File: Meballs.jpg (47 KB, 996x282)
47 KB JPG
>>109464919
>current state of my balls
>>
>>109464923
>It's insane. It just follows your prompt, no fighting back, no struggling, no refusals. It just does what you tell it, which is so nice for a change
I'm still shocked they gave us peasants a top tier model. This is something you keep locked away behind a paywall. This will sustain us for years to come.
>>
>>109464826
>will smith thread
>intelligent discussions about optimizing local diffusion generations
>link spam threads
>tranny meltdowns and BBC spam
i like this one better
>>
>>109464945
>>109463530
>>
>>109464945
>please let the schizo win or else he'll do meltdowns and he'll bully us
you're just giving up anon, congrats on being a pussy who takes the knee I guess
>>
>i dont like the thread when i shit it up
what could be the solution here?
>>
>>109464907
>I don't see the point of gens longer than 10 secs desu
you don't have to worry about losing context if you try doing clip extensions. it depends on your usecase i suppose
>>
>>109464954
/adt/ anons just kept ignoring him until he left back to /ldg/
>>
>>109464935
Oh if you want zoomzoom gens just go on TikTok and search for backrooms
>>
>>109464954
looks like the schizo did win to me. nobodys even bothering to make real threads to challenge the trollbakes
>>
can you use image references on fl2va model?
>>
Why is the troll baker crying in his own thread?
IDGI
>>
>>109464977
i haven't tried it but someone here claimed to be doing it. i am guessing he puts in the starting frame and prompts it to immediately change to a different scene. that's exactly how i was mimicking reference images in ltx
>>
>>109464974
>the schizo
why are you talking about yourself in the third person though?
>>
>>109464892
The thing is models are getting good enough where it's just becoming a tool instead of a standalone novelty.
>>
>>109464940
They know Flux 3 will be better than their API option and needed a reason to stay relevant. It was panic decision likely by the higher ups at their company. What's great is that now the guys who made Kling can go fuck themselves.
>>
>>109464966
I am not a zoomer but I will admit my first gen using H3 was a Backrooms one, my excuse was I picked an image in my downloads and thought it would be funny with Tony Soprano shooting up the walls and it collapsing. And if this was maybe a few weeks ago I might have genned something to do with Nolan and The Odyssey.
>>
>>109464974
35 stars status?
>>
>>109464985
Yeah, the cutting edge models. But H3 is not Seedance 2.5. H3 is a toy that you use to make Seinfeld edits with that they released publicly because nobody would have used it otherwise.
>>
>>109464981
That, and you can even use multiple characters in one image that way, and tell it something like char on left is X, char on right is Y, and use them in your prompt afterward.
>>
https://files.catbox.moe/as1cwo.mp4
>>
her mouth dry af
>>
>>109465005
good one I love making these too
>>
>>109465000
nice FUD
>>
>>109464986
LOL, Eurocucks can't bake for shit
>>
>>109465005
how did you have it not mangle the dick?
>>
>>109464986
>They know Flux 3 will be better than their API option and needed a reason to stay relevant. It was panic decision likely by the higher ups at their company.
nice fanfiction BFL employee
>>
>>109465015
my workflow should be in the file?
>>
File: grifterss.png (260 KB, 1513x1247)
260 KB PNG
https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot/discussions/7

AI Jeets will ruin the fun for the rest of us because of their excessive snakeoil
>>
>get the clanker to gen a prompt for me based on my feelslop
>don't even read it
>paste
>run
>>
>>109465025
isn't this just snake oil
>>
>>109465025
>he's even going to nuke snake oils
ITS FUCKING OVER
>>
>>109465025
Back in the day, license fags would at least have the decency to post with a sock puppet
>>109465031
Yeah
>>
>>109465023
My apologies
>>
>>109465025
why the fuck is he trying to use his licence on a qwen model?? he doesn't own Alibaba wtf??
>>
https://files.catbox.moe/8cu6xa.mp4
Except for res_multistep, every other sampler causes some weird audio glitches. Not sure if this was intended or just a comfy issue.
>>
>>109465040
parfect for gorgous looks ser
>>
>>109465011
Yeah it's pretty next level to be able to take a picture of your dick and have a woman do whatever you want to it.
>>
>>109465047
>fucked up sound
you used EasyCache right?
>>
>>109465023
Where's the mystery lora from?
>>
>subgraphs no longer show previews
COMFY
>>
>>109465013
>>109465022
They actually cooked and are on the verge of open sourcing it. It may be cucked, but even cucked at the level they're giving us would've been significantly better than a gated H3, and even with their own neighbors (Seedance) giving them intense competition, there wasn't much else they could do but release those open weights, and they knew it.
>>
>>109465062
>EasyCache
Nope, only a retard would use that.
>>
>>109465047
Ok. Here's the real fix for audio glitches. you gotta bump the audio sigma shift. the default 3 is too low. bump that shit to 6.
>>
if I use a floorplan as a reference will the model just figure it out
>>
>>109465076
kek
>>
EasyCache doesn't even trigger on dpmpp btw.
>>
>>109465068
It got deleted. Let me upload it.
Its uploading real slow. I'll post it in a moment.
>>
>>109465089
Very based of you.
>>
>>109465087
real
the only time it worked for me it was on res multistep and on shit where the video was basically still
>>
>>109465092
https://gofile.io/d/8RUWhT
so the original creator said to use this at 0.5 strength due to it being trained on distilled. also I'm pretty sure it only works on the pruned version. minimax_h3_fl2va_pruned_int8_convrot
>>
>tfw I only got room for 6 gens in an hour.
>>
>>109465108
and there's no trigger words? or recommended prompt?
>>
>>109465120
no he didn't say anything about that. I think its just trained on porn in general. which is why I guess it got deleted off of huggingface.
>>
>>109464528
So much of this kind of stuff was too annoying to bother with a year ago and now you just tell AI what you want. When I was having problems with comfyui dynamic vram slop I just let AI fuck around with it until it reported back that the gen times were fixed.
>>
Save Video doesn't let you set the bitrate and I found out the video quality is horrible after I switched. That's ridiculous. Users will be waiting 5 minutes to an hour for a video that's a few megabytes either way and it looks like shit
>>
https://www.reddit.com/r/StableDiffusion/comments/1vfwijz/minimax_are_issuing_takedowns_on_decensorexplicit/
That's why I want BFL to win
>>
>>109465129
how is that comfortable in the long run?
>>
>>109465025
I don't get it, doesn't the model itself already violate the community license agreement? It's already uncensored, what could that poster be trying to hide?
>>
>>109465132
use "Save Video" from that node
https://github.com/kosinkadink/ComfyUI-VideoHelperSuite
>>
>>109465135
nigga they will release a safety cucked model on top of issuing takedowns
>>
>>109463498
>>>/gif/30996965
I told H3 to animate a single image. This is just 0.2 MP so quality isn't the greatest.
>>
>>109465146
>>109465146
>>
File: 1759865015371532.png (921 KB, 1080x1080)
921 KB PNG
>>109465140
they have the right to train their model on porn, but not you, make it make sense
>>
>>109465123
Thanks for sharing. will report back on if it improved my gens.
>>
>>109465140
would be funny as fuck if the company gave us the wrong model and we were meant to get a censored one instead of their API model
>>
>>109465149
it's probably a farce to justify releasing such an uncensored model
>look, we totally care about safety!
>>
>>109465132
or use this, crf value is quality value
>>
>>109465172
How?
>>
>>109465138
wdym exactly? Because it just works. I have no software issues anymore, I just use AI. There's nothing more comfy than that.
>>
>>109465182
Ask chatgpt how to git clone video helper suite.
>>
File: 197953898.mp4 (2.95 MB, 1504x640)
2.95 MB
2.95 MB MP4
>>109464914
It's not very good with flips.
>>
>>109463530
i thought the blade runner one was funny
>>
>gen shit
>`no ice lool` text in one of the corners in the gen
>no such thing in the prompt anywhere
the aliens are communicating
>>
File: file.png (2.04 MB, 1600x900)
2.04 MB PNG
>>109464904
>>109464924
I'm guessing it's trained on a lot of GTA RP streams which results in it looking like the los santos freeway
>>
>>109465015
>how did you have it not mangle the dick?
>dick is quite clearly mangled
>>
>>109465149
I mean cause thee is not the people who are gonna get their ass sued in this case. It's just covering their ass and saying they don't allow it. Even his tweet said to keep it under the radar if you do make that stuff.
>>
File: file.png (461 KB, 450x806)
461 KB PNG



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.