[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1786242760918661.webm (3.96 MB, 960x784)
3.96 MB
3.96 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109502333

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Bleeds thread
>>
minimax h3 video of the year award goes to #8
>>
make some sleepy gens
>>
Blessed Bake of Teamwork.
>>
real thread here
>>109503609
>>109503609
>>109503609
>>
Will minimax even release their super secret magical upscaler that would solve the shit faces we have right now for H3?
>>
>>109503683
xD
>>
>>109503684
If we get stood up again like the chinese tend to i'm gonna declare jihad on the 100 acre wood
>>
Why does debo the idiot keep trying? Even though he keeps failing?
>>
>>109503683
haha
no
>>
>>109503703
its the pokemon spammer
>>
>>109503683
You fail at this every fucking day. Next to self soothe your autism you'll spam news
>>109503703
He has nothing else to live for
>>
File: output.mp4 (2.49 MB, 1056x608)
2.49 MB
2.49 MB MP4
>>109503584
>>109503551
>>
>>109503706
It's not working debo. You're not smart enough to change your posting style.
>>
File: attention ldg.mp4 (1.84 MB, 960x544)
1.84 MB
1.84 MB MP4
>>109503683
>>
>>109503692
I wouldn't be surprised, it wouldn't be the first time just the last little thing is withheld.
>>
>>109503703
it's not debo, it's whoever decided to shit on the thread this day
>>
>>109503703
It really would be better if he just died. Just committed suicide. The world will be a better place when he finally dies.
>>
>>109503719
I want him to live a long life but lose access to the internet.
>>
Any way to make H3 work in other aspect ratios without simply genning in 16:9 then chopping the video up?
>>
Animabros...
>>
File: brugh.png (64 KB, 663x775)
64 KB PNG
>>109503738
>>
File: H3_resolution_master.jpg (107 KB, 667x1083)
107 KB JPG
>>109503738
I am using custom resolution just fine.
>>
>>109503743
What in the shit is that UI horror
>>
>>109503743
>making 1girls in ideogram
>>
>>109503741
>>109503743
Thanks, I guess my early tests were just bad luck. Figured it was something forced by the training.
>>
File: 1771285338479684.mp4 (1.22 MB, 1350x1086)
1.22 MB
1.22 MB MP4
>>109503738
>>
File: seitjfeswioutfjgesw.png (719 KB, 956x525)
719 KB PNG
>>109503708
>no quality degredation
>sd1.5 lookinass
>>
>>109503761
>smear frame posting
back to /a/
>>
>>109503761
Eh, I'll take the occasional mediocre in-between frame if it cuts my runtime in half.
>>
praise be to op
>>
>>109503739
So anima is a meme then?
>>
thanks for bakering
>>
debo is the god of ldg
>>
>>109503793
benchod
>>
The latest on /vp/:

>>>/vp/59492200
>>>/vp/59492256
>>
>>109503793
Always funny seeing you pathetic morons trying to test your VPN
>>
>>109503752
nta but Ideogram is still unmatched in photorealism
>>
Hello?
>>
>mfw Resource news

08/08/2026

>Kijai: MiniMax H3 Ref Lora Rank 256 bf16
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)
https://github.com/Amduraznak/minimax-h3-fp16-fix

>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopoints
https://github.com/fengwang/Cosmos3-Nano-WebUI

>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patches
https://github.com/charlie12345/R9700AIProComfyUIPatch

>MiniMax-H3-Pruned-GGUF
https://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient VidGen
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>>
>>109503813
among other, not so favourable things
>>
>mfw Research news

08/08/2026

>Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
https://arxiv.org/abs/2608.04902

>Coherence-Oriented Dream Scene Visualisation
https://arxiv.org/abs/2608.05233

>Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
https://arxiv.org/abs/2608.05210

>GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
https://arxiv.org/abs/2608.03083

>IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Images
https://arxiv.org/abs/2608.03539

>Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
https://arxiv.org/abs/2608.00663

>A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
https://arxiv.org/abs/2608.05260

>Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
https://zhangbo135.github.io/EviSelect

>Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries
https://arxiv.org/abs/2607.29156

>GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration
https://arxiv.org/abs/2608.03923

>WorldClaw: Agentic 3D Open-World Generation at Scale
https://arxiv.org/abs/2608.05248

>UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
https://arxiv.org/abs/2608.03817

>Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
https://arxiv.org/abs/2608.03867

>Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models
https://arxiv.org/abs/2608.03160

>Attention is Case-Sensitive
https://arxiv.org/abs/2608.03711

>In-Context Collapse in Vision-Language Models and How to Mitigate it?
https://arxiv.org/abs/2608.02830
>>
>gen looks good
>"oh but what about this added detail"
at some point the prompt will just overloaded and stuff will start to fall out, but we haven't gotten there yet
>>
>>109503820
I hope the minimax imagegen model can bring the best of all worlds (non-autistic prompting, H3's uncensoredness, Krea 2 knowledge of pop culture in terms of characters and celebrities, Ideogram's photorealism, Flux 2 editing)
>>
>>109503839
>Krea 2 knowledge of pop culture
the video model seems way worse than krea2 when it comes to knowledge and characters
>>
>>109503708
miyako... my love
>>
File: debo_tt_k2_00082_.png (2.61 MB, 1872x1007)
2.61 MB PNG
>>109503792
true
>>
>>109503819
>>109503827
thanks!
>>
>>109503835
why did you put two quotation marks next to each other?
>>
>>109503846
References work even with the text model you can do a style transfer too and it will know the character or pose, it's not as good as the other model but for my task it's been fine.
>>109503858
You're a waste of space and are gradually being gate kept out of this hobby, once next gen hits you'll be priced out.
>>
>>109503863
autist anon...
The top one is their impression of the gen, the bottom is them saying that to themselves.
Not that anon by the way
>>
>>109503864
yeah the reference is quite nice. if the image model can take references its not too bad then
>>
File: debo_tt_k2_00085_.png (1.78 MB, 1872x1007)
1.78 MB PNG
>>109503860
:)

>>109503864
nah
>>
>>109503871
what do you mean "their"? who is they? i only saw one person
>>
turbo lora and sage attention are all I need
imagine having more than that
>>
>>109503739
Shiina best girl...
>>
File: getAjob.webm (2.11 MB, 1143x2048)
2.11 MB
2.11 MB WEBM
>>109503873
I wish krea2 could do that desu, even the style can be well preserved
>>109503877
I guess you're going to suck your way to a new gpu. Fair enough, you could try getting a job but....
>>
File: Screenshot_3547.png (14 KB, 1461x122)
14 KB PNG
>cucking yourself with snake oil instead of being patient
couldnt be me
>>
So I actually read the schizo rentries and 99% of them is complete fabrications. care to explain author schizo?

also who replaced the collage generator with lowfps vibe coded slop?
>>
>>109503890
based patient anon
>>
>>109503879
Saw? I'm talking about an internet post.
>>
two pass workflow
first pass
>0.2mp
>err_sde / beta57 15 steps
>cope free

second pass
>RTX upscale to 0.8mp
> eular / beta - 0.4 denoise - 4 steps
>turbo lora + sol

total runtime: 4min16s on rtx 3090

https://h.uguu.se/rUbOiawR.mp4
>>
>>109503904
wow, that's pretty good
>>
>0.2mp
guillotine
>>
File: MiniMax_H3_00033_.mp4 (533 KB, 928x672)
533 KB
533 KB MP4
>Watashi, kirei?
Translation: Am I pretty?
https://files.catbox.moe/1evygl.mp4
>>
>>109503904
cool but sadly stereo is still a bit of a ways off, doesn't work well
>>
File: cw1.jpg (21 KB, 1077x33)
21 KB JPG
>>109503890
My brother. 0.8mp 15secs absolutely raw not even sage on a 3090. Feels good to be pure.
>>
>>109503921
honestly you can sage without worry, it's visually the same and way faster
>>
>>109503924
>honestly you can sage without worrty, it's visually the same and way faster
this is true and if you're still skeptical read the sageattention2 paper. sageattention3 is lossy however
>>
>>109503904
0.5 denoise might be even better
https://n.uguu.se/TrkDYPcK.mp4
>>
File: 1701311567875653.jpg (140 KB, 1080x1080)
140 KB JPG
sage + mem eff sage + easy cache + turbo + torch compile + radial att + teacache + spectrum + mangekyo sharingan = izanagi
>>
i cant stand that i need to reopen comfy every so often because it decided to just consume ram while doing nothing
>>
File: H3_webm_00001.webm (3.84 MB, 1680x660)
3.84 MB
3.84 MB WEBM
>>109503541

>KJ sage vs caches mixture

Huh. The cope caches mixture.... are pretty good? Both are cold start from boot up. 5090, ref model, 1MP
>>
>>109503957
where else would comfycloud get their hardware???
>>
>>109503890
>>109503921
>when you're so afraid of breaking something by updating comfy you started to become patient.
>>
>>109503961
>tests it on tranime
every time
>>
repost.
Sonnet II continues to be repainted.

>>
Anonymous 08/08/26(Sat)21:34:12 No.109503664▶
>>109503648 (You)
reminder this week's gens
india diss:
https://files.catbox.moe/dioqb9.mp3
sonnet III
https://files.catbox.moe/sddmtd.mp3
sonnet I
https://files.catbox.moe/w2ysvt.mp3
>>
File: MiniMax_H3_00034_.mp4 (360 KB, 928x672)
360 KB
360 KB MP4
>>109503913
https://files.catbox.moe/nk3udp.mp4
>>
>>109503971
i pee pee'd my trans panties. you owe me a new pair
>>
>>109503971
wow, eye contact. Don't ever get that shit.
>>
>>109503975
calm down and listen to thread music:
>>109503970
>>
>>109503962
bro...
i'm gonna cry
>>
>>109503971
im sorry, asians are just not creepy.
>>
>>109503975
what makes trans panties different from regular ones? the bulge pocket?
>>
>>109503990
I beg to differ
>>
>>109503994
has a dilator built in
>>
>>109503971
House 2 lookin good!
>>
>>109503949
yeah sage 3 isn't that interesting, even if it's fast, it's too lossy
>>
Why do you guys hate sexy minimax gens so much?
>>
>>109504000
whats creepy about this?
>>
>>109503970
oh yeah and another india diss:
https://files.catbox.moe/92cery.mp3
>>
might be on to something with this 2pass wf

first pass
https://d.uguu.se/TfnamihY.mp4

second pass
https://n.uguu.se/ppDynuxh.mp4

I never had this gen look this smooth before.
>>
Total /ldg/ Victory?
>>
>>109504056
>mp
:^)

>4
*jazz music stops
HARAM!!!! ABSOLUTELY HA''M
>>
>>109503739
it really is a mess the more i use it. i wish he would've picked a better base. i hope we get a krea 'tune or something, it's hard to keep using when the prompt comprehension is so bad compared to krea.
>>
>>109504064
a small victory in combat does not mean the war is won
>>
>>109504070
are you disdaining the tillage of thy husbandry?
>>
is pinned memory just flat out broken? it's jumping to OOM on 1s/0.1MP. I know just disable it, but I've got tonnes of ram just doing nothing and my gpu is only ever half full
>>
>>109503924
>>109503949
guess I will taste the forbidden fruit. should I use the kjnode or just add it to the bat file?
>>
>>109503739
Who made this purposefully disingenuous "comparison"? kek
>>
>>109503913
>>109503971
nice, was wondering when we'd finally start seeing the horror gens.
>>
forsen

https://files.catbox.moe/ca03s7.mp4
>>
>>109504109
Both do the same thing, kj is just more convenient since you can disable it without rebooting comfy.
I personally have it on all the time so don't use the kijai node.
>>
>>109504149
sure feels like summer.
>>
>>109504149
how to get the layout/box:

Use <Picture 1> as the reference image for Forsen. Use <Audio 1> as the voice reference for Forsen's dialogue.

LAYOUT:
The output must look like a live broadcast stream layout. In the top-right corner of the frame, there is a small, sharp, rectangular webcam box displaying <Picture 1> reacting live.

SCENE & DIALOGUE:
The background features an anime style Hatsune Miku walking in Tokyo, holding a green leek vegetable with her hand. Inside the square webcam box, the streamer Forsen maintains a completely relaxed, calm, and low-energy demeanor. Forsen moves their mouth with a casual, effortless talking cadence.

Forsen says: "Wow what a great link, totally worth the 20 dollars, right chat? What a great use of your hard earned money."
>>
File: 1663125987808951.jpg (182 KB, 576x512)
182 KB JPG
Help. How can I generate creative thoughts?
>>
>>109504161
Read some books, be interested in things, stop to think, imagine, dream, hope
>>
>>109504109
i mean if you're lazy you can just add it to the bat file but kjnodes has a mem optimized one that you should probably use to savce vram
>>
so things like <Picture 123> is going to always point to frame 123?
>>
File: MiniMax_H3_00037_.mp4 (2.25 MB, 928x672)
2.25 MB
2.25 MB MP4
>>109503990
The video is not meant to be scary, it's an artform
https://files.catbox.moe/pemgwk.mp4
>>
>>109504171
I was waiting for the scary part....
>>
>>109503904
>>109503953
Wait, you can use LTX upscale in Minimax now ?? Source ???
>>
>>109504180
it's not LTX, the WF is in the files, using turbo lora as upscaler
>>
File: 1779820828042099.jpg (73 KB, 660x574)
73 KB JPG
>>109504189
No i mean Minimax dont use two pass gen type in the first place. Is that two pass gen ?
>>
>>109504161
if you can't do it personally, you probably can with an LLM
>>
>>109504202
>Minimax dont use two pass gen
says who? is it a rule? did I break the law?
>>
>>109504161
scroll civitai red
it may be porn but it will spark ideas
>>
>>109504202
nta Is just akin to a hi-res pass just with video gens. Is not really specific to LTX, you can do with any of the models. LTX just had it's own infrastructure around it
>>
>>109504209
Youre the first guy who did it. Share your tech to everyone
>>
>>109504217
it's in the workflow!!!
>>
>>109504176
same i expected someone to be behind her, would be a perfect use of AI
>>
hahaha

first time trying a non english voice clone gen.

https://files.catbox.moe/64h6np.mp4
>>
File: MiniMax_H3_00038_.mp4 (1.36 MB, 928x672)
1.36 MB
1.36 MB MP4
>>109504149
Kek, the voice is so realistic I thought it was really him commentating an AI game sarcastically for a second
https://files.catbox.moe/wkf1rj.mp4
>>
of course it's hitler
next it'll be floyd
>>
File: 1558174675225.jpg (10 KB, 240x200)
10 KB JPG
>mfw it consistently replaces physique, clothing and hair, but not the fucking face, despite explicit, correctly formatted instructions
>>
What vm are you guys using for cumfart? Just docker right?
>>
>>109504258
maybe it'll be miku and hitler dancing together
>>
>>109504176
>>109504224
Anon I'm just testing how well it can do 1girl vlogs for gooner purposes, though it would possibly would need a LoRA to make her masturbate and say dirty things
>>
>>109504236
just plug english in an english to german translator, use "says in german", plug in a downfall hitler voice sample in ref audio 0, and voila:

Use <Picture 1> as the reference image for Adolf. Use <Audio 1> as the voice reference for Adolf's dialogue.

https://files.catbox.moe/6si4d6.mp4
>>
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main

Anyone tried Step600 ??
How was it ?
>>
kino

https://files.catbox.moe/na4mn9.mp4
>>
>>109504310
>4
looks like it's not an mp3. Maybe it's some kind of mistake, I'll allow you to resubmit.
>>
Should there be any quality difference between int8 and nvfp4 CLIPs?
>>
>>109504289
step 600 ema is the best one to use currently.
>>
What's the best image edit model for anime right now? I need to start building last frames.
>>
Using H3, I am going to personally revive the 90s late night sleazy thriller movies. It will be a new golden age of boobs and violence!
>>
>>109504324
Not that I've seen. I now exclusively use nvfp4.
>>
File: asset_00001_.png (2.08 MB, 2112x1216)
2.08 MB PNG
I hope you boy's are behaving yourselves in here.
>>
>>109504385
bomboclat
>>
>>109504330
Whenever H3 releases their model...
>>
>>109504330
Anima.
Ignore the schizo.
>>
>>109504400
I use anima for generating the first frames already, it's great for that, but as far as I know there's no good tune for image editing.
>>
>>109504414
Then use Qwen Image Edit 2511
>>
>>109504330
local edit is absolute garbage, just like video was until h3. people will cope and tell you klein or qwen but they're just freetarded.
>>
asuka test (ref)

https://files.catbox.moe/m6l8zp.mp4
>>
>>109504417
Thanks, I'll try

>>109504438
I'll see I guess
>>
File: debo_tt_k2_00090_.png (1.97 MB, 1872x1007)
1.97 MB PNG
>>109504444
checked
>>
time flies so fast when genning its not even funny
>>
File: 1737772903328659.gif (2.07 MB, 415x498)
2.07 MB GIF
>>109504262
Well its a quirk from the model, at some point people will figure out how to train character loras. At this point after autistically tadwrangling it for 5 days i finally got it to be able to actually use the video reference for motion and initial positions. I still get hit by generic face and some anatomically horror beyond man comprehension here and there even with 9 reference pics from different angles so at this point i made peace with it and just gonna wait and focus on preparing the data sets to train the loras when there enough research about how to do them properly If you are extremely autistic remember that video are just a gorillion images so you can just use klein to faceswap every frame with some masking nodes
>>
>>109504438
Sadly true. Also lacking a good music model.
>>
sometimes i get "lazy" gens that only animate the specific most talked-about part of the i2v prompt, especially when the input image is an art style you wouldn't often see animated (hand-drawn, high quality art etc that it's rarely practical to animate), anyone found reliable workarounds for this? experimenting with "the whole image is animated" etc at the moment but maybe it's better to just explicitly spell out lots and lots of parts' motions and eventually it'll be animating half the image and decide to just animate the lot?
>>109504330
>>109504414
try qwen image edit or flux klein 9b, but yes, >>109504438 is basically right, edit's very hard. if it's SFW you can use api models of course, they have basically perfect edit, eg gpt-image-2
>>
>>109503713
The greatest output to ever come from LDG
>>
So, Today summary is only ComfyUI 0.31 which is support Kijai int8 VAE and this >>109504289
huh
>>
>>109504480
>anyone found reliable workarounds for this?
Other than a reroll and a prayer? No, I fucking wish.
>>
pretty impressive considering all I plugged in was baka shinji, 2 seconds only:

https://files.catbox.moe/p5vegl.mp4
>>
sorry guys i was too busy grilling up some steaks. kino will be generated soon
>>
There's a bug :(
>>
>>109504521
your wires make me sad. the bug is between the user and the screen
>>
>>109504523
i think you mean between the chair and keyboard retard-chama
>>
>>109504523
try it, minimax doesn't work on a ksampler if you put
>add_noise disable
>>
File: 1777374430678291.png (1.84 MB, 1024x1224)
1.84 MB PNG
>animeArtDiffusionXL_alpha3
>>
how come describing facial features is shit in every single AI model
You'd think being able to describe the shape of the eyes, nose, mouth and face would matter a bit more
>>
Did they messed something up again because my gen got slower with ComfyUI 0.31
>>
probably
>>
>>109504257
>east asian feet
the best there is
>>
File: HPPYIxgakAAkBys.png (738 KB, 1238x1376)
738 KB PNG
Based!

>YES WE WILL KEEP OPEN UNTIL AGI ARRIVES.

https://xcancel.com/MiniMax_AI/status/2086253065657790895
>>
>>109504438
why can't h3 be used as an image edit model?
>>
>>109504585
They've already said they're going to make an image model and edit model.
>>
ace step in the oven. is everyone else an indian in a cybercafe with only free horrible headphones available?
>>
>>109504580
agi doesn't even mean anything at this point
>>
>>109504480
>>109504438
Yeah, qwen image edit kinda sucks
Trying to create anything remotely pornographic just doesn't work, and I can't find any finetunes for it. What a shame.
>>
>>109504571
Nevermind it was the turbol ora fault.
larryvrh fuck something up again

>H3TURBO fwd full/bypass] call#1 qkv_proj.forward_owner=BypassForwardHook (BypassForwardHook => lora ACTIVE; else => BASE ONLY!) is_injected=True timestep=1000.00 video_rms=1.0000 audio_rms=1.0048 dtype=torch.bfloat16
>>
>>109504588
Someone should get ahead of them and make it.
>>
>>109504580
we are literally living in the best time
the time between when ai is first starting out and we still have full control, and the great filter in the next 5-10 years
>>
>>109504585
you can, set length at 5 and extract the first frame. voila.
>>
>>109504611
>great filter
I have faith, we'll wing it like we did the nuke filter
>>
>>109504496
rei test:

https://files.catbox.moe/c326bl.mp4
>>
>>109504623
saving this for your notes alone. thanks nerd
>>
>>109504623
there we go, close shot of Rei worked nice.

Use <Picture 1> as the reference image for Rei. Use <Audio 1> as the voice reference for Rei's dialogue.

LAYOUT:
Close shot of Rei. Rei is standing on a sunny beach in Tokyo. Rei says in Japanese "(put english to japanese test text here)".

https://files.catbox.moe/vtcrlc.mp4
>>
File: MiniMax_H3_00041_.mp4 (2.12 MB, 928x672)
2.12 MB
2.12 MB MP4
>>109504257
>>109504574
If only it could always spell right kek, I would regen but then I have to wait another 10 mins

https://files.catbox.moe/2o2uxa.mp4
>>
I wonder, if there was an extension to batch process gens with a script, in theory you could generate a show while afk, even.
>sequence of 15s clips
>stitch them all together
>>
>>109504580
>until AGI
Do they mean that they commit to releasing their hypothetical AGI model, or that they'll stop sharing as soon as they reach it?
>>
File: 1784302500084393.webm (3.91 MB, 576x1024)
3.91 MB
3.91 MB WEBM
https://files.catbox.moe/5xtze0.mp4
>>
Use <Picture 1> as the reference image for Rei. Use <Audio 1> as the voice reference for Rei's dialogue.

LAYOUT:
Close shot of Rei. Rei is standing on a sunny beach in Tokyo. Rei says in Japanese "Totemo sutekina katada to omoimasu. Motto ohanashi shimashou!". the girl from <Picture 2> walks in from the right and says "Miku, dayooo".

https://files.catbox.moe/dy4wax.mp4
>>
alright, it's high time i stopped lurking and posted a vid
https://files.catbox.moe/1rldy7.mp4
>>
>>109504663
Especially since you can use fl or ref and just indicate the last frame of last video is first frame of new video
>>
File: vlad.png (102 KB, 182x346)
102 KB PNG
>>109504672
long shower before + bag over head during
>>
>>109504679
light chuckle
>>
>>109504679
lole
>>
>>109504392
https://files.catbox.moe/mw4ybn.mp4

You shall not pass.
>>
>>109504699
>claude, write me an edgy prompt.
>>
a big anime for you:

https://files.catbox.moe/xyrqhm.mp4
>>
File: mmh34_video_edit.webm (1.13 MB, 1600x544)
1.13 MB
1.13 MB WEBM
H3 can do video editing.
>>
>>109504723
which is which
>>
>>109504615
Why not set the length at 1 frame?
>>
>>109504737
nta but you can't it's just how the model works
>>
>>109504723
You missed the shadow
>>
>It doesn't understand "bandaid underwear"
>Show it a reference
>It gets it the very next gen
man I love living in the future
well, well worth 25% longer gens over the first/last frame model
>>
>>109504723
It can't remove/swap minor details at specific timestamps.
>>
>>109504664
The latter
>>
>>109504611
no, it's terrible time
>>
>>109504743
>nta
why do redditors always say this? just answer his question without saying that, nigga
>>
>>109504754
Oh, and I tried describing it about 30 different ways, exact placement, super literal descriptions, nothing worked. A reference instantly solved it.
>>
What should I generate?
>>
>>109504580
>until AGI
translation: until it's competitive with seedance, and I have no doubt minimax h4 will be competitive, so we'll never get a model ever again from those fags
>>
>>109504767
>redditors
>not that anon
hmm
>>
>>109504672
pajeet
>>
>>109504672
this isn't ai
>>
>>109504785
what?
>>
>>109504580
china is so fucking cool seriously, do you see a single western company screenshotting a "The Rock" reaction image and posting it officially on twitter??
>>
>>109504792
correct. both the webm and the catbox are real.
>>
>>109504779
If the "Seedance is 200b" rumors are true, they might just have to train a Large version of H3 and put it behind API
>>
File: nojoy.jpg (13 KB, 314x302)
13 KB JPG
>>109504679
>>
>>>/wsg/6210720

Ref model with an audio clip and just prompting something like "Dance to the music of <Audio 1>" works pretty well.
>>
>>109504653
Total tomboy enjoyer victory
>>
>>109504580
I still remember back when anons were talking about how Kling or Minimax will never be open-source,, all we had was Wan
>>
>>109504679
this is why I pay for the internet!
>>
>All development slowed down today
Guess we all running out of semen huh
>>
>>109504793
you're the redditor if you think "nta" is a reddit thing, it has a different meaning on 4chan, newfren
>>
File: mikusalute.webm (412 KB, 736x576)
412 KB
412 KB WEBM
>>109504723
Yeah but it's a major pain in the ass to get it working though. You have to write a hugeass prompt depending on the complexity, and get a good enough video where the model can figure out where is what.

I could only do the picrel (out of the MGS3 big boss gif) with an LLM-written prompt based on the official docs
>>
>>109504816
>I still remember back when anons were talking about how Kling or Minimax will never be open-source
I said that, I never expected Minimax to make such a move, not that I'm complaining though!
>>
Ugh. Really hate it when i got low framerate animation in Anime style
>>
my pc just segfaulted right at the end of a 20 min gen...
>>
>>109504672
gross, that’s indian
>>
File: 1765429766860517.mp4 (924 KB, 736x704)
924 KB
924 KB MP4
>>
>>109504830
>you're the redditor if you think "nta" is a reddit thing, it has a different meaning on 4chan, newfren
no, you're the redditor if you feel the need to deanonymize yourself just to give a simple answer to a question. fact.
>>
>>109504815
It's all that gay stuff. Just admit gayness or get a real woman.
>>
File: 1774062405213297.gif (2.26 MB, 498x498)
2.26 MB GIF
>>109504580
>The answer is YES: we plan to open-source a unified text-to-image and general image-editing model derived directly from the H3 lineage. It is currently in post-training refinement.
HOLY SHIT WERE SO FUCKING BACK WTF
>>
File: MiniMax_H3_00070.mp4 (3.82 MB, 1440x816)
3.82 MB
3.82 MB MP4
>>
>>109504841
https://civitai.red/models/2843112/astrowitch-cinematic-comic-style-lora-by-astroburner-ai?modelVersionId=3209714

>Please pay to access this model goy

Fuck off
>>
Why did you guys write off wan fun-VACE off so quickly? I think you are all overrating minmax and underrating wan's true capabilities.
>>
>>109504828
This is a civilization destroying technology
Ultimate dopamine machines
People are probably getting addicted to genning
As more barriers are removed it will become worse
>>
>>109504653
idk, doesn't seem to have the same crazy errors as krea does.

I'm not too happy with the feet, but they are not quite goofy Gnome feet.
>>
>>109504858
dont make me feel old you son of a bitch
>>
>>109504580
surely a CEO would never lie
>>
>>109504854
oh ok, sorry for misinterpreting your point, i disagree with it but you're entitled to your opinion
>>
>>109504856
nahh, tomboys are the ultimate expression of femininity, if a woman is still beautiful with short hair that means she's naturally way more feminine than your average woman who has to get her hair long to be attractive
>>
>make a unique workflow with help from claude to solve the errors I'm getting
>spend like a week on it and finalize it
>works amazing
>update comfy
>workflow still works but now has blur in all of the gens

NIGGERS
>>
>>109504870
he doesn't need to lie here, his post is so vague it can be interpreted in any way he wants, "AGI" can mean anything
>>
>>109504871
thank you
>>
i click on every video because i know anon took a lot of time to plan it and generate the output
>>
>>109504857
>>109504580
so we'll probably get Minimax edit before Krea 2 edit huh? interesting
>>
>>109504887
I love you
>>
>>109504679
fucking slop so is everything here.
>>
If Minimax always force you to get low FPS on anime, theres no reason to have 24fps genning render
>>
>>109504860
its all fucking slop all of it why would you care? every krea lora is fucking slop everything is slop all fucking slop.
>>
>>109504909
>>109504895

>>>/v/
>>
i deleted 2 TB of slop and i feel better, i am free.
>>
>>109504909
Don't fight the slop..embrace the slop. The slop is your friend
>>
I hope the Minimax devs use what they already have for H3 to make an audiogen model too. Shouldn't be expensive since they would essentially just have to fine-tune it to generate long audio. One audio model that does everything: TTS, sound effects, music. Bonus points if you are not forced to autistically add key and bpm into the prompt like AceStep requires.
>>
>>109504919
damn, anima krea and h3 really take up that much space?
>>
File: be_not_afraid_snafu.jpg (287 KB, 1920x1080)
287 KB JPG
be not afraid (snafu)
https://youtu.be/JvjxOLRNLeI
https://suno.com/s/zMaaz8vjHjZ1Nam1
>>
Seems like the hype winded down, threads were def. faster the last few days
>>
>>109504931
pretty harsh on the ears no?
>>
>>109504929
No Anima, Krea, H3, XL. I only have left my one and only friend SD 1.5
>>
>>109504946
1.5 is just 1.4 but slopped
>>
>read h3 loras comments on civitai
>most of them is just users complaining that they get body horrors and they don't work
>>
File: lastframe_00002_.png (775 KB, 1056x608)
775 KB PNG
>>109504929
>>
the scalping solution:

https://files.catbox.moe/s6m3s9.mp4
>>
>>109504953
loras are gonna be hard to train, they only gave us the guidance distilled model
>>
INT8 Video VAE saves gen time by 30 secs
Nice
>>
>>109504965
also fuckstarts the quality
>>
>>109504893
I would not mind for them to tease us a bit with some image showcases, if they said they almost finished it then it probably looks good already
>>
>>109504944
cum get thinned out bro. Need days to recover
>>
File: 542759506583435.mp4 (3.46 MB, 1280x736)
3.46 MB
3.46 MB MP4
>>
>>109504876
>>make a unique workflow with help from claude to solve the errors I'm getting
what errors could you possibly be talking about
>>
>>109504965
i would not do that if i were you.
>>
>>109504873
It's mere homosexuality.

Like having women on scotus or whatever. Mere homosexual tendencies.

Real men put women under strict rules. Certainly they don't vote.
>>
>>109504953
Keep in mind that users on civitai are extra retarded.
>>
>>109504992
lmao what
>>
>>109505006
>anon always thinking about queers
you're gay
>>
Just one more gen
one more gen
one more
one
more
>>
>>109505006
you also hate women another red flag or indication that you are gay.
>>
>>109504960
better scalping solution

https://files.catbox.moe/5672u9.mp4
>>
the `text generate` node got a bit more retarded with comfy update. impressive
>>
>>109504925
You can audio gen with H3 by generating a video with the minimal resolution (32x32) and a long length (~1 min).

>>>/wsg/6209660
>>
>>109505019
Never sleep
Only gen
>>
File: MiniMax_H3_00661__1.webm (1.95 MB, 896x704)
1.95 MB
1.95 MB WEBM
>>
>>109505009
If your lora cant be used by retards then its badly trained
>>
>>109505079
only a retard would say that
>>
>>109505040
>swap to the next `text generate` node available in cumfart
>same settings, thinking enabled
>doesnt start off mid-sentence, doesnt copypaste the same prompt i put in 3 times, doesnt cuck the gen
>everything just works
someone on the inside is trying to sabotage comfy
>>
>>109505019
Just one more gen and I'm going to bed.
One more.
No the preview of that one's not good enough. Cancelled.
Just one more gen and I'm going to bed.
One more.
>>
>>109505036
That scalper shouldn't have killed his dog!
>>
>>109504619
>we'll wing it
there's no "we" here. we're going to die to a "black hole in the garage" event where someone unleashes a paperclip maximizer and everyone gets turned into a mcdonalds hamburger

>>109504679
kek her facve in the last frame
>>
R2V turbo lora when? Please, I need it.
>>
>>109505024
Yeah, everyone hated women before women had the vote, just like in the reddit moooovie slop.
>>
>>109505107
The power behind r2v is incredible. Imagine when we get an upscaler, loras, any official updates from Minimax - it's gonna be impossible to do anything but sit here and gen
>>
>>109505122
sadly thats not gonna happen, the team behind h3 will be hired for a private company and those models will never be released, just like happened with z-edit
>>
File: 1781194100541878.png (35 KB, 1480x168)
35 KB PNG
>not a meme
>>
I need to a sprite sheet for a project, I dont care so much about the style or being blown away by how great the visuals are. I just want it to be functional/consistent between sprites. What would be the best model for this?
>>
>>109505122
>it's gonna be impossible to do anything but sit here and gen
im already getting bored
>>
>official updates from Minimax
>
>>
>>109505137
>z-edit
Z-image edit status??
https://files.catbox.moe/sf09k4.webm
>>
File: 1785966070138357.mp4 (3.08 MB, 736x1120)
3.08 MB
3.08 MB MP4
>your cpu
>your gpu
>your ram
>how youre feeling about it
go
>>
>>109505148
you're gonna get maxxed out by Tyrone in jail, pedo
>>
>>109504887
I use JavaScript to delete all posts with videos from the rendered HTML on 4chan and Reddit via tampermonkey.

It's the closest thing to a genocide of these people, and it's the best experience.
>>
>>109505122
h3 already hit plateau, anything complex will filter 90% of lazy AI jeets, and if you hadn't already noticed, anons like >>109505157 got out of ideas and they just repost their unoriginal funny slop, oh its that another Seinfeld but they don't look anything like Seinfeld h3 slop? or is that another the big bang theory h3 clip
>>
>>109505157
can he do the heart rip one?
>>
>>
I continue to do repaint work with ace step 1.5 xl base, on Shakespeare Sonnet II
>>
that's one gigantic ass (adult woman:-1) workflow kek
>>
>>109505112
no you are gay. you do not understand human biology.
>>
>>109505178
the reference model can make anything look 1:1, there are no limits with the ref model. image/voice/whatever.

https://files.catbox.moe/na4mn9.mp4
>>
>>109505148
Total Janny Death
>>
>>109505143
good.
>>
>>109505195
all deviation from the Eve ideal are faggotry. I'll die on this hill bring it. If you don't want Eve, YOUR GAY.

Eve: basically Venus.
>>
File: h3.png (148 KB, 349x303)
148 KB PNG
>>109505197
wow yeah, it looks amazing, gj anon
>>
>>109505157
https://www.youtube.com/watch?v=0oMEuyhBkRo
>>
>>109505201
Nah, pretty sure they deleted it themself. The workflow from the catbox has some leftover prompts in it lmao.
>>
>>109505208
yes it's a 0.3mp test render ofc it wont be high fidelity
>>
>>109505208
>the ai learned to have the same blur as movies
idk man
>>
>>109505178
not my gen i just saved it in case i need to post one that isnt mine for a pointless post
>>
>>109505215
>they
>themself
what do you mean? did multiple people type out that single post?
>>
>>109504672
some poltards actually think this shit is real
>>>540459779
looks like my work today is done
>>
and of course it doesn't resemble the German prophet, but that's clearly a feature to avoid profaning him by creation of a graven image.
>>
>>109505223
uh oh, esl melty
>>
>>109505208
name another model that can make AVGN like this.

https://files.catbox.moe/xqsl9a.mp4
>>
>>109505228
idgi
>>
>>109505197
>https://files.catbox.moe/na4mn9.mp4
KEKKKKDD
>>
>>109505157
Two pcs:
96gb ram
5090
9950x3d

96gb
5090
5950x3d

Feels pretty good.
>>
File: v.jpg (152 KB, 1051x1200)
152 KB JPG
>>109505148
those bypassed nodes aren't chinese cartoons..
>>
>"she is talking animatedly"
>every seed is just two hands raised and mouth open
>>
>>109505238
nice
>>
>>109505254
>one for waifu
>one for tiger mom life coach
>one for vibecodez
>one for gaymer
>one for blender
>one for music diffusion
>one for image diffusion
>one for video diffusion

you don't have enough.
>>
the harsh true is that h3 turbo loras sucks, they only look good in 2d animations or close-up shots, any complex motions looks like shit or stiff, light2xv has better motion of they have been using the same synthetic dataset since their first turbo lora, so it changes the face of the subjects
>>
okay. NOW the scalpers are finished.

https://files.catbox.moe/tmdb5o.mp4
>>
>>109505270
There aren't any complete turbo loras yet thoughbeit
>>
>>109505238
the problem with these more static gen styles is it's not exactly pushing the capabilities of H3, it's the kind of thing LTX could do - that is not to say LTX would be as good but it could do the voice and the emotion - you could literally do the soundtrack you wanted in a t2s model and that would be even better.
>>
File: 1768505046165492.png (128 KB, 498x311)
128 KB PNG
>>109505270
anon, wait for them to finish the cook, it's only been a week
>>
File: mmh3_video_edit_stop.webm (1.1 MB, 1600x544)
1.1 MB
1.1 MB WEBM
>>109504760
Well, I did try, but if the cyclist do not stop, the sign must insist.
>>
File: 1646333888461.png (340 KB, 787x720)
340 KB PNG
I haven't been genning minmax all day. I've been studying and brainstorming the capabilities of R2V.
I'm not bothering to post any of my test gens here, but I am making progress. ChatGPT 5.6 Sol has been very helpful through this.
>>
does this suffer the same problem as SCAIL where the character swapped matches the body of the target video?
>>
>>109505297
kek
>>
>>109505288
The reference control pushes it leagues beyond ltx, even with more static gens.With ltx you could generate a start frame with an image model and then pray it works as you wanted it, but with H3 you can have whatever characters doing whatever whenever with whatever audio thanks to references.
>>
>>109505279
no one fucking cares man, get some imagination, like really get into knowing how to imagine a world and create it, start simple. let the model give you ideas, don't use speech prompts just say "2 men talk about thing" and let the model fill it in, it will be mostly gibberish but it will include some English about that thing but the simlish is what inspires.
>>
>>109505306
no i don't think it does but that method is hard to prompt and takes longer to process.
>>
File: this.png (711 KB, 640x741)
711 KB PNG
>>109505297
ngl this goes hard, like the sign won't stop following you unless you fucking stop
>>
>>109505311
ltx could do references too
>>
>turbo lora
>shift enabled
>spectrum enabled
>dont change steps from 20
shit whoops
>>
>>109505314
anon, we are full throttle into retard/autistic h3 cringe slop, their creativity and unfunny shit just replicates how they are in real life
>>
spectrum was absolutely shit in WAN. why should i trust it'd not be shit in H3
>>
>>109505328
Oh huh. I spent like 3 days using it, generated nearly a thousand videos, and only liked one of them so I stopped. I genuinely had no idea it could do references.
>>
>>109505338
and your constant bitching replicates how you are in real life?
>>
>>109505250
add english subtitle
>>
>>109505142
anyone? :(
>>
>>109505311
That is a different thing. If you are basically doing a guy sitting still from that source image then it's not really going to be different unless you start do cutscenes which is kind of weird because why would a guy reviewing shit be doing editing? Yes I realise he is a parody e-celeb doing "reviews" for comedy value. Not that I have ever watched him anyway so I don't' know if his delivery style sounds so amateur as usual e-celebs have a better vocal delivery and are more professional - he sounds like he has never done anything in front of a camera before.
>>
>>109505360
krea2
>>
>>109505346
you tell me
>>
>>109505367
ok ty
>>
>>109505360
>>109505142
probably gpt image if its sfw. if its nsfw i dunno
>>
bakin
>>
>>109505342
it kind of works like the FL2V model. you put your characters into the starting frames and you can also put audio at the start as well. then you start the prompt for the scene you want
>>
>>109505142
>>109505360

Unfortunately it's SaaS / not local, but check this:
https://www.autosprite.io/
>>
>>109505342
it really is somewhat similar to what H3 is doing but obv it isn't as good because the LTX model isn't as good. But LTX used within it's limits works pretty well. Use a short audio ref (say a voce) and it will replicate that voice just like you can in H3 with the talking parts in the prompt. I used to use an enhanced prompt so that the speech was varied around a subject matter for them to talk about.
>>
>>109505389
make sure to put my gens in the fagollage or else i WILL sperg out
>>
File: ed5.png (131 KB, 680x1112)
131 KB PNG
>make sure to put my gens in the fagollage or else i WILL sperg out
>>
>>109505402
how I'm supposed to know what gens are yours though? kek
>>
File: 1782998240737885.png (30 KB, 600x600)
30 KB PNG
>>
>poorfags unable to animate their favourite jaks
>>
File: kekekekekkke.png (1.17 MB, 864x1184)
1.17 MB PNG
>favourite
>>
Minimax doesn't know teto so I had to use the reference model
https://files.catbox.moe/aswsua.mp4
>>
https://www.reddit.com/r/StableDiffusion/comments/1vj5scs/deroping_minimax_h3_fast_motion_to_reduce/

https://github.com/matlowai/ComfyUI-MAINodes

>H3 smears bursty motion: backflips, fast sword arcs, whip-fast reversals. The cause is structural. One latent token spans four pixel frames, and at high motion speed those four frames need four distinct poses that a single token can't hold. Re-denoising the affected region doesn't help, because the missing poses were never generated in the first place.

>This pipeline works around that at inference time. It re-generates the clip as a slowed-down version of itself, seeded from the original. Frames where motion is too fast get held (repeated) so the model has more temporal room, the result is generated video-to-video from that retimed init at partial denoise, and the original frame rate is recovered afterward by dropping the held frames. The oracle that decides where to slow down reads the clip's own latent. No extra model, no training.
>>
tv is saved

https://files.catbox.moe/eksdh5.mp4
>>
>>109505420
mine are the good ones obviously
>>
>>109505402
im not actually bakin
>>
>>109505444
people cried about the final season so much I decided to never watch this show, thanks for beta testing this for me fags!
>>
>>109505443
>oracle
demon
>>
>>109505397
>>109505379
hmm if i cant get usable results with local models ill give this a go, thanks
>>
>>109505444
checked 10/10

thanks for keeping the file lmao
>>
>>109505443
sounds amazing and also that the computer has to do some extra work
>>
Blessed thread of frenship
>>
>>109505444
>>109505470
unironically the first video gen in the world I've saved afaik
>>
>>109505449
no one wants to run the malware i take it? cause i'm not bloody trusting it.
>>
>>109505484
i refuse to make these threads on the premise of i dont care i got other shit to worry about
>>
>>109505489
>>109505489
>>
>>109505484
ask claude if the code is sus or not, that's all
>>
>>109505492
that benchod can't even go through a 3kb file at high effort without telling you to upgrade
>>
>>109505492
still won't trust it, it could change any time, id rather do something like that myself, it could be knocked up pretty fast i'm just too lazy.
>>
>>109505500
>it could change any time
what
>>
>>109505500
i mean i had similar bash scripts using ffmpeg to make a canvas so that different size videos could be merged. similar can be done with the collage. I do not trust code i can't read sorry.
>>
>>109505505
famous last words people.
>No how could they have done this?!?
>>
>>109505505
if i do not understand the code i won't use it.
>>
>>109505530
they couldn't, it has no update URL and the one dependency you can just copy paste into the script
>>
>>
MOOOOOOOOOOOOOOOOOOOOOODS
>>
File: altMiniMax_H3_00001_.mp4 (1.21 MB, 640x640)
1.21 MB
1.21 MB MP4
I'm trying MinMax H3 for the first time, it takes 15 minutes on a 4060ti for a 10 second video at 480p, at default settings (slow but relatively acceptable although not at WAN levels with 4-step LORAs). I wonder if there are already 4-step LORAs for MinMax.

https://files.catbox.moe/jdugu9.mp4
>>
>>109507221
there are scuffed ones but we are waiting for official ones
>>
Playing around with the T2V H3 model now for a bit and really struggling to get truly dark scenes. Even if I prompt stuff like "dark room, darkness, dimly lit, video shows a very dark room, chiaroscuro" etc. it always wants to add studio lights on the people in the shot. Any anon with tips here?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.