[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109483964

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109485064
fax
>>
>>109485123
Damn i got error every time i tried to update my comfyui to nightly
>>
File: 1767712655107740.mp4 (1.16 MB, 512x800)
1.16 MB
1.16 MB MP4
>>
>>109485128
>collage
finally, thanks for the bake anon

https://files.catbox.moe/xbzkob.mp4
>>>/wsg/6209573
>>
most kino collage in years
>>
>mfw Resource news

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo – OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view

>OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
https://xin1u.github.io/OminiVR_PAGE

>DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
https://github.com/Zhong-Chenchen/DIVE.git

>Multi-View Face and Gesture Animation with Dynamic Gaussians
https://dfki-av.github.io/MVFGA

>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
https://empaava.top

>Context-Anchored Tile Refine
https://github.com/Blakeem/ComfyUI-ContextAnchoredTileRefine

>ComfyUI Video Tiler
https://github.com/maDcaDDie2000/comfyui-video-tiler

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE
>>
BLESSED COLLAGE THREAD
>>
>mfw Research news

08/06/2026

>When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions
https://arxiv.org/abs/2608.04820

>HelloWorld: Enabling Socially Interactive Characters in Video World Models
https://arxiv.org/abs/2608.05070

>OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
https://arxiv.org/abs/2608.05049

>ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
https://guoxu1233.github.io/ContextMaster

>STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
https://arxiv.org/abs/2608.04887

>Simile Understanding in Text-to-Image Models: An Evaluation Framework
https://arxiv.org/abs/2608.04750

>ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
https://arxiv.org/abs/2608.04436

>CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
https://arxiv.org/abs/2608.04302

>Rethinking Pixel Mean Flows via Interval Denoiser
https://arxiv.org/abs/2608.04818

>Persistent Object Narratives for Token-Efficient Video Language Models
https://arxiv.org/abs/2608.04866

>Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
https://arxiv.org/abs/2608.04349

>Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
https://arxiv.org/abs/2608.04483

>Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models
https://arxiv.org/abs/2608.04454

>When does training on downscaled images yield the same gradients?
https://arxiv.org/abs/2608.04448

>Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection
https://arxiv.org/abs/2608.04935
>>
>>109485143
funny
>>
for me its 4 and 10
>>
Nope, deleting my comfy install. It's too addictive, this is a dead end.
>>
>>109485138
did you ask claude or chatgpt to fix it?
>>
ok I think I'll try ref model
>>
for me it's 2 and 12
>>
excluding mine,
1 and 3
>>
How do I trick H3 to generate a single image?
>>
>>109485158
I get it. I uninstalled all of my shit earlier this year and deleted all of my gens. Came back though, vid gen just too tempting rn :'(
>>
reminder that with local AI, even if the entire world ended tomorrow, you would have infinite entertainment at your disposal.
>>
>>109485173
Try ref image size on max mode.
I hear it's better trying to test it to see if it's true
>>
File: 1782917296440420.mp4 (638 KB, 544x800)
638 KB
638 KB MP4
>>
>>109485201
is animate inanimate a fetish of yours? because it's patrician tier shit you're making here pal.
>>
File: MiniMax_H3_00269_.mp4 (2.77 MB, 736x576)
2.77 MB
2.77 MB MP4
>>
>>109485188
>describe frame you want
> set video length to like, 2 frames

>>109485201
damn, it does stop motion quite well
>>
why didnt they do the same frame injection system as ltx? that would solve the problem of seamless video continuations
>>
>>109485208
I thought it was supposed to be good at anime
>>
>>109485156
>for me its 4 and 10
same, assuming you're talking about ages
>>
File: MiniMax_H3_00034.mp4 (2.32 MB, 1296x720)
2.32 MB
2.32 MB MP4
>>109485225
Seems fine.
>>
File: 1771207741977560.mp4 (1.95 MB, 640x640)
1.95 MB
1.95 MB MP4
>>109485205
>is animate inanimate a fetish of yours?
no i just like how it can do stop motion. but what would be an example of something less patrician in that fetish?

>>109485215
indeed https://pastebin.com/WBbYXvPE
>>
File: testmax.webm (2.16 MB, 1550x2048)
2.16 MB
2.16 MB WEBM
max might be the way to go
>>
File: 0000000000000.png (121 KB, 1631x1045)
121 KB PNG
I've got a question for anyone who knows the ins and outs of generating loras.
When compiling a big collection of images to train the lora on, do I want to crop it to remove all empty space?
>>
>>109485225
I never claimed it was a good prompt.
>>
Time to sleep.
I will gen less tomorrow. Time to do my work
>>
>>109485253
No.
>>
he's not wrong though, euler sucks!
https://files.catbox.moe/61dif8.mp4
>>
>>109485247
If you want her to suck it like one of those gravure videos dont mention that ice cream as food
>>
you weren't kidding, ref is like 30% slower
>>
>>109485268
>the incoherent rambly gibberish when they start fighting like toddlers
lol
>>
>>109485268
funny
>>
>>109485270
I intended for her to eat it because it's in a public space.
I should change the icy shape if I want to do that.
>>109485277
change the pixel setting to max too and you get more speed losses
>>109485268
I've had great success with euler, what's wrong with it?
>>
File: debo_sc_k2_00042_.png (2.3 MB, 1872x1007)
2.3 MB PNG
>>
>>109485295
good gen debo.
>>
File: MiniMax_H3_00014_.mp4 (2.42 MB, 1280x736)
2.42 MB
2.42 MB MP4
>>
File: h3_00008.mp4 (3.49 MB, 736x576)
3.49 MB
3.49 MB MP4
>>109485208
same prompt but with ref model and the reference panels, (gemini already shat me a ref prompt)
>>109485288
>change the pixel setting to max too and you get more speed losses
The 30% was already with max. I'm ok with it being slower.
>>
>>109485318
holy kino, can you share it with sound?
>>
File: MiniMax_H3_00049.mp4 (2.51 MB, 832x1248)
2.51 MB
2.51 MB MP4
https://files.catbox.moe/0n93fm.mp4
>>
>>109485318
She got paid $10k by the Olympics for just showing up btw.
>>
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/blob/main/minimax_h3_turbo_4step_ckpt850_pruned_comfyui.safetensors
it's overcooked, so go for a strength of ~0.75 it's still better than 500 steps at strength 1
>>
>>109485268
kek, nice one.
>>109485225
It is
>>
>>109485327
https://streamable.com/cdjsv4
>>
>>109485340
It's better if you throw it in the trash and stick to 20 steps with optimisation nodes.
>>
>>109485347
kek, nice
>>
File: debo_sc_k2_00041_.png (2.35 MB, 1872x1007)
2.35 MB PNG
>>109485299
tyty

>>109485318
my goat
>>
>>109485340
i find that my gens become a bit overcooked as they go past ~7 seconds
>>
the latest comfy update fucked lots of things and is causing lots of glitches
>random noise widget gets stuck and using the same seed.
>preview gets stuck
>image loader hangs until you refresh the tab
>job queue gets stuck at the end of a gen, causing the run button not to work
the fuck did comfy do to mess it this badly in one update?
>>
>>109485340
are we really supposed to use 4 steps for these 4 step loras?

the lora from the other day where it was recommended to use 10 steps I think was producing better videos.

i mean i guess it's obvious that 10 steps would be better than 4 steps?
>>
for once. I'm not pooling.
>>
>>109485396
for the moment not really, I go for 8
>>
File: MiniMax_H3_00045.mp4 (3.16 MB, 1296x720)
3.16 MB
3.16 MB MP4
LLMs help with prompts a lot if you give them the prompt guide docs

https://files.catbox.moe/ul1krv.mp4
>>
File: 1771621847761873.mp4 (1.12 MB, 736x544)
1.12 MB
1.12 MB MP4
>>
yeah sorry tranime pedos are not allowed to comment on the turbo lora. it's complete dogshit for anything live action and immediately gives it that "AI slop" look
>>
>>109485208
>>109485326
What's the manga/series? I've seen (presumably) your gens in the past and I like the style.
>>
Is unloading and loading models often a bad thing for your video card?
>>
>>109485427
i love clown girls so this speaks to me on a deep level
>>
>>109485427
me irl
>>
>sarrs live action i need the marvel superhero 1 bob 1 vagene family meetup kindly sir
>>
>>109485427
>pc exploded because my malicious comfyui node blew it up when it detected you forgot asian in the prompt
>>
>>109485427
I look like this and I do this!
>>
I just love H<3
>https://files.catbox.moe/z2fygc.mp4
>>
>>109485443
>it detected you forgot masterpiece, best quality, absurdres, highres, realistic, ultra realistic, score_10,
>>
>>109485454
Negative: Bad hands, deformed hands, no more than 5 fingers, no less than 5 fingers
>>
>>109485452
that is fucking completely insane what the fuck
this concept is so far out there can any other model even pull this off?
will have a cheeky wank to this later anyway.
>>
File: goofy ahh smiley.gif (3.24 MB, 640x640)
3.24 MB GIF
>>109485452
>H<3
aktually H < 3 would be H2 and H2 sucks
>>
>>109485452
she shouldve sucked the little one up inside her cookie
>>
File: MiniMax_H3_00046.mp4 (3.01 MB, 1296x720)
3.01 MB
3.01 MB MP4
T2VA is pretty nice with an LLM's help.
20steps res_multi/simple on 306012gb in under 5mins with just sageattention and spectrum
>>
File: MiniMax_H3_00272_.mp4 (2.27 MB, 736x576)
2.27 MB
2.27 MB MP4
>>109485432
It's the first time I do gens like this, but I think I know the anon too.
manga is BLAME! highly recommend it.

ok ref model is insane.
>>
>>109485470
sus
gotta be a turbo lora in there
>>
>>109485475
>BLAME!
Very cool gens, anon. Thanks. I'll check it out tomorrow.
>>
Aiiie someone (kijai probably) renamed or moved time shift slope in the minimax model implementation and it broke like every cache/attention/optimization node. A couple nodes have been updated but not all.
>>
File: ComfyUI_temp_ccsls_00001_.png (1.34 MB, 1152x1024)
1.34 MB PNG
>>109485475
reference
with sound >>>/wsg/6209600
>>
is there a consensus on the best sage attention node and the best cache node?
>>
Minimax h3h3
https://files.catbox.moe/xjrnx7.mp4
>>
>>109485459
none whatsoever, even image models struggle with such niche concepts and here we have a blown video model doing crazy concepts with simple prompting.
>>109485464
this is only the start
>>
cozy breas
>>
https://huggingface.co/SexGod1979/PinkCherry_MiniMax-H3
>>
>>109485519
>Trained for excellent rabbit motion, along with pink flowers with nectars glistening
BASED
>>
>high quality furry rabbits, rainbows and cherry trees (pink flowers open). Trained for excellent rabbit motion, along with pink flowers with nectars glistening
>>
>>109485478
generating at 0.4mp and using rtx upscaler as well
>>
>>109485519
>a whole ass checkpoint
actually, not based at all wtf? just do a lora hot damn.
>>
>>109485519
Why did vidrel suddenly pop up in my browser when i opened that link?
https://www.youtube.com/watch?v=mJmjljQP3oY&pp=ygUYaSBsb3ZlIGJlaWppbmcgdGlhbmFubWVu
>>
File: that's right.png (6 KB, 407x102)
6 KB PNG
>>109485524
>BASED
he's based in China after all
>>
>>109485475
mappa on suicide watch
>>
>>109485492
[beta] utilscollection :: MiniMax H3 Cache @ 0.05 / 0.15 / 0.90 / 2 / auto / true
>Skipped 7/15 block-stack executions (1.88x theoretical block-stack speedup).

easycache 0.10 / 0.30 /0.70
>EasyCache - skipped 0/15 steps (1.00x speedup).

easycache 0.30 / 0.20 / 0.90
>EasyCache - skipped 4/15 steps (1.36x speedup).
>speed wise just a tiny bit slower than the MiniMax H3 Cache outcome, about 5%, quality mostly the same between them both, though i prefer the first one
>>
>>109485533
https://huggingface.co/SexGod1979/PinkFluffyBunny-MiniMax-H3/tree/main
>>
>>109485574
so it's just the same lora as the one he has on civit?
>>
>>109485519
>linking this in the threads lee checks daily for anything to send license infractions to
You're going to hurt the Chinese person's credit score at this rate.
>>
>>109485092
Isn't ref the one that can do everything fl can do, but with more image inputs?
>>
>>109485580
It is just for generating better bunnies
>>
>>109485581
FL can do multiple references too, just feed it multiple characters in one image and prompt for the scene to change after the first few frames to whatever you want.
>>
>minimax_h3_turbo_4step_ema_ckpt850.safetensors
>recommended — current final checkpoint (time-averaged EMA, sharp at 4 steps)
Is this in a usable state?
>>
>>109485592
yea but dont use comfy converted, use originals with https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
>>
File: 1777078554062244.mp4 (1.33 MB, 768x544)
1.33 MB
1.33 MB MP4
>>
Wan animate 2 is comming (lol)
https://github.com/Comfy-Org/ComfyUI/commit/a464ac33588ae182f81a090d910cfbf21e255b73
>>
>>109485607
hey you just reminded me one of the first things i genned with ltx 2.3 was boxxy going "MY NAME IS BOXXY AAND.. THESE ARE MY TITS" and then lifting her shirt

gonna remake that tomorrow with h3.

>>109485611
lmao
>>
File: AnimateDiff_00178.mp4 (3.96 MB, 1280x672)
3.96 MB
3.96 MB MP4
>>
File: AYAAAAAAAA.png (855 KB, 1280x720)
855 KB PNG
>>109485611
>based on Wan2.1-I2V-14B
>>
Can the reference given to H3 be a voice? (Can't test shit until the weekend).
>>
>>109485620
yes, you can use audio clips as references
>>
https://files.catbox.moe/00a1o2.mp4
>>
>>109485128
Link to Morrigan/the succubus please? I didn't find it in the old thread and can't tell if it is one of the expired link or bake issues.
>>
>>109485627
kino
>>
>>109485628
It's in one of the last three threads if you follow the previous thread link.
>>
What sampler should I use for h3 then?
>>
>>109485470
>20steps res_multi/simple on 306012gb in under 5mins with just sageattention and spectrum
you should try kijais new low vram and feed forward nodes after your sage attention patch and before spectrum and see if you get any speed improvement (there is no quality loss)
>>
>>109485626
Very nice.
>>
https://files.catbox.moe/l0qcq6.mp4
>>
>>109485492
>is there a consensus on the best sage attention node and the best cache node?
best sage attention is kijai's mem eff sage attention afaik

best cache node is either spectrum or the utilscollection node. no clear thread favorite yet
>>
>8 steps

>turbo lora 850 strength 0.65
https://files.catbox.moe/5hn9vh.mp4

>turbo lora 500 strength 1
https://files.catbox.moe/t8ts0v.mp4

I don't like it, they make the video slopped af
>>
>>109485499
i chuckled
>>
>>109485654
I'm gonna become a master lip reader by the end of the H3 era. I was able to get "read the fucking" from walt and "memes" at the end from jesse kek
>>
>>109485654
It's true.
>>
>>109485599
>MiniMax-H3 denoises the video and audio streams on two different flow schedules (video shift 12, audio shift 3). ComfyUI's stock samplers step both streams on one schedule, which is fine at ~20 steps but badly over-steps the audio at 4 steps — the audio comes out distorted or blown out. This sampler steps each stream on its own schedule, so audio stays clean at 4 steps. If you load the LoRA and use a stock sampler at 4 steps and the audio is broken, this is why.
This is not true though???? Changing audio shift while keeping everything else the same changes audio.
>>
>>109485699
true or not it works. Use either the original or 500 one though, the 850 is too overcooked
>>
>>109485639
>three
Ah thanks, gave up too early.
>>
>>109485711
its in here >>109480220
>>
I would not bother with the turbo lora yet. let it cook.
>>
>>109485423
The original repo even gives skill files, but those seem to be for specific styles. I wonder how many users know about either of these things.
>>
>>109485423
>>109485727
lol it's immediately obvious from the markdown that LLMs were used to write the prompt guide
if you're writing prompts by hand for H3 that's pretty crazy
>>
>>109485741
I assume that LLMs have no concept of pacing, and I like having well paced dialogue so I do it myself
>>
8 steps for 850 step lora seems good
https://files.catbox.moe/wfcxps.mp4
>>
File: sad.gif (415 KB, 220x220)
415 KB GIF
>>109485741
>if you're writing prompts by hand for H3 that's pretty crazy
haha... yeah...
>>
>>109485758
what lora strength did you go for?
>>
>>109485705
Ema or non-ema for 500 steps one?
>>
>>109485741
>>109485759
writing a prompt to hand off to an llm right
>>
>>109485756
>I assume that LLMs have no concept of pacing, and I like having well paced dialogue so I do it myself
from my testing it seems to be alright. gotta remember to let it know the duration of videos you're interested in so it doesnt make prompts that are too long or too short
>>
I think the only way Black Forest Labs can win now is if they forget current Flux 3 Dev and distill a new one from their best model, into at most 22b params not including TE/VAE, add NSFW/Copyrighted data to the dataset, and do their best to actually make a good model to publish in order to compete with H3, everything less than that will just be DOA.
>>
>>109485763
1.0, ema
>>
patiently waiting for things to chill the fuck out. in the mean time any ideas on what i should gen in the mean time?
>>
>>109485777
thanks anon, did you test both non-ema and ema?
>>
>>109485778
1girl, asian, huge breasts
>>
850 EMA
4 vs 6 steps
https://files.catbox.moe/f7089e.mp4

>>109485782
ema is far better
>>
Has anyone has done any tests on how many languages the model know and how accurate it is at each language?
I wonder if it can do like Albanian and stuff like that.
>>
>>109485785
>ema is far better
cool, because I've read in the model card that ema was worse at the begining of the training, but I guess 850 steps is the moment when ema takes the cake
>>
>>109485741
>>109485770
You vastly overestimate the IQ, patience, and skill level of the average prompter.
>>
>>109485756
You know you can prompt for the pacing too right? anyways, it's not like you're not allowed to go and edit the timestamps it came up with yourself.
>>
File: 1767279756036306.mp4 (233 KB, 704x608)
233 KB
233 KB MP4
you get what i was going for at least
>>
>>109485785
How'd you make that cool overlay?
I'll touch tips with you if you tell me
>>
First time trying this on my 10 year old laptop. It's actually not that bad. It's fun to keep genning until I got a style I like.
>>
>>109485794
it can do portuguese, so it can probably do most
>>
just deleted ltx and wan
AMA
>>
File: MiniMax_H3_00276_.mp4 (2.46 MB, 736x576)
2.46 MB
2.46 MB MP4
>wondering why the gen didn't use the reference.
>I gave it the wrong image.
>>>/wsg/6209623
>>
>>109485828
looks like GUTS or GANTS whatevert the fuck that animo was named
>>
>>109485828
This style is a crazy hack because the gen artifacts just end up looking like they're meant to be there.
>>
>>109485811
genuinely cute
love it anon
>>
https://www.reddit.com/r/StableDiffusion/comments/1vhloyz/walter_white_and_the_minimax_h3_official/
>>
File: MiniMax_H3_00277_.mp4 (1.89 MB, 736x576)
1.89 MB
1.89 MB MP4
>>109485828
with the actual reference.
>>>/wsg/6209626

>>109485833
GANTZ? I could see that. this is BLAME! tho. tbtbh one or the other could have easily been influenced by the other.
>>
>>109485826
>just deleted ltx and wan
based >>>/wsg/6208736
>>
>>109485870
All right, Damn Mr Anon... I just wanted to generate memes :(
>>
my head hurts from how good this model is
https://voca.ro/1lmz6OUgFbal
>>
File: f.mp4 (882 KB, 864x480)
882 KB
882 KB MP4
>>
>>109485919
link to inappropriate girls in question?
>>
>>109485929
aww that's cute
>>
>>109485929
>shadman
>>
>>109485930
>link to inappropriate girls in question?
i'm rerunning the prompt i changed it a bit because the LLM embellished with some visual stuff that didnt look good e.g. leaving a lipstick mark
i also made their outfits sluttier
>>
>>109485919
i don't know your source audio but i am actually probably expecting even a bit better from omnivoice/fish/moss/...

it's still nice tho, did you get this from h3?
>>
>>109485981
i had no source audio, that was from a text to video h3 gen
>>
Flux 3 max VS Minimax H3 comparison
https://streamable.com/l47iwi
>>
File: 1782469763978869.mp4 (1.94 MB, 800x544)
1.94 MB
1.94 MB MP4
>>109485616
i wanted the camera to stay affixed to her helmet but gave up after two attempts.
>>
>start reading the manual
>its starting to make sense
oh no, its fucking over for me isnt it
>>
https://files.catbox.moe/l5laxv.mp4
>>
File: zzz_001.png (2.13 MB, 1920x1080)
2.13 MB PNG
https://files.catbox.moe/81r3tu.mp4
>>
>tfw I'm still running cuda 12
Fuck why didn't you tell me anon
>>
>>109486006
just wait until your teacher covers writing, that's gonna go crazy in the special ed classroom
>>
>>109486014
>Fuck why didn't you tell me anon
GPT 5.6 luna which costs 20 cents per million tokens told me and did it all for me
>>
File: MiniMax_H3_00230_.webm (3.54 MB, 960x544)
3.54 MB
3.54 MB WEBM
I have so much to meme I feel overwhelmed..
>>
File: IMG_2030.jpg (49 KB, 983x461)
49 KB JPG
Should I be using this node? I’ve been on an ealy H3 workflow and it was fine without. And why do people have different values on this node?
>>
>>109486014
Cu130 has been out for a year anon. Of course newly trained models and designed opts would need or at minimum benefit from it anon.
Do you want a reminder that the sky is blue as well?
>>
>>109486029
kek
>>
>>109486016
i have never read a book in my life, not even a children's book. i still dont know how the hungry caterpillar ends
on the other hand i read documentation and manuals back to back if its autistic enough of a device.
>>
>>109486029
Lel
>>109486031
If you have to ask remove it and forget about it.
>>
>>109486000
really good gen
>>
>>109486032
>Do you want a reminder that the sky is blue as well?
I guess I need one :(
>>
>>109486031
This node is a curse because once you know what it does you'll be tweaking the values for every single gen.
>>
>>109486036
>i still dont know how the hungry caterpillar ends
he eats the entire universe
>>
Any good video upscaling workflows?
I saw comfyui at openmodeldatabase and want to find a simple video workflow but they aren't around.
>>
https://files.catbox.moe/khcen8.webm
Video references is the best human invention in history!
>>
>>>/wsg/6209638
>>>/wsg/6209639
>>
>>109486048
i approve any and all doro posting
>>
>>109486048
kek quicker than me to post my own shit well done
>>
>>109486045
i refuse to believe you and i will not read it to find out if youre bullshitting me or not
>>
>>109486031
It's really hard to tune. Only need to mess with it if you're doing low step stuff.
>>
>>109486013
>saliva trail
kino
>>
my uncle works at BFL and said they just installed suicide nets on the sides of the building.
>>
File: MiniMax_H3_00231_.webm (3.84 MB, 640x832)
3.84 MB
3.84 MB WEBM
Interesting, good way to test how it interprets multiple characters in split image. Motion is pretty much identical. I guess I need a much more in depth prompt, very basic one.
>>
>>109486014
>>109486032
what are the benefits of cuda 13 concretely?
>>
>>109486067
whats the benefit of running newer software?
>>
>>109486073
More tracking and data collection?
>>
File: s.mp4 (1.67 MB, 864x480)
1.67 MB
1.67 MB MP4
>>109485990
oh. that's very convenient then!

>>109485934
ty. needed to give the cube a better ending.
>>
Is there a turbo lora + lora weight + step count + scheduler + shift combo that eliminates the blur without producing artifacts or make it extremely strongly opinionated?
I can't get the turbo lora to work properly.
>>
>>109486062
kekd
>>
https://civitaiarchive.com/models/2839513?modelVersionId=3205117

The wait is over boys. Finally we have the LoRA we needed to gen actual good shit
>>
>>109486031
bypass, forget about it. just enabling it is gonna fuck up the pace of your videos because it does not actually default to 12. using a lower value will cause the actions to speed up in the vid and the longer your vid the more fucked up it gets in the end
>>
>>109486097
For a second I thought she was May Li
>>
Why doesn't the OP have a rentry listing all H3 optimizations? Why are you niggers so lazy? They are not like his in /lmg/
>>
>>109486150
>They are not like his
https://youtu.be/H58vbez_m4E?t=111
>>
>>109486150
No one is stopping you from writing a rentry. We'll add it to the OP if it's good.
>>
File: MiniMax_H3_00234_.webm (3.77 MB, 640x832)
3.77 MB
3.77 MB WEBM
Zoom in and out test for recognizing subjects and memory, very impressive.
>>
File: IMG_2031.png (259 KB, 1810x1695)
259 KB PNG
Anyone here tried using either of these new director nodes? Or should I stick with basic proompting
>>
>>109486167
Those are useless and retarded.
Takes more time for something I can literally just write.
More so when 90% of my prompts are written by my local LLM.
>>
keek
https://files.catbox.moe/ft7mhy.mp4
>>>/wsg/6209659
>>
>>109486186
Quality on par with the latest seasons desu
>>
File: 1779139654322459.gif (1.92 MB, 404x404)
1.92 MB GIF
>>109486167
the point of AI is to not do work anon
>>
For AceStep anon
>>>/wsg/6209660
>>>/wsg/6209661
>>
File: file.png (399 KB, 1099x763)
399 KB PNG
Is this overkill?
>>
>>109486217
Post output and references when you're done and we'll see
>>
>>109486223
I'm talking about all the optimizer nodes at the top. I got the from some dude on reddit but it seems like overkill and they could be working against each other on some level.
>>
>>109486217
No is not.
A good varied character video ref at high res is way better than pics and basically perfect but it also tanks performance like crazy, x2 or more.
>>
>>109486225
Was mostly curious how good it took several references desu. Haven't tried with more than 2
>>
File: optimying.png (20 KB, 1003x312)
20 KB PNG
big speed gains inc
>>
>>109486217
can't wait to see the results, thank you for exploring the frontier
>>
With Krea2, is it possible to gen anime-style images with both a man and a woman where the man has the same or a slightly lighter skin tone? I mean pasty white guy X white woman, not white guy on black girl.
Basically what I'm asking is whether there's a way to prompt around the inherent bias in the dataset where the guy usually has darker skin when he's shown touching the woman.
>>
VRAMlet (8GB) and RAMlet (16GB) status? Will I be able to gen funny memes?
>>
File: MiniMax_H3_00237_.webm (3.67 MB, 832x640)
3.67 MB
3.67 MB WEBM
Adding a new character while retaining the original style and lighting of the input frame, very impressed.
>>
>>109486256
>--yuse-sage -attinitin
kek
>>
>>109486271
>very impressed.
yeah it's really good at inserting a character in the same style as the one from the input image >>>/wsg/6207411
>>
>>109486271
Real video of Me in my Apartment
>>
Is the turbo H3 lora only for fl2v model or also supports ref2v?
Or it’s shit either way and I shouldn’t use it
>>
File: MiniMax_H3_00238_.webm (3.81 MB, 544x960)
3.81 MB
3.81 MB WEBM
So fucking cool. I know sd2 etc can do this shit, but now it's local.

>>109486285
That's with the reference model? Regular irl photo of costanza, prompted to use the style of another?
>>
File: h.mp4 (1.65 MB, 1280x720)
1.65 MB
1.65 MB MP4
>>109486217
you decide if you need it, but it can work. there were 9 ref image + prompt demos even with the early access users
>>
sloooooow moooootttiiiooonnnnnn
>>
>>109486031
i think a shift_video value of 8.0 makes the movement more rapid, while 12.0 makes the movement more chill.

being too rapid isn't a good thing because it can look unnatural and shitty, but being slow can be boring. I've set it to 10.0 now
>>
so whats the duration limit exactly
>>
>>109486323
>That's with the reference model?
no, I2V, the model already knows costanza
>>
>>109486335
15 seconds officially, I've had it up to 30 but it loses all coherency after 40.
>>
File: 4u65u4652.jpg (24 KB, 249x427)
24 KB JPG
Erm... Can't really tell you what's up other than I have a 2 minute reference video for my r2v for one of them, and two parallel r2v happenings.
>>
>>109486347
>I have a 2 minute reference video
I hope you have at least a pro 6000 and the video it's scaled to 240p
>>
File: 1780275608256187.mp4 (3.32 MB, 544x736)
3.32 MB
3.32 MB MP4
>>
>>109486370
Indeed, took a while to chunk, it was about to be a 100 minute wait, so I stopped that, yes shorter vids.
>>
File: MiniMax_H3_00243_.webm (3.82 MB, 672x800)
3.82 MB
3.82 MB WEBM
Which of all the social medias is good for farming ai memes for views?
>>
>>109486312
Last time I tried it made the output on ref look worse than without it, so probably. Though it's weird that nobody specifies it.
>>
>>109486346
so what youre saying is i need to chain several videos together
>>
>>109486390
>MiniMax_H3_00243_.webm
ah fuck thats a really good one kek
>>
File: trump bibisi.png (461 KB, 736x432)
461 KB PNG
>>109486390
Truth social
>>
>>109486393
that depends, but a single shot can't be more than 30 or so seconds. I'm sure it's simple to stitch shots together, character consistency is a solved problem
>>
>>109486390
x is good for it.
For bigger accounts to farm you for millions of impressions while you get none that is.
>>
I don't get it, I'm on 16gb vram/32gb ram, and everyone seems to be running the int8 h3 model. Meanwhile when I run it something hits swap and gen speed falls off a cliff
>>
>>109486439
Does your hardware accelerate int8?
If this is intel/amd, did comfy merge dynamic memory shit for them?
Do you have any fancy cli args that might interfere?
Lastly you are a LinuxGOD and using ZRAM, right anon?
>>
File: MiniMax_H3_00245_.webm (3.73 MB, 640x832)
3.73 MB
3.73 MB WEBM
Huh, it doesn't know hitler, shame. Easily fixed with ref model.

>>109486407
Good point, watermark my shit.
>>
File: 1781684288046439.mp4 (3.58 MB, 1920x1088)
3.58 MB
3.58 MB MP4
>>109485930
>link to inappropriate girls in question?
i hope you now see why i only posted the audio the first time
this is a pretty common fantasy for highschool girls too
>>
Thread based beyond belief
>>
>>109486454
>>109486458
these need to go on /r/stablediffusion asap
>>
This model has brought gooning to 2D to a whole new level.
>>
>>109485128
How did this stuff get so good
>>
>>109486453
>Does your hardware accelerate int8?
I think this might be it, apparently the venv runs on an older cuda build that runs int8 in software. Thanks for that.

Other than that amd/nvidia, no and I really should be
>>
File: MiniMax_H3_00246_.webm (3.62 MB, 608x864)
3.62 MB
3.62 MB WEBM
>>109486467
omg yaaaas
>>
assuming youre not using the turbo shit, and arent using sage attention, whats the expected gen time for 4:3 0.4mp thats 15 seconds long?
>>
>>109486485
thank the chinese
>>
File: image_2026-08-07_10-30-41.png (895 KB, 3631x1672)
895 KB PNG
>>109486315
ok so it seems to work, and pretty fast too.
3 ref images, 0.4mp (768x640) 15 seconds clip takes 230s on my 5080
>>
Patch Sage attention never works for me. Am I retarded? Flash attention v2 works for me. I'm on AMD
>>
wow, h3 is able to reconstruct impossible 2D hentai body proportion into 3D, which WAN fails to do.
https://files.catbox.moe/ovigqf.mp4
>>
>>109486512
oops, meant for >>109486392

no sage, no torch patching
>>
https://files.catbox.moe/5kyksf.mp4
>>>/wsg/6209683
>>
Just a heads up I got kino in the oven.
>>
File: MiniMax_H3_00023_.mp4 (495 KB, 576x736)
495 KB
495 KB MP4
Feels like my gens are especially low res with ref2v, what do you think?
0.4 mp, sage attn 2, res_multistep, 20 steps. nothing else special, just seems like everyone elses 0.4's are better looking.
>>
>>109486585
Try beta scheduler
>>
File: 1754660081300794.webm (2.18 MB, 704x704)
2.18 MB
2.18 MB WEBM
im dying
>>
>not tapping
respect to that alpha male
>>
>>109486603
one of best gens i've seen so far
topkek
>>
File: MiniMax_H3_00024_.mp4 (305 KB, 576x736)
305 KB
305 KB MP4
>>109486589
thanks, posting in case anyone else wants to judge the difference.
>>
>>109486523
>the thigh pubes
>>
Anyone using this wf?
https://civitai.com/models/2834514/minimax-h3-t2v-i2v-ref2v-advanced-filmmaking-workflow-or-all-speedups-qol-features

getting a 5sec gen done at 437 secs
3090
6 steps
1.4mp
using turbo lora at 1.40 and another lora

UPDATE
enabled sol-attn and the gen was done at 384.93

noticed that the sound is very bad.. I added a nsfw lora after the turbo lora, unsure if related
>>
File: MiniMax_H3_00255_.webm (3.81 MB, 736x736)
3.81 MB
3.81 MB WEBM
>>
>>109486665
Why is demon Putin giving me weed for free?
>>
Any tricks to getting dialogue timing right? Usually the dialogue starts after the actions that should happen after it in the prompt
>>
>>109486679
And apparently he has no midsection.
>>
while i like the speed, turbo lora makes skin look diseased and also adds moles and stuff
shift 12/4, euler/beta, 8steps
>>
>>109486703
are you using time codes for both?
>>
>>109486709
To be fair that's selling the demon look more than a starving African with glued-on deer horns
>>
>>109486735
So far I'm only using time codes for shots. I guess I'll try to manually time literally everything
>>
>>109485128
Where's that ass video? Can somebody reupload it?
>>
Can you extend videos with h3?
>>
best way to get anima to generate JUST pictures of clothing items? It always wants to put a person in it.
>>
File: 1763549322288977.webm (3.85 MB, 864x480)
3.85 MB
3.85 MB WEBM
This is fucking insane. Literally grok tier.
>>
It feels like my gens got slower today. Like I'm not loading as many layers of the model into VRAM as before.
>>
>>109486825
Mine definitely are after I realized yesterday that my GPU was dying. The temperatures were relatively okay but I hadn’t noticed that the hotspot was going over 100°C during generations because HWiNFO wasn’t showing it for my card. I had to undervolt it and adjust the power and fan settings quite a bit.
>>
File: MiniMax_H3_00257_R.mp4 (1.7 MB, 672x1216)
1.7 MB
1.7 MB MP4
testing turbo lora. it seems like the prompt adherence is weaker with the turbo.
>>
>>109486876
lol 100c hotspot is fine
a 5090 has a 130c hotspot at 50% tdp
>>
>>109486886
I had some other symptoms before that as well like crashing which didn’t seem to be an OOM issue because it went away immediately after I power-limited the GPU.
>>
>>109486217
https://i.4cdn.org/wsg/1786097011798158.mp4
>>
>>109486901
that's a blitzball sphere
>>
File: MiniMax_H3_00116.mp4 (1.76 MB, 736x416)
1.76 MB
1.76 MB MP4
>>
Also, gimi
>>
File: 1757664439663152.png (345 KB, 864x480)
345 KB PNG
>>109486913
The sphere was one of the reference images and sometimes it genned it correctly, sometimes not, in all gens it was referenced directly in the prompt.

The magic of seed numbers i guess.
>>
is there a proper dialogue format? when I add some "dialogue line" the character then won’t shut up and continues blabbering nonsense
>>
I'd like to use another browser to launch comfyui than my default.
Anyone? The edited "Main.py" from Github isn't doing it.
>>
>>109487029
there is an easier way to do it but i don't remember what it is.
>>
>>109487009
I'm experiencing the same problem. Lots of gibberish. Also hard to control which caracter says what.
>>
File: 1777681174005447.jpg (63 KB, 574x688)
63 KB JPG
>https://civitai.red/models/2840051/speculum?modelVersionId=3205793
https://civitai.red/models/2840051/speculum?modelVersionId=3205793
>https://civitai.red/models/2840051/speculum?modelVersionId=3205793
https://civitai.red/models/2840051/speculum?modelVersionId=3205793
IMAGINE THE 20 SECOND LONG VIDEOS MADE AFTER
>>
>>109487029
just type http://localhost:8188 in the other browser
>>
>>109487046
yeah, gemini solved it.
I am sad, at one point we won't have anything to say to each other if we keep asking to fucking robots.
>>
>>109487029
Just copy the address and paste it in another browser.
>>
>>109487009
Is your video long enough to fit the dialog?

Ref video is better for this with a very specific syntax.
>>
>>109486497
Stop using "megapixels", no one ever used that term in computer graphics except r-eddit tier normies when they got their first digital cameras.
Comfyanonymous (he has literal zero real world experience in digital image manipulation) and his merry folk of slop retards are so incredibly stupid for trying to normalise anything like this.
>>
>start saving my dogshit qwen enhanced prompts
>3 million characters in length
shan't scroll past or read any of that
>>
>>109487087
*group
>>
>>109487083
Yes but how do i generate a reference video that contains the correct dialogue?
>>
>>109487094
Read the guide!!

You have to list all the speakers in the video and assign them ids in the speaker section then when prompting them to speak in the video description you have to put
<Subject #> (S#) <d>[English] TEXT </d>
>>
>>109487104
Okay, I will read the guide, but shit's hard to remember. Lots of rules in that guide.
>>
fuuuck I think GGUF can reach a permanent state where it doesn't load into vram in ComfyUI (instead wants to load into ram).
I tried a thousand different things: multiple vram cleaners, vram debuggers, closing lots of programs, restarting the workflow, cutting away at the workflow, restarting the browser, and the ONLY thing that worked was restarting the PC.

It's only a 16GB gguf file while I have 24GB 3090 RTX.

The reason I'm willing to blame gguf is I've heard the guy in charge of ComfyUI is a little bitch about gguf and ignoring it, so it likely has these kinds of bugs.
>>
>>109486390
Instagram reels 100%
>>
>>109487104
will try that, thanks. but I think most of this is placebo as QwenVL 32b should be smart enough to handle your garbage prompt with any structured format
>>
Would upgrading 64gb to 128 increase speed or just get rid of OOMs?
>>
Sigma shift? Are we still doing that?
>>
>>109487169
>getting rid of OOMs
lol. Nah. I think there's memory management issues that throwing more RAM at won't solve.
>>
>>109487183
We have been doing that non-stop last two years.
>>
>>109487131
GGUF is much slower than int4convrot from my experience.
>>
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>writing boomer prompts gives the best results
kek
>>
File: 1760638631156925.webm (3.9 MB, 1044x1034)
3.9 MB
3.9 MB WEBM
>>109487305
>>
anons update spectrum if you use it cause it's better now

New default settings
degree = 1
warmup_steps = 1
bootstrap_first_forecast = true
tail_actual_steps = 1
>>
>>109487104
>>109487149
Reporting back. This format works, no more additional gobbledygook
>>
>>109487337
>soulful video vs corposlop (ours)
>>
File: MiniMax_H3_00066_.mp4 (2.23 MB, 864x480)
2.23 MB
2.23 MB MP4
>>
>>109487339
heres a test:

https://files.catbox.moe/u8bua0.mp4
>>
Updated Comfyui and sol-attn and now I just get OOMs trying to generate 17 seconds video using references.
>[WARNING] [MiniMaxH3-SolAttn] kernel 执行失败 (OutOfMemoryError: Allocation on device 0 would exceed allowed memory. (out of memory)
>>
>>109487340
What is (S#)?
>>
>>109485998
So flux is just sloppa generation, all here is slop like fuck and have the same flux girl face
>>
For some reason, gen speed heavily depends on reference resolution even dough resolution of actual gens remains the same
>>
>>109487393
# = number so S1... S2 etc which is the Subject
>>
>>109485998
Okay but can flux do bobs and begana when it's not even local?
>>
How do I use reference videos? Just the load video node doesn't seem to work, unless it can't take .mp4s or something
>>
>>109487422
Use the get video componements node
>>
>>109487385
Both Sol implementations are still highly experimental.
Can't say much besides roll both comfy and the extension back or wait until it gets patched.
>>
what a time to be alive. 15s in 218 seconds with the new spectrum update and patch sage kj (auto), 0.3mp, no turbo

The setting is the TV show South Park.

0 to 5s: the 4 characters Cartman, Kyle, Stan, and Kenny are standing together at a bus stop, in front of a widescreen TV that is off. Cartman says "guys, why do THEY control everything? You know what I mean Kyle.". Kyle, in his classic green hat says "no, I don't.".
5 to 10s: Cartman turns on the widescreen TV that shows a chart that says "Jew ownership in Blackrock", with a pie chart that shows "95% Jewish". Cartman says "see THIS is why I can't have good video games, Kyle."
10s to 15s: Cartman presses a button and the TV shows a new chart saying "Shekel Shekelberg", with a South Park style rabbi beside the text. Cartman says "they are ruining it all Kyle, I just want to play my videogames."

https://files.catbox.moe/fo2psg.mp4
>>
>>109487339
>>109487481
how much faster?
>>
>>109487337
>>109487349
for real lol, they just made a sloppifier LoRA
>>
>fresh copy of comfy
>python -m pip install triton-windows
>3.13.14 python libs and include inside comfy
>sageattention 2.2.0 cu130torch2,10.0andhigher.post6
did I install it correctly? I'm on 3090
>>
Can someone redpill me on sol diffusion? Is it better than sage? Can or should they be used together? Any side effects?
>>
>>109487490
"In that first post I released the MiniMax H3 Spectrum integration and was getting around 34% lower Euler sampling time and 30% lower RES sampling time with the more conservative settings I was using at the time."

ive only tested a couple gens but it does seem faster than the old one, update/git pull it if using spectrum, it has new default settings, pick vram if you have 12gb+
>>
>>109483968
>>109483968
>>109483968
>>
No collage no migrate. Kys, troll.
>>
>>109487536
don't migrate to the spaghetti



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.