[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Previous: >>109485128 >>109488892
https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/07/2026

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo – OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view

>OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
https://xin1u.github.io/OminiVR_PAGE
>>
>gens getting noticeably worse as people pile on the latest cope nodes
>>
>mfw Research news

08/07/2026

>Vorch-Omni: Multi-Task Orchestration of Sight and Sound
https://vorch-project.github.io/Vorch-Omni-project

>Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
https://vorch-project.github.io/Vorch-Streamer-project

>Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
https://vorch-project.github.io/Vorch-Director-project

>Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
https://vorch-project.github.io/Vorch-IR-project

>In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
https://arxiv.org/abs/2608.05237

>Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
https://arxiv.org/abs/2608.05976

>EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
https://arxiv.org/abs/2608.06231

>MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
https://arxiv.org/abs/2608.05878

>Wan-Animate-2: Pushing the Application Boundaries of Character Animation
https://humanaigc.github.io/wan-animate-2

>StyleComposer: Training-Free Multi-Reference Style Composition
https://lexxsh.github.io/StyleComposer

>Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
https://arxiv.org/abs/2608.06125

>Adapting Vision Foundation Models with Cascaded Semantics
https://xixiaouab.github.io/Cascaded-Semantics

>Learning visual representations for compositional analysis of artworks and photographs
https://arxiv.org/abs/2608.06142

>MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
https://arxiv.org/abs/2608.05450

>Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
https://arxiv.org/abs/2605.16411
>>
>kijai
>>
>>109490719
Let anon experiment. Eventually the dust will settle and there will be one community preferred standard, as well as thirty five other voodoo cope schizo workflows that one anon swears is better. Natural progression of thread.
>>
File: skyla_00016_.png (1.28 MB, 1536x1024)
1.28 MB PNG
https://d.uguu.se/zcKFDUuB.mp4
>>
>>109490719
it's just experiments, I guess you just learned how that thing goes, no experiment goes linearly and smoothly towards better quality, sometimes there's bumps, you need that to improve
>>
>>109490718
>>109490723
fuck off loser
>>
File: 847737372.jpg (187 KB, 952x614)
187 KB JPG
please help anons. surely this cannot be the right speed. I'm on 3090 + 64gb ram, and this is the speed I get for 1mp 10 sec video with default workflow + sage + spectrum. I don't see any errors in comfy, and everything is installed properly. Fallback is also disabled in nvidia
>>
>>109490713
finally a proper collage bake.
>>
>>109490713
thanks for the bake anon, godspeed
>>
thanks for bakering
>>
File: AnimateDiff_00218.mp4 (1.82 MB, 672x1280)
1.82 MB
1.82 MB MP4
>>109490732
All I see its progress, video extension, story telling, editing, composition

Some anons are just salty because they cant generate good frames or images references, let alone have resources to experiment, at the end your videos are good as your image gens
>>
So, without the turbo lora, how many steps should I use?
>>
>>109490737
Can you make milk come out too with this model?
>>
>>109490761
>turbo lora face
>>
>>109490763
Oh, and which scheduler
>>
I would post my kino H3 gens here, but they've gotten to the point where they're too kino to share, and I don't want to waste compute on meme gens just for some (you)s. Not sure what the solution is
>>
>>109490763
20 min, 30 is better, 50 is api
>>
>>109490774
>50 is api
it's actually 40 (and 720p) but yeah, not many differences between 40 and 50
>>
>>109490774
api makes it overkill tho.
>>
File: 1762559211524500.png (84 KB, 498x305)
84 KB PNG
>>109490771
it's ok anon, we understand
>>
>>109490774
Okay, I'll try 20 first because it's probably gonna take real fucking log
what about scheduler?
>>
>>109490794
just go for the regular comfyui official workflow, it works fine
>>
>>109490763
I'm happy with 15
>>
>>109490771
same, but mine also have copyright shit and i REALLY don't want to upload anything which could possibly be used as ammo for regulations and bullshit
>>
File: i2v_MiniMax_H3_00060_.mp4 (2.07 MB, 704x1056)
2.07 MB
2.07 MB MP4
https://files.catbox.moe/nbmpes.mp4
>>
>>109490794
res multi
simple (some say beta)
if you switch to turbo, use euler because res multi will create artifacts
>>
File: AnimateDiff_00222.mp4 (2.05 MB, 672x1280)
2.05 MB
2.05 MB MP4
>>
ext week diffNews start!

Forget Reddit, 4chan, dozens of Discord servers, and GitHub repositories.
A site that uses AI to monitor all these channels and, through a “human-in-the-loop” process, publishes high-quality news, tips, and tutorials online the very same day.

* I'd write that if I weren't so lazy when it comes to business
>>
You can just say your gens are shit and you're embarrassed to post them you don't have to make up some story about whatever as an excuse
>>
>>109490718
>>109490723
Stop spamming this shit
>>
>>109490809
>Lie-nix
>>
File: 1758556974640742.jpg (495 KB, 1920x1088)
495 KB JPG
>>109490761
>>109490812
>at the end your videos are good as your image gens
try to do some pure text to video gens with your girls anon, i am interested in seeing how H3 does
>>
>>109490761
Oh nice it can do some facial expressions. Quite a leap from wan
>>
>>109490825
it got it right half the time
>>
>>109490812
>>109490761
>proceeds to keep posting overcooked ugly 1girl walking gens
>>
Beautiful
https://files.catbox.moe/r9g8kc.mp4
>>
>>109490809
you guys have finally your smoke animation maker.gg
>>
>>109490809
so this is how you pronounce linux huh?
>>
>>109490834
>can do accurate 3d camera movements on pixel style
this model is really amazing
>>
>>109490809
quick tip, but for words it can't pronounce correctly you can do stuff like write "Lee Nux"
>>
File: orodSh_00231_.jpg (1.11 MB, 1776x2560)
1.11 MB JPG
>>
File: 1758010040666855.png (35 KB, 678x102)
35 KB PNG
>>109490861
https://files.catbox.moe/rqkt4g.mp4

why stop there? this works
>>
https://litter.catbox.moe/b0il5trp973hr06e.mp4
>>
>>109490877
kek
>>
>>109490872
I use that phonetic shit for vibe voice for a few words it consistently fucks up. It works really well.
>>
>>109490872
technologia
>>
https://files.catbox.moe/dnyg17.mp4
>>
>>109490745
1mp is actually a pretty large res to gen at, what are your speeds at 0.4mp?
>>
File: MiniMaxH3_00016.png (1.21 MB, 1200x1200)
1.21 MB PNG
fuck it, posting the gens even if i can't stop the gibberish.

https://files.catbox.moe/zz23wg.mp4
https://files.catbox.moe/j0vpnk.mp4
https://files.catbox.moe/08r533.mp4
>>
STOP USING THE TURBO LORA
>>
>>109490885
The gen is pretty good, but at least offer some minor insight on what you are doing.
>>
File: SS_02572.png (58 KB, 1240x649)
58 KB PNG
ref_image_0 is always <Picture 1> when prompting, right? And ref_audio_0 is <Audio 1>?
>>
>>109490889
that's how he talks just after a plastic surgery so that's accurate kek
>>
>>109490895
I think spectrum is better cause there is less of a quality and prompt understanding dropoff
>>
>>109490898
Yes
>>
>>109490898
yes
>>
>>109490895
What are all the other speedup options?
>>
>>109490908
use spectrum or minimax cache.
>>
>>109490888
around 2 minutes for 0.4mp and 5 sec. But some random ass redditors can gen 1.0 and 10 sec fine in like 8 minutes with the exact same setup as mine.
https://www.reddit.com/r/StableDiffusion/comments/1vhyl33/minimax_h3_compare/
>>
File: 1774884150283064.webm (2.92 MB, 864x1344)
2.92 MB
2.92 MB WEBM
Works well with 6 step.
But i cant solve those Choppy anime animation with flat anime shading. The more 3d looking your image is, the less chance it got choppy animation.
>>
File: MiniMax_H3_00056_.webm (3.8 MB, 1376x768)
3.8 MB
3.8 MB WEBM
I've been playing around a ton with sd2 and having a pretty big fraction of its power in local is amazing.
H3 followed every part of my prompt and it genned it pretty much exactly as I had envisioned it.
>>
>>109490809
make her do the Kramer thing where she drinks while she still has the cigarette in her mouth
>>
>>109490895
>STOP USING THE TURBO LORA
I wanted to put the "stop having fun" meme to clown you but yeah the lora is pretty bad at its current state, they've only begun the training
>>
>>109490921
>H3 followed every part of my prompt and it genned it pretty much exactly as I had envisioned it.
I guess using a 32b text encoder is useful yeah lol
>>
so are first worlders still not allowed to use H3?
>>
Why should I use this :
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

instead of my usual gemma 31B?
>>
>>109490897
emmm. r2v model with the images of the character + image of a stand + PoV image as a guidance for style and the place of the video + prompt describing what's going on in the video. It actually listens if you tell it to only use images as a guide, not to directly copy them.
https://files.catbox.moe/cdsnzt.mp4
>>
>>109490929
The thing is the alternatives are just as fast so it's not like you're giving up anything.
>>
>>109489307
gonna need more like that
>>
File: 1765561576960228.webm (2.39 MB, 864x1344)
2.39 MB
2.39 MB WEBM
>>109490918
lol
>>
File: file.png (20 KB, 535x167)
20 KB PNG
>>109490939
fucking useless
>>
>>109490939
You absolutely shouldn't. Gemma is leagues better for prompt writing.
>>
>>109490939
>he needs a model to write his prompt
>>
File: file.png (4 KB, 277x57)
4 KB PNG
>>109490940
i was there that year
>>
>he doesn't use a model to write his prompt
>>
>>109490921
pretty
>>
File: 1784993469498710.webm (3.92 MB, 544x800)
3.92 MB
3.92 MB WEBM
What is the best way of upscaling/refining minimax? I'm new to videogen, skipped basically the entire wan era.
>>
>>109490944
I forget to mention. The larger your image resolution is. The better the gen results. I gen this with Anima and upscale it to 2K resolution
>>
>>109490948
OK thought so, thanks anon.

>>109490951
Of course I do, I convert my messy thoughts into best practices prompting.
>>
>>109490916
>around 2 minutes for 0.4mp and 5 sec
Sounds about right without speedup nodes, I wouldn't necessarily believe that's the only thing different between that ledditor and your setup.
>>
File: 1768382862349188.webm (3.92 MB, 544x800)
3.92 MB
3.92 MB WEBM
>>
>>109490941
nah, if they manage to make a great turbo lora that works as expected (4 steps) it'll be sooo fast
>>
>>109490961
RTX Super Resolution
Seedvr2
Topaz
>>
>>109490944
>>109490962
I still don't get how this works
so the ref image goes into the ref_image_0 or first_frame input without any downscaling done to it? and the width/height inputs are used from the same image but after downscaling it?
>>
>>109490981
>sneedvr2
>>
Does H3 know what a fart is?
>>
>>109490998
I'm gonna find out.
>>
I think H3 has some kind of distorted training on porn. I'm trying to gen a video of a girl grinding on a dude's underpants, and it keeps producing strange meat that looks vaguely like a penis. I'm not exactly sure how to prompt it out because I don't know what the model thinks it is.
>>
>>109490970
but I don't know what could be wrong... Sage is working, because when I disable it, it gens slower. I even tried a fresh comfy installation, and it didn't help. Can something outside of comfy affect it?
>>
>>109490961
rtx super resolution + starlight precise if you maximum quality
>>
>>109490980
yeah. I'm not saying never use it. just use it when it's actually better than the other options.
>>
>>109491009
yes if you prompt for penis it'll generate a mangled meat sausage, it is known
>>
>>109491009
"He takes his tube worm and inserts it into her conch shell"
>>
File: 1762242536500331.jpg (1.02 MB, 1248x1824)
1.02 MB JPG
>>109490981
>>109491012
Thanks.
>>
File: 1784123826264942.mp4 (1.48 MB, 544x960)
1.48 MB
1.48 MB MP4
>>
>>109491011
Only thing I can think of is if you're having to load the model in and out of memory constantly, maybe a lowvram flag is that slows you down is enabled. Shouldn't be happening with 64gb of ram and a 3090, although it will fill them up as much as possible
>>
>>109490921
kino
>>
>>109491022
Well, I did not prompt for a penis.
>>
>>109491027
that is insanely fucking hot
>>
>>109491012
>starlight precise
Aren't these cloud only, or did people share them?
>>
there is nothing like the reference model desu

https://files.catbox.moe/fb9791.mp4
>>
>>109491039
yeah but you prompted for a situation where a penis would show up more often and very likely has remnants in the model's activations when prompting for a girl grinding on a guy
>>
I didn't know. What the fuck is a --fast-disk? It seems it just stops comfy from eating all my 64gigs of RAM. Same gen speeds btw.
Anything I'm missing?
>>
do any of the h3 turbo loras work for the reference 2 video model?
>>
the Floyd shit is so fucking unfunny jfc
>>
Man I need to get back to image gen. I've been genning h3 videos from my library of past gens and collected official art, but I'm beginning to run out of good candidates (or rather I can't be bothered to keep searching).
Yes, I'm going to start generating base images specifically for video gen. All I need is the girl.

https://www.youtube.com/watch?v=6sMwvcJVKmQ
>>
>>109491073
yeah it's pretty repetitive
>>
>>109491059
looks like I2V, how did you manage to get an image as the first frame on the reference model?
>>
>>109491059
if only they also trained a turbo lora for the reference model
>>
>>109491059
can you post new gens instead of the same ones over and over again?
>>
>>109491049
with audio
>>>/wsg/6209947
>>
Starting to think it might be for the best if a real NSFW checkpoint of H3 doesn't come out. I might die of dehydration.
>>
>>109491065
>What the fuck is a --fast-disk? It seems it just stops comfy from eating all my 64gigs of RAM.
it means it's using your SSD instead of your ram, don't do that
>>
>>109491073
you will get 6 more years of rehashed floyd memes and you will be unhappy
>>
>>109491065
https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/cli_args.py#L181C176-L181C176
>Prefer disk-backed dynamic loading and offload over unpinned RAM. Can be faster for users with fast NVME disks.
>>
>>109491090
>>109491049
why are you replying to yourself?
>>
>>109491073
george droid was the last time it was funny, and even then it got run into the ground
>>
>>109491090
gaht damn those are some heavy sounding titties. thank you for this gift my brother.

gonna throw her into an edit model, make her nude, and remake this.
(and give her bangs)
>>
>>109491022
Also the word crotch can summon it as well.
>>
223 seconds with spectrum at 0.3mp (fine for south park)

pretty good even without timestamps kek

https://files.catbox.moe/9ymlbr.mp4
>>
>>109491094
>>109491098
Well... I guess I shouldn't rape my shitsung nvme. Thanks.
>>
>>109491099
you got me
>>
>>109491099
you also got me
>>
>>109491079
I tried this, might be revised slightly from that one though


the setting is the forest from <Picture 2>, both characters are sitting in the wooden cart from <Picture 2>.

0 to 3s: the black man in <Picture 1> falls from the sky into the wooden cart.

4 to 10s: the blonde character in <Picture 2> says "by Talos...where did this dark skinned elf come from?" in a swedish accent. the black man in <Picture 1> says "ah sheet, where is the fent at, nordic brotha?" in a black man's accent. A dragon from above shoots fire at the cart and lights <Picture 1> on fire. <Picture 1> says "I cant breathe!"
>>
>>109491090
this is Wan or LTX tier tho, static in frame girl, not exactly pushing the capabilities of H3
>>
File: AnimateDiff_00228.mp4 (2.52 MB, 1216x896)
2.52 MB
2.52 MB MP4
>>
>>109491126
BOO HOO NIGGA nobody cares, i'm about to go into a cum coma over here.
>>
>>109491132
>animatediff
into the trash
>>
File: 865378472727.jpg (254 KB, 957x299)
254 KB JPG
>>109491034
>maybe a lowvram flag is that slows you down is enabled
nope. Only using sage attention flag.
>>
>>109491138
does she have downs?
>>
>>109491148
girls with downs fuck the best. just look at greta thunberg.
>>
>>109491132
i loved that show
>>
>>109491148
if she's downs i'm downs.
>>
Can you run h3 as a two pass, I tried it and the nodes doesn't seem made for it yet. I get so much better results at a lower mp, it's annoying. Running it through as video ref keeps all of the artifacts.

>>109491132
Holy shit anon, we had the same idea.
>>
>>109491164
last frame to video?
>>
>>109491126
You're not getting that kind of car motion and audio with those
>>
>>109491164
I didn't expect that. Nice.
>>
>>109491150
>greta thunberg
yeh well she is an exception to the rule regarding that, plus she has the background to make good gens with.
>>
>>109490713
Anyone on a 5090 track gen time variation between gens? I can get 12 sec at 1.0 in 15 minutes, then 8 then 7, then 15 minutes again.
I just did a regen using the seed of a job that was 7 minutes and I got the same speed if gen as the original
>>
>>109491109
>0.3mp
that's the most impressive part of the model imo, it's surprisignly coherent at really low resolutions
>>
>>109491179
Ref model is incredibly powerful. But still a lot of rng with words and subtitles.
>>
so the latest comfy update breaks h3 cache, spectrum, and sol attention? why
>>
sulphur up to 8k funding now.
>>
>>109491191
Sounds like you might be spilling into page sometimes?
>>
File: AnimateDiff_00230.mp4 (2.28 MB, 1216x896)
2.28 MB
2.28 MB MP4
>>109491141
sorry, using other people WF makes me wanna vomit, same as those vibecoded h3 custom nodes, like h3 director, seems to me that some users just take advantage of the hype to push out their useless crap and inflate their egos
>>
File: image.png (358 KB, 1628x1583)
358 KB PNG
>>109491210
I cant figure it atm. 12 sec in 7-8 minutes is great, but it randomly doubling sometimes is weird.
>>
Alright guys, I'm testing R2V to see how well it can synchronize dance movement to an Audio track.

https://n.uguu.se/DmlKPfvt.webm
https://n.uguu.se/RHvEtqJJ.webm

What do you think? Does the subject have a sense of rhythm?
>>
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

apparently this helps also?
>>
>>109491242
Its good. Just needs hot glue
>>
Having fun using Minimax to animate lewds https://n.uguu.se/utKycChU.mp4
>>
man I get used to better shit way too quickly
I'm already like
>UGH when do we get a model where I can just use natural language instead of this <subject 1> and <picture 1> shit...
>>
>>109491132

I remember your old renderings, those were great. Those about that 90s show with the goth girl.
>>
>>109490872
https://files.catbox.moe/5o0fah.mp4

That gave me an idea. Opening lines of the Iliad in what the clanker thinks is close to old Greek IPA.
>Sing, O goddess, the destructive wrath of Achilles, son of Peleus, which brought countless woes upon the Greeks
Ofc i have no idea on how to judge this.
>>
>>109491254
why would you animate censored garbage
>>
>>109491245
after some reading, it works in tandem with spectrum:
>Why it Complements the Spectrum NodeYou might wonder why you need this if you already have the Spectrum Node running. They actually target two completely different bottlenecks:Spectrum Node: Uses Chebyshev forecasting to skip entire sampling steps completely.FirstBlockCache Node: Speeds up the steps that you do calculate by stripping away their heaviest internal layers.When stacked together, they form a highly effective optimization tag-team.
>>
>>109491242
needs some cum
>>
>my 3090 running at consistent 99c for days on end now
I can't be doing that in this market
>>
>>109491260
the model already is excellent at understanding, reference model has those tags cause it works differently, once you define stuff you prompt as usual

also no other model can do what the reference one can do. it essentially replaces loras, kinda.
>>
>>109491277
Show us a good gen then anon :)
>>
>>109491260
One of the many reasons flux 3 dev is going to mog minimax
>>
>>109491226
i get the same thing, i've never found a reliable solution. i think comfy mismanages system ram occasionally and either gets stuck in a loop trying to sort itself out or just says fuck it and dumps to the page file. it generally happens to me after long sessions, i can manually flush my system ram and it will act normally for awhile.
>>
>>109491283
Use msi afterburner, put a temp limit, learn how to undervolt, jesus anon, don't treat such an expensive card like shit...
https://youtu.be/SIlXT32fOMk?t=76
>>
>>109491283
repaste it anon, I know it's annoying but 3090 are a pain in the ass to obtain with current prices
>>
>>109491286
you literally have the tools to uncensor and you animate a shitty censored image from some jap artist
>>
>>109491287
flux 3 dev will be a slopped cucked piece of shit, like every fluxs before
>>
>>109491260
what do you propose as the "better" natural language alternative to that that is just as good at reducing ambiguity?
>>
>>109491254
more norma. add nangong also.
>>
>>109491277
Why would I want to see uncensored cock? I'm not a faggot.
>>
>>109491287
>censorship company allowing a reference model where you can do anything
nope.
>>
>>109491286
bruh there's gozillions of fan made hentai images in rule34 or danbooru, you have plenty of choices
>>
I'm not getting my hopes up for flux 3. I didn't have any hope for h3 and it turned out great.
>>
Damn, almost managed to do a fat hitler.
>>
>>109491296
Why so focused on dicks bro?
>>
>>109491306
>I'm not a faggot.
You most definitely are
>>109491314
I'm not the one who animated a pixel cock.
>>
It's too hot to gen
>>
>>109491295
why should anon repaste that card when it was build to sustain heavy workloads? he could mess it
>>
>>109491313
Lmao
>>
>>109491306
seeing any kind of cocks makes you a faggot desu, that's why real heterosexual men only watch yuri
>>>/wsg/6208025
>>
>>109491202
oh wow
>>
>>109491323
that paste might be 5 years old now
>>
>>109491313
Hitler of the Norfland
>>
>>109491302
I don't care about reducing ambiguity I just want it to "get" it like if I say "make the blonde girl run around the room while singing" and I provide a picture of a blonde girl, a picture of a room, and an audio file that contains singing then it should just figure it out
>>
File: 1771160113704554.png (619 KB, 1134x2662)
619 KB PNG
>>109491313
the character face consistency is on another level, I really hope their image edit model will be as good, if they do, they will have solved local, I wouldn't have the need to search for more
>>
>>109491313
steamed hams fresh out of the oven
>>
best model for promptmaxxing?
>>
>>109491349
i just jumped back into this tab on desktop in the middle of a gen purely to call you a dumb fucking retard mouthbreathing knuckledragging mongoloid
most of us are literally prompting and giving h3 its references exactly like that, and it's working as intended.
its literally such a robust model you can go out of your way not to follow the instructions and it'll still do better than both ltx 2.3 and wan 2.2 what the fuck more could you want? they can't make it more retarded than that.
>>
>>109491349
deal, your text encoder is now 700B
>>
>>109491359
gemma 31b
>>
>>109491359
why can't we use the text encoder (qwen 32b) to rewrite our prompts?
>>
I remember some people saying the audio issues with the turbo lora were fixed, what was the fix again?
>>
>>109491375
>what was the fix again?
update comfyui (nightly)
>>
>>109491373
you can. nobody stops you
>>
File: Clipboard01dwa.jpg (44 KB, 540x90)
44 KB JPG
>>109491283
>>109491291
>>109491295
Also have a 3090. Not sure if this is ok or not. Didn't undervolt but I did set a temp limit.
>>
https://files.catbox.moe/g8ptyn.mp4
>>
>>109491385
I do, I stop anon
>>
File: 415.png (234 KB, 2878x1256)
234 KB PNG
>>109490718
>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental
How do I load this shit?
>>
File: AnimateDiff_00232.mp4 (2.77 MB, 1216x896)
2.77 MB
2.77 MB MP4
>>
>>109491290
Im trying a full flush each time, but yeah maybe comfy just skips it sometimes
>>
>>109491390
ive noticed with longer gens if you give timestamps (ie 0 to 5s:) it works even better, I wasnt specific enough and used no timestamps here.
>>
>>109491385
I don't think it's possible though, it's using a truncaturated version of qwen 32b (50/62 layers), not the full thing
>>
>>109491394
regular vae loader with updated comfy
>>
File: ComfyUI_temp_fomiu_00091_.png (3.02 MB, 1728x1344)
3.02 MB PNG
>>109491395
I couldnt generate the effect I wanted with this, I wanted a long exposure fast motion crowd in the back while real speed in the foreground with the woman, but couldn't after few tries, maybe another anon can otherwise, a lora has to be trained
>>
File: MiniMax_H3_00054.mp4 (2.52 MB, 832x1248)
2.52 MB
2.52 MB MP4
https://files.catbox.moe/eah3tm.mp4

:/
>>
File: 1757562907106927.mp4 (378 KB, 736x576)
378 KB
378 KB MP4
https://files.catbox.moe/ua0gmz.mp4
>>
>>109491410
That's an experimental model, not vae. I want to experiment
>>
File: AnimateDiff_00234.mp4 (2.06 MB, 1216x672)
2.06 MB
2.06 MB MP4
>>
>>109491417
2girls > grass desu
>>
For people using topaz video, what version do you use?
>>
>>109490998
Well, you asked. https://i.4cdn.org/gif/1786133998021099.mp4
>>
>>109491403
see, like this. now it came out much nicer. 5-10s doesnt really need shot descriptions but it's nice if you want something very specific.

https://files.catbox.moe/94zihc.mp4
>>
File: debo-no.png (1.5 MB, 1056x992)
1.5 MB PNG
>>109490742
>>109490817
>>
File: MiniMax_H3_00462.mp4 (3.73 MB, 1440x1088)
3.73 MB
3.73 MB MP4
>>
>>109491219
why are you so ass-burger
>>
File: 1771373403716183.png (424 KB, 405x720)
424 KB PNG
>>109491422
this shit works so well it knows exactly what kind of animation to adapt to the specific drawing style, Minimax is my new god IDC
>>
>>109491439
yummy
>>
>>109491434
i think even shit fetishists would look at this and be disgusted
thank you for taking one for the team pal
>>
>>109491164
You can use LTX's separateAVlatent and concatAVlatent with it, so you can sorta do a hi-res pass with it similar to image gens.
>>
I got H3 installed last night, I saw some anons posting about using LLMs to produce prompts from prompt guide and there's ever "skills" somewhere? Have links for a retard?
>>
>>109491434
I don't think people realize how big of a deal this is, this will replace all the rich furfags that commision their art 200 dollars each... oh wait, BASED ACTUALLY
>>
>>109491439
nice until the balloon tits
>>
>>109491439
perfect including the balloon tits.
>>
So does comfy nightly break everything or not
>>
>>109491456
he is actually planning to train it lol
>>
>>109491434
>i think even shit fetishists would look at this and be disgusted
lol
lmao
>>
>>109491313
How to prompting a reference for each character? Theres no guide of it in the https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
>>109491465
no surprise, he's training all the models in the world without finishing any kek
>>
>>109491470
he still has krea and his pixelspace model training. He did give up on zimage though since its just a way worse base.
>>
>>109491466
For a few seconds, before they resume fapping.
>>
>>109491434
Well... good to know I suppose. Crazy we can prompt this shit but can't do some basic genitalia.
>>
>>109491375
Slightly better quality but its still hallucinate often
>>
Nice to have a video model that can handle a bit of violence.
>>
What chatbot are you using to make h3 prompts? Grok was working fine but i'm hitting the limit instantly. i tried deepseek but the prompts arent very good.
>>
File: AnimateDiff_00236.mp4 (1.86 MB, 672x1216)
1.86 MB
1.86 MB MP4
>>109491434
lool, thank you for cropping out the dick btw, I noticed the balls at the end
>>
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

1.20x speedup on top of spectrum, aggressive setting only skipped 2 frames at no quality loss.
>>
>>109490972
>webm
kino, what was your process like for this? looks trippy. base image/generation prompt?
>>
>>109491490
androjuniarto/gassytank never even drew her horse dick, just the backsack
>>
What is the minimax meta currently? The speed, spectrum, easy cache? Quants? Pruned seems fine, but still a hog. Qwen32b seems to be my biggest issue speaking of hogs.
>>
the LDG mascot has a show!

https://files.catbox.moe/4wkl29.mp4
>>
turbo lora is the only optimization worth it. Even sage attention hurts motion
>>
>>109491504
how did you manage to get her voice?
>>
>>109491488
Gemma 4, run a low quant on another computer.
>>
File: 1744644695985404.png (50 KB, 320x320)
50 KB PNG
I'm having so much fun but am stuck on 12gb vram. is it worth paying like $1-1.5k just to get a 24gb card?

card chads can you chime in?
>>
File: 1759193211760884.png (447 KB, 2155x889)
447 KB PNG
>>109491502
having success with this, spectrum > easycache
>>
>>109491517
absolutely, VRAM is king
>>
>>109491510
I just asked grok for a structured 15 second prompt for a show called miku LDG live

0 to 5s:
A vibrant animated ad opens with Hatsune Miku appearing on stage in her classic turquoise twin-tails and futuristic outfit, waving to a roaring crowd under bright holographic lights. Bold text flashes across the screen: “Miku LDG Live!” while upbeat electronic music plays.5 to 10s:
The ad cuts to dynamic clips of Miku performing high-energy songs, dancing with glowing effects and virtual band members as confetti and light beams fill the arena. A cheerful announcer voice declares: “Experience the ultimate virtual concert – Miku LDG Live! Tickets on sale now!”10 to 15s:
The screen shows Miku striking a final pose with the full stage exploding in colorful lights and fireworks. The ad ends with the concert logo “Miku LDG Live!” and ticket info glowing brightly as the crowd cheers.
>>
>>109491313
>>
>>109491517
even 16 is enough desu, 4080 here, a 5070ti has 16 too but more vram is better ofc, system ram helps if you dont have lots of vram
>>
>>109491504
the music is genuinely good, better than fucking acestep lmao
>>
>>109491494
>no quality loss.
I've noticed every speed up node I added destroyed the quality of 2D video but not 3D. On the flip side, I have to crank up the megapixels or step count just so the 2d video doesn't have any artifacts. Even with my 5090, this can't have both high MP and long video time. My dream of making my own HQ anime is expensive.
>>
File: AnimateDiff_00238.mp4 (2.32 MB, 672x1216)
2.32 MB
2.32 MB MP4
>>109491501
sorry, I’m not well versed into furry degenerate art
>>
>>109491519
I was finding that too, spectrum is really good compared to te speed or easycache. Turbo lora seems shit as of now.
>>
>>109491517
Regular ram is way cheaper, no?
>>
>>109491517
what 24GB card can $1-1.5k get you?
>>
>>109491517
Wait for the super refresh in 6 months.
>>
>>109491550
A used 3090, they've gone up in price.
>>
>>109491504
same prompt but bumped to 0.4mp, big improvement:

https://files.catbox.moe/6id1zj.mp4
>>
>>109491517
honestly my advice, if you're on fast enough 12gb or 16gb of vram, just get at least 64gb of ddr4 system ram to pair it with. Even if you have 24 or 32gb of vram, you're still coping with boosts like the rest of the fellas in this thread. There's never enough vram, even chads on 64gb are laughing at 32 and 24. Imagine spending over $1000 on 24-32gb of vram and then still having to cope after that. FUUUCK THAAAT.
>>
>>109491548
if the model is in RAM the GPU needs to wait for it to be loaded in the VRAM
>>
>>109491313
He fused with Goering.
>>
>>109490745
sucks to be a ramlet. 64gb is new ramlet when it comes to this model. You should be using wan2gp is you want optimization pertaining to your pc build.
>>
>>109491517
Also anons, what will happen to all the enterprise GPUs AI datacenters are hoarding when they're obsolete and need to be replaced?
>>
>>109491562
this, having enough ram is more important with offloading. Even 64GB is getting too low
>>
genning at 4k native. see you guys in 3 years
>>
>>109490921
>sd2
Stable Diffusion 2.0? That's kinda old and unpopular.
>>
does h3 really need to be in multiples of 32?
>>
>>109491498
The base image was created with z-image turbo and some kind of a 80s movie lora i honestly has already forgotten about, from 6 months ago. The prompt is literally what you see on screen, a bundle of snakes carrying a big apple in a garden with an eye in the sky.
>>
>>109491567
>sucks to be a ramlet. 64gb is new ramlet when it comes to this model.
I have 64gb and I feel like this is far from enough, 128 is really the sweet spot for video models but I definitely don't have the money for that lol
>>
>>109491578
unfortunately yeah, it's not that much of a big deal but I wish they would have gone for 16 instead
>>
>>109491576
Seedance. I think your brain needs more vram.
>>
im a mega ramlet. 16gb. h3 werks pretty good. 50-100/it s
>>
>>109491567
>>109491572
>>109491583
this is why i say system is more important than video now. you can get WAY more system ram for the same money 16gb gets you in video, and that goes a further way from preventing paging which is way slower than ram offloading.
thank fuck i didn't pay more than MSRP for my 16gb 5060 though, i think i'd be angry enough to hunt down jensen myself for that 128bit bus alone.
>>
File: MiniMax_H3_00464.mp4 (2.95 MB, 1184x1600)
2.95 MB
2.95 MB MP4
>>
File: 1366823474973.jpg (32 KB, 384x402)
32 KB JPG
>tfw 64GB with my old 32GB sticks just collecting dust because the motherboard doesn't boot with 4 and I don't wanna buy a new one because i'd have to disassemble everything
>>
>>109491567
>>109491583
I'm on 96GB and have no issues
>>
>>109491571
That's in 10 years.
>>
>>109491434
Holy shit thank you anon
Downloading comfy and H3 right now LET'S GO
>>
File: file.png (6 KB, 320x89)
6 KB PNG
>>109491588
if i live, ill let you know
>>
commercial 2!

https://files.catbox.moe/rl7any.mp4
>>
>>109491618
miku on me pen0r
>>
>>109491504
>the LDG mascot
Talk for yourself. How this unoriginal waifu who -barely- has anything to do with AI is the thread's mascot?
Do you happen to be the /lmg/ guy who has been forcing this for years?
>>
>>109491618
audio quality is way worse
>>
are 10 seconds genns now the norm or just
>last frame=starting frame
>>
>>109491571
>>109491609
I don't even know why they bother hoarding consumer gpus when they have access to workstation gpus with terabytes of vram
>>
File: 1762609352288823.png (354 KB, 500x500)
354 KB PNG
>>109491597
>this is why i say system is more important than video now.
Jensen went so jewish with the VRAM that people had no choice but to perfectly optimize the RAM offloading kek
>>
>>109491634
they hate gamers because they're all chuds
>>
File: 1775421488832275.png (1.06 MB, 1023x732)
1.06 MB PNG
>>109491618
>me knowing I have an i2v/t2v and reference model to make literally anything
>>
>>109491628
On top of that adopting something as a carbon copy for a gen AI mascot. lol
>>
is it possible to use audio input with H3? I had fun generating text -> audio clips and using those to generate image -> video with LTX, but I haven't seen any of that with H3 yet
>>
We're getting to the point where you can do anything, so I have no idea what I actually want to do.
>>
>>109491629
yea im gonna bypass the blockcache node and retest

also is sigma shift still 12/3 default for no turbo?
>>
>>109491634
>I don't even know why they bother hoarding consumer gpus
who is they?
data centers are not hoarding consumer GPUs
>>
>mfw I built a PC with a 4090 and 128GB at the beginning of the year because I knew prices would keep going up.
>>
>>109491654
yeah, audio reference track
>>
>>109491628
Do you happen to be the /lmg/ guy who has been seething about this for years?
>>
>>109491606
Same.
>>
>>109491652
clearly Seinfeld is the mascot as irritating as that might be
>>
>>109491671
he's my second favorite jew.
>>
>>109491666
yeah it's probably the same schizo who spams that jart miku image lol
>>
3090 and 4090 owners are the most uppity, and i don't know why. you both got scammed
>>
>>109491662
>i didnt go for 128gb because im autistic and want 2 stick dimms
>wanted as low cl as i could get
>wanted those cl26 lexar modules
>ali express was the only vendor at the time and due to tariffs wouldnt ship to america
>had to settle for fury kits at cl28 (2x32)
>have a spare of cl30, (2x32)
>could just be on 128gb right now
>but my slop box is a threadripper on ddr4 so it doesnt even matter but i have 64gb of ecc ddr4 anyway so it doesnt matter
ive made strange choices in life but they seemingly pay off somehow
>>
>>109491604
>I have a 5090 I got 6 months ago with a new pc + 96GB ram on zen 5
>I'm too lazy so I'm still use the 3090 one on zen 3
I feel ashamed
>>
>>109491631
H3 can go all the way to like 30 seconds and remain coherent so long as you can fit it into memory. It takes a LONG time though past 20 seconds. Like over an hour for a RTX 6000 PRO at 1MP.
>>
>>109491681
Heh... >>109491678
>>
here's a very thorough comparison between defaut and turbo loras
https://jo-nike.github.io/h3-turbo-eval/
>>
>>109491658
>no more excuses for subpar gens
it is funny how easier it is to use tech how the result of using it is more harshly judged. If something could be done in Wan or LTX then it isn't going to impress anyone
>>
>>109491094
I thought this options meant that models will be loaded from disk instead keeping then in ram, not that it writes temporary data. Are you sure?
>>
File: x.png (23 KB, 749x113)
23 KB PNG
>15s
>1.0mp
feelsgoodman
wish I got another 64gb of ram when it was cheap though
>>
>>109491683
Scammed how?
>>
File: 1768643253495787.png (1.87 MB, 1040x1520)
1.87 MB PNG
Hit a wall with Turbo lora. No matter what i prompt Pic related hair always turned into twintails
>>
>>109491634
They're not anon, it's just that it's the vram/ram itself that doesn't have enough production for now, regardless if it's consumer or enterprise.
It should get better during 2027 as there is a massive scrambling to build new production lines, but until then it's gonna be shitty.
>>
>>109491652
I get that using Miku to test models is valid (to evaluate pop culture knowledge, accuracy etc), but calling it "the waifu of AI threads" is a stretch since it's a meme that got forced by one chronically online guy (who some people say it's a janny)

>>109491666
I doubt I am the only one lol

Like, people can easily make an OC waifu nowadays and reuse it with AI, like the guys at /v/ and maybe other boards made OC for themselves. But some guy out there is a massive Hatsune Miku fan and wants to shove his personal waifu as the face of every AI thread
>>
>>109491697
>it is funny how easier it is to use tech how the result of using it is more harshly judged.
yeah, that's how it work, democratization always means higher standards, and that's always a good thing
>>
>>109491705
Sounds like you need a negative node.
>>
>>109491700
nta. Based on the option description, it only reads from disk. Use a scratch disk if you're worried, but it should be fine.
>>
I want a prompt helper for my Minimax gens. Right now I'm using ChatGPT but want to use an uncensored local model for prompts that are too spicy for OpenAI.
I tried Gemma 4 Heretic 31B, but I only have an RTX 5070 Ti. Even the Q4/nvfp4 version is too big and too slow for my shit, especially since it's a dense model. 4t/s, s mh..
Do you guys have any other suggestions? For lewd prompting, what would you recommend between Gemma4 12B and Gemma4 26B (which of course is a mixture of experts model) quantized to 4bit? Are MoE models any good for lewd proompts? How will it fare compared to the 12B dense?
>>
>>109491094
>it means it's using your SSD instead of your ram
it is writes that kill SSDs, not reads
>>
god damn, minimax understands things out of the box that I tried for months to get ltx 2.3 to render, but failed even with physics loras.
>>
>>109491705
Just add more references
>>
>>109491720
Gemma 4 26B A4B heretic'd
>>
>>109491736
Can you run Minimax reference node with I2V model ?
>>
>>109491720
12B is smarter
>>
>>109491731
Really feels like an early Christmas gift, isn't it? I am shocked how it rarely misses the point in prompts and rarely does something noticeably mangled
I saw some normies on X saying that it follows prompts better than Seedance, even
>>
File: orodSh_00071_.jpg (1.14 MB, 1792x2304)
1.14 MB JPG
>>109491720
Gemma4-12B-QAT
>>
>>109491743
You can ask the reference model to start on <Picture 1> to basically use it as I2V. Doesn't always work, but does sometimes
>>
>>109491768
thanks anon. Gonna try it.
>>
>>109491708
>OC waifu
I don't mind Miku either way. It could have been any character. What happens with "OC" is that they start using it as an avatar for themselves.
>>
>>109491720
I find 26B to be smarter, way faster too. Use an ablit.
>>
>>109491755
>rarely does something noticeably mangled
And even when it does, it feels like it happened because of the lowres (lower resolution than it was trained on), not because the model inherently fucked up
>>
>>109491690
>30 seconds
oh damn
>Like over an hour for a RTX 6000 PRO at 1MP.
holy shit
>>
>>109491780
That one Miku poster from /lmg/ (who probably posts here as well) does seem to be avatarfagging, though
In many times I went there he was replying people with Miku or Teto pics
>>
>>109491540
hot
>>
>>109491690
I am doing 20 secs which is the max without needing a sliding window which kicks off another gen before they are spliced together automatically. So really LTX level, I can't see the point in short gens unless the yare needed for that reason.
>>
>>109491794
avatarfagging is a rule but I have never not once seen it enforced.
>>
>>109491755
It really does. I'm amazed at what it can do on consumer hardware. It almost feels too good to be true. Like I think it might actually be a Chinese weapon to disrupt productivity.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.