[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: awij.gif (2.1 MB, 319x320)
2.1 MB GIF
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109475019

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
based
>>
zased
>>
>>109476286
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE

>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
https://github.com/yizzz927/CAPE-T2V

>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
https://github.com/jd-opensource/JoyAI-Video-Edit

>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
https://github.com/YangYangGirl/ParVL

>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
https://huggingface.co/JamesZar/OliveGemma-3B

08/04/2026

>stable-diffusion.cpp adds support for MiniMax-H3
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md

>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES time
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
https://github.com/IntMeGroup/MIEScore

>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
https://rathgrith.github.io/PeCA

>Kandinsky WM 1.0: A family of models for Physical AI
https://github.com/kandinskylab/kandinsky-wm

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
>>109476286
will smith is so stronk the jannies can't do anything against him!
>>
>mfw Research news

08/05/2026

>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval
https://arxiv.org/abs/2608.03120

>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Plane
https://arxiv.org/abs/2608.03422

>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
https://arxiv.org/abs/2608.03082

>Can T2I Models Draw from the Right Frame of Reference?
https://arxiv.org/abs/2608.03357

>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0

>Self-Supervised Representation-Guided Generative Dataset Distillation
https://arxiv.org/abs/2608.03218

>Latent Reward Registers for Diffusion Preference Alignment
https://arxiv.org/abs/2608.03929

>UniWorld-Design: From Pixel Generation to Layer-Native Design
https://arxiv.org/abs/2608.03971

>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
https://arxiv.org/abs/2608.03708

>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
https://arxiv.org/abs/2608.03059

>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
https://arxiv.org/abs/2608.03269

>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
https://arxiv.org/abs/2608.03135

>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
https://arxiv.org/abs/2608.03057

>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMs
https://arxiv.org/abs/2608.03112

>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
https://arxiv.org/abs/2608.03875

>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
https://qwen-3d.github.io

>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
https://arxiv.org/abs/2608.03649
>>
File: x21.jpg (15 KB, 447x447)
15 KB JPG
me rn
>>
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

they say its not fully trained yet, but something to test, the full one will obviously be better
>>
the nice thing about ltx is that all the little micro camera movements were implicit. with h3, i have to put timestamps all over the place telling it to look here and there
>>
>>109476297
interesting read.
this "animanon" seems like a bit of a faggot.
>>
and suddenly its as if all the "newfrens" have vanished
quite the coincidence
>>
File: AniStudio.jpg (2.05 MB, 1389x2316)
2.05 MB JPG
>>109476331
totally
>>
>>109476346
I'm still here
>>
File: migu.mp4 (944 KB, 864x480)
944 KB
944 KB MP4
>>
>>109476346
vague king
>>
>>109476353
that is impressive
>>
>>109476346
im totally new here and im wondering why anon is spamming off topic links??? >>109476297 >>109476297 >>109476297
how are those links related to local diffusion?????? seems to me like they just defame a literal saint and programming god who will destroy comfyui am i right??
>>
>>109476357
>replying to yourself
>>
I was too lazy to test H3, maybe on weekend..
>>
>>109476368
You are missing out
>>
>>109476368
it's a nothing burger, don't bother
>>
File: ok that's goo.png (270 KB, 500x500)
270 KB PNG
>>109476357
>a literal saint and programming god
>>
>>109476297
>>109476331
>>109476347
Samefag.
>>
>>109476357
How much is tr(ani) paying you
>>
File: cosplayed_robot_abuse.mp4 (602 KB, 864x480)
602 KB
602 KB MP4
>>109476368
kek why? it's an amazing model and comfyui has templates, wan2gp has it preconfigured.

well worth your time if you're ever entertained by such models I'd say
>>
>>109476386
>no u
>>
>>109476286
Can someone give me a good prompt for minimax H3 ref2v workflow. I just want to change the character in the video with the character in the reference image.
>>
>>109476405
You're not attempting to do anything untoward, are you?
>>
>>109476405
character swap is a shot in the dark
>>
>>109476386
proof?
>>
>>109476346
almost as if "someone" isnt trying to hit bump limit as quick as possible anymore
huh
>>
>>109476389
idk, tired in general
>>
File: 1781769189748967.mp4 (488 KB, 544x800)
488 KB
488 KB MP4
>>
>>109476405
You want to use the Load Video -> Get Video Components nodes and align them to the R2V workflow. Then in the prompt, you want to specify that you want to replace <Subject 1> in <Photo 1> with whoever from <Video 1>.

Need to test this in the morning to see if it actually works, but let me know anon.
>>
File: migu.mp4 (1.13 MB, 864x480)
1.13 MB
1.13 MB MP4
>>109476356
agreed. that was a simple prompt (draw graffiti, followed by pose)

looks like we could probably even get her to change cans if i fully described each color she should pick up/set down? not that I will do that rn
>>
>>109476444
good trips, very good post, shit gen but it gets the point across just right
>>
Is the 4-step lora only for the fl model? https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>>
I understand the desire to go fast but is it truly worth quality loss to use a turbo lora? Sure I understand when its wait a minute or two vs half an hour, but even then surely its not needed?
>>
>>109476467
ignorance is bliss
>>
>>109476467
you use it if you are tweaking a prompt. the motions are mostly correct with the lora so you can crank it up when it seems right
>>
kino, reference works well once you figure it out

<Subject 1> is the character in <Picture 1>.

A high-detail cinematic tracking shot of <Picture 1>. He is walking through a dimly lit office with a "Halo Studios" sign on the wall. A man with purple hair and a tshirt saying "they/them" walks up to <Picture 1> and says in a high-pitched, frantic, youthful voice "you didn't respect my pronouns!". The camera focuses on <Picture 1> as he speaks in a deep, gravelly voice: "I don't give a fuck". His tone, pace, and vocal style perfectly match <Audio 1>, with precise lip-syncing. <Picture 1> fires his halo battle rifle at the man with purple hair standing 5 feet away, and they fall to the ground. <Picture 1> walks out the doors of the office.

https://files.catbox.moe/lveb7c.mp4
>>
>>109476485
that's pretty simple, two of the four basic headers
>>
>>109476485
audio tip: 5-10 seconds but not over 15 for cloning, then use that
>>109476489
true but you also need to specify voice traits for the other characters or they will both get the voice, for more complex stuff google ai mode already knows the minimax manual stuff.
>>
>>109476485
Any luck with using it as an image edit model?
Currently trying to do a 0.5 second video where given a base image, the character is replaced with a reference image but no luck so far. It partially works but not as good as nanobanana would normally.
>>
>>109476485
oh you know what you could do? >>>/v/744650971
this could be some real kino
>>
Can you specify height of characters
>>
File: 1772656483394142.mp4 (544 KB, 544x800)
544 KB
544 KB MP4
>>109476444
yeah stop motion looks much better
>>
File: 1758414551456084.jpg (226 KB, 1617x787)
226 KB JPG
I know you guys are all into your minmax vids but I'm having issues with just generating a simple image. I haven't messed around with any of this stuff since A1111 in 2023. Is this a sampler issue, a steps issue, cfg? Is there some sort of pre-made workflow for realistic z-image gens?
>>
>>109476496
10/10 I'll take a dozen
>>
File: x.mp4 (798 KB, 864x480)
798 KB
798 KB MP4
>>109476435
stretching out the internet shitposting is not da wae to get motivation, nap/sleep and do it after you wake up again
>>
>>109476495
<subject 2> a massive man twice the height of <subject 1>
>>
anyone else been running some of their shit through starlight precise?
>>
>>109476496
New Robot Chicken ep just dropped
>>
KINO

https://files.catbox.moe/u060cd.mp4
>>
>>109476500
Use some official workflow. That weight dtype you set for the model may be the cause as well.
>>
File: x.mp4 (1.18 MB, 864x480)
1.18 MB
1.18 MB MP4
>>109476508
probably not. because for what purpose even?
>>
>do a final fantasy i2v from a game screenshot
>it does the broken screen space reflections properly as the main subject moves
holy fuck
>>
this is actually a realistic scenario btw:

https://files.catbox.moe/ya0wtt.mp4
>>
File: 1764410662372557.mp4 (1.16 MB, 640x640)
1.16 MB
1.16 MB MP4
>>
>>109476517
I did that but it puts everything into a single node and doesn't let me use loras. Also the weight-dtype was just me fucking around, I used default for everything previously.
>>
>>109476529
fuck
>>
6 steps euler/beta with turbo lora. Audio seems fine.
https://files.catbox.moe/e157fm.mp4
>>
>>109476530
>it puts everything into a single node
enter the subgraph (it's a button on one of the top corners of the subgraph node, forgot if left or right), or unpack it (right click menu)
>>
File: 1775584922120735.mp4 (1.24 MB, 544x800)
1.24 MB
1.24 MB MP4
>>
File: 2far.webm (3.73 MB, 1728x960)
3.73 MB
3.73 MB WEBM
>>109476453
>>
Anons, H3 has been out for literal days. What's the strategy to get it to generate actual sexy shit and not weird SD 1.0-style body-horror?
>>
>>109476346
No Im just lurking after spending a day trying to get my gens sharp. Trawling through other peoples workflows atm
>>
>minimax_h3_video_vae_int8_convrot.safetensors
I get black screen on final video with this vae
>>
>>109476569
did you update comfyui retard-chama
>>
>>109476297
>https://rentry.org/animanon
Is it true that this guy is one of H3 devs? I've heard that he flew to Japan to strike a deal or something. The name makes sense
>>
>>109476569
update comfy / CU130+ / comfykitchen
>>
>>109476529
It can do stop motion? Unedited?
>>
>>109476519
Better quality and less compute but its probably not even necessary if you're a cartoonfag.
>>
>>109476574
No, Ani, your name is because you used to make dogshit gifs in the SD 1 days. You have nothing to do with h3, absolutely zero model baking skills, and you flew to Japan to beg for money only to get laughed in the face because of your /d/ post history.

https://desuarchive.org/d/search/filename/anistudio%2A/end/2026-08-06/
>>
>>109476585
Meds? I am not "Ani". Seems like you have some sort of personal beef with this guy.
>>
>>109476591
Keep trying. You haven't picked up a single user in 2 years but with a few more years of samefagging I'm sure at least one person will use AniStudio. You aren't being stealthy, by the way.
>>
File: 669.gif (3.86 MB, 400x532)
3.86 MB GIF
>>
reference to video model is SO good man. all I needed was a 10s reference clip and a static image.

https://files.catbox.moe/qh1arn.mp4
>>
File: 1784327066074499.mp4 (721 KB, 576x736)
721 KB
721 KB MP4
>>109476581
>Unedited?
yeah its just the prompt
https://pastebin.com/GRVkKB6y
>>
>>109476453
why has so much artifact this is shit are you using the turbo lora experimental?
>>
>>109476598
>I eard U leik the big p-ajeet meme?
>>
>>109476553
>>109476603
You're killing me with this shit, I love it.
>>
>>109476554
anon... I...
>>
File: 1766614739779912.mp4 (1.07 MB, 544x800)
1.07 MB
1.07 MB MP4
>>
>prompt a 130cm tall girl
>still comes out adult-height
>>
>>109476553
kek how did you get the barbie to stay in static doll form, thats good
>>
based
>>
>>109476643
relative to something else, like as tall as the doorframe
>>
you can add this to get better preview
>>
I've been out since the release, so anons, what optimizations are considered lossless for h3?
I know sage attention has kind of been, but what about sol attention?
is the 22B version of the model kijai really the same quality as the original?
any other optimization?
>>
>>109476656
Just buy an RTX Pro 6000
>>
kino!

https://files.catbox.moe/58zirn.mp4
>>
>>109476645
https://pastebin.com/dgDVmrJq
>>
>>109476655
thanks anon, how much does it reflect the actual end output?
>>
File: windblows_mem.jpg (1.16 MB, 1920x1080)
1.16 MB JPG
Finally minimized pagefile and ssd thrashing, we cooking now boys.
>>
BREAKING
using "..." in your dialogue makes it sound all moany and asmr-like
>>
>>109476690
elipses, dash works the same
>>
OY SHUT IT DOWN!

https://files.catbox.moe/5vzfa1.mp4
>>
File: MiniMax_H3c_00023.mp4 (1.65 MB, 1250x654)
1.65 MB
1.65 MB MP4
>>109476685
>>
>>109475799
same
/wsg/ or fuck off
>>
>>109476705
That's damn nice.
>>
>>109476655
at the cost of extra 2gb of vram
>>
>>109476705
OK that's pretty good, is it:
https://github.com/simsim9-stack/ComfyUI-MiniMaxH3-PreviewOverride
?
>>
bruh, i can't get the enemy tank to shoot. it's always my own tank that shoots no matter what
>>
>>109476099
is this ai
>>
>>109476651
Thanks, I prompted "short as a child" and it worked
>>
>>109476717
try "tank in the distance shoots" idk
>>
File: he's so funny kek.png (298 KB, 399x443)
298 KB PNG
>>109476574
>Is it true that this guy is one of H3 devs?
yes "anon", ani is chinese and works in china
>>
>>109476726
>le black loud man
https://files.catbox.moe/yb9wns.mp4
>>
>>109476716
yea, but KJ also made his version u can try that also
>>
>>109476720
i2v
>>
>>109476732
>entertainers must be as intelligent as Einstein
I don't really agree with that, intelligent people are meant to do smart shit like maths and stuff, and you want them to be clowns?
>>
>>109476735
bs, I don't believe you
>>
>>109476574
>I've heard that he flew to Japan to strike a deal or something.
and he completly failed, sounds like that showing a 35 stars github project as your best achievement is a bit light as an argument to get a million dollar grant, color me shocked lol
>>
>>109476742
that guy in particular is a subhuman, all he can do is act like a monkey complain about a guy like Messi who isn't a worthless dindu.
>>
File: 1778373567822483.mp4 (1.23 MB, 800x544)
1.23 MB
1.23 MB MP4
>>
>>109476757
that said, there are many harmless "content creators" who dont act like apes in public. I would euthanize every shitty streamer that is a nuisance in public.
>>
>>109476757
>that guy is a subhuman unlike my heckin wholesome nigball kicker
>>
File: snitches get stitches.png (365 KB, 399x501)
365 KB PNG
>>109476757
>complain about a guy like Messi who isn't a worthless dindu
bro have you seen how Messi acted in the finals? he's also a fucking subhuman
>>
>>109476765
shouldnt have been racist, chud
>>
>>109476734
oh yeah I'd rather minimize the number of extra repos
>>
>>109476767
I know you're joking, but there's people who really believe a white guy has been racist towards another white guy as if that's possible kek
>>
>>109476774
no, its very easy to not be racist, you should try it some time.
>>
File: et2.png (1.17 MB, 864x1184)
1.17 MB PNG
>>109476774
>white guy
>>
updated turbo lora is out. audio is much better.
https://litter.catbox.moe/9iu2isopewmh3ptg.mp4
>>
>>109476792
https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
>>
>>109476787
>Messi is brown
tell me more frog
>>
>>109476792
Thanks, anon.
>>
>>109476796
>a custom node just to load a lora
I really don't understand why comfy doesn't want to support the diffusers loras in the first place
>>
>>109476799
>we are white senoooooor
>>
>>109476813
wait for someone to convert it
>>
File: 1778559072211804.png (705 KB, 960x960)
705 KB PNG
>>109476814
Messi is genuinely white though, he lives in Argentina but he's ethnically italien, your race doesn't change because you go live to a non-white country, or else Elon Musk would be a black person juste because he was born in Africa lol
>>
>>109476825
Musk is half chink half jew
>>
>>109476820
i think someone posted a conversion script, i'll try it
>>
>>109476829
Xi Long Muskstein
>>
>>109476832
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main
>>
File: jej.png (269 KB, 432x621)
269 KB PNG
>>109476825
>italien
>white guy
>>
>>109476829
He's half chink half nigger
His ancestors were Song dynasty Chinese maritime merchants that settled in African coasts
This was from a 23andme leak
>>
File: I cri everitim.png (1.49 MB, 887x1774)
1.49 MB PNG
>>109476837
b-but we european like you :'(
>>
This is 4 steps using the dual clock euler node and turbo
https://files.catbox.moe/ep9ueg.mp4
>>
*yawn*
>>
>>109476846
>>109476834
https://github.com/shuaixn/ComfyUI-MiniMaxH3DualClockSampler
>>
>>109476844
kek
>>
>>109476792
Straight up doesn't work for me, it doesn't matter what amount of steps I use it just takes even more time that the full 20 steps with no lora.
Do I have to remove all the cache nodes and stuff for it to work properly?
>>
>>109476844
3 doesn't happen. All the others are accurate though
>>
>>109476853
duh
>>
>>109476834
thank you
>>
>>109476844
Mexicans don't look like this
>>
File: uhjgugiuh.png (289 KB, 1520x1044)
289 KB PNG
this is the best WF
>>
>>109476863
i knew a mexican that looked exactly like that
>>
>>109476867
Check under your foreskin
>>
>>109476748
Wasn't Ani one of the first devs ever who told everyone about the potential of video generation? I have never seen anyone promoting AI videos before him. Even if Ani didn't work on H3, I think he's still partly responsible for local video revolution. He was the first to believe in it.
>>
>>109476867
Get rid of easycache
>>
File: file.png (188 KB, 1234x485)
188 KB PNG
>>109476867
i thought these dont stack, u only need 1?
>>
File: uhhhhhhhhhhhhhh.jpg (78 KB, 642x641)
78 KB JPG
>download a couple loras off of civit just to test the lora loader
>none of them work
>ask GLM why
>turns out they're all for some pruned version of the model and incompatible with bf16
holy poorfagerinos
>>
>I looooooooooveee dogshit audio and fried quality
Just wait till lightx or someone actually competent makes an real lora.
>>
>>109476883
yea they do stack and kijai said to use them together
>>
>>109476885
You should train your own loras with your richfag hardware then
>>
>>109476834
this is worse than the original turbo lora
>>
File: ITS TOO SLOW.png (52 KB, 220x220)
52 KB PNG
>>109476886
>Just wait
NOOO I CANT WAIT 10 MN FOR ONE VIDEO ANYMORE THATS A TORTURE
>>
File: 1758186479777615.mp4 (651 KB, 800x544)
651 KB
651 KB MP4
>>
>>109476895
lol its 10x better
>>109476792
>>109476846
>>
>>109476892
kek
>>
>>109476901
and this with only 4 steps and the lora trained to 500 steps so far
>>
>>109476892
>literally 1 (ONE) GPU
>richfag hardware
I don't know man
>>
>>109476878
>the first devs ever who told everyone about the potential of video generation?
Wow, that's some next level delusion of grandeur. All you did was make janky, dogshit gifs that no one in the industry even knows exist. I can't believe you're typing this with a straight face and not realizing the humiliation you inflict yourself.
>>
>>109476886
lol, lightx has horrible distill loras that kill motion, they can't even get it right for their own LTX model
>>
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

its not fully done, but to try
>>
imagine needing poorfagtimizations
>>
>>109476901
it's very blocky for me using the same settings as the first turbo lora
>>
File: 1772627308876279.png (75 KB, 1399x395)
75 KB PNG
>>109476912
So I download the minimax_h3_turbo_4step_ckpt500_pruned_comfyui.safetensors one right?
>>
this is new turbo lora at 8 steps
https://files.catbox.moe/9bnw18.mp4
>>
File: MiniMax_H3_00026.webm (1.24 MB, 864x480)
1.24 MB
1.24 MB WEBM
5mins at 20steps on a 3060 12gb for a 5sec video with audio
this is pretty nice

https://files.catbox.moe/d4u7hf.webm
>>
>>109476911
Dunno man my guess is that at least is going to be significantly better than the stuff from literal whos we getting right now.
They are on it anyways so we'll see https://github.com/ModelTC/LightX2V/commits/main/
>>
>>109476918
I think so, thats what im trying, havent genned stuff yet though
>>
>>109476931
based
>>
File: image (16).jpg (36 KB, 520x253)
36 KB JPG
>>109476931
I hope you have 8x 5090s
>>
>>109476926
do u switch off the easy cache / h3 cache when u use turbo lora?
>>
>>109476941
they need that to make those turbo loras, jesus I didn't expect the requirements to be this expensive
>>
>>109476943
dont use easycache with turbo lora unless you are doing some high steps still. 20 steps with distill lora + it would prob be like doing 50 steps though if you want the quality
>>
so normally Comfy will offload to cpu/ram, would adding a second GPU to your rig speed it up then? would it offload to the second GPU first, and only then to CPU?
>>
>>109476941
You don't?
>>
>>109476949
they dont say they are making a turbo lora, just that they are doing sparse attention along with better 8x card communication for better speeds when running it on 8x 5090s
>>
>>109476953
no
>>
>>109476959
why do you need 8 5090 cards to run minimax?? do you really need 32*8 = 256gb of vram to load everything? lol
>>
6 steps with new turbo lora works well enough
https://files.catbox.moe/zmxtkk.mp4
>>
>>109476834
>https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main
don't go for 4 steps lol
https://files.catbox.moe/vimzug.mp4
>>
Why is it that LTX 2.3 works better the more you talk in the prompt like a complete jeet?
>>
>>109476969
its parallelism, you can have them all running the model at once in shards making it gen much faster
>>
>>109476983
so why >>109476964
why can't 2 gpus do parallel?
>>
>>109476996
layers, or something
>>
>>109476996
you can with NVlink. Without it they can't communicate with each other faster than they run alone
>>
>>109476979
I smoked weed on an empty stomach once and the same exact thing happened
>>
>>109476979 (You)
8 steps still gives you terrible audio, they haven't finished cooking
https://files.catbox.moe/zskodl.mp4
>>
>>109476981
Mechanic Jeet dataset
>>
File: file.png (86 KB, 1615x438)
86 KB PNG
4060ti, ref2va with one image

with pic related row of optimization
175 seconds https://files.catbox.moe/jy3sxu.mp4

with no optimizations what so ever
485 seconds https://files.catbox.moe/7p6lbt.mp4
>>
>>109477023
are you using the dual clock sampler?
>>
>>109477003
how come this works then >>109476969 5090 doesn't nvlink
>>
>>109477050
cause they are developing a way to make it work. That is the whole point. Can't you read?
>>
>>109476889
proofs?
>>
so with that lora you can gen with 10 steps (or half), so far. not fully trained but it works: didnt specify dialogue so he's mumbling

https://files.catbox.moe/oj7vzb.mp4
>>
>>109476899

Sweet nice loop, what was the prompt?
>>
>>109477063
don't make excuses for it, gibberish while amusing is the bane of this model
>>
>>109477063
The loras makes my gens slower even at 4 steps, I'm gonna assume they are not compatible with the reference model.
>>
File: ugu.png (287 KB, 577x433)
287 KB PNG
>>109477033
tf is that, why can't it work like normal?
>>
>>109477073
im gonna wait till it's properly finished cause with sageattn/spectrum the gens are high quality and not too long at 0.3mp.
>>
do you need something special to use the 500 step turbo lora? it looks bad for me
>>
>>109477078
cause the audio gets negatively effected otherwise
>>
File: MiniMax_H3_00027.webm (2.13 MB, 864x480)
2.13 MB
2.13 MB WEBM
30 steps with res_multistep and simple scheduler

Anyone figured out a better sampler/scheduler combo?

https://files.catbox.moe/3caocl.webm
>>
Almost 1 week since H3 release and still no cock/pussy LoRA?
>>
>>109477100
sulphur plans to train soon. He already raised 4k of 10K goal in one day
>>
>>109477100
there's some early attempts on civit but they all likely suck
>>
>>109476689
how
>>
>>109477095
>simple
Some people say beta works better for them but it produces horrible results for me
>>
>>109477055
sorry, I did miss the point.
so for now a 2 GPU build is only good for 2 separate runs at the same time then? not sure that's worth as much as a speedup of 1 run would have been
>>
h3 50 steps does dramatically improve the motion. It feels soulful and emotionally charged.
>>
>>109477127
running two separate generations is faster. there will be more latency to have to manage the memory of multiple devices for a single one
>>
>>109477134
try the distill lora at 20 steps. Its close to 50 steps
>>
>>109477134
this but 500 steps
>>
stardew valley, middle east edition

https://files.catbox.moe/0n9a1a.mp4
>>
File: MiniMax_H3_00029.webm (2.32 MB, 864x480)
2.32 MB
2.32 MB WEBM
>>109477114
I set my pagefile manually to 24gb and disabled sysmem fallback in nvidia control panel.
Also have this in my .bat for comfy

set COMFYUI_CACHE_RAM=True
set CUDA_MANAGED_FORCE_DEVICE_ALLOC=1
set PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:64

.\python_embeded\python.exe -s ComfyUI\main.py --fast fp16_accumulation --fp8_e4m3fn-text-enc

Not sure if that fp8 text encoder flag is helping though.

euler_a a shit
>>
>>109477142
That must be something. The leap from 10 to 50 is already consequent.

Also the more local models evolve, the more I hate myself for not getting a Blackwell at $10k
>>
>>109477124
Are you using the reference model?
>>
this is why they opensourced
>>
>>109477134
does more steps improve the grainy artifacts of fast moving stuff?
>>
>>109477163
it falls off over time but yes, it will keep improving
>>
File: MiniMax_H3_00208.mp4 (1.37 MB, 704x736)
1.37 MB
1.37 MB MP4
>>
>>109477163
Image quality looks bound to the output resolution a lot.
>>
we can fix hollywood now.

https://files.catbox.moe/g9zznk.mp4
>>
File: 1778207483959798.jpg (1.01 MB, 1248x1824)
1.01 MB JPG
>>
>>>/wsg/6209063
Still getting the hang of the reference model. Gemma31bchan has been a big help with prompts
>>
>>109477175
t2v result: pretty good desu

https://files.catbox.moe/ahb7p5.mp4
>>
File deleted.
>>109477030
sadly this setup seems to produce errors
>>
>>109477201
free vacation
>>
Are there any options for GPU renting that could handle H3? I'm stuck on a gpuless laptop for a while and H3 is probably the first video model that seemed legit interesting to me.
>>
mental note
[96768, 2688] - bf16
[96768, 8] - pruned
>>
>>109477217
Note: not some apishit service but something that would let me run a ComfyUI instance. I know people used paperspace way back in the day.
>>
>>109477201
try disabling the nodes one by one. see which one is the weak link.
>>
File: lambda_08.jpg (141 KB, 637x1282)
141 KB JPG
>>109477217
The problem with being a rentcuck is the thing you want to rent is always taken first.
>>
File: Pagefile_Is_Maximum.png (38 KB, 755x857)
38 KB PNG
A 10 second clip might have been asking a bit much...
>>
lmao, I t2v'd a gen with bethesda in it and it got a game aesthetic, I was pretty vague, still good though.

https://files.catbox.moe/mwrrxg.mp4
>>
>>109477275
gonna make breakfast and do a longer slightly more hq gen. this is fun.
>>
File: 1766985014439598.png (72 KB, 791x613)
72 KB PNG
>Minimax's stock is up 66% since the launch of H3
they deserve it
>>
>>109477292
lmao that's awesome to see really
>>
>>109477292
but did it pay for the training costs?
>>
>>109477292
They took a gamble going open weights and it paid off
>>
>>109477159
>this is why they opensourced
i don't know why more companies don't go open source. it's basically free money, and you avoid dealing with payment processors or trying to make the model profitable
>>
>>109477304
>i don't know why more companies don't go open source.
big companies like disney can sue you and say that open sourcing the model helps into facilitating the copyright infringement or some satananistic big corporation mumbo jumbo shit
>>
>>109477304
>and you avoid dealing with payment processors
that doesn't make sense. they are already dealing with payment processors as a business
>>
>>109477228 >>109477217
maybe still vast.ai or something. ask in the non-local threads or look up that and alternatives, I'm not up to date on API shit.

alternatively, give comfyui and minimax some business via the SaaS API offer as reward for giving you the option to freely switch whenever you finally want/can build your own machine? they're at least not holding your ass hostage.
>>
any fix for sage attention doing nothing? I installed the right versions, ran the tests, no errors during gens, but the times stay the same. 3090 rtx
>>
What turbo lora to use for the int8 convrot versions of the i2v and ref2v h3 models? I'm seeing pruned everywhere.
>>
>>109477320
>ask in the non-local threads or look up that and alternatives
"Non-local" threads are mostly using proprietary models or API services. I'm interested in spinning up an instance with fully-fledged ComfyUI/whatever. I know it's not local, but it's not really "apishit" either. I'll look into vast.ai, thanks, I think I've seen it mentioned before.
>give comfyui and minimax some business via the SaaS API offer as reward for giving you the option to freely switch whenever you finally want/can build your own machine? they're at least not holding your ass hostage.
I'll have a look but I'm not really interested in a restricted service. If it's any good I'll try it though
>>
>>109477292
they do deserve it good for them
so far, testing this model, i found that error prompting rate is low if one follows the format and quality is consistent

not sneeddance tier but at least 80% of quality indeed
>>
>>109477195
It's crazy to me that shit just werks. No hacky loras or weird checkpoints. You just give it an image and it uses it seamlessly
>>
File: Screenshot_38.jpg (56 KB, 803x413)
56 KB JPG
>>109477348
Not natively enabled for H3 by comfyUI. Needs to have "--use-sage-attention" batch start flag and KJ nodes.
>>
>>109477376
Pretty sure the patch sage attention node does nothing if you use the --use-sage-attention flag. It's specifically for people who don't use that launch flag but still have sage attention installed
>>
>>109477353
if the file size of your h3 model is 22gb, then you are using the pruned one
>>
LLMs have made setting things up so complicated that you need another LLM just to set it up for you. not all of us are advanced computer scientists
>>
HAHAHA

praise China for this model.

https://files.catbox.moe/ul7310.mp4
>>
File: Screenshot_39.jpg (99 KB, 803x329)
99 KB JPG
>>109477384
For your sake, I hope you're right. Cause it works on my machine.
>>
Holy shit dpmpp_2m_sde sounds so clean and looks so good by Kijai's example.
Glad samplers are now working properly so now we can gen even better looking shit.
>>
>>109477403
That's the other node, you always want that one enabled. I'm talking about the "Patch Sage Attention KJ" node which does nothing with the --use-sage-attention flag
>>
anyone tried training with ostris experimental adapter yet? i'm gonna give it a go. but he said he'll train another better version
>>
File: absolute cinema.png (122 KB, 498x311)
122 KB PNG
>>109477402
>oh no my pronouns were ignored
>>
>>109477357
Maybe this? https://www.thundercompute.com/cloud
>>
>>109477369
I notice that if there's an object at say 9s that is entirely new to the scene at that time but is say hidden, like glasses for example, even though they are not being acted upon (unfolded and placed on the head) they appear in the scene before the 9s mark in the persons hands.
I'd not noticed this before.
Nifty.
>>
Hope you fuckers are not genning CP with the model because it's very capable
>>
>>109477402
version 2:

https://files.catbox.moe/gk35st.mp4
>>
>>109477393
if i could do it, you should be able to as well
>>
>>109477357
>I'll have a look but I'm not really interested in a restricted service. If it's any good I'll try it though
Sure, I leave it to you to try. If you ultimately toss them a few bucks, IDK if it will incentivize future models but they certainly deserve it for this.
>>
>>109477391
The convrot int 8 is 33gb, so no. There really isn't a turbo lora for that yet, why?
>>
tried timestamps. almost fully worked, I could refine the prompt a bit.

https://files.catbox.moe/65phjn.mp4
>>
The new lora is fucked at whatever settings people were saying to use before, it's seems ok at a lower strength? 1.2/1.3 is way too much
>>
File: 7203644421.mp4 (3.61 MB, 864x864)
3.61 MB
3.61 MB MP4
>>
>>109477498
i think you're looking at the text encoder, not the diffusion model. the int8 for the diffusion model is about 22gb
>>
>>109477503
Most of the time I'm retarded, but not this time I think.
>>
>>109476485
you didn't even do it right, retard. You're calling the character <Picture 1> instead of <Subject 1>
>>
>>109477501
Do you want to know how I know you're a jeet?
>>
>>109477513
how is he retard if it worked?
>>
>>109477509
oops, sorry. download the pruned versions, there is no reason to use the original ones unless you are training things
>>
>>109477421
The prices look nice at first glance, thanks.
>>109477466
It's kinda fun seeing Comfy grow so big that his corp is running their own SaaS and funding the big models. I mostly remember him from the SD 1.5 days when he felt like an asshole shilling his weird UI but the dude really was serious about this stuff.
And yeah, if the comfy service is not anal about horny prompts I'll probably go for it.
>>
>>109476147
>Maybe specify "no dialogue" in your prompt
Tried, doesn't work.
>>
holy shit, t2v not even reference audio. I want a data list of known characters just out of curiosity.

https://files.catbox.moe/mbvnro.mp4
>>
>>109477541
>t2v
yawn
>>
Anon you did get claude to make you a minimax prompting tool that can connect to your LLM that's running on your network so you can greatly increase the quality of your H3 gens, RIGHT?
>>
>>109477543
r2v is my favorite model by far cause you can plug in any character or any audio, im just testing both models
>>
>>109477547
my pc is the most powerful on this network, I do have a spare 12gb card but any LLM that can fit on it won't be worth a damn
>>
woah 36 stars now
holy shit
>>
>>109477550
Might be able to use a quant of gemma 4.
>>
With video continuation, is the meta to cut and re-encode the video with a lower frame rate, right?

I tried continuing a 9-second gen last night, using the video + audio track as reference and it took 1400+ seconds. The result was good but damn that's a long time.
Then I cut the video length in half (trimmed the start) and encoded at 20 fps (could've probably gone less), tried running the same prompt with the re-encode as input, and it took ~700 seconds.
>>
>>109477543
the video is in the style of a cinematic Star Wars movie with Yoda as the star.

0 to 5s: Yoda is standing beside a tree in front of a pond, in the forest. Yoda says "blacks problem yes, remove blacks you must do!"

5 to 10s: close up of yoda who says "otherwise, bike steal, they do." the camera zooms out and a regular bicycle near him is force pushed away by Yoda.

https://files.catbox.moe/la18am.mp4
>>
>>109477534
exactly why Ani is destined to win. comfy still won't fix shit that people were complaining about for years.
>>
Nvidia_RTX_Nodes

gonna try this with 0.3 or 0.4 gens
>>
How do I prompt petite body but with huge breasts? Every time I do this the model gives me a landwhale
>>
>>109477569
Who win what?
>>
>>109477601
UI wars
>>
Can one of you cucks explain to me how the cloud providers will be able to sustain their video generation?
Sora already had to shut down because it wasn't profitable. Now that Minimax H3 is out, and it's outstanding, why would anyone pay for shit like VEO 3 and Kling? And wan, what are alibaba going to do there? Absolutely nobody will pay for wan at this point.
>>
>>109477609
sora 2 gave away 30 generations free a day to millions for months
>>
>>109477609
>>109477614
also sora 2 was making a bit on having companies pay to license with them. But they got denied. And what's funny is now Disney has teamed up with bytedance to use seedance
>>
>>109477614
Sora 2 wasn't good
It was worse than H3/Seedance 2.0
>>
>>109477609
Whoever's able to monetize NSFW saas vidgen despite is gonna make billions.
>>
>>109477620
ehh... it still had magic that even seedance 2.5 does not. it knew SO MUCH. It was clearly a BIG model
>>
>>109477547
Close, I made claude make me a workflow for an interactive avatar connected to an LLM. I send it questions or requests and it sends me video answers, it speaks and grabs props. It uses a picture for reference (and it can look at its own clothes) and an audio reference for voice.
>>
>>109477620
H3 is amazing but come on, Sora was the best video gen model BY FAR quality-wise.
>>
>>109477624
you technically can't because of the agreement you need to have with the company but this is reality so there is probably some Russian telegram you can gen porn on
>>
File: 1783455013465978.png (106 KB, 861x639)
106 KB PNG
if on nvidia, try this as you can gen at lower res (faster) and get better outputs, essentially dlss for comfy

the nvidia file was only 400mb
>>
>>109477609
>Sora already had to shut down because it wasn't profitable
it wasn't profitable when compared to using the same hardware for making and selling tokens for programmers
>>
>>109477639
*change scale to 1 for 1:1 video
>>
File: MiniMax_H3_00027.mp4 (1.17 MB, 864x480)
1.17 MB
1.17 MB MP4
Goddamn, it's really good.

https://files.catbox.moe/bxj2nv.mp4
>>
>>109477624
you cant monetize nsfw vidgen because of payment processors. crypto was supposed to solve this but it just became a meme instead
>>
>>109477556
Continuation doesn't work properly. Even when you very clearly state in the prompt that the target video is a direct continuation of <Video 1> and set all the labels and parameters right, it still won't be a seamless continuation and will feel like a jumpcut. Think you just have to manually save the final frame and use that as the first frame, kind of like with wan chained continuation.
>>
>>109477645
Proven false because H3 is a very strong model compared to frontier video models and can run on shitboxes, while the best model /lmg/ can run still trails frontier LLMs
>>
File: 428163429.mp4 (1.66 MB, 736x1024)
1.66 MB
1.66 MB MP4
>>109477514
You've made me curious.
>>
File: 655356.png (310 KB, 935x1327)
310 KB PNG
>>109477547
no i've been writing this thing for almost 2 days now
>>
>>109477661
>crypto was supposed to solve this but it just became a meme instead
wouldn't it "kind of" still work here? not that I'm going to be doing this, but in principle it should be possible to get some very small crowd of cryptofags to send you their coins no?
>>
rtx upscale at 1.0 test, same yoda:

https://files.catbox.moe/16kiy7.mp4
>>
>>109477547
I just use SillyTavern and dump the whole two prompt instruction docs into the context and then tell it what I want after that.
>>
>>109477686
>>109477686
>>
>>109477609

What percentage of chinks has a good GPU? Last time I heard from a chink friend, most of them still goes to gaming cafes. Also third world browns.
>>
>>109477678
>very small crowd of cryptofags to send you their coins no
That wouldn't be enough to justify the costs probably. Most people are too scared of crypto.
>>
>>109477678
DLsite does it so it could work here, too.
>>
>>109477627
Stop sucking Emliy Youcis's cock. Even she is on board with H3 being the best
>>
messed up my wf or comfyui
it is taking double the time to gen a video with minimax h3!
tried installing sage 2.2, reverted back to sage 1 but comfy is performing worse now!

need maybe a better workflow for ref! anyone care to halp please?
>>
>>109477534
definitely put in a lot of work since.
>>
>>109477748
for references, video input can make it a lot longer, pictures and audio not so much
>>
>>109477664
yeah it runs but on medium-high local hardware you can gen a 10 second H3 video or about 50k tokens of current best-easily-runnable LLM, the latter feels a lot easier to package up and sell for a given price plus you get free training data off the input people send you, which potentially is novel fully human-authored text but at the very least is their AI-generated but curated codebase which they as a human assessed as being worth continuing to improve and work on. i generally assume that's why the OpenAI/Anthropic personal coding plans are so well priced relative to the API $/token businesses have to pay rather than them just pricing solo individuals out of the market; we're worth it for the authored input variety
>>
>>109477666
Bollywood slow mo.
>>
>>109477159
>>109477304
>i don't know why more companies don't go open source.
they only got a stock jump because they published literal fucking global sota video gen model (given seedance got censored/nerfed), not really because they went open source, although that did help since it guarantess that it wont get censored, that people can fine tune it, that it gives you full privacy etc
>>
>>109477952
I gotcha, makes sense I suppose.
>>
>>109478108
truth is there was a pump yesterday, check the indexes
>>
>>109477614
>sora 2 gave away 30 generations free a day
amazing
>>
Hi anons. Anyone tried running this on 3090 and 32gb ram? is it painfully slow? is 5090 and 64gb ram the minimum for good time with minimax H3? I want to atleast be able to watch youtube ishowspeed while i wait for the prompt to finish. Maybe also play a match of Dota 2 while i queue up some prompts.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.