[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Americans Aren't Allowed To Use The Model Edition

Discussion and Development of Local Image, Video, and Music Models and Software

Previous: >>109443073

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>LTX-2.3
https://huggingface.co/collections/Lightricks/ltx-23

>Wan
https://github.com/Wan-Video/Wan2.2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/02/2026

>MiniMax H3
https://huggingface.co/MiniMaxAI/MiniMax-H3

>MiniMax H3: Repackaged model files for ComfyUI
https://huggingface.co/Comfy-Org/MiniMax-H3

>MiniMax-H3-INT8-CONVROT
https://huggingface.co/Gluttony10/MiniMax-H3-INT8-CONVROT

>MiniMax H3 COMMUNITY LICENSE AGREEMENT
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE

>LoRA Dataset Studio: LoRA workflow in one tab
https://github.com/perfectgf/lora-dataset-studio

>comfyui-vram-tracker
https://github.com/PuppetMasterAI/comfyui-vram-tracker

>Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
https://nvlabs.github.io/Sana/Sol-Attn

08/01/2026

>EU to get power to enforce rules on AI starting today
https://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786

>FameGrid Auto Color for ComfyUI
https://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui

07/31/2026

>One-take Creation, Flexible Referencing: Introducing Seedance 2.5
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5

>ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
https://github.com/avaxiao/ReToken

>RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
https://github.com/liuxiaobo66/RefineSVG

>ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
https://github.com/H-EmbodVis/ROAD

>PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
https://physomni.github.io

>DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
https://github.com/anonyme610/dinolizer

>Inline Studio v1.2.6 - Flux 2 & Minimax H3 API
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.6

>SAM 3.1 Multiplex
https://huggingface.co/Sparknight/sam3.1-int8-int4-convrot
>>
THE REAL THREAD IS HERE
>>109440861
>>109440861
>>109440861
>>
>>109444042
>The chinks are releasing such a powerful model to destroy the American movie industry.
>>
>109444064
eternal sneed trani
>>
bodied
>>
god I have so much shit to try, haven't touched the ref model yet, I give this 2 weeks before I'm bored
>>
I AM USING MINIMAX H3 IN VIOLATION OF THE LICENSE AGREEMENT
>>
>>109444064
>(Dead)
finally the jannies did their job, based
>>
>>109444052
>>LTX-2.3
>https://huggingface.co/collections/Lightricks/ltx-23
>
>>Wan
>https://github.com/Wan-Video/Wan2.2


Do the new model replace these ?
>>
File: keeeeeeeeeeeeeeeeeeeeek.png (1.17 MB, 864x1184)
1.17 MB PNG
>>109444064
>>
>have to gym, shower, eat big dinnet and rest before I can play with h3

Fuck.
>>
>>109444073
I genuinely wondering if I'm conditioned to want to use the i2v model more but would actually get more juice out of the ref model just because of how consistent it seems.
>>
File: 1770200728560654.png (19 KB, 620x192)
19 KB PNG
Another week, another comfyui humiliation on their own preddit for the same issues that people have been complaining about for months and months on end, kek
>>
>>109444064
huh? where? I don't see anything. can you show me where the real thread is?
>>
File: f.mp4 (519 KB, 864x480)
519 KB
519 KB MP4
>>109444064
it just got deleted? ah well reposting

I tried the UI type prompt with a command and conquer ui screenshot: >>109443518

this is so damn good

>>109444046
no doubt, and for many other things - storyboarding with image/video/audio references and an edit model that "understands" how to pick elements/characters/... by prompt is fucking powerful
>>
For those who missed it kek
https://streamable.com/6sdsoc
>>
>>109444078
early days yet, but I have genuinely no reason to touch ltx again
>>
>>
>>109444082
always use reference images. the reason why the ltx fail posters were screwing up is because input images don't have the type of motion blur that a video would have. it's better if you let the model create the starting frame
>>
https://litter.catbox.moe/o2l0wqp1glmcb5s4.mp4
>>
https://litter.catbox.moe/a5k6olstccoii55h.mp4
impressive
>>
>>109444086
dumbass newfren kek
>>
Yeah there's no way they didn't release this model to disrupt various industries.
I can pretty much do all of the stuff I needed api models for currently.
It's to the point where I'm having doomer thoughts in terms of regulation. This, Kimi, Deepseek, etc...
>>
File: f.mp4 (1.38 MB, 864x480)
1.38 MB
1.38 MB MP4
>>109444078
oh yes, pretty much entirely.

this is far more powerful with the references, better audio, much more seconds of coherent video (without SCAIL reference+continuation type tricks only like on wan) etc.

well yes we still will need some loras again etc. but this is hands down and easy to call better
>>
i deleted ltxv-2.3 and all the loras i had for it. hope israel gets bombed into oblivion just like their outdated shit model.
>>
>>109444099
>This, Kimi, Deepseek, etc...
Qwen too lol >>109442634
it's like China wants to show who's the boss now, BASED
>>
RUN, WARWICK, RUN!
https://files.catbox.moe/z5e98o.mp4
>>
>>109444099
I've been worried they will outright ban GPUs since the SDXL days.
>>
File: zeev.png (1.59 MB, 1000x1000)
1.59 MB PNG
>>109444107
you will be crawling back when ltx 3.0 gets released... but i won't let you have it, stupid goy cattle
>>
>>109444113
why is he speaking non american
>>
>>109444099
>I can pretty much do all of the stuff I needed api models for currently.
it's insane to have a model that's like top 5 ever (compared to API), Minimax is my new god, I fucking love communism wtf!!
>>
>>109444099
everything is about power
>>
reminder that flux 3, wan 3 and LTX3 are apparently soonish as well. And now they have to at least match H3 or basically die as companies
>>
>>109444113
midge
>>
>>109444099
>to disrupt various industries.
tell me a single industry level use case this model can deliver on
>did you see the janky 15 second clip anon posted??
just because its the best local video model doesn't mean its useful. you still can't make anything professional
>>
kek
https://files.catbox.moe/vd3t2u.mp4
>>
File: 2309498385.jpg (43 KB, 1000x591)
43 KB JPG
what are you listening to right now, bros?
https://www.youtube.com/watch?v=OjNpRbNdR7E
>>
>>109444088
comfy ads are cringe
>>
>>109444134
later wan iterations were all mid/shit, they probably cant top h3, ltx3 is not until end of year and they wont top h3
>>
https://files.catbox.moe/vqzha4.mp4
How to get jumpcuts working. I wanted an abrupt cut to her eating like a pig
>Jump cut to the Korean woman messily eating fried chicken. Slurping and chomping loudly
>>
Someone on bandoco discord said they are gonna work on a 4 step lora
>>
vramlets, how you holding up?
please call in every hour
>>
https://files.catbox.moe/5c75mm.mp4
it do be like that
>>
>>109444144
https://www.youtube.com/watch?v=MrA1gKoVKSc
>>
>>109444155
16 gb no problem
>>
File: MiniMax_H3_00007_.mp4 (2.63 MB, 576x736)
2.63 MB
2.63 MB MP4
>>
>>109444113
is he speaking midget?
>>
I lost the link but can whoever did that outfit transform post prompt
>>
breast friend of threadship
>>
>>109444155
4gb. I'll try it tomorrow. I used to wait 30min per acestep gen but I still dunno if it'll run at all
>>
>>109444156
LMAO
KINO
>>
File: i_00264_.png (1.47 MB, 768x1376)
1.47 MB PNG
so how long it takes you niggas to gen and whats your hardware?
any way to make this shit generate faster?
>>
Is generating images worth going to prison over?
>>
>>109444156
so it can do text? based if so
>>
>>109444156
lel. This feels well directed. What model was it? i2v or ref?
>>
>>109444179
it's amazing at text, I genuinely cannot understand why they decided to give such a powerful model, this is insane
>>
>>109444175
depends how long I'm willing to wait, but I'm settling for 5 minutes at the moment to test the knowledge base
>>
>>109444158
This one's actually good, but the guy is Taiwanese.
>>
File: f.mp4 (1.61 MB, 864x480)
1.61 MB
1.61 MB MP4
>>109444118
i wish ltx good success, i do like how they gave us the last model.

but h3 is like basically ALL the issues I had in mind with ltx already fixed. the (pretty much always) odd audio noises being extremely obvious but also i had already wished before that prompt adherence and storyboarding worked more like wan... h3 is pretty to very good there as far as I can telel AND it also does nice stuff on a simple prompt that isn't storyboarded in detail as far as I can tell.

it's so. damn. nice.
>>
File: 57467.png (130 KB, 2414x565)
130 KB PNG
>>109444182
arghhh wake up nigga!!!
>>
>>109444175
looks like I spoke a little too soon. when you gen longer videos the gen times go up drastically.
>>
>>109444175
7 seconds at .5mp is like 11 minutes on a 3090.

Needs a turbo LoRA now.
>>
low res + RTX upscale node is much faster and nearly as good as doing high res
>>
File: 1773464664719039.png (71 KB, 745x785)
71 KB PNG
>>
>>109444196
its half that on my 3090. use sage attention
>>
File: 1778600678637553.png (70 KB, 729x713)
70 KB PNG
>>109444200
>chatgpt give a corpo response to this, but in lower case so it seems like i typed it out and care.
>>
>>109444200
Just like, unpack the subgraph lol?
It's the first thing I did when I downloaded the WF.
>>
>>109444205
jfc just copy-pasted from llm lol
>>
>>109444203
Oh shit I forgot about that. It's been so long since I've used a model that sage didn't fuck up.
>>
Thank you for keeping us updated on the Plebbit situation, Anon!
>>
File: 1778356522465765.png (30 KB, 692x335)
30 KB PNG
>>109444205
>>
>>109444205
wtf lol?
>>
>>109444144
https://www.youtube.com/watch?v=u5t65IZ-fZI
>>
does small resolution mean small audio?
>>
https://files.catbox.moe/285140.mp4
>>
https://github.com/kohya-ss/musubi-tuner/pull/1018
>>
>>109444231
>Qing Long Qing Long Sheng Zhe
>>
So can it do porn? That's all I care about and can Americans and Europeans actually use it? Signing a waiver is super suspicious
>>
File: Untitled.png (182 KB, 351x252)
182 KB PNG
>>109444230
>https://files.catbox.moe/285140.mp4
mfw
>>
>>109444231
cool but he should give it at least a day or two to enjoy, if anyone's clamouring for loras already then fuck em
>>
File: 542335432.mp4 (2.98 MB, 896x608)
2.98 MB
2.98 MB MP4
>>109444156
kek, they had their time under the sun o7
>>
>>109444238
for personal use you don't have to. It's only commercial
>>
>>103416132
>>103416182
did we lose our soul along the way?
>>
>>109444156
skipped some of these niggas
>>
File: 1761092097028015.png (58 KB, 218x239)
58 KB PNG
>>109444230
I look like this
>>
https://litter.catbox.moe/nhbvnw7e4dto1osy.mp4
wow, it can do rimjob decently out of box
probably need lora for specific movement.
>>
they said the 2K upscaler model will be opensourced soon, that it was more difficult to make work. That apparently clears up all the blurriness of low res gens
>>
>>109444175
>>109444194
I have a 3090ti and 64GB and going from 5 to 10 seconds of 480p was almost a linear increase, 8 to 17 minutes.
>>
>>109444268
frames are a linear increase, res is quadratic. Hopefully we get 4 step lora soonish
>>
>>109444258
lmao
>>
File: MiniMax_H3_00008_.mp4 (2.56 MB, 864x480)
2.56 MB
2.56 MB MP4
>>
>>109444230
>slime girls out of the box

what is my opinion on this
>>
File: dfgdfgdfe.jpg (26 KB, 798x117)
26 KB JPG
they say they have a way to make it much faster soonish
>>
Any workflows?
>>
>>109444283
we're in the dark ages, only way is up
>>
https://litter.catbox.moe/kyn2qpo3pigmctzh.mp4
>>
>>109444156
Kek shit on sora and grok too
>>
>>109444252
There's no way it doesn't know Noir.
>>
>>109444118
It's over baldie! Just like your model!
https://streamable.com/v1o0rs
>>
>>109444283
>sparse attention inference
if it's faster than sage and the quality is still good I'll definitely take it
>>
>>109444285
IMO the ones in the comfyui templates are fine. I only changed the save node for now (to use VHS video combine).
>>
>>109444155
10gb its working and good
>>
File: 1778094800444549.png (424 KB, 405x720)
424 KB PNG
>>109444295
LMAOOOOOOOO I KNEEL
>>
>>109444295
wow how did you get the amateur gay porn aesthetic down pat
>>
>>109444295
>didnt make chang say anything
>didnt make the jews face become nervous as he approached
what could have been
>>
>>109444308
he's an artist.
>>
>>109444295
really impressed by the text, jesus this model has no weaknesses
>>
>>109444308
Model is smart. Context alone.
>>
>>109444283
The china man is the man of the future, I kneel.
>>
>>109444314
>this model has no weaknesses
Gen times
>>
So it can do porn. Oh boy.
>>
>>109444295
>https://streamable.com/v1o0rs
lol
>>
My GPU is taking a pounding
>>
flux can still win. j-just you wait
>>
File: perfect.png (122 KB, 498x311)
122 KB PNG
>>109444295
>oy vey
>>
>>109444333
Okay ZITChad. Take your meds..
>>
>using 90% of 64gb ram and most of my vram with a 4080 to gen
...is this normal?
>>
>>109444341
no, but on meth it is
>>
File: 492761030.png (834 KB, 512x1024)
834 KB PNG
>>109444341
yes

https://streamable.com/as7eoc
>>
>>109444341
21GB model + 26GB TE + 5GB video VAE + 0.5GB Audio Vae + the latent which is prob like 5-10GB
>>
>>109444333
I see no universe where they can compete, if I were them I would not release anything locally, Flux 3 max already looks like a Sora 2 replacement, they'll get their niche
>>
protip: set the megapixels to 0.2 and it gens super fast, good for testing concepts. 2.67s/it for 20 steps. of course, it will be potato resolution, but still.
>>
I have a 5060 Ti 16GB, 64gb ram. Using H3 (the default comfy template models, 41gb total size), it's taking over 40 fucking minutes. Left everything default, so 480p for 5 seconds, and my input image is about 480p too.

What the fuck am I doing wrong? I am a retard when it comes to this shit and just wanna make memes. It's stuck on "SamplerCustomAdvanced". My instructions were simple too, less than the given template.
>>
>>109444352
you shut your filthy whore mouth. I need local sora 2
>>
>>109444362
you won't get local sora 2, "dev" is always significantly inferior to "max"
>>
>>109444356
0.4 + RTX upscale node is good enough

>>109444361
you are overfilling vram with latent likely, use smaller res or use FNN chunking
>>
>>109444356
see, it works:

https://files.catbox.moe/f35eol.mp4
>>
>>109444347
>it can also do toe flexing
damn, I could settle for this model for the rest of my life and not care if there won't be anything better, it's already perfect for my use cases
>>
>>109444366
actually dev has always been like 95% of max, just with no commercial use
>>
>>109444033
https://files.catbox.moe/v83m0l.mp4
>>
>>109444333
>>109444352
they can compete if they redistill a new model without cucking it, but they wont do that
>>
Has someone tested if you write the sound in english but tell the model to make the characters speak japanese or whatever language works? Or do you have to go to the effort to translate all the speech?
>>
need some biblical event kino
>>
>>109444382
hate when that happens
>>
>>109444333
H3 is currently illegal to use in the USA and Europe. They've already won
>>
>>109444369
>use FNN chunking

How do I do that? No clue where to even begin fucking around with the nodes, this is way above my iq, but I know my specs are good enough that the speed is way off
>>
>>109444394
Fud. You can use it locally. Just can't use it commercially. Stop spreading fud, ZITFag
>>
>>109444394
their terms of use aren't law
>>
>>109444400
>everyone I don't like is a zit enjoyer
obsessed
>>
>>109444400
he really does this to every model that he can't run. it's a pathetic existence
>>
>>109444278
no that was i2v
>>
>>109444394
>illegal
didn't know Minimax was a government
>>
File: 998021776023818.png (2.03 MB, 1344x1339)
2.03 MB PNG
>>109444361
>>109444369
Maybe you have other stuff filling VRAM, cuz I got the same card and this 7s vid took ~420 seconds.

https://files.catbox.moe/c2cods.mp4
>>
>>109444341
You paid for that shit, software better load it
>>
t2v test, 0.2 megapixels (4080 16gb/64 ram), 75s gen time. KINO.

https://files.catbox.moe/icxpke.mp4
>>
>>109444413
you asked her to eat it? kek
>>
Wonder what's lowest step count that still gives acceptable quality
>>
Also the BF16 is base model, not distilled.
>>
>>109444422
25
>>
>>109444416
Hold up, Migu is spitting facts!!
>>
>>109444422
lower the megapixels to test ideas fast, then raise it to bump quality
>>
https://litter.catbox.moe/x7ff2ujxvhxhrmcy.mp4
>>
>>109444433
typo in the prompt
>>
>>109444423
there's no base model, we only got the guidance distilled models
>>
>>109444382
kek
>>109444389
real life continuous unedited go-pro footage. first person eye perspective. modern combat footage. high resolution. impressive audio. very shaky and dynamic movements. lots of motion blur. our perspective is an american soldier on a street in new york city. the city is a war zone. missiles are launching from the city ground and flying into the clouds. they leave behind thin smoke trails. our assault rifle is held lowered at the bottom of the frame. battlefield noises coming from the city.
there is a loud shrieking noise as a rocket hits across the street, causing a big loud dusty explosion and the camera to violently shake as the shockwave covers the camera with debris and leaves a deep crater in the road. the camera becomes frantic. dust fills the street. american soldiers are in the background screaming in agonizing pain.
we run to the sidewalk and lean against a building with our rifle lowered at the bottom of the frame. noisy footsteps. very active combat. the street is full of dust. constant background gunfire noise
more dusty explosions detonate along the street, shaking the camera and leaving deep craters in the road.
the sky becomes very bright. loud trumpets reverberate from high in the sky, startling us and causing the camera to shake. we look up to see sunbeams coming down from the clouds. a large biblically accurate angel is flying high above the clouds. it has a large eyeball in the center with a ring of multiple large feathered wings surrounding it. the wings are spinning in a circle at lightning speed which creates a lot of motion blur. a deep voice booms with echo and reverb from the heavens and says "BE NOT AFRAID".
we make a distressed whimper noise while breathing. our fingers obstruct the lens.
>>
>>109444418
Nah, just lick, but I did prompt "ASMR sound" which is probably why it did that lol
>>
>>109444433
Her fingers appear to be fusing into her body.

ZIT wins again!!
>>
>>109444416
10 second test same prompt, 178s. I got a Migu anime concert, kek

https://files.catbox.moe/wdk0xu.mp4
>>
File: best model ever.png (560 KB, 1080x1080)
560 KB PNG
>>109444433
looks like it can do all fetishes right?
>>
>>109444433
So far the audio has been really impressive.
>>
i'll wait until you plebs make this shit run on 8gb vram
>>
File: 1774901250958731.jpg (542 KB, 832x1216)
542 KB JPG
>>109444382
hate when that happens
>>
>>109444446
>So far the audio has been really impressive.
that's really surprising because it's a really small audio model, they could have gone for something bigger desu
>>
>>109444436
devs said its not distilled, just was not trained with CFG
>>
File: 1776768630813403.png (759 KB, 1080x2294)
759 KB PNG
>>109444436
>>109444460
HOLD UP, WHAT?
>>
>>109444277
glorious
>>109444347
that didn't stay up long lol
>>
>>109444460
>not trained with CFG
didn't know that was possible, dunno if I have to find it cool or not
>>
>>109444444
holy GET
>>
lmao, I need to work out the quirks but wait for it.

https://files.catbox.moe/2vcaj7.mp4
>>
oh, and it will take about the same vram to train as LTX but will be about 2-4x slower per step. Still worth of course
>>
>>109444347
>>109444468
wadahell, it's not even anything bad

https://files.catbox.moe/584ncx.mp4
>>
>2 hours to download a simple Qwen text encoder
Kek, DOA
>>
>>109444433
https://files.catbox.moe/90sojz.mp4
worth it
>>
>>109444476
don't get my hopes up
>>
>>109444476
though that is WITHOUT the possible sparse attention training the devs mentioned
>>
https://files.catbox.moe/x9kcgk.mp4
Zoomer brainrot bullshit test, I saw a bunch of people genning shit like this with Tiktok's AI cast (which is actually just seedance)
>>
>>109444474
LowTrumpGod
>>
Can H3 do smut?
>>
Where gguf
>>
>>109444484
https://github.com/AkaneTendo25/musubi-tuner/issues/106
>>
Gen time is a little too painful for me so I will probably be waiting for optimizations/sparse attention/step loras. Model was amazing in the few tests I did though. In addition to overall quality, was surprised by how uncensored it is how good the prompt adherence seemed. Local somehow won.
>>
>>109444491
No. Stick with LTX
>>
>>109444490
>Give your money to Israel NOW!
>>
where my big rig celebfags at? how is H3?
>>
File: 7446854.webm (2.99 MB, 448x256)
2.99 MB
2.99 MB WEBM
>>109444440
just as an example
>>
>>109444505
/r/stablediffusion says it's uncensored though?
>>
https://streamable.com/62crmp
it's a bit long to make a video, but once there's a turbo version of it I can see myself spending my whole days doing this shit, amazing model
>>
>>109444474
okay, you get jibberish if you dont state specifics. I made a super basic prompt with 2 dialogue lines. this is closer:

https://files.catbox.moe/16p6wc.mp4
>>
File: MM_00015_.mp4 (567 KB, 640x1152)
567 KB
567 KB MP4
10 steps seems to produce usable outputs.
>>
>>109444522
and no motion
>>
>>109444516
>He browses Reddit
>>
https://files.catbox.moe/fjgtga.mp4
>>
Where's the Q4/nvfp4 quants???
>>
>>109444533
what
>>
>>109444535
kijai says he is working on a new type of quant that is a better quality int4 convrot
>>
>>109444507
It knows the big names at like 60% likeness. Worse than Krea
>>
>>109444507
not here probably? but seeing how the anime waifu and other characters can be freely referenced...
>>
>>109444413
Fucking weird. I restarted my computer and still it's slow as. You using the default template?

I'm using "image to video" not "reference to video" if that might be the difference
>>
>>109444361
Are you using --disable-pinned-memory and --vram-headroom 1
I have the same setup and it's working normally with 2:3 @ 0.4mp
>>
lmao, we're entering a new meme age. and this is at the lowest preset (0.2). 70 seconds:

https://files.catbox.moe/7edill.mp4
>>
>>109444546
Worse than an already shit model?
You know what that means?
D
O
A
>>
File: Jen Snake.webm (3.95 MB, 710x1280)
3.95 MB
3.95 MB WEBM
>>109444369
Just putting RTX VSR in the workflow slows down my generation time bu a lot (this took 13m39s on a 4090). This is from 0.7mp to 2x ultra and back down to 720p for posting (it looks better after supersampling down).
>>
>>109444558
10 seconds. im dying.

https://files.catbox.moe/5uqrue.mp4
>>
>>109444558
lmaoooo, absolute kino
>>
>>109444580
Sage attention works fine for me, on a 5090. Each gen takes 20 seconds longer than the last though, until I restart it. Might still be some bugs to work out.
>>
remember when ltx had trouble with who was delivering what dialogue?
does h3h3 do that
>>
>>109444586
>>109444558
>it even has the DBZ sound
this is so good
>>
IT KNOWS GUNDAMS

kinooooooo

https://files.catbox.moe/cgmsnx.mp4
>>
File: 48567638.mp4 (3.57 MB, 544x768)
3.57 MB
3.57 MB MP4
>>109444553
Default workflow, also i2v. Maybe check if you have some comfy launch arguments that are fucking it up?
>>
>>109444613
that's why I'm a bit sad there's not a single unified video, you could combine i2v with references that would be so amazing
>>
>>109444613
Try making a kung fu fight scene with them, using better images of them as references (if you're using the reference model that is)
>>
>>109444617
can't you use the story board thing to do the same thing?
>>
>>109444617
? Don't you just prompt the reference model that one of your reference images is the starting image. I haven't tried it, but that's what I assumed.
>>
>>109444522
Where are you changing the amount of steps?
>>
>>109444614
Nothing I can see, but I checked the logs and sw this:

[WARNING] WARNING: You need pytorch with cu130 or higher to use optimized CUDA operations.
[WARNING] Unsupported Pytorch detected. DynamicVRAM support requires Pytorch version 2.8 or later. Falling back to legacy ModelPatcher. VRAM estimates may be unreliable especially on Windows

Guessing this might be an issue?
>>
https://files.catbox.moe/ty6pm6.mp4
first attempt (i2v), it's a bit long but it's doable, it's gonna be hella fast on turbo model, really can't wait
>>
>>109444630
no, reference images are just there to add any characters that you want but it's still a t2v process
>>
I wonder how much copyright content it was trained on.

https://files.catbox.moe/006lii.mp4
>>
So flva stands for first and/or last frame.
Ref2va is for reference images.
If I want just t2i which one do I get?
>>
>>109444637
seems reasonably good at preserving the identity compared to ltx.
>>
>>109444654
fl2va (first/last to video/audio)
>>
>>109444654
if you want regular t2i, you can go for flva without adding any frame, if you want t2i and add characters, go for ref2va
>>
>>109444654
flva, it supports t2i just fine and AFAIK that's what they intended t2i to be used on
>>
>>109444614
Yup, that'll do it.
>>
>>109444665 meant for >>109444636
>>
>>109444654
first last if you want it to actually use the frame. References is just that, refrences
>>
File: 599-5991841.png (153 KB, 860x602)
153 KB PNG
i just deleted 500gb of kinos. you will be missed
>>
File: Reference Images.jpg (2.68 MB, 2670x2000)
2.68 MB JPG
>>109444617
Just use reference images and prompt... see: >>109444579, you tell the model what the images are and how they're interacting.
>>
>>109444646
Yeah, but have you tried it? I don't see any reason it wouldn't work.
>>
considering no optimizing has been done yet it's still really good, audio and prompt coherence seems great, it will only get better.
>>
lmao even via i2v the video can somehow make a realistic chat

https://files.catbox.moe/bunz5z.mp4
>>
It's really good for ENF stuff, huh
>>
>>109444682
likeness seems pretty bad
>>
>>109444682
use t2v model with reference image input?
>>
>>109444665
Thanks for the help. Wouldn't have thought to look in that place without you prompting me. Strange how it didn't make a big fuss about it though
>>
>>109444440
>>109444512
https://files.catbox.moe/4tlibf.mp4
This one turned out horrible kek. The cod 4 running animation makes me laugh though.
You're definitely going to have to rework your prompts for H3
>>
>>109444696
also silly me, didnt change it from square to landscape. good for 1024x1024 base gens though.
>>
>>109444698
I bet it would turn out better if he included a close-up of her ugly face desu
>>
>>109444382
sweet
>>
>>109444698
likeness was the only thing that didn't impress me in my tests (first frame).
>>
File: screen.png (106 KB, 1404x162)
106 KB PNG
I can do 480p on my 8GB 3050 with 16GB of RAM

I'm satisfied and reluctantly surprised. Through some sort of meme magic (async offloading?) peak RAM usage during genning is 11GB and 7.5GB VRAM with the pruned int8 weights.It's not fast by any means, but waiting ~20 mins for a 5s clip is fine if I can still use my system for other stuff as it hasn't brought it to a crashing halt of leaking into pagefile/swap and leaving me with no physical RAM left. Doing ref model with multiple inputs (including video/audio input) slows it down a bit compared to t/i2v but not by that much.

Also learning the prompt syntax for complicated ref2vid tasks is kind of a lot. Spend 10-20 minutes just writing the prompt out.
>>
>>109444713
yeah, you're probably right
>>
>>109444704
hmm yeah it looks like they made it very good at following the visual descriptions perfectly. i guess i don't need to write things in an esoteric way like with ltx. thanks again for the kino
>>
>>109444698
Must be anon's workflow, the likeness with the reference to video is insanely good, don't know about lower than 1mp res tho.
>>
>>109444701
No worries brother, godspeed.
>>
Are you guys hermits or what? Can it do porn?
>>
Everyone is using the pruned model, right?
>>
how do i stop the noise seed from changing
>>
Reference model is so fucking good that if the character has like a pendant, clothing parts or anything that has to move it has realistic physics without even prompting anything, based xi.
>>
>>109444727
yep, it has the same level of quality as the full model
>>
>>109444727
Of course. You think I'm made of vram?
>>
>>109444730
i might use it to create different perspectives of a character so i can make better reference sheets
>>
File: MiniMax_H3_00013.mp4 (2.06 MB, 576x1056)
2.06 MB
2.06 MB MP4
>>
>>109444726
Try reading any of the thread.
>>
>>109444739
fucking reading and fuck you
>>
File: 8rly9x.png (336 KB, 405x720)
336 KB PNG
I kneel to H3, 5060 Ti 16 gb, 480p @ 5 seconds in 5 minutes, really good output.
10 seconds at a slightly higher resolution was 24 minutes though.
>>
>>109444748
If 4 steps loras drop relatively soon this is going to be beyond insane.
>>
>>109444748
Install sage attention. I got 16gb vram 64gb system ram and >>109444704
540p 10 seconds took 14.5 minutes.
>>
>>109444698
Yeah, it's about 50/50. I could probably pull the higher res original (which I think is 2880x4320) and crop in tighter on the face and add it to the reference images to fix it up. I'm still working on how to better create prompts from my images though (it's a little clunky putting it all together right now).

>>109444699
Just vanilla ref2va.
>>
>>109444506
You lost, trannie.
>>
Do 15 fps and you get faster gens. it still looks good
>>
>>109444780
Does it fuck with the sound?
>>
Any tips on prompting?
>>
>>109444413
>ai can only make disgusting roastie pussies
lmao
>>
File: 1764818570787799.png (33 KB, 899x618)
33 KB PNG
Let's do this shit
>>
>>109444786
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

But use gemini
>>
A man is in a Walmart store. The man is grabbing lots of boxes of Pokemon cards, and Arnold Schwarzenegger walks in and shoots the man with a shotgun, and he drops to the ground and all the pokemon card boxes also fall.

https://files.catbox.moe/zyow31.mp4
>>
>>109444413
Umm sweetie this violates the ToS
>>
>>109444791
that was at lowest size. now i'll try 0.9mp.
>>
where's kijai
KIJAI
SAVE US
>>
File: image (13).jpg (134 KB, 1257x1002)
134 KB JPG
it can do image editing
>>
>>109444799
He already saved us
>>
>>109444790
Thank you, I've never been able to master a T2V model..
>>
will light tricks make a speed lora for a model that's likely wiped them off the map?
>>
>>109444780
How?
I don't see any exposed fps settings other than when actually saving to video.
At least in the reference workflow.
>>
>>109444803
flux bros... we fricking lost...
>>
>>109444815
in the subgraph
>>
File: tumble.webm (931 KB, 736x576)
931 KB
931 KB WEBM
>>
>>109444803
Can you edit an image with multiple references?
>>
>>109444799
but kj boss contributed a lot already and it probably works unless you have one of the very worst machines in the bread? just try it.
>>
should i bother downloading at 8+32 RAM? seems like it might not fit
>>
>>109444821
I don't see it.
You mean the math expression node?
>>
File: MiniMax_H3_00014.mp4 (2.53 MB, 608x1088)
2.53 MB
2.53 MB MP4
>>
>>109444828
It's worth a shot.
>>
>>109444803
I look like the girl on the right
>>
https://litter.catbox.moe/wfle0fthl1wm2bp1.mp4
>>
File: 1770743150954657.png (42 KB, 978x184)
42 KB PNG
>>109444828
A 3060 could run it
>>
>>109444791
and now at 0.9mp, it took 500 seconds but the quality is WAY higher.

https://files.catbox.moe/s6dqtp.mp4
>>
File: MiniMax_H3n_00007.mp4 (2.15 MB, 608x1088)
2.15 MB
2.15 MB MP4
couldn't get completely smooth armpit
>>
File: 1785733317159405.png (3.86 MB, 1920x2410)
3.86 MB PNG
Would one of you be so kind as to take this anime girl and make an image of her crying in the bathroom mirror and about to cut one of her ears off with a pair of scissors?
Thanks if anybody does.
>>
what quant is that model?
>>
What's the longest it can be, before it takes an hour to generate? Ideally under 20 minutes.
>>
>>109444778
>being unironically antisemitic
Zion don doesn't approve of your opinion, migatard.
>>
im better than that no talent retard nolan

https://files.catbox.moe/m7ahnb.mp4
>>
>>109444860
lmao, this is good
>>
Should I go for in8 or pruned int8?
>>
>>109444914
pruned, the quality is the same and it's a 20b model instead of the 33b model
>>
so this isn't even distilled? we will get faster speed than ltx once that happens?
>>
Yeah, this model is gonna go a long way.
>>
0.3mp, this is great.

https://files.catbox.moe/ih8g53.mp4
>>
>>109444922
it seems like it's a "base" model and they trained it without CFG, didn't know that was possible but there you go
>>
File: 1769295774036476.png (391 KB, 621x473)
391 KB PNG
>>109444929
HE CANT KEEP GETTING AWAY WITH IT!!
>>
All it needs is a turbo Lora. The slow gen time is the only bad thing.
>>
>>109444931
>trained it without CFG
That's the default though? You can always add negative prompt at the cost of gen speed
>>
>>109444931
hope they do a distilled one. i don't want to wait a long time for generations considering it takes 2 minutes for distilled ltx to make a 20 second 480p video
>>
The GPU screams in agony as I attempt to insert 120 extra frames into it's VRAM
>>
Can H3 do T2I? I don't care about animation
>>
>>109444938
at 0.2 mp scaling you can make stuff fairly fast, for now it's great for fast gens

the anime character says "I came here to laugh at you." in japanese. He points and laughs at the camera.

https://files.catbox.moe/p3f277.mp4
>>
H3 is like a human reading the prompt and immediately understanding what you meant, it's insane.
I've been very unclear with my prompts and it still gives you an ok result of exactly what you asked that you can then add more details too.
It works how I expected video gen to work when I found out about it years ago.
And I'm on a 3060 12gb.
>>
>>109444948
Yes.
>>
didn't realise it was gonna be a scorcher today, fuck
I guess I can sleep though it, it's 9am kek
>>
>>109444948
>Can H3 do T2I?
yeah, you go for 1 frame



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.