[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: collage_1785808372.webm (3.2 MB, 2048x894)
3.2 MB
3.2 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109449543

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
fucking finally.
>>
thx 4 bake
>>
>>109454191
Finally! Thanks a lot baker.
>>
Blessed thread of frenship
>>
File: MiniMax_H3_00024_.mp4 (1.94 MB, 736x416)
1.94 MB
1.94 MB MP4
>>
breast thred of frenshit
>>
TYVM for the proper bake

Now I am more inclined to post my T2V gens, which I shall do every ten minutes or so
>>
Damn, 3 gens in the collage...
Feels good to be the king...
>>
>>109454191
>skipping the will smith schizo bakes on Previous
how can a man be this based?
>>
>>109454200 (me)
holy shit my collage is a fail, much rather have yours op thx
>>
where jencon
>>
>>109454200
when does she start moving?
>>
Read this if you want to do ref 2 vid
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
>baking before bump limit
how did he buck break this general so hard?
>>
>>109454210
>holy shit my collage is a fail
wide images/videos tend to make the layouts odd it seems
>>
>>109454218
why do you talk about yourself in the third person?
>>
>>109454218
you kinda cute when you mad bby
post some gens for me girl
>>
>>109454207
>which I shall do every ten minutes or so
kek'd
>>
>>109454218
willbur smith is the literal definition of buckbroken.
no one wants to post in a gay cuckold general.
>>
>>109454231
i did, it's in the collage
>>
I wish H3 would generate single frame image like Wan did
The minimum seems to be 5 frames (it can only generate 17k+5)
>>
File: MM_00048_.mp4 (1.24 MB, 960x544)
1.24 MB
1.24 MB MP4
I thought it would know more fast food mascots.
>>
>>109454218
> baking before bump limit
30 seconds after actually
pretty good reaction
>>
>>109454218
bump limit was reached literal hours ago
>>
File: MiniMax_H3_00030_.mp4 (1.69 MB, 576x736)
1.69 MB
1.69 MB MP4
>>
>mfw Resource news

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

>Raylight 1.7.2 Adds MiniMax Support, 2x Speedup
https://github.com/komikndr/raylight/releases/tag/1.7.2

>Scaling Properties of Text Conditioning in Visual Generation
https://heheyas.github.io/context-scaling

>Retrieval-Driven Training-Free AI-Generated Video Attribution
https://github.com/renxi-seu/Video_Attribution

>A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
https://github.com/zfu006/SSG

>ComfyUI MiniMax H3 Image Studio (Experimental)
https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio

>MiniMax H3 — NVFP4 (Blackwell)
https://huggingface.co/lilcheaty/MiniMax-H3-NVFP4

08/02/2026

>MiniMax H3
https://huggingface.co/MiniMaxAI/MiniMax-H3

>MiniMax H3: Repackaged model files for ComfyUI
https://huggingface.co/Comfy-Org/MiniMax-H3

>MiniMax-H3-INT8-CONVROT
https://huggingface.co/Gluttony10/MiniMax-H3-INT8-CONVROT

>MiniMax H3 COMMUNITY LICENSE AGREEMENT
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE

>LoRA Dataset Studio: LoRA workflow in one tab
https://github.com/perfectgf/lora-dataset-studio

>comfyui-vram-tracker
https://github.com/PuppetMasterAI/comfyui-vram-tracker

>Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
https://nvlabs.github.io/Sana/Sol-Attn

08/01/2026

>EU to get power to enforce rules on AI starting today
https://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786

>FameGrid Auto Color for ComfyUI
https://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui

07/31/2026

>Introducing Seedance 2.5
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5

>ReToken: One Token to Improve VLMs for Visual Retrieval
https://github.com/avaxiao/ReToken
>>
>>>/wsg/6207600
Upscaling with seedvr2 works really well.
>>
>>109454256
>EU to get power to enforce rules on AI starting today
kek
all regulation, no innovation
>>
>>109454245
>>109454252
was it 30 seconds or several hours? typical incoherence from a schizo
>>
>>109454243
it knows Goatstanza that's all that matter to me >>>/wsg/6207386
>>
>mfw Research news

08/03/2026

>MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
https://orange-3dv-team.github.io/MoRoute

>WaiT for the Signal: Simple Frequency-Aware Flow-Matching
https://arxiv.org/abs/2607.28760

>Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
https://arxiv.org/abs/2607.29025

>Visual Distribution Anchoring for Efficient Prompt Tuning
https://arxiv.org/abs/2607.28967

>MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
https://arxiv.org/abs/2607.29180

>RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
https://arxiv.org/abs/2607.28974

>SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation
https://arxiv.org/abs/2607.29367

>When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
https://arxiv.org/abs/2607.29240

>Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs
https://arxiv.org/abs/2607.29412

>Domain-Adaptive Deep Joint Source-Channel Coding for Image Classification
https://arxiv.org/abs/2607.28907

>Explaining AI-Image Detection: What the Heatmap Actually Shows
https://arxiv.org/abs/2607.29581
>>
>>109454262
I don't care what you say since I wasted your Christmas. :)
>>
>>109454220
>wide images/videos tend to make the layouts odd it seems
vibecode in a single check that collages have to be 1:1 square aspect ratio or wider
>>
>>109454256
>>109454266
thanks!
>>
>>109454262
Schizoposting isn't going to make Yan take you back.
>>
>>109454243
>including Ronald McDonald, Wendy, The Burger King, Colonel Sanders, and Gustavo Fring from Los Pollos Hermanos
>jollibee holding a machine gun
>>
File: 3565.png (28 KB, 1852x125)
28 KB PNG
wow anon was right. use the FL model if you only care about t2v. it cut my time by 3-4x
>>
> >109454272
> >109454266
> >109454256
>nigbo resorting to samefagging again
>>
Is Minimax H3 compatible with smaller qwen text encoders or was anon in the last thread running inference with a different video model?
>>
File: 64757.png (22 KB, 954x271)
22 KB PNG
>>109454280
yes
>>
File: 74546.webm (815 KB, 320x320)
815 KB
815 KB WEBM
>>
File: 00025-3789773617.jpg (454 KB, 2880x1920)
454 KB JPG
nobody cares about krea2 anymore?
>>
>>109454286
? these are just different quants of same model? it doesn't say how many parameters
>>
>>109454236
Prove it :]
>>
is the honeymoon period nearing the tail end yet?
>>
>>109454296
you aren't entilted to anything, you have to understand that
>>
>>109454293
krea 2 h3
>>
>>109454293
sasori please lurk for three more years before you ever post here again
>>
>>109454296
he's just gonna post my catbox and claim it's his just like he used to do to boost his shitty trollbakes

he's never made a gen that isn't his avatarfag mascot in his life
>>
>>109454293
H3 is a more groundbreaking model than Krea
>>
>>109454293
i'm still genning kreano
>>
>>109454293
it's my initial frame generator
>>
>h3
>k2
>anima

The perfect trio.
>>
>>109454309
krea + guano = kreano
>>
>>109454297
it's only just begun
>>
>>109454295
oh, i didn't know what you meant. i have no idea, i assume not since people would be using those instead of the big ones. most of the generation time is spent in the denoising so i don't think there is much point in using a smaller text encoder if you can run it already
>>
>>109454103
Oh yeah for sure. I saw and played a bunch of rather barebones games made with Opus, Fable, etc. And they play rather well. they aren't *great* in the sense of being an actual thought out game with loads and loads of things to do but they the beginning of something. However all the 3D assets those games uses are pure trash. But I guess that will be better with time.
>>
>>109454297
not even close, for the moment I can only make videos while waiting more than 10 mn, once the turbo lora comes out the spam will be 10x stronger
>>
>>109454315
uhhhh actually it is a portmanteau of krea and KINO
>>
>>109454262
it could be before bump limit actually
the previous thread at the moment is between
22:27:11 >>109454263 and 22:25:41 >>109454249 on >>>/g/
so clearly not several hours
>>
>>109454314
>The perfect trio.
i remember when people used to merge 70b local LLMs with themselves to make goliath-120b monsters that were like 5% better than the 70b models themselves. wonder if thats possible for h3 and some 96gb vram card chads can get closer to sora2 at home than anyone else. wonder if fable (or a local model like kimi k3) can code that
>>
>>109454308
what ground does it break
>>
>>109454331
sora is already destroyed by this
>>
>>109454331
are vidgen models even remotely like textgen LLMs?
>>
>>109454331
>merge 70b local LLMs with themselves to make goliath-120b monsters
I don't know how it works in text land but in image land you can only merge models derived from the same arc. And only if all their components are the same, so no "merging multiple models to get a bigger one".

You merge two models together and it comes out the same size as a single model.
>>
LLMs have a lot of layers and a lot of that stupid fuckery involved adding new layers or wiring layers back to prior layers or something like that, do video/image gen models even have layers?
>>
File: output.mp4 (1.31 MB, 576x896)
1.31 MB
1.31 MB MP4
>>
File: miguthanks.mp4 (801 KB, 864x480)
801 KB
801 KB MP4
>>
cozy breas
>>
we need an ejaculation lora
>>
How do I stop H3H3 doing slow motion? Add more stuff happening?
>>
Getting SamplerCustomAdvanced errors on ComfyUI, what am I doing wrong? Using the img2video workflow I found on huggingface.
>>
>>109454336
>>109454343
Here's what the #9 model in the world (Deepseek-V4-Flash-0731) had to say

Yes, the Goliath trick transfers to diffusion/video models in principle, but in a weaker form. Merging is weight-space math (SLERP/averaging), not architecture-stacking, so you only get a blend, never a bigger net. The method is mainstream in the SDXL world and backed by papers like Diffusion Soup, and it gives the same modest "soup" uplifts LLMs get. Crucially, it only works between fine-tunes of the same base model with matching architecture, tokenizer, and latent space.
For MiniMax-H3 specifically, it's not practical right now: it's a single 33B DiT (nothing to stack into a "120B"), it ships only two different-task checkpoints (FL2VA and Ref2VA) rather than parallel same-task fine-tunes, it's CFG-distilled, and its quality-driving Context-IR/2K modules aren't open-sourced. The realistic way to get the soup effect on video would be to fine-tune H3 for a few styles, then blend those checkpoints.
>>
File: 00035-3049170554.png (2.67 MB, 1280x1920)
2.67 MB PNG
>>
>>109454384
Yeah that first part is exactly what I said.
>>
>>>/wsg/6207618
https://files.catbox.moe/78uhfu.mp4
>>
>>109454278
what is FL model?
>>
>584s for 10s vid
>191s for 5s vid
why
>>
>>109454413
exponential weight
>>
>>109454412
first and last frame model. the one that doesn't let you use references
>>
>>109454413
exponential not linear
>>
>>109454420
why would it not be linear?
>>
>>109454418
>>109454420
I see :/
>>
>>109454418
>>109454420
>>109454426
>exponential
>linear
attention is quadratic you dumb fucks
>>
>>109454381
It doesn't do slow motion though? Did you change the FPS?
>>
Attention is linear for human
>>
Can I handle minimax? My sissyclitty only has a 3070 super and 32GB ram.
>>
>>109454443
They ran their demo on 3060
>>
>>109454426
Jews.
>>
>>109454433
>attention is quadratic you dumb fucks
then get it a wheel chair. no need to be RUDE.
>>
>i just realized i've been using the RF model for FL
how the fuck was i s'posed to know the damned model needs two separate checkpoints to know how to use more than one image at a time
anyone also curious why there's a speed difference between the two; you might be making my mistake
>>
>>109454449
which had almost 2x the vram kek
>>
Whatever happened to that new super duper attention that was like rotating and I think logarithmic or some shit everyone was hyping around the time of Wan 2.2?
>>
>>109454352
Fucking throat goat
>>
>>109454459
>anyone also curious why there's a speed difference between the two; you might be making my mistake
because it has to constantly sample whatever you are inputting as reference through out the whole denoising process, even if it's just random noise
>>
>>109454459
>he didn't read the README.md
Let me guess, you signed the Terms of Service without reading it too?
>>
File: chad horse.jpg (35 KB, 736x971)
35 KB JPG
>>109454472
>he signed the terms of service
you sir might be more stupid than I
>>
>>109454419
>first and last frame model. the one that doesn't let you use references
so you're saying if you only care about t2v you use the first and last frame model without a first or last frame and you just get a faster t2v model?
>>
>>109454472
I didn't sign anything? Whatchu talm bout nigga
>>
>>109454480
based criminal
>>
Summer is dreadful
>>
File: MiniMax_H3_00032_.mp4 (540 KB, 576x736)
540 KB
540 KB MP4
>>
>>109454467
you mean radial att? it was a meme that only worked for tv2 and had really strict requirements so nobody used it.
>>
>minimax_h3_fl2va_pruned_int8_convrot.safetensors

so this is bad for i2v?
>>
>>109454481
yes the FL model doesn't require you to use any frames. you can use pure text
>>
>>109454486
buy air conditioner
>>
>>109454492
it's good
i2v is not r2v
>>
>>109454486
wrong
https://www.youtube.com/watch?v=ayE6Shlv598
>>
>>109454498
>air conditioner
does that help get rid of all the new faggotry
>>
File: generadeded.mp4 (3.25 MB, 1344x768)
3.25 MB
3.25 MB MP4
>>109454492
no you just use an image as the first frame. dont need a last one
>>
>>109454498
I mean it doesn't matter if you have an air con or not if the thermostat is in another room
>>
https://github.com/BobJohnson24/ComfyUI-INT8-Fast

use this instead of the regular load model
>>
>>109454492
A bit slower. If you only want T2V or I2V, stick to the First Last Frame mode.
>>
File: 1774970658106479.webm (1.99 MB, 960x544)
1.99 MB
1.99 MB WEBM
this shit is incredible, t2v only, but how do i prompt to have no subtitles?
>>
>>109454507
i've been using the comfy int8 convrot i'm gonna try some in4 convrot now, dunno how much worse it's gonna be but if it's faster it's gonna be nice
>>
>>109454507
there is no reason to use this if you're on the most up to date comfy.
>>
>>109454506
>thermostat
?
I have 5 air conditioners in my house
>>
>>109454507
>use this instead of the regular load model
why?
>>
>>109454507
>https://github.com/BobJohnson24/ComfyUI-INT8-Fast
What do you set the model type as? I saw a fork earlier that had a development branch with minimax support supposedly but I refreshed and the branch was gone. I figured it didn't support minimax.
>>
>>109454514
still have to try it

https://www.youtube.com/watch?v=5ruR0omB9zc
>>
>>109454513
Just put no subtitles....
>>
>>109454514
int4 convrot is way worse quality. not worth it.
>>
where gguf
>>
>>109454531
i've seen some guy report 2x speed with sage attention too.
wonder if you get 4x by doing both.
>>
>>109454535
make it yourself. there's scripts out there
>>
>>109454539
int8 convrot (dont use the custom node, its built into comfy now) + sage attention patcher + sol attention for the start steps + easycache = fast as fuck
>>
>>109454549
sol attention?
>>
File: i_00183_.png (1.21 MB, 768x1376)
1.21 MB PNG
>>109454549
can some based nigga post the workflow pls
>>
>>109454549
>its built into comfy now
what's the node name, can't seem to find it.
>sol attention
?
>easycache
i don't know that either.
>>
>>109454507
so git clone that URL to custom nodes, then add it in place of the subgraph load model node, including easycache/sage attention kj nodes like before.

need to test more but it is working, but idk yet how much faster it is, gonna try a 10s gen.
>>
File: coooomin.mp4 (669 KB, 576x896)
669 KB
669 KB MP4
>>109454368
It won't take much.
>>
>>109454556
you are not entitled to anything
>>
>>109454563
bandoco discord, its custom nodes you need
>>
>>109454566
you could probably prompt the consistency and color/translucency to make it look less solid like that.
honestly probably doesn't need a lora
>>
File: 1769070411849236.png (3.14 MB, 4088x4088)
3.14 MB PNG
it's good but it's not seedance good, that's a stretch
>>
>>109454563
its not a node, just regular model loader, it does int8. "fast-int8" custom node is slower now
>>
for any faggot that uses api videogen aside from sneedance 2, how does H3 compare?
>>
https://files.catbox.moe/ppvro0.mp4
>>>/wsg/6207628
>>
>>109454580
Well yeah that's why its two points below seedance you dummy baka anon
>>
>>109454580
its better at far more styles and allows IP
>>
>>109454580
It lives inside my vram so it better
>>
>>109454591
>he doesn't know about the statistical significance
oh boy, you're pretty retarded aren't you?
>>
>>109454581
oh ok cool thanks.
what about sol and easycache?
>>
>>109454599
if you're so smart name three statistical significances
>>
>>109454505
lmao
>>
WARNING: the math expression node has hard values baked in.
max(5, round(a * 24)) + (5 - (max(5, round(a * 24)) % 17)) % 17

the 24 is the fps. if you use a different frame rate, you must change this.
>>
>>109454580
>just two points behind seedance
local is eating unimaginably good
>>
>>109454591
>>109454592
>>109454598
It's comparing aganist Seedance 2.0
Seedance's current version is 2.5
>>
File: fa.mp4 (297 KB, 352x608)
297 KB
297 KB MP4
>>109454580
i'm sure this benchmark is precise and objective to that degree
>>
>>109454614
I did change the math but I don't think the audio likes it when you change the framerate.
>>
>>109454617
Still man Seedance 2.0 at home is a huge jump. Wan 2.2, the last best open source model isn't even on that list. LTX ain't on there either.
>>
File: 00058-4132071726.png (2.63 MB, 1920x1280)
2.63 MB PNG
>>109454580
lets not kid ourselves here. h3 is nowhere near the quality of seedance 2.0 even for nsfw but its the best video model for local open source.
>>
>>109454580
remember when seedance 1.5 and veo 3 seemed impossibly out of reach for local? i member
>>
>>
how many steps do you guys use for H3?
>>
>>109454629
>h3 is nowhere near the quality of seedance 2.0 even for nsfw
true, seedance can't even do nsfw
>>
>>109454632
The best API model (seedance 2.5) is still pretty out of reach though, don't get me wrong, local improved a lot with H3, but it's still not there
>>
>>109454617
>>109454629
holy cope batman
>>
>>109454629
I look like this irl
>>
>>109454629
lol gen at 2K res and It IS seedance 2 quality
>>
>>109454530
it doesn't work.........
>>
>most uncensored model to date
>literal peanus in vageen on day 0
>"eh its not that good"
LMAO
>>
>>109454665
it's seething bfl employees
>>
>>109454649
no one can do that locally
>>
>>109454673
>sneedance shill trying to deflect blame from api cucks
>>
>>109454665
I'm running it on my toaster at .2 megapixels and using single line prompts and its bad. I happen to be an expert on this subject.
>>
>>109454637
>Wan2.2_i2v_00003__prob4-2x-RIFE-RIFE4.0-32fps
u copied the video combine node from your old wan wf?
>>
>>109454675
lol, 10 mins on 5090, with turbo lora it will be quick
>>
>>109454640
yes it can. seedance 2.0 can do nudity and very super detailed nudity.
https://litter.catbox.moe/4kxxbu.mp4
https://litter.catbox.moe/uggd15.mp4
https://litter.catbox.moe/hqo14i.mp4
https://litter.catbox.moe/2c48pm.mp4
>>
File: 1766273751959007.png (296 KB, 1716x738)
296 KB PNG
>>109454565
the reddit post about it noted people with 3090s noticed an increase, I also do on a 4080.

138 seconds (10s 0.3mp) with the int8 fast node vs 173.65 seconds original (default load model, no int8 optimization), using node -> patch sage kj -> easycache

second run with new node: 153.43 seconds, and then genning again, 135 seconds.

give it a try anons, it's like downloading free RAM.
>>
File: MM_00058_.mp4 (805 KB, 544x960)
805 KB
805 KB MP4
jake paul, at home?
>>>/wsg/6207634
>>
>>109454680
I made it with wan2.2, I just got into this shit the other day so I haven't even tried h3 yet
>>
so on my 4090 with sage attenion i get about 8.8s per iteration, with 0.4 megapixels, what about you guys?
>>
>>109454677
kek
>>
>>109454693
Do you happen to be twelve years old?
>>
>>109454693
>>109454637
now we are full ai cringe slop flood, bandoco h3 gens channel is somehow worst
>>
>>109454699
i don't know what a mega pixel is
>>
File: 187183.mp4 (3.75 MB, 640x832)
3.75 MB
3.75 MB MP4
>>109454629
For action it's not as good, but for most other things it's actually pretty close.
>>
>>109454644
not telling you spend loads on money for seedance 2.0 on some third party website but your insane if think h3 in its current state is better than seedance 2.0.
https://www.reddit.com/r/VeniceAI_uncensored/comments/1vddlyh/seedance_can_do_so_much/

https://www.reddit.com/r/VeniceAI_uncensored/comments/1v7jt5y/seedance_cumshots_and_nsfw_solved_sort_of/
>>
Is there an accepted best set of comfy params for a 3090 running H3, or is it just "let the backend figure it out"?
>>
>>109454726
for example a 480p resolution video is 854*480 = 409920 = 0.409 MP
>>
>>109454726
its bigger than a kilopixel
>>
>>109454699
5090. 0.4 MP is about 2it/s with sage
>>
File: 98247968802435.jpg (32 KB, 360x360)
32 KB JPG
>>109454726
>>
File: ai_00051_.jpg (927 KB, 1712x2456)
927 KB JPG
>>
>>109454685
ComfyUI-INT8-Fast test: the anime girl is holding a slice of toast in her mouth and runs very fast towards her local highschool in Japan with the toast in her mouth.

https://files.catbox.moe/czs3df.mp4
>>
>>109454732
what a waste of money
>>
>>109454732
why do you always feel the need to shill seedance in the local general?
>>
>>109454739
wait, i think the total length affect the gen speed as well.
how long for a 10s gen?
>>
can't wait for the NSFW lora's so you can do thrusting better with H3
>>
>>109454762
I'm genning some hi res right now. Can't answer that.
>>
i want to open civitai but i know it will freeze my puter if i do
>>
>MiniMax is soon releasing sparse attention
When exactly though? I need that speed it's too long...
>>
File: Screenshot_3539.png (16 KB, 1478x130)
16 KB PNG
maybe 8s gens is the sweet spot. 5s is a bit too short and 10s takes over an hour to gen
>>
>>109454754
that was 133 seconds, this is 199 with easycache off, skipping less steps leads to a clear quality bump: in any case, the int8 fast node seems to be a speed increase vs default.

https://files.catbox.moe/z1xc7k.mp4
>>
>>109454787
the key is to gen at 0.2 or 0.3mp as it's super fast relative to 1.0. ive had gens at 0.2 even that have looked high quality, it depends.
>>
>>109454772
no worries.
so any good reason to use the ref2va instead of the fl2va or vice versa?

also how do you use easycache, after or before attention sage?
>>
>>109454796
im genning at 0.4...
>>
>>109454797
>also how do you use easycache, after or before attention sage?
the order never mattered.
>>
>>109454739
wtf this is faster than genning a single modestly big image on a 3090 with anima
>>
>>109454433
No one cares NERD
>>
>>109454803
thanks, do you use the default values for easycache?
also what is sol attention?
>>
File: MiniMax_H3_00002_u.mp4 (3.74 MB, 832x640)
3.74 MB
3.74 MB MP4
Second try of H3
25:44 on RTX PRO 6000
BF16
>>
>>109454800
use patch sage kijai node, use easycache, every bit helps desu
>>
>>109454797
I am not using EasyCache at the moment cause I don't want blurriness that it may introduce. Use Ref2V if you need a reference motion or audio. Otherwise, stick to FL2V
>>
File: 54678.png (251 KB, 1869x880)
251 KB PNG
>>109454787
10 second gens here
>>
>>109454810
>no sound
don't put that into waste anon, share this on a catbox or here >>>/wsg/6198872
>>
>>109454804
Yeah. This shit is indistinguishable from magic.
>>
>>109454501
>Sublime Summertime
Every girl under the age of 25 knows this song as a lana del rey song (It's not a bad cover)
>>
lmao the reference model has so many possibilities, I didnt say the full name of Ash so maybe thats why it didnt show him.
>the man in <Picture 1> appears from the ball

https://files.catbox.moe/7qth9b.mp4
>>
honestly 1MP seems like the best, at 7s or lower depending on your prompt. it gets everything right for me from specific details to motion.
>>
>>109454827
lmaoooooo
>>
>>109454827
OMG it worked. Ash Ketchum is the correct prompt. reference model btw.

The setting is an empty park in Japan. Ash Ketchum from the anime Pokemon throws a pokemon ball and says "niggamon, I choose you!". The ball hits the ground and the black man in <Picture 1> appears and says "fent! fent!".

https://files.catbox.moe/uail2j.mp4
>>
>>109454814
damn what the hell man...
>>
>>109454813
alright, thanks
are you still using attention sol?
>>
>>109454827
>>109454835
Actually burst out laughing
>>
>>109454507
Does this offer any benefit with 50xx cards?
>>
>>109454840
I am satisfied with the speed/qualify, so didn't try sol attention.
>>
the fuck is a convrot
>>
omg he even sounds like a pokemon.

https://files.catbox.moe/r42xqm.mp4
>>
File: 49848333.mp4 (3.63 MB, 672x448)
3.63 MB
3.63 MB MP4
>>109454810
Dope.
>>
File: finally some good shit.png (339 KB, 546x562)
339 KB PNG
>>109454827
>>
Is there a noticeable quality loss using qwen3vl nvfp4 vs int8convrot when using Minimax ?
>>
now replace him with ryan gosling and then kys
>>
I haven't found that easycache degrades quality.
>>
>>109454827
kekkus maximus
>>
>>109454866
Then you haven't searched very far.
>>
>>109454873
not him but how do you set it up optimally? skipping steps will lower quality to some degree
>>
>>109454814
Damn, what are you running on?
>>
>>109454827
gem
>>
>>109454815
Sent to /wsg/, catbox seems to be struggling at the moment
>>
>>109454879
12gb card
>>
>>109454796
gen low mp and then use topaz video and flowframes to upscale and increase fps
>>
This is the fate for those who refuse to share their workflows
https://files.catbox.moe/0jzf60.mp4
>>>/wsg/6207656
>>
>>109454885
How did you gen 30s? Officially it says 15 sec, so it can go over 15s?
>>
lmao it used the game UI, im dying.

The setting is an empty park in Japan. Ash Ketchum from the anime Pokemon throws a pokemon ball and says "niggamon, I choose you!". The ball hits the ground and the black man in <Picture 1> appears and says "fent! fent!". Ash Ketchum says "niggamon, use steal attack!". The black man says "fent!" then runs forward very fast. The Pokemon battle music is playing.

reference model/workflow.

https://files.catbox.moe/1tfevi.mp4
>>
>>109454873
I've A/B tested several gens on the same seed.
>>
>>109454898
keek
>>
File: 74514233.png (28 KB, 443x271)
28 KB PNG
>>109454878
Don't know, I just tried it with KJs settings.
>>
>>109454892
>t. redditor.
>>
A golden age of /ldg/
>>
this reference workflow is so fun and im only using images so far. you can plug in anything. it also can take audio AND video content as references. I assume you can character swap videos too.
>>
Why did chinks censor seedbox?
>>
no prompt with a nsfw image and it instictively does the nasty
>>
File: joan.jpg (217 KB, 1920x1080)
217 KB JPG
>>
>>109454922
seedance*
>>
so to recap speed increases:
>int8 fast load model node
>easycache
>patch sageattn kijai
all work together btw, also another test:

The setting is an empty park in Japan. Ash Ketchum from the anime Pokemon throws a pokemon ball and says "niggamon, I choose you!". The ball hits the ground and the black man in <Picture 1> appears and says "fent! fent!". Brock from the anime Pokemon walks by and says "you know he's gonna steal your bike, right Ash?"

https://files.catbox.moe/yzb860.mp4
>>
>>109454897
I typed 30 in the field and hit go. It might be VRAM limited. It stalled without producing a error when trying a higher resolution.
>>
GET THE MODELS NOW
jews are gonna get this pulled VERY FAST
>>
File: built_for_bgc.mp4 (770 KB, 576x896)
770 KB
770 KB MP4
>>
File: 1757066949548013.png (171 KB, 485x242)
171 KB PNG
>>>/wsg/6207665
https://files.catbox.moe/n83mi3.mp4
>>
>>109454948
wow he's literally me
>>
>>109454926
just tried this and it wasn't pretty
>>
File: FUCK.png (237 KB, 498x380)
237 KB PNG
>>109454948
>>
>>109454955
Where'd you get this video of me
>>
>i can only afford an RX 7900XTX
>can't get any image gen to work on windows
>>
>>109454947
r u dum?
>>
>>109454947
Its already too late nigga
>>
>>109454942
>>int8 fast load model node
its the same node as the one provided in default wf?
>>
>>109454976
nope try this one

https://github.com/BobJohnson24/ComfyUI-INT8-Fast
>>
>>109454942
>int8 fast load model node
You don't need this if you comfy and comfy kitchen is up to date.
>>
>>109454948
Wholesome and not sexual at all
>>
I hope you used a VPN to download the models if you are residing in one of the Excluded Territories.
>>
ok enough fent man, heres the ldg mascot in pokemon.

https://files.catbox.moe/ax558r.mp4
>>
>>109454984
Oh fuck are they comin for me?:!
>>
>>109454984
I'm from India (i've been hit by trains 3 times). Do I am in big trouble?
>>
File: MM_00061_.mp4 (1.81 MB, 544x960)
1.81 MB
1.81 MB MP4
>>
>>109454990
She should kick him in the balls
>>
File: improvements.png (576 KB, 814x578)
576 KB PNG
reposting. NOT AI

>>109454165
oc
>>
>>109454994
NA walls
>>
>>109454586
Holy shit, when you think this model can't wow me anymore lol
>>
File: file.png (244 KB, 554x322)
244 KB PNG
>>109454947
>PULL IT
>>
>>109454994
is that purely text to image or did you use a starting frame?
>>
File: 00010-3171941328.jpg (605 KB, 1920x2880)
605 KB JPG
>>109454948
i uploaded the leina vance lora a few minutes ago civitai
https://civitai.red/models/2832568/leina-vance-queens-blade-krea-2-lora
>>
>>109455009
purely text to video*
>>
Local is truly back if that api anon can only compare H3 with Seedance.
>It's not good as Seedance 2.0 blah blah
The fact we have something this close is giving apikeks extreme amounts of melty.
You can pretty much create whatever you imagine with H3 if you have more than half a brain, whereas with any api model you can't.
>>
>>109455009
pure T2V
>>
>>109454586
I actually find myself counting people's (real life) fingers.

kind of funny
>>
>TFW you waited 10 minutes for a H3 gen and you forgot to fix the aspect ratio.

Is there like some node that will fucking find the nearest supported aspect ratio and crop the image to fit it, this is so annoying. I guess I could vibecode it.
>>
File: 89464.webm (964 KB, 448x256)
964 KB
964 KB WEBM
>>109455024
can you share the prompt? i am still struggling to make it look like real life footage. the camera motions are stiff like video game footage and putting negative quality terms into the prompt doesn't really work either
dusk time. thin clouds. low quality microphone. bad quality go-pro footage. motion blur. blurry video. haze on lens. chromatic aberration. cockpit of a very large military jet. we are wearing a green flight suit. reflective glass canopy. our hands are on the joystick and throttle. the flight instruments are in front of our legs. the screens on the instruments have a green tint. green holographic sight over the dashboard. rows of buttons and switches on the sides of the cockpit. the nose of the plane is visible over the cockpit. our wings are not in frame.
we are flying very fast and close to the ground over a mountainous forest. snowy mountains in the distance. flat horizon. highly dynamic and natural shaky unstabilized camera motion. very suddenly, the dashboard becomes full of red warning lights that flash on and off constantly. loud beeping alarm. the perspective is startled and frantically looks out the side of the canopy. modern swept fuselage comes into view. in the distance, a surface to air missile is launched from a launch site in the forest which flies up into the sky while emitting a trail of smoke. the plane dives down to the ground while doing a fast evasive maneuver. the missile directly hits our cockpit and fills the frame with an explosion of fire and debris. violent shaking and blur.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.