[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Hopefully Janny Fine With This Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109492953

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 1755097654104992.png (87 KB, 953x252)
87 KB PNG
>>109493880
>This is a lie, it reduces output coherence slightly.
is that true? do you have a comparison video to show the quality's change?
>>
b t o' f
>>
>inb4 nigbo malware
>>
REAL THREAD HERE
>>109480220
>>109480220
>>109480220
>>109480220
>>
>>109493914
as usual thanks for the bake anon
>>
>>109493924
It technically does, but sage attention 2 is what is the most approaching from a free speedup you'll ever find in optimization for image and video generations.
>>
>>109493930
fuck off debo
>>
File: 5019299603.png (2.2 MB, 1122x1402)
2.2 MB PNG
gn ai frens
>>
>>109493943
y isnt this moving
>>
File: w a y g.png (215 KB, 500x500)
215 KB PNG
>>109493943
gn to you too I guess...
>>
File: debo.webm (2.11 MB, 1143x2048)
2.11 MB
2.11 MB WEBM
>>109493930
>>
File: Untitled.png (274 KB, 3296x912)
274 KB PNG
if you're doing low megapixels (0.3, 0.4, etc) then add an upscaler to get a big quality boost for only like an extra +30 seconds. It may depend on the upscale model though, 2x_Ani4K_Compact_35000.pth is really fast.

Not seeing any Out Of Memory issues.
>>
>>109493924
same exact seed and settings just with/without that node.
last frame.
>>
>>109493947
I am afraid of going through the trouble of setting up h3 and being let down by it
>>
>>109493955
Workflow?
>>
File: 1765152083770262.png (156 KB, 1546x661)
156 KB PNG
>>109493949
For the anons saying quality issues are tied to the vae, it's not.
then what is it? I guess even the minimax engineers don't have the answer
>>
>>109493961
>being let down by it
You will not be
>>
>>109493961
just look at what this model can do, if you feel this looks good then go for it
>>>/wsg/6209322
>>
>>109493963
That's engineering speak for
>it's normal and we won't do anything about it.
>>
>>109493955
why would i use this over seedvr?
>>
>>109493962
i have too many local custom nodes

Just add a new node for:
Upscale Image (using Model)
and
Load Upscale Model

then put them after the vae decode, but before the Create Video
>>
File: 1786090733508194.jpg (230 KB, 832x1216)
230 KB JPG
>>
how to continue a scene with ref mode? is it just as simple as saying use <picture 3> as your first frame?
>>
>>109493976
sneedvr is bait
>>
oops, I made a mistake prompting. and it turned out good, still.

<Picture 1> is the image reference file for Marciana. <Picture 2> is the image reference for 2B. HatsuneMiku and 2B are sitting on a couch watching a movie. 2B grabs the breasts of Marciana and smiles.

spot the error!

https://files.catbox.moe/xi91nc.mp4
>>
Babe wake up a turbo "v4" 600 steps got released
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/main
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main
>>
Why is this happening? I mean the distortion with hexagonal patterns. Do I need more resolution?
>>
how do i stop myself from making lewds of irl people
>>
>>109493997
>Do I need more resolution?
>284x360
fucking duh
>>
>>109493997
>284x360
>Do I need more resolution?
What do you think?
>>
>>109493989
and fixed with the proper name.

https://files.catbox.moe/ute39t.mp4
>>
>>109493998
that's probably the lesser of two evils, just cover your tracks
>>
File: 622393383188127.mp4 (3.66 MB, 736x576)
3.66 MB
3.66 MB MP4
>>
File: 1755499091266218.png (215 KB, 1701x889)
215 KB PNG
>>109493990
cool, I'll try that since the lightvx fags still haven't finished their job
>>
>>109493997
>Do I need more resolution?
Yes, that's the far most likely cause.

Well yes it could also be partly some of the cache turbo step distil sigma whatever speedup stuff.
>>
Stop being poor.
STOP IT.
>>
>>109494004
and since they are basically static variables you can just swap pictures to get a diff result.

it's a new world out there with this model!

https://files.catbox.moe/s7lbfs.mp4
>>
>>109493999
>>109494003
That's a crop. I tested with 0.4 and 0.6 mp. The problem happens specially in fast moving parts. I tested also removing spectrum (no lora here) and increasing the steps to 30.
>>
>>109494029
one more, but you get the idea. essentially, you dont need lora training with this ref model. you can even use pictures to copy styles (like a lora).

https://files.catbox.moe/8zx8vs.mp4
>>
File: 80sdarkfantasy.webm (2.78 MB, 864x480)
2.78 MB
2.78 MB WEBM
Tried doing the 80s dark fantasy aesthetic using reference images (reference model) but the results aren't too great in my opinion. Also prompting for VHS effects does not work. This may be a case where I2V is preferable.
>>
>>109494039
>you dont need lora training with this ref model.
that's definitely the future, a model that is insane at references won't need character loras at all
>>
File: file.png (2 KB, 318x45)
2 KB PNG
>>109493861
holy shit i went from 20-70sec/it to 3-1it/s with that node
how much is this going to fuck up the quality
>>
>>109494051
>holy shit i went from 20-70sec/it to 3-1it/s with that node
don't scream victory too soon, it just skipped some steps but it won't do that all the time kek
>>
>>109494030
i don't think this should happen with any regularity.

just in case update comfyui and custom nodes.
>>
we need to increase our foreign aid to israel so LTX 3.0 can save us with 1 minute 720p videos on 12gb of vram
>>
if you use spectrum reminder to update, works well
>>
>>109494010
nice.
>>
>>109494058
that's how it will go yeah, there's definitely room for improvement, the only way to surpass minimax is to get similar quality but smaller
>>
>>109493976
A lot more lightweight, seedvr2 is good but it's an extra 24 minutes for me. I'ma give this a try when I wake.
>>
File: 1785516873072940.webm (1.37 MB, 1172x720)
1.37 MB
1.37 MB WEBM
>>109494054
gotta start somewhere. skipping steps or not, thats the biggest speedup in any capacity ive ever had for the last 5 years doing this shit. christ it sucks shit being an amd user
>>
>>109494056
Since I was going for a retro feeling I changed to 4:3 and I think it is better, but I still think that the model (the video vae in particular) has limitations very fast movement.
>>
<Picture 1> is the image reference file for Bocchi. <Picture 2> is the image reference for Axl. The setting is a Guns N Roses rock concert, anime style Bocchi is on stage playing a guitar while photorealistic Axl holds a microphone. Bocchi is doing a rock guitar solo while Axl points at Bocchi and smiles.

low res test but you get the idea.

https://files.catbox.moe/72r0h8.mp4
>>
i like frogs, i wish anons did more pepe gens
>>
>>109494085
it shouldn't look this bad, wtf is wrong with your workflow??
>>
i know i should check the llms generated prompt before i submit it to h3 but i dont wanna

>>109494046
its almost there
>>
>>109494085
you probably listened to the dumbasses who told you to use beta57
or you're using 4 steps because "it's a 4step turbo lora" like an idiot
>>
>>109494085
i had a nice giggle
>>
Why does Comfy only use half of my VRAM and all of my RAM?
>>
>can't do vhs effects
>can't do alarms
any other epic fails?
>>
>>109493997
>>109494003
>>109494030
>>109494085
>>109494105
this is why if you request assistance without providing a catbox then you should kys
>>
>>109494095
0.6mp, now you can see the detail a lot more:

https://files.catbox.moe/uegs2v.mp4
>>
>>109494105
>the dumbasses who told you to use beta57
err_sde / beta57 at 10 step is totally viable and doesn't look like that.
>>
/ldg/ mogged by /lmg/ btw >>109494062
>>
>>109494115
serious question, has anyone tried prompting to H3 in chinese? I remember some old chinese models were only responsive in some cases if you prompted in chinese
>>
>>109494116
turbo lora deep fry.
>>
File: 1754849619888881.png (10 KB, 414x388)
10 KB PNG
>>109494108
same, first it reaches the peak, then it's like "oh I reached a limit, then I'll fucking go for 19gb of vram usage instead of 22 for example!!"
>>
>>109494125
I'm actually using er_sde right now, but it works better with 'simple' than with beta57 at lower steps in my comparisons.
>>
>>109494126
is this your first time browsing ldg?
>>
>>109494126
>1 girl dancing
>mogged
pepe with red question mark over his head
>>
File: 177306861.png (430 KB, 800x582)
430 KB PNG
>>109494126
>>109494101
truth nuke, total Chinese victory
>>
>>109494132
i tried translating things like "loud alarm beeping" to simplified chinese and it still didn't work. i wouldn't know if it was an accurate translation though
>>
>>109494126
Anon posts (real looking) pizza here lmg cannot compare
>>
>>109494126
why is an lmg anon using cope speedups tho
>>
File: comfy_shot.png (271 KB, 1605x795)
271 KB PNG
>>109494105
>>109494119
Nothing too extreme:
https://files.catbox.moe/mklyfa.mp4

Sage attentions is enabled with --use-sage-attention.
>>
File: MiniMax_H3_00529.webm (3.76 MB, 832x640)
3.76 MB
3.76 MB WEBM
>>
>>109494126
By the look of it, hairs haven't grown on your balls yet. And yet you flex with a anime girl.
Pathetic.
>>
>>109493990
>>109494013
the v4 step 600 ema is really impressive at 8 steps, the sound issues are solved and the skin isn't plastic anymore, didn't expect them to make such a solid turbo lora in less than a week desu, not that I'm complaining!
https://files.catbox.moe/qjpfcb.mp4
>>
>>109494199
She looks 23
>>
>>109494174
why are the faces so terrible
>>
>>109494182
amazing!
what kind of workflow are you using? do you use a prompt enhancer?
>>
>>109494174
you're not showing the number of steps
>>
do i use the EMA turbo lora or the non-EMA?
>>
>>109494182
can you share the video with sound? it's good
>>
>>109494219
ema >>109494193 >>109493990
>>
>>109494193
yeah I've been testing it and the video quality is really good in comparison to lightvx, however I think it has less movement than lightvx. I generated a pussy that didn't look like an alien.

I think it kinda feels like these loras are trying to balance and pick between movement and quality, like they can only have one.
>>
qrd on the 50 steps meme
>>
Can I put a generated video as a reference and tell it to keep going from there with my next prompt?
>>
>>109494224
ok i trust you
>>
>>109493924
>mfw Resource news

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo – OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view
>>
>>109493627
>>109493781
Coolest thing I've seen in years. Please make more of these
>>
>mfw Research news

08/07/2026

>Vorch-Omni: Multi-Task Orchestration of Sight and Sound
https://vorch-project.github.io/Vorch-Omni-project

>Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
https://vorch-project.github.io/Vorch-Streamer-project

>Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
https://vorch-project.github.io/Vorch-Director-project

>Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
https://vorch-project.github.io/Vorch-IR-project

>In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
https://arxiv.org/abs/2608.05237

>Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
https://arxiv.org/abs/2608.05976

>EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
https://arxiv.org/abs/2608.06231

>MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
https://arxiv.org/abs/2608.05878

>Wan-Animate-2: Pushing the Application Boundaries of Character Animation
https://humanaigc.github.io/wan-animate-2

>StyleComposer: Training-Free Multi-Reference Style Composition
https://lexxsh.github.io/StyleComposer

>Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
https://arxiv.org/abs/2608.06125

>Adapting Vision Foundation Models with Cascaded Semantics
https://xixiaouab.github.io/Cascaded-Semantics

>Learning visual representations for compositional analysis of artworks and photographs
https://arxiv.org/abs/2608.06142

>MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
https://arxiv.org/abs/2608.05450

>Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
https://arxiv.org/abs/2605.16411
>>
>>109494241
trying to get the colors consistent across gens.
>>
bocchi season 2:

https://files.catbox.moe/844ibw.mp4
>>
>>109494235
Yes, you can absolutely do that. In the context of the Full-Reference Mode described in your guide, this falls under the task type video continuation.
>>
>>109494247
https://www.blackmagicdesign.com/ca/products/davinciresolve
>>
anyone tried multiple video continuations to see if it degrades?
>>
File: debo_tt_k2_00023_.png (2.23 MB, 1872x1007)
2.23 MB PNG
>>109494046
>the results aren't too great
I disagree. looks great
>>
>>109494225
It might have worse prompt adherence than base+spectrum, or I might just be hitting seed variance. My first gen with the lora missed a lot of action and detail that the one without it caught. More testing tk
>>
>>109494251
the sound is really good, can't believe we have a local model that btfo all API video models except the best
>>
>>109494241
>>109494247
This is great! Keep going anon until we have the first anime made by 4chan (not counting very old attempt by /a/).
>>
>>109494257
no, no, AI must do everything now. Hollywood is dead.
>>
>>109492642
Prompted the Flux 3 API with just "anime opening" for comparison

It's really good and the sound is lossless
https://files.catbox.moe/rf653a.mp4

Let's hope that Dev is this good
>>
>>109494235
See >>109492990
>>
>>109494272
does it count as "made by 4chan" if he's feeding in a manga page by page?
>>
>>109494283
anon you realize how close you are to committing several crimes right?
>>
>>109494287
BLAME! animation made by 4chan?
pretty sure animations count as a whole thing.
>>
>>109494288
oy vey you can't watch a girl eat food
>>
File: sd152022.png (855 KB, 512x704)
855 KB PNG
if i have a rtx5090 do i even need the turbo lora? i can generate a 8 second video in like a min in a half. I can do less than a min with the turbo but the turbo seems to make it look like slop (turbo lora deep fry) any other 5090 bros wanna share their h3 workflow?

pic not related
>>
>>109494300
>ah ah
>I'm merely *pretending* to be a pedophile!
>>
Anyone have luck prompting different voice styles or timbres? I'd like to have better control over what exact kind of voice my 1girls have.
>>
Have they said what the hardware reqs for Flux 3 will be? Will it be heavier to run than H3?
>>
>>109494305
you can give it a voice sample as one of the inputs
>>
anyone here use sd.cpp? what's a good ui for it?
>>
>>109494283
>also i'm trying out different ages in prompts.
which age is the most difficult to get?
>>
>>109494301
I'm on a RTX PRO 6000 and I use the turbo lora for memes/shitposts, it allows you to push a higher res for less time. Also kinda good for quick tests of prompts before doing a "real" run although it's not going to turn out the same.
>>
>>109494313
248 year olds
>>
>>109494300
People get bagged off of 4chan all the time
Do what you want, but don't say you weren't warned when you're facing 60+ years
>>
>>109494301
Why isn't this moving?
>>
>>109494279
>Let's hope that Dev is this good
That would be nice but zero chances of that, they'll probably retrain the model for reinforced safety, which will destroy useful concepts too.
>>
https://files.catbox.moe/9t1hdu.mp4
>>
>>109494323
only if you dare to challenge emperor chitwood
>>
>>109494300
anon im a reasonable man, you know damn well admitting to genning anything close to realistic and going younger is playing with more than just fire, especially since there are in fact federal laws about this now. nobody is going to stop you, but theres a tremendous difference between the shitheads on /b/ spamming cunny slop, and someone posting any realistic equivalent to it even as a test or tease
>>
>>109494336
lmao, I like that.
>>
File: 1776164806541237.png (1.06 MB, 1280x739)
1.06 MB PNG
>>109494193
meanwhile in the bizzaro world...
https://files.catbox.moe/zqp9ln.mp4
>>>/wsg/6210102
>>
>>109494336
holy shit
>>
>>109494336
you're literally 400 years old
>>
>>109494346
>>>>/wsg/6210102
desu shouldve just done an actual scene from that film but with miku as bateman
>>
>>109494336
i was not expecting that at all
>>
File: MiniMax_H3_00535.mp4 (2.56 MB, 832x640)
2.56 MB
2.56 MB MP4
>>109494212
Just the kijai workflow with the turbo lora and spectrum node. I use gpt sol to write the prompts
>>109494220
The audio is bad on that gen because i haven't updated comfy with the fix yet
>>
File: 1762767853294304.mp4 (3.69 MB, 640x800)
3.69 MB
3.69 MB MP4
>>
I'm calling for a complete and total shutdown of turbo loras until we can figure out what the hell is going on.
>>
>>109494363
>yoga wand
is this a real thing?
>>
How do I use a video reference? Comfy doesn't let me connect the "Load video" node to the "ref_video_0" point.
>>
>>109494369
local gods won
>>
>>109494369
cool stuff
>>
>>109494301
>i can generate a 8 second video in like a min in a half
at what megapixel?
>>
File: 78247110607884.mp4 (2.66 MB, 576x736)
2.66 MB
2.66 MB MP4
>>109494072
thx
>>
>>109494390
this activates the monkey brain.
>>
im gonna go on a limb and ask for someone to be brave, and try to recreate skibidi biden and see if it turns out less shit than the real deal
>>
>>109494379
the load video node needs to output as image
>>
>>109494346
it has slopified the original image, look at the first and second frame, that's a shame
>>
File: 1774405008403077.mp4 (597 KB, 640x800)
597 KB
597 KB MP4
>>
https://files.catbox.moe/j4q6nj.mp4
>>
>>109494413
my dads girlfriend used to do his with my dad, I said it was gross and she laughed
>>
File: 900262738109811.mp4 (3.91 MB, 544x768)
3.91 MB
3.91 MB MP4
>>
>>109494379
>>109494404
Ok, I found the answer: the "Load Video" connect to "Get Video Components" and from there "images" go to "ref_video_0" and "audio" to "ref_video_audio_0". Now genning to see if it works.
>>
the solution to scalping:

https://files.catbox.moe/6gstw2.mp4
>>
>>109494414
no ai model in the world can capture how freakish she actually looks
>>
>>109494373
But im already solved whats going on. We just need more datas
>>
>>109494429
>he kills them all with just one bullet
bro thinks he's Chuck Norris :skull: fr fr
>>
There sure are a lot of weird fetishes going on here. This general is a walking DSM
>>
>>109494429
the solution to scalpers v2, on camera edition:

https://files.catbox.moe/4jf94i.mp4
>>
>>109494459
Fresh off the boat?
>>
>>109494429
>>109494460
why even use that faggot arnold for this? why not someone cool like stallone?
>>
>>109494460
can't you prompt for him with an Uzi 9mm and machine gun them all? Nice to see Arnie's terrible acting is replicated tho
>>
>>109494471
>stallone
>cool
>>
>>109494471
Arnold was a bodybuilder champion, what has Stallone achieved?
>>
>>109494471
tbf it is the kind of faggy thing Arnie would do, he's a morality fag in his older years
>>
File: 1731356720111808.webm (3.49 MB, 480x854)
3.49 MB
3.49 MB WEBM
I experimented with generating video last year, but each one took a super long time, and the results were ok at best.
Minimax has got me interested in trying again, but not if it's going to take ages to gen a single video.
If I have an RTX 3080 (12GB of VRAM) is it going to take a super long time to generate a 10-second 24fps video clip at a resolution like 512x512?
>>
You ARE using a LLM with the Official MiniMax skill markdown files to prompt... right, anon?
>>
>>109494479
>he uses the 3 seas shells
ngmi
>>
>>109494488
OHHH MAMACITA
>>
File: file.png (185 KB, 1244x615)
185 KB PNG
>>109494387
ok i lied its basically 2min after the vae decoding
>at what megapixel?
whatever 864x480 is, so probably less than half of one. im using pixaroma nodes

https://files.catbox.moe/v9c5dc.mp4
>>
>>109494481
>what has Stallone achieved?
not becoming this in his old age >>109494486

stallone aint ever told me fuck my freedoms, or called me a nazi.
>>
>>109494481
I don't think pumping steroids and also being a pedo is cool.
>>
>>109494508
>pumping steroids
his body his choice
>>
>>109494505
>that catbox
dan, its time to stop
>>
File: 1761195566996668.mp4 (571 KB, 640x832)
571 KB
571 KB MP4
>>
>>109494471
well you are in luck, it knows rocky.

https://files.catbox.moe/7qi1q0.mp4
>>
>>109494507
>stallone aint ever told me fuck my freedoms, or called me a nazi.
what? but Arnold is a conservative
>>
>>109494517
>>109494460
so you're doing some I2V + references or does the model knows who those actors are?
>>
>>109494519
like Hillary is conservative
>>
is there anyway to stop comfyui from having it's ui get gradually slower and slower? It starts taking 7GB of ram doing nothing at all in Firefox, and that's with purge nodes and shit. Is it a browser issue or just Comfyui being Comfyui?
>>
>>109494526
that was i2v, it knows.
>>
>>109494517
damn it even looked like he enjoyed blasting those fatties
>>
>>109494519
>supporting trannies is conservative
>>
File: ComfyUI_temp_qvmpq_00001_.png (1.67 MB, 1664x1248)
1.67 MB PNG
ref model as edit model... hmmm?
>>
>gave shitty llm i had here for erp the instructions for writing h3 prompts
>it actually did a good job minus a few rerolls i needed to do
f-fuck i actually really do need that 64 gigs of system ram, especially for a better LLM. Isn't gemma supposed to be good? people were saying the Q4 of that is usable. But i suppose this is a convo for /lmg/.
either way, i think the quality of the prompt you give h3 matters even more than resolution and other settings.
>>
>>109494526
model knows a lot of people whether it's actors, politicians, TV hosts presenters. I case of just realising what it won't and being surprised when it knows them and being disappointed it doesn't. Not that there won't be celeb and character Loras made anyway
>>
>>109494534
I didn't know he worshipped troons, is that real?
>>
>>109494548
what makes this model special is you can use the reference one to just add a picture and a voice clip and you can gen what the model doesnt know natively.
>>
>>109494539
is that his dih
>>
>>109494539
would be cool to try a disassembled body parts and see if it can fuse them all together in various formats, real life, cartoon etc. Could make some good stuff especially horror
>>
>>109494556
its his dingaloid
>>
File: ComfyUI_temp_qvmpq_00002_.png (2.26 MB, 1664x1248)
2.26 MB PNG
>>109494539
hmmm ok ok ok
>>109494556
I hope not...
>>
>>109494550
you live under a rock
>>
>>109494561
klein can't do this btw, it would rape the style.
>>
THATS HIS PEE PEE
>>
>>109494553
yeh I have to retry the ref model, didn't really know what I was doing and went back to the main one. Certainly is for the more creative people out there as after all you can simply just so t2v with a small prompt and "create" something.
>>
>>109494539
>>109494561
how 2 single image output from h3?
>>
>>109494562
He likely just fucks them. That's what being on test does to a mf. Mbappe is also famous and publicly dated a tranny despite success and fame.
>>
>>109494570
>set length to 5 (minimum)
>generate
>>
File: 1764392564528655.jpg (104 KB, 572x621)
104 KB JPG
you guys lied. i keep running out of memory when i try to use h3 with 24gb vram and 32gb ram.
>>
>>109494517
boxing day for scalpers:

https://files.catbox.moe/jo2oj7.mp4
>>
>>109494580
there's a bug, for the moment use those flags --vram-headroom 1 --disable-pinned-memory
https://github.com/Comfy-Org/ComfyUI/issues/15255#issue-5050874759
>>
>>109494336
art

>>109494488
fairly long but it will do it
>>
>>109494584
Too sloppy anon, try increasing res and using more cinematic terms in the prompt
>>
>>109494580
do you have nvme2? Set up a swap memory there
I set a 64gb swap and never got any OOMs so far
>>
Feels good running comfy on a separate LINUX box so I can use ALL my VRAM for gen.
>>
>>109494584
>>109494596
probably also have the individuals named (like subject 1 etc) so he can punch each one or just a few as it looks like he hits one and they all fall over for no reason
>>
>Bane from the movie Batman the dark knight rises, walks in from the left and says "those cards are too big, for you."
I guess he was too big for the i2v model, and will need the reference treatment

https://files.catbox.moe/ozy9j0.mp4
>>
https://files.catbox.moe/9gpu75.mp4
>>>/wsg/6210120
>>
Why don't you guys talk about your discord requests like /vp/?

>>>/vp/59489955
>Been getting other suggestions from people on Discord. For example, swapping Acerola's outfit with Neptune's (from the game of the same name). I'm surprised how well this worked out. I'm gonna say Neptune did her hair up more like Acerola's for her shots.
>>
>>109494620
nice, you used the music as a reference right?
>>
>>109494621
what is a "discord request"?
>>
File: MiniMax_H3_00543.webm (1.62 MB, 832x640)
1.62 MB
1.62 MB WEBM
>>
>>109494621
>being amazed by "1girl, character a, character b (cosplay)"
>>
kek
https://files.catbox.moe/i8qxqn.mp4
>>
>>109494513
I love bouncy boobs/ass/thighs but this is ridiculous lol
>>
File: 1768631487070178.png (419 KB, 736x416)
419 KB PNG
>the man in picture 1 is on a 747 jet flying in the sky
I was not specific enough!
>>
>>109494645
it's insane how good it is at anime
>>
>>109494570
>>109494578
https://github.com/tori29umai0123/ComfyUI-MiniMaxH3-SingleFrame
>>
>>109494662
yeah and it does real animation, not some 3d rotoscoped 60 fps slop
>>
>>109494653
It was great until the music cut out.
>>
>>109494334
I hope not, it may be local's only salvation for anything involving high speed movement.
>>
>>109494460
Now get to the chopper.
50 meters spread no sound.
>>
Has anyone done a comparison between the int8_convrot and nvfp4 text encoders to see if there's a significant drop in quality from going with the latter?
I'm using a 3090 so being able to fit it on vram would be nice if the quality tradeoff isn't too bad
>>
>>109494413
if you are trying to impress us with violates of the social code you're going to be disappointed.

Why even?
>>
File: Index x Touma.webm (1.61 MB, 672x960)
1.61 MB
1.61 MB WEBM
>>
>>109494658
Bane?
>>
File: MiniMax_H3_00562_.webm (858 KB, 768x960)
858 KB
858 KB WEBM
>>
>>109494689
Nice. Now have him do that to Kuroko
>>
>>109494686
Haven't tested it since int8 doesn't fit on my PC as 64GB of RAMlet, but generally the rule of thumb is, the larger the model, the more accurate fp4 gets, so at huge model sizes like this, the difference is negligible compared to smaller models.
>>
File: MiniMax_H3_00576_.webm (1.67 MB, 672x960)
1.67 MB
1.67 MB WEBM
>>
>>109494702
>not eating them
fucking coward
>>
File: debo_tt_k2_00026_.png (2.19 MB, 1872x1007)
2.19 MB PNG
>>109494620
fire
>>
>>109494620
excellent
but there's no way it generated the music
>>
>>109494689
>no one complaining about pedos for this one
curious.
>>
File: ComfyUI_temp_cbapf_00013_.png (2.18 MB, 1920x1088)
2.18 MB PNG
oy vey
>>
>>109494653
a bit too random but I liked the dancing comfy doodle
>>
>>109494711
>but there's no way it generated the music
it didn't. the stereo audio feature is a meme that doesn't work
>>
>>109494689
Need some Touma x Kuroko.
>>
open source vs closed source

https://files.catbox.moe/yar5a2.mp4
>>
>>109494689
nice, I forgot the nun was this cute
>>
>>109494395
why even fucking post it? are you fucking retarded? all points made by those anons are valid but have fun pretending you aren't being watched I guess. Do you even know where you are mate? This is fed controlled shit you are posting on.
>>
File: erin.jpg (85 KB, 320x320)
85 KB JPG
>>109494645
>I'll kill them all
>>
File: 1658896026863193.png (628 KB, 976x850)
628 KB PNG
>>109494689
>went straight for the tummy first
>>
File: ComfyUI_temp_cbapf_00015_.png (1.78 MB, 1920x1088)
1.78 MB PNG
ok.
>>
>>109494739
he needs to get warmed up
>>
File: 1760351403154577.png (548 KB, 978x1019)
548 KB PNG
Was WAN in the first day of Turbo lora are this ? Unstable and such. I dont follow WAN, only LTX
>>
>>109494726
>>109494698
while biribiri watches...
>>
>>109494731
>mate
found the issue
my brain genuinely thanks random chance once a month im not in a country with english culture that has made me a slave cuck to my government hot damn
>>
>>109494755
based noticer
>>
>>109494749
I don't get your question.
If you mean people excited, sure, as much as now for wan21.
If you mean new discoveries and optimizations, it was slower back then, it was one of the first or the first model that didn't completely suck.
>>
>>109494745
is that Pam?
>>
What LLM do you use to make H3 Prompts? I tried Gemma and Qwen but they keep getting things wrong
>>
File: that's the joke.png (330 KB, 728x412)
330 KB PNG
>>109494745
>those 4 women really look alike...
>>
>>109494763
Its still sucks today. Even with Turbo its still much slower than current Minimax
>>
>>109494764
yes.
>>
>>109494776
which gemma
>>
>>109494755
You can't tell me that video was genned to be a genuine tasteful video of a kid. There's the difference, i hope you get fucking v& because scruffs like you give AI such bad press.
>>
>>109494776
I dont use it often but :

<subject>
<picture 1>
integrated_multimodal_description:
overall_soundscape :
non_diegetic_music:

Works for me
>>
>>109494776
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf
but you need a vision model like:
mmproj-Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-f16.gguf
to understand what's in the image while generating the prompt
>>
File: test.jpg (265 KB, 2368x1344)
265 KB JPG
>>109494745
what are the odds some other anon is testing out making reference sheets using ref2video at the same time as me

unless you're not /minimax/ing?
>>
>>109494794
e4 hau
>>109494798
I have a larger model that works good but it is so incredible slow.
>>109494804
I have tried with 3.5, is 3.6 a major upgrade?
>>
>>109494776
Grok for NSFW shit, Gemini for everything else
>>
>>109494810
fml saved the wrong frame lol
the quality is better than this, and not yellow
>>
>>109494810
Yes we are doing the same thing
>>109494815
>e4 hau
my guy you need to use 31B, e4 is WAY too small for the amount of info it needs to remember about the prompt structure.
>>
>>109494818
Could you share the edit prompt plox?
>>
File: google.png (207 KB, 1770x1462)
207 KB PNG
>>109494815
it's a step up, there's no reason to use 3.5
>>
Guys, I think I'm finally suffering burnout from Minimax H3.

I'm going to go to sleep, and hopefully when I wake up, I won't have the itch to gen again. As great as the model is, it's draining my life (and my semen).
>>
>>109494825
>unironically pasting the google AI overview
>>
>>109494826
I have no idea how I will be bored of this model, it can literally do anything, if your imagination is strong enough you won't have enough of your lifetime to make all the videos you had in mind
>>
I have a 3060. I will remake GoT S08
Cya in half a decade
>>
>>109494831
so? he could've just googled it himself, he's the retard for asking questions instead of googling them. I'm the useful one here acting as google.
>>
>>109494835
I wouldn't say I'm bored, but feeling burned out rn. 768 gens and counting. And I know I haven't even scratched the surface of it's true capability with R2V.

The thing is, when I go to sleep and wake up, I'll probably feel like genning again...
>>
File: 1772228754269970.jpg (89 KB, 479x793)
89 KB JPG
I miss LTX first pass
>>
>>109494823
picture 1 denotes a picture of the outfit/body whatever, picture 2 denotes a face image.

An orthographic reference of the woman in <Picture 1>, with full accuracy. She's on a white background. There's a front view, side view, and rear view. On the left, there are two imposed Face Close-up's. One has her smiling, the other has her neutral. She stands perfectly still in a relaxed standing pose.

(Optional: Her prominent cameltoe and big butt are perfectly reproduced.)

Beautiful Face, similar to <Picture 2>
High Quality, Professional Reference Video, in 4K. In focus, High Detail, No Special Effects. Still Frame.
>>
https://files.catbox.moe/h7hqh0.mp4
>>
>>109494839
>he's the retard for asking questions instead of googling them
Be kind Anon.
>>
>>109494848
shit, the first sentence isn't part of the prompt, the prompt starts at, "An orthographic"
my bad, just wanted to clarify a bit
>>
fancy dinner

https://files.catbox.moe/ooapdn.mp4
>>
>>109494849
not a terrible groove
>>
Are the newfags still here or did they finally fuck off
>>
What do you guys think? Does her dancing sync up with the audio reference?

https://n.uguu.se/hcQwmCUs.webm
https://h.uguu.se/dnqwjVdK.webm

(ignore the frame rate drop and audio breaking at the end. not sure what's causing that)
>>
>>109494871
why would you want them to leave?
>>
>>109494872
why do you keep making the same thing over and over. are you insane?
>>
>>109494872
could be tighter but it works kinda?
>>
>>109494881
He doesn't get enough (yous) on e and vp so he has to spam here too
>>
>>109494872
>Does her dancing sync up with the audio reference?
no it looks unrelated to the song
>>
>>109494881
ngmi
>>
>>109494881
>>109494886
I understand that you're a terminally online autist, but I'm just asking for feedback (i didn't get it last time I asked).
>>
>>109494891
bro I watched two movies and ate dinner. you've been making the same thing with different songs since then.
>>
you know what this thread needs? more will smith kinos
>>
>>109494895
>you've been making the same thing with different songs since then
No I haven't you laughable little retard.
Typical schizo conflating a bunch of people together. Being terminally online is very bad for your mental health anon.
>>
honeymoon period over already?
>>
>>109494902
stop projecting, retard. I watched TWO movies.
Mimic and Evil Dead Burn. You're still making the fucking pokemon pvc figure dancing shit since before I watched those movies.
>>
File: file.png (66 KB, 758x383)
66 KB PNG
>>109494902
>Typical schizo conflating a bunch of people together.
Couldn't agree more desu.
>>
>>109494901
the thread is almost over, are you retarded or something?
>>
>>109494909
idk, pretty sure H3 is the new thread meta. Why gen images anymore when videos are so accessible?
>>
>>109494902
>Being terminally online is very bad for your mental health anon.
proof?
>>
>>109494917
ok, EVERYONE GO TO THE FRESH THREAD
>>109480220
>>109480220
>>109480220
>>109480220
>>
>>109494911
>>109494913
Anon why are you using a VPN to double post like that?
>>
>>109494928
stinks like shit, why don't YOU go touch grass
>>
>>109494911
Doesn't matter what you say, you're on an anonymous imageboard, moron. Show me all these other similar gens I've made with different songs, and then show your data to prove it's "me".
You're a dime a dozen schizophrenic, terminally online little cretin.
>>
>>109494919
Yep it's a whole new world now. We're going to see videogen everywhere just like we saw SD sloppa posted everywhere.
>>
>>109494331
>Why isn't this moving?
my bad
https://files.catbox.moe/auc6y2.mp4
>>
>>109494927
nevermind
>>109494928
>>
He shut up but i'll warn him anyway.
>>109494755
My cucked country you say?

Probable cause
In United States criminal law, probable cause is the legal standard by which police authorities have reason to obtain a warrant for the arrest of a suspected criminal and for a court's issuing of a search warrant.

fafo

The girl in the video appeared apparently topless, does not matter if her chest was not exposed, the content was gross focusing on her mouth probably for sicko like you to jack off to. Then they wonder what else you might be generating.
Probable cause...
>>
>>109494935
put on a trip so I can filter your retard ass.
>>
>>109494849
good song. video appears unrelated
>>
>>109494919
because you need references and inputs for the first frame.
>>
>>109494950
>he doesn't write a good t2v prompt
ngmi
>>
>>109494945
I'm just going to keep correctly labeling you what you are: an autistic, low-IQ little cretin. You need to go outside a bit more, my terminally online friend.
>>
>two autists dukin it out
>>
>>109494954
deadass not interested in t2v.
>>
>>109494958
>terminally online
You calling someone else that when you're on minimum 3 boards spamming the same thing and gets upset when people doesn't reply to you lol
>>
>>109494971
>>109494971
>>
>>109494967
>You calling someone else that when you're on minimum 3 boards spamming the same thing and gets upset when people doesn't reply to you lol
ironic since you start throwing a temper tantrum when a video of someone eating spaghetti gets posted
>>
>>109494872
second one does
people who don't dance think people dance to the beat, but they usually dance to the phrasing. her movements catch the phrase shifts well
zoomer tiktok kpop dancing is a different thing though so if thats the intention, you didn't get there. zoomer tiktok kpop dancing is very literal and autistic. those retards will practice a dance for a dozen hours for a 10s clip
>>
Bros I'm using the default workflow for minimax. I got a 4080 and 32gb of ram and this shit just doesn't seem worth it. for a 864 x 480 5 second vid takes like 30 minutes. I just wanna goon.
>>
>>109495050
why is your 4080 slower than my 3060
>>
>>109495001
Image to video will always be superior.
>>109495050
I'm using a 4090, for that resolution it takes me like 2-3 minutes. I'm using a jewtuber's workflow.
>>
File: file.png (110 KB, 1421x464)
110 KB PNG
>>109495067
make sure you are using using cuda130 (have an up to date comfy ui), install sageattention and triton, and use these nodes for optimization. using comfyui through stabilitymatrix makes the first two things easier
>>
>>109495083
>>109495050
was supposed to be a reply to you. also those nodes connect to your basic guider (model)
>>
>>109494668
Lot's interesting. But it is really needed a separated node for that?

>>109494686
I am using a 3090 and tested both. I am unsure if there is difference in quality, but the gen time was almost the same, so I am using int8 now.
>>
File: MiniMax_H3_00117.mp4 (2.44 MB, 960x480)
2.44 MB
2.44 MB MP4
H3 can generated videos in LR SBS format. It's not completely perfect but it is stereo 3D. I'm curious if the minor issues are resolution related.
>>
Alright, it looks like the pedophile just got purged and the new thread got swept in the crossfire.

Someone bake a new. Not using debo's thread.
>>
y do
>>
>anon went full retard and got banned
>new thread got deleted as well
jesus christ i know this general sucks ass but this is pretty ridiculous
>>
>>109495243
>>109495247
neither of you are baking so why don't you shut the fuck up
>>
>>109495243
>>109495247
It wasn't the baker >>109494972 post is still up
>>
New non-debo thread:

>>109495264
>>109495264
>>109495264
>>
is there a way to use the optimizer nodes and rtx super resolution on minimax i2v?
>>
>>109493955
bro ... just get the rtx vsr node
takes 10 secs tops
git clone https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI
...
python_embeded\python.exe -m pip install wheel-stub
python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\Nvidia_RTX_Nodes_ComfyUI\requirements.txt
>>
>>109494046
it looks great as others said. low res is the issue. need them megapixels



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.