[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1759523346459527.mp4 (2.86 MB, 832x640)
2.86 MB
2.86 MB MP4
Previous: >>109491842

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109492953
thanks for the bake
>>
>mfw Resource news

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo – OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view
>>
>>109492953
I need the wf for this, directing uncanny anime scenes this is incredible
>>
>mfw Research news

08/07/2026

>Vorch-Omni: Multi-Task Orchestration of Sight and Sound
https://vorch-project.github.io/Vorch-Omni-project

>Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
https://vorch-project.github.io/Vorch-Streamer-project

>Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
https://vorch-project.github.io/Vorch-Director-project

>Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
https://vorch-project.github.io/Vorch-IR-project

>In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
https://arxiv.org/abs/2608.05237

>Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
https://arxiv.org/abs/2608.05976

>EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
https://arxiv.org/abs/2608.06231

>MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
https://arxiv.org/abs/2608.05878

>Wan-Animate-2: Pushing the Application Boundaries of Character Animation
https://humanaigc.github.io/wan-animate-2

>StyleComposer: Training-Free Multi-Reference Style Composition
https://lexxsh.github.io/StyleComposer

>Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
https://arxiv.org/abs/2608.06125

>Adapting Vision Foundation Models with Cascaded Semantics
https://xixiaouab.github.io/Cascaded-Semantics

>Learning visual representations for compositional analysis of artworks and photographs
https://arxiv.org/abs/2608.06142

>MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
https://arxiv.org/abs/2608.05450

>Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
https://arxiv.org/abs/2605.16411
>>
REAL THREAD IS HERE
>>109480220
>>109480220
>>109480220
>>
File: 1755303398777984.webm (794 KB, 720x1296)
794 KB
794 KB WEBM
Hopefully janny fine with this
>>
>>109492965
there's no magic wf anon, you take a real anime screenshot, and you describe what they will be doing, and that's all there is to it, minimax is that good >>>/wsg/6209484
>>
>>109492977
oh debo, you sad little cretin
>>
Director's cut is unreal. How'd they get away with it?
https://files.catbox.moe/pyuwgy.mp4
>>
>>109492978
>Makoto with huge boobs
ruined!!
>>
>>109492953
Thanks for bake
>>
>>109492899
>local music gen
music gen*

>>109492886
For the most part you need to train LoRAs to get really good outputs. It's possible to get non-LoRA gens but they will just be slopped (ACEStep without a LoRA will sound too much like synthetic-like slop with instruments, the voice is okay though). You need to train LoRAs yourself, it's not hard, made a rentry here
https://rentry.co/s8fg8ber#captioning

Remember, like with image models, you can train LoRAs on subsets of genres you want instead of just artists and it'll do fine. For instance, one on 80s city pop, another on metal, and so on... After having several, it will beat using API models because you're not going to necessarily prompt API models outside of a few genres you like.

>>109492886
>I'm assuming no one wants to get sued
Yep, though it's a safetensors file, doesn't really have music in it, and one workaround is releasing genre based LoRAs like the super eurobeats one I trained here https://civitai.com/models/2702491/super-eurobeats-acestep-15-xl
>>
>>109492978
failed the conservation of mass test
>>
>>109492990
>How'd they get away with it?
THEY CANT KEEP GETTING AWAY WITH IT
https://files.catbox.moe/4t9m1p.webm
>>
>>109492990
this is insane...
>>
>>109492990
It's pretty much seamless, fuck
>>
>>109492990
Spicy.
>>
HAHAHA reference model is amazing, can literally do anything.

Use Image 1 as the exact character identity reference for JC Denton, keeping his signature black sunglasses and stoic features. Extract the vocal timbre, frequency, and speech pattern from Audio 1 to build a voice clone. JC Denton is in an office answering phones in a call center. JC Denton picks up the phone and the female caller says "I need help with my router." JC Denton says "bitch, I work for UNATCO, i'm not your tech support. Go kill yourself.".
overall_soundscape: Cold, flat, monotone spoken dialogue with completely deadpan delivery. The voice print has an electronic, slightly raspy, low-intonation pitch mirroring the robotic timbre of Audio 1. No vocal inflections or emotional excitement.

https://files.catbox.moe/trc29w.mp4
>>
>>109492977
Imagine having so little going for you in life that the highlight of your day is just sameposting in your shitty thread just so it doesn't get get deleted.
>>
File: 4gh5.png (11 KB, 644x800)
11 KB PNG
>Imagine having so little going for you in life that the highlight of your day is just sameposting in your shitty thread just so it doesn't get get deleted.
>>
Hit a little too close to home?
>>
Everything already perfect....
..... Except Vagina/Anus/Penis is still shit........
>>
>>109493071
what if you use close ups of vaginas and dicks as a reference image lol
>>
File: 34.jpg (34 KB, 600x600)
34 KB JPG
>109493056
its (you)
>>
>>109493076
>What if you try this really specific thing I want and link it here with a workflow lol
>>
>>109493076
That does work and in i2v
>>
>>109493082
exactly, now you get it
>>
>>109493089
i2v is really tricky because unless you start with everything fully visible it'll just morph into abominations.
>>
>>109492958
>>109492968
fuck off wheelchairbo
>>
>>109493071
Use a reference or wait a few weeks for a proper NSFW checkpoint or lora.
You can't really expect companies to drop 100% NSFW capable models on the market at this point, but H3 is remarkably permissive for a base video model. It's amazingly permissive compared to what we've had access to before given how well references work.
>>
File: MiniMax_H3_00661__1.webm (1.95 MB, 896x704)
1.95 MB
1.95 MB WEBM
https://n.uguu.se/OcKAJMdV.mp4
>>
gonna go buy some snacks and I'll do some more
manga to animation conversions.
>>
>ani implements stealth png in the base ui
based
>>
>>109493101
now we're talking
>>
H3 ref2va might be one of the biggest tactical nukes dropped on the local community in a while
>>
>>109493116
>in a while
more like "never", it's the first time we got something this powerful
>>
>>109493112
qrd?
>>
Would you guys say H3 is the biggest leap forward in the local scene since 2020?
>>
>>109493126
no, that's Krea 2
>>
>>109493126
I'd say the leap is as big from LTX to Minimax than SDXL was to Flux.1 dev
>>
>>109493126
feels like she should have got this or something on par with it a while ago but the big names went to api, I could have run this easily like a year ago even before the dynamic vram stuff, so in a roundabout way yes
>>
>>109493124
>>109492934
you'll be able to see the metadata on 4chan without having to download the image using the catboxanon script
>>
>>109493141
>she
damn prompt muscle memory
>>
>>109492990
What gets me is how it continued the performance perfectly. Pitch perfectly. Like 4 second input video and the way the AI delivered the lines feels and sounds 100% in line with how Nicholson might have delivered it.
>>
>>109492978
western tourist fandom in shambles
>>
>>109493147
didnt ask
>>
>>109493112
>>109493124
>>109493147
>>
>>109493071
even with references?
>>
>>109493126
It's big but I don't think I have powerful enough GPU to gen at the quality I'd want to with it without making it too slow.
>>
>>109493147
>forces a metadata without the user's consent
cringe, that's why I create my gens in comfyui, I know there won't be schizo talking to me since the metadata is stripped off by 4chan
>>
>>109493160
this anon did >>109493124
>>
>>109493177
>doesnt know what qrd means
retard
>>
What's the current meta, lads? What am I missing from my model stack?
>>
https://files.catbox.moe/9nw1np.mp4
https://files.catbox.moe/usfkha.mp4
https://files.catbox.moe/ldyqwx.mp4
https://files.catbox.moe/yukt7r.mp4
>>
words of wisdom

https://files.catbox.moe/84d90m.mp4
>>
>>109493186
I think you're good, the turbo loras we have currently suck, but it's just a matter of days before they get good enough to replace spectrum
>>
File: screenshot.1786146677.jpg (157 KB, 1051x916)
157 KB JPG
I stole anon's prompt enhancer idea and told claude to build it based on the screenshot and it works perfectly, LOL
>>
File: tard.mp4 (378 KB, 736x576)
378 KB
378 KB MP4
>>109493180
SO waht?
>>
>>109493195
Did you try lightx2v yet?
>>
>>109493186
You don't patch sage twice. You use one or the other. They both do the same thing differently.
You're missing Patch Sol-Attn if you've got a 40/50 series.
>>
>>109493192
fixed start, I needed more dialogue.

https://files.catbox.moe/0yz28v.mp4
>>
>>109493204
I did, it's mushy and the sound is really bad, not as slopped as the other turbo lora but the quality takes a hit that's too big, not surprising it's just v0.1
>>
>>109493205
>40/50 series
It works on 30 series now too, I recommend turning off INT8 mode for that though as it uses more vram.
>>
>>109493174
what do you mean? comfyorg steals your metadata and owns the rights to it
>>
>>109493186
I don't use sage attention yet, where do I get that model?
>>
File: you know how it goes ani.jpg (477 KB, 2560x1920)
477 KB JPG
>>109493219
>comfyorg steals your metadata and owns the rights to it
>>
File: 1502979408671.jpg (7 KB, 236x190)
7 KB JPG
>>109493205
>You don't patch sage twice. You use one or the other. They both do the same thing differently.
Was under the impression it didn't make too much of a difference? I have use-sage-attention in my startup script, but Kijai's node let's me pick an option with triton and allow_compile in it, and i've been told that's good. (>mfw). Does the Mem Eff Sage Attention really do the same thing?

>You're missing Patch Sol-Attn if you've got a 40/50 series.
Well I do have an RTX 5070 Ti. What can I expect from adding this node? shorter gen time? Will it play nice with sage-attn and spectrum?
>>
>>109493224
startup script
>>
>>109493228
read their tos
>>
>>109493231
>shorter gen time?
yes
>Will it play nice with sage-attn and spectrum?
yes. place it after sage in the chain
>>
>>109493126
It's the biggest leap for video gen yes, maybe second after wan 2.1 which itself was a revolution when it got released.
>>
>>109493244
I did, where did it say that?
>>
>>109493231
Half right. You only use the mem eff patch if you need the vram. Cost is apparently a slight drop in gen speed.
>These two (Patch Sage Attention KJ + MiniMax H3 Mem Eff Sage Attention Patch) aren't really competing with Sol-Attn, but they do somewhat compete with each other. Kijai: one node activates sage attention while the other modifies the model's forward pass to reduce peak VRAM usage. They're meant to be usable together, though the plain KJ Sage patch is a bit faster, and the VRAM-focused one only really helps if you're memory-starved.
>>
>>109493147
so basically he made a spyware and he's proud of that? damn lmao
>>
>>109492953
>>109492990
what the fuck? are local models better than cloud now?
>>
>>109493254
https://www.reddit.com/r/StableDiffusion/comments/1u3br9y/psa_comfyui_wont_train_on_your_images_but_if_you/
>>
>>109493281
It's closer than it's ever been, and local has an edge in that you can do shit like that without being refused/censored.
>>
>>109493281
It's close in technical terms but it's heavily uncensored so it's better.
>>
>>109493281
>now
has been for months now
>>
>>109493281
Always were, chud.
>>
>>109493283
>if you use Cloud
kek
>>
>>109493205
Kijai Turbo Lora workflow use both of those.
>>
>>109493292
>>109493281
>has been for months now
stop the cap, Seedance 2.5 is still the king of videogen
>>
>>109493281
no paid service can do what reference is capable of cause it would be blocked.
>>
>>109493294
yeah the BPD schizo is always grasping at straws
kinda sad
>>
>>109493298
see >>109493274
If you've got 24GB+, you don't need it, and it's actually hurting your s/it
>>
Surprising still no one make local videos of Gugu Gaga and Doro yet
>>
h3 doesn't have as much sovl when it comes to older style music videos compared to ltx
>>
>>109493306
Well i have 16gb of VRAM (5070ti) like that frog anon
>>
>be me
>5070 owner
>update from cuda 12 to 13
>gen speeds stay basically the same if not slightly worse
wtf anon why did you lie
>>
>>109493307
there is one of doro that I found
>>>/wsg/6209178
>>
Tried to turn a manga panel into an video but got this cursed shit instead https://h.uguu.se/zRJCkMoX.mp4
>>
File: 1765626768682779.png (5 KB, 206x183)
5 KB PNG
Why is h3 filling up my RAM when there's available VRAM?
>>
>>109493186
you can drop the patch sage attention node but keep the minimax h3 mem eff sage attention, you will gain some speed doing this.
>>
>>109493294
who is to say the workflows that ended up getting sent over aren't marked as cloud?
>>
stop disney

https://files.catbox.moe/qvvsp9.mp4
>>
>>109493319
>tfw custom nodes
>>
File: 526.jpg (72 KB, 785x847)
72 KB JPG
>>109493274
>KJ Sage patch is a bit faster
aren't they... both sage patches? not sure which node this is referring to. i mean one is called "sage attention patch" and the other "patch sage attention"
>and the VRAM-focused one only really helps if you're memory-starved
Well, i mean, i'm on an RTX 5070Ti, so I don't have anywhere near enough vram to fit the model, the text encoder and the fat vae on my vram, and offload a lot to system ram. So which node should I be using?
>>
>>109493325
if you have any proof of that I'm all ears, you're a C++ dev you know how to read code, go dig it
>>
>>109493274
oh really, nta but i disabled to the patch node since that other node does activate sage attention it says so on the fucking nodes title and requires sage attention 2.2.0 and i gained 1 minute reduction in gen time.
>>
>>109493337
I have no idea how c++ works. Who do you think I am?
>>
File: spoon.png (91 KB, 1442x711)
91 KB PNG
>>109493335
>>
>>109492990
>Director's cut
What's this? Where can I get it?
>>
File: SEE YA IN COURT.png (204 KB, 1311x1136)
204 KB PNG
>>109493327
what are you doing anon??
https://files.catbox.moe/u6mlet.mp4
>>
File: 1781662220546949.mp4 (499 KB, 960x544)
499 KB
499 KB MP4
Those who think this won't be used to produce anime are delusional.
>>
how the FUCK do i get overly bouncy titties without them looking like they're spazzing
>>
>>109493340
You have --use-sage-attention enabled globally, dumbo
>>
>>109493357
DON'T use the word jiggle. Usually body interactions are best. If you directly say breasts bounce and jiggle it can make em go nuts.
>>
>>109493311
That's one of the reasons why I am still looking forward to Flux 3 Dev
>>
>>109493351
I'm still a little confused about the second node. If the models are already using 100% of my vram and offloading to system ram is this supposed to help?
I guess I'll test gen times myself.
>>
>>109493363
have any in context examples? i know cause & effect is the best way but it still seems pretty rigid in it's view of when tits bounce lol
>>
>>109493356
Has anyone made a coherent short episode yet?
>>
>>109493356
this is a nuclear weapon in terms of arts
>>
>>109493368
>I am still looking forward to Flux 3 Dev
you know it's gonna be slopped, censored and shit anon, stop coping :(
>>
so what's the deal with sol attn? Are the defaults fine or should I be changing parameters?
>>
hlky is really throwing tantrums lately
>>
>>109493382
I dunno how Dev is going to be, but the API one is the least slopped model to date
>>
>>109493376
Anyways, if Flux 3 delivers, I think that will be used more (except for hentai creation, at least out of the box)
>>
Crazy how much of a hub for AI /ldg/ is
>>
>>109493376
sorry, anon only has the motivation for "what if this character said nigger" and "what if miku"
maybe reddit will show more ambition
>>
>>109493392
>he API one is the least slopped model to date
sora 2 at its very begining was less slopped, but yeah I'm not gonna pretend flux 3 max doesn't have sovl, it has a lot, that's why it's gonna be sad to see dev being completly lobotomized
>>
File: 1757321705536267.jpg (440 KB, 2551x2480)
440 KB JPG
Well i turned off my Sage attention node but keep the Minimax mem eff Sage attention and genning time increased from 210 second to 230 seconds.

Is the quality of gen increased at least ??
>>
File: 1785930147803604.png (258 KB, 2240x763)
258 KB PNG
infinite node works
>>
>>109493351
This is wrong. I had Fable audit the nodes. It said that the MiniMax H3 Mem Eff Sage Attention Patch is self contained. Seem to be right in that there's a performance cost of some kind, need to do tests. Also, if you apply both, whichever one is last in the chain is the one that runs.
>>
>>109493399
Those are good tests of the tech though. We have to make sure it can hold up under heavy loads.
>>
>>109493410
how much is the quailty hit though, at some point all those stacking takes a tool
>>
File: 473444.webm (3.9 MB, 832x1184)
3.9 MB
3.9 MB WEBM
>>109493368
same prompt and resolution for both models
>>
>>109493410
do NOT do this, it creates mustard gas
>>
I have to go up to 0.6 . I dohn't care if it takes 2 extra minutes per gen. I can't handle these faces I wish I had a 5090 fuck
>>
The key to prompting is to describe literally what you want to see in the frame, not what is happening in the imaginary world inside the frame.
>>
>>109493423
shockingly, it has become better and better, blockcache no quality hit cause it rarely skips frames, and it's mostly math fixing stuff
>>
extending Flux 3 Dev gens with H3 is gonna be so kino
>>
File: Untitled.mp4 (2.94 MB, 672x800)
2.94 MB
2.94 MB MP4
Holy fuck H3 is leagues beyond LTX2, it follows instructions way better and has vastly better motion. At last we can leave behind slow motion molasses crap.
>>
>>109493431
>blockcache no quality hit cause it rarely skips frames,
so it rarely gives you any speed increase?
>>
>>109493425
>same prompt
which is...?
>>
>>109493399
anon's vision of the tech :
- what if nigger
- what if floyd
- what if nigger x floyd
>>
>>109493435
>At last we can leave behind slow motion molasses crap.
Oh ho ho anon. There's a whole new class of slowmo fags now. Unfortunately.
>>
>>109493405
Sage attention (2++, NOT 3) is one of the most painless optimization to apply.
>>
>>109493351
>less vram at the cost of speed
wtf why did no one tell me? I just saved 20sec. which likely translates to a minute at higher res.
>>
>>109493447
>slowmo fags
*shlomo fags
>>
so here is the game industry plan.

https://files.catbox.moe/ypdh6a.mp4
>>
>>109493399
Kekd
>>
>>109493441
continuation of a music video for a 1980s pop song. rolling bass, synth brass chorus and fast gated snare drums. smooth frame rate. high resolution. high focal length. bloom. film grain. VHS distortion. realistic balanced colors. low contrast. beautiful summer. the sky and clouds have a faint blue-pink tint to it. slightly dim.
the singer is Katie. she is a young adult. she has fluffy long orange brown hair with bangs. her ears and cheeks are covered by hair. she has dull blue eyes, short cat eyelashes, and freckles.
other young adults are near by. they have black, brown, or blonde hair. long, medium, or short hair. straight, curly, or wavy styles. most of the girls are european, some have darker skin. they have rough skin with minor imperfections.
everyone has the same black sailor uniform. tall waist short pleated skirt, white shirt, long socks, bow, and beret. everyone is dancing with fast footwork and arms swaying. the camera is dynamic and unstable as it is moving around the scene and matching the beat of the song. a few angled close ups of Katie and the other dancers. the song is not over
>>
>>109493376
Working on it.
>>
>>109493443
I think maybe 1 or 2 anons (who are Americans) are obsessed with blacks and prompt for that though
>>
>>109493436
when it does work it's 20%, but if it cant optimize any more it's 0%. spectrum and sageattn/sol are doing the heavy lifting more or less
>>
where the FUCK

are the 80s fantasy kinos?
>>
>>109493405
some useless faggot was wrong then I guess. I'll have to test with the patch node on again and see if it improves my gen time. But I honestly don't know at this point now, there isn't much to go off other than what others say.
>>
>>109493425
Which one is H3 and which one is Flux 3?
>>
Big bang theory if it wasn't mid
https://files.catbox.moe/77wykj.mp4
>>>/wsg/6210058
>>
>>109493472
I think I have some 80s dark fantasy images around, let me see if the reference model can do something interesting
>>
>>109493410
infinite what? will it do gens longer than 15 seconds with out OOM?
>>
>>109493477
h3 is on top and ltx is on bottom
>>
File: really.png (28 KB, 134x144)
28 KB PNG
>>109493357
>when they randomly start bouncing after the action that caused them to bounce has stopped

>>109493435
could I perhaps acquire the workflow?
>>
>>109493472
>are the 80s fantasy kinos?
Ikr, those are peak AI sovl
https://www.tiktok.com/@knightcore___/video/7659820863612538143
>>
flux 3 = simple prompts work better than the same simple prompts in h3
h3 = precise control via the ref model, tagging system and prompting language will prove unbeatable for anything beyond 'make pepe hard with 1girl tits bouncing'
>>
It seems very difficult to get "every frame is a new drawing / painting" style of animation. It really wants to do regular animation instead of sovlful trad stuff... :-(
>>
>>109493452
What about that vs kj patch sage?
>>
I SWEAR THIS IS THE LAST DAY ILL DO THIS
TOMORROW I WILL GO TO WORK
>>
>>109493492
so h3 is like ideogram, and we all know how that went since no one here uses it anymore
>>
attention: dont use sol with spectrum as it slowed my stuff down (bypassed was faster)

idk why but yeah.
>>
>>109493472
>>109493485
>>109493491
Someone also needs to test Harry Potter but 80s high fashion runway.

What the fuck has anon been doing?
>>
>>109493494
They do the same thing.
>>
>>109492328
>>109492311
>What would even be a good place to get the voices from other than skimming the anime and having to deal with background noises.
https://sounds.spriters-resource.com/

>having to deal with background noises.
I use this SAM Audio Lite script to save VRAM; it isolates speech from background music and sound effects:
https://github.com/0x0funky/audioghost-ai
Next, I feed the SAM output file into Resemble Enhance to clean up any remaining background noise:
https://github.com/resemble-ai/resemble-enhance
Then, I use the Acon Digital DeVerberate 3 plugin in Audacity to reduce reverb on the Resemble Enhance output file.
https://rutracker.org/forum/viewtopic.php?t=6118812

Example
Input Audio:
https://vocaroo.com/12owejruutr0
Input audio fed into the SAM Audio Lite script:
https://vocaroo.com/1i5a1MAJ6lbZ
SAM Audio Lite output fed into Resemble Enhance:
https://vocaroo.com/1dIgDeD6sXVF
Resemble Enhance output with the Acon Digital DeVerberate 3 plugin applied:
https://vocaroo.com/1dY1tzvQQhRZ
>>
>>109493492
>flux 3 = simple prompts work better than the same simple prompts in h3
With BFL making it? lol
>>
day 6 without video continuation
>>
>>109493495
>TOMORROW I WILL GO TO WORK
did you actually take days off lol
>>
>>109493488
Oh, ok
Btw, if some kind anon could try Flux 3 on the API to compare with H3, that would be appreciated
>>
>>109493497
h3 is ideogram if ideogram existed back in 2022
>>
>>109493497
they both have their own use cases
>>
>>109493489
unfortunately i think a bouncy titty lora is needed (again)
>>
>>109493510
nta but i took half a week off for the original Flux Release. And another half a week for Anima.
>>
>>109493509
what? it's already here >>>/wsg/6210036
>>
>>109493495
LIIIIIIEEEES!
>>
>>109493482
wasn't expecting to kek
>>
h3 is so big for us fartfags it's not even funny anymore lmao
>>
File: neet life best life.png (708 KB, 840x1014)
708 KB PNG
>>109493495
not my problem
>>
>>109493519
I respect the dedication to the hobby
>>
>>109493506
>https://sounds.spriters-resource.com/
Holy 2000 website
>>
File: 17974535427.png (84 KB, 511x675)
84 KB PNG
>>109493527
>>
>>109493435
did you prompt the slight breasts physics? it's perfect
>>
File: 03198-anima_baseV10.png (1000 KB, 896x1024)
1000 KB PNG
>>109493489
It's as dumb as dumb can be. I just stole it from here:
https://civitai.red/models/2663838/plaguekind-minimax-h3-ltx23-workflow-ease-of-use-eros-or-sulphur-compatible-or-faceid?modelVersionId=3198279
Also here's a catbox link that hopefully has the metadata. https://files.catbox.moe/9r0cwe.mp4. And the image which I just genned in anima if you want to experiment with it.
>>
>>109493485
pls make them kino and not "camera panning slowly through scene with unmoving characters" slop
>>
File: MiniMax_H3_00510.mp4 (3.21 MB, 832x640)
3.21 MB
3.21 MB MP4
>>109492965
Most of the prompt is written by 5.6 sol using the H3 prompt writing skill.
https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/SKILL.md
I also use a custom ekonte skill to improve natural pacing and scene logic. From there, i iterate on the results and direct how each scene should play out
>>
>>109493527
>h3 is so big for us fartfags it's not even funny anymore lmao
do you have a prompt to share/ any fart wisdom? i haven't gotten to that fetish yet
>>
>>109493491
I got you. I was too lazy to prompt more of an 80s look other than film grain.

https://files.catbox.moe/jmi75u.mp4
>>
>>109493548
we are so back it's unbelievable
>>
>>109493540
I didn't even prompt for jiggle physics, it just does it out of the box lel. You can see the prompt in the metadata of the catbox version
>>
>>109493559
soulless
>>
>not using fable ultracode to generate the prompt
Dumb fag
Dumb poor fag
>>
>>109493548
Anime studios should just swallow their pride and buy RTX PRO 6000 already to make animes
>>
>>109493548
amazing
>>
>>109493568
>I condier that you're not poor if you pay 20 dollars a month
are you living in Somalia or something??
>>
>>109493548
>5.6 sol
It's fine writing smut/ecchi?
>>
File: 1759636145969364.png (559 KB, 624x722)
559 KB PNG
>>109493548
>she knocks the other girl unconscious
>her first reflex is to stare at her ass instead of calling an ambulance
BASED???
>>
If I have access to an RTX Pro 6000, is there any reason to go above int8-convrot and 4bit text encoder for h3?
>>
>>109493548
>custom ekonte skill
que?
>>
>>109493577
OAI models can do anything so long as its not explicit. not even vanilla.
>>
>>109493583
>unconscious
well she seemed to be moving OK with her ass up in the air, stunned more like
>>
>>109493589
OK so describing a sexy scene is ok, but if it goes into actual sex it complains?
>>
>>109492953
2
>>
>>109493586
>int8-convrot
No.

>4bit text encoder
Go at least fp8.
>>
>>109493597
so long as its pure ecchi, yup
>>
>>109493548
Wrong shot at the end, the view is at the side of the bed instead of in front of the bed
>>
>>109493616
True. I'll delete the model immediately and throw my gpu out the window. Thanks, anon.
>>
no model on the market can compete with the reference model cause you LITERALLY have no limits on what you can do. it's literally like an evolution of how loras work, with no training needed.
>>
File: MiniMax_H3_00306_.mp4 (1.9 MB, 832x640)
1.9 MB
1.9 MB MP4
hmm hmm...
>>
>>109493616
real anime makes similar errors so its fine
>>
>>109493564
Why faggot. I like it.
>>
>>109493606
OK good to know. Last cloud model I tried was claude and it was unbelievably prude.
>>
>>109493622
It's still good, just my autism
>>
>>109493625
>you LITERALLY have no limits on what you can do
It can't make it actually feel like my cock is getting guzzled by my highschool crush, it can only make it look like that.
>>
>>109493638
a great way to make me sad
I'd rather gen impossible beauties
>>
>>109493602
Thanks, I'm guessing there's no reason to even consider the pruned weights for this either.
>>
>>109493638
>highschool crush
it's over
>>
>still caring about his highschool crush
how? i don't really remember anyone from there and i only graduated a few years ago
>>
File: MiMx-IMG_00054.webm (1.58 MB, 832x1248)
1.58 MB
1.58 MB WEBM
>gen time increases with every gen
Is this an issue with video models, or is it just on my end?
>>
>>109493602
>>109493647
>fp8.
this shit is deprecated
>>
why are u gae?

https://files.catbox.moe/y1j1iw.mp4
>>
>>109493647
You should go pruned, there is no difference in quality.
>>
>>109493654
I was really close to getting with her is the thing. I could've if I wasn't such a pussy as a young man.

But now thanks to AI and Minimax H3, I can live out my wildest fantasies and missed opportunities! :D
>>
>>109493659
windows ram management
>>
>>109493666
>a transginger
who the fuck would want to become a ginger??
https://www.youtube.com/watch?v=CbuYb6lLHX8
>>
>>109493670
>I could've if I wasn't such a pussy as a young man.
you still are
>>
>>109493670
>I could've if I wasn't such a pussy as a young man
if you learned from that experience you wouldn't be making posts like this
>>
i have never coomed so much
>>
BEST MODEL EVER

it can copy the captain alex MC, vj emmie.

https://files.catbox.moe/u0j770.mp4
>>
>>109493693
>why aren't you generating miku?
because he's gae
>>
>>109493693
source audio: clipped from here

https://youtu.be/d-P07qJ0NpI?t=75
>>
-.-
>>
o_o
>>
X_X
>>
File: h3-image.png (48 KB, 1593x276)
48 KB PNG
I wonder if this will dethrone Krea 2.
>>
>[reference generation] The face, chest and waist of <Subject 1>
>[Shot 1] Capture the face, chest and slim waist of <Subject 1>
>detailed_description:
Soft indoor lighting, realistic skin texture, focused on face.
Why the fuck are all my gens close-ups of her fucking tits...
>>
it's over
>>
>>109493711
I need to see how to use the speech tags for emotion more, but the voice is vj emmie

https://files.catbox.moe/n4oplq.mp4
>>
>>109493742
either way my cock is eating good this summer
>>
>>109493742
it will destroy it, minimax h3 has an amazing reference character consistency, if the image model is on the same level we literally have NBP at home
>>
>>109493742
I love the free market. Thank you China very cool.
>>
I'm just waiting for the first hitpiece on H3 to drop and the subsequent media outrage. Luckily everyone's too busy right now reporting on marketable security incidents where some llm escaped its badly secured testing environment and hacked nasa or something.
>>
>>109493762
to that end i could see it replacing edit models, but for raw image gen, idk. it's a huge fucking model, even if they cut down the vae.
could be worth it, who knows.
>>
>>109493766
I was told communism = bad.
>>
>>109493768
jeet sirs can't afford the API or the hardware to run it so i don't think anything will happen
>>
>>109493768
journalists don't care because there is no one company to blame or attack, twitter or big AI company bad gets clicks
>>
>>109493773
chyna is capitalist doe
>>
35 mins to generate this 15 second 720p (1 megapixel) video at 6 steps with the turbo lora locally on my 3090.

I cannot get the crows so do what I want though. They're supposed to be flying around Catherine (the gothic alien robot lady thing)

>>>/wsg/6210074
https://files.catbox.moe/fsty4w.mp4
>>
File: h3_00004.mp4 (2.92 MB, 832x640)
2.92 MB
2.92 MB MP4
>>>/wsg/6210073
need to got those colors constant across gens.
I guess maybe feed the last frame as first frame of next gen.
>>
File: debo_tt_k2_00019_.png (2.43 MB, 1872x1007)
2.43 MB PNG
>>109493742
krea is very good but its also very small. certainly can be toppled
>>
>>109493773
>free market = you have the choice to open source or not
>communism = you're forced to open source
not the same thing at all
>>
>>109493782
>krea is very good
but also very bad
>>
>>109493773
there is nothing communist with the way these companies compete with each others and with western companies
>>
>>109493757
commando.

https://files.catbox.moe/7spnmz.mp4
>>
Crowd made /g/ anime when?
>>
>>109493780
Nice gen
>>
>>109493742
If they go with a giant text encoder again, it could be an amazing model, but who knows.
Also a vae that doesn't suck.
>>
>>109493783
>corporatism = the government or disney can decide your model violates the rules
>>
>>109493784
>but also very bad
blame the cunt that told jeets about seedvr for that
>>
>>109493780
the music is really good
>>
is a man not entitled to the sweat of his brow
>>
>>109493795
did it happen though? Disney sued Minimax and Minimax was still able to open source their models
>>
>>109493797
i blame the vae and whatever ai-generated training data they used
>>
is there any practical way to run comfyui with clustered strix halo shit? id genuinely love to try doing one if i could FUCKING AFFORD IT
>>
File: MiniMax_H3_00103_.mp4 (2.32 MB, 1008x1392)
2.32 MB
2.32 MB MP4
>>
>>109493806
if you can't make krea 2 look good, then you aren't going to make minimax look any better
>uhhhh actually h3 is bad because they made the vae small for the image model! that's why my gens look bad.. trust me.
>>
>>109493742
they said they'll be using the same vae for the image model than on the video model, so Idk, what level is this vae? is this qwen image tier? flux 1 tier? flux 2 tier??
>>
>>109493813
good
>>
>>109493814
do you gen anime? ideogram 4 set the bar for realism and krea 2 comes nowhere close.
>>
https://litter.catbox.moe/hun4u6rrmglav58h.png
>>
>>109493780
really nice one anon
>>
>catbox.moe rangebanned brazil
lol, even r34 did it, big sad
>>
>>109493822
post one of your kino realistic ideo gens.
>>
https://civitai.red/models/2841940/h3-innie-pussy?modelVersionId=3208228

you're welcome, fags. you don't want to know how many days i spent on this shit
>>
>>109493833
slop, dont care
>>
>>109493833
thanks ill reupload it without credit and just specify anonymous like it should be
>>
File: kek.png (20 KB, 1147x65)
20 KB PNG
>>109493833
>you don't want to know how many days i spent on this shit
less than 6 days that's for sure
>>
>>109493833
pedotroon
>>
>>109493711
>watching some eceleb gaymer react to someone basically react to a video
that absolute state of zoomies and millennials
>>
get me the quick rundown on
>current h3 model to use that isnt nvfp4
>which turbo should be used or not used
>which custom nodes are needed for the above if applicable
>should i care for solattn or not
too much shit is going on at the same time too quickly i would wait but waiting is clearly a mistake right now too
>>
kinosol
>>
>>109493853
>but waiting is clearly a mistake right now too
why? this model isn't going to vanish
>>
ani just archived another milestone >>109492934
>>
>>109493833
anime?
>>
>>109493853
>>current h3 model to use that isnt nvfp4
int8 or bust, that shit gives you Q8 quailty and 2x speed over fp8 >>109493662
>which turbo should be used or not used
none, they're all undertained at the moment
>which custom nodes are needed for the above if applicable
spectrum + 20 steps
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
>should i care for solattn or not
Idk I haven't tested it out yet
>>
File: 1755116952915227.png (2.36 MB, 1328x1333)
2.36 MB PNG
>>109493831
>>
>>109493842
true, maybe like 1.5 days

>>109493843
nigga this model is so untrainable that you can't teach it how a real vagina looks like. maybe someone will figure it out hopefully..
>>
File: 16746.png (162 KB, 502x477)
162 KB PNG
wan2gp still doesn't support the int8 vae
>>
>>109493833
Finally, thanks anon, I hate outie roast beefs.
>>
This is a lie, it reduces output coherence slightly.
>>
>>109493873
>nigga this model is so untrainable
why those fags didn't gave us the base model... reeeeeeeee
>>
>>109493843
mutt
>>
>>109493868
Look! we have a Salgado who likes colors here!
>>
>>109493876
>using wan2gp
lol
>not vibecoding the feature yourself
LOL
>not vibecoding your own interface and using comfy as a backend
LOLE
>>
>>109493858
the longer i wait the more behind i get

>>109493861
>int8 or bust
thanks
>none
thanks
>spectrum + 20 steps
nice, any particular tip on this one's usage?
>idk
works for me, i notice no change with sage attention 1, and sage attention 2 i havent even tried yet, im just stuck on flash attention so i dont know what to even do on this front. my goal is to reduce time that doesnt need to be wasted
im already trying to make prompts but i feel like im going insane while reading the manual. it makes sense to me with how im supposed to make these prompts but all of the timing and lack of creativity on my part isnt making this easier. i just tried making a 15 second 1280x720 gen and it did work but it took nearly an hour and a half
>>
FRESH!!!!!!!!!!
>>109480220
>>109480220
>>109480220
>>109480220
>>
>>109493896
stop it with the stinky debo bread anon
>>
>>109493896
debo
>>
>>109493868
and what is stopping you from doing that with korea?
>>
File: 1784191419534805.png (942 KB, 784x784)
942 KB PNG
>>109493905
because krea cant do skin texture
>>
How do I get generated videos to look fast? They always come out looking slow, but I want the illusion of speed
>>
>>109493914
>>109493914
>>
>>109493915
i seem to get slightly better results out of it.
>>
File: h3-vae.png (138 KB, 1546x661)
138 KB PNG
For the anons saying quality issues are tied to the vae, it's not.
>>
>>109493899
this looks like a child.
>>
File: 1289528227949.jpg (261 KB, 1152x864)
261 KB JPG
brothers where in the H3 workflow should "Load LoRA" go? So far the results are ass, so I must be doing something wrong. Is that even the right node for H3?
>>
>>109494214
probably just load it right after the model unless it's a very specialized lora?
>>
How do I use a video reference? Comfy doesn't let me connect the "Load Video" note to the "ref_video_0" point.
>>
>>109493570
just delete yourself.
>>
>>109494285
Yeah I tried both that, and after Spectrum Apply, and I get nightmare fuel. Tried a few different loras too, I'll play around with it a bit more.
>>
>>109494214
>>109494285
>>109494325
Turns out it was the lora strength. Placing "Load LoRA" immediately after "Load Diffusion Model" and 0.2 strength had excellent results.
>>
>>109493192
>these kinos get 0 yous
lmao at the offended browns
>>
>>109493126
depends how you look at it, wan 2.1 vs most things before was huge, zit realism, size, speed and resolution ootb was huge compared to the previous slow chroma, h3 is also huge

i think zit was the biggest outlier since there is nothing similar that came out that optimized and improved all the things i mentioned as well as zit. wan 2.1 was slightly better than hunyuan and so it won out, h3 is truly insane and better than most proprietary models too but a big factor is local not having a proper wan successor for 1.5 years so h3 had the time to cook with a lot more optimizations and features, insane seedance 2.0 model to train on, and proper youtube and internet data scraping pipelines to use.

i think in the ai space the biggest "novelty" jumps that werent just things that incrementally improved until one of the models hit a "milestone" were:
1. zit
2. first mixtral 8x7
.
.
5. noobai vpred colors

otherwise, in terms of general capability jumps i think h3 is one of the top if not the top model ever (if we dont include going from nothing to something which would always "win" these)
>>
>>109495479
true zimage and and this minmax model are big time changers
yet klein is no mentioned and it does fit into this due to low gen speed and requirements while having quality
>>
>>109495969
klein was an good jump in most ways but its still not pixelspace, not hugely better than last QIE, it just didnt fully solve any problems like consistency, very deep knowledge, no color/pixel shift etc
>>
Anyone would happen to have it saved please?
>>109492070



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.