[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1777297867747972.mp4 (961 KB, 576x736)
961 KB
961 KB MP4
Previous: >>109482228
https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Blessed thread of frenship
>>
Forget the turbo lora. Turbo is too slow. We need ludicrous speed lora.
>>
File: 1757267630921245.png (143 KB, 1080x989)
143 KB PNG
Let's FUCKING GO

ACCELERATE! ACCELERATE! ACCELERATE!
>>
File: 907811248090899.mp4 (3.56 MB, 800x512)
3.56 MB
3.56 MB MP4
>>
>>109484153
/g/ is back, baby
>>
>>109484827
Buy high sell low!
>>
r/Midjourney mogs this place so hard
>>
>>109484153
I commend porn autists for single-handedly keeping /g/ alive
>>
>>109484922
true, I use that subreddit to find cool images so that I can do some I2V videos, there's still no models that managed to capture so much sovl like Midjourney
>>
>>109485064
faxx
>>
Trying out max ref images for 15 seconds at 1mp it ends up being 13 minutes with sage attention node. Let's see how it goes
>>
File: MiniMax_H3_00049.mp4 (2.51 MB, 832x1248)
2.51 MB
2.51 MB MP4
https://files.catbox.moe/0n93fm.mp4
>>
>>109485233
close up talkies, ltx's forté
>>
good to see you absolute retards still haven't managed to properly delegate baking since the 2022 /sdg/ days
>>
>>109485289
hey

fuck you
>>
>>109485301
>t. makes 3 threads at the end of every thread
>>
>>109484922
midjourney is literally just press x to awesome in AI form, butt he flavor of awesome is always the same
>>
Are people seeing meaningful improvements with the chunk and low vram nodes or is it purely to prevent errors? I just ran a gen that was fully crashing my machine before but it took 42 minutes, and I'm not really seeing improvements on lower res (.3mp at 15 sec type) gens.
>>
>>109485233
nice to see it can do shitty Grok quality with the audio as well, would be good to troll the grok thread on /gif/
>>
>>109485344
I can increase the mp though it'd take longer, I think the audio is OK
>>
File: 1768154422354680.jpg (993 KB, 1616x3008)
993 KB JPG
>>109484922
geg
calarts subreddit
>>
>>109485358
I wasn't dissing the actual quality really, more the style of how Grok videos are, they just have that AI feel and the audio has the same AI feel to it.
>>
>>109485092
Isn't ref the one that can do everything fl can do, but with more image inputs?
>>
.
>>
/g/ - Neural Networks
>>
Solved my PrompExecuter cache reset problem.
It was the multigpu crap node that I haven't used forever.
Don't be like me anons, clean your nodes from time to time.
>>
we can actually make a south park the creators are afraid to make now.

https://files.catbox.moe/q8gt32.mp4
>>
>>109485380
Is that a male or female?
>>
How do you slop longer videos? Just outline what's happening by second with an LLM and chain first/last frames?
>>
>>109487559
95% pie chart is killing me
>>
Can someone redpill me on sol attention? Is it better than sage? Can or should they be used together? Any side effects?

If I should use it, should I be using the kijai version or this?(which has more options):

https://github.com/Saganaki22/ComfyUI-sol-attn
>>
I have no idea i just add them all
>>
File: MiniMax_H3_00265.mp4 (2.77 MB, 672x1216)
2.77 MB
2.77 MB MP4
>>
File: 1761000612864614.webm (3.15 MB, 864x1344)
3.15 MB
3.15 MB WEBM
Finally did this on Minimax.
I use the sexgod loras and prompt the ice stick as penis
>>
>h3 supports multiple keyframes for I2V, but comfy devs were too lazy to add support for it in the default node
reeeeee
>>
File: 1760577324456909.png (163 KB, 1884x549)
163 KB PNG
>>109487588
south park anon, having success with this so far:
>>
>>109487603
You mean similar to LTX Director node ?
>>
>>109487575
>Can someone redpill me on sol attention?
AI models have a lot of weights. Not all weights are important for everything. If we can identify/guess most important weights for given task, keep them as is, approximate the rest, it's possible that we can still get good results with a significant speed up.
>Is it better than sage?
Sage is better in the sense that it's more reliable and mature. There are a lot of parameters and open questions as to how to implement Sol without any consensus.
>Can or should they be used together?
In theory they can since they modify different things. In practice I was unable to. As an Ampere cuck I don't have access to full featureset of both. Maybe your 4000 or 5000 GPU can.
>Any side effects?
It's still highly experimental. You will likely need a lot of parameter tuning for good results.
>should I be using the kijai version or this
Kijai's version also runs on older GPUs.
I can't run the other one because it needs >=sm90, which again I don't have.
>>
File: 1774018982002883.webm (2.45 MB, 640x880)
2.45 MB
2.45 MB WEBM
>>109487601
LTX comparison
>>
>fresh copy of comfy
>python -m pip install triton-windows
>3.13.14 python libs and include inside comfy
>sageattention 2.2.0 cu130torch2,10.0andhigher.post6
did I install it correctly? I'm on 3090
>>
>>109487604
There is no point in Sampling node if you are not going to add it to the scheduler.
>>
File: 1772478213737236.png (800 KB, 1323x691)
800 KB PNG
>>109487404
I'm trying this now, my source (2) images were 2x-3x the output resolution and i did note the time was very long comparitively but didnt think much of it as id not done many gens with source images.
Yes it's definately quicker, a few low single digit % in my use case, allowing for a bit of randomness.
773s vs 813s

https://i.4cdn.org/wsg/1786105376122066.mp4
>>
>>109487601
Could you please upload it with your workflow included? That's some really good progress.
>>
HAHAHA

Praise China for this open source model. I need to refine it but it actually worked. 222s at 0.3mp/15s.

https://files.catbox.moe/k8sq86.mp4
>>
>>109487621
Yes
>>
>>109487635
The setting is the TV show South Park.

0 to 5s: Cartman is flying a f16 jet over the middle east. The camera shows an overhead view of a middle east town with a large sign that says "Gaza". Cartman says "Captain Cartman here, permission to engage over.".

5 to 10s: view of Cartman from the front of the jet. Cartman says "No 72 virgins for you, Mohammed", and then presses a button. External view of the jet dropping bombs that are falling towards teh city.

10s to 15s: the bombs make a large explosion in the city as many middle eastern people run in a panic. One of them shout "Oh no, my 8 year old wife!"
>>
File: max_optimization.jpg (715 KB, 3224x2127)
715 KB JPG
(((this one simple trick changed my life)))
>>
>>109487641
It's getting very hard to follow at this point
>>
>>109487603
Is that with the i2v model? You can definitely do it with the ref model with their provided nodes/workflow.
>>
Bros, is minmax unironically better at generating music than all the dedicated music generation models? I'm starting to feel like it is.
>>
>download more RAM
>it literally can't run at full speed because you have a shitty dell mobo
>gen times actually get worse

I'm in hell. Speaking of, have vramlets seen actual speed increases from the chunk+low vram nodes, or is it just supposed to prevent OOM
>>
slightly higher res (not too relevant for south park gens)

https://files.catbox.moe/ooe3as.mp4
>>
So many """""optimizations""""" it gonna make your results hallucinate more often
>>
>>109487652
make every gen count if you are in such a hurry
>>
>>109487652
>have vramlets seen actual speed increases from the chunk+low vram nodes, or is it just supposed to prevent OOM
I want to know this too.
So many optimizations to try.
At least better than being forced to wait forever like Wan I guess.
>>
>>109487670
>greenscreen explosion
simply sovl. southpark perfectly imitated
>>
>>109487670
shouldn't he say 6?
>>
Do you guys have a good LLM prompt to enhance simple prompts to run better with the model?
>>
>>109487693
Hello Sir here pdf reference to do prompts do the needful and make gorgeous looks for the following:
>>
>>109487693
I made a tool that allows you to prompt an llm for each relevant section, it also feeds the guide and the current version of the prompt as well as any pictures I want to include with the instruction. I think its best to do it piecemeal instead of trying to have an LLM build the whole thing.
>>
>>109487641
you can use cache node with spectrum? I thought you can only pick one.
>>
How many images of a specific art style do you need in order to "train" a model for that output????? Also how should I go about doing this?
>>
File: 1778566419581117.png (1.38 MB, 1874x1324)
1.38 MB PNG
>>109487641
I was joking about the clattering bones thing, by the way
>>
>>109487706
>If you don't want to type multiple paragraphs any time you want to gen something you are Indian
Ok.
>>
How do you guys feel about the /trash/ diffusion thread?
>>
>>109487636
weird. Compared to other anons, my gens are slower than expected...
>>
File: annDxWd9_700w_0.jpg (42 KB, 600x600)
42 KB JPG
>>109487601
>she's still opening her mouth everytime to do the bite motion
we need better nsfw loras asap
>>
So I actually started getting better looking gen's when I stopped using my upscaled images as references. It seems obvious now but yeah I think the down scaling process was ruining my outputs
>>
Hello sirs what upscaler to use??
>>
Thoughts on the new spectrum update? Did it make things better?
>>
>both spectrum and h3 cache are now throwing errors when starting gen
back to easycache I guess
>>
File: 1781059738236384.png (174 KB, 1749x976)
174 KB PNG
>>109487604
did you try the recommended order?
>>
https://files.catbox.moe/1lv264.mp4
>>
https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/
>>
>>109487820
did u prompt breast jiggle?
>>
File: MiniMax_H3_00270.mp4 (2.72 MB, 672x1216)
2.72 MB
2.72 MB MP4
>>
>>109487834
I just wanted them to sway and hang naturally. For some reason, minimax decided that they want to be excited
>>
File: file.png (448 KB, 526x355)
448 KB PNG
How do you stop faces from looking fucked up like that in so many gens?
>>
kek, another seed

https://files.catbox.moe/3ffn0s.mp4
>>
>>109485233
Did you prompt for Tifa to slightly look at the other girl in the middle, because it's so natural it's impressive.
>>
>>109487820
nice, prompt?
>>
>>109487820
convention gens are kino
>>
anons, how do you prompt for bouncy boobs/ass, it seems like the concept isn't associated with these words
>>
>>109487794
yep I have that and the sage attn kijai node after spectrum on auto
>>
>>109487784
It's faster as claimed but the delta is also a lot higher.
Old spectrum would make little changes to even vaguest t2v prompts, now the difference is larger.
I am testing, I don't know if the quality is high enough to call just seed variance yet.
>>
HOW DO I KEEP THE FACES AND EYS FROM GETTING FUCKED!?! HELP!!
>>
>>109487876
>anons, how do you prompt for bouncy boobs/ass, it seems like the concept isn't associated with these words
no need to prompt it, it does jiggle physics by itself
>>
>>109487890
gen at a higher resolution
>>
>>109487892
1mp is fucking it up too
>>
>>109487889
What about I2V and R2V? Are you saying the output is noticeably different now?
>>
File: 1783810583440658.jpg (55 KB, 600x600)
55 KB JPG
I dont get it. Why i get an error when trying to update comfyUI to nightly ?????????????????????? Did i do something wrong ? Should i reinstall comfyui portable again because installing SageAttention is a PAIN in the ASS
>>
File: 054.png (947 KB, 2637x899)
947 KB PNG
Used some anon's starting frame here, don't get mad.
It's so simple bros. Could this be the end of offmodel R34 garbage?

https://files.catbox.moe/kkq83h.mp4
>>
>>109487890
Stop using optimizations. Try to gen with default minimax workflow
>>
Man bong tangent is so fucking crisp and good, too bad is like a x2 increase in gen time.
Why everything good has to cost so much...
>>
>>>/wsg/6209775

I hope the hair isn't done with brahmin poop.
>>
>>109487933
but gen times :/
>>
Would using seedvr2 somehow unfuck the faces?
>>
>>109487929
the shittier ones maybe, but the few quality ones not really, some animators have insane talent
>>
>>109487933
I think they still got fucked on day one, when we had none of it. At least at the resos I could gen at (up to 1.2 MP).

I don't know if anyone has mentioned it here, but you can try the LTX upscaler for H3.
>>
>>109487943
no I get that, smaller faces will be fucked. its the plastic look as well. the textures in your gen are great. thats what Im lacking
>>
>>109487929
>https://files.catbox.moe/kkq83h.mp4
That third pic is also off-model though
>>
yeah the main issue with h3 is the smaller faces being fucked up (and generally smaller details)
I wonder if their secret sauce upscaler they didn't release yet would solve that
>>
okay, now we're in the middle east. problem with the partial jibberish was a 5 second block (0 to 5s:) but brief dialogue and not being specific enough. NOW it's good.

https://files.catbox.moe/ooe3as.mp4
>>
>>109487955
how many steps are you doing?
>>
>>109487923
try this. remember to backup before trying
>>
>>109487876
simply say "breasts bounce wildly" or gently or however much you like. it will do it.
>>
>>109487960
Well do we have a madman with a H100 that can run the full model at high res to see how that helps.
>>
Is August officially the christmas period of video gen? Last year we got wan2.2 which was a huge upgrade over wan2.1. This August we got Minimax H3, which is the greatest leap in progress for local media generation since Stable Diffusion released.
>>
>>109487962

20 when I wasnt using the turbo lora. at 0.7mp

10 when I used the turbo lora at 1mp
>>
>>109487970
Isn't the model maximum "recommended" resolution 720p, and anything above is just their upscaler model even in API?
>>
>>109487929
>Could this be the end of offmodel R34 garbage?
I'm afraid it's almost over for 3D artists and animators except maybe those who already have an established career.
>>
>Queue up a 20 second gen at 0.5 MP, expect it to take like 10-15 minutes
>Come back 20 minutes later and it's still running, realize I forgot to change it from 1MP, gen is only like 30% done.

Goddammit.
>>
>>109487923
just use easy comfy. best shit ever
>>
>>109487972
yeah, this is why doomers are always wrong with this, they always say "it's the last local model", and every fucking time they were wrong
people still listen to them for some reason
>>
File: MiniMax_H3_00217_R(1).mp4 (3.39 MB, 1088x1920)
3.39 MB
3.39 MB MP4
>>109487960
>smaller faces being fucked up
gen at 2mp
>>
>>109487973
try without the turbo lora, its known to ruin the texture
>>
>>109487967
Jiggle sometimes seems to make them go nuts. I just want a little tiny bit o' jiggle and they start flying all over the place.
>>
>>109487929
bricked the file, reuploaded
https://files.catbox.moe/o3rwsz.webm

>>109487956
Pretty close to the desired so that's okay, H3 alone wasn't adapting the face too well without a close ref
>>
>>109487929
shitty badly animated ones : yeah they're fucked or at least they have to release for free
good ones : no they can do very intentional details
>>
grok is pretty good at formatting prompt ideas fast. but the text encoder is so good that just telling it what you want and what shot also works well. but for elaborate stuff, can just use AI to get a fast block of text.

ie: Make a video prompt for Cartman from South Park making fun of AI slop. Format it in 3 parts with timestamps, 0 to 5s:, 5 to 10s:, and 10 to 15s.

(gen in progress)
>>
Anyone have a good workflow which works well for 24GB VRAM (3090)? I want to fuck around with Minimax but I've been waiting for the speed optimisations.
>>
>>109487983
what gpu do you have
I don't think I can gen at 1080p even with my 5090, unless I missed amazing optimizations
>>
>>109487820
Why is she wagging her tits out of excitement
>>
>>109487914
I only tested I2V once and that had basically no difference..
I am leaning towards "quality is still good" but I am not certain yet.
>>
Is h3 able to make characters talk in Japanese without it sounding gibberish?
>>
Can anyone explain to me what us plebs are missing out on with the non-pruned model?
The pruned model is amazing but I'm curious as to how much more capable non-pruned is, or what kind of advantage it has specifically. Pruned cuts the size down a lot so I assume we're missing out on something.
>>
>>109488012
Yes
>>
>>109488000
5090, no turbo lora, that gen took 22mins
>>
>>109488012
It can talk in languages with a few million speakers, albeit with strange accents at times. Pretty sure it can do Japanese quite well.
>>
>>109488017
OK that's amazing, can you share the workflow? unless it's already in the video (not home so can't check)
>>
>>109487992
kino!

https://files.catbox.moe/2cgc74.mp4
>>
>>109488012
Yes. Here is my gen with KAORI's cloned voice:

https://d.uguu.se/BHnFnapu.webm
>>
https://civitai.red/models/2834417/hmnsfw-aio-sex-lora
hearmeman all in one sex lora v2 for h3
>>
untested aio lora for all you h3 gooners https://civitai.red/models/2834417/hmnsfw-aio-sex-lora
>>
>>109488012
Yeah >>>/wsg/6209192 You can even slip English words in seamlessly
>>
>>109488016
>>109488021
OK that's pretty cool.
And I thought we'd never go beyond ltx2.3 horrible voice quality, boy was I wrong.
>>
>>109488025
the voice is spot on, did you give it a reference?
>>
File: 1531721121424.gif (2.44 MB, 400x304)
2.44 MB GIF
>>109488027
holy FUCK
this model is way too fucking dangerous, jesus christ
>>
>>109488042
t2v. it knows south park natively so it works very well. reference workflow can input any voice which is amazing too.
>>
is krea2 significantly better than zbase in photoreal?
>>
>>109488027
OK finally my dream of the cutest seiyuus' voices saying whatever I want to hear is coming true!
>>
>>109488027
How long of a voice reference did you give it? One clip or multiple?
>>
>>109488049
Emad and company were right, image models really are unsafe and they kill you through dehydration.
>>
>>109488027
god damn, the intent and inflection is perfect too
>>
>>109488072
>Emad
Rereading how he self sabotaged after all these years is kind of comical.
>>
>Lightx2v first 4step lora already dropped
lel
>>
>>109488060
I just do what I've always been doing with tts voice models: compile a 10-15 second clip from multiple clips of clean dialogue with minimal noise/ambience then encode to .wav.
With >>109488027 I ripped May speaking clips from the Pokemon Jirachi movie. Movies are always better because there are more scenes with silent ambience and only talking, whereas the TV show has constant background music.
>>
>>109483968
This looks worse than kiggers
>>
HOLY SHIT
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main
>>
>>109488094
Great, now I gotta test this shit and compare it. I hate this part of getting a new model. I'm exhausted.
>>
>>109488094
>v0.1
>>
>>109488094
Going to need weight settings before I waste my time using it.
If it produces a slow-motion effect like it does with wan, no thanks. Minmax's native motion handling is one of the best parts of the model.
>>
>>109488094
pls... I can't take any more turbo loras...
>>
>>109488094
please be good please be good please be good
>>
>>109488094
this will prolly need to be convert to comfy format :(
>>
File: 1767490825176230.png (7 KB, 967x34)
7 KB PNG
>pull latest FE
>this is in the logs
UHMMMM BROS?!?!?!?
WAN3?????
>>
Low Vram attention node just slightly slows it down for me.
Maybe it is good if it prevents offloading but seems useless if you are offloading a massive amount already.
>>
>>109488101
same. too much stuff everywhere.
>>
>>109488122
yeah, via API
>>
>>109488109
No slowmo it seems so that's a good sign.
>>
>>109488136
What strength/step count you testing?
>>
>>109488085
>>
https://d.uguu.se/EajtPteG.webm

cannot even begin to tell you how long i've spent trying to make something like this on wan. then minimax just does it easily.
>>
>>109488122
APIslop
>>
>>109488094
Get ready for slowmotion feast.
>>
>>109488122
Partner nodes API only:)
>>
>>109488122
someone posted a tweet the other day with some wan 3.0 videos. looks like fucking shit.
>>
Is the soon to be locally released FLUX 3 Video good or just "look at this animals and landscapes", while completely failing at anything human beyond stock footage?
>>
Any /trash/ anons here? What is the difference between /sdg/ and /slop/?
>>
>>109488143
I will never understand you people but I'm glad your dreams came true.
>>
>>109488143
you missed the occasion for her to say "nii-sama"
>>
>>109488161
Bro you're gonna be able to gen so many astronauts riding unicorns on the moon.
>>
>>109488161
someone posted comparisons between H3 and Flux 3 Max. Flux 3 max is maybe 5-10% better but not worth dealing with at all if it's any more annoying to use in any way compared to H3
>>
>>109488163
the link is different
>>
>>109488161
It's the go to choice for the safety conscious genner
>>
>>109488094
having to juggle all the solattns, sageattns, memeff patches, chunk feedforwards, spectrums, caches AND make them work with different turbo loras
i'm tired boss, I just wanna gen...
>>109488105
this
>>
>>109487729
shiver me timbers and rattle my bones you fucking cad
that said this does further add credence to the theory that this model works better when you DON'T prompt for certain details you want. tit jiggle, succ sounds, wet pussy sounds, that's three i've noticed already.
>>
>>109488027
How did you manage to emulate her voice ?
>>
>>109488168
lol, I expect it to have zero knowledge about anything potentially unsafe or copyrighted (and knowing them they'll make a specific finetune to change anything nsfw to sfw)
>>
File: 1570969333863.jpg (41 KB, 575x603)
41 KB JPG
>>109488164
>I will never understand you people but I'm glad your dreams came true.
>>
>>109488170
going to be a hard sell unless the model is faster, smaller, trains easier, and has nsfw without body horror.
>>
>>109488188
nta but you just use a 10-15 second clip of the character speaking as reference audio and add <Audio 1> is the voice-timbre reference for <Subject 1> to your prompt
>>
File: MiniMax_H3_00229.mp4 (2.98 MB, 672x1216)
2.98 MB
2.98 MB MP4
>>
>>109488167
This was I2V. If I was going to use dialogue then I would do R2V, but that's a lot more effort setting up the proompt and clooning the vooce.
>>
>>109488182
>that said this does further add credence to the theory that this model works better when you DON'T prompt for certain details you want. tit jiggle, succ sounds, wet pussy sounds, that's three i've noticed already.
I want a finetune or even lora where these are understood and steerable, instead of being random.
I'm glad the eros guy is making a nsfw finetune but I hope it doesn't destroy what the model can already do and replace it with stock 2000s porn.
>>
grok test: redo the prompt but the joke is Cartman is at the DMV trying to say he is a transgender woman to get cheaper insurance rates.

lmao

https://files.catbox.moe/q32li1.mp4
>>
>>109488168
>>109488170
>>109488174
OK, I guess I'll expect nothing out of it.
>>
>>109488188
Refer to >>109488083
I had a lot of practice doing this with stuff like index-tts2. Minmax handles it the same way (sometimes there are problems referencing audio though).
>>
>>109488170
user friendliness is ultimately what's winning everyone over with h3. not just its quality, right now 99% of the gens have the exact same low quality noise patterns and compression, but there's hope enough to go around to wait for the upscaler/further improvements.
If it weren't literally braindead easy to use it'd be DOA because of the spooky scary parameters.

>>109488203
It's not random, it's ((context sensitive)).
>>
>>109488193
it's a good message, better than the usual tourettes one
>>
>>109488194
>going to be a hard sell unless the model is faster, smaller, trains easier, and has nsfw without body horror.
has BFL ever released a model that fits into any of these categories?
>>
File: 1766754829222131.png (61 KB, 600x535)
61 KB PNG
>>109488209
>>109488083
Im too brainlet to understand this
>>
Will Flux Video be able to compete with Minimax on boob jiggle? Hell, will it even know what a nipple is?
>>
I'm heading to my local police station to turn myself in.
>>
>>109488221
>I'm heading to my local police station to turn myself in.
thanks for the prompt inspiration
>>
>>109488060
>>109488027
I've tried it with Meili, it's pretty amazing how well it matches the way she speaks.
https://n.uguu.se/jdipTgei.mp4
>>
>>109488218
>Will Flux Video be able to compete with Minimax on boob jiggle
I'm pretty sure they'll make a pass just to replace any jiggle with perfect unmoving bricks.
>>
>>109488218
It would be a real SD3 moment if they pruned the dataset from any busty woman moving.
>>
>>109488170
flux will 100000% be censored and have far less copyright training data so it will be worse.

also, the reference minimax model makes anything possible.
>>
>>109488216
lol uhhh can i interest you in an astronaut riding a unicorn through the misty hills of scotland?
>>
>>109488194
if Flux3-dev uses a configurable guidance distillation like the others did, then you will be able to set guidance to 1 and train loras on it. you can't do this with minimax, it's like a flux-dev model but with guidance locked at 4

Assuming flux3 isn't some fuckhuge model, I think quality doesn't really matter. trainability will determine which one wins, and minimax is honestly kind of retarded for not using a configurable guidance, or failing that, release an undistilled base model
>>
>>109488228
Voice and tone is ok, but sound quality is kind of bad, almost ltx like, is it because it's the turbo lora?
>>
>>109488237
i would honestly be ok with a cucked model if it can crank out kino sfw shit. h3 can be the coom king, flux can be its dainty prudish wife.
>>
>>109488232
>the reference minimax model makes anything possible
i think this is the actual biggest thing about the H3 release. with H3, we are now over the hill, and the actual only limit to what we can create is our creativity and time
>>
>>109488241
Nah, it's because I didn't had a voice sample of Meili at hand so I ripped it off youtube and used stem extraction because it had a BGM. It would probably sound way better if I had better audio.
>>
>redo the prompt but the joke is Cartman watching The Odyssey and is upset Helen of Troy is black

0 to 5s:
Eric Cartman from South Park sits on his couch in the living room, remote in hand, watching a big-screen TV showing a modern adaptation of The Odyssey. His eyes widen in shock as Helen of Troy appears on screen as a Black woman. He freezes mid-bite of a Cheesy Poof and stammers in his high-pitched voice: “Wait… what the hell is this?!”5 to 10s:
Cartman jumps to his feet, pointing furiously at the TV as the scene continues. His face turns red with rage while he yells: “Helen of Troy is supposed to be a beautiful Greek chick, not some Black lady! This is pure woke garbage! They’re ruining the classics again!”10 to 15s:
Cartman throws the remote at the screen, stomps in place, and faces the camera with pure disgust, still ranting: “First they race-swapped everything else, now Helen of Troy?! Screw this PC Odyssey crap—I’m out!” He storms off as the TV keeps playing behind him.

lmao

https://files.catbox.moe/4ucym4.mp4
>>
>>109488245
what are you retarded? just use H3 for sfw shit if you want it. that's its main purpose. you don't realize you're already living in an amish paradise.
>>
>>109488217
Here I'll break it down simply:

1. Find a voice you want to clone.
2. Open a video file where said voice plays
3. Find all the examples of the voice speaking with no other background noises. Just speaking, no sound effects.
4. Extract every clip you can fine using something like XMedia Recode or Handbrake til you have about 10 seconds worth of audio clips.
5. Download Tenacity, then combine all those clips together in one single long clip.
6. In ComfyUI R2VA workflow, load the single long clip in a "Load Audio" node, and plug it into ref_audio_0
7. Follow the prompt guide to make it work in the video properly: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
>>109488218
>BFL runs a safety tuning process on all of its models before release. The models are initially trained on NSFW content, but the subsequent safety tuning attempts to make them unlearn or suppress that knowledge. This is why you can end up with body horror-like anatomical failures across BFL models: sexual positions and ordinary, non-sexual poses can share similar learned representations. When the safety tuning tries to suppress one set of representations, it can inadvertently interfere with the other, degrading the model’s broader understanding of human anatomy and posing.
>>
>>109488231
They don't even need to prune it, all they need is to train "jiggling breasts" concept as meaning "unmoving thing".
That's what they probably do with many nsfw concepts.
>>
>>109488245
H3 can do sfw fine, my issue with flux video is that if they didn't train on porn, I feel like anatomy itself will be inferior to h3, even purely on sfw outputs.
>>
File: chad horse.jpg (35 KB, 736x971)
35 KB JPG
>i can animate full music videos for my favorite songs using my waifus now
>shit they can be nude too
holy fuck i'm not gonna get SHIT done this weekend
>>
>>109488256
Wait, Minimax can Reference Audio as well ??? Goddamn......
>>
>>109488278
thats my biggest concern.
>>
>>109488260
>>109488266
It's kind of insane when you think about it: they'd rather fuck up sfw anatomy just to suppress nsfw.
Even API models like sora or kling and so on had nsfw understanding because it's useful data, they just rejected nsfw prompts or outputs.
>>
>>109488278
I don't think lack of NSFW is this reason for raped flux anatomy.
I think it is caused by safety-troon post training shit they do, like training on synthetic images with Barbie doll genitals.
>>
>>109488279
Have you figured out how to prompt for replacing the character in an existing video yet?
>>
>>109488292
i get not training suck and fuck, but basic nudity is something every model should be trained on.
>>
File: 3261108054768120.png (1.3 MB, 1152x640)
1.3 MB PNG
>>109487973
I'm using the ema_ckpt850 turbo lora at 0.7mp and 8 steps. Textures are not as good as no turbo, but they seem fine to me.

https://files.catbox.moe/n263r3.mp4
>>
File: Cynthia_Onsen_015.png (1.03 MB, 640x1536)
1.03 MB PNG
https://d.uguu.se/pYZwRisr.mp4
>>
>>109488292
>>109488302
genuine pondering, what the FUCK made them think all of these extra steps were necessary? Who where they trying to impress? The US gov? I don't think they give a shit either way.
It's such a funny extra twenty steps way to burn their goodwill to the ground and look like retards while doing it.

>>109488308
that was the first test i ran for when i grabbed h3 days ago, it went half-fine, and that was my first try.
as soon as my sleepybrain clears and caffeine kicks in i'm giving a really simple clip a shot to see how it goes.
one major plan i have is grabbing four of my original a.i characters and putting them into one scene where they're walking side by side.
shit i should do that first.
>>
>>109488281
Yes, read the guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

It can reference audio in multiple ways. Copying, referencing, and weak referencing.
The R2V model + workflow is like an extremely powerful toolset.
>>
>>109488281
yes, here is master chief for example.

>The camera focuses on <Picture 1> as he speaks in a deep, gravelly voice: "I don't give a fuck". His tone, pace, and vocal style perfectly match <Audio 1>, with precise lip-syncing.

https://files.catbox.moe/ya0wtt.mp4
>>
>>109488328
oops, I changed the dialogue since that note in the workflow. but it works.
>>
>>109488313
nice
>>
>>109488313
holy cum
>>
File: MiniMax_H3NoAudio_00014_.mp4 (3.13 MB, 736x1280)
3.13 MB
3.13 MB MP4
>>
>>109488143
is it censored because ai still can't gen peni and vagina?
>>
I've found h3's weakness. it cannot gen smooth armpit. it always have some armpit hair stubble
>>
>>109488342
love the 80s/90s anime aesthetic
>>
>>109488354
good good... very good..
this means it's capable of pubic hair. and treasure trails..
>>
>>109488354
smooth brain issue
>>
>>109488354
steak too juicy, lobster too buttery?
>>
>>109488354
>I've found h3's weakness. it cannot gen smooth armpit. it always have some armpit hair stubble
if you give it a reference of a hairless armpit it will do it. there's no weakness anymore
>>
kek

The setting is the TV show South Park.

0 to 5s:
Eric Cartman from South Park sits in a comfy chair in the living room, watching a modern adaptation of "The Odyssey" on a big screen tv. He’s munching Cheesy Poofs when the tv screen shows Leonidas from 300 charging into frame in full Spartan armor, spear in hand, as he collides with Elliot Page. Cartman’s eyes go wide in surprise.

5 to 10s:
Leonidas stops short, points the spear accusingly at Elliot Page, and bellows in a deep, commanding voice: “You’re not Greek!” Cartman leans forward in the comfy chair, grinning and cheering, while Elliot Page looks shocked on the tv screen.

10 to 15s:
Leonidas thrusts the spear forward and stabs Elliot Page. Cartman jumps up from the comfy chair, pumping his fists and laughing maniacally in his high-pitched voice as the tv screen shows the aftermath, yelling “Yeah! Get ’em, Leonidas!”

https://files.catbox.moe/wq75p0.mp4
>>
File: MiniMax_H3_00276_R.mp4 (3.96 MB, 672x1216)
3.96 MB
3.96 MB MP4
>>
>>109488251
got to include
"Screw you guys, I'm going home!"
>>
>>109488395
I'm no pits guy by any means, but mama mia.. H3 you dog..
>>
>>109488345
Minimax absolutely can do pussy. It doesn't even need a reference, although it's not fully supported.
It can generate the front view of a vulva with slit, but don't expect detailed labia or anything.
Pussy from the back? : Forget it.
I haven't tried using R2V using a pussy reference image and telling the model to apply it.
>>
>>109488308
You use the reference model with the video edit mode. Give it a picture and the video and say something like replace the <describe character in video> in <Video 1> with the Character in <Picture 1>. You can also add in a subject reference to the picture.

For me the hardest part has been getting it to recognize the right character to replace if there's a few characters in a scene, like the same prompt will replace different characters across generations, even if I'm being pretty specific (i.e. the girl on the left with red hair and white dress). I'm probably doing something wrong.
>>
yeah lightx2v's turbo lora doesn't work for me, not sure what's going on. i'll just wait a few hours
>>
https://huggingface.co/Kijai/MiniMax-H3_comfy/blob/main/loras/minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors

have not tried it yet
>>
>>109488014
Anyone?
>>
>>109488424
did you get the comfy one?
>>
>>109488425
why would you bother, none of them are ready but I'm glad there's competition to be the defacto turbo lora
>>
File: MiniMax_H3NoAudio_00015_.mp4 (3.85 MB, 1120x832)
3.85 MB
3.85 MB MP4
>>109488357
Looks more authentic at 4:3 aspect ratio
>>
light 2x lora is amazing at 6 steps. Deleting the other turbo loras
>>
>>109488453
its nice the model can do a wide range of styles, they trained this on seemingly everything.
>>
>>109488433
>did you get the comfy one?
what comfy one? there is one safetensors file in the repo
i tried to convert it using Kimi but it was still fucked
>>
Are the turbo loras usable on the ref model?
>>
>>109488459
Strength?
>>
>>109488463
>>109488425
>>
>>109488464
It specifically has fl2v on it so would lean towards a no
>>
>>109488419
I've tried a pussy reference, it kind of tries maybe but it doesn't really work
>>
>>109488321
>genuine pondering, what the FUCK made them think all of these extra steps were necessary? Who where they trying to impress? The US gov? I don't think they give a shit either way.
>It's such a funny extra twenty steps way to burn their goodwill to the ground and look like retards while doing it.
They're basically obsessed with safety to the point of sabotaging their own model, it is what it is.
>>
>>109488468
1
>>
>>109488334
so did it invent part of the dialogue or is it exactly what you wrote
>>
https://x.com/tapehead_Lab/status/2085304883847258179
how the fuck do you prompt something like this out of H3, looks professional
>>
>>109488459
Kijai distilled or og lightx2v?

which one do I use??
>>
>>109488354
lora will smooth out everything
it's also shit at innie pussy, which is sad as I don't like outies
>>
>>109488459
no shitty slowmo?
>>
>>109488495

damn thats crazy
and also you should see what @ingi_erlingsson used to do on wan... mad
>>
>>109488496
kijai's is the one converted to comfy

>>109488502
no
>>
>>109488459
Link?
>>
still rapes the audio like the other turbo loras
>>
>>109488495
>random jap bullshit
*yawn*
>>
>>109488514
nooooooo
>>
>>109488518
*yawn*



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.