[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: hg38yg2.mp4 (1.1 MB, 834x480)
1.1 MB
1.1 MB MP4
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109455104

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
I want to know how much copyright content the model knows. t2v test:

Hatsune Miku is standing and facing the camera. 0 to 2s: Miku is dancing. 2 to 8s: The character is having a fist fight with Vegeta from Dragonball Z in the style of Dragonball Z. 8 to 10s: Miku punches Vegeta far away sending him flying into a wall that makes a large explosion.

https://files.catbox.moe/cklzvm.mp4
>>
Is there seriously no way to do a latent preview wile H3 is genning? This was the greatest thing ever with WAN, because I could abort bad gens without having to go through the whole thing.
>>
>>109456162
It works for me
>>
>>109456162
i believe some people had it working in previous threads
>>
File: file.png (17 KB, 1259x188)
17 KB PNG
>>109456171
>>109456175
What the hell am I doing wrong?
>>
>>109456130
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109456130
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
OP is a shit stirrer, false flagger, same fagger, spammer, botter, fudder troll.
DO NOT INTERACT
>>
Why does /ldg/ make new threads before the bump and image limit while /lmg/ is cruising on page 8 and will do so for a few more hours because /g/ is a slow board?
The image limit is 150. There's 5 /ldg/ threads in the catalog and each has 50 images.
>>
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 63446.webm (1.21 MB, 448x256)
1.21 MB
1.21 MB WEBM
time to start figuring out the ejection prompt
>>
>>109456130
What the fuck happened to wan2.2 in the OP?
>>
>>109456178
I have it as taesd but I didn't have to change anything it kinda always worked.
>>
>>109456188
One anon wants to remove something in the OP that everyone else wants to stay thus he bakes these troll threads.
>>
>>109456188
Because the thread schizo keeps trying to hijack the OP because his dream was to get rich off of samefagging himself into relevance on 4chan. The links that he keeps removing from the OP prevent him from doing that.
>>
>>109456195
bruh no one use dat anymore unc :skull:
>>
>>109456195
Obsolete model got removed.
Need to do the same for Qwen and Chr-ma too.
>>
>chroma
>in 2000+26
kek
>>
>>109456178
You need VideoHelperSuite nodes installed
>>
Do Alibaba now need to release wan2.5 to the community asap?
>>
>>109456231
You mean Wan 3?
>>
>>109456188
>There's 5 /ldg/ threads in the catalog
you missed last year when >>109456130 would bake multiple at the same time just to stop two links at the end of op from staying. there was like a dozen of them before a janny woke up.
>>
>>109456231
This is better than Wan 2.5.
They need to release a newer model.
>>
>>109456223
what's the direct upgrade path for porn
>>
>>109456178
>>109456198
It looks like it's doing something in the subgraph, but it's only showing a frozen first frame. Tried both options.
>>109456224
Same issue.
>>
>>109456195
>What the fuck happened to wan2.2 in the OP?
schizo keeps hijacking this thread, next time he'll put his 35 stars C++ wrapper in there kek
>>
>>109456247
so the previous thread was made by a schizo? >>109455104
>>
>>109456231
all wans after 2.2 were barely better if at all, they need to release wan 4 to compete
>>
>>109456247
I removed Wan and Ltx from OP as a bit of housekeeping. I don't mind putting them back. I really don't care I just thought when I add a section to the OP I should take one out to avoid reaching character limit. FWIW I agree with anon's suggestion to take Qwen and Chroma out.
>>
File: file.png (10 KB, 1045x117)
10 KB PNG
>>109456224
>>109456244
Nevermind. Had to also enable this after installing. It's working now. Rentry folks, pls add this.
>>
wtf I didnt even mention rdr2 and it did arthurs voice but I didnt specify anything, I just used it in dialogue.

The man says "Arthur, I just got a new bike and i'm afraid all these so called doctors and lawyers will steal it. You need to protect my bike, my boy."

https://files.catbox.moe/fmy7ru.mp4
>>
>109456257
>falls flat on his face
>again
>>
Chr-ma should have been taken out months ago and we also could also have benefited from another quality control rentry to warn people about Chr-macord shilling here.
>>
ldg so juicy
always something habbenin :D
>>
>>109456195
>What the fuck happened to wan2.2 in the OP?
name one usecase that isnt covered by h3
>>
>>109456298
uuuummm it makes him feel special for it to stay in the OP
>>
File: MiniMax_H3_00077.mp4 (2.82 MB, 768x1376)
2.82 MB
2.82 MB MP4
h3 understand height difference, WAN would never.
>>
>>109456303
I look like this irl
>>
>>109456306
What was it like meeting Bayonetta?
>>
>>109456298
Explicit sex.
H3 is still limited for porn. Wan2.2 still has far more lora support.
LTX-2 is 100% obsolete now, and /ldg/ should hold an obituary.
>>
>>109456290
what is the upgrade to it for porn
>>
>>109456306
yea, i bet you like the short man
>>
>>109456315
>Realism
SNOFS lora for Klein, maybe BigASP 3 if it finishes RL and turns out good, Krea2 also has a growing NSFW catalog, just need a reliable method for consistent better realism
>Cartoon/Anime
Anima
>Animal fucking like its 'cord intended
Don't know, don't care
>>
File: MiniMax_H3_00078.mp4 (3.85 MB, 768x1376)
3.85 MB
3.85 MB MP4
>>
>>109456303
>married man taking selfies with a cosplayer
not very realistic
>>
>>109456330
yeah I'm talking about realism obviously. I had really bad results with all other models but last I checked was about a year ago, so I'll look into this, thanks
>>
I ran it with vram headroom 1 once yesterday and now it permanently stopped using 5 gigs of my vram??? speed is shit???? even with arg removed????? what should I do??????
>>
>>109456372
Quant? Total Vram? Verify how much VRAM is actually in use with an external tool instead of just looking at Comfy report.
>>
>>109456393
task manager shows 18.5 out of 24. the model itself is 20 and it used 23 before I tried the arg once
>>
seems like lower shift scale values are better if you want realistic camera shaking
>>
With Scale Image to Total Pixels in H3, should I be using nearest-exact or lanczos for the upscale method?
>>
took a while to warm up but I like this <Person> lark, it's much easier to define the subjects and then just call them instead of referencing them every time and hoping the damn thing remembers what's supposed to be goin on
>>
>>109456372
>>109456393
In process I discovered that by default the reference video is not resealed, so my hd reference resulted in 60it/s and after making the reference 256p it became 10it/s
>>
>>109456485
>resealed
rescaled
>>
File: MiniMax_H3_00083_R.mp4 (2.97 MB, 768x1376)
2.97 MB
2.97 MB MP4
>>
Can someone please explain to me how it all went wrong for LTX?
It was supposed to be the big new SOTA video model, the chosen one to dethrone wan. Despite numerous updates, it could never really take off, and has now been completely superseded by the new chinese model Minimax H3.

Where did the LTX developers go wrong?
>>
https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/

less degradation than easycache, but may be slightly slower, try testing both.

Use the Spectrum node in this order:

MiniMax H3 model loader
LoRA and other model patches
MiniMax H3 Sigma Shift
Spectrum Apply MiniMax H3
guider and sampler
>>
>>109456512
I get an error while trying to use this node, so I'm gonna guess it can't be used together with easy cache?
>>
File: file.png (77 KB, 542x466)
77 KB PNG
>minimax_h3_fl2va_bf16.safetensors
>minimax_h3_fl2va_int8_convrot.safetensors
>minimax_h3_fl2va_pruned_int8_convrot.safetensors
>minimax_h3_fl2va_pruned_fp8_scaled.safetensors
Do I need all those if I'm a 12gb vramlet?
>>
>>109456526
you only need
>minimax_h3_fl2va_pruned_int8_convrot.safetensors
>>
What's the difference between euler and res multistep samplers for minimax gens?
>>
File: file.png (166 KB, 1765x736)
166 KB PNG
>>109456512
I get longer gen time after adding Sigma shift and Spectrum. does the order matter?
>>
>>109456541
nice, thanks
I downloaded everything but ~380 GB space for something I'll use a few times a month seems excessive
>>
>>109456061
>15 second 540p videos in 6 minutes on 16gb/64gb
How many samples? Can you post a complete workflow so I can compare? I'm using sageattention, the kj node, and easycache and I still need >10 minutes for only 10 seconds on blackwell 6000 at 20 samples. I'm using the reference workflow. I must be doing something wrong.
>>
>>109456553
>for something I'll use a few times a month
i get it ;^)
>>
>>109456555
Resolution?
RTX Pro 6000 or just RTX 6000 ?
>>
File: thats_what_you_are.png (2.16 MB, 1920x1080)
2.16 MB PNG
that's what you are
https://youtu.be/rHeaoBAbBhI
https://suno.com/s/GRTExLgYEjLtodxk
>>
>>109456572
0.5mp so that should be anon's 540p
There's no non-pro blackwell 6000 card.
>>
>>109456555
Picrel workflow or drop catbox into comfy
https://files.catbox.moe/vlr1oh.mp4
>>
where
are
the fucking
loras
>>
might switch to qwen q4. don't want to need swap space enabled to use int8
>>
File: 1761208770400949.jpg (62 KB, 620x624)
62 KB JPG
>>109456130

Minimax following your prompt my ass.
I said in the prompt "NO TALKING" and yet she still talking gibberish.

Thats it, im going to go back to LTX. 10eros 1.5 is god tier and you people never uses it
>>
>>109456592
4 shekels have been deposited into your account
>>
>>109456590
What for? Just give it a ref concept, or several
>>
>>109456605
for t2v
>>
have third parties ever released a step distilled model before? i am worried that if minimax doesn't do it then no one will
>>
>>109456582
>0.5mp, 10 sec, 20 steps
>RTX pro 6000
something is wrong. my 5090 at 1mp, 10 sec, 20 steps takes about 8 mins
>>
>>109456618
no but light tricks have made 4 step lightning loras for third party models before
>>
>>109456586
Can you post it on other than catbox ? that website blocked in here

put it in here https://uguu.se/
>>
>>109456118
what did you notice? used same seed I assume??
>>
>>109456648
I dunno what that is, I'm not clicking it.
The full workflow is in the image. It's just the template workflow with the 3 nodes added at the top, literally nothing special
>>
>>109456661
newfag
>>
Guys, I need help. I'm using Forge Neo Classic and whenever I use Anima, as the image is rendering, everything works but by the time it's finished, it just gives me a blank gray canvas.

For example, if I prompt "apple on a table, indoors," the image will be that as it renders and when it finishes, it doesn't give me what I prompt. Just a blank colored canvas. Idk wtf I'm doing wrong.
>>
I fucking hate Catbox so much
>>
File: final_design1.png (1.97 MB, 1024x1536)
1.97 MB PNG
>>
>have H3 make a girl take her top off
>she has brown nipples
>have it make her say a sentence in english and take her top off
>she now has pink nipples
lmao that's pretty funny
>>
The camera zooms far out from the man in the cowboy hat, and the man starts running forward, as Hatsune Miku fires her revolver at the man. the man falls down and then "wasted" pops up on the screen in the style of GTA 5.

https://files.catbox.moe/v7sr97.mp4
>>
I got face drift on Minimax
>>
>>109456592
lets see some of that godtier proof
>>
The camera zooms far out from the man in the cowboy hat, and the man starts running forward, as Hatsune Miku fires her revolver at the man. the man falls down and then "wasted" pops up in the center of the screen in the style of GTA 5, as Hatsune Miku dances in celebration.

ok, that worked.

https://files.catbox.moe/4oe75e.mp4
>>
>>109456675
A-anybody...? ;_;
>>
>non_diegetic_music: N/A
just discovered this at the end of a prompt will remove any music you dont specify.
>>
File: MiniMax_H3_00087.mp4 (3.91 MB, 1760x608)
3.91 MB
3.91 MB MP4
>>
File: MiniMax_H3_00054_.webm (675 KB, 720x816)
675 KB
675 KB WEBM
>>
>>109456675
>>109456732
>I'm using Forge Neo Classic
>>
>>109456586
Where is your Patch Sage Attention node?
>>
I'm hungry. https://streamable.com/tocn2f
>>
>>109456750
I don't want to use Comfy because I hate cluttered, node-based UIs.
>>
Beta sampler looks considerably better with a ton less smearing "filament" motion problem, the problem is that you lose some micro details and stuff like tattoo accuracy when using the ref model, also has a lot more color saturation for some reason.
Holy shit this model has so much potential is beyond mind boggling.
The next couple of months gonna be insane.
>>
any fix for the easycache audio degrading yet?
>>
>>109456765
Beta sigma*
>>
File: full.png (69 KB, 714x574)
69 KB PNG
>>109456765
>tattoo
>>
>>109456766
try the spectrum node thing maybe
>>
>>109456757
I just use the --use-sage-attention flag instead
>>
The average Diffusion user lacks self-confidence and has a small dick.
You're free - you can put all your creativity to use. And all they do is produce attention-grabbing TikTok trash that lasts 5 seconds.
Anyone who doesn't do that is practically a higher form of life - Gooners included.

It's as if a nigga had solved a 20-piece kids' puzzle and was driving through his college neighborhood in his 300-ps BMW, honking the horn to let everyone know he was the one who did it.
>>
>>109456740
death to choppy animation. Bring back smooth anime animation
>>
>>109456586
>160 seconds
I guess the ref2va model is just that much slower. Or it depends on the number of references. I'll have to test more. Thanks.
The other reply isn't me. It took me a while to download the model.
>>
>>109456330
Wait, so Chroma is good for bestiality?
>>
Can Minimax append audio to existing silent videos?
>>
>>109456795
>I guess the ref2va model is just that much slower
it is. my generation times got cut in half when i switched
>>
>>109456800
yeah you can tell it to not change anything about the video and then describe the audio you want
it'll still degrade the video tho so just use the original video source and add the generated audio to it
>>
>>109456798
very good
>>
What's better, EasyCache or Spectrum?
>>
I don't see H3 sigma shift in the official Comfy workflow. WTF is that?
>>
t2v, trying easycache + the sigma shift node, 0.3mp/10s

Arthur from the game Red Dead Redemption 2 starts running forward, as Hatsune Miku wearing a cowboy hat says "you came to the wrong town, baka!", and then fires her revolver at the man. the man falls to the ground, and then Miku runs away very fast.

https://files.catbox.moe/7p3o1v.mp4
>>
How should Sigma Shift be configured? Should the default of 12.0 for video and 3.0 for audio be changed?
>>
i think we need separate generals for image diffusion and video diffusion
>>
Whatever the fuck comfy did with managing h3 seems to be good. God bless an actual competent dev, fucking love blonde fennecs
>>
>>109456546
Not specifically for minimax but:
Euler is the simplest sampler out there, but it still just works. Often eclipsed by other samplers in terms of quality BUT it's super fucking flexible if you have sub-optimal step count or sigma distribution, it can still provide decent quality in cases where other samplers shit themselves. If nothing else it's a good starting point.
Res multistep, I know that it is is a minor upgrade over dpmpp 2m. I don't recall the precise details of dpmpp 2m but it's one of the great samplers.
>>
KINO, i2v worked even better.

https://files.catbox.moe/qxr89i.mp4
>>
>>109456675
Browser issue? Do you have aggressive privacy settings, or privacy extensions like canvas blocker?
Disable enhanced protections or whatever equivalent your browser has for 127.0.0.1
>>
>>109456840
seems to be good for me at 12 and 3 so far. easycache is a bigger speed boost than the spectrum node thing.
>>
File: 1777015060208822.jpg (440 KB, 2551x2480)
440 KB JPG
Its just good at making slop.
Not porn unfortunately. Yes even if its uncensored
>>
File: h1c95vcd4bhh1.jpg (213 KB, 1170x1871)
213 KB JPG
You did fill in the MiniMax Formal Authorization form before deploying H3 to your GPU's vram, right?
>>
>>109456868
loras should fix it but where are the fucking loras
>>
File: MiniMax_H3_00089.mp4 (1.39 MB, 960x1088)
1.39 MB
1.39 MB MP4
>>
okay, more lore accurate for miku.

https://files.catbox.moe/ntctip.mp4
>>
>>109456869
>Organizations
what about individuals tho?
>>
>>109456872
you use the reference model cause you can plug in literally anything.

https://files.catbox.moe/gyjo9a.mp4
>>
>>109456869
ltx doesn't have this problem
>>
File: 1768983670767703.jpg (38 KB, 708x708)
38 KB JPG
My Point still stands.
Slops. Not Porns
>>
>>109456764
Use App mode then.. There's no excuses anymore
>>
https://huggingface.co/lilcheaty/MiniMax-H3-NVFP4

Is it ok to use the 12GB pruned nvfp4 model? Is there noticeable quality degradation?
The description says:
>pruned_nvfp4 is doubly quantized (bf16 -> int8_convrot by Comfy-Org -> NVFP4 here). Error from both passes compounds. It held up in testing, but that is a real caveat.
>>
File: 1767210439350221.webm (576 KB, 220x516)
576 KB
576 KB WEBM
Im into Live2d and i dont like the direction Minimax is going.

LTX is suprisingly well at doing Live2d like animations
>>
>>109456882
Same applies.
>>
>>109456890
>Slops. Not Porns
>>109456907
>post slops
>>
i still haven't deleted my LTX folder. it feels like deleting pictures of an old friend who died, even though i barely ever used it
>>
>>109456925
Keep the gens delete the models
>>
Getting the speed boost from Sigma+Spectrum.
>>
File: MiniMax_H3_00091.mp4 (1.63 MB, 960x1088)
1.63 MB
1.63 MB MP4
>>
>>109456906
NVFP4 is a 4bit quant, the highest quality 4bit quant out there by far, but it's still a 4bit quant at its core. So yes there is a noticeable delta between bf16 and it. As for whether it is still good or not I don't know, I never used it.
It's also only fast if you have 5000 series GPU.
>>
>>109456925
I still keep the shitty pony loras, most I never used, on an external drive. Deleting loras feel irrationally tough to do for me too.
>>
>>109456950
I just cleared up hundreds of GBs worth of space from all the pony era shit I still had lying around
I had a whole bunch of personal merges that produced really unique outputs because I merged in blocks from regular SDXL models, it hurt a little to let go of them ngl
>>
>>109456938
> elbows
pic from >>109456770
>>
>>109456925
i'm still keeping ltx since i still haven't figured out the correct prompting for h3 yet
>>
kek the chinese trained this on a bunch of games too.

Arthur from the game Red Dead Redemption 2 is riding a horse down the road, as Hatsune Miku wearing a cowboy hat and also riding a horse says "Hey Arthur, we gotta rob the bank so I can buy more green onions!". Arthur says "Will do Miku, will do.".

https://files.catbox.moe/vgsb8e.mp4
>>
File: MiniMax_H3_00092.mp4 (3.57 MB, 896x1152)
3.57 MB
3.57 MB MP4
>>
Did we not agree that Easy Cache and its variants were massive copes that heavily degraded quality for other video models? Is anything actually changed with H3 or are we in the copium phase again?
>>
>>109456974
I don't know if I should spoonfeed a retard like you who can't prompt the easiest to prompt model in existence but:
Go to the minimaxh3 huggingface open the prompting guide doc for the model you gonna use, write a shitty simple as shit prompt in natural language and which subjects there is and what is it about in very simple terms.
Feed your prompt + the guide to any llm like deepseek 4 flash or qwen 3.6 27b if you have a local setup.
Tell the llm to fit your prompt for the proper minimaxh3 way using the guide and improve it with details etc.
Enjoy a insanely high quality prompt adherence, the model can quite literally do almost anything.
>>
>>109457009
KJ is working on a turbo lora, should be out in a couple days
>>
>>109456939
> the highest quality 4bit quant out there by far
on par with int4 convrot
>>
>>109457014
Bless his ass
>>
File: MiniMax_H3_00021_na.mp4 (1.23 MB, 960x544)
1.23 MB
1.23 MB MP4
Has anyone been able to use a reference image for style transfer? I've had no luck convincing it to use the reference's style.
>>
>>109457019
Source?
>>
>>109457026
make sure you use the reference model for the reference workflow specifically or it wont work properly, the other one will work but it doesnt transfer styles as effectively
>>
>>109457014
sizeable if factual
>>
File: 1463179529127.png (16 KB, 840x773)
16 KB PNG
Does anyone else get a bit emotional following the release of Minimax H3? It almost brings a tear to my eye knowing we still get local models that are this good and such a big upgrade on what we had before.
>>
>>109457012
>Enjoy a insanely high quality prompt adherence, the model can quite literally do almost anything.
it doesn't follow the audio part of my prompt. i read the official guide and it didn't have much information for sound effects
>>
>generated normally on minmax yesterday
>getting out of memory today
????
>>
>>109457055
take one of your workflows from yesterday and run it again
>>
File: 1755002130234059.webm (1.78 MB, 768x1280)
1.78 MB
1.78 MB WEBM
>>109456925
Keep 10eros v1.4 or v.1.5 You dont even need 80% of LTX Loras. I still have faith in LTX for making anime sex.
>>
>>109457055
You didn't fill in the Minimax H3 license form and now your vram has been confiscated
>>
Just a reminder LTX 2.5 will out on december :)
>>
File: 1500124111077.jpg (8 KB, 200x200)
8 KB JPG
I just deleted my AI mommy tiktok profile, all the pics, my info folder of prompts and commands, workflows etc because I was spending too long generating images and chatting to old horny men online. Now all I want to do is make another one lel, why is this so fucking addicting? The amount of attention I get online for simply existing as a woman is crazy.
>>
>>109457087
proof?
>>
>>109457089
I can imagine.
>>
>>109457089
>not even scamming them
>>
>>109456740
https://files.catbox.moe/glbina.mp4
>>
>>109457029
https://old.reddit.com/r/StableDiffusion/comments/1uimp1j/so_is_int8convrot_the_new_hot_thing/ouuyc12/
>>
File: MiniMax_H3_00025_na.mp4 (2.15 MB, 960x544)
2.15 MB
2.15 MB MP4
>>109457037
I'm using the reference model and example workflow, but it seems to strongly latch onto what it assumes is the appropriate style for the content rather than how the reference was actually rendered.

If using I2V instead, it usually quickly morphs away from the style of the input frame unless it already matches what it's expecting.
>>
>>109457107
Hawt
>>
File: 364455.webm (990 KB, 448x256)
990 KB
990 KB WEBM
>>
Which of these are meant to be the better choice?
>>
>>109457126
don't worry about it just download a gguf
>>
>>109456939
>>109457019
i got that quant since my card will not be able to handle higher...and handful of generation so far with default workflow where some guy is supposed to be running on the roofs has generated talking cats (indian accent no joke), x2 gens of two old people conversing, one vid of three asians making noodles, one landscape shot and just now another gen of two old people speaking in some madeup language.

wtf is this. could be that quant encoder.
>>
>>109457110
Are you some stupid LLM bot searching keywords?
NOTHING here says "[nvfp4 is] on par with int4 convrot"
>>
>>109457100
Surprisingly only 1 indian (who was also the most unhinged of course), the majority is just horny old white guys.

>>109457104
I have just enough empathy not to do that as easy as it would be, but not enough that I'm uncomfortable deceiving them into jacking off to computer generated imagery for my own enjoyment.
>>
>>109457126
always take the rot
>>
>>109457146
As long as you're all enjoying yourselves.
>>
>>109457112
just use the tags then prompt as if it was t2v

like, the girl in <Picture 1> is doing (whatever)

with picture 1 being your asuka source
>>
>>109457143
> as a W4A4 method, ConvRot is slightly behind Nunchaku’s SVDQuant in terms of accuracy
>>
>>109457126
The bottom one has significantly higher quality than top and accelerated on more hardware.
FP8 is deprecated.
>>
File: MiniMax_H3_00094_R.mp4 (3.64 MB, 896x1152)
3.64 MB
3.64 MB MP4
>>
>>109457129
>no convrot
Into the garbage it goes.
>>
>>109457160
*image 1

if you want a specific style you could use a second image and maybe say "in the style of image 2" or whatever.
>>
>>109457166
> FP8 is deprecated
it's for training
>>
File: brainlet soyjak.png (16 KB, 600x800)
16 KB PNG
>>109457164
?????????
I will humor you with a serious response if you can articulate how the two statements are supposed to be related.
>>
Do you guys use the prompt structure in your prompts? Like do you write yours like the guide suggests?:

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description:

overall_soundscape:

non_diegetic_music: N/A
>>
>>109457173
pretty good except for the tail.
>>
>>109457178
>*image 1
both picture and image works from my experience
>>
>>109457159
Until I disappear forever in the middle of the night lol
>>
>>109457184
Alright.
*FP8 is deprecated for inference.
>>
>>109457190
yes, I don't do it but gemma-chan does it for me when I tell her what I want to prompt
>>
>>109457178
Yeah, those are the default instructions, but it doesn't work. It knows what kind of media Asuka is from and really dislikes the idea of changing it to the unconventional one indicated by the reference.
>>
>>109457203
Do you load Gemma alongside H3? Or do you unload her during video gen?
>>
>>109457198
do they know it's AI or are you actually fishing
>>
File: 1756918313175861.jpg (27 KB, 985x156)
27 KB JPG
Which one should i delete
>>
>>109457213
I keep her loaded, I have a lot of vram and ram to spare because I built it for bigger llms. I use the 26BA4B version because it's faster and has never messed up any complex prompts yet; you could run it entirely in ram at an acceptable speed if you want
>>
>>109457220
all of them except flux
>>
>>109457198
Life goes on. I think I've met multiple people like you.
>>109457219
You get addicted to the attention, it's not about the images.
>>
>>
>>109457229
no I mean what do they get. do they actually think they're talking to a real girl, or do you mark it as "this is AI content" or whatever
>>
File: 1646373884548.png (558 KB, 1106x1012)
558 KB PNG
how much longer until minimax will be able to generate pusy?
>>
>>109457190
I just tewll my gemma 31B to write a prompt and give her a guides from huggingface to input. works each time.
>>
>>109457220
transfer to another HDD
>>
spectrum and sage attention work together?
>>
What is sigma shift for?
>>
>>109457219
>>109457232
They don't know, my whole schtick is generating images that look like shit quality front camera smartphone selfies. AI does 90% of the work, I do 1-2 min of editing and boom. No AI content marking. They always want to end up meeting me IRL.

>>109457229
I guarantee there's a lot out there, the obvious fake AI keep the heat off for now.
>>
>>109457251
sigma balls
>>
anime at 0.3 works just fine, using the spectrum/sigma thing.

https://files.catbox.moe/w90wa2.mp4
>>
>>109457232
I don't do shit. I just get bombarded with stuff that's clearly made specifically for me, and no, I don't think they're real girls. I feel pretty much the same about it as I would if a buddy were sending me porn pics. Some of it is weirdly emotional though.
>>
>>109457240
and what if it's an nsfw prompt?
>>
>>109456274
When I enable this and have VideoHelperSuite installed it says it's not supported for H3. Do you need something else too?
>>
File: 00460-2653269664.png (2.22 MB, 1536x1024)
2.22 MB PNG
>>
>>109457234
Wait... H3 don't do i2v ?
>>
File: keekekekekkkekek.png (1.17 MB, 864x1184)
1.17 MB PNG
>gay ahh ngas catfishing unc
pfftt hahah oh no no no wat da heck yxll doing bruh JSID already keeeeeeeeeeeeeeeeeek
>>
>>109457271
gemma is mostly uncensored and piss easy to jailbreak if you do get refusals
>>
>>109457220
all of them except ltx
>>
File: preview-broken.png (56 KB, 633x391)
56 KB PNG
anyone know why my preview is broken? specifically on the subgraph. it works fine in the sampler
>>
>>109457278
no, it does.
when i say generate pusy, i mean when there is no visible pusy in the input image, and the model has to generate one.
i've tried and the pusy is wan2.2 tier (thankfully not wan2.1...). need loras.
>>
File: mischievous.png (142 KB, 220x220)
142 KB PNG
>>109457280
>>
>>>/wsg/6207798
>>
>>109457284
Subgraphs always have something broken with them
>>
>>109457287
Yeah similar experience with wan2.2. Pure horror show.
I had reasonable results with mask inpaints.
>>
>>109457284
just unpack the subgraph please
>>
>>109457291
So wait, you're not using an input image? The model knows who George is and what he sounds like?
>>
File: 0804195819046-i2eHPAojcbX.png (1.44 MB, 2503x1611)
1.44 MB PNG
>>109457284
just move them out of the sub-graph instead of waiting for fix.
>>
>>109457262
and this time, misty was not upset: still i2v

https://files.catbox.moe/hznq6i.mp4
>>
None of you fuckers know what a real vagina looks like anyways.
>>
>>109457186
logical chains are too hard for you I see
>>
>>109457251
higher shift > more high sigmas > better performance for anything newer than SD3
But too much shift results in blur or deformities in low level detail
It makes the timestep distribution more optimal, to a point.
>>
>>109457324
Bullshit I've studied the pics and videos very hard.
>>
>>109457329
based examiner
>>
>>109457327
Performance as in speed or as in quality?
What are sigmas and why are high sigmas good?
>>
>>109457319
Yeah. It knows all the Seinfeld main cast
>>
Who the fuck said that Spectrum thing same shit as EasyCache? Spectrum doesn't even make the output worse. It looks ~ the same.
>>
How is Minimax's boob physics so much better than Wan2.2 with loras already? Why did the chinks train it on boob jiggle?
>>
>>109457342
Blue board, anon
>>
Fuck i posted on the wrong board
>>
>>109457350
It's a big model trained on lots of data, and boobs naturally jiggle. To make it NOT have good boob physics you have to develop and elaborate censorship regime. Minimax decided to just train on data and not do that.
>>
>>109457358
See you in three days king
>>
>>109457350
they clearly used a lot of synthetic 3d render data for physics prompt pairs. half my stuff look like video game footage
>>
>Error: You must wait longer before deleting this post.
FUCK
>>
how can i get better quality with minimax t2v? the animations are nice and all, but the video quality is too low
>>
File: MiniMax_H3_00097.mp4 (1.96 MB, 992x1056)
1.96 MB
1.96 MB MP4
i give up, i couldn't make the hand go in between the girl's cleavage.
>>
>>109457364
why couldn't they fucking do it for the holes too is my question. retarded
>>
>>109457342
Migu !
>>
>>109457342
uhhhhh...
...!!!!
>>
File: 35uy.png (269 KB, 432x621)
269 KB PNG
>>109457370
>>
File: incomprehensibly retarded.png (703 KB, 1024x1024)
703 KB PNG
>>109457326
You blew your chance for a serious response, sorry tard-kun
>>
File: 1772961089715389.jpg (203 KB, 474x444)
203 KB JPG
>>109457370
>>
I got into media generation with Hailuo AI in June-July 2025. Back then I was just using their cloud model for 6 second gens on the free plan (across multiple accounts). This was before switching to Wan2.1 (then 2.2 when it dropped)

How does MiniMax H3 compare to Hailuo back then?
>>
>>109457370
When that happens I usually can no longer delete it at all. Anon you should be using 4chan X which shows you the deletion cooldowns.
>>
>h3 can do shirt lifts ootb
this is going to get ai regulated as fuck isn't it? maybe this was china's plan
>>
>>109457364
that's a bit strange because base wan2.2 generates nipples but no jiggle.
one time wan2.2 even generated me a pussy that wasn't in the input image, and this was without loras.
>>
the fact that the prompt processing step takes like 0.1 milisecond and its prompt following is miles better than ltx and definitely better than wan makes me wonder what the fuck those retards were doing before
>>
>>109457346
>Performance as in speed or as in quality?
Performance as in quality
>What are sigmas
Basically the amount of noise injected to and removed from image at each step, contributing to how the model diffuses pure noise into a coherent image.
>why are high sigmas good
High sigmas correspond to structurally difficult problems the model needs to solve. More time spent on them = better image. But you still get diminishing returns after a point and if you steal too much time from lower sigmas you get performance regressions as I mentioned because you still need some amount of lower sigmas for a coherent image.
If you need to know more ask the chatbot of your choice.
>>
>>109457403
No. Over the past 12 hours of using H3 I have come to realize that China's actual plan was to destroy our birth rates, and it's going to work.
>>
What was the bigger upgrade, Wan2.1>2.2 or Wan2.2>MinimaxH3?
>>
well, it can do dancing asian girls, i guess new model is alright
https://files.catbox.moe/1qh8t7.mp4
>>
>>109457423
But China's own birthrates are on the decline?
China hates India, but India's birthrates keep going up, and they're beginning to heavily saturate the land China wants to conquer, like Australia.
>>
>>109457342
I look like this irl
>>
>>109457414
it's the fault of whoever made the text encoder for ltx
>>
>>109457426
ppl will say minimax, but the actual answer is it's close. ppl forget how horrible 2.1 was, especially with censorship.
>>
>>109457426
>Wan2.2>MinimaxH3
Obviously. Hunyuan video was a massive improvement over what came before it (cog video, mochi, etc.) it's been incremental improvements since then until H3 which is an enormous leap
>>
>>>/b/952240408
>>
reminder, timestamps work and can give you exactly what you want in a longer prompt, like 10s!

i2v: 0 to 3s: the man is standing in a doorway with a white jacket. 4 to 6s: the man is driving a car in a cyberpunk neon city at night. 7 to 10s: the man is standing, looking up at a giant holographic Hatsune Miku in a cyberpunk city at night.

https://files.catbox.moe/1tvqbd.mp4
>>
>>109457475
very good to know thank you
>>
All those LORA and Edited Checkpoint makers are getting cucked by Minimax license
>>
I miss the krea gens desu honest
>>
reference model test:

Use Image 1 as the exact character identity reference for JC Denton, keeping his signature black sunglasses and stoic features. Extract the vocal timbre, frequency, and speech pattern from Audio 1 to build a voice clone. JC Denton says "I am here to generate more Hatsune Miku videos with this AI technology, from the illuminati. We need to protect open source from faggots like Sam Altman and Closed AI. An evil corporation.". overall_soundscape: Cold, flat, monotone spoken dialogue with completely deadpan delivery. The voice print has an electronic, slightly raspy, low-intonation pitch mirroring the robotic timbre of Audio 1. No vocal inflections or emotional excitement.non_diegetic_music: N/A"

so it can clone voices and do stuff with it too. image source was default JC.

https://files.catbox.moe/o5z8qh.mp4
>>
>>109457495
I mean nobody but third worlders can currently "legally" make anything with it desu
>>
>>109457089
>deleting data
KEEEEEEEEEEEEEK, truly only something a tranime agp faggot does
>>
>>109456675
Do you use AMD? I recall something that having the VAE at fp32 should make it work.
>>
>>109457220
all of them
>>
File: 202608030645.png (270 KB, 3569x432)
270 KB PNG
reminder of this factnvke
>>
Can everyone just calm down? H3 is nice and all but it's not THAT good. Can we stop pretending local is suddenly saved and just be realistic for a moment?
>>
>>109457527
There's actually been a lot of kino
>>
>>109457400
We have a large number of posters not even remotely educated in tech in this thread
>>
File: MiniMax_H3n_00061.mp4 (1.68 MB, 1376x768)
1.68 MB
1.68 MB MP4
>>
I booted into linux, asked the deepseek flash sir to do the needful and it configured a comfy setup that doubles h3 performance compared to my earlier win11 attempts at the very least.
What a great time to be retarded, didn't even have to ask anyone to spoonfeed me
>>
>>109457426
H3 is the same kind of leap for video that ZiT was for image generation if we were still stuck on SD 1.5. That's how far behind we've been in video generation
>>
Some one had this issue?
The size of tensor a (75480) must match the size of tensor b (1824768) at non-singleton dimension 2
I using 124 frames and 32 multiple for the resolution so is not that,
>>
>>109457434
>India's birthrates keep going up
They're not. Some states are already below replacement. And deluge of easily generated bobs and vegana is going to hit them the hardest out of any demographic.
>>
File: i_00018_.png (1.18 MB, 720x1280)
1.18 MB PNG
>>109457546
top kek
>>
>>109457555
minimax h3 model btw
>>
>>109457535
the model is a huge jump worthy of most of the hype, although its indeed a bit too much, its from a lot of newfags and vramlets who either werent there, or just didnt have 24+64 rig to gen at max quality wan 2.2 with a heavily optimized non-raped-quality workflow, and from people who are impressed by the simplest of meme videos instead of incredibly coherent, detailed and candid ones. same reason krea 2 was posted as much as it was, a lot are just failed normie tier retards that get excited at their cartoon character / celebrity obsession already being in knowledge of the model and don't care about the quality of the gen itself as much,
>>
Bros, is it time for me to delete all my wan models and loras? it's been a good ride, but I can't see any usecase to have them taking up storage on my drive. maybe nsfw? i don't know how long it will be until minimax can do proper nsfw.
>>
haha

you can be VERY specific in t2v/i2v prompts, here's another good example of timestamping to get the desired result.

0 to 3s: the man with glasses holds up a black remote. 4 to 6s: the dog on the left glows with a blue electric aura and says "not this time, faggot". 7 to 10s: the dog fires a thunderbolt at the man with glasses, causing him to get electrified and shake rapidly, with smoke emitting from his body.

https://files.catbox.moe/wvjic8.mp4
>>
>>109457573
why? archive it into your 16tb hard drive
>>
>>109457575
why does kaya sound like a man
>>
>>109457583
intense shock training and the will to surpass dog limits
>>
>>109457580
Anon, you might be a hoarder, but I do not see a valid reason to keep garbage on my hard drive.
>>
File: 1770633891655559.jpg (87 KB, 675x1200)
87 KB JPG
None of you are actually make a good stuff from this. Its just sloppa after sloppa.
>>
>>109457592
the fun is that you can make anything and you pay scam altman $0 to make it.
>>
Please stop posting still images in my video general, thank you
>>
>>109457591
why did you ask the question then, nigga?
>>
Blackwell user here.

Should I use `fl2va_pruned_nvfp4_convrot_int8` over `fl2va_pruned_int8_convrot`.
The nvfp4 version is 19.6GB, the non-nvfp4 is 20.4GB.
>>
>>109457592
what does that anime picture have to do with your post?
>>
lmao I tried to get the dog to transform, not quite and expected given there is no ssj dog in dbz.

https://files.catbox.moe/1r6ww9.mp4
>>
does H3 have the same problem as wan with undesired mouth flapping/talking?
>>
>>109457597
yeah but its still bad for sex.
Handjob and Blowjob is surprisingly fine though. but not 10eros v1.4 level of good.
>>
>>109457536
Like what?
>>
tran
be quick
or I will bake with the black legend
and remove your links
again
>>
>>109457342
Nice
>>
>>109457630
Nope
>>
>>109457592
I do not "make", I do not design, I do not create, I SLOP. And if I don't love it, I don't prompt.
>>
>>109457629
Amazing.
>>
>>109457662
>>109457662
>>
https://files.catbox.moe/ea4e14.mp4
>>
FRESH
>>109457666
>>109457666
>>109457666
>>
>>109457629
>inside of me, there are two saiyans
>>
>>109457664
BASED
>>109457669
You lost trani
>>
>>109457272
I found my problem. It was ComfyUI-bleh being installed silently hijacking the preview prep.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.