[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1775278511210573.mp4 (3.57 MB, 1280x498)
3.57 MB
3.57 MB MP4
Discussion and Development of Local Image, Video, and Music Models

collage edition

Previous:>>109609572

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Blessed thread of frenship
>>
that was messy but I guess we got there in the end....hopefully
>>
File: image.png (254 KB, 707x660)
254 KB PNG
>>
>>109611941
Based
Fuck the lolcow wrapper "developer"
>>
File: image.png (27 KB, 375x146)
27 KB PNG
>>109611959
Yes Julien, there's absolutely no reason why you got physically banned from the ComfyOrg offices. You are a saint, an angel even.
>>
I would like to remind you that this is the most important gen of the year.

https://d.uguu.se/zSeEyUlg.webm
>>
>>109611965
> Julien
So obsessed.
>>
File: MiniMax-H3-00092.mp4 (3.54 MB, 864x480)
3.54 MB
3.54 MB MP4
>>
>>109611941
you are a ban evader so we should probably use a different non schizo bake
>>
>>109611941
thanks for the collage bake
>>
>>109611993
I'm not your boogeyman
>>
>>109611941
>video collage
>rentries of frenship
Blessed bread indeed, thank you anon
>>
File: 1748351684074420.png (217 KB, 716x659)
217 KB PNG
why all this shitty discussion? you're supposed to be sharing gens and shit. we already have good image and video models. you guys are really api fags...
>>
>>109611941
Arigatou anon-kun
>>
>>109612042
you want a melty?
>>
File: MiniMax_H3_00363_.webm (2.15 MB, 928x672)
2.15 MB
2.15 MB WEBM
>>
>>109612078
Best Zelda
>>
>>109612078
blows my mind how much better this is than wan
>>
File: H3_00375_.webm (1.62 MB, 928x672)
1.62 MB
1.62 MB WEBM
>>
>>109612122
kek
>>
https://www.youtube.com/watch?v=QDK1KxsBYzM

chemo
>>
>>109612130
why is no one animating this
>>
https://www.youtube.com/watch?v=MzV_frMdecg
>>
File: gemmaprompt2.png (383 KB, 1774x1348)
383 KB PNG
Really digging my new GemmaPrompt fork, R2V prompts are so much more consistent now. Using Qwen3.8 Abliterated.

Soon I'll be able to craft an actual story.
>>
>>109612137
we aren't mentally obsessed troons
>>
>>109612137
Because Julien is xer and only xer.
>>
Can i request some animated juliens? It's on topic really
>>
>>109612130
who? there is no incentive to animoot just let me get high man.
>>
>>109612151
What kind of sauce do you need for local Qwen3.8? Mind sharing the repo for this?
>>
>>109612228
sauce? I just select it from the model selector. this is possible even in the old abandoned version. give it a shot: https://github.com/whp199/GemmaPrompt
Pretty sure you can choose a model in all the comfy promptwriter meme nodes and extensions too.
Id share mine but i dont have a github account and i'll get accused of spreading a virus if i share it as a zip
>>
File: 1759413369504217.png (1 KB, 196x92)
1 KB PNG
FFFFFFFFFFFFFFFFFFUCKING DYNAMIC VRAAAAAAAAAAAAAAAAAAM
>>
>>109612254
Oh sorry, I mean how much VRAM to run it at an ok quant. I'll ask my toaster.
>>
>>109611941
ty for bake.
>>
Are realistic gens with krea2 supposed to vary so much in style depending on the prompt? Some prompts will give me decent looking gens, but others will literally look like fake slop. What lora do you guys use?
>>
File: 290384567324.png (70 KB, 1539x387)
70 KB PNG
Will my PC explode?
>>
>>109612374
No, only burn.
>>
>>109612345
I have an RTX 5070 Ti an 64GB dram and the Q4_K_S gguf runs fine. Decent pace at 8.14 t/s.
Gemma4 31B and Glimmer 30B are far slower.
>>
>>109612328
To be fair, windows just chokes out the driver
>>
>>109612398
but only when dynamic vram is enabled, lmao
>>
>>109612386
why?
why is this one file unsafe lmao?
>>
>>109612400
maybe ani was right all along
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/performance.md
>>
>>
>>109612425
I heard on /adg/ he able to decently gen videos...... FROM HARD DISK

The fuck is wrong with ComfyUI. With every update its getting worse
>>
File: 27839456234.png (40 KB, 732x661)
40 KB PNG
>>109610327
>>109610351
Hi im the anon from last thread.
So which ones do I really need vs which ones can I skip? Or is this a true "You need them all man" moment?
>>
File: 1762122277446493.png (1005 KB, 2419x1105)
1005 KB PNG
Comfy is the most disgusting program ever made.
>>
>>109612443
>my name is not important...
>>
Are you just going to talk to yourself?
>>
Would kill myself if I was still a windozefag LMAO
>>
>>109612500
i have friends to play slop with and dont want to dualboot multiple times a week
>>
>>109612463
>install kijai nodes
>put get/set on shit that has to be reused often
>hope the execution order of the wf doesnt break
>????
>profit
>>
File: Larry.mp4 (2.65 MB, 1000x768)
2.65 MB
2.65 MB MP4
>>
>>109612500
True
The only reason to run windows is if you want to play games with kernel level anticheat (literal malware btw) or if your workplace forces you to run windows.
I have yet to find a single game (except kernel leven anticheat), setup or program which i couldnt run on linux using wine / proton / winboat
>>
>>109612443
why did you make two catjacks though?
>>
>>109612443
kino
>>
>>109612508
>i have friends
everyone point and laugh at this fag
>>
>>109612459
minimum to get going is one of the reference models from the first section (ref2va), one of the text encoders, the ref2v turbo lora, and both VAE files.
"int8_convrot" options are usually preferred, but for the text encoder at least you probably want to use the nvfp4 version
>>
>>109612459
you need diffusion models, a fl2va for straight vide, ref2 for reference videos. probably convrot or pruned.
you need a text encoder.
turbo lora is optional but you probably want one.
both of the vaes.
>>
>>109612374
>safetensors
>is unsafe
uhh tensorbros?? i thought we left the pickle rick era??
>>
>>109612443
kekd
>>
>>109612519
even better, get the anything anywhere node.
>>
File: 1925376814587.gif (1.98 MB, 480x270)
1.98 MB GIF
>>109612586
I am being genuine in that ive downloaded alot of safetensors off of civitai and huggingface and yet ive actually NEVER seen that warning.

I want to assume its safe, because literally everyone has clearly downloaded it. On the otherhand, computer aids.
>>
>download .safetensor
>not safe.
>>
>made it into the OP
I'm a little happier now. https://files.catbox.moe/jkjz9x.webm
>>
>>109612459
For the text encored I recommend using the biggest that doesn't cause OOM. It gets loaded before the actual video generation starts and then unloaded. The impact on the total generation time is small. In exchange, you get considerably better prompt adherence.
>>
Am I dumb or why does my latent upscale workflow make all the lines wobbly (like they're hand-drawn) in the upscaled result
>>
>>109612254
>week old
>abandoned
wut
>>
>>109612671
latent upscaling requires higher denoising in the 2nd pass than pixel space upscaling does or else that happens
>>
>>109612687
Just a normal day in vibecode land.
>>
>>109612687
>no new feature or security commits in DAYS
It's abandonware bro.
>>
>>109612463
I hope one day someone will save us from this piece of shit (or AI gets good enough to do it by itself).
>>
>>109612671
I would not recommend latent upscaling on any model that is older than SDXL
>>
>>109612706
iirc he mentioned in /lmg/ that he was waiting for his claude quota to reset. Never used it before so no idea how long that takes.
>>
If your gen looks like plastic slop. please, have better standards and fix your workflow.
>>
>>109612691
I am using around 0.35 - 0.4 for the second ksampler, should I go even higher? Also should I match the sampler and cfg setting for both ksamplers or does it not matter?

>>109612720
I'm trying it with anima
>>
thanks for bakering
>>
>>109612754
Yeah. Dit needs a lot higher denoise than unets
>>
>>109612754
anima can already gen at very high resolution. you should try just genning at your target resolution.
>>
>>109612754
in general i had weirdly bad results trying to upscale with anima no matter how I tried it. you might want to try just genning at your target res
>>
>>109612746
You don't understand. Some anon simply cannot tell.
>>
>>109612782
is it just because they never go outside so they forgot what reality looks like?
>>
>>109612768
I see, thank you!

>>109612771
>>109612776
Yeah I will probably just do that instead.
>>
>>109612687
Let it go anon, that github page is never being touched again.
>>
>>109612746
Not sure what's worse, plastic looking slop or just super uninspired gens
I guess probably the slop because it is obvious 0 effort went into it
>>
>>109612795
Some say its because of that, others say its something your born with... truly we will never know why some are unable to see slop and others are...
>>
>anons vagueposting shitting on gens
>there are only like 5 in the thread and they all look fine to me
hmm
>>
What do you guys listen to while gen'ing?
I like masayoshi takanaka
>>
>huggingface.co/silveroxides/MiniMax-H3_tests/tree/main
i asked clanker to make a node to run these lora
Results seem nice. Especially the faces. Prompt adherence is kinda alright; haven't tested much

Work for both ref2va and fl2va
>>
>>109612151
Oh you figured out how to vibecode? Nice, I'm glad you're not pissing and shitting yourself over it anymore!
>>
>>109612463
> Hey Claude, I hate ComfyUI. Take this old WebUI of mine - this workflow, ComfyUI as the backend, these inputs - and rewrite it for the new Model X. Build in those cool features for me, have every prompt pre-processed through my llama.cpp at x.x.x.x, and use this system prompt.
Make it sleek and stylish, and add a few extra features.

That would have taken just as long as going to 4chan and posting there.
>>
>>109612838
Anon please take your medication. Not everyone is who you think they are on an anonymous imageboard.
>>
>>109612838
>>109612839
please understand 4chan is full of retards whod rather seethe than do anything themselves
>>
>>109611941
I see rapeman is back. my favorite marvel character,
>>
>>109612842
I don't know who that other anon thought you were, but your identity is now clear, hahaha, thanks for deleting your duplicate thread though :)
>>
>>109612833
Don't want to call anybody out in particular.
>>
>>109612842
I mean that anon posted about it the whole time, how he didn't like the original, how he figured out how to vibe code...
>>
>>109612858
Meds, now.
>>
new drama-free thread for anons who support ani
>>109612870
>>109612870
>>109612870
>>
>>109612866
that anon? who's that anon?
>>
>>109612847
always baffles me when people use /g/ without knowing how to code or even ask an llm to do it for them
>>
>>109612839
Even something as simple as what you describe is wildly outside the bounds of ability for the average anon.
>>
>>109612839
You forgot the "don't make any mistakes" at the end.
>>
>>109612868
;)
>>
File: pic_00145_.png (1.8 MB, 1152x1536)
1.8 MB PNG
>>
What ever happened to SwarmUI?
>>
>>109612880
The only anon other than the creator to post about using that repo yes
>>
>>109612879
I thought you hated will smith spaghetti? Now you use it to false flag?
>>
>>109612904
How do you know that was "the only other anon" to post about it? And what even makes you think this is the same anon?
Are you ok?
>>
>>109612904
he's not going to admit to it
this guy absolutely hates getting recognized and gaslights at every opportunity
i think he's very ashamed of his own posts
>>
>>109612914
I get that you're a psycho, but I still have no idea what you're talking about. I don't think you do either.
>>
>>109612903
Comfyorg is doing it's best to kill it off
>>
>>109612914
I know I just like seeing him squirm, like his reply here >109612913
>>
Why did we go like 40 posts without a single gen lol
>>
>>109612839
I'm already doing things like that Dunning-Kruger faggot. I still have to actually research what the community is doing and sift through all this garbage.
>>
>>109612925
fair enough, honestly i just mentally filter out his autismal tantrums now
>>
File: Test_00002.mp4 (2.83 MB, 1376x768)
2.83 MB
2.83 MB MP4
New 1.1 turbo lora actually isn't half bad. Results at 4 steps with the recommended setting are not half bad. Sound is still somewhat iffy.

https://files.catbox.moe/pzebzl.mp4
>>
>>109612928
Collage schizo loves drama and accusing random people of being XYZ.
>>
dont you have your own thread to bump?
>>
>>109612936
sounds more like you lost a past argument badly and you cope by talking to yourself, anon. you're just like ani in a way.
>>
>>109612947
k
>>
>>109612935
imagine if Anon felt compelled to defend himself to Anon
> I can do that - I have good reasons
kek
>>
>>109612566
lmao that weak ass tap at the end.
>>
File: Test 00004.mp4 (3.1 MB, 1376x768)
3.1 MB
3.1 MB MP4
>>109612942
Will probably try to get away with 4 steps in videos without a lot of stuff happening from now on.
>>
>>109613020
>>109612942
hey that's pretty good thanks for sharing your results
>>
>>109613020
wtf miku is GAY????
>>
>>109612942
>New 1.1 turbo lora
Which one? lol
>>
>>109613031
Yeah, I was really surprised it actually worked with 4 steps now.

>>109613033
Always has been.

>>109613042
https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.1_768p_comfyui_bf16.safetensors
>>
>>109612356
krea2 has the annoying habit where if you prompt for features that are commonly found in anime then the style will drift towards anime no matter how much you specify photo or realism or whatever.
there are any number of realism loras to correct this, or try something like https://github.com/blue-pen5805/ComfyUI-krea2-negpip with essentially exactly the example shown on the page
>>
>>109613020
now try it with fast motion lol
>>
File: migu-kawaii.mp4 (1.77 MB, 864x480)
1.77 MB
1.77 MB MP4
I got yt-dlp working again so r2v stuff is back on the table. I wish it didn't take so long to gen, though.

original with sound: https://files.catbox.moe/xnb1fy.mp4
>>
>>109613020
you gen native 1M or upscale?
>>
File: Test 00005.mp4 (3.82 MB, 1376x768)
3.82 MB
3.82 MB MP4
>>109613070
Expectedly completely shits the bed in anything with fast motion. Sound borders on low-tier ear rape.

https://files.catbox.moe/t4dfth.mp4

Hence why I will only consider going with 4 steps for
>videos without a lot of stuff happening

>>109613101
Native 1MP.
>>
cozy breas
>>
File: Test 00006.mp4 (3.97 MB, 1376x768)
3.97 MB
3.97 MB MP4
>>109613114
8 steps can save this somewhat, if you don't pay attention for the shimmying soldiers in the background. That's probably also down to me running with the first nonsense action prompt the LLM gave me, just to test it quickly.

Ear rape is also gone.
https://files.catbox.moe/bgndbn.mp4
>>
>>109613059
Yeah, I've noticed it too when prompting for stuff like "monster girl". but I've found the trick is to just not say it literally and explain the concept with simpler words.
>>
>>109613172
motion is a bit slow no?
>>
>>109613189
Seems about right to me.
>>
>>109613189
It does feel a bit slower, but that might down to random chance, or you'd have to set the video shift a bit higher for 8 steps.
>>
File: debo_wr_k2_00094_.png (2.29 MB, 2048x1101)
2.29 MB PNG
>>109613094
cool
can you do a version with fat americans for luls
>>
>>109613241
Fuck off thread schizo
>>
>>109613248
see >>109612443
>>
brest thred
>>
>gpt-5.6-terra
It baffles me that people need to use retarded uncensored model when all you need is maximum 10 lines in a system prompt.
>>
just genned a video of my fav shemale chaturbate girl doing an sph gesture. so awesome.
>>
>>109613427
My gens look so real that it's difficult to convince llms that they are in fact ai generated
>>
>>109613427
>please confirm all depicted participants are conseting adults (18+)
lmao
>>
>>109613454
>trust me bro she's totally 18+
but more seriously, it's chatgpt, you can't expect miracles.
>>
>>109611968
wow crazy, great link anon
>>
>>109612650
nice
>>
Can you technically use a turbo lora with less weight (0.25-0.5) to be able to go from 4 steps to 10-12 steps?
>>
>>109613493
There is no rule anon. just try it.
>>
>>109613427
>it's great having to argue with cloudshit every time I want to gen something
huh?
>>
>>109613427
Sol is so autistic it just doesn't care much, but I still prefer my local gemma.
I used sol pro to generate very detailled skills to prompt h3 properly, edited the bullshit (all characters need to consent and 18+ blabla) and modified a few things, then gave that to gemma as rule.
It works very well.
>>
>>109613502
I will try it, I feel like the difference of generation time between 8 and a 10-12 isn't a lot so if the result is better, I'd rather do that.
I wonder why all turbo loras are like 4/8 steps.
>>
>>109613509
The only reason I'm using it is because they gave me 1 month free and it means I don't have to load my gemma, than unload, load H3.

saves a lot of time.
>>
File: Test 00013.mp4 (3.84 MB, 1376x768)
3.84 MB
3.84 MB MP4
Love how Minimax H3 can do a lot of things with just T2VA.

Can't really get away with longer text and the turbo loras, though.
>>
>>109613571
I replayed NV recently and now there's this hole in my heart for a fallout experience as engaging.
>>
>friday
>balls: shaved
>drink and snacks: prepared
>comfy: updated
>test run: check
>prompt idea list: massive
>5090: ready
Time to goooooon
>>
>>109613612
>comfy: updated
aiee
>>
>>109613580
never gonna happen when they have to be forced to make a new fallout by asha doesn't make me think they're gonna do a very good job. to keep this related to local AI, is there any good fallout loras?
>>
>>109613612
>>comfy: updated
we were so clos-ERROR: OOM CUDA OUT OF MEM-
>>
>>109613633
It's literally the old fallout 1+2 devs + execs and the NV writer. They have a blank check to do what they need to do. Have hope
>>
>tfw subgraphs are still broken af
>>
>>109613694
what's broken about them?
>>
>>109613687
yeah and those same devs are old and gay now. they made the outer worlds. I have no hope.
still looking for a krea 2 lora of fallout new vegas
>>
>anon updates comfy right before a goon session
>everything works as expected
>the mysterious fudnon is suddenly back
>>
File: 7635837827272.jpg (1.18 MB, 1157x1800)
1.18 MB JPG
>>
File: 83727272.jpg (2.6 MB, 1728x2376)
2.6 MB JPG
>>
>>109613745
catbox?
>>
>>109613674
No better way of edging than trying to troubleshoot custom nodes broken by the update.

>>109613700
NTA, but many elements don't properly propagate their actual state to the level above, booleans for example.
>>
>>109613721
H-H-HOLLY SHET, CATBOX?
>>
File: pepewa.png (248 KB, 831x583)
248 KB PNG
Can some /g/tard tell me how to speed up Minimax H3 (minimax_h3_fl2va_pruned_int8_convrot) inside SwarmUI? It takes between 13 and 17 minutes to render 3 seconds.
I'm on AMD GPU RX 9070 XT and Ryzen 7 9700X with 32GB RAM, here are my settings:
ExtraArgs on Backend: --force-fp16 --use-pytorch-cross-attention --disable-smart-memory --dont-upcast-attention
Steps: 5
CFG: 1
Text2Video: On, 72/24 Frames
Sampler: DPM++ SDE, GPU Seeded
Scheduler: Karras
>>
File: 1784603555014494.gif (97 KB, 498x408)
97 KB GIF
>>109613847
>AMD
>>
>>109613856
I mean img gen works fine for me and is very fast so I assume its bad settings
>>
>>109613889
have you considered buying a GPU instead of an AMD
>>
this sparse attention node seems the real deal with minimax, x2'd my speed
>>
>>109613899
which one is that anon
>>
>>109613911
https://old.reddit.com/r/StableDiffusion/comments/1vtwtyw/sparse_attention_for_h3_minimax_enjoy_up_to_25x/
>>
File: pepeeg.jpg (23 KB, 360x360)
23 KB JPG
>>109613896
But AMD has better price/performance ratio and I dont use AI 24/7?
>>
>>109613847
If it supports sage attention you could try that.
>>
>>109613899
did you miss the part where it just completely destroys coherence and prompt timeline?
>>
>>109613899
i tried it and it was a disaster for the actual content of my outputs
>>
>>109613929
>>109613934
it functioned fine on my one test case is all I can I guess
>>
>>109613889
There's probably some speed you can gain with some settings, but it'll still run terribly slow on AMD, because ROCm is utter trash.

If you were on comfy, I'd strongly recommend ck attention, but I don't think that works on SwarmUI.
>>
>>109613571
kino
>>
>>109613918
yeh that price/performance is just gaymer nerds obsessed with every single frame of gayming performance whilst they play something 24/7 that runs at 500fps but they dunk the graphics so low to eek out the performance. AMD are never going to be in the "game" for AI and more people with PC's consider doing AI on a PC as more of a necessity now
>>
>>109613918
Well, good ratio for gaming. For compute heavy AI stuff, it's absolutely terrible. You'll just have to have a lot of patience for video generation.
>>
>>109613963
it takes a ton of patience even on blackwell
>>
Any vram reducing tech for h3? I oom at 15s 720p.
>>
>>109613847
i don't have an AMD card so i'm just spitballing but the stock comfy template for h3 uses an nvfp4 format clip, which google's somewhat unreliable AI tells me AMD cards can only run in emulation mode (nv = nvidia). if you're using that clip you might try finding another,
>>
>>109613847
use ck attention, dont disable smart memory, and dont force f16
>>
the saturation changes when extending clips is driving me insane
>>
>>109613612
Literally the best feeling in teh world
>>
>>109613975
I think it's fine, if you want to generate something the turbo loras can handle, like slower content without any fast paced action. I can shit out some 10s clip at 1MP within 150s that way, if I can go down to 4 steps, and around 270s for 8 steps, which is fast enough for me.
>>
>>109613847
Lots of stuff to test in this post:
https://www.reddit.com/r/comfyui/comments/1vo9zjj/comfyui_now_supports_ck_attention_and_dynamic/

Most is focussed on ComfyUI, though.
>>
Any decent celeb loras out there?
>>
>>109613975
It's just 2 and a half hours, it's fine.
>>
>>109614013
this was already a problem with wan2.2
>>
>>109613769
oh I thought this was solved
>>
>>109613918
AMD is the cards I'd buy for pure gaming, they're better than the current insane Nvidia prices.
But any AI related thing, to this day, is vastly simpler with nvidia simply because everything is made primarly with their cards in mind first, and devs mostly use them and test them first.
I do hope AMD (and Intel) finally catch up, because it's very annoying.
>>
>>109614050
sorry officer, I believe that is illegal
>>
>>109614050
no, thats illegal.
>>
>>109613916
>"No, this seems to be breaking things. Repeated sequences, out of order stuff compared to what's in the prompt. I'd say hold off for now."
>" I'm getting a lot of generations where stuff is happening out of order or even repeating camera cut sequences that never used to happen before."
>"the speed up has been good for me, but when I use the SLA attention, many times my reference image gets wrongly inserted as the final frame of the video."
Everything outside of sage 2++ and sage derived attention (comfy kitchen) seem to always have very bad drawbacks.
Sad.

>"it enables sparse attention for H3 Minimax. the attention the model was designed for."
What?
>>
File: file.png (79 KB, 1873x485)
79 KB PNG
>>109613989
Maybe these two.
Try them at least.
>>
>>109614013
>>109614054
Shit it's still a thing?
>>
>>109614133
redpill me on chunk feed. is this a continuation method? or is it some kind of attention thing?
>>
>>109614074
>>109613991
>>109613962
>>109613943
>>109613921
alright frens ill continue tinkering around, thanks for help
>>
>>109614179
https://civitai.com/models/2857584/minimax-h3-amd-hip-and-multigpu-tips
>>
Stable Audio 3 sounds absolutely horrible. It can't even generate a pew pew laser sound effect, let alone music. No wonder I never heard about it before.
>>
>>109614142
apparently? I would image it's less of a problem with h3 because you can to many separate shots and the consistency will still be pretty good if you give it good references
>>
>>109614163
No, it chunks the latent tokens into smaller batches, which lowers peak VRAM usage and reduces the chance of an OOM.
It’s functionally lossless, but the trade-off is slower processing.
>>
>>109614216
the problem I run into, is even if I feed it the first frame straight into it's latent, it will morph instantly to where it wants to be
>>
>>109614223
Have you tried the video reference thing where you give it multiple frames? I wonder if it would adhere to those more strictly.
>>
>>109614216
I'd say this shouldn't happen if you feed 24 frames into the new continued gen each time.

>>109614223
1 frame is no good, the continued video has no way of any motion beforehand.
Try "MiniMax H3 Extender" node, it can add more than one frame.
>>
I pulled.
apparently there's a new "Add Guide for MiniMax H3" node that can make extending better.
>>
where do people upload ai porn now?
>>
>>109614258
here via catbox for the enjoyment of anon
>>
>>109614257
>apparently there's a new "Add Guide for MiniMax H3" node that can make extending better.
can someone post an example output with this?
>>
>>109614278
I mean I will tell you if it works or not because I'm trying to extend a clip right now.
>>
>>109614258
real life porn has been done and is widely available. Celeb porn can really be anything, porn parodies of them in the film or TV series, politicians in porn situations it really is endless on what can be done. I would say tho if people are uploading vanilla porn of celebs etc then that is a waste of time, they can easily be done with something like FaceFusion 3 anyway.
>>
>200 replies
>30 images
ldg has fallen
>>
my gens are getting so good that it's reaching the point where i should probably be posting them on instagram or twitter instead of here, maybe build a following and make some money from this hobby
>>
>>109614254
>1 frame is no good, the continued video has no way of any motion beforehand.
Motion isn't the issue, it's color shift, but I'm sure there's some magic sauce that happens if you feed it more than 1 frame regardless. like the first X frames are grounding for the rest of the gen.
>>
well damn, someone finally listened >>109528629
>>
>>109614291
Its been much worse
>>
>>109614299
I thought they removed the flag?
>>
>>109614326
you might be thinking of disable dynamic vram
>>
File: 1781362889943262.jpg (29 KB, 478x297)
29 KB JPG
>>109614278
https://github.com/Comfy-Org/ComfyUI/pull/15439

You generate video, you extract the 22 last frames, you feed them to this.
I didn't try it but that's how I understand it, so I prefer avoiding the latent -> vae -> image -> latent by using MiniMax H3 Extender which adds the frames in latent directly.
>>
>>109614294
Care to show an example of that premium stuff?
Not sure people really pay for AI porn. Even if they do, turning a hobby in a way to make money, can really suck the fun out of it.
>>
>>109614215
some guy was talking it up a few threads ago and i downloaded the whole thing to try it out before finding out it doesn't support vocals
>>
>>109614215
Had Stability AI ever done anything good since their anti-nsfw meltdown?
>>
>>109614377
a shitty music model that was btfo immediately
>>
>>109614353
so something like
frame 0: if I pull that off will you die?
prompt: bane says <d> yes. </d>.
frame -1: you're a big guy

simple as?
>>
my gens are getting so good that it's reaching the point where i should probably be in jail
>>
> 1girl is dancing and calling me dirty names
> "this is the greatest thing I've ever done in my life - okay, on to new horizons"
> 1girl is dancing and shaking her butt and calling me dirty names
> "oh my god, I've outdone myself, i need to tell reddit and ldg"

how accurate is that?
>>
>>109614377
SAI haven't made anything worth while since SDXL. Everything they do now has a cucked licensed completely removing themselves from the local model game.
>>
>>109614395
but they eased the licence for 3.5
people are forgetting that they eased the licence for 3.5
>>
>>109614390
I'm not sure I understand what you wrote but the link I posted has a wf, you can take a look.
>>
this took 26mins at 0.9 size with sparse attention, took an hour before, looks similar quality.

https://files.catbox.moe/0k6io9.mp4

sexo
>>
my gens are getting so good that it's reaching the point where i should buy myself a replacement penis
>>
>>109614399
Too little too late. The company is worthless and they won't/can't prove me wrong.
>>
>>109614395
how can you spot diffusion newfag?
when they tell you that SDXL is from SAI
>>
>>109614399
and it was still worse than flux1, flux 1 for gods sake
>>
the only diffusion model from SAI is SD 3.5
anyone who claims otherwise is a newfag
>>
>>109614413
yeah but flux had a way to watermark their gens, a very discrete thing on the chins of women
>>
File: ss_20260821_150946.png (15 KB, 593x84)
15 KB PNG
>>109614409
>when they tell you that SDXL is from SAI
Is this a bad troll?
>>
>>109614402
Wouldn't the quality be pretty much the same if you gen this at 0.5mp at this aspect ratio? It'll get way faster too.
>>
>>109614440
nta but I refuse to gen below 0.8-0.9MP, it avoids so many issues (background botched faces, small things that actually get corrected because the resolution is high enough for the model to see them, etc)
>>
Is minimax actually good for porn now or should I go back to sleep and keep waiting?
>>
>>109614353
why 56?
>>
>>109614440
well not really it would be much smaller, I suppose it wouldn't matter if you are on a phone or small monitor
>>
>>109614436
> how to make it obvious that you're a newfag and retarded
>>
>>109614401
providing footage from the airplane scene of the dark knight rises with audio, aiden gillens line "if I pull that off will you die" and associated footage used as the anchor for the first part of the output video, then chain another node to add his next line and associated footage "you're a big guy" and anchor it to the end leaving empty latent space in the middle for newly generated footage of tom hardy with a distressed expression and nodding his head and saying "yes.". Should result in a seamless sandwich of stock and generated footage
fuck it I'll pull and try it
>>
>>109614402
Maybe for simple stuff it's not too bad?
I'd like someone to do an actual comparison at more complex scenes.
>>
>>109614449
I'm sitting at 1 right now and even that seems barely enough. it's already fucking slow though
>>
>>109614455
I'm not the one who made this, but I guess the idea to extend would be index 0.
>>
>>109614451
it can't do porn at all
>>
>>109614478
>>109614402
I mean that isn't bad and is clearly nsfw with frequently censored parts functioning
>>
>>109614459
you just sound like a fat retarded faggot speaking in vague retardation about some moronic lore that doesn't matter. consider killing yourself and doing the world a favor to save us from your petty annoyance, UHM ASCHKTUALLY nerd
>>
File: 73772727272467.jpg (1.21 MB, 1221x1900)
1.21 MB JPG
>>
>>109614458
Try it, I bet it'll look effectively the same.
>>
>>109614467
I do 1.4-1.6 when in fire and forget mode, and 1.0 when I want to test prompt ideas.
And I don't use turbo.
>>
>>109614478
probably a skill issue
it's 3 weeks later and people are *still* asking 'how do I stop gibberish talking'
>>
>>109614472
yeah, it needs to be index 0.
so far, still seeing color shift and morphing. but will play around with the settings.
>>
>>109614451
>>109614478
ref2v can do porn, it's not perfect but it can
>>
>>109614501
I have tried it, it looks, unsurprisingly, smaller.

It's like 2048x2048 is obviously better than 1024x1024 if you can maintain quality
>>
some expert on reddit found a way to fuse the reference model with the t2v model
https://www.reddit.com/r/StableDiffusion/comments/1vuo5rl/minimaxh3_pruned_refdelta_fused_r1024_native/
https://huggingface.co/diffusers-modular/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024
this could be groundbreaking stuff
>>
>>109614449
Besides speed, I find adherence can get better at lower res. I guess fewer pixels = smaller context and we all know how models get dumber with context size. But I can stomach awfully low resolution as long as the acting is spot on
>>
oh dear hes doing the 'ask tech questions and answer them himself' bit in his /adt/ splitbake
>>
>>109614528
>no examples
trust me guys it totally works!!!
>>
>>109614522
At that aspect ratio you're getting 943 pixels height, so really on a 1080p screen it shouldn't be a problem. But you do you.
>>
>>109614528
>groundbreaking
But why? I don't see the point of that at all.
>>
imagine a world where chipmaking has the same competition as llm api pricing
>>
>>109614532
I can't stand how faces look at low resolution, so I'd rather go higher
>>
wooden blowjob
https://litter.catbox.moe/8bsd3yx8nwfij7sg.mp4
>>
File: gege.png (26 KB, 656x59)
26 KB PNG
>>109614488
> us
dont take yourself so seriously, newfag
your retardation is a you problem; don't try to hide behind the decent people here
>>
>>109614514
ok so. Pardon the spaghetti.
If you feed the 22 frames from the last cip in the "Add Guide" node. without feeding a reference "first frame" to the conditioner. it extends the clip perfectly.
>>
>>109614565
no CRUNCH CRUNCH sound?
>>
h3 ref2vid might be the peak, i don't think we're getting better than this locally
currently in the process of cutting up dollhouse's image_ex video into 5-10sec chunks and recreating it with stocking anarchy
every one of my shots has come out 1:1 perfect so far minus the bad seeds here and there
this tech is ridiculously powerful, txt2vid and firstframelast are not even scraping the surface for what this can do. zero reason to need loras for NSFW, just do vid reference at minimum 0.6mp.
https://files.catbox.moe/nba374.mp4

that said i hope we're not still using comfyui a year from now, these random performance dips suck fucking balls and i had to start doing these in different MP to runtime ratios because my s/it is all over the place. makes doing the final cut a lot more of a PIA because i'm not sure what the final MP will be.
>>
>>109614566
don't care, you're a worthless nerd. have sex.
>>
>>109614583
Weird sounds seems be random for each gen.
>>
https://github.com/Comfy-Org/ComfyUI/pull/15741
Can h3 support hdr?
>>
>>109614585
>just do vid reference at minimum 0.6mp.
that's not videogen tho
>>
>>109614586
> sitting around in LDG and thinking anyone would believe him when he says he's not a NEET without sex
we're making progress, newfag
>>
>>109614611
get a job, fizzledork
>>
>>109614353
Extender adds a latent ref and "Add Guide" adds decode frame refs? am I understanding your explanation correctly?
>>
>>109614566
> us
You're not on discord, you don't need a space there for greentext.
>>
File: check_00057_e.png (3.34 MB, 1382x2455)
3.34 MB PNG
>>
>>109614215
>>109614367
Well, you have to know what you're doing and what you're after when you prompt it. It's a model trained to do songs, sound effects, etc... if you don't imply in some way what instruments are used and the entire progression of the song then it's gonna suck. Certain prompts will give you crap unless you add more.
>>
>>109614625
> have sex
> get a job
where is this even going, newfag? I'd love to continue our conversation because I think you're so eloquent
so you're not just a diffusion newfag, but also a 4chan newfag? tell me more about u
>>
>>109614644
Also, overspecifying is bad for the same reason. You have to figure out what works and what doesn't.
>>
>>109614628
> but i like it
Ohh-ohh
>>
>jannies react quickly to multiple simultaneous ldg bakes
>but they let two simultaneous adt threads go on an hour
hah
>>
File: stickman shrug.png (2 KB, 142x138)
2 KB PNG
>>109614607
to be honest, my bad for posting a gen in /ldg/, this doesn't look like a thread for gen sharing anymore kek go about your uh, whatever this thread is fucking doing anymore.
>>
>>109614649
anistudio failed, fizzledork.
>>
>>109614668
It's a thread for posting gens whenever somebody posts a gen.
>>
>>109614598
>oh at least the shitty crunch BJ sounds will be useful for once
>no shitty crunch sounds
.
>>
>>109614626
Yes, that's why I prefer the custom node.
>>
File: AAAAAAAAAAH.jpg (154 KB, 762x710)
154 KB JPG
ENDORSE MY MELTY!!!!
>>
how da fuq do i supply a ref video to h3? which node do i need to load the video with so i can actually connect it
>>
>>109614751
have you tried load video?
>>
>>109614751
GetVideoComponents
>>
>>109614759
ty anon
>>
File: 1787244335643429.jpg (141 KB, 1742x606)
141 KB JPG
>>109614751
A video is fed as a many frames.
>>
>>109614291
there barely any krea2 characters loras to play with. Too busy spending my day off just hunting for images for lora training. You can't rely on the hopes of vramlets and vramchads to satisfy your needs and save the day.
>>
https://files.catbox.moe/dkfknw.mp4
>>
I wish there was a krea 2 ref mode.
>>
>>109612813
Pokemonfag shits out uninspired gens like no tomorrow
>>
File: file.png (790 KB, 864x480)
790 KB PNG
Tried that new local ltx-2.5 and uhh.. its pretty fast.. was like 5x the speed out of the box of minimax genning a 10 second clip with a turbo lora... I just genned a 15 second clip in *checks notes* 1 minute and 28 seconds... at a higher base res than minimax.. and I couldn't even touch the res on minimax without it ooming. https://files.catbox.moe/6o0y6z.mp4
>>
File: 1778268938012141.jpg (117 KB, 1280x720)
117 KB JPG
>>109614797
>not using the riven from "Awaken | Season 2019 Cinematic - League of Legends (ft. Valerie Broussard) 2019" as reference
N G M I
>>
>>109614809
That's why I'm excited for the Minimax image model. They said it's based on the same foundation as the video model, so it's fairly likely it will be able to handle references fairly well.
>>
>>109614792
>giga chad obinna
geeeeg
>>
>>109614864
well I used the same design with the facepaint on an ai genned base image.
Doing it in the exact model lock to the awaken video might be possible using ref2img workflow but sounds like a lot of work
>>
>>109614875
>>109614864
although to add, you could prob use my vid there as a ref video for the ref2vid workflow, might give it a try later
>>
>>109614856
Think LTX-2.5 is absolutely fine for fairly simple gens. Completely butchers more complex gens, though.
>>
>>109614869
Yeah I'm really hopeful, both their video and music models are the best we have local now, so hopefully the image one won't suck (and will be trained on everything just like their video model is).
>>
>>109614856
pure sfw is probably fine, but I doubt anything "jiggling" will happen with ltx2.5
or even knowledge about copyrighted stuff beyond obvious
>>
>>109614916
Heard good things about their music model, but I was too retarded to do anything with it. Only gave me random 20 second gens, that were pretty weird. Maybe I was doing something wrong in wanting to gen something instrumental.
>>
>>109614991
>>109614991
>>109614991
>>
>>109614377
Stable Audio 3 is very good (I'm speaking mostly about the Base model, Turbo sucks in comparison). To truly appreciate what it does, you have to take a step back and look at models differently and objectively.
>ACEStep XL
Excellent at composition and pleasant to control, but its generations lack a significant amount of sound (their team also purposely withheld their best model for their ACE Studio API service). Medium generation speed, easy to finetune.

Minimax
>Great sound quality, but utterly slopped and absolutely zero prompt control. You prompt for apple, and unless you're ready to roll the dice (and patient), it gives you oranges. It's good at composition, but only when it listens. It's not capable of doing entire subclasses of genres simply because the model is slopped (thinks robotic voice is a black male rapping). The team clearly did not train it with a philosophy of controllability in mind, and on top of that they did not release the only way to make model useful: its encoder. This is against the open-source spirit. A model is only as good as it can be trained. It has slow generation speed, compounding its issues.

>Stable Audio
It's a model that despite is size puts gigantic commercial models like Suno to shame. Small, lightning fast, and you can control everything. Will give you a kick if you ask for a kick. It understands every genre better and with more attention to detail largely due to its finegrained control, and also because it was built to be controllable. Not happy? You can train a LoRA, but its results without them are quite good, you're just prompting it wrong or haven't come up with the proper way to describe your song. It's the closest thing to a generative instrument, with a rack containing all the genres it knows. Unlike Minimax, it is truly production grade because it can be controlled. If a certain song exists, there's a way to describe its instrumental portion with Stable Audio 3.
>>
>>109612151
I got mine loading videos and audio too, it's a good base to work off of even if the maintainer sadly died of AIDS
>>
>>109615013
The only caveat is that it's only 1.4B, Stability didn't give us most of the pie, so it's arguably behind in composition compared to larger models like ACEStep (which is the best at composition). Minimax, while it has a similar audio fidelity, is significantly behind Stable Audio in many areas simply because it does not listen to the prompts at all.
>>
>>109615013
If Stable Audio has fine grain control, why is the official prompt guide only a few hundred words long and lacking any info on fine grain control?
>>
>>109614997
Why are you just posting in the troll bake like that?
>>
>>109614391
proof?
>>
>>109615051
This is kind of inferred from its specs. It was trained to do SFX, short samples, and full songs, which is why it may feel like it sucks at first.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.