[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1785010577476917.webm (3.24 MB, 2048x1166)
3.24 MB
3.24 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109617657

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
When do you think minimax will share their upscaler model? And their sparse attention implementation?
>>
>mfw Resource news

08/22/2026

>ComfyUI-H3-AudioRefine
https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine

>Anima-3.8B with Qwen-3.5 4B
https://huggingface.co/lylogummy/Anima-3.8B

08/21/2026

>MiniMax H3 Super Acceleration fast draft generation and high-resolution refinement, powered by Sol Engine
https://nvlabs.github.io/Sana/Sol-Engine/H3-Super-Acceleration

>4DAnyone: Create Anyone in 4D from a Casual Monocular Video
https://4danyone.github.io

>H3 Prompt Composer Version 5.37.1
https://github.com/BMB12d3/minimax-h3-prompt-composer

>LTX-2.5 Gemma-4 12B NVFP4 for ComfyUI
https://huggingface.co/Deadshot699/ltx-2.5-gemma4-12b-comfy-nvfp4

>MiniMax-H3 Pruned Ref-Delta Fused r1024 — ComfyUI Single File
https://huggingface.co/xmarre/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-ComfyUI

>Phosphene 4.6.0 Adds Video Editor, H3, LTX 2.5
https://github.com/mrbizarro/Phosphene/releases/tag/v4.6.0

>H3 Optimizations: Standalone production optimization nodes for MiniMax H3 in ComfyUI
https://github.com/Zironic/H3-Optimizations

>MiniMax H3 known character lists
https://huggingface.co/datasets/malcolmrey/various

08/20/2026

>MiniMax-H3 Turbo 4-step v1.1 by Lightx2v
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

>MiniMax-H3 Turbo-SLA
https://huggingface.co/lightx2v/Minimax-h3-Turbo-SLA

>MiniMax-H3 Single-Frame VAE 500K
https://huggingface.co/iamkaikai/MiniMax-H3-Single-Frame-VAE-500K

>MiniMax-H3 RAVEN Streaming LoRA
https://huggingface.co/mvp-lab/MiniMax-H3-RAVEN-Streaming-LoRA

>Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
https://pardistaghavi.github.io/SparsePR-website

>VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
https://github.com/ShareLab-SII/VA-Judger

>Frozen DINO Localizes Image Edits Without a Localizer
https://github.com/VishalJ99/trail-image-edit-localization
>>
>mfw Research news

08/22/2026

>RigidBench: Evaluating Rigid-Body Physics in Video Generation Models
https://doi.org/10.5281/zenodo.21649156

>UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models
https://arxiv.org/abs/2608.15238

>Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation
https://arxiv.org/abs/2608.16384

>Does Marginal Coverage Guarantee Class-Conditional Safety for Zero-Shot VLMs Under Shift?
https://arxiv.org/abs/2608.19376

>Core-KAN: Continuous Vision Kernels with Kolmogorov-Arnold Networks
https://arxiv.org/abs/2608.19817

>Anchor-Regularized Adaptation for Generalizable AI-Generated Image Detection with DINOv3
https://arxiv.org/abs/2608.15196

>TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
https://arxiv.org/abs/2608.19737

>Discovery and Spatial Characterisation of Multiple Shortcut Groups for Auditing Vision Model Bias
https://arxiv.org/abs/2608.14051

>Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation
https://arxiv.org/abs/2608.13602
>>
> >109621408
> >109621412
Fuck off
>>
>>109621408
>>>ComfyUI-H3-AudioRefine
>https://github.com/Adudeguyman/ComfyUI-H3-AudioRefine
That sounds like a cool idea.
>>
What is the best configuration for speed/quality in mini max? I was using spectrum, there is something better with new updates spectrum has less quality with the turbo lora.
>>
File: 416207362236395.mp4 (3.57 MB, 960x544)
3.57 MB
3.57 MB MP4
>>
>>109621447
Kinda wish the author would've used a more egregious example of bad audio. Would've liked to see how it cleans up the ear rapey music of more intense turbo gens.
>>
File: Test 00030.mp4 (3.9 MB, 1376x768)
3.9 MB
3.9 MB MP4
>>109621513
Certainly much better than mine. I think everything that avoids too many profile shots, and even dares to have some intense camera movement avoids looking like cheap animation.
>>
>>109621359
>>Maintain Thread Quality
>https://rentry.org/debo
>https://rentry.org/animanon
It's a lot funnier to bake without these and make the lolcow shit his pants and get banned (again). Remember that for next time
>>
File: 416207362236395_2.mp4 (2.92 MB, 960x544)
2.92 MB
2.92 MB MP4
>>109621513
that's too much CGI for me
dynamic dropouts would look better, but I don't feel like messing around with the settings.
>>
>>109621447
ask when it's going to be in the ggml versions. cumfart just sounds like ass to begin with
>>
>>109621689
nvm it's h3 not the music model
>>
File: 854939554061271.mp4 (3.9 MB, 960x544)
3.9 MB
3.9 MB MP4
>>109621662
Yeah, one of the things that sets H3 apart from the previous gen local video models is the ability to do multiple shots/cuts in one clip.
>>109621675
It's all CGI I suppose lol, but I catch your meaning.
>>
Did Anon ever post the catbox?
>>
What are some good prompts to try?
>>
Reposting.
>>
>>109621883
Appropriate. Rentry obsessed baking troons still don't get it
>>
File: 1780596690435582.jpg (265 KB, 1462x1100)
265 KB JPG
>>109621492
I use this but it's more "quality I like with a reasonable amount of optimization", I don't mind waiting as I hate bad quality gens.
>>
>>109621513
kissing is so much more erotic on h3 vs ltx 2.3, but man the tongues always seem to merge lol
>>
>>109621924
> man tongues
>>
File: StraightLinks.png (36 KB, 347x182)
36 KB PNG
>>109621914
I love the straight links in theory, but I think I'll stick with spline until they've fixed bullshit like this.
>>
>>109621934
TONGUE(1)

NAME
tongue : manipulate, deploy, and occasionally bite the muscular hydrostat installed in the oral cavity

SYNOPSIS
tongue [OPTION]... [TARGET]

tongue --speak [LANGUAGE]

tongue --taste [OBJECT]

tongue --kiss [TARGET]

DESCRIPTION
tongue provides a general-purpose interface to speech, taste, swallowing, licking and kissing.

The tongue operates in close cooperation with mouth(1), teeth(1), and brain(7). Running without brain(7) may produce undefined speech.
>>
>>109622002
I largely prefer that to curved spaghetti, and usually I have my nodes in a way so it's not that confusing, but I get it.
>>
wake me up when we can 3d print life size synthetic 1girls. women need to be replaced.
>>
>>109622075
seriously someone has to train the model for sexy sfx because this crunchy sound is everywhere
>>
>>109622017
kek
>>
>>109622086
finetunes of video models tend to take a while
>>
>>109621408
>>109621412
thanks!
>>
>>109622064
nice prompt idea
>>
>>109621883
i still dont get why there's two catjacks
>>
Good afternoon anons. I'm a complete beginner except for the fact that I had stable diffusion installed on my laptop few years ago and could never work it properly.

I installed swarm ui on my PC, and I need a model recommendation. i think i just want to generate waifu images (portrait, selfie, lewd, etc) for now. pls help.
>>
>>109622215
Wanschizo is the other spammer obviously. Catjack is black irl which explains a lot
>>
File: 1781355290621741.jpg (85 KB, 800x800)
85 KB JPG
>years pass
>have a hundred thousand gens
>forgot a lot of them
>built app to show me a random image
>relive seeing those great images again
>build app feature to automatically send random high-rated prompts with a random seed to comfyui to generate more similar images
>get more great unique images with 1 click
>>
What is the best uncensored qwen model for minimax?
>>
>>109622243
https://civitai.com/models/2458426/anima
https://civitai.com/models/833294/noobai-xl-nai-xl
>>109622272
https://www.reddit.com/r/StableDiffusion/comments/1vmdxzk/psa_im_the_creator_of_heretic_and_i_advise_you_to/
>>
>>
>>109622293
I wont trust the heretic faggot. He is just afraid of the Chinese government
>>
File: 478964704597027.mp4 (3.5 MB, 832x640)
3.5 MB
3.5 MB MP4
>original, 1girl, angel, angel wings, closed eyes, dress, flower, from side, full body, halo, lying, on ground, on stomach, outdoors, red dress, short hair, solo, traditional halo, white hair, wings

hmm
>>
>>
>>109622408
love the pose
>>
>>109622447
Thank you I love you too
>>
File: 51036428917253.mp4 (3.5 MB, 832x640)
3.5 MB
3.5 MB MP4
>>
File: MiniMax_H3__00281.mp4 (1.52 MB, 448x672)
1.52 MB
1.52 MB MP4
>>
I can't wait until this stuff gets good enough to make actual movies.
>>
>>109622490
I just want real fucking time 480p generation on a 3090
>>
>>109622293
I forgot to mention i installed Z-image turbo because apparently its better. i didn't get stable diffusion
>>
what exactly is the difference between /ldg/, /adg/, and /sdg/? why have 3 generals for a similar topic? is one of these a containment general?
>>
>>109622490
If it's good enough, everyone will start making movies, and then it'll become uninteresting again.
People do things to get attention from others; if others can do it too and they don't pay any attention to your AI slop, you'll lose interest in it and move on to the next thing.
Maybe you'll start juggling burning monkeys while jumping over burning tires on a burning skateboard.
>>
Am I the only one who feels like prompt adherence for MiniMax H3 has degraded in ComfyUI?
>>
>>109622559
Compared to previous versions of ComfyUI, or to API?
>>
>>109622559
You can easily compare by fetching an older commit.
>>
File: 253404239031963.mp4 (3.79 MB, 1184x800)
3.79 MB
3.79 MB MP4
>>
>>109622529
> everyone will start making movies, and then it'll become uninteresting again
Doubt it. Most people are extremely uncreative and don't have what it takes to direct a movie.
>>
>>109622592
Yeah but the problem is there are other moving parts like the cope nodes.
>>
>>109622648
You can do the same with the nodes.
>>
>>109622644
indeed, some people are just more creatively gifted
>>
>>109622644
and because they're so uncreative, they don't even recognize well-made videos and are satisfied with any piece of shit
Why do people produce scripted reality shows, or why is 90% of every anime season absolute trash?
Because the majority is satisfied with it.
>>
>>109622644
Anyone that watches anime made after 2010 regularly all fall into this category.
>>
File: 419192876859793.mp4 (3.44 MB, 960x544)
3.44 MB
3.44 MB MP4
>>
>>109622674
>creatively gifted
Some, sure, but you don't need natural talent to craft a good story. Most people simply lack the discipline and taste.
>>
>>109622604
hot
>>
god bless shitty low requirement vidya games that still run while you're making a h3 gen in the background
>>
>>109622746
yeah.. its good to have idle games and whatnot goin while u gen
>>
>>109622746
>tfw have a Steam Deck
Playing Mega Man 2 and generating coom videos is extremely comfy.
>>
>>109622505
i know you fuckers saw my post, fucking answer me
>>
>>109621673
>It's a lot funnier to bake with these and make the lolcow shit his pants and get banned (again). Remember that for next time
Fixed your post for you, Julien
>>
>>109622778
eat my ass
>>
File: 847919202305269.mp4 (3.4 MB, 960x544)
3.4 MB
3.4 MB MP4
>>
>>109622490
its already good enough with seedance 2.5
for local its still a couple of years away
>>
>>109622801
nword
>>
>>109622820
Don't use cloudshit. What sets it apart from H3 that makes it good enough for movies?
>>
What should I generate?
>>
>>109622666
How far back should I go? Should I revert all the cope nodes? You make it sound so easy.
>>
File: 1762939567511376.mp4 (1.44 MB, 640x960)
1.44 MB
1.44 MB MP4
>>
I was testing comfyui with 16gb card, now i am using 12gb card and i've noticed that my pc is lagging pretty bad when generating. I was able to do other stuff on previous card in the meantime.
Is it caused by vram?
>>
i really like providing last frame instead of first frame, means you can gen with Krea 2 or something and you write the prompt of how it gets to that point
but H3 tends to freeze the motion at last 2 seconds with last frame, idk if i need to add some motion blur to the input image or something or prompt differently
>>
>>109622837
No idea, you're the one having the issue, so you should know when you started to have it and revert back nodes and the program to the date when it felt faster for you.
If you don't see a difference it was all in your head, if you see one you can start debugging to see if it's comfyui itself or one of the custom nodes.
>>
>>109622861
Not necessarily, I for example never lagged while genning then realized windows 11 raped me with an update and resetted my pagefile settings, creating 200GB pagefiles over and over again on my main ssd. I switched the pagefile over to a useless and slow dummy ssd and ever since then I also lag while genning. So there are literally a thousand reason why you could be lagging, but I don't think it's your vram and it probably has to do with ram and memory management instead
>>
>>109622834
PAWG oil massage
>>
>>109622908
nta but solid advice
>>
>>109622828
>What sets it apart from H3
the speed, the image quality, the prompt adherence, the quality of the acting, the quality of the audio, the amount of references you can use, the ease of prompting it
>>
>>109622896
I get 100% usage on my hdd for pagefile.sys
There's nothing related to comfy on my hdd and system is running off of different drive so i don't know why the lag is happening.
>>
wildcard prompts are a whole new level of dopamine hit
>>
>>109622928
any recommended list?
I wanna gacha
>>
>swaps to harddrive
heh... nothing personal
>>
>>109622926
>I get 100% usage on my hdd for pagefile.sys
So it's fully utilized and that is causing your lag, unless you meant you get 0% usage?
>>
holy... my gens are already clearing by a mile, but if i had a sensible clip extension node made by someone who's not a complete fucking vibecoding retard id become a household name
>>
one of the best things I did with image and video gen was to switch to linux, it's so much more stable
>>
>>109622962
I get 100% usage and it seems like it's causing the lag and it's beyond annoying.
>>
>>109622926
>>109622962
I also got another theory, the current comfyui has a cuda bug or something like that. It's especially noticeable if you run multiple gpu's in your system causing it to give you OOM errors unless you start with the command --cuda-device "0/1/2/etc". So it could literally just be comfyui and you could try and download any version before 0.3v to see if the slowdowns still persist, though you'd lose many native functionalities for certain things like MinimaxH3 and Krea2 I think
>>
>>109622983
which distro are you running? and are you running as a workstation or headless?
>>
>>109622994
headless ubuntu server 24.04
>>
>>109622926
Your system is dumping RAM content into your HDD to make space for whatever the active task requires. Video gens take up tons of memory because each frame is written to RAM between node outputs (Comfy caches node outputs so the more [Image] output nodes you have, the more RAM it will eat).
If your pagefile is on your system drive, your system will freeze up until it's done writing.
>>
>>109622967
Show your gens, then.
>>
anyone tried this?
https://github.com/ttulttul/ComfyUI-Minimax-H3-Continuation
>>
>>109622991
Also, also, you could try to disable the pagefile completely (win+r ->sysdm.cpl) then after restarting your computer see if it lets you gen without an OOM error, this could solve your problems potentially, but from my experience comfyui always uses a pagefile even if you have 192gb of ram and 80gb vram.
So yeah either try one of the portables before 0.3v or play around with your pagefile
>>
>>109623021
My gens are too powerful for you, anon.
>>
File: 1763177630162729.mp4 (1.03 MB, 640x960)
1.03 MB
1.03 MB MP4
>>109622994
i agree with anon, nixos workstation
>>
>>109623030
Yeah, thought so.
>>
File: 623621335211008.mp4 (3.39 MB, 640x832)
3.39 MB
3.39 MB MP4
>>
>>109622834
woman swimming underwater and it tracks her ass
>>
>>109622834
woman swimming underwater and it tracks her penis
>>
>>109623021
i cant its pornographic in nature, I will be sanctioned! My previous gen was in the collage last thread and I'm suspected it caused the OP image to be deleted!
>>
>>109619514
>Also everyone is using Comfy now, it seems. Is that for any good reason, or did it just shake out that way? And do I need to copy them?
The current meta is vibecoding your own front end for ComfyUI. It is, in fact, very comfy.
>>
File: 821874655206592.mp4 (3.35 MB, 640x832)
3.35 MB
3.35 MB MP4
>>
>>109623076
>>>/gif/vdg
>>
its a shame h3 doesnt have more training on VHS tapes
>>
>>109623054
oops i came
>>
>>109623020
Why it wasn't happening with previous card?
>>
File: 1064916834175441.mp4 (3.6 MB, 608x864)
3.6 MB
3.6 MB MP4
>>
>>109623165
I would enjoy this subnautica even more
>>
File: 1775989447782025.mp4 (1.04 MB, 640x960)
1.04 MB
1.04 MB MP4
>>
>>109623153
You went from 16 to 12 VRAM. Model can't fit in VRAM, model pages to RAM. Not enough RAM for offloaded model and all of Comfy's bullshit? Pagefile.
>>
File: 1763834962283668.mp4 (346 KB, 864x480)
346 KB
346 KB MP4
Was supposed to be night time and much spookier but I'm retarded and didn't specify that to Gemma, nor did I catch it when proofreading the prompt.
>>
>>109623025
have you tried it? or other extender nodes?
>>
>>109623282
>proofreading the prompt.
i hate doing this. i never checked my dataset captions, either.
>>
>>109623282
Also I just realized it switched the hand that opened the door. Fug.
>>
>>109623020
How can Ani use his HDD and still get good speeds?
>>
File: 100747207440926.mp4 (3.53 MB, 608x864)
3.53 MB
3.53 MB MP4
>>
>>109623282
what does this makes me think of an actual horror movie with this exact scene
>>
>>109623025
>>109623284
It's the same as https://github.com/tritant/ComfyUI_MiniMax_H3_Extender

On paper, extender has more options, they do the exact same continuation under the hood.
>>
>>109623330
what's with the flash
>>
>>109623329
his gens look like straight dookie shit poopy brown turd excrement.
funny the two people who have a rentry on them also seem to be the worst at genning
>>
Are the modded 2080Ti 22GB or 3080 20GB models worth picking up if I'm on a strict budget?
I'd go for a 3090 but they're horribly priced in my local market.
>>
I do not care for animation
>>
File: file.png (127 KB, 1436x785)
127 KB PNG
This Upscale node never seems to progress - there's nothing in the console and the GPU is working hard. It's a 7 second video, like 720p. Did I set it up wrong?
>>
>>109623347
sdcpp can only get better, comfyui can only get more annoying
>>
>>109623352
>modded 2080Ti 22GB
no, gen too old

>3080 20GB
yes, ampere is fine

>3090
best non modded card to get for cheap if you can currently
>>
if I start a gofundme to buy me a 6000 pro, you guys will contribute, right?
>>
>>109623363
great now wake me up when a good ui is available for sdcpp. they're all abysmal dog shit.
>>
After two years I have re-enabled generation previews.
>>
>>109622834
Something that no one has generated before
>>
Is there any reason to keep vanilla Klein 9B around if I have the KV version?
>>
File: 1763885991060836.mp4 (2.64 MB, 640x960)
2.64 MB
2.64 MB MP4
>>
File: 67524895801712.mp4 (3.54 MB, 864x608)
3.54 MB
3.54 MB MP4
>>109623346
It's I2V, but it cuts to a new shot immedietly, probably cuz of the prompt.
>>
>>109623412
this must be when kusanagi retired to start her onlyfans.
>>
File: 979306593105224.mp4 (3.64 MB, 864x608)
3.64 MB
3.64 MB MP4
>>109623425
She'd be killing it too.
>>
https://files.catbox.moe/pqnux8.mp4
https://litter.catbox.moe/loqyxm.mp4
https://files.catbox.moe/d2kbp5.mp4
https://litter.catbox.moe/qm8rbn.mp4
>>
>>109623335
reminds me of the 80s poltergeist movie
>>
>>109623335
klee-shay
>>
>>109623473
anime with full phonetic mouth animation looks weird
>>
>>109623473
i'll save everyone the time. she doesn't get hit by the train or get into a car crash
>>
finally a generic video swap prompt that works for minimax.

Use <Video_1> as the master performance and scene and audio. Replace the character in <Video_1> with the character shown in <Image_1>.

Match the new character's face, hair, outfit, colors, body proportions, and visible design details throughout the clip. Use its front, and close-up views as one identity reference.

Inherit the original performer's movement, position, pose, rotation, speed, and timing frame by frame. Preserve background, camera path, framing, focus, motion blur, lighting, shadows, and original edit from <Video_1>.

Maintain correct occlusion when the hands, arms, pass in front of the body. Keep the face, hair, clothing, hands, and body shape consistent. Do not add new movement, objects, cuts, text, accessories, or camera motion.
>>
>>109623473
2nd one is comfy.
4th one has good driving discipline.
>>
File: 350761356482833.mp4 (3.63 MB, 864x608)
3.63 MB
3.63 MB MP4
>>
>>109623498
Thanks anon, I will try that one.
>>
>>109623498
><Video_1>
><Image_1>
That's doesn't fit the official documentation, but I guess if it works it works.
>>
>>109623498
this goes in
summary
?
>>
>>109623498
> Replace the character in <Video_1> with the character shown in <Image_1>.
Let me guess, this simple prompt shits the bed if there is more than one character?
>>
File: 1761245645601899.mp4 (362 KB, 864x480)
362 KB
362 KB MP4
>>109623282
Take 2. It turned the basement door into a front door kek
>>
>>109623562
are you editing the prompt or just rerolling? I've found it to be surprisingly good at nuance, it even understands negatives sometimes
>>
>>109623566
For this one I just edited the prompt to say that it's night and that she opens the door with her left hand.
>>
File: flavortown_noaudio.mp4 (1.76 MB, 1184x672)
1.76 MB
1.76 MB MP4
I'm not very creative, I just like redoing funny stuff from the internet as videos.

https://files.catbox.moe/oy5jo4.mp4
>>
>>109623571
that's a start
>>
>>109623498
Doesn't work.
>>
Does anyone else have problems with the ref model and color preservation? When I use an image as the last or first frame, it completely degrades the colors, darkening the original image. I don't know if it's something with the prompt, or Comfy Kitchen, or something else.
>>
Bro the comfyui devs HAVE to be personally bankrolled by Nvidia's Jensen Huang himself
>have the option to add vulkan support
>would make every single gpu on the planet capable of diffusion
>literally refuse to do so for ((((mysterious))))) reasons
>insist on using CUDA for everything and throwing ROCm a bone once in a while when it gets too obvious you're being bribed by Nvidia
There is literally NO REASON for them to behave like this when stable-diffusion.cpp already exists, the blue prints ARE LITERALLY RIGHT THERE
>>
>>109623587
yes, in every model ever even paid ones
>>
>>109623601
fork it and make gpt add vulkan
>>
>>109623601
wouldn't really be a problem if amd weren't in cahoots with nvidia and just made cards that didn't suck dick for genning
>>
>>109623601
It's not just comfy, everything is like this. Nobody voluntarily uses Vulkan in llama.cpp when they could use CUDA either. I don't know much about gpu programming but either Vulkan is shit for compute or somehow everyone is using it wrong (including amd who keeps trying to make rocm happen)
>>
>>109623498
example (just add the music back after):

https://files.catbox.moe/m5vsh3.mp4
>>
>>109623601
didn't read but just set claude fable or openai sol onto your local comfyui to do updates now, don't bother pulling
>>
https://files.catbox.moe/ffn2nz.mp4

k good it works

Use <Video_1> as the master performance and scene and audio. Replace the character in <Video_1> with the character shown in <Image_1>.

Match the new character's face, hair, outfit, colors, body proportions, and visible design details throughout the clip. Use its front, and close-up views as one identity reference.

Inherit the original performer's movement, position, pose, rotation, speed, and timing frame by frame. Preserve background, camera path, framing, focus, motion blur, lighting, shadows, and original edit from <Video_1>.

Maintain correct occlusion when the hands, arms, pass in front of the body. Keep the face, hair, clothing, hands, and body shape consistent. Do not add new movement, objects, cuts, text, accessories, or camera motion.
>>
File: file.png (9 KB, 540x108)
9 KB PNG
bless Gemma's innocent heart
>>
File: ComfyUI_00010_.png (1.53 MB, 1448x1088)
1.53 MB PNG
>109623662

release the pc build already retard
>>
>>109623645
With Vulkan you're mostly missing out on vendor specific optimizations, and there's no reason forgo them, if you have the correct hardware. Only reason Vulkan is or used to be a decent option for AMD people is ROCm being complete garbage.

That's also why you'll probably never see any more official pytorch Vulkan wheels. The few people who could benefit from that don't really matter to anyone, and the performance will most likely be subpar anyway.
>>
Is there any specific way to prompt for krea? I'd like to make an enhancer like I did for h3.
>>
>>109623360
>upscale with model
>upscale method: nearest exact
Huh? This makes no sense. Is it doing a 2x upscale twice? Just use the regular upscale with model node.
>>
>>109623377
>if you can currently
That's the thing, for the cost of a used 3090 I could get two 20GB 3080s from China.
>>
>>109623662
hmm, I wonder if it has to be a full body shot, a headshot didnt work for a diff gen.
>>
>>109623761
you might as well wait until early next year and the nvidia refresh in that case
>>
>>109623767
*update, I think it's because the leek guy was just on a white background, maybe it fucks up when the character to swap isnt isolated.
>>
>>109623761
My risk appetite wouldn't be high enough to order one or possibly two bubba'd GPUs from overseas.
>>
>>109623791
it does that when i background remove instead of mask, too
>>
File: 83585387358272.jpg (991 KB, 1200x1800)
991 KB JPG
>>
>>109623816
they look delicious
>>
>>109623837
lotta steaks on those heifers
>>
File: 1775629224825501.jpg (23 KB, 678x452)
23 KB JPG
wtf now it works again

so the key for video swap is...masking?

https://files.catbox.moe/9re5wq.mp4
>>
Last I left off I was using illustrious models. Is Anima genuinely better, eg the same community eg: Nova IL vs Nova Anima for example when its similar from the same author?

IL is like 6-7GB and Anima is 3.9GB, I have a hard time thinking it could be better

Krea 2 has 1 anime model - Cat Tower from what I can see?
>>
File: 999999888.jpg (1007 KB, 1133x1700)
1007 KB JPG
>>
>>109623856
Just read up on how to use it and try it out for yourself.
>>
>>109623865
shes going to fuck up the leather
>>
>>109623856
I have fully switched over to anima from IL. Has an easier time with anatomy and perspective. Not perfect, though. The ability to describe more complex scenes in natural language also helps with getting what you want.

Just give it a try. It's fairly small download, even with the seperate VAE and clip.
>>
AnimaXL when
>>
File: kuro_dress.png (2.87 MB, 1900x1964)
2.87 MB PNG
Does anyone know how corporate image editing models like Gemini and Grok work, and if there are any local alternatives?
It's actually impressive how you can give them a sketch or lineart and get a fully-rendered piece in return, and how they can change details small or big, including poses, or just take a reference image and replicate the character and/or artstyle.
Last time I messed with imagegen, making changes like that would require some heavy manual editing before inpainting, but with these you just tell them what to do and get your magic image edit within seconds. The only problems are the resolution limits, the decaying quality if you make too many subsequent edits, and the fact that they're filtered, so I'm wondering if there's anything similar that you can use on your own machine.
>>
Judging by the lack of loras, I'm guessing H3 is not easy to train compared to WAN/LTX
>>
>>109623936
the secret is chaining language models and image models together, be it cloud or local. thats it
>>
>>109623946
95% of loras are made by the poorest 10% of this community
they are too poor to run h3 so there is nobody left who can do a lora
>>
>>109623936
https://huggingface.co/collections/black-forest-labs/flux2

>>109623946
It's brand new
>>
>>109623954
huh, im pretty sure all the big lora creators on civitai use runpod or rent gpu's anyway
>>
File: 1756878309118011.png (164 KB, 1104x618)
164 KB PNG
>>109623846
there. now we are cooking. a node for everything.
>>
cozy breas
>>
>>109623865
share a wf anon
>>
File: 9498838383.jpg (1.38 MB, 1900x1068)
1.38 MB JPG
>>109623898
you are right. Next time, she won't be allowed inside
>>
Is the kinoplexitorium open?
>>
>>109623964
also 193s cause I had to download the model kek
>>
>>109623846
>https://files.catbox.moe/9re5wq.mp4
This will take us to $50 million rupee marketcap
>>
File: output.mp4 (1.5 MB, 864x480)
1.5 MB
1.5 MB MP4
>>
>>109624008
>doesn't rip a fart
in the trash
>>
her good sirs
have they figured out how to fix H3 vagene yet?
>>
File: file.png (37 KB, 324x361)
37 KB PNG
>>109621492
>Sage attention + H3 Cache
H3 Cache comes from the UtilsCollection by silveroxide.
Sage attention you'll probably have to build yourself. You need at least 2.2.0

I have yet to find anything faster that allows you to turn it off and get a better version as long as you keep the same steps/resolution
>>
>>109623473
>humming a melody on top of the song
That's cool
>>
>>109623969

this is /ldg/. this is pretty evidently not possible with local diffusion
>>
I updated comfy and H3addguide node is gone?
>>
File: 358658.webm (2.73 MB, 320x320)
2.73 MB
2.73 MB WEBM
did space shuttle anon finish his kino?
>>
>>109621359
How's AMD GPU local AI coming alone?
Asking as a complete tech illiterate when it comes to this shit. There's some... uhh... e-experimenting I want to do with a certain photo.
>>
>>109624050
>>109621492
>Sage attention
Yeah it's great if you don't care at all about quality to the point where you might as well just lower the resolution to save time.
Just use comfy kitchen attention instead, it doesn't destroy the video and still reduces prompt times by a total 35%.
And turbo 600 lora at 8 steps for semi decent audio.
Also don't use any cache/memory patcher nodes, they also suck massive dick turning the video into pixelated garbage
>>
>>109621492
I currently use comfy kitchen attention, H3 sparse attention
Turbo lora if you want to run at low step count
>>
>>109624100
bait
>>
kek it worked with ash

https://files.catbox.moe/ekmbce.mp4
>>
File: MiniMax_H3_00304_.mp4 (2.51 MB, 864x480)
2.51 MB
2.51 MB MP4
>>109624026
>>
well this one turned out...mostly ok

https://files.catbox.moe/04cdd4.mp4
>>
File: jeeeeeeeeeeej.png (269 KB, 432x621)
269 KB PNG
>>109624146
>>
>>109623946
H3 does well with a start frame, and can be good with reference 2 video. So you can generate what you want from another model like IL and animate it in H3
>>
>>109624146
11/10
>>
>>109624146
collage worthy
>>
>>109624146
drop the json
drop the prompt
drop somethin bro onegai
>>
File: lopunny.webm (3.02 MB, 864x480)
3.02 MB
3.02 MB WEBM
This looks like it'll work well once I get into it properly
>>
>>109624175
I know that. But sometimes you don't want to rely on reference videos because that kills creativity. A lora ensures identity consistency and/or concept materialization.
>>
File: sj-sr_00001_.webm (557 KB, 736x416)
557 KB
557 KB WEBM
>>
>>109623812
I haven't had an issue with AliExpress in the past. For some godforsaken reason all the US GPU sellers on eBay charge $300+ shipping to my country, so buying the standard stock of shit from there (plus the exchange rate) is just uneconomical.
>>
>>109624185
https://docs.comfy.org/built-in-nodes/MiniMaxH3AddGuide

It's a messy workflow, based on this node
you feed 22 last frame of the video at frame_idx = 0 to extend video
>>
>>109624191
Reference frame not video, you can create the reference, or you can just use it loosely from anything to push into a certain direction for consistency. Omni Flash seems to have the best understanding of identity and concepts with or without a reference, really solid for that, but its online only and kvetch's a lot.
>>
>>109624146
>>109624197
>>
File: 6345277272.jpg (133 KB, 1087x534)
133 KB JPG
>>109623968
after I fix the gore
>>109624060
I'm not using api nodes. Just krea
>>
comfy niggas be like
>>
I envy that ani only has to deal with his own bullshit while we have to deal with comfy's bullshit
>>
>>109624187
I'd like to see those hyena's necks snapping when she kicks them. Also blood splashes from their mouths,
Post again once you have fixed it.
>>
>>109622250
>which explains a lot
What does it explain?
>>
>>109624350
obsession with anis bwc
>>
>>109624350
Buck broken
>>
>>109624267
I hate nodes so fucking much.
>>
>>109624375
why do you use them then?
>>
hurry the FUCK up ani
>>
>>109624267
not 100% on what i'm looking at here. edit lora>gen>edit lora inpaint?
>>
>
>>
>>109624399
spaghetti because half of this shit is not needed. Workflow is simply gen - upscale - clean up noise - upscale faces. I wanna add inpainting too
>>
File: debo_wr_k2_00007_.png (2.92 MB, 2048x1101)
2.92 MB PNG
>>
>>109624443
agh.. what is that thing
>>
File: ooh.gif (954 KB, 498x370)
954 KB GIF
>>109624429
>spaghetti
>>
>>109624387
Because every other option right now sucks even more
>>
>>109624485
my time with wan2gp has been nothing but rainbows and sunshine



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.