[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: collage_1785905006.webm (3.16 MB, 2048x1228)
3.16 MB
3.16 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109460446

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
gm saars
>>
https://www.reddit.com/r/StableDiffusion/comments/1vfwijz/minimax_are_issuing_takedowns_on_decensorexplicit/
Why Minimax, WHY????
>>
blsd trd
>>
we need a proper ejaculation lora
https://files.catbox.moe/uegmjm.mp4
>>
Has anyone had success with animating manga panels?
>>>/gif/30996965
>>
Blessed thread of frenship
>>
>>109465146
Thanks for the bake anon
>>
>>109465168
I don't want to be friends with the schizo rentey author
>>
thanks for bakering
>>
File: kek.png (2.79 MB, 2944x1648)
2.79 MB PNG
>>109465164
that's Pam?
>>
File: oilrod.jpg (209 KB, 1920x1080)
209 KB JPG
finally a good bake
>>
>>109465181
THIS
>>
>>109465168
>>109465174
>>109465184
I don't think we should thank schizo trolls
>>
No still images in the video general please
>>
>mfw Resource news

08/04/2026

>stable-diffusion.cpp adds support for MiniMax-H3
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md

>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES time
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
https://github.com/IntMeGroup/MIEScore

>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
https://rathgrith.github.io/PeCA

>Kandinsky WM 1.0: A family of models for Physical AI
https://github.com/kandinskylab/kandinsky-wm

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

>Raylight 1.7.2 Adds MiniMax Support, 2x Speedup
https://github.com/komikndr/raylight/releases/tag/1.7.2

>Scaling Properties of Text Conditioning in Visual Generation
https://heheyas.github.io/context-scaling

>Retrieval-Driven Training-Free AI-Generated Video Attribution
https://github.com/renxi-seu/Video_Attribution

>A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
https://github.com/zfu006/SSG

>ComfyUI MiniMax H3 Image Studio (Experimental)
https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio

>MiniMax H3 — NVFP4 (Blackwell)
https://huggingface.co/lilcheaty/MiniMax-H3-NVFP4

08/02/2026

>MiniMax H3
https://huggingface.co/MiniMaxAI/MiniMax-H3

>MiniMax H3: Repackaged model files for ComfyUI
https://huggingface.co/Comfy-Org/MiniMax-H3

>MiniMax-H3-INT8-CONVROT
https://huggingface.co/Gluttony10/MiniMax-H3-INT8-CONVROT

>MiniMax H3 COMMUNITY LICENSE AGREEMENT
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE

>LoRA Dataset Studio: LoRA workflow in one tab
https://github.com/perfectgf/lora-dataset-studio

>comfyui-vram-tracker
https://github.com/PuppetMasterAI/comfyui-vram-tracker
>>
>>109465181
>>109465192
>exactly one minute apart
your samefagging is more and more sloppy anifart kek
>>
>>109465186
yeah its what I prompted. made her older than the show, though. it also gave me a weird lump where my pubes are for some reason.
>>
>>109465187
schizo troll rentry bake is never good
>>
>mfw Research news

08/04/2026

>Token Radius Attention for Efficient Video Generation
https://arxiv.org/abs/2608.02504

>CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
https://hanxjing.github.io/CultureVidBench

>EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
https://arxiv.org/abs/2608.02474

>Investigating Social Bias in Narrative Image Generation
https://arxiv.org/abs/2608.01780

>CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds
https://arxiv.org/abs/2608.00674

>Diagnosing Under-Development of Irreversible Processes in Video Generation
https://arxiv.org/abs/2608.00617

>Where Does Generative Difficulty Reside? An Empirical Study of Target Representations
https://arxiv.org/abs/2608.00626

>UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction
https://arxiv.org/abs/2608.01298

>One-Sided Quantile Coupling for Flow Matching
https://arxiv.org/abs/2608.00978

>MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning
https://arxiv.org/abs/2608.01521

>ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport
https://arxiv.org/abs/2608.00769

>CoT-Edit: Let CoT Guide Instruction Video Editing
https://arxiv.org/abs/2608.01113

>Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
https://arxiv.org/abs/2608.00716

>A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2
https://arxiv.org/abs/2608.01258

>CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
https://arxiv.org/abs/2608.01644

>DiffPrune: differentiable information throttling for token pruning in vision-language models
https://arxiv.org/abs/2608.01985
>>
>>109465201
>gave me
>>
>>109465164
I'm really close to trying out h3
>>
File: 669.gif (3.86 MB, 400x532)
3.86 MB GIF
>>109465200
>anifart
>>
>>109465209
yeah its my dick nigga
>>
>>109465200
thanking yourself for baking the schizo rentries is underage
>>
>mfw MORE Research news

>IDraw: Artist Verification from Digital Drawing Images
https://arxiv.org/abs/2608.01737

>Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation
https://arxiv.org/abs/2608.01978

>UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
https://tanliming-daniel.github.io/UniMoCa

>MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration
https://arxiv.org/abs/2608.01829

>SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
https://arxiv.org/abs/2608.01306

>Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation
https://arxiv.org/abs/2608.00562

>Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution
https://arxiv.org/abs/2608.01823

>Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression
https://arxiv.org/abs/2608.02109

>SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching
https://arxiv.org/abs/2608.01990

>DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models
https://arxiv.org/abs/2608.01821

>Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
https://arxiv.org/abs/2608.01848

>SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation
https://arxiv.org/abs/2608.01977

>Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
https://arxiv.org/abs/2608.01314

>ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models
https://arxiv.org/abs/2608.01067

>RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation
https://arxiv.org/abs/2607.09757
>>
please put that news in a rentry omg
>>
>>109465212
It's his boyfriend he loves posting about. very tsundere of him
>>
File: keekekekekkkekek.png (1.17 MB, 864x1184)
1.17 MB PNG
>>109465213
>yeah its my dick nigga
oh hell naw cuh what the heeeeeeeeeeeeeeeeelll what bruh doing nga keeeekekek nah JSID already this place is cooked fr
>>
>>109465219
why don't you? you have time to go through desuarchive and screencap schizo allegations already
>>
What do you want to name the new general?
>>
every time we get a good bake it's the same thing. ani just starts baiting everyone just so we get to bump limit faster.

Literally just stop fucking engaging with bait.
>>
>>109465226
Yeah, my dick isn't brown I don't have anything to hide.
https://files.catbox.moe/01prq3.mp4
>>
>>109465169
>it's probably a farce to justify releasing such an uncensored model
>look, we totally care about safety!
let's hope he's virtue signaling for just a few day and let he let us make some coom kinos on civitai in peace then
>>
idk, is Pam sucking your dick? I think anon is more based than us right now.
>>
>>109465237
>every time we get a good bake
I can't remember the last time we had one. It's either a rentry obsessed schizo or a troll that bakes for over a year
>>
>>109465187
I appreciate that she has the strong JenCon brow.
>>
>>109465238
>my dick isn't brown
it is circumsed though
>>
>>109465253
I am for medical reasons.
>>
>>109465253
Don't remind me.
>>
>>109465195
>no comfy screenshots
oh no
>>
>>109465228
i would if i was the one who posted the news links
>you have time to go through desuarchive and screencap schizo allegations already
?
>>
File: 1782283950636296.png (319 KB, 1513x1247)
319 KB PNG
is he allowed to do takedowns on models that aren't his? like he doesn't own Qwen it's Alibaba's model
https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot/discussions/7
>>
>>109465262
it's probably automatic since it had minimax in the name
>>
>>109465262
If someone had the power to take it down they would not make a measly request
>>
>>109465260
the only one who complains about the news is the schizo. literally no reason to complain about two posts you can scroll past
>>
>>109465270
sometimes it's good to stay polite for good PRs
>>
>>109465262
at this point is just looks automated?
>>
Very nice, Minimax is a beast
https://files.catbox.moe/1qf7wy.mp4
>>>/wsg/6208249
>>
>>109465272
>two posts
you mean three full character posts
no need to start calling people names :]
>>
>>109465286
>last frame
I look like this
>>
>>109465262
You think they're regretting releasing the open weights kek?
>>
>>109465301
no why would they?
>>
>>109465301
what did they expect seriously? they made a model that was trained on fucking porn lool
>>
File: debo_sc_k2_00004_.png (1.82 MB, 1872x1007)
1.82 MB PNG
>>109465287
three posts is rare in the summer. there was a lot of preprints today for some reason. they weren't full character though
>>
When you should have written it in Rust
>>>/wsg/6208338
>>
anyone here using the Q4 quant for the text encoder? wondering how good it is for prompt adherence
>>
>>109465309
why did you get so upset over anon asking you to use a rentry link? at least the lmg frens are nice enough to keep it short and in OP
>>
File: 1782847496525849.png (2.15 MB, 1957x1057)
2.15 MB PNG
>>109465322
>only 2 rows of letters
poor Spongebob, hard to code like that :d
>>
>>109465081
Thanks anon, shifting the audio sigmas did the trick as you said.
>>
>>109465304
>>
>>109465332
never go under 8bit for the TE it destroys everything
>>
>>109465348
let's not he's not that retarded and is just virtue signaling to not be destroyed by the media or something lol
>>
So what is the meta for fast gens now?
>>
>>109465348
lol.
>>
>>109465363
Spectrum is the only one that doesn't destroy the quality imo
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
>>
>>109465363
wan2gp
>>
File: Angry-Face-Emoji.png (45 KB, 1200x1200)
45 KB PNG
>>109465348
kek
>>
>>109465348
Might as well do what he says and pretend that we are using some other cloud only APIs to ruin their reputation and kill cloud services for good.
>>
/ldg/

THE place for image and video gen
>>
>>109465351
okay.... i might still try to see what happens ;P
>>
>>109465348
>please don't use the gun I made to kill people
when did that even work?
>>
>>109465386
By larping as terrorists like a certain 3 letter agency?
>>
File: 1783028080052563.png (132 KB, 1606x493)
132 KB PNG
https://xcancel.com/RyanLeeMiniMax/status/2084558818902753487#m
looool
>>
>>109465397
my hero
>>
>>109465397
>your countries are cucked so we have to cuck the license
based
>>
Did I miss any useful NSFW tune/lora/whatever in the pasts 10 hours?
>>
>>109465397
I don't get him, why won't he simply go for Alibaba Wan's licence? that one allows for coom loras on civitai, they're from the same country!! If alibaba can do it, he can do it
>>
>>109465421
VERBOTEN
>>
>>109465397
>this license is nonsense because in order to do any of this you must violate every ip and copyright law possible and disregard the opinions of anyone
>however, this license explicitly states you cannot use mickey mouse as per disney copyright and ip ownership, and as per the same license, you are forbidden from generating mickey mouse with a gun as he shoots his cock off with remarkably comical detail as whinnie the pooh watches speaking fluent mandarin impersonating president xi, as china outlaws such depictions
>>
>>109465421
>>109465108
>>
>>109465146
>80 post in 41 minutes
>at this hour

>>109465237
>just so we get to bump limit faster.
that part at least is true
>>
>>109465348
so kind of like don't go over the speed limit in this car that can go well over the speed limit. Legal and PR thing they have to say.
>>
>>109465421
model was released yesterday and this nigga already out of gooning, he already needs his new fix
>>
>>109465435
Thanks for sharing/helping me catch up.
>>
>>109465445
glad i am a no fapper
>>
I'll try it today, hopefully that comfyui portable works
>>
>>109465421
nah
i gave the pen lora a shot and it fucked everything up
>>
>>109465445
yeah, and?
>>
>>109465348
Civitai.red has Minimax porn loras up already, if they were serious about shutting NSFW down that would be the first place they contact

This is just scaremongering
>>
File: debo_sc_k2_00008_.png (2.44 MB, 1872x1007)
2.44 MB PNG
>>
>>109465397
Okay, this guy is definitely salty at closed source models and media companies trying to control everything.
If this isn't a signal for the community to act right during these times, then we deserve to lose access to future OS frontier models.
Unfortunately a great majority of this community can't read the room.
>>
pretty convinced we'll get good, minute long porn gens coming out very soon. No fetish will be spared.
>>
>>109465482
Make sure to tag all your porn gens Flux 3
>>
>>109465466
They stance is clear, go have fun but don't point at us.
>>
>>109465425
1. They are not as big, powerful, nor rich as alibaba
2. Alibaba is not being sued at this very moment over ai video gen by pedowood
3. H3 is a huge step forwards, normgroidgoyim need time to get used to new ai capability when it comes out otherwise they will shortcircut about "misuse" of making an image of a woman she posted online herself take off her bikini but in better quality now, thus RAPING the woman
>>
>>109465498
>1. They are not as big, powerful, nor rich as alibaba
I thought every chinese companies were protected by the CCP or something
>>
>>109465485
based
>>
how much closer are we to making a long-form movie?
>>
>>109465485
I will do my part.
>>
>>109465503
probably a case of them all being equal but some more than others
>>
>>109465509
Nolan style? Probably a few days, just gotta stich 5 sec gens together
>>
>>109465509
you can already do that but it's annoying to wait a long time to look for good seeds for each scene you want to do
>>
>>109465348
They knew exactly what they were doing when they released it, it's not an acccident that the model is so good at NSFW kek.
>>
>>109465509
If you're serious, you can stitch shots together. Most movie scenes are less than 15 seconds each on average.
>>
>>109465514
>>109465522
>>109465538
how do you make gens naturally follow each other? there's something better than flf?
>>
turbo lora eta? 20s/step is killing me
>>
How long until someone removes brie from captain marvel and replaces her with Britta?
>>
>They're taking down NSFW Loras for minimax
Kek. Get fucked coomers
>>
>>109465541
>how do you make gens naturally follow each other?
with images, you do first frame, last frame, and then you use edit model to go onto another scene or something, that's why he did to make that movie
https://www.youtube.com/watch?v=fyZhC2TXgcs
>>
>>109465541
directly continuing a clip? not sure if anyone figured that out with h3 yet. someone has to add frame injection support like with ltx
>>
>>109465551
doesn't matter, we have it.
>>
>>109465561
omg where can we find it
>>
usecase for H3?
>>
File: MiniMax_H3_00139.mp4 (1001 KB, 832x480)
1001 KB
1001 KB MP4
>>
>>109465566
Brie Larson?
>>
The few Ideogram nsfw loras survive on civit by being uploaded under other intead of the model tag. Maybe H3 loras can do the same
>>
File: H3_Combine.jpg (195 KB, 1470x888)
195 KB JPG
>>109465541
Continuously combine First/Last frame scenes together with transition shots. I am using Ref model, but the F/L model should work too.

https://streamable.com/up4knd
>>
>>109465551
I stand behind this opinion.
>>
>>109465565
nobody even takes your bait bro why spam more than trAni at this point?
>>
Sorry brainlett ,in h3 can you use several references , is it just a simple case of adding another box in comfy ? Thanks
>>
>>109465581
Your punctuation is giving me a fucking stroke.
>>
>>109465570
damn this is good, remind me of the LTX days
https://files.catbox.moe/xc6ta3.mp4
>>
>>109465588
Sorry indian , english not my first language also am dalit
>>
>>109465570
The transitions are obvious but that's really really impressive
>>
>>109465591
I don't even care if you're lying or trolling or telling the truth. You being Indian from the start is my expectation.
>>
>>109465591
rape Your Mother, your sister. Brother . you bloody
>>
>sharing my catbox with indians
absolutely fucking not
>>
>>109465561
Until they sue your ass. If you want to coom to celebs, fine. Just don't upload it.
>>
>>109465434
> this license is nonsense because in order to do any of this you must violate every ip and copyright law possible and disregard the opinions of anyone
not really the case. many countries will decide that LEARNING from something is not relevant to whether anything is copyrightable at all

how far the right to copyright-able characters then goes in *production/publication* and whether there are "fair use" / educational / archival / fan art / other official limits to the scope of copyright was usually the actual fight?
>>
>>109465581
Sorry working , can I add additional references in h3 videos? Do I just add more nodes.
>>
>>109465592
I need to work on passing transition. No, not a as tranny. lol
>>
>>109465570
>mouth hanging open for way too long
>smooth to on 2s to smooth
>>
>>109465498
>1. They are not as big, powerful, nor rich as alibaba
I'd say they're pretty rich, you have to if you want to make such a complete dataset that was used to train Minimax
>>
>>109465601
Heres my response to you from the stance of someone who has spent most of my life pirating, and violating copyright law.
Nobody cares, people will do it, it would be really nice if the profiting of stuff was penalized rather than just giving the finger to companies but unfortunately that is not the world we live in. We have tremendously large companies getting away with training ai shit off shitloads of pirated or stolen material without authorization, and some trying to do so with proper approvals. In the end it doesnt matter, it is impossible for any of this shit to be governed in any real way unless every single possible thing going in was properly approved for this use and derivative use and that simply wont happen for any quality result.

What I'm saying is anything from these companies involved with (Anthropic, OpenAI, Google, Apple, Elon/Grok) should be free, open, and not restricted even if it realistically requires datacenter tier hardware infrastructure to use it.
>>
>>109465619
enough time passed that youtube scraping pipelines are there
>>
>>109465581
indeed sir, I gen myself and my bride making the love sir
>>
>>109465623
not just the scraping, but they made really accurate captions, the model is a beast at prompt following, there's no way everything was automated they had to hire people to do this
>>
>>109465632
huge and much better TE and VLLMS
>>
https://www.reddit.com/r/LocalLLaMA/comments/1vfujnc/chinas_openweight_models_will_be_spared_us_safety/
>>
in comfyui is there any way to see your scheduler preview again after it goes away from a tab change/ refresh?
>>
>>109465632
>there's no way everything was automated they had to hire people to do this
didn't alibaba showcase their qwen vision model being able to fully understand the world through a camera? they must have used that for captioning
>>
>>109465645
Stop using chrome like a gay faggot
>>
>>109465627
Can you add several references my bastard ?
>>
>>109465646
bruh alibaba's model aren't nowhere as good on prompt following nor do they have all those IPs, what Minimax did is a far superior job
>>
>>109465642
based orange man, I voted for that!
>>
>>109465649
calm thy tits im on firefox
>>
File: 0241325.mp4 (3.41 MB, 1120x832)
3.41 MB
3.41 MB MP4
The happy ending.
>>
>>109465642
>>109465657
zion don knows nothing about tech but the globohomo advisors know that if they ban companies from using chinese ai that wont stop all regular people and many companies doing it internally/in secret, fucking over only american companies who now have to pay premium for ai while the rest of the world will just use the chinese one, theres nothing benevolent in this, banning chinese ai is not enforcable and will only fuck their own tech sector.
>>
>>109465642
that's definitely a good news, but the US is nerfing themselves by doing that, now only the US models will have to respect a threshold of cuckoldery
>>
File: 0.mp4.webm (806 KB, 640x480)
806 KB
806 KB WEBM
12 months to the day
T2V wasn't wans strongest feature
>>
>>109465375
I'm using mini claude to generate my vids
>>
>>109465670
the headline says china but its open weights in general. This does fuck over openai / anthropic though
>>
>>109465262
this int8 meme is for old cards or something? those that cant do modern quants?
>>
>>109465684
int8 convrot is the best quant but abliterated / uncensored TE's are a useless meme. They only hurt it.
>>
>>109465642
>we cant force them to safety test
roru rumao
>>
File: 1764410790491399.jpg (2.86 MB, 5726x3072)
2.86 MB JPG
>>109465684
wait what, no, we have a new thing, int8 convrot, that thing is the same quality of Q8 while being 2x faster than fp16
https://github.com/BobJohnson24/ComfyUI-INT8-Fast/blob/main/Metrics.md
>>
>>109465697
its not same as q8 but the loss is negligible
>>
>>109465262
I want the snake oil salesmen that push that uncensored TE BS to suffer more than I want people to not be able to enforce that
>>
File: 1766410401149554.png (278 KB, 2144x1432)
278 KB PNG
>>109465701
on some models int8 convrot beats Q8
>>
>>109465697
>>109465692
better than fp8 and bf16?
i though i saw on comfy repo that it was made for amd and old cards that are slow when using common quants
>>
what kind of speeds are you guys getting in Minimax?
I'm on a 4070, 12GB VRAM, 32GB RAM--
takes about 20 minutes to generate 10 seconds of video at 0.5MP, 16:9 Landscape.
>>
>>109465668
>Congratulations on your transition
>>
>>109465707
hmm, surely q8 wasnt created with a proper recipe then, it cant be that much worse, especially than any fp8
>>
>>109465711
Get sage and you will cut that down to half
>>
>>109465697
>>109465707
just clicked on the images
kk see it
>>
>>109465708
>better than fp8
yep, better quality, 2x faster
>bf16?
bf16 is the original lossless model you can't beat that
>>
>>109465717
use SOL AND sage and it will be even faster. This is faster than LTX was now. And we should get 4 step lora AND sparse attention prob this week
>>
https://files.catbox.moe/h3j8lu.mp4
>>>/wsg/6208362
>>
>>109465719
>yep, better quality, 2x faster
only on 30xx which doesnt have native fp8, although int8 is faster than fp8 on 40xx 50xx also but much less
>>
>>109465711
2:30 for 5 seconds at 0.4 on 10gb vram 32gb ram
>>
>>109465719
saw the images got it ty
if it is possible to convert FP8s into INT would be great
i will look into it seems very cool
>>
>>109465731
>only on 30xx which doesnt have native fp8
and the 20xx too? so basically 90% of users lol
>>
File: MiniMax_H3_00148.mp4 (1.57 MB, 1152x640)
1.57 MB
1.57 MB MP4
>>
>>109465736
i would assume most have 30xx at least by now but yes
>>
>>109465731
Even on 5000 series its 40% faster. Its also MUCH higher quality.

If you DONT get such a speedup that means you have outdated pytorch+cuda+comfykitchen / sageattention2.2.0+
>>
>>109465734
>if it is possible to convert FP8s into INT would be great
no, you need the bf16 model to convert into int8 convrot
>>
File: fr.png (132 KB, 550x535)
132 KB PNG
just say the resolution bruh. ion no what mega pixels is
>>
I still can't get fucking face swapping to work with a reference video. I followed prompting example shown in the docs.
>>
>>109465751
it's not that complicated anon, for example a 1920x1080 = 2073600 = 2.07 MP, it's better to talk about the total of pixels, that's the relevant metric to use to evaluate how many time you're gonna wait
>>
>>109465717
>>109465721
what kind of speed increase / quality decrease are we lookin at?
>>
>>109465758
say 1080p. no one here generating unstandard resolutions
>>
File: rss.png (17 KB, 373x303)
17 KB PNG
>>109465751
>>109465758
lol dont give out such bullshit explication, you all using megapixels because of the comfy default WF lol
>>
>>109465768
0.5 MP 124 frames is 1 min on 4090
>>
>>109465756
i gave up, it's basically rolling the dice and you have 5% success rate.
>>
>>109465774
oh, i was wondering why all you redditors were saying that. CHADS are using our own workflow
>>
>You can make some nice short movies, trailers, music videos, or animations with Minimax and get a shit ton of views for quality ones
>But its license means you can't actually monetize it outside of China, so the model is practically useless
>>
File: realniggas.png (26 KB, 725x291)
26 KB PNG
>>109465781
You can tell right away who are the newbies are here lool
>>
>>109465787
1. u can, unless your an amerifat or european union cuck, in which case u need a license
2. just dont tell anyone which model ur using theyre all similar lmao
>>
>>109465787
Also from what they've shown, they have every intention to be predatory and take down any content and seek compensation from anything that generates revenue and violates the license.
>>
>>109465781
imagine wasting time making a custom workflow, when workflows consist of the exact same nodes and shit 99.9% of the time
>>
>>109465797
My point is that even API models have less restrictions than this "open" model. It's "open" only in disguise, more like a trap.
>>
>>109465793
I'm not a newbie and I switched to the Resolution Selector node
>>
you guys are using ai to generate prompts while i manually tweak and re-generate all day. i'm like a x86 assembly programmer
>>
>>109465808
they are just covering their ass legally while being sued bro, if they wanted to lock shit down the model would have been actually safetycucked, the license is literally not affecting you or most people at all.
>>
>>109465787
>>109465798
They have no power to actually take down your content but you're also a massive loser if monetization is your criterion for model utility
>>
>>109465810
>xe doesn't diffuse his prompt in latent space in his head to gen his 1girls
ngmi
>>
>>109465820
I know my own imagination and I'm tired of it
>>
>>109465238
I wonder why it fucked up the hand
>>
>>109465787

>I invest in AI stocks early
>The more people consoom and goonmax, the more I profit.
>He didn't invest in AI stocks
>He has to help other men cum for beta bux.

ngmi
>>
>>109465742
"eat shit" ?
bad at lip reading kek
>>
>>109465814
How are they getting sued if they have no business in America?

>>109465819
Models that cease being a toy should be primarily used for monetization. You don't actually use the models for coom, do you anon?
>>
File: 087520784072508.png (203 KB, 1275x553)
203 KB PNG
>>109465793
i2v chad here
>>
>>109465853
>comfy
chad status revoked
>>
Hustler saars who can only enjoy something if they can make money from it deserve to be killed
>>
>>109465857
why wish death on comfy like that?
>>
>>109465831
There are no such thing as AI stock, ClosedAI never even started out as a for-profit corp
>>
>>109465853
mother of based
>>
>>109465856
i can't use anything other than comfy, it's too good.
i wish i could put more shit into comfy, i wish comfyui was an operating system.
>>
>>109465856
What are people supposed to use? Forge?
>>
>>109465857
>>109465863
Why don't you work on anistudio instead of shitposting all day on 4chan and then claiming you don't have the time to work on your projects? Lol
How's the e-begging going by the way? Did tagexplorer's "associate" reach out yet? (lololol)
>>
>>109465879
Automatic1111
retvrn to tradition
>>
File: .png (123 KB, 1385x1203)
123 KB PNG
>>109465872
>>109465879
you're talking to animanon, the schizo who FUDs comfy all day because he thinks it'll make people use his dogshit sdcpp wrapper
>>
>>109465882
yeah good luck with that, fartface.
>>
>>109465621
> should be free, open, and not restricted
good idea for new rules i suppose? will your country pass them?

if you thought remembering mickey mouse without having gotten a license was a "life of violating copyright law" that would however be wrong no matter how exactly you remember even every last detail, it's just not copyrightable or w/e.
>>
>>109465570
Try to add the two last frames as a reference instead of one to see if it improves the flow.
I am using only references (not frames from the clips) and I don't think I will need very long shots
>>
>>109465721
Sol Engine or just sol-attn?
>>
>>109465865
Holy retard
>>
>>109465902
Still learning Hhw ref model behaves. Some seams are totally unnoticable, but others are bit jerky.
>>
>>109465721
can you post the entire model loader chain (or does order not matter?)
>>
I wonder how much of an effect using a model already pruned by kijai has on using sparsity based opts but I won't download the unpruned version to test that.
>>
>>109465879
For video stuff Wan2GP is actually pretty good but you lose the ability to make custom workflows.
>>
>>109465164
I'm not gay i just want to s>>109465165
ee a good pussy
>>
>>109465932
i'm personally just going to guess that int8 convrot and/or gguf q8 are very close to the unquantified versions regardless the usage. which yes, i suppose is unproven.

but there are too many possible "speedup" nodes with too many settings to try already.
>>
>>109465721
what repo?
> https://github.com/KingGore/ComfyUI_sol-attn_Blackwell
is for blackwell only
>>
File: h3_00070_.mp4.webm (1.59 MB, 896x1184)
1.59 MB
1.59 MB WEBM
she did that on purpose
>>
>>109465894
if youre going to violate veery fundamental copyright and ip law possible, then the resulting tools and data or model or whatever the fuck they make from it should be open and accessible to anyone, even if theyre too financially poor to utilize it, even if it ultimately takes a huge shit on the same laws violated to make the models to begin with.
im holding the stance that today's handling of copyright and ip law is a joke that should not exist, but at the same time desire FOSS shit (free as in freedom or literal) of the resulting process. in the current handling of american law you can just get your ass handed to you if disney feels like you used the mouse in the wrong way and thats horse shit
>>
>>109465964
why she wear boot sir?
>>
https://civitai.com/models/2834514/minimax-h3-ref2va-advanced-filmmaking-workflow-or-all-speedups-qol-features

interesting reference workflow
>>
i'm trying to create that one scene from the movie Fury
>>
>>109465958
I was referring to stuff like Sol-Attn rather than quantization with sparsity optimization part. I don't have concerns about the quant.
And yes so much stuff, I am still not halfway done all different parameters and combinations.
>>
>>109465963
https://github.com/kijai/ComfyUI-SolAttn_triton
btw i suck cock
>>
>>109465977
looks like a bunch of snake oils I don't need
>>
>>109465466
Hugginface is a research and academic web. The chinese goverment is against porn but they are socialist, this is the clasical chinese of do whatever but not in public
>>
File: MiniMax_H3_00156.mp4 (1.2 MB, 1152x864)
1.2 MB
1.2 MB MP4
>>
File: manykek.webm (665 KB, 640x640)
665 KB
665 KB WEBM
what is required to preserve distant faces
>>
https://files.catbox.moe/kqk8gf.mp4
>>
>>109466005
something in the foreground to distract
>>
>>109466005
More pixels
Higher resolution
>>
>>109466005
>what is required to preserve distant faces
higher res :(
>>
File: 11111.png (115 KB, 739x616)
115 KB PNG
ok im pissed
how do i get out ofthis comfy update launch install fucking loop
2 days ago i bypassed by renaming my comfy folders, letting it update, then deleting the new folders and reverting to my old ones
but now its looping again
>>
>>109466017
>>109466015
ty
at least low res can be used for face reference gens

one shot btw, therefore very good
imagine the finetunes and loras
>>
>>109466030
start by using portable version instead
>>
>>109466030
use portable. now and always.
>>
>>109466030
>desktop version
anon i...
>>
https://civitai.red/models/2835126/minimax-h3-penis-by-coachbate?modelVersionId=3199605
>>
reference model:

A scene in the show south park. <Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

<Subject 1> is standing near <Subject 2> and says "how is austria today, man?". <Subject 2> says "kinda hot, got something in the oven". <Subject 1> says "true and real and true and real."

https://files.catbox.moe/ucjyfh.mp4
>>
Can h3 take reference video and lipsync it with reference audio?
>>
Is there a point downscaling the images or can you just feed H3 full 2k/4k images without any quality disavantages?
>>
https://files.catbox.moe/ty4kcd.mp4
>>>/wsg/6208388
>>
>>109465981
ah ok. if you have candidates for something that seems to work pretty well, let us know.

>>109466030
the most common is to just use the portable version or pulled from git installed like any dev would

install via stability matrix or pinokio may also work. note the earlier has a folder structure of configured/symlinked folders it uses that many people here won't know even if some other stuff is exactly the same... but it's not super complex.
>>
reference...sorta worked, kek

A scene in the show south park. <Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

<Subject 1> is drawn as a south park character and standing near <Subject 2> and says "I need some ideas for my stream, man.". <Subject 2> is drawn as a south park character and says "here's a good one, play games and stop reacting to politics, faggot. When I was your age I was piling up schlomos sky high.". <Subject 1> then walks away.

https://files.catbox.moe/oh2ejt.mp4
>>
>>109466064
cringe as fuck, I literally felt embarassment
>>
my prompt is too big. i need to use a dedicated text editor to keep working on it
>>
>>109466062
yes. (as a basic statement). i didn't try it much yet and IDK how well it behaves if you pass many other references or with prompts etc. but it can do it.
>>
>>109466102
The more stuff you have in the prompt the more likely it is something will get ignored.
>>
>>109466099
gonna test this vs my non easycache workflow (the base template, with the spectrum node stuff)
>>
>>109466063
haven't noticed a problem but since all my actually large images are jxl and they actually seemed to crash the software i haven't tried anything particularly large

maybe it just downscales anyhow?
>>
>>109466110
i think that's only true if your frame count is too short
>>
>>109466102
use timestamps.

0 to 3s: the man grabs a soda. 4 to 7s: the man drinks the soda and throws it in a trash can.

etc, if you want specific directions this can make it work better
>>
>>109466116
also if you have a 10s prompt and brief dialogue, it will generate jibberish speech if your prompt is too vague. if nothing happens for 8 seconds, specify what happens.
>>
>>109466116
i had to remove the timestamps since it caused frame cuts
>>
>>109466112
and here we go, I think spectrum wins with the more steps and similar time to gen

https://files.catbox.moe/v61ghz.mp4
>>
>>109466122
you also have to specify the shot, retard
>>
>>109466127
yo mama
>>
>>109466128
im givin u a helpin hand but u insult? faggot brownoid behavior saar
>>
File: 1781460379686882.png (119 KB, 1628x471)
119 KB PNG
Localchads got the VIP pass. APIkeks got TSA.
>>
>>109466090
Spectrum is the only thing I can confidently vouch for now. No idea how well it combines but it's very good on its own. Didn't bother with parameter tuning that since the default profile works well for me.
>>
>>109466136
we won so hard
>>
>>109466136
>30 day review
for insider trading purposes, what a fucking joke.
>>
so its time to upgrade pytorch? :(
>>
File: output.png (449 KB, 736x432)
449 KB PNG
>>109466136
>>
okay now I am getting the hang of reference model prompts. ie: <Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

then just reference them as <subject 1> etc.
>>
>>109466064
Awesome as fuck, I literally felt enlightenment
>>
>>109466131
>u insult?
>>109466127
>retard
>>
>>109466152
proof, got this KINO fight with this simple prompt.

<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "I can't breathe!". <Subject 1> says "i'm literally me." and puts on black sunglasses.

https://files.catbox.moe/jf9sdq.mp4
>>
>>109466131
you insulted me first, sir
>>
>>109466139
so for example no sage attention 2/3 implicitly or explicitly via nodes?
>>
>>109466154
thanks anon
>>
>>109466152
TIP: copypaste the prompting guide into your favorite AI assistant (e.g. Fagle)

https://huggingface.co/MiniMaxAI/MiniMax-H3/raw/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

then ask it "How do I properly prompt to replace the skateboarding guy in video 1 with the anime girl in picture 1?"
>>
>>109466157
also this is with spectrum node thing and at 0.3mp for speed.
>>
>>109466157
>https://files.catbox.moe/jf9sdq.mp4
he's literally me
>>
>>109466157
round 2, what a time to be alive.

https://files.catbox.moe/4mlbug.mp4
>>
>>109466161
I have only tested sage briefly.
I am on Ampere so I get the boring version. It gives around 10-15% speed up to me.
It changes the output noticeably, but whether that is degradation, or seed variation, I haven't tested it enough to say confidently.
Kijai also made a specific sage patch for H3, I haven't compared it directly with standard sage but it's worth checking out.
>>
https://xcancel.com/bdsqlsz/status/2084627424311116006?sort=Likes#r
>Wan 3.0
this shit genuinely looks worse than Minimax, feelsgoodman
>>
>>109466030
what is this? Am I the only one who just cloned the repo and set up a venv?
>>
>>109466207
based me too
>>
>>109466200
*10-20% or 15-20% should be more accurate
Nevertheless
>>
>>109466144
upgraded to latest pytorch, in case you guys have problems with torchaudio, use the one from this repo:
https://huggingface.co/ussoewwin/torchaudio-built-on-cu132-for-windows/tree/main
>>
Insane. H3 knows the difference between left and right when a person is facing the camera. So the right hand is actually on the left side of the image, and it knows it.
>>
>working this hard to hit bump
lole
>>
>>109466207
yes you are the first
>>
>>109466160
gm
mine was a friendly insult while giving advice
be of thankgins
>>
>>109466204
She's walking in the air.
>>
holy shit, anime works with real life people too. and well, even.

<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "I can't breathe!". <Subject 1> says "i'm 2B and you're just 1 bitch.".

https://files.catbox.moe/9ubxgp.mp4
>>
>>109466215
turns out that going for a giant text encoder has giant consequences
>>
>>109466228
the reference model has way more fuzzy artifacts that the normal model, that's a shame
>>
Best tool that I can use to remove speech bubbles? Krea2 edit was a let down.
I'm open to not use Comfy if something else has better in paint
>>
>>109466238
im genning at 0.3 just for testing, at higher res it's pretty crisp in general
>>
>>109466239
klein edit 9b distilled int8 convrot (as good as q8 but faster)

either is fine, 4 steps
>>
File: 1770928337876409.png (137 KB, 1209x353)
137 KB PNG
so this is the peak setup right?
>>
>>109466228
asuka version:

https://files.catbox.moe/3gmizm.mp4
>>
>>109466251
the "kv" or non kv version? took first page on google.
>>
>>109466259
I skipped a day. What is this new tech?
>>
>>109466259
Sol attention doesn't do anything for me. Sage + Spectrum did decrease gen time by 1 minute.
>>
>>109466251
>>109466262
Adding tot that can I use the built in comfy work flow or should I get a different one?
>>
>>109466262
non kv is fine in general

im using flux-2-klein-9b-int8-ConvRot-comfyui.safetensors in the default 2 image workflow (bypass 1 if doing 1 image edits)
>>
>>109466267
sol attn should only work if you have triton and recent kernels... and only from the 2nd gen onwards (triton needs to compile shit).
Im testing timing with sol attn now (on 16gb vram), ill do sageattn + patch after to see if I get better timings and report back
>>
hahaha

<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.

the setting is New York City during the day. <Subject 1> plays a rock song on her guitar, then hits <Subject 2> in the head with the guitar, breaking it. <Subject 1> grabs a large speaker and throws it on top of <Subject 2>

https://files.catbox.moe/zmvljb.mp4
>>
>>109466275
I do have one colored image of the manga, I would really like to color it all while removing the speech bubbles, does the order of the images matter? Or is this just too advanced?
>>
>>109466283
keek
>>
File: MiniMax_H3_00161.mp4 (1.26 MB, 1152x864)
1.26 MB
1.26 MB MP4
>>
>>109466284
just prompt "remove the speech bubbles in the image" and it should work, also you can do "make it monochrome" or black and white, should work, it can do lots of edit stuff.
>>
>>109466291
I meant that I use the cover as a color reference and it colors the images as it removes the bubbles
>>
>>109466284
klein 9b is not going to consistently color across all pages but each page on its own might more or less work
>>
File: ws.jpg (161 KB, 1796x1026)
161 KB JPG
a dark force looms over the land
>>
>>109466297
>>109466297
>>
>>109466207
Yeah i had it on my Iphone 17 pro max before, i didn't know you could put it on one of those old fashioned "PC's"
>>
>>109466294
can remove the bubbles in one pass then do whatever you like with the new image on a new gen?
>>
File: 1761464529526887.png (22 KB, 228x236)
22 KB PNG
>>109466279
aight results are in
this is on a 4080 super
1st gen solattn triton: slowest
2nd gen solattn triton: fast
3rd gen sage+eff patch: fast
4th gen sage+eff patch: fast
I'll just use sage attn I guess
>>
>>109466283
good lord :DD
>>
>>109466200 >>109466213
sage seemed like the biggest speedup for the least noticeable downsides so far, that's why i asked
>>
>>109465348
use your brains guys keep the nsfw shit off the internet, why the fuck do people even need to share their degenerate shit anyway? All its doing is drawing bad attention from media cucks who will in days complain about muh safety.
>>
File: mini.png (66 KB, 638x840)
66 KB PNG
is it necessary to download all of these files or are those just multiple quants of the same thing? I have 12GB VRAM, and I'd rather not download anything I can't run anyway.
>>
>>109466446
diffusion models
One fl2va .safetensor for t2v/i2v
One ref2va .safetensor for reference2v
Pruned int8 convrot is the go to

text encoder
nvfp4 is fine

and both vaes
>>
>>109466446
You will be able to run the int8 or fp8 pruned versions just don't expect to be able to do a full 15 seconds at higher resolutions. lower res like 0.5 megapixel you could do full 15 seconds. The model is flexible thought from what I've heard so you could do smaller length and then extend the videos and merge them at higher res
>>
>>109466469
thanks bro
>>
>>109466005
>what is required to preserve distant faces
vramlet holocaust
>>
>>109466030
>kekstop ver
>>
>>109466484
thanks a lot. I am coming back to playing with local models after a long break, so I appreciate all the intel. it's first time I am touching any video models.
>>
>>109466536
> it's first time I am touching any video models.
you came back at the right time, we didn't have this until just now

simple prompts are better, complex prompts are better, couldn't use that many references but now we can and so on
>>
What do you think about integrating RTX Super Resolution in your minimax workflows? The video looks good to me, but I don't know why it failed to reintegrate the audio.
>>
>>109466136
because plump does not know that every open sauce, when you ask it about plump, responds with 'he is turbo gay'
>>109466506
default workflow i did not even see that 0.3 megapixels thing was active

if everyone does like few minutes of gen for some script we can make blockbuster movies now.
and since my vramlet situation is indeed somewhat there, i can make a logo for the intro
>>
File: ai.png (171 KB, 1406x600)
171 KB PNG
>>109466484
>>109466563
I was able to run it with the default i2v PC mouse job, and it worked fine. 78 minutes though... kinda brutal. I am on 4070 (12 GB) and 128 GB RAM. GPU and VRAM use was at 100% for the whole time.
>>
>>109466313
follow kijai's setup. Also are you on blackwell? Also also, sol gets better at higher res / frame count
>>
Having success but also trouble with audio references. Adding a reference for voice timbre works, but then the output video doesn't create any other audio except the background noise from the reference, despite explicit prompting. Anyone had more luck with using audio reference to create only one element in a video?

Actually, for that matter, anyone having good luck with audio prompting in general?
>>
wish I could test this new video model, im AI burnt out



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.