[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Lurk More Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109454191

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
gm saars
>>
blesd agin
>>
>>109455085
Mikudayooo
>>
gn oomfs
breast thred of frenshit
>>
File: MiniMax_H3_00038_.mp4 (308 KB, 384x544)
308 KB
308 KB MP4
>>
is it finally okay to admit that ltx was always hot garbage that we tolerated because it was all we had?
>>
Why is it so good at Seinfeld clips?
>>
>>109455113
Yes
In fact its been that time for a while
>>
>new model drops
>new fren posting increases 10 fold
uhg
>>
>>109455113
My LTX gens didn't have this ugly noise tho.
>>
low res test but it does work, reference model again

https://files.catbox.moe/ms868n.mp4
>>
>does porn out of the box
why is seeddance so censored?
>>
>>109455113
I think everyone admits that. There was like 1 ltx poster here posting micro gens
>>
>>109455121
>forsen says "hey forsen"
kek
>>
>>109455120
yeah i noticed the dots are very prevalent with h3. ltx was better at smoothing it out but at the expense of small details being smudged
>>
>mfw Resource news

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md

>Raylight 1.7.2 Adds MiniMax Support, 2x Speedup
https://github.com/komikndr/raylight/releases/tag/1.7.2

>Scaling Properties of Text Conditioning in Visual Generation
https://heheyas.github.io/context-scaling

>Retrieval-Driven Training-Free AI-Generated Video Attribution
https://github.com/renxi-seu/Video_Attribution

>A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
https://github.com/zfu006/SSG

>ComfyUI MiniMax H3 Image Studio (Experimental)
https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio

>MiniMax H3 — NVFP4 (Blackwell)
https://huggingface.co/lilcheaty/MiniMax-H3-NVFP4

08/02/2026

>MiniMax H3
https://huggingface.co/MiniMaxAI/MiniMax-H3

>MiniMax H3: Repackaged model files for ComfyUI
https://huggingface.co/Comfy-Org/MiniMax-H3

>MiniMax-H3-INT8-CONVROT
https://huggingface.co/Gluttony10/MiniMax-H3-INT8-CONVROT

>MiniMax H3 COMMUNITY LICENSE AGREEMENT
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE

>LoRA Dataset Studio: LoRA workflow in one tab
https://github.com/perfectgf/lora-dataset-studio

>comfyui-vram-tracker
https://github.com/PuppetMasterAI/comfyui-vram-tracker

>Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
https://nvlabs.github.io/Sana/Sol-Attn

08/01/2026

>EU to get power to enforce rules on AI starting today
https://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786

>FameGrid Auto Color for ComfyUI
https://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui

07/31/2026

>Introducing Seedance 2.5
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5

>ReToken: One Token to Improve VLMs for Visual Retrieval
https://github.com/avaxiao/ReToken
>>
>>109455113
it always was, i never said it wasn't
>>
>mfw Research news

08/03/2026

>MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
https://orange-3dv-team.github.io/MoRoute

>WaiT for the Signal: Simple Frequency-Aware Flow-Matching
https://arxiv.org/abs/2607.28760

>Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
https://arxiv.org/abs/2607.29025

>Visual Distribution Anchoring for Efficient Prompt Tuning
https://arxiv.org/abs/2607.28967

>MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
https://arxiv.org/abs/2607.29180

>RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
https://arxiv.org/abs/2607.28974

>SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation
https://arxiv.org/abs/2607.29367

>When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
https://arxiv.org/abs/2607.29240

>Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs
https://arxiv.org/abs/2607.29412

>Domain-Adaptive Deep Joint Source-Channel Coding for Image Classification
https://arxiv.org/abs/2607.28907

>Explaining AI-Image Detection: What the Heatmap Actually Shows
https://arxiv.org/abs/2607.29581
>>
>>109455113
>is it finally okay to admit that ltx was always hot garbage that we tolerated because it was all we had?
of course, it was fun to have sound for the first time but I always got horror shit
>>
File: 1772169418678905.png (190 KB, 588x1297)
190 KB PNG
Ostris H3 training support.
>>
btw anyone getting a speed up from that int8 custom node instead of the regular loader is doing so cause your not using CU130+ with comfykitchen. Update and use regular cause it uses less memory
>>
dont think ive been this excited since the original NAI leak
>>
File: 1232131232.png (77 KB, 521x713)
77 KB PNG
Help /g/. Am I doing something wrong? I am getting horrible artifacts using the default workflows with the mixed int version
>>
>>109455123
it's for normies, retards and indians, if you give them access to an nsfw model they will flood the internet with csam. just look at grok lol.
>>
no german accent but the reference picture worked:

https://files.catbox.moe/44apvk.mp4
>>
>>>/wsg/6207687
>>
>>109455173
when was the last time you defragmented your tensor cores
>>
>>109455167
my pytorch version: 2.9.1+cu130
do i need to update?
>>
>>109455179
and catbox is dying cause the link didnt work.

https://litter.catbox.moe/66im8dl12tje8rzl.mp4
>>
>>109455176
>if you give them access to an nsfw model they will flood the internet with csam
so this is the natural human desire
>>
>>109455181
the gabagool is gonna hunt down tony like the fat slob he is
>>
>>109455173
Classic tensorcore fragmentation issue
>>
File: MiniMax_H3_00005_na.mp4 (1.99 MB, 832x640)
1.99 MB
1.99 MB MP4
4:03 for 10 seconds
12:19 for 20 seconds
25:44 for 30 seconds
@ BF16, RTX 6000
>>
>>109455204
if you consider indians human, then yes, i guess it is.
>>
>>109455212
I dont but you mentioned normies
>>
I consider them bio weapons. They should not have access to any form of AI.
>>
File: 1774453609983844.png (345 KB, 640x465)
345 KB PNG
hollywood jews are going to seethe SO fucking hard about this kek
>>
>>109455222
normies and twitter users are not human either in case that wasn't obvious
>>
>>109455201
kek
>>
>>109455173
Increase mega pixels.
>>
>>109455210
>4:03 for 10 seconds
>12:19 for 20 seconds
so it's 3x slower if you go for a video 2x longer? damn
>>
>>109455226
Why? They'll just use AI cheapen production and maximize profits. You need to think like a Jew.
>>
File: 27456124.mp4 (3.62 MB, 448x576)
3.62 MB
3.62 MB MP4
>>109455210
Have you tested how long it takes on INT8 convrot?
>>
>>109455226
I doubt hollywood is going to be toppled by 15 second slop gens
>>
Fuck I'm up like 2 hours later than I should be. Fucking AI addiction. I hate myself.

https://files.catbox.moe/64rcus.mp4
>>
>>109455226
>Disney sues Minimax
>Minimax's answer is to make a model full of IP and releasing it locally
BASED??
>>109455239
but they wanted the cheap AI production only for themselves, if everyone has the power Disney is in a big trouble, they can't gatekeep movies and series anymore
>>
>>109455239
>>109455242
they already had this, now WE have it and it doesn't give a FUCK about jewish copyright or muh protected politicians or whatever

if you think the jews aren't seething that any chud can make
>>109455201
you don't know much about them
>>
>>109455242
disney must be suing minimax for kicks then
>>
Local Diffusion?
>>
>>109455245
>>109455246
You have no idea how Jews think. They only seethe when they have to pay taxes. AI is not a threat on their radar.
>>
File: 2538070835286395853.jpg (18 KB, 335x597)
18 KB JPG
>>109455222
>it's for normies, retards and indians
you are on 4chan, so probably scratch normie off the list. i guess that narrows it down to retard or indian.
>>
>>109455246
>any chud can make
make what? make 15 second slop vids? maybe when they add a new "best 15 second movie" category to the oscars, you can compete with hollywood
>>
he CANT keep getting away with it!

https://files.catbox.moe/njzbe0.mp4
>>
>>109455201
amazing
do you have to give it reference material to generate this? or just the concept and the dialogue? (or even the dialog)?
>>
File: 1784977980946392.png (1.76 MB, 1547x1696)
1.76 MB PNG
>>109455254
>AI is not a threat on their radar.

>Major studios like Disney and Paramount quickly accused ByteDance of copyright infringement but concerns about the technology run deeper than legal issues.
nah, they're terrified of that, the goyim shouldn't have the power to make Holywood tier level of videos, that would ruin their business
>>
>>109455267
They sue for sport. This is not them seething; they're having fun.
>>
The EU bans random argie soccer memes from twitter, they're seething about this
think of all the chud propaganda people wil make
>>
So what sage attention version do I need to use dat KJ EfficientSageAttentionPatch node
>>
>>109455259
this joke is so Seinfeld, vague but understable, you nailed that shit anon
>>
>>109455269
if you're gonna pull shit out your ass then have fun with it
suing is healthy, healing, they sue because they love
>>
>>109455269
>>109455277
if they don't sue, it means your model is shit, Disney suing you means that you made a good model, it's like a seal of quality
>>
Is 20 steps enough? Would 40 improve the result?
>>
>>109455277
Why would I pull shit out of my ass? My bowels work perfectly fine.
>>
File: ComfyUI_00105_.jpg (87 KB, 800x800)
87 KB JPG
>>109455254
>You have no idea how Jews think.
>>
>>109455271
2.2.0+
https://github.com/woct0rdho/SageAttention/releases
>>
>anon has never experienced the joy of digging in his ass for treasure
>>
>>109455285
people get decent results with 10-15
>>
ryan gosching?

https://files.catbox.moe/al9fa1.mp4
>>
File: 36847693132.mp4 (3.51 MB, 672x480)
3.51 MB
3.51 MB MP4
>>
>>109455162
Holy shit. If it's actually going to work for images, we'll be able to pump all the booru tags into it, along with all the artist styles.
>>
>>109455315
I don't think it will, and I don't think the video loras will train how people expect them to
>>
the volk be mp4-ing, but I'm not.

I'm an image guy.
>>
File: MiniMax_H3_00039_.mp4 (336 KB, 384x544)
336 KB
336 KB MP4
>>
>ImportError: cannot import name 'guidedFilter' from 'cv2.ximgproc'
Anyone run into this error?

I went to the the thread the console gave me https://github.com/chflame163/ComfyUI_LayerStyle/issues/5 says to run some commands in "the plugin directory" What is this vague plugin directory? I have like 3 plugin directories inside comfy.
>>
>>109455244
I can't gen on weekdays for that reason. It's too dangerous.
>>
File: HunyuanVideo_00043.mp4 (497 KB, 640x480)
497 KB
497 KB MP4
>prompt for vagina
>get front hole instead
>>
So I've gotten used to most of the H3 reference model syntax, but still haven't really figured out how to use image/video input as a style reference, like art style. attribute_transfer in subject rentention_analysis doesn't seem to influence much. I've managed motion/theme transfer and subject definition stuff, but I'm stumped on actual style reference.
>>
>>109455323
the ubermensch webm
>>
I'm having hard time understanding what models my 5070 can handle, is it just models that are <12gb in size?
>>
File: 198122579543.mp4 (3.57 MB, 672x480)
3.57 MB
3.57 MB MP4
>>
the reference model/workflow are funny cause you can plug in ANYTHING as a reference and you use reference tags to have the prompt interact with it. <Picture 1> <Video 1> etc.
>>
>>109455352
nice
>>
>>109455352
based
>>
>>109455333
can it do this where she turns around and is doing a photo shoot (glamour)

prompt:
>A beautiful Spanish woman with an unusual nose is posing for photographs with a professional photographer. There are assistants tossing colored liquid, power strobes in flash umbrellas on flash stands, and a laptop recording the photos
>>
>>109455351
No, it will offload to RAM and storage.
>>
>>109455113
I tried all the old shit and it was trash, we have something thats actually usable now
Took awhile but we're finally here
>>
>>109455361
>unusual
that is a non descriptive word, it describes nothing but your own preconceptions
>>
>>109455361
yeah
>>
File: 1780520963860672.png (65 KB, 1244x496)
65 KB PNG
don't use the INT8 Fast node anymore it's deprecated
>>
>>109455325
nvm just figured it out
>>
>>109455378
comfyanon propaganda.
>>
>>109455361
Don't you have a bullet to eat?
>>
>>109455378
it's literally faster for me than load model.
>>
>>109455382
no, the opposite is retard propaganda that never update their shit and wonder why stuff sucks
>>
>>109455390
>t. gguf chud
>>
>>109455386
you dont have up to date torch/cu130+/comfykitchen
>>
https://xcancel.com/Hailuo_AI/status/2084504994636783979#m
>the official minimax account posts an AI video based on the office
they're scared of no one... I fucking kneel
>>
>>109455378
yep. he hasnt updated repo in 2 months. for me native load diffusion is way faster
>>
>>109455394
I updated cu130 for convrot before and for whatever reason this new node is faster for me
>>
>>109455396
where do i submit my application to work for the ccp in exchange for access to latest model weights and state mandated bugwaifu?
>>
>>109455401
and torch and comfykitchen?
>>
BUY TODDS GAME!

https://files.catbox.moe/9n256z.mp4
>>
>>109455167
Can confirm. I'm retarded. Updating all the cuda shit made it even faster than the int8 custom node
>>
>>109455414
too bad seinfeld stole the lines of todd the idea was great lol
>>
>>109455378
There should be an updated fork around, can't remember where though but had to pull it up for something
>>
>>109455396
HOLY FUCKING BASED
>>
i'm looking at the prompt guide and i see "overall_soundscape". does it properly divide the sound actions across the whole video if i put them into a single line for this parameter?
>>
WE'RE GONNA NEED
FIVE TIMES THE VRAM
>>
>>109455374
cool
>>
>>109455418
how do update sageattention shit? all the guides are "how to install", but they don't said how to update it.
>>
>>109455396
>>
File: MiniMax_H3_00006_na.mp4 (2.37 MB, 832x640)
2.37 MB
2.37 MB MP4
>>109455241
This took 11:38 on the pruned INT8 convrot, 44 seconds less.
>>
>>109455454
ask chatgpt?
>>
>>109455396
it's like trying to load a jpeg in the fucking 90s.. wtf is that trash ass site
>>
>>109455462
sorry Elon, but I'm not gonna make an account just to read comments, if you want to go for x.com then replace xcancel by x
>>
>>109455396
no no no NO THAT'S COPYRIGHTED YOU CAN'T DO THAT
>>
>comfyui still not having a simple vibecoded dialog to autoupdate/change torch/cuda/sage
do they not care about normgroids using their ui?
>>
File: file.png (4 KB, 469x27)
4 KB PNG
what am I in for?
>>
I tried, kind of, to do the 80's metal.

this is the result.

https://vocaroo.com/1031490awBh3

MASTERPIECE
A
S
T
E
R
P
I
E
C
E
>>
>>109455454
for me it was
python_embeded\python.exe -m pip uninstall sageattention -y
python_embeded\python.exe -m pip install https://github.com/woct0rdho/SageAttention/releases/download/v2.2.0-windows.post6/sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl
>>
>>109455470
>>109455361
My /r/
>>
>>109455466
>autoupdating a system-wide cuda and torch
I can think of few things more terrifying
>>
File: H3_mute__00004.webm (3.71 MB, 944x736)
3.71 MB
3.71 MB WEBM
It's now much easier to stitch together multiple cuts for a music video locally. I am going to attempt a full song.

>>>/wsg/6199339
>>
>>109455457
Huh, I expected a more significant speedup.
>>
I'm the llm. you're the user.
>>
>>109455483
the speedup is for the 3090 and lower cards, big cards like the rtx 6000 already has a lot of optimizations so int8 won't move the needle that much
>>
File: MiniMax_H3_00043_.mp4 (199 KB, 576x384)
199 KB
199 KB MP4
>>
>>109455494
thanks xi
>>
>>109455493
It's blackwell right? Consumer blackwell gpus get a much larger speedup.
>>
File: lul.png (986 KB, 3840x2162)
986 KB PNG
>>109455494
after playing around with minimax, my conclusion is that maybe communism isn't bad after all...
>>
Okay so what is the optimization meta for a 3090? SageAttention, EasyCache. What other nodes and tricks can I use for quick efficiency gains?
>>
File: MiniMax_H3_00044_.mp4 (383 KB, 448x448)
383 KB
383 KB MP4
>>
Just tried updating comfyui and getting "Cannot connect to comfyregistry." in stabiltiymatrix console at it won't launch. What gives?
>>
File: 1766744375776218.png (3.12 MB, 1759x1539)
3.12 MB PNG
Here's a node to transform H3 into an image/edit model
https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio
>>
File: MiniMax_H3_00068.mp4 (2.45 MB, 960x1088)
2.45 MB
2.45 MB MP4
body inspection
>>
>>109455514
lmao
>>
>>109455524
This is just official FF7 remake footage.
>>
>>109455517
try "remove the whore and replace her with a cool computer"
>>
>>109455310
extremely hot if that's just I2V
>>
>>109455532
>Woke Enix catering to the coomers
hah, I wish...
>>
no other model can do this
https://litter.catbox.moe/32bgtqka5nccg322.mp4
>>
this came out more nsfw than I prompted for
https://files.catbox.moe/lbczff.mp4
>>
>>109455549
a good dataset is all you need, it's that simple, and Minimax applied that to the T
>>
>>109455555
checked.

A guy who liked to say gigo that I kind of knew is now in prison for rape.

I think about that every time the topic comes up, and it's pretty depressing.
>>
>>109455509
I think INT8 convrot like other quantization methods is mainly to save VRAM, so if you're not running out of it, you are compute limited
>>
>>109455568
no, its 2x faster than FP16 with almost no visible quality loss
>>
The setting is a diner from the show Seinfeld. Jerry Seinfeld from the show Seinfeld is sitting at a booth with with George Costanza and Jerry says "George, you know I want a new PC game to play. I'm so bored!". The man from <Picture 1> walks in the diner and says "Have you considered the great award winning game Skyrim?", and the audience laughs.

https://files.catbox.moe/muzf1l.mp4
>>
File: 00022-901612745.jpg (573 KB, 2880x1920)
573 KB JPG
>>
>>109455572
I don't know why but I would love Emplemon to make a video explaining why people love to make memes with Seinfeld so much, there's something special about that series and I love it but I can't put my finger on why it's the perfect series for modern memes
>>
File: 1760482682327618.png (62 KB, 202x268)
62 KB PNG
>me realizing the jews can never put the genie back in the bottle
all we need is voice cloning from audio samples
>>
>>109455555
epic get for epic truths.
>>
>>109455544
It's just the state mandated mandatory NTR shit pushed in every jap media. You're allowed to pander to cookers, but only in that specific way.
>>
>Costanza is a localkek and Seinfeld is an APICuck
sounds about right
https://files.catbox.moe/d0ftu7.mp4
>>
>>109455585
you can use audio as input in the reference model, it can clone too I think.
>>
>>109455585
They can always ban huggingface and/or delete its domain name
>>
File: spongebob laughing.jpg (109 KB, 938x528)
109 KB JPG
I never bothered with LTX turd and I am proud of it.
LMFAO at all those coping retards here shitting themselves and having tendie meltdowns at me whenever I mentioned their gens looked like shit and worse than even what Wan 2.1 could output.
>B-b-but it had hecking audio!
L bozos, don't throw yourselves at any random model you see. Have some self-respect. See? Patience pays for itself in the end.
>>
>>109455592
At least animate it...
>>
>>109455591
they know it's useless, people will go for modelscope or torrent, they won't move a dent and they will pass as totalitarian, nothing worse than getting bad PR + no results
>>
>>109455598
Bad PR hasn't stopped (((them)))
Iran war is bad PR and look at where we are now
>>
>>109455602
I'm sure Trump thought he was gonna get some results though, yeah he's retarded what can I say
>>
File: 1769867051187676.png (256 KB, 1594x792)
256 KB PNG
neat, google ai mode can help with the new model/prompting stuff:
>>
>>109455589
we can literally generate full seinfeld episodes with modern content and jokes they would never allow on TV now.
>>
>>109455605
He did though. Russia lost its main source of cheap high quality oil and both it and and china lost their source of cheap drones. Also Iran's sponsorship of terror has cut down by like 99% as most groups that used to habor them turned on them now. Also it cuts down iran's weapons they had been building for decades to eventually try and make a move in the region.
>>
File: 1774691527108247.png (114 KB, 640x640)
114 KB PNG
>>109455592
I had my fun with ltx, yeah it looked like shit, like the ps1 looked like shit but I was witnessing history and I participated in it
>>
>>109455378
It still allows on the fly quantization and baking lora weights into the model before quantization for maximum quality (pre-load), but besides these yeah not too many reasons to use it for general use.
>>
>>109455592
why do they live rent free in your mind?
>>
>>109455597
>At least animate it...
Neither him nor >>109455619 are able to generate locally.
>>
>>109455572
>hold on buddy, you can't take 87th! that's THE street.
>kramer, they don't have a street.
>i'm tell you jerry, bob sac-
>yeah yeah, "bob sacamano, said they have a designated street." that's nonsense kramer, you can't designate a whole street to that.
george enters the apartment
>george what happened to your shoes?
>87th street happened jerry.... THE WHOLE STREET! it was like a minefield.
kramer clicks his teeth and points his finger at jerry
>i told ya!
>>
>>109455615
Ikr, it comes full circle, the chinks aren't just giving us a good model, they are giving us the power to express ourselves like the good old times
>>
>>109455584
People wish they hadn't married.
>>
File: rent free cope.jpg (60 KB, 900x425)
60 KB JPG
>>109455623
>rent free cope
Nonnie it's time to admit you just fucked up in the past and move on.
>>
>>109455591
>They can always ban huggingface and/or delete its domain name
they can't delete modelscope.cn
>>
>>109455641
>you just fucked up
fucked up? I just tested a model and found out it wasn't that good, big deal, fucking something up is like getting into a car accident and breaking your legs, if you're so mentally fragile you think testing a bad model takes a tool on you there's not much I can say anon, it's a bit sad really
>>
File: 16746.png (162 KB, 502x477)
162 KB PNG
>>109455641
>Nonnie it's time to admit you just fucked up in the past and move on.
so i made a mistake by running sd 1.4 because krea 2 exists now?
>>
>>109455592
based as fuck. i generated ONE good LTX gen ever out of 30 tries and said fuck no. my only regret is dooming so hard because i predicted the sora 2 killer by july instead of by august
>>
>using Krea 2 when Z Image Turbo exists
>>
>>109455653
ltx was never SOTA in any field, SD was
>>
>>109455660
no, dalle was better than SD at release
>>
>>109455583
looks like melztube
>>
can it do breast expansion
>>
File: 1758570595326679.jpg (16 KB, 462x169)
16 KB JPG
are there any seed nodes like this that can go back to the previous seed when necessary so it doesnt trigger new actions? they currently dont work with subgraphs
>>
>license forbids americans, euros and koreans from using it
kek
>>
>>109455589
top kek
reminder those jews seethed about a twitter account making up fake seinfeld episodes
this is gold
>>
>>109455665
You couldn't run Dall-e on your computer so as far as LOCAL diffusion general (/sdg/ back then) was concerned it was SOTA.
>>
>>109455592
i clearly remember posting several ltx 2 gens back in November of last year and several anons liking my gens. No need to larp as a old fag. The model was heavily flawed but the audio and talking heads was it's selling point beause we weren't getting wan 2.5. Local had lack shit at the time besides endless coping with high noise and low noise tweaking bullshit that was wan2.2.
https://litter.catbox.moe/qhktbi.mp4
https://litter.catbox.moe/dli326.mp4
https://litter.catbox.moe/hmc175.mp4
https://litter.catbox.moe/9540ml.mp4
https://litter.catbox.moe/60wkw0.mp4
https://litter.catbox.moe/5bxjps.mp4
>>
>>109455687
>this is gold
https://www.youtube.com/watch?v=CF7OnW4XDck
>>
tips for time based events:

Image-to-Video (I2V) Timing Example

For I2V, Image 1 is your starting frame. The prompt's job is to tell H3 exactly when and how to break that static image into motion.Prompt:"Using Image 1 as the exact starting frame.0 to 3s: The character in Image 1 smiles gently and blinks naturally while looking directly at the camera.3 to 7s: The character suddenly turns their head to look sharply toward the right side of the frame.7 to 10s: They lift their right hand into the frame and wave goodbye.Camera: Static shot for the first 3 seconds, then a subtle camera pan right to follow their gaze.Sound: Ambient wind rustling throughout the entire clip."

so just do "0 to 5s: Miku runs to the store. 5s to 10s: Miku grabs a juice" for example.
>>
>>109455689
ltx was the best local model that we had, so what's the problem here?
>>
training a test lora but man, its about 4x slower than ltx to train per step
>>
>>109455692
>close up, slow zooms, just talking
obviously ltx can do that, would be a disaster if it can't
>>
>>109455697
>so what's the problem here?
The problem is the preceding statement being objectively incorrect.
>>
>>109455706
ok, if you say so
>>
File: MiniMax_H3_00069.mp4 (911 KB, 960x1088)
911 KB
911 KB MP4
action just get fast forward if you set shorter duration
>>
The nightmare of updating pytorch and reinstalling triton/sage done. Back to kinos
>>
File: 74453455.webm (1.16 MB, 448x256)
1.16 MB
1.16 MB WEBM
yooooo this is kino fr
>>
what are some good hello world gens? i'm kind of blanking. i need something i can eval quality vs performance
>>
>>109455745
1girl, huge breasts
>>
>>109455745
Astronaut riding a horse on the moon
>>
>>109455745
1girl, gigantic_breasts, breasts_bigger_than_head
>>
I have 2015 pc specs I definitely won't be able to run a local model. What's the best option on the market rn?
>>
>>109455755
Impossible to gen because there aren't any boob bigger than head pics in the training set
>>
>update_comfyui_and_python_dependencies.bat
this actually updates it for me. be sure to backup before trying.
>>
what does the shift scale parameter do?
>>
why the fuck does comfyui not support pooling vram
>>
>>109455768
A single model spread across multiple cards? Does any non llm engine do that?
>>
>>109455768
you need nvlink for it to be worth doing otherwise it will be too slow to be worth any possible speed gains. Otherwise parallel inference / higher batches is the way to go
>>
timestamp test:

the anime girl is standing wearing a japanese highschool uniform. 0 to 3s: The character in puts a slice of toast in their mouth to hold it. 3 to 7s: The character suddenly runs very fast with the toast in their mouth down the street. 7 to 10s: They jump high into the air and throw the toast at the camera.

https://files.catbox.moe/1uov1x.mp4
>>
>>109455758
literally spend money on a prebuilt gaming pc with 64gb of ram. microcenter has 5090 prebuilts for $5,200.
https://www.microcenter.com/search/search_results.aspx?fq=Subcategory:Gaming+PCs,Total+Memory:32GB+and+greater&Ntt=prebuilt+pc+gaming&myStore=true
>>
Does python 3.14 really have issues with comfyui on linux?
How do I install an older version without messing with the rest of my system?
>>
>>109455813
docker.
>>
>>109455793
pretty much spot on.
>>
holy shit, kino

the anime girl is standing wearing a japanese highschool uniform. 0 to 3s: The character in puts a slice of toast in their mouth to hold it. 3 to 7s: The character is having a fist fight with Hatsune Miku in the style of Dragonball Z. 7 to 10s: the anime girl and Miku leap away from each other.

https://files.catbox.moe/hvk1uv.mp4
>>
>>109455817
I use docker for open web-ui, how do I use it to set up comfy properly?
Which version of python should I use?
>>
>>>/g/sqt
>>
>>109455817
is uv not enough?
>>
>>109455482
cute
>>
>>109455834
uv will do just fine.
>>
>>109455827
the anime girl is standing wearing a japanese highschool uniform. 0 to 3s: The character in puts a slice of toast in their mouth to hold it. 3 to 6s: The character is having a fist fight with Hatsune Miku in the style of Dragonball Z. 6 to 10s: the anime girl with blue hair powers up and transforms into a female Super Saiyan.

https://files.catbox.moe/42nwxc.mp4
>>
>>109455715
it's like he knows he needs to hurry
>>
>>109455847
i am retarded
what do i do exactly
>>
>>109455482
reference model? how do you swap characters? or did you i2v the dance
>>
>>109455853
Reference model with 1st frame.

summary:
[reference generation + audio reuse] The target video depicts a music video with the girl as the main subject.
>>
File: 325344.webm (1.29 MB, 448x256)
1.29 MB
1.29 MB WEBM
okay i converted my prompt to the official format. i think that marking the 00:00:00 timestamp after i describe the scene composition was the fix. it's kino time
>>
File: 1xip7i.jpg (15 KB, 400x300)
15 KB JPG
Sooo how long until we get a porn finetune?
>>
>>109455861
Yeah that looks way better
>>
>>109455861
integrated_multimodal_description: [Shot 1] dusk time. thin clouds. snowy mountains in the distance. flat horizon. live-action, raw found-footage style. bad quality combat footage. blurry video. motion blur. smudged lense. chromatic aberration. cockpit of a very large military jet. we are wearing a green flight suit with black fabric gloves. reflective glass canopy. our left hand is on the throttle lever, and our right hand is on the flight joystick. the flight instruments are in front of our legs. the screens on the instruments have a green tint. green holographic sight over the dashboard. rows of buttons and switches on the sides of the cockpit. the nose of the plane is visible over the cockpit. our wings are not in frame. tilted and off center framing. at 00:00.000 we are flying very fast and close to the ground over a mountainous forest. highly dynamic and natural shaky unstabilized camera motion. very suddenly at 00:03.000, the dashboard becomes full of red warning lights that flash on and off constantly. the perspective is startled and frantically looks out the side of the canopy. modern swept fuselage comes into view. at 00:06.000 in the distance, a surface to air missile is launched from a launch site in the forest which flies up into the sky while emitting a trail of smoke. very fast and frantic camera movements as look back and forth between our flight instruments and the missile. the plane dives down to the ground while doing a fast evasive maneuver. at 00:09.000 we look to the side one final time at the incoming missile. the missile directly hits our cockpit and fills the frame with an explosion of fire and debris.

overall_soundscape: low quality go-pro audio. muffled jet engine noise with rumbling airframe. pilot breathing muffled by oxygen mask. alarms start beeping from the dashboard. hyperventilating. distant misisle shrieking noise approaches and creates a loud boom when it hits the cockpit.

non_diegetic_music: N/A
>>
>>109455864
that's just a formality
>>
>>109455864
I finetuned your mom last night, I need a bit of rest.
>>
>MiniMax H3 is a multimodal "omni" model, meaning it generates video frames and audio tracks at the exact same time. By default, if a prompt is about a cool action scene, the AI will automatically try to be helpful and generate an epic background music track to match it. Writing non_diegetic_music: N/A completely disables that specific musical background-score layer.

useful if you dont want random music
>>
Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8- is out

anyone tried it?

https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-MiniMax-H3-ComfyUI-INT8-ConvRot/tree/main

Snek oil?
>>
>>109455888
definitely a snake oil, the model is already fully uncensored it doesn't have any gay filters or refusals
>>
Heretic removes alignment from LLMs, what does it do on a model that makes porn, copyrighted characters, and porn of copyrighted characters out of the box?
>>
>>109455885
yep very nice feature. ltx had the problem of random music
>>
>>109455893
>>109455888
never use a different TE than the one the model was trained with. Divergence = worse results. And the TE being censored does not effect its use as a TE at all, its basically just a translator, it has no say
>>
https://files.catbox.moe/1w9lz6.mp4
It knows actors voices somewhat. This is insane kek
>>
>>109455852
# Install uv, then restart your shell
curl -LsSf https://astral.sh/uv/install.sh | sh

git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI

# Downloads Python 3.13 and creates an isolated environment
uv venv --python 3.13

# NVIDIA
uv pip install torch torchvision torchaudio \
--extra-index-url https://download.pytorch.org/whl/cu130

# Remaining dependencies
uv pip install -r requirements.txt

# Start
uv run python main.py
>>
>>109455885
Hatsune Miku is walking down the street. 0 to 2s: Miku is walking to the right saying "what a nice day!". 2 to 8s: The character is having a fist fight with Vegeta from Dragonball Z in the style of Dragonball Z. 8 to 10s: Miku punches Vegeta far away sending him flying into a wall that makes a large explosion. non_diegetic_music: N/A

https://files.catbox.moe/m883bu.mp4
>>
>>109455900
>somewhat
the model has been trained on those actors and with the right captions, those chinks don't give a single fuck, absolute basado
>>
File: MiniMax_H3_00007_.mp4 (1.02 MB, 864x480)
1.02 MB
1.02 MB MP4
>>
>>109455900
404
>>
>>109455911
nothing to see here, citizen
>>
>>109455897
I will test your claim using both TEs.
>>
>sees MMH3 using qwen3 model
>krea2 uses a 4b version of the same model
>You are an expert prompt engineer for text-to-image models. Your task is to expand the user's prompt into a highly effective I2VA prompt using the following md file and user's input.
>copy paste prompt guide from VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>User's Input:
>output prompt is 1:1 of what the guide suggests using the same terminology
>video output quality improves with more smoothness
I think I get it. Use the LLM to make good prompts that H3 was trained with because the chinese used qwen3vl to write the english caption pairs.
>>
>>109455923
go ahead, so tired of schizos posting this nonsense here and there with literally worse / simply divergent results
>>
using reference can minimax replace the character in the video like scail?
>>
oh shit I just realized I was using the fl2 model in the reference workflow not the ref2va one when I added the int8 fast node. will test with image and video inputs now.
>>
t2v miku fight:

https://files.catbox.moe/7h4we7.mp4
>>
>>109455929
Minimax can't do shit, don't bother. Biggest psyop of the past year
>>
>>109455929
yes it can
>>109455951
sorry (((Zeev))) but you lost
>>
>>109455929
depends on the video, works 33% of the time.
>>
>>109455923
>>109455926
fully uncensored te can and does prevent anomalies due to censored te interpretation of the prompt
and no it has nothing to do with prawns and what not it simply follows prompt as user intends it to be and one shotting it

if it is a qwen it is retarded. one should never use default qwens since they are too restricted for ordinary use let alone as a text encoder unit
>>
>>109455961
that is not how it works and I'm not gonna explain why to a drooling retard
>>
>>109455954
>t. rombach
>>
>>>/wsg/6207737
Still can't believe we got this model bros
>>
>>109455592
based WAITGOD
>>109455692
we already had multiple "talking heads" projects based on wan 2.1 much before ltx, ltx audio was clown tier

wan 2.1 multitalk from more than a year ago:
https://files.catbox.moe/s8796o.mp4
>>
>>109455902
what if i have a 9070XT
>>
File: 1757492368797616.jpg (532 KB, 1448x1448)
532 KB JPG
>>109455583



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.