[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models and Software

Previous: >>109433015

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>LTX-2.3
https://huggingface.co/collections/Lightricks/ltx-23

>Wan
https://github.com/Wan-Video/Wan2.2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
gm saars
>>
https://github.com/huggingface/diffusers/pull/14355
>33b
>26b text encoder
>no base model
>2 transformer files (one for t2v and another for reference shit)
Jesus... I thought Minimax was a serious company for a second
>>
File: 1757583687239419.png (464 KB, 1147x1108)
464 KB PNG
>>109434230
How is sulfur gonna make a coom finetune without base?
>>
Reminder the comfyui doesnt want to support GGUF despite scenarios like with these video models benefiting the most from GGUF, since Q4 Q5 Q6 are infinitely better than int4 in quality, and faster than int8 which needs to offload to RAM.
>>
literal whos
>>
>>109434238
no gguf
not allowed
>>
File: 052428CUI_00001_.png (963 KB, 1216x832)
963 KB PNG
>>
>>109434236
>How is sulfur gonna make a coom finetune without base?
won't be better with BFL, we won't get a flux 3 base model either, it's grim
>>
>>109434230
>>2 transformer files (one for t2v and another for reference shit)
boooooooooooooo
>>
>>109434249
can distill, but gonna wait till we see what the best of the bunch ends up being. Also LTX could end up redeeming themselves as they said they would always release the base already
>>
>>109434230
if they only gave us the guidance distilled version, does that mean that the API is using the full version?
>>
>>109434267
with jews, you win
>>
Chinese Culture
>>
>>109434279
literally lol
>>
>>109434230
>>109434249
If flux 3 dev ends up the same size or like within 20B id rather undistill that (if that's even possible...)
>>
>>109434238
>Q4 [] infinitely better than int4 in quality
Citation needed, especially for convrot int4
>and faster than int8
Also citation needed. In my experience accelerated quant with offloading still beats non-accelerated quant that fits into VRAM, and by big margins too.
Provide evidence for the claims you are shitting up the thread with (You won't because you are a raped schizo).
>>
I hope it knows Michael Jackson
>>
>>109434293
at this point it's more likely to get LTX 3.0 than being able to undistil a video model kek
>>
>>109434212
>mfw Resource news

08/01/2026

>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
https://explorative-modeling.github.io

>EU to get power to enforce rules on AI starting today
https://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786

>FameGrid Auto Color for ComfyUI
https://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui

07/31/2026

>One-take Creation, Flexible Referencing: Introducing Seedance 2.5
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5

>ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
https://github.com/avaxiao/ReToken

>RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
https://github.com/liuxiaobo66/RefineSVG

>ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
https://github.com/H-EmbodVis/ROAD

>PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
https://physomni.github.io

>DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
https://github.com/anonyme610/dinolizer

>Inline Studio v1.2.6 - Flux 2 & Minimax H3 API
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.6

>SAM 3.1 Multiplex
https://huggingface.co/Sparknight/sam3.1-int8-int4-convrot

>Suno Loses AI Copyright Lawsuit to German Music Rights Society GEMA
https://variety.com/2026/digital/news/suno-loses-ai-lawsuit-gema-1236825010

>AI labels to be compulsory on authentic-looking content under EU rules
https://www.theguardian.com/technology/2026/jul/31/ai-labels-to-be-compulsory-on-authentic-looking-content-under-eu-rules

07/30/2026

>AnimeGen: AI Models for Anime Video Generation
https://huggingface.co/collections/aidealab/animegen
>>
Dedistilling is a massive fucking cope anyway.
It completely destroys model quality. You need massive, millions of images/videos massive finetuning to recover from that. There are like only 3 people currently training at that scale: nutbutter, td_russell, a certain furry lolcow
>>
File: 053708CUI_00001_.png (867 KB, 1216x832)
867 KB PNG
>>
>mfw Research news

08/01/2026

>VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
https://arxiv.org/abs/2607.23472

>Layering Virtual Try-On
https://arxiv.org/abs/2607.22924

>What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
https://stevencylu.github.io/PeakPatch

>S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
https://arxiv.org/abs/2607.28164

>Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
https://arxiv.org/abs/2607.24731

>FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
https://arxiv.org/abs/2607.27800

>UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
https://umi3d-project.github.io

>UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
https://arxiv.org/abs/2607.23373

>Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
https://arxiv.org/abs/2607.26411

>LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
https://arxiv.org/abs/2607.25962

>MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
https://arxiv.org/abs/2607.28589

>What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
https://arxiv.org/abs/2607.28526

>Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
https://arxiv.org/abs/2607.24440

>Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
https://arxiv.org/abs/2607.22872

>Unifying Adversarially Robust Model Experts in Vision-Language Models
https://arxiv.org/abs/2607.27897
>>
>>109434312
most people don't use undistilled models, and most people don't train models, and ostris will have a training adapter for it tomorrow afternoon anyway.
seems like a completely pointless fud.
>>
>>109434324
None of them ever trained a video model, much less a massive one btw.
It's joever if no base.
>>
>>109434248
>>109434328
Whatever this artstyle is, I hate it
>>
>>109434329
STOP SPAM
>>
>>109434334
He can't yap about creating the most ebinest training adapter ever as much as he wants but the loras won't be in good quality.
>>
>>109434334
>ostris will have a training adapter for it tomorrow afternoon anyway.
the same ostris who made an adapter for flux.1, and it worked like as
>most people don't train models
the guy that made sulfur was waiting for Minimax so that he could train it as well, he won't be able to do it because those chinese fucks don't respect us enough to give us the base model, say what you want about the LTX jews, but they gave us a base model at least
>>
>>109434345
*can yap
>>
>>
File: spam.png (35 KB, 881x263)
35 KB PNG
>>109434322
>>FameGrid Auto Color for ComfyUI

Do you even check what you share on your spam? This is just a vibecoded slop custom node made to advertise some grifter AI site
>>
File: AniStudio-00431.png (1.36 MB, 1024x1408)
1.36 MB PNG
POST GENS
>>
>>109434360
the fuck is this schizo doing here, get the fuck back to your containment zone
>>
File: debo_pc_k2_00071_.png (2.32 MB, 1920x1280)
2.32 MB PNG
>>109434357
so don't use it? not everything is for you
>>
>>109434345
>>109434347
>you can't train z-image turbo! literally no one has every trained z-image turbo, and z-image turbo is bad!
again, seems like a pointless fud.
>>
>>109434357
It's a vibecoded bot that probably reads top posts on r/StableDiffusion and searches a few key words on arxiv.
The malware spammer faggot in fact, doesn't check what he spams.
>>
>>109434364
>you can't train z-image turbo
you literally can't yeah, name one succesful serious finetune of zimage turbo
>>
>>109434300
You'd think in a thread about AI people like you would know that you can simply use AI to verify whether a statement about AI is true or not.
>>
File: 054635CUI_00001_.png (899 KB, 1216x832)
899 KB PNG
>>109434338
it's ludo doe
>>
>>109434371
he could just put it in a rentry instead of three huge walls of text if sharing news was the goal but its obvious its not
>>
>>109434364
Lots of ZIT loras commonly give body horror. They are also nightmare to combine. (Inb4 esoteric block swap combo cope) It's needlessly difficult to train, I sincerely doubt you trained any, if you call concerns about training on distilled models a "fud".
>>
File: 9f5fa7.gif (768 KB, 240x240)
768 KB GIF
can h3 do this?
>>
you already scrolled past anon. the news can't hurt you anymore
>>
>>109434376
What are you rambling about faggot, post your proof or get the fuck out.
If you mean you asked bunch of loaded questions to Gemini Flash to confirm your priors, you are Indian.
>>
>>109434363
just like when you shared literal malware for weeks?
>>
>>109434360
based gguf supporter
>>
>>109434392
Don't see you providing proof to the contrary. Not-so-subtle way of admitting defeat, bud.
>>
>>109434379
i call general meltdowns anytime a new model pops up a fud.
for 4 weeks straight "someone" shat up this general all day everyday about how bad krea is and how goated zit was, literally the best model, super easy to train, best finetunes, super based!
now zit is slop trash that is untrainable because a new video model is distilled.
>>
>>109434397
>getting mad at text you can simply ignore on 4chinz
>>
>>109434406
>how goated zit was
for inference yes, nothing was said about training, maybe you're just too stupid to understand the nuance
>>
total namefag holocaust
>>
>>109434406
its always the same guy. i dont know why you retards keep replying to him since his typing style is obvious
>>
>>109434406
ZIT mogs Krea was one schizo spamming the general while posting next to no gens
And ZIT was in fact difficult to train.
How are these supposed to be contradictory?
>>
File: ComfyUI_hgdf_04105_.jpg (570 KB, 1792x2304)
570 KB JPG
>>
>>109434423
>And ZIT was in fact difficult to train.
what the fuck are you on about? its the model that instantly had a fuck ton of loras basically day 2
>>
So apparently Q8 > FP8 according to this guy using Wan2.2 and LTX for comparison. And use Q4_K_M for efficiency. Any lower is just coping.

https://www.youtube.com/watch?v=octq1gJ0dkM
>>
>>109434434
>loras
are you the anima cope anon as well?
>>
>>109434404
>Make nonsensical claim that goes against established facts
>Asked to provide any evidence whatsoever to back up said the claim
>Nuh-uh! It's your job to provide evidence on the contrary to disprove my claims that I provided zero reasons for you to believe that they are true chud!
I can take my time to provide the numbers but why would I bother to do so with a schizo or troll? I know int8 runs faster than q quants on my system. You are just some random fag on the internet and nowhere as important as you think you are.
>>
Does the cover mode in ace step 1.5 actually do anything?
>>
>>109434446
are u implying the word "training" u used meant not what everyone does (train loras) but the 0.00000000001% of people are doing, full finetunes?
and even for those why would people do full finetunes of a model that is turbo, after base was announced for it?
>>
whats the difference between Q8 and INT8?
>>
>>109434465
Only chuds use GGUF.
>>
>>109434446
now you are just being silly
>>
>>109434434
>Lots of ZIT loras commonly give body horror.
>They are also nightmare to combine.
The difficulty was training loras that didn't suffer from these issues due to training on a distilled base you pedantic midwit.
>>
>>109434465
Q8 is an illegal format DO NOT use
>>
>>109434458

At 0.5 Cover strength the outpost songs and singing is similar to input song and singing. So it does a cover of a song, duh?
>>
>>109434469
>chuds
*chads, sorry for the typo
>>
>>109434465
Q8 is unaccelerated and slow.
Int8 is accelerated (on Turing and newer) and fast.
Naive int8 has lower quality than q8.
Convrot int8 has higher quality than q8 (And also faster).
>>
>>109434465
Q8 trims layers, int8 converts to integers. The CPU has a better time with ints so it's better when offloading but if you can fit the whole thing in vram them who cares
>>
>>109434475
>Lots of ZIT loras commonly give body horror.
someone claiming one of the easiest models to train, zit, was actually 'hard to train' is enough to spot a retard trolling but now with another low iq troll lie like that ill just take your concession.
>>
>>109434485
>Convrot int8 has higher quality than q8
citation needed.
>>
>>109434485
>Convrot int8 has higher quality than q8
can you show a grid or something? It always looks worse for me
>>
File: faggot ass liar.png (180 KB, 1055x1245)
180 KB PNG
>>109434485
>Convrot int8 has higher quality than q8
*loud incorrect buzzer*
>>
>>109434487
so i should use int if i am using mixing gpu and cpu memory?
>>
>>109434499
can we see the images?
>>
>>109434504
yes unless you are at capacity even with that then you would use Q4
>>
>>109434505
nyo
>>
>>109434455
>i can but i won't
tee~hee same here :3
>>
>>109434499
Link to repo?
>>
File: ComfyUI_Krea2_turbo_00006.jpg (2.97 MB, 3840x2160)
2.97 MB JPG
I can do a direct comparison with Krea2 between int8 convrot and q4_k_m gguf
>>
>>109434512
thanks
>>
>>109434528
No one would dispute int8 would have higher quality in that case.
Tell us how much RAM and VRAM you have and then post inference times for both though. (Total and per step time.)
>>
>>109434499
Liar faggot.
https://github.com/BobJohnson24/ComfyUI-INT8-Fast/blob/main/Metrics.md
Only Chroma and ZIT had better quality on GGUF. HiDream, Qwen, Klein, Anima all had better quality with Convrot.
That's why you didn't put the link.
RAPE AND GAPE ALL GGUF FAGS
>>
>>109434445
Yes and GGUF Q8 being better than FP8 is the case on almost any model quantized from a more precise one - video, image, llm.
>>
File: Ideogram__00225_.jpg (1.81 MB, 2048x2048)
1.81 MB JPG
>>
>>109434561
>Only Chroma and ZIT had better quality on GGUF.
and you don't care about that subhuman?
>>
>>109434535
int8_convrot (original turbo model)
50s inference time on a 3060 12gb with 24gb of sys memory and a i7 6700
8 steps, euler ancestral, simple scheduler, cfg 1.0
3:4 ratio at 2mp (1216x1664)
6.2 seconds per iteration
>>
>>109434567
Nyt you prompt separately for those stickers?
>>
>>109434580
how much memory did it use?
>>
File: 1_00077_.jpg (2.53 MB, 4064x2284)
2.53 MB JPG
>>
>>109434571
Idc about Chroma, and I consider ZIT obsolete at this point but nevertheless bigger issue is that it doesn't debunk my claim (FACT) that convrot int8 has better quality q8 on average. And your cherry picked cope post was pathetic.
>>
>>109434583
They don't have their own bboxes. They're part of the bike bounding box prompt and high level description
>>
>>109434580
And Q4?
>>
>>109434567
nice
>>
>>109434580
Q4_k_m gguf of original turbo model (from realrebelai)

same settings

1m9s inference time at 8.7 seconds per iteration

>>109434587
int8?
all 11.5gb of vram and about 50% of sys memory with 16gb paged (not max)
>>
>>109434419
Cope. Your model sucks.
>>
I said 5 minutes per second, I was off by a factor of 2.5
no, the other direction
>>
>>109434617
and how much did q4 use?
>>
>>109434617
Yes, even q quant half the size is still slower due to need to dequant.
Even if you were spilling into system memory with int8 it would still be faster (up to a point)
>>
>>109434607
Being against ZIT is like being antisemitic, transphobic, racist, and homophobic, all ZIT posters are trans women that you should respect, sweaty.
>>
>>109434617
q4_k_m gguf uses about 9.5gb of vram and 35% sys memory during inference and 65% after vae decode.
>>
not that ever seems to care any more (or ever did) but what's the word on flux 3, no ETA?
>>
>>109434655
DOA
>>
God I love blonde fennecs, I will make one my wife.
Fucking hate silver cats though, will behead with the swiftest of ease
>>
>>109434609
No loras?
>>
>>109434655
I may know things that others do not but I cannot say exactly
>>
>>109434655
Within the next few weeks and months because they're "rigorously safety testing" it.
>>
>>109434655
Expectations are low, it'll be huge and it'll be pozzed to hell and back
>>
File: comfy plebbit h3.png (434 KB, 763x702)
434 KB PNG
>Minimum requirements for 480p video on this model is a 3060 with 12GB vram, 32GB of system ram and a good nvme SSD. We tested generating a 5 second (124 frames) 480p (864x480) video on this system and it took a bit less than 9 minutes end to end (20 steps). I can pretty much guarantee it will also work on 8GB vram too but we did not test that.
480p 5 seconds 20 steps is half a minute per step. This is almost certainly 8 bit quant too, I dread it if this was int4, so we will need a good step distill lora for sane inference speed/resolution on mid range cards.
>>
>>
>>109434680
low res will be garbage regardless, it always is
>>
>>109434637
>Even if you were spilling into system memory with int8 it would still be faster (up to a point)
Absolutely fucking not. The moment you spill even 10% of the model into RAM your speed craters unless you are running 9000 MT/s quad-channel RAM on a Threadripper setup. Spending a little compute on dequanting is a cheap and quick. Moving assloads of data over PCIE is not.
>>
>>109434655
Black Forest Labs seeing how much they can mindrape, censor, and lobotomize a model while it still outputs something ok is hard work, anon, it takes time.
>>
>>109434680
he really needs to stop posting ads in the main sd reddit
>>
File: file.png (408 KB, 1870x937)
408 KB PNG
>>109434212
>>109434212 [Posting this here in case if the thread got archived.]

This is my battlestation setup right now about Stable Diffusion. [sd-webui-forge-neo]
Are there new updates I need to download for SD? Or new .safetensors files for new and better default generations? Or is the IllustriousXL still the best at this time?
>>
>>109434680
3090 is 3x the speed of 3060 at least so im good
>>
>>109434687
That's not how any of this shit works retard. The slowdowns are fairly linear to the percentage offloaded. It's a cool bonus of how fused ops work in diffusion. What you said only applies to auto-regressive models like LLMs.
BF16 ZIT with circa 2gb offloading was still faster than q8 ZIT fitting fully on my system.
>>
>>109434680
just text to video?
>>
>>109434702
>slowdowns are fairly linear to the percentage offloaded
>linear
lmao retard
>>
File: Ideogram__00226_.jpg (1.62 MB, 2048x2048)
1.62 MB JPG
>>109434668
My own lora trained off 1000+ amateur photos from 2009-2011.
An attempt to try take the overly cinematic edge look off Ideogram. Picrel is with no lora
>>
Minimax is gonna be shit.
LTX Is still the GOAT.
The ZIT of video genning.
>>
>>109434707
Your schizophrenic sense of humor does not make the statement incorrect.
>>
>>109434687
the int8 model is over 12gb, so I am definitely "spilling" over to system memory
but it's still faster and better quality than the q4 gguf
>>
>>109434707
Hi ani
>>
>>109434655
since when BFL ever deliver? it's gonna be a bloated piece of shit like flux 2 dev, and because they never give base models we won't even be able to make serious finetunes of it (if it turns out to be good and small enough by some miracle...)
>>
>>109434718
>better quality
By what metric? Your two images are practically identical, and what differences there are might as well be seed variations (because empty latents will be different with different models even with the same seed value).
>>
>>109434718
that is what we were saying the whole time. Q8 has higher quality, int8 better speed, q4 absolute lowest
>>
File: 070512CUI_00001_.png (831 KB, 1216x832)
831 KB PNG
>>
>>109434698
but you'll probably want to bump the resolution so that's at least a 1.5x slowdown
>>
>>109434680
>Don't be scared to try our default template
Will it be 70% useless bloat like all the recent templates
>>
>>109434760
yes, but it's all hidden in a single group node



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.