[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Smell Your GPU Edition

Discussion and Development of Local Image, Video, and Music Models and Software

Previous: >>109430870

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>LTX-2.3
https://huggingface.co/collections/Lightricks/ltx-23

>Wan
https://github.com/Wan-Video/Wan2.2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
What the fuck is this collage?
>>
why no kino in collage?
>>
Blessed thread of frenship
>>
File: i_00171_.png (1.47 MB, 768x1376)
1.47 MB PNG
so in 14 hours it comes out?
>>
>>109433028
this goes hard
much love from iraq
>>
>>109433015
what the fuck is that faggollage
>>
>>109433028
>so in 14 hours it comes out?
yep, glad that I will sleep soon, I know I'll wake up with a kino present
https://modelscope.cn/models/MiniMax/MiniMax-H3
>>
File: 1708024613797546.jpg (1.18 MB, 1504x1200)
1.18 MB JPG
>>109433015
>>109432905
>>109432910
>>109432959
Once again I'm asking you for a wallpaper that I can use in normie space without getting shot or cancelled.

Pic rel is current pape.
>>
>>109433045
you're gonna get hilariously trolled and owned and made fun of by normies if they see that shit on your computer
it's so blatantly AISLOPPA.
>>
another trollbake, yawn
>>
File: 1781297662682168.jpg (2.64 MB, 3840x2025)
2.64 MB JPG
>>
sometimes you need a strict baker to clean anons palette
>>
File: 1785607126521214.jpg (6 KB, 250x240)
6 KB JPG
>>109433066
open your mouth so i can clear your palette
>>
File: i_00237_.png (1.33 MB, 768x1376)
1.33 MB PNG
>>
>it's even better than acestep at making music
kek
https://litter.catbox.moe/00k8tl4h39e0p7nd.mp4
>>
>5 1girls, asian, sing
get new material, white incels.
>>
>>109433100
now the question is whether you can get consistent music if you string together a bunch of clips
>>
File: Cumfy_00304_.png (1.2 MB, 840x1256)
1.2 MB PNG
>>
File: Krea2_turbo_01119_.png (1.51 MB, 1368x768)
1.51 MB PNG
>>109433045
>>
>>109433124
>gaming chair
>>
File: obligatory.jpg (239 KB, 1920x1080)
239 KB JPG
>>109431232
>>
>she doesnt own a Gaming(tm) brand chair
>>
>>109433145
>turned an oven dodger into a rice cooker
i dig it
>>
Which krea mix are the cool kids using these days
>>
File: i_00107_.png (1.27 MB, 768x1376)
1.27 MB PNG
>>
how is the chroma krea2 lora?
>>
>>109433242
I recommend it
>>
File: i_00250_.png (1.43 MB, 768x1376)
1.43 MB PNG
>>
>>109433205
flashback to when we had a schizo spamming george floyd
>>
>>109433028
then in 48 hours the comfy nodes will be ready
>>
>>109433258
Fake AI
>>
File: i_00140_.png (1.28 MB, 768x1376)
1.28 MB PNG
>>109433260
krea2 knows about George Floyd too
>>
>>109433260
I'm sure he'll come back to try H3 out
>>
>>109433258
He has the smile of a king and gentle person
>>
File: Cumfy_00327_.png (1.15 MB, 768x1368)
1.15 MB PNG
>>
It is a shame they are releasing the model on Sunday, the day set aside to worship the Lord. (I almost said the Lord's day but that would have been a mistake as every day is the Lord's).
>>
>>109433318
chang cares little of God and has historically released on the sabbath
>>
File: 07318-3655656521.png (1.85 MB, 1280x1280)
1.85 MB PNG
>>
File: 47556542.webm (968 KB, 320x320)
968 KB
968 KB WEBM
what's the first thing you will generate when h3 drops?
>>
>>109433332
ur mom lol
>>
>>109433332
Did they say how big the model is?
My poor 3070 still has life left in it
>>
>>109433332
pornography of ur mom
>>
File: LTX-2_00057-2.webm (3.93 MB, 1920x1056)
3.93 MB
3.93 MB WEBM
Skateboarding is always a good test
>>
>>109433332
some yugioh dueling kino
>>
>>109433332
ur mom lol but probably princess peach undressing or something basic to start with
>>
>>109433340
no, i don't know why they are keeping it a secret
>>
>>109433364
>we know it can run on an 8gb 3060

>but still no parameter/file size
very strange indeed
>>
>>109433369
we cant know the filesize until the moment of release due to the quantum degradation and bit entropy
>>
>>109433344
can you make one where she dismembers a luddite?
>>
>>109433376
you scare me a bit not gonna lie
>>
File: Cumfy_00318_.png (991 KB, 768x1368)
991 KB PNG
>>
>>109433384
her farts probably STINK like SHIT also elongated torso
>>
File: 00202-2411629685re.png (2.99 MB, 1920x1273)
2.99 MB PNG
>>
3090 status: ready
feed me h3
>>
>>109433369
Wait I have a chance?
>>
>>109433369
hmm, 38gb for fp8
god knows on the TE, maybe 20gb
>>
>>109433340
no

but we've had quite viable VRAM->RAM swap offloading for most models so far, most likely that works again in some way or another.

if the model was completely too large for almost all consumer cards, I think they'd have warned.
>>
>>109433015
Haven't been keeping up, will MiniMax H3 be able to do porn or is it useless?
>>
>mfw Resource news

08/01/2026

>FaceFusionYu: ComfyUI integration of the FaceFusion 3.5.4 processing engine
https://github.com/thenotrealuser/facefusionuy

>Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
https://explorative-modeling.github.io

>EU to get power to enforce rules on AI starting today
https://www.taipeitimes.com/News/front/archives/2026/08/02/2003861786

>FameGrid Auto Color for ComfyUI
https://github.com/ultramuseart/famegrid-auto-color#famegrid-auto-color-for-comfyui

07/31/2026

>One-take Creation, Flexible Referencing: Introducing Seedance 2.5
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5

>ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
https://github.com/avaxiao/ReToken

>RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
https://github.com/liuxiaobo66/RefineSVG

>ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
https://github.com/H-EmbodVis/ROAD

>PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
https://physomni.github.io

>DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization
https://github.com/anonyme610/dinolizer

>Inline Studio v1.2.6 - Flux 2 & Minimax H3 API
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.6

>SAM 3.1 Multiplex
https://huggingface.co/Sparknight/sam3.1-int8-int4-convrot

>Suno Loses AI Copyright Lawsuit to German Music Rights Society GEMA
https://variety.com/2026/digital/news/suno-loses-ai-lawsuit-gema-1236825010

>AI labels to be compulsory on authentic-looking content under EU rules
https://www.theguardian.com/technology/2026/jul/31/ai-labels-to-be-compulsory-on-authentic-looking-content-under-eu-rules

07/30/2026

>AnimeGen: AI Models for Anime Video Generation
https://huggingface.co/collections/aidealab/animegen
>>
if ur gpu has never genned 24/7 for at least 2 weeks, lower ur tone when speaking here
>>
only oldfags remember flux3
>>
>mfw Research news

08/01/2026

>VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation
https://arxiv.org/abs/2607.23472

>Layering Virtual Try-On
https://arxiv.org/abs/2607.22924

>What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
https://stevencylu.github.io/PeakPatch

>S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
https://arxiv.org/abs/2607.28164

>Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
https://arxiv.org/abs/2607.24731

>FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
https://arxiv.org/abs/2607.27800

>UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
https://umi3d-project.github.io

>UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
https://arxiv.org/abs/2607.23373

>Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
https://arxiv.org/abs/2607.26411

>LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
https://arxiv.org/abs/2607.25962

>MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
https://arxiv.org/abs/2607.28589

>What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
https://arxiv.org/abs/2607.28526

>Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
https://arxiv.org/abs/2607.24440

>Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
https://arxiv.org/abs/2607.22872

>Unifying Adversarially Robust Model Experts in Vision-Language Models
https://arxiv.org/abs/2607.27897
>>
>>109433440
>>109433446
instead of spamming this trash every time with duplicate info that most dont care about, make a rentry and link it at the top and stfu
>>
>>109433440
>>109433446
fuck off debo
>>
>>109433440
>>FaceFusionYu: ComfyUI integration of the FaceFusion 3.5.4 processing engine
>https://github.com/thenotrealuser/facefusionuy
>404
>>
File: Krea2_turbo_01142_.png (1.37 MB, 1024x1024)
1.37 MB PNG
>>
>>109433439
probably not extensively on release. chinese company expecting public attention.
>>
File: debo_pc_k2_00066_.png (2.21 MB, 1920x1280)
2.21 MB PNG
>>109433482
weird. they must have pulled it
I'll remove the link, thanks
>>
https://github.com/Comfy-Org/ComfyUI/pull/15210
lmao, the PR for minimax is here, but he closed it, he probably made a mistake, let's see
>It uses Qwen3-VL-32B as the encoder (50 layers of it)
>There is a message in this commit: "model must be a split MiniMax H3 transformer"

>Minmax H3 is about 28B-33B dense model
>50 layers.
>Main DiT blocks:
>4 Attn, 5376 x (56 heads, 128 dim)
>2 MLPs, 5376 x 14336.
>AdaLN 6Ă—(2688Ă—5376Ă—3)
bruh...
>>
predict how big the shitstorm tomorrow will be. i already expect anons to upload shitty ltx gens claiming theyre h3
>>
File: DOA.png (57 KB, 360x360)
57 KB PNG
>>109433504
I can't run that wtf
>>
File: 1766513239124149.png (99 KB, 1700x298)
99 KB PNG
>>109433504
yeah I confirmed the sizes with Claude
>>
>>109433509
ever heard of krea?
or qwen image
that, again
>>
File: 7808357930589745.jpg (958 KB, 1149x2056)
958 KB JPG
>>109433504
yay local
>>
>>109433504
My 3060 (8gb) is going to crush this shit
>>
>>109433504
nah that's too big I'm out, back to flux 3 waiting room
>>
File: 1719435836295.gif (3.47 MB, 500x500)
3.47 MB GIF
just got hit by the hardest oom yet
shit shut down my computer
>>
krea2 boys eating good :)
https://civitai.red/models/1055425/ilulu-miss-kobayashis-dragon-maid-illustrious-krea-2?modelVersionId=3189920
>>
>>109433531
Is that dynamic enough for you, pal?
>>
>>109433531
ltx in comfy makes my pc bluescreen, no idea why
>>
>>109433536
kek
>>
>>109433532
look at that fucking trigger
>>
>>109433536
lmao
>>
>>109433504
Get multimodal'ed, Flux 3 dev will be as big too btw!
>>
File: 0_00080_.png (1.66 MB, 1024x1024)
1.66 MB PNG
>>
>>109433332
definitely porn.
>>
>>109433504
that's too big, I doubt sulfur will make a finetune of that, would be too expensive
>>
>>109433553
sar please understand we must having use trigger
>>
>>109433504
I can probably run q4 with 24gb vram and 128gb ram
wonder how that will be
>>
File: fuck this hobby.gif (405 KB, 220x293)
405 KB GIF
>>109433504
>>
>>109433586
>q4
the quality will be so ass though, I kinda wanted the model to be small enough so that I could run it on int8 and get that nice x2 speed increase
>>
i have hope
>>
>>109433504
"the Z-image turbo of videos" they said, it's gonna be small and optimized they said...
>>
>>109433553
it's almost like people read no tutorials on how to train a lora. Might as well just forego the lora and just use the trigger as prompt.
>>
is comfy still giving out 300 credits per month? how much is a h3 video?
>>
>>109433600
>"the Z-image turbo of videos" they said
Who?
>>
File: ayaaaaaaa.png (489 KB, 976x549)
489 KB PNG
>>109433504
>28B-33B
Looks like LTX and Wan will never die kek.
>>
>>109433332
A girl getting punched in the face.
>>
>>109433602
nope and they upped the cloud price
>>
File: h3 cloud.png (96 KB, 685x265)
96 KB PNG
>>109433504
LOL BODIED YOU RETARDS. only the brownest of localkeks fail to see it
>>
>>109433609
Wan 2.2 is 14Bx2
>>
>>109433332
>what's the first thing you will generate when h3 drops?
nothing, I can't run that big boi after all, oh well >>109433504
>>
>>109433124
Workflow pls
>>
>>109433615
yeah, that means you can put only 14b to your memory, unload it, then put another 14b, you can't do that when it's a unique 33b model
>>
>>109433504
If I were them I wouldn't release anything, what's the point? who can run that??
>>
>>109433622
why haven't they figured out how to load groups of layers at a time?
>>
>>109433614
>>109428169
>he meant you can use ur 3060 to render the nodes 2.0 ui at 30 fps with browser accel while the video generates in comfy cloud
>>
>>109433504
Why is everyone dooming? The only thing that matters is the VAE compression ratio. If it's the same as LTX, then the model is almost as fast as LTX. You stream the weights off RAM / SSD using dynamic VRAM (which LTX also requires on most setups).
>>
>>109433596
even at that quant it would still be leagues ahead of ltx and wan considering how big of an improvement it is
>>
>>109433614
Comfy is so funni, next time a 100b model will appear he will say it can be run on a gtx1600 (thanks to their MAGIC optimisation which is... putting everything in the fucking cpu) lmao
>>
>>109433614
the browns are the only ones who will have a problem.
they will run to youtube or reddit and download pardeeps retarded as fuck workflow with 40 custom nodes, and a special seedvr upscale polish pass at the end to really slop the tits off of it. then they will complain that they are ooming "their" h100s.
>>
>>109433634
yep that's what i'm thinking so far
>>
Reminder the comfyui doesnt want to support GGUF despite scenarios like with these video models benefiting the most from GGUF, since Q4 Q5 Q6 are infinitely better than int4, or int8 with offload.
>>
>>109433632
>You stream the weights off RAM / SSD using dynamic VRAM (which LTX also requires on most setups).
dynamic vram doesn't work, I always OOM with that shit, I think I'm gonna take a break, ComfyUi is pissing me off, and the model we have so far have all been disappointing
>>
>>109433504
Are 5090chads safe or ...?
>>
my pagefile is ready
>>
>>109433636
Going to have to start measuring /its in hours
>>
>>109433647
I don't think so, if it's a 33b model then on int8 it's probably 34gb just to load it, now you account the memory to make the video and you overflow your VRAM space
>>
>>109433657
back to slopping on wan
>>
>>109433504
looking at the code, it's probably a non-distilled model, how many steps would we need? 30? jesus...
>>
>>109433402
is this a k2 lora?
>>
Let's not start doomposting. LTX 2.3 requirements look hefty but it's still pretty speedy on modest hardware
>>
Can tomorrow get here faster please?
https://files.catbox.moe/sxv4wt.mp4

>>109433674
It will be 1/2 the speed of LTX but about the same size.
>>
File: 44477.png (95 KB, 1840x824)
95 KB PNG
>>109433674
this. the fudders don't want you to have hope for the new age of kinos
>>
File: Ideogram__00215_.jpg (776 KB, 1536x2048)
776 KB JPG
>>
File: vramlet.png (24 KB, 938x156)
24 KB PNG
once again, local hardware stagnation continues to hold local back. these requirements are quite small for a 2026 model. 96GB cards should be the norm by now, but there is a coordinated effort to kill off consumer hardware to force reliance on centralized API models.
>>
>>109433504
>you need a RTX6000 to fully load this model on 8bit
kek, I'll pass, that model is good, but I'm not spending 13000 dollars for this
>>
>>109433622
yes you can retard. Its called offloading. Comfy does it automaticlly.
>>
Its about LTX size model wise, but it has half the temporal and spatial compression so it will be about half as fast. It will be worth it though.
>>
>>109433697
>Comfy does it automaticlly.
it doesn't work subhuman, his automatic offloading shit always OOM my card, Comfy has no idea how to make a good memory managment system, I used to use multigpu to manually do that and control that shit but I can't anymore because it doesn't work with dynamic vram anymore, FUCK THIS PIECE OF SHIT OF A SOFTWARE, AND FUCK YOU
>>
>>109433712
uhh, chuddie, stop using UNSUPPORTED GGUF NODES!
>>
>>109433712
stop being poor. You need 64GB+ ram for local video. Vram 12GB+
>>
holy meli
>>
>>109433504
I really didn't expect the model to be so big, I thought they were smart enough to understand that 20b is the absolute limit, do they really believe there will be a lot of people that will be able to make loras from a 33b model? lmao
>>
>>109433712
wan2gp chads can't stop winning
>>
>>109433720
I have 64gb of ram, the ram isn't the issue, it always completly fills my vram space and I fucking OOM crash, "automatic offloading" my ass
>>
Kroma RL, seems like he is actually doing a RL himself this time. Seems to massively increase quality
https://huggingface.co/lodestones/Kroma/blob/main/kroma-v0.1-rl-mild.safetensors
>>
>>109433726
then you are doing something dumb cause it works fine for me. I can do full fp16 wan2.2 easily without enough ram to fit both / te at once
>>
>>109433722
Why should they give a shit? Doesn't matter for API inference powered by B200 or whatever the fuck. They release model's as ads basically. If people can't run it despite being open weights and have to paypig for API, that's even better. They both farm cred and still get money this way.
>>
>>109433720
>>109433726
u need 96gb+ for video if you dont want to touch the disk
>>
File: 1760684441361907.png (354 KB, 500x500)
354 KB PNG
>>109433684
>96GB cards should be the norm by now
yeah but we don't live in a normal world, we are at the mercy of Nvidia's monopoly so...
>>
>>109433733
>If people can't run it despite being open weights and have to paypig for API, that's even better.
what?
>>
>>109433727
His schizo training might actually be salvaged this time if he is going to do proper RL, but I wouldn't get my hopes up.
>>
>>109433684
there's very little any of us can do about that
>>
File: Kroma_RL_test.jpg (1.29 MB, 3072x1960)
1.29 MB JPG
he is testing a RL atm at different settings
>>
>>109433504
BFL please don't be as retarded as them, please save us...
>>
>>109433657
There is no reason not to run Q4 these days. It's like 90% of the quality of Q8. Especially in cases where "quality loss" is abstract or subjective like in image diffusion you won't even notice the difference.
>>
>>109433727
Does anyone in his furcord here know the difference between rl and rl mild btw?
I am assuming more and less aggressive rl, but a proper clarification on what these precisely entail would be nice.
>>
>>109433749
regular left, full RL mid, RL at lower strength right
>>
>>109433749
>the skin is even more plastic
good job kekestone!
>>
>>109433727
apparently there's another one that is just "rl" one hour earlier.

>>109433743
he basically always managed to train a checkpoint in the end
>>
>>109433710
># frames 17k+5 <-> latents 5k+2, 16x spatial
From the PR. It's a weird ass VAE that maps 17 frames to 5 latent frames, with 16x spatial compression. So somewhere in between the compression of Wan (4x8x8) and LTX (8x16x16), but closer to LTX. It will be several times slower than LTX but probably still faster than Wan per step.
>>
>>109433759
are you blind. The RL has far far better skin
>>
File: it's genuinely over.jpg (2.04 MB, 7961x2897)
2.04 MB JPG
>>109433752
>Especially in cases where "quality loss" is abstract or subjective like in image diffusion you won't even notice the difference.
Q4 literally stops listening to your prompts but go off king
>>
>>109433671
Which part?
>>
>>109433710
>Its about LTX size model wise
not at all, LTX is 23b, that model is 33b
>>
>>109433750
>filename
Desu it would be fine by me if they also give us a kino size distill like Klein again.
I am not sure if they would though.
>>
>>109433773
no the TE is. People are guessing on the actual model
>>
>>109433769
that image is 2 years old
>>
>>109433752
>>109433769
I bet int4 convrot mogs it while running at least double the speed.
>>
>>109433778
>no the TE is.
the te is 26b, it's a truncatured version of qwen 3 32b (50/67 layers) >>109433522
>>
File: kroma_RL_test_2.jpg (1.21 MB, 3072x1960)
1.21 MB JPG
left regular, mid full RL, right RL at lower strength
Its a massive improvement
>>
>>109433779
And nothing changed about Q quants.
They are ass and dead end for diffusion.
>>
>tfw nochekaiser881 going full hard on krea2.
tick-tock anima bros...you guy is moving on hehehe.
https://civitai.red/user/nochekaiser881/models?periodMode=published&period=AllTime&baseModels=Krea+2
>>
>>109433774
how long did it take for them to release klein after flux 2 was released?
>>
>>109433786
These comparisons would be far more meaningful if you also include images without any lora
>>
>>109433791
2 months
>>
>>109433769
Out of all of them, Q4 and Q5 are the only ones who adhered to the prompt. Everything else put her in Tokyo and gave her dark skin instead of light black (grey).
>>
>>109433752
holy vramlet cope
>>
>>109433786
Wow all three are slop !
>>
>>109433791
A month I think?
>>
>>109433793
the left one
>>
>>109433769
anything to do with AI is way too jeeted at this point to take anything at face value.
>>109433786
what a fucking mess
>>
>>109433797
Sorry I only have 32G and I refuse to offload for marginal gains.
>>
>>109433774
>>109433791
the problem is that when they made that flux 2 announcment, they already said on their blog that they were gonna make a klein version, the flux 3 blog only has "dev", we're fucked lol
>>
>>109433786
The middle one is legit amazing. This seems to have fixed krea's skin texture issues
>>
File: quants.png (18 KB, 886x157)
18 KB PNG
kek, aged like wine
>>
>>109433683
Why do people keep trying to realismmaxx Krea when Ideogram is right there
>>
>>109433808
its over
>>
>>109433816
ideogram knows a fraction as much / is boring nsfw wise. Krea can do booru stuff with realism style
>>
who does this fudder work for? he does his localkek spam literally every time a new model releases which deprecates the api ones
>>
>>109433789
>the model that already knows all the characters doesn't get as many character loras
>guy known for training character loras is training them on a model that actually needs them
as you can clearly see this means that anima is shit and a dead end
>>
>>109433808
>>109433818
dev can be of resonable size though, dev 1 was 12b, dev 2 was 32b, I doubt they'll gonna go for ~30b again, they got clowned so hard they probably learned their lesson
>>
>>109433826
i don't see any localkek spam?
>every time a new model releases which deprecates the api ones
oh that must be why, because that never happened
>>
>>109433801
Add kroma lora without the lr into the mix too then
>>
H3 is gonna be amazing
https://files.catbox.moe/pclaeg.mp4
>>
>>109433834
klein deprecated nana banana and now these new models will btfo seedance
>>
File: mfw.png (1001 KB, 1280x720)
1001 KB PNG
>>109433504
I always refuse to go under 8bit for diffusion models, so I think I'm gonna pass on this one
>>
>>109433786
>SD1.5 vs realityDreamshaper1.5 vs perfectbodyMysticAnime1.5
>>
>>109433848
>klein deprecated nana banana
now that's a quality bait
>>
>>109433861
cope. we will be getting videos while you get fired from your call center because no one cares about your fud anymore
>>
the chinese really don't care lol. Its not gonna take much to train this into a hentai model
https://files.catbox.moe/q5pdbt.mp4
https://files.catbox.moe/sxv4wt.mp4
https://files.catbox.moe/wz7x8g.mp4
https://files.catbox.moe/7yr24t.mp4
>>
>>109433848
true, klein is the top ranked image edit model in the world
..on the plastic safetymaxxed leaderboard!
>>
>>109433868
I'm sure the finetuners will be delighted to spend tens of thousands of dollars to train a model that can be correctly run only by 3 people.
>>
File: kroma rl.png (36 KB, 947x151)
36 KB PNG
Never mind he is doing some incomprehensible bizarre nonsense again.
>>
>>109433877
those are just different ranks it looks like
>>
>>109433860
It brings me great paint to see trainers with such poor aesthetic sensibilities. It's all so ugly.
>>
>>109433848
what even happened with nbp? for awhile there was the "saar you can use to make ai influcers!"
i feel like now the only time i hear someone mention nbp is in this thread saying it's actualllllly still super sota!
>>
>>109433876
30B is easily trainable. I was ready for a 100B. Its called offloading.
>>
>>109433877
peak pseudbabble slop. actual competent people would just called it "test version 2"
>>
>>109433876
nah. This should train much faster than LTX did and on 16GB vram with offloading. LTX had to be taught from the ground up since it knew so little
>>
>>109433891
>Its called offloading.
by the time you finish your training we'll be on Flux 7 already kek
>>
>>109433904
1 block on device vs all blocks on device is a 5% slow down with DDR5.
https://files.catbox.moe/mp08fi.mp4
>>
>>109433889
>what even happened with nbp?
still the best model, I don't know why people glaze GPT Image 2 so much, it's slopped and the noise is so uncanny, it's so noisy how can people not notice that??
>>
>>109433904
at least we'll have krea chroma partial epoch 2 around then (512x512 hires training run)
>>
>>109433258
Happy for him he totally deserves :)
>>
>>109433910
this is for lora training btw. Its like a 20% slowdown for full finetuning. Still not much of a deal
>>
>>109433847
ayy lmao
>>
>>109433877
I lurk the chroma discord. He is truly doing bizarre nonsense again. His "RL" is training 2 models off of the base, one on good data and one on bad, and then taking the difference and applying that delta to base. He made some schizoposts about how that's contrastive learning and that's all RL really is (it's not).
>>
File: ight.png (314 KB, 480x356)
314 KB PNG
>>109433504
I guess I'm stuck with LTX and Wan for the rest of my life lol, looks like all the new video models will be giant one, oh well, I think it's a sign that tells me that this hobby isn't for me anymore.
>>
>>109433912
after looking up some recent nbp gens, i'm less than impressed.
it looks like shit compared to ideo and krea.
>>
>>109433932
>nbp looks like shit compared to krea
(You)
>>
>>109433927
"just train on the entirety of civitai and put it in the negative prompt" logic
>>
>>109433927
that works well. You can do RL training in either direction. Doing both is even better.
>>
>>109433927
Lmao, thanks for the info anon.
Another month, another doomed kekstone failbake.
>>
>>109433939
it's cool that you can use your phone to gen stuff, but that doesn't make it good.
>>
>>109433504
I miss the days when I was optimistic about the future of diffusion models, Z-image turbo showed us that you could make a small but powerful model, looks like they didn't get the memo, they catched the LLM virus, stacking moare and more layers is all you need!
>>
>>109433927
>>109433945
stop responding to your own posts
>>
File: 042108CUI_00001_.png (1.13 MB, 1216x832)
1.13 MB PNG
why do people say the russians fixed anima
>>
>>109433951
they managed to finetune it to produce good art instead of vectorized slop
>>
>>109433950
>>
File: 042346CUI_00001_.png (1.16 MB, 1216x832)
1.16 MB PNG
>>109433956
where did you see the outputs from their model
>>
>>109433504
they should do like Kimi K3, train their model on 4bit, so that they can go for big model but it won't be memory expensive for inference, still doing some bf16 in the year of our lord 2026 is sooo lazy
>>
Wait, is H3 a all in one model? qwenVL as a video + audio + TE? So its 32B all together?
>>
>>109433504
Do we know if it's distilled or not?
>>
>>109433968
no, H3 is video + audio, and then there's a separate TE that is qwen 3 VL 32b
>>
File: 1754794012959349.gif (97 KB, 498x408)
97 KB GIF
>>109433968
>>
>>109433968
No LMAO.
It's using 32b version of qwen3vl as text encoder.
So in total probably like 60 billion weights.
So like 120gb when all weights are in bf16.
>>
>>109433970
all will be revealed soon™.
looking forward to seeing what ostris says about how easy it is to train.
>>
>>109433644
shit like this makes me want to root for an alternative, but there's nothing unfortunately...
>>
comfy said 32GB of ram + 12GB of vram would be enough
>>
>>109433982
wonder if that is with int4 convrot or something though
>>
>>109433982
>comfy said
>>
>>109433982
10 minutes for 480p + 5 seconds kek
>>
File: 1784977756471788.png (35 KB, 1385x218)
35 KB PNG
Reminder for chuds itt.
>>
>>109433982
comfy also said that GGUFs are a meme so I don't tend to take this retard's opinions too seriously
>>
>>109433964
it was shared in a previous thread, it's in a telegram link
>>
>>109433988
worth it if it one shots 5 seconds of kino
>>
>>109433994
GGUFs ARE a meme. Just offload and int8 convrot is both faster and higher quality
>>
>>109433982
to scroll around a 9-node workflow at 10FPS maybe
>>
>>109433998
>Just offload
doesn't work, it OOM, and this retard killed multigpu (the only node that lets you manually offload) with his dynamic vram meme
>int8
GGUF has 6bits, 5bits, 4bits, and they're good quality, int8 doesn't solve anything, and you know that, stop being disingenuous for a second will ya?
>higher quality
Q8 is superior
>>
File: 042952CUI_00001_.png (830 KB, 1216x832)
830 KB PNG
kino transparency

>>109433995
I'll check it out, thanks.
>>
>>109433982
who gives a fuck about what Comfy says, he's literally paid to shill and implement those models, of course he's gonna find excuses
>>
>>109433982
sounds great, did he elaborate for what? surely it's not so good that it's essentially enough for 25s at 2k?
>>
>>109433998
ggufs are small, thats a big boon when you are minmaxing a new workflow and trying to do other shit on the side.
>>
>>109434004
it does not OOM for me meaning its something on your end / a custom node of yours fucking it
>>
how many of you faggots are in lodestone's furry discord? everything that happens there just happens to appear in this thread minutes later like
>>109433749
>>109433786
>>
>>109434013
>it does not OOM for me
good for you, but your experience isn't anyone's experience
>>
>>109434016
we keep our finger on the pulse of the local AI community
>>
>>109434018
its anyone who does not have a bad custom node it seems
>>
>>109434004
>int8 doesn't solve anything
i don't agree with the other anon that gguf are useless but int8 convrot is actually quite obviously quite effective for quantization too.
>>
>>109433504
more info here
https://github.com/huggingface/diffusers/pull/14355
>>
H3 is 26B
>>
File: 044000CUI_00001_.png (978 KB, 1216x832)
978 KB PNG
>>
File: 1773910804714858.png (179 KB, 1929x808)
179 KB PNG
>>109434030
it's not an unified model, you have one transformer for the t2v and i2v process, and another transformer for the reference process, that's really disappointing, I expected a unique model that does it all
>>
>>109433982
Comfy hyped the fuck of SD3 and look where it is right now
>>
it's so funny seeing so much FUD about a model that isn't even out yet.

If the model is good people will use it, if it's bad they wont. there's nothing else to it.
>>
>>109434034
no
>>109434030
>One packed token sequence carries text, conditioning, audio and video rows through a shared 33B transformer
>>
File: SD3.jpg (156 KB, 640x640)
156 KB JPG
>>109434042
Levels of kino seen since
>>
>>
>>109434046
>If the model is good people will use it, if it's bad they wont.
Flux 2 dev (32b) and HunyuanImage 3.0 Instruct (80b) were good models, and they are dead, try to guess why
>>
>>109434048
*not seen since
>>
>>109434042
iirc there were also <added safety> reasons why it sucked MORE at release than a week before

either way this doesn't mean hardware specs are completely wrong?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.