[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109864856

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109867541
First for BWC.
>>
>inb4 ffaze
>>
Next bread replace klein in OP with nuqwen?
>>
File: welovebwc.jpg (484 KB, 1440x1440)
484 KB JPG
>>109867558
>>
>>
File: riceburner.jpg (463 KB, 1440x1440)
463 KB JPG
>>
It feels like we are back in the SD 1.5 days with these schizo negative prompts just for the model to function.
>>
Why do retards think more wrinkles and dirty pores equals higher quality? I swear it's impossible to find genuine advice from people that know what they're talking about.
>>
>>109867626
You really need to stop misusing the term schizo it makes you look fucking retarded. This is how 90% of models functioned up until turbo models
>>
>masterpiece
>trending on artstation
those were the times
>>
File: horsecocksucker.png (985 KB, 689x777)
985 KB PNG
>Do high res pass with qwen
>generation time is 7 minutes even on top end hardware
It's....doesn't even take this long genning at 4mp on a single pass, it better be worth it
>>
File: welove.jpg (435 KB, 1632x1632)
435 KB JPG
>>
File: Qwen_image_2.1_00073.png (2.81 MB, 1248x1664)
2.81 MB PNG
See you in 7 minutes hope it's worth it vs just natively doing this 4mp which the model can handle
>>
>>109867688
correction 11 minutes, this is obscene I need to use a sage attention node
>>
File: condom-factory.jpg (523 KB, 1632x1632)
523 KB JPG
>>
>you need to
make me , faggot.
>>
>>109867719
Sorry that you keep coming off as retarded by saying that.
>>
>>109867633
monkey see "airbrushed plastic skin bad! aislop!"
monkey do "dirty skin, pores, and wrinkles! realism! no longer aislop!"
>>
and you need to get a job
>>
File: you need to.jpg (523 KB, 1408x1856)
523 KB JPG
>>109867719
Autism is an unfortunate disease of the mind, and its sufferers a subject to an ever-increasing alienation from regular social interaction.
>>
>>109867739
ok (I wont)
>>
I'm not the experienced with image gen, I only check out new models when they are out, see that they are not good enough for me and then wait for new one.
It does seem like stable-diffusion.cpp is popular now, is it better that ComfyUI? Maybe faster?
>>
File: lilrat.jpg (667 KB, 1408x1856)
667 KB JPG
Listen here, Jack...
>>
File: the rat of the fountain.jpg (857 KB, 1408x1856)
857 KB JPG
>>
>>109867688
>>109867692
It's broken seems to be some regression or something with this family of models
>>
File: QwenT2I_0058.png (3.65 MB, 1408x1856)
3.65 MB PNG
>>
>>109867568
nah
>>
File: Qwen_image_2.1_00191.png (2.18 MB, 1024x1504)
2.18 MB PNG
>>
>>109867568
why wouldn't you just re-add the separate Qwen column that used to exist but was removed for no particular reason? The current OP has a like a gorillion fewer models than it used to as it is
>>
>>109867734
the real cancer is the absolute faggots who insist that "muh smartphone photo that looks far lower quality than a photo any actual smartphone released in the past 15 years would ever take under any circumstances" is somehow "realism"
>>
File: migu.png (1.51 MB, 1024x1024)
1.51 MB PNG
>>
you kinda have to use image resize to 1024 for edit otherwise large images make it take forever
>>
>>109867936
nm, can just use custom size toggle
>>
>>109867936
> non-turbo models are slow as shit
wow thanks for that
>>
>>109867894
>just re-add the separate Qwen column that used to exist
Would any of the previous qwen models be worth readding or should it just be the new one?
>>
File: 1774551604745750.png (1.73 MB, 1024x1024)
1.73 MB PNG
the character in <image1> sitting at the table holding a knife and fork, is wearing the hat and beige business suit of the character in <image2>.
>>
>>109867964
I mean the entire idea of it being some competition and that the OP should be shrunk to like less than five models total was done by some faggot troll OP in the first place IIRC, but presumably you'd just have Qwen 2.1 probably if you were gonna do a new Qwen row
>>
HOW DOES ONE UPSCALE ALL MICRO DETAILS WHEN GENERATING ULTRA-4K-QUANTUM-QUAD-WIDTH DESKTOP WALLPAPERS??????
>>
>>109867975
I don't care as long as the maintain thread quality links stay
>>
>>109867986
those are the only things that have always been there, everything else got fucked up within the last few months. Whoever the non-retarded consistent baker used to be got replaced with some dramanigger at some point fairly recently
>>
>>109867975
We need an anon to make an automated OP already so the deranged schizo is barred from baking altogether
>>
>>109867991
I think all models should be shown, I think what often happens the split baker fucks things up and the bakers that fix it often focus on the rentry links thinking that's the only thing they vandalized. Let's fix it for next thread.
>>
no one even used the old qwen models tbdesu
>>
>>109867991
It happened when the animanon schizo rentry was added and it's been progressively fucked ever since
>>
>>109868001
>all models
Really, the only missing model is the new Qwen.
>>
>>109868021
Listen bros figure it out
>>
i love when changes to the OP are brought up and the trolls always move straight to removing the rentrys kek cry more
>>
qwen 1.2 > krea
>>
>>109867889
oh my booba
>>
>>109868023
Well, how do we get rid of the schizo? Once you figure that out, then the problems with the OP disappear
>>
>>109868025
They will seethe about it until they get sent to a care facility.
>>
>>109868025
>>109868042
If that's all you care about this general isn't for you
>>
>>109868046
We were discussing adding to the OP as needed, what the fuck are you babbling about?
How can you seethe here all day get banned constantly and think your opinion means anything
>>
>>109868051
You are ban evading right now, schizojack
>>
>>109867975
>I mean the entire idea of it being some competition and that the OP should be shrunk to like less than five models total
Because there are only so many base/foundational models, retard. Not every jeetmix deserves a place in OP.
>>
>>109868061
Just make a master list rentry retard. Then it's one link
>>
File: blaargh.jpg (811 KB, 1536x2048)
811 KB JPG
post gens or be quiet, ty!
>>
why this model has 3 text_encoders?
https://huggingface.co/Comfy-Org/Qwen-Image-2.1/tree/main/text_encoders
>>
>>109868075
because you touch yourself at night
>>
File: MiniMax_H3__00749.mp4 (1.5 MB, 576x896)
1.5 MB
1.5 MB MP4
>>
>>109868075
I assume it's for the vram poor
>>
>>109868084
Too early

But kinda cool
>>
>>109868028
1.2 you say
>>
>>109868075
it don't, you pick one
>>
so is klein edit 9b still a better edit model?
>>
File: Qwen_image_2.1_00077.jpg (874 KB, 2176x2896)
874 KB JPG
It's crazy how this guy makes shit up in his head, gets exposed constantly and still thinks he's welcome here. I never seen someone seethe so hard and lack the ability to understand why he's not wanted here or in the other thread.
>>
>>109868007
people kinda did for a good while. But like at one time we had shit like NetaYume in there and nobody cared, there's definitely been a significant increase in people with random agendas pruning the OP in arbitrary ways over time
>>
>>109868101
I think so personally especially given no turbo option exists for Qwen 2.1 ATM, so it's a significant speed difference vs 9B Distilled
>>
everyanon wants to dig their grubby little fingers into the golden LDG OP pot...
>>
>>109868110
convrot model isnt too slow but the klein edit 9b convrot model is super fast.
>>
>>109867254
this literally looks like 3 pass DLSS 5 kek
it outputs the same faces when skin intensity isn't dragged down
>>
File: MiniMax_H3__00750.mp4 (3.48 MB, 576x896)
3.48 MB
3.48 MB MP4
>>
~~ Summoning Miku Floidchad to test out Qwen 2.1 ~~
>>
>>109868084
I-I'm scared...
>>
>>109868081
i do that but how do they know
>>
>>109868157
Anon...I did a DLSS5 6-pass on this...

I'm sorry!
>>
>>109868111
Better than letting a schizo shit in it for years
>>
>>109868165
Cool!
>>
I've been messing around with adding five images as style references. It works pretty decently well. It's like a miniature style lora.
>>
>>109868217
Same seed and prompt except with references removed.
>>
Great, we now have a regression to about a year ago. Alibaba should be ashamed of themselves.
>>
File: hijabber.jpg (1.19 MB, 2176x2880)
1.19 MB JPG
>>109868103
>>
File: the nazi elf.jpg (286 KB, 1024x1024)
286 KB JPG
>>109867971
>>
>>109868274
nice gpt face
>>
File: bikiner.jpg (566 KB, 1248x1664)
566 KB JPG
>>109867688
>>
File: Qwen_image_2.1_00125.png (2.07 MB, 928x1376)
2.07 MB PNG
>>
>>109867652
>>
>>109868292
if you can, try running this with 1.5MP, euleur + linear quadradic
>>
>>109868246
>>109868288
>>109868301
Kek this edit model is amazing
>>
>>109868310
I'm having my fun with it, I hope other anons will too!

Even if the results are sometimes sub-par GPT-facey, it still follows instructions amazingly well, and FAST!
>>
>>109868310
What is kek worthy about it?
>>
File: MiniMax_H3__00752.mp4 (2.16 MB, 736x736)
2.16 MB
2.16 MB MP4
>>
File: QwenImageEdit2dot1_0113.png (3.17 MB, 1152x1728)
3.17 MB PNG
>>109868330
Where did you get this video of me prompting for Toph gens?
>>
File: Qwen_image_2.1_00088.jpg (596 KB, 1536x2720)
596 KB JPG
I think this model is one finetune away from being top tier imo, it's not as pretty as krea but it's understanding of elements are just as good if not better in some cases
>>109868317
Exactly, the model is worth it's weight in gold for the edit ability, I fucked up the aspect ratio of this image and it still hit the right notes, going to try the new text encoders, I don't know why comfy doesn't have the fp 16 ones
>>
>>109868317
>and FAST!
only if you don't use negative
>>
>>109868338
I can never tell with your images because they are all shit and usually have something to do with your mentally ill obsession. So far, most images look like overly detailed slop or gpt slop
>>
File: future.jpg (232 KB, 800x450)
232 KB JPG
The world if minimax knew what genitals looked like
>>
File: the wizpep.jpg (729 KB, 1152x1728)
729 KB JPG
>>109868358
Fair point, I usually avoid putting negs, if I need to put an art style or ethnicity, it'll go into the positives
>>
Hit dog will holler and all of that, every time
>>
>>109868374
Show the class a good gen, I bet you won't
>>
File: fuhrerwizard.jpg (619 KB, 1152x1728)
619 KB JPG
>>109868392
NTA, but I hope this will satiate your hunger <3
>>
>>109868382
How hard did ani hit catjack to holler for years?
>>
>>109868404
I want him to show something since he cries all day.
>>
File: Qwen_image_2.1_00193.png (2.02 MB, 928x1376)
2.02 MB PNG
>>109868309
what am I looking at?
>>
>>109868404
Looks like a bad shoop. Based but the quality is pretty bad anon. Normally happens with edit models anyways
>>
>>109868410
Anon is letting you know subtly that you are a subhuman mutt, catjack
>>
File: animegril.jpg (3.51 MB, 4032x1728)
3.51 MB JPG
Qwen 2.1 can't even remove the watermark from this Anima gen without fucking around with random shit. Look at her eye: in the original there's a smaller white line coming out of the bottom left-ish area of her kinda tilted plus sign shaped pupil. Klein maintains that as it should, Qwen tries (and only partally succeeds) to erase it, for no reason, and also slightly shifts the entire image left. Why the fuck would I use a full CFG non-distilled model that literally produces worse results than a 4-step distilled one on the same prompt?

I fed the original 1344x1728 Anima gen to both Klein and Qwen directly with no intermediate resizing of any kind to be clear. Prompt for both was just: `Remove the text that reads "@pellas panix233" in the bottom-left corner of anime image 1.`
>>
>>109868421
You're still afraid to post or do anything other than seethe "anon". Does it hurt knowing that you have no power here and never had it?
>>
File: QwenImageEdit2dot1_0124.png (3.75 MB, 1792x1664)
3.75 MB PNG
>>109868412
I was just curious, thank you anon. I (think) that there's better micro-detail, but again, I could be wrong.

Have a downvote hitlerjak as thanks <3
>>
>>109868426
You need to read the guide and properly tag each object
>>
File: QwenImageEdit2dot1_0125.png (2.34 MB, 1088x1344)
2.34 MB PNG
>>
>>109868436
gr8 b8 m8
>>
>>109868374
Agreed, I'm fine with "looks like gpt slop" on the txt2img side (or overly transformative img-edit request) though if the edit aspect works well. Klein9b I don't think I bothered with txt2img even a single time after assessing it to be worse than some combo of illu/zit/even fucking pony loras for NSFW or else API models if zit wasn't cutting it for sfw realism, but i used it as much as all those others combined for editing images where its base incompetence didn't really matter. The edit workflow stuff I've seen (replace this outfit, add this item etc rather than make this a photo, which appears to work fine but with not so good aesthetics) seems great. Does it do genitals okay with a reference, if so how good relative to h3 (which usually does inflating balloons or accordion dicks but occasionally gets it perfect)?
>>
File: why crying.jpg (1.13 MB, 1920x2016)
1.13 MB JPG
this got a laff out of me
>>
File: Q2.1Edit.jpg (2.17 MB, 1920x2160)
2.17 MB JPG
>>109868426
Uhh... you could complete that operation in Photoshop in like two seconds, what are you doing? I don't usually use edit models, but it seems fairly competent in my testing right now. Not that slow either.
>>
>>109868497
He's looking for a reason to complain
>>
I tried to edit the previously edited gwen image and the image was fried
why
>>
>>109868451
Why does it have that gpt like wash over?
>>
>>109868432
pretty accurate depiction of reddit i must say
>>
File: MiniMax_H3__00754.mp4 (2.84 MB, 736x736)
2.84 MB
2.84 MB MP4
>>
File: QwenImageEdit2dot1_0162.png (703 KB, 704x1120)
703 KB PNG
>>109868518
no clue
>>
>>109868497
wat. the fact I COULD do it with no edit model doesn't change the fact that Q2.1 is worse and slower than Distilled Klein 9B
>>
>>109868501
> nooo you have to agree that the slower, visibly worse model is great just because it's new, or some other inexplicable reason
>>
>>
>>109868518
because the model was very very clearly trained on six gorillion GPT outputs, it's generic chinkslop
>>
File: comfyui2.png (26 KB, 457x409)
26 KB PNG
>comfyui
nice UI you got there lol
>>
>>109868559
>>
>>109867541
I can't wait for a turbo lora/model so I can make this work on my large krea 2 images.
Protip bump the cfg above 5 and steps above 50 to reduce chatgpt effects
>>
>>109868559
as opposed to amerifat slop
>>
File: Qwen_image_2.1_00094.jpg (907 KB, 1792x2368)
907 KB JPG
Forgot image fuck
>>
File: QwenImageEdit2dot1_0173.png (2.3 MB, 1536x1152)
2.3 MB PNG
>>
File: QwenT2I_0073.png (1.34 MB, 1024x1024)
1.34 MB PNG
>>
>>109868570
why is it 480x768
>>
>>109868075
Oh, I get it.
PE text encoder are not encoder. they are LLM to enhance prompts,
but it's censored
>>
>>109868575
more like opposed to WE DON'T DO THAT IN GERMANY slop
>>
>>109868613
reference image was this size
>>
File: QwenT2I_0080.png (1.42 MB, 1024x1024)
1.42 MB PNG
>>
File: MiniMax_H3__00756.mp4 (2.01 MB, 736x736)
2.01 MB
2.01 MB MP4
>>109868615
if you keep doing the racism your face will get stuck like this
>>
File: MiniMax_H3__00757.mp4 (2.65 MB, 736x736)
2.65 MB
2.65 MB MP4
>>
>>109868629
wow i'm glad we can FINALLY generate images like this, at 1024x1024.
>>
>>109868609
>>109868629
I choose to believe these are the same scene
>>
>>109868629
Backwards physics on the pant leg
>>
The edits stick too closely to the inputs so they don't mix very well style-wise
>>
>>109867541
Do you guys use diffusion models for edits of existing images? Say i wanted to change the shirt of a person on a photo, would diffusion models be the go to? Claude tells me the meta is using qwen image edit or others like it, but sometimes it doesn't adhere to the prompt and it has a plastic texture to skin



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.