[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


With Fries Edition

Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109886022

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://animastyles.thetacursed.com
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

09/23/2026

>Qwen-Image-2.1-viggle-turbo — v0.2 (preview)
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

>Ming-Image-0.1-Design: 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs
https://huggingface.co/inclusionAI/Ming-Image-0.1-Design

>MiniMax-H3-Fun-Controlnet-Union-2.0
https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0

>Streaming Video Editing with Easy Adaptation
https://github.com/YujiaHu1109/SVEET

>WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
https://drexubery.github.io/WorldCrafter

>GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
https://jiah-cloud.github.io/GAE.github.io

>Visual Jev: Accurate and Efficient Decisions from Shared Visual Context
https://github.com/guanxuyu-sv/Visual-Jev

>Comfy Router: One API for Frontier Media Models
https://blog.comfy.org/p/introducing-comfy-router-one-api

>qwen image studio: Command line and a local web studio for Qwen-Image-2.1
https://github.com/janishar/qwen-image-2.1-studio

>OpenAI: Priorities and principles for effective third party assessments
https://openai.com/index/priorities-principles-third-party-assessments

09/22/2026

>CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation
https://zshyang.github.io/CoaG

>AniPrO: Interpretable Anime Image Provenance Detection via Multi-Dimensional Semantic Reasoning
https://github.com/YAN-LIU05/AniPrO

>AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport
https://github.com/51xOne/Alignmorph

>Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
https://sunyangtian.github.io/Mira-Scene-web

>Rethinking Vision Architectures with Gated Linear Attention and KAN
https://github.com/mehizelali/linear-kan-transformer

>SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted VLMs
https://github.com/aliathar1401/SK-Stars-shroom-visions-2026
>>
File: Krea2_turbo_03964_.jpg (1.4 MB, 2368x1776)
1.4 MB JPG
>>109890959
>>
>mfw Research news

09/23/2026

>LoRango: It Takes Two LoRAs to Unlock Hidden Behaviors in Diffusion Models
https://arxiv.org/abs/2609.25884

>Code Plans, Diffusion Renders: Open-Ended Generative World Modeling
https://becauseimbatman0.github.io/CoDeR

>QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation
https://arxiv.org/abs/2609.26425

>TRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection
https://arxiv.org/abs/2609.25775

>Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models
https://arxiv.org/abs/2609.25444

>StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
https://tt-day.github.io/StableVQ

>PixelDiT2: Representation-Grounded Pixel Diffusion Transformers
https://arxiv.org/abs/2609.24919

>PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning
https://arxiv.org/abs/2609.24768

>VideoGen-Agent: Reinforcing Video Generation Agents
https://arxiv.org/abs/2609.24997

>EMERGE: Resolution-Agnostic Point Cloud Generation with Equivariant Graph-Based Diffusion
https://arxiv.org/abs/2609.26039

>RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models
https://arxiv.org/abs/2609.25492

>RULER: Instance-aware Rubric Rewards for SVG Generation
https://hangyuran.github.io/RULER

>KwaiMind Technical Report
https://arxiv.org/abs/2609.26375

>ImIR: Image-Instruction Tuning for All-in-One Image Restoration
https://arxiv.org/abs/2609.25267

>SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models
https://arxiv.org/abs/2609.24875

>Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs
https://arxiv.org/abs/2609.25635

>What Drives Hierarchy-Aware Image Retrieval? Taxonomy Alignment, Objective Choice, and Geometry
https://arxiv.org/abs/2609.25638

>Qwen3.8-Omni: Towards Native Omni-Modal Agents
https://arxiv.org/abs/2609.25611
>>
>This gen >>109888662 not in the collage
Yep.
>>
>>109890983
It was included in a previous faggollage the first time it was posted
>>
With Minimax H3 R2V, is it basically essential to add a dedicated face reference for subject reference to prevent face drift? I'm running anime gens but I'm definitely noticing a tendency for unique faces to become more like generic anime faces from the mid-2000s.

I'm guessing a first frame anchor with the face visible would be even better.
>>
>>109890959
>>109890966
thanks bro
>>
File: debo_ms_k2_00047_.png (2.74 MB, 1664x1069)
2.74 MB PNG
>>109891048
I got u senpai
>>
File: Qwen_image_2.1_00140.png (1.37 MB, 1056x992)
1.37 MB PNG
>>109891048
>>109891057
>it thinks it belongs here
>>
>>109890959
>>109890966
Your posts are probably the only reason I still check this schizo shithole.
>>
>>109890959
>>109890966
Fuck off unemployed loser
>>
>>109891057
it's like you can't decide between first person and over the shoulder and you're too much of a promptlet to get either right so shit is all fucked all of the time
>>
>>109891096
woman, come on, this is a GUY place. We like computers here!
>>
It's up!
https://huggingface.co/black-forest-labs/flux-3-action-base
>>
>>109891114
Zero hype
>>
:(

I don't want a robotics model.

>astra
>want a ROBOT

NO!

:(((
>>
>>
>>109891182
Nice
>>
>>109891114
but is there an action to make the user coom?
>>
>>109891182
damn he thicc af
>>
These sad attempts at trolling and derailing bore me. Dance better, monkey.
>>
File: file.png (3.77 MB, 1728x1152)
3.77 MB PNG
>>
>>
why is this guy posting 2022 era API slop?
>>
>>109891238
takai, takai!
>>
>>109891249
Why are you not posting gens instead of bitching?
>>
File: YogaLatina1080.png (1.31 MB, 1296x1080)
1.31 MB PNG
I AM THE COOM MACHINE NOW
>>
>>109891256
No need to cry anon he's just asking a question
>>
This is why we have two rentry links btw
He gets very upset when called disabled, I wish people would stop triggering him.
>>
>>109891249
why do you care someone is posting images? not your place to hall monitor. post images and get on topic. pic related is you when someone does things you dislike and pretend you have a say in it. again get on topic.
>>
>>109891256
because this bread is based and doesnt require anon to attach an identifiable gen to be taken seriously like some kinda namefag hugbox
>>
What is the current optimized, minimal quality loss meta for Minimax H3? I'm currently using sageattention, spectrum and 22 steps.
Should I be looking at shit like SolAttention, SparseAttention? Whatever else?
I prefer not to use turbo because the quality loss is too noticeable, especially with any kind of motion like hand movement.
>>
>>109890959
>>109890966
thanks!
>>
>>109891299
2022 era cloud gens are not on topic in the Local Diffusion General sperg
>>
>>109891317
prove they are cloud gens?
>>
109891317
>reads the idiots babble and opinions
stop trying to hall monitor others. get on topic or fuck off back to >>>/b/.
you're ass annoying as that idiot from /sdg/ showing up and larping he's in charge here. piss off wacko jacko
>>
>continues crying and posting off topic images
They're not sending their best are they
>>
>>109891329
he can't prove it. he's doing what some of the resident wacko jacko's do and fabricate things and run with it as fact and than get melty upset when their comical delusions are called out or mocked. just ignore him he's one of the residents who sniffed his farts too much.
>>
Didn't you just get banned for doing this yesterday?
Funny how I was right about you.
>>
>>
How the fuck is HiDream O1 near the top of literally any benchmark chart?
>>
>>109891344
I like that instead of posting a catbox to prove it's a local gen you choose to sperg out like you usually do. You're just trying to derail this thread correct? Just answer yes or no please.
>>
>>109891307
If you care about quality, I wouldn't touch any of the stuff you mentioned, and maybe swap sage attention out for CK attention.

I don't mind turbo loras at all. They're pretty great for shitting out some quick gens to post. Even better if you don't need audio anyway.
>>
>>109891372
"I like that instead of posting a catbox to prove it's a local gen you choose to sperg out like you usually do."
implying I prove anything to idiots and people acting like hall monitors. you prove you're authority and own the site or piss off and stop derailing the thread with your obsessions you're god and dictate to others.
"you choose to sperg out like you usually do."
top tier story telling anon. I love how you threw in choose to really amp up your story time you made up.
"You're just trying to derail this thread correct?"
No and your lies and insistence your lies are valid are mocked openly.
this is all you get as I do not answer to any of the resident posters nor are any of you authority. be glad I even graced you with this reply and dont just keep mocking you and your lies with the others.
>>
>>109891329
you werent around way back when? pretty obviously cloud slop or some lora trained on old gemini or whatever desu. but you can tell hes upset at being called out based on his replies >>109891344 and >>109891393 so obviously hes just trying to troll for some reason
>>
>>109891393
I rate this shitpost a solid 3/10.
>>
Why do the anons who post non local gens here always get so upset? The guy here right now and then the other guy yesterday who got wiped by mods. IDGI
>>
"shitpost"
"he's trying to troll"
"hes upset at being called out"
lol, lmao stop derailing in your melty posts
>>
>>109891369
Arena benchmarks. A surprising amount of people seem to like the slop look.
>>
comfyniggers really ruined this place for good man
>>
File: 190994.png (2.46 MB, 1398x1824)
2.46 MB PNG
>>
>>109891410
I think he's severely autistic and doesn't understand how his posts come across. He typically has meltdowns that last for like half the thread.
>>
File: 1761222939238812.jpg (310 KB, 1024x1024)
310 KB JPG
>>
How does Qwen respond to LoRA training?
>>
>>109891444
He's taking the piss, nobody is subhuman enough to reply and use "(you)" this way to reply.
It's just extremely low effort shitposting, might be an automated LLM, now with omarchy the reddit OS being mainstream even the most retarded person can use hermes as automated agent to shitpost
>>
Blessed thread of frenship
>>
>>109891114
brb getting a degree in engineering
>>
Debo and his ilk can't troll they just come off as disabled which is why /sdg/ is dead
>>
>>109891114
Based on the shitty example video where they're scared to show a single output alone I can tell this model is ass
>>
File: 1759258963053809.jpg (228 KB, 1024x1024)
228 KB JPG
>>
>>109891469
I hope it's automated or else I'd feel bad that someone spends their life like that. He's been popping in and out with his sperging for months at this point.
>>
GET A ROPE
>>
>>109891547
Are you new to this general?
This is done manually mostly by the rentry schizos
>>
File: 1762565112524378.jpg (274 KB, 1024x1024)
274 KB JPG
>>
anon can tell you changed the filenames to random while continuing to post cloud slop
>>
>>109891612
>continuing to post cloud slop
proof?
>>
>>109891536
It's not for generating normal videos anyways
>>
File: QwenImageEdit2dot1_0283.png (1.45 MB, 1056x992)
1.45 MB PNG
>>109891363
I'm European and I look like this
>>
File: QwenImageEdit2dot1_0284.png (1.08 MB, 1024x1024)
1.08 MB PNG
>>109891452
>>
>>109891587
qwen? looks like dall-e
>>
File: debo_ms_k2_00050_.png (2.57 MB, 1664x1069)
2.57 MB PNG
>>109891793
I think it missed the assignment on this one
>>
File: QwenImageEdit2dot1_0286.png (2.72 MB, 1664x1056)
2.72 MB PNG
>>109891834
I tried many times to get the suit to disappear, but I'm too much of a brainlet to prompt this correctly, sadly.

Have another gemmie instead :)
>>
>>109891847
Crazy how you made that shitty image make sense within seconds.
Generational talent desu
>>
File: QwenImageEdit2dot1_0287.png (1.43 MB, 1024x1024)
1.43 MB PNG
>>109891537
>>
File: debo_ms_k2_00051_.png (2.74 MB, 1664x1069)
2.74 MB PNG
>>109891847
based
>>
>>109891866
It just...makes sense.
>>
File: QwenImageEdit2dot1_0289.png (2.98 MB, 1664x1056)
2.98 MB PNG
>>109891872
Target spotted.
>>
File: QwenImageEdit2dot1_0290.png (1.91 MB, 1024x1024)
1.91 MB PNG
>>109891299
RAAAAAAAAAAAAAAAH!
>>
caveman frog is my favorite LDG meme ^_^
>>
https://huggingface.co/collections/black-forest-labs/flux-3-action

flux 3 released not yet
>>
File: QwenT2I_0111.png (1.39 MB, 896x1152)
1.39 MB PNG
>>109891975
Unreleased prototype caveman
>>
>Debo on his knees in the main thread.
I guess it sucks to only samefag in your containment
>>
>>109891933
saved. thanks anon
>>
File: 458780467937579846984.jpg (3.84 MB, 3360x5040)
3.84 MB JPG
>forget to bypass prompt i was working on
>queue up a bunch of misc fantasy slop
>end up with dozens of victorian era tranny orcs, trolls, goblins, etc.
>mfw
>>
File: QwenImageEdit2dot1_0293.jpg (1.15 MB, 1664x2528)
1.15 MB JPG
>>109892013
>>
File: debo_ms_k2_00052_.png (2.96 MB, 1664x1069)
2.96 MB PNG
>>109891975
hard agree
although competition is light because people dont post gens here xD
>>
File: QwenImageEdit2dot1_0294.png (2.81 MB, 1664x1056)
2.81 MB PNG
>>109892115
>>
>>109892115
>>109891975
Glad you guys like them, I'll churn them out while I study Chinese
>>
>>109891975
>>109892115
>>109892137
Why continue to post here if you don't like this thread?
>>
>>109892155
He's disabled and only lives to grief us
You know what to do in order to fix this
>>
>>109892155
He never did mention that he dislikes the thread, only that people rarely post gens, anon.
>>
>>109892183
He actually hates this place if you were actually here when the thread was first made. He's a legitimately disabled and has a long history of faggotry that can be seen in the OP
>>
>>
File: QwenImageEdit2dot1_0295.png (3.72 MB, 1536x1600)
3.72 MB PNG
>>
>>
>>109892155
its the only imgvidgen thread on g that ever talks about the tech at all so of course trolls will come here to see whats new
>>
You're now being a avatarfag
>>
>>109892183
How long have you lurked in LDG for?
>>
>>109892115
>although competition is light
i pity you
>>
>>109892288
A couple of years, never posting, occasionally taking multi-month breaks, fren
>>
>>109892301
Really? So not very much would you say?
>>
>>109892288
Remember he likes to pretend to be a newfag which is how our disabled friend got a rentry. Also kind of odd the API fag that got banned did the same thing.
>>
File: QwenImageEdit2dot1_0298.png (3.21 MB, 1248x1664)
3.21 MB PNG
>>109892316
That depends on your definition of much
>>
>>109892338
You're officially a avatarfag
You know what to do boys
>>
File: file.png (3.52 MB, 1216x1664)
3.52 MB PNG
>set aspect ratio to portrait
>keep samplers and settings the same
>fine quality improves
what the hell, not enough landscape images in the dataset perhaps?
>>
>>109892338
That's why I asked you, fren. So what would you say then?
>>
>>109892372
Like I said, a couple of years, but please rest assured that I'm not one of the resident schizos - you're free to believe otherwise, but if you were able to pull my IP, it would show <3
>>
File: QwenImageEdit2dot1_0301.png (2.17 MB, 1216x1216)
2.17 MB PNG
>>
>>109892384
Ah okay, I understand now. So why do you continue to post here if you don't like this thread?
>>
>>109892316
>>109892372
>role playing he decides who is allowed based on some time constrainst he larps is valid
>>109892331
>still obsessing over things he makes up

not seeing any images posted just a lot of idiots trying to play "detective" and role playing "hall monitor".
>>
>>109892407
Feel free to keep prodding, but I actually never made a post where I said I dislike this thread, fren

I enjoy posting gens, and so should you <3
>>
>>109892410
You just got banned for spamming yesterday and you're ban evading just to cry?
You /sdg/ faggots are pathetic
>>
>>109892417
nta but the garbage dump is this way >>>/g/sdg
>>
File: 1788114563537105.jpg (1.54 MB, 1776x2368)
1.54 MB JPG
>>109892425
oh ho ho such filthy lies from you.
I never got banned so you lying I did and making up I am ban evading is being called out. You can keep lying and making up bullshit but it wont go how you think. so stop confusing me with someone else you dog dumb idiot. I dont post in /sdg/ as the looping crazy in that thread is tedious. I havent posted there in a good 5 months, so you need to cease this lie too. get on topic kid
>>
You know how to remove the spam, tear it all down and send a message
>>109892465
Why are you reposting gens from other anons?
>>
File: 1788120685572289.jpg (1.49 MB, 2368x1776)
1.49 MB JPG
>>109892477
funny story anon but people have been saving images from the internet for decades. its a thing everyone does including you. so you trying to play some card that posting images is not valid and saving images is not valid is called out and exposed. you are not going to win by constantly trying to change "rules" you made up that arent valid as decades of saved images and files from people exposes you're grasping at nothing valid thinking you can get some "win". get on topic kid.
>>
File: QwenT2I_0112.png (1.48 MB, 896x1152)
1.48 MB PNG
>>
How can one anon cause so much seething?
We get spammed daily by these fags and it always comes back to one anon that they hate.
>>
>>109892517
obsession is a hell of a drug
>>
File: thegrab.jpg (325 KB, 1248x1248)
325 KB JPG
>>
File: tsunade-party.jpg (236 KB, 832x1248)
236 KB JPG
>>
>>109892527
They spend everyday trying to get him but they keep failing at it and only come out worse.
>>
File: selfie-time.jpg (409 KB, 1056x1584)
409 KB JPG
>>
It's strange how he cycles through annoying personas once he gets called out. Can you imagine being this pathetic?
>>
qwn 2.1 is so slow, man
any lossless speed up?
>>
File: the running creature.jpg (324 KB, 880x1184)
324 KB JPG
>>109892622
What personas would those be?
Not that guy, just curious which ones you're referring to.
>>
File: QwenT2I_0114.png (1.62 MB, 896x1152)
1.62 MB PNG
>>
>>109892368
Could also be your prompt idea simply works better in one aspect ratio vs another.

>>109892622
I mean earlier he admitted to assuming all anons are retarded so obviously he doesn't think he needs to try very hard. You can tell from his reply to your post as well.
>>
>>109892642
a better computer :(
>>
>>109892465
>I dont post in /sdg/ as the looping crazy in that thread is tedious.
Are you trying to say you never had those long drawn out discussions with debo in sdg about your "chatbots"? You really underestimate the intelligence and memory of the average anon.
>>
File: QwenT2I_0115.png (1.52 MB, 896x1152)
1.52 MB PNG
>>
>>109892642
? qwen is quicker than krea
>>
Already tuckered out? Bored of spamming low effort slop? Shame...
>>
>>109892795
Try to use >cfg 3.0 and a minimum 30 steps and report back
The default workflow is wrong
>>
>>109892690
>>109892713
Just another anon chiming in. It's all so repetitive these people really has some sort of AI psychosis and they take it out on us for some reason. They are not intelligent and they assume everyone is as dumb as they are.
>>
is krea 2 identity edit better than qwen 2.1?
>>
>>109892825
I think qwen 2.1 is better because it's not censored and will write whatever you want it to with text
>>
>>109892809
i do use cfg and high steps. i must be using a lower res and not realizing it then. if anon posted a direct comparison of speed between the two then i missed it
>>
>>109892809
CFG 1 is the official recommendation, IDK why they did it as like a sorta-distilled model that can do both CFG 1 and higher though
>>
If only qwen didn't use a subpar vae!!!!! How has vaeless not become mainstream yet?
>>
>>109892831
It's objectively wrong I use between 6-10 50+ steps
>>
>>109892831
>IDK why they did it
retardation who knows
>>
>>109892830
What is your speed for 1M edits?
I only got 1.5it/s. For the fact that it requires >30 steps to work properly, it's kinda slow
>>
>>109892817
>they take it out on us for some reason.
They take it out on /ldg/ because they feel it "stole" the discussion from their home thread. Which is true but also pretty funny.
>>
>>109892555
RIP Lindsay Graham
>>
>>109892858
i havent tried editing yet. any quirks different than t2i other than changing the default sampling params or is it pretty much plug and play for editing?
>>
>>109892925
Yes, it is, except it uses the original picture's resolution as output as well
>>
Cozy half hour
>>
>>109892858
>I only got 1.5it/s
really lmao?

I run models at 20 s / it minimum. Usually <1m / it, and then it just depends how many steps. Because my rationale is, if it's not annoying a little bit, I'm not using my resources.
>>
animooo
>>
i wish the local llms were just as good as gemini 3.7 and 3.8 at captioning images.
>>
>>109892947
>I run models at 20 s / it minimum. Usually <1m / it, and then it just depends how many steps. Because my rationale is, if it's not annoying a little bit, I'm not using my resources.
I kneel to you, patience chad. You're crazy.
>>
>>109892966
>gemini
is it really that much better than for example claude? despite all the bullshit with anthropic it seems to me thier newest model is great at captioning.
>>
>>109892860
They can't even deal with discussion amongst themselves. Fuck them, they are worthless dregs that hate their own company.
>>
>>109892966
the 1 gorillion parameters are, but you cant run them can you?
>>
>>109892955
cute :3
>>
>>
>>109892966
A good local harness helps. This is often neglected.
>>
why cumrag hasn't released the fixed qwen vae yet?
>>
>>109892984
3.8 is "not ai" - in how it feels. Less "ai" by a LOT than Gemma 4.

It's just ... messed up. It has some deep knowledge, but it's hit or miss, and based on another's hint, you don't want to keep talking with it, make a new prompt, it can't shift the tone.
>>
>>109893045
>the fixed qwen vae yet?
What was fixed? Link?
>>
>>109892984
never used claude, i would imagine they're ultra safety cucks. Do the anthropic models even admit "nude, naked, nudity" captions?
>>109893012
i don't like the hallucinations rates with gemma4 and the web search functions are ass. Its a night and day difference unfortunately.
>>
>>109893069
>gemma4
>big model
lol
>>
why does anon constantly make up stuff about comfyui to make it seem evil or bad
>>
>>109893099
Comfy the Dev wouldn't blow out his back.
>>
>>109893037
i use this shit. It was recommended by /lmg/.
https://github.com/lostruins/koboldcpp
if you have a better front end harness than this that allows vision multimodal support for images, audio and videos. please share the github repository or link.
>>
File: cZWDH7L1xdvFX0-SdzFea.jpg (430 KB, 3812x917)
430 KB JPG
>>109893056
Qwen VAE has a built-in grid pattern.
You see it when you zoom in
I thought someone would train a better LoRA at this point
>>
>>109893115
>x zoom

what are you doing
>>
do one is talking about flux 3 action which can make video game bots?
>>
>>109893113
>koboldcpp
I prefer ollama, but a harness is more than what backend you choose. It's incredibly easy to vibe your own frontend with whatever tools and skills you wish. I highly recommend doing that.

>>109893115
>I thought someone would train a better LoRA at this point
It's not that simple. You essentially have to retrain the entire model. I highly doubt anyone will sink money into doing that for Qwen.
>>
>>109893115
they both look like they have the same artifacts?
>>
>>109893126
You want me to show a 48x48px picture to get my point across, huh retard?
>>
>>109893045
Why did you post this like comfy alluded to having a better vae? Don't get my hopes up like that.
>>
>>109893174
times isn't the best unit.
>>
>low effort spammers who claim to enjoy this thread and claim to not post off topic stop posting their gens
>on topic discussion of the tech suddenly swells
Loooole
>>
>>109893113
>>109893155
In other words, just do it yourself
>>
>>109893146
desu wet fart release boring model
>>
>>109893228
wet farts are exciting thoughie
>>
I just realized, why am I spending all this time trying to get models to generate the perfect image or video when I can just use my brain and imagination and generate something to masturbate to and there are zero content policy violations or bad gens and I don't have to wait 10 minutes to see it I can just see the butthole getting fucked right now in my head as I'm typing this with no increase to my electricity bill.
>>
>>109893270
only works with shape rotator brains
>>
>>109893270
Once you get to a certain point, the hallucinations of the model are the interesting parts that you want to see. But you must reach a very high skill level to understand this.
>>
>>109893281
I have no idea what that means
>>109893289
I'm sorry bro top of torso bent around to look around at the dude fucking her doggy style is not hot it looks demonic and evil
>>
File: ComfyUI_00134.jpg (2.2 MB, 3552x4736)
2.2 MB JPG
>>109893115
Can you point out the grid pattern in this image and prove it on the same image that has not been touched by the model?
>>
>>109893300
I did say you have to be of a very high skill level to understand it so...
>>
>>109893300
>he doesnt get it
low IQ
>>
>>109893045
>>109893115
The real issue is that it's not as "high quality" as other VAEs. It's not completely unsably shit but for the trained eye it is subpar.
>>
>>109893314
Ok now prove it on
>>109893337
>>
File: QwenT2I_0116.png (1.45 MB, 896x1152)
1.45 MB PNG
>>109893202
let's goooo jack!
>>
File: QwenImageEdit2dot1_0299.png (2.18 MB, 1216x1216)
2.18 MB PNG
>>
>admits to trolling
>goes back go spamming
Like clockwork
>>
File: KREA_01075.png (2.27 MB, 720x720)
2.27 MB PNG
>>
File: ComfyUI_0520.png (2.32 MB, 1152x1792)
2.32 MB PNG
>>
>>109893385
Mentally ill seething, just knowing it hurts them makes it worth it, still waiting for the anon that claims to have a trained eye actually prove his point instead of having a skill issue and ignoring anons that advised against using the shitty defaults
>>
good on you for changing your filename and resolution so anon doesnt think its the same person
>>
>>109893385
I just had some time again, can't stay on the board all day fren
>>
>>
File: QwenImageEdit2dot1_0246.png (3.25 MB, 1600x1248)
3.25 MB PNG
>>
>>109893431
personally, I'm quite fond of this one
>>
Debo was originally a pepe spammer before he went to his demon fag persona. I feel like he forgets that many of us predate him.
>>
File: discojoe.jpg (230 KB, 1024x1024)
230 KB JPG
>>
File: da-mei-mei.jpg (404 KB, 992x1488)
404 KB JPG
>>
>>109893407
So you realize you're spamming low effort slop in a thread that doesn't want it? Weird.
>>
File: attack on winey anons.jpg (1.27 MB, 3056x1712)
1.27 MB JPG
>>
>>109893314
I guess he can't back up his claims.....
>>
>spams this thread so often to try to drive anons out
>thread still has more tech discussion than the other place
>>
File: sick migu.jpg (266 KB, 832x1248)
266 KB JPG
>>
>>109893512
It's a pathetic group of losers. Who just seethe 24/7
>>
>>109893321
>>109893323
I might be retarded but miku is literally bouncing on my cock right now in my imagination while you are a virgin.
>>
>>109893447
was he the nigger pic spammer or it was someone else?
>>
>>109893526
um okay, mr. sex haver. since when did 4chan turn into normie central?
>>
>>109893531
He posted BBC in a sfw thread so yes that's him. It's odd how all of these schizos have an obsession with dick too
>>
>>
File: KREA_01376.png (3.33 MB, 1184x880)
3.33 MB PNG
>>
qwen 2.1 is pretty okay. just batch cleaned a ton of work.
>>
I'm seriously. Sdg is more for posting many images one after another without much text. Why is he not doing it over there?
>>
the based image poster vs the cringe no-gen seether
>>
>>109893638
They are competently ass mad that everyone split from them to make /ldg/ and that it's a dead schizo shit hole where a rentry exist to aware newfags of their leader's behavior.
>>
>patting yourself on the back
>>
File: dlss5 image.jpg (270 KB, 832x1248)
270 KB JPG
The advantages of running a DLSS5 pass on some images seem great, have you tried this yet?

https://github.com/criso2hd-alt/DLSS5-Image-Converter
>>
>>109893665
*seems*
I'm retarged
>>
>>109893661
*Completely
Most anons post in that thread only to tell debo to go back into his tard cage and it makes them all seethe
>>
>>109893665
does it work for non-rtx50xx cards?
>>
File: SUPERSHART.jpg (266 KB, 1024x1024)
266 KB JPG
>>
>>109893681
>does it work for non-rtx50xx cards?

It should, as long as you use an RTX card
>>
gunna have to do better than mentioning months old tech to convince anon you arent trolling kek
>>
>>109893681
>>109893699
I forgot to mention, you need to bring your own DLSS5 files, but I can catbox them if needed
>>
File: riceburner.jpg (463 KB, 1440x1440)
463 KB JPG
>>
>>109893736
are there linux versions?
>>
>>109893749
Not that I'm aware of anon, and if not, you may be able to vibecode it
>>
File: cleavage-dawg.jpg (566 KB, 2240x960)
566 KB JPG
>>
>>109893753
i don't know how to do that, and i have a feeling it would be a waste of money since it is a driver-level feature
>>
File: the-heist.jpg (861 KB, 2304x1792)
861 KB JPG
reposting this one because it's gold
>>
*yawn*
>>
File: nun-server.jpg (1.18 MB, 2624x1984)
1.18 MB JPG
>>
>>109893661
Yeah I just can't think of another reason why someone would spam post images with little text here when sdg is quite literally the thread that encourages it
>>
File: Screenshot 2026-09-24.jpg (67 KB, 882x486)
67 KB JPG
i tell glm 5.3 to rewrite the generate text node to use the official qwen image prompt enhancer models
, and it already outputs 2x faster token/s than the one they use in cumfart. It's still not as fast as llama.cpp though
>>
>>109893790
It's one severely autistic faggot pretending to be multiple anons and a small number of anons that don't interact there's a reason why the thread falls off to page 10 when the main schizo is not present
>>
File: study-session.jpg (456 KB, 1344x1344)
456 KB JPG
Let's talk, LDG.

What's your favourite model as of this week?

Mine is krea2 turbo for images and qwen 2.1 for edits :D
>>
>ai image thread (you can create anything you want)
>it's all reposts
MIMS
>>
>>109893838
what do you plan to do with that information
>>
>>109893854
Nothing, desu, just curious, I want to try other models if anons recommend
>>
>somehow manages to lower the quality of discussion even further
Impressive, low effort spammer
>>
Report him
>>
File: 4557845437894.png (18 KB, 900x806)
18 KB PNG
where is kino?
>>
https://github.com/awdqwdasdg/Comfyui-Spectrum-Qwen2.1
i run 50 steps with spectrum faster than base 25 step
maybe this model is usable after all
>>
I'm happy to announce that I have taken my seat in the kinoplexatorium
>>
File: Bielsko-Biała,_Kinoplex.jpg (3.06 MB, 2560x1920)
3.06 MB JPG
>>109893934
we are seated
>>
>>109893931
Bookmarked for now, will check out soon
>>
>>109893931
No phree lunch
>>
>>109893931
im spectriiiing
>>
>account created for the sole purpose of putting sneed into the kinoplex wikipedia article
it was one of you guys, wasn't it?
https://en.wikipedia.org/wiki/Special:Contributions/~2026-42926-92
>>
>>109893959
>turn sloppa into sloppu
they are all the same
>>
>>109893996
Not really
>>
>>109893923
idk
>>
File: this app is cooked.webm (951 KB, 832x1080)
951 KB
951 KB WEBM
https://files.catbox.moe/zwdoqj.mp4
>>
>>109888224
>Can make some good edits at a fraction of the time i.e 30 seconds at 150 steps on a 4mp image
What secret magic does your node have that makes it quicker? Pretty cool.
>>
>search civitai for qwen2.1 loras
>only two exist
lame
>one is an age slider
based
>>
>>109894116
It grabs a piece of the image instead of the full image and blends it with the rest of the image.
>>
File: 1762032132098599.png (3.19 MB, 1152x1728)
3.19 MB PNG
>>
File: 2657.png (3.43 MB, 1034x1102)
3.43 MB PNG
ok maybe qwen is ok
>>
>>109894103
JSID
>>
i can't believe my covfefe didn't make the collage, fuck this site
>>
>>109894145
this is how i sex
>>
File: QwenT2I_0118.png (1.86 MB, 896x1152)
1.86 MB PNG
>>109894157
>>
File: 1764420585152003.png (2.23 MB, 1408x1408)
2.23 MB PNG
>99% legible and correct text but deepfried
>dropping CFG does nothing
>>
>>109894219
what did you prompt for?
>>
File: QwenImageEdit2dot1_0306.png (1.96 MB, 1120x1600)
1.96 MB PNG
hate this one, hope you all will like it
>>
>>109894238
Exactly what you see in the image https://pastebin.com/WwCCb0Ri
>>
>>109894253
oh, cool. i was wondering if you tried running a screenshot through it
>>
File: QwenImageEdit2dot1_0308.png (2.16 MB, 992x1376)
2.16 MB PNG
>>
>>109894149
its bretty gud
>>
>>109894279
What does the onion represent?
>>
>want to make anime slop
>audio is absolutely fucked no matter what until i add samples
>no issue with this but im lazy
>try to gen without audio samples and let he models do their thing
>absolutely fucked nonsense babble the whole time
>even between explicitly defined dialog
>try several finetunes
>same problem with varying results
>trying to stay away from turbo models
am i just going to need to accept defeat on this one?
>>
File: QwenImageEdit2dot1_0310.png (3.06 MB, 1152x1728)
3.06 MB PNG
>>109894319
My many layers, and also how I cry when I fail a generation
>>
>>109894345
if you are using a prompt bloater, then that might be the problem
>>
File: Screenshot 2026-09-24.jpg (45 KB, 393x506)
45 KB JPG
>>109893808
wait. I'm actually retarded
Just tell it to make a llama.cpp bridge node instead.
>>
>>109893115
>>109893511
So you just gave up and couldn't find it?
I will post this whenever you post stupid shit like that again
>>
>>109894377
im not good at prompting to be fair. its a lot of shit to remember for each model. im using https://github.com/Adudeguyman/ComfyUI-Fantastic-MiniMaxH3-PromptBuilder becasuse its convenient enough for what it offers, but i hate multiple things about it too. i can understand the references, how they link to a point, swapping shit, cropping shit, describing it all, but i cannot grasp where the underlying issue ultimately comes from especially when i prompt to fix it too
ive even tried dumping that prompting altogether and trying a generic alternative that still works but still has the same issue

is it because im using ref2v and not fl2v?
>>
>>109894409
i heard the reference model has problems with it. if it's not too important, it would quit trying to use a voice reference and just directly prompt the model to produce anime voices
>>
>>10989427
dat trunk
>>
>>109894387
share that node pls
>>
File: Qwen_image_2.1_00005.png (1.81 MB, 1024x1024)
1.81 MB PNG
>>
File: Qwen_image_2.1_00007.png (3.27 MB, 1440x1440)
3.27 MB PNG
>>
File: Qwen_image_2.1_00008.png (3.01 MB, 1440x1440)
3.01 MB PNG
>>
>>109894431
>mediafire.com/file/munyfqi7zqxr7x7/ComfyUI-LlamaPEBridge-v2.zip/file

It's a one-shot with zero debug checks.
since i use a separate folder for Llama.ccp., ell LLM to fix the folder path file (prepare_external_folder.bat)

>https://github.com/ggml-org/llama.cpp
Get your llama.cpp and cuda.dll here.
Extract to correct folder.

i use Q5(anything would work) + mmproj file to test its edit. Copy both files to the same I2I models folder
>huggingface.co/prithivMLmods/Qwen-Image-2.1-PE-I2I-GGUF/tree/main
>>
>>109894487
thx
>>
>>109894487
Don't forget to paste the official qwen image system prompt:

>https://github.com/QwenLM/Qwen-Image-2.1/tree/main/prompt_rewrite/prompts

into system prompt text slot
>>
File: Qwen_image_2.1_00010.png (3.56 MB, 1440x1440)
3.56 MB PNG
>>
qwen image 2.1 kinda sucks a bit.. sameface with group pics, mangled hands still.. like.. come on
>>
File: Qwen_image_2.1_00011.png (2.81 MB, 1440x1440)
2.81 MB PNG
>>
File: Qwen_image_2.1_00290.png (1.26 MB, 800x1216)
1.26 MB PNG
>>
File: Qwen_image_2.1_00015.png (3.12 MB, 1440x1440)
3.12 MB PNG
me and the senpai
>>
File: Qwen_image_2.1_00016.png (2.81 MB, 1440x1440)
2.81 MB PNG
>>
File: Qwen_image_2.1_00018.png (3.03 MB, 1440x1440)
3.03 MB PNG
>>
soon. trust the plan
>>
File: ComfyUI_07303_.png (1.66 MB, 1920x1080)
1.66 MB PNG
>>
File: 1771358287689912.png (1.62 MB, 896x1184)
1.62 MB PNG
>>
>>109894253
did you type all of that by hand?
>>
>>109894499
Does it work?
Do you see miner.exe in your Task Manager?
>>
File: 1771843584103126.png (2.07 MB, 1408x1088)
2.07 MB PNG
>>
>>109893037
how does harness help?
>>
File: 1780022957905640.png (2.48 MB, 1024x1536)
2.48 MB PNG
>>
File: ComfyUI_temp_xvgvj_00003_.png (3.22 MB, 1200x1920)
3.22 MB PNG
Why is that reddit AI nsfw porn content videos are just the same POV videos of pseudo-incest, old-guy young woman, bdsm, underage cartoons, always with the same shots, doggy, blowjob and cumshot, you can tell right away that most of them don't have sex
>>
>>109894828
if i wanted to talk about reddit, then i would be on reddit
>>
File: 1787651308227582.png (2.18 MB, 1888x1056)
2.18 MB PNG
so cooked
>>
>>109894828
because they dont know how2kino like you and i
>>
>>109894696
i came buckets
>>
>>109894828
who the fuck cares
>>
>>109894845
>>109894892
you guys are so fucking boring, no wonder these threads are dead
>>
>>109894904
>>109894904
>>109894904
>>109894904
>>
>>109894900
aww gee nobody wants to talk about reddit with you? QQ
>>
File: file.mp4 (681 KB, 576x736)
681 KB
681 KB MP4
>>109890948
Been mapping my characters over vidyas with minimax, pretty neat stuff.
"static camera view.
[Shot 1] Strict scene takes in picture 2 featuring blue anthro character with messy blue fur looking tired and barely able to keep eyes open from picture 4 seen sitting swaddled in soft white blanket on bed in picture 2 with big mug of hot coco steaming continuously forever in front of her on plate on wooden trayboard on bed. anthro character perfectly mimics character movements from video 1 in sleepy bleary eyed tired manner of character from video 1 in perfect sync in perfect loop.
dialogue: none
music: none
White bold lowercase thick kookie font text with black outline droopshadow centered at bottom says: good night i go slep"
>>
>>109894436
can notice is not 100%
and that reference as it not?

i tested qwen with face swaps.
it is south park copypasat mspaint tier.

waste of download. flux is better.
>>
>>109895790
>reference as it
***is ref is it not
>>
Been a while since used local AI.

Is ZIT still the best model for me? I have 12gb vram and 16gb ram.

For TE I use Qwen 3 4b.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.