[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109894904

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Anima
https://huggingface.co/circlestone-labs/Anima
https://animastyles.thetacursed.com
https://tagexplorer.github.io/
https://animadex.net

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>
>my stick figure made it into the op
neat. i still feel hopeless but glad to know it found a home
>>
Is Qwen Image Edit 2.1 worth downloading? It's hard to tell from these recent threads since most of the images are just inane spam.
>>
>>109899090
its pretty smart and has some good pockets in its training data and some bad pockets. overall though its preddy gud.
>>
>>109899078
to be fair that gen is kino
has just the right speed too imho
>>
File: cs2_comfyUI.png (28 KB, 815x29)
28 KB PNG
>>109899111
I was pretty exited about Klein 9B too but after the initial enthusiasm, I haven't really done anything with it.
Also this guy was in my CS2 lobby just a moment ago. Very curious.
>>
>>109899122
i wanted the movement and speed to be more violent for liftoff but i shouldnt complain i suppose. space wise i didnt like the spiral turning into an arc either but that isnt that out of the norm either
>>
>>109899165
maybe i'm too high but it has some kind of "i'm leaving the mortal bounds"-energy
>>
>>109899158
Well Qwen2.1 is better than Klein hands down. And anon might not like it for t2i but from my testing it's pretty alright for a base model.
>>
>>109899006
>mfw Resource news

09/24/2026

>Making the MiniMax H3 Video VAE 2x Faster
https://blog.comfy.org/p/making-the-minimax-h3-video-vae-2x

>Qwen-Image-2.1 Text Encoder (Heretic) — GGUF · FP8 · bf16
https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF

>unsloth/Qwen-Image-2.1-GGUF
https://huggingface.co/unsloth/Qwen-Image-2.1-GGUF

>Ming-Image-0.1-Design GGUF
https://huggingface.co/realrebelai/Ming-Image_GGUFs

>MiMo-V2.6 series: Frontier intelligence, all the modalities, built in public
https://mimo.xiaomi.com/mimo-v2-6

>ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP
https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP

>Qwen-Image-2.1-Fun-Controlnet-Union
https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union

>Latent evolving World Action Model
https://github.com/XuejiFang/LeWAM

>Prompt Studio: Multimodal prompt studio for ComfyUI
https://github.com/tngklp/ComfyUI-Prompt-Studio

09/23/2026

>Qwen-Image-2.1-viggle-turbo — v0.2 (preview)
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

>Ming-Image-0.1-Design: 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs
https://huggingface.co/inclusionAI/Ming-Image-0.1-Design

>MiniMax-H3-Fun-Controlnet-Union-2.0
https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0

>Streaming Video Editing with Easy Adaptation
https://github.com/YujiaHu1109/SVEET

>WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
https://drexubery.github.io/WorldCrafter

>GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
https://jiah-cloud.github.io/GAE.github.io

>Visual Jev: Accurate and Efficient Decisions from Shared Visual Context
https://github.com/guanxuyu-sv/Visual-Jev

>Comfy Router: One API for Frontier Media Models
https://blog.comfy.org/p/introducing-comfy-router-one-api

>qwen image studio
https://github.com/janishar/qwen-image-2.1-studio
>>
>>109899218
>mfw Research news

09/24/2026

>InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation
https://arxiv.org/abs/2609.27620

>ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming
https://jiayi-hit.github.io/ZoomDiff.github.io

>MotionSpec: Spectral Trajectory Supervision for Motion-Consistent Video Generation
https://arxiv.org/abs/2609.28095

>The Past Frames the Future: Memory for Autoregressive Video Generation
https://arxiv.org/abs/2609.28466

>Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality
https://arxiv.org/abs/2609.27493

>All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
https://arxiv.org/abs/2609.27901

>Gender Bias in Vision-Language In-Context Learning
https://arxiv.org/abs/2609.27682

>On the Diffusibility of High-Dimensional Latents
https://cfeng16.github.io/on_the_diffusibility

>Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges
https://arxiv.org/abs/2609.27110

>ASAP: Visual Analytics for Identifying and Analyzing Image Patterns in AI-generated Images
https://arxiv.org/abs/2609.27371

>Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows
https://arxiv.org/abs/2609.27327

>DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video
https://arxiv.org/abs/2609.27470

>WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
https://arxiv.org/abs/2609.27033

>What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit
https://arxiv.org/abs/2609.27408

>From PyTorch to the NPU: LLM-Agent-Driven Model Conversion Across Heterogeneous Inference Runtimes
https://arxiv.org/abs/2609.27249
>>
>debo the idiot still posting his spyware links
>>
>>109899197
I guess it's worth the ComfyUI update then...
>>
File: 1789260714840166.jpg (1.29 MB, 2368x1776)
1.29 MB JPG
>>109899221
>>109899218
>>
>>109899218
>>109899221
Fuck off unemployed cripple
>>
>>109899221
>>109899218
thanks for the news!
>>
File: TestE_00017.png (2.01 MB, 1024x1536)
2.01 MB PNG
>>109899090
It's certainly not perfect, but editing is a lot of fun.
>>
Best trainer for Qwen?
>>
>>109899314
autussy
>>
Currently streamlining the qwen 2.1 inpainting node to be more concise and doing general cleanup. I should have it up on github by the weekend at the latest
>>
Is there anything specific anon would like for me to generate?
>>
File: mai_00008_.png (1.33 MB, 1536x1536)
1.33 MB PNG
>>
File: 698860840.png (2.46 MB, 992x1472)
2.46 MB PNG
>>
>>109899355
Prof. Oak
>>
>>109899454
no
>>
File: 2029.png (3.32 MB, 1120x1664)
3.32 MB PNG
>>109899540
yes
>>
>>109899582
i do
>>
>>109899582
>>109899605
no.
>>
how to make an image look like an mspaint drawing? certain prompt or lora?
>>
https://huggingface.co/madebyollin/texture-fix-vae-for-qwen-image-2.1
>>
>>109899712
Nice
>>
>>109899728
>>109899712
the official vae is tiny, very weird.
>>
File: Qwen_image_2.1_00045.png (2.93 MB, 1440x1440)
2.93 MB PNG
>>
File: check_00006_.png (662 KB, 1256x840)
662 KB PNG
>>109899701
There's a lot of different styles you could do with mspaint. Maybe give Krea2 a try.
>>
>>109899827
Forgot the prompt:
>A sloppy Microsoft Paint drawing of Asuka Langley Soryu and Misato Katsuragi sitting at a table eating. It looks like someone drew it fast in Paint.exe and never cleaned it up: shaky freehand outlines that often fail to close, gaps where the bucket fill leaked onto the white canvas, leftover pencil strokes that miss the corners, no anti-aliasing, no gradients. Asuka on the left has long orange twin tails and a red shirt, drawn with uneven loops instead of boxes. Misato on the right has long wavy dark-purple hair and a yellow top, same loose linework. The table is a crooked brown shape with two roughly round plates and a few messy food scribbles. Faces are simple but not geometric. Empty pale canvas around them. Static scene, no motion lines, unfinished MS Paint look.
>>
>>109899701
Give it a reference image of what a good MS Paint drawing looks like; tell it to copy the artstyle only
>>
>>109899827
>>109899847
>>109899852
thanks, I'll give it a try
>>
I'm not a trainer, but does it seem they include more nsfw images in the dataset nowadays?
The base model understands the concept right away.
>>
I'm glad people are finally seeing the potential of qwen 2.1
>>
>>109899982
it feels no different than nano banana. well at least what I remember from it, haven't used it in a while
>>
i can't believe niggas willing to pay 1000$ for 5060 ti
>>
>>109900049
I doubt they are desu
>>
I don't want to be bullied. If I were to post an image, what are some allowed subjects here?
>>
>>109900079
Loli/shota is encouraged
>>
109900089
>who would do that
>go on the internet and tell filthy lies
>>
>>109899827
>>
File: 8.jpg (685 KB, 1792x2400)
685 KB JPG
>>
File: debo_cr_k2_00004_.png (2.84 MB, 1664x1069)
2.84 MB PNG
>>
>>109900215
Like always it looks like shit, go back to your tard cage
>>
>>109900215
lonely again?
>>
>>109900089
But in terms of art and all that?
>>
Any model for de-censoring Japanese pixel crotches easily?
>>
>>109900223
I think it looks pretty cool. Better than your gens.
>>
>>109900115
Yet another model that can't drawl.
>>
>>109900215
The people should be blurry too.
>>
>>109900079
cool stuff. artistic.

>>109900214
>>109900089
no
>>
>>109900317
Anti-art and 'shart' posting is favoured here.
>>
>>109900303
Pure distilled cope
>>
File: I Was a Mover.jpg (149 KB, 575x1024)
149 KB JPG
>>109900359
>>
>>109900324
nope, we like artsy art-fart
>>
File: 97.jpg (426 KB, 1632x2176)
426 KB JPG
>>
>>109900377
>we
Are you schizophreniac?
>>
File: Qwen_image_2.1_00046.png (3.3 MB, 1440x1440)
3.3 MB PNG
sometimes you just gotta think about how the fuck you got here
>>
File: steve1.jpg (146 KB, 573x1011)
146 KB JPG
>>
Have there been any major improvements to H3 in the past couple weeks? Either speed-wise or generation quality?
>>
>>109900410
>those pant legs staying up through sheer will force
AGI soon
>>
>>109900410
no

>>109900411
I'm white, we say we.
>>
>>109900450
?
>>
>>109900432
More cope vae shilling but that's about it
>>
I think this might be a good sample image for demonstrating the inpaint node
>>
File: Krea2_turbo_04399_.jpg (1.39 MB, 1776x2368)
1.39 MB JPG
>>109900478
fucking hell
Thinking of adding a speech bubble example, shirt change example and giving her glasses
>>
>>109900486
What is your art lora?
>>
>>109900486
can it make the text painterly?

I'm trying to get it running on my machine, on my own vibe fork of stable diffusion cpp
>>
>>109900495
This is base krea 2 the edits will be with qwen though
>>109900496
I'll find out give me a few minutes
>>
File: ComfyUI_01257.jpg (910 KB, 1800x1200)
910 KB JPG
>>
>>109900496
>>
best scheduler for qwen image?
>>
File: 09626-1322039961.png (875 KB, 768x1152)
875 KB PNG
>>109900500
a few minutes?... bro... you've got seconds.
>>
File: ComfyUI_01258.jpg (828 KB, 1800x1200)
828 KB JPG
>>
File: debo_cr_k2_00010_.png (2.56 MB, 1664x1069)
2.56 MB PNG
>>
>>109900543
eurler simpl
>>
>>109900543
what
>>109900572
Said

Glasses example
>>
>>109900583
T-shirt change, I could prompt it better but whatever
>>
File: ComfyUI_01236.jpg (716 KB, 1800x1200)
716 KB JPG
>>
File: 00207-2557236228.png (1.62 MB, 1344x768)
1.62 MB PNG
>>
File: artwork.png (1.82 MB, 1248x832)
1.82 MB PNG
>>
>>109900643
Do you quantize this image futher? If I wasn't still used gimp I would try some form a pixelation in Photoshop.
>>
File: ComfyUI_01266.jpg (782 KB, 1800x1200)
782 KB JPG
Last one... I hope I won't get banned.
Next one is about Cliver Barker's book... I doubt this model can do anything.
>>
File: Qwen_image_2.1_00327.png (3.01 MB, 1920x1088)
3.01 MB PNG
>>
>>109900712
final straw.. get him bois
>>
File: debo_cr_k2_00014_.png (2.24 MB, 1664x1069)
2.24 MB PNG
>>
File: ComfyUI_01268_.jpg (759 KB, 1800x1200)
759 KB JPG
>>109900759
Oh, I am afraid of you.
>>
>>109900534
Neat! (I actually meant on her shirt lol)
>>
>>109900789
give me that 2^3 rubix. I can solve it!
>>
>>109900827
Let's see yours first.
>>
File: artwork2.png (1.97 MB, 1248x832)
1.97 MB PNG
>>
MOM! A GUY ON THE INTERNET WANTS TO SEE MY RUBIX CUBE
>>
>>109900860
Everything what you do with this thread. Is like someone's kid party.
Narcissism.
>>
File: 4557845437894.png (18 KB, 900x806)
18 KB PNG
where is kino?
>>
>>109900879
Animanon's computer
>>
File: Output.jpg (823 KB, 1800x1200)
823 KB JPG
>>
>>109900763
You’re quite a pushy sort, aren’t you?
>>
>>109900937
This is what they get.
>>
>>109900944
You don't get anything.
>>
Anon yearns dearly for my kinosovl
>>
>>109900958
I was trained to be a mental health carer.
>>
File: Tree.jpg (1 MB, 1800x1200)
1 MB JPG
In this tree.
You can see...
>>
>>109900975
please
>>
krea2 is total dogshit
>>
>>109900993
So everything about 'jesus' is wrong.
>>
File: Qwen_image_2.1_00331.png (490 KB, 1275x602)
490 KB PNG
>>
File: 0.jpg (482 KB, 1120x2112)
482 KB JPG
>>
File: 1781437177645116.gif (60 KB, 640x640)
60 KB GIF
Is it feasible to gen video on x2 rtx2070s (8GB total VRAM) yet? Any rentrys I can use to get started

I want to animate album covers on my home media server, Spotify style
>>
>>
>>109901002
You are quite stupid.
>>
>>109901036
The problem is that - I have the doors open, you don't have them open.
>>
>>109901027
>I want to animate album covers on my home media server, Spotify style
Neat idea I'm stealing it.

>Is it feasible to gen video on x2 rtx2070s (8GB total VRAM) yet?
Hm... good question
>>
>>109900993
>>109901036
>>109901039
All me btw
>>
>>109901041
>good question
Thanks! I thought of it myself!
>>
jesus saves
>>
>>109900879
novram esl:
>>109900870
>>
>>109901084
spam bot is not kino
>>
File: 65331593.png (2.73 MB, 960x1216)
2.73 MB PNG
kek
>>
>>109901027
Sharding seems to be possible. Look up raylight maybe. I think most of the sharding stuff is for more... modern cards thoever.
>>
>>109901119
you might be a faggot but you have good taste anon
>>
Anyone know what to to use for pure pose transfer in Krea2? Not just extracting the pose from one image, but applying it to other image with as little other changes as possible.
>>
>>109901244
yes
>>
anyone run ComfyUI MCP with claude agent?
Is it worth it?
>>
can you take an image and change where the viewer/camera is with qwen 2.1? for example, say there's a bed on the left and a window on the right. and i want the viewer to be with the window on his back, looking at the bed. but with the other objects in the scene relatively consistent (i realize things are gonna look slightly off etc)
>>
>>109901256
I generate prompts with language models but I don't like giving it direct control.
>>
>>109901119
no.
>>
File: 00042-1793547339.png (1.61 MB, 832x1216)
1.61 MB PNG
>>109900664
idk, i kinda like this level of pixelization.
>>
>>109901262
This is what ai does to your brain. Imagine not being able to formulate a coherent sentence because you're so used to spamming ai with absolute schizo talk that you now write in the same way to people
>>
>>109901297
here's my apology to you: https://litter.catbox.moe/nyjlhbgt116cakkt.gif
>>
>>109901119
not a fan of that hair
>>
>>109901324
>me slapping you for speaking Ai-bonics
apology accepted
>>
>>109901262
Draw a pointer where the camera is pointing at, then use LLM with vision to write the prompt
>>
>>109901284
sexo
>>
>>109901080
put LDG in this one desu
>>
File: debo_cr_k2_00031_.png (2.39 MB, 1664x1069)
2.39 MB PNG
>>
>post gen of cute grill to one of the jailbait threads on /b/
>5 replies all asking for more not realizing its a gen
>>
>nigbo
>>
>Being able to fool a coomer is akin to fooling the elderly
>>
the world is full of coomers and the elderly
>>
It's up
https://github.com/CatJak/qwen2.1_inpaint
Let me know what you think
>>
>>109900664
Thank you for you energy. I received it.
>>
>>109901509
Hype
>>
>>109901509
looks cool, will try it out when i can
>>
>>109901509
I vibed my own interface.
>>
>>109901509
looks like vibeslopped shit
>>
>>109901548
what software isnt these days
>>
>>109901551
then don't share it if you know it's shit
>>
>>109901551
:^) Cursor Grok 4.7 extra high vibed me my own interface - AND support for NAG in Flux 2 dev.

no relation to the gay furry guy.
>>
>i will have my computer generate slop images
>but i would NEVER EVER have my computer generate slop code
>>
>>109901548
There's nothing hiding it read the acknowledgements.
>>
>>109901509
I wonder if there's a cheeky way to get a model to generate the mask instead of requiring the user to do it. I'm sure there is.
>>
>>109901574
I think that would require SAM or something like that, the problem is having enough vram to do it desu, might need to load and unload the llm model which would take time
>>
>>109901555
Please tell me it's not webshit at least
>>
>>109901596
hmm how odd of you to use that term
>>
>>109901602
It's the one thing I agree with the schizo on, the one thing
>>
File: Krea2_Turbo_00076_.png (1.26 MB, 1024x1024)
1.26 MB PNG
LoRA Training seems much easier on Krea2, I struggled to get my dragonborn down in SD1.5
>>
>>109901607
ani was never the schizo
>>
>>109901548
>>109901619
calm down schizo
>>
>>109901626
Try again :) here is a hint: I use rust
>>
>here is a hint: im trans
>>
>>109901596
Of course it uses a web browser. Web browsers are great, but they shouldn't be used with the Internet.
>>
>>109901632
I don't think you understand the other user isn't trying to peg you as one person, he's making fun of you for sperging out over pedantic shit
>>
>>109901608
why...
>>
>>109901646
Because I wanted to make campaign art for our D&D campaign and can now easily train a LoRA for each party member. Why not? With local its easy.
>>
>>109901654
I don't think some people realize how fast and easy it is to do these things if you have strong enough hardware
>>
https://www.youtube.com/watch?v=szagdbR4-NY
>>
File: 1650126689019.jpg (32 KB, 267x274)
32 KB JPG
now civitai is asking me to wait a day, for a lora training. are they based now in india or something?
>>
>>109901661
Curating the commissioned artwork and a few SD1.5 gens took way longer than the actual rest of the process. The lora training took like 1.5 hours on a 4090 while I was eating dinner.
>>
>>109901662
That's Frentch FFL.
>>
>>109901509
https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union
>>
>>109901689
Oh that's cool!
>>
Just to be clear I made this for myself and anons asked, I don't really mind what happens I'm really interested in that controlnet repo though, might make something for myself to use that too.
>>
>>109901727
You samefag all your compliments catjack
>>
can someone explain to me why its bad for a certain group of people to avatar in /ldg/, but its okay for this guy to avatar in here???
>>
>>109901731
If it helps you sleep at night
>>
how is qwen 2.1 for anime content, i really care about accuracy and how much brain damage i'll need to endure to get it right
>>
>>109901760
Meh without a reference
>>
>>109901509
What's the point of this and why would I use it over the native inpaint workflow that already exists and is compatible with every model in existence?
>>
>>109901741
There's just nothing anyone can do about it.
>>
>>109901741
sorry your avatar isnt as cool
>>
>>109901791
i dont have an avatar THOUGHIE im a truenon
>>
>>109901773
I made an all in one that functions like the webui interface, it was a personal project for myself. It's also compact and self contained. I didn't like the native comfy method
>>
>>109901760
Just use Anima
>>
>>109901832
janima amirite
>>
File: 1769673352850736.png (2.92 MB, 1632x1216)
2.92 MB PNG
So slopped sometimes.
>>
>>109901741
Can you point out the avatar?
>>
File: CouldBeAnyoneNigga.jpg (1.16 MB, 1776x2368)
1.16 MB JPG
>>109901877
I too wonder who he is talking about
>>
>>109901877
the avatar of the nametroon
>>
>>109901887
>Doesn't understand what a avatarfag is
/sdg/ is not sending it's best
>>
>36 stars behind his arch nemesis
ngmi promptlet
>>
>>109901654
:^) did you know you can do accurate iching with 6 d6, one for each line? (or re-roll, obviously).
>>
>>109894391
>>109899712
Your response?
>>
>>109901916
>So you just gave up and couldn't find it?
Nothing changes
>>
>>109901933
Post something with clean texture in a high-quality format
Yours is rigged with bad JPEG artifacts and painting texture
I don't even know you used qwen vae to begin with
>>
>>109901946
You couldn't give a good response at the time and acted like a sperg seething at everyone in the thread. I'm sorry you have anger issues, I must of left quite an impression on you.
You're dismissed
>>
>>109901916
thank you for reminding me to download it
>>
this is a great blog
>>
File: debo_cr_k2_00032_.png (2.63 MB, 1664x1069)
2.63 MB PNG
>>
File: Screenshot 2026-09-25.jpg (93 KB, 925x503)
93 KB JPG
here you go nigger
you can't out smart me with your beggar's trick

https://files.catbox.moe/fmgc6l.png
>>
File: 1778399904448645.jpg (41 KB, 463x618)
41 KB JPG
I give up with minimax
I only want it for porn and the results for that tend to be pretty bad
>>
File: Krea2_turbo_04407_.jpg (1.42 MB, 2368x1776)
1.42 MB JPG
Just let it go bro your cortisol is too high.
You lost your chance and are just salty now, the image had one masked segment that used the qwen vae and you couldn't find it, you fail here because you're regenerating the whole image when that wasn't the ask.
>>
>>109902058
what are you niggas trying to prompt for that h3 cant do?
>>
File: 1779876287261400.png (2.38 MB, 1408x1408)
2.38 MB PNG
>>
Why isn't there a good NSFW lora for qwen 2.1 yet?
>>
>>109902100
its a brand new model
>>
>>109902104
It's been at least 2 days. That's long enough.
Qwen2.1 is DOA
>>
File: 1764477428108889.png (3.31 MB, 1216x1632)
3.31 MB PNG
>>
>>109902116
is that some dandruff topping? i love dandruff topping! especially on my dogs
>>
>>109900879
>109900879
dunno
>>
>>109902083
the entire point of local AI is porn
for non porn you're better using online services with jewgle datacenters
>>
>>109902060
I love the images of these big booby babes. I save them, make images of it and chatbot it up. festive stuff. the harem grows.
>>
>>109902058
...have you not looked at the /gif/ ai threads yet?
>>
can someone please share a krea 2 identity workflow that actually works?
>>
File: QwenImageEdit2dot1_0311.png (2.98 MB, 1664x1056)
2.98 MB PNG
>>109900215
>>
>>109902149
he said the h3 porn is bad. i find that hard to believe
>>
File: debo_cr_k2_00036_.png (2.49 MB, 1664x1069)
2.49 MB PNG
>>109902088
very mid/late 90s electronica/trance type of album cover
>>
File: QwenImageEdit2dot1_0310.png (3.06 MB, 1152x1728)
3.06 MB PNG
>>
Oh we're shifting to this because the retard stumbled again
>>
File: debo_cr_k2_00045_.png (2.62 MB, 1664x1069)
2.62 MB PNG
>>109902166
huge missed opportunity for him not to be wielding the chicken like the sword

>>109902173
>mfw
>>
I just trained a krea2 lora on two drawings I made for the fun of it and... this should be illegal. Took less than 2 minutes and it can faithfully recreate the art style of those two images on everything now. And it's so accurate holy shit
>>
File: QwenImageEdit2dot1_0312.png (2.3 MB, 960x1216)
2.3 MB PNG
>>109901119
>>
>>109902194
approved.
>>
File: QwenImageEdit2dot1_0313.png (2.95 MB, 1664x1056)
2.95 MB PNG
>>109902183
ended up re-rendering it, so that the sword is gone...behold!
>>
Did any of the NSFW finetunes of h3 ever end up worthwhile?
>>
>>109899090
my initial impression was that it is smarter than klein 9b, but more prone to just copy+pasting the two images into each other when you give it two sources, which is very annoying. haven't tried it that much with just an instruction on one source, maybe it's better at that, and maybe i need an LLM to enhance my shit into a boomerprompt to get the blending happening properly
>>
File: QwenImageEdit2dot1_0316.png (1.33 MB, 1344x768)
1.33 MB PNG
>>109901541
Good vibes only <3

>>109902233
It's pretty easy to get the 'blend' going better, but usually it takes me one or two takes to specify what exactly to blend - a general prompt works well enough.
>>
File: 238047.jpg (1.28 MB, 2896x1632)
1.28 MB JPG
>>109902229
eros v5 works pretty well as long as you can live with a little image fry. And base h3 is great for vanilla stuff. All fl2va models of course, not ref2va
>>
File: Krea2_turbo_04422_.jpg (1.46 MB, 2368x1776)
1.46 MB JPG
>>
how do we deal with the avatar problem?
>>
>>109902253
so ref2va bad, but fl2va good. hows the audio though since thats always been fucked hasnt it?
>>
>>109902275
Just report the pepe spammer for avatar spam
>>
File: 423432423.jpg (1.22 MB, 2560x1440)
1.22 MB JPG
>>109902286
Audio is okish if you're at or above 32 steps. Though if you're just prompting voice and dialogue and no sfx then I recommend a node that converts the audio signal using tts to get a voice that doesn't sound like shit or monotone. Though base H3 does have a lot of problem with porn audio, so sucking literally sounds like someone sucking on a straw and that kind of thing. But I think eros fixed it a bit? Not entirely sure
>>
>>109902229
Base already knows basic concept, just prompt properly & add mystic_xxx lora on top of it
>>109902253
Eros is not a finetune
Most finetune rn focus on speed and quality. None of them really add new concept
>>
>>109902296
what about the other avatarers?
>>
>>109902312
kewl style anon
>>
>>109902316
Only one faggot is spamming the same character right now
>>
>>109902312
>above 32 steps
so my steps were always just too low realistically, capping them at 20. i hate that most of the resources are using turbo shit and all sorts of other nodes to fry it to fuck just to make it gen faster, but i want more quality. downside is the vram is rough, 24gb is just enough to get it to load but not much past that without swapping to system resources.
>>
>>109902331
what do i do when the other avatarers start doing it? like that avatar that has a paper bag over his head
>>
>>109902355
Can you not read the previous post retard kun?
>>
File: 1786694930739035.png (1.09 MB, 1888x1056)
1.09 MB PNG
>>
>>109902345
Don't discount turboshit, from my experience turbo at 2.0MP absolutely mogs 50 steps 1.0MP. Except for the audio of course, that sounds like shit. But there are also some tricks I found that lets you improve turboshit audio massively
>>
>>109902364
which one?
>>
File: Krea2_turbo_04423_.jpg (1.45 MB, 2368x1776)
1.45 MB JPG
>>
im bored, post nudes
>>
File: MiniMax_H3__00805.mp4 (1.38 MB, 576x896)
1.38 MB
1.38 MB MP4
>>
>>109902403
Does that box belong to Ani?
>>
>>
>>109902412
i dont know, who is that
>>
>>109902372
care to elaborate a little more on that? im somewhat interested mostly because i want to both stop wasting time, but also get the quality result i desire. audio will be fucked forever im sure unless i use a secondary method to handle it, but still
>>
Can i do video if i cant rotate an apple in my head?
>>
File: debo_cr_k2_00061_.png (2.72 MB, 1664x1069)
2.72 MB PNG
>>109902275
>the avatar problem
the fire nation could help

>>109902465
generate a video of a rotating apple and then think about it
>>
there we go.
>>
File: 34754838.jpg (337 KB, 1928x1088)
337 KB JPG
>>109902463
Elaborate on what exactly? The audio thing or?
>>
File: MiniMax_H3__00811.mp4 (729 KB, 896x576)
729 KB
729 KB MP4
>>109902369
>>
File: 1763198536834650.png (2.35 MB, 1280x1568)
2.35 MB PNG
>>
>>109902058
lern2fetish prompt. you don't need to jerk off to just boobs and dick thrusting
>>
>>109902493
tricks that can make turbo usable despite the audio being fucked
and then elaborate on any tips for audio separately

my current strategy has been ref2va because i use 4+ references on average, so i2va and fl2va are not enough for me so i need to really figure out how to do this ideally
>>
>>109902493
Kino
>>
>>109902512
it's disguting lol, but still qwen?
>>
File: MiniMax_H3__00816.mp4 (3.13 MB, 576x896)
3.13 MB
3.13 MB MP4
>>109902512
>>
>>109902526
I can honestly only help you out with the fl2va model and my strategies for quick but high quality turbo videos, let me know if you wanna know more before I waste my time writing up a storm. But if you're looking for ref2va advice then I only have one, use 32 steps and spectrum, there are no shortcuts for ref2va, it kind of sucks which is why I was hoping they'd release H3 Max but alas, here we are
>>
>>109902550
keeek
>>
>>109902550
whered you get this video of me
>>
>>109902551
any tip helps, small or large. just drop whatever you feel has the most value, i dont need all of it. i do have some value higher than room temperature IQ so i can fill in and test on my own but just some points to start at should suffice
>ref2va
>spectrum and 32 steps
works for me already.

for fl2va just like 1 or 2 major suggestions will work. im already annoyed as fuck with how hard it is to go past 15 seconds but i can save that for another time
>>
File: 1763306934164909.png (1.97 MB, 1408x1408)
1.97 MB PNG
>>109902536
yessum https://files.catbox.moe/t5bd25.png
>>
>>109902558
hidden camera
>>
File: 777438884.jpg (364 KB, 1844x1032)
364 KB JPG
>>109902563
I was going to explain how it works but realized how difficult the theory behind it is to explain, so I just spill the gatekept secret sauce.
>lightfx 4-step 1.2 lora
>spectrum at default settings
>8-9 steps depending on whether it starts to do two steps at once from the first or second iteration step (just about saving time, when in doubt just do 9)
>0.98MP minimum, 2MP for best results
>euler, simple
>ref image size max
>image reference set at 2MP (lower doesn't save much time, higher doesn't increase quality)
>Always 16:9 for both the video and reference if you want optimal quality and prompt adherence
>look for a good seed and lock it in (might be placebo)
>comfykitchen att. for free speedup
>Set turbo lora strength up to 1.1 if it ever bugs, otherwise leave at 1.0
>Switch to eros5 when doing anything that's more than a handjob
>hmcumshot lora at 0.6 if you want a cumshot video
That's about it anon
>>
can we just get an image model on level of krea2 or qwen image 2.1 but with ability to do image references well like minimax and with optional bbox prompting. fuck loras, minimax does references quite well, and refmods are very convenient
>>
>>109902666
you didnt need to go that far into the detail but i appreciate it regardless. good invader zim gen too
>>
just woke from a qwen 2.1 release data coma, was the text 2 image workflow ever fixed so its not dogshit anymore?
>>
>>109902694
release date
>>
>>109902692
Well it doesn't work if you don't do everything as I wrote up. And the difference in quality to what I've seen from other peoples turbo gens is massive, since it doesn't have those motion artifacts and shitty hair that regular turbo would have
>>
File: MiniMax_H3__00822.mp4 (3.47 MB, 640x832)
3.47 MB
3.47 MB MP4
>>109902512
i really like this pic.. can be so many things
>>
>>109902708
i never said i wasnt going to do everything just that the specifics were more than expected. i can work with it regardless though. i have a particular aversion to fl2va
>>
>>109902710
much more pleasant.
>>
File: 7544576.jpg (222 KB, 1368x768)
222 KB JPG
>>109902723
Don't forget to post kinos here, should you make one
>>
>>109902768
see >>109899078
>>
can local diffusion fix my dental health?
>>
>>109902806
Yes
>>
ok one more
>>
File: Qwen_image_2.1_00047.png (3.74 MB, 1440x1440)
3.74 MB PNG
>>109902806
>>
File: Qwen_image_2.1_00048.png (3.4 MB, 1440x1440)
3.4 MB PNG
>>
File: debo_cr_k2_00064_.png (2.37 MB, 1664x1069)
2.37 MB PNG
>>109902768
is this a new oc?
>>
>>109902823
eheheheh dem fangers
>>
*
>>
>>109902836
No, I just have a preference
>>
File: 5413.jpg (1.78 MB, 2324x4096)
1.78 MB JPG
>>
first qwen 2.1

(just 4 steps, because I'm fixing stuff in my local diffusion jackler)

*a jankler is a clanker production with little to no planning, by the seat of the pants, and *exactly* how it works may be rube goldbergish.
>>
>>109902871
here!!!
>>
>>109902843
those are bad to have, i'd get rid of it if i were you
>>
File: ComfyUI_temp_kpied_00356_.jpg (457 KB, 1008x1600)
457 KB JPG
Are there any workflows you would recommend for faceswaps? is that even a thing?
>>
>look at the paid-only loras on civit ai
>it's all dogshit
>>
gemini flash 3.8 has a suggestion.
>>
bitches don't know bout my plane proof buildings
>>
>>109902885
Some are into men and others have a preference
>>
>>109902919
no wonder everyone on that plane died, that thing has almost no crumple zone, the impulse is hella through the roof
>>
>>109902898
slopsite
>>
>>109902930
No, they have seatbelts.
>>
i wish anon made the space shuttle explosion kino
>>
>>109902874
suggestions? it put the lady there...

>The woman from image 1 is driving a car with a frown. Beside her is a fat man who has a beard, long hair, and glasses, with smudges on his face. He's smiling. The text says "state sponsored gf acquired #neetlife"
>>
>>109902930
the point wasn't to save the passengers of course.. it's to save the property
>>
hm...
>>
>>109902710
bratwurst sausage
>>
Indeed
>>
>>109902994
>>109902994
>>109902994
>>109902994



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.