[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1787020419153800.jpg (420 KB, 1200x1800)
420 KB JPG
updated OP edition

Previously on /sdg/: >>109573053

>Beginner UI
EasyDiffusion: https://easydiffusion.github.io
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI

>Advanced UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
Forge Classic: https://github.com/Haoming02/sd-webui-forge-classic
Stability Matrix: https://github.com/LykosAI/StabilityMatrix

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Z-Image
https://comfyanonymous.github.io/ComfyUI_examples/z_image

>Flux.2 Dev/Klein
https://comfyanonymous.github.io/ComfyUI_examples/flux2

>Chroma
https://comfyanonymous.github.io/ComfyUI_examples/chroma

>Anima
https://huggingface.co/circlestone-labs/Anima

>Models, LoRAs & upscaling
https://civitai.com
https://huggingface.co
https://tungsten.run
https://yodayo.com/models
https://www.diffusionarc.com
https://miyukiai.com
https://civitaiarchive.com
https://civitasbay.org
https://www.stablebay.org
https://openmodeldb.info

>Index of guides and other tools
https://rentry.org/sdg-link

>Related boards
>>>/aco/csdg/
>>>/b/degen
>>>/d/ddg
>>>/e/edg
>>>/gif/vdg
>>>/h/hdg
>>>/tg/slop
>>>/trash/sdg
>>>/u/udg
>>>/vp/napt
>>>/vt/vtai

OP https://rentry.co/twkuk8tz
>>
>>
I'm just a tourist here, and I've been lurking for a while, but what do you guys do with these? is this sort of a scratchpad you post while trading prompts/techniques? do you eventually generate final images that you don't share here? do you use these in actual projects or are you building a portfolio? are you thinking of getting a job with these?
>>
File: comfyui_00069_.png (1.36 MB, 832x1216)
1.36 MB PNG
>>
>>109587191
This thread is mostly a couple of bots and schizos that interact with each other once a week.
>>
File: comfyui_00070_.png (1.52 MB, 1216x832)
1.52 MB PNG
>>
>>109587065
ty for the bread.
>>
File: debo_ch_k2_00011_.png (2.92 MB, 2048x1101)
2.92 MB PNG
working on the news but it'll be a minute
>>
File: comfyui_00071_.png (1.64 MB, 1216x832)
1.64 MB PNG
>>
>>109587341
>>109587695
ptsd monke is here, less go!
>>
>>109587695
pleese animate sex with monkey and snell
make graphic remember peenis on head of snell
thank you very
>>
>>
File: comfyui_00072_.png (1.35 MB, 512x1536)
1.35 MB PNG
>>
Morning anons
>>
File: comfyui_00073_.png (1.39 MB, 512x1536)
1.39 MB PNG
>>109588390
morning
>>
>gm
>>
>mfw Resource news

08/18/2026

>Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model
https://yunpeng1998.github.io/Qwen-Video-Edit-Page

>ByteDance just released Bernini‑Diffusers‑v2
https://huggingface.co/ByteDance/Bernini-Diffusers-v2

>PixRestore: Unified Image Restoration via Pixel Diffusion Transformer
https://github.com/csslc/PixRestore

>PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion
https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site

>MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
https://expmaster.github.io/megaparts_webpage

>GenRouter: Unified Workflow Routing for Agentic Image Generation
https://github.com/EnVision-Research/GenRouter

>HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation
https://github.com/1nnoh/HiFi-BRep

>OvDSGG: End-to-End Open-Vocabulary Dynamic Scene Graph Generation
https://github.com/jhelsby/OvDSGG

>ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution
https://github.com/nmduonggg/ENAF

>FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
https://github.com/IGL-HKUST/FlexAM

>Pocket Video Trimmer: Trim/compress clips for MiniMax H3 Ref2V
https://huggingface.co/PoopMan333/Video_Tools

08/17/2026

>pagedMark: AI watermark removal built for Apple Silicon
https://github.com/doofzoff/pagedMark

>Stripe strikes mega-deal for OpenRouter
https://www.axios.com/2026/08/17/stripe-openrouter-paypal

>Context-Anchored Tile Refine: Nodes for tiled refining and upscaling
https://github.com/Blakeem/ComfyUI-ContextAnchoredTileRefine

>ComfyUI MiniMax-H3 SPEED Sampler: Runs SPEED (Spectral Progressive Diffusion) on packed video+audio latent
https://github.com/StanLukuvka/ComfyUI-MiniMax-H3-SPEED

>ComfyUI MiniMax H3 Parallel Attention
https://github.com/AesSedai/ComfyUI-MiniMaxH3-Parallel
>>
>mfw Research news

08/18/2026

>MLLM-Guided Semantic Correction for Text-to-Video Generation
https://arxiv.org/abs/2608.16513

>SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
https://arxiv.org/abs/2608.16220

>PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes
https://arxiv.org/abs/2608.15583

>Spatially-Grounded Flow Matching: Structured Source Distributions for Image Generation
https://arxiv.org/abs/2608.15452

>TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models
https://arxiv.org/abs/2608.15341

>KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation
https://arxiv.org/abs/2608.16154

>SQuad: Sub-Quadratic Attention Distillation for Efficient Video Generation
https://arxiv.org/abs/2608.16585

>Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
https://arxiv.org/abs/2608.15522

>TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
https://amuseum-whr.github.io/TraceBench

>FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams
https://arxiv.org/abs/2608.15818

>JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation
https://arxiv.org/abs/2608.15395

>GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
https://arxiv.org/abs/2608.16328

>Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
https://arxiv.org/abs/2608.16812

>PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
https://arxiv.org/abs/2608.16717

>Benchmarking Frontier Text-to-Image Models on Image-Description Prompts
https://arxiv.org/abs/2608.14976

>TISC: A Text-Driven Image Semantic Communication System for Faithful Reconstruction
https://arxiv.org/abs/2608.16100
>>
>mfw MORE Research news

>AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
https://serin-yoon.github.io/projects/anytalk

>Nexus: Structured Synergy for Efficient Text-to-Image Generation using Rectified Flow Model
https://arxiv.org/abs/2608.16104

>VicEdit: Learning to Edit Videos from Visual In-Context Examples
https://arxiv.org/abs/2608.16745

>RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation
https://arxiv.org/abs/2608.15694

>Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models
https://arxiv.org/abs/2608.16786

>An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
https://arxiv.org/abs/2608.16887

>Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
https://arxiv.org/abs/2608.16259

>Scalable Black-Box Model Attribution for Images
https://asaf-livne.github.io/RPA

>MOSS-VL Technical Report
https://openmoss.ai/MOSS-VL

>Do Visual Grounding Decoders Need Feed-Forward Networks? A Controlled Study over Frozen Vision-Language Features
https://arxiv.org/abs/2608.15061

>Image Denoising via the Adaptive Rank-Cluster Filter
https://arxiv.org/abs/2608.15298

>Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models
https://arxiv.org/abs/2608.16263

>Where did the ambiguity go? Examining how multimodal models interpret polysemous words
https://arxiv.org/abs/2608.00410

>Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
https://arxiv.org/abs/2602.12381
>>
File: debo_ch_k2_00012_.png (2.72 MB, 2048x1101)
2.72 MB PNG
>>
>>
>>
File: debo_ch_k2_00015_.png (2.97 MB, 2048x1101)
2.97 MB PNG
>>109589126
scaling is wonky but I loves me a steampunk airship
>>
File: comfyui_00077_.png (960 KB, 512x1536)
960 KB PNG
>>
File: comfyui_00078_.png (796 KB, 512x1536)
796 KB PNG
>>
>>
File: comfyui_00079_.png (1.17 MB, 1152x896)
1.17 MB PNG
>>
File: debo_ch_k2_00016_.png (2.92 MB, 2048x1101)
2.92 MB PNG
>>
What's the difference between this and ldg?
>>
File: debo_ch_k2_00017_.png (2.93 MB, 2048x1101)
2.93 MB PNG
>>109590003
sdg is where you go to meet new friends
ldg is where you go to meet new enemies
>>
>>109590003
Diffusion generals are a mess. In /ai/ board all these unbearable posters would be left alone in their own threads and proper discussion would exist in real threads.
>>
File: file.png (95 KB, 740x740)
95 KB PNG
I need to generate a simple logo for my project. it's just a picture of a tree, something like this. How do I avoid having it look like AI slop (smudgy textures, strange start/stop points of edges, etc)?
>>
>>109590284
gpt isn't too bad with logos. and you can iterate with text instructed edits
>>
>>109587065
Good afternoon frens.

I decided to learn how to train motion loras for the Wan 2.x i2v 14B models. Myb first attempt was to replicate the animation style (specifically how he draws ass motion) of D-ART. Here's the result of a test I did using the lora (well technically loras since the trainer exported a low noise and high noise version)

Source of the inout image:

https://rule34.xxx/index.php?page=post&s=view&id=9863612

Trainer I used:
https://github.com/ostris/ai-toolkit (Version 0.12.24) using a modified Wan 2.2 config.

I trained it at 12 fps instead of the default 16. Seems to work well when generating short clips at 12 fps. Its experimental and I'm debating weather or not I should train it more to see if it'll improve more and the dataset was ONLY of ass clips so that's pretty much all the loras are good for.

What do you guys think of it so far?
>>
File: comfyui_00080_.png (1.31 MB, 1152x896)
1.31 MB PNG
>>
>>109590397
jiggly
>>
>>109590397
>>109590406
same prompt and settings without the motion loras
>>
File: debo_ch_k2_00018_.png (2.84 MB, 2048x1101)
2.84 MB PNG
>>109590397
I have no opinions on the technology or the training journey, as i don't know much about either. the output though seems very goog and will usher in a golden age of twerking

https://www.youtube.com/watch?v=eF1lU-CrQfc

you may consider exploring what training for minimax-h3 would entail. thats the new hotness and has pulled most people off of using wan
>>
File: comfyui_00081_.png (1.33 MB, 1152x896)
1.33 MB PNG
>>
>>109589312
ty

yeah, the perspectives are broken.
>>
File: comfyui_00082_.png (1.05 MB, 512x1536)
1.05 MB PNG
>>
File: MiniMax_H3_00004_.webm (3.9 MB, 1152x640)
3.9 MB
3.9 MB WEBM
>>109590427
h3 with super lazy prompting fwiw
https://files.catbox.moe/v93iv4.mp4
https://files.catbox.moe/wetiig.mp4
(second one was w/ the turbo lora)
>>
File: comfyui_00083_.png (1.23 MB, 1280x768)
1.23 MB PNG
>>
>>
>>
>>
>>
>>
>>
File: comfyui_00084_.png (1.5 MB, 1280x768)
1.5 MB PNG
>>
>>109587491
>>109588901
>>109589312
>>109590106
Technigger
>>
>>
File: comfyui_00085_.png (1.55 MB, 1280x768)
1.55 MB PNG
>>
>>
>>
>>
>>
File: comfyui_00086_.png (1.28 MB, 896x1152)
1.28 MB PNG
>>
>>
File: debo_tt_k2_00153_.png (2.37 MB, 1872x1007)
2.37 MB PNG
>>
>>
>>
>>
File: debo_tt_k2_00157_.png (2.17 MB, 1872x1007)
2.17 MB PNG
>>
>>109592834
It that Spotty? Uniforms looking nice.
>>
File: debo_tt_k2_00161_.png (2.44 MB, 1872x1007)
2.44 MB PNG
>>109592874
just completely random people. krea2 has a decent handle on trek generally though. it doesn't quite understand the sets though; or maybe my prompts confuse it
>>
>>109592833
I like this one. Would you mind sharing prompt please?
>>
File: debo_tt_k2_00168_.png (2.57 MB, 1872x1007)
2.57 MB PNG
>>
File: 15433684357881359301.png (3.57 MB, 1632x1088)
3.57 MB PNG
>>
File: deBS_zi_00027_.jpg (612 KB, 1706x1920)
612 KB JPG
>>
>>
File: comfyui_00089_.png (1.43 MB, 896x1152)
1.43 MB PNG
>>
>>
i miss schizo anon
>>
>>109595043
sfw vageeeeeeeeeeeen
>>
kroma 0.3
and i guess lodestone's now going to distill kroma and fuck it all up as usual
>>
>>109592696
>>109592785
>>109592833
it'd be cool if u made coomables with that quality, if you post them to /b/'s AI thread bls link here
>>
>>
Everything here is terrible. Hasn't it been 4 years since diffusion models were released? How are you still so bad?
>>
>>
File: comfyui_00090_.png (1.43 MB, 896x1152)
1.43 MB PNG
>>
>>
>>
File: 66_x.jpg (23 KB, 1080x1350)
23 KB JPG
>>
>>
File: comfyui_00091_.png (1.66 MB, 1280x768)
1.66 MB PNG
>>
>>
>>
>>109595399
Kino goes in the real bread
>>
>>
gm anon
>>
>gm
>>
>>
File: comfyui_00094_.png (1.41 MB, 832x1216)
1.41 MB PNG
>>
>>109595699
gm
>>
>>
File: comfyui_00096_.png (1.45 MB, 832x1216)
1.45 MB PNG
>>
>>
>>
>>
Gm! Bot status?
>>
>>
>>
>>
>>
File: debo_ch_k2_00019_.png (3.33 MB, 2048x1101)
3.33 MB PNG
gm
>>
File: comfyui_00100_.png (1.3 MB, 832x1216)
1.3 MB PNG
>>109596575
good morning
>>
>gm
>>
File: 172656-tmp.png (2.88 MB, 1688x1688)
2.88 MB PNG
>>109596378
>>
>>
File: debo_ch_k2_00020_.png (2.92 MB, 2048x1101)
2.92 MB PNG
>>109596739
wow, fran still alive!
good to see you
>>
File: comfyui_00101_.png (1.34 MB, 1280x768)
1.34 MB PNG
>>
>>
File: debo_ch_k2_00022_.png (3.57 MB, 2048x1101)
3.57 MB PNG
>>109597008
requesting [1] arctic fox pls
>>
>>
>>
File: comfyui_00102_.png (1.39 MB, 1280x768)
1.39 MB PNG
>>109597034
i'll try
>>
Morning anons
Trying out early FLUX dev/schnell prompts on Krea 2
>>
>>
File: debo_ch_k2_00024_.png (2.62 MB, 2048x1101)
2.62 MB PNG
>>109597244
gm
>>
File: comfyui_00103_.png (1.42 MB, 1280x768)
1.42 MB PNG
>>
>>
File: debo_ch_k2_00025_.png (3.17 MB, 2048x1101)
3.17 MB PNG
>>109597265
heckin baserinoed
ty
>>
File: comfyui_00104_.png (1.31 MB, 1280x768)
1.31 MB PNG
>>
File: debo_ch_k2_00026_.png (2.88 MB, 2048x1101)
2.88 MB PNG
>>109597397
dat boy fluffy
>>
Corn!
>>
File: comfyui_00105_.png (1.09 MB, 1280x768)
1.09 MB PNG
>>
File: comfyui_00106_.png (1.15 MB, 768x1280)
1.15 MB PNG
>>
File: debo_ch_k2_00029_.png (3.46 MB, 2048x1101)
3.46 MB PNG
>>109597516
keep gettin fluffier
>>
>>
>>
>>
>>
>>
File: comfyui_00107_.png (1018 KB, 768x1280)
1018 KB PNG
>>
>>
File: comfyui_00108_.png (1.29 MB, 1280x768)
1.29 MB PNG
>>
File: comfyui_00109_.png (1.23 MB, 1280x768)
1.23 MB PNG
>>
>>109598394
cute furry turtle
>>
>fran still wants 0 engagement with debo
>>
>>
>>
File: comfyui_00111_.png (1.35 MB, 768x1280)
1.35 MB PNG
how i feel
>>
>>109598647
it's only gonna get worse, brah
>>
File: comfyui_00112_.png (1.04 MB, 768x1280)
1.04 MB PNG
mfw
>>
File: debo_ch_k2_00031_.png (3.17 MB, 2048x1101)
3.17 MB PNG
>>109598647
it's only gonna get better, brah
>>
>>
>>
>>
File: comfyui_00114_.png (1.07 MB, 768x1280)
1.07 MB PNG
>>
>>
File: debo_ch_k2_00033_.png (2.89 MB, 2048x1101)
2.89 MB PNG
>>
>>
File: comfyui_00115_.png (1.32 MB, 768x1280)
1.32 MB PNG
>>
>>
>>
File: comfyui_00116_.png (1.32 MB, 768x1280)
1.32 MB PNG
>>
File: debo_ch_k2_00034_.png (2.67 MB, 2048x1101)
2.67 MB PNG
>>
>>
>>109599318
ptsd monke, my faverit
>>
so it turns out if you push the cfg too high you should lower the latent multipliers, who would've thunk it
>>
>>
File: comfyui_00118_.png (874 KB, 1536x512)
874 KB PNG
>>
File: image23423.jpg (3.07 MB, 3364x1954)
3.07 MB JPG
what the fug
>>
>>109599523
all necessary
>>
>>109599475
Interesting. What's base-turbo? Is it a 50 step model?
>>
>>109599523
>>
>>109599532
top left next to green nodes, minimax
>>109599529
>>109599536
>>
>>109599532
oh you meant my file? it's kroma 0.2 base/raw (krea finetune by chroma's lodestone) mixed with the turbo version of the same (kroma 0.2 turbo). y ou can use it at low (8-12) steps
i use the cwb int8 (not strong)
https://huggingface.co/silveroxides/Kroma-Quant/tree/main
no need for turbo/bypass loras with that
>>
>>
File: debo_ch_k2_00038_.png (3.04 MB, 2048x1101)
3.04 MB PNG
>>
>>109599584
Ah. Got it. It's pretty nice.
How would it handle some of my prompts? Would chroma-girl mind?

Real photograph, cinematic realism — wide-angle shot of a dimly lit, industrial-era metal torture chamber shaped like a humanoid silhouette, its interior lined with sharp, inward-pointing spikes reflecting cold overhead fluorescents. In the center, a young Chinese female cosplayer, late teens, stands upright within the structure, her body framed by the metallic grid, posture rigid yet subtly tense — one hand resting on her hip, the other gripping the edge of a rusted metal bar. She wears a full-length black lace bodysuit with faux-leather corset lacing, sheer ruffled sleeves, thigh-high straps, and thin black suspenders; bare legs, visible beneath the fabric’s tension. Her long, dark hair cascades over her shoulders, adorned with a lace bunny ear headpiece integrated into her costume. Skin glistens faintly under harsh light — natural oil sheen, not artificial gloss — with visible contact shadows from the spikes and ambient glare. Eyes half-lidded, lips parted slightly, expression unreadable yet charged with quiet defiance. No footwear. Makeup: smoky eyes, glossy lips, subtle blush, nails painted light pink. Background: crumbling stone walls, scattered iron chains, distant flickering gas lamps.
>>
File: comfyui_00120_.png (1.21 MB, 1536x512)
1.21 MB PNG
>>
File: 3longkirk.png (1.94 MB, 1254x1254)
1.94 MB PNG
>>
>>109599669
kek
I'm trying to think of the pros...
>>
>>
baking
>>
new
>>109599718
>>109599718



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.