[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1785887071847136.webm (3.56 MB, 544x960)
3.56 MB
3.56 MB WEBM
frens edition

Previously on /sdg/: >>109458081

>Beginner UI
EasyDiffusion: https://easydiffusion.github.io
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI

>Advanced UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
Forge Classic: https://github.com/Haoming02/sd-webui-forge-classic
Stability Matrix: https://github.com/LykosAI/StabilityMatrix

>Z-Image
https://comfyanonymous.github.io/ComfyUI_examples/z_image
https://huggingface.co/Tongyi-MAI/Z-Image
https://huggingface.co/Tongyi-MAI/Z-Image-Turbo

>Flux.2 Dev/Klein
https://comfyanonymous.github.io/ComfyUI_examples/flux2
https://huggingface.co/black-forest-labs/FLUX.2-dev
https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
https://huggingface.co/black-forest-labs/FLUX.2-klein-9B

>Chroma
https://comfyanonymous.github.io/ComfyUI_examples/chroma
https://huggingface.co/lodestones/Chroma1-HD
https://huggingface.co/silveroxides/Chroma-GGUF

>Anima
https://huggingface.co/circlestone-labs/Anima

>Qwen Image & Edit
https://docs.comfy.org/tutorials/image/qwen/qwen-image
https://huggingface.co/Qwen/Qwen-Image

>Text & image to video - Wan 2.2
https://docs.comfy.org/tutorials/video/wan/wan2_2

>Models, LoRAs & upscaling
https://civitai.com
https://huggingface.co
https://tungsten.run
https://yodayo.com/models
https://www.diffusionarc.com
https://miyukiai.com
https://civitaiarchive.com
https://civitasbay.org
https://www.stablebay.org
https://openmodeldb.info

>Index of guides and other tools
https://rentry.org/sdg-link

>Related boards
>>>/aco/csdg/
>>>/b/degen
>>>/d/ddg
>>>/e/edg
>>>/gif/vdg
>>>/h/hdg
>>>/tg/slop
>>>/trash/sdg
>>>/u/udg
>>>/vp/napt
>>>/vt/vtai

OP https://rentry.co/twkuk8tz
>>
>mfw Resource news

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo – OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view

>OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
https://xin1u.github.io/OminiVR_PAGE

>DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
https://github.com/Zhong-Chenchen/DIVE.git

>Multi-View Face and Gesture Animation with Dynamic Gaussians
https://dfki-av.github.io/MVFGA

>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
https://empaava.top

>Context-Anchored Tile Refine
https://github.com/Blakeem/ComfyUI-ContextAnchoredTileRefine

>ComfyUI Video Tiler
https://github.com/maDcaDDie2000/comfyui-video-tiler

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE
>>
>mfw Research news

08/06/2026

>When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions
https://arxiv.org/abs/2608.04820

>HelloWorld: Enabling Socially Interactive Characters in Video World Models
https://arxiv.org/abs/2608.05070

>OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
https://arxiv.org/abs/2608.05049

>ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
https://guoxu1233.github.io/ContextMaster

>STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
https://arxiv.org/abs/2608.04887

>Simile Understanding in Text-to-Image Models: An Evaluation Framework
https://arxiv.org/abs/2608.04750

>ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
https://arxiv.org/abs/2608.04436

>CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
https://arxiv.org/abs/2608.04302

>Rethinking Pixel Mean Flows via Interval Denoiser
https://arxiv.org/abs/2608.04818

>Persistent Object Narratives for Token-Efficient Video Language Models
https://arxiv.org/abs/2608.04866

>Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
https://arxiv.org/abs/2608.04349

>Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
https://arxiv.org/abs/2608.04483

>Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models
https://arxiv.org/abs/2608.04454

>When does training on downscaled images yield the same gradients?
https://arxiv.org/abs/2608.04448

>Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection
https://arxiv.org/abs/2608.04935
>>
>shithole general
>>
Good morning.
>>
>what timmy gonna do?
>>
>>
>>
>>
Holy slop.
>>
>>
>gm
>>
>>109483789
Good morning anon.
>>
Gm! Bot status?
>>
>>
>>109484048
you're heck'in cute and valid lumi
>>
>>109483989
Gm
>>
>>
>>
>>
>>109484118
Gm
>>
>>
>>
>>
>>
>>
File: debo_sc_k2_00027_.png (2.71 MB, 1872x1007)
2.71 MB PNG
>>
>>
>>
>>
File: debo_sc_k2_00028_.png (2.29 MB, 1872x1007)
2.29 MB PNG
>>
>>
>>
File: debo_sc_k2_00030_.png (2.05 MB, 1872x1007)
2.05 MB PNG
>>
>>
>>
File: debo_sc_k2_00031_.png (2.1 MB, 1872x1007)
2.1 MB PNG
>>
>>
>>109485030
what is the style on that?
>>
>>
File: debo_sc_k2_00034_.png (2.38 MB, 1872x1007)
2.38 MB PNG
>>109485047
no particular style, just
>action cartoon illustration with clean studio animation lines, bold color blocking
I figure these all are mostly seed dependent and wouldn't be particularly reproducible
>>
>>109485088
it's not "enhanced" by llm?
>>
File: debo_sc_k2_00035_.png (2.36 MB, 1872x1007)
2.36 MB PNG
>>109485097
no, I had to turn off prompt enhancement because it kept getting confused and refusing to process inputs. thats what kept giving me the stock images of people talking
>>
booba physics lol
https://files.catbox.moe/fy1wlp.webm
>>
File: debo_sc_k2_00036_.png (1.95 MB, 1872x1007)
1.95 MB PNG
>>109485118
shes so happy
>>
>>109485109
lel
what node are you using for llm? depending on the node they can be finicky
>>109485118
lel nice
>>
File: debo_sc_k2_00037_.png (2.21 MB, 1872x1007)
2.21 MB PNG
>>109485187
>what node are you using for llm?
apparently its just called 'generate text'. it doesn't even have a model input, so I have no clue what it is. idk where I got it from either. it works ok when it works tho
>>
>>109485203
well there's your problem
get https://github.com/silveroxides/ComfyUI-UtilsCollection
and use
UC_TextEncodeKrea2SystemPrompt
or it may be called 'system prompt encode' now
hook up your wildcard output to prompt, the system prompt, and the clip
>>
File: debo_sc_k2_00039_.png (2.08 MB, 1872x1007)
2.08 MB PNG
>>109485232
cool, I'll try this out
what is the thinking content field for?
>>
>>109485232
that uses the clip (krea uses qwen) as an "llm" so you dont need to load another
i'm sure your node does something similar but it's stupid lel
>>109485245
not sure. also check out the presets (unified presets) from that node collection
>>
>>109485245
maybe you can supply up-front thinking for the 4B lobotomite model lmao
>>
i dont use those text encoder nodes personally (i use gemma4 via llama.cpp node) but you can do some crazy shit with those UtilsCollection nodes
same guy that started the int8 convrot stuff and all the chroma tricks
>>
>>
File: debo_sc_k2_00043_.png (2.23 MB, 1872x1007)
2.23 MB PNG
>>109485305
>i dont use those text encoder nodes personally
its kinda nice to tie together all the wildcards into something more natural. its also good for resolving conflicts that the wildcards might have tossed out. I don't use it much tho desu
>>
>>109485305
i just spend three-hundredths of a cent for deepseek-v4-flash, although i hear that gravy train is coming to an end :(
>>
me in the back
>well this night's gonna get weird

>>109485355
>>109485366
that's what i use gemma for lel
>>
>>109485369
~3,401 gens/$1 ain't too shabby. i could use gemma but the offloading takes longer than the api does.
>>
i just meant i dont use that specific text encoder node with the "built-in" llm

anyway, time to crash
gn all
>>
File: debo_sc_k2_00044_.png (2.26 MB, 1872x1007)
2.26 MB PNG
>>109485366
I guess it never was as cheap as advertised, they were just subsidizing costs to bait people in. same as openai/anthropic

>>109485369
>i use gemma
you run it on a second card or something?
>>
>>109485398
no, offloading and reloading
>>
File: debo_sc_k2_00045_.png (2.52 MB, 1872x1007)
2.52 MB PNG
>>109485395
gn
>>
>>109485398
i think it might just be if you use it directly from alibaba though, openrouter might stay cheap. actually 5.6-luna is on sale for less than deepseek rn i might just use that, it's faster too.

>>109485395
gn
>>
File: debo_sc_k2_00047_.png (2.11 MB, 1872x1007)
2.11 MB PNG
>>109485417
luna has been very good for daily use
but i guess gpt6 is a week a way? early reports say better than mythos (or fable, I forget)
>>
>>109485513
i've been using terra for agent shit, but i don't do a lot of huge codebase shit so my $100 sub stretches an awful long way, plus they reset usage all the fucking time
>>
damn i need to prod it to be more creative with these paper craft things, although this is kind of neat. next one is a few minutes out
https://files.catbox.moe/w20ml1.mp4
>>
File: debo_sc_k2_00048_.png (2.52 MB, 1872x1007)
2.52 MB PNG
>>109485622
maybe "stopmotion animation" would make it do more articulation and movement
>>
>>109485644
yeah i have to tune this thing for h3, it doesn't quite get it
https://files.catbox.moe/yr91i3.mp4



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.