[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109902994

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3
https://neta.art/use-cases/en/h3-1000-prompt-list

>Anima
https://huggingface.co/circlestone-labs/Anima
https://animastyles.thetacursed.com
https://tagexplorer.github.io/
https://animadex.net

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
it's so over
>>
>>109908496
Damn, still baking schizo rentries. Anons need to try harder to get the schizo to kill herself
>>
>>109908496
Yuck
>>
>mfw Resource news

09/25/2026

>Fizgig H3 Tweaks: Training-free tweaks for MiniMax H3 in ComfyUI
https://github.com/shootthesound/ComfyUI-Fizgig-H3-Tweaks

>Krea 2 inpaint edit
https://huggingface.co/Cierpliwy/krea2-inpaint-edit

>Pruna-Qwen-Image-2.1: Few-step LoRA adapters for Qwen-Image-2.1
https://huggingface.co/PrunaAI/Pruna-Qwen-Image-2.1

>AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
https://github.com/zhiyuxu03/AV-GRPO

>Kijai Qwen Image 2.1 Fun_controlnet_union (BF16/int8_convrot)
https://huggingface.co/Kijai/QwenImage_experimental/tree/main/model_patches

09/24/2026

>Making the MiniMax H3 Video VAE 2x Faster
https://blog.comfy.org/p/making-the-minimax-h3-video-vae-2x

>Qwen-Image-2.1 Text Encoder (Heretic) — GGUF · FP8 · bf16
https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF

>unsloth/Qwen-Image-2.1-GGUF
https://huggingface.co/unsloth/Qwen-Image-2.1-GGUF

>Ming-Image-0.1-Design GGUF
https://huggingface.co/realrebelai/Ming-Image_GGUFs

>MiMo-V2.6 series: Frontier intelligence, all the modalities, built in public
https://mimo.xiaomi.com/mimo-v2-6

>ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP
https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP

>Qwen-Image-2.1-Fun-Controlnet-Union
https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union

>Latent evolving World Action Model
https://github.com/XuejiFang/LeWAM

>Prompt Studio: Multimodal prompt studio for ComfyUI
https://github.com/tngklp/ComfyUI-Prompt-Studio

09/23/2026

>Qwen-Image-2.1-viggle-turbo — v0.2 (preview)
https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

>Ming-Image-0.1-Design: 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs
https://huggingface.co/inclusionAI/Ming-Image-0.1-Design

>MiniMax-H3-Fun-Controlnet-Union-2.0
https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union-2.0
>>
Blessed thread of frenship
>>
>mfw Research news

09/25/2026

>ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
https://arxiv.org/abs/2609.28923

>WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
https://arxiv.org/abs/2609.30221

>ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
https://arxiv.org/abs/2609.29225

>SALI: Shot-Aware Late Interaction for Cross-Shot Relation Matching in Text-to-Video Retrieval using Film-Grammar Knowledge
https://arxiv.org/abs/2609.29721

>OmniFabric: Coherent UV Space Texture Synthesis for 3D Garment Reconstruction
https://humansensinglab.github.io/OmniFabric

>CARE: Condition-Aware Representation Regularization for Diffusion Models
https://arxiv.org/abs/2609.28561

>EIB-Net: Entropy-Guided Information Bottleneck for Generalizable AI-Generated Image Detection
https://arxiv.org/abs/2609.29064

>Accelerating Video Diffusion via Training-Free Trajectory Routing
https://arxiv.org/abs/2609.30096

>TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution
https://arxiv.org/abs/2609.29240

>Spectral Amplitude Purification in Distribution Matching for Diffusion Distillation
https://arxiv.org/abs/2609.29116

>Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models
https://arxiv.org/abs/2609.29358

>Training Object Permanence in World Models
https://object-permanence.world

>CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
https://arxiv.org/abs/2609.28813

>Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
https://shamanthak-hegde.github.io/where-hallucinations-live

>Pistis Technical Report
https://arxiv.org/abs/2609.28554

>JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
https://arxiv.org/abs/2606.14777
>>
>>109908533
Friends don't let you schizo out in the OP
>>
>>109908533
nigger
>>
>>109908496
Thank you for baking this thread, anon
>>109908533
Thank you for blessing this thread, anon
>>
>>109908550
We already know he is one
>>
File: gun and trench.jpg (617 KB, 1024x1024)
617 KB JPG
>>
File: 6553.png (2.54 MB, 1024x1568)
2.54 MB PNG
>>109908512
its just getting started
>>
need kino
>>
Thanks so much to catjack specifically for making this a trans safe space! You are valid sister!
>>
lilbro is crashing out over the bake *skull*
>>
>>109908585
The OP has had schizo coped baked in since the beginning
>>
>>109908604
Catjack doesn't know how filters work or how ignoring people works so she has a shitty tantrum and ruins everything. It's the only logical endgame for /ldg/ and catjack should just stay in the containment thread /sdg/ for namefags like him
>>
i think diffusion models are slop and can never not be slop
i think ultimately the role of diffusion will be a rendering layer for llm's to orchestrate
but on its own the diffusion model is just not capable of producing kino
>>
>>109908661
good thing we pivoted to flow based models, then
>>
>>109908661
very deep anon, I came twice
>>
>>109908672
qrd
>>
>>109908672
flow matching is also slop
more specifically the ways that we interface with these models is slop
text, image, control etc
all these methods are bad and awful
to truly ascend you need and entity can just play the latent space like a musician plays and instrument
we will never be able to do this
>>
>>109908585
It's funny that he replied to your post continuing to crash out keeeeeek
>>
Every time someone says "ai will never do X" some months later it starts doing X. Rookie thinking.
>>
i stayed up all night making kinos again
>>
>>109908648
I thought troonjack started /ldg/.
>>
File: Allears.jpg (131 KB, 573x865)
131 KB JPG
>>109908714
Well anon, I'm all ears
>>
gettin real sick of your shit around here
>>
>>109908687
play me like one of your french trumpets
>>
>>109908773
kek
>>
>>109908698
AI you will never be a woman
>>
>>109908801
AI will never be my girlfriend
>>
krea's censorship is starting to annoy me, even more sfw but ecchi concepts like holding a cup between breasts it will just refuse
Is there a best recommend lora that can fix that while not impacting the intelligence too much and not just a straight up porn lora that makes everything nude
>>
>>109908808
the textfusion refusal lora but just use qwen instead
>>
>>109908773
i don't get it
>>
>>109908760
No, anon likes to make stuff up to trick newfrens.
>>
>>109908828
lol what a prankster :P
>>
File: 66444.jpg (26 KB, 289x403)
26 KB JPG
>>109908808
Skeelshoe?
>>
File: schizo1.png (1.79 MB, 1254x1254)
1.79 MB PNG
every schizo argument requires multiple sides
>>
File: ErnieBase_Output_37373.jpg (3.17 MB, 1344x1728)
3.17 MB JPG
TIL Ernie Base is a gorillion times better than the Turbo version
>>
>>109908922
>Ernie
get a load of this oldfag
>>
File: 6.png (2.03 MB, 832x1056)
2.03 MB PNG
>>
>>109908937
Kek nice
>>
cozy breas



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.