[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109881361

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
gm saars
>>
>mfw Resource news

09/22/2026

>CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation
https://zshyang.github.io/CoaG

>AniPrO: Interpretable Anime Image Provenance Detection via Multi-Dimensional Semantic Reasoning
https://github.com/YAN-LIU05/AniPrO

>AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport
https://github.com/51xOne/Alignmorph

>Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
https://sunyangtian.github.io/Mira-Scene-web

>Rethinking Vision Architectures with Gated Linear Attention and KAN
https://github.com/mehizelali/linear-kan-transformer

>SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models
https://github.com/aliathar1401/SK-Stars-shroom-visions-2026

>Accurate Motion Estimation with Bézier Control Point for Efficient Frame Interpolation
https://github.com/SHH-Han/ABC-Inter

>Qwen Image 2.1 for intel Mac's with AMD GPU
https://github.com/haseebeqx/qwen-image-intel-mac

>China's Alibaba unveils new powerful chip and ambitious AI model plans
https://tech.yahoo.com/ai/articles/chinas-alibaba-unveils-powerful-chip-071601752.html

>Photoshoot: Build a person once, then shoot a whole series
https://github.com/ralksta/ComfyUI-Photoshoot

09/21/2026

>Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
https://qwen.ai/blog?id=qwen-image-2.1

>Qwen Image 2.1 Prompt Enhancer — ComfyUI Custom Node
https://github.com/benjiyaya/ComfyUI-Qwen-Image-2.1-Prompt-Enhancer

>Qwen-Image 2.1: ComfyUI Repack
https://huggingface.co/Comfy-Org/Qwen-Image-2.1

>Qwen-Image-2.1 GGUF quantized files
https://huggingface.co/leejet/Qwen-Image-2.1-GGUF

>Spectrum for Qwen2.1 (>2x Speedup on a 3060 12gb)
https://github.com/awdqwdasdg/Comfyui-Spectrum-Qwen2.1

>Supra2-IMG: Text-To-Image • 100M Parameters
https://huggingface.co/SupraLabs/Supra2-IMG
>>
>mfw Research news

09/22/2026

>Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance
https://arxiv.org/abs/2609.24215

>IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
https://arxiv.org/abs/2609.24228

>An Efficient and Effective Watermarking Scheme for the Protection of the Intellectual Property Rights of Video Generative Models
https://arxiv.org/abs/2609.23586

>VISTA: Video-Injected Stylized Text-to-Animation
https://arxiv.org/abs/2609.23817

>SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265 $ Single-GPU Acceleration of Visual Generation
https://arxiv.org/abs/2609.23153

>PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining
https://arxiv.org/abs/2609.22789

>RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
https://arxiv.org/abs/2609.22947

>Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
https://arxiv.org/abs/2609.23658

>PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models
https://arxiv.org/abs/2609.23600

>Rethinking Diffusion Segmentation: When Does It Rely on Its Noisy State, and Does Diffusion Matter?
https://arxiv.org/abs/2609.23967

>Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation
https://arxiv.org/abs/2609.22916

>Hierarchical Prompt Learning for Hyperbolic Vision-Language Models
https://arxiv.org/abs/2609.24276

>Classifier-Free Guidance in Flow Matching: Non-Autonomous Potentials, Overshoot, and Posterior-Mean Control
https://arxiv.org/abs/2609.24287

>LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba
https://arxiv.org/abs/2609.24337
>>
>sdg: schizo image gens
>ldg: schizos arguing in text
>>
File: debo_bf_k2_00003_.png (2.68 MB, 1664x1069)
2.68 MB PNG
>>109886072
>>
>>109885340
>>109885587
still hoping for answers on this
>>
>>109883871
This is great, probably the last "free lunch" we're gonna see from H3. Personally this will cut 40 seconds from each of my gens. It makes me excited enough to try out some 1366x768 stuff again tomorrow

>>109886017
>japs are fucked in the head
There are 300,000 self hating Japanese in Japan that hate themselves because they can't understand and come to terms with why their wartime atrocities were so much more fucked up than anyone else's. Like of course they deny shoving bamboo sticks up Chinese 3 year olds to dilate them so the Japanese army could rape them, who wouldn't? And it's ended up as a unique self-hating ideology, similar but distinct to Brazil's mongrel complex

But to Japan's credit they seem to only rape others mostly (yeah that gang raped for 34 days and murdered girl but that happened like once ever) and their crime statistics are enviable (yeah I know it's weird with most crimes being solved by confession etc but still) so maybe we're the fucked in the head ones where we obviously can't handle some forms of artistic expression but they can


This is related to local diffusion models because Comfy Inc is in Japan
>>
What would anon wish for me to generate?
>>
what is the current status on video models?

t. hasn't been genning in almost 6 months.
>>
>>109886086
Thoughts on qwen image 2.1 deeby? (Sorry if you already voiced them)
>>
>>109886112
kinos
>>
>>109886113
Minimax H3 is better than Sora 2. Hopefully you have at least 16gb of vram and 64gb of ram
>>109886112
>What would anon wish for me to generate
Non-sexual but interesting POV stuff, like working on fixing a satellite on a space station while in orbit, or being one of the Muslims walking around the black cube
>>
File: debo_bf_k2_00005_.png (2.62 MB, 1664x1069)
2.62 MB PNG
>>109886114
I haven't touched it personally. looks like a great edit model based on the stuff people have been posting though
>>
>>109886120
>Minimax H3 is better than Sora 2.
fucking seriously? how? I have 24 GB of VRAM and 128 GB of ram.
>>
>>109886133
unironically yes
>>
>>109886087
Set seed to fixed. Change one thing in the prompt at a time so you can see how it changes. Change scheduler and sampler individually to see how those affect the output. Each model has its quirks so you need to isolate the variables by using a fixed seed until you know approximately what prompts work and which ones don't.
>>
File: MiniMax_H3__00787-small.mp4 (3.76 MB, 1024x1024)
3.76 MB
3.76 MB MP4
>>109886120
sorry best i can do is 1girl
>>
Someone else bake before the schizo baker gets banned again
>>
File: may_00007_.webm (1.42 MB, 928x672)
1.42 MB
1.42 MB WEBM
>>
>>109885630
I don't think anon realizes how good that specific gen is.

>>109886174
The eyes are really pleasant.
>>
>>109886141
right but as far as the literal prompting how verbose does h3 need to really be? does it need to be as literal and autistically crafted with an llm like krea and qwen or does broad stuff work? to me its felt like a coinflip
>>
>>109886212
I don't do t2i h3. too weird and variable. Always first frame ref image. Much easier to control.
>>
>>109886133
>fucking seriously? how?
They generated videos filling in the knowledge gaps with Seedance 2.5 and cloud video models and it mostly worked out.
>>109886133
>I have 24 GB of VRAM and 128 GB of ram.
You're in for a treat then. Text to video is actually good now. Use the FL2VA model without any image references for text to video. Don't expect faces to look good at resolutions smaller than 720p though
>>109886148
Wtf bruh why even ask then. At least do pov featuring a 1girl you should still get turned on by that

>>109886212
>right but as far as the literal prompting how verbose does h3 need to really be
You're done with manual prompting. You are supposed to feed the official prompt guide (not the one with "ref" in the name, that's for the reference model, the other one) to an AI and ask it to make T2VA prompts featuring xyz that follow the structure
>>
>>109886057
>>109886064
thanks!
>>
>>109886236
>>109886242
im not doing t2v for h3, im doing ref2v and its still a fucking mess
>>
>>109886278
ref2v is slow and wonky. Only advice I can give is to put a preview node in your workflow so you can abort failed gens early. Give it 3-4 iterations and if it hasn't unfucked itself, abort. Use Model Preview Override node in the KJNode pack.
>>
>>109886278
>im not doing t2v for h3, im doing ref2v
Ok stop doing that. Pretty sure using fl2va with the ref Lora from kijai is better anyways
>>
>>
File: infinite-truck.gif (1.86 MB, 498x274)
1.86 MB GIF
>>109886331
>>
Thoughts?

https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/wip
>>
The latest benchmark report for MiniConstruct H3 story authoring:

Gemma behaved almost exactly as before: mystery passed, quiet passed, interpersonal failed on the same structured-output family of problem. This time the invalid value was again:
tension: "low"

The prose itself was reasonable. So I would treat that as confirming our conclusion: the remaining Gemma failure is a Tone/Performance enum-compliance issue, not a richer-creativeRequest problem.
Qwen is more interesting. It went 3/3 structurally valid, and its interpersonal story is substantially stronger than Gemma's: the disagreement actually becomes personal, the sisters trade meaningful dialogue, moving the radio reveals the father's station notes, and the discovery visibly resolves the conflict. That's very good evidence that the Story policy generalizes beyond Gemma.

But Qwen also demonstrates the opposite weakness: it is far more verbose than it needs to be.

Case Gemma total tokens Qwen total tokens Qwen request time
Mystery 7,401 11,980 ~14.4 min
Interpersonal failed before metrics 11,133 ~12.8 min
Quiet 7,045 9,088 ~9.2 min

For mystery, Qwen used about 2.4× Gemma's completion tokens and took about 2.4× as long. For quiet it used about 1.7× the completion tokens and took about 1.74× as long.
And some of that verbosity is genuinely too much for a six-second Generation. Qwen's mystery G2, for example, has her checking locker numbers, opening one, examining a ticket, turning it over, evaluating the handwriting and date, putting it back, closing the locker, then discovering a hidden scrawl. That's good storytelling, but realistically too much visual action for six seconds.

So I would characterize them like this:
Gemma: more economical, generally good clip-sized decomposition, but unreliable about strict enums.
Qwen: stronger connective storytelling and excellent interpersonal reasoning, but prone to over-authoring and much heavier reasoning/output use.
>>
File: lol who cares.png (103 KB, 336x359)
103 KB PNG
>>109886381



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.