[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109474113

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109475019
animal abuse :(
>>
>>109475019
Any way to make VAE decode faster ?? It took betweek 30 seconds and 1 minute :(
>>
>>109475037
those blacks are working hard. it's not abuse
>>
>>109475043
>Any way to make VAE decode faster ??
good news for you anon
https://github.com/Comfy-Org/ComfyUI/pull/15334
>>
good morning sars, turbo is out? is it good, we back?
>>
>>109475037
>mfw Resource news

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE

>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
https://github.com/yizzz927/CAPE-T2V

>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
https://github.com/jd-opensource/JoyAI-Video-Edit

>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
https://github.com/YangYangGirl/ParVL

>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
https://huggingface.co/JamesZar/OliveGemma-3B

08/04/2026

>stable-diffusion.cpp adds support for MiniMax-H3
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md

>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES time
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
https://github.com/IntMeGroup/MIEScore

>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
https://rathgrith.github.io/PeCA

>Kandinsky WM 1.0: A family of models for Physical AI
https://github.com/kandinskylab/kandinsky-wm

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
>>109475028
typical retard judging a model from only a single output
>Still trying to make a case for wan over h3.
youll learn its pointless to try to beat that kind of thing into anon
if you know its (being any model) is good then you know itll eventually proliferate
>>
>mfw Research news

08/05/2026

>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval
https://arxiv.org/abs/2608.03120

>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Plane
https://arxiv.org/abs/2608.03422

>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
https://arxiv.org/abs/2608.03082

>Can T2I Models Draw from the Right Frame of Reference?
https://arxiv.org/abs/2608.03357

>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0

>Self-Supervised Representation-Guided Generative Dataset Distillation
https://arxiv.org/abs/2608.03218

>Latent Reward Registers for Diffusion Preference Alignment
https://arxiv.org/abs/2608.03929

>UniWorld-Design: From Pixel Generation to Layer-Native Design
https://arxiv.org/abs/2608.03971

>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
https://arxiv.org/abs/2608.03708

>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
https://arxiv.org/abs/2608.03059

>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
https://arxiv.org/abs/2608.03269

>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
https://arxiv.org/abs/2608.03135

>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
https://arxiv.org/abs/2608.03057

>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMs
https://arxiv.org/abs/2608.03112

>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
https://arxiv.org/abs/2608.03875

>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
https://qwen-3d.github.io

>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
https://arxiv.org/abs/2608.03649
>>
>>109475055
saar no cumfy saarport yet
>>
>>109475052
I hope comfy is paying kjGOD well
>>
>>109475055
>>109475066
samefag
>>
how do i merge the lora into the model so i dont have to keep reloading it?
>>
>>109475071
I actually made every post in the past 3 threads, saar.
>>
>>109475019
we are so close to working realtime fmv games.
>>
8/6 and the default 12/3 sigma shift makes the subjects act like crackheads, like completely off the rails adhd mode. together w/ the turbo lora (the HF one) at 8 steps multires and a boomer prompt.
fuck that shit, bypassed.
>>
>>109475066
what's this then? https://huggingface.co/QrusherZA/H3_Turbo_ComfyUI
just tried it, it works, dunno how good tho
>>
>>109475052
>>109475092
>I put swapped the current VAE for this one, but my vids just returned as black
is it because of sageattention?
>>
>>109475090
first I see of this saar, thank you may blue goddess give you many handjob
>>
How make a the sex with the big boobed anime woamn with "H 3" ?
>>
>>109475101
wait until brahmin sir wakes up and spoonfeeds us
>>
Is camera shake broken in Minimax? Whenever i've prompted for an unstable camera/shaky camera, it's more like a vibrating camera with very fast jitter. I just want a camera that's like a person holding a handheld camera. The documentation is useless at providing information on this.
>>
>>109475123
using sigma shift? it seems to speed up certain things
>>
>>109475123
Prompt for Michael J Fox holding the camera.
>>
fucking hate comfyui so much, everytime i have to update for new models my old reliable workflows break and i have to spend an hour fixing it
>>
>>109475123
werks for me
>>
Tom cruise as a vampire goes hard ngl
https://files.catbox.moe/hx5ka2.mp4
>>>/wsg/6208955
>>
>>109475132
>using sigma shift?
Yeah.

I guess I'll try it without sigma.
>>
>>109475135
beg ani to make a new release
>>
Minimax is really Seed dependent. Dont waste your seeds if you got the good one
>>
>>109475153
wrong
>>
>>109475153
probably not correct
>>
Still using sigma shift with the turbo lora?
>>
>>109475052
>vae decoding speedup
wow thanks for saving me the 2s after my 15 minute gen
>>
>>109475147
Buffy is 5' 4"
Tom Cruise is a midget.
>>
>>109475175
Movie magic
>>
>>109475153
even if that's true, that's a good thing. nobody wants a boring model
>>
>>109475174
the cherry on top is more artifacts! :D
>>
>>109475174
Yeah the VAE is super important I don't know why that's the thing people want to butcher for nominal speed increases.
>>
>>109475153
If youre not chase the gap that exists you no longer a racing genner
>>
>>109475193
people are retarded, you have to assume most things posted went through at least 4 different quality raping settings
>>
Turbo lora fucking RAPES the audio
>>
>>109475206
it does, it's an unfinished lora, we have to let the poor lad finish the job
>>
>>109475206
true. but it's WIP, right?
>>
File: debo_sc_k2_00019_.png (2.48 MB, 1872x1007)
2.48 MB PNG
>>
Welp, i tried running H3 Int8 on a 3060 12gb and 24gb of sys memory.
Was paging hard and thrashing my ssd at 64% system memory usage, so not sure if I want to keep using it
>>
bored.
flux3 waiting room
>>
>>109475239
>24gb of sys memory
don't you need 32?
>>
Can i finally make on the spot cnc videos of abi shapiro or is that still a pipedream
>>
>>109475239
I have less system ram than you and have no issues. you must be on winblows.
>>
>>109475244
2x8, 2x4
>>
>>109475244
I was told 640k was enough for everybody.
>>
>>109475244
no, I have a 3060 and 16gb of ram. works fine for me including the non-pruned version.
>>
>>109475256
>16gb of ram. works fine for me
neat.
>>
so no video extension ala ltx yet
that's fine, I can wait
>>
File: 1774541635756140.mp4 (1.11 MB, 608x512)
1.11 MB
1.11 MB MP4
https://files.catbox.moe/f9ncl2.mp4
haha cool
>>
H3 is """"usable"""" on a wide range of hardware. It just depends how low your standards are
>>
>>109475250
Yup, that sounds about right.
Been struggling with comfyui not using system memory and using the pagefile instead for a while now. Thought i fixed the issue with krea2 but now it's happening again...
>>
>>109475239
...why does it use the ssd? my ram is at 80%. is comfy fucking retarded?
>>
>>109475283
I mean its less about standards and more about paitence. its not too bad for me. 10 minutes for 10s? not ideal but hey it works well
>>
pokemonGOD anon, what are your settings?
>>
>>109475256
Oh yup. I bet if you check task manager, the drive with your pagefile (assuming windblows) will be getting thrashed during inference.

If not, I would like to know what black magic you are using to run the int8 model with 16gb.
>>
I feel like people blame comfyui when its likely just windows being a giant steaming piece of shit

>>109475307
>If not, I would like to know what black magic you are using to run the int8 model with 16gb.
no black magic I just use linux. it just works I think mostly due to dynamic vram which everyone shits on here or some reason. I use zram swap
>>
I hope that how you guys pick up girls.
https://files.catbox.moe/kadplw.mp4
>>
>>109475328
>masterpiece
>>
File: i_00029_.png (1.72 MB, 768x1376)
1.72 MB PNG
so would you dudes who tested minimax h3 say that this is good enough to produce real kino?
like do you think this is good enough that people can produce their own short movies, tv shows and such?
also any other AI tools you'd recommend that could help with such?
krea2 model also seems decent enough that it can produce something consistent enough.
>>
>8gb vram
>64gb ddr5
>loonix
Can I run it?
>>
>>109475328
score_9
>>
>>109475322
ok? why can't comfy make it work properly on piece of shit windows then
>>
File: debo_sc_k2_00022_.png (1.83 MB, 1872x1007)
1.83 MB PNG
>>109475301
>...why does it use the ssd?
I'm assuming it needs 32gb to swap the whole model in memory, so 24gb means it has to constantly move layers in and out between ram and disk
>>
>>109475345
>like do you think this is good enough that people can produce their own short movies, tv shows and such?
I can see that yes >>>/wsg/6208850
>>
>>109475354
you answered your own question so why are you asking
>>
File: i_00023_.png (1.16 MB, 768x1376)
1.16 MB PNG
>>109475351
>8gb vram
why?
>>
>>109475322
>just windows being a giant steaming piece of shit
100% Windows fault. it's what made me completely drop it. It just can't handle heavy memory workloads. Like it's broken at it's core.
>>
>>109475358
i mean i assume it doesn't happen on wan2gp
>>
>>109475356
maybe a good technical showcase but painful to watch
>>
File: file.png (23 KB, 250x213)
23 KB PNG
Idk if it's something I'm doing wrong but this node rapes FLF generations.
>>
>>109475365
this is what happens when the entire OS is scaffolded off of code written in the 90s
>>
>>109475359
laptop, please understand saar
>>
>>109475365
windows is made for normies. it doesn't need to be anything else other than that.
>>109475368
blame whatever you want. the model, comfy, or anything else other than the root cause of the issue. no skin of my back.
>>
>>109475372
it does, I prefer to wait a bit more and use Spectrum but the quality is there >>109474992
>>
>>109475379
off*
anyways gn saars
>>
>>109475368
wangp is more comfy than comfy tbqhdesu
>>
vid gen has made me realize i am not very creative
>>
>>109475390
fact
>>
some say the text encoder being censored or uncensored doesn't matter-- are there any comparisons? Idk what to believe
>>
>>109475351
Maybe you can but it's not gonna be worth it. IMO genning is only fun if it's fast. Otherwise it's a slog and a waste of time.
>>
>>109475403
yes has been done to death at this point man...
>>
>>109475396
im just making my ocs moan n shit
>>
>>109475403
"Uncensored" text encoders are a meme and do literally nothing, has always been the case, stop giving attention to retards.
>>
>>109475422
proof?
>>
File: 1773019187552274.mp4 (581 KB, 640x480)
581 KB
581 KB MP4
>>
It doesn't know that a penis is supposed to be rigid and instead makes it move like it's some kind of sea cucumber.
>>
>>109475430
for my eyes only, chuddie
>>
>>109475429
I'm something of an uncensored text encoder my myself, reggin.
>>
>>109475396
we got so used on using shit models that can only do 1girl that we don't know what to do anymore when we finally get a model that can do anything



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.