[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 37o4u.gif (1.34 MB, 192x192)
1.34 MB GIF
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109478939

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
First for based Will Smith gen
>>
>mfw Resource news

08/06/2026

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view

>OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
https://xin1u.github.io/OminiVR_PAGE

>DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
https://github.com/Zhong-Chenchen/DIVE.git

>Multi-View Face and Gesture Animation with Dynamic Gaussians
https://dfki-av.github.io/MVFGA

>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
https://empaava.top

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE

>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
https://github.com/yizzz927/CAPE-T2V

>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
https://github.com/jd-opensource/JoyAI-Video-Edit

>ParVL: Parallel Scaling and Expandable Compute Allocation for MLLMs
https://github.com/YangYangGirl/ParVL

>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
https://huggingface.co/JamesZar/OliveGemma-3B

08/04/2026

>stable-diffusion.cpp adds support for MiniMax-H3
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md
>>
>>109480220
kino
>>
based
>>
>mfw Research news

08/06/2026

>When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions
https://arxiv.org/abs/2608.04820

>HelloWorld: Enabling Socially Interactive Characters in Video World Models
https://arxiv.org/abs/2608.05070

>OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
https://arxiv.org/abs/2608.05049

>ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
https://guoxu1233.github.io/ContextMaster

>STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
https://arxiv.org/abs/2608.04887

>Simile Understanding in Text-to-Image Models: An Evaluation Framework
https://arxiv.org/abs/2608.04750

>ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
https://arxiv.org/abs/2608.04436

>CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
https://arxiv.org/abs/2608.04302

>Rethinking Pixel Mean Flows via Interval Denoiser
https://arxiv.org/abs/2608.04818

>Persistent Object Narratives for Token-Efficient Video Language Models
https://arxiv.org/abs/2608.04866

>Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models
https://arxiv.org/abs/2608.04349

>Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
https://arxiv.org/abs/2608.04483

>Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models
https://arxiv.org/abs/2608.04454

>When does training on downscaled images yield the same gradients?
https://arxiv.org/abs/2608.04448

>Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection
https://arxiv.org/abs/2608.04935
>>
>>109480214

https://old.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/
>>
>>109480228
finally
>>
>>109480247
kek
>>
>>109480236
>>109480242
thanks!

>>109480230
kys
>>
>literal tranitor bake
looooole
>>
theres way too much coomer stuff in the data set. random penis in the video., what th heck?
>>
File: download.jpg (46 KB, 600x603)
46 KB JPG
i hope you all behave yourselves
>>
https://streamable.com/idwah1
>>
I don't get it, Gen works great at 0.4mp, bump to 0.7, prompt adherence jumps out the window.
>>
>>109480220

Real thread
>>109480276
>>109480276
>>109480276
>>
>>109480259
Please don't dig into him too hard, if you do he'll start crying and pissing himself
>>
File: 1754633788891113.png (1.9 MB, 2474x1448)
1.9 MB PNG
>>109480236
>EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot
>>
>>109480280
> trolling outside of /b/
>>
File: gegy].png (1.17 MB, 864x1184)
1.17 MB PNG
nah JSID bruh its over
>>
>>109480287
>ID 190, how are you today?
>i am very good sir, thank you for asking. and how are you today?
>i'm great, so i was wonde-
>you can open bobs?
>excuse me?
>show milk! open bobs!
>>
>>109480220
>>109480236
>>109480242
Fuck off debo
>>
>>109480276
>>109480276
>>109480276
>>
Ignore debo thread.

Everyone here:

>>109480276
>>109480276
>>109480276
>>
non-debo bread
>>109480276
>>109480276
>>109480276
>>
File: 45432.webm (3.77 MB, 320x576)
3.77 MB
3.77 MB WEBM
the turbo lora makes better audio for war kinos surprisingly. still can't make alarm sounds sadly. something is wrong with the h3 dataset
>>
>>109480378
>still can't make alarm sounds sadly.
such a strange limitation. There is so much free stock audio of alarms
>>
who is ready for flux3?
>>
>>109480392
i might have to use the video2audio feature with ltx since it makes really nice realistic alarms and other combat stuff
>>
>>109480412
isn't there foley models that can make cool sound effects or do you just want to skip over editing?
>>
Was there some prompt guide for H3?
>>
>>109480276
>>109480276
>>109480276
>>
>>109480378
The parachute was pretty bad lol, I think they only trained it on gliders.
>>
>>109480436
i don't want to put that much work into it. if you want it to be realistic, then you have to do a lot of editing on the sounds themselves so they match the acoustics of the video
>>
>>109480455
what do you mean? the forward momentum would be from ejecting out of a fast moving jet. i do agree that it isn't really realistic if you see how a real ejection seat works. i tried to prompt something similar but it kept glitching out
>>
>>109480244
Someone go to reddit and tell me if they say anything interesting.
>>
>>109480550
we got:
>thank you cocksucking
>cumfart cocksucking
>industry plugs
nothing of note
>>
the other thread feels like a bot that loops through the same convo
>>
File: debo-no.png (1.5 MB, 1056x992)
1.5 MB PNG
>>109480327
>>
>>109480622
What one?
>>
>>109480634
based
>>
File: debo_ac_k2_00055_.png (2.26 MB, 1872x1007)
2.26 MB PNG
>>109480647
I couldn't quite get the eyes right and gave up lol. close enough
>>
>>109480657
have you still never got into vidgen?
>>
File: debo_ac_k2_00056_.png (2.37 MB, 1872x1007)
2.37 MB PNG
>>109480729
I might play around with h3 when a decent turbo lora drops. I have a 4070 w/ 32gb ram so I'm just barely on the edge of workable hardware
>>
>>109480763
that should be enough with the 8int diffusion and text encoder if you enable disk memory
>>
>>109480763
just leave genning and go goon or something.
Once I see h3 got the idea semi-right, queue like 10 clips and fuck off.

when you're back, it's done.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.