[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1769747953087517.webm (3.16 MB, 2048x826)
3.16 MB
3.16 MB WEBM
Previous: >>109540195

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
balloon boobs
>>
File: 11111.mp4 (2.02 MB, 1056x608)
2.02 MB
2.02 MB MP4
>>
blessed thread of frenship
>>
animators status?
>>
>>109542154
RAPED
>>
Can someone contact minimaxAI and tell them to train their image model on danbooru? Better yet, make an anime finetune for it in like a day or two.
>>
>>109542150
Nah, 10s is plenty, 12s is great, 15s is perfect. 20 is too much right now
>>
>>109542158
It can do anime just fine?
>>
Now that the dust settled, has anyone else started feeling H3’s 10-15s length limit being just too short?
LTX spoiled me with 25s
>>
>>109542171
>>109542159
>>
>>109542171
>>109542159
>>
File: H3_noAudio__00007_.mp4 (3.89 MB, 1664x1216)
3.89 MB
3.89 MB MP4
https://litter.catbox.moe/fty2n9bs1c6d06hl.mp4
>>
>>109542159
What are you generally genning?
>>
>>109542154
Im become jobless soon. What should i do ?
>>
idk why I'm getting filtered so hard by this gen.
>>
>>109542144
>mfw Resource news

08/12/2026

>LTX-2.5 22B IC-LoRA Pixel Spatial Upscaler
https://huggingface.co/Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler

>LTX-2.5 22B Distilled — NVFP4, ComfyUI-ready
https://huggingface.co/BennyDaBall/LTX-2.5-22b-distilled-nvfp4-comfy

>LTX-2.5 22B — GGUF
https://huggingface.co/realrebelai/LTX-2.5_GGUFs

>ComfyUI NVIDIA RTX VSR Pro
https://github.com/whmc76/ComfyUI-NVIDIA-RTX-VSR-Pro

>Stable Layers: Decomposing Images into Editable RGBA Layers
https://huggingface.co/StabilityLabs/Stable-Layers

>PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders
https://github.com/manmanTAT/PEAK

>Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
https://github.com/aiimaginglab/PCFlow

>MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding
https://shuaiwang97.github.io/MMArt

>ComfyUI H3 Studio: Video editor for MiniMax H3 inside a single ComfyUI node
https://github.com/shootthesound/ComfyUI-H3Studio

>MINIMAX H3 Prompt Studio: Build structured MiniMax H3 video-generation prompts
https://github.com/lololerigolo60/Minimax-H3-prompt-studio/tree/main

>ComfyUI Image Conveyor v1.4 — now with MiniMax H3 multi-reference support
https://github.com/xmarre/ComfyUI-Image-Conveyor/releases/tag/v1.4.0

>VPIPE: Real-time multimodal AI pipelines on Apple Silicon
https://github.com/tgo-app-dev/vpipe

08/11/2026

>LTX-2.5
https://ltx.io/model/ltx-2-5

>LIGHTX2V v1.0 4-step/8-step Turbo Minimax H3 loras
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

>ComfyUI-DoRA-Dynamic-LoRA-Loader
https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader

>DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models
https://github.com/chenhsiu48/DeepFreqMark

>Staying True to the Origin: Continuous Image Stylization with Smooth Transitions
https://reychiaro.github.io/StyleController
>>
>>109542171
you can go further with h3. i haven't tested what length it starts to fall apart since i don't have enough vram
>>
>mfw Research news

08/12/2026

>UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
https://research.nvidia.com/labs/par/uniprobe

>SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis
https://arxiv.org/abs/2608.10519

>Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning
https://arxiv.org/abs/2608.11013

>Beyond Pixels: From Video Priors to 4D Worlds
https://hayd-zju.github.io/Beyond-Pixels

>Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation
https://joseph-lin-tech.github.io/BridgeEventDiT-VFI

>Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation
https://arxiv.org/abs/2608.10439

>NullEdit: Stealthy Image Protection via VLM Condition Redirection
https://arxiv.org/abs/2608.10870

>Where To Look? : Causal Tracing of Vision Encoders in VLM
https://arxiv.org/abs/2608.10758

>Rethinking Text-Based Image Retrieval in Specific Domain
https://arxiv.org/abs/2608.10524

>Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers
https://arxiv.org/abs/2608.10989

>Human versus Computer Vision
https://arxiv.org/abs/2608.10181

>Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
https://arxiv.org/abs/2608.10525

>When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models
https://arxiv.org/abs/2608.11024

>Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs
https://arxiv.org/abs/2608.10959

>Meshy T2: Fast Native Mesh Generation with Flow Matching
https://arxiv.org/abs/2607.28675

>FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing
https://arxiv.org/abs/2509.23452
>>
>>109542171
I can 20 secs with 8step turbo lora. It took 15 minutes though
>>
Gotta try how it looks if I spin a character around with just frontal view ref.
My imagen model doesn't have that many face shots from multiple angles, at least bot ones that look alike
>>
>>109542200
>>109542194
Not with any decent res.
>>
>>109542188
porn
>>
>>109542200
resolution and gpu?
>>
File: 1448800539597.png (660 KB, 1106x1012)
660 KB PNG
I'm just generating sexy anime videos.
>>
>>109542211
>>109542216

20 secs 0.7mp + RTX upscale + Lightx2v 8step turbo lora. took 15 minutes on my 5070ti
>>
>>109542220
just as god intended
>>
most of the game is about genning spank bank material
the true endgame, however, is turning my LLM ERP's into movies/OVA's. with lots of sex in them obviously.
>>
File: 993713440631932.mp4 (3.64 MB, 992x736)
3.64 MB
3.64 MB MP4
>>
>>109542223
my 5080 gpu farts and freezes when I go for 20s 0.7mp 8 steps (128 ram)
maybe one of the cope nodes is bad, could you share the wf?
>>
>>109542162
it can't do stuff that finetuned anime model would.
>>
>>109542235
I only use sage attention
>>
>>109542158
>>109542239
>Can someone contact minimaxAI and tell them to train their image model on danbooru?
Is this just referring to H3? Is there some sort of image workflow for it?
>>
>>109542241
startup flag or the node? all up to date?
>>
why do a 20s gen when you can do 2 10s gens
>>
My uncles boyfriend told me that my uncle, a researcher at Lightricks, participated in a company-wide suicide pact yesterday.
>>
>>109542251
one long video with a story line is more kino than multiple short ones
>>
>he can't tell a story across two gens
ngmi
>>
is it able to not sound like sand paper whenever there's fucking?
>>
Why does the shit from civitai generator look way better than what I do locally with the same models/loras? What's their secret sauce?
>>
>>109542263
describe the sound as bones breaking no im not kidding
>>
>>109542261
They think they're going to destroy Hollywood but can't even make two gens. Rough.
>>
which turbo lora do I use for H3?
there are like 20
>>
File: h3-1786592334411 (1).mp4 (892 KB, 576x768)
892 KB
892 KB MP4
im glad to be here with all of you
who remembers 2023 when todd was letting us use his model for free chats on the tavern
>>
File: h3_00001.webm (942 KB, 1192x880)
942 KB
942 KB WEBM
>>
>>109542278
Todd was the best model I’ve ever tried, still no clue what it was though
>>
>>109542260
do you think when they make a movie, they film the whole thing in one take?
>>
>>109542276
Theres only two of them Larry and Lightx2v
>>
File: 1758323099012685.jpg (328 KB, 2560x1286)
328 KB JPG
crazy how clankers can make you basic ass video editors in fucking html

https://files.catbox.moe/oymtau.webm
>>
>>109542269
I went hunting on civ, the best sounding videos had this
>overall_soundscape:
>Slimy squelching, wet fleshy slaps, soft female moans.

otherwise I'll try
>loud sound of bones breaking, with wet sounds of blood gushing and gore splattering
>>
>>109542288
depends on the movie
https://www.youtube.com/watch?v=ucspfmRM7vI&list=PLb4iXJOYkKszmT5dwMzuf38hzCxeqjCgJ&index=6
>>
>>109542297
good style, good gen.
>>
>>109542288
the best ones do
>>
>>109542283
What an evil wizard
>>
>>109542297
It’s actually kind of creepy how many mundane tasks I’ve offloaded to llms. I don’t have to think that far back to when I’d have to convince myself not to spend a hundred or so bucks on software I’m sure I absolutely need. Now it’s totally possible to ask an llm to make a basic functional piece of software and to do the exact job you need and then forget about it entirely. I think once it fully sinks in that a majority of software can be completely replaced by a text prompt and a few minutes, there will be some upheaval.
>>
>>109542288
Russian Ark did it
>>
>>109542297
?????
>>
>>109542297
>>109542315
Can you guys recommend local prompt builder for H3? I’ve been writing manually

Also do local models have vision? need for r2v
>>
>>109542239
which are?
>>
File: H3_noAudio__00008_.mp4 (2.21 MB, 1248x896)
2.21 MB
2.21 MB MP4
>>
File: 772663288412600.mp4 (3.61 MB, 832x640)
3.61 MB
3.61 MB MP4
>>109542339
nice lol
>>
Does Torch compile work with H3 ?
>>
>overall_soundscape:
>Slimy squelching, wet fleshy slaps, soft female moans.
>loud sound of bones breaking, with wet sounds of blood gushing and gore splattering

https://files.catbox.moe/hdvgue.mp4

I think my model is cursed.
>>
File: 1761327137413319.mp4 (946 KB, 640x640)
946 KB
946 KB MP4
>>
File: iChads.jpg (225 KB, 1846x648)
225 KB JPG
>>109542235
>>109542223
it'll be slower but I'll try on the mac ultras and see if I can do full res over 20 seconds. I have the vram
>>
>>109542246
https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/comment/p29u3x8/
>Regarding single-frame image generation, we are deriving a dedicated image model from a common ancestor in the H3 model lineage, and we expect to make it available to the community.

It will use the same VAE encoder as H3. Since the temporal encoder is causal, we can obtain a 2D VAE encoder through weight slicing. We also plan to provide a dedicated VAE decoder specifically designed for image generation.
>>109542335
most anime concepts? 99% of anime characters you can find on danbooru? 99.9% of artist artstyles? Come on anon.
>>
>>109542401
the VAE is ass though
>>
>>109542365
why this looks so sloppa
>>
>>109542412
the concept and its execution are so good that i can forgive the slopped realism
>>
>>109542339
oddly enough I have trouble getting krea and h3 to do dark rooms that dont have an imaginary studio light
>>
>>109542412
he probably used a fried image for it
>>
>>109542412
has turbo lora stink on it
>>
File: 1782085692555059.gif (473 KB, 640x640)
473 KB GIF
This is the endgame of Local Video gen right because i cant see this get improved much more. And theres no way we can afford a new GPU with the current prices
>>
>>109542334
>Also do local models have vision?
yes. most do. and it's typically in the info exposed on either the UI/CLI or on hf or wherever you get your models.

> Can you guys recommend local prompt builder for H3?
See if you're already running a decently powerful LLM with vision then you can actually hand it the prompting guide https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md and tell it to write a prompt where x y happens with z.
>>
yeah ok. I'm pretty convinced we'll need a sex sounds lora.
>>
>>109542449
>can’t see this get improved
I agree but I want more than 15 seconds of slop
>>
>the original DiT paper was published four years ago
>the first model to use this new arc was pixart alpha
>only recently has DiT became the norm for the average user
Incredible



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.