[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: renak1-ns.webm (2.12 MB, 640x992)
2.12 MB
2.12 MB WEBM
Discussion and Development of Local Image, Video, and Music Models and Software

Previous: >>109510784

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Real bread
>>
>mfw Resource news

08/09/2026

>Kroma v0.2 — Krea 2 fine-tune (full model)
https://huggingface.co/lodestones/Kroma

>krea2-turbo-bbox
https://huggingface.co/jimmycarter/krea2-turbo-bbox

>Kroma v0.2 Quant
https://huggingface.co/silveroxides/Kroma-Quant/tree/main

>Spectrum for Ideogram 4
https://github.com/Nif00/ComfyUI-Spectrum-Ideogram4

>ClipProj — MiniMax H3 conditioning from a Qwen3-VL-4B
https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3

>ComfyUI-SigmaSync-LoRA: Sigma-aware model-only LoRA strength scheduling
https://github.com/capitan01R/ComfyUI-SigmaSync-LoRA

>NexusBTA v0.2.44 adds MiniMax H3 support
https://github.com/JpAndreBTA/Nexus-BTA/releases/tag/v0.2.44

>Experimental MiniMax H3 single-image VAE
https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE

>MiniMax H3 REF2VA w4a8
https://huggingface.co/realrebelai/Rebels_w4a8s

08/08/2026

>Kijai: MiniMax H3 Ref Lora Rank 256 bf16
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)
https://github.com/Amduraznak/minimax-h3-fp16-fix

>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopoints
https://github.com/fengwang/Cosmos3-Nano-WebUI

>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patches
https://github.com/charlie12345/R9700AIProComfyUIPatch

>MiniMax-H3-Pruned-GGUF
https://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
>>
I'm a bit confused anon, so which model is the best between t2v/i2v and ref2v for h3?
>>
File: 1774816237429770.webm (2.18 MB, 1184x896)
2.18 MB
2.18 MB WEBM
>>109512077
You should've used >>109511826
>>
>>109512091
>ClipProj — MiniMax H3 conditioning from a Qwen3-VL-4B
This is ultra retarded.
>>
>mfw Research news

08/09/2026

>MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations
https://arxiv.org/abs/2608.00736

>Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
https://arxiv.org/abs/2608.03284

>Test-Time Curriculum for Open-Set AIGC Detection
https://arxiv.org/abs/2608.00559

>Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
https://arxiv.org/abs/2608.00094

>Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition
https://arxiv.org/abs/2608.04623

>Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
https://arxiv.org/abs/2608.04394

>Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models
https://arxiv.org/abs/2608.05834

>EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuning
https://arxiv.org/abs/2608.01142

>DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation
https://arxiv.org/abs/2608.01343

>Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning
https://arxiv.org/abs/2608.00994

>MiniWorld: Democratizing the Training of Video World Models from Scratch
https://arxiv.org/abs/2608.01127

>GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression
https://arxiv.org/abs/2608.03517

>Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
https://arxiv.org/abs/2608.03471

>Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression
https://arxiv.org/abs/2608.02134

>OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
https://arxiv.org/abs/2608.03812
>>
File: inner peace.jpg (667 KB, 1024x1024)
667 KB JPG
im gonna find inner peace by buying a 32gb vram card, downloading qwen image 20b and using my hyper specific art style Loras from my favorite artists and goon my brains out
>>
>>109512093
they both have different usecases, you can't use references on t2v/i2v
>>
>>109512098
Why are you reading the schizo links?
>>
>>109512098
Yeah
>We finally have an unretarded model with a powerful text encoder that lets it understand what the user is trying to do.
>Quick, make the text encoder retarded! It could save up to 2% gen time for the text encoding step that only runs once! This will surely help.
>>
>>109512119
Are they not automated?
>>
If I used the quantized heretic text encoder, will Minimax know what sex is?
>>
>>109512132
there's only one way to find out
>>
>>109512129
No he does it manually by hand, all of his thread rituals are manual endeavors.
There's a reason why he's in OP and he does it in multiple threads.
>>
kino, you can do so much with references.

The entire video must forcefully match the exact art direction, PlayStation 1 graphics, fixed-camera style, and pre-rendered 2D backdrop appearance of <Picture 1>

Hatsune Miku is holding a silver pistol, and running around a room in a mansion, and walks up the staircase.

img source is re1 original.

https://files.catbox.moe/w5vslt.mp4
>>
>>109512077
Why did this cuck of an OP generate a man groping his waifu?
>>
breast thred
>>
/a/ autists are filled with so much hatred. feels good.
one of them even picked up on one of ani's boogeymen labels 'wanschizo', probably without even knowing what wan is.
>>
>>109512149
that's exactly what I want for their upcoming image model, to use a reference and perfectly transfer it to another image's style
>>
>>109512112
well you can use an audio and of course images but of course they don't 'work the same way, the audio tho works like LTX in that you can provide a small sample of a voice and it uses that sample as a refence for the talking in the prompt. And unlike LTX if that is used for someone not known (some voices are bad even if supposed to be the person) then that can be used with people in the prompt it does know
>>
>>109512151
We're not using your thread, debo.
>>
>>109512163
look man I love my aislop but you gotta stop being obsessed like this
>>
File: 1615829958599.png (16 KB, 506x606)
16 KB PNG
>>109512171
>look man I love my aislop but you gotta stop being obsessed like this
>>
i wish i could hook my computer into my brain to give it more memory to generate stuff
>>
>thread is ending
>2 or 3 bakes
>enter "non-troll bake"
>walls of garbage links
>someone starts complaining about /a/
>100 posts in
>finally some discussion begins
tiresome
>>
holy shit, almost like the game itself minus the gun sounds (could prompt those)

The entire video must forcefully match the exact art direction, PlayStation 1 graphics, fixed-camera style, and pre-rendered 2D backdrop appearance of <Picture 1>

Hatsune Miku is holding a silver pistol. A nearby door is broken open and a group of human zombies walks through. Miku shoots the zombies with her pistol and they fall down.

https://files.catbox.moe/mgz89e.mp4
>>
File: 669.gif (3.86 MB, 400x532)
3.86 MB GIF
>>109512163
>>
>>109512188
am I in that book?
>>
>>109512185
>100 posts in
*29
>>
File: debo_tt_k2_00132_.png (2.53 MB, 1872x1007)
2.53 MB PNG
>>109512098
what makes it ultra retarded?
>>
https://d.uguu.se/uBArvTUD.mp4
>>
>>109512195
who?
>>
>>109512091
>>109512101
go back to your containment general
>>
>>109512201
Can you even run H3?
>>
File deleted.
>>109510922
Oh, sorry I missed this. I took a "break" from what I was making to generate something safer to share for you.

This is safe, right?
https://files.catbox.moe/mu5hz2.webm
>>
Has anyone tried to take the missle knows where it is transcript and gen something with it yet? id be curious to know what h3 does
>>
File: 1766421093048035.jpg (105 KB, 800x737)
105 KB JPG
What are games that really light enough to play while genning
>>
>>109512186
yes. the sound often is the weakest. it already seems to have some capabilities to match externally provided sounds to video tho.

so perhaps even before finetunes that make it better you could already supply audio made/generated elsewhere.
>>
>>109512244
Slay the spire on you igpu
>>
https://files.catbox.moe/dtra7d.mp4
anons... Something went wrong...
>>
>warning
hmmm
>>
>>109512255
What do you mean? How do you typically float around? Head first? That's dangerous.
>>
>another bake with no collage and bakers shitty gen.
Grim times for /ldg/
>>
>>109512276
Guess you shoulda baked then, huh?
>>
How many of you guys are using the heretic text encoder for H3?
>>
>>109512250
I could even feed game sounds as a reference, there is so much shit this model can do
>>
>>109512280
You're only slightly less worse than the rentry schizos.
>>
>>109512276
worse, the OP video is like 4 days old https://desuarchive.org/g/thread/109477686#109478380
>>
>>109512293
>>
>>109512282
snake oil.
>>
>>109512280
>heh.... YOU should have baked then...
>NOOOOOO WHAT THE FRICK WHY DID YOU BAKE THAT?
>SPAM SPAM SPAM SPAM!!!!!!!!
tiresome
>>
>>109512276
I gave you 3 minutes and you didn't bake, dumbshit autist.
>>
Anyone ever find out why cumfart only uses like 60% of vram??
>>
>>109512313
I wasn't there, and you're pretty delusional if you think it's just 1 person that cares about the collage.

Like the other anon said, Why do you insist on not baking with one?
>>
>>109512326
I don't care how many people care about it. You're all the same brand of boring autist.
I don't give a fuck about your shitty collages and only care about having a non-debo thread active. If nobody else bakes, I do as the failsafe.
Don't like it? Cry about it. Learn to realize there are bigger problems in life than /ldg/ not having a fkin collage OP lol.
>>
>>109512326
go and join debo in the other thread anon.
>>
>>109512345
I can smell the cheeto dust coming from this post
>>
>>109512345
Fedora energy
>>
>>109512355
RED 40 BABY
>>
>>109512355
>>109512358
So head on over to the debo thread. What's stopping you?
>>
File: 1766905838115164.png (135 KB, 474x532)
135 KB PNG
>>109512355
>>109512358
>I can smell the cheeto dust coming from this post
>Fedora energy
>>
>>109512106
good luck, I got two 5090s last year, and now with the same prices I paid, I can get 80% of one
>>
>>109512244
>https://slither.io/
I'm going through my switch backlog.
>>
>>109512369
really wish i could use AI to filter out all basedjak/jak posting. Even frogposting was never this fucking annoying.
>>
>>109512345
uh oh melty
>>
voicecloners ww@?

https://n.uguu.se/NXNuXaNp.webm

still need to figure out how to get rid of that subtle oscillating undertone. maybe it's just the compressed sample, but not sure.
I wonder if you can adjust the audio quality via the prompt.
>>
to the anon earlier maybe someone else already implemented the multi image loader to your satisfaction:
https://github.com/Deno2026/comfyui-deno-custom-nodes#deno-minimax-h3-multi-reference-image-loader
>>
>>109512345
based
>>
>>109512345
>I only care about having a non-debo thread active
>Learn to realize there are bigger problems in life than /ldg/
are you a woman or a tranny?
>>
>>109512385
>english
dropped
>>
wake up samurai

https://files.catbox.moe/372yob.mp4
>>
>>109512385
if you don't succeed to your satisfaction (i couldn't really, maybe it's early support or maybe the audio model just isn't so close to SOTA) perhaps use another voice cloning TTS https://github.com/diodiogod/TTS-Audio-Suite and supply the audio.
>>
>>109512399
eh, someone on /vp/ requested the english voice so that's what i used. normally i'd go japanese myself, and did with this one:

https://n.uguu.se/tdFbLerH.webm

Can recreate May's original voice very authentically, even with a seductive tone. Looking forward to doing more voice cloning in H3.
>>
>>109512414
not gonna lie this video is way better
no idea why the subtle tease of her just opening the top was enough for me but holy by god i want more
(also i do prefer may's english purely because of nostalgia)
>>
>>109512399
agreed, jp va are just that good

>>109512414
better
>>
>>109512407
I've actually got a decent amount of experience with tts models, and actually downloaded a bunch right before minimax launched and blew my dick up.
They're really much of a muchness. For video synchronization there's certainly no reason to use a dedicated tts model over minimax, which is trained to sync with lip movement.
I'm sure some nerd could figure out how to optimize it in post running a pass thru Adobe Audition or Tenacity or something.
>>
>>109512379
must mean its working if you're getting this angry
>>
https://github.com/jpietek/PenguinBurner

Something I recommend anyone to do : powerlimit the card to like 80-90% then run penguinburner to undervolt it, performance will be almost the same and the card will be more stable in general.
>>
>>109512437
You can undervolt in MSI Afterburner. Why would you install this just to do that?
>>
>>109512441
It's for Linux
>>
Does anyone make 3D stereo images?

I prompted krea 2 to make "3D side-by-side binocular stereogram stereographic image"

The result was...it looked like it was going to work. There was clearly depth to the image but it would vary between being correct for cross-eyed viewing and parallel viewing at different parts of the image. My guess is the only reason it doesn't work is because the data wasn't properly tagged between these two for training.

There are programs to convert images to 3d but they are based on lidar which can't do stereo correct transparency, refraction, reflection, etc. Plus finding a one with good inpainting for occluded areas has been a pain. It would be nice if I could just make stuff 3D in one shot.
>>
>>109512163
Stay here in your containment thread wanschizo. You can slop and slop and slop to your heart's content. No need to spread like a cancer trying to metastasize
>>
>>109512437
No thanks I did it in lact
>>
>>109512441
I'm genning on a headless linux.
>>
>>109512452
Anon, you do realize you're both stupid AND autistic, right?
>>
File: 1786319922.jpg (38 KB, 928x297)
38 KB JPG
Witcher 4 showcase.
>>>/wsg/6211190
>>
>>109512427
>I'm sure some nerd could figure out how to optimize it in post running a pass thru Adobe Audition or Tenacity or something.
but do you need to with all the tools that are available from local AI and conventional sound libs even in that custom node pack? https://github.com/diodiogod/TTS-Audio-Suite#features ... of course use any others but this is really quite nice rn since almost all the best stuff is in there with more or less the configuration you'd want.
>>
>>109512460
Someone should make a checklist of all your insults
>>
>>109512431
whats working is zoomer fags constantly forcing shit that is unfunny reddit dog shit.
>>
>>109512467
Anon, you don't know anything about the dynamics of /ldg/. You're just a low-IQ autist from /a/ who started copying ani's boogeyman labels.
There's a rentry link about ani in the OP of this thread. Perhaps you should enlighten yourself before exposing yourself as an even bigger moron ;)
>>
>>109512474
Actually I take it back, it should be a bingo card. Yeah, that would work much better.
>>
>>109512480
Relax anon, we can all see you're quite stupid.
>>
>>109512462
How did you make 44 secs vid ?
>>
this is how we do action in Uganda

https://files.catbox.moe/vqcm7y.mp4
>>
>>109512487
It's stitched. Supposedly you can just gen much longer videos than 15 but i don't have the VRAM for that.
>>
>>109512485
Who is "we"
>>
>>
>>109512495
This is what was promised to us by mark zuckerberg
>>
File: NO IM NOT JEALOUS.png (629 KB, 640x628)
629 KB PNG
>>109512495
holy shit this dude is living his best life
>>
>>109512495
kys
>>
Any implied sex h3 gens?
>>
>>109512196
told you
>>
>>109512495
oh my
>>
>>109512499
lets see your gens, if you're so great
>>
>>109512495
>you can see the woman walk all the way up from the water to the viewer
impressive
>>
File: gfdgfthy.png (578 KB, 796x802)
578 KB PNG
>>109512507
>>
>>109512493
I have a 5060ti 16gb and 64gb ram and do 20 sec gens which are just that bit longer to avoid sped up talking to try to fir it in. But really it is all about the prompt fitting to the length of the gen length anyway, so timestamps would make it easier to do it, the prompt enhancer (or using a another LLM even online) can on the most part do it all for you via a basic prompt
>>
>>109512512
>>109512499
>>
>>109512495
What Epstein could have been if he wasn't the way he was
>>
>>109512516
mad
>>
>>109512499
>>109512522
>>
File: PREVIEW_00001.mp4 (1.74 MB, 1376x768)
1.74 MB
1.74 MB MP4
>>109512345
>>>/wsg/6211203
>>
>>109512530
keeeeeeeeeeeeek
>>
minimax does a good vj emmie clone and I dont even have the best cropped clip

https://files.catbox.moe/yv4udn.mp4
>>
>>109512530
my sides
>>
>>109512532
impressive, I need to get in on that
>>
File: file.png (700 KB, 565x542)
700 KB PNG
>>109512531
>keeeeeeeeeeeeek
>>
Adding a 9 second video reference triples the generation time...
>>
>>109512495
this is fucking amazing
>>
File: 1773600390985886.mp4 (777 KB, 672x640)
777 KB
777 KB MP4
i forgot to mention he should be screaming the words, not just screaming.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.