[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109752182

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

09/07/2026

>Sol-H3 Speed-of-Light MiniMax-H3 on an 8× NVIDIA B300 Blackwell System
https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3

>MageTrail: Danbooru/E621 Full-Finetune of MageFlow 4B
https://huggingface.co/RicemanT/MageTrail

>Learning 3D Editing without Paired Supervision via Generative Prior Distillation
https://github.com/thiamine128/PriorEdit3D

>UniMate: One Unified Model to Animate Diverse Skeletons
https://linzhanmou.com/unimate

>VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
https://github.com/FlamieZhu/Vicinal-Consistency-Alignment

>WorldSculpt: Generating Compositional Worlds from Grounded Videos
https://alaya-lab.github.io/WorldSculpt

>AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
https://zunwang1.github.io/AnchorWeave

>ComfyUI-InpaintCanvas: Krita-style inpainting without leaving ComfyUI
https://github.com/DenRakEiw/ComfyUI-InpaintCanvas

>ComfyUI-Ref2VA-VSA: Ultra-Fast Character Video Generation
https://github.com/Kablex/ComfyUI-Ref2VA-VSA

09/06/2026

>ComfyUI-Viggle-Animate-H3
https://github.com/Saganaki22/ComfyUI-Viggle-Animate-H3

>uncomfymcp: MCP server for generating images on ComfyUI from a chat client.
https://github.com/aschet/uncomfymcp

>MiniMax H3 Semantic Bridge
https://huggingface.co/speach1sdef178/MiniMax-H3-Semantic-Bridge

09/05/2026

>Intern Lumina U2: Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation
https://internlm.github.io/InternLumina-U2

>Musk’s xAI loses court bid to block Minnesota's AI ‘nudification’ ban
https://www.reuters.com/legal/litigation/musks-xai-loses-court-bid-block-minnesotas-ai-nudification-ban-2026-09-04

>Add Sparse Attention node
https://github.com/Comfy-Org/ComfyUI/pull/16072

>MiniMax-H3 FL2V 8-Step Motion Enhancer
https://huggingface.co/rzgar/minimax-h3_fl2v_8Step_motion_enhancer

>Qwen3.8-Flash-Next-NVFP4
https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4
>>
>mfw Research news

09/07/2026

>Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
https://arxiv.org/abs/2609.04283

>ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features
https://arxiv.org/abs/2609.04649

>RefDiT: Local Attribute Guidance in Reference-Based Image Generation
https://arxiv.org/abs/2609.04976

>Importance-Aware Low-Rank Distillation of Diffusion Transformers
https://vislearn.github.io/SVDtrunc

>Measured Sliders: Learning Continuous Controls from Differentiable Image Measurements
https://arxiv.org/abs/2609.05234

>PAPT++: Risk-Aware Adversarial Tuning and Generation for Single Domain Generalization
https://arxiv.org/abs/2609.04837

>AngelFingerprint: A Traceable, Explainable, and White-Box Stealthy Watermark for Text-Guided Image Editing
https://arxiv.org/abs/2609.04709

>Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation
https://arxiv.org/abs/2609.04282

>LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering
https://arxiv.org/abs/2609.04939

>Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval
https://arxiv.org/abs/2609.04800

>WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing
https://arxiv.org/abs/2609.05171

>FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
https://arxiv.org/abs/2609.04276

>Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
https://arxiv.org/abs/2609.05275
>>
>>109757421
>>109757414
>>
File: stfu.webm (955 KB, 736x1120)
955 KB
955 KB WEBM
>>109757430
https://files.catbox.moe/sdeh8d.mp4
>>
File: biggiefattie.webm (633 KB, 928x928)
633 KB
633 KB WEBM
https://files.catbox.moe/6khu5e.mp4
>>
>>109757430
You've never noticed it generated the wrong type of keyboard?
>>
Any YTP Minimax LoRAs?
>>
>>109757571
This is from the qwen prompt but you have to know the esoteric lore that debo was a failed DJ before going full NEET so it stabs extra deep.
>>
>>109757589
>failed DJ
>gens black characters
debo is 100% black. no way this guy is white.
>>
I think I fucked up. the SaveVideo node in comfyui isn't saving any output on the server from clients genning on the local lan (with --listen). filename_prefix was the same for all of them, from the workflow.
there wasn't a warning in the logs that it couldn't write or anything, and genning from a browser on the host works fine. is this because an output directory wasn't set at launch?
>>
>>109757608
the SaveVideo node ignores your output directory and write to the comfyui root for some reason
>>
File: output_150_89mb.mp4 (3.69 MB, 1552x2048)
3.69 MB
3.69 MB MP4
>>109757597
No he's a sheltered Californian NEET
>>
>>109757613
What the fuck
>>
>>109757617
whats with the wheelchair
>>
>>109757631
Don't click random catbox links anon
>>
File: output_smal.mp4 (3.92 MB, 2048x1130)
3.92 MB
3.92 MB MP4
>>109757639
A anon asked if he was disabled due to the amount of time he spends in the general and I made a wheelchair gen to make fun of him and he flipped the fuck out and threatened to kill me while posting base XL gens at default resolution.
I then posted it after the migration and he also flipped the fuck out so now I just post it because it hurts him.
It does make sense seeing how often he's here and how he acts.
>>
>>109757615
oh jesus. it looks like it created a single file in the base output directory and then overwrote it instead of incrementing the filenames. I guess the clients couldn't read the path on the host?
how do I fix this
>>
File: output_no_audio_lossless.mp4 (3.36 MB, 1056x608)
3.36 MB
3.36 MB MP4
It's kind of sad both the rentry schizos have threatened physical harm to me over bantz. Not once have I ever called for violence to either of them but they are so fragile gens and simply questioning their behavior pushed them that far.
Too sensitive for the internet desu
>>
>>109757641
>>109757631
For him to delete it, it must have been bad. What was it?
>>
>>109757716
Just something that doesn't belong on a Blue Board. The other one is still up, lazy jannies.
>>
I am using ref 2 video , i have done one video with background music now making part 2 , what is the best way to make the audio so it is seamless? or is this better after i made the videos i need
>>
>>109757716
wrong board.
>>
File deleted.
3 fucking days of failed gens.
I should go back to LTX at this rate
>>
>>109757744
Based. Model?
>>
>>109757673
use a different node ig?
>>
File: Jeralt.webm (2.33 MB, 1120x736)
2.33 MB
2.33 MB WEBM
https://files.catbox.moe/amorcq.mp4
>>
>>109757869
I was waiting for him to slay the monster but this is fine, too.
>>
>>109757761
fixed it, I'm just a retard. the prefix was set to %date:yyyy-MM-dd%-user/ instead of %date:yyyy-MM-dd%/user.
I changed it on the host early on and didn't ask anyone to refresh.
>>
File: file.mp4 (1.38 MB, 864x480)
1.38 MB
1.38 MB MP4
https://files.catbox.moe/h60qz5.mp4 https://files.catbox.moe/g1p627.mp4
>>
>>109758018
beautiful..
>>
life must be tough and hard being a vramlet and ramlet in q3 of 2026. i hope your hardware is ready for flux3.
>>
>>109758071
>>
>>109758078
>All the VRAM
>20B Pruned
Genuinely, why? Speed? Fitting the whole thing in one GPU? I'm curious, not criticizing.
>>
File: Ohfug.webm (1.39 MB, 1216x672)
1.39 MB
1.39 MB WEBM
https://files.catbox.moe/844nyk.mp4
>>
>>109758083
as opposed to what?
>>
File: file.png (1 KB, 103x28)
1 KB PNG
>>109758112
As opposed to the full un-pruned complete 33B model.
>>
>>109758117
why would you want to use that?
>>
>>109758125
It produces consistently higher quality results.
>>
>>109758133
no. the people that pruned it found that the weights removed had very little to do with the generation side of things. you would only want to use the unpruned model if you are trying to do further training
>>
File: 1440952787112.jpg (52 KB, 750x705)
52 KB JPG
latest miniconstruct is ready but i'm literally being filtered by git semantics
>>
File: file.png (374 KB, 563x728)
374 KB PNG
>>109758154
I've been lied to.
>>
>>109758180
note that if you are saving seeds, the pruned one will make a different generation from the same seed that you used on the 33b one
>>
>>109757250
https://rentry.org/neo_collage
I updated anon's collage script. It should now work better on lower end systems and smoother on higher ends, 5 minute max video length, less jank, a fresh interface, and some other changes.

The original script recorded canvas playback in real time while this one decodes source frames, composites them, and encodes explicit timestamps. It should no longer drop frames and slow FPS.
>>
>>109758292
For editing and auditing purposes, please please @require mediabunny from a CDN and make the userscript code unminified.
>>
File: WHAT IS THE PROMPT.png (2.09 MB, 1184x896)
2.09 MB PNG
>>109757250
OK, so this is probably an obvious idea but has any company actually done the following?

Take an already trained model and a diverse data set of various images (photos, paintings, illustrations, ...) and have the AI prompt itself to recreate the image from scratch (i.e. not I2I) using itself. Given that its prompted output image can easily be compared mathematically against the original to determine how close it came to matching it, this should allow it to quickly get better like neural nets are wont to at these kind of optimization tasks. By the end of this automated training you should have a model that can take any input image and give you prompts that, whether they make "sense" to a human or not, result in output that matches the original.

So what you get out of that is an AI prompt engineer or sorts that can look at picture A and help you figure out the kind of prompts that would get picture B to look like it (style-wise or composition-wise or whatever).
>>
>>109758292
The op is proud of his low effort slopscript and will probably ignore this
>>
>>109758489
*of sorts
>>
File: krea2.png (3.42 MB, 1280x1600)
3.42 MB PNG
I like Krea 2
>>
>>109758425
Done.
>>109758492
OP and the author of the script in OP are clearly two different people.
>>
Brainlet here. I'm getting decent 5 second clips with coomer loras on H3 so now I'm taking a shot at extending them. Do I need to write individual prompts for every single five second video extension?

I'm using wangp, I assume comfy could let me automate this 100% for a 45 second clip without me writing these in between prompts
>>
>>109758526
>Do I need to write individual prompts for every single five second video extension?
You can do either. Separate prompts are recommended when the action changes over time; one reused prompt works well for an ongoing activity. You have a setting, "How to Process each Line of the Text Prompt". With "All the Lines are Part of the Same Prompt", the same prompt conditions each window, with overlap carrying context forward. Good for someone walking through a forest, dancing, a long sustained landscape shot, that sort of thing. With "Each Paragraph Separated by an Empty line will be used for a new Sliding Window of the same Video Generation", you still paste everything into one prompt box and run one job, but you package it differently. One complete H3 prompt per window, no blank lines inside a window’s prompt, one blank line between window prompts, include ALL required H3 fields/sections in each window. Sliding window generation just automates the process, you can reproduce the same results without using sliding window generation at all, but it saves a lot of effort.
>>
>>109758519
Your script is broken at least on Tampermonkey due to the way it bundles requires. You need to add a leading semicolon on L24.

;(() => {


minor nits:
- I'd suggest making the "collage -" button green to make it visually more obvious the image was selected than the + turning into a -
- I'd suggest making selected gens have a green background in the review UI to make selected more obvious than a ticked checkbox

Otherwise it looks good to me,
I deprecated https://rentry.org/ldgcollage_v2
>>
>>109758593
>Your script is broken at least on Tampermonkey due to the way it bundles requires. You need to add a leading semicolon on L24.
Copy the new version I just changed. I ran into the same issue it should work now.
>minor nits:
Good ideas, the first one I was thinking of as well. I'll add those soon. Maybe more stuff.
>>
I have to fix the aspect ratio stuff as well. Hold on.
>>
>>109758601
>>109758607
Cool, no rush.
>>
>>109758526
Comfy or Wan can automate that 100%, and both are fucking terrible at it. They work by rewriting your prompt, that's how it can be automated, you're letting the clanker take your prompt and decide how to break it up into several smaller ones, which they are very bad at in my experience. A frontier model with a real H3 prompting skill can do it, but I don't expect most people to use Astra to write smut. Though I do recommend it, because she's a dirty whore.
>>
File: Oh-My-Yes-2733482897.png (1.05 MB, 1440x1080)
1.05 MB PNG
Good news everyone: MiniConstruct 0.4.0 is out. It is a MONSTER of a release.

https://github.com/InsertSpice/MiniConstruct

What's new:
>New 'creative interpretation' toggle - you can adjust how much creativity the LM is allowed to have when generating the prompt.
>Extensive benchmarking and strengthening of the system prompt, ensuring the model can understand exactly how to incorporate references into the output prompt
>this now includes reference sheets for subject identities, saving you from using multiple images to support all angles of a subject
>Projects have officially transitioned from IndexDB to the filesystem.
>New story planner: rather than generating single prompts, plan and generate a larger story across a sequence of generations, each with their own incoming state, outgoing state, camera shot/intent.
>Don't want to fill in all that boilerplate? No problem, just click a button to have the LM generate it
>Iterative revision features for story control. Don't like how a certain generation is composed? Just tick it's box and revise.
>ComfyUI integration is now all systems go. You can now use MiniConstruct as a front-end for generating videos in ComfyUI.
>>
idc
>>
>>109758648
how does connecting to comfy work? does the output just get routed to the string node?
>>
>>109758639
Just go binary, final destination.
>>
>>109758620
Changed some of the layout and rendering, minor nits, and removed the check boxes all together. Just click on the cards to select - you can still click the thumbnails to preview. Aspect ratio accepts different formats btw but I think that logic is still weird. Other than that, the newest version should be g2g.

I'm open to other changes in the future. I might make more updates.
>>
>>109758620
is this downscaled?
>>
>>109758706
You export the Comfy workflow as (API), load it in MiniConstruct comfy settings, then bind the assets to the comfy nodes from the exported API workflow. Then you can generate directly from miniconstruct so long as you binded correctly.
>>
>>109758292
Here is a test of every candidate in this thread (the duplicates are from anons posting catbox links of the same video) done on my low spec laptop. As you can see, no dropped frames, stable FPS, no weirdness.
Pretty good.
>>
>>109758746
You caught me, smart anon
>>
anyone else got OOM with new model sparse attention node?
>>
>>109758585
>>109758644
thanks im tinkering some more
im finishing up a 30 second clip and ive noticed that artifacts are appearing in the later gens, like little white spots or blocky artifacts, so after 20 seconds it looks a little weird. how do i fix/denoise these, im gonna try to start over at the 10 second mark
>>
File: dat night feel.png (2.45 MB, 1360x1024)
2.45 MB PNG
>>109758781
WTF? Did Mika-chan lose weight? What are those gens?
>>
We're going to have to change the name of the general again soon, because now there's new ways to make AI videos and that's using LLMs with Blender harnesses
https://litter.catbox.moe/2k4so2qzwa3bg0j6.mp4
Vidrel was Astra from some dude on /vcg/, but local Chinese models are only 3-6 months away so there will be a(nother) new way to locally gen passable video very soon that isn't strictly diffusion
>>
I thought i had good sense of humor but I got no (you) from my gen
>>
File: Krea2_turbo_00158_.png (2.08 MB, 1448x1448)
2.08 MB PNG
>>109758946
Sir, this is a schizo general, if you're not here to discuss catjack or animanon then we're going to have to ask you to leave.
>>
>>109759237
that isn't new, it's basically just a digital plate and people have been doing it since wan 2.1.
>>
File: Test 00006(3).mp4 (3.78 MB, 832x1248)
3.78 MB
3.78 MB MP4
>>109758648
Neat, might give it a try, once I'm too tired to deal with groks occasional bullshit. Wish I had some second PC I could host the LLM on, but I guess switching models in VRAM isn't too slow.

>>109759392
Happens to the best of us.
>>
>>109759533
Show me someone making a video like that with WAN 2.1 in a single prompt you retarded fucking gorilla, at least spend the tokens to invoke the tool call to look at the video if you're going to burn compute replying to me
>>
>>109759702
>in a single prompt
i like how you had to move the goalpost just in case.
could you explain how you think that video was created?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.