[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


Career Opportunities Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109637805

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/24/2026

>MiniMax-H3-Fun-Controlnet-Union
https://huggingface.co/alibaba-pai/MiniMax-H3-Fun-Controlnet-Union

>MiniMax-H3-Longvideos: Long (up to ~120s) MiniMax-H3 video + synchronised audio from a single prompt
https://huggingface.co/Smite79/MiniMax-H3-Longvideos

>DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
https://github.com/KLMAV-CUC/DiGS-Avatar

>Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair
https://github.com/oceanflowlab/AESR

>OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank
https://github.com/Wenyang-hong/OccluRank

>Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
https://github.com/ernie-research/AVIOT

>CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
https://github.com/CubicSplat/repo

>Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
https://github.com/SWUFE-DB-Group/Vis-Poison

>Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization
https://github.com/oceanflowlab/EDD.git

>ArtiMo: Agent-Driven Articulated Mesh Animation
https://zou-2004.github.io/ArtiMo

>MiniMax-H3 × Z-Image — spatial detail graft (comfy-native)
https://huggingface.co/joeygambino/MiniMax-H3-x-Z-Image-native

>MiniMax-H3 4-Step LoRA (FlashGen)
https://huggingface.co/Beidouqixing/minimax-h3-4step-lora-flashgen

08/23/2026

>Krea 2 Turbo — 4-Step Distillation LoRA
https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA

>Alibaba to issue US$10 billion in new shares for huge AI push
https://www.scmp.com/tech/big-tech/article/3364957/alibaba-issue-hk80-billion-new-shares-global-ai-push

>Nvidia Customers Notified About AI-Related Price Hikes Above 15%
https://www.bloomberg.com/news/articles/2026-08-22/nvidia-customers-notified-about-ai-related-price-hikes-above-15
>>
File: ComfyUI_temp_bzkmi_00022_.png (3.14 MB, 1920x1200)
3.14 MB PNG
>>
Blessed thread of frenship
>>
>mfw Research news

08/24/2026

>DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer
https://arxiv.org/abs/2608.20515

>MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control
https://multi-cube.github.io

>GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization
https://arxiv.org/abs/2608.20929

>Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation
https://research.nvidia.com/labs/amri/projects/grounded-exo2ego

>ES-VP : Energy-Shaped Dynamic Visual Prompting for Efficient Model Adaptation
https://arxiv.org/abs/2608.21194

>Scaling Muon for Diffusion Transformers
https://arxiv.org/abs/2608.20818

>InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
https://arxiv.org/abs/2608.20910

>Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers
https://arxiv.org/abs/2608.21229

>Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
https://orarl.github.io

>Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation
https://arxiv.org/abs/2608.20691

>When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference
https://amughrabi.github.io/MomentAux

>Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
https://arxiv.org/abs/2608.21170

>Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores
https://arxiv.org/abs/2608.20725

>LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic Residuals
https://arxiv.org/abs/2608.20882

>Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
https://arxiv.org/abs/2608.21134
>>
first for lolcow baker has no employment opportunities
>>
>>109640479
>>109640482
>>
The advert imagegens from last thread were extremely aesthetic, what model+lora combo is that
>>
File: ComfyUI_temp_bzkmi_00024_.png (2.69 MB, 1920x1200)
2.69 MB PNG
>>
>>109640487
>>109640497
>>
File: ComfyUI_temp_bzkmi_00025_.png (2.84 MB, 1920x1200)
2.84 MB PNG
>>
File: f4micom.jpg (1.02 MB, 3632x2416)
1.02 MB JPG
>>
File: ComfyUI_temp_bzkmi_00027_.png (3.83 MB, 1120x2080)
3.83 MB PNG
>>
>>109640524
>>109640504
>>109640499
>>109640471
This is less fappable than niggers please switch or stop thanks
>>
File: 1765506982315915.png (956 KB, 976x1296)
956 KB PNG
>Posted in the dead thread award
>>109640289
I was very impressed by anima when I was playing around with it, but I couldn't get good controlnets. I want to make a visual novel/WEGslop. I've got a couple style loras trained for illustrious, I'm not sure how necessary controlnets are for what I want, but I had a feeling I'd need them at some point.
>>
File: ComfyUI_temp_bzkmi_00026_.png (2.82 MB, 1920x1200)
2.82 MB PNG
>>109640598
thats because you're indian
>>
File: ComfyUI_temp_bzkmi_00028_.png (3.22 MB, 1120x2080)
3.22 MB PNG
>>
>>109640657
That makes no sense it would be easier to self insert if I was Indian kill yourself
>>
>>109640598
>>109640684
Calm down, Rakesh.
>>
File: ComfyUI_temp_bzkmi_00031_.png (2.73 MB, 1280x1840)
2.73 MB PNG
>>109640684
But you're Indian
>>
>>109640709
>>109640702
it's okay to be Indian
>>
File: ComfyUI_temp_bzkmi_00032_.png (3.19 MB, 1280x1840)
3.19 MB PNG
>>
>>109640714
What about black?
>>
can't go to sleep until i queue up all my gens
>>
>>109640723
not okay
>>
>>109640725
why do you let one bake the threads?
>>
File: ComfyUI_temp_bzkmi_00033_.png (2.91 MB, 1280x1840)
2.91 MB PNG
>>
>>109640746
sniff
>>
>>109639981
Aren't those both Strix Halo? You loaded bastard. Try them out, they should be usable. Either download a portable AMD build of ComfyUI from the latter's Github, or for manual AMD setup steps, see >>109625036 and >>109625041.
>>
File: 1785640564334027.png (879 KB, 1077x816)
879 KB PNG
posting on this website:

https://litter.catbox.moe/90ryru8cmb5elkwh.mp4
>>
>>109640775
oops, disregard image that was from the game test gen.
>>
>>109640775
ummm, why didn't she get TOS banned?
>>
lets see if this works:

https://litter.catbox.moe/x815nhkdvmziq3ps.mp4
>>
>>109640823
off by one
>>
>>109640823
better check (im not going to keep trying, but the point worked this time.)

https://litter.catbox.moe/ausd7dv6anwmmd7v.mp4
>>
>>109640340
Maybe someone will figure out how to train a lora but it does not affect voice path . lol
>>
>>>/gif/31084754
which one of you is this
>>
how much h3 latent space is 40gb of ram worth? or how many seconds at 1mp is it
>>
which lora is better?
https://twinlens.app/compare?share=30e5ab353574
>>
>>109640993
>>109641018

i dont know



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.